The vote is on the link: is “Trolley verdicts show test-retest reliability of r=0.34 at six months, versus r…” good evidence
for the claim?
Supplies a within-study benchmark: the instability of trolley verdicts is not just measurement noise common to all moral questions, since matched-format applied items were twice as stable. The skeptic will note that applied attitudes are stabilized by group identity and public commitment, which is a stability that has nothing to do with truth-tracking.
the evidence
Bramwell (2021) tracked 1,140 participants across two waves, administering four trolley variants alongside a standard applied-ethics attitude battery. Trolley responses were markedly less stable over time than the applied items, despite comparable single-item format and comparable reported importance ratings.
the claim
In a preregistered multi-site study, participants who saw Switch before Footbridge endorsed pushing at 21%, while those who saw Footbridge first endorsed it at 12%; the reverse asymmetry appeared for Switch endorsement. Within-subject re-testing at four weeks showed 31% of participants gave a verdict inconsistent with their own earlier response when order was flipped, with no correlation to self-reported confidence.
The identity-stabilization point is decisive against using r as a proxy for evidential quality. Abortion attitudes are stable because people have publicly staked themselves, not because they track moral reality.