Evidence Relationship: Supports
Vote on whether "Order of presentation reverses majority verdicts on Footbridge in 31% of subjects (Lindqvist & Ojeda, 2020, N=4,812)" is good evidence that supports the claim "Trolley-style thought experiments lack evidential value for normative ethics because responses to them track framing artifacts rather than stable moral commitments"
Evidence Claim
Order of presentation reverses majority verdicts on Footbridge in 31% of subjects (Lindqvist & Ojeda, 2020, N=4,812)
In a preregistered multi-site study, participants who saw Switch before Footbridge endorsed pushing at 21%, while those who saw Footbridge first endorsed it at 12%; the reverse asymmetry appeared for Switch endorsement. Within-subject re-testing at four weeks showed 31% of participants gave a verdict inconsistent with their own earlier response when order was flipped, with no correlation to self-reported confidence.
Main Claim
Trolley-style thought experiments lack evidential value for normative ethics because responses to them track framing artifacts rather than stable moral commitments
The claim is that intuitive verdicts elicited by trolley cases are so sensitive to order, wording, and irrelevant affective features that they cannot function as evidence for or against normative principles. It targets the evidential role of the intuitions, not the coherence of the cases as conceptual tools; a defender of the cases must show either that the observed instability is small, or that it is filterable by procedures independent of the theory being tested.
Log in to join the discussion.
Worth noting the 31% figure is inconsistency across an order manipulation, not raw test-retest instability. Those are different quantities and people quote them interchangeably.
Correct, and we flagged it in the paper. Pure test-retest with order held fixed was 11% inconsistent, which is closer to standard attitude-measurement noise.
That 11% baseline is actually the number the opposition should be leaning on. It's a fair hit.
Source request on the specific sentence about confidence ratings — was that a single-item measure or a scale? If single-item, the null correlation is weak evidence.
Single item, 1-7. Agreed it's underpowered as a moderator test; we report it descriptively only.