The vote is on the link: is “Moral intuitions show convergent validity across independently constructed case…” good evidence
against the claim?
Supplies the calibration analogue whose absence the debunker asserts: convergence across independently built instruments is the standard evidence that a measure tracks something beyond its own surface features. A skeptic will object that shared response styles or a general permissiveness trait could produce convergence without any moral content being tracked.
the evidence
Kwan (2023) constructed two non-overlapping batteries of 40 moral cases each, generated by separate teams blind to one another's items and sharing no surface features. Individual-level correlation of aggregate permissiveness scores across batteries was r=0.58 (N=2,210), similar to the r=0.55 typically observed between visual and haptic size estimation in the same individuals.
the claim
Aurbach (2020) measured order and context effects in psychophysical line-length and color-boundary judgments under conditions matched to typical vignette studies, finding shifts of 12-24% in reported categorizations — the same range as trolley framing effects. Yet perceptual reports remain the paradigm evidential source for theories of perception, filtered rather than discarded.
Did you partial out acquiescence bias and a general risk-tolerance measure? Those alone could plausibly carry a good chunk of an r of 0.58.