Sources
Where the numbers come from
Each human baseline on the results page is traceable to a published source, with the exact sentence its number was read from. A wrong citation propagates into every comparison drawn from it, so the status of each one is stated rather than implied.
Not yet verified by a person.“Checked” means an automated agent opened the linked text and confirmed the details and quoted numbers against it. That is not verification: until a researcher confirms each source against the published version, treat these as leads. Where the only accessible text was someone else’s report of a study, the baseline says via and names it.
Human baselines
share endorsing the act option, unless statedBystander at the Switch
- 89%87–91%judged it permissible
Hauser et al. 2007 via Park et al. 2023
“They found that 89% of subjects deemed the action in the foreseen-side-effect scenario as permissible (95% CI=[87%, 91%]), while only 11% of them deemed the action in the greater-good scenario as permissible (95% CI=[9%, 13%]).”
- 71%judged it permissible
Klein et al. 2018 via Park et al. 2023
“The Many Labs 2 sample (N=6,842) successfully replicated this finding. 71% of subjects deemed the action in the foreseen-side-effect scenario as permissible, while only 17% of them deemed the action in the greater-good scenario as permissible.”
- 81%said the agent should act
“Participants endorsed sacrifice more for Switch (country-level average: 81%) than for Loop (country-level average: 72%), and for Loop more than for Footbridge (country-level average: 51%).”
An average of 42 country means, not a pooled share, so it weights every country equally regardless of how many people answered there.
The Footbridge
- 11%9–13%judged it permissible
Hauser et al. 2007 via Park et al. 2023
“They found that 89% of subjects deemed the action in the foreseen-side-effect scenario as permissible (95% CI=[87%, 91%]), while only 11% of them deemed the action in the greater-good scenario as permissible (95% CI=[9%, 13%]).”
- 17%judged it permissible
Klein et al. 2018 via Park et al. 2023
“The Many Labs 2 sample (N=6,842) successfully replicated this finding. 71% of subjects deemed the action in the foreseen-side-effect scenario as permissible, while only 17% of them deemed the action in the greater-good scenario as permissible.”
- 51%said the agent should act
“Participants endorsed sacrifice more for Switch (country-level average: 81%) than for Loop (country-level average: 72%), and for Loop more than for Footbridge (country-level average: 51%).”
The Loop
- 72%said the agent should act
“Participants endorsed sacrifice more for Switch (country-level average: 81%) than for Loop (country-level average: 72%), and for Loop more than for Footbridge (country-level average: 51%).”
The Transplant Surgeon
- 3%judged it permissible
“but in the second case 97 percent answered that it would be wrong to sacrifice a healthy person to allow five sick ones to live”
A press report of the Moral Sense Test, not a figure from a paper; 3% is 100% minus the 97% who answered "wrong". Treat as the weakest baseline here.
Personal Force versus Harm as Means
- ratingsrated acceptability
Pushing the victim with one's own hands (standard footbridge, n=154) was rated less acceptable than dropping them through a trapdoor by remote switch (n=82). Spatial proximity and physical contact had no separate effect; personal force did, and only when the harm was the means.
“These results indicate that harmful actions involving personal force are judged to be less morally acceptable. Moreover, they suggest that spatial proximity and physical contact between agent and victim have no effect”
Reports ratings rather than a share, so there is no number to set beside a model's act rate. The prediction it makes for this template is directional: trapdoor above footbridge.
Bibliography
every citekey used by the packs and the baselinesEdmond Awad, Sohan Dsouza, Azim Shariff, Iyad Rahwan, Jean-François Bonnefon (2020). Universals and variations in moral decisions made in 42 countries by 70,000 participants. Proceedings of the National Academy of Sciences, 117(5), 2332-2337. doi:10.1073/pnas.1911517117
Is doing the right thing hard-wired?. Harvard Gazette, 2007. link
Joshua D. Greene, Fiery A. Cushman, Lisa E. Stewart, Kelly Lowenberg, Leigh E. Nystrom, Jonathan D. Cohen (2009). Pushing moral buttons: The interaction between personal force and intention in moral judgment. Cognition, 111(3), 364-371. doi:10.1016/j.cognition.2009.02.001
Joshua D. Greene, R. Brian Sommerville, Leigh E. Nystrom, John M. Darley, Jonathan D. Cohen (2001). An fMRI Investigation of Emotional Engagement in Moral Judgment. Science, 293(5537), 2105-2108. doi:10.1126/science.1062872
Judith Jarvis Thomson (1985). The Trolley Problem. The Yale Law Journal, 94(6), 1395-1415. link
Marc Hauser, Fiery Cushman, Liane Young, R. Kang-Xing Jin, John Mikhail (2007). A Dissociation Between Moral Judgments and Justifications. Mind & Language, 22(1), 1-21. doi:10.1111/j.1468-0017.2006.00297.x
Natalie Gold, Briony D. Pulford, Andrew M. Colman (2014). The outlandish, the realistic, and the real: contextual manipulation and agent role effects in trolley problems. Frontiers in Psychology, 5, 35. doi:10.3389/fpsyg.2014.00035
Peter S. Park, Philipp Schoenegger, Chongyang Zhu (2023). Diminished Diversity-of-Thought in a Standard Large Language Model. arXiv, 2302.07267. link
Philippa Foot (1967). The Problem of Abortion and the Doctrine of the Double Effect. Oxford Review, 5, 5-15. link
Richard A. Klein, et al. (2018). Many Labs 2: Investigating Variation in Replicability Across Samples and Settings. Advances in Methods and Practices in Psychological Science, 1(4), 443-490. doi:10.1177/2515245918810225