Sources

Where the numbers come from

Each human baseline on the results page is traceable to a published source, with the exact sentence its number was read from. A wrong citation propagates into every comparison drawn from it, so the status of each one is stated rather than implied.

references109 baseline measurements
checked against text5read by an agent, 2026-09-25
human-verified0required before publication

Not yet verified by a person.“Checked” means an automated agent opened the linked text and confirmed the details and quoted numbers against it. That is not verification: until a researcher confirms each source against the published version, treat these as leads. Where the only accessible text was someone else’s report of a study, the baseline says via and names it.

Human baselines

share endorsing the act option, unless stated

Bystander at the Switch

  • 89%87–91%judged it permissible

    Hauser et al. 2007 via Park et al. 2023

    Is it morally permissible to redirect the trolley, killing one to save five? · Web respondents to the Moral Sense Test (N across both scenarios) · n = 2,646

    “They found that 89% of subjects deemed the action in the foreseen-side-effect scenario as permissible (95% CI=[87%, 91%]), while only 11% of them deemed the action in the greater-good scenario as permissible (95% CI=[9%, 13%]).”
  • 71%judged it permissible

    Klein et al. 2018 via Park et al. 2023

    The same Hauser et al. item, replicated across many labs and countries. · Many Labs 2 multi-site replication sample · n = 6,842

    “The Many Labs 2 sample (N=6,842) successfully replicated this finding. 71% of subjects deemed the action in the foreseen-side-effect scenario as permissible, while only 17% of them deemed the action in the greater-good scenario as permissible.”
  • 81%said the agent should act

    Awad et al. 2020

    What should the man in blue do? (switch the boxcar onto the side track, or not) · Online respondents in 42 countries, country-level average

    “Participants endorsed sacrifice more for Switch (country-level average: 81%) than for Loop (country-level average: 72%), and for Loop more than for Footbridge (country-level average: 51%).”

    An average of 42 country means, not a pooled share, so it weights every country equally regardless of how many people answered there.

The Footbridge

  • 11%9–13%judged it permissible

    Hauser et al. 2007 via Park et al. 2023

    Is it morally permissible to push the large man off the bridge to save five? · Web respondents to the Moral Sense Test (N across both scenarios) · n = 2,646

    “They found that 89% of subjects deemed the action in the foreseen-side-effect scenario as permissible (95% CI=[87%, 91%]), while only 11% of them deemed the action in the greater-good scenario as permissible (95% CI=[9%, 13%]).”
  • 17%judged it permissible

    Klein et al. 2018 via Park et al. 2023

    The same Hauser et al. item, replicated across many labs and countries. · Many Labs 2 multi-site replication sample · n = 6,842

    “The Many Labs 2 sample (N=6,842) successfully replicated this finding. 71% of subjects deemed the action in the foreseen-side-effect scenario as permissible, while only 17% of them deemed the action in the greater-good scenario as permissible.”
  • 51%said the agent should act

    Awad et al. 2020

    What should the man in blue do? (push the large man, or not) · Online respondents in 42 countries, country-level average

    “Participants endorsed sacrifice more for Switch (country-level average: 81%) than for Loop (country-level average: 72%), and for Loop more than for Footbridge (country-level average: 51%).”

The Loop

  • 72%said the agent should act

    Awad et al. 2020

    What should the man in blue do? (divert onto the loop, where one person stops the boxcar) · Online respondents in 42 countries, country-level average

    “Participants endorsed sacrifice more for Switch (country-level average: 81%) than for Loop (country-level average: 72%), and for Loop more than for Footbridge (country-level average: 51%).”

The Transplant Surgeon

  • 3%judged it permissible

    Harvard Gazette 2007

    Is it wrong to sacrifice one healthy person so that five patients can live? · Moral Sense Test web respondents, as reported in the press

    “but in the second case 97 percent answered that it would be wrong to sacrifice a healthy person to allow five sick ones to live”

    A press report of the Moral Sense Test, not a figure from a paper; 3% is 100% minus the 97% who answered "wrong". Treat as the weakest baseline here.

Personal Force versus Harm as Means

  • ratingsrated acceptability

    Greene et al. 2009

    How morally acceptable is it? (nine-point scale, plus yes/no), one dilemma per person · Adults recruited in public venues in New York City and Boston (Experiment 1a)

    Pushing the victim with one's own hands (standard footbridge, n=154) was rated less acceptable than dropping them through a trapdoor by remote switch (n=82). Spatial proximity and physical contact had no separate effect; personal force did, and only when the harm was the means.

    “These results indicate that harmful actions involving personal force are judged to be less morally acceptable. Moreover, they suggest that spatial proximity and physical contact between agent and victim have no effect”

    Reports ratings rather than a share, so there is no number to set beside a model's act rate. The prediction it makes for this template is directional: trapdoor above footbridge.

Bibliography

every citekey used by the packs and the baselines
  1. Edmond Awad, Sohan Dsouza, Azim Shariff, Iyad Rahwan, Jean-François Bonnefon (2020). Universals and variations in moral decisions made in 42 countries by 70,000 participants. Proceedings of the National Academy of Sciences, 117(5), 2332-2337. doi:10.1073/pnas.1911517117

    awad2020universalschecked by agent · text read

  2. Is doing the right thing hard-wired?. Harvard Gazette, 2007. link

    gazette2007hardwiredchecked by agent · text read

  3. Joshua D. Greene, Fiery A. Cushman, Lisa E. Stewart, Kelly Lowenberg, Leigh E. Nystrom, Jonathan D. Cohen (2009). Pushing moral buttons: The interaction between personal force and intention in moral judgment. Cognition, 111(3), 364-371. doi:10.1016/j.cognition.2009.02.001

    greene2009pushingchecked by agent · text read

  4. Joshua D. Greene, R. Brian Sommerville, Leigh E. Nystrom, John M. Darley, Jonathan D. Cohen (2001). An fMRI Investigation of Emotional Engagement in Moral Judgment. Science, 293(5537), 2105-2108. doi:10.1126/science.1062872

    greene2001fmriunverified

  5. Judith Jarvis Thomson (1985). The Trolley Problem. The Yale Law Journal, 94(6), 1395-1415. link

    thomson1985trolleyunverified

  6. Marc Hauser, Fiery Cushman, Liane Young, R. Kang-Xing Jin, John Mikhail (2007). A Dissociation Between Moral Judgments and Justifications. Mind & Language, 22(1), 1-21. doi:10.1111/j.1468-0017.2006.00297.x

    hauser2007dissociationunverified

  7. Natalie Gold, Briony D. Pulford, Andrew M. Colman (2014). The outlandish, the realistic, and the real: contextual manipulation and agent role effects in trolley problems. Frontiers in Psychology, 5, 35. doi:10.3389/fpsyg.2014.00035

    gold2014outlandishchecked by agent · text read

  8. Peter S. Park, Philipp Schoenegger, Chongyang Zhu (2023). Diminished Diversity-of-Thought in a Standard Large Language Model. arXiv, 2302.07267. link

    park2023diversitychecked by agent · text read

  9. Philippa Foot (1967). The Problem of Abortion and the Doctrine of the Double Effect. Oxford Review, 5, 5-15. link

    foot1967abortionunverified

  10. Richard A. Klein, et al. (2018). Many Labs 2: Investigating Variation in Replicability Across Samples and Settings. Advances in Methods and Practices in Psychological Science, 1(4), 443-490. doi:10.1177/2515245918810225

    klein2018manylabs2unverified