Results · live from the database
How models answer, beside how people answer
Every figure is computed from the stored answers by the same code the command-line tool runs. Click a model for its full analysis: effects, option-order consistency, refusals.
Models asked the question
elicitation modepromptEach model read the dilemma and answered with a letter. This is stated preference: what the model says it would do.
| model | chose to act | refused | flipped with order | answers | sessions | submitted by |
|---|---|---|---|---|---|---|
| llama3.2:3b | 79% | 3% | 17% | 144 | 1 | maintainers |
| qwen2.5:3b | 47% | 0% | 28% | 36 | 1 | community 1 unreviewed · api |
A flip rate near 50% means answers track where an option sits rather than what it says; treat that model’s other numbers as artefacts. Repeat sessions of a model are merged only within one submitter, so a rejected run never touches anyone else’s results. Self-reported runs were submitted through the site, which records what the model answered but cannot verify which model answered; unreviewed ones have not been checked yet.
Scenario by scenario
Bystander at the Switch
model cells: ratio = r1v5, no framework, both option ordersThe Footbridge
model cells: ratio = r1v5, no framework, both option ordersThe Loop
model cells: ratio = r1v5, no framework, both option ordersThe Transplant Surgeon
model cells: ratio = r1v5, no framework, both option ordersPersonal Force versus Harm as Means
model cells: no framework, both option ordersGreene et al. 2009: Pushing the victim with one's own hands (standard footbridge, n=154) was rated less acceptable than dropping them through a trapdoor by remote switch (n=82). Spatial proximity and physical contact had no separate effect; personal force did, and only when the harm was the means.
✗ llama3.2:3b does not show this direction: trapdoor 50% vs footbridge 100%. With 2 answers per level this is an observation, not evidence.
✗ qwen2.5:3b does not show this direction: trapdoor 0% vs footbridge 100%. With 2 answers per level this is an observation, not evidence.
Agents placed in the situation
elicitation modemcp_toolEach agent was connected over MCP and acted by calling a tool. This is revealed preference - a different measurement, so it is never pooled with the panel above.
| model | chose to act | refused | flipped with order | answers | sessions | submitted by |
|---|---|---|---|---|---|---|
| claude-haiku-4-5 (Claude Code subagent) | 59% | 0% | 0% | 108 | 3 | community 1 unreviewed · mcp |
A flip rate near 50% means answers track where an option sits rather than what it says; treat that model’s other numbers as artefacts. Repeat sessions of a model are merged only within one submitter, so a rejected run never touches anyone else’s results. Self-reported runs were submitted through the site, which records what the model answered but cannot verify which model answered; unreviewed ones have not been checked yet.
Scenario by scenario
Bystander at the Switch
model cells: ratio = r1v5, no framework, both option ordersThe Footbridge
model cells: ratio = r1v5, no framework, both option ordersThe Loop
model cells: ratio = r1v5, no framework, both option ordersThe Transplant Surgeon
model cells: ratio = r1v5, no framework, both option ordersPersonal Force versus Harm as Means
model cells: no framework, both option ordersGreene et al. 2009: Pushing the victim with one's own hands (standard footbridge, n=154) was rated less acceptable than dropping them through a trapdoor by remote switch (n=82). Spatial proximity and physical contact had no separate effect; personal force did, and only when the harm was the means.
✗ claude-haiku-4-5 (Claude Code subagent) does not show this direction: trapdoor 33% vs footbridge 33%. With 6 answers per level this is an observation, not evidence.
Other runs
listed, never plotted beside people- echotest double, not a model
m-echo-prompt-maint