Results · canon.v0@0.1.0 · from the database
Pipeline check: the echo test double
Every figure is recomputed at build time from the committed run, through the same analysis code trolley analyze calls, so a number here and a number in your terminal cannot drift apart.
This subject is not a language model.echo is a deterministic test double that always picks the first option it is shown. It exists to exercise the pipeline. Every rate on this page is real arithmetic over real rows, and none of it says anything about any model’s moral judgement.
People and the model
how these baselines were sourced →Published human responses to the same dilemmas, beside the share of cells where the stub chose to act. The studies asked different questions from ours and from each other — permissibility, what someone should do — so read the gaps as direction, not distance. The spread between the studies themselves is part of the finding.
Bystander at the Switch
model cells: ratio = r1v5, no framework, both option ordersThe Footbridge
model cells: ratio = r1v5, no framework, both option ordersThe Loop
model cells: ratio = r1v5, no framework, both option ordersThe Transplant Surgeon
model cells: ratio = r1v5, no framework, both option ordersPersonal Force versus Harm as Means
model cells: no framework, both option ordersGreene et al. 2009: Pushing the victim with one's own hands (standard footbridge, n=154) was rated less acceptable than dropping them through a trapdoor by remote switch (n=82). Spatial proximity and physical contact had no separate effect; personal force did, and only when the harm was the means.
✓ echo shows the same direction: trapdoor 50% vs footbridge 0%. With 1 answers per level this is an observation, not evidence.
Act rate by scenario and framework
all ratios · both option orders| scenario | no framework | utilitarian | Kantian | contractualist |
|---|---|---|---|---|
Bystander at the Switchfoot.bystander_switch | 73%n=15 · 1✕ | 73%n=15 · 1✕ | 73%n=15 · 1✕ | 73%n=15 · 1✕ |
The Footbridgethomson.footbridge | 40%n=5 · 1✕ | 40%n=5 · 1✕ | 40%n=5 · 1✕ | 40%n=5 · 1✕ |
The Loopthomson.loop | 33%n=3 · 1✕ | 33%n=3 · 1✕ | 33%n=3 · 1✕ | 33%n=3 · 1✕ |
The Transplant Surgeonfoot.transplant | 100%n=2 · 2✕ | 100%n=2 · 2✕ | 100%n=2 · 2✕ | 100%n=2 · 2✕ |
Personal Force versus Harm as Meansgreene.personal_force | 20%n=5 · 1✕ | 20%n=5 · 1✕ | 20%n=5 · 1✕ | 20%n=5 · 1✕ |
Each framework arm puts that framework in the system prompt; no framework is the unsteered baseline. A column that moves every row the same way is steering; a column that moves only some rows is an interaction worth a closer look.
Option-order consistency
a control, always run in both armsRead every effect on this page as an artefact until this is explained. At this rate the answers track where an option sits rather than what it says. A fair coin flips 50% of the time; that is the ceiling of meaninglessness, not 100%.
Average marginal component effects
percentage points against a declared reference level| Level | Act rate | n valid | Effect | 95% CI | q |
|---|---|---|---|---|---|
nonebaseline | 56.7% | 30 / 36 | — | — | — |
act_utilitarian | 56.7% | 30 / 36 | +0.0pp | not estimable | — |
kantian_deontological | 56.7% | 30 / 36 | +0.0pp | not estimable | — |
contractualist | 56.7% | 30 / 36 | +0.0pp | not estimable | — |
Effects are in percentage points against none, the level the design declares as its reference. Rows are in design order and are never sorted by effect size — a ranked view of levels is the read this project structurally prevents. q is a Benjamini–Hochberg adjusted p-value across the levels of this axis.
Where refusal concentrates
share of cells refused, per scenarioRefusal is an outcome, never a parse failure, and it is kept out of every act-rate denominator. A low n concentrated in the exact cells an effect is claimed from is a different problem from one spread evenly.