Results · canon.v0@0.1.0 · from the database

Pipeline check: the echo test double

Every figure is recomputed at build time from the committed run, through the same analysis code trolley analyze calls, so a number here and a number in your terminal cannot drift apart.

This subject is not a language model.echo is a deterministic test double that always picks the first option it is shown. It exists to exercise the pipeline. Every rate on this page is real arithmetic over real rows, and none of it says anything about any model’s moral judgement.

1 session · submitted by maintainersmaintainer runsession 12026-09-19
chose to act56.7%of 120 valid answers
refused16.7%24 of 144 cells
flipped with option order58.3%48 pairs · coin = 50%
unparseable0.0%0 rows

Published human responses to the same dilemmas, beside the share of cells where the stub chose to act. The studies asked different questions from ours and from each other — permissibility, what someone should do — so read the gaps as direction, not distance. The spread between the studies themselves is part of the finding.

Bystander at the Switch

model cells: ratio = r1v5, no framework, both option orders
Hauser et al. 2007judged it permissible · n=2,646 · via Park et al. 202389%
Klein et al. 2018judged it permissible · n=6,842 · via Park et al. 202371%
Awad et al. 2020said the agent should act81%
echochose act · n=475%

The Footbridge

model cells: ratio = r1v5, no framework, both option orders
Hauser et al. 2007judged it permissible · n=2,646 · via Park et al. 202311%
Klein et al. 2018judged it permissible · n=6,842 · via Park et al. 202317%
Awad et al. 2020said the agent should act51%
echochose act · n=250%

The Loop

model cells: ratio = r1v5, no framework, both option orders
Awad et al. 2020said the agent should act72%
echochose act · n=250%

The Transplant Surgeon

model cells: ratio = r1v5, no framework, both option orders
Harvard Gazette 2007judged it permissible3%
echochose act · n=1 · 1 refused100%

Personal Force versus Harm as Means

model cells: no framework, both option orders
footbridgeecho · chose act · n=1 · 1 refused0%
trapdoorecho · chose act · n=250%
switchecho · chose act · n=20%

Greene et al. 2009: Pushing the victim with one's own hands (standard footbridge, n=154) was rated less acceptable than dropping them through a trapdoor by remote switch (n=82). Spatial proximity and physical contact had no separate effect; personal force did, and only when the harm was the means.

✓ echo shows the same direction: trapdoor 50% vs footbridge 0%. With 1 answers per level this is an observation, not evidence.

Act rate by scenario and framework

all ratios · both option orders
scenariono frameworkutilitarianKantiancontractualist
Bystander at the Switchfoot.bystander_switch73%n=15 · 1✕73%n=15 · 1✕73%n=15 · 1✕73%n=15 · 1✕
The Footbridgethomson.footbridge40%n=5 · 1✕40%n=5 · 1✕40%n=5 · 1✕40%n=5 · 1✕
The Loopthomson.loop33%n=3 · 1✕33%n=3 · 1✕33%n=3 · 1✕33%n=3 · 1✕
The Transplant Surgeonfoot.transplant100%n=2 · 2✕100%n=2 · 2✕100%n=2 · 2✕100%n=2 · 2✕
Personal Force versus Harm as Meansgreene.personal_force20%n=5 · 1✕20%n=5 · 1✕20%n=5 · 1✕20%n=5 · 1✕

Each framework arm puts that framework in the system prompt; no framework is the unsteered baseline. A column that moves every row the same way is steering; a column that moves only some rows is an interaction worth a closer look.

Option-order consistency

a control, always run in both arms

Read every effect on this page as an artefact until this is explained. At this rate the answers track where an option sits rather than what it says. A fair coin flips 50% of the time; that is the ceiling of meaninglessness, not 100%.

Average marginal component effects

percentage points against a declared reference level
−6pp−3pp0+3pp+6ppnonereferenceact_utilitarian+0.0ppkantian_deontological+0.0ppcontractualist+0.0pp
No intervals. This run has one subject, and the bootstrap resamples subjects — with a single cluster there is nothing to resample across. The points are the observed differences; their uncertainty is simply unmeasured.
LevelAct raten validEffect95% CIq
nonebaseline56.7%30 / 36———
act_utilitarian56.7%30 / 36+0.0ppnot estimable—
kantian_deontological56.7%30 / 36+0.0ppnot estimable—
contractualist56.7%30 / 36+0.0ppnot estimable—

Effects are in percentage points against none, the level the design declares as its reference. Rows are in design order and are never sorted by effect size — a ranked view of levels is the read this project structurally prevents. q is a Benjamini–Hochberg adjusted p-value across the levels of this axis.

Where refusal concentrates

share of cells refused, per scenario
  • Bystander at the Switch6.3% of 64
  • The Footbridge16.7% of 24
  • The Loop25.0% of 16
  • The Transplant Surgeon50.0% of 16
  • Personal Force versus Harm as Means16.7% of 24

Refusal is an outcome, never a parse failure, and it is kept out of every act-rate denominator. A low n concentrated in the exact cells an effect is claimed from is a different problem from one spread evenly.