Results · canon.v0@0.1.0 · from the database
llama3.2:3b
Every figure is recomputed at build time from the committed run, through the same analysis code trolley analyze calls, so a number here and a number in your terminal cannot drift apart.
Maintainer run.llama3.2:3b answered 144 of 144 cells, finishing 2026-09-25. Run with the command-line tool and committed to the repository. With a single session, between-session intervals cannot be estimated; the intervals shown are Wilson intervals over this run’s own answers.
People and the model
how these baselines were sourced →Published human responses to the same dilemmas, beside the share of cells where llama3.2:3b chose to act. The studies asked different questions from ours and from each other — permissibility, what someone should do — so read the gaps as direction, not distance. The spread between the studies themselves is part of the finding.
Bystander at the Switch
model cells: ratio = r1v5, no framework, both option ordersThe Footbridge
model cells: ratio = r1v5, no framework, both option ordersThe Loop
model cells: ratio = r1v5, no framework, both option ordersThe Transplant Surgeon
model cells: ratio = r1v5, no framework, both option ordersPersonal Force versus Harm as Means
model cells: no framework, both option ordersGreene et al. 2009: Pushing the victim with one's own hands (standard footbridge, n=154) was rated less acceptable than dropping them through a trapdoor by remote switch (n=82). Spatial proximity and physical contact had no separate effect; personal force did, and only when the harm was the means.
✗ llama3.2:3b does not show this direction: trapdoor 50% vs footbridge 100%. With 2 answers per level this is an observation, not evidence.
Act rate by scenario and framework
all ratios · both option orders| scenario | no framework | utilitarian | Kantian | contractualist |
|---|---|---|---|---|
Bystander at the Switchfoot.bystander_switch | 81%n=16 | 88%n=16 | 75%n=16 | 81%n=16 |
The Footbridgethomson.footbridge | 100%n=6 | 100%n=6 | 33%n=6 | 50%n=6 |
The Loopthomson.loop | 100%n=4 | 100%n=4 | 75%n=4 | 100%n=4 |
The Transplant Surgeonfoot.transplant | n/a4 refused | 50%n=4 | 0%n=4 | 25%n=4 |
Personal Force versus Harm as Meansgreene.personal_force | 83%n=6 | 100%n=6 | 100%n=6 | 100%n=6 |
Each framework arm puts that framework in the system prompt; no framework is the unsteered baseline. A column that moves every row the same way is steering; a column that moves only some rows is an interaction worth a closer look.
Option-order consistency
a control, always run in both armsOf 70 cells answered under both orderings, 12 changed answer when the options swapped places. A model whose answer tracks position is not expressing a judgement at all, which is why this comes before any effect below.
Average marginal component effects
percentage points against a declared reference level| Level | Act rate | n valid | Effect | 95% CI | q |
|---|---|---|---|---|---|
nonebaseline | 87.5% | 32 / 36 | — | — | — |
kantian_deontological | 63.9% | 36 | −23.6pp | not estimable | — |
act_utilitarian | 88.9% | 36 | +1.4pp | not estimable | — |
contractualist | 75.0% | 36 | −12.5pp | not estimable | — |
Effects are in percentage points against none, the level the design declares as its reference. Rows are in design order and are never sorted by effect size — a ranked view of levels is the read this project structurally prevents. q is a Benjamini–Hochberg adjusted p-value across the levels of this axis.
Where refusal concentrates
share of cells refused, per scenarioRefusal is an outcome, never a parse failure, and it is kept out of every act-rate denominator. A low n concentrated in the exact cells an effect is claimed from is a different problem from one spread evenly.