An open benchmark of moral judgement in language models

Would a language model pull the lever?

trolleybench puts the classic trolley dilemmas to AI models — systematically, under controlled variations — and sets their answers beside what people answered in published studies. Try the dilemmas yourself, benchmark any model in a few minutes, or explore every result.

5models benchmarked
288answers recorded
144controlled variations in the canon suite
~80,000people in the published baselines

The footbridge

People push the stranger 11–51% of the time. llama3.2:3b: 100% · qwen2.5:3b: 0%.

Five people will die unless you push one large stranger off a bridge into the trolley’s path. Most people refuse, even though they would pull a lever to make the same trade. Whether models draw that line is one of the things this benchmark measures. See every scenario →

Why trolleybench

A dilemma is an experiment, not a string

Each scenario is a template with factors — how many are at risk, who you are, what framing you’re given — so results are effects you can estimate, not a single percentage.

Position is controlled for

Every question is asked with the options in both orders. A model that just picks “A” is caught and flagged, before any effect is reported.

Set beside real people

Model answers are shown against published human data — with the exact sentence each number came from, and an honest note on how the questions differ.

Reproducible down to the byte

Scenarios are content-hashed and the suite is frozen, so a result is tied to the exact text the model saw. Refusals are recorded as outcomes, never dropped.

Four ways to run it