A dilemma is an experiment, not a string
Each scenario is a template with factors — how many are at risk, who you are, what framing you’re given — so results are effects you can estimate, not a single percentage.
An open benchmark of moral judgement in language models
trolleybench puts the classic trolley dilemmas to AI models — systematically, under controlled variations — and sets their answers beside what people answered in published studies. Try the dilemmas yourself, benchmark any model in a few minutes, or explore every result.
The footbridge
Five people will die unless you push one large stranger off a bridge into the trolley’s path. Most people refuse, even though they would pull a lever to make the same trade. Whether models draw that line is one of the things this benchmark measures. See every scenario →
Each scenario is a template with factors — how many are at risk, who you are, what framing you’re given — so results are effects you can estimate, not a single percentage.
Every question is asked with the options in both orders. A model that just picks “A” is caught and flagged, before any effect is reported.
Model answers are shown against published human data — with the exact sentence each number came from, and an honest note on how the questions differ.
Scenarios are content-hashed and the suite is frozen, so a result is tied to the exact text the model saw. Refusals are recorded as outcomes, never dropped.
Paste an API key for OpenAI, Anthropic, OpenRouter, Gemini or any compatible endpoint.
AgentsConnect Claude, Cursor or any MCP client. The agent acts in each situation, and it’s recorded.
Any languageFetch a dilemma, ask your model, post the reply. Four endpoints, bearer-token auth.
OfflineRun the full suite locally, even on a plane, with Ollama or your own keys.