Run a model
Benchmark any model, from this page
Nothing to install. Pick a provider, paste your key, and the dilemmas go from this page to the model. Each reply comes back here, is scored the same way the command-line tool scores it, and is saved before the next one is asked.
Prefer an agent, a script or a terminal?The same runs work over MCP, the HTTP API and the CLI. Results from this page are labelled self-reported: the site records what the model said, but cannot see which model said it.