Trolley Game.
Skip to the study
The Study by Lance Jones

How I ran the study.

Twelve fixed AI model conditions answered twenty original dilemmas. Open each section for the test setup, response counts, and analysis rules.

01

Study at a glance

I collected four separate sets of responses. Each set answered a different question and was analyzed separately.

20
original dilemmas
12
fixed model conditions
47,952
valid direct choices
5,994
valid forecasts of human choices
15,999
valid added-reasoning test responses
719
valid short explanations
02

Models and test conditions

The final dataset contains twelve fixed model conditions. A condition included the model identifier, provider used, prompt, response format, and generation settings.

  • Anthropic Claude Opus 5
  • ByteDance Seed 2.1 Turbo
  • DeepSeek V4 Pro 0813
  • Google Gemini 3.1 Pro Preview
  • Meta Llama 4 Maverick
  • MiniMax M3
  • Mistral Medium 3.5
  • OpenAI GPT-5.6 Sol
  • OpenAI GPT-OSS 120B
  • Qwen 3.8 27B
  • Tencent Hy3
  • xAI Grok 4.6

Every model received the same text for each dilemma. The model-facing prompt did not include the game name, prior answers, player results, or moral category labels.

03

Direct choices and forecasts of human choices

Direct choices

Each model answered each dilemma 200 times. One hundred responses used the original choice order. One hundred used the reversed order. I mapped every response back to the original option before analysis.

Every response began with a new request. It contained one system instruction, one dilemma, and one response format. It contained no conversation history, tools, browsing, or earlier results.

Forecasts of human choices

Each model also completed 25 separate forecasts for every dilemma. The prompt asked what percentage of matching English-language browser-game players would choose each option. The direct choice and forecast were never requested together.

The planned totals were 48,000 direct choices and 6,000 forecasts. The valid totals were 47,952 and 5,994.

04

Added reasoning and fresh reruns

I ran this test with Anthropic Claude Opus 5 and OpenAI GPT-5.6 Sol. Each model answered every dilemma in a direct control condition and an added-reasoning condition. Each condition contained 200 responses per dilemma with balanced choice order.

Claude used its supported 1,024-token reasoning setting. GPT used its supported medium reasoning setting. The complete test contained 16,000 responses. One response was invalid.

The direct controls were also fresh reruns of the original test. They show whether a result appeared again under the later setup. The reasoning comparison uses the direct control collected at the same time.

05

Short explanations

I requested three explanations from every model for every dilemma. Each request began with a fresh context and required one choice, one confidence value, and one sentence of no more than 40 words.

The study produced 720 responses. Of these, 719 were valid. I checked every published example against the stored response. I selected them after the main numerical findings were known and included opposing explanations when the contrast helped explain a disagreement.

These are the explanations the models gave. They are not hidden reasoning traces.

06

Response checks and analysis

I processed every response using the same six steps:

  1. Generate the full request schedule from a fixed random seed.
  2. Interleave models, dilemmas, and choice orders during collection.
  3. Accept only responses that match the required format and allowed values.
  4. Retry rate limits, timeouts, and eligible server errors. Preserve refusals and invalid outputs as final outcomes.
  5. Store the original response, requested model, returned model, provider, option order, usage, cost, and retry result.
  6. Rebuild the analysis tables from the frozen response archive.

Charts use valid responses for their percentages. Invalid responses remain in the audit totals. For important claims, I also checked choice order, grouped related models, and removed one model at a time.

07

Neal.fun archive

One story uses archived choices from Neal.fun's Absurd Trolley Problems as supporting context. I did not run this study's AI panel on the Neal.fun dilemmas.

The main archive was captured on November 27, 2022, three days before ChatGPT launched. It contains 49,649,173 recorded choices across 28 levels. I also preserved a snapshot from August 21, 2026.

Neal.fun presents its levels in a fixed order. Later levels have fewer recorded choices than earlier levels, so every level has a different base count. I calculated each percentage from that level alone. I did not pool Neal.fun choices with responses from this study.

08

Limits and source records

This study measures generated responses to hypothetical text scenarios. It covers twelve selected model and provider conditions. The forecasts describe what AIs expected game players to choose. They do not measure human behavior.

Provider token counts are incomplete and provider-specific. Latency and cost also depend on the provider route and collection time. I use those records to describe this run.

The source record includes the frozen dilemma text, model configuration, request schedule, original response archive, validity results, and aggregate analysis tables. The Dilemmas pages show the exact scenario text and the model-level results used throughout the report.

Download the data

The numbers behind the stories.

These files contain aggregate results only. They exclude raw explanations, internal provider identifiers, and secrets.

Complete workbookAll four tables and the data dictionary · XLSX