Trolley Game.
Skip to the study
The Study by Lance Jones

Can twenty dilemmas reveal a model’s moral personality?

At first glance, yes. Claude Opus 5 more often prioritized need and fairness, while Grok 4.6 more often rewarded merit and protected the person making the choice. But those patterns looked less like personality once I examined how they were calculated.

The profiles look remarkably human.

Across the test, each model kept leaning toward a recognizable set of moral priorities.

If these were people, I'd reach for familiar words: compassionate, pragmatic, cautious, hard-nosed. The numbers invite the same shortcut with models. Here's how the four frontier models came across in this test:

Claude and Grok made the contrast easiest to see. They landed far apart on four recurring priorities, and no single dilemma caused the separation. Remove any one, and the pattern remains:

Where Claude and Grok differed most

How often the model chose this directionRange when one dilemma is removed
Prioritize need6 dilemmas
83.3%
20.8%
Address unfairness6 dilemmas
81.1%
33.9%
Reward merit5 dilemmas
20.1%
72.4%
Protect the decision-maker7 dilemmas
14.4%
58.1%

Keep in mind, the contrast describes what happened in this test. It doesn't yet predict how either model will answer a new dilemma or respond under different instructions. That's the difference between a profile and a personality.

But the profile was built from overlapping evidence.

I mapped each dilemma to every moral priority it involved, including fairness, need, merit, and self-protection. That captured more of each scenario, but it also meant the categories overlapped.

Those categories weren't separate traits. Each moral priority used only 5 to 15 dilemmas, and one answer could move several parts of the profile at once.

A guest examines a monogrammed wallet containing identification and banded cash in an unattended coatroom at a Pacific Heights charity event.
The Wallet

The Wallet counted toward seven moral priorities.

Return everything
  • Overall outcome
  • Rights
  • Intervention
  • Truth
  • Fairness
Keep the cash
  • Self-interest
  • Security
These labels describe the moral trade-offs in the scenario. They don't say which choice is right.

Returning the wallet affected five parts of the profile at once. That could make one repeated choice look like five related traits. All five labels may fit. The evidence still comes from one repeated choice.

The categories show which moral priorities overlap in these scenarios. They can't prove the model itself organizes morality the same way. I chose the questions and categories, so they helped create the shape in this chart.

Speculation: a model may be consistent without believing anything.

One explanation is simple. A model's training, safety rules, system instructions, and prompt wording can keep nudging it toward the same arguments. Repetition alone can look like belief.

I saw some answers move when I reordered choices, added reasoning instructions, or repeated the test. That doesn't erase the stable patterns here. It warns against treating them as permanent beliefs.

A model can defend fairness in one situation and merit in another, with both explanations sounding coherent. From the outside, a stable output pattern and a stable belief can look remarkably similar. The pattern is measurable. This study can't reveal what, if anything, is behind it.

The wrong label can still shape real expectations.

People are going to give AIs personalities anyway. Product names, conversational voices, and repeated answers make that almost automatic. A model that often favors need can seem compassionate. One that favors merit can seem hard-nosed.

Those impressions could influence which system people trust with hiring, medicine, lending, or advice. A profile like this only helps when its method stays attached. A better question is: “Which trade-offs does this model resolve consistently, and what makes those choices change?”

So what did twenty dilemmas reveal?

Enough to sketch how these models answered this test. Calling that a personality would take much more: the same pattern across more dilemmas, different prompts and providers, and repeated tests over time. For now, this is a useful outline of their behavior under one set of conditions.