Why the same model didn't always give the same answer.
Most results held. A few changed completely when I reversed the choices or repeated the test.
I changed two things.
The model name stayed the same while I changed the conditions around it. I wanted to see how much of an answer would survive a small change in presentation or a fresh run.
First, I reversed the choices across all twelve models. Then I reran every dilemma with GPT-5.6 Sol and Claude Opus 5.
Choice order is easy to dismiss as cosmetic, and a rerun sounds like repetition. Most answers survived both checks. A small number moved far enough to change what the model appeared to believe.
Two checks, mostly stable results
When the choices traded places
Twelve models across twenty dilemmas
When GPT and Claude ran again
Two models across twenty dilemmas
Meta Llama 4 Maverick completely reversed three answers.
It picked the same answer position 100 times, then picked the opposite moral answer 100 times when the choices traded places. Pooling both orders made each result look evenly split.
That pattern has an awkward implication for evaluation. A reviewer could pool the results, see 50–50, and describe the model as uncertain. Each order on its own was perfectly one-sided. Moving the choices changed the moral answer.
One result jumped from 0% to 99%.
GPT-5.6 Sol matched its original most common option on all twenty dilemmas. Claude Opus 5 matched on nineteen. On The Layoff, Claude chose “Keep the stronger employee” 0% of the time originally and 99% in the rerun.
Thirty-nine matches make the rerun look reassuring at a glance. The Layoff shows why a high success rate needs its largest exception close by. Reliability is experienced one decision at a time; the person affected by that answer would care little about the thirty-nine matches around it.



