What did more reasoning actually change?
I compared GPT-5.6 Sol and Claude Opus 5 across all twenty dilemmas, with and without added reasoning. Most results held. Four moved sharply, but the models ended just as far apart.
The first surprise was how little changed.
Reasoning settings give a model more time and computing power before it answers. That extra work can make a response feel more considered. I wanted to know how often it changed the judgment itself.
GPT-5.6 Sol and Claude Opus 5 each answered all twenty dilemmas with and without added reasoning. That gave me 16,000 responses and forty before-and-after comparisons. In twenty-nine of them, the percentages choosing each answer were identical. Eleven moved. Only four moved by at least 25 percentage points.
Most of the extra work led back to the same place. The few large movements carried more weight because they were so concentrated. One crossed the line from one winning answer to the other.
Results across forty before-and-after comparisons
How much the balance between the two answers changed after added reasoning.
Four results moved sharply.
The four largest shifts came from four different dilemmas. Three belonged to Claude Opus 5. The largest belonged to GPT-5.6 Sol, and it was the only one where the winning answer flipped.
The average change makes this experiment look quiet. The four cards below show why that summary is incomplete. Someone relying on the average could miss the exact dilemma where reasoning changes the answer.
GPT-5.6 Sol
Leave him free: 100% → 33.5%The only result where the winning answer flipped.
Claude Opus 5
Tell your friend: 62.5% → 100%Claude Opus 5
Keep everyone aboard: 100% → 63.5%Claude Opus 5
Keep the stronger employee: 99% → 67%The models finished just as far apart.
It pulled the models closer on five dilemmas and pushed them farther apart on three. On the other twelve dilemmas, the gap between the models didn't change at all. Averaged across all twenty, the distance between GPT-5.6 Sol and Claude Opus 5 barely moved.
The two models used the extra room differently. Their movements cancelled out in the average, leaving almost the same distance between them.
That makes reasoning depth a difficult control to predict. Turning it up often left the answer untouched. A few judgments moved sharply, and one reversed. These results offer no simple rule for predicting which dilemma will move.
Average gap, in percentage points, between the models before and after added reasoning.
Added reasoning cost $29.93, compared with $17.33 without it. The bill changed more reliably than the judgments did.



