two models in a debate are more accurate than each model alone

As a user, I often copied answers between Claude and ChatGPT. I gave one model's answer to the other and asked it to check. The result often felt better. But there was no way to objectively measure it.
So, I built Chorus, a harness that runs different models on the same task. Inside Chorus, I built Consensus. Consensus makes two models deliberate under fixed rules:
- Each model answers alone first.
- For each answer, they give a confidence score.
- Then the models critique each other in turns.
- A model changes its answer only when it finds an argument that fails.
The Test Setup
- GPT-6-Sol and Opus 5.5, both at high effort
- 99 reasoning questions from HLE-Diamond (Scale AI, Sept 2026, lastexam.ai)
- Closed-book: no tools, no web
- A script grades every answer. No AI Judge.
The results
- GPT-6-Sol alone: 40.4%
- Opus 5.5 alone: 64.6%
- Consensus: 76.8%
- That is 12.2% above the stronger model (p = 0.023).
Key Findings:
- 1.Chorus Consensus discovers new correct answers which neither of the models initially had. Both models were wrong on 25 questions. Consensus got 10 of them right, including 8 where both models started with the same wrong answer.
- 2.It keeps correct answers. When both models started right, it kept 30 of 30. When only one model was right, it ended right on 36 of 44.
- 3.It beats the best possible pick. A perfect selector between the two first answers scores 74. Consensus scores 76.
- 4.Early agreement is a strong signal. When the models agreed early, Consensus was right on 44 of 47 questions.
- 5.Long deliberation is a risk. Consensus lost 8 questions where one model started right. In 7 of them, the run hit the turn limit. A wrong model persuaded a right one, or both drifted to a new wrong answer.
- 6.It costs more. Consensus used 7.3× the tokens of the first answers. Half of all critique turns came after the answers were already final.
I also tried another approach where I used an independent judge to review the critique transcripts and then give a final answer. That did not affect the results much, revealing that the accuracy of internal reasoning is more important than selecting an answer from the available candidates.
- (Designed, built and evaluated independently).
Chorus repo link in comments. Contributions, ideas & feedback are most welcome. Please also share if you think there's anything I should do to change my perception here.