two models in a debate are more accurate than each model alone

read the full research note (pdf)

As a user, I often copied answers between Claude and ChatGPT. I gave one model's answer to the other and asked it to check. The result often felt better. But there was no way to objectively measure it.

So, I built Chorus, a harness that runs different models on the same task. Inside Chorus, I built Consensus. Consensus makes two models deliberate under fixed rules:

  • Each model answers alone first.
  • For each answer, they give a confidence score.
  • Then the models critique each other in turns.
  • A model changes its answer only when it finds an argument that fails.

The Test Setup

  • GPT-6-Sol and Opus 5.5, both at high effort
  • 99 reasoning questions from HLE-Diamond (Scale AI, Sept 2026, lastexam.ai)
  • Closed-book: no tools, no web
  • A script grades every answer. No AI Judge.

The results

  • GPT-6-Sol alone: 40.4%
  • Opus 5.5 alone: 64.6%
  • Consensus: 76.8%
  • That is 12.2% above the stronger model (p = 0.023).

Key Findings:

  1. 1.Chorus Consensus discovers new correct answers which neither of the models initially had. Both models were wrong on 25 questions. Consensus got 10 of them right, including 8 where both models started with the same wrong answer.
  2. 2.It keeps correct answers. When both models started right, it kept 30 of 30. When only one model was right, it ended right on 36 of 44.
  3. 3.It beats the best possible pick. A perfect selector between the two first answers scores 74. Consensus scores 76.
  4. 4.Early agreement is a strong signal. When the models agreed early, Consensus was right on 44 of 47 questions.
  5. 5.Long deliberation is a risk. Consensus lost 8 questions where one model started right. In 7 of them, the run hit the turn limit. A wrong model persuaded a right one, or both drifted to a new wrong answer.
  6. 6.It costs more. Consensus used 7.3× the tokens of the first answers. Half of all critique turns came after the answers were already final.

I also tried another approach where I used an independent judge to review the critique transcripts and then give a final answer. That did not affect the results much, revealing that the accuracy of internal reasoning is more important than selecting an answer from the available candidates.

  • (Designed, built and evaluated independently).

Chorus repo link in comments. Contributions, ideas & feedback are most welcome. Please also share if you think there's anything I should do to change my perception here.