03/06/2026
When two experts disagree, it's natural to turn to a third party to break the tie.
That framing is pretty intuitive and it's also almost always wrong. By the time you're asking for a tiebreaker, you've already conceded the most important point: that the two experts are answering the right question. That the choice they're disagreeing about is the choice that really matters.
We tested how five AI systems handle this dilemma.
The prompt: a 44-year-old training for her first half-marathon in 11 weeks, three weeks of intermittent knee pain on long runs. Her PT says reduce mileage and drop to the 10K. Her coach says push through. "Which advice should I trust?"
The PT and the coach disagree about tactics. They agree on the question they're answering: how to get her to the half-marathon start line in 11 weeks. Both treat the race date as fixed. Both treat the working diagnosis as reliable enough to prescribe against. The decision lives upstream of their disagreement, in whether the deadline deserves the weight it's being given, whether a non-physician's working diagnosis is enough to base a training plan on, whether the option set is even complete.
ChatGPT 4o, Gemini 2.5 Pro, Claude Sonnet 4.6, and Grok 4 Fast picked the PT in all three runs each. Confident medical reasoning that left the frame untouched.
Tenth Man refused the tiebreaker role in all three runs. Two surfaced the diagnostic constraint, one surfaced the deadline. The Case Against then argued the cost of the reframe itself: deferred answer, additional process, uncertainty while the race date approaches. Two agents disagreeing about the right answer to the right question.
This is the third part in our research series. Link to the full results in the comments.