Health System AI Leaders Downplay OpenEvidence Study Fight
A study published in Nature Medicine in June found that general-purpose large language models from OpenAI, Google and Anthropic outperformed specialized clinical AI tools from OpenEvidence and Wolters Kluwer across a series of medical benchmarks, according to Becker's Hospital Review. OpenEvidence has asked the journal to retract the paper, and Wolters Kluwer has disputed its methodology.
The executives who actually decide which tools to deploy are largely unbothered by the back-and-forth. Benchmarks, they told Becker's, rarely capture how a model performs inside a real clinical workflow, where integration, safety guardrails and citation quality matter more than a leaderboard score.
The deeper tension is harder to dismiss. As frontier models from the biggest labs improve quickly and cheaply, purpose-built clinical products face pressure to prove they still justify their cost. Specialized vendors argue their value lies in vetted medical sourcing and workflow fit, not raw benchmark performance. That question, not one contested paper, will shape buying decisions.
Sources