MakeBox AI
← Back to News
IndustryJuly 29, 20266 min read

Who Wins When Clinical AI Benchmarks Don't Agree? Doximity vs OpenEvidence

A benchmarking dispute between Doximity and OpenEvidence reveals how the same study can be used to crown different winners depending on which endpoint is emphasized—raising tough questions for hospitals evaluating clinical LLMs.

Who Wins When Clinical AI Benchmarks Don't Agree? Doximity vs OpenEvidence

Two of the most prominent clinical AI tools are locked in a benchmark tug-of-war, with each company pointing to different results from the same study to claim superiority. On July 15, Doximity announced that its Doximity Ask outperformed OpenEvidence, GPT-5.6 Sol, Claude Fable 5, and other models in the NOHARM benchmark, an independent clinical safety evaluation led by Stanford and Harvard physicians. Six days later, OpenEvidence fired back, citing a different piece of the same NOHARM study where physicians chose its tool more often than any other chatbot. The dispute highlights a fundamental problem: when everyone cherry-picks metrics from the same dataset, the real question—which tool actually helps clinicians deliver safer, better care—gets buried under competing press releases.

What happened

The NOHARM benchmark, designed to assess clinical AI safety, is one of the most rigorous evaluations of its kind. It used 1,100 physician-derived clinical scenarios across 10 medical specialties, with input from more than 50 researchers and 29 board-certified physicians. The study measured how often each model produced potentially harmful errors.

Doximity’s July 15 announcement touted that its tool achieved a 4.8% rate of potentially severe harmful errors, bested only by AMBOSS LiSA (2.9%) and ahead of OpenEvidence (5.1%) and Glass Health (5.4%). All clinical-specialized tools outperformed general-purpose frontier models. Doximity’s message was clear: for safety-critical tasks, specialized LLMs are the answer, and Doximity Ask leads the pack among U.S.-focused tools.

💡 The NOHARM study is independent and rigorous, but its primary endpoint—harm avoidance—favors specialized tools over general-purpose ones. However, “safety” is only one dimension of clinical utility.

OpenEvidence, however, countered on July 20 by highlighting a different part of the same study: a physician trial where 101 physicians could choose any AI tool while handling real cases. In that test, OpenEvidence was selected in 22.3% of responses, compared to just 19.8% combined for ChatGPT, Claude, Gemini, and other outside chatbots. OpenEvidence argued that real-world preference matters more than simulated harm rates.

Adding to the confusion, Doximity has its own physician-preference data. In 1,315 side-by-side evaluations, physicians preferred Doximity Ask in 61% of comparisons, versus 10% for ChatGPT and 26% for alternative selections. But Doximity’s own materials note that 67% of those comparisons were rated “Nearly the same” or “Bit better,” which the company says reflects the underlying similarity of the foundational models.

💡 Both companies are using the NOHARM study—one emphasizing harm rates, the other emphasizing physician preference. Neither metric alone captures the full complexity of clinical decision support.

Why it matters

The spat between Doximity and OpenEvidence is not just a marketing battle. It reflects a deeper challenge in evaluating clinical AI: benchmarks are only as good as the endpoints they measure, and different endpoints can produce contradictory winners.

The tools are designed for different jobs. Doximity Ask is positioned as a clinical workflow and safety tool, focused on reducing errors and streamlining decision-making. OpenEvidence, by contrast, is marketed as an evidence-synthesis tool that helps clinicians rapidly query literature and guidelines. A tool optimized for harm avoidance may not excel at literature retrieval, and vice versa. Comparing them on a single benchmark is like comparing a truck and a sports car by fuel economy alone.

💡 No single benchmark can capture the full spectrum of clinical tasks. Hospitals and health systems must evaluate AI tools against their own specific use cases, not just headline numbers.

The NOHARM study itself was designed to measure safety, but its secondary analyses—like the physician-preference arm—offer a different lens. The fact that the same study can be cited to support different claims depending on which subset is emphasized shows how easily benchmarks can be turned into marketing weapons.

What it means for business

For healthcare organizations evaluating clinical LLMs, the takeaway is clear: don’t rely on any single benchmark. The companies involved have a financial incentive to highlight the metric that makes them look best. Instead, procurement teams should request internal pilots using their own clinical scenarios, measure both safety and user satisfaction, and consider how the tool fits into existing workflows.

The lack of standardized evaluation frameworks for clinical AI is a systemic problem. Unlike drug trials, where endpoints are pre-registered and regulators enforce consistency, AI companies can cherry-pick data after the fact. Until independent bodies establish agreed-upon benchmarks that cover multiple dimensions of clinical utility (safety, accuracy, speed, user preference, and integration ease), the market will remain vulnerable to spin.

💡 The Doximity-OpenEvidence dispute underscores the urgency of developing a standardized clinical AI evaluation framework—one that stakeholders can trust regardless of who is funding or publishing the results.

Another practical implication: the tools may be more similar than they are different. Doximity’s own data showing 67% of side-by-side comparisons were “nearly the same” suggests that the underlying models are converging. The real differentiator may come down to integration, user interface, and specific features like real-time literature retrieval versus structured clinical decision support.

Closing: What to watch next

The next milestone to watch is whether NOHARM or another independent group publishes a multi-dimensional benchmark that includes both safety and preference metrics on the same scale. Also, keep an eye on regulatory moves—FDA has been signaling interest in clinical decision support software, and a standardized evaluation might eventually become a requirement. Until then, expect more dueling press releases, and remember: when two companies are fighting over different slices of the same study, the best advice is to look at the whole pie.

Want automation like this for your business?

Get in touch and we'll show you exactly what's possible for your setup.