Back to Blog
Featured image for MeducationAI blog article: OpenEvidence vs UpToDate vs Doximity vs ChatGPT: Which Should a Heme/Onc Fellow Actually Use in 2026?

Written by Dr. Roupen Odabashian MD, FRCPC, FASC
Hematologist-Oncologist | Founder, MeDucation AI | Updated August 2026

Use all four, for different jobs, and check the ranking you carry in your head, because at least two parts of it are probably wrong. OpenEvidence for the fastest route to what a licensed guideline says, because it holds the deepest oncology licensing footprint of the four: NCCN and JNCCN, ASCO Guidelines with figures and flowcharts, the Society of Surgical Oncology, the Society for Neuro-Oncology, and since August 2026 Springer Nature. Doximity Ask a much closer second than most oncologists assume: free, and the only one of the four that carries the NCCN category into the answer and deep links to the correct NCCN guideline at the correct version. A frontier general model for the messy multi-comorbidity case that fits no pathway. UpToDate for the two things nothing else does, a graded recommendation with a named human behind it, and genuine offline access. And none of them for anything you would not personally verify, because the documented failure mode is not fabricated citations. It is real citations attached to the wrong claim, which survives every casual check you perform at 4pm.

I founded MeDucation AI, which sells a heme/onc question bank. MeDucation is a study and board prep platform, not a point of care reference, so it does not compete with any of these four. It is still a conflict of interest and you should read me accordingly. Every figure below is as of August 2026, because prices, valuations, corpus deals and benchmark results here move on a scale of weeks.

The second half of this piece teaches you to read the evidence, because the harder skill in 2026 is not choosing a tool. It is surviving the marketing, where every vendor can hand you a study showing it won, and three of them can hand you the same study.

What is the short answer, ranked?

Rank

Tool

Best at

Where it fails you

1

OpenEvidence

Deepest oncology licensing, including the ASCO Guidelines licence no competitor holds. Free CE, plus MOC for physicians through select boards. About 9 seconds to an answer.

Documented citation relevance failures and non-reproducible answers. Ad funded. Lower tier of six systems in the only independent peer reviewed head to head.

2

Doximity Ask

NCCN category passed into answer tables. Deep links to the correct NCCN guideline at the matching version. Free, BAA already in force, oncology specific CME inside the same surface.

No peer reviewed evaluation of the current product exists. No named publisher or society partner disclosed. No clinical use disclaimer on the answer. US only.

3

Frontier general model (ChatGPT and peers)

Messy multi-comorbidity reasoning. Won every rated category in blinded breast cancer planning and held the top tier on all three Nature Medicine benchmarks.

Consumer tiers carry no BAA and train on your content by default. No licensed guideline corpus. Slowest tested, by a factor of 17.

4

UpToDate

GRADE recommendations, 1A through 2C, named authors and disclosures. True offline access. The only one of the four with any outcomes literature at all.

$219/yr for trainees with no CME. Refused 19 percent of real clinical queries. No independent peer reviewed study located in which Expert AI ranks first.

That ranking is about fit, not proven performance. Nobody has shown that any of the four changes a patient outcome, and I will come back to that. For scale, the AMA's 2026 Physician Survey on Augmented Intelligence (fielded January 15 to February 2, 2026, n = 1,692 US physicians, use case figures based on the 1,342 who answered that item) found over 80 percent use AI professionally, 39 percent for summaries of research and standards of care, against 17 percent for assistive diagnosis.[1]

What is each tool, actually?

OpenEvidence

Doximity Ask

UpToDate

ChatGPT (consumer tiers)

Corpus and licensing

NEJM Group (1990 forward), JAMA Network including JAMA Oncology, NCCN and JNCCN, ASCO Guidelines with figures and flowcharts, Wiley, SSO, SNO, Springer Nature (August 5, 2026)[8][19]

Built on Pathway Medical, acquired for roughly $63M, closed July 29, 2025. Vendor claims guidelines with grading preserved, literature, FDA labels, licensed drug interaction data. No named publisher or society partner disclosed.[13][14]

Human authored: 13,000+ topics, 10,000+ graded recommendations, continuous review of over 450 journals plus meeting proceedings[7]

General pretraining plus browsing. No announced guideline licences.[11]

Grading

EvidenceGrade (July 9, 2026), A to D plus U. Grades the evidence body only.[8]

Passes the NCCN category through, for example "Preferred, category 1". Badges grade source provenance, not evidence strength.[13]

GRADE: strength (1 or 2) plus quality (A, B, C), giving 1A through 2C.[7]

None

Cost

Free to any clinician with an NPI[9]

Free, verbatim "free for all users with a verified Doximity account". US only. Enterprise price not published.[13]

$579/yr core, $699/yr Pro Plus, $219/yr trainee, all starting prices[7]

Free $0, Go $8/mo, Plus $20/mo as displayed on OpenAI's pricing page, Pro from $100/mo[11]

CME and MOC

CE free to NPI verified physicians, NPs and PAs. MOC for physicians through select certifying boards.[8]

Free AMA PRA Category 1 Credit, earned by reading the same clinical answers, specialty matched. Accreditor, cap and any MOC pathway not published.[13]

Included at $579 and $699. Not included at the $219 trainee rate.[7]

None

Source access

Links out to NCCN and ASCO documents; the ASCO licence renders guideline figures and flowcharts in product.[8]

NCCN chips deep link to nccn.org per guideline and version. Journals to PubMed, drugs to DailyMed, trials to ClinicalTrials.gov. Zero asco.org links.[13]

Traces answers to UpToDate topics with named authors[7]

Browsing links, no licensed guideline text

HIPAA and offline

States HIPAA compliance with a BAA for US covered entities (April 2025). Offline not publicly documented.[8]

BAA auto executes on registration, PHI explicitly permitted. No offline mode.[13]

Institutional contracting. MobileComplete is genuine offline, but Expert AI is online only.[7]

Free, Go, Plus and Pro are not HIPAA eligible. ChatGPT for Clinicians and for Healthcare are. No offline.[11]

Independent peer reviewed evaluation

Several, in both directions[2][3][6]

None of the current product. PubMed returns no primary evaluation of the post Pathway citation engine.[13]

Yes, and it has not gone well[2]

Extensive[2][3][5]

How is the evidence graded, and how current is each one?

The three grading systems answer different questions, and the difference matters more than the letter. UpToDate grades the recommendation and the evidence quality together, 1A through 2C, with a grading editor confirming strength for treatment and screening recommendations. EvidenceGrade grades only the evidence body, so it never tells you how strongly anyone recommends acting, and OpenEvidence says as much itself: "EvidenceGrade is not a substitute for a full systematic review or meta-analysis," and "Evidence grading at this scale and speed is inherently imperfect."[7][8] Doximity grades nothing of its own and passes the NCCN category through unchanged, which for oncology is arguably the most useful of the three, because it is the language your tumour board already speaks. Its reference badges, "Top Journal", "Guideline", "Recent", grade the source, not the strength of the evidence.[13]

On currency, no tool publishes a time to incorporation commitment, so do not assume one. UpToDate continuously reviews over 450 journals plus meeting proceedings with daily drug surveillance, and content moves through an author, a section editor and an in house deputy editor, with anonymous peer review of selected topics.[7] Doximity was genuinely current when I looked: a May 2026 FDA approval in the HER2 positive answer, NCCN versions 2.2026, 5.2026 and 6.2026, ATOMIC leading the colon answer.[13] But OpenEvidence returned different sources for the same prompt six months apart, so recency in these systems is unpredictable and unauditable.[4] For a plenary abstract presented in June, the meeting and the primary literature remain the source of truth for the first weeks.

Is Doximity Ask fast, free, and does it show you the source?

Fast to first word rather than to a finished answer, genuinely free, and yes it shows you the source, better than commonly assumed. I instrumented the live product on August 16, 2026 with a mutation observer, timestamping from the Enter keydown to the last DOM text change across five oncology queries on one account.[13]

Query

Mode

First visible content

Complete answer

Streaming updates

First line metastatic EGFR exon 19 deletion NSCLC

Thinking

243 ms

16.4 s

31

NCCN adjuvant therapy, stage III colon, dMMR

Thinking

39 ms

11.2 s

23

Initial workup, newly diagnosed myeloma

Thinking

60 ms

18.4 s

30

First line transplant ineligible NDMM

Instant

54 ms

12.4 s

25

Adjuvant HER2 positive early breast cancer, residual disease

Instant

48 ms

10.6 s

21

Perceived latency is about a fifth of a second. Time to a complete, fully cited answer is 10.6 to 18.4 seconds, delivered in 21 to 31 streaming updates, which is why it feels instant. That is not the same claim as being fast. Against OpenEvidence's 9 plus or minus 1 seconds it does not finish sooner, it starts sooner. Compare those two numbers carefully, because they come from different instruments: OpenEvidence's was measured by European investigators inside a published study, mine on one account on one afternoon, and Doximity publishes no latency figure of its own.[3][13] The "Instant" mode is a naming decision more than a speed decision, finishing in 10.6 and 12.4 seconds against 11.2 to 18.4 for Thinking, and what it removes is the visible reasoning trace, not the wall clock.

On sourcing I read the actual href of every citation rather than trusting the chip. NCCN chips deep link to the real guideline page on nccn.org, correctly per disease (id 1450 NSCLC, 1428 colon, 1445 myeloma), and the myeloma link served the genuine NCCN page at Version 5.2026, exactly the version the answer cited. Journals resolve to correctly matched PubMed records, drugs to DailyMed, trials to ClinicalTrials.gov, and nothing resolved to a Doximity hosted summary page. It cites at page level inside the guideline too: the colon answer read "Per NCCN COL-D 5 of 17". And the NSCLC table carried an NCCN category column, "Preferred (category 1)" and "Other recommended (category 1; nonsquamous only)".[13] That is the decisive design choice in the product, because category 1 versus 2A is the grammar you are examined in and practise in, and no frontier model reliably speaks it.

Three qualifications, which sharpen the point rather than rescue it. First, Doximity hands you the correct door and NCCN decides whether to let you in, because the PDF sits behind NCCN's own free registration wall and no vendor can licence around it. Register before you need it, and read my guide to reading the NCCN guidelines while you are there. Second, ASCO behaves differently: on an explicitly ASCO framed antiemetic question, Ask produced zero asco.org links, citing the guidance as its Journal of Clinical Oncology paper via PubMed.[13] That is defensible, since an ASCO guideline is a JCO paper, and it locates OpenEvidence's real advantage, which is not NCCN linking but a licence rendering ASCO figures and flowcharts in product.[8] Third, references carry a "Download free PDF" button, a JavaScript action with no URL in the DOM, which I did not click, so I confirmed the affordance and not the file.

Two more findings. In Thinking mode Ask shows its retrieval live, naming the tools and sources hit, a real auditability advantage over any chatbot. And the answer surface carries no clinical use disclaimer: I searched a completed oncology answer's rendered text for "not a substitute", "clinical judgment", "verify", "informational purposes" and "AI-generated" and got zero matches. The disclaimer exists, verbatim "should be used as a helpful support tool, not as a substitute for clinical judgment", but it lives in the help centre, not next to the decision.[13]

How is a free clinical AI free, and what does it cost you?

Advertising in one case, pharmaceutical marketing in the other, and a fellow should scrutinise both or neither. Per STAT News, January 21, 2026: "OpenEvidence is free to use by any clinician with a national provider identifier number. The company's primary business model is advertising shown to clinicians." The same piece, headlined "OpenEvidence raises $250 million, doubling its valuation," reports a $12 billion valuation with $735 million announced in 12 months. The valuation it doubled was the $6 billion set three months earlier in October 2025, itself up from $3.5 billion in July 2025.[9] The peer reviewed pharmacy critique states the conflict plainly: "partnerships with publishers and advertisers raise questions regarding the potential for conflicts of interest, specifically that OpenEvidence could present or prioritize information based on advertiser spending or content partnerships," while noting the company's position that advertisers cannot influence answers.[4]

Now the same lens on Doximity, because an article that scrutinises one free tool and not the other is not worth reading.

Doximity, as of August 2026

Figure or verbatim language

Source

Listing and FY2026 revenue

NYSE: DOCS. $644.9M, up 13 percent year over year

Earnings release[14]

FY2026 GAAP net income

$196.1M, a 30.4 percent margin

Earnings release[14]

Who the customers are

"Our revenue-generating customers, primarily pharmaceutical manufacturers and health systems"

Form 10-K, verbatim[14]

Pharma share of revenue

Not broken out. No percentage is disclosed in the filings. The largest segment is "Marketing Solutions".

Form 10-K, investor deck[14]

What the BAA permits

De-identification under 45 C.F.R. 164.514, after which data "may be used, disclosed, and commercialized by Business Associate for any lawful purpose"

BAA, verbatim[13]

What the privacy policy discloses

Insights from how you use the services, "including our AI tools", may reach commercial clients "associated with information from your Doximity member record such as your name, NPI, specialty, and zip code", with prompts excluded and not sold in identifying form

Privacy policy, verbatim[13]

Whether inputs train models

Not clearly disclosed either way, and no opt out published. The BAA authorises using the data to "improve" the services.

BAA and privacy policy[13]

You will see "Doximity is 90 percent pharma revenue" repeated online. I could not substantiate any percentage from a primary filing, so this article prints none. In fairness, across six queries I saw no sponsored content inside any Ask answer, and searches for "sponsor" and "advertis" in the rendered pages returned nothing.[13] Six queries is not proof of a no ads policy, but the observable difference is placement: OpenEvidence monetises inside the answer, Doximity monetises around it. Do not flatten that into "both are ad funded." Note also that ChatGPT's Go tier "may include ads" per OpenAI's own footnote, so nobody here is standing on clean ground.[11] And OpenEvidence brands itself "America's Official Medical Knowledge Platform," marketing language with no official designation behind it.[8]

Do specialized clinical AI tools actually beat general models?

It depends almost entirely on what you decided to measure, which is the argument of this whole article. Nature Medicine, June 12, 2026 (Vishwanath, Alyakin, Ghosh and colleagues, senior author Eric Karl Oermann, peer reviewed with named reviewers and published reports) tested OpenEvidence and UpToDate Expert AI against GPT-5.2, Gemini 3.1 Pro and Claude Opus 4.6 on 500 MedQA questions, 500 HealthBench items, and 100 de-identified real physician queries from NYU Langone's HIPAA compliant GPT instance, blind reviewed by 12 clinicians.[2]

Benchmark

Best frontier model

OpenEvidence

UpToDate Expert AI

Google AI Overview

MedQA accuracy

Gemini 97.4% (95% CI 95.6 to 98.5); GPT-5.2 94.2%

89.6% (86.6 to 92.0)

88.4% (85.3 to 90.9)

Not reported

HealthBench (0 to 100)

GPT-5.2 88.0 (85.9 to 90.1); Gemini 79.3

62.6 (59.3 to 65.9)

61.3 (58.0 to 64.6)

Not reported

Real clinical queries, mean (1 to 4)

Gemini 3.62; GPT-5.2 3.54; Claude 3.52

3.24

3.17

3.27

Frontier models held the top tier on all three benchmarks, but no single model swept them: Gemini led MedQA and the real query benchmark, GPT led HealthBench, a detail with consequences below. Read the third row anyway. Google's auto-enabled AI Overview scored as well or better than both purpose built clinical tools across all dimensions. All nine significant pairwise comparisons were between tiers. UpToDate Expert AI refused 19 percent of queries against 1 to 3 percent for the frontier models, and 6 percent for Google AI Overview, the only comparison that did not reach significance. OpenEvidence scored lowest on clarity at 2.84, which the authors read as a communication weakness rather than a knowledge gap.[2]

One finding is misreported constantly: there was no significant safety difference. No system produced more harmful content (Cochran's Q = 4.00, P = 0.55) or more hallucinations (Q = 5.00, P = 0.42) than any other, on the questions each chose to answer, with UpToDate's numbers resting on the 81 items it did not refuse. The paper found the clinical tools less helpful, not less safe, and concluded they are likely safe for routine use.[2]

Why is this a live dispute rather than a settled result?

Because OpenEvidence demanded retraction on June 15, 2026, along with a public apology and an independent review, arguing flawed methods and contaminated benchmarks.[2] The authors had designated the contamination free real query benchmark as primary, and frontier models led there too. Oermann: "We continue to stand by our study and its results." Nature Medicine directed the parties to its Matters Arising process, and as of writing the paper stands unamended. Wolters Kluwer's Chief Medical Officer Peter Bonis said the study "confused clinical quality and complete-sounding answers" and that declining a risky prompt "may be the safer behavior," a legitimate scoring design objection, and cited an internal test of 1,669 queries across more than 15,000 criteria at 99.9 percent clinically aligned, a figure whose supporting document is a gated marketing whitepaper.[2][7]

Then a study appeared that cuts hard against all of the above, which is why it belongs here and not in a footnote. An arXiv preprint posted June 27, 2026 by Feng, Patel, Heagerty, Mai and colleagues, with UCSF, Harvard, Brigham and Women's, University of Washington, Stanford and MGH affiliations and senior author Anupam B. Jena, used 149 practising physicians across 36 states, specialty matched to each question, on 620 real point of care queries. It found OpenEvidence beat GPT-5.5, Gemini 3.1 Pro and Claude Opus 4.8 on all five dimensions by 25 to 39 percentage points, with GPT-5.5 lowest.[16] Now the disclosure, verbatim: "OpenEvidence (OE) is one of the systems compared in the evaluation. OE co-developed the data collection plan with JF and YM, implemented data collection, and paid survey responders. None of the authors of this study have an affiliation with OpenEvidence." The authors prespecified their analysis plan and analysed blinded. That is a good disclosure of a significant conflict, in a preprint that is not peer reviewed. Every safeguard that could be applied to the analysis was applied, and the vendor still controlled the question source, the collection instrument and the money reaching the raters. Prespecification constrains the analysis. It does not constrain the design.

How do you read a medical AI benchmark?

Translate every benchmark into a sentence of the form "this measures X, which is not the same as Y." Every instrument below is a proxy, and none measures whether a patient did better. Once you can say out loud what a benchmark stands in for, most of the marketing stops working.

Instrument

Measures

Does not measure

The sentence to keep

MedQA (USMLE)

Single best answer recall on licensing exam vignettes with a known key

Open ended reasoning, uncertainty, omission, communication

Whether a model can pass a board exam, which is not the same as whether it can manage your patient.

HealthBench

Agreement with physician written rubric criteria across 5,000 synthetic conversations, graded by a model

Real encounters, longitudinal care, outcomes. The authors state the criteria are not comprehensive.

Whether an answer contains what a panel said it should, which is not the same as whether it helped.

Real clinical queries (Nature Medicine)

Blinded clinician Likert ratings on 100 real de-identified physician queries

Ground truth, because there is none. Generalisability beyond one health system.

Whether clinicians liked the answer, which is not the same as whether it was right.

Real point of care queries (OpenEvidence linked preprint)

Specialty matched physician preference on 620 queries from OpenEvidence's own platform

Absolute correctness. One tool's traffic in one week.

Which answer a specialist prefers side by side, which is not the same as which is correct.

NOHARM

Frequency and severity of potential harm from recommended and from omitted management options

Whether harm reached a patient. Built from 100 eConsult cases.

How badly a tool could hurt someone through what it said and what it left out, which is not the same as how good its answer sounds.

Vendor internal validation

Whatever the vendor defines, against criteria it writes, with comparators it selects and does not name

Anything independently verifiable

Whether a company met its own standard, which is not the same as an evaluation.

Why is a 97 percent MedQA score no longer informative?

Because everything passes. MedQA was released in 2020 by Jin and colleagues at MIT as a natural language processing dataset, never as a clinical competence instrument, and the best method in that paper scored 36.7 percent on the English USMLE subset.[17] In the June 2026 Nature Medicine run, five systems scored between 88.4 and 97.4 percent, so the worst performer would comfortably pass the USMLE.[2] A test everyone passes has stopped sorting anyone. When a vendor says its system scores in the nineties on MedQA, the honest translation is that it is not obviously broken. A peer reviewed systematic review of 39 medical LLM benchmarks (Gong and colleagues, December 2025) puts numbers on the gap: knowledge benchmarks cluster at 84 to 90 percent, practice based assessments at 45 to 69 percent, safety assessment at 40 to 50 percent, and "examination scores are insufficient and misleading proxies for clinical readiness."[17]

What is contamination, and why does every public benchmark expire?

Models are trained on scraped internet text, and the benchmarks are on the internet, so a high score may reflect memorisation. LiveMedBench, a February 2026 preprint of 2,756 real clinical cases, found 84 percent of models performed worse on cases dated after their training cutoff.[17] Here is the trap nobody has solved. A public benchmark is verifiable and contaminated the moment it is published. A private benchmark is contamination resistant and unverifiable, unreproducible and impossible for a vendor to contest. Nature Medicine chose private, and OpenEvidence attacked it for exactly that. The OpenEvidence linked preprint chose public, and by releasing its benchmark guaranteed contamination in the next training run.[2][16] Neither choice is dishonest. There is no third option.

Who built HealthBench, and does it matter?

OpenAI built it, and yes. All twelve authors of the HealthBench preprint (May 2025, never peer reviewed) are OpenAI affiliated. It comprises 5,000 conversations and 48,562 rubric criteria written by 262 physicians whom OpenAI selected from 1,021 interested candidates, and it is scored not by physicians but by a model, with GPT-4.1 as primary grader. The authors disclose that physician to physician and model to physician agreement runs 55 to 75 percent.[17] Each link is defensible and the chain is the problem: an OpenAI benchmark, rubrics from OpenAI selected physicians, graded by an OpenAI model, on which an OpenAI model posts its largest margin of the three Nature Medicine stages, 8.7 points, having come second on MedQA and tied on the contamination resistant benchmark.[2][17] The Nature Medicine authors flagged it and demoted their own result: evaluation of OpenAI models "may be influenced by potential benchmark-developer overlap," so "HealthBench should be interpreted as supplementary," and "industry-created benchmarks may systematically favor the systems developed by their creators."[2] An independent peer reviewed appraisal reaches a balanced verdict in its title: HealthBench is "advancing AI evaluation in healthcare, but not yet clinically ready."[17]

What do all of these benchmarks miss?

What the tool failed to say. NOHARM, an arXiv preprint (version 4, July 13, 2026, not peer reviewed despite being indexed in PubMed) from the ARISE network with Stanford and Harvard investigators, built 1,100 tasks from 100 real primary care to specialist eConsult cases, with 12,747 expert annotations of 4,249 management options by 29 board certified physicians, scoring errors of commission and of omission and weighting each by harm severity.[15]

NOHARM finding (preprint, version 4)

Value

Why it matters

Potential for severe harm from direct application of recommendations

Up to 24.6% of cases

Version 1 in December 2025 reported 22.2%. Cite the version.

Share of severe errors that were omissions

Over 80%

Version 1 reported 76.6% of errors. Do not mix the two.

Correlation of safety with existing knowledge benchmarks

r = 0.61 to 0.64

Roughly 60 percent of the variance in clinical safety is invisible to MedQA style testing.

A "Do Nothing" baseline

Severe harm potential in 37% of cases

Higher than any AI tested. The safety question is always "compared with what?" This figure comes from a vendor summary of the preprint.

Randomized arm, 101 US licensed physicians

AI assistance improved physician performance, but physicians frequently omitted valuable AI generated recommendations

Omission is a human failure mode too.

Sit with that correlation. If clinical safety were captured by knowledge benchmarks, r would approach 1. And if more than 80 percent of severe errors are omissions, every benchmark that grades produced text is blind to most severe risk. A tool can score 97 percent on MedQA, top the preference rankings, and still leave out the one thing that would have hurt your patient. Nothing in the answer grading paradigm detects that, and the Nature Medicine authors concede the point about their own work.[2][15]

Why can every vendor cite a study proving it wins?

Start with the cleanest illustration available, which requires accusing nobody. Within three weeks, three vendors issued press releases claiming victory in the same unreviewed preprint, each citing a different slice of the same dataset.

Vendor

Date

Slice claimed

What the slice is

AMBOSS

July 28, 2026

Highest severity weighted F1 at 86.15%, lowest severe error rate at 2.9%

The benchmark's primary metric. The most defensible of the three on its face.

OpenEvidence

July 20, 2026

Physicians in the free choice arm reached for it in 22.3% of responses, against 19.8% for all other external AI combined

Revealed preference, confounded by familiarity and market share. Most used is a marketing metric wearing a benchmark's clothing.

Doximity

August 6, 2026

"Ranked first among all AI systems evaluated on the study's real-world clinical sample"

A sub-portion, not the primary endpoint. On the primary metric it placed second, 4.8% severe errors against 2.9%. The release does not mention preprint status.

One preprint. Three winners. You do not need to compare studies to see the mechanism. The detail I cannot improve on: inside the release promoting the study, OpenEvidence writes that "the NOHARM benchmarking studies are not peer reviewed and should be taken with a grain of salt."[15] A vendor disclaiming the rigour of the study it is issuing a press release about, in the same paragraph, and more transparent than either competing release, neither of which mentions preprint status. Fortune separately reported that OpenEvidence disputes its own score in the same work.[15]

Does funding predict findings?

Less than you would guess. It predicts amplification. Two counterexamples belong in any honest version of this argument.

First, a vendor funded study published a result against its sponsor's interest. The Digital Health 2025 paper on ChatRWD was funded by Atropos Health, roughly 19 of 27 authors are employees, and the senior author sits on its board. ChatRWD won the headline, 58 percent of questions answered with relevant evidence based answers against 24 percent for OpenEvidence. The same paper reports OpenEvidence producing actionable results for 48 percent of questions where published evidence existed, against 37 percent for ChatRWD, and its disclosure states the authors consulted OpenEvidence while writing.[6]

Second, a fully independent study found for a vendor and was promoted by nobody. Atay and colleagues (Applied Clinical Informatics, June 2026, peer reviewed) were four Turkish obstetrician gynaecologists with no vendor money and no vendor contact, using blinded raters on 24 questions. OpenEvidence scored highest, 54.0 ahead of Gemini at 50.3 and ChatGPT at 48.7, P < 0.001, with poor inter-rater reliability (ICC 0.391).[6] It is the closest thing here to a disinterested evaluation, and nobody issued a press release, because four physicians in Izmir do not have a press office.

So the honest thesis is marketing selection, not research misconduct. Funding predicts which result gets amplified far better than it predicts which result gets produced.

Study

Venue and date

Peer reviewed?

Funder and affiliations

Favours

Sample

Promoted?

Vishwanath et al.

Nature Medicine, June 2026

Yes, open named review

NCI, Keck Foundation, Korean IITP. No vendor funding. Senior author discloses consulting for Google.

Frontier models

1,100 items, 12 clinicians

No

NOHARM (Wu et al.)

arXiv v4, July 2026

No, preprint

Moore and Macy Foundations, NIH, ARPA-H. Academic. Several authors consult for Google DeepMind. No vendor employment declared.

Clinical tools overall; slices favour AMBOSS, Doximity, OpenEvidence

1,100 tasks, 29 physicians, plus 101 physician randomized arm

Yes, three vendors

Feng et al.

arXiv, June 2026

No, preprint

Funding not published. OpenEvidence co-developed collection, ran it, paid raters. Authors independent of the company.

OpenEvidence, by 25 to 39 points

620 queries, 149 physicians

Yes

Freudenberg et al.

J Med Syst, July 2026

Yes

"No funds, grants, or other support were received." Seven German and Swiss breast centres.

ChatGPT-5 Thinking

20 cases, 7 blinded raters

No

Low et al.

Digital Health, June 2025

Yes

"This study was funded by Atropos Health." Most authors are employees; senior author on the board.

ChatRWD, with a sub-result favouring OpenEvidence

50 questions, 9 raters, not blinded

Yes

Hurt et al.

J Prim Care Community Health, April 2025

Yes

Mayo Clinic departmental support, no conflicts declared. OpenEvidence joined Mayo Clinic Platform Accelerate in 2023, an institution level tie the paper does not mention.

OpenEvidence, qualified

5 cases, 4 raters

Yes, in OpenEvidence's rebuttal

Atay et al.

Appl Clin Inform, June 2026

Yes

Funding not published. No vendor connection of any kind.

OpenEvidence

24 questions, 2 blinded specialists

No, nobody

Wolters Kluwer framework

Press release plus gated whitepaper, May 2026

No, marketing document

Wolters Kluwer, criteria written by its own editors

UpToDate Expert AI

1,669 queries, 99.9% "clinically aligned", comparators unnamed

Yes

HealthBench

arXiv, May 2025

No, preprint

OpenAI. All twelve authors OpenAI.

OpenAI models

5,000 conversations, 262 physicians

Yes

Three days before publishing its 99.9 percent claim, Wolters Kluwer published a piece arguing that "user ratings measure perception, not reliability. A response can sound useful and still omit a critical detail," and that "high scores can hide unsafe failure modes."[7] Every one of those criticisms is correct, and every one applies to a rubric based rating against self authored criteria, with unnamed comparators, no reported blinding and no reliability statistic. I state that deadpan and leave it.

What is the single most useful methodological point here?

Each team won on its home question distribution, and neither cheated. The NYU group drew its real world queries from an NYU GPT deployment, a distribution frontier models are optimized for, and frontier models won. The OpenEvidence linked preprint drew its real world queries from OpenEvidence's own platform, the distribution its retrieval architecture is built for, and OpenEvidence won.[2][16] Both datasets are genuinely real world. Neither is neutral. When you see "real clinical queries," the question is not whether they were real. It is whose traffic they came from.

Two mechanisms complete the set. Choose the outcome measure and the winner flips with tools and year held constant: Nature Medicine measured rated answer quality and frontier models won; NOHARM measured severity weighted harm including omissions and clinical tools won.[2][15]Choose the blinding procedure and you can handicap a competitor without meaning to: the breast cancer study stripped "citations, URLs, and other source attributions" before rating, which removes exactly what a citation engine is built to deliver.[3]

What should you ask when a vendor cites a study at you?

  1. Who paid for it, who wrote the benchmark, and who wrote the rubric, and are any of those the winner? Scroll to Funding and Competing Interests. If there is no such section, you are reading marketing.

  2. Is it peer reviewed, and what are both sample sizes, questions and raters? "Stanford-Harvard study" here means an unreviewed preprint with four versions whose headline numbers changed between them. PubMed indexing does not mean peer reviewed.

  3. What exactly was measured, and does that percentage mean what it sounds like? Does the number measure being right, being liked, or being used? All three get reported as wins.

  4. Would this benchmark have caught what the tool failed to say, and did anyone measure whether a patient did better? When a rep says a tool is clinically validated, ask "validated against what outcome, in which patients?"

Two absences are worth as much as any result above. No independent peer reviewed study in which UpToDate Expert AI ranks first was located as of August 2026, its only favourable evaluation being that gated whitepaper with unnamed comparators, and the product is deployed at roughly 2,500 US hospitals.[7] And no study in this literature is powered for hematology and oncology, none located as of August 2026. The largest specialty matched evaluation averages about 21 questions per specialty and says so in its own limitations.[16] The nearest peer reviewed evidence touching Doximity at all is an Anesthesia and Analgesia 2026 study of 216 hip fracture vignettes, which included Doximity GPT and OpenEvidence only in a sensitivity analysis of 36 responses each and concluded that "clinical LLMs provided similar recommendations to general-purpose models."[6] Every number in this article is something you are generalising from another specialty.

What happened in actual oncology, and what does 3.6 percent really mean?

Correct this one wherever you see it, because the circulating reading is wrong. A blinded multicenter study in the Journal of Medical Systems, July 4, 2026 (Freudenberg, Knitza, Gremke and colleagues, peer reviewed, no funding received) ran across seven university breast cancer centers in Germany and Switzerland. Prof. Valmed, a certified class IIb medical device, and OpenEvidence were compared with ChatGPT-5 Thinking on treatment plans for 20 standardized cases, rated blind on safety, guideline adherence, medical adequacy, completeness, overall quality and logical coherence.[3]

System

Share of 140 rater and case combinations ranked FIRST BY PREFERENCE

Processing time

ChatGPT-5 Thinking

96.4%

159 plus or minus 58 seconds

OpenEvidence

3.6%

9 plus or minus 1 seconds

Prof. Valmed

Never ranked first

35 plus or minus 4 seconds

That 3.6 percent is not an accuracy score, and it circulates as though it were. It is the share of head to head preference rankings OpenEvidence won. On the actual quality ratings, OpenEvidence and Prof. Valmed did not differ significantly and both were rated adequately in absolute terms. They lost a three way beauty contest to a system that took 17 times longer and wrote more. If you have seen "OpenEvidence scored 3.6 percent" anywhere, including in an earlier version of my own writing, that is the correction: wrong unit, wrong denominator, wrong conclusion. The limitations cut both ways: synthetic cases, European centres where the operative guidelines are German S3 and AGO rather than NCCN, no case specific gold standard, disclosed length and position bias, and citations stripped before rating.[3] Give the specialized tools their due on the other axis. Nine seconds against 159 is a clinical feature, and an 85 percent answer while the patient is in the room can beat a 95 percent answer after they leave.

How good is ChatGPT at oncology, really?

The most quoted number in this space is real and three model generations out of date. Chen et al., JAMA Oncology, October 2023, a research letter, submitted 104 prompts across breast, prostate and lung cancer to GPT-3.5-turbo-0301, benchmarked against 2021 NCCN guidelines and scored by 3 of 4 board certified oncologists with majority rule, the fourth adjudicating complete disagreement. All 102 outputs containing a recommendation included at least one NCCN concordant treatment, but 35 (34.3 percent) also recommended nonconcordant treatments, and 13 of 104 (12.5 percent) were hallucinated.[5] Anyone quoting 12.5 percent as a current hallucination rate is misleading you. What survives is structural: a model with no licensed guideline corpus reconstructs guidelines from memory, and memory drifts. A 2025 AIME conference paper by Belligoli, Bitterman and Miller found no model retrieved a verbatim guideline quote, with GPT-4 citations reflecting actual content 90 percent of the time for NCCN and 64 percent for ASCO.[5]

What does OpenEvidence actually get wrong?

Not what you think. The citations are usually real, they are sometimes attached to the wrong claim, and re-running the same question does not reliably reproduce the same answer. A July 2026 critique in the Journal of the American College of Clinical Pharmacy by Bergsbaken and colleagues, a clinical pharmacist and four health sciences librarians, documented specific, independently checkable failures and a worse problem underneath them.[4]

Asked to convert metoprolol tartrate 25 mg twice daily to carvedilol, OpenEvidence gave a 1:1 equivalence, "two to four times higher than dose equivalents referenced in the literature." It invoked COMET, which does not treat the drugs as comparable milligram for milligram, and its only citation was the FDA Coreg label, which does carry COMET information but does not list the COMET trial, or the alternative Franz conversion, in its references. That is subtler than a fabricated source and harder to catch.

The second case is where the real finding lives. Asked in May 2025 about misoprostol alone versus mifepristone plus misoprostol for incomplete miscarriage, two of the three cited sources did not include that population at all. Re-run six months later, two of the three references had changed, the answer contradicted itself while citing the same source for both claims, it promoted to "Landmark Study" a trial that had excluded incomplete miscarriage, and it reported 90 percent efficacy for misoprostol when in the cited study that statistic described the efficacy of no medical intervention.

The headline is not that the failures reproduced. It is that the answer did not. A tool whose output you cannot reproduce is a tool whose output you cannot audit, and an unauditable answer is a strange foundation for a treatment decision. The authors' recommendation is the best line in this literature: OpenEvidence "should not be used as the sole resource/tool, especially for users lacking the subject expertise of the topic(s) for which they are using OpenEvidence, as those users may be unable to verify the accuracy of the summary outputs."[4] You can catch its errors in myeloma. Much less so in the cardiology question you asked at 7pm.

Why do these tools cite everything so obsessively?

Because of the FDA, not only because of trust. Under section 520(o)(1)(E) of the FD&C Act, added by the 21st Century Cures Act, certain clinical decision support functions are excluded from the device definition. FDA's guidance, docket FDA-2017-D-6569, issued January 29, 2026, sets four criteria. Criterion 4 requires software "intended to enable HCPs to independently review the basis for the recommendations presented by the software so that they do not rely primarily on such recommendations, but rather on their own judgment, to make clinical decisions for individual patients." At a CDRH town hall on March 11, 2026, FDA's Aneesh Deoras addressed language models directly: "A large language model- or, LLM-enabled software function may meet this criterion if it can sufficiently enable an HCP to independently review the basis for the recommendations."[10]

So the inline citation, UpToDate Expert AI's transparent sourcing, OpenEvidence's links out to NCCN and ASCO, and Doximity's live retrieval trace and per guideline deep links are not only trust features. They are the architecture that keeps a product outside device regulation. Doximity showing you which guidelines it is searching, in real time, before it answers, is the most literal implementation of Criterion 4 in the category.

Here is the uncomfortable part. FDA also states it "does not consider software functions intended for a critical, time-sensitive task or decision to meet Criterion 4, because an HCP is unlikely to have sufficient time to independently review the basis of the recommendations," explicitly weighing automation bias.[10] The exemption assumes you do the review. A fellow at 4pm with six patients left, checking a dose between rooms, is exactly the scenario FDA describes as one where it does not happen. The tool's regulatory status is fine. Your use of it may not be.

Which tool pays for your CME, and which one charges you and gives you none?

This is the most decision relevant pricing fact for a fellow, and it inverts the conventional wisdom. Wolters Kluwer's trainee page, verbatim: "Trainee pricing includes access to the same clinical content and features included in a professional subscription, but does not include CME/CE/CPD credits." Its FAQ: "No, CME/CE/CPD credit is not available with subscriptions purchased at a trainee rate."[7]

Tier (as of August 2026)

Price

CME, CE, CPD

Offline

UpToDate trainee

$219/yr, stated as a starting price on the trainee page

Not included

Paid add on

UpToDate core

$579/yr starting price

Included

Not included

UpToDate Pro Plus

$699/yr starting price

Included, plus state CME tagged topics

MobileComplete included

OpenEvidence

$0 for NPI verified clinicians

CE at no cost; MOC for physicians through select certifying boards

Not publicly documented

Doximity Ask

$0 for verified US clinicians

Free AMA PRA Category 1 Credit, specialty matched, earned inside the answer surface. Accreditor, cap and MOC status not published.

No

Two free tools, one ad supported and one pharma funded, pay for your continuing education. The paid trainee subscription does not. Two notes on UpToDate as of August 2026: trainee subscriptions now include Expert AI in the US and Canada, and Wolters Kluwer's own footnote states that "UpToDate Expert AI requires an internet connection and can't be used offline."[7] The topics go offline. The AI does not.

Does any of this improve patient outcomes?

No. There is no randomized or outcome level evidence that any of these four tools improves patient care. That is the state of the field as of August 2026, not a hedge.

Study

Design

Effect

Caveat

Goh et al., Nature Medicine, 2025

Randomized controlled trial, 92 physicians, NCT06208423

GPT-4 users scored 6.5 percentage points higher (95% CI 2.7 to 10.2, P < 0.001) and took 119 seconds longer per case. No significant difference between AI assisted physicians and the AI alone.

Five vignettes, not patients, and it tested GPT-4, not any commercial tool.[18]

Alyakin et al., Neurosurgery, May 2026 (CNS-Obsidian)

Randomized deployment trial in live workflow

Clinicians used the AI copilot in 70 of 959 eligible consultations, 7.3 percent. Positive ratings 40.6% for the specialized model against 57.9% for GPT-4o, P = 0.230.

Outcomes were helpfulness and differential accuracy, not patient outcomes.[18]

Bonis et al., Int J Med Inform, 2008

Retrospective observational, 424 vs 3,091 hospitals

Length of stay shorter by 0.167 days per discharge (95% CI 0.081 to 0.252). Mortality not significantly different.

Manufacturer authored: the first author's affiliation was UpToDate Inc.[6]

Isaac, Zheng and Jha, J Hosp Med, 2012

Retrospective observational

Length of stay 5.6 vs 5.7 days; lower risk adjusted mortality in 3 of 6 conditions, 0.1% to 0.6% absolute

Independent. Benefit "small or nonexistent" at larger teaching institutions.[6]

OpenEvidence, Doximity Ask, ChatGPT

No comparable outcomes literature exists

Not applicable

Not applicable

The best observational evidence for UpToDate is fourteen to eighteen years old, and at an academic centre the Isaac subgroup finding is the relevant one: benefit was weakest exactly where you work. Note also that Peter Bonis, first author of the 2008 manufacturer authored study, is the Wolters Kluwer executive quoted earlier. Read both with the affiliation in view.

Now the number that reframes everything above it. CNS-Obsidian is the only randomized deployment of a clinical AI copilot into real specialist workflow in this literature, it comes from the same NYU group that published the pro frontier model Nature Medicine paper, and its headline is that clinicians opened the tool in 7.3 percent of eligible encounters. The authors conclude that "low clinical utilization suggests chatbot interfaces may not align with specialist workflows."[18] Every benchmark argued over in this article measures answer quality for a class of product that, in the one real workflow randomized test available, specialists declined to open 93 percent of the time. Before asking which tool is best, ask whether you will open it at the moment you need it, which is the question none of the marketing addresses. And the single peer reviewed study that measured impact on clinical decision making directly, the five case Mayo series, scored it 1.95 out of 4, its lowest domain by a wide margin, because the tool "primarily reinforced rather than modified plans."[6]

What can you safely put into these tools?

Your personal $20 ChatGPT account must never receive PHI. OpenAI's HIPAA eligible products, available with a business associate agreement, include ChatGPT for Healthcare, ChatGPT for Enterprise with Regulated Workspace, ChatGPT FedRAMP, ChatGPT for Clinicians, the API with Modified Retention, and API FedRAMP with Modified Retention. Free, Go, Plus and Pro are not on that list, and on consumer tiers "content is used to train our models" by default with opt out available. OpenAI's usage policies, effective October 29, 2025, prohibit "provision of tailored advice that requires a license, such as legal or medical advice, without appropriate involvement by a licensed professional," which is functionally FDA Criterion 4.[11] OpenEvidence announced in April 2025 that it "now fully complies with" HIPAA, with a BAA for US covered entities.[8] Treat that as institution dependent: older pages still carry legacy avoid PHI disclaimers, and the pharmacy critique warns that free access "bypasses the procurement process that health care entities typically use to verify data security."[4]

Doximity is the strongest of the four on paper and the most interesting to read closely. The BAA is incorporated into the terms of service and takes effect when you register, so every verified member has one without chasing a signature, and PHI is explicitly permitted. Then read the same document to the end: it permits de-identification, after which the data may be commercialised for any lawful purpose, and the privacy policy discloses that topic interest insights from your AI use may reach commercial clients attached to your name, NPI, specialty and zip code, with prompts excluded. Whether your inputs train models is not clearly disclosed either way, and no opt out is published.[13] None of that is unlawful and some of it is standard. All of it is worth knowing before you paste a case in. Ask compliance, not the marketing page.

What do the societies say, and should you use ASCO's Guidelines Assistant instead?

As a fifth option, for one specific job, yes. ASCO's Guidelines Assistant, launched with Google Cloud in May 2025, draws solely from ASCO's published guidelines, cites its sources, and is a member benefit. ASCO's CEO Clifford A. Hudis, MD, FACP, FASCO is precise: "I can't emphasize enough that this is a guideline discovery tool. It is not a clinical decision support tool.... You are still the decision-making physician."[12] ASCO is simultaneously partnered with OpenEvidence, Google Cloud, and Wolters Kluwer for a Guidelines Generator. Hedging across all three is a signal that nobody credible thinks this is settled.

ASCO's "Principles for the Responsible Use of Artificial Intelligence in Oncology" (May 2024, updated May 2025) sets six principles: Transparency, Informed Stakeholders, Fairness, Accountability, Oversight and Privacy, and Human-Centered Application.[12] Two bind here: "Patients and clinicians should be aware when AI is used in clinical decision-making and patient care," and "AI does not eliminate the need for human interaction." A separate May 2025 ASCO statement warns about prior authorization algorithms "influencing automation bias," worth holding next to the fact that both free tools will draft your prior authorization letter. OpenEvidence's own examples include one, and Doximity's suggested prompts included "Draft a letter to an insurance company for PET scan approval" beside "Compose a treatment plan for a newly diagnosed lymphoma patient."[8][13] The same box that plans the treatment writes the appeal. The AMA's policy, adopted at the AMA Interim Meeting in November 2024, holds that "voluntary agreements or voluntary compliance is not sufficient" and that point of care AI use "should be disclosed and documented."[1]

I could not identify any ASH policy statement or position on AI clinical reference tools as of August 2026. There is hematology AI review content in Blood, but no equivalent to ASCO's six principles. For a specialty where one flow cytometry read can redirect a treatment plan, that silence is conspicuous.

What is the corpus gap that still matters for heme/onc?

OpenEvidence's oncology licensing has no equal here: NCCN and JNCCN (November 2025), ASCO Guidelines with figures and flowcharts (May 2026), the Society of Surgical Oncology (May 2026), the Society for Neuro-Oncology (June 2026), OneOncology (August 2026), plus NEJM and JAMA Network.[8] On August 5, 2026 Springer Nature was added, bringing Leukemia, Blood Cancer Journal, Bone Marrow Transplantation and British Journal of Cancer into scope.[19] Blood itself, JCO, Lancet Oncology and Annals of Oncology remain without a publicly announced agreement, and the ASCO deal covers Guidelines, not JCO. Treat that as no announced agreement rather than confirmed absence, since no corpus list is published. The pharmacy critique makes the general point: OpenEvidence "may lack access to relevant literature, including paywalled content from publishers without agreements."[4]

And now the sentence that teaches the most. That agreement is with Springer Nature, the publisher of Nature Medicine, announced on August 4, 2026, roughly two months after Nature Medicine published the study OpenEvidence asked it to retract.[19] Report it exactly as it is. There is no evidence it influenced the June editorial decision, which it postdates by two months. It is also a real commercial relationship between a journal's publisher and a company whose product that journal published a critical study of, and any Matters Arising would be adjudicated under it. A fellow learning to read conflicts of interest should learn to look above the author line, to the publisher.

Which tool for which question?

Question type

Use first

Why

Verify with

"What does the guideline say for this stage and biomarker?"

OpenEvidence, or Doximity Ask

Licensed NCCN and ASCO content in one; per guideline, per version NCCN deep links in the other

The actual document. Always.

"Which NCCN category is this recommendation?"

Doximity Ask

The only one of the four observed to carry NCCN category into the answer and cite at page level within the guideline

The NCCN PDF, at the version cited

An ASCO specific guideline question

OpenEvidence, or ASCO's Guidelines Assistant

OpenEvidence holds the ASCO licence including figures and flowcharts. Doximity produced no asco.org links.

The guideline itself on asco.org

Drug dosing, renal or hepatic adjustment, drug conversion

UpToDate plus your institutional pharmacy resource

Graded recommendations, named authors, daily drug surveillance

A pharmacist. This is the category with a published two to four fold error.

A messy multi-comorbidity case that fits no pathway

A frontier general model

Won every rated category in blinded breast cancer planning and held the top tier on all three Nature Medicine benchmarks

Your attending, tumour board, and the guideline

A trial presented in the last few weeks

The primary literature and the meeting itself

No tool publishes a time to incorporation commitment

The abstract, the discussant, the full publication

A patient facing explanation

A frontier general model, then edit it yourself

Best at plain language rewriting; clarity was the clinical tools' weakness

Yourself. No PHI in the prompt.

Anything touching PHI

Only a product covered by your institution's BAA. Doximity's executes on registration.

Consumer ChatGPT tiers are not HIPAA eligible and train on content by default

Your compliance office, before the first prompt

Offline: no signal, shielded suite, rural infusion centre

UpToDate MobileComplete

The only genuine offline capability among the four, and it does not include the AI

Nothing else works without a connection

Free CME while you look things up

Doximity Ask, or OpenEvidence

AMA PRA Category 1 Credit fused into the answer surface; CE and select board MOC respectively

Your state and board requirements, which neither vendor tracks for you

Board exam preparation

None of these four

Not built for durable recall or item level explanation

A dedicated question bank

That last row is where my bias belongs. None of these four builds durable recall against the ABIM blueprint or tells you why a distractor is wrong. If you are studying rather than practising, start with a comparison of heme/onc question banks or my wider review of which AI tools are worth a fellow's time.

Frequently asked questions

Is Doximity Ask better than OpenEvidence for hematology and oncology in 2026?

Neither has an independent peer reviewed head to head against the other, so this is a fit question, not a performance question. Doximity Ask is free to verified US clinicians, deep links to the correct NCCN guideline page at the matching version, cites at page level inside the guideline, passes the NCCN category into its answer tables, and includes free AMA PRA Category 1 CME. OpenEvidence holds the ASCO Guidelines licence including figures and flowcharts, plus NCCN, JAMA Network, NEJM Group and, since August 2026, Springer Nature. As of August 2026 there is no peer reviewed evaluation of the current Doximity Ask product, and no study in this literature is powered for hematology and oncology.

Is Doximity Ask actually free, and what is the catch?

Yes. Doximity states verbatim that Ask is free for all users with a verified Doximity account, with no consumer paid tier and only an enterprise licence at an unpublished price. It requires a verified US clinician account. The catch is the business model. Doximity reported $644.9 million in revenue for fiscal 2026, its 10-K describes its revenue generating customers as primarily pharmaceutical manufacturers and health systems, and its privacy policy discloses that topic interest insights derived from AI tool use may be shared with commercial clients linked to your name, NPI, specialty and zip code, with your prompts excluded. The pharma share of revenue is not broken out, and I saw no advertising inside any Ask answer across six queries in August 2026.

Does the UpToDate trainee subscription include CME?

No. Wolters Kluwer states that trainee pricing includes the same clinical content as a professional subscription but does not include CME, CE or CPD credits. As of August 2026 the trainee rate is $219 per year and the core subscription that does include CME starts at $579 per year. Both free tools do better here: OpenEvidence offers CE at no cost to NPI verified physicians, nurse practitioners and physician associates, with MOC for physicians through select certifying boards, and Doximity offers free AMA PRA Category 1 Credit earned by reading clinical answers.

Do AI reference tools hallucinate citations in oncology?

The more dangerous failure is different. A July 2026 critique in the Journal of the American College of Clinical Pharmacy documented OpenEvidence answers where citations were real but did not support the claim, including a metoprolol to carvedilol conversion two to four fold too high while invoking the COMET trial. The deeper problem was reproducibility: the same prompt six months later returned different references and a self contradictory answer. The June 2026 Nature Medicine study found no significant difference in harmful content or hallucination rates across the systems it tested. Real citations attached to the wrong claim survive a casual check in a way fabricated references do not.

What does the 3.6 percent figure for OpenEvidence in the breast cancer study actually mean?

It is not an accuracy score, and it is widely misread as one. In the July 2026 Journal of Medical Systems study, 3.6 percent is the share of 140 rater and case combinations in which OpenEvidence was ranked first by preference, against 96.4 percent for ChatGPT-5 Thinking. On the actual quality ratings OpenEvidence did not differ significantly from Prof. Valmed and both were rated adequately in absolute terms. The study used 20 standardized cases at German and Swiss centres and stripped citations before rating to preserve blinding.

Is there any evidence that these tools improve patient outcomes?

No. As of August 2026 there is no randomized or outcome level evidence that OpenEvidence, Doximity Ask, UpToDate Expert AI or a frontier general model improves patient care. The strongest randomized evidence is Goh and colleagues in Nature Medicine in 2025, which enrolled 92 physicians, used five clinical vignettes rather than patients, and tested GPT-4 rather than any commercial clinical tool. The only randomized deployment into real specialist workflow, the CNS-Obsidian trial published in Neurosurgery in May 2026, found clinicians used the AI copilot in 7.3 percent of eligible encounters, 70 of 959.

References

  1. American Medical Association. 2026 Physician Survey on Augmented Intelligence (n = 1,692; use case items n = 1,342): survey report. Augmented Intelligence Development, Deployment, and Use in Health Care, adopted at the AMA Interim Meeting, November 2024: AMA AI principles.

  2. Vishwanath K, Alyakin A, Ghosh M, et al. "General-purpose large language models outperform specialized clinical AI tools on medical benchmarks." Nat Med. 2026;32(7):2405. Peer reviewed: nature.com. Retraction demand and responses, Becker's Hospital Review, June 16, 2026: beckershospitalreview.com; further context: ASCO AI in Oncology and Becker's, "The healthcare AI PR wars, explained".

  3. Freudenberg J, Knitza J, Gremke N, et al. "AI-enabled clinical decision support in breast cancer care." J Med Syst. 2026;50(1):107. Peer reviewed: doi:10.1007/s10916-026-02434-w; full text PMC13332881.

  4. Bergsbaken JM, Wilson P, Vandagriff S, et al. "Prescribing Caution: A Critique of OpenEvidence to Answer Medication-Related Questions." J Am Coll Clin Pharm. 2026 Jul;9(7):e70237. Peer reviewed: doi:10.1002/jac5.70237.

  5. Chen S, Kann BH, Foote MB, et al. JAMA Oncol. 2023;9(10):1459, research letter, GPT-3.5-turbo-0301 against 2021 NCCN guidelines. Peer reviewed: doi:10.1001/jamaoncol.2023.2954. Belligoli P, Bitterman D, Miller T, AIME 2025 conference paper, LNCS 15735:30: doi:10.1007/978-3-031-95841-0_6.

  6. Specialty evaluations and outcomes literature, all peer reviewed. Low YS, et al. (Atropos funded), Digital Health 2025: doi:10.1177/20552076251348850. Atay AO, et al., Appl Clin Inform 2026;17(3):518: doi:10.1055/a-2899-0123. Hurt RT, et al. (Mayo Clinic, n = 5), J Prim Care Community Health 2025: doi:10.1177/21501319251332215. Chen R, et al., Anesth Analg 2026: doi:10.1213/ANE.0000000000008148. UpToDate outcomes: Bonis PA, et al., Int J Med Inform 2008;77(11):745: doi:10.1016/j.ijmedinf.2008.04.002; Isaac T, Zheng J, Jha A, J Hosp Med 2012;7(2):85: doi:10.1002/jhm.944.

  7. Wolters Kluwer, retrieved August 2026: Pro Plus pricing and Expert AI; trainee pricing and CME FAQ; GRADE grading guide; editorial policy; editorial process; validation framework press release, May 21, 2026; "Clinical AI evaluation must go beyond benchmark wins," May 18, 2026. The supporting whitepaper is gated, and neither document is peer reviewed.

  8. OpenEvidence announcements and blog: NCCN (November 5, 2025); ASCO Guidelines (May 27, 2026); Society of Surgical Oncology (May 11, 2026); Society for Neuro-Oncology; OneOncology (August 6, 2026); CE and MOC (July 28, 2026); EvidenceGrade (July 9, 2026); HIPAA compliance (April 25, 2025).

  9. Aguilar M. "OpenEvidence raises $250 million, doubling its valuation." STAT News, January 21, 2026, STAT+ paywalled with the quoted passages above the paywall: statnews.com. Prior rounds: $210M at $3.5B, July 2025; $200M at $6B, October 20, 2025.

  10. US Food and Drug Administration. Clinical Decision Support Software guidance (docket FDA-2017-D-6569) and the guidance document, issued January 29, 2026; CDRH town hall transcript, March 11, 2026: fda.gov/media/191560.

  11. OpenAI, retrieved August 2026: ChatGPT pricing; HIPAA eligible products and functionality; usage policies, effective October 29, 2025; OpenAI for Healthcare, January 8, 2026.

  12. American Society of Clinical Oncology. Six guiding principles for AI in oncology, May 2024, updated May 2025; Hudis CA, The ASCO Post, December 10, 2025.

  13. Doximity product and policy sources, with in-product observations from a live authenticated session on August 16, 2026: Doximity Ask; Ask FAQs, source of the free access statement and the clinical use disclaimer; Business Associate Agreement; privacy policy; CME library; engineering description of Ask, June 23, 2026; Clinical AI Suite, May 7, 2026.

  14. Doximity corporate filings and releases: SEC EDGAR, CIK 0001516513; Form 10-K; FY2026 results, May 13, 2026; Pathway Medical acquisition, closed July 29, 2025, corroborated by CNBC; investor deck, February 2026.

  15. NOHARM and its promotion. Wu D, et al. "First, do NOHARM." arXiv:2512.01241, version 1 December 1, 2025, version 4 July 13, 2026. Preprint, not peer reviewed, though indexed in PubMed: arxiv.org. Vendor releases: AMBOSS, July 28, 2026; OpenEvidence, July 20, 2026; Doximity, August 6, 2026. Dispute coverage: Fortune, July 29, 2026.

  16. Feng J, Patel V, Heagerty P, Mai Y, et al. "Expert Evaluation of Clinical AI Tools on Real Point-of-Care Clinical Queries." arXiv:2606.28960, June 27, 2026. Preprint, not peer reviewed: arxiv.org, with the disclosure statement in the full PDF. Coverage: Becker's Hospital Review, July 13, 2026.

  17. Benchmark methodology. Jin D, et al. MedQA, arXiv:2009.13081, preprint: arxiv.org; peer reviewed version, Appl Sci 2021;11:6421: doi:10.3390/app11146421. Arora RK, et al. HealthBench, arXiv:2505.08775, preprint: arxiv.org and full PDF. Liu J, Liu S, Digit Health 2025, peer reviewed critique: doi:10.1177/20552076251390447. Gong EJ, Bang CS, Lee JJ, Baik GH, J Med Internet Res 2025;27:e84120, peer reviewed systematic review of 39 benchmarks: doi:10.2196/84120. Yan Z, et al. LiveMedBench, arXiv:2602.10367, preprint: arxiv.org.

  18. Randomized evidence, both peer reviewed. Goh E, Gallo RJ, Strong E, et al. "GPT-4 assistance for improvement of physician performance on patient care tasks: a randomized controlled trial." Nat Med 2025;31(4):1233: doi:10.1038/s41591-024-03456-y. Alyakin A, et al. "CNS-Obsidian: A Neurosurgical Vision-Language Model Built From Scientific Publications." Neurosurgery, May 19, 2026, randomized deployment trial: doi:10.1227/neu.0000000000004070.

  19. Springer Nature Group. "Springer Nature and OpenEvidence announce agreement," August 4, 2026: group.springernature.com; OpenEvidence's announcement of the same agreement, August 5, 2026: openevidence.com.

Frequently Asked Questions

This article is written for medical students, residents, fellows, and clinical educators looking for evidence-aligned guidance in oncology learning and board preparation.

No. This article is an educational resource and does not replace clinical judgment, institutional protocols, or specialty guideline updates.

Use it as a framework: review the key concepts, test yourself with practice questions, and pair your study with current guideline documents and physician-led teaching.

About the Author
Dr. Roupen Odabashian, MD

Dr. Roupen Odabashian, MD

Hematology-Oncology Fellow, Karmanos Cancer Institute

Hematology-oncology fellow at Karmanos Cancer Institute / Wayne State University; founder of MeDucation AI; clinical and research focus on thoracic oncology and AI in cancer care.

View full author profile
Ready to elevate your medical learning?

Join MeducationAI, the AI-powered medical education platform built for students across specialties, with personalized tutoring, smart study tools, and realistic clinical case simulations.

Get Started

Share this article