GPT-6 Astra arrived September 3, two days after Claude Fable 5.1. Gemini 3.8 Flash landed in between. Here’s how they compare on health tasks, plus a catch-up on the last few weeks.
The new generation, on the same scoreboard
HealthBench Professional evaluates clinical consults, medical research, and writing tasks drawn from real clinician conversations. That makes it useful for understanding medical reasoning, though it doesn’t directly test how well a patient can manage their own care. For the background, see my earlier HealthBench explainer.
Source: OpenAI’s launch comparison, Science and Health table and footnote 11. These are OpenAI-run results, including its evaluations of competitors. Fable 5.1’s score includes Opus 5 fallback on provider refusals.
Where Astra improves
Source: GPT-6 Astra System Card, §6.1, Table 6.
The bigger gains are on Professional and Hard. These are rubric-score gains, not proportional improvements in diagnostic accuracy.
The length adjustment matters, too. Longer answers get more chances to hit the rubric’s criteria. The adjustment tries to reduce that advantage. A model shouldn’t win just because it gives you a pamphlet when you asked a question. Method
Fable’s practical change may be less visible in the scores: Anthropic says its revised safeguards intervene 85% less often on benign biology and medical questions than the original Fable 5 safeguards. Research requests can still fall back to Opus. The model you select and the model answering your health question may differ. Anthropic
For patients, I’m interested in whether these models can turn a pile of records into something useful before an appointment. A practical task to try:
Using only these records, build a dated timeline. Cite the file and page for each event. Flag contradictions and missing information. Then list five questions for my next appointment. Separate what the records say from your interpretation.
Check the timeline against the documents. A benchmark can help choose a model; the receipts tell you whether it did this job correctly.
A few other things worth catching up on
Better transcription for the parts of the visit you forgot
Google’s Gemini 3.5 Transcribe recorded-audio API supports speaker attribution and word-level timestamps. That could help turn recorded visits into notes for doctors or patients, with each speaker labeled. It’s a developer preview; medication names, doses, and instructions still need checking against the audio.
Gemini can help find the appointment, too
Google announced Zocdoc among its new connected apps on August 12. Its current support table lists appointment search and booking for U.S. adults using personal accounts, in English, through Gemini chat on web and mobile. This is the kind of unglamorous capability I care about: getting from a question to an appointment. Confirm the provider, coverage, and slot before booking.
ChatGPT connects to Epic
OpenAI’s September 1 announcement adds Epic integration and structured public healthcare sources. The Epic connection requires an eligible organization; it isn’t available to individual accounts. Potentially useful for the people treating you, but this release doesn’t give patients a new self-service route into their charts.



