AI in Medicine — What Doctors Actually Use It For Today
TL;DRAI's biggest win in medicine isn't diagnosis — it's the note. Ambient scribes give doctors back hours a week, and that's the fastest-spreading use by far. Imaging AI is the most mature and most regulated: it triages and flags, a human still signs. Early-warning models watch patients between rounds. Paperwork AI handles coding and inbox replies. And the whole field is being reshaped less by any single model than by capacity — the same doctors seeing more patients with less clerical drag. Every specialty has its own version: auto-contouring gives radiation oncologists back an hour a day, endoscopy gets polyp detection, endocrinology already runs genuinely autonomous closed-loop insulin. And what is actually coming — individualised mRNA cancer vaccines, pharmacogenomics, digital twins — is closer in some areas and much further in others than the headlines suggest. The failure modes are specific and knowable: models that degrade at a new hospital, automation bias, gaps in the training data, and confident text that isn't true.
…
Will AI replace doctors?
Not on any credible near-term path. Almost every deployed clinical AI is regulated as an assistive device — it flags, ranks, drafts or measures, and a clinician signs. What is genuinely changing is the mix of a doctor's day: less typing and clerical work, more decisions per hour. The realistic risk isn't replacement, it's a doctor who stops checking the machine's output.
What is the single most useful AI tool for a practising doctor today?
An ambient documentation assistant — it listens to the visit and drafts the note. It's the one category with broad real-world adoption because the benefit is immediate and the failure mode is visible: you read the draft before you sign it. Nothing else currently gives back that much time per week.
Is AI diagnosis accurate enough to trust?
In narrow, well-defined tasks with good data — flagging a large-vessel stroke on a CT, grading diabetic retinopathy, measuring an ejection fraction — it performs at a clinically useful level and is regulated for exactly that scope. The failure is when a model trained at one hospital is used at another with different scanners, populations or workflows, where accuracy can silently drop. Performance is a property of the model plus the setting, never the model alone.
What should a clinic check before buying a clinical AI tool?
Five things: what regulatory clearance it holds and for exactly which indication; whether it was validated on patients who resemble yours; what happens to accuracy when your scanner, lab or population changes; who is liable when it's wrong; and where the patient data goes. If a vendor can't answer the second and third clearly, that's the answer.
Which medical specialties use AI the most?
Radiology by a wide margin — most authorised AI devices are imaging. Then pathology, cardiology (especially ECG and echo), ophthalmology and radiation oncology, where auto-contouring saves real hours daily. But the specialty with the biggest time return is primary care, through ambient documentation rather than any diagnostic model. Endocrinology is furthest along the autonomy curve: closed-loop insulin delivery genuinely doses without asking each time.
What are individualised cancer vaccines?
You sequence a patient's tumour, identify the mutations producing neoantigens unique to that specific cancer, and manufacture an mRNA vaccine training their immune system against those targets — not a vaccine for melanoma, a vaccine for their melanoma. It is in randomised trials, most visibly in melanoma alongside immunotherapy. AI is load-bearing because choosing which of hundreds of candidate mutations will actually provoke a response is a prediction problem no human does by inspection.
Is AI going to make personalised medicine real?
Partly, and unevenly. Pharmacogenomics is the closest — knowing before you prescribe that a patient metabolises a drug poorly is already actionable and guideline-supported for a set of drugs, and under-used mainly for workflow reasons. Polygenic risk scores are further out, and were derived disproportionately from European-ancestry cohorts, which limits validity elsewhere. Digital twins remain research with real pilots in cardiac and orthopaedic modelling.
Where does AI in medicine still fail?
Distribution shift, where a model that worked in validation degrades at a new site. Bias, where training data under-represents some patients. Automation bias, where clinicians stop scrutinising a confident output. And with language models specifically, fluent text that is simply wrong — which is why they're used for drafting and summarising under review, not for unsupervised clinical claims.
Share
Read the headlines and you'd think medicine had already been handed over to the machines. Sit in a clinic on a Tuesday and you'll see something much less cinematic and much more useful: a doctor finishing a note in forty seconds instead of six minutes, a radiologist's worklist reordered so the bleed is at the top, an alert that fires at 3am because a patient's numbers drifted in a pattern nobody was in the room to notice.
That gap between the headline and the exam room is the whole story. AI is genuinely changing medicine — just not mainly where the marketing points. Here's what it's actually doing, what you can put to work this week, and the specific ways it goes wrong.
The five jobs AI is really doing
1. It writes the note
This is the big one, and it's not close.
Ambient documentation tools listen to the consultation and draft the clinical note — history, exam, assessment, plan — in the format you already use. You read it, correct it, sign it. That's the whole loop.
It's the fastest-spreading clinical AI in the world for three unglamorous reasons. The benefit is immediate and personal: documentation is the single largest non-clinical time sink in a doctor's day, and a chunk of it happens after hours, which is a direct line to burnout. The failure mode is visible: a wrong draft is a wrong draft on your screen, before it becomes anything else. And it doesn't require you to trust a machine's judgement — only its transcription and structuring.
If you adopt exactly one thing, adopt this one.
2. It reads the image
Imaging is where clinical AI is most mature, most regulated, and most misunderstood.
The large majority of AI-enabled medical devices authorised by regulators are in radiology, and the pattern is consistent: they triage, flag, measure and prioritise. A model scans incoming studies and pushes the suspected large-vessel occlusion to the top of the list, cutting the time from scan to intervention in stroke. It flags a possible pulmonary embolism. It measures ejection fraction, or lung nodule volume across two scans, more consistently than a human eyeballing it twice.
Two things follow from "triage and flag":
A human still signs. Almost all of these are assistive by design and by regulation. The clinical value is often in ordering the queue rather than making the call — getting the urgent study seen sooner.
Consistency is the underrated win. Humans are inconsistent on measurement tasks — between two radiologists, and between the same radiologist on a Monday and a Friday. A model measures the same way every time. For anything you track over time, that reproducibility matters more than a headline accuracy number.
Beyond radiology, the same shape appears in digital pathology, dermatology, and ophthalmology — where autonomous screening for diabetic retinopathy became one of the first AI systems cleared to produce a result without a specialist reading it, precisely because the task is narrow, the images are standardised, and the alternative is patients not being screened at all.
3. It watches the patient between rounds
Nobody is at the bedside continuously. Models are.
Early-warning systems read the monitoring and lab stream and flag deterioration — sepsis, decompensation, a patient sliding toward an ICU transfer — sometimes hours before it becomes obvious. In principle it's the highest-leverage use in the hospital: the alert arrives during the window where acting is still cheap.
In practice this category also produced clinical AI's most instructive public failure. A widely deployed proprietary sepsis model, when independently validated at hospitals outside its development setting, performed substantially worse than its marketing implied — missing many cases while generating a large volume of alerts. The lesson wasn't "prediction doesn't work." It was that a model's accuracy is a property of the model plus the setting, and a vendor's internal numbers tell you very little about your hospital.
4. It does the paperwork
Unglamorous, enormous.
Clinical coding, prior authorisation packets, referral letters, discharge summaries, and — increasingly — first drafts of replies to the patient portal inbox, which has quietly become one of the heaviest uncompensated loads in outpatient medicine. A model drafts, a clinician edits and sends.
There's a subtle benefit beyond time here. Patients frequently report that AI-drafted replies read as more thorough and more empathetic than the rushed two-line answer a doctor types between patients — not because the machine cares, but because it isn't in a hurry and doesn't get compassion-fatigued at 6pm. The draft gives the doctor something to soften and personalise rather than compose from nothing.
5. It compresses the search
Two very different scales.
At the bench: protein structure prediction went from a multi-year experimental problem to something you can often get computationally in an afternoon, which changes what's worth attempting in target identification. Models screen enormous candidate libraries, propose molecules, and predict properties before anything is synthesised. This doesn't eliminate the decade-long slog of trials — the binding constraint on a new drug is clinical, not computational — but it reshapes the front end.
At the bedside: retrieval over the literature and the patient's own record. "What did this patient's cardiology note say about the aortic root in 2023" is a question that used to cost ten minutes of scrolling. Trial matching is the same shape — surfacing the three studies a specific patient might be eligible for out of thousands.
What a doctor can actually do this week
Concretely, in rough order of effort:
Put an ambient scribe in your consultations. Biggest time return, lowest trust requirement. Read every draft before signing — that's not a temporary precaution, it's the design.
Let it draft your inbox replies. Same rule: it drafts, you send.
Use it to summarise a long record before a complex visit. A five-year chart into a paragraph you verify against the source for anything that will change your management.
Use it as a second read, deliberately, after you've formed your own view. The order matters — asking first invites anchoring. Forming your impression, then asking, turns it into a check rather than a suggestion.
Ask it to argue against you. "What else could this be, and what would distinguish them?" Differential-broadening is one of the genuinely good uses of a language model in clinical thinking, precisely because it isn't attached to your first impression.
Don't paste identifiable patient data into a consumer chatbot. Use whatever your institution has approved and covered by an agreement. This is the boring rule that ends careers when ignored.
Find your specialty
The five jobs above are the shape. What they look like depends entirely on what you do. Find yours.
A note on how to read this: cleared means regulators have authorised a product for that use and it's in hospitals today. In trials means real patients, not yet routine. Research means promising and not yet yours to use.
Radiology
The deepest deployment of any specialty, and the most cleared products by a wide margin. Detection and triage on stroke, intracranial haemorrhage, pulmonary embolism, pneumothorax and lung nodules; automatic quantification of nodule growth between scans; worklist reordering so the time-critical study surfaces first. Report drafting from findings is arriving fast. The quiet growth area is opportunistic screening — a CT ordered for abdominal pain also contains a coronary calcium score and a vertebral fracture assessment, and a model reads what the human wasn't looking for.
Pathology
Whole-slide imaging made the specialty computational, and the wins are where human attention fails: finding micrometastases in lymph nodes (a needle-in-haystack task where models genuinely outperform tired eyes), counting mitoses, and grading — Gleason scoring is more reproducible from a model than between two pathologists. The frontier is predicting molecular status from morphology alone — inferring things like microsatellite instability from a routine H&E slide, which could triage who needs expensive molecular testing at all.
Cardiology
ECG is the standout: models detect atrial fibrillation, and — more surprisingly — infer structural disease like reduced ejection fraction from a plain 12-lead, spotting things the human eye was never able to read off that trace. Echo automation measures EF and strain consistently instead of operator-dependently. Consumer wearables now push AF notifications to patients, which changed who walks into your clinic and why. CT coronary plaque quantification is cleared and growing.
Emergency medicine
Triage and acuity scoring, imaging prioritisation so the intracranial bleed is read first, and detection models that shave minutes off door-to-needle in stroke. ECG interpretation at the front door, sepsis flags, and department-level crowding and admission forecasting for staffing. The high-value clinical bit is time-to-decision — in this specialty the queue is the outcome.
Critical care
Deterioration and sepsis prediction, ventilator weaning support, sedation titration, and organ-failure prediction. Also a genuine target for reducing alarms: models that suppress the false alarms driving alert fatigue are as valuable as models that generate new alerts. Treat vendor performance claims here with the most scepticism of any specialty — this is where external validation has most often disappointed.
Anaesthesiology
Intraoperative hypotension prediction minutes before it happens, depth-of-anaesthesia monitoring, and closed-loop infusion control under supervision. Preop risk stratification and difficult-airway prediction from patient data and images. Theatre scheduling and turnover optimisation is an unglamorous but real win for a department judged on throughput.
Surgery
Preoperative 3D planning and segmentation from imaging; intraoperative guidance overlaying anatomy; and surgical video analysis — automatically recognising the phases of an operation, flagging critical views, and giving objective feedback on technique. Video-based skill assessment is quietly one of the most consequential uses in the field, because surgical education has never had objective measurement at scale. Robotic platforms remain surgeon-controlled; autonomy exists only in narrow research settings.
Gastroenterology and endoscopy
Computer-aided detection of polyps during colonoscopy is among the most-studied clinical AI anywhere, with consistent increases in adenoma detection rate, plus characterisation tools that predict histology in real time. It's also the specialty that produced the clearest deskilling signal: evidence that unassisted detection can decline in endoscopists who become accustomed to the assist. That finding is a warning to every specialty, not just this one.
Neurology
Large-vessel-occlusion detection with automated alerting to the stroke team is the flagship — it compresses the interval that determines disability. Beyond stroke: EEG interpretation and seizure detection, MS lesion tracking across serial MRI, and objective quantification of movement disorders from video or wearables, replacing a subjective scale with a measurement.
Oncology
Trial matching against a patient's molecular profile, treatment-response assessment on serial imaging, and molecular tumour board support that surfaces the evidence for a variant. Toxicity and complication prediction is developing. The eventual promise is treatment selection informed by a genuinely multimodal picture — imaging, pathology, genomics and clinical course read together rather than in separate silos.
Radiation oncology
Possibly the specialty where AI has changed the day-to-day workflow most concretely. Auto-segmentation of organs at risk turns hours of manual contouring into minutes of review — that's a real hour of a clinician's day, every day. Plan optimisation and adaptive replanning, where the plan adjusts to anatomy changing over a course of treatment, is moving from research into practice.
Ophthalmology
Home to one of the first autonomous diagnostic systems ever cleared: diabetic retinopathy screening that produces a result without a specialist reading it, in primary-care settings where no ophthalmologist exists. Beyond that: OCT segmentation and fluid quantification for injection decisions, glaucoma progression detection, and retinopathy-of-prematurity screening.
Dermatology
Lesion triage and teledermatology workflows, with a caveat this specialty must state loudly: training sets have historically under-represented darker skin tones, so performance is not uniform across your patients. Treat any dermatology model's headline accuracy as unproven for the patients least represented in its training data until shown otherwise.
Primary care and family medicine
The specialty with the largest time return rather than the largest diagnostic one. Ambient documentation, inbox reply drafting, and pre-visit chart summarisation. Then population-level work: risk stratification to find who needs to be called in, screening-gap detection, and referral triage. A generalist's constraint is breadth and time, and this is aimed squarely at both.
Psychiatry
Documentation relief matters more here than most places, because sessions are long and notes are heavy. Digital therapeutics deliver structured CBT between appointments. Risk stratification from record data exists but demands unusual care — a suicidality flag is a clinical and ethical event, not a dashboard metric. Speech and language biomarkers for depression and psychosis remain research, and should be described as such to patients.
Obstetrics and gynaecology
Fetal biometry and anomaly detection on ultrasound, reducing operator dependence in a famously operator-dependent modality. Cardiotocography interpretation support, where inter-observer disagreement is well documented. In fertility, embryo selection from time-lapse imaging is in real clinical use. Preeclampsia and preterm-birth risk models are developing.
Paediatrics and neonatology
Growth and development tracking, neonatal sepsis prediction, retinopathy-of-prematurity screening, and weight-based dosing safety checks. The core caveat: children are not small adults, and a model trained on adults may fail on them in ways that aren't obvious. Ask what age range a tool was validated on — every time.
Pulmonology
Nodule detection and volumetric tracking, pulmonary function test interpretation, COPD exacerbation prediction, and sleep study scoring — the last being a high-volume, highly repetitive task that automation suits almost perfectly.
Endocrinology
The quiet success story of autonomous clinical AI: closed-loop insulin delivery. A continuous glucose monitor and pump running an algorithm that doses without asking permission each time is genuinely autonomous, genuinely regulated, and genuinely routine. Nothing else in medicine is as far along the autonomy curve — worth remembering when people claim autonomy is impossible.
Nephrology
Acute kidney injury prediction from labs and medication exposure, dialysis parameter optimisation, and progression modelling in chronic disease. AKI is a good fit for prediction because the signal is in data you already collect continuously.
Infectious disease and microbiology
Antimicrobial resistance prediction, stewardship support suggesting narrower therapy sooner, automated plate and gram-stain reading, and outbreak detection from surveillance data. This is one of the clearer public-good uses: better antibiotic selection is a benefit that outlives the patient in front of you.
Orthopaedics
Fracture detection — particularly the subtle ones missed on a busy shift, where a second read genuinely changes outcomes. Implant sizing and preoperative templating, bone age assessment, and post-operative recovery tracking from wearables.
Urology
Prostate MRI lesion detection and PI-RADS support, pathology grading integration, and stone characterisation on CT.
Genetics and rare disease
Variant interpretation and prioritisation — reducing a list of thousands of variants to a handful worth a human's attention — plus phenotype matching, including facial phenotyping for syndrome recognition. For rare disease this attacks the diagnostic odyssey directly: the problem is rarely that the answer is unknowable, it's that no single clinician has seen enough cases.
Pharmacy and clinical pharmacology
Interaction and contraindication checking that accounts for the actual patient rather than a generic warning, dose optimisation from pharmacokinetics, and adverse-event signal detection across populations. Reducing alert fatigue is as much the goal as adding checks.
What's actually transforming the field
Step back from individual tools and the systemic change is mostly about capacity.
Health systems everywhere face the same arithmetic: more patients, older patients, and not enough clinicians — a shortfall no country is training its way out of quickly. You can respond by working people harder, which is what produced the current burnout numbers, or by removing work that never needed a medical degree. Almost all of AI's real value so far sits in that second category. The note, the letter, the code, the queue, the summary.
Three second-order effects follow:
Access changes shape. Autonomous screening where there's no specialist to do it — retinal screening in a primary-care clinic, triage in a rural hospital with no on-site radiologist — is a different proposition to AI in a well-staffed academic centre. The comparator isn't "AI versus expert." It's "AI versus nobody."
Some specialties change more than others. The prediction that imaging specialists would be automated away got the direction wrong. Demand for imaging keeps rising, and the work is shifting toward the parts a model can't do: integrating the image with the patient, handling ambiguity, procedures, and being accountable. What changes is the mix of the job, not its existence.
The evidence bar is rising. Early clinical AI was sold on retrospective accuracy. That's no longer enough, and regulators have moved: the mature question is now whether the tool improves an outcome prospectively, and how it's monitored after deployment — including a defined plan for what happens when the vendor updates the model underneath you.
What's coming — and how far away it actually is
Every item below is real work by serious people. They are not equally close, so each one says how far.
Individualised cancer vaccines — in trials, genuinely
The most striking near-term idea in medicine. Sequence a patient's tumour, identify the mutations that produce neoantigens unique to that cancer, and manufacture an mRNA vaccine that trains their immune system against those specific targets. Not a vaccine for melanoma — a vaccine for your melanoma.
This has moved from concept into randomised trials, most visibly in melanoma alongside immunotherapy, with work extending to pancreatic and other cancers. AI is load-bearing in it: choosing which of hundreds of candidate mutations will actually be presented and provoke a response is a prediction problem no human can do by inspection.
What to watch: whether the survival benefit holds at scale, and whether manufacturing a bespoke product per patient in weeks can survive contact with real health systems and their budgets.
Precision medicine that finally uses the genome routinely — partly here
Pharmacogenomics is the closest piece: knowing before you prescribe that a patient metabolises a drug poorly is already actionable, already guideline-supported for a set of drugs, and still under-used mostly for workflow reasons rather than scientific ones.
Polygenic risk scores are further out clinically. They're real predictors at population level, but they've been derived disproportionately from European-ancestry cohorts, which limits validity elsewhere — a fairness problem that is also a straightforward accuracy problem.
Medical foundation models — arriving, unevenly
Instead of one model per narrow task, large models trained across imaging, notes, labs and waveforms that can be adapted to many tasks. The appeal is obvious: most clinical questions are multimodal, and today's tools mostly aren't.
The honest position is that generality and safety pull against each other. A model cleared for one indication has a bounded claim you can evaluate. A general model that does many things has a much harder story to tell a regulator — and the regulatory frameworks for exactly this, including how a model may keep changing after approval, are being written right now.
Digital twins — research, with real pilots
A computational model of an individual patient — a specific heart, say — that you can simulate against. Test an ablation strategy or a device setting in software before doing it in a person. Cardiac and orthopaedic modelling are furthest along. Cross-organ, whole-patient twins remain a research programme, not a product.
AI-designed drugs — in the clinic, verdict pending
Molecules discovered or designed with heavy computational involvement have entered human trials. The front end of drug discovery genuinely compressed: target identification, candidate generation, property prediction, and — since protein structure prediction became routine — structural insight that used to take years.
The sober part: the attrition that kills drug programmes is mostly clinical, in efficacy and safety in real humans. Compressing discovery from years to months does not compress a phase III trial. Expect faster starts, not instant medicines. The first genuinely AI-originated approved drug will be a milestone worth marking — and it will still have taken years to get there.
De novo protein design — research, moving fast
Designing proteins that don't exist in nature: binders, enzymes, potential therapeutics. This is one of the areas where progress has outpaced most predictions. Therapeutic use is still early.
Ambient clinical intelligence beyond the note — early
Documentation tools currently listen and transcribe. The next step is a system that also notices — the guideline that applies to this patient, the medication that conflicts with the new prescription, the follow-up nobody booked — surfacing it during the visit rather than in an audit six months later. Technically close; the hard part is doing it without becoming another alert nobody reads.
Hospital-at-home and continuous monitoring — deploying now
Passive sensing plus prediction to run acute-level care in a patient's home. This is less about a clever model and more about capacity: it's one of the few credible ways to add hospital beds without building hospitals.
Supervised surgical autonomy — research
Narrow, well-defined steps of an operation performed autonomously under supervision, demonstrated in animal and bench work. Meaningful clinical autonomy is far off, and the regulatory and liability questions are harder than the technical ones.
Federated learning — infrastructure, quietly important
Training models across many hospitals without moving patient data between them. Unglamorous, but it's the plumbing that would let a model see enough diversity to work for everyone — which is the fix for the bias problem that keeps appearing throughout this article.
The pattern worth noticing
Look down that list and the near-term items share a shape: they compress work humans already do, or they extend capacity that doesn't exist. The far-off items are the ones asking a machine to take clinical responsibility.
That ordering isn't an accident, and it's a decent guide for reading any claim you meet in the next few years.
Where it breaks
This is the part worth reading twice, because these failures are quiet.
Distribution shift. A model trained at one hospital meets different scanners, different labs, a different population, a different threshold for ordering the test — and degrades. It doesn't announce this. It returns confident outputs that are slightly more wrong than they used to be. Any tool that isn't monitored after go-live is being trusted on faith.
Automation bias. Once a system is usually right, humans stop checking it properly. This is well documented across aviation and medicine, and it's the mechanism by which a tool that improves average performance can still cause a specific harm — the missed finding the model didn't flag and nobody looked for. It's worst precisely when the tool is good, because that's when scrutiny relaxes.
Deskilling. A related, slower version: if the machine always finds the lesion, the human's own detection practice erodes. This is an active concern in procedural specialties, and it argues for deliberately preserving unassisted practice rather than assuming the skill persists.
Bias in the data. A model learns the population it was trained on. Where that under-represents a group — different skin tones in dermatology, different body habitus, women in cardiac datasets, non-English speakers in note-derived models — performance drops for exactly the patients already worst served. Historical inequity in the data becomes automated inequity in the output.
Fluent wrongness. Language models produce confident, well-formed text regardless of whether it's true. In medicine that's a specific hazard: a plausible drug interaction that doesn't exist, a citation to a paper that was never written, a summary that quietly drops the one abnormal value. This is why the safe uses are drafting, summarising and retrieval under review — and why "it sounded right" is not verification.
Liability and consent. When an assisted decision goes wrong, responsibility currently lands on the clinician who signed. That's a real asymmetry: you carry the risk of a system you didn't build and can't inspect. Know what your indemnity actually covers.
How to evaluate a tool before it touches a patient
Five questions. If a vendor is vague on the middle three, that is your answer.
What is it cleared for, exactly? Not "AI for radiology" — the specific indication, population and role. A triage clearance is not a diagnostic one.
Who was it validated on? Numbers from patients who resemble yours, ideally from sites that aren't the developer's.
What happens when our setting differs? Different scanner, different lab assay, different case mix. Ask what monitoring exists to detect degradation after go-live, and who watches it.
What does it do to the workflow? A tool that adds clicks dies regardless of accuracy. A tool that fires too many alerts trains people to dismiss it — and alert fatigue is itself a patient-safety problem.
Where does the data go, and who's liable? Storage, training rights, and the answer to "if this is wrong and a patient is harmed, what happens."
The honest summary
AI in medicine is neither the revolution of the press release nor the hype of the sceptics. It is a set of genuinely useful tools that mostly do clerical and perceptual work extremely well, are regulated as assistive for good reason, and fail in specific, knowable ways.
The doctors getting the most out of it right now aren't the ones with the most advanced model. They're the ones who moved documentation off their evenings, use the machine as a second opinion rather than a first, and stayed in the habit of checking it.
That last habit is the whole ballgame. The tool that's usually right is the one you'll stop reading.