Artificial intelligence
Doing it for themselves: GenAI is the front door to a personalized health system
August 28, 2026
Patients are doing it for themselves, and generative AI is becoming their front door to the health system. The question is how the formal system can help them get good information, or at least not get in the way.
Patient use of generative AI is very high and still climbing. In January, OpenAI reported that 230 million people ask ChatGPT health questions every week, 40 million of them daily, and that seven in ten of those conversations happen outside clinic hours. The same report counted nearly 600,000 messages a week from US “hospital deserts,” locations more than a 30-minute drive from a hospital. By last week the count was 300 million. Let me briefly review what we now know.
Microsoft’s health team has published a four-part series on Copilot use. An April paper in Nature Health classified 617,827 health conversations into a clinician-validated taxonomy. By day, the machine serves the clerk and the researcher; by night, the worried patient. One user in seven is asking not for themselves but for a child or an aging parent. The system closes at five. Worry does not.
July’s installment (also in Nature Health) extends the analysis to 1.7 million de-identified conversations across 109 countries. Two key findings. Where people distrust hospitals, they ask the machine more. Where the state has built a structured system, they ask it differently: universal health coverage is the strongest predictor in the analysis, and what it predicts is paperwork queries. Volume measures confidence lost, as people turn to AI when they can’t get answers; content measures structure built, as users shift from clinical questions to navigational and bureaucratic ones.
OpenAI and Anthropic have published parallel work; all three labs now run usage epidemiology on their own platforms. All of it is vendor-produced, but credit those who publish, especially in peer-reviewed journals. Most of the best data is American. We need good Canadian data, now.
The closest we have is a CMA survey: 48 percent of Canadians have used AI for health information; only 27 percent trust it. Given the decline in trust in US public health, this should be of immediate concern. We can no longer rely on a search of US assets like the CDC and academic centres to give us trustworthy, Canadian-relevant answers.
And this is no longer just usage data. On July 23, OpenAI launched ChatGPT Health for every American adult: it connects Apple Health and pulls medical records directly from Epic and Oracle Health patient portals. The mini personal health record early adopters were building by hand is now a product feature. It is US-only; no Canadian launch has been announced.
There will be edge cases: A July report from the UN’s Independent International Scientific Panel on AI put patient safety back on the table. It treats chatbots as emerging health infrastructure and lists the possible harms, with almost no usage data behind them.
That is frustrating: the harms cited are mostly qualitative or anecdotal, which is no basis for quality control and measurement. But the underlying point stands. The UN is right to point at the edge cases and to ask how we will monitor the quality of the machine.
As a reality check, any machine used hundreds of millions of times a week will have edge cases. Cars are widely used and routinely injurious to users and bystanders alike. As an adult, I once almost killed myself at a parking garage gate. It still gives me nightmares, and I always put my vehicle in park now. You should too.
With a technology less than four years old, the number of errors may be meaningful. How high, we are still trying to get our arms around, and it is not an easy question. The leading foundation models already beat human performance on exam questions such as the US medical licensing exam. But exams were always a bad way to measure performance, and even an AI that scores 95 percent may not be good enough at these volumes. People will approach a question from an unexpected path, as I did with the parking gate.
Even at 99.99 percent “performance” (however defined) OpenAI’s post-launch numbers imply roughly 30,000 “bad” answers a week on that platform alone. That is a lot of bad answers for some system to be liable for. The first lawsuit, over a chatbot’s suggestion not to consult a doctor, was filed the day before ChatGPT Health launched.
The really uncomfortable part: we do not know how to measure performance, or safety, or bias, for consumer health AI. And nobody measures routinely. The closest thing is OpenAI’s HealthBench; credit the vendor for publishing, but a vendor-run bench is not independent quality control.
Someone in Canada should run every new model against our own bank of a thousand synthetic consumer health queries – limited to trusted sources, provincially aware, bilingual, validated against clinical guidelines – and publish rated answers, at least to a Consumer Reports level.
Should we just stop? No. And we could not if we tried. These models are filling a real need; that is why the usage is widespread. Can we make the models better? A surprisingly difficult question. Most people think better means safer, but whenever I unpack that word in discussions, safer usually means not answering certain questions.
Guardrails of that kind degrade performance: we make a model safe by making it less honest and less good. That works for bioterrorism and nuclear secrets. But do I really want a model that is less good on purpose because it is “safer” for my healthcare needs?
Another personal and very practical example. As a 60-plus male with minor cardiac issues, I regularly use Claude and ChatGPT to discuss my blood pressure, medications, diet and exercise. That also affects my damaged foot, back and nerve issues, all of which affects my yoga and pickleball.
GenAI advice is incredibly useful to me. Many colleagues and friends have similar stories, dropping their information in to build that mini personal health record by hand. Last time I was in my NP’s office I grabbed a screenshot of my OLIS record and rebuilt my lab trends for the last 17 years.
For the record, no, I am not adjusting a dosage without talking to my MD and NP first.
But is this OK at a society level? You will not get anything close to 99.99 percent safety on cardiac medication conversations with 60-plus Canadian men on the current GenAI platforms. I would be surprised if the error rate were not closer to 5 or 10 percent, and this is non-trivial advice.
There is no easy way for a health system to take responsibility and liability at that error rate. So, if asked, we would likely stop the machine from answering my questions. That is the wrong answer. It is a nanny state answer.
My answer is caveat emptor: consumers need to exercise discretion. People need to act like adults and not do silly things with new GenAI tools, and we need to find ways to help them.
Can it be made safer for me? Yes. The simple idea is harnesses that raise the quality of what the AI draws on rather than shrink what it will discuss: source control and citations.
Eventually these harnesses will become elaborate and may become medical devices. ChatGPT Health is the first commercial harness at scale, though its terms still say it is not intended for diagnosis or treatment.
For the moment, much of this can be do-it-yourself with better prompts. The figure shows the skeleton of a proto-RAG prompt: which types of website to trust, in what order, with citations required and 811 as the escalation. Paste it into any major AI and the machine changes character. Add the institutions you want checked, or ask your AI to suggest local sites relevant to your condition. Here is the scaffolding.
Who should publish the lists, and fuller prompts like this? My doctor. Disease societies such as Heart & Stroke and Diabetes Canada. Caregiver groups like the Ontario Caregivers Organization. Universities and academic health centres. Governments, provincial first, because health information is local: services, coverage and referral paths differ by province.
Every public institution has the counterpart obligation: make your information RAG-ready. Make it structured, citable and machine-retrievable. Chatbots are already a front door, running an order of magnitude beyond 811. If patients are doing it for themselves, the least the system can do is publish the sources and the prompts that make the machine safer.
This space will move fast in the next 12 to 18 months. Patient-side scribes are already emerging. The language translation opportunities are a huge potential equity gain for our multicultural society.
As an individual, consider keeping your own set of health and wellness prompts. DIY and RAG-ready will supplement, not replace, testing and reporting on both foundation models and harnesses. The Americans just got their records wired into the machine. Canada should answer our way: published sources, published prompts, and an independent bench keeping score.
Want one built for you? Drop this article into Claude Code or OpenAI Codex, or even just plain Claude or ChatGPT, and ask the AI to build a version of the prompt for you and your conditions. Try it (caveat emptor).
My GenAI is the front door to My Health System.
Will Falk is a policy fellow at the CSA Public Policy Group, the C.D. Howe Institute, Rotman, and WiHV.