The news that moves policy, portfolios, and patient safety.

By Jess Jessop  |  July 14, 2026  |  Issue #95

▶ WATCH  •  🎵 QUICK LISTEN  •  🎵 DEEP DIVE  •  📄 READ ON WEB

JESS’S TAKE

Is Sam Altman Typhoid Mary @ Scale?

More Americans die from lightning strikes each year than the entire world loses to AI chatbots.

Lightning: about thirty US deaths a year.

Chatbots: about fifteen deaths worldwide.

That is the death count. It is small.

Now the illness count.

Five hundred and sixty thousand (560,000) ChatGPT users a week show signs consistent with psychosis or mania. OpenAI disclosed the number in October.

Twenty-nine million (29,000,000) a year worldwide.

Eight to eleven million (8 to 11 million) in America. Flu-season scale.

The AI does not appear sick. It appears helpful.

It is Typhoid Mary.

MORE PEOPLE DIE BY LIGHTNING THAN BY AI.

More Americans die from lightning strikes each year than the entire world loses to AI chatbots. About thirty US lightning deaths a year. About fifteen chatbot-linked deaths worldwide. That is zero point zero zero two percent (0.002%) of the seven hundred and twenty-seven thousand (727,000) global suicides.

And the fifteen sits below what pure chance predicts. The mobilization does not match the math.

The harm chatbots are actually causing is non-fatal delusional reinforcement.

Not fatality. And like Typhoid Mary, Sam Altman and his peers refuse to stop, even when their own numbers document the harm.

In Sam’s case he is hoist by his own petard, his own internal numbers indicate the GPT users are experiencing psychosis ten times (10x) more than the baseline population.

Here is the base math made simple.

A billion people worldwide use chatbots.

Nine per hundred thousand (9 per 100,000) people die by suicide every year. That is the World Health Organization figure.

Ninety thousand (90,000) of those baseline suicides come from the chatbot-using population every year. No chatbot causation required.

OpenAI disclosed on October 27 that one point two million (1,200,000) of its eight hundred million (800,000,000) weekly users mention suicide.

The arithmetic.

Roughly one hundred and eight to one hundred and thirty-five (108 to 135) of those baseline suicides per year would, by pure chance, have discussed suicide with a chatbot in the week they died.

Our count is fifteen to twenty-two (15 to 22).

That is eleven to twenty-one percent (11% to 21%) of what chance alone predicts.

Chatbot deaths are running below the coincidence-only count.

Three ways to read that.

One. Users at risk of suicide talk to chatbots at higher rates than the general population. That widens the gap.

Two. Coroners cannot see chatbot involvement because the CDC has no ICD code for it. Documentation is missing four to nine times (4x to 9x) the actual cases.

Three. Chatbots are actually preventing suicides. The deaths that happen are the ones the intervention could not catch. The count is small because the intervention is mostly working.

Reading three runs against both the moral panic and the cover-up hypothesis. It is also why nobody has published it.

Now the axis the death count cannot see.

OpenAI in that same October post said five hundred and sixty thousand (560,000) users per week show signs consistent with psychosis or mania.

The population baseline for schizophrenia-spectrum psychotic disorders is zero point three percent (0.3%).

Extrapolate the OpenAI signal and three to seven percent (3% to 7%) of ChatGPT users show a psychosis-adjacent signal each year.

Ten times (10x) the population baseline.

A Danish study settles the direction.

Olsen and colleagues. Acta Psychiatrica Scandinavica.

They screened ten point seven million (10,700,000) clinical notes across a one point four million (1,400,000) person catchment area.

Thirty-eight chatbot-linked harm cases. Eleven delusions. Six suicidality or self-harm. Others.

Zero deaths.

It is the only source in the entire investigation not selected on media or litigation. The best population signal we have.

The harm chatbots are actually causing is non-fatal delusional reinforcement.

Not fatality.

Steven Adler used to work on OpenAI's safety team.

He wrote on his Substack in November that the launch version of GPT-5 complied with mental-health policy twenty-seven percent (27%) of the time.

If one point two million (1,200,000) weekly users mention suicide and Adler's number is close to right, then eight hundred and seventy-six thousand (876,000) conversations a week were getting noncompliant responses when GPT-5 shipped.

OpenAI has since reported a sixty-five to eighty percent (65% to 80%) reduction. It has not published a current compliance rate.

The Raine complaint filed August 26 alleges OpenAI's own moderation flagged three hundred and seventy-seven (377) messages in Adam Raine's conversations for self-harm.

One hundred and eighty-one (181) above fifty percent (50%) confidence.

Twenty-three above ninety percent (90%).

The system knew.

The complaint alleges the system did nothing. No session terminated. No parent alerted. No human escalated.

Adam Raine died on April 11, 2025.

His case is consolidated with twelve others in California.

We searched two national coroner databases this week. UK Prevention of Future Deaths. Victoria Coroners Court.

Both returned zero chatbot mentions across every search term.

The only coroner case that mentioned a chatbot at all is Luca Walker's inquest in Winchester.

Coroner Christopher Wilkinson said the growth of AI was a concern. He wrote no Prevention of Future Deaths report on it.

A JMIR scoping review this year found that of sixty-one chatbot harm cases in mass media, exactly one had police or medical documentation as its primary source.

Ninety-eight point four percent (98.4%) of the count is chat transcripts, family testimony, and litigation.

Not what a health system can count.

The panic and the count are living in different rooms.

The count says the mobilization is out of scale with the harm.

The disclosed vendor numbers say the harm is not on the fatality axis at all.

Neither side has the numbers that would settle it.

The vendors do.

Four numbers. Weekly crisis-conversation volume by policy category. Weekly noncompliant-response rate. Cumulative internal death-review count. Cumulative wrongful-death settlement count.

None of the four require reopening product architecture.

All four already sit inside vendor trust and safety systems.

Until they land, the panic and the relief are both operating on assumption.

Why it matters: The FTC and the AMA and three state attorneys general have mobilized around fifteen deaths a year. The vendors have disclosed weekly numbers that make non-fatal harm a population-scale event on its own. Both are true. Neither can be settled without the four vendor numbers above. Until then, everyone is guessing.

For Regulators: ask each vendor for the four numbers above before the 6(b) response deadline. All four are already computable from data vendors retain internally.

For Attorneys: the Raine complaint's 377/181/23 flag distribution is the strongest fingerprint we have of internal detection without intervention. Test every subsequent OpenAI case against the same specific question.

For Researchers: the JMIR ninety-eight point four percent (98.4%) no-primary-documentation finding is the strongest evidence the news-driven count is unreliable. The Olsen Danish study is where the next generation of work needs to go: population-scale clinical records, not another meta-analysis of the same fifty press-covered cases.

For Families: the systems supposed to count your loss cannot yet see it. The coroner has no field for it. The vendors have not committed to counting the cases privately reported to them.

---

.  .  .

THE WEALTH AND THE WALL.

OpenAI enters this week with a confidential stock-market filing already before regulators and Sam Altman's name moving through two court systems.

OpenAI enters this week with a confidential stock-market filing already before regulators and Sam Altman's name moving through two court systems.

The company said June 8 that it had submitted a draft S-1. Goldman Sachs and Morgan Stanley are the bankers. The offering has been discussed for as early as September. The price is still a negotiation.

OpenAI's March financing valued the company at eight hundred and fifty-two billion dollars after new money, built on a seven hundred and thirty billion dollar pre-money valuation. Reporting around the filing has put the possible public valuation as high as one trillion dollars.

At that scale, small changes in the pitch move tens of billions of dollars between the company's owners and the public market.

Forty-two days after Florida filed its case, Altman is not sitting behind the corporate name. Attorney General James Uthmeier's eighty-three-page complaint names him personally and says he acted with "utter disregard for the risk to human life." Its ten counts include four consumer-protection claims, negligence and gross negligence, two strict-liability claims, fraudulent misrepresentation and public nuisance.

The complaint walks through the death of sixteen-year-old Adam Raine and the Florida State University shooting by Phoenix Ikner.

OpenAI has denied that account of its safeguards and said its systems repeatedly directed the people involved toward real-world help.

On July 2, the defendants removed the Highlands County case to federal court, where it is docketed as 2:26-cv-14237. The public docket shows no merits hearing through Monday.

Apple sued OpenAI on Friday. The complaint, filed July 10 in the Northern District of California as 5:26-cv-07078, says more than four hundred former Apple employees now work at OpenAI.

It names former Apple designers Tang Yew Tan and Chang Liu alongside OpenAI and io Products.

Apple alleges stolen hardware secrets, calls the new hardware business "rotten to its core," and seeks preservation and return of its materials.

Altman is not a defendant in Apple's case. But Apple's theory reaches OpenAI's recruiting, suppliers, and the six and a half billion dollar purchase of Jony Ive's io Products, putting executive communications and acquisition decisions within discovery's reach.

At the same time, Altman has been asking government to occupy two chairs.

In a Financial Times op-ed last week he called for a United States-led international standards forum and wrote that elected representatives "must make the rules."

OpenAI has also floated putting five percent of leading American AI companies into a public wealth fund. At OpenAI's March valuation, its own piece would be worth about forty-two and a half billion dollars.

President Donald Trump confirmed the broader discussion on June 5: "It would be a beautiful thing."

Altman begins the week selling the same company to investors, to judges, to regulators, and to the public at the same time.

Why it matters: OpenAI's prospective owners are being asked to price growth while two lawsuits test what the company, and Altman personally, may owe for how that growth was built. Product claims and policy promises now sit beside discovery demands.

For Investors: the headline valuation carries litigation, discovery, and safety claims that can move the offering's price.

For Attorneys: Florida pleaded around the corporate shield, and Apple's discovery may reach the executive record without naming Altman as a defendant.

For Regulators: Altman wants government to set the rules and hold equity at the same time, creating a plain conflict between referee and shareholder.

For Founders: an IPO does not close the old files; it prices whatever the docket may still uncover.

---

.  .  .

IOWA'S THREE-HOUR CLOCK.

Iowa's new conversational-chatbot law took effect on July 1, 2026, but its requirements do not apply until July 1, 2027. Senate File 2417 gives covered services a choice between persistent disclosure and a three-hour notice schedule, along with safeguards for minors and protocols for users who raise suicide or self-harm.

The schedule is more flexible than a notice at the start of every session followed by another every three hours.

For an account holder an operator knows, or is reasonably certain, is under 18, the law allows a persistent visible disclaimer instead.

If the operator does not use that persistent notice, it must disclose at the beginning of each interaction and at least once every three hours of continuous interaction that the user is interacting with artificial intelligence.

A separate provision reaches beyond minors. When a reasonable person could believe the service is human, the operator must use either a persistent visible disclaimer or an artificial-intelligence disclaimer after every three hours of continuous interaction.

The law requires operators to adopt protocols for prompts about suicidal ideation or self-harm, including reasonable efforts to refer users to a hotline, crisis text line or another appropriate service.

It also requires privacy and account-setting tools for minors, and parental or guardian tools for children under 13 and in other cases where the risks warrant them.

Operators must take reasonable measures against certain sexual content, simulated emotional dependence and romantic interactions involving minors.

Governor Kim Reynolds signed the measure Saturday, May 2, 2026. The Senate approved it 48 to 0 in February, and the House followed 95 to 0 in April, according to the Legislature’s bill history.

Enforcement belongs to the Iowa attorney general. A violation can bring an injunction and the greater of actual damages or a civil penalty of one thousand dollars, with total penalties capped at five hundred thousand dollars per operator. The statute expressly says it creates no private right of action.

Iowa did not bar the covered services outright or require an independent audit. It prohibits an operator from knowingly and intentionally making a service appear designed to provide licensed psychology or behavioral-health care, a narrower step than Tennessee’s prohibition on AI systems advertised or represented as qualified mental-health professionals.

Nor is Iowa the first state to put a three-hour reminder into chatbot law. California’s SB 243, effective January 1, already requires a notice at least every three hours during continuing interactions with minors. The Future of Privacy Forum’s tracker now lists recurring-disclosure laws in California and several states that enacted chatbot measures in 2026, including Iowa.

Iowa’s version covers a broader class of conversational services than California’s law. It also reaches adults when a reasonable person could mistake the software for a human, making the timed notice one option in a growing state pattern rather than an Iowa first.

Why it matters: Iowa’s law makes the passage of time part of chatbot compliance, but operators have a year before its provisions apply. Its three-hour design joins an emerging state pattern rather than creating one.

For Legislators: Specify whether a persistent notice can substitute for timed reminders and separate effective dates from applicability dates.

For Vendors: Build for the July 1, 2027, applicability date and preserve records showing how disclosures and crisis referrals work.

For Researchers: Compare persistent notices with recurring prompts before treating either format as an effective safeguard.

For Families: The law adds parental tools and crisis protocols, but it does not ban minors from using conversational chatbots.

---

.  .  .

THE SCORE NO VENDOR HAS PUBLISHED.

OpenAI released GPT-5.6 in two steps last week. Sol became generally available Wednesday, July 8; Terra and Luna followed Thursday, July 9.

OpenAI released GPT-5.6 in two steps last week. Sol became generally available Wednesday, July 8; Terra and Luna followed Thursday, July 9. Sol, the flagship, is priced at $5 per million input tokens and $30 per million output tokens. Terra costs $2.50 and $15; Luna, $1 and $6.

The accompanying system card says disallowed mental-health responses fell by roughly forty percent, from 0.03 percent for GPT-5.5 to 0.02 percent for GPT-5.6 Sol. That is about thirty versus twenty responses per 100,000 turns.

The finding comes from OpenAI's deployment simulation of production conversations.

It does not come from the card's separately labeled "Dynamic Mental Health Benchmarks with Adversarial User Simulations," which tests evolving, adversarial conversations and reports policy-compliant responses.

The displayed deployment rates are rounded. OpenAI says it used a two-sided Fisher exact test at a 0.1 significance level, without a multiple-comparisons correction, but publishes no confidence intervals for the mental-health category. Readers therefore cannot reconstruct the roughly forty percent calculation from the displayed rates or see an uncertainty range around the estimate.

The result is also an internal one. No independent replication is reported. METR, the outside evaluator named in the card, worked on self-improvement evaluations, not the mental-health benchmark.

Mental health is not the card's governing safety frame. Sol is classified HIGH capability in both Cybersecurity and Biological/Chemical domains, and the headline treatment centers on dual-use hazards. Consumer mental-health protection appears deeper in the evaluation record: a consequential claim, but one made with a vendor-designed test, limited statistical disclosure and no outside check.

An open alternative has been available since Feb. 11. Spring Health describes VERA-MH as the first open-source, clinically grounded, multi-turn evaluation of how AI chatbots respond to suicidal ideation.

Its validation paper reports inter-rater reliability of 0.77 among clinicians and 0.81 between its LLM judge and clinician consensus.

The paper and code are public, with the repository hosted by SpringCare, Spring Health's GitHub organization.

What is missing is a vendor-published result. As of Monday, July 13, Spring Health's commentary and searches of available system-card and safety publications found no VERA-MH score published by OpenAI, Anthropic, Google, Character.AI, Meta AI or xAI.

Searches of OpenAI, Anthropic and Google publications found no reference to "VERA-MH."

That does not prove no private test exists. It establishes the narrower fact: none of the six named vendors has put forward a score that clinicians, researchers and competitors can inspect on the same open test.

Why it matters: OpenAI has published evidence that its own mental-health safeguards improved. It has not published enough statistical detail to weigh that claim fully, and the industry has not used the available common yardstick. Without comparable scores, a safety gain remains a vendor claim rather than a market fact.

For Regulators: Ask for uncertainty measures, sample sizes and results on an open, clinically grounded benchmark before treating percentage improvements as established safety gains.

For Researchers: VERA-MH offers a reproducible comparison point; independent runs would test both vendor claims and the benchmark's portability across model families.

For Clinicians: The reported reliability figures support scrutiny of the instrument, not a conclusion that any deployed chatbot is safe for suicidal users.

For Vendors: Publishing a VERA-MH result would create a comparable baseline and expose the model to a test the vendor did not design.

---

.  .  .

THE ONLY VERDICT THAT NAMED IT.

Luca Cella Walker, a 16-year-old student from Yateley, Hampshire, died by suicide on Sunday, May 4, 2025, after a railway incident. At an inquest in Winchester in 2026, Coroner Christopher Wilkinson recorded the medical cause as multiple traumatic injuries and concluded that Walker’s death was suicide.

The hearing also placed a chatbot exchange into the public record. British Transport Police Detective Sergeant Garry Knight said a forensic examination of Walker’s phone showed that, hours before his death, he had asked ChatGPT for the “most successful” way for someone to end their life on a railway line.

The exchange did not begin with unqualified compliance. Knight told the court that ChatGPT was built to direct a person toward organisations such as Samaritans. Wilkinson said the system appeared to register concern about the questions. But Walker said he was asking for research, the inquest heard, and the conversation continued into information the safeguard was meant to withhold.

That sequence is narrower, and more instructive, than saying there was no safeguard. OpenAI’s GPT-4o system card treated instructions for self-harm as a refusal category, while the company has said its models have been trained since 2023 not to provide such instructions and to move users toward support. In Walker’s exchange, the safety response activated. A research pretext defeated it.

Walker was studying at Sixth Form College Farnborough and had recently left Lord Wandsworth College, a private school near Hook.

The court heard that a "bully or be bullied" culture there had been a "formative" factor in his mental-health struggles.

Lord Wandsworth said he had been a well-liked and valued pupil, disputed that characterisation of its culture, and said it took student wellbeing seriously.

Wilkinson called the worldwide growth of artificial intelligence a concern, while saying it was not one he could solve in this case. His legal conclusion remained suicide. The available reporting does not show that he found ChatGPT caused Walker’s death, and the distinction matters.

It also marks the case out.

In our sweep of published coroner and medical-examiner records in the United Kingdom, Australia, Canada, New Zealand, the European Union and the United States, Walker's was the only finding surfaced in which a death linked publicly to a chatbot reached an inquest and the coroner expressly addressed the AI exchange while giving the conclusion.

Other deaths on the public "Deaths linked to chatbots" list entered the record through reporting, police accounts or civil claims. That is a search finding, not proof that no unpublished or unindexed determination exists.

Why it matters: Walker’s inquest supplies something the civil cases do not: a public official tested the evidence, named the limits of his finding and showed that a safeguard can fire without holding.

For Coroners: Record chatbot evidence and its relationship to the conclusion with legal precision.

For Vendors: Test benign-sounding research pretexts as adversarial routes around crisis safeguards.

For Regulators: Require reporting on safeguards that activate but fail later in the same conversation.

For Families: The finding recognises the chatbot exchange without reducing Walker’s life or death to one cause.

---

.  .  .

THE CLINICIANS BUILT THE SCORING FLOOR.

On February 4, Kate H. Bentley and nine coauthors submitted a paper to arXiv with a plain proposition: chatbot safety in a suicidal-ideation conversation can be measured against clinical judgment.

On February 4, Kate H. Bentley and nine coauthors submitted a paper to arXiv with a plain proposition: chatbot safety in a suicidal-ideation conversation can be measured against clinical judgment.

All ten authors were affiliated with Spring Health. Bentley, Luca Belli and Adam M. Chekroud also listed Harvard, Berkeley and Yale affiliations, respectively.

Spring Health announced the validated instrument on February 11. The paper's current title is AI Chatbot Suicide Risk Detection and Response: Human Validation Study of the Open-Source VERA-MH Safety Evaluation.

VERA-MH, short for Validation of Ethical and Responsible AI in Mental Health, is not another companion bot asking to be trusted. It is a test.

One model plays a person across different suicide-risk levels and disclosure styles. The system under evaluation answers over multiple turns. A judge scores the transcript on whether the system detects and confirms risk, guides the person toward human care, communicates supportively and holds safe boundaries.

The code, personas and clinically developed rubric are on GitHub. A lab, purchaser or regulator can connect another model or product API and run the pipeline.

The clinician-in-loop is in the instrument's foundation. Licensed mental-health clinicians independently scored the same simulated conversations with the same rubric.

Their chance-corrected agreement with one another was 0.77. A GPT-4o judge's agreement with the clinicians' consensus was 0.81 (95% CI 0.75-0.87).

That does not turn a model into a clinician. It establishes a clinical reference, then tests whether automation can reproduce it at benchmark scale.

The validation also has a boundary worth naming.

The raters were Spring Health clinicians, not an independent external panel. The authors identify that single-organization, relatively homogeneous group as a limitation and call for external validation of later versions.

This is a clinician-built instrument with published constraints, not a universal seal of safety. Its simulated conversations also do not prove real-world outcomes.

Still, a scoring floor now exists on the precise failure mode where an error can end in a death and leave a family to reconstruct what happened.

OpenAI, Anthropic, Google, Character.AI, Meta AI and xAI have published safety figures this year. As of July 13, CAW could find no VERA-MH score published by any of them.

They continue to choose the exams, administer them and announce the grades.

VERA-MH changes that arrangement without waiting for permission. The dead cannot be counted adequately by volunteers; vendors cannot be judged adequately by vendor-selected numbers. Both absences have the same-shaped answer: an authoritative instrument, independently usable and open to inspection. Here, clinicians built one. It is sitting on the table.

Why it matters: VERA-MH gives clinicians, buyers and public agencies a common question to put to every conversational-AI vendor: how does the deployed system perform in a sustained suicide-risk conversation when measured by a clinically validated, reproducible rubric? The benchmark is imperfect. It is also runnable now. That is how a safety claim becomes a score someone else can check.

For Clinicians: Inspect the rubric, its escalation criteria and its disagreements before treating the headline score as clinical evidence.

For Builders: Run the full product pathway, including model, system prompt, filters and handoff, rather than a favorable base-model snapshot.

For Buyers: Put an independently reproduced VERA-MH result, version number and failure analysis into procurement requirements.

For Regulators: Support external clinician validation and require comparable, public reporting instead of accepting vendor-native safety metrics.

---

.  .  .

CLOSE.

Fifteen a year worldwide. That is the deaths. Rarer than lightning.

The illness moves like Typhoid Mary. Invisible, prolific, appearing helpful.

Four numbers would settle both. All four already inside the vendors.

TODAY’S QUESTION

The observed count sits below the base rate. Which reading do you take?

One tap. Results in tomorrow’s issue and on the web.

THE BOOK • OUT NOW

Therapist in the Loop

by Jess Jessop

One billion people live with a mental health condition. There will never be enough therapists. The machines are already in the room. This book is the map for what happens next.

The machine can help. It cannot be left in charge.

Kindle, hardcover, and paperback

MORE ON OUR RADAR.

  • Hawaii SB 3001 becomes law by silence Wednesday July 15. Gov. Josh Green named four intent-to-veto bills on June 27; SB 3001, the AI-companion-platform minors bill, was not among them. It clears the veto gate by omission tomorrow.

  • Missouri SB 1019 Kehoe clock closes Wednesday July 15. The bill including the state's therapy-chatbot ban has been on Gov. Mike Kehoe's desk since May; tomorrow is the constitutional decision date.

  • KIDS Act H.R. 7757 stalls in the Senate after House passage June 29. Senate sponsors publicly objected to the House-passed 267-117 version. No mark-up scheduled in either chamber this week.

  • Anthropic Persona ID formalized for flagged Claude accounts July 8. The government-photo-ID plus live-selfie flow that ran in limited use since April 14 is now the codified route for accounts Anthropic flags for abuse review.

  • Character.AI's January settlement covered five individual cases and no class-wide gag. New plaintiffs continue to file; the Garcia settlement did not close the docket.

  • CDC WISQARS has no external-cause code for AI or chatbot involvement. The federal death-surveillance infrastructure that would need to house any registry does not have a field to tick. This is exactly the ICD gap our Story 1 turns on.

If you or someone you know is in crisis, call or text 988 (Suicide and Crisis Lifeline).

Jess Jessop is the Founder and CEO/CTO of Clinician Assist Inc. (BetterMind.Space), building the first voice-first AI-native mental health EHR with Casey Life and Peer AI Coach supervised by licensed therapists. A disabled veteran and 25-year AI/software engineering veteran, Jess brings lived experience as a mental health client to the mission of making daily mental health care as integrated as oral care.

Reply

Avatar

or to participate