Dev.to · 11 min read

We Counted AI Hallucination Disclosures in Every 10-K. Pharma Outnumbered Software 29 to 10.

We Counted AI Hallucination Disclosures in Every 10-K. Pharma Outnumbered Software 29 to 10.

Thirty public companies told the Securities and Exchange Commission about "hallucinations" in their annual reports in 2022. ChatGPT was released in November of that year, near the end of it, so for most of 2022 the product that made the word famous did not exist. Thirty filings, in a year that mostly predates the thing everyone now means by the term. That number should feel wrong, and the wrongness is the whole story. Here is the setup. If you wanted to measure how fast corporate America started formally admitting, in the one place it is legally obligated to be candid, that its AI makes things up, there is an obvious way to do it. Search SEC filings for the word "hallucinations" and count them by year. The data is public, free, searchable through EDGAR's full-text index, and it produces a clean rising line: 30 filings in 2022, then 32, then 42, then 54, and then 166 so far in 2026. Five and a half times growth. You could put that on a slide. You should not. The line is confidently, checkably wrong, and the first correction anyone would reach for is wrong too, in the opposite direction. This is a small investigation into how a number that looks like a measurement can be nothing of the kind, and why the failure is one you have almost certainly shipped yourself. The word was taken Search EDGAR for 10-K filings containing "hallucinations" and read who shows up, and the mystery dissolves immediately. Many of them are drug companies. "Hallucinations" is not, to a pharmaceutical filer, a property of a language model. It is a clinical symptom, an indication, the thing their product treats. The cleanest specimen is ACADIA Pharmaceuticals, ticker ACAD, whose 10-K was filed on February 26, 2026. ACADIA's lead product is pimavanserin, approved for hallucinations and delusions associated with Parkinson's disease psychosis. Their annual report says "hallucinations" repeatedly, across many pages, because hallucinations are the entire commercial reason the company exists. It has nothing to do with artificial intelligence. It is a filing about a brain, not a model. ACADIA is not alone in the sample. KALA BIO, Madrigal Pharmaceuticals, Esperion Therapeutics, Harmony Biosciences, Repligen, and Black Diamond Therapeutics all surface the same way: real companies, real filings, the right keyword, the wrong referent. The naive count is not measuring AI-risk disclosure at all in its early years. It is mostly measuring the base rate of a neurological symptom in the annual reports of biotech firms, a rate that has nothing to do with the technology story and was chugging along at roughly thirty filings a year before the technology story began. That is why 2022 reads 30 and not zero. A rising line that had started at zero would have sailed straight through every review, because a zero baseline is what a genuine new phenomenon looks like. The nonzero baseline before the thing existed is the tell, and it is the only reason the contamination was ever caught. The obvious fix makes a different mistake So you correct it. The natural move is to narrow the search: require the filing to say "hallucinations" and also say "artificial intelligence." Now you are scoping to documents that are at least AI-adjacent, and the pharma noise should drop out. It feels like a fix. Run it, and the count for 2026 comes back at 159, barely below the naive 166. And when you classify who those 159 are, by the industry codes EDGAR attaches to every filing, the single largest category is still pharmaceutical preparations, at 29 filings against 10 for prepackaged software. A search built specifically to measure AI confabulation returns, as its most common industry, companies that treat literal hallucinations. The correction did almost nothing. The reason it did almost nothing is worth sitting with, because it is the transferable part. Requiring "artificial intelligence" only filters if mentioning artificial intelligence is discriminating. In 2026 it is not. Nearly every 10-K of any size now mentions AI somewhere across its several hundred pages, in a risk factor or a strategy paragraph or a boilerplate sentence about the competitive landscape. ACADIA mentions it. So the conjunction search pulls ACADIA, and pimavanserin, and Parkinson's psychosis, right back in. A filter stops filtering the exact moment its discriminator becomes universal, and nothing in the output announces that it happened. The query still runs. It still returns a tidy number. The number is just no longer about what you think. The three lines, side by side Here is the whole dataset, from EDGAR's own response counts, retrieved on August 11, 2026. Each cell is the number of 10-K filings the search returned for that year, not an estimate and not a model's summary of anything. Year "hallucinations" + "artificial intelligence" exact phrase "AI hallucinations" 2022 30 2 0 2023 32 5 0 2024 42 22 2 2025 54 36 3 2026 166 159 38 Two honest caveats about that bottom row, and neither is allowed to do quiet work. First, 2026 is a partial year: the window runs from January 1 to August 11, so it is not like-for-like with the four completed years above it. Second, it is less partial than that sounds, because most calendar-year companies file their 10-K in February and March, so the window already captures the bulk of the annual cohort rather than a random third of it. Both facts are true, and the piece needs both, because with just one of them a reader would either over-trust or over-discount the same number. The growth rates tell the real story. The naive series grows 5.5 times from 2022 to 2026. The conjunction series grows about eightyfold, from 2 to 159. The exact-phrase series grows from zero to 38. Those are three wildly different pictures of the same underlying event, and the gap between them is not noise. It is the shape of the measurement error. Two errors, and the second one hides inside the first Name what went wrong precisely, because the names are the reusable part. The first error is a contaminant that inverts over time. In 2022, the pharmaceutical noise was roughly ninety percent of the naive "hallucinations" count. By 2026, with 166 filings, that same base rate of clinical usage is a small minority, maybe four percent of the total. A contaminant whose absolute size holds steady while the real signal explodes does not merely add error to the series. It manufactures a fake growth curve out of the shrinking proportion, and at the same time it flattens the true one, because the inflated early baseline makes the later growth look smaller than it was. The naive line shows five and a half times growth. The thing the line is trying to measure grew something like eighty times. The contaminant stole most of the slope. The second error is subtler, and it lives inside the correction rather than the original. Adding "artificial intelligence" to the query is a filter whose discriminating power decayed to nearly nothing over the same window, for the same reason the signal grew: AI went from niche to universal in corporate disclosure. A filter is only as good as the rarity of the thing it filters on. The moment the discriminator is present in almost every document, the filter passes almost everything, and it does so silently. There is no error message for "your keyword is now boilerplate." The 2022 conjunction count of 2 and the 2026 conjunction count of 159 are not measuring the same construct, because "mentions AI" meant something specific in 2022 and means nearly nothing in 2026. So what is the real number A range, with a named reason for each edge, because a point estimate here would be its own kind of lie. The floor is 38: the filings that use the exact phrase "AI hallucinations" in 2026. That referent is unambiguous. Nobody writes "AI hallucinations" to mean Parkinson's psychosis. But 38 undercounts, because the more common drafting style never uses the tidy phrase. Lawyers write "the model may produce inaccurate, incomplete, or fabricated outputs," or "hallucinations, inaccuracies, or errors," and every one of those filings is a real AI-risk disclosure that the exact-phrase search misses. The ceiling is about 97: the 159 conjunction hits, minus the pharma-class share. That share has to be estimated rather than counted, and here is the limit I have to state plainly. EDGAR's full-text search returns at most 100 results per page, and there are 159 hits, so the industry breakdown above is a sample of 100, not a census. In that sample, pharmaceutical and health-related industry codes account for 39 of 100 filings. Extrapolated across all 159, that is roughly 60 pharma-class filings, leaving about 97 that plausibly refer to the AI meaning. So the honest answer to "how many public companies disclosed AI hallucination risk in 2026" is: somewhere between about 38 and 97, and the width of that band is the actual finding. What would tighten it is reading the sentence around each of the 159 hits to classify its referent directly, or a proximity search that EDGAR does not offer. I did not do the former. The classification here is by industry code, which is a proxy for referent and not a reading of each document, and the essay is worth less if it pretends otherwise, given that the essay is about exactly this kind of pretending. Why this is your problem too None of this is really about pharmaceuticals or the SEC. It is about a failure mode that is completely generic, and if you have ever built an eval, a dashboard, a log alert, an incident tagger, or any chart whose title is "mentions of X over time," you have this failure mode in your stack right now. The mechanism is a keyword whose referent silently changes across the corpus. The word "hallucination" carries two unrelated meanings, and a corpus that shifts from mostly-one to mostly-the-other over a few years will produce a confident trend line that measures the shift in mix rather than the growth of either thing. The tell, the single cheapest diagnostic, is a nonzero baseline before the phenomenon existed. If your "AI incidents" counter reads meaningfully above zero in a year before you had any AI, you are not counting AI incidents. You are counting a homograph, and the number will keep lying to you in a direction you cannot predict from the number alone. I will give the last example against myself, because it happened, and because inventing a cleaner one would violate the whole point. This essay went through a routine duplication check against our own archive before it was written, a tool that flags when a new piece overlaps an old one by shared named entities. It flagged a match. The overlapping essay was about induced demand in highway construction, and the shared entity that triggered the alert was "Parkinson's." In the highway essay, "Parkinson's" is Parkinson's Law, the 1955 observation that work expands to fill the time available. In this one, "Parkinson's" is Parkinson's disease, the indication for the drug in the specimen filing. Same string, two referents with nothing in common, a confident match that means nothing. Our own tooling, on the essay documenting the homograph problem, committed the homograph problem, through the very word that carries the ambiguity. I did not stage that. I checked the match rather than waving it through, because a scan's output is a finding and not a verdict, and if "Parkinson's" had turned out to be the disease in both places it would have been a real collision worth stopping for. It was not. It was the same error the drug companies' filings produce, in miniature, live, in-house, on the same afternoon. Which is the most honest evidence I can offer that this is not a story about the SEC. It is a story about what happens when you count words and trust the total, and it is running quietly inside more of your measurements than you would like. Sources. All filing counts from SEC EDGAR full-text search (efts.sec.gov), form type 10-K, by calendar-year window, retrieved 2026-08-11; each count is EDGAR's own returned total for the query. Specimen: ACADIA Pharmaceuticals Inc. (NASDAQ: ACAD, CIK 0001070494), 10-K filed 2026-02-26; product and indication per ACADIA's own disclosures. Industry classification by SEC SIC code from the same EDGAR results (a sample of 100 of 159 hits; a proxy for referent, not a per-document reading). The 2026 figures cover January 1 to August 11, 2026, a partial but front-loaded year.

This is a summary aggregated from Dev.to. Read the complete article on the original site:

Read full article at Dev.to

More AI & Machine Learning News