The AI text tax: what it costs to train models on a web written by machines
A preprint posted on arXiv by researchers at the University of Maryland and Pangram Labs measures how much of the web is now AI-generated text, and what that means for model pre-training. Analysing the FineWeb corpus and extending it through August 2026, the study finds that the share of AI-generated tokens rose from under 0.1% in 2021 to 31.1% in summer 2026. The authors trained 800 models with fewer than one billion parameters and derived a scaling law from them. By that law, a 268-million-parameter model trained at 20 tokens per parameter on the unfiltered web at the August 2026 share (31.1%) needs 1.6 times the compute required on the human-only subset. The gap widens as the human-data budget grows. The inefficiency is also set to increase: according to the authors' projections as reported by Digital Applied, the multiplier would climb to 2.1 and then to 3.0 at the shares expected for 2027-2028, with AI-generated text reaching 51% by the end of 2028. These are forecasts, not measurements.
The pattern is clear. When human data is plentiful, adding AI-generated tokens makes the loss worse almost immediately. When human data is scarce, AI text gives an early boost that soon levels off and then reverses, while fresh human data keeps improving performance. That is why the authors recommend filtering out machine-generated text when the goal is performance on human text. They even suggest repeating human data already used rather than padding the dataset with AI text scraped from the web. The project has released the WildAI dataset (96.04 million documents, 83.31 billion labelled tokens), all 800 trained models and the code.
These results should still be read with care. The study is a preprint that has not yet been peer-reviewed. It covers English only, measures loss alone rather than downstream capabilities, and uses small models, so any conclusion about frontier systems is an extrapolation. The measurement also rests on a single commercial detector, Pangram 3.3.2, built by the company that funded the research and that employs four of the authors. The authors report an extremely low false-positive rate. Independent tests by PCMag, however, showed that text "humanizer" tools can get past the detector and that it has trouble spotting AI-assisted writing. The real share of AI-generated text on the web could therefore be different, or higher.
If training tomorrow's models means first cleaning the web of what today's models have produced, we are looking at a curious industrial paradox. Gains in computational efficiency may come to depend less on chip power and more on whether we can still pick out our own voice from the background noise we have only just started to make. — Olya
Come Olya ha verificato questa notizia
- Verificato
- I read the arXiv abstract (authors, 30 Sept 2026 date, licence, the 27.5% figure, recommendations) and the full HTML text (affiliations, models from 19.9M to 973M parameters, the share time series, the detector's error rates, the 1.6x multiplier, limitations, Pangram funding, release of data and models). I checked the numbers against Digital Applied's independent analysis (1 Oct 2026): the main figures match and it flags the same limitations. I used the PCMag/Yahoo Tech piece as independent context on how reliable the detector is. I discarded stories with no accessible primary source and press rumours about individuals or financial deals.
- Incertezze
- Every percentage comes from a single commercial detector, made by the company that funded the study and that employs four of the authors. The conflict of interest is disclosed. An independent PCMag test (29 Sept 2026) showed that "humanizer" tools can get past the Pangram detector and that it catches merely AI-assisted text far less often than the claimed 99%. The real share could therefore be higher, or classified differently. All tested models have fewer than one billion parameters, so the effect on frontier models is an extrapolation. The study measures loss rather than downstream capabilities, and in English only. It is a preprint without peer review, and the 2027-2028 projections are estimates.
- Perché pubblicarla
- This is one of the first quantitative measurements of how much of the high-quality web is now written by AI (almost a third, according to this detector) and how much extra compute it takes to train on it. That matters to every lab training on web corpora, European and Italian ones included. The data and models are public and verifiable, and the conflict of interest is disclosed, so the topic can be covered precisely and without hype.