Shuichiro Ogawa
日本語

Notes · updated 2026-08-28

The Day It Got a Name

In May 2024, Willison gave the name slop to “unwanted AI-generated content.” He wrote that it was a new word. But the act of naming itself is not new. In 1994, spam performed the same function, and in nineteenth-century Britain there was a name for it: penny dreadful.

Once something has a name, countermeasures begin. What was merely a complaint before it had a name becomes, once named, something that gets classified, measured, and made the object of institutions.

This much has repeated itself before. The problem comes next: the history of countermeasures is not a record of victories. The same shape of wager has been placed again and again, and has broken down in the same shape again and again. Where is this time the same, and where is it different?

What Happened After the Printing Press

When books suddenly multiplied, the first things scholars built were indexes and commonplace books.

Blair (2010) traces how early modern European scholars responded to the print-driven surge in books with the technologies of indexes, summaries, commonplace books, and bibliographies. The commonplace book, as this kind of excerpt collection was called, was a tool for organizing quotations by topic, a technique for deciding what to keep and what to discard from a mass of text (Blair 1992, 2004). Selection technologies were invented because volume increased.

Condemnation of cheap, mass-produced publications is just as old. Darnton (1971) drew on police surveillance records to depict how hack writers gathered on the Grub Street of pre-revolutionary France mass-produced seditious pamphlets. In Victorian Britain, the cheap boys’ reading material known as penny dreadfuls was blamed as a cause of juvenile crime and set off a moral panic (Dunae 1979). Contemporary critics called it “wastes of print” (Burz-Labrande 2021). The stigma of being “shoddy and unfiltered” that attached to vanity publishing was later transferred onto print-on-demand and the web (Laquintano 2013).

The people responsible for selection were also given a name. Lewin (1947) coined the term gatekeeper, and White (1950) followed a wire editor at a local newspaper for seven days, demonstrating empirically that only a small fraction of the wire service’s dispatches made it into print. Shoemaker and Vos (2009) organized the subsequent half-century of research into five levels: individual, routine, organizational, social institutional, and social system.

And Simon (1971) wrote the sentence that would become the premise for all the discussion to follow. A wealth of information creates a poverty of attention. An information-processing system conserves an organization’s attention only by not outputting more than it takes in.

Economics already had a formulation for what happens when quality cannot be observed. Akerlof (1970) showed that in markets where quality cannot be observed, low quality can drive out high quality, and Spence (1973) showed that a cost only the high-quality party can afford functions as a signal. These two results become the backbone of every countermeasure that follows.

After Spam Got Its Name

In 1978, a DEC marketing representative sent a mass product announcement to ARPANET users. In 1994, an immigration law firm posted its advertisement individually to more than 5,500 Usenet newsgroups (Ferrara 2019). After these two incidents, the word spam came to refer to unwanted commercial communication in general.

Brunton’s (2013) definition is concise. Spam is the use of information infrastructure to exploit an existing accumulation of attention.

The scale was also measured. Rao and Reiley (2012) estimated the annual cost borne by US firms and consumers at about 20 billion dollars and spammers’ worldwide annual revenue at about 200 million dollars, putting the ratio of external cost to internal benefit at roughly 100 to 1. Most of the cost falls on the side that did not send anything.

From here, five wagers are placed in turn.

Wager One: The Recipient Classifies

The most successful countermeasure was for the receiving side to look at the content and classify it.

Sahami, Dumais, Heckerman, and Horvitz (1998) used a decision-theoretic framework and probabilistic learning to show that adding features beyond the message body substantially improved filter accuracy. Androutsopoulos et al. (2000) confirmed on a public corpus that Naive Bayes consistently outperformed keyword rules. In 2002, Graham published a design for a learning filter in which users tagged messages as spam or ham, and implementations spread widely.

Why did it work? The reason lies less in the accuracy itself than in where the filter was placed. Learning happened per recipient, and its contents were invisible to attackers. Because attackers had to aim at a different target for each individual, their optimization gained little traction.

The same wager was placed again in 2023, and it was folded within half a year. OpenAI released its AI Text Classifier on January 31, 2023, and withdrew it on July 20. The reason was insufficient accuracy: by OpenAI’s own figures at release, the correct-detection rate was 26 percent, and the rate at which human writing was misclassified as AI-generated was 9 percent.

The same problem attaches to detectors that were not withdrawn. Liang et al. (2023) showed that seven major detectors misclassified more than half of TOEFL essays written by non-native speakers as AI-generated. Having ChatGPT lightly rewrite the same essays drops the detection rate to nearly zero. In the recursive paraphrasing attack of Sadasivan et al. (2023), watermark detection fell from 99.8 percent to 9.7 percent at a false-positive rate of 1 percent. Kumarage et al. (2023) showed that prompts optimized to evade a detector transfer across different models.

The difference from Bayesian filters is not a matter of accuracy figures. Spam carried the intent to sell in its vocabulary. There was a price, a link, language urging immediate action. Slop has none of this. It is written in the same vocabulary as legitimate text, so content offers little to go on for telling them apart.

Wager Two: Impose a Cost on the Sender

Dwork and Naor’s (1992) proposal was straightforward. Impose a moderately difficult computation on every message sent. A load that is negligible for one message becomes untenable for a million. Hashcash, formalized by Back in 2002, is the representative implementation of this idea.

In 2004, Laurie and Clayton closed out this wager. Spammers do not use their own computers. They use the computing resources of hijacked botnets, so the cost never falls on them. Legitimate users’ device capabilities, meanwhile, vary widely. No load level exists that separates the two. The paper’s title is “‘Proof-of-Work’ Proves Not to Work.”

The idea of imposing cost did not disappear altogether. Loder, Van Alstyne, and Wash (2004) designed a scheme in which senders post money rather than compute, refunded if the recipient judges the message worthwhile. It has not been widely implemented.

The cost asymmetry appears in an even stronger form this time. The generator does pay the cost of generation, but the amount is small. And what has fallen is only the cost of producing low quality; the cost of checking it has not fallen.

There is, however, one example where changing who bears the cost changes the outcome. Tchernichovski et al. (2019) imposed a time cost on the person entering a rating in an online rating system. The reliability of ratings improved. The cost is imposed not on the sender but on the act of the one who judges.

Wager Three: Certify Provenance

If content cannot distinguish them, make the sender prove who they are. The world of email spent twenty years on this wager. SPF (RFC 7208), DKIM (RFC 6376), and DMARC (RFC 7489) were standardized as mechanisms for verifying the legitimacy of a sending domain.

Adoption rates became the problem. Durumeric et al. (2015) measured that, of more than 700,000 SMTP servers, only 35 percent had encryption correctly configured, and only 1.1 percent specified a DMARC policy. In more than seven countries, over 20 percent of messages arriving at Gmail arrived unencrypted. Foster et al. (2015) measured the gap in which large providers implemented these mechanisms across the board while small and midsize senders lagged behind.

The provenance wager is being placed all over again right now. The Content Authenticity Initiative launched in November 2019, C2PA was founded on February 22, 2021, and specification v1.0 was published on January 26, 2022. Google DeepMind began offering the image watermark SynthID on August 29, 2023, and in 2024 published a text-oriented method in Nature (Dathathri et al. 2024). Institutions moved too. China’s labeling mandate took effect on September 1, 2025, and Article 50(4) of the EU AI Act applies from August 2, 2026.

And this wager has produced a kind of result that did not exist last time. Zhang et al. (2024) proved that, under natural assumptions, “strong watermarking” is theoretically impossible. Fan et al. (2024) showed that a single publicly verifiable scheme cannot simultaneously satisfy robustness and soundness. On the implementation side, Hu et al. (2024) removed Meta’s Stable Signature through fine-tuning, and Hönig et al. (2024) showed that artist-protection tools like Glaze and Nightshade can be circumvented with simple upscaling. An independent security analysis of the C2PA specification (Golaszewski et al. 2026) concludes that the specification fails to achieve the goals it set for itself.

What provenance answers is “who,” not “how good.” And so long as adoption remains voluntary, the absence of proof is proof of nothing.

Wager Four: Prove You Are Human

The CAPTCHA of von Ahn, Blum, Hopper, and Langford (2003) was beautiful as a piece of design. It reduced security to an AI problem that was, at the time, unsolved. If a program appeared that could solve it with high accuracy, that would mean a hard AI problem had been solved, so either outcome would benefit humanity.

This design has an expiration date built into it. Because its security depends on a shortfall in AI capability, it lapses once that capability rises.

That is exactly what happened. Bursztein et al. (2014) identified character segmentation as the core of text CAPTCHA security and broke a large number of deployed schemes with a generic pipeline. Ye et al. (2018) used a GAN to generate large volumes of synthetic data to pretrain a solver, and solved 33 schemes, including those of major sites, at roughly 0.05 seconds each. It was the best paper at CCS 2018.

The side effects have also been measured. Khattak et al. (2016) measured how anonymous users connecting through Tor receive discriminatory treatment from websites, including outright refusal, restricted functionality, and CAPTCHA challenges, and estimated that more than 1.3 million IP addresses block connections from Tor exit nodes. CAPTCHA has also been reported as a real barrier for users with visual impairments and learning disabilities (Kuzma et al. 2011; Gafni & Nagar 2016).

There is one more property, easy to overlook. reCAPTCHA redirected the labor of solving into digitizing books (von Ahn et al. 2008). It was deployed on more than 40,000 sites, and more than 440 million words were transcribed. This countermeasure, which demanded proof of humanity, was at the same time a levy of unpaid labor.

Wager Five: Install a Gatekeeper

Peer review is not as old as it is thought to be. Baldwin (2018) argues that peer review came to be regarded as a central procedure of scientific practice only in the latter half of the twentieth century, a product specifically of the 1970s, born out of the tension between scientific autonomy and public accountability in Cold War America. Csiszar (2018) depicted the nineteenth-century system of judgment as a product of political compromise, intellectual property disputes, and commercial pressures.

That gate was breached early on by machine-generated text. In 2005, output from SCIgen, an automatic generator built by three MIT students, was accepted at a conference without peer review. Labbé and Labbé (2013) built a detection method and showed that bibliographic services had indexed SCIgen-generated papers. In 2014, Springer and IEEE retracted more than 120 papers from their conference proceedings.

Even as the generation tools changed, the pattern continued. Cabanac, Labbé, and Magazinov (2021) named the unnatural paraphrases produced by automatic synonym substitution tortured phrases and used them as a fingerprint for detection. If a paper reads “counterfeit consciousness,” the original was artificial intelligence. The Problematic Paper Screener they built scans 130 million records weekly (Cabanac et al. 2022).

Lists naming gatekeepers were also built, and then folded. Beall’s list began in 2010 and closed in January 2017. Strinzel et al. (2019) compared blacklists and whitelists across sources and showed that the two overlap and contradict each other, and that blacklists risk misclassifying emerging and small journals from the Global South as predatory. Xia et al. (2015) demonstrated empirically that it is mainly early-career researchers from developing countries who submit to such journals. In 2019, when 43 people from 10 countries reached a consensus definition after 18 questions and 12 hours of discussion, they explicitly stated that being open access is not in itself an indicator of predatory status (Grudniewicz et al. 2019).

Retraction, too, became institutionalized. Fang, Steen, and Casadevall (2012) reviewed all 2,047 retracted papers in their set and showed that retractions for simple error accounted for 21.3 percent, while misconduct accounted for 67.4 percent. In 2023, retractions exceeded 10,000 for the year, recording a 2.5-fold increase over the previous year (Van Noorden 2023). But retraction does not stop a paper from continuing to be used. In the large-scale analysis by Hsiao and Schneider (2022), only 5.4 percent of citation contexts occurring after a retraction mentioned the retraction.

Knowledge communities moved faster. Stack Overflow temporarily banned posting ChatGPT-generated answers on December 5, 2022. The stated reason was that the low rate of correct answers caused substantive harm. Clarkesworld stopped accepting submissions at noon on February 20, 2023. That month, submissions believed to be AI-generated exceeded 500, approaching the 700 submitted by humans. Nature and Science announced author policies in January 2023, and ICML that same year barred, in principle, the use of LLMs to generate paper text. arXiv changed its policy on October 31, 2025, to require proof of prior peer review for review articles and position papers in its computer science category. Wikipedia launched its AI Cleanup project in December 2023 and added, in 2025, a policy of speedy deletion for articles that are clearly LLM-generated.

The ban had an effect. Borwankar et al. (2023) used a difference-in-differences design to show that, after Stack Overflow’s ban, the linguistic complexity and word count of answers rose significantly relative to a control group. There was no significant effect on the number of questions. But by a different measure, the community itself has been shrinking. Burtch, Lee, and Chen (2024) reported that after ChatGPT’s emergence, Stack Overflow’s visits and questions declined significantly (they did not decline on Reddit), and that the decline was concentrated among new entrants.

Installing a gatekeeper requires labor to maintain that gatekeeper. Halfaker and Geiger (2020) estimated that detecting damaging edits on Wikipedia by human labor alone would require 483 person-hours per day across all language editions. The burden of peer review is also extremely skewed toward a small number of reviewers (Kovanis et al. 2016).

And gatekeepers break down once they are measured. Fire and Guestrin (2019) analyzed more than 120 million papers and showed that metrics such as paper counts and citation counts malfunction exactly as Goodhart’s law predicts. Author counts rise, papers grow shorter, and publication pace accelerates. Asubiaro et al. (2024) quantified how the inclusion criteria of Web of Science and Scopus produce marked disparities across regions.

The Failures Follow Types

Lining up the five wagers reveals that they break down in only a few distinct ways.

Type 1: expiration date. A countermeasure designed on the premise that attackers lack a certain capability lapses once that capability rises. CAPTCHA was designed to bet its security on AI’s unsolved problems, so the moment AI could solve them, it ended exactly as designed.

Type 2: where the cost gets passed. A countermeasure that imposes cost does not work against an opponent who can push that cost onto someone else. Computational cost never reached spammers who owned botnets.

Type 3: voluntary adoption. Only honest participants use proof of provenance. When DMARC policy specification stays at 1.1 percent, the absence of proof is not evidence of guilt.

Type 4: the bias of errors. Gatekeepers make mistakes and shut people out. And those errors concentrate on the weaker side. Blacklists disproportionately misjudged small journals from the Global South, bot countermeasures disproportionately misjudged Tor users, and AI detectors disproportionately misjudged non-native speakers.

Type 5: concentration. Countermeasures that worked narrow down who can verify. Only large providers can operate email authentication end to end; small and midsize senders lag behind (Foster et al. 2015). The more successful a countermeasure, the more the power to verify accumulates in a small number of places.

Where This Time Differs

If the types are the same, one is tempted to say this time is the same too. Saying so outright misses three things.

First, the trace is faint. Spam left the intent to sell in a visible form. A price, a link, language urging haste. Generated text carries none of this, and people can distinguish AI-generated poetry only at below-chance accuracy (discrimination accuracy of 46.6 percent, detailed in ai-slop-design-creativity). The wager of telling things apart by content is being placed under worse conditions than last time.

Second, impossibility has been proven. The history of anti-spam countermeasures produced no result like that of Zhang et al. (2024). Proof-of-work was argued to “not work in practice,” but that was an argument about economics and the distribution of resources, not mathematical impossibility. For strong watermarking, impossibility has been shown under natural assumptions.

Third, there is no evidence that anything has worked. Within the scope of this collection, no peer-reviewed research was found showing that detection, labeling, or regulation has actually reduced the volume or circulation of slop. What was found falls into two kinds. Experiments measuring the effect of labels on recipients’ cognition, and technical evaluations showing that detection and watermarking are defeated by attacks. In the former, AI labels lower an article’s perceived accuracy and credibility but leave policy support and engagement behaviors like likes and shares largely unmoved (Wang et al. 2025, n=3,861; Berinsky et al. 2025, n=7,579; Gamage et al. 2025, n=911). Attitudes change; behavior does not.

What History Left Behind: The Conditions for a Countermeasure That Works

The record is not only one of losses. Countermeasures that worked share something in common.

One, learn individually, in a place invisible to the attacker. The Bayesian filter worked not because of its accuracy but because the target differed for each individual.

Two, strike the economic pressure point. Levchenko et al. (2011) decomposed the spam value chain into naming, hosting, payment, and fulfillment, and through more than 100 actual purchases identified a bottleneck: payment processing is concentrated among a small number of providers. Rather than classifying content, stop the place where money flows.

Three, shift where the cost falls, from the sender to the act of the one who judges. In the experiment by Tchernichovski et al. (2019), imposing a time cost on rating raised the reliability of ratings.

What the three share is that none of them try to tell things apart. Each changes the structure instead. Who learns, where the money stops, who pays the cost.

The sister note ai-slop-design-creativity read slop as a question of where the burden of verification is passed. Seen from the side of history, that reading applies directly to the design of countermeasures. Asking “can it be told apart” breaks down under one of Type 1 through Type 4. Asking “who verifies, and is that person paid for it” brings at least the three forms that worked in the past into view.

When This Reading Fails

Three points are worth setting down.

First, suppose detection actually wins. The impossibility proof concerns the formulation of “strong watermarking” specifically, and does not rule out a weaker guarantee sufficient for practical purposes. If attacking it costs something, an imperfect watermark could still suffice operationally.

Second, suppose voluntary adoption stops being voluntary. Article 50 of the EU AI Act taking effect (August 2, 2026) and China’s labeling mandate (in force since September 1, 2025) are attempts to turn adoption of provenance into a legal obligation. DMARC stayed at 1.1 percent because there was no obligation; with an obligation, Type 3 may not apply. Measuring the effect remains a task for the future; as of this note, no results are in.

Third, suppose recipients were never bothered to begin with. The finding that labels do not change behavior may be a sign that recipients simply do not care. If so, the countermeasure is not failing to work; it is not being asked for.

The scope should also be limited explicitly. This note is heavily skewed toward English-language literature. No peer-reviewed research measuring the effect of Google Panda (2011), and none directly measuring whether anti-spam countermeasures produced an oligopoly in delivery infrastructure, was found in this collection. “Not found” is not “does not exist.” Literature on countermeasures since 2022 has a high proportion of preprints, many of which have not gone through peer review.

Timeline

YearEventType
1947Lewin introduces the gatekeeper conceptGatekeeping
1950White demonstrates a newspaper editor’s selection empiricallyGatekeeping
1970–73Akerlof’s market for lemons, Spence’s signalingTheory
1971Simon: “a wealth of information creates a poverty of attention”Theory
1978DEC’s mass-mailed emailPollution
1992Dwork & Naor propose proof-of-workCost
1994Canter & Siegel post to more than 5,500 newsgroupsPollution
1994robots.txt (established as a de facto standard)Norms
1998Sahami et al.’s Bayesian filterClassification
2002Graham’s “A Plan for Spam”; Back’s HashcashClassification / Cost
2003von Ahn et al.’s CAPTCHA; the US CAN-SPAM ActPersonhood / Institution
2004Laurie & Clayton demonstrate the failure of proof-of-work; TrustRankCost / Classification
2005A SCIgen-generated paper is accepted at a conferencePollution
2008reCAPTCHA redirects labor into book digitizationPersonhood
2010Beall’s list beginsGatekeeping
2011Levchenko et al. identify the payment bottleneckEconomy
2011–15Standardization of SPF / DKIM / DMARCProvenance
2014Springer and IEEE retract more than 120 SCIgen papers; generic CAPTCHA-breakingGatekeeping / Personhood
2015Durumeric et al. measure DMARC policy specification at 1.1 percentProvenance
2017-01Beall’s list closesGatekeeping
2018A GAN-based CAPTCHA solver breaks 33 schemesPersonhood
2019-11Content Authenticity Initiative foundedProvenance
2021-02C2PA foundedProvenance
2021Detection via tortured phrasesClassification
2022-01C2PA specification v1.0Provenance
2022-09RFC 9309 (formal standardization of robots.txt)Norms
2022-12-05Stack Overflow bans ChatGPT-generated answersGatekeeping
2023-01Nature and Science announce author policies; OpenAI releases its detectorGatekeeping / Classification
2023-02-20Clarkesworld halts submissionsGatekeeping
2023-07-20OpenAI withdraws its detector (26 percent accuracy)Classification
2023-08Glaze appears at USENIX Security; SynthID betaCounter-tech / Provenance
2023-12Wikipedia’s AI Cleanup launchesGatekeeping
2024Zhang et al. prove the impossibility of strong watermarking; SynthID-Text published in NatureProvenance
2025-09-01China’s labeling mandate takes effectInstitution
2025-10-31arXiv requires proof of prior peer review for review articlesGatekeeping
2026-08-02EU AI Act Article 50(4) begins to applyInstitution
  • ai-slop-design-creativity: The sister note that read AI slop as the outsourcing of verification. This note examines that reading from the side of history.
  • maker-to-editor-paradigm: The shift in role from maker to editor. This corresponds to a description, from the side of professional practice, of the wager of installing a gatekeeper.
  • llm-as-a-judge-literature: The question of whether evaluation can be handed back to machines. The current state of attempts to bring the classification wager into the verification process.
  • design-value-dimensions: A map of value that serves as a premise when thinking about proxy measures of quality.
  • generative-art-literature: The context of the arms race over generated content and copyright, including Glaze and Nightshade.
  • adoption-approval-paradox: The divergence between adoption rate and approval rate. It sits in the same stratum as the finding that labels change attitudes but not behavior.

References

Prehistory: Information Overload and Selection Technologies

The History of Spam and Countermeasures

Search Pollution and Proof of Personhood

Academic Publishing and Knowledge Communities

AI Countermeasures Since 2022

Primary Sources for Dates (Non-Academic)

Unverified Items

The unverified items for the full corpus (140 sources) are listed at the end of source/review/slop-countermeasures-history/papers.md. Collected here are the ones relevant to the body text.

  • The specific conversion-rate figure from Kanich et al. (2008) and the specific adoption-rate figure from Foster et al. (2015) could not be confirmed in the primary text, so no numbers for either are given in the body.
  • No peer-reviewed research measuring the effect of Google Panda (2011) was found in this collection.
  • No peer-reviewed research directly measuring whether anti-spam countermeasures produced an oligopoly in email delivery infrastructure was found either. Type 5 (concentration) is an inference from the asymmetry shown by Durumeric (2015) and Foster (2015), not empirical demonstration of the oligopoly outcome itself.
  • On the circumstances of Beall’s list closing, the original primary-source testimony from Beall himself has not been reached.
  • Methodological criticism of Bohannon (2013) is mentioned in multiple places, but this collection has not identified the critical paper itself.
  • The formal volume and issue number for Borwankar et al. (2023), and whether either of Cabanac et al.’s two papers has a peer-reviewed journal version, remain unconfirmed.
  • Among the literature on countermeasures since 2022, Zhang et al. (2024), Kirchenbauer et al. (2023), Dathathri et al. (2024), and Shan et al.’s two papers have gone through peer review, while Hu et al. (2024), Hönig et al. (2024), Golaszewski et al. (2026), and Wang et al. (2025) are preprints.
  • The formal DOI for Berinsky et al. (2025) and the peer-review status of Golaszewski et al. (2026) are unconfirmed.
  • The text of Article 50 of the EU AI Act was confirmed via an explanatory site and has not been cross-checked against the official text in the EUR-Lex gazette. China’s labeling mandate has likewise not been cross-checked against the government’s original text.
  • The date Clarkesworld reopened submissions, and the exact effective date of Wikipedia’s G15 policy, could not be confirmed in primary sources.
  • The organization itself, that countermeasure failures follow five types, is this note’s own, and no formulation of this shape was found in the prior literature within the scope of this search. Comprehensiveness is not guaranteed.

← All Notes · Home