Shuichiro Ogawa
日本語

Notes · updated 2026-08-28

The Same Picture Becomes Something Else the Moment It Is Sent

Someone generates a picture for their own enjoyment. That is not called slop. The same picture becomes slop when it arrives as a deliverable, sent without the person who generated it ever having looked at it. Not one pixel of the image has changed. What changed is only who bears the burden of checking it.

The account that first fixed this word in place says exactly this. In May 2024, Willison introduced slop as a name for “unwanted AI-generated content,” placing its core not in poor quality but in behavior. Foisting on someone else something you have not verified yourself. That is the principle by which it counts as rude.

If so, slop is not a property that objects have. It is a distribution of the burden of verification between maker and recipient.

This reading touches design directly. Design is the work of constructing a relationship with a recipient. But touching it directly does not make it correct. What follows checks how far the academic literature supports this reading, and where that support gives out.

The Word Still Has No Settled Referent

A word that emerged from the developer community in 2024 appeared in a lead editorial in Science by 2026 (Thorp 2026). The journal presented mandatory disclosure of AI use in draft writing and a ban on AI-generated figures as “resisting AI slop.” It is a record of the word being absorbed into institutions.

But the definition has not settled. Madsen and Puyt (2025) built a typology along seven axes: volume, velocity, variety, value, verification, visibility, and virality. That verification is one of these seven axes matters for the argument that follows. Galip and Lupinacci (2026) capture slop and brainrot through the concept of the “gimmick,” in which excess and insufficiency coexist, pointing simultaneously to labor-saving deskilling and to excessive computation aimed at capturing attention.

There are also positions that defend slop. Kommers et al. (2026) treat slop not as pollution but as “a supply-side solution to a demand for content that exceeds what humans can supply,” crediting it with its own aesthetic value and a function in collective meaning-making. A response to this (Nishal, Sax & Kieslich 2026) criticizes that argument for ignoring sociotechnical context. The dispute itself is ongoing and unresolved.

The same holds on the measurement side. Shaib et al. (2025) start from the premise that no agreed definition or measurement instrument for AI slop exists, and build a taxonomy of quality-evaluation dimensions from expert interviews. The binary judgment of “is this slop” swings with subjectivity, but correlates with sub-factors such as coherence and relevance. That is where things currently stand.

Neither Hallucination Nor Bullshit

The distinction from neighboring terms is worth clearing away first.

Hicks, Humphries, and Slater (2024) argue that calling LLM misoutput “hallucination” is inaccurate. Hallucination implies a failure of perception, but LLMs are not failing at aiming for truth. The accurate description, they argue, is bullshit in Frankfurt’s sense: indifference to truth. Hannigan, McCarthy, and Spicer (2024) add the human side to this. Chatbot output is prediction, not knowledge, and they call it botshit when people bring that output uncritically into their work.

Both of these terms are concerned with the epistemic status of the output. Slop is different. It is established on a single point, prior to any question of whether the output is correct or wrong: that no one has checked it. Even accurate content is slop if it is foisted on someone unverified.

That is why slop is not a problem that improved accuracy will make disappear.

Individual Work Improves While the Collective Converges

Empirical creativity research has accumulated results that point in one direction.

Doshi and Hauser (2024) showed that short stories by writers exposed to LLM-generated ideas were rated more highly on creativity, polish, and enjoyment alike. The effect was larger for writers whose creativity had originally been rated lower. The same experiment produced a second result. AI-assisted works resembled one another closely, and collective novelty fell.

This asymmetry has been reproduced repeatedly. Moon, Green, and Kushlev (2025) analyzed 2,200 college admissions essays across three pre-registered studies, measured with a newly constructed metric. The novelty that one additional human essay adds to a pool of ideas was 2 to 8 times what one additional GPT-4 essay adds. Homogenization persisted even when prompts and parameters were adjusted to try to raise diversity.

Padmakumar and He (2024) isolate where the homogenization occurs. A significant drop in diversity in collaborative op-ed writing appeared only in the condition using RLHF-tuned InstructGPT; the base GPT-3 showed no significant difference. The decline stems mainly from the homogeneity of the text the model offers, and the portions users wrote themselves were unaffected. Wenger and Kenett (2026) further report that the resemblance among LLMs to one another is far stronger than the resemblance among humans to one another.

Large-scale observational data from Zhou and Lee (2024) shows quantity and quality at once. On a platform with over 4 million works and more than 50,000 users, the output of authors who adopted text-to-image generation rose 50% immediately and doubled by the following month. Ratings rose too. And both the average novelty of the content and its visual novelty fell.

Why Every Generated Screen Wears the Same Face

What design practitioners notice first is not the statistics but the look.

Simonen et al. (2026) computationally analyzed over 750,000 Midjourney-generated images and showed that large numbers of mutually unrelated prompts converge on strikingly similar images. The authors call these “default images.” Declos (2026), writing from aesthetics, lists five biases that appear systematically in generative AI images. Cultural bias, beautification of appearance, spectacle, exaggerated saturation and contrast, and kitsch symmetry. The argument is that this bias threatens cultural diversity and creates an aesthetic bubble.

Homogeneity of appearance feeds back into judgment as well. Kurosu and Kashimura (1995), in an experiment with 252 participants, showed that ratings of apparent usability correlate more strongly with aesthetic impression than with actual usability. Tuch et al. (2012) show that visual complexity and prototypicality determine aesthetic judgment within 50 milliseconds of presentation. Generated output tends to lean toward highly prototypical appearances. So the first impression of “this looks good” passes easily, and that impression diverges from a judgment of quality.

The production process itself is affected too. Wadinambiarachchi et al. (2024), in a between-subjects experiment (N=60), reported that visual ideation using AI image generation tools strengthened fixation and lowered fluency, diversity, and originality alike, compared with a condition that did not use them. The design fixation that Jansson and Smith (1991) identified, the tendency to be bound by examples seen earlier, is being reinforced by generative tools.

Everything so far is a story about output resembling other output. Why does that require a separate name, slop?

The Workmanship of Risk as a Coordinate

Pye (1968) divided making into two kinds.

The workmanship of risk is work in which the outcome depends on judgment, skill, and attention exercised during production, and can still fail up to the last moment. The carving knife can slip. If it slips, the maker bears the loss.

The workmanship of certainty is work in which the outcome is fixed before work begins. Molds, jigs, and machines constrain what will result in advance. Mass production belongs here, and quality assurance is correspondingly embedded in the process itself.

Production by generative AI is neither. It is not workmanship of certainty, because the output is not fixed in advance. Nor is it workmanship of risk. The prompt shapes the outcome to some degree, but it is not judgment applied continuously through the course of production that produces the result.

What remains is a structure in which the outcome stays uncertain while whoever would bear that risk is not standing in the maker’s position. Handed over unverified, the loss falls to the recipient. This structure is exactly what Willison’s definition was pointing at.

Borrowing Schön’s (1992) words makes the situation more concrete. Designing was an act of constructing problem and solution together, in conversation with the unexpected backtalk that materials return. Generative models return responses too. But a response becomes a conversation only once it is read. If it is passed along unread, what comes back is not material but inventory.

When Sennett (2008) described skill as embodied knowledge acquired through repetition and resistance, and Ingold (2013) described making as a correspondence between maker and material, both presupposed that the maker takes on the resistance. What is new is production that passes that resistance downstream.

Hernández-Ramírez and Ferreira (2024) reread claims that generative AI ends design labor as a critique of managerialism, arguing that design labor cannot be reduced to procedure and automation. Recast as a question of who takes on the risk, that argument becomes narrower and easier to test.

Only Production Got Cheaper

Seen from the economics side, the asymmetry becomes explicit.

Zhang and Zhang (2025), using a general equilibrium model, showed that LLMs push down the marginal cost of low-quality synthetic content asymmetrically, while the cost of high-quality production remains unchanged. This creates a structural incentive toward pollution. The cost of verification does not fall either. If anything, it rises in proportion to the volume that needs checking.

Measured data is consistent with this prediction.

Wu et al. (2026) cross-referenced metadata for 256 million Spotify tracks against a Deezer detector and reported that AI-generated music as a share of new releases rose from under 1% in January 2024 to over 40% in November 2025. 92.7% of AI tracks have fewer than 1,000 plays (versus 67.5% for human tracks). An experiment actually uploading to 11 streaming services found that disclosure policies were not functioning.

Chakrabarty et al. (2026) ran full-text detection on 14,419 self-published genre novels on Amazon and reported that the sales share of non-AI books, nearly 100% in early 2023, had fallen to about 60% by Q2 2026. Dolezal et al. (2026), using Internet Archive snapshots, judge that by mid-2025 about 35% of newly published websites were AI-generated or AI-assisted. The same study also states that while the rise of AI-generated text correlates with a decline in semantic diversity, an adverse effect on factual accuracy was not statistically supported. Volume and accuracy are separate problems, this result suggests.

One case put a number on exactly who bears the cost of verification. curl’s security team reported that in 2025 about 20% of bug bounty submissions were AI-generated slop (Stenberg 2025). The confirmation rate for genuine vulnerabilities fell from over 15% historically to under 5%. Over six years, not one genuine vulnerability has ever been found from a report generated by AI alone. In January 2026, the project announced the suspension of its bug bounty program altogether (Stenberg 2026). These figures are the maintainer’s own self-reported tally and have not undergone third-party verification. Even so, as a case in which the transfer of verification costs shut down an institution, it is hard to replace.

The same structure appears in academic publishing. Tang and Cai (2025) examined 200,000 systematic reviews from 2013 to 2024 and found that 299 of them, 0.15%, had incorporated retracted paper-mill articles into their evidence synthesis. 124 citations occurred even after retraction.

And the pollution loops back into the input. Shumailov et al. (2024) demonstrated model collapse: mixing generated output indiscriminately into the next generation’s training data causes the tails of the distribution to be lost. Yu, Kim, and Kim (2026) formalize the structure by which search and RAG consume AI-generated content as a source, as retrieval collapse. What Tredinnick and Laybats (2025) call epistemic decay is exactly this recursion, in which today’s synthetic output becomes tomorrow’s training data.

There is also an experiment that brings Akerlof’s market for lemons to bear on the choice of AI systems. Erlei et al. (2026), in a between-subjects design with N=330, showed that partial disclosure of accuracy alone significantly improves the efficiency of delegation decisions, while even full disclosure leaves reliance on superior systems underprovided. In a market where quality is not visible, high-quality supply is not readily rewarded.

Indistinguishable, Yet Marked Down Once Disclosed

On the recipient’s side sit two facts that, at first glance, do not fit together.

First: people cannot tell AI-generated work apart. In Porter and Machery’s (2024) experiment, discrimination accuracy for AI-generated poetry was 46.6%, below chance. Participants tended to judge AI poems as human-written, and high ratings for rhythm and beauty contributed to that error. Köbis and Mossink (2021), in an experiment with 830 participants, showed that discrimination fails when a human has curated the output but succeeds in the uncurated condition. Whether a curator was present, not the source itself, determined whether people could tell.

Second: people still rate work lower once told AI made it. Raj, Berg, and Seamans (2026), across 16 pre-registered experiments with 27,491 participants, showed that disclosing AI involvement in creative writing consistently lowers ratings. The effect is mediated by perceived authenticity and persists across different rating measures, different contexts, and different content types. Most mitigation strategies they tested, reframing perspective, anthropomorphizing, or framing the work as collaboration, did not work. Ansani et al. (2025) presented the identical musical performance falsely labeled as human or AI and confirmed that ratings fell in the condition believed to be AI. Magni, Park, and Chao (2024), across four experiments with 2,039 participants, showed that the same generated output receives lower creativity ratings once attributed to AI, positioning human judges as gatekeepers of creativity.

The two facts do not contradict each other. What is being evaluated is not the work but the provenance label.

What makes the label take effect is also reasonably well understood. Kruger et al.’s (2004) effort heuristic is foundational research showing that people rate a work more highly when told it involved more labor to produce. Heimstad, Wien, and Gaustad (2025), in two pre-registered experiments, confirmed a pathway in which an “AI-made” label lowers creativity ratings by way of perceived low effort. Newman and Bloom (2012) decompose the mechanism by which originals are valued above copies into the narrative of a unique creative act and physical contact with the creator. Arielli (2024), writing from aesthetics, argues that the decline in the investment of labor, time, and skill required for production implicitly drives negative evaluation of AI-generated work.

For design practice, this carries a concrete consequence. “Made by a person” is beginning to be used as a guarantee of quality. But that is not a judgment of quality; it is a proxy for one. The proxy works only as long as provenance is honestly declared. Rijsbosch, van Dijck, and Kollnig (2026) surveyed the state of watermarking implementation under the EU AI Act and report a gap between machine-readable marking and visible disclosure. Regulatory requirements and technical feasibility are not yet aligned.

Why “Become an Editor” Is Not Enough of an Answer

The shift in roles has already been observed.

Jang (2026), studying 34 undergraduates, compared workflows using generative AI with conventional digital production and examined the cognitive trade-offs involved as the designer’s role shifts from maker to curator of output. Hall (2025) assigns designers in the generative AI era four roles. Advocate, curator, orchestrator, and emotional mediator. Tsao et al. (2025), in a scoping review of 57 empirical studies from 2022 to 2025, conclude that creative professionals are developing strategies of using new capabilities while retaining control, neither simple acceptance nor simple rejection. Fields that place weight on embodied practice show stronger resistance, they also note.

But “becoming an editor” does not, by itself, stop the outsourcing of verification.

Anderson, Shah, and Kreminski’s (2024) comparative experiment with N=36 showed that participants using ChatGPT produced more, more detailed ideas, while their sense of ownership over the output fell. Moving into the position of selector and taking on responsibility for what one selects are not automatically connected.

Disclosure norms in the field bear this out as well. What Hwang et al. (2026) describe is the reality that freelance workers adopt a strategy of not disclosing AI use in advance and only addressing it after problems arise. The deliverable circulates without it being visible who verified how much, or how.

And whether the role of taking on verification will be priced remains unsettled. Demirci, Hannane, and Zhu (2025) reported that writing and coding job postings fell 21% within 8 months of ChatGPT’s release, and image-production job postings fell 17% following the release of image-generation AI. The postings that remain are more complex and command higher rates. Hui, Reshef, and Zhou (2024) present suggestive evidence that among freelancers in strongly affected occupations, employment and earnings fell, and those who had held higher ratings suffered more. In Teutloff et al.’s (2025) analysis of over 3 million job postings, demand for substitutable skills fell 20% to 50%, while demand for complementary skills rose.

Verification sits on the side of complementary skills. Whether that shows up as a price is what will determine whether the outsourcing stops.

When This Reading Fails

Three conditions under which the reading developed here, treating slop as outsourced verification, would fail are worth setting down.

First, if the cost of verification actually falls. If detectors and provenance technology work cheaply and reliably, confirmation becomes automated and foisting no longer holds. That is not where things stand now. Watermarking implementation shows a gap (Rijsbosch et al. 2026), and the reliability of detectors themselves remains an open research question (Sun et al. 2025). But as long as this is a technical problem, the possibility of solving it remains.

Second, if recipients are not asking for verification. If Kommers et al.’s (2026) position is correct, reading slop as a supply-side solution to a gap between supply and demand, then outsourcing is not harm but a transaction. It would describe recipients choosing not to check, in exchange for receiving content cheaply and in volume. The response to that position (Nishal et al. 2026) demands attention to context, and the dispute has not closed. At minimum, it is premature to assume that every recipient wants verification.

Third, if homogenization can be solved on the model side. Given that Padmakumar and He (2024) found no significant difference in the base model, at least part of homogenization may originate in post-training adjustment. If that can be fixed through tuning, the decline in diversity is handled as a problem of model design rather than a problem of design itself.

The scope of this note should also be bounded. No measurement standard for slop yet exists (Shaib et al. 2025). Much of the research used to estimate volume consists of pre-peer-review preprints and depends on detector accuracy. Because the word itself only coalesced in 2024, the accumulation of peer-reviewed empirical work remains thin. Data from the Japanese-language sphere falls outside the scope of this search.

The remaining question narrows to one. How does one put a price on the work of taking on verification? That, probably, is what the design profession will answer next.

  • maker-to-editor-paradigm: A note analyzing the shift in roles from maker to editor through 19 academic and 15 industry sources. This note’s “Why ‘Become an Editor’ Is Not Enough of an Answer” section supplements that one.
  • subtractive-creativity-design: An argument that treats subtraction as a generative operation symmetric to addition. Connects here through where not-choosing is positioned.
  • hallucination-as-creativity-resource-literature: A literature map of arguments that treat hallucination as a resource for creativity. This note separates hallucination and slop as distinct problems.
  • generative-art-literature: A map of 94 academic studies on generative art. The reception evidence showing that AI labels lower ratings is also treated there.
  • design-value-dimensions: An account of design value across multiple dimensions. It underlies how proxy indicators for quality judgment are considered here.
  • llm-as-a-judge-literature: A genealogy of research on having LLMs perform evaluation. Connects to the question of whether verification cost can be handed back to machines.
  • vibe-coding-design-production: An industry-side observation that the bottleneck has shifted from “making” to “judging.”
  • adoption-approval-paradox: The gap between a 90% adoption rate and 10% approval. Sits in the same stratum as the disclosure penalty.
  • slop-countermeasures-history: A sister note tracing the cycle of low-quality output and its countermeasures from the printing press to 2026. It tests this note’s thesis from the side of history.
  • slop-design-history: A sister note revisiting the same question from within the discipline of design history. It traces how design as a profession was itself institutionalised as the first countermeasure against slop, and shows that this note’s definition has the advantage of bypassing judgments of taste.

References

Concept and Etymology

Creativity and Homogenization

Design Theory and Craft

  • Pye, D. (1968). The Nature and Art of Workmanship. Cambridge University Press. (reprint: Bloomsbury, 2022, ed. Ezra Shales, ISBN 9781350004528)
  • Schön, D. A. (1992). Designing as Reflective Conversation with the Materials of a Design Situation. Knowledge-Based Systems 5(1), 3–14. https://doi.org/10.1016/0950-7051(92)90020-G
  • Schön, D. A. (1983). The Reflective Practitioner: How Professionals Think in Action. Basic Books. ISBN 9780465068784
  • Sennett, R. (2008). The Craftsman. Yale University Press. ISBN 9780300119091
  • Ingold, T. (2013). Making: Anthropology, Archaeology, Art and Architecture. Routledge. ISBN 9780415567237
  • Nelson, H. G. & Stolterman, E. (2012). The Design Way: Intentional Change in an Unpredictable World, 2nd ed. MIT Press. ISBN 9780262526708
  • Cross, N. (2001). Designerly Ways of Knowing: Design Discipline Versus Design Science. Design Issues 17(3), 49–55. https://doi.org/10.1162/074793601750357196
  • Hernández-Ramírez, R. & Ferreira, J. B. (2024). The Future End of Design Work. She Ji 10(4), 414–440. https://doi.org/10.1016/j.sheji.2024.11.002
  • Kurosu, M. & Kashimura, K. (1995). Apparent Usability vs. Inherent Usability. CHI ‘95 Conference Companion. https://doi.org/10.1145/223355.223680
  • Tuch, A. N., Presslaber, E. E., Stöcklin, M., Opwis, K., & Bargas-Avila, J. A. (2012). The Role of Visual Complexity and Prototypicality Regarding First Impression of Websites. International Journal of Human-Computer Studies 70(11), 794–811. https://doi.org/10.1016/j.ijhcs.2012.06.003

Evaluation and Provenance

  • Raj, M., Berg, J. M., & Seamans, R. (2026). The Artificial Intelligence Disclosure Penalty: Humans Persistently Devalue AI-Generated Creative Writing. Journal of Experimental Psychology: General 155(4), 896–915. https://doi.org/10.1037/xge0001889
  • Porter, B. & Machery, E. (2024). AI-generated poetry is indistinguishable from human-written poetry and is rated more favorably. Scientific Reports 14, 26133. https://doi.org/10.1038/s41598-024-76900-1
  • Köbis, N. & Mossink, L. D. (2021). Artificial intelligence versus Maya Angelou. Computers in Human Behavior 114, 106553. https://doi.org/10.48550/arXiv.2005.09980 (the published version’s DOI has not been directly confirmed)
  • Ansani, A., Koehler, F., Giombini, L., Hämäläinen, M., Meng, C., Marini, M., & Saarikallio, S. (2025). AI Performer Bias. Empirical Studies of the Arts 43(2), 1137–1161. https://doi.org/10.1177/02762374241308807
  • Magni, F., Park, J., & Chao, M. M. (2024). Humans as Creativity Gatekeepers: Are We Biased Against AI Creativity? Journal of Business and Psychology 39(3), 643–656. https://doi.org/10.1007/s10869-023-09910-x
  • Kruger, J., Wirtz, D., Van Boven, L., & Altermatt, T. W. (2004). The Effort Heuristic. Journal of Experimental Social Psychology 40, 91–98. https://psycnet.apa.org/record/2003-11126-009 (DOI unconfirmed)
  • Heimstad, S. B., Wien, A., & Gaustad, T. (2025). Machine heuristic in algorithm aversion. Computers in Human Behavior: Artificial Humans 7, 100190. https://doi.org/10.1016/j.chbah.2025.100190
  • Newman, G. E. & Bloom, P. (2012). Art and authenticity: The importance of originals in judgments of value. Journal of Experimental Psychology: General 141(3), 558–569. https://doi.org/10.1037/a0026035
  • Arielli, E. (2024). Effort in Aesthetic Appreciation: from Avant-Garde to AI. Proceedings of the European Society for Aesthetics 16. https://www.eurosa.org/wp-content/uploads/Arielli.pdf
  • Hitsuwari, J., Ueda, Y., Yun, W., & Nomura, M. (2023). Does human–AI collaboration lead to more creative art? Computers in Human Behavior 139, 107502. https://doi.org/10.1016/j.chb.2022.107502

Information Environment and Cost Structure

Labor and Roles

Unverified Items

This collects items whose bibliography remains unconfirmed, in the body or the reference list, or whose full text was not reached. Unverified items for the full corpus (92 sources) appear at the end of source/review/ai-slop/papers.md.

  • The DOI for Hannigan et al. (2024) has not been directly confirmed on the publisher’s page (the bibliographic details themselves agree across multiple sources).
  • The published version’s DOI for Köbis and Mossink (2021) has not been directly confirmed.
  • The DOI number for Kruger et al. (2004) has not been confirmed through DOI resolution.
  • The author names for Heimstad et al. (2025) were confirmed only from search-result snippets.
  • The full text of Jang (2026) was not reached due to a paywall, and conditions other than the 34 participants have not been confirmed.
  • The DOI number for Tsao et al. (2025) is unconfirmed.
  • curl’s figures (about 20% of submissions, confirmation rate falling from over 15% to under 5%, zero in six years) are the maintainer’s own self-reported tally and have not undergone third-party verification.
  • Wu et al. (2026), Chakrabarty et al. (2026), Dolezal et al. (2026), Zhang and Zhang (2025), Erlei et al. (2026), and Hwang et al. (2026) are all pre-peer-review preprints or pre-proceedings versions. All volume estimates depend on detector accuracy.
  • The model names and count of LLMs compared by Wenger and Kenett (2026) have not been confirmed.
  • No prior literature explicitly stating the formulation “slop equals outsourced verification” was found within the scope of this search. The completeness of the search is not guaranteed.

← All Notes · Home