Notes — Research Log
Notes.
A log of research, experiments, and reflections. Ongoing investigations into AI applications and design processes, accumulated and published in an LLM Wiki format.
Want to ask about the notes? Ask the Wiki →
-
AI and Design Weekly Watch (2026-08-24 to 08-31)
2026-08-31An integrated summary of 'AI and design' developments over the past seven days, collected in three tiers: T1v vendor primary sources, T2 public institutions and standards, and T3 expert opinion. Cursor enabled starting cloud agents without an SCM connection, Figma turned its agent chat panel into a standalone window, and Simon Willison explained the architecture of OpenAI's ChatGPT Work. In the same week, security researcher Johann Rehberger reported a 60-80% success rate attacking Claude Code Opus 5's Auto Mode, raising questions about how this squares with the 0.00% figure from Anthropic's third-party evaluation. Stability AI raised $76 million from EA, Sony Music, Universal Music, and Warner Music, while Kaelig Deloumeau-Prigent's survey found that open-source design systems are rapidly adopting support for MCP and agent skills, even as Figma Code Connect support remains at only 10%.
Read → -
Was the 'Absence of a Science of Design' Asserted Three Times? Rereading the Observations of 1978, 2008, and 2023 in Their Sources
2026-08-31A rereading of the problem statement repeatedly invoked in design research in Japan, the 'absence of a science of design', against the three statements it rests on: Kimimasa Abe (1978), Shutaro Mukai, and Kazaru Yaegashi and colleagues (2023). The three are not saying the same thing. Abe asks about the standing of design as a discipline, whether it is no more than 'applied art'; Mukai judges that only the Ulm School of Design possessed a practical theory of designing, while leaving its reconstruction unfinished; Yaegashi and colleagues diagnose that the pursuit of scientization arrived at the paradox of 'the impossibility of scientization'. The question moves from standing to construction, and from construction to method. Then in 2025, Nigel Cross himself, on whom Yaegashi and colleagues rest, wrote that 'design is now recognized as an academic discipline', and declared in a co-authored paper that 'we have emerged'. As a problem statement, absence is no longer current as of 2026.
Read → -
The Genealogy of Design Research in Japan: Three Origins, and the University That Passed the Presidency Around
2026-08-31An integrative summary mapping design research in Japan across four layers: the founding genealogy of institutions (1876–2026), the branching of learned societies, theoretical schools, and the structure of faculty hiring. It draws on 94 sources: peer-reviewed papers and primary materials from universities, learned societies, and public agencies. The institutional forms of design education split into three origins (an engineering line, an aesthetics line, and a geijutsu kogaku line), each of which has its own learned society. The Japan Society for the Science of Design publishes no list of past presidents; reconstructing all 14 from the colophons of its journal shows that 6 of the 12 whose affiliation could be established were at Chiba University, and that a 1986 paper lists three future presidents as co-authors from the same laboratory. On faculty hiring, by contrast, the self-institution hiring rate for Japanese universities as a whole (32.6%, FY2013) is published, yet no figure specific to design exists anywhere. To fill that gap, the alma maters of faculty at three major institutions were tallied here: the self-institution rate ranges from 12.5% to 66.7% depending on the institution, and at the same time, at the institution with the highest rate, none of the faculty whose credentials could be confirmed hold a doctorate.
Read → -
Does Seeming Common Online Mean It Is Common?
2026-08-30Is "I see it a lot" a reliable cue for "a lot of people do it"? This note gathers 303 academic sources bearing on that question (29 on frequency cues and prevalence estimation, 26 on high-volume output by few actors, 27 on the generative-AI-specific information environment, 26 on self-fulfillment and norm dynamics, 27 on the representativeness of digital trace data, 117 from an additional collection covering the last three years, and 51 from a differential check against echo chamber research). The phenomenon of frequency and headcount coming apart has been measured since well before generative AI. Weaver et al. (2007), across 6 experiments, showed that people who heard one person repeat the same opinion three times estimated "this opinion is widely supported" almost as strongly as people who heard three different people each say it once, and the effect did not disappear even when participants were explicitly told the speaker was a single person. In Pew Research Center's (2019) probability-sample survey, the top 10% of U.S. adult Twitter users produced 80% of all tweets, and van Mierlo (2014) measured that 1.3% of users wrote 74.7% of posts across four health social networks. Rao and Reiley (2012), noting that the marginal cost of spam is close to zero, formalized the structure by which send volume decouples from the number of senders, and Lerman, Yan, and Wu (2016) formalized how skew in a network's degree distribution alone can make a globally rare attribute look like a majority under local observation. Over the last three years, the picture has shifted substantially. The path by which exposure moves the judgment of "what is normal" has now been directly demonstrated. In the third experiment of Glickman and Sharot (2025, *Nature Human Behaviour*, N=1,401), merely showing three Stable-Diffusion-generated images of a "financial manager," 1.5 seconds each, raised the share of participants choosing a white man as the most manager-like from 32.36% to 38.20% (p=0.04; the control group showed no significant difference). AlDahoul et al. (2025) showed that exposure to AI-generated faces moves perceptions of race and gender, and does so regardless of whether the images are disclosed as AI-generated. Brady et al. (2026, a *Nature* preregistered report) followed 2,000 people on Bluesky for eight weeks and confirmed, on a real platform, that an engagement-optimizing feed lowers the accuracy of social-norm perception while a diversifying algorithm improves it. Geber and Stahel (2026, N=1,021), however, reported that the type of recommendation had no significant effect on norm perception, and Liu et al. (2025, *PNAS*), across four YouTube experiments (roughly 9,000 participants total), could not detect a short-term polarization effect from manipulating filter-bubble conditions. Measurement of volume has advanced as well. Liang et al. (2024) estimated that 6.5% to 16.9% of the body text of peer reviews at AI conferences had been substantially altered by an LLM; Allaham and Diakopoulos (2026) measured that about 16% of generative search's cited sources are already AI-generated content; and Lee et al. (2025, *PNAS Nexus*) showed that estimates of the share of harmful posters overstate the measured value by up to roughly 100-fold. The detecting side has not kept up. A watermark-removal attack succeeded against 7 schemes at near-100% (Cheng et al. 2025, at a cost of $0.88 per million tokens); C2PA was concluded, under independent evaluation, not to meet its stated security goals (Golaszewski et al. 2026); and bot detection loses up to 29.6% performance to LLM-based evasion (Feng et al. 2024). Naming has arrived first as well: Schroeder et al. (2026) put forward the term "synthetic consensus" in a *Science* policy forum (unaccompanied by empirical evidence). The measuring instruments themselves have been contaminated too: Westwood (2025, *PNAS*) showed that an autonomous LLM agent can evade a survey's quality controls in 99.8% of 6,000 trials. What remains as a gap has narrowed to a single point. Even when Daikeler et al. (2025) systematically reviewed 58 data-quality frameworks, no error type existed for the case where the sender lies outside the population, and it was not even named as a gap. Alsalti et al. (2026), in a section on the coverage error of the total-survey-error framework, put this premise into words for the first time, but it has not reached formalization as a subtype. The differential check against echo chamber research, the closest neighboring concept, was also carried out, adding 51 sources. Within echo chamber research, the causal locus itself is contested (Sunstein places it in the receiver's own selection, Pariser in the algorithm, Cinelli et al. 2021 in homophily, and Bakshy et al. 2015 separate the two, finding that individual clicking accounts for a roughly 70% reduction that exceeds the algorithm's roughly 15% reduction), yet the premise that a real human being stands behind each observed item is never made explicit across the 24 sources reviewed — this follows from the structure of the methodology itself, since a node being human is the starting point. The accumulation of disconfirming evidence is substantial as well: Guess (2021) measured the overlap in partisan media diets at 50% to 65%, and four studies from Meta's 2020 election research showed that even substantially moving exposure left attitude measures unmoved (Nyhan et al. 2023, *Nature*; two papers by Guess et al. 2023, *Science*; González-Bailón et al. 2023, *Science*; at a scale of 208 million people). The lineage of misestimated composition ratios carries the same structure: Ahler and Sood (2018) showed that respondents estimate the share of LGBT people among Democrats at 32% (the true figure is 6%) and the share of high earners among Republicans at 38% (the true figure is 2%), yet in both cases the estimated target is a group that actually exists, and the remedy is to present the true base rate (in Ahler 2014, an intervention moderated opinion by 8 to 13 percentage points). The effect size of selective exposure itself remains modest, at d=0.36 (Hart et al.'s 2009 meta-analysis). The difference sorts into three points: where the distortion is located (in the observer's sample, or in whether the supplied item corresponds to a human); what the remedy presupposes (a real distribution of opinion, or the existence of a true ratio); and the direction of the error (toward one's own side, toward a salient minority, or independent of either). A rebuttal holds, however, if Guay et al.'s (2025, *PNAS*) account of "rescaling under uncertainty" is correct: since ratio estimation regresses toward 0.5 regardless of input, there may be little room left for contaminated input to make it worse. This rebuttal bears on the task of ratio estimation, but it does not directly touch the task Glickman and Sharot (2025) moved — judging "who is typical."
Read → -
Choosing Between Frontier and Cheaper Models: Claude Fable 5 and GPT-5.6 Sol
2026-08-29Organized around Claude Fable 5 and OpenAI GPT-5.6 Sol, this note sorts the practice of choosing between top-tier and cheaper models into three tiers of evidence: vendor first-party documentation, neutral data from public bodies, and the opinions of individually verifiable technical authorities. Both vendors' official guides now say the same thing, that tuning effort is often a better lever than switching models, and the unit of judgment has moved from price per request to cost per completed task. In Anthropic's own measurements, dropping effort to low on research work cost 1 to 3 points of accuracy for a third to a half of the price; on long-horizon coding, Claude Opus 5 at medium gave up about 2 points for half the cost and about 8 points at low for a quarter. Running everything at low and re-running only the failures at the default effort reached a higher pass rate than running everything at the default, for half the money. Mixing an expensive model with cheap ones pays only when there is bulk to hand off that no single context window could hold (55% below the frontier model alone); when the work is one dependent chain, the coordinating model alone at lower effort won in every case measured.
Read → -
Slop Seen Through Design History: The Profession Began as a Countermeasure to Shoddy Mass Production
2026-08-28This note reconsiders AI slop through the discipline of design history, drawing on 116 sources (24 on nineteenth-century British design reform, 24 on the debate over machines and authenticity, 21 on kitsch and the politics of taste, 24 on the methodology of design historiography, and 23 on planned obsolescence and homogenization). The core finding is that the design profession itself was institutionalized as the first countermeasure against mass-produced shoddiness. The British parliamentary Select Committee of 1835–36 deliberated an industrial crisis, that British manufactured goods' designs were inferior to foreign ones, and the Government School of Design was established in 1837. In 1852, Henry Cole and Richard Redgrave opened an exhibition, known as the Chamber of Horrors, that named and displayed 87 products they judged to be in bad taste, but it closed after two weeks under pressure from the manufacturers named in it (Suga 2004). The effectiveness of the education was also doubtful: the Stourbridge School of Art failed to connect with local industry (Measell 2020). The debate over machines and authenticity repeated in the same form from Ruskin through Loos, the Werkbund's Typenstreit, Benjamin, and Gute Form, yet Morris & Co. actually used both machines and handwork together (Harvey & Press 1991), and Pye (1968) argued that the distinction between handwork and machine work is itself technically meaningless. Design history has further shown that judgments of bad taste also functioned as instruments of class and gender (Greenberg 1939, Bourdieu 1979, Buckley 1986, Sparke 1995, Attfield 2000, Venturi et al. 1972). But the critique of taste is not simple either: a study visiting 160 households concluded that neither Veblen, nor the Frankfurt School, nor Bourdieu applies in any straightforward way (Halle 1993). What is decisive is that homogenization was already being measured before AI. Web design became significantly more similar after 2007, and the average distance between page layouts shrank by more than 30% (Goree et al., CHI 2021). The factors were shared libraries, standardized color palettes, and mobile responsiveness, not generative AI. From this, what remains as genuinely new this time is threefold: the way of talking about something as looking AI-made and tacky can repeat the 1852 exhibition, the sister note's definition (externalizing verification) has the advantage of bypassing taste judgments, and designers have moved from the side of countermeasures to the side of production.
Read → -
The History of Slop and Countermeasures: Why the Same Wager Keeps Failing
2026-08-28This note traces the cycle in which low-quality generated content pollutes the environment and countermeasures for selection are built in response, across 140 sources (21 on prehistory, 28 on email spam, 25 on search and bots, 29 on academic publishing, 37 on AI countermeasures since 2022) spanning the era after print to 2026. Historically, countermeasures fall into roughly five types: the recipient classifies (Bayesian filters), cost is imposed on the sender (proof-of-work), provenance is certified (SPF/DKIM/DMARC, C2PA, watermarking), personhood is proven (CAPTCHA), and a gatekeeper is placed (peer review, blacklists, community norms). And there are five types in how they break down. (1) A design that presumes the attacker lacks a certain capability becomes invalid once that capability improves (CAPTCHA was a design that bet its security on AI's unsolved problems, and a GAN solver broke 33 schemes in 0.05 seconds). (2) A countermeasure that imposes cost does not work against an opponent who can pass that cost onto someone else (Laurie & Clayton 2004 demonstrated the failure of proof-of-work). (3) When provenance is voluntary, only honest participants use it (in Durumeric 2015, DMARC policy specification stood at 1.1%). (4) Gatekeeping errors concentrate on the weaker side (blacklists misclassifying Global South journals, discrimination against Tor users, false positives against non-native speakers). (5) Countermeasures that worked concentrated verification capability among large actors. Three things are new this time only: generated content and the genuine article cannot be told apart, the impossibility of strong watermarking has been proven (Zhang et al. ICML 2024), and no peer-reviewed research within the scope of this collection was found showing that any countermeasure actually reduced the volume or circulation of slop. What the countermeasures that worked historically shared was not "telling apart" but "changing the structure."
Read → -
AI Slop: Reading It as Outsourced Verification, Not Low Quality
2026-08-28This note examines "AI slop" from the perspective of design and creativity through 92 academic sources (22 on conceptual history, 24 on empirical creativity research, 27 on design theory, 19 on industry and labor). The word's originating definition (Willison 2024) placed its core not in poor quality but in "foisting something on others that you haven't verified yourself," which makes slop not a property of the object but a distribution of the burden of verification. Overlaying this reading on Pye's (1968) distinction between the "workmanship of risk" and the "workmanship of certainty" situates generative output as a third mode: not workmanship of certainty, since the outcome is not fixed in advance, and not workmanship of risk either, since judgments made during production do not determine the outcome. It is a mode of work in which risk is passed downstream. Three lines of evidence support this reading. (1) Homogenization: individual work is rated more highly even as collective novelty falls (Doshi & Hauser 2024); the novelty added by one human essay is 2 to 8 times that added by one GPT-4 essay (Moon et al. 2025, N=2,200); and in Midjourney, unrelated prompts converge on the same default images (Simonen et al. 2026, 750,000 images). (2) Cost asymmetry: only the marginal cost of low-quality generation falls, while verification cost does not (Zhang & Zhang 2025). AI music as a share of new Spotify releases rose from under 1% to over 40% (Wu et al. 2026), and curl's confirmation rate fell from over 15% to under 5%, leading it to suspend its bug bounty (Stenberg 2026). (3) Provenance-dependent evaluation: people cannot distinguish AI-generated work (46.6% discrimination accuracy, Porter & Machery 2024), yet consistently rate it lower once told it is AI-generated (16 experiments, N=27,491, Raj et al. 2026). The note specifies three conditions under which this reading fails, and limits its scope to include the fact that no measurement standard for slop yet exists (Shaib et al. 2025).
Read → -
From This Is Heavy to That Person Is Incompetent
2026-08-24A research design that examines the claim that people who raise others' cognitive load are incompetent, not by measuring it but by describing how the judgment is made. The prior literature review found that raters split into three groups whose division is explained neither by organization nor by demographics. With no prospect that measurement will settle the matter, the question moves from whether the claim is true to how the judgment comes about. Cognitive load is not scored on any instrument, self-regulation is not defined as a variable, and records are not coded into categories. Blumer's sensitizing concepts govern how concepts are held, Sacks's membership categorization governs how competence talk is treated, and Katz's argument that description carries inference governs where analysis sits. What the study goes to see is the moment when the experience that work is heavy turns into a judgment that a person is incompetent.
Read → -
Can Incompetence Be Defined as Raising Other People's Cognitive Load?
2026-08-24A test of the claim that incompetence in modern knowledge work means raising other people's cognitive load, and that raw processing ability barely matters, checked against 58 academic sources and against industry practice. The claim splits into three parts with different logical characters: a definition, a causal pathway, and an exclusion. The definition cannot be falsified, no empirical study within the search range measures the causal pathway directly, and only the exclusion clause is testable. On that clause, the validity of cognitive ability tests for job performance was revised down from .51 to .31 and fell behind structured interviews and job knowledge tests, while 504 raters split into three clusters (task-weighted, counterproductive-weighted, and both equally) that are explained neither by organization nor by demographics. From cognitive load theory, load is a relational quantity arising from the pairing of a sender's presentation with a receiver's prior knowledge, which leaves no basis for attributing it to an individual.
Read →