Notes
Research, experiments, and reflections on AI and design, published along with the process of making them.
Published notes: 151. Built as an LLM wiki. How it works → Ask the notes → See the citation network (how notes share sources) →
August 2026
-
The Genealogy of Design Research in Japan: Three Origins, and the University That Passed the Presidency Around
2026-08-31
An integrative summary mapping design research in Japan across four layers: the founding genealogy of institutions (1876–2026), the branching of learned societies, theoretical schools, and the structure of faculty hiring. It draws on 94 sources: peer-reviewed papers and primary materials from universities, learned societies, and public agencies. The institutional forms of design education split into three origins (an engineering line, an aesthetics line, and a geijutsu kogaku line), each of which has its own learned society. The Japan Society for the Science of Design publishes no list of past presidents; reconstructing all 14 from the colophons of its journal shows that 6 of the 12 whose affiliation could be established were at Chiba University, and that a 1986 paper lists three future presidents as co-authors from the same laboratory. On faculty hiring, by contrast, the self-institution hiring rate for Japanese universities as a whole (32.6%, FY2013) is published, yet no figure specific to design exists anywhere. To fill that gap, the alma maters of faculty at three major institutions were tallied here: the self-institution rate ranges from 12.5% to 66.7% depending on the institution, and at the same time, at the institution with the highest rate, none of the faculty whose credentials could be confirmed hold a doctorate.
-
AI and Design Weekly Watch (2026-08-24 to 08-31)
2026-08-31
An integrated summary of 'AI and design' developments over the past seven days, collected in three tiers: T1v vendor primary sources, T2 public institutions and standards, and T3 expert opinion. Cursor enabled starting cloud agents without an SCM connection, Figma turned its agent chat panel into a standalone window, and Simon Willison explained the architecture of OpenAI's ChatGPT Work. In the same week, security researcher Johann Rehberger reported a 60-80% success rate attacking Claude Code Opus 5's Auto Mode, raising questions about how this squares with the 0.00% figure from Anthropic's third-party evaluation. Stability AI raised $76 million from EA, Sony Music, Universal Music, and Warner Music, while Kaelig Deloumeau-Prigent's survey found that open-source design systems are rapidly adopting support for MCP and agent skills, even as Figma Code Connect support remains at only 10%.
-
Was the 'Absence of a Science of Design' Asserted Three Times? Rereading the Observations of 1978, 2008, and 2023 in Their Sources
2026-08-31
A rereading of the problem statement repeatedly invoked in design research in Japan, the 'absence of a science of design', against the three statements it rests on: Kimimasa Abe (1978), Shutaro Mukai, and Kazaru Yaegashi and colleagues (2023). The three are not saying the same thing. Abe asks about the standing of design as a discipline, whether it is no more than 'applied art'; Mukai judges that only the Ulm School of Design possessed a practical theory of designing, while leaving its reconstruction unfinished; Yaegashi and colleagues diagnose that the pursuit of scientization arrived at the paradox of 'the impossibility of scientization'. The question moves from standing to construction, and from construction to method. Then in 2025, Nigel Cross himself, on whom Yaegashi and colleagues rest, wrote that 'design is now recognized as an academic discipline', and declared in a co-authored paper that 'we have emerged'. As a problem statement, absence is no longer current as of 2026.
-
Does Seeming Common Online Mean It Is Common?
2026-08-30
Is "I see it a lot" a reliable cue for "a lot of people do it"? This note gathers 303 academic sources bearing on that question (29 on frequency cues and prevalence estimation, 26 on high-volume output by few actors, 27 on the generative-AI-specific information environment, 26 on self-fulfillment and norm dynamics, 27 on the representativeness of digital trace data, 117 from an additional collection covering the last three years, and 51 from a differential check against echo chamber research). The phenomenon of frequency and headcount coming apart has been measured since well before generative AI. Weaver et al. (2007), across 6 experiments, showed that people who heard one person repeat the same opinion three times estimated "this opinion is widely supported" almost as strongly as people who heard three different people each say it once, and the effect did not disappear even when participants were explicitly told the speaker was a single person. In Pew Research Center's (2019) probability-sample survey, the top 10% of U.S. adult Twitter users produced 80% of all tweets, and van Mierlo (2014) measured that 1.3% of users wrote 74.7% of posts across four health social networks. Rao and Reiley (2012), noting that the marginal cost of spam is close to zero, formalized the structure by which send volume decouples from the number of senders, and Lerman, Yan, and Wu (2016) formalized how skew in a network's degree distribution alone can make a globally rare attribute look like a majority under local observation. Over the last three years, the picture has shifted substantially. The path by which exposure moves the judgment of "what is normal" has now been directly demonstrated. In the third experiment of Glickman and Sharot (2025, *Nature Human Behaviour*, N=1,401), merely showing three Stable-Diffusion-generated images of a "financial manager," 1.5 seconds each, raised the share of participants choosing a white man as the most manager-like from 32.36% to 38.20% (p=0.04; the control group showed no significant difference). AlDahoul et al. (2025) showed that exposure to AI-generated faces moves perceptions of race and gender, and does so regardless of whether the images are disclosed as AI-generated. Brady et al. (2026, a *Nature* preregistered report) followed 2,000 people on Bluesky for eight weeks and confirmed, on a real platform, that an engagement-optimizing feed lowers the accuracy of social-norm perception while a diversifying algorithm improves it. Geber and Stahel (2026, N=1,021), however, reported that the type of recommendation had no significant effect on norm perception, and Liu et al. (2025, *PNAS*), across four YouTube experiments (roughly 9,000 participants total), could not detect a short-term polarization effect from manipulating filter-bubble conditions. Measurement of volume has advanced as well. Liang et al. (2024) estimated that 6.5% to 16.9% of the body text of peer reviews at AI conferences had been substantially altered by an LLM; Allaham and Diakopoulos (2026) measured that about 16% of generative search's cited sources are already AI-generated content; and Lee et al. (2025, *PNAS Nexus*) showed that estimates of the share of harmful posters overstate the measured value by up to roughly 100-fold. The detecting side has not kept up. A watermark-removal attack succeeded against 7 schemes at near-100% (Cheng et al. 2025, at a cost of $0.88 per million tokens); C2PA was concluded, under independent evaluation, not to meet its stated security goals (Golaszewski et al. 2026); and bot detection loses up to 29.6% performance to LLM-based evasion (Feng et al. 2024). Naming has arrived first as well: Schroeder et al. (2026) put forward the term "synthetic consensus" in a *Science* policy forum (unaccompanied by empirical evidence). The measuring instruments themselves have been contaminated too: Westwood (2025, *PNAS*) showed that an autonomous LLM agent can evade a survey's quality controls in 99.8% of 6,000 trials. What remains as a gap has narrowed to a single point. Even when Daikeler et al. (2025) systematically reviewed 58 data-quality frameworks, no error type existed for the case where the sender lies outside the population, and it was not even named as a gap. Alsalti et al. (2026), in a section on the coverage error of the total-survey-error framework, put this premise into words for the first time, but it has not reached formalization as a subtype. The differential check against echo chamber research, the closest neighboring concept, was also carried out, adding 51 sources. Within echo chamber research, the causal locus itself is contested (Sunstein places it in the receiver's own selection, Pariser in the algorithm, Cinelli et al. 2021 in homophily, and Bakshy et al. 2015 separate the two, finding that individual clicking accounts for a roughly 70% reduction that exceeds the algorithm's roughly 15% reduction), yet the premise that a real human being stands behind each observed item is never made explicit across the 24 sources reviewed — this follows from the structure of the methodology itself, since a node being human is the starting point. The accumulation of disconfirming evidence is substantial as well: Guess (2021) measured the overlap in partisan media diets at 50% to 65%, and four studies from Meta's 2020 election research showed that even substantially moving exposure left attitude measures unmoved (Nyhan et al. 2023, *Nature*; two papers by Guess et al. 2023, *Science*; González-Bailón et al. 2023, *Science*; at a scale of 208 million people). The lineage of misestimated composition ratios carries the same structure: Ahler and Sood (2018) showed that respondents estimate the share of LGBT people among Democrats at 32% (the true figure is 6%) and the share of high earners among Republicans at 38% (the true figure is 2%), yet in both cases the estimated target is a group that actually exists, and the remedy is to present the true base rate (in Ahler 2014, an intervention moderated opinion by 8 to 13 percentage points). The effect size of selective exposure itself remains modest, at d=0.36 (Hart et al.'s 2009 meta-analysis). The difference sorts into three points: where the distortion is located (in the observer's sample, or in whether the supplied item corresponds to a human); what the remedy presupposes (a real distribution of opinion, or the existence of a true ratio); and the direction of the error (toward one's own side, toward a salient minority, or independent of either). A rebuttal holds, however, if Guay et al.'s (2025, *PNAS*) account of "rescaling under uncertainty" is correct: since ratio estimation regresses toward 0.5 regardless of input, there may be little room left for contaminated input to make it worse. This rebuttal bears on the task of ratio estimation, but it does not directly touch the task Glickman and Sharot (2025) moved — judging "who is typical."
-
Choosing Between Frontier and Cheaper Models: Claude Fable 5 and GPT-5.6 Sol
2026-08-29
Organized around Claude Fable 5 and OpenAI GPT-5.6 Sol, this note sorts the practice of choosing between top-tier and cheaper models into three tiers of evidence: vendor first-party documentation, neutral data from public bodies, and the opinions of individually verifiable technical authorities. Both vendors' official guides now say the same thing, that tuning effort is often a better lever than switching models, and the unit of judgment has moved from price per request to cost per completed task. In Anthropic's own measurements, dropping effort to low on research work cost 1 to 3 points of accuracy for a third to a half of the price; on long-horizon coding, Claude Opus 5 at medium gave up about 2 points for half the cost and about 8 points at low for a quarter. Running everything at low and re-running only the failures at the default effort reached a higher pass rate than running everything at the default, for half the money. Mixing an expensive model with cheap ones pays only when there is bulk to hand off that no single context window could hold (55% below the frontier model alone); when the work is one dependent chain, the coordinating model alone at lower effort won in every case measured.
-
Slop Seen Through Design History: The Profession Began as a Countermeasure to Shoddy Mass Production
2026-08-28
This note reconsiders AI slop through the discipline of design history, drawing on 116 sources (24 on nineteenth-century British design reform, 24 on the debate over machines and authenticity, 21 on kitsch and the politics of taste, 24 on the methodology of design historiography, and 23 on planned obsolescence and homogenization). The core finding is that the design profession itself was institutionalized as the first countermeasure against mass-produced shoddiness. The British parliamentary Select Committee of 1835–36 deliberated an industrial crisis, that British manufactured goods' designs were inferior to foreign ones, and the Government School of Design was established in 1837. In 1852, Henry Cole and Richard Redgrave opened an exhibition, known as the Chamber of Horrors, that named and displayed 87 products they judged to be in bad taste, but it closed after two weeks under pressure from the manufacturers named in it (Suga 2004). The effectiveness of the education was also doubtful: the Stourbridge School of Art failed to connect with local industry (Measell 2020). The debate over machines and authenticity repeated in the same form from Ruskin through Loos, the Werkbund's Typenstreit, Benjamin, and Gute Form, yet Morris & Co. actually used both machines and handwork together (Harvey & Press 1991), and Pye (1968) argued that the distinction between handwork and machine work is itself technically meaningless. Design history has further shown that judgments of bad taste also functioned as instruments of class and gender (Greenberg 1939, Bourdieu 1979, Buckley 1986, Sparke 1995, Attfield 2000, Venturi et al. 1972). But the critique of taste is not simple either: a study visiting 160 households concluded that neither Veblen, nor the Frankfurt School, nor Bourdieu applies in any straightforward way (Halle 1993). What is decisive is that homogenization was already being measured before AI. Web design became significantly more similar after 2007, and the average distance between page layouts shrank by more than 30% (Goree et al., CHI 2021). The factors were shared libraries, standardized color palettes, and mobile responsiveness, not generative AI. From this, what remains as genuinely new this time is threefold: the way of talking about something as looking AI-made and tacky can repeat the 1852 exhibition, the sister note's definition (externalizing verification) has the advantage of bypassing taste judgments, and designers have moved from the side of countermeasures to the side of production.
-
The History of Slop and Countermeasures: Why the Same Wager Keeps Failing
2026-08-28
This note traces the cycle in which low-quality generated content pollutes the environment and countermeasures for selection are built in response, across 140 sources (21 on prehistory, 28 on email spam, 25 on search and bots, 29 on academic publishing, 37 on AI countermeasures since 2022) spanning the era after print to 2026. Historically, countermeasures fall into roughly five types: the recipient classifies (Bayesian filters), cost is imposed on the sender (proof-of-work), provenance is certified (SPF/DKIM/DMARC, C2PA, watermarking), personhood is proven (CAPTCHA), and a gatekeeper is placed (peer review, blacklists, community norms). And there are five types in how they break down. (1) A design that presumes the attacker lacks a certain capability becomes invalid once that capability improves (CAPTCHA was a design that bet its security on AI's unsolved problems, and a GAN solver broke 33 schemes in 0.05 seconds). (2) A countermeasure that imposes cost does not work against an opponent who can pass that cost onto someone else (Laurie & Clayton 2004 demonstrated the failure of proof-of-work). (3) When provenance is voluntary, only honest participants use it (in Durumeric 2015, DMARC policy specification stood at 1.1%). (4) Gatekeeping errors concentrate on the weaker side (blacklists misclassifying Global South journals, discrimination against Tor users, false positives against non-native speakers). (5) Countermeasures that worked concentrated verification capability among large actors. Three things are new this time only: generated content and the genuine article cannot be told apart, the impossibility of strong watermarking has been proven (Zhang et al. ICML 2024), and no peer-reviewed research within the scope of this collection was found showing that any countermeasure actually reduced the volume or circulation of slop. What the countermeasures that worked historically shared was not "telling apart" but "changing the structure."
-
AI Slop: Reading It as Outsourced Verification, Not Low Quality
2026-08-28
This note examines "AI slop" from the perspective of design and creativity through 92 academic sources (22 on conceptual history, 24 on empirical creativity research, 27 on design theory, 19 on industry and labor). The word's originating definition (Willison 2024) placed its core not in poor quality but in "foisting something on others that you haven't verified yourself," which makes slop not a property of the object but a distribution of the burden of verification. Overlaying this reading on Pye's (1968) distinction between the "workmanship of risk" and the "workmanship of certainty" situates generative output as a third mode: not workmanship of certainty, since the outcome is not fixed in advance, and not workmanship of risk either, since judgments made during production do not determine the outcome. It is a mode of work in which risk is passed downstream. Three lines of evidence support this reading. (1) Homogenization: individual work is rated more highly even as collective novelty falls (Doshi & Hauser 2024); the novelty added by one human essay is 2 to 8 times that added by one GPT-4 essay (Moon et al. 2025, N=2,200); and in Midjourney, unrelated prompts converge on the same default images (Simonen et al. 2026, 750,000 images). (2) Cost asymmetry: only the marginal cost of low-quality generation falls, while verification cost does not (Zhang & Zhang 2025). AI music as a share of new Spotify releases rose from under 1% to over 40% (Wu et al. 2026), and curl's confirmation rate fell from over 15% to under 5%, leading it to suspend its bug bounty (Stenberg 2026). (3) Provenance-dependent evaluation: people cannot distinguish AI-generated work (46.6% discrimination accuracy, Porter & Machery 2024), yet consistently rate it lower once told it is AI-generated (16 experiments, N=27,491, Raj et al. 2026). The note specifies three conditions under which this reading fails, and limits its scope to include the fact that no measurement standard for slop yet exists (Shaib et al. 2025).
-
From This Is Heavy to That Person Is Incompetent
2026-08-24
A research design that examines the claim that people who raise others' cognitive load are incompetent, not by measuring it but by describing how the judgment is made. The prior literature review found that raters split into three groups whose division is explained neither by organization nor by demographics. With no prospect that measurement will settle the matter, the question moves from whether the claim is true to how the judgment comes about. Cognitive load is not scored on any instrument, self-regulation is not defined as a variable, and records are not coded into categories. Blumer's sensitizing concepts govern how concepts are held, Sacks's membership categorization governs how competence talk is treated, and Katz's argument that description carries inference governs where analysis sits. What the study goes to see is the moment when the experience that work is heavy turns into a judgment that a person is incompetent.
-
Can Incompetence Be Defined as Raising Other People's Cognitive Load?
2026-08-24
A test of the claim that incompetence in modern knowledge work means raising other people's cognitive load, and that raw processing ability barely matters, checked against 58 academic sources and against industry practice. The claim splits into three parts with different logical characters: a definition, a causal pathway, and an exclusion. The definition cannot be falsified, no empirical study within the search range measures the causal pathway directly, and only the exclusion clause is testable. On that clause, the validity of cognitive ability tests for job performance was revised down from .51 to .31 and fell behind structured interviews and job knowledge tests, while 504 raters split into three clusters (task-weighted, counterproductive-weighted, and both equally) that are explained neither by organization nor by demographics. From cognitive load theory, load is a relational quantity arising from the pairing of a sender's presentation with a receiver's prior knowledge, which leaves no basis for attributing it to an individual.
-
AI and Design Weekly Watch (2026-08-18 to 08-24)
2026-08-24
An integrated summary of 'AI and design' developments over the past seven days, collected in three tiers: T1v vendor primary sources, T2 public institutions and research, and T3 expert opinion. Anthropic made browser use and the Skills API generally available, Vercel connected v0-built apps to more than a hundred external services, and Cursor shipped always-on agents that wake on PR and Slack subscriptions. As the reach of agents acting on a user's behalf widened in three directions in a single week, Nielsen surfaced research on the 'design theater' gap between what generative UI tools explain and what they implement, Willison argued that verification is the core skill, and design leaders quoted in Figma's newsletter raised checkpoints and accountability as the open questions. The regulatory side was quiet: the only in-window official event was a change of the UK minister responsible for IP.
-
Novice Group Work with a Shared AI Agent: A Literature Map of CSCL and AI-Supported Collaborative Learning (2026)
2026-08-22
Best practices for introducing a single shared AI agent into group work, organized from 48 academic publications spanning CSCL foundations, the ACLS (adaptive collaborative learning support) lineage, SSRL x AI, and LLM-era empirical studies: intervention timing, addressee design, AI-centered interaction drift and its condition-dependence, and mitigation of over-reliance and free-riding.
-
AI and Design Weekly Watch (2026-08-10 to 08-17)
2026-08-17
An integrated summary of 'AI and design' developments over the past seven days, collected in three tiers: T1v vendor primary sources, T2 public institutions and research, and T3 expert opinion. Figma's Weave tool became directly callable by external agents over MCP, and Lovable added an MCP registry, both repositioning design tools from standalone apps into components other agents can invoke. The same week, Anthropic revealed that Claude Cowork's browser integration inserts a separate verification step before 'consequential actions,' and Nielsen surfaced research on the 'delegation regret' users feel toward agents that act without a preview. The EU AI Act's code of practice finally disclosed the breakdown of its 82 Section 1 signatories, but left a new discrepancy with the total count unresolved.
-
When Is a Disposition Being Measured? Measurement as Conditional Tendency, and Measurement as Co-occurrence
2026-08-15
A literature map of the measurement paradigm that treats a disposition not as a fixed possession but as a conditional tendency of person by situation. It places Mischel & Shoda's if-then behavioral signatures, Fleeson's density distributions of states, Steyer's latent state-trait theory, experience-sampling methods, dynamic SEM and the RI-CLPM, Molenaar's critique on ergodicity, the DIAMONDS taxonomy of situations, and situational judgment and conditional reasoning tests. Alongside these it sets out the structural limits of self-report Likert scales (the reference-group effect, cultural differences in response style, causal validity), the measurement controversies over mindset and grit, emic/etic in indigenous psychology, inference from behavioral traces, and the analysis of co-occurrence by epistemic network analysis. Two of the descriptions that served as the starting point could not be traced to their attributed sources. That the "2025 revised latent state-trait theory" does not exist and that the only Revised version confirmed is from 2015, and that no review matching "a recent review positions ENA as a move away from code and counting" could be identified while the framing in question appears in a 2018 empirical comparison study, are both stated in the text. 76 references.
-
Why Heisei Retro Caught On: Building and Testing a Scale of Aspiration toward Gyaru-Mind
2026-08-15
An account, drawn from the full 36-page text, of Motode Rino's "An Examination of the Effects of the 'Heisei Gal' Mindset on Contemporary Women" (Bulletin of the School of Sociology, Kwansei Gakuin University 141, 2023). A content analysis of ten gyaru magazines defines gyaru-mind in four factors; a scale is built on 594 respondents and hypotheses tested on 893 women aged 15 to 59. Alphas run from .703 to .904 across a combined sample of 1,487. Two of the four factors came out against expectation in the criterion-related validity check, and one of those was rebuilt and re-tested. Hypothesis 1 — that nostalgia proneness mediates the appeal of Shōwa retro — was not supported; Hypothesis 2 was partly supported, with dissatisfaction under COVID-19 shown to raise the appeal of Heisei retro through aspiration toward gyaru-mind. The paper also sets aside an earlier 2011 scale from Dentsu's Gal Labo after a critical review. With references.
-
How Might Gyaru-Mind Be Measured: What the Two Existing Scales Leave Out
2026-08-15
An examination, from the side of measurement theory, of what the two prior studies that measured gyaru-mind (Motode Rino's four factors and forty-four items, 2023; Ikegami Momoka and colleagues' eight-factor GYARU-MIDX, 2025–2026) are in fact measuring, together with a proposed design for measuring it as a conditional tendency. All forty-four of Motode's items take the form 'I admire...' or 'I want to...', so the scale measures aspiration, exactly as its name says. Ikegami and colleagues measure the style of speech in a single fifteen-minute conversation. Since neither carries a description of the situation, neither yields an answer to 'how does this person respond, and in which situation'. The reference-group effect bears on the generational comparison, and the if-then signature bears on the interpretation of the two factors that came out against expectation. On that basis, a five-stage measurement design that begins by collecting situations from the people concerned is proposed, along with four cheap tests worth running first. Measured against the procedure for turning an indigenous concept into a scale (Church and Katigbak's 1988 four stages), both scales stop at the second stage and have not undergone a discriminant-validity test against established scales. The latter half of this note is a proposed design, not a verified finding, and says so.
-
Difference Does Not Become a Criterion: An Adversarial Review of a Novelty Claim about Gyaru Practice
2026-08-15
A record of testing, against competing prior work, the claim that gyaru practice is theoretically novel in that "a difference that was an object evaluated by existing criteria turns instead into the source of new criteria of evaluation." Thornton's subcultural capital, Fraser's subaltern counterpublics, Collins's self-definition and self-valuation, Agha's enregisterment, Lamont's valuation studies, Costanza-Chock's design justice, and Dorst's frame creation mean that every part of the claim already has strong prior work behind it. What remains is only the line that distinguishes revaluation of difference from criterion formation, and demonstrates the recursive sequence running from amplification to evaluative terms, rules of judgement, and feedback into making. Assessing the causal chain stage by stage, difference through sharing can be supported, while criterion formation and the updating of making are not. The contemporary case the material offered as grounds for criterion formation (SHIBUYA109 lab. × CGO dot com's 'Gyaru-Style Self-Analysis') could not be traced at first pass and was judged a conflation of two entities, but a renewed search on 2026-08-16 confirmed that it **does exist** (a correction to the initial judgement). The vocabulary was prepared by the organizers, however, so it is not evidence of criterion formation by the people themselves, and criterion formation remains unproven. Seven alternative explanations are set out, including the reverse causation suggested by Matsui Takeshi's research on social signs (that the namer may have been the media), together with the falsification conditions that ought to be fixed in advance. No study that separates category terms, evaluative terms, and metapragmatic terms and tracks the first appearance and users of each word was found across five search routes. 36 references.
-
"Making It Your Own" Is Not Enough: The Ten Dimensions of Tojisha-sei, and How to Measure Them
2026-08-15
A literature map setting the Japanese *tojisha-sei* against the English-language traditions of agency, autonomy, self-efficacy, sense of agency, psychological ownership, empowerment, political efficacy, participation, lay expertise, and co-production. Tojisha-sei corresponds to none of these singly; it has to be handled across ten dimensions, from affectedness to stewardship. Nakanishi Shoji and Ueno Chizuko's tojisha shuken, the theory of needs introduced by Doba Manabu, and the tojisha-kenkyu of Bethel's House in Urakawa are set alongside Bandura, Emirbayer & Mische, Haraway, Epstein, Arnstein, and Scandinavian Participatory Design. On measurement, no standard scale exists, and measurement has to run across four layers — subjective, behavioural, outcome, and structural authority — set out here together with the psychometrics of existing scales. The argument that the further LLMs lower the cost of implementation, the further the bottleneck moves toward problem framing, verification, and governance was checked against two empirical studies from 2026. Checking the source report's descriptions against primary APIs found six places that did not agree with the originals, and identified figures that could not be verified; both are stated in the body. That the three studies measuring gyaru-mind place the people concerned nowhere is also named through this framework. 26 references.
-
Re-searching the Gaps in Gyaru Research: 71 Additions and the One Absence That Survived a Change of Route
2026-08-15
An integrative summary of a 71-source supplementary corpus, produced by re-searching along seven different routes the four areas that the previous literature map (61 sources) reported as 'absent within the search range'. It shows that two of those areas — ganguro/yamanba and gyaru-o/gyaru circles — were never absent but simply missed by the earlier search route; that Reiwa gyaru appears not as a topic but as a comparison point in other fields; that Korean-language scholarship forms an independent line of research; and that only 'critical discourse analysis of the gyaru mind' still returns zero results after five changes of route. 71 references at the end.
-
End-User Development (EUC/EUD): Genealogy and Limits — A Map of 152 Sources
2026-08-15
An integrative summary mapping the history, social demands, and limits of end-user computing (EUC) and end-user development (EUD) across 152 sources, from Licklider in 1960 to the LLM era of 2026. It shows how EUC, born as a framework of managerial control, and EUD, born as a framework of emancipation, have described the same phenomenon in opposite directions; how four distinct social demands were folded into the single word 'democratization'; how the limits take a different shape in accounting, medicine, and scientific computing; and how the limits the field documented for itself are being replayed in the same form in the LLM era.
-
How Gyaru Have Been Studied: A Map of 61 Works and the Gap Around "Gyaru Mind"
2026-08-15
A synthesis of 61 scholarly works on Japanese gyaru culture (kogal, ganguro, gyaru-o, gyaru circles, Reiwa gyaru) and its associated mindset, mapped across ten clusters. Thirty years of scholarship addressed fashion, language, the body, and media representation, yet few works take the ethos itself as their object, and most of those stand almost disconnected from the humanities and social sciences. Peer-reviewed research on the 2020s revival, and any critical analysis of the self-affirmation discourse, remain absent. Corrected 2026-08-15: the first version held that only four works since 2022, all from design studies and HCI, had taken the ethos as their object; this was wrong, as Motode Rino (2023, Kwansei Gakuin University) preceded them by two years with a four-factor scale validated on 1,487 respondents. 61 references.
-
What Has Generative AI Changed About Making for Oneself? A Map of 59 Sources
2026-08-14
An integrated summary mapping, through 59 sources, how generative AI (2023–2026) has changed making that starts from one's own need and serves oneself. It fills the one intersection absent from the two sister corpora (337 works on EUD history and learning outcomes): making-for-oneself × generative AI. Evidence that the entry barrier to making has dropped is accumulating across four areas (personal software, self-data, self-made assistive technology, hobbyist creation), but 32 of the 59 sources are unreviewed preprints and measurement is skewed toward immediate evaluation. Of the four gaps identified by the prior reviews, attachment to self-made artifacts has begun to move via IKEA-effect RCTs on AI co-created text, while delayed-test longitudinal tracking, tracking of private personal tools, and motive comparison remain unfilled in the generative AI era.
-
Does Being Able to Build Mean Being Able to Learn? Rereading the Learning Outcomes of End-User Development Through Motive
2026-08-14
An integrative summary that rereads end-user development (EUD) along two axes — who builds and why, and what is learned there — and maps it with 185 sources. It sets out that the users EUD research assumes are 'non-programmers who are nonetheless domain experts', a different kind of learner from the novice; that far transfer from building has not been supported since 1984; that rises in self-efficacy do not correspond to actual competence; that studies measuring learning outcomes under AI assistance with delayed tests are all but absent; that civic tech brings in the condition of unpaid public activity and separates the location of domain knowledge from the builder; and that EUD research has not dealt with the most basic motive of all, building for one's own need.
-
Recent Academic Topics in Design Research (2024–2026): A Literature Map and This Wiki's Position
2026-08-09
An integrative summary mapping 39 peer-reviewed publications (2024–2026) in design research / design studies into eight clusters: AI and design, methodology, the normative turn (sustainability / decoloniality / more-than-human), education, the profession, and the field's self-diagnosis, identifying where this wiki's holdings overlap with the field and where they leave gaps.
-
Writing Material for a Review Article: The Assumption Ledger and Nearest-Neighbour Literature Left by 17 Novelty Audits
2026-08-09
A note that reworks, as writing material for a review article, the record of narrowing 20 paper candidates extracted from a 117-item wiki corpus down to 5 qualitative studies, then generating 12 challenges by assumption reversal (problematization), and subjecting all 17 to external novelty audits. The audits returned 0 items intact, 12 that survive with qualifications, and 5 that are substantively already published, confirming that the novelty of an empirical paper lies not in an unclaimed finding but only in empirical execution within a specific field. Following this verdict, the centre of gravity shifts to a review article, and for the three writable proposals (a scoping review of LLM-as-a-Judge and creativity evaluation, a critical review of the assumptions underlying research on AI and the creative professions, and an integrative review bundling the conditions under which harm occurs) it sets out the research questions, where the contribution is placed, the actual size of the existing corpus, methodological conventions, remaining work, and target venues. As an analytical framework it presents an assumption ledger of 45 items across four fields (AI and design practice, generative AI and learning, the evaluation and measurement of creativity, and design process research and methodology) in five layers, and collects by field the nearest-neighbour literature identified in the audits as evidence for positioning the review. 88 references (with DOI or URL; 7 of them explicitly not verified against the primary source).
-
Issues for Design Education (Not Craft Education): From the Evidence on Failure Design and Cognitive Offloading
2026-08-08
Derives seven issues that arise when this session's accumulated evidence (failure-embedded learning design and its debate, generative-AI cognitive-offloading harms, the robustness assessment of technology-education harms, GBS and scenario learning) is transferred to design education, finalized after a critical review by an academic critic on the design-education canon. The core structure: what generative AI offloads by default (alternative generation, exploration, evaluative judgment) coincides with design education's learning objectives — yet the underlying premise that craft and design judgment can be severed is itself an unresolved issue challenged by the educational canon (Dewey, Lave & Wenger). Most of the evidence is extrapolated from well-defined domains; validation in ill-structured design problems remains an open gap
-
Patterns of Technological Harm to Education, and Their Refutations: A Robustness Assessment Across 25 Works
2026-08-08
Typologized claims of technology harming education (including the Google effect and cognitive offloading) into six mechanism-based patterns and checked 12 original works against 13 peer-reviewed refutations, replications, and boundary-condition reports. Many famous harms (the Google-effect Stroop experiment, laptop note-taking harm, media-multitasking harm, the smartphone mere-presence effect, screen-time harm) have failed replication or shrunk drastically in peer-reviewed work, while classroom attention distraction (RCTs) and passive answer-providing AI remain comparatively robust. Organizes both which harms a learning service should exclude and which refuted harms it should not over-fit to
-
Does Early Use of Generative AI Inhibit the Formation of Thought? A Literature Map of Cognitive Offloading and Learning
2026-08-08
Organized the academic evidence on the claim that early generative-AI use preempts the thinking that would otherwise have emerged from learners, across 16 works: 7 experimental reports of harm, 2 correlational/self-report studies, 3 theoretical frameworks, and 3 counter/conditioning works (plus one retraction record). Harm concentrates under three conditions: answer-providing unguarded AI, novice or struggling learners, and outcomes measured as retention/transfer without AI. The apparent contradiction with positive meta-analyses dissolves as a difference in what is measured. The claim is empirically supported when restricted to early use of unguarded, answer-providing AI
-
The Debate Over Designing Failure Into Learning: Four Lineages Versus Cognitive Load Theory
2026-08-07
Based on a 27-work corpus (14 response works added), organizes the debate between the four failure-design lineages (productive failure, desirable difficulties, impasse-driven learning, error management training) and the cognitive load theory camp (Kirschner, Sweller & Clark) into five contention points (timing of instruction, the validity of the 'minimal guidance' framing, measurement of effects, prior knowledge, the role of affect), with each side's claims, empirical results, and the unresolved conflicts in a contention table
-
Designing Failure Into Learning: From Expectation Failure to Productive Failure and Industry Implementations
2026-08-07
Organized the lineage of deliberately under-instructed learning design across two streams: 13 core works from four learning-science theories (productive failure, desirable difficulties, impasse-driven learning, error management training) and 13 technical learning services whose makers explicitly state intentional difficulty, identifying the industry-side thinness of post-failure consolidation as the gap against theory
-
Scenario-Based System Design Learning: An Industry Survey in the Context of Goal-Based Scenarios
2026-08-07
Collected 28 data points on industry services that teach system design with narrative, role, and mission, and checked them against Schank's seven Goal-Based Scenario components, identifying the scarcity of design-as-operations exercises and the absence of the resources component (just-in-time expert stories)
-
Tools for Learning System Design: Industry Intelligence
2026-08-07
An industry research note organizing tools for learning correct architecture design into four types: review (missing-element detection), exercise-with-feedback, reference, and notation
-
Pre-Production Validation and Learning for System Architecture: Industry Intelligence
2026-08-07
An industry research note organizing the means of validating and learning system architecture before real operation, across standards, independent surveys, developer testimony, and labor-market data
-
Prior Research on Goal-Based Scenarios: A 31-Item Literature Map from Schank's Theory to Components, Evidence, and Design Contexts
2026-08-07
A literature map of prior research on Roger Schank's Goal-Based Scenarios (GBS), organized into two streams, general theory and design contexts (31 items; 40 rows explored, 9 merged as duplicates, 0 excluded). (1) Theoretical background: the lineage from Dynamic Memory's case-based reasoning (a memory theory in which expectation failure drives learning) as the cognitive foundation, to the formulation of GBS as an implementation theory of learning by doing in the 1994 Journal of the Learning Sciences papers. (2) Components: the seven elements (learning goals, mission, cover story, role, scenario operations, resources, feedback; Schank, Berman & Macpherson 1999) and their design procedure. (3) Technical requirements: the ILS implementation lineage of multimedia simulation environments and the GBS Builder. (4) Evidence: application cases and quasi-experiments in statistics education, computing fundamentals, elementary programming, and VR training, identifying the thinness of evidence, with controlled studies concentrated in 2020s Turkish and Taiwanese quasi-experiments and no confirmable meta-analysis. (5) Comparisons with PBL, anchored instruction, Learning by Design, and Merrill's first principles. (6) Design contexts: canonization in instructional design, scenario-centred curricula in engineering education, and PBL practice in HCI studios, while identifying the absence of published research explicitly applying GBS to design studios as a gap.
-
Can the Forward Deployed Designer Stand as a Role?
2026-08-04
An examination that measures the responsibilities of FDEs (Forward Deployed Engineers), strategy consultants, and designers through job postings and primary discourse, then re-situates Nielsen's proposed FDD (Forward Deployed Designer) on top of those differences. The FDD's job content largely overlaps with existing service-design skills; what is new is the previously uncombined pairing of that work with the FDE-style placement (vendor-side employment, responsibility for running artifacts, and a feedback loop into the product). A vacancy exists in the workplace division of labor, but as a job title the market has not yet materialized: the two postings we could confirm have both closed. The watershed for its emergence is whether this placement settles as an equal division of labor rather than subordination to the FDE.
-
AI and Design Monthly Scholarly Watch (July–August 2026)
2026-08-04
A monthly scholarly watch collecting 25 peer-reviewed papers and preprints on 'AI and design' from July through early August 2026. The 11 papers from ACM C&C 2026 shift the center of gravity from 'aligning output' to participation, craft, and governance, and DRS 2026 practice studies partially fill the gap on professional-skill change flagged last month. Meanwhile, new empirical evidence that AI ideation support compresses collective diversity leaves the conflict from the previous watch unresolved.
-
What Is Generative Art For in Education? Seven Purpose Types and a Literature Map of 72 Studies
2026-08-02
A literature map of 72 academic studies (86 explored, 14 merged as duplicates, 0 excluded) that use generative art (generative art / creative coding / computational art) as educational material, organized by the purpose of educational use. Seven purpose types emerge: (1) motivating and contextualizing introductory programming (media computation, Processing/p5.js); (2) inclusion of diverse learners (women, non-majors, low-income communities); (3) developing and assessing computational thinking (CT assessment of Scratch projects); (4) STEAM integration and constructionist making; (5) expanding expressive techniques in art, design, and music education; (6) creativity education; and (7) AI literacy and critical media literacy after LLMs (2023–2026). In terms of educational stages, CT development, STEAM, and inclusion dominate K-12; motivational introductions dominate university CS; and expressive expansion dominates art schools and higher education in architecture and music. Since 2023, 'education about writing code' has been rapidly joined by 'education using generative models as material for critical understanding' (one review reports that empirical studies of generative AI in art education grew from 2 in 2023 to 14 in 2025). All 72 items carry DOIs/URLs.
-
Adversarial Review of Two Generative Art Surveys: Examining the Literature Map and the Citation Strategy Against the Design-Theory Canon
2026-08-02
An adversarial examination of two notes — a generative art literature map (94 items) and a bibliometric analysis ranking citation-earning angles for a review article — through the design-theory lens (Simon/Schön/Cross/Dorst & Cross/Buchanan/Rittel & Webber/Krippendorff/Costanza-Chock). The literature map's four main weaknesses: an entanglement in which AI's ontological status (tool / co-designer / environment / evaluation apparatus) shifts from chapter to chapter; the unsorted concept of 'generation'; a sampling-inference leap that infers a disciplinary disconnect from absence in the 94-item corpus; and a compression of conceptual history that runs the generative Ästhetik of 1965 and the text-to-image of the 2020s through a single definition. The citation-strategy side's five main weaknesses: a slide from citation prediction to research value (a Schön-style critique of delegating problem-setting to the citation market); age bias in the citation-velocity metric and extrapolation to oneself; an exhaustiveness slide in which the qualified gap judgment loses its qualification in the No. 1 verdict; survivorship bias from looking only at highly cited surveys; and the implication leap of the 'cumulative 3,850-citation re-citation node.' At the same time, steelmanning established that the substance of the gap discovery (the unconnected state of computational-creativity evaluation and LLM empirical work) and the limited usefulness as circulation prediction cannot be overturned. All criticisms are tied to sources in an Evidence Ledger, and disagreements are left as a disagreement table rather than folded. The full critiques are stored in source/review/generative-art-adversarial-review/.
-
Which Angle for a Generative-Art Review Paper Attracts Citations? An Empirical Analysis Using Bibliographic Data
2026-08-02
This note empirically identifies, using public bibliographic data from OpenAlex and Semantic Scholar, the angles most likely to attract citations when writing a review paper on generative-art research. (1) Measured citations of 43 existing reviews show that 2023–2025 empirical studies (Zhou & Lee, 191.5 citations/year) and large technical surveys (Yang 2023, 288.2 citations/year) are fastest, while theoretical work holds a stable base of 14–30 citations per year. (2) In theme-level growth, human-AI co-creation is fastest at +607% from 2024→2025, followed by copyright at +160% and LLM creativity evaluation at +115%. (3) Highly cited surveys all include taxonomy and open-problems sections (7/7), and only the top-cited tier maintains GitHub repositories. Systematic reviews are cited more than other research designs (6.6 per year on average; JIF explains R²=0.59). (4) Verification of gap areas shows that a review connecting computational-creativity evaluation frameworks (Boden/SPECS/Lovelace) to the evaluation of LLM/diffusion-model outputs is most promising (a re-citation node for prior theory with over 3,850 cumulative citations, a 15.8-fold increase in primary studies, and the nearest review explicitly stating the connection is absent). The label-effect meta-analysis is no longer a gap, as De Rooij 2025 has already been published. Integrating these, the note ranks the seven most citation-promising angles. All figures carry sources and retrieval dates.
-
The Scholarly Lineage of Generative Art: A 94-Item Literature Map from Information Aesthetics to the Post-LLM Era
2026-08-02
A literature map of scholarly research on generative art, organized into four streams (94 items; 108 rows explored, 8 merged as duplicates, 0 excluded) corresponding to the chapters of a review article. (1) Historical origins: the lineage from Bense's and Moles's information aesthetics as the theoretical foundation, through the coining of generative Ästhetik at the 1965 Nees exhibition and the 'rot' magazine, the 1968 Cybernetic Serendipity and Zagreb New Tendencies, to the art world's rejection (Taylor) and the definitional debate (Galanter's autonomous-system definition). (2) Technical genealogy: the generational succession from L-systems and evolutionary computation (Sims, IEC) through GAN/CAN, StyleGAN, and diffusion models (DDPM, CLIP, Latent Diffusion) to LLM multimodal generation. (3) Theoretical frameworks: Boden's three types and their formalization (Wiggins, Jordanous's SPECS), the authorship debate (Hertzmann's tool argument versus McCormack's four concepts), and empirical reception studies showing that AI labels lower evaluations (Ragot, Bellaiche). (4) The post-LLM era (2024–2026, the thickest chapter): empirical text-to-image studies (AI adoption raises productivity +25% while average novelty declines), LLM code generation for creative coding (Spellburst, GenP5), human-AI co-creation, the arms race between copyright-protection tools and circumvention (Glaze/Nightshade vs. bypass attacks), and split empirical findings on labor effects (no short-term income decline vs. five-year longitudinal reports of job loss). Identified gaps include the disconnection between computational-creativity evaluation frameworks and post-LLM empirical research.