Shuichiro Ogawa
日本語

Notes

Research, experiments, and reflections on AI and design, published along with the process of making them.

Published notes: 151. Built as an LLM wiki. How it works → Ask the notes → See the citation network (how notes share sources) →

July 2026

  • The Genesis of LLM-as-a-Judge: A Confluence of Three Lineages and Four Demands

    2026-07-29

    A genesis history of LLM-as-a-Judge, the concept of using LLMs as evaluators, organized through an extended corpus of 38 academic publications (C1–C38) and 10 vendor primary sources (P1–P10). The concept came into being in 2023 as the confluence of three lineages: (1) the crisis of NLG evaluation (the loss of validity of the surface metrics BLEU/ROUGE, the impossibility of standardizing human evaluation and its reproducibility crisis, and the turn toward learned metrics that made evaluation learnable), (2) the technical foundation of preference learning (RLHF reward models, and AI feedback via Constitutional AI and RLAIF), and (3) the normative demand of scalable oversight. Practice (GPT-4 grading in the Vicuna blog of 2023-03) preceded the term (the title phrase of the MT-Bench paper of 2023-06), and four demands pushed the concept forward: the absence of correct answers in open-ended generation, benchmark saturation and contamination, the supply-demand gap in evaluation, and the need for alignment evaluation. After its establishment, institutionalization advanced through surveys, meta-evaluation benchmarks, and documentation by public institutions (NIST AI 800-2), but the validity problem of surface metrics persists in altered form as the judge bias problem, and the reproducibility problem of human evaluation as the judge meta-evaluation problem.

  • Recent Currents in LLM-as-a-Judge: A Literature Map of 53 Core Studies and a Standalone Chapter on Creativity Evaluation

    2026-07-29

    A literature map of research on LLM-as-a-Judge (using LLMs as evaluators), based on an academic corpus of 53 items (58 explored, 5 merged as duplicates, 0 excluded) and an industry corpus of 39 items (18 T1v vendor primary sources, 9 T2 public institutions, 12 T3 practitioner opinions). On the academic side, it organizes (1) a three-generation lineage from the establishment of prompted judges (MT-Bench, G-Eval) to fine-tuned judges (Prometheus, JudgeLM) to RL-based reasoning judges (JudgeLRM, J1); (2) empirical demonstrations of position bias, verbosity bias, and self-preference bias, together with meta-evaluation benchmarks (on JudgeBench, GPT-4o performs near-randomly on hard problems); and (3) a standalone chapter on the evaluation of creative artifacts (creative writing, creativity tests, images, UI, design). In creativity evaluation, human correlation is high for formalizable rubric dimensions (AUT r=.81, metaphor r=.72), while judgments on originality dimensions and expert-quality design critique remain out of reach; this divergence is confirmed. On the industry side, productization splits into two types, managed services with general-purpose model judges and specialized fine-tuned judges, and both vendor official documents and public institutions (NIST, UK AISI, J-AISI) state calibration against human evaluation as a precondition, converging with the academic bias findings.

  • How Novelty Is Made and How It Is Claimed: A Typology of Establishing Strategies

    2026-07-28

    A literature note that collects, across fields, the strategies for establishing novelty, starting from the question of what operations can actually be moved in response to a demand to produce novelty. It organizes 243 items across eight traditions: problematization research in management studies, heuristics of discovery in the philosophy of science, scientometrics, rhetoric and genre analysis, contribution taxonomies in HCI and design research, the extension and falsification of theory, analogy and boundary crossing, and idea generation in the age of LLMs. It first establishes that strategies split into two kinds—generation (making novelty) and claiming (asserting novelty)—and that the two have been conflated. On the generation side it shows that the assumptions open to challenge divide into five layers (Alvesson & Sandberg), that naming one's ignorance and strategically choosing one's materials are different strategies (Merton's specified ignorance and strategic research materials), that mechanism discovery reduces to decomposition and localization (Bechtel & Richardson), that there is a path by which a tool becomes a theory outright (Gigerenzer), and that atypical combinations work only when built on a base of high conventionality (Uzzi et al.). On the claiming side it shows that Move 2 of CARS is exhausted by four steps—counter-claiming, gap-indicating, question-raising, and continuing a tradition (Swales)—and that hype words have roughly doubled over fifty years and appear in about 97% of successful NIH applications (Hyland & Jiang; Millar et al.). Few strategies have had their effects verified quantitatively, and many of those that have collide with the inverted-U preference of evaluators (Criscuolo et al.) and the distance penalty (Boudreau et al.). 96 references.

  • Where Do Novel Research Questions Come From? The Four Loci and Their Combinations, Examined Against the Literature

    2026-07-28

    A literature note that examines a practitioner's hypothesis—that the only places where a paper's novelty can reside are Related Work, Problem Definition, Materials, Methods, and Results, and that securing two or more of the four loci left after Related Work makes a paper strong—against 135 items across five traditions: problematization research in management studies, scientometrics, philosophy and sociology of science, contribution taxonomies in HCI and design research, and genre analysis of academic English. No scholarly formulation that normatively maps IMRaD sections onto types of contribution was found, so the note first establishes that the four-locus classification stands on its own as practitioner knowledge. It then shows that each locus has an independent lineage (Problem Definition = problematization and specified ignorance; Materials = strategic research materials and experimental systems; Methods = tools-to-theories), that Related Work is not a locus of generation but a locus of claiming corresponding to Move 2 of CARS, and that at the Results locus a design that wins by scale cannot be conflated with serendipity. On the prescription of two or more loci, it sets Uzzi's atypical combinations, Shi and Evans's contents and contexts, and Arts and Veugelers's familiarity and novelty on the supporting side alongside Boudreau's distance penalty, Criscuolo's preference for the moderate, Johnson and Proudfoot's variance in evaluation, and Zhao et al.'s 2026 disconfirming measurement (results-based novelty alone is cited more than holding all three dimensions), and makes explicit, as a gap, that no peer-reviewed evidence directly measuring combinations of loci exists. 78 references.

  • An Academic Map of Research That Treats Hallucination as a Resource for Creativity

    2026-07-21

    A literature map that surveys peer-reviewed research and preprints positioning generative-AI hallucination (confabulation) not as a defect to be suppressed but as a starting point for ideation and creativity. It connects, cluster by cluster: the positive reframing of hallucination (the value of confabulation in Sui et al., valuable hallucination in Chen & Wang, the creativity-perspective survey by Jiang et al., the decoding-layer tradeoff in He et al., the usefulness in drug discovery in Yuan et al.); LLM creativity evaluation and output diversity (Stevenson et al., Bellemare-Pepin et al., the homogeneity of Wenger & Kenett, diversity reduction under RLHF in Kirk et al., diversity collapse at the SFT stage in Karouzos et al., min-p sampling in Nguyen et al.); ideation support and design fixation (AI-augmented brainwriting in Shaer et al., design fixation in Wadinambiarachchi et al., CST in Shneiderman); scientific hypothesis generation (Shahhosseini et al.); idea homogenization and prompt-driven diversity recovery (Girotra et al., Meincke et al.); and the classics of creativity psychology (Guilford, Mednick, Beaty & Kenett). A neutral prior-work review. Thirty-one references (bibliographies confirmed reachable; caveats aggregated at the end). The view that treats the deviations of older, lower-precision models as a resource is thinly represented in explicit research, and is treated here as an implication drawn from research on the performance/diversity tradeoff.

  • Gaps in AI Research, a Fourth Time: Seen from Outside the Inversion Family

    2026-07-21

    An essay that rereads the three sibling notes that searched for gaps in AI research (the seven frames, the A–G operational sheet, the game board of Problem Definition) from the vantage of the full set of problem-setting methods now assembled. It shows that all three generators belong to a single family of operations—binary inversion—and that they had been converging on that family's fixed points (an endogenous scaling law, the institution of recording). On that basis it applies operations that lie outside the family. Morphological analysis (GMA) makes the board's constants orthogonal as dimensions and yields an unvisited intersection cell (a single system that is recursive and ecological and longitudinal at once); TRIZ turns the incompatibilities that cannot be inverted (cultural diversity versus universal human rights; individual novelty versus collective homogenization) into design problems solved by separation; PSM/CATWOE restores the demoted cognitive diversity and non-Western epistemologies as 'another board' and points to the absence of a method for adjudicating the plurality of boards; and Cynefin and representational change diagnose the infinite regress all three fall into not as a defect but as a misclassification of the type of problem situation. On the single point that convergence may be a property of the family of operations more than of the object, it weakens the earlier three notes' argument from 'the convergence of independent generators.' The three gaps that the operations outside the family pointed to were checked against the coverage of recent (2024–2026) prior work, and each was judged partial (a residue can be narrowed against a dense body of prior work). References augmented across method, evidence, and coverage-confirmation studies.

  • A Genealogy of Practitioner Methods for Reframing Problems: Who Made Them, Traced to Their Origins

    2026-07-21

    An industry-side map of practitioner methods for problem setting, problem framing, and reframing, narrowed to their originators, primary sources, and standing in practice. It covers convergence/divergence process methods (Design Council's Double Diamond, IDEO/d.school's Design Thinking, GV's Design Sprint), question-posing/reframing methods (How Might We, de Bono's lateral thinking and Six Thinking Hats, First Principles, SCAMPER), structured-decomposition methods (TRIZ, morphological analysis, Toyota's Five Whys, McKinsey's seven-step process, Conn & McLean's Bulletproof, Minto's SCQA), and the three lineages of Jobs-to-be-Done (Ulwick/Christensen/Moesta). Where a method's origin is disputed, the note lists the competing accounts rather than asserting one, and treats promotional effectiveness figures as partial with their methodology attached. The gaps between academia and industry (IDEO is a propagator, not the originator, of HMW; First Principles offers no decision criterion; GMA and TRIZ have not penetrated practice, etc.) are organized at the end. A neutral practitioner review. References carry primary sources and official explanations.

  • An Academic Map of Methods for Reframing Problems: From Abduction-2 to Problem Structuring

    2026-07-21

    A literature map that surveys the methods of problem setting, problem framing, and reframing across peer-reviewed research and canon from design methodology, cognitive science, and systems thinking. Starting from Dorst's abduction-2 (frame creation), it connects, by lineage: functional fixedness and representational change in cognitive science (Duncker, Ohlsson, Kaplan & Simon, Chi); the psychology of problem finding and problem construction (Getzels & Csikszentmihalyi, Reiter-Palmon, Mumford); C-K theory and problem-solution co-evolution (Hatchuel & Weil, Maher & Tang, Wiltschnig); wicked problems and problem structuring methods (Rittel & Webber, Checkland SSM, Eden SODA, Rosenhead PSM, Ackoff); Cynefin (Kurtz & Snowden); morphological analysis (Zwicky, Ritchey); the theorization of TRIZ (Altshuller, Cavallucci); lateral thinking (de Bono); and empirical work on reframing in design (Valkenburg & Dorst, Paton & Dorst, Stompff). A neutral prior-work review. 38 references (bibliography reachability-confirmed and ISBNs fixed; remaining reservations collected at the end).

  • AI Research Gaps, a Third Time: Rereading Them as Rewritings of the Game Board

    2026-07-20

    An essay that rereads the two sister notes that searched for gaps in AI research (the catalogue of seven frames, and the revisit via the A-through-G worksheet) with a third generator: Problem Definition (operations on the game board). It writes out AI research's game board, with sources, in seven entries (certification of success, measurement, unit, endpoint, position of observation, noise, record-keeping institution) and selects the three most unnatural asymmetries from what was pushed off the board. It maps each of the seven frames to the board operation it was, showing that symmetry inversion is one species of board rewriting (an operation specialized to an asymmetric pair). It retries the two candidates demoted in the revisit note (cognitive diversity and non-Western epistemologies, and the self-mapping of blind spots), makes explicit that the demotion was a certification by the A-through-G worksheet's own success-certifying apparatus (falsifiability and the minimal observable), and reinstates them as questions strong under a different type while upholding the determination of kind. It confirms that applying the conversion table to scaling laws comes out at the same void as the endogenous scaling law, taking the convergence of independent generators as weak evidence that the void is not an artifact. Finally it generates two problem statements and RQs whose units are the success-certifying apparatus and the record-keeping institution (both gap candidates whose coverage by prior work is unchecked), and closes with the division of labor between generators (Problem Definition generates; A-through-G vets). 17 references.

  • Thinking in Problem Definition: Rereading the Game Board of Existing Research

    2026-07-20

    Relocates the strength of research from 'the ability to produce good answers' to 'the ability to see through what is being counted as a problem,' and organizes the craft of Problem Definition, which operates on the observation apparatus, evaluation axes, time horizon, units, and boundaries (the game board) that existing research implicitly fixes. Presents ten patterns easy to transfer, a 14-row meta-conversion table that maps existing research onto questions, a five-step generation protocol that lowers a description of the board into a problem statement and an RQ, and five habits of thought, connecting them to the lineage of the contrast between gap-spotting and problematization (Sandberg & Alvesson), problem setting (Schön), problem finding (Getzels & Csikszentmihalyi), ill-structured problems (Simon), and wicked problems (Rittel & Webber). 10 references.

  • AI Research Gaps, Revisited: Sorting Symmetry-Completion from Coverage, and Digging Out an Endogenous Scaling Law

    2026-07-20

    A revisit of the seven frames catalogued earlier, using a sharpened discipline (a procedure that lowers a premise, inverts a symmetry, and operationalizes it across an A-through-G worksheet). The same discipline is turned on its own past product to select, strengthen, and generate blind spots. Selection yields three groups: the single strongest candidate whose every slot fills (the relational ontology of capability), four that survive as symmetry-completion frames, and two that fail the symmetry test and are demoted (cognitively diverse users with non-Western epistemologies, and self-mapping of blind spots). Demotion is not disqualification but a determination of kind: legitimate fairness-and-coverage concerns that are not symmetry-completion blind spots. The discipline then generates one new blind spot. It inverts the open-loop assumption that a scaling law treats its input distribution as exogenously fixed, and formalizes a closed loop in which a model's capability produces its future data, an endogenous scaling law. Novelty is limited strictly. Model collapse, self-consuming loops, and scaling laws parameterized by the synthetic-data fraction are all prior work; what remains new is only inverting the exogeneity of the input distribution itself, so the data fraction becomes an endogenous state variable that closes the loop. 24 references (21 with confirmed bibliography, 3 flagged for primary verification).

  • The Vanishing of Knowledge: Undiscovery as the Dual of Discovery

    2026-07-19

    A note that inverts, with the symmetry-completing move by which de Broglie derived matter waves, the asymmetry whereby science and its meta-science measure only the production of knowledge and treat loss as accident or noise. The inversion says that the vanishing of knowledge (undiscovery) is a process dual to discovery, with its own mechanisms and a measurable rate of loss. Agnotology, the study of deliberately manufactured ignorance, serves as the 'light quantum' shown on one side; the completion extends loss into a general, rate-bearing process and binds production and loss into a single ledger. This symmetrization is offered not as an established fact but as a falsifiable hypothesis, with two predictions: a measurable undiscovery rate and a loss-adjusted progress metric. Novelty is strictly limited to the symmetrization itself and the loss-adjusted meta-science; the description of individual loss phenomena is prior art and is credited by name. 22 references (20 confirmed, 2 with unreached full text flagged for primary verification).

  • Is Expertise a One-Way Street? Unlearning and the Blind Spot of Reversible Mastery

    2026-07-19

    A note that takes one asymmetry quietly assumed by the learning sciences and expertise research, namely that development is a monotonic climb from novice to expert while unlearning and regression are degradations to be minimized, and inverts it with the same move by which de Broglie derived matter waves: completing a symmetry. The inversion says that unlearning is a developmental process symmetric to learning, and that expertise is not a one-way ladder but a reversible, frame-relative state. The Einstellung effect and design fixation are the 'light quanta,' the side where this asymmetry has already been shown to break empirically. The inversion is offered not as an established fact but as a falsifiable hypothesis: a controlled experiment can test whether deliberate unlearning interventions outperform 'learning more' for adaptation to a changed environment. Novelty is strictly limited to the ontological inversion that treats unlearning as a developmental stage symmetric to learning, and to the reversibility of expertise. 17 references (16 with confirmed bibliography, 1 pending primary verification).

  • The Creativity of Subtraction: The Additive Bias, and the Hypothesis of Designing Absence

    2026-07-19

    A single deep argument that inverts one asymmetry creativity and design quietly assume, namely that addition (multiplying options, adding elements) is generative while subtraction (removing, withholding, not-making) is a secondary work of convergence, using the same symmetry-completion move by which de Broglie derived matter waves. The inversion says subtraction is a first-class generative operation symmetric to addition, and that absence and the un-made carry structure not reducible to filtering. This is offered not as a settled fact but as a falsifiable hypothesis, testable by comparative experiments and tool audits. The additive bias has already been demonstrated on the problem-solving side (Adams, Converse, Hales & Klotz 2021, Nature); the core proposal treats that as the 'photon shown on one side' and moves the target to the symmetry of generative operations in design theory. Novelty is strictly limited to two points: the theoretical symmetrization that raises subtraction to a first-class generative operation, and the methodological audit of the additive bias in creativity metrics and authoring tools. 24 references (20 with confirmed bibliography, 4 flagged for primary verification).

  • Is Capability Inside the Model? A Relational Ontology of AI Evaluation and a Test for Contextuality

    2026-07-19

    A single deep argument that inverts one asymmetry AI evaluation quietly assumes, namely that a model 'has' a capability while measurement merely 'reveals' it, using the same symmetry-completion move by which de Broglie derived matter waves. The inversion says there is no context-independent 'true capability': measurement does not reveal capability, it brings it into being. This is offered not as a settled fact but as a falsifiable hypothesis, testable by whether capability measurement exhibits non-classical contextuality. The sheaf-theoretic apparatus has already been applied at scale to LLM word meaning with non-classicality measured empirically (Lo, Sadrzadeh & Mansfield 2025, Proc. R. Soc. A); the core proposal is to move the target from meaning to capability. Novelty is strictly limited to three points: the ontological inversion that drops realism, the transfer of the contextuality test from meaning to capability, and the unification of scattered phenomena under one ontology. 22 references (19 with confirmed bibliography, 3 flagged for primary verification).

  • Where Are the Gaps in AI Research? Seven Voids Found by Reframing

    2026-07-19

    An essay that searches for AI's unexplored territory not through the open problems surveys enumerate (deductive gap-finding) but through the leap of reframing, inverting a field's taken-for-granted premises (the abduction-2 of Dorst and Peirce). It raises seven frames—stationarity, research reflexivity, ecology, the asymmetry of non-knowledge, cognitive diversity, the monism of epistemology, and the field's own blind spots—and presents the gap each frame reveals as a conditional hypothesis. In every section it names existing partial research (model collapse, emergent norms in LLM populations, decolonial AI, the systematization of the reproducibility crisis) to narrow each gap precisely, from 'wholly unexplored' to 'the absence of an integrative methodology' or 'an asymmetry of institutionalization.' 45 references (36 with confirmed bibliography, 9 carrying a primary-source-verification caveat for partial bibliography or unreached full text).

  • Methods and Validity in Design Research: A Literature Map of Case Study, Research through Design, and Mixed Methods

    2026-07-19

    A literature map organizing the methodological foundations of knowledge generation through artifacts and practice-based research, drawing on established public sources. Five themes—case study design (Eisenhardt 1989, Yin 2014), Research through Design and constructive design research (Frayling 1993, Zimmerman et al. 2007/2010, Koskinen et al. 2011), securing validity and reliability (Cook & Campbell 1979, Lincoln & Guba 1985), generalizing from cases (Flyvbjerg 2006, Yin's analytic generalization), and integrating quality and quantity in mixed methods (Tashakkori & Teddlie 2010, Creswell & Plano Clark 2011)—are described not to advocate one position but to show how these frameworks divide the labor of answering. 10 references (all in English).

  • Scaffolding Learning and the Design of Interventions: A Literature Map of Cognitive Apprenticeship, Legitimate Peripheral Participation, and Theories of Mediation

    2026-07-19

    A single map of the published literature on how to design interventions that support learning and the development of expertise. It runs from scaffolding and the zone of proximal development (Wood, Bruner & Ross 1976; Vygotsky 1978), through cognitive apprenticeship and fading (Collins, Brown & Newman 1989), legitimate peripheral participation (Lave & Wenger 1991), reflective practice (Schön 1983, 1987), collaborative learning and intrinsic motivation (Dillenbourg 1999; Johnson & Johnson 1989; Deci & Ryan 1985; Ryan & Deci 2000), the general theories that treat artifacts as mediators (Latour 2005; Callon 1984), to the psychometric scales and interaction analyses that measure interventions (Dennis & Vander Wal 2010; McLain 2009; Sacks, Schegloff & Jefferson 1974; Fairclough 2003). 16 references (6 verified, 10 to be confirmed).

  • Justifying the Democratization of Design: A Literature Map of Public Goods, the Capability Approach, and Theories of Justice

    2026-07-19

    A neutral literature map that answers the normative question of why design should be democratized, using only established published sources. The Nordic lineage of participatory design shows process democratization as practice; Arrow and Hess & Ostrom's work on knowledge as a public good gives it an economic foundation. The capability approach of Sen and Nussbaum, Rawls's distributive justice, and Fraser's theory of recognition ground democratization at the level of desirability. Finally, Fischer's meta-design and Wenger's communities of practice show the redefinition of the expert's role and the conditions for sustained participation. 12 references.

  • Design and the Professions: A Literature Map of Design's Position Seen Through the Theory of Professions

    2026-07-19

    Is design a profession in the same sense as medicine or law? This note maps that question neutrally against the sociological theory of professions. It builds on Larson's (1977) account of professionalization as market closure, Abbott's (1988) theory of jurisdiction and inter-professional competition, and Freidson's (2001) professionalism as a third logic, and it positions Schön's (1983) reflective practice, Cross's (2001) disciplining of design, Buchanan's (1992) wicked problems, Star & Griesemer's (1989) boundary objects, and Fischer & Scharff's (2000) meta-design. Design carries practical knowledge and a disciplinary basis, yet it has reached the present without the institutional conditions of entry licensure and jurisdictional monopoly. 8 references.

  • The Many Dimensions of Design Value: A Literature Map from Instrumental to Relational and Recognitive Value

    2026-07-19

    A neutral survey, drawing only on established published literature, of the claim that design produces value across several dimensions rather than one. It ranges from instrumental usability (Norman 1988; Nielsen 1993), through learning from experience and meaning-making (Kolb 1984; Bruner 1990), the distinctive stance of designing (Cross 1982; Michlewski 2008), and relational value in communities of practice (Lave & Wenger 1991; Wenger 1998), to the ontological dimension of recognition (Honneth 1995). Practice theory (Bourdieu 1977; Schatzki 2002; Reckwitz 2002) offers a lens that reads these dimensions as socially formed practices rather than individual traits. 12 references.

  • Creativity, Situated Cognition, and Environment: A Literature Map of 4E Cognition, Affordances, and Situated Learning

    2026-07-19

    A survey of published scholarship that treats creativity and cognition not as inner individual traits but as distributed across body, environment, situation, and relations with others. It runs from Gibson's affordances (1979) and Norman's adaptation for design (1988), through the embodied mind of Varela, Thompson, and Rosch (1991) and the extended mind of Clark and Chalmers (1998), to Lave and Wenger's legitimate peripheral participation (1991), Hutchins's distributed cognition (1995), Vygotsky's zone of proximal development (1978), and Wood, Bruner, and Ross's scaffolding (1976), and on to a critical view of measuring ability apart from environment (Ross 1977, Merton 1968, Bourdieu & Passeron 1977, Sandel 2020). 14 references.

  • The Epistemological Premises and Methodological Foundations of Design: A Literature Map of Wicked Problems, Abduction, and Set-Based Exploration

    2026-07-19

    The problems design confronts have no clear stopping rule, and the act of designing cannot avoid value judgment. Starting from this epistemological premise, this note maps reductionism and its limits, methodological pluralism, reflective practice and abduction, and set-based exploration through established published literature. Organized around Rittel & Webber's wicked problems (1973), it places Simon, Dewey, Churchman, Ackoff, Buchanan, Dorst, Feyerabend, Cross, Peirce, Schön, and Ward. 15 references.

  • Research Currents in AI and the Study of Art and Culture: Aesthetics, Media Studies, Computational Creativity (2024–2026)

    2026-07-19

    A synthesis of 16 scholarly works on the intersection of generative AI and the study of art and culture, across aesthetics, computational creativity, media/cultural studies, musicology, and critical cultural theory. It traces how authorship and aura are being redefined, how creativity evaluation is being questioned, and how the politics of representation and labour moved to the centre. 16 references.

  • Research Currents in AI, Education, and the Learning Sciences: Evidence on Generative AI and Learning (2024–2026)

    2026-07-19

    A lightweight scoping synthesis of 39 peer-reviewed studies at the intersection of generative AI and general education and learning sciences. Organized into four clusters: large RCTs and meta-analyses, metacognitive laziness and cognitive offloading, assessment and academic integrity, and equity and the changing role of teachers. 39 references (plus one retracted paper noted separately).

  • AI, Psychology, and Cognitive Science Research Trends: Machine Psychology and the Cognitive Modeling of LLMs (2024–2026)

    2026-07-19

    An integrative summary of research trends at the intersection of large language models and psychology/cognitive science in 2024–2026, drawn from 30 peer-reviewed papers and major preprints. It tracks four clusters: machine psychology, LLMs as cognitive models, the psychology of human-AI interaction, and the methodological turn in cognitive science. 30 references.

  • Research Currents in AI and Economics: Productivity, Labor Markets, and Methods (2024–2026)

    2026-07-19

    A synthesis of 16 studies at the intersection of generative AI and economics that gained depth in 2024–2026, drawn from peer-reviewed journals and NBER/SSRN working papers. Organized into five clusters: productivity field experiments, labor-market exposure and employment, macro growth theory, LLMs as a research tool, and algorithmic pricing. 16 references.

  • Research Trends in AI Law and Governance: Regulation, Accountability, and AI Governance (2024–2026)

    2026-07-19

    A literature review of AI law and governance research from 2024–2026, drawn from 33 primary legal sources and peer-reviewed law articles. Covers the staged application of the EU AI Act and its GPAI Code of Practice, the reversal of U.S. executive orders, the multi-layered governance of NIST/OECD/UN/Council of Europe, the institutionalization of algorithmic audits, gaps in AI liability, and the litigation and scholarship over copyright and LLM training. 33 references.

  • AI and the Humanities: Digital Humanities and Large Language Models (2024–2026)

    2026-07-19

    A literature review of the intersection of AI (LLMs and generative AI) and the humanities, drawn from 19 peer-reviewed articles and major preprints published between 2024 and mid-2026. It traces five currents: the LLM turn in digital humanities, the reversal in transcription and OCR accuracy, the philosophy of AI around meaning and understanding, the disputes over authorship and disclosure, and critical AI studies. 19 references.

  • AI and Social Science Research Trends: Computational Social Science and LLM Agents (2024–2026)

    2026-07-19

    An integrated scoping summary of research currents that emerged after 2023 at the intersection of AI (especially LLMs and generative AI) and the social sciences, drawn from peer-reviewed papers and highly cited preprints. It surveys four clusters: the LLM turn in computational social science, LLMs as research tools and the validity critiques of that use, silicon sampling and generative agents, and AI as a substantive object of study. 29 references.

  • Analyzing the Making Process and Its Contextual Supports: A Literature Map of the Critical Incident Technique, Reflective Writing, and Design Education History

    2026-07-19

    Organized around the question of how to record and analyze the making process and reflection on it, this note reviews three methodological lineages neutrally: the critical incident technique (Flanagan 1954), the design of structured reflection (Ash & Clayton's 2009 DEAL model), and the quality assessment of reflective writing (Ullmann 2019), attending to differences in collection design and unit of analysis. As contextual support it adds the quality assurance of literature reviews (Boote & Beile 2005 and others) and historical cases of environments that cultivate an exploratory disposition (the Bauhaus preliminary course, organizational slack-time programs). The organizational slack-time cases are handled with their low academic rigor and success bias made explicit. 11 reference footnotes (the Bauhaus note bundles 3 works).

  • The Democratization of Design and Ontological Designing: A Literature Map of Pluralizing the Paths of Value Realization

    2026-07-19

    Organized around the question of who designs and along which paths the value of that design is realized, this note maps the literature on the democratization of design as a neutral review. It brings together Manzini's design for social innovation (2015), ontological designing (the circle in which design creates a world that in turn creates us), traceable to Willis and to Winograd & Flores, Escobar's designs for the pluriverse (2018), Honneth's theory of recognition (1995), and Tronto's caring democracy (2013), under the single question of pluralizing the paths of value realization. 12 references.

  • Participatory Design and the Design of Collaboration: A Literature Map of Boundary Objects, Infrastructuring, and the Conditions of Participation

    2026-07-19

    A single map of the literature on how to design the participation of non-experts in design. It runs from participatory design rooted in Nordic labor movements (Ehn 1988; Robertson & Simonsen 2013), through the sociotechnics of an age when everybody designs (Manzini 2015), the generative tools and probes that engage non-designers (Sanders & Stappers 2008, 2012), the boundary objects that translate knowledge across groups (Star & Griesemer 1989; Carlile 2002, 2004), infrastructuring and agonism as the slow cultivation of collaborative ground (Ehn 2008; Karasti 2014; Björgvinsson et al. 2012), to design justice and the capability approach that interrogate the inequality of participation (Costanza-Chock 2020; Oosterlaken 2015). 17 references (0 Japanese, 17 English).

  • Is the Primacy of Qualitative Methods in Design Research Academically Mainstream?

    2026-07-15

    The conclusion of the preceding note (qualitative-quantitative-design-research) -- that the exploratory character of design is epistemologically consonant with qualitative research -- occupies a mainstream position in design research since Frayling (1993), Cross (2006), and Dorst (2011). A bibliometric analysis of fifteen years of Design Studies (Chai & Xiao 2012) corroborates the predominance of qualitative methods. Three countervailing tensions nevertheless persist: the quantitative orientation of HCI (the experimental norms of CHI), the quantitative tradition of evidence-based design, and the call by Gaver (2012) and Koskinen et al. (2011) to transcend the qualitative-quantitative dichotomy altogether. A research agenda centered on 'how evaluation affects people's willingness to try again' and 'designing conditions that enable a second attempt' harbors a methodological tension: qualitative methods are needed to describe those conditions, yet some form of empirical verification is required to assess whether altered conditions produce the intended effects. This tension cannot be bridged by mixed-methods pragmatism alone; a mechanism-oriented epistemology such as critical realism (Bhaskar 1975) emerges as a candidate framework.

  • Qualitative and Quantitative Research: In the Context of Design

    2026-07-15

    An overview of the epistemological foundations, strengths, and limitations of qualitative and quantitative research, along with guidelines for methodological choice in design research. Design is an exploratory activity that envisions and realizes what does not yet exist (Simon 1969; Schon 1983; Cross 2006). Because of this character, qualitative methods (ethnography, protocol analysis, case studies, research through design) are called for at the exploratory stage, while quantitative methods (usability testing, surveys, experiments) are needed at the evaluative stage. Mixed methods research provides a framework that methodologically secures the continuity between these two stages within a single study.

  • AI and Design Weekly Watch (2026-07-06 to 07-13)

    2026-07-13

    An integrated summary of 'AI and design' developments over the past seven days, collected in three tiers: T1v vendor primary sources, T2 public institutions and research, and T3 expert opinion. This was a week in which the design shift toward 'treating images as collections of objects,' shown separately by Adobe and Wroblewski, and the regulatory move by Korea's intellectual property authority to require records of human contribution in design applications lined up as the technical and institutional faces of the same movement: decomposing artifacts into elements and giving them units.

  • AI and Design Monthly Scholarly Watch (June–July 2026)

    2026-07-12

    A monthly scholarly watch collecting 30 peer-reviewed papers and preprints on 'AI and design' from June through early July 2026. ACM DIS 2026 alone accounts for 13 papers addressing the alignment of generative AI with design intent, while two empirical studies from the same period report that generative AI narrows divergent thinking and lowers functional success rates. This conflict is not resolved within the present corpus.

  • AI and Design Weekly Watch (2026-07-05 to 07-12)

    2026-07-12

    An integrated summary collecting the past 7 days of 'AI and design' developments in three tiers: T1v vendor primary sources, T2 public institutions and surveys, and T3 expert commentary. This week's three observation points: the commoditization of execution seen in GPT-5.6's same-week integration into tools, experts converging on judgment, taste, and critique as the scarce resources, and regulation moving toward mandatory disclosure of AI-generated content.

  • Competency Measurement Premised on AI Use — The Aptitude x AI Skill Interaction and Measurement Frameworks

    2026-07-10

    Analyzes human competency measurement premised on AI use, drawing on 28 academic sources and 16 industry sources. Three structural findings: (1) AI compresses the productivity distribution (43% improvement for low-skill workers vs 17% for high-skill workers, Dell'Acqua 2026 N=758), yet the higher the task complexity, the more existing expertise is amplified; (2) articulation ability and proactiveness indirectly determine the effectiveness of AI use (Power Users experiment 68% more frequently and persist 30% more often after failure, Microsoft 2024 N=31,000); (3) existing AI literacy scales show divergence between self-report and objective assessment (Zhang et al. 2026). What should be measured is not 'whether one can use AI' but 'which cognitive functions are amplified through collaboration with AI.'

  • Design Systems as AI's Foundation Layer — Structured Design Knowledge Determines Agent Accuracy

    2026-07-09

    Analyzes the structural transformation whereby design systems (component libraries, design tokens, structured specifications) function as foundational infrastructure for AI agents, drawing on 27 academic papers and first-party information from 5 vendors. Provision of formal specifications reduces agent navigation by 33-44% and achieves 100% accuracy (Jin 2026); leveraging Figma JSON metadata outperforms pixel-only conversion (Gui 2026, ICLR); 5 companies simultaneously implement design systems as AI context layers. The meaning of investing in design systems has shifted from 'maintaining consistency' to 'determining the accuracy ceiling of AI agents.'

  • AI Adaptation in Design Education — The Current State and Structural Challenges of Curriculum Reform

    2026-07-09

    Analyzes AI-era design education curriculum reform drawing on 22 academic publications and information from 8 educational institutions plus public bodies. Three structural findings: (1) research concentrates on creativity development (35.9%), assessment (27.6%), and curriculum design (22.4%), with 66.7% not reporting outcomes (Musiienko 2026); (2) four newly required skills are identified (agency, domain knowledge, imagination, and aesthetic judgment); (3) no update to NASAD's AI standards has been confirmed. Reflective practice is being reconceived as 'learning by co-doing.'

  • From Maker to Editor: A Structural Analysis of the Designer Role Transition in the Age of AI

    2026-07-09

    A cross-sectional analysis of the phenomenon in which the designer's role shifts from maker to editor/curator, drawing on 19 academic sources and 15 industry sources. Three structural findings: (1) the transition is empirically confirmed, with 71% spending more time on evaluation/curation than original production (Rivera & Russi 2026, 217 practitioners across 43 countries); (2) new judgment typologies have emerged alongside the transition (agency allocation judgment, trustworthiness judgment); (3) behind the efficiency narrative, 'AI management labor' has surfaced as a new form of cognitive burden. Industry data reveal a structural divergence between 90% adoption and only 10% approval.

  • EU AI Act Article 50 and Design Practice — A Structural Analysis on the Eve of Enforcement

    2026-07-09

    Analyzes the impact of EU AI Act Article 50 (transparency obligations), effective August 2, 2026, on design practice through a corpus of 30 academic publications. Three structural problems are identified: (1) the scope of disclosure obligations is ambiguous, leaving the boundary between design editing and regulated manipulation undefined; (2) watermark and labeling implementation rates fall far short of regulatory requirements (38% / 18%); (3) greater disclosure granularity paradoxically erodes trust rather than improving user judgment. Designers confront a 'disclosure dilemma' — disclose AI use and lose client trust, or conceal it and incur legal risk — while gaps in copyright protection undermine the basis for billing.

  • 90% Adoption x 10% Approval — The Paradox of AI Tool Diffusion and Evaluative Divergence

    2026-07-09

    Analyzes the structural divergence in creative industries where AI tool usage rates reach 90% while only 10% view the impact on their industry positively, drawing on 27 academic sources and 15 industry surveys. Three structural mechanisms are identified: (1) mandatory adoption generates symbolic adoption — using without endorsing (Heidenreich & Talke 2020); (2) AI triggers professional identity threat and psychological reactance (Mirbabaie et al. 2022; Jussupow et al. 2022); (3) technostress negatively moderates continuance intention across all pathways (Lu & Hu 2025, n=443 designers). A parallel trend appears among developers (Stack Overflow: trust declining from 40% to 29%, favorability from 72% to 60%). 'Using it but not endorsing it' is not individual irrationality but a structural collision between organizational pressure for efficiency and professional identity.

  • Activating Reframing in Design: Cognitive Mechanisms and Practical Methods

    2026-07-09

    Reframing -- the act of restructuring a problem's frame to open new solution spaces -- is a core competency of design cognition. Centering on Dorst's Frame Innovation (2015), this note examines the cognitive mechanisms of reframing (abduction, problem-solution co-evolution), eleven practical methods for activating reframing (analogical reasoning, How Might We, assumption surfacing, TRIZ contradiction analysis, speculative design, the 9-step Frame Creation process, and others), inhibiting factors (design fixation, anchoring, functional fixedness), and conditions at the team and organizational level (boundary objects, strategic use of surprise). Reframing is not an innate creative gift; it can be deliberately activated through structured methods and environmental design.

  • Field Deploy Engineers and Designers: How Field Knowledge Shapes Design Judgment

    2026-07-09

    Discusses how the professional experience of a field deploy engineer providing global on-site support for superconducting and cryogenic systems functions as an epistemological foundation for design, through career transitions into service designer, business designer, and design consultant. The trajectory in which the object of deployment abstracts from physical systems -> software -> learning experiences -> design methodology can be read as a concrete instance of Schon's reflective practice and Suchman's situated action. As designers' work shifts from 'making' to 'judging' in the AI era, the note shows through one practitioner's career that the sense of field deployment (the judgment to make a designed solution function in the context of the field) becomes a core capability of design.

  • The Nexus of Design and Social Science: A Genealogy of Methods, Theories, and Institutions

    2026-07-09

    This note traces the genealogy through which design has absorbed methods and theories from the social sciences to form its own distinctive mode of knowledge. Beginning with Simon's sciences of the artificial (1969), continuing through Schon's reflective practice, Cross's designerly ways of knowing, and Buchanan's wicked problems, the trajectory extends to the importation of ethnography, participatory design, practice theory, actor-network theory, and activity theory. HCI, design anthropology, and STS have served as bridging fields, while contemporary movements such as design justice, transition design, and more-than-human design are reshaping the boundary between design and social science. Design's distinctiveness lies not in the appropriation of social-scientific methods but in the integrative judgment that envisions artifacts not yet in existence.

  • AI and Design Weekly Watch (2026-06-29 to 07-06)

    2026-07-06

    A consolidated summary of the past seven days of 'AI and design' developments, collected in three tiers: T1v vendor primary, T2 public institutions and surveys, and T3 expert commentary. The main observations are same-day integration of new models into tools, the divergence between 90% usage rates and evaluations, and the role shift from makers to editors.

  • Knowledge Management Methods for the LLM Era (Mid-2026 Status Report)

    2026-07-06

    A mid-2026 survey of context and knowledge management methods designed for LLMs. Covers Google OKF (Open Knowledge Format), the Karpathy LLM Wiki pattern, the lineage of Context Engineering, the convergence of in-repo knowledge files (CLAUDE.md / AGENTS.md / Cursor Rules), tool connectivity via MCP, integration with PKM tools (Obsidian / Notion / Logseq / Tana), GraphRAG, and Fabric.

  • Model Tiering Patterns for Fable

    2026-07-05

    Organizes five patterns for tiering models within the Claude family (router, cascade, orchestrator + worker, task classification, cache stacking) to curb Claude Fable 5's token consumption. Applying findings from RouteLLM, FrugalGPT, and Anthropic's multi-agent work to Fable operations, cost reductions can be expected ranging from 40-70% for a single lever to 70-85% with all levers combined.

  • What Is the Frontier Model Premium Buying? (A Debate via Historical Analogy)

    2026-07-03

    Debates the future returns of premium payments for the highest-performing AI models through historical analogies (dynamo/AlexNet/Bloomberg/SGI/Concorde), productivity RCTs, and disruption from the good-enough argument. Deliberately non-convergent, with a table of unresolved disagreements as the primary deliverable

  • AI and Design Weekly Watch (2026-07-02)

    2026-07-02

    Organizes the latest developments in "AI and design" centered on 2026-06-25 to 07-02 into three tiers: T1v vendor primary sources, T2 public institutions and surveys, and T3 trusted individual commentary. Covers the integration race immediately following Figma Config 2026, Canva Grow 2.0, Webflow's ChatGPT integration, EU/US regulation of AI-generated content, surveys on AI adoption in design practice, and the axis of expert disagreement over floor and ceiling.

  • Agentic Coding: The Current State of Orchestration Patterns (2026)

    2026-07-01

    Organizes orchestration patterns in agentic coding as of mid-2026. Contrasts, with sources, five types of inter-agent communication (hierarchical / handoff / shared state / message passing / context isolation), seven orchestration architectures, the design philosophies of eight major frameworks (Claude Code / OpenAI Agents SDK / LangGraph / CrewAI / MS Agent Framework / Google ADK / Strands / Mastra), and converging best practices and anti-patterns.