Notes — Research Log
Notes.
A log of research, experiments, and reflections. Ongoing investigations into AI applications and design processes, accumulated and published in an LLM Wiki format.
Want to ask about the notes? Ask the Wiki →
-
Pre-Production Validation and Learning for System Architecture: Industry Intelligence
2026-08-07An industry research note organizing the means of validating and learning system architecture before real operation, across standards, independent surveys, developer testimony, and labor-market data
Read → -
Prior Research on Goal-Based Scenarios: A 31-Item Literature Map from Schank's Theory to Components, Evidence, and Design Contexts
2026-08-07A literature map of prior research on Roger Schank's Goal-Based Scenarios (GBS), organized into two streams, general theory and design contexts (31 items; 40 rows explored, 9 merged as duplicates, 0 excluded). (1) Theoretical background: the lineage from Dynamic Memory's case-based reasoning (a memory theory in which expectation failure drives learning) as the cognitive foundation, to the formulation of GBS as an implementation theory of learning by doing in the 1994 Journal of the Learning Sciences papers. (2) Components: the seven elements (learning goals, mission, cover story, role, scenario operations, resources, feedback; Schank, Berman & Macpherson 1999) and their design procedure. (3) Technical requirements: the ILS implementation lineage of multimedia simulation environments and the GBS Builder. (4) Evidence: application cases and quasi-experiments in statistics education, computing fundamentals, elementary programming, and VR training, identifying the thinness of evidence, with controlled studies concentrated in 2020s Turkish and Taiwanese quasi-experiments and no confirmable meta-analysis. (5) Comparisons with PBL, anchored instruction, Learning by Design, and Merrill's first principles. (6) Design contexts: canonization in instructional design, scenario-centred curricula in engineering education, and PBL practice in HCI studios, while identifying the absence of published research explicitly applying GBS to design studios as a gap.
Read → -
Can the Forward Deployed Designer Stand as a Role?
2026-08-04An examination that measures the responsibilities of FDEs (Forward Deployed Engineers), strategy consultants, and designers through job postings and primary discourse, then re-situates Nielsen's proposed FDD (Forward Deployed Designer) on top of those differences. The FDD's job content largely overlaps with existing service-design skills; what is new is the previously uncombined pairing of that work with the FDE-style placement (vendor-side employment, responsibility for running artifacts, and a feedback loop into the product). A vacancy exists in the workplace division of labor, but as a job title the market has not yet materialized: the two postings we could confirm have both closed. The watershed for its emergence is whether this placement settles as an equal division of labor rather than subordination to the FDE.
Read → -
AI and Design Monthly Scholarly Watch (July–August 2026)
2026-08-04A monthly scholarly watch collecting 25 peer-reviewed papers and preprints on 'AI and design' from July through early August 2026. The 11 papers from ACM C&C 2026 shift the center of gravity from 'aligning output' to participation, craft, and governance, and DRS 2026 practice studies partially fill the gap on professional-skill change flagged last month. Meanwhile, new empirical evidence that AI ideation support compresses collective diversity leaves the conflict from the previous watch unresolved.
Read → -
What Is Generative Art For in Education? Seven Purpose Types and a Literature Map of 72 Studies
2026-08-02A literature map of 72 academic studies (86 explored, 14 merged as duplicates, 0 excluded) that use generative art (generative art / creative coding / computational art) as educational material, organized by the purpose of educational use. Seven purpose types emerge: (1) motivating and contextualizing introductory programming (media computation, Processing/p5.js); (2) inclusion of diverse learners (women, non-majors, low-income communities); (3) developing and assessing computational thinking (CT assessment of Scratch projects); (4) STEAM integration and constructionist making; (5) expanding expressive techniques in art, design, and music education; (6) creativity education; and (7) AI literacy and critical media literacy after LLMs (2023–2026). In terms of educational stages, CT development, STEAM, and inclusion dominate K-12; motivational introductions dominate university CS; and expressive expansion dominates art schools and higher education in architecture and music. Since 2023, 'education about writing code' has been rapidly joined by 'education using generative models as material for critical understanding' (one review reports that empirical studies of generative AI in art education grew from 2 in 2023 to 14 in 2025). All 72 items carry DOIs/URLs.
Read → -
Adversarial Review of Two Generative Art Surveys: Examining the Literature Map and the Citation Strategy Against the Design-Theory Canon
2026-08-02An adversarial examination of two notes — a generative art literature map (94 items) and a bibliometric analysis ranking citation-earning angles for a review article — through the design-theory lens (Simon/Schön/Cross/Dorst & Cross/Buchanan/Rittel & Webber/Krippendorff/Costanza-Chock). The literature map's four main weaknesses: an entanglement in which AI's ontological status (tool / co-designer / environment / evaluation apparatus) shifts from chapter to chapter; the unsorted concept of 'generation'; a sampling-inference leap that infers a disciplinary disconnect from absence in the 94-item corpus; and a compression of conceptual history that runs the generative Ästhetik of 1965 and the text-to-image of the 2020s through a single definition. The citation-strategy side's five main weaknesses: a slide from citation prediction to research value (a Schön-style critique of delegating problem-setting to the citation market); age bias in the citation-velocity metric and extrapolation to oneself; an exhaustiveness slide in which the qualified gap judgment loses its qualification in the No. 1 verdict; survivorship bias from looking only at highly cited surveys; and the implication leap of the 'cumulative 3,850-citation re-citation node.' At the same time, steelmanning established that the substance of the gap discovery (the unconnected state of computational-creativity evaluation and LLM empirical work) and the limited usefulness as circulation prediction cannot be overturned. All criticisms are tied to sources in an Evidence Ledger, and disagreements are left as a disagreement table rather than folded. The full critiques are stored in source/review/generative-art-adversarial-review/.
Read → -
Which Angle for a Generative-Art Review Paper Attracts Citations? An Empirical Analysis Using Bibliographic Data
2026-08-02This note empirically identifies, using public bibliographic data from OpenAlex and Semantic Scholar, the angles most likely to attract citations when writing a review paper on generative-art research. (1) Measured citations of 43 existing reviews show that 2023–2025 empirical studies (Zhou & Lee, 191.5 citations/year) and large technical surveys (Yang 2023, 288.2 citations/year) are fastest, while theoretical work holds a stable base of 14–30 citations per year. (2) In theme-level growth, human-AI co-creation is fastest at +607% from 2024→2025, followed by copyright at +160% and LLM creativity evaluation at +115%. (3) Highly cited surveys all include taxonomy and open-problems sections (7/7), and only the top-cited tier maintains GitHub repositories. Systematic reviews are cited more than other research designs (6.6 per year on average; JIF explains R²=0.59). (4) Verification of gap areas shows that a review connecting computational-creativity evaluation frameworks (Boden/SPECS/Lovelace) to the evaluation of LLM/diffusion-model outputs is most promising (a re-citation node for prior theory with over 3,850 cumulative citations, a 15.8-fold increase in primary studies, and the nearest review explicitly stating the connection is absent). The label-effect meta-analysis is no longer a gap, as De Rooij 2025 has already been published. Integrating these, the note ranks the seven most citation-promising angles. All figures carry sources and retrieval dates.
Read → -
The Scholarly Lineage of Generative Art: A 94-Item Literature Map from Information Aesthetics to the Post-LLM Era
2026-08-02A literature map of scholarly research on generative art, organized into four streams (94 items; 108 rows explored, 8 merged as duplicates, 0 excluded) corresponding to the chapters of a review article. (1) Historical origins: the lineage from Bense's and Moles's information aesthetics as the theoretical foundation, through the coining of generative Ästhetik at the 1965 Nees exhibition and the 'rot' magazine, the 1968 Cybernetic Serendipity and Zagreb New Tendencies, to the art world's rejection (Taylor) and the definitional debate (Galanter's autonomous-system definition). (2) Technical genealogy: the generational succession from L-systems and evolutionary computation (Sims, IEC) through GAN/CAN, StyleGAN, and diffusion models (DDPM, CLIP, Latent Diffusion) to LLM multimodal generation. (3) Theoretical frameworks: Boden's three types and their formalization (Wiggins, Jordanous's SPECS), the authorship debate (Hertzmann's tool argument versus McCormack's four concepts), and empirical reception studies showing that AI labels lower evaluations (Ragot, Bellaiche). (4) The post-LLM era (2024–2026, the thickest chapter): empirical text-to-image studies (AI adoption raises productivity +25% while average novelty declines), LLM code generation for creative coding (Spellburst, GenP5), human-AI co-creation, the arms race between copyright-protection tools and circumvention (Glaze/Nightshade vs. bypass attacks), and split empirical findings on labor effects (no short-term income decline vs. five-year longitudinal reports of job loss). Identified gaps include the disconnection between computational-creativity evaluation frameworks and post-LLM empirical research.
Read → -
The Genesis of LLM-as-a-Judge: A Confluence of Three Lineages and Four Demands
2026-07-29A genesis history of LLM-as-a-Judge, the concept of using LLMs as evaluators, organized through an extended corpus of 38 academic publications (C1–C38) and 10 vendor primary sources (P1–P10). The concept came into being in 2023 as the confluence of three lineages: (1) the crisis of NLG evaluation (the loss of validity of the surface metrics BLEU/ROUGE, the impossibility of standardizing human evaluation and its reproducibility crisis, and the turn toward learned metrics that made evaluation learnable), (2) the technical foundation of preference learning (RLHF reward models, and AI feedback via Constitutional AI and RLAIF), and (3) the normative demand of scalable oversight. Practice (GPT-4 grading in the Vicuna blog of 2023-03) preceded the term (the title phrase of the MT-Bench paper of 2023-06), and four demands pushed the concept forward: the absence of correct answers in open-ended generation, benchmark saturation and contamination, the supply-demand gap in evaluation, and the need for alignment evaluation. After its establishment, institutionalization advanced through surveys, meta-evaluation benchmarks, and documentation by public institutions (NIST AI 800-2), but the validity problem of surface metrics persists in altered form as the judge bias problem, and the reproducibility problem of human evaluation as the judge meta-evaluation problem.
Read → -
Recent Currents in LLM-as-a-Judge: A Literature Map of 53 Core Studies and a Standalone Chapter on Creativity Evaluation
2026-07-29A literature map of research on LLM-as-a-Judge (using LLMs as evaluators), based on an academic corpus of 53 items (58 explored, 5 merged as duplicates, 0 excluded) and an industry corpus of 39 items (18 T1v vendor primary sources, 9 T2 public institutions, 12 T3 practitioner opinions). On the academic side, it organizes (1) a three-generation lineage from the establishment of prompted judges (MT-Bench, G-Eval) to fine-tuned judges (Prometheus, JudgeLM) to RL-based reasoning judges (JudgeLRM, J1); (2) empirical demonstrations of position bias, verbosity bias, and self-preference bias, together with meta-evaluation benchmarks (on JudgeBench, GPT-4o performs near-randomly on hard problems); and (3) a standalone chapter on the evaluation of creative artifacts (creative writing, creativity tests, images, UI, design). In creativity evaluation, human correlation is high for formalizable rubric dimensions (AUT r=.81, metaphor r=.72), while judgments on originality dimensions and expert-quality design critique remain out of reach; this divergence is confirmed. On the industry side, productization splits into two types, managed services with general-purpose model judges and specialized fine-tuned judges, and both vendor official documents and public institutions (NIST, UK AISI, J-AISI) state calibration against human evaluation as a precondition, converging with the academic bias findings.
Read →