Notes — Research Log
Notes.
A log of research, experiments, and reflections. Ongoing investigations into AI applications and design processes, accumulated and published in an LLM Wiki format.
Want to ask about the notes? Ask the Wiki →
-
Recent Currents in LLM-as-a-Judge: A Literature Map of 53 Core Studies and a Standalone Chapter on Creativity Evaluation
2026-07-29A literature map of research on LLM-as-a-Judge (using LLMs as evaluators), based on an academic corpus of 53 items (58 explored, 5 merged as duplicates, 0 excluded) and an industry corpus of 39 items (18 T1v vendor primary sources, 9 T2 public institutions, 12 T3 practitioner opinions). On the academic side, it organizes (1) a three-generation lineage from the establishment of prompted judges (MT-Bench, G-Eval) to fine-tuned judges (Prometheus, JudgeLM) to RL-based reasoning judges (JudgeLRM, J1); (2) empirical demonstrations of position bias, verbosity bias, and self-preference bias, together with meta-evaluation benchmarks (on JudgeBench, GPT-4o performs near-randomly on hard problems); and (3) a standalone chapter on the evaluation of creative artifacts (creative writing, creativity tests, images, UI, design). In creativity evaluation, human correlation is high for formalizable rubric dimensions (AUT r=.81, metaphor r=.72), while judgments on originality dimensions and expert-quality design critique remain out of reach; this divergence is confirmed. On the industry side, productization splits into two types, managed services with general-purpose model judges and specialized fine-tuned judges, and both vendor official documents and public institutions (NIST, UK AISI, J-AISI) state calibration against human evaluation as a precondition, converging with the academic bias findings.
Read → -
How Novelty Is Made and How It Is Claimed: A Typology of Establishing Strategies
2026-07-28A literature note that collects, across fields, the strategies for establishing novelty, starting from the question of what operations can actually be moved in response to a demand to produce novelty. It organizes 243 items across eight traditions: problematization research in management studies, heuristics of discovery in the philosophy of science, scientometrics, rhetoric and genre analysis, contribution taxonomies in HCI and design research, the extension and falsification of theory, analogy and boundary crossing, and idea generation in the age of LLMs. It first establishes that strategies split into two kinds—generation (making novelty) and claiming (asserting novelty)—and that the two have been conflated. On the generation side it shows that the assumptions open to challenge divide into five layers (Alvesson & Sandberg), that naming one's ignorance and strategically choosing one's materials are different strategies (Merton's specified ignorance and strategic research materials), that mechanism discovery reduces to decomposition and localization (Bechtel & Richardson), that there is a path by which a tool becomes a theory outright (Gigerenzer), and that atypical combinations work only when built on a base of high conventionality (Uzzi et al.). On the claiming side it shows that Move 2 of CARS is exhausted by four steps—counter-claiming, gap-indicating, question-raising, and continuing a tradition (Swales)—and that hype words have roughly doubled over fifty years and appear in about 97% of successful NIH applications (Hyland & Jiang; Millar et al.). Few strategies have had their effects verified quantitatively, and many of those that have collide with the inverted-U preference of evaluators (Criscuolo et al.) and the distance penalty (Boudreau et al.). 96 references.
Read → -
Where Do Novel Research Questions Come From? The Four Loci and Their Combinations, Examined Against the Literature
2026-07-28A literature note that examines a practitioner's hypothesis—that the only places where a paper's novelty can reside are Related Work, Problem Definition, Materials, Methods, and Results, and that securing two or more of the four loci left after Related Work makes a paper strong—against 135 items across five traditions: problematization research in management studies, scientometrics, philosophy and sociology of science, contribution taxonomies in HCI and design research, and genre analysis of academic English. No scholarly formulation that normatively maps IMRaD sections onto types of contribution was found, so the note first establishes that the four-locus classification stands on its own as practitioner knowledge. It then shows that each locus has an independent lineage (Problem Definition = problematization and specified ignorance; Materials = strategic research materials and experimental systems; Methods = tools-to-theories), that Related Work is not a locus of generation but a locus of claiming corresponding to Move 2 of CARS, and that at the Results locus a design that wins by scale cannot be conflated with serendipity. On the prescription of two or more loci, it sets Uzzi's atypical combinations, Shi and Evans's contents and contexts, and Arts and Veugelers's familiarity and novelty on the supporting side alongside Boudreau's distance penalty, Criscuolo's preference for the moderate, Johnson and Proudfoot's variance in evaluation, and Zhao et al.'s 2026 disconfirming measurement (results-based novelty alone is cited more than holding all three dimensions), and makes explicit, as a gap, that no peer-reviewed evidence directly measuring combinations of loci exists. 78 references.
Read → -
An Academic Map of Research That Treats Hallucination as a Resource for Creativity
2026-07-21A literature map that surveys peer-reviewed research and preprints positioning generative-AI hallucination (confabulation) not as a defect to be suppressed but as a starting point for ideation and creativity. It connects, cluster by cluster: the positive reframing of hallucination (the value of confabulation in Sui et al., valuable hallucination in Chen & Wang, the creativity-perspective survey by Jiang et al., the decoding-layer tradeoff in He et al., the usefulness in drug discovery in Yuan et al.); LLM creativity evaluation and output diversity (Stevenson et al., Bellemare-Pepin et al., the homogeneity of Wenger & Kenett, diversity reduction under RLHF in Kirk et al., diversity collapse at the SFT stage in Karouzos et al., min-p sampling in Nguyen et al.); ideation support and design fixation (AI-augmented brainwriting in Shaer et al., design fixation in Wadinambiarachchi et al., CST in Shneiderman); scientific hypothesis generation (Shahhosseini et al.); idea homogenization and prompt-driven diversity recovery (Girotra et al., Meincke et al.); and the classics of creativity psychology (Guilford, Mednick, Beaty & Kenett). A neutral prior-work review. Thirty-one references (bibliographies confirmed reachable; caveats aggregated at the end). The view that treats the deviations of older, lower-precision models as a resource is thinly represented in explicit research, and is treated here as an implication drawn from research on the performance/diversity tradeoff.
Read → -
Gaps in AI Research, a Fourth Time: Seen from Outside the Inversion Family
2026-07-21An essay that rereads the three sibling notes that searched for gaps in AI research (the seven frames, the A–G operational sheet, the game board of Problem Definition) from the vantage of the full set of problem-setting methods now assembled. It shows that all three generators belong to a single family of operations—binary inversion—and that they had been converging on that family's fixed points (an endogenous scaling law, the institution of recording). On that basis it applies operations that lie outside the family. Morphological analysis (GMA) makes the board's constants orthogonal as dimensions and yields an unvisited intersection cell (a single system that is recursive and ecological and longitudinal at once); TRIZ turns the incompatibilities that cannot be inverted (cultural diversity versus universal human rights; individual novelty versus collective homogenization) into design problems solved by separation; PSM/CATWOE restores the demoted cognitive diversity and non-Western epistemologies as 'another board' and points to the absence of a method for adjudicating the plurality of boards; and Cynefin and representational change diagnose the infinite regress all three fall into not as a defect but as a misclassification of the type of problem situation. On the single point that convergence may be a property of the family of operations more than of the object, it weakens the earlier three notes' argument from 'the convergence of independent generators.' The three gaps that the operations outside the family pointed to were checked against the coverage of recent (2024–2026) prior work, and each was judged partial (a residue can be narrowed against a dense body of prior work). References augmented across method, evidence, and coverage-confirmation studies.
Read → -
A Genealogy of Practitioner Methods for Reframing Problems: Who Made Them, Traced to Their Origins
2026-07-21An industry-side map of practitioner methods for problem setting, problem framing, and reframing, narrowed to their originators, primary sources, and standing in practice. It covers convergence/divergence process methods (Design Council's Double Diamond, IDEO/d.school's Design Thinking, GV's Design Sprint), question-posing/reframing methods (How Might We, de Bono's lateral thinking and Six Thinking Hats, First Principles, SCAMPER), structured-decomposition methods (TRIZ, morphological analysis, Toyota's Five Whys, McKinsey's seven-step process, Conn & McLean's Bulletproof, Minto's SCQA), and the three lineages of Jobs-to-be-Done (Ulwick/Christensen/Moesta). Where a method's origin is disputed, the note lists the competing accounts rather than asserting one, and treats promotional effectiveness figures as partial with their methodology attached. The gaps between academia and industry (IDEO is a propagator, not the originator, of HMW; First Principles offers no decision criterion; GMA and TRIZ have not penetrated practice, etc.) are organized at the end. A neutral practitioner review. References carry primary sources and official explanations.
Read → -
An Academic Map of Methods for Reframing Problems: From Abduction-2 to Problem Structuring
2026-07-21A literature map that surveys the methods of problem setting, problem framing, and reframing across peer-reviewed research and canon from design methodology, cognitive science, and systems thinking. Starting from Dorst's abduction-2 (frame creation), it connects, by lineage: functional fixedness and representational change in cognitive science (Duncker, Ohlsson, Kaplan & Simon, Chi); the psychology of problem finding and problem construction (Getzels & Csikszentmihalyi, Reiter-Palmon, Mumford); C-K theory and problem-solution co-evolution (Hatchuel & Weil, Maher & Tang, Wiltschnig); wicked problems and problem structuring methods (Rittel & Webber, Checkland SSM, Eden SODA, Rosenhead PSM, Ackoff); Cynefin (Kurtz & Snowden); morphological analysis (Zwicky, Ritchey); the theorization of TRIZ (Altshuller, Cavallucci); lateral thinking (de Bono); and empirical work on reframing in design (Valkenburg & Dorst, Paton & Dorst, Stompff). A neutral prior-work review. 38 references (bibliography reachability-confirmed and ISBNs fixed; remaining reservations collected at the end).
Read → -
AI Research Gaps, a Third Time: Rereading Them as Rewritings of the Game Board
2026-07-20An essay that rereads the two sister notes that searched for gaps in AI research (the catalogue of seven frames, and the revisit via the A-through-G worksheet) with a third generator: Problem Definition (operations on the game board). It writes out AI research's game board, with sources, in seven entries (certification of success, measurement, unit, endpoint, position of observation, noise, record-keeping institution) and selects the three most unnatural asymmetries from what was pushed off the board. It maps each of the seven frames to the board operation it was, showing that symmetry inversion is one species of board rewriting (an operation specialized to an asymmetric pair). It retries the two candidates demoted in the revisit note (cognitive diversity and non-Western epistemologies, and the self-mapping of blind spots), makes explicit that the demotion was a certification by the A-through-G worksheet's own success-certifying apparatus (falsifiability and the minimal observable), and reinstates them as questions strong under a different type while upholding the determination of kind. It confirms that applying the conversion table to scaling laws comes out at the same void as the endogenous scaling law, taking the convergence of independent generators as weak evidence that the void is not an artifact. Finally it generates two problem statements and RQs whose units are the success-certifying apparatus and the record-keeping institution (both gap candidates whose coverage by prior work is unchecked), and closes with the division of labor between generators (Problem Definition generates; A-through-G vets). 17 references.
Read → -
Thinking in Problem Definition: Rereading the Game Board of Existing Research
2026-07-20Relocates the strength of research from 'the ability to produce good answers' to 'the ability to see through what is being counted as a problem,' and organizes the craft of Problem Definition, which operates on the observation apparatus, evaluation axes, time horizon, units, and boundaries (the game board) that existing research implicitly fixes. Presents ten patterns easy to transfer, a 14-row meta-conversion table that maps existing research onto questions, a five-step generation protocol that lowers a description of the board into a problem statement and an RQ, and five habits of thought, connecting them to the lineage of the contrast between gap-spotting and problematization (Sandberg & Alvesson), problem setting (Schön), problem finding (Getzels & Csikszentmihalyi), ill-structured problems (Simon), and wicked problems (Rittel & Webber). 10 references.
Read → -
AI Research Gaps, Revisited: Sorting Symmetry-Completion from Coverage, and Digging Out an Endogenous Scaling Law
2026-07-20A revisit of the seven frames catalogued earlier, using a sharpened discipline (a procedure that lowers a premise, inverts a symmetry, and operationalizes it across an A-through-G worksheet). The same discipline is turned on its own past product to select, strengthen, and generate blind spots. Selection yields three groups: the single strongest candidate whose every slot fills (the relational ontology of capability), four that survive as symmetry-completion frames, and two that fail the symmetry test and are demoted (cognitively diverse users with non-Western epistemologies, and self-mapping of blind spots). Demotion is not disqualification but a determination of kind: legitimate fairness-and-coverage concerns that are not symmetry-completion blind spots. The discipline then generates one new blind spot. It inverts the open-loop assumption that a scaling law treats its input distribution as exogenously fixed, and formalizes a closed loop in which a model's capability produces its future data, an endogenous scaling law. Novelty is limited strictly. Model collapse, self-consuming loops, and scaling laws parameterized by the synthetic-data fraction are all prior work; what remains new is only inverting the exogeneity of the input distribution itself, so the data fraction becomes an endogenous state variable that closes the loop. 24 references (21 with confirmed bibliography, 3 flagged for primary verification).
Read →