Notes · updated 2026-07-19
How to analyze what leaves no record
When you set out to analyze a making process, the first thing you run into is that the process itself is not at hand. The finished artifact remains. But where someone stumbled, what they tried, and where they changed course all fade unless the person writes them down.
There are two ways in to reconstruct a process after the fact: observing and recording the activity from the outside, and having the person write a retrospective account and reading it. The first calls for a design of observation; the second for a design of reflection and of how the written account is to be read.
This literature map places three lineages standing at these two entrances side by side. The critical incident technique (CIT) is the classic of the observation side, collecting outstandingly effective or ineffective events systematically. Guided reflection is a framework on the side that designs the act of reflection deliberately as a task. Reflective writing analysis is a method on the side that breaks written reflection into elements and assesses them. All three give different answers to one question: how to record and analyze the making process and reflection on it.
Introducing the methods is not enough. Research that analyzes a process rests on the very conditions that make the process possible. So the second half of this map places, as contextual support, the perspective of how to audit the quality of a literature review itself, and historical cases of how environments that cultivate an exploratory disposition have been designed.
The critical incident technique and incident analysis
The classic method for recording a process from the outside is the critical incident technique, formulated by the psychologist John Flanagan in 1954. Starting from wartime U.S. Army Air Forces pilot selection and studies of combat leader behavior, Flanagan accumulated job analyses across many occupations and afterward abstracted the shared procedural principles from them1.
At the center of the method are two definitions. An incident is any observable human activity sufficiently complete in itself to permit inferences and predictions about the person performing the act. For it to be critical, the purpose of the act must seem fairly clear to the observer, and its consequences must be definite enough to leave little doubt about its effects (Flanagan 1954, p.327). What Flanagan urges collecting is not average behavior. It is extreme cases, “outstandingly effective or outstandingly ineffective” (ibid., p.338). Extreme cases are easier to judge for their contribution to the aim, and their descriptions come out more concrete.
The critical incident technique is not a single fixed procedure. Flanagan himself states that it is “a flexible set of principles which must be modified and adapted to meet the specific situation at hand” (ibid., p.335). The whole procedure decomposes into five steps: determining the general aims, plans and specifications, collecting the data, analyzing the data, and interpreting and reporting (ibid., pp.335–346). Yet the choice of category system by which the collected incidents are classified (the frame of reference) has no theoretical basis and, Flanagan concedes, must remain inductive and subjective according to its use (ibid., p.344).
The method has a character built into its collection design. In the critical incident technique, observers are trained, what counts as important is defined in advance at the point of collection, and primary data are then newly gathered through interviews, questionnaires, or recording forms. What counts as “critical” also depends heavily on a binary judgment of success versus failure. So the method transfers cleanly only to research where the criteria can be controlled before collection. Where already-written, unstructured text is analyzed after the fact, the premise of the collection design does not match, and the design has to be rebuilt even while borrowing part of the principle.
The side that designs reflection
Where the critical incident technique records action from the outside, education has accumulated designs on the side that has people write their reflection. In 2009, Sarah Ash and Patti Clayton presented a design theory of reflection known as the DEAL model2.
The starting point is that experience itself is not a good teacher. Reflecting on it is what produces learning, and reflection that is not designed can become “hazardous, accidental, and superficial” (Ash & Clayton 2009, p.27, citing Stanton 1990). Ash and Clayton frame critical reflection as having three functions: it generates, deepens, and documents learning (ibid., p.27). On that basis they argue that effective reflection follows three principles. Decide the desired learning outcomes first, design the reflection to achieve those outcomes, and integrate formative and summative assessment into the reflective process (ibid., p.28).
The DEAL model gives this principle a concrete structure. First describe the experience objectively and in detail (Describe), then examine that experience in light of specific learning goals (Examine), and finally articulate what was learned, how, why it matters, and what to do next (Articulate Learning) (ibid., p.41). The quality of reflection is measured on two independent scales. One is the level of reasoning by Bloom’s taxonomy (six levels), that is, how deeply one has thought. The other is the quality of reasoning by critical thinking standards (clarity, accuracy, relevance, and so on), that is, how accurate and fair it is (ibid., p.32, p.37).
The framework has a premise. DEAL takes as its object of analysis a response to a guiding prompt that the instructor has designed in advance. The unit of analysis is always “a student’s response to a reflection task the instructor designed,” never free writing produced without a structured prompt. Ash and Clayton themselves state that this paper is a normative design proposal not accompanied by empirical validation (ibid., p.28). Its standards for measuring the quality of reflection are also borrowed from other fields, educational-objective taxonomy and critical-thinking education, and no validation specific to DEAL is given within the paper.
How to read written reflection
Even with a framework for designing reflection in place, someone must judge the quality of the written account. That judgment has long depended on hand coding by content analysis. It is work that costs both time and resources. In 2019, Thomas Ullmann tested how far this judgment can be automated3.
Ullmann synthesized 24 existing reflection models into a system of eight categories for detecting reflection (Ullmann 2019, pp.221–222). The system splits into two dimensions. The depth dimension is a binary judgment of whether a sentence is reflective or descriptive. The breadth dimension looks at which of seven elements a sentence contains: experience, feeling, belief, difficulty, perspective, learning, and intention.
The validation proceeded in two stages. First, 76 student essays across health, business, and engineering, totaling 5,080 sentences, were annotated sentence by sentence by crowd workers, and inter-rater reliability was confirmed for all eight categories (Cohen’s κ = 0.78–0.98, ibid., p.235). Then a machine learning classifier was trained on 80 percent of these and tested on the remaining 20 percent; five of the eight categories could be detected automatically with high reliability and the remaining three with moderate reliability (κ = 0.53–0.85, ibid., p.217). Automated analysis was on average about 10 percent less accurate than hand coding. Even so, the co-occurrence of reflection (depth) with each of the seven breadth categories was statistically significant throughout, empirically supporting the construct the system aims to measure (ibid., pp.235–236).
The character of this method lies in judging the presence or absence of elements. The eight categories each check for existence independently, rather than judging any temporal linkage among the elements. Ullmann also names limitations. The corpus is limited to English-language academic student essays, and generalization to other languages or to other kinds of text such as blogs or dialogue records is unvalidated (ibid., p.248). Because the unit of analysis is fixed at the sentence, reliability at a larger unit such as whole-text has not been checked (ibid., p.248). A separate lineage grades the depth of reflection on a single ordinal scale, but that differs in conception from Ullmann’s element detection4.
Setting the three methods side by side reveals that, under the same banner of “analyzing process and reflection,” the collection designs and units of analysis diverge. The critical incident technique fixes criteria at the point of collection and gathers new primary data. DEAL takes as its object responses to instructor-designed prompts. Ullmann’s system breaks written reflection into elements and judges their presence. Which method to choose is settled by what kind of record is at hand and what one wants to ask.
Process data and the recording environment
All three methods handle written text or reports. Against this stands a lineage that records the making process directly as a time-ordered operation log. With a dedicated recording environment that captures operation histories and dialogue logs, the actual process can be traced without relying on later recall.
The weakness of relying on recall was noted from the critical incident side as well. Flanagan reports that recalled incident reports drop out substantially compared with daily records. In one case, about 80 percent of observed events were lost in recall two weeks later (Flanagan 1954, p.331). Whether a recording environment can be provided governs how far this dropout can be prevented.
Yet the presence or absence of a recording environment determines the object of analysis itself. Research with a dedicated log environment can analyze the actual time-ordered process. In ordinary classes or making settings where no log exists, one can only analyze after-the-fact self-reports or existing accounts. This difference is not one of merit but of conditions: how much record an environment leaves behind.
Quality assurance of the literature review
Even after choosing a method and recording the process, one more inspection remains. It is a basis for auditing, non-arbitrarily, how a study has drawn on prior work. What is needed is not the impressionistic demand to “add more references” but a perspective grounded in established methodology.
A widely used framework for scoring dissertation literature reviews is the rubric David Boote and Penny Beile presented in 20055. They set as scoring axes whether the review justifies its inclusion and exclusion criteria (Coverage), distinguishes what has been done from what should be done and situates the work in a broad scholarly context (Synthesis), connects methodology with theory (Methodology), and justifies its practical and scholarly significance (Significance). On the systematic-review side, Andrea Tricco and colleagues presented the scoping-review reporting standard PRISMA-ScR in 20186. It requires, as 20 essential items, making the search scope, search strategy, and inclusion criteria explicit and reproducible.
Techniques for mechanically catching what has been missed have also been formalized. In 2014, Claes Wohlin organized as guidelines for systematic reviews the backward snowballing that traces the reference lists of adopted works, and the forward snowballing that traces later works citing them7. A stopping criterion for how much is enough has also been discussed. Theoretical saturation regards as the point of saturation the moment when adding new literature no longer changes theoretical understanding, judging comprehensiveness not by “all the literature” but by “the representativeness of the constructs” (Saunders et al. 2018)8.
The review has pitfalls that quality assurance alone cannot capture. In 2011, Jörgen Sandberg and Mats Alvesson analyzed how research questions are constructed and pointed out that the most common type is gap-spotting, finding a gap in the literature and filling it9. But influential theory, they argued, tends to arise from problematization, questioning the assumptions of existing literature, whereas gap-filling stays mediocre because it preserves the assumptions. This distinction bears on choosing a method for analyzing process too. Is one merely filling a gap in existing methods, or questioning the very assumption those methods rest on? Placing the critical incident technique, DEAL, and Ullmann’s system side by side was also a way to make the differences in their assumptions visible.
Historical cases of environments that promote exploration
Even with a method for analyzing process and an audit for reviews in place, unless there is a place where people can try, stumble, and start over, the process to be analyzed never arises. How have environments that cultivate an exploratory disposition been designed? For historical reference points, there is one case each from design education history and from organizational institutions.
The Bauhaus preliminary course
The case cited repeatedly in design education history is the Bauhaus preliminary course (Vorkurs), which ran from 1919 to 193310. A required foundation course taken by all students regardless of specialization, it was taught in turn by Johannes Itten (1919–1923), László Moholy-Nagy (1923–1928), and Josef Albers (from 1928), each bringing a different pedagogical philosophy.
The course is said to have placed practice before theory. Students engaged directly with materials such as paper, metal, cloth, and light, folding, gluing, and assembling them, learning through experience the properties each material holds. The teacher’s role is described as designing an environment for students to discover on their own, rather than drilling in styles and techniques. This design ethos, treating classroom, furniture, tools, and technology themselves as integral to pedagogy, is said to have carried over to Black Mountain College, where Albers moved, and to later sites of design education.
Citing the Bauhaus as “an environment that cultivated exploration” calls for several reservations.
Admission to the Bauhaus was not open to anyone; there was a selection process, and it was limited to those who could bear the tuition.
Women students were admitted, but opportunity was skewed by gender, with many concentrated in the weaving workshop.
The telling leans toward success stories that produced figures like Paul Klee and Wassily Kandinsky, and records of students who dropped out or could not adapt are scarce.
And the Bauhaus was closed by the Nazis in 1933.
That closure shows that an environment cultivating exploration is fragile, politically and economically alike.
The primary sources for this course (Wingler 1969, Itten 1964, Moholy-Nagy 1947) and the pages of direct quotations are unverified in this note [requires source verification].
Organizational slack-time programs
The organizational institutions cited are 3M’s 15% rule and Google’s 20% time11. Both are introduced as programs that let employees devote part of their paid working time to projects of their own choosing. 3M’s program is said to have been introduced in the 1940s by the then executive William McKnight, and the anecdote that the Post-it note was born from this time is repeated often. Google’s program was announced in its 2004 IPO letter, and Gmail and Google Maps are frequently said to have come out of this 20%.
These cases cannot be handled with academic rigor.
The sources in this area are not peer-reviewed papers but company announcements and business-media articles, and no controlled experiments or sample-based empirical studies are found [requires source verification].
The telling leans toward success stories, and data on projects that spent the time and produced nothing are not published.
The causal attribution that a product “was born thanks to slack time” rests on anecdotal evidence and is hard to falsify.
That the creator of Gmail is reported to have said “that was not a 20% project” indicates the fragility of this causal attribution [requires source verification].
There is also a gap between the existence of the program and its actual use.
At Google, only about 10 percent of engineers are reported to have used 20% time consistently [requires source verification].
Even where the program exists, under an evaluation structure that prioritizes the primary job, using slack time can become effectively disadvantageous.
The reported criticism to the effect that “the real 20% time is 120% time” points to the way the program worked as a burden stacked on top of the main job [requires source verification].
In other words, guaranteeing time as an institution alone does not promote exploration; it works only when the evaluation structure and organizational culture align, a reading that must carry conditions.
The slack-time cases show the exploration conditions of elite workers, not conditions open to all workers. That highly educated, highly skilled employees can explore on the basis of stable employment is one thing; to whom that opportunity is distributed is another. The same point as the selectivity of Bauhaus access appears on the organizational side as well.
The question this map leaves open
The three methodologies divide the kinds of record they can handle by their differences in collection design and unit of analysis. None is the single correct one; what kind of trace of the process is at hand narrows the choosable methods first. Research that can design a recording environment analyzes the time-ordered process; research with only existing accounts analyzes after-the-fact self-reports, each under its own constraints.
The historical cases on the environment side showed that conditions cultivating exploration can be designed and, at the same time, are selective in access and fragile both politically and economically. The Bauhaus’s admissions selection and closure, and the low usage rate of organizational slack time, tell in separate scenes that in speaking of the conditions that make exploration possible, one cannot leave unasked to whom those conditions are open.
The methodological question of how to record and analyze process and the environmental question of how to make exploration possible look separate but are continuous. The process to be analyzed does not arise without an environment that permits it. This map only places the two side by side; under which epistemology to bind them into a single research program remains on the side of the methodological tension addressed in design-research-methodology-mainstream and qualitative-quantitative-design-research.
References
Unverified items
- The primary sources for the Bauhaus preliminary course (Wingler 1969, Itten 1964, Moholy-Nagy 1947) and the pages of direct quotations are unverified; the account rests on secondary sources
[requires source verification]. - The figures on 3M’s 15% rule and Google’s 20% time (roughly 30% of annual revenue from new products, about 10% usage, the “120% time” criticism, and the creator’s testimony about Gmail’s origin) derive from company announcements and business media, lack peer-reviewed verification, and are unverified against primary sources
[requires source verification]. - Kember et al. (2008) is confirmed at the bibliographic level only; the full text has not been read closely
[requires source verification].
Footnotes
-
Flanagan, J. C. (1954). The Critical Incident Technique. Psychological Bulletin, 51(4), 327–358. (public domain) ↩
-
Ash, S. L., & Clayton, P. H. (2009). Generating, Deepening, and Documenting Learning: The Power of Critical Reflection in Applied Learning. Journal of Applied Learning in Higher Education, 1, 25–48. https://files.eric.ed.gov/fulltext/EJ1188550.pdf ↩
-
Ullmann, T. D. (2019). Automated Analysis of Reflection in Writing: Validating Machine Learning Approaches. International Journal of Artificial Intelligence in Education, 29(2), 217–257. https://doi.org/10.1007/s40593-019-00174-2 ↩
-
Kember, D., McKay, J., Sinclair, K., & Wong, F. K. Y. (2008). A four-category scheme for coding and assessing the level of reflection in written work. Assessment & Evaluation in Higher Education, 33(4), 369–379. https://doi.org/10.1080/02602930701293355 Bibliographic details only; the full text has not been read closely
[requires source verification]. ↩ -
Boote, D. N., & Beile, P. (2005). Scholars Before Researchers: On the Centrality of the Dissertation Literature Review in Research Preparation. Educational Researcher, 34(6), 3–15. https://doi.org/10.3102/0013189X034006003 ↩
-
Tricco, A. C., et al. (2018). PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and Explanation. Annals of Internal Medicine, 169(7), 467–473. https://doi.org/10.7326/M18-0850 ↩
-
Wohlin, C. (2014). Guidelines for snowballing in systematic literature studies and a replication in software engineering. Proceedings of the 18th International Conference on Evaluation and Assessment in Software Engineering (EASE ‘14). https://doi.org/10.1145/2601248.2601268 ↩
-
Saunders, B., et al. (2018). Saturation in qualitative research: exploring its conceptualization and operationalization. Quality & Quantity, 52, 1893–1907. https://doi.org/10.1007/s11135-017-0574-8 ↩
-
Sandberg, J., & Alvesson, M. (2011). Ways of constructing research questions: gap-spotting or problematization? Organization, 18(1), 23–44. https://doi.org/10.1177/1350508410372151 ↩
-
For the Bauhaus preliminary course, the main secondary references are Wingler, H. (1969). The Bauhaus: Weimar, Dessau, Berlin, Chicago. MIT Press; Itten, J. (1964). Design and Form: The Basic Course at the Bauhaus. Reinhold; and Moholy-Nagy, L. (1947). Vision in Motion. Paul Theobald. This note has not directly verified the pages of the originals, and the account rests on secondary sources
[requires source verification]. ↩ -
The account of 3M’s 15% rule and Google’s 20% time rests on secondary information synthesizing company announcements and business-media articles. Peer-reviewed empirical studies have not been confirmed, and the figures and anecdotes are all unverified against primary sources
[requires source verification]. Candidate primary sources include Page, L., & Brin, S. (2004). Google IPO Letter; and Bock, L. (2015). Work Rules! Twelve. ↩