Notes · updated 2026-10-10
Does Prototyping Widen Design Exploration? Fixation on Early Examples, Parallel Prototyping, Rapid Prototyping with Generative AI, and Documentation in Research through Design
This note examines whether making prototypes widens or narrows design exploration, drawing on 49 academic sources on prototyping, design fixation, prototyping with generative AI, tools for exploring design spaces, and process documentation in Research through Design (RtD).
Contents (13)
- Does making early lead to finding early?
- Scope and method
- What are prototypes made for?
- What improves when designers make and test?
- Held by the first example
- Does making reduce fixation or increase it?
- What holds designers when they prototype with generative AI?
- Wider for each person, narrower for the group
- What supports the operations that move a design space?
- What does externalizing the process make into research?
- Does prototyping with generative AI leave a process to document?
- How should the benefits and harms be read?
- Gaps
Does making early lead to finding early?
Design practice often repeats the advice to make early and fail early. Dow, Heddleston, and Klemmer (2009) cite it in the form of a design adage: enlightened trial and error outperforms the planning of flawless intellect. Making something and trying it reveals problems that remain invisible in the designer’s head. Many practitioners know this from experience.
However, whatever is made also becomes an example. Since 1991, experiments have studied design fixation, the tendency of designers who have seen an example to repeat its features. If making widens exploration, being held by what has been made may also narrow it. The question is which conditions decide between the two.
This question takes a different form now that generative AI has nearly removed the effort of prototyping. A single line of text produces an image or a working user interface within seconds. If designers held on to one idea because making it had been costly, that attachment should disappear when the cost disappears. If it does not, something other than effort is holding designers in place.
Experiments on the benefits of prototyping, experiments on fixation, experiments on prototyping with generative AI, and studies of tools for exploring design spaces each answer this question from a different angle. The literature on Research through Design (RtD) is also needed. RtD has used making as a method of research and has treated the documentation of the making process as a basis for knowledge. If generative AI skips the middle of that process, the basis is also affected.
Scope and method
- Scope: peer-reviewed journal articles and peer-reviewed conference papers on the benefits and fidelity of prototyping, design fixation and sunk cost, prototyping with generative AI and its effects on fixation and diversity, tools that support design space exploration, and evaluation criteria and process documentation in RtD. The period covers 1991 to 2026, with generative AI studies from 2023 to 2026.
- Collection: 98 candidates were gathered across academic databases (Crossref, OpenAlex, arXiv). Retractions, corrections, and predatory venues were checked through Crossref (none found), and duplicates and off-topic items were removed. Studies released on arXiv were replaced by their peer-reviewed versions (ASME IDETC, CHIWORK, SEMISH).
- Verification: For 34 of the 49 references, the full text was obtained and numbers and claims were checked against the relevant passages. Two (Jansson and Smith 1991; Youmans 2011) could not be obtained and are reported through their citation in other reviews. Twelve are reported within the scope of their abstracts and attributed to their authors. For one (Meincke et al. 2025), only the opening paragraph of a subscription page was checked.
- Presentation: For each study, the note separates the participants and what they made, the conditions compared, and the direction of the effect. Effect sizes are reported only where the numbers were checked in the full text.
What are prototypes made for?
Lim, Stolterman, and Tenenberg (2008) described prototypes from two sides. A prototype is a filter that isolates the qualities of a design idea that the designer wants to examine and leaves out the rest. It is also a manifestation that gives the idea form through choices of material, resolution, and scope. From this view, the authors derived an economic principle: the best prototype is the one that shows the possibilities and limitations of an idea in the simplest and most efficient way. Their position is that a prototype is a means of traversing a design space before it is a means of evaluation.
Buchenau and Fulton Suri (2000) discussed experience prototyping using commercial projects at IDEO and identified three uses: understanding existing experiences, exploring design ideas, and communicating design concepts. Camburn et al. (2017) reviewed prototyping studies from engineering, management, and architecture and summarized the findings in a table. That table lists both “fast prototyping reduces fixation” and “feedback may induce corrections but also increase fixation.” The benefits and harms of prototyping already sit next to each other in the same review.
For fidelity, there is evidence that low fidelity is sufficient for some purposes. According to the abstract of Virzi, Sokolov, and Karis (1996), two experiments found substantially the same usability problems with paper prototypes and with high-fidelity prototypes or actual products. If the goal is only to find problems, elaborate prototypes are not required.
What improves when designers make and test?
The benefits of making and testing were examined experimentally in a series of studies by Klemmer and colleagues at Stanford. Dow, Heddleston, and Klemmer (2009) compared participants who could test and revise a vessel that protects a raw egg from a fall with participants who could only build. Even under tight time constraints, the vessels of the iterating participants protected the egg from greater heights. Iterating participants without prior experience of the task performed as well as non-iterating participants with prior experience.
The next question is how many designs to make, and in what order. Dow et al. (2010) asked novices to design web advertisements and compared a serial condition, in which participants received critique after each prototype, with a parallel condition, in which participants created several prototypes before receiving critique. According to the abstract of the published version, the parallel ads outperformed the serial ads in click-through data and expert ratings, independent raters judged the parallel prototypes to be more diverse, and parallel participants reported a larger increase in task-specific self-confidence. The abstract also states the premise that iteration improves ideas but can also produce fixation, continuously refining one option without considering others.
Dow et al. (2011) examined the stage at which designs are shown to others. In a study with 84 participants working in pairs, they compared pairs who shared three ads, pairs who shared the best of three, and pairs who made and shared one ad. Pairs who shared three ads did better in outcome, exploration, and rapport, incorporated more of their partner’s ideas into their own designs, and achieved a higher click-through rate (χ2 = 4.72, p < 0.05). Neeley et al. (2013) gave 39 engineering students a design-and-assemble task and compared participants who made five prototypes in the first stage with participants who made one. Participants who made five prototypes reported greater time pressure and dissatisfaction, but their prototypes performed better and improved more between iterations.
The number of designs also matters for the people who evaluate them. Tohidi, Buxton, Baecker, and Sellen (2006) asked 48 participants to use paper prototypes and compared participants who saw one design with participants who saw three designs side by side. Participants who saw a single design gave it significantly higher ratings and were more reluctant to criticize it than participants who saw the same design among three. Tohidi et al. concluded that presenting alternatives makes ratings less prone to inflation and gives rise to more and stronger criticism. However, they also reported that usability testing elicited few suggestions for improvement in either condition, describing it as a means to identify problems rather than to provide solutions.
Across these experiments, results improved when designers delayed narrowing down to one design and compared several designs. The number of designs made and the timing of narrowing them down distinguished the outcomes more than making itself.
Held by the first example
Fixation research began by asking what happens when designers are shown an example. Jansson and Smith (1991) asked students to design a car-mounted bicycle rack, a measuring cup for blind people, and a spill-proof coffee cup, and showed some participants pictures of existing solutions. According to the summary by Crilly and Cardoso (2017), participants who saw the pictures repeated their features, even when instructed to avoid particular features. Jansson and Smith defined this as “a blind adherence to a set of ideas or concepts limiting the output of conceptual design.”
Expertise does not remove fixation. Linsey et al. (2010) assigned engineering design faculty to a control group, a group shown an example, and a group shown an example together with materials to mitigate fixation. The faculty also fixated on the example, and they only partially perceived its effects. Providing mitigation materials reduced fixation.
Fixation research has expanded, but its results are not uniform. Vasconcelos and Crilly (2016) reviewed the experimental studies and showed that results vary with the modality, fidelity, number, and problem proximity of examples, the timing of exposure, and the experience of participants, and that methods and measures differ from study to study. On fidelity, they summarized the following studies. Industrial designers who saw real photographs (high fidelity) produced less novel ideas than those who saw line drawings (low fidelity). Designers who saw only partial photographs of products developed more original solutions than those who saw the full photographs. These results suggest a tendency for designers to depart less from examples that look more complete.
The assumption that fixation is always harmful has also been questioned. Crilly and Cardoso (2017) redefined fixation as a state in which someone engaged in a design task undertakes a restricted exploration of the design space because of an unconscious bias from prior experience, knowledge, or assumptions, and reframed it as a trade-off between flexibility and commitment. Their research agenda asks how much sticking with an idea counts as “too much” and what not sticking with an idea “enough” would look like. It also asks whether a narrow but deep exploration can become the foundation for a broader one, and whether fixation can occur at the level of a group or community.
Does making reduce fixation or increase it?
In most fixation experiments, participants sketch ideas but do not build and test them. In their section on testing ideas, Vasconcelos and Crilly (2016) showed that studies of building and testing disagree. Youmans (2011) reported that building physical models and testing them against requirements made original and useful solutions more likely and reduced fixation. By contrast, other studies reported that exposure to prototypes inhibits distant analogies and that designers become attached to initial ideas in which they have invested effort.
Two experiments by Viswanathan and Linsey offered one explanation for this disagreement. According to the abstract of Viswanathan and Linsey (2012), two experiments with novice designers found that building physical models increased the share of ideas that satisfied all requirements, did not change novelty or variety, and produced no evidence of fixation. This result contradicted earlier observational studies that had reported fixation. Viswanathan and Linsey (2013) then compared conditions that included building with low-cost materials, building with high-cost materials, and sketching only. According to the abstract, the more time and effort building required, the lower the novelty and variety of ideas, and the authors concluded that the sunk cost of building, rather than physical models themselves, produces fixation. They recommended building early models with materials that require minimal time, cost, and effort.
Vasconcelos and Crilly (2016) noted that the sunk-cost explanation is closer to designers’ reluctance to change their own ideas, related to psychological ownership, than to fixation on a supplied example. This reading suggests at least two routes to fixation. One route runs through the examples that designers see. The other runs through ideas in which designers have invested effort and that they are reluctant to give up.
The timing of building also matters. Jang and Schunn (2012) observed 43 engineering design teams over a semester. According to the abstract, successful teams used physical prototypes consistently throughout the design process, whereas unsuccessful teams adopted physical prototypes late.
What holds designers when they prototype with generative AI?
Prototyping with generative AI reduces the effort of making to a few seconds. Under the sunk-cost explanation, attachment to one’s own idea should weaken when that effort disappears. The first experiment to test this directly produced the opposite result.
Wadinambiarachchi, Kelly, Pareek, Zhou, and Velloso (2024) asked 60 participants with experience in visual work to sketch ideas for a chatbot avatar. All participants saw an example avatar, and they were assigned to one of three conditions: no support, Google Image Search, or the image generator Midjourney. In their Bayesian model, the Midjourney group repeated the most features of the example, and the probability that its fixation exceeded that of the no-support group was estimated at 100% (98% for the image-search group). Fluency, variety, and originality tended to be lower with either form of support than without support. For variety, the effect of image search almost disappeared once the number of sketches was taken into account, whereas a negative effect remained only for the AI group, with weak evidence (Bayes factor 2.64). Of the images Midjourney produced, 44% (206 of 468) portrayed humanoid robots similar to the example avatar.
A notable observation in this study is what the authors called fixation displacement. Participants who began with prompts unrelated to the example, such as “goddess,” went on to draw sketches that did not resemble the example but closely resembled the Midjourney images. These sketches scored low on fixation, which was measured as overlap with the example, yet the object of fixation had only moved from the example to the AI images.
Designers have also raised the appearance of generated output as a problem. From an exchange session with a product design team, Lin et al. (2025) derived the design challenge that generated images appear “too complete” to build on and can lead to fixation. The authors responded with Inkspire, a tool that turns generated images back into low-fidelity sketch scaffolding for the designer to draw over, and reported in a within-subjects comparison with ControlNet that it supported more inspiration and exploration. The abstract of Ranscombe et al. (2026) analyzes 1,292 prompts from 13 student designers, discusses the tension between instantaneous high-fidelity output and ambiguous, iterative sketching, and identifies fixation alongside broad exploration and reinterpretation in the students’ use of generative AI. The abstract of Song et al. (2025) states that, in an experiment with ten designers, fixation appeared at the level of design features in the output of generative AI itself.
Similar phenomena have been reported when the prototype is text or functionality. Subramonyam et al. (2025) studied prompt-based prototyping with 39 industry professionals and identified overfitting of the output to the examples in a prompt as a challenge. One participant was surprised at how much the model could produce from a single example but pointed out that the generated examples all resembled the first one.
Generative AI does not always make ideas poorer. Ge and Hou (2025) compared text-to-image and image-to-image tools with 84 environmental design students. The text-to-image group scored higher on originality and other measures, and also higher on repetition of features from the AI images (3.84 vs. 3.00, p < 0.001). This comparison is between types of AI tools, and the study did not include a condition without AI. Wadinambiarachchi, Waycott, and Wadley (2026) asked students to use a tool that accepts sketches instead of text as prompts and reported that sketching tended to increase fluency, that differences in variety, originality, and quality were inconclusive, and that students strongly preferred text prompts.
Wider for each person, narrower for the group
The effect of generative AI changes direction depending on whether one looks at an individual creator or at a group of creators. Doshi and Hauser (2024) asked 293 writers to write short stories with no access to LLM ideas, access to one idea, or access to up to five ideas, and had 600 evaluators read the stories. Stories by writers with access to up to five ideas were rated 8.1% higher in novelty and 9.0% higher in usefulness than stories by writers without access. At the same time, stories written with AI ideas were more similar to each other, and the increase in similarity corresponded to 10.7% (one idea) and 8.9% (five ideas) of the range of similarity scores in the human-only condition. The authors compared this pattern to a social dilemma in which individuals are better off while the group produces a narrower range of novel content.
Anderson, Shah, and Kreminski (2024) gave 36 participants divergent ideation tasks and compared ChatGPT with Oblique Strategies, a deck of cards that prompts ideas. ChatGPT users produced more ideas and more detailed ideas, but ideas from different users were semantically closer, and users felt less responsible for their ideas. The authors interpreted this homogenization as arising from the LLM providing different users with similar ideas rather than from increased individual fixation, and suggested that output that looks close to a finished product may contribute to it. Meincke, Nave, and Terwiesch (2025) reanalyzed data from experiments on brainstorming with ChatGPT and state that ChatGPT enhances the creativity of individual ideas while reducing the diversity of ideas in a pool. Padmakumar and He (2024) showed that co-writing argumentative essays with an LLM made the writing of different authors more similar, and attributed the effect to the text contributed by the model, while the text written by users remained unaffected (the effect occurred with InstructGPT but not with GPT-3).
Studies of actual creative output show the same pattern. Zhou and Lee (2024) analyzed more than 4 million artworks by more than 50,000 artists and estimated that adopting text-to-image AI increased productivity by 25% and the likelihood of receiving a favorite per view by 50%. Peak content novelty increased while average content novelty declined, and both peak and average visual novelty declined. The authors interpreted the content results as an expanding but inefficient idea space.
For user interface prototypes, one study reports that generated designs lean toward convention. Romero et al. (2026) asked 92 participants to evaluate, without knowing the authorship, prototypes generated with five generative AI tools and prototypes created by people. The AI-generated prototypes were rated positively on pragmatic qualities such as usability and efficiency and neutral to negative on hedonic qualities such as originality, and the authors interpreted this as AI prototypes reinforcing conventional interface patterns. In interviews with 22 members of product teams by Li et al. (2026), participants described vibe coding (asking AI in natural language to produce working prototypes and code) as moving them toward working prototypes before polished mockups. The authors argue that there is a tension between efficiency-driven prototyping and reflection on values and alternatives, and that without such reflection, design cycles risk premature convergence and homogenization.
What supports the operations that move a design space?
These results indicate that increasing the number of ideas and widening exploration are different things. This raises the question of what widening exploration consists of.
Design research has described exploration as movement in two spaces. Based on protocol analysis, Dorst and Cross (2001) stated that the problem space and the solution space both remain unstable and evolve, and that a creative event occurs when a “bridge” emerges that temporarily fixes a problem-solution pair. The abstract of Woodbury and Burrow (2006) describes a design space as the network structure of related designs visited in an exploration process. In this view, exploration operations include not only adding ideas but also comparing them, connecting them, and reframing the problem. Lim et al. (2008) called prototypes filters that traverse a design space in this sense of operational tools.
Tool research has tried to support the operation of placing several designs side by side and comparing them. Hartmann et al. (2008) built Juxtapose, an environment for editing and running alternatives of user interfaces in parallel, and showed in a study with 18 participants that designers could survey more options faster. For generative AI tools, the starting problem is that the default interaction returns a single answer. Suh et al. (2024) argued that current LLM interfaces guide users toward rapid convergence on a limited set of ideas, and built Luminate, which first generates dimensions relevant to the task and then arranges outputs along those dimensions. In a user study with 14 professional writers, support for exploration received the highest rating, although the study had no comparison condition. Zamfirescu-Pereira et al. (2025) took as their problem that code-generating LLMs by default deliver code representing one particular point solution, and built an IDE that surfaces alternatives, requirements, and implicit decisions. The 11 participants combined and parallelized design phases to explore a broader design space, but struggled to keep up with LLM-originated changes and with information overload. DynEx by Ma et al. (2025) explores design options in a matrix before generating code step by step, and the 10 participants rated it higher than a Claude Artifact baseline in supporting design exploration.
Increasing diversity mechanically does not necessarily improve exploration. According to the abstract of Cai et al. (2023), DesignAID, which generates images from verbal ideas, was rated by 87 designers as more inspirational than image search. However, generating highly diverse ideas was only somewhat better in inspiration and no better in other dimensions for image generation, and it was worse in all dimensions for image search.
Most evaluations of these tools are short user studies with about 10 to 20 participants, and the comparison conditions differ from tool to tool. Even so, several studies take the same design decision. Instead of returning generated output as a single answer, they turn it back into material for exploration operations in the form of dimensions, alternatives, or low-fidelity scaffolding.
What does externalizing the process make into research?
RtD treats making as a method of research. Frayling (1993) distinguished research into art and design, research through art and design, and research for art and design. As an example of research through art and design, he described a research diary that reports a practical studio experiment step by step and a report that places it in context, and wrote that the diary and the report communicate the results, which is what separates research from practice. From the beginning, RtD placed the externalization of the process as a condition for counting as research.
Zimmerman, Forlizzi, and Evenson (2007), who brought RtD into HCI, proposed four lenses for evaluating a contribution: process, invention, relevance, and extensibility. On process, they did not expect that reproducing the process would produce the same results, but required researchers to document enough detail that the process could be reproduced and to give a rationale for the methods selected, so that the rigor of the work could be judged. The abstract of Zimmerman, Stolterman, and Forlizzi (2010) critiques RtD practice based on interviews with 12 HCI design researchers and three historical RtD projects that the interviewees repeatedly mentioned.
At the same time, discomfort with documentation modeled on scientific standards has also been expressed. Gaver (2012) argued that the theories RtD produces are provisional, contingent, and aspirational, and proposed treating theory as annotation of realized design examples, particularly of portfolios of related work. The abstract of Bowers (2012) proposes annotated portfolios as a means of capturing family resemblances in a collection of artifacts while respecting the particularity of each design. Löwgren (2013) summarized annotated portfolios as selecting a collection of designs, re-presenting them in an appropriate medium, and adding brief textual annotations, and placed them as knowledge between individual cases and general theory. Gaver (2011) described design workbooks, collections of design proposals, as creating a design space rather than only describing it, and argued that keeping early proposals provisional and vague invites viewers to interpret and respond to them.
Practical difficulties of documentation have also been reported. According to the abstract of Dalsgaard and Halskov (2012), using a documentation tool in cases lasting nine to thirteen months raised challenges concerning roles and responsibilities, a lack of routines, what to document, and the right level of detail. Some benefits of documentation did not appear until the research was written up. The abstract of Bardzell et al. (2016) states that RtD documentation is essential for translating design knowledge into academic knowledge but has received little sustained attention. The authors present a framework for planning and evaluating documentation that addresses three concerns: the medium of documentation, the performativity of documentation, and equal support for research and design. Odom et al. (2016) distinguished artifacts made to investigate questions through long-term use from prototypes, called them research products, and defined them by four qualities: inquiry-driven, finish, fit, and independent. Within RtD, documentation that keeps provisional proposals and the placement of finished artifacts are distinguished as serving different roles.
Does prototyping with generative AI leave a process to document?
Prototyping with generative AI may reduce the intermediate representations that documentation records. In the interviews by Li et al. (2026), participants described a shift from workflows that began with polished mockups and role-specific handoffs to workflows that move directly to working prototypes. In the same study, one participant said that AI should be treated like a co-author, with prompts recorded and the places where AI intervened and the edit history documented. Wadinambiarachchi, Waycott, and Wadley (2026) wrote that many students no longer wrestle with visualizing their ideas and instead race to construct prompts that fill their screens with highly refined images, and argued for bringing reflection through sketching back into designer-AI interaction.
Some studies use RtD to engage with generative AI. Benjamin et al. (2024) conducted RtD with a primary school to develop a constructionist curriculum around generative AI and argued that RtD can serve as a rapid response methodology for fast-changing technologies such as generative AI. However, generative AI is the subject of inquiry in that study, not the means of making prototypes. In this collection, no peer-reviewed study was found that directly examines what remains in RtD documentation (diaries, workbooks, annotated portfolios) and what disappears when prototypes are made with generative AI.
How should the benefits and harms be read?
Within the range supported by the literature, three points can be made.
First, the benefits and harms of prototyping depend on the conditions under which prototypes are made. Creating several designs before comparing them and delaying the decision to narrow down improved results and the diversity of designs (Dow et al. 2010, 2011; Neeley et al. 2013). When only one example was shown, designers repeated its features (Jansson and Smith 1991), and evaluators rated a single design higher and criticized it less (Tohidi et al. 2006). The effect of physical prototyping differed across studies, and the effort of building (Viswanathan and Linsey 2013) and the timing of building (Jang and Schunn 2012) have been proposed as candidates to explain the difference.
Second, for prototyping with generative AI, there is evidence both that it strengthens individual fixation on examples and that it improves individual output. Designers who saw an example and AI images repeated features of both (Wadinambiarachchi et al. 2024), and stories written with AI ideas were rated more highly (Doshi and Hauser 2024). These findings do not contradict each other. The evaluations measured the quality of individual works, whereas the fixation study measured how far ideas departed from an example. At the collective level, every study that measured it reported results in the direction of works becoming more similar or average novelty declining (Doshi and Hauser 2024; Anderson et al. 2024; Meincke et al. 2025; Padmakumar and He 2024; average novelty in Zhou and Lee 2024).
Third, research on tools that support exploration distinguishes increasing output from widening exploration. Generating diverse ideas automatically did not raise their value (Cai et al. 2023), and the tools for which positive effects were reported gave generated output scaffolding for comparison and reframing, such as dimensions, alternatives, or low-fidelity underlays (Suh et al. 2024; Zamfirescu-Pereira et al. 2025; Lin et al. 2025; Ma et al. 2025).
The following is this note’s interpretation, which combines the literature; no study was found that tests it directly. If fixation is divided into a route through examples that designers see and a route through ideas in which they have invested effort, generative AI may act on the two routes in opposite directions. Reducing the effort of making to seconds creates conditions that weaken the sunk-cost route. At the same time, presenting many finished-looking outputs in a short time creates conditions that strengthen the example route. The fixation displacement observed by Wadinambiarachchi et al. (2024) and the “too complete” problem identified by Lin et al. (2025) are consistent with the second route being at work. The finding summarized by Vasconcelos and Crilly (2016), that higher-fidelity examples lead to less novel ideas, points in the same direction. If this reading is correct, what matters when prototyping with generative AI is not a further reduction in the effort of making but ways of keeping generated output from looking finished and procedures for laying out and comparing several designs. Testing this would require experiments that manipulate the effort of making and the fidelity of generated output separately and measure both fixation and collective diversity.
Gaps
- Separating sunk cost from example exposure: No experiment was found that separately manipulates the effort of making and the appearance (fidelity) of generated output in prototyping with generative AI. The interpretation above has not been tested.
- Longer periods and real design settings: Experiments on generative AI and fixation mainly use a single task of about 20 minutes (Wadinambiarachchi et al. 2024). No study was found that follows fixation or homogenization over the course of a project.
- Homogenization in teams and organizations: Reduced collective diversity has been measured among individuals working independently. Apart from the artwork data of Zhou and Lee (2024), no study has examined whether the range of designs narrows in design teams using the same tool or in industries where the same tool has spread. This work has not yet been connected to the collective fixation that Crilly and Cardoso (2017) proposed as a research question.
- Comparing exploration tools: Most generative AI tools that expose design spaces were evaluated in user studies with about 10 to 20 participants, with comparison conditions that differ from tool to tool. No study compared such tools on the same task.
- RtD documentation and generative AI: No peer-reviewed study was found on how the process of making with generative AI can be kept as RtD documentation (diaries, workbooks, annotated portfolios) or on what is no longer kept. Calls to record prompts and edit histories appear so far in practitioner interviews (Li et al. 2026).
- Related notes: AI Slop: Reading It as Outsourced Verification, Not Low Quality reads Wadinambiarachchi et al. (2024) within a discussion of homogenization in generative AI and creativity, Vibe Coding and UI Agents: Natural-Language Production and Its Boundaries covers the industry side of vibe coding, and Asking an LLM for Design: Briefs, Skills, and a Process That Keep UIs from Converging on the Obvious covers the convergence of LLM-generated user interfaces on generic forms.
Unverified items
- No major claim remains unverified.
- The full texts of Jansson and Smith (1991) and Youmans (2011) could not be obtained, and these studies are reported within the summaries in the full texts of Crilly and Cardoso (2017), Vasconcelos and Crilly (2016), and Wadinambiarachchi et al. (2024).
- Twelve sources are reported within the scope of their abstracts because their full texts could not be obtained (Virzi et al. 1996; Viswanathan and Linsey 2012, 2013; Jang and Schunn 2012; Song et al. 2025; Ranscombe et al. 2026; Woodbury and Burrow 2006; Cai et al. 2023; Zimmerman et al. 2010; Bowers 2012; Dalsgaard and Halskov 2012; Bardzell et al. 2016). In the text, these are attributed to their authors with “according to the abstract” or “the abstract states.” For Meincke et al. (2025), only the opening paragraph of the subscription page was checked.
- The full text of Dow et al. (2010) could be read only in the submitted version, so the note reports no numbers from it and stays within the abstract of the published version.
References
All URLs were accessed on 2026-10-09.
Benefits and fidelity of prototypes
- Lim, Y.-K., Stolterman, E., & Tenenberg, J. (2008). The anatomy of prototypes: Prototypes as filters, prototypes as manifestations of design ideas. ACM Transactions on Computer-Human Interaction, 15(2), 1–27. https://doi.org/10.1145/1375761.1375762
- Buchenau, M., & Fulton Suri, J. (2000). Experience prototyping. In Proceedings of the 3rd Conference on Designing Interactive Systems (DIS ‘00) (pp. 424–433). ACM. https://doi.org/10.1145/347642.347802
- Camburn, B., Viswanathan, V., Linsey, J., Anderson, D., Jensen, D., Crawford, R., Otto, K., & Wood, K. (2017). Design prototyping methods: State of the art in strategies, techniques, and guidelines. Design Science, 3, e13. https://doi.org/10.1017/dsj.2017.10
- Virzi, R. A., Sokolov, J. L., & Karis, D. (1996). Usability problem identification using both low- and high-fidelity prototypes. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems (CHI ‘96) (pp. 236–243). ACM. https://doi.org/10.1145/238386.238516
- Tohidi, M., Buxton, W., Baecker, R., & Sellen, A. (2006). Getting the right design and the design right. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems (CHI ‘06) (pp. 1243–1252). ACM. https://doi.org/10.1145/1124772.1124960
- Dow, S. P., Heddleston, K., & Klemmer, S. R. (2009). The efficacy of prototyping under time constraints. In Proceedings of the Seventh ACM Conference on Creativity and Cognition (C&C ‘09) (pp. 165–174). ACM. https://doi.org/10.1145/1640233.1640260
Parallel prototyping and alternatives
- Dow, S. P., Glassco, A., Kass, J., Schwarz, M., Schwartz, D. L., & Klemmer, S. R. (2010). Parallel prototyping leads to better design results, more divergence, and increased self-efficacy. ACM Transactions on Computer-Human Interaction, 17(4), 1–24. https://doi.org/10.1145/1879831.1879836
- Dow, S. P., Fortuna, J., Schwartz, D., Altringer, B., Schwartz, D. L., & Klemmer, S. R. (2011). Prototyping dynamics: Sharing multiple designs improves exploration, group rapport, and results. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems (CHI ‘11) (pp. 2807–2816). ACM. https://doi.org/10.1145/1978942.1979359
- Neeley, W. L., Jr., Lim, K., Zhu, A., & Yang, M. C. (2013). Building fast to think faster: Exploiting rapid prototyping to accelerate ideation during early stage design. In Proceedings of the ASME 2013 International Design Engineering Technical Conferences and Computers and Information in Engineering Conference, Vol. 5. ASME. https://doi.org/10.1115/detc2013-12635
- Hartmann, B., Yu, L., Allison, A., Yang, Y., & Klemmer, S. R. (2008). Design as exploration: Creating interface alternatives through parallel authoring and runtime tuning. In Proceedings of the 21st Annual ACM Symposium on User Interface Software and Technology (UIST ‘08) (pp. 91–100). ACM. https://doi.org/10.1145/1449715.1449732
Design fixation
- Jansson, D. G., & Smith, S. M. (1991). Design fixation. Design Studies, 12(1), 3–11. https://doi.org/10.1016/0142-694X%2891%2990003-F
- Linsey, J. S., Tseng, I., Fu, K., Cagan, J., Wood, K. L., & Schunn, C. (2010). A study of design fixation, its mitigation and perception in engineering design faculty. Journal of Mechanical Design, 132(4), 041003. https://doi.org/10.1115/1.4001110
- Vasconcelos, L. A., & Crilly, N. (2016). Inspiration and fixation: Questions, methods, findings, and challenges. Design Studies, 42, 1–32. https://doi.org/10.1016/j.destud.2015.11.001
- Crilly, N., & Cardoso, C. (2017). Where next for research on fixation, inspiration and creativity in design? Design Studies, 50, 1–38. https://doi.org/10.1016/j.destud.2017.02.001
- Youmans, R. J. (2011). The effects of physical prototyping and group work on the reduction of design fixation. Design Studies, 32(2), 115–138. https://doi.org/10.1016/j.destud.2010.08.001
- Viswanathan, V. K., & Linsey, J. S. (2012). Physical models and design thinking: A study of functionality, novelty and variety of ideas. Journal of Mechanical Design, 134(9), 091004. https://doi.org/10.1115/1.4007148
- Viswanathan, V. K., & Linsey, J. S. (2013). Role of sunk cost in engineering idea generation: An experimental investigation. Journal of Mechanical Design, 135(12), 121002. https://doi.org/10.1115/1.4025290
- Jang, J., & Schunn, C. D. (2012). Physical design tools support and hinder innovative engineering design. Journal of Mechanical Design, 134(4), 041001. https://doi.org/10.1115/1.4005651
Prototyping with generative AI and fixation
- Wadinambiarachchi, S., Kelly, R. M., Pareek, S., Zhou, Q., & Velloso, E. (2024). The effects of generative AI on design fixation and divergent thinking. In Proceedings of the CHI Conference on Human Factors in Computing Systems (CHI ‘24) (pp. 1–18). ACM. https://doi.org/10.1145/3613904.3642919
- Song, Y., Zheng, C., Jing, Q., Hansen, P., Sun, L., & Chen, L. (2025). Understanding design fixation in generative artificial intelligence. In Proceedings of the ASME 2025 International Design Engineering Technical Conferences and Computers and Information in Engineering Conference, Vol. 4. ASME. https://doi.org/10.1115/detc2025-168630
- Ranscombe, C., Tan, L., Zhang, W., & Afif, N. (2026). Friend and foe: Characterising (counter)productive prompting of generative AI to supplement exploratory sketching. In DRS2026: Edinburgh. Design Research Society. https://doi.org/10.21606/drs.2026.1595
- Ge, W., & Hou, G. (2025). The effects of generative AI model type and visual stimuli type on design creativity. Artificial Intelligence for Engineering Design, Analysis and Manufacturing, 39, e17. https://doi.org/10.1017/S0890060425100061
- Wadinambiarachchi, S., Waycott, J., & Wadley, G. (2026). Reviving reflection-in-action: Instilling designerly thinking in AI-supported ideation through multimodal prompting. In Proceedings of the 2026 Conference on Creativity and Cognition (C&C ‘26) (pp. 359–376). ACM. https://doi.org/10.1145/3803784.3807524
- Subramonyam, H., Thakkar, D., Ku, A., Dieber, J., & Sinha, A. K. (2025). Prototyping with prompts: Emerging approaches and challenges in generative AI design for collaborative software teams. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems (CHI ‘25) (pp. 1–22). ACM. https://doi.org/10.1145/3706598.3713166
Individual output and collective diversity
- Doshi, A. R., & Hauser, O. P. (2024). Generative AI enhances individual creativity but reduces the collective diversity of novel content. Science Advances, 10(28), eadn5290. https://doi.org/10.1126/sciadv.adn5290
- Anderson, B. R., Shah, J. H., & Kreminski, M. (2024). Homogenization effects of large language models on human creative ideation. In Proceedings of the 16th Conference on Creativity & Cognition (C&C ‘24) (pp. 413–425). ACM. https://doi.org/10.1145/3635636.3656204
- Meincke, L., Nave, G., & Terwiesch, C. (2025). ChatGPT decreases idea diversity in brainstorming. Nature Human Behaviour, 9(6), 1107–1109. https://doi.org/10.1038/s41562-025-02173-x
- Padmakumar, V., & He, H. (2024). Does writing with language models reduce content diversity? In The Twelfth International Conference on Learning Representations (ICLR 2024). https://arxiv.org/abs/2309.05196
- Zhou, E., & Lee, D. (2024). Generative artificial intelligence, human creativity, and art. PNAS Nexus, 3(3), pgae052. https://doi.org/10.1093/pnasnexus/pgae052
- Romero, J., Wiese, I., Balancieiri, R., Leal, G., & Guerino, G. (2026). Usable but conventional: An empirical study on the UX of AI-generated interface prototypes. In Anais do LIII Seminário Integrado de Software e Hardware (SEMISH 2026) (pp. 830–841). SBC. https://doi.org/10.5753/semish.2026.21904
- Li, J., Hou, Y., Lin, L., Zhu, R., Cao, H., & El Ali, A. (2026). Vibe coding in product teams: Reconfiguring AI-assisted workflows, prototyping, and collaboration. In Proceedings of the 5th Annual Symposium on Human-Computer Interaction for Work (CHIWORK ‘26) (pp. 1–16). ACM. https://doi.org/10.1145/3808045.3808062
Design spaces and exploration tools
- Dorst, K., & Cross, N. (2001). Creativity in the design process: Co-evolution of problem–solution. Design Studies, 22(5), 425–437. https://doi.org/10.1016/S0142-694X%2801%2900009-6
- Woodbury, R. F., & Burrow, A. L. (2006). Whither design space? Artificial Intelligence for Engineering Design, Analysis and Manufacturing, 20(2), 63–82. https://doi.org/10.1017/S0890060406060057
- Suh, S., Chen, M., Min, B., Li, T. J.-J., & Xia, H. (2024). Luminate: Structured generation and exploration of design space with large language models for human-AI co-creation. In Proceedings of the CHI Conference on Human Factors in Computing Systems (CHI ‘24) (pp. 1–26). ACM. https://doi.org/10.1145/3613904.3642400
- Zamfirescu-Pereira, J. D., Jun, E., Terry, M., Yang, Q., & Hartmann, B. (2025). Beyond code generation: LLM-supported exploration of the program design space. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems (CHI ‘25) (pp. 1–17). ACM. https://doi.org/10.1145/3706598.3714154
- Ma, J. G., Sreedhar, K., Liu, V., Perez, P. A., Wang, S., Sahni, R., & Chilton, L. B. (2025). DynEx: Dynamic code synthesis with structured design exploration for accelerated exploratory programming. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems (CHI ‘25) (pp. 1–27). ACM. https://doi.org/10.1145/3706598.3714115
- Lin, D. C.-E., Kang, H. B., Martelaro, N., Kittur, A., Chen, Y.-Y., & Hong, M. K. (2025). Inkspire: Supporting design exploration with generative AI through analogical sketching. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems (CHI ‘25) (pp. 1–18). ACM. https://doi.org/10.1145/3706598.3713397
- Cai, A., Rick, S. R., Heyman, J. L., Zhang, Y., Filipowicz, A., Hong, M., Klenk, M., & Malone, T. (2023). DesignAID: Using generative AI and semantic diversity for design inspiration. In Proceedings of the ACM Collective Intelligence Conference (CI ‘23) (pp. 1–11). ACM. https://doi.org/10.1145/3582269.3615596
Research through Design and process documentation
- Frayling, C. (1993). Research in art and design. Royal College of Art Research Papers, 1(1). https://researchonline.rca.ac.uk/384/
- Zimmerman, J., Forlizzi, J., & Evenson, S. (2007). Research through design as a method for interaction design research in HCI. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems (CHI ‘07) (pp. 493–502). ACM. https://doi.org/10.1145/1240624.1240704
- Zimmerman, J., Stolterman, E., & Forlizzi, J. (2010). An analysis and critique of Research through Design: Towards a formalization of a research approach. In Proceedings of the 8th ACM Conference on Designing Interactive Systems (DIS ‘10) (pp. 310–319). ACM. https://doi.org/10.1145/1858171.1858228
- Gaver, W. (2011). Making spaces: How design workbooks work. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems (CHI ‘11) (pp. 1551–1560). ACM. https://doi.org/10.1145/1978942.1979169
- Gaver, W. (2012). What should we expect from research through design? In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems (CHI ‘12) (pp. 937–946). ACM. https://doi.org/10.1145/2207676.2208538
- Bowers, J. (2012). The logic of annotated portfolios: Communicating the value of ‘research through design’. In Proceedings of the Designing Interactive Systems Conference (DIS ‘12) (pp. 68–77). ACM. https://doi.org/10.1145/2317956.2317968
- Löwgren, J. (2013). Annotated portfolios and other forms of intermediate-level knowledge. Interactions, 20(1), 30–34. https://doi.org/10.1145/2405716.2405725
- Dalsgaard, P., & Halskov, K. (2012). Reflective design documentation. In Proceedings of the Designing Interactive Systems Conference (DIS ‘12) (pp. 428–437). ACM. https://doi.org/10.1145/2317956.2318020
- Bardzell, J., Bardzell, S., Dalsgaard, P., Gross, S., & Halskov, K. (2016). Documenting the research through design process. In Proceedings of the 2016 ACM Conference on Designing Interactive Systems (DIS ‘16) (pp. 96–107). ACM. https://doi.org/10.1145/2901790.2901859
- Odom, W., Wakkary, R., Lim, Y.-K., Desjardins, A., Hengeveld, B., & Banks, R. (2016). From research prototype to research product. In Proceedings of the 2016 CHI Conference on Human Factors in Computing Systems (CHI ‘16) (pp. 2549–2561). ACM. https://doi.org/10.1145/2858036.2858447
- Benjamin, J. J., Lindley, J., Edwards, E., Rubegni, E., Korjakow, T., Grist, D., & Sharkey, R. (2024). Responding to generative AI technologies with research-through-design: The Ryelands AI Lab as an exploratory study. In Proceedings of the 2024 ACM Designing Interactive Systems Conference (DIS ‘24) (pp. 1823–1841). ACM. https://doi.org/10.1145/3643834.3660677
Related notes
- AI Slop: Reading It as Outsourced Verification, Not Low Quality: Reads Wadinambiarachchi et al. (2024) and Jansson and Smith (1991) within a discussion of homogenization and judgment in generative AI. This note rereads the same experiment in terms of the conditions under which prototypes are made.
- Vibe Coding and UI Agents: Natural-Language Production and Its Boundaries: Covers the industry side of vibe coding and the practitioner view that the bottleneck moves from making to judging. It corresponds to the interviews by Li et al. (2026) in this note.
- Asking an LLM for Design: Briefs, Skills, and a Process That Keep UIs from Converging on the Obvious: Covers the convergence of LLM-generated user interfaces on generic forms and the briefs and procedures for avoiding it. It is the practitioner counterpart of the section “Wider for each person, narrower for the group.”
Author: Shuichiro Ogawa (Design Researcher / Consultant) About me →