Shuichiro Ogawa
日本語

Notes · updated 2026-10-05

AI and Design Monthly Scholarly Watch (September–October 2026)

The fourth installment of the monthly scholarly watch, collecting 12 peer-reviewed papers and preprints at the intersection of AI and design for September 8 through October 5, 2026.

Contents (6)
  1. In practice, AI use falls off in later phases, and time savings split by profession
  2. Tool research decomposes design knowledge, pairs it with speech, and shares it with clients
  3. In education, positive effects and negative signs measure different things
  4. A methodological proposal: making HCI design knowledge machine-readable
  5. Gaps
  6. Notes on reading

Related: AI and Design Monthly Scholarly Watch (August–September 2026) (previous monthly watch) / AI and Design Monthly Scholarly Watch (July–August 2026) (the one before) / AI and Design Monthly Scholarly Watch (June–July 2026) (first watch) / AI in the Design Industry: An Academic Review (2026) (literature review of AI use across the design field) / AI Adaptation in Design Education — The Current State and Structural Challenges of Curriculum Reform (design education curriculum reform) / Designer Careers That Will Generate Value over the Next Five Years — Academic Review (2026) (designers’ professional value).

The previous window ended with 7 papers and two threads running in parallel: tool research on controllability and agency, and classroom reports of negative effects. This window (September 8 to October 5, 2026) covers 12 papers, and the question is how those threads move once practice data and a meta-analysis enter. The items divide into three on practice, three on tools, five on education, and one methodological proposal.

In practice, AI use falls off in later phases, and time savings split by profession

Tajarmakan and Hojjati Emami surveyed 443 designers in 43 countries on AI use across the four Double Diamond phases (Discover, Define, Develop, Deliver) (2026). A total of 353 respondents (79.7%) reported confirmed use in at least one phase. Adopters used AI in a mean of 2.89 of the four phases, and tended to keep the earliest phases while dropping the latest. In the barrier table, only low perceived usefulness differed between groups after Holm correction (16.4% of adopters versus 40.5% of non-adopters, p < .001). Cost and job-security concerns were nominally significant before correction, did not survive it, and were endorsed more often by adopters. The authors report that lower use intensity reflects fewer designers engaging, not less use among those who engage.

Stewart and colleagues asked the same question about time with a randomized trial (2026). Fifty product designers and 50 product managers (PMs) were assigned access or no access to Figma Make for three standardized tasks. Among participants who completed the tasks, completion time was about 20% shorter when averaged across tasks (average marginal effect 10:53, p = 0.0014). The gains were larger for PMs, whose total time fell by 35%. For product designers, the reduction was significant only on Task 3 (26%), and the overall 17% reduction was marginally significant. The authors suggest that prompt-to-design tools may widen PMs’ participation in design work, while benefits for designers may depend on the task. On this single metric, the gains followed professional lines.

In architecture, Yiannoudes scored 67 cases from 65 peer-reviewed papers on a six-indicator Practice Engagement Index (International Journal of Architectural Computing, 2026). According to the abstract, many cases meet realistic constraints (62.6%) and common building types (58.2%), while 17.9% score high on workflow accessibility and 50.7% score low. The full text was not accessible, so these figures rest on the abstract. The abstract’s diagnosis, that technical novelty does not reach routine tools, points the same way as the late-phase drop-off in the designer survey.

Tool research decomposes design knowledge, pairs it with speech, and shares it with clients

Yang and colleagues’ Scanvas (2026) targets the tendency of LLM ideation tools toward feature blending. It decomposes seed ideas into components, behaviors, surpluses, and issues, then searches with three operators: strengthening goals, turning weaknesses into resources, and sharing components across functions. In a within-subject study with 12 professionals, expert ratings were higher than the baseline on all four rating dimensions (p ≤ .001). Where LegoUI staged the generation process so that people could trace it, Scanvas builds a design theory (synergy) into the search operators. The object of staging has moved from the provenance of outputs to the structure of ideation.

Shi and colleagues’ CommSketch (2026) compared sketching alone with sketching plus concurrent speech in a between-subjects study (N = 24). Perceived alignment with the AI was significantly higher with speech (p < .05), which the authors read as better shared understanding of design intent. This is a work-in-progress paper with a small sample.

Zhang and Wang studied 11 architect-client pairs who used generative image AI over video conference to produce early renderings of the client’s dream house (2026). Images served as concrete material for shared understanding, and clients took a more active part in setting direction. Two problems appeared: bias toward particular styles, and unpredictable shifts in direction caused by variation across generations. Earlier tool studies addressed designers’ controllability; this one addresses agreement between designers and non-designers. Read with Stewart and colleagues’ PM result, both studies concern settings where AI enters a pairing of designers and non-designers.

In education, positive effects and negative signs measure different things

Chiang’s meta-analysis (International Journal of Technology and Design Education, 2026) pooled 27 controlled studies (N = 2,847) and reports an overall effect of d = 0.60 (95% CI [0.55, 0.65]). According to the abstract, instructor-student collaboration gave the largest usage-mode estimate (d = 0.89) and creative ability the largest outcome estimate (d = 0.71), with no detected differences by tool type, instructional phase, or design discipline. The author suggests that pedagogical configuration may matter more than tool type, while cautioning about small subgroups and short interventions. The full text was not accessible, so this rests on the abstract.

Xiao and colleagues ran an undergraduate architectural design studio of 27 students in 7 teams, in which 4 teams used GenARch (generative AI with extended reality) and 3 did not (2026). Self-efficacy confidence fell more for the GenARch condition (difference-in-differences −1.675, p = 0.021). Panel-rated presentation outcomes did not differ significantly between conditions. The teamwork subscale for managing conflict had a significant positive difference-in-differences estimate (β = 0.479, p = 0.014), and no other subscale differed significantly. Interviews reported that generative AI supported externalizing ideas and XR supported spatial evaluation. The authors ask that the decline be read together with implementation difficulties, namely students learning unfamiliar interaction methods while judging generated outputs. With 4 teams against 3, the comparison cannot separate the effect of AI itself from those difficulties.

The meta-analytic gain in creative ability and the studio’s drop in confidence are not contradictory. The first is an average of learning outcomes across controlled studies, and the second is a self-assessment in one studio. Neither measures impacts on decision-making, dependence, or self-perception with a common instrument.

Two studies address those dimensions. An and colleagues’ mixed-methods study (International Journal of Technology and Design Education, 2026) surveyed 317 design students and reports, per the abstract, that more frequent use was associated with stronger dependency on all three dimensions of reading, writing, and creativity, with the creative dimension having the greatest explanatory power. Huh and colleagues surveyed 112 design students (Education Sciences, 2026) and report, per the abstract, higher use for concept development, research, and written articulation than for visual refinement, technical documentation, and later stages, with use growing more selective as projects progressed. Instructor expectations, peer influence, and concern about negative judgment were associated with students’ decisions, and the authors interpret this selectivity as boundary-setting around how AI participates in design. The decline of AI use in later phases appears both among professionals in the Tajarmakan survey and among students here, but the populations and phase divisions differ, so the two cannot be treated as one mechanism.

Ahmed’s mixed-methods study (Frontiers in Education, 2026) combined a survey of 270 participants (265 valid for exploratory factor analysis) with interviews of 12 design academics and practitioners (650 meaning units) in graphic design education, and proposes the Bot-Haus Curriculum Architecture. The author states that the quantitative orientations are not validated latent constructs, so the framework is best read as an empirically informed design proposal.

A methodological proposal: making HCI design knowledge machine-readable

Conklin and colleagues (2026) start from the fact that LLMs can now generate working interfaces from natural-language task descriptions. They propose encoding classical design knowledge, such as Nielsen’s heuristics, Norman’s affordances, and WCAG success criteria, as skill.md files that the generating agent loads at runtime. This is a framework proposal, not an empirical study. It answers the question this watch has followed, where to place controllability, by making design knowledge a property of the generative process instead of the finished artifact.

Gaps

First, the positive and negative education findings are still not compared on one scale. The full text of Chiang’s meta-analysis was not accessible, so the outcome measures of its 27 studies could not be checked. A study that measures self-efficacy, dependence, and decision quality in one design, split by usage mode, is needed.

Second, both practice figures are limited: one is a self-reported survey and the other an experiment on a single product, Figma Make. The survey identifies only low perceived usefulness as a differing barrier, so the mechanism behind the late-phase drop-off remains open.

Third, DRS 2026 papers and the ACM Digital Library C&C ‘26 listing, carried over from the last watch, were not searched this time.

Notes on reading

No major conference session falls inside this window, so items were collected from arXiv cs.HC and from peer-reviewed articles with a Crossref publication date from 2026-09-08 to 2026-10-05. Seven of the 12 are arXiv preprints (Tajarmakan, Stewart, Yang, Shi, Zhang, Xiao, Conklin), all unreviewed. The journal articles by Chiang, An, Yiannoudes, and Huh could not be retrieved in full text, so they are described within their abstracts. Ahmed’s full text was retrieved, but the results are described only to the extent of the abstract and figure titles. Details of source tracking are recorded in the Provenance and verification-evidence sections of the corpus.

参照文献

すべて2026年10月5日にアクセス。

  • Tajarmakan, S., Hojjati Emami, K. “AI Tools Adoption across the Double Diamond Workflow: Phase, Mode, and Barriers in Designer Practice.” arXiv(未査読). https://arxiv.org/abs/2609.34655
  • Stewart, R., Anise, O., Hogan, A., Griffin, A. “Does AI Save Time on Product Design? A Randomized Controlled Experiment of AI Prompt-to-Design Workflows.” arXiv(未査読). https://arxiv.org/abs/2609.26725
  • Yang, Y., Kittur, A., Wang, H. H., Martelaro, N., Klenk, M., Chen, Y.-Y., Hong, M. K. “Scanvas: Discovering and Developing Synergistic Opportunities in Generative Design Spaces.” arXiv(未査読). https://arxiv.org/abs/2609.34062
  • Shi, W., Lim, D., Quek, G., Choo, K. T. W. “CommSketch: How Speaking while Sketching Steers Human-AI Design Ideation.” arXiv(未査読). https://arxiv.org/abs/2609.37813
  • Zhang, C., Wang, W. “Exploring the Affordances of Generative Image AI for Supporting Early-stage Architect-client Communication.” arXiv(未査読). https://arxiv.org/abs/2609.34118
  • Xiao, Y., Chen, M., Li, Y., Powers, N., Wiesenfeld, M., Smith, G., Farzin, S., Liu, S. “Generative AI and Extended Reality in Collaborative Architectural Design Education: An Exploratory Studio Study.” arXiv(未査読). https://arxiv.org/abs/2609.13494
  • Huh, M. B., Tural, A., Tural, E. “Students’ Generative AI Use in Studio-Based Design Education: Selective Use, Design Judgment, and Studio Context.” Education Sciences, 16(10), 1634. https://doi.org/10.3390/educsci16101634
  • Chiang, S.-B. “Effectiveness of generative AI in higher education design pedagogy: a meta-analysis of usage modes.” International Journal of Technology and Design Education. https://doi.org/10.1007/s10798-026-10122-6
  • An, B., Zhu, C. W., Dong, H. “Understanding design students’ dependency on generative AI tools: a mixed-methods investigation across reading, writing, and creative dimensions.” International Journal of Technology and Design Education. https://doi.org/10.1007/s10798-026-10113-7
  • Ahmed, M. K. M. “Reconfiguring graphic design education for generative AI in higher education: a human-agency and adaptive-curriculum framework.” Frontiers in Education, 11, 1972794. https://doi.org/10.3389/feduc.2026.1972794
  • Yiannoudes, S. “Practice readiness in generative and AI-driven architectural design: A systematic assessment.” International Journal of Architectural Computing. https://doi.org/10.1177/14780771261488679
  • Conklin, N., Capra, M., North, C. “Automating the Application of HCI Principles: Skills for On-Demand UI Construction, the Human-AI Space to Think, and the Future of HCI.” arXiv(未査読). https://arxiv.org/abs/2610.02369

Author: Shuichiro Ogawa (Design Researcher / Consultant) About me →