Notes · updated 2026-08-07
Scope and Method
This note organizes, by contention point, the debate between the four lineages of deliberately designed failure identified in Designing Failure Into Learning (productive failure (PF), desirable difficulties, impasse-driven learning, error management training (EMT)) and the opposing school of cognitive load theory (CLT, the camp centered on Kirschner, Sweller & Clark).
To the existing 13 works, 14 were added covering the debate’s responses, direct confrontations, and syntheses (the 2007 journal response symposium, CLT’s empirical foundations and direct experiments, arbitrating meta-analyses, and the recent integration controversy), for a 27-work corpus (source/review/failure-driven-learning/papers.md).
Each side’s claims are reported as the papers’ claims, not asserted by this note.
The Shape of the Debate (Chronology)
The debate has two peaks. As a forerunner, Mayer (2004) reviewed three generations of pure discovery learning research, all losing to guided discovery, laying CLT’s groundwork. The first wave began when Kirschner, Sweller & Clark (2006; hereafter KSC) lumped constructivist, discovery, problem-based, and inquiry learning together as “minimal guidance” and pronounced it ineffective; in the same issue of Educational Psychologist (42(2), 2007), Schmidt et al., Hmelo-Silver et al., and Kuhn responded, and Sweller et al. issued a rejoinder, making it a de facto journal symposium. The controversy crystallized in the Tobias & Duffy (2009) edited volume, where the main figures of both camps engaged chapter by chapter. The second wave is a direct confrontation of experiments and meta-analyses. The PF side’s meta-analysis of 53 studies (Sinha & Kapur 2021) showed a moderate significant effect for problem-solving-first (PS-I), while the CLT side’s randomized trials (Ashman et al. 2020) reported the opposite. Since 2022, an integration controversy continues: against Zhang, Kirschner, Cobern & Sweller (2022, unverified) claiming general superiority of direct instruction, thirteen authors led by de Jong (2023) responded that combining inquiry and direct instruction is optimal and the oppositional framing itself is mistaken.
Contention 1: Timing of Instruction (Problem Solving First, or Instruction First?)
This is the most direct empirical contention.
The failure-design claim is that unsupported problem solving (in which learners fail) generates awareness of knowledge gaps and activates diverse solutions, preparing the reception of subsequent instruction (Kapur 2008; Kapur & Bielaczyc 2012). Empirically, Schwartz & Martin (2004) reported that invention activities before formula instruction enhance learning from that instruction, and Sinha & Kapur’s (2021) meta-analysis (53 studies, 166 comparisons) reported a moderate significant effect for PS-I, increasing with PF design fidelity.
The CLT claim is that exploring high element-interactivity content without instruction overloads working memory and impedes learning. The empirical base is the worked-example effect (Sweller & Cooper 1985: studying worked solutions cut test errors to roughly one fifth compared with problem-solving practice), and Ashman et al. (2020) directly tested the PS-I design in randomized trials, reporting that with high element interactivity, instruction-first (I-PS) was superior on both similar and transfer problems.
The empirical state is squarely split: a meta-analysis (PS-I superior) and RCTs (I-PS superior) coexist. Loibl et al.’s (2017) boundary-condition review offers the mediating position: PS-I works when the prior problem solving generates contrasting cases and the instruction builds on learners’ solutions, and the effect vanishes otherwise. Glogger-Frey et al. (2015) showed that invention and worked examples prepare different things (the former, gap awareness and curiosity; the latter, transfer and self-efficacy), beginning to decompose the either-or question itself.
Contention 2: Is the “Minimal Guidance” Framing Valid? (The Definitional Contention)
The crux of KSC (2006) was lumping constructivist, discovery, PBL, and inquiry approaches together as minimal guidance, ineffective for novices. The respondents’ claim is that this lumping is a straw man. Schmidt et al. (2007) argued that PBL is in fact heavily scaffolded instruction compatible with CLT’s account of cognitive architecture, and Hmelo-Silver et al. (2007) countered with the richness of scaffolding in inquiry learning and empirical achievement evidence. Kuhn (2007) shifted the question itself: direct instruction may transmit declarative knowledge, but does not answer the educational goal of developing the capacity to inquire.
Sweller et al.’s rejoinder (2007): if PBL provides scaffolds, it has merely moved toward the guidance CLT recommends; the matter should be settled by randomized trials.
The empirical state has Alfieri et al.’s (2011) two meta-analyses in the arbitrating position: unassisted discovery is inferior to explicit instruction (d = −0.38), while enhanced discovery with feedback and scaffolding is superior (d = +0.30). Both KSC’s claim (unsupported discovery fails) and the respondents’ claim (discovery works when scaffolded) receive partial support, and the contention has effectively shifted to the dosage question of how much support is enough.
Contention 3: What Counts as an Effect? (The Measurement Contention)
This contention explains why the same experimental results are read to opposite conclusions.
The desirable-difficulties claim is that performance during and immediately after training misleads as an index of learning. Soderstrom & Bjork’s (2015) integrative review organized the dissociation between learning (durable change) and performance (momentary execution), with evidence that difficult conditions depress performance while enhancing retention and transfer. The PF side makes the parallel claim: across Kapur’s reports, procedural knowledge equals direct instruction while conceptual understanding and transfer favor PF, implying PF’s advantage is invisible on immediate similar-task measures.
The CLT side measures both immediate and transfer outcomes, and Ashman et al. (2020) reported I-PS superiority on transfer problems too, answering the measurement objection empirically.
The empirical state: both camps increasingly acknowledge that measurement choices (immediate versus delayed, similar versus transfer, procedural versus conceptual) partly explain the direction of results, but the conflict between “I-PS wins even on transfer” (Ashman et al.) and “PS-I wins precisely on transfer” (Schwartz & Martin; Sinha & Kapur) remains unresolved.
Contention 4: Learners’ Prior Knowledge (Difficult for Whom?)
This contention is the closest to convergence.
The CLT claim is the expertise reversal effect (Kalyuga et al. 2003): strong guidance such as worked examples helps novices but becomes redundant and then harmful as knowledge grows, so problem solving becomes effective for the more expert. CLT itself thus predicts failure-laden problem solving for experts.
The failure-design response is Kapur’s (2016) 2×2 framework (success/failure × productive/unproductive): failure is productive only when learners have sufficient prior resources (intuitions, everyday knowledge) to explore the problem space and when post-failure consolidation is designed, explicitly conditionalizing the claim and partially absorbing CLT’s critique. The EMT meta-analysis (Keith & Frese 2008) likewise reports moderators (larger effects for complex tasks and transfer measures) that do not support uniform application to simple tasks or novices.
The empirical state: that prior knowledge changes the optimal instructional form is common ground. The remaining conflict narrows to whether designed failure works for novices lacking canonical prior knowledge: PF reports that intuitions and everyday knowledge suffice; CLT reports that with high element interactivity they do not (the same root as Contention 1).
Contention 5: The Role of Affect (The Cost of Failure and Motivation)
The EMT claim is that the emotional cost of failure can and must be managed by design. Keith & Frese’s (2008) meta-analysis identified explicit emotion-regulation messaging (“errors are informative”) and metacognitive activity as mediators of EMT’s effect (d ≈ 0.44). The PF side has recently absorbed this: Kapur’s (2024) 4A model places Affect among the design elements. Glogger-Frey et al. (2015) reported an affective differentiation (invention raises curiosity and interest; worked examples raise self-efficacy), so affective effects vary by activity type.
The CLT side carries a theoretical concern about frustration and motivational loss, but no systematic empirical work focused on this was confirmed in this collection. Affect is the contention with the thinnest evidence on both sides.
Contention Table (Unresolved Conflicts)
| Contention | Failure-design claim (source) | CLT claim (source) | Empirical state | Evidence needed to adjudicate |
|---|---|---|---|---|
| Timing of instruction | The failure phase prepares reception of instruction and improves transfer (Kapur 2008; Sinha & Kapur 2021 meta-analysis) | With high element interactivity, instruction-first is superior (Sweller & Cooper 1985; Ashman et al. 2020 RCTs) | Meta-analysis and RCTs conflict; boundary conditions (Loibl et al. 2017) are the mediating candidate | Pre-registered RCTs jointly manipulating element interactivity and design fidelity |
| The “minimal guidance” framing | PBL/inquiry are richly scaffolded, not minimal (Schmidt et al. 2007; Hmelo-Silver et al. 2007) | Admitting scaffolds concedes movement toward guided instruction (Sweller et al. 2007) | Alfieri et al. 2011 partially supports both (unassisted inferior, assisted superior) | An operational definition of “sufficient support” and dose-response evidence |
| Measurement of effects | Immediate performance misleads; assess retention and transfer (Soderstrom & Bjork 2015) | Instruction-first can win even on transfer (Ashman et al. 2020) | Measurement choice partly explains result direction; the transfer conflict is unresolved | Experiments fully crossing immediate/delayed and similar/transfer in one design |
| Application to novices | Failure is productive given intuitions plus designed consolidation (Kapur 2016) | Novices without canonical knowledge need guidance (Kalyuga et al. 2003; KSC 2006) | Prior-knowledge moderation is common ground; the novice threshold is unresolved | Interaction estimates with continuously measured prior knowledge |
| Role of affect | Designed emotion regulation mediates the effect (Keith & Frese 2008; Kapur 2024) | Concern about frustration and motivational loss (no systematic evidence confirmed) | Thinnest evidence on both sides; affective differentiation by activity type (Glogger-Frey et al. 2015) is the lead | Experiments manipulating emotion-regulation support in failure designs; dropout measurement |
Where the Debate Stands: From Which Is Right to Designing the Integration
The debate is moving away from the discovery-versus-instruction dichotomy. de Jong et al. (2023) argue for combining inquiry and direct instruction and reject the oppositional framing (the CLT response to this is unverified). PF places post-failure consolidation at the core of its design (Kapur & Bielaczyc 2012), and CLT predicts problem solving for the more expert (Kalyuga et al. 2003), so the two camps’ practical prescriptions overlap substantially. The remaining essential conflict condenses to the single point shared by Contentions 1 and 4: whether designed failure-first works for novices without canonical knowledge. On that point, the meta-analysis and the randomized trials give opposite answers, and the matter is not settled.
Where the industry services (those explicitly stating deliberate under-instruction) sit in this debate is organized in Designing Failure Into Learning; the implementation reality that most services lack post-failure consolidation connects directly to the implications of Contentions 1 and 4 (failure designs that fail the consolidation and prior-knowledge conditions lose their theoretical backing).
Gaps
- No direct debate literature was found between impasse-driven learning (VanLehn et al. 2003) and CLT; the impasse hypothesis is invoked as a mechanism for PF, with no confirmed direct rebuttal from CLT.
- No direct confrontation between EMT and CLT was found either; EMT’s arena is workplace training and CLT’s is schooling, with no bibliographic engagement.
- Zhang, Kirschner, Cobern & Sweller (2022) and the Sweller-side response to de Jong et al. (2023) are not yet in this corpus (candidates for a follow-up addendum on the latest round).
- No empirical work focused on affect (Contention 5) in failure designs was confirmed from either camp.
Unverified Items
- Bibliography of Zhang, Kirschner, Cobern & Sweller (2022) (reportedly DOI: 10.1007/s10648-021-09646-1; full text unconfirmed)
- DOI and content of the Sweller-side response to de Jong et al. (2023)
- Chapter contents of Tobias & Duffy (2009) (requires access to the book)
- Full text of Sweller & Cooper (1985) (paywalled; confirmed via bibliography and abstract)
- Chapter pages of Bjork (1994) (no DOI; confirmed by ISBN; carried over from the parent note)
References
All accessed 2026-08-07. For ledger details, see source/review/failure-driven-learning/papers.md. For the foundational works of the four lineages, see also the references of Designing Failure Into Learning.
Debate responses, direct confrontations, and syntheses (addendum)
- Schmidt, H. G., Loyens, S. M. M., Van Gog, T., & Paas, F. (2007). Problem-Based Learning is Compatible with Human Cognitive Architecture. Educational Psychologist, 42(2), 91–97. https://doi.org/10.1080/00461520701263350
- Hmelo-Silver, C. E., Duncan, R. G., & Chinn, C. A. (2007). Scaffolding and Achievement in Problem-Based and Inquiry Learning. Educational Psychologist, 42(2), 99–107. https://doi.org/10.1080/00461520701263368
- Kuhn, D. (2007). Is Direct Instruction an Answer to the Right Question? Educational Psychologist, 42(2), 109–113. https://doi.org/10.1080/00461520701263376
- Sweller, J., Kirschner, P. A., & Clark, R. E. (2007). Why Minimally Guided Teaching Techniques Do Not Work: A Reply to Commentaries. Educational Psychologist, 42(2), 115–121. https://doi.org/10.1080/00461520701263426
- Sweller, J., & Cooper, G. A. (1985). The Use of Worked Examples as a Substitute for Problem Solving in Learning Algebra. Cognition and Instruction, 2(1), 59–89. https://doi.org/10.1207/s1532690xci0201_3
- Kalyuga, S., Ayres, P., Chandler, P., & Sweller, J. (2003). The Expertise Reversal Effect. Educational Psychologist, 38(1), 23–31. https://doi.org/10.1207/S15326985EP3801_4
- Mayer, R. E. (2004). Should There Be a Three-Strikes Rule Against Pure Discovery Learning? American Psychologist, 59(1), 14–19. https://doi.org/10.1037/0003-066X.59.1.14
- Ashman, G., Kalyuga, S., & Sweller, J. (2020). Problem-solving or Explicit Instruction: Which Should Go First When Element Interactivity Is High? Educational Psychology Review, 32, 229–247. https://doi.org/10.1007/s10648-019-09500-5
- Glogger-Frey, I., Fleischer, C., Grüny, L., Kappich, J., & Renkl, A. (2015). Inventing a Solution and Studying a Worked Solution Prepare Differently for Learning from Direct Instruction. Learning and Instruction, 39, 72–87. https://doi.org/10.1016/j.learninstruc.2015.05.001
- Alfieri, L., Brooks, P. J., Aldrich, N. J., & Tenenbaum, H. R. (2011). Does Discovery-Based Instruction Enhance Learning? Journal of Educational Psychology, 103(1), 1–18. https://doi.org/10.1037/a0021017
- Kapur, M. (2016). Examining Productive Failure, Productive Success, Unproductive Failure, and Unproductive Success in Learning. Educational Psychologist, 51(2), 289–299. https://doi.org/10.1080/00461520.2016.1155457
- Soderstrom, N. C., & Bjork, R. A. (2015). Learning Versus Performance: An Integrative Review. Perspectives on Psychological Science, 10(2), 176–199. https://doi.org/10.1177/1745691615569000
- de Jong, T., Lazonder, A. W., Chinn, C. A., et al. (2023). Let’s Talk Evidence – The Case for Combining Inquiry-Based and Direct Instruction. Educational Research Review, 39, 100536. https://doi.org/10.1016/j.edurev.2023.100536
- Tobias, S., & Duffy, T. M. (Eds.) (2009). Constructivist Instruction: Success or Failure? Routledge. ISBN 9780415994231. (non-peer)
Foundational works of the four lineages (main items from the existing corpus)
- Kapur, M. (2008). Productive Failure. Cognition and Instruction, 26(3), 379–424. https://doi.org/10.1080/07370000802212669
- Kapur, M., & Bielaczyc, K. (2012). Designing for Productive Failure. Journal of the Learning Sciences, 21(1), 45–83. https://doi.org/10.1080/10508406.2011.591717
- Sinha, T., & Kapur, M. (2021). When Problem Solving Followed by Instruction Works: Evidence for Productive Failure. Review of Educational Research, 91(5), 761–798. https://doi.org/10.3102/00346543211019105
- Bjork, R. A. (1994). Memory and metamemory considerations in the training of human beings. In Metacognition: Knowing about Knowing (pp. 185–205). MIT Press. ISBN 0-262-13298-2.
- VanLehn, K., et al. (2003). Why Do Only Some Events Cause Learning During Human Tutoring? Cognition and Instruction, 21(3), 209–249. https://doi.org/10.1207/s1532690xci2103_01
- Schwartz, D. L., & Martin, T. (2004). Inventing to Prepare for Future Learning. Cognition and Instruction, 22(2), 129–184. https://doi.org/10.1207/s1532690xci2202_1
- Keith, N., & Frese, M. (2008). Effectiveness of Error Management Training: A Meta-Analysis. Journal of Applied Psychology, 93(1), 59–69. https://doi.org/10.1037/0021-9010.93.1.59
- Kirschner, P. A., Sweller, J., & Clark, R. E. (2006). Why Minimal Guidance During Instruction Does Not Work. Educational Psychologist, 41(2), 75–86. https://doi.org/10.1207/s15326985ep4102_1
- Loibl, K., Roll, I., & Rummel, N. (2017). Towards a Theory of When and How Problem Solving Followed by Instruction Supports Learning. Educational Psychology Review, 29(4), 693–715. https://doi.org/10.1007/s10648-016-9379-x
- Kapur, M. (2024). Productive Failure. Jossey-Bass. https://onlinelibrary.wiley.com/doi/book/10.1002/9781394308712 (non-peer)