Notes · updated 2026-07-03
What Is the Frontier Model Premium Buying? (A Debate via Historical Analogy)
The Question and the Method
The question is this. Does paying to use the highest-performing models available right now ($100–200 per month for individuals, premium APIs for enterprises) yield future returns?
The method is a five-round debate. R1 presents historical cases where paying paid off, R2 presents cases where it did not, and each round extracts the underlying mechanisms. R3 examines which analogy the quantitative evidence from productivity RCTs and price dynamics supports. R4 confronts the disruption from the good-enough argument with practitioners’ personal views, and R5 synthesizes. No convergence is forced; the primary deliverable is a table of disagreements with resolution conditions. Related notes: ai-economics-and-design, design-pricing-vs-ai-commoditization, ai-in-design-industry, fable-model-tiering-patterns.
TL;DR
- Premium payment buys not a single good called “performance” but three distinct things: early access, the opportunity to learn frontier-premised workflows, and time.
- The effects in productivity RCTs reverse in sign depending on task and skill level, and benchmark differences do not translate directly into outcome differences.
- The cost of reaching a given level of performance falls by 50x to 200x per year, while the price of the frontier itself remains high.
- The subscription structure carries no resale loss, so the SGI- or DynaTAC-style risk of “melting assets” is structurally absent. Losses are limited to opportunity costs.
- The question therefore reduces not to “will I lose big” but to “will the learning compound,” and this point remains unresolved.
R1: Histories Where Paying Paid Off, and Their Mechanisms
Bloomberg Terminal (1982–present) has sustained premium pricing of $24,000–30,000 per year for over 40 years, with roughly 325,000 terminals in operation as of 2022. Because the traders who use it generate revenues of hundreds of thousands to millions of dollars per year, the return relative to the fee is orders of magnitude larger. Mechanism: network effects (IB chat) and lock-in, plus a fee that is minuscule relative to users’ revenues.
In early GPU adoption and deep learning (2009–2012), Krizhevsky and Sutskever trained AlexNet on two NVIDIA GTX 580s costing about $500 each and won the 2012 ImageNet competition by a 9.8pt margin over second place (170,000+ citations). The cohort of early-adopting researchers became the core talent of the subsequent AI industry. Mechanism: the hardware depreciated, but the skills and tacit knowledge of combining GPUs with the new paradigm did not depreciate and instead compounded (a first-mover advantage of 2–3 years).
Early professional adoption of DTP (1985–1995) brought phototypesetting and paste-up processes in-house for an investment of a few thousand dollars, allowing one person to complete work that had spanned multiple specialized trades. The first-mover advantage of the “designer who can do DTP” vanished within 5–10 years, but on the displaced side, prepress technicians plummeted from 34,600 in 2016 to about 5,871 in 2023 (BLS). Mechanism: workflow redesign and complementary skills (combining DTP with aesthetic judgment and client relationships) were the source of the advantage.
In factory electrification (1890s–1920s), for the two decades after the introduction of electric motors, factories that merely replaced their steam engines without changing their layouts saw no productivity gains. Factories that redesigned their layouts around unit drive in the 1920s achieved explosive gains, and electrification explains half of the productivity growth in 1920s manufacturing (David’s dynamo paper). Mechanism: a general-purpose technology does not pay off on its own; it requires complementary investment that redesigns the entire workflow.
Netflix’s early adoption of AWS (2008–2016) was one of the first large-scale cases of full cloud migration, involving a complete redesign around microservices and enabling deployment across 190 countries without in-house infrastructure. Mechanism: what proved decisive was not the cloud spending itself but complementary investment in cloud-native organizational capabilities (Chaos Engineering and the like).
The mechanisms reduce to three types. The Bloomberg type, where the fee is so small relative to revenue that the question disappears; the AlexNet type, where investment in skills and a paradigm compounds; and the dynamo type, which pays off only when paired with complementary investment and workflow redesign.
R2: Histories Where Paying Did Not Pay Off, and Their Mechanisms
In SGI and Sun workstations (1995–2009), SGI’s market capitalization fell from over $7B in 1995 to $120M in 2005 (a 98% decline) under the catch-up of commodity GPUs and Linux clusters, and the company went bankrupt in 2009. The highest performance still belonged to SGI, but lock-in collapsed through software portability (such as the porting of Maya), and the performance gap failed to convert into an outcome gap.
Concorde (1969–2003) saw development costs reach five times the initial estimate (£12–16.7B in 2024 terms), and only 14 aircraft entered service against projected orders of 150–500. The cost structure of fuel consumption 5–7 times that of the 747 was irreversible by the laws of physics; speed, the dimension of top performance, failed to reach the willingness to pay of a sufficient number of buyers, and no market formed.
Motorola DynaTAC (1983–1989) cost $3,995 ($12,800 in 2025 terms), weighed about 900g, and offered 30 minutes of talk time. It was superseded within five years, and because the hardware was purchased outright, depreciation from technological progress became a direct loss.
3D television (2010–2017) peaked at 41 million units shipped in 2012, and every manufacturer had withdrawn by 2017. At household screen sizes the performance difference did not convert into an experience difference, the complementary good of 3D content was lacking, and the glasses degraded the existing experience.
High-end PCs under Moore’s law (1990–2010s) were routinely caught up to by the next generation’s midrange within 12–18 months of purchase. For general use the midrange was sufficient, and only specific specialized uses could exploit the performance gap.
The mechanisms reduce to four. The good-enough catch-up, structures in which performance gaps do not convert into outcome gaps, the absence of complementary goods, and rapid depreciation.
R3: Which Analogy Does the Quantitative Evidence Support?
The results of productivity RCTs do not point in one direction (the two lineages debating this asymmetry of effects are set out in Does AI Make the Smart Smarter?).
- Noy & Zhang (2023, Science): ChatGPT reduced task time by 40% and raised quality ratings by 18%. Gains were larger for lower-skilled workers (preregistered RCT of 453 professionals, limited to writing tasks).
- Brynjolfsson, Li & Raymond (2025, QJE): conversational AI assistance raised productivity by 14% on average, 34% for novice workers, and only marginally for experienced workers (quasi-experiment with 5,179 support agents).
- Peng et al. (2023): Copilot made an HTTP server implementation task 55.8% faster (95% CI: 21–89%; conflict of interest from GitHub-affiliated co-authors).
- Cui et al. (2025/2026, Management Science): pooled RCTs across three companies with 4,867 developers showed a 26.08% increase in tasks completed per week (SE 10.3%; conflict of interest from Microsoft Research-affiliated co-authors; peer-reviewed).
- METR (2025): in an RCT with 16 experienced OSS developers and 246 real tasks, tasks where AI use was permitted were completed 19% slower (significant). The developers themselves predicted a 24% speedup beforehand and still misperceived a 20% speedup afterward (limited to mature large codebases; small sample).
- Dell’Acqua et al. (2025, Organization Science): in an RCT with 758 BCG consultants, tasks within AI’s capabilities saw +12.2% completion and 25.1% faster work, while tasks outside them saw a 19% drop in correctness. The sign of the effect reverses with task type (conflict of interest from BCG co-authors; peer-reviewed and preregistered).
The relationship between model compute and human outcomes is measured by Merali’s two papers. Merali (2024), in a preregistered RCT with 300 professional translators, 13 LLMs, and 1,800 tasks, showed a sublinear relationship: per 10x of compute, speed +12.3%, quality +0.18SD, and earnings per minute +16.1% (gains 4x larger for lower-skilled workers; preprint not yet peer-reviewed). Merali (2025) showed, for consulting, analysis, and management tasks (500+ participants, 13 LLMs), an 8% reduction in task time per model-generation year. While AI-alone quality improves nearly linearly with compute, human+AI quality stayed flat as model generations advanced (speed rises but quality plateaus; not yet peer-reviewed).
Price dynamics split into two layers. According to Epoch AI, the cost of reaching a given performance level is falling at 40x per year for GPT-4-equivalent performance, 50x per year at the median across six benchmarks, and at a pace of 200x per year since 2024 (independent nonprofit). By a16z’s tally, GPT-3-class performance fell 1,000x in three years ($60->$0.06/M; only the figures are adopted here given the VC conflict of interest). Meanwhile, the price of the frontier itself remains high: OpenAI o1’s output price of $60/M tokens is on par with GPT-3 at launch.
Enterprise ROI surveys tilt toward the skeptical side. MIT Project NANDA (2025) found that only about 5% of GenAI pilots achieved rapid revenue growth, while 95% had almost no measurable business impact (against $30–40B in investment; not peer-reviewed, definition of “success” and raw data undisclosed, and NANDA itself has been flagged for structural conflicts of interest). In McKinsey’s State of AI (2025), organizations reporting material financial returns (over 5% of EBIT) were 5.5% (109 of 1,993 companies; self-selected panel, interests of a consulting vendor). Gartner (2025) forecast 2025 GenAI spending at $644B (+76.4% year over year, roughly 80% toward hardware) while placing GenAI in the trough of disillusionment on its hype cycle (calculation method undisclosed).
In sum, the quantitative evidence supports neither analogy on its own. Effects at the individual task level are real but depend strongly on task and skill level (consistent with the dynamo type’s “conditional on complements”), while enterprise-level failure rates and plunging prices are consistent with the good-enough analogy.
R4: The Good-Enough Attack and Where Practitioners Stand Now
The good-enough argument attacks, in turn, four implicit premises of the case for paying: that performance gaps convert into outcome gaps, that performance gaps persist, that skills compound, and that users can perceive the difference.
Compression of the performance gap. The closed-versus-open gap on MMLU shrank from 17.5pt at the end of 2023 to 0.3pt within a year, and by early 2026 it was effectively zero. The UK AISI reports that the lag for open models to catch up to closed frontier performance has shortened to 4–8 months; on this view, the essence of premium payment is a right of early access. The provider’s own projections point the same way: ChatGPT’s free-to-paid conversion rate remains at 5–6%, and OpenAI projects Plus subscribers to fall 80% from 44 million in 2025 to 9 million in 2026 (a shift to a low-priced ad-supported plan, with ARPU going from about $23 to under $12).
Depreciation of skills. Prompt engineering has all but disappeared as a standalone occupation (Fast Company, May 2025 reporting; 68% of companies have made it standard training across all roles), and because optimal practice differs by model, the skill is model-specific and depreciates. On this view, skills are consumables, not compounding assets.
As parallels: camera shipments fell 94% from 2010 levels as smartphones became good enough; high-end audio faces diminishing returns in which perceptible differences nearly vanish at the top price range; and LexisNexis, its general market eroded by free Google search, retreated to a niche (exclusively licensed content).
That said, the good-enough side itself lists the conditions under which its attack does not hold. Professionals whose hourly rates are high enough that the value of early access exceeds $2,400 per year; occupations whose core work rests on reasoning or agent pipelines that only function at the frontier; domains where a single quality difference translates directly into order-of-magnitude revenue differences; and enterprise Team/Enterprise plans purchased for security and compliance rather than performance (12-month retention of 88% for Enterprise versus 59% for Plus).
Practitioners’ personal views need to be read as changing over time. Ethan Mollick said on 2024-10-03 that “the bigger model always wins” (an extrapolation from the single case of BloombergGPT underperforming GPT-4), but on 2025-06-26 revised his position to “it isn’t about the best model, it is about the best overall system.” At the same time, he holds that free tiers are “demos, not tools” and that the paid $20/month is the effective floor, and on 2025-10-19 he concluded that complex technical and coding uses require about $200/month. Simon Willison, as of 2025-12-31, wrote that local lightweight models have far exceeded his expectations but that he has “not yet encountered” a local model he can trust for coding-agent use, rated the flat $200/month plan a “substantial discount” relative to metered API pricing, and said he plans to move from the $100 plan to $200 himself (his “I hear from many practitioners equally happy to pay” is hearsay). Paul Gauthier, on 2025-01-26, offered an example of a heuristic that is not model-specific (one that transfers): adherence degrades across multiple models once context exceeds roughly 25,000–30,000 tokens. swyx, on 2026-06-09, cited the result that the top model Opus 4.8 reaches only 13% on the hardest subset of Cognition’s FrontierCode benchmark, presenting it as “coding is not solved.” If the ceiling is still low, room for frontier improvement remains.
Both Mollick and Willison vary their positions by period and use case, and the debate cannot be reduced to a simple binary of “always use the best” versus “cheap is enough.”
R5: Synthesis (The Three Distinct Things Payment Is Buying)
Premium payment buys not a single good called “performance” but three distinct things.
- Early access: frontier performance descends to the free tier within 4–8 months. What is being bought is the time lag.
- Learning opportunity: the chance to learn, ahead of others, workflows that only function at the frontier. Whether this compounds AlexNet-style depends on whether the skills transfer across model generations.
- Time: a sublinear speedup of +12.3% per 10x of compute.
The crux of the historical analogy lies in the difference in purchase form. Unlike the era of hardware purchases, subscriptions carry no resale loss, and losses are limited to opportunity costs. The question therefore reduces not to “will I lose big” but to “will the learning in (2) compound.”
The dynamo lesson applies here. (2) pays off only when paired with complementary investment (workflow redesign). The Bloomberg lesson concerns the relativity of the fee: if the fee is 1–2% or less of income, the question itself disappears.
Points of Agreement
Only the points on which all perspectives agreed are listed.
- Benchmark differences do not translate directly into outcome differences (grounded in Dell’Acqua’s sign reversal and METR’s negative result).
- The cost of reaching a given performance level keeps plunging (Epoch AI’s and a16z’s estimates differ in magnitude but point the same way).
- The effects of AI use depend strongly on task and skill level (a finding common across the RCTs).
- The subscription structure carries less downside risk than the historical hardware-purchase cases (because there is no resale loss and losses are limited to opportunity costs).
Unresolved Disagreements (Not Smoothed Over; the Primary Deliverable)
| Issue | Position A | Position B | Resolution Condition |
|---|---|---|---|
| (1) Are skills compounding assets or consumables? | AlexNet type: paradigm mastery does not depreciate and compounds. Some knowledge transfers, like Gauthier’s heuristic | Prompting techniques are model-specific and become obsolete, and get abstracted into libraries as the baseline | Longitudinal data on skill persistence across model generations |
| (2) Who captures the value of early access? | Limited to high-hourly-rate professionals for whom the value exceeds $2,400 per year | Broadly distributed as a learning opportunity (consistent with the RCTs showing larger gains for lower-skilled workers) | Effect measurement by occupation and by income |
| (3) Will frontier prices stay high? | As o1’s price on par with GPT-3 at launch shows, the top tier has pricing power | The open-model catch-up (4–8 months) erodes that pricing power | Continued observation of the frontier tier’s price time series |
| (4) Is the 95% failure rate of enterprise adoption due to “insufficient models” or “lack of complementary investment”? | Dynamo analogy: deployments fail because workflow redesign is missing | Good-enough argument: there are structures in which adding models does not convert into outcome differences | Comparison of adoption outcomes stratified by presence or absence of complementary investment |
| (5) Does the plateau in human+AI quality render the premium differential meaningless? | Merali (2025): if quality plateaus, the differential buys only speed and is of little value | The plateau is a constraint of current human-side workflows and will be lifted by agentification | Replication decomposing the cause of the plateau (human-side bottleneck or model-side) |
Conclusions Not to Adopt
- “History repeats, so it will definitely pay off (or definitely be wasted)”: historical cases are material for extracting mechanisms, not sources from which to substitute conclusions.
- “Benchmark difference = outcome difference”: the RCTs show reversals even in the sign of the effect.
- “95% fail, so paying is a waste”: NANDA’s definition of “success” is undisclosed, and conflicts of interest have been flagged.
- “It’s a subscription, so zero risk”: there is no resale loss, but opportunity costs remain.
Items to Verify
- Whether empirical data on skill transferability (longitudinal studies across model generations) exists.
- The peer-review outcomes of Merali (2024, 2025).
- The disclosure status of MIT NANDA’s definition of “success” and its raw data.
- Continued observation of AISI’s catch-up lag (4–8 months).
- The actual figures against OpenAI’s Plus subscriber projection (an 80% decline).
References
Academic Literature
- David, P. A. (1990). The Dynamo and the Computer: An Historical Perspective on the Modern Productivity Paradox. American Economic Review, 80(2), 355-361. http://www.jstor.org/stable/2006600
- Noy, S. & Zhang, W. (2023). Experimental Evidence on the Productivity Effects of Generative Artificial Intelligence. Science. https://pubmed.ncbi.nlm.nih.gov/37440646/
- Brynjolfsson, E., Li, D. & Raymond, L. (2025). Generative AI at Work. Quarterly Journal of Economics. https://www.nber.org/papers/w31161
- Peng, S. et al. (2023). The Impact of AI on Developer Productivity: Evidence from GitHub Copilot. arXiv. https://arxiv.org/abs/2302.06590
- Cui, Z. et al. (2025/2026). The Effects of Generative AI on High-Skilled Work: Evidence from Three Field Experiments with Software Developers. Management Science. https://pubsonline.informs.org/doi/10.1287/mnsc.2025.00535
- METR (2025). Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity. https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/ / https://arxiv.org/abs/2507.09089
- Dell’Acqua, F. et al. (2025). Navigating the Jagged Technological Frontier. Organization Science. https://pubsonline.informs.org/doi/10.1287/orsc.2025.21838
- Merali, A. (2024). Scaling Laws for Economic Productivity: Experimental Evidence in LLM-Assisted Translation. arXiv. https://arxiv.org/abs/2409.02391
- Merali, A. (2025). Scaling Laws for Economic Productivity: Experimental Evidence in Consulting, Data Analyst, and Management Tasks. arXiv. https://arxiv.org/abs/2512.21316
- Flamm, K. (2018). Measuring Moore’s Law: Evidence from Price, Cost, and Quality Indexes. NBER Working Paper 24553. https://www.nber.org/system/files/working_papers/w24553/w24553.pdf
- Finding News Stories: A Comparison of Searches Using LexisNexis and Google News. https://www.researchgate.net/publication/241655410_Finding_News_Stories_A_Comparison_of_Searches_Using_Lexisnexis_and_Google_News
Surveys and Reports
- Epoch AI. LLM inference price trends. https://epoch.ai/data-insights/llm-inference-price-trends
- a16z. LLMflation: LLM inference cost is going down fast. https://a16z.com/llmflation-llm-inference-cost/
- MIT Project NANDA (2025). State of AI in Business 2025. https://mlq.ai/media/quarterly_decks/v0.1_State_of_AI_in_Business_2025_Report.pdf
- McKinsey (2025). The State of AI. https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai
- Gartner (2025). Gartner Forecasts Worldwide GenAI Spending to Reach $644 Billion in 2025. https://www.gartner.com/en/newsroom/press-releases/2025-03-31-gartner-forecasts-worldwide-genai-spending-to-reach-644-billion-in-2025
- UK AI Security Institute. Frontier AI Trends Report. https://www.aisi.gov.uk/frontier-ai-trends-report
- U.S. Bureau of Labor Statistics. OES: Prepress Technicians and Workers. https://www.bls.gov/oes/current/oes515111.htm
Web Articles
- Bloomberg Terminal — Wikipedia (https://en.wikipedia.org/wiki/Bloomberg_Terminal)
- Bloomberg’s 7 Powers and Why the Terminal Endures — The Terminalist (https://theterminalist.substack.com/p/bloombergs-7-powers-and-why-the-terminal)
- AlexNet Source Code — IEEE Spectrum (https://spectrum.ieee.org/alexnet-source-code)
- CHM Releases AlexNet Source Code — Computer History Museum (https://computerhistory.org/blog/chm-releases-alexnet-source-code/)
- Desktop publishing — Wikipedia (https://en.wikipedia.org/wiki/Desktop_publishing)
- The history of prepress — Prepressure (https://www.prepressure.com/prepress/history)
- Netflix AWS Migration — Bacancy Technology (https://www.bacancytechnology.com/blog/netflix-aws-migration)
- Silicon Graphics — Wikipedia (https://en.wikipedia.org/wiki/Silicon_Graphics)
- The Rise and Fall of Silicon Graphics — Hackaday (https://hackaday.com/2024/04/08/the-rise-and-fall-of-silicon-graphics/)
- Why Did Supersonic Airliners Fail? — Construction Physics (https://www.construction-physics.com/p/why-did-supersonic-airliners-fail)
- Concorde — Wikipedia (https://en.wikipedia.org/wiki/Concorde)
- DynaTAC — Wikipedia (https://en.wikipedia.org/wiki/DynaTAC)
- 3D TV is finally, officially dead — Digital Journal (https://www.digitaljournal.com/tech-science/3d-tv-is-finally-officially-dead/article/484327 , now 404. Archive: https://web.archive.org/web/20260121200512/https://www.digitaljournal.com/tech-science/3d-tv-is-finally-officially-dead/article/484327 )
- Moore’s law — Wikipedia (https://en.wikipedia.org/wiki/Moore%27s_law)
- AI Price Index — tokencost (https://tokencost.app/blog/ai-price-index)
- Falling LLM Token Prices and What They Mean for AI Companies — The Batch, DeepLearning.AI (https://www.deeplearning.ai/the-batch/falling-llm-token-prices-and-what-they-mean-for-ai-companies)
- Open Source vs Closed LLMs: Choosing the Right Model in 2026 — Let’s Data Science (https://letsdatascience.com/blog/open-source-vs-closed-llms-choosing-the-right-model-in-2026)
- ChatGPT Statistics — Second Talent (https://www.secondtalent.com/resources/chatgpt-statistics/)
- OpenAI Projects ChatGPT Plus Subscriptions to Drop by 80% — Where’s Your Ed At (https://www.wheresyoured.at/openai-projects-chatgpt-plus-subscriptions-to-drop-by-80-from-44-million-in-2025-to-9-million-in-2026-made-up-using-cheaper-subscriptions-somehow/)
- Context Engineering Guide 2026 — Jobs by Culture (https://jobsbyculture.com/blog/context-engineering-guide-2026)
- The Camera Market Is Shrinking, But That’s Not the Story — Fstoppers (https://fstoppers.com/originals/camera-market-shrinking-thats-not-story-902590)
- The Value Proposition of Diminishing Returns in Audio — Audiophile Review (https://audiophilereview.com/audiophile/the-value-proposition-of-diminishing-returns-in-audio/) [requires source verification: original URL is 404, new location unconfirmed. Wayback archive: https://web.archive.org/web/20251212133850/https://audiophilereview.com/audiophile/the-value-proposition-of-diminishing-returns-in-audio/]
- Interview with Ethan Mollick — Evident Insights (https://evidentinsights.com/bankingbrief/interview-with-ethan-mollick/)
- Using AI Right Now: A Quick Guide — One Useful Thing, Ethan Mollick (https://www.oneusefulthing.org/p/using-ai-right-now-a-quick-guide)
- An Opinionated Guide to Using AI — One Useful Thing, Ethan Mollick (https://www.oneusefulthing.org/p/an-opinionated-guide-to-using-ai)
- The Year in LLMs — Simon Willison (https://simonwillison.net/2025/Dec/31/the-year-in-llms/)
- Paul Gauthier on Long Context — Simon Willison (https://simonwillison.net/2025/Jan/26/paul-gauthier/)
- AINews: FrontierCode Benchmarking — Latent Space, swyx (https://www.latent.space/p/ainews-frontiercode-benchmarking)