AI performance is getting dramatically cheaper, reports say

The cost of achieving a given level of AI performance is falling at an unusually fast pace, according to two recent analyses. That shift could make AI a routine component of software tools rather than a separately priced service, affecting everything from content workflows to search, security, and internal business systems.

Alex Tabarrok writes for Marginal Revolution that an Epoch AI report by Emberson and Roodman estimates that the cost of a fixed level of AI performance has declined by an average of 47 percent per quarter over the past three years. That would equal an approximately 13-fold annual reduction, though such estimates depend on the benchmarks and models selected.

The report’s example illustrates the scale of the claimed change. Tabarrok says OpenAI’s o3 model cost about 30 cents per question to reach 75 percent on the GPQA Diamond benchmark in early 2025. A newer model, GPT-5.6 Luna, reportedly reached a similar score for $0.0004 per question by mid-2026. That is a roughly 725-fold decline in less than 18 months.

The key point is not simply that models improve. Older levels of capability are also becoming much cheaper to deliver. For organizations using AI, that could change purchasing decisions. A task that once required a premium model may increasingly be handled by a smaller, less expensive model with comparable results.

Several forces are reducing AI costs

In an analysis published on jyn.dev, developer jyn514 argues that lower costs come from several improvements that reinforce one another:

  • More efficient hardware: GPUs are improving their performance per unit of energy. The article cites a pace of roughly doubling energy efficiency every two years.
  • Better models: Model developers are improving the cost of completing a task, not only the price of individual tokens. More capable models may need fewer attempts, corrections, or reasoning steps.
  • Improved inference software: The systems that run AI models on servers are becoming more efficient. The article cites gains of roughly 10 to 50 percent per year for some inference engines.
  • New model architectures: Mixture-of-Experts designs activate only parts of a model for each request, reducing the computation needed for many tasks.

The jyn.dev analysis estimates that model cost efficiency per task has improved by around 100 times over the past year, while hardware and serving software add further gains. It also points to models designed for narrow classification tasks. Such systems can return structured judgments, such as whether a command appears risky, without generating long text responses. Their costs can be far lower than general-purpose chat models.

Cost will not be the only limit

Lower inference prices do not automatically mean all AI work becomes inexpensive or reliable. Quality, access to proprietary frontier models, data protection, integration effort, and human review remain important constraints. The jyn.dev article also notes that specialized systems may require technical setup or fine-tuning before they perform well in a specific environment.

Both analyses expect cheaper intelligence to increase overall demand for computing resources. That could favor cloud providers and companies operating large AI infrastructure, even if the price of an individual request keeps falling. It could also put pressure on software vendors whose products offer functions that users can increasingly reproduce with AI-generated, customized tools.

The immediate implication is practical: organizations may need to evaluate AI less as a scarce premium service and more as a low-cost building block. The differentiators may shift toward trustworthy outputs, useful workflows, proprietary data, security, and the ability to turn abundant AI capacity into dependable results.

Sources

Stay up to date

AI for content creation: the latest tools, tips and trends. Every two weeks in your inbox:

More info …

About the author

Related posts:

Advertisement

×