Reports indicate that OpenAI engineers have found a way to reduce the cost of AI inference by more than half.



A challenge for companies that provide AI models on a large scale is

the inference cost involved in the process of how AI models generate output in response to user input. Now, The Information reports that engineers at OpenAI have said they have found a way to reduce inference costs by more than half.

OpenAI Discovers New Way to Cut Inference Costs in Half — The Information
https://www.theinformation.com/newsletters/ai-agenda/openai-discovers-new-way-cut-inference-costs-half

OpenAI engineers say they've more than halved inference costs | AI Weekly
https://aiweekly.co/alerts/openai-engineers-say-theyve-more-than-halved-inference-costs

Inference costs are expenses incurred each time an AI model is actually used, and they constitute a significant portion of the operational costs for AI development companies that provide chat AI and coding AI at scale. Therefore, finding ways to reduce inference costs is a crucial challenge in providing AI models at a lower cost and with higher profit margins.

Edward Zitron, a technology industry expert, estimates that OpenAI will spend more than $5 billion (approximately 814 billion yen) on inference costs in the first half of 2025 alone. This amount is said to significantly exceed OpenAI's projected revenue.

How expensive is OpenAI's inference cost? - GIGAZINE



The Information reported, citing sources familiar with the issue, that in early June 2026, OpenAI engineers told colleagues that they had found a way to reduce inference costs by more than half using a new optimization technique.

While the specific optimization method used is unclear, an OpenAI representative reportedly stated that 'when the optimization method was applied to ChatGPT guest users, the number of NVIDIA GPUs required for part of the processing was reduced to just around 200.'

While training a cutting-edge AI model is a one-time event, inference costs are incurred at every step of the AI ​​agent's operation, such as responding to chat messages or making API calls. Therefore, significantly reducing the number of GPUs used in the free tier simply by changing the software can lead to operational cost reductions that cannot be achieved through hardware contract optimization alone.

Furthermore, given that the reduction in inference costs is attributed to improved utilization efficiency of existing servers, AI Weekly, an AI-related media outlet, speculates that possible methods include 'smarter batch processing,' 'improved cache reusability,' ' quantization ,' and 'routing simpler queries to less expensive models.' However, they point out that it is unclear whether the same methods are available to ChatGPT users with free or paid accounts who are not guest users.

AI Weekly stated, 'If this can be generalized, OpenAI could either lower prices, expand free access, or absorb more agent workloads without purchasing additional chips. The last option is particularly noteworthy because, as the entire industry races to build AI data centers, it is the cheapest way to maintain profit margins.'



in AI, Posted by log1h_ik