IT Brief New Zealand - Technology news for CIOs & IT decision-makers
New Zealand
OpenAI cuts GPT-5.6 prices in push for wider access

OpenAI cuts GPT-5.6 prices in push for wider access

Tue, 4th Aug 2026 (Today)
Mark Tarre
MARK TARRE News Chief

OpenAI has cut prices for two GPT-5.6 models and outlined a full-stack strategy to expand access to advanced AI. The approach is meant to make useful AI cheaper and more widely used.

The changes include an 80 per cent price cut for GPT-5.6 Luna and a 20 per cent cut for GPT-5.6 Terra. Luna now costs USD $0.20 per million input tokens and USD $1.20 per million output tokens. Terra is priced at USD $2 per million input tokens and USD $12 per million output tokens.

OpenAI also said GPT-5.6 Sol now has a Fast mode that delivers up to 2.5 times the speed of standard processing at twice the price, with no change in model intelligence. It presented the move as part of a broader effort to give customers more flexibility on cost, speed and model choice.

The announcement also set out the economics behind that strategy. OpenAI argued that lower costs make more work commercially viable, while stronger models increase the value of that work. That creates a cycle in which broader adoption supports further spending on research and infrastructure.

Sarah Friar, Chief Financial Officer at OpenAI, said the focus should be on the practical cost of outcomes rather than the sticker price of tokens alone.

"Customers do not buy tokens for their own sake. They want the support issue resolved, the software shipped, the contract reviewed, or the scientific question answered. The right measure is the cost of a successful outcome, including the time, retries, oversight, and errors required to get there. A stronger model that completes the work correctly and efficiently can, at the end, be more economical than a cheaper model that requires repeated attempts or extensive human intervention. Conversely, a lower-cost model can dramatically expand access when it meets the same quality bar. The opportunity is to apply the maximum useful intelligence at the right price," Friar said.

Efficiency gains

Beyond pricing, OpenAI detailed internal engineering work aimed at lowering the cost of serving models. GPT-5.6 Sol helped optimise the production software used to run its services, cutting end-to-end serving costs by 20 per cent.

OpenAI added that the same model also improved speculative decoding, lifting token-generation efficiency by more than 15 per cent. It said these changes are part of a wider effort to make each unit of compute more productive rather than relying only on building more data centre capacity.

It also pointed to gains from changes around the model rather than in the model itself. According to OpenAI, improvements in retained reasoning and context management raised GPT-5.6 Sol's score on the public ARC-AGI-3 task set from 13.3 per cent to 38.3 per cent while using six times fewer output tokens.

OpenAI said the result shows how software routing, context handling and product design can materially change performance and cost. In its view, those system-level adjustments compound over time because more capable models can help engineers find further operational savings.

Scale of use

OpenAI used the update to underline the scale of its current user base and commercial reach. Its models now reach more than one billion active users and more than two million businesses.

It also said engagement rises after initial adoption. Six months after signing up, users send roughly 50 per cent more messages each day and use ChatGPT for about twice as many kinds of work, according to the figures disclosed.

In workplace settings, adoption often starts with one team or workflow before spreading across functions as quality improves and economics become more attractive, OpenAI said. It also said agentic work through Codex now accounts for 99.8 per cent of weekly output tokens across OpenAI, with Finance among the internal teams using such tools as a primary part of their work.

Investment discipline

OpenAI said AI infrastructure requires long-term planning because data centre and compute investments must be made years before they are fully needed, while customer demand and product development move much faster. That mismatch, it said, makes commercial discipline important.

It bases investment decisions on user growth, workload growth, enterprise commitments, API consumption, utilisation, revenue, and progress in model performance and efficiency. Technical and commercial milestones help determine when projects move forward, it added.

OpenAI also said it does not need to own every asset in its infrastructure chain, arguing that ownership, partnerships and purchasing can all play a role depending on economics and customer need. What matters is coordinating learning across infrastructure, models, platforms and products.

Friar said that principle is central to OpenAI's broader view of growth and cost control.

"The objective is not to build the most infrastructure. It is to deploy the right capacity, at the right time, against credible demand. For me, the key questions are straightforward: How quickly does new capacity become productive? How efficiently is it used? What customer demand does it support? How rapidly can technical progress lower the cost of delivering useful intelligence? Those questions connect long-term ambition to operating discipline," Friar said.