OpenAI's Codex 0.144.6 sets 272,000-token context metadata for GPT-5.6 Sol, Terra, and Luna after one Sol user documented a brief 372,000-token setting, renewing a dispute over how Codex accounts for input, output, and compaction headroom.
OpenAI has made 272,000 the stable Codex context value for all three GPT-5.6 models after one Sol user documented a short-lived 372,000-token setting. The change is real in the product's metadata and in that user's runtime records. What it means is less settled: the available evidence does not show whether users lost part of the models' total context capacity or whether Codex corrected an oversized input allowance.

OpenAI’s company-reported hotfix summary says Codex 0.144.6 retains 272,000-token context metadata for GPT-5.6 Sol, Terra and Luna while reverting unrelated catalog changes. Source: OpenAI Codex pull request 34009.
The Codex 0.144.6 release, published July 18, says OpenAI refreshed bundled instructions for GPT-5.6 Sol, Terra, and Luna and “corrected” their context windows to 272,000 tokens. It does not describe the change as a capacity cut or explain the accounting behind the number.
The associated hotfix record makes the engineering scope unusually precise. A broader generated-catalog refresh had introduced changes that were not needed in the stable 0.144 branch. OpenAI reverted those unrelated changes while keeping four fields for each GPT-5.6 model: base instructions, an instruction template, context_window, and max_context_window. That left exactly 12 values changed from the stated baseline.
This establishes that 272,000 is intentional stable metadata. It does not establish that 372,000 was a promised product limit, how long it was available, or why OpenAI considers the lower value correct.
A July 13 account based on posts by OpenAI head of core products Thibault Sottiaux described the context setting as being reduced from 372,000 to 272,000 to curb usage. That is the clearest retained claim about motive, but it comes through a secondary report; neither the release nor the hotfix record says cost or demand caused the correction.
A user-filed Codex issue provides the most detailed before-and-after observation. In Codex CLI 0.144.3 on Linux with a paid ChatGPT-authenticated subscription, the user reported that the server-delivered Sol catalog changed as follows:
| Measure | Earlier catalog | Later catalog | Change |
|---|---|---|---|
Raw context_window | 372,000 | 272,000 | -100,000 (-26.9%) |
| Reported effective runtime window at 95% | 353,400 | 258,400 | -95,000 (-26.9%) |
The user said old and new values briefly overlapped across existing and newly opened threads, consistent with a phased or cached rollout. This is a reproducible account from one environment, not an OpenAI measurement of the entire user base. It also does not show how frequently real coding tasks reach the limit or how task quality changes after compaction.
The provisional interpretation—that Codex simply removed 100,000 tokens from every session—misses an important accounting issue. In a January discussion about earlier Codex models, an OpenAI collaborator explained the 272,000 figure as an input allowance inside a 400,000-token total: 128,000 tokens were reserved for a possible model output, because exceeding the total would terminate the session. The collaborator said raising the input allowance would create too many overflow errors.
That explanation predates GPT-5.6 and is not a direct statement about this rollout. It nevertheless supplies a concrete alternative to the cost-control theory because GPT-5.6 Sol also has a published 128,000-token maximum output. OpenAI's new release still calls 272,000 the model's “context window,” however, while the user report applies another 95% factor and displays 258,400. The retained sources do not reconcile those terms.

OpenAI’s company-reported GPT-5.6 Sol API specifications list a 1,050,000-token context window and 128,000 maximum output tokens. Source: OpenAI GPT-5.6 Sol model documentation.
OpenAI's API documentation for GPT-5.6 Sol lists a 1,050,000-token context window and a 128,000-token maximum output. The API therefore exposes a much larger advertised envelope than the default Codex metadata, but it is a different access and billing surface.
Standard API prices are $5 per million input tokens and $30 per million output tokens. When a prompt exceeds 272,000 input tokens, OpenAI says the full request is billed at twice the input rate and 1.5 times the output rate—$10 and $45 per million tokens, respectively.
The shared 272,000 threshold is notable, but it does not prove that API pricing caused the Codex change. The earlier Codex explanation ties the same number to output reservation and overflow prevention. The defensible conclusion is narrower: API customers can pay to submit longer inputs, while the retained Codex documentation offers no comparable opt-in route for ChatGPT-authenticated use.
The context correction landed during a wider effort to make GPT-5.6 usage more available. The July 13 account said Sottiaux temporarily removed a five-hour restriction for Plus, Business, and Pro subscribers and attributed to him an inference optimization expected to yield about 10% more Sol usage within subscription allowances. It also said OpenAI briefly tested lower reasoning effort in multi-agent work before restoring the prior setting.
A separate account of Sottiaux's posts said OpenAI gave one banked weekly reset to 500,000 ChatGPT Work and Codex accounts and made the reset available on web and mobile, not only desktop. Sottiaux also said fewer than 10% of people who tried a reset during a two-hour period did not receive it. Because OpenAI could not precisely identify the affected accounts, it granted another reset to everyone who pressed the button during that interval.
Those measures concern how often subscribers can use the models. They do not change how much input Codex retains before compaction, and the reset failure is evidence of a rollout problem rather than proof of a broader infrastructure shortage. Treating quota relief as compensation for the context change would go beyond the sources.
The immediate software decision is complete: stable Codex 0.144.6 uses 272,000 for both context fields across Sol, Terra, and Luna. The unresolved product decision is whether OpenAI will tell users exactly what that value represents and whether longer inputs will become an option inside Codex.
Three disclosures would resolve most of the dispute: whether 372,000 was erroneous or experimental; whether the 272,000 value for GPT-5.6 is an input allowance that reserves its 128,000-token maximum output; and how the additional 95% effective-window factor relates to compaction headroom.
Performance evidence matters more than the catalog label. Before-and-after measurements of compaction frequency, task completion, error rates, and total token use on long coding workflows would show whether the 95,000-token reduction in the reported effective window materially harms work or mainly restores a safety margin. Until OpenAI supplies that explanation and evidence, both “context cut” and “metadata correction” describe only part of what users experienced.
Get concise AI news and useful context from the Magica team.
Read the newsletterKimi says a 48-hour demand surge pushed its GPU capacity close to the limit, forcing a pause in new subscriptions as a separate Kimi Code plan prepares to change how coding access is bundled and rationed.
Ant Digital Technologies has expanded Agentar with 200 preconfigured job-role templates and a multi-agent management pitch, but it has not disclosed pricing, customer use or performance data—and governance is becoming a market-wide requirement rather than a distinctive feature.
Apple is reportedly testing an opt-in tool that transcribes and summarizes Genius Bar appointments. Its current safeguards are clear, but its accuracy, retention rules and uses after the pilot are not.
Moonshot AI is seeking investor approval for a Hong Kong IPO within six months while an unfinished private round could value it above $30 billion. The pitch pairs rapid reported recurring-revenue growth with Kimi K3, but neither audited financials nor enough independent model and deployment evidence is public yet.
Andy Serkis says machine learning has a narrow role in The Hunt for Gollum’s de-aging work. That boundary remains a production claim, because the actors, shots, tools, data, labor effects and likeness terms have not been disclosed.
Nebius Group revenue rose 684% to $399 million in the first quarter of 2026 while purchases of property, equipment and intangible assets reached $2.47 billion. Microsoft and Meta reduce demand risk, but options and an unsold-capacity backstop leave delivery, financing and unit economics unresolved.
Chinese-developed models have overtaken U.S. rivals in token volume on OpenRouter, where low prices and token-heavy agent workloads favor DeepSeek V4. The crossover covers a small, platform-specific slice of AI use and does not establish leadership in revenue, enterprise demand or infrastructure control.
OpenAI said it had identified the cause of elevated ChatGPT errors and was applying mitigations, but the retained status update did not disclose the technical failure, measure the impact or confirm full recovery.
Alibaba has opened hosted access to Qwen3.8-Max-Preview and says the 2.4-trillion-parameter model will be released with downloadable weights, but the license, architecture, benchmark evidence, release date and model-level credit economics remain undisclosed.
Anthropic has put Bun’s Rust port into Claude Code ahead of Bun 1.4’s general release, creating a real but tightly controlled proving ground for an AI-led migration whose total cost and broader reliability remain unsettled.
Zhipu reportedly reached $1 billion in annual recurring revenue in July, roughly four times a March estimate, but the unconfirmed run rate is not annual sales and still sits far ahead of recognized cloud revenue while margins remain thin.
An account of PNC transaction data puts household-paid generative AI near 2%, while a separate user survey finds much broader paid access when employer-funded plans count. The gap shows why card charges alone cannot settle whether consumer AI is becoming a mass subscription business.
Kimi K3 reduces attention traffic, but Moonshot recommends deploying it across at least 64 accelerators. SemiAnalysis says expert routing will more than erase the bandwidth savings; until the promised weights are deployed independently, that remains a hardware thesis rather than a measured result.
Morgan Stanley raised its Micron fiscal-2027 gross-margin estimate to 89.3%, but Micron’s results show the forecast depends chiefly on exceptional memory pricing, customer contracts and delayed supply rather than a disclosed HBM4 margin advantage.
New Mexico’s land commissioner refused to reconsider state-land crossings for a pipeline serving Project Jupiter, preserving a fuel-supply obstacle for the planned Oracle data center. But an analyst’s 2029 forecast predates that decision and remains at odds with Oracle’s first-half-2027 delivery statement.
A reported CIA mission examined whether an influential Emirati sheikh could be trusted with sensitive U.S. technology. The public export rule that followed gives G42 and Core42 a narrow, temporary exception, but it does not connect the intelligence operation to that decision or show how compliance will be tested.
Tracebit says a guardrail-triggering string in one decoy AWS secret sharply reduced five AI agents' success in a 152-run cyber range, but the company-run test did not cover uncensored models, adaptive attackers or production deployments.
SenseTime’s U1 Pro preview combines a claimed native 8K ceiling with a multi-step image-creation loop, but its August API will need to disclose dimensions, latency, pricing and repeatable results before buyers can compare the cost of a usable asset.
Alibaba Cloud has begun invite-only testing of a 64-card Zhenwu M890 supernode instance, extending its in-house chip and infrastructure stack to outside users while leaving the service's price, benchmark methodology and customer economics undisclosed.
Nvidia's 616.00 driver and CUDA 13.4 preview let developers begin native Windows Arm64 work for RTX Spark, but the release is an ecosystem-building step—not evidence of final performance, compatibility or pricing.