Fresh stories

OpenAI fixes Codex long-session usage accounting
OpenAI says it fixed inefficient usage accounting in long Codex sessions and reset affected accounts. Some users report that business accounts or active sessions did not receive the reset.
Study finds instructions account for 60.5% of coding-agent reading
A study of 557 coding-agent sessions finds instruction files and working notes account for most of what agents read. Related work puts installed skills' standing prompt cost at 50–280 tokens, while Backpass turns past sessions into reviewable AGENTS.M files.


Tests link Ox Alpha to Zhipu GLM API routes and error codes
Researchers say malformed Ox Alpha requests exposed Zhipu-specific routes, error codes, and an internal class name. Independent vision comparisons also argue against speculation that the stealth model is Gemini.

OpenAI fixes Codex long-session usage accounting
OpenAI says it fixed inefficient usage accounting in long Codex sessions and reset affected accounts. Some users report that business accounts or active sessions did not receive the reset.

Practitioners propose a standard harness for agent benchmarks
Practitioners argue that coding-agent results depend heavily on the evaluation harness, including tools, execution control, compaction, and token handling. They propose stable common harnesses rather than vendor-specific setups.

Study finds instructions account for 60.5% of coding-agent reading
A study of 557 coding-agent sessions finds instruction files and working notes account for most of what agents read. Related work puts installed skills' standing prompt cost at 50–280 tokens, while Backpass turns past sessions into reviewable AGENTS.M files.

Together reports GLM-5.3 solves 87.6% of DeepSWE work for about $16
Together reports GLM-5.3 solved 87.6% of DeepSWE after four attempts for about $16, versus Fable 5 at 69.7% for $21.63. It estimates equal $100 budgets yield roughly 17 solved tasks for GLM-5.3 and three for Fable 5.
Anthropic says serving test remapped Claude Code effort settings
Tests link Ox Alpha to Zhipu GLM API routes and error codes
Independent DeepSWE retest puts Ox Alpha at about 63%
Qwen 3.8 27B reaches 3,200 TPM at 262K context on two RTX 3090s
Top storiesthis week
Study finds CLI-first agents cost 5–28x less than MCP agents
A study across seven agents and five models found CLI-first agents were as reliable on mature software tasks as MCP-enabled agents. The CLI setups cost 5–28 times less in the reported experiments.


OpenAI grants Codex customers a banked usage reset
OpenAI gave paid ChatGPT Work and Codex users a banked usage reset and said Codex has reached 20 million active users. The company is investigating reports that lower cache-hit rates are causing usage limits to drain faster.

OpenAI cuts GPT-5.6 Sol API rates to $4 input and $20 output
OpenAI cut GPT-5.6 Sol API input and output rates from $5 and $30 to $4 and $20 per million tokens for three months. Subscription usage remains unchanged.

Marin starts training open 535B-A23B model on 18.75T tokens
Marin has begun an open training run for a 535B-parameter mixture-of-experts model with 23B active parameters. The team plans to train on 18.75T tokens across 11 GB200 NVL72 systems over about three months.

NVIDIA AVO reportedly reaches 100% on ARC-AGI-3 demo tasks
NVIDIA's AVO coding-agent harness reportedly raised Claude Opus 5 from a 30% baseline to 100% on ARC-AGI-3's public demonstration set. ARC-AGI's creator says the result is not a full benchmark score.




