Fresh stories

Practitioners propose a standard harness for agent benchmarks
Practitioners argue that coding-agent results depend heavily on the evaluation harness, including tools, execution control, compaction, and token handling. They propose stable common harnesses rather than vendor-specific setups.

OpenAI fixes Codex long-session usage accounting
OpenAI says it fixed inefficient usage accounting in long Codex sessions and reset affected accounts. Some users report that business accounts or active sessions did not receive the reset.

Practitioners propose a standard harness for agent benchmarks
Practitioners argue that coding-agent results depend heavily on the evaluation harness, including tools, execution control, compaction, and token handling. They propose stable common harnesses rather than vendor-specific setups.

Study finds instructions account for 60.5% of coding-agent reading
A study of 557 coding-agent sessions finds instruction files and working notes account for most of what agents read. Related work puts installed skills' standing prompt cost at 50–280 tokens, while Backpass turns past sessions into reviewable AGENTS.M files.
Top storiesthis week
Anthropic says serving test remapped Claude Code effort settings
Anthropic says a test serving configuration mapped Claude Code’s numeric effort settings differently, allowing “high” to display as 10. The company says evaluations found no regression and the underlying models were unchanged.


Tests link Ox Alpha to Zhipu GLM API routes and error codes
Researchers say malformed Ox Alpha requests exposed Zhipu-specific routes, error codes, and an internal class name. Independent vision comparisons also argue against speculation that the stealth model is Gemini.

multiPL-E regex bug corrupts MBPP benchmark language variants
An audit found that multiPL-E's MBPP subset replaced every occurrence of "py" rather than the word "python," creating malformed language names. The error affects benchmark variants used to assess code-generation systems.

Claude Security adds Mythos 5 scans for GitHub repositories
Claude Security's public beta now uses Mythos 5 to scan GitHub repositories for enterprise customers. Findings include CWE labels, severity, confidence, and suggested patches that can open in Claude Code on the web.

Study finds CLI-first agents cost 5–28x less than MCP agents
A study across seven agents and five models found CLI-first agents were as reliable on mature software tasks as MCP-enabled agents. The CLI setups cost 5–28 times less in the reported experiments.




