Fresh stories

DataSpace finds harnesses shift data-task accuracy by 15 points
Across 410 cross-source data tasks, DataSpace found fixed-model accuracy ranged from 30.98% to 46.34% across harnesses. Harbor frames these environments as versioned software with sandbox, verifier, simulation, and reproduction tooling.
ASI-Bench finds full procedures lift research-agent scores to 50.91
Across 60 research projects, ASI-Bench found full procedures averaged 50.91, versus 29.10 for prompts naming only a method. Other evaluations similarly measure whether procedural skills improve execution rather than merely adding more instructions.


OpenAI fixes Codex long-session usage accounting
OpenAI says it fixed inefficient usage accounting in long Codex sessions and reset affected accounts. Some users report that business accounts or active sessions did not receive the reset.

OpenRouter says Ox Alpha reaches 8 trillion daily tokens
OpenRouter says its free coding and sustained-agent model reached 8 trillion daily tokens within five days of launch. OpenCode separately reported processing 26 trillion Ox Alpha tokens over four days.

DataSpace finds harnesses shift data-task accuracy by 15 points
Across 410 cross-source data tasks, DataSpace found fixed-model accuracy ranged from 30.98% to 46.34% across harnesses. Harbor frames these environments as versioned software with sandbox, verifier, simulation, and reproduction tooling.

NVIDIA puts Groq 3 LPX into Vera Rubin production
Groq 3 LPX adds dedicated token generation to NVIDIA Vera Rubin systems, with Groq and Nebius among planned deployers. Artificial Analysis measured about 3,400 output tokens per second on Gemma 4 31B.

OpenAI cuts GPT-5.6 Sol API prices by up to 33%
GPT-5.6 Sol now costs $4 per million input tokens and $10 per million output tokens. Benchmark comparisons place Sol at 72.7% on DeepSWE for $6.47 per task, while OpenAI and AWS report lower successful-task costs for Terra in Kiro.
ASI-Bench finds full procedures lift research-agent scores to 50.91
WAN 3.0 launches on API platforms with 30-second video output
Practitioners propose a standard harness for agent benchmarks
OpenAI fixes Codex long-session usage accounting
Top storiesthis week
Anthropic says serving test remapped Claude Code effort settings
Anthropic says a test serving configuration mapped Claude Code’s numeric effort settings differently, allowing “high” to display as 10. The company says evaluations found no regression and the underlying models were unchanged.


Tests link Ox Alpha to Zhipu GLM API routes and error codes
Researchers say malformed Ox Alpha requests exposed Zhipu-specific routes, error codes, and an internal class name. Independent vision comparisons also argue against speculation that the stealth model is Gemini.

Independent DeepSWE retest puts Ox Alpha at about 63%
A practitioner’s larger DeepSWE run reported roughly 63% for Ox Alpha, revising an earlier result near 80% from a 10-task subset. Testers report capable long-task work but note dead code and incomplete fixes.

Qwen 3.8 27B reaches 3,200 TPM at 262K context on two RTX 3090s
Community tests report Qwen 3.8 27B handling coding, OCR, and long-context workloads locally. One vLLM setup reached 3,200 tokens per minute at 262K context on two RTX 3090s without NVLink.

OpenAI grants Codex customers a banked usage reset
OpenAI gave paid ChatGPT Work and Codex users a banked usage reset and said Codex has reached 20 million active users. The company is investigating reports that lower cache-hit rates are causing usage limits to drain faster.



