Fresh stories
Claude Security adds Mythos 5 scans for GitHub repositories
Claude Security's public beta now uses Mythos 5 to scan GitHub repositories for enterprise customers. Findings include CWE labels, severity, confidence, and suggested patches that can open in Claude Code on the web.

OpenAI cuts GPT-5.6 Sol API rates to $4 input and $20 output
OpenAI cut GPT-5.6 Sol API input and output rates from $5 and $30 to $4 and $20 per million tokens for three months. Subscription usage remains unchanged.

Marin starts training open 535B-A23B model on 18.75T tokens
Marin has begun an open training run for a 535B-parameter mixture-of-experts model with 23B active parameters. The team plans to train on 18.75T tokens across 11 GB200 NVL72 systems over about three months.


Claude Security adds Mythos 5 scans for GitHub repositories
Claude Security's public beta now uses Mythos 5 to scan GitHub repositories for enterprise customers. Findings include CWE labels, severity, confidence, and suggested patches that can open in Claude Code on the web.

Study finds CLI-first agents cost 5–28x less than MCP agents
A study across seven agents and five models found CLI-first agents were as reliable on mature software tasks as MCP-enabled agents. The CLI setups cost 5–28 times less in the reported experiments.

OpenAI grants Codex customers a banked usage reset
OpenAI gave paid ChatGPT Work and Codex users a banked usage reset and said Codex has reached 20 million active users. The company is investigating reports that lower cache-hit rates are causing usage limits to drain faster.

OpenAI cuts GPT-5.6 Sol API rates to $4 input and $20 output
OpenAI cut GPT-5.6 Sol API input and output rates from $5 and $30 to $4 and $20 per million tokens for three months. Subscription usage remains unchanged.
multiPL-E regex bug corrupts MBPP benchmark language variants
Users report Qwen 3.8 27B agents vary sharply by harness
NVIDIA AVO reportedly reaches 100% on ARC-AGI-3 demo tasks
Marin starts training open 535B-A23B model on 18.75T tokens
Top storiesthis week
Meta publishes Muse Spark 1.2 multimodal evaluation results
Meta published Muse Spark 1.2 evaluations covering tool-based web page and game creation, robotics planning, and audio-visual tasks. A separate result places it first on Design Arena’s video-to-website benchmark.


ARC Prize verifies Gemini 3.7 Flash at 84.6% on ARC-AGI-2
ARC Prize verified Gemini 3.7 Flash at 84.6% on ARC-AGI-2, at a reported $0.25 per task. Artificial Analysis’ AnalystAgent benchmark placed it at 60%, ahead of Claude Opus 5 and GPT-5.5.

Transluce trains 8B–1.1T activation-reading oversight models
Transluce trained 8B to 1.1T parameter oversight models to inspect other models’ activations. It reports results improve with additional training, though its oracle still has room to improve on a reward-hacking evaluation.

Nous launches Hermes Agent with managed remote computers
Nous Research launched Hermes Agent with managed remote computers, provider and model choice, and local or hosted execution. Hermes Cloud idle instances start at 3 cents per day, according to the company.

OpenRouter tests free Ox Alpha with a 1M-token context window
OpenRouter is testing Ox Alpha, a free stealth model with a 1M-token context window. The model accepts text, image, and video inputs and is offered with zero data retention during the test.






