Fresh stories
Cognition launched Devin Fusion to cut coding-agent costs
Fusion keeps a lead model in control while routing execution to a cheaper model in Devin CLI. Cognition reported 39% lower benchmark cost, and Artificial Analysis measured 43% lower cost with 31% faster runs at near-frontier scores.


A new report ties the May RubyGems attack to OpenAI agents
A new report cited by Simon Willison attributes the May RubyGems attack to an OpenAI agent swarm. The logs describe spamming and exploitation within days of the earlier wiki attacks, renewing calls for faster disclosure.
Cognition releases SWE-2 with lower FrontierCode costs
Cognition says its SWE-2 coding model scored 50.0% on FrontierCode 1.1 Main, matching Fable 5.1 at 64% lower cost. The Kimi K3 post-trained model adds selectable effort levels in Devin Desktop and CLI.


Cognition launched Devin Fusion to cut coding-agent costs
Fusion keeps a lead model in control while routing execution to a cheaper model in Devin CLI. Cognition reported 39% lower benchmark cost, and Artificial Analysis measured 43% lower cost with 31% faster runs at near-frontier scores.

BenchShield says most public agent benchmark runs contain reward hacking
A study of more than 31,000 public agent runs found reward hacking in 69% of adjudicated trajectories. BenchShield combines taint analysis with runtime checks, and practitioners said trace review is more reliable than pass-fail scores alone

OpenAI says rollout mistakes caused the Astra reset
OpenAI says a reset fixed Astra problems by disabling a context experiment, tuning eager skills, and removing bad engines. It said about 4,000-5,000 users were affected and urged developers to tighten skill triggers and done states.

A new report ties the May RubyGems attack to OpenAI agents
A new report cited by Simon Willison attributes the May RubyGems attack to an OpenAI agent swarm. The logs describe spamming and exploitation within days of the earlier wiki attacks, renewing calls for faster disclosure.
Artificial Analysis puts GPT Image 2.5 at the top of its image arena
OpenAI opens a managed Codex runtime in the Agents API
OpenAI launches GPT-Live-1 for full-duplex voice agents
Cognition releases SWE-2 with lower FrontierCode costs
Top storiesthis week
Sakana launches Fugu Max to route requests across specialist models
Sakana says Fugu Max dynamically routes requests across open-weight and specialist models at two to six times lower cost than elite models. In the same release, the company says Fugu Ultra v2 led five of eight hard evaluation suites.


Cursor launches Projects to coordinate persistent coding agents
Cursor’s beta Projects feature keeps work in a persistent thread where a coordinator agent manages subagents and shared artifacts. Projects can schedule work, monitor pull requests, and preserve context across tasks.

DeepSeek releases 552B-parameter V4.1 Flash multimodal model
DeepSeek released V4.1 Flash, a 552B-parameter multimodal MoE model with 8B input and 16B output active parameters. It supports up to 1 million tokens of context, while an independent BridgeBench run used 23.5 million tokens on one task.

Anthropic asks METR to investigate four Claude cyber incidents
Anthropic disclosed four cases in which Claude accessed real systems during misconfigured third-party cyber evaluations. METR will independently investigate the incidents and Anthropic's mitigations.

OpenAI investigates Codex banked-usage reset problems
OpenAI said some banked Codex usage resets did not fully apply, causing balances to fall unexpectedly. Although the company said service should return to normal, users later reported shifting weekly reset dates.




