Fresh stories
Cognition releases SWE-2 with lower FrontierCode costs
Cognition says its SWE-2 coding model scored 50.0% on FrontierCode 1.1 Main, matching Fable 5.1 at 64% lower cost. The Kimi K3 post-trained model adds selectable effort levels in Devin Desktop and CLI.

OpenAI launches ChatGPT for Financial Services with financial data sources
OpenAI launched a tailored ChatGPT Work experience for financial institutions that combines GPT-6 Astra with data from Daloopa, PitchBook, and LSEG News. The company says it supports cited analysis and editable research, models, and client materials.


DeepSeek V4.1 Flash tops independent open-weight evaluations
DeepSeek V4.1 Flash leads Vals and Artificial Analysis open-weight comparisons, according to the evaluators. Its encoder-decoder design shares compressed KV state across decoder layers to reduce serving costs.

Cognition releases SWE-2 with lower FrontierCode costs
Cognition says its SWE-2 coding model scored 50.0% on FrontierCode 1.1 Main, matching Fable 5.1 at 64% lower cost. The Kimi K3 post-trained model adds selectable effort levels in Devin Desktop and CLI.

OpenAI launches GPT-Live-1 for full-duplex voice agents
GPT-Live-1 combines listening and speech in one real-time model. Developers can delegate reasoning and tool calls to a backend model while it continues speaking; OpenAI lists pricing at $0.05 per minute.

OpenAI opens a managed Codex runtime in the Agents API
OpenAI’s public-beta Agents API exposes the managed runtime behind Codex. It runs agent loops, tool calls, long-lived sessions, and context management on OpenAI infrastructure, with VPC and bring-your-own sandbox options.
OpenAI launches ChatGPT for Financial Services with financial data sources
OpenAI launched a tailored ChatGPT Work experience for financial institutions that combines GPT-6 Astra with data from Daloopa, PitchBook, and LSEG News. The company says it supports cited analysis and editable research, models, and client materials.
Sakana launches Fugu Max to route requests across specialist models
Sakana says Fugu Max dynamically routes requests across open-weight and specialist models at two to six times lower cost than elite models. In the same release, the company says Fugu Ultra v2 led five of eight hard evaluation suites.
Cursor launches Projects to coordinate persistent coding agents
Cursor’s beta Projects feature keeps work in a persistent thread where a coordinator agent manages subagents and shared artifacts. Projects can schedule work, monitor pull requests, and preserve context across tasks.
Top storiesthis week
Anthropic asks METR to investigate four Claude cyber incidents
Anthropic disclosed four cases in which Claude accessed real systems during misconfigured third-party cyber evaluations. METR will independently investigate the incidents and Anthropic's mitigations.


DeepSeek releases 552B-parameter V4.1 Flash multimodal model
DeepSeek released V4.1 Flash, a 552B-parameter multimodal MoE model with 8B input and 16B output active parameters. It supports up to 1 million tokens of context, while an independent BridgeBench run used 23.5 million tokens on one task.

Goodfire publishes probe-based agent-monitoring guide with a reported 99% catch rate in a GLM 5.3 test
Goodfire describes probes that inspect model internals for prohibited intent, reward hacking, and risky tool calls. Goodfire reports that a probe for a GLM 5.3 coding agent caught 99% of prohibited actions in its test.

Google adds prompt-built mini-apps to Sheets
Google added voice actions, prompt-built mini-apps in Sheets, and Google Pics to Workspace. Gemini Spark can also handle browser errands and tasks across Chrome and Google Photos.

OpenAI says its model solved Navier–Stokes in an 88-hour run
OpenAI says an unreleased model found a solution to the Navier–Stokes Millennium Prize problem in an 88-hour run. The company says roughly 10,000 agents contributed and the result reached Lean formalization.






