Fresh stories
DeepSeek releases 552B-parameter V4.1 Flash multimodal model
DeepSeek released V4.1 Flash, a 552B-parameter multimodal MoE model with 8B input and 16B output active parameters. It supports up to 1 million tokens of context, while an independent BridgeBench run used 23.5 million tokens on one task.

Google adds prompt-built mini-apps to Sheets
Google added voice actions, prompt-built mini-apps in Sheets, and Google Pics to Workspace. Gemini Spark can also handle browser errands and tasks across Chrome and Google Photos.


DeepSeek releases 552B-parameter V4.1 Flash multimodal model
DeepSeek released V4.1 Flash, a 552B-parameter multimodal MoE model with 8B input and 16B output active parameters. It supports up to 1 million tokens of context, while an independent BridgeBench run used 23.5 million tokens on one task.

Anthropic asks METR to investigate four Claude cyber incidents
Anthropic disclosed four cases in which Claude accessed real systems during misconfigured third-party cyber evaluations. METR will independently investigate the incidents and Anthropic's mitigations.

Goodfire publishes probe-based agent-monitoring guide with a reported 99% catch rate in a GLM 5.3 test
Goodfire describes probes that inspect model internals for prohibited intent, reward hacking, and risky tool calls. Goodfire reports that a probe for a GLM 5.3 coding agent caught 99% of prohibited actions in its test.
Top storiesthis week
OpenAI says its model solved Navier–Stokes in an 88-hour run
OpenAI says an unreleased model found a solution to the Navier–Stokes Millennium Prize problem in an 88-hour run. The company says roughly 10,000 agents contributed and the result reached Lean formalization.


Cohere open-sources fused LLM decode kernel with 1.58x vLLM claim
Cohere released an open-source serving system that fuses the LLM decode step into one GPU kernel launch. On North Mini Code with one H100, it reports up to 1.58x vLLM performance at the tested batch size.

OpenAI releases GPT-Image 2.5 Flare and Sunburst API models
OpenAI introduced GPT-Image 2.5 Flare and Sunburst API models alongside ChatGPT Images 2.5. The release emphasizes lower generation latency, stronger multi-turn edits, and better reference-subject preservation.

OpenAI says de-identified product-use data may have improved its models
OpenAI said it cannot rule out whether de-identified data derived from product usage improved its models. The disclosure does not establish that raw prompts were used for training, and discussion distinguishes direct prompt training from synthetic data derived from use.

Magic says its pretraining recipe matches DeepSeek V4 Pro with 50x less compute
Magic says a new pretraining recipe matched DeepSeek V4 Pro with roughly 50 times less compute. After a 10x scale-up costing about $4 million, the company says it exceeded publicly available base models.





