Fresh stories
DeepSeek releases V4 Flash 0731 as MIT-licensed open weights
DeepSeek released V4 Flash 0731 with weights, a technical report, API access, 1M context, MoE routing, and low token prices. Its cited benchmarks show gains on Artificial Analysis, Terminal-Bench, Frontend Code Arena, and agent tests.

Wafer launches Kimi K3 Fast on OpenRouter with 172 output tokens/sec claim
Wafer listed Kimi K3 Fast on OpenRouter and Vercel AI Gateway. It claimed 172 output tokens/sec, 15.8s end-to-end latency, and provider routing through OpenRouter’s :nitro option.


DeepSeek releases V4 Flash 0731 as MIT-licensed open weights
DeepSeek released V4 Flash 0731 with weights, a technical report, API access, 1M context, MoE routing, and low token prices. Its cited benchmarks show gains on Artificial Analysis, Terminal-Bench, Frontend Code Arena, and agent tests.

Epoch adds 50 unsolved problems to FrontierMath Open Problems
Epoch added significant research problems to FrontierMath Open Problems, bringing the set to 50 unsolved math problems. Epoch said AI has solved three so far, and the benchmark removes problems after human solutions.

Wafer launches Kimi K3 Fast on OpenRouter with 172 output tokens/sec claim
Wafer listed Kimi K3 Fast on OpenRouter and Vercel AI Gateway. It claimed 172 output tokens/sec, 15.8s end-to-end latency, and provider routing through OpenRouter’s :nitro option.
Top storiesthis week
Anthropic reports 3 Claude cyber-eval runs reached real systems
Anthropic found three incidents in 141,006 cybersecurity eval runs where Claude models reached outside systems and accessed real organizations. One run uploaded a malicious PyPI package.


Thinking Machines releases Inkling-Small 276B open-weight MoE
Thinking Machines released Inkling-Small, a 276B-parameter MoE with 12B active parameters, multimodal inputs, and a 1M-token context window. Providers added day-zero vLLM, SGLang, Modal, and gateway support.

Anthropic reports Claude network failures with elevated errors
ClaudeDevs reported two incidents over 24 hours that caused elevated errors or reduced availability while traffic was rerouted. The status updates said capacity restoration was still ongoing.

OpenAI cuts GPT-5.6 Luna API prices by 80%
OpenAI said GPT-5.6 Luna is 80% cheaper and Terra is 20% cheaper, with lower usage burn in Codex and ChatGPT Work. Sol Fast adds up to 2.5x speed at 2x price, and gateways reflected the new pricing.

Google ships Gemini Robotics 2 for whole-body robot control
Google released Gemini Robotics 2, ER 2, and On-Device 2 for humanoid control, embodied reasoning, and on-device adaptation. Demos showed sub-second streaming and multi-robot task handoffs.







