Fresh stories

DeepSeek V4 Flash benchmarks at 61.4% on ARC-AGI-2 for $0.04 per task
ARC Prize verified DeepSeek V4 Flash at 61.4% on ARC-AGI-2 for $0.04 per task. Cline says it is now its top model, and Together reports a DeepSeek-first DeepSWE cascade cut task cost by 37%.
LangChain opens Managed Deep Agents public beta with sandboxes and LangSmith deploys
LangChain opened Managed Deep Agents in public beta for scaffolding agents with channels, sandboxes, memory, and identity. The agents deploy on LangSmith-managed infrastructure with evals and lifecycle controls.

OpenAI updates GPT-Live with continuous audio and one-round-trip WebRTC startup
OpenAI says GPT-Live can listen while speaking using continuous audio, async reasoning and tool use, one-round-trip WebRTC startup, and async context compaction. Staff said the rebuilt stack removes a separate turn detector.


DeepSeek V4 Flash benchmarks at 61.4% on ARC-AGI-2 for $0.04 per task
ARC Prize verified DeepSeek V4 Flash at 61.4% on ARC-AGI-2 for $0.04 per task. Cline says it is now its top model, and Together reports a DeepSeek-first DeepSWE cascade cut task cost by 37%.

Claude Code makes auto permissions default for Pro, Max, and Team on August 14
Anthropic says Claude Code auto mode becomes the default for Pro, Max, and Team users on August 14. Its tool-call classifier caught 89% of dangerous commands in a 1,053-tester study, versus 14% for manual approval.

OpenAI says Astra crossed Critical cyber-risk threshold
OpenAI says internal evaluations put Astra in the Critical tier of its Preparedness Framework. Reports say broad release is slowing while OpenAI adds isolated tests, tool restrictions, monitoring, and sandboxing.

LangChain opens Managed Deep Agents public beta with sandboxes and LangSmith deploys
LangChain opened Managed Deep Agents in public beta for scaffolding agents with channels, sandboxes, memory, and identity. The agents deploy on LangSmith-managed infrastructure with evals and lifecycle controls.
Databricks reports coding-agent token spend is rising exponentially
Magnitude launches open-source offline coding agent for local models
Seedance 2.5 ships in ComfyUI with 30-second runs and timeline shot control
OpenAI updates GPT-Live with continuous audio and one-round-trip WebRTC startup
Top storiesthis week
DeepSeek V4 Flash adds Baseten and Together AI serving with 1M-token context
Baseten and Together AI added DeepSeek V4 Flash with a 1M-token context window, reasoning-effort controls, and DSpark decoding. ValsAI ranked it the cheapest model above 60 on its index, and Nous promoted a short 90% discount.


Cursor users report hard-to-audit coding-agent runs and hidden routing
Cursor users say AI IDE agent runs are hard to audit and hard to constrain. Reddit threads cite hidden model routing, cache charges, destructive SQL migrations, rules folders, runbooks, and context-trimming pipelines.

Hermes Agent ships Herald with Ironproxy secrets lockdown
Teknium shipped the Herald release of Hermes Agent as a desktop-agent and runtime update. The release adds voice chats, desktop plugins, Agent2Agent/webhooks, productivity skills, grounded research, Ironproxy secrets lockdown, and trace-driven token-efficiency work from 250,000 conversations.

MiniMax H3 releases Hugging Face weights and fal video endpoints
MiniMax H3 now has Hugging Face weights, fal endpoints, AI Toolkit LoRA support, and reported single-RTX-5090 local runs. MiniMax also said deployment in the US, EU, UK, and South Korea is available through formal authorization.

Cline raises free DeepSeek Flash quota 3x for coding agents
Cline tripled its free quota while Nous discounted Flash 0731 by 90% for a week and OpenHands offered free cloud use. OpenCode reported 8T Flash tokens on Aug. 1.




