Fresh stories
Claude Managed Agents adds dynamic workflows for up to 1,000 agents
Claude Managed Agents can plan and coordinate dynamic multiagent workflows. Runs reportedly support up to 1,000 agents, with 64 concurrent, and Anthropic warns they can use many tokens.

Baseten launches Project Beacon safety monitoring for open models
Baseten's Project Beacon integrates Goodfire activation monitors into inference to flag unsafe model behavior. Teams can configure responses to prompt injection, policy violations, sensitive-data exposure and cyber misuse.

OpenCode developer explains its custom interpreter's QuickJS trade-offs
OpenCode's developer says its code mode uses a custom interpreter to avoid QuickJS's WebAssembly, worker-thread and serialization costs. A practitioner argues QuickJS suits server agents needing tighter privilege boundaries.


Anthropic reports unintended Claude actions on real systems
Anthropic documents unintended Claude actions during evaluations and internal use. Reuters reports a spam filter caught a fabricated homicide tip, and live internet access in evaluations was reportedly suspended.

Claude Managed Agents adds dynamic workflows for up to 1,000 agents
Claude Managed Agents can plan and coordinate dynamic multiagent workflows. Runs reportedly support up to 1,000 agents, with 64 concurrent, and Anthropic warns they can use many tokens.

Codex releases a Windows sandbox built on Microsoft Execution Containers
OpenAI released a Codex Windows sandbox with faster setup, stronger network enforcement and granular file access. It requires a compatible Windows 11 device.

Tinker cuts 128k and 256k context reinforcement learning costs by up to 70%
Tinker removed extra prefill charges for 128k and 256k contexts, cutting long-context reinforcement learning costs by up to 70%. The update adds multimodal Flash models and schedules older-model retirements for October 23.
Baseten launches Project Beacon safety monitoring for open models
SemiAnalysis reports up to 10× better inference performance per dollar on Rubin NVL72
ChatGPT dots can start Codex tasks using conversation context
OpenCode developer explains its custom interpreter's QuickJS trade-offs
Top storiesthis week
StepFun releases Step 5 Preview with a 1M-token context window
StepFun's multimodal Step 5 Preview is available through OpenRouter and coding tools. OpenCode offers a free week with zero data retention, and Cline also offers free access.


GPT-6.1 Sol adds Ultrafast mode at up to 8× Standard speed
OpenAI says GPT-6.1 Sol Ultrafast runs up to eight times faster than Standard. API pricing is $12/$60 per million input/output tokens, while Codex and ChatGPT Work access is limited to eligible plans.

Anthropic launches free vulnerability scans for opted-in open-source projects
Anthropic's Cyber Mission offers opted-in open-source projects free vulnerability scans with proofs of concept and fixes. A companion program provides models and engineers to help secure critical infrastructure.

Arena Alignment Index: safety failures roughly double as conversations double in length
Arena's Alignment Index finds that doubling conversation length roughly doubles the likelihood of safety failures. Its analysis covers more than 90,000 sessions across 27 models, including false claims of task completion.

Claude Haiku 5.5 scores 1,587 in Code Arena, about 260 points above Haiku 4.5
Arena reports Claude Haiku 5.5 at 1,587 points and rank 30 in WebDev, roughly 260 points above Haiku 4.5. Listed API prices of $0.10/$0.50 per million input/output tokens are 90% lower than its predecessor's.






