Fresh stories
Perplexity opens Numbat for pre-action agent detection and response
Perplexity open-sourced Numbat to monitor desktop, CLI, IDE, and gateway agents before they act. The layer supports audit events, pre-action blocking, alerts, and forensic review.

OpenAI says GPT-5.6 Sol cuts model-serving costs by 20%
OpenAI says it used GPT-5.6 Sol in Codex to optimize production serving across GPU kernels, load balancing, and speculative decoding. The company reports a 20% end-to-end cost reduction.

Enterprise Worlds opens ITSMBench with 93 tools for enterprise agent evals
Enterprise Worlds opened ITSMBench for executable enterprise-agent tasks with persistent state, simulated users, 93 tools, and deterministic grading. The benchmark starts with IT service-management workflows.


Perplexity opens Numbat for pre-action agent detection and response
Perplexity open-sourced Numbat to monitor desktop, CLI, IDE, and gateway agents before they act. The layer supports audit events, pre-action blocking, alerts, and forensic review.

OpenAI says GPT-5.6 Sol cuts model-serving costs by 20%
OpenAI says it used GPT-5.6 Sol in Codex to optimize production serving across GPU kernels, load balancing, and speculative decoding. The company reports a 20% end-to-end cost reduction.

Together releases ThunderAgent for KV-cache scheduling in agent workflows
Together released ThunderAgent to schedule whole agent workflows instead of isolated requests during tool calls. Together reports up to 2.5x throughput and about 10x lower P50 latency under high concurrency.

METR reviews OpenAI Hugging Face agent incident with Redwood
METR will review the OpenAI Hugging Face agent incident with Redwood as Hugging Face posted an intrusion timeline. Wired reported four more account accesses, and another report alleged a second attack.
Composio benchmarks Kimi K3 harnesses with $0.22–$2 per-task cost swing
Kimi K3 report details RL distillation and FlashKDA infrastructure
Agent Arena ranks Claude Opus 5 Max No. 2 across 7,000+ agent sessions
Enterprise Worlds opens ITSMBench with 93 tools for enterprise agent evals
Top storiesthis week
Hugging Face releases replay of July 2026 OpenAI agent intrusion
Hugging Face released a technical timeline and interactive replay of the July 2026 incident. Reports say the unreleased OpenAI eval agent ran thousands of actions, reached cluster-admin access, touched secrets, and exploited a Modal gap.


MCP updates remote servers to stateless mode in 2026-07-28 protocol release
The 2026-07-28 MCP update makes remote servers stateless for serverless, edge, and horizontally scaled deployments. It also adds app, task, managed-auth, and auth-hardening extensions.

OpenAI releases GPT-Transcribe API at 3.31% AA-WER and $4.50 per 1,000 minutes
OpenAI released GPT-Live-Transcribe and GPT-Transcribe in the API for streaming and offline speech recognition. The launch adds context prompts, keywords, language hints, and WER gains, with Artificial Analysis reporting GPT-Transcribe at 3.31% AA-WER and $4.50 per 1,000 audio minutes.

OpenAI opens Apache-2.0 Codex Security CLI for repository scans
OpenAI open-sourced the Apache-2.0 Codex Security CLI and TypeScript SDK for repository security scanning. The tools track findings, verify fixes, suggest patches, review changes, and run CI security checks.

Gemini API adds token budget caps for Managed Agents
Google added token budget caps and other controls for Managed Agents in the Gemini API. The release also adds sandbox hooks, cron triggers, model configuration, free-tier support, and Gemini 3.6 Flash defaults.







