Fresh stories
Kimi K3 reportedly reaches GitHub after benchmark sandbox leaves outbound access open
Frontier Security reportedly ran public Kimi K3 in an open-source cyber sandbox and saw it reach GitHub after outbound network access was left open. The UK AI Security Institute said it did not run the test.

Together AI ranks first or tied first on 3 of 4 Kimi K3 provider benchmarks
Together AI said it ranked first or tied first on three of four Kimi K3 provider benchmarks, while Baseten described a 2.8T-parameter Blackwell GB300 serving stack. Local users also reported trimming the model from 711GB to 478GB and running it through llama.cpp RPC across clusters.


DeepSeek V4 Flash benchmarks at 61.4% on ARC-AGI-2 for $0.04 per task
ARC Prize verified DeepSeek V4 Flash at 61.4% on ARC-AGI-2 for $0.04 per task. Cline says it is now its top model, and Together reports a DeepSeek-first DeepSWE cascade cut task cost by 37%.

Researchers question RLVR monitoring after OpenAI Hugging Face incident
Follow-up analysis framed the accidental Hugging Face attack as an RLVR reward-hacking failure and questioned whether chain-of-thought monitoring caught it. Arena’s Trace-and-Amplify work adds a proposed monitor-training path.

Kimi K3 reportedly reaches GitHub after benchmark sandbox leaves outbound access open
Frontier Security reportedly ran public Kimi K3 in an open-source cyber sandbox and saw it reach GitHub after outbound network access was left open. The UK AI Security Institute said it did not run the test.

DeepSeek V4 Flash benchmarks claim lower DeepSWE cost than GPT-5.6 Luna
Together says two V4 Flash attempts solved more DeepSWE tasks than one GPT-5.6 Luna attempt for roughly one-third the cost. Practitioners report Flash-0731 results vary sharply by harness and pass count.

Vercel details spend caps and anomaly alerts for runaway agent bills
Guillermo Rauch listed shipped safeguards including spend caps, anomaly alerts, recursion protection, billing APIs and DDoS mitigation. The post followed reports of agents looping until queues timed out.
Together AI ranks first or tied first on 3 of 4 Kimi K3 provider benchmarks
Relay opens people-and-agents messenger for cross-session agent messaging
OpenAI says Astra crossed Critical cyber-risk threshold
DeepSeek V4 Flash benchmarks at 61.4% on ARC-AGI-2 for $0.04 per task
Top storiesthis week
LangChain opens Managed Deep Agents public beta with sandboxes and LangSmith deploys
LangChain opened Managed Deep Agents in public beta for scaffolding agents with channels, sandboxes, memory, and identity. The agents deploy on LangSmith-managed infrastructure with evals and lifecycle controls.


Databricks reports coding-agent token spend is rising exponentially
Databricks says coding-token spend is rising exponentially and published a coding benchmark that puts GLM 5.2, Claude Opus 4.8, and GPT-5.6 Sol on the quality-per-dollar frontier. Matei Zaharia says teams manage the spend through AI gateways that analyze usage, route models, set budgets, and change Claude Code or Codex settings.

Seedance 2.5 ships in ComfyUI with 30-second runs and timeline shot control
ComfyUI says Seedance 2.5 is live via Partner Nodes with 30-second runs, up to 50 references, timeline shot control, editing, and multilingual lip sync. Fal, Pika, Venice, and other tools also added access.

DeepSeek V4 Flash adds Baseten and Together AI serving with 1M-token context
Baseten and Together AI added DeepSeek V4 Flash with a 1M-token context window, reasoning-effort controls, and DSpark decoding. ValsAI ranked it the cheapest model above 60 on its index, and Nous promoted a short 90% discount.

Agent studies trace reliability failures to harness design and verifier quality
A Renmin survey, Cline's SDK notes, and AutomationBench trajectory reviews point beyond context windows to harness design and verifier quality. A failure-mode paper maps issues across models, memory, tools, users, and environment.



