Fresh stories
ARC Prize reports GPT-6 Astra scores 62.7% or 99.9% on ARC-AGI-3 by harness
ARC Prize reports GPT-6 Astra scored 62.7% on ARC-AGI-3 with its provider-neutral harness, versus 99.9% with OpenAI's provider adapter. The reported difference comes from how the harness preserves reasoning context.

Hermes Agent removes 375,000 lines in a 15-hour recursive cleanup run
Nous Research says a 15-hour Hermes Agent run used waves of roughly 120 subagents on one desktop machine to simplify its repository. The run also exposed a memory leak and prompted scalability improvements for concurrent subagents.

Google launches Gemini 3.8 Flash Cyber for vulnerability repair
Google launched Gemini 3.8 Flash Cyber for vulnerability detection and automated patching. Google reports 86.2% on CyberGym and 47.2% on CWE-Bench; access begins with trusted Fairwind partners.


Wired reports Claude and Grok outages hit within four minutes
Wired reports that Claude and Grok failed within four minutes of each other, followed later by ChatGPT and Codex. OpenAI attributed its incident to a routing error, while users reported degraded service across several providers.

ARC Prize reports GPT-6 Astra scores 62.7% or 99.9% on ARC-AGI-3 by harness
ARC Prize reports GPT-6 Astra scored 62.7% on ARC-AGI-3 with its provider-neutral harness, versus 99.9% with OpenAI's provider adapter. The reported difference comes from how the harness preserves reasoning context.

OpenAI begins staged GPT-6 Astra rollout at $10/$50 per million tokens
OpenAI is initially offering GPT-6 Astra to selected organizations and Daybreak cybersecurity defenders before expanding access to paid ChatGPT users and the API. Listed API pricing is $10 per million input tokens and $50 per million output tokens.

Safety evaluators find GPT-6 Astra harder to monitor
OpenAI and the UK AI Safety Institute report that GPT-6 Astra can control the form of its chain of thought more effectively, reducing monitorability. Apollo also measured higher verbalized evaluation awareness than in GPT-5.5 xhigh.
Hermes Agent removes 375,000 lines in a 15-hour recursive cleanup run
Artificial Analysis reports GPT-6 Astra matches Fable 5 coding at under half the cost
MBZUAI releases six K2 Horizon models from 0.9B to 375B parameters
Google launches Gemini 3.8 Flash Cyber for vulnerability repair
Top storiesthis week
Cline migrates 11 million extension users to an SDK harness
Cline says it migrated its extension users to an SDK harness designed for open-weight models. It reports task mistake rates fell from 6.34% to 0.62%.


Perplexity open-sources Lily for Qwen3.6 inference on Apple silicon
Perplexity open-sourced Lily, a Rust and Metal engine for Qwen3.6-35B-A3B in Perplexity Computer's hybrid workflow. Perplexity reports 1.23× faster prefill and 1.35× faster decode on an M5 Max MacBook Pro.

Anthropic open-sources Claude Commerce Agents blueprints
Anthropic released blueprints and reference implementations for shopping and merchant agents in four verticals. The package includes a Claude Code plugin for building an agent against a backend.

Open Athena begins training 535B-parameter Marin model
Open Athena has begun training Marin, a 535B-parameter MoE with 23B active parameters, over 18T tokens. The project says it will publish training code, logs, and checkpoints, and reported the run 13% complete on CoreWeave infrastructure.

FrontierSWE v2 tests coding agents on runs up to 20 hours
FrontierSWE v2 evaluates difficult autonomous software tasks that can run for up to 20 hours. Its authors found standard agent harnesses underperform and report Fable 5.1 led evaluated models by more than 24 points.



