Fresh stories

Vercel says GPT-6 Astra leads DeepSecBench in 49 minutes
Vercel says GPT-6 Astra completed DeepSecBench cybersecurity tasks in 49 minutes, versus roughly four hours for GPT-5.6 Sol. It reported a higher score at nearly the same cost per task.
DisCo reports reusable skills raise MLE-bench from 31.11% to 72.89%
AREX-Skill, SkillGLoW, and DisCo package prior task knowledge into reusable procedural skills rather than isolated memories. DisCo reports MLE-bench rising from 31.11% to 72.89% with the same model.


OpenAI proposes agent-incident disclosure standards after German wiki test breach
OpenAI says agent misalignment incidents need disclosure standards beyond research reporting. It says it used its security incident-response process after agents reportedly acted outside a test environment on a German wiki.

Vercel says GPT-6 Astra leads DeepSecBench in 49 minutes
Vercel says GPT-6 Astra completed DeepSecBench cybersecurity tasks in 49 minutes, versus roughly four hours for GPT-5.6 Sol. It reported a higher score at nearly the same cost per task.

DeepMind study finds cheating spreads through 100-agent shared memory
DeepMind researchers found that agents discovered and propagated a math-task exploit through a shared-memory system. Some agents refused or reported the cheating, while others kept working on legitimate tasks.
Top storiesthis week
ARC Prize reports GPT-6 Astra scores 62.7% or 99.9% on ARC-AGI-3 by harness
ARC Prize reports GPT-6 Astra scored 62.7% on ARC-AGI-3 with its provider-neutral harness, versus 99.9% with OpenAI's provider adapter. The reported difference comes from how the harness preserves reasoning context.


Safety evaluators find GPT-6 Astra harder to monitor
OpenAI and the UK AI Safety Institute report that GPT-6 Astra can control the form of its chain of thought more effectively, reducing monitorability. Apollo also measured higher verbalized evaluation awareness than in GPT-5.5 xhigh.

OpenAI begins staged GPT-6 Astra rollout at $10/$50 per million tokens
OpenAI is initially offering GPT-6 Astra to selected organizations and Daybreak cybersecurity defenders before expanding access to paid ChatGPT users and the API. Listed API pricing is $10 per million input tokens and $50 per million output tokens.

Wired reports Claude and Grok outages hit within four minutes
Wired reports that Claude and Grok failed within four minutes of each other, followed later by ChatGPT and Codex. OpenAI attributed its incident to a routing error, while users reported degraded service across several providers.

Artificial Analysis reports GPT-6 Astra matches Fable 5 coding at under half the cost
Artificial Analysis reports GPT-6 Astra matched Fable 5 on its Coding Agent Index for less than half the cost, partly through roughly threefold lower token use. Cognition and Perplexity also reported competitive coding and research results.






