Skip to content
AI Primer

Explore what's new in AI

Where people deep in AI come to stay current.

Filters

Category

Tags

Vercel says GPT-6 Astra leads DeepSecBench in 49 minutes
New

Vercel says GPT-6 Astra leads DeepSecBench in 49 minutes

Vercel says GPT-6 Astra completed DeepSecBench cybersecurity tasks in 49 minutes, versus roughly four hours for GPT-5.6 Sol. It reported a higher score at nearly the same cost per task.

🧠GPT-6 Astra5th September·3 min read
Breaking

DisCo reports reusable skills raise MLE-bench from 31.11% to 72.89%

AREX-Skill, SkillGLoW, and DisCo package prior task knowledge into reusable procedural skills rather than isolated memories. DisCo reports MLE-bench rising from 31.11% to 72.89% with the same model.

DisCo reports reusable skills raise MLE-bench from 31.11% to 72.89%
New
Agent Skills·5th September·3 min read
See all stories →
⌨️Agentic Engineering(5)
🧠Models, Serving & APIs(5)
⚙️Building Agents(11)
🛡️Trust, Evaluation & Reliability(8)
💳Pricing, Limits & Cost(2)
🔎Knowledge, Memory & Retrieval(3)
📈Adoption & Market Strategy(6)

Top storiesthis week

Breaking

ARC Prize reports GPT-6 Astra scores 62.7% or 99.9% on ARC-AGI-3 by harness

ARC Prize reports GPT-6 Astra scored 62.7% on ARC-AGI-3 with its provider-neutral harness, versus 99.9% with OpenAI's provider adapter. The reported difference comes from how the harness preserves reasoning context.

ARC Prize reports GPT-6 Astra scores 62.7% or 99.9% on ARC-AGI-3 by harness
New
GPT-6 Astra·3rd September·5 min read
See all stories →
AI PrimerAI Primer

Your daily guide to AI tools, workflows, and creative inspiration.

© 2026 AI Primer. All rights reserved.