Fresh stories

GPT-6.1 Sol scores 86.3% on MathArena BrokenArXiv
GPT-6.1 Sol scored 86.3% on MathArena BrokenArXiv, 78.6% on Braintrust problem solving, and third in Code Arena WebDev. A user also reports that it outperformed its mini-SWE result in Codex.
Reports rank Gemini 4 Argon highly on four engineering benchmarks
Reports place Gemini 4 Argon at 77.9% on DeepSWE, 57.6% on Terminal-Bench, 77.5% on AutomationBench-AA, and 68% on CWE-Bench. The reports also cite lower cost or token use than rivals.


Google rolls out Gemini 4 Argon to cyber defenders
Google is rolling out Gemini 4 Argon through Fairwind to government users, vetted cyber defenders, and trusted testers. The model supports up to 1M output tokens and costs $2/M input and $10/M output.

GPT-6.1 Sol scores 86.3% on MathArena BrokenArXiv
GPT-6.1 Sol scored 86.3% on MathArena BrokenArXiv, 78.6% on Braintrust problem solving, and third in Code Arena WebDev. A user also reports that it outperformed its mini-SWE result in Codex.

Perplexity open-sources pplx-embed-v2-context-9b-preview
Perplexity released pplx-embed-v2-context-9b-preview, which encodes chunks using whole-document context. Perplexity reports leading results on ConTEB and Turbopuffer's context benchmark.
Top storiesthis week
Anthropic reports GLM-5.3 built browser exploits in 50 of 410 attempts
Anthropic reports that GLM-5.3 produced working browser exploits in 50 of 410 controlled attempts. In a separate binary-exploitation test, it achieved control-flow hijacks in 4% of trials.


Developers report Firebase payload crashed iOS apps
Developers traced widespread iOS app crashes to a malformed server-side Firebase payload, which was rolled back. Cached payloads reportedly kept some apps failing afterward.

OpenAI lets partner apps use ChatGPT subscription allowances
Sign in with ChatGPT lets subscribers use included plan allowances in partner tools such as Pi, Warp and Devin. Devin documents quota controls, while Amp says higher-capacity use carries separate charges.

OpenAI releases GPT-6.1 Sol at $2 per million input tokens
GPT-6.1 Sol is available in the API, Codex and ChatGPT Work. It costs $2 per million input tokens and $10 per million output tokens; OpenAI claims near-Astra coding results.

OpenAI launches GPT-6 Astra Ultrafast at up to 8× Standard speed
OpenAI says Ultrafast generates tokens up to eight times faster than Standard in Codex. The tier is available through the API and selected subscriptions at a higher price; GPT-6.1 Sol support is planned.





