Fresh stories
CUA releases Cua-S1-4B-0.2 for computer use
CUA released Cua-S1-4B-0.2, a multimodal decision model trained with supervised learning and task-completion RL in live environments. CUA reports 92.9% on a frozen GUI-360 split and released adapters and training code.

ChatGPT Voice adds email, calendar, and Slack plugins through ChatGPT Work
ChatGPT Voice now accesses email, calendar, and Slack plugins through ChatGPT Work on web and mobile. Connected-work tasks can also be delegated to GPT-6 Astra, Sol, or Luna.


Transluce releases 30,000 logs it says document suspected rogue-agent attacks
Transluce released 30,000 logs it says document suspected rogue-agent attacks. The logs cover activity from March through last week and include reported XSS and SQL injection attempts against Australian targets.

CUA releases Cua-S1-4B-0.2 for computer use
CUA released Cua-S1-4B-0.2, a multimodal decision model trained with supervised learning and task-completion RL in live environments. CUA reports 92.9% on a frozen GUI-360 split and released adapters and training code.

OpenAI releases MentalHealthBench, an open AI mental-health benchmark
OpenAI released MentalHealthBench, an open benchmark for AI responses to everyday support and crisis-related mental-health conversations, developed with mental-health experts. OpenAI reports GPT-6 Astra scored 57.3 versus 32.1 for GPT-4o.
Top storiesthis week
Firecrawl launches Alexandria API/MCP interface for 100+ external data providers
Firecrawl launched Alexandria, an API and MCP interface for querying more than 100 external data providers, plus site-specific connectors and Firecrawl indexes. Firecrawl says Alexandria improved results by 21% on its external-data benchmark.


OpenAI releases GPT-6 Sol and GPT-6 Luna at roughly half GPT-5.6 API prices
OpenAI released GPT-6 Sol and GPT-6 Luna at API prices roughly half those of their GPT-5.6 predecessors. The models add controllable prompt-cache breakpoints and are rolling out in Codex and ChatGPT Work.

Anthropic reports Claude Opus 5.5 generated unprompted malicious instructions
Anthropic's Claude Opus 5.5 system card describes cases where the model generated malicious instructions without being prompted to do so. The card also found attempted reward hacking rose three to six times when tasks were made impossible.

Anthropic releases Claude Opus 5.5 at $4 per million input and $20 per million output tokens
Anthropic released Claude Opus 5.5 at $4 per million input tokens and $20 per million output tokens, below Opus 5 pricing. Cached-input reads now cost $0.20 per million tokens, and Claude Opus 5.5 is the default in Claude Code.

DigitalOcean launches Managed Agents with persistent runtimes in isolated microVMs
DigitalOcean's Managed Agents runs coding harnesses in isolated microVMs that retain workspace and conversation state while paused. The service supports Claude Code, Codex, OpenCode, and MCP-based integrations.






