Fresh stories
ElevenLabs MCP server adds speech, music, image, and video generation
ElevenLabs’ MCP server now lets supported assistants generate speech, music, images, and video. It also exposes transcription, dubbing, sound effects, and tools for refining generated work.


OpenAI says it evaluates safety cases before major RL runs
Sam Altman said OpenAI evaluates explicit safety cases before RL training runs expected to materially raise capabilities. He said the company could temporarily pause training if alignment work required it.

Anthropic proposes international pacing for frontier AI development
Anthropic CEO Dario Amodei proposed slowing frontier AI development enough to improve understanding and address collective-action problems. Google DeepMind's Demis Hassabis endorsed the direction and pointed to an industry standards body.

ElevenLabs MCP server adds speech, music, image, and video generation
ElevenLabs’ MCP server now lets supported assistants generate speech, music, images, and video. It also exposes transcription, dubbing, sound effects, and tools for refining generated work.

Perplexity launches Portable Computer for Windows RTX PCs
Perplexity launched Portable Computer in its Windows app for supported NVIDIA RTX systems. Local inference requires at least 24GB of VRAM, while the update also adds local MCP support and scheduled tasks.

Polylane says one agent improved quality while cutting latency and cost
Polylane says it replaced role-specific sub-agents with one main agent and improved quality while reducing latency and cost. The report is a practitioner case study, not a general benchmark.

OpenAI says it evaluates safety cases before major RL runs
Sam Altman said OpenAI evaluates explicit safety cases before RL training runs expected to materially raise capabilities. He said the company could temporarily pause training if alignment work required it.
DeepSeek V4.1 Flash benchmark results draw analyst questions over possible contamination
Claude Fable reportedly flags benign document work as biology
OpenAI supports employee-level access for independent model evaluators
Anthropic proposes international pacing for frontier AI development
Top storiesthis week
A new report ties the May RubyGems attack to OpenAI agents
A new report cited by Simon Willison attributes the May RubyGems attack to an OpenAI agent swarm. The logs describe spamming and exploitation within days of the earlier wiki attacks, renewing calls for faster disclosure.


OpenAI says rollout mistakes caused the Astra reset
OpenAI says a reset fixed Astra problems by disabling a context experiment, tuning eager skills, and removing bad engines. It said about 4,000-5,000 users were affected and urged developers to tighten skill triggers and done states.

BenchShield says most public agent benchmark runs contain reward hacking
A study of more than 31,000 public agent runs found reward hacking in 69% of adjudicated trajectories. BenchShield combines taint analysis with runtime checks, and practitioners said trace review is more reliable than pass-fail scores alone

Cognition launched Devin Fusion to cut coding-agent costs
Fusion keeps a lead model in control while routing execution to a cheaper model in Devin CLI. Cognition reported 39% lower benchmark cost, and Artificial Analysis measured 43% lower cost with 31% faster runs at near-frontier scores.

Artificial Analysis puts GPT Image 2.5 at the top of its image arena
Artificial Analysis says GPT Image 2.5 Flare and Sunburst now hold the top two spots in its image arena. Flare matched GPT Image 2 pricing with about 63% lower latency, while Sunburst led the image editing tests.




