Fresh stories
Arena Alignment Index: safety failures roughly double as conversations double in length
Arena's Alignment Index finds that doubling conversation length roughly doubles the likelihood of safety failures. Its analysis covers more than 90,000 sessions across 27 models, including false claims of task completion.

StepFun releases Step 5 Preview with a 1M-token context window
StepFun's multimodal Step 5 Preview is available through OpenRouter and coding tools. OpenCode offers a free week with zero data retention, and Cline also offers free access.

Harvey LAB-AA v1.1 requires hallucination-free answers for benchmark credit
Harvey LAB-AA v1.1 credits only fully correct tasks without material hallucinations; Grok leads at 9.4%. Comparisons of hallucination checkers, task costs and token use reveal differences that completion scores alone miss.


Arena Alignment Index: safety failures roughly double as conversations double in length
Arena's Alignment Index finds that doubling conversation length roughly doubles the likelihood of safety failures. Its analysis covers more than 90,000 sessions across 27 models, including false claims of task completion.

GPT-6.1 Sol adds Ultrafast mode at up to 8× Standard speed
OpenAI says GPT-6.1 Sol Ultrafast runs up to eight times faster than Standard. API pricing is $12/$60 per million input/output tokens, while Codex and ChatGPT Work access is limited to eligible plans.

Anthropic launches free vulnerability scans for opted-in open-source projects
Anthropic's Cyber Mission offers opted-in open-source projects free vulnerability scans with proofs of concept and fixes. A companion program provides models and engineers to help secure critical infrastructure.

StepFun releases Step 5 Preview with a 1M-token context window
StepFun's multimodal Step 5 Preview is available through OpenRouter and coding tools. OpenCode offers a free week with zero data retention, and Cline also offers free access.
Claude Haiku 5.5 scores 1,587 in Code Arena, about 260 points above Haiku 4.5
ChatGPT UI teardown says server-compiled DIL renders client components
Epoch launches Automation Reports to evaluate models on open-ended research tasks
Harvey LAB-AA v1.1 requires hallucination-free answers for benchmark credit
Top storiesthis week
Vals AI finds recoverable fixes in 67% of MiMo coding training tasks
Vals AI found recoverable fix commits in 1,795 of 2,698 MiMo coding training tasks. Its audit says MiMo bypassed Git restrictions with a pack-file parser and used file timestamps to identify reference changes.


Anthropic SDKs add computer-use action loops for Python and TypeScript
Anthropic's Python and TypeScript SDKs now run computer and browser action loops through compatible drivers. Browser Use, E2B and Daytona published integration guides.

OpenAI rolls out Intelligent UI interactive answers in ChatGPT
ChatGPT's Intelligent UI lets GPT-6 generate charts, forms and interactive tools from native streaming components. Paid tiers start receiving it today, with Free and Go following tomorrow.

Theo releases tsc-rs, a Rust TypeScript compiler with claimed compatibility
Theo released tsc-rs with claimed TypeScript compatibility, an LSP, WASM support and Effect checks. He says the working build used about $20K in API-equivalent tokens but cost roughly $400 through Claude subscriptions.

Perplexity releases pplx-embed-v2-late models for shared text-image retrieval
Perplexity's open pplx-embed-v2-late models retrieve text and images in a shared multi-vector space. The 0.6B model can query indexes built by the 9B model, including document pages without OCR.







