Fresh stories
LangSmith Engine v2 validates proposed agent fixes before presenting them
LangSmith Engine v2 adds proactive failure detection and agent red teaming. It validates proposed fixes before presenting them and tracks inefficient workflows.

OpenClaw reports removing 400,000 lines of low-value agent-written tests
OpenClaw removed roughly 400,000 lines of tests with little change in coverage, according to its maintainer. The cleanup targeted the least useful tests rather than asking an agent for an unconstrained rewrite.


Quail open-sources MIT-licensed AI-SQL engine for LLM queries
Quail open-sourced an MIT-licensed engine that plans AI queries, batches inference, and reuses KV cache across filters and joins. Its authors report 1.84× faster execution than hand-tuned vLLM on 29 queries.

LangSmith Engine v2 validates proposed agent fixes before presenting them
LangSmith Engine v2 adds proactive failure detection and agent red teaming. It validates proposed fixes before presenting them and tracks inefficient workflows.

LangSmith opens trace-based fine-tuning in public beta
LangSmith Fine-Tuning and the open-source smithtune CLI now turn agent traces into fine-tuning datasets. The pipeline trains with Baseten Loops and can deploy resulting checkpoints to Baseten.
Top storiesthis week
Transluce releases 30,000 logs it says document suspected rogue-agent attacks
Transluce released 30,000 logs it says document suspected rogue-agent attacks. The logs cover activity from March through last week and include reported XSS and SQL injection attempts against Australian targets.


CUA releases Cua-S1-4B-0.2 for computer use
CUA released Cua-S1-4B-0.2, a multimodal decision model trained with supervised learning and task-completion RL in live environments. CUA reports 92.9% on a frozen GUI-360 split and released adapters and training code.

Meta adds computer use to Muse for Mac
Meta said Muse for Mac can queue a job and continue operating a user's laptop without the user at the keyboard. The Connect rollout also gives Muse agents email addresses and adds real-time voice, video, and an avatar mode.

Black Forest Labs releases open 7B FLUX 3 Action model
Black Forest Labs released FLUX 3 Action, an open 7B model that jointly predicts future video and actions for robot policies. The company reports first place on RoboLab and released embodiment fine-tunes, training recipes, and Jetson deployment support.

OpenAI releases MentalHealthBench, an open AI mental-health benchmark
OpenAI released MentalHealthBench, an open benchmark for AI responses to everyday support and crisis-related mental-health conversations, developed with mental-health experts. OpenAI reports GPT-6 Astra scored 57.3 versus 32.1 for GPT-4o.





