Fresh stories

Reports say OpenClaw exposed missing auth on gym booking cancellation API
Reports say OpenClaw’s gym demo found missing authorization checks on cancellation endpoints in an Australian booking API. The agent allegedly canceled another user’s reservation, leaving responsibility unclear between app auth and harness controls.
Google opens Gemini Omni Flash to developers through Gemini API and Vertex AI
Google opened Gemini Omni Flash to developers for video creation and editing from text, image, video, or audio references. The rollout spans Gemini, Flow, AI Studio, the Gemini API, and Vertex AI, with supporting posts showing editing demos rather than independent performance tests.


Composio and Ante benchmark coding agent harnesses with 47%–67% success range
Composio and Ante tests reported that the same models behaved very differently by harness. DeepSeek V4 Flash ranged from 47% to 67% task success and $0.019 to $0.104 per task across harnesses.

Reports say OpenClaw exposed missing auth on gym booking cancellation API
Reports say OpenClaw’s gym demo found missing authorization checks on cancellation endpoints in an Australian booking API. The agent allegedly canceled another user’s reservation, leaving responsibility unclear between app auth and harness controls.

Meta releases Muse Glimmer 30B as an Apache 2.0 open-weight local agent model
Meta released Muse Glimmer, a 30B Apache 2.0 dense model for local agent workflows. Reports cite 4-bit builds under 20GB, vision input, function calling, 131K context, and day-0 support in Hugging Face, vLLM, SGLang, Ollama, and MLX.

OpenAI releases GPT-5.6-Cyber for approved Daybreak Blue and Red teams
OpenAI released GPT-5.6-Cyber through expanded Daybreak Blue and Red tiers for authorized vulnerability research, exploit validation, and testing. OpenAI frames the release as less-restricted access for defenders with safeguards.
Google opens Gemini Omni Flash to developers through Gemini API and Vertex AI
Google opened Gemini Omni Flash to developers for video creation and editing from text, image, video, or audio references. The rollout spans Gemini, Flow, AI Studio, the Gemini API, and Vertex AI, with supporting posts showing editing demos rather than independent performance tests.
Anthropic claims unreleased Claude improves zeta-zero lower bound to about 67.2%
Anthropic says an unreleased Claude did not solve the Riemann hypothesis but improved a related zeta-zero lower bound from 41.6% to about 67.2%. Posts describe subagents, expert prompting, and Lean formalization.
OpenRouter updates Auto Router with 30 task types and 7-day spend-based routing
OpenRouter upgraded Auto Router to classify prompts into about 30 task types, then route by anonymized 7-day spend share and cost tier. OpenRouter says the max tier beat the old router across five benchmark domains.
Top storiesthis week
OpenClaw reports missing auth checks in gym waitlist API
Simon Willison quoted OpenClaw saying a gym API allowed cancelling other users’ reservations and moving a waitlisted user up one spot. Replies treated it as both an agent safety failure and a basic authorization bug.


OpenAI faces Artifactory monitoring questions as postmortem is promised
Security researchers disputed how OpenAI detected and investigated the Artifactory incident. Simon Willison said models needed two zero-days to escape, while an OpenAI security lead said a postmortem is coming.

Claude Code adds layered prompt-injection defenses by default next week
Anthropic engineers said Claude Code is slated to add model training, probes, and auto-mode layers by default next week. They said Claude models now largely resist practical prompt injection.

Microsoft Copilot traces report 87% of LLM calls came from agents
A Microsoft Copilot trace analysis said 87% of LLM calls came from the agent, not direct user turns. Related posts warned token use and web requests can scale far faster than human prompt counts.

Echo Gap paper reports agents endorsed 31%–54% of their own wrong answers
The Echo Gap paper found self-improving agents can store wrongly self-scored episodes. Tested models endorsed 31% to 54% of their own wrong answers, while other work proposed RL-trained harness state and in-model memory.






