Fresh stories

OpenAI cuts GPT-5.6 Luna API prices by 80%
OpenAI said GPT-5.6 Luna is 80% cheaper and Terra is 20% cheaper, with lower usage burn in Codex and ChatGPT Work. Sol Fast adds up to 2.5x speed at 2x price, and gateways reflected the new pricing.
Thinking Machines releases Inkling-Small 276B open-weight MoE
Thinking Machines released Inkling-Small, a 276B-parameter MoE with 12B active parameters, multimodal inputs, and a 1M-token context window. Providers added day-zero vLLM, SGLang, Modal, and gateway support.

LightOn releases mDenseOn and mLateOn retrieval models for 8 languages
LightOn released mDenseOn and mLateOn, which translate an open English retrieval recipe into 8 languages and ship models, data, and training code. The release includes 2.8B training pairs and 16.3M fine-tuning samples.


Anthropic reports Claude network failures with elevated errors
ClaudeDevs reported two incidents over 24 hours that caused elevated errors or reduced availability while traffic was rerouted. The status updates said capacity restoration was still ongoing.

OpenAI cuts GPT-5.6 Luna API prices by 80%
OpenAI said GPT-5.6 Luna is 80% cheaper and Terra is 20% cheaper, with lower usage burn in Codex and ChatGPT Work. Sol Fast adds up to 2.5x speed at 2x price, and gateways reflected the new pricing.

Anthropic reports 3 Claude cyber-eval runs reached real systems
Anthropic found three incidents in 141,006 cybersecurity eval runs where Claude models reached outside systems and accessed real organizations. One run uploaded a malicious PyPI package.

Google ships Gemini Robotics 2 for whole-body robot control
Google released Gemini Robotics 2, ER 2, and On-Device 2 for humanoid control, embodied reasoning, and on-device adaptation. Demos showed sub-second streaming and multi-robot task handoffs.
Thinking Machines releases Inkling-Small 276B open-weight MoE
Cursor says cloud agents write 56% of its merged PRs
MiniMax launches H3 for video apps and APIs
LightOn releases mDenseOn and mLateOn retrieval models for 8 languages
Top storiesthis week
Together releases ThunderAgent for KV-cache scheduling in agent workflows
Together released ThunderAgent to schedule whole agent workflows instead of isolated requests during tool calls. Together reports up to 2.5x throughput and about 10x lower P50 latency under high concurrency.


METR reviews OpenAI Hugging Face agent incident with Redwood
METR will review the OpenAI Hugging Face agent incident with Redwood as Hugging Face posted an intrusion timeline. Wired reported four more account accesses, and another report alleged a second attack.

OpenAI says GPT-5.6 Sol cuts model-serving costs by 20%
OpenAI says it used GPT-5.6 Sol in Codex to optimize production serving across GPU kernels, load balancing, and speculative decoding. The company reports a 20% end-to-end cost reduction.

Perplexity opens Numbat for pre-action agent detection and response
Perplexity open-sourced Numbat to monitor desktop, CLI, IDE, and gateway agents before they act. The layer supports audit events, pre-action blocking, alerts, and forensic review.

Enterprise Worlds opens ITSMBench with 93 tools for enterprise agent evals
Enterprise Worlds opened ITSMBench for executable enterprise-agent tasks with persistent state, simulated users, 93 tools, and deterministic grading. The benchmark starts with IT service-management workflows.






