Fresh stories
Xiaomi releases open-weight MiMo-V2.6 models with 1M-token context
Xiaomi released Pro and Flash MiMo-V2.6 mixture-of-experts models with open weights and a 1M-token context window. The release includes an RL dashboard and day-one vLLM support, while training artifacts are planned.

xAI releases Grok 4.7 through coding tools and APIs
xAI released Grok 4.7 through Grok Build, APIs, Cursor, and other gateways. Early evaluations report stronger coding and knowledge-work results than Grok 4.6, with mixed results across individual coding benchmarks.


LangSmith adds Jev as a production trace judge
LangSmith now lets teams score production traces with Jev and trigger automated responses. Tests found Jev fast and competitive for groundedness, but weaker than reasoning models on math and code.

Xiaomi releases open-weight MiMo-V2.6 models with 1M-token context
Xiaomi released Pro and Flash MiMo-V2.6 mixture-of-experts models with open weights and a 1M-token context window. The release includes an RL dashboard and day-one vLLM support, while training artifacts are planned.

OpenAI says an internal model solved more than 100 long-standing math problems
OpenAI says an internal model resolved more than 100 long-standing mathematical problems, including work on the Navier–Stokes problem. OpenAI also announced an independent mathematicians' group to review emerging results, while outside discussion questioned the training and verification process.
Top storiesthis week
Qwen releases Qwen-Image-2.1 with open weights
Qwen-Image-2.1 is a 7B model that combines image generation and editing, including native RGBA output and support for up to 10 reference images. It has day-one support in ComfyUI, Diffusers, vLLM-Omni, and Ostris, but its license is non-com


Kev releases open decision models built on Qwen3
Kev is an Apache-2.0 family of 0.6B, 4B, and 8B decision models compatible with TypeSafe System One APIs. Its author reports that the 8B model reached 79.6% out-of-domain accuracy versus Jev’s 85.7%, while the 4B model runs on a 32 GB Mac.

Developers test Jev as a low-cost judge for agent evaluations
Practitioners are testing Jev as a fast semantic verifier for online evaluations and reinforcement-learning trajectories. A field analysis found it useful for progress and completion estimates, but warned against using it to detect harmful

Teknium reports Jev compaction can increase token costs
Teknium’s public evaluation says a Jev compaction strategy removes tool calls and eventually stops yielding savings. Repeated compaction can invalidate caches and increase total token costs, according to the critique.

Jev cuts Stagehand's median Act latency from 1.97s to 0.46s, Stagehand says
Stagehand says adding Jev to its browser primitives cut median Act latency from 1.97 seconds to 0.46 seconds. Jev handles page-level choices and falls back to an LLM when uncertain.






