Fresh stories
OpenAI reports Jalapeño delivers 1.5–1.9× more work per watt
OpenAI says Jalapeño delivered 1.5–1.9× more work per watt and 1.7–3.6× lower end-to-end latency than NVIDIA systems in its tests. The company plans to deploy the inference chip in its compute infrastructure by year-end.

Vercel Connect reaches GA with authenticated MCP for 100+ services
Vercel Connect is generally available for connecting apps and agents to more than 100 services, including Slack, Linear, GitHub, and Notion. It uses short-lived scoped tokens, RBAC, and audit logs for delegated service access.


DataSpace finds harnesses shift data-task accuracy by 15 points
Across 410 cross-source data tasks, DataSpace found fixed-model accuracy ranged from 30.98% to 46.34% across harnesses. Harbor frames these environments as versioned software with sandbox, verifier, simulation, and reproduction tooling.

Perplexity launches Portable Computer on DGX Spark with a 27B model
Perplexity’s Portable Computer runs its orchestrator, subagents, and harness locally on NVIDIA DGX Spark with a post-trained 27B model. Frontier-model escalation requires user approval and flags PII before text is sent externally.

OpenAI reports Jalapeño delivers 1.5–1.9× more work per watt
OpenAI says Jalapeño delivered 1.5–1.9× more work per watt and 1.7–3.6× lower end-to-end latency than NVIDIA systems in its tests. The company plans to deploy the inference chip in its compute infrastructure by year-end.

ChatGPT browser adds WebMCP support for sites as tools
ChatGPT’s desktop browser and ChatGPT Sites can now use WebMCP-compatible websites as tools for ChatGPT and Codex. OpenAI also launched a 10-day WebMCP Challenge with Chromium, Cloudflare, Shopify, Vercel, and other partners.

Prime Intellect publishes Prime Agent report with 7-day Factorio evaluation
Prime Intellect’s report describes a self-improving long-horizon agent harness with persistent memory, skills, prompts, and subagent specifications. Its Factorio evaluation ran for seven days using 23.4 million output tokens across 633 trajectories.
Vercel Connect reaches GA with authenticated MCP for 100+ services
OpenAI launches $100 ChatGPT Business Premium seats
Figure launches Index robot-training dataset with 16M video uploads
DataSpace finds harnesses shift data-task accuracy by 15 points
Top storiesthis week
NVIDIA puts Groq 3 LPX into Vera Rubin production
Groq 3 LPX adds dedicated token generation to NVIDIA Vera Rubin systems, with Groq and Nebius among planned deployers. Artificial Analysis measured about 3,400 output tokens per second on Gemma 4 31B.


ASI-Bench finds full procedures lift research-agent scores to 50.91
Across 60 research projects, ASI-Bench found full procedures averaged 50.91, versus 29.10 for prompts naming only a method. Other evaluations similarly measure whether procedural skills improve execution rather than merely adding more instructions.

WAN 3.0 launches on API platforms with 30-second video output
WAN 3.0 is now available through Replicate, OpenRouter, and ComfyUI for text-, image-, and reference-driven video generation up to 1080p. Pika says the model supports up to 20 references and 30-second output, while quality and price comparisons remain vendor-reported.

Study finds instructions account for 60.5% of coding-agent reading
A study of 557 coding-agent sessions finds instruction files and working notes account for most of what agents read. Related work puts installed skills' standing prompt cost at 50–280 tokens, while Backpass turns past sessions into reviewable AGENTS.M files.

Practitioners propose a standard harness for agent benchmarks
Practitioners argue that coding-agent results depend heavily on the evaluation harness, including tools, execution control, compaction, and token handling. They propose stable common harnesses rather than vendor-specific setups.



