Skip to content
AI Primer

Explore what's new in AI

Where people deep in AI come to stay current.

Filters

Category

Tags

OpenAI cuts GPT-5.6 Luna API prices by 80%
New

OpenAI cuts GPT-5.6 Luna API prices by 80%

OpenAI said GPT-5.6 Luna is 80% cheaper and Terra is 20% cheaper, with lower usage burn in Codex and ChatGPT Work. Sol Fast adds up to 2.5x speed at 2x price, and gateways reflected the new pricing.

🧠GPT30th July·7 min read
Breaking

Thinking Machines releases Inkling-Small 276B open-weight MoE

Thinking Machines released Inkling-Small, a 276B-parameter MoE with 12B active parameters, multimodal inputs, and a 1M-token context window. Providers added day-zero vLLM, SGLang, Modal, and gateway support.

Thinking Machines releases Inkling-Small 276B open-weight MoE
New
Inkling·30th July·7 min read
Breaking

LightOn releases mDenseOn and mLateOn retrieval models for 8 languages

LightOn released mDenseOn and mLateOn, which translate an open English retrieval recipe into 8 languages and ship models, data, and training code. The release includes 2.8B training pairs and 16.3M fine-tuning samples.

LightOn releases mDenseOn and mLateOn retrieval models for 8 languages
New
RAG·30th July·6 min read
See all stories →
⌨️Agentic Engineering(7)
🧠Models, Serving & APIs(7)
⚙️Building Agents(10)
🛡️Trust, Evaluation & Reliability(6)
🔎Knowledge, Memory & Retrieval(6)
📈Adoption & Market Strategy(12)

Top storiesthis week

Breaking

Together releases ThunderAgent for KV-cache scheduling in agent workflows

Together released ThunderAgent to schedule whole agent workflows instead of isolated requests during tool calls. Together reports up to 2.5x throughput and about 10x lower P50 latency under high concurrency.

Together releases ThunderAgent for KV-cache scheduling in agent workflows
New
KV Cache·29th July·6 min read
See all stories →
AI PrimerAI Primer

Your daily guide to AI tools, workflows, and creative inspiration.

© 2026 AI Primer. All rights reserved.