Home/Archive/13 August 2026
Daily digest · 13 August 2026

Grok 4.6 matches the frontier at half the price

Today was about squeezing more out of the same money, at both ends of the market. SpaceXAI's Grok 4.6 matched a frontier index score while holding its predecessor's price and claiming half the cost of rivals, and on the open side a 2-bit build of NVIDIA's Nemotron 3.5 Lightning ran ten minutes of continuous tool use inside 22GB of VRAM. Between those poles, DeepMind did something rarer than a price cut, shipping a sign-language keyboard input straight into Gboard on Pixel, while Cohere, Runway and Hugging Face all pushed on the connective tissue around models rather than the models themselves. The through-line is that raw capability is increasingly a given, and the contest has moved to price, footprint and how easily a model plugs into where work already happens.

xAI (SpaceXAI)

Grok 4.6 lands at the same price as 4.5 and undercuts rivals by half

If you are paying for frontier tokens, this is the most consequential pricing move of the day. Grok 4.6 lands at 2 dollars per million input tokens and 6 dollars per million output tokens, which SpaceXAI describes as half the cost of comparable frontier models, and it does so without raising the price of the tier it replaces. The stated focus is long-running agents and more ambitious interactive and visual work, meaning tasks that run across many steps rather than single-shot answers, and independent index scoring puts it level with GPT-5.6 Sol at 61 on the AA Intelligence Index.

It is available immediately in Grok Build, Cursor, Grok Bot and the API, with double the included usage inside Grok Build and Cursor for the first week. For anyone running agent workloads where token spend scales with step count rather than prompt count, that combination of parity performance and halved pricing is worth a re-benchmark this week rather than next quarter.

Sources
SpaceXAI, 12 Aug [Direct] SpaceXAI on X, 12 Aug [Direct] Unite.AI, 12 Aug
Google DeepMind

SL2T puts sign-language input into a shipping phone keyboard

This is the rare research-to-product jump that changes what a device can do rather than how fast it does it. SL2T is a sign-language-to-text model that now powers signing input in Gboard and Live Transcribe on Pixel 11, starting with American Sign Language to English, so a Deaf user can sign at the camera to run a search, draft a message or query Gemini instead of typing. DeepMind says the model was trained on more than 100,000 hours of signing across over 50 sign languages, and it translates rather than transcribes, skipping the gloss step that made earlier systems stilted.

Privacy is handled by running pose tracking on-device with MediaPipe Holistic so only geometric coordinates leave the handset, a design pattern worth noting if you are building anything camera-based. The wider signal for builders is that Google is now comfortable shipping a novel input modality straight into a default keyboard, and expansion to further sign languages is stated as the roadmap.

Sources
Google DeepMind, 12 Aug [Direct] DeepMind on X, 12 Aug [Direct] Engadget, 12 Aug
Cohere

North Micro Vision ships as a 2.4B open vision model for documents

Small open vision models are where a lot of practical automation actually gets built, and this one is aimed squarely at document work. North Micro Vision Instruct is a 2.4-billion-parameter vision-language model released under Apache 2.0, with native-resolution image handling that preserves aspect ratio and fine detail rather than squashing everything to a fixed square. Cohere claims it outperforms Gemma 4 E2B and Ministral 3 3B across visual understanding benchmarks, strongest in document understanding and visual question answering, with a 128K context backbone although the validated multimodal range is 8K.

Crucially for anyone who wants to deploy it, the launch came with a fine-tuning recipe on Axolotl, MLX support for Apple silicon, and an NVIDIA AutoModel recipe, so you can train and serve it without assembling the toolchain yourself. If you have been paying per page for OCR and document-extraction APIs, a permissively licensed model at this size is a genuine cost lever.

Sources
Cohere on X, 12 Aug [Direct] Cohere Labs, 12 Aug Hugging Face, 12 Aug
Runway

Runway Agent gains Figma, Dropbox and Notion connectors

The friction in creative AI tooling has stopped being model quality and started being asset logistics. Runway has connected Agent directly to Figma, Dropbox and Notion, so designs, files and documents sync in rather than being exported, renamed and re-uploaded every time a job runs. It is available across all paid plans rather than gated to enterprise, which matters if you are a small studio or a solo operator whose whole workflow already lives in those three tools.

The broader pattern is that every serious agent product is converging on the same conclusion: an agent without access to where your work already sits is a demo rather than a tool. Expect connector coverage, not raw generation quality, to become the thing people compare when choosing between creative agents.

Sources
Runway on X, 12 Aug [Direct]
Also noted
[Direct] Runway also added SpaceXAI's Grok Imagine Image 2.0 to its model picker, continuing its positioning as a neutral surface rather than a single-model studio. You can now compare it against the other image and video models in the same project without moving assets or holding a second subscription, which reinforces the aggregator play where the value sits in the editing surface and the model roster rather than in owning the weights. (Source: Runway on X, 12 Aug)
NVIDIA

Nemotron 3.5 Lightning runs 10 minutes of tool calls in 22GB of VRAM

The interesting number here is not the benchmark, it is the memory footprint. A 2-bit quantised build of NVIDIA's open 30B Nemotron 3.5 Lightning ran tool calls continuously for ten minutes on 22GB of VRAM, citing more than 80 websites, executing code and searching for real-world locations in a single session. GGUF weights are on Hugging Face via Unsloth with a matching training guide, which means a single high-end consumer card is now enough to run a genuinely agentic open model locally rather than just chat with one.

NVIDIA also spent the day highlighting partners post-training the model for their own domains and tools, which is the part that determines whether an open release becomes an ecosystem or a download. If local, always-on agents are on your roadmap, this is the configuration to test against before committing to a hosted API.

Sources
NVIDIA AI on X, 12 Aug [Direct] Unsloth on Hugging Face, 12 Aug
IBM

IBM and Together AI sign a 240 million dollar inference deal

This is an infrastructure story with a direct read-through for anyone serving open models in production. IBM and Together AI have committed to a multi-year, 240 million dollar build on IBM Cloud, starting with roughly 2,000 NVIDIA Blackwell-generation chips in HGX B300 systems with Spectrum-X networking, all of it aimed at inference rather than training. Together AI's business is helping enterprises run open models, so the capacity is effectively a bet that companies want a serving layer not owned by OpenAI, Anthropic or Google.

If you have been holding off on open-weight deployments because production-grade inference capacity felt scarce or expensive, the supply side is moving in your favour. Expect inference pricing on open models to keep drifting downwards as this capacity lands.

Sources
Tech Startups, 12 Aug
Also noted
Two smaller items rounded out the day. Hugging Face's Gradio 6.24 now saves and replays every run of an app in browser local storage with no code changes, a useful win for evaluation and live demos (Gradio changelog). And Qwen-Image-3.0 became selectable through OpenArt, the quiet mechanism by which Chinese open models keep reaching Western users who will never touch the Alibaba Cloud console (Alibaba Qwen on X, 12 Aug). Microsoft Research also published MindTopo, a benchmark for topological reasoning aimed at robotics and navigation (MSFT Research on X, 12 Aug).
Quiet in the last 24 hours
OpenAI · Anthropic Nothing significant. OpenAI's ChatGPT desktop app for Linux, the Codex preview and Daybreak on AWS all landed on 11 August, just outside the window; Anthropic's most recent posts predate the cutoff.
Meta · Mistral Nothing significant. Meta's Muse Glimmer open-weight release was 10 August, and Mistral's European compute and regional-inference news posted on 11 August.
DeepSeek · Perplexity · ElevenLabs Nothing significant in products or releases in the last 24 hours from any of the three.

Industry themes

The clearest pattern today is price and footprint compression at the top of the market. Grok 4.6 matched a frontier index score while holding its predecessor's price and claiming half the cost of rivals, and on the open side a 2-bit Nemotron 3.5 Lightning ran ten minutes of continuous tool use inside 22GB of VRAM. The question for buyers is shifting from which model is smartest to which one is cheap enough, and small enough, to run at the scale you actually need.

The competitive edge in agent products has moved decisively from model quality to connective tissue. Runway wired Agent into Figma, Dropbox and Notion on every paid plan, while Cohere shipped its new vision model with Axolotl, MLX and NVIDIA recipes on day one. An agent that cannot reach your files, and a model you cannot deploy without a week of plumbing, are both losing ground to ones that can.

The background context is scale rather than capability, with Google confirming the Gemini app has passed a billion monthly active users and Hugging Face reporting Transformers.js crossing ten million monthly downloads. Both say more about where inference will physically run than about what any single model can do. It is also worth noting that Grok 4.6, SL2T and Cohere's North Micro Vision all broke on the companies' own X channels hours before the trade press settled on a framing.

← 12 August 2026 All issues →

Every day I sift through the noise, cut out the hype, and serve up the AI updates that actually matter for your business. Straight to your inbox before 8am. No fluff, no jargon, no faff.

👩👨👩👨+
Join 2,400+ business professionals already subscribed

You're in! First issue lands tomorrow morning.