Home/Archive/14 August 2026
Daily digest · 14 August 2026

Gemini 3.7 Flash and DeepSeek V4-Pro race the mid-tier down

The frontier stayed still while the mid-tier got cheaper and faster. Google's Gemini 3.7 Flash halved its token price against 3.6, and DeepSeek pushed V4-Pro to general availability alongside an MIT-licensed agent framework called Harness that passed 33,000 GitHub stars within hours. OpenAI's contribution was speed rather than a new model, previewing an Ultrafast mode that runs GPT-5.6 Sol some fourteen times quicker on Cerebras silicon, plus an opt-in Computer History feature that gives the desktop app a memory of your work, though not in the UK. Around those, Grok 4.6 kept propagating into Perplexity and Pi barely a day after launch, and Perplexity trimmed its search costs again. The theme is unmistakable: with quality converging, the fight is now over price, latency and how cleanly a model wires into real work.

Google DeepMind

Gemini 3.7 Flash lands at half the price of 3.6 Flash

If you run anything at volume on Gemini, this is the release that changes your unit economics. Introductory pricing is 75 cents per million input tokens and 3.75 dollars per million output tokens, exactly half what 3.6 Flash cost, and Google is pitching it squarely at coding, knowledge work and web development. The model card is candid that this is a refinement of 3.6 Flash rather than a fresh pretraining run, so the gains come from algorithmic work on the reasoning core rather than raw scale.

That matters in practice because you are getting a faster, cheaper model whose behaviour should feel familiar if you have already tuned prompts against 3.6. The obvious move is to rerun your evals on 3.7 Flash before your next billing cycle, because a halved token price on an unchanged workload is free margin. It also puts real pressure on the mid-tier of every rival API, which is where most production traffic actually sits.

Sources
Google DeepMind, 13 Aug [Direct] Reuters, 13 Aug DeepMind on X, 13 Aug [Direct]
DeepSeek

DeepSeek-V4-Pro leaves preview and goes generally available

DeepSeek shipped V4-Pro to app, web and API, and the headline is agents rather than chat. Thinking effort is now a three-way dial on both V4-Pro and V4-Flash, low, high and max, so you can spend cheap tokens on simple lookups and reserve deep reasoning for genuinely hard multi-step work instead of paying a flat premium on everything. The API now speaks the OpenAI Responses format natively and ships with Codex integration, so swapping DeepSeek in behind an existing OpenAI-shaped integration is closer to a config change than a rewrite.

The catch is pricing: rates rise on 16 August, with peak-hour output moving to 3.96 dollars per million from the current flat 0.87 dollars, alongside a new peak and off-peak split that halves the cost outside busy hours. If you have batch workloads with no latency requirement, moving them to off-peak windows before Sunday is the cheapest optimisation available to you this week.

Sources
DeepSeek on X, 13 Aug [Direct] VentureBeat, 13 Aug Tech Times, 13 Aug

DeepSeek Harness opens as an MIT-licensed agent framework

Announced on DeepSeek's own X account before the press caught up, Harness is an open-source agent harness built on the Cordis meta-framework around a single idea: everything is a plugin. Models, tools, skills, sessions, sandboxes, filesystems, loops, orchestration and the UI are all mounted the same way, so extending it means adding a plugin rather than patching a privileged core. It ships with four presets, from a full coding agent with filesystem, shell, web search, subagents and plan mode, down to a Minimal profile with only bash and a string-replace editor.

The repo passed 33,000 GitHub stars within hours of going live, which tells you how hungry developers are for an agent stack they can actually take apart. If you have been locked into a closed coding agent and resenting it, this is the first credible MIT-licensed alternative worth a weekend.

Sources
DeepSeek on X, 13 Aug [Direct] The New Stack, 13 Aug GitHub, 13 Aug
OpenAI

Ultrafast mode previews GPT-5.6 Sol at up to 14 times the speed

This is a latency play, not an intelligence one, and that is precisely why it matters. Ultrafast runs the same GPT-5.6 Sol you already use at up to 750 output tokens per second, roughly fourteen times standard throughput, on Cerebras hardware rather than the usual stack. Anything where a human is waiting on the model becomes a different product at that speed: live incident response, customer support, market analysis, e-commerce flows that currently feel sluggish.

Access is limited to a select group of API customers for now, widening as capacity allows, so treat it as a signal about where inference is heading rather than something you can switch on today. The strategic read is that OpenAI is willing to route frontier traffic through third-party silicon to win on speed, which quietly weakens the assumption that model quality alone decides who wins a workload.

Sources
OpenAI, 13 Aug [Direct] TechCrunch, 13 Aug OpenAI on X, 13 Aug [Direct]

Computer History gives the ChatGPT desktop app a memory

Announced on OpenAI's X account rather than the newsroom, Computer History is an opt-in macOS feature that turns your activity across apps and websites into a timeline that ChatGPT and Codex can draw on. Crucially it reads interaction events through macOS accessibility rather than taking screenshots, a meaningfully less invasive design than the Chronicle feature it replaces. You choose which apps and sites contribute, you can pause collection, and you can review or delete the history whenever you like.

The practical payoff is less re-explaining yourself at the start of every session, which is the single biggest tax on using an assistant for real work. One important caveat for readers here: the rollout covers Pro, Business and Enterprise globally but explicitly excludes the UK, the EEA and Switzerland at launch, so this is one to watch rather than one to try.

Sources
OpenAI on X, 13 Aug [Direct] OpenAI docs, 13 Aug 9to5Mac, 13 Aug
Also noted
OpenAI also published The builder's guide to GPT-5.6, a practical reference for teams choosing between the 5.6 variants, and appointed Dali Rajic as Chief Revenue Officer. (Source: OpenAI Newsroom, 13 Aug)
Perplexity

Grok 4.6 arrives in Perplexity as search costs fall again

Perplexity has added xAI's newest model, Grok 4.6, to both its assistant and its Computer agent product, framed entirely around cost per unit of quality. On Perplexity's own WANDR benchmark, Grok 4.6 sits on the Pareto frontier of performance against cost, matching Fable 5 results at over 60 per cent lower cost per task. It shows how quickly frontier models now propagate into third-party surfaces, barely a day after xAI shipped 4.6 itself.

Separately, Perplexity tuned its Search as Code pipeline, which lets a model write Python that calls the search stack directly instead of looping through tool calls one at a time, cutting cost per task by nearly a tenth. That is not a headline number, but it compounds: on a deep-research task the original architecture already cut token use by around 85 per cent against a conventional tool-calling stack. The lesson for anyone building retrieval is that the expensive part is rarely the search itself, it is the round trips.

Sources
Perplexity on X, 13 Aug [Direct] Perplexity Research, Aug
xAI

Grok 4.6 reaches Pi as a Wix plugin lands in Grok Build

A day after shipping Grok 4.6, xAI confirmed the model is now available in Pi, extending the distribution of a release built for long-running agent tasks at 2 dollars per million input and 6 dollars per million output. The pattern is worth noting: xAI is pushing 4.6 into every surface it can reach, Cursor, Grok Build, Grok Bot, the API and now Pi, within roughly forty-eight hours of launch, and since the price did not move there are fewer excuses to stay on 4.5.

Grok Build also gained a Wix plugin, letting builders create apps and sites, connect any frontend to Wix business services through Wix Headless, and work with the Wix MCP, all from the Grok CLI. It is a small announcement with a clear direction of travel: agent CLIs are becoming the place where you administer a real business platform, not just scaffold code. If you build or maintain client sites on Wix, an agent that can drive the platform end to end from a terminal changes the shape of routine maintenance work.

Sources
SpaceXAI on X, 13 Aug [Direct]
Also noted · NVIDIA
NVIDIA's AI account confirmed that agents in Cursor can now pull from over 300 NVIDIA skills spanning more than thirty products, including CUDA, NeMo and RAG workflows. Skills are the portable packaging format that took hold across coding agents this year, and NVIDIA putting its full product surface into that format removes a lot of documentation-hunting from GPU, inference and robotics work. (Source: NVIDIA AI on X, 13 Aug)
Quiet in the last 24 hours
Anthropic · Meta AI Nothing significant. Anthropic's most recent newsroom entry remains the 7 August Fable 5 biology-safeguards update, and Meta's stands at the 10 August Muse Glimmer open-weight release.
Microsoft · Mistral · Qwen Nothing significant. Microsoft Research last posted the MindTopo benchmark on 12 August, Mistral's last item was the 11 August European inference roadmap, and Qwen's recent releases predate the window.
Hugging Face · ElevenLabs Nothing significant in products or releases in the last 24 hours from either.

Industry themes

Three of today's biggest items were price cuts dressed as launches. Gemini 3.7 Flash halved its token cost, Perplexity shaved nearly a tenth off Search as Code, and Grok 4.6 landed in Perplexity claiming Fable 5 quality at over 60 per cent lower cost per task. The frontier is no longer where the competition is happening for most working developers; the mid-tier is, and it is getting cheaper roughly every week.

The second pattern is that speed and plumbing are becoming the product. OpenAI's Ultrafast mode buys nothing but latency, DeepSeek's V4-Pro sells a three-position thinking dial and native OpenAI Responses compatibility, and DeepSeek Harness gives away the agent scaffolding entirely under MIT. When models converge on quality, the differentiator moves to how fast and how cheaply you can wire them into something real.

Worth flagging that half of today's material never reached the newsroom. ChatGPT's Computer History, Grok 4.6 in Pi and Perplexity, the Wix plugin for Grok Build and Perplexity's Search as Code optimisations all surfaced on official company X accounts with no press release behind them. If you follow this industry through headlines alone, you missed most of what shipped today.

← 13 August 2026 All issues →

Every day I sift through the noise, cut out the hype, and serve up the AI updates that actually matter for your business. Straight to your inbox before 8am. No fluff, no jargon, no faff.

👩👨👩👨+
Join 2,400+ business professionals already subscribed

You're in! First issue lands tomorrow morning.