Home/Archive/17 September 2026
Daily digest ยท 17 September 2026

ChatGPT ads talk back, and Grok Build starts to remember

Today's news was about memory, money and megawatts rather than models. No frontier lab shipped anything new, yet the product launches all pointed at persistence and distribution: OpenAI turned ChatGPT ads into two-way conversations and wired them into Shopify and HubSpot, xAI gave its Grok Build coding agent memory that survives between sessions, and Mistral got its open-weight models put behind Firefox's built-in assistant. Underneath, the money and the power told the same story, with Perplexity moving its whole model lifecycle onto Crusoe Cloud and NVIDIA posting agentic benchmark gains alongside a new alliance to unblock the grid connections those workloads need. The through-line is a maturing market: the near-term advantage is going to whoever owns the surface a user already sits in, not whoever tops a benchmark.

OpenAI

ChatGPT ads become two-way conversations, and land inside Shopify and HubSpot

If you sell anything online, this is the day ChatGPT stopped being only a place people research and became a place people can be sold to directly. OpenAI is testing Sponsored Agents, which let a user click an ad and then open a separate, clearly labelled conversation with a business-owned agent that answers follow-up questions before handing them off to the brand's site. Alongside that, advertisers can write natural-language prompts in ChatGPT to create, update and analyse campaigns through an Ads Manager plugin, and opt into AI-generated copy, imagery and automatic translation of headlines.

The practical change for small businesses is distribution: US Shopify merchants can install a ChatGPT Ads app from today with their product catalogue already wired in, and HubSpot users can run and attribute campaigns without leaving their CRM. International Shopify availability follows on 23 September, so anyone outside the US who wants a head start has roughly a week to get catalogue and landing pages in order. It builds on the billion-dollar ad business OpenAI opened up this month, and a conversational ad unit is a genuinely new format rather than a repackaged one.

Sources
OpenAI, 16 Sep Search Engine Land, 16 Sep The Next Web, 16 Sep
Also noted
OpenAI published a framework setting out when and how it will report model misalignment, including criteria and timelines for disclosing behaviour it has not yet fully explained or mitigated. It is a governance document rather than a product change, but it is the mechanism that will decide how much you get told about a model you already depend on. OpenAI, 16 Sep.
xAI (SpaceXAI)

Grok Build now remembers your project between sessions

Persistent memory is the difference between a coding agent you have to brief every morning and one that already knows your stack. Grok Build now writes notes in the background as you work, capturing conventions, decisions and project facts, then reads them back before it touches related code in a later session. xAI has exposed this as commands rather than magic, so you can inspect what it has stored with /memory, force a note with /remember, and flush or consolidate what it holds, which matters because an agent that silently remembers a wrong decision is worse than one that remembers nothing.

If you have been running Claude Code, Codex or Cursor alongside Grok Build and found the context reset annoying, this closes a real gap. It applies to new sessions from today, so existing long-running projects will need a little seeding before the benefit shows up.

Sources
xAI, 16 Sep Unite.AI, 16 Sep

Grok Voice arrives on fal for low-latency voice agents

Voice agents have been held back less by model quality than by the plumbing needed to run speech-to-speech in real time, and this removes a chunk of it. Grok Voice is now served through fal, the inference platform, so developers can build low-latency voice agents that call tools and handle real customer queries without standing up their own realtime stack. The commercial angle is the interesting one: metered per second of audio, an hour of continuous conversation lands at a few dollars, which puts always-on voice support in range for businesses that could not justify it a year ago.

If you have been evaluating ElevenLabs agents or OpenAI's realtime API for a support or booking flow, there is now a third serious option to price against. Coming the day after Gemini 3.8 Live topped the speech quality tables, it confirms voice has become a properly contested market. This one surfaced on xAI's own X account before any significant press coverage.

Sources
SpaceXAI on X, 16 Sep [Direct] xAI, 16 Sep
Mistral AI

Firefox's AI assistant is now running on Mistral models

This is the first time a mainstream browser has put an open-weights European model behind its built-in assistant, and it gives ordinary users an AI experience that is not routed through OpenAI, Google or Anthropic. Mozilla's Firefox Smart Window, currently in beta, uses Mistral models to handle messy searches, recall something you clicked away from, and pull together information across your open tabs. Availability covers France and North America today, with the United Kingdom and Germany expected later this year, so UK readers can see the feature coming but cannot test it yet.

The privacy terms are the part worth reading: conversations are not saved on Mozilla's servers by default and Mistral has agreed to zero data retention, which is a materially different default from most assistant integrations. For anyone weighing up which assistant to recommend to privacy-conscious clients or family, Firefox has just become a credible answer.

Sources
Mistral, 16 Sep Mozilla Blog, 16 Sep Mistral on X, 16 Sep [Direct]
Perplexity

Perplexity moves its whole model lifecycle onto Crusoe Cloud

Perplexity has signed a multi-year deal to train and serve its models on a single provider, running frontier training on Crusoe's NVIDIA GB300 NVL72 clusters and production inference through Crusoe's managed service. Consolidating training and serving on one platform is usually a bet on cost and latency rather than capability, so the thing to watch over the next few months is whether Perplexity's response times and free-tier limits shift.

The deal runs both ways, with Crusoe rolling Perplexity Enterprise Pro and Max out to its own eighteen hundred staff, a reminder that enterprise search deals are increasingly bundled into infrastructure contracts. For anyone comparing Perplexity against ChatGPT search or Gemini for research, this is a signal about the economics underneath the product rather than the product itself.

Sources
AIwire, 16 Sep
NVIDIA

Vera Rubin NVL72 posts its first MLPerf numbers, agentic workloads the headline

Benchmark posts are easy to skim past, but this one contains a number that explains where inference pricing is heading. In its MLPerf Inference v6.1 debut, Vera Rubin NVL72 delivered up to 3.7 times the throughput of GB300 NVL72 on the Qwen3-VL benchmark and around 2.5 times on DeepSeek-R1, while pure software optimisations added up to 1.6 times over the previous round. The striking figure is on SemiAnalysis AgentX, a benchmark built to reflect agentic workloads, where the preview system reported roughly thirty times the performance of its predecessor.

If that holds outside preview conditions, the cost of running long-horizon agents rather than single prompts falls sharply, extending the tokens-per-megawatt argument from earlier this week. These are peer-reviewed preview results rather than shipping hardware, so treat the multiples as direction of travel.

Sources
NVIDIA Blog, 16 Sep MLCommons, 16 Sep

NVIDIA, Google and Emerald AI form an alliance for grid-flexible data centres

Power, not silicon, is the constraint most likely to slow AI capacity over the next two years, and this is the industry organising around it. The AI Energy Management Alliance brings together around twenty companies, including Anthropic, National Grid, AES, Constellation, NRG and RWE, to push data centres that can vary their draw from the grid by shifting workloads, discharging storage or leaning on paired generation. The stated aim is to speed up grid interconnection for new sites, the bottleneck that currently adds years to a build.

For users, the downstream effect is capacity and price: compute that cannot be connected does not become tokens you can buy. It is a coalition rather than a product, so nothing changes this week, but it tells you where the next round of capacity announcements will come from.

Sources
NVIDIA Blog, 16 Sep Axios, 16 Sep
Quiet in the last 24 hours
Anthropic ยท Google DeepMind Nothing significant. Anthropic's most recent entry is its 10 September threat report and Google's is Gemini 3.8 Live on 15 September, both outside the window; both appear as founding members of the AI Energy Management Alliance above.
Meta ยท Microsoft ยท DeepSeek ยท Qwen Nothing qualifying in the window. DeepSeek's V4.1-Flash on 10 September remains the most recent release from either Chinese lab.
Hugging Face ยท ElevenLabs ยท Others Nothing significant from Hugging Face, ElevenLabs, IBM, Snowflake, Cohere, Amazon Bedrock, Runway, Stability AI or Scale AI in the last 24 hours.

Industry themes

Today's releases were about memory, money and megawatts rather than models. Not one frontier lab shipped a new model, yet three of the five product announcements were about persistence and distribution: Grok Build remembering your project, ChatGPT ads remembering the conversation, and Mistral reaching consumers through someone else's browser. That is what a maturing market looks like, and it suggests the near-term advantage goes to whoever owns the surface a user already sits in, not whoever tops a benchmark.

The second pattern is that the interesting numbers have moved from model quality to agent economics. NVIDIA's roughly thirtyfold gain on an agentic inference benchmark, and the energy alliance formed the same day to unblock the power those workloads need, are two halves of the same story: the industry is retooling for long-running agents rather than single prompts, and it has worked out that the constraint is grid connection rather than chips.

Worth noting on method: the Grok Voice launch on fal surfaced on xAI's own X account well ahead of meaningful press coverage, and the Mistral and Mozilla announcement appeared on the company channel at the same moment as the blog post. Anyone relying only on tech press aggregators would have picked up both a day late.

โ† 16 September 2026 All issues โ†’

Every day I sift through the noise, cut out the hype, and serve up the AI updates that actually matter for your business. Straight to your inbox before 8am. No fluff, no jargon, no faff.

๐Ÿ‘ฉ๐Ÿ‘จ๐Ÿ‘ฉ๐Ÿ‘จ+
Join 2,400+ business professionals already subscribed

You're in! First issue lands tomorrow morning.