Home/Archive/26 August 2026
Daily digest · 26 August 2026

Perplexity moves the AI agent off the cloud

The interesting work in AI has quietly moved off the model and onto everything around it. The day's biggest release is Perplexity's Portable Computer, which runs a full agent, orchestrator and models on hardware you already own so that local steps cost nothing in billing credits, changing the economics of agentic work rather than just the privacy story. Around it, NVIDIA shipped the kind of unglamorous plumbing that decides whether an inference API feels dependable during a bad hour, along with a stable Python route into CUDA, while OpenAI folded workspace administration into the ChatGPT chat window and published the first benchmarks for its own Jalapeno inference chip. Nobody released a new frontier model, but the through-line is clear: the next round of cost pressure is coming from where inference runs and how gracefully it fails, and per-token pricing is becoming something you can design around rather than simply absorb.

Perplexity

Portable Computer runs the whole agent on hardware you already own

If you have watched per-token costs climb on anything agentic, this is the most consequential release of the day. Portable Computer runs Perplexity's agent harness, orchestrator and models entirely on local hardware, starting with NVIDIA's DGX Spark and Linux machines carrying RTX GPUs. Every task begins on the device by default, local steps consume no billing credits, and the system asks permission before routing any single step to a frontier model in the cloud. You choose between Qwen 3.8 27B and PPLX 27B, Perplexity's own post-trained variant, with NVIDIA Nemotron 3.5 Lightning listed as coming soon.

It is live now for Pro and Max subscribers rather than sitting behind a waitlist, though the hardware requirement keeps it out of reach for most people today. Read alongside the rest of the week, it is the sharpest version yet of the argument that agent economics are a hardware decision, not a model one, and the first credible answer to spiralling agentic token bills that does not involve a smaller cloud model. Treat it as a strong signal of where pricing is heading rather than something you will run this afternoon.

Sources
Perplexity, 25 Aug VentureBeat, 25 Aug Perplexity on X, 25 Aug [Direct]
NVIDIA

Shadow engine recovery cuts inference restarts from minutes to seconds

This one matters if you self-host models, and indirectly if you buy inference from anyone who does. When an LLM engine process crashes, the standard fix is a cold restart that reloads weights, recompiles kernels and recaptures CUDA graphs, which NVIDIA measured at 283 seconds on a two-worker GLM-5.2 deployment. Shadow engine recovery keeps a fully initialised standby engine parked on the same GPUs, sharing a single physical copy of the weights, and promotes it in 7.3 seconds when the active engine dies. The user-facing effect is the part worth noting: median time to first token after the fault fell from 23.8 seconds to 1.3 seconds, and requests dropping below 20 tokens per second went from 226 to zero.

It is an opt-in preview that needs Kubernetes 1.34 with Dynamic Resource Allocation and works primarily with vLLM, so it is evaluation-grade rather than production-ready. Even so, this is exactly the sort of reliability plumbing that sits underneath the generation-speed race between inference clouds, and it decides whether a model API feels dependable when something goes wrong.

Sources
NVIDIA Technical Blog, 25 Aug NVIDIA AI on X, 26 Aug [Direct]

CUDA Python 1.0 gives Python developers a stable route into CUDA

Shipping with CUDA 13.3, this is the first CUDA Python release to commit to semantic versioning across cuda.core, cuda.compute, cuda.bindings, nvmath-python and cuda-pathfinder. Breaking changes now land only in major releases, and anything scheduled for removal gets deprecated first with a stated replacement. If you have avoided writing GPU code in Python because the bindings kept moving under you, that objection has largely been answered.

It also consolidates several overlapping projects onto one foundation, which should make GPU libraries interoperate more predictably and brings features such as green contexts and process checkpointing within reach from Python. Relevant mainly if you maintain anything that touches CUDA without going through PyTorch.

Sources
NVIDIA Technical Blog, 25 Aug
OpenAI

The Admin plugin pulls workspace management into the chat

Anyone administering a ChatGPT Work or Codex deployment has been living in a separate settings console, toggling between analytics and controls. The new Admin plugin folds that work into the chat itself, so an admin can review credit usage across the workspace, spot members approaching their limits, add and remove people, diagnose a permissions problem, and approve or deny spending requests in one thread. The important detail is that it carries out authorised changes rather than only reporting on them, which makes it a genuine agentic surface rather than a dashboard with a chat box bolted on.

If a meaningful share of your week goes on onboarding, offboarding and access requests, this removes a lot of routine clicking. Expect other vendors to copy the pattern, because admin work suits a conversational interface with explicit confirmation steps rather well.

Sources
OpenAI, 25 Aug 9to5Mac, 25 Aug

Jalapeno's first benchmarks land, and OpenAI's own silicon looks quick

OpenAI published the first public results for Jalapeno, its custom inference chip, tested on the SemiAnalysis InferenceX suite against leading commercial systems using GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5 1T. The claimed gains are 1.7x to 3.6x lower end-to-end latency, 2.1x to 4.1x better interactivity for agent-style workloads, and 1.5x to 1.9x more work per watt at peak throughput. None of this is something you can buy or run today, so the practical relevance is indirect: it is the reason to expect ChatGPT and Codex responses to get faster and cheaper over the next two years rather than the next two months.

Deployment into OpenAI's own infrastructure starts by year end in small volumes, with a second generation already in development. The broader read is that inference cost is where the next round of pricing pressure comes from, and OpenAI is now less exposed to a single supplier for it.

Sources
OpenAI, 25 Aug TechCrunch, 25 Aug OpenAI on X, 25 Aug [Direct]
Quiet in the last 24 hours
Anthropic · Google DeepMind · Meta Nothing significant. Anthropic's one post in the window covered funding for wellbeing evaluations, which is research funding rather than a product change, and Google DeepMind and Meta had nothing qualifying.
xAI · Microsoft · Mistral Nothing significant in the window. xAI's most recent product post remains Grok Bot's wider plan availability on 21 August; Microsoft and Mistral were quiet.
DeepSeek · Qwen · Others Nothing qualifying from DeepSeek, Qwen, IBM, Snowflake, Cohere, Hugging Face or ElevenLabs in the last 24 hours.

Industry themes

The clearest pattern today is that the interesting work has moved from the model to everything around it. Perplexity's Portable Computer and NVIDIA's shadow engine recovery are both bets that the next gains come from where inference runs and how gracefully it fails, not from another benchmark point, and both surfaced on the companies' own X accounts before the trade press picked them up.

Alongside that, OpenAI's Jalapeno results and the same company's Admin plugin show the two ends of one squeeze: drive the cost of a token down at the silicon layer, then make the seats and credits easier to govern at the top. Between them, per-token pricing is turning into a variable you can design around, whether by running smaller models locally or by managing spend inside the workspace itself.

The open question is whether Anthropic and Google answer the local-first move. A credible on-device agent from either would shift the pricing conversation quickly, and for anyone building on these tools the practical takeaway is to start treating inference cost as something you architect around rather than simply absorb.

← 25 August 2026 All issues →

Every day I sift through the noise, cut out the hype, and serve up the AI updates that actually matter for your business. Straight to your inbox before 8am. No fluff, no jargon, no faff.

👩👨👩👨+
Join 2,400+ business professionals already subscribed

You're in! First issue lands tomorrow morning.