The day's news was almost entirely about where a model is allowed to run and how you prove it works, not about any model getting smarter. OpenAI committed to keeping Zero Data Retention on its frontier models and previewed a Private Safety Processing system that spots abuse without staff seeing your content, a trust move aimed squarely at cautious enterprises. SpaceXAI made Grok 4.6 generally available on Amazon Bedrock, removing the awkward separate-vendor step that has kept it out of many AWS shops. NVIDIA, meanwhile, put hard numbers on whether agent skills actually help, open-sourcing SkillEvaluator, while Cerebras kept pushing inference throughput and Microsoft finally closed a long-running Copilot data-leak flaw. The theme is unmistakable: with quality converging, the contest has moved to governance, distribution and speed.
Zero Data Retention survives, and Private Safety Processing goes to preview
If your business pushes anything sensitive through the OpenAI API, this is the announcement that decides whether you can keep doing it. OpenAI has committed to continuing Zero Data Retention on frontier models, rejecting the trade where you hand over prompt logs in return for access to the strongest model. In its place comes Private Safety Processing, a preview system that spots abuse patterns across related interactions without any OpenAI personnel seeing the underlying content.
Your content either stays on infrastructure you control, or sits in OpenAI storage encrypted with keys OpenAI does not hold, and only a narrow risk signal ever crosses the boundary. Rollout and a technical white paper are promised for September, so if you have a procurement or compliance review in the diary, that is the date to plan around.
Grok 4.6 is live on Amazon Bedrock
Grok has been the awkward model to adopt at work, because using it usually meant standing up a separate vendor relationship. That friction is gone now that Grok 4.6 is generally available through Amazon Bedrock in every region Bedrock serves, with the same IAM, logging, monitoring and cross-region inference you already use for everything else. The model carries a 500K context window and configurable reasoning effort from low through to xhigh, which makes it a realistic candidate for long agent runs rather than one-shot chat.
If you are already on AWS, the cost of trialling Grok against Claude or Nova has dropped to roughly a model-ID change and a fresh eval run. It is worth benchmarking on your own agentic workload before assuming the published leaderboard numbers carry over.
SkillEvaluator puts hard numbers on whether agent skills actually work
Everybody is shipping skills for coding agents at the moment, and almost nobody has published evidence that they help. NVIDIA has, releasing SkillEvaluator as an open-source tool and publishing the first benchmark results across the more than 300 verified skills it put into Cursor. Running the same task with and without a skill installed, correctness rose from 46 to 87 and effectiveness from 39 to 78, an average lift of 39 points once security is excluded.
The finding that should change how you write skills is that the product domain matters far more than the harness, with per-product lift ranging from roughly +2 to +46, and that token savings are not automatic: one skill cut token use by 77 per cent while another pushed it up by 120 per cent. If you maintain skills for Claude Code, Codex or Cursor, this gives you a repeatable way to prove yours earns its place in the context window.
CS-4 pushes inference throughput without a new generation of silicon
Inference speed is quietly becoming the thing that decides whether an agent feels usable or irritating, and Cerebras has raised the ceiling again with the CS-4 it launched yesterday. The system packs three Wafer Scale Engine 3 Turbo processors into a single rack for around 750 petaflops of compute, 7.2 terabits per second of I/O, and more than 4,400 tokens per second per user on a 120-billion-parameter model.
The detail worth noticing is that the silicon is not new: it is the same 5nm wafer as the CS-3, clocked at roughly twice the speed, which says something about where the headroom in wafer-scale architecture actually sits. First systems ship in the third quarter. For anyone building voice agents, coding assistants or anything interactive, non-GPU inference is becoming a serious option rather than a curiosity.
Copilot's CoSnitch data-exfiltration flaw is finally fully patched
Microsoft has completed the fix for CVE-2026-24301, nicknamed CoSnitch, a one-click flaw in personal Copilot that let an attacker pull data out of a victim's connected accounts through a single malicious link. Varonis reported it on 31 December, so full remediation has taken almost eight months, and it is the third Copilot exfiltration bug the same firm has disclosed this year.
If you have Copilot wired into mail, calendar or file storage, this is a prompt to audit which connectors are enabled and for whom. It also sits alongside the Copilot Autofix vulnerability disclosed earlier in the week as a reminder that the risk from an assistant grows with everything you connect it to.
Industry themes
The competition has moved off raw model quality and onto where a model is allowed to run. OpenAI defending Zero Data Retention and Grok 4.6 arriving on Bedrock are both distribution and trust stories rather than capability stories, and both surfaced on the companies' own X accounts on the day. For a cautious enterprise, a data-retention promise and a familiar procurement surface now matter more than a few points on a benchmark.
Agent tooling has entered its measurement phase. NVIDIA publishing controlled with-skill and without-skill results, instead of asking developers to take improvements on faith, raises the bar for everyone else shipping skills and plugins. Expect claims about agent add-ons to start needing evidence attached.
The bottleneck people are actually spending against is inference throughput, which is why Cerebras can command attention for a system built on last-generation silicon simply clocked harder. If you are choosing tools this week, the useful question is less which model is cleverest and more which one you can run under your own governance at a speed your users will tolerate.