The Agentic Era Arrived This Week - PM Status Report, 25 May 2026
- Google I/O Delivered the Biggest Product Shift in Years
- Anthropic Spent $45 Billion on Compute
- OpenAI Disproved a 75-Year-Old Conjecture and Pushed Codex Onto Your Phone
- Open-Source Just Got a Frontier-Class Option
- What This Means for Project Environments
The biggest week in AI this year just wrapped up, and it wasn’t about one announcement. It was about a dozen of them landing at once. Google I/O. Anthropic’s $45 billion compute deal. A coding agent that disproved a 75-year-old maths conjecture. Open-source frontier models that run on two GPUs.
For practitioners thinking about AI project management tools in 2026, this week’s releases aren’t abstract signals - they’re capability upgrades arriving in tools your teams are already using.
Google I/O Delivered the Biggest Product Shift in Years
The Google I/O keynote on 19 May was the clearest single statement yet of where AI is heading: out of the chat window and into the workflow.
Gemini 3.5 Flash is now the default model across Google’s entire product surface - Search, Workspace, the Gemini app, and AI Studio. It runs roughly four times faster than comparable frontier models, with benchmark scores that beat Google’s previous Pro model on agentic and coding tasks. For anyone using Gemini in Workspace, this is a meaningful step up without any configuration change required.
Gemini Spark is the one to watch for project teams. It’s a 24/7 personal agent that runs on Google’s servers - no laptop open required - with direct access to Gmail, Docs, Sheets, Slides, and third-party apps including HubSpot and QuickBooks. The framing from Google was explicit: this isn’t an assistant you chat with, it’s an agent that works while you’re doing something else. Rolling out to Ultra subscribers first, broader access to follow.
Gemini Omni extends that agentic logic into video: any combination of text, image, audio, and video in - edited, consistent, SynthID-watermarked video out. Most practical near-term use for project teams is probably presentation and stakeholder communication assets, not filmmaking. But the shift from “generate an image” to “edit this footage to match this brief” is significant for anyone producing delivery artefacts with visual content.
Anthropic Spent $45 Billion on Compute
No new flagship model this week from Anthropic - by deliberate choice. What they shipped instead was the infrastructure for everything coming next.
The big number: $1.25 billion per month through May 2029, totalling roughly $45 billion, to access xAI’s Colossus GPU cluster. This is the largest AI compute contract ever disclosed, revealed through SpaceX’s IPO filing on 20 May. The underlying story is that Claude Code demand grew 80× year-on-year in Q1, and Anthropic needed capacity to match. They wrote the cheque.
KPMG also joined the enterprise rollout this week - all 276,000 employees getting Claude access, embedded into KPMG’s Digital Gateway. In the same week, PwC announced its own expanded Claude deployment toward 364,000 professionals globally. Both commitments landed simultaneously, and together they represent the most significant week of enterprise AI adoption announcements the industry has seen. The pattern is now unmistakable: the major professional services firms are standardising on Claude as their enterprise AI platform. If your organisation uses any of these firms, you’ll be working with Claude-powered delivery workflows whether you planned for it or not.
OpenAI Disproved a 75-Year-Old Conjecture and Pushed Codex Onto Your Phone
OpenAI didn’t launch a new model this week either, but it had two announcements worth noting.
The mathematical result is genuinely unusual. An internal reasoning model produced a disproof of Erdős Problem #90 - a conjecture from 1946 about planar unit distances - with a companion paper co-authored by Fields Medallist Timothy Gowers and researchers from Princeton and elsewhere. Whether this signals a new level of frontier reasoning capability or a carefully selected demonstration isn’t clear yet. But Gowers calling it “a milestone in AI mathematics” is a signal worth tracking, particularly for anyone using AI on complex analytical or modelling work.
More immediately practical: Codex is now in the ChatGPT mobile app. You can start a task on your desktop, review progress on your phone mid-commute, approve a decision or redirect the agent, and come back to finished work. That’s async agent supervision on a handset. The pattern - set it running, come back at decision points - is the operating model for agent-augmented project work. If you haven’t built that discipline on lower-stakes tasks yet, mobile Codex is a reasonable place to start.
The content provenance push deserves a mention, too. OpenAI and Google DeepMind coordinated on 19 May to embed SynthID watermarking into both platforms’ image outputs, with a public verification tool at openai.com/verify. This is quiet infrastructure - but if you’re producing any AI-generated content that needs to withstand scrutiny (proposals, client-facing documents, reports), knowing how to use provenance tools is becoming relevant.
Open-Source Just Got a Frontier-Class Option
Cohere released Command A+ on 20 May under a full Apache 2.0 licence. This matters for any organisation that has been held back from frontier-class AI by data residency requirements, compliance constraints, or the risk of vendor lock-in.
Independent benchmarking puts it first on non-hallucination metrics among all assessed models. Native citation generation ties every claim to a source document - relevant for risk registers, audit trails, and any PM output where provenance matters.
The Cohere CEO framed it plainly: “organisations can own their intelligence layer instead of renting it.” For teams in regulated industries - financial services, healthcare, government - or in regions with strict data sovereignty requirements, Command A+ is the first serious open-source option that doesn’t require a compromise on capability.
What This Means for Project Environments
Three things to carry into your planning conversations this week.
The agent supervision model is now standard, not experimental. Auto Mode in Claude Code, Gemini Spark, Codex on mobile - every major lab shipped an async supervision pattern this week. For project teams, the question is no longer whether to use agents on delivery tasks. It’s how to define the checkpoints where human judgement is required. Drafting a status pack, scanning a risk register for threshold breaches, triaging change requests - all of these can run unattended if the scope is defined clearly enough. The PM skill is checkpoint design, not prompt writing.
The professional services firms are getting ahead of you. KPMG and PwC are rolling out Claude to a combined 640,000 people. If your organisation has advisory or consulting relationships, the AI tooling is arriving through those relationships whether or not your internal programme has caught up.
Sovereign AI is now a viable option. Cohere’s Command A+ means that “we can’t send data to a third-party API” is no longer the end of the conversation. A frontier-class model that runs inside your own infrastructure, under your own licence, with native citation grounding - that’s a credible option for regulated environments. If you’re leading projects in sectors where this has been the blocker, the capability case just changed.
Frequently Asked Questions
What is Gemini Spark and is it useful for project managers? Gemini Spark is Google’s new 24/7 personal agent that runs on Google’s servers and integrates directly with Gmail, Docs, Sheets, and Slides. For project managers, the practical use case is delegating recurring tasks - status summary drafts, calendar-to-plan reconciliation, action item tracking - to an agent that works in the background without a laptop open. It’s rolling out to Google AI Ultra subscribers first.
What are the best AI project management tools available in May 2026? The landscape has split into three tiers: cloud-hosted agents (Gemini Spark, Claude Managed Agents, Codex) for teams comfortable with third-party data processing; Microsoft 365 Copilot with Claude now on by default for enterprises already on Microsoft; and self-hosted open-weight options (Cohere Command A+) for regulated or data-sovereign environments. The right choice depends on your data requirements, your existing toolchain, and whether your use cases need frontier reasoning or reliable structured output.
What does Anthropic’s $45 billion compute deal mean for Claude users? It means Claude’s capacity constraints - which caused rate-limit issues for many teams earlier in 2026 - are being addressed at a significant scale. Anthropic raised rate limits for Claude Code and Opus API in May, and the SpaceX capacity deal extends through 2029. If you scaled back Claude usage due to capacity, it’s worth reassessing your limits and re-establishing your workflow baselines.
Should my project team care about the open-source AI model Cohere released? If your team operates in a regulated industry or has data residency requirements that prevent using cloud-based AI APIs, yes. Cohere’s Command A+ is the first Apache 2.0-licensed frontier-class model from a Western lab, designed to run on a single enterprise GPU. It scores first on non-hallucination benchmarks - relevant for any PM output that needs to be accurate and traceable.
The week that just closed didn’t change what project managers need to be good at. It changed how quickly the tooling around those skills is evolving, and how narrow the window is between “early adopter” and “catching up.”
Which of this week’s releases is most relevant to a project you’re running right now - and what would it take to actually put it to work?
Yes, AI helped me to write this :)