Google is at War with Itself
Inside Google’s Dramatic Exits, Gemini Delays and More
It takes time to create work that’s clear, independent, and genuinely useful. If you’ve found value in this newsletter, consider becoming a paid subscriber. It helps me dive deeper into research, reach more people, stay free from ads/hidden agendas, and supports my crippling chocolate milk addiction. We run on a “pay what you can” model—so if you believe in the mission, there’s likely a plan that fits (over here).
Every subscription helps me stay independent, avoid clickbait, and focus on depth over noise, and I deeply appreciate everyone who chooses to support our cult.
PS – Supporting this work doesn’t have to come out of your pocket. If you read this as part of your professional development, you can use this email template to request reimbursement for your subscription.
Every month, the Chocolate Milk Cult reaches over a million Builders, Investors, Policy Makers, Leaders, and more. If you’d like to meet other members of our community, please fill out this contact form here (I will never sell your data nor will I make intros w/o your explicit permission)- https://forms.gle/Pi1pGLuS1FmzXoLr6
A lot of people have been caught off guard by Google’s sudden string of high-profile exits, with Google’s stock reacting very negatively to this news —
On paper, Google should be the runaway winner of the agentic AI race. They have the complete stack, reputation, talent, and the cash to fight costly wars of attrition against Anthropic and OpenAI (their 2 biggest competitors). They have a strong research culture that spawns some of the most promising ideas in all of tech.
So…why are the Lionhearts at Google in a fetal position, while all the people that don’t match up as well are wailing on them (if you get this reference, we should be buddies)? And did this sudden collapse really come out of nowhere?
If you’re a regular reader at the Chocolate Milk Cult, the answer to the 2nd question is a clear no: we’ve been talking both about the deeper structural issues at DeepMind/Google and about how you should expect some very high profile exits because of that. For instance, take our livestream back in February, when Gemini 3.1 Pro dropped, and everyone was staring at the benchmark numbers — 77.1% on ARC-AGI-2, #1 on Artificial Analysis. I was less excited about the whole thing —
Six months later, here are some news items that are worth paying attention to:
The flagship model is stalled. Gemini 3.5 Pro missed its target window, with internal reports pointing to persistent failures in complex coding workflows.
Redundant teams are building overlapping tools. DeepMind, Google Cloud, and Android spent the last year building competing AI coding assistants. Gemini CLI and Code Assist are now being forcibly merged into Antigravity just to stop the internal cannibalization.
Internal friction is leaking. DeepMind and Google Cloud’s have had a mean tug of war, with executive battles even playing out in the tech press.
Key talent is walking. Senior researchers who built the core Transformer and Gemini foundations are leaving for lean, single-product AI startups. What about Google seems to alienate their high agency builders (this pattern is much deeper than just the celebrities).
Leadership had to step in. Google just restructured DeepMind’s operating leadership with a very high-profile organizational reshuffle.
All of these stories point to the same underlying cause: Google as an institution is being torn apart from the inside by multiple competing philosophies on what it should be. This means that while individual teams and engineers might do great work, it’s hard for Google to move towards a meaningful long-term goal due to a lack of a deeper underlying identity.
This article will be one of the most honest and in-depth looks at Google and where things keep going wrong for a company that should have everything going for it. While I do have access to a lot of insider comments from high-level Google execs in my network (it’s how I was able to tell y’all that Demis’s faction at DeepMind wasn’t as “clean” as some engineers liked to imagine and that they were trying to do some hardcore lobbying to position Demis as the next Google CEO), I’ll keep most of the article focused on easily verifiable public facts and reports to avoid getting into an unproductive “he said, she said”.
Executive Highlights (tl;dr of the article)
DeepMind, Cloud, Android, and Workspace have spent years building toward different goals, often duplicating tools and fighting over product ownership and TPU access.
Google repeatedly fails to compound its own advantages. It buried multimodal embeddings, kept TPUs unnecessarily difficult to adopt, fragmented its coding agents, and fired the engineer behind a viral Workspace CLI while another team rebuilt the same thing.
The same incentives distort Gemini. Benchmarks offer research teams autonomy and easy internal recognition; reliable agents require messy cooperation across models, tools, permissions, infrastructure, and products. So Google keeps producing intelligence that looks better on charts than it feels in actual workflows.
Those incentives also drive out high-agency builders. Google creates extraordinary researchers, frustrates them when they try to ship, then watches them build competitors — or pays billions to bring them back.
The recent consolidation around Antigravity and Koray Kavukcuoglu suggests Google has decided to make Google Cloud it’s primary Champion. That could work. The real test will be whether this centralization of power enables deeper alignment, or if it will turn Google’s brightest into automatons.
I put a lot of work into writing this newsletter. To do so, I rely on you for support. If a few more people choose to become paid subscribers, the Chocolate Milk Cult can continue to provide high-quality and accessible education and opportunities to anyone who needs it. If you think this mission is worth contributing to, please consider a premium subscription. You can do so for less than the cost of a Netflix Subscription (pay what you want here).
Want access to a repository containing all of our research? 300+ files containing our notes of various experiments, discussions with cutting-edge teams, and insights into where the industry is headed next. Get a Founding Member Subscription to AI Made Simple. Want to talk to me for details/get my insights into the tech ecosystem? Reach out to me through any of my socials over here or reply to this email.
Every month, the Chocolate Milk Cult reaches over a million Builders, Investors, Policy Makers, Leaders, and more. If you’d like to meet other members of our community, please fill out this contact form here (I will never sell your data nor will I make intros w/o your explicit permission)- https://forms.gle/Pi1pGLuS1FmzXoLr6
Google Keeps Building Overlapping Products Instead of Compounding Value
In business, moats are anything that deters new competitors from entering your business stream. For example, if I wanted to start a soft-drink competitor to Coke, I would have to set up my own distribution partnerships and find my own set of poor people to steal water from. Moats provide incumbents with a lot of stability and time to react to any competition; consequently, investors and product managers love moats.
As an organization, Google has a very unique skill of ruining its own work in building these moats (outside of its standard search and ads space). The big G loves to turn what should be strong iterative loops of development, testing, and improvements into Jackson Pollock-esque clusterfucks.
Our first case study on this matter will be the coding agent mess.
Consider the following — Both OpenAI and the big G launched their own CLI products to try and capture some of Claude Code’s smash-hit success. Both Codex and Gemini CLI were vastly inferior products to CC. But they ended up with very divergent paths.
Today, Codex has bridged the gap against Claude Code and has had viral growth of its own. Additionally, the same core Codex engine follows a developer everywhere: inside the CLI, embedded in the IDE, in the web interface, and across cloud environments. Every fix, every state-tracking enhancement, and every context-window optimization compounds into the exact same platform layer.
Google ran the opposite play. Take a look at this passage from here —
“Google co-founder Sergey Brin and others were advocating for Google to move faster to seize opportunities in AI coding, but their efforts were slowed by competing factions within the company, two former employees said. Cloud computing unit Google Cloud, research lab Google DeepMind and the team behind the Android operating system are all building AI coding tools for developers, with involvement from some consumer product teams, too, people familiar with the work said.”
To stop the internal friction, Google is now consolidating those parallel workstreams into Antigravity so that state-tracking improvements, tool-use optimizations, and context management developed by one team can finally flow across all interfaces.
Which… yes. Obviously you would want that. It’s why we teach first-year engineers that software is a team sport. But it somehow took Google months to get there. (To be clear, at their scale, it’s understandable that you had multiple teams start the agent buildout simultaneously; what’s not acceptable is that they took this long to converge).
This reveals a fundamental difference between Google and leaner AI labs. OpenAI and Anthropic make bad product decisions, but when an architectural pattern works, they collapse organizational resources behind it. Google lets competing implementations survive, duplicate, and fragment.
Sometimes that fragmentation produces redundant product SKUs. Other times it produces structurally bizarrely self-destructive edge cases.
Google fired him.
Internal reports indicated that while individual engineering teams wanted to adopt his tool, other divisions raised concerns about external branding and unapproved release channels. Two days prior to his termination, Google publicly announced an official Workspace CLI was coming anyway.
Compare this response to OpenAI’s when they heard about JP getting fired —
A simple appreciation of high agency and an invitation to talk further (from the creator of OpenClaw, someone who was himself scouted by OpenAI).
So, a builder inside the company identifies a massive friction point in agent execution, builds a working primitive, and validates user demand in production. Instead of absorbing the tool, standardizing its security boundaries, and scaling it, Google’s corporate structure treated the initiative as a governance breach — firing the engineer while a separate team worked to replicate the exact same feature set from scratch.
That, my friends, is a sickness. A sickness of apathy and decay where the company fails to recognize and work on the exceptional work being done internally.
Cue story number 3: multimodal embeddings.
Google had a multimodal embedding model in Vertex AI that could embed text, images, and video into the same semantic space. Video embeddings were generally available by early 2024. You could actually take video, generate embeddings from it, and do things like semantic video retrieval.
That was not a typo. Google did functional video embeddings early 2024. These embeddings weren’t great (they were a nightmare to use since they were gated behind Vertex and couldn’t do audio), but this should have been a big deal regardless. The cutting-edge models at this time were Gemini 1.5, GPT -4 Turbo (do you even remember this one), and Claude 3. Some additional grounding — GPT-4o wasn’t even released then.
So why did almost no one talk about this? Simply put, it was hard to get to. The feature sat inside Vertex AI without clear integration into Google’s primary developer frameworks, developer marketing, or higher-level SDKs. Their dev-rel guys wrote some technical documentation, published sample notebooks, and then left it buried inside Cloud’s platform menu as a feature developers had to stumble across by accident. And actually using it was even worse, since you had to jump through 5 hoops and gluck gluck your nearest Google Cloud support rep b/c of how hard setting everything up properly was.
The result? Most people didn’t even know this existed, Google lost a LOT of money as other embedding models took center stage for semantic retrieval systems, and they had to pay a 100 M to license tech from the overpriced and overrated Contextual AI just to try to have a presence in this space(which has so far also failed).
It really didn’t have to be this way. With some marketing, easier accessibility, and an overall better rollout + integration into their ecosystem, they could have made a lot of waves in both the direct embeddings space and the indirect follow-on products from it (OCR, bulk document/video ingestion, or bundling it up with their knowledge graphs services and some Gen AI for selling a serious file QA contender (which is a market being eaten by Codex and Claude Cowork rn)).
This isn’t easy, and it could have still failed. But at least they would have failed attempting to make progress towards capturing a valuable slice of the world. Not failed because they were thrashing around with no overarching goal.
In all cases, the pattern is eerie: Google tends to overlook promising work being done in their own org. Sometimes it’s because it’s done by a little team that doesn’t know how to get attention; sometimes it’s because the work done with whatever is aligned with Google’s agenda (which is itself often dictated by competitors), and thus it suffers from attentional blindness. Sometimes it’s just because Google genuinely does a lot of good work, so people don’t know where to focus, and they tend to default to whatever is the default it doesn’t help that their team is a weird mix b/w PhDs who don’t study the industry and think about the ecosystems and MBAs who don’t think about much at all).
In all cases, it leads to very promising ideas being overlooked for eons before someone eventually gets that market and Google realizes that they have to play catch-up there. To really drill this pattern home, here is the most egregious example of this: TPUs.
Across the industry, engineering teams evaluate AI infrastructure almost exclusively in terms of NVIDIA hardware, often treating Google’s Custom Silicon (TPU v5p and Trillium) as an internal Google detail rather than a viable cloud hardware platform. That market perception stems from how Google organized hardware access internally. For years, Google Cloud did not fully control the supply of one of its most important products. The group selling TPUs sat inside Google’s core engineering organization, while most chips were reserved for internal use. In 2022, Thomas Kurian successfully lobbied to move that group into Cloud. Until then, Cloud reportedly needed approval from another part of Google before it could offer TPU capacity to customers.
This extremely delayed reorg had some pretty disastrous consequences for Google’s market position. TPUs were originally engineered to run on Google’s proprietary frameworks (JAX and XLA) while the broader machine learning community standardized on PyTorch. Because TPUs were largely internal, Google delayed deep, native integration between PyTorch and the TPU compiler stack.
Why does any of this matter? Because due to this delay, developers chose to pay NVIDIA’s hardware premium rather than rewrite their model code for JAX during the biggest chip demand boom ever.
Take a second to think about how stupid this is. Google built TPUs more than a decade ago b/c they realized that all current hardware was inefficient for Neural Networks. They are one of the people who helped make Gen AI mainstream. And at no point did any tubelight think of making their TPUs easily accessible to customers. Their teams were too busy patting themselves on the back by writing about their efficiency gains in the appendices of their papers. That is a generational, Lukaku first-touch level of fumble, and it reflects in the business growth of Nvidia and Google Cloud (keep in mind the image below is Nvidia Data Center vs ALL of Google Cloud)—
Caption: Google began offering Cloud TPUs externally in early 2018. Yet NVIDIA’s Data Center business produced $75.2 billion in its latest quarter, compared with $24.8 billion for the entirety of Google Cloud — not merely TPUs, but GCP, Workspace, and other enterprise products. Google TPU history, NVIDIA results, Alphabet results.
By late 2025, Google was forced to allocate heavy engineering resources to TorchTPU specifically to fix the software friction that prevented customers from adopting their hardware. This will do wonders for their hardware adoption
Even with structural adjustments, those internal friction points persist. Google Cloud executives continue to negotiate directly with DeepMind leadership over the allocation of constrained TPU clusters. The hardware supply sits inside the same company, yet different business units negotiate internal transfer prices and compute access as if they were third-party vendors. There is also a very large bloc of researchers who push for TPUs to stay in-house so that they can use them for experiments. This creates a lot of unnecessary conflict.
This conflict is the driver behind Google’s recent rupture. And in case you haven’t caught it, it contains Google’s two shining aspirations —
The Messy Relationship Between Google Cloud and DeepMind
The TPU fight points to a larger problem: DeepMind and Google Cloud aren’t trying to win the same game. DeepMind wants Gemini to be the undisputed best model on the planet. Google Cloud wants Google Cloud to win the enterprise market.
Those two goals have a strong overlap, but they aren’t strictly the same thing. Cloud doesn’t need Gemini to top every leaderboard, and Kurian doesn’t even strictly need his enterprise customers to use Gemini at all. If a Fortune 500 bank wants Anthropic’s Claude, Cloud will happily sell them Claude through Vertex AI, host the workload on Google hardware, plug it into BigQuery, Cloud Spanner, and Google IAM, and collect the margin on the entire surrounding platform.
That last sentence might sound like Heresy, but I’m citing Kurian directly here. Cloud aggressively positions itself as a neutral multi-model platform and has hired senior leadership directly from Anthropic to drive that push.
As you might imagine, this strategy doesn’t go over well with some of the more purist elements at DeepMind. But while this would be problematic enough as is, there is another deeper rupture that was causing an internal arms race. It isn’t as bombastic, so it’s been largely overlooked by the media/discussions, but imo it’s far more interesting since it hints at the future of model development and research. To see how, let’s go back to Gemini<>Cloud dynamics.
Theoretically, while Gemini isn’t needed to make Cloud a lot of money, it can still be exceptionally useful. Running an in-house model like Gemini on proprietary TPUs eliminates the third-party margin tax, lets Cloud undercut competitors on inference pricing, and captures economics that OpenAI or Anthropic can’t touch. In this world, the relationship b/w Google Cloud and DeepMind should still be tight. So why isn’t it?
It’s because for this vision to work, Google doesn’t need Gemini to be the best. Or even the 3rd best. It just needs it to be good enough at a really good price. Essentially, it needs to turn Gemini’s training focus from intelligence to intelligence per dollar.
If you found a way to have Gemini give up 5% on an abstract reasoning benchmark but in return it connects cleanly to BigQuery, handles enterprise permissions without hallucinating access tokens, and executes deterministic API calls across 20 turns, Kurian will build you a temple in the GCP HQ.
DeepMind, however, will drive up the prices of voodoo dolls out of their passion for you. Their focus and internal culture optimize for frontier capabilities, either by topping benchmarks or in doing cool and next-gen research. By itself, this isn’t a bad thing, but it creates an extreme problem when this culture (as it often is with research-heavy orgs) becomes hostile to the plumbing/boring engineering work needed to make the model boringly reliable within an org. At that point, there are only two ways forward:
You go all in on research, à la peak Bell Labs. Google has the free cash flow, core business moat, and the talent density (also diversity; once you look beyond the standard publications in the LLM research space, they actually have some insane projects floating).
You turn the research into a way to improve your product/feed your business (Amazon does this very well). Here Cloud would function as DeepMind’s massive real-world testing ground. When Gemini mangles a SQL query inside BigQuery or fails an enterprise IAM boundary, that failure telemetry should flow straight back to London to guide the next post-training checkpoint.
Both these directions can work. But both require commitment, and a clearly defined order of priorities. There can only be one king at a time, and this wasn’t a decision that Google was willing to make so far. So, the two divisions have spent years in bureaucratic turf wars over developer tooling, product boundaries, and TPU allocations. The Information documented how AI Studio was bounced between Cloud and DeepMind. Reuters reported that Cloud executives openly welcomed Koray Kavukcuoglu taking operational control of DeepMind because he is commercially minded and actually focuses on product execution.
From that lens, the leadership shakeup is a clear signal that Google has seen DeepMind’s struggle to really establish a dominant lead, factored in Cloud’s 82% quarterly + 390% backlog growth, and decided that Google Cloud is the right horse to back.
Hopefully, this consolidation addresses one of Google’s biggest problems: the company has spent much of the AI boom optimizing for the kind of intelligence that photographs well but disappoints in person. Why is that putting aside Gemini 2.5 Pro, Gemini has failed to capture the attention of the larger community in the same way that Claude and OpenAI have?
Why Does Google Keep Optimizing for the Wrong Kind of Intelligence?
To be clear, Gemini isn’t a bad model. It routinely matches or beats OpenAI and Anthropic on math, abstract reasoning, and multimodal benchmarks. Every few months Google drops a checkpoint alongside a small rainforest of charts showing that the line went up again. Think back to 3.1 Pro’s release, and how it was meant to be a slam-dunk against the rest of the industry. The numbers were compelling, influencers went wild, and the entire ecosystem stood on attention. Then people used it for a sec, and pretty much everyone went back to Claude Code or Codex to get actual work done.
Why would this happen? And why have Google’s releases failed to inspire the same level of awe as the others, even when the benchmarks tell us that we should love the models?
If you’re familiar with the industry, you’d know that the answer is benchmark-maxxing (we even broke it down extensively here). The more interesting question is why does Google seem to be particularly badly hit by benchmark-maxxing.
In an org like Google, benchmarks are the ultimate political currency. A research team can train a model, run an eval script, post the score on X, all w/o needing to navigate internal politics and negotiate access with the mail/drive teams. In other words, optimizing around a benchmark provides two holy grail outcomes in corporate environments —
Autonomy: the team can work by themselves with no blockers.
Legibility: middle management can point to the benchmark wins and thus continue to justify their existence. If they have a particularly good result, then they even get visibility into the ecosystem.
Almost by definition, this is anti-thetical to what agentic systems need. Models that thrive in agentic loop have to be able to coordinate different tool schemas/outputs, have to be extremely reliable so that you they don’t produce wildly different outputs due to slight variations in outputs, and has to preserve complex workflows/state across 40 steps. In case of Google, this work cuts across Cloud, Workspace, security, infrastructure, and product teams. Since one teams failures can reflect badly on everyone (and credit is also shared, with the model team the highest profile receipient), loss aversion keeps teams from coordinating properly.
(You don’t have to take my word on this, their own customers have complained about this experince)
Sources: Google Drive Help (2024); Gemini Help (2026); Google Workspace user report (2025); BigQuery user report (2024); Google’s BigQuery warning (2026); Google Cloud command failures (2024); “Gemini Code Assist is a mess(2025)”.
Antigravity and the operational promotion of Koray Kavukcuoglu is likely a push to bring DeepMind to finally yield to commercial pressure and optimize for completed work, not great benchmark scores.
This will be useful, but alongside this, Google needs fix their deeper human incentive alignment issue. Inside a bureaucracy, the safest career move is to stay in your lane, collect cross-functional approvals, and let every VP touch the steering wheel. High-agency builders do the opposite: they see a broken primitive, bypass the committee, and ship.
Think back to Justin Poehnelt (Google engineer that built the viral Workspace CLI and then was fired for it). A less high profile version of his profile plays out a lot more than you’d think. Let’s see some case studies:
All eight authors of the original Transformer paper left to start or lead competitors like Character.AI, Cohere, Sakana AI, Essential AI, Inceptive, and Adept.
Noam Shazeer left in 2021 when Google refused to ship the chatbot he built, founded Character.AI, and came back in 2024 via a $2.7 billion licensing and talent deal to co-lead Gemini. Less than two years later, he left for OpenAI.
AlphaFold co-creator and Nobel laureate John Jumper left DeepMind for Anthropic, followed by senior personnel departures across the core Gemini teams.
Google routinely attracts the people who push the frontiers, frustrates the builders who want to productionize them, and then buys their startups back at massive valuations when the frustrated builders eventually leave. This has two problems:
Obviously, losing the high agency people is frustrating. They leave with ideas, prototypes, and a lot of institutional knowledge. These people will either go to more aggressive competitors (very bad) or start their own startup. In the latter case, their brands allow them to raise at high valuations, so Google will have to buy them back at a very high number.
The people left are the types who are more comfortable with bureaucracy (either the political types or the complacent types). As their numbers grow (and they find themselves in more senior positions), they will deepen the babu-giri culture at the Big G.
It’s no coincidence that Google tends to be the biggest poaching ground for AI Labs (Google is the big tech company most likely to lose it’s seniors to AI Labs).
(This chart doesn’t include DeepMind, which would increase Google’s number even more.)
In other words, Google rewards the same qualities in both models and employees: locally measurable output, procedural compliance, and organizational legibility. This is where Google’s next battle will be fought.
Conclusion: What’s Next For Google.
When Napoleon met the Third Coalition at Austerlitz, he faced a larger army with experienced soldiers and far more artillery. He did not win because every French officer stopped thinking and waited for instructions. His corps could move independently, improvise, and fight on their own — but they were all moving towards the same objective. The coalition’s commanders were coordinating different armies, different priorities, and different chains of command. This meant that the Napoleonic Corps could all take the best actions for themselves individually, all while being aligned on the larger vision (while the Coalition was always paralyzed by multiple chains of command).
That is the real test for Google. Bringing DeepMind, Cloud, and the product teams together only matters if it creates one direction without requiring everyone to think the same way. Too little alignment, and Google returns to the turf wars that got it here; too much control, and its best builders either stop taking risks or leave. If it gives Cloud, DeepMind, and Google’s builders a shared direction while preserving their freedom to act, the company may finally make its extraordinary advantages compound.
In other words, the question is whether Google becomes one army — or merely one larger committee. Personally, I’m optimistic based on my internal conversations with the company.
Thank you for being here, and I hope you have a wonderful day,
Dev <3
If you liked this article and wish to share it, please refer to the following guidelines.
That is it for this piece. I appreciate your time. As always, if you’re interested in working with me or checking out my other work, my links will be at the end of this email/post. And if you found value in this write-up, I would appreciate you sharing it with more people. It is word-of-mouth referrals like yours that help me grow. The best way to share testimonials is to share articles and tag me in your post so I can see/share it.
Reach out to me
Use the links below to check out my other content, learn more about tutoring, reach out to me about projects, or just to say hi.
Small Snippets about Tech, AI and Machine Learning over here
AI Newsletter- https://artificialintelligencemadesimple.substack.com/
My grandma’s favorite Tech Newsletter- https://codinginterviewsmadesimple.substack.com/
My (imaginary) sister’s favorite MLOps Podcast-
Check out my other articles on Medium. :
https://machine-learning-made-simple.medium.com/
My YouTube: https://www.youtube.com/@ChocolateMilkCultLeader/
Reach out to me on LinkedIn. Let’s connect: https://www.linkedin.com/in/devansh-devansh-516004168/
My Instagram: https://www.instagram.com/iseethings404/
My Twitter: https://twitter.com/Machine01776819




















Is that an anthony smith reference I spy????