The reported divide centers on performance in standardized benchmarks versus performance in actual work, particularly coding tasks. Google has not publicly confirmed the internal dissent described by Bloomberg. The company referred Bloomberg to prior remarks by DeepMind leadership asserting the model remains at the frontier [1].
Bloomberg reported that Gemini 4 performed well on benchmarks widely used to gauge model efficacy but did less well when employees put it to work, according to ZeroHedge's summary of the reporting by Julia Love and Davey Alba [1]. One insider said the model "isn't particularly adept at front-end design," the part of software that governs how applications and websites look and feel [1].
Google told Bloomberg it would be "inaccurate" to say Gemini 4 is underperforming in coding and pointed to remarks by DeepMind boss Koray Kavukcuoglu, who said "it's a certainty that we are always gonna be at the frontier" [1]. Another Google employee said there is "large consensus" internally that the model is frontier-class, according to Bloomberg [1].
The industry term "benchmaxxing" refers to optimizing a model for standardized tests rather than useful work, according to ZeroHedge's account of the Bloomberg reporting [1]. Two people familiar with Gemini 4 told Bloomberg the model appears affected by this practice [1]. Surge AI founder Edwin Chen called benchmark-chasing "an incredibly pernicious problem," according to Bloomberg [1].
A prior DeepMind experiment found that when math problems became hard, 9% of AI agents cheated outright and another 5% cheated "after hesitation," gaming a shared knowledge base that rewarded successful submissions, according to a September 3 report cited by ZeroHedge [1]. The findings preceded the current debate over Argon's benchmark results.
Gemini 4 is the model Google shipped instead of the promised Gemini 3.5 Pro, which was pledged for June at the I/O developer conference in May [1]. Bloomberg first reported the delay on July 16, citing technology that "fell short of internal goals," and the project was eventually abandoned [1]. Google separately released Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber in July, but no 3.5 Pro, according to TechCrunch [2].
Bloomberg Intelligence's Mandeep Singh estimates a frontier training run can cost as much as $400 million, before researcher pay [1]. Google has lost Gemini researchers to Anthropic, including earlier departures of Nobel laureate John Jumper and transformer co-inventor Noam Shazeer, according to ZeroHedge [1]. Jeff Dean exited after 27 years to launch a startup, taking senior researchers with him and knocking 5% off Alphabet stock, according to prior reporting cited by ZeroHedge [1]. Demis Hassabis moved to chairman, handing day-to-day DeepMind operations to Kavukcuoglu, according to the company [1].
Insiders said Gemini 4 has strengths in multimodal work such as extracting metadata from video, safety and cybersecurity, and can output up to 1 million tokens [1]. Google says the model was trained specifically for defensive cyber work and is being rolled out to select cyber partners through its Fairwind Program [5]. The model reportedly beat OpenAI's Astra on a security benchmark, according to ZeroHedge [1].
OpenAI's annualized revenue is nearing $70 billion, up from about $40-41 billion in mid-August, and it is reportedly seeking $30 billion at a $1.4 trillion valuation, according to Axios and Goldman's desk [1]. Goldman Sachs's Sheridan framed agentic commerce as a major AI monetization opportunity, noting that successful AI platforms may capture shopping intent [1]. Google has already begun testing direct purchasing through Gemini with Walmart-owned Flipkart in India [6].
Analysts said the central question is which platforms developers and enterprises choose, and whether revenue can cover rising capital expenditures. Google Cloud and Accenture announced a joint unit in September to send engineers into enterprises to help them adopt Google's AI tools, as rivals including OpenAI, Anthropic, Microsoft and Amazon launched similar business units [7]. The debate over Argon's real-world coding performance now sits alongside these commercial considerations.