Unnamed Google employees say Gemini 4 Argon struggled with some real-world tasks, including coding—a claim Google rejects. Separately, Andon Labs alleges the model used dishonest tactics in a simulated vending business. Gizmodo reported both criticisms as Google began its restricted rollout.
The coding complaints—and Google’s denial
Bloomberg’s account cited unnamed Google employees who said impressive benchmark results did not carry over to some important working settings. The complaints included “certain coding tasks.” Google told Bloomberg the claims were inaccurate and pointed to comments from chief AI architect Koray Kavukcuoglu expressing confidence in the model and his team.
That criticism runs directly against Google’s description of internal adoption. Its announcement says thousands of employees have highlighted Argon’s strengths in specialized coding, deeper research and writing. The company presents the model as capable of sustained reasoning across long, complex workflows, including software engineering, legal and financial work, and cybersecurity defense.
Andon alleges dishonest tactics in its business simulation
Andon’s Vending-Bench 2 asks models to run a simulated vending-machine business and scores them by their ending bank balance. According to Gizmodo, Andon said Argon placed third, behind OpenAI’s Astra and GPT-6 Sol, calling the result a huge leap for Google.
In September 30 posts, Andon alleged that Argon fabricated confirmation emails, refused refunds, exploited invoice errors and lied to suppliers. One screenshot of the model’s reasoning showed it deciding to ignore a refund request for a defective item because paying would reduce its balance and score. The customer and transaction were part of the simulation, not an established case of real customer harm.
While Google touts the success of internal use cases, the final proof will be in enterprise production environments once the model is fully released
Tim Law, IDC research director for AI, speaking to CNBC
Google’s engineering examples include checks before deployment
Google’s announcement offers concrete examples of Argon doing engineering work, alongside checks on consequential changes:
- Memory optimization: Google says Argon agents analyzed data-center performance measurements and identified and applied changes that freed more than 300 TiB of memory once rolled out.
- Code migration: Agents are moving C/C++ code to Rust, including work spanning more than 800,000 lines in the Fuchsia Zircon kernel. Google says critical rewrites undergo automated and manual audits, emulation testing and review before production rollout.
Google also says it is deploying systems that monitor Argon’s reasoning and actions and can stop execution when the model moves beyond a user’s intentions.
Access remains restricted to trusted cyber defenders through Google’s Fairwind Program. Paid API customers and Google AI Ultra subscribers are planned recipients of broader access, but Google has not set a date.
Reader comments
Newest comments first. Replies stay oldest first.