Google demonstrates autonomous LLM fine-tuning loops on TPUs
Google Developers Blog details autofinetune, an agent-driven system that automates hyperparameter tuning for SFT and RL post-training on Cloud TPUs.
In a post on the Google Developers Blog in September 2026, engineers described a new approach to large language model optimization called autofinetune. This system uses autonomous AI agents to manage the repetitive cycle of supervised fine-tuning and reinforcement learning, running experiments on Google’s Cloud TPUs without constant human supervision.
What happened
Traditional post-training workflows require developers to manually hypothesize changes, edit scripts, launch jobs, and monitor results. This process is slow and prone to human error. The new autofinetune project automates this entire loop by leveraging Google’s full AI stack, including Tunix, Gemma models, and Cloud TPUs, orchestrated via the Antigravity CLI and Gemini Flash 3.7. The goal is to allow engineers to define high-level constraints and let an agent discover optimal configurations overnight.
The team demonstrated this capability through two distinct case studies. The first focused on Supervised Fine-Tuning (SFT) using the google/functiongemma-270m-it model on the google/mobile-actions dataset. Running on a single Cloud TPU v5e, the agent executed 20 automated experiments in just a couple of hours. It adjusted variables such as LoRA rank, learning rates, and optimizers while keeping the dataset and architecture fixed. The agent successfully improved accuracy in function call generation by iteratively refining these parameters.
The second case study tackled Reinforcement Learning via Group Relative Policy Optimization (GRPO), a more complex and unstable training method. Using Gemma 3 1B on the GSM8K math reasoning dataset, the agent ran 40 experiments over two to three days on a Cloud TPU v6e. Despite the longer iteration times and sensitivity of RL, the agent identified better LoRA configurations and rollout temperatures. This resulted in a roughly 10% improvement in the combined metric of numerical and format accuracy.
How it works
The system operates on a paradigm shift from manual tuning to an autonomous research loop. Humans define the arena by creating a program.md file that sets boundary conditions, evaluation criteria, and constraints. They also provide a clean, self-contained execution script called run.py. The AI agent then takes over, reading the instructions in program.md to modify run.py, launch training jobs, and measure target metrics.

If an experiment yields improvements, the agent commits the changes to Git and logs the results in a results.tsv file. If a change causes regression, the agent reverts it. This closed-loop system allows for continuous hill-climbing toward better performance without manual intervention. The agent explores a defined search space, such as allowed optimizers or batch sizes, while respecting disallowed changes like altering the core model architecture.
Key details
- The project uses Google’s Tunix library, Gemma models, and Cloud TPUs, orchestrated with Antigravity CLI and Gemini Flash 3.7.
- In the SFT case study, the agent ran 20 experiments in a few hours on a Cloud TPU v5e-1, optimizing LoRA ranks and learning rates.
- For the GRPO case study, the agent conducted 40 experiments over 2–3 days on a Cloud TPU v6e-1, improving total reward by approximately 10%.
- Human operators define constraints in
program.md, specifying allowed parameters like batch size and disallowed changes like dataset swaps. - The agent automatically manages version control, committing successful configurations to Git and logging metrics in
results.tsv. - The objective metric for the RL experiment was a simplified sum of numerical accuracy and format accuracy.
Why it matters
For software teams building AI products, the bottleneck often lies not in model architecture but in the tedious process of hyperparameter tuning. Manual experimentation is resource-intensive and difficult to scale. By automating this loop, engineers can reclaim time spent on monitoring loss curves and managing spreadsheets. This shift allows teams to focus on higher-level problem definition and data quality rather than the mechanical aspects of training runs.

Furthermore, autonomous agents can explore configuration spaces more thoroughly than humans might attempt due to fatigue or time constraints. The ability to run dozens of experiments overnight means faster iteration cycles and potentially better-performing models. This is particularly valuable for reinforcement learning, where instability makes manual tuning especially challenging. The integration with existing tools like Git ensures that successful experiments are tracked and reproducible, maintaining engineering rigor.
What you can do
- Review the
autofinetuneGitHub repository to access the code, sample runs, andprogram.mdtemplates provided by the team. - Define clear boundary conditions and evaluation metrics for your own fine-tuning tasks before attempting automation.
- Start with simple supervised fine-tuning experiments to understand how the agent interacts with your training scripts.
- Use the Tunix library and Cloud TPUs to replicate the hardware environment described in the case studies.
- Monitor the
results.tsvlogs to understand the agent’s decision-making process and identify patterns in successful configurations. - Experiment with different constraint sets in
program.mdto balance exploration speed with computational cost.


