Wagtail's single-model coding challenge: lessons from 2B tokens
A Wagtail developer attempted to use only GLM 5.3 Flash for a month of coding. Infrastructure limits and prototype costs forced a switch to other models halfway through.
In September 2026, a core developer at Wagtail CMS attempted to restrict all coding tasks to a single efficient open-weight model, GLM 5.3 Flash. The experiment consumed 2 billion tokens but failed to maintain the single-model constraint due to infrastructure bottlenecks and high-cost prototyping.
What happened
The goal was straightforward: spend the entire month using only GLM 5.3 Flash for development work. For the first half of September, the strategy held firm. Usage remained within a tight budget of $68, consuming approximately 4kWh of energy and generating 365 grams of carbon emissions. This initial success demonstrated that lean, efficient models could handle standard engineering tasks without breaking the bank or the environment.
However, the second half of the month saw a significant deviation. Roughly 1 billion tokens were spent on alternative models, splitting the total usage evenly between the target model and others. The developer noted that while the challenge technically failed to meet its strict single-model criteria, the data gathered provided crucial insights into the practical limitations of relying on one specific inference provider or model architecture in a production-like setting.
Several factors contributed to this shift. The team encountered unexpected infrastructure availability issues with their chosen providers. Because GLM 5.3 Flash sits high on the Pareto frontier for their specific workload, it became a popular choice among other users as well. The smaller inference providers lacked the GPU capacity of major labs, leading to performance degradation. To maintain productivity, the developer switched to comparable alternatives like DeepSeek V4.1 Flash and Qwen 3.8 Flash, which were available in European data centers.
How it works
The experiment relied on monitoring tools like AgentsView to track token distribution, cost, and energy consumption across different models. The workflow involved using the AI assistant for various tasks, including UI implementation, documentation writing, and running evaluations on the Wagtail codebase. The selected model, GLM 5.3 Flash, offers a 1 million token context window and vision support, allowing it to process screenshots and handle extended coding sessions.
A significant portion of the resource drain came from "vibe coding" an experimental Model Context Protocol (MCP) server. This approach prioritizes rapid prototyping over optimized code structure. The developer selected a suboptimal model configuration for this prototype, resulting in a sudden spike of 450 million tokens, costing $150 and consuming 5kWh of energy in a single night. While the prototype succeeded functionally, it highlighted how agentic patterns and poor model selection can drastically inflate costs compared to more deliberate engineering approaches.
Key details
- Total consumption: The experiment processed 2 billion tokens in September 2026.
- Budget adherence: The first half stayed within a $68 budget, but the total month exceeded initial efficiency targets.
- Energy impact: Total energy use reached approximately 35 kWh, significantly higher than the projected 10 kWh for a purely efficient workflow.
- Infrastructure limits: Performance degradation on GLM 5.3 Flash forced a switch to DeepSeek V4.1 Flash and Qwen 3.8 Flash.
- Prototype cost: A single MCP server prototype consumed 450 million tokens and $150 overnight due to inefficient model selection.
- Model capabilities: GLM 5.3 Flash was praised for its 1M context window, vision support, and multi-provider availability.
Why it matters
For software teams adopting AI-assisted development, this case study underscores the difference between theoretical efficiency and operational reality. While flash-tier models are cost-effective for routine tasks, they are not immune to supply chain constraints. Relying on a single model or provider creates a single point of failure when demand spikes. Engineers must recognize that "open" models still depend on physical infrastructure, which may be limited compared to proprietary giants.
Furthermore, the distinction between production coding and research and development (R&D) is critical for budgeting. Experimental work, particularly when using agentic workflows or rapid prototyping techniques, consumes resources at a much higher rate. Without separate budgets and monitoring for R&D, these experiments can derail cost-saving initiatives. Teams need to measure success not just in tokens generated, but in energy used and concrete outcomes achieved.
What you can do
- Implement local usage measurement tools to track tokens, energy use, and spend in real time.
- Create separate budgets for day-to-day engineering tasks and experimental R&D projects.
- Evaluate multi-agent techniques, such as separating orchestrator, scout, and reviewer roles, to reduce redundant token usage.
- Maintain a fallback list of compatible models, such as DeepSeek or Qwen variants, to handle infrastructure outages.
- Prioritize models with broad provider availability to leverage market competition and ensure uptime.
- Focus on efficiency metrics like energy per task rather than raw token counts when assessing model performance.


