Building a private cloud on bare metal with Kubernetes and Talos Linux
A technical guide details how to construct a self-hosted cloud platform using Kubernetes, Talos Linux, and GitOps tools to manage bare metal infrastructure.
In a post on the Kubernetes Blog in April 2024, Andrei Kvapil from Ænix outlined a method for building a private cloud infrastructure using only open-source technologies. The guide focuses on preparing bare metal servers to host managed Kubernetes clusters, moving away from traditional virtualization layers like OpenStack.
What happened
Kvapil described the process of creating Cozystack, a platform designed to run tenant Kubernetes clusters directly on physical hardware. The approach challenges the common practice of using OpenStack to manage bare metal servers before installing Kubernetes on top. Instead, it leverages Kubernetes itself to handle the complexity of infrastructure management, aiming to reduce the number of complex systems in the ecosystem.
The article serves as the first part of a series detailing this architecture. It covers the initial groundwork required to prepare data centers, including running virtual machines, isolating networks, and setting up fault-tolerant storage. The goal is to provision full-featured Kubernetes clusters that support dynamic volume provisioning, load balancers, and autoscaling without relying on external cloud providers.
How it works
The core distinction lies in how Kubernetes operates in the cloud versus on bare metal. In public clouds, services like persistent volumes, load balancers, and node provisioning are handled externally by the provider. This allows nodes to be treated as ephemeral utilities that can be deleted and recreated easily. On bare metal, these services must run inside the cluster, making updates and maintenance significantly more complex because physical servers cannot be simply deleted and replaced like virtual machines.
To address this, the guide recommends using Talos Linux, a specialized operating system designed for Kubernetes. Talos allows the entire system configuration to be defined in a single file. This declarative approach enables updates to kernel modules and Kubernetes components without requiring full node recreation or service migration. The team at Ænix uses this method to bundle necessary kernel modules, such as ZFS and DRBD, into a custom system image.
For deployment, the process utilizes PXE booting. Temporary DHCP and PXE servers run inside containers to deliver the custom Talos Linux image to physical nodes. A bootstrap script then initializes the nodes and establishes the initial Kubernetes control plane. Once the base cluster is running, GitOps tools like FluxCD are used to install and manage system components, ensuring the cluster maintains its desired state through declarative Helm charts.
Key details
- The guide advocates for replacing OpenStack with a Kubernetes-native approach to reduce ecosystem complexity.
- Talos Linux is used as the base operating system due to its immutable, declarative configuration model.
- Custom system images are built using Docker to include specific kernel modules like ZFS, DRBD, and OpenvSwitch.
- PXE booting is employed to deliver the OS image to bare metal servers during initial provisioning.
- FluxCD is recommended over ArgoCD for managing system components and maintaining cluster uniformity via GitOps.
- The initial bootstrap process can deploy a functional Kubernetes cluster on bare metal in approximately five minutes.
Why it matters
For engineering teams managing their own hardware, this approach offers a path to cloud-like agility without the overhead of maintaining separate virtualization platforms. By treating the infrastructure as code and using immutable operating systems, teams can reduce the risk of configuration drift and simplify the update process. This is particularly relevant for organizations that need strict control over their data and hardware but still want the operational benefits of Kubernetes.
However, this method shifts significant responsibility to the infrastructure team. Unlike managed cloud services, engineers must handle networking, storage, and security patching themselves. The complexity moves from managing multiple disparate systems to deeply understanding the interplay between Kubernetes, the underlying OS, and the physical hardware. This requires a higher level of expertise but can result in a more streamlined and cost-effective infrastructure for large-scale deployments.
What you can do
- Evaluate Talos Linux as a base OS for your bare metal Kubernetes clusters to simplify node management.
- Implement GitOps practices using FluxCD to declaratively manage system components and Helm charts.
- Create custom OS images that include all necessary kernel modules and firmware for your specific hardware.
- Set up a PXE boot environment to automate the initial installation of the operating system on physical servers.
- Design your infrastructure to treat physical nodes as immutable as possible, minimizing in-place updates.
- Start with a small pilot cluster to test the bootstrap process and validate the networking and storage configurations.


