This blog post is the first part of a series covering the evolution of the network stacks behind Hetzner Cloud and their architecture. In this post, we will look into the development we’ve done until today, and how the current Open vSwitch-based network stack works. Although the network stack is similar for Cloud servers and Load Balancers, we will focus on Cloud servers and virtualization hosts (VM hosts), as they are providing more services and features.
For brevity, we will keep the Cloud Control Plane, which manages all servers, Load Balancers, Private Networks, etc. out of scope, and will only look into the network stack on the hosts.
How we roll
All cloud products require versatile and reliable connectivity options, so that our customers and potentially their customers can reach the services running in our network.
At Hetzner, we generally believe in solutions we own and can operate ourselves, and in having building blocks, where we can further build upon. The network, and especially the network connectivity for our cloud offerings, is no different.
Although technologies changed over the years, we’ve always preferred solutions, where the moving parts are on the hosts, rather than having a smart and complex physical network setup. As a result, we designed the setup in a way, that the physical network is providing stable IP connectivity to the hosts, and the hosts offer higher level network functions. These include Firewalls, Private Networks, and auxiliary infrastructure services.
By running these services and network functions on the hosts, we gain a lot of freedom and flexibility when designing and building our product, and are not limited to what hardware vendors offer. This way, we have a clear separation of concerns between teams and technologies, and can contain the blast radius to a single host, in case of (most) errors. This also means that we can and must fully own the solution and adjust it to our needs, but that’s in our DNA anyway.
A bit of history
Historically, Hetzner has been using different technologies to provide connectivity for its cloud products.
In the early days, the Hetzner vServers launched in 2011 started out with HDDs, native Linux bridges, and static routes. It provided dual-stack connectivity with up to 1 Gb/s per host. In the end, there were approximately 25,000 instances running with this setup.
The follow-up product launched in September 2015 with a hyper-converged Ceph-based setup. It switched to dynamic routing via BGP and a Linux bridge setup using 1:1 NAT for IPv4 and fully routed IPv6 prefixes. This enabled better IP mobility and host maintenance without customer impact. In 2016, the links to the hosts were upgraded to 2x 10 Gb/s. This setup lasted until 2018 and grew to about 50,000 instances.
When Hetzner Cloud launched in 2018, we dropped NAT and moved to direct IPv4 routing. In July 2019, we introduced an Open vSwitch-based data plane, and added support for private cloud networks based on VXLAN encapsulation. In November 2020, we started migrating the public network part over to Open vSwitch, allowing us to launch stateful cloud firewalls based on Open vSwitch flows and netfilter in March 2021.
Ingredients of a VM host
We’ve already established, that important parts of the network stack are running on the VM hosts; but what do they have to provide for a cloud server to operate?
For the majority of servers, Internet reachability is required, so the system and its services can be accessed from home or abroad. Some servers use Private Networks, either instead of, or alongside public network , which connect servers, Load Balancers, and possibly dedicated servers via a private and virtual network that only your systems can access. Private network traffic is encapsulated using Virtual eXtensible LAN (VXLAN) and a unique Virtual Network Identifier (VNI) per private network. The host network stack ensures that only members attached to the respective private network can send or receive traffic.
To control access to your services available from the Internet, we provide the cloud Firewall feature. It allows controlling access to certain ports or protocols of the server, or can limit which traffic can leave the server. This should happen as close to the server as possible, hence on the host. Building stateful firewalls centrally, would require additional hardware and network overlays to transport the “clean” traffic to the servers, and add much complexity for highly available, central connection tracking.
By default, our cloud server images are configured to use the Dynamic Host Configuration Protocol (DHCP) for dynamic configuration of IPv4 addresses and routes. Hence, we have to provide a DHCP server, so servers can request their IPv4 networking configuration for public and private network interfaces.
Within most images, cloud-init is used to configure some parts of the system, for example configure IPv6 networking, trigger DHCP for private network interfaces, attach cloud Volumes, etc. For these tasks, cloud-init has to query information about the local system or user-specific configuration bits from our metadata server available at http://169.254.169.254/. Besides cloud-init, the metadata server is used by many integrations, including but not limited to our CSI driver, hc-utils, distribution tooling like Flatcar Linux's Afterburn and Ignition, or Talos Linux.
Both the DHCP and metadata servers are running locally on each VM host for maximum resiliency and availability.
Besides these user-facing services, the VM hosts run a bunch of internal services, which make sure your servers are up and running, the network stack gets the correct information to configure connectivity, and our monitoring knows what’s happening.
In this post, however, we’ll focus on the network stack and will have a deeper look into its inner workings.
Open vSwitch
Most of the network data path is realized by Open vSwitch, or in short OVS, an Open Source multilayer virtual switch. Its data path is included in the Linux Kernel for a long time, and packages for the user space / control plane are available for all major Linux distributions.
It is the default network stack for some virtualization environments, and many off-the-shelf solutions come with OVS support. This includes but is not limited to Proxmox VE, Open Stack, and oVirt, to name a few popular ones. We however, don’t use any of these, but rather manage KVM and all surrounding components using custom-built tooling.
Open vSwitch is a versatile networking solution, offering a vast amount of features, of which we’re only using a subset for our use cases and setup. It uses OpenFlow, an open protocol built to configure the data plane of network devices and denote how packets should be forwarded. For each direction of communication, there has to be a flow specifying how packets should be handled, and where they need to be sent.
This can be exact matches, e.g. for communication with the gateway via ARP, NDP, ICMP, etc., traffic destined for other members of a private network, or connections to a service on the Internet.
Open vSwitch orchestration - Flusskrebs
To provide network connectivity for our servers and Load Balancers, and to configure any Firewall(s) applied to a server, we have to install the required flows for Open vSwitch to provide these connections and features.
A common solution to orchestrate Open vSwitch would be the Open Virtual Network (OVN) control plane project, which was started alongside OVS itself and is intended for maintaining a fleet of nodes running Open vSwitch. It provides an abstraction layer, where one can configure logical routes and switches, which then get translated into OpenFlow.
At the time when Open vSwitch was introduced into the Hetzner Cloud stack, we decided to not go down this road, and rather built our own custom solution called Flusskrebs (literally translates to “river crab” or “flow crab” 🦀). It is written in Python, provides a REST API used by the host-local part of our Cloud control plane, and configures the flows each VM or Load Balancer needs to work properly. This includes flows to implement any Firewall rules configured by the user or our backend, e.g., for blocking egress SMTP by default.
The controller also acts as the central DHCP server for private network interfaces on each host.
VM host network stack
With the general architecture ideas and these building blocks in mind, let's have a look at our VM host network stack in more detail. The following figure shows a simplified overview of the inner working of the network stack of a VM host.
Most hosts are using two 10 Gb/s uplink interface, which are bundled into a link aggregation group (LAG) or bond on Linux using LACP, connected to two physical switches forming a virtual chassis for increased resilience and bandwidth. To increase resilience even further, we’ve started migrating to two native routed uplinks using BGP, connected to two independent switches earlier this year. Note, that the uplink interfaces are not directly connected to the OVS bridge, but the Linux system routes between the uplinks and OVS. This way, the host reachability is independent of OVS, and it eases host installations and maintenance.
A central Open vSwitch bridge acts as the main component delivering network connectivity and features, and all cloud servers on the host are connected to it. Each server can have one public network interface, and additionally up to three private network interfaces. The following templates are used to render flows allowing egress traffic from the server towards the network:
Ingress traffic towards the server is forwarded using the following flows, one for each IPv4 or IPv6 prefix respectively. The host network stack identifies itself to the VM using a well-known “virtual gateway MAC” which is always d2:74:7f:6e:37:e3.
Infrastructure services, like DHCP and metadata servers are also connected to Open vSwitch and are running alongside each server inside a dedicated Linux Network Namespace per server.
Note that the udhcpd DHCP server shown in the drawing above is only responsible for any server’s public interface. As described above, the Flusskrebs controller contains a central DHCP server handling all private network interfaces. The flows to connect the public DHCP server are equally straightforward:
Alongside Open vSwitch, the Linux netfilter connection tracking is being used to realize the stateful cloud Firewall feature. With a global shared connection tracking table, we have to prevent it from overflowing and ensure fair use for all servers and customers. This is done by ctcount, another internal project, which keeps track of connection tracking entries, and enforces the limit of up to 80,000 active, concurrent connections per server. If a server reaches this limit, no additional connections can be opened, until a previous one has been closed.
Status Quo
This Open vSwitch-based data plane has served us well and mostly still does today, even with over a million cloud servers. However, we have reached some limitations and wanted to improve scalability, resiliency, and flexibility with a more specialized network stack. Over the last years, our SDN team has been building said solution that's tailored to our needs while being easy to operate and being a strong foundation for new features, such as IPv6 for private networks.
We’ll cover the next steps of our journey, our latest fully self-built network stack, and how it works, in the next part of this series. So stay tuned!
At Hetzner Summit, we had a session "The new Hetzner Cloud network stack - A technical deep dive into how we deliver your packets", where we talked about our current stack and what we’ve built. The recordings should be available soon on the Summit page.
Maximilian Wilhelm
Teamlead SDN











