Infrastructure

Proxmox VE

We design, deploy and run Proxmox VE clusters in production. We are an official Silver Partner.

What Proxmox VE is

A mature virtualization platform, in development since 2008. It runs virtual machines and containers, and manages whole clusters with shared storage and automatic restart on failure, from a single interface. And it is open source.

Built on
Debian GNU/Linux
Virtualization
KVM and LXC
Storage
Ceph, ZFS, NFS, iSCSI, LVM and others
Snapshots
On every storage type, with or without machine memory
Management
One web interface for the whole cluster, with no separate management server to install and licence
Live migration
Machines move between nodes while running, with no interruption
Backup
Built in and schedulable on every storage type. With Proxmox Backup Server it becomes incremental and deduplicated: after the first run only what changed is copied
Automation
A REST API covering every function, plus a command line
Networking
VLAN, bonding, Open vSwitch, virtual networks (SDN)
Firewall
Built in, at cluster, node and machine level
High availability
From three nodes up
Cost
Free. You pay only for support, if you want it
Licence
GNU AGPLv3: free software

High availability: a node fails, the work carries on

Node 1Node 2Node 3Node 3DOWNVMVMVMVMVMVM
If a node fails, its virtual machines restart on the other nodes in the cluster. With no manual intervention.

On clusters we did not install, too

It is the common case. You do not have to explain how it was put together or why: that is ours to work out. Inherited clusters turn up often, with no documentation and nobody left who built them.

Get in touch

How it can be built

A single node

NodeVMVMDisksExternal storageBackup

One server, its virtual machines, and backups on a second system. It is where many start: nodes are added later with nothing to rebuild.

Two nodes with a QDevice and replication

Node 1VMZFSNode 2VMZFSQDeviceExternal storageBackup

Disks are copied to the second node on a schedule, so a failure does not stop work for long. Said plainly: the copy is periodic, so whatever happened since the last run is lost. It is a quick restart, not continuity: and a third vote is needed for the cluster to keep deciding.

Hyper-converged cluster

Node 1VMDisksNode 2VMDisksNode 3VMDisksExternal storageBackup

Three nodes doing compute and storage together with Ceph. Here high availability is real: a node fails and the machines restart elsewhere without losing data. It is what we recommend to anyone who cannot stop.

Cluster with external storage

Node 1VMNode 2VMNode 3VMExternal storageBackup

The nodes use a SAN or NAS you already have, over iSCSI or NFS. It fits when the storage is a recent investment and replacing it to change hypervisor makes no sense.

What we do

Design

Node sizing, storage choice, networking and redundancy. How many machines you actually need and how they should be built, before the hardware is bought.

Deployment

Cluster installation and configuration, distributed storage, backup, firewall, high availability. Handed over working, not half-finished.

Ongoing support

Updates, monitoring, and someone to call when something breaks: a person who knows your infrastructure rather than a support line.

Version upgrades

Moving from an old Proxmox VE release to the current one, even across several versions. Planned, one node at a time, with the option to stop halfway.

Infrastructure recovery

Stalled clusters, upgrades gone wrong, inherited installations with no documentation. We work on systems built by other people too.

Custom development

Components and automation for cases Proxmox VE does not cover out of the box. That is how the cv4pve suite started.

We train the people who will run it

Two levels, from a first installation to a highly available cluster. In our classroom, at your site or online, with the programme built on what your engineer already knows.

Proxmox VE course

What we work with

  • Official Proxmox partner: see our entry on the Proxmox site
  • Ceph in production
  • Red Hat Certified Specialist: Ceph Cloud Storage

The tools we were missing, we wrote

cv4pve is the suite of tools we build for Proxmox VE: automatic snapshots, cluster management, a VDI client, diagnostics, monitoring, and API libraries for five languages. Public since 2016, and running on our clients’ clusters every day.

Explore cv4pve

Coming from VMware?

VMware migration

Questions we get

Is Proxmox VE ready for production?

Yes, and it has been for a long time. Underneath are the same technologies that carry much of the world’s datacentre capacity, the Linux kernel and KVM, used by people running services at scale for years. What differs from closed products is not the solidity: it is that you can see how it is built and do not depend on a single vendor to keep it running.

What support is there if something goes wrong?

Two levels, and they do not exclude each other. There is the official subscription, which puts you in touch with the people who develop Proxmox VE. And there is us: we are an official Silver Partner, we answer, we know your cluster, and we work on what you actually have rather than a generic case.

Do we need a subscription? What does it cost?

The software is free and complete: no features are locked behind a licence, and you can install and run it as it is. The subscription gives you the tested update channel and support from the people who build it: in production we recommend it: and it is priced per socket per year, with public pricing. Nothing per virtual machine, no core counting.

Will it run on the hardware we already have?

Almost always, and it is one of the reasons people arrive here. Proxmox VE has no approved hardware list: it runs on whatever runs Linux, including servers others have stopped certifying. We check first, and if something will not work we say so. It is often the starting point of a migration from VMware.

What happens if a node fails?

The virtual machines that were running on it restart by themselves on another node in the cluster, with no one having to step in. They restart rather than carry on: the node is gone, so there is the time of a reboot. It needs at least three nodes and shared storage, and it is the part to test before going into production, not on the day of the failure.

Have a cluster to design, one giving you trouble, or an emergency?

We can look at how it stands today and what is worth doing. Even just for an opinion.

Get in touch