Infrastructure engineering for East Africa's operators, platforms and regulators See the work →
CloudSpinx · Proxmox Consulting

Proxmox Setup Consultants in Kenya

CloudSpinx plans, builds, secures and runs Proxmox VE clusters for businesses in Kenya and across East Africa: high availability, Ceph and ZFS storage, verified backups, and migrations off VMware, Hyper-V, plain KVM and OpenStack. The software itself carries no licence fee, so what you pay for is the hardware, the network and the engineering.

Free 30-min consultation No lock-in contracts Local on-site engineers

Who we build for

  • 14organisations, from ISPs and payment platforms to a national regulator
  • 6flagship engagements published in full, with the numbers counted
  • 4thof all contributors to the open-source payment switch national systems run on
See the engineering record →
Certified engineers 24/7 support
KES 0 Proxmox licence cost
99.99% HA uptime SLA
3 nodes Where real HA starts
What's Included

Everything in Our Proxmox Consulting Service

Every engagement covers the full scope: no hidden extras, no upselling.

Infrastructure planning

We start from your workload, not a parts list. VM inventory, IOPS and memory ceilings, growth over three years, then the node count and spec that carries it with one node offline.

Proxmox installation and setup

Installation, cluster join, storage pools, VLAN and bridge layout, templates and cloud-init images, resource pools, and a naming scheme your team can actually reason about.

Proxmox HA cluster design

Quorum that holds, corosync on its own physical rings, watchdog fencing, HA groups and restart priorities, plus honest numbers on how long a recovery really takes.

Storage architecture

Ceph across the nodes, ZFS with scheduled replication, or your existing SAN over iSCSI and NFS. We size the failure domain first and pick the technology second.

Backup and disaster recovery

Proxmox Backup Server with deduplicated incrementals, verify jobs, an offsite copy, and a restore rehearsal on the calendar. Retention matched to what the business can lose.

VMware and Hyper-V migration

Wave-based moves off vSphere, ESXi and Hyper-V. Drivers prepared in advance, disks converted, test boots before cutover, and the old estate left standing until you are sure.

KVM and OpenStack consolidation

Bringing scattered libvirt hosts, oVirt, XenServer or an over-built OpenStack estate onto one cluster with a management plane your team will use.

Networking and SDN

Bonds, VLANs, the built-in firewall, and SDN zones with VXLAN or EVPN where you need tenant separation or routing between sites.

Hardening and access

Active Directory or LDAP realms, per-pool role assignments, two-factor for admins, API tokens instead of shared passwords, and the management plane kept off the guest network.

Monitoring integration

Metrics into Prometheus and Grafana or your existing Zabbix, alerts routed to email, Slack or WhatsApp, and dashboards that show cluster capacity rather than pretty graphs.

Automation and IaC

Terraform and Ansible against the Proxmox API, golden images built with Packer, cloud-init for first boot. New VMs stop being a ticket and become a commit.

Managed Proxmox support

Patch windows, capacity reviews, backup verification, and an engineer who already knows your cluster when something breaks at 21:00 on a Friday.

Technologies we use

Proxmox VE 9.2Proxmox Backup Server 4.2CephZFSCorosyncQEMU/KVMLXCProxmox Datacenter ManagerTerraformAnsiblePackercloud-initPrometheusGrafanaZabbixVLAN / VXLAN / EVPN
Reference Architectures

How we design a Proxmox cluster

Four shapes we deploy for Kenyan clients: the cluster and its network fabric, what a node failure actually looks like, how we move you off VMware or Hyper-V, and where the backups live. Sizing, addressing and tuning come out of your workload during the design phase.

01

Three-node Proxmox HA cluster with Ceph

The smallest cluster that survives losing a node and still holds quorum. Every node runs monitors, managers and OSDs, so storage stays available while one box is down. Guest traffic and storage traffic ride bonded uplinks into two switches; corosync gets its own physical rings because it cares about latency, not bandwidth, and a busy storage network will happily starve it into a false failure.

Three-node Proxmox cluster topology Two internet links into a firewall pair, two data switches, three Proxmox nodes each running Ceph monitors and OSDs, and two separate corosync rings. ISP A ISP B Edge firewall pair failover WAN · site VPN · north-south rules Switch A 25 GbE · VLAN trunk Switch B 25 GbE · VLAN trunk pve-01 MON · MGR Guests: VMs and LXC containers OSD OSD OSD NVMe 2 × 25 GbE bond · 2 × 1 GbE rings out-of-band management port pve-02 MON · MGR Guests: VMs and LXC containers OSD OSD OSD NVMe replica of every object lives elsewhere losing this node costs no data pve-03 MON · MGR Guests: VMs and LXC containers OSD OSD OSD NVMe three monitors, so two can still agree quorum survives one dead node Corosync ring 0 dedicated 1 GbE · no other traffic Corosync ring 1 second path · survives a switch reboot quorum = 2 of 3 nodes no single switch in the path
  • Bonded data uplinks: guest, management and Ceph VLANs
  • Corosync ring 0
  • Corosync ring 1
  • WAN uplinks
02

What a node failure actually looks like

Proxmox HA restarts guests, it does not keep them running. A failed node self-fences through its watchdog, the cluster manager waits for that to be certain, then the guests boot again on the survivors. Budget a couple of minutes, not zero. If a workload cannot take a restart, it needs clustering inside the application as well, and we will say so during design rather than after the first outage.

Proxmox HA failover sequence Before and after states of a three-node cluster losing one node, and the four stages of recovery. Before: node 2 stops responding After: guests running again on the survivors pve-01 quorate pve-02 no heartbeat pve-03 quorate pve-01 +1 recovered pve-02 fenced and out of the cluster awaiting hands pve-03 +1 recovered Recovery sequence 1 Heartbeat stops Surviving nodes still hold quorum and keep serving. 2 The node fences itself A hardware watchdog resets it. No split brain, no writes. 3 Guests are rescheduled HA groups and priorities decide who lands where. 4 Guests boot on survivors Shared storage means the disks are already there.
03

Migrating VMware, Hyper-V or OpenStack to Proxmox

We build the new cluster beside the old one and move workloads in waves, quietest first. Nothing is deleted until the last wave has run for a full month, so the way back is always open. Windows guests get their drivers sorted before the move rather than during it, which is where most rushed migrations lose their weekend.

Wave-based migration to Proxmox The existing estate, then six numbered stages: inventory, landing cluster, pilot wave, business wave, critical wave and decommission, with a rollback path back to the old platform. Estate today VMware vSphere / ESXi Hyper-V clusters KVM, oVirt, XenServer OpenStack Physical servers 01 Inventory and dependency map 02 Landing cluster built and proven empty 03 Pilot wave low risk guests first 04 Business wave apps, files, print, AD 05 Critical wave databases and ERP 06 Decommission after a month of quiet Rollback stays open until the old estate is switched off Every wave: drivers prepared · disks converted · test boot · cutover · verify
04

Proxmox Backup Server, verified and offsite

Proxmox Backup Server takes changed blocks only, deduplicates them across every guest, and can prove a snapshot is readable without a human watching. One copy on site for fast restores, a second copy somewhere else entirely, and a restore rehearsal on the calendar. A backup nobody has restored from is a hope, not a plan.

Proxmox Backup Server topology Cluster backing up to an on-site Proxmox Backup Server, which syncs to a second site and to object storage, with a retention ladder and quarterly restore drills. Proxmox cluster guests and host config changed blocks Backup server on site deduplicated datastore verify job proves it restores prune and garbage collection file-level and full-VM restores Second site backup server pull sync · different building, different power Object storage tier S3 compatible target for long retention Retention ladder, tuned to what the business can actually lose daily weekly monthly yearly Restore drill every quarter, into an isolated network

Which shape do you actually need?

The honest version, including what each one costs you. Most Kenyan businesses we scope land on three nodes; a fair number should not be buying a cluster at all yet.

1 node

Single node, backed up properly

One well-specified box with local ZFS and a separate backup server. Ten to fifteen guests for a small office, at a fraction of the cost of a cluster.

Costs you: a hardware failure is a restore, measured in hours, not a failover measured in minutes.

2 nodes + witness

Two nodes with a QDevice

A third vote from a small machine elsewhere on the network keeps quorum alive when one node dies. Storage is ZFS, replicated between the pair on a schedule.

Costs you: a failover loses everything written since the last replication run, and the witness has to sit somewhere genuinely independent.

3 nodes

Three nodes with Ceph

The standard production build and where most clients land. Every node sees every volume, quorum survives one failure, and a dead disk heals itself without anyone driving in.

Costs you: fast switching and enough disks per node. Three is Ceph's floor, and it only breathes properly above it.

5+ nodes

Five or more, split roles

Compute and storage separated, more failure domains, room to lose a node during a patch window and not think about it. This is where Ceph performance gets interesting.

Costs you: real hardware spend, rack space, and someone who owns capacity planning as a job rather than a favour.

These are the shapes, not the build sheet. Node count, disk layout, network sizing and the failure policy all come out of your workload, and that is the first thing we do together.

What a Proxmox setup actually costs in Kenya

The software is free and there is no catch in that sentence. Proxmox VE is open source, so there is no per-socket licence to buy, no core-count audit, and no renewal that doubles because a vendor changed its mind. A support subscription is optional and runs from about EUR 120 per CPU socket per year for the stable enterprise repository up to EUR 1,100 per socket for the top support tier, straight off the published Proxmox price list. We build on Proxmox VE 9.2 with Backup Server 4.2 behind it. On production nodes we take at least the entry tier, and we tell you plainly when the higher ones are not worth it. Everything else is hardware, the network under it, and the engineering. Our engagements are scoped rather than listed, because a single node carrying a 12-VM office and a three-node cluster carrying an ERP are different pieces of work.

  • Node count: an HA cluster is sized to run with one node missing, so the spare capacity is part of the price, not waste
  • CPU sockets and cores: they set the hardware bill and, if you take a subscription, the subscription too
  • Memory: the real ceiling on VM density, and it binds long before the cores do
  • Storage design: Ceph wants more disks and a faster network than local ZFS, and buys you shared storage in return
  • Network: switching, bonded uplinks and the separate corosync paths, which we plan alongside your LAN and firewall design
  • Windows and SQL licensing: follows the physical host cores whichever hypervisor you run, and it is usually the largest line on the quote
  • Migration effort: how many guests, how much of the estate is Windows, and how many cutover windows you can tolerate
  • Managed or self-run: the ongoing service if you would rather not carry it in-house

Designing a Proxmox HA cluster that survives a node failure

High availability in Proxmox is a democracy with a strict rule: more than half the nodes must be able to see each other before anything is allowed to run. That single fact drives most of the design. Two nodes cannot hold a vote when one of them dies, which is why the smallest cluster we recommend for real HA is three. It is also why we refuse to put corosync on the same wire as backup traffic, because a saturated link looks exactly like a dead node from the outside.

Quorum, and the two-node question

Clients often want two nodes because two feels like redundancy. It is not, at least not on its own. A two-node cluster loses quorum the moment one node fails, and the survivor will refuse to start guests rather than risk both nodes writing to the same disks. The fix is a third vote, either a real third node or a small QDevice on a low-power box or VM elsewhere on the network. We will build the two-node-plus-witness shape when budget demands it, and we will tell you exactly which failure modes it does not cover. How many nodes a cluster actually needs goes through the quorum arithmetic and the real failover timings.

The guest has to be able to start somewhere else

HA moves the guest, not its disk. That only works when the disk is reachable from the other nodes, so the storage decision and the HA decision are the same decision. Ceph gives every node access to every volume. ZFS replication gives you a recent copy on a partner node instead, which means a failover can cost you the minutes since the last replication run. Both are legitimate. Choosing without knowing which one you picked is not.

Fencing is not optional and it is not instant

Before a guest restarts elsewhere, the cluster has to be certain the original is really gone. A failed node resets itself through a hardware watchdog, and only then does recovery begin. That is roughly a minute of waiting followed by however long your VM takes to boot. We design HA groups and restart priorities around it so the database comes back before the application that depends on it, and we say up front which workloads need clustering inside the application because a restart is not good enough.

Proxmox storage: Ceph, ZFS or the SAN you already own

There is no default answer here, and anyone who gives you one before asking about your disks is selling. Ceph is the right call when you want every node to see every volume and you can put 25 GbE (or a full mesh of direct links between three nodes, which saves a switch pair) behind it. ZFS is the right call for smaller clusters, tight budgets and workloads that can tolerate losing the last few minutes of writes on a failure. Reusing an existing SAN is often the right call when it is only three years old and already paid for.

  • Ceph: three nodes minimum, more disks per node, a fast dedicated network, and self-healing when a disk or a node dies. It rewards you at five nodes and above
  • ZFS with replication: cheap, fast on local NVMe, snapshots and checksums for free, and an honest recovery point of one replication interval
  • Existing SAN or NAS: iSCSI or NFS keeps a paid-for array in service, at the cost of a shared failure domain the array vendor controls
  • Single node with good backups: for a small office, this beats a badly built cluster every single time

VMware to Proxmox migration, and Hyper-V, KVM and OpenStack

Most of the Proxmox work we are asked to quote in Nairobi starts with a renewal quote that arrived and shocked somebody. Perpetual VMware licences are gone, everything is a subscription bundle priced per core, and even after Broadcom walked back the 72-core minimum in 2025 the renewals landing on Kenyan finance directors are multiples of what they used to be. Proxmox is the VMware alternative most of them land on, because one open-source stack covers what vSphere, vCenter and vSAN were doing, with no licence bill attached. Hyper-V shops usually arrive for a different reason: the Windows estate outgrew a single host and Failover Clustering with shared storage costs more than the workload deserves. Either way, the answer is not a big-bang weekend. We build the new cluster next to the old one and move in waves, quietest first, with the old platform left running until the last wave has been quiet for a month. The detail of what that migration costs and how long it takes, including what does not carry across, is written up separately.

What actually happens to each guest

Linux guests generally move without complaint. Windows guests need their storage and network drivers sorted before the move rather than during it, which is where rushed migrations lose their weekend to a machine that boots to a blue screen at 02:00. We prepare drivers on the source side, convert disks, boot the guest on the new cluster in an isolated network, prove the application works, and only then take the cutover window. Your Windows Server estate keeps its identity, its licensing position and its AD trust through the move.

The licensing trap nobody mentions

Windows Server licences are tied to the physical cores of the host they run on, and a Standard licence can generally only be reassigned to different hardware every 90 days. In a cluster where guests move freely between nodes, that maths bites. We work it out before you buy hardware, not after, because the answer sometimes changes the node count.

OpenStack and scattered KVM hosts

Not every migration is off a commercial hypervisor. We regularly consolidate a handful of standalone libvirt or oVirt boxes, or an OpenStack deployment that turned out to be five times the platform the team needed, onto one cluster with a management plane people will actually log into. The gain is rarely raw performance. It is that patching, backups and capacity finally happen in one place. When the opposite is true and you genuinely need tenants, quotas and a self-service API, we say so and build the OpenStack private cloud instead.

Proxmox Backup Server, restore testing and disaster recovery

A cluster without a tested restore is a single point of failure with extra steps. Proxmox Backup Server sends only changed blocks, deduplicates across every guest so twenty similar Windows VMs cost far less than twenty times one, and can verify a snapshot is readable without anyone watching. We keep one copy on site for fast restores and push a second copy somewhere else entirely, whether that is a second Nairobi facility, your branch office, or an S3 compatible object store now that Backup Server 4.2 supports it directly. Where your backups may legally live is a real constraint here, and the Data Protection Act rules on cross-border storage are part of the design conversation, not an afterthought.

  • Retention laddered daily, weekly, monthly and yearly, set against what the business can actually afford to lose
  • Verify jobs that prove a backup restores, because the alternative is finding out during an incident
  • File-level restore out of a VM backup, so a deleted spreadsheet does not need a full recovery
  • A restore rehearsal every quarter into an isolated network, with the timings written down
  • Host and cluster configuration backed up alongside the guests, so a rebuild is an afternoon rather than a week

Proxmox monitoring, patching and day-2 support

The build is the easy part. What separates a cluster that is still healthy in year three from one that quietly rotted is whether anybody watched it. Proxmox pushes metrics out to an external time-series database natively, which we usually pair with Grafana, or feed into the Zabbix you already run. Alerts route to email, Slack or WhatsApp through the notification targets rather than sitting unread in a web console. On top of that we set a patch cadence, review capacity against real growth every quarter, and check that backups are still verifying. If you would rather not carry any of that, managed Proxmox support folds it into our wider managed IT cover with a response commitment attached.

  • Cluster, node, guest and Ceph health metrics with alert thresholds that mean something
  • Disk and pool capacity trending, because Ceph gets unhappy long before it gets full
  • Backup and verify job outcomes surfaced as alerts, not as a log nobody opens
  • Scheduled patching with a rolling reboot through the cluster and no downtime for HA guests
  • Quarterly capacity and design review as your VM count grows

Running Proxmox on-premises in Nairobi

Plenty of Kenyan businesses have good reasons to keep the platform in their own building: data they would rather not move, an ERP that talks to machines on the factory floor, or an internet link that cannot carry the workload. Those clusters work well, but the constraints are physical and local. Grid outages are frequent enough that your UPS runtime and generator start time become part of the storage design, since Ceph does not enjoy an unplanned group reboot. Dust and heat kill more disks in an unfiltered comms room than anything in the software. And when a node dies at midnight, someone has to physically be there. We are equally happy putting the cluster in a Nairobi colocation facility or on hardware we source and host for you, and the trade-offs between the two come down to who is on site at 2am rather than to price.

When we tell clients not to use Proxmox

We install Proxmox for a living and we still talk people out of it several times a year. Saying so is the whole point of hiring a consultant rather than a box shifter.

  • Your line-of-business vendor only certifies their appliance on ESXi and will drop support the moment you move. Argue with the vendor first, then decide
  • You are deep into NSX, vSAN policy-driven storage or SRM, and the replacement effort dwarfs the licence saving
  • Nobody on the team is comfortable with Linux and there is no budget for support. Proxmox is honest software, and honest software expects an operator
  • You run four VMs for one small office. A single well-backed-up node or managed cloud will serve you better than a cluster you cannot afford to staff
  • The workload genuinely cannot survive a two-minute restart and there is no application-level clustering available. That is a different architecture, not a different hypervisor

What you own when we hand over

No lock-in, and nothing that only works while we hold the password. Every engagement ends with the documentation an engineer who has never met us could pick up: an as-built diagram of the cluster you saw sketched above, the IP and VLAN plan, the storage layout and why it is that way, the HA policy, and a written restore procedure that has been run at least once. We train your team on the console and the API, hand over the automation repository if we built one, and keep supporting the Linux platform underneath only if you want us to. If you decide to run it yourself from day one, that is a perfectly good outcome and we scope for it.

Scope Your Cluster

Tell us what you are virtualising

Answer what you know and skip what you do not. A senior engineer reads every one of these and comes back with a sizing view and a scoped price, usually within one business day. If the honest answer is that you do not need a cluster, we will tell you that instead.

Free, no obligation, and you get a technical answer rather than a sales call.

Next step

Ready to discuss Proxmox Consulting?

A 30-minute scoping call, free, and it commits you to nothing.

Our Process

How Every Proxmox Consulting Engagement Starts

01

Workload discovery

A free session where we inventory what runs today, what it needs, what your recovery targets really are, and what the renewal calendar looks like.

02

Design and sizing

A written design: topology, node spec, network plan, storage maths, HA policy and backup targets, with the reasoning behind each choice and one scoped price.

03

Build, migrate, prove

We build the cluster, move you in waves, pull the plug on a node in front of you to prove failover, and run a real restore before anyone calls it done.

04

Handover or management

Documentation, runbooks and training, then either you run it or we do, with monitoring, patching and a named engineer.

FAQ

Common Questions

What does Proxmox cost in Kenya?
The Proxmox VE software itself is free and open source, so there is no licence to buy. An optional support subscription runs from roughly EUR 120 per CPU socket per year to EUR 1,100 for the highest tier. Your real spend is hardware, network and the engineering to design it. We quote per engagement rather than publishing tiers, because a single node for a small office and a three-node cluster carrying an ERP are different jobs. Send your requirements through the form and we come back with a sizing view and one figure.
Do I need a Proxmox subscription?
Not to run it, and yes for production. The subscription buys access to the stable enterprise repository, which is the package stream we want on a business cluster, plus a vendor to escalate to. We usually put the entry tier on production nodes and skip it on lab and test hardware.
How many nodes do I need for high availability?
Three. A cluster needs more than half its votes present to run anything, so two nodes cannot survive losing one. If three nodes are out of budget, a two-node cluster plus a small QDevice witness elsewhere on the network gives you the third vote, and we will explain which failure modes that shape does not cover.
Can you migrate our VMware environment to Proxmox?
Yes, and it is the most common reason clients call us. We build the new cluster alongside vSphere, inventory dependencies, prepare drivers on the Windows guests, then move workloads in waves with a test boot before every cutover. The old environment stays up until the last wave has been stable for a month, so there is always a way back.
Is Proxmox a real alternative to VMware?
For most business workloads, yes. Proxmox VE covers what vSphere and vCenter did for server virtualization, clustering and live migration, and Ceph stands in for vSAN on shared storage, with no per-core licence behind any of it. Where it is not a like-for-like swap is the wider ecosystem: NSX micro-segmentation, SRM orchestration and vSAN storage policies have no direct equivalent, and a line-of-business vendor that certifies its appliance only on ESXi is a genuine blocker. We tell you which of those applies to you before you commit to anything.
How long is the downtime for each VM during a migration?
For most guests it is a single short window while the final disk sync completes and the machine boots on the new cluster, typically minutes rather than hours. Large database VMs get a planned window agreed with the business, usually overnight or at the weekend. We give you the per-wave timings in the design document before anything moves.
Can Proxmox run our ERP, SQL Server and Active Directory?
Yes. It runs on KVM with the same virtio devices most Linux-based hypervisors use, so Windows Server, SQL Server, Sage, Odoo and Active Directory all run normally. The work is in sizing memory and storage correctly and in getting the Windows licensing position right, since licences follow the physical host cores no matter which hypervisor you choose.
What happens if a node fails?
The remaining nodes keep serving. The failed node resets itself through its watchdog so it cannot corrupt shared storage, then the cluster restarts its HA-managed guests on the survivors, which takes roughly a minute of waiting plus the guest boot time. Guests without HA enabled stay down until you start them, which is deliberate and something we set per workload.
Can you use our existing servers?
Often, yes. Proxmox is not fussy about hardware brands, but Ceph and ZFS both want disks presented directly rather than hidden behind a RAID controller, so an older host may need its controller reflashed or replaced. We audit what you have and tell you what is reusable, what should become a backup target, and what is honestly finished.
Do you offer managed Proxmox support in Kenya?
If you want us to. Managed Proxmox support includes patching with rolling reboots, monitoring and alerting, backup verification, capacity reviews and a named engineer reachable on WhatsApp and phone. Plenty of clients take a build-and-train engagement instead and run it themselves, which is a good outcome too.
Where can the cluster be hosted?
Your own server room, a Nairobi colocation facility, or hardware we source and host for you. Kenyan power and cooling realities matter more than most people expect for on-premises clusters, so we size UPS runtime and cooling into the design and tell you honestly when colocation is the better answer.
WhatsApp