Proxmox Setup Consultants in Kenya
CloudSpinx plans, builds, secures and runs Proxmox VE clusters for businesses in Kenya and across East Africa: high availability, Ceph and ZFS storage, verified backups, and migrations off VMware, Hyper-V, plain KVM and OpenStack. The software itself carries no licence fee, so what you pay for is the hardware, the network and the engineering.
Who we build for
- 14organisations, from ISPs and payment platforms to a national regulator
- 6flagship engagements published in full, with the numbers counted
- 4thof all contributors to the open-source payment switch national systems run on
Everything in Our Proxmox Consulting Service
Every engagement covers the full scope: no hidden extras, no upselling.
Infrastructure planning
We start from your workload, not a parts list. VM inventory, IOPS and memory ceilings, growth over three years, then the node count and spec that carries it with one node offline.
Proxmox installation and setup
Installation, cluster join, storage pools, VLAN and bridge layout, templates and cloud-init images, resource pools, and a naming scheme your team can actually reason about.
Proxmox HA cluster design
Quorum that holds, corosync on its own physical rings, watchdog fencing, HA groups and restart priorities, plus honest numbers on how long a recovery really takes.
Storage architecture
Ceph across the nodes, ZFS with scheduled replication, or your existing SAN over iSCSI and NFS. We size the failure domain first and pick the technology second.
Backup and disaster recovery
Proxmox Backup Server with deduplicated incrementals, verify jobs, an offsite copy, and a restore rehearsal on the calendar. Retention matched to what the business can lose.
VMware and Hyper-V migration
Wave-based moves off vSphere, ESXi and Hyper-V. Drivers prepared in advance, disks converted, test boots before cutover, and the old estate left standing until you are sure.
KVM and OpenStack consolidation
Bringing scattered libvirt hosts, oVirt, XenServer or an over-built OpenStack estate onto one cluster with a management plane your team will use.
Networking and SDN
Bonds, VLANs, the built-in firewall, and SDN zones with VXLAN or EVPN where you need tenant separation or routing between sites.
Hardening and access
Active Directory or LDAP realms, per-pool role assignments, two-factor for admins, API tokens instead of shared passwords, and the management plane kept off the guest network.
Monitoring integration
Metrics into Prometheus and Grafana or your existing Zabbix, alerts routed to email, Slack or WhatsApp, and dashboards that show cluster capacity rather than pretty graphs.
Automation and IaC
Terraform and Ansible against the Proxmox API, golden images built with Packer, cloud-init for first boot. New VMs stop being a ticket and become a commit.
Managed Proxmox support
Patch windows, capacity reviews, backup verification, and an engineer who already knows your cluster when something breaks at 21:00 on a Friday.
Technologies we use
How we design a Proxmox cluster
Four shapes we deploy for Kenyan clients: the cluster and its network fabric, what a node failure actually looks like, how we move you off VMware or Hyper-V, and where the backups live. Sizing, addressing and tuning come out of your workload during the design phase.
Three-node Proxmox HA cluster with Ceph
The smallest cluster that survives losing a node and still holds quorum. Every node runs monitors, managers and OSDs, so storage stays available while one box is down. Guest traffic and storage traffic ride bonded uplinks into two switches; corosync gets its own physical rings because it cares about latency, not bandwidth, and a busy storage network will happily starve it into a false failure.
- Bonded data uplinks: guest, management and Ceph VLANs
- Corosync ring 0
- Corosync ring 1
- WAN uplinks
What a node failure actually looks like
Proxmox HA restarts guests, it does not keep them running. A failed node self-fences through its watchdog, the cluster manager waits for that to be certain, then the guests boot again on the survivors. Budget a couple of minutes, not zero. If a workload cannot take a restart, it needs clustering inside the application as well, and we will say so during design rather than after the first outage.
Migrating VMware, Hyper-V or OpenStack to Proxmox
We build the new cluster beside the old one and move workloads in waves, quietest first. Nothing is deleted until the last wave has run for a full month, so the way back is always open. Windows guests get their drivers sorted before the move rather than during it, which is where most rushed migrations lose their weekend.
Proxmox Backup Server, verified and offsite
Proxmox Backup Server takes changed blocks only, deduplicates them across every guest, and can prove a snapshot is readable without a human watching. One copy on site for fast restores, a second copy somewhere else entirely, and a restore rehearsal on the calendar. A backup nobody has restored from is a hope, not a plan.
Which shape do you actually need?
The honest version, including what each one costs you. Most Kenyan businesses we scope land on three nodes; a fair number should not be buying a cluster at all yet.
Single node, backed up properly
One well-specified box with local ZFS and a separate backup server. Ten to fifteen guests for a small office, at a fraction of the cost of a cluster.
Costs you: a hardware failure is a restore, measured in hours, not a failover measured in minutes.
Two nodes with a QDevice
A third vote from a small machine elsewhere on the network keeps quorum alive when one node dies. Storage is ZFS, replicated between the pair on a schedule.
Costs you: a failover loses everything written since the last replication run, and the witness has to sit somewhere genuinely independent.
Three nodes with Ceph
The standard production build and where most clients land. Every node sees every volume, quorum survives one failure, and a dead disk heals itself without anyone driving in.
Costs you: fast switching and enough disks per node. Three is Ceph's floor, and it only breathes properly above it.
Five or more, split roles
Compute and storage separated, more failure domains, room to lose a node during a patch window and not think about it. This is where Ceph performance gets interesting.
Costs you: real hardware spend, rack space, and someone who owns capacity planning as a job rather than a favour.
These are the shapes, not the build sheet. Node count, disk layout, network sizing and the failure policy all come out of your workload, and that is the first thing we do together.
What a Proxmox setup actually costs in Kenya
The software is free and there is no catch in that sentence. Proxmox VE is open source, so there is no per-socket licence to buy, no core-count audit, and no renewal that doubles because a vendor changed its mind. A support subscription is optional and runs from about EUR 120 per CPU socket per year for the stable enterprise repository up to EUR 1,100 per socket for the top support tier, straight off the published Proxmox price list. We build on Proxmox VE 9.2 with Backup Server 4.2 behind it. On production nodes we take at least the entry tier, and we tell you plainly when the higher ones are not worth it. Everything else is hardware, the network under it, and the engineering. Our engagements are scoped rather than listed, because a single node carrying a 12-VM office and a three-node cluster carrying an ERP are different pieces of work.
- Node count: an HA cluster is sized to run with one node missing, so the spare capacity is part of the price, not waste
- CPU sockets and cores: they set the hardware bill and, if you take a subscription, the subscription too
- Memory: the real ceiling on VM density, and it binds long before the cores do
- Storage design: Ceph wants more disks and a faster network than local ZFS, and buys you shared storage in return
- Network: switching, bonded uplinks and the separate corosync paths, which we plan alongside your LAN and firewall design
- Windows and SQL licensing: follows the physical host cores whichever hypervisor you run, and it is usually the largest line on the quote
- Migration effort: how many guests, how much of the estate is Windows, and how many cutover windows you can tolerate
- Managed or self-run: the ongoing service if you would rather not carry it in-house
Designing a Proxmox HA cluster that survives a node failure
High availability in Proxmox is a democracy with a strict rule: more than half the nodes must be able to see each other before anything is allowed to run. That single fact drives most of the design. Two nodes cannot hold a vote when one of them dies, which is why the smallest cluster we recommend for real HA is three. It is also why we refuse to put corosync on the same wire as backup traffic, because a saturated link looks exactly like a dead node from the outside.
Quorum, and the two-node question
Clients often want two nodes because two feels like redundancy. It is not, at least not on its own. A two-node cluster loses quorum the moment one node fails, and the survivor will refuse to start guests rather than risk both nodes writing to the same disks. The fix is a third vote, either a real third node or a small QDevice on a low-power box or VM elsewhere on the network. We will build the two-node-plus-witness shape when budget demands it, and we will tell you exactly which failure modes it does not cover. How many nodes a cluster actually needs goes through the quorum arithmetic and the real failover timings.
The guest has to be able to start somewhere else
HA moves the guest, not its disk. That only works when the disk is reachable from the other nodes, so the storage decision and the HA decision are the same decision. Ceph gives every node access to every volume. ZFS replication gives you a recent copy on a partner node instead, which means a failover can cost you the minutes since the last replication run. Both are legitimate. Choosing without knowing which one you picked is not.
Fencing is not optional and it is not instant
Before a guest restarts elsewhere, the cluster has to be certain the original is really gone. A failed node resets itself through a hardware watchdog, and only then does recovery begin. That is roughly a minute of waiting followed by however long your VM takes to boot. We design HA groups and restart priorities around it so the database comes back before the application that depends on it, and we say up front which workloads need clustering inside the application because a restart is not good enough.
Proxmox storage: Ceph, ZFS or the SAN you already own
There is no default answer here, and anyone who gives you one before asking about your disks is selling. Ceph is the right call when you want every node to see every volume and you can put 25 GbE (or a full mesh of direct links between three nodes, which saves a switch pair) behind it. ZFS is the right call for smaller clusters, tight budgets and workloads that can tolerate losing the last few minutes of writes on a failure. Reusing an existing SAN is often the right call when it is only three years old and already paid for.
- Ceph: three nodes minimum, more disks per node, a fast dedicated network, and self-healing when a disk or a node dies. It rewards you at five nodes and above
- ZFS with replication: cheap, fast on local NVMe, snapshots and checksums for free, and an honest recovery point of one replication interval
- Existing SAN or NAS: iSCSI or NFS keeps a paid-for array in service, at the cost of a shared failure domain the array vendor controls
- Single node with good backups: for a small office, this beats a badly built cluster every single time
VMware to Proxmox migration, and Hyper-V, KVM and OpenStack
Most of the Proxmox work we are asked to quote in Nairobi starts with a renewal quote that arrived and shocked somebody. Perpetual VMware licences are gone, everything is a subscription bundle priced per core, and even after Broadcom walked back the 72-core minimum in 2025 the renewals landing on Kenyan finance directors are multiples of what they used to be. Proxmox is the VMware alternative most of them land on, because one open-source stack covers what vSphere, vCenter and vSAN were doing, with no licence bill attached. Hyper-V shops usually arrive for a different reason: the Windows estate outgrew a single host and Failover Clustering with shared storage costs more than the workload deserves. Either way, the answer is not a big-bang weekend. We build the new cluster next to the old one and move in waves, quietest first, with the old platform left running until the last wave has been quiet for a month. The detail of what that migration costs and how long it takes, including what does not carry across, is written up separately.
What actually happens to each guest
Linux guests generally move without complaint. Windows guests need their storage and network drivers sorted before the move rather than during it, which is where rushed migrations lose their weekend to a machine that boots to a blue screen at 02:00. We prepare drivers on the source side, convert disks, boot the guest on the new cluster in an isolated network, prove the application works, and only then take the cutover window. Your Windows Server estate keeps its identity, its licensing position and its AD trust through the move.
The licensing trap nobody mentions
Windows Server licences are tied to the physical cores of the host they run on, and a Standard licence can generally only be reassigned to different hardware every 90 days. In a cluster where guests move freely between nodes, that maths bites. We work it out before you buy hardware, not after, because the answer sometimes changes the node count.
OpenStack and scattered KVM hosts
Not every migration is off a commercial hypervisor. We regularly consolidate a handful of standalone libvirt or oVirt boxes, or an OpenStack deployment that turned out to be five times the platform the team needed, onto one cluster with a management plane people will actually log into. The gain is rarely raw performance. It is that patching, backups and capacity finally happen in one place. When the opposite is true and you genuinely need tenants, quotas and a self-service API, we say so and build the OpenStack private cloud instead.
Proxmox Backup Server, restore testing and disaster recovery
A cluster without a tested restore is a single point of failure with extra steps. Proxmox Backup Server sends only changed blocks, deduplicates across every guest so twenty similar Windows VMs cost far less than twenty times one, and can verify a snapshot is readable without anyone watching. We keep one copy on site for fast restores and push a second copy somewhere else entirely, whether that is a second Nairobi facility, your branch office, or an S3 compatible object store now that Backup Server 4.2 supports it directly. Where your backups may legally live is a real constraint here, and the Data Protection Act rules on cross-border storage are part of the design conversation, not an afterthought.
- Retention laddered daily, weekly, monthly and yearly, set against what the business can actually afford to lose
- Verify jobs that prove a backup restores, because the alternative is finding out during an incident
- File-level restore out of a VM backup, so a deleted spreadsheet does not need a full recovery
- A restore rehearsal every quarter into an isolated network, with the timings written down
- Host and cluster configuration backed up alongside the guests, so a rebuild is an afternoon rather than a week
Proxmox monitoring, patching and day-2 support
The build is the easy part. What separates a cluster that is still healthy in year three from one that quietly rotted is whether anybody watched it. Proxmox pushes metrics out to an external time-series database natively, which we usually pair with Grafana, or feed into the Zabbix you already run. Alerts route to email, Slack or WhatsApp through the notification targets rather than sitting unread in a web console. On top of that we set a patch cadence, review capacity against real growth every quarter, and check that backups are still verifying. If you would rather not carry any of that, managed Proxmox support folds it into our wider managed IT cover with a response commitment attached.
- Cluster, node, guest and Ceph health metrics with alert thresholds that mean something
- Disk and pool capacity trending, because Ceph gets unhappy long before it gets full
- Backup and verify job outcomes surfaced as alerts, not as a log nobody opens
- Scheduled patching with a rolling reboot through the cluster and no downtime for HA guests
- Quarterly capacity and design review as your VM count grows
Running Proxmox on-premises in Nairobi
Plenty of Kenyan businesses have good reasons to keep the platform in their own building: data they would rather not move, an ERP that talks to machines on the factory floor, or an internet link that cannot carry the workload. Those clusters work well, but the constraints are physical and local. Grid outages are frequent enough that your UPS runtime and generator start time become part of the storage design, since Ceph does not enjoy an unplanned group reboot. Dust and heat kill more disks in an unfiltered comms room than anything in the software. And when a node dies at midnight, someone has to physically be there. We are equally happy putting the cluster in a Nairobi colocation facility or on hardware we source and host for you, and the trade-offs between the two come down to who is on site at 2am rather than to price.
When we tell clients not to use Proxmox
We install Proxmox for a living and we still talk people out of it several times a year. Saying so is the whole point of hiring a consultant rather than a box shifter.
- Your line-of-business vendor only certifies their appliance on ESXi and will drop support the moment you move. Argue with the vendor first, then decide
- You are deep into NSX, vSAN policy-driven storage or SRM, and the replacement effort dwarfs the licence saving
- Nobody on the team is comfortable with Linux and there is no budget for support. Proxmox is honest software, and honest software expects an operator
- You run four VMs for one small office. A single well-backed-up node or managed cloud will serve you better than a cluster you cannot afford to staff
- The workload genuinely cannot survive a two-minute restart and there is no application-level clustering available. That is a different architecture, not a different hypervisor
What you own when we hand over
No lock-in, and nothing that only works while we hold the password. Every engagement ends with the documentation an engineer who has never met us could pick up: an as-built diagram of the cluster you saw sketched above, the IP and VLAN plan, the storage layout and why it is that way, the HA policy, and a written restore procedure that has been run at least once. We train your team on the console and the API, hand over the automation repository if we built one, and keep supporting the Linux platform underneath only if you want us to. If you decide to run it yourself from day one, that is a perfectly good outcome and we scope for it.
Tell us what you are virtualising
Answer what you know and skip what you do not. A senior engineer reads every one of these and comes back with a sizing view and a scoped price, usually within one business day. If the honest answer is that you do not need a cluster, we will tell you that instead.
Ready to discuss Proxmox Consulting?
A 30-minute scoping call, free, and it commits you to nothing.
How Every Proxmox Consulting Engagement Starts
Workload discovery
A free session where we inventory what runs today, what it needs, what your recovery targets really are, and what the renewal calendar looks like.
Design and sizing
A written design: topology, node spec, network plan, storage maths, HA policy and backup targets, with the reasoning behind each choice and one scoped price.
Build, migrate, prove
We build the cluster, move you in waves, pull the plug on a node in front of you to prove failover, and run a real restore before anyone calls it done.
Handover or management
Documentation, runbooks and training, then either you run it or we do, with monitoring, patching and a named engineer.