Managed Kubernetes Clusters in Kenya
CloudSpinx designs, builds and runs Kubernetes clusters for businesses in Kenya and across East Africa, on EKS, AKS and GKE or on hardware you own in a Nairobi facility. What you buy is the cluster, the monitoring, the upgrade calendar and an engineer who picks up at 03:00, rather than a platform only we know how to operate. Build work is quoted per cluster, running it afterwards is a monthly figure driven by node count and cover hours, and where the honest answer is that you do not need Kubernetes at all, you get told that before you sign anything.
Who we build for
- 14organisations, from ISPs and payment platforms to a national regulator
- 6flagship engagements published in full, with the numbers counted
- 4thof all contributors to the open-source payment switch national systems run on
Everything in Our Managed Kubernetes Service
Every engagement covers the full scope: no hidden extras, no upselling.
Kubernetes consulting and cluster design
The six decisions that are painful to reverse: control-plane shape and failure domains, the CNI, the storage layer, how traffic gets in, how tenants are separated, and what a node pool is for. Sold as a fixed piece of work whether or not we build it.
Kubernetes installation and setup
kubeadm, RKE2, k3s or Talos Linux on your own hardware, or EKS, AKS and GKE built with Terraform or OpenTofu. Everything in your repository, nothing clicked together in a console we then have to remember.
Managed Kubernetes operations
We hold the pager. Patching, certificate rotation, capacity headroom, etcd health, node replacement and the boring weekly work that decides whether the cluster is still healthy in year three.
On-premises and bare-metal Kubernetes
A real cluster in your rack or a Nairobi facility, with MetalLB or kube-vip for load balancing and Ceph, Rook or Longhorn for storage, because none of that arrives free the way it does in a cloud.
EKS, AKS and GKE
Managed control planes done properly: private endpoints, IRSA or workload identity instead of long-lived keys, node groups sized from measurement, and Karpenter or Cluster Autoscaler where the load genuinely varies.
Kubernetes version upgrades
Upstream patches three minor versions and each gets about a year, so this is a calendar item rather than a project. We upgrade control plane then nodes, in a window, with the API deprecations checked against your manifests first.
Kubernetes monitoring and alerting
Prometheus, Grafana, Alertmanager and Loki, wired to what your users actually feel: pods that will not schedule, a PersistentVolume filling, certificates about to expire, etcd latency. Not another CPU graph nobody acts on.
Cluster networking and ingress
Cilium or Calico, network policy that defaults to deny rather than to allow, ingress-nginx or Gateway API, and cert-manager issuing and renewing TLS so nobody is renewing certificates by hand at midnight.
Storage and stateful workloads
CSI drivers, StorageClasses that reflect what the workload needs, volume snapshots, and a straight answer on whether your database belongs in the cluster at all. Frequently it does not.
Kubernetes security hardening
RBAC that is not cluster-admin for everyone, Pod Security Standards, admission policy with Kyverno, image scanning in the pipeline, secrets from Vault or External Secrets Operator, and a CIS benchmark run you can hand to an auditor.
Backup and cluster disaster recovery
Velero for namespaces and volumes, scheduled etcd snapshots held off the cluster, and a restore rehearsed on a calendar rather than assumed. A backup nobody has restored is a belief, not a backup.
24/7 Kubernetes support
P1 response within an hour, around the clock, whether the cluster is on your metal or in somebody else's cloud. Remote first, and on site the same day in Nairobi when the problem is physical.
Technologies we use
What we actually draw before a cluster gets built
Four of the figures that come out of a Kubernetes design engagement: who is accountable for each layer, the three places the cluster can physically live, what a failure really costs you in seconds, and the upgrade calendar nobody budgets for. Node counts, address plans and resource sizing come out of your workloads.
Kubernetes cluster architecture, by who is accountable for each layer
Almost every argument about a cluster is really an argument about this table. A managed service moves one layer off your plate and leaves five. The two layers with an accent bar have no safe default: somebody has to choose, and both are close to unchangeable once production workloads are sitting on them.
- The cloud provider's problem
- Ours, and what the monthly figure buys
- Yours, and it stays yours
Where a Kubernetes cluster can run from Kenya, and what each shape costs
No hyperscaler runs a region in this country, so a managed control plane answers from Cape Town or Johannesburg and every call to it crosses a border. For most workloads that is a non-issue. It stops being one the moment a regulator, a client contract or the Data Protection Act attaches conditions to where the data sits.
- Traffic from your users
- Link between the two halves
What actually happens when a Kubernetes node dies
The figure people expect us not to draw. Kubernetes does not notice a dead node for 50 seconds and then waits out a 300 second toleration before it moves anything, so a single-replica workload is gone for the best part of six minutes on stock defaults. Both numbers are documented defaults, not a measurement of a bad day.
- Timer the cluster is waiting on
- Time your users can see
The Kubernetes upgrade calendar, and what falling behind costs
Kubernetes patches three minor versions at a time and each one gets about a year, so one upgrade a year is the floor rather than the plan. The two clouds that publish a number for lateness charge the same six times over, which makes this the rare piece of infrastructure economics you can check yourself before we quote.
Which of these are you actually buying?
None of them is marked as the recommended one, because the answer comes out of your workloads rather than out of a preference. The last card is a real outcome of a design conversation and we have written it more than once.
EKS, AKS or GKE, run properly
Right when nothing forces the data to stay in the country and you would rather the control plane were somebody else's pager. We design it, build it as code and operate it.
Costs you: a per-cluster fee that multiplies with environments, and a control plane a border away.
A cluster in a Nairobi facility
Right when a regulator or a client contract fixes where the data lives, or the workload is steady and large enough that renting it costs more than owning it.
Costs you: etcd, load balancing and storage all become real jobs. Usually ours, under the monthly figure.
Take over an existing cluster
The commonest call we get. Somebody built it, left, and it has not been upgraded since. We read it first and tell you what we found before anybody signs anything.
Costs you: a review, and occasionally the news that a rebuild is cheaper than the rescue.
You do not need Kubernetes
Three applications, one team, steady traffic. Kubernetes would add an operating burden you then have to staff, and the cluster becomes the thing that breaks.
Costs you: nothing, which is the point. Containers on a couple of servers, and a conversation again in a year.
These are shapes, not build sheets. Node counts, CNI and storage choices, resource requests and the failure timings you actually want instead of the defaults above all come out of your workloads during design, which is the first piece of work rather than a free extra bolted onto a build.
What managed Kubernetes actually costs, and where the money goes
The cluster is almost never the expensive part, which is the single most useful thing to know before you budget for one. Two of the three big clouds publish a per-cluster fee and it is small: Amazon charges USD 0.10 per cluster per hour for an EKS control plane and Google charges the same flat rate for GKE. That is roughly USD 73 a month. The nodes underneath it cost several times more, and the people cost more again.
What we quote on
Design is a fixed piece of work. A build is priced per cluster: how many clusters and environments, whether it is managed cloud or your own hardware, how much of the platform layer you need on day one, and whether anything stateful is going in. Running it afterwards is a monthly figure driven by node count, cluster count, whether you want out-of-hours cover, and whether we operate it or support your engineers operating it. We do not price against your cloud bill, because a supplier who earns more when your spend grows should not be the one advising you on node sizing.
- Node capacity, and how much of it you waste. Requests set from a guess rather than from measurement is the biggest line on almost every cluster we inherit. Half-empty nodes cost the same as full ones.
- How many clusters you decided to have. The per-cluster fee is per cluster. Separate production, staging and development clusters triple it before a single pod runs, and there are good reasons to do it anyway. It should be a decision, not an accident.
- Falling behind on versions. Both AWS and Google move an out-of-support cluster onto extended support automatically and charge USD 0.60 per cluster per hour for it, six times the standard rate, around USD 438 a month. This is the cost nobody models.
- Whether anything is stateful. Stateless web workloads are cheap to run and cheap to recover. Databases, queues and anything with a PersistentVolume bring storage, backup and a much harder failure conversation.
- Who carries the pager. A cluster nobody operates degrades quietly for months and then fails loudly. Either you resource that properly in house or you buy it, and both are real numbers.
Kubernetes cluster design: the decisions made before anything is installed
Installing Kubernetes is a solved problem and takes an afternoon. Everything that makes a cluster survivable is decided before that afternoon, and a fair number of those decisions are close to unchangeable once real workloads are running. This is why we sell design as its own piece of work, and why you are welcome to take it and have somebody else build.
Design without the build
A good number of clients take the design and hand it to their own platform team, or to whoever already holds the contract. That is priced to stand alone deliberately. A design that only works if we are the ones building it is a sales document, and you can usually tell which one you were given by about page ten. Everything we produce is written to be executed by any competent engineer, the same reason our advisory work ends in artefacts rather than in a retainer. If you are still choosing that engineer, we have written up what to ask before you hire one.
The ceilings are published, and they are higher than you need
Kubernetes documents its own tested limits: no more than 110 pods per node, 5,000 nodes and 150,000 pods in a single cluster. Almost nobody in this market is anywhere near those, which is worth saying plainly, because a great deal of Kubernetes advice on the internet is written for organisations operating at a size that has nothing to do with yours. If you are running twelve services on six nodes, most of what you read about multi-cluster federation is noise.
- The control plane and its failure domains. Three nodes is the minimum that tolerates losing one. Putting all three in one rack, on one switch or on one power feed means you built a three-node cluster with a one-node failure domain, which is the most common mistake we find.
- The CNI. Cilium, Calico or the cloud provider's own. This decides how network policy works, how much observability you get for free, and how many pods actually fit on a node. Swapping it later on a live cluster is a migration, not a change.
- The storage layer. Which CSI driver, which StorageClasses, whether volumes can move between nodes, and how snapshots work. Get this wrong and the first stateful workload is the one that finds out.
- How traffic gets in. A cloud load balancer, or MetalLB and kube-vip if you are on your own metal. On-premises clusters do not get a load balancer handed to them, and this surprises people.
- Isolation between teams and environments. Namespaces with quotas and network policy, separate node pools, or genuinely separate clusters. All three are defensible; only one is right for your organisation and your auditors.
- What a node pool is for. Sizes, taints, and which workloads are allowed where. Note that pod capacity does not track instance size the way people assume: on the AWS VPC CNI, an m5.xlarge and an m5.2xlarge both cap at 58 pods, because the limit comes from network interface arithmetic rather than from CPU or memory.
On-premises Kubernetes: bare metal, your rack, your rules
Self-managed Kubernetes on your own hardware is a genuine option here and we build it regularly, usually for one of two reasons: the data is not allowed to leave the country, or the workload is steady and large enough that renting compute costs more than owning it. What changes is the amount of the platform you now own. A cloud hands you a load balancer, a storage backend and a managed control plane; on your own metal all three become yours, and pretending otherwise is how on-premises clusters end up half-built.
- The distribution matters less than the operator. kubeadm for a plain upstream cluster, RKE2 where a hardened default and a security profile are wanted, k3s for small or edge footprints, Talos Linux where an immutable API-managed node OS is worth the change in habits. We have production experience with all four and no religion about it.
- Load balancing is now a component you install. MetalLB in layer 2 or BGP mode, or kube-vip for the control-plane endpoint. Neither is difficult and both are frequently forgotten until the first Service of type LoadBalancer sits in pending forever.
- Storage is the real work. Ceph through Rook where the estate justifies it, Longhorn for something smaller, or an existing SAN through its CSI driver. This is the layer that decides whether a node can be rebooted casually.
- The control plane is yours to keep alive. Three nodes across three failure domains, etcd on fast disks and nowhere near the workload, snapshots off the cluster, and certificate expiry on a calendar. Cluster certificates expiring at the one-year mark is a real outage we have been called in to fix.
- The hardware underneath still has to be right. Sometimes the honest answer is a Kubernetes cluster on top of a Proxmox platform so that nodes are rebuildable virtual machines rather than pets, and sometimes it is straight onto the metal for the performance.
EKS, AKS and GKE: what the cloud actually manages for you
A managed control plane removes exactly one of the six layers a cluster is made of, and it is not the layer that generates most of your incidents. Amazon, Microsoft and Google run the API server, etcd, the scheduler and the controllers, and they do it well. Your node pools, your CNI configuration, your ingress, your RBAC, your resource requests, your upgrades and your monitoring all remain yours. The gap between "we are on EKS" and "we have a cluster somebody is responsible for" is where most of our rescue work comes from.
The tiers each vendor actually sells
AWS charges a flat control-plane fee and nothing else at that layer. Google charges the same flat fee, with a monthly free-tier credit that covers one zonal or Autopilot cluster outright, so a first cluster can genuinely be free. Azure is the odd one: AKS cluster management is free on the Free tier, which caps at 1,000 nodes and carries no financially backed API server SLA, and you pay for the Standard tier to get that SLA and a higher ceiling, or Premium for two years of long-term support on one version. Choose the tier from the SLA you need, not from the sticker.
The Kenyan constraint on all three
None of the three runs a region in Kenya. The nearest are Cape Town and Johannesburg, so your API server, your etcd and your nodes are all a border away, and every kubectl call and every control-plane operation is a cross-border round trip. For a normal web application that is unremarkable and the latency does not matter. It matters when a regulator, a client contract or the Data Protection Act attaches conditions to where personal data lives, and at that point a managed cloud cluster has no answer to give you. The same constraint drives how we approach cloud architecture generally, and it is the reason on-premises Kubernetes is a live option here rather than a nostalgia.
The pod-density trap on EKS specifically
The AWS VPC CNI gives every pod a real VPC address, which is elegant until you count them. The ceiling per node comes from network interface arithmetic: an m5.large tops out at 29 pods and an m5.xlarge at 58, and doubling again to an m5.2xlarge still gives you 58. Teams size nodes on CPU and memory, then discover the cluster will not schedule while the nodes look half idle. Prefix delegation or a different CNI both solve it, and both are much easier to choose at design time than to retrofit.
Kubernetes upgrades are a calendar item, not a project
This is the part that quietly decides whether a cluster is an asset or a liability in three years. Upstream Kubernetes maintains the three most recent minor releases, and each one gets roughly a year of patches. Releases land about three times a year. So a cluster nobody touches falls out of support in about twelve months, and every month after that it is running a version where a published CVE simply never gets a fix.
Rescuing a cluster that is already years behind
We get this call often enough that it has a shape. Somebody built the cluster, left, and it has not been upgraded since. You cannot jump minor versions, so getting current means a sequence of upgrades, each with its own deprecation check, and on a badly neglected cluster a rebuild alongside with a workload migration is genuinely cheaper and safer than the climb. We will tell you which of the two you are looking at after reading the cluster, and we say so before quoting rather than after.
- The cloud makes it a bill. AWS and Google both move a lapsed cluster onto extended support automatically and charge six times the standard rate for it. Per cluster. Multiply that by every environment you have.
- Your own hardware makes it a risk. Nobody invoices you. Nobody patches you either, and eventually you are running a control plane with known holes in it because the upgrade got harder every quarter you deferred it.
- The work is in the manifests, not the cluster. Removed API versions are what break upgrades. We check your workloads against the deprecation list before touching anything, which is why our upgrades are dull and why the ones that go wrong are the ones done in a hurry.
- Control plane first, then nodes. Nodes may run one minor version behind the control plane, never ahead, so the order is fixed. On a well-built cluster the node half is a rolling replacement nobody notices.
Kubernetes monitoring, alerting and who gets woken up
Kubernetes is very good at hiding degradation. Pods restart, replicas cover for each other, and a cluster can be quietly unhealthy for weeks while every dashboard stays green. What we install is Prometheus and Grafana with Alertmanager and usually Loki for logs, and then the part that takes actual judgement: deciding what deserves to wake a person.
An alert that nobody answers is not monitoring
Every alert we configure routes to a person and has a documented action. If neither exists, we delete the alert rather than leave it, because a channel full of noise trains a team to ignore the one that mattered. Where you want us carrying that pager rather than your engineers, it comes through the same cover model as our managed support, and where your team carries it, we build the runbooks they will be reading at 03:00.
- Pods that cannot schedule. Pending pods mean you are out of capacity, out of pod slots, or your requests are wrong. Users feel this before any CPU graph moves.
- Restart loops and OOM kills. A CrashLoopBackOff that has been running for a fortnight is a fault somebody stopped noticing.
- PersistentVolumes filling up. The most boring outage in Kubernetes, and one of the most common. It has a slope, so it can always be caught early.
- Certificate expiry. Both cluster certificates and application TLS. cert-manager handles the second; the first is on us and on a calendar.
- etcd latency and quorum health. The earliest honest warning that a control plane is in trouble, and the one almost nobody has an alert for.
Kubernetes support in Kenya: who answers, and how fast
Some of this is engineering and some of it is geography, and the geography decides more architectures than people expect. If the cluster is in a cloud, it is in Cape Town or Johannesburg and your team administers it over the internet, which is fine and which we do daily. If the cluster is in a Nairobi facility, latency to your users is a couple of milliseconds, peering through KIXP keeps local traffic local, and somebody has to be able to reach the rack. That last one is not a small detail at 02:00 on a Sunday.
Power, and what a node reboot really means here
A Nairobi facility with real generator and UPS cover is a different proposition from a cluster in a server room behind a building's power supply. Kubernetes will happily reschedule around a node that lost power, which is precisely the problem: it hides the fault until the day two nodes go at once and you discover all three control-plane members were on the same feed. Where the hardware sits in your own building, that assessment is part of the design, and where it should not, colocated hardware is usually the cheaper fix.
Who is actually on site
Remote is where nearly all of this work happens, including upgrades, incident response and the monitoring. On-site in Nairobi is same-day when the problem is physical: a failed disk, a switch, a node that will not come back. We deliver across East Africa and we engineer remotely beyond it, and we would rather tell you that a cluster in a city we cannot reach quickly needs a local pair of hands on retainer than pretend distance does not exist.
When we tell clients not to run Kubernetes
We build and operate clusters for a living, which is exactly why this section is here. Kubernetes is the right answer less often than the industry implies, and the wrong cluster is more expensive than no cluster.
- Three applications, one team, steady traffic. Containers on a couple of well-run servers, or a managed container service, will be cheaper to run and far cheaper to staff. You get most of the benefit and none of the operating burden.
- Nobody will own it. A cluster is not a product you install once. If no one is going to hold the pager, watch capacity and run the upgrades, it decays, and a decayed cluster is worse than the virtual machines it replaced.
- Your problem is actually delivery. Plenty of teams who ask for Kubernetes want repeatable deploys and a way back from a bad release. That is a pipeline problem, it is a fraction of the cost, and it works fine without a cluster.
- One large stateful database. A serious database usually belongs outside the cluster on hardware or a managed service. Putting it in Kubernetes for consistency's sake buys complexity and no resilience you did not already have.
- You are moving to Kubernetes to fix reliability. It will not. Kubernetes restarts things; it does not make an application that falls over under load stop falling over. Fix that first, then decide whether you still want a cluster.
What you own at handover
The Terraform or OpenTofu repository and its state, the Helm charts and manifests, the monitoring and alert configuration, the runbooks, the architecture drawings and the upgrade calendar with the next dates already in it. Cloud accounts are created under your own organisation with your billing relationship, never resold through ours. Access is your identity provider and your RBAC, so there is no CloudSpinx credential to revoke on the day you decide to run it yourselves. That day is a decision you get to make again every year, and a platform you cannot take back is not one we are willing to sell.
Tell us what you are trying to run
Enough for a sizing view and a real figure rather than a range. Every field has an escape hatch, so answer what you know and leave the rest to the assessment.
Ready to discuss Managed Kubernetes?
A 30-minute scoping call, free, and it commits you to nothing.
How Every Managed Kubernetes Engagement Starts
Read what you have
Workload inventory, an honest look at any existing cluster, where the data is allowed to live, and whether Kubernetes is the right shape at all. This is the free part and it changes the plan more often than not.
Design and cost it
Control-plane shape and failure domains, CNI, storage, ingress, isolation model and node pools, plus one figure for the build and a modelled monthly run rate. Yours to keep whoever builds it.
Build it as code
Cluster provisioned with Terraform or OpenTofu in your repository, platform layer installed, workloads migrated in waves, and the old environment kept standing until a full business cycle has passed.
Run it, or hand it over
Monitoring live, alerts routed, upgrade dates on the calendar, runbooks written. Then either we hold the pager or your engineers do, and the handover pack is identical either way.