Disaster Recovery and Business Continuity in Kenya
CloudSpinx designs, builds, tests and runs disaster recovery for businesses in Kenya and across East Africa: a business impact analysis that produces recovery numbers a board can approve, a second-site DR design, replication and immutable backup, runbooks somebody has actually run, and the evidence pack a CBK or SASRA examiner asks for. We work from the instruments rather than the commentary, so you get told what genuinely binds you and what does not. Cost follows your recovery target, your data volume and how much of the estate has to come back first, which is why we scope every engagement instead of publishing tiers.
Who we build for
- 14organizations, from ISPs and payment platforms to a national regulator
- 6flagship engagements published in full, with the numbers counted
- 4thof all contributors to the open-source payment switch national systems run on
Everything in Our Disaster Recovery Service
Every engagement covers the full scope: no hidden extras, no upselling.
Business impact analysis
Which activities are genuinely mission-critical, what each one depends on, what an hour of it being down costs, and who owns it. PG/14 requires recovery solutions to be based on BIA information, and this is the document that produces every number after it.
RTO and RPO definition
A working session that turns "as fast as possible" into a recovery time and a tolerable data loss per critical activity, written down, costed, and taken to the board for approval before anyone buys replication technology.
DR site design and selection
Where the recovery facility sits, on which grid, on which telecom path, and why. We shortlist Nairobi and upcountry colocation options against the geography rule and your latency budget rather than against a brochure.
Replication architecture
Storage-level, hypervisor-level or database-level replication, chosen against what your core application vendor supports in writing. Recovery point measured, not assumed, and the failback path designed at the same time.
Immutable and offsite backup
Object storage with a retention lock nobody can shorten, verified restores, a copy outside the primary blast radius, and retention laddered against what the business and the regulator each require you to keep.
Recovery runbooks
The declaration threshold, who is allowed to invoke, the call tree, the promotion sequence per system, the verification checks, and the failback. Written so the person on duty at 02:00 can follow it without ringing the architect.
DR testing and live failover drills
Walkthroughs, technical tests and at least one live run a year, with the pre and post test reports PG/14 requires as the actual deliverable. We break something on purpose while you watch, in a window you choose.
Regulatory evidence pack
The BIA, board-approved recovery targets, test reports, restoration records naming who restored what, the incident register that demonstrates the 24-hour clock was met, and the annual independent BCP review.
CBK outsourcing approval support
PG/16 makes buying DR a material outsourcing requiring prior CBK approval. We draft the rationale, the risk matrix, the retained-control description and the provider detail that the submission has to contain.
Cloud disaster recovery design
AWS Cape Town, Azure South Africa North or Google Johannesburg as the third copy and as recovery compute where the application vendor supports it, with the cross-border transfer paperwork handled rather than ignored.
Managed DR and drill retainer
We hold the runbooks current, watch the replication and verify jobs, run the drills on a calendar, and produce the evidence every cycle. Somebody has to keep doing this, and it is the part that quietly stops.
Ransomware recovery readiness
The recovery path that assumes your admin credentials are already compromised: locked copies, an isolated rebuild network, a clean-room restore order, and the honest answer on how long a full rebuild takes.
Technologies we use
What a defensible DR design looks like
Three of the drawings that come out of a continuity engagement: the two-facility shape that satisfies both the geography rule and the in-country rule at once, the anatomy of a recovery clock, and the failure a replicated pair does not survive. Site addresses, circuit sizing and your actual recovery numbers come out of the business impact analysis.
A two-facility DR site design in Nairobi
The shape that satisfies CBK/PG/14 geography and CII regulation 28 in the same build, which is why we recommend it more often than anything else. The detail that decides whether it works is not the replication technology. It is whether the second facility sits on a different power grid and a different telecom path, because a recovery site that fails with the primary is a line item, not a control.
- replication and backup path
- second provider circuit
- user path after a declaration
What RTO is made of, and where the RPO loss sits
Recovery time is four intervals, and only one of them is technology. The interval almost nobody measures is the second one, the wait while somebody senior decides to declare a disaster, and it is routinely the longest. Put a named decision-maker and a declaration threshold in the runbook and you shorten the recovery more than any replication product will.
Why replication is not a backup
The figure we draw earliest in most engagements, because it is the misunderstanding that costs the most. Replication protects you from a dead building. It does nothing about a bad write, a dropped table or an encryption run, all of which arrive at the standby on schedule. The control that survives those is a copy nobody, including us and including whoever holds your admin credentials, can delete before its lock expires. The number that matters on that copy is retention, not RPO.
Which recovery shape do you actually need?
Four shapes, in ascending order of what they cost and what they defend. Most Kenyan organizations we assess are on the first one and believe they are on the third. Read the bottom line of each card first, because that is the part an examiner reads.
Tape, or a drive in a drawer
A nightly job to removable media kept in the same building, or in a branch cupboard. Nobody has completed a full restore in living memory.
Fails on three counts: geography, tested recovery, and BIA-derived numbers. It is a filing habit rather than a control.
One site, plus an immutable offsite copy
Production stays where it is. The defensible part is a locked copy somewhere else with a written, rehearsed restore procedure against it.
Recovery in days, honestly stated. The cheapest shape we will put our name to, and the right answer for plenty of businesses.
Two facilities, replicated
Primary in one Nairobi facility, standby in a second outside the CBD on another grid and another telecom path, immutable copy behind both.
Satisfies PG/14 and reg 28 together. It also survives a designation event without a redesign, which is the part people miss.
Cloud as the third copy
Cape Town or Johannesburg for the offsite copy, and for recovery compute where your core application vendor supports it in writing.
Fine as a copy, wrong as a primary for a system that may be designated: that turns an engineering choice into a Form CMCA 3 application.
One thing decides more of this design than any regulation: your core application vendor's supported recovery topology. A beautiful architecture the vendor will not support is an audit finding waiting to happen, so we ask for that document in writing before drawing anything. If they cannot produce it, that is itself worth knowing early.
What disaster recovery actually costs in Kenya
Almost none of the cost is the technology. What you pay for is a second place to run, the link between the two, and the people who keep proving it works. That is why we scope rather than publish tiers: a lender whose teller network must be back inside an hour and a distributor who can trade off paper for a day are buying different things even when the servers look identical. Give us the recovery target and the data volume and we come back with one figure and the reasoning behind it.
- Recovery time: the single biggest driver. Hours are affordable, minutes cost multiples, and seconds usually means changing the application rather than buying more infrastructure
- Recovery point: how much data you can lose sets the replication design, and continuous replication over a long link costs more than an hourly one
- How much of the estate comes back first: recovering four critical systems is a fraction of the price of recovering ninety, which is exactly what the BIA is for
- Second-site choice: rack, power and cross-connect in a second facility, or space you already hold in a branch or a regional office
- The link between the sites: replication bandwidth is often the line nobody budgeted, and we size it against your real change rate, alongside your WAN and circuit design
- Data volume and retention: the immutable copy is priced per terabyte per month, and the retention ladder decides that number
- Licensing at the recovery site: passive failover rights vary by vendor and Windows or SQL Server standby is frequently the largest single line
- Testing and evidence: annual drills and the report pack are real recurring work, and they are what an examiner actually inspects
Does CBK require your data to stay in Kenya?
No. There is no in-country storage requirement anywhere in the Prudential Guidelines, the 2017 cybersecurity guidance note or the 2019 guideline for payment service providers. We searched the full text of all three, and you can too from the CBK legislation index. What CBK imposes is different and stricter in one respect: prior approval for material outsourcing, and unimpeded access for the regulator and your auditors wherever the workload physically sits. The test in PG/16 4.7.5 is supervisability, not geography. Most of the residency advice circulating in this market argues from the Data Protection Act alone, which is an incomplete map and produces the wrong architecture.
Where the in-country rule actually lives
Regulation 28 of the Critical Information Infrastructure and Cybercrime Management Regulations 2024 says an owner of critical information infrastructure shall ensure the infrastructure on which critical information is domiciled is located in Kenya, and anyone wanting it elsewhere applies on Form CMCA 3, with the Committee consulting the National Security Council before deciding. That is the real constraint, it is not CBK, and it is published in full on Kenya Law. Banking and savings services sit among the gazetted critical sectors, so the question applies to the financial sector as a whole. Whether a specific institution's systems are designated is a question of fact to confirm with NC4 and your regulator, and we will not tell you otherwise across a table.
The Data Protection Act asks for less than people think
Regulation 26 of the Data Protection (General) Regulations 2021 requires in-country processing for a short list of state-interest purposes, and one of its limbs catches systems protected under the Computer Misuse and Cybercrimes Act. Where it does bite, it is satisfiable by keeping at least one serving copy in a data center in Kenya rather than by moving the primary, and that distinction is worth real money on an architecture. When personal data does leave, the ODPC treats cloud processing abroad as a cross-border transfer and expects a DPIA, transfer documentation and a sub-processor chain that survives inspection. We cover the same ground for non-regulated buyers in the data protection rules across East Africa.
Where your DR site can sit: what CBK/PG/14 requires
PG/14 is more prescriptive about geography than most people expect, and it is the requirement Kenyan institutions most commonly fail on paper. Office, data center or server room recovery must not be in the same building or close to normal business operations. If the head office or back office is inside the city limits, the recovery facility should be outside the city center, and the guideline expects a second power grid and separate telecom infrastructure so that one outage cannot take both. Recovery sites need generator, UPS and adequate fuel. Shared recovery sites need SLAs that spell out the terms, and a syndicated solution must not be shared between institutions whose normal operations sit close to each other, because a city-wide incident takes them all at once.
- A second facility across town on a different grid satisfies the rule and keeps replication latency low enough for synchronous options to stay on the table
- Upcountry or Mombasa gives you real geographic separation and a longer round trip, which pushes the design toward asynchronous replication and a larger recovery point
- A branch office is legitimate for the immutable copy and rarely adequate as a recovery site, because power, cooling and physical access are not controlled
- Two facilities in Nairobi that peer at the same exchange avoid hairpinning replication traffic through Europe, which is a real failure mode for badly chosen pairs
- Whichever you pick, the DR facility needs the same access control and monitoring as production, and the security controls travel with the data rather than staying at head office
Setting an RTO and RPO an auditor will accept
The common failure here is buying the technology first and back-filling the numbers, which produces evidence an examiner rejects even when the technology works perfectly. PG/14 requires an RTO and an RPO per mission-critical activity, derived from a business impact analysis, with dependencies, system owners and the named technical person responsible recorded alongside them. So the order is fixed: analyze the business, agree the numbers, get them approved, then design. A vendor who arrives with an architecture before those answers exist is guessing, and any experienced IT manager can tell.
Numbers the business will actually stand behind
Ask a department head how long they can be down and the answer is always zero. Ask what they would do for the first two hours without the system, and you get a real answer, usually involving paper, WhatsApp and a phone call to a supervisor. That workaround is what sets the honest recovery time, and writing it into the plan is what makes the number survive a board discussion. We run the session, put a cost against each hour of downtime per activity, and let the difference in price between one hour and four make the decision visible.
The interval nobody measures
Recovery time is four intervals: noticing, deciding to declare, promoting the standby, and verifying the application really works. Only the third is technology. The second is a human being with the authority to say this is a disaster, and in most organizations we test, it is the longest of the four by a distance. Naming that person, writing the declaration threshold, and rehearsing the phone call cuts more time off a real recovery than any product upgrade.
Replication, immutable backup and what each one defends
These solve different problems and buying one instead of the other is the most expensive mistake in this field. Replication defends you against a dead building, a dead node or a cut circuit. It does nothing about a bad write, a dropped table or an encryption run, because all three arrive at the standby inside the replication interval and now exist in two places. The control that survives those is a copy under a retention lock that nobody can shorten, including your domain administrator and including us. On that copy the number that matters is retention, not recovery point.
- Synchronous replication: no data loss, a hard limit on distance, and a primary that slows down or stalls when the link degrades
- Asynchronous replication: survives distance and a poor link, and costs you whatever was written since the last shipped change
- Application-level replication: usually the only option your core application vendor will actually support, which is why we ask for their document before choosing anything
- Immutable object copy: the ransomware answer and the audit answer at once, because reg 42 of the CII Regulations wants restoration events recorded, naming who restored what
- Verified restores: a backup that has never been restored is a hypothesis, and the verification job is what turns it into a control
Cloud disaster recovery from Kenya, and what latency does to it
There is still no hyperscaler region live in Kenya, so cloud DR from here means Cape Town, Johannesburg or further. Round trip from Nairobi to Cape Town measures around 84 ms on our own tests from a Nairobi connection, and Johannesburg lands near 81 ms. That is comfortable for replication, batch and object storage, and it becomes uncomfortable for interactive teller and branch work once an application makes several round trips per screen. Treat the cloud as the third copy and as recovery compute where your vendor supports it, and be careful with anyone whose latency table shows a hyperscaler region answering from Nairobi in single-digit milliseconds, because they measured an anycast front door rather than a region. The wider picture is in our comparison of the three clouds from a Kenyan starting point.
What we would not propose
Primary core banking for a Kenyan lender in Europe or the United States. That is 190 to 210 ms of round trip, a reg 28 application with a National Security Council consultation attached, and a supervisability argument we would be making on your behalf to your own regulator. Cape Town for the offsite copy and for recovery compute is a different and defensible proposition, and we will say which one you are actually asking for. Where the honest answer is a private platform in Nairobi instead, that is the cloud and infrastructure design conversation rather than a DR one.
DR testing and the evidence pack an examiner asks for
The deliverable in this field is evidence, not uptime. PG/14 sets a minimum of one BCM test a year, scaled up by criticality, and names technical tests, walkthroughs, live runs, simulations and integrated tests across dependent departments among the expected methods. Documented pre and post test reports are to be completed for all recovery testing, call trees tested at least quarterly, the plan reviewed and approved by the board annually, an independent review each year reported to the board, and activation of the plan reported to CBK within 24 hours. Design that paper trail into the runbook and it produces itself. Reconstruct it during an inspection and it shows.
- The business impact analysis, with dependencies, owners and the named technical person per critical activity
- Board-approved RTO and RPO per mission-critical activity, dated by the minute of the meeting that approved them
- Pre and post test reports for every recovery test, including the things that failed and what was changed afterward
- Restoration records naming who restored what and when, which regulation 42(2)(e) of the CII Regulations asks for and almost nobody holds
- An incident register that can demonstrate the 24-hour notification clock was met, and the 2-hour clock where a payment system is systemically important
- The annual independent review of the plan, with findings reported to the board rather than filed
Outsourcing disaster recovery is itself a regulated act
This surprises project teams late, and late is expensive. PG/16 lists business continuity and disaster recovery, along with data centers, facilities management and information system management, as material outsourcing. No institution shall outsource material activities without prior CBK approval, and the proposal goes in well in advance of the intended start date. The submission has to contain a rationale, the provider details, a risk matrix of risks and mitigations, the draft agreement, and a description of how you retain the ability to control and monitor the function. That is a document somebody has to write, so it belongs on the plan as a dated deliverable ahead of procurement, not as a form filled in after the contract is signed. Payment service providers have a tighter rule again: notify CBK of an intention to outsource at least 30 days before the agreement is executed, and give both the PSP and CBK the right to audit the provider.
- Materiality has a bright line in PG/16: an activity counts as material if it accounts for at least five percent of the institution's revenues or costs
- Accountability does not transfer. Outsourcing does not diminish your legal obligations, and you remain responsible for the provider's actions and for customer confidentiality
- Write the regulator and auditor access clause into the first draft of the contract, because retro-fitting it is what kills a signed cloud arrangement at compliance review
- Expect due diligence on us, and we are set up to survive it, including our own ODPC registration as a data processor and a data processing agreement
- PG/14 and PG/16 bind institutions licensed under the Banking Act by their own terms. For a deposit-taking microfinance bank the binding force is supervisory practice, and we will tell you which of the two you are dealing with rather than blurring it
When we tell clients not to build a DR site
We build these for a living and we still talk organizations out of one several times a year. Saying so is the difference between a consultant and a reseller with a diagram.
- Nobody has ever completed a restore from your existing backups. Fix that first. A second site protecting data you cannot recover is an expensive way to duplicate a problem
- Your recovery target is a day and your data fits comfortably in an immutable copy. Buy the locked copy and a rehearsed procedure, and put the difference into the support that stops incidents happening
- Your core application vendor will not support a replicated topology, and nobody has asked them. Ask, in writing, before spending anything
- The real exposure is a single internet circuit or one aging server, not a lost building. That is a resilience problem with a much cheaper answer, and we would rather sell you the cheaper answer
- You want the second site mainly to satisfy a checklist. A DR site with no tested runbook fails an inspection more visibly than having no site at all, because it proves you knew
What you own when we hand over
Everything, and nothing that only works while we hold a password. The engagement ends with the business impact analysis, the approved recovery targets, an as-built design of the shape you saw drawn above, the replication and retention configuration with the reasoning behind each choice, the runbooks, and the test reports from a drill your own people ran with us watching rather than the other way round. We train the duty rota on the declaration path, hand over the automation repository, and keep running the platform underneath only if you want us to. Plenty of clients take the build, the first drill and the documentation, then run the calendar themselves, which is a good outcome and we scope for it openly.
Tell us what has to come back first
Enough for a recovery view and a real figure rather than a range. Every field has an escape hatch, so answer what you know and leave the rest to the impact analysis.
Ready to discuss Disaster Recovery?
A 30-minute scoping call, free, and it commits you to nothing.
How Every Disaster Recovery Engagement Starts
Impact analysis and sequencing
A free first session, then the BIA: what is genuinely critical, what it depends on, what an hour costs. In parallel we confirm your license class and what actually binds you, because that order saves redesigns.
Targets, then design
Recovery numbers agreed and taken for board approval, the vendor's supported topology obtained in writing, and only then a written design: sites, replication, retention, runbooks and one scoped price.
Build, then break it
We build the recovery side, then prove it in front of you with a live failover and a real restore, in a window you choose. The pre and post test reports come out of that drill as deliverables.
Drills and evidence
Handover and training, then either you run the calendar or we do: replication watched, restores verified, drills run, runbooks kept current, and the evidence pack produced every cycle.