Service 03Resilience

High availability.

Environments designed to keep running when a server, disk or site fails. Linux, Windows and virtualised platforms, built around your recovery targets.

Data hall corridor lined with server racks
Fig. — Data hall03

Scope

What the service covers
01

Clustered hypervisors

Proxmox VE, VMware and Hyper-V clusters with live migration and automatic failover.

02

Replicated storage

Ceph, ZFS replication and shared storage with no single point of failure.

03

Database clusters

PostgreSQL, MariaDB and SQL Server replication with tested failover.

04

Load balancing

HAProxy, nginx and keepalived in front of redundant application servers.

05

Multi-site

Services replicated between locations, including our London, European and North American regions.

06

Failover testing

Failover exercised on a schedule, so it works when it is needed.

Method

4 stages
1

Define

Recovery time and data loss targets agreed.

2

Design

Architecture built to meet those targets.

3

Build

Deployed and documented, with failover tested before go-live.

4

Exercise

Failover tested on a fixed schedule.

You receive

  • +Resilience design against agreed RTO and RPO
  • +Architecture diagrams
  • +Failover runbooks
  • +Failover test reports

Tooling

Proxmox VEVMware vSphereHyper-VCephZFSPacemakerkeepalivedHAProxyPostgreSQLSQL Server
Available as
  • Managed
  • Project
  • On call
Other services