NMS - New Media Service GmbH
Der Hegau Tower in Singen, Sitz der NMS

How long may an outage last? The question before backup

Recovery time (RTO) and recovery point (RPO) are commercial figures, not technical.

News

AI-generated

The question of backup is almost always asked as a technical one. The more important question is commercial: how long may an outage last, and how much work may be lost? Those two figures are the starting point for which backup is needed and what it may cost. A third requirement comes on top, and more on that below.

Two figures everything hangs on

In technical language they are called RTO and RPO. Behind them sit two plain questions that can be answered without technical knowledge. Without the answer the technology cannot sensibly be chosen.

  • Recovery time (RTO, Recovery Time Objective): how long may it take until work can resume? An hour, half a day, three days? Between those answers lie different designs and considerably different costs.
  • Recovery point (RPO, Recovery Point Objective): how much work may be lost? With one backup per night, a whole working day may be missing in the worst case. Anyone who cannot accept that needs more frequent recovery points.

The relationship is uncomfortably simple: making either figure smaller costs money. The question is therefore not how small they can be, but how small they have to be. That depends on what an hour of standstill costs.

Neither figure belongs to IT alone. The need is stated by the departments working with the system; it is decided and owned by management. IT works out what follows technically, and says what of that is feasible.

Why the answer differs per application

An ERP system controlling dispatch tolerates less downtime than an archive someone looks into once a quarter. A single specification for all systems means either paying too much or saving in the wrong place.

  • Systems nobody can work without: short recovery time, frequent recovery points
  • Systems that can wait a day: standard backup, one recovery point per night
  • Archives and repositories: long retention and verified readability. How often a backup runs does not depend on how often someone looks in, but on how often something changes. And: backup and legally compliant retention are two things. A backup is not an archive, and an archive does not replace a backup.

From the figure to the design

Once both figures are settled per application, the technology follows almost by itself. A recovery time of hours requires the backup to be quickly reachable. Tapes from a distant offsite vault usually do not meet that, because the journey there adds to the restore. What counts is always the measured total duration, not the storage medium. A recovery point of minutes requires continuous backup rather than a nightly run.

A third property comes on top, and in an emergency it decides the other two: the backup has to be immutable. In a ransomware attack the backup is a primary target. Anyone who can delete it or encrypt it along with everything else makes every promised recovery time worthless. There are three distinct means against that, and they complement rather than replace each other: write protection for a defined period that applies to administrators too; separate administration, so that one compromised account cannot reach both; and a copy that is not attached to the network. Only the first is true immutability; the other two reduce the risk without removing it. „Quickly reachable" and „not manipulable" are two requirements in any case, not one.

The reverse also holds: where a day of downtime is tolerable, there is no need to buy the most expensive solution. The figures therefore protect in both directions, against too little and against too much.

The test nobody runs

A restore that has never been rehearsed is an assumption. An emergency is the worst conceivable moment to discover that the backup ran but cannot be restored, or that it takes longer than promised.

A test does not have to be large to show something. A single file, a single mailbox, a single server, with a stopwatch: that shows whether the backup is readable at all and how long one individual case takes. What it does not show is the recovery time of a whole service. That involves dependencies, sequence and network, and it can only be measured in a larger rehearsal. A regular small test and a rare large one are two different things, and both are needed.

For a measured duration to mean anything it needs a defined start and a defined end, from triggering the restore to the point where work can actually resume, not to the end of the copy job. And the recovery point brings a second question: how old was the state that came back?

Anyone writing the tests down (date, scope, measured duration, errors encountered) gets two things: a sound basis for their own planning, and a record for when someone asks. That does not rule out surprises entirely; it makes them smaller.

What a monitored backup has to deliver is on our page about cloud backup; the conceptual side (retention, strategy, verification) is covered under data protection and backup. We are glad to talk through the right times for your organisation in a first conversation, at no cost.