Backup and Business ContinuityMay 27, 2026Serdar YAMAN5 min read

Scenario Analysis: a Disk Failure in Tax-Filing Week — What RTO and RPO Mean in Practice

Scenario Analysis: a Disk Failure in Tax-Filing Week — What RTO and RPO Mean in Practice

TL;DR: RTO is management's accepted answer to "how fast must we be back"; RPO to "how many hours of data can we afford to lose". In this representative case the same failure cost 3 days plus a full day of data in the target-less year — and 2.5 hours plus a 45-minute window in the year with targets. RAID is not a backup, monitoring is early warning, and drills turn recovery time from a guess into a measurement.

This article is a composite, anonymised account of recurring real-world cases.

Picture a twenty-person accounting practice: filing week, everyone on the system, the accounting software's database on the office server. Tuesday 09:20 — the application freezes, and a steady beep rises from the server chassis: disk failure. We will watch this scenario in two different years of the same firm — because the difference between them is not hardware but the existence of two acronyms: RTO and RPO.

Year One: the Day Without Targets

That year, "we have backups" was true but incomplete: a nightly backup went to a single NAS beside the server and had never seen a restore test. The chronology was merciless: failure in the morning, "maybe it recovers" attempts until noon, a hunt for a new disk in the afternoon. The next day the disk went in and the restore began — first surprise: the backup file's last three nights were corrupt; the healthy copy was four days old. Second surprise: nobody knew how many hours the restore would take; at that database size, it took eight. When the system opened on the third afternoon, four days of records had to be re-keyed by hand — in filing week, after hours. The total bill: three business days of downtime, days of overtime, and one lost client.

Two Questions at the Management Meeting

In the review that followed, the technical team put two questions in front of management — the RTO/RPO framework itself:

  • "When this system goes down, within how many hours must we be back?" The answer was debated and written: 4 hours normally, 2 hours in filing periods (RTO).
  • "At most how many hours of data entry can we write off?" The answer: 4 hours normally, 1 hour in peak periods (RPO) — because re-keying one day of records meant two people's full shifts in this office.

Those two numbers became a technical shopping list: an arrangement capable of hourly backups, a second copy under the 3-2-1 rule, a database-native backup method, monitoring that flags disk health in advance, and a quarterly drill that measures the restore time.

Year Two: the Same Failure, a Different Day

TimeEventWhat made the difference
09:20Disk failure; server slows, alert firesMonitoring reported the fault before user complaints did
09:25Plan active: last hourly backup verified (the 09:00 copy is sound)RPO window: at most 20 minutes of entries at risk
09:40Decision: restore to standby hardware without waiting on repairsWith a 2-hour RTO there is no "maybe it recovers" waiting
11:10Database opened on the standby server; application connections switchedThe 80-minute restore measured in drills finished in 90
11:50Team working; the 09:00–09:20 entries completed against a checklistDowntime: 2.5 hours · data window: 45 minutes

Same office, same software, a comparable failure. The difference: two numbers in writing, and an infrastructure built to meet them.

Four Lessons from the Scenario

  • RAID is not a backup: year two's server had RAID as well, but the plan did not lean on it — RAID tolerates a disk failure, not deletion, corruption or a double failure. Indeed the scenario's failure turned serious when a second disk signalled during the RAID rebuild.
  • An untested backup is hope: year one's "last three nights corrupt" surprise cannot happen under a regular drill regime — the corruption shows in the first drill.
  • RTO/RPO are seasonal: an accountant's January and July are not the same. A calendared rule that tightens backup frequency in peak periods avoids needless cost the rest of the year.
  • Hardware age is a plan input: the failed server was six years old; with refresh signals watched, the failure could have become a planned replacement. Year two's monitoring regime was that lesson's product.

Set Your Own Numbers

The template applies to any business: list your critical systems, answer the two questions with management for each, and measure with one drill whether your current backup arrangement actually delivers those answers. In most businesses the first measurement reveals an uncomfortable gap between target and reality — and seeing that gap before failure day is this article's entire purpose.

Yamanlar Bilişim's Approach

Our backup projects begin with those two numbers, before any technology choice; the installation is sized to the RTO/RPO targets, quarterly drills measure and report the durations, and seasonal criticality — filing periods, high season, year-end — is written into the calendar. For maintenance-agreement customers, disk health and backup success are daily monitoring items.

FAQ

Frequently Asked Questions

Who sets RTO and RPO — IT or management?

Together: management knows the business cost of downtime and data loss; IT knows which target is achievable at which budget. A target set unilaterally is either unmeetable or needlessly expensive; these two numbers are, at heart, a business decision.

Won't hourly backups strain the server?

Modern methods (incremental backups, snapshots) copy changed blocks, not full images; correctly configured, hourly backups go unfelt. What strains a system is not frequency but the wrong method — databases in particular need a database-native mechanism, not file copies.

Is standby hardware (a second server) mandatory?

It depends on your RTO: a 2-hour target needs a ready restore destination — a second physical server, capacity in a virtualisation cluster, or a warm copy in the cloud. At a 24-hour RTO, even "procure hardware on failure day" can be defensible — provided it is written down and tested.

Our developer says "I copy the database every night" — enough?

Test it with three questions: is the copy on a different device; when was a restore last attempted; and a 23:00 copy means an 18-hour loss at a 17:00 failure the next day — does that window fit your business? If even one answer is weak, "I copy it nightly" is not assurance; it is year one's story.

Doesn't a drill put the production system at risk?

A proper drill never touches production: the backup is restored into a separate environment (a test server or VM), opened and verified. The risk is not the drill — it is a first-ever restore attempted on the day of a real failure.

Share:
SY

Author

Serdar YAMAN

Yamanlar Bilişim Expert

Writes content on IT infrastructure, cybersecurity, and digital transformation at Yamanlar Bilişim. Get in touch for any questions.

Professional Support

Get help on this topic

Let's design the Backup and Business Continuity solution you need together. Our experts get back to you within 1 business day.

support@yamanlarbilisim.com · Response time: 1 business day