Software Development

Website Backup Strategy: From File Copies to Restore Testing

Test your website's recovery plan with RPO, RTO, database consistency, independent copies and an isolated restore drill.

Website Backup Strategy: From File Copies to Restore Testing

The archive opens, the database imports and the home page loads. Even so, the product photos are missing, the administrator cannot sign in and order notifications do not fire. Restoring a site is a broader problem than keeping an archive intact.

The real output of a website backup plan is not a folder full of copies. It is a recovery process you can actually run when you need it. You have to decide together which data is protected, how much loss is acceptable and who brings the system back up.

This article proposes a workable plan using a hypothetical order-taking site. The durations given are calculation examples only; they are not a performance commitment for any particular infrastructure, nor a WebWizz customer result.

Define acceptable loss and acceptable downtime separately

RPO expresses, in units of time, how much data loss is acceptable after recovery. RTO is the target for returning the service within an acceptable period. One being small does not mean the other is small too. AWS's recovery planning guidance covers how these targets should be set from business requirements. AWS: Disaster Recovery Planning.

Suppose our example business accepts losing at most one hour of order data. A backup taken once a day cannot meet that target on its own. Yet writing an hourly schedule is not enough either: failed copies and unverified transfers can leave the last usable recovery point older than the schedule suggests.

In another assumption the downtime target might be four hours. Fit reaching the right person, preparing a clean environment, downloading the backup, restoring, verifying and reopening traffic inside that window. Measuring only the time it takes to unpack an archive is misleading.

Discuss the target with the business owner using real records

“One hour of data loss” can sound abstract. Work out what it would mean in that business's own records: how many orders re-entered, which applications lost, which transactions reconciled by hand. Do not present estimates as precise costs before you have measured them.

It is perfectly normal for a content site and a high-transaction order system not to need the same schedule. Writing targets per service makes the trade-off visible between applying the most expensive protection to all data and under-protecting the data that matters.

Inventory everything a restore depends on

Having the application repository does not mean the site's whole state is protected. User-uploaded photos, database records and configuration can live in entirely different places.

In the example inventory, track these items separately:

  • Databases, roles and any extensions the application requires.
  • User uploads, documents and their access permissions.
  • Application version, dependency definitions and database schema.
  • Securely stored configuration and decryption keys.
  • Scheduled jobs, queues and third-party service connections.
  • Domain, DNS and the procedure for reaching certificate management.

For each item, record its owner, where it is kept and its position in the restore order. Do not paste secrets into this inventory in plain text; describe how an authorised person reaches the secure source instead.

If the key to an encrypted backup exists only on the server you lost, the copy may be unusable in practice. Plan key recovery separately and limit access to the people whose role requires it.

Keep files and the database consistent with each other

An order's database record surviving while its attached document is missing shows how two individually successful backups can still be broken together. Choose a consistency approach that matches how the application writes data.

In PostgreSQL, pg_dump can take a consistent view of a single database while work continues. However, global objects such as roles and the application's file storage are not automatically part of that output. Separate dumps taken from different databases should not be assumed to share one common point in time either. PostgreSQL: SQL Dump.

Where more frequent recovery points are required, a base backup combined with WAL archiving and point-in-time recovery can be considered. That depends on preserving the chain of records it needs; keeping only a base backup does not provide the same capability. PostgreSQL: Continuous Archiving.

The point here is not to impose the same database method on every site. Choose the official backup mechanism of the system you run, document how it relates to your files, and test it by restoring both together.

Assess how independent your copies really are

When the live site and its backups share the same administrator account, the same disk and the same delete permission, you have created a common failure risk. One account being misused or compromised can affect both sides.

As you build the plan, weigh protections such as a different storage location, separate access rights and, where supported, copies that cannot be altered. Consider these in terms of the event you want protection from, not as a single button or product name.

For example, a site administrator might be able to delete content but not change the backup retention policy. The recovery owner's access should be testable on its own. Making sure the account and contact details you would need do not live only inside one provider's system, in case that provider is unreachable, is part of the operational plan too.

Run the restore drill in an isolated environment

Do not run recovery attempts directly on top of the live system. Prepare an isolated environment that sends no email, payment request or webhook to real customers. Disabling external integrations, or pointing them at safe test endpoints, keeps the drill from becoming an incident of its own.

AWS's backup validation guidance recommends testing both the data and the process through periodic restores. For our example site, that approach translates into the following running order. AWS: Periodic Recovery Testing.

  1. Record the drill's target recovery point and its start time.
  2. Verify authorised access and that the required keys are available.
  3. Install a compatible application version and the data into the isolated environment.
  4. Compare baseline record counts, relationships and sample files.
  5. Exercise administrator sign-in, product display and a test order flow.
  6. Report the gaps, the time spent and the procedures that need correcting.

Verifying a file's checksum is useful for transfer integrity; on its own it does not show that business records are meaningful or that the application runs. Opening a sample order together with its line items, totals and attached documents gives you a different kind of check.

Keep failures visible

A backup job starting and a usable copy existing are two different states. Define monitoring for the timestamp of the last successful copy, unexpected changes in size and files that could not be transferred. Include whether the alert actually reaches the responsible person in your testing.

Do not set retention by disk usage alone. Weigh how long a problem might go unnoticed, the nature of the data and your organisation's retention obligations together. The same corrupted data being backed up over and over is exactly why older, sound recovery points matter.

After a security incident, opening any copy quickly is not enough. Without identifying a clean recovery point and a trustworthy working environment, you can restore the same problem. The recovery plan has to run alongside the incident investigation and any access changes it requires.

Implementation checklist

  • Set RPO and RTO targets separately, with the business owner.
  • Complete the inventory of files, databases, versions and keys.
  • Examine the failure points your backups share with the live system.
  • Choose the consistency method that fits the infrastructure you use.
  • Track the age of the last usable copy.
  • Disable real external notifications in the test environment.
  • Verify records and core user journeys together.
  • Feed drill results back into the recovery procedure.

To review your current backup plan with WebWizz, tell us about your infrastructure and your acceptable downtime target. We can start the conversation with the business process that has to come back, rather than with storage capacity.

Frequently asked questions

Is a server snapshot enough on its own?

It depends on the consistency and recovery characteristics of the infrastructure. There may be external files or services the snapshot does not cover. Judge its sufficiency by restoring your full inventory, not by the product name.

Does a “backup successful” message prove recovery works?

No. The message reports the outcome of a specific job. Key access, application compatibility and data coming back in working order all have to be tested separately.

How many days of backups should be kept?

There is no single number that holds for every site. Time to notice a fault, how often data changes, your organisation's obligations and cost all have to be weighed together.

How often should the drill be repeated?

Set a regular interval based on your risk level, and retest the plan after any significant change to the database, encryption, storage or application version. An old successful drill does not automatically validate a new setup.

Comments (0)

Join the discussion

You must be logged in to post a comment and interact with this post.

Log In

No comments yet. Be the first to share your thoughts!

WebWizz Newsletter

Keep up with what we build

Get newly published blog posts and projects by email. You will only receive the update types you select.

Update preferences