There are two kinds of restore. The one you do ten times a year — a deleted file, an earlier version — and the one you do once every three years, under pressure, with a managing director asking every ten minutes when it will be back.
The second cannot be improvised. Here is what you need to have checked beforehand.
The scenario, in the order it really happens
A server will not boot. Before even thinking about restoring, three questions arise, and the order matters:
- Is the hardware usable? If the disk is dead, the target will be another disk, or a virtual machine.
- Do you need to come back identically, or somewhere else? Rebuilding on the same hardware is the simplest; rebuilding into a VM is often the fastest when the hardware is unavailable.
- What is an acceptable point to return to? Last night's backup, or an earlier one if the incident has a software origin.
It is at the third question that you find out whether retention was thought through.
What actually gets in the way
Booting
Restoring a system's files is not enough to make it start. You need the system partition, the EFI partition or the boot sector depending on the mode, and a coherent partition table.
This is where "files only" backups show their limit: they contain the data, not the machine. A server rebuilt from a file backup requires a complete reinstallation of the operating system before you can recover anything at all — reckon on a day, not an hour.
Logical partitions
A disk partitioned in MBR with an extended partition contains a chain of EBR blocks: each logical partition points to the next. Restoring the partitions one by one without recomputing that chain produces a disk the operating system cannot read.
This is the kind of detail you only see on the day you need it. A restore tool must handle it on its own.
4Kn disks
Disks with 4096-byte sectors are becoming standard on recent servers. Restoring an image taken from a 512-byte disk onto a 4Kn disk — or the other way round — is not a simple copy. If your tool says nothing about it, test before you need it.
Resizing
The replacement disk is rarely exactly the same size. Larger, and it has to be extended; smaller, and you have to check the data fits. A serious restore plan refuses an impossible resize and explains why, rather than attempting it and leaving a corrupt volume behind.
The case where the hardware is not available
This is the most common situation in a small business: the server died on a Friday, the replacement arrives on Tuesday, and the business cannot wait.
The answer is P2V — restoring a physical machine into a virtual one. If you have an ESXi host with spare resources, the server comes back up there until the hardware arrives, then moves back to physical afterwards.
What to check before counting on it:
- Can the tool create the target VM, or does it have to be prepared by hand?
- Is the disk-by-disk mapping between source and target explicit?
- Is the VM stopped and restarted automatically during the operation?
- Can all of this be driven remotely, or do you have to be in front of the hypervisor?
The boot environment
When the machine will not start, the backup agent installed on its system will not start either. So you need a rescue environment: on Windows, WinPE.
Two questions to put to your tool:
- Does the agent run under WinPE, or is a separate recovery medium needed, to be built and maintained?
- Does the same backup storage serve, or do you have to have prepared a specific image?
The answer determines whether the restore starts within ten minutes or within three hours.
The checklist, to run once per client
To go through calmly, not on the day of the incident.
| Point | Checked |
|---|---|
| The system backup includes the partition table | ☐ |
| A full restore has been tested, and timed | ☐ |
| The measured time is compatible with what the client expects | ☐ |
| The recovery key is kept somewhere other than the machine being backed up | ☐ |
| The boot medium or environment is available and up to date | ☐ |
| A virtual fallback target exists, with spare resources | ☐ |
| The procedure is written down, and readable by somebody other than its author | ☐ |
The last line is the one people skip. It is nonetheless decisive: on the day of the incident, the person who set up the backup is on holiday.
The test that is worth every audit
Take a test workstation. Wipe it. Restore it entirely, from the backup, without consulting the vendor's documentation. Time it.
In half a day you will learn what no product sheet says: the real speed, the steps that jam, and the quality of support when you call because something is not going as planned.
BeBackup covers the file, the partition, the whole disk, P2V to ESXi and booting under WinPE, from the same backup and the same console. See the restore page or ask for a demo.