04 66 72 51 22

Monitoring fifty clients without spending your days on it

A provider who starts backing up their clients goes through three phases. The first five clients: all is well, the reports get read. The next twenty: they get skimmed. Beyond that: they stop being read at all, and you find out on the day a client asks for a restore from a backup that stopped running six weeks ago.

This is not a discipline problem. It is a design problem.

Why the daily email report does not hold up

The default model in many solutions: every backup sends a report, you receive them all, you look.

At fifty clients and three machines each, that is a hundred and fifty emails a day. A few reflexes appear naturally: a filter rule files them in a folder, then you consult them "when there is time", then not at all.

The design flaw is right there: the daily report tells you what happened, not what failed to happen. A backup that fails sends an error email. A backup that does not start at all — machine switched off, agent uninstalled during a hardware swap, service stopped — sends nothing. Silence is indistinguishable from calm.

The alert that counts: the run that did not happen

The turning point comes when you invert the logic: instead of waiting for an error signal, you watch for the absence of an expected signal.

A useful "did not start" alert has three properties:

It is set per machine, not globally. A laptop that backs up three times a week and a server that backs up every night do not share a threshold. A single threshold produces either noise or blind spots.

It takes multiple servers into account. If you replicate, or if you have several backup servers, a machine that backed up elsewhere must not raise an alert. A solution that queries the other servers before alerting saves you from false alarms — and false alarms are what kills monitoring, because you end up ignoring all of them.

It can be consulted, not only sent. An email gets lost. A dashboard that shows current states is read in thirty seconds in the morning.

The dashboard as a starting point

The aim in the morning is not to check everything: it is to know from one screen whether there is anything to do.

In concrete terms, a status that distinguishes between situations: error, warning, did not start, locked, backup running, restore running, replication running, integrity check, Microsoft 365 backup. Each calls for a different action — or for none.

Filtering by scope matters as much as the statuses themselves. A technician looking after ten clients should not see all fifty: the noise makes the view unusable.

Permissions per technician

This is a subject people put off and that ends up blocking them.

As long as you are alone, everybody is an administrator. With three technicians, the questions arrive: should a first-line technician be able to restart a backup? Yes, that is the point. Change a retention rule? No, that can delete versions. Restore at a client's site? It depends.

Fine granularity — of the order of fifty named permissions — lets you answer those questions once, then hire without reopening the debate. The right question to ask: what is the most destructive action this profile can trigger?

The reports that go out to the client

Two realities coexist: what you monitor, and what your client sees.

The client does not need the hundred and fifty daily lines. They need, at regular intervals, a readable page saying what is protected, over what depth, and whether all is well. Under your name, not the vendor's — you are the one providing the service.

It is also the document you bring out at the annual review when somebody asks what the "backup" line on the invoice is for.

The restore test, to be scheduled as a task

Everything above monitors that the backups are happening. Nothing guarantees they can be restored.

The restore test is the only check that counts, and it is also the first one abandoned for lack of time. Two ways to sustain it:

  • Automatic integrity checking on the server side, which reads the data back and compares the fingerprints after each backup or according to a per-machine rule. It does not replace a restore, but it detects silent corruption without occupying anybody.
  • One real restore per client per year, scheduled as a maintenance task. A file, a folder, and once every two or three years a full system at the critical clients.

The second point has a commercial benefit that is often underestimated: a dated restore report is the strongest argument you have against a competitor who will produce none.

A routine that holds

FrequencyWhat you do
Every morningA glance at the dashboard, filtered on the anomalies
Every weekDealing with silent machines and recurring warnings
Every monthA review of volumes, quotas and remaining disk space
Every quarterOne control restore, documented
Every yearA scope review with the client: what has changed in their IT

The last line is the one that pays. A client who has added a server, a line-of-business application or a share without telling you has a hole in their protection — and discovering it during a review, rather than during a disaster, is how you stay their provider.


BeBackup was designed by an IT service provider for their own estate: multi-client console, per-machine "did not start" alerts, fine-grained permissions and white-label reports. See the providers page or ask for a demo.

Worth reading too

Contact the BeBackup team

Would you like to know more about our BeBackup backup solution?