Your Backup Job Ran. Did It Actually *Work*?

Backup, Snapshotting & Disaster Recovery

Your Backup Job Ran. Did It Actually *Work*?

Technical Briefing | 9/29/2026

Every sysadmin worth their salt sets up backups. We write the scripts, schedule the cron jobs, configure the `rsync` or `tar` commands, and then, most of the time, we let them run. The logs are green, `exit 0` greets us every morning. We tell ourselves it’s fine. It’s a checkbox, right? And then one day, the call comes, the server’s down, and you discover the horrifying truth: your ‘working’ backups are completely useless.

The Lie We Tell Ourselves About Backup Logs

We’ve all done it. You see `rsync` completed, or `tar` compressed everything, and you move on. But that success message only tells you the *tool* ran without errors. It says absolutely nothing about the *integrity* of the data it copied. Think about it: `rsync` will happily copy a corrupted file. A `mysqldump` might complete, but if your database was in an inconsistent state during the dump, you’ve just backed up garbage. I’ve seen this silently fail in production environments more times than I care to admit. The backup was there, it was the right size, but when we tried to restore, it choked on missing foreign keys or malformed config. That’s when you learn the hard way that a broken backup is often worse than no backup at all; it’s a false sense of security that burns you when you need it most.

Beyond ‘Success’: Real Verification

So, how do you actually know your backups are good? You test them. Not just once a year in a full-blown DR drill, but regularly, automatically, with real restore operations. You don’t need to rebuild your entire infrastructure daily, but you absolutely need to verify critical data points. The simplest approach for filesystem backups is to grab a critical file, restore it to a temporary location, and run some basic checks. Does it exist? Is it non-zero in size? Can an application parser validate it? Here’s a very basic script concept that gets you started.

#!/bin/bash
# Simple script to verify a critical file's integrity from a backup
#
# Usage: ./verify_backup.sh <backup_root_directory> <original_file_path_relative_to_root>
# Example: ./verify_backup.sh /mnt/backups/serverX_daily /etc/nginx/nginx.conf


BACKUP_ROOT="$1"
ORIGINAL_FILE_PATH="$2"


if [ -z "$BACKUP_ROOT" ] || [ -z "$ORIGINAL_FILE_PATH" ]; then
echo "ERROR: Usage: $0 <backup_root_directory> <original_file_path_relative_to_root>"
echo "Example: $0 /mnt/backups/prodserver_daily /etc/nginx/nginx.conf"
exit 1
fi


FULL_BACKUP_PATH="${BACKUP_ROOT}${ORIGINAL_FILE_PATH}"
TEST_RESTORE_PATH="/tmp/backup_verify_$(basename "$ORIGINAL_FILE_PATH")_$$"


if [ ! -f "${FULL_BACKUP_PATH}" ]; then
echo "ERROR: Expected file '${FULL_BACKUP_PATH}' not found in backup. Is the backup path correct?"
exit 1
fi


# Simulate a restore: Copy the file to a temporary location
cp "${FULL_BACKUP_PATH}" "${TEST_RESTORE_PATH}"
if [ $? -ne 0 ]; then
echo "ERROR: Failed to copy '${FULL_BACKUP_PATH}' to '${TEST_RESTORE_PATH}' for verification." >&2
exit 1
fi


# Basic sanity check: file exists and isn't empty
if [ ! -s "${TEST_RESTORE_PATH}" ]; then
echo "ERROR: Verified file '${TEST_RESTORE_PATH}' is empty or missing content." >&2
rm -f "${TEST_RESTORE_PATH}"
exit 1
fi


# Add application-specific validation here for true confidence.
# For example, for an Nginx config:
# sudo nginx -t -c "${TEST_RESTORE_PATH}"
# For a common text config, maybe grep for a known directive:
# if ! grep -q "listen 80" "${TEST_RESTORE_PATH}"; then
# echo "WARNING: Nginx config '${TEST_RESTORE_PATH}' missing 'listen 80'." >&2
# fi


echo "SUCCESS: Backup of '${ORIGINAL_FILE_PATH}' from '${BACKUP_ROOT}' verified successfully."
rm -f "${TEST_RESTORE_PATH}"
exit 0

That example is basic, of course. For databases, verification gets trickier. A logical dump from `pg_dump` or `mysqldump` isn’t transactionally consistent unless you take specific steps, like locking tables or using a read replica. The right approach here is often to use filesystem-level snapshots (LVM, ZFS) to get an atomic point-in-time copy, then back that up. Or, even better, leverage database-specific tools like Percona XtraBackup for MySQL or `pg_basebackup` for PostgreSQL, which understand transactional consistency. Once you’ve restored, even to a small canary VM, run some application-level smoke tests: hit an API endpoint, log into a dashboard, or run a simple `SELECT COUNT(*)` to ensure data is present and queryable. This is where automation pays off, giving you a daily, concrete check rather than just hopeful thinking.

Ultimately, your backup strategy isn’t complete just because data is being copied. It’s only complete when you can confidently say you can restore that data, quickly and reliably. This isn’t about avoiding a headache next year; it’s about minimizing the impact of the outage that’s *going* to happen, eventually. Spend a little time on verification now. Your future self, hip-deep in a production emergency, will silently salute you for it.

Linux Admin Automation  |  © www.ngelinux.com  |  9/29/2026

0 0 votes
Article Rating
Subscribe
Notify of
guest

0 Comments
Newest
Oldest Most Voted