Tuning Proxmox Backup Server for space and disk writes
Deduplication and prune rules squeeze 2.24 TiB of nightly guest backups into 164 GiB, while cache exclusions and journal caps spare the disk
Overview#
The single-box homelab backs up every guest, every night, into Proxmox Backup Server. The datastore sits on the utank ZFS pool, and the garbage collection run that closed out the tuning reports where all of that lands:
| Metric | Value |
|---|---|
| Original data indexed | 2.24 TiB |
| Chunks on disk | 164.25 GiB (7.15% of the original) |
| Deduplication factor | 13.99 |
| Chunk count | 86,243 |
This post is the tuning behind those numbers. The trigger was growth: media and photo libraries that only gain weight, guests that cache and download on the side, and nightly backups that scale with all of it, all fighting for the same 14 TB Exos. One spinning disk wears with every write, and whatever the backups keep has to be stored and pushed offsite too, so this round treats the write budget, the space budget, and the offsite budget as one. What follows: what deduplication was already doing on its own, what prune schedules and garbage collection actually reclaim, which inputs I stopped backing up, and two write-wasters I removed from the backup container itself.
What deduplication does for free#
PBS does not store archive files. Every backup is split into content-addressed chunks, named by their SHA-256 digest and stored once in the chunk store; a snapshot is just an index of digests. At the end of the tuning, 86,243 chunks totalling 164.25 GiB served 420 indexes referencing 2.24 TiB of data.
The factor is almost 14 because the fleet is thirty-odd LXC guests cloned from a handful of Debian, Ubuntu and Alpine bases. Identical system files produce identical digests across guests and across nights, so each night’s snapshot mostly writes a fresh index plus a thin layer of changed chunks. Nothing had to be configured for this; it falls out of running many similar guests.
Prune, then collect the garbage#
Deleting a snapshot only deletes its index. The chunks stay, because other snapshots still reference them, so an untuned datastore grows forever. The nightly pipeline runs in three steps, in this order:
- the backup job dumps every guest into PBS (the schedule lives in the homelab post)
- a prune job applies
keep-last 2,keep-daily 3,keep-weekly 1 - garbage collection sweeps the chunk store
Garbage collection walks the chunk store, marks every chunk no index points at, and removes the orphans. The subtle part is the 24-hour grace rule: freshly unreferenced chunks are reported as pending removals and survive until the next day’s run. If a backup is uploading while an overlapping snapshot is pruned, both can legitimately reference the same chunks, and the grace period guarantees the running backup can never have a chunk pulled out from under it. The first cycle after the retune left 47.98 GiB in 20,198 chunks pending; a follow-up run removed 2.55 GB in 1,240 chunks, found zero bad chunks, and finished in 255 seconds.
The before and after of tuning the prune job: 315 snapshots (288 CT, 27 VM) holding 262 GB of chunks, down to 210 snapshots (192 CT, 18 VM) over 35 guests in 164.25 GiB. Restore depth is shallower, and that is the point. Every kept snapshot is chunks the disk writes, holds, and later cleans up, and payload the offsite leg carries too. Two hot copies, three dailies and a weekly is the window I actually restore from, so the retention stops there.
Shrinking the input with .pxarexclude#
Deduplication keeps one copy of everything, but a copy is still a copy. PBS walks the entire root disk of each LXC, so hidden caches and downloaded files end up in the chunk store no matter how similar they are. I swept the guests with ncdu -x / (the -x stays on the root filesystem and ignores bind mounts) and found the usual suspects:
| Guest | Found |
|---|---|
| a media downloader | 29 GB of downloaded files |
| a download manager | 6 GB of uv package cache |
| the photo library | 6 GB of node_modules and pnpm stores |
Each rule is one line in /.pxarexclude at the container root:
echo "/home/downloader/.cache/uv/*" >> /.pxarexcludebashThe files stay live on ZFS; the backup walk just steps over them, so nothing breaks and the chunk store stops growing to track downloads and package caches. The real application state lives in bind mounts under /data with its own restic coverage via Backrest, which is why skipping these paths costs nothing. One deliberate exception: the git server has no exclude file. Its 45 GB of repositories and SQLite databases is exactly the kind of data the backups exist for.
Two write-wasters inside the backup container#
Space is half the story. The PBS container’s root disk lives on the NVMe ZFS mirror, and ZFS is copy-on-write, so every redundant journal line is amplified into metadata writes on the pool that also hosts every guest root disk.
The first waster was unbounded systemd-journald. The PBS container alone had generated 1.0 GB of journal, and a network scanner was past 200 MB, fed by constant ARP broadcast chatter. The fix is a vacuum now and a cap forever, deployed by a host-level loop over every running LXC that actually runs systemd, skipping the OpenRC Alpine guests:
journalctl --vacuum-size=50Mbash[Journal]
SystemMaxUse=50M
MaxRetentionSec=1weekiniAfter the sweep the PBS container’s journal sat at 36.9 MB, and the cap holds it there.
The second was quieter and nastier. zfs-zed.service ships enabled with the PBS install, but the ZFS Event Daemon cannot reach block devices from an unprivileged LXC, so it crash-looped forever. The restart counter was at 128,238 when I found it, and every restart wrote more log lines to the same pool. The Proxmox host handles ZFS events natively; the copy inside the container is pure noise:
systemctl disable --now zfs-zed
systemctl mask zfs-zedbashChecking the disk is still healthy#
Everything above lands on one physical disk, so it gets checked. SMART on the Exos after 21,893 power-on hours:
| Attribute | Value |
|---|---|
| Overall health | PASSED |
| Reallocated / pending / uncorrectable sectors | 0 / 0 / 0 |
| Temperature | 33 C |
| Data on the BKP dataset | 243 GB |
Both pools get monthly scrubs, and the last garbage collection run reported zero bad chunks. The honest limitation from the homelab post stands: the backups share the box with the data they protect. The offsite leg runs on restic as the policy layer and rclone as the backend, two separate providers, everything encrypted before it leaves the box; that wiring is its own post.
Closing#
Deduplication carries most of the load once the guests share base images. The wins that needed actual work were the prune schedule cutting 315 snapshots to 210, exclusions that stop caches and downloads from entering the chunk store, and write hygiene in the containers themselves, a journald cap and a masked zed whose crash loop had written a small novel into the journal. The result is a backup system protecting 2.24 TiB of guests in 164.25 GiB of disk, with writes only where they earn their place.