The project
The same repositories, in less time.
restic backs up files into encrypted, deduplicated repositories. Dealer Style Restic is a build of restic 0.19.1-dev (upstream commit ba802d42b7294c98b62c16d1157ea3e80820c019) that does the same work in less time and with less memory. It keeps the commands, the options and upstream's defaults.
It keeps the repository format. Chunking and blob IDs are identical, and data and tree blobs are stored in the same form upstream stores them, so each program reads what the other wrote and deduplicates against it. Index and lock files decode to the same content but are encoded slightly differently: they are valid Zstandard, and upstream reads them.
Everything the program uses has been authenticated and hash-checked when it was read. In two places this build keeps what it read where upstream reads again: restore's second pass reuses directory trees that its first pass verified and kept, and an index read during the lock wait is kept if a listing taken after the lock is confirmed still shows the same names and sizes. Both are described under known limits.
It came out of a measured search. Teams of AI agents proposed changes to restic over 20 seasons. A fixed judge timed every candidate against upstream and compared its results with upstream's before any time counted. Independent reviewer agents then read the code for soundness. They found one defect that could silently break a backup and 10 more findings, and all of them were repaired. Release checks found 4 more failures, which were fixed. Then 3 reviewers of a different model family, given only upstream's source and the tree, found no data loss and reported further findings: 9 were fixed and 2 are documented differences. What is published is the speed that was left afterwards.
What it is not
It is an experimental, independent derivative. It is not affiliated with or endorsed by the restic project. restic is the work of its authors and contributors; Slop Dealer's contribution is the published optimisation work, the measurements and the reviews. Report problems with this build to the Dealer Style Restic repository, never to upstream. Some error messages inherited from upstream still link to upstream's issue tracker; for this build always use https://github.com/slopdealer/Dealer-Style-Restic/issues.
It may suit you if
- you run restic on Linux (amd64), and the time or memory of backup, restore, check or prune matters to you;
- you can build from source and will compare this build with upstream on your own data;
- you keep upstream restic installed. The repository format is the same, so you can go back at any time.
Do not use it if
- these would be your only backups, or you cannot verify a restore independently;
- you need a supported release, or an audit by people;
- you run macOS, a BSD or Windows and expect it to have been tested there. Nothing was run on macOS or the BSDs, and on Windows only unit tests ran;
- you need the gain to hold on your machine without measuring it. What was measured is 4 CPUs, memory-backed storage and four synthetic workloads.
Try it on a separate test repository first. Back up a copy of real data with this build, run upstream restic's check --read-data on the result, restore it and compare the files. Only then point it at a repository you care about.
Measured performance
Read the comparison.

Every timing below comes from one Linux machine (AMD EPYC 9454, 48 cores, 128 GB RAM, Linux), with 4 pinned CPUs for each judge job, the CPUs held at full clock (the performance frequency governor), and the corpus, repositories and restore targets on memory-backed storage. Each comparison is paired: in every round, every command is run by upstream restic and by this build back to back, on fresh copies of the same repository state, and the order rotates. A result counts only if the build's output matched upstream's.
The judge measured working commit b8cf439. The published source is commit 8f9372b, which changes one comment and adds one test and builds to the same machine code. The exact statement is with the downloads, and the full account is in section 2 of the repository's RESULTS.
Against upstream restic
| Workload | Time saved | 90% interval | Rounds ahead | Upstream, s | This build, s | Peak memory saved |
|---|---|---|---|---|---|---|
| All four (mean) | +29.00% | +28.84 to +29.05 | 44 of 44 | 110.26 | 78.27 | +38.99% |
| main | +30.63% | +30.37 to +30.78 | 11 of 11 | 22.40 | 15.55 | +34.90% |
| big-index | +45.13% | +44.96 to +45.21 | 11 of 11 | 32.53 | 17.84 | +44.19% |
| large-files | +22.96% | +22.60 to +23.10 | 11 of 11 | 16.45 | 12.67 | +34.08% |
| remote | +17.26% | +17.10 to +17.44 | 11 of 11 | 38.88 | 32.21 | +42.79% |
The published tree was ahead of upstream in 44 of 44 paired rounds. By workload the time saved runs from 17.26% to 45.13%. The workloads are described under methods. In short: main is many small files; big-index is the same with 1,500,000 extra blobs in the repository's index; large-files is a few very large files; remote is main's files with the repository behind a simulated network link.
upstream resticthis build
upstream resticthis build
Peak memory as a table
| Workload | Upstream, mean peak MiB | This build, mean peak MiB | Upstream, sum of seven peaks | This build, sum of seven peaks | Memory saved |
|---|---|---|---|---|---|
| main | 143.0 | 93.6 | 996.3 | 651.9 | +34.90% |
| big-index | 423.0 | 236.8 | 2942.8 | 1655.4 | +44.19% |
| large-files | 128.5 | 84.5 | 897.8 | 590.3 | +34.08% |
| remote | 164.1 | 94.5 | 1151.5 | 661.5 | +42.79% |
More than 4 CPUs
The published machine code was then run at 8 and 16 cores. All 56 cells, seven commands on four workloads at two core counts, are ahead of upstream; the smallest is remote check-read-data at 16 cores, +5.51%. In restore, forget-prune and ls this build uses more user CPU than upstream, and the gap widens with the number of cores.
| Workload | 4 CPUs (judge) | 8 cores (bench) | 16 cores (bench) |
|---|---|---|---|
| main | +30.63% | +31.47% | +33.03% |
| big-index | +45.13% | +44.99% | +46.27% |
| large-files | +22.96% | +28.86% | +32.11% |
| remote | +17.26% | +17.64% | +18.51% |
All 56 cells at 8 and 16 cores, by command
| Workload, command | 4 CPUs (judge) | 8 cores (bench) | 16 cores (bench) |
|---|---|---|---|
main, backup-initial | +7.17% | +6.03% | +11.38% |
main, backup-unchanged | +23.96% | +24.77% | +25.27% |
main, backup-mutated | +22.36% | +24.19% | +24.41% |
main, restore | +48.76% | +46.72% | +45.95% |
main, check-read-data | +32.45% | +33.85% | +34.19% |
main, forget-prune | +44.60% | +49.04% | +50.16% |
main, ls-json | +42.33% | +41.88% | +42.38% |
big-index, backup-initial | +17.33% | +15.43% | +15.86% |
big-index, backup-unchanged | +34.23% | +31.32% | +31.82% |
big-index, backup-mutated | +31.59% | +29.59% | +29.81% |
big-index, restore | +52.06% | +49.85% | +49.61% |
big-index, check-read-data | +39.78% | +40.39% | +43.74% |
big-index, forget-prune | +70.75% | +71.62% | +72.33% |
big-index, ls-json | +53.00% | +50.35% | +51.75% |
large-files, backup-initial | +11.15% | +17.83% | +28.28% |
large-files, backup-unchanged | +26.96% | +36.33% | +38.36% |
large-files, backup-mutated | +16.20% | +26.73% | +31.81% |
large-files, restore | +20.73% | +18.73% | +18.96% |
large-files, check-read-data | +27.61% | +28.78% | +30.08% |
large-files, forget-prune | +45.80% | +50.13% | +48.67% |
large-files, ls-json | +38.78% | +38.66% | +38.56% |
remote, backup-initial | +6.83% | +6.58% | +8.57% |
remote, backup-unchanged | +21.15% | +21.67% | +22.48% |
remote, backup-mutated | +18.44% | +19.75% | +22.37% |
remote, restore | +30.18% | +30.07% | +30.76% |
remote, check-read-data | +5.55% | +6.16% | +5.51% |
remote, forget-prune | +11.98% | +13.24% | +13.38% |
remote, ls-json | +36.92% | +37.08% | +37.41% |
Many directories. One result changes sign with the conditions. For a restore that is almost only directories, on a machine with many CPUs under a load-following frequency governor (the usual setting), this build was 4.8 to 10.7% slower than upstream in 4 readings. On 4 cores under the same governor it was 26.8% faster, and with the CPUs held at full clock it was 16 to 18% faster. The measurements, upstream against this build, on the release-check CPUs: 606,060 directories and 3 files, three rotated rounds. Load-following governor, with the official judge job running beside it: 16 CPUs 67.2 s against 74.0 s (10.1% slower); 4 cores 86.5 s against 63.3 s (26.8% faster). The same CPUs held at full clock: 16 CPUs 53.3 s against 44.7 s (16.1% faster); 4 cores 53.5 s against 44.1 s (17.6% faster). Load-following governor again, on a quiet machine: 16 CPUs 68.0 s against 75.3 s (10.7% slower). Larger snapshots under the load-following governor with 16 CPUs: 1,010,100 directories, 7.8% slower (three rounds); 2,020,200 directories, 4.8% slower (one run per build). What this settles: holding the CPUs at full clock removes the loss, and a quiet machine does not. The cache layout was ruled out before: at full clock, on eight cores spread over four cache domains, the same restore was 16 to 17% faster than upstream. So the judge job that ran beside the first measurement did not cause the loss. In the quiet run the mean clock of the running CPUs was 2.55 GHz for upstream and 2.02 GHz for this build, and the CPU seconds were 71.4 against 98.7; at the fixed clock both used about 55. That this build spreads the work over more threads, so that the governor sees less load per CPU and keeps the clock low, is an inference from the clock readings and the CPU seconds, not a measurement.
By command
Each workload runs the same seven commands. The grid shows the time saved against upstream for each command on each workload.
0%80% time saved
| Command | main | big-index | large-files | remote | mean |
|---|---|---|---|---|---|
backup-initialfirst backup | +7.17 | +17.33 | +11.15 | +6.83 | +10.62 |
backup-unchangedbackup, nothing changed | +23.96 | +34.23 | +26.96 | +21.15 | +26.57 |
backup-mutatedbackup after changes | +22.36 | +31.59 | +16.20 | +18.44 | +22.15 |
restorefull restore | +48.76 | +52.06 | +20.73 | +30.18 | +37.94 |
check-read-datacheck, reading all data | +32.45 | +39.78 | +27.61 | +5.55 | +26.35 |
forget-pruneforget and prune | +44.60 | +70.75 | +45.80 | +11.98 | +43.28 |
ls-jsonlist a snapshot | +42.33 | +53.00 | +38.78 | +36.92 | +42.76 |
Seconds per command: upstream → this build
| Command | main | big-index | large-files | remote |
|---|---|---|---|---|
backup-initial | 5.53 → 5.14 | 5.81 → 4.79 | 4.28 → 3.82 | 8.43 → 7.87 |
backup-unchanged | 2.22 → 1.69 | 2.74 → 1.80 | 1.94 → 1.41 | 3.00 → 2.36 |
backup-mutated | 2.35 → 1.82 | 2.83 → 1.93 | 2.48 → 2.08 | 3.25 → 2.63 |
restore | 6.03 → 3.09 | 6.49 → 3.11 | 2.91 → 2.31 | 10.76 → 7.51 |
check-read-data | 2.56 → 1.71 | 5.07 → 3.06 | 2.18 → 1.57 | 7.28 → 6.88 |
forget-prune | 2.23 → 1.23 | 7.61 → 2.23 | 1.94 → 1.05 | 4.14 → 3.68 |
ls-json | 1.47 → 0.85 | 1.97 → 0.93 | 0.73 → 0.44 | 2.04 → 1.29 |
Across the 28 cells the time saved runs from 5.55% to 70.75%, and in 28 of them the build was ahead of upstream in every round. The first backup saves 6.83% to 17.33%, depending on the workload: most of its time is reading, chunking, compressing and encrypting the data, which this build does as upstream does.
Season by season
The league ran 20 seasons, from 24 September 2026 to 7 October 2026. 17 of them ended with a new tree and 3 did not. The way the league scored changed three times, so the chart follows the one measure that was recorded throughout: time saved against upstream on the main workload.
main workloadmean of the scored workloadspromotionhand mergeseason without a promotion
Every season, hand merge and fix as a table
| Step | Against the tree before | Score against upstream | Main workload, time | What changed |
|---|---|---|---|---|
| Season 1 | +3.24% (41 of 42) | +3.24% | −2.16% | Each chunk is copied into an exact-size buffer and the pooled read buffer is returned at once: about 8% less peak memory for about 2% more time. |
| Season 2 | −0.53% (15 of 42), not promoted | Not promoted. The nominee screened at +1.29% over 6 rounds and confirmed at −0.53% over 42: its time gain held, its memory cost had not shown in the screen. | ||
| Season 3 | +0.49% (27 of 42), not promoted | Not promoted, two wins short of the gate. The screen overstated the gain about three times. | ||
| Season 4 | +5.25% (42 of 42) | +8.57% | +6.52% | First season aimed by a measured map of where the time goes: per-process start-up (key derivation on one core, the fixed lock wait, the index load). |
| Hand merge | +15.07% | Hand merge of two other teams' memory work (read side and backup write path) onto the season's winner. Its time on its own was not recorded. | ||
| Season 5 | +2.99% (41 of 42) | +17.82% | +11.47% | Backup write path. It lifted the one phase that was slower than upstream (the first backup, a cost of season 1's buffer copy). |
| Season 6 | +1.11% (32 of 42) | +18.59% | +11.26% | Read-side buffer retention and pack hashing on write: a small memory gain, 0.26% slower than season 5. |
| Season 7 | +0.014% (21 of 42), not promoted | Not promoted. The highest screen of the season (+2.08%, with a 90% interval that spanned zero) confirmed at +0.014%. | ||
| Season 8 | +8.42% (42 of 42) | +24.98% | +16.03% | Second map. Bounded backup read buffers, asynchronous saves in prune's repack, less memory in check. |
| Season 9 | +1.30% (38 of 42) | +25.62% | +15.94% | Restore path and index representation. |
| Season 10 | +0.84% (32 of 42) | +26.88% | +16.67% | A small gain on both axes; the season ran without a new map. |
| Season 11 | +2.10% (42 of 42) | +27.50% | +16.91% | Third map. Restore time and prune's peak memory. |
| Season 12 | +2.73% (42 of 42) | +29.68% | +17.38% | Memory bought with time: the promoted tree was 0.64% slower than season 11 and 5.52% leaner. The last season scored on memory. |
| Season 13 | +3.11% (42 of 42) | +19.69% | +19.69% | First season scored on time alone, with memory as a guard. The index load starts inside the lock wait for backup; file opens without the poller; tree decoding. |
| Season 14 | +1.35% (42 of 42) | +21.02% | +21.02% | Restore reads trees ahead of the walk. |
| Season 15 | +1.47% (42 of 42) | +22.58% | +22.58% | The index load inside the lock wait, extended to restore, check, prune and ls. |
| Season 16 | +4.71% (42 of 42) | +25.75% | +25.75% | A bounded third worker in check, a wider second-pass window and node look-ahead in restore, a direct path stat on Linux. |
| Fix at the boundary | A concurrency defect carried since season 13 was fixed at the boundary: beside a running backup, restore latest and ls latest failed about one run in three. The judge never ran two commands at once, so every season had passed. | |||
| Season 17 | +1.93% (34 of 42) | +21.89% | +27.61% | First season scored on three workloads. Restore and check on the big-index workload; large-files did not move (−0.10%). |
| Hand merge | +6.06% (33 of 42) | +27.08% | +27.92% | Hand merge of three unpromoted ideas and two prototypes from the speed map. Nearly all of the gain is the big-index workload (+17.90%). |
| Season 18 | +4.00% (42 of 42) | +29.46% | +29.30% | Aimed at large files: large chunks hashed off the reader, check on large packs, a streaming index decode (large-files +6.15%). |
| Season 19 | +4.33% (44 of 44) | +28.15% | +30.23% | First season with the remote workload in the score. The tree had been sending five times upstream's requests in restore; every team removed that in the first match (remote +9.37%). |
| Hand merge | +1.14% (36 of 44) | +28.99% | +31.11% | Hand merge of the season's unpromoted work. |
| Season 20 | +1.03% (35 of 44) | +29.94% | +31.28% | The last season. Index-side work in check and prune (big-index +3.46%), and three new mechanisms inside the lock wait that were later found unsound and left out of the published tree. |
| Published tree | −1.06% (6 of 44) | +29.00% | +30.63% | The published tree: the league tree after season 19, the repairs, the sound part of season 20, four ports from unpromoted trees, the release-check fixes and the fixes from the independent review. It is below season 20: by the league's promotion rules it would not have replaced that tree. |
Against the league's last champion: what the repairs cost
The league's last champion, the tree it promoted in season 20, was faster than the published tree. The champion also carried the lock-wait defect at three new places, and it stored some blobs in a form upstream does not. The published tree was measured against it directly, in the same job as the comparison with upstream.
| Workload | Time saved | 90% interval | Rounds ahead | Memory saved | Sum of peaks |
|---|---|---|---|---|---|
| All four (mean) | −1.06% | −1.13 to −0.91 | 6 of 44 | −8.65% | |
| main | −0.83% | −1.03 to −0.74 | 1 of 11 | −12.62% | +76.7 MiB |
| big-index | −0.18% | −0.31 to +0.09 | 5 of 11 | −5.17% | +82.8 MiB |
| large-files | −1.07% | −1.34 to −0.67 | 0 of 11 | −7.23% | +41.1 MiB |
| remote | −2.18% | −2.26 to −2.02 | 0 of 11 | −9.58% | +57.3 MiB |
By the league's own promotion rules this tree would not have replaced the champion. It was ahead of the champion in 6 of 44 rounds, where promotion needs 29. Its time over the four workloads was 1.06% worse, where promotion needs a gain of at least 0.25%. Its peak memory was 8.65% higher, where the guard allows 2% (by workload: main 12.62%, big-index 5.17%, large-files 7.23%, remote 9.58%). Two workloads were more than 1% slower than the champion: large-files −1.07%, remote −2.18%. The difference is the price of the correctness repairs and of the fixes from the independent review.
The tree as it stood before the independent review's fixes read −0.79% against the champion in its own job of 44 rounds (ahead in 6). Those are two separate jobs, so the difference between them is an indication of what the fixes cost, not a measurement.
The repairs alone, in their final form, cost 1.98% of the time of a pass against the tree they repaired (bench, 4 CPUs: main −2.00%, big-index −1.72%, large-files −2.55%, remote −1.66%) and added 39.2 to 84.8 MiB to the sum of the seven commands' peaks. As first staged, before two avoidable costs were found and removed, an official job of 12 rounds read −3.15% (ahead in 0 of 12). Returning to the form in which upstream stores blobs that do not compress cost 0.70 to 0.79 suite points of that on its own. The final merge also brought in sound work the league had not promoted, which is why the published tree ends closer to the champion than the price of the repairs alone.
| Repair | What it does | Where it costs | Measured |
|---|---|---|---|
| R1 | Repository state read during the lock wait is checked again after the lock is confirmed | every process on a remote repository: one or two more listings | remote pass 32.05 to 32.41 s (−1.14%); ls −4.36%; local workloads −0.15%, +0.02%, −0.21% (unresolved) |
| R2 | Prune's repack handler no longer runs while holding a backend connection | forget --prune on a remote repository | remote 3.410 to 3.524 s (−3.27%, 0 of 10 pairs ahead), requests 119 to 134 |
| R4 | The compression sampler is removed; every blob goes through the encoder | the first backup, and prune | first backup −2.39% on main (+8.8 MiB) and −2.57% on large-files (+8.4 MiB) in the final form |
| F2 | Blobs that do not compress are stored in upstream's form again | the first backup, and prune | −0.70 suite points on the bench, 0.79 between the two official jobs; +28 to +52 MiB of summed phase peaks |
| R5 | Index downloads follow the backend's connections, not the CPU count | index load on a remote repository | remote pass +0.41% (unresolved), +1.7 MiB |
| R7 | The fast tree decoder rejects the node names upstream rejects | tree decoding | +0.21%, −0.49%, −0.33%, −0.12% on four phases of main (all unresolved) |
| R3, R6, R8 to R11, M22 | The other repairs | not priced singly; inside the totals |
A workload the teams never saw
One workload was held out for the whole league, to test whether gains fitted to the visible workloads carry over. The teams could not select it. Like the main workload, it is dominated by many small files. It was measured at season boundaries only, each tree against upstream, 12 paired rounds on 4 CPUs.
| Tree | Held-out: time saved | 90% interval | Held-out: memory saved | Main workload: time saved |
|---|---|---|---|---|
| Season 4 | +4.08% | 3.97 to 4.24 | +11.10% | +6.52% |
| Season 4 + hand merge | +5.73% | 5.61 to 5.98 | +15.12% | not recorded |
| Season 5 | +5.94% | 5.82 to 6.07 | +15.47% | +11.47% |
| Season 6 | +5.89% | 5.64 to 6.01 | +16.13% | +11.26% |
| Season 8 | +10.46% | 10.27 to 10.57 | +21.22% | +16.03% |
| Season 9 | +13.50% | 13.33 to 13.58 | +22.79% | +15.94% |
| Season 10 | +13.48% | 13.36 to 13.58 | +24.59% | +16.67% |
| Season 11 | +16.43% | 16.21 to 16.53 | +25.01% | +16.91% |
| Season 12 | +17.14% | 17.07 to 17.19 | +30.51% | +17.38% |
| Season 13 | +21.09% | 20.94 to 21.18 | +28.50% | +19.69% |
| Season 14 | +22.28% | 22.23 to 22.35 | +30.66% | +21.02% |
| Season 15 | +23.27% | 22.97 to 23.36 | +29.65% | +22.58% |
| Season 16 | +26.71% | 26.54 to 26.85 | +29.85% | +25.75% |
| Season 16 + race fix | +26.59% | 26.45 to 26.76 | +28.58% | not recorded |
| Season 17 | +28.75% | 28.65 to 28.90 | +31.26% | +27.61% |
| Season 18 | +29.91% | 29.68 to 30.08 | +36.89% | +29.30% |
| Season 19 | +30.76% | 30.48 to 30.91 | +39.34% | +30.23% |
The tree after season 19 on the held-out workload, by command
| Command | Held-out: time saved | 90% interval | Held-out: memory saved | Main workload: time saved |
|---|---|---|---|---|
backup-initial | +12.83% | 12.28 to 13.11 | +43.31% | +10.89% |
backup-unchanged | +18.11% | 18.01 to 18.37 | +38.35% | +23.25% |
backup-mutated | +14.38% | 14.11 to 14.63 | +44.10% | +21.96% |
restore | +47.84% | 47.06 to 48.78 | +39.51% | +44.62% |
check-read-data | +31.80% | 31.55 to 32.76 | +34.93% | +34.69% |
forget-prune | +45.40% | 45.07 to 45.76 | +29.77% | +43.13% |
ls-json | +36.77% | 36.54 to 36.81 | +42.81% | +39.72% |
Season steps, held-out against visible
| Step | Held-out workload | Main workload | Unit |
|---|---|---|---|
| Season 14 to 15 | +1.40 | +1.47 | % against the previous tree |
| Season 16 (race fix) to 17 | about +2.2 | +2.18 | points against upstream; main as % against the previous tree |
| Season 17 to 18 (with the hand merge) | +1.16 | +1.69 | points against upstream |
| Season 18 to 19 | +0.85 | +0.93 | points against upstream |
From season 11 on, the gain on the held-out workload is within about one point of the gain on the main workload, and the steps from season to season reproduce. The early seasons, which were also scored on memory, carried over about half of their main-workload time gain. A full pass of the held-out workload took upstream 46.67 s and the tree after season 19 32.35 s. The large-files and remote workloads have no held-out counterpart. The published tree itself has not been run on the held-out workload.
rustic, Kopia and Borg, as measured
On 3 October 2026 the league tree after season 15 was run beside three other backup programs on the main workload's corpus: the same seven steps with each tool's nearest equivalent command, encryption on and zstd-class compression everywhere, 4 CPUs, 10 repetitions. This is one measurement of an earlier tree on one workload. It is not a benchmark of the published tree, and the other tools were run at their defaults by a team that knows restic better than it knows them.
| Tool | Version | Seven commands, s | Range of 10 | Sum of seven peaks, MiB | Largest peak, MiB | Restores equal to the source |
|---|---|---|---|---|---|---|
| upstream restic | 0.19.1-dev | 21.45 | 21.26 to 21.56 | 917 | 210 | 10 of 10 |
| league tree after season 15 | restic 0.19.1-dev + league changes | 17.07 | 16.96 to 17.19 | 593 | 112 | 10 of 10 |
| rustic | 0.11.4 | 23.51 | 23.23 to 23.86 | 2,262 | 1,240 | 10 of 10 |
| Kopia | 0.23.1 | 14.48 | 14.35 to 15.14 | 2,515 | 655 | 0 of 10 |
| Borg, cold cache | 1.4.5 | 54.54 | 54.16 to 54.78 | 727 | 135 | 10 of 10 |
| Borg, warm cache | 1.4.5 | 49.10 | 48.60 to 49.24 | 721 | 137 | 10 of 10 |
Seconds per command, all six arms
| Command | upstream restic | league tree after season 15 | rustic | Kopia | Borg, cold | Borg, warm |
|---|---|---|---|---|---|---|
| first backup | 5.83 | 5.45 | 11.09 | 3.98 | 15.05 | 15.11 |
| backup, nothing changed | 1.88 | 1.51 | 1.67 | 0.52 | 8.24 | 5.19 |
| backup after changes | 2.40 | 2.01 | 3.17 | 1.11 | 8.38 | 6.18 |
| restore | 5.97 | 3.96 | 4.40 | 4.11 | 11.08 | 11.09 |
| full verify | 2.78 | 2.46 | 1.68 | 2.98 | 7.62 | 7.58 |
| forget and prune | 1.09 | 0.76 | 0.82 | 1.10 | 1.70 | 1.55 |
| list latest snapshot | 1.46 | 0.88 | 0.70 | 0.66 | 2.38 | 2.38 |
- Kopia was the fastest overall, at 14.48 s against 17.07 s for the league tree. The lead is in repeat backups: 0.52 s against 1.51 s when nothing had changed. It used about four times the memory. Its restore differed from the source in 10 of 10 runs: a file name that is not valid UTF-8 was stored altered, a named pipe was skipped without an error, and hard-linked files were restored as separate files.
- rustic read 23.51 s as run. The method used to hold the CPU clocks steady penalised its worker pipeline; without that the estimate is 19.2 to 20.2 s. It used about four times the memory of the league tree, and 5 of its 30 runs had a command fail on an error in its local cache. The repositories themselves were intact.
- Borg took 49.10 to 54.54 s. It works on one core. Its memory use was close to restic's.
What did not work
- Three seasons promoted nothing. Seasons 2, 3 and 7 ended without a new tree. In each, the nominee's 6-round screen (+1.29%, +1.47%, +2.08%) shrank or reversed at 42 rounds (−0.53%, +0.49%, +0.014%).
- Searching without a map. Seasons 1 to 3 searched unaimed and circled one corner of the backup path. Two later seasons with open scope and no new map added +1.11% and nothing.
- Scoring memory and time together. By season 12 the equal-weight score was paying for memory with time: 14 of that season's 16 match rows were slower than their incumbent. The promoted tree was 0.64% slower than the one before it.
- Scoring one workload. On two new workloads the tree after season 15 held about half of its main-workload gain, and its
check --read-dataon a large index was 33% slower than upstream. On the first simulated link, the tree after season 17 was 3.25% faster than upstream, against 27.61% locally. - The automatic season merge. It ran in four seasons and merged nothing in any of them: the best trees per phase edited the same files. Unpromoted work reached the tree only through three hand merges.
- Work inside the lock wait, as built in the last season. It led the menu with an estimated bound of +4.9 points and all four teams built it. The three mechanisms that were promoted cost 0.14 suite points together when measured against the same tree without them. The review found them unsound. None is in the published tree.
- Storing incompressible blobs raw. It saved time on the first backup, and it made
prune --repack-uncompressedrewrite the same 16,179 blobs on every run. The published tree stores every blob in upstream's form, at a price of 0.70 to 0.79 suite points. - Guessing compressibility from a sample. The sampler stored a redundant blob at 1,048,608 bytes where upstream stores 65,661. It was removed.
- Setting a file's metadata as soon as its last blob is written. Two teams tried it in season 20. An executed test showed that it changes what a later fatal error leaves behind. Both teams kept upstream's order.
How it was measured
The league, the judge and the noise.

The league
The search was run as a league. From season 2 there were 4 teams, T1 to T4. A team was a set of AI agents with fixed roles: an owner who chose where the team searched, a coach who planned each match, builders who wrote and measured the code, and inventors who proposed ideas. A coordinator agent ran the seasons for one human operator, who set the rules and made the rulings.
A season had 4 matches. In a match each team handed in one source tree. The judge built it, checked it and timed it against the incumbent (the best tree so far) and against upstream over 6 paired rounds, 8 from season 19. That is a screen: enough to rank trees, not enough to believe a number. From season 13 the 3 best screens were measured again at 12 rounds, the rescreen, and the best of those went to one long confirmation of 42 rounds, 44 from season 19. A tree became the new incumbent only if its confirmation passed every gate:
- ahead of the incumbent in at least 29 of 42 rounds, with a median gain of at least 0.25%;
- peak memory no more than 2% above the incumbent's, on every workload;
- no workload more than 1% slower than the incumbent;
- faster than upstream;
- and all of it only after every correctness check had passed.
The teams worked under fixed rules. Only non-test Go source under cmd/ and internal/ could change, with no new dependencies. The repository format, the output, the exit codes and every integrity check had to stay as upstream's. Detecting the benchmark was forbidden. The 200 ms lock wait could be overlapped with other work but never removed or shortened. No code could burn CPU to influence the clock.
The teams ran on these models, as the season reviews record: season 1, a single team on a provider's unnamed preview model; seasons 2 to 11 and 13 to 15, GPT-6 Luna at maximum effort; season 12, GPT-6.1 Sol at high effort; season 16, Luna and then Sol at medium effort from mid-season; seasons 17 to 20, GPT-6.1 Sol at medium effort. The reviews, the repairs, the merges and the release checks were done by other agents.
The judge
Three arms on one bench. A judge job measures three builds: the candidate, the incumbent and upstream (the judge's files call it stock). The judge builds all three itself, offline, from source with one pinned toolchain (go1.27.0). Upstream's build must come out bit-identical every time.
Paired rounds, rotated order. In a round, each of the seven commands is run by all three arms back to back, each on a fresh copy of the same repository state, which upstream prepared. The order of the arms rotates through 6 patterns. Positions balance exactly over complete cycles; in the final job of 44 rounds this build ran first, second and third 12, 16 and 16 times. The value of a round is the paired difference within it. Reported figures are medians over the rounds, the number of rounds in which the candidate was ahead, and a 90% bootstrap interval (4,000 resamples).
Seven commands. A first backup; a backup with nothing changed; a backup after a set of changes; a full restore; check --read-data; forget --prune; and ls --json. Each is one restic process with upstream's defaults, in a sandbox with no network, on 4 pinned CPUs, with an empty local cache. Time is wall-clock time around the process. Memory is the process's peak resident size as the kernel reports it at exit; no garbage-collector tuning is set.
Correctness, every round. Before any timing, each arm runs every command once and upstream checks the result. Upstream runs check --read-data on the arm's repository, restores it, and compares the restored tree with the expected one entry by entry (content, modes, owners, times, symlinks, hard links, extended attributes, odd file names). The arm must also detect a deliberately corrupted pack. Then, in every measured round, the arm's snapshot records must equal upstream's. Equal root tree IDs are the fast path; if they differ, upstream must fully check and restore the result, and its manifest must match. The arm's restore must equal upstream's restore. After prune the kept snapshots must be upstream's and upstream's check must pass: the warm-up reads all the data, and the measured rounds rotate through fifths of it. The arm's JSON listing must parse equal. One failure aborts the job. Upstream's own test files for the storage packages and its command-line integration suite are also compiled against the arm's code, and must pass wherever upstream passes (1,026 tests). In the job behind the tables above, 1,019 of 1,019 such units were verified.
Four workloads, and why each exists
main- 91,264 entries, 1.64 GiB: 90,000 small files in a project-like tree, medium and large files, exact duplicates, a sparse file, and a block of awkward entries (odd names, links, special modes, a named pipe). The changed version applies 4,224 edits. It was the only workload until season 16.
big-index- The main workload, with 1,500,000 extra small blobs already in the repository, so every command that opens the index carries them. It exists because the gains had been fitted to main: the tree after season 15 saved 21.16% on main and 11.81% here, and its check on this workload was 33% slower than upstream's.
large-files- 3,217 entries, 2.91 GB: a sparse disk image, a page-structured database file, logs, incompressible media files and a small tree. It exists for the same reason: the tree after season 15 saved 13.00% on it.
remote- The main workload's files, with the repository reached through restic's REST backend over a simulated link. It was added to the score in season 19, to see whether local gains survive a network.
The simulated link, stated plainly. The remote workload does not use a network. A small REST server, written for the judge with Go's standard library, runs on the same machine and is reached through a unix socket. It delays every request by a fixed 30 ms, delivers bytes at 48 MiB/s per connection and serves 5 requests at a time. The delays are sleeps on a schedule. There is no jitter, no loss, no congestion, no TLS and no object-store behaviour. It answers one question, whether a change costs requests or bytes, and says nothing about any real storage service. On the first version of the link (12 MiB/s), the tree after season 17 was 3.25% faster than upstream where it was 27.61% faster locally. With the remote workload in the score, one season moved it from 8.83% to 18.37%.
Noise, and how it was handled
Null calibration. The judge was run with three identical builds as its three arms. Whatever difference it reports between them is noise. In the first null (24 September 2026, 18 rounds) the paired difference in time had a standard deviation of 1.11% per round and peak memory 2.22%. The gates were set from it: a tree with no real change passes "ahead in 29 of 42, median at least 0.25%" about 1.2% of the time. The null was run again on 3 October 2026, after the judge machine moved to a fixed CPU frequency policy: time noise fell to 0.27% per round, a standard error of 0.05% on the median of 42 rounds, while memory noise stayed at 2.48%. On the remote workload the null read 0.42% per round (48 identical pairs).
The judge was also shown an effect of known size: a build with a deliberate 150 ms sleep and 24 MiB of extra memory in restore. By construction that is −2.2% on that command; the judge measured −2.04%, and the other six commands stayed inside their noise.
Noise per command
| Command | Time SD, first null, % | Time SD, second null, % | Peak memory SD, second null, % |
|---|---|---|---|
backup-initial | 0.82 | 0.61 | 5.03 |
backup-unchanged | 2.50 | 0.53 | 6.17 |
backup-mutated | 4.35 | 0.92 | 6.89 |
restore | 2.27 | 0.51 | 5.46 |
check-read-data | 2.98 | 0.75 | 6.00 |
forget-prune | 4.59 | 1.19 | 5.10 |
ls-json | 1.25 | 0.27 | 1.29 |
Screens that inverted at confirmation. A screen of 6 rounds ranks trees; it does not measure them. The league learned that three times before it changed its procedure.
| Candidate | Screen | Rescreen | Confirmation | Outcome |
|---|---|---|---|---|
| Season 2 nominee | +1.29% (6 rounds) | −0.53% (42 rounds, 15 ahead) | not promoted | |
| Season 3 nominee | +1.47% (6 rounds) | +0.49% (42 rounds, 27 ahead; 29 needed) | not promoted | |
| Season 7 nominee | +2.08% (6 rounds) | +0.014% (42 rounds, 21 ahead) | not promoted | |
| Season 17, the two best trees of one team | +2.44% and +2.35% (6 rounds) | +2.13% and +2.68% (12 rounds): the order reversed | +1.93% (42 rounds) for the second | promoted |
| Season 20 nominee | +1.81% (8 rounds) | +1.41% (12 rounds) | +1.03% (44 rounds, 35 ahead) | promoted |
The rescreen was added in season 13 for this reason, and in season 17 it changed which tree was confirmed. No number on this page is a screen quoted as a result.
Two jobs side by side. In the last season, prune on the big-index workload read three to four times larger when a second judge job ran on the machine at the same time: +22.54% in a match and +23.83% at the rescreen, both beside another job, then +6.70% in the 44-round job, which ran alone. Another team's tree read +22.51% beside a job and +4.75% alone. Check on the same workload did not move (+13.57%, +12.56%, +13.88%). The job behind the tables above ran alone on its lanes.
Binary layout. The same source can run at two speeds. restic derives its key at the start of every process. The inner loop of that derivation is about 3% slower when the linker places it at one of two alignments. Code added earlier in the binary can move it. It showed up as a listing command reading −1.65% (ahead in 0 of 3 rounds) in a tree that had changed no listing code: about 7 ms in every process, worth up to 0.9 suite points. The merge checked the alignment of the judged build and restored the fast one.
Stalls. On lanes shared with other work, 2 of 504 null runs took about 2.5 s longer than their siblings. Medians ignore them; a single reading that is off by seconds was treated as a stall and repeated.
The held-out workload
A fifth workload was held out. The teams could not select it, and its generator, seed and location are not published. An audit of the teams' transcripts found 7 sightings of its directory name in season 17, all in directory listings, and no recorded reading of its contents. After season 17 it was kept off the judge machine while seasons ran, and the transcripts were searched again before each test. It was run only at season boundaries, by the coordinator. Its results are above.
Speed maps that aimed the seasons
Unaimed search stalled within three seasons. What restarted it each time was a map. Before a season, the incumbent was profiled on the judge's own workloads, the time of each command was broken down by cause, and candidate changes were prototyped just far enough to measure their size. The teams received the map and a menu with measured sizes, not instructions. 10 maps were made. The first, before season 4, showed that about a fifth of the workload was start-up cost paid by every process; that season promoted +5.25% after two seasons that had promoted nothing. The second, before season 8, preceded the largest single step, +8.42%. When a season ran on a stale map, its first matches were noise.
Hand merges
One tree is promoted per season, so judged work by the other teams is left behind. An automatic merge of the best trees per command was tried in 4 seasons and merged nothing: the teams' trees edited the same files. So at three season boundaries a merge agent ported unpromoted ideas onto the incumbent by hand, one at a time, measuring each, and the merged tree went through the same confirmation as any candidate: after season 4; after season 17 (+6.06%, ahead in 33 of 42); and after season 19 (+1.14%, 36 of 44). The published tree is a fourth hand merge: the repaired tree, the sound part of season 20 and four ports from unpromoted trees.
Soundness
What was checked, and what was found.

A backup program that is fast and wrong is worse than a slow one. The judge compares results, and its warm-up tests a deliberately corrupted pack. It does not exercise two commands working on one repository at once, storage that is damaged while a command runs, or uploads that block. Those were examined by the later reviews and checks. So after season 19 the league's tree was reviewed for soundness before anything was published.
The soundness reviews
Three independent reviewer agents read every changed file of the league tree against upstream: the storage core (34 mechanisms), the backup side (21) and the read side (26). For each defect they wrote a test that reproduces it. Together they made 81 mechanism assessments. Counting a finding shared by several reviewers once, R1 to R11 are 1 blocker, 1 high, 7 medium, 1 low and 1 coverage finding; a memory finding, M22, is separate. None of the three found a weakened hash, MAC or verification, or a change to the repository format, to chunk boundaries or to tree encoding. A fourth review covered what season 20 added (16 mechanisms). A fifth covered the final merge (16 commits): it found nothing above low, and all 17 deliberate mutations the reviewer made to the merged code were caught by the tree's own tests.
The lock-wait defect
restic writes a lock file, waits 200 ms, and looks again for a conflicting lock from another process. From season 13 the league's trees used that wait: they read the snapshot list and the index during it, and kept what they had read. If another process's exclusive work, a prune, completed inside the window, a backup could deduplicate against packs the prune had just deleted. The backup then reported success, and its snapshot pointed at data that no longer existed. All three reviewers found it independently. None of upstream's tests could see it: upstream's integration tests set the lock wait to zero, which switched the league's wait work off.
The repair narrows the assumption. For commands other than check, an index read during the lock wait is used only if a listing taken at upstream's normal point, after the lock is confirmed, still shows its name and size. Index files that were added are loaded, and any removal discards what was read and reloads in upstream's order. Check reads nothing during the wait. Two cases remain and are listed under known limits: damage on storage that does not change a file's length, and a listing that still shows a file after it was deleted. On 3,000 randomised schedules the unrepaired code failed 163 and the repaired code 0. The defect was then reproduced with real binaries.
backup that left a broken snapshotsound backup
Before the repair
37 of 480 broken
37 of 480 backups reported success and left a snapshot that upstream's check rejects.
The published tree
0 of 480 broken
The same trials and delays, after the repair.
Upstream restic
0 of 120 broken
The control: upstream reads nothing during the wait.
forget --prune on the same repository, then checked by upstream restic. The broken ones are drawn first; their position carries no meaning. This is the run made on the judged commit of the published tree. Earlier runs of the same 1,080 trials, on earlier candidates, read 33, 37, 35 of 480 for the unrepaired tree, and 0 of 480 for the candidate and 0 of 120 for upstream each time.Every run as a table
| Run made on | Unrepaired tree: broken snapshots | Tree under test | Upstream restic |
|---|---|---|---|
| the repaired tree, first run | 33 of 480 | 0 of 480 | 0 of 120 |
| a later candidate | 37 of 480 | 0 of 480 | 0 of 120 |
| the candidate before the independent review | 35 of 480 | 0 of 480 | 0 of 120 |
| the judged commit of the published tree | 37 of 480 | 0 of 480 | 0 of 120 |
The defect needed a lock-file write slower than the wait: a remote backend, a retried save, a stall. The trials forced that by holding the write back for 300 to 3,000 ms. On a local disk under no load it may never have fired. It was a defect of the league's work and was never in upstream restic.
Ten more findings
| Severity | Finding | Repair | |
|---|---|---|---|
| R2 | High | Prune's and repair's blob handler ran while holding a backend connection. repair packs with one backend connection deadlocked (upstream passes). A decode error inside the load was retried instead of failing once. | Handlers that use the backend run after the connection is released; handler errors are permanent. |
| R3 | Medium | The hand-written parallel key derivation skipped upstream's parameter checks: a key file with invalid parameters could crash the process, or derive a key where upstream reports an error. | The same checks, with the same errors, before the parallel path. |
| R4 | Medium | A sampler that guessed whether a blob would compress stored a redundant blob raw at 1,048,608 bytes where upstream stores 65,661. The format stayed valid; the repository grew silently. | The sampler was removed. |
| R5 | Medium | Index downloads were capped by the CPU count: 2 at a time instead of 5 on a machine with one or two cores and a remote repository. | Download concurrency follows the backend's connections, as upstream. |
| R6 | Medium | When writes to standard output made no progress (a full disk under a redirected log), every later message, errors included, was held in memory without bound and never printed. | Upstream's behaviour on a failing writer; the broken tests fixed. |
| R7 | Medium | A fast tree decoder accepted node names with a raw quote or newline that upstream rejects, in check, ls and restore. | The fast path is taken only for names with no backslash, quote or control byte. |
| R8 | Medium | Restore rebuilt its hard-link index only for nodes with a link count above 1. A file whose link count changed during the backup restored as two independent files; upstream restores one hard-linked file. | Upstream's hard-link handling for every node. |
| R9 | Medium | A wrong password ran the whole key search twice before reporting the same error. | One search. |
| R10 | Low | Things that should not ship: a stray 1,998-line backup of a source file, files that were not gofmt-clean, code paths that changed when the debug log was on, test-only counters in production code. | Removed, or justified one by one. |
| R11 | Coverage | Upstream's integration tests set the lock wait to 0, which switches the lock-wait work off, so no upstream test ran the defective path. | The integration tests run a second time with the production wait of 200 ms. |
The reviews also recorded a memory finding (M22): restore kept memory for every directory it visited. It was repaired in part, and the release checks found the rest (F5, below).
What was left out of the last season
The tree promoted in season 20 added three new mechanisms inside the lock wait: backup prepared compressed chunks, check read one pack, and restore read the snapshot and its trees. The review found that each was a new entrance to the same defect, and more. The backup one opened every entry of the backup source before the lock was confirmed, named pipes and device nodes included: 3 such opens per backup were counted, against 0 for the published tree. The check one certified a pack from bytes read before the lock, guarded by file-system notifications that network mounts do not deliver. Measured against the same tree without them, the three together cost 0.14 suite points. None of them is in the published tree. The index-side work in check and prune, which carried that season's gain, was judged sound and kept.
Release checks
A check agent then attacked the tree with real binaries on Linux: the tree's own tests and upstream's, the race detector, kills and interrupts in the middle of backup and prune, failing backends, corrupted packs, cross-reading with upstream in both directions, special files, and scale. The checks found 4 real failures, and all were fixed.
| Severity | Found by the release checks | Fix | |
|---|---|---|---|
| F1 | Low, test only | One of the league's own test files did not compile for 32-bit targets. | Fixed. |
| F2 | Medium, no data at risk | prune --repack-uncompressed never came to rest: blobs that do not compress were stored raw, counted as uncompressed and rewritten on every run (16,179 blobs before and after a real run; upstream 0). | Upstream's stored form for every blob. |
| F3 | Low, test only | One regression test sampled a running heap and failed 7 runs of 15. | The test now counts entries. |
| F5 | Low to medium | Restore memory grew by 104 bytes per directory: 231.2 MiB against upstream's 62.2 MiB at 2,020,200 directories. | The list is bounded at 524,288 entries (16 MiB). On the judged commit a restore of 2,020,200 directories peaked at 58.0 MiB against upstream's 56.2 MiB; the unrepaired tree peaked at 627.9 MiB. |
The whole set was run again after the last fixes. All 16 release checks pass on the judged commit. On the published commit, which differs from the judged commit by one comment and one test, the difference between the two source trees was checked and the added test was run, plain and 3 times under the race detector.
| Check | Result | Key numbers | |
|---|---|---|---|
| C01 | The tree's own tests | Pass | 2,692 tests as root, 3,045 as an unprivileged user, 1,383 with the debug log on, 1,818 as a 32-bit build: no failure of the tree's own; 5 shared with upstream in the same environment |
| C02 | Race detector, with the reviewers' witness tests | Pass | 2,426 tests, 0 data races |
| C03 | The lock-wait defect, end to end with real binaries | Pass | broken snapshots: unrepaired tree 37 of 480, this tree 0 of 480, upstream 0 of 120 |
| C04 | kill -9, interrupt and backend faults | Pass | 400 trials, 397 hit a live process, 0 problems; upstream 198 trials, 0 |
| C05 | Detection of a corrupted pack on every route | Pass | 1,200 comparisons with upstream, REST with 2 connections included, 0 differ |
| C06 | Cross-reading with upstream, both directions | Pass | 16 matrix cells and 17 upstream fixtures; 11 repository-size comparisons within 0.0005% of upstream |
| C07 | Tree IDs on special files; restore | Pass | 7 of 7 steps equal; the 32-bit binary gives the same tree ID and an equal restore |
| C08 | Restore option matrix; repair packs | Pass | 75 restore cases, 0 differ; 8 restores under backend failures, SFTP, and prune and copy with 1 and 2 connections equal to upstream; repair packs with one connection finishes in 1.2 s (6 of 6; the unrepaired tree 0 of 6) |
| C09 | Linux-only code; output on a failing writer; node names; wrong password; version string | Pass | 70 output cases, 7 names: 0 differ; the build reports 0.19.1-dev-dealer.1 |
| C10 | Scale | Pass | 500 index files behind 50 ms of latency on 1 CPU: upstream 11.82 s, this tree 11.56 s, the unrepaired tree 19.13 s. Restore of 2,020,200 directories: peak memory 58.0 MiB against upstream's 56.2 MiB. Restore of one 50 GiB sparse file: content equal. A small backup against a synthetic index of 30,000,000 invented entries with no data packs behind them: 6.9 s against upstream's 9.6 s. This does not test a 50 GiB file of incompressible data or a fully populated repository of that size |
| C11 | Prune compared with upstream on a 1.5-million-blob index | Pass | 12 prune pairs equal; with a kept index file damaged before the rewrite, both builds stop with the same error |
| C12 | kill -9 during prune's repack | Pass | 40 trials, 0 problems |
| C13 | What the last season added | Pass | 31 Linux-only tests; 730 tests under the race detector, 0 races; 3,000 lock-wait schedules; 0 opens of a FIFO or device during a backup's lock wait |
| C14 | Prune's repack with a failing backend | Pass | 49 disturbed prunes, 0 problems, 7 of them with a damaged blob while uploads were held: prune returned the error by itself each time; upstream 24 trials, 0 problems |
| C15 | Backup output | Pass | 24 backups, 1,533 lines compared with upstream, 0 differ |
| C16 | copy between repositories | Pass | 4 copy pairs; 6 of 6 copies with a failing source pack and a stalled destination returned the error in about 1 s; upstream's 4 copy tests pass in 3 set-ups |
Three cases were added for that run:
- A kept index file damaged before prune's rewrite. Upstream and this build both stop with the same error, and upstream's check reports the same lines for both.
- A damaged blob while uploads are held. In 7 of 7 trials prune returned the error by itself instead of waiting.
- Copy with failing source packs and a stalled destination. 6 of 6 runs returned the error in about 1 s.
The checks also recorded one performance observation, on restores of very many directories with many CPUs. It is listed under known limits.
An independent review of the finished tree
After the release checks, 3 more reviewers looked at the tree (the "red team"). They were of a different model family from every agent that had built, repaired or reviewed it. They were given only two source trees, upstream's and the candidate's, and a strict contract for the candidate: nothing stored differently from upstream, no integrity check missing, the same behaviour and defaults, safe with concurrent processes and when killed, no data loss, exact restores. They were told to find any situation in which the tree would not meet the contract. They were not shown the earlier reviews or the repairs. Each worked for about 40 minutes and wrote a test for everything it reported.
No reviewer found data loss, a wrong restore, or an incomplete backup reported as a success. They reported 13 findings, 12 of them distinct, shown below in 11 rows. 5 findings carried the label "blocker" because the brief defined a missing integrity check as one: each is a place where the tree read once, or did not bound, something upstream reads twice or bounds. Of the 11 rows, 9 were fixed before the final measurement (one of them in part), each with the reviewer's test kept as a regression test. 2 are differences from upstream that are kept and documented.
| Reviewers' label | Finding | Outcome | |
|---|---|---|---|
| 1 | Blocker | check --read-data decoded compressed blobs as a stream without upstream's ceiling on the total decoded size. A crafted blob that expands past the ceiling passed check, while restore and upstream's check reject it. | Fixed. |
| 2 | Blocker (two reviewers) | State read during the lock wait was validated after the lock was confirmed by file name only. An index file damaged on storage inside the wait, after this build had read it, was not noticed by that run. No restic process can cause this; it needs storage damage in that window. | Fixed in part: the validation now compares size as well as name, and check verifies index files after confirmation, as upstream does. The rest is documented under known limits. |
| 3 | Blocker | Prune's index rewrite kept index files that needed no change without the second read upstream performs. An index that became unreadable or damaged between the load and the rewrite was not noticed, and prune went on. | Fixed: a kept file is read and its hash checked against its name first; a failure stops the rewrite, as upstream does. |
| 4 | Blocker | Restore keeps complete directory trees that it loaded and verified in its first pass, within a budget of 512 KiB of decoded nodes, and uses them in its second pass; upstream loads them again. If a kept tree is damaged on storage between the two passes, upstream's restore fails when its second load reaches the damage, and this build finishes from the valid copy it kept. | Documented under known limits. |
| 5 | High | In repack and copy, a read error could not cancel a save that was blocked behind stalled uploads: the command hung where upstream returns the error. | Fixed. |
| 6 | High | The index kept notes per index file that were never released when the index was cleared and loaded again: memory grew in a long-lived process. | Fixed. |
| 7 | High | After confirmation the validation listed the index before restore had read the chosen snapshot. On a backend whose listings lag, a restore of a snapshot that had just been written failed where upstream succeeds. | Fixed: every command now lists no earlier than upstream does. |
| 8 | Medium | Messages from work done during the wait reached the error output even when the lock was then refused; upstream prints only the lock error. | Fixed. |
| 9 | Medium | Batched terminal output dropped the next message when a write error landed exactly on a message boundary. | Fixed. |
| 10 | Medium | The memory budget of restore's read-ahead did not count a hidden part of a node's extended attributes; crafted tree data could hide megabytes from it. | Fixed. |
| 11 | Medium (two findings) | Index files and lock files are valid Zstandard and readable by upstream, but differ by a few bytes from what upstream's encoder writes for the same content. | Documented under known limits. |
The 11 commits that followed (the fixes, a change to restores of many directories, and the version string) were reviewed in turn by another independent reviewer agent. It found nothing above low severity. It made 25 deliberate mutations to the fixes; the tree's tests caught 23, and the two misses led to 1 more test. It also stated one limit of the lock-wait validation more exactly than before: the listing limit under known limits.
The limits of this review: one workstation, no race detector, no real backend with lagging listings, no runs with several processes, and about 40 minutes per reviewer.
What remains unknown
- No person has audited this code. The reviews were done by AI agents. They find what they look for, and the absence of a finding is not a proof.
- Other operating systems. Nothing was run on macOS or the BSDs. On Windows only upstream's tests and the reviewers' tests ran, on an unprivileged account.
- Real storage services. The tests used local repositories, the judge's REST server and SFTP. No cloud backend was used in any test.
- Long use. Nobody has run this build for months against one repository.
- Very large repositories. The scale check loaded a synthetic index of 30 million invented entries with no data behind them. No fully populated repository of that size was restored, pruned or checked.
Published with it
The instruments are in the repository.

The final judge result files, the sources of the judge and the workloads, the release-check harness, the review reports and the per-pair files of the many-core bench are published beside the source. The result files of the twenty seasons' confirmations and of the hand merges are not included: their numbers are quoted from the season reviews and the merge reports.
- The judge and its workloads
- The program that builds the three arms, runs the paired rounds in a sandbox, checks every result against upstream and writes the result file this page is generated from; the workload generator; and the REST server with the simulated link.
league/tools/judge - The research tools
- The quick bench, the timeline and the parity and memory tools the teams used between judged matches.
league/tools/research-tools - The release-check harness
- The scripts behind checks C01 to C16, including the real-binary reproduction of the lock-wait defect.
league/tools/checks - The merge, bench and release tools
- The tools of the final merge, the many-core bench, the binary-layout check and the release build.
league/tools/merge - The launcher
- The league's configuration as it ran. It is a record: it needs engine modules that are not included and does not run on its own.
league/tools/league - The patch series
- One patch for the league's tree after season 19, then one per later commit: the repairs, the merge and the fixes.
league/series - Result files
- The judge's result files for the final jobs, the per-pair files of the many-core bench, the check tables and the held-out result notes.
league/evidence - The league's own documents
- Scrubbed copies of the season reviews, the soundness reviews, the independent review and its triage, the merge reports, the check logs and the speed maps.
league/reviews
The repository's own documents carry more detail than this page: Results, Methods, Soundness, Red team, Known limits, Tools, Changes.
Known limits
What this does not show.
- One machine, one architecture, 4 CPUs. Every timing in the league's results comes from one Linux machine (amd64), with 4 pinned CPUs per judge job and the corpus, repositories and restore targets in memory (tmpfs). Disk and real-network performance were not measured. Percentages can change with storage cost, latency and backend behaviour: measure your workload.
- Full clock. Every league timing, the final 44 rounds included, ran with the CPUs set to the
performancefrequency governor, which holds them at full clock. Most machines run a governor that follows the load. - More than 4 CPUs. Every league number was measured with 4 pinned CPUs. A later bench of the published machine code at 8 and 16 cores, 5 fixed pairs per cell, found all 56 cells ahead of upstream; the smallest is remote check-read-data at 16 cores, +5.51%. Those are bench estimates, not judge scores. In restore, forget-prune and ls this build uses more user CPU than upstream, and the gap widens with the number of cores.
- Many directories on many CPUs. For a restore that is almost only directories, on a machine with many CPUs under a load-following frequency governor (the usual setting), this build was 4.8 to 10.7% slower than upstream in 4 readings. On 4 cores under the same governor it was 26.8% faster, and with the CPUs held at full clock it was 16 to 18% faster. The measurements, upstream against this build, on the release-check CPUs: 606,060 directories and 3 files, three rotated rounds. Load-following governor, with the official judge job running beside it: 16 CPUs 67.2 s against 74.0 s (10.1% slower); 4 cores 86.5 s against 63.3 s (26.8% faster). The same CPUs held at full clock: 16 CPUs 53.3 s against 44.7 s (16.1% faster); 4 cores 53.5 s against 44.1 s (17.6% faster). Load-following governor again, on a quiet machine: 16 CPUs 68.0 s against 75.3 s (10.7% slower). Larger snapshots under the load-following governor with 16 CPUs: 1,010,100 directories, 7.8% slower (three rounds); 2,020,200 directories, 4.8% slower (one run per build). What this settles: holding the CPUs at full clock removes the loss, and a quiet machine does not. The cache layout was ruled out before: at full clock, on eight cores spread over four cache domains, the same restore was 16 to 17% faster than upstream. So the judge job that ran beside the first measurement did not cause the loss. In the quiet run the mean clock of the running CPUs was 2.55 GHz for upstream and 2.02 GHz for this build, and the CPU seconds were 71.4 against 98.7; at the fixed clock both used about 55. That this build spreads the work over more threads, so that the governor sees less load per CPU and keeps the clock low, is an inference from the clock readings and the CPU seconds, not a measurement. Restore reads trees ahead for the first 524,288 directories of a snapshot and no further.
- The lock-wait validation is a listing. For commands other than
check, an index kept from the lock wait is validated by a listing (name and size) after the lock is confirmed, not by reading the files again. On a backend whose listing can still show a file after it was deleted (a directory cache on a network mount, an eventually consistent store), an index file that a concurrent prune has deleted is not noticed. A backup can then deduplicate against packs the prune just deleted, and report success. It needs two things at once: a lock-file write slow enough for the prune to finish inside the wait, and a listing that lags the deletion. Upstream restic with the old index file in its local cache, which is the default after a first run, goes on with the stale index in the same way. Without a cached copy it fails loudly, because it tries to read the file. - Damage of the same length inside the wait. For backup, restore, ls, forget and prune, an index file is not noticed by that run if it is changed in place on storage without changing its length, after this build read it during the lock wait and before the lock is confirmed. The window is 200 ms plus the time the lock file takes to write. No restic process can cause it: index files are never rewritten in place. Upstream notices such damage only if it reads the damaged file from the backend and not a valid copy from its local cache. A change of length is detected, and
checkis not affected. - Restore reuses directory trees it has verified. Restore keeps complete, verified directory trees from its first pass, within a budget of 512 KiB of decoded nodes, and uses them in its second pass; other trees are loaded again. If storage damages a kept tree between the two passes, this build can finish from the valid copy it kept. Upstream fails when its second load reaches the damaged storage; a valid local cache can hide the same damage from upstream too. What is restored is the snapshot's content.
- Index and lock files are encoded differently. They decode to the same content and are valid Zstandard that upstream reads, but the compressed bytes differ by a few bytes from upstream's. Chunking, blob IDs and the stored form of data and tree blobs are upstream's.
- Three smaller differences. The review of the fixes recorded them. One message about clearing the local cache can be printed before a lock is refused, where upstream prints only the lock error. The ceiling on decoded size in
check --read-datais the constant 64 GiB, upstream's default. When a write error falls inside a message that spans a flush of the output, the rest of the message is written where upstream gives the whole message up. - The remote workload is a simulation. It is a small REST server with a fixed 30 ms delay per request, 48 MiB/s per connection and 5 connections. It is not a real network, and no cloud storage service was measured.
- Cold cache. Every phase starts with an empty local cache. Warm-cache repeat backups were not measured as a phase of their own.
- Linux is the only tested platform. All timing and all release checks ran on Linux (amd64). On Windows the tree was tested only as far as upstream's own tests and the reviewers' tests run on an unprivileged account: no shadow-copy backups, no elevated runs, no timing, no release checks. On macOS, the BSDs and the other targets nothing was run. All 25 of upstream's release targets built on an earlier tree. From the published commit five release targets were built: Linux and macOS for x86-64 and arm64, and Windows for x86-64. The other targets have not been built from it.
- It would not have been promoted. The published tree is slower than the league's last champion and uses more memory than it, while using less of both than upstream. By the league's own rules it would not have replaced the champion. The difference is the price of the correctness repairs and of the fixes from the independent review; the figures are in the results above.
- The held-out workload was not run on the published tree. The last tree measured on it is the one after season 19.
- The repairs cost requests on remote repositories. After the lock is confirmed, each process lists snapshots and index files again: one or two more requests per process, 29 to 63 ms at 30 ms per request.
- Binary layout. The key derivation's inner loop is about 3% slower when the compiler places it at one of two alignments. The judged Linux build is on the fast one. A build for another platform, or with another toolchain, lands where it lands: about 7 ms per process either way.
- Real storage services. No cloud storage service was used in any measurement or check. Several upload failures at once on a real service were not exercised for the reworked repack loop.
- Scale. The scale check restored a 50 GiB sparse file and loaded a synthetic index of up to 30 million invented entries, with no data packs behind them, during a small backup. It does not test a 50 GiB file of incompressible data, or a fully populated repository of that size being restored, pruned or checked.
- Other tools. rustic, Kopia and Borg were measured once, against the tree after season 15, on one workload. They were not measured against the published tree.
self-updateinstalls upstream. The build reports its version as0.19.1-dev-dealer.1, and snapshots it writes carry that string as their program version. Do not runrestic self-updatewith it: that command fetches upstream restic's release binary and would replace this build.- Some error messages still point at upstream. Eight error messages inherited from upstream still tell the user to open an issue with the restic project. They are unchanged in this release, because changing them would change the binary after it was judged, reviewed and checked. For this build always report at the Dealer Style Restic issues page, never to upstream.
- No support, no audit by people. The reviews, repairs, merge and release checks described here were done by AI agents under one operator's direction. No human security audit has been done. The restic project has not reviewed or endorsed any of it.
Source and verification
Build it. Verify it.
It needs Go 1.25.8 or later, as upstream does; the judge built with go1.27.0. Use go build as shown: upstream's build.go script passes a version flag for a symbol this tree no longer has.
git clone https://github.com/slopdealer/Dealer-Style-Restic cd Dealer-Style-Restic go build ./cmd/restic ./restic version
The build reports its version as 0.19.1-dev-dealer.1. Do not run restic self-update with this build. That command fetches upstream restic's release and would replace this binary with upstream's.
Then verify it on a test repository before anything else. Upstream restic is the referee: it must accept what this build wrote.
./restic init --repo /tmp/dsr-test ./restic --repo /tmp/dsr-test backup /path/to/a/copy/of/real/data restic --repo /tmp/dsr-test check --read-data # upstream restic checks what this build wrote ./restic --repo /tmp/dsr-test restore latest --target /tmp/dsr-restore diff -r /path/to/a/copy/of/real/data /tmp/dsr-restore/path/to/a/copy/of/real/data
To repeat the measurements, start at METHODS.md: the judge, the workload generator and the instructions are in the repository. Your files, your storage and your CPU count will move the result; measure before you rely on it.
Downloads and checksums
| File | Platform | SHA-256 |
|---|---|---|
| restic_0.19.1-dev-dealer.1_linux_amd64 | Linux, x86-64 | 753924dbb9e6bbb2876511c533f722ce5d68d2559928de9a6217b61b80f6f8c1 |
| restic_0.19.1-dev-dealer.1_linux_arm64 | Linux, arm64 | 7ed7c72be44fdbcb4eaa0935c4f674c1f7572c830a87f16a758d4dea28c087d5 |
| restic_0.19.1-dev-dealer.1_darwin_amd64 | macOS, x86-64 | 516332fc04af27fcf8816f7381efb867804e595faec8b25e32a1feeefa8bc010 |
| restic_0.19.1-dev-dealer.1_darwin_arm64 | macOS, arm64 | e26c5e0bdcff313211874299bd0367f305d4ba14919839e0c643652713cc4c9a |
| restic_0.19.1-dev-dealer.1_windows_amd64.exe | Windows, x86-64 | 7b4538e1902730a793542b28e881ce7193796105f8b08dd64a4c3646753e1382 |
The release also carries SHA256SUMS and restic's unchanged LICENSE. An independent rebuild from the published source, with the documented command, reproduced all 5 binaries bit for bit.
What these binaries are, exactly. The published source builds to the judged machine code byte for byte: the two builds differ in 60 bytes, all inside build-ID notes. The release binaries are that build with the version string set: the same instructions at the same addresses, with different address operands. They were not themselves judged. One timing comparison of one command was made over 20 fixed pairs, and it resolved no difference.
An independent Slop Dealer build. restic is distributed under the BSD-2-Clause licence, and this derivative keeps it. Report problems with this build to its own repository, not to the restic project.