mirror of
https://github.com/borgbackup/borg.git
synced 2026-09-02 14:43:21 +02:00
Add a new "fastcdc" content-defined chunker selectable via --chunker-params.
It uses the FastCDC Gear rolling hash (fp = (fp << 1) + Gear[byte]), which is
window-less and cheaper per byte than buzhash's cyclic-polynomial update, so it
chunks noticeably faster (see "borg benchmark cpu" output), while producing
the same chunk-size distribution and deduplication.
The Gear table is keyed: it is derived from the repo id key via CSPRNG (own
"fastcdc" domain), exactly like the buzhash64 table, so chunk cut points stay
unpredictable without the key (anti-fingerprinting). It implements the same
FastCDC techniques as buzhash64 (sub-minimum skipping, normalized chunking with
a required nc_level, min/max clamping); the mask uses the high bits of the hash
(Gear accumulates entropy there).
chunker-params: "fastcdc,chunk_min,chunk_max,chunk_mask,nc_level" - there is no
window field, because Gear is window-less. e.g. fastcdc,19,23,21,2
Also: borg benchmark cpu now measures the fastcdc chunker; tests in
borg.testsuite.chunkers (golden vector, size distribution, keyed gear table,
param parsing, slow fuzz); docs and changelog.
Benchmarks (scripts/chunker_bench.py, buzhash64 vs fastcdc, both nc_level=2,
incompressible data unless noted):
5 GiB, 2 MiB target (default params):
buzhash64: CV 0.294, 1011 MB/s
fastcdc: CV 0.295, 1313 MB/s (+30%)
64 MiB, 64 KiB target:
buzhash64: CV 0.374, shift-resilience 0.9928, 963 MB/s
fastcdc: CV 0.359, shift-resilience 0.9929, 1331 MB/s (+38%)
Re-backup of a 2.5 GiB file after scattered single-byte edits (dedup ratio,
0.5 = v2 fully deduplicated, lower is better):
64 edits: buzhash64 0.5237, fastcdc 0.5236
320 edits: buzhash64 0.6133, fastcdc 0.6161
borg benchmark cpu, 1 GB: fastcdc 3.80s, buzhash 4.36s, buzhash64 8.13s,
fixed 0.56s.
Chunk-size distribution, deduplication and shift-resilience match buzhash64
within noise; fastcdc is consistently faster.
Also: fix bug when computing the mask, one needs to use 1ULL instead of
1, so the shifting computation is done in a uint64, not in a 32bit int.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
30 lines
1.5 KiB
C++
30 lines
1.5 KiB
C++
.. highlight:: bash
|
|
.. |package_dirname| replace:: borgbackup-|version|
|
|
.. |package_filename| replace:: |package_dirname|.tar.gz
|
|
.. |package_url| replace:: https://pypi.org/project/borgbackup/#files
|
|
.. |git_url| replace:: https://github.com/borgbackup/borg.git
|
|
.. _github: https://github.com/borgbackup/borg
|
|
.. _issue tracker: https://github.com/borgbackup/borg/issues
|
|
.. _deduplication: https://en.wikipedia.org/wiki/Data_deduplication
|
|
.. _AES: https://en.wikipedia.org/wiki/Advanced_Encryption_Standard
|
|
.. _HMAC-SHA256: https://en.wikipedia.org/wiki/HMAC
|
|
.. _SHA256: https://en.wikipedia.org/wiki/SHA-256
|
|
.. _PBKDF2: https://en.wikipedia.org/wiki/PBKDF2
|
|
.. _argon2: https://en.wikipedia.org/wiki/Argon2
|
|
.. _ACL: https://en.wikipedia.org/wiki/Access_control_list
|
|
.. _libacl: https://savannah.nongnu.org/projects/acl/
|
|
.. _libattr: https://savannah.nongnu.org/projects/attr/
|
|
.. _liblz4: https://github.com/Cyan4973/lz4
|
|
.. _libzstd: https://github.com/facebook/zstd
|
|
.. _OpenSSL: https://www.openssl.org/
|
|
.. _`Python 3`: https://www.python.org/
|
|
.. _Buzhash: https://en.wikipedia.org/wiki/Buzhash
|
|
.. _FastCDC: https://www.usenix.org/conference/atc16/technical-sessions/presentation/xia
|
|
.. _msgpack: https://msgpack.org/
|
|
.. _`msgpack-python`: https://pypi.org/project/msgpack-python/
|
|
.. _llfuse: https://pypi.org/project/llfuse/
|
|
.. _mfusepy: https://pypi.org/project/mfusepy/
|
|
.. _pyfuse3: https://pypi.org/project/pyfuse3/
|
|
.. _userspace filesystems: https://en.wikipedia.org/wiki/Filesystem_in_Userspace
|
|
.. _Cython: https://cython.org/
|
|
.. _virtualenv: https://pypi.org/project/virtualenv/
|