What's wrong
FileScanner.WalkOptions (FileDeduplicator/FileScanner.cs) correctly skips symlinks via AttributesToSkip = FileAttributes.ReparsePoint (with a doc comment explaining exactly why: "a symlink is not a second copy, so deleting one reclaims nothing"). But there is no equivalent check anywhere in the scanner or Deduplicator for hardlinks — two directory entries sharing the same inode/file ID. Nothing inspects link count or file identity (e.g. via GetFileInformationByHandle/stat.st_ino), so two hardlinked paths hash identically and are grouped as ordinary duplicates.
Concrete failure scenario
A user runs the deduplicator over a backup tree produced by a snapshot tool that intentionally uses hardlinks to save space (e.g. rsync --link-dest, cp -al, rsnapshot-style rotating backups). Each snapshot directory has its own directory entry for a file, but multiple snapshots' entries point at the same inode. The tool sees these as N duplicate copies, keeps only the one SelectFileToKeep prefers, and deletes the directory entries for every other snapshot via File.Delete. That file is now silently gone from every other snapshot's listing, defeating the point of keeping separate snapshots — with no warning that the "duplicate" was actually a hardlink rather than an independent copy.
A secondary, same-root-cause issue: the reported "bytes reclaimed" figures multiply file size by (copies - 1) regardless of whether copies are hardlinks. Deleting a hardlink frees no disk space at all (the data blocks are still referenced by the surviving link), so the reported reclaimed-space total is inaccurate whenever hardlinks are involved.
Suggested fix
Before grouping, dedupe by (device, file ID) as well as by hash — either skip additional hardlinks to an already-seen inode entirely (treat as a single file, matching the existing symlink-skip policy), or at minimum flag hardlinked members of a duplicate group distinctly in the keep/delete listing and exclude their size from "bytes reclaimed" totals.
What's wrong
FileScanner.WalkOptions(FileDeduplicator/FileScanner.cs) correctly skips symlinks viaAttributesToSkip = FileAttributes.ReparsePoint(with a doc comment explaining exactly why: "a symlink is not a second copy, so deleting one reclaims nothing"). But there is no equivalent check anywhere in the scanner orDeduplicatorfor hardlinks — two directory entries sharing the same inode/file ID. Nothing inspects link count or file identity (e.g. viaGetFileInformationByHandle/stat.st_ino), so two hardlinked paths hash identically and are grouped as ordinary duplicates.Concrete failure scenario
A user runs the deduplicator over a backup tree produced by a snapshot tool that intentionally uses hardlinks to save space (e.g.
rsync --link-dest,cp -al, rsnapshot-style rotating backups). Each snapshot directory has its own directory entry for a file, but multiple snapshots' entries point at the same inode. The tool sees these as N duplicate copies, keeps only the oneSelectFileToKeepprefers, and deletes the directory entries for every other snapshot viaFile.Delete. That file is now silently gone from every other snapshot's listing, defeating the point of keeping separate snapshots — with no warning that the "duplicate" was actually a hardlink rather than an independent copy.A secondary, same-root-cause issue: the reported "bytes reclaimed" figures multiply file size by
(copies - 1)regardless of whether copies are hardlinks. Deleting a hardlink frees no disk space at all (the data blocks are still referenced by the surviving link), so the reported reclaimed-space total is inaccurate whenever hardlinks are involved.Suggested fix
Before grouping, dedupe by (device, file ID) as well as by hash — either skip additional hardlinks to an already-seen inode entirely (treat as a single file, matching the existing symlink-skip policy), or at minimum flag hardlinked members of a duplicate group distinctly in the keep/delete listing and exclude their size from "bytes reclaimed" totals.