What's wrong
FileScanner.ScanForFiles (FileDeduplicator/FileScanner.cs) skips only FileAttributes.ReparsePoint while walking. A Linux bind mount, a share mounted twice (NFS or SMB), or a macOS directory hardlink is not a reparse point, so the walk goes into both paths and lists one file twice. Both paths hash the same, so they end up in one duplicate group. Deduplicate then re-hashes the keeper, which passes, and deletes the "other copy".
That other copy is the same directory entry as the keeper, so deleting it deletes the keeper too.
This is worse than #132 (hardlinks). With a hardlink, the other link keeps the data alive, so what you lose is a name. Here nothing else holds the data, so the content is gone. The keeper re-hash, which is the one safeguard against data loss before deleting, can't catch it, because the keeper is the file being deleted.
Reproduction (Linux, as root)
mkdir -p t1/real t1/view
echo "only copy of important data" > t1/real/important.txt
mount --bind t1/real t1/view
printf 'y\n' | dotnet ktsu.FileDeduplicator.dll Deduplicate -p t1
Output:
KEEP: .../t1/real/important.txt
DELETE: .../t1/view/important.txt
Deleted: .../t1/view/important.txt
Deleted 1 file(s). Reclaimed 28 B of disk space.
After that, ls -R t1 shows both real/ and view/ empty. The only copy of the file has been deleted, and the tool reports it as reclaimed space.
Expected: the tool sees that the two paths are one file and leaves it alone. Nothing should be deleted or reported as reclaimed.
Suggested fix
Acceptance criteria
- A tree where one directory is visible at two paths ends a
Deduplicate run with every file still present and readable. The run reports 0 deleted and 0 bytes reclaimed for those files.
Scan, DryRun and Stats do not report a single file seen at two paths as a duplicate group.
What's wrong
FileScanner.ScanForFiles(FileDeduplicator/FileScanner.cs) skips onlyFileAttributes.ReparsePointwhile walking. A Linux bind mount, a share mounted twice (NFS or SMB), or a macOS directory hardlink is not a reparse point, so the walk goes into both paths and lists one file twice. Both paths hash the same, so they end up in one duplicate group.Deduplicatethen re-hashes the keeper, which passes, and deletes the "other copy".That other copy is the same directory entry as the keeper, so deleting it deletes the keeper too.
This is worse than #132 (hardlinks). With a hardlink, the other link keeps the data alive, so what you lose is a name. Here nothing else holds the data, so the content is gone. The keeper re-hash, which is the one safeguard against data loss before deleting, can't catch it, because the keeper is the file being deleted.
Reproduction (Linux, as root)
Output:
After that,
ls -R t1shows bothreal/andview/empty. The only copy of the file has been deleted, and the tool reports it as reclaimed space.Expected: the tool sees that the two paths are one file and leaves it alone. Nothing should be deleted or reported as reclaimed.
Suggested fix
(st_dev, st_ino)on Unix, and(VolumeSerialNumber, FileIndexHigh/Low)fromGetFileInformationByHandleon Windows. Keep one path per identity. This also fixes Hardlinked files are deduplicated like independent copies, risking silent data loss in snapshot/backup trees #132.SkippedFile. That way no path to the keeper's own data is ever deleted, whatever the scan produced.Acceptance criteria
Deduplicaterun with every file still present and readable. The run reports 0 deleted and 0 bytes reclaimed for those files.Scan,DryRunandStatsdo not report a single file seen at two paths as a duplicate group.