What's wrong
Scan, DryRun, Deduplicate and Stats (FileDeduplicator/Verbs/*.cs, around line 50 in each) pass the full scanned file list straight to FileHasher.HashFiles. FileHasher.cs:21-27 computes a full SHA-256 of every file, and files are then grouped by hash alone (Deduplicator.GroupByHash). Two files can only be duplicates if they are the same length, but the tool never checks sizes first. So every file with a size unique in the tree is read end to end for nothing.
Repro
Create 40 random files of 20,000,001 to 20,000,040 bytes (763 MB, no two the same size), then run ktsu.FileDeduplicator Scan -p <dir>.
- Observed: 40
Hashed: lines, meaning all 763 MB were read, followed by "No duplicate files found."
- Expected: nothing is hashed, because no two files can match.
This took 1.9 s here only because the files were already in the page cache. On a cold disk, a USB drive or a network share, which are typical deduplication targets, the cost is a full read of the whole tree. In a media library most large files have unique sizes, so most of the I/O is wasted.
Suggested fix / acceptance criteria
- Add one helper that all four verbs use. It groups
ScanForFiles results by new FileInfo(path).Length, drops groups with a single file, and hashes only the files that remain.
- Change
Stats so that files skipped because their size is unique still count toward "Unique files". It currently derives that figure from hashGroups.Count.
- Optional: hash the first 4–64 KB of each file as a second filter before the full hash.
- Keep the existing behavior where zero-byte files are duplicates of each other (
EmptyFilesAreDuplicatesOfEachOther).
- Add a test that a tree of files with unique sizes produces no hashing calls.
What's wrong
Scan,DryRun,DeduplicateandStats(FileDeduplicator/Verbs/*.cs, around line 50 in each) pass the full scanned file list straight toFileHasher.HashFiles.FileHasher.cs:21-27computes a full SHA-256 of every file, and files are then grouped by hash alone (Deduplicator.GroupByHash). Two files can only be duplicates if they are the same length, but the tool never checks sizes first. So every file with a size unique in the tree is read end to end for nothing.Repro
Create 40 random files of 20,000,001 to 20,000,040 bytes (763 MB, no two the same size), then run
ktsu.FileDeduplicator Scan -p <dir>.Hashed:lines, meaning all 763 MB were read, followed by "No duplicate files found."This took 1.9 s here only because the files were already in the page cache. On a cold disk, a USB drive or a network share, which are typical deduplication targets, the cost is a full read of the whole tree. In a media library most large files have unique sizes, so most of the I/O is wasted.
Suggested fix / acceptance criteria
ScanForFilesresults bynew FileInfo(path).Length, drops groups with a single file, and hashes only the files that remain.Statsso that files skipped because their size is unique still count toward "Unique files". It currently derives that figure fromhashGroups.Count.EmptyFilesAreDuplicatesOfEachOther).