Tolerate a file that vanishes during the hash pass - #156
Merged
Merged
Conversation
After hashing, FindDuplicates sized each group with FileInfo.Length on its first file, and Stats summed FileInfo.Length over every hashed file. A file removed or renamed during a multi-minute hash pass made either throw an uncaught FileNotFoundException, losing the whole DryRun, Deduplicate, Scan or Stats run -- often over an unrelated file. FindDuplicates now sizes each copy tolerantly, drops a copy that has vanished, and drops a group left with fewer than two copies. DuplicateGroup takes its size from that pass instead of reading the disk itself. Stats sums sizes through the same tolerant lookup, counting nothing for a vanished file. Fixes #141 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015LsGhhi4Vq9pNq2fr5T75T
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015LsGhhi4Vq9pNq2fr5T75T
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.



Fixes #141
Problem
After the hash pass, two places read file sizes from disk again without guarding against a file that had since disappeared:
DuplicateGroupconstructor readnew FileInfo(files[0]).Length. It runs insideFindDuplicates, which Scan, DryRun and Deduplicate all use.new FileInfo(f).Lengthover every hashed file.If a file was removed or renamed during a multi-minute hash pass, an uncaught
FileNotFoundExceptionthrew away the whole run. The missing file was often unrelated to any duplicate.Change
FindDuplicatesnow sizes each copy with a tolerant lookup,TryGetSize, which catchesIOException(including not-found) andUnauthorizedAccessException.DuplicateGrouptakes its size as a constructor argument instead of reading the disk itself.Deduplicator.TotalSize, which counts nothing for a vanished file.DeleteDuplicatesis unchanged. It already re-hashes each copy before deleting it and skips any that can't be confirmed.I didn't take the issue's other suggestion of recording sizes during hashing. The acceptance criteria require grouping to notice a copy that has vanished, which recorded sizes can't tell you. The existence check at grouping time handles both halves and leaves
HashFilesand its roughly 20 test call sites unchanged.Tests
New tests in
DeduplicatorTests:files[0]crash path).TotalSizeskips a vanished file.I checked that these catch the bug by temporarily making the size lookup unguarded, as the old code was: all four test cases failed, and they pass with the fix. Full suite: 59 passed and 2 skipped locally. The two skips are existing permission tests that can't run as root.
🤖 Generated with Claude Code
https://claude.ai/code/session_015LsGhhi4Vq9pNq2fr5T75T
Generated by Claude Code