What's wrong
FileDeduplicator/FileScanner.cs:38-43 sets AttributesToSkip = FileAttributes.ReparsePoint. The remarks above it (lines 23-26) say the intent is to avoid descending into directory symlinks that point back at an ancestor, which is the fix for #125.
On Windows, FILE_ATTRIBUTE_REPARSE_POINT is not limited to links. It is also set on:
- OneDrive / Files-On-Demand (Cloud Files API) files and folders, which carry
IO_REPARSE_TAG_CLOUD_* tags. This includes files already hydrated locally, which is why Get-ChildItem lists them with mode l.
- Files on Windows Server Data Deduplication volumes (
IO_REPARSE_TAG_DEDUP).
- Some other filter-driver placeholders, such as HSM and backup/archive stubs.
FileSystemEnumerator applies AttributesToSkip before it decides whether to recurse. A skipped cloud folder is therefore never entered.
Failure scenario
On a Windows machine with OneDrive:
- Run
Scan -p C:\Users\me\OneDrive\Pictures.
- The tool reports far fewer files than the folder holds, possibly zero.
- It prints no warning, so the result looks like a clean scan with no duplicates.
A OneDrive photo library is one of the most likely targets for this tool, and the README promises that it "Recursively scans all files".
How this was verified
This comes from tracing the code against documented Windows and .NET behaviour. It has not been reproduced: this review ran on Linux, where only symlinks carry the attribute.
Suggested fix
Skip only reparse points that are real links, meaning name surrogates such as symlinks and junctions:
- Replace the plain
EnumerationOptions walk with a FileSystemEnumerable<string>.
- In both
ShouldIncludePredicate and ShouldRecursePredicate, reject an entry only when it has FileAttributes.ReparsePoint and entry.ToFileSystemInfo().LinkTarget != null. LinkTarget is null for reparse points that are not links, such as cloud placeholders.
Hashing a placeholder that is not hydrated forces a download. Either accept that, or skip entries that have FileAttributes.Offline or FILE_ATTRIBUTE_RECALL_ON_DATA_ACCESS (0x400000) and report how many were skipped.
Acceptance criteria
What's wrong
FileDeduplicator/FileScanner.cs:38-43setsAttributesToSkip = FileAttributes.ReparsePoint. The remarks above it (lines 23-26) say the intent is to avoid descending into directory symlinks that point back at an ancestor, which is the fix for #125.On Windows,
FILE_ATTRIBUTE_REPARSE_POINTis not limited to links. It is also set on:IO_REPARSE_TAG_CLOUD_*tags. This includes files already hydrated locally, which is whyGet-ChildItemlists them with model.IO_REPARSE_TAG_DEDUP).FileSystemEnumeratorappliesAttributesToSkipbefore it decides whether to recurse. A skipped cloud folder is therefore never entered.Failure scenario
On a Windows machine with OneDrive:
Scan -p C:\Users\me\OneDrive\Pictures.A OneDrive photo library is one of the most likely targets for this tool, and the README promises that it "Recursively scans all files".
How this was verified
This comes from tracing the code against documented Windows and .NET behaviour. It has not been reproduced: this review ran on Linux, where only symlinks carry the attribute.
Suggested fix
Skip only reparse points that are real links, meaning name surrogates such as symlinks and junctions:
EnumerationOptionswalk with aFileSystemEnumerable<string>.ShouldIncludePredicateandShouldRecursePredicate, reject an entry only when it hasFileAttributes.ReparsePointandentry.ToFileSystemInfo().LinkTarget != null.LinkTargetis null for reparse points that are not links, such as cloud placeholders.Hashing a placeholder that is not hydrated forces a download. Either accept that, or skip entries that have
FileAttributes.OfflineorFILE_ATTRIBUTE_RECALL_ON_DATA_ACCESS(0x400000) and report how many were skipped.Acceptance criteria