Skip to content

Import accepts upper-case hashes and null fields: the entries never match a Scan (duplicate descriptions) or crash Search/Stats/Export on every later run #163

Description

@matt-edmondson

What's wrong

Import.MergeEntries (ImageDescriber/Verbs/Import.cs ~88-130) validates an imported entry only with ImageHasher.IsValidHash, then stores it as-is under entry.Hash. Two gaps follow from that.

(a) Upper-case hashes are accepted but can never match

  • IsValidHash (ImageHasher.cs:52) uses char.IsAsciiHexDigit, so it accepts upper-case hex.
  • ComputeHash (ImageHasher.cs:61) always produces lower-case (Convert.ToHexStringLower).
  • PersistentState.Descriptions is a default, case-sensitive Dictionary<string, ImageDescription>.

The result:

  • An imported entry with hash AB12… is never found by Scan's descriptions.TryGetValue(hash, …) (Scan.cs:84). The image is described again and stored as a second entry.
  • Importing the same hash once upper-case and once lower-case gives two entries. MergeEntries returns (1, 0, 0) both times.

Upper-case SHA-256 is what several external tools produce (for example PowerShell Get-FileHash), so hand-built or tool-built import files hit this.

(b) JSON null fields are saved and poison the store

An entry like {"Hash":"<64 hex>","Description":null, ...} or "KnownPaths":null is accepted and saved. After that:

  • Search throws (Search.cs:29, .Description.Contains(...) on null).
  • Stats throws a NullReferenceException (Stats.cs:63).
  • CSV export throws (Export.cs:77). A "SuggestedFileName": null alone is enough to break it.
  • Scan with KnownPaths: null throws a NullReferenceException at Scan.cs:87, outside any try, the next time it meets that image.

These failures repeat on every run until someone edits the JSON by hand. It is the same kind of permanently broken store as #151, reached through different fields.

Both cases were reproduced with a scratch MSTest probe. It calls Import.MergeEntries with such entries and then runs Search, Stats and Export.BuildCsv, which fail exactly as described above.

Suggested fix

  • Normalize the hash with entry.Hash = entry.Hash.ToLowerInvariant() before validating and storing it, and use that as the dictionary key. Alternatively, give Descriptions StringComparer.OrdinalIgnoreCase, but normalizing also keeps the stored data canonical.
  • Repair or reject entries whose KnownPaths, Description, SuggestedFileName or Model are null. Defaulting them to empty, or skipping the entry with a message as is done for invalid hashes, would both work.

Acceptance criteria

  • If you import an entry with an upper-case hash and then scan the matching file, the path is added to that entry and no new entry is created.
  • Importing the same hash in upper and lower case gives one entry.
  • After importing entries with null fields, Search, Stats, Export (CSV and JSON) and Scan all still work.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't workingreadyFully specified; implement as written

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions