Skip to content

Recover named references and truncated HTML text - #15

Merged
rogerchappel merged 3 commits into
mainfrom
agent/oss-7366ad4ca305-html-recovery
Aug 28, 2026
Merged

rogerchappel merged 3 commits into
mainfrom
agent/oss-7366ad4ca305-html-recovery

Conversation

@rogerchappel

Copy link
Copy Markdown
Owner

Summary

  • decode semicolon-terminated HTML Latin-1 named references while leaving semicolonless and unknown names unchanged
  • omit comments that run to end-of-input and recognize only syntactically tag-shaped markup, preserving ordinary </> comparisons
  • add focused unit coverage, an executable recovery fixture, and aligned API/safety documentation

Commits

  • 0fa1dab test: define deterministic HTML recovery
  • 24fb68e fix: recover common HTML text edge cases
  • 940aaa8 docs: exercise supported HTML recovery

Verification

  • npm run release:check (36 tests, syntax, build, CLI smoke, package smoke, release contract)
  • bash scripts/validate.sh
  • git diff --check origin/main...HEAD

All commits are authored and committed by Roger Chappel miscanalysis@gmail.com.

@rogerchappel

Copy link
Copy Markdown
Owner Author

Automated merge note

Triage class: auto-merge

Summary: deterministic local HTML-to-text recovery for named entities, truncated comments, and text comparisons (8 files, +46/-7).

Checks run: CI test completed SUCCESS; required review/branch protection re-read.

Rebased/CI-repaired: no.

Verified head SHA: 940aaa8b74aa9bfd91e1fe22a6599f2205d0f9ed.

@rogerchappel
rogerchappel merged commit 1eed950 into main Aug 28, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant