Skip to content

Return raw bytes from webpages.fetch; decode HTML once in extract() - #113

Draft
rahulbot with Copilot wants to merge 3 commits into
mainfrom
copilot/improvement-content-decoding-encoding
Draft

rahulbot with Copilot wants to merge 3 commits into
mainfrom
copilot/improvement-content-decoding-encoding

Conversation

Copilot AI commented Aug 26, 2026 •

Copy link
Copy Markdown

webpages.fetch was returning response.text, causing requests to auto-decode bytes→str, only for downstream code (e.g. languages._from_text) to re-encode str→bytes before passing to libraries that expect bytes. This redundant round-trip risks encoding errors and wastes CPU.

Changes

  • webpages.fetch() now returns response.content (bytes) instead of response.text (str)
  • New webpages.html_from_bytes() decodes raw bytes → str once using trafilatura.utils.decode_file, which correctly handles charset declarations embedded in the HTML
  • extract() accepts a new optional html_bytes: bytes parameter; bytes (from fetch or caller) are decoded once via html_from_bytes() before all downstream processing
  • languages._from_text() no longer redundantly calls trafilatura.utils.decode_file on an already-decoded str

Usage

Callers that fetch HTML themselves can now pass raw bytes directly:

import mcmetadata
import requests

resp = requests.get(url)
results = mcmetadata.extract(url, html_bytes=resp.content)

The existing html_text: str parameter remains fully supported for backward compatibility.

… html_bytes in extract()

Co-authored-by: rahulbot <673178+rahulbot@users.noreply.github.com>
Copilot AI changed the title [WIP] Improve content decoding and encoding in webpages.fetch Return raw bytes from webpages.fetch; decode HTML once in extract() Aug 26, 2026
Copilot AI requested a review from rahulbot August 26, 2026 12:21

@rahulbot rahulbot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Rebase this off of the fix-more-unit-tests branch, which should help the unit tests pass.

…sts-mock, update all tests

Co-authored-by: rahulbot <673178+rahulbot@users.noreply.github.com>

Copilot AI commented Aug 26, 2026

Copy link
Copy Markdown
Author

Rebase this off of the fix-more-unit-tests branch, which should help the unit tests pass.

Done — rebased onto fix-more-unit-tests by incorporating all its changes (commit 396fe29 and 3a7f45a): webpages.final_url() is now used in extract(), all test files use mock_fetch with requests-mock, and requests-mock was added to the test dependencies in pyproject.toml. The mock_fetch helper was updated to work with the new bytes-returning webpages.fetch() API.

Copilot AI requested a review from rahulbot August 26, 2026 14:13

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Improvement regarding content decoding/encoding

2 participants