Skip to content

perf: bound Python lexer buffering and preserve JSON conversion - #183

Closed
RyaliNvidia wants to merge 1 commit into
ICRAR:masterfrom
RyaliNvidia:codex/streaming-lexer-memory
Closed

RyaliNvidia wants to merge 1 commit into
ICRAR:masterfrom
RyaliNvidia:codex/streaming-lexer-memory

Conversation

@RyaliNvidia

@RyaliNvidia RyaliNvidia commented Oct 1, 2026 •

Copy link
Copy Markdown

Large documents containing many strings can retain already-consumed prefixes in the Python lexer's buffer. Compact completed prefixes at a bounded threshold while preserving unfinished tokens and absolute error offsets. A subprocess regression increases generated input from 16 MiB to 256 MiB and requires peak RSS growth of at most 48 MiB.

The Python backend also gains an explicit allow_float_overflow option requiring use_float=True. It preserves Python float conversion for valid JSON exponent overflow while keeping arbitrary-size integer conversion and existing defaults. Exact JSON number syntax and whitespace checks prevent Python's permissive numeric conversions or Unicode whitespace matching from accepting malformed input.

Tests cover chunk boundaries, escaped strings and surrogate identities, offsets, overflow/underflow, malformed syntax, and memory growth. No parser fallback or runtime patch is introduced.

Validation: the upstream test suite passed with 621 tests and 18 skips. Independent chunk/offset and standard-library differential checks also passed.

Summary by Sourcery

Bound Python lexer memory usage and make JSON numeric conversion behavior explicit and standards-compliant.

New Features:

  • Add an explicit Python-backend option to permit float overflow conversion while retaining arbitrary-size integer handling and requiring float conversion.

Bug Fixes:

  • Bound Python lexer buffer growth while preserving partial tokens and absolute error offsets.
  • Enforce exact JSON number syntax and ASCII-only JSON whitespace without weakening string Unicode handling.

Tests:

  • Add streaming, chunk-boundary, numeric conversion, malformed-input, Unicode, offset, and bounded-memory regression coverage.

@sourcery-ai

sourcery-ai Bot commented Oct 1, 2026 •

Copy link
Copy Markdown

Reviewer's Guide

The PR bounds Python lexer buffering through periodic compaction, preserves streaming token and error-offset semantics, adds an explicit opt-in for Python float overflow conversion, and enforces exact JSON number and whitespace syntax with comprehensive regression and memory tests.

Sequence diagram for Python float overflow handling

sequenceDiagram
    participant Caller
    participant Backend as PythonBackend
    participant Parser as parse_value
    participant Converter as to_number

    Caller->>Backend: basic_parse_basecoro(use_float, allow_float_overflow)
    alt allow_float_overflow and not use_float
        Backend-->>Caller: ValueError
    else valid configuration
        Backend->>Parser: parse_value(..., use_float, allow_float_overflow)
        Parser->>Parser: JSON_NUMBER_RE.fullmatch(symbol)
        Parser->>Converter: to_number(symbol)
        alt float overflow and option disabled
            Parser-->>Caller: JSONError
        else float overflow explicitly allowed
            Parser-->>Caller: Converted Python float
        else valid number
            Parser-->>Caller: JSON number event
        end
    end
Loading

Flow diagram for bounded Python lexer buffering

flowchart LR
    Chunks[Input chunks] --> Buffer[Lexer buffer]
    Buffer --> Tokens[Completed lexemes]
    Buffer --> Partial[Unfinished token]
    Tokens --> Compact{Position reaches threshold}
    Compact -->|yes| Discard[Discard consumed prefix]
    Compact -->|no| Buffer
    Discard --> Buffer
    Partial --> Buffer
Loading

Flow diagram for exact JSON number validation

flowchart TD
    Lexeme[Numeric lexeme] --> Match{JSON_NUMBER_RE.fullmatch}
    Match -->|no| Error[Reject malformed JSON number]
    Match -->|yes| Convert[to_number]
    Convert --> Overflow{Float infinity?}
    Overflow -->|yes and not allowed| Error
    Overflow -->|no or explicitly allowed| Event[Emit numeric value]
Loading

File-Level Changes

Change Details Files
Bounds lexer buffer growth while preserving token continuity and absolute error positions.
  • Compacts consumed buffer data at a 64 KiB threshold.
  • Retains unfinished strings and numbers across chunks.
  • Maintains discarded-byte accounting for error offsets.
src/ijson/backends/python.py
tests/test_python_streaming.py
Adds opt-in Python float overflow behavior without changing defaults or integer conversion.
  • Introduces allow_float_overflow and enforces use_float=True.
  • Allows overflow results to match Python/json conversion only when explicitly enabled.
  • Keeps default overflow errors and arbitrary-size integer handling.
src/ijson/backends/python.py
tests/test_python_streaming.py
Tightens Python-backend validation to exact JSON lexical rules.
  • Validates numbers with a strict JSON number pattern before conversion.
  • Restricts lexer whitespace to JSON ASCII whitespace.
  • Adds coverage for malformed numbers, keywords, whitespace, strings, and chunk boundaries.
src/ijson/backends/python.py
tests/test_python_streaming.py
Adds regression coverage for memory scaling and streaming correctness.
  • Measures RSS while parsing 16 MiB and 256 MiB generated string arrays.
  • Verifies bounded memory growth, event counts, offsets, escapes, surrogates, and conversion types.
tests/test_python_streaming.py
Documents the new backend behavior and options.
  • Updates project documentation with the Python backend configuration/details.
README.rst

Tips and commands

Interacting with Sourcery

  • Trigger a new review: Comment @sourcery-ai review on the pull request.
  • Continue discussions: Reply directly to Sourcery's review comments.
  • Generate a GitHub issue from a review comment: Ask Sourcery to create an
    issue from a review comment by replying to it. You can also reply to a
    review comment with @sourcery-ai issue to create an issue from it.
  • Generate a pull request title: Write @sourcery-ai anywhere in the pull
    request title to generate a title at any time. You can also comment
    @sourcery-ai title on the pull request to (re-)generate the title at any time.
  • Generate a pull request summary: Write @sourcery-ai summary anywhere in
    the pull request body to generate a PR summary at any time exactly where you
    want it. You can also comment @sourcery-ai summary on the pull request to
    (re-)generate the summary at any time.
  • Generate reviewer's guide: Comment @sourcery-ai guide on the pull
    request to (re-)generate the reviewer's guide at any time.
  • Resolve all Sourcery comments: Comment @sourcery-ai resolve on the
    pull request to resolve all Sourcery comments. Useful if you've already
    addressed all the comments and don't want to see them anymore.
  • Dismiss all Sourcery reviews: Comment @sourcery-ai dismiss on the pull
    request to dismiss all existing Sourcery reviews. Especially useful if you
    want to start fresh with a new review - don't forget to comment
    @sourcery-ai review to trigger a new review!

Customizing Your Experience

Access your dashboard to:

  • Enable or disable review features such as the Sourcery-generated pull request
    summary, the reviewer's guide, and others.
  • Change the review language.
  • Add, remove or edit custom review instructions.
  • Adjust other review settings.

Getting Help

@sourcery-ai sourcery-ai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hey - I've found 2 issues

Prompt for AI Agents
Please address the comments from this code review:

## Individual Comments

### Comment 1
<location path="src/ijson/backends/python.py" line_range="229" />
<code_context>
+                        raise ValueError("Invalid JSON number")
                     number = to_number(symbol)
-                    if number == inf:
+                    if number == inf and not allow_float_overflow:
                         raise common.JSONError("float overflow: %s" % (symbol,))
                 except:
</code_context>
<issue_to_address>
**issue (bug_risk):** Negative floating-point overflow is not rejected by the default path because `float('-1e400')` is `-inf`, which does not equal the positive `inf` sentinel. The Python backend therefore emits `-inf` when `use_float=True` and `allow_float_overflow` is left at its default.

**Triggers:** When valid JSON contains a sufficiently large negative exponent value such as `-1e400`.

**Suggested fix:** Use an infinity check that covers both signs, such as `math.isinf(number)`, before applying `allow_float_overflow`.
</issue_to_address>

### Comment 2
<location path="tests/test_python_streaming.py" line_range="67-69" />
<code_context>
+
+def test_compact_long_string_array_has_bounded_peak_memory(tmp_path):
+    code = '''
+import json, resource, sys
+from ijson.backends import python as backend
+with open(sys.argv[1], 'rb') as stream:
</code_context>
<issue_to_address>
**issue (testing):** The RSS regression subprocess imports the Unix-only `resource` module unconditionally, so the test raises `ModuleNotFoundError` and fails on Windows before measuring memory.

**Triggers:** When the test suite runs on Windows.

**Suggested fix:** Skip this test when `resource` is unavailable, or use a platform-independent process-memory measurement.

```suggestion
def test_compact_long_string_array_has_bounded_peak_memory(tmp_path):
    try:
        import resource
    except ImportError:
        pytest.skip('resource module is unavailable')
    code = '''
import json, resource, sys
```
</issue_to_address>

Sourcery assessment

Approval pending. 2 findings to address first.

Blocking findings: src/ijson/backends/python.py:229, tests/test_python_streaming.py:69


Sourcery is free for open source - if you like our reviews please consider sharing them ✨

raise ValueError("Invalid JSON number")
number = to_number(symbol)
if number == inf:
if number == inf and not allow_float_overflow:

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

issue (bug_risk): Negative floating-point overflow is not rejected by the default path because float('-1e400') is -inf, which does not equal the positive inf sentinel. The Python backend therefore emits -inf when use_float=True and allow_float_overflow is left at its default.

Triggers: When valid JSON contains a sufficiently large negative exponent value such as -1e400.

Suggested fix: Use an infinity check that covers both signs, such as math.isinf(number), before applying allow_float_overflow.

Comment on lines +67 to +69
def test_compact_long_string_array_has_bounded_peak_memory(tmp_path):
code = '''
import json, resource, sys

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

issue (testing): The RSS regression subprocess imports the Unix-only resource module unconditionally, so the test raises ModuleNotFoundError and fails on Windows before measuring memory.

Triggers: When the test suite runs on Windows.

Suggested fix: Skip this test when resource is unavailable, or use a platform-independent process-memory measurement.

Suggested change
def test_compact_long_string_array_has_bounded_peak_memory(tmp_path):
code = '''
import json, resource, sys
def test_compact_long_string_array_has_bounded_peak_memory(tmp_path):
try:
import resource
except ImportError:
pytest.skip('resource module is unavailable')
code = '''
import json, resource, sys

@RyaliNvidia RyaliNvidia closed this Oct 1, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant