Skip to content

[HTTP interoperability] Bound stacked Content-Encoding chains without budget bypass #326

Description

@seonghobae

Buyer-visible gap

The current originweave-http first slice intentionally accepts at most one HTTP content coding (identity, gzip, or deflate). That is safe and explicit, but RFC 9110 permits an ordered list of content codings applied to a representation. A governed browser/runtime can therefore reject otherwise valid responses such as a representation encoded through two supported layers.

This is also a negotiation-consistency gap in the current first slice: request serialization advertises Accept-Encoding: gzip, deflate, while response decoding accepts either coding alone but rejects a second Content-Encoding member. This is fail-closed behavior, not an unsafe decoder bypass.

A fresh RFC 9110 review also found a narrower interoperability defect in the proposed repair contract itself: Content-Encoding uses the #content-coding list rule, and §5.6.1.2 requires recipients to parse and ignore a reasonable number of empty list elements. Therefore an implementation that categorically rejects empty comma members would be non-conforming. OriginWeave already bounds individual field values, field count, and the complete header section, so empty-element tolerance can remain linear and byte-bounded without turning an empty-member run into unbounded work.

This is an interoperability gap, not a claim that current #37 is protocol-unsafe. Keep #37/#9 bounded and mergeable; implement this only after the canonical HTTP authority is stable in its parent lineage or through an ordinary-forward successor that preserves every current test/evidence contract.

Standards / semantics authority

RFC 9110 defines Content-Encoding as #content-coding and states that codings are listed in the order in which they were applied. Decoding therefore proceeds in reverse application order. Repeated field lines with the same field name are combined in received order; the accepted coding chain must therefore preserve both field-line encounter order and comma-member order.

RFC 9110 §5.6.1.1 says senders MUST NOT generate empty list elements, but §5.6.1.2 separately says recipients MUST parse and ignore a reasonable number of empty list elements. Empty elements do not contribute to list cardinality. For this bounded client, non-empty content-coding count is therefore the depth authority; empty elements are ignored only within the already enforced header-value/header-section byte budgets and must not create decoder work. An all-empty Content-Encoding list has zero non-empty codings and is treated as no content coding, rather than as an unsupported coding token.

RFC 9110 reserves identity for its special role in Accept-Encoding and says it SHOULD NOT be included in Content-Encoding; a sole explicit identity may remain a compatibility leniency, but identity combined with another non-empty coding is not an admitted stack.

RFC 1952 permits gzip data as a concatenated series of members; that intra-gzip member rule is distinct from HTTP-level stacking. The existing raw RFC 1951 fallback for a single deflate layer is a separately evidenced compatibility path and must not become an unbounded retry mechanism inside a coding chain. RFC 9530 digest semantics remain bound to their current content/representation byte domains; adding coding layers must not silently move integrity verification to post-decoding bytes.

References:

Current RFC Editor metadata lists the RFC 1951/1952 author as P. Deutsch. IETF Datatracker synchronized that metadata on 20 May 2026; the immutable RFC text retains the historical full name “L. Peter Deutsch.” This provenance note prevents a metadata-normalization change from being mistaken for a standards or semantic change.

Ownership / DDD boundary

This remains inside OriginWeave HTTP semantics. It must not move content-decoding policy into Chromium, MCP/driver adapters, contextual-orchestrator, EgressWeave, Wardnet, or a model decision. Browser/domain policy continues to consume typed HTTP evidence; adapters do not mint success.

Model the accepted coding chain as a bounded value rather than an unbounded string list. Evidence must preserve the exact admitted wire-order coding sequence and each layer's actual decoder outcome without retaining response content.

RED → repair acceptance

  1. Start with a realistic regression proving the current exact implementation rejects a standards-valid two-layer response composed only of already-supported codings. Use a real authenticated loopback TLS exchange, not only a unit fixture. The captured request must prove the client actually advertised Accept-Encoding: gzip, deflate before the test attributes buyer impact to negotiation inconsistency.
  2. Use reviewed fixed/default max_content_coding_depth = 2 for this first slice while request serialization remains fixed to Accept-Encoding: gzip, deflate. Do not expose a caller-lowerable depth of 1 under that fixed advertisement. If a lower depth is required later, introduce one bounded accepted-coding/advertisement policy and derive both request advertisement and decoder admission from it, or narrow the request advertisement to one coding. A future increase above 2 is a separate resource/security decision.
  3. Parse the complete repeated-field/combined field-value sequence first, preserving received field-line and member order in a bounded ContentCodingChain. Ignore empty list elements as RFC 9110 §5.6.1.2 requires, within the existing bounded header bytes; count only non-empty codings toward max_content_coding_depth. Reject unknown non-empty codings, mixed identity, and non-empty depth overflow before any decoder runs. Decode admitted layers in reverse application order. Preserve a sole explicit Content-Encoding: identity only as the existing compatibility behavior; an all-empty list is zero codings/no-op.
  4. Enforce decoded-byte and expansion limits cumulatively across the complete chain. Every layer output must be checked against the original transfer-decoded coded-content byte count; a layer transition must not reset the ratio denominator or allow ratio^depth amplification. Use checked or fail-closed arithmetic and fail before observable success.
  5. Preserve gzip multi-member semantics within a single gzip layer. Do not confuse concatenated RFC 1952 members with multiple HTTP content-coding layers.
  6. Preserve the raw RFC 1951 DEFLATE fallback as a bounded compatibility decision for that exact layer only. Size/ratio failures must never trigger fallback, and nested-layer failures must not cause combinatorial retry paths.
  7. Realistic hostile/interoperability cases must cover: valid gzip, deflate and deflate, gzip chains; the equivalent repeated-field-line form in encounter order; RFC 9110 empty-list-element tolerance such as gzip, , deflate, plus an all-empty list; malformed outer and inner layers; truncated members; trailing non-coding bytes; non-empty layer-depth overflow; cumulative expansion attack; zero-length input; and a chain containing an unsupported non-empty coding. Empty elements must not create decoder invocations or count toward depth.
  8. Content-Digest/Repr-Digest semantics must remain code-current. In the current full-200 slice, integrity validation occurs on transfer-decoded, still-content-coded bytes before content decoding; stacked decoding must not silently move that byte-domain boundary or collapse distinct digest semantics into one generic post-decode result.
  9. Evidence must retain the exact admitted wire-order chain of non-empty codings plus each layer's actual decoder outcome (standard or raw-DEFLATE compatibility where applicable). Audit evidence need not preserve ignored empty list syntax; wire-order semantics must never be reconstructed from the final decoder result.
  10. Update ADR/design/doctoring/TRACEABILITY/CHANGELOG and the canonical product-gap baseline owner lane with problem, alternatives, selected limit, rejected alternatives, risk/effect, rollback and exact evidence. Keep ADR Proposed until exact-head repository/security/browser-relevant evidence exists.
  11. Owned production rustdoc, unit/integration/edge-case coverage remain 100%; exact-head CI/SAST/Security/CodeQL and eligible independent review must execute. Queued/skipped/predecessor/status-only evidence is not acceptance.

Performance / resource contract

Measure the buyer path with representative compressed payloads. The implementation should stream layer output through bounded sinks rather than materialize an unbounded intermediate per layer. If the current Vec-based decoder is retained for the first repair, every intermediate remains bounded by the same decoded-content ceiling and the coarse strict-default payload-buffer peak is approximately 80 MiB: original coded body up to 16 MiB + one prior intermediate up to 32 MiB + one new output up to 32 MiB, excluding allocator, decoder and TLS overhead. Record that explicitly; do not call the first Vec-based implementation streaming.

Empty list-element handling must remain linear in already admitted header bytes and must not allocate or invoke a decoder per empty element. The existing per-field value, field-count, and complete-header-section budgets are the outer DoS bounds; if implementation profiling shows a separate list-element cap is required, add it as a named policy with a documented RFC 9110 “reasonable number” rationale rather than silently rejecting every empty element.

Non-goals

  • adding Brotli or Zstandard in the same change;
  • increasing current maximum encoded/decoded size or expansion ratio to make tests pass;
  • changing transfer-coding support;
  • connection pooling/reuse;
  • browser rendering or file persistence;
  • LLM/model adjudication of content safety;
  • copying mutable source from another CWL owner.

Depends on the canonical HTTP authority tracked by #9/#37. Keep this issue open until stacked supported codings have exact-head realistic evidence and normal protected integration.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions