Skip to content

feat: admit free review routes only on per-call zero-cost evidence - #2384

Draft
seonghobae wants to merge 3 commits into
cursor/experiential-labs-sidecar-d7d5from
feat/free-now-launcher
Draft

seonghobae wants to merge 3 commits into
cursor/experiential-labs-sidecar-d7d5from
feat/free-now-launcher

Conversation

@seonghobae

@seonghobae seonghobae commented Sep 26, 2026 •

Copy link
Copy Markdown
Contributor

Stacked on #2377: a36c9de5, then ac62a025 (#2377 review notes), fbd78b0e, 5d198582
(review fixes), and 9cd8aee8 (billing-exposure wording; comment and changelog text only).
This PR uses the contextual-orchestrator signal from
ContextualWisdomLab/contextual-orchestrator#1260 (feat/free-serving-cost-evidence). Under the
current pin 098ea168, which predates that signal, it degrades safely.

Why

A zero catalog price nominates a route; it doesn't prove the next call is free.
Experiential Labs bills past a per-organization free allowance once "credits overflow" is
on (it turns on automatically at the first real payment). A static credential exclusion can't
follow an account that moves between free and paid, so orchestrator/free admission now
follows the orchestrator's per-call cost evidence.

What

  • Candidates (_nominated_free_models): the orchestrator's free_discovered_models
    (zero token price) plus rows it nominated from Experiential Labs' public, keyless
    GET /api/models promotions[] catalog (free_promotion), deduplicated by route.
    • Candidates are filtered to general-chat, text-output routes before probing, so no
      probe is spent on a route that can never be selected.
    • If the orchestrator's promotions fetch fails, nothing is nominated (fail-closed; this
      happens on the orchestrator side).
  • Evidence (_free_now_models): the pinned orchestrator's
    free_serving_evidence.probe_free_candidates sends at most
    REVIEW_FREE_EVIDENCE_MAX_PROBES = 4 16-token proxy_send_once probes per run.
    • Only nominated, probe-due, evidence-required routes are probed.
    • A route is admitted only through free_serving_admitted: a JSON-number
      usage.cost == 0 and is_byok explicitly false.
    • These withhold the route from the free pool, with the reason recorded:
      • cost_paid: a positive cost.
      • cost_exhausted: a 429 free_limit_reached / insufficient_quota, or an HTTP 402.
      • cost_unknown: a missing or unparseable cost (a string "0" included), or a failed
        probe.
      • no_evidence: no probe budget left, or probing skipped.
    • A withheld route stays eligible as priced.
    • On the orchestrator side (contextual-orchestrator#1260), a demotion holds until the
      allowance reset at 00:05 UTC (09:05 KST: 00:00 UTC plus the 5-minute
      ALLOWANCE_RESET_SKEW_SECONDS) even if later calls are unknown. FREE evidence expires
      at the reset, and re-admission needs a fresh FREE probe. Idempotency-Key requests
      neither promote nor demote. That ledger (FREE_SERVING_LEDGER) is in-process memory,
      so all of this holds only within one launcher run; see the trade-off below.
  • ZDR runs never probe: with --require-zdr (private repositories), no probe is
    sent. The artifact records probe_skipped: "require_zdr", and stderr logs
    probe_skipped=require_zdr. Experiential Labs has no ZDR scope, so its routes would be
    dropped by the ZDR filter anyway. Without a recorded FREE verdict they are withheld as
    no_evidence.
  • Report and policy: rows admitted on per-call evidence carry
    free_evidence: "per_call_zero_cost" and no token prices.
    contextual_orchestrator_review_policy.parse_discovery_report classifies them free
    with non_token_price_evidence = {"source": "usage.cost", "price": 0.0, "unit": "per_call"}.
    • This applies only to experiential_labs rows with is_free: true.
    • A marker never launders a conflicting list price on another provider.
  • Pin compatibility:
    • If the pinned orchestrator has no free_serving_evidence, or its version has no
      probe_free_candidates (today's pin 098ea168), evidence-required providers and
      promotion-only rows are withheld as no_signal (signal: "unavailable").
    • If the signal has an unexpected shape (missing FREE_SERVING_LEDGER, a
      signature TypeError, an AttributeError, or a non-mapping probe report), the
      sidecar is no longer aborted. Only those same routes are withheld as no_signal,
      with signal: "incompatible" and signal_error: "<ExceptionType>".
    • Both cases match the pin's own _is_free_agent, so no dead route takes a free slot.
      FREE_POOL_CREDENTIAL_NAMES is unchanged.
  • The evidence-required provider set and the per_call_zero_cost marker are now defined
    once, in contextual_orchestrator_review_policy, and the launcher imports them. At
    runtime the launcher also unions in the orchestrator's COST_EVIDENCE_REQUIRED_PROVIDERS.
  • The discovery artifact's free_now block records the signal, the probe count, the
    probed routes, probe_skipped, and the withheld routes with reasons. stderr gets
    free_now_signal signal=… probes=… probe_skipped=… admitted=… withheld=….
  • The cost is read from the documented body field usage.cost. No per-call cost header is
    documented; the orchestrator keeps REPORTED_COST_HEADER = None as an optional slot.

Unverified assumption (read this)

The FREE rule, that a promotional free-tier call reports exactly usage.cost: 0 with
usage.is_byok: false, is inferred from the public docs
(https://platform.experientiallabs.ai/docs/cost-api,
https://platform.experientiallabs.ai/llms.txt) and unverified. No authenticated call
was made. If the real response differs, Experiential routes stay out of the free pool
(fail-closed).

Accepted trade-off (owner's decision)

A probe sent after the organization's free allowance is spent bills one 16-token call.
This happens only when Experiential Labs' org-wide credits overflow is on, and overflow
turns on automatically at the organization's first real payment. With overflow off,
the provider answers 429 and nothing is billed.

FREE_SERVING_LEDGER lives in one process's memory, and every launcher run is a new
process, so a demotion never carries over to the next run: two runs that each probe a
positive-cost route each send (and pay for) the probe. The exposure is therefore one
billed call per nominated route per sidecar run
(OpenCode, Noema and Strix sidecars),
capped at REVIEW_FREE_EVIDENCE_MAX_PROBES = 4 probes per run; today only jev-latest
is nominated. Within a run, the billed response reports cost > 0 and demotes the route
for the rest of that run. The repository owner accepted this explicitly, and the launcher
documents it next to REVIEW_FREE_EVIDENCE_MAX_PROBES. --require-zdr runs never probe.

Notes / follow-ups (not changed here)

  • Runtime preflight calls every admitted route once more after the cost probe. So an
    admitted Experiential route gets two calls per run: the probe plus the preflight.
  • Probes are bounded by count, not wall-clock time. An explicit probe timeout was
    requested, but ADR 0003 (superseding ADR 0005) forbids fixed wall-clock timeouts for
    inference, preflight and DNS/TLS setup in the review sidecars. The contract test
    test_preflight_transport_has_no_inference_timeout_and_is_provider_neutral enforces
    that. Adding a probe timeout needs an ADR amendment first. The probe client uses the
    same timeout-free settings as the preflight.

Merge note

#2382 and #2363 each change the single line REVIEW_DISPATCH_BLOB_SHA in
tests/test_pr_review_autofix_nvidia_nim_contract.py, which #2377 (the base of this
stack) also changes. Whichever merges second needs a one-line rebase. It is not
pre-resolved here.

Validation

  • tests/test_contextual_orchestrator_free_now_signal.py: 16 tests in total. 8 of them
    were already in fbd78b0e; 8 cases are new in 5d198582, and all 8 fail on
    fbd78b0e. The 8 original tests cover fail-closed behaviour without the signal or with
    a pin that lacks probe_free_candidates; promotion nomination and dedupe;
    FREE/PAID/EXHAUSTED/UNKNOWN outcomes; the probe budget and reuse of existing verdicts;
    per-call rows dropping list prices and parsing as free; and the policy honouring the
    marker only for free experiential_labs rows. One of them (the no-signal test) now also
    asserts probe_skipped: None, so it fails on fbd78b0e as well.

    The 8 new cases cover:

    • probe_skip_reason (no probe sent, and no_evidence);
    • four unexpected pin shapes (missing ledger, signature TypeError, AttributeError,
      wrong report shape), each withholding only the evidence routes as no_signal;
    • the single provider-set definition;
    • two main()-level tests with a stubbed vendored orchestrator: --require-zdr sends
      no probe, and a public run sends exactly one probe through a client with no
      wall-clock timeout. Removing the require_zdr wiring fails the first.
  • There is no launcher test that admits a string "0" cost. The fake evidence module
    records verdicts directly, and the real parser now rejects strings (tested in
    contextual-orchestrator#1260).

  • Related files (11: free-now, autofix NIM contract, sidecar contract, free credential
    admission, review policy, zdr policy, Noema workflow, required-workflow queue, Strix
    contract, Bytez catalog, runtime preflight), re-run at 9cd8aee8: 380 passed, vs 372
    on fbd78b0e (the difference is the 8 new cases).

  • Full suite with --continue-on-collection-errors, re-run at 9cd8aee8 against base
    ac62a025 in the same venv: the failure and error sets are identical, and this
    branch adds exactly the 16 passes of the free-now file (which doesn't exist on the
    base). Absolute totals depend on the environment and aren't quoted. In this venv the
    shared failures are the missing cargo (test_materialize_base_rust_dependencies.py),
    the broken venv pip, and Noema modules that need defusedxml, which isn't installed.

  • ruff check on the launcher and the new test file: clean, with no new findings (same
    command before and after). The review fixes change no shell or workflow files, so
    bash -n and shellcheck results are unchanged.

A zero catalog price nominates a route; it does not prove the next call is
free. Experiential Labs bills past a per-organization free allowance once
credits overflow is on, so the launcher now consults the pinned
orchestrator's free_serving_evidence signal before admitting a route to
orchestrator/free:

- Candidates are zero-priced rows plus rows the orchestrator nominated
  from Experiential Labs' public keyless promotions[] catalog
  (free_promotion), limited to general-chat text routes before probing.
- probe_free_candidates sends at most REVIEW_FREE_EVIDENCE_MAX_PROBES (4)
  16-token probes per run; a route is admitted only through
  free_serving_admitted (usage.cost == 0 with is_byok false). Positive
  cost, a 429 free-quota error, a missing cost, or a failed probe withhold
  it from the free pool.
- Admitted per-call rows carry free_evidence "per_call_zero_cost" and no
  token prices; the policy classifies them free with usage.cost
  non-token evidence (experiential_labs only).
- A pin without the module (or without probe_free_candidates) keeps
  evidence-required providers paid (fail-closed).

The FREE rule (cost 0 with is_byok false) is inferred from the provider's
public docs and has not been observed on a live response.
@coderabbitai

coderabbitai Bot commented Sep 26, 2026

Copy link
Copy Markdown
Contributor

Important

Draft PR not reviewed

Draft PRs are not automatically reviewed by default.

  • Trigger a manual review

To automatically review draft PRs, update your CodeRabbit configuration:

reviews:
  auto_review:
    drafts: true

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

Copy link
Copy Markdown
Contributor Author

Review (head fbd78b0e): changes needed; fixes are in progress on this branch.

  • R1: probes run before the ZDR filter (probe ~L1303, filter ~L1341), so REQUIRE_ZDR private-repo runs probe Experiential routes that get dropped anyway. Fix: skip probing when require_zdr is set. The single billed probe after the allowance runs out stays an accepted trade-off (the maintainer's decision). The docstring will say so explicitly and note that it only applies when org-wide credits overflow is on.
  • Non-blocking, also being fixed: an unexpected pin shape (missing FREE_SERVING_LEDGER, or a signature TypeError) should withhold only the evidence-required routes instead of aborting the sidecar, and probes need an explicit timeout.
  • Noted: preflight calls admitted routes a second time, and the provider set is defined in 3 places.

Verified: the old pin 098ea168 fails closed, the probe cap of 4 holds, cost 0/>0/missing map to admitted/cost_paid/cost_unknown against the real #1260 module, no keys appear in logs, and the policy marker is honoured only for free experiential_labs rows.

… shapes

- --require-zdr runs send no cost probe (Experiential Labs has no ZDR
  scope, so probed routes would be dropped anyway); the artifact records
  probe_skipped.
- Document the accepted one-billed-probe trade-off (owner's decision; only
  with org-wide credits overflow on) next to the probe budget.
- An unexpected free_serving_evidence shape (missing ledger, TypeError,
  AttributeError, ...) withholds only evidence-required routes as
  no_signal instead of aborting the sidecar.
- Probes stay count-bounded with no wall-clock timeout (ADR 0003).
- Reuse the policy's provider set and per-call marker in the launcher.
The launcher comment next to REVIEW_FREE_EVIDENCE_MAX_PROBES and the
changelog fragment claimed a probe past the free allowance is billed
"ONCE" and the route is "not probed or served free again that day".
FREE_SERVING_LEDGER is in-process memory and every launcher run is a new
process, so a demotion never carries over between runs.

State the real exposure instead: with Experiential Labs credits overflow
on (it turns on automatically at the first real payment) and the free
allowance used up, one billed call per nominated route per sidecar run
(OpenCode, Noema and Strix sidecars; at most 4 probes per run; today only
jev-latest is nominated). Correct the reset time to 00:05 UTC (00:00 UTC
plus the orchestrator's 5-minute ALLOWANCE_RESET_SKEW_SECONDS).

Comment and changelog text only; no behavior change.

Copy link
Copy Markdown
Contributor Author

Blocking governance review at exact head 9cd8aee8020e5be87b8f6cbc587358129093c8ab (stacked on ac62a025688567e64e12a302682bf103050aea62):

The implementation currently permits a probe to become a billed call after the organization's free allowance is exhausted. The PR body explicitly accepts “one billed call per nominated route per sidecar run,” up to four probes per run. That is incompatible with the current review-routing contract: GitHub Actions review work is fixed to orchestrator/free; missing/unknown free capability must fail closed, and there is no paid fallback or paid discovery/probe bypass.

This is not cured by withholding the route after a positive-cost response: cost has already been incurred at the trust boundary. The unverified assumption that promotional calls report numeric usage.cost == 0 and is_byok == false also cannot authorize a potentially paid probe.

Required owner-path correction before Ready:

  1. Keep the CO 🛡️ Sentinel: [MEDIUM] Fix missing shell=False in sandboxed_web_e2e.py #1260 signal contract as an owner prerequisite only after protected-main/release evidence.
  2. Remove any consumer-triggered probe that can incur cost, or require independently authenticated pre-call zero-cost entitlement evidence.
  3. Treat absent, expired, string-typed, unknown, or incompatible evidence as unavailable and fail closed without sending the candidate request.
  4. Add RED/GREEN contracts proving zero provider calls when zero-cost entitlement is not established before dispatch.
  5. Preserve --require-zdr no-probe behavior and the existing no-timeout model contract.

The current Draft/Proposed lifecycle is therefore correct. Hosted checks are queued, current-head independent approval is absent, and the canonical owner PR remains open. No merge, bypass, rerun, or duplicate consumer implementation is appropriate.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant