feat: admit free review routes only on per-call zero-cost evidence - #2384
seonghobae wants to merge 3 commits into
Conversation
A zero catalog price nominates a route; it does not prove the next call is free. Experiential Labs bills past a per-organization free allowance once credits overflow is on, so the launcher now consults the pinned orchestrator's free_serving_evidence signal before admitting a route to orchestrator/free: - Candidates are zero-priced rows plus rows the orchestrator nominated from Experiential Labs' public keyless promotions[] catalog (free_promotion), limited to general-chat text routes before probing. - probe_free_candidates sends at most REVIEW_FREE_EVIDENCE_MAX_PROBES (4) 16-token probes per run; a route is admitted only through free_serving_admitted (usage.cost == 0 with is_byok false). Positive cost, a 429 free-quota error, a missing cost, or a failed probe withhold it from the free pool. - Admitted per-call rows carry free_evidence "per_call_zero_cost" and no token prices; the policy classifies them free with usage.cost non-token evidence (experiential_labs only). - A pin without the module (or without probe_free_candidates) keeps evidence-required providers paid (fail-closed). The FREE rule (cost 0 with is_byok false) is inferred from the provider's public docs and has not been observed on a live response.
|
Important Draft PR not reviewedDraft PRs are not automatically reviewed by default.
To automatically review draft PRs, update your CodeRabbit configuration: reviews:
auto_review:
drafts: trueThanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
|
Review (head
Verified: the old pin |
… shapes - --require-zdr runs send no cost probe (Experiential Labs has no ZDR scope, so probed routes would be dropped anyway); the artifact records probe_skipped. - Document the accepted one-billed-probe trade-off (owner's decision; only with org-wide credits overflow on) next to the probe budget. - An unexpected free_serving_evidence shape (missing ledger, TypeError, AttributeError, ...) withholds only evidence-required routes as no_signal instead of aborting the sidecar. - Probes stay count-bounded with no wall-clock timeout (ADR 0003). - Reuse the policy's provider set and per-call marker in the launcher.
The launcher comment next to REVIEW_FREE_EVIDENCE_MAX_PROBES and the changelog fragment claimed a probe past the free allowance is billed "ONCE" and the route is "not probed or served free again that day". FREE_SERVING_LEDGER is in-process memory and every launcher run is a new process, so a demotion never carries over between runs. State the real exposure instead: with Experiential Labs credits overflow on (it turns on automatically at the first real payment) and the free allowance used up, one billed call per nominated route per sidecar run (OpenCode, Noema and Strix sidecars; at most 4 probes per run; today only jev-latest is nominated). Correct the reset time to 00:05 UTC (00:00 UTC plus the orchestrator's 5-minute ALLOWANCE_RESET_SKEW_SECONDS). Comment and changelog text only; no behavior change.
|
Blocking governance review at exact head The implementation currently permits a probe to become a billed call after the organization's free allowance is exhausted. The PR body explicitly accepts “one billed call per nominated route per sidecar run,” up to four probes per run. That is incompatible with the current review-routing contract: GitHub Actions review work is fixed to This is not cured by withholding the route after a positive-cost response: cost has already been incurred at the trust boundary. The unverified assumption that promotional calls report numeric Required owner-path correction before Ready:
The current Draft/Proposed lifecycle is therefore correct. Hosted checks are queued, current-head independent approval is absent, and the canonical owner PR remains open. No merge, bypass, rerun, or duplicate consumer implementation is appropriate. |
Stacked on #2377:
a36c9de5, thenac62a025(#2377 review notes),fbd78b0e,5d198582(review fixes), and
9cd8aee8(billing-exposure wording; comment and changelog text only).This PR uses the contextual-orchestrator signal from
ContextualWisdomLab/contextual-orchestrator#1260 (
feat/free-serving-cost-evidence). Under thecurrent pin
098ea168, which predates that signal, it degrades safely.Why
A zero catalog price nominates a route; it doesn't prove the next call is free.
Experiential Labs bills past a per-organization free allowance once "credits overflow" is
on (it turns on automatically at the first real payment). A static credential exclusion can't
follow an account that moves between free and paid, so
orchestrator/freeadmission nowfollows the orchestrator's per-call cost evidence.
What
_nominated_free_models): the orchestrator'sfree_discovered_models(zero token price) plus rows it nominated from Experiential Labs' public, keyless
GET /api/modelspromotions[]catalog (free_promotion), deduplicated by route.probe is spent on a route that can never be selected.
happens on the orchestrator side).
_free_now_models): the pinned orchestrator'sfree_serving_evidence.probe_free_candidatessends at mostREVIEW_FREE_EVIDENCE_MAX_PROBES = 416-tokenproxy_send_onceprobes per run.free_serving_admitted: a JSON-numberusage.cost == 0andis_byokexplicitlyfalse.cost_paid: a positive cost.cost_exhausted: a 429free_limit_reached/insufficient_quota, or an HTTP 402.cost_unknown: a missing or unparseable cost (a string"0"included), or a failedprobe.
no_evidence: no probe budget left, or probing skipped.allowance reset at 00:05 UTC (09:05 KST: 00:00 UTC plus the 5-minute
ALLOWANCE_RESET_SKEW_SECONDS) even if later calls are unknown. FREE evidence expiresat the reset, and re-admission needs a fresh FREE probe. Idempotency-Key requests
neither promote nor demote. That ledger (
FREE_SERVING_LEDGER) is in-process memory,so all of this holds only within one launcher run; see the trade-off below.
--require-zdr(private repositories), no probe issent. The artifact records
probe_skipped: "require_zdr", and stderr logsprobe_skipped=require_zdr. Experiential Labs has no ZDR scope, so its routes would bedropped by the ZDR filter anyway. Without a recorded FREE verdict they are withheld as
no_evidence.free_evidence: "per_call_zero_cost"and no token prices.contextual_orchestrator_review_policy.parse_discovery_reportclassifies themfreewith
non_token_price_evidence = {"source": "usage.cost", "price": 0.0, "unit": "per_call"}.experiential_labsrows withis_free: true.free_serving_evidence, or its version has noprobe_free_candidates(today's pin098ea168), evidence-required providers andpromotion-only rows are withheld as
no_signal(signal: "unavailable").FREE_SERVING_LEDGER, asignature
TypeError, anAttributeError, or a non-mapping probe report), thesidecar is no longer aborted. Only those same routes are withheld as
no_signal,with
signal: "incompatible"andsignal_error: "<ExceptionType>"._is_free_agent, so no dead route takes a free slot.FREE_POOL_CREDENTIAL_NAMESis unchanged.per_call_zero_costmarker are now definedonce, in
contextual_orchestrator_review_policy, and the launcher imports them. Atruntime the launcher also unions in the orchestrator's
COST_EVIDENCE_REQUIRED_PROVIDERS.free_nowblock records the signal, the probe count, theprobed routes,
probe_skipped, and the withheld routes with reasons. stderr getsfree_now_signal signal=… probes=… probe_skipped=… admitted=… withheld=….usage.cost. No per-call cost header isdocumented; the orchestrator keeps
REPORTED_COST_HEADER = Noneas an optional slot.Unverified assumption (read this)
The FREE rule, that a promotional free-tier call reports exactly
usage.cost: 0withusage.is_byok: false, is inferred from the public docs(https://platform.experientiallabs.ai/docs/cost-api,
https://platform.experientiallabs.ai/llms.txt) and unverified. No authenticated call
was made. If the real response differs, Experiential routes stay out of the free pool
(fail-closed).
Accepted trade-off (owner's decision)
A probe sent after the organization's free allowance is spent bills one 16-token call.
This happens only when Experiential Labs' org-wide credits overflow is on, and overflow
turns on automatically at the organization's first real payment. With overflow off,
the provider answers 429 and nothing is billed.
FREE_SERVING_LEDGERlives in one process's memory, and every launcher run is a newprocess, so a demotion never carries over to the next run: two runs that each probe a
positive-cost route each send (and pay for) the probe. The exposure is therefore one
billed call per nominated route per sidecar run (OpenCode, Noema and Strix sidecars),
capped at
REVIEW_FREE_EVIDENCE_MAX_PROBES = 4probes per run; today onlyjev-latestis nominated. Within a run, the billed response reports
cost > 0and demotes the routefor the rest of that run. The repository owner accepted this explicitly, and the launcher
documents it next to
REVIEW_FREE_EVIDENCE_MAX_PROBES.--require-zdrruns never probe.Notes / follow-ups (not changed here)
admitted Experiential route gets two calls per run: the probe plus the preflight.
requested, but ADR 0003 (superseding ADR 0005) forbids fixed wall-clock timeouts for
inference, preflight and DNS/TLS setup in the review sidecars. The contract test
test_preflight_transport_has_no_inference_timeout_and_is_provider_neutralenforcesthat. Adding a probe timeout needs an ADR amendment first. The probe client uses the
same timeout-free settings as the preflight.
Merge note
#2382 and #2363 each change the single line
REVIEW_DISPATCH_BLOB_SHAintests/test_pr_review_autofix_nvidia_nim_contract.py, which #2377 (the base of thisstack) also changes. Whichever merges second needs a one-line rebase. It is not
pre-resolved here.
Validation
tests/test_contextual_orchestrator_free_now_signal.py: 16 tests in total. 8 of themwere already in
fbd78b0e; 8 cases are new in5d198582, and all 8 fail onfbd78b0e. The 8 original tests cover fail-closed behaviour without the signal or witha pin that lacks
probe_free_candidates; promotion nomination and dedupe;FREE/PAID/EXHAUSTED/UNKNOWN outcomes; the probe budget and reuse of existing verdicts;
per-call rows dropping list prices and parsing as free; and the policy honouring the
marker only for free
experiential_labsrows. One of them (the no-signal test) now alsoasserts
probe_skipped: None, so it fails onfbd78b0eas well.The 8 new cases cover:
probe_skip_reason(no probe sent, andno_evidence);TypeError,AttributeError,wrong report shape), each withholding only the evidence routes as
no_signal;main()-level tests with a stubbed vendored orchestrator:--require-zdrsendsno probe, and a public run sends exactly one probe through a client with no
wall-clock timeout. Removing the
require_zdrwiring fails the first.There is no launcher test that admits a string
"0"cost. The fake evidence modulerecords verdicts directly, and the real parser now rejects strings (tested in
contextual-orchestrator#1260).
Related files (11: free-now, autofix NIM contract, sidecar contract, free credential
admission, review policy, zdr policy, Noema workflow, required-workflow queue, Strix
contract, Bytez catalog, runtime preflight), re-run at
9cd8aee8: 380 passed, vs 372on
fbd78b0e(the difference is the 8 new cases).Full suite with
--continue-on-collection-errors, re-run at9cd8aee8against baseac62a025in the same venv: the failure and error sets are identical, and thisbranch adds exactly the 16 passes of the free-now file (which doesn't exist on the
base). Absolute totals depend on the environment and aren't quoted. In this venv the
shared failures are the missing
cargo(test_materialize_base_rust_dependencies.py),the broken venv
pip, and Noema modules that needdefusedxml, which isn't installed.ruff checkon the launcher and the new test file: clean, with no new findings (samecommand before and after). The review fixes change no shell or workflow files, so
bash -nandshellcheckresults are unchanged.