Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
50 changes: 50 additions & 0 deletions CHANGELOG.d/20260926-launcher-free-now-cost-signal.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,50 @@
### Review sidecar admits free routes only when servable free now

- `scripts/ci/contextual_orchestrator_review_launcher.py` no longer treats a
zero catalog price alone as `orchestrator/free` admission. Candidates are
the orchestrator's zero-priced rows plus rows it nominated from Experiential
Labs' public keyless `promotions[]` catalog (`free_promotion`), limited to
general-chat, text-output routes before any probe. `_free_now_models`
consults the pinned orchestrator's `free_serving_evidence` signal (the same
predicate its `general_free_serving_candidates` and
`TaskOrchestrator._is_free_agent` use): `probe_free_candidates` sends at
most `REVIEW_FREE_EVIDENCE_MAX_PROBES = 4` 16-token probes per run to
evidence-required routes (Experiential Labs), and a route is admitted only
when a response reported `usage.cost == 0` with `is_byok: false`. A positive
cost, a 429 `free_limit_reached` / `insufficient_quota`, a missing or
unparseable cost, or a failed probe keeps the route out of the free pool (it
stays eligible as priced). That a free-tier call reports exactly
`cost: 0` with `is_byok: false` is inferred from the provider docs and not
yet observed on a live response.
- Such rows reach the policy with `free_evidence: "per_call_zero_cost"` and no
token prices; `contextual_orchestrator_review_policy.parse_discovery_report`
classifies them free with `non_token_price_evidence`
`{"source": "usage.cost", "price": 0.0, "unit": "per_call"}`, only for
`experiential_labs` rows marked `is_free: true`.
- When the pinned orchestrator predates that module (or lacks
`probe_free_candidates`), those providers are treated as paid
(fail-closed), matching the pin's own serving selector so no dead route
occupies a free slot. The discovery artifact records the `free_now` signal,
probe count, probed routes, and withheld routes with reasons.
- A probe sent after the organization's free allowance is used up is
billed; this happens only when Experiential Labs' org-wide credits overflow
is on (it turns on automatically at the first real payment; otherwise the
provider answers 429 and nothing is billed). The orchestrator's
`FREE_SERVING_LEDGER` is in-process memory and every launcher run is a new
process, so the exposure is one billed call per nominated route per sidecar
run (OpenCode, Noema and Strix sidecars; at most 4 probes per run; today
only `jev-latest` is nominated). A billed response demotes the route only
for the rest of that run; the orchestrator's allowance reset is 00:05 UTC
(00:00 UTC plus a 5-minute clock-skew margin). This is an accepted
trade-off (the repository owner's decision). Runs with
`--require-zdr` send no probe at all, since Experiential Labs has no ZDR
scope and its routes would be dropped by the ZDR filter anyway; the
artifact records `probe_skipped: "require_zdr"`. Probes stay bounded by
count only; per ADR 0003 they carry no fixed wall-clock timeout.
- An unexpected shape of the pinned signal (a missing `FREE_SERVING_LEDGER`,
a changed signature, any other error from it) no longer aborts the
sidecar: only evidence-required and promotion-only routes are withheld as
`no_signal`, and the artifact records `signal: "incompatible"` with the
error type. The evidence-required provider set and the per-call marker are
defined once in `contextual_orchestrator_review_policy` and reused by the
launcher.
277 changes: 257 additions & 20 deletions scripts/ci/contextual_orchestrator_review_launcher.py
Original file line number Diff line number Diff line change
Expand Up @@ -36,6 +36,8 @@

from scripts.ci.contextual_orchestrator_review_policy import (
FREE_POOL_CREDENTIAL_NAMES,
PER_CALL_COST_EVIDENCE_PROVIDERS,
PER_CALL_FREE_EVIDENCE,
provider_account,
)

Expand Down Expand Up @@ -144,6 +146,38 @@
# ranking (higher priority first; catalog priorities are 0..-11) never places a
# deferred route ahead of a ready one.
REVIEW_PREFLIGHT_DEFERRED_PRIORITY_PENALTY = 1000
# Providers whose free status the pinned orchestrator can only prove from
# per-call cost evidence (``contextual_orchestrator.free_serving_evidence``).
# One definition shared with the policy's per-call marker check. Used on its
# own only when the pinned orchestrator predates (or exposes an unexpected
# shape of) that module: without the signal those rows are treated as paid
# (fail-closed), matching the pinned ``TaskOrchestrator._is_free_agent`` so no
# dead route occupies a free slot.
FREE_EVIDENCE_REQUIRED_PROVIDERS_FALLBACK = PER_CALL_COST_EVIDENCE_PROVIDERS
# At most this many one-shot cost-evidence probes per run. Each probe is the
# same 16-token plain-chat request as the preflight.
#
# Accepted trade-off (the repository owner's explicit decision): a probe sent
# after the organization's free allowance is used up is billed. That happens
# only when Experiential Labs' org-wide "credits overflow" is on (it turns on
# automatically at the organization's first real payment); otherwise the
# provider answers 429 ``free_limit_reached`` and nothing is billed. The
# orchestrator's ``FREE_SERVING_LEDGER`` lives in one process's memory and
# every launcher run is a new process, so nothing carries a demotion from one
# run to the next: the exposure is one billed call per nominated route per
# sidecar run (OpenCode, Noema and Strix sidecars), at most this many probes
# per run (today only ``jev-latest`` is nominated). Within a run, the billed
# response reports ``usage.cost > 0``, which demotes the route. The
# orchestrator holds a demotion until the allowance reset at 00:05 UTC
# (00:00 UTC plus its 5-minute ``ALLOWANCE_RESET_SKEW_SECONDS``), but the
# ledger ends with the process. Runs with ``--require-zdr`` never probe
# (Experiential Labs has no ZDR scope, so its routes would be dropped by the
# ZDR filter anyway).
#
# Probes are bounded by this count, not by a wall-clock timeout: ADR 0003 keeps
# inference, preflight, and DNS/TLS setup free of fixed timeouts, and the
# probe uses the same timeout-free ModelClient settings as the preflight.
REVIEW_FREE_EVIDENCE_MAX_PROBES = 4


class ReviewPreflightError(RuntimeError):
Expand Down Expand Up @@ -220,8 +254,138 @@ def _route_identity(model: object) -> tuple[str, str]:
)


def _nominated_free_models(routable: list[object], free_models: list[object]) -> list[object]:
"""Return zero-priced rows plus provider-promotion nominees, deduplicated.

``free_models`` is the pinned orchestrator's ``free_discovered_models``
(zero token price). A row with ``free_promotion`` (set by the pinned
orchestrator from Experiential Labs' public ``promotions[]`` catalog) is a
candidate too; it is admitted only by per-call cost evidence.
"""
nominated = list(free_models)
seen = {_route_identity(model) for model in nominated}
for model in routable:
identity = _route_identity(model)
if getattr(model, "free_promotion", False) is True and identity not in seen:
nominated.append(model)
seen.add(identity)
return nominated


def _evidence_required_providers(evidence: Any | None) -> frozenset[str]:
"""Return the providers whose free status needs per-call cost evidence."""
if evidence is None:
return FREE_EVIDENCE_REQUIRED_PROVIDERS_FALLBACK
return frozenset(getattr(evidence, "COST_EVIDENCE_REQUIRED_PROVIDERS", ())) | (
FREE_EVIDENCE_REQUIRED_PROVIDERS_FALLBACK
)


def _withhold_evidence_required(
free_models: list[object],
) -> tuple[list[object], list[dict[str, str]]]:
"""Fail-closed split without a usable signal: withhold evidence-required rows."""
required = _evidence_required_providers(None)
kept: list[object] = []
withheld: list[dict[str, str]] = []
for model in free_models:
provider, model_id = _route_identity(model)
if provider in required or not getattr(model, "is_free", False):
withheld.append({"provider": provider, "model": model_id, "reason": "no_signal"})
continue
kept.append(model)
return kept, withheld


def _free_now_models(
free_models: list[object],
*,
evidence: Any | None,
probe: Callable[[object], object] | None,
probe_skip_reason: str | None = None,
) -> tuple[list[object], dict[str, object]]:
"""Keep only nominated routes the orchestrator can serve free *right now*.

A zero catalog price or a provider free promotion nominates a route; it
does not prove the next call is free (Experiential Labs bills past a
per-org free allowance once credits overflow is on). ``evidence`` is the
pinned orchestrator's ``free_serving_evidence`` module -- the same signal
its discovery selector and ``TaskOrchestrator._is_free_agent`` consult --
or ``None`` when the pin predates it. ``probe`` sends one bounded
plain-chat request; the pinned ``ModelClient`` records that response's
reported ``usage.cost`` (or ``EXHAUSTED`` for a 429 free-quota error) in
the shared ledger.

* Signal unavailable (or a pin without ``probe_free_candidates``):
evidence-required providers are withheld (paid, fail-closed).
* Signal available: ``evidence.probe_free_candidates`` sends at most
``REVIEW_FREE_EVIDENCE_MAX_PROBES`` probes to nominated, probe-due
evidence-required routes. Every route is then admitted only through
``evidence.free_serving_admitted``, so a route demoted by a positive
cost or an exhausted allowance is dropped too.
* ``probe_skip_reason`` set (``"require_zdr"``): no probe is sent;
admission still reads the ledger, so evidence-required routes without
a recorded ``FREE`` verdict are withheld.
* Unexpected pin shape (a missing ``FREE_SERVING_LEDGER``, a changed
signature, or any other error from the signal): the run continues with
only evidence-required (and promotion-only) routes withheld as
``no_signal``; the sidecar is never aborted by the free-now signal.

Returns:
The admitted models and a secret-free report for the discovery artifact.
"""
compatible = evidence is not None and callable(
getattr(evidence, "probe_free_candidates", None)
)
report: dict[str, object] = {
"signal": "free_serving_evidence" if compatible else "unavailable",
"probes": 0,
"probed": [],
"probe_skipped": probe_skip_reason,
"withheld": [],
}
if not compatible:
kept, report["withheld"] = _withhold_evidence_required(free_models)
return kept, report

try:
if probe is not None and probe_skip_reason is None:
probe_report = evidence.probe_free_candidates(
free_models, probe=probe, max_probes=REVIEW_FREE_EVIDENCE_MAX_PROBES
)
report["probes"] = int(probe_report.get("probes", 0))
report["probed"] = [str(route) for route in probe_report.get("probed", [])]
ledger = evidence.FREE_SERVING_LEDGER
kept = []
withheld: list[dict[str, str]] = []
for model in free_models:
provider, model_id = _route_identity(model)
if evidence.free_serving_admitted(
provider, model_id, catalog_free=bool(getattr(model, "is_free", False))
):
kept.append(model)
continue
verdict = ledger.verdict(provider, model_id)
withheld.append(
{
"provider": provider,
"model": model_id,
"reason": f"cost_{verdict.value}" if verdict is not None else "no_evidence",
}
)
except Exception as exc: # noqa: BLE001 - an unexpected pin shape must not abort the sidecar
report["signal"] = "incompatible"
report["signal_error"] = type(exc).__name__
kept, report["withheld"] = _withhold_evidence_required(free_models)
return kept, report
report["withheld"] = withheld
return kept, report


def _report_rows(
discovered: list[object], free_route_identities: frozenset[tuple[str, str]]
discovered: list[object],
free_route_identities: frozenset[tuple[str, str]],
per_call_free_identities: frozenset[tuple[str, str]] = frozenset(),
) -> list[dict[str, object]]:
"""Convert in-process discovered models into price-evidenced report rows.

Expand All @@ -231,9 +395,17 @@ def _report_rows(
read from the discovered model when present and otherwise falls back to the
org ZDR policy table (``scripts/ci/zdr_policy.py``).

A route in ``per_call_free_identities`` was admitted by a recorded
per-call ``usage.cost == 0`` verdict rather than a zero token price (for
example an Experiential Labs promotion on a list-priced model). Its row
carries ``free_evidence: "per_call_zero_cost"`` and no token prices, so
the policy classifies it free on that evidence without reinterpreting a
list price.

Args:
discovered: Selected ``discover_all_models()`` result.
free_route_identities: Routes the orchestrator attested as zero-priced.
free_route_identities: Routes the orchestrator attested as free now.
per_call_free_identities: Free routes whose evidence is per-call cost.

Returns:
Price-evidenced rows shaped for
Expand All @@ -254,20 +426,26 @@ def _report_rows(
auth_scheme = str(
getattr(model, "auth_scheme", None) or zdr_policy.PROVIDER_AUTH_SCHEMES[provider]
)
rows.append(
{
"provider": provider,
"model": model_id,
"agent_id": str(getattr(model, "agent_id", None) or f"{provider}_{model_id}"),
"is_free": (provider, model_id) in free_route_identities,
"prompt_price_per_1k": getattr(model, "prompt_price_per_1k", None),
"completion_price_per_1k": getattr(model, "completion_price_per_1k", None),
"currency_code": getattr(model, "currency_code", None),
"base_url": base_url,
"credential_key": credential_key,
"auth_scheme": auth_scheme,
}
)
per_call_free = (provider, model_id) in per_call_free_identities
row: dict[str, object] = {
"provider": provider,
"model": model_id,
"agent_id": str(getattr(model, "agent_id", None) or f"{provider}_{model_id}"),
"is_free": (provider, model_id) in free_route_identities,
"prompt_price_per_1k": None
if per_call_free
else getattr(model, "prompt_price_per_1k", None),
"completion_price_per_1k": None
if per_call_free
else getattr(model, "completion_price_per_1k", None),
"currency_code": None if per_call_free else getattr(model, "currency_code", None),
"base_url": base_url,
"credential_key": credential_key,
"auth_scheme": auth_scheme,
}
if per_call_free:
row["free_evidence"] = PER_CALL_FREE_EVIDENCE
rows.append(row)
return rows


Expand Down Expand Up @@ -1080,7 +1258,11 @@ def main(argv: list[str] | None = None) -> int:

from contextual_orchestrator.credentials import get_credential
from contextual_orchestrator.chat_capability import is_general_chat_agent_model_id
from contextual_orchestrator.model_discovery import discover_all_models, free_discovered_models
from contextual_orchestrator.model_discovery import (
agent_from_discovered,
discover_all_models,
free_discovered_models,
)
from contextual_orchestrator.orchestrator import ModelClient, TaskOrchestrator, load_agents
from contextual_orchestrator.review_gateway import (
REVIEW_AUTH_CREDENTIAL_NAME,
Expand Down Expand Up @@ -1128,8 +1310,63 @@ def main(argv: list[str] | None = None) -> int:
raise SystemExit(f"review sidecar discovery failed: {exc}") from exc
_log_discovery_errors(discovery_errors)
routable_discovered = _routable_discovered_models(discovered)
free_models = list(free_discovered_models(routable_discovered)) if routable_discovered else []
# Only general-chat, text-output nominees can serve review traffic; filter
# before probing so no cost probe is spent on a route that is never selected.
free_models = [
model
for model in _nominated_free_models(
routable_discovered,
list(free_discovered_models(routable_discovered)) if routable_discovered else [],
)
if is_general_chat_agent_model_id(getattr(model, "model_id", ""))
and _has_text_output(model)
]
try:
from contextual_orchestrator import free_serving_evidence
except ImportError: # pinned orchestrator predates the per-call cost signal
free_serving_evidence = None
evidence_client = ModelClient(
max_output_tokens=REVIEW_PREFLIGHT_BASE_TOKENS,
max_retries=0,
temperature=REVIEW_TEMPERATURE,
)

def _probe_free_evidence(model: object) -> object:
return evidence_client.proxy_send_once(
agent_from_discovered(model),
"chat/completions",
{
"model": getattr(model, "model_id", ""),
"messages": [{"role": "user", "content": "Reply with just 'OK'."}],
"temperature": REVIEW_TEMPERATURE,
"max_tokens": REVIEW_PREFLIGHT_BASE_TOKENS,
"stream": False,
},
)

# Private-repository runs (--require-zdr) keep only ZDR-attested routes;
# evidence-required providers have no ZDR scope, so probing them would
# spend requests (possibly one billed call) on routes that are dropped.
# Note: runtime preflight later calls every admitted route once more.
free_models, free_evidence_report = _free_now_models(
free_models,
evidence=free_serving_evidence,
probe=_probe_free_evidence,
probe_skip_reason="require_zdr" if args.require_zdr else None,
)
print(
"free_now_signal "
f"signal={free_evidence_report['signal']} probes={free_evidence_report['probes']} "
f"probe_skipped={free_evidence_report['probe_skipped'] or 'no'} "
f"admitted={len(free_models)} withheld={len(free_evidence_report['withheld'])}",
file=sys.stderr,
flush=True,
)
free_route_identities = frozenset(_route_identity(model) for model in free_models)
evidence_required = _evidence_required_providers(free_serving_evidence)
per_call_free_identities = frozenset(
identity for identity in free_route_identities if identity[0] in evidence_required
)
selected_models = []
for model in routable_discovered:
model_id = getattr(model, "model_id", "")
Expand All @@ -1143,8 +1380,8 @@ def main(argv: list[str] | None = None) -> int:
f"review sidecar discovered no eligible models; orchestrator/{args.pool} would fail closed"
)

rows = _report_rows(selected_models, free_route_identities)
_write_json(args.discovery_out, {"models": rows})
rows = _report_rows(selected_models, free_route_identities, per_call_free_identities)
_write_json(args.discovery_out, {"models": rows, "free_now": free_evidence_report})
zdr_endpoints = _load_zdr_endpoints(args.zdr_endpoints)
normalized_rows = parse_discovery_report({"models": rows})
free_rows = [
Expand Down
Loading
Loading