From f146991c9b8f088cb3d6898753a1b5ea2d821966 Mon Sep 17 00:00:00 2001 From: kailash Date: Mon, 21 Sep 2026 17:51:18 +0000 Subject: [PATCH 01/27] Propose YAML-driven model deployments --- docs/deployment-yaml-design.md | 246 +++++++++++++++++++++++++++++++++ 1 file changed, 246 insertions(+) create mode 100644 docs/deployment-yaml-design.md diff --git a/docs/deployment-yaml-design.md b/docs/deployment-yaml-design.md new file mode 100644 index 0000000..d49eba8 --- /dev/null +++ b/docs/deployment-yaml-design.md @@ -0,0 +1,246 @@ +# Proposed YAML deployments for Lilo + +Status: design only. None of the YAML fields, CLI commands, URL layouts, or new Python APIs below are implemented by this document. It does not change existing deployments. Inspected Lilo main at `67f21ee` and upstream Miles at `12754e9507e64d5e537288da17793246e913c525` on September 21, 2026. + +## Decision + +Use one declarative specification per deployment. It identifies a real base model, training mode, context limit, trainer resources and backend options, and inference resources and backend options. The same specification builds shared and scoped deployments. Lilo ships editable presets for tested model/context combinations. A user-provided specification is sufficient to add a model; no model-specific Lilo Python module, central import, or catalog edit is required. + +A running deployment necessarily has a concrete configuration for a concrete model. It does not follow that Lilo must ship a configuration for every possible model. Templates reduce duplication, and the user selects which deployments to run. + +The initial version provisions only explicitly applied specifications. A plain Tinker create request does not choose hardware, build an image, or provision an arbitrary unconfigured model. Automatic provisioning from a default hardware profile can be added later as a separate policy. + +## What Miles Tinker does + +Upstream `serve_tinker.py` parses startup arguments, checks that trainer and inference use the same frozen HF base, initializes the inference controller and trainer, and builds a gateway whose public model name is `args.tinker_base_model or args.hf_checkpoint`. Its service checks incoming model names against that single configured value. Its capabilities endpoint advertises that one name. Each training client gets an adapter on those base weights; it does not select a different architecture. + +Sources at the inspected commit: + +- [Startup and gateway model name](https://github.com/radixark/miles/blob/12754e9507e64d5e537288da17793246e913c525/serve_tinker.py#L59) +- [Training model-name check](https://github.com/radixark/miles/blob/12754e9507e64d5e537288da17793246e913c525/miles/tinker/core/service.py#L133) +- [Capabilities](https://github.com/radixark/miles/blob/12754e9507e64d5e537288da17793246e913c525/miles/tinker/server/app.py#L92) + +Lilo currently invokes Miles directly as a backend, rather than forwarding requests to Miles' standalone Tinker server. Keep this arrangement: Lilo owns sessions, futures, checkpoints and provisioning; Miles owns trainer construction and operations. These upstream observations do not imply that Lilo's pinned runtime implements every feature on Miles main. + +## One complete deployment file + +Illustrative configuration based on the existing Qwen3.5-9B-Base 16K deployment. Native option spelling below uses argparse destination names, generally underscores. Exact accepted options are validated against the selected runtime image. + +```yaml +api_version: lilo/v1 +name: qwen35-9b-lora-16k + +model: + id: Qwen/Qwen3.5-9B-Base + revision: main # Resolved to a commit before application. + parameterization: lora + max_context_length: 16384 + +deployment: + mode: shared # Or scoped, tied to lilo.run lifetime. + modal: + environment: dev + region: us-west + secrets: + api: lilo-api + sampler_proxy: lilo-proxy + huggingface: huggingface # Optional; secret reference, never a token. + storage: + assets: lilo-model-assets + checkpoints: lilo-checkpoints + bulletin: lilo-snapshot-bulletin + +trainer: + backend: miles + image: + preset: miles # Resolved to a concrete build/runtime revision. + resources: + gpu: H100:4 + cpu: 16 + memory_mib: 65536 + timeout_s: 86400 + scaling: + min_instances: 0 + max_instances: 1 + engine: + max_clients_per_instance: 6 + sampler_persistence_concurrency: 8 + miles: + model_args: qwen3.5-9B # Backend architecture preset; not a Lilo model entry. + options: + tensor_model_parallel_size: 4 + context_parallel_size: 1 + expert_model_parallel_size: 1 + expert_tensor_parallel_size: 1 + multi_lora_n_adapters: 6 + lora_rank: 32 # Allocated maximum; clients may request supported lower ranks. + lora_alpha: 32 + lora_dropout: 0.0 + target_modules: + - linear_qkv + - linear_proj + - linear_fc1 + - linear_fc2 + - output_layer + max_tokens_per_gpu: 16384 + recompute_granularity: full + recompute_method: uniform + recompute_num_layers: 1 + env: + PYTORCH_CUDA_ALLOC_CONF: expandable_segments:True + +inference: + backend: sglang + image: + preset: sglang + resources: + gpu: H200:1 + cpu: 8 + memory_mib: 32768 + scaling: + min_replicas: 0 + max_replicas: 8 + target_concurrency: 16 + scaledown_window_s: 300 + sglang: + options: + tp_size: 1 + mem_fraction_static: 0.8 + max_running_requests: 32 + max_queued_requests: 8 + max_loaded_loras: 64 + max_loras_per_batch: 8 + schedule_policy: lpm + +lifecycle: + session_idle_timeout_s: 300 + pool_idle_timeout_s: 300 + sweep_interval_s: 300 +``` + +The numbers describe a deployment choice, not a guarantee of optimal performance. In particular, zero inference minimum is an intentional departure from presets that keep workers warm. `lora_rank`, trainer slots, simultaneous adapters in an inference batch, retained adapter versions, and HTTP concurrency are different capacities. + +The deployment spec owns alpha and allocated rank. A Tinker client supplies its supported rank, seed and trainable-module selection. The resolved module selection must be compatible with the allocation and the exported adapter schema. Optimizer parameters and training batches remain per-client Tinker operations, rather than deployment-wide training-loop configuration. + +## Presets, overrides and resolution + +Proposed commands: + +```bash +lilo config init --preset qwen35-9b-lora-16k > deployment.yaml +lilo config validate deployment.yaml +lilo config resolve deployment.yaml --output deployment.resolved.yaml +lilo deploy deployment.yaml +lilo deployment check qwen35-9b-lora-16k --gpu +``` + +`config init` writes the full editable YAML. This is the simplest supported workflow. As a convenience, allow a single `extends: builtin:qwen35-9b-lora-16k` or `extends: ./base.yaml`. Local relative references resolve against the containing file. Reject cycles and duplicate YAML keys. The resolver records hashes of referenced content so later preset edits cannot silently alter an existing deployment. + +Resolution order: schema defaults, parent preset, current YAML, explicitly provided CLI overrides. Maps merge recursively; lists replace; null clears only optional fields and is otherwise rejected. Changing a field to false must actually disable it, including when enabled by a preset. Show the final resolved values and their origins; never execute shell interpolation in YAML. + +Changing `model.id` does not imply that a copied architecture preset remains valid. Require successful backend validation; incompatible explicit architecture dimensions must be rejected. Initial architecture sources are a Miles model-argument preset or explicitly supplied architecture options. Automatic HF-to-backend inference is used only where the selected backend provides a reliable implementation; do not build a second Lilo model-name table that guesses upstream recipe names. + +A normalized `ResolvedDeploymentSpec` contains the pinned model revision, rendered backend arguments, concrete image/runtime revisions, resources, storage references and module mapping. Persist it before provisioning. Redact secrets from all rendered output; specifications contain secret names only. + +## Routing from Tinker's base_model + +The primary workflow uses the URL returned for that deployment: + +```python +service = tinker.ServiceClient(base_url=deployment_url, api_key=api_key) +training = service.create_lora_training_client( + base_model="Qwen/Qwen3.5-9B-Base", rank=32, +) +``` + +The URL selects `qwen35-9b-lora-16k`; `base_model` confirms the real model. Another URL can expose the same model with 64K context. This needs no SDK change and no synthetic Hugging Face model IDs. These URLs can be paths on one gateway, for example `/deployments/qwen35-9b-lora-16k`, mounted so SDK `/api/v1/...` requests work below that prefix; they need not each be a separate CPU service. Routing remains authenticated. + +For an optional shared root URL, register deployments automatically on apply. Generate the index `(canonical_model_id, parameterization) -> deployed configurations`. One candidate routes directly. Multiple candidates require an explicit default for that model/mode; applying a second default is an error. Without a default, fail with the available deployment URLs rather than choosing the shortest context, newest deployment, or lowest GPU price. + +A shared gateway can configure defaults using deployment names. This is the only manual routing choice required for ambiguous variants. It is not a model allowlist and never requires a Python edit. The canonical model ID is read from each deployment spec, not repeated in that defaults file. + +Initially, prefer deployment URLs over names such as `Qwen/model@64k` in `base_model`. Tinker and surrounding libraries can use model names for tokenizers and metadata. If aliases are later added, resolve them at the gateway and consistently return the canonical model in model/tokenizer metadata. + +At a deployment URL, capabilities reports its model and context length. The shared root advertises only unambiguous/default routes; a separate Lilo deployment-list endpoint reports every variant, its status and URL. Sampling-only requests follow the same selection rules. Training-derived sampling clients and checkpoint restores inherit the selected deployment generation; they must not be rerouted through the current default mid-run. + +## Backend option passthrough + +Use a strict schema for Lilo-owned settings. Backend-native options are open maps, validated by their selected backend version, rather than a hand-maintained exhaustive list in Lilo. + +| Source | Destination / rule | +| --- | --- | +| `model.id`, resolved revision | Prepare one exact HF snapshot; trainer and sampler get its path. | +| `model.max_context_length` | Advertised context and supported trainer/sampler sequence limits. Packing/token budgets remain separate. | +| `trainer.resources` | Modal trainer function resources; derive actual world size from provisioned GPUs. | +| `trainer.engine` | Lilo admission, scheduling and persistence settings. | +| `trainer.miles.model_args` | Load the selected backend architecture preset in its runtime image. | +| `trainer.miles.options` | Native Miles/Megatron options, after architecture defaults. | +| `inference.resources/scaling` | Modal serving resources and autoscaling. | +| `inference.sglang.options` | Native SGLang server options. | +| Backend-exported adapter schema | Validate/derive serving target names and maximum rank. | +| `env` | Role-local environment; managed `LILO_*` variables cannot be overridden. | + +The adapter first obtains the actual parser/schema in the selected image, normalizes argument aliases, merges architecture defaults with user options, and validates/serializes the final configuration. Handle booleans, negative flags, multi-value/repeated options and comma-separated values according to that parser. Do not implement a naive `--key str(value)` loop. Unsupported encoding or unknown native options fail with the YAML path and backend error. Build argv lists; do not evaluate shell strings. Backends whose schema requires a GPU receive structural checks in preflight and full checks at initialization. + +Some settings are owned by Lilo because they determine integration behavior: model/checkpoint paths, launch world size, communication addresses, managed storage paths, trainer-only mode, and dispatch/publication hooks. Reject attempts to redefine these through passthrough, even under a CLI alias. For example, reject a Miles rollout allocation because Lilo provisions the serving pool itself. Show generated managed arguments in resolved output so this is inspectable. + +Context and rank limits are configured once. Reject conflicting native values rather than silently overriding them. Preserve hard integration constraints independently of upstream argument acceptance: an option accepted by Miles is not evidence that Lilo implements its scheduling, export, or distributed layout. Enabling PP or changing transfer mode must not bypass these checks. + +The current `_PEFT_TARGETS` mapping is not a universal model compatibility solution. Prefer backend-resolved export names and validate SGLang support. Allow explicit mapping/provider configuration when necessary, with a startup check. A custom provider must already be installed in the selected image; accepting YAML does not install arbitrary new runtime dependencies. + +## Provisioning and validation + +`lilo deploy` compiles the YAML into generic trainer and sampler definitions, deploys/registers them, and returns the URL and generation ID. Model-specific Python source is not generated or imported. Modal resource declarations are built during application construction, when GPU and image choices are known. Passing a new `gpu` value to an already-deployed function invocation cannot change its allocation. + +Use dedicated generated function/app identities per deployment generation. The current single-use trainer container behavior stays intact. A serving container is permanently associated with one resolved base/model configuration; a warm container cannot pick up another model because a registry pointer changed. A finite set of generic resource pools could be an optimization later, but is unnecessary for this design. + +With minimum capacity zero, applying a specification registers resources without warming model GPUs. The first client starts capacity; an explicit warm/check command performs initialization earlier. Nonzero configured minima intentionally reserve capacity. + +Validation has three stages: + +1. Local schema checks: types, required values, context/rank relationships, supported Lilo integration features, duplicate names/defaults, resources and topology. +2. Preflight in selected runtime images: resolve model config/revision and tokenizer assets, parse native options, resolve architecture/provider and adapter export mapping. No successful preflight should claim to prove GPU memory fit. +3. First startup (or explicit GPU check): load trainer and sampler, validate an adapter through export/load/generation, and test a small forward/backward operation on a disposable slot where appropriate. Discard probe state so no user optimizer or checkpoint is changed. Reuse the initialized resources for real work. Validate the requested allocation, but do not imply that a tiny probe proves every advertised maximum batch fits; boundary-capacity checks are a separate explicit test. + +A lightweight base-generation probe alone is insufficient: verify that the exported adapter contains expected tensors and the serving request actually selects that adapter. Cache successful verification against exact model/runtime/configuration revisions, not a moving model name. Every new container still performs normal load/readiness checks. + +Creation remains asynchronous through Tinker's existing future. Internally expose `resolving`, `provisioning`, `initializing`, `ready`, and `failed`, with timestamps and a structured failure cause. At model creation, require readiness of the trainer and the serving compatibility check for the training+sampling deployment. A valid checkpoint or active trainer must not later be destroyed solely because an individual sampling request fails. + +Startup errors must reach the control plane even if the backend dies before `accept_model` becomes available. Persist operation/attempt IDs and progress before launch, watch deployment/function completion, and complete the creation future with the original diagnostic. Retry capacity/network failures with bounds; do not repeatedly provision a permanently invalid configuration. Deduplicate simultaneous creates for the same generation. Use recoverable provisioning leases with ownership checks, not permanent claim markers. + +On failure/cancellation, release resources created solely by the failed request when no other client needs them. Preserve shared resources and existing client jobs. Cleanup is idempotent and reconciles after process death. It must include pending provisioning, not only already placed models. + +## Identities, upgrades and storage + +Keep four distinct identifiers: + +- Deployment name: user-facing, stable (`qwen35-9b-lora-16k`). +- Generation ID: hash of the normalized model/runtime/execution configuration; used for containers and compatibility verification. +- Model ID: individual Tinker client's adapter/training state. +- Publication version: a particular set of that client's adapter weights. + +Compute compatibility fingerprints from exact base revision, tokenizer/architecture settings, parameterization, export schema, parallelism and runtime revisions. Store a separate deployment-policy revision for scaling limits and timeouts so changing replica count does not invalidate model weights. Image/environment settings affecting numerical behavior belong to the execution fingerprint, not only the scaling policy. + +Assets are keyed by full repository ID and revision, not only basename. Checkpoints record canonical model, exact revision, adapter schema and topology alongside the existing metadata. Reopening a checkpoint uses recorded information and explicit compatibility checks, not the currently selected model default. Changing replica count should not make a checkpoint unloadable. Older metadata remains readable through existing compatibility rules; unknown old revisions are not silently treated as a new revision. + +Applying changed execution settings creates a new generation. New sessions use it only after deployment/preflight succeeds; old sessions remain pinned and drain normally. Startup failure on a lazily warmed generation is reported without rewriting existing sessions. Explicit warm-before-switch can provide stronger rollout guarantees. Retain old configuration records while referenced by clients, pools, futures or checkpoints. Removal disables new admission; stopping active jobs requires an explicit operation. + +The current `_lose_undefined_models` behavior must be replaced with checks against durable deployment records. A removed Python module or a restarted control plane must not invalidate a YAML deployment. The same recorded spec drives trainer reconciliation, LoRA pool cleanup, FFT latest/pinned pools and scoped teardown. + +## Shared/scoped deployment parity + +The shared CLI and a proposed `lilo.run(config="deployment.yaml")` load the same schema, resolver and builders. Lifecycle ownership differs: shared applications outlive the invoking command; scoped applications follow their owner. Preserve existing checkpoint-volume, proxy-auth and pinned-pool behavior in both modes. Secret references and deployment management belong to the operator; ordinary Tinker API keys do not grant configuration-management access. + +Multi-node, alternate trainer backends, multimodal training and different weight-transfer mechanisms are extensions behind backend adapters. They are not implicitly enabled because a YAML option exists. V1 preserves the execution layouts Lilo can actually support and rejects unsupported combinations explicitly. + +## Implementation plan and acceptance criteria + +1. Introduce schema/resolver and a normalized specification. Translate existing definitions into presets with identical values, including warm minima, adapter targets and historical GPU layouts. Snapshot-test resolved settings. +2. Introduce a deployment registry and generic builders. Keep old definition IDs as compatibility aliases while migrating references. Remove runtime Python-module imports and source-file hashing from pool resolution. +3. Add CLI validate/resolve/deploy and deployment-specific URL routing. Implement explicit shared-gateway defaults and test ambiguity. Keep the ordinary Tinker call unchanged. +4. Add native backend parser adapters, managed-option collision checks, image-based preflight, and durable startup failure reporting. Replace the deleted-module cleanup assumption before allowing dynamic entries. +5. Add startup compatibility checks, generation-aware update/drain behavior, and shared/scoped parity. Remove the old hand-maintained imports/catalog after migration coverage passes. + +Required tests: preset parity; merge/list/false override semantics; unknown options and protected aliases; backend parser errors; two contexts for one model; base versus instruct model IDs; sampling-only requests; tokenizer metadata; checkpoint restore after default changes; simultaneous cold creates; interrupted provisioning; failed trainer/sampler initialization; old-client draining; idle teardown and scoped owner loss. Run real GPU smoke tests on an existing dense LoRA preset and at least one backend-supported model absent from Lilo's old catalog, plus FFT and a supported MoE configuration before claiming those migrations complete. + +Success means a user can copy a YAML, set a new supported model and appropriate backend/hardware options, deploy it, and use an unchanged Tinker script. No edits to Lilo's Python catalog are required. Failures identify the exact unsupported setting or runtime operation instead of reporting merely that the model name is missing. From 27b1c398406ca62192d05e433f5c7c6e936e0e5e Mon Sep 17 00:00:00 2001 From: kailash Date: Mon, 21 Sep 2026 20:10:00 +0000 Subject: [PATCH 02/27] Route models through a shared Tinker frontend --- docs/deployment-yaml-design.md | 40 ++++++++++++++++++++++++---------- 1 file changed, 29 insertions(+), 11 deletions(-) diff --git a/docs/deployment-yaml-design.md b/docs/deployment-yaml-design.md index d49eba8..e2e6b7c 100644 --- a/docs/deployment-yaml-design.md +++ b/docs/deployment-yaml-design.md @@ -1,11 +1,13 @@ # Proposed YAML deployments for Lilo -Status: design only. None of the YAML fields, CLI commands, URL layouts, or new Python APIs below are implemented by this document. It does not change existing deployments. Inspected Lilo main at `67f21ee` and upstream Miles at `12754e9507e64d5e537288da17793246e913c525` on September 21, 2026. +Status: design only. None of the YAML fields, CLI commands, routing extensions, or new Python APIs below are implemented by this document. It does not change existing deployments. Inspected Lilo main at `67f21ee` and upstream Miles at `12754e9507e64d5e537288da17793246e913c525` on September 21, 2026. ## Decision Use one declarative specification per deployment. It identifies a real base model, training mode, context limit, trainer resources and backend options, and inference resources and backend options. The same specification builds shared and scoped deployments. Lilo ships editable presets for tested model/context combinations. A user-provided specification is sufficient to add a model; no model-specific Lilo Python module, central import, or catalog edit is required. +All applied deployments register behind one model-agnostic Tinker frontend. Its URL selects the service, while `base_model` selects the model. The routing table is generated from applied YAML specifications, without a separate model-to-file map. + A running deployment necessarily has a concrete configuration for a concrete model. It does not follow that Lilo must ship a configuration for every possible model. Templates reduce duplication, and the user selects which deployments to run. The initial version provisions only explicitly applied specifications. A plain Tinker create request does not choose hardware, build an image, or provision an arbitrary unconfigured model. Automatic provisioning from a default hardware profile can be added later as a separate policy. @@ -36,7 +38,11 @@ model: parameterization: lora max_context_length: 16384 +routing: + default: true # Default for this model and training mode. + deployment: + frontend: lilo-dev # Register with this shared frontend. mode: shared # Or scoped, tied to lilo.run lifetime. modal: environment: dev @@ -144,24 +150,36 @@ A normalized `ResolvedDeploymentSpec` contains the pinned model revision, render ## Routing from Tinker's base_model -The primary workflow uses the URL returned for that deployment: +The service URL identifies the shared Tinker frontend, not a model or a deployment. A single ServiceClient can create training clients for different configured models: ```python -service = tinker.ServiceClient(base_url=deployment_url, api_key=api_key) +service = tinker.ServiceClient(base_url=lilo_url, api_key=api_key) training = service.create_lora_training_client( base_model="Qwen/Qwen3.5-9B-Base", rank=32, ) +other_training = service.create_lora_training_client( + base_model="organization/model-b", rank=16, +) ``` -The URL selects `qwen35-9b-lora-16k`; `base_model` confirms the real model. Another URL can expose the same model with 64K context. This needs no SDK change and no synthetic Hugging Face model IDs. These URLs can be paths on one gateway, for example `/deployments/qwen35-9b-lora-16k`, mounted so SDK `/api/v1/...` requests work below that prefix; they need not each be a separate CPU service. Routing remains authenticated. +Applying a YAML registers its deployment name, canonical `model.id`, parameterization, context limit, routing preference and generation with the frontend selected by `deployment.frontend`. Lilo generates the index `(canonical_model_id, parameterization) -> deployed configurations`. No hand-maintained Python catalog, separate model-to-YAML map, model-specific URL, or model-specific API path is required. Deployment names are namespaced by frontend. + +Route each creation request as follows: + +1. Look up `base_model` and the requested training mode in the frontend's deployment registry. +2. With one eligible deployment, select it automatically. With multiple eligible deployments, select the one with `routing.default: true`. +3. Without an explicit default for an ambiguous model/mode, return an actionable ambiguity error listing deployment names and context limits. Unknown model IDs return an unconfigured-model error. Do not guess from rank, prompt length, context length, GPU price or registration order. +4. Bind the created model ID to that deployment generation. All later training operations, futures and training-derived sampling use that binding, regardless of subsequent default changes. + +`routing.default` defaults to false. At most one applied deployment per model/mode may declare true. Validate and publish registry changes atomically so concurrent applies cannot introduce two defaults. A default switch is an explicit registry transaction that clears the previous preference and sets the new one. Availability or a full trainer does not silently reroute clients to another configuration; capacity management operates within the selected deployment. -For an optional shared root URL, register deployments automatically on apply. Generate the index `(canonical_model_id, parameterization) -> deployed configurations`. One candidate routes directly. Multiple candidates require an explicit default for that model/mode; applying a second default is an error. Without a default, fail with the available deployment URLs rather than choosing the shortest context, newest deployment, or lowest GPU price. +This is the only routing preference users need for multiple variants. For example, both 16K and 64K YAMLs can contain `model.id: Qwen/Qwen3.5-9B-Base`; marking the 16K deployment as default makes ordinary Tinker calls select it. Deploying both does not change the frontend URL. -A shared gateway can configure defaults using deployment names. This is the only manual routing choice required for ambiguous variants. It is not a model allowlist and never requires a Python edit. The canonical model ID is read from each deployment spec, not repeated in that defaults file. +An optional Lilo client extension can later add an explicit deployment selector to creation requests while retaining the same frontend URL and real `base_model`. The server must check that the selected deployment matches the requested model and training mode. Ordinary Tinker clients require no change and use the configured default. Do not introduce `Qwen/model@64k` names: model names can also be used for tokenizer lookup and metadata. -Initially, prefer deployment URLs over names such as `Qwen/model@64k` in `base_model`. Tinker and surrounding libraries can use model names for tokenizers and metadata. If aliases are later added, resolve them at the gateway and consistently return the canonical model in model/tokenizer metadata. +The Tinker capabilities endpoint advertises canonical model names and the effective context limit of the selected default. If LoRA and FFT defaults for one model have different limits, advertise the minimum in a model-level field that cannot represent separate modes; expose the exact limits in Lilo's deployment-list endpoint. Models with ambiguous routing are not advertised as unambiguously creatable. The Lilo endpoint lists every configuration, generation, mode, context limit, default status and readiness state, without model-specific service URLs. -At a deployment URL, capabilities reports its model and context length. The shared root advertises only unambiguous/default routes; a separate Lilo deployment-list endpoint reports every variant, its status and URL. Sampling-only requests follow the same selection rules. Training-derived sampling clients and checkpoint restores inherit the selected deployment generation; they must not be rerouted through the current default mid-run. +Sampling-only creation also resolves `base_model` through the registry. With no training mode in that request, use the default configurations that expose sampling: select a unique one, or require an explicit `routing.sampling_default: true` across modes when more than one remains. This optional boolean defaults to false and is unique per canonical model in the frontend. Do not pick LoRA versus FFT arbitrarily. Training-derived sampling clients and checkpoint restores inherit their recorded configuration and compatibility requirements; they must not be rerouted through the current default mid-run. ## Backend option passthrough @@ -190,7 +208,7 @@ The current `_PEFT_TARGETS` mapping is not a universal model compatibility solut ## Provisioning and validation -`lilo deploy` compiles the YAML into generic trainer and sampler definitions, deploys/registers them, and returns the URL and generation ID. Model-specific Python source is not generated or imported. Modal resource declarations are built during application construction, when GPU and image choices are known. Passing a new `gpu` value to an already-deployed function invocation cannot change its allocation. +`lilo deploy` compiles the YAML into generic trainer and sampler definitions, deploys/registers them, and returns the shared frontend URL, deployment name and generation ID. Creating or applying another model deployment leaves the frontend URL unchanged. Model-specific Python source is not generated or imported. Modal resource declarations are built during application construction, when GPU and image choices are known. Passing a new `gpu` value to an already-deployed function invocation cannot change its allocation. Use dedicated generated function/app identities per deployment generation. The current single-use trainer container behavior stays intact. A serving container is permanently associated with one resolved base/model configuration; a warm container cannot pick up another model because a registry pointer changed. A finite set of generic resource pools could be an optimization later, but is unnecessary for this design. @@ -237,10 +255,10 @@ Multi-node, alternate trainer backends, multimodal training and different weight 1. Introduce schema/resolver and a normalized specification. Translate existing definitions into presets with identical values, including warm minima, adapter targets and historical GPU layouts. Snapshot-test resolved settings. 2. Introduce a deployment registry and generic builders. Keep old definition IDs as compatibility aliases while migrating references. Remove runtime Python-module imports and source-file hashing from pool resolution. -3. Add CLI validate/resolve/deploy and deployment-specific URL routing. Implement explicit shared-gateway defaults and test ambiguity. Keep the ordinary Tinker call unchanged. +3. Add CLI validate/resolve/deploy and shared-frontend routing generated from the deployment registry. Implement explicit defaults, atomic default switches, sampling-only selection and ambiguity errors. Keep the ordinary Tinker call and frontend URL unchanged across models. 4. Add native backend parser adapters, managed-option collision checks, image-based preflight, and durable startup failure reporting. Replace the deleted-module cleanup assumption before allowing dynamic entries. 5. Add startup compatibility checks, generation-aware update/drain behavior, and shared/scoped parity. Remove the old hand-maintained imports/catalog after migration coverage passes. -Required tests: preset parity; merge/list/false override semantics; unknown options and protected aliases; backend parser errors; two contexts for one model; base versus instruct model IDs; sampling-only requests; tokenizer metadata; checkpoint restore after default changes; simultaneous cold creates; interrupted provisioning; failed trainer/sampler initialization; old-client draining; idle teardown and scoped owner loss. Run real GPU smoke tests on an existing dense LoRA preset and at least one backend-supported model absent from Lilo's old catalog, plus FFT and a supported MoE configuration before claiming those migrations complete. +Required tests: preset parity; merge/list/false override semantics; unknown options and protected aliases; backend parser errors; one ServiceClient creating clients for multiple models; two contexts for one model; concurrent default changes; existing-client pinning after default changes; base versus instruct model IDs; sampling-only requests with LoRA/FFT ambiguity; tokenizer metadata; checkpoint restore after default changes; simultaneous cold creates; interrupted provisioning; failed trainer/sampler initialization; old-client draining; idle teardown and scoped owner loss. Run real GPU smoke tests on an existing dense LoRA preset and at least one backend-supported model absent from Lilo's old catalog, plus FFT and a supported MoE configuration before claiming those migrations complete. Success means a user can copy a YAML, set a new supported model and appropriate backend/hardware options, deploy it, and use an unchanged Tinker script. No edits to Lilo's Python catalog are required. Failures identify the exact unsupported setting or runtime operation instead of reporting merely that the model name is missing. From a0d4d295231f0adb164c2634586032752e1d020e Mon Sep 17 00:00:00 2001 From: kailash Date: Mon, 21 Sep 2026 20:57:52 +0000 Subject: [PATCH 03/27] Implement YAML-configured shared Modal trainer and rollout apps --- README.md | 2 + docs/deployment-yaml-design.md | 299 ++++++---------- examples/codeforces-codegolf/uv.lock | 4 + pyproject.toml | 8 + .../megatron_runtime/fft/checkpoint.py | 8 +- src/lilo/backends/miles_config.py | 9 +- src/lilo/backends/miles_lora.py | 7 + src/lilo/backends/miles_runtime/runtime.py | 17 +- src/lilo/control_plane/deployments.py | 85 +++++ src/lilo/control_plane/http.py | 62 ++-- src/lilo/control_plane/service.py | 8 +- src/lilo/deployment_cli.py | 232 +++++++++++++ src/lilo/deployments.py | 303 ++++++++++++++++ src/lilo/inference/native_sglang.py | 33 ++ src/lilo/native_options.py | 90 +++++ src/lilo/presets/qwen35-4b-fft-64k.yaml | 47 +++ src/lilo/presets/qwen35-9b-lora-16k.yaml | 54 +++ src/lilo/presets/qwen35-9b-lora-64k.yaml | 13 + src/lilo/providers/modal/app.py | 146 +++++--- src/lilo/providers/modal/fft_pool.py | 9 +- .../providers/modal/image_dependencies.py | 1 + src/lilo/providers/modal/kv.py | 1 + src/lilo/providers/modal/lora_pool.py | 11 +- src/lilo/providers/modal/recipe.py | 193 +++++++++++ src/lilo/providers/modal/serve.py | 9 + src/lilo/providers/modal/yaml_apps.py | 323 ++++++++++++++++++ src/lilo/providers/modal/yaml_pool_app.py | 25 ++ tests/backends/test_megatron_fft.py | 25 ++ tests/backends/test_miles.py | 19 ++ tests/providers/test_modal_serve.py | 28 ++ tests/providers/test_yaml_apps.py | 212 ++++++++++++ tests/test_deployment_cli.py | 98 ++++++ tests/test_deployments.py | 277 +++++++++++++++ uv.lock | 4 + 34 files changed, 2378 insertions(+), 284 deletions(-) create mode 100644 src/lilo/control_plane/deployments.py create mode 100644 src/lilo/deployment_cli.py create mode 100644 src/lilo/deployments.py create mode 100644 src/lilo/inference/native_sglang.py create mode 100644 src/lilo/native_options.py create mode 100644 src/lilo/presets/qwen35-4b-fft-64k.yaml create mode 100644 src/lilo/presets/qwen35-9b-lora-16k.yaml create mode 100644 src/lilo/presets/qwen35-9b-lora-64k.yaml create mode 100644 src/lilo/providers/modal/recipe.py create mode 100644 src/lilo/providers/modal/yaml_apps.py create mode 100644 src/lilo/providers/modal/yaml_pool_app.py create mode 100644 tests/providers/test_yaml_apps.py create mode 100644 tests/test_deployment_cli.py create mode 100644 tests/test_deployments.py diff --git a/README.md b/README.md index 4e958aa..0bd7ee4 100644 --- a/README.md +++ b/README.md @@ -48,6 +48,8 @@ training = service.create_lora_training_client( ## Shared deployment quick start +For the draft YAML-based provider path, see [YAML deployments](docs/deployment-yaml-design.md). + Install Lilo into your own Python project, deploy it once to Modal, then call its API from your training scripts. The commands below work in Bash or Zsh. diff --git a/docs/deployment-yaml-design.md b/docs/deployment-yaml-design.md index e2e6b7c..8a9b929 100644 --- a/docs/deployment-yaml-design.md +++ b/docs/deployment-yaml-design.md @@ -1,264 +1,177 @@ -# Proposed YAML deployments for Lilo +# YAML deployments -Status: design only. None of the YAML fields, CLI commands, routing extensions, or new Python APIs below are implemented by this document. It does not change existing deployments. Inspected Lilo main at `67f21ee` and upstream Miles at `12754e9507e64d5e537288da17793246e913c525` on September 21, 2026. +This draft implements an opt-in YAML path for shared Modal deployments. A file specifies the base model, trainer backend, resources, context length, adapter capacity, and inference settings. Adding a backend-supported model does not require a new Python definition or a catalog entry. -## Decision +The implementation has CPU tests. No applications have been redeployed and no GPU compatibility or capacity tests have been run for this change. The existing Python deployment and scoped-run paths remain available. -Use one declarative specification per deployment. It identifies a real base model, training mode, context limit, trainer resources and backend options, and inference resources and backend options. The same specification builds shared and scoped deployments. Lilo ships editable presets for tested model/context combinations. A user-provided specification is sufficient to add a model; no model-specific Lilo Python module, central import, or catalog edit is required. +## The provider and Modal app structure -All applied deployments register behind one model-agnostic Tinker frontend. Its URL selects the service, while `base_model` selects the model. The routing table is generated from applied YAML specifications, without a separate model-to-file map. +Start with [`yaml_apps.py`](../src/lilo/providers/modal/yaml_apps.py). It contains the two generic builders and the trainer entrypoint: -A running deployment necessarily has a concrete configuration for a concrete model. It does not follow that Lilo must ship a configuration for every possible model. Templates reduce duplication, and the user selects which deployments to run. +```python +trainer_app, trainer_function = build_trainer_app(resolved) +pool_app, server_class = build_rollout_app(resolved, pool) +``` -The initial version provisions only explicitly applied specifications. A plain Tinker create request does not choose hardware, build an image, or provision an arbitrary unconfigured model. Automatic provisioning from a default hardware profile can be added later as a separate policy. +Both receive the resolved configuration as data. Neither imports a model-specific definition module. + +```mermaid +flowchart TD + YAML[Complete set of deployment YAMLs] --> CLI[lilo deploy] + CLI --> Registry[Modal Dict: saved configurations and apply lock] + CLI --> Frontend[Shared Modal app: deployment.frontend] + Frontend --> HTTP[Tinker HTTP server] + Frontend --> Assets[prepare_model_assets: CPU] + Frontend --> Reconcile[trainer_reconciler and idle sweep: CPU] + Frontend --> Sampling[execute_sample: CPU] + Frontend --> TrainerA[Generated trainer function A: GPU] + Frontend --> TrainerB[Generated trainer function B: GPU] + Frontend --> Ensure[ensure_lora_pool / ensure_fft_pool: CPU] + Ensure --> Pool[Separate Modal rollout-pool app] + Pool --> Replica[Server replicas: GPU] + Replica --> Sidecar[LoRA or FFT sidecar on port 8000] + Sidecar --> SGLang[SGLang on port 8001] +``` -## What Miles Tinker does +`app.py` reads a resolved manifest from `LILO_DEPLOYMENT_MANIFEST`. For each configuration it calls `definition_from_spec`, which constructs the routing metadata and a trainer app. `app.include` places those trainer functions inside the shared frontend app. With no manifest, `app.py` uses the existing Python definitions. -Upstream `serve_tinker.py` parses startup arguments, checks that trainer and inference use the same frozen HF base, initializes the inference controller and trainer, and builds a gateway whose public model name is `args.tinker_base_model or args.hf_checkpoint`. Its service checks incoming model names against that single configured value. Its capabilities endpoint advertises that one name. Each training client gets an adapter on those base weights; it does not select a different architecture. +The trainer function's GPU type/count, CPU, RAM, timeout, maximum instances, secrets and mounted volumes come from YAML. Containers remain single-use. `run_trainer` reloads the prepared asset volume, constructs the backend configuration, and calls the existing `run_engine_with_backend` launcher. Miles uses one controller process that manages its GPU workers; FFT launches one process per allocated GPU. Client admission and sampler-persistence concurrency are configured separately. -Sources at the inspected commit: +`ensure_lora_pool` and `ensure_fft_pool` use the existing pool deployment and cleanup machinery. For YAML definitions, their deployment subprocess imports `yaml_pool_app.py` and receives the saved configuration through `LILO_POOL_DEPLOYMENT`. The resulting `Server` class captures that configuration. Startup launches SGLang with native options, waits for its health endpoint, then starts the appropriate sidecar and process supervisor. Shutdown terminates both children. -- [Startup and gateway model name](https://github.com/radixark/miles/blob/12754e9507e64d5e537288da17793246e913c525/serve_tinker.py#L59) -- [Training model-name check](https://github.com/radixark/miles/blob/12754e9507e64d5e537288da17793246e913c525/miles/tinker/core/service.py#L133) -- [Capabilities](https://github.com/radixark/miles/blob/12754e9507e64d5e537288da17793246e913c525/miles/tinker/server/app.py#L92) +A LoRA pool is shared by clients using the same deployment configuration and base weights. FFT retains its existing per-client latest pools and pinned-version/base pools. GPU resources and autoscaling settings are attached to each generated `Server` class when its app is constructed. -Lilo currently invokes Miles directly as a backend, rather than forwarding requests to Miles' standalone Tinker server. Keep this arrangement: Lilo owns sessions, futures, checkpoints and provisioning; Miles owns trainer construction and operations. These upstream observations do not imply that Lilo's pinned runtime implements every feature on Miles main. +| File | Purpose | +| --- | --- | +| [`deployments.py`](../src/lilo/deployments.py) | Schema, YAML inheritance, revision pinning and configuration identifiers | +| [`deployment_cli.py`](../src/lilo/deployment_cli.py) | Operator commands, saved manifests and serialized applies | +| [`recipe.py`](../src/lilo/providers/modal/recipe.py) | Translate configuration into Miles/Megatron and SGLang settings | +| [`yaml_apps.py`](../src/lilo/providers/modal/yaml_apps.py) | Declare trainer functions and rollout server classes; start their processes | +| [`app.py`](../src/lilo/providers/modal/app.py) | Register generated trainers with the existing shared app | +| [`yaml_pool_app.py`](../src/lilo/providers/modal/yaml_pool_app.py) | Construct a rollout app in the pool deployment subprocess | +| [`control_plane/deployments.py`](../src/lilo/control_plane/deployments.py) | Select a deployment from `base_model` and training mode | +| [`native_options.py`](../src/lilo/native_options.py) | Apply YAML values through each backend's argparse schema | -## One complete deployment file +## Configuration and commands -Illustrative configuration based on the existing Qwen3.5-9B-Base 16K deployment. Native option spelling below uses argparse destination names, generally underscores. Exact accepted options are validated against the selected runtime image. +Generate a complete editable file: -```yaml -api_version: lilo/v1 -name: qwen35-9b-lora-16k +```bash +lilo config init --preset qwen35-9b-lora-16k > deployment.yaml +lilo config validate deployment.yaml +lilo config resolve deployment.yaml --output deployment.resolved.json +``` -model: - id: Qwen/Qwen3.5-9B-Base - revision: main # Resolved to a commit before application. - parameterization: lora - max_context_length: 16384 +`validate` is offline and does not contact Modal. `resolve` resolves the HF revision and Miles runtime revision; it may contact Hugging Face and GitHub but does not allocate GPUs. Use an exact HF commit and `LILO_MILES_COMMIT` to avoid moving branch references. The resolved JSON is inspectable deployment data; `deploy` takes YAML files and resolves them again. For repeatable later applies, put those exact revisions in the YAML/environment. + +The packaged presets are: -routing: - default: true # Default for this model and training mode. +- [`qwen35-9b-lora-16k.yaml`](../src/lilo/presets/qwen35-9b-lora-16k.yaml): Qwen3.5-9B-Base, rank 32, six clients per H100:4 trainer, H200:1 inference replicas. +- [`qwen35-9b-lora-64k.yaml`](../src/lilo/presets/qwen35-9b-lora-64k.yaml): a larger-context example using H200:8 training. +- [`qwen35-4b-fft-64k.yaml`](../src/lilo/presets/qwen35-4b-fft-64k.yaml): the existing 4B FFT topology expressed as YAML. +These are starting configurations. The 16K and FFT backend settings are based on the existing definitions; inference minima are explicitly zero and the trainer maximum is one. The new 64K example has not been GPU-validated. + +You can instead keep a small override file: + +```yaml +extends: builtin:qwen35-9b-lora-16k +name: my-9b-16k +model: + id: Qwen/Qwen3.5-9B-Base + revision: main deployment: - frontend: lilo-dev # Register with this shared frontend. - mode: shared # Or scoped, tied to lilo.run lifetime. + frontend: my-lilo-yaml modal: environment: dev region: us-west secrets: api: lilo-api sampler_proxy: lilo-proxy - huggingface: huggingface # Optional; secret reference, never a token. - storage: - assets: lilo-model-assets - checkpoints: lilo-checkpoints - bulletin: lilo-snapshot-bulletin - + huggingface: huggingface-secret trainer: - backend: miles - image: - preset: miles # Resolved to a concrete build/runtime revision. resources: - gpu: H100:4 - cpu: 16 - memory_mib: 65536 - timeout_s: 86400 - scaling: - min_instances: 0 - max_instances: 1 + gpu: H200:8 engine: - max_clients_per_instance: 6 - sampler_persistence_concurrency: 8 + max_clients_per_instance: 12 miles: - model_args: qwen3.5-9B # Backend architecture preset; not a Lilo model entry. options: - tensor_model_parallel_size: 4 - context_parallel_size: 1 - expert_model_parallel_size: 1 - expert_tensor_parallel_size: 1 - multi_lora_n_adapters: 6 - lora_rank: 32 # Allocated maximum; clients may request supported lower ranks. - lora_alpha: 32 - lora_dropout: 0.0 - target_modules: - - linear_qkv - - linear_proj - - linear_fc1 - - linear_fc2 - - output_layer - max_tokens_per_gpu: 16384 - recompute_granularity: full - recompute_method: uniform - recompute_num_layers: 1 - env: - PYTORCH_CUDA_ALLOC_CONF: expandable_segments:True - + tensor_model_parallel_size: 8 + multi_lora_n_adapters: 12 inference: - backend: sglang - image: - preset: sglang - resources: - gpu: H200:1 - cpu: 8 - memory_mib: 32768 scaling: min_replicas: 0 - max_replicas: 8 - target_concurrency: 16 - scaledown_window_s: 300 - sglang: - options: - tp_size: 1 - mem_fraction_static: 0.8 - max_running_requests: 32 - max_queued_requests: 8 - max_loaded_loras: 64 - max_loras_per_batch: 8 - schedule_policy: lpm - -lifecycle: - session_idle_timeout_s: 300 - pool_idle_timeout_s: 300 - sweep_interval_s: 300 + max_replicas: 6 ``` -The numbers describe a deployment choice, not a guarantee of optimal performance. In particular, zero inference minimum is an intentional departure from presets that keep workers warm. `lora_rank`, trainer slots, simultaneous adapters in an inference batch, retained adapter versions, and HTTP concurrency are different capacities. - -The deployment spec owns alpha and allocated rank. A Tinker client supplies its supported rank, seed and trainable-module selection. The resolved module selection must be compatible with the allocation and the exported adapter schema. Optimizer parameters and training batches remain per-client Tinker operations, rather than deployment-wide training-loop configuration. - -## Presets, overrides and resolution +A local `extends: ./base.yaml` also works. Maps merge recursively, lists replace, and `false` overrides `true`. Duplicate YAML keys, unknown Lilo fields and inheritance cycles are rejected. YAML contains secret names; credentials stay in Modal secrets. The API and proxy secret contents are the same as in the [shared deployment setup](../README.md#2-configure-modal-and-secrets-once). -Proposed commands: +When ready to deploy, supply the **complete active set** of files for one frontend: ```bash -lilo config init --preset qwen35-9b-lora-16k > deployment.yaml -lilo config validate deployment.yaml -lilo config resolve deployment.yaml --output deployment.resolved.yaml -lilo deploy deployment.yaml -lilo deployment check qwen35-9b-lora-16k --gpu +lilo deploy model-a.yaml model-b.yaml model-a-64k.yaml ``` -`config init` writes the full editable YAML. This is the simplest supported workflow. As a convenience, allow a single `extends: builtin:qwen35-9b-lora-16k` or `extends: ./base.yaml`. Local relative references resolve against the containing file. Reject cycles and duplicate YAML keys. The resolver records hashes of referenced content so later preset edits cannot silently alter an existing deployment. - -Resolution order: schema defaults, parent preset, current YAML, explicitly provided CLI overrides. Maps merge recursively; lists replace; null clears only optional fields and is otherwise rejected. Changing a field to false must actually disable it, including when enabled by a preset. Show the final resolved values and their origins; never execute shell interpolation in YAML. +This command builds/deploys the shared app; it is not a validation command. All files must agree on frontend, Modal environment/region, secrets, storage and shared lifecycle settings. The preset frontend name is `lilo-yaml`. The CLI refuses to overwrite a pre-existing application without a YAML registry, so migration of a legacy frontend must be handled separately. -Changing `model.id` does not imply that a copied architecture preset remains valid. Require successful backend validation; incompatible explicit architecture dimensions must be rejected. Initial architecture sources are a Miles model-argument preset or explicitly supplied architecture options. Automatic HF-to-backend inference is used only where the selected backend provides a reliable implementation; do not build a second Lilo model-name table that guesses upstream recipe names. +Trainer minimum capacity is zero. Rollout apps are created on first demand, and their configured inference minimum applies once the pool exists. A pool with a nonzero minimum will keep that many workers warm until it is stopped by the existing idle cleanup. -A normalized `ResolvedDeploymentSpec` contains the pinned model revision, rendered backend arguments, concrete image/runtime revisions, resources, storage references and module mapping. Persist it before provisioning. Redact secrets from all rendered output; specifications contain secret names only. +## Routing and existing clients -## Routing from Tinker's base_model - -The service URL identifies the shared Tinker frontend, not a model or a deployment. A single ServiceClient can create training clients for different configured models: +The frontend URL stays the same across models: ```python service = tinker.ServiceClient(base_url=lilo_url, api_key=api_key) -training = service.create_lora_training_client( +a = service.create_lora_training_client( base_model="Qwen/Qwen3.5-9B-Base", rank=32, ) -other_training = service.create_lora_training_client( - base_model="organization/model-b", rank=16, +b = service.create_lora_training_client( + base_model="organization/another-configured-model", rank=16, ) ``` -Applying a YAML registers its deployment name, canonical `model.id`, parameterization, context limit, routing preference and generation with the frontend selected by `deployment.frontend`. Lilo generates the index `(canonical_model_id, parameterization) -> deployed configurations`. No hand-maintained Python catalog, separate model-to-YAML map, model-specific URL, or model-specific API path is required. Deployment names are namespaced by frontend. - -Route each creation request as follows: - -1. Look up `base_model` and the requested training mode in the frontend's deployment registry. -2. With one eligible deployment, select it automatically. With multiple eligible deployments, select the one with `routing.default: true`. -3. Without an explicit default for an ambiguous model/mode, return an actionable ambiguity error listing deployment names and context limits. Unknown model IDs return an unconfigured-model error. Do not guess from rank, prompt length, context length, GPU price or registration order. -4. Bind the created model ID to that deployment generation. All later training operations, futures and training-derived sampling use that binding, regardless of subsequent default changes. - -`routing.default` defaults to false. At most one applied deployment per model/mode may declare true. Validate and publish registry changes atomically so concurrent applies cannot introduce two defaults. A default switch is an explicit registry transaction that clears the previous preference and sets the new one. Availability or a full trainer does not silently reroute clients to another configuration; capacity management operates within the selected deployment. - -This is the only routing preference users need for multiple variants. For example, both 16K and 64K YAMLs can contain `model.id: Qwen/Qwen3.5-9B-Base`; marking the 16K deployment as default makes ordinary Tinker calls select it. Deploying both does not change the frontend URL. - -An optional Lilo client extension can later add an explicit deployment selector to creation requests while retaining the same frontend URL and real `base_model`. The server must check that the selected deployment matches the requested model and training mode. Ordinary Tinker clients require no change and use the configured default. Do not introduce `Qwen/model@64k` names: model names can also be used for tokenizer lookup and metadata. - -The Tinker capabilities endpoint advertises canonical model names and the effective context limit of the selected default. If LoRA and FFT defaults for one model have different limits, advertise the minimum in a model-level field that cannot represent separate modes; expose the exact limits in Lilo's deployment-list endpoint. Models with ambiguous routing are not advertised as unambiguously creatable. The Lilo endpoint lists every configuration, generation, mode, context limit, default status and readiness state, without model-specific service URLs. - -Sampling-only creation also resolves `base_model` through the registry. With no training mode in that request, use the default configurations that expose sampling: select a unique one, or require an explicit `routing.sampling_default: true` across modes when more than one remains. This optional boolean defaults to false and is unique per canonical model in the frontend. Do not pick LoRA versus FFT arbitrarily. Training-derived sampling clients and checkpoint restores inherit their recorded configuration and compatibility requirements; they must not be rerouted through the current default mid-run. - -## Backend option passthrough +The active YAMLs generate the lookup from `(model.id, parameterization)` to deployed configurations. A unique match is selected automatically. If both 16K and 64K configurations exist for the same model and mode, exactly one can set `routing.default: true`. Without a default, creation returns an ambiguity error listing the alternatives. Capacity pressure does not change which configuration is selected. -Use a strict schema for Lilo-owned settings. Backend-native options are open maps, validated by their selected backend version, rather than a hand-maintained exhaustive list in Lilo. +Sampling-only requests select the model's default configuration. If both LoRA and FFT remain eligible, set `routing.sampling_default: true` on the desired one. Training-derived sampling uses the training client's saved definition. -| Source | Destination / rule | -| --- | --- | -| `model.id`, resolved revision | Prepare one exact HF snapshot; trainer and sampler get its path. | -| `model.max_context_length` | Advertised context and supported trainer/sampler sequence limits. Packing/token budgets remain separate. | -| `trainer.resources` | Modal trainer function resources; derive actual world size from provisioned GPUs. | -| `trainer.engine` | Lilo admission, scheduling and persistence settings. | -| `trainer.miles.model_args` | Load the selected backend architecture preset in its runtime image. | -| `trainer.miles.options` | Native Miles/Megatron options, after architecture defaults. | -| `inference.resources/scaling` | Modal serving resources and autoscaling. | -| `inference.sglang.options` | Native SGLang server options. | -| Backend-exported adapter schema | Validate/derive serving target names and maximum rank. | -| `env` | Role-local environment; managed `LILO_*` variables cannot be overridden. | - -The adapter first obtains the actual parser/schema in the selected image, normalizes argument aliases, merges architecture defaults with user options, and validates/serializes the final configuration. Handle booleans, negative flags, multi-value/repeated options and comma-separated values according to that parser. Do not implement a naive `--key str(value)` loop. Unsupported encoding or unknown native options fail with the YAML path and backend error. Build argv lists; do not evaluate shell strings. Backends whose schema requires a GPU receive structural checks in preflight and full checks at initialization. - -Some settings are owned by Lilo because they determine integration behavior: model/checkpoint paths, launch world size, communication addresses, managed storage paths, trainer-only mode, and dispatch/publication hooks. Reject attempts to redefine these through passthrough, even under a CLI alias. For example, reject a Miles rollout allocation because Lilo provisions the serving pool itself. Show generated managed arguments in resolved output so this is inspectable. - -Context and rank limits are configured once. Reject conflicting native values rather than silently overriding them. Preserve hard integration constraints independently of upstream argument acceptance: an option accepted by Miles is not evidence that Lilo implements its scheduling, export, or distributed layout. Enabling PP or changing transfer mode must not bypass these checks. - -The current `_PEFT_TARGETS` mapping is not a universal model compatibility solution. Prefer backend-resolved export names and validate SGLang support. Allow explicit mapping/provider configuration when necessary, with a startup check. A custom provider must already be installed in the selected image; accepting YAML does not install arbitrary new runtime dependencies. - -## Provisioning and validation +`get_server_capabilities` advertises canonical model names and the selected context limits. Where the model-level response must cover both FFT and LoRA, it reports the smaller default context limit. Ambiguous models are omitted. The authenticated `GET /api/v1/lilo/deployments` endpoint lists active deployment names, saved definition identifiers, context limits, modes and defaults. -`lilo deploy` compiles the YAML into generic trainer and sampler definitions, deploys/registers them, and returns the shared frontend URL, deployment name and generation ID. Creating or applying another model deployment leaves the frontend URL unchanged. Model-specific Python source is not generated or imported. Modal resource declarations are built during application construction, when GPU and image choices are known. Passing a new `gpu` value to an already-deployed function invocation cannot change its allocation. +Each client records its selected definition identifier. Changing defaults affects new clients. Old trainer definitions remain registered, so existing models, sampling pools and checkpoint restores can keep referring to them. Explicit definition IDs retain the existing compatibility lookup; they are not model names advertised to ordinary clients. -Use dedicated generated function/app identities per deployment generation. The current single-use trainer container behavior stays intact. A serving container is permanently associated with one resolved base/model configuration; a warm container cannot pick up another model because a registry pointer changed. A finite set of generic resource pools could be an optimization later, but is unnecessary for this design. +## Backend options and model support -With minimum capacity zero, applying a specification registers resources without warming model GPUs. The first client starts capacity; an explicit warm/check command performs initialization earlier. Nonzero configured minima intentionally reserve capacity. +`trainer.miles.model_args` names an architecture preset inside Miles. It can be omitted when the YAML supplies explicit architecture options. Changing `model.id` does not make an inherited architecture preset compatible; Miles still validates the HF configuration at startup. -Validation has three stages: +`trainer.miles.options` and `inference.sglang.options` use argparse destination names such as `tensor_model_parallel_size` and `max_running_requests`. The Miles and SGLang entrypoints consult their real parsers before initialization. Boolean flags, scalar values and ordinary list arguments are supported. Both spellings of opposing boolean flags are replaced when they share a destination. Unknown options and custom/repeated argparse actions fail explicitly. Backend validation still runs after these overrides. -1. Local schema checks: types, required values, context/rank relationships, supported Lilo integration features, duplicate names/defaults, resources and topology. -2. Preflight in selected runtime images: resolve model config/revision and tokenizer assets, parse native options, resolve architecture/provider and adapter export mapping. No successful preflight should claim to prove GPU memory fit. -3. First startup (or explicit GPU check): load trainer and sampler, validate an adapter through export/load/generation, and test a small forward/backward operation on a disposable slot where appropriate. Discard probe state so no user optimizer or checkpoint is changed. Reuse the initialized resources for real work. Validate the requested allocation, but do not imply that a tiny probe proves every advertised maximum batch fits; boundary-capacity checks are a separate explicit test. +Lilo checks settings that affect its integration locally: GPU counts and parallelism, maximum clients versus adapter slots, managed model paths, context/rank configuration, communication endpoints and trainer-only mode. Passthrough cannot override these managed values. Megatron FFT uses its existing `EngineModelConfig` and provider overrides. -A lightweight base-generation probe alone is insufficient: verify that the exported adapter contains expected tensors and the serving request actually selects that adapter. Cache successful verification against exact model/runtime/configuration revisions, not a moving model name. Every new container still performs normal load/readiness checks. +Inference adapter targets are derived from the existing Miles-to-PEFT mapping unless `lora_target_modules` is explicitly supplied. This mapping does not establish support for every architecture. A model still needs compatible training, adapter export and SGLang loading implementations in the selected images. -Creation remains asynchronous through Tinker's existing future. Internally expose `resolving`, `provisioning`, `initializing`, `ready`, and `failed`, with timestamps and a structured failure cause. At model creation, require readiness of the trainer and the serving compatibility check for the training+sampling deployment. A valid checkpoint or active trainer must not later be destroyed solely because an individual sampling request fails. +Local validation cannot establish memory fit or prove that an unfamiliar model works. Native parser validation occurs inside the runtime images on startup. This draft does not implement a separate image-preflight command or a GPU export/load/generation probe. -Startup errors must reach the control plane even if the backend dies before `accept_model` becomes available. Persist operation/attempt IDs and progress before launch, watch deployment/function completion, and complete the creation future with the original diagnostic. Retry capacity/network failures with bounds; do not repeatedly provision a permanently invalid configuration. Deduplicate simultaneous creates for the same generation. Use recoverable provisioning leases with ownership checks, not permanent claim markers. +## Saved configurations, failures and updates -On failure/cancellation, release resources created solely by the failed request when no other client needs them. Preserve shared resources and existing client jobs. Cleanup is idempotent and reconciles after process death. It must include pending provisioning, not only already placed models. +The CLI stores configurations in the Modal Dict `-yaml-deployments`, scoped to the chosen Modal environment. A single apply lock serializes registry changes. A pending manifest is written before deployment; only successful deployment replaces the committed manifest. The next attempt retains pending configurations too, covering an interruption after Modal accepted a deployment but before the CLI saved its result. -## Identities, upgrades and storage +Applying a changed YAML produces a new definition identifier. The hash includes the pinned model revision, normalized settings and implementation fingerprint. Routing preferences are excluded, so switching a default does not change trainer identity. Old configurations are retained as inactive entries. Scaling changes currently also create a new identifier; a separate scaling-policy revision is future work. -Keep four distinct identifiers: +The implementation fingerprint includes shipped Lilo source, declared dependencies and the selected Miles commit. The existing image recipes supply the other backend source revisions. This is not a fully pinned Python/container dependency lock. This draft rejects applies that would rebuild retained configurations with a different implementation fingerprint or different shared storage/lifecycle settings; use a separate frontend for those upgrades. Automatic pruning of historical configurations and migration across code/image versions are not implemented. -- Deployment name: user-facing, stable (`qwen35-9b-lora-16k`). -- Generation ID: hash of the normalized model/runtime/execution configuration; used for containers and compatibility verification. -- Model ID: individual Tinker client's adapter/training state. -- Publication version: a particular set of that client's adapter weights. +HF asset paths hash the full repository name and exact revision. The trainer and sampler use that same directory. Miles checkpoints and FFT native-resume metadata record the base revision and reject a mismatched or unknown revision when resuming into a pinned deployment. Legacy deployments retain their existing behavior when neither side records a revision. FFT portable weights-only loading retains its existing compatibility checks. -Compute compatibility fingerprints from exact base revision, tokenizer/architecture settings, parameterization, export schema, parallelism and runtime revisions. Store a separate deployment-policy revision for scaling limits and timeouts so changing replica count does not invalidate model weights. Image/environment settings affecting numerical behavior belong to the execution fingerprint, not only the scaling policy. +A backend startup failure is recorded for that YAML definition. Waiting creation futures receive an error with the failed instance identifier, and further trainer launches are blocked for that definition. Detailed backend stderr is available in Modal call logs. Existing placed jobs are not invalidated. Once an operator fixes a transient cause, they can explicitly retry: -Assets are keyed by full repository ID and revision, not only basename. Checkpoints record canonical model, exact revision, adapter schema and topology alongside the existing metadata. Reopening a checkpoint uses recorded information and explicit compatibility checks, not the currently selected model default. Changing replica count should not make a checkpoint unloadable. Older metadata remains readable through existing compatibility rules; unknown old revisions are not silently treated as a new revision. - -Applying changed execution settings creates a new generation. New sessions use it only after deployment/preflight succeeds; old sessions remain pinned and drain normally. Startup failure on a lazily warmed generation is reported without rewriting existing sessions. Explicit warm-before-switch can provide stronger rollout guarantees. Retain old configuration records while referenced by clients, pools, futures or checkpoints. Removal disables new admission; stopping active jobs requires an explicit operation. - -The current `_lose_undefined_models` behavior must be replaced with checks against durable deployment records. A removed Python module or a restarted control plane must not invalidate a YAML deployment. The same recorded spec drives trainer reconciliation, LoRA pool cleanup, FFT latest/pinned pools and scoped teardown. - -## Shared/scoped deployment parity - -The shared CLI and a proposed `lilo.run(config="deployment.yaml")` load the same schema, resolver and builders. Lifecycle ownership differs: shared applications outlive the invoking command; scoped applications follow their owner. Preserve existing checkpoint-volume, proxy-auth and pinned-pool behavior in both modes. Secret references and deployment management belong to the operator; ordinary Tinker API keys do not grant configuration-management access. +```bash +lilo deployment retry --frontend my-lilo-yaml --env dev yaml_NAME_GENERATION +``` -Multi-node, alternate trainer backends, multimodal training and different weight-transfer mechanisms are extensions behind backend adapters. They are not implicitly enabled because a YAML option exists. V1 preserves the execution layouts Lilo can actually support and rejects unsupported combinations explicitly. +This clears the recorded failure and requests reconciliation; it may start GPU trainers if demand remains. Changing an invalid configuration produces a new definition instead. Failures before the trainer process starts, such as image-build failures, still rely on Modal's deployment diagnostics. Rollout initialization errors use the existing pool/sampling error path; trainer creation does not yet wait for an adapter-generation compatibility probe. -## Implementation plan and acceptance criteria +A killed CLI can leave its apply lock behind. Confirm that the original apply has stopped before running `lilo deployment unlock --frontend NAME --env ENV`. The lock is not automatically stolen while a slow deployment may still be running. -1. Introduce schema/resolver and a normalized specification. Translate existing definitions into presets with identical values, including warm minima, adapter targets and historical GPU layouts. Snapshot-test resolved settings. -2. Introduce a deployment registry and generic builders. Keep old definition IDs as compatibility aliases while migrating references. Remove runtime Python-module imports and source-file hashing from pool resolution. -3. Add CLI validate/resolve/deploy and shared-frontend routing generated from the deployment registry. Implement explicit defaults, atomic default switches, sampling-only selection and ambiguity errors. Keep the ordinary Tinker call and frontend URL unchanged across models. -4. Add native backend parser adapters, managed-option collision checks, image-based preflight, and durable startup failure reporting. Replace the deleted-module cleanup assumption before allowing dynamic entries. -5. Add startup compatibility checks, generation-aware update/drain behavior, and shared/scoped parity. Remove the old hand-maintained imports/catalog after migration coverage passes. +## Current scope and validation -Required tests: preset parity; merge/list/false override semantics; unknown options and protected aliases; backend parser errors; one ServiceClient creating clients for multiple models; two contexts for one model; concurrent default changes; existing-client pinning after default changes; base versus instruct model IDs; sampling-only requests with LoRA/FFT ambiguity; tokenizer metadata; checkpoint restore after default changes; simultaneous cold creates; interrupted provisioning; failed trainer/sampler initialization; old-client draining; idle teardown and scoped owner loss. Run real GPU smoke tests on an existing dense LoRA preset and at least one backend-supported model absent from Lilo's old catalog, plus FFT and a supported MoE configuration before claiming those migrations complete. +The implemented YAML path supports shared, single-node Miles LoRA and Megatron FFT deployments with the existing runtime images. The inference GPU allocation must match tensor parallelism. Scoped `lilo.run(config=...)`, DP-attention layouts, custom image selection, automatic provisioning of unknown `base_model` values, automatic runtime upgrades, and GPU compatibility probes remain follow-up work. Unsupported schema choices are rejected rather than treated as implemented features. -Success means a user can copy a YAML, set a new supported model and appropriate backend/hardware options, deploy it, and use an unchanged Tinker script. No edits to Lilo's Python catalog are required. Failures identify the exact unsupported setting or runtime operation instead of reporting merely that the model name is missing. +CPU coverage exercises configuration loading and validation, typed native overrides, routing several models through one HTTP service, ambiguity handling, preserved client definitions, interrupted/concurrent applies, trainer resources and executor settings, LoRA/FFT pool startup and shutdown, and startup-error handling. Existing backend, provider, HTTP and scoped-run tests also run. GPU smoke tests are still required before recommending this draft for production deployments. diff --git a/examples/codeforces-codegolf/uv.lock b/examples/codeforces-codegolf/uv.lock index 8cca0e9..66a9bd1 100644 --- a/examples/codeforces-codegolf/uv.lock +++ b/examples/codeforces-codegolf/uv.lock @@ -497,11 +497,13 @@ source = { editable = "../../" } dependencies = [ { name = "fastapi" }, { name = "httpx" }, + { name = "huggingface-hub" }, { name = "modal" }, { name = "opentelemetry-exporter-otlp-proto-http" }, { name = "opentelemetry-sdk" }, { name = "protobuf" }, { name = "pydantic" }, + { name = "pyyaml" }, { name = "stitch" }, { name = "tinker" }, { name = "uvicorn" }, @@ -513,11 +515,13 @@ dependencies = [ requires-dist = [ { name = "fastapi", specifier = ">=0.141.1" }, { name = "httpx", specifier = ">=0.28.1" }, + { name = "huggingface-hub", specifier = ">=0.34" }, { name = "modal", specifier = ">=1.5.3" }, { name = "opentelemetry-exporter-otlp-proto-http", specifier = ">=1.39,<2" }, { name = "opentelemetry-sdk", specifier = ">=1.39,<2" }, { name = "protobuf", specifier = ">=5.29" }, { name = "pydantic", specifier = ">=2.13.4" }, + { name = "pyyaml", specifier = ">=6.0.2" }, { name = "stitch", git = "https://github.com/modal-projects/stitch.git?rev=375a9396a7b05770dc4ed9cc5fe34fc4d5a472d5" }, { name = "tinker", specifier = ">=0.24.1,<0.25" }, { name = "uvicorn", specifier = ">=0.52.0" }, diff --git a/pyproject.toml b/pyproject.toml index d39e2a8..8ea6096 100644 --- a/pyproject.toml +++ b/pyproject.toml @@ -11,8 +11,10 @@ dependencies = [ "opentelemetry-sdk>=1.39,<2", "fastapi>=0.141.1", "httpx>=0.28.1", + "huggingface-hub>=0.34", "modal>=1.5.3", "protobuf>=5.29", + "pyyaml>=6.0.2", "pydantic>=2.13.4", "stitch @ git+https://github.com/modal-projects/stitch.git@375a9396a7b05770dc4ed9cc5fe34fc4d5a472d5", "tinker>=0.24.1,<0.25", @@ -31,3 +33,9 @@ testpaths = ["tests"] dev = [ "pytest>=9.1.1", ] + +[project.scripts] +lilo = "lilo.deployment_cli:main" + +[tool.setuptools.package-data] +lilo = ["presets/*.yaml"] diff --git a/src/lilo/backends/megatron_runtime/fft/checkpoint.py b/src/lilo/backends/megatron_runtime/fft/checkpoint.py index 4aa0f14..dd8ff09 100644 --- a/src/lilo/backends/megatron_runtime/fft/checkpoint.py +++ b/src/lilo/backends/megatron_runtime/fft/checkpoint.py @@ -56,9 +56,13 @@ class FFTCheckpointMetadata: optimizer_config: dict[str, Any] has_optimizer: bool optimizer_state_format: str | None + base_model_revision: str | None = None def to_dict(self) -> dict[str, Any]: - return asdict(self) + value = asdict(self) + if self.base_model_revision is None: + value.pop("base_model_revision") + return value def identity(self) -> str: encoded = json.dumps( @@ -82,6 +86,7 @@ def create_fft_checkpoint_metadata( version=CHECKPOINT_FORMAT_VERSION, checkpoint_id=checkpoint_id, base_model=base_model, + base_model_revision=os.environ.get("LILO_BASE_MODEL_REVISION"), world_size=world_size, tensor_model_parallel_size=config.tensor_model_parallel_size, pipeline_model_parallel_size=config.pipeline_model_parallel_size, @@ -280,6 +285,7 @@ def load_fft_training_checkpoint( "format", "version", "base_model", + "base_model_revision", "world_size", "tensor_model_parallel_size", "pipeline_model_parallel_size", diff --git a/src/lilo/backends/miles_config.py b/src/lilo/backends/miles_config.py index 00ba8e5..dfedcbb 100644 --- a/src/lilo/backends/miles_config.py +++ b/src/lilo/backends/miles_config.py @@ -1,6 +1,6 @@ from __future__ import annotations -from dataclasses import dataclass +from dataclasses import dataclass, field from pathlib import Path from typing import Any @@ -44,6 +44,7 @@ class MilesBackendConfig: ) max_tokens_per_gpu: int = 8192 extra_args: tuple[str, ...] = () + native_options: dict[str, Any] = field(default_factory=dict) @property def world_size(self) -> int: @@ -81,8 +82,10 @@ def validate(self) -> None: raise ValueError(f"{name} must be at least 1") if not self.hf_checkpoint: raise ValueError("hf_checkpoint is required") - if not self.model_type: - raise ValueError("model_type is required") + if not self.model_type and not self.native_options: + raise ValueError( + "model_type or explicit native architecture options are required" + ) if not self.target_modules: raise ValueError("target_modules must not be empty") if ( diff --git a/src/lilo/backends/miles_lora.py b/src/lilo/backends/miles_lora.py index 0e44452..4921ebf 100644 --- a/src/lilo/backends/miles_lora.py +++ b/src/lilo/backends/miles_lora.py @@ -276,6 +276,7 @@ def capture_checkpoint( "backend": "miles", "miles_revision": self.runtime.revision, "base_model": self.base_model, + "base_model_revision": os.environ.get("LILO_BASE_MODEL_REVISION"), "engine_definition_id": os.environ.get("LILO_DEFINITION_ID"), "parameterization": {"type": "lora"}, "lora_config": { @@ -590,6 +591,12 @@ def _validate_checkpoint( raise ValueError("unsupported Miles checkpoint format") if metadata.get("base_model") != self.base_model: raise ValueError("checkpoint base model does not match the deployment") + if metadata.get("base_model_revision") != os.environ.get( + "LILO_BASE_MODEL_REVISION" + ): + raise ValueError( + "checkpoint base model revision does not match the deployment" + ) lora = metadata.get("lora_config") or {} if ( int(lora.get("rank", 0)) != state.rank diff --git a/src/lilo/backends/miles_runtime/runtime.py b/src/lilo/backends/miles_runtime/runtime.py index 169979e..7a4d070 100644 --- a/src/lilo/backends/miles_runtime/runtime.py +++ b/src/lilo/backends/miles_runtime/runtime.py @@ -192,9 +192,22 @@ async def _start(self) -> None: _configure_actor_spec(train_specs) _allow_context_parallel_multi_lora() - architecture = shlex.split(load_model_args(self.config.model_type)) + architecture = ( + shlex.split(load_model_args(self.config.model_type)) + if self.config.model_type + else [] + ) with _temporary_argv([*architecture, *self.config.miles_arguments()]): - args = parse_args(entry="serve") + if self.config.native_options: + from lilo.native_options import apply_defaults + + def configure(parser): + apply_defaults(parser, self.config.native_options, sys.argv) + return parser + + args = parse_args(add_custom_arguments=configure, entry="serve") + else: + args = parse_args(entry="serve") args.use_dynamic_global_batch_size = True args.delay_split_train_data_by_dp = True configure_logger(args, source=MainProcessIdentity()) diff --git a/src/lilo/control_plane/deployments.py b/src/lilo/control_plane/deployments.py new file mode 100644 index 0000000..7c371ed --- /dev/null +++ b/src/lilo/control_plane/deployments.py @@ -0,0 +1,85 @@ +"""Model-name routing for legacy definitions and YAML deployment generations.""" + + +class DeploymentRoutes: + def __init__(self, definitions): + self.all = tuple(definitions) + self.visible = tuple(d for d in self.all if d.CATALOG_VISIBLE) + for model, mode in {(d.MODEL_NAME, d.PARAMETERIZATION) for d in self.visible}: + defaults = [ + d + for d in self.visible + if d.MODEL_NAME == model + and d.PARAMETERIZATION == mode + and getattr(d, "ROUTING_DEFAULT", False) + ] + if len(defaults) > 1: + raise ValueError(f"multiple defaults for {model} ({mode})") + + def select(self, model, mode): + explicit = [ + d + for d in self.all + if d.DEFINITION_ID == model and d.PARAMETERIZATION == mode + ] + if explicit: + return explicit[0] + matches = [ + d + for d in self.visible + if d.MODEL_NAME == model and d.PARAMETERIZATION == mode + ] + if len(matches) <= 1: + return next(iter(matches), None) + defaults = [d for d in matches if getattr(d, "ROUTING_DEFAULT", False)] + if len(defaults) == 1: + return defaults[0] + choices = ", ".join( + f"{getattr(d, 'DEPLOYMENT_NAME', d.DEFINITION_ID)} ({d.MAX_CONTEXT_LENGTH} tokens)" + for d in matches + ) + raise ValueError( + f"ambiguous {mode} deployment for {model}; configure routing.default: {choices}" + ) + + def sampling(self, model): + matches = [d for d in self.visible if d.MODEL_NAME == model] + if not any(hasattr(d, "ROUTING_DEFAULT") for d in matches): + return self.select(model, "full") or self.select(model, "lora") + explicit = [d for d in matches if getattr(d, "SAMPLING_DEFAULT", False)] + if len(explicit) == 1: + return explicit[0] + if len(explicit) > 1: + raise ValueError(f"multiple sampling defaults for {model}") + selected = [ + d + for mode in {d.PARAMETERIZATION for d in matches} + if (d := self.select(model, mode)) is not None + ] + if len(selected) > 1: + raise ValueError( + f"ambiguous sampling deployment for {model}; configure routing.sampling_default" + ) + return next(iter(selected), None) + + def capabilities(self): + result = [] + for model in dict.fromkeys(d.MODEL_NAME for d in self.visible): + try: + selected = [ + self.select(model, mode) + for mode in { + d.PARAMETERIZATION + for d in self.visible + if d.MODEL_NAME == model + } + ] + except ValueError: + continue + result.append( + { + "model_name": model, + "max_context_length": min(d.MAX_CONTEXT_LENGTH for d in selected), + } + ) + return result diff --git a/src/lilo/control_plane/http.py b/src/lilo/control_plane/http.py index 9fbba5b..8450cd5 100644 --- a/src/lilo/control_plane/http.py +++ b/src/lilo/control_plane/http.py @@ -165,28 +165,13 @@ def create_control_plane_app( definition for definition in all_definitions if definition.CATALOG_VISIBLE ) - def definition_for( - model_name: str, - parameterization: Parameterization, - ) -> str | None: - explicit = [ - d - for d in all_definitions - if d.DEFINITION_ID == model_name and d.PARAMETERIZATION == parameterization - ] - if explicit: - return explicit[0].DEFINITION_ID - matches = [ - definition.DEFINITION_ID - for definition in definitions - if definition.MODEL_NAME == model_name - and definition.PARAMETERIZATION == parameterization - ] - if len(matches) > 1: - raise ValueError( - f"duplicate {parameterization} definition for {model_name}" - ) - return matches[0] if matches else None + from .deployments import DeploymentRoutes + + routes = DeploymentRoutes(all_definitions) + + def definition_for(model_name, parameterization): + selected = routes.select(model_name, parameterization) + return selected.DEFINITION_ID if selected else None def supports_model(model_name: str) -> bool: return any( @@ -283,20 +268,23 @@ async def healthz() -> dict[str, str]: @app.get("/api/v1/get_server_capabilities") async def get_server_capabilities() -> dict[str, object]: + return {"supported_models": routes.capabilities()} + + @app.get("/api/v1/lilo/deployments") + async def list_deployments(): return { - "supported_models": [ + "deployments": [ { - "model_name": name, - "max_context_length": min( - definition.MAX_CONTEXT_LENGTH - for definition in definitions - if definition.MODEL_NAME == name - ), + "name": getattr(d, "DEPLOYMENT_NAME", d.DEFINITION_ID), + "generation": d.DEFINITION_ID, + "base_model": d.MODEL_NAME, + "parameterization": d.PARAMETERIZATION, + "max_context_length": d.MAX_CONTEXT_LENGTH, + "default": getattr(d, "ROUTING_DEFAULT", False), + "sampling_default": getattr(d, "SAMPLING_DEFAULT", False), } - for name in dict.fromkeys( - definition.MODEL_NAME for definition in definitions - ) - ], + for d in definitions + ] } @app.post("/api/v1/client/config") @@ -395,12 +383,8 @@ async def create_sampling_session( status_code=400, detail="base_model or model_path is required", ) - definition_id = ( - definition_for(body.base_model, "full") - or definition_for(body.base_model, "lora") - if body.base_model - else None - ) + selected = routes.sampling(body.base_model) if body.base_model else None + definition_id = selected.DEFINITION_ID if selected else None if body.model_path is None and definition_id is None: raise HTTPException( status_code=400, diff --git a/src/lilo/control_plane/service.py b/src/lilo/control_plane/service.py index d3b895f..47f699e 100644 --- a/src/lilo/control_plane/service.py +++ b/src/lilo/control_plane/service.py @@ -137,6 +137,7 @@ def __init__( list_checkpoints: CheckpointListing | None = None, delete_checkpoint: Callable[[str], Awaitable[None]] | None = None, checkpoint_root: str = "/checkpoints", + creation_error: Callable[[str], Awaitable[str | None]] | None = None, ) -> None: self.kv = kv self.engines = engines @@ -153,6 +154,7 @@ def __init__( self.list_checkpoints = list_checkpoints self.delete_checkpoint = delete_checkpoint self.checkpoint_root = checkpoint_root + self.creation_error = creation_error async def create_session( self, @@ -1140,6 +1142,10 @@ async def _retrieve_creation( ) -> FutureResolution: try: placement = await self._place(model) + if placement is None and self.creation_error is not None: + error = await self.creation_error(model.engine_definition_id) + if error: + raise ValueError(error) except ModelLost: return FutureResolution( request_id, @@ -1162,7 +1168,7 @@ async def _retrieve_creation( if placement is None: return FutureResolution(request_id, FutureResolutionStatus.PENDING) try: - instance = await self._live_instance(placement) + await self._live_instance(placement) except ModelLost: return FutureResolution( request_id, diff --git a/src/lilo/deployment_cli.py b/src/lilo/deployment_cli.py new file mode 100644 index 0000000..50aa23f --- /dev/null +++ b/src/lilo/deployment_cli.py @@ -0,0 +1,232 @@ +"""Operator CLI for YAML deployments. Only `deploy` provisions Modal resources.""" + +from __future__ import annotations + +import argparse +import hashlib +import json +import os +from pathlib import Path +import re +import subprocess +import sys +import uuid + +import yaml + +from lilo.deployments import ( + ResolvedDeployment, + load, + preset_path, + resolve, + validate_frontend, +) + + +def implementation_fingerprint(miles_commit: str | None) -> str: + """Hash shipped source and dependency declarations, independent of Git checkout.""" + from importlib.metadata import requires + + root = Path(__file__).parent + digest = hashlib.sha256() + for path in sorted(root.rglob("*.py")): + digest.update(str(path.relative_to(root)).encode()) + digest.update(path.read_bytes()) + digest.update(json.dumps([miles_commit, sorted(requires("lilo") or [])]).encode()) + return digest.hexdigest() + + +def compile_configs(paths): + specs = [load(path) for path in paths] + validate_frontend(specs) + from lilo.providers.modal.miles_revision import resolve_miles_commit + + miles_commit = resolve_miles_commit() + implementation = implementation_fingerprint(miles_commit) + resolved = [] + for spec in specs: + revision = spec.model.revision + if not re.fullmatch(r"[0-9a-f]{40,64}", revision): + from huggingface_hub import HfApi + + revision = HfApi().model_info(spec.model.id, revision=revision).sha + resolved.append( + resolve( + spec, + revision=revision, + implementation=implementation, + miles_commit=miles_commit, + ) + ) + return resolved + + +def retain_generations(previous, desired): + """The command supplies the complete active set; old generations remain usable.""" + validate_frontend([row.spec for row in desired]) + expected = desired[0] + for row in previous: + if row.implementation != expected.implementation: + raise ValueError( + "This draft cannot rebuild retained generations with different Lilo/runtime code. Use a separate frontend for a code upgrade; YAML-only changes can retain existing generations." + ) + if ( + row.spec.deployment != expected.spec.deployment + or row.spec.lifecycle != expected.spec.lifecycle + ): + raise ValueError( + "Cannot change shared storage, secrets, region, or lifecycle while retaining generations; use a separate frontend." + ) + ids = {row.definition_id for row in desired} + return [ + *desired, + *( + row.model_copy(update={"active": False}) + for row in previous + if row.definition_id not in ids + ), + ] + + +def deploy(desired): + """Serialize operator applies and retain interrupted attempts for safe recovery.""" + import modal + from lilo.providers.modal.yaml_apps import MANIFEST_ENV + + settings = desired[0].spec.deployment + registry = modal.Dict.from_name( + f"{settings.frontend}-yaml-deployments", + create_if_missing=True, + environment_name=settings.modal.environment, + ) + owner = uuid.uuid4().hex + if not registry.put("apply_lock", owner, skip_if_exists=True): + raise ValueError( + f"An apply owns {settings.frontend}. If it was interrupted, confirm it has stopped before running lilo deployment unlock --frontend {settings.frontend}." + ) + try: + rows = registry.get("manifest", []) + if not rows and not registry.get("pending", []): + try: + modal.App.lookup( + settings.frontend, environment_name=settings.modal.environment + ) + except modal.exception.NotFoundError: + pass + else: + raise ValueError( + "The frontend already exists without a YAML registry. Choose a new frontend name; automatic migration of legacy deployments is not implemented." + ) + # A killed deploy may already have updated Modal. Keep its functions on retry. + rows = {r["generation"]: r for r in [*rows, *registry.get("pending", [])]} + manifest = retain_generations( + [ResolvedDeployment.model_validate(row) for row in rows.values()], desired + ) + data = [row.model_dump(mode="json") for row in manifest] + env = { + **os.environ, + MANIFEST_ENV: json.dumps(data), + "LILO_APP_NAME": settings.frontend, + } + if desired[0].miles_commit: + env["LILO_MILES_COMMIT"] = desired[0].miles_commit + command = [ + sys.executable, + "-m", + "modal", + "deploy", + "-m", + "lilo.providers.modal.app", + ] + if settings.modal.environment: + command += ["--env", settings.modal.environment] + registry.put("pending", data) + subprocess.run(command, check=True, env=env) + registry.put("manifest", data) + registry.pop("pending", None) + finally: + if registry.get("apply_lock") == owner: + registry.pop("apply_lock", None) + + +def parser(): + result = argparse.ArgumentParser(prog="lilo") + commands = result.add_subparsers(dest="command", required=True) + config = commands.add_parser("config").add_subparsers(dest="action", required=True) + init = config.add_parser("init") + init.add_argument("--preset", required=True) + for name in ("validate", "resolve"): + cmd = config.add_parser(name) + cmd.add_argument("files", nargs="+") + if name == "resolve": + cmd.add_argument("--output") + apply = commands.add_parser( + "deploy", help="Deploy the complete active YAML set behind one frontend" + ) + apply.add_argument("files", nargs="+") + management = commands.add_parser("deployment").add_subparsers( + dest="action", required=True + ) + for name in ("unlock", "retry"): + cmd = management.add_parser(name) + cmd.add_argument("--frontend", required=True) + cmd.add_argument("--env") + if name == "retry": + cmd.add_argument("definition_id") + return result + + +def main(argv=None): + cli = parser() + args = cli.parse_args(argv) + try: + if args.command == "config": + if args.action == "init": + print( + yaml.safe_dump( + load(preset_path(args.preset)).model_dump(mode="json"), + sort_keys=False, + ), + end="", + ) + elif args.action == "validate": + specs = [load(path) for path in args.files] + validate_frontend(specs) + print( + f"Validated {len(specs)} deployment(s). Native backend options are checked at startup in their runtime images." + ) + else: + output = ( + json.dumps( + [ + row.model_dump(mode="json") + for row in compile_configs(args.files) + ], + indent=2, + ) + + "\n" + ) + if args.output: + Path(args.output).write_text(output) + else: + print(output, end="") + elif args.command == "deploy": + deploy(compile_configs(args.files)) + else: + import modal + + if args.action == "unlock": + registry = modal.Dict.from_name( + f"{args.frontend}-yaml-deployments", environment_name=args.env + ) + registry.pop("apply_lock", None) + else: + modal.Function.from_name( + args.frontend, "clear_deployment_failure", environment_name=args.env + ).remote(args.definition_id) + except (ValueError, OSError, subprocess.CalledProcessError) as exc: + cli.exit(1, f"{exc}\n") + + +if __name__ == "__main__": + main() diff --git a/src/lilo/deployments.py b/src/lilo/deployments.py new file mode 100644 index 0000000..fd9440d --- /dev/null +++ b/src/lilo/deployments.py @@ -0,0 +1,303 @@ +"""Deployment specifications. Loading YAML never contacts Modal or allocates GPUs.""" + +from __future__ import annotations + +import hashlib +import json +import re +from importlib.resources import files +from pathlib import Path +from typing import Any, Literal + +import yaml +from pydantic import BaseModel, ConfigDict, Field, model_validator + + +class StrictModel(BaseModel): + model_config = ConfigDict(extra="forbid") + + +class Model(StrictModel): + id: str = Field(min_length=1) + revision: str = "main" + parameterization: Literal["lora", "full"] = "lora" + max_context_length: int = Field(gt=0) + + +class Routing(StrictModel): + default: bool = False + sampling_default: bool = False + + +class Resources(StrictModel): + gpu: str + cpu: float = Field(default=8, gt=0) + memory_mib: int = Field(default=32768, gt=0) + timeout_s: int = Field(default=86400, gt=0, le=86400) + + @property + def gpu_count(self) -> int: + if not re.fullmatch(r"[A-Za-z0-9-]+(?::[1-9][0-9]*)?", self.gpu): + raise ValueError(f"invalid GPU resource: {self.gpu}") + return int(self.gpu.split(":")[1]) if ":" in self.gpu else 1 + + +class TrainerScaling(StrictModel): + min_instances: Literal[0] = 0 + max_instances: int = Field(default=1, gt=0) + + +class InferenceScaling(StrictModel): + min_replicas: int = Field(default=0, ge=0) + max_replicas: int = Field(default=8, gt=0) + target_concurrency: int = Field(default=16, gt=0) + scaledown_window_s: int = Field(default=300, gt=0) + + @model_validator(mode="after") + def ordered(self): + if self.min_replicas > self.max_replicas: + raise ValueError("min_replicas must not exceed max_replicas") + return self + + +class EngineOptions(StrictModel): + max_clients_per_instance: int = Field(default=1, gt=0) + sampler_persistence_concurrency: int = Field(default=8, gt=0) + + +class MilesOptions(StrictModel): + model_args: str | None = None + options: dict[str, Any] = Field(default_factory=dict) + + +class NativeOptions(StrictModel): + options: dict[str, Any] = Field(default_factory=dict) + + +class Trainer(StrictModel): + backend: Literal["miles", "megatron"] = "miles" + resources: Resources + scaling: TrainerScaling = Field(default_factory=TrainerScaling) + engine: EngineOptions = Field(default_factory=EngineOptions) + miles: MilesOptions | None = None + megatron: NativeOptions | None = None + env: dict[str, str] = Field(default_factory=dict) + + +class Inference(StrictModel): + backend: Literal["sglang"] = "sglang" + resources: Resources + scaling: InferenceScaling = Field(default_factory=InferenceScaling) + sglang: NativeOptions = Field(default_factory=NativeOptions) + env: dict[str, str] = Field(default_factory=dict) + + +class Secrets(StrictModel): + api: str = "lilo-api" + sampler_proxy: str = "lilo-proxy" + huggingface: str | None = "huggingface-secret" + + +class Storage(StrictModel): + assets: str = "lilo-model-assets" + checkpoints: str = "lilo-checkpoints" + bulletin: str = "lilo-snapshot-bulletin" + + +class ModalSettings(StrictModel): + environment: str | None = None + region: str = "us-west" + + +class Deployment(StrictModel): + frontend: str = Field( + default="lilo-yaml", pattern=r"^[a-zA-Z0-9][a-zA-Z0-9_-]{0,46}$" + ) + mode: Literal["shared"] = "shared" + modal: ModalSettings = Field(default_factory=ModalSettings) + secrets: Secrets = Field(default_factory=Secrets) + storage: Storage = Field(default_factory=Storage) + + +class Lifecycle(StrictModel): + session_idle_timeout_s: int = Field(default=300, gt=0) + pool_idle_timeout_s: int = Field(default=300, gt=0) + sweep_interval_s: int = Field(default=300, gt=0) + + +class DeploymentSpec(StrictModel): + api_version: Literal["lilo/v1"] = "lilo/v1" + name: str = Field(pattern=r"^[a-zA-Z0-9][a-zA-Z0-9_-]{0,63}$") + model: Model + routing: Routing = Field(default_factory=Routing) + deployment: Deployment = Field(default_factory=Deployment) + trainer: Trainer + inference: Inference + lifecycle: Lifecycle = Field(default_factory=Lifecycle) + + @model_validator(mode="after") + def compatible(self): + from lilo.providers.modal.recipe import backend_config, serving_options + + for role in (self.trainer, self.inference): + role.resources.gpu_count + if any(k.startswith("LILO_") for k in role.env): + raise ValueError("LILO_ environment variables are managed by Lilo") + if self.trainer.backend == "miles": + if ( + self.model.parameterization != "lora" + or self.trainer.miles is None + or self.trainer.megatron is not None + ): + raise ValueError( + "Miles requires lora parameterization and trainer.miles" + ) + elif ( + self.model.parameterization != "full" + or self.trainer.megatron is None + or self.trainer.miles is not None + ): + raise ValueError( + "Megatron YAML deployments require full parameterization and trainer.megatron" + ) + if ( + self.model.parameterization == "full" + and self.trainer.engine.max_clients_per_instance != 1 + ): + raise ValueError("FFT trainers admit one client per instance") + if ( + self.trainer.backend == "megatron" + and self.trainer.engine.sampler_persistence_concurrency != 1 + ): + raise ValueError("Megatron requires sampler_persistence_concurrency: 1") + try: + backend_config(self) + serving_options(self) + except TypeError as exc: + raise ValueError(f"invalid backend option types: {exc}") from exc + return self + + +class ResolvedDeployment(StrictModel): + spec: DeploymentSpec + implementation: str + miles_commit: str | None = None + generation: str + active: bool = True + + @property + def definition_id(self) -> str: + return f"yaml_{self.spec.name}_{self.generation[:16]}" + + @property + def asset_path(self) -> str: + digest = hashlib.sha256( + f"{self.spec.model.id}@{self.spec.model.revision}".encode() + ).hexdigest() + return f"/assets/{digest}" + + +def resolve( + spec: DeploymentSpec, + *, + revision: str, + implementation: str, + miles_commit: str | None = None, +) -> ResolvedDeployment: + if not re.fullmatch(r"[a-fA-F0-9]{40,64}", revision): + raise ValueError("model revision must resolve to an exact commit") + value = spec.model_dump() + value["model"]["revision"] = revision + pinned = DeploymentSpec.model_validate(value) + # Include resource and lifecycle policy: each applied version remains self-contained. + # Checkpoint compatibility uses model/topology metadata, not this generation hash. + identity = pinned.model_dump(exclude={"routing"}) + digest = hashlib.sha256( + json.dumps([implementation, identity], sort_keys=True).encode() + ).hexdigest() + return ResolvedDeployment( + spec=pinned, + implementation=implementation, + generation=digest, + miles_commit=miles_commit, + ) + + +class UniqueLoader(yaml.SafeLoader): + pass + + +def _mapping(loader, node): + result = {} + for key_node, value_node in node.value: + key = loader.construct_object(key_node) + if not isinstance(key, str) or key in result: + raise ValueError(f"duplicate or non-string YAML key: {key!r}") + result[key] = loader.construct_object(value_node) + return result + + +UniqueLoader.add_constructor(yaml.resolver.BaseResolver.DEFAULT_MAPPING_TAG, _mapping) + + +def merge(parent: dict, child: dict) -> dict: + return { + key: merge(parent[key], value) + if isinstance(parent.get(key), dict) and isinstance(value, dict) + else value + for key, value in (parent | child).items() + } + + +def preset_path(name: str) -> Path: + if not re.fullmatch(r"[a-zA-Z0-9_-]+", name): + raise ValueError("invalid preset name") + return Path(str(files("lilo").joinpath("presets", name + ".yaml"))) + + +def load(path: str | Path, _seen: tuple[Path, ...] = ()) -> DeploymentSpec: + return DeploymentSpec.model_validate(_load(Path(path), _seen)) + + +def _load(path: Path, seen: tuple[Path, ...]) -> dict: + path = path.resolve() + if path in seen: + raise ValueError(f"cyclic extends: {path}") + data = yaml.load(path.read_text(), Loader=UniqueLoader) + if not isinstance(data, dict): + raise ValueError("deployment YAML must be a mapping") + parent = data.pop("extends", None) + if parent is not None: + if not isinstance(parent, str): + raise ValueError("extends must be a path or builtin:preset") + parent_path = ( + preset_path(parent[8:]) + if parent.startswith("builtin:") + else path.parent / parent + ) + data = merge(_load(parent_path, (*seen, path)), data) + return data + + +def validate_frontend(specs: list[DeploymentSpec]) -> None: + if not specs: + raise ValueError("at least one deployment is required") + if len({s.name for s in specs}) != len(specs): + raise ValueError("duplicate deployment name") + first = specs[0] + for spec in specs: + if spec.deployment != first.deployment or spec.lifecycle != first.lifecycle: + raise ValueError( + "deployments on one frontend must share deployment and lifecycle settings" + ) + defaults, sampling = set(), set() + for spec in specs: + key = (spec.model.id, spec.model.parameterization) + if spec.routing.default: + if key in defaults: + raise ValueError(f"multiple defaults for {key}") + defaults.add(key) + if spec.routing.sampling_default: + if spec.model.id in sampling: + raise ValueError(f"multiple sampling defaults for {spec.model.id}") + sampling.add(spec.model.id) diff --git a/src/lilo/inference/native_sglang.py b/src/lilo/inference/native_sglang.py new file mode 100644 index 0000000..12c193a --- /dev/null +++ b/src/lilo/inference/native_sglang.py @@ -0,0 +1,33 @@ +"""SGLang entrypoint for resolved YAML serving options.""" + +import argparse +import json +import logging +import os +import sys + +from lilo.native_options import apply_defaults + + +def main(): + from sglang.srt.server_args import ServerArgs + from sglang.launch_server import run_server + from sglang.srt.utils import kill_process_tree + from sglang.srt.plugins import load_plugins + + load_plugins() + parser = argparse.ArgumentParser() + ServerArgs.add_cli_args(parser) + argv = ["--model-path", sys.argv[1], "--host", "127.0.0.1", "--port", "8001"] + apply_defaults(parser, json.loads(sys.argv[2]), argv) + raw = parser.parse_args(argv) + logging.basicConfig(level=getattr(logging, raw.log_level.upper())) + args = ServerArgs.from_cli_args(raw) + try: + run_server(args) + finally: + kill_process_tree(os.getpid(), include_parent=False) + + +if __name__ == "__main__": + main() diff --git a/src/lilo/native_options.py b/src/lilo/native_options.py new file mode 100644 index 0000000..c250562 --- /dev/null +++ b/src/lilo/native_options.py @@ -0,0 +1,90 @@ +"""Apply typed YAML values through a backend's own argparse schema.""" + +from __future__ import annotations + +import argparse + + +def apply_defaults( + parser: argparse.ArgumentParser, options: dict, argv: list[str] +) -> None: + """Override preset arguments after parser construction, before backend validation. + + Values use argparse destination names. Ordinary scalar/list options and boolean + flags are supported. Custom argparse actions fail explicitly instead of being + bypassed by set_defaults. + """ + by_name = {} + for action in parser._actions: + by_name.setdefault(action.dest, []).append(action) + flags, defaults = {}, {} + booleans = ( + argparse._StoreTrueAction, + argparse._StoreFalseAction, + argparse.BooleanOptionalAction, + ) + for key, value in options.items(): + actions = by_name.get(key, []) + if not actions or not all(action.option_strings for action in actions): + raise ValueError(f"unknown backend option: {key}") + if value is None: + raise ValueError(f"backend option {key} cannot be null") + if all(isinstance(action, booleans) for action in actions): + if not isinstance(value, bool): + raise ValueError(f"backend option {key} requires a boolean") + else: + if len(actions) != 1 or type(actions[0]) is not argparse._StoreAction: + raise ValueError( + f"backend option {key} uses an unsupported argparse action" + ) + action = actions[0] + multiple = action.nargs in ("+", "*") or isinstance(action.nargs, int) + if multiple and not isinstance(value, list): + raise ValueError(f"backend option {key} requires a list") + if not multiple and isinstance(value, (dict, list, bool)): + raise ValueError(f"backend option {key} requires a scalar") + values = value if multiple else [value] + if action.nargs == "+" and not values: + raise ValueError(f"backend option {key} requires a nonempty list") + if isinstance(action.nargs, int) and len(values) != action.nargs: + raise ValueError(f"backend option {key} requires {action.nargs} values") + cast = action.type or (lambda x: x) + try: + values = [cast(v) for v in values] + except (ValueError, TypeError, argparse.ArgumentTypeError) as exc: + raise ValueError( + f"invalid value for backend option {key}: {exc}" + ) from exc + if action.choices is not None and any( + v not in action.choices for v in values + ): + raise ValueError(f"invalid choice for backend option {key}") + value = values if multiple else values[0] + defaults[key] = value + for action in actions: + action.required = False + for flag in action.option_strings: + flags[flag] = action + known = {flag for action in parser._actions for flag in action.option_strings} + kept, i = [], 0 + while i < len(argv): + token = argv[i] + action = flags.get(token.split("=", 1)[0]) + if action is None: + kept.append(token) + i += 1 + continue + i += 1 + if "=" in token or action.nargs == 0: + continue + remaining = action.nargs if isinstance(action.nargs, int) else 1 + while i < len(argv): + if argv[i].split("=", 1)[0] in known or argv[i].startswith("--"): + break + i += 1 + if action.nargs not in ("+", "*"): + remaining -= 1 + if remaining == 0: + break + argv[:] = kept + parser.set_defaults(**defaults) diff --git a/src/lilo/presets/qwen35-4b-fft-64k.yaml b/src/lilo/presets/qwen35-4b-fft-64k.yaml new file mode 100644 index 0000000..b8015a7 --- /dev/null +++ b/src/lilo/presets/qwen35-4b-fft-64k.yaml @@ -0,0 +1,47 @@ +api_version: lilo/v1 +name: qwen35-4b-fft-64k +model: + id: Qwen/Qwen3.5-4B + revision: main + parameterization: full + max_context_length: 65536 +routing: + default: true +deployment: + frontend: lilo-yaml +trainer: + backend: megatron + resources: + gpu: H100:4 + engine: + max_clients_per_instance: 1 + sampler_persistence_concurrency: 1 + megatron: + options: + tensor_model_parallel_size: 2 + context_parallel_size: 2 + sequence_parallel: true + micro_batch_size: 1 + max_tokens_per_microbatch: 65536 + defer_fp32_logits: true + fp32_lm_head: true + use_distributed_optimizer: true + provider_overrides: + mtp_num_layers: 0 + recompute_granularity: full + recompute_method: uniform + recompute_num_layers: 1 + optimizer: + lr: 0.0001 + min_lr: 0.0001 + loss_scale: 1.0 +inference: + resources: + gpu: H100:1 + sglang: + options: + tp_size: 1 + mem_fraction_static: 0.85 + max_running_requests: 32 + max_queued_requests: 4 + cpu_weight_cache_max_compile_group_gb: 16 diff --git a/src/lilo/presets/qwen35-9b-lora-16k.yaml b/src/lilo/presets/qwen35-9b-lora-16k.yaml new file mode 100644 index 0000000..cf9ee63 --- /dev/null +++ b/src/lilo/presets/qwen35-9b-lora-16k.yaml @@ -0,0 +1,54 @@ +api_version: lilo/v1 +name: qwen35-9b-lora-16k +model: + id: Qwen/Qwen3.5-9B-Base + revision: main + parameterization: lora + max_context_length: 16384 +routing: + default: true +deployment: + frontend: lilo-yaml + mode: shared +trainer: + backend: miles + resources: + gpu: H100:4 + cpu: 16 + memory_mib: 65536 + scaling: + max_instances: 1 + engine: + max_clients_per_instance: 6 + sampler_persistence_concurrency: 8 + miles: + model_args: qwen3.5-9B + options: + tensor_model_parallel_size: 4 + multi_lora_n_adapters: 6 + lora_rank: 32 + lora_alpha: 32 + target_modules: [linear_qkv, linear_proj, linear_fc1, linear_fc2, output_layer] + max_tokens_per_gpu: 16384 + recompute_granularity: full + recompute_method: uniform + recompute_num_layers: 1 + env: + PYTORCH_CUDA_ALLOC_CONF: expandable_segments:True + TORCHINDUCTOR_COMPILE_THREADS: '1' +inference: + backend: sglang + resources: + gpu: H200:1 + scaling: + min_replicas: 0 + max_replicas: 8 + target_concurrency: 16 + sglang: + options: + tp_size: 1 + mem_fraction_static: 0.8 + max_running_requests: 32 + max_queued_requests: 8 + max_loaded_loras: 64 + max_loras_per_batch: 8 diff --git a/src/lilo/presets/qwen35-9b-lora-64k.yaml b/src/lilo/presets/qwen35-9b-lora-64k.yaml new file mode 100644 index 0000000..a434fdd --- /dev/null +++ b/src/lilo/presets/qwen35-9b-lora-64k.yaml @@ -0,0 +1,13 @@ +extends: builtin:qwen35-9b-lora-16k +name: qwen35-9b-lora-64k +model: + max_context_length: 65536 +routing: + default: false +trainer: + resources: + gpu: H200:8 + miles: + options: + tensor_model_parallel_size: 8 + max_tokens_per_gpu: 65536 diff --git a/src/lilo/providers/modal/app.py b/src/lilo/providers/modal/app.py index df55dbc..781083f 100644 --- a/src/lilo/providers/modal/app.py +++ b/src/lilo/providers/modal/app.py @@ -20,21 +20,6 @@ ModalCheckpointStorage, checkpoint_volume, ) -from .definitions import ( - qwen3_5_4b_full_64k, - qwen3_5_9b_base_miles_lora_2k, - qwen3_5_9b_base_miles_lora_16k, - qwen3_5_9b_base_miles_lora_16k_single, - qwen3_5_9b_full_64k, - qwen3_5_9b_miles_lora_16k, - qwen3_5_9b_miles_lora_16k_dp2, - qwen3_5_35b_a3b_full_64k, - qwen3_6_27b_full_64k, - qwen3_6_35b_a3b_full_64k, - qwen3_8_27b_miles_lora_16k, - qwen3_8_27b_miles_lora_64k, - qwen3_8_27b_miles_lora_128k, -) from .deployment import ( trainer_deployment_env, trainer_max_containers, @@ -47,7 +32,7 @@ proxy_auth_headers, stop_pool, ) -from .image_dependencies import STITCH_PACKAGE +from .image_dependencies import CORE_PACKAGES, STITCH_PACKAGE, TINKER_PACKAGE from .kv import ( ModalSessionKeyValueStores, fft_pool_kv, @@ -66,6 +51,12 @@ stop_pool as stop_lora_pool, ) from .sampling import ModalSamplingTaskPlatform +from .yaml_apps import ( + MANIFEST_ENV, + definition_from_spec, + frontend_settings, + manifest_from_env, +) APP_NAME = os.environ.get("LILO_APP_NAME", "lilo") ROUTING_REGION = "us-west" @@ -81,23 +72,58 @@ _lora_pool_gateways: dict[str, tuple[float, str]] = {} _lora_pool_checks: dict[str, asyncio.Lock] = {} -DEFINITIONS = ( - qwen3_5_4b_full_64k, - qwen3_5_9b_full_64k, - qwen3_5_9b_base_miles_lora_2k, - qwen3_5_9b_base_miles_lora_16k, - qwen3_5_9b_base_miles_lora_16k_single, - qwen3_5_9b_miles_lora_16k, - qwen3_5_9b_miles_lora_16k_dp2, - qwen3_5_35b_a3b_full_64k, - qwen3_6_27b_full_64k, - qwen3_6_35b_a3b_full_64k, - qwen3_8_27b_miles_lora_16k, - qwen3_8_27b_miles_lora_64k, - qwen3_8_27b_miles_lora_128k, -) + +YAML_SETTINGS = frontend_settings() +if YAML_SETTINGS is not None: + APP_NAME = YAML_SETTINGS.deployment.frontend + ROUTING_REGION = YAML_SETTINGS.deployment.modal.region + SESSION_IDLE_TIMEOUT = YAML_SETTINGS.lifecycle.session_idle_timeout_s + FFT_POOL_IDLE_TIMEOUT = LORA_POOL_IDLE_TIMEOUT = ( + YAML_SETTINGS.lifecycle.pool_idle_timeout_s + ) + SWEEP_PERIOD = modal.Period(seconds=YAML_SETTINGS.lifecycle.sweep_interval_s) + CHECKPOINT_VOLUME_NAME = YAML_SETTINGS.deployment.storage.checkpoints + checkpoint_volume = modal.Volume.from_name( + CHECKPOINT_VOLUME_NAME, create_if_missing=True, version=2 + ) + DEFINITIONS = tuple(definition_from_spec(row) for row in manifest_from_env()) +else: + from .definitions import ( + qwen3_5_4b_full_64k, + qwen3_5_9b_base_miles_lora_2k, + qwen3_5_9b_base_miles_lora_16k, + qwen3_5_9b_base_miles_lora_16k_single, + qwen3_5_9b_full_64k, + qwen3_5_9b_miles_lora_16k, + qwen3_5_9b_miles_lora_16k_dp2, + qwen3_5_35b_a3b_full_64k, + qwen3_6_27b_full_64k, + qwen3_6_35b_a3b_full_64k, + qwen3_8_27b_miles_lora_16k, + qwen3_8_27b_miles_lora_64k, + qwen3_8_27b_miles_lora_128k, + ) + + DEFINITIONS = ( + qwen3_5_4b_full_64k, + qwen3_5_9b_full_64k, + qwen3_5_9b_base_miles_lora_2k, + qwen3_5_9b_base_miles_lora_16k, + qwen3_5_9b_base_miles_lora_16k_single, + qwen3_5_9b_miles_lora_16k, + qwen3_5_9b_miles_lora_16k_dp2, + qwen3_5_35b_a3b_full_64k, + qwen3_6_27b_full_64k, + qwen3_6_35b_a3b_full_64k, + qwen3_8_27b_miles_lora_16k, + qwen3_8_27b_miles_lora_64k, + qwen3_8_27b_miles_lora_128k, + ) TRAINER_MAX_CONTAINERS = trainer_max_containers() TRAINER_DEPLOYMENT_ENV = trainer_deployment_env() +if YAML_SETTINGS is not None: + TRAINER_DEPLOYMENT_ENV[MANIFEST_ENV] = os.environ[MANIFEST_ENV] + TRAINER_DEPLOYMENT_ENV["LILO_APP_NAME"] = APP_NAME app = modal.App(APP_NAME) for definition in DEFINITIONS: app.include(definition.app) @@ -124,13 +150,23 @@ async def _delete_checkpoint(uri: str) -> None: image = ( modal.Image.debian_slim(python_version="3.11") .apt_install("git") - .pip_install_from_pyproject("pyproject.toml") + .pip_install(*CORE_PACKAGES, TINKER_PACKAGE) .pip_install(STITCH_PACKAGE, "huggingface-hub") .add_local_python_source("lilo") + .env(TRAINER_DEPLOYMENT_ENV) +) +model_assets = modal.Volume.from_name( + YAML_SETTINGS.deployment.storage.assets if YAML_SETTINGS else "lilo-model-assets", + create_if_missing=True, +) +API_SECRET_NAME = YAML_SETTINGS.deployment.secrets.api if YAML_SETTINGS else "lilo-api" +HF_SECRET_NAME = ( + YAML_SETTINGS.deployment.secrets.huggingface + if YAML_SETTINGS + else "huggingface-secret" ) -model_assets = modal.Volume.from_name("lilo-model-assets", create_if_missing=True) proxy_secret = modal.Secret.from_name( - "lilo-proxy", + YAML_SETTINGS.deployment.secrets.sampler_proxy if YAML_SETTINGS else "lilo-proxy", required_keys=["MODAL_PROXY_TOKEN_ID", "MODAL_PROXY_TOKEN_SECRET"], ) @@ -138,7 +174,7 @@ async def _delete_checkpoint(uri: str) -> None: @app.function( image=image, volumes={MODEL_ASSET_ROOT: model_assets}, - secrets=[modal.Secret.from_name("huggingface-secret")], + secrets=[modal.Secret.from_name(HF_SECRET_NAME)] if HF_SECRET_NAME else [], timeout=4 * 60 * 60, max_containers=1, retries=2, @@ -153,7 +189,10 @@ def prepare_model_assets(definition_id: str) -> None: or checkpoint == MODEL_ASSET_ROOT ): raise ValueError(f"invalid model asset path: {checkpoint}") - snapshot_download(repo_id=definition.MODEL_NAME, local_dir=checkpoint) + kwargs = {} + if hasattr(definition, "MODEL_REVISION"): + kwargs["revision"] = definition.MODEL_REVISION + snapshot_download(repo_id=definition.MODEL_NAME, local_dir=checkpoint, **kwargs) model_assets.commit() @@ -235,7 +274,7 @@ async def _ready_lora_pool(spec: LoraPoolSpec) -> str: min_containers=0, timeout=60 * 60, retries=2, - secrets=[proxy_secret, modal.Secret.from_name("lilo-api")], + secrets=[proxy_secret, modal.Secret.from_name(API_SECRET_NAME)], ) @modal.concurrent(max_inputs=128) async def execute_sample(task: dict) -> dict: @@ -355,11 +394,13 @@ async def trainer_reconciler(delay_seconds: float = 0.0) -> None: async def run(definition_id: str, token: str) -> None: parameterization = parameterization_for(definition_id) - if parameterization is None: + if parameterization is None or await deployment_error(definition_id): await complete_reconcile(definition_id, token) return module = module_for(definition_id) - maximum_instances = TRAINER_MAX_CONTAINERS + maximum_instances = getattr( + module, "TRAINER_MAX_CONTAINERS", TRAINER_MAX_CONTAINERS + ) try: await reconcile_trainers( shared_kv(), @@ -405,11 +446,29 @@ async def spawn(delay_seconds: float) -> str: async def _spawn_engine(definition_id: str, instance_id: str) -> str: + if error := await deployment_error(definition_id): + raise ValueError(error) engine = module_for(definition_id).ENGINE_FUNCTION call = await engine.spawn.aio(instance_id) return call.object_id +async def deployment_error(definition_id: str) -> str | None: + if not definition_id.startswith("yaml_"): + return None + record = await shared_kv().get(f"deployment_failure:{definition_id}") + if record: + return f"Trainer startup failed for {definition_id}: {record['error']}. See Modal call logs for instance {record['instance_id']}; after correcting the cause, run lilo deployment retry." + return None + + +@app.function(image=image) +async def clear_deployment_failure(definition_id: str) -> None: + module_for(definition_id) + await shared_kv().delete(f"deployment_failure:{definition_id}") + await kick_trainer_reconciler(definition_id) + + def _plane(): from lilo.control_plane import ControlPlane @@ -437,6 +496,8 @@ async def ensure_pool(session) -> None: definition_id = session.engine_definition_id parameterization = parameterization_for(definition_id) if parameterization == "lora": + if definition_id.startswith("yaml_") and session.model_id is None: + await prepare_model_assets.remote.aio(definition_id) await _ready_lora_pool(LoraPoolSpec(definition_id)) return if parameterization != "full": @@ -468,7 +529,9 @@ async def kick_trainers(definition_id: str) -> bool: await kick_trainer_reconciler(definition_id) if not trainer_autoscaling(definition_id): return False - maximum = TRAINER_MAX_CONTAINERS + maximum = getattr( + module_for(definition_id), "TRAINER_MAX_CONTAINERS", TRAINER_MAX_CONTAINERS + ) if maximum is None: return True instances = [ @@ -487,6 +550,7 @@ async def kick_trainers(definition_id: str) -> bool: session_idle_timeout=SESSION_IDLE_TIMEOUT, ensure_sampling_pool=ensure_pool, prepare_model=prepare_model, + creation_error=deployment_error, sampling_task_stores=task_stores, read_checkpoint_metadata=_read_checkpoint_metadata, list_checkpoints=_list_checkpoints, @@ -504,7 +568,7 @@ async def kick_trainers(definition_id: str) -> bool: routing_region=ROUTING_REGION, timeout=20 * 60, volumes={CHECKPOINT_ROOT: checkpoint_volume}, - secrets=[modal.Secret.from_name("lilo-api", required_keys=["TINKER_API_KEY"])], + secrets=[modal.Secret.from_name(API_SECRET_NAME, required_keys=["TINKER_API_KEY"])], ) @modal.concurrent(max_inputs=128) @modal.asgi_app(requires_proxy_auth=False) diff --git a/src/lilo/providers/modal/fft_pool.py b/src/lilo/providers/modal/fft_pool.py index 8155fb7..ee67089 100644 --- a/src/lilo/providers/modal/fft_pool.py +++ b/src/lilo/providers/modal/fft_pool.py @@ -131,12 +131,17 @@ def deploy_pool(spec: FFTPoolSpec) -> str: modal_cli = shutil.which("modal") if modal_cli is None: raise RuntimeError("modal CLI is unavailable") - env = {**os.environ, **spec.env()} + from .yaml_apps import pool_environment + + recipe_env = pool_environment(spec.definition_id) + env = {**os.environ, **spec.env(), **recipe_env} command = [ modal_cli, "deploy", "-m", - "lilo.providers.modal.fft_pool_app", + "lilo.providers.modal.yaml_pool_app" + if recipe_env + else "lilo.providers.modal.fft_pool_app", "--name", spec.app_name, ] diff --git a/src/lilo/providers/modal/image_dependencies.py b/src/lilo/providers/modal/image_dependencies.py index ec39f81..e9c73a7 100644 --- a/src/lilo/providers/modal/image_dependencies.py +++ b/src/lilo/providers/modal/image_dependencies.py @@ -5,6 +5,7 @@ "opentelemetry-exporter-otlp-proto-http==1.43.0", "modal>=1.5.3", "protobuf>=5.29", + "pyyaml>=6.0.2", "pydantic>=2.13.4", "uvicorn>=0.52.0", "xxhash>=3.8.1", diff --git a/src/lilo/providers/modal/kv.py b/src/lilo/providers/modal/kv.py index 16af6c2..c37b1ae 100644 --- a/src/lilo/providers/modal/kv.py +++ b/src/lilo/providers/modal/kv.py @@ -42,6 +42,7 @@ "trainer_demand": "models", "engine_instance": "engines", "lora_pool": "engines", + "deployment_failure": "engines", "engine_call": "engines", "engine_current": "engines", "trainer_plan": "engines", diff --git a/src/lilo/providers/modal/lora_pool.py b/src/lilo/providers/modal/lora_pool.py index 6c58b8c..a46376f 100644 --- a/src/lilo/providers/modal/lora_pool.py +++ b/src/lilo/providers/modal/lora_pool.py @@ -66,18 +66,23 @@ def deploy_pool(spec: LoraPoolSpec) -> str: modal_cli = shutil.which("modal") if modal_cli is None: raise RuntimeError("modal CLI is unavailable") + from .yaml_apps import pool_environment + + recipe_env = pool_environment(spec.definition_id) command = [ modal_cli, "deploy", "-m", - "lilo.providers.modal.lora_pool_app", + "lilo.providers.modal.yaml_pool_app" + if recipe_env + else "lilo.providers.modal.lora_pool_app", "--name", spec.app_name, ] environment = os.environ.get("MODAL_ENVIRONMENT") if environment: command.extend(["--env", environment]) - subprocess.run(command, env={**os.environ, **spec.env()}, check=True) + subprocess.run(command, env={**os.environ, **spec.env(), **recipe_env}, check=True) return pool.gateway_url() @@ -99,6 +104,8 @@ def stop_pool(spec: LoraPoolSpec) -> None: def _implementation_revision(definition_id: str) -> str: + if definition_id.startswith("yaml_"): + return definition_id.rsplit("_", 1)[-1] here = Path(__file__) files = ( here, diff --git a/src/lilo/providers/modal/recipe.py b/src/lilo/providers/modal/recipe.py new file mode 100644 index 0000000..65c95cf --- /dev/null +++ b/src/lilo/providers/modal/recipe.py @@ -0,0 +1,193 @@ +"""Translate a deployment recipe into backend and serving configurations.""" + +from __future__ import annotations + +from dataclasses import asdict + +from lilo.backends.miles_config import MilesBackendConfig + +MILES_MANAGED = { + "hf_checkpoint", + "load", + "pretrained_checkpoint", + "train_backend", + "actor_num_nodes", + "actor_num_gpus_per_node", + "rollout_num_gpus", + "debug_train_only", + "megatron_to_hf_mode", + "seq_length", + "pipeline_model_parallel_size", + "virtual_pipeline_model_parallel_size", + "colocate", + "custom_actor", + "sglang_model_path", + "use_dynamic_global_batch_size", + "delay_split_train_data_by_dp", + "use_dynamic_batch_size", + "optimizer", + "gradient_accumulation_fusion", + "save", + "save_interval", + "ckpt_step", + "rollout_num_gpus_per_engine", + "num_gpus_per_node", +} +SGLANG_MANAGED = { + "model_path", + "model", + "host", + "port", + "context_length", + "enable_lora", + "max_lora_rank", + "enable_cpu_weight_cache", + "api_key", + "dp_size", + "pp_size", + "lora_paths", + "enable_dp_attention", + "dist_init_addr", + "nnodes", + "node_rank", + "tokenizer_path", + "tokenizer_revision", + "revision", + "grpc_mode", + "smg_grpc_mode", + "encoder_only", + "use_ray", + "disaggregation_mode", + "skip_tokenizer_init", +} + + +def native_options(values, protected): + result = {} + for key, value in values.items(): + if key.startswith("-") or key.replace("_", "").isalnum() is False: + raise ValueError(f"native option must use its underscore name: {key}") + if key in protected: + raise ValueError(f"option {key} is managed by Lilo") + result[key] = value + return result + + +def backend_config(spec, asset_path="/assets/pending"): + trainer = spec.trainer + if trainer.backend == "miles": + options = native_options(trainer.miles.options, MILES_MANAGED) + names = { + "tensor_model_parallel_size": "tensor_model_parallel_size", + "context_parallel_size": "context_parallel_size", + "expert_model_parallel_size": "expert_model_parallel_size", + "expert_tensor_parallel_size": "expert_tensor_parallel_size", + "multi_lora_n_adapters": "max_lora_slots", + "lora_rank": "max_lora_rank", + "lora_alpha": "default_lora_alpha", + "lora_dropout": "lora_dropout", + "target_modules": "target_modules", + "max_tokens_per_gpu": "max_tokens_per_gpu", + } + settings = { + dest: options.pop(source) + for source, dest in names.items() + if source in options + } + for name, value in settings.items(): + if ( + name not in {"target_modules", "default_lora_alpha", "lora_dropout"} + and type(value) is not int + ): + raise ValueError(f"trainer.miles option {name} must be an integer") + if "target_modules" in settings: + value = settings["target_modules"] + value = value.split(",") if isinstance(value, str) else value + if not isinstance(value, list) or not all( + isinstance(item, str) and item for item in value + ): + raise ValueError( + "target_modules must be a nonempty list of module names" + ) + settings["target_modules"] = tuple(value) + config = MilesBackendConfig( + hf_checkpoint=asset_path, + model_type=trainer.miles.model_args or "", + actor_num_gpus_per_node=trainer.resources.gpu_count, + native_options=options, + extra_args=("--seq-length", str(spec.model.max_context_length)), + **settings, + ) + config.validate() + if config.world_size % ( + config.expert_model_parallel_size * config.expert_tensor_parallel_size + ): + raise ValueError( + "expert parallel sizes must divide the trainer GPU allocation" + ) + if trainer.engine.max_clients_per_instance > config.max_lora_slots: + raise ValueError("max_clients_per_instance exceeds multi_lora_n_adapters") + return {"miles": asdict(config), "checkpoint_dir": "/checkpoints"} + from lilo.backends.megatron_runtime.common.config import ( + EngineModelConfig, + OptimizerConfig, + ) + + options = native_options( + trainer.megatron.options, + {"hf_checkpoint", "seq_length", "max_lora_slots", "max_lora_rank"}, + ) + overrides = options.get("provider_overrides", {}) + protected = { + "tensor_model_parallel_size", + "pipeline_model_parallel_size", + "virtual_pipeline_model_parallel_size", + "context_parallel_size", + "expert_model_parallel_size", + "expert_tensor_parallel_size", + "sequence_parallel", + "variable_seq_lengths", + "params_dtype", + "seq_length", + } + native_options(overrides, protected) + try: + optimizer = OptimizerConfig(**options.pop("optimizer", {})) + except TypeError as exc: + raise ValueError(f"invalid trainer.megatron optimizer options: {exc}") from exc + try: + config = EngineModelConfig( + hf_checkpoint=asset_path, + seq_length=spec.model.max_context_length, + optimizer=optimizer, + **options, + ) + except TypeError as exc: + raise ValueError(f"invalid trainer.megatron options: {exc}") from exc + config.validate(trainer.resources.gpu_count) + return {"megatron": asdict(config), "checkpoint_dir": "/checkpoints"} + + +def serving_options(spec): + options = native_options(spec.inference.sglang.options, SGLANG_MANAGED) + tp = options.get("tp_size", spec.inference.resources.gpu_count) + ep = options.get("ep_size", 1) + if not isinstance(tp, int) or tp != spec.inference.resources.gpu_count: + raise ValueError( + "sglang.tp_size must equal the replica GPU allocation; DP-attention is not yet supported in YAML" + ) + if not isinstance(ep, int) or ep < 1 or tp % ep: + raise ValueError("sglang.ep_size must divide the replica GPU allocation") + for key in ( + "max_loaded_loras", + "max_loras_per_batch", + "max_running_requests", + "max_queued_requests", + ): + if key in options and (not isinstance(options[key], int) or options[key] < 1): + raise ValueError(f"sglang.{key} must be positive") + if not 0 < options.get("mem_fraction_static", 0.8) < 1: + raise ValueError("sglang.mem_fraction_static must be between zero and one") + if options.get("max_loaded_loras", 64) < options.get("max_loras_per_batch", 8): + raise ValueError("max_loaded_loras must be >= max_loras_per_batch") + return options diff --git a/src/lilo/providers/modal/serve.py b/src/lilo/providers/modal/serve.py index 9aba19f..76d573b 100644 --- a/src/lilo/providers/modal/serve.py +++ b/src/lilo/providers/modal/serve.py @@ -119,6 +119,7 @@ def run_engine_with_backend( startup_timeout: float = BACKEND_STARTUP_TIMEOUT, operation_timeout: float = BACKEND_OPERATION_TIMEOUT, notify_reconciler: bool = True, + on_startup_error: Callable[[Exception], Awaitable[None]] | None = None, ) -> None: if sampler_persistence_concurrency > 1 and nproc != 1: raise ValueError( @@ -168,7 +169,10 @@ def signal_backend(sig: signal.Signals) -> None: on_transport_error=lambda: signal_backend(signal.SIGKILL), ) + ready = False + async def make_server() -> Engine: + nonlocal ready try: async with asyncio.timeout(startup_timeout): while True: @@ -178,6 +182,7 @@ async def make_server() -> Engine: ) try: if (await executor.http.get("/healthz")).is_success: + ready = True return Engine( executor, max_models=max_models, @@ -210,6 +215,10 @@ async def serve_until_backend_exits() -> None: await serving return raise RuntimeError(f"backend exited with code {backend.returncode}") + except Exception as exc: + if not ready and on_startup_error is not None: + await on_startup_error(exc) + raise finally: serving.cancel() await asyncio.gather(serving, return_exceptions=True) diff --git a/src/lilo/providers/modal/yaml_apps.py b/src/lilo/providers/modal/yaml_apps.py new file mode 100644 index 0000000..3cacbc3 --- /dev/null +++ b/src/lilo/providers/modal/yaml_apps.py @@ -0,0 +1,323 @@ +"""Modal app builders shared by all YAML model deployments. + +Trainer functions live in the frontend app. Rollout pools are separate apps, +created on demand with the existing LoRA/FFT pool lifecycle. +""" + +from __future__ import annotations + +import json +import os +from types import SimpleNamespace + +import modal + +from lilo.deployments import ResolvedDeployment, validate_frontend +from .recipe import backend_config, serving_options + +MANIFEST_ENV = "LILO_DEPLOYMENT_MANIFEST" +POOL_CONFIG_ENV = "LILO_POOL_DEPLOYMENT" + + +def manifest_from_env(): + data = os.environ.get(MANIFEST_ENV) + return ( + [ResolvedDeployment.model_validate(row) for row in json.loads(data)] + if data + else [] + ) + + +def frontend_settings(): + deployments = manifest_from_env() + if not deployments: + return None + active = [row.spec for row in deployments if row.active] + validate_frontend(active) + if len({row.definition_id for row in deployments}) != len(deployments): + raise ValueError("duplicate deployment generation") + commits = {row.miles_commit for row in deployments if row.miles_commit} + if len(commits) > 1: + raise ValueError("manifest contains different Miles runtimes") + if commits: + os.environ["LILO_MILES_COMMIT"] = commits.pop() + return active[0] + + +def image_for(backend): + if not modal.is_local(): + return modal.Image.debian_slim() + if backend == "miles": + from .miles_image import image + elif backend == "megatron": + from .megatron_image import image + else: + from .rollout_image import image + return image + + +def volumes_for(spec): + storage = spec.deployment.storage + return { + "/assets": modal.Volume.from_name(storage.assets, create_if_missing=True), + "/checkpoints": modal.Volume.from_name( + storage.checkpoints, create_if_missing=True, version=2 + ), + "/bulletin": modal.Volume.from_name( + storage.bulletin, create_if_missing=True, version=2 + ), + } + + +def secrets_for(spec, *, training=False): + names = spec.deployment.secrets + result = [modal.Secret.from_name(names.api, required_keys=["TINKER_API_KEY"])] + if training: + result.append( + modal.Secret.from_name( + names.sampler_proxy, + required_keys=["MODAL_PROXY_TOKEN_ID", "MODAL_PROXY_TOKEN_SECRET"], + ) + ) + if names.huggingface: + result.append(modal.Secret.from_name(names.huggingface)) + return result + + +def build_trainer_app(resolved: ResolvedDeployment, *, image=None): + spec = resolved.spec + app = modal.App(f"{spec.deployment.frontend}-{resolved.definition_id}") + resource = spec.trainer.resources + from .deployment import trainer_deployment_env + + env = { + **trainer_deployment_env(), + **spec.trainer.env, + "LILO_APP_NAME": spec.deployment.frontend, + } + + @app.function( + name=resolved.definition_id, + serialized=True, + image=image if image is not None else image_for(spec.trainer.backend), + gpu=resource.gpu, + region=spec.deployment.modal.region, + cpu=resource.cpu, + memory=resource.memory_mib, + timeout=resource.timeout_s, + max_containers=spec.trainer.scaling.max_instances, + min_containers=0, + single_use_containers=True, + volumes=volumes_for(spec), + secrets=secrets_for(spec, training=True), + env=env, + ) + def trainer(instance_id: str): + run_trainer(resolved, instance_id) + + return app, trainer + + +def run_trainer(resolved, instance_id): + from modal.config import config + from .kv import shared_kv + from .serve import run_engine_with_backend + + spec = resolved.spec + settings = backend_config(spec, resolved.asset_path) + # Assets are prepared by the frontend before demand is registered. Reload once + # on startup to see the committed exact snapshot; never race a trainer download. + volumes_for(spec)["/assets"].reload() + env = { + **spec.trainer.env, + "LILO_APP_NAME": spec.deployment.frontend, + "LILO_BACKEND_CONFIG": json.dumps(settings), + "LILO_BASE_MODEL": spec.model.id, + "LILO_BASE_MODEL_REVISION": spec.model.revision, + "LILO_DEFINITION_ID": resolved.definition_id, + "LILO_CHECKPOINT_VOLUME": spec.deployment.storage.checkpoints, + "LILO_BULLETIN_ROOT": "/bulletin", + "LILO_BULLETIN_VOLUME": spec.deployment.storage.bulletin, + "LILO_DEFINITION_REVISION": resolved.generation, + } + executor = ( + "lilo.backends.miles_lora:build_executor" + if spec.trainer.backend == "miles" + else "lilo.backends.megatron_fft:build_executor" + ) + + async def failed(error): + await shared_kv().put( + f"deployment_failure:{resolved.definition_id}", + {"error": str(error), "instance_id": instance_id}, + ) + + run_engine_with_backend( + shared_kv(), + executor, + definition_id=resolved.definition_id, + revision=config["image_id"], + instance_id=instance_id, + backend_env=env, + nproc=1 + if spec.trainer.backend == "miles" + else spec.trainer.resources.gpu_count, + max_models=spec.trainer.engine.max_clients_per_instance, + sampler_persistence_concurrency=spec.trainer.engine.sampler_persistence_concurrency, + on_startup_error=failed, + ) + + +def definition_from_spec(resolved, *, register_trainer=True, image=None): + spec = resolved.spec + native = serving_options(spec) + definition = SimpleNamespace( + DEFINITION_ID=resolved.definition_id, + MODEL_NAME=spec.model.id, + MODEL_REVISION=spec.model.revision, + HF_CHECKPOINT=resolved.asset_path, + PARAMETERIZATION=spec.model.parameterization, + CATALOG_VISIBLE=resolved.active, + ROUTING_DEFAULT=spec.routing.default, + SAMPLING_DEFAULT=spec.routing.sampling_default, + DEPLOYMENT_NAME=spec.name, + RESOLVED=resolved, + MAX_CONTEXT_LENGTH=spec.model.max_context_length, + TRAINER_MODELS_PER_INSTANCE=spec.trainer.engine.max_clients_per_instance, + TRAINER_MAX_CONTAINERS=spec.trainer.scaling.max_instances, + ROLLOUT_GPUS=spec.inference.resources.gpu_count, + ROLLOUT_TENSOR_PARALLEL_SIZE=native.get( + "tp_size", spec.inference.resources.gpu_count + ), + ) + if register_trainer: + definition.app, definition.ENGINE_FUNCTION = build_trainer_app( + resolved, image=image + ) + return definition + + +def pool_deployment(definition_id): + for row in manifest_from_env(): + if row.definition_id == definition_id: + return row + return None + + +def pool_environment(definition_id): + resolved = pool_deployment(definition_id) + if resolved is None: + if definition_id.startswith("yaml_"): + raise ValueError(f"missing recorded deployment: {definition_id}") + return {} + return {POOL_CONFIG_ENV: resolved.model_dump_json()} + + +def build_rollout_app(resolved, pool, *, image=None): + """Create one frozen-base LoRA pool or one FFT latest/pinned/base pool.""" + spec = resolved.spec + lora = spec.model.parameterization == "lora" + if pool.definition_id != resolved.definition_id: + raise ValueError("pool generation does not match deployment") + app = modal.App(pool.app_name) + resources, scaling = spec.inference.resources, spec.inference.scaling + options = { + "context_length": spec.model.max_context_length, + "tp_size": resources.gpu_count, + "mem_fraction_static": 0.8, + "max_running_requests": 32, + "weight_loader_disable_mmap": True, + **serving_options(spec), + } + if lora: + config = backend_config(spec)["miles"] + from lilo.backends.miles_config import MilesBackendConfig + + targets = MilesBackendConfig(**config).peft_target_modules + options.update(enable_lora=True, max_lora_rank=config["max_lora_rank"]) + options.setdefault("lora_target_modules", list(targets)) + options.setdefault("max_loaded_loras", 64) + options.setdefault("max_loras_per_batch", 8) + else: + options["enable_cpu_weight_cache"] = True + minimum = getattr(pool, "min_containers", None) + maximum = getattr(pool, "max_containers", None) + window = getattr(pool, "scaledown_window", None) + + # Serialized class captures the spec; no model-specific module is imported. + @app.server( + name="Server", + serialized=True, + image=image if image is not None else image_for("sglang"), + gpu=resources.gpu, + cpu=resources.cpu, + memory=resources.memory_mib, + volumes=volumes_for(spec), + secrets=secrets_for(spec), + env=spec.inference.env, + min_containers=scaling.min_replicas if minimum is None else minimum, + max_containers=scaling.max_replicas if maximum is None else maximum, + target_concurrency=scaling.target_concurrency, + scaledown_window=scaling.scaledown_window_s if window is None else window, + startup_timeout=1200, + exit_grace_period=300, + port=8000, + routing_region=spec.deployment.modal.region, + compute_region=spec.deployment.modal.region, + ) + class Server: + @modal.enter() + def start(self): + import subprocess + import sys + from lilo.inference.serving import ( + start_lora_sidecar, + start_fft_sidecar, + supervise_children, + terminate, + wait_http, + ) + + self.sglang = subprocess.Popen( + [ + sys.executable, + "-m", + "lilo.inference.native_sglang", + resolved.asset_path, + json.dumps(options), + ], + start_new_session=True, + ) + try: + wait_http("http://127.0.0.1:8001/health", self.sglang, 1200) + kwargs = dict( + port=8000, + sglang_port=8001, + bulletin_root="/bulletin", + bulletin_volume=spec.deployment.storage.bulletin, + ) + self.sidecar = ( + start_lora_sidecar(**kwargs) + if lora + else start_fft_sidecar( + **kwargs, + model_path=resolved.asset_path, + run_id=pool.model_id, + pinned_version=None if pool.latest else pool.version, + ) + ) + self.supervisor = supervise_children(self.sglang, self.sidecar) + wait_http("http://127.0.0.1:8000/health", self.sidecar, 1200) + except BaseException: + terminate(getattr(self, "sidecar", None)) + terminate(self.sglang) + raise + + @modal.exit() + def stop(self): + from lilo.inference.serving import terminate + + terminate(getattr(self, "sidecar", None)) + terminate(getattr(self, "sglang", None)) + + return app, Server diff --git a/src/lilo/providers/modal/yaml_pool_app.py b/src/lilo/providers/modal/yaml_pool_app.py new file mode 100644 index 0000000..65bfea5 --- /dev/null +++ b/src/lilo/providers/modal/yaml_pool_app.py @@ -0,0 +1,25 @@ +"""Generic rollout app constructed in the pool deployment subprocess.""" + +import os + +from lilo.deployments import ResolvedDeployment +from .fft_pool import FFTPoolSpec +from .lora_pool import LoraPoolSpec +from .yaml_apps import POOL_CONFIG_ENV, build_rollout_app + +resolved = ResolvedDeployment.model_validate_json(os.environ[POOL_CONFIG_ENV]) +if resolved.spec.model.parameterization == "lora": + pool = LoraPoolSpec(resolved.definition_id, revision=resolved.generation[:16]) +else: + pool = FFTPoolSpec( + definition_id=resolved.definition_id, + model_id=os.environ["LILO_FFT_POOL_MODEL_ID"], + latest=os.environ["LILO_FFT_POOL_LATEST"] == "1", + version=int(os.environ["LILO_FFT_POOL_VERSION"]), + **{ + key: int(os.environ[f"LILO_FFT_POOL_{key.upper()}"]) + for key in ("min_containers", "max_containers", "scaledown_window") + if f"LILO_FFT_POOL_{key.upper()}" in os.environ + }, + ) +app, Server = build_rollout_app(resolved, pool) diff --git a/tests/backends/test_megatron_fft.py b/tests/backends/test_megatron_fft.py index d54c37c..4225707 100644 --- a/tests/backends/test_megatron_fft.py +++ b/tests/backends/test_megatron_fft.py @@ -903,3 +903,28 @@ def _impl(self, **kwargs): assert calls[0]["sequence_parallel"] is True assert model.output_layer.weight.grad is not None assert not hasattr(model.decoder, "_forward_impl") + + +def test_native_resume_rejects_different_or_unknown_base_revision( + tmp_path, monkeypatch +): + config = EngineModelConfig(hf_checkpoint="/model") + monkeypatch.setenv("LILO_BASE_MODEL_REVISION", "a" * 40) + metadata = fft_checkpoint.create_fft_checkpoint_metadata( + config, + checkpoint_id="snapshot", + base_model=BASE_MODEL, + include_optimizer=True, + world_size=1, + ) + assert metadata.base_model_revision == "a" * 40 + for revision in ("b" * 40, None): + saved = metadata.to_dict() + saved["base_model_revision"] = revision + (tmp_path / fft_checkpoint.CHECKPOINT_METADATA_FILENAME).write_text( + json.dumps(saved) + ) + with pytest.raises(ValueError, match="base_model_revision"): + fft_checkpoint.load_fft_training_checkpoint( + str(tmp_path), config, base_model=BASE_MODEL, world_size=1 + ) diff --git a/tests/backends/test_miles.py b/tests/backends/test_miles.py index 8c33d30..beb9b6a 100644 --- a/tests/backends/test_miles.py +++ b/tests/backends/test_miles.py @@ -709,3 +709,22 @@ def test_optimizer_worker_failure_is_fatal(tmp_path, outcome): runtime.optim_step = lambda parameters: {0: outcome} with pytest.raises(BackendFailed): backend.optim_step(("a",), AdamParams(learning_rate=1e-4)) + + +def test_checkpoint_rejects_different_or_unknown_pinned_base_revision( + tmp_path, monkeypatch +): + backend = _backend(tmp_path) + backend.accept_model("model-a", _spec()) + monkeypatch.setenv("LILO_BASE_MODEL_REVISION", "a" * 40) + backend.capture_checkpoint( + "model-a", "capture", destination="step-1", include_optimizer=False + ) + uri = Path(backend.persist_checkpoint("capture", "step-1")) + metadata = json.loads((uri / "metadata.json").read_text()) + assert metadata["base_model_revision"] == "a" * 40 + backend._validate_checkpoint(metadata, backend.jobs["model-a"], False) + for revision in ("b" * 40, None): + metadata["base_model_revision"] = revision + with pytest.raises(ValueError, match="base model revision"): + backend._validate_checkpoint(metadata, backend.jobs["model-a"], False) diff --git a/tests/providers/test_modal_serve.py b/tests/providers/test_modal_serve.py index 3f3ff92..54afc01 100644 --- a/tests/providers/test_modal_serve.py +++ b/tests/providers/test_modal_serve.py @@ -88,3 +88,31 @@ async def serving(kv, make_server, **kwargs): assert signals and all(sig == signal.SIGKILL for sig in signals) assert posted.count("/execute_forward_backward_batch") == 1 assert closed == [True] + + +def test_backend_startup_failure_is_reported_before_health_ready(monkeypatch): + class Backend: + pid = 12345 + returncode = 1 + + def poll(self): + return self.returncode + + backend = Backend() + errors = [] + monkeypatch.setattr(serve.subprocess, "Popen", lambda *args, **kwargs: backend) + monkeypatch.setattr(serve.os, "killpg", lambda *args: None) + + async def record(error): + errors.append(str(error)) + + with pytest.raises(RuntimeError, match="backend exited with code 1"): + serve.run_engine_with_backend( + None, + "unused:executor", + definition_id="test", + revision="test", + instance_id="test", + on_startup_error=record, + ) + assert errors == ["backend exited with code 1"] diff --git a/tests/providers/test_yaml_apps.py b/tests/providers/test_yaml_apps.py new file mode 100644 index 0000000..83eb039 --- /dev/null +++ b/tests/providers/test_yaml_apps.py @@ -0,0 +1,212 @@ +import asyncio +import json +from types import SimpleNamespace + +import modal +import pytest + +from lilo.deployments import load, preset_path, resolve +from lilo.providers.modal import yaml_apps +from lilo.providers.modal.fft_pool import FFTPoolSpec +from lilo.providers.modal.lora_pool import LoraPoolSpec + + +def deployment(preset="qwen35-9b-lora-16k"): + return resolve( + load(preset_path(preset)), revision="a" * 40, implementation="runtime" + ) + + +class App: + def __init__(self, name): + self.name = name + self.functions = {} + self.servers = {} + + def function(self, **settings): + def decorate(fn): + self.functions[settings["name"]] = (settings, fn) + return fn + + return decorate + + def server(self, **settings): + def decorate(cls): + self.servers[settings["name"]] = (settings, cls) + return cls + + return decorate + + +@pytest.fixture +def builders(monkeypatch): + monkeypatch.setattr(modal, "App", App) + monkeypatch.setattr(modal, "enter", lambda: lambda fn: fn) + monkeypatch.setattr(modal, "exit", lambda: lambda fn: fn) + + +def test_trainer_declaration_and_executor_configuration(builders, monkeypatch): + row = deployment() + image = object() + app, trainer = yaml_apps.build_trainer_app(row, image=image) + declaration, _ = app.functions[row.definition_id] + assert declaration["gpu"] == "H100:4" + assert declaration["region"] == "us-west" + assert declaration["max_containers"] == 1 + assert declaration["single_use_containers"] is True + assert declaration["image"] is image + calls = [] + reloaded = [] + from lilo.providers.modal import serve, kv + + monkeypatch.setattr(kv, "shared_kv", lambda: "store") + monkeypatch.setattr( + yaml_apps, + "volumes_for", + lambda spec: {"/assets": SimpleNamespace(reload=lambda: reloaded.append(True))}, + ) + monkeypatch.setattr( + serve, + "run_engine_with_backend", + lambda *args, **kwargs: calls.append((args, kwargs)), + ) + trainer("instance-a") + args, kwargs = calls[0] + assert args == ("store", "lilo.backends.miles_lora:build_executor") + assert kwargs["max_models"] == 6 + assert kwargs["nproc"] == 1 + assert kwargs["backend_env"]["LILO_BASE_MODEL_REVISION"] == "a" * 40 + config = json.loads(kwargs["backend_env"]["LILO_BACKEND_CONFIG"]) + assert config["miles"]["hf_checkpoint"] == row.asset_path + assert reloaded == [True] + + +@pytest.mark.parametrize("kind", ["lora", "base", "latest", "pinned"]) +def test_pool_starts_native_server_and_correct_sidecar(builders, monkeypatch, kind): + row = deployment("qwen35-9b-lora-16k" if kind == "lora" else "qwen35-4b-fft-64k") + pool = ( + LoraPoolSpec(row.definition_id) + if kind == "lora" + else ( + FFTPoolSpec.base(row.definition_id) + if kind == "base" + else FFTPoolSpec(row.definition_id, "job", kind == "latest", 7) + ) + ) + app, server = yaml_apps.build_rollout_app(row, pool, image="test-image") + settings, _ = app.servers["Server"] + assert app.name == pool.app_name + assert settings["gpu"] == row.spec.inference.resources.gpu + assert settings["min_containers"] == 0 + assert settings["target_concurrency"] == 16 + assert settings["compute_region"] == "us-west" + from lilo.inference import serving + import subprocess + + calls, commands, stops = [], [], [] + process = object() + monkeypatch.setattr( + subprocess, "Popen", lambda argv, **kw: (commands.append(argv) or process) + ) + monkeypatch.setattr(serving, "wait_http", lambda *args: None) + monkeypatch.setattr(serving, "supervise_children", lambda *args: None) + monkeypatch.setattr( + serving, + "start_lora_sidecar", + lambda **kw: (calls.append(("lora", kw)) or process), + ) + monkeypatch.setattr( + serving, + "start_fft_sidecar", + lambda **kw: (calls.append(("fft", kw)) or process), + ) + monkeypatch.setattr(serving, "terminate", stops.append) + replica = server() + replica.start() + assert commands[0][2] == "lilo.inference.native_sglang" + assert commands[0][3] == row.asset_path + native = json.loads(commands[0][4]) + assert native["context_length"] == row.spec.model.max_context_length + if kind == "lora": + assert native["enable_lora"] is True + assert native["max_lora_rank"] == 32 + assert native["lora_target_modules"] == [ + "q_proj", + "k_proj", + "v_proj", + "o_proj", + "gate_proj", + "up_proj", + "down_proj", + "lm_head", + ] + else: + assert native["enable_cpu_weight_cache"] is True + assert calls[0][1]["pinned_version"] == ( + None if kind == "latest" else 0 if kind == "base" else 7 + ) + replica.stop() + assert stops == [process, process] + + +def test_pool_subprocess_receives_recorded_generation(monkeypatch): + row = deployment() + monkeypatch.setenv(yaml_apps.MANIFEST_ENV, json.dumps([row.model_dump()])) + env = yaml_apps.pool_environment(row.definition_id) + assert json.loads(env[yaml_apps.POOL_CONFIG_ENV])["generation"] == row.generation + with pytest.raises(ValueError, match="missing recorded"): + yaml_apps.pool_environment("yaml_missing_123") + assert yaml_apps.pool_environment("legacy") == {} + + +def test_startup_failure_is_visible_and_blocks_new_spawns(monkeypatch): + import importlib + from lilo.providers.local import InMemoryKeyValueStore + + app = importlib.import_module("lilo.providers.modal.app") + store = InMemoryKeyValueStore() + row = deployment() + monkeypatch.setattr(app, "shared_kv", lambda: store) + + async def run(): + await store.put( + f"deployment_failure:{row.definition_id}", + {"error": "backend exited with code 1", "instance_id": "failed-instance"}, + ) + with pytest.raises(ValueError, match="Trainer startup failed.*failed-instance"): + await app._spawn_engine(row.definition_id, "new-instance") + assert await app.deployment_error("legacy") is None + + asyncio.run(run()) + + +def test_real_modal_app_constructs_from_manifest_without_legacy_catalog(monkeypatch): + import os + import subprocess + import sys + + row = deployment() + env = {**os.environ, yaml_apps.MANIFEST_ENV: json.dumps([row.model_dump()])} + result = subprocess.run( + [ + sys.executable, + "-c", + """ +import importlib, sys, modal +from lilo.providers.modal import yaml_apps +yaml_apps.image_for = lambda backend: modal.Image.debian_slim() +app = importlib.import_module('lilo.providers.modal.app') +assert len(app.DEFINITIONS) == 1 +assert app.APP_NAME == 'lilo-yaml' +assert app.DEFINITIONS[0].ENGINE_FUNCTION is not None +assert not any(name.startswith('lilo.providers.modal.definitions.') for name in sys.modules) +print('constructed') +""", + ], + env=env, + capture_output=True, + text=True, + timeout=30, + ) + assert result.returncode == 0, result.stderr + assert "constructed" in result.stdout diff --git a/tests/test_deployment_cli.py b/tests/test_deployment_cli.py new file mode 100644 index 0000000..3bd2301 --- /dev/null +++ b/tests/test_deployment_cli.py @@ -0,0 +1,98 @@ +import json +import subprocess + +import modal +import pytest + +from lilo import deployment_cli as cli +from lilo.deployments import load, preset_path, resolve +from lilo.providers.modal.yaml_apps import MANIFEST_ENV + + +class Registry(dict): + def put(self, key, value, skip_if_exists=False): + if skip_if_exists and key in self: + return False + self[key] = value + return True + + +def deployment(): + return resolve( + load(preset_path("qwen35-9b-lora-16k")), + revision="a" * 40, + implementation="test", + ) + + +@pytest.fixture +def registry(monkeypatch): + value = Registry() + monkeypatch.setattr(modal.Dict, "from_name", lambda *args, **kwargs: value) + + def missing(*args, **kwargs): + raise modal.exception.NotFoundError("not deployed") + + monkeypatch.setattr(modal.App, "lookup", missing) + return value + + +def test_deploy_is_serialized_and_commits_only_after_success(registry, monkeypatch): + row = deployment() + seen = [] + + def run(command, **kwargs): + assert "apply_lock" in registry + assert "pending" in registry + assert "manifest" not in registry + assert "lilo.providers.modal.app" in command + seen.extend(json.loads(kwargs["env"][MANIFEST_ENV])) + + monkeypatch.setattr(subprocess, "run", run) + cli.deploy([row]) + assert seen == registry["manifest"] + assert "pending" not in registry and "apply_lock" not in registry + registry["apply_lock"] = "other-operator" + with pytest.raises(ValueError, match="An apply owns"): + cli.deploy([row]) + assert registry["apply_lock"] == "other-operator" + + +def test_failed_apply_keeps_pending_generations_for_next_attempt(registry, monkeypatch): + row = deployment() + + def fail(*args, **kwargs): + raise subprocess.CalledProcessError(1, args[0]) + + monkeypatch.setattr(subprocess, "run", fail) + with pytest.raises(subprocess.CalledProcessError): + cli.deploy([row]) + assert registry["pending"][0]["generation"] == row.generation + assert "manifest" not in registry and "apply_lock" not in registry + new_spec = row.spec.model_copy(deep=True) + new_spec.trainer.scaling.max_instances = 2 + new = resolve(new_spec, revision="a" * 40, implementation="test") + monkeypatch.setattr(subprocess, "run", lambda *args, **kwargs: None) + cli.deploy([new]) + assert [(r["generation"], r["active"]) for r in registry["manifest"]] == [ + (new.generation, True), + (row.generation, False), + ] + + +def test_refuse_overwriting_legacy_frontend(registry, monkeypatch): + monkeypatch.setattr(modal.App, "lookup", lambda *args, **kwargs: object()) + with pytest.raises(ValueError, match="already exists without a YAML registry"): + cli.deploy([deployment()]) + assert "pending" not in registry + + +def test_validate_never_resolves_or_deploys(monkeypatch, capsys): + monkeypatch.setattr( + cli, "compile_configs", lambda *args: pytest.fail("unexpected resolution") + ) + monkeypatch.setattr( + cli, "deploy", lambda *args: pytest.fail("unexpected deployment") + ) + cli.main(["config", "validate", str(preset_path("qwen35-9b-lora-16k"))]) + assert "Validated 1 deployment" in capsys.readouterr().out diff --git a/tests/test_deployments.py b/tests/test_deployments.py new file mode 100644 index 0000000..56c1884 --- /dev/null +++ b/tests/test_deployments.py @@ -0,0 +1,277 @@ +import argparse +import asyncio +from types import SimpleNamespace + +import httpx +import pytest +import yaml + +from lilo.deployments import ( + DeploymentSpec, + load, + preset_path, + resolve, + validate_frontend, +) +from lilo.deployment_cli import retain_generations +from lilo.control_plane.deployments import DeploymentRoutes +from lilo.providers.modal.recipe import backend_config +from lilo.providers.modal.yaml_apps import definition_from_spec +from lilo.native_options import apply_defaults + + +def recipe(preset="qwen35-9b-lora-16k", **changes): + data = load(preset_path(preset)).model_dump() + for path, value in changes.items(): + keys = path.split("__") + target = data + for key in keys[:-1]: + target = target[key] + target[keys[-1]] = value + return DeploymentSpec.model_validate(data) + + +def resolved(spec=None, **changes): + return resolve( + spec or recipe(**changes), revision="a" * 40, implementation="test-runtime" + ) + + +def definition(value): + return definition_from_spec(value, register_trainer=False) + + +def test_presets_context_topology_and_backend_options(): + spec = recipe() + config = backend_config(spec)["miles"] + assert ( + config["actor_num_gpus_per_node"], + config["tensor_model_parallel_size"], + config["max_lora_slots"], + ) == (4, 4, 6) + assert config["native_options"]["recompute_num_layers"] == 1 + assert config["extra_args"] == ("--seq-length", "16384") + large = recipe("qwen35-9b-lora-64k") + assert large.model.max_context_length == 65536 + assert backend_config(large)["miles"]["actor_num_gpus_per_node"] == 8 + fft = backend_config(recipe("qwen35-4b-fft-64k"))["megatron"] + assert (fft["tensor_model_parallel_size"], fft["context_parallel_size"]) == (2, 2) + assert fft["provider_overrides"]["recompute_granularity"] == "full" + + +def test_no_model_catalog_required(): + spec = recipe( + model__id="my-org/new-model", + trainer__miles__model_args=None, + trainer__miles__options={ + "num_layers": 12, + "hidden_size": 768, + "num_attention_heads": 12, + }, + ) + assert definition(resolved(spec)).MODEL_NAME == "my-org/new-model" + assert backend_config(spec)["miles"]["model_type"] == "" + + +def test_extends_false_and_lists_replace(tmp_path): + path = tmp_path / "child.yaml" + path.write_text( + yaml.safe_dump( + { + "extends": "builtin:qwen35-9b-lora-16k", + "routing": {"default": False}, + "trainer": {"miles": {"options": {"target_modules": ["linear_qkv"]}}}, + } + ) + ) + spec = load(path) + assert spec.routing.default is False + assert spec.trainer.miles.options["target_modules"] == ["linear_qkv"] + assert spec.trainer.miles.options["lora_rank"] == 32 + + +def test_duplicate_keys_and_cycles(tmp_path): + path = tmp_path / "bad.yaml" + path.write_text("name: first\nname: second\n") + with pytest.raises(ValueError, match="duplicate"): + load(path) + path.write_text("extends: bad.yaml\n") + with pytest.raises(ValueError, match="cyclic"): + load(path) + + +@pytest.mark.parametrize( + "changes,match", + [ + ({"trainer__miles__options": {"hf_checkpoint": "other"}}, "managed"), + ({"trainer__miles__options": {"pipeline_model_parallel_size": 2}}, "managed"), + ({"inference__sglang__options": {"model_path": "other"}}, "managed"), + ({"inference__sglang__options": {"tp_size": 2}}, "replica GPU"), + ({"trainer__engine__max_clients_per_instance": 7}, "multi_lora_n_adapters"), + ({"trainer__env": {"LILO_BACKEND_CONFIG": "oops"}}, "managed"), + ( + { + "inference__sglang__options": { + "max_loaded_loras": 2, + "max_loras_per_batch": 8, + } + }, + "max_loaded_loras", + ), + ], +) +def test_invalid_integrations_fail_locally(changes, match): + with pytest.raises(ValueError, match=match): + recipe(**changes) + + +def test_generation_and_asset_paths_include_exact_base(): + a = resolved() + assert a.generation == resolved(recipe(routing__default=False)).generation + assert a.generation != resolved(recipe(trainer__resources__gpu="H200:4")).generation + b = resolved(recipe(model__id="other/Qwen3.5-9B-Base")) + assert a.asset_path != b.asset_path + assert ( + a.asset_path + != resolve(a.spec, revision="b" * 40, implementation="test-runtime").asset_path + ) + with pytest.raises(ValueError, match="exact commit"): + resolve(a.spec, revision="main", implementation="test-runtime") + + +def test_frontend_defaults_and_retained_generations(): + small = resolved() + large = resolved(recipe("qwen35-9b-lora-64k")) + routes = DeploymentRoutes([definition(small), definition(large)]) + assert ( + routes.select(small.spec.model.id, "lora").DEFINITION_ID == small.definition_id + ) + switched = retain_generations( + [small, large], [resolved(recipe("qwen35-9b-lora-64k", routing__default=True))] + ) + routes = DeploymentRoutes(map(definition, switched)) + assert ( + routes.select(small.spec.model.id, "lora").DEFINITION_ID == large.definition_id + ) + # Saved model/checkpoint records continue using their original definition. + assert ( + routes.select(small.definition_id, "lora").DEFINITION_ID == small.definition_id + ) + assert routes.capabilities()[0]["max_context_length"] == 65536 + with pytest.raises(ValueError, match="different Lilo/runtime"): + retain_generations( + [small.model_copy(update={"implementation": "old"})], [large] + ) + with pytest.raises(ValueError, match="multiple defaults"): + validate_frontend( + [small.spec, recipe("qwen35-9b-lora-64k", routing__default=True)] + ) + + +def test_ambiguous_model_does_not_get_random_configuration(): + rows = [ + resolved(recipe(routing__default=False)), + resolved(recipe("qwen35-9b-lora-64k")), + ] + routes = DeploymentRoutes(map(definition, rows)) + with pytest.raises(ValueError, match="ambiguous.*16k.*64k"): + routes.select(rows[0].spec.model.id, "lora") + assert routes.capabilities() == [] + + +def test_sampling_requires_default_across_training_modes(): + lora = resolved() + fft = resolved(recipe("qwen35-4b-fft-64k", model__id=lora.spec.model.id)) + routes = DeploymentRoutes(map(definition, [lora, fft])) + with pytest.raises(ValueError, match="sampling_default"): + routes.sampling(lora.spec.model.id) + fft.spec.routing.sampling_default = True + assert ( + DeploymentRoutes(map(definition, [lora, fft])) + .sampling(lora.spec.model.id) + .DEFINITION_ID + == fft.definition_id + ) + + +def test_native_false_list_aliases_and_scalar_overrides(): + parser = argparse.ArgumentParser() + parser.add_argument("--use-feature", action="store_true") + parser.add_argument("--layers", nargs="+", type=int) + parser.add_argument("--tp", "--tensor-parallel-size", dest="tp_size", type=int) + parser.add_argument("--unchanged") + argv = ["--use-feature", "--layers", "1", "2", "--tp=4", "--unchanged", "keep"] + apply_defaults(parser, {"use_feature": False, "layers": [3], "tp_size": 8}, argv) + parsed = parser.parse_args(argv) + assert vars(parsed) == { + "use_feature": False, + "layers": [3], + "tp_size": 8, + "unchanged": "keep", + } + with pytest.raises(ValueError, match="unknown backend option"): + apply_defaults(parser, {"typo": 1}, []) + with pytest.raises(ValueError, match="boolean"): + apply_defaults(parser, {"use_feature": "false"}, []) + with pytest.raises(ValueError, match="list"): + apply_defaults(parser, {"layers": "1,2"}, []) + + +def test_multiple_models_same_http_service_and_old_binding_survives_switch(): + from lilo.control_plane import ControlPlane, create_control_plane_app + from lilo.providers.local import InMemoryKeyValueStore + from lilo.control_plane.keys import model_key + + async def run(): + store = InMemoryKeyValueStore() + plane = ControlPlane(store, SimpleNamespace()) + first = resolved() + other = resolved(recipe(name="other", model__id="org/other-model")) + app = create_control_plane_app( + plane, list(map(definition, [first, other])), api_key="test" + ) + async with httpx.AsyncClient( + transport=httpx.ASGITransport(app=app), + base_url="http://one-frontend", + headers={"x-api-key": "test"}, + ) as client: + session = (await client.post("/api/v1/create_session", json={})).json()[ + "session_id" + ] + for seq, row in enumerate([first, other]): + response = await client.post( + "/api/v1/create_model", + json={ + "session_id": session, + "model_seq_id": seq, + "base_model": row.spec.model.id, + "lora_config": {"rank": 32}, + }, + ) + assert response.status_code == 200, response.text + record = await store.get(model_key(response.json()["model_id"])) + assert record["engine_definition_id"] == row.definition_id + configs = await client.get("/api/v1/lilo/deployments") + assert len(configs.json()["deployments"]) == 2 + client.headers.clear() + assert (await client.get("/api/v1/lilo/deployments")).status_code == 401 + + asyncio.run(run()) + + +def test_native_boolean_opposite_flags_and_optional_value(): + parser = argparse.ArgumentParser() + parser.add_argument("--bias", dest="bias", action="store_true") + parser.add_argument("--no-bias", dest="bias", action="store_false") + parser.add_argument("--optional", nargs="?") + parser.add_argument("--keep", action="store_true") + argv = ["--bias", "--optional", "--keep"] + apply_defaults(parser, {"bias": False, "optional": "supplied"}, argv) + assert vars(parser.parse_args(argv)) == { + "bias": False, + "optional": "supplied", + "keep": True, + } + parser.add_argument("--custom", action="append") + with pytest.raises(ValueError, match="unsupported argparse action"): + apply_defaults(parser, {"custom": [1]}, []) diff --git a/uv.lock b/uv.lock index 2f3031b..6b608a4 100644 --- a/uv.lock +++ b/uv.lock @@ -476,11 +476,13 @@ source = { editable = "." } dependencies = [ { name = "fastapi" }, { name = "httpx" }, + { name = "huggingface-hub" }, { name = "modal" }, { name = "opentelemetry-exporter-otlp-proto-http" }, { name = "opentelemetry-sdk" }, { name = "protobuf" }, { name = "pydantic" }, + { name = "pyyaml" }, { name = "stitch" }, { name = "tinker" }, { name = "uvicorn" }, @@ -497,11 +499,13 @@ dev = [ requires-dist = [ { name = "fastapi", specifier = ">=0.141.1" }, { name = "httpx", specifier = ">=0.28.1" }, + { name = "huggingface-hub", specifier = ">=0.34" }, { name = "modal", specifier = ">=1.5.3" }, { name = "opentelemetry-exporter-otlp-proto-http", specifier = ">=1.39,<2" }, { name = "opentelemetry-sdk", specifier = ">=1.39,<2" }, { name = "protobuf", specifier = ">=5.29" }, { name = "pydantic", specifier = ">=2.13.4" }, + { name = "pyyaml", specifier = ">=6.0.2" }, { name = "stitch", git = "https://github.com/modal-projects/stitch.git?rev=375a9396a7b05770dc4ed9cc5fe34fc4d5a472d5" }, { name = "tinker", specifier = ">=0.24.1,<0.25" }, { name = "uvicorn", specifier = ">=0.52.0" }, From 13d2a310f51c8fdcd5e681040aeab7c986dc194c Mon Sep 17 00:00:00 2001 From: kailash Date: Wed, 23 Sep 2026 06:16:42 +0000 Subject: [PATCH 04/27] Fix YAML runtime deployment and preserve retained trainer functions --- scripts/yaml_deployment_smoke.py | 193 ++++++++++++++++++++++++++ src/lilo/deployment_cli.py | 4 + src/lilo/native_options.py | 4 +- src/lilo/providers/modal/app.py | 4 +- src/lilo/providers/modal/yaml_apps.py | 12 +- tests/providers/test_yaml_apps.py | 16 +++ tests/test_deployment_cli.py | 11 ++ tests/test_deployments.py | 15 ++ 8 files changed, 254 insertions(+), 5 deletions(-) create mode 100644 scripts/yaml_deployment_smoke.py diff --git a/scripts/yaml_deployment_smoke.py b/scripts/yaml_deployment_smoke.py new file mode 100644 index 0000000..2fa851e --- /dev/null +++ b/scripts/yaml_deployment_smoke.py @@ -0,0 +1,193 @@ +"""Small real-GPU training/publication/sampling check for a YAML deployment. + +Run from an authenticated operator environment. Results are written to --output; +credentials are read from TINKER_API_KEY and never included in the report. +""" + +from __future__ import annotations + +import argparse +import json +import math +import os +from pathlib import Path +import time + +import httpx +import modal +import tinker +from tinker import types + +from lilo.client import create_full_training_client +from lilo.providers.modal.kv import app_store_name + + +def write(path, report): + temp = path.with_suffix(".tmp") + temp.write_text(json.dumps(report, indent=2) + "\n") + temp.replace(path) + + +def trainer_record(app_id, model_id): + placement = modal.Dict.from_name(app_store_name("lilo-models", app_id)).get( + f"placement:{model_id}" + ) + if not placement: + raise RuntimeError(f"No trainer placement for {model_id}") + row = modal.Dict.from_name(app_store_name("lilo-engines", app_id)).get( + f"engine_instance:{placement['engine_instance_id']}" + ) + return { + key: row.get(key) + for key in ( + "instance_id", + "definition_id", + "state", + "boot_id", + "call_id", + "revision", + ) + } + + +def step(training, tokens): + datum = types.Datum( + model_input=types.ModelInput.from_ints(tokens[:-1]), + loss_fn_inputs={ + "target_tokens": tokens[1:], + "weights": [1.0] * (len(tokens) - 1), + }, + ) + start = time.monotonic() + result = training.forward_backward([datum], "cross_entropy").result(timeout=3600) + if len(result.loss_fn_outputs[0]["logprobs"].data) != len(tokens) - 1: + raise AssertionError("wrong forward/backward output length") + if not all( + math.isfinite(float(x)) for x in result.loss_fn_outputs[0]["logprobs"].data + ): + raise AssertionError("non-finite trainer logprobs") + train_seconds = time.monotonic() - start + optim = training.optim_step(types.AdamParams(learning_rate=1e-4)).result( + timeout=3600 + ) + sample = training.save_weights_and_get_sampling_client() + published = time.monotonic() + output = sample.sample( + prompt=types.ModelInput.from_ints(tokens[:16]), + num_samples=1, + sampling_params=types.SamplingParams(max_tokens=16, temperature=0.0, seed=42), + ).result(timeout=3600) + sequence = output.sequences[0] + assert len(sequence.tokens) > 0 + assert sequence.logprobs is not None and len(sequence.logprobs) == len( + sequence.tokens + ) + assert all(math.isfinite(float(v)) for v in sequence.logprobs) + return { + "train_seconds": train_seconds, + "step_seconds": time.monotonic() - start, + "sample_seconds": time.monotonic() - published, + "train_metrics": result.metrics, + "optim_metrics": optim.metrics, + "generated_tokens": len(sequence.tokens), + "sample_tokens": sequence.tokens, + "sample_logprobs": sequence.logprobs, + } + + +def run(args): + app = modal.App.lookup(args.frontend) + server = modal.Function.from_name(args.frontend, "server") + url = server.get_web_url() + headers = {"X-API-Key": os.environ["TINKER_API_KEY"]} + rows = modal.Dict.from_name(f"{args.frontend}-yaml-deployments").get("manifest") + row = next(r for r in rows if r["active"] and r["spec"]["name"] == args.name) + definition_id = f"yaml_{args.name}_{row['generation'][:16]}" + report = { + "frontend": args.frontend, + "app_id": app.app_id, + "name": args.name, + "generation": row["generation"], + "spec": row["spec"], + "status": "creating", + "steps": [], + } + output = Path(args.output) + output.parent.mkdir(parents=True, exist_ok=True) + write(output, report) + service = tinker.ServiceClient(base_url=url, api_key=os.environ["TINKER_API_KEY"]) + training = None + try: + start = time.monotonic() + training = ( + service.create_lora_training_client(base_model=definition_id, rank=32) + if row["spec"]["model"]["parameterization"] == "lora" + else create_full_training_client(service, definition_id) + ) + report.update( + model_id=training.model_id, + create_seconds=time.monotonic() - start, + trainer_before=trainer_record(app.app_id, training.model_id), + status="training", + ) + write(output, report) + tokenizer = training.get_tokenizer() + text = "The sum of one and one is two. The sum of two and two is four.\n" * 40 + tokens = tokenizer.encode(text, add_special_tokens=True)[:513] + for index in range(args.steps): + report["steps"].append(step(training, tokens)) + write(output, report) + print(f"{args.name}: completed step {index + 1}", flush=True) + if args.continue_file: + report["status"] = "waiting_for_redeploy" + write(output, report) + deadline = time.monotonic() + 3600 + while not Path(args.continue_file).exists(): + if time.monotonic() > deadline: + raise TimeoutError("redeploy gate did not open") + time.sleep(5) + report["after_redeploy_step"] = step(training, tokens) + report["trainer_after"] = trainer_record(app.app_id, training.model_id) + assert ( + report["trainer_before"]["boot_id"] + == report["trainer_after"]["boot_id"] + ), "unchanged trainer restarted" + report["status"] = "passed" + except BaseException as exc: + report.update(status="failed", error=f"{type(exc).__name__}: {exc}") + raise + finally: + write(output, report) + if training is not None: + try: + with httpx.Client(base_url=url, headers=headers, timeout=60) as client: + response = client.post( + "/api/v1/unload_model", json={"model_id": training.model_id} + ) + response.raise_for_status() + request = response.json()["request_id"] + deadline = time.monotonic() + 180 + while time.monotonic() < deadline: + response = client.post( + "/api/v1/retrieve_future", json={"request_id": request} + ) + if response.status_code != 408: + response.raise_for_status() + break + except Exception as exc: + report["cleanup_error"] = str(exc) + write(output, report) + + +def main(): + parser = argparse.ArgumentParser() + parser.add_argument("--frontend", required=True) + parser.add_argument("--name", required=True) + parser.add_argument("--steps", type=int, default=3) + parser.add_argument("--output", required=True) + parser.add_argument("--continue-file") + run(parser.parse_args()) + + +if __name__ == "__main__": + main() diff --git a/src/lilo/deployment_cli.py b/src/lilo/deployment_cli.py index 50aa23f..4f9fc3f 100644 --- a/src/lilo/deployment_cli.py +++ b/src/lilo/deployment_cli.py @@ -93,6 +93,10 @@ def deploy(desired): import modal from lilo.providers.modal.yaml_apps import MANIFEST_ENV + if sys.version_info[:2] != (3, 12): + raise ValueError( + "YAML deployment requires Python 3.12 to match the serialized GPU runtime images" + ) settings = desired[0].spec.deployment registry = modal.Dict.from_name( f"{settings.frontend}-yaml-deployments", diff --git a/src/lilo/native_options.py b/src/lilo/native_options.py index c250562..e220a71 100644 --- a/src/lilo/native_options.py +++ b/src/lilo/native_options.py @@ -50,7 +50,9 @@ def apply_defaults( raise ValueError(f"backend option {key} requires {action.nargs} values") cast = action.type or (lambda x: x) try: - values = [cast(v) for v in values] + # argparse type callbacks receive command-line text, including + # custom converters such as SGLang human_readable_int. + values = [cast(str(v)) for v in values] except (ValueError, TypeError, argparse.ArgumentTypeError) as exc: raise ValueError( f"invalid value for backend option {key}: {exc}" diff --git a/src/lilo/providers/modal/app.py b/src/lilo/providers/modal/app.py index 781083f..195d41b 100644 --- a/src/lilo/providers/modal/app.py +++ b/src/lilo/providers/modal/app.py @@ -148,12 +148,12 @@ async def _delete_checkpoint(uri: str) -> None: image = ( - modal.Image.debian_slim(python_version="3.11") + modal.Image.debian_slim(python_version="3.12" if YAML_SETTINGS else "3.11") .apt_install("git") .pip_install(*CORE_PACKAGES, TINKER_PACKAGE) .pip_install(STITCH_PACKAGE, "huggingface-hub") - .add_local_python_source("lilo") .env(TRAINER_DEPLOYMENT_ENV) + .add_local_python_source("lilo") ) model_assets = modal.Volume.from_name( YAML_SETTINGS.deployment.storage.assets if YAML_SETTINGS else "lilo-model-assets", diff --git a/src/lilo/providers/modal/yaml_apps.py b/src/lilo/providers/modal/yaml_apps.py index 3cacbc3..18c26e4 100644 --- a/src/lilo/providers/modal/yaml_apps.py +++ b/src/lilo/providers/modal/yaml_apps.py @@ -12,7 +12,7 @@ import modal -from lilo.deployments import ResolvedDeployment, validate_frontend +from lilo.deployments import ResolvedDeployment, Routing, validate_frontend from .recipe import backend_config, serving_options MANIFEST_ENV = "LILO_DEPLOYMENT_MANIFEST" @@ -85,7 +85,15 @@ def secrets_for(spec, *, training=False): def build_trainer_app(resolved: ResolvedDeployment, *, image=None): + # Admission metadata must not change the serialized function for a running + # configuration when a default switches or an older configuration drains. + resolved = resolved.model_copy(deep=True) + resolved.active = True + resolved.spec.routing = Routing() spec = resolved.spec + config_json = json.dumps( + resolved.model_dump(mode="json"), sort_keys=True, separators=(",", ":") + ) app = modal.App(f"{spec.deployment.frontend}-{resolved.definition_id}") resource = spec.trainer.resources from .deployment import trainer_deployment_env @@ -113,7 +121,7 @@ def build_trainer_app(resolved: ResolvedDeployment, *, image=None): env=env, ) def trainer(instance_id: str): - run_trainer(resolved, instance_id) + run_trainer(ResolvedDeployment.model_validate_json(config_json), instance_id) return app, trainer diff --git a/tests/providers/test_yaml_apps.py b/tests/providers/test_yaml_apps.py index 83eb039..88e5180 100644 --- a/tests/providers/test_yaml_apps.py +++ b/tests/providers/test_yaml_apps.py @@ -210,3 +210,19 @@ def test_real_modal_app_constructs_from_manifest_without_legacy_catalog(monkeypa ) assert result.returncode == 0, result.stderr assert "constructed" in result.stdout + + +def test_admission_changes_preserve_serialized_trainer(builders): + from modal._serialization import serialize + + first = deployment() + old_bytes = serialize(yaml_apps.build_trainer_app(first, image="test")[1]) + changed = first.model_copy(deep=True) + changed.active = False + changed.spec.routing.default = not first.spec.routing.default + changed.spec.routing.sampling_default = True + new_bytes = serialize(yaml_apps.build_trainer_app(changed, image="test")[1]) + assert new_bytes == old_bytes + assert first.active is True and first.spec.routing.default is True + changed.spec.trainer.resources.gpu = "H200:4" + assert serialize(yaml_apps.build_trainer_app(changed, image="test")[1]) != old_bytes diff --git a/tests/test_deployment_cli.py b/tests/test_deployment_cli.py index 3bd2301..803b6fa 100644 --- a/tests/test_deployment_cli.py +++ b/tests/test_deployment_cli.py @@ -96,3 +96,14 @@ def test_validate_never_resolves_or_deploys(monkeypatch, capsys): ) cli.main(["config", "validate", str(preset_path("qwen35-9b-lora-16k"))]) assert "Validated 1 deployment" in capsys.readouterr().out + + +def test_deploy_rejects_python_mismatch_before_remote_changes(monkeypatch): + monkeypatch.setattr(cli.sys, "version_info", (3, 11, 0)) + monkeypatch.setattr( + modal.Dict, + "from_name", + lambda *a, **k: pytest.fail("must reject before touching Modal"), + ) + with pytest.raises(ValueError, match="requires Python 3.12"): + cli.deploy([deployment()]) diff --git a/tests/test_deployments.py b/tests/test_deployments.py index 56c1884..0245a66 100644 --- a/tests/test_deployments.py +++ b/tests/test_deployments.py @@ -275,3 +275,18 @@ def test_native_boolean_opposite_flags_and_optional_value(): parser.add_argument("--custom", action="append") with pytest.raises(ValueError, match="unsupported argparse action"): apply_defaults(parser, {"custom": [1]}, []) + + +def test_native_type_callbacks_receive_text(): + def readable_int(value): + return int(value.strip().removesuffix("k")) * ( + 1000 if value.endswith("k") else 1 + ) + + parser = argparse.ArgumentParser() + parser.add_argument("--context-length", type=readable_int) + parser.add_argument("--sizes", nargs="+", type=readable_int) + apply_defaults(parser, {"context_length": 65536, "sizes": [32, "2k"]}, []) + args = parser.parse_args([]) + assert args.context_length == 65536 + assert args.sizes == [32, 2000] From dfa697074fd595b0f0ae1d986a56824240714fe4 Mon Sep 17 00:00:00 2001 From: kailash Date: Wed, 23 Sep 2026 06:44:53 +0000 Subject: [PATCH 05/27] Document GPU YAML validation and shared-app redeploy isolation --- docs/deployment-yaml-design.md | 10 ++-- docs/deployment-yaml-validation.md | 77 ++++++++++++++++++++++++++++++ 2 files changed, 83 insertions(+), 4 deletions(-) create mode 100644 docs/deployment-yaml-validation.md diff --git a/docs/deployment-yaml-design.md b/docs/deployment-yaml-design.md index 8a9b929..f0a993a 100644 --- a/docs/deployment-yaml-design.md +++ b/docs/deployment-yaml-design.md @@ -2,7 +2,7 @@ This draft implements an opt-in YAML path for shared Modal deployments. A file specifies the base model, trainer backend, resources, context length, adapter capacity, and inference settings. Adding a backend-supported model does not require a new Python definition or a catalog entry. -The implementation has CPU tests. No applications have been redeployed and no GPU compatibility or capacity tests have been run for this change. The existing Python deployment and scoped-run paths remain available. +The implementation has CPU tests and live training/sampling checks; see the [validation report](deployment-yaml-validation.md) for tested configurations and shared-app redeploy results. The existing Python deployment and scoped-run paths remain available. ## The provider and Modal app structure @@ -70,7 +70,7 @@ The packaged presets are: - [`qwen35-9b-lora-64k.yaml`](../src/lilo/presets/qwen35-9b-lora-64k.yaml): a larger-context example using H200:8 training. - [`qwen35-4b-fft-64k.yaml`](../src/lilo/presets/qwen35-4b-fft-64k.yaml): the existing 4B FFT topology expressed as YAML. -These are starting configurations. The 16K and FFT backend settings are based on the existing definitions; inference minima are explicitly zero and the trainer maximum is one. The new 64K example has not been GPU-validated. +These are starting configurations. The 16K and FFT backend settings are based on the existing definitions; inference minima are explicitly zero and the trainer maximum is one. All three presets passed short GPU training/sampling checks; see the [validation report](deployment-yaml-validation.md). Full-context memory capacity was not tested. You can instead keep a small override file: @@ -106,6 +106,8 @@ inference: A local `extends: ./base.yaml` also works. Maps merge recursively, lists replace, and `false` overrides `true`. Duplicate YAML keys, unknown Lilo fields and inheritance cycles are rejected. YAML contains secret names; credentials stay in Modal secrets. The API and proxy secret contents are the same as in the [shared deployment setup](../README.md#2-configure-modal-and-secrets-once). +Run `lilo deploy` with Python 3.12, matching the serialized trainer and rollout images. The YAML frontend image also uses Python 3.12 because it launches rollout deployment subprocesses. + When ready to deploy, supply the **complete active set** of files for one frontend: ```bash @@ -154,7 +156,7 @@ Local validation cannot establish memory fit or prove that an unfamiliar model w The CLI stores configurations in the Modal Dict `-yaml-deployments`, scoped to the chosen Modal environment. A single apply lock serializes registry changes. A pending manifest is written before deployment; only successful deployment replaces the committed manifest. The next attempt retains pending configurations too, covering an interruption after Modal accepted a deployment but before the CLI saved its result. -Applying a changed YAML produces a new definition identifier. The hash includes the pinned model revision, normalized settings and implementation fingerprint. Routing preferences are excluded, so switching a default does not change trainer identity. Old configurations are retained as inactive entries. Scaling changes currently also create a new identifier; a separate scaling-policy revision is future work. +Applying a changed YAML produces a new definition identifier. The hash includes the pinned model revision, normalized settings and implementation fingerprint. Routing preferences are excluded, so switching a default does not change trainer identity. Old configurations are retained as inactive entries. Trainer functions capture their settings as consistently ordered JSON. The captured copy uses fixed admission and routing flags, so retiring a configuration or changing a default does not change its trainer function. Scaling changes currently also create a new identifier; a separate scaling-policy revision is future work. The implementation fingerprint includes shipped Lilo source, declared dependencies and the selected Miles commit. The existing image recipes supply the other backend source revisions. This is not a fully pinned Python/container dependency lock. This draft rejects applies that would rebuild retained configurations with a different implementation fingerprint or different shared storage/lifecycle settings; use a separate frontend for those upgrades. Automatic pruning of historical configurations and migration across code/image versions are not implemented. @@ -174,4 +176,4 @@ A killed CLI can leave its apply lock behind. Confirm that the original apply ha The implemented YAML path supports shared, single-node Miles LoRA and Megatron FFT deployments with the existing runtime images. The inference GPU allocation must match tensor parallelism. Scoped `lilo.run(config=...)`, DP-attention layouts, custom image selection, automatic provisioning of unknown `base_model` values, automatic runtime upgrades, and GPU compatibility probes remain follow-up work. Unsupported schema choices are rejected rather than treated as implemented features. -CPU coverage exercises configuration loading and validation, typed native overrides, routing several models through one HTTP service, ambiguity handling, preserved client definitions, interrupted/concurrent applies, trainer resources and executor settings, LoRA/FFT pool startup and shutdown, and startup-error handling. Existing backend, provider, HTTP and scoped-run tests also run. GPU smoke tests are still required before recommending this draft for production deployments. +CPU coverage exercises configuration loading and validation, typed native overrides, routing several models through one HTTP service, ambiguity handling, preserved client definitions, interrupted/concurrent applies, trainer resources and executor settings, LoRA/FFT pool startup and shutdown, and startup-error handling. Existing backend, provider, HTTP and scoped-run tests also run. The [GPU validation report](deployment-yaml-validation.md) covers short training/sampling requests and shared-app redeploy continuity. Maximum-context capacity and sustained-load testing remain separate checks. diff --git a/docs/deployment-yaml-validation.md b/docs/deployment-yaml-validation.md new file mode 100644 index 0000000..3cce0b9 --- /dev/null +++ b/docs/deployment-yaml-validation.md @@ -0,0 +1,77 @@ +# YAML deployment validation + +These checks exercise the shared-app deployment path in PR #55. Each deployment contains the frontend and generated trainer functions; inference pools are separate apps started on demand. + +## GPU results + +All checks passed on 2026-09-23. A step includes forward/backward, an optimizer update, publication of sampling weights, and a sample with a 16-token cap. + +| Configuration | Region | Completed steps | Result | +| --- | --- | ---: | --- | +| 9B LoRA, 16K preset | us-west | 3 before deployment updates + 1 after | Passed; original trainer and inference container survived | +| 9B LoRA, changed 16K YAML | us-west | 3 on a fresh client | Passed; client selected the new configuration | +| 4B FFT, 64K preset | us-west | 4 across deployment updates | Passed; original trainer survived | +| 9B LoRA, 64K preset, 8×H200 | us-east | 3 | Passed; functional check in a separate frontend | + +Every step returned 512 finite training log probabilities and 16 generated tokens with finite log probabilities. The test sends 512 training tokens and a 16-token sampling prompt; it does **not** fill a 16K or 64K context. It tests the configured processes and interfaces, not maximum-context capacity or learning quality. + +All completed smoke-test clients unloaded without reported cleanup errors. The dedicated test frontends and their inference pools were stopped after validation. + +## Deployment isolation checked on 2026-09-23 + +Source commit: `13d2a31`. CPU suite: **554 passed, 1 skipped**. The isolated shared app was `lilo-yaml-pr55-smoke-v3` in `modal-labs/kailash-dev`, initially deployed with all three packaged presets and pinned model revisions. + +Modal reported 7.314 seconds for an unchanged deployment. A second apply changed only `qwen35-9b-lora-16k`'s `inference.sglang.options.max_running_requests` from 32 to 24; Modal reported 4.138 seconds for that deployment. The shared frontend was redeployed on both applies. + +Both applies preserved the running 16K LoRA and FFT trainer boot IDs. Existing inference app IDs were unchanged, and the running LoRA inference container survived both applies. The original LoRA client completed three initial steps and another complete training/publication/sampling step after both applies, on the same trainer boot ID. + +The changed YAML received a new definition; the previous definition was retained as inactive for existing clients. The unchanged FFT and 64K definitions retained their function IDs. The 64K trainer was still queued during these applies, so these observations do not establish continuity of an already-running 8×H200 trainer. + +The FFT client also completed four training/publication/sampling steps on its original trainer boot ID. Its first sample was waiting for inference startup during the deployments, then completed successfully. A fresh 16K LoRA client selected the new configuration and passed three steps with the 24-request inference limit recorded in its resolved configuration. + +These results establish continuity of these existing jobs. They do not mean `lilo deploy` skips deploying the shared app, or that an unrelated frontend request can never be retried during deployment. + +## Runtime revisions and hardware + +| Preset | Model revision | Trainer GPU request | Inference GPU request per replica | +| --- | --- | --- | --- | +| `qwen35-9b-lora-16k` | `68c46c4b3498877f3ef123c856ecfde50c39f404` | 4×H100, TP4 | H200 | +| `qwen35-9b-lora-64k` | `68c46c4b3498877f3ef123c856ecfde50c39f404` | 8×H200, TP8 | H200 | +| `qwen35-4b-fft-64k` | `851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a` | 4×H100, TP2/CP2 | H100 | + +Miles was pinned to `f6d0257d83b7c14f4ce43ecfcd71d955112f6e0e`. GPU types in the table are requests: Modal can fulfill an H100 request with H200 hardware, which occurred during this test. Inference retained the presets' zero minimum and eight maximum replicas; these small requests did not exercise scaling to eight replicas. + +The western 8×H200 trainer remained queued without a container for approximately 20 minutes. Its session and queued call were cancelled before retrying the same hardware/model configuration in a separate `us-east` test frontend. That retry passed three steps and is a functional check of the 64K configuration, separate from the western shared-app isolation test. + +## Reproduce the smoke test + +Use Python 3.12 and an authenticated Modal environment. Set `TINKER_API_KEY` locally to the value in the deployment's API secret. Generate the preset YAMLs, give them the same isolated `deployment.frontend`, and pin model revisions plus `LILO_MILES_COMMIT` before deploying. + +```bash +lilo deploy qwen35-9b-lora-16k.yaml qwen35-9b-lora-64k.yaml qwen35-4b-fft-64k.yaml +python scripts/yaml_deployment_smoke.py \ + --frontend YOUR_TEST_FRONTEND \ + --name qwen35-9b-lora-16k \ + --steps 3 \ + --output /tmp/lora16-smoke.json \ + --continue-file /tmp/continue-after-redeploy +``` + +Run the script once per configuration. Each step performs cross-entropy forward/backward on 512 tokens, an Adam update, publication of sampling weights, and generation with a 16-token cap. It checks output lengths and finite trainer/sampler log probabilities. Miles does not support per-client seeds, so the test does not request one. + +The optional `--continue-file` keeps the client alive after its initial steps. Apply the complete unchanged YAML set, record trainer boot IDs, then change one YAML and apply the complete set again. Create the file only after both applies finish: + +```bash +touch /tmp/continue-after-redeploy +``` + +Each waiting client runs another training/publication/sampling step and checks that its trainer boot ID is unchanged. Run a fresh client against the changed configuration as well. Inspect inference app IDs and container IDs before and after applies to check pool continuity. The script unloads its model on exit; stop the dedicated frontend and its pools after completing the test. + +This checks short-request functionality and continuity across deployment updates. It does not establish maximum-context memory fit, convergence, or performance under sustained load. The JSON reports are local test artifacts and are not committed to the documentation. + +## Bugs found during live validation + +- The frontend image applied environment settings after `add_local_python_source`; Modal rejected the image build. Environment settings now precede local source mounting. +- The YAML frontend used Python 3.11 to launch rollout deployment subprocesses while serialized GPU images used Python 3.12. The YAML frontend and deployment CLI now require the matching Python version. +- SGLang's integer parser expects strings and calls `.strip()`. YAML integers now go through backend type converters as command-line text would. +- Retiring a configuration changed data captured by its trainer function. The function now captures consistently ordered JSON with fixed admission/routing flags. A regression test checks that retiring a configuration preserves the serialized function and changing GPU resources changes it. From 89e3d4850e295c3c9b14123770184459be806f7e Mon Sep 17 00:00:00 2001 From: kailash Date: Wed, 23 Sep 2026 16:26:34 +0000 Subject: [PATCH 06/27] Add an explicit YAML deployment list and deployment script --- README.md | 2 +- deployments/qwen35-4b-fft-64k.yaml | 3 +++ deployments/qwen35-9b-lora-16k.yaml | 3 +++ deployments/qwen35-9b-lora-64k.yaml | 3 +++ docs/deployment-yaml-design.md | 9 +++++++++ scripts/deploy_models.sh | 25 +++++++++++++++++++++++++ 6 files changed, 44 insertions(+), 1 deletion(-) create mode 100644 deployments/qwen35-4b-fft-64k.yaml create mode 100644 deployments/qwen35-9b-lora-16k.yaml create mode 100644 deployments/qwen35-9b-lora-64k.yaml create mode 100755 scripts/deploy_models.sh diff --git a/README.md b/README.md index 0bd7ee4..401e1f8 100644 --- a/README.md +++ b/README.md @@ -48,7 +48,7 @@ training = service.create_lora_training_client( ## Shared deployment quick start -For the draft YAML-based provider path, see [YAML deployments](docs/deployment-yaml-design.md). +For the draft YAML-based provider path, see [YAML deployments](docs/deployment-yaml-design.md). Keep the active YAML list in [scripts/deploy_models.sh](scripts/deploy_models.sh); run it with `--check` to validate locally, or without arguments to deploy. Install Lilo into your own Python project, deploy it once to Modal, then call its API from your training scripts. The commands below work in Bash or Zsh. diff --git a/deployments/qwen35-4b-fft-64k.yaml b/deployments/qwen35-4b-fft-64k.yaml new file mode 100644 index 0000000..bb6102c --- /dev/null +++ b/deployments/qwen35-4b-fft-64k.yaml @@ -0,0 +1,3 @@ +extends: builtin:qwen35-4b-fft-64k +model: + revision: 851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a diff --git a/deployments/qwen35-9b-lora-16k.yaml b/deployments/qwen35-9b-lora-16k.yaml new file mode 100644 index 0000000..66912d9 --- /dev/null +++ b/deployments/qwen35-9b-lora-16k.yaml @@ -0,0 +1,3 @@ +extends: builtin:qwen35-9b-lora-16k +model: + revision: 68c46c4b3498877f3ef123c856ecfde50c39f404 diff --git a/deployments/qwen35-9b-lora-64k.yaml b/deployments/qwen35-9b-lora-64k.yaml new file mode 100644 index 0000000..b152564 --- /dev/null +++ b/deployments/qwen35-9b-lora-64k.yaml @@ -0,0 +1,3 @@ +extends: builtin:qwen35-9b-lora-64k +model: + revision: 68c46c4b3498877f3ef123c856ecfde50c39f404 diff --git a/docs/deployment-yaml-design.md b/docs/deployment-yaml-design.md index f0a993a..3cf6c97 100644 --- a/docs/deployment-yaml-design.md +++ b/docs/deployment-yaml-design.md @@ -114,6 +114,15 @@ When ready to deploy, supply the **complete active set** of files for one fronte lilo deploy model-a.yaml model-b.yaml model-a-64k.yaml ``` +For a checked-in deployment list, use [`scripts/deploy_models.sh`](../scripts/deploy_models.sh). Its `deployment_files` array lists the editable YAMLs under [`deployments/`](../deployments), initially covering the three presets above with the model revisions used in GPU validation. These inherit the `lilo-yaml` frontend and the presets' shared settings. + +```bash +./scripts/deploy_models.sh --check # Validate the full list locally. +./scripts/deploy_models.sh # Deploy the full list. +``` + +To add a model, create its YAML in `deployments/`, add its path to `deployment_files`, then run the script. The YAMLs must agree on the frontend and shared settings. Keep existing entries to keep those configurations available to new clients; removing an entry retires it from new-client selection on the next deployment. The script works from any working directory and uses `lilo` from your active Python 3.12 environment. Keep the same pinned `LILO_MILES_COMMIT` across applies, as with the direct CLI. + This command builds/deploys the shared app; it is not a validation command. All files must agree on frontend, Modal environment/region, secrets, storage and shared lifecycle settings. The preset frontend name is `lilo-yaml`. The CLI refuses to overwrite a pre-existing application without a YAML registry, so migration of a legacy frontend must be handled separately. Trainer minimum capacity is zero. Rollout apps are created on first demand, and their configured inference minimum applies once the pool exists. A pool with a nonzero minimum will keep that many workers warm until it is stopped by the existing idle cleanup. diff --git a/scripts/deploy_models.sh b/scripts/deploy_models.sh new file mode 100755 index 0000000..263b46f --- /dev/null +++ b/scripts/deploy_models.sh @@ -0,0 +1,25 @@ +#!/usr/bin/env bash +# The complete set of configurations available to new clients on this frontend. +# Add a YAML under deployments/ and add its path to this list. +set -euo pipefail + +repo_root="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")/.." && pwd)" +cd -- "$repo_root" + +deployment_files=( + deployments/qwen35-9b-lora-16k.yaml + deployments/qwen35-9b-lora-64k.yaml + deployments/qwen35-4b-fft-64k.yaml +) + +case "${1-}" in + "") command_args=(deploy) ;; + --check) command_args=(config validate) ;; + *) echo "Usage: $0 [--check]" >&2; exit 2 ;; +esac +if (( $# > 1 )); then + echo "Usage: $0 [--check]" >&2 + exit 2 +fi + +exec lilo "${command_args[@]}" "${deployment_files[@]}" From d7b4c7430af169316edc403fc930e9f072deea1c Mon Sep 17 00:00:00 2001 From: kailash Date: Wed, 23 Sep 2026 16:48:02 +0000 Subject: [PATCH 07/27] Simplify the deployment script to a YAML list and deploy command --- README.md | 2 +- docs/deployment-yaml-design.md | 3 +-- scripts/deploy_models.sh | 22 ++++++---------------- 3 files changed, 8 insertions(+), 19 deletions(-) diff --git a/README.md b/README.md index 401e1f8..dca9f55 100644 --- a/README.md +++ b/README.md @@ -48,7 +48,7 @@ training = service.create_lora_training_client( ## Shared deployment quick start -For the draft YAML-based provider path, see [YAML deployments](docs/deployment-yaml-design.md). Keep the active YAML list in [scripts/deploy_models.sh](scripts/deploy_models.sh); run it with `--check` to validate locally, or without arguments to deploy. +For the draft YAML-based provider path, see [YAML deployments](docs/deployment-yaml-design.md). Keep the active YAML list in [scripts/deploy_models.sh](scripts/deploy_models.sh); run it to deploy the complete list. Install Lilo into your own Python project, deploy it once to Modal, then call its API from your training scripts. The commands below work in Bash or Zsh. diff --git a/docs/deployment-yaml-design.md b/docs/deployment-yaml-design.md index 3cf6c97..64811e5 100644 --- a/docs/deployment-yaml-design.md +++ b/docs/deployment-yaml-design.md @@ -117,8 +117,7 @@ lilo deploy model-a.yaml model-b.yaml model-a-64k.yaml For a checked-in deployment list, use [`scripts/deploy_models.sh`](../scripts/deploy_models.sh). Its `deployment_files` array lists the editable YAMLs under [`deployments/`](../deployments), initially covering the three presets above with the model revisions used in GPU validation. These inherit the `lilo-yaml` frontend and the presets' shared settings. ```bash -./scripts/deploy_models.sh --check # Validate the full list locally. -./scripts/deploy_models.sh # Deploy the full list. +./scripts/deploy_models.sh ``` To add a model, create its YAML in `deployments/`, add its path to `deployment_files`, then run the script. The YAMLs must agree on the frontend and shared settings. Keep existing entries to keep those configurations available to new clients; removing an entry retires it from new-client selection on the next deployment. The script works from any working directory and uses `lilo` from your active Python 3.12 environment. Keep the same pinned `LILO_MILES_COMMIT` across applies, as with the direct CLI. diff --git a/scripts/deploy_models.sh b/scripts/deploy_models.sh index 263b46f..3b8218a 100755 --- a/scripts/deploy_models.sh +++ b/scripts/deploy_models.sh @@ -1,25 +1,15 @@ #!/usr/bin/env bash -# The complete set of configurations available to new clients on this frontend. -# Add a YAML under deployments/ and add its path to this list. -set -euo pipefail +set -e -repo_root="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")/.." && pwd)" -cd -- "$repo_root" +# Run from the repository root so the YAML paths below work from any directory. +cd "$(dirname "$0")/.." +# Add a model by creating its YAML in deployments/ and adding it here. +# Keep every configuration that should be available to new clients in this list. deployment_files=( deployments/qwen35-9b-lora-16k.yaml deployments/qwen35-9b-lora-64k.yaml deployments/qwen35-4b-fft-64k.yaml ) -case "${1-}" in - "") command_args=(deploy) ;; - --check) command_args=(config validate) ;; - *) echo "Usage: $0 [--check]" >&2; exit 2 ;; -esac -if (( $# > 1 )); then - echo "Usage: $0 [--check]" >&2 - exit 2 -fi - -exec lilo "${command_args[@]}" "${deployment_files[@]}" +lilo deploy "${deployment_files[@]}" From 964c0192dc1d53b05ca78004cae729ff9f990a0b Mon Sep 17 00:00:00 2001 From: kailash Date: Wed, 23 Sep 2026 17:04:52 +0000 Subject: [PATCH 08/27] Make YAML the only shared deployment definition path --- README.md | 20 +- docs/deployment-yaml-design.md | 12 +- docs/deployment-yaml-validation.md | 8 + docs/design.md | 45 ++--- docs/observability.md | 2 +- docs/profiling.md | 2 +- scripts/e2e_engine_definition.py | 88 +++++---- src/lilo/control_plane/deployments.py | 7 +- src/lilo/deployment_cli.py | 2 +- src/lilo/presets/qwen35-35b-a3b-fft-64k.yaml | 63 +++++++ src/lilo/presets/qwen35-9b-fft-64k.yaml | 52 ++++++ .../qwen35-9b-instruct-lora-16k-dp2.yaml | 8 + .../presets/qwen35-9b-instruct-lora-16k.yaml | 59 ++++++ .../presets/qwen35-9b-lora-16k-single.yaml | 7 + src/lilo/presets/qwen35-9b-lora-2k.yaml | 61 ++++++ src/lilo/presets/qwen36-27b-fft-64k.yaml | 54 ++++++ src/lilo/presets/qwen36-35b-a3b-fft-64k.yaml | 9 + src/lilo/presets/qwen38-27b-lora-128k.yaml | 22 +++ src/lilo/presets/qwen38-27b-lora-16k.yaml | 59 ++++++ src/lilo/presets/qwen38-27b-lora-64k.yaml | 17 ++ src/lilo/providers/modal/app.py | 123 ++++--------- .../providers/modal/definitions/__init__.py | 0 .../definitions/qwen3_5_35b_a3b_full_64k.py | 152 --------------- .../modal/definitions/qwen3_5_4b_full_64k.py | 144 --------------- .../qwen3_5_9b_base_miles_lora_16k.py | 169 ----------------- .../qwen3_5_9b_base_miles_lora_16k_single.py | 54 ------ .../qwen3_5_9b_base_miles_lora_2k.py | 155 ---------------- .../modal/definitions/qwen3_5_9b_full_64k.py | 144 --------------- .../definitions/qwen3_5_9b_miles_lora_16k.py | 170 ----------------- .../qwen3_5_9b_miles_lora_16k_dp2.py | 170 ----------------- .../modal/definitions/qwen3_6_27b_full_64k.py | 143 -------------- .../definitions/qwen3_6_35b_a3b_full_64k.py | 152 --------------- .../qwen3_8_27b_miles_lora_128k.py | 174 ------------------ .../definitions/qwen3_8_27b_miles_lora_16k.py | 172 ----------------- .../definitions/qwen3_8_27b_miles_lora_64k.py | 174 ------------------ src/lilo/providers/modal/deployment.py | 17 -- src/lilo/providers/modal/fft_pool.py | 4 +- src/lilo/providers/modal/fft_pool_app.py | 144 --------------- src/lilo/providers/modal/lora_pool.py | 47 +---- src/lilo/providers/modal/lora_pool_app.py | 109 ----------- src/lilo/providers/modal/recipe.py | 14 +- src/lilo/providers/modal/yaml_apps.py | 22 +-- tests/control_plane/test_http.py | 5 +- tests/providers/conftest.py | 18 ++ tests/providers/test_checkpoint_storage.py | 80 ++------ tests/providers/test_deployment_presets.py | 88 +++++++++ tests/providers/test_lora_pool.py | 35 +--- tests/providers/test_miles_definition.py | 56 ------ tests/providers/test_modal_app.py | 39 +++- tests/providers/test_modal_deployment.py | 39 +--- tests/providers/test_qwen36_definitions.py | 33 ---- tests/providers/test_yaml_apps.py | 67 ++++++- tests/providers/test_yaml_e2e_helper.py | 29 +++ 53 files changed, 827 insertions(+), 2712 deletions(-) create mode 100644 src/lilo/presets/qwen35-35b-a3b-fft-64k.yaml create mode 100644 src/lilo/presets/qwen35-9b-fft-64k.yaml create mode 100644 src/lilo/presets/qwen35-9b-instruct-lora-16k-dp2.yaml create mode 100644 src/lilo/presets/qwen35-9b-instruct-lora-16k.yaml create mode 100644 src/lilo/presets/qwen35-9b-lora-16k-single.yaml create mode 100644 src/lilo/presets/qwen35-9b-lora-2k.yaml create mode 100644 src/lilo/presets/qwen36-27b-fft-64k.yaml create mode 100644 src/lilo/presets/qwen36-35b-a3b-fft-64k.yaml create mode 100644 src/lilo/presets/qwen38-27b-lora-128k.yaml create mode 100644 src/lilo/presets/qwen38-27b-lora-16k.yaml create mode 100644 src/lilo/presets/qwen38-27b-lora-64k.yaml delete mode 100644 src/lilo/providers/modal/definitions/__init__.py delete mode 100644 src/lilo/providers/modal/definitions/qwen3_5_35b_a3b_full_64k.py delete mode 100644 src/lilo/providers/modal/definitions/qwen3_5_4b_full_64k.py delete mode 100644 src/lilo/providers/modal/definitions/qwen3_5_9b_base_miles_lora_16k.py delete mode 100644 src/lilo/providers/modal/definitions/qwen3_5_9b_base_miles_lora_16k_single.py delete mode 100644 src/lilo/providers/modal/definitions/qwen3_5_9b_base_miles_lora_2k.py delete mode 100644 src/lilo/providers/modal/definitions/qwen3_5_9b_full_64k.py delete mode 100644 src/lilo/providers/modal/definitions/qwen3_5_9b_miles_lora_16k.py delete mode 100644 src/lilo/providers/modal/definitions/qwen3_5_9b_miles_lora_16k_dp2.py delete mode 100644 src/lilo/providers/modal/definitions/qwen3_6_27b_full_64k.py delete mode 100644 src/lilo/providers/modal/definitions/qwen3_6_35b_a3b_full_64k.py delete mode 100644 src/lilo/providers/modal/definitions/qwen3_8_27b_miles_lora_128k.py delete mode 100644 src/lilo/providers/modal/definitions/qwen3_8_27b_miles_lora_16k.py delete mode 100644 src/lilo/providers/modal/definitions/qwen3_8_27b_miles_lora_64k.py delete mode 100644 src/lilo/providers/modal/fft_pool_app.py delete mode 100644 src/lilo/providers/modal/lora_pool_app.py create mode 100644 tests/providers/conftest.py create mode 100644 tests/providers/test_deployment_presets.py delete mode 100644 tests/providers/test_miles_definition.py delete mode 100644 tests/providers/test_qwen36_definitions.py create mode 100644 tests/providers/test_yaml_e2e_helper.py diff --git a/README.md b/README.md index dca9f55..245981d 100644 --- a/README.md +++ b/README.md @@ -48,7 +48,7 @@ training = service.create_lora_training_client( ## Shared deployment quick start -For the draft YAML-based provider path, see [YAML deployments](docs/deployment-yaml-design.md). Keep the active YAML list in [scripts/deploy_models.sh](scripts/deploy_models.sh); run it to deploy the complete list. +Shared deployments are defined in YAML. See [YAML deployments](docs/deployment-yaml-design.md). Keep the active YAML list in [scripts/deploy_models.sh](scripts/deploy_models.sh); run it to deploy the complete list. Install Lilo into your own Python project, deploy it once to Modal, then call its API from your training scripts. The commands below work in Bash or Zsh. @@ -62,7 +62,7 @@ clients do not need Modal deployment credentials or sampler proxy tokens. With [uv](https://docs.astral.sh/uv/) installed: ```bash -uv init my-lilo-project +uv init --python 3.12 my-lilo-project cd my-lilo-project uv add 'lilo @ git+https://github.com/modal-projects/lilo.git' ``` @@ -131,15 +131,17 @@ uv run modal secret create lilo-proxy \ ### 3. Deploy the installed package -Deploying the entire Tinker server can be done with a single modal deploy command: +Create a configuration from a preset, validate it, and deploy it with Python 3.12: ```bash -uv run modal deploy -m lilo.providers.modal.app +uv run lilo config init --preset qwen35-9b-lora-16k > deployment.yaml +uv run lilo config validate deployment.yaml +uv run lilo deploy deployment.yaml ``` -This deploys the control plane and bundled model definitions, then prints the -`server` URL to use in step 4. Reuse the deployment across training runs and -redeploy after updating Lilo. +This deploys the shared app and prints its `server` URL. Add more YAML files to the same command to serve more configurations. Always supply the complete active set. Pin model revisions and `LILO_MILES_COMMIT` for repeatable applies; see [YAML deployments](docs/deployment-yaml-design.md). + +From a repository checkout, maintain the list in `scripts/deploy_models.sh` and run that script. `lilo deploy` supplies the saved configuration to Modal; importing the shared app directly without a manifest is no longer a deployment entrypoint. Deploying the server doesn't allocate any GPUs; rather, this allocation for both the training and sampling sides are done on demand. See [cold starts and capacity configuration](docs/full-fine-tunes.md#performance-and-behavior-considerations) before running a larger workload. @@ -154,8 +156,8 @@ has finished in the Modal dashboard or list apps with: uv run modal app list ``` -To tear down the deployment, stop its `lilo-fft-...` sampler apps, then `lilo`, -using `uv run modal app stop `. Stopping `lilo` does not stop sampler apps. +To tear down the deployment, stop its `lilo-fft-...` sampler apps, then the frontend named in your YAML (`lilo-yaml` by default), +using `uv run modal app stop `. Stopping the frontend does not stop sampler apps. ## Next steps diff --git a/docs/deployment-yaml-design.md b/docs/deployment-yaml-design.md index 64811e5..a49023f 100644 --- a/docs/deployment-yaml-design.md +++ b/docs/deployment-yaml-design.md @@ -1,8 +1,8 @@ # YAML deployments -This draft implements an opt-in YAML path for shared Modal deployments. A file specifies the base model, trainer backend, resources, context length, adapter capacity, and inference settings. Adding a backend-supported model does not require a new Python definition or a catalog entry. +Shared Modal deployments are defined through YAML. A file specifies the base model, trainer backend, resources, context length, adapter capacity, and inference settings. Adding a backend-supported model does not require a new Python definition or a catalog entry. -The implementation has CPU tests and live training/sampling checks; see the [validation report](deployment-yaml-validation.md) for tested configurations and shared-app redeploy results. The existing Python deployment and scoped-run paths remain available. +The implementation has CPU tests and live training/sampling checks; see the [validation report](deployment-yaml-validation.md) for tested configurations and shared-app redeploy results. The separate scoped `lilo.run(...)` API remains available. ## The provider and Modal app structure @@ -33,7 +33,7 @@ flowchart TD Sidecar --> SGLang[SGLang on port 8001] ``` -`app.py` reads a resolved manifest from `LILO_DEPLOYMENT_MANIFEST`. For each configuration it calls `definition_from_spec`, which constructs the routing metadata and a trainer app. `app.include` places those trainer functions inside the shared frontend app. With no manifest, `app.py` uses the existing Python definitions. +`app.py` reads a resolved manifest from `LILO_DEPLOYMENT_MANIFEST`. For each configuration it calls `definition_from_spec`, which constructs the routing metadata and a trainer app. `app.include` places those trainer functions inside the shared frontend app. A missing manifest is an error with instructions to use `lilo deploy`; there is no Python model-catalog fallback. The trainer function's GPU type/count, CPU, RAM, timeout, maximum instances, secrets and mounted volumes come from YAML. Containers remain single-use. `run_trainer` reloads the prepared asset volume, constructs the backend configuration, and calls the existing `run_engine_with_backend` launcher. Miles uses one controller process that manages its GPU workers; FFT launches one process per allocated GPU. Client admission and sampler-persistence concurrency are configured separately. @@ -70,6 +70,8 @@ The packaged presets are: - [`qwen35-9b-lora-64k.yaml`](../src/lilo/presets/qwen35-9b-lora-64k.yaml): a larger-context example using H200:8 training. - [`qwen35-4b-fft-64k.yaml`](../src/lilo/presets/qwen35-4b-fft-64k.yaml): the existing 4B FFT topology expressed as YAML. +Additional packaged YAML presets preserve the earlier Python recipes for 2K and single-client LoRA, the 9B instruct model, Qwen3.5/3.6 FFT variants, and Qwen3.8 LoRA at 16K/64K/128K. They are optional templates, not automatically deployed models. The deployment script still lists only the three configurations above. Migrated presets use the YAML defaults of at most one trainer and zero-to-eight inference replicas; raise these limits in your YAML when needed. Their model/context/parallelism settings are covered by CPU tests, not new GPU validation. + These are starting configurations. The 16K and FFT backend settings are based on the existing definitions; inference minima are explicitly zero and the trainer maximum is one. All three presets passed short GPU training/sampling checks; see the [validation report](deployment-yaml-validation.md). Full-context memory capacity was not tested. You can instead keep a small override file: @@ -122,7 +124,7 @@ For a checked-in deployment list, use [`scripts/deploy_models.sh`](../scripts/de To add a model, create its YAML in `deployments/`, add its path to `deployment_files`, then run the script. The YAMLs must agree on the frontend and shared settings. Keep existing entries to keep those configurations available to new clients; removing an entry retires it from new-client selection on the next deployment. The script works from any working directory and uses `lilo` from your active Python 3.12 environment. Keep the same pinned `LILO_MILES_COMMIT` across applies, as with the direct CLI. -This command builds/deploys the shared app; it is not a validation command. All files must agree on frontend, Modal environment/region, secrets, storage and shared lifecycle settings. The preset frontend name is `lilo-yaml`. The CLI refuses to overwrite a pre-existing application without a YAML registry, so migration of a legacy frontend must be handled separately. +This command builds/deploys the shared app; it is not a validation command. All files must agree on frontend, Modal environment/region, secrets, storage and shared lifecycle settings. The preset frontend name is `lilo-yaml`. The CLI refuses to overwrite a pre-existing application without a YAML registry, so it cannot accidentally replace an unrelated app. Trainer minimum capacity is zero. Rollout apps are created on first demand, and their configured inference minimum applies once the pool exists. A pool with a nonzero minimum will keep that many workers warm until it is stopped by the existing idle cleanup. @@ -182,6 +184,6 @@ A killed CLI can leave its apply lock behind. Confirm that the original apply ha ## Current scope and validation -The implemented YAML path supports shared, single-node Miles LoRA and Megatron FFT deployments with the existing runtime images. The inference GPU allocation must match tensor parallelism. Scoped `lilo.run(config=...)`, DP-attention layouts, custom image selection, automatic provisioning of unknown `base_model` values, automatic runtime upgrades, and GPU compatibility probes remain follow-up work. Unsupported schema choices are rejected rather than treated as implemented features. +The implemented YAML path supports shared, single-node Miles LoRA and Megatron FFT deployments with the existing runtime images. The inference GPU allocation must match SGLang’s total `tp_size`. Data-parallel attention uses explicit `dp_size` and `enable_dp_attention` options, with `dp_size` dividing the allocation. Scoped `lilo.run(config=...)`, custom image selection, automatic provisioning of unknown `base_model` values, automatic runtime upgrades, and GPU compatibility probes remain follow-up work. Unsupported schema choices are rejected rather than treated as implemented features. CPU coverage exercises configuration loading and validation, typed native overrides, routing several models through one HTTP service, ambiguity handling, preserved client definitions, interrupted/concurrent applies, trainer resources and executor settings, LoRA/FFT pool startup and shutdown, and startup-error handling. Existing backend, provider, HTTP and scoped-run tests also run. The [GPU validation report](deployment-yaml-validation.md) covers short training/sampling requests and shared-app redeploy continuity. Maximum-context capacity and sustained-load testing remain separate checks. diff --git a/docs/deployment-yaml-validation.md b/docs/deployment-yaml-validation.md index 3cce0b9..a288748 100644 --- a/docs/deployment-yaml-validation.md +++ b/docs/deployment-yaml-validation.md @@ -2,6 +2,14 @@ These checks exercise the shared-app deployment path in PR #55. Each deployment contains the frontend and generated trainer functions; inference pools are separate apps started on demand. +## YAML-only deployment cleanup + +After the live checks below, the shared deployment's Python-catalog fallback and model-specific trainer/pool modules were removed. Existing recipes are available as YAML presets. The deployment script still selects only its three listed configurations; additional presets are opt-in. + +Validation of this cleanup: **575 CPU tests passed, 1 skipped**; Ruff and whitespace checks passed. A wheel built with all 14 YAML presets and none of the removed Python definition modules or pool launchers. Coverage includes required manifests, configuration-ID sampling, both generic executors, custom checkpoint storage, generic pool deployment, migrated parallelism settings, and the deployed-configuration E2E helper. + +No apps were redeployed for this cleanup. The GPU results below apply to the earlier source commit `13d2a31`; the additional migrated presets and DP-attention settings have CPU coverage only. + ## GPU results All checks passed on 2026-09-23. A step includes forward/backward, an optimizer update, publication of sampling weights, and a sample with a 16-token cap. diff --git a/docs/design.md b/docs/design.md index 6070b53..81647f6 100644 --- a/docs/design.md +++ b/docs/design.md @@ -122,51 +122,34 @@ sampling scales according to rollout traffic. We best-effort sticky-route groups - [`providers/modal/fft_pool.py`](../src/lilo/providers/modal/fft_pool.py): creates, finds, wakes, and stops each FFT model's sampling service using Modal flash proxy -## Adding a new model deployment +## Adding a new model deployment -"Model deployment" in this context refers to a particular training configuration for a base model, defined by its parameterization (full parameter, LoRA, etc.), desired context/sampling length, parallelism, quantization, and so forth. To add a new deployment that can be spun up by the control plane: +A deployment YAML specifies the base model, training mode, context length, GPUs, parallelism, and inference settings. Shared deployments are defined only through these files. -1. Existing model definitions are in [`providers/modal/definitions`](../src/lilo/providers/modal/definitions) (one per file). Create a new file with the desired configuration details (model, checkpoint, context-length, GPU, and parallelism settings). -2. Keep the module filename, `DEFINITION_ID`, and engine function name the same. -3. Import the module in - [`providers/modal/app.py`](../src/lilo/providers/modal/app.py) and append it - to `DEFINITIONS`. +1. Create a YAML under `deployments/`, optionally extending a packaged preset. +2. Add its path to the list in [`scripts/deploy_models.sh`](../scripts/deploy_models.sh). +3. Run the script to apply the complete list to the shared frontend. -Every definition exports: +There is no model-specific Python module or catalog registration to update. The generic builders in [`yaml_apps.py`](../src/lilo/providers/modal/yaml_apps.py) construct trainer functions and inference apps from the resolved YAML. See [YAML deployments](deployment-yaml-design.md) for the configuration schema and app structure. -- `DEFINITION_ID`, `MODEL_NAME`, `PARAMETERIZATION`, and `CATALOG_VISIBLE`; -- `TRAINER_MODELS_PER_INSTANCE`; -- the model asset Volume and backend configuration; -- a Modal `app` and engine function that calls - [`run_engine_with_backend`](../src/lilo/providers/modal/serve.py); and -- `ENGINE_FUNCTION`, referencing that engine function. +Set each configuration's trainer limit with `trainer.scaling.max_instances` and its inference limits with `inference.scaling`. Trainer limits are read from YAML; `LILO_TRAINER_MAX_CONTAINERS` is no longer used. -The existing model definitions show the complete template for parameterization-specific settings. FOr example, FFT definitions include rollout resources + Stitch bulletin volume, whereas LoRA definitions include adapter rank, slot capacity, and adapter storage. Our FFT engines are currently only capable of hosting one model (but if multiple FFT experiments are submitted to the control plane, it will spin up as many engine replicas as necessary to support these concurrently). LoRA engines are multi-lora and so can host multiple adapters. - -Adding the module to `DEFINITIONS` automatically includes its Modal -sub-application and makes it available to model lookup, trainer provisioning, -parameterization lookup, and the public catalog. The registry tests also cover -the new definition automatically. To run the tests before deploying: +To check training, publication, and sampling against a deployed configuration: ```bash -uv run pytest tests/providers/test_definition_registry.py -uv run modal deploy -m lilo.providers.modal.app +uv run python scripts/yaml_deployment_smoke.py \ + --frontend lilo-yaml --name qwen35-9b-lora-16k \ + --output /tmp/lilo-smoke.json ``` -Trainer container limits are deployment-specific. Set -`LILO_TRAINER_MAX_CONTAINERS` to a positive integer when deploying to apply the -same limit to every definition. Leaving it unset makes trainer containers -unlimited. - -The following script tests e2e deployment for one model definition, launching the Modal app, creating the particular model, running forward_backward + optim_step operations, and sampler weight publication/rollouts: +For the longer FFT checkpoint and sampler-recovery checks: ```bash uv run python scripts/e2e_engine_definition.py \ - --definition-id + --frontend lilo-yaml --name qwen35-4b-fft-64k --checkpoint-only ``` -This validates all steps of the training cycle for the particular model definition. For an FFT checkpoint round trip, add -`--checkpoint-only`. +Both scripts read the deployed configuration rather than importing model-specific Python definitions. ## Adding a new backend diff --git a/docs/observability.md b/docs/observability.md index 62b9646..fa5db3f 100644 --- a/docs/observability.md +++ b/docs/observability.md @@ -45,7 +45,7 @@ Deploy Lilo after updating the secret: ```bash MODAL_PROFILE=your-workspace MODAL_ENVIRONMENT=your-environment \ - uv run modal deploy -m lilo.providers.modal.app + uv run lilo deploy deployment.yaml ``` The configuration applies to the control plane, sampling workers, and new diff --git a/docs/profiling.md b/docs/profiling.md index 5136d9d..2102d16 100644 --- a/docs/profiling.md +++ b/docs/profiling.md @@ -14,7 +14,7 @@ Pick the step you want to trace and set one environment variable in the shell you deploy from. The trainer inherits it: ```bash -LILO_TORCH_PROFILE_STEP=2 uv run modal deploy -m lilo.providers.modal.app +LILO_TORCH_PROFILE_STEP=2 uv run lilo deploy deployment.yaml ``` Then run your training loop as usual. Steps are counted from 0, so `2` traces diff --git a/scripts/e2e_engine_definition.py b/scripts/e2e_engine_definition.py index 39a22ff..feba8b1 100644 --- a/scripts/e2e_engine_definition.py +++ b/scripts/e2e_engine_definition.py @@ -7,26 +7,49 @@ import random import time import uuid -from contextlib import nullcontext from pathlib import Path from typing import Any +from types import SimpleNamespace import httpx import modal import tinker from tinker import types -from lilo.providers.modal.app import app, cleaner, module_for, server +from lilo.deployments import ResolvedDeployment +from lilo.providers.modal.recipe import backend_config -DEFAULT_DEFINITION = "qwen3_5_4b_full_64k" TIMEOUT = 3 * 60 * 60 -def _definition(definition_id: str) -> tuple[Any, str]: - module = module_for(definition_id) - if not module.CATALOG_VISIBLE: - raise ValueError(f"definition is not cataloged: {definition_id}") - return module, module.PARAMETERIZATION +def _definition(frontend: str, name: str) -> tuple[Any, str]: + rows = modal.Dict.from_name(f"{frontend}-yaml-deployments").get("manifest", []) + matches = [ + ResolvedDeployment.model_validate(row) + for row in rows + if row["active"] and row["spec"]["name"] == name + ] + if len(matches) != 1: + raise ValueError( + f"expected one active YAML configuration named {name} in {frontend}" + ) + resolved = matches[0] + spec = resolved.spec + settings = backend_config(spec)[spec.trainer.backend] + definition = SimpleNamespace( + DEFINITION_ID=resolved.definition_id, + MODEL_NAME=spec.model.id, + PARAMETERIZATION=spec.model.parameterization, + MAX_CONTEXT_LENGTH=spec.model.max_context_length, + MAX_TOKENS_PER_MICROBATCH=settings.get( + "max_tokens_per_microbatch", settings.get("max_tokens_per_gpu") + ), + MICRO_BATCH_SIZE=settings.get("micro_batch_size", 1), + GPU_TYPE=spec.trainer.resources.gpu.split(":")[0], + GPUS=spec.trainer.resources.gpu_count, + LORA_RANK=settings.get("max_lora_rank"), + ) + return definition, definition.PARAMETERIZATION def _timestamped(path: Path) -> Path: @@ -205,9 +228,7 @@ def _checkpoint_roundtrip( load_finished = time.perf_counter() restore_error = _max_error(restored, before) if restore_error > 1e-5: - raise RuntimeError( - f"checkpoint restore max logprob error: {restore_error}" - ) + raise RuntimeError(f"checkpoint restore max logprob error: {restore_error}") resumed_step = _forward_step(resumed, [datum], [length], trained_tokens) continued = _forward_logprobs(resumed, datum, trained_tokens) @@ -417,9 +438,9 @@ def _create_training(service, module, parameterization: str): if parameterization == "full": from lilo.client import create_full_training_client - return create_full_training_client(service, module.MODEL_NAME) + return create_full_training_client(service, module.DEFINITION_ID) return service.create_lora_training_client( - base_model=module.MODEL_NAME, + base_model=module.DEFINITION_ID, rank=module.LORA_RANK, ) @@ -595,14 +616,14 @@ def _correctness( def _run(args: argparse.Namespace, base_url: str, output: Path) -> dict: - module, parameterization = _definition(args.definition_id) + module, parameterization = _definition(args.frontend, args.name) api_key = os.environ["TINKER_API_KEY"] context_length = int(module.MAX_CONTEXT_LENGTH) packed_capacity = int(module.MAX_TOKENS_PER_MICROBATCH) report: dict[str, Any] = { "status": "running", "definition": { - "definition_id": args.definition_id, + "definition_id": module.DEFINITION_ID, "base_model": module.MODEL_NAME, "parameterization": parameterization, "context_length": context_length, @@ -668,7 +689,7 @@ def _run(args: argparse.Namespace, base_url: str, output: Path) -> dict: sampling, warmup = _warm( training, tokenizer, - args.definition_id, + module.DEFINITION_ID, ) report["phases"]["warmup"] = warmup _write(output, report) @@ -705,7 +726,8 @@ def _run(args: argparse.Namespace, base_url: str, output: Path) -> dict: def main() -> None: parser = argparse.ArgumentParser() - parser.add_argument("--definition-id", default=DEFAULT_DEFINITION) + parser.add_argument("--frontend", required=True) + parser.add_argument("--name", required=True) parser.add_argument("--base-url") parser.add_argument("--skip-max-context", action="store_true") parser.add_argument("--checkpoint-only", action="store_true") @@ -717,30 +739,30 @@ def main() -> None: default=Path("scripts/results/engine_definition_e2e.json"), ) args = parser.parse_args() - if sum( - ( - args.checkpoint_only, - args.hf_roundtrip_only, - args.sampler_recovery_only, + if ( + sum( + ( + args.checkpoint_only, + args.hf_roundtrip_only, + args.sampler_recovery_only, + ) ) - ) > 1: + > 1 + ): parser.error( "--checkpoint-only, --hf-roundtrip-only, and " "--sampler-recovery-only are exclusive" ) output = _timestamped(args.output) - context = nullcontext(args.base_url) - if args.base_url is None: - context = app.run(name=f"tinker-e2e-{uuid.uuid4().hex[:12]}") try: - with modal.enable_output(), context: - base_url = args.base_url or server.get_web_url() - if not base_url: - raise RuntimeError("Modal did not provide a control-plane URL") - report = _run(args, base_url, output) - if args.base_url is None: - cleaner.remote() + base_url = ( + args.base_url + or modal.Function.from_name(args.frontend, "server").get_web_url() + ) + if not base_url: + raise RuntimeError("Modal did not provide a control-plane URL") + report = _run(args, base_url, output) finally: print(output) summary = { diff --git a/src/lilo/control_plane/deployments.py b/src/lilo/control_plane/deployments.py index 7c371ed..38a63da 100644 --- a/src/lilo/control_plane/deployments.py +++ b/src/lilo/control_plane/deployments.py @@ -1,4 +1,4 @@ -"""Model-name routing for legacy definitions and YAML deployment generations.""" +"""Model-name routing for shared YAML deployments and scoped engines.""" class DeploymentRoutes: @@ -43,9 +43,10 @@ def select(self, model, mode): ) def sampling(self, model): + requested = [d for d in self.all if d.DEFINITION_ID == model] + if len(requested) == 1: + return requested[0] matches = [d for d in self.visible if d.MODEL_NAME == model] - if not any(hasattr(d, "ROUTING_DEFAULT") for d in matches): - return self.select(model, "full") or self.select(model, "lora") explicit = [d for d in matches if getattr(d, "SAMPLING_DEFAULT", False)] if len(explicit) == 1: return explicit[0] diff --git a/src/lilo/deployment_cli.py b/src/lilo/deployment_cli.py index 4f9fc3f..c3a1cf2 100644 --- a/src/lilo/deployment_cli.py +++ b/src/lilo/deployment_cli.py @@ -119,7 +119,7 @@ def deploy(desired): pass else: raise ValueError( - "The frontend already exists without a YAML registry. Choose a new frontend name; automatic migration of legacy deployments is not implemented." + "The frontend already exists without a YAML registry. Choose a new frontend name; an app with no deployment registry cannot be safely updated." ) # A killed deploy may already have updated Modal. Keep its functions on retry. rows = {r["generation"]: r for r in [*rows, *registry.get("pending", [])]} diff --git a/src/lilo/presets/qwen35-35b-a3b-fft-64k.yaml b/src/lilo/presets/qwen35-35b-a3b-fft-64k.yaml new file mode 100644 index 0000000..fa5761c --- /dev/null +++ b/src/lilo/presets/qwen35-35b-a3b-fft-64k.yaml @@ -0,0 +1,63 @@ +api_version: lilo/v1 +name: qwen35-35b-a3b-fft-64k +model: + id: Qwen/Qwen3.5-35B-A3B + parameterization: full + max_context_length: 65536 +routing: + default: true +trainer: + backend: megatron + resources: + gpu: H200:8 + engine: + max_clients_per_instance: 1 + sampler_persistence_concurrency: 1 + megatron: + options: + tensor_model_parallel_size: 4 + pipeline_model_parallel_size: 1 + context_parallel_size: 2 + expert_model_parallel_size: 8 + expert_tensor_parallel_size: 1 + sequence_parallel: true + micro_batch_size: 1 + max_tokens_per_microbatch: 65536 + bf16: true + fp16: false + gpu_memory_fraction: 0.9 + use_distributed_optimizer: true + provider_overrides: + mtp_num_layers: 0 + recompute_granularity: selective + moe_layer_recompute: true + moe_token_dispatcher_type: alltoall + moe_router_fusion: true + moe_permute_fusion: true + moe_grouped_gemm: true + moe_shared_expert_overlap: false + moe_aux_loss_coeff: 0.0 + optimizer: + optimizer: adam + lr: 0.0001 + min_lr: 0.0001 + env: + PYTORCH_CUDA_ALLOC_CONF: expandable_segments:True + TORCHINDUCTOR_COMPILE_THREADS: '1' +inference: + resources: + gpu: H200:4 + scaling: + min_replicas: 0 + max_replicas: 8 + target_concurrency: 16 + sglang: + options: + tp_size: 4 + ep_size: 4 + mem_fraction_static: 0.9 + max_running_requests: 32 + max_queued_requests: 4 + cpu_weight_cache_max_compile_group_gb: 32 + dp_size: 4 + enable_dp_attention: true diff --git a/src/lilo/presets/qwen35-9b-fft-64k.yaml b/src/lilo/presets/qwen35-9b-fft-64k.yaml new file mode 100644 index 0000000..9c324f5 --- /dev/null +++ b/src/lilo/presets/qwen35-9b-fft-64k.yaml @@ -0,0 +1,52 @@ +api_version: lilo/v1 +name: qwen35-9b-fft-64k +model: + id: Qwen/Qwen3.5-9B + parameterization: full + max_context_length: 65536 +routing: + default: true +trainer: + backend: megatron + resources: + gpu: H200:4 + engine: + max_clients_per_instance: 1 + sampler_persistence_concurrency: 1 + megatron: + options: + tensor_model_parallel_size: 2 + context_parallel_size: 2 + sequence_parallel: true + micro_batch_size: 1 + max_tokens_per_microbatch: 65536 + defer_fp32_logits: true + fp32_lm_head: true + use_distributed_optimizer: true + provider_overrides: + mtp_num_layers: 0 + recompute_granularity: full + recompute_method: uniform + recompute_num_layers: 1 + optimizer: + lr: 0.0001 + min_lr: 0.0001 + loss_scale: 1.0 + env: + PYTORCH_CUDA_ALLOC_CONF: expandable_segments:True + TORCHINDUCTOR_COMPILE_THREADS: '1' +inference: + resources: + gpu: H200:1 + scaling: + min_replicas: 0 + max_replicas: 8 + target_concurrency: 16 + sglang: + options: + tp_size: 1 + ep_size: 1 + mem_fraction_static: 0.85 + max_running_requests: 32 + max_queued_requests: 4 + cpu_weight_cache_max_compile_group_gb: 16 diff --git a/src/lilo/presets/qwen35-9b-instruct-lora-16k-dp2.yaml b/src/lilo/presets/qwen35-9b-instruct-lora-16k-dp2.yaml new file mode 100644 index 0000000..bfc36b0 --- /dev/null +++ b/src/lilo/presets/qwen35-9b-instruct-lora-16k-dp2.yaml @@ -0,0 +1,8 @@ +extends: builtin:qwen35-9b-instruct-lora-16k +name: qwen35-9b-instruct-lora-16k-dp2 +routing: + default: false +trainer: + miles: + options: + tensor_model_parallel_size: 4 diff --git a/src/lilo/presets/qwen35-9b-instruct-lora-16k.yaml b/src/lilo/presets/qwen35-9b-instruct-lora-16k.yaml new file mode 100644 index 0000000..083c193 --- /dev/null +++ b/src/lilo/presets/qwen35-9b-instruct-lora-16k.yaml @@ -0,0 +1,59 @@ +api_version: lilo/v1 +name: qwen35-9b-instruct-lora-16k +model: + id: Qwen/Qwen3.5-9B + parameterization: lora + max_context_length: 16384 +routing: + default: true +trainer: + backend: miles + resources: + gpu: H100:8 + engine: + max_clients_per_instance: 6 + sampler_persistence_concurrency: 8 + miles: + model_args: qwen3.5-9B + options: + tensor_model_parallel_size: 8 + target_modules: + - linear_qkv + - linear_proj + - linear_fc1 + - linear_fc2 + max_tokens_per_gpu: 16384 + recompute_granularity: full + recompute_method: uniform + recompute_num_layers: 1 + multi_lora_n_adapters: 6 + lora_rank: 32 + lora_alpha: 32 + env: + PYTORCH_CUDA_ALLOC_CONF: expandable_segments:True + TORCHINDUCTOR_COMPILE_THREADS: '1' +inference: + resources: + gpu: H200:1 + scaling: + min_replicas: 0 + max_replicas: 8 + target_concurrency: 16 + sglang: + options: + tp_size: 1 + ep_size: 1 + mem_fraction_static: 0.8 + max_running_requests: 32 + max_queued_requests: 8 + max_loaded_loras: 256 + max_loras_per_batch: 8 + lora_target_modules: + - q_proj + - k_proj + - v_proj + - o_proj + - gate_proj + - up_proj + - down_proj + schedule_policy: lpm diff --git a/src/lilo/presets/qwen35-9b-lora-16k-single.yaml b/src/lilo/presets/qwen35-9b-lora-16k-single.yaml new file mode 100644 index 0000000..0dd5c24 --- /dev/null +++ b/src/lilo/presets/qwen35-9b-lora-16k-single.yaml @@ -0,0 +1,7 @@ +extends: builtin:qwen35-9b-lora-16k +name: qwen35-9b-lora-16k-single +routing: + default: false +trainer: + engine: + max_clients_per_instance: 1 diff --git a/src/lilo/presets/qwen35-9b-lora-2k.yaml b/src/lilo/presets/qwen35-9b-lora-2k.yaml new file mode 100644 index 0000000..9d1bdef --- /dev/null +++ b/src/lilo/presets/qwen35-9b-lora-2k.yaml @@ -0,0 +1,61 @@ +api_version: lilo/v1 +name: qwen35-9b-lora-2k +model: + id: Qwen/Qwen3.5-9B-Base + parameterization: lora + max_context_length: 2048 +routing: + default: false +trainer: + backend: miles + resources: + gpu: H200:4 + engine: + max_clients_per_instance: 4 + sampler_persistence_concurrency: 8 + miles: + model_args: qwen3.5-9B + options: + tensor_model_parallel_size: 4 + target_modules: + - linear_qkv + - linear_proj + - linear_fc1 + - linear_fc2 + - output_layer + max_tokens_per_gpu: 2048 + recompute_granularity: full + recompute_method: uniform + recompute_num_layers: 1 + multi_lora_n_adapters: 4 + lora_rank: 32 + lora_alpha: 32 + env: + PYTORCH_CUDA_ALLOC_CONF: expandable_segments:True + TORCHINDUCTOR_COMPILE_THREADS: '1' +inference: + resources: + gpu: H200:1 + scaling: + min_replicas: 0 + max_replicas: 8 + target_concurrency: 16 + sglang: + options: + tp_size: 1 + ep_size: 1 + mem_fraction_static: 0.8 + max_running_requests: 32 + max_queued_requests: 8 + max_loaded_loras: 32 + max_loras_per_batch: 8 + lora_target_modules: + - q_proj + - k_proj + - v_proj + - o_proj + - gate_proj + - up_proj + - down_proj + - lm_head + schedule_policy: lpm diff --git a/src/lilo/presets/qwen36-27b-fft-64k.yaml b/src/lilo/presets/qwen36-27b-fft-64k.yaml new file mode 100644 index 0000000..6bea537 --- /dev/null +++ b/src/lilo/presets/qwen36-27b-fft-64k.yaml @@ -0,0 +1,54 @@ +api_version: lilo/v1 +name: qwen36-27b-fft-64k +model: + id: Qwen/Qwen3.6-27B + parameterization: full + max_context_length: 65536 +routing: + default: true +trainer: + backend: megatron + resources: + gpu: H200:8 + engine: + max_clients_per_instance: 1 + sampler_persistence_concurrency: 1 + megatron: + options: + tensor_model_parallel_size: 4 + pipeline_model_parallel_size: 1 + context_parallel_size: 2 + sequence_parallel: true + micro_batch_size: 1 + max_tokens_per_microbatch: 65536 + bf16: true + fp16: false + gpu_memory_fraction: 0.9 + use_distributed_optimizer: true + provider_overrides: + mtp_num_layers: 0 + recompute_granularity: full + recompute_method: uniform + recompute_num_layers: 1 + optimizer: + optimizer: adam + lr: 0.0001 + min_lr: 0.0001 + env: + PYTORCH_CUDA_ALLOC_CONF: expandable_segments:True + TORCHINDUCTOR_COMPILE_THREADS: '1' +inference: + resources: + gpu: H200:4 + scaling: + min_replicas: 0 + max_replicas: 8 + target_concurrency: 16 + sglang: + options: + tp_size: 4 + ep_size: 1 + mem_fraction_static: 0.9 + max_running_requests: 32 + max_queued_requests: 4 + cpu_weight_cache_max_compile_group_gb: 32 diff --git a/src/lilo/presets/qwen36-35b-a3b-fft-64k.yaml b/src/lilo/presets/qwen36-35b-a3b-fft-64k.yaml new file mode 100644 index 0000000..03a9ae7 --- /dev/null +++ b/src/lilo/presets/qwen36-35b-a3b-fft-64k.yaml @@ -0,0 +1,9 @@ +extends: builtin:qwen35-35b-a3b-fft-64k +name: qwen36-35b-a3b-fft-64k +model: + id: Qwen/Qwen3.6-35B-A3B +inference: + sglang: + options: + dp_size: 1 + enable_dp_attention: false diff --git a/src/lilo/presets/qwen38-27b-lora-128k.yaml b/src/lilo/presets/qwen38-27b-lora-128k.yaml new file mode 100644 index 0000000..f6abba9 --- /dev/null +++ b/src/lilo/presets/qwen38-27b-lora-128k.yaml @@ -0,0 +1,22 @@ +extends: builtin:qwen38-27b-lora-16k +name: qwen38-27b-lora-128k +model: + max_context_length: 131072 +routing: + default: false +trainer: + miles: + options: + tensor_model_parallel_size: 2 + context_parallel_size: 4 + max_tokens_per_gpu: 32768 +inference: + resources: + gpu: H200:2 + scaling: + max_replicas: 4 + target_concurrency: 4 + sglang: + options: + tp_size: 2 + max_running_requests: 8 diff --git a/src/lilo/presets/qwen38-27b-lora-16k.yaml b/src/lilo/presets/qwen38-27b-lora-16k.yaml new file mode 100644 index 0000000..118a28a --- /dev/null +++ b/src/lilo/presets/qwen38-27b-lora-16k.yaml @@ -0,0 +1,59 @@ +api_version: lilo/v1 +name: qwen38-27b-lora-16k +model: + id: Qwen/Qwen3.8-27B + parameterization: lora + max_context_length: 16384 +routing: + default: true +trainer: + backend: miles + resources: + gpu: H200:8 + engine: + max_clients_per_instance: 6 + sampler_persistence_concurrency: 8 + miles: + model_args: qwen3.8-27B + options: + tensor_model_parallel_size: 4 + target_modules: + - linear_qkv + - linear_proj + - linear_fc1 + - linear_fc2 + max_tokens_per_gpu: 16384 + recompute_granularity: full + recompute_method: uniform + recompute_num_layers: 1 + multi_lora_n_adapters: 6 + lora_rank: 32 + lora_alpha: 32 + env: + PYTORCH_CUDA_ALLOC_CONF: expandable_segments:True + TORCHINDUCTOR_COMPILE_THREADS: '1' +inference: + resources: + gpu: H200:1 + scaling: + min_replicas: 0 + max_replicas: 8 + target_concurrency: 16 + sglang: + options: + tp_size: 1 + ep_size: 1 + mem_fraction_static: 0.8 + max_running_requests: 32 + max_queued_requests: 8 + max_loaded_loras: 256 + max_loras_per_batch: 8 + lora_target_modules: + - q_proj + - k_proj + - v_proj + - o_proj + - gate_proj + - up_proj + - down_proj + schedule_policy: lpm diff --git a/src/lilo/presets/qwen38-27b-lora-64k.yaml b/src/lilo/presets/qwen38-27b-lora-64k.yaml new file mode 100644 index 0000000..4a9f127 --- /dev/null +++ b/src/lilo/presets/qwen38-27b-lora-64k.yaml @@ -0,0 +1,17 @@ +extends: builtin:qwen38-27b-lora-16k +name: qwen38-27b-lora-64k +model: + max_context_length: 65536 +routing: + default: false +trainer: + miles: + options: + context_parallel_size: 2 + max_tokens_per_gpu: 32768 +inference: + scaling: + target_concurrency: 8 + sglang: + options: + max_running_requests: 16 diff --git a/src/lilo/providers/modal/app.py b/src/lilo/providers/modal/app.py index 195d41b..2a3254e 100644 --- a/src/lilo/providers/modal/app.py +++ b/src/lilo/providers/modal/app.py @@ -16,14 +16,9 @@ from .checkpoint_storage import ( CHECKPOINT_ROOT, - CHECKPOINT_VOLUME_NAME, ModalCheckpointStorage, - checkpoint_volume, -) -from .deployment import ( - trainer_deployment_env, - trainer_max_containers, ) +from .deployment import trainer_deployment_env from .engines import ModalEnginePlatform from .fft_pool import ( FFTPoolSpec, @@ -58,72 +53,29 @@ manifest_from_env, ) -APP_NAME = os.environ.get("LILO_APP_NAME", "lilo") -ROUTING_REGION = "us-west" +SETTINGS = frontend_settings() +APP_NAME = SETTINGS.deployment.frontend +ROUTING_REGION = SETTINGS.deployment.modal.region MODEL_ASSET_ROOT = "/assets" -SESSION_IDLE_TIMEOUT = 300.0 -FFT_POOL_IDLE_TIMEOUT = 300.0 -LORA_POOL_IDLE_TIMEOUT = 300.0 +SESSION_IDLE_TIMEOUT = SETTINGS.lifecycle.session_idle_timeout_s +FFT_POOL_IDLE_TIMEOUT = LORA_POOL_IDLE_TIMEOUT = SETTINGS.lifecycle.pool_idle_timeout_s FFT_POOL_TOUCH_INTERVAL = 60.0 LORA_POOL_CHECK_INTERVAL = 60.0 -SWEEP_PERIOD = modal.Period(minutes=5) +SWEEP_PERIOD = modal.Period(seconds=SETTINGS.lifecycle.sweep_interval_s) CHECKPOINT_READ_LOCK = asyncio.Lock() _pool_touches: dict[str, float] = {} _lora_pool_gateways: dict[str, tuple[float, str]] = {} _lora_pool_checks: dict[str, asyncio.Lock] = {} - - -YAML_SETTINGS = frontend_settings() -if YAML_SETTINGS is not None: - APP_NAME = YAML_SETTINGS.deployment.frontend - ROUTING_REGION = YAML_SETTINGS.deployment.modal.region - SESSION_IDLE_TIMEOUT = YAML_SETTINGS.lifecycle.session_idle_timeout_s - FFT_POOL_IDLE_TIMEOUT = LORA_POOL_IDLE_TIMEOUT = ( - YAML_SETTINGS.lifecycle.pool_idle_timeout_s - ) - SWEEP_PERIOD = modal.Period(seconds=YAML_SETTINGS.lifecycle.sweep_interval_s) - CHECKPOINT_VOLUME_NAME = YAML_SETTINGS.deployment.storage.checkpoints - checkpoint_volume = modal.Volume.from_name( - CHECKPOINT_VOLUME_NAME, create_if_missing=True, version=2 - ) - DEFINITIONS = tuple(definition_from_spec(row) for row in manifest_from_env()) -else: - from .definitions import ( - qwen3_5_4b_full_64k, - qwen3_5_9b_base_miles_lora_2k, - qwen3_5_9b_base_miles_lora_16k, - qwen3_5_9b_base_miles_lora_16k_single, - qwen3_5_9b_full_64k, - qwen3_5_9b_miles_lora_16k, - qwen3_5_9b_miles_lora_16k_dp2, - qwen3_5_35b_a3b_full_64k, - qwen3_6_27b_full_64k, - qwen3_6_35b_a3b_full_64k, - qwen3_8_27b_miles_lora_16k, - qwen3_8_27b_miles_lora_64k, - qwen3_8_27b_miles_lora_128k, - ) - - DEFINITIONS = ( - qwen3_5_4b_full_64k, - qwen3_5_9b_full_64k, - qwen3_5_9b_base_miles_lora_2k, - qwen3_5_9b_base_miles_lora_16k, - qwen3_5_9b_base_miles_lora_16k_single, - qwen3_5_9b_miles_lora_16k, - qwen3_5_9b_miles_lora_16k_dp2, - qwen3_5_35b_a3b_full_64k, - qwen3_6_27b_full_64k, - qwen3_6_35b_a3b_full_64k, - qwen3_8_27b_miles_lora_16k, - qwen3_8_27b_miles_lora_64k, - qwen3_8_27b_miles_lora_128k, - ) -TRAINER_MAX_CONTAINERS = trainer_max_containers() -TRAINER_DEPLOYMENT_ENV = trainer_deployment_env() -if YAML_SETTINGS is not None: - TRAINER_DEPLOYMENT_ENV[MANIFEST_ENV] = os.environ[MANIFEST_ENV] - TRAINER_DEPLOYMENT_ENV["LILO_APP_NAME"] = APP_NAME +CHECKPOINT_VOLUME_NAME = SETTINGS.deployment.storage.checkpoints +checkpoint_volume = modal.Volume.from_name( + CHECKPOINT_VOLUME_NAME, create_if_missing=True, version=2 +) +DEFINITIONS = tuple(definition_from_spec(row) for row in manifest_from_env()) +TRAINER_DEPLOYMENT_ENV = { + **trainer_deployment_env(), + MANIFEST_ENV: os.environ[MANIFEST_ENV], + "LILO_APP_NAME": APP_NAME, +} app = modal.App(APP_NAME) for definition in DEFINITIONS: app.include(definition.app) @@ -148,7 +100,7 @@ async def _delete_checkpoint(uri: str) -> None: image = ( - modal.Image.debian_slim(python_version="3.12" if YAML_SETTINGS else "3.11") + modal.Image.debian_slim(python_version="3.12") .apt_install("git") .pip_install(*CORE_PACKAGES, TINKER_PACKAGE) .pip_install(STITCH_PACKAGE, "huggingface-hub") @@ -156,17 +108,13 @@ async def _delete_checkpoint(uri: str) -> None: .add_local_python_source("lilo") ) model_assets = modal.Volume.from_name( - YAML_SETTINGS.deployment.storage.assets if YAML_SETTINGS else "lilo-model-assets", + SETTINGS.deployment.storage.assets, create_if_missing=True, ) -API_SECRET_NAME = YAML_SETTINGS.deployment.secrets.api if YAML_SETTINGS else "lilo-api" -HF_SECRET_NAME = ( - YAML_SETTINGS.deployment.secrets.huggingface - if YAML_SETTINGS - else "huggingface-secret" -) +API_SECRET_NAME = SETTINGS.deployment.secrets.api +HF_SECRET_NAME = SETTINGS.deployment.secrets.huggingface proxy_secret = modal.Secret.from_name( - YAML_SETTINGS.deployment.secrets.sampler_proxy if YAML_SETTINGS else "lilo-proxy", + SETTINGS.deployment.secrets.sampler_proxy, required_keys=["MODAL_PROXY_TOKEN_ID", "MODAL_PROXY_TOKEN_SECRET"], ) @@ -189,10 +137,11 @@ def prepare_model_assets(definition_id: str) -> None: or checkpoint == MODEL_ASSET_ROOT ): raise ValueError(f"invalid model asset path: {checkpoint}") - kwargs = {} - if hasattr(definition, "MODEL_REVISION"): - kwargs["revision"] = definition.MODEL_REVISION - snapshot_download(repo_id=definition.MODEL_NAME, local_dir=checkpoint, **kwargs) + snapshot_download( + repo_id=definition.MODEL_NAME, + local_dir=checkpoint, + revision=definition.MODEL_REVISION, + ) model_assets.commit() @@ -398,18 +347,14 @@ async def run(definition_id: str, token: str) -> None: await complete_reconcile(definition_id, token) return module = module_for(definition_id) - maximum_instances = getattr( - module, "TRAINER_MAX_CONTAINERS", TRAINER_MAX_CONTAINERS - ) + maximum_instances = module.TRAINER_MAX_CONTAINERS try: await reconcile_trainers( shared_kv(), ModalEnginePlatform(shared_kv(), _spawn_engine), definition_id, revision=None, - maximum_instances=( - int(maximum_instances) if maximum_instances is not None else None - ), + maximum_instances=maximum_instances, models_per_instance=module.TRAINER_MODELS_PER_INSTANCE, scale_up=trainer_autoscaling(definition_id), ) @@ -454,8 +399,6 @@ async def _spawn_engine(definition_id: str, instance_id: str) -> str: async def deployment_error(definition_id: str) -> str | None: - if not definition_id.startswith("yaml_"): - return None record = await shared_kv().get(f"deployment_failure:{definition_id}") if record: return f"Trainer startup failed for {definition_id}: {record['error']}. See Modal call logs for instance {record['instance_id']}; after correcting the cause, run lilo deployment retry." @@ -496,7 +439,7 @@ async def ensure_pool(session) -> None: definition_id = session.engine_definition_id parameterization = parameterization_for(definition_id) if parameterization == "lora": - if definition_id.startswith("yaml_") and session.model_id is None: + if session.model_id is None: await prepare_model_assets.remote.aio(definition_id) await _ready_lora_pool(LoraPoolSpec(definition_id)) return @@ -529,11 +472,7 @@ async def kick_trainers(definition_id: str) -> bool: await kick_trainer_reconciler(definition_id) if not trainer_autoscaling(definition_id): return False - maximum = getattr( - module_for(definition_id), "TRAINER_MAX_CONTAINERS", TRAINER_MAX_CONTAINERS - ) - if maximum is None: - return True + maximum = module_for(definition_id).TRAINER_MAX_CONTAINERS instances = [ instance for instance in await engines.list_instances() diff --git a/src/lilo/providers/modal/definitions/__init__.py b/src/lilo/providers/modal/definitions/__init__.py deleted file mode 100644 index e69de29..0000000 diff --git a/src/lilo/providers/modal/definitions/qwen3_5_35b_a3b_full_64k.py b/src/lilo/providers/modal/definitions/qwen3_5_35b_a3b_full_64k.py deleted file mode 100644 index 95016a6..0000000 --- a/src/lilo/providers/modal/definitions/qwen3_5_35b_a3b_full_64k.py +++ /dev/null @@ -1,152 +0,0 @@ -from __future__ import annotations - -import os - -import modal - -from ..checkpoint_storage import ( - CHECKPOINT_ROOT, - CHECKPOINT_VOLUME_NAME, - checkpoint_volume, -) -from ..deployment import trainer_max_containers - -MODEL_NAME = "Qwen/Qwen3.5-35B-A3B" -HF_CHECKPOINT = "/assets/Qwen3.5-35B-A3B" -MICRO_BATCH_SIZE = 1 -MAX_CONTEXT_LENGTH = 65_536 -MAX_TOKENS_PER_MICROBATCH = MAX_CONTEXT_LENGTH -DEFINITION_ID = "qwen3_5_35b_a3b_full_64k" -PARAMETERIZATION = "full" -CATALOG_VISIBLE = True -TRAINER_MODELS_PER_INSTANCE = 1 -GPU_TYPE = "H200" -GPUS = 8 -TENSOR_MODEL_PARALLEL_SIZE = 4 -PIPELINE_MODEL_PARALLEL_SIZE = 1 -CONTEXT_PARALLEL_SIZE = 2 -DATA_PARALLEL_SIZE = 1 -EXPERT_MODEL_PARALLEL_SIZE = 8 -EXPERT_TENSOR_PARALLEL_SIZE = 1 -SEQUENCE_PARALLEL = True -GPU_MEMORY_FRACTION = 0.90 -MTP_NUM_LAYERS = 0 -ROLLOUT_GPU_TYPE = "H200" -ROLLOUT_GPUS = 4 -ROLLOUT_TENSOR_PARALLEL_SIZE = 1 -ROLLOUT_EXPERT_PARALLEL_SIZE = 4 -ROLLOUT_EXPERT_TENSOR_PARALLEL_SIZE = 1 -ROLLOUT_MEMORY_FRACTION = 0.90 -ROLLOUT_CPU_WEIGHT_CACHE_MAX_COMPILE_GROUP_GB = 32 -ROLLOUT_MAX_RUNNING_REQUESTS = 32 -ROLLOUT_MAX_QUEUED_REQUESTS = 4 -ROLLOUT_TARGET_CONCURRENCY = 16 -BULLETIN_ROOT = "/bulletin" -BULLETIN_VOLUME_NAME = "lilo-snapshot-bulletin" -app = modal.App(f"lilo-{DEFINITION_ID}") - -if modal.is_local(): - from ..megatron_image import image -else: - image = modal.Image.debian_slim() - -assets = modal.Volume.from_name("lilo-model-assets", create_if_missing=True) -bulletin = modal.Volume.from_name( - BULLETIN_VOLUME_NAME, - create_if_missing=True, - version=2, -) -TRAINER_VOLUMES = { - "/assets": assets, - BULLETIN_ROOT: bulletin, - CHECKPOINT_ROOT: checkpoint_volume, -} -api_secret = modal.Secret.from_name( - "lilo-api", - required_keys=["TINKER_API_KEY"], -) -proxy_secret = modal.Secret.from_name( - "lilo-proxy", - required_keys=["MODAL_PROXY_TOKEN_ID", "MODAL_PROXY_TOKEN_SECRET"], -) - - -@app.function( - image=image, - gpu=f"{GPU_TYPE}:{GPUS}", - volumes=TRAINER_VOLUMES, - secrets=[api_secret, proxy_secret], - timeout=86_400, - max_containers=trainer_max_containers(), - single_use_containers=True, -) -def qwen3_5_35b_a3b_full_64k(instance_id: str) -> None: - import json - - from huggingface_hub import snapshot_download - from modal.config import config - - from lilo.providers.modal.kv import shared_kv - from lilo.providers.modal.serve import run_engine_with_backend - - if not os.path.exists(HF_CHECKPOINT): - snapshot_download(repo_id=MODEL_NAME, local_dir=HF_CHECKPOINT) - assets.commit() - backend_config = { - "megatron": { - "hf_checkpoint": HF_CHECKPOINT, - "tensor_model_parallel_size": TENSOR_MODEL_PARALLEL_SIZE, - "pipeline_model_parallel_size": PIPELINE_MODEL_PARALLEL_SIZE, - "context_parallel_size": CONTEXT_PARALLEL_SIZE, - "expert_model_parallel_size": EXPERT_MODEL_PARALLEL_SIZE, - "expert_tensor_parallel_size": EXPERT_TENSOR_PARALLEL_SIZE, - "sequence_parallel": SEQUENCE_PARALLEL, - "micro_batch_size": MICRO_BATCH_SIZE, - "max_tokens_per_microbatch": MAX_TOKENS_PER_MICROBATCH, - "seq_length": MAX_CONTEXT_LENGTH, - "bf16": True, - "fp16": False, - "gpu_memory_fraction": GPU_MEMORY_FRACTION, - "use_distributed_optimizer": True, - "provider_overrides": { - "mtp_num_layers": MTP_NUM_LAYERS, - "recompute_granularity": "selective", - "moe_layer_recompute": True, - "moe_token_dispatcher_type": "alltoall", - "moe_router_fusion": True, - "moe_permute_fusion": True, - "moe_grouped_gemm": True, - "moe_shared_expert_overlap": False, - "moe_aux_loss_coeff": 0.0, - }, - "optimizer": { - "optimizer": "adam", - "lr": 1e-4, - "min_lr": 1e-4, - }, - }, - "checkpoint_dir": CHECKPOINT_ROOT, - } - run_engine_with_backend( - shared_kv(), - "lilo.backends.megatron_fft:build_executor", - definition_id=DEFINITION_ID, - revision=config["image_id"], - instance_id=instance_id, - backend_env={ - "PYTORCH_CUDA_ALLOC_CONF": "expandable_segments:True", - "TORCHINDUCTOR_COMPILE_THREADS": "1", - "LILO_BACKEND_CONFIG": json.dumps(backend_config), - "LILO_BASE_MODEL": MODEL_NAME, - "LILO_DEFINITION_ID": DEFINITION_ID, - "LILO_CHECKPOINT_VOLUME": CHECKPOINT_VOLUME_NAME, - "LILO_BULLETIN_ROOT": BULLETIN_ROOT, - "LILO_BULLETIN_VOLUME": BULLETIN_VOLUME_NAME, - "LILO_DEFINITION_REVISION": config["image_id"], - }, - nproc=GPUS, - max_models=1, - ) - - -ENGINE_FUNCTION = qwen3_5_35b_a3b_full_64k diff --git a/src/lilo/providers/modal/definitions/qwen3_5_4b_full_64k.py b/src/lilo/providers/modal/definitions/qwen3_5_4b_full_64k.py deleted file mode 100644 index 9ac0f84..0000000 --- a/src/lilo/providers/modal/definitions/qwen3_5_4b_full_64k.py +++ /dev/null @@ -1,144 +0,0 @@ -from __future__ import annotations - -import os - -import modal - -from ..checkpoint_storage import ( - CHECKPOINT_ROOT, - CHECKPOINT_VOLUME_NAME, - checkpoint_volume, -) -from ..deployment import trainer_max_containers - -MODEL_NAME = "Qwen/Qwen3.5-4B" -HF_CHECKPOINT = "/assets/Qwen3.5-4B" -MICRO_BATCH_SIZE = 1 -MAX_CONTEXT_LENGTH = 65_536 -MAX_TOKENS_PER_MICROBATCH = MAX_CONTEXT_LENGTH -DEFINITION_ID = "qwen3_5_4b_full_64k" -PARAMETERIZATION = "full" -CATALOG_VISIBLE = True -TRAINER_MODELS_PER_INSTANCE = 1 -GPU_TYPE = "H100" -GPUS = 4 -TENSOR_MODEL_PARALLEL_SIZE = 2 -CONTEXT_PARALLEL_SIZE = 2 -DATA_PARALLEL_SIZE = GPUS // (TENSOR_MODEL_PARALLEL_SIZE * CONTEXT_PARALLEL_SIZE) -MTP_NUM_LAYERS = 0 -RECOMPUTE_GRANULARITY = "full" -RECOMPUTE_METHOD = "uniform" -RECOMPUTE_NUM_LAYERS = 1 -LOSS_SCALE = 1.0 -DEFER_FP32_LOGITS = True -FP32_LM_HEAD = True -ROLLOUT_GPU_TYPE = "H100" -ROLLOUT_GPUS = 1 -ROLLOUT_TENSOR_PARALLEL_SIZE = 1 -ROLLOUT_EXPERT_PARALLEL_SIZE = 1 -ROLLOUT_EXPERT_TENSOR_PARALLEL_SIZE = 1 -ROLLOUT_MEMORY_FRACTION = 0.85 -ROLLOUT_MAX_RUNNING_REQUESTS = 32 -ROLLOUT_MAX_QUEUED_REQUESTS = 4 -ROLLOUT_TARGET_CONCURRENCY = 16 -ROLLOUT_CPU_WEIGHT_CACHE_MAX_COMPILE_GROUP_GB = 16 -BULLETIN_ROOT = "/bulletin" -BULLETIN_VOLUME_NAME = "lilo-snapshot-bulletin" -app = modal.App(f"lilo-{DEFINITION_ID}") - -if modal.is_local(): - from ..megatron_image import image -else: - image = modal.Image.debian_slim() - -assets = modal.Volume.from_name("lilo-model-assets", create_if_missing=True) -bulletin = modal.Volume.from_name( - BULLETIN_VOLUME_NAME, - create_if_missing=True, - version=2, -) -TRAINER_VOLUMES = { - "/assets": assets, - BULLETIN_ROOT: bulletin, - CHECKPOINT_ROOT: checkpoint_volume, -} -api_secret = modal.Secret.from_name( - "lilo-api", - required_keys=["TINKER_API_KEY"], -) -proxy_secret = modal.Secret.from_name( - "lilo-proxy", - required_keys=["MODAL_PROXY_TOKEN_ID", "MODAL_PROXY_TOKEN_SECRET"], -) - - -@app.function( - image=image, - gpu=f"{GPU_TYPE}:{GPUS}", - volumes=TRAINER_VOLUMES, - secrets=[api_secret, proxy_secret], - timeout=86_400, - max_containers=trainer_max_containers(), - single_use_containers=True, -) -def qwen3_5_4b_full_64k(instance_id: str) -> None: - import json - - from huggingface_hub import snapshot_download - from modal.config import config - - from lilo.providers.modal.kv import shared_kv - from lilo.providers.modal.serve import run_engine_with_backend - - if not os.path.exists(HF_CHECKPOINT): - snapshot_download(repo_id=MODEL_NAME, local_dir=HF_CHECKPOINT) - assets.commit() - backend_config = { - "megatron": { - "hf_checkpoint": HF_CHECKPOINT, - "tensor_model_parallel_size": TENSOR_MODEL_PARALLEL_SIZE, - "context_parallel_size": CONTEXT_PARALLEL_SIZE, - "sequence_parallel": True, - "micro_batch_size": MICRO_BATCH_SIZE, - "max_tokens_per_microbatch": MAX_TOKENS_PER_MICROBATCH, - "seq_length": MAX_CONTEXT_LENGTH, - "defer_fp32_logits": DEFER_FP32_LOGITS, - "fp32_lm_head": FP32_LM_HEAD, - "use_distributed_optimizer": True, - "provider_overrides": { - "mtp_num_layers": MTP_NUM_LAYERS, - "recompute_granularity": RECOMPUTE_GRANULARITY, - "recompute_method": RECOMPUTE_METHOD, - "recompute_num_layers": RECOMPUTE_NUM_LAYERS, - }, - "optimizer": { - "lr": 1e-4, - "min_lr": 1e-4, - "loss_scale": LOSS_SCALE, - }, - }, - "checkpoint_dir": CHECKPOINT_ROOT, - } - run_engine_with_backend( - shared_kv(), - "lilo.backends.megatron_fft:build_executor", - definition_id=DEFINITION_ID, - revision=config["image_id"], - instance_id=instance_id, - backend_env={ - "LILO_BACKEND_CONFIG": json.dumps(backend_config), - "LILO_BASE_MODEL": MODEL_NAME, - "LILO_DEFINITION_ID": DEFINITION_ID, - "LILO_CHECKPOINT_VOLUME": CHECKPOINT_VOLUME_NAME, - "LILO_BULLETIN_ROOT": BULLETIN_ROOT, - "LILO_BULLETIN_VOLUME": BULLETIN_VOLUME_NAME, - "LILO_DEFINITION_REVISION": config["image_id"], - "PYTORCH_CUDA_ALLOC_CONF": "expandable_segments:True", - "TORCHINDUCTOR_COMPILE_THREADS": "1", - }, - nproc=GPUS, - max_models=1, - ) - - -ENGINE_FUNCTION = qwen3_5_4b_full_64k diff --git a/src/lilo/providers/modal/definitions/qwen3_5_9b_base_miles_lora_16k.py b/src/lilo/providers/modal/definitions/qwen3_5_9b_base_miles_lora_16k.py deleted file mode 100644 index ccd1e6a..0000000 --- a/src/lilo/providers/modal/definitions/qwen3_5_9b_base_miles_lora_16k.py +++ /dev/null @@ -1,169 +0,0 @@ -from __future__ import annotations - -import os - -import modal - -from ..checkpoint_storage import ( - CHECKPOINT_ROOT, - CHECKPOINT_VOLUME_NAME, - checkpoint_volume, -) -from ..deployment import trainer_deployment_env, trainer_max_containers - -MODEL_NAME = "Qwen/Qwen3.5-9B-Base" -HF_CHECKPOINT = "/assets/Qwen3.5-9B-Base" -DEFINITION_ID = "qwen3_5_9b_base_miles_lora_16k" -PARAMETERIZATION = "lora" -CATALOG_VISIBLE = True -MAX_CONTEXT_LENGTH = 16_384 - -GPU_TYPE = "H100" -GPUS = 4 -TENSOR_MODEL_PARALLEL_SIZE = 4 -MAX_LORA_SLOTS = 6 -MAX_LORA_RANK = 32 -DEFAULT_LORA_ALPHA = 32 -TARGET_MODULES = ( - "linear_qkv", - "linear_proj", - "linear_fc1", - "linear_fc2", - "output_layer", -) -TRAINER_MODELS_PER_INSTANCE = MAX_LORA_SLOTS - -ROLLOUT_GPU_TYPE = "H200" -ROLLOUT_GPUS = 1 -ROLLOUT_TENSOR_PARALLEL_SIZE = 1 -ROLLOUT_EXPERT_PARALLEL_SIZE = 1 -ROLLOUT_EXPERT_TENSOR_PARALLEL_SIZE = 1 -ROLLOUT_MEMORY_FRACTION = 0.8 -ROLLOUT_MAX_RUNNING_REQUESTS = 32 -ROLLOUT_MAX_QUEUED_REQUESTS = 8 -ROLLOUT_TARGET_CONCURRENCY = 16 -# Bound retained publication versions per replica; older versions reload on demand. -ROLLOUT_MAX_LOADED_LORAS = 64 -ROLLOUT_MIN_CONTAINERS = 8 -ROLLOUT_MAX_CONTAINERS = 8 -ROLLOUT_MAX_LORAS_PER_BATCH = 8 -ROLLOUT_LORA_TARGET_MODULES = ( - "q_proj", - "k_proj", - "v_proj", - "o_proj", - "gate_proj", - "up_proj", - "down_proj", - "lm_head", -) - -BULLETIN_ROOT = "/bulletin" -BULLETIN_VOLUME_NAME = "lilo-snapshot-bulletin" -app = modal.App(f"lilo-{DEFINITION_ID}") - -if modal.is_local(): - from ..miles_image import image -else: - image = modal.Image.debian_slim() - -assets = modal.Volume.from_name("lilo-model-assets", create_if_missing=True) -bulletin = modal.Volume.from_name( - BULLETIN_VOLUME_NAME, - create_if_missing=True, - version=2, -) -TRAINER_VOLUMES = { - "/assets": assets, - BULLETIN_ROOT: bulletin, - CHECKPOINT_ROOT: checkpoint_volume, -} -api_secret = modal.Secret.from_name( - "lilo-api", - required_keys=["TINKER_API_KEY"], -) -proxy_secret = modal.Secret.from_name( - "lilo-proxy", - required_keys=["MODAL_PROXY_TOKEN_ID", "MODAL_PROXY_TOKEN_SECRET"], -) - - -@app.function( - image=image, - gpu=f"{GPU_TYPE}:{GPUS}", - volumes=TRAINER_VOLUMES, - secrets=[api_secret, proxy_secret], - timeout=86_400, - max_containers=trainer_max_containers(), - env=trainer_deployment_env(), - single_use_containers=True, -) -def qwen3_5_9b_base_miles_lora_16k(instance_id: str) -> None: - run_trainer(instance_id) - - -def run_trainer( - instance_id: str, - *, - definition_id: str = DEFINITION_ID, - max_models: int = MAX_LORA_SLOTS, -) -> None: - import json - - from huggingface_hub import snapshot_download - from modal.config import config - - from lilo.providers.modal.kv import shared_kv - from lilo.providers.modal.serve import run_engine_with_backend - - if not os.path.exists(HF_CHECKPOINT): - snapshot_download(repo_id=MODEL_NAME, local_dir=HF_CHECKPOINT) - assets.commit() - backend_config = { - "miles": { - "hf_checkpoint": HF_CHECKPOINT, - "model_type": "qwen3.5-9B", - "actor_num_gpus_per_node": GPUS, - "tensor_model_parallel_size": TENSOR_MODEL_PARALLEL_SIZE, - "max_lora_slots": MAX_LORA_SLOTS, - "max_lora_rank": MAX_LORA_RANK, - "default_lora_alpha": DEFAULT_LORA_ALPHA, - "target_modules": TARGET_MODULES, - "max_tokens_per_gpu": MAX_CONTEXT_LENGTH, - "extra_args": ( - "--seq-length", - str(MAX_CONTEXT_LENGTH), - "--recompute-granularity", - "full", - "--recompute-method", - "uniform", - "--recompute-num-layers", - "1", - ), - }, - "checkpoint_dir": CHECKPOINT_ROOT, - } - run_engine_with_backend( - shared_kv(), - "lilo.backends.miles_lora:build_executor", - definition_id=definition_id, - revision=config["image_id"], - instance_id=instance_id, - backend_env={ - "LILO_BACKEND_CONFIG": json.dumps(backend_config), - "LILO_BASE_MODEL": MODEL_NAME, - "LILO_DEFINITION_ID": definition_id, - "LILO_CHECKPOINT_VOLUME": CHECKPOINT_VOLUME_NAME, - "LILO_BULLETIN_ROOT": BULLETIN_ROOT, - "LILO_BULLETIN_VOLUME": BULLETIN_VOLUME_NAME, - "LILO_DEFINITION_REVISION": config["image_id"], - "PYTORCH_CUDA_ALLOC_CONF": "expandable_segments:True", - "TORCHINDUCTOR_COMPILE_THREADS": "1", - }, - nproc=1, - max_models=max_models, - sampler_persistence_concurrency=8, - ) - - -ENGINE_FUNCTION = qwen3_5_9b_base_miles_lora_16k diff --git a/src/lilo/providers/modal/definitions/qwen3_5_9b_base_miles_lora_16k_single.py b/src/lilo/providers/modal/definitions/qwen3_5_9b_base_miles_lora_16k_single.py deleted file mode 100644 index bf3c9ed..0000000 --- a/src/lilo/providers/modal/definitions/qwen3_5_9b_base_miles_lora_16k_single.py +++ /dev/null @@ -1,54 +0,0 @@ -"""Isolated single-tenant option; identical hardware and backend to the shared trainer.""" -# ruff: noqa: F401 -import modal -from ..deployment import trainer_deployment_env, trainer_max_containers -from .qwen3_5_9b_base_miles_lora_16k import ( - MODEL_NAME, - HF_CHECKPOINT, - PARAMETERIZATION, - MAX_CONTEXT_LENGTH, - GPU_TYPE, - GPUS, - MAX_LORA_SLOTS, - MAX_LORA_RANK, - DEFAULT_LORA_ALPHA, - TARGET_MODULES, - TENSOR_MODEL_PARALLEL_SIZE, - BULLETIN_ROOT, - BULLETIN_VOLUME_NAME, - TRAINER_VOLUMES, - assets, - bulletin, - image, - api_secret, - proxy_secret, - run_trainer, - ROLLOUT_GPU_TYPE, - ROLLOUT_GPUS, - ROLLOUT_TENSOR_PARALLEL_SIZE, - ROLLOUT_EXPERT_PARALLEL_SIZE, - ROLLOUT_EXPERT_TENSOR_PARALLEL_SIZE, - ROLLOUT_MEMORY_FRACTION, - ROLLOUT_MAX_RUNNING_REQUESTS, - ROLLOUT_MAX_QUEUED_REQUESTS, - ROLLOUT_TARGET_CONCURRENCY, - ROLLOUT_MAX_LOADED_LORAS, - ROLLOUT_MIN_CONTAINERS, - ROLLOUT_MAX_CONTAINERS, - ROLLOUT_MAX_LORAS_PER_BATCH, - ROLLOUT_LORA_TARGET_MODULES, -) - -DEFINITION_ID = "qwen3_5_9b_base_miles_lora_16k_single" -CATALOG_VISIBLE = False -TRAINER_MODELS_PER_INSTANCE = 1 -app = modal.App(f"lilo-{DEFINITION_ID}") - -@app.function(image=image, gpu=f"{GPU_TYPE}:{GPUS}", volumes=TRAINER_VOLUMES, - secrets=[api_secret, proxy_secret], timeout=86_400, - max_containers=trainer_max_containers(), env=trainer_deployment_env(), - single_use_containers=True) -def qwen3_5_9b_base_miles_lora_16k_single(instance_id: str) -> None: - run_trainer(instance_id, definition_id=DEFINITION_ID, max_models=1) - -ENGINE_FUNCTION = qwen3_5_9b_base_miles_lora_16k_single diff --git a/src/lilo/providers/modal/definitions/qwen3_5_9b_base_miles_lora_2k.py b/src/lilo/providers/modal/definitions/qwen3_5_9b_base_miles_lora_2k.py deleted file mode 100644 index d5a5d1a..0000000 --- a/src/lilo/providers/modal/definitions/qwen3_5_9b_base_miles_lora_2k.py +++ /dev/null @@ -1,155 +0,0 @@ -from __future__ import annotations - -import os - -import modal - -from ..checkpoint_storage import ( - CHECKPOINT_ROOT, - CHECKPOINT_VOLUME_NAME, - checkpoint_volume, -) -from ..deployment import trainer_deployment_env, trainer_max_containers - -MODEL_NAME = "Qwen/Qwen3.5-9B-Base" -HF_CHECKPOINT = "/assets/Qwen3.5-9B-Base" -DEFINITION_ID = "qwen3_5_9b_base_miles_lora_2k" -PARAMETERIZATION = "lora" -CATALOG_VISIBLE = False -MAX_CONTEXT_LENGTH = 2048 - -GPU_TYPE = "H200" -GPUS = 4 -TENSOR_MODEL_PARALLEL_SIZE = 4 -MAX_LORA_SLOTS = 4 -MAX_LORA_RANK = 32 -DEFAULT_LORA_ALPHA = 32 -TARGET_MODULES = ( - "linear_qkv", - "linear_proj", - "linear_fc1", - "linear_fc2", - "output_layer", -) -TRAINER_MODELS_PER_INSTANCE = MAX_LORA_SLOTS - -ROLLOUT_GPU_TYPE = "H200" -ROLLOUT_GPUS = 1 -ROLLOUT_TENSOR_PARALLEL_SIZE = 1 -ROLLOUT_EXPERT_PARALLEL_SIZE = 1 -ROLLOUT_EXPERT_TENSOR_PARALLEL_SIZE = 1 -ROLLOUT_MEMORY_FRACTION = 0.8 -ROLLOUT_MAX_RUNNING_REQUESTS = 32 -ROLLOUT_MAX_QUEUED_REQUESTS = 8 -ROLLOUT_TARGET_CONCURRENCY = 16 -ROLLOUT_MAX_LOADED_LORAS = 32 -ROLLOUT_MAX_LORAS_PER_BATCH = 8 -ROLLOUT_LORA_TARGET_MODULES = ( - "q_proj", - "k_proj", - "v_proj", - "o_proj", - "gate_proj", - "up_proj", - "down_proj", - "lm_head", -) - -BULLETIN_ROOT = "/bulletin" -BULLETIN_VOLUME_NAME = "lilo-snapshot-bulletin" -app = modal.App(f"lilo-{DEFINITION_ID}") - -if modal.is_local(): - from ..miles_image import image -else: - image = modal.Image.debian_slim() - -assets = modal.Volume.from_name("lilo-model-assets", create_if_missing=True) -bulletin = modal.Volume.from_name( - BULLETIN_VOLUME_NAME, - create_if_missing=True, - version=2, -) -TRAINER_VOLUMES = { - "/assets": assets, - BULLETIN_ROOT: bulletin, - CHECKPOINT_ROOT: checkpoint_volume, -} -api_secret = modal.Secret.from_name( - "lilo-api", - required_keys=["TINKER_API_KEY"], -) -proxy_secret = modal.Secret.from_name( - "lilo-proxy", - required_keys=["MODAL_PROXY_TOKEN_ID", "MODAL_PROXY_TOKEN_SECRET"], -) - - -@app.function( - image=image, - gpu=f"{GPU_TYPE}:{GPUS}", - volumes=TRAINER_VOLUMES, - secrets=[api_secret, proxy_secret], - timeout=86_400, - max_containers=trainer_max_containers(), - env=trainer_deployment_env(), - single_use_containers=True, -) -def qwen3_5_9b_base_miles_lora_2k(instance_id: str) -> None: - import json - - from huggingface_hub import snapshot_download - - from lilo.providers.modal.kv import shared_kv - from lilo.providers.modal.serve import run_engine_with_backend - from modal.config import config - - if not os.path.exists(HF_CHECKPOINT): - snapshot_download(repo_id=MODEL_NAME, local_dir=HF_CHECKPOINT) - assets.commit() - backend_config = { - "miles": { - "hf_checkpoint": HF_CHECKPOINT, - "model_type": "qwen3.5-9B", - "actor_num_gpus_per_node": GPUS, - "tensor_model_parallel_size": TENSOR_MODEL_PARALLEL_SIZE, - "max_lora_slots": MAX_LORA_SLOTS, - "max_lora_rank": MAX_LORA_RANK, - "default_lora_alpha": DEFAULT_LORA_ALPHA, - "target_modules": TARGET_MODULES, - "max_tokens_per_gpu": MAX_CONTEXT_LENGTH, - "extra_args": ( - "--recompute-granularity", - "full", - "--recompute-method", - "uniform", - "--recompute-num-layers", - "1", - ), - }, - "checkpoint_dir": CHECKPOINT_ROOT, - } - run_engine_with_backend( - shared_kv(), - "lilo.backends.miles_lora:build_executor", - definition_id=DEFINITION_ID, - revision=config["image_id"], - instance_id=instance_id, - backend_env={ - "LILO_BACKEND_CONFIG": json.dumps(backend_config), - "LILO_BASE_MODEL": MODEL_NAME, - "LILO_DEFINITION_ID": DEFINITION_ID, - "LILO_CHECKPOINT_VOLUME": CHECKPOINT_VOLUME_NAME, - "LILO_BULLETIN_ROOT": BULLETIN_ROOT, - "LILO_BULLETIN_VOLUME": BULLETIN_VOLUME_NAME, - "LILO_DEFINITION_REVISION": config["image_id"], - "PYTORCH_CUDA_ALLOC_CONF": "expandable_segments:True", - "TORCHINDUCTOR_COMPILE_THREADS": "1", - }, - nproc=1, - max_models=MAX_LORA_SLOTS, - sampler_persistence_concurrency=8, - ) - - -ENGINE_FUNCTION = qwen3_5_9b_base_miles_lora_2k diff --git a/src/lilo/providers/modal/definitions/qwen3_5_9b_full_64k.py b/src/lilo/providers/modal/definitions/qwen3_5_9b_full_64k.py deleted file mode 100644 index 28dc5ac..0000000 --- a/src/lilo/providers/modal/definitions/qwen3_5_9b_full_64k.py +++ /dev/null @@ -1,144 +0,0 @@ -from __future__ import annotations - -import os - -import modal - -from ..checkpoint_storage import ( - CHECKPOINT_ROOT, - CHECKPOINT_VOLUME_NAME, - checkpoint_volume, -) -from ..deployment import trainer_max_containers - -MODEL_NAME = "Qwen/Qwen3.5-9B" -HF_CHECKPOINT = "/assets/Qwen3.5-9B" -MICRO_BATCH_SIZE = 1 -MAX_CONTEXT_LENGTH = 65_536 -MAX_TOKENS_PER_MICROBATCH = MAX_CONTEXT_LENGTH -DEFINITION_ID = "qwen3_5_9b_full_64k" -PARAMETERIZATION = "full" -CATALOG_VISIBLE = True -TRAINER_MODELS_PER_INSTANCE = 1 -GPU_TYPE = "H200" -GPUS = 4 -TENSOR_MODEL_PARALLEL_SIZE = 2 -CONTEXT_PARALLEL_SIZE = 2 -DATA_PARALLEL_SIZE = GPUS // (TENSOR_MODEL_PARALLEL_SIZE * CONTEXT_PARALLEL_SIZE) -MTP_NUM_LAYERS = 0 -RECOMPUTE_GRANULARITY = "full" -RECOMPUTE_METHOD = "uniform" -RECOMPUTE_NUM_LAYERS = 1 -LOSS_SCALE = 1.0 -DEFER_FP32_LOGITS = True -FP32_LM_HEAD = True -ROLLOUT_GPU_TYPE = "H200" -ROLLOUT_GPUS = 1 -ROLLOUT_TENSOR_PARALLEL_SIZE = 1 -ROLLOUT_EXPERT_PARALLEL_SIZE = 1 -ROLLOUT_EXPERT_TENSOR_PARALLEL_SIZE = 1 -ROLLOUT_MEMORY_FRACTION = 0.85 -ROLLOUT_MAX_RUNNING_REQUESTS = 32 -ROLLOUT_MAX_QUEUED_REQUESTS = 4 -ROLLOUT_TARGET_CONCURRENCY = 16 -ROLLOUT_CPU_WEIGHT_CACHE_MAX_COMPILE_GROUP_GB = 16 -BULLETIN_ROOT = "/bulletin" -BULLETIN_VOLUME_NAME = "lilo-snapshot-bulletin" -app = modal.App(f"lilo-{DEFINITION_ID}") - -if modal.is_local(): - from ..megatron_image import image -else: - image = modal.Image.debian_slim() - -assets = modal.Volume.from_name("lilo-model-assets", create_if_missing=True) -bulletin = modal.Volume.from_name( - BULLETIN_VOLUME_NAME, - create_if_missing=True, - version=2, -) -TRAINER_VOLUMES = { - "/assets": assets, - BULLETIN_ROOT: bulletin, - CHECKPOINT_ROOT: checkpoint_volume, -} -api_secret = modal.Secret.from_name( - "lilo-api", - required_keys=["TINKER_API_KEY"], -) -proxy_secret = modal.Secret.from_name( - "lilo-proxy", - required_keys=["MODAL_PROXY_TOKEN_ID", "MODAL_PROXY_TOKEN_SECRET"], -) - - -@app.function( - image=image, - gpu=f"{GPU_TYPE}:{GPUS}", - volumes=TRAINER_VOLUMES, - secrets=[api_secret, proxy_secret], - timeout=86_400, - max_containers=trainer_max_containers(), - single_use_containers=True, -) -def qwen3_5_9b_full_64k(instance_id: str) -> None: - import json - - from huggingface_hub import snapshot_download - from modal.config import config - - from lilo.providers.modal.kv import shared_kv - from lilo.providers.modal.serve import run_engine_with_backend - - if not os.path.exists(HF_CHECKPOINT): - snapshot_download(repo_id=MODEL_NAME, local_dir=HF_CHECKPOINT) - assets.commit() - backend_config = { - "megatron": { - "hf_checkpoint": HF_CHECKPOINT, - "tensor_model_parallel_size": TENSOR_MODEL_PARALLEL_SIZE, - "context_parallel_size": CONTEXT_PARALLEL_SIZE, - "sequence_parallel": True, - "micro_batch_size": MICRO_BATCH_SIZE, - "max_tokens_per_microbatch": MAX_TOKENS_PER_MICROBATCH, - "seq_length": MAX_CONTEXT_LENGTH, - "defer_fp32_logits": DEFER_FP32_LOGITS, - "fp32_lm_head": FP32_LM_HEAD, - "use_distributed_optimizer": True, - "provider_overrides": { - "mtp_num_layers": MTP_NUM_LAYERS, - "recompute_granularity": RECOMPUTE_GRANULARITY, - "recompute_method": RECOMPUTE_METHOD, - "recompute_num_layers": RECOMPUTE_NUM_LAYERS, - }, - "optimizer": { - "lr": 1e-4, - "min_lr": 1e-4, - "loss_scale": LOSS_SCALE, - }, - }, - "checkpoint_dir": CHECKPOINT_ROOT, - } - run_engine_with_backend( - shared_kv(), - "lilo.backends.megatron_fft:build_executor", - definition_id=DEFINITION_ID, - revision=config["image_id"], - instance_id=instance_id, - backend_env={ - "LILO_BACKEND_CONFIG": json.dumps(backend_config), - "LILO_BASE_MODEL": MODEL_NAME, - "LILO_DEFINITION_ID": DEFINITION_ID, - "LILO_CHECKPOINT_VOLUME": CHECKPOINT_VOLUME_NAME, - "LILO_BULLETIN_ROOT": BULLETIN_ROOT, - "LILO_BULLETIN_VOLUME": BULLETIN_VOLUME_NAME, - "LILO_DEFINITION_REVISION": config["image_id"], - "PYTORCH_CUDA_ALLOC_CONF": "expandable_segments:True", - "TORCHINDUCTOR_COMPILE_THREADS": "1", - }, - nproc=GPUS, - max_models=1, - ) - - -ENGINE_FUNCTION = qwen3_5_9b_full_64k diff --git a/src/lilo/providers/modal/definitions/qwen3_5_9b_miles_lora_16k.py b/src/lilo/providers/modal/definitions/qwen3_5_9b_miles_lora_16k.py deleted file mode 100644 index 1739fdf..0000000 --- a/src/lilo/providers/modal/definitions/qwen3_5_9b_miles_lora_16k.py +++ /dev/null @@ -1,170 +0,0 @@ -from __future__ import annotations - -import os - -import modal - -from ..checkpoint_storage import ( - CHECKPOINT_ROOT, - CHECKPOINT_VOLUME_NAME, - checkpoint_volume, -) -from ..deployment import trainer_max_containers - -MODEL_NAME = "Qwen/Qwen3.5-9B" -HF_CHECKPOINT = "/assets/Qwen3.5-9B" -DEFINITION_ID = "qwen3_5_9b_miles_lora_16k" -PARAMETERIZATION = "lora" -CATALOG_VISIBLE = True -MAX_CONTEXT_LENGTH = 16_384 - -GPU_TYPE = "H100" -GPUS = 8 -TENSOR_MODEL_PARALLEL_SIZE = 8 -MAX_LORA_SLOTS = 6 -MAX_LORA_RANK = 32 -DEFAULT_LORA_ALPHA = 32 -TARGET_MODULES = ( - "linear_qkv", - "linear_proj", - "linear_fc1", - "linear_fc2", -) -TRAINER_MODELS_PER_INSTANCE = MAX_LORA_SLOTS - -ROLLOUT_GPU_TYPE = "H200" -ROLLOUT_GPUS = 1 -ROLLOUT_TENSOR_PARALLEL_SIZE = 1 -ROLLOUT_EXPERT_PARALLEL_SIZE = 1 -ROLLOUT_EXPERT_TENSOR_PARALLEL_SIZE = 1 -ROLLOUT_MEMORY_FRACTION = 0.8 -ROLLOUT_MAX_RUNNING_REQUESTS = 32 -ROLLOUT_MAX_QUEUED_REQUESTS = 8 -ROLLOUT_TARGET_CONCURRENCY = 16 -ROLLOUT_MAX_LOADED_LORAS = 256 -ROLLOUT_MIN_CONTAINERS = 8 -ROLLOUT_MAX_CONTAINERS = 8 -ROLLOUT_MAX_LORAS_PER_BATCH = 8 -ROLLOUT_LORA_TARGET_MODULES = ( - "q_proj", - "k_proj", - "v_proj", - "o_proj", - "gate_proj", - "up_proj", - "down_proj", -) - -BULLETIN_ROOT = "/bulletin" -BULLETIN_VOLUME_NAME = "lilo-snapshot-bulletin" -app = modal.App(f"lilo-{DEFINITION_ID}") - -if modal.is_local(): - from ..miles_image import image -else: - image = modal.Image.debian_slim() - -assets = modal.Volume.from_name("lilo-model-assets", create_if_missing=True) -bulletin = modal.Volume.from_name( - BULLETIN_VOLUME_NAME, - create_if_missing=True, - version=2, -) -TRAINER_VOLUMES = { - "/assets": assets, - BULLETIN_ROOT: bulletin, - CHECKPOINT_ROOT: checkpoint_volume, -} -api_secret = modal.Secret.from_name( - "lilo-api", - required_keys=["TINKER_API_KEY"], -) -proxy_secret = modal.Secret.from_name( - "lilo-proxy", - required_keys=["MODAL_PROXY_TOKEN_ID", "MODAL_PROXY_TOKEN_SECRET"], -) - - -@app.function( - image=image, - gpu=f"{GPU_TYPE}:{GPUS}", - volumes=TRAINER_VOLUMES, - secrets=[api_secret, proxy_secret], - timeout=86_400, - max_containers=trainer_max_containers(), - single_use_containers=True, -) -def qwen3_5_9b_miles_lora_16k(instance_id: str) -> None: - run_trainer(instance_id) - - -def run_trainer( - instance_id: str, - *, - definition_id: str = DEFINITION_ID, - max_models: int = MAX_LORA_SLOTS, - deterministic_training: bool = False, -) -> None: - import json - - from huggingface_hub import snapshot_download - from modal.config import config - - from lilo.providers.modal.kv import shared_kv - from lilo.providers.modal.serve import run_engine_with_backend - - if not os.path.exists(HF_CHECKPOINT): - snapshot_download(repo_id=MODEL_NAME, local_dir=HF_CHECKPOINT) - assets.commit() - backend_config = { - "miles": { - "hf_checkpoint": HF_CHECKPOINT, - "model_type": "qwen3.5-9B", - "actor_num_gpus_per_node": GPUS, - "tensor_model_parallel_size": TENSOR_MODEL_PARALLEL_SIZE, - "max_lora_slots": MAX_LORA_SLOTS, - "max_lora_rank": MAX_LORA_RANK, - "default_lora_alpha": DEFAULT_LORA_ALPHA, - "target_modules": TARGET_MODULES, - "max_tokens_per_gpu": MAX_CONTEXT_LENGTH, - "extra_args": ( - "--seq-length", - str(MAX_CONTEXT_LENGTH), - "--recompute-granularity", - "full", - "--recompute-method", - "uniform", - "--recompute-num-layers", - "1", - ), - }, - "checkpoint_dir": CHECKPOINT_ROOT, - } - if deterministic_training: - backend_config["miles"].update( - tp_reduce_precision="float64", deterministic_attention=True - ) - run_engine_with_backend( - shared_kv(), - "lilo.backends.miles_lora:build_executor", - definition_id=definition_id, - revision=config["image_id"], - instance_id=instance_id, - backend_env={ - "LILO_BACKEND_CONFIG": json.dumps(backend_config), - "LILO_BASE_MODEL": MODEL_NAME, - "LILO_DEFINITION_ID": definition_id, - "LILO_CHECKPOINT_VOLUME": CHECKPOINT_VOLUME_NAME, - "LILO_BULLETIN_ROOT": BULLETIN_ROOT, - "LILO_BULLETIN_VOLUME": BULLETIN_VOLUME_NAME, - "LILO_DEFINITION_REVISION": config["image_id"], - "PYTORCH_CUDA_ALLOC_CONF": "expandable_segments:True", - "TORCHINDUCTOR_COMPILE_THREADS": "1", - }, - nproc=1, - max_models=max_models, - sampler_persistence_concurrency=8, - ) - - -ENGINE_FUNCTION = qwen3_5_9b_miles_lora_16k diff --git a/src/lilo/providers/modal/definitions/qwen3_5_9b_miles_lora_16k_dp2.py b/src/lilo/providers/modal/definitions/qwen3_5_9b_miles_lora_16k_dp2.py deleted file mode 100644 index 0a2e569..0000000 --- a/src/lilo/providers/modal/definitions/qwen3_5_9b_miles_lora_16k_dp2.py +++ /dev/null @@ -1,170 +0,0 @@ -from __future__ import annotations - -import os - -import modal - -from ..checkpoint_storage import ( - CHECKPOINT_ROOT, - CHECKPOINT_VOLUME_NAME, - checkpoint_volume, -) -from ..deployment import trainer_max_containers - -MODEL_NAME = "Qwen/Qwen3.5-9B" -HF_CHECKPOINT = "/assets/Qwen3.5-9B" -DEFINITION_ID = "qwen3_5_9b_miles_lora_16k_dp2" -PARAMETERIZATION = "lora" -CATALOG_VISIBLE = False -MAX_CONTEXT_LENGTH = 16_384 - -GPU_TYPE = "H100" -GPUS = 8 -TENSOR_MODEL_PARALLEL_SIZE = 4 -MAX_LORA_SLOTS = 6 -MAX_LORA_RANK = 32 -DEFAULT_LORA_ALPHA = 32 -TARGET_MODULES = ( - "linear_qkv", - "linear_proj", - "linear_fc1", - "linear_fc2", -) -TRAINER_MODELS_PER_INSTANCE = MAX_LORA_SLOTS - -ROLLOUT_GPU_TYPE = "H200" -ROLLOUT_GPUS = 1 -ROLLOUT_TENSOR_PARALLEL_SIZE = 1 -ROLLOUT_EXPERT_PARALLEL_SIZE = 1 -ROLLOUT_EXPERT_TENSOR_PARALLEL_SIZE = 1 -ROLLOUT_MEMORY_FRACTION = 0.8 -ROLLOUT_MAX_RUNNING_REQUESTS = 32 -ROLLOUT_MAX_QUEUED_REQUESTS = 8 -ROLLOUT_TARGET_CONCURRENCY = 16 -ROLLOUT_MAX_LOADED_LORAS = 256 -ROLLOUT_MIN_CONTAINERS = 8 -ROLLOUT_MAX_CONTAINERS = 8 -ROLLOUT_MAX_LORAS_PER_BATCH = 8 -ROLLOUT_LORA_TARGET_MODULES = ( - "q_proj", - "k_proj", - "v_proj", - "o_proj", - "gate_proj", - "up_proj", - "down_proj", -) - -BULLETIN_ROOT = "/bulletin" -BULLETIN_VOLUME_NAME = "lilo-snapshot-bulletin" -app = modal.App(f"lilo-{DEFINITION_ID}") - -if modal.is_local(): - from ..miles_image import image -else: - image = modal.Image.debian_slim() - -assets = modal.Volume.from_name("lilo-model-assets", create_if_missing=True) -bulletin = modal.Volume.from_name( - BULLETIN_VOLUME_NAME, - create_if_missing=True, - version=2, -) -TRAINER_VOLUMES = { - "/assets": assets, - BULLETIN_ROOT: bulletin, - CHECKPOINT_ROOT: checkpoint_volume, -} -api_secret = modal.Secret.from_name( - "lilo-api", - required_keys=["TINKER_API_KEY"], -) -proxy_secret = modal.Secret.from_name( - "lilo-proxy", - required_keys=["MODAL_PROXY_TOKEN_ID", "MODAL_PROXY_TOKEN_SECRET"], -) - - -@app.function( - image=image, - gpu=f"{GPU_TYPE}:{GPUS}", - volumes=TRAINER_VOLUMES, - secrets=[api_secret, proxy_secret], - timeout=86_400, - max_containers=trainer_max_containers(), - single_use_containers=True, -) -def qwen3_5_9b_miles_lora_16k_dp2(instance_id: str) -> None: - run_trainer(instance_id) - - -def run_trainer( - instance_id: str, - *, - definition_id: str = DEFINITION_ID, - max_models: int = MAX_LORA_SLOTS, - deterministic_training: bool = False, -) -> None: - import json - - from huggingface_hub import snapshot_download - from modal.config import config - - from lilo.providers.modal.kv import shared_kv - from lilo.providers.modal.serve import run_engine_with_backend - - if not os.path.exists(HF_CHECKPOINT): - snapshot_download(repo_id=MODEL_NAME, local_dir=HF_CHECKPOINT) - assets.commit() - backend_config = { - "miles": { - "hf_checkpoint": HF_CHECKPOINT, - "model_type": "qwen3.5-9B", - "actor_num_gpus_per_node": GPUS, - "tensor_model_parallel_size": TENSOR_MODEL_PARALLEL_SIZE, - "max_lora_slots": MAX_LORA_SLOTS, - "max_lora_rank": MAX_LORA_RANK, - "default_lora_alpha": DEFAULT_LORA_ALPHA, - "target_modules": TARGET_MODULES, - "max_tokens_per_gpu": MAX_CONTEXT_LENGTH, - "extra_args": ( - "--seq-length", - str(MAX_CONTEXT_LENGTH), - "--recompute-granularity", - "full", - "--recompute-method", - "uniform", - "--recompute-num-layers", - "1", - ), - }, - "checkpoint_dir": CHECKPOINT_ROOT, - } - if deterministic_training: - backend_config["miles"].update( - tp_reduce_precision="float64", deterministic_attention=True - ) - run_engine_with_backend( - shared_kv(), - "lilo.backends.miles_lora:build_executor", - definition_id=definition_id, - revision=config["image_id"], - instance_id=instance_id, - backend_env={ - "LILO_BACKEND_CONFIG": json.dumps(backend_config), - "LILO_BASE_MODEL": MODEL_NAME, - "LILO_DEFINITION_ID": definition_id, - "LILO_CHECKPOINT_VOLUME": CHECKPOINT_VOLUME_NAME, - "LILO_BULLETIN_ROOT": BULLETIN_ROOT, - "LILO_BULLETIN_VOLUME": BULLETIN_VOLUME_NAME, - "LILO_DEFINITION_REVISION": config["image_id"], - "PYTORCH_CUDA_ALLOC_CONF": "expandable_segments:True", - "TORCHINDUCTOR_COMPILE_THREADS": "1", - }, - nproc=1, - max_models=max_models, - sampler_persistence_concurrency=8, - ) - - -ENGINE_FUNCTION = qwen3_5_9b_miles_lora_16k_dp2 diff --git a/src/lilo/providers/modal/definitions/qwen3_6_27b_full_64k.py b/src/lilo/providers/modal/definitions/qwen3_6_27b_full_64k.py deleted file mode 100644 index 4b4067c..0000000 --- a/src/lilo/providers/modal/definitions/qwen3_6_27b_full_64k.py +++ /dev/null @@ -1,143 +0,0 @@ -from __future__ import annotations - -import os - -import modal - -from ..checkpoint_storage import ( - CHECKPOINT_ROOT, - CHECKPOINT_VOLUME_NAME, - checkpoint_volume, -) -from ..deployment import trainer_max_containers - -MODEL_NAME = "Qwen/Qwen3.6-27B" -HF_CHECKPOINT = "/assets/Qwen3.6-27B" -MICRO_BATCH_SIZE = 1 -MAX_CONTEXT_LENGTH = 65_536 -MAX_TOKENS_PER_MICROBATCH = MAX_CONTEXT_LENGTH -DEFINITION_ID = "qwen3_6_27b_full_64k" -PARAMETERIZATION = "full" -CATALOG_VISIBLE = True -TRAINER_MODELS_PER_INSTANCE = 1 -GPU_TYPE = "H200" -GPUS = 8 -TENSOR_MODEL_PARALLEL_SIZE = 4 -PIPELINE_MODEL_PARALLEL_SIZE = 1 -CONTEXT_PARALLEL_SIZE = 2 -DATA_PARALLEL_SIZE = 1 -SEQUENCE_PARALLEL = True -GPU_MEMORY_FRACTION = 0.90 -MTP_NUM_LAYERS = 0 -ROLLOUT_GPU_TYPE = "H200" -ROLLOUT_GPUS = 4 -ROLLOUT_TENSOR_PARALLEL_SIZE = 4 -ROLLOUT_EXPERT_PARALLEL_SIZE = 1 -ROLLOUT_EXPERT_TENSOR_PARALLEL_SIZE = 1 -ROLLOUT_MEMORY_FRACTION = 0.90 -ROLLOUT_CPU_WEIGHT_CACHE_MAX_COMPILE_GROUP_GB = 32 -ROLLOUT_MAX_RUNNING_REQUESTS = 32 -ROLLOUT_MAX_QUEUED_REQUESTS = 4 -ROLLOUT_TARGET_CONCURRENCY = 16 -BULLETIN_ROOT = "/bulletin" -BULLETIN_VOLUME_NAME = "lilo-snapshot-bulletin" -app = modal.App(f"lilo-{DEFINITION_ID}") - -if modal.is_local(): - from ..megatron_image import image -else: - image = modal.Image.debian_slim() - -assets = modal.Volume.from_name("lilo-model-assets", create_if_missing=True) -bulletin = modal.Volume.from_name( - BULLETIN_VOLUME_NAME, - create_if_missing=True, - version=2, -) -TRAINER_VOLUMES = { - "/assets": assets, - BULLETIN_ROOT: bulletin, - CHECKPOINT_ROOT: checkpoint_volume, -} -api_secret = modal.Secret.from_name( - "lilo-api", - required_keys=["TINKER_API_KEY"], -) -proxy_secret = modal.Secret.from_name( - "lilo-proxy", - required_keys=["MODAL_PROXY_TOKEN_ID", "MODAL_PROXY_TOKEN_SECRET"], -) - - -@app.function( - image=image, - gpu=f"{GPU_TYPE}:{GPUS}", - volumes=TRAINER_VOLUMES, - secrets=[api_secret, proxy_secret], - timeout=86_400, - max_containers=trainer_max_containers(), - single_use_containers=True, -) -def qwen3_6_27b_full_64k(instance_id: str) -> None: - import json - - from huggingface_hub import snapshot_download - from modal.config import config - - from lilo.providers.modal.kv import shared_kv - from lilo.providers.modal.serve import run_engine_with_backend - - if not os.path.exists(HF_CHECKPOINT): - snapshot_download(repo_id=MODEL_NAME, local_dir=HF_CHECKPOINT) - assets.commit() - backend_config = { - "megatron": { - "hf_checkpoint": HF_CHECKPOINT, - "tensor_model_parallel_size": TENSOR_MODEL_PARALLEL_SIZE, - "pipeline_model_parallel_size": PIPELINE_MODEL_PARALLEL_SIZE, - "context_parallel_size": CONTEXT_PARALLEL_SIZE, - "sequence_parallel": SEQUENCE_PARALLEL, - "micro_batch_size": MICRO_BATCH_SIZE, - "max_tokens_per_microbatch": MAX_TOKENS_PER_MICROBATCH, - "seq_length": MAX_CONTEXT_LENGTH, - "bf16": True, - "fp16": False, - "gpu_memory_fraction": GPU_MEMORY_FRACTION, - "use_distributed_optimizer": True, - "provider_overrides": { - "mtp_num_layers": MTP_NUM_LAYERS, - "recompute_granularity": "full", - "recompute_method": "uniform", - "recompute_num_layers": 1, - }, - "optimizer": { - "optimizer": "adam", - "lr": 1e-4, - "min_lr": 1e-4, - }, - }, - "checkpoint_dir": CHECKPOINT_ROOT, - } - run_engine_with_backend( - shared_kv(), - "lilo.backends.megatron_fft:build_executor", - definition_id=DEFINITION_ID, - revision=config["image_id"], - instance_id=instance_id, - backend_env={ - "PYTORCH_CUDA_ALLOC_CONF": "expandable_segments:True", - "TORCHINDUCTOR_COMPILE_THREADS": "1", - "LILO_BACKEND_CONFIG": json.dumps(backend_config), - "LILO_BASE_MODEL": MODEL_NAME, - "LILO_DEFINITION_ID": DEFINITION_ID, - "LILO_CHECKPOINT_VOLUME": CHECKPOINT_VOLUME_NAME, - "LILO_BULLETIN_ROOT": BULLETIN_ROOT, - "LILO_BULLETIN_VOLUME": BULLETIN_VOLUME_NAME, - "LILO_DEFINITION_REVISION": config["image_id"], - }, - nproc=GPUS, - max_models=1, - ) - - -ENGINE_FUNCTION = qwen3_6_27b_full_64k diff --git a/src/lilo/providers/modal/definitions/qwen3_6_35b_a3b_full_64k.py b/src/lilo/providers/modal/definitions/qwen3_6_35b_a3b_full_64k.py deleted file mode 100644 index ea585ca..0000000 --- a/src/lilo/providers/modal/definitions/qwen3_6_35b_a3b_full_64k.py +++ /dev/null @@ -1,152 +0,0 @@ -from __future__ import annotations - -import os - -import modal - -from ..checkpoint_storage import ( - CHECKPOINT_ROOT, - CHECKPOINT_VOLUME_NAME, - checkpoint_volume, -) -from ..deployment import trainer_max_containers - -MODEL_NAME = "Qwen/Qwen3.6-35B-A3B" -HF_CHECKPOINT = "/assets/Qwen3.6-35B-A3B" -MICRO_BATCH_SIZE = 1 -MAX_CONTEXT_LENGTH = 65_536 -MAX_TOKENS_PER_MICROBATCH = MAX_CONTEXT_LENGTH -DEFINITION_ID = "qwen3_6_35b_a3b_full_64k" -PARAMETERIZATION = "full" -CATALOG_VISIBLE = True -TRAINER_MODELS_PER_INSTANCE = 1 -GPU_TYPE = "H200" -GPUS = 8 -TENSOR_MODEL_PARALLEL_SIZE = 4 -PIPELINE_MODEL_PARALLEL_SIZE = 1 -CONTEXT_PARALLEL_SIZE = 2 -DATA_PARALLEL_SIZE = 1 -EXPERT_MODEL_PARALLEL_SIZE = 8 -EXPERT_TENSOR_PARALLEL_SIZE = 1 -SEQUENCE_PARALLEL = True -GPU_MEMORY_FRACTION = 0.90 -MTP_NUM_LAYERS = 0 -ROLLOUT_GPU_TYPE = "H200" -ROLLOUT_GPUS = 4 -ROLLOUT_TENSOR_PARALLEL_SIZE = 1 -ROLLOUT_EXPERT_PARALLEL_SIZE = 4 -ROLLOUT_EXPERT_TENSOR_PARALLEL_SIZE = 1 -ROLLOUT_MEMORY_FRACTION = 0.90 -ROLLOUT_CPU_WEIGHT_CACHE_MAX_COMPILE_GROUP_GB = 32 -ROLLOUT_MAX_RUNNING_REQUESTS = 32 -ROLLOUT_MAX_QUEUED_REQUESTS = 4 -ROLLOUT_TARGET_CONCURRENCY = 16 -BULLETIN_ROOT = "/bulletin" -BULLETIN_VOLUME_NAME = "lilo-snapshot-bulletin" -app = modal.App(f"lilo-{DEFINITION_ID}") - -if modal.is_local(): - from ..megatron_image import image -else: - image = modal.Image.debian_slim() - -assets = modal.Volume.from_name("lilo-model-assets", create_if_missing=True) -bulletin = modal.Volume.from_name( - BULLETIN_VOLUME_NAME, - create_if_missing=True, - version=2, -) -TRAINER_VOLUMES = { - "/assets": assets, - BULLETIN_ROOT: bulletin, - CHECKPOINT_ROOT: checkpoint_volume, -} -api_secret = modal.Secret.from_name( - "lilo-api", - required_keys=["TINKER_API_KEY"], -) -proxy_secret = modal.Secret.from_name( - "lilo-proxy", - required_keys=["MODAL_PROXY_TOKEN_ID", "MODAL_PROXY_TOKEN_SECRET"], -) - - -@app.function( - image=image, - gpu=f"{GPU_TYPE}:{GPUS}", - volumes=TRAINER_VOLUMES, - secrets=[api_secret, proxy_secret], - timeout=86_400, - max_containers=trainer_max_containers(), - single_use_containers=True, -) -def qwen3_6_35b_a3b_full_64k(instance_id: str) -> None: - import json - - from huggingface_hub import snapshot_download - from modal.config import config - - from lilo.providers.modal.kv import shared_kv - from lilo.providers.modal.serve import run_engine_with_backend - - if not os.path.exists(HF_CHECKPOINT): - snapshot_download(repo_id=MODEL_NAME, local_dir=HF_CHECKPOINT) - assets.commit() - backend_config = { - "megatron": { - "hf_checkpoint": HF_CHECKPOINT, - "tensor_model_parallel_size": TENSOR_MODEL_PARALLEL_SIZE, - "pipeline_model_parallel_size": PIPELINE_MODEL_PARALLEL_SIZE, - "context_parallel_size": CONTEXT_PARALLEL_SIZE, - "expert_model_parallel_size": EXPERT_MODEL_PARALLEL_SIZE, - "expert_tensor_parallel_size": EXPERT_TENSOR_PARALLEL_SIZE, - "sequence_parallel": SEQUENCE_PARALLEL, - "micro_batch_size": MICRO_BATCH_SIZE, - "max_tokens_per_microbatch": MAX_TOKENS_PER_MICROBATCH, - "seq_length": MAX_CONTEXT_LENGTH, - "bf16": True, - "fp16": False, - "gpu_memory_fraction": GPU_MEMORY_FRACTION, - "use_distributed_optimizer": True, - "provider_overrides": { - "mtp_num_layers": MTP_NUM_LAYERS, - "recompute_granularity": "selective", - "moe_layer_recompute": True, - "moe_token_dispatcher_type": "alltoall", - "moe_router_fusion": True, - "moe_permute_fusion": True, - "moe_grouped_gemm": True, - "moe_shared_expert_overlap": False, - "moe_aux_loss_coeff": 0.0, - }, - "optimizer": { - "optimizer": "adam", - "lr": 1e-4, - "min_lr": 1e-4, - }, - }, - "checkpoint_dir": CHECKPOINT_ROOT, - } - run_engine_with_backend( - shared_kv(), - "lilo.backends.megatron_fft:build_executor", - definition_id=DEFINITION_ID, - revision=config["image_id"], - instance_id=instance_id, - backend_env={ - "PYTORCH_CUDA_ALLOC_CONF": "expandable_segments:True", - "TORCHINDUCTOR_COMPILE_THREADS": "1", - "LILO_BACKEND_CONFIG": json.dumps(backend_config), - "LILO_BASE_MODEL": MODEL_NAME, - "LILO_DEFINITION_ID": DEFINITION_ID, - "LILO_CHECKPOINT_VOLUME": CHECKPOINT_VOLUME_NAME, - "LILO_BULLETIN_ROOT": BULLETIN_ROOT, - "LILO_BULLETIN_VOLUME": BULLETIN_VOLUME_NAME, - "LILO_DEFINITION_REVISION": config["image_id"], - }, - nproc=GPUS, - max_models=1, - ) - - -ENGINE_FUNCTION = qwen3_6_35b_a3b_full_64k diff --git a/src/lilo/providers/modal/definitions/qwen3_8_27b_miles_lora_128k.py b/src/lilo/providers/modal/definitions/qwen3_8_27b_miles_lora_128k.py deleted file mode 100644 index 192af5e..0000000 --- a/src/lilo/providers/modal/definitions/qwen3_8_27b_miles_lora_128k.py +++ /dev/null @@ -1,174 +0,0 @@ -from __future__ import annotations - -import os - -import modal - -from ..checkpoint_storage import ( - CHECKPOINT_ROOT, - CHECKPOINT_VOLUME_NAME, - checkpoint_volume, -) -from ..deployment import trainer_deployment_env, trainer_max_containers - -MODEL_NAME = "Qwen/Qwen3.8-27B" -HF_CHECKPOINT = "/assets/Qwen3.8-27B" -DEFINITION_ID = "qwen3_8_27b_miles_lora_128k" -PARAMETERIZATION = "lora" -CATALOG_VISIBLE = False -MAX_CONTEXT_LENGTH = 131_072 - -GPU_TYPE = "H200" -GPUS = 8 -TENSOR_MODEL_PARALLEL_SIZE = 2 -CONTEXT_PARALLEL_SIZE = 4 -MAX_LORA_SLOTS = 6 -MAX_LORA_RANK = 32 -DEFAULT_LORA_ALPHA = 32 -TARGET_MODULES = ( - "linear_qkv", - "linear_proj", - "linear_fc1", - "linear_fc2", -) -TRAINER_MODELS_PER_INSTANCE = MAX_LORA_SLOTS - -ROLLOUT_GPU_TYPE = "H200" -ROLLOUT_GPUS = 2 -ROLLOUT_TENSOR_PARALLEL_SIZE = 2 -ROLLOUT_EXPERT_PARALLEL_SIZE = 1 -ROLLOUT_EXPERT_TENSOR_PARALLEL_SIZE = 1 -ROLLOUT_MEMORY_FRACTION = 0.8 -ROLLOUT_MAX_RUNNING_REQUESTS = 8 -ROLLOUT_MAX_QUEUED_REQUESTS = 8 -ROLLOUT_TARGET_CONCURRENCY = 4 -ROLLOUT_MAX_LOADED_LORAS = 256 -ROLLOUT_MIN_CONTAINERS = 4 -ROLLOUT_MAX_CONTAINERS = 4 -ROLLOUT_MAX_LORAS_PER_BATCH = 8 -ROLLOUT_LORA_TARGET_MODULES = ( - "q_proj", - "k_proj", - "v_proj", - "o_proj", - "gate_proj", - "up_proj", - "down_proj", -) - -BULLETIN_ROOT = "/bulletin" -BULLETIN_VOLUME_NAME = "lilo-snapshot-bulletin" -app = modal.App(f"lilo-{DEFINITION_ID}") - -if modal.is_local(): - from ..miles_image import image -else: - image = modal.Image.debian_slim() - -assets = modal.Volume.from_name("lilo-model-assets", create_if_missing=True) -bulletin = modal.Volume.from_name( - BULLETIN_VOLUME_NAME, - create_if_missing=True, - version=2, -) -TRAINER_VOLUMES = { - "/assets": assets, - BULLETIN_ROOT: bulletin, - CHECKPOINT_ROOT: checkpoint_volume, -} -api_secret = modal.Secret.from_name( - "lilo-api", - required_keys=["TINKER_API_KEY"], -) -proxy_secret = modal.Secret.from_name( - "lilo-proxy", - required_keys=["MODAL_PROXY_TOKEN_ID", "MODAL_PROXY_TOKEN_SECRET"], -) -huggingface_secret = modal.Secret.from_name("huggingface-secret") - - -@app.function( - image=image, - gpu=f"{GPU_TYPE}:{GPUS}", - volumes=TRAINER_VOLUMES, - secrets=[api_secret, proxy_secret, huggingface_secret], - env=trainer_deployment_env(), - timeout=86_400, - max_containers=trainer_max_containers(), - single_use_containers=True, -) -def qwen3_8_27b_miles_lora_128k(instance_id: str) -> None: - run_trainer(instance_id) - - -def run_trainer( - instance_id: str, - *, - definition_id: str = DEFINITION_ID, - max_models: int = MAX_LORA_SLOTS, - deterministic_training: bool = False, -) -> None: - import json - - from huggingface_hub import snapshot_download - from modal.config import config - - from lilo.providers.modal.kv import shared_kv - from lilo.providers.modal.serve import run_engine_with_backend - - if not os.path.exists(HF_CHECKPOINT): - snapshot_download(repo_id=MODEL_NAME, local_dir=HF_CHECKPOINT) - assets.commit() - backend_config = { - "miles": { - "hf_checkpoint": HF_CHECKPOINT, - "model_type": "qwen3.8-27B", - "actor_num_gpus_per_node": GPUS, - "tensor_model_parallel_size": TENSOR_MODEL_PARALLEL_SIZE, - "context_parallel_size": CONTEXT_PARALLEL_SIZE, - "max_lora_slots": MAX_LORA_SLOTS, - "max_lora_rank": MAX_LORA_RANK, - "default_lora_alpha": DEFAULT_LORA_ALPHA, - "target_modules": TARGET_MODULES, - "max_tokens_per_gpu": MAX_CONTEXT_LENGTH // CONTEXT_PARALLEL_SIZE, - "extra_args": ( - "--seq-length", - str(MAX_CONTEXT_LENGTH), - "--recompute-granularity", - "full", - "--recompute-method", - "uniform", - "--recompute-num-layers", - "1", - ), - }, - "checkpoint_dir": CHECKPOINT_ROOT, - } - if deterministic_training: - backend_config["miles"].update( - tp_reduce_precision="float64", deterministic_attention=True - ) - run_engine_with_backend( - shared_kv(), - "lilo.backends.miles_lora:build_executor", - definition_id=definition_id, - revision=config["image_id"], - instance_id=instance_id, - backend_env={ - "LILO_BACKEND_CONFIG": json.dumps(backend_config), - "LILO_BASE_MODEL": MODEL_NAME, - "LILO_DEFINITION_ID": definition_id, - "LILO_CHECKPOINT_VOLUME": CHECKPOINT_VOLUME_NAME, - "LILO_BULLETIN_ROOT": BULLETIN_ROOT, - "LILO_BULLETIN_VOLUME": BULLETIN_VOLUME_NAME, - "LILO_DEFINITION_REVISION": config["image_id"], - "PYTORCH_CUDA_ALLOC_CONF": "expandable_segments:True", - "TORCHINDUCTOR_COMPILE_THREADS": "1", - }, - nproc=1, - max_models=max_models, - sampler_persistence_concurrency=8, - ) - - -ENGINE_FUNCTION = qwen3_8_27b_miles_lora_128k diff --git a/src/lilo/providers/modal/definitions/qwen3_8_27b_miles_lora_16k.py b/src/lilo/providers/modal/definitions/qwen3_8_27b_miles_lora_16k.py deleted file mode 100644 index 34ed6d0..0000000 --- a/src/lilo/providers/modal/definitions/qwen3_8_27b_miles_lora_16k.py +++ /dev/null @@ -1,172 +0,0 @@ -from __future__ import annotations - -import os - -import modal - -from ..checkpoint_storage import ( - CHECKPOINT_ROOT, - CHECKPOINT_VOLUME_NAME, - checkpoint_volume, -) -from ..deployment import trainer_deployment_env, trainer_max_containers - -MODEL_NAME = "Qwen/Qwen3.8-27B" -HF_CHECKPOINT = "/assets/Qwen3.8-27B" -DEFINITION_ID = "qwen3_8_27b_miles_lora_16k" -PARAMETERIZATION = "lora" -CATALOG_VISIBLE = True -MAX_CONTEXT_LENGTH = 16_384 - -GPU_TYPE = "H200" -GPUS = 8 -TENSOR_MODEL_PARALLEL_SIZE = 4 -MAX_LORA_SLOTS = 6 -MAX_LORA_RANK = 32 -DEFAULT_LORA_ALPHA = 32 -TARGET_MODULES = ( - "linear_qkv", - "linear_proj", - "linear_fc1", - "linear_fc2", -) -TRAINER_MODELS_PER_INSTANCE = MAX_LORA_SLOTS - -ROLLOUT_GPU_TYPE = "H200" -ROLLOUT_GPUS = 1 -ROLLOUT_TENSOR_PARALLEL_SIZE = 1 -ROLLOUT_EXPERT_PARALLEL_SIZE = 1 -ROLLOUT_EXPERT_TENSOR_PARALLEL_SIZE = 1 -ROLLOUT_MEMORY_FRACTION = 0.8 -ROLLOUT_MAX_RUNNING_REQUESTS = 32 -ROLLOUT_MAX_QUEUED_REQUESTS = 8 -ROLLOUT_TARGET_CONCURRENCY = 16 -ROLLOUT_MAX_LOADED_LORAS = 256 -ROLLOUT_MIN_CONTAINERS = 8 -ROLLOUT_MAX_CONTAINERS = 8 -ROLLOUT_MAX_LORAS_PER_BATCH = 8 -ROLLOUT_LORA_TARGET_MODULES = ( - "q_proj", - "k_proj", - "v_proj", - "o_proj", - "gate_proj", - "up_proj", - "down_proj", -) - -BULLETIN_ROOT = "/bulletin" -BULLETIN_VOLUME_NAME = "lilo-snapshot-bulletin" -app = modal.App(f"lilo-{DEFINITION_ID}") - -if modal.is_local(): - from ..miles_image import image -else: - image = modal.Image.debian_slim() - -assets = modal.Volume.from_name("lilo-model-assets", create_if_missing=True) -bulletin = modal.Volume.from_name( - BULLETIN_VOLUME_NAME, - create_if_missing=True, - version=2, -) -TRAINER_VOLUMES = { - "/assets": assets, - BULLETIN_ROOT: bulletin, - CHECKPOINT_ROOT: checkpoint_volume, -} -api_secret = modal.Secret.from_name( - "lilo-api", - required_keys=["TINKER_API_KEY"], -) -proxy_secret = modal.Secret.from_name( - "lilo-proxy", - required_keys=["MODAL_PROXY_TOKEN_ID", "MODAL_PROXY_TOKEN_SECRET"], -) -huggingface_secret = modal.Secret.from_name("huggingface-secret") - - -@app.function( - image=image, - gpu=f"{GPU_TYPE}:{GPUS}", - volumes=TRAINER_VOLUMES, - secrets=[api_secret, proxy_secret, huggingface_secret], - env=trainer_deployment_env(), - timeout=86_400, - max_containers=trainer_max_containers(), - single_use_containers=True, -) -def qwen3_8_27b_miles_lora_16k(instance_id: str) -> None: - run_trainer(instance_id) - - -def run_trainer( - instance_id: str, - *, - definition_id: str = DEFINITION_ID, - max_models: int = MAX_LORA_SLOTS, - deterministic_training: bool = False, -) -> None: - import json - - from huggingface_hub import snapshot_download - from modal.config import config - - from lilo.providers.modal.kv import shared_kv - from lilo.providers.modal.serve import run_engine_with_backend - - if not os.path.exists(HF_CHECKPOINT): - snapshot_download(repo_id=MODEL_NAME, local_dir=HF_CHECKPOINT) - assets.commit() - backend_config = { - "miles": { - "hf_checkpoint": HF_CHECKPOINT, - "model_type": "qwen3.8-27B", - "actor_num_gpus_per_node": GPUS, - "tensor_model_parallel_size": TENSOR_MODEL_PARALLEL_SIZE, - "max_lora_slots": MAX_LORA_SLOTS, - "max_lora_rank": MAX_LORA_RANK, - "default_lora_alpha": DEFAULT_LORA_ALPHA, - "target_modules": TARGET_MODULES, - "max_tokens_per_gpu": MAX_CONTEXT_LENGTH, - "extra_args": ( - "--seq-length", - str(MAX_CONTEXT_LENGTH), - "--recompute-granularity", - "full", - "--recompute-method", - "uniform", - "--recompute-num-layers", - "1", - ), - }, - "checkpoint_dir": CHECKPOINT_ROOT, - } - if deterministic_training: - backend_config["miles"].update( - tp_reduce_precision="float64", deterministic_attention=True - ) - run_engine_with_backend( - shared_kv(), - "lilo.backends.miles_lora:build_executor", - definition_id=definition_id, - revision=config["image_id"], - instance_id=instance_id, - backend_env={ - "LILO_BACKEND_CONFIG": json.dumps(backend_config), - "LILO_BASE_MODEL": MODEL_NAME, - "LILO_DEFINITION_ID": definition_id, - "LILO_CHECKPOINT_VOLUME": CHECKPOINT_VOLUME_NAME, - "LILO_BULLETIN_ROOT": BULLETIN_ROOT, - "LILO_BULLETIN_VOLUME": BULLETIN_VOLUME_NAME, - "LILO_DEFINITION_REVISION": config["image_id"], - "PYTORCH_CUDA_ALLOC_CONF": "expandable_segments:True", - "TORCHINDUCTOR_COMPILE_THREADS": "1", - }, - nproc=1, - max_models=max_models, - sampler_persistence_concurrency=8, - ) - - -ENGINE_FUNCTION = qwen3_8_27b_miles_lora_16k diff --git a/src/lilo/providers/modal/definitions/qwen3_8_27b_miles_lora_64k.py b/src/lilo/providers/modal/definitions/qwen3_8_27b_miles_lora_64k.py deleted file mode 100644 index 417ee85..0000000 --- a/src/lilo/providers/modal/definitions/qwen3_8_27b_miles_lora_64k.py +++ /dev/null @@ -1,174 +0,0 @@ -from __future__ import annotations - -import os - -import modal - -from ..checkpoint_storage import ( - CHECKPOINT_ROOT, - CHECKPOINT_VOLUME_NAME, - checkpoint_volume, -) -from ..deployment import trainer_deployment_env, trainer_max_containers - -MODEL_NAME = "Qwen/Qwen3.8-27B" -HF_CHECKPOINT = "/assets/Qwen3.8-27B" -DEFINITION_ID = "qwen3_8_27b_miles_lora_64k" -PARAMETERIZATION = "lora" -CATALOG_VISIBLE = False -MAX_CONTEXT_LENGTH = 65_536 - -GPU_TYPE = "H200" -GPUS = 8 -TENSOR_MODEL_PARALLEL_SIZE = 4 -CONTEXT_PARALLEL_SIZE = 2 -MAX_LORA_SLOTS = 6 -MAX_LORA_RANK = 32 -DEFAULT_LORA_ALPHA = 32 -TARGET_MODULES = ( - "linear_qkv", - "linear_proj", - "linear_fc1", - "linear_fc2", -) -TRAINER_MODELS_PER_INSTANCE = MAX_LORA_SLOTS - -ROLLOUT_GPU_TYPE = "H200" -ROLLOUT_GPUS = 1 -ROLLOUT_TENSOR_PARALLEL_SIZE = 1 -ROLLOUT_EXPERT_PARALLEL_SIZE = 1 -ROLLOUT_EXPERT_TENSOR_PARALLEL_SIZE = 1 -ROLLOUT_MEMORY_FRACTION = 0.8 -ROLLOUT_MAX_RUNNING_REQUESTS = 16 -ROLLOUT_MAX_QUEUED_REQUESTS = 8 -ROLLOUT_TARGET_CONCURRENCY = 8 -ROLLOUT_MAX_LOADED_LORAS = 256 -ROLLOUT_MIN_CONTAINERS = 8 -ROLLOUT_MAX_CONTAINERS = 8 -ROLLOUT_MAX_LORAS_PER_BATCH = 8 -ROLLOUT_LORA_TARGET_MODULES = ( - "q_proj", - "k_proj", - "v_proj", - "o_proj", - "gate_proj", - "up_proj", - "down_proj", -) - -BULLETIN_ROOT = "/bulletin" -BULLETIN_VOLUME_NAME = "lilo-snapshot-bulletin" -app = modal.App(f"lilo-{DEFINITION_ID}") - -if modal.is_local(): - from ..miles_image import image -else: - image = modal.Image.debian_slim() - -assets = modal.Volume.from_name("lilo-model-assets", create_if_missing=True) -bulletin = modal.Volume.from_name( - BULLETIN_VOLUME_NAME, - create_if_missing=True, - version=2, -) -TRAINER_VOLUMES = { - "/assets": assets, - BULLETIN_ROOT: bulletin, - CHECKPOINT_ROOT: checkpoint_volume, -} -api_secret = modal.Secret.from_name( - "lilo-api", - required_keys=["TINKER_API_KEY"], -) -proxy_secret = modal.Secret.from_name( - "lilo-proxy", - required_keys=["MODAL_PROXY_TOKEN_ID", "MODAL_PROXY_TOKEN_SECRET"], -) -huggingface_secret = modal.Secret.from_name("huggingface-secret") - - -@app.function( - image=image, - gpu=f"{GPU_TYPE}:{GPUS}", - volumes=TRAINER_VOLUMES, - secrets=[api_secret, proxy_secret, huggingface_secret], - env=trainer_deployment_env(), - timeout=86_400, - max_containers=trainer_max_containers(), - single_use_containers=True, -) -def qwen3_8_27b_miles_lora_64k(instance_id: str) -> None: - run_trainer(instance_id) - - -def run_trainer( - instance_id: str, - *, - definition_id: str = DEFINITION_ID, - max_models: int = MAX_LORA_SLOTS, - deterministic_training: bool = False, -) -> None: - import json - - from huggingface_hub import snapshot_download - from modal.config import config - - from lilo.providers.modal.kv import shared_kv - from lilo.providers.modal.serve import run_engine_with_backend - - if not os.path.exists(HF_CHECKPOINT): - snapshot_download(repo_id=MODEL_NAME, local_dir=HF_CHECKPOINT) - assets.commit() - backend_config = { - "miles": { - "hf_checkpoint": HF_CHECKPOINT, - "model_type": "qwen3.8-27B", - "actor_num_gpus_per_node": GPUS, - "tensor_model_parallel_size": TENSOR_MODEL_PARALLEL_SIZE, - "context_parallel_size": CONTEXT_PARALLEL_SIZE, - "max_lora_slots": MAX_LORA_SLOTS, - "max_lora_rank": MAX_LORA_RANK, - "default_lora_alpha": DEFAULT_LORA_ALPHA, - "target_modules": TARGET_MODULES, - "max_tokens_per_gpu": MAX_CONTEXT_LENGTH // CONTEXT_PARALLEL_SIZE, - "extra_args": ( - "--seq-length", - str(MAX_CONTEXT_LENGTH), - "--recompute-granularity", - "full", - "--recompute-method", - "uniform", - "--recompute-num-layers", - "1", - ), - }, - "checkpoint_dir": CHECKPOINT_ROOT, - } - if deterministic_training: - backend_config["miles"].update( - tp_reduce_precision="float64", deterministic_attention=True - ) - run_engine_with_backend( - shared_kv(), - "lilo.backends.miles_lora:build_executor", - definition_id=definition_id, - revision=config["image_id"], - instance_id=instance_id, - backend_env={ - "LILO_BACKEND_CONFIG": json.dumps(backend_config), - "LILO_BASE_MODEL": MODEL_NAME, - "LILO_DEFINITION_ID": definition_id, - "LILO_CHECKPOINT_VOLUME": CHECKPOINT_VOLUME_NAME, - "LILO_BULLETIN_ROOT": BULLETIN_ROOT, - "LILO_BULLETIN_VOLUME": BULLETIN_VOLUME_NAME, - "LILO_DEFINITION_REVISION": config["image_id"], - "PYTORCH_CUDA_ALLOC_CONF": "expandable_segments:True", - "TORCHINDUCTOR_COMPILE_THREADS": "1", - }, - nproc=1, - max_models=max_models, - sampler_persistence_concurrency=8, - ) - - -ENGINE_FUNCTION = qwen3_8_27b_miles_lora_64k diff --git a/src/lilo/providers/modal/deployment.py b/src/lilo/providers/modal/deployment.py index a91ab02..b95b4ef 100644 --- a/src/lilo/providers/modal/deployment.py +++ b/src/lilo/providers/modal/deployment.py @@ -1,26 +1,9 @@ import os -TRAINER_MAX_CONTAINERS_ENV = "LILO_TRAINER_MAX_CONTAINERS" APP_NAME_ENV = "LILO_APP_NAME" -def trainer_max_containers() -> int | None: - value = os.environ.get(TRAINER_MAX_CONTAINERS_ENV) - if value is None: - return None - try: - limit = int(value) - except ValueError as exc: - raise ValueError( - f"{TRAINER_MAX_CONTAINERS_ENV} must be a positive integer" - ) from exc - if limit < 1: - raise ValueError(f"{TRAINER_MAX_CONTAINERS_ENV} must be a positive integer") - return limit - - FORWARDED_DEPLOYMENT_ENVS = ( - TRAINER_MAX_CONTAINERS_ENV, APP_NAME_ENV, "LILO_TORCH_PROFILE_STEP", "LILO_TORCH_PROFILE_DIR", diff --git a/src/lilo/providers/modal/fft_pool.py b/src/lilo/providers/modal/fft_pool.py index ee67089..e1b1865 100644 --- a/src/lilo/providers/modal/fft_pool.py +++ b/src/lilo/providers/modal/fft_pool.py @@ -139,9 +139,7 @@ def deploy_pool(spec: FFTPoolSpec) -> str: modal_cli, "deploy", "-m", - "lilo.providers.modal.yaml_pool_app" - if recipe_env - else "lilo.providers.modal.fft_pool_app", + "lilo.providers.modal.yaml_pool_app", "--name", spec.app_name, ] diff --git a/src/lilo/providers/modal/fft_pool_app.py b/src/lilo/providers/modal/fft_pool_app.py deleted file mode 100644 index 036b5ad..0000000 --- a/src/lilo/providers/modal/fft_pool_app.py +++ /dev/null @@ -1,144 +0,0 @@ -from __future__ import annotations - -import importlib -import os - -import modal - -from lilo.inference.serving import ( - start_fft_sidecar, - start_sglang, - supervise_children, - terminate, - wait_http, -) - -from .rollout_image import image - -APP_NAME = os.environ["LILO_FFT_POOL_APP_NAME"] -DEFINITION_ID = os.environ["LILO_FFT_POOL_DEFINITION_ID"] -MODEL_ID = os.environ["LILO_FFT_POOL_MODEL_ID"] -LATEST = os.environ["LILO_FFT_POOL_LATEST"] == "1" -VERSION = int(os.environ["LILO_FFT_POOL_VERSION"]) -definition = importlib.import_module( - f"lilo.providers.modal.definitions.{DEFINITION_ID}" -) -ROLLOUT_GPU_TYPE = definition.ROLLOUT_GPU_TYPE -ROLLOUT_GPUS = definition.ROLLOUT_GPUS -ROLLOUT_TENSOR_PARALLEL_SIZE = definition.ROLLOUT_TENSOR_PARALLEL_SIZE -ROLLOUT_EXPERT_PARALLEL_SIZE = definition.ROLLOUT_EXPERT_PARALLEL_SIZE -ROLLOUT_EXPERT_TENSOR_PARALLEL_SIZE = definition.ROLLOUT_EXPERT_TENSOR_PARALLEL_SIZE -ROLLOUT_MEMORY_FRACTION = definition.ROLLOUT_MEMORY_FRACTION -ROLLOUT_MAX_RUNNING_REQUESTS = definition.ROLLOUT_MAX_RUNNING_REQUESTS -ROLLOUT_MAX_QUEUED_REQUESTS = definition.ROLLOUT_MAX_QUEUED_REQUESTS -ROLLOUT_TARGET_CONCURRENCY = definition.ROLLOUT_TARGET_CONCURRENCY - - -def _pool_setting(name: str, default: int | None) -> int | None: - value = os.environ.get(f"LILO_FFT_POOL_{name}") - return default if value is None else int(value) - - -ROLLOUT_MIN_CONTAINERS = _pool_setting( - "MIN_CONTAINERS", - getattr(definition, "ROLLOUT_MIN_CONTAINERS", None), -) -ROLLOUT_MAX_CONTAINERS = _pool_setting( - "MAX_CONTAINERS", - getattr(definition, "ROLLOUT_MAX_CONTAINERS", None), -) -ROLLOUT_SCALEDOWN_WINDOW = _pool_setting( - "SCALEDOWN_WINDOW", - getattr(definition, "ROLLOUT_SCALEDOWN_WINDOW", 5 * 60), -) -ROLLOUT_EXIT_GRACE_PERIOD = getattr( - definition, - "ROLLOUT_EXIT_GRACE_PERIOD", - 5 * 60, -) -ROLLOUT_CPU_WEIGHT_CACHE_MAX_COMPILE_GROUP_GB = ( - definition.ROLLOUT_CPU_WEIGHT_CACHE_MAX_COMPILE_GROUP_GB -) -SGLANG_PORT = 8001 -SIDECAR_PORT = 8000 - -pool_secret = modal.Secret.from_dict( - { - key: value - for key, value in os.environ.items() - if key.startswith("LILO_FFT_POOL_") - } -) -api_secret = modal.Secret.from_name( - "lilo-api", - required_keys=["TINKER_API_KEY"], -) -app = modal.App(APP_NAME) - - -@app.server( - image=image, - gpu=f"{ROLLOUT_GPU_TYPE}:{ROLLOUT_GPUS}", - volumes={ - "/assets": definition.assets, - definition.BULLETIN_ROOT: definition.bulletin, - }, - secrets=[api_secret, pool_secret], - target_concurrency=ROLLOUT_TARGET_CONCURRENCY, - min_containers=ROLLOUT_MIN_CONTAINERS, - max_containers=ROLLOUT_MAX_CONTAINERS, - scaledown_window=ROLLOUT_SCALEDOWN_WINDOW, - startup_timeout=20 * 60, - exit_grace_period=ROLLOUT_EXIT_GRACE_PERIOD, - port=SIDECAR_PORT, - routing_region="us-west", -) -class Server: - @modal.enter() - def start(self) -> None: - self.sglang = start_sglang( - definition.HF_CHECKPOINT, - port=SGLANG_PORT, - context_length=definition.MAX_CONTEXT_LENGTH, - max_loras_per_batch=1, - max_loaded_loras=1, - max_lora_rank=1, - max_running_requests=ROLLOUT_MAX_RUNNING_REQUESTS, - max_queued_requests=ROLLOUT_MAX_QUEUED_REQUESTS, - tensor_parallel_size=ROLLOUT_TENSOR_PARALLEL_SIZE, - expert_parallel_size=ROLLOUT_EXPERT_PARALLEL_SIZE, - expert_tensor_parallel_size=ROLLOUT_EXPERT_TENSOR_PARALLEL_SIZE, - parallel_world_size=ROLLOUT_GPUS, - enable_lora=False, - enable_cpu_weight_cache=True, - cpu_weight_cache_max_compile_group_gb=( - ROLLOUT_CPU_WEIGHT_CACHE_MAX_COMPILE_GROUP_GB - ), - memory_fraction=ROLLOUT_MEMORY_FRACTION, - schedule_policy="lpm", - ) - wait_http( - f"http://127.0.0.1:{SGLANG_PORT}/health", - self.sglang, - 20 * 60, - ) - self.sidecar = start_fft_sidecar( - port=SIDECAR_PORT, - sglang_port=SGLANG_PORT, - model_path=definition.HF_CHECKPOINT, - bulletin_root=definition.BULLETIN_ROOT, - bulletin_volume=definition.BULLETIN_VOLUME_NAME, - run_id=MODEL_ID, - pinned_version=None if LATEST else VERSION, - ) - self.supervisor = supervise_children(self.sglang, self.sidecar) - wait_http( - f"http://127.0.0.1:{SIDECAR_PORT}/health", - self.sidecar, - 20 * 60, - ) - - @modal.exit() - def stop(self) -> None: - terminate(getattr(self, "sidecar", None)) - terminate(getattr(self, "sglang", None)) diff --git a/src/lilo/providers/modal/lora_pool.py b/src/lilo/providers/modal/lora_pool.py index a46376f..3270636 100644 --- a/src/lilo/providers/modal/lora_pool.py +++ b/src/lilo/providers/modal/lora_pool.py @@ -1,6 +1,5 @@ from __future__ import annotations -import ast import hashlib import os import shutil @@ -73,9 +72,7 @@ def deploy_pool(spec: LoraPoolSpec) -> str: modal_cli, "deploy", "-m", - "lilo.providers.modal.yaml_pool_app" - if recipe_env - else "lilo.providers.modal.lora_pool_app", + "lilo.providers.modal.yaml_pool_app", "--name", spec.app_name, ] @@ -104,42 +101,6 @@ def stop_pool(spec: LoraPoolSpec) -> None: def _implementation_revision(definition_id: str) -> str: - if definition_id.startswith("yaml_"): - return definition_id.rsplit("_", 1)[-1] - here = Path(__file__) - files = ( - here, - here.with_name("lora_pool_app.py"), - here.with_name("rollout_image.py"), - here.with_name("image_dependencies.py"), - *_definition_sources(here.with_name("definitions") / f"{definition_id}.py"), - here.parents[2] / "inference" / "bulletin.py", - here.parents[2] / "inference" / "lora_sidecar.py", - here.parents[2] / "inference" / "serving.py", - ) - digest = hashlib.sha256() - for path in files: - digest.update(path.name.encode()) - digest.update(path.read_bytes()) - return digest.hexdigest() - - -def _definition_sources(path: Path): - """Include inherited sibling definitions without importing deployment code.""" - pending, seen = [path], set() - while pending: - source = pending.pop() - if source in seen: - continue - seen.add(source) - yield source - for node in ast.walk(ast.parse(source.read_text())): - if not isinstance(node, ast.ImportFrom) or node.level != 1: - continue - modules = ( - [node.module] if node.module else [alias.name for alias in node.names] - ) - for module in modules: - sibling = source.parent / (module.replace(".", "/") + ".py") - if sibling.is_file(): - pending.append(sibling) + if not definition_id.startswith("yaml_"): + raise ValueError(f"expected a YAML deployment id: {definition_id}") + return definition_id.rsplit("_", 1)[-1] diff --git a/src/lilo/providers/modal/lora_pool_app.py b/src/lilo/providers/modal/lora_pool_app.py deleted file mode 100644 index 9320998..0000000 --- a/src/lilo/providers/modal/lora_pool_app.py +++ /dev/null @@ -1,109 +0,0 @@ -from __future__ import annotations - -import importlib -import os - -import modal - -from lilo.inference.serving import ( - start_lora_sidecar, - start_sglang, - supervise_children, - terminate, - wait_http, -) - -from .rollout_image import image - -APP_NAME = os.environ["LILO_LORA_POOL_APP_NAME"] -DEFINITION_ID = os.environ["LILO_LORA_POOL_DEFINITION_ID"] -definition = importlib.import_module( - f"lilo.providers.modal.definitions.{DEFINITION_ID}" -) -SGLANG_PORT = 8001 -SIDECAR_PORT = 8000 - -api_secret = modal.Secret.from_name( - "lilo-api", - required_keys=["TINKER_API_KEY"], -) -pool_secret = modal.Secret.from_dict( - { - key: value - for key, value in os.environ.items() - if key.startswith("LILO_LORA_POOL_") - } -) -app = modal.App(APP_NAME) - - -@app.server( - image=image, - gpu=f"{definition.ROLLOUT_GPU_TYPE}:{definition.ROLLOUT_GPUS}", - volumes={ - "/assets": definition.assets, - definition.BULLETIN_ROOT: definition.bulletin, - }, - secrets=[api_secret, pool_secret], - target_concurrency=definition.ROLLOUT_TARGET_CONCURRENCY, - min_containers=getattr(definition, "ROLLOUT_MIN_CONTAINERS", 0), - max_containers=getattr(definition, "ROLLOUT_MAX_CONTAINERS", None), - scaledown_window=getattr(definition, "ROLLOUT_SCALEDOWN_WINDOW", 5 * 60), - startup_timeout=20 * 60, - exit_grace_period=getattr( - definition, - "ROLLOUT_EXIT_GRACE_PERIOD", - 5 * 60, - ), - port=SIDECAR_PORT, - routing_region="us-west", -) -class Server: - @modal.enter() - def start(self) -> None: - self.sglang = start_sglang( - definition.HF_CHECKPOINT, - port=SGLANG_PORT, - context_length=definition.MAX_CONTEXT_LENGTH, - max_loras_per_batch=getattr( - definition, - "ROLLOUT_MAX_LORAS_PER_BATCH", - 8, - ), - max_loaded_loras=definition.ROLLOUT_MAX_LOADED_LORAS, - max_lora_rank=definition.MAX_LORA_RANK, - max_running_requests=definition.ROLLOUT_MAX_RUNNING_REQUESTS, - max_queued_requests=definition.ROLLOUT_MAX_QUEUED_REQUESTS, - tensor_parallel_size=definition.ROLLOUT_TENSOR_PARALLEL_SIZE, - expert_parallel_size=definition.ROLLOUT_EXPERT_PARALLEL_SIZE, - expert_tensor_parallel_size=( - definition.ROLLOUT_EXPERT_TENSOR_PARALLEL_SIZE - ), - parallel_world_size=definition.ROLLOUT_GPUS, - lora_target_modules=definition.ROLLOUT_LORA_TARGET_MODULES, - enable_lora=True, - memory_fraction=definition.ROLLOUT_MEMORY_FRACTION, - schedule_policy="lpm", - ) - wait_http( - f"http://127.0.0.1:{SGLANG_PORT}/health", - self.sglang, - 20 * 60, - ) - self.sidecar = start_lora_sidecar( - port=SIDECAR_PORT, - sglang_port=SGLANG_PORT, - bulletin_root=definition.BULLETIN_ROOT, - bulletin_volume=definition.BULLETIN_VOLUME_NAME, - ) - self.supervisor = supervise_children(self.sglang, self.sidecar) - wait_http( - f"http://127.0.0.1:{SIDECAR_PORT}/health", - self.sidecar, - 20 * 60, - ) - - @modal.exit() - def stop(self) -> None: - terminate(getattr(self, "sidecar", None)) - terminate(getattr(self, "sglang", None)) diff --git a/src/lilo/providers/modal/recipe.py b/src/lilo/providers/modal/recipe.py index 65c95cf..98c18ff 100644 --- a/src/lilo/providers/modal/recipe.py +++ b/src/lilo/providers/modal/recipe.py @@ -43,10 +43,8 @@ "max_lora_rank", "enable_cpu_weight_cache", "api_key", - "dp_size", "pp_size", "lora_paths", - "enable_dp_attention", "dist_init_addr", "nnodes", "node_rank", @@ -173,11 +171,17 @@ def serving_options(spec): tp = options.get("tp_size", spec.inference.resources.gpu_count) ep = options.get("ep_size", 1) if not isinstance(tp, int) or tp != spec.inference.resources.gpu_count: - raise ValueError( - "sglang.tp_size must equal the replica GPU allocation; DP-attention is not yet supported in YAML" - ) + raise ValueError("sglang.tp_size must equal the replica GPU allocation") if not isinstance(ep, int) or ep < 1 or tp % ep: raise ValueError("sglang.ep_size must divide the replica GPU allocation") + dp = options.get("dp_size", 1) + dp_attention = options.get("enable_dp_attention", False) + if type(dp) is not int or dp < 1 or tp % dp: + raise ValueError("sglang.dp_size must divide the replica GPU allocation") + if not isinstance(dp_attention, bool): + raise ValueError("sglang.enable_dp_attention must be a boolean") + if dp > 1 and not dp_attention: + raise ValueError("sglang.dp_size > 1 requires enable_dp_attention") for key in ( "max_loaded_loras", "max_loras_per_batch", diff --git a/src/lilo/providers/modal/yaml_apps.py b/src/lilo/providers/modal/yaml_apps.py index 18c26e4..f43b11b 100644 --- a/src/lilo/providers/modal/yaml_apps.py +++ b/src/lilo/providers/modal/yaml_apps.py @@ -21,17 +21,18 @@ def manifest_from_env(): data = os.environ.get(MANIFEST_ENV) - return ( - [ResolvedDeployment.model_validate(row) for row in json.loads(data)] - if data - else [] - ) + if not data: + raise ValueError( + "Missing deployment manifest. Use lilo deploy with your YAML files." + ) + rows = json.loads(data) + if not isinstance(rows, list) or not rows: + raise ValueError("Deployment manifest must be a nonempty list") + return [ResolvedDeployment.model_validate(row) for row in rows] def frontend_settings(): deployments = manifest_from_env() - if not deployments: - return None active = [row.spec for row in deployments if row.active] validate_frontend(active) if len({row.definition_id for row in deployments}) != len(deployments): @@ -196,7 +197,8 @@ def definition_from_spec(resolved, *, register_trainer=True, image=None): ROLLOUT_GPUS=spec.inference.resources.gpu_count, ROLLOUT_TENSOR_PARALLEL_SIZE=native.get( "tp_size", spec.inference.resources.gpu_count - ), + ) + // (native.get("dp_size", 1) if native.get("enable_dp_attention") else 1), ) if register_trainer: definition.app, definition.ENGINE_FUNCTION = build_trainer_app( @@ -215,9 +217,7 @@ def pool_deployment(definition_id): def pool_environment(definition_id): resolved = pool_deployment(definition_id) if resolved is None: - if definition_id.startswith("yaml_"): - raise ValueError(f"missing recorded deployment: {definition_id}") - return {} + raise ValueError(f"missing recorded deployment: {definition_id}") return {POOL_CONFIG_ENV: resolved.model_dump_json()} diff --git a/tests/control_plane/test_http.py b/tests/control_plane/test_http.py index 22df0eb..402c2b4 100644 --- a/tests/control_plane/test_http.py +++ b/tests/control_plane/test_http.py @@ -283,13 +283,14 @@ async def run() -> None: asyncio.run(run()) -def test_base_sampling_session_prefers_full_definition() -> None: +def test_base_sampling_session_uses_explicit_sampling_default() -> None: async def run() -> None: plane = ControlPlane( InMemoryKeyValueStore(), LocalEnginePlatform(DEFINITION, EchoExecutor), ) - app = create_control_plane_app(plane, DEFINITIONS, api_key=None) + definitions = [SimpleNamespace(**vars(d), SAMPLING_DEFAULT=d.PARAMETERIZATION == "full") for d in DEFINITIONS] + app = create_control_plane_app(plane, definitions, api_key=None) client = httpx.AsyncClient( base_url="http://control-plane", transport=httpx.ASGITransport(app=app), diff --git a/tests/providers/conftest.py b/tests/providers/conftest.py new file mode 100644 index 0000000..41b846f --- /dev/null +++ b/tests/providers/conftest.py @@ -0,0 +1,18 @@ +"""Shared provider tests construct the app from an explicit offline manifest.""" + +import json +import os + +from lilo.deployments import load, preset_path, resolve + +os.environ.setdefault( + "LILO_DEPLOYMENT_MANIFEST", + json.dumps( + [ + resolve( + load(preset_path(name)), revision="a" * 40, implementation="tests" + ).model_dump(mode="json") + for name in ("qwen35-9b-fft-64k", "qwen35-9b-lora-16k") + ] + ), +) diff --git a/tests/providers/test_checkpoint_storage.py b/tests/providers/test_checkpoint_storage.py index c9f4da2..072669f 100644 --- a/tests/providers/test_checkpoint_storage.py +++ b/tests/providers/test_checkpoint_storage.py @@ -1,4 +1,3 @@ -import ast import runpy from pathlib import Path from unittest.mock import patch, sentinel @@ -8,8 +7,6 @@ from lilo.providers.modal.app import DEFINITIONS from lilo.providers.modal.checkpoint_storage import ( CHECKPOINT_ROOT, - CHECKPOINT_VOLUME_NAME, - checkpoint_volume, ) FULL_DEFINITIONS = tuple( @@ -17,18 +14,6 @@ ) -def source_tree(definition) -> ast.Module: - return ast.parse(Path(definition.__file__).read_text()) - - -def string_dict_entries(node: ast.Dict) -> dict[str, ast.expr]: - return { - key.value: value - for key, value in zip(node.keys, node.values, strict=True) - if isinstance(key, ast.Constant) and isinstance(key.value, str) - } - - def test_checkpoint_storage_creates_one_v2_volume_without_live_lookup() -> None: storage_path = Path(__file__).parents[2] / ( "src/lilo/providers/modal/checkpoint_storage.py" @@ -48,56 +33,22 @@ def test_checkpoint_storage_creates_one_v2_volume_without_live_lookup() -> None: assert storage["CHECKPOINT_ROOT"] == "/checkpoints" -def test_all_definitions_share_checkpoint_storage() -> None: - assert len(FULL_DEFINITIONS) == 5 - assert CHECKPOINT_VOLUME_NAME == "lilo-checkpoints" - assert CHECKPOINT_ROOT == "/checkpoints" +def test_yaml_definitions_use_configured_checkpoint_storage(): + from lilo.providers.modal.yaml_apps import volumes_for + from lilo.providers.modal.recipe import backend_config for definition in DEFINITIONS: - assert definition.TRAINER_VOLUMES[CHECKPOINT_ROOT] is checkpoint_volume - assert ( - definition.TRAINER_VOLUMES[definition.BULLETIN_ROOT] is definition.bulletin + spec = definition.RESOLVED.spec + with patch.object( + modal.Volume, "from_name", side_effect=lambda name, **kwargs: (name, kwargs) + ): + volumes = volumes_for(spec) + assert volumes[CHECKPOINT_ROOT] == ( + spec.deployment.storage.checkpoints, + {"create_if_missing": True, "version": 2}, ) - assert definition.bulletin is not checkpoint_volume - - -def test_full_definitions_configure_checkpoint_dir_and_environment() -> None: - for definition in FULL_DEFINITIONS: - tree = source_tree(definition) - config_assignments = [ - node - for node in ast.walk(tree) - if isinstance(node, ast.Assign) - and any( - isinstance(target, ast.Name) and target.id == "backend_config" - for target in node.targets - ) - and isinstance(node.value, ast.Dict) - ] - assert len(config_assignments) == 1 - config = string_dict_entries(config_assignments[0].value) - checkpoint_dir = config["checkpoint_dir"] - assert isinstance(checkpoint_dir, ast.Name) - assert checkpoint_dir.id == "CHECKPOINT_ROOT" - - engine_calls = [ - node - for node in ast.walk(tree) - if isinstance(node, ast.Call) - and isinstance(node.func, ast.Name) - and node.func.id == "run_engine_with_backend" - ] - assert len(engine_calls) == 1 - backend_env = next( - keyword.value - for keyword in engine_calls[0].keywords - if keyword.arg == "backend_env" - ) - assert isinstance(backend_env, ast.Dict) - env = string_dict_entries(backend_env) - volume_name = env["LILO_CHECKPOINT_VOLUME"] - assert isinstance(volume_name, ast.Name) - assert volume_name.id == "CHECKPOINT_VOLUME_NAME" + assert volumes["/bulletin"][0] == spec.deployment.storage.bulletin + assert backend_config(spec)["checkpoint_dir"] == CHECKPOINT_ROOT def test_same_checkpoint_name_isolated_by_model(tmp_path, monkeypatch) -> None: @@ -106,8 +57,10 @@ def test_same_checkpoint_name_isolated_by_model(tmp_path, monkeypatch) -> None: app = importlib.import_module("lilo.providers.modal.app") from lilo.providers.modal.checkpoint_storage import _scan_checkpoints + def scan(model_id): return _scan_checkpoints(str(tmp_path), model_id) + monkeypatch.setattr(app, "CHECKPOINT_ROOT", str(tmp_path)) for relative in ("final/run-a", "final/run-b"): checkpoint = tmp_path / relative @@ -117,7 +70,8 @@ def scan(model_id): (tmp_path / "notes.txt").write_text("notes") entries = scan(None) assert {(e["model_id"], e["name"]) for e in entries} == { - ("run-a", "final"), ("run-b", "final") + ("run-a", "final"), + ("run-b", "final"), } assert len(entries) == 2 assert scan("run-a")[0]["path"] == str(tmp_path / "final/run-a") diff --git a/tests/providers/test_deployment_presets.py b/tests/providers/test_deployment_presets.py new file mode 100644 index 0000000..2a3af35 --- /dev/null +++ b/tests/providers/test_deployment_presets.py @@ -0,0 +1,88 @@ +import pytest + +from lilo.deployments import load, preset_path, resolve +from lilo.providers.modal.recipe import backend_config, serving_options +from lilo.providers.modal.yaml_apps import ( + definition_from_spec, + frontend_settings, + manifest_from_env, +) + + +@pytest.mark.parametrize( + "path", + sorted(preset_path("qwen35-9b-lora-16k").parent.glob("*.yaml")), + ids=lambda p: p.stem, +) +def test_all_packaged_recipes_validate_offline(path): + spec = load(path) + config = backend_config(spec) + assert config[spec.trainer.backend]["hf_checkpoint"] == "/assets/pending" + serving_options(spec) + + +@pytest.mark.parametrize("preset", ["qwen35-35b-a3b-fft-64k", "qwen36-35b-a3b-fft-64k"]) +def test_moe_recipes_preserve_trainer_expert_parallelism(preset): + config = backend_config(load(preset_path(preset)))["megatron"] + assert config["tensor_model_parallel_size"] == 4 + assert config["context_parallel_size"] == 2 + assert config["expert_model_parallel_size"] == 8 + assert config["provider_overrides"]["moe_token_dispatcher_type"] == "alltoall" + + +def test_moe_rollout_preserves_attention_data_parallelism(): + spec = load(preset_path("qwen35-35b-a3b-fft-64k")) + options = serving_options(spec) + assert options["tp_size"] == options["dp_size"] == options["ep_size"] == 4 + assert options["enable_dp_attention"] is True + definition = definition_from_spec( + resolve(spec, revision="a" * 40, implementation="test"), register_trainer=False + ) + assert definition.ROLLOUT_GPUS == 4 + assert definition.ROLLOUT_TENSOR_PARALLEL_SIZE == 1 + + +@pytest.mark.parametrize("context,cp", [(16384, 1), (65536, 2), (131072, 4)]) +def test_qwen38_context_parallel_token_budget(context, cp): + spec = load(preset_path(f"qwen38-27b-lora-{context // 1024}k")) + config = backend_config(spec)["miles"] + assert config["context_parallel_size"] == cp + assert config["max_tokens_per_gpu"] == context // cp + assert config["actor_num_gpus_per_node"] == 8 + + +def test_single_client_recipe_keeps_shared_backend_capacity(): + shared = load(preset_path("qwen35-9b-lora-16k")) + single = load(preset_path("qwen35-9b-lora-16k-single")) + assert single.trainer.engine.max_clients_per_instance == 1 + assert single.trainer.resources == shared.trainer.resources + assert backend_config(single) == backend_config(shared) + + +@pytest.mark.parametrize("value", [None, "", "[]", "{}"]) +def test_missing_manifest_has_no_python_catalog_fallback(monkeypatch, value): + if value is None: + monkeypatch.delenv("LILO_DEPLOYMENT_MANIFEST", raising=False) + else: + monkeypatch.setenv("LILO_DEPLOYMENT_MANIFEST", value) + with pytest.raises(ValueError, match="manifest"): + manifest_from_env() + with pytest.raises(ValueError, match="manifest"): + frontend_settings() + + +@pytest.mark.parametrize( + "options", + [ + {"dp_size": 3, "enable_dp_attention": True}, + {"dp_size": 2, "enable_dp_attention": False}, + {"dp_size": 2, "enable_dp_attention": "true"}, + ], +) +def test_invalid_attention_parallelism_is_rejected(options): + from lilo.deployments import DeploymentSpec + + data = load(preset_path("qwen35-35b-a3b-fft-64k")).model_dump() + data["inference"]["sglang"]["options"].update(options) + with pytest.raises(ValueError, match="sglang"): + DeploymentSpec.model_validate(data) diff --git a/tests/providers/test_lora_pool.py b/tests/providers/test_lora_pool.py index 902b4e6..e1dc9f5 100644 --- a/tests/providers/test_lora_pool.py +++ b/tests/providers/test_lora_pool.py @@ -2,7 +2,7 @@ def test_lora_pool_is_shared_by_every_adapter_for_definition() -> None: - first = LoraPoolSpec("qwen3_5_9b_base_miles_lora_2k") + first = LoraPoolSpec("yaml_example_0123456789abcdef") second = LoraPoolSpec.from_dict(first.as_dict()) assert first == second @@ -34,30 +34,15 @@ def test_stop_already_stopped_lora_pool_succeeds_but_real_failure_propagates( lora_pool.stop_pool(spec) -def test_pool_revision_tracks_inherited_settings_and_bulletin(monkeypatch): - from pathlib import Path +def test_pool_revision_comes_from_resolved_generation(): + first = LoraPoolSpec("yaml_example_0123456789abcdef") + changed = LoraPoolSpec("yaml_example_fedcba9876543210") + assert first.revision == "0123456789abcdef" + assert first.app_name != changed.app_name - original = Path.read_bytes - changed = None - def read(path): - value = original(path) - return value + b"\n# changed\n" if path.name == changed else value - - monkeypatch.setattr(Path, "read_bytes", read) - name = "qwen3_5_9b_base_miles_lora_16k_single" - original_pool = LoraPoolSpec(name).app_name - for changed in ("qwen3_5_9b_base_miles_lora_16k.py", "bulletin.py"): - assert LoraPoolSpec(name).app_name != original_pool - changed = "qwen3_5_4b_full_64k.py" - assert LoraPoolSpec(name).app_name == original_pool - - -def test_definition_dependency_scan_handles_cycles_and_relative_imports(tmp_path): - from lilo.providers.modal.lora_pool import _definition_sources +def test_python_definition_cannot_choose_a_pool_revision(): + import pytest - (tmp_path / "child.py").write_text("from .parent import CONFIG\n") - (tmp_path / "parent.py").write_text("from . import shared\n") - (tmp_path / "shared.py").write_text("from .child import CONFIG\n") - sources = list(_definition_sources(tmp_path / "child.py")) - assert [path.name for path in sources] == ["child.py", "parent.py", "shared.py"] + with pytest.raises(ValueError, match="YAML deployment id"): + LoraPoolSpec("qwen3_5_9b_base_miles_lora_16k") diff --git a/tests/providers/test_miles_definition.py b/tests/providers/test_miles_definition.py deleted file mode 100644 index a4a3342..0000000 --- a/tests/providers/test_miles_definition.py +++ /dev/null @@ -1,56 +0,0 @@ -import ast -from pathlib import Path - -from lilo.backends.miles_config import MILES_REF -from lilo.providers.modal import miles_image -from lilo.providers.modal.definitions import qwen3_5_9b_base_miles_lora_2k as definition - - -def test_miles_definition_uses_one_lilo_driver_for_all_ray_workers() -> None: - assert definition.PARAMETERIZATION == "lora" - assert definition.TRAINER_MODELS_PER_INSTANCE == definition.MAX_LORA_SLOTS - assert definition.GPUS == 4 - assert MILES_REF == "main" - assert len(miles_image.MILES_COMMIT) == 40 - - tree = ast.parse(Path(definition.__file__).read_text()) - calls = [ - node - for node in ast.walk(tree) - if isinstance(node, ast.Call) and isinstance(node.func, ast.Name) and node.func.id == "run_engine_with_backend" - ] - assert len(calls) == 1 - call = calls[0] - assert isinstance(call.args[1], ast.Constant) - assert call.args[1].value == "lilo.backends.miles_lora:build_executor" - keywords = {keyword.arg: keyword.value for keyword in call.keywords} - assert isinstance(keywords["nproc"], ast.Constant) - assert keywords["nproc"].value == 1 - assert isinstance(keywords["max_models"], ast.Name) - assert keywords["max_models"].id == "MAX_LORA_SLOTS" - - -def test_long_context_miles_definition_bounds_retained_adapter_versions(): - from lilo.providers.modal.definitions import qwen3_5_9b_base_miles_lora_16k as long - - assert long.MAX_CONTEXT_LENGTH >= 8192 + 2048 - assert long.MAX_LORA_SLOTS >= 6 - assert long.ROLLOUT_MAX_LOADED_LORAS == 64 - assert long.ROLLOUT_MAX_LOADED_LORAS >= long.ROLLOUT_MAX_LORAS_PER_BATCH - assert long.ROLLOUT_MAX_CONTAINERS == 8 - assert long.CATALOG_VISIBLE - assert not definition.CATALOG_VISIBLE - - - -def test_single_tenant_definition_preserves_shared_hardware_and_backend(): - from lilo.providers.modal.definitions import qwen3_5_9b_base_miles_lora_16k as shared - from lilo.providers.modal.definitions import qwen3_5_9b_base_miles_lora_16k_single as single - - assert single.TRAINER_MODELS_PER_INSTANCE == 1 - assert not single.CATALOG_VISIBLE - assert single.run_trainer is shared.run_trainer - for name in ("MODEL_NAME", "GPUS", "GPU_TYPE", "MAX_CONTEXT_LENGTH", - "TENSOR_MODEL_PARALLEL_SIZE", "MAX_LORA_SLOTS", "MAX_LORA_RANK", - "ROLLOUT_MIN_CONTAINERS", "ROLLOUT_MAX_CONTAINERS", "ROLLOUT_GPU_TYPE"): - assert getattr(single, name) == getattr(shared, name) diff --git a/tests/providers/test_modal_app.py b/tests/providers/test_modal_app.py index 055cc64..7dd8fb7 100644 --- a/tests/providers/test_modal_app.py +++ b/tests/providers/test_modal_app.py @@ -10,8 +10,17 @@ from lilo.providers.modal.fft_pool import FFTPoolSpec from lilo.providers.modal.lora_pool import LoraPoolSpec -FULL_DEFINITION = "qwen3_5_9b_full_64k" -LORA_DEFINITION = "qwen3_5_9b_base_miles_lora_2k" +from lilo.deployments import load, preset_path, resolve + + +def definition_id(preset): + return resolve( + load(preset_path(preset)), revision="a" * 40, implementation="tests" + ).definition_id + + +FULL_DEFINITION = definition_id("qwen35-9b-fft-64k") +LORA_DEFINITION = definition_id("qwen35-9b-lora-16k") @pytest.fixture(autouse=True) @@ -21,10 +30,11 @@ def reset_lora_pool_cache(monkeypatch): monkeypatch.setattr(modal_app, "_lora_pool_checks", {}) -def test_definitions_exclude_stale_128k_definition() -> None: +def test_definitions_come_only_from_the_configured_manifest() -> None: modal_app = importlib.import_module("lilo.providers.modal.app") - assert "qwen3_5_9b_full_128k" not in { - definition.DEFINITION_ID for definition in modal_app.DEFINITIONS + assert {definition.DEFINITION_ID for definition in modal_app.DEFINITIONS} == { + FULL_DEFINITION, + LORA_DEFINITION, } @@ -60,7 +70,9 @@ async def kick(_definition_id: str) -> None: monkeypatch.setattr(modal_app, "ModalSessionKeyValueStores", SimpleNamespace) monkeypatch.setattr(modal_app, "ModalEnginePlatform", lambda *args: engines) monkeypatch.setattr(modal_app, "kick_trainer_reconciler", kick) - monkeypatch.setattr(modal_app, "TRAINER_MAX_CONTAINERS", "1") + monkeypatch.setattr( + modal_app.module_for(LORA_DEFINITION), "TRAINER_MAX_CONTAINERS", 1 + ) plane = modal_app._plane() assert asyncio.run(plane.reconcile_trainers(LORA_DEFINITION)) is available @@ -174,8 +186,8 @@ def test_prepare_model_assets_validates_snapshot_before_commit(monkeypatch) -> N modal_app = importlib.import_module("lilo.providers.modal.app") events = [] - def download(*, repo_id: str, local_dir: str) -> None: - events.append(("download", repo_id, local_dir)) + def download(*, repo_id: str, local_dir: str, revision: str) -> None: + events.append(("download", repo_id, local_dir, revision)) monkeypatch.setattr("huggingface_hub.snapshot_download", download) monkeypatch.setattr( @@ -188,7 +200,12 @@ def download(*, repo_id: str, local_dir: str) -> None: definition = modal_app.module_for(FULL_DEFINITION) assert events == [ - ("download", definition.MODEL_NAME, definition.HF_CHECKPOINT), + ( + "download", + definition.MODEL_NAME, + definition.HF_CHECKPOINT, + definition.MODEL_REVISION, + ), ("commit",), ] @@ -834,7 +851,9 @@ async def deploy(record): "ensure_lora_pool", SimpleNamespace(remote=SimpleNamespace(aio=deploy)), ) - session = SimpleNamespace(engine_definition_id=LORA_DEFINITION) + session = SimpleNamespace( + engine_definition_id=LORA_DEFINITION, model_id="existing-model" + ) async def run(): plane = modal_app._plane() diff --git a/tests/providers/test_modal_deployment.py b/tests/providers/test_modal_deployment.py index 164378f..8562f95 100644 --- a/tests/providers/test_modal_deployment.py +++ b/tests/providers/test_modal_deployment.py @@ -1,37 +1,6 @@ -import pytest +from lilo.providers.modal.deployment import trainer_deployment_env -from lilo.providers.modal.deployment import ( - TRAINER_MAX_CONTAINERS_ENV, - trainer_deployment_env, - trainer_max_containers, -) - -def test_trainer_max_containers_is_configured_for_deployment(monkeypatch) -> None: - monkeypatch.setenv(TRAINER_MAX_CONTAINERS_ENV, "3") - - assert trainer_max_containers() == 3 - assert trainer_deployment_env() == {TRAINER_MAX_CONTAINERS_ENV: "3"} - - -def test_trainer_max_containers_is_unlimited_when_unset(monkeypatch) -> None: - monkeypatch.delenv(TRAINER_MAX_CONTAINERS_ENV, raising=False) - - assert trainer_max_containers() is None - assert trainer_deployment_env() == {} - - -@pytest.mark.parametrize( - "value", - [ - "invalid", - "0", - "-1", - "1.5", - ], -) -def test_trainer_max_containers_rejects_invalid_config(monkeypatch, value) -> None: - monkeypatch.setenv(TRAINER_MAX_CONTAINERS_ENV, value) - - with pytest.raises(ValueError, match=TRAINER_MAX_CONTAINERS_ENV): - trainer_max_containers() +def test_capacity_is_not_forwarded_from_legacy_environment(monkeypatch): + monkeypatch.setenv("LILO_TRAINER_MAX_CONTAINERS", "3") + assert "LILO_TRAINER_MAX_CONTAINERS" not in trainer_deployment_env() diff --git a/tests/providers/test_qwen36_definitions.py b/tests/providers/test_qwen36_definitions.py deleted file mode 100644 index ee33a1d..0000000 --- a/tests/providers/test_qwen36_definitions.py +++ /dev/null @@ -1,33 +0,0 @@ -from lilo.providers.modal.app import ( - DEFINITIONS, - module_for, - parameterization_for, -) -from lilo.providers.modal.definitions import ( - qwen3_6_27b_full_64k, - qwen3_6_35b_a3b_full_64k, -) - - -def test_qwen36_full_definitions_are_registered() -> None: - for definition in ( - qwen3_6_27b_full_64k, - qwen3_6_35b_a3b_full_64k, - ): - assert definition in DEFINITIONS - assert module_for(definition.DEFINITION_ID) is definition - assert parameterization_for(definition.DEFINITION_ID) == "full" - - -def test_qwen36_35b_reuses_qwen35_moe_topology() -> None: - assert qwen3_6_35b_a3b_full_64k.TENSOR_MODEL_PARALLEL_SIZE == 4 - assert qwen3_6_35b_a3b_full_64k.CONTEXT_PARALLEL_SIZE == 2 - assert qwen3_6_35b_a3b_full_64k.EXPERT_MODEL_PARALLEL_SIZE == 8 - assert qwen3_6_35b_a3b_full_64k.ROLLOUT_EXPERT_PARALLEL_SIZE == 4 - - -def test_qwen36_27b_uses_dense_tp4_topology() -> None: - assert qwen3_6_27b_full_64k.TENSOR_MODEL_PARALLEL_SIZE == 4 - assert qwen3_6_27b_full_64k.CONTEXT_PARALLEL_SIZE == 2 - assert qwen3_6_27b_full_64k.ROLLOUT_TENSOR_PARALLEL_SIZE == 4 - assert qwen3_6_27b_full_64k.ROLLOUT_EXPERT_PARALLEL_SIZE == 1 diff --git a/tests/providers/test_yaml_apps.py b/tests/providers/test_yaml_apps.py index 88e5180..40a90fe 100644 --- a/tests/providers/test_yaml_apps.py +++ b/tests/providers/test_yaml_apps.py @@ -45,8 +45,18 @@ def builders(monkeypatch): monkeypatch.setattr(modal, "exit", lambda: lambda fn: fn) -def test_trainer_declaration_and_executor_configuration(builders, monkeypatch): - row = deployment() +@pytest.mark.parametrize( + "preset,backend,clients,nproc", + [ + ("qwen35-9b-lora-16k", "miles_lora", 6, 1), + ("qwen35-4b-fft-64k", "megatron_fft", 1, 4), + ], +) +def test_trainer_declaration_and_executor_configuration( + builders, monkeypatch, preset, backend, clients, nproc +): + row = deployment(preset) + row.spec.deployment.storage.checkpoints = "test-custom-checkpoints" image = object() app, trainer = yaml_apps.build_trainer_app(row, image=image) declaration, _ = app.functions[row.definition_id] @@ -72,12 +82,14 @@ def test_trainer_declaration_and_executor_configuration(builders, monkeypatch): ) trainer("instance-a") args, kwargs = calls[0] - assert args == ("store", "lilo.backends.miles_lora:build_executor") - assert kwargs["max_models"] == 6 - assert kwargs["nproc"] == 1 + assert args == ("store", f"lilo.backends.{backend}:build_executor") + assert kwargs["max_models"] == clients + assert kwargs["nproc"] == nproc + assert kwargs["backend_env"]["LILO_CHECKPOINT_VOLUME"] == "test-custom-checkpoints" assert kwargs["backend_env"]["LILO_BASE_MODEL_REVISION"] == "a" * 40 config = json.loads(kwargs["backend_env"]["LILO_BACKEND_CONFIG"]) - assert config["miles"]["hf_checkpoint"] == row.asset_path + assert config[row.spec.trainer.backend]["hf_checkpoint"] == row.asset_path + assert config["checkpoint_dir"] == "/checkpoints" assert reloaded == [True] @@ -156,7 +168,8 @@ def test_pool_subprocess_receives_recorded_generation(monkeypatch): assert json.loads(env[yaml_apps.POOL_CONFIG_ENV])["generation"] == row.generation with pytest.raises(ValueError, match="missing recorded"): yaml_apps.pool_environment("yaml_missing_123") - assert yaml_apps.pool_environment("legacy") == {} + with pytest.raises(ValueError, match="missing recorded"): + yaml_apps.pool_environment("unconfigured-python-definition") def test_startup_failure_is_visible_and_blocks_new_spawns(monkeypatch): @@ -226,3 +239,43 @@ def test_admission_changes_preserve_serialized_trainer(builders): assert first.active is True and first.spec.routing.default is True changed.spec.trainer.resources.gpu = "H200:4" assert serialize(yaml_apps.build_trainer_app(changed, image="test")[1]) != old_bytes + + +@pytest.mark.parametrize("kind", ["lora", "full"]) +def test_pool_launch_uses_only_generic_yaml_app(monkeypatch, kind): + from lilo.providers.modal import fft_pool, lora_pool + + row = deployment("qwen35-9b-lora-16k" if kind == "lora" else "qwen35-4b-fft-64k") + monkeypatch.setenv(yaml_apps.MANIFEST_ENV, json.dumps([row.model_dump()])) + module = lora_pool if kind == "lora" else fft_pool + spec = ( + LoraPoolSpec(row.definition_id) + if kind == "lora" + else FFTPoolSpec(row.definition_id, "model", True, 0) + ) + calls = [] + + class Pool: + def __init__(self, *args): + self.lookups = 0 + + def gateway_url(self): + self.lookups += 1 + if self.lookups == 1: + raise modal.exception.NotFoundError("not deployed") + return "https://pool" + + monkeypatch.setattr(module, "ModalFlashPool", Pool) + monkeypatch.setattr(module.shutil, "which", lambda _: "/bin/modal") + monkeypatch.setattr( + module.subprocess, + "run", + lambda command, **kwargs: calls.append((command, kwargs)), + ) + assert module.deploy_pool(spec) == "https://pool" + command, kwargs = calls[0] + assert command[command.index("-m") + 1] == "lilo.providers.modal.yaml_pool_app" + assert ( + json.loads(kwargs["env"][yaml_apps.POOL_CONFIG_ENV])["generation"] + == row.generation + ) diff --git a/tests/providers/test_yaml_e2e_helper.py b/tests/providers/test_yaml_e2e_helper.py new file mode 100644 index 0000000..126ac14 --- /dev/null +++ b/tests/providers/test_yaml_e2e_helper.py @@ -0,0 +1,29 @@ +from pathlib import Path +import runpy +from types import SimpleNamespace + +import modal +import pytest + +from lilo.deployments import load, preset_path, resolve + + +@pytest.mark.parametrize("preset", ["qwen35-9b-lora-16k", "qwen35-4b-fft-64k"]) +def test_e2e_helper_reads_active_deployed_configuration(monkeypatch, preset): + helper = runpy.run_path( + str(Path(__file__).parents[2] / "scripts/e2e_engine_definition.py") + ) + row = resolve(load(preset_path(preset)), revision="a" * 40, implementation="test") + retired = row.model_copy(update={"active": False, "generation": "b" * 64}) + registry = SimpleNamespace( + get=lambda *args: [retired.model_dump(), row.model_dump()] + ) + monkeypatch.setattr(modal.Dict, "from_name", lambda name: registry) + definition, mode = helper["_definition"]("test-frontend", row.spec.name) + assert definition.DEFINITION_ID == row.definition_id + assert definition.MAX_CONTEXT_LENGTH == row.spec.model.max_context_length + assert definition.GPUS == row.spec.trainer.resources.gpu_count + assert mode == row.spec.model.parameterization + assert definition.MAX_TOKENS_PER_MICROBATCH > 0 + with pytest.raises(ValueError, match="one active YAML"): + helper["_definition"]("test-frontend", "missing") From 32506a0d0cf5cb16ef71be79589115ecc3bac866 Mon Sep 17 00:00:00 2001 From: kailash Date: Wed, 23 Sep 2026 17:24:55 +0000 Subject: [PATCH 09/27] Separate deployment orchestration from native backend configuration --- docs/deployment-yaml-design.md | 77 ++++++- docs/deployment-yaml-validation.md | 8 + scripts/e2e_engine_definition.py | 2 +- src/lilo/backend_options.py | 14 ++ src/lilo/backends/deployment.py | 28 +++ src/lilo/backends/megatron_deployment.py | 87 ++++++++ .../megatron_runtime/common/config.py | 2 + .../megatron_runtime/common/modeling.py | 2 + .../megatron_runtime/fft/checkpoint.py | 44 ++-- src/lilo/backends/miles_deployment.py | 90 ++++++++ src/lilo/deployments.py | 42 +--- src/lilo/inference/sglang_deployment.py | 60 ++++++ src/lilo/presets/qwen35-35b-a3b-fft-64k.yaml | 57 +++-- src/lilo/presets/qwen35-4b-fft-64k.yaml | 35 ++-- src/lilo/presets/qwen35-9b-fft-64k.yaml | 43 ++-- .../qwen35-9b-instruct-lora-16k-dp2.yaml | 2 +- .../presets/qwen35-9b-instruct-lora-16k.yaml | 43 ++-- src/lilo/presets/qwen35-9b-lora-16k.yaml | 30 +-- src/lilo/presets/qwen35-9b-lora-2k.yaml | 45 ++-- src/lilo/presets/qwen35-9b-lora-64k.yaml | 2 +- src/lilo/presets/qwen36-27b-fft-64k.yaml | 43 ++-- src/lilo/presets/qwen36-35b-a3b-fft-64k.yaml | 7 +- src/lilo/presets/qwen38-27b-lora-128k.yaml | 9 +- src/lilo/presets/qwen38-27b-lora-16k.yaml | 43 ++-- src/lilo/presets/qwen38-27b-lora-64k.yaml | 7 +- src/lilo/providers/modal/recipe.py | 197 ------------------ src/lilo/providers/modal/yaml_apps.py | 2 +- tests/backends/test_megatron_fft.py | 17 +- tests/backends/test_native_megatron_config.py | 61 ++++++ tests/providers/test_checkpoint_storage.py | 2 +- tests/providers/test_deployment_presets.py | 4 +- tests/test_deployments.py | 85 +++++++- 32 files changed, 733 insertions(+), 457 deletions(-) create mode 100644 src/lilo/backend_options.py create mode 100644 src/lilo/backends/deployment.py create mode 100644 src/lilo/backends/megatron_deployment.py create mode 100644 src/lilo/backends/miles_deployment.py create mode 100644 src/lilo/inference/sglang_deployment.py delete mode 100644 src/lilo/providers/modal/recipe.py create mode 100644 tests/backends/test_native_megatron_config.py diff --git a/docs/deployment-yaml-design.md b/docs/deployment-yaml-design.md index a49023f..8c1c9d7 100644 --- a/docs/deployment-yaml-design.md +++ b/docs/deployment-yaml-design.md @@ -45,7 +45,7 @@ A LoRA pool is shared by clients using the same deployment configuration and bas | --- | --- | | [`deployments.py`](../src/lilo/deployments.py) | Schema, YAML inheritance, revision pinning and configuration identifiers | | [`deployment_cli.py`](../src/lilo/deployment_cli.py) | Operator commands, saved manifests and serialized applies | -| [`recipe.py`](../src/lilo/providers/modal/recipe.py) | Translate configuration into Miles/Megatron and SGLang settings | +| [`backends/deployment.py`](../src/lilo/backends/deployment.py) | Dispatch configuration to backend-owned readers; no Modal dependency | | [`yaml_apps.py`](../src/lilo/providers/modal/yaml_apps.py) | Declare trainer functions and rollout server classes; start their processes | | [`app.py`](../src/lilo/providers/modal/app.py) | Register generated trainers with the existing shared app | | [`yaml_pool_app.py`](../src/lilo/providers/modal/yaml_pool_app.py) | Construct a rollout app in the pool deployment subprocess | @@ -96,7 +96,7 @@ trainer: gpu: H200:8 engine: max_clients_per_instance: 12 - miles: + config: options: tensor_model_parallel_size: 8 multi_lora_n_adapters: 12 @@ -152,11 +152,78 @@ Each client records its selected definition identifier. Changing defaults affect ## Backend options and model support -`trainer.miles.model_args` names an architecture preset inside Miles. It can be omitted when the YAML supplies explicit architecture options. Changing `model.id` does not make an inherited architecture preset compatible; Miles still validates the HF configuration at startup. +The deployment schema reads model identity, GPU resources, scaling, routing, storage and lifecycle settings. Each trainer and inference definition has a `backend` name and an opaque `config` mapping. The schema does not enumerate backend options. Backend-specific readers live beside the backend code, and the Modal app builder consumes their output. There is no Modal `recipe.py`. -`trainer.miles.options` and `inference.sglang.options` use argparse destination names such as `tensor_model_parallel_size` and `max_running_requests`. The Miles and SGLang entrypoints consult their real parsers before initialization. Boolean flags, scalar values and ordinary list arguments are supported. Both spellings of opposing boolean flags are replaced when they share a destination. Unknown options and custom/repeated argparse actions fail explicitly. Backend validation still runs after these overrides. +For Miles: -Lilo checks settings that affect its integration locally: GPU counts and parallelism, maximum clients versus adapter slots, managed model paths, context/rank configuration, communication endpoints and trainer-only mode. Passthrough cannot override these managed values. Megatron FFT uses its existing `EngineModelConfig` and provider overrides. +```yaml +trainer: + backend: miles + resources: + gpu: H100:4 + config: + model_args: qwen3.5-9B + options: + tensor_model_parallel_size: 4 + recompute_granularity: full + recompute_method: uniform + recompute_num_layers: 1 +``` + +`config.model_args` names an architecture preset inside Miles. It can be omitted when `config.options` supplies explicit architecture options. Changing `model.id` does not make an inherited architecture preset compatible; Miles still validates the HF configuration at startup. Miles options use its argparse destination names. The integration also reads parallelism, adapter limits and targets because Lilo uses those values for admission and adapter export. Other options reach Miles's parser without a Lilo allowlist. + +For Megatron FFT: + +```yaml +trainer: + backend: megatron + resources: + gpu: H100:4 + engine: + sampler_persistence_concurrency: 1 + config: + runtime: + tensor_model_parallel_size: 2 + context_parallel_size: 2 + sequence_parallel: true + max_tokens_per_microbatch: 65536 + use_distributed_optimizer: true + provider: + recompute_granularity: full + recompute_method: uniform + recompute_num_layers: 1 + optimizer: + lr: 0.0001 + loss_scale: 1.0 + distributed: + grad_reduce_in_fp32: true +``` + +- `runtime` configures Lilo's training loop, packing and parallelism. It has a finite schema because Lilo implements these settings. +- `provider` sets attributes on the model provider returned by Megatron Bridge. Names are checked against that actual provider at worker startup. Values, including lists and mappings, are preserved. +- `optimizer` accepts Megatron optimizer constructor options. Existing Lilo optimizer and scheduling fields (such as `lr` and `min_lr`) remain available to the loop; additional fields pass directly to Megatron's `OptimizerConfig`. Each Tinker `optim_step` still supplies the request's Adam parameters. +- `distributed` passes additional fields directly to Megatron's `DistributedDataParallelConfig`. + +New native Megatron provider, optimizer or distributed options do not require deployment-parser changes. The installed backend rejects unsupported options on worker startup. Optimizer and distributed overrides are recorded in FFT checkpoint metadata and checked on exact resume. + +For SGLang, `inference.config` directly contains its options: + +```yaml +inference: + backend: sglang + resources: + gpu: H200:1 + config: + tp_size: 1 + mem_fraction_static: 0.8 + max_running_requests: 32 +``` + +Miles and SGLang use their real argument parsers before initialization. Boolean flags, scalar values and ordinary list arguments are supported. Both spellings of opposing boolean flags are replaced when they share a destination. Unknown options and custom/repeated argparse actions fail explicitly. These are parser-backed options, so arbitrary Python objects and custom actions are not supported. + +Lilo still checks settings that affect its integration locally: GPU counts and parallelism, maximum clients versus adapter slots, managed model paths, context/rank configuration, communication endpoints and trainer-only mode. The backend readers reject conflicting values. For Megatron, set parallelism and shared distributed-optimizer controls under `runtime` so both Lilo and Megatron receive the same values. Tinker training requires an Adam optimizer. Native passthrough does not make other training protocols or unsupported process layouts work automatically. + +All built-in YAMLs use this structure. The earlier draft's `trainer.miles`, `trainer.megatron` and `inference.sglang` sections have been removed; user YAMLs overriding those sections must move them under `config` as shown above. This schema change is part of the draft and has not been deployed. Existing source-fingerprint checks continue to prevent applying a different runtime implementation over running trainers. Inference adapter targets are derived from the existing Miles-to-PEFT mapping unless `lora_target_modules` is explicitly supplied. This mapping does not establish support for every architecture. A model still needs compatible training, adapter export and SGLang loading implementations in the selected images. diff --git a/docs/deployment-yaml-validation.md b/docs/deployment-yaml-validation.md index a288748..340b806 100644 --- a/docs/deployment-yaml-validation.md +++ b/docs/deployment-yaml-validation.md @@ -2,6 +2,14 @@ These checks exercise the shared-app deployment path in PR #55. Each deployment contains the frontend and generated trainer functions; inference pools are separate apps started on demand. +## Backend configuration passthrough + +The deployment schema now uses `trainer.backend` / `trainer.config` and `inference.backend` / `inference.config`. Backend readers interpret those mappings outside the Modal provider. Megatron exposes native provider, optimizer and distributed-training settings alongside its Lilo runtime settings. + +Validation: **588 CPU tests passed, 1 skipped**. Ruff passed for changed Python files, and whitespace checks passed. A comparison against the previous commit confirmed that all 14 migrated presets produce the same resolved backend and inference settings, apart from the new empty native-override fields. Tests cover native values reaching Megatron constructors, preserving nested values and false booleans through serialization, rejecting conflicting integration settings, and rejecting exact checkpoint resume when native optimizer/distributed settings differ. + +Megatron constructor tests use CPU stubs; they verify forwarding and error propagation, not acceptance by the GPU image's installed Megatron version. The native backend remains responsible for validating its options at worker startup. This refactor has not been deployed or GPU-tested. + ## YAML-only deployment cleanup After the live checks below, the shared deployment's Python-catalog fallback and model-specific trainer/pool modules were removed. Existing recipes are available as YAML presets. The deployment script still selects only its three listed configurations; additional presets are opt-in. diff --git a/scripts/e2e_engine_definition.py b/scripts/e2e_engine_definition.py index feba8b1..b013ff7 100644 --- a/scripts/e2e_engine_definition.py +++ b/scripts/e2e_engine_definition.py @@ -17,7 +17,7 @@ from tinker import types from lilo.deployments import ResolvedDeployment -from lilo.providers.modal.recipe import backend_config +from lilo.backends.deployment import backend_config TIMEOUT = 3 * 60 * 60 diff --git a/src/lilo/backend_options.py b/src/lilo/backend_options.py new file mode 100644 index 0000000..831f8f6 --- /dev/null +++ b/src/lilo/backend_options.py @@ -0,0 +1,14 @@ +"""Helpers shared by backend-owned deployment configuration readers.""" + + +def native_options(values, protected): + if not isinstance(values, dict): + raise ValueError("backend configuration must be a mapping") + result = {} + for key, value in values.items(): + if not isinstance(key, str) or not key.isidentifier(): + raise ValueError(f"native option must use its underscore name: {key}") + if key in protected: + raise ValueError(f"option {key} is managed by Lilo") + result[key] = value + return result diff --git a/src/lilo/backends/deployment.py b/src/lilo/backends/deployment.py new file mode 100644 index 0000000..b4752f6 --- /dev/null +++ b/src/lilo/backends/deployment.py @@ -0,0 +1,28 @@ +"""Dispatch opaque YAML configuration to its backend-owned integration. + +These readers are CPU-only. Native libraries validate their options in workers. +""" + +from importlib import import_module + +TRAINERS = { + "miles": "lilo.backends.miles_deployment", + "megatron": "lilo.backends.megatron_deployment", +} +INFERENCE = {"sglang": "lilo.inference.sglang_deployment"} + + +def _reader(registry, backend): + try: + module = registry[backend] + except KeyError: + raise ValueError(f"unknown deployment backend: {backend}") from None + return import_module(module).build_config + + +def backend_config(spec, asset_path="/assets/pending"): + return _reader(TRAINERS, spec.trainer.backend)(spec, asset_path) + + +def serving_options(spec): + return _reader(INFERENCE, spec.inference.backend)(spec) diff --git a/src/lilo/backends/megatron_deployment.py b/src/lilo/backends/megatron_deployment.py new file mode 100644 index 0000000..c6dfce0 --- /dev/null +++ b/src/lilo/backends/megatron_deployment.py @@ -0,0 +1,87 @@ +"""Build Lilo loop settings while preserving native Megatron configuration.""" + +from dataclasses import asdict, fields + +from lilo.backend_options import native_options +from .megatron_runtime.common.config import EngineModelConfig, OptimizerConfig + +# These values also control packing, collectives, and checkpoint metadata in Lilo. +# Configure them once under runtime so the provider and training loop agree. +PROVIDER_MANAGED = { + "tensor_model_parallel_size", + "pipeline_model_parallel_size", + "virtual_pipeline_model_parallel_size", + "context_parallel_size", + "expert_model_parallel_size", + "expert_tensor_parallel_size", + "sequence_parallel", + "variable_seq_lengths", + "params_dtype", + "seq_length", +} +DISTRIBUTED_MANAGED = { + "use_distributed_optimizer", + "overlap_grad_reduce", + "overlap_param_gather", + "align_param_gather", +} +OPTIMIZER_MANAGED = { + "bf16", + "fp16", + "params_dtype", + "use_distributed_optimizer", + "overlap_param_gather", +} + + +def build_config(spec, asset_path): + trainer = spec.trainer + if spec.model.parameterization != "full": + raise ValueError("Megatron YAML deployments require full parameterization") + if trainer.engine.sampler_persistence_concurrency != 1: + raise ValueError("Megatron requires sampler_persistence_concurrency: 1") + sections = native_options(trainer.config, set()) + unknown = sections.keys() - {"runtime", "provider", "optimizer", "distributed"} + if unknown: + raise ValueError(f"unknown Megatron config sections: {sorted(unknown)}") + runtime = native_options( + sections.get("runtime", {}), + { + "hf_checkpoint", + "seq_length", + "max_lora_slots", + "max_lora_rank", + "optimizer", + "provider_overrides", + "optimizer_overrides", + "distributed_overrides", + }, + ) + provider = native_options(sections.get("provider", {}), PROVIDER_MANAGED) + optimizer = native_options(sections.get("optimizer", {}), OPTIMIZER_MANAGED) + distributed = native_options(sections.get("distributed", {}), DISTRIBUTED_MANAGED) + if optimizer.get("optimizer", "adam") != "adam": + raise ValueError("Tinker optim_step requires an Adam optimizer") + # Keep scheduling and per-request optimizer settings available to Lilo. Every + # other optimizer field goes to the installed Megatron constructor unchanged. + loop_fields = {field.name for field in fields(OptimizerConfig)} + loop_optimizer = { + key: value for key, value in optimizer.items() if key in loop_fields + } + native_optimizer = { + key: value for key, value in optimizer.items() if key not in loop_fields + } + try: + config = EngineModelConfig( + hf_checkpoint=asset_path, + seq_length=spec.model.max_context_length, + optimizer=OptimizerConfig(**loop_optimizer), + provider_overrides=provider, + optimizer_overrides=native_optimizer, + distributed_overrides=distributed, + **runtime, + ) + except TypeError as exc: + raise ValueError(f"invalid Megatron runtime options: {exc}") from exc + config.validate(trainer.resources.gpu_count) + return {"megatron": asdict(config), "checkpoint_dir": "/checkpoints"} diff --git a/src/lilo/backends/megatron_runtime/common/config.py b/src/lilo/backends/megatron_runtime/common/config.py index 785852a..8ac08cf 100644 --- a/src/lilo/backends/megatron_runtime/common/config.py +++ b/src/lilo/backends/megatron_runtime/common/config.py @@ -61,6 +61,8 @@ class EngineModelConfig: defer_fp32_logits: bool = False fp32_lm_head: bool = False provider_overrides: dict[str, object] = field(default_factory=dict) + optimizer_overrides: dict[str, object] = field(default_factory=dict) + distributed_overrides: dict[str, object] = field(default_factory=dict) overlap_grad_reduce: bool = False align_grad_reduce: bool = True diff --git a/src/lilo/backends/megatron_runtime/common/modeling.py b/src/lilo/backends/megatron_runtime/common/modeling.py index 97cdbf7..3f16ab8 100644 --- a/src/lilo/backends/megatron_runtime/common/modeling.py +++ b/src/lilo/backends/megatron_runtime/common/modeling.py @@ -61,6 +61,7 @@ def distributed_model( overlap_grad_reduce=config.overlap_grad_reduce, overlap_param_gather=config.overlap_param_gather, align_param_gather=config.align_param_gather, + **config.distributed_overrides, ), bf16=config.bf16, fp16=config.fp16, @@ -87,4 +88,5 @@ def optimizer_config( params_dtype=dtype, use_distributed_optimizer=distributed_optimizer, overlap_param_gather=config.overlap_param_gather, + **config.optimizer_overrides, ) diff --git a/src/lilo/backends/megatron_runtime/fft/checkpoint.py b/src/lilo/backends/megatron_runtime/fft/checkpoint.py index dd8ff09..68e0ea9 100644 --- a/src/lilo/backends/megatron_runtime/fft/checkpoint.py +++ b/src/lilo/backends/megatron_runtime/fft/checkpoint.py @@ -7,7 +7,7 @@ import os import shutil from copy import deepcopy -from dataclasses import asdict, dataclass +from dataclasses import asdict, dataclass, field from pathlib import Path from tempfile import NamedTemporaryFile from typing import Any @@ -57,11 +57,16 @@ class FFTCheckpointMetadata: has_optimizer: bool optimizer_state_format: str | None base_model_revision: str | None = None + native_optimizer_config: dict[str, Any] = field(default_factory=dict) + native_distributed_config: dict[str, Any] = field(default_factory=dict) def to_dict(self) -> dict[str, Any]: value = asdict(self) if self.base_model_revision is None: value.pop("base_model_revision") + for name in ("native_optimizer_config", "native_distributed_config"): + if not value[name]: + value.pop(name) return value def identity(self) -> str: @@ -100,11 +105,15 @@ def create_fft_checkpoint_metadata( precision="bf16" if config.bf16 else "fp16" if config.fp16 else "fp32", use_distributed_optimizer=config.use_distributed_optimizer, optimizer_config=asdict(config.optimizer), + native_optimizer_config=dict(config.optimizer_overrides), + native_distributed_config=dict(config.distributed_overrides), has_optimizer=include_optimizer, optimizer_state_format=( DISTRIBUTED_OPTIMIZER_STATE_FORMAT if include_optimizer and config.use_distributed_optimizer - else REGULAR_OPTIMIZER_STATE_FORMAT if include_optimizer else None + else REGULAR_OPTIMIZER_STATE_FORMAT + if include_optimizer + else None ), ) @@ -181,8 +190,7 @@ def restore_fft_optimizer_state( ) if not isinstance(state, dict) or state.get("format") != expected_format: raise ValueError( - "checkpoint optimizer state format mismatch: " - f"expected {expected_format!r}" + f"checkpoint optimizer state format mismatch: expected {expected_format!r}" ) state_dict = state.get("state_dict") if not isinstance(state_dict, dict | list): @@ -297,6 +305,8 @@ def load_fft_training_checkpoint( "precision", "use_distributed_optimizer", "optimizer_config", + "native_optimizer_config", + "native_distributed_config", "optimizer_state_format", ) for name in comparable_fields: @@ -369,7 +379,9 @@ def synchronize_checkpoint_preflight( def _materialize_local_sharded_state(value): """Detach Megatron's local torch_dist representation from live optimizer buffers.""" if isinstance(value, ShardedTensorFactory): - raise TypeError("distributed optimizer state contains an unresolved tensor factory") + raise TypeError( + "distributed optimizer state contains an unresolved tensor factory" + ) if isinstance(value, ShardedTensor): if value.data is None: raise ValueError("distributed optimizer sharded tensor has no local data") @@ -380,8 +392,7 @@ def _materialize_local_sharded_state(value): return _materialize_local_sharded_state(value.unwrap()) if isinstance(value, dict): return { - key: _materialize_local_sharded_state(item) - for key, item in value.items() + key: _materialize_local_sharded_state(item) for key, item in value.items() } if isinstance(value, list): return [_materialize_local_sharded_state(item) for item in value] @@ -394,7 +405,9 @@ def _validate_distributed_optimizer_state_dict(state_dict: dict | list) -> None: """Require every DistributedOptimizer leaf to contain its local tensor shards.""" if isinstance(state_dict, list): if not state_dict: - raise ValueError("distributed optimizer checkpoint has no optimizer entries") + raise ValueError( + "distributed optimizer checkpoint has no optimizer entries" + ) for item in state_dict: if not isinstance(item, dict | list): raise TypeError("distributed optimizer entry must be a dict or list") @@ -403,17 +416,22 @@ def _validate_distributed_optimizer_state_dict(state_dict: dict | list) -> None: if state_dict.get("param_state_sharding_type") == "dp_reshardable": if not isinstance(state_dict.get("optimizer"), dict): - raise ValueError("distributed optimizer checkpoint has no optimizer metadata") - if not isinstance(state_dict.get("param_state"), dict) or not state_dict[ - "param_state" - ]: + raise ValueError( + "distributed optimizer checkpoint has no optimizer metadata" + ) + if ( + not isinstance(state_dict.get("param_state"), dict) + or not state_dict["param_state"] + ): raise ValueError("distributed optimizer checkpoint has no parameter state") return if state_dict and all(isinstance(key, int) for key in state_dict): for item in state_dict.values(): if not isinstance(item, dict | list): - raise TypeError("chained distributed optimizer entry must be a dict or list") + raise TypeError( + "chained distributed optimizer entry must be a dict or list" + ) _validate_distributed_optimizer_state_dict(item) return diff --git a/src/lilo/backends/miles_deployment.py b/src/lilo/backends/miles_deployment.py new file mode 100644 index 0000000..caa22bd --- /dev/null +++ b/src/lilo/backends/miles_deployment.py @@ -0,0 +1,90 @@ +"""Miles deployment integration. Native options are validated by Miles at startup.""" + +from dataclasses import asdict + +from lilo.backend_options import native_options +from .miles_config import MilesBackendConfig + +MILES_MANAGED = { + "hf_checkpoint", + "load", + "pretrained_checkpoint", + "train_backend", + "actor_num_nodes", + "actor_num_gpus_per_node", + "rollout_num_gpus", + "debug_train_only", + "megatron_to_hf_mode", + "seq_length", + "pipeline_model_parallel_size", + "virtual_pipeline_model_parallel_size", + "colocate", + "custom_actor", + "sglang_model_path", + "use_dynamic_global_batch_size", + "delay_split_train_data_by_dp", + "use_dynamic_batch_size", + "optimizer", + "gradient_accumulation_fusion", + "save", + "save_interval", + "ckpt_step", + "rollout_num_gpus_per_engine", + "num_gpus_per_node", +} + + +def build_config(spec, asset_path): + trainer = spec.trainer + if spec.model.parameterization != "lora": + raise ValueError("Miles requires lora parameterization") + values = native_options(trainer.config, set()) + unknown = values.keys() - {"model_args", "options"} + if unknown: + raise ValueError(f"unknown Miles config sections: {sorted(unknown)}") + options = native_options(values.get("options", {}), MILES_MANAGED) + names = { + "tensor_model_parallel_size": "tensor_model_parallel_size", + "context_parallel_size": "context_parallel_size", + "expert_model_parallel_size": "expert_model_parallel_size", + "expert_tensor_parallel_size": "expert_tensor_parallel_size", + "multi_lora_n_adapters": "max_lora_slots", + "lora_rank": "max_lora_rank", + "lora_alpha": "default_lora_alpha", + "lora_dropout": "lora_dropout", + "target_modules": "target_modules", + "max_tokens_per_gpu": "max_tokens_per_gpu", + } + settings = { + dest: options.pop(source) for source, dest in names.items() if source in options + } + for name, value in settings.items(): + if ( + name not in {"target_modules", "default_lora_alpha", "lora_dropout"} + and type(value) is not int + ): + raise ValueError(f"Miles option {name} must be an integer") + if "target_modules" in settings: + value = settings["target_modules"] + value = value.split(",") if isinstance(value, str) else value + if not isinstance(value, list) or not all( + isinstance(item, str) and item for item in value + ): + raise ValueError("target_modules must be a nonempty list of module names") + settings["target_modules"] = tuple(value) + config = MilesBackendConfig( + hf_checkpoint=asset_path, + model_type=values.get("model_args") or "", + actor_num_gpus_per_node=trainer.resources.gpu_count, + native_options=options, + extra_args=("--seq-length", str(spec.model.max_context_length)), + **settings, + ) + config.validate() + if config.world_size % ( + config.expert_model_parallel_size * config.expert_tensor_parallel_size + ): + raise ValueError("expert parallel sizes must divide the trainer GPU allocation") + if trainer.engine.max_clients_per_instance > config.max_lora_slots: + raise ValueError("max_clients_per_instance exceeds multi_lora_n_adapters") + return {"miles": asdict(config), "checkpoint_dir": "/checkpoints"} diff --git a/src/lilo/deployments.py b/src/lilo/deployments.py index fd9440d..ae9d96c 100644 --- a/src/lilo/deployments.py +++ b/src/lilo/deployments.py @@ -65,30 +65,20 @@ class EngineOptions(StrictModel): sampler_persistence_concurrency: int = Field(default=8, gt=0) -class MilesOptions(StrictModel): - model_args: str | None = None - options: dict[str, Any] = Field(default_factory=dict) - - -class NativeOptions(StrictModel): - options: dict[str, Any] = Field(default_factory=dict) - - class Trainer(StrictModel): - backend: Literal["miles", "megatron"] = "miles" + backend: str = "miles" resources: Resources scaling: TrainerScaling = Field(default_factory=TrainerScaling) engine: EngineOptions = Field(default_factory=EngineOptions) - miles: MilesOptions | None = None - megatron: NativeOptions | None = None + config: dict[str, Any] = Field(default_factory=dict) env: dict[str, str] = Field(default_factory=dict) class Inference(StrictModel): - backend: Literal["sglang"] = "sglang" + backend: str = "sglang" resources: Resources scaling: InferenceScaling = Field(default_factory=InferenceScaling) - sglang: NativeOptions = Field(default_factory=NativeOptions) + config: dict[str, Any] = Field(default_factory=dict) env: dict[str, str] = Field(default_factory=dict) @@ -137,39 +127,17 @@ class DeploymentSpec(StrictModel): @model_validator(mode="after") def compatible(self): - from lilo.providers.modal.recipe import backend_config, serving_options + from lilo.backends.deployment import backend_config, serving_options for role in (self.trainer, self.inference): role.resources.gpu_count if any(k.startswith("LILO_") for k in role.env): raise ValueError("LILO_ environment variables are managed by Lilo") - if self.trainer.backend == "miles": - if ( - self.model.parameterization != "lora" - or self.trainer.miles is None - or self.trainer.megatron is not None - ): - raise ValueError( - "Miles requires lora parameterization and trainer.miles" - ) - elif ( - self.model.parameterization != "full" - or self.trainer.megatron is None - or self.trainer.miles is not None - ): - raise ValueError( - "Megatron YAML deployments require full parameterization and trainer.megatron" - ) if ( self.model.parameterization == "full" and self.trainer.engine.max_clients_per_instance != 1 ): raise ValueError("FFT trainers admit one client per instance") - if ( - self.trainer.backend == "megatron" - and self.trainer.engine.sampler_persistence_concurrency != 1 - ): - raise ValueError("Megatron requires sampler_persistence_concurrency: 1") try: backend_config(self) serving_options(self) diff --git a/src/lilo/inference/sglang_deployment.py b/src/lilo/inference/sglang_deployment.py new file mode 100644 index 0000000..fe5ef3b --- /dev/null +++ b/src/lilo/inference/sglang_deployment.py @@ -0,0 +1,60 @@ +"""SGLang settings that must agree with Lilo replica orchestration.""" + +from lilo.backend_options import native_options + +SGLANG_MANAGED = { + "model_path", + "model", + "host", + "port", + "context_length", + "enable_lora", + "max_lora_rank", + "enable_cpu_weight_cache", + "api_key", + "pp_size", + "lora_paths", + "dist_init_addr", + "nnodes", + "node_rank", + "tokenizer_path", + "tokenizer_revision", + "revision", + "grpc_mode", + "smg_grpc_mode", + "encoder_only", + "use_ray", + "disaggregation_mode", + "skip_tokenizer_init", +} + + +def build_config(spec): + options = native_options(spec.inference.config, SGLANG_MANAGED) + tp = options.get("tp_size", spec.inference.resources.gpu_count) + ep = options.get("ep_size", 1) + if not isinstance(tp, int) or tp != spec.inference.resources.gpu_count: + raise ValueError("sglang.tp_size must equal the replica GPU allocation") + if not isinstance(ep, int) or ep < 1 or tp % ep: + raise ValueError("sglang.ep_size must divide the replica GPU allocation") + dp = options.get("dp_size", 1) + dp_attention = options.get("enable_dp_attention", False) + if type(dp) is not int or dp < 1 or tp % dp: + raise ValueError("sglang.dp_size must divide the replica GPU allocation") + if not isinstance(dp_attention, bool): + raise ValueError("sglang.enable_dp_attention must be a boolean") + if dp > 1 and not dp_attention: + raise ValueError("sglang.dp_size > 1 requires enable_dp_attention") + for key in ( + "max_loaded_loras", + "max_loras_per_batch", + "max_running_requests", + "max_queued_requests", + ): + if key in options and (not isinstance(options[key], int) or options[key] < 1): + raise ValueError(f"sglang.{key} must be positive") + if not 0 < options.get("mem_fraction_static", 0.8) < 1: + raise ValueError("sglang.mem_fraction_static must be between zero and one") + if options.get("max_loaded_loras", 64) < options.get("max_loras_per_batch", 8): + raise ValueError("max_loaded_loras must be >= max_loras_per_batch") + return options diff --git a/src/lilo/presets/qwen35-35b-a3b-fft-64k.yaml b/src/lilo/presets/qwen35-35b-a3b-fft-64k.yaml index fa5761c..5b64a25 100644 --- a/src/lilo/presets/qwen35-35b-a3b-fft-64k.yaml +++ b/src/lilo/presets/qwen35-35b-a3b-fft-64k.yaml @@ -13,8 +13,11 @@ trainer: engine: max_clients_per_instance: 1 sampler_persistence_concurrency: 1 - megatron: - options: + env: + PYTORCH_CUDA_ALLOC_CONF: expandable_segments:True + TORCHINDUCTOR_COMPILE_THREADS: '1' + config: + runtime: tensor_model_parallel_size: 4 pipeline_model_parallel_size: 1 context_parallel_size: 2 @@ -27,23 +30,20 @@ trainer: fp16: false gpu_memory_fraction: 0.9 use_distributed_optimizer: true - provider_overrides: - mtp_num_layers: 0 - recompute_granularity: selective - moe_layer_recompute: true - moe_token_dispatcher_type: alltoall - moe_router_fusion: true - moe_permute_fusion: true - moe_grouped_gemm: true - moe_shared_expert_overlap: false - moe_aux_loss_coeff: 0.0 - optimizer: - optimizer: adam - lr: 0.0001 - min_lr: 0.0001 - env: - PYTORCH_CUDA_ALLOC_CONF: expandable_segments:True - TORCHINDUCTOR_COMPILE_THREADS: '1' + provider: + mtp_num_layers: 0 + recompute_granularity: selective + moe_layer_recompute: true + moe_token_dispatcher_type: alltoall + moe_router_fusion: true + moe_permute_fusion: true + moe_grouped_gemm: true + moe_shared_expert_overlap: false + moe_aux_loss_coeff: 0.0 + optimizer: + optimizer: adam + lr: 0.0001 + min_lr: 0.0001 inference: resources: gpu: H200:4 @@ -51,13 +51,12 @@ inference: min_replicas: 0 max_replicas: 8 target_concurrency: 16 - sglang: - options: - tp_size: 4 - ep_size: 4 - mem_fraction_static: 0.9 - max_running_requests: 32 - max_queued_requests: 4 - cpu_weight_cache_max_compile_group_gb: 32 - dp_size: 4 - enable_dp_attention: true + config: + tp_size: 4 + ep_size: 4 + mem_fraction_static: 0.9 + max_running_requests: 32 + max_queued_requests: 4 + cpu_weight_cache_max_compile_group_gb: 32 + dp_size: 4 + enable_dp_attention: true diff --git a/src/lilo/presets/qwen35-4b-fft-64k.yaml b/src/lilo/presets/qwen35-4b-fft-64k.yaml index b8015a7..2de6f4e 100644 --- a/src/lilo/presets/qwen35-4b-fft-64k.yaml +++ b/src/lilo/presets/qwen35-4b-fft-64k.yaml @@ -16,8 +16,8 @@ trainer: engine: max_clients_per_instance: 1 sampler_persistence_concurrency: 1 - megatron: - options: + config: + runtime: tensor_model_parallel_size: 2 context_parallel_size: 2 sequence_parallel: true @@ -26,22 +26,21 @@ trainer: defer_fp32_logits: true fp32_lm_head: true use_distributed_optimizer: true - provider_overrides: - mtp_num_layers: 0 - recompute_granularity: full - recompute_method: uniform - recompute_num_layers: 1 - optimizer: - lr: 0.0001 - min_lr: 0.0001 - loss_scale: 1.0 + provider: + mtp_num_layers: 0 + recompute_granularity: full + recompute_method: uniform + recompute_num_layers: 1 + optimizer: + lr: 0.0001 + min_lr: 0.0001 + loss_scale: 1.0 inference: resources: gpu: H100:1 - sglang: - options: - tp_size: 1 - mem_fraction_static: 0.85 - max_running_requests: 32 - max_queued_requests: 4 - cpu_weight_cache_max_compile_group_gb: 16 + config: + tp_size: 1 + mem_fraction_static: 0.85 + max_running_requests: 32 + max_queued_requests: 4 + cpu_weight_cache_max_compile_group_gb: 16 diff --git a/src/lilo/presets/qwen35-9b-fft-64k.yaml b/src/lilo/presets/qwen35-9b-fft-64k.yaml index 9c324f5..79ae97c 100644 --- a/src/lilo/presets/qwen35-9b-fft-64k.yaml +++ b/src/lilo/presets/qwen35-9b-fft-64k.yaml @@ -13,8 +13,11 @@ trainer: engine: max_clients_per_instance: 1 sampler_persistence_concurrency: 1 - megatron: - options: + env: + PYTORCH_CUDA_ALLOC_CONF: expandable_segments:True + TORCHINDUCTOR_COMPILE_THREADS: '1' + config: + runtime: tensor_model_parallel_size: 2 context_parallel_size: 2 sequence_parallel: true @@ -23,18 +26,15 @@ trainer: defer_fp32_logits: true fp32_lm_head: true use_distributed_optimizer: true - provider_overrides: - mtp_num_layers: 0 - recompute_granularity: full - recompute_method: uniform - recompute_num_layers: 1 - optimizer: - lr: 0.0001 - min_lr: 0.0001 - loss_scale: 1.0 - env: - PYTORCH_CUDA_ALLOC_CONF: expandable_segments:True - TORCHINDUCTOR_COMPILE_THREADS: '1' + provider: + mtp_num_layers: 0 + recompute_granularity: full + recompute_method: uniform + recompute_num_layers: 1 + optimizer: + lr: 0.0001 + min_lr: 0.0001 + loss_scale: 1.0 inference: resources: gpu: H200:1 @@ -42,11 +42,10 @@ inference: min_replicas: 0 max_replicas: 8 target_concurrency: 16 - sglang: - options: - tp_size: 1 - ep_size: 1 - mem_fraction_static: 0.85 - max_running_requests: 32 - max_queued_requests: 4 - cpu_weight_cache_max_compile_group_gb: 16 + config: + tp_size: 1 + ep_size: 1 + mem_fraction_static: 0.85 + max_running_requests: 32 + max_queued_requests: 4 + cpu_weight_cache_max_compile_group_gb: 16 diff --git a/src/lilo/presets/qwen35-9b-instruct-lora-16k-dp2.yaml b/src/lilo/presets/qwen35-9b-instruct-lora-16k-dp2.yaml index bfc36b0..4c2d8f0 100644 --- a/src/lilo/presets/qwen35-9b-instruct-lora-16k-dp2.yaml +++ b/src/lilo/presets/qwen35-9b-instruct-lora-16k-dp2.yaml @@ -3,6 +3,6 @@ name: qwen35-9b-instruct-lora-16k-dp2 routing: default: false trainer: - miles: + config: options: tensor_model_parallel_size: 4 diff --git a/src/lilo/presets/qwen35-9b-instruct-lora-16k.yaml b/src/lilo/presets/qwen35-9b-instruct-lora-16k.yaml index 083c193..dee8021 100644 --- a/src/lilo/presets/qwen35-9b-instruct-lora-16k.yaml +++ b/src/lilo/presets/qwen35-9b-instruct-lora-16k.yaml @@ -13,7 +13,10 @@ trainer: engine: max_clients_per_instance: 6 sampler_persistence_concurrency: 8 - miles: + env: + PYTORCH_CUDA_ALLOC_CONF: expandable_segments:True + TORCHINDUCTOR_COMPILE_THREADS: '1' + config: model_args: qwen3.5-9B options: tensor_model_parallel_size: 8 @@ -29,9 +32,6 @@ trainer: multi_lora_n_adapters: 6 lora_rank: 32 lora_alpha: 32 - env: - PYTORCH_CUDA_ALLOC_CONF: expandable_segments:True - TORCHINDUCTOR_COMPILE_THREADS: '1' inference: resources: gpu: H200:1 @@ -39,21 +39,20 @@ inference: min_replicas: 0 max_replicas: 8 target_concurrency: 16 - sglang: - options: - tp_size: 1 - ep_size: 1 - mem_fraction_static: 0.8 - max_running_requests: 32 - max_queued_requests: 8 - max_loaded_loras: 256 - max_loras_per_batch: 8 - lora_target_modules: - - q_proj - - k_proj - - v_proj - - o_proj - - gate_proj - - up_proj - - down_proj - schedule_policy: lpm + config: + tp_size: 1 + ep_size: 1 + mem_fraction_static: 0.8 + max_running_requests: 32 + max_queued_requests: 8 + max_loaded_loras: 256 + max_loras_per_batch: 8 + lora_target_modules: + - q_proj + - k_proj + - v_proj + - o_proj + - gate_proj + - up_proj + - down_proj + schedule_policy: lpm diff --git a/src/lilo/presets/qwen35-9b-lora-16k.yaml b/src/lilo/presets/qwen35-9b-lora-16k.yaml index cf9ee63..8c1df86 100644 --- a/src/lilo/presets/qwen35-9b-lora-16k.yaml +++ b/src/lilo/presets/qwen35-9b-lora-16k.yaml @@ -21,21 +21,26 @@ trainer: engine: max_clients_per_instance: 6 sampler_persistence_concurrency: 8 - miles: + env: + PYTORCH_CUDA_ALLOC_CONF: expandable_segments:True + TORCHINDUCTOR_COMPILE_THREADS: '1' + config: model_args: qwen3.5-9B options: tensor_model_parallel_size: 4 multi_lora_n_adapters: 6 lora_rank: 32 lora_alpha: 32 - target_modules: [linear_qkv, linear_proj, linear_fc1, linear_fc2, output_layer] + target_modules: + - linear_qkv + - linear_proj + - linear_fc1 + - linear_fc2 + - output_layer max_tokens_per_gpu: 16384 recompute_granularity: full recompute_method: uniform recompute_num_layers: 1 - env: - PYTORCH_CUDA_ALLOC_CONF: expandable_segments:True - TORCHINDUCTOR_COMPILE_THREADS: '1' inference: backend: sglang resources: @@ -44,11 +49,10 @@ inference: min_replicas: 0 max_replicas: 8 target_concurrency: 16 - sglang: - options: - tp_size: 1 - mem_fraction_static: 0.8 - max_running_requests: 32 - max_queued_requests: 8 - max_loaded_loras: 64 - max_loras_per_batch: 8 + config: + tp_size: 1 + mem_fraction_static: 0.8 + max_running_requests: 32 + max_queued_requests: 8 + max_loaded_loras: 64 + max_loras_per_batch: 8 diff --git a/src/lilo/presets/qwen35-9b-lora-2k.yaml b/src/lilo/presets/qwen35-9b-lora-2k.yaml index 9d1bdef..23f91e4 100644 --- a/src/lilo/presets/qwen35-9b-lora-2k.yaml +++ b/src/lilo/presets/qwen35-9b-lora-2k.yaml @@ -13,7 +13,10 @@ trainer: engine: max_clients_per_instance: 4 sampler_persistence_concurrency: 8 - miles: + env: + PYTORCH_CUDA_ALLOC_CONF: expandable_segments:True + TORCHINDUCTOR_COMPILE_THREADS: '1' + config: model_args: qwen3.5-9B options: tensor_model_parallel_size: 4 @@ -30,9 +33,6 @@ trainer: multi_lora_n_adapters: 4 lora_rank: 32 lora_alpha: 32 - env: - PYTORCH_CUDA_ALLOC_CONF: expandable_segments:True - TORCHINDUCTOR_COMPILE_THREADS: '1' inference: resources: gpu: H200:1 @@ -40,22 +40,21 @@ inference: min_replicas: 0 max_replicas: 8 target_concurrency: 16 - sglang: - options: - tp_size: 1 - ep_size: 1 - mem_fraction_static: 0.8 - max_running_requests: 32 - max_queued_requests: 8 - max_loaded_loras: 32 - max_loras_per_batch: 8 - lora_target_modules: - - q_proj - - k_proj - - v_proj - - o_proj - - gate_proj - - up_proj - - down_proj - - lm_head - schedule_policy: lpm + config: + tp_size: 1 + ep_size: 1 + mem_fraction_static: 0.8 + max_running_requests: 32 + max_queued_requests: 8 + max_loaded_loras: 32 + max_loras_per_batch: 8 + lora_target_modules: + - q_proj + - k_proj + - v_proj + - o_proj + - gate_proj + - up_proj + - down_proj + - lm_head + schedule_policy: lpm diff --git a/src/lilo/presets/qwen35-9b-lora-64k.yaml b/src/lilo/presets/qwen35-9b-lora-64k.yaml index a434fdd..0f6e9f8 100644 --- a/src/lilo/presets/qwen35-9b-lora-64k.yaml +++ b/src/lilo/presets/qwen35-9b-lora-64k.yaml @@ -7,7 +7,7 @@ routing: trainer: resources: gpu: H200:8 - miles: + config: options: tensor_model_parallel_size: 8 max_tokens_per_gpu: 65536 diff --git a/src/lilo/presets/qwen36-27b-fft-64k.yaml b/src/lilo/presets/qwen36-27b-fft-64k.yaml index 6bea537..9dd7446 100644 --- a/src/lilo/presets/qwen36-27b-fft-64k.yaml +++ b/src/lilo/presets/qwen36-27b-fft-64k.yaml @@ -13,8 +13,11 @@ trainer: engine: max_clients_per_instance: 1 sampler_persistence_concurrency: 1 - megatron: - options: + env: + PYTORCH_CUDA_ALLOC_CONF: expandable_segments:True + TORCHINDUCTOR_COMPILE_THREADS: '1' + config: + runtime: tensor_model_parallel_size: 4 pipeline_model_parallel_size: 1 context_parallel_size: 2 @@ -25,18 +28,15 @@ trainer: fp16: false gpu_memory_fraction: 0.9 use_distributed_optimizer: true - provider_overrides: - mtp_num_layers: 0 - recompute_granularity: full - recompute_method: uniform - recompute_num_layers: 1 - optimizer: - optimizer: adam - lr: 0.0001 - min_lr: 0.0001 - env: - PYTORCH_CUDA_ALLOC_CONF: expandable_segments:True - TORCHINDUCTOR_COMPILE_THREADS: '1' + provider: + mtp_num_layers: 0 + recompute_granularity: full + recompute_method: uniform + recompute_num_layers: 1 + optimizer: + optimizer: adam + lr: 0.0001 + min_lr: 0.0001 inference: resources: gpu: H200:4 @@ -44,11 +44,10 @@ inference: min_replicas: 0 max_replicas: 8 target_concurrency: 16 - sglang: - options: - tp_size: 4 - ep_size: 1 - mem_fraction_static: 0.9 - max_running_requests: 32 - max_queued_requests: 4 - cpu_weight_cache_max_compile_group_gb: 32 + config: + tp_size: 4 + ep_size: 1 + mem_fraction_static: 0.9 + max_running_requests: 32 + max_queued_requests: 4 + cpu_weight_cache_max_compile_group_gb: 32 diff --git a/src/lilo/presets/qwen36-35b-a3b-fft-64k.yaml b/src/lilo/presets/qwen36-35b-a3b-fft-64k.yaml index 03a9ae7..839bf2b 100644 --- a/src/lilo/presets/qwen36-35b-a3b-fft-64k.yaml +++ b/src/lilo/presets/qwen36-35b-a3b-fft-64k.yaml @@ -3,7 +3,6 @@ name: qwen36-35b-a3b-fft-64k model: id: Qwen/Qwen3.6-35B-A3B inference: - sglang: - options: - dp_size: 1 - enable_dp_attention: false + config: + dp_size: 1 + enable_dp_attention: false diff --git a/src/lilo/presets/qwen38-27b-lora-128k.yaml b/src/lilo/presets/qwen38-27b-lora-128k.yaml index f6abba9..aae6b67 100644 --- a/src/lilo/presets/qwen38-27b-lora-128k.yaml +++ b/src/lilo/presets/qwen38-27b-lora-128k.yaml @@ -5,7 +5,7 @@ model: routing: default: false trainer: - miles: + config: options: tensor_model_parallel_size: 2 context_parallel_size: 4 @@ -16,7 +16,6 @@ inference: scaling: max_replicas: 4 target_concurrency: 4 - sglang: - options: - tp_size: 2 - max_running_requests: 8 + config: + tp_size: 2 + max_running_requests: 8 diff --git a/src/lilo/presets/qwen38-27b-lora-16k.yaml b/src/lilo/presets/qwen38-27b-lora-16k.yaml index 118a28a..7215db1 100644 --- a/src/lilo/presets/qwen38-27b-lora-16k.yaml +++ b/src/lilo/presets/qwen38-27b-lora-16k.yaml @@ -13,7 +13,10 @@ trainer: engine: max_clients_per_instance: 6 sampler_persistence_concurrency: 8 - miles: + env: + PYTORCH_CUDA_ALLOC_CONF: expandable_segments:True + TORCHINDUCTOR_COMPILE_THREADS: '1' + config: model_args: qwen3.8-27B options: tensor_model_parallel_size: 4 @@ -29,9 +32,6 @@ trainer: multi_lora_n_adapters: 6 lora_rank: 32 lora_alpha: 32 - env: - PYTORCH_CUDA_ALLOC_CONF: expandable_segments:True - TORCHINDUCTOR_COMPILE_THREADS: '1' inference: resources: gpu: H200:1 @@ -39,21 +39,20 @@ inference: min_replicas: 0 max_replicas: 8 target_concurrency: 16 - sglang: - options: - tp_size: 1 - ep_size: 1 - mem_fraction_static: 0.8 - max_running_requests: 32 - max_queued_requests: 8 - max_loaded_loras: 256 - max_loras_per_batch: 8 - lora_target_modules: - - q_proj - - k_proj - - v_proj - - o_proj - - gate_proj - - up_proj - - down_proj - schedule_policy: lpm + config: + tp_size: 1 + ep_size: 1 + mem_fraction_static: 0.8 + max_running_requests: 32 + max_queued_requests: 8 + max_loaded_loras: 256 + max_loras_per_batch: 8 + lora_target_modules: + - q_proj + - k_proj + - v_proj + - o_proj + - gate_proj + - up_proj + - down_proj + schedule_policy: lpm diff --git a/src/lilo/presets/qwen38-27b-lora-64k.yaml b/src/lilo/presets/qwen38-27b-lora-64k.yaml index 4a9f127..67128f4 100644 --- a/src/lilo/presets/qwen38-27b-lora-64k.yaml +++ b/src/lilo/presets/qwen38-27b-lora-64k.yaml @@ -5,13 +5,12 @@ model: routing: default: false trainer: - miles: + config: options: context_parallel_size: 2 max_tokens_per_gpu: 32768 inference: scaling: target_concurrency: 8 - sglang: - options: - max_running_requests: 16 + config: + max_running_requests: 16 diff --git a/src/lilo/providers/modal/recipe.py b/src/lilo/providers/modal/recipe.py deleted file mode 100644 index 98c18ff..0000000 --- a/src/lilo/providers/modal/recipe.py +++ /dev/null @@ -1,197 +0,0 @@ -"""Translate a deployment recipe into backend and serving configurations.""" - -from __future__ import annotations - -from dataclasses import asdict - -from lilo.backends.miles_config import MilesBackendConfig - -MILES_MANAGED = { - "hf_checkpoint", - "load", - "pretrained_checkpoint", - "train_backend", - "actor_num_nodes", - "actor_num_gpus_per_node", - "rollout_num_gpus", - "debug_train_only", - "megatron_to_hf_mode", - "seq_length", - "pipeline_model_parallel_size", - "virtual_pipeline_model_parallel_size", - "colocate", - "custom_actor", - "sglang_model_path", - "use_dynamic_global_batch_size", - "delay_split_train_data_by_dp", - "use_dynamic_batch_size", - "optimizer", - "gradient_accumulation_fusion", - "save", - "save_interval", - "ckpt_step", - "rollout_num_gpus_per_engine", - "num_gpus_per_node", -} -SGLANG_MANAGED = { - "model_path", - "model", - "host", - "port", - "context_length", - "enable_lora", - "max_lora_rank", - "enable_cpu_weight_cache", - "api_key", - "pp_size", - "lora_paths", - "dist_init_addr", - "nnodes", - "node_rank", - "tokenizer_path", - "tokenizer_revision", - "revision", - "grpc_mode", - "smg_grpc_mode", - "encoder_only", - "use_ray", - "disaggregation_mode", - "skip_tokenizer_init", -} - - -def native_options(values, protected): - result = {} - for key, value in values.items(): - if key.startswith("-") or key.replace("_", "").isalnum() is False: - raise ValueError(f"native option must use its underscore name: {key}") - if key in protected: - raise ValueError(f"option {key} is managed by Lilo") - result[key] = value - return result - - -def backend_config(spec, asset_path="/assets/pending"): - trainer = spec.trainer - if trainer.backend == "miles": - options = native_options(trainer.miles.options, MILES_MANAGED) - names = { - "tensor_model_parallel_size": "tensor_model_parallel_size", - "context_parallel_size": "context_parallel_size", - "expert_model_parallel_size": "expert_model_parallel_size", - "expert_tensor_parallel_size": "expert_tensor_parallel_size", - "multi_lora_n_adapters": "max_lora_slots", - "lora_rank": "max_lora_rank", - "lora_alpha": "default_lora_alpha", - "lora_dropout": "lora_dropout", - "target_modules": "target_modules", - "max_tokens_per_gpu": "max_tokens_per_gpu", - } - settings = { - dest: options.pop(source) - for source, dest in names.items() - if source in options - } - for name, value in settings.items(): - if ( - name not in {"target_modules", "default_lora_alpha", "lora_dropout"} - and type(value) is not int - ): - raise ValueError(f"trainer.miles option {name} must be an integer") - if "target_modules" in settings: - value = settings["target_modules"] - value = value.split(",") if isinstance(value, str) else value - if not isinstance(value, list) or not all( - isinstance(item, str) and item for item in value - ): - raise ValueError( - "target_modules must be a nonempty list of module names" - ) - settings["target_modules"] = tuple(value) - config = MilesBackendConfig( - hf_checkpoint=asset_path, - model_type=trainer.miles.model_args or "", - actor_num_gpus_per_node=trainer.resources.gpu_count, - native_options=options, - extra_args=("--seq-length", str(spec.model.max_context_length)), - **settings, - ) - config.validate() - if config.world_size % ( - config.expert_model_parallel_size * config.expert_tensor_parallel_size - ): - raise ValueError( - "expert parallel sizes must divide the trainer GPU allocation" - ) - if trainer.engine.max_clients_per_instance > config.max_lora_slots: - raise ValueError("max_clients_per_instance exceeds multi_lora_n_adapters") - return {"miles": asdict(config), "checkpoint_dir": "/checkpoints"} - from lilo.backends.megatron_runtime.common.config import ( - EngineModelConfig, - OptimizerConfig, - ) - - options = native_options( - trainer.megatron.options, - {"hf_checkpoint", "seq_length", "max_lora_slots", "max_lora_rank"}, - ) - overrides = options.get("provider_overrides", {}) - protected = { - "tensor_model_parallel_size", - "pipeline_model_parallel_size", - "virtual_pipeline_model_parallel_size", - "context_parallel_size", - "expert_model_parallel_size", - "expert_tensor_parallel_size", - "sequence_parallel", - "variable_seq_lengths", - "params_dtype", - "seq_length", - } - native_options(overrides, protected) - try: - optimizer = OptimizerConfig(**options.pop("optimizer", {})) - except TypeError as exc: - raise ValueError(f"invalid trainer.megatron optimizer options: {exc}") from exc - try: - config = EngineModelConfig( - hf_checkpoint=asset_path, - seq_length=spec.model.max_context_length, - optimizer=optimizer, - **options, - ) - except TypeError as exc: - raise ValueError(f"invalid trainer.megatron options: {exc}") from exc - config.validate(trainer.resources.gpu_count) - return {"megatron": asdict(config), "checkpoint_dir": "/checkpoints"} - - -def serving_options(spec): - options = native_options(spec.inference.sglang.options, SGLANG_MANAGED) - tp = options.get("tp_size", spec.inference.resources.gpu_count) - ep = options.get("ep_size", 1) - if not isinstance(tp, int) or tp != spec.inference.resources.gpu_count: - raise ValueError("sglang.tp_size must equal the replica GPU allocation") - if not isinstance(ep, int) or ep < 1 or tp % ep: - raise ValueError("sglang.ep_size must divide the replica GPU allocation") - dp = options.get("dp_size", 1) - dp_attention = options.get("enable_dp_attention", False) - if type(dp) is not int or dp < 1 or tp % dp: - raise ValueError("sglang.dp_size must divide the replica GPU allocation") - if not isinstance(dp_attention, bool): - raise ValueError("sglang.enable_dp_attention must be a boolean") - if dp > 1 and not dp_attention: - raise ValueError("sglang.dp_size > 1 requires enable_dp_attention") - for key in ( - "max_loaded_loras", - "max_loras_per_batch", - "max_running_requests", - "max_queued_requests", - ): - if key in options and (not isinstance(options[key], int) or options[key] < 1): - raise ValueError(f"sglang.{key} must be positive") - if not 0 < options.get("mem_fraction_static", 0.8) < 1: - raise ValueError("sglang.mem_fraction_static must be between zero and one") - if options.get("max_loaded_loras", 64) < options.get("max_loras_per_batch", 8): - raise ValueError("max_loaded_loras must be >= max_loras_per_batch") - return options diff --git a/src/lilo/providers/modal/yaml_apps.py b/src/lilo/providers/modal/yaml_apps.py index f43b11b..45385f1 100644 --- a/src/lilo/providers/modal/yaml_apps.py +++ b/src/lilo/providers/modal/yaml_apps.py @@ -13,7 +13,7 @@ import modal from lilo.deployments import ResolvedDeployment, Routing, validate_frontend -from .recipe import backend_config, serving_options +from lilo.backends.deployment import backend_config, serving_options MANIFEST_ENV = "LILO_DEPLOYMENT_MANIFEST" POOL_CONFIG_ENV = "LILO_POOL_DEPLOYMENT" diff --git a/tests/backends/test_megatron_fft.py b/tests/backends/test_megatron_fft.py index 4225707..7abb965 100644 --- a/tests/backends/test_megatron_fft.py +++ b/tests/backends/test_megatron_fft.py @@ -416,9 +416,7 @@ def sharded_state_dict(self, **kwargs): "metadata": {"distrib_optim_sharding_type": "dp_reshardable"}, } ] - assert captured["format"] == ( - fft_checkpoint.DISTRIBUTED_OPTIMIZER_STATE_FORMAT - ) + assert captured["format"] == (fft_checkpoint.DISTRIBUTED_OPTIMIZER_STATE_FORMAT) tensors = captured["state_dict"]["param_state"][0]["float32"][0][0] assert tensors["param"].tolist() == [1.25] assert tensors["exp_avg"].tolist() == [2.5] @@ -619,6 +617,19 @@ def test_fft_backend_hf_load_never_reads_native(monkeypatch) -> None: ), "optimizer_config", ), + ( + lambda metadata: replace( + metadata, + native_optimizer_config={"use_precision_aware_optimizer": True}, + ), + "native_optimizer_config", + ), + ( + lambda metadata: replace( + metadata, native_distributed_config={"grad_reduce_in_fp32": True} + ), + "native_distributed_config", + ), ], ) def test_native_metadata_mismatch_prevents_model_mutation( diff --git a/tests/backends/test_native_megatron_config.py b/tests/backends/test_native_megatron_config.py new file mode 100644 index 0000000..f9339b9 --- /dev/null +++ b/tests/backends/test_native_megatron_config.py @@ -0,0 +1,61 @@ +"""CPU checks for the boundary between YAML settings and Megatron constructors.""" + +from types import SimpleNamespace +from unittest.mock import Mock + +import pytest +from runtime_stubs import backend_runtime_imports + +from lilo.backends.deployment import backend_config +from lilo.backends.megatron_config import parse_backend_config +from lilo.deployments import DeploymentSpec, load, preset_path + +with backend_runtime_imports(): + from lilo.backends.megatron_runtime.common import modeling + + +def test_yaml_native_values_reach_megatron(monkeypatch): + data = load(preset_path("qwen35-4b-fft-64k")).model_dump() + data["trainer"]["config"]["optimizer"]["native_optimizer_setting"] = False + data["trainer"]["config"]["distributed"] = {"native_ddp_setting": 123} + data["trainer"]["config"]["provider"]["native_provider_setting"] = [1, 2] + config, _ = parse_backend_config( + backend_config(DeploymentSpec.model_validate(data)) + ) + # These stand in for an installed upstream version with extra fields. The + # deployment reader must not need its own list of those fields. + provider = SimpleNamespace( + native_provider_setting=None, + mtp_num_layers=0, + recompute_granularity=None, + recompute_method=None, + recompute_num_layers=None, + provide_distributed_model=Mock(return_value="model"), + ) + bridge = SimpleNamespace(to_megatron_provider=lambda: provider) + monkeypatch.setattr( + modeling, + "AutoBridge", + SimpleNamespace(from_hf_pretrained=lambda *a, **k: bridge), + ) + monkeypatch.setattr(modeling, "parameter_dtype", lambda c: "bf16") + optimizer_constructor = Mock(side_effect=lambda **kwargs: kwargs) + ddp_constructor = Mock(side_effect=lambda **kwargs: kwargs) + monkeypatch.setattr(modeling, "MCoreOptimizerConfig", optimizer_constructor) + monkeypatch.setattr(modeling, "DistributedDataParallelConfig", ddp_constructor) + _, actual_provider, _ = modeling.model_provider(config) + assert actual_provider.native_provider_setting == [1, 2] + assert actual_provider.tensor_model_parallel_size == 2 + optimizer = modeling.optimizer_config(config, "bf16", distributed_optimizer=True) + assert optimizer["native_optimizer_setting"] is False + assert optimizer["lr"] == 0.0001 + assert ( + modeling.distributed_model(provider, config, distributed_optimizer=True) + == "model" + ) + assert ddp_constructor.call_args.kwargs["native_ddp_setting"] == 123 + assert ddp_constructor.call_args.kwargs["use_distributed_optimizer"] is True + # An unsupported native field is the installed backend's error at startup. + config.provider_overrides["unknown_field"] = True + with pytest.raises(ValueError, match="unknown Megatron provider override"): + modeling.model_provider(config) diff --git a/tests/providers/test_checkpoint_storage.py b/tests/providers/test_checkpoint_storage.py index 072669f..33ea0e9 100644 --- a/tests/providers/test_checkpoint_storage.py +++ b/tests/providers/test_checkpoint_storage.py @@ -35,7 +35,7 @@ def test_checkpoint_storage_creates_one_v2_volume_without_live_lookup() -> None: def test_yaml_definitions_use_configured_checkpoint_storage(): from lilo.providers.modal.yaml_apps import volumes_for - from lilo.providers.modal.recipe import backend_config + from lilo.backends.deployment import backend_config for definition in DEFINITIONS: spec = definition.RESOLVED.spec diff --git a/tests/providers/test_deployment_presets.py b/tests/providers/test_deployment_presets.py index 2a3af35..676917c 100644 --- a/tests/providers/test_deployment_presets.py +++ b/tests/providers/test_deployment_presets.py @@ -1,7 +1,7 @@ import pytest from lilo.deployments import load, preset_path, resolve -from lilo.providers.modal.recipe import backend_config, serving_options +from lilo.backends.deployment import backend_config, serving_options from lilo.providers.modal.yaml_apps import ( definition_from_spec, frontend_settings, @@ -83,6 +83,6 @@ def test_invalid_attention_parallelism_is_rejected(options): from lilo.deployments import DeploymentSpec data = load(preset_path("qwen35-35b-a3b-fft-64k")).model_dump() - data["inference"]["sglang"]["options"].update(options) + data["inference"]["config"].update(options) with pytest.raises(ValueError, match="sglang"): DeploymentSpec.model_validate(data) diff --git a/tests/test_deployments.py b/tests/test_deployments.py index 0245a66..e8462f7 100644 --- a/tests/test_deployments.py +++ b/tests/test_deployments.py @@ -15,7 +15,7 @@ ) from lilo.deployment_cli import retain_generations from lilo.control_plane.deployments import DeploymentRoutes -from lilo.providers.modal.recipe import backend_config +from lilo.backends.deployment import backend_config from lilo.providers.modal.yaml_apps import definition_from_spec from lilo.native_options import apply_defaults @@ -62,8 +62,8 @@ def test_presets_context_topology_and_backend_options(): def test_no_model_catalog_required(): spec = recipe( model__id="my-org/new-model", - trainer__miles__model_args=None, - trainer__miles__options={ + trainer__config__model_args=None, + trainer__config__options={ "num_layers": 12, "hidden_size": 768, "num_attention_heads": 12, @@ -80,14 +80,14 @@ def test_extends_false_and_lists_replace(tmp_path): { "extends": "builtin:qwen35-9b-lora-16k", "routing": {"default": False}, - "trainer": {"miles": {"options": {"target_modules": ["linear_qkv"]}}}, + "trainer": {"config": {"options": {"target_modules": ["linear_qkv"]}}}, } ) ) spec = load(path) assert spec.routing.default is False - assert spec.trainer.miles.options["target_modules"] == ["linear_qkv"] - assert spec.trainer.miles.options["lora_rank"] == 32 + assert spec.trainer.config["options"]["target_modules"] == ["linear_qkv"] + assert spec.trainer.config["options"]["lora_rank"] == 32 def test_duplicate_keys_and_cycles(tmp_path): @@ -103,15 +103,15 @@ def test_duplicate_keys_and_cycles(tmp_path): @pytest.mark.parametrize( "changes,match", [ - ({"trainer__miles__options": {"hf_checkpoint": "other"}}, "managed"), - ({"trainer__miles__options": {"pipeline_model_parallel_size": 2}}, "managed"), - ({"inference__sglang__options": {"model_path": "other"}}, "managed"), - ({"inference__sglang__options": {"tp_size": 2}}, "replica GPU"), + ({"trainer__config__options": {"hf_checkpoint": "other"}}, "managed"), + ({"trainer__config__options": {"pipeline_model_parallel_size": 2}}, "managed"), + ({"inference__config": {"model_path": "other"}}, "managed"), + ({"inference__config": {"tp_size": 2}}, "replica GPU"), ({"trainer__engine__max_clients_per_instance": 7}, "multi_lora_n_adapters"), ({"trainer__env": {"LILO_BACKEND_CONFIG": "oops"}}, "managed"), ( { - "inference__sglang__options": { + "inference__config": { "max_loaded_loras": 2, "max_loras_per_batch": 8, } @@ -290,3 +290,66 @@ def readable_int(value): args = parser.parse_args([]) assert args.context_length == 65536 assert args.sizes == [32, 2000] + + +def test_native_sections_survive_serialization_without_allowlist(): + import json + from lilo.backends.megatron_config import parse_backend_config + + spec = recipe("qwen35-4b-fft-64k") + data = spec.model_dump() + data["trainer"]["config"]["provider"]["future_provider_option"] = { + "layers": [1, 4], + "enabled": False, + } + data["trainer"]["config"]["optimizer"]["future_optimizer_option"] = 0.125 + data["trainer"]["config"]["distributed"] = {"future_ddp_option": False} + spec = DeploymentSpec.model_validate(data) + settings = backend_config(spec, "/assets/pinned") + config, _ = parse_backend_config(json.loads(json.dumps(settings))) + assert config.hf_checkpoint == "/assets/pinned" + assert config.seq_length == spec.model.max_context_length + assert config.provider_overrides["future_provider_option"] == { + "layers": [1, 4], + "enabled": False, + } + assert config.optimizer_overrides == {"future_optimizer_option": 0.125} + assert config.distributed_overrides == {"future_ddp_option": False} + assert config.optimizer.lr == 0.0001 + assert spec.model_dump() == data # Building does not consume or mutate YAML. + + +@pytest.mark.parametrize( + "section,options,match", + [ + ("provider", {"context_parallel_size": 4}, "managed"), + ("optimizer", {"bf16": False}, "managed"), + ("distributed", {"use_distributed_optimizer": False}, "managed"), + ("runtime", {"optimizer_overrides": {}}, "managed"), + ("runtime", {"misspelled_loop_option": 1}, "runtime options"), + ("provider", [], "mapping"), + ("optimizer", {"optimizer": "sgd"}, "Adam"), + ], +) +def test_megatron_native_options_preserve_integration_contract(section, options, match): + with pytest.raises(ValueError, match=match): + recipe("qwen35-4b-fft-64k", **{f"trainer__config__{section}": options}) + + +def test_backend_dispatch_rejects_unknown_backend(): + with pytest.raises(ValueError, match="unknown deployment backend"): + recipe(trainer__backend="missing") + + +def test_new_miles_and_sglang_options_need_no_deployment_schema_change(): + from lilo.backends.deployment import serving_options + + spec = recipe( + trainer__config__options__future_miles_option=[1, 2], + inference__config__future_sglang_option=False, + ) + assert backend_config(spec)["miles"]["native_options"]["future_miles_option"] == [ + 1, + 2, + ] + assert serving_options(spec)["future_sglang_option"] is False From 45fd06ec887dcf1b411481df474c032e76f19487 Mon Sep 17 00:00:00 2001 From: kailash Date: Wed, 23 Sep 2026 17:40:34 +0000 Subject: [PATCH 10/27] Keep backend checks out of YAML model construction --- docs/deployment-yaml-design.md | 4 +- docs/deployment-yaml-validation.md | 6 ++ src/lilo/backends/megatron_deployment.py | 2 + src/lilo/deployment_cli.py | 2 +- src/lilo/deployments.py | 63 ++++++++---------- src/lilo/providers/modal/yaml_apps.py | 13 +++- tests/providers/test_deployment_presets.py | 5 +- tests/test_deployments.py | 77 ++++++++++++++++++++-- 8 files changed, 124 insertions(+), 48 deletions(-) diff --git a/docs/deployment-yaml-design.md b/docs/deployment-yaml-design.md index 8c1c9d7..45326a7 100644 --- a/docs/deployment-yaml-design.md +++ b/docs/deployment-yaml-design.md @@ -221,12 +221,14 @@ inference: Miles and SGLang use their real argument parsers before initialization. Boolean flags, scalar values and ordinary list arguments are supported. Both spellings of opposing boolean flags are replaced when they share a destination. Unknown options and custom/repeated argparse actions fail explicitly. These are parser-backed options, so arbitrary Python objects and custom actions are not supported. -Lilo still checks settings that affect its integration locally: GPU counts and parallelism, maximum clients versus adapter slots, managed model paths, context/rank configuration, communication endpoints and trainer-only mode. The backend readers reject conflicting values. For Megatron, set parallelism and shared distributed-optimizer controls under `runtime` so both Lilo and Megatron receive the same values. Tinker training requires an Adam optimizer. Native passthrough does not make other training protocols or unsupported process layouts work automatically. +When preparing a trainer or inference pool, its backend reader checks settings that affect the integration: GPU counts and parallelism, maximum clients versus adapter slots, managed model paths, context/rank configuration, communication endpoints and trainer-only mode. The backend readers reject conflicting values. For Megatron, set parallelism and shared distributed-optimizer controls under `runtime` so both Lilo and Megatron receive the same values. Tinker training requires an Adam optimizer. Native passthrough does not make other training protocols or unsupported process layouts work automatically. All built-in YAMLs use this structure. The earlier draft's `trainer.miles`, `trainer.megatron` and `inference.sglang` sections have been removed; user YAMLs overriding those sections must move them under `config` as shown above. This schema change is part of the draft and has not been deployed. Existing source-fingerprint checks continue to prevent applying a different runtime implementation over running trainers. Inference adapter targets are derived from the existing Miles-to-PEFT mapping unless `lora_target_modules` is explicitly supplied. This mapping does not establish support for every architecture. A model still needs compatible training, adapter export and SGLang loading implementations in the selected images. +YAML loading validates document fields and inheritance only; it does not call backend readers. Backend errors surface when their trainer or pool settings are built, without a Pydantic wrapper. The loader walks the inheritance chain and merges it from parent to child before constructing `DeploymentSpec`; partial parent files are supported. + Local validation cannot establish memory fit or prove that an unfamiliar model works. Native parser validation occurs inside the runtime images on startup. This draft does not implement a separate image-preflight command or a GPU export/load/generation probe. ## Saved configurations, failures and updates diff --git a/docs/deployment-yaml-validation.md b/docs/deployment-yaml-validation.md index 340b806..27552dc 100644 --- a/docs/deployment-yaml-validation.md +++ b/docs/deployment-yaml-validation.md @@ -2,6 +2,12 @@ These checks exercise the shared-app deployment path in PR #55. Each deployment contains the frontend and generated trainer functions; inference pools are separate apps started on demand. +## YAML loader simplification + +`DeploymentSpec.compatible()` and `_load()` were removed. `load()` reads the inheritance chain, merges parent to child, and constructs the deployment fields once. Loading, deserializing and resolving a specification do not call backend readers. Backend integration checks run when building backend settings; the Modal provider checks reserved environment variables when configuring apps. + +Validation: **592 CPU tests passed, 1 skipped**; Ruff passed for changed Python files and whitespace checks passed. Coverage includes partial parents, inheritance cycles, intermediate replacements, and ensuring loading/resolution never calls backend parsers. No apps were redeployed. + ## Backend configuration passthrough The deployment schema now uses `trainer.backend` / `trainer.config` and `inference.backend` / `inference.config`. Backend readers interpret those mappings outside the Modal provider. Megatron exposes native provider, optimizer and distributed-training settings alongside its Lilo runtime settings. diff --git a/src/lilo/backends/megatron_deployment.py b/src/lilo/backends/megatron_deployment.py index c6dfce0..ac420f1 100644 --- a/src/lilo/backends/megatron_deployment.py +++ b/src/lilo/backends/megatron_deployment.py @@ -38,6 +38,8 @@ def build_config(spec, asset_path): trainer = spec.trainer if spec.model.parameterization != "full": raise ValueError("Megatron YAML deployments require full parameterization") + if trainer.engine.max_clients_per_instance != 1: + raise ValueError("FFT trainers admit one client per instance") if trainer.engine.sampler_persistence_concurrency != 1: raise ValueError("Megatron requires sampler_persistence_concurrency: 1") sections = native_options(trainer.config, set()) diff --git a/src/lilo/deployment_cli.py b/src/lilo/deployment_cli.py index c3a1cf2..a68ca2b 100644 --- a/src/lilo/deployment_cli.py +++ b/src/lilo/deployment_cli.py @@ -197,7 +197,7 @@ def main(argv=None): specs = [load(path) for path in args.files] validate_frontend(specs) print( - f"Validated {len(specs)} deployment(s). Native backend options are checked at startup in their runtime images." + f"Validated {len(specs)} deployment(s). Backend integration settings are checked when preparing trainers and pools; native options are checked at worker startup." ) else: output = ( diff --git a/src/lilo/deployments.py b/src/lilo/deployments.py index ae9d96c..3452f46 100644 --- a/src/lilo/deployments.py +++ b/src/lilo/deployments.py @@ -125,26 +125,6 @@ class DeploymentSpec(StrictModel): inference: Inference lifecycle: Lifecycle = Field(default_factory=Lifecycle) - @model_validator(mode="after") - def compatible(self): - from lilo.backends.deployment import backend_config, serving_options - - for role in (self.trainer, self.inference): - role.resources.gpu_count - if any(k.startswith("LILO_") for k in role.env): - raise ValueError("LILO_ environment variables are managed by Lilo") - if ( - self.model.parameterization == "full" - and self.trainer.engine.max_clients_per_instance != 1 - ): - raise ValueError("FFT trainers admit one client per instance") - try: - backend_config(self) - serving_options(self) - except TypeError as exc: - raise ValueError(f"invalid backend option types: {exc}") from exc - return self - class ResolvedDeployment(StrictModel): spec: DeploymentSpec @@ -223,28 +203,37 @@ def preset_path(name: str) -> Path: return Path(str(files("lilo").joinpath("presets", name + ".yaml"))) -def load(path: str | Path, _seen: tuple[Path, ...] = ()) -> DeploymentSpec: - return DeploymentSpec.model_validate(_load(Path(path), _seen)) - - -def _load(path: Path, seen: tuple[Path, ...]) -> dict: - path = path.resolve() - if path in seen: - raise ValueError(f"cyclic extends: {path}") - data = yaml.load(path.read_text(), Loader=UniqueLoader) - if not isinstance(data, dict): - raise ValueError("deployment YAML must be a mapping") - parent = data.pop("extends", None) - if parent is not None: +def load(path: str | Path) -> DeploymentSpec: + """Read and merge YAML inheritance, then construct the deployment fields. + + Backend configuration is interpreted when preparing its trainer or pool. + Parent files may be partial; only the fully merged document is constructed. + """ + path = Path(path).resolve() + seen = set() + documents = [] + while True: + if path in seen: + raise ValueError(f"cyclic extends: {path}") + seen.add(path) + data = yaml.load(path.read_text(), Loader=UniqueLoader) + if not isinstance(data, dict): + raise ValueError("deployment YAML must be a mapping") + parent = data.pop("extends", None) + documents.append(data) + if parent is None: + break if not isinstance(parent, str): raise ValueError("extends must be a path or builtin:preset") - parent_path = ( + path = ( preset_path(parent[8:]) if parent.startswith("builtin:") else path.parent / parent - ) - data = merge(_load(parent_path, (*seen, path)), data) - return data + ).resolve() + merged = {} + for document in reversed(documents): + merged = merge(merged, document) + return DeploymentSpec(**merged) def validate_frontend(specs: list[DeploymentSpec]) -> None: diff --git a/src/lilo/providers/modal/yaml_apps.py b/src/lilo/providers/modal/yaml_apps.py index 45385f1..e411cd3 100644 --- a/src/lilo/providers/modal/yaml_apps.py +++ b/src/lilo/providers/modal/yaml_apps.py @@ -85,6 +85,13 @@ def secrets_for(spec, *, training=False): return result +def deployment_env(values): + """Keep user environment overrides separate from Lilo's deployment wiring.""" + if any(key.startswith("LILO_") for key in values): + raise ValueError("LILO_ environment variables are managed by Lilo") + return dict(values) + + def build_trainer_app(resolved: ResolvedDeployment, *, image=None): # Admission metadata must not change the serialized function for a running # configuration when a default switches or an older configuration drains. @@ -101,7 +108,7 @@ def build_trainer_app(resolved: ResolvedDeployment, *, image=None): env = { **trainer_deployment_env(), - **spec.trainer.env, + **deployment_env(spec.trainer.env), "LILO_APP_NAME": spec.deployment.frontend, } @@ -138,7 +145,7 @@ def run_trainer(resolved, instance_id): # on startup to see the committed exact snapshot; never race a trainer download. volumes_for(spec)["/assets"].reload() env = { - **spec.trainer.env, + **deployment_env(spec.trainer.env), "LILO_APP_NAME": spec.deployment.frontend, "LILO_BACKEND_CONFIG": json.dumps(settings), "LILO_BASE_MODEL": spec.model.id, @@ -262,7 +269,7 @@ def build_rollout_app(resolved, pool, *, image=None): memory=resources.memory_mib, volumes=volumes_for(spec), secrets=secrets_for(spec), - env=spec.inference.env, + env=deployment_env(spec.inference.env), min_containers=scaling.min_replicas if minimum is None else minimum, max_containers=scaling.max_replicas if maximum is None else maximum, target_concurrency=scaling.target_concurrency, diff --git a/tests/providers/test_deployment_presets.py b/tests/providers/test_deployment_presets.py index 676917c..2b45fa8 100644 --- a/tests/providers/test_deployment_presets.py +++ b/tests/providers/test_deployment_presets.py @@ -84,5 +84,8 @@ def test_invalid_attention_parallelism_is_rejected(options): data = load(preset_path("qwen35-35b-a3b-fft-64k")).model_dump() data["inference"]["config"].update(options) + from lilo.backends.deployment import serving_options + + spec = DeploymentSpec.model_validate(data) with pytest.raises(ValueError, match="sglang"): - DeploymentSpec.model_validate(data) + serving_options(spec) diff --git a/tests/test_deployments.py b/tests/test_deployments.py index e8462f7..5909f01 100644 --- a/tests/test_deployments.py +++ b/tests/test_deployments.py @@ -108,7 +108,6 @@ def test_duplicate_keys_and_cycles(tmp_path): ({"inference__config": {"model_path": "other"}}, "managed"), ({"inference__config": {"tp_size": 2}}, "replica GPU"), ({"trainer__engine__max_clients_per_instance": 7}, "multi_lora_n_adapters"), - ({"trainer__env": {"LILO_BACKEND_CONFIG": "oops"}}, "managed"), ( { "inference__config": { @@ -120,9 +119,13 @@ def test_duplicate_keys_and_cycles(tmp_path): ), ], ) -def test_invalid_integrations_fail_locally(changes, match): +def test_invalid_integrations_fail_when_building_backend_settings(changes, match): + from lilo.backends.deployment import serving_options + + spec = recipe(**changes) with pytest.raises(ValueError, match=match): - recipe(**changes) + backend_config(spec) + serving_options(spec) def test_generation_and_asset_paths_include_exact_base(): @@ -332,13 +335,15 @@ def test_native_sections_survive_serialization_without_allowlist(): ], ) def test_megatron_native_options_preserve_integration_contract(section, options, match): + spec = recipe("qwen35-4b-fft-64k", **{f"trainer__config__{section}": options}) with pytest.raises(ValueError, match=match): - recipe("qwen35-4b-fft-64k", **{f"trainer__config__{section}": options}) + backend_config(spec) def test_backend_dispatch_rejects_unknown_backend(): + spec = recipe(trainer__backend="missing") with pytest.raises(ValueError, match="unknown deployment backend"): - recipe(trainer__backend="missing") + backend_config(spec) def test_new_miles_and_sglang_options_need_no_deployment_schema_change(): @@ -353,3 +358,65 @@ def test_new_miles_and_sglang_options_need_no_deployment_schema_change(): 2, ] assert serving_options(spec)["future_sglang_option"] is False + + +def test_load_merges_partial_parents_before_constructing_spec(tmp_path): + parent = tmp_path / "parent.yaml" + parent.write_text("trainer:\n config:\n options:\n custom_option: 1\n") + middle = tmp_path / "middle.yaml" + middle.write_text("extends: parent.yaml\ntrainer:\n config:\n options:\n custom_option: 2\n") + data = recipe().model_dump() + data["extends"] = "middle.yaml" + data["trainer"]["config"]["options"]["other_option"] = False + child = tmp_path / "child.yaml" + child.write_text(yaml.safe_dump(data)) + spec = load(child) + assert spec.trainer.config["options"]["custom_option"] == 2 + assert spec.trainer.config["options"]["other_option"] is False + + +def test_loading_and_resolving_do_not_interpret_backend_config(tmp_path, monkeypatch): + import lilo.backends.deployment as backends + + def unexpected(*args, **kwargs): + pytest.fail("YAML construction must not interpret backend configuration") + + monkeypatch.setattr(backends, "backend_config", unexpected) + monkeypatch.setattr(backends, "serving_options", unexpected) + data = recipe().model_dump() + data["trainer"]["backend"] = "unknown-until-startup" + path = tmp_path / "deployment.yaml" + path.write_text(yaml.safe_dump(data)) + spec = load(path) + assert resolve(spec, revision="a" * 40, implementation="test").spec == spec.model_copy( + update={"model": spec.model.model_copy(update={"revision": "a" * 40})} + ) + + +def test_fft_capacity_is_checked_by_backend_setup(): + spec = recipe("qwen35-4b-fft-64k", trainer__engine__max_clients_per_instance=2) + with pytest.raises(ValueError, match="FFT trainers admit one client"): + backend_config(spec) + + +def test_reserved_environment_is_checked_by_modal_setup(): + from lilo.providers.modal.yaml_apps import deployment_env + + spec = recipe(trainer__env={"LILO_BACKEND_CONFIG": "oops"}) + with pytest.raises(ValueError, match="managed"): + deployment_env(spec.trainer.env) + assert deployment_env({"MY_SETTING": "value"}) == {"MY_SETTING": "value"} + + +def test_inheritance_keeps_intermediate_replacements(tmp_path): + parent = recipe().model_dump() + parent["trainer"]["config"]["options"]["custom"] = {"old": 1} + (tmp_path / "parent.yaml").write_text(yaml.safe_dump(parent)) + (tmp_path / "middle.yaml").write_text( + "extends: parent.yaml\ntrainer:\n config:\n options:\n custom: null\n" + ) + child = tmp_path / "child.yaml" + child.write_text( + "extends: middle.yaml\ntrainer:\n config:\n options:\n custom:\n new: 2\n" + ) + assert load(child).trainer.config["options"]["custom"] == {"new": 2} From a6b343aae8e04caa329c271f4192e2472277be7e Mon Sep 17 00:00:00 2001 From: kailash Date: Wed, 23 Sep 2026 17:45:04 +0000 Subject: [PATCH 11/27] Simplify saved deployment records without reparsing YAML --- docs/deployment-yaml-design.md | 12 +++++ docs/deployment-yaml-validation.md | 6 +++ scripts/e2e_engine_definition.py | 4 +- src/lilo/deployment_cli.py | 19 ++++--- src/lilo/deployments.py | 59 ++++++++++++---------- src/lilo/providers/modal/yaml_apps.py | 8 +-- src/lilo/providers/modal/yaml_pool_app.py | 4 +- tests/providers/conftest.py | 4 +- tests/providers/test_deployment_presets.py | 4 +- tests/providers/test_modal_app.py | 4 +- tests/providers/test_yaml_apps.py | 4 +- tests/providers/test_yaml_e2e_helper.py | 4 +- tests/test_deployment_cli.py | 33 ++++++++++-- tests/test_deployments.py | 46 ++++++++++++++--- 14 files changed, 148 insertions(+), 63 deletions(-) diff --git a/docs/deployment-yaml-design.md b/docs/deployment-yaml-design.md index 45326a7..1c019d7 100644 --- a/docs/deployment-yaml-design.md +++ b/docs/deployment-yaml-design.md @@ -231,6 +231,18 @@ YAML loading validates document fields and inheritance only; it does not call ba Local validation cannot establish memory fit or prove that an unfamiliar model works. Native parser validation occurs inside the runtime images on startup. This draft does not implement a separate image-preflight command or a GPU export/load/generation probe. +## From YAML to a saved deployment + +`load()` returns a `DeploymentSpec`: the merged YAML fields. It does not interpret backend options. + +`deployment_cli.compile_configs()` resolves a model branch or tag to a Hugging Face commit. A revision that is already a commit needs no lookup. It also identifies the Lilo code and Miles revision being deployed. + +`DeploymentRecord.create()` copies the specification, sets that pinned revision, and hashes the code identity and configuration. It does not serialize and reparse the specification or call backend readers. The standalone `resolve()` function is gone. + +The separate record holds information used by the deployment manifest: the specification, code identity, Miles revision, configuration hash, and whether new clients can select it. These fields are not user YAML options. The hash keeps jobs attached to the configuration that created them; the active flag lets an older configuration keep serving existing jobs after it is removed from the deployment list. `definition_id` and the model asset path are derived properties. Stored JSON is decoded when crossing a registry or worker-process boundary. + +The class is named `DeploymentRecord` to describe that role. Its serialized fields and hash format are unchanged from the earlier `ResolvedDeployment` name. + ## Saved configurations, failures and updates The CLI stores configurations in the Modal Dict `-yaml-deployments`, scoped to the chosen Modal environment. A single apply lock serializes registry changes. A pending manifest is written before deployment; only successful deployment replaces the committed manifest. The next attempt retains pending configurations too, covering an interruption after Modal accepted a deployment but before the CLI saved its result. diff --git a/docs/deployment-yaml-validation.md b/docs/deployment-yaml-validation.md index 27552dc..c9e2afb 100644 --- a/docs/deployment-yaml-validation.md +++ b/docs/deployment-yaml-validation.md @@ -2,6 +2,12 @@ These checks exercise the shared-app deployment path in PR #55. Each deployment contains the frontend and generated trainer functions; inference pools are separate apps started on demand. +## Deployment record simplification + +The saved manifest wrapper is now named `DeploymentRecord`. Its factory copies the parsed specification and computes the configuration hash without a dictionary-to-model round trip. Model tag resolution and validation of the Hugging Face result stay in the CLI. The standalone `resolve()` function was removed. Serialized record fields and hash format are unchanged. + +Validation: **595 CPU tests passed, 1 skipped**; Ruff passed for changed Python files and whitespace checks passed. Added coverage verifies model tag pinning, no lookup for an existing commit, record isolation from subsequent mutations, JSON round trips, and unchanged configuration hashes. No apps were redeployed. + ## YAML loader simplification `DeploymentSpec.compatible()` and `_load()` were removed. `load()` reads the inheritance chain, merges parent to child, and constructs the deployment fields once. Loading, deserializing and resolving a specification do not call backend readers. Backend integration checks run when building backend settings; the Modal provider checks reserved environment variables when configuring apps. diff --git a/scripts/e2e_engine_definition.py b/scripts/e2e_engine_definition.py index b013ff7..b4f7031 100644 --- a/scripts/e2e_engine_definition.py +++ b/scripts/e2e_engine_definition.py @@ -16,7 +16,7 @@ import tinker from tinker import types -from lilo.deployments import ResolvedDeployment +from lilo.deployments import DeploymentRecord from lilo.backends.deployment import backend_config TIMEOUT = 3 * 60 * 60 @@ -25,7 +25,7 @@ def _definition(frontend: str, name: str) -> tuple[Any, str]: rows = modal.Dict.from_name(f"{frontend}-yaml-deployments").get("manifest", []) matches = [ - ResolvedDeployment.model_validate(row) + DeploymentRecord.model_validate(row) for row in rows if row["active"] and row["spec"]["name"] == name ] diff --git a/src/lilo/deployment_cli.py b/src/lilo/deployment_cli.py index a68ca2b..7b81d31 100644 --- a/src/lilo/deployment_cli.py +++ b/src/lilo/deployment_cli.py @@ -15,10 +15,9 @@ import yaml from lilo.deployments import ( - ResolvedDeployment, + DeploymentRecord, load, preset_path, - resolve, validate_frontend, ) @@ -43,22 +42,26 @@ def compile_configs(paths): miles_commit = resolve_miles_commit() implementation = implementation_fingerprint(miles_commit) - resolved = [] + records = [] for spec in specs: revision = spec.model.revision - if not re.fullmatch(r"[0-9a-f]{40,64}", revision): + if not re.fullmatch(r"[0-9a-fA-F]{40,64}", revision): from huggingface_hub import HfApi revision = HfApi().model_info(spec.model.id, revision=revision).sha - resolved.append( - resolve( + if not revision or not re.fullmatch(r"[0-9a-fA-F]{40,64}", revision): + raise ValueError( + f"Hugging Face did not return a commit for {spec.model.id}" + ) + records.append( + DeploymentRecord.create( spec, revision=revision, implementation=implementation, miles_commit=miles_commit, ) ) - return resolved + return records def retain_generations(previous, desired): @@ -124,7 +127,7 @@ def deploy(desired): # A killed deploy may already have updated Modal. Keep its functions on retry. rows = {r["generation"]: r for r in [*rows, *registry.get("pending", [])]} manifest = retain_generations( - [ResolvedDeployment.model_validate(row) for row in rows.values()], desired + [DeploymentRecord.model_validate(row) for row in rows.values()], desired ) data = [row.model_dump(mode="json") for row in manifest] env = { diff --git a/src/lilo/deployments.py b/src/lilo/deployments.py index 3452f46..08b11ac 100644 --- a/src/lilo/deployments.py +++ b/src/lilo/deployments.py @@ -126,13 +126,44 @@ class DeploymentSpec(StrictModel): lifecycle: Lifecycle = Field(default_factory=Lifecycle) -class ResolvedDeployment(StrictModel): +class DeploymentRecord(StrictModel): + """Saved deployment metadata around a YAML specification. + + The CLI resolves the model revision before creating this record. Its hash + binds jobs to their original code and configuration across later deploys; + active controls whether new clients can select it. + """ + spec: DeploymentSpec implementation: str miles_commit: str | None = None generation: str active: bool = True + @classmethod + def create( + cls, + spec: DeploymentSpec, + *, + revision: str, + implementation: str, + miles_commit: str | None = None, + ) -> DeploymentRecord: + """Record an already-resolved revision without reparsing the YAML.""" + pinned = spec.model_copy(deep=True) + pinned.model.revision = revision + # Changing routing defaults should not restart an existing trainer. + identity = pinned.model_dump(exclude={"routing"}) + generation = hashlib.sha256( + json.dumps([implementation, identity], sort_keys=True).encode() + ).hexdigest() + return cls( + spec=pinned, + implementation=implementation, + generation=generation, + miles_commit=miles_commit, + ) + @property def definition_id(self) -> str: return f"yaml_{self.spec.name}_{self.generation[:16]}" @@ -145,32 +176,6 @@ def asset_path(self) -> str: return f"/assets/{digest}" -def resolve( - spec: DeploymentSpec, - *, - revision: str, - implementation: str, - miles_commit: str | None = None, -) -> ResolvedDeployment: - if not re.fullmatch(r"[a-fA-F0-9]{40,64}", revision): - raise ValueError("model revision must resolve to an exact commit") - value = spec.model_dump() - value["model"]["revision"] = revision - pinned = DeploymentSpec.model_validate(value) - # Include resource and lifecycle policy: each applied version remains self-contained. - # Checkpoint compatibility uses model/topology metadata, not this generation hash. - identity = pinned.model_dump(exclude={"routing"}) - digest = hashlib.sha256( - json.dumps([implementation, identity], sort_keys=True).encode() - ).hexdigest() - return ResolvedDeployment( - spec=pinned, - implementation=implementation, - generation=digest, - miles_commit=miles_commit, - ) - - class UniqueLoader(yaml.SafeLoader): pass diff --git a/src/lilo/providers/modal/yaml_apps.py b/src/lilo/providers/modal/yaml_apps.py index e411cd3..478a1d9 100644 --- a/src/lilo/providers/modal/yaml_apps.py +++ b/src/lilo/providers/modal/yaml_apps.py @@ -12,7 +12,7 @@ import modal -from lilo.deployments import ResolvedDeployment, Routing, validate_frontend +from lilo.deployments import DeploymentRecord, Routing, validate_frontend from lilo.backends.deployment import backend_config, serving_options MANIFEST_ENV = "LILO_DEPLOYMENT_MANIFEST" @@ -28,7 +28,7 @@ def manifest_from_env(): rows = json.loads(data) if not isinstance(rows, list) or not rows: raise ValueError("Deployment manifest must be a nonempty list") - return [ResolvedDeployment.model_validate(row) for row in rows] + return [DeploymentRecord.model_validate(row) for row in rows] def frontend_settings(): @@ -92,7 +92,7 @@ def deployment_env(values): return dict(values) -def build_trainer_app(resolved: ResolvedDeployment, *, image=None): +def build_trainer_app(resolved: DeploymentRecord, *, image=None): # Admission metadata must not change the serialized function for a running # configuration when a default switches or an older configuration drains. resolved = resolved.model_copy(deep=True) @@ -129,7 +129,7 @@ def build_trainer_app(resolved: ResolvedDeployment, *, image=None): env=env, ) def trainer(instance_id: str): - run_trainer(ResolvedDeployment.model_validate_json(config_json), instance_id) + run_trainer(DeploymentRecord.model_validate_json(config_json), instance_id) return app, trainer diff --git a/src/lilo/providers/modal/yaml_pool_app.py b/src/lilo/providers/modal/yaml_pool_app.py index 65bfea5..7da18d7 100644 --- a/src/lilo/providers/modal/yaml_pool_app.py +++ b/src/lilo/providers/modal/yaml_pool_app.py @@ -2,12 +2,12 @@ import os -from lilo.deployments import ResolvedDeployment +from lilo.deployments import DeploymentRecord from .fft_pool import FFTPoolSpec from .lora_pool import LoraPoolSpec from .yaml_apps import POOL_CONFIG_ENV, build_rollout_app -resolved = ResolvedDeployment.model_validate_json(os.environ[POOL_CONFIG_ENV]) +resolved = DeploymentRecord.model_validate_json(os.environ[POOL_CONFIG_ENV]) if resolved.spec.model.parameterization == "lora": pool = LoraPoolSpec(resolved.definition_id, revision=resolved.generation[:16]) else: diff --git a/tests/providers/conftest.py b/tests/providers/conftest.py index 41b846f..ba0ffd8 100644 --- a/tests/providers/conftest.py +++ b/tests/providers/conftest.py @@ -3,13 +3,13 @@ import json import os -from lilo.deployments import load, preset_path, resolve +from lilo.deployments import load, preset_path, DeploymentRecord os.environ.setdefault( "LILO_DEPLOYMENT_MANIFEST", json.dumps( [ - resolve( + DeploymentRecord.create( load(preset_path(name)), revision="a" * 40, implementation="tests" ).model_dump(mode="json") for name in ("qwen35-9b-fft-64k", "qwen35-9b-lora-16k") diff --git a/tests/providers/test_deployment_presets.py b/tests/providers/test_deployment_presets.py index 2b45fa8..68f3bd0 100644 --- a/tests/providers/test_deployment_presets.py +++ b/tests/providers/test_deployment_presets.py @@ -1,6 +1,6 @@ import pytest -from lilo.deployments import load, preset_path, resolve +from lilo.deployments import load, preset_path, DeploymentRecord from lilo.backends.deployment import backend_config, serving_options from lilo.providers.modal.yaml_apps import ( definition_from_spec, @@ -36,7 +36,7 @@ def test_moe_rollout_preserves_attention_data_parallelism(): assert options["tp_size"] == options["dp_size"] == options["ep_size"] == 4 assert options["enable_dp_attention"] is True definition = definition_from_spec( - resolve(spec, revision="a" * 40, implementation="test"), register_trainer=False + DeploymentRecord.create(spec, revision="a" * 40, implementation="test"), register_trainer=False ) assert definition.ROLLOUT_GPUS == 4 assert definition.ROLLOUT_TENSOR_PARALLEL_SIZE == 1 diff --git a/tests/providers/test_modal_app.py b/tests/providers/test_modal_app.py index 7dd8fb7..22de02c 100644 --- a/tests/providers/test_modal_app.py +++ b/tests/providers/test_modal_app.py @@ -10,11 +10,11 @@ from lilo.providers.modal.fft_pool import FFTPoolSpec from lilo.providers.modal.lora_pool import LoraPoolSpec -from lilo.deployments import load, preset_path, resolve +from lilo.deployments import load, preset_path, DeploymentRecord def definition_id(preset): - return resolve( + return DeploymentRecord.create( load(preset_path(preset)), revision="a" * 40, implementation="tests" ).definition_id diff --git a/tests/providers/test_yaml_apps.py b/tests/providers/test_yaml_apps.py index 40a90fe..546f4ca 100644 --- a/tests/providers/test_yaml_apps.py +++ b/tests/providers/test_yaml_apps.py @@ -5,14 +5,14 @@ import modal import pytest -from lilo.deployments import load, preset_path, resolve +from lilo.deployments import load, preset_path, DeploymentRecord from lilo.providers.modal import yaml_apps from lilo.providers.modal.fft_pool import FFTPoolSpec from lilo.providers.modal.lora_pool import LoraPoolSpec def deployment(preset="qwen35-9b-lora-16k"): - return resolve( + return DeploymentRecord.create( load(preset_path(preset)), revision="a" * 40, implementation="runtime" ) diff --git a/tests/providers/test_yaml_e2e_helper.py b/tests/providers/test_yaml_e2e_helper.py index 126ac14..7c3feb3 100644 --- a/tests/providers/test_yaml_e2e_helper.py +++ b/tests/providers/test_yaml_e2e_helper.py @@ -5,7 +5,7 @@ import modal import pytest -from lilo.deployments import load, preset_path, resolve +from lilo.deployments import load, preset_path, DeploymentRecord @pytest.mark.parametrize("preset", ["qwen35-9b-lora-16k", "qwen35-4b-fft-64k"]) @@ -13,7 +13,7 @@ def test_e2e_helper_reads_active_deployed_configuration(monkeypatch, preset): helper = runpy.run_path( str(Path(__file__).parents[2] / "scripts/e2e_engine_definition.py") ) - row = resolve(load(preset_path(preset)), revision="a" * 40, implementation="test") + row = DeploymentRecord.create(load(preset_path(preset)), revision="a" * 40, implementation="test") retired = row.model_copy(update={"active": False, "generation": "b" * 64}) registry = SimpleNamespace( get=lambda *args: [retired.model_dump(), row.model_dump()] diff --git a/tests/test_deployment_cli.py b/tests/test_deployment_cli.py index 803b6fa..678c140 100644 --- a/tests/test_deployment_cli.py +++ b/tests/test_deployment_cli.py @@ -5,7 +5,7 @@ import pytest from lilo import deployment_cli as cli -from lilo.deployments import load, preset_path, resolve +from lilo.deployments import load, preset_path, DeploymentRecord from lilo.providers.modal.yaml_apps import MANIFEST_ENV @@ -18,7 +18,7 @@ def put(self, key, value, skip_if_exists=False): def deployment(): - return resolve( + return DeploymentRecord.create( load(preset_path("qwen35-9b-lora-16k")), revision="a" * 40, implementation="test", @@ -71,7 +71,7 @@ def fail(*args, **kwargs): assert "manifest" not in registry and "apply_lock" not in registry new_spec = row.spec.model_copy(deep=True) new_spec.trainer.scaling.max_instances = 2 - new = resolve(new_spec, revision="a" * 40, implementation="test") + new = DeploymentRecord.create(new_spec, revision="a" * 40, implementation="test") monkeypatch.setattr(subprocess, "run", lambda *args, **kwargs: None) cli.deploy([new]) assert [(r["generation"], r["active"]) for r in registry["manifest"]] == [ @@ -107,3 +107,30 @@ def test_deploy_rejects_python_mismatch_before_remote_changes(monkeypatch): ) with pytest.raises(ValueError, match="requires Python 3.12"): cli.deploy([deployment()]) + + +@pytest.mark.parametrize("revision,lookups", [("main", 1), ("a" * 40, 0)]) +def test_compile_pins_revision_at_external_boundary( + tmp_path, monkeypatch, revision, lookups +): + from types import SimpleNamespace + from unittest.mock import Mock + import huggingface_hub + import yaml + from lilo.providers.modal import miles_revision + + data = load(preset_path("qwen35-9b-lora-16k")).model_dump() + data["model"]["revision"] = revision + path = tmp_path / "model.yaml" + path.write_text(yaml.safe_dump(data)) + lookup = Mock(return_value=SimpleNamespace(sha="a" * 40)) + monkeypatch.setattr(huggingface_hub.HfApi, "model_info", lookup) + monkeypatch.setattr(miles_revision, "resolve_miles_commit", lambda: "b" * 40) + monkeypatch.setattr(cli, "implementation_fingerprint", lambda _: "runtime") + (row,) = cli.compile_configs([path]) + assert row.spec.model.revision == "a" * 40 + assert lookup.call_count == lookups + if lookups: + lookup.return_value.sha = None + with pytest.raises(ValueError, match="did not return a commit"): + cli.compile_configs([path]) diff --git a/tests/test_deployments.py b/tests/test_deployments.py index 5909f01..beb273d 100644 --- a/tests/test_deployments.py +++ b/tests/test_deployments.py @@ -7,10 +7,10 @@ import yaml from lilo.deployments import ( + DeploymentRecord, DeploymentSpec, load, preset_path, - resolve, validate_frontend, ) from lilo.deployment_cli import retain_generations @@ -32,7 +32,7 @@ def recipe(preset="qwen35-9b-lora-16k", **changes): def resolved(spec=None, **changes): - return resolve( + return DeploymentRecord.create( spec or recipe(**changes), revision="a" * 40, implementation="test-runtime" ) @@ -136,10 +136,10 @@ def test_generation_and_asset_paths_include_exact_base(): assert a.asset_path != b.asset_path assert ( a.asset_path - != resolve(a.spec, revision="b" * 40, implementation="test-runtime").asset_path + != DeploymentRecord.create( + a.spec, revision="b" * 40, implementation="test-runtime" + ).asset_path ) - with pytest.raises(ValueError, match="exact commit"): - resolve(a.spec, revision="main", implementation="test-runtime") def test_frontend_defaults_and_retained_generations(): @@ -364,7 +364,9 @@ def test_load_merges_partial_parents_before_constructing_spec(tmp_path): parent = tmp_path / "parent.yaml" parent.write_text("trainer:\n config:\n options:\n custom_option: 1\n") middle = tmp_path / "middle.yaml" - middle.write_text("extends: parent.yaml\ntrainer:\n config:\n options:\n custom_option: 2\n") + middle.write_text( + "extends: parent.yaml\ntrainer:\n config:\n options:\n custom_option: 2\n" + ) data = recipe().model_dump() data["extends"] = "middle.yaml" data["trainer"]["config"]["options"]["other_option"] = False @@ -388,7 +390,9 @@ def unexpected(*args, **kwargs): path = tmp_path / "deployment.yaml" path.write_text(yaml.safe_dump(data)) spec = load(path) - assert resolve(spec, revision="a" * 40, implementation="test").spec == spec.model_copy( + assert DeploymentRecord.create( + spec, revision="a" * 40, implementation="test" + ).spec == spec.model_copy( update={"model": spec.model.model_copy(update={"revision": "a" * 40})} ) @@ -420,3 +424,31 @@ def test_inheritance_keeps_intermediate_replacements(tmp_path): "extends: middle.yaml\ntrainer:\n config:\n options:\n custom:\n new: 2\n" ) assert load(child).trainer.config["options"]["custom"] == {"new": 2} + + +def test_record_creation_copies_without_reparsing(monkeypatch): + import hashlib + import json + + spec = recipe() + original = spec.model_dump() + # Creating a record must not reconstruct an already-parsed specification. + monkeypatch.setattr( + DeploymentSpec, "model_validate", lambda *a, **k: pytest.fail("reparse") + ) + row = DeploymentRecord.create(spec, revision="a" * 40, implementation="test") + assert spec.model_dump() == original + assert row.spec.model.revision == "a" * 40 + # Keep the existing manifest fields and hash format stable. + expected = original | {"model": original["model"] | {"revision": "a" * 40}} + expected.pop("routing") + assert ( + row.generation + == hashlib.sha256( + json.dumps(["test", expected], sort_keys=True).encode() + ).hexdigest() + ) + row.spec.trainer.config["options"]["lora_rank"] = 64 + assert spec.trainer.config["options"]["lora_rank"] == 32 + saved = row.model_dump_json() + assert DeploymentRecord.model_validate_json(saved) == row From 9c8d3c51f6a211ac5f2a4befc49e9c4bd29beea2 Mon Sep 17 00:00:00 2001 From: kailash Date: Wed, 23 Sep 2026 18:24:18 +0000 Subject: [PATCH 12/27] Replace YAML deployment presets with Python dataclass configs --- README.md | 12 +- deployments/qwen35-4b-fft-64k.yaml | 3 - deployments/qwen35-9b-lora-16k.yaml | 3 - deployments/qwen35-9b-lora-64k.yaml | 3 - deployments/qwen35_4b_fft_64k.py | 8 + deployments/qwen35_9b_lora_16k.py | 8 + deployments/qwen35_9b_lora_64k.py | 9 + docs/deployment-configs.md | 143 ++++++++++ ...validation.md => deployment-validation.md} | 17 +- docs/deployment-yaml-design.md | 270 ------------------ docs/design.md | 10 +- docs/observability.md | 2 +- docs/profiling.md | 2 +- pyproject.toml | 4 - scripts/deploy_models.sh | 10 +- ...eployment_smoke.py => deployment_smoke.py} | 2 +- src/lilo/backends/deployment.py | 2 +- src/lilo/backends/megatron_deployment.py | 2 +- src/lilo/configs/__init__.py | 1 + src/lilo/configs/qwen35_35b_a3b_fft_64k.py | 81 ++++++ src/lilo/configs/qwen35_4b_fft_64k.py | 68 +++++ src/lilo/configs/qwen35_9b_fft_64k.py | 70 +++++ .../configs/qwen35_9b_instruct_lora_16k.py | 81 ++++++ .../qwen35_9b_instruct_lora_16k_dp2.py | 11 + src/lilo/configs/qwen35_9b_lora_16k.py | 81 ++++++ src/lilo/configs/qwen35_9b_lora_16k_single.py | 11 + src/lilo/configs/qwen35_9b_lora_2k.py | 83 ++++++ src/lilo/configs/qwen35_9b_lora_64k.py | 14 + src/lilo/configs/qwen36_27b_fft_64k.py | 72 +++++ src/lilo/configs/qwen36_35b_a3b_fft_64k.py | 12 + src/lilo/configs/qwen38_27b_lora_128k.py | 19 ++ src/lilo/configs/qwen38_27b_lora_16k.py | 81 ++++++ src/lilo/configs/qwen38_27b_lora_64k.py | 15 + src/lilo/deployment_cli.py | 29 +- src/lilo/deployments.py | 236 +++++++-------- src/lilo/native_options.py | 2 +- src/lilo/presets/qwen35-35b-a3b-fft-64k.yaml | 62 ---- src/lilo/presets/qwen35-4b-fft-64k.yaml | 46 --- src/lilo/presets/qwen35-9b-fft-64k.yaml | 51 ---- .../qwen35-9b-instruct-lora-16k-dp2.yaml | 8 - .../presets/qwen35-9b-instruct-lora-16k.yaml | 58 ---- .../presets/qwen35-9b-lora-16k-single.yaml | 7 - src/lilo/presets/qwen35-9b-lora-16k.yaml | 58 ---- src/lilo/presets/qwen35-9b-lora-2k.yaml | 60 ---- src/lilo/presets/qwen35-9b-lora-64k.yaml | 13 - src/lilo/presets/qwen36-27b-fft-64k.yaml | 53 ---- src/lilo/presets/qwen36-35b-a3b-fft-64k.yaml | 8 - src/lilo/presets/qwen38-27b-lora-128k.yaml | 21 -- src/lilo/presets/qwen38-27b-lora-16k.yaml | 58 ---- src/lilo/presets/qwen38-27b-lora-64k.yaml | 16 -- src/lilo/providers/modal/app.py | 2 +- .../{yaml_apps.py => deployment_apps.py} | 4 +- ...aml_pool_app.py => deployment_pool_app.py} | 2 +- src/lilo/providers/modal/fft_pool.py | 4 +- src/lilo/providers/modal/lora_pool.py | 4 +- tests/backends/test_native_megatron_config.py | 11 +- tests/providers/conftest.py | 4 +- tests/providers/test_checkpoint_storage.py | 2 +- ...t_yaml_apps.py => test_deployment_apps.py} | 40 +-- ...elper.py => test_deployment_e2e_helper.py} | 4 +- tests/providers/test_deployment_presets.py | 24 +- tests/providers/test_modal_app.py | 4 +- tests/test_deployment_cli.py | 27 +- tests/test_deployments.py | 177 ++++++------ uv.lock | 2 - 65 files changed, 1178 insertions(+), 1129 deletions(-) delete mode 100644 deployments/qwen35-4b-fft-64k.yaml delete mode 100644 deployments/qwen35-9b-lora-16k.yaml delete mode 100644 deployments/qwen35-9b-lora-64k.yaml create mode 100644 deployments/qwen35_4b_fft_64k.py create mode 100644 deployments/qwen35_9b_lora_16k.py create mode 100644 deployments/qwen35_9b_lora_64k.py create mode 100644 docs/deployment-configs.md rename docs/{deployment-yaml-validation.md => deployment-validation.md} (87%) delete mode 100644 docs/deployment-yaml-design.md rename scripts/{yaml_deployment_smoke.py => deployment_smoke.py} (98%) create mode 100644 src/lilo/configs/__init__.py create mode 100644 src/lilo/configs/qwen35_35b_a3b_fft_64k.py create mode 100644 src/lilo/configs/qwen35_4b_fft_64k.py create mode 100644 src/lilo/configs/qwen35_9b_fft_64k.py create mode 100644 src/lilo/configs/qwen35_9b_instruct_lora_16k.py create mode 100644 src/lilo/configs/qwen35_9b_instruct_lora_16k_dp2.py create mode 100644 src/lilo/configs/qwen35_9b_lora_16k.py create mode 100644 src/lilo/configs/qwen35_9b_lora_16k_single.py create mode 100644 src/lilo/configs/qwen35_9b_lora_2k.py create mode 100644 src/lilo/configs/qwen35_9b_lora_64k.py create mode 100644 src/lilo/configs/qwen36_27b_fft_64k.py create mode 100644 src/lilo/configs/qwen36_35b_a3b_fft_64k.py create mode 100644 src/lilo/configs/qwen38_27b_lora_128k.py create mode 100644 src/lilo/configs/qwen38_27b_lora_16k.py create mode 100644 src/lilo/configs/qwen38_27b_lora_64k.py delete mode 100644 src/lilo/presets/qwen35-35b-a3b-fft-64k.yaml delete mode 100644 src/lilo/presets/qwen35-4b-fft-64k.yaml delete mode 100644 src/lilo/presets/qwen35-9b-fft-64k.yaml delete mode 100644 src/lilo/presets/qwen35-9b-instruct-lora-16k-dp2.yaml delete mode 100644 src/lilo/presets/qwen35-9b-instruct-lora-16k.yaml delete mode 100644 src/lilo/presets/qwen35-9b-lora-16k-single.yaml delete mode 100644 src/lilo/presets/qwen35-9b-lora-16k.yaml delete mode 100644 src/lilo/presets/qwen35-9b-lora-2k.yaml delete mode 100644 src/lilo/presets/qwen35-9b-lora-64k.yaml delete mode 100644 src/lilo/presets/qwen36-27b-fft-64k.yaml delete mode 100644 src/lilo/presets/qwen36-35b-a3b-fft-64k.yaml delete mode 100644 src/lilo/presets/qwen38-27b-lora-128k.yaml delete mode 100644 src/lilo/presets/qwen38-27b-lora-16k.yaml delete mode 100644 src/lilo/presets/qwen38-27b-lora-64k.yaml rename src/lilo/providers/modal/{yaml_apps.py => deployment_apps.py} (99%) rename src/lilo/providers/modal/{yaml_pool_app.py => deployment_pool_app.py} (93%) rename tests/providers/{test_yaml_apps.py => test_deployment_apps.py} (86%) rename tests/providers/{test_yaml_e2e_helper.py => test_deployment_e2e_helper.py} (90%) diff --git a/README.md b/README.md index 245981d..986cd59 100644 --- a/README.md +++ b/README.md @@ -48,7 +48,7 @@ training = service.create_lora_training_client( ## Shared deployment quick start -Shared deployments are defined in YAML. See [YAML deployments](docs/deployment-yaml-design.md). Keep the active YAML list in [scripts/deploy_models.sh](scripts/deploy_models.sh); run it to deploy the complete list. +Shared deployments use Python dataclasses. See [Python deployment configs](docs/deployment-configs.md). Keep the active Python config list in [scripts/deploy_models.sh](scripts/deploy_models.sh); run it to deploy the complete list. Install Lilo into your own Python project, deploy it once to Modal, then call its API from your training scripts. The commands below work in Bash or Zsh. @@ -134,12 +134,12 @@ uv run modal secret create lilo-proxy \ Create a configuration from a preset, validate it, and deploy it with Python 3.12: ```bash -uv run lilo config init --preset qwen35-9b-lora-16k > deployment.yaml -uv run lilo config validate deployment.yaml -uv run lilo deploy deployment.yaml +uv run lilo config init --preset qwen35-9b-lora-16k > deployment.py +uv run lilo config validate deployment.py +uv run lilo deploy deployment.py ``` -This deploys the shared app and prints its `server` URL. Add more YAML files to the same command to serve more configurations. Always supply the complete active set. Pin model revisions and `LILO_MILES_COMMIT` for repeatable applies; see [YAML deployments](docs/deployment-yaml-design.md). +This deploys the shared app and prints its `server` URL. Add more Python config files to the same command to serve more configurations. Always supply the complete active set. Pin model revisions and `LILO_MILES_COMMIT` for repeatable applies; see [Python deployment configs](docs/deployment-configs.md). From a repository checkout, maintain the list in `scripts/deploy_models.sh` and run that script. `lilo deploy` supplies the saved configuration to Modal; importing the shared app directly without a manifest is no longer a deployment entrypoint. @@ -156,7 +156,7 @@ has finished in the Modal dashboard or list apps with: uv run modal app list ``` -To tear down the deployment, stop its `lilo-fft-...` sampler apps, then the frontend named in your YAML (`lilo-yaml` by default), +To tear down the deployment, stop its `lilo-fft-...` sampler apps, then the frontend named in your config (`lilo-yaml` by default), using `uv run modal app stop `. Stopping the frontend does not stop sampler apps. ## Next steps diff --git a/deployments/qwen35-4b-fft-64k.yaml b/deployments/qwen35-4b-fft-64k.yaml deleted file mode 100644 index bb6102c..0000000 --- a/deployments/qwen35-4b-fft-64k.yaml +++ /dev/null @@ -1,3 +0,0 @@ -extends: builtin:qwen35-4b-fft-64k -model: - revision: 851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a diff --git a/deployments/qwen35-9b-lora-16k.yaml b/deployments/qwen35-9b-lora-16k.yaml deleted file mode 100644 index 66912d9..0000000 --- a/deployments/qwen35-9b-lora-16k.yaml +++ /dev/null @@ -1,3 +0,0 @@ -extends: builtin:qwen35-9b-lora-16k -model: - revision: 68c46c4b3498877f3ef123c856ecfde50c39f404 diff --git a/deployments/qwen35-9b-lora-64k.yaml b/deployments/qwen35-9b-lora-64k.yaml deleted file mode 100644 index b152564..0000000 --- a/deployments/qwen35-9b-lora-64k.yaml +++ /dev/null @@ -1,3 +0,0 @@ -extends: builtin:qwen35-9b-lora-64k -model: - revision: 68c46c4b3498877f3ef123c856ecfde50c39f404 diff --git a/deployments/qwen35_4b_fft_64k.py b/deployments/qwen35_4b_fft_64k.py new file mode 100644 index 0000000..911d983 --- /dev/null +++ b/deployments/qwen35_4b_fft_64k.py @@ -0,0 +1,8 @@ +from dataclasses import dataclass +from lilo.configs.qwen35_4b_fft_64k import Config as ParentConfig + + +@dataclass(kw_only=True) +class Config(ParentConfig): + def __post_init__(self): + self.model.revision = "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a" diff --git a/deployments/qwen35_9b_lora_16k.py b/deployments/qwen35_9b_lora_16k.py new file mode 100644 index 0000000..42f7060 --- /dev/null +++ b/deployments/qwen35_9b_lora_16k.py @@ -0,0 +1,8 @@ +from dataclasses import dataclass +from lilo.configs.qwen35_9b_lora_16k import Config as ParentConfig + + +@dataclass(kw_only=True) +class Config(ParentConfig): + def __post_init__(self): + self.model.revision = "68c46c4b3498877f3ef123c856ecfde50c39f404" diff --git a/deployments/qwen35_9b_lora_64k.py b/deployments/qwen35_9b_lora_64k.py new file mode 100644 index 0000000..161f28d --- /dev/null +++ b/deployments/qwen35_9b_lora_64k.py @@ -0,0 +1,9 @@ +from dataclasses import dataclass +from lilo.configs.qwen35_9b_lora_64k import Config as ParentConfig + + +@dataclass(kw_only=True) +class Config(ParentConfig): + def __post_init__(self): + super().__post_init__() + self.model.revision = "68c46c4b3498877f3ef123c856ecfde50c39f404" diff --git a/docs/deployment-configs.md b/docs/deployment-configs.md new file mode 100644 index 0000000..ec45ca8 --- /dev/null +++ b/docs/deployment-configs.md @@ -0,0 +1,143 @@ +# Python deployment configs + +Each deployment is a Python file exporting a `Config` dataclass that inherits from `BaseConfig`. The class contains the model, trainer, inference and Modal settings. There is no YAML loader or model catalog. + +The layout follows the Python recipe approach used by [training-gym](https://github.com/modal-labs/training-gym) and the [multinode training guide](https://github.com/modal-labs/multinode-training-guide/blob/main/nemo-rl/configs/llama3_1_8b_math_2node.py). Lilo's config classes use standard-library dataclasses; importing either project is not required. + +## Define a deployment + +Start with an example: + +```bash +lilo config init --preset qwen35-9b-lora-16k > deployments/my_model.py +``` + +The generated file imports a packaged config and subclasses it. Customize it using ordinary Python: + +```python +from dataclasses import dataclass +from lilo.configs.qwen35_9b_lora_16k import Config as ParentConfig + + +@dataclass(kw_only=True) +class Config(ParentConfig): + name: str = "my-9b-64k" + + def __post_init__(self): + self.model.max_context_length = 65536 + self.trainer.resources.gpu = "H200:8" + self.trainer.config["options"]["tensor_model_parallel_size"] = 8 + self.trainer.config["options"]["max_tokens_per_gpu"] = 65536 + self.inference.scaling.max_replicas = 6 +``` + +When a parent defines `__post_init__`, call `super().__post_init__()` before your changes. Mutable defaults use `field(default_factory=...)`, so modifying one instance does not change another config. New deployments can also inherit directly from `BaseConfig` and supply `Model`, `Trainer` and `Inference` fields; the packaged [16K example](../src/lilo/configs/qwen35_9b_lora_16k.py) shows the complete structure. + +Python imports provide reuse. There is no `extends` key or implicit dictionary merge. Backend dictionaries support normal Python operations such as `update()` or `|`. Config files execute as Python when loaded; keep provisioning and training calls outside them. Sibling imports are available while loading a file. + +All 14 previous presets are available under [`src/lilo/configs/`](../src/lilo/configs), with the same model, resource and backend settings. The [64K config](../src/lilo/configs/qwen35_9b_lora_64k.py) inherits from the 16K config. The three files under [`deployments/`](../deployments) pin the model revisions used in the earlier GPU checks. + +## Deploy the complete active set + +Use Python 3.12, matching the serialized trainer and inference images: + +```bash +lilo config validate deployments/my_model.py +lilo config resolve deployments/my_model.py --output /tmp/deployment.json +lilo deploy deployments/model_a.py deployments/model_b.py +``` + +`validate` loads the Python classes and checks shared frontend settings. It does not start backend libraries or prove that the model fits in GPU memory. `resolve` additionally pins model revisions and emits the deployment records as JSON; it does not provision compute. + +For a checked-in list, edit [`scripts/deploy_models.sh`](../scripts/deploy_models.sh). Add a Python config file and its path to the `deployment_files` array, then run: + +```bash +./scripts/deploy_models.sh +``` + +Supply every configuration that should remain available to new clients. Omitted configurations are retained for existing jobs but removed from new-client selection. All files in the list must agree on shared frontend, region, secrets, storage and lifecycle settings. Pin `LILO_MILES_COMMIT` for repeatable deployments. Credentials remain in Modal secrets; the config contains only secret names. + +## Code path + +```text +lilo deploy config.py + deployment_cli.main() + compile_configs() + deployments.load() → execute config.py → Config() + resolve model tag to a Hugging Face commit + DeploymentRecord.create() → copy settings and compute hash + deploy() + retain existing job configurations + save/pass manifest JSON + modal deploy -m lilo.providers.modal.app +``` + +[`load()`](../src/lilo/deployments.py) executes the file and instantiates its exported `Config` subclass. The returned dataclass goes directly to the orchestration code. There is no dictionary-to-config conversion on this path. + +`DeploymentRecord` adds the code identity, pinned Miles revision, configuration hash and active status. JSON is used only to store records and pass them to other processes. Pydantic reconstructs the standard dataclasses when reading those records; it does not import or run the user's config file in a GPU worker. Records preserve the full computed settings rather than a reference to the original Python file. + +| File | Responsibility | +| --- | --- | +| [`deployments.py`](../src/lilo/deployments.py) | Dataclasses, Python file loader, saved record and shared-frontend checks | +| [`deployment_cli.py`](../src/lilo/deployment_cli.py) | Revision lookup, manifest updates and Modal deployment | +| [`deployment_apps.py`](../src/lilo/providers/modal/deployment_apps.py) | Trainer functions, inference server classes and process startup | +| [`app.py`](../src/lilo/providers/modal/app.py) | Shared frontend and `app.include()` for generated trainers | +| [`deployment_pool_app.py`](../src/lilo/providers/modal/deployment_pool_app.py) | Construct an inference pool from its saved record | +| [`backends/deployment.py`](../src/lilo/backends/deployment.py) | Select the backend configuration reader | + +`build_trainer_app(record)` configures resources, secrets, storage and limits. The shared app includes its generated trainer function. When that function starts, `run_trainer()` obtains backend settings and passes them as `LILO_BACKEND_CONFIG` to the existing Miles or Megatron executor. + +Inference pools are separate apps created on demand. `build_rollout_app()` constructs a server from the saved configuration. Startup launches SGLang with native options, waits for its health endpoint, then starts the LoRA or FFT sidecar. + +## Backend options + +`trainer.config` and `inference.config` remain open dictionaries. Adding an upstream option does not require adding a deployment dataclass field. Backend readers check values used by Lilo's integration, while installed backend libraries check native options at worker startup. + +For Miles, `trainer.config` contains: + +```python +{ + "model_args": "qwen3.5-9B", + "options": { + "tensor_model_parallel_size": 4, + "multi_lora_n_adapters": 6, + "recompute_granularity": "full", + }, +} +``` + +`model_args` selects a Miles architecture preset; it can be omitted when explicit architecture options are provided. Miles's actual argument parser validates native options. Lilo also reads parallelism, adapter slots/rank and targets for admission and adapter export. + +For Megatron, the [FFT example](../src/lilo/configs/qwen35_4b_fft_64k.py) separates: + +- `runtime`: Lilo's training loop, packing and parallelism settings. +- `provider`: attributes assigned to the actual Megatron Bridge model provider. +- `optimizer`: optimizer settings, including additional native Megatron constructor options. +- `distributed`: additional `DistributedDataParallelConfig` constructor options. + +These section names are owned by Lilo. Native provider/optimizer/distributed fields do not need a Lilo allowlist. Shared precision and parallelism controls remain protected so Lilo's packing and collectives agree with Megatron. Tinker optimizer steps use Adam parameters supplied by the client. Native optimizer/distributed settings are recorded and checked for exact FFT checkpoint resume. + +`inference.config` contains SGLang options directly, such as `tp_size`, `mem_fraction_static` and `max_running_requests`. The real SGLang parser checks them during startup. Miles and SGLang support ordinary scalar, boolean and list arguments; unsupported custom/repeated argparse actions fail explicitly. + +Backend readers still check managed paths, GPU topology, adapter capacity and communication settings. Python configs do not bypass backend compatibility, adapter export/load requirements or available GPU memory. + +## Routing and updates + +The frontend URL is shared across models. Clients select the model through `base_model`: + +```python +service = tinker.ServiceClient(base_url=lilo_url, api_key=api_key) +training = service.create_lora_training_client( + base_model="Qwen/Qwen3.5-9B-Base", rank=32, +) +``` + +If multiple configurations serve the same model and training mode, select one with `routing.default=True`. Otherwise client creation reports the ambiguity. `routing.sampling_default` resolves sampling-only selection across LoRA and FFT configurations. Training-derived sampling remains attached to the client's saved configuration. + +Every client stores its selected definition ID. Changing routing defaults affects new clients. Changing compute or backend settings creates a new configuration ID, and old configurations remain registered for existing jobs and checkpoints. The CLI serializes applies and retains interrupted deployments for recovery. + +The shared app is deployed on each apply. Existing trainer functions use stable captured configuration JSON, and unchanged inference pools retain their apps. This is the shared-app architecture, not independent trainer-app deployment. Source/runtime upgrades still require a separate frontend in this draft. Config files outside the installed Lilo package can change without changing the runtime source fingerprint. + +Existing resource names (`lilo-yaml`, `*-yaml-deployments`, and the `yaml_` definition prefix) are retained so this authoring change does not rename saved resources. They no longer indicate a YAML ingestion path. PyYAML is not a direct Lilo dependency; other installed libraries may depend on it. + +See [validation results](deployment-validation.md) for CPU coverage and the earlier GPU smoke tests. The Python-config migration has not been redeployed. diff --git a/docs/deployment-yaml-validation.md b/docs/deployment-validation.md similarity index 87% rename from docs/deployment-yaml-validation.md rename to docs/deployment-validation.md index c9e2afb..65ab4e5 100644 --- a/docs/deployment-yaml-validation.md +++ b/docs/deployment-validation.md @@ -1,7 +1,16 @@ -# YAML deployment validation +# Deployment configuration validation These checks exercise the shared-app deployment path in PR #55. Each deployment contains the frontend and generated trainer functions; inference pools are separate apps started on demand. +## Python dataclass configuration + +Python `Config` subclasses replace the YAML presets and loader. All 14 configs were compared with the previous specifications: model, resources, backend options and inference options are unchanged. The three checked-in deployment files retain their pinned model revisions. Inheritance uses normal dataclass defaults and `__post_init__`; tests verify mutable defaults are independent. + +Validation: **597 CPU tests passed, 1 skipped**; Ruff passed for changed Python files and whitespace/shell-syntax checks passed. The CLI generated a Python config and validated it. A clean wheel contains all 14 Python configs and no YAML presets or old YAML app modules; every config was loaded from that wheel, its backend settings built, and its record round-tripped through JSON. PyYAML was removed from Lilo's direct dependencies; the lockfile otherwise preserves package versions and sources. + +This migration has not been redeployed or GPU-tested. The GPU results below describe the earlier YAML-based source at `13d2a31`, before the subsequent CPU-only refactors. Earlier cleanup counts below are historical. + + ## Deployment record simplification The saved manifest wrapper is now named `DeploymentRecord`. Its factory copies the parsed specification and computes the configuration hash without a dictionary-to-model round trip. Model tag resolution and validation of the Hugging Face result stay in the CLI. The standalone `resolve()` function was removed. Serialized record fields and hash format are unchanged. @@ -73,11 +82,11 @@ The western 8×H200 trainer remained queued without a container for approximatel ## Reproduce the smoke test -Use Python 3.12 and an authenticated Modal environment. Set `TINKER_API_KEY` locally to the value in the deployment's API secret. Generate the preset YAMLs, give them the same isolated `deployment.frontend`, and pin model revisions plus `LILO_MILES_COMMIT` before deploying. +Use Python 3.12 and an authenticated Modal environment. Set `TINKER_API_KEY` locally to the value in the deployment's API secret. Generate Python config files from the packaged examples, give them the same isolated `deployment.frontend`, and pin model revisions plus `LILO_MILES_COMMIT` before deploying. ```bash -lilo deploy qwen35-9b-lora-16k.yaml qwen35-9b-lora-64k.yaml qwen35-4b-fft-64k.yaml -python scripts/yaml_deployment_smoke.py \ +lilo deploy deployments/qwen35_9b_lora_16k.py deployments/qwen35_9b_lora_64k.py deployments/qwen35_4b_fft_64k.py +python scripts/deployment_smoke.py \ --frontend YOUR_TEST_FRONTEND \ --name qwen35-9b-lora-16k \ --steps 3 \ diff --git a/docs/deployment-yaml-design.md b/docs/deployment-yaml-design.md deleted file mode 100644 index 1c019d7..0000000 --- a/docs/deployment-yaml-design.md +++ /dev/null @@ -1,270 +0,0 @@ -# YAML deployments - -Shared Modal deployments are defined through YAML. A file specifies the base model, trainer backend, resources, context length, adapter capacity, and inference settings. Adding a backend-supported model does not require a new Python definition or a catalog entry. - -The implementation has CPU tests and live training/sampling checks; see the [validation report](deployment-yaml-validation.md) for tested configurations and shared-app redeploy results. The separate scoped `lilo.run(...)` API remains available. - -## The provider and Modal app structure - -Start with [`yaml_apps.py`](../src/lilo/providers/modal/yaml_apps.py). It contains the two generic builders and the trainer entrypoint: - -```python -trainer_app, trainer_function = build_trainer_app(resolved) -pool_app, server_class = build_rollout_app(resolved, pool) -``` - -Both receive the resolved configuration as data. Neither imports a model-specific definition module. - -```mermaid -flowchart TD - YAML[Complete set of deployment YAMLs] --> CLI[lilo deploy] - CLI --> Registry[Modal Dict: saved configurations and apply lock] - CLI --> Frontend[Shared Modal app: deployment.frontend] - Frontend --> HTTP[Tinker HTTP server] - Frontend --> Assets[prepare_model_assets: CPU] - Frontend --> Reconcile[trainer_reconciler and idle sweep: CPU] - Frontend --> Sampling[execute_sample: CPU] - Frontend --> TrainerA[Generated trainer function A: GPU] - Frontend --> TrainerB[Generated trainer function B: GPU] - Frontend --> Ensure[ensure_lora_pool / ensure_fft_pool: CPU] - Ensure --> Pool[Separate Modal rollout-pool app] - Pool --> Replica[Server replicas: GPU] - Replica --> Sidecar[LoRA or FFT sidecar on port 8000] - Sidecar --> SGLang[SGLang on port 8001] -``` - -`app.py` reads a resolved manifest from `LILO_DEPLOYMENT_MANIFEST`. For each configuration it calls `definition_from_spec`, which constructs the routing metadata and a trainer app. `app.include` places those trainer functions inside the shared frontend app. A missing manifest is an error with instructions to use `lilo deploy`; there is no Python model-catalog fallback. - -The trainer function's GPU type/count, CPU, RAM, timeout, maximum instances, secrets and mounted volumes come from YAML. Containers remain single-use. `run_trainer` reloads the prepared asset volume, constructs the backend configuration, and calls the existing `run_engine_with_backend` launcher. Miles uses one controller process that manages its GPU workers; FFT launches one process per allocated GPU. Client admission and sampler-persistence concurrency are configured separately. - -`ensure_lora_pool` and `ensure_fft_pool` use the existing pool deployment and cleanup machinery. For YAML definitions, their deployment subprocess imports `yaml_pool_app.py` and receives the saved configuration through `LILO_POOL_DEPLOYMENT`. The resulting `Server` class captures that configuration. Startup launches SGLang with native options, waits for its health endpoint, then starts the appropriate sidecar and process supervisor. Shutdown terminates both children. - -A LoRA pool is shared by clients using the same deployment configuration and base weights. FFT retains its existing per-client latest pools and pinned-version/base pools. GPU resources and autoscaling settings are attached to each generated `Server` class when its app is constructed. - -| File | Purpose | -| --- | --- | -| [`deployments.py`](../src/lilo/deployments.py) | Schema, YAML inheritance, revision pinning and configuration identifiers | -| [`deployment_cli.py`](../src/lilo/deployment_cli.py) | Operator commands, saved manifests and serialized applies | -| [`backends/deployment.py`](../src/lilo/backends/deployment.py) | Dispatch configuration to backend-owned readers; no Modal dependency | -| [`yaml_apps.py`](../src/lilo/providers/modal/yaml_apps.py) | Declare trainer functions and rollout server classes; start their processes | -| [`app.py`](../src/lilo/providers/modal/app.py) | Register generated trainers with the existing shared app | -| [`yaml_pool_app.py`](../src/lilo/providers/modal/yaml_pool_app.py) | Construct a rollout app in the pool deployment subprocess | -| [`control_plane/deployments.py`](../src/lilo/control_plane/deployments.py) | Select a deployment from `base_model` and training mode | -| [`native_options.py`](../src/lilo/native_options.py) | Apply YAML values through each backend's argparse schema | - -## Configuration and commands - -Generate a complete editable file: - -```bash -lilo config init --preset qwen35-9b-lora-16k > deployment.yaml -lilo config validate deployment.yaml -lilo config resolve deployment.yaml --output deployment.resolved.json -``` - -`validate` is offline and does not contact Modal. `resolve` resolves the HF revision and Miles runtime revision; it may contact Hugging Face and GitHub but does not allocate GPUs. Use an exact HF commit and `LILO_MILES_COMMIT` to avoid moving branch references. The resolved JSON is inspectable deployment data; `deploy` takes YAML files and resolves them again. For repeatable later applies, put those exact revisions in the YAML/environment. - -The packaged presets are: - -- [`qwen35-9b-lora-16k.yaml`](../src/lilo/presets/qwen35-9b-lora-16k.yaml): Qwen3.5-9B-Base, rank 32, six clients per H100:4 trainer, H200:1 inference replicas. -- [`qwen35-9b-lora-64k.yaml`](../src/lilo/presets/qwen35-9b-lora-64k.yaml): a larger-context example using H200:8 training. -- [`qwen35-4b-fft-64k.yaml`](../src/lilo/presets/qwen35-4b-fft-64k.yaml): the existing 4B FFT topology expressed as YAML. - -Additional packaged YAML presets preserve the earlier Python recipes for 2K and single-client LoRA, the 9B instruct model, Qwen3.5/3.6 FFT variants, and Qwen3.8 LoRA at 16K/64K/128K. They are optional templates, not automatically deployed models. The deployment script still lists only the three configurations above. Migrated presets use the YAML defaults of at most one trainer and zero-to-eight inference replicas; raise these limits in your YAML when needed. Their model/context/parallelism settings are covered by CPU tests, not new GPU validation. - -These are starting configurations. The 16K and FFT backend settings are based on the existing definitions; inference minima are explicitly zero and the trainer maximum is one. All three presets passed short GPU training/sampling checks; see the [validation report](deployment-yaml-validation.md). Full-context memory capacity was not tested. - -You can instead keep a small override file: - -```yaml -extends: builtin:qwen35-9b-lora-16k -name: my-9b-16k -model: - id: Qwen/Qwen3.5-9B-Base - revision: main -deployment: - frontend: my-lilo-yaml - modal: - environment: dev - region: us-west - secrets: - api: lilo-api - sampler_proxy: lilo-proxy - huggingface: huggingface-secret -trainer: - resources: - gpu: H200:8 - engine: - max_clients_per_instance: 12 - config: - options: - tensor_model_parallel_size: 8 - multi_lora_n_adapters: 12 -inference: - scaling: - min_replicas: 0 - max_replicas: 6 -``` - -A local `extends: ./base.yaml` also works. Maps merge recursively, lists replace, and `false` overrides `true`. Duplicate YAML keys, unknown Lilo fields and inheritance cycles are rejected. YAML contains secret names; credentials stay in Modal secrets. The API and proxy secret contents are the same as in the [shared deployment setup](../README.md#2-configure-modal-and-secrets-once). - -Run `lilo deploy` with Python 3.12, matching the serialized trainer and rollout images. The YAML frontend image also uses Python 3.12 because it launches rollout deployment subprocesses. - -When ready to deploy, supply the **complete active set** of files for one frontend: - -```bash -lilo deploy model-a.yaml model-b.yaml model-a-64k.yaml -``` - -For a checked-in deployment list, use [`scripts/deploy_models.sh`](../scripts/deploy_models.sh). Its `deployment_files` array lists the editable YAMLs under [`deployments/`](../deployments), initially covering the three presets above with the model revisions used in GPU validation. These inherit the `lilo-yaml` frontend and the presets' shared settings. - -```bash -./scripts/deploy_models.sh -``` - -To add a model, create its YAML in `deployments/`, add its path to `deployment_files`, then run the script. The YAMLs must agree on the frontend and shared settings. Keep existing entries to keep those configurations available to new clients; removing an entry retires it from new-client selection on the next deployment. The script works from any working directory and uses `lilo` from your active Python 3.12 environment. Keep the same pinned `LILO_MILES_COMMIT` across applies, as with the direct CLI. - -This command builds/deploys the shared app; it is not a validation command. All files must agree on frontend, Modal environment/region, secrets, storage and shared lifecycle settings. The preset frontend name is `lilo-yaml`. The CLI refuses to overwrite a pre-existing application without a YAML registry, so it cannot accidentally replace an unrelated app. - -Trainer minimum capacity is zero. Rollout apps are created on first demand, and their configured inference minimum applies once the pool exists. A pool with a nonzero minimum will keep that many workers warm until it is stopped by the existing idle cleanup. - -## Routing and existing clients - -The frontend URL stays the same across models: - -```python -service = tinker.ServiceClient(base_url=lilo_url, api_key=api_key) -a = service.create_lora_training_client( - base_model="Qwen/Qwen3.5-9B-Base", rank=32, -) -b = service.create_lora_training_client( - base_model="organization/another-configured-model", rank=16, -) -``` - -The active YAMLs generate the lookup from `(model.id, parameterization)` to deployed configurations. A unique match is selected automatically. If both 16K and 64K configurations exist for the same model and mode, exactly one can set `routing.default: true`. Without a default, creation returns an ambiguity error listing the alternatives. Capacity pressure does not change which configuration is selected. - -Sampling-only requests select the model's default configuration. If both LoRA and FFT remain eligible, set `routing.sampling_default: true` on the desired one. Training-derived sampling uses the training client's saved definition. - -`get_server_capabilities` advertises canonical model names and the selected context limits. Where the model-level response must cover both FFT and LoRA, it reports the smaller default context limit. Ambiguous models are omitted. The authenticated `GET /api/v1/lilo/deployments` endpoint lists active deployment names, saved definition identifiers, context limits, modes and defaults. - -Each client records its selected definition identifier. Changing defaults affects new clients. Old trainer definitions remain registered, so existing models, sampling pools and checkpoint restores can keep referring to them. Explicit definition IDs retain the existing compatibility lookup; they are not model names advertised to ordinary clients. - -## Backend options and model support - -The deployment schema reads model identity, GPU resources, scaling, routing, storage and lifecycle settings. Each trainer and inference definition has a `backend` name and an opaque `config` mapping. The schema does not enumerate backend options. Backend-specific readers live beside the backend code, and the Modal app builder consumes their output. There is no Modal `recipe.py`. - -For Miles: - -```yaml -trainer: - backend: miles - resources: - gpu: H100:4 - config: - model_args: qwen3.5-9B - options: - tensor_model_parallel_size: 4 - recompute_granularity: full - recompute_method: uniform - recompute_num_layers: 1 -``` - -`config.model_args` names an architecture preset inside Miles. It can be omitted when `config.options` supplies explicit architecture options. Changing `model.id` does not make an inherited architecture preset compatible; Miles still validates the HF configuration at startup. Miles options use its argparse destination names. The integration also reads parallelism, adapter limits and targets because Lilo uses those values for admission and adapter export. Other options reach Miles's parser without a Lilo allowlist. - -For Megatron FFT: - -```yaml -trainer: - backend: megatron - resources: - gpu: H100:4 - engine: - sampler_persistence_concurrency: 1 - config: - runtime: - tensor_model_parallel_size: 2 - context_parallel_size: 2 - sequence_parallel: true - max_tokens_per_microbatch: 65536 - use_distributed_optimizer: true - provider: - recompute_granularity: full - recompute_method: uniform - recompute_num_layers: 1 - optimizer: - lr: 0.0001 - loss_scale: 1.0 - distributed: - grad_reduce_in_fp32: true -``` - -- `runtime` configures Lilo's training loop, packing and parallelism. It has a finite schema because Lilo implements these settings. -- `provider` sets attributes on the model provider returned by Megatron Bridge. Names are checked against that actual provider at worker startup. Values, including lists and mappings, are preserved. -- `optimizer` accepts Megatron optimizer constructor options. Existing Lilo optimizer and scheduling fields (such as `lr` and `min_lr`) remain available to the loop; additional fields pass directly to Megatron's `OptimizerConfig`. Each Tinker `optim_step` still supplies the request's Adam parameters. -- `distributed` passes additional fields directly to Megatron's `DistributedDataParallelConfig`. - -New native Megatron provider, optimizer or distributed options do not require deployment-parser changes. The installed backend rejects unsupported options on worker startup. Optimizer and distributed overrides are recorded in FFT checkpoint metadata and checked on exact resume. - -For SGLang, `inference.config` directly contains its options: - -```yaml -inference: - backend: sglang - resources: - gpu: H200:1 - config: - tp_size: 1 - mem_fraction_static: 0.8 - max_running_requests: 32 -``` - -Miles and SGLang use their real argument parsers before initialization. Boolean flags, scalar values and ordinary list arguments are supported. Both spellings of opposing boolean flags are replaced when they share a destination. Unknown options and custom/repeated argparse actions fail explicitly. These are parser-backed options, so arbitrary Python objects and custom actions are not supported. - -When preparing a trainer or inference pool, its backend reader checks settings that affect the integration: GPU counts and parallelism, maximum clients versus adapter slots, managed model paths, context/rank configuration, communication endpoints and trainer-only mode. The backend readers reject conflicting values. For Megatron, set parallelism and shared distributed-optimizer controls under `runtime` so both Lilo and Megatron receive the same values. Tinker training requires an Adam optimizer. Native passthrough does not make other training protocols or unsupported process layouts work automatically. - -All built-in YAMLs use this structure. The earlier draft's `trainer.miles`, `trainer.megatron` and `inference.sglang` sections have been removed; user YAMLs overriding those sections must move them under `config` as shown above. This schema change is part of the draft and has not been deployed. Existing source-fingerprint checks continue to prevent applying a different runtime implementation over running trainers. - -Inference adapter targets are derived from the existing Miles-to-PEFT mapping unless `lora_target_modules` is explicitly supplied. This mapping does not establish support for every architecture. A model still needs compatible training, adapter export and SGLang loading implementations in the selected images. - -YAML loading validates document fields and inheritance only; it does not call backend readers. Backend errors surface when their trainer or pool settings are built, without a Pydantic wrapper. The loader walks the inheritance chain and merges it from parent to child before constructing `DeploymentSpec`; partial parent files are supported. - -Local validation cannot establish memory fit or prove that an unfamiliar model works. Native parser validation occurs inside the runtime images on startup. This draft does not implement a separate image-preflight command or a GPU export/load/generation probe. - -## From YAML to a saved deployment - -`load()` returns a `DeploymentSpec`: the merged YAML fields. It does not interpret backend options. - -`deployment_cli.compile_configs()` resolves a model branch or tag to a Hugging Face commit. A revision that is already a commit needs no lookup. It also identifies the Lilo code and Miles revision being deployed. - -`DeploymentRecord.create()` copies the specification, sets that pinned revision, and hashes the code identity and configuration. It does not serialize and reparse the specification or call backend readers. The standalone `resolve()` function is gone. - -The separate record holds information used by the deployment manifest: the specification, code identity, Miles revision, configuration hash, and whether new clients can select it. These fields are not user YAML options. The hash keeps jobs attached to the configuration that created them; the active flag lets an older configuration keep serving existing jobs after it is removed from the deployment list. `definition_id` and the model asset path are derived properties. Stored JSON is decoded when crossing a registry or worker-process boundary. - -The class is named `DeploymentRecord` to describe that role. Its serialized fields and hash format are unchanged from the earlier `ResolvedDeployment` name. - -## Saved configurations, failures and updates - -The CLI stores configurations in the Modal Dict `-yaml-deployments`, scoped to the chosen Modal environment. A single apply lock serializes registry changes. A pending manifest is written before deployment; only successful deployment replaces the committed manifest. The next attempt retains pending configurations too, covering an interruption after Modal accepted a deployment but before the CLI saved its result. - -Applying a changed YAML produces a new definition identifier. The hash includes the pinned model revision, normalized settings and implementation fingerprint. Routing preferences are excluded, so switching a default does not change trainer identity. Old configurations are retained as inactive entries. Trainer functions capture their settings as consistently ordered JSON. The captured copy uses fixed admission and routing flags, so retiring a configuration or changing a default does not change its trainer function. Scaling changes currently also create a new identifier; a separate scaling-policy revision is future work. - -The implementation fingerprint includes shipped Lilo source, declared dependencies and the selected Miles commit. The existing image recipes supply the other backend source revisions. This is not a fully pinned Python/container dependency lock. This draft rejects applies that would rebuild retained configurations with a different implementation fingerprint or different shared storage/lifecycle settings; use a separate frontend for those upgrades. Automatic pruning of historical configurations and migration across code/image versions are not implemented. - -HF asset paths hash the full repository name and exact revision. The trainer and sampler use that same directory. Miles checkpoints and FFT native-resume metadata record the base revision and reject a mismatched or unknown revision when resuming into a pinned deployment. Legacy deployments retain their existing behavior when neither side records a revision. FFT portable weights-only loading retains its existing compatibility checks. - -A backend startup failure is recorded for that YAML definition. Waiting creation futures receive an error with the failed instance identifier, and further trainer launches are blocked for that definition. Detailed backend stderr is available in Modal call logs. Existing placed jobs are not invalidated. Once an operator fixes a transient cause, they can explicitly retry: - -```bash -lilo deployment retry --frontend my-lilo-yaml --env dev yaml_NAME_GENERATION -``` - -This clears the recorded failure and requests reconciliation; it may start GPU trainers if demand remains. Changing an invalid configuration produces a new definition instead. Failures before the trainer process starts, such as image-build failures, still rely on Modal's deployment diagnostics. Rollout initialization errors use the existing pool/sampling error path; trainer creation does not yet wait for an adapter-generation compatibility probe. - -A killed CLI can leave its apply lock behind. Confirm that the original apply has stopped before running `lilo deployment unlock --frontend NAME --env ENV`. The lock is not automatically stolen while a slow deployment may still be running. - -## Current scope and validation - -The implemented YAML path supports shared, single-node Miles LoRA and Megatron FFT deployments with the existing runtime images. The inference GPU allocation must match SGLang’s total `tp_size`. Data-parallel attention uses explicit `dp_size` and `enable_dp_attention` options, with `dp_size` dividing the allocation. Scoped `lilo.run(config=...)`, custom image selection, automatic provisioning of unknown `base_model` values, automatic runtime upgrades, and GPU compatibility probes remain follow-up work. Unsupported schema choices are rejected rather than treated as implemented features. - -CPU coverage exercises configuration loading and validation, typed native overrides, routing several models through one HTTP service, ambiguity handling, preserved client definitions, interrupted/concurrent applies, trainer resources and executor settings, LoRA/FFT pool startup and shutdown, and startup-error handling. Existing backend, provider, HTTP and scoped-run tests also run. The [GPU validation report](deployment-yaml-validation.md) covers short training/sampling requests and shared-app redeploy continuity. Maximum-context capacity and sustained-load testing remain separate checks. diff --git a/docs/design.md b/docs/design.md index 81647f6..b431a8d 100644 --- a/docs/design.md +++ b/docs/design.md @@ -124,20 +124,20 @@ sampling scales according to rollout traffic. We best-effort sticky-route groups ## Adding a new model deployment -A deployment YAML specifies the base model, training mode, context length, GPUs, parallelism, and inference settings. Shared deployments are defined only through these files. +A Python deployment dataclass specifies the base model, training mode, context length, GPUs, parallelism, and inference settings. Shared deployments are defined only through these files. -1. Create a YAML under `deployments/`, optionally extending a packaged preset. +1. Create a Python `Config` subclass under `deployments/`, optionally inheriting from a packaged config. 2. Add its path to the list in [`scripts/deploy_models.sh`](../scripts/deploy_models.sh). 3. Run the script to apply the complete list to the shared frontend. -There is no model-specific Python module or catalog registration to update. The generic builders in [`yaml_apps.py`](../src/lilo/providers/modal/yaml_apps.py) construct trainer functions and inference apps from the resolved YAML. See [YAML deployments](deployment-yaml-design.md) for the configuration schema and app structure. +No central model catalog registration is needed. The generic builders in [`deployment_apps.py`](../src/lilo/providers/modal/deployment_apps.py) construct trainer functions and inference apps from the saved config. See [Python deployment configs](deployment-configs.md) for the configuration schema and app structure. -Set each configuration's trainer limit with `trainer.scaling.max_instances` and its inference limits with `inference.scaling`. Trainer limits are read from YAML; `LILO_TRAINER_MAX_CONTAINERS` is no longer used. +Set each configuration's trainer limit with `trainer.scaling.max_instances` and its inference limits with `inference.scaling`. Trainer limits are read from the config; `LILO_TRAINER_MAX_CONTAINERS` is no longer used. To check training, publication, and sampling against a deployed configuration: ```bash -uv run python scripts/yaml_deployment_smoke.py \ +uv run python scripts/deployment_smoke.py \ --frontend lilo-yaml --name qwen35-9b-lora-16k \ --output /tmp/lilo-smoke.json ``` diff --git a/docs/observability.md b/docs/observability.md index fa5db3f..0196f66 100644 --- a/docs/observability.md +++ b/docs/observability.md @@ -45,7 +45,7 @@ Deploy Lilo after updating the secret: ```bash MODAL_PROFILE=your-workspace MODAL_ENVIRONMENT=your-environment \ - uv run lilo deploy deployment.yaml + uv run lilo deploy deployment.py ``` The configuration applies to the control plane, sampling workers, and new diff --git a/docs/profiling.md b/docs/profiling.md index 2102d16..dfac5ed 100644 --- a/docs/profiling.md +++ b/docs/profiling.md @@ -14,7 +14,7 @@ Pick the step you want to trace and set one environment variable in the shell you deploy from. The trainer inherits it: ```bash -LILO_TORCH_PROFILE_STEP=2 uv run lilo deploy deployment.yaml +LILO_TORCH_PROFILE_STEP=2 uv run lilo deploy deployment.py ``` Then run your training loop as usual. Steps are counted from 0, so `2` traces diff --git a/pyproject.toml b/pyproject.toml index 8ea6096..26e74d7 100644 --- a/pyproject.toml +++ b/pyproject.toml @@ -14,7 +14,6 @@ dependencies = [ "huggingface-hub>=0.34", "modal>=1.5.3", "protobuf>=5.29", - "pyyaml>=6.0.2", "pydantic>=2.13.4", "stitch @ git+https://github.com/modal-projects/stitch.git@375a9396a7b05770dc4ed9cc5fe34fc4d5a472d5", "tinker>=0.24.1,<0.25", @@ -36,6 +35,3 @@ dev = [ [project.scripts] lilo = "lilo.deployment_cli:main" - -[tool.setuptools.package-data] -lilo = ["presets/*.yaml"] diff --git a/scripts/deploy_models.sh b/scripts/deploy_models.sh index 3b8218a..5aa3df4 100755 --- a/scripts/deploy_models.sh +++ b/scripts/deploy_models.sh @@ -1,15 +1,15 @@ #!/usr/bin/env bash set -e -# Run from the repository root so the YAML paths below work from any directory. +# Run from the repository root so the Python config paths below work from any directory. cd "$(dirname "$0")/.." -# Add a model by creating its YAML in deployments/ and adding it here. +# Add a model by creating its Python config in deployments/ and adding it here. # Keep every configuration that should be available to new clients in this list. deployment_files=( - deployments/qwen35-9b-lora-16k.yaml - deployments/qwen35-9b-lora-64k.yaml - deployments/qwen35-4b-fft-64k.yaml + deployments/qwen35_9b_lora_16k.py + deployments/qwen35_9b_lora_64k.py + deployments/qwen35_4b_fft_64k.py ) lilo deploy "${deployment_files[@]}" diff --git a/scripts/yaml_deployment_smoke.py b/scripts/deployment_smoke.py similarity index 98% rename from scripts/yaml_deployment_smoke.py rename to scripts/deployment_smoke.py index 2fa851e..ef94c55 100644 --- a/scripts/yaml_deployment_smoke.py +++ b/scripts/deployment_smoke.py @@ -1,4 +1,4 @@ -"""Small real-GPU training/publication/sampling check for a YAML deployment. +"""Small real-GPU training/publication/sampling check for a Python-configured deployment. Run from an authenticated operator environment. Results are written to --output; credentials are read from TINKER_API_KEY and never included in the report. diff --git a/src/lilo/backends/deployment.py b/src/lilo/backends/deployment.py index b4752f6..baec4f0 100644 --- a/src/lilo/backends/deployment.py +++ b/src/lilo/backends/deployment.py @@ -1,4 +1,4 @@ -"""Dispatch opaque YAML configuration to its backend-owned integration. +"""Dispatch backend configuration to its backend-owned integration. These readers are CPU-only. Native libraries validate their options in workers. """ diff --git a/src/lilo/backends/megatron_deployment.py b/src/lilo/backends/megatron_deployment.py index ac420f1..32d8fe7 100644 --- a/src/lilo/backends/megatron_deployment.py +++ b/src/lilo/backends/megatron_deployment.py @@ -37,7 +37,7 @@ def build_config(spec, asset_path): trainer = spec.trainer if spec.model.parameterization != "full": - raise ValueError("Megatron YAML deployments require full parameterization") + raise ValueError("Megatron deployments require full parameterization") if trainer.engine.max_clients_per_instance != 1: raise ValueError("FFT trainers admit one client per instance") if trainer.engine.sampler_persistence_concurrency != 1: diff --git a/src/lilo/configs/__init__.py b/src/lilo/configs/__init__.py new file mode 100644 index 0000000..4c40641 --- /dev/null +++ b/src/lilo/configs/__init__.py @@ -0,0 +1 @@ +"""Example deployment dataclasses; import and subclass any Config to customize it.""" diff --git a/src/lilo/configs/qwen35_35b_a3b_fft_64k.py b/src/lilo/configs/qwen35_35b_a3b_fft_64k.py new file mode 100644 index 0000000..23508f4 --- /dev/null +++ b/src/lilo/configs/qwen35_35b_a3b_fft_64k.py @@ -0,0 +1,81 @@ +from dataclasses import dataclass, field +from lilo.deployments import ( + BaseConfig, + EngineOptions, + Inference, + InferenceScaling, + Model, + Resources, + Routing, + Trainer, +) + + +@dataclass(kw_only=True) +class Config(BaseConfig): + name: str = "qwen35-35b-a3b-fft-64k" + model: Model = field( + default_factory=lambda: Model( + id="Qwen/Qwen3.5-35B-A3B", parameterization="full", max_context_length=65536 + ) + ) + routing: Routing = field(default_factory=lambda: Routing(default=True)) + trainer: Trainer = field( + default_factory=lambda: Trainer( + backend="megatron", + resources=Resources(gpu="H200:8"), + engine=EngineOptions( + max_clients_per_instance=1, sampler_persistence_concurrency=1 + ), + env={ + "PYTORCH_CUDA_ALLOC_CONF": "expandable_segments:True", + "TORCHINDUCTOR_COMPILE_THREADS": "1", + }, + config={ + "runtime": { + "tensor_model_parallel_size": 4, + "pipeline_model_parallel_size": 1, + "context_parallel_size": 2, + "expert_model_parallel_size": 8, + "expert_tensor_parallel_size": 1, + "sequence_parallel": True, + "micro_batch_size": 1, + "max_tokens_per_microbatch": 65536, + "bf16": True, + "fp16": False, + "gpu_memory_fraction": 0.9, + "use_distributed_optimizer": True, + }, + "provider": { + "mtp_num_layers": 0, + "recompute_granularity": "selective", + "moe_layer_recompute": True, + "moe_token_dispatcher_type": "alltoall", + "moe_router_fusion": True, + "moe_permute_fusion": True, + "moe_grouped_gemm": True, + "moe_shared_expert_overlap": False, + "moe_aux_loss_coeff": 0.0, + }, + "optimizer": {"optimizer": "adam", "lr": 0.0001, "min_lr": 0.0001}, + }, + ) + ) + inference: Inference = field( + default_factory=lambda: Inference( + resources=Resources(gpu="H200:4"), + scaling=InferenceScaling( + min_replicas=0, max_replicas=8, target_concurrency=16 + ), + config={ + "tp_size": 4, + "ep_size": 4, + "mem_fraction_static": 0.9, + "max_running_requests": 32, + "max_queued_requests": 4, + "cpu_weight_cache_max_compile_group_gb": 32, + "dp_size": 4, + "enable_dp_attention": True, + }, + ) + ) diff --git a/src/lilo/configs/qwen35_4b_fft_64k.py b/src/lilo/configs/qwen35_4b_fft_64k.py new file mode 100644 index 0000000..44f9578 --- /dev/null +++ b/src/lilo/configs/qwen35_4b_fft_64k.py @@ -0,0 +1,68 @@ +from dataclasses import dataclass, field +from lilo.deployments import ( + BaseConfig, + Deployment, + EngineOptions, + Inference, + Model, + Resources, + Routing, + Trainer, +) + + +@dataclass(kw_only=True) +class Config(BaseConfig): + name: str = "qwen35-4b-fft-64k" + model: Model = field( + default_factory=lambda: Model( + id="Qwen/Qwen3.5-4B", + revision="main", + parameterization="full", + max_context_length=65536, + ) + ) + routing: Routing = field(default_factory=lambda: Routing(default=True)) + deployment: Deployment = field( + default_factory=lambda: Deployment(frontend="lilo-yaml") + ) + trainer: Trainer = field( + default_factory=lambda: Trainer( + backend="megatron", + resources=Resources(gpu="H100:4"), + engine=EngineOptions( + max_clients_per_instance=1, sampler_persistence_concurrency=1 + ), + config={ + "runtime": { + "tensor_model_parallel_size": 2, + "context_parallel_size": 2, + "sequence_parallel": True, + "micro_batch_size": 1, + "max_tokens_per_microbatch": 65536, + "defer_fp32_logits": True, + "fp32_lm_head": True, + "use_distributed_optimizer": True, + }, + "provider": { + "mtp_num_layers": 0, + "recompute_granularity": "full", + "recompute_method": "uniform", + "recompute_num_layers": 1, + }, + "optimizer": {"lr": 0.0001, "min_lr": 0.0001, "loss_scale": 1.0}, + }, + ) + ) + inference: Inference = field( + default_factory=lambda: Inference( + resources=Resources(gpu="H100:1"), + config={ + "tp_size": 1, + "mem_fraction_static": 0.85, + "max_running_requests": 32, + "max_queued_requests": 4, + "cpu_weight_cache_max_compile_group_gb": 16, + }, + ) + ) diff --git a/src/lilo/configs/qwen35_9b_fft_64k.py b/src/lilo/configs/qwen35_9b_fft_64k.py new file mode 100644 index 0000000..26e0af4 --- /dev/null +++ b/src/lilo/configs/qwen35_9b_fft_64k.py @@ -0,0 +1,70 @@ +from dataclasses import dataclass, field +from lilo.deployments import ( + BaseConfig, + EngineOptions, + Inference, + InferenceScaling, + Model, + Resources, + Routing, + Trainer, +) + + +@dataclass(kw_only=True) +class Config(BaseConfig): + name: str = "qwen35-9b-fft-64k" + model: Model = field( + default_factory=lambda: Model( + id="Qwen/Qwen3.5-9B", parameterization="full", max_context_length=65536 + ) + ) + routing: Routing = field(default_factory=lambda: Routing(default=True)) + trainer: Trainer = field( + default_factory=lambda: Trainer( + backend="megatron", + resources=Resources(gpu="H200:4"), + engine=EngineOptions( + max_clients_per_instance=1, sampler_persistence_concurrency=1 + ), + env={ + "PYTORCH_CUDA_ALLOC_CONF": "expandable_segments:True", + "TORCHINDUCTOR_COMPILE_THREADS": "1", + }, + config={ + "runtime": { + "tensor_model_parallel_size": 2, + "context_parallel_size": 2, + "sequence_parallel": True, + "micro_batch_size": 1, + "max_tokens_per_microbatch": 65536, + "defer_fp32_logits": True, + "fp32_lm_head": True, + "use_distributed_optimizer": True, + }, + "provider": { + "mtp_num_layers": 0, + "recompute_granularity": "full", + "recompute_method": "uniform", + "recompute_num_layers": 1, + }, + "optimizer": {"lr": 0.0001, "min_lr": 0.0001, "loss_scale": 1.0}, + }, + ) + ) + inference: Inference = field( + default_factory=lambda: Inference( + resources=Resources(gpu="H200:1"), + scaling=InferenceScaling( + min_replicas=0, max_replicas=8, target_concurrency=16 + ), + config={ + "tp_size": 1, + "ep_size": 1, + "mem_fraction_static": 0.85, + "max_running_requests": 32, + "max_queued_requests": 4, + "cpu_weight_cache_max_compile_group_gb": 16, + }, + ) + ) diff --git a/src/lilo/configs/qwen35_9b_instruct_lora_16k.py b/src/lilo/configs/qwen35_9b_instruct_lora_16k.py new file mode 100644 index 0000000..a156970 --- /dev/null +++ b/src/lilo/configs/qwen35_9b_instruct_lora_16k.py @@ -0,0 +1,81 @@ +from dataclasses import dataclass, field +from lilo.deployments import ( + BaseConfig, + EngineOptions, + Inference, + InferenceScaling, + Model, + Resources, + Routing, + Trainer, +) + + +@dataclass(kw_only=True) +class Config(BaseConfig): + name: str = "qwen35-9b-instruct-lora-16k" + model: Model = field( + default_factory=lambda: Model( + id="Qwen/Qwen3.5-9B", parameterization="lora", max_context_length=16384 + ) + ) + routing: Routing = field(default_factory=lambda: Routing(default=True)) + trainer: Trainer = field( + default_factory=lambda: Trainer( + backend="miles", + resources=Resources(gpu="H100:8"), + engine=EngineOptions( + max_clients_per_instance=6, sampler_persistence_concurrency=8 + ), + env={ + "PYTORCH_CUDA_ALLOC_CONF": "expandable_segments:True", + "TORCHINDUCTOR_COMPILE_THREADS": "1", + }, + config={ + "model_args": "qwen3.5-9B", + "options": { + "tensor_model_parallel_size": 8, + "target_modules": [ + "linear_qkv", + "linear_proj", + "linear_fc1", + "linear_fc2", + ], + "max_tokens_per_gpu": 16384, + "recompute_granularity": "full", + "recompute_method": "uniform", + "recompute_num_layers": 1, + "multi_lora_n_adapters": 6, + "lora_rank": 32, + "lora_alpha": 32, + }, + }, + ) + ) + inference: Inference = field( + default_factory=lambda: Inference( + resources=Resources(gpu="H200:1"), + scaling=InferenceScaling( + min_replicas=0, max_replicas=8, target_concurrency=16 + ), + config={ + "tp_size": 1, + "ep_size": 1, + "mem_fraction_static": 0.8, + "max_running_requests": 32, + "max_queued_requests": 8, + "max_loaded_loras": 256, + "max_loras_per_batch": 8, + "lora_target_modules": [ + "q_proj", + "k_proj", + "v_proj", + "o_proj", + "gate_proj", + "up_proj", + "down_proj", + ], + "schedule_policy": "lpm", + }, + ) + ) diff --git a/src/lilo/configs/qwen35_9b_instruct_lora_16k_dp2.py b/src/lilo/configs/qwen35_9b_instruct_lora_16k_dp2.py new file mode 100644 index 0000000..4c44452 --- /dev/null +++ b/src/lilo/configs/qwen35_9b_instruct_lora_16k_dp2.py @@ -0,0 +1,11 @@ +from dataclasses import dataclass +from lilo.configs.qwen35_9b_instruct_lora_16k import Config as ParentConfig + + +@dataclass(kw_only=True) +class Config(ParentConfig): + name: str = "qwen35-9b-instruct-lora-16k-dp2" + + def __post_init__(self): + self.routing.default = False + self.trainer.config["options"]["tensor_model_parallel_size"] = 4 diff --git a/src/lilo/configs/qwen35_9b_lora_16k.py b/src/lilo/configs/qwen35_9b_lora_16k.py new file mode 100644 index 0000000..0e25106 --- /dev/null +++ b/src/lilo/configs/qwen35_9b_lora_16k.py @@ -0,0 +1,81 @@ +from dataclasses import dataclass, field +from lilo.deployments import ( + BaseConfig, + Deployment, + EngineOptions, + Inference, + InferenceScaling, + Model, + Resources, + Routing, + Trainer, + TrainerScaling, +) + + +@dataclass(kw_only=True) +class Config(BaseConfig): + name: str = "qwen35-9b-lora-16k" + model: Model = field( + default_factory=lambda: Model( + id="Qwen/Qwen3.5-9B-Base", + revision="main", + parameterization="lora", + max_context_length=16384, + ) + ) + routing: Routing = field(default_factory=lambda: Routing(default=True)) + deployment: Deployment = field( + default_factory=lambda: Deployment(frontend="lilo-yaml", mode="shared") + ) + trainer: Trainer = field( + default_factory=lambda: Trainer( + backend="miles", + resources=Resources(gpu="H100:4", cpu=16, memory_mib=65536), + scaling=TrainerScaling(max_instances=1), + engine=EngineOptions( + max_clients_per_instance=6, sampler_persistence_concurrency=8 + ), + env={ + "PYTORCH_CUDA_ALLOC_CONF": "expandable_segments:True", + "TORCHINDUCTOR_COMPILE_THREADS": "1", + }, + config={ + "model_args": "qwen3.5-9B", + "options": { + "tensor_model_parallel_size": 4, + "multi_lora_n_adapters": 6, + "lora_rank": 32, + "lora_alpha": 32, + "target_modules": [ + "linear_qkv", + "linear_proj", + "linear_fc1", + "linear_fc2", + "output_layer", + ], + "max_tokens_per_gpu": 16384, + "recompute_granularity": "full", + "recompute_method": "uniform", + "recompute_num_layers": 1, + }, + }, + ) + ) + inference: Inference = field( + default_factory=lambda: Inference( + backend="sglang", + resources=Resources(gpu="H200:1"), + scaling=InferenceScaling( + min_replicas=0, max_replicas=8, target_concurrency=16 + ), + config={ + "tp_size": 1, + "mem_fraction_static": 0.8, + "max_running_requests": 32, + "max_queued_requests": 8, + "max_loaded_loras": 64, + "max_loras_per_batch": 8, + }, + ) + ) diff --git a/src/lilo/configs/qwen35_9b_lora_16k_single.py b/src/lilo/configs/qwen35_9b_lora_16k_single.py new file mode 100644 index 0000000..c60b684 --- /dev/null +++ b/src/lilo/configs/qwen35_9b_lora_16k_single.py @@ -0,0 +1,11 @@ +from dataclasses import dataclass +from lilo.configs.qwen35_9b_lora_16k import Config as ParentConfig + + +@dataclass(kw_only=True) +class Config(ParentConfig): + name: str = "qwen35-9b-lora-16k-single" + + def __post_init__(self): + self.routing.default = False + self.trainer.engine.max_clients_per_instance = 1 diff --git a/src/lilo/configs/qwen35_9b_lora_2k.py b/src/lilo/configs/qwen35_9b_lora_2k.py new file mode 100644 index 0000000..f2bc407 --- /dev/null +++ b/src/lilo/configs/qwen35_9b_lora_2k.py @@ -0,0 +1,83 @@ +from dataclasses import dataclass, field +from lilo.deployments import ( + BaseConfig, + EngineOptions, + Inference, + InferenceScaling, + Model, + Resources, + Routing, + Trainer, +) + + +@dataclass(kw_only=True) +class Config(BaseConfig): + name: str = "qwen35-9b-lora-2k" + model: Model = field( + default_factory=lambda: Model( + id="Qwen/Qwen3.5-9B-Base", parameterization="lora", max_context_length=2048 + ) + ) + routing: Routing = field(default_factory=lambda: Routing(default=False)) + trainer: Trainer = field( + default_factory=lambda: Trainer( + backend="miles", + resources=Resources(gpu="H200:4"), + engine=EngineOptions( + max_clients_per_instance=4, sampler_persistence_concurrency=8 + ), + env={ + "PYTORCH_CUDA_ALLOC_CONF": "expandable_segments:True", + "TORCHINDUCTOR_COMPILE_THREADS": "1", + }, + config={ + "model_args": "qwen3.5-9B", + "options": { + "tensor_model_parallel_size": 4, + "target_modules": [ + "linear_qkv", + "linear_proj", + "linear_fc1", + "linear_fc2", + "output_layer", + ], + "max_tokens_per_gpu": 2048, + "recompute_granularity": "full", + "recompute_method": "uniform", + "recompute_num_layers": 1, + "multi_lora_n_adapters": 4, + "lora_rank": 32, + "lora_alpha": 32, + }, + }, + ) + ) + inference: Inference = field( + default_factory=lambda: Inference( + resources=Resources(gpu="H200:1"), + scaling=InferenceScaling( + min_replicas=0, max_replicas=8, target_concurrency=16 + ), + config={ + "tp_size": 1, + "ep_size": 1, + "mem_fraction_static": 0.8, + "max_running_requests": 32, + "max_queued_requests": 8, + "max_loaded_loras": 32, + "max_loras_per_batch": 8, + "lora_target_modules": [ + "q_proj", + "k_proj", + "v_proj", + "o_proj", + "gate_proj", + "up_proj", + "down_proj", + "lm_head", + ], + "schedule_policy": "lpm", + }, + ) + ) diff --git a/src/lilo/configs/qwen35_9b_lora_64k.py b/src/lilo/configs/qwen35_9b_lora_64k.py new file mode 100644 index 0000000..077f58a --- /dev/null +++ b/src/lilo/configs/qwen35_9b_lora_64k.py @@ -0,0 +1,14 @@ +from dataclasses import dataclass +from lilo.configs.qwen35_9b_lora_16k import Config as ParentConfig + + +@dataclass(kw_only=True) +class Config(ParentConfig): + name: str = "qwen35-9b-lora-64k" + + def __post_init__(self): + self.model.max_context_length = 65536 + self.routing.default = False + self.trainer.resources.gpu = "H200:8" + self.trainer.config["options"]["tensor_model_parallel_size"] = 8 + self.trainer.config["options"]["max_tokens_per_gpu"] = 65536 diff --git a/src/lilo/configs/qwen36_27b_fft_64k.py b/src/lilo/configs/qwen36_27b_fft_64k.py new file mode 100644 index 0000000..857e69f --- /dev/null +++ b/src/lilo/configs/qwen36_27b_fft_64k.py @@ -0,0 +1,72 @@ +from dataclasses import dataclass, field +from lilo.deployments import ( + BaseConfig, + EngineOptions, + Inference, + InferenceScaling, + Model, + Resources, + Routing, + Trainer, +) + + +@dataclass(kw_only=True) +class Config(BaseConfig): + name: str = "qwen36-27b-fft-64k" + model: Model = field( + default_factory=lambda: Model( + id="Qwen/Qwen3.6-27B", parameterization="full", max_context_length=65536 + ) + ) + routing: Routing = field(default_factory=lambda: Routing(default=True)) + trainer: Trainer = field( + default_factory=lambda: Trainer( + backend="megatron", + resources=Resources(gpu="H200:8"), + engine=EngineOptions( + max_clients_per_instance=1, sampler_persistence_concurrency=1 + ), + env={ + "PYTORCH_CUDA_ALLOC_CONF": "expandable_segments:True", + "TORCHINDUCTOR_COMPILE_THREADS": "1", + }, + config={ + "runtime": { + "tensor_model_parallel_size": 4, + "pipeline_model_parallel_size": 1, + "context_parallel_size": 2, + "sequence_parallel": True, + "micro_batch_size": 1, + "max_tokens_per_microbatch": 65536, + "bf16": True, + "fp16": False, + "gpu_memory_fraction": 0.9, + "use_distributed_optimizer": True, + }, + "provider": { + "mtp_num_layers": 0, + "recompute_granularity": "full", + "recompute_method": "uniform", + "recompute_num_layers": 1, + }, + "optimizer": {"optimizer": "adam", "lr": 0.0001, "min_lr": 0.0001}, + }, + ) + ) + inference: Inference = field( + default_factory=lambda: Inference( + resources=Resources(gpu="H200:4"), + scaling=InferenceScaling( + min_replicas=0, max_replicas=8, target_concurrency=16 + ), + config={ + "tp_size": 4, + "ep_size": 1, + "mem_fraction_static": 0.9, + "max_running_requests": 32, + "max_queued_requests": 4, + "cpu_weight_cache_max_compile_group_gb": 32, + }, + ) + ) diff --git a/src/lilo/configs/qwen36_35b_a3b_fft_64k.py b/src/lilo/configs/qwen36_35b_a3b_fft_64k.py new file mode 100644 index 0000000..0c157bd --- /dev/null +++ b/src/lilo/configs/qwen36_35b_a3b_fft_64k.py @@ -0,0 +1,12 @@ +from dataclasses import dataclass +from lilo.configs.qwen35_35b_a3b_fft_64k import Config as ParentConfig + + +@dataclass(kw_only=True) +class Config(ParentConfig): + name: str = "qwen36-35b-a3b-fft-64k" + + def __post_init__(self): + self.model.id = "Qwen/Qwen3.6-35B-A3B" + self.inference.config["dp_size"] = 1 + self.inference.config["enable_dp_attention"] = False diff --git a/src/lilo/configs/qwen38_27b_lora_128k.py b/src/lilo/configs/qwen38_27b_lora_128k.py new file mode 100644 index 0000000..889b7bf --- /dev/null +++ b/src/lilo/configs/qwen38_27b_lora_128k.py @@ -0,0 +1,19 @@ +from dataclasses import dataclass +from lilo.configs.qwen38_27b_lora_16k import Config as ParentConfig + + +@dataclass(kw_only=True) +class Config(ParentConfig): + name: str = "qwen38-27b-lora-128k" + + def __post_init__(self): + self.model.max_context_length = 131072 + self.routing.default = False + self.trainer.config["options"]["tensor_model_parallel_size"] = 2 + self.trainer.config["options"]["context_parallel_size"] = 4 + self.trainer.config["options"]["max_tokens_per_gpu"] = 32768 + self.inference.resources.gpu = "H200:2" + self.inference.scaling.max_replicas = 4 + self.inference.scaling.target_concurrency = 4 + self.inference.config["tp_size"] = 2 + self.inference.config["max_running_requests"] = 8 diff --git a/src/lilo/configs/qwen38_27b_lora_16k.py b/src/lilo/configs/qwen38_27b_lora_16k.py new file mode 100644 index 0000000..237f2b1 --- /dev/null +++ b/src/lilo/configs/qwen38_27b_lora_16k.py @@ -0,0 +1,81 @@ +from dataclasses import dataclass, field +from lilo.deployments import ( + BaseConfig, + EngineOptions, + Inference, + InferenceScaling, + Model, + Resources, + Routing, + Trainer, +) + + +@dataclass(kw_only=True) +class Config(BaseConfig): + name: str = "qwen38-27b-lora-16k" + model: Model = field( + default_factory=lambda: Model( + id="Qwen/Qwen3.8-27B", parameterization="lora", max_context_length=16384 + ) + ) + routing: Routing = field(default_factory=lambda: Routing(default=True)) + trainer: Trainer = field( + default_factory=lambda: Trainer( + backend="miles", + resources=Resources(gpu="H200:8"), + engine=EngineOptions( + max_clients_per_instance=6, sampler_persistence_concurrency=8 + ), + env={ + "PYTORCH_CUDA_ALLOC_CONF": "expandable_segments:True", + "TORCHINDUCTOR_COMPILE_THREADS": "1", + }, + config={ + "model_args": "qwen3.8-27B", + "options": { + "tensor_model_parallel_size": 4, + "target_modules": [ + "linear_qkv", + "linear_proj", + "linear_fc1", + "linear_fc2", + ], + "max_tokens_per_gpu": 16384, + "recompute_granularity": "full", + "recompute_method": "uniform", + "recompute_num_layers": 1, + "multi_lora_n_adapters": 6, + "lora_rank": 32, + "lora_alpha": 32, + }, + }, + ) + ) + inference: Inference = field( + default_factory=lambda: Inference( + resources=Resources(gpu="H200:1"), + scaling=InferenceScaling( + min_replicas=0, max_replicas=8, target_concurrency=16 + ), + config={ + "tp_size": 1, + "ep_size": 1, + "mem_fraction_static": 0.8, + "max_running_requests": 32, + "max_queued_requests": 8, + "max_loaded_loras": 256, + "max_loras_per_batch": 8, + "lora_target_modules": [ + "q_proj", + "k_proj", + "v_proj", + "o_proj", + "gate_proj", + "up_proj", + "down_proj", + ], + "schedule_policy": "lpm", + }, + ) + ) diff --git a/src/lilo/configs/qwen38_27b_lora_64k.py b/src/lilo/configs/qwen38_27b_lora_64k.py new file mode 100644 index 0000000..d50e7a5 --- /dev/null +++ b/src/lilo/configs/qwen38_27b_lora_64k.py @@ -0,0 +1,15 @@ +from dataclasses import dataclass +from lilo.configs.qwen38_27b_lora_16k import Config as ParentConfig + + +@dataclass(kw_only=True) +class Config(ParentConfig): + name: str = "qwen38-27b-lora-64k" + + def __post_init__(self): + self.model.max_context_length = 65536 + self.routing.default = False + self.trainer.config["options"]["context_parallel_size"] = 2 + self.trainer.config["options"]["max_tokens_per_gpu"] = 32768 + self.inference.scaling.target_concurrency = 8 + self.inference.config["max_running_requests"] = 16 diff --git a/src/lilo/deployment_cli.py b/src/lilo/deployment_cli.py index 7b81d31..2898d4d 100644 --- a/src/lilo/deployment_cli.py +++ b/src/lilo/deployment_cli.py @@ -1,4 +1,4 @@ -"""Operator CLI for YAML deployments. Only `deploy` provisions Modal resources.""" +"""Operator CLI for Python deployments. Only `deploy` provisions Modal resources.""" from __future__ import annotations @@ -12,12 +12,11 @@ import sys import uuid -import yaml from lilo.deployments import ( DeploymentRecord, load, - preset_path, + config_path, validate_frontend, ) @@ -71,7 +70,7 @@ def retain_generations(previous, desired): for row in previous: if row.implementation != expected.implementation: raise ValueError( - "This draft cannot rebuild retained generations with different Lilo/runtime code. Use a separate frontend for a code upgrade; YAML-only changes can retain existing generations." + "This draft cannot rebuild retained generations with different Lilo/runtime code. Use a separate frontend for a code upgrade; Config-only changes can retain existing generations." ) if ( row.spec.deployment != expected.spec.deployment @@ -94,11 +93,11 @@ def retain_generations(previous, desired): def deploy(desired): """Serialize operator applies and retain interrupted attempts for safe recovery.""" import modal - from lilo.providers.modal.yaml_apps import MANIFEST_ENV + from lilo.providers.modal.deployment_apps import MANIFEST_ENV if sys.version_info[:2] != (3, 12): raise ValueError( - "YAML deployment requires Python 3.12 to match the serialized GPU runtime images" + "Python deployment requires Python 3.12 to match the serialized GPU runtime images" ) settings = desired[0].spec.deployment registry = modal.Dict.from_name( @@ -122,7 +121,7 @@ def deploy(desired): pass else: raise ValueError( - "The frontend already exists without a YAML registry. Choose a new frontend name; an app with no deployment registry cannot be safely updated." + "The frontend already exists without a deployment registry. Choose a new frontend name; an app with no deployment registry cannot be safely updated." ) # A killed deploy may already have updated Modal. Keep its functions on retry. rows = {r["generation"]: r for r in [*rows, *registry.get("pending", [])]} @@ -168,7 +167,7 @@ def parser(): if name == "resolve": cmd.add_argument("--output") apply = commands.add_parser( - "deploy", help="Deploy the complete active YAML set behind one frontend" + "deploy", help="Deploy the complete active Python config set behind one frontend" ) apply.add_argument("files", nargs="+") management = commands.add_parser("deployment").add_subparsers( @@ -189,12 +188,16 @@ def main(argv=None): try: if args.command == "config": if args.action == "init": + module = config_path(args.preset).stem + if not config_path(args.preset).is_file(): + raise ValueError(f"unknown example config: {args.preset}") print( - yaml.safe_dump( - load(preset_path(args.preset)).model_dump(mode="json"), - sort_keys=False, - ), - end="", + "from dataclasses import dataclass\n" + f"from lilo.configs.{module} import Config as ParentConfig\n\n\n" + "@dataclass(kw_only=True)\n" + "class Config(ParentConfig):\n" + " # Override fields or customize nested settings in __post_init__.\n" + " pass" ) elif args.action == "validate": specs = [load(path) for path in args.files] diff --git a/src/lilo/deployments.py b/src/lilo/deployments.py index 08b11ac..ae6f998 100644 --- a/src/lilo/deployments.py +++ b/src/lilo/deployments.py @@ -1,39 +1,41 @@ -"""Deployment specifications. Loading YAML never contacts Modal or allocates GPUs.""" +"""Python deployment configs and the records saved by the deployment CLI.""" from __future__ import annotations +from copy import deepcopy +from dataclasses import asdict, dataclass, field import hashlib -import json -import re from importlib.resources import files +import json from pathlib import Path +import re +import runpy +import sys from typing import Any, Literal -import yaml -from pydantic import BaseModel, ConfigDict, Field, model_validator - - -class StrictModel(BaseModel): - model_config = ConfigDict(extra="forbid") +from pydantic import BaseModel, ConfigDict -class Model(StrictModel): - id: str = Field(min_length=1) +@dataclass(kw_only=True) +class Model: + id: str + max_context_length: int revision: str = "main" parameterization: Literal["lora", "full"] = "lora" - max_context_length: int = Field(gt=0) -class Routing(StrictModel): +@dataclass(kw_only=True) +class Routing: default: bool = False sampling_default: bool = False -class Resources(StrictModel): +@dataclass(kw_only=True) +class Resources: gpu: str - cpu: float = Field(default=8, gt=0) - memory_mib: int = Field(default=32768, gt=0) - timeout_s: int = Field(default=86400, gt=0, le=86400) + cpu: float = 8 + memory_mib: int = 32768 + timeout_s: int = 86400 @property def gpu_count(self) -> int: @@ -42,99 +44,110 @@ def gpu_count(self) -> int: return int(self.gpu.split(":")[1]) if ":" in self.gpu else 1 -class TrainerScaling(StrictModel): +@dataclass(kw_only=True) +class TrainerScaling: min_instances: Literal[0] = 0 - max_instances: int = Field(default=1, gt=0) + max_instances: int = 1 -class InferenceScaling(StrictModel): - min_replicas: int = Field(default=0, ge=0) - max_replicas: int = Field(default=8, gt=0) - target_concurrency: int = Field(default=16, gt=0) - scaledown_window_s: int = Field(default=300, gt=0) +@dataclass(kw_only=True) +class InferenceScaling: + min_replicas: int = 0 + max_replicas: int = 8 + target_concurrency: int = 16 + scaledown_window_s: int = 300 - @model_validator(mode="after") - def ordered(self): + def __post_init__(self): if self.min_replicas > self.max_replicas: raise ValueError("min_replicas must not exceed max_replicas") - return self -class EngineOptions(StrictModel): - max_clients_per_instance: int = Field(default=1, gt=0) - sampler_persistence_concurrency: int = Field(default=8, gt=0) +@dataclass(kw_only=True) +class EngineOptions: + max_clients_per_instance: int = 1 + sampler_persistence_concurrency: int = 8 -class Trainer(StrictModel): - backend: str = "miles" +@dataclass(kw_only=True) +class Trainer: resources: Resources - scaling: TrainerScaling = Field(default_factory=TrainerScaling) - engine: EngineOptions = Field(default_factory=EngineOptions) - config: dict[str, Any] = Field(default_factory=dict) - env: dict[str, str] = Field(default_factory=dict) + backend: str = "miles" + scaling: TrainerScaling = field(default_factory=TrainerScaling) + engine: EngineOptions = field(default_factory=EngineOptions) + config: dict[str, Any] = field(default_factory=dict) + env: dict[str, str] = field(default_factory=dict) -class Inference(StrictModel): - backend: str = "sglang" +@dataclass(kw_only=True) +class Inference: resources: Resources - scaling: InferenceScaling = Field(default_factory=InferenceScaling) - config: dict[str, Any] = Field(default_factory=dict) - env: dict[str, str] = Field(default_factory=dict) + backend: str = "sglang" + scaling: InferenceScaling = field(default_factory=InferenceScaling) + config: dict[str, Any] = field(default_factory=dict) + env: dict[str, str] = field(default_factory=dict) -class Secrets(StrictModel): +@dataclass(kw_only=True) +class Secrets: api: str = "lilo-api" sampler_proxy: str = "lilo-proxy" huggingface: str | None = "huggingface-secret" -class Storage(StrictModel): +@dataclass(kw_only=True) +class Storage: assets: str = "lilo-model-assets" checkpoints: str = "lilo-checkpoints" bulletin: str = "lilo-snapshot-bulletin" -class ModalSettings(StrictModel): +@dataclass(kw_only=True) +class ModalSettings: environment: str | None = None region: str = "us-west" -class Deployment(StrictModel): - frontend: str = Field( - default="lilo-yaml", pattern=r"^[a-zA-Z0-9][a-zA-Z0-9_-]{0,46}$" - ) +@dataclass(kw_only=True) +class Deployment: + frontend: str = "lilo-yaml" mode: Literal["shared"] = "shared" - modal: ModalSettings = Field(default_factory=ModalSettings) - secrets: Secrets = Field(default_factory=Secrets) - storage: Storage = Field(default_factory=Storage) + modal: ModalSettings = field(default_factory=ModalSettings) + secrets: Secrets = field(default_factory=Secrets) + storage: Storage = field(default_factory=Storage) -class Lifecycle(StrictModel): - session_idle_timeout_s: int = Field(default=300, gt=0) - pool_idle_timeout_s: int = Field(default=300, gt=0) - sweep_interval_s: int = Field(default=300, gt=0) +@dataclass(kw_only=True) +class Lifecycle: + session_idle_timeout_s: int = 300 + pool_idle_timeout_s: int = 300 + sweep_interval_s: int = 300 -class DeploymentSpec(StrictModel): - api_version: Literal["lilo/v1"] = "lilo/v1" - name: str = Field(pattern=r"^[a-zA-Z0-9][a-zA-Z0-9_-]{0,63}$") +@dataclass(kw_only=True) +class BaseConfig: + """Subclass in a config file and override defaults or use __post_init__.""" + + name: str model: Model - routing: Routing = Field(default_factory=Routing) - deployment: Deployment = Field(default_factory=Deployment) trainer: Trainer inference: Inference - lifecycle: Lifecycle = Field(default_factory=Lifecycle) + api_version: Literal["lilo/v1"] = "lilo/v1" + routing: Routing = field(default_factory=Routing) + deployment: Deployment = field(default_factory=Deployment) + lifecycle: Lifecycle = field(default_factory=Lifecycle) -class DeploymentRecord(StrictModel): - """Saved deployment metadata around a YAML specification. +class DeploymentRecord(BaseModel): + """Saved deployment metadata around a Python configuration. The CLI resolves the model revision before creating this record. Its hash binds jobs to their original code and configuration across later deploys; active controls whether new clients can select it. """ - spec: DeploymentSpec + model_config = ConfigDict(extra="forbid") + + spec: BaseConfig implementation: str miles_commit: str | None = None generation: str @@ -143,17 +156,18 @@ class DeploymentRecord(StrictModel): @classmethod def create( cls, - spec: DeploymentSpec, + spec: BaseConfig, *, revision: str, implementation: str, miles_commit: str | None = None, ) -> DeploymentRecord: - """Record an already-resolved revision without reparsing the YAML.""" - pinned = spec.model_copy(deep=True) + """Record an already-resolved revision without reparsing the configuration.""" + pinned = deepcopy(spec) pinned.model.revision = revision # Changing routing defaults should not restart an existing trainer. - identity = pinned.model_dump(exclude={"routing"}) + identity = asdict(pinned) + identity.pop("routing") generation = hashlib.sha256( json.dumps([implementation, identity], sort_keys=True).encode() ).hexdigest() @@ -176,72 +190,34 @@ def asset_path(self) -> str: return f"/assets/{digest}" -class UniqueLoader(yaml.SafeLoader): - pass - - -def _mapping(loader, node): - result = {} - for key_node, value_node in node.value: - key = loader.construct_object(key_node) - if not isinstance(key, str) or key in result: - raise ValueError(f"duplicate or non-string YAML key: {key!r}") - result[key] = loader.construct_object(value_node) - return result - - -UniqueLoader.add_constructor(yaml.resolver.BaseResolver.DEFAULT_MAPPING_TAG, _mapping) - - -def merge(parent: dict, child: dict) -> dict: - return { - key: merge(parent[key], value) - if isinstance(parent.get(key), dict) and isinstance(value, dict) - else value - for key, value in (parent | child).items() - } - - -def preset_path(name: str) -> Path: +def config_path(name: str) -> Path: + """Locate an installed example config without maintaining a model catalog.""" if not re.fullmatch(r"[a-zA-Z0-9_-]+", name): - raise ValueError("invalid preset name") - return Path(str(files("lilo").joinpath("presets", name + ".yaml"))) - + raise ValueError("invalid config name") + return Path(str(files("lilo").joinpath("configs", name.replace("-", "_") + ".py"))) -def load(path: str | Path) -> DeploymentSpec: - """Read and merge YAML inheritance, then construct the deployment fields. - Backend configuration is interpreted when preparing its trainer or pool. - Parent files may be partial; only the fully merged document is constructed. - """ +def load(path: str | Path) -> BaseConfig: + """Execute a Python config file and instantiate its exported Config class.""" path = Path(path).resolve() - seen = set() - documents = [] - while True: - if path in seen: - raise ValueError(f"cyclic extends: {path}") - seen.add(path) - data = yaml.load(path.read_text(), Loader=UniqueLoader) - if not isinstance(data, dict): - raise ValueError("deployment YAML must be a mapping") - parent = data.pop("extends", None) - documents.append(data) - if parent is None: - break - if not isinstance(parent, str): - raise ValueError("extends must be a path or builtin:preset") - path = ( - preset_path(parent[8:]) - if parent.startswith("builtin:") - else path.parent / parent - ).resolve() - merged = {} - for document in reversed(documents): - merged = merge(merged, document) - return DeploymentSpec(**merged) - - -def validate_frontend(specs: list[DeploymentSpec]) -> None: + if path.suffix != ".py": + raise ValueError("deployment configs must be Python .py files") + # Let a config import sibling modules using normal Python imports. + original_path = sys.path[:] + sys.path.insert(0, str(path.parent)) + try: + namespace = runpy.run_path(str(path)) + config_class = namespace.get("Config") + if not isinstance(config_class, type) or not issubclass( + config_class, BaseConfig + ): + raise ValueError(f"{path} must export a Config subclass of BaseConfig") + return config_class() + finally: + sys.path[:] = original_path + + +def validate_frontend(specs: list[BaseConfig]) -> None: if not specs: raise ValueError("at least one deployment is required") if len({s.name for s in specs}) != len(specs): diff --git a/src/lilo/native_options.py b/src/lilo/native_options.py index e220a71..b29c5d0 100644 --- a/src/lilo/native_options.py +++ b/src/lilo/native_options.py @@ -1,4 +1,4 @@ -"""Apply typed YAML values through a backend's own argparse schema.""" +"""Apply typed config values through a backend's own argparse schema.""" from __future__ import annotations diff --git a/src/lilo/presets/qwen35-35b-a3b-fft-64k.yaml b/src/lilo/presets/qwen35-35b-a3b-fft-64k.yaml deleted file mode 100644 index 5b64a25..0000000 --- a/src/lilo/presets/qwen35-35b-a3b-fft-64k.yaml +++ /dev/null @@ -1,62 +0,0 @@ -api_version: lilo/v1 -name: qwen35-35b-a3b-fft-64k -model: - id: Qwen/Qwen3.5-35B-A3B - parameterization: full - max_context_length: 65536 -routing: - default: true -trainer: - backend: megatron - resources: - gpu: H200:8 - engine: - max_clients_per_instance: 1 - sampler_persistence_concurrency: 1 - env: - PYTORCH_CUDA_ALLOC_CONF: expandable_segments:True - TORCHINDUCTOR_COMPILE_THREADS: '1' - config: - runtime: - tensor_model_parallel_size: 4 - pipeline_model_parallel_size: 1 - context_parallel_size: 2 - expert_model_parallel_size: 8 - expert_tensor_parallel_size: 1 - sequence_parallel: true - micro_batch_size: 1 - max_tokens_per_microbatch: 65536 - bf16: true - fp16: false - gpu_memory_fraction: 0.9 - use_distributed_optimizer: true - provider: - mtp_num_layers: 0 - recompute_granularity: selective - moe_layer_recompute: true - moe_token_dispatcher_type: alltoall - moe_router_fusion: true - moe_permute_fusion: true - moe_grouped_gemm: true - moe_shared_expert_overlap: false - moe_aux_loss_coeff: 0.0 - optimizer: - optimizer: adam - lr: 0.0001 - min_lr: 0.0001 -inference: - resources: - gpu: H200:4 - scaling: - min_replicas: 0 - max_replicas: 8 - target_concurrency: 16 - config: - tp_size: 4 - ep_size: 4 - mem_fraction_static: 0.9 - max_running_requests: 32 - max_queued_requests: 4 - cpu_weight_cache_max_compile_group_gb: 32 - dp_size: 4 - enable_dp_attention: true diff --git a/src/lilo/presets/qwen35-4b-fft-64k.yaml b/src/lilo/presets/qwen35-4b-fft-64k.yaml deleted file mode 100644 index 2de6f4e..0000000 --- a/src/lilo/presets/qwen35-4b-fft-64k.yaml +++ /dev/null @@ -1,46 +0,0 @@ -api_version: lilo/v1 -name: qwen35-4b-fft-64k -model: - id: Qwen/Qwen3.5-4B - revision: main - parameterization: full - max_context_length: 65536 -routing: - default: true -deployment: - frontend: lilo-yaml -trainer: - backend: megatron - resources: - gpu: H100:4 - engine: - max_clients_per_instance: 1 - sampler_persistence_concurrency: 1 - config: - runtime: - tensor_model_parallel_size: 2 - context_parallel_size: 2 - sequence_parallel: true - micro_batch_size: 1 - max_tokens_per_microbatch: 65536 - defer_fp32_logits: true - fp32_lm_head: true - use_distributed_optimizer: true - provider: - mtp_num_layers: 0 - recompute_granularity: full - recompute_method: uniform - recompute_num_layers: 1 - optimizer: - lr: 0.0001 - min_lr: 0.0001 - loss_scale: 1.0 -inference: - resources: - gpu: H100:1 - config: - tp_size: 1 - mem_fraction_static: 0.85 - max_running_requests: 32 - max_queued_requests: 4 - cpu_weight_cache_max_compile_group_gb: 16 diff --git a/src/lilo/presets/qwen35-9b-fft-64k.yaml b/src/lilo/presets/qwen35-9b-fft-64k.yaml deleted file mode 100644 index 79ae97c..0000000 --- a/src/lilo/presets/qwen35-9b-fft-64k.yaml +++ /dev/null @@ -1,51 +0,0 @@ -api_version: lilo/v1 -name: qwen35-9b-fft-64k -model: - id: Qwen/Qwen3.5-9B - parameterization: full - max_context_length: 65536 -routing: - default: true -trainer: - backend: megatron - resources: - gpu: H200:4 - engine: - max_clients_per_instance: 1 - sampler_persistence_concurrency: 1 - env: - PYTORCH_CUDA_ALLOC_CONF: expandable_segments:True - TORCHINDUCTOR_COMPILE_THREADS: '1' - config: - runtime: - tensor_model_parallel_size: 2 - context_parallel_size: 2 - sequence_parallel: true - micro_batch_size: 1 - max_tokens_per_microbatch: 65536 - defer_fp32_logits: true - fp32_lm_head: true - use_distributed_optimizer: true - provider: - mtp_num_layers: 0 - recompute_granularity: full - recompute_method: uniform - recompute_num_layers: 1 - optimizer: - lr: 0.0001 - min_lr: 0.0001 - loss_scale: 1.0 -inference: - resources: - gpu: H200:1 - scaling: - min_replicas: 0 - max_replicas: 8 - target_concurrency: 16 - config: - tp_size: 1 - ep_size: 1 - mem_fraction_static: 0.85 - max_running_requests: 32 - max_queued_requests: 4 - cpu_weight_cache_max_compile_group_gb: 16 diff --git a/src/lilo/presets/qwen35-9b-instruct-lora-16k-dp2.yaml b/src/lilo/presets/qwen35-9b-instruct-lora-16k-dp2.yaml deleted file mode 100644 index 4c2d8f0..0000000 --- a/src/lilo/presets/qwen35-9b-instruct-lora-16k-dp2.yaml +++ /dev/null @@ -1,8 +0,0 @@ -extends: builtin:qwen35-9b-instruct-lora-16k -name: qwen35-9b-instruct-lora-16k-dp2 -routing: - default: false -trainer: - config: - options: - tensor_model_parallel_size: 4 diff --git a/src/lilo/presets/qwen35-9b-instruct-lora-16k.yaml b/src/lilo/presets/qwen35-9b-instruct-lora-16k.yaml deleted file mode 100644 index dee8021..0000000 --- a/src/lilo/presets/qwen35-9b-instruct-lora-16k.yaml +++ /dev/null @@ -1,58 +0,0 @@ -api_version: lilo/v1 -name: qwen35-9b-instruct-lora-16k -model: - id: Qwen/Qwen3.5-9B - parameterization: lora - max_context_length: 16384 -routing: - default: true -trainer: - backend: miles - resources: - gpu: H100:8 - engine: - max_clients_per_instance: 6 - sampler_persistence_concurrency: 8 - env: - PYTORCH_CUDA_ALLOC_CONF: expandable_segments:True - TORCHINDUCTOR_COMPILE_THREADS: '1' - config: - model_args: qwen3.5-9B - options: - tensor_model_parallel_size: 8 - target_modules: - - linear_qkv - - linear_proj - - linear_fc1 - - linear_fc2 - max_tokens_per_gpu: 16384 - recompute_granularity: full - recompute_method: uniform - recompute_num_layers: 1 - multi_lora_n_adapters: 6 - lora_rank: 32 - lora_alpha: 32 -inference: - resources: - gpu: H200:1 - scaling: - min_replicas: 0 - max_replicas: 8 - target_concurrency: 16 - config: - tp_size: 1 - ep_size: 1 - mem_fraction_static: 0.8 - max_running_requests: 32 - max_queued_requests: 8 - max_loaded_loras: 256 - max_loras_per_batch: 8 - lora_target_modules: - - q_proj - - k_proj - - v_proj - - o_proj - - gate_proj - - up_proj - - down_proj - schedule_policy: lpm diff --git a/src/lilo/presets/qwen35-9b-lora-16k-single.yaml b/src/lilo/presets/qwen35-9b-lora-16k-single.yaml deleted file mode 100644 index 0dd5c24..0000000 --- a/src/lilo/presets/qwen35-9b-lora-16k-single.yaml +++ /dev/null @@ -1,7 +0,0 @@ -extends: builtin:qwen35-9b-lora-16k -name: qwen35-9b-lora-16k-single -routing: - default: false -trainer: - engine: - max_clients_per_instance: 1 diff --git a/src/lilo/presets/qwen35-9b-lora-16k.yaml b/src/lilo/presets/qwen35-9b-lora-16k.yaml deleted file mode 100644 index 8c1df86..0000000 --- a/src/lilo/presets/qwen35-9b-lora-16k.yaml +++ /dev/null @@ -1,58 +0,0 @@ -api_version: lilo/v1 -name: qwen35-9b-lora-16k -model: - id: Qwen/Qwen3.5-9B-Base - revision: main - parameterization: lora - max_context_length: 16384 -routing: - default: true -deployment: - frontend: lilo-yaml - mode: shared -trainer: - backend: miles - resources: - gpu: H100:4 - cpu: 16 - memory_mib: 65536 - scaling: - max_instances: 1 - engine: - max_clients_per_instance: 6 - sampler_persistence_concurrency: 8 - env: - PYTORCH_CUDA_ALLOC_CONF: expandable_segments:True - TORCHINDUCTOR_COMPILE_THREADS: '1' - config: - model_args: qwen3.5-9B - options: - tensor_model_parallel_size: 4 - multi_lora_n_adapters: 6 - lora_rank: 32 - lora_alpha: 32 - target_modules: - - linear_qkv - - linear_proj - - linear_fc1 - - linear_fc2 - - output_layer - max_tokens_per_gpu: 16384 - recompute_granularity: full - recompute_method: uniform - recompute_num_layers: 1 -inference: - backend: sglang - resources: - gpu: H200:1 - scaling: - min_replicas: 0 - max_replicas: 8 - target_concurrency: 16 - config: - tp_size: 1 - mem_fraction_static: 0.8 - max_running_requests: 32 - max_queued_requests: 8 - max_loaded_loras: 64 - max_loras_per_batch: 8 diff --git a/src/lilo/presets/qwen35-9b-lora-2k.yaml b/src/lilo/presets/qwen35-9b-lora-2k.yaml deleted file mode 100644 index 23f91e4..0000000 --- a/src/lilo/presets/qwen35-9b-lora-2k.yaml +++ /dev/null @@ -1,60 +0,0 @@ -api_version: lilo/v1 -name: qwen35-9b-lora-2k -model: - id: Qwen/Qwen3.5-9B-Base - parameterization: lora - max_context_length: 2048 -routing: - default: false -trainer: - backend: miles - resources: - gpu: H200:4 - engine: - max_clients_per_instance: 4 - sampler_persistence_concurrency: 8 - env: - PYTORCH_CUDA_ALLOC_CONF: expandable_segments:True - TORCHINDUCTOR_COMPILE_THREADS: '1' - config: - model_args: qwen3.5-9B - options: - tensor_model_parallel_size: 4 - target_modules: - - linear_qkv - - linear_proj - - linear_fc1 - - linear_fc2 - - output_layer - max_tokens_per_gpu: 2048 - recompute_granularity: full - recompute_method: uniform - recompute_num_layers: 1 - multi_lora_n_adapters: 4 - lora_rank: 32 - lora_alpha: 32 -inference: - resources: - gpu: H200:1 - scaling: - min_replicas: 0 - max_replicas: 8 - target_concurrency: 16 - config: - tp_size: 1 - ep_size: 1 - mem_fraction_static: 0.8 - max_running_requests: 32 - max_queued_requests: 8 - max_loaded_loras: 32 - max_loras_per_batch: 8 - lora_target_modules: - - q_proj - - k_proj - - v_proj - - o_proj - - gate_proj - - up_proj - - down_proj - - lm_head - schedule_policy: lpm diff --git a/src/lilo/presets/qwen35-9b-lora-64k.yaml b/src/lilo/presets/qwen35-9b-lora-64k.yaml deleted file mode 100644 index 0f6e9f8..0000000 --- a/src/lilo/presets/qwen35-9b-lora-64k.yaml +++ /dev/null @@ -1,13 +0,0 @@ -extends: builtin:qwen35-9b-lora-16k -name: qwen35-9b-lora-64k -model: - max_context_length: 65536 -routing: - default: false -trainer: - resources: - gpu: H200:8 - config: - options: - tensor_model_parallel_size: 8 - max_tokens_per_gpu: 65536 diff --git a/src/lilo/presets/qwen36-27b-fft-64k.yaml b/src/lilo/presets/qwen36-27b-fft-64k.yaml deleted file mode 100644 index 9dd7446..0000000 --- a/src/lilo/presets/qwen36-27b-fft-64k.yaml +++ /dev/null @@ -1,53 +0,0 @@ -api_version: lilo/v1 -name: qwen36-27b-fft-64k -model: - id: Qwen/Qwen3.6-27B - parameterization: full - max_context_length: 65536 -routing: - default: true -trainer: - backend: megatron - resources: - gpu: H200:8 - engine: - max_clients_per_instance: 1 - sampler_persistence_concurrency: 1 - env: - PYTORCH_CUDA_ALLOC_CONF: expandable_segments:True - TORCHINDUCTOR_COMPILE_THREADS: '1' - config: - runtime: - tensor_model_parallel_size: 4 - pipeline_model_parallel_size: 1 - context_parallel_size: 2 - sequence_parallel: true - micro_batch_size: 1 - max_tokens_per_microbatch: 65536 - bf16: true - fp16: false - gpu_memory_fraction: 0.9 - use_distributed_optimizer: true - provider: - mtp_num_layers: 0 - recompute_granularity: full - recompute_method: uniform - recompute_num_layers: 1 - optimizer: - optimizer: adam - lr: 0.0001 - min_lr: 0.0001 -inference: - resources: - gpu: H200:4 - scaling: - min_replicas: 0 - max_replicas: 8 - target_concurrency: 16 - config: - tp_size: 4 - ep_size: 1 - mem_fraction_static: 0.9 - max_running_requests: 32 - max_queued_requests: 4 - cpu_weight_cache_max_compile_group_gb: 32 diff --git a/src/lilo/presets/qwen36-35b-a3b-fft-64k.yaml b/src/lilo/presets/qwen36-35b-a3b-fft-64k.yaml deleted file mode 100644 index 839bf2b..0000000 --- a/src/lilo/presets/qwen36-35b-a3b-fft-64k.yaml +++ /dev/null @@ -1,8 +0,0 @@ -extends: builtin:qwen35-35b-a3b-fft-64k -name: qwen36-35b-a3b-fft-64k -model: - id: Qwen/Qwen3.6-35B-A3B -inference: - config: - dp_size: 1 - enable_dp_attention: false diff --git a/src/lilo/presets/qwen38-27b-lora-128k.yaml b/src/lilo/presets/qwen38-27b-lora-128k.yaml deleted file mode 100644 index aae6b67..0000000 --- a/src/lilo/presets/qwen38-27b-lora-128k.yaml +++ /dev/null @@ -1,21 +0,0 @@ -extends: builtin:qwen38-27b-lora-16k -name: qwen38-27b-lora-128k -model: - max_context_length: 131072 -routing: - default: false -trainer: - config: - options: - tensor_model_parallel_size: 2 - context_parallel_size: 4 - max_tokens_per_gpu: 32768 -inference: - resources: - gpu: H200:2 - scaling: - max_replicas: 4 - target_concurrency: 4 - config: - tp_size: 2 - max_running_requests: 8 diff --git a/src/lilo/presets/qwen38-27b-lora-16k.yaml b/src/lilo/presets/qwen38-27b-lora-16k.yaml deleted file mode 100644 index 7215db1..0000000 --- a/src/lilo/presets/qwen38-27b-lora-16k.yaml +++ /dev/null @@ -1,58 +0,0 @@ -api_version: lilo/v1 -name: qwen38-27b-lora-16k -model: - id: Qwen/Qwen3.8-27B - parameterization: lora - max_context_length: 16384 -routing: - default: true -trainer: - backend: miles - resources: - gpu: H200:8 - engine: - max_clients_per_instance: 6 - sampler_persistence_concurrency: 8 - env: - PYTORCH_CUDA_ALLOC_CONF: expandable_segments:True - TORCHINDUCTOR_COMPILE_THREADS: '1' - config: - model_args: qwen3.8-27B - options: - tensor_model_parallel_size: 4 - target_modules: - - linear_qkv - - linear_proj - - linear_fc1 - - linear_fc2 - max_tokens_per_gpu: 16384 - recompute_granularity: full - recompute_method: uniform - recompute_num_layers: 1 - multi_lora_n_adapters: 6 - lora_rank: 32 - lora_alpha: 32 -inference: - resources: - gpu: H200:1 - scaling: - min_replicas: 0 - max_replicas: 8 - target_concurrency: 16 - config: - tp_size: 1 - ep_size: 1 - mem_fraction_static: 0.8 - max_running_requests: 32 - max_queued_requests: 8 - max_loaded_loras: 256 - max_loras_per_batch: 8 - lora_target_modules: - - q_proj - - k_proj - - v_proj - - o_proj - - gate_proj - - up_proj - - down_proj - schedule_policy: lpm diff --git a/src/lilo/presets/qwen38-27b-lora-64k.yaml b/src/lilo/presets/qwen38-27b-lora-64k.yaml deleted file mode 100644 index 67128f4..0000000 --- a/src/lilo/presets/qwen38-27b-lora-64k.yaml +++ /dev/null @@ -1,16 +0,0 @@ -extends: builtin:qwen38-27b-lora-16k -name: qwen38-27b-lora-64k -model: - max_context_length: 65536 -routing: - default: false -trainer: - config: - options: - context_parallel_size: 2 - max_tokens_per_gpu: 32768 -inference: - scaling: - target_concurrency: 8 - config: - max_running_requests: 16 diff --git a/src/lilo/providers/modal/app.py b/src/lilo/providers/modal/app.py index 2a3254e..633f468 100644 --- a/src/lilo/providers/modal/app.py +++ b/src/lilo/providers/modal/app.py @@ -46,7 +46,7 @@ stop_pool as stop_lora_pool, ) from .sampling import ModalSamplingTaskPlatform -from .yaml_apps import ( +from .deployment_apps import ( MANIFEST_ENV, definition_from_spec, frontend_settings, diff --git a/src/lilo/providers/modal/yaml_apps.py b/src/lilo/providers/modal/deployment_apps.py similarity index 99% rename from src/lilo/providers/modal/yaml_apps.py rename to src/lilo/providers/modal/deployment_apps.py index 478a1d9..566e7bb 100644 --- a/src/lilo/providers/modal/yaml_apps.py +++ b/src/lilo/providers/modal/deployment_apps.py @@ -1,4 +1,4 @@ -"""Modal app builders shared by all YAML model deployments. +"""Modal app builders shared by all Python-configured model deployments. Trainer functions live in the frontend app. Rollout pools are separate apps, created on demand with the existing LoRA/FFT pool lifecycle. @@ -23,7 +23,7 @@ def manifest_from_env(): data = os.environ.get(MANIFEST_ENV) if not data: raise ValueError( - "Missing deployment manifest. Use lilo deploy with your YAML files." + "Missing deployment manifest. Use lilo deploy with your Python config files." ) rows = json.loads(data) if not isinstance(rows, list) or not rows: diff --git a/src/lilo/providers/modal/yaml_pool_app.py b/src/lilo/providers/modal/deployment_pool_app.py similarity index 93% rename from src/lilo/providers/modal/yaml_pool_app.py rename to src/lilo/providers/modal/deployment_pool_app.py index 7da18d7..55dd5b5 100644 --- a/src/lilo/providers/modal/yaml_pool_app.py +++ b/src/lilo/providers/modal/deployment_pool_app.py @@ -5,7 +5,7 @@ from lilo.deployments import DeploymentRecord from .fft_pool import FFTPoolSpec from .lora_pool import LoraPoolSpec -from .yaml_apps import POOL_CONFIG_ENV, build_rollout_app +from .deployment_apps import POOL_CONFIG_ENV, build_rollout_app resolved = DeploymentRecord.model_validate_json(os.environ[POOL_CONFIG_ENV]) if resolved.spec.model.parameterization == "lora": diff --git a/src/lilo/providers/modal/fft_pool.py b/src/lilo/providers/modal/fft_pool.py index e1b1865..45d6bdd 100644 --- a/src/lilo/providers/modal/fft_pool.py +++ b/src/lilo/providers/modal/fft_pool.py @@ -131,7 +131,7 @@ def deploy_pool(spec: FFTPoolSpec) -> str: modal_cli = shutil.which("modal") if modal_cli is None: raise RuntimeError("modal CLI is unavailable") - from .yaml_apps import pool_environment + from .deployment_apps import pool_environment recipe_env = pool_environment(spec.definition_id) env = {**os.environ, **spec.env(), **recipe_env} @@ -139,7 +139,7 @@ def deploy_pool(spec: FFTPoolSpec) -> str: modal_cli, "deploy", "-m", - "lilo.providers.modal.yaml_pool_app", + "lilo.providers.modal.deployment_pool_app", "--name", spec.app_name, ] diff --git a/src/lilo/providers/modal/lora_pool.py b/src/lilo/providers/modal/lora_pool.py index 3270636..e082ede 100644 --- a/src/lilo/providers/modal/lora_pool.py +++ b/src/lilo/providers/modal/lora_pool.py @@ -65,14 +65,14 @@ def deploy_pool(spec: LoraPoolSpec) -> str: modal_cli = shutil.which("modal") if modal_cli is None: raise RuntimeError("modal CLI is unavailable") - from .yaml_apps import pool_environment + from .deployment_apps import pool_environment recipe_env = pool_environment(spec.definition_id) command = [ modal_cli, "deploy", "-m", - "lilo.providers.modal.yaml_pool_app", + "lilo.providers.modal.deployment_pool_app", "--name", spec.app_name, ] diff --git a/tests/backends/test_native_megatron_config.py b/tests/backends/test_native_megatron_config.py index f9339b9..eecdcd9 100644 --- a/tests/backends/test_native_megatron_config.py +++ b/tests/backends/test_native_megatron_config.py @@ -1,4 +1,7 @@ -"""CPU checks for the boundary between YAML settings and Megatron constructors.""" +"""CPU checks for native config forwarding into Megatron constructors.""" + +from dataclasses import asdict +from pydantic import TypeAdapter from types import SimpleNamespace from unittest.mock import Mock @@ -8,19 +11,19 @@ from lilo.backends.deployment import backend_config from lilo.backends.megatron_config import parse_backend_config -from lilo.deployments import DeploymentSpec, load, preset_path +from lilo.deployments import BaseConfig, load, config_path with backend_runtime_imports(): from lilo.backends.megatron_runtime.common import modeling def test_yaml_native_values_reach_megatron(monkeypatch): - data = load(preset_path("qwen35-4b-fft-64k")).model_dump() + data = asdict(load(config_path("qwen35-4b-fft-64k"))) data["trainer"]["config"]["optimizer"]["native_optimizer_setting"] = False data["trainer"]["config"]["distributed"] = {"native_ddp_setting": 123} data["trainer"]["config"]["provider"]["native_provider_setting"] = [1, 2] config, _ = parse_backend_config( - backend_config(DeploymentSpec.model_validate(data)) + backend_config(TypeAdapter(BaseConfig).validate_python(data)) ) # These stand in for an installed upstream version with extra fields. The # deployment reader must not need its own list of those fields. diff --git a/tests/providers/conftest.py b/tests/providers/conftest.py index ba0ffd8..2e516ba 100644 --- a/tests/providers/conftest.py +++ b/tests/providers/conftest.py @@ -3,14 +3,14 @@ import json import os -from lilo.deployments import load, preset_path, DeploymentRecord +from lilo.deployments import load, config_path, DeploymentRecord os.environ.setdefault( "LILO_DEPLOYMENT_MANIFEST", json.dumps( [ DeploymentRecord.create( - load(preset_path(name)), revision="a" * 40, implementation="tests" + load(config_path(name)), revision="a" * 40, implementation="tests" ).model_dump(mode="json") for name in ("qwen35-9b-fft-64k", "qwen35-9b-lora-16k") ] diff --git a/tests/providers/test_checkpoint_storage.py b/tests/providers/test_checkpoint_storage.py index 33ea0e9..9516507 100644 --- a/tests/providers/test_checkpoint_storage.py +++ b/tests/providers/test_checkpoint_storage.py @@ -34,7 +34,7 @@ def test_checkpoint_storage_creates_one_v2_volume_without_live_lookup() -> None: def test_yaml_definitions_use_configured_checkpoint_storage(): - from lilo.providers.modal.yaml_apps import volumes_for + from lilo.providers.modal.deployment_apps import volumes_for from lilo.backends.deployment import backend_config for definition in DEFINITIONS: diff --git a/tests/providers/test_yaml_apps.py b/tests/providers/test_deployment_apps.py similarity index 86% rename from tests/providers/test_yaml_apps.py rename to tests/providers/test_deployment_apps.py index 546f4ca..85cca79 100644 --- a/tests/providers/test_yaml_apps.py +++ b/tests/providers/test_deployment_apps.py @@ -5,15 +5,15 @@ import modal import pytest -from lilo.deployments import load, preset_path, DeploymentRecord -from lilo.providers.modal import yaml_apps +from lilo.deployments import load, config_path, DeploymentRecord +from lilo.providers.modal import deployment_apps from lilo.providers.modal.fft_pool import FFTPoolSpec from lilo.providers.modal.lora_pool import LoraPoolSpec def deployment(preset="qwen35-9b-lora-16k"): return DeploymentRecord.create( - load(preset_path(preset)), revision="a" * 40, implementation="runtime" + load(config_path(preset)), revision="a" * 40, implementation="runtime" ) @@ -58,7 +58,7 @@ def test_trainer_declaration_and_executor_configuration( row = deployment(preset) row.spec.deployment.storage.checkpoints = "test-custom-checkpoints" image = object() - app, trainer = yaml_apps.build_trainer_app(row, image=image) + app, trainer = deployment_apps.build_trainer_app(row, image=image) declaration, _ = app.functions[row.definition_id] assert declaration["gpu"] == "H100:4" assert declaration["region"] == "us-west" @@ -71,7 +71,7 @@ def test_trainer_declaration_and_executor_configuration( monkeypatch.setattr(kv, "shared_kv", lambda: "store") monkeypatch.setattr( - yaml_apps, + deployment_apps, "volumes_for", lambda spec: {"/assets": SimpleNamespace(reload=lambda: reloaded.append(True))}, ) @@ -105,7 +105,7 @@ def test_pool_starts_native_server_and_correct_sidecar(builders, monkeypatch, ki else FFTPoolSpec(row.definition_id, "job", kind == "latest", 7) ) ) - app, server = yaml_apps.build_rollout_app(row, pool, image="test-image") + app, server = deployment_apps.build_rollout_app(row, pool, image="test-image") settings, _ = app.servers["Server"] assert app.name == pool.app_name assert settings["gpu"] == row.spec.inference.resources.gpu @@ -163,13 +163,13 @@ def test_pool_starts_native_server_and_correct_sidecar(builders, monkeypatch, ki def test_pool_subprocess_receives_recorded_generation(monkeypatch): row = deployment() - monkeypatch.setenv(yaml_apps.MANIFEST_ENV, json.dumps([row.model_dump()])) - env = yaml_apps.pool_environment(row.definition_id) - assert json.loads(env[yaml_apps.POOL_CONFIG_ENV])["generation"] == row.generation + monkeypatch.setenv(deployment_apps.MANIFEST_ENV, json.dumps([row.model_dump()])) + env = deployment_apps.pool_environment(row.definition_id) + assert json.loads(env[deployment_apps.POOL_CONFIG_ENV])["generation"] == row.generation with pytest.raises(ValueError, match="missing recorded"): - yaml_apps.pool_environment("yaml_missing_123") + deployment_apps.pool_environment("yaml_missing_123") with pytest.raises(ValueError, match="missing recorded"): - yaml_apps.pool_environment("unconfigured-python-definition") + deployment_apps.pool_environment("unconfigured-python-definition") def test_startup_failure_is_visible_and_blocks_new_spawns(monkeypatch): @@ -199,15 +199,15 @@ def test_real_modal_app_constructs_from_manifest_without_legacy_catalog(monkeypa import sys row = deployment() - env = {**os.environ, yaml_apps.MANIFEST_ENV: json.dumps([row.model_dump()])} + env = {**os.environ, deployment_apps.MANIFEST_ENV: json.dumps([row.model_dump()])} result = subprocess.run( [ sys.executable, "-c", """ import importlib, sys, modal -from lilo.providers.modal import yaml_apps -yaml_apps.image_for = lambda backend: modal.Image.debian_slim() +from lilo.providers.modal import deployment_apps +deployment_apps.image_for = lambda backend: modal.Image.debian_slim() app = importlib.import_module('lilo.providers.modal.app') assert len(app.DEFINITIONS) == 1 assert app.APP_NAME == 'lilo-yaml' @@ -229,16 +229,16 @@ def test_admission_changes_preserve_serialized_trainer(builders): from modal._serialization import serialize first = deployment() - old_bytes = serialize(yaml_apps.build_trainer_app(first, image="test")[1]) + old_bytes = serialize(deployment_apps.build_trainer_app(first, image="test")[1]) changed = first.model_copy(deep=True) changed.active = False changed.spec.routing.default = not first.spec.routing.default changed.spec.routing.sampling_default = True - new_bytes = serialize(yaml_apps.build_trainer_app(changed, image="test")[1]) + new_bytes = serialize(deployment_apps.build_trainer_app(changed, image="test")[1]) assert new_bytes == old_bytes assert first.active is True and first.spec.routing.default is True changed.spec.trainer.resources.gpu = "H200:4" - assert serialize(yaml_apps.build_trainer_app(changed, image="test")[1]) != old_bytes + assert serialize(deployment_apps.build_trainer_app(changed, image="test")[1]) != old_bytes @pytest.mark.parametrize("kind", ["lora", "full"]) @@ -246,7 +246,7 @@ def test_pool_launch_uses_only_generic_yaml_app(monkeypatch, kind): from lilo.providers.modal import fft_pool, lora_pool row = deployment("qwen35-9b-lora-16k" if kind == "lora" else "qwen35-4b-fft-64k") - monkeypatch.setenv(yaml_apps.MANIFEST_ENV, json.dumps([row.model_dump()])) + monkeypatch.setenv(deployment_apps.MANIFEST_ENV, json.dumps([row.model_dump()])) module = lora_pool if kind == "lora" else fft_pool spec = ( LoraPoolSpec(row.definition_id) @@ -274,8 +274,8 @@ def gateway_url(self): ) assert module.deploy_pool(spec) == "https://pool" command, kwargs = calls[0] - assert command[command.index("-m") + 1] == "lilo.providers.modal.yaml_pool_app" + assert command[command.index("-m") + 1] == "lilo.providers.modal.deployment_pool_app" assert ( - json.loads(kwargs["env"][yaml_apps.POOL_CONFIG_ENV])["generation"] + json.loads(kwargs["env"][deployment_apps.POOL_CONFIG_ENV])["generation"] == row.generation ) diff --git a/tests/providers/test_yaml_e2e_helper.py b/tests/providers/test_deployment_e2e_helper.py similarity index 90% rename from tests/providers/test_yaml_e2e_helper.py rename to tests/providers/test_deployment_e2e_helper.py index 7c3feb3..2c8bad4 100644 --- a/tests/providers/test_yaml_e2e_helper.py +++ b/tests/providers/test_deployment_e2e_helper.py @@ -5,7 +5,7 @@ import modal import pytest -from lilo.deployments import load, preset_path, DeploymentRecord +from lilo.deployments import load, config_path, DeploymentRecord @pytest.mark.parametrize("preset", ["qwen35-9b-lora-16k", "qwen35-4b-fft-64k"]) @@ -13,7 +13,7 @@ def test_e2e_helper_reads_active_deployed_configuration(monkeypatch, preset): helper = runpy.run_path( str(Path(__file__).parents[2] / "scripts/e2e_engine_definition.py") ) - row = DeploymentRecord.create(load(preset_path(preset)), revision="a" * 40, implementation="test") + row = DeploymentRecord.create(load(config_path(preset)), revision="a" * 40, implementation="test") retired = row.model_copy(update={"active": False, "generation": "b" * 64}) registry = SimpleNamespace( get=lambda *args: [retired.model_dump(), row.model_dump()] diff --git a/tests/providers/test_deployment_presets.py b/tests/providers/test_deployment_presets.py index 68f3bd0..a06f9bd 100644 --- a/tests/providers/test_deployment_presets.py +++ b/tests/providers/test_deployment_presets.py @@ -1,8 +1,10 @@ +from dataclasses import asdict +from pydantic import TypeAdapter import pytest -from lilo.deployments import load, preset_path, DeploymentRecord +from lilo.deployments import load, config_path, DeploymentRecord from lilo.backends.deployment import backend_config, serving_options -from lilo.providers.modal.yaml_apps import ( +from lilo.providers.modal.deployment_apps import ( definition_from_spec, frontend_settings, manifest_from_env, @@ -11,7 +13,7 @@ @pytest.mark.parametrize( "path", - sorted(preset_path("qwen35-9b-lora-16k").parent.glob("*.yaml")), + sorted(config_path("qwen35-9b-lora-16k").parent.glob("qwen*.py")), ids=lambda p: p.stem, ) def test_all_packaged_recipes_validate_offline(path): @@ -23,7 +25,7 @@ def test_all_packaged_recipes_validate_offline(path): @pytest.mark.parametrize("preset", ["qwen35-35b-a3b-fft-64k", "qwen36-35b-a3b-fft-64k"]) def test_moe_recipes_preserve_trainer_expert_parallelism(preset): - config = backend_config(load(preset_path(preset)))["megatron"] + config = backend_config(load(config_path(preset)))["megatron"] assert config["tensor_model_parallel_size"] == 4 assert config["context_parallel_size"] == 2 assert config["expert_model_parallel_size"] == 8 @@ -31,7 +33,7 @@ def test_moe_recipes_preserve_trainer_expert_parallelism(preset): def test_moe_rollout_preserves_attention_data_parallelism(): - spec = load(preset_path("qwen35-35b-a3b-fft-64k")) + spec = load(config_path("qwen35-35b-a3b-fft-64k")) options = serving_options(spec) assert options["tp_size"] == options["dp_size"] == options["ep_size"] == 4 assert options["enable_dp_attention"] is True @@ -44,7 +46,7 @@ def test_moe_rollout_preserves_attention_data_parallelism(): @pytest.mark.parametrize("context,cp", [(16384, 1), (65536, 2), (131072, 4)]) def test_qwen38_context_parallel_token_budget(context, cp): - spec = load(preset_path(f"qwen38-27b-lora-{context // 1024}k")) + spec = load(config_path(f"qwen38-27b-lora-{context // 1024}k")) config = backend_config(spec)["miles"] assert config["context_parallel_size"] == cp assert config["max_tokens_per_gpu"] == context // cp @@ -52,8 +54,8 @@ def test_qwen38_context_parallel_token_budget(context, cp): def test_single_client_recipe_keeps_shared_backend_capacity(): - shared = load(preset_path("qwen35-9b-lora-16k")) - single = load(preset_path("qwen35-9b-lora-16k-single")) + shared = load(config_path("qwen35-9b-lora-16k")) + single = load(config_path("qwen35-9b-lora-16k-single")) assert single.trainer.engine.max_clients_per_instance == 1 assert single.trainer.resources == shared.trainer.resources assert backend_config(single) == backend_config(shared) @@ -80,12 +82,12 @@ def test_missing_manifest_has_no_python_catalog_fallback(monkeypatch, value): ], ) def test_invalid_attention_parallelism_is_rejected(options): - from lilo.deployments import DeploymentSpec + from lilo.deployments import BaseConfig - data = load(preset_path("qwen35-35b-a3b-fft-64k")).model_dump() + data = asdict(load(config_path("qwen35-35b-a3b-fft-64k"))) data["inference"]["config"].update(options) from lilo.backends.deployment import serving_options - spec = DeploymentSpec.model_validate(data) + spec = TypeAdapter(BaseConfig).validate_python(data) with pytest.raises(ValueError, match="sglang"): serving_options(spec) diff --git a/tests/providers/test_modal_app.py b/tests/providers/test_modal_app.py index 22de02c..5a87921 100644 --- a/tests/providers/test_modal_app.py +++ b/tests/providers/test_modal_app.py @@ -10,12 +10,12 @@ from lilo.providers.modal.fft_pool import FFTPoolSpec from lilo.providers.modal.lora_pool import LoraPoolSpec -from lilo.deployments import load, preset_path, DeploymentRecord +from lilo.deployments import load, config_path, DeploymentRecord def definition_id(preset): return DeploymentRecord.create( - load(preset_path(preset)), revision="a" * 40, implementation="tests" + load(config_path(preset)), revision="a" * 40, implementation="tests" ).definition_id diff --git a/tests/test_deployment_cli.py b/tests/test_deployment_cli.py index 678c140..dbe735b 100644 --- a/tests/test_deployment_cli.py +++ b/tests/test_deployment_cli.py @@ -1,3 +1,4 @@ +from copy import deepcopy import json import subprocess @@ -5,8 +6,8 @@ import pytest from lilo import deployment_cli as cli -from lilo.deployments import load, preset_path, DeploymentRecord -from lilo.providers.modal.yaml_apps import MANIFEST_ENV +from lilo.deployments import load, config_path, DeploymentRecord +from lilo.providers.modal.deployment_apps import MANIFEST_ENV class Registry(dict): @@ -19,7 +20,7 @@ def put(self, key, value, skip_if_exists=False): def deployment(): return DeploymentRecord.create( - load(preset_path("qwen35-9b-lora-16k")), + load(config_path("qwen35-9b-lora-16k")), revision="a" * 40, implementation="test", ) @@ -69,7 +70,7 @@ def fail(*args, **kwargs): cli.deploy([row]) assert registry["pending"][0]["generation"] == row.generation assert "manifest" not in registry and "apply_lock" not in registry - new_spec = row.spec.model_copy(deep=True) + new_spec = deepcopy(row.spec) new_spec.trainer.scaling.max_instances = 2 new = DeploymentRecord.create(new_spec, revision="a" * 40, implementation="test") monkeypatch.setattr(subprocess, "run", lambda *args, **kwargs: None) @@ -82,7 +83,7 @@ def fail(*args, **kwargs): def test_refuse_overwriting_legacy_frontend(registry, monkeypatch): monkeypatch.setattr(modal.App, "lookup", lambda *args, **kwargs: object()) - with pytest.raises(ValueError, match="already exists without a YAML registry"): + with pytest.raises(ValueError, match="already exists without a deployment registry"): cli.deploy([deployment()]) assert "pending" not in registry @@ -94,7 +95,7 @@ def test_validate_never_resolves_or_deploys(monkeypatch, capsys): monkeypatch.setattr( cli, "deploy", lambda *args: pytest.fail("unexpected deployment") ) - cli.main(["config", "validate", str(preset_path("qwen35-9b-lora-16k"))]) + cli.main(["config", "validate", str(config_path("qwen35-9b-lora-16k"))]) assert "Validated 1 deployment" in capsys.readouterr().out @@ -116,13 +117,17 @@ def test_compile_pins_revision_at_external_boundary( from types import SimpleNamespace from unittest.mock import Mock import huggingface_hub - import yaml from lilo.providers.modal import miles_revision - data = load(preset_path("qwen35-9b-lora-16k")).model_dump() - data["model"]["revision"] = revision - path = tmp_path / "model.yaml" - path.write_text(yaml.safe_dump(data)) + path = tmp_path / "model.py" + path.write_text( + "from dataclasses import dataclass\n" + "from lilo.configs.qwen35_9b_lora_16k import Config as ParentConfig\n" + "@dataclass(kw_only=True)\n" + "class Config(ParentConfig):\n" + " def __post_init__(self):\n" + f" self.model.revision = {revision!r}\n" + ) lookup = Mock(return_value=SimpleNamespace(sha="a" * 40)) monkeypatch.setattr(huggingface_hub.HfApi, "model_info", lookup) monkeypatch.setattr(miles_revision, "resolve_miles_commit", lambda: "b" * 40) diff --git a/tests/test_deployments.py b/tests/test_deployments.py index beb273d..690d9fe 100644 --- a/tests/test_deployments.py +++ b/tests/test_deployments.py @@ -1,34 +1,35 @@ +from dataclasses import asdict +from pydantic import TypeAdapter import argparse import asyncio from types import SimpleNamespace import httpx import pytest -import yaml from lilo.deployments import ( DeploymentRecord, - DeploymentSpec, + BaseConfig, load, - preset_path, + config_path, validate_frontend, ) from lilo.deployment_cli import retain_generations from lilo.control_plane.deployments import DeploymentRoutes from lilo.backends.deployment import backend_config -from lilo.providers.modal.yaml_apps import definition_from_spec +from lilo.providers.modal.deployment_apps import definition_from_spec from lilo.native_options import apply_defaults def recipe(preset="qwen35-9b-lora-16k", **changes): - data = load(preset_path(preset)).model_dump() + data = asdict(load(config_path(preset))) for path, value in changes.items(): keys = path.split("__") target = data for key in keys[:-1]: target = target[key] target[keys[-1]] = value - return DeploymentSpec.model_validate(data) + return TypeAdapter(BaseConfig).validate_python(data) def resolved(spec=None, **changes): @@ -73,33 +74,6 @@ def test_no_model_catalog_required(): assert backend_config(spec)["miles"]["model_type"] == "" -def test_extends_false_and_lists_replace(tmp_path): - path = tmp_path / "child.yaml" - path.write_text( - yaml.safe_dump( - { - "extends": "builtin:qwen35-9b-lora-16k", - "routing": {"default": False}, - "trainer": {"config": {"options": {"target_modules": ["linear_qkv"]}}}, - } - ) - ) - spec = load(path) - assert spec.routing.default is False - assert spec.trainer.config["options"]["target_modules"] == ["linear_qkv"] - assert spec.trainer.config["options"]["lora_rank"] == 32 - - -def test_duplicate_keys_and_cycles(tmp_path): - path = tmp_path / "bad.yaml" - path.write_text("name: first\nname: second\n") - with pytest.raises(ValueError, match="duplicate"): - load(path) - path.write_text("extends: bad.yaml\n") - with pytest.raises(ValueError, match="cyclic"): - load(path) - - @pytest.mark.parametrize( "changes,match", [ @@ -300,14 +274,14 @@ def test_native_sections_survive_serialization_without_allowlist(): from lilo.backends.megatron_config import parse_backend_config spec = recipe("qwen35-4b-fft-64k") - data = spec.model_dump() + data = asdict(spec) data["trainer"]["config"]["provider"]["future_provider_option"] = { "layers": [1, 4], "enabled": False, } data["trainer"]["config"]["optimizer"]["future_optimizer_option"] = 0.125 data["trainer"]["config"]["distributed"] = {"future_ddp_option": False} - spec = DeploymentSpec.model_validate(data) + spec = TypeAdapter(BaseConfig).validate_python(data) settings = backend_config(spec, "/assets/pinned") config, _ = parse_backend_config(json.loads(json.dumps(settings))) assert config.hf_checkpoint == "/assets/pinned" @@ -319,7 +293,7 @@ def test_native_sections_survive_serialization_without_allowlist(): assert config.optimizer_overrides == {"future_optimizer_option": 0.125} assert config.distributed_overrides == {"future_ddp_option": False} assert config.optimizer.lr == 0.0001 - assert spec.model_dump() == data # Building does not consume or mutate YAML. + assert asdict(spec) == data # Building does not consume or mutate YAML. @pytest.mark.parametrize( @@ -360,43 +334,6 @@ def test_new_miles_and_sglang_options_need_no_deployment_schema_change(): assert serving_options(spec)["future_sglang_option"] is False -def test_load_merges_partial_parents_before_constructing_spec(tmp_path): - parent = tmp_path / "parent.yaml" - parent.write_text("trainer:\n config:\n options:\n custom_option: 1\n") - middle = tmp_path / "middle.yaml" - middle.write_text( - "extends: parent.yaml\ntrainer:\n config:\n options:\n custom_option: 2\n" - ) - data = recipe().model_dump() - data["extends"] = "middle.yaml" - data["trainer"]["config"]["options"]["other_option"] = False - child = tmp_path / "child.yaml" - child.write_text(yaml.safe_dump(data)) - spec = load(child) - assert spec.trainer.config["options"]["custom_option"] == 2 - assert spec.trainer.config["options"]["other_option"] is False - - -def test_loading_and_resolving_do_not_interpret_backend_config(tmp_path, monkeypatch): - import lilo.backends.deployment as backends - - def unexpected(*args, **kwargs): - pytest.fail("YAML construction must not interpret backend configuration") - - monkeypatch.setattr(backends, "backend_config", unexpected) - monkeypatch.setattr(backends, "serving_options", unexpected) - data = recipe().model_dump() - data["trainer"]["backend"] = "unknown-until-startup" - path = tmp_path / "deployment.yaml" - path.write_text(yaml.safe_dump(data)) - spec = load(path) - assert DeploymentRecord.create( - spec, revision="a" * 40, implementation="test" - ).spec == spec.model_copy( - update={"model": spec.model.model_copy(update={"revision": "a" * 40})} - ) - - def test_fft_capacity_is_checked_by_backend_setup(): spec = recipe("qwen35-4b-fft-64k", trainer__engine__max_clients_per_instance=2) with pytest.raises(ValueError, match="FFT trainers admit one client"): @@ -404,7 +341,7 @@ def test_fft_capacity_is_checked_by_backend_setup(): def test_reserved_environment_is_checked_by_modal_setup(): - from lilo.providers.modal.yaml_apps import deployment_env + from lilo.providers.modal.deployment_apps import deployment_env spec = recipe(trainer__env={"LILO_BACKEND_CONFIG": "oops"}) with pytest.raises(ValueError, match="managed"): @@ -412,32 +349,14 @@ def test_reserved_environment_is_checked_by_modal_setup(): assert deployment_env({"MY_SETTING": "value"}) == {"MY_SETTING": "value"} -def test_inheritance_keeps_intermediate_replacements(tmp_path): - parent = recipe().model_dump() - parent["trainer"]["config"]["options"]["custom"] = {"old": 1} - (tmp_path / "parent.yaml").write_text(yaml.safe_dump(parent)) - (tmp_path / "middle.yaml").write_text( - "extends: parent.yaml\ntrainer:\n config:\n options:\n custom: null\n" - ) - child = tmp_path / "child.yaml" - child.write_text( - "extends: middle.yaml\ntrainer:\n config:\n options:\n custom:\n new: 2\n" - ) - assert load(child).trainer.config["options"]["custom"] == {"new": 2} - - -def test_record_creation_copies_without_reparsing(monkeypatch): +def test_record_creation_copies_without_reparsing(): import hashlib import json spec = recipe() - original = spec.model_dump() - # Creating a record must not reconstruct an already-parsed specification. - monkeypatch.setattr( - DeploymentSpec, "model_validate", lambda *a, **k: pytest.fail("reparse") - ) + original = asdict(spec) row = DeploymentRecord.create(spec, revision="a" * 40, implementation="test") - assert spec.model_dump() == original + assert asdict(spec) == original assert row.spec.model.revision == "a" * 40 # Keep the existing manifest fields and hash format stable. expected = original | {"model": original["model"] | {"revision": "a" * 40}} @@ -452,3 +371,71 @@ def test_record_creation_copies_without_reparsing(monkeypatch): assert spec.trainer.config["options"]["lora_rank"] == 32 saved = row.model_dump_json() assert DeploymentRecord.model_validate_json(saved) == row + + +def test_python_config_inheritance_and_independent_defaults(tmp_path): + from dataclasses import is_dataclass + + path = tmp_path / "model.py" + path.write_text( + "from dataclasses import dataclass\n" + "from lilo.configs.qwen35_9b_lora_64k import Config as ParentConfig\n" + "@dataclass(kw_only=True)\n" + "class Config(ParentConfig):\n" + " name: str = 'custom'\n" + " def __post_init__(self):\n" + " super().__post_init__()\n" + " self.trainer.config['options']['new_backend_option'] = False\n" + ) + first, second = load(path), load(path) + assert is_dataclass(first) + assert first.name == "custom" + assert first.model.max_context_length == 65536 + assert first.trainer.config["options"]["new_backend_option"] is False + first.trainer.config["options"]["target_modules"].append("extra") + assert "extra" not in second.trainer.config["options"]["target_modules"] + assert ( + "extra" + not in load(config_path("qwen35-9b-lora-16k")).trainer.config["options"][ + "target_modules" + ] + ) + + +def test_loading_python_config_does_not_call_backend_readers(monkeypatch): + import lilo.backends.deployment as backends + + monkeypatch.setattr( + backends, "backend_config", lambda *a: pytest.fail("backend read") + ) + monkeypatch.setattr( + backends, "serving_options", lambda *a: pytest.fail("serving read") + ) + spec = load(config_path("qwen35-9b-lora-16k")) + record = DeploymentRecord.create(spec, revision="a" * 40, implementation="test") + assert record.spec.model.revision == "a" * 40 + assert spec.model.revision == "main" + + +@pytest.mark.parametrize("source", ["Config = {}", "class Config: pass", "value = 1"]) +def test_config_file_must_export_config_subclass(tmp_path, source): + path = tmp_path / "model.py" + path.write_text(source) + with pytest.raises(ValueError, match="Config subclass"): + load(path) + + +def test_config_import_error_preserves_traceback_and_restores_path(tmp_path): + import sys + + path = tmp_path / "model.py" + path.write_text("raise RuntimeError('bad user config')") + before = list(sys.path) + with pytest.raises(RuntimeError, match="bad user config"): + load(path) + assert sys.path == before + + +def test_no_yaml_config_ingestion(tmp_path): + with pytest.raises(ValueError, match="Python .py"): + load(tmp_path / "old.yaml") diff --git a/uv.lock b/uv.lock index 6b608a4..3d76c80 100644 --- a/uv.lock +++ b/uv.lock @@ -482,7 +482,6 @@ dependencies = [ { name = "opentelemetry-sdk" }, { name = "protobuf" }, { name = "pydantic" }, - { name = "pyyaml" }, { name = "stitch" }, { name = "tinker" }, { name = "uvicorn" }, @@ -505,7 +504,6 @@ requires-dist = [ { name = "opentelemetry-sdk", specifier = ">=1.39,<2" }, { name = "protobuf", specifier = ">=5.29" }, { name = "pydantic", specifier = ">=2.13.4" }, - { name = "pyyaml", specifier = ">=6.0.2" }, { name = "stitch", git = "https://github.com/modal-projects/stitch.git?rev=375a9396a7b05770dc4ed9cc5fe34fc4d5a472d5" }, { name = "tinker", specifier = ">=0.24.1,<0.25" }, { name = "uvicorn", specifier = ">=0.52.0" }, From 3d4bded488011d9681351c8ccf01540c4b1c098a Mon Sep 17 00:00:00 2001 From: kailash Date: Wed, 23 Sep 2026 18:29:09 +0000 Subject: [PATCH 13/27] Keep deployment definitions in one config directory --- deployments/qwen35_4b_fft_64k.py | 8 ----- deployments/qwen35_9b_lora_16k.py | 8 ----- deployments/qwen35_9b_lora_64k.py | 9 ----- docs/deployment-configs.md | 14 ++++---- docs/deployment-validation.md | 8 ++++- docs/design.md | 2 +- scripts/deploy_models.sh | 8 ++--- src/lilo/configs/qwen35_4b_fft_64k.py | 2 +- src/lilo/configs/qwen35_9b_lora_16k.py | 2 +- src/lilo/deployment_cli.py | 5 ++- src/lilo/providers/modal/app.py | 4 ++- .../providers/modal/image_dependencies.py | 5 +++ src/lilo/providers/modal/megatron_image.py | 4 ++- src/lilo/providers/modal/miles_image.py | 4 ++- src/lilo/providers/modal/rollout_image.py | 4 ++- tests/test_deployment_cli.py | 33 ++++++++++++++++++- tests/test_deployments.py | 3 +- 17 files changed, 77 insertions(+), 46 deletions(-) delete mode 100644 deployments/qwen35_4b_fft_64k.py delete mode 100644 deployments/qwen35_9b_lora_16k.py delete mode 100644 deployments/qwen35_9b_lora_64k.py diff --git a/deployments/qwen35_4b_fft_64k.py b/deployments/qwen35_4b_fft_64k.py deleted file mode 100644 index 911d983..0000000 --- a/deployments/qwen35_4b_fft_64k.py +++ /dev/null @@ -1,8 +0,0 @@ -from dataclasses import dataclass -from lilo.configs.qwen35_4b_fft_64k import Config as ParentConfig - - -@dataclass(kw_only=True) -class Config(ParentConfig): - def __post_init__(self): - self.model.revision = "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a" diff --git a/deployments/qwen35_9b_lora_16k.py b/deployments/qwen35_9b_lora_16k.py deleted file mode 100644 index 42f7060..0000000 --- a/deployments/qwen35_9b_lora_16k.py +++ /dev/null @@ -1,8 +0,0 @@ -from dataclasses import dataclass -from lilo.configs.qwen35_9b_lora_16k import Config as ParentConfig - - -@dataclass(kw_only=True) -class Config(ParentConfig): - def __post_init__(self): - self.model.revision = "68c46c4b3498877f3ef123c856ecfde50c39f404" diff --git a/deployments/qwen35_9b_lora_64k.py b/deployments/qwen35_9b_lora_64k.py deleted file mode 100644 index 161f28d..0000000 --- a/deployments/qwen35_9b_lora_64k.py +++ /dev/null @@ -1,9 +0,0 @@ -from dataclasses import dataclass -from lilo.configs.qwen35_9b_lora_64k import Config as ParentConfig - - -@dataclass(kw_only=True) -class Config(ParentConfig): - def __post_init__(self): - super().__post_init__() - self.model.revision = "68c46c4b3498877f3ef123c856ecfde50c39f404" diff --git a/docs/deployment-configs.md b/docs/deployment-configs.md index ec45ca8..d9661f0 100644 --- a/docs/deployment-configs.md +++ b/docs/deployment-configs.md @@ -9,7 +9,7 @@ The layout follows the Python recipe approach used by [training-gym](https://git Start with an example: ```bash -lilo config init --preset qwen35-9b-lora-16k > deployments/my_model.py +lilo config init --preset qwen35-9b-lora-16k > src/lilo/configs/my_model.py ``` The generated file imports a packaged config and subclasses it. Customize it using ordinary Python: @@ -35,16 +35,18 @@ When a parent defines `__post_init__`, call `super().__post_init__()` before you Python imports provide reuse. There is no `extends` key or implicit dictionary merge. Backend dictionaries support normal Python operations such as `update()` or `|`. Config files execute as Python when loaded; keep provisioning and training calls outside them. Sibling imports are available while loading a file. -All 14 previous presets are available under [`src/lilo/configs/`](../src/lilo/configs), with the same model, resource and backend settings. The [64K config](../src/lilo/configs/qwen35_9b_lora_64k.py) inherits from the 16K config. The three files under [`deployments/`](../deployments) pin the model revisions used in the earlier GPU checks. +All 14 previous presets are available under [`src/lilo/configs/`](../src/lilo/configs), with the same model, resource and backend settings. The [64K config](../src/lilo/configs/qwen35_9b_lora_64k.py) inherits from the 16K config. The 9B LoRA and 4B FFT configs directly pin the model revisions used in the earlier GPU checks; derived configs inherit them. There is no separate `deployments/` wrapper directory. + +`model.revision` is the Hugging Face commit or branch containing the base weights and tokenizer. You can omit it to use `"main"`; the CLI resolves that branch to an exact commit before deployment. A fixed commit makes repeated deployments use the same files even if the repository's main branch changes. This is separate from training steps and published adapter versions. ## Deploy the complete active set Use Python 3.12, matching the serialized trainer and inference images: ```bash -lilo config validate deployments/my_model.py -lilo config resolve deployments/my_model.py --output /tmp/deployment.json -lilo deploy deployments/model_a.py deployments/model_b.py +lilo config validate src/lilo/configs/my_model.py +lilo config resolve src/lilo/configs/my_model.py --output /tmp/deployment.json +lilo deploy src/lilo/configs/model_a.py src/lilo/configs/model_b.py ``` `validate` loads the Python classes and checks shared frontend settings. It does not start backend libraries or prove that the model fits in GPU memory. `resolve` additionally pins model revisions and emits the deployment records as JSON; it does not provision compute. @@ -136,7 +138,7 @@ If multiple configurations serve the same model and training mode, select one wi Every client stores its selected definition ID. Changing routing defaults affects new clients. Changing compute or backend settings creates a new configuration ID, and old configurations remain registered for existing jobs and checkpoints. The CLI serializes applies and retains interrupted deployments for recovery. -The shared app is deployed on each apply. Existing trainer functions use stable captured configuration JSON, and unchanged inference pools retain their apps. This is the shared-app architecture, not independent trainer-app deployment. Source/runtime upgrades still require a separate frontend in this draft. Config files outside the installed Lilo package can change without changing the runtime source fingerprint. +The shared app is deployed on each apply. Existing trainer functions use stable captured configuration JSON, and unchanged inference pools retain their apps. This is the shared-app architecture, not independent trainer-app deployment. Source/runtime upgrades still require a separate frontend in this draft. The `src/lilo/configs/` directory is excluded from runtime fingerprints and image source mounts. Its computed values are stored and hashed in each deployment record, so editing a config does not count as a runtime-code upgrade. Backend code changes still do. Existing resource names (`lilo-yaml`, `*-yaml-deployments`, and the `yaml_` definition prefix) are retained so this authoring change does not rename saved resources. They no longer indicate a YAML ingestion path. PyYAML is not a direct Lilo dependency; other installed libraries may depend on it. diff --git a/docs/deployment-validation.md b/docs/deployment-validation.md index 65ab4e5..9ce13c4 100644 --- a/docs/deployment-validation.md +++ b/docs/deployment-validation.md @@ -2,6 +2,12 @@ These checks exercise the shared-app deployment path in PR #55. Each deployment contains the frontend and generated trainer functions; inference pools are separate apps started on demand. +## One config directory + +Removed the three `deployments/` wrappers. The deployment script now lists `src/lilo/configs/` directly, and the validated model revisions live in the 9B LoRA and 4B FFT configs. Derived configs inherit those revisions. Authoring configs are excluded from runtime source fingerprints and shared-app/trainer/inference source mounts; workers receive computed settings as JSON. + +Validation: **599 CPU tests passed, 1 skipped**; Ruff, whitespace and deployment-script syntax checks passed. The CLI loaded the three selected configs together. Regression tests verify that editing config source leaves the runtime fingerprint unchanged, editing backend source changes it, and worker source filtering retains runtime code while excluding authoring files. No apps were redeployed. + ## Python dataclass configuration Python `Config` subclasses replace the YAML presets and loader. All 14 configs were compared with the previous specifications: model, resources, backend options and inference options are unchanged. The three checked-in deployment files retain their pinned model revisions. Inheritance uses normal dataclass defaults and `__post_init__`; tests verify mutable defaults are independent. @@ -85,7 +91,7 @@ The western 8×H200 trainer remained queued without a container for approximatel Use Python 3.12 and an authenticated Modal environment. Set `TINKER_API_KEY` locally to the value in the deployment's API secret. Generate Python config files from the packaged examples, give them the same isolated `deployment.frontend`, and pin model revisions plus `LILO_MILES_COMMIT` before deploying. ```bash -lilo deploy deployments/qwen35_9b_lora_16k.py deployments/qwen35_9b_lora_64k.py deployments/qwen35_4b_fft_64k.py +lilo deploy src/lilo/configs/qwen35_9b_lora_16k.py src/lilo/configs/qwen35_9b_lora_64k.py src/lilo/configs/qwen35_4b_fft_64k.py python scripts/deployment_smoke.py \ --frontend YOUR_TEST_FRONTEND \ --name qwen35-9b-lora-16k \ diff --git a/docs/design.md b/docs/design.md index b431a8d..169d6da 100644 --- a/docs/design.md +++ b/docs/design.md @@ -126,7 +126,7 @@ sampling scales according to rollout traffic. We best-effort sticky-route groups A Python deployment dataclass specifies the base model, training mode, context length, GPUs, parallelism, and inference settings. Shared deployments are defined only through these files. -1. Create a Python `Config` subclass under `deployments/`, optionally inheriting from a packaged config. +1. Create a Python `Config` subclass under `src/lilo/configs/`, optionally inheriting from a packaged config. 2. Add its path to the list in [`scripts/deploy_models.sh`](../scripts/deploy_models.sh). 3. Run the script to apply the complete list to the shared frontend. diff --git a/scripts/deploy_models.sh b/scripts/deploy_models.sh index 5aa3df4..2b74006 100755 --- a/scripts/deploy_models.sh +++ b/scripts/deploy_models.sh @@ -4,12 +4,12 @@ set -e # Run from the repository root so the Python config paths below work from any directory. cd "$(dirname "$0")/.." -# Add a model by creating its Python config in deployments/ and adding it here. +# Add a model by creating its Python config in src/lilo/configs/ and adding it here. # Keep every configuration that should be available to new clients in this list. deployment_files=( - deployments/qwen35_9b_lora_16k.py - deployments/qwen35_9b_lora_64k.py - deployments/qwen35_4b_fft_64k.py + src/lilo/configs/qwen35_9b_lora_16k.py + src/lilo/configs/qwen35_9b_lora_64k.py + src/lilo/configs/qwen35_4b_fft_64k.py ) lilo deploy "${deployment_files[@]}" diff --git a/src/lilo/configs/qwen35_4b_fft_64k.py b/src/lilo/configs/qwen35_4b_fft_64k.py index 44f9578..3a1b45e 100644 --- a/src/lilo/configs/qwen35_4b_fft_64k.py +++ b/src/lilo/configs/qwen35_4b_fft_64k.py @@ -17,7 +17,7 @@ class Config(BaseConfig): model: Model = field( default_factory=lambda: Model( id="Qwen/Qwen3.5-4B", - revision="main", + revision="851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", parameterization="full", max_context_length=65536, ) diff --git a/src/lilo/configs/qwen35_9b_lora_16k.py b/src/lilo/configs/qwen35_9b_lora_16k.py index 0e25106..637a8cb 100644 --- a/src/lilo/configs/qwen35_9b_lora_16k.py +++ b/src/lilo/configs/qwen35_9b_lora_16k.py @@ -19,7 +19,7 @@ class Config(BaseConfig): model: Model = field( default_factory=lambda: Model( id="Qwen/Qwen3.5-9B-Base", - revision="main", + revision="68c46c4b3498877f3ef123c856ecfde50c39f404", parameterization="lora", max_context_length=16384, ) diff --git a/src/lilo/deployment_cli.py b/src/lilo/deployment_cli.py index 2898d4d..409e629 100644 --- a/src/lilo/deployment_cli.py +++ b/src/lilo/deployment_cli.py @@ -28,7 +28,10 @@ def implementation_fingerprint(miles_commit: str | None) -> str: root = Path(__file__).parent digest = hashlib.sha256() for path in sorted(root.rglob("*.py")): - digest.update(str(path.relative_to(root)).encode()) + relative = path.relative_to(root) + if relative.parts[0] == "configs": + continue # Config values are hashed separately in DeploymentRecord. + digest.update(str(relative).encode()) digest.update(path.read_bytes()) digest.update(json.dumps([miles_commit, sorted(requires("lilo") or [])]).encode()) return digest.hexdigest() diff --git a/src/lilo/providers/modal/app.py b/src/lilo/providers/modal/app.py index 633f468..bb0d3a1 100644 --- a/src/lilo/providers/modal/app.py +++ b/src/lilo/providers/modal/app.py @@ -7,6 +7,8 @@ from dataclasses import asdict import modal + +from .image_dependencies import ignore_config_source from stitch.pools.modal_flash import ModalFlashPool from lilo.providers.contracts import ( @@ -105,7 +107,7 @@ async def _delete_checkpoint(uri: str) -> None: .pip_install(*CORE_PACKAGES, TINKER_PACKAGE) .pip_install(STITCH_PACKAGE, "huggingface-hub") .env(TRAINER_DEPLOYMENT_ENV) - .add_local_python_source("lilo") + .add_local_python_source("lilo", ignore=ignore_config_source) ) model_assets = modal.Volume.from_name( SETTINGS.deployment.storage.assets, diff --git a/src/lilo/providers/modal/image_dependencies.py b/src/lilo/providers/modal/image_dependencies.py index e9c73a7..6d8c071 100644 --- a/src/lilo/providers/modal/image_dependencies.py +++ b/src/lilo/providers/modal/image_dependencies.py @@ -1,3 +1,8 @@ +def ignore_config_source(path): + """Workers receive computed configs as JSON; authoring files stay local.""" + return path.suffix != ".py" or path.parts[0] == "configs" + + CORE_PACKAGES = ( "fastapi>=0.141.1", "httpx>=0.28.1", diff --git a/src/lilo/providers/modal/megatron_image.py b/src/lilo/providers/modal/megatron_image.py index 6105296..7ab4621 100644 --- a/src/lilo/providers/modal/megatron_image.py +++ b/src/lilo/providers/modal/megatron_image.py @@ -1,5 +1,7 @@ import modal +from .image_dependencies import ignore_config_source + from .image_dependencies import ( CORE_PACKAGES, MEGATRON_RUNTIME_CHECK, @@ -44,5 +46,5 @@ .pip_install(*CORE_PACKAGES, STITCH_PACKAGE) .pip_install(*MEGATRON_RUNTIME_PACKAGES) .run_commands(MEGATRON_RUNTIME_CHECK) - .add_local_python_source("lilo") + .add_local_python_source("lilo", ignore=ignore_config_source) ) diff --git a/src/lilo/providers/modal/miles_image.py b/src/lilo/providers/modal/miles_image.py index be46765..fd66aba 100644 --- a/src/lilo/providers/modal/miles_image.py +++ b/src/lilo/providers/modal/miles_image.py @@ -1,5 +1,7 @@ import modal +from .image_dependencies import ignore_config_source + from .miles_revision import MILES_REPOSITORY, resolve_miles_commit from .image_dependencies import ( @@ -65,5 +67,5 @@ "('load_slot', 'unload_slot', 'forward_backward', " "'forward_only', 'optim_step', 'save_slot', 'export_slot'))\"", ) - .add_local_python_source("lilo") + .add_local_python_source("lilo", ignore=ignore_config_source) ) diff --git a/src/lilo/providers/modal/rollout_image.py b/src/lilo/providers/modal/rollout_image.py index 0aaa9ff..3b3f3ff 100644 --- a/src/lilo/providers/modal/rollout_image.py +++ b/src/lilo/providers/modal/rollout_image.py @@ -1,5 +1,7 @@ import modal +from .image_dependencies import ignore_config_source + from .image_dependencies import CORE_PACKAGES, STITCH_PACKAGE, TINKER_PACKAGE SGLANG_IMAGE = "lmsysorg/sglang:v0.5.17" @@ -240,5 +242,5 @@ "SGLANG_DISABLE_CUDNN_CHECK": "1", } ) - .add_local_python_source("lilo") + .add_local_python_source("lilo", ignore=ignore_config_source) ) diff --git a/tests/test_deployment_cli.py b/tests/test_deployment_cli.py index dbe735b..9f5434d 100644 --- a/tests/test_deployment_cli.py +++ b/tests/test_deployment_cli.py @@ -83,7 +83,9 @@ def fail(*args, **kwargs): def test_refuse_overwriting_legacy_frontend(registry, monkeypatch): monkeypatch.setattr(modal.App, "lookup", lambda *args, **kwargs: object()) - with pytest.raises(ValueError, match="already exists without a deployment registry"): + with pytest.raises( + ValueError, match="already exists without a deployment registry" + ): cli.deploy([deployment()]) assert "pending" not in registry @@ -139,3 +141,32 @@ def test_compile_pins_revision_at_external_boundary( lookup.return_value.sha = None with pytest.raises(ValueError, match="did not return a commit"): cli.compile_configs([path]) + + +def test_config_edits_do_not_change_runtime_fingerprint(tmp_path, monkeypatch): + import importlib.metadata + + root = tmp_path / "lilo" + (root / "configs").mkdir(parents=True) + runtime = root / "runtime.py" + runtime.write_text("runtime = 1") + config = root / "configs" / "example.py" + config.write_text("gpu = 'H100:4'") + monkeypatch.setattr(cli, "__file__", str(root / "deployment_cli.py")) + monkeypatch.setattr(importlib.metadata, "requires", lambda _: []) + before = cli.implementation_fingerprint("miles-commit") + config.write_text("gpu = 'H200:8'") + assert cli.implementation_fingerprint("miles-commit") == before + runtime.write_text("runtime = 2") + assert cli.implementation_fingerprint("miles-commit") != before + + +def test_worker_source_mount_excludes_authoring_configs(): + from pathlib import Path + from lilo.providers.modal.image_dependencies import ignore_config_source + + assert ignore_config_source(Path("configs/example.py")) + assert ignore_config_source(Path("configs/__init__.py")) + assert ignore_config_source(Path("data.json")) + assert not ignore_config_source(Path("deployments.py")) + assert not ignore_config_source(Path("backends/miles_config.py")) diff --git a/tests/test_deployments.py b/tests/test_deployments.py index 690d9fe..64022e2 100644 --- a/tests/test_deployments.py +++ b/tests/test_deployments.py @@ -412,9 +412,10 @@ def test_loading_python_config_does_not_call_backend_readers(monkeypatch): backends, "serving_options", lambda *a: pytest.fail("serving read") ) spec = load(config_path("qwen35-9b-lora-16k")) + original_revision = spec.model.revision record = DeploymentRecord.create(spec, revision="a" * 40, implementation="test") assert record.spec.model.revision == "a" * 40 - assert spec.model.revision == "main" + assert spec.model.revision == original_revision @pytest.mark.parametrize("source", ["Config = {}", "class Config: pass", "value = 1"]) From e018a7904ecafa94636a0865b7faadcd5ae1df7e Mon Sep 17 00:00:00 2001 From: kailash Date: Wed, 23 Sep 2026 18:34:23 +0000 Subject: [PATCH 14/27] Simplify config recipes with class defaults and inherited overrides --- docs/deployment-configs.md | 30 +++-- docs/deployment-validation.md | 6 + src/lilo/configs/qwen35_35b_a3b_fft_64k.py | 116 ++++++++--------- src/lilo/configs/qwen35_4b_fft_64k.py | 92 ++++++------- src/lilo/configs/qwen35_9b_fft_64k.py | 94 ++++++-------- .../configs/qwen35_9b_instruct_lora_16k.py | 118 ++++++++--------- .../qwen35_9b_instruct_lora_16k_dp2.py | 12 +- src/lilo/configs/qwen35_9b_lora_16k.py | 112 +++++++--------- src/lilo/configs/qwen35_9b_lora_16k_single.py | 9 +- src/lilo/configs/qwen35_9b_lora_2k.py | 122 ++++++++---------- src/lilo/configs/qwen35_9b_lora_64k.py | 18 ++- src/lilo/configs/qwen36_27b_fft_64k.py | 98 +++++++------- src/lilo/configs/qwen36_35b_a3b_fft_64k.py | 14 +- src/lilo/configs/qwen38_27b_lora_128k.py | 28 ++-- src/lilo/configs/qwen38_27b_lora_16k.py | 118 ++++++++--------- src/lilo/configs/qwen38_27b_lora_64k.py | 20 ++- src/lilo/deployment_cli.py | 8 +- src/lilo/deployments.py | 52 +++++++- tests/test_deployment_cli.py | 5 +- tests/test_deployments.py | 57 +++++++- 20 files changed, 562 insertions(+), 567 deletions(-) diff --git a/docs/deployment-configs.md b/docs/deployment-configs.md index d9661f0..9b51c43 100644 --- a/docs/deployment-configs.md +++ b/docs/deployment-configs.md @@ -1,8 +1,8 @@ # Python deployment configs -Each deployment is a Python file exporting a `Config` dataclass that inherits from `BaseConfig`. The class contains the model, trainer, inference and Modal settings. There is no YAML loader or model catalog. +Each deployment is a Python file exporting a `Config` class that inherits from `BaseConfig`. The class contains the model, trainer, inference and Modal settings. There is no YAML loader or model catalog. -The layout follows the Python recipe approach used by [training-gym](https://github.com/modal-labs/training-gym) and the [multinode training guide](https://github.com/modal-labs/multinode-training-guide/blob/main/nemo-rl/configs/llama3_1_8b_math_2node.py). Lilo's config classes use standard-library dataclasses; importing either project is not required. +The layout follows the Python recipe approach used by [training-gym](https://github.com/modal-labs/training-gym) and the [multinode training guide](https://github.com/modal-labs/multinode-training-guide/blob/main/nemo-rl/configs/llama3_1_8b_math_2node.py). Lilo uses dataclasses for the resulting settings, but config files need no decorators or default factories. Importing either project is not required. ## Define a deployment @@ -15,25 +15,27 @@ lilo config init --preset qwen35-9b-lora-16k > src/lilo/configs/my_model.py The generated file imports a packaged config and subclasses it. Customize it using ordinary Python: ```python -from dataclasses import dataclass from lilo.configs.qwen35_9b_lora_16k import Config as ParentConfig -@dataclass(kw_only=True) class Config(ParentConfig): - name: str = "my-9b-64k" - - def __post_init__(self): - self.model.max_context_length = 65536 - self.trainer.resources.gpu = "H200:8" - self.trainer.config["options"]["tensor_model_parallel_size"] = 8 - self.trainer.config["options"]["max_tokens_per_gpu"] = 65536 - self.inference.scaling.max_replicas = 6 + name = "my-9b-64k" + overrides = { + "model.max_context_length": 65536, + "trainer.resources.gpu": "H200:8", + "trainer.config.options.tensor_model_parallel_size": 8, + "trainer.config.options.max_tokens_per_gpu": 65536, + "inference.scaling.max_replicas": 6, + } ``` -When a parent defines `__post_init__`, call `super().__post_init__()` before your changes. Mutable defaults use `field(default_factory=...)`, so modifying one instance does not change another config. New deployments can also inherit directly from `BaseConfig` and supply `Model`, `Trainer` and `Inference` fields; the packaged [16K example](../src/lilo/configs/qwen35_9b_lora_16k.py) shows the complete structure. +New deployments can inherit directly from `BaseConfig` and declare ordinary class defaults such as `model = Model(...)` and `trainer = Trainer(...)`. The [16K example](../src/lilo/configs/qwen35_9b_lora_16k.py) shows the complete structure. No `@dataclass`, `field(default_factory=...)`, or `__post_init__` is needed in config files. -Python imports provide reuse. There is no `extends` key or implicit dictionary merge. Backend dictionaries support normal Python operations such as `update()` or `|`. Config files execute as Python when loaded; keep provisioning and training calls outside them. Sibling imports are available while loading a file. +`BaseConfig` copies defaults for each instance, then applies each parent's overrides before its child's. Dotted paths traverse dataclass fields and dictionary keys. A value replaces the selected field/key, including whole lists and dictionaries; other keys remain unchanged. Native backend dictionaries can receive new option names. Misspelled dataclass fields and missing intermediate paths raise an error naming the override. Constructor keywords, when supplied, replace top-level fields last. + +Copying happens inside the base class, so edits to a config's nested lists or dictionaries do not change its parent, another instance, or the class's override dictionary. Only the resulting settings enter the saved deployment record; workers do not apply inheritance again. + +Python imports provide reuse. There is no YAML loader or `extends` key. Config files execute as Python when loaded; keep provisioning and training calls outside them. Sibling imports are available while loading a file. All 14 previous presets are available under [`src/lilo/configs/`](../src/lilo/configs), with the same model, resource and backend settings. The [64K config](../src/lilo/configs/qwen35_9b_lora_64k.py) inherits from the 16K config. The 9B LoRA and 4B FFT configs directly pin the model revisions used in the earlier GPU checks; derived configs inherit them. There is no separate `deployments/` wrapper directory. diff --git a/docs/deployment-validation.md b/docs/deployment-validation.md index 9ce13c4..9d1f383 100644 --- a/docs/deployment-validation.md +++ b/docs/deployment-validation.md @@ -2,6 +2,12 @@ These checks exercise the shared-app deployment path in PR #55. Each deployment contains the frontend and generated trainer functions; inference pools are separate apps started on demand. +## Simple class defaults and inherited overrides + +Config files now declare ordinary class defaults and dotted `overrides` dictionaries. They contain no dataclass decorators, default factories or `__post_init__` methods. `BaseConfig` copies values per instance and applies parent defaults/overrides before child defaults/overrides. The resulting settings and saved-record format are unchanged for all 14 configs. + +Validation: **603 CPU tests passed, 1 skipped**. Added tests cover inherited overrides, child field replacement, false values, list/dictionary replacement, independent mutable values and override typo errors. All 14 computed configs were compared with the previous version and round-tripped through saved JSON. The CLI generated and validated the simplified config. Ruff and whitespace checks passed. No apps were redeployed. + ## One config directory Removed the three `deployments/` wrappers. The deployment script now lists `src/lilo/configs/` directly, and the validated model revisions live in the 9B LoRA and 4B FFT configs. Derived configs inherit those revisions. Authoring configs are excluded from runtime source fingerprints and shared-app/trainer/inference source mounts; workers receive computed settings as JSON. diff --git a/src/lilo/configs/qwen35_35b_a3b_fft_64k.py b/src/lilo/configs/qwen35_35b_a3b_fft_64k.py index 23508f4..52c653a 100644 --- a/src/lilo/configs/qwen35_35b_a3b_fft_64k.py +++ b/src/lilo/configs/qwen35_35b_a3b_fft_64k.py @@ -1,4 +1,3 @@ -from dataclasses import dataclass, field from lilo.deployments import ( BaseConfig, EngineOptions, @@ -11,71 +10,62 @@ ) -@dataclass(kw_only=True) class Config(BaseConfig): - name: str = "qwen35-35b-a3b-fft-64k" - model: Model = field( - default_factory=lambda: Model( - id="Qwen/Qwen3.5-35B-A3B", parameterization="full", max_context_length=65536 - ) + name = "qwen35-35b-a3b-fft-64k" + model = Model( + id="Qwen/Qwen3.5-35B-A3B", parameterization="full", max_context_length=65536 ) - routing: Routing = field(default_factory=lambda: Routing(default=True)) - trainer: Trainer = field( - default_factory=lambda: Trainer( - backend="megatron", - resources=Resources(gpu="H200:8"), - engine=EngineOptions( - max_clients_per_instance=1, sampler_persistence_concurrency=1 - ), - env={ - "PYTORCH_CUDA_ALLOC_CONF": "expandable_segments:True", - "TORCHINDUCTOR_COMPILE_THREADS": "1", + routing = Routing(default=True) + trainer = Trainer( + backend="megatron", + resources=Resources(gpu="H200:8"), + engine=EngineOptions( + max_clients_per_instance=1, sampler_persistence_concurrency=1 + ), + env={ + "PYTORCH_CUDA_ALLOC_CONF": "expandable_segments:True", + "TORCHINDUCTOR_COMPILE_THREADS": "1", + }, + config={ + "runtime": { + "tensor_model_parallel_size": 4, + "pipeline_model_parallel_size": 1, + "context_parallel_size": 2, + "expert_model_parallel_size": 8, + "expert_tensor_parallel_size": 1, + "sequence_parallel": True, + "micro_batch_size": 1, + "max_tokens_per_microbatch": 65536, + "bf16": True, + "fp16": False, + "gpu_memory_fraction": 0.9, + "use_distributed_optimizer": True, }, - config={ - "runtime": { - "tensor_model_parallel_size": 4, - "pipeline_model_parallel_size": 1, - "context_parallel_size": 2, - "expert_model_parallel_size": 8, - "expert_tensor_parallel_size": 1, - "sequence_parallel": True, - "micro_batch_size": 1, - "max_tokens_per_microbatch": 65536, - "bf16": True, - "fp16": False, - "gpu_memory_fraction": 0.9, - "use_distributed_optimizer": True, - }, - "provider": { - "mtp_num_layers": 0, - "recompute_granularity": "selective", - "moe_layer_recompute": True, - "moe_token_dispatcher_type": "alltoall", - "moe_router_fusion": True, - "moe_permute_fusion": True, - "moe_grouped_gemm": True, - "moe_shared_expert_overlap": False, - "moe_aux_loss_coeff": 0.0, - }, - "optimizer": {"optimizer": "adam", "lr": 0.0001, "min_lr": 0.0001}, + "provider": { + "mtp_num_layers": 0, + "recompute_granularity": "selective", + "moe_layer_recompute": True, + "moe_token_dispatcher_type": "alltoall", + "moe_router_fusion": True, + "moe_permute_fusion": True, + "moe_grouped_gemm": True, + "moe_shared_expert_overlap": False, + "moe_aux_loss_coeff": 0.0, }, - ) + "optimizer": {"optimizer": "adam", "lr": 0.0001, "min_lr": 0.0001}, + }, ) - inference: Inference = field( - default_factory=lambda: Inference( - resources=Resources(gpu="H200:4"), - scaling=InferenceScaling( - min_replicas=0, max_replicas=8, target_concurrency=16 - ), - config={ - "tp_size": 4, - "ep_size": 4, - "mem_fraction_static": 0.9, - "max_running_requests": 32, - "max_queued_requests": 4, - "cpu_weight_cache_max_compile_group_gb": 32, - "dp_size": 4, - "enable_dp_attention": True, - }, - ) + inference = Inference( + resources=Resources(gpu="H200:4"), + scaling=InferenceScaling(min_replicas=0, max_replicas=8, target_concurrency=16), + config={ + "tp_size": 4, + "ep_size": 4, + "mem_fraction_static": 0.9, + "max_running_requests": 32, + "max_queued_requests": 4, + "cpu_weight_cache_max_compile_group_gb": 32, + "dp_size": 4, + "enable_dp_attention": True, + }, ) diff --git a/src/lilo/configs/qwen35_4b_fft_64k.py b/src/lilo/configs/qwen35_4b_fft_64k.py index 3a1b45e..35fd111 100644 --- a/src/lilo/configs/qwen35_4b_fft_64k.py +++ b/src/lilo/configs/qwen35_4b_fft_64k.py @@ -1,4 +1,3 @@ -from dataclasses import dataclass, field from lilo.deployments import ( BaseConfig, Deployment, @@ -11,58 +10,49 @@ ) -@dataclass(kw_only=True) class Config(BaseConfig): - name: str = "qwen35-4b-fft-64k" - model: Model = field( - default_factory=lambda: Model( - id="Qwen/Qwen3.5-4B", - revision="851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", - parameterization="full", - max_context_length=65536, - ) + name = "qwen35-4b-fft-64k" + model = Model( + id="Qwen/Qwen3.5-4B", + revision="851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", + parameterization="full", + max_context_length=65536, ) - routing: Routing = field(default_factory=lambda: Routing(default=True)) - deployment: Deployment = field( - default_factory=lambda: Deployment(frontend="lilo-yaml") - ) - trainer: Trainer = field( - default_factory=lambda: Trainer( - backend="megatron", - resources=Resources(gpu="H100:4"), - engine=EngineOptions( - max_clients_per_instance=1, sampler_persistence_concurrency=1 - ), - config={ - "runtime": { - "tensor_model_parallel_size": 2, - "context_parallel_size": 2, - "sequence_parallel": True, - "micro_batch_size": 1, - "max_tokens_per_microbatch": 65536, - "defer_fp32_logits": True, - "fp32_lm_head": True, - "use_distributed_optimizer": True, - }, - "provider": { - "mtp_num_layers": 0, - "recompute_granularity": "full", - "recompute_method": "uniform", - "recompute_num_layers": 1, - }, - "optimizer": {"lr": 0.0001, "min_lr": 0.0001, "loss_scale": 1.0}, + routing = Routing(default=True) + deployment = Deployment(frontend="lilo-yaml") + trainer = Trainer( + backend="megatron", + resources=Resources(gpu="H100:4"), + engine=EngineOptions( + max_clients_per_instance=1, sampler_persistence_concurrency=1 + ), + config={ + "runtime": { + "tensor_model_parallel_size": 2, + "context_parallel_size": 2, + "sequence_parallel": True, + "micro_batch_size": 1, + "max_tokens_per_microbatch": 65536, + "defer_fp32_logits": True, + "fp32_lm_head": True, + "use_distributed_optimizer": True, }, - ) - ) - inference: Inference = field( - default_factory=lambda: Inference( - resources=Resources(gpu="H100:1"), - config={ - "tp_size": 1, - "mem_fraction_static": 0.85, - "max_running_requests": 32, - "max_queued_requests": 4, - "cpu_weight_cache_max_compile_group_gb": 16, + "provider": { + "mtp_num_layers": 0, + "recompute_granularity": "full", + "recompute_method": "uniform", + "recompute_num_layers": 1, }, - ) + "optimizer": {"lr": 0.0001, "min_lr": 0.0001, "loss_scale": 1.0}, + }, + ) + inference = Inference( + resources=Resources(gpu="H100:1"), + config={ + "tp_size": 1, + "mem_fraction_static": 0.85, + "max_running_requests": 32, + "max_queued_requests": 4, + "cpu_weight_cache_max_compile_group_gb": 16, + }, ) diff --git a/src/lilo/configs/qwen35_9b_fft_64k.py b/src/lilo/configs/qwen35_9b_fft_64k.py index 26e0af4..563512f 100644 --- a/src/lilo/configs/qwen35_9b_fft_64k.py +++ b/src/lilo/configs/qwen35_9b_fft_64k.py @@ -1,4 +1,3 @@ -from dataclasses import dataclass, field from lilo.deployments import ( BaseConfig, EngineOptions, @@ -11,60 +10,51 @@ ) -@dataclass(kw_only=True) class Config(BaseConfig): - name: str = "qwen35-9b-fft-64k" - model: Model = field( - default_factory=lambda: Model( - id="Qwen/Qwen3.5-9B", parameterization="full", max_context_length=65536 - ) + name = "qwen35-9b-fft-64k" + model = Model( + id="Qwen/Qwen3.5-9B", parameterization="full", max_context_length=65536 ) - routing: Routing = field(default_factory=lambda: Routing(default=True)) - trainer: Trainer = field( - default_factory=lambda: Trainer( - backend="megatron", - resources=Resources(gpu="H200:4"), - engine=EngineOptions( - max_clients_per_instance=1, sampler_persistence_concurrency=1 - ), - env={ - "PYTORCH_CUDA_ALLOC_CONF": "expandable_segments:True", - "TORCHINDUCTOR_COMPILE_THREADS": "1", + routing = Routing(default=True) + trainer = Trainer( + backend="megatron", + resources=Resources(gpu="H200:4"), + engine=EngineOptions( + max_clients_per_instance=1, sampler_persistence_concurrency=1 + ), + env={ + "PYTORCH_CUDA_ALLOC_CONF": "expandable_segments:True", + "TORCHINDUCTOR_COMPILE_THREADS": "1", + }, + config={ + "runtime": { + "tensor_model_parallel_size": 2, + "context_parallel_size": 2, + "sequence_parallel": True, + "micro_batch_size": 1, + "max_tokens_per_microbatch": 65536, + "defer_fp32_logits": True, + "fp32_lm_head": True, + "use_distributed_optimizer": True, }, - config={ - "runtime": { - "tensor_model_parallel_size": 2, - "context_parallel_size": 2, - "sequence_parallel": True, - "micro_batch_size": 1, - "max_tokens_per_microbatch": 65536, - "defer_fp32_logits": True, - "fp32_lm_head": True, - "use_distributed_optimizer": True, - }, - "provider": { - "mtp_num_layers": 0, - "recompute_granularity": "full", - "recompute_method": "uniform", - "recompute_num_layers": 1, - }, - "optimizer": {"lr": 0.0001, "min_lr": 0.0001, "loss_scale": 1.0}, + "provider": { + "mtp_num_layers": 0, + "recompute_granularity": "full", + "recompute_method": "uniform", + "recompute_num_layers": 1, }, - ) + "optimizer": {"lr": 0.0001, "min_lr": 0.0001, "loss_scale": 1.0}, + }, ) - inference: Inference = field( - default_factory=lambda: Inference( - resources=Resources(gpu="H200:1"), - scaling=InferenceScaling( - min_replicas=0, max_replicas=8, target_concurrency=16 - ), - config={ - "tp_size": 1, - "ep_size": 1, - "mem_fraction_static": 0.85, - "max_running_requests": 32, - "max_queued_requests": 4, - "cpu_weight_cache_max_compile_group_gb": 16, - }, - ) + inference = Inference( + resources=Resources(gpu="H200:1"), + scaling=InferenceScaling(min_replicas=0, max_replicas=8, target_concurrency=16), + config={ + "tp_size": 1, + "ep_size": 1, + "mem_fraction_static": 0.85, + "max_running_requests": 32, + "max_queued_requests": 4, + "cpu_weight_cache_max_compile_group_gb": 16, + }, ) diff --git a/src/lilo/configs/qwen35_9b_instruct_lora_16k.py b/src/lilo/configs/qwen35_9b_instruct_lora_16k.py index a156970..28399ec 100644 --- a/src/lilo/configs/qwen35_9b_instruct_lora_16k.py +++ b/src/lilo/configs/qwen35_9b_instruct_lora_16k.py @@ -1,4 +1,3 @@ -from dataclasses import dataclass, field from lilo.deployments import ( BaseConfig, EngineOptions, @@ -11,71 +10,62 @@ ) -@dataclass(kw_only=True) class Config(BaseConfig): - name: str = "qwen35-9b-instruct-lora-16k" - model: Model = field( - default_factory=lambda: Model( - id="Qwen/Qwen3.5-9B", parameterization="lora", max_context_length=16384 - ) + name = "qwen35-9b-instruct-lora-16k" + model = Model( + id="Qwen/Qwen3.5-9B", parameterization="lora", max_context_length=16384 ) - routing: Routing = field(default_factory=lambda: Routing(default=True)) - trainer: Trainer = field( - default_factory=lambda: Trainer( - backend="miles", - resources=Resources(gpu="H100:8"), - engine=EngineOptions( - max_clients_per_instance=6, sampler_persistence_concurrency=8 - ), - env={ - "PYTORCH_CUDA_ALLOC_CONF": "expandable_segments:True", - "TORCHINDUCTOR_COMPILE_THREADS": "1", - }, - config={ - "model_args": "qwen3.5-9B", - "options": { - "tensor_model_parallel_size": 8, - "target_modules": [ - "linear_qkv", - "linear_proj", - "linear_fc1", - "linear_fc2", - ], - "max_tokens_per_gpu": 16384, - "recompute_granularity": "full", - "recompute_method": "uniform", - "recompute_num_layers": 1, - "multi_lora_n_adapters": 6, - "lora_rank": 32, - "lora_alpha": 32, - }, - }, - ) - ) - inference: Inference = field( - default_factory=lambda: Inference( - resources=Resources(gpu="H200:1"), - scaling=InferenceScaling( - min_replicas=0, max_replicas=8, target_concurrency=16 - ), - config={ - "tp_size": 1, - "ep_size": 1, - "mem_fraction_static": 0.8, - "max_running_requests": 32, - "max_queued_requests": 8, - "max_loaded_loras": 256, - "max_loras_per_batch": 8, - "lora_target_modules": [ - "q_proj", - "k_proj", - "v_proj", - "o_proj", - "gate_proj", - "up_proj", - "down_proj", + routing = Routing(default=True) + trainer = Trainer( + backend="miles", + resources=Resources(gpu="H100:8"), + engine=EngineOptions( + max_clients_per_instance=6, sampler_persistence_concurrency=8 + ), + env={ + "PYTORCH_CUDA_ALLOC_CONF": "expandable_segments:True", + "TORCHINDUCTOR_COMPILE_THREADS": "1", + }, + config={ + "model_args": "qwen3.5-9B", + "options": { + "tensor_model_parallel_size": 8, + "target_modules": [ + "linear_qkv", + "linear_proj", + "linear_fc1", + "linear_fc2", ], - "schedule_policy": "lpm", + "max_tokens_per_gpu": 16384, + "recompute_granularity": "full", + "recompute_method": "uniform", + "recompute_num_layers": 1, + "multi_lora_n_adapters": 6, + "lora_rank": 32, + "lora_alpha": 32, }, - ) + }, + ) + inference = Inference( + resources=Resources(gpu="H200:1"), + scaling=InferenceScaling(min_replicas=0, max_replicas=8, target_concurrency=16), + config={ + "tp_size": 1, + "ep_size": 1, + "mem_fraction_static": 0.8, + "max_running_requests": 32, + "max_queued_requests": 8, + "max_loaded_loras": 256, + "max_loras_per_batch": 8, + "lora_target_modules": [ + "q_proj", + "k_proj", + "v_proj", + "o_proj", + "gate_proj", + "up_proj", + "down_proj", + ], + "schedule_policy": "lpm", + }, ) diff --git a/src/lilo/configs/qwen35_9b_instruct_lora_16k_dp2.py b/src/lilo/configs/qwen35_9b_instruct_lora_16k_dp2.py index 4c44452..00b460a 100644 --- a/src/lilo/configs/qwen35_9b_instruct_lora_16k_dp2.py +++ b/src/lilo/configs/qwen35_9b_instruct_lora_16k_dp2.py @@ -1,11 +1,9 @@ -from dataclasses import dataclass from lilo.configs.qwen35_9b_instruct_lora_16k import Config as ParentConfig -@dataclass(kw_only=True) class Config(ParentConfig): - name: str = "qwen35-9b-instruct-lora-16k-dp2" - - def __post_init__(self): - self.routing.default = False - self.trainer.config["options"]["tensor_model_parallel_size"] = 4 + name = "qwen35-9b-instruct-lora-16k-dp2" + overrides = { + "routing.default": False, + "trainer.config.options.tensor_model_parallel_size": 4, + } diff --git a/src/lilo/configs/qwen35_9b_lora_16k.py b/src/lilo/configs/qwen35_9b_lora_16k.py index 637a8cb..a885981 100644 --- a/src/lilo/configs/qwen35_9b_lora_16k.py +++ b/src/lilo/configs/qwen35_9b_lora_16k.py @@ -1,4 +1,3 @@ -from dataclasses import dataclass, field from lilo.deployments import ( BaseConfig, Deployment, @@ -13,69 +12,58 @@ ) -@dataclass(kw_only=True) class Config(BaseConfig): - name: str = "qwen35-9b-lora-16k" - model: Model = field( - default_factory=lambda: Model( - id="Qwen/Qwen3.5-9B-Base", - revision="68c46c4b3498877f3ef123c856ecfde50c39f404", - parameterization="lora", - max_context_length=16384, - ) + name = "qwen35-9b-lora-16k" + model = Model( + id="Qwen/Qwen3.5-9B-Base", + revision="68c46c4b3498877f3ef123c856ecfde50c39f404", + parameterization="lora", + max_context_length=16384, ) - routing: Routing = field(default_factory=lambda: Routing(default=True)) - deployment: Deployment = field( - default_factory=lambda: Deployment(frontend="lilo-yaml", mode="shared") - ) - trainer: Trainer = field( - default_factory=lambda: Trainer( - backend="miles", - resources=Resources(gpu="H100:4", cpu=16, memory_mib=65536), - scaling=TrainerScaling(max_instances=1), - engine=EngineOptions( - max_clients_per_instance=6, sampler_persistence_concurrency=8 - ), - env={ - "PYTORCH_CUDA_ALLOC_CONF": "expandable_segments:True", - "TORCHINDUCTOR_COMPILE_THREADS": "1", - }, - config={ - "model_args": "qwen3.5-9B", - "options": { - "tensor_model_parallel_size": 4, - "multi_lora_n_adapters": 6, - "lora_rank": 32, - "lora_alpha": 32, - "target_modules": [ - "linear_qkv", - "linear_proj", - "linear_fc1", - "linear_fc2", - "output_layer", - ], - "max_tokens_per_gpu": 16384, - "recompute_granularity": "full", - "recompute_method": "uniform", - "recompute_num_layers": 1, - }, + routing = Routing(default=True) + deployment = Deployment(frontend="lilo-yaml", mode="shared") + trainer = Trainer( + backend="miles", + resources=Resources(gpu="H100:4", cpu=16, memory_mib=65536), + scaling=TrainerScaling(max_instances=1), + engine=EngineOptions( + max_clients_per_instance=6, sampler_persistence_concurrency=8 + ), + env={ + "PYTORCH_CUDA_ALLOC_CONF": "expandable_segments:True", + "TORCHINDUCTOR_COMPILE_THREADS": "1", + }, + config={ + "model_args": "qwen3.5-9B", + "options": { + "tensor_model_parallel_size": 4, + "multi_lora_n_adapters": 6, + "lora_rank": 32, + "lora_alpha": 32, + "target_modules": [ + "linear_qkv", + "linear_proj", + "linear_fc1", + "linear_fc2", + "output_layer", + ], + "max_tokens_per_gpu": 16384, + "recompute_granularity": "full", + "recompute_method": "uniform", + "recompute_num_layers": 1, }, - ) + }, ) - inference: Inference = field( - default_factory=lambda: Inference( - backend="sglang", - resources=Resources(gpu="H200:1"), - scaling=InferenceScaling( - min_replicas=0, max_replicas=8, target_concurrency=16 - ), - config={ - "tp_size": 1, - "mem_fraction_static": 0.8, - "max_running_requests": 32, - "max_queued_requests": 8, - "max_loaded_loras": 64, - "max_loras_per_batch": 8, - }, - ) + inference = Inference( + backend="sglang", + resources=Resources(gpu="H200:1"), + scaling=InferenceScaling(min_replicas=0, max_replicas=8, target_concurrency=16), + config={ + "tp_size": 1, + "mem_fraction_static": 0.8, + "max_running_requests": 32, + "max_queued_requests": 8, + "max_loaded_loras": 64, + "max_loras_per_batch": 8, + }, ) diff --git a/src/lilo/configs/qwen35_9b_lora_16k_single.py b/src/lilo/configs/qwen35_9b_lora_16k_single.py index c60b684..459e2d9 100644 --- a/src/lilo/configs/qwen35_9b_lora_16k_single.py +++ b/src/lilo/configs/qwen35_9b_lora_16k_single.py @@ -1,11 +1,6 @@ -from dataclasses import dataclass from lilo.configs.qwen35_9b_lora_16k import Config as ParentConfig -@dataclass(kw_only=True) class Config(ParentConfig): - name: str = "qwen35-9b-lora-16k-single" - - def __post_init__(self): - self.routing.default = False - self.trainer.engine.max_clients_per_instance = 1 + name = "qwen35-9b-lora-16k-single" + overrides = {"routing.default": False, "trainer.engine.max_clients_per_instance": 1} diff --git a/src/lilo/configs/qwen35_9b_lora_2k.py b/src/lilo/configs/qwen35_9b_lora_2k.py index f2bc407..ba84d3e 100644 --- a/src/lilo/configs/qwen35_9b_lora_2k.py +++ b/src/lilo/configs/qwen35_9b_lora_2k.py @@ -1,4 +1,3 @@ -from dataclasses import dataclass, field from lilo.deployments import ( BaseConfig, EngineOptions, @@ -11,73 +10,64 @@ ) -@dataclass(kw_only=True) class Config(BaseConfig): - name: str = "qwen35-9b-lora-2k" - model: Model = field( - default_factory=lambda: Model( - id="Qwen/Qwen3.5-9B-Base", parameterization="lora", max_context_length=2048 - ) + name = "qwen35-9b-lora-2k" + model = Model( + id="Qwen/Qwen3.5-9B-Base", parameterization="lora", max_context_length=2048 ) - routing: Routing = field(default_factory=lambda: Routing(default=False)) - trainer: Trainer = field( - default_factory=lambda: Trainer( - backend="miles", - resources=Resources(gpu="H200:4"), - engine=EngineOptions( - max_clients_per_instance=4, sampler_persistence_concurrency=8 - ), - env={ - "PYTORCH_CUDA_ALLOC_CONF": "expandable_segments:True", - "TORCHINDUCTOR_COMPILE_THREADS": "1", - }, - config={ - "model_args": "qwen3.5-9B", - "options": { - "tensor_model_parallel_size": 4, - "target_modules": [ - "linear_qkv", - "linear_proj", - "linear_fc1", - "linear_fc2", - "output_layer", - ], - "max_tokens_per_gpu": 2048, - "recompute_granularity": "full", - "recompute_method": "uniform", - "recompute_num_layers": 1, - "multi_lora_n_adapters": 4, - "lora_rank": 32, - "lora_alpha": 32, - }, - }, - ) - ) - inference: Inference = field( - default_factory=lambda: Inference( - resources=Resources(gpu="H200:1"), - scaling=InferenceScaling( - min_replicas=0, max_replicas=8, target_concurrency=16 - ), - config={ - "tp_size": 1, - "ep_size": 1, - "mem_fraction_static": 0.8, - "max_running_requests": 32, - "max_queued_requests": 8, - "max_loaded_loras": 32, - "max_loras_per_batch": 8, - "lora_target_modules": [ - "q_proj", - "k_proj", - "v_proj", - "o_proj", - "gate_proj", - "up_proj", - "down_proj", - "lm_head", + routing = Routing(default=False) + trainer = Trainer( + backend="miles", + resources=Resources(gpu="H200:4"), + engine=EngineOptions( + max_clients_per_instance=4, sampler_persistence_concurrency=8 + ), + env={ + "PYTORCH_CUDA_ALLOC_CONF": "expandable_segments:True", + "TORCHINDUCTOR_COMPILE_THREADS": "1", + }, + config={ + "model_args": "qwen3.5-9B", + "options": { + "tensor_model_parallel_size": 4, + "target_modules": [ + "linear_qkv", + "linear_proj", + "linear_fc1", + "linear_fc2", + "output_layer", ], - "schedule_policy": "lpm", + "max_tokens_per_gpu": 2048, + "recompute_granularity": "full", + "recompute_method": "uniform", + "recompute_num_layers": 1, + "multi_lora_n_adapters": 4, + "lora_rank": 32, + "lora_alpha": 32, }, - ) + }, + ) + inference = Inference( + resources=Resources(gpu="H200:1"), + scaling=InferenceScaling(min_replicas=0, max_replicas=8, target_concurrency=16), + config={ + "tp_size": 1, + "ep_size": 1, + "mem_fraction_static": 0.8, + "max_running_requests": 32, + "max_queued_requests": 8, + "max_loaded_loras": 32, + "max_loras_per_batch": 8, + "lora_target_modules": [ + "q_proj", + "k_proj", + "v_proj", + "o_proj", + "gate_proj", + "up_proj", + "down_proj", + "lm_head", + ], + "schedule_policy": "lpm", + }, ) diff --git a/src/lilo/configs/qwen35_9b_lora_64k.py b/src/lilo/configs/qwen35_9b_lora_64k.py index 077f58a..7420a6a 100644 --- a/src/lilo/configs/qwen35_9b_lora_64k.py +++ b/src/lilo/configs/qwen35_9b_lora_64k.py @@ -1,14 +1,12 @@ -from dataclasses import dataclass from lilo.configs.qwen35_9b_lora_16k import Config as ParentConfig -@dataclass(kw_only=True) class Config(ParentConfig): - name: str = "qwen35-9b-lora-64k" - - def __post_init__(self): - self.model.max_context_length = 65536 - self.routing.default = False - self.trainer.resources.gpu = "H200:8" - self.trainer.config["options"]["tensor_model_parallel_size"] = 8 - self.trainer.config["options"]["max_tokens_per_gpu"] = 65536 + name = "qwen35-9b-lora-64k" + overrides = { + "model.max_context_length": 65536, + "routing.default": False, + "trainer.resources.gpu": "H200:8", + "trainer.config.options.tensor_model_parallel_size": 8, + "trainer.config.options.max_tokens_per_gpu": 65536, + } diff --git a/src/lilo/configs/qwen36_27b_fft_64k.py b/src/lilo/configs/qwen36_27b_fft_64k.py index 857e69f..e8e7c0f 100644 --- a/src/lilo/configs/qwen36_27b_fft_64k.py +++ b/src/lilo/configs/qwen36_27b_fft_64k.py @@ -1,4 +1,3 @@ -from dataclasses import dataclass, field from lilo.deployments import ( BaseConfig, EngineOptions, @@ -11,62 +10,53 @@ ) -@dataclass(kw_only=True) class Config(BaseConfig): - name: str = "qwen36-27b-fft-64k" - model: Model = field( - default_factory=lambda: Model( - id="Qwen/Qwen3.6-27B", parameterization="full", max_context_length=65536 - ) + name = "qwen36-27b-fft-64k" + model = Model( + id="Qwen/Qwen3.6-27B", parameterization="full", max_context_length=65536 ) - routing: Routing = field(default_factory=lambda: Routing(default=True)) - trainer: Trainer = field( - default_factory=lambda: Trainer( - backend="megatron", - resources=Resources(gpu="H200:8"), - engine=EngineOptions( - max_clients_per_instance=1, sampler_persistence_concurrency=1 - ), - env={ - "PYTORCH_CUDA_ALLOC_CONF": "expandable_segments:True", - "TORCHINDUCTOR_COMPILE_THREADS": "1", + routing = Routing(default=True) + trainer = Trainer( + backend="megatron", + resources=Resources(gpu="H200:8"), + engine=EngineOptions( + max_clients_per_instance=1, sampler_persistence_concurrency=1 + ), + env={ + "PYTORCH_CUDA_ALLOC_CONF": "expandable_segments:True", + "TORCHINDUCTOR_COMPILE_THREADS": "1", + }, + config={ + "runtime": { + "tensor_model_parallel_size": 4, + "pipeline_model_parallel_size": 1, + "context_parallel_size": 2, + "sequence_parallel": True, + "micro_batch_size": 1, + "max_tokens_per_microbatch": 65536, + "bf16": True, + "fp16": False, + "gpu_memory_fraction": 0.9, + "use_distributed_optimizer": True, }, - config={ - "runtime": { - "tensor_model_parallel_size": 4, - "pipeline_model_parallel_size": 1, - "context_parallel_size": 2, - "sequence_parallel": True, - "micro_batch_size": 1, - "max_tokens_per_microbatch": 65536, - "bf16": True, - "fp16": False, - "gpu_memory_fraction": 0.9, - "use_distributed_optimizer": True, - }, - "provider": { - "mtp_num_layers": 0, - "recompute_granularity": "full", - "recompute_method": "uniform", - "recompute_num_layers": 1, - }, - "optimizer": {"optimizer": "adam", "lr": 0.0001, "min_lr": 0.0001}, + "provider": { + "mtp_num_layers": 0, + "recompute_granularity": "full", + "recompute_method": "uniform", + "recompute_num_layers": 1, }, - ) + "optimizer": {"optimizer": "adam", "lr": 0.0001, "min_lr": 0.0001}, + }, ) - inference: Inference = field( - default_factory=lambda: Inference( - resources=Resources(gpu="H200:4"), - scaling=InferenceScaling( - min_replicas=0, max_replicas=8, target_concurrency=16 - ), - config={ - "tp_size": 4, - "ep_size": 1, - "mem_fraction_static": 0.9, - "max_running_requests": 32, - "max_queued_requests": 4, - "cpu_weight_cache_max_compile_group_gb": 32, - }, - ) + inference = Inference( + resources=Resources(gpu="H200:4"), + scaling=InferenceScaling(min_replicas=0, max_replicas=8, target_concurrency=16), + config={ + "tp_size": 4, + "ep_size": 1, + "mem_fraction_static": 0.9, + "max_running_requests": 32, + "max_queued_requests": 4, + "cpu_weight_cache_max_compile_group_gb": 32, + }, ) diff --git a/src/lilo/configs/qwen36_35b_a3b_fft_64k.py b/src/lilo/configs/qwen36_35b_a3b_fft_64k.py index 0c157bd..8d29075 100644 --- a/src/lilo/configs/qwen36_35b_a3b_fft_64k.py +++ b/src/lilo/configs/qwen36_35b_a3b_fft_64k.py @@ -1,12 +1,10 @@ -from dataclasses import dataclass from lilo.configs.qwen35_35b_a3b_fft_64k import Config as ParentConfig -@dataclass(kw_only=True) class Config(ParentConfig): - name: str = "qwen36-35b-a3b-fft-64k" - - def __post_init__(self): - self.model.id = "Qwen/Qwen3.6-35B-A3B" - self.inference.config["dp_size"] = 1 - self.inference.config["enable_dp_attention"] = False + name = "qwen36-35b-a3b-fft-64k" + overrides = { + "model.id": "Qwen/Qwen3.6-35B-A3B", + "inference.config.dp_size": 1, + "inference.config.enable_dp_attention": False, + } diff --git a/src/lilo/configs/qwen38_27b_lora_128k.py b/src/lilo/configs/qwen38_27b_lora_128k.py index 889b7bf..f6af875 100644 --- a/src/lilo/configs/qwen38_27b_lora_128k.py +++ b/src/lilo/configs/qwen38_27b_lora_128k.py @@ -1,19 +1,17 @@ -from dataclasses import dataclass from lilo.configs.qwen38_27b_lora_16k import Config as ParentConfig -@dataclass(kw_only=True) class Config(ParentConfig): - name: str = "qwen38-27b-lora-128k" - - def __post_init__(self): - self.model.max_context_length = 131072 - self.routing.default = False - self.trainer.config["options"]["tensor_model_parallel_size"] = 2 - self.trainer.config["options"]["context_parallel_size"] = 4 - self.trainer.config["options"]["max_tokens_per_gpu"] = 32768 - self.inference.resources.gpu = "H200:2" - self.inference.scaling.max_replicas = 4 - self.inference.scaling.target_concurrency = 4 - self.inference.config["tp_size"] = 2 - self.inference.config["max_running_requests"] = 8 + name = "qwen38-27b-lora-128k" + overrides = { + "model.max_context_length": 131072, + "routing.default": False, + "trainer.config.options.tensor_model_parallel_size": 2, + "trainer.config.options.context_parallel_size": 4, + "trainer.config.options.max_tokens_per_gpu": 32768, + "inference.resources.gpu": "H200:2", + "inference.scaling.max_replicas": 4, + "inference.scaling.target_concurrency": 4, + "inference.config.tp_size": 2, + "inference.config.max_running_requests": 8, + } diff --git a/src/lilo/configs/qwen38_27b_lora_16k.py b/src/lilo/configs/qwen38_27b_lora_16k.py index 237f2b1..f70cf24 100644 --- a/src/lilo/configs/qwen38_27b_lora_16k.py +++ b/src/lilo/configs/qwen38_27b_lora_16k.py @@ -1,4 +1,3 @@ -from dataclasses import dataclass, field from lilo.deployments import ( BaseConfig, EngineOptions, @@ -11,71 +10,62 @@ ) -@dataclass(kw_only=True) class Config(BaseConfig): - name: str = "qwen38-27b-lora-16k" - model: Model = field( - default_factory=lambda: Model( - id="Qwen/Qwen3.8-27B", parameterization="lora", max_context_length=16384 - ) + name = "qwen38-27b-lora-16k" + model = Model( + id="Qwen/Qwen3.8-27B", parameterization="lora", max_context_length=16384 ) - routing: Routing = field(default_factory=lambda: Routing(default=True)) - trainer: Trainer = field( - default_factory=lambda: Trainer( - backend="miles", - resources=Resources(gpu="H200:8"), - engine=EngineOptions( - max_clients_per_instance=6, sampler_persistence_concurrency=8 - ), - env={ - "PYTORCH_CUDA_ALLOC_CONF": "expandable_segments:True", - "TORCHINDUCTOR_COMPILE_THREADS": "1", - }, - config={ - "model_args": "qwen3.8-27B", - "options": { - "tensor_model_parallel_size": 4, - "target_modules": [ - "linear_qkv", - "linear_proj", - "linear_fc1", - "linear_fc2", - ], - "max_tokens_per_gpu": 16384, - "recompute_granularity": "full", - "recompute_method": "uniform", - "recompute_num_layers": 1, - "multi_lora_n_adapters": 6, - "lora_rank": 32, - "lora_alpha": 32, - }, - }, - ) - ) - inference: Inference = field( - default_factory=lambda: Inference( - resources=Resources(gpu="H200:1"), - scaling=InferenceScaling( - min_replicas=0, max_replicas=8, target_concurrency=16 - ), - config={ - "tp_size": 1, - "ep_size": 1, - "mem_fraction_static": 0.8, - "max_running_requests": 32, - "max_queued_requests": 8, - "max_loaded_loras": 256, - "max_loras_per_batch": 8, - "lora_target_modules": [ - "q_proj", - "k_proj", - "v_proj", - "o_proj", - "gate_proj", - "up_proj", - "down_proj", + routing = Routing(default=True) + trainer = Trainer( + backend="miles", + resources=Resources(gpu="H200:8"), + engine=EngineOptions( + max_clients_per_instance=6, sampler_persistence_concurrency=8 + ), + env={ + "PYTORCH_CUDA_ALLOC_CONF": "expandable_segments:True", + "TORCHINDUCTOR_COMPILE_THREADS": "1", + }, + config={ + "model_args": "qwen3.8-27B", + "options": { + "tensor_model_parallel_size": 4, + "target_modules": [ + "linear_qkv", + "linear_proj", + "linear_fc1", + "linear_fc2", ], - "schedule_policy": "lpm", + "max_tokens_per_gpu": 16384, + "recompute_granularity": "full", + "recompute_method": "uniform", + "recompute_num_layers": 1, + "multi_lora_n_adapters": 6, + "lora_rank": 32, + "lora_alpha": 32, }, - ) + }, + ) + inference = Inference( + resources=Resources(gpu="H200:1"), + scaling=InferenceScaling(min_replicas=0, max_replicas=8, target_concurrency=16), + config={ + "tp_size": 1, + "ep_size": 1, + "mem_fraction_static": 0.8, + "max_running_requests": 32, + "max_queued_requests": 8, + "max_loaded_loras": 256, + "max_loras_per_batch": 8, + "lora_target_modules": [ + "q_proj", + "k_proj", + "v_proj", + "o_proj", + "gate_proj", + "up_proj", + "down_proj", + ], + "schedule_policy": "lpm", + }, ) diff --git a/src/lilo/configs/qwen38_27b_lora_64k.py b/src/lilo/configs/qwen38_27b_lora_64k.py index d50e7a5..2bb796d 100644 --- a/src/lilo/configs/qwen38_27b_lora_64k.py +++ b/src/lilo/configs/qwen38_27b_lora_64k.py @@ -1,15 +1,13 @@ -from dataclasses import dataclass from lilo.configs.qwen38_27b_lora_16k import Config as ParentConfig -@dataclass(kw_only=True) class Config(ParentConfig): - name: str = "qwen38-27b-lora-64k" - - def __post_init__(self): - self.model.max_context_length = 65536 - self.routing.default = False - self.trainer.config["options"]["context_parallel_size"] = 2 - self.trainer.config["options"]["max_tokens_per_gpu"] = 32768 - self.inference.scaling.target_concurrency = 8 - self.inference.config["max_running_requests"] = 16 + name = "qwen38-27b-lora-64k" + overrides = { + "model.max_context_length": 65536, + "routing.default": False, + "trainer.config.options.context_parallel_size": 2, + "trainer.config.options.max_tokens_per_gpu": 32768, + "inference.scaling.target_concurrency": 8, + "inference.config.max_running_requests": 16, + } diff --git a/src/lilo/deployment_cli.py b/src/lilo/deployment_cli.py index 409e629..c846d3c 100644 --- a/src/lilo/deployment_cli.py +++ b/src/lilo/deployment_cli.py @@ -170,7 +170,8 @@ def parser(): if name == "resolve": cmd.add_argument("--output") apply = commands.add_parser( - "deploy", help="Deploy the complete active Python config set behind one frontend" + "deploy", + help="Deploy the complete active Python config set behind one frontend", ) apply.add_argument("files", nargs="+") management = commands.add_parser("deployment").add_subparsers( @@ -195,12 +196,9 @@ def main(argv=None): if not config_path(args.preset).is_file(): raise ValueError(f"unknown example config: {args.preset}") print( - "from dataclasses import dataclass\n" f"from lilo.configs.{module} import Config as ParentConfig\n\n\n" - "@dataclass(kw_only=True)\n" "class Config(ParentConfig):\n" - " # Override fields or customize nested settings in __post_init__.\n" - " pass" + " overrides = {}" ) elif args.action == "validate": specs = [load(path) for path in args.files] diff --git a/src/lilo/deployments.py b/src/lilo/deployments.py index ae6f998..40e1b00 100644 --- a/src/lilo/deployments.py +++ b/src/lilo/deployments.py @@ -3,7 +3,7 @@ from __future__ import annotations from copy import deepcopy -from dataclasses import asdict, dataclass, field +from dataclasses import MISSING, asdict, dataclass, field, fields import hashlib from importlib.resources import files import json @@ -11,7 +11,7 @@ import re import runpy import sys -from typing import Any, Literal +from typing import Any, ClassVar, Literal from pydantic import BaseModel, ConfigDict @@ -123,9 +123,9 @@ class Lifecycle: sweep_interval_s: int = 300 -@dataclass(kw_only=True) +@dataclass(kw_only=True, init=False) class BaseConfig: - """Subclass in a config file and override defaults or use __post_init__.""" + """Declare class defaults and dotted overrides; each instance owns its values.""" name: str model: Model @@ -136,6 +136,50 @@ class BaseConfig: deployment: Deployment = field(default_factory=Deployment) lifecycle: Lifecycle = field(default_factory=Lifecycle) + overrides: ClassVar[dict[str, Any]] = {} + + def __init__(self, **kwargs): + definitions = {item.name: item for item in fields(BaseConfig)} + unknown = kwargs.keys() - definitions.keys() + if unknown: + raise TypeError(f"unknown config fields: {sorted(unknown)}") + values = {} + for name, item in definitions.items(): + if item.default_factory is not MISSING: + values[name] = item.default_factory() + elif item.default is not MISSING: + values[name] = deepcopy(item.default) + # Apply each parent's defaults and overrides before its child's. Copy at + # every assignment so instances never mutate class defaults or parents. + for cls in reversed(type(self).__mro__): + for name in definitions.keys() & vars(cls).keys(): + values[name] = deepcopy(vars(cls)[name]) + for path, value in vars(cls).get("overrides", {}).items(): + parts = path.split(".") + target = values + try: + for part in parts[:-1]: + target = ( + target[part] + if isinstance(target, dict) + else getattr(target, part) + ) + if isinstance(target, dict): + if len(parts) == 1 and parts[0] not in definitions: + raise KeyError(parts[0]) + target[parts[-1]] = deepcopy(value) + else: + # Reject misspelled dataclass fields. + getattr(target, parts[-1]) + setattr(target, parts[-1], deepcopy(value)) + except (KeyError, AttributeError) as exc: + raise ValueError(f"unknown config override: {path}") from exc + values.update(deepcopy(kwargs)) + missing = definitions.keys() - values.keys() + if missing: + raise TypeError(f"missing config fields: {sorted(missing)}") + self.__dict__.update(values) + class DeploymentRecord(BaseModel): """Saved deployment metadata around a Python configuration. diff --git a/tests/test_deployment_cli.py b/tests/test_deployment_cli.py index 9f5434d..8884cd5 100644 --- a/tests/test_deployment_cli.py +++ b/tests/test_deployment_cli.py @@ -123,12 +123,9 @@ def test_compile_pins_revision_at_external_boundary( path = tmp_path / "model.py" path.write_text( - "from dataclasses import dataclass\n" "from lilo.configs.qwen35_9b_lora_16k import Config as ParentConfig\n" - "@dataclass(kw_only=True)\n" "class Config(ParentConfig):\n" - " def __post_init__(self):\n" - f" self.model.revision = {revision!r}\n" + f" overrides = {{'model.revision': {revision!r}}}\n" ) lookup = Mock(return_value=SimpleNamespace(sha="a" * 40)) monkeypatch.setattr(huggingface_hub.HfApi, "model_info", lookup) diff --git a/tests/test_deployments.py b/tests/test_deployments.py index 64022e2..781bed8 100644 --- a/tests/test_deployments.py +++ b/tests/test_deployments.py @@ -378,14 +378,10 @@ def test_python_config_inheritance_and_independent_defaults(tmp_path): path = tmp_path / "model.py" path.write_text( - "from dataclasses import dataclass\n" "from lilo.configs.qwen35_9b_lora_64k import Config as ParentConfig\n" - "@dataclass(kw_only=True)\n" "class Config(ParentConfig):\n" - " name: str = 'custom'\n" - " def __post_init__(self):\n" - " super().__post_init__()\n" - " self.trainer.config['options']['new_backend_option'] = False\n" + " name = 'custom'\n" + " overrides = {'trainer.config.options.new_backend_option': False}\n" ) first, second = load(path), load(path) assert is_dataclass(first) @@ -440,3 +436,52 @@ def test_config_import_error_preserves_traceback_and_restores_path(tmp_path): def test_no_yaml_config_ingestion(tmp_path): with pytest.raises(ValueError, match="Python .py"): load(tmp_path / "old.yaml") + + +def test_overrides_inherit_replace_and_copy_values(): + from lilo.configs.qwen35_9b_lora_16k import Config as Example + from lilo.deployments import Trainer, Resources + + class Parent(Example): + overrides = { + "trainer.config.options.target_modules": ["parent"], + "trainer.config.options.future_option": {"enabled": True}, + "inference.config.max_running_requests": 24, + } + + class Child(Parent): + name = "child" + overrides = { + "trainer.config.options.target_modules": ["child"], + "trainer.config.options.future_option": {"enabled": False}, + } + + child = Child() + assert child.inference.config["max_running_requests"] == 24 + assert child.trainer.config["options"]["target_modules"] == ["child"] + assert child.trainer.config["options"]["future_option"] == {"enabled": False} + child.trainer.config["options"]["target_modules"].append("changed") + child.trainer.config["options"]["future_option"]["enabled"] = True + assert Child.overrides["trainer.config.options.target_modules"] == ["child"] + assert Child().trainer.config["options"]["future_option"] == {"enabled": False} + assert Parent().trainer.config["options"]["target_modules"] == ["parent"] + assert "overrides" not in asdict(child) + assert Child(name="keyword").name == "keyword" + + class Replacement(Parent): + trainer = Trainer(resources=Resources(gpu="H200:8"), config={"options": {}}) + overrides = {"trainer.config.options.new_option": 1} + + # A child's complete field replacement wins over its parent's dotted edits. + assert Replacement().trainer.config == {"options": {"new_option": 1}} + + +@pytest.mark.parametrize("path", ["model.typo", "trainer.missing.value", "typo"]) +def test_override_typos_fail_with_the_path(path): + from lilo.configs.qwen35_9b_lora_16k import Config as Example + + class Config(Example): + overrides = {path: 1} + + with pytest.raises(ValueError, match=path): + Config() From 4f349337160d1f3d223d67558f61738886f8eecf Mon Sep 17 00:00:00 2001 From: kailash Date: Wed, 23 Sep 2026 18:43:48 +0000 Subject: [PATCH 15/27] Replace deployment section templates with plain dictionaries --- docs/deployment-configs.md | 12 +- docs/deployment-validation.md | 6 + scripts/e2e_engine_definition.py | 14 +- src/lilo/backends/deployment.py | 4 +- src/lilo/backends/megatron_deployment.py | 14 +- src/lilo/backends/miles_deployment.py | 12 +- src/lilo/configs/qwen35_35b_a3b_fft_64k.py | 47 ++-- src/lilo/configs/qwen35_4b_fft_64k.py | 49 ++-- src/lilo/configs/qwen35_9b_fft_64k.py | 47 ++-- .../configs/qwen35_9b_instruct_lora_16k.py | 47 ++-- src/lilo/configs/qwen35_9b_lora_16k.py | 59 ++--- src/lilo/configs/qwen35_9b_lora_2k.py | 47 ++-- src/lilo/configs/qwen36_27b_fft_64k.py | 47 ++-- src/lilo/configs/qwen38_27b_lora_16k.py | 47 ++-- src/lilo/deployment_cli.py | 21 +- src/lilo/deployments.py | 246 +++++++----------- src/lilo/inference/sglang_deployment.py | 8 +- src/lilo/providers/modal/app.py | 20 +- src/lilo/providers/modal/deployment_apps.py | 118 ++++----- .../providers/modal/deployment_pool_app.py | 2 +- tests/providers/test_checkpoint_storage.py | 4 +- tests/providers/test_deployment_apps.py | 16 +- tests/providers/test_deployment_e2e_helper.py | 8 +- tests/providers/test_deployment_presets.py | 6 +- tests/test_deployment_cli.py | 4 +- tests/test_deployments.py | 88 ++++--- 26 files changed, 448 insertions(+), 545 deletions(-) diff --git a/docs/deployment-configs.md b/docs/deployment-configs.md index 9b51c43..331a598 100644 --- a/docs/deployment-configs.md +++ b/docs/deployment-configs.md @@ -2,7 +2,7 @@ Each deployment is a Python file exporting a `Config` class that inherits from `BaseConfig`. The class contains the model, trainer, inference and Modal settings. There is no YAML loader or model catalog. -The layout follows the Python recipe approach used by [training-gym](https://github.com/modal-labs/training-gym) and the [multinode training guide](https://github.com/modal-labs/multinode-training-guide/blob/main/nemo-rl/configs/llama3_1_8b_math_2node.py). Lilo uses dataclasses for the resulting settings, but config files need no decorators or default factories. Importing either project is not required. +The layout follows the Python recipe approach used by [training-gym](https://github.com/modal-labs/training-gym) and the [multinode training guide](https://github.com/modal-labs/multinode-training-guide/blob/main/nemo-rl/configs/llama3_1_8b_math_2node.py). Only the outer BaseConfig is a dataclass; its sections are ordinary dictionaries. Config files need no decorators or default factories. Importing either project is not required. ## Define a deployment @@ -29,9 +29,11 @@ class Config(ParentConfig): } ``` -New deployments can inherit directly from `BaseConfig` and declare ordinary class defaults such as `model = Model(...)` and `trainer = Trainer(...)`. The [16K example](../src/lilo/configs/qwen35_9b_lora_16k.py) shows the complete structure. No `@dataclass`, `field(default_factory=...)`, or `__post_init__` is needed in config files. +New deployments can inherit directly from `BaseConfig` and declare ordinary class defaults such as `model = {"id": "...", "max_context_length": 16384}` and `trainer = {"resources": {"gpu": "H100:4"}, "config": {...}}`. The [16K example](../src/lilo/configs/qwen35_9b_lora_16k.py) shows the complete structure. No `@dataclass`, `field(default_factory=...)`, or `__post_init__` is needed in config files. -`BaseConfig` copies defaults for each instance, then applies each parent's overrides before its child's. Dotted paths traverse dataclass fields and dictionary keys. A value replaces the selected field/key, including whole lists and dictionaries; other keys remain unchanged. Native backend dictionaries can receive new option names. Misspelled dataclass fields and missing intermediate paths raise an error naming the override. Constructor keywords, when supplied, replace top-level fields last. +`BaseConfig` copies defaults for each instance, then applies each parent's overrides before its child's. Dotted paths select a section and traverse its dictionary keys. A value replaces the selected field/key, including whole lists and dictionaries; other keys remain unchanged. Native backend dictionaries can receive new option names. Unknown top-level sections and missing intermediate paths raise an error naming the override. Section keys are open; their consumers read the options they need. Constructor keywords, when supplied, replace top-level fields last. + +Omitted orchestration settings come from one defaults dictionary in `deployments.py`. Backend options have no schema there. Replacing an entire section fills its omitted orchestration defaults; use dotted overrides to retain the parent’s other settings. Copying happens inside the base class, so edits to a config's nested lists or dictionaries do not change its parent, another instance, or the class's override dictionary. Only the resulting settings enter the saved deployment record; workers do not apply inheritance again. @@ -78,11 +80,11 @@ lilo deploy config.py [`load()`](../src/lilo/deployments.py) executes the file and instantiates its exported `Config` subclass. The returned dataclass goes directly to the orchestration code. There is no dictionary-to-config conversion on this path. -`DeploymentRecord` adds the code identity, pinned Miles revision, configuration hash and active status. JSON is used only to store records and pass them to other processes. Pydantic reconstructs the standard dataclasses when reading those records; it does not import or run the user's config file in a GPU worker. Records preserve the full computed settings rather than a reference to the original Python file. +`DeploymentRecord` adds the code identity, pinned Miles revision, configuration hash and active status. JSON is used only to store records and pass them to other processes. Pydantic reconstructs BaseConfig with its dictionary sections when reading those records; it does not import or run the user's config file in a GPU worker. Records preserve the full computed settings rather than a reference to the original Python file. | File | Responsibility | | --- | --- | -| [`deployments.py`](../src/lilo/deployments.py) | Dataclasses, Python file loader, saved record and shared-frontend checks | +| [`deployments.py`](../src/lilo/deployments.py) | BaseConfig defaults and overrides, Python file loader, saved record and shared-frontend checks | | [`deployment_cli.py`](../src/lilo/deployment_cli.py) | Revision lookup, manifest updates and Modal deployment | | [`deployment_apps.py`](../src/lilo/providers/modal/deployment_apps.py) | Trainer functions, inference server classes and process startup | | [`app.py`](../src/lilo/providers/modal/app.py) | Shared frontend and `app.include()` for generated trainers | diff --git a/docs/deployment-validation.md b/docs/deployment-validation.md index 9d1f383..75e5e54 100644 --- a/docs/deployment-validation.md +++ b/docs/deployment-validation.md @@ -2,6 +2,12 @@ These checks exercise the shared-app deployment path in PR #55. Each deployment contains the frontend and generated trainer functions; inference pools are separate apps started on demand. +## Plain dictionary sections + +Config files import only `BaseConfig`. Model, trainer, inference, resource, scaling, routing and deployment sections are plain dictionaries, consumed directly by backend and Modal code. The nested template classes have been removed. One defaults dictionary supplies omitted orchestration settings; inherited dotted overrides still work. + +Validation: **604 CPU tests passed, 1 skipped**. All 14 computed configs match their previous settings, build backend/inference options and round-trip through saved JSON. Tests cover independent mutable values, omitted defaults, inherited overrides and native backend option forwarding. Ruff and whitespace checks passed. No apps were redeployed. + ## Simple class defaults and inherited overrides Config files now declare ordinary class defaults and dotted `overrides` dictionaries. They contain no dataclass decorators, default factories or `__post_init__` methods. `BaseConfig` copies values per instance and applies parent defaults/overrides before child defaults/overrides. The resulting settings and saved-record format are unchanged for all 14 configs. diff --git a/scripts/e2e_engine_definition.py b/scripts/e2e_engine_definition.py index b4f7031..e400b66 100644 --- a/scripts/e2e_engine_definition.py +++ b/scripts/e2e_engine_definition.py @@ -16,7 +16,7 @@ import tinker from tinker import types -from lilo.deployments import DeploymentRecord +from lilo.deployments import DeploymentRecord, gpu_count from lilo.backends.deployment import backend_config TIMEOUT = 3 * 60 * 60 @@ -35,18 +35,18 @@ def _definition(frontend: str, name: str) -> tuple[Any, str]: ) resolved = matches[0] spec = resolved.spec - settings = backend_config(spec)[spec.trainer.backend] + settings = backend_config(spec)[spec.trainer['backend']] definition = SimpleNamespace( DEFINITION_ID=resolved.definition_id, - MODEL_NAME=spec.model.id, - PARAMETERIZATION=spec.model.parameterization, - MAX_CONTEXT_LENGTH=spec.model.max_context_length, + MODEL_NAME=spec.model['id'], + PARAMETERIZATION=spec.model['parameterization'], + MAX_CONTEXT_LENGTH=spec.model['max_context_length'], MAX_TOKENS_PER_MICROBATCH=settings.get( "max_tokens_per_microbatch", settings.get("max_tokens_per_gpu") ), MICRO_BATCH_SIZE=settings.get("micro_batch_size", 1), - GPU_TYPE=spec.trainer.resources.gpu.split(":")[0], - GPUS=spec.trainer.resources.gpu_count, + GPU_TYPE=spec.trainer['resources']['gpu'].split(":")[0], + GPUS=gpu_count(spec.trainer['resources']), LORA_RANK=settings.get("max_lora_rank"), ) return definition, definition.PARAMETERIZATION diff --git a/src/lilo/backends/deployment.py b/src/lilo/backends/deployment.py index baec4f0..f2b9a98 100644 --- a/src/lilo/backends/deployment.py +++ b/src/lilo/backends/deployment.py @@ -21,8 +21,8 @@ def _reader(registry, backend): def backend_config(spec, asset_path="/assets/pending"): - return _reader(TRAINERS, spec.trainer.backend)(spec, asset_path) + return _reader(TRAINERS, spec.trainer["backend"])(spec, asset_path) def serving_options(spec): - return _reader(INFERENCE, spec.inference.backend)(spec) + return _reader(INFERENCE, spec.inference["backend"])(spec) diff --git a/src/lilo/backends/megatron_deployment.py b/src/lilo/backends/megatron_deployment.py index 32d8fe7..34a13af 100644 --- a/src/lilo/backends/megatron_deployment.py +++ b/src/lilo/backends/megatron_deployment.py @@ -1,7 +1,9 @@ """Build Lilo loop settings while preserving native Megatron configuration.""" + from dataclasses import asdict, fields +from lilo.deployments import gpu_count from lilo.backend_options import native_options from .megatron_runtime.common.config import EngineModelConfig, OptimizerConfig @@ -36,13 +38,13 @@ def build_config(spec, asset_path): trainer = spec.trainer - if spec.model.parameterization != "full": + if spec.model["parameterization"] != "full": raise ValueError("Megatron deployments require full parameterization") - if trainer.engine.max_clients_per_instance != 1: + if trainer["engine"]["max_clients_per_instance"] != 1: raise ValueError("FFT trainers admit one client per instance") - if trainer.engine.sampler_persistence_concurrency != 1: + if trainer["engine"]["sampler_persistence_concurrency"] != 1: raise ValueError("Megatron requires sampler_persistence_concurrency: 1") - sections = native_options(trainer.config, set()) + sections = native_options(trainer["config"], set()) unknown = sections.keys() - {"runtime", "provider", "optimizer", "distributed"} if unknown: raise ValueError(f"unknown Megatron config sections: {sorted(unknown)}") @@ -76,7 +78,7 @@ def build_config(spec, asset_path): try: config = EngineModelConfig( hf_checkpoint=asset_path, - seq_length=spec.model.max_context_length, + seq_length=spec.model["max_context_length"], optimizer=OptimizerConfig(**loop_optimizer), provider_overrides=provider, optimizer_overrides=native_optimizer, @@ -85,5 +87,5 @@ def build_config(spec, asset_path): ) except TypeError as exc: raise ValueError(f"invalid Megatron runtime options: {exc}") from exc - config.validate(trainer.resources.gpu_count) + config.validate(gpu_count(trainer["resources"])) return {"megatron": asdict(config), "checkpoint_dir": "/checkpoints"} diff --git a/src/lilo/backends/miles_deployment.py b/src/lilo/backends/miles_deployment.py index caa22bd..e6ebb70 100644 --- a/src/lilo/backends/miles_deployment.py +++ b/src/lilo/backends/miles_deployment.py @@ -1,7 +1,9 @@ """Miles deployment integration. Native options are validated by Miles at startup.""" + from dataclasses import asdict +from lilo.deployments import gpu_count from lilo.backend_options import native_options from .miles_config import MilesBackendConfig @@ -36,9 +38,9 @@ def build_config(spec, asset_path): trainer = spec.trainer - if spec.model.parameterization != "lora": + if spec.model["parameterization"] != "lora": raise ValueError("Miles requires lora parameterization") - values = native_options(trainer.config, set()) + values = native_options(trainer["config"], set()) unknown = values.keys() - {"model_args", "options"} if unknown: raise ValueError(f"unknown Miles config sections: {sorted(unknown)}") @@ -75,9 +77,9 @@ def build_config(spec, asset_path): config = MilesBackendConfig( hf_checkpoint=asset_path, model_type=values.get("model_args") or "", - actor_num_gpus_per_node=trainer.resources.gpu_count, + actor_num_gpus_per_node=gpu_count(trainer["resources"]), native_options=options, - extra_args=("--seq-length", str(spec.model.max_context_length)), + extra_args=("--seq-length", str(spec.model["max_context_length"])), **settings, ) config.validate() @@ -85,6 +87,6 @@ def build_config(spec, asset_path): config.expert_model_parallel_size * config.expert_tensor_parallel_size ): raise ValueError("expert parallel sizes must divide the trainer GPU allocation") - if trainer.engine.max_clients_per_instance > config.max_lora_slots: + if trainer["engine"]["max_clients_per_instance"] > config.max_lora_slots: raise ValueError("max_clients_per_instance exceeds multi_lora_n_adapters") return {"miles": asdict(config), "checkpoint_dir": "/checkpoints"} diff --git a/src/lilo/configs/qwen35_35b_a3b_fft_64k.py b/src/lilo/configs/qwen35_35b_a3b_fft_64k.py index 52c653a..0269890 100644 --- a/src/lilo/configs/qwen35_35b_a3b_fft_64k.py +++ b/src/lilo/configs/qwen35_35b_a3b_fft_64k.py @@ -1,32 +1,23 @@ -from lilo.deployments import ( - BaseConfig, - EngineOptions, - Inference, - InferenceScaling, - Model, - Resources, - Routing, - Trainer, -) +from lilo.deployments import BaseConfig class Config(BaseConfig): name = "qwen35-35b-a3b-fft-64k" - model = Model( - id="Qwen/Qwen3.5-35B-A3B", parameterization="full", max_context_length=65536 - ) - routing = Routing(default=True) - trainer = Trainer( - backend="megatron", - resources=Resources(gpu="H200:8"), - engine=EngineOptions( - max_clients_per_instance=1, sampler_persistence_concurrency=1 - ), - env={ + model = { + "id": "Qwen/Qwen3.5-35B-A3B", + "parameterization": "full", + "max_context_length": 65536, + } + routing = {"default": True} + trainer = { + "backend": "megatron", + "resources": {"gpu": "H200:8"}, + "engine": {"max_clients_per_instance": 1, "sampler_persistence_concurrency": 1}, + "env": { "PYTORCH_CUDA_ALLOC_CONF": "expandable_segments:True", "TORCHINDUCTOR_COMPILE_THREADS": "1", }, - config={ + "config": { "runtime": { "tensor_model_parallel_size": 4, "pipeline_model_parallel_size": 1, @@ -54,11 +45,11 @@ class Config(BaseConfig): }, "optimizer": {"optimizer": "adam", "lr": 0.0001, "min_lr": 0.0001}, }, - ) - inference = Inference( - resources=Resources(gpu="H200:4"), - scaling=InferenceScaling(min_replicas=0, max_replicas=8, target_concurrency=16), - config={ + } + inference = { + "resources": {"gpu": "H200:4"}, + "scaling": {"min_replicas": 0, "max_replicas": 8, "target_concurrency": 16}, + "config": { "tp_size": 4, "ep_size": 4, "mem_fraction_static": 0.9, @@ -68,4 +59,4 @@ class Config(BaseConfig): "dp_size": 4, "enable_dp_attention": True, }, - ) + } diff --git a/src/lilo/configs/qwen35_4b_fft_64k.py b/src/lilo/configs/qwen35_4b_fft_64k.py index 35fd111..c0df8ac 100644 --- a/src/lilo/configs/qwen35_4b_fft_64k.py +++ b/src/lilo/configs/qwen35_4b_fft_64k.py @@ -1,32 +1,21 @@ -from lilo.deployments import ( - BaseConfig, - Deployment, - EngineOptions, - Inference, - Model, - Resources, - Routing, - Trainer, -) +from lilo.deployments import BaseConfig class Config(BaseConfig): name = "qwen35-4b-fft-64k" - model = Model( - id="Qwen/Qwen3.5-4B", - revision="851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", - parameterization="full", - max_context_length=65536, - ) - routing = Routing(default=True) - deployment = Deployment(frontend="lilo-yaml") - trainer = Trainer( - backend="megatron", - resources=Resources(gpu="H100:4"), - engine=EngineOptions( - max_clients_per_instance=1, sampler_persistence_concurrency=1 - ), - config={ + model = { + "id": "Qwen/Qwen3.5-4B", + "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", + "parameterization": "full", + "max_context_length": 65536, + } + routing = {"default": True} + deployment = {"frontend": "lilo-yaml"} + trainer = { + "backend": "megatron", + "resources": {"gpu": "H100:4"}, + "engine": {"max_clients_per_instance": 1, "sampler_persistence_concurrency": 1}, + "config": { "runtime": { "tensor_model_parallel_size": 2, "context_parallel_size": 2, @@ -45,14 +34,14 @@ class Config(BaseConfig): }, "optimizer": {"lr": 0.0001, "min_lr": 0.0001, "loss_scale": 1.0}, }, - ) - inference = Inference( - resources=Resources(gpu="H100:1"), - config={ + } + inference = { + "resources": {"gpu": "H100:1"}, + "config": { "tp_size": 1, "mem_fraction_static": 0.85, "max_running_requests": 32, "max_queued_requests": 4, "cpu_weight_cache_max_compile_group_gb": 16, }, - ) + } diff --git a/src/lilo/configs/qwen35_9b_fft_64k.py b/src/lilo/configs/qwen35_9b_fft_64k.py index 563512f..1ace764 100644 --- a/src/lilo/configs/qwen35_9b_fft_64k.py +++ b/src/lilo/configs/qwen35_9b_fft_64k.py @@ -1,32 +1,23 @@ -from lilo.deployments import ( - BaseConfig, - EngineOptions, - Inference, - InferenceScaling, - Model, - Resources, - Routing, - Trainer, -) +from lilo.deployments import BaseConfig class Config(BaseConfig): name = "qwen35-9b-fft-64k" - model = Model( - id="Qwen/Qwen3.5-9B", parameterization="full", max_context_length=65536 - ) - routing = Routing(default=True) - trainer = Trainer( - backend="megatron", - resources=Resources(gpu="H200:4"), - engine=EngineOptions( - max_clients_per_instance=1, sampler_persistence_concurrency=1 - ), - env={ + model = { + "id": "Qwen/Qwen3.5-9B", + "parameterization": "full", + "max_context_length": 65536, + } + routing = {"default": True} + trainer = { + "backend": "megatron", + "resources": {"gpu": "H200:4"}, + "engine": {"max_clients_per_instance": 1, "sampler_persistence_concurrency": 1}, + "env": { "PYTORCH_CUDA_ALLOC_CONF": "expandable_segments:True", "TORCHINDUCTOR_COMPILE_THREADS": "1", }, - config={ + "config": { "runtime": { "tensor_model_parallel_size": 2, "context_parallel_size": 2, @@ -45,11 +36,11 @@ class Config(BaseConfig): }, "optimizer": {"lr": 0.0001, "min_lr": 0.0001, "loss_scale": 1.0}, }, - ) - inference = Inference( - resources=Resources(gpu="H200:1"), - scaling=InferenceScaling(min_replicas=0, max_replicas=8, target_concurrency=16), - config={ + } + inference = { + "resources": {"gpu": "H200:1"}, + "scaling": {"min_replicas": 0, "max_replicas": 8, "target_concurrency": 16}, + "config": { "tp_size": 1, "ep_size": 1, "mem_fraction_static": 0.85, @@ -57,4 +48,4 @@ class Config(BaseConfig): "max_queued_requests": 4, "cpu_weight_cache_max_compile_group_gb": 16, }, - ) + } diff --git a/src/lilo/configs/qwen35_9b_instruct_lora_16k.py b/src/lilo/configs/qwen35_9b_instruct_lora_16k.py index 28399ec..fdeb869 100644 --- a/src/lilo/configs/qwen35_9b_instruct_lora_16k.py +++ b/src/lilo/configs/qwen35_9b_instruct_lora_16k.py @@ -1,32 +1,23 @@ -from lilo.deployments import ( - BaseConfig, - EngineOptions, - Inference, - InferenceScaling, - Model, - Resources, - Routing, - Trainer, -) +from lilo.deployments import BaseConfig class Config(BaseConfig): name = "qwen35-9b-instruct-lora-16k" - model = Model( - id="Qwen/Qwen3.5-9B", parameterization="lora", max_context_length=16384 - ) - routing = Routing(default=True) - trainer = Trainer( - backend="miles", - resources=Resources(gpu="H100:8"), - engine=EngineOptions( - max_clients_per_instance=6, sampler_persistence_concurrency=8 - ), - env={ + model = { + "id": "Qwen/Qwen3.5-9B", + "parameterization": "lora", + "max_context_length": 16384, + } + routing = {"default": True} + trainer = { + "backend": "miles", + "resources": {"gpu": "H100:8"}, + "engine": {"max_clients_per_instance": 6, "sampler_persistence_concurrency": 8}, + "env": { "PYTORCH_CUDA_ALLOC_CONF": "expandable_segments:True", "TORCHINDUCTOR_COMPILE_THREADS": "1", }, - config={ + "config": { "model_args": "qwen3.5-9B", "options": { "tensor_model_parallel_size": 8, @@ -45,11 +36,11 @@ class Config(BaseConfig): "lora_alpha": 32, }, }, - ) - inference = Inference( - resources=Resources(gpu="H200:1"), - scaling=InferenceScaling(min_replicas=0, max_replicas=8, target_concurrency=16), - config={ + } + inference = { + "resources": {"gpu": "H200:1"}, + "scaling": {"min_replicas": 0, "max_replicas": 8, "target_concurrency": 16}, + "config": { "tp_size": 1, "ep_size": 1, "mem_fraction_static": 0.8, @@ -68,4 +59,4 @@ class Config(BaseConfig): ], "schedule_policy": "lpm", }, - ) + } diff --git a/src/lilo/configs/qwen35_9b_lora_16k.py b/src/lilo/configs/qwen35_9b_lora_16k.py index a885981..1c17dc6 100644 --- a/src/lilo/configs/qwen35_9b_lora_16k.py +++ b/src/lilo/configs/qwen35_9b_lora_16k.py @@ -1,39 +1,26 @@ -from lilo.deployments import ( - BaseConfig, - Deployment, - EngineOptions, - Inference, - InferenceScaling, - Model, - Resources, - Routing, - Trainer, - TrainerScaling, -) +from lilo.deployments import BaseConfig class Config(BaseConfig): name = "qwen35-9b-lora-16k" - model = Model( - id="Qwen/Qwen3.5-9B-Base", - revision="68c46c4b3498877f3ef123c856ecfde50c39f404", - parameterization="lora", - max_context_length=16384, - ) - routing = Routing(default=True) - deployment = Deployment(frontend="lilo-yaml", mode="shared") - trainer = Trainer( - backend="miles", - resources=Resources(gpu="H100:4", cpu=16, memory_mib=65536), - scaling=TrainerScaling(max_instances=1), - engine=EngineOptions( - max_clients_per_instance=6, sampler_persistence_concurrency=8 - ), - env={ + model = { + "id": "Qwen/Qwen3.5-9B-Base", + "revision": "68c46c4b3498877f3ef123c856ecfde50c39f404", + "parameterization": "lora", + "max_context_length": 16384, + } + routing = {"default": True} + deployment = {"frontend": "lilo-yaml", "mode": "shared"} + trainer = { + "backend": "miles", + "resources": {"gpu": "H100:4", "cpu": 16, "memory_mib": 65536}, + "scaling": {"max_instances": 1}, + "engine": {"max_clients_per_instance": 6, "sampler_persistence_concurrency": 8}, + "env": { "PYTORCH_CUDA_ALLOC_CONF": "expandable_segments:True", "TORCHINDUCTOR_COMPILE_THREADS": "1", }, - config={ + "config": { "model_args": "qwen3.5-9B", "options": { "tensor_model_parallel_size": 4, @@ -53,12 +40,12 @@ class Config(BaseConfig): "recompute_num_layers": 1, }, }, - ) - inference = Inference( - backend="sglang", - resources=Resources(gpu="H200:1"), - scaling=InferenceScaling(min_replicas=0, max_replicas=8, target_concurrency=16), - config={ + } + inference = { + "backend": "sglang", + "resources": {"gpu": "H200:1"}, + "scaling": {"min_replicas": 0, "max_replicas": 8, "target_concurrency": 16}, + "config": { "tp_size": 1, "mem_fraction_static": 0.8, "max_running_requests": 32, @@ -66,4 +53,4 @@ class Config(BaseConfig): "max_loaded_loras": 64, "max_loras_per_batch": 8, }, - ) + } diff --git a/src/lilo/configs/qwen35_9b_lora_2k.py b/src/lilo/configs/qwen35_9b_lora_2k.py index ba84d3e..eb0de08 100644 --- a/src/lilo/configs/qwen35_9b_lora_2k.py +++ b/src/lilo/configs/qwen35_9b_lora_2k.py @@ -1,32 +1,23 @@ -from lilo.deployments import ( - BaseConfig, - EngineOptions, - Inference, - InferenceScaling, - Model, - Resources, - Routing, - Trainer, -) +from lilo.deployments import BaseConfig class Config(BaseConfig): name = "qwen35-9b-lora-2k" - model = Model( - id="Qwen/Qwen3.5-9B-Base", parameterization="lora", max_context_length=2048 - ) - routing = Routing(default=False) - trainer = Trainer( - backend="miles", - resources=Resources(gpu="H200:4"), - engine=EngineOptions( - max_clients_per_instance=4, sampler_persistence_concurrency=8 - ), - env={ + model = { + "id": "Qwen/Qwen3.5-9B-Base", + "parameterization": "lora", + "max_context_length": 2048, + } + routing = {"default": False} + trainer = { + "backend": "miles", + "resources": {"gpu": "H200:4"}, + "engine": {"max_clients_per_instance": 4, "sampler_persistence_concurrency": 8}, + "env": { "PYTORCH_CUDA_ALLOC_CONF": "expandable_segments:True", "TORCHINDUCTOR_COMPILE_THREADS": "1", }, - config={ + "config": { "model_args": "qwen3.5-9B", "options": { "tensor_model_parallel_size": 4, @@ -46,11 +37,11 @@ class Config(BaseConfig): "lora_alpha": 32, }, }, - ) - inference = Inference( - resources=Resources(gpu="H200:1"), - scaling=InferenceScaling(min_replicas=0, max_replicas=8, target_concurrency=16), - config={ + } + inference = { + "resources": {"gpu": "H200:1"}, + "scaling": {"min_replicas": 0, "max_replicas": 8, "target_concurrency": 16}, + "config": { "tp_size": 1, "ep_size": 1, "mem_fraction_static": 0.8, @@ -70,4 +61,4 @@ class Config(BaseConfig): ], "schedule_policy": "lpm", }, - ) + } diff --git a/src/lilo/configs/qwen36_27b_fft_64k.py b/src/lilo/configs/qwen36_27b_fft_64k.py index e8e7c0f..60bb5c4 100644 --- a/src/lilo/configs/qwen36_27b_fft_64k.py +++ b/src/lilo/configs/qwen36_27b_fft_64k.py @@ -1,32 +1,23 @@ -from lilo.deployments import ( - BaseConfig, - EngineOptions, - Inference, - InferenceScaling, - Model, - Resources, - Routing, - Trainer, -) +from lilo.deployments import BaseConfig class Config(BaseConfig): name = "qwen36-27b-fft-64k" - model = Model( - id="Qwen/Qwen3.6-27B", parameterization="full", max_context_length=65536 - ) - routing = Routing(default=True) - trainer = Trainer( - backend="megatron", - resources=Resources(gpu="H200:8"), - engine=EngineOptions( - max_clients_per_instance=1, sampler_persistence_concurrency=1 - ), - env={ + model = { + "id": "Qwen/Qwen3.6-27B", + "parameterization": "full", + "max_context_length": 65536, + } + routing = {"default": True} + trainer = { + "backend": "megatron", + "resources": {"gpu": "H200:8"}, + "engine": {"max_clients_per_instance": 1, "sampler_persistence_concurrency": 1}, + "env": { "PYTORCH_CUDA_ALLOC_CONF": "expandable_segments:True", "TORCHINDUCTOR_COMPILE_THREADS": "1", }, - config={ + "config": { "runtime": { "tensor_model_parallel_size": 4, "pipeline_model_parallel_size": 1, @@ -47,11 +38,11 @@ class Config(BaseConfig): }, "optimizer": {"optimizer": "adam", "lr": 0.0001, "min_lr": 0.0001}, }, - ) - inference = Inference( - resources=Resources(gpu="H200:4"), - scaling=InferenceScaling(min_replicas=0, max_replicas=8, target_concurrency=16), - config={ + } + inference = { + "resources": {"gpu": "H200:4"}, + "scaling": {"min_replicas": 0, "max_replicas": 8, "target_concurrency": 16}, + "config": { "tp_size": 4, "ep_size": 1, "mem_fraction_static": 0.9, @@ -59,4 +50,4 @@ class Config(BaseConfig): "max_queued_requests": 4, "cpu_weight_cache_max_compile_group_gb": 32, }, - ) + } diff --git a/src/lilo/configs/qwen38_27b_lora_16k.py b/src/lilo/configs/qwen38_27b_lora_16k.py index f70cf24..2cb36a6 100644 --- a/src/lilo/configs/qwen38_27b_lora_16k.py +++ b/src/lilo/configs/qwen38_27b_lora_16k.py @@ -1,32 +1,23 @@ -from lilo.deployments import ( - BaseConfig, - EngineOptions, - Inference, - InferenceScaling, - Model, - Resources, - Routing, - Trainer, -) +from lilo.deployments import BaseConfig class Config(BaseConfig): name = "qwen38-27b-lora-16k" - model = Model( - id="Qwen/Qwen3.8-27B", parameterization="lora", max_context_length=16384 - ) - routing = Routing(default=True) - trainer = Trainer( - backend="miles", - resources=Resources(gpu="H200:8"), - engine=EngineOptions( - max_clients_per_instance=6, sampler_persistence_concurrency=8 - ), - env={ + model = { + "id": "Qwen/Qwen3.8-27B", + "parameterization": "lora", + "max_context_length": 16384, + } + routing = {"default": True} + trainer = { + "backend": "miles", + "resources": {"gpu": "H200:8"}, + "engine": {"max_clients_per_instance": 6, "sampler_persistence_concurrency": 8}, + "env": { "PYTORCH_CUDA_ALLOC_CONF": "expandable_segments:True", "TORCHINDUCTOR_COMPILE_THREADS": "1", }, - config={ + "config": { "model_args": "qwen3.8-27B", "options": { "tensor_model_parallel_size": 4, @@ -45,11 +36,11 @@ class Config(BaseConfig): "lora_alpha": 32, }, }, - ) - inference = Inference( - resources=Resources(gpu="H200:1"), - scaling=InferenceScaling(min_replicas=0, max_replicas=8, target_concurrency=16), - config={ + } + inference = { + "resources": {"gpu": "H200:1"}, + "scaling": {"min_replicas": 0, "max_replicas": 8, "target_concurrency": 16}, + "config": { "tp_size": 1, "ep_size": 1, "mem_fraction_static": 0.8, @@ -68,4 +59,4 @@ class Config(BaseConfig): ], "schedule_policy": "lpm", }, - ) + } diff --git a/src/lilo/deployment_cli.py b/src/lilo/deployment_cli.py index c846d3c..a77408c 100644 --- a/src/lilo/deployment_cli.py +++ b/src/lilo/deployment_cli.py @@ -46,14 +46,14 @@ def compile_configs(paths): implementation = implementation_fingerprint(miles_commit) records = [] for spec in specs: - revision = spec.model.revision + revision = spec.model["revision"] if not re.fullmatch(r"[0-9a-fA-F]{40,64}", revision): from huggingface_hub import HfApi - revision = HfApi().model_info(spec.model.id, revision=revision).sha + revision = HfApi().model_info(spec.model["id"], revision=revision).sha if not revision or not re.fullmatch(r"[0-9a-fA-F]{40,64}", revision): raise ValueError( - f"Hugging Face did not return a commit for {spec.model.id}" + f"Hugging Face did not return a commit for {spec.model['id']}" ) records.append( DeploymentRecord.create( @@ -104,21 +104,22 @@ def deploy(desired): ) settings = desired[0].spec.deployment registry = modal.Dict.from_name( - f"{settings.frontend}-yaml-deployments", + f"{settings['frontend']}-yaml-deployments", create_if_missing=True, - environment_name=settings.modal.environment, + environment_name=settings["modal"]["environment"], ) owner = uuid.uuid4().hex if not registry.put("apply_lock", owner, skip_if_exists=True): raise ValueError( - f"An apply owns {settings.frontend}. If it was interrupted, confirm it has stopped before running lilo deployment unlock --frontend {settings.frontend}." + f"An apply owns {settings['frontend']}. If it was interrupted, confirm it has stopped before running lilo deployment unlock --frontend {settings['frontend']}." ) try: rows = registry.get("manifest", []) if not rows and not registry.get("pending", []): try: modal.App.lookup( - settings.frontend, environment_name=settings.modal.environment + settings["frontend"], + environment_name=settings["modal"]["environment"], ) except modal.exception.NotFoundError: pass @@ -135,7 +136,7 @@ def deploy(desired): env = { **os.environ, MANIFEST_ENV: json.dumps(data), - "LILO_APP_NAME": settings.frontend, + "LILO_APP_NAME": settings["frontend"], } if desired[0].miles_commit: env["LILO_MILES_COMMIT"] = desired[0].miles_commit @@ -147,8 +148,8 @@ def deploy(desired): "-m", "lilo.providers.modal.app", ] - if settings.modal.environment: - command += ["--env", settings.modal.environment] + if settings["modal"]["environment"]: + command += ["--env", settings["modal"]["environment"]] registry.put("pending", data) subprocess.run(command, check=True, env=env) registry.put("manifest", data) diff --git a/src/lilo/deployments.py b/src/lilo/deployments.py index 40e1b00..f961936 100644 --- a/src/lilo/deployments.py +++ b/src/lilo/deployments.py @@ -3,7 +3,7 @@ from __future__ import annotations from copy import deepcopy -from dataclasses import MISSING, asdict, dataclass, field, fields +from dataclasses import asdict, dataclass, fields import hashlib from importlib.resources import files import json @@ -11,176 +11,118 @@ import re import runpy import sys -from typing import Any, ClassVar, Literal +from typing import Any, ClassVar from pydantic import BaseModel, ConfigDict -@dataclass(kw_only=True) -class Model: - id: str - max_context_length: int - revision: str = "main" - parameterization: Literal["lora", "full"] = "lora" - - -@dataclass(kw_only=True) -class Routing: - default: bool = False - sampling_default: bool = False - - -@dataclass(kw_only=True) -class Resources: - gpu: str - cpu: float = 8 - memory_mib: int = 32768 - timeout_s: int = 86400 - - @property - def gpu_count(self) -> int: - if not re.fullmatch(r"[A-Za-z0-9-]+(?::[1-9][0-9]*)?", self.gpu): - raise ValueError(f"invalid GPU resource: {self.gpu}") - return int(self.gpu.split(":")[1]) if ":" in self.gpu else 1 - - -@dataclass(kw_only=True) -class TrainerScaling: - min_instances: Literal[0] = 0 - max_instances: int = 1 - - -@dataclass(kw_only=True) -class InferenceScaling: - min_replicas: int = 0 - max_replicas: int = 8 - target_concurrency: int = 16 - scaledown_window_s: int = 300 - - def __post_init__(self): - if self.min_replicas > self.max_replicas: - raise ValueError("min_replicas must not exceed max_replicas") - - -@dataclass(kw_only=True) -class EngineOptions: - max_clients_per_instance: int = 1 - sampler_persistence_concurrency: int = 8 - - -@dataclass(kw_only=True) -class Trainer: - resources: Resources - backend: str = "miles" - scaling: TrainerScaling = field(default_factory=TrainerScaling) - engine: EngineOptions = field(default_factory=EngineOptions) - config: dict[str, Any] = field(default_factory=dict) - env: dict[str, str] = field(default_factory=dict) - - -@dataclass(kw_only=True) -class Inference: - resources: Resources - backend: str = "sglang" - scaling: InferenceScaling = field(default_factory=InferenceScaling) - config: dict[str, Any] = field(default_factory=dict) - env: dict[str, str] = field(default_factory=dict) - - -@dataclass(kw_only=True) -class Secrets: - api: str = "lilo-api" - sampler_proxy: str = "lilo-proxy" - huggingface: str | None = "huggingface-secret" - - -@dataclass(kw_only=True) -class Storage: - assets: str = "lilo-model-assets" - checkpoints: str = "lilo-checkpoints" - bulletin: str = "lilo-snapshot-bulletin" - - -@dataclass(kw_only=True) -class ModalSettings: - environment: str | None = None - region: str = "us-west" - - -@dataclass(kw_only=True) -class Deployment: - frontend: str = "lilo-yaml" - mode: Literal["shared"] = "shared" - modal: ModalSettings = field(default_factory=ModalSettings) - secrets: Secrets = field(default_factory=Secrets) - storage: Storage = field(default_factory=Storage) - - -@dataclass(kw_only=True) -class Lifecycle: - session_idle_timeout_s: int = 300 - pool_idle_timeout_s: int = 300 - sweep_interval_s: int = 300 +# Shared orchestration defaults. Backend option dictionaries have no schema here. +_DEFAULTS = { + "api_version": "lilo/v1", + "model": {"revision": "main", "parameterization": "lora"}, + "routing": {"default": False, "sampling_default": False}, + "trainer": { + "backend": "miles", + "resources": {"cpu": 8, "memory_mib": 32768, "timeout_s": 86400}, + "scaling": {"min_instances": 0, "max_instances": 1}, + "engine": {"max_clients_per_instance": 1, "sampler_persistence_concurrency": 8}, + "config": {}, + "env": {}, + }, + "inference": { + "backend": "sglang", + "resources": {"cpu": 8, "memory_mib": 32768, "timeout_s": 86400}, + "scaling": { + "min_replicas": 0, + "max_replicas": 8, + "target_concurrency": 16, + "scaledown_window_s": 300, + }, + "config": {}, + "env": {}, + }, + "deployment": { + "frontend": "lilo-yaml", + "mode": "shared", + "modal": {"environment": None, "region": "us-west"}, + "secrets": { + "api": "lilo-api", + "sampler_proxy": "lilo-proxy", + "huggingface": "huggingface-secret", + }, + "storage": { + "assets": "lilo-model-assets", + "checkpoints": "lilo-checkpoints", + "bulletin": "lilo-snapshot-bulletin", + }, + }, + "lifecycle": { + "session_idle_timeout_s": 300, + "pool_idle_timeout_s": 300, + "sweep_interval_s": 300, + }, +} + + +def _with_defaults(defaults, values): + """Fill omitted orchestration settings; explicit values win.""" + if not isinstance(defaults, dict) or not isinstance(values, dict): + return deepcopy(values) + result = deepcopy(defaults) + for key, value in values.items(): + result[key] = _with_defaults(defaults.get(key), value) + return result @dataclass(kw_only=True, init=False) class BaseConfig: - """Declare class defaults and dotted overrides; each instance owns its values.""" + """Declare plain section dictionaries and optional inherited overrides.""" name: str - model: Model - trainer: Trainer - inference: Inference - api_version: Literal["lilo/v1"] = "lilo/v1" - routing: Routing = field(default_factory=Routing) - deployment: Deployment = field(default_factory=Deployment) - lifecycle: Lifecycle = field(default_factory=Lifecycle) - + model: dict[str, Any] + trainer: dict[str, Any] + inference: dict[str, Any] + api_version: str + routing: dict[str, Any] + deployment: dict[str, Any] + lifecycle: dict[str, Any] overrides: ClassVar[dict[str, Any]] = {} def __init__(self, **kwargs): - definitions = {item.name: item for item in fields(BaseConfig)} - unknown = kwargs.keys() - definitions.keys() + names = {item.name for item in fields(BaseConfig)} + unknown = kwargs.keys() - names if unknown: raise TypeError(f"unknown config fields: {sorted(unknown)}") - values = {} - for name, item in definitions.items(): - if item.default_factory is not MISSING: - values[name] = item.default_factory() - elif item.default is not MISSING: - values[name] = deepcopy(item.default) - # Apply each parent's defaults and overrides before its child's. Copy at - # every assignment so instances never mutate class defaults or parents. + values = deepcopy(_DEFAULTS) for cls in reversed(type(self).__mro__): - for name in definitions.keys() & vars(cls).keys(): - values[name] = deepcopy(vars(cls)[name]) + for name in names & vars(cls).keys(): + values[name] = _with_defaults(_DEFAULTS.get(name), vars(cls)[name]) for path, value in vars(cls).get("overrides", {}).items(): parts = path.split(".") + if parts[0] not in names: + raise ValueError(f"unknown config override: {path}") target = values try: for part in parts[:-1]: - target = ( - target[part] - if isinstance(target, dict) - else getattr(target, part) - ) - if isinstance(target, dict): - if len(parts) == 1 and parts[0] not in definitions: - raise KeyError(parts[0]) - target[parts[-1]] = deepcopy(value) - else: - # Reject misspelled dataclass fields. - getattr(target, parts[-1]) - setattr(target, parts[-1], deepcopy(value)) - except (KeyError, AttributeError) as exc: + target = target[part] + target[parts[-1]] = deepcopy(value) + except (KeyError, TypeError) as exc: raise ValueError(f"unknown config override: {path}") from exc - values.update(deepcopy(kwargs)) - missing = definitions.keys() - values.keys() + for name, value in kwargs.items(): + values[name] = _with_defaults(_DEFAULTS.get(name), value) + missing = names - values.keys() if missing: raise TypeError(f"missing config fields: {sorted(missing)}") self.__dict__.update(values) +def gpu_count(resources): + gpu = resources["gpu"] + if not re.fullmatch(r"[A-Za-z0-9-]+(?::[1-9][0-9]*)?", gpu): + raise ValueError(f"invalid GPU resource: {gpu}") + return int(gpu.split(":")[1]) if ":" in gpu else 1 + + class DeploymentRecord(BaseModel): """Saved deployment metadata around a Python configuration. @@ -208,7 +150,7 @@ def create( ) -> DeploymentRecord: """Record an already-resolved revision without reparsing the configuration.""" pinned = deepcopy(spec) - pinned.model.revision = revision + pinned.model["revision"] = revision # Changing routing defaults should not restart an existing trainer. identity = asdict(pinned) identity.pop("routing") @@ -229,7 +171,7 @@ def definition_id(self) -> str: @property def asset_path(self) -> str: digest = hashlib.sha256( - f"{self.spec.model.id}@{self.spec.model.revision}".encode() + f"{self.spec.model['id']}@{self.spec.model['revision']}".encode() ).hexdigest() return f"/assets/{digest}" @@ -274,12 +216,12 @@ def validate_frontend(specs: list[BaseConfig]) -> None: ) defaults, sampling = set(), set() for spec in specs: - key = (spec.model.id, spec.model.parameterization) - if spec.routing.default: + key = (spec.model["id"], spec.model["parameterization"]) + if spec.routing["default"]: if key in defaults: raise ValueError(f"multiple defaults for {key}") defaults.add(key) - if spec.routing.sampling_default: - if spec.model.id in sampling: - raise ValueError(f"multiple sampling defaults for {spec.model.id}") - sampling.add(spec.model.id) + if spec.routing["sampling_default"]: + if spec.model["id"] in sampling: + raise ValueError(f"multiple sampling defaults for {spec.model['id']}") + sampling.add(spec.model["id"]) diff --git a/src/lilo/inference/sglang_deployment.py b/src/lilo/inference/sglang_deployment.py index fe5ef3b..3242e83 100644 --- a/src/lilo/inference/sglang_deployment.py +++ b/src/lilo/inference/sglang_deployment.py @@ -1,5 +1,7 @@ """SGLang settings that must agree with Lilo replica orchestration.""" + +from lilo.deployments import gpu_count from lilo.backend_options import native_options SGLANG_MANAGED = { @@ -30,10 +32,10 @@ def build_config(spec): - options = native_options(spec.inference.config, SGLANG_MANAGED) - tp = options.get("tp_size", spec.inference.resources.gpu_count) + options = native_options(spec.inference["config"], SGLANG_MANAGED) + tp = options.get("tp_size", gpu_count(spec.inference["resources"])) ep = options.get("ep_size", 1) - if not isinstance(tp, int) or tp != spec.inference.resources.gpu_count: + if not isinstance(tp, int) or tp != gpu_count(spec.inference["resources"]): raise ValueError("sglang.tp_size must equal the replica GPU allocation") if not isinstance(ep, int) or ep < 1 or tp % ep: raise ValueError("sglang.ep_size must divide the replica GPU allocation") diff --git a/src/lilo/providers/modal/app.py b/src/lilo/providers/modal/app.py index bb0d3a1..b65338b 100644 --- a/src/lilo/providers/modal/app.py +++ b/src/lilo/providers/modal/app.py @@ -56,19 +56,19 @@ ) SETTINGS = frontend_settings() -APP_NAME = SETTINGS.deployment.frontend -ROUTING_REGION = SETTINGS.deployment.modal.region +APP_NAME = SETTINGS.deployment['frontend'] +ROUTING_REGION = SETTINGS.deployment['modal']['region'] MODEL_ASSET_ROOT = "/assets" -SESSION_IDLE_TIMEOUT = SETTINGS.lifecycle.session_idle_timeout_s -FFT_POOL_IDLE_TIMEOUT = LORA_POOL_IDLE_TIMEOUT = SETTINGS.lifecycle.pool_idle_timeout_s +SESSION_IDLE_TIMEOUT = SETTINGS.lifecycle['session_idle_timeout_s'] +FFT_POOL_IDLE_TIMEOUT = LORA_POOL_IDLE_TIMEOUT = SETTINGS.lifecycle['pool_idle_timeout_s'] FFT_POOL_TOUCH_INTERVAL = 60.0 LORA_POOL_CHECK_INTERVAL = 60.0 -SWEEP_PERIOD = modal.Period(seconds=SETTINGS.lifecycle.sweep_interval_s) +SWEEP_PERIOD = modal.Period(seconds=SETTINGS.lifecycle['sweep_interval_s']) CHECKPOINT_READ_LOCK = asyncio.Lock() _pool_touches: dict[str, float] = {} _lora_pool_gateways: dict[str, tuple[float, str]] = {} _lora_pool_checks: dict[str, asyncio.Lock] = {} -CHECKPOINT_VOLUME_NAME = SETTINGS.deployment.storage.checkpoints +CHECKPOINT_VOLUME_NAME = SETTINGS.deployment['storage']['checkpoints'] checkpoint_volume = modal.Volume.from_name( CHECKPOINT_VOLUME_NAME, create_if_missing=True, version=2 ) @@ -110,13 +110,13 @@ async def _delete_checkpoint(uri: str) -> None: .add_local_python_source("lilo", ignore=ignore_config_source) ) model_assets = modal.Volume.from_name( - SETTINGS.deployment.storage.assets, + SETTINGS.deployment['storage']['assets'], create_if_missing=True, ) -API_SECRET_NAME = SETTINGS.deployment.secrets.api -HF_SECRET_NAME = SETTINGS.deployment.secrets.huggingface +API_SECRET_NAME = SETTINGS.deployment['secrets']['api'] +HF_SECRET_NAME = SETTINGS.deployment['secrets']['huggingface'] proxy_secret = modal.Secret.from_name( - SETTINGS.deployment.secrets.sampler_proxy, + SETTINGS.deployment['secrets']['sampler_proxy'], required_keys=["MODAL_PROXY_TOKEN_ID", "MODAL_PROXY_TOKEN_SECRET"], ) diff --git a/src/lilo/providers/modal/deployment_apps.py b/src/lilo/providers/modal/deployment_apps.py index 566e7bb..d6bccad 100644 --- a/src/lilo/providers/modal/deployment_apps.py +++ b/src/lilo/providers/modal/deployment_apps.py @@ -12,7 +12,7 @@ import modal -from lilo.deployments import DeploymentRecord, Routing, validate_frontend +from lilo.deployments import DeploymentRecord, gpu_count, validate_frontend from lilo.backends.deployment import backend_config, serving_options MANIFEST_ENV = "LILO_DEPLOYMENT_MANIFEST" @@ -58,30 +58,30 @@ def image_for(backend): def volumes_for(spec): - storage = spec.deployment.storage + storage = spec.deployment["storage"] return { - "/assets": modal.Volume.from_name(storage.assets, create_if_missing=True), + "/assets": modal.Volume.from_name(storage["assets"], create_if_missing=True), "/checkpoints": modal.Volume.from_name( - storage.checkpoints, create_if_missing=True, version=2 + storage["checkpoints"], create_if_missing=True, version=2 ), "/bulletin": modal.Volume.from_name( - storage.bulletin, create_if_missing=True, version=2 + storage["bulletin"], create_if_missing=True, version=2 ), } def secrets_for(spec, *, training=False): - names = spec.deployment.secrets - result = [modal.Secret.from_name(names.api, required_keys=["TINKER_API_KEY"])] + names = spec.deployment["secrets"] + result = [modal.Secret.from_name(names["api"], required_keys=["TINKER_API_KEY"])] if training: result.append( modal.Secret.from_name( - names.sampler_proxy, + names["sampler_proxy"], required_keys=["MODAL_PROXY_TOKEN_ID", "MODAL_PROXY_TOKEN_SECRET"], ) ) - if names.huggingface: - result.append(modal.Secret.from_name(names.huggingface)) + if names["huggingface"]: + result.append(modal.Secret.from_name(names["huggingface"])) return result @@ -97,31 +97,31 @@ def build_trainer_app(resolved: DeploymentRecord, *, image=None): # configuration when a default switches or an older configuration drains. resolved = resolved.model_copy(deep=True) resolved.active = True - resolved.spec.routing = Routing() + resolved.spec.routing = {"default": False, "sampling_default": False} spec = resolved.spec config_json = json.dumps( resolved.model_dump(mode="json"), sort_keys=True, separators=(",", ":") ) - app = modal.App(f"{spec.deployment.frontend}-{resolved.definition_id}") - resource = spec.trainer.resources + app = modal.App(f"{spec.deployment['frontend']}-{resolved.definition_id}") + resource = spec.trainer["resources"] from .deployment import trainer_deployment_env env = { **trainer_deployment_env(), - **deployment_env(spec.trainer.env), - "LILO_APP_NAME": spec.deployment.frontend, + **deployment_env(spec.trainer["env"]), + "LILO_APP_NAME": spec.deployment["frontend"], } @app.function( name=resolved.definition_id, serialized=True, - image=image if image is not None else image_for(spec.trainer.backend), - gpu=resource.gpu, - region=spec.deployment.modal.region, - cpu=resource.cpu, - memory=resource.memory_mib, - timeout=resource.timeout_s, - max_containers=spec.trainer.scaling.max_instances, + image=image if image is not None else image_for(spec.trainer["backend"]), + gpu=resource["gpu"], + region=spec.deployment["modal"]["region"], + cpu=resource["cpu"], + memory=resource["memory_mib"], + timeout=resource["timeout_s"], + max_containers=spec.trainer["scaling"]["max_instances"], min_containers=0, single_use_containers=True, volumes=volumes_for(spec), @@ -145,20 +145,20 @@ def run_trainer(resolved, instance_id): # on startup to see the committed exact snapshot; never race a trainer download. volumes_for(spec)["/assets"].reload() env = { - **deployment_env(spec.trainer.env), - "LILO_APP_NAME": spec.deployment.frontend, + **deployment_env(spec.trainer["env"]), + "LILO_APP_NAME": spec.deployment["frontend"], "LILO_BACKEND_CONFIG": json.dumps(settings), - "LILO_BASE_MODEL": spec.model.id, - "LILO_BASE_MODEL_REVISION": spec.model.revision, + "LILO_BASE_MODEL": spec.model["id"], + "LILO_BASE_MODEL_REVISION": spec.model["revision"], "LILO_DEFINITION_ID": resolved.definition_id, - "LILO_CHECKPOINT_VOLUME": spec.deployment.storage.checkpoints, + "LILO_CHECKPOINT_VOLUME": spec.deployment["storage"]["checkpoints"], "LILO_BULLETIN_ROOT": "/bulletin", - "LILO_BULLETIN_VOLUME": spec.deployment.storage.bulletin, + "LILO_BULLETIN_VOLUME": spec.deployment["storage"]["bulletin"], "LILO_DEFINITION_REVISION": resolved.generation, } executor = ( "lilo.backends.miles_lora:build_executor" - if spec.trainer.backend == "miles" + if spec.trainer["backend"] == "miles" else "lilo.backends.megatron_fft:build_executor" ) @@ -176,10 +176,12 @@ async def failed(error): instance_id=instance_id, backend_env=env, nproc=1 - if spec.trainer.backend == "miles" - else spec.trainer.resources.gpu_count, - max_models=spec.trainer.engine.max_clients_per_instance, - sampler_persistence_concurrency=spec.trainer.engine.sampler_persistence_concurrency, + if spec.trainer["backend"] == "miles" + else gpu_count(spec.trainer["resources"]), + max_models=spec.trainer["engine"]["max_clients_per_instance"], + sampler_persistence_concurrency=spec.trainer["engine"][ + "sampler_persistence_concurrency" + ], on_startup_error=failed, ) @@ -189,21 +191,21 @@ def definition_from_spec(resolved, *, register_trainer=True, image=None): native = serving_options(spec) definition = SimpleNamespace( DEFINITION_ID=resolved.definition_id, - MODEL_NAME=spec.model.id, - MODEL_REVISION=spec.model.revision, + MODEL_NAME=spec.model["id"], + MODEL_REVISION=spec.model["revision"], HF_CHECKPOINT=resolved.asset_path, - PARAMETERIZATION=spec.model.parameterization, + PARAMETERIZATION=spec.model["parameterization"], CATALOG_VISIBLE=resolved.active, - ROUTING_DEFAULT=spec.routing.default, - SAMPLING_DEFAULT=spec.routing.sampling_default, + ROUTING_DEFAULT=spec.routing["default"], + SAMPLING_DEFAULT=spec.routing["sampling_default"], DEPLOYMENT_NAME=spec.name, RESOLVED=resolved, - MAX_CONTEXT_LENGTH=spec.model.max_context_length, - TRAINER_MODELS_PER_INSTANCE=spec.trainer.engine.max_clients_per_instance, - TRAINER_MAX_CONTAINERS=spec.trainer.scaling.max_instances, - ROLLOUT_GPUS=spec.inference.resources.gpu_count, + MAX_CONTEXT_LENGTH=spec.model["max_context_length"], + TRAINER_MODELS_PER_INSTANCE=spec.trainer["engine"]["max_clients_per_instance"], + TRAINER_MAX_CONTAINERS=spec.trainer["scaling"]["max_instances"], + ROLLOUT_GPUS=gpu_count(spec.inference["resources"]), ROLLOUT_TENSOR_PARALLEL_SIZE=native.get( - "tp_size", spec.inference.resources.gpu_count + "tp_size", gpu_count(spec.inference["resources"]) ) // (native.get("dp_size", 1) if native.get("enable_dp_attention") else 1), ) @@ -231,14 +233,14 @@ def pool_environment(definition_id): def build_rollout_app(resolved, pool, *, image=None): """Create one frozen-base LoRA pool or one FFT latest/pinned/base pool.""" spec = resolved.spec - lora = spec.model.parameterization == "lora" + lora = spec.model["parameterization"] == "lora" if pool.definition_id != resolved.definition_id: raise ValueError("pool generation does not match deployment") app = modal.App(pool.app_name) - resources, scaling = spec.inference.resources, spec.inference.scaling + resources, scaling = spec.inference["resources"], spec.inference["scaling"] options = { - "context_length": spec.model.max_context_length, - "tp_size": resources.gpu_count, + "context_length": spec.model["max_context_length"], + "tp_size": gpu_count(resources), "mem_fraction_static": 0.8, "max_running_requests": 32, "weight_loader_disable_mmap": True, @@ -264,21 +266,21 @@ def build_rollout_app(resolved, pool, *, image=None): name="Server", serialized=True, image=image if image is not None else image_for("sglang"), - gpu=resources.gpu, - cpu=resources.cpu, - memory=resources.memory_mib, + gpu=resources["gpu"], + cpu=resources["cpu"], + memory=resources["memory_mib"], volumes=volumes_for(spec), secrets=secrets_for(spec), - env=deployment_env(spec.inference.env), - min_containers=scaling.min_replicas if minimum is None else minimum, - max_containers=scaling.max_replicas if maximum is None else maximum, - target_concurrency=scaling.target_concurrency, - scaledown_window=scaling.scaledown_window_s if window is None else window, + env=deployment_env(spec.inference["env"]), + min_containers=scaling["min_replicas"] if minimum is None else minimum, + max_containers=scaling["max_replicas"] if maximum is None else maximum, + target_concurrency=scaling["target_concurrency"], + scaledown_window=scaling["scaledown_window_s"] if window is None else window, startup_timeout=1200, exit_grace_period=300, port=8000, - routing_region=spec.deployment.modal.region, - compute_region=spec.deployment.modal.region, + routing_region=spec.deployment["modal"]["region"], + compute_region=spec.deployment["modal"]["region"], ) class Server: @modal.enter() @@ -309,7 +311,7 @@ def start(self): port=8000, sglang_port=8001, bulletin_root="/bulletin", - bulletin_volume=spec.deployment.storage.bulletin, + bulletin_volume=spec.deployment["storage"]["bulletin"], ) self.sidecar = ( start_lora_sidecar(**kwargs) diff --git a/src/lilo/providers/modal/deployment_pool_app.py b/src/lilo/providers/modal/deployment_pool_app.py index 55dd5b5..a4348ec 100644 --- a/src/lilo/providers/modal/deployment_pool_app.py +++ b/src/lilo/providers/modal/deployment_pool_app.py @@ -8,7 +8,7 @@ from .deployment_apps import POOL_CONFIG_ENV, build_rollout_app resolved = DeploymentRecord.model_validate_json(os.environ[POOL_CONFIG_ENV]) -if resolved.spec.model.parameterization == "lora": +if resolved.spec.model["parameterization"] == "lora": pool = LoraPoolSpec(resolved.definition_id, revision=resolved.generation[:16]) else: pool = FFTPoolSpec( diff --git a/tests/providers/test_checkpoint_storage.py b/tests/providers/test_checkpoint_storage.py index 9516507..c02778c 100644 --- a/tests/providers/test_checkpoint_storage.py +++ b/tests/providers/test_checkpoint_storage.py @@ -44,10 +44,10 @@ def test_yaml_definitions_use_configured_checkpoint_storage(): ): volumes = volumes_for(spec) assert volumes[CHECKPOINT_ROOT] == ( - spec.deployment.storage.checkpoints, + spec.deployment['storage']['checkpoints'], {"create_if_missing": True, "version": 2}, ) - assert volumes["/bulletin"][0] == spec.deployment.storage.bulletin + assert volumes["/bulletin"][0] == spec.deployment['storage']['bulletin'] assert backend_config(spec)["checkpoint_dir"] == CHECKPOINT_ROOT diff --git a/tests/providers/test_deployment_apps.py b/tests/providers/test_deployment_apps.py index 85cca79..11859c4 100644 --- a/tests/providers/test_deployment_apps.py +++ b/tests/providers/test_deployment_apps.py @@ -56,7 +56,7 @@ def test_trainer_declaration_and_executor_configuration( builders, monkeypatch, preset, backend, clients, nproc ): row = deployment(preset) - row.spec.deployment.storage.checkpoints = "test-custom-checkpoints" + row.spec.deployment['storage']['checkpoints'] = "test-custom-checkpoints" image = object() app, trainer = deployment_apps.build_trainer_app(row, image=image) declaration, _ = app.functions[row.definition_id] @@ -88,7 +88,7 @@ def test_trainer_declaration_and_executor_configuration( assert kwargs["backend_env"]["LILO_CHECKPOINT_VOLUME"] == "test-custom-checkpoints" assert kwargs["backend_env"]["LILO_BASE_MODEL_REVISION"] == "a" * 40 config = json.loads(kwargs["backend_env"]["LILO_BACKEND_CONFIG"]) - assert config[row.spec.trainer.backend]["hf_checkpoint"] == row.asset_path + assert config[row.spec.trainer['backend']]["hf_checkpoint"] == row.asset_path assert config["checkpoint_dir"] == "/checkpoints" assert reloaded == [True] @@ -108,7 +108,7 @@ def test_pool_starts_native_server_and_correct_sidecar(builders, monkeypatch, ki app, server = deployment_apps.build_rollout_app(row, pool, image="test-image") settings, _ = app.servers["Server"] assert app.name == pool.app_name - assert settings["gpu"] == row.spec.inference.resources.gpu + assert settings["gpu"] == row.spec.inference['resources']['gpu'] assert settings["min_containers"] == 0 assert settings["target_concurrency"] == 16 assert settings["compute_region"] == "us-west" @@ -138,7 +138,7 @@ def test_pool_starts_native_server_and_correct_sidecar(builders, monkeypatch, ki assert commands[0][2] == "lilo.inference.native_sglang" assert commands[0][3] == row.asset_path native = json.loads(commands[0][4]) - assert native["context_length"] == row.spec.model.max_context_length + assert native["context_length"] == row.spec.model['max_context_length'] if kind == "lora": assert native["enable_lora"] is True assert native["max_lora_rank"] == 32 @@ -232,12 +232,12 @@ def test_admission_changes_preserve_serialized_trainer(builders): old_bytes = serialize(deployment_apps.build_trainer_app(first, image="test")[1]) changed = first.model_copy(deep=True) changed.active = False - changed.spec.routing.default = not first.spec.routing.default - changed.spec.routing.sampling_default = True + changed.spec.routing['default'] = not first.spec.routing['default'] + changed.spec.routing['sampling_default'] = True new_bytes = serialize(deployment_apps.build_trainer_app(changed, image="test")[1]) assert new_bytes == old_bytes - assert first.active is True and first.spec.routing.default is True - changed.spec.trainer.resources.gpu = "H200:4" + assert first.active is True and first.spec.routing['default'] is True + changed.spec.trainer['resources']['gpu'] = "H200:4" assert serialize(deployment_apps.build_trainer_app(changed, image="test")[1]) != old_bytes diff --git a/tests/providers/test_deployment_e2e_helper.py b/tests/providers/test_deployment_e2e_helper.py index 2c8bad4..35db6d8 100644 --- a/tests/providers/test_deployment_e2e_helper.py +++ b/tests/providers/test_deployment_e2e_helper.py @@ -5,7 +5,7 @@ import modal import pytest -from lilo.deployments import load, config_path, DeploymentRecord +from lilo.deployments import load, config_path, DeploymentRecord, gpu_count @pytest.mark.parametrize("preset", ["qwen35-9b-lora-16k", "qwen35-4b-fft-64k"]) @@ -21,9 +21,9 @@ def test_e2e_helper_reads_active_deployed_configuration(monkeypatch, preset): monkeypatch.setattr(modal.Dict, "from_name", lambda name: registry) definition, mode = helper["_definition"]("test-frontend", row.spec.name) assert definition.DEFINITION_ID == row.definition_id - assert definition.MAX_CONTEXT_LENGTH == row.spec.model.max_context_length - assert definition.GPUS == row.spec.trainer.resources.gpu_count - assert mode == row.spec.model.parameterization + assert definition.MAX_CONTEXT_LENGTH == row.spec.model['max_context_length'] + assert definition.GPUS == gpu_count(row.spec.trainer['resources']) + assert mode == row.spec.model['parameterization'] assert definition.MAX_TOKENS_PER_MICROBATCH > 0 with pytest.raises(ValueError, match="one active YAML"): helper["_definition"]("test-frontend", "missing") diff --git a/tests/providers/test_deployment_presets.py b/tests/providers/test_deployment_presets.py index a06f9bd..8587c8e 100644 --- a/tests/providers/test_deployment_presets.py +++ b/tests/providers/test_deployment_presets.py @@ -19,7 +19,7 @@ def test_all_packaged_recipes_validate_offline(path): spec = load(path) config = backend_config(spec) - assert config[spec.trainer.backend]["hf_checkpoint"] == "/assets/pending" + assert config[spec.trainer['backend']]["hf_checkpoint"] == "/assets/pending" serving_options(spec) @@ -56,8 +56,8 @@ def test_qwen38_context_parallel_token_budget(context, cp): def test_single_client_recipe_keeps_shared_backend_capacity(): shared = load(config_path("qwen35-9b-lora-16k")) single = load(config_path("qwen35-9b-lora-16k-single")) - assert single.trainer.engine.max_clients_per_instance == 1 - assert single.trainer.resources == shared.trainer.resources + assert single.trainer["engine"]["max_clients_per_instance"] == 1 + assert single.trainer["resources"] == shared.trainer["resources"] assert backend_config(single) == backend_config(shared) diff --git a/tests/test_deployment_cli.py b/tests/test_deployment_cli.py index 8884cd5..4b1535c 100644 --- a/tests/test_deployment_cli.py +++ b/tests/test_deployment_cli.py @@ -71,7 +71,7 @@ def fail(*args, **kwargs): assert registry["pending"][0]["generation"] == row.generation assert "manifest" not in registry and "apply_lock" not in registry new_spec = deepcopy(row.spec) - new_spec.trainer.scaling.max_instances = 2 + new_spec.trainer["scaling"]["max_instances"] = 2 new = DeploymentRecord.create(new_spec, revision="a" * 40, implementation="test") monkeypatch.setattr(subprocess, "run", lambda *args, **kwargs: None) cli.deploy([new]) @@ -132,7 +132,7 @@ def test_compile_pins_revision_at_external_boundary( monkeypatch.setattr(miles_revision, "resolve_miles_commit", lambda: "b" * 40) monkeypatch.setattr(cli, "implementation_fingerprint", lambda _: "runtime") (row,) = cli.compile_configs([path]) - assert row.spec.model.revision == "a" * 40 + assert row.spec.model["revision"] == "a" * 40 assert lookup.call_count == lookups if lookups: lookup.return_value.sha = None diff --git a/tests/test_deployments.py b/tests/test_deployments.py index 781bed8..b45a948 100644 --- a/tests/test_deployments.py +++ b/tests/test_deployments.py @@ -53,7 +53,7 @@ def test_presets_context_topology_and_backend_options(): assert config["native_options"]["recompute_num_layers"] == 1 assert config["extra_args"] == ("--seq-length", "16384") large = recipe("qwen35-9b-lora-64k") - assert large.model.max_context_length == 65536 + assert large.model["max_context_length"] == 65536 assert backend_config(large)["miles"]["actor_num_gpus_per_node"] == 8 fft = backend_config(recipe("qwen35-4b-fft-64k"))["megatron"] assert (fft["tensor_model_parallel_size"], fft["context_parallel_size"]) == (2, 2) @@ -121,14 +121,16 @@ def test_frontend_defaults_and_retained_generations(): large = resolved(recipe("qwen35-9b-lora-64k")) routes = DeploymentRoutes([definition(small), definition(large)]) assert ( - routes.select(small.spec.model.id, "lora").DEFINITION_ID == small.definition_id + routes.select(small.spec.model["id"], "lora").DEFINITION_ID + == small.definition_id ) switched = retain_generations( [small, large], [resolved(recipe("qwen35-9b-lora-64k", routing__default=True))] ) routes = DeploymentRoutes(map(definition, switched)) assert ( - routes.select(small.spec.model.id, "lora").DEFINITION_ID == large.definition_id + routes.select(small.spec.model["id"], "lora").DEFINITION_ID + == large.definition_id ) # Saved model/checkpoint records continue using their original definition. assert ( @@ -152,20 +154,20 @@ def test_ambiguous_model_does_not_get_random_configuration(): ] routes = DeploymentRoutes(map(definition, rows)) with pytest.raises(ValueError, match="ambiguous.*16k.*64k"): - routes.select(rows[0].spec.model.id, "lora") + routes.select(rows[0].spec.model["id"], "lora") assert routes.capabilities() == [] def test_sampling_requires_default_across_training_modes(): lora = resolved() - fft = resolved(recipe("qwen35-4b-fft-64k", model__id=lora.spec.model.id)) + fft = resolved(recipe("qwen35-4b-fft-64k", model__id=lora.spec.model["id"])) routes = DeploymentRoutes(map(definition, [lora, fft])) with pytest.raises(ValueError, match="sampling_default"): - routes.sampling(lora.spec.model.id) - fft.spec.routing.sampling_default = True + routes.sampling(lora.spec.model["id"]) + fft.spec.routing["sampling_default"] = True assert ( DeploymentRoutes(map(definition, [lora, fft])) - .sampling(lora.spec.model.id) + .sampling(lora.spec.model["id"]) .DEFINITION_ID == fft.definition_id ) @@ -221,7 +223,7 @@ async def run(): json={ "session_id": session, "model_seq_id": seq, - "base_model": row.spec.model.id, + "base_model": row.spec.model["id"], "lora_config": {"rank": 32}, }, ) @@ -285,7 +287,7 @@ def test_native_sections_survive_serialization_without_allowlist(): settings = backend_config(spec, "/assets/pinned") config, _ = parse_backend_config(json.loads(json.dumps(settings))) assert config.hf_checkpoint == "/assets/pinned" - assert config.seq_length == spec.model.max_context_length + assert config.seq_length == spec.model["max_context_length"] assert config.provider_overrides["future_provider_option"] == { "layers": [1, 4], "enabled": False, @@ -345,7 +347,7 @@ def test_reserved_environment_is_checked_by_modal_setup(): spec = recipe(trainer__env={"LILO_BACKEND_CONFIG": "oops"}) with pytest.raises(ValueError, match="managed"): - deployment_env(spec.trainer.env) + deployment_env(spec.trainer["env"]) assert deployment_env({"MY_SETTING": "value"}) == {"MY_SETTING": "value"} @@ -357,7 +359,7 @@ def test_record_creation_copies_without_reparsing(): original = asdict(spec) row = DeploymentRecord.create(spec, revision="a" * 40, implementation="test") assert asdict(spec) == original - assert row.spec.model.revision == "a" * 40 + assert row.spec.model["revision"] == "a" * 40 # Keep the existing manifest fields and hash format stable. expected = original | {"model": original["model"] | {"revision": "a" * 40}} expected.pop("routing") @@ -367,8 +369,8 @@ def test_record_creation_copies_without_reparsing(): json.dumps(["test", expected], sort_keys=True).encode() ).hexdigest() ) - row.spec.trainer.config["options"]["lora_rank"] = 64 - assert spec.trainer.config["options"]["lora_rank"] == 32 + row.spec.trainer["config"]["options"]["lora_rank"] = 64 + assert spec.trainer["config"]["options"]["lora_rank"] == 32 saved = row.model_dump_json() assert DeploymentRecord.model_validate_json(saved) == row @@ -386,13 +388,13 @@ def test_python_config_inheritance_and_independent_defaults(tmp_path): first, second = load(path), load(path) assert is_dataclass(first) assert first.name == "custom" - assert first.model.max_context_length == 65536 - assert first.trainer.config["options"]["new_backend_option"] is False - first.trainer.config["options"]["target_modules"].append("extra") - assert "extra" not in second.trainer.config["options"]["target_modules"] + assert first.model["max_context_length"] == 65536 + assert first.trainer["config"]["options"]["new_backend_option"] is False + first.trainer["config"]["options"]["target_modules"].append("extra") + assert "extra" not in second.trainer["config"]["options"]["target_modules"] assert ( "extra" - not in load(config_path("qwen35-9b-lora-16k")).trainer.config["options"][ + not in load(config_path("qwen35-9b-lora-16k")).trainer["config"]["options"][ "target_modules" ] ) @@ -408,10 +410,10 @@ def test_loading_python_config_does_not_call_backend_readers(monkeypatch): backends, "serving_options", lambda *a: pytest.fail("serving read") ) spec = load(config_path("qwen35-9b-lora-16k")) - original_revision = spec.model.revision + original_revision = spec.model["revision"] record = DeploymentRecord.create(spec, revision="a" * 40, implementation="test") - assert record.spec.model.revision == "a" * 40 - assert spec.model.revision == original_revision + assert record.spec.model["revision"] == "a" * 40 + assert spec.model["revision"] == original_revision @pytest.mark.parametrize("source", ["Config = {}", "class Config: pass", "value = 1"]) @@ -440,7 +442,6 @@ def test_no_yaml_config_ingestion(tmp_path): def test_overrides_inherit_replace_and_copy_values(): from lilo.configs.qwen35_9b_lora_16k import Config as Example - from lilo.deployments import Trainer, Resources class Parent(Example): overrides = { @@ -457,26 +458,28 @@ class Child(Parent): } child = Child() - assert child.inference.config["max_running_requests"] == 24 - assert child.trainer.config["options"]["target_modules"] == ["child"] - assert child.trainer.config["options"]["future_option"] == {"enabled": False} - child.trainer.config["options"]["target_modules"].append("changed") - child.trainer.config["options"]["future_option"]["enabled"] = True + assert child.inference["config"]["max_running_requests"] == 24 + assert child.trainer["config"]["options"]["target_modules"] == ["child"] + assert child.trainer["config"]["options"]["future_option"] == {"enabled": False} + child.trainer["config"]["options"]["target_modules"].append("changed") + child.trainer["config"]["options"]["future_option"]["enabled"] = True assert Child.overrides["trainer.config.options.target_modules"] == ["child"] - assert Child().trainer.config["options"]["future_option"] == {"enabled": False} - assert Parent().trainer.config["options"]["target_modules"] == ["parent"] + assert Child().trainer["config"]["options"]["future_option"] == {"enabled": False} + assert Parent().trainer["config"]["options"]["target_modules"] == ["parent"] assert "overrides" not in asdict(child) assert Child(name="keyword").name == "keyword" class Replacement(Parent): - trainer = Trainer(resources=Resources(gpu="H200:8"), config={"options": {}}) + trainer = {"resources": {"gpu": "H200:8"}, "config": {"options": {}}} overrides = {"trainer.config.options.new_option": 1} # A child's complete field replacement wins over its parent's dotted edits. - assert Replacement().trainer.config == {"options": {"new_option": 1}} + assert Replacement().trainer["config"] == {"options": {"new_option": 1}} -@pytest.mark.parametrize("path", ["model.typo", "trainer.missing.value", "typo"]) +@pytest.mark.parametrize( + "path", ["model.missing.value", "trainer.missing.value", "typo"] +) def test_override_typos_fail_with_the_path(path): from lilo.configs.qwen35_9b_lora_16k import Config as Example @@ -485,3 +488,22 @@ class Config(Example): with pytest.raises(ValueError, match=path): Config() + + +def test_plain_sections_fill_defaults_without_sharing_values(): + class Config(BaseConfig): + name = "plain" + model = {"id": "example/model", "max_context_length": 2048} + trainer = {"resources": {"gpu": "H100:4"}, "config": {"future_option": False}} + inference = {"resources": {"gpu": "H200"}} + + first, second = Config(), Config() + assert type(first.model) is dict + assert type(first.trainer) is dict + assert first.model["revision"] == "main" + assert first.trainer["resources"]["cpu"] == 8 + assert first.trainer["config"] == {"future_option": False} + first.inference["scaling"]["max_replicas"] = 2 + first.trainer["env"]["CUSTOM"] = "value" + assert second.inference["scaling"]["max_replicas"] == 8 + assert second.trainer["env"] == {} From 7134f73a680361ae1ba7d8ff4ed84a1fa720e958 Mon Sep 17 00:00:00 2001 From: kailash Date: Wed, 23 Sep 2026 19:25:41 +0000 Subject: [PATCH 16/27] Deploy trainer and inference workers independently with explicit runtime versions --- docs/deployment-configs.md | 57 +++++- docs/deployment-validation.md | 10 +- src/lilo/control_plane/deployments.py | 2 +- src/lilo/deployment_cli.py | 93 ++++++--- src/lilo/deployments.py | 57 +++++- src/lilo/providers/modal/app.py | 29 +-- src/lilo/providers/modal/deployment_apps.py | 87 ++++++--- .../providers/modal/deployment_worker_app.py | 16 ++ src/lilo/providers/modal/fft_pool.py | 15 +- src/lilo/providers/modal/lora_pool.py | 21 ++- tests/providers/conftest.py | 2 +- tests/providers/test_deployment_apps.py | 177 ++++++++++++++++-- tests/providers/test_deployment_e2e_helper.py | 8 +- tests/providers/test_deployment_presets.py | 4 +- tests/providers/test_lora_pool.py | 2 +- tests/providers/test_modal_app.py | 2 +- tests/test_deployment_cli.py | 136 +++++++++++--- tests/test_deployments.py | 50 +++-- 18 files changed, 607 insertions(+), 161 deletions(-) create mode 100644 src/lilo/providers/modal/deployment_worker_app.py diff --git a/docs/deployment-configs.md b/docs/deployment-configs.md index 331a598..d7b71f9 100644 --- a/docs/deployment-configs.md +++ b/docs/deployment-configs.md @@ -71,29 +71,31 @@ lilo deploy config.py compile_configs() deployments.load() → execute config.py → Config() resolve model tag to a Hugging Face commit - DeploymentRecord.create() → copy settings and compute hash + DeploymentRecord.create() → copy settings and compute config hash deploy() retain existing job configurations + deploy new trainer/inference apps; skip existing apps save/pass manifest JSON modal deploy -m lilo.providers.modal.app ``` [`load()`](../src/lilo/deployments.py) executes the file and instantiates its exported `Config` subclass. The returned dataclass goes directly to the orchestration code. There is no dictionary-to-config conversion on this path. -`DeploymentRecord` adds the code identity, pinned Miles revision, configuration hash and active status. JSON is used only to store records and pass them to other processes. Pydantic reconstructs BaseConfig with its dictionary sections when reading those records; it does not import or run the user's config file in a GPU worker. Records preserve the full computed settings rather than a reference to the original Python file. +`DeploymentRecord` adds the pinned Miles revision, configuration hash and active status. JSON is used only to store records and pass them to other processes. Pydantic reconstructs BaseConfig with its dictionary sections when reading those records; it does not import or run the user's config file in a GPU worker. Records preserve the full computed settings rather than a reference to the original Python file. | File | Responsibility | | --- | --- | | [`deployments.py`](../src/lilo/deployments.py) | BaseConfig defaults and overrides, Python file loader, saved record and shared-frontend checks | | [`deployment_cli.py`](../src/lilo/deployment_cli.py) | Revision lookup, manifest updates and Modal deployment | -| [`deployment_apps.py`](../src/lilo/providers/modal/deployment_apps.py) | Trainer functions, inference server classes and process startup | -| [`app.py`](../src/lilo/providers/modal/app.py) | Shared frontend and `app.include()` for generated trainers | +| [`deployment_apps.py`](../src/lilo/providers/modal/deployment_apps.py) | Independent trainer apps, inference provisioners, server classes and process startup | +| [`app.py`](../src/lilo/providers/modal/app.py) | Shared frontend referencing deployed trainer functions | +| [`deployment_worker_app.py`](../src/lilo/providers/modal/deployment_worker_app.py) | Entrypoint for deploying one trainer or inference provisioner | | [`deployment_pool_app.py`](../src/lilo/providers/modal/deployment_pool_app.py) | Construct an inference pool from its saved record | | [`backends/deployment.py`](../src/lilo/backends/deployment.py) | Select the backend configuration reader | -`build_trainer_app(record)` configures resources, secrets, storage and limits. The shared app includes its generated trainer function. When that function starts, `run_trainer()` obtains backend settings and passes them as `LILO_BACKEND_CONFIG` to the existing Miles or Megatron executor. +`build_trainer_app(record)` configures resources, secrets, storage and limits. The CLI deploys it as its own app. The frontend looks up its `trainer` function by app name. When that function starts, `run_trainer()` obtains backend settings and passes them as `LILO_BACKEND_CONFIG` to the existing Miles or Megatron executor. -Inference pools are separate apps created on demand. `build_rollout_app()` constructs a server from the saved configuration. Startup launches SGLang with native options, waits for its health endpoint, then starts the LoRA or FFT sidecar. +Inference pools are separate apps created on demand by a deployed `provision` function. That function keeps the source used when its inference app was deployed, including when an idle pool needs to be recreated after a frontend upgrade. `build_rollout_app()` constructs a server from the saved configuration. Startup launches SGLang with native options, waits for its health endpoint, then starts the LoRA or FFT sidecar. ## Backend options @@ -142,8 +144,47 @@ If multiple configurations serve the same model and training mode, select one wi Every client stores its selected definition ID. Changing routing defaults affects new clients. Changing compute or backend settings creates a new configuration ID, and old configurations remain registered for existing jobs and checkpoints. The CLI serializes applies and retains interrupted deployments for recovery. -The shared app is deployed on each apply. Existing trainer functions use stable captured configuration JSON, and unchanged inference pools retain their apps. This is the shared-app architecture, not independent trainer-app deployment. Source/runtime upgrades still require a separate frontend in this draft. The `src/lilo/configs/` directory is excluded from runtime fingerprints and image source mounts. Its computed values are stored and hashed in each deployment record, so editing a config does not count as a runtime-code upgrade. Backend code changes still do. +## Hashes and independent updates -Existing resource names (`lilo-yaml`, `*-yaml-deployments`, and the `yaml_` definition prefix) are retained so this authoring change does not rename saved resources. They no longer indicate a YAML ingestion path. PyYAML is not a direct Lilo dependency; other installed libraries may depend on it. +There is no source fingerprint and no check that a config matches the current Lilo checkout. Hashes identify settings; they do not certify model support or successful training. + +| Identifier | Inputs | Used for | +| --- | --- | --- | +| `generation` | Complete computed config except routing defaults, plus pinned Miles commit | Saved job/checkpoint configuration and routing | +| `trainer_hash` | Deployment name, model, trainer section, shared deployment settings, Miles commit | Independently deployed trainer app name | +| `inference_hash` | Deployment name, model, inference section, shared deployment settings, adapter rank and target modules | Independently deployed inference provisioner app name | +| Asset hash | Model repository and resolved model commit | Download directory | + +The hashes use SHA-256 over sorted JSON. Trainer/inference app names use the first 24 hex characters. Definition IDs use the first 16 characters of `generation`. Code is retained by the deployed Modal apps, rather than reconstructed from a source fingerprint. + +Both `trainer` and `inference` have a `runtime_version` setting, defaulting to `"1"`. To deploy a code-only trainer update, change `trainer.runtime_version` for the configurations that should use it: + +```python +overrides = { + "trainer.runtime_version": "2", +} +``` + +An inference code update uses `inference.runtime_version` instead. These are operator-selected release labels, not source hashes or validation requirements. Editing source alone does not update an existing worker app. If shared worker code changes, bump each affected role's version. The frontend itself is redeployed on every apply. + +Examples: + +- Change inference concurrency: deploy a new inference provisioner; keep the existing trainer app. +- Change trainer batch settings or trainer runtime version: deploy a new trainer app; keep the inference provisioner. +- Change adapter rank or target modules: update both because inference must load the changed adapters. +- Change routing defaults: keep both worker apps. +- Change one Miles configuration: other Miles configurations and Megatron apps remain deployed as they were. + +The CLI records each successfully deployed worker app before updating the frontend. Retries skip completed apps and recover a worker deployment that succeeded before its registry write. Retained configurations reference old apps; the CLI does not rebuild those apps with new source. Old apps are retained, with GPU trainers scaling to zero when idle. Automatic deletion of unused worker app definitions is not implemented. + +Trainer capacity is enforced per saved definition by control-plane admission and reconciliation. The reusable Modal trainer function has no additional global container cap, so retained jobs do not prevent new definitions from starting. Old and new definitions can consume their configured capacity simultaneously during an update. + +Pools remain scoped to the full deployment definition (and FFT job/version). A trainer update can therefore cause a new job to obtain a separate pool even when it uses the same inference provisioner. Existing pools are not redeployed. + +Backend startup checks and optional smoke tests remain separate from these identifiers. Worker/frontend protocol changes still require compatible APIs or an explicit migration; removing the source fingerprint does not guarantee arbitrary old and new versions interoperate. + +This changes the draft's deployment-record format and replaces its earlier shared-app trainers. Existing deployments from that draft need a fresh frontend/registry or an explicit migration; this change does not silently convert running shared-app trainers. + +Existing resource names (`lilo-yaml`, `*-yaml-deployments`, and the `yaml_` definition prefix) are retained for naming continuity. They no longer indicate a YAML ingestion path. PyYAML is not a direct Lilo dependency; other installed libraries may depend on it. See [validation results](deployment-validation.md) for CPU coverage and the earlier GPU smoke tests. The Python-config migration has not been redeployed. diff --git a/docs/deployment-validation.md b/docs/deployment-validation.md index 75e5e54..d6e6cb3 100644 --- a/docs/deployment-validation.md +++ b/docs/deployment-validation.md @@ -1,6 +1,14 @@ # Deployment configuration validation -These checks exercise the shared-app deployment path in PR #55. Each deployment contains the frontend and generated trainer functions; inference pools are separate apps started on demand. +These checks exercise PR #55. The current implementation deploys trainers and inference provisioners independently; the shared frontend references them by app name. Earlier sections record validation of the previous shared-app implementation. + +## Independent worker apps and config hashes + +Removed the global source fingerprint and its code-upgrade rejection. The config hash identifies saved job settings; separate trainer and inference hashes identify worker apps. Runtime upgrades use explicit per-role `runtime_version` labels. The CLI skips existing worker apps, deploys changed ones before updating the frontend, and records successful workers for retry. Inference provisioners retain their source in an image so idle pools can restart using their original code. + +CPU tests cover inference-only changes, Miles-only updates beside Megatron, runtime-version changes, adapter-shape changes, retained configurations, interrupted deploy recovery, frontend trainer references, remote spawn arguments, saved inference provisioners, and construction of both real Modal worker entrypoints without deploying. Trainer capacity remains enforced per saved definition by the existing control plane; there is no function-wide cap blocking new definitions behind retained jobs. + +Validation: **615 CPU tests passed, 1 skipped**. Ruff and whitespace checks passed. Deployment-command isolation is tested with mocked Modal calls; no apps were redeployed and this architecture has not yet had a live GPU/deployment test. Existing manifests from the earlier shared-app draft require migration or a fresh frontend/registry. Worker API compatibility across future releases and cleanup of unused worker apps remain explicit operational concerns. ## Plain dictionary sections diff --git a/src/lilo/control_plane/deployments.py b/src/lilo/control_plane/deployments.py index 38a63da..1b9fab2 100644 --- a/src/lilo/control_plane/deployments.py +++ b/src/lilo/control_plane/deployments.py @@ -1,4 +1,4 @@ -"""Model-name routing for shared YAML deployments and scoped engines.""" +"""Model-name routing for configured deployments and scoped engines.""" class DeploymentRoutes: diff --git a/src/lilo/deployment_cli.py b/src/lilo/deployment_cli.py index a77408c..673d3c8 100644 --- a/src/lilo/deployment_cli.py +++ b/src/lilo/deployment_cli.py @@ -3,7 +3,6 @@ from __future__ import annotations import argparse -import hashlib import json import os from pathlib import Path @@ -21,29 +20,16 @@ ) -def implementation_fingerprint(miles_commit: str | None) -> str: - """Hash shipped source and dependency declarations, independent of Git checkout.""" - from importlib.metadata import requires - - root = Path(__file__).parent - digest = hashlib.sha256() - for path in sorted(root.rglob("*.py")): - relative = path.relative_to(root) - if relative.parts[0] == "configs": - continue # Config values are hashed separately in DeploymentRecord. - digest.update(str(relative).encode()) - digest.update(path.read_bytes()) - digest.update(json.dumps([miles_commit, sorted(requires("lilo") or [])]).encode()) - return digest.hexdigest() - - def compile_configs(paths): specs = [load(path) for path in paths] validate_frontend(specs) from lilo.providers.modal.miles_revision import resolve_miles_commit - miles_commit = resolve_miles_commit() - implementation = implementation_fingerprint(miles_commit) + miles_commit = ( + resolve_miles_commit() + if any(spec.trainer["backend"] == "miles" for spec in specs) + else None + ) records = [] for spec in specs: revision = spec.model["revision"] @@ -59,8 +45,9 @@ def compile_configs(paths): DeploymentRecord.create( spec, revision=revision, - implementation=implementation, - miles_commit=miles_commit, + miles_commit=miles_commit + if spec.trainer["backend"] == "miles" + else None, ) ) return records @@ -71,10 +58,6 @@ def retain_generations(previous, desired): validate_frontend([row.spec for row in desired]) expected = desired[0] for row in previous: - if row.implementation != expected.implementation: - raise ValueError( - "This draft cannot rebuild retained generations with different Lilo/runtime code. Use a separate frontend for a code upgrade; Config-only changes can retain existing generations." - ) if ( row.spec.deployment != expected.spec.deployment or row.spec.lifecycle != expected.spec.lifecycle @@ -93,6 +76,10 @@ def retain_generations(previous, desired): ] +def worker_apps_ready(row, deployed): + return row.trainer_app_name in deployed and row.inference_app_name in deployed + + def deploy(desired): """Serialize operator applies and retain interrupted attempts for safe recovery.""" import modal @@ -127,8 +114,14 @@ def deploy(desired): raise ValueError( "The frontend already exists without a deployment registry. Choose a new frontend name; an app with no deployment registry cannot be safely updated." ) - # A killed deploy may already have updated Modal. Keep its functions on retry. - rows = {r["generation"]: r for r in [*rows, *registry.get("pending", [])]} + # Only complete worker pairs could have been exposed by a pending frontend. + deployed = set(registry.get("worker_apps", [])) + pending = [ + row + for row in registry.get("pending", []) + if worker_apps_ready(DeploymentRecord.model_validate(row), deployed) + ] + rows = {r["generation"]: r for r in [*rows, *pending]} manifest = retain_generations( [DeploymentRecord.model_validate(row) for row in rows.values()], desired ) @@ -138,8 +131,6 @@ def deploy(desired): MANIFEST_ENV: json.dumps(data), "LILO_APP_NAME": settings["frontend"], } - if desired[0].miles_commit: - env["LILO_MILES_COMMIT"] = desired[0].miles_commit command = [ sys.executable, "-m", @@ -151,6 +142,50 @@ def deploy(desired): if settings["modal"]["environment"]: command += ["--env", settings["modal"]["environment"]] registry.put("pending", data) + # Deploy each worker app once. Retained apps keep their original code. + deployed = set(registry.get("worker_apps", [])) + for row in manifest: + for role, app_name in ( + ("trainer", row.trainer_app_name), + ("inference", row.inference_app_name), + ): + if app_name in deployed: + continue + if not row.active: + raise ValueError( + f"Retained worker {app_name} is missing; restore its original deployment." + ) + # Recover a crash after Modal succeeded but before the registry write. + try: + modal.App.lookup( + app_name, environment_name=settings["modal"]["environment"] + ) + except modal.exception.NotFoundError: + pass + else: + deployed.add(app_name) + registry.put("worker_apps", sorted(deployed)) + continue + worker_env = { + **os.environ, + "LILO_WORKER_DEPLOYMENT": row.model_dump_json(), + "LILO_WORKER_ROLE": role, + } + if row.miles_commit: + worker_env["LILO_MILES_COMMIT"] = row.miles_commit + worker_command = [ + sys.executable, + "-m", + "modal", + "deploy", + "-m", + "lilo.providers.modal.deployment_worker_app", + ] + if settings["modal"]["environment"]: + worker_command += ["--env", settings["modal"]["environment"]] + subprocess.run(worker_command, check=True, env=worker_env) + deployed.add(app_name) + registry.put("worker_apps", sorted(deployed)) subprocess.run(command, check=True, env=env) registry.put("manifest", data) registry.pop("pending", None) diff --git a/src/lilo/deployments.py b/src/lilo/deployments.py index f961936..e3651e5 100644 --- a/src/lilo/deployments.py +++ b/src/lilo/deployments.py @@ -23,6 +23,7 @@ "routing": {"default": False, "sampling_default": False}, "trainer": { "backend": "miles", + "runtime_version": "1", "resources": {"cpu": 8, "memory_mib": 32768, "timeout_s": 86400}, "scaling": {"min_instances": 0, "max_instances": 1}, "engine": {"max_clients_per_instance": 1, "sampler_persistence_concurrency": 8}, @@ -31,6 +32,7 @@ }, "inference": { "backend": "sglang", + "runtime_version": "1", "resources": {"cpu": 8, "memory_mib": 32768, "timeout_s": 86400}, "scaling": { "min_replicas": 0, @@ -123,18 +125,22 @@ def gpu_count(resources): return int(gpu.split(":")[1]) if ":" in gpu else 1 +def settings_hash(settings: dict) -> str: + """Stable identifier for settings, not a claim that they have been validated.""" + return hashlib.sha256(json.dumps(settings, sort_keys=True).encode()).hexdigest() + + class DeploymentRecord(BaseModel): """Saved deployment metadata around a Python configuration. The CLI resolves the model revision before creating this record. Its hash - binds jobs to their original code and configuration across later deploys; + binds jobs to their original configuration across later deploys; active controls whether new clients can select it. """ model_config = ConfigDict(extra="forbid") spec: BaseConfig - implementation: str miles_commit: str | None = None generation: str active: bool = True @@ -145,7 +151,6 @@ def create( spec: BaseConfig, *, revision: str, - implementation: str, miles_commit: str | None = None, ) -> DeploymentRecord: """Record an already-resolved revision without reparsing the configuration.""" @@ -154,16 +159,54 @@ def create( # Changing routing defaults should not restart an existing trainer. identity = asdict(pinned) identity.pop("routing") - generation = hashlib.sha256( - json.dumps([implementation, identity], sort_keys=True).encode() - ).hexdigest() + generation = settings_hash({"config": identity, "miles_commit": miles_commit}) return cls( spec=pinned, - implementation=implementation, generation=generation, miles_commit=miles_commit, ) + @property + def trainer_hash(self) -> str: + """Identify trainer settings and the operator-selected runtime version.""" + return settings_hash( + { + "name": self.spec.name, + "model": self.spec.model, + "trainer": self.spec.trainer, + "deployment": self.spec.deployment, + "miles_commit": self.miles_commit, + } + ) + + @property + def inference_hash(self) -> str: + """Identify inference settings, including the adapter shape it must load.""" + adapter = {} + if self.spec.model["parameterization"] == "lora": + options = self.spec.trainer["config"].get("options", {}) + adapter = { + "lora_rank": options.get("lora_rank"), + "target_modules": options.get("target_modules"), + } + return settings_hash( + { + "name": self.spec.name, + "model": self.spec.model, + "inference": self.spec.inference, + "deployment": self.spec.deployment, + "adapter": adapter, + } + ) + + @property + def trainer_app_name(self) -> str: + return f"lilo-trainer-{self.trainer_hash[:24]}" + + @property + def inference_app_name(self) -> str: + return f"lilo-inference-{self.inference_hash[:24]}" + @property def definition_id(self) -> str: return f"yaml_{self.spec.name}_{self.generation[:16]}" diff --git a/src/lilo/providers/modal/app.py b/src/lilo/providers/modal/app.py index b65338b..4942cd4 100644 --- a/src/lilo/providers/modal/app.py +++ b/src/lilo/providers/modal/app.py @@ -56,19 +56,21 @@ ) SETTINGS = frontend_settings() -APP_NAME = SETTINGS.deployment['frontend'] -ROUTING_REGION = SETTINGS.deployment['modal']['region'] +APP_NAME = SETTINGS.deployment["frontend"] +ROUTING_REGION = SETTINGS.deployment["modal"]["region"] MODEL_ASSET_ROOT = "/assets" -SESSION_IDLE_TIMEOUT = SETTINGS.lifecycle['session_idle_timeout_s'] -FFT_POOL_IDLE_TIMEOUT = LORA_POOL_IDLE_TIMEOUT = SETTINGS.lifecycle['pool_idle_timeout_s'] +SESSION_IDLE_TIMEOUT = SETTINGS.lifecycle["session_idle_timeout_s"] +FFT_POOL_IDLE_TIMEOUT = LORA_POOL_IDLE_TIMEOUT = SETTINGS.lifecycle[ + "pool_idle_timeout_s" +] FFT_POOL_TOUCH_INTERVAL = 60.0 LORA_POOL_CHECK_INTERVAL = 60.0 -SWEEP_PERIOD = modal.Period(seconds=SETTINGS.lifecycle['sweep_interval_s']) +SWEEP_PERIOD = modal.Period(seconds=SETTINGS.lifecycle["sweep_interval_s"]) CHECKPOINT_READ_LOCK = asyncio.Lock() _pool_touches: dict[str, float] = {} _lora_pool_gateways: dict[str, tuple[float, str]] = {} _lora_pool_checks: dict[str, asyncio.Lock] = {} -CHECKPOINT_VOLUME_NAME = SETTINGS.deployment['storage']['checkpoints'] +CHECKPOINT_VOLUME_NAME = SETTINGS.deployment["storage"]["checkpoints"] checkpoint_volume = modal.Volume.from_name( CHECKPOINT_VOLUME_NAME, create_if_missing=True, version=2 ) @@ -79,8 +81,6 @@ "LILO_APP_NAME": APP_NAME, } app = modal.App(APP_NAME) -for definition in DEFINITIONS: - app.include(definition.app) async def _read_checkpoint_metadata(uri: str) -> dict[str, object]: @@ -110,13 +110,13 @@ async def _delete_checkpoint(uri: str) -> None: .add_local_python_source("lilo", ignore=ignore_config_source) ) model_assets = modal.Volume.from_name( - SETTINGS.deployment['storage']['assets'], + SETTINGS.deployment["storage"]["assets"], create_if_missing=True, ) -API_SECRET_NAME = SETTINGS.deployment['secrets']['api'] -HF_SECRET_NAME = SETTINGS.deployment['secrets']['huggingface'] +API_SECRET_NAME = SETTINGS.deployment["secrets"]["api"] +HF_SECRET_NAME = SETTINGS.deployment["secrets"]["huggingface"] proxy_secret = modal.Secret.from_name( - SETTINGS.deployment['secrets']['sampler_proxy'], + SETTINGS.deployment["secrets"]["sampler_proxy"], required_keys=["MODAL_PROXY_TOKEN_ID", "MODAL_PROXY_TOKEN_SECRET"], ) @@ -395,8 +395,9 @@ async def spawn(delay_seconds: float) -> str: async def _spawn_engine(definition_id: str, instance_id: str) -> str: if error := await deployment_error(definition_id): raise ValueError(error) - engine = module_for(definition_id).ENGINE_FUNCTION - call = await engine.spawn.aio(instance_id) + definition = module_for(definition_id) + engine = definition.ENGINE_FUNCTION + call = await engine.spawn.aio(instance_id, definition.RESOLVED.model_dump_json()) return call.object_id diff --git a/src/lilo/providers/modal/deployment_apps.py b/src/lilo/providers/modal/deployment_apps.py index d6bccad..ceb8c6a 100644 --- a/src/lilo/providers/modal/deployment_apps.py +++ b/src/lilo/providers/modal/deployment_apps.py @@ -1,6 +1,6 @@ """Modal app builders shared by all Python-configured model deployments. -Trainer functions live in the frontend app. Rollout pools are separate apps, +Trainers and inference provisioners are independently deployed apps. Pools are created on demand with the existing LoRA/FFT pool lifecycle. """ @@ -37,11 +37,6 @@ def frontend_settings(): validate_frontend(active) if len({row.definition_id for row in deployments}) != len(deployments): raise ValueError("duplicate deployment generation") - commits = {row.miles_commit for row in deployments if row.miles_commit} - if len(commits) > 1: - raise ValueError("manifest contains different Miles runtimes") - if commits: - os.environ["LILO_MILES_COMMIT"] = commits.pop() return active[0] @@ -93,16 +88,9 @@ def deployment_env(values): def build_trainer_app(resolved: DeploymentRecord, *, image=None): - # Admission metadata must not change the serialized function for a running - # configuration when a default switches or an older configuration drains. - resolved = resolved.model_copy(deep=True) - resolved.active = True - resolved.spec.routing = {"default": False, "sampling_default": False} spec = resolved.spec - config_json = json.dumps( - resolved.model_dump(mode="json"), sort_keys=True, separators=(",", ":") - ) - app = modal.App(f"{spec.deployment['frontend']}-{resolved.definition_id}") + trainer_hash = resolved.trainer_hash + app = modal.App(resolved.trainer_app_name) resource = spec.trainer["resources"] from .deployment import trainer_deployment_env @@ -113,7 +101,7 @@ def build_trainer_app(resolved: DeploymentRecord, *, image=None): } @app.function( - name=resolved.definition_id, + name="trainer", serialized=True, image=image if image is not None else image_for(spec.trainer["backend"]), gpu=resource["gpu"], @@ -121,15 +109,20 @@ def build_trainer_app(resolved: DeploymentRecord, *, image=None): cpu=resource["cpu"], memory=resource["memory_mib"], timeout=resource["timeout_s"], - max_containers=spec.trainer["scaling"]["max_instances"], + # Admission/reconciliation caps each definition. A function-wide cap + # would block new definitions behind retained jobs sharing this app. + max_containers=None, min_containers=0, single_use_containers=True, volumes=volumes_for(spec), secrets=secrets_for(spec, training=True), env=env, ) - def trainer(instance_id: str): - run_trainer(DeploymentRecord.model_validate_json(config_json), instance_id) + def trainer(instance_id: str, config_json: str): + record = DeploymentRecord.model_validate_json(config_json) + if record.trainer_hash != trainer_hash: + raise ValueError("trainer settings do not match the deployed app") + run_trainer(record, instance_id) return app, trainer @@ -210,8 +203,10 @@ def definition_from_spec(resolved, *, register_trainer=True, image=None): // (native.get("dp_size", 1) if native.get("enable_dp_attention") else 1), ) if register_trainer: - definition.app, definition.ENGINE_FUNCTION = build_trainer_app( - resolved, image=image + definition.ENGINE_FUNCTION = modal.Function.from_name( + resolved.trainer_app_name, + "trainer", + environment_name=spec.deployment["modal"]["environment"], ) return definition @@ -247,9 +242,9 @@ def build_rollout_app(resolved, pool, *, image=None): **serving_options(spec), } if lora: - config = backend_config(spec)["miles"] from lilo.backends.miles_config import MilesBackendConfig + config = backend_config(spec)["miles"] targets = MilesBackendConfig(**config).peft_target_modules options.update(enable_lora=True, max_lora_rank=config["max_lora_rank"]) options.setdefault("lora_target_modules", list(targets)) @@ -338,3 +333,51 @@ def stop(self): terminate(getattr(self, "sglang", None)) return app, Server + + +def provision_pool(record, pool): + """Ask the saved inference app to create a pool using its original code.""" + provision = modal.Function.from_name( + record.inference_app_name, + "provision", + environment_name=record.spec.deployment["modal"]["environment"], + ) + return provision.remote(record.model_dump_json(), pool.as_dict()) + + +def build_inference_app(record, *, image=None): + """Freeze pool-building code so idle pools can restart after frontend upgrades.""" + from .image_dependencies import ( + CORE_PACKAGES, + STITCH_PACKAGE, + TINKER_PACKAGE, + ignore_config_source, + ) + + if image is None: + image = ( + modal.Image.debian_slim(python_version="3.12") + .apt_install("git") + .pip_install( + *CORE_PACKAGES, STITCH_PACKAGE, TINKER_PACKAGE, "huggingface-hub" + ) + .add_local_python_source("lilo", copy=True, ignore=ignore_config_source) + ) + app = modal.App(record.inference_app_name) + inference_hash = record.inference_hash + + @app.function(name="provision", image=image, serialized=True, timeout=1800) + def provision(config_json: str, pool_data: dict): + from .lora_pool import LoraPoolSpec, deploy_pool as deploy_lora + from .fft_pool import FFTPoolSpec, deploy_pool as deploy_fft + + saved = DeploymentRecord.model_validate_json(config_json) + if saved.inference_hash != inference_hash: + raise ValueError("inference settings do not match the deployed app") + if pool_data["definition_id"] != saved.definition_id: + raise ValueError("pool definition does not match deployment") + if saved.spec.model["parameterization"] == "lora": + return deploy_lora(LoraPoolSpec.from_dict(pool_data), record=saved) + return deploy_fft(FFTPoolSpec.from_dict(pool_data), record=saved) + + return app, provision diff --git a/src/lilo/providers/modal/deployment_worker_app.py b/src/lilo/providers/modal/deployment_worker_app.py new file mode 100644 index 0000000..ed8fac8 --- /dev/null +++ b/src/lilo/providers/modal/deployment_worker_app.py @@ -0,0 +1,16 @@ +"""Deploy one trainer or inference provisioner; the frontend references its name.""" + +import os + +from lilo.deployments import DeploymentRecord +from .deployment_apps import build_trainer_app, build_inference_app + +record = DeploymentRecord.model_validate_json(os.environ["LILO_WORKER_DEPLOYMENT"]) +if record.miles_commit: + os.environ["LILO_MILES_COMMIT"] = record.miles_commit +if os.environ["LILO_WORKER_ROLE"] == "trainer": + app, trainer = build_trainer_app(record) +elif os.environ["LILO_WORKER_ROLE"] == "inference": + app, provision = build_inference_app(record) +else: + raise ValueError("unknown worker role") diff --git a/src/lilo/providers/modal/fft_pool.py b/src/lilo/providers/modal/fft_pool.py index 45d6bdd..8c57c86 100644 --- a/src/lilo/providers/modal/fft_pool.py +++ b/src/lilo/providers/modal/fft_pool.py @@ -119,7 +119,7 @@ async def pool_gateway(spec: FFTPoolSpec) -> str: return await ModalFlashPool(spec.app_name, "Server").gateway_url_async() -def deploy_pool(spec: FFTPoolSpec) -> str: +def deploy_pool(spec: FFTPoolSpec, *, record=None) -> str: pool = ModalFlashPool(spec.app_name, "Server") try: return pool.gateway_url() @@ -128,12 +128,19 @@ def deploy_pool(spec: FFTPoolSpec) -> str: if not isinstance(exc, modal.exception.NotFoundError): raise + if record is None: + from .deployment_apps import pool_deployment, provision_pool + + saved = pool_deployment(spec.definition_id) + if saved is None: + raise ValueError(f"missing recorded deployment: {spec.definition_id}") + return provision_pool(saved, spec) modal_cli = shutil.which("modal") if modal_cli is None: raise RuntimeError("modal CLI is unavailable") - from .deployment_apps import pool_environment + from .deployment_apps import POOL_CONFIG_ENV - recipe_env = pool_environment(spec.definition_id) + recipe_env = {POOL_CONFIG_ENV: record.model_dump_json()} env = {**os.environ, **spec.env(), **recipe_env} command = [ modal_cli, @@ -143,7 +150,7 @@ def deploy_pool(spec: FFTPoolSpec) -> str: "--name", spec.app_name, ] - environment = os.environ.get("MODAL_ENVIRONMENT") + environment = record.spec.deployment["modal"]["environment"] if environment: command.extend(["--env", environment]) subprocess.run(command, env=env, check=True) diff --git a/src/lilo/providers/modal/lora_pool.py b/src/lilo/providers/modal/lora_pool.py index e082ede..42635f0 100644 --- a/src/lilo/providers/modal/lora_pool.py +++ b/src/lilo/providers/modal/lora_pool.py @@ -22,7 +22,7 @@ def __post_init__(self) -> None: object.__setattr__( self, "revision", - _implementation_revision(self.definition_id), + _definition_revision(self.definition_id), ) @classmethod @@ -53,7 +53,7 @@ async def pool_gateway(spec: LoraPoolSpec) -> str: return await ModalFlashPool(spec.app_name, "Server").gateway_url_async() -def deploy_pool(spec: LoraPoolSpec) -> str: +def deploy_pool(spec: LoraPoolSpec, *, record=None) -> str: pool = ModalFlashPool(spec.app_name, "Server") try: return pool.gateway_url() @@ -62,12 +62,19 @@ def deploy_pool(spec: LoraPoolSpec) -> str: if not isinstance(exc, modal.exception.NotFoundError): raise + if record is None: + from .deployment_apps import pool_deployment, provision_pool + + saved = pool_deployment(spec.definition_id) + if saved is None: + raise ValueError(f"missing recorded deployment: {spec.definition_id}") + return provision_pool(saved, spec) modal_cli = shutil.which("modal") if modal_cli is None: raise RuntimeError("modal CLI is unavailable") - from .deployment_apps import pool_environment + from .deployment_apps import POOL_CONFIG_ENV - recipe_env = pool_environment(spec.definition_id) + recipe_env = {POOL_CONFIG_ENV: record.model_dump_json()} command = [ modal_cli, "deploy", @@ -76,7 +83,7 @@ def deploy_pool(spec: LoraPoolSpec) -> str: "--name", spec.app_name, ] - environment = os.environ.get("MODAL_ENVIRONMENT") + environment = record.spec.deployment["modal"]["environment"] if environment: command.extend(["--env", environment]) subprocess.run(command, env={**os.environ, **spec.env(), **recipe_env}, check=True) @@ -100,7 +107,7 @@ def stop_pool(spec: LoraPoolSpec) -> None: result.check_returncode() -def _implementation_revision(definition_id: str) -> str: +def _definition_revision(definition_id: str) -> str: if not definition_id.startswith("yaml_"): - raise ValueError(f"expected a YAML deployment id: {definition_id}") + raise ValueError(f"expected a configured deployment id: {definition_id}") return definition_id.rsplit("_", 1)[-1] diff --git a/tests/providers/conftest.py b/tests/providers/conftest.py index 2e516ba..5400f8a 100644 --- a/tests/providers/conftest.py +++ b/tests/providers/conftest.py @@ -10,7 +10,7 @@ json.dumps( [ DeploymentRecord.create( - load(config_path(name)), revision="a" * 40, implementation="tests" + load(config_path(name)), revision="a" * 40 ).model_dump(mode="json") for name in ("qwen35-9b-fft-64k", "qwen35-9b-lora-16k") ] diff --git a/tests/providers/test_deployment_apps.py b/tests/providers/test_deployment_apps.py index 11859c4..4acd304 100644 --- a/tests/providers/test_deployment_apps.py +++ b/tests/providers/test_deployment_apps.py @@ -12,9 +12,7 @@ def deployment(preset="qwen35-9b-lora-16k"): - return DeploymentRecord.create( - load(config_path(preset)), revision="a" * 40, implementation="runtime" - ) + return DeploymentRecord.create(load(config_path(preset)), revision="a" * 40) class App: @@ -56,13 +54,13 @@ def test_trainer_declaration_and_executor_configuration( builders, monkeypatch, preset, backend, clients, nproc ): row = deployment(preset) - row.spec.deployment['storage']['checkpoints'] = "test-custom-checkpoints" + row.spec.deployment["storage"]["checkpoints"] = "test-custom-checkpoints" image = object() app, trainer = deployment_apps.build_trainer_app(row, image=image) - declaration, _ = app.functions[row.definition_id] + declaration, _ = app.functions["trainer"] assert declaration["gpu"] == "H100:4" assert declaration["region"] == "us-west" - assert declaration["max_containers"] == 1 + assert declaration["max_containers"] is None assert declaration["single_use_containers"] is True assert declaration["image"] is image calls = [] @@ -80,7 +78,7 @@ def test_trainer_declaration_and_executor_configuration( "run_engine_with_backend", lambda *args, **kwargs: calls.append((args, kwargs)), ) - trainer("instance-a") + trainer("instance-a", row.model_dump_json()) args, kwargs = calls[0] assert args == ("store", f"lilo.backends.{backend}:build_executor") assert kwargs["max_models"] == clients @@ -88,7 +86,7 @@ def test_trainer_declaration_and_executor_configuration( assert kwargs["backend_env"]["LILO_CHECKPOINT_VOLUME"] == "test-custom-checkpoints" assert kwargs["backend_env"]["LILO_BASE_MODEL_REVISION"] == "a" * 40 config = json.loads(kwargs["backend_env"]["LILO_BACKEND_CONFIG"]) - assert config[row.spec.trainer['backend']]["hf_checkpoint"] == row.asset_path + assert config[row.spec.trainer["backend"]]["hf_checkpoint"] == row.asset_path assert config["checkpoint_dir"] == "/checkpoints" assert reloaded == [True] @@ -108,7 +106,7 @@ def test_pool_starts_native_server_and_correct_sidecar(builders, monkeypatch, ki app, server = deployment_apps.build_rollout_app(row, pool, image="test-image") settings, _ = app.servers["Server"] assert app.name == pool.app_name - assert settings["gpu"] == row.spec.inference['resources']['gpu'] + assert settings["gpu"] == row.spec.inference["resources"]["gpu"] assert settings["min_containers"] == 0 assert settings["target_concurrency"] == 16 assert settings["compute_region"] == "us-west" @@ -138,7 +136,7 @@ def test_pool_starts_native_server_and_correct_sidecar(builders, monkeypatch, ki assert commands[0][2] == "lilo.inference.native_sglang" assert commands[0][3] == row.asset_path native = json.loads(commands[0][4]) - assert native["context_length"] == row.spec.model['max_context_length'] + assert native["context_length"] == row.spec.model["max_context_length"] if kind == "lora": assert native["enable_lora"] is True assert native["max_lora_rank"] == 32 @@ -165,7 +163,9 @@ def test_pool_subprocess_receives_recorded_generation(monkeypatch): row = deployment() monkeypatch.setenv(deployment_apps.MANIFEST_ENV, json.dumps([row.model_dump()])) env = deployment_apps.pool_environment(row.definition_id) - assert json.loads(env[deployment_apps.POOL_CONFIG_ENV])["generation"] == row.generation + assert ( + json.loads(env[deployment_apps.POOL_CONFIG_ENV])["generation"] == row.generation + ) with pytest.raises(ValueError, match="missing recorded"): deployment_apps.pool_environment("yaml_missing_123") with pytest.raises(ValueError, match="missing recorded"): @@ -232,13 +232,16 @@ def test_admission_changes_preserve_serialized_trainer(builders): old_bytes = serialize(deployment_apps.build_trainer_app(first, image="test")[1]) changed = first.model_copy(deep=True) changed.active = False - changed.spec.routing['default'] = not first.spec.routing['default'] - changed.spec.routing['sampling_default'] = True + changed.spec.routing["default"] = not first.spec.routing["default"] + changed.spec.routing["sampling_default"] = True new_bytes = serialize(deployment_apps.build_trainer_app(changed, image="test")[1]) assert new_bytes == old_bytes - assert first.active is True and first.spec.routing['default'] is True - changed.spec.trainer['resources']['gpu'] = "H200:4" - assert serialize(deployment_apps.build_trainer_app(changed, image="test")[1]) != old_bytes + assert first.active is True and first.spec.routing["default"] is True + changed.spec.trainer["resources"]["gpu"] = "H200:4" + assert ( + serialize(deployment_apps.build_trainer_app(changed, image="test")[1]) + != old_bytes + ) @pytest.mark.parametrize("kind", ["lora", "full"]) @@ -272,10 +275,148 @@ def gateway_url(self): "run", lambda command, **kwargs: calls.append((command, kwargs)), ) - assert module.deploy_pool(spec) == "https://pool" + assert module.deploy_pool(spec, record=row) == "https://pool" command, kwargs = calls[0] - assert command[command.index("-m") + 1] == "lilo.providers.modal.deployment_pool_app" + assert ( + command[command.index("-m") + 1] == "lilo.providers.modal.deployment_pool_app" + ) assert ( json.loads(kwargs["env"][deployment_apps.POOL_CONFIG_ENV])["generation"] == row.generation ) + + +def test_frontend_uses_deployed_trainer_without_building_it(monkeypatch): + row = deployment() + calls = [] + monkeypatch.setattr( + deployment_apps, + "build_trainer_app", + lambda *a, **k: pytest.fail("frontend must not rebuild trainer"), + ) + monkeypatch.setattr( + modal.Function, + "from_name", + lambda *a, **k: calls.append((a, k)) or "remote-trainer", + ) + definition = deployment_apps.definition_from_spec(row) + assert definition.ENGINE_FUNCTION == "remote-trainer" + assert calls[0][0] == (row.trainer_app_name, "trainer") + + +@pytest.mark.parametrize("kind", ["lora", "full"]) +def test_missing_pool_uses_saved_provisioner(monkeypatch, kind): + from lilo.providers.modal import fft_pool, lora_pool + + row = deployment("qwen35-9b-lora-16k" if kind == "lora" else "qwen35-4b-fft-64k") + monkeypatch.setenv(deployment_apps.MANIFEST_ENV, json.dumps([row.model_dump()])) + module = lora_pool if kind == "lora" else fft_pool + spec = ( + LoraPoolSpec(row.definition_id) + if kind == "lora" + else FFTPoolSpec.base(row.definition_id) + ) + + class MissingPool: + def __init__(self, *args): + pass + + def gateway_url(self): + raise modal.exception.NotFoundError("not deployed") + + monkeypatch.setattr(module, "ModalFlashPool", MissingPool) + calls = [] + monkeypatch.setattr( + modal.Function, + "from_name", + lambda name, function, **kw: ( + calls.append((name, function)) + or SimpleNamespace(remote=lambda record, pool: "https://saved-runtime") + ), + ) + monkeypatch.setattr( + module.subprocess, "run", lambda *a, **k: pytest.fail("frontend rebuilt pool") + ) + assert module.deploy_pool(spec) == "https://saved-runtime" + assert calls == [(row.inference_app_name, "provision")] + + +def test_provisioner_rejects_wrong_settings_and_uses_saved_record( + builders, monkeypatch +): + from lilo.providers.modal import lora_pool + + row = deployment() + app, provision = deployment_apps.build_inference_app(row, image="test") + assert app.name == row.inference_app_name + calls = [] + monkeypatch.setattr( + lora_pool, + "deploy_pool", + lambda pool, *, record: calls.append((pool, record)) or "https://pool", + ) + pool = LoraPoolSpec(row.definition_id) + assert provision(row.model_dump_json(), pool.as_dict()) == "https://pool" + assert calls[0][1].inference_hash == row.inference_hash + changed = row.model_copy(deep=True) + changed.spec.inference["runtime_version"] = "2" + with pytest.raises(ValueError, match="inference settings"): + provision(changed.model_dump_json(), pool.as_dict()) + + +@pytest.mark.parametrize("role", ["trainer", "inference"]) +def test_real_worker_entrypoint_constructs_offline(monkeypatch, role): + import os + import subprocess + import sys + + row = deployment() + env = { + **os.environ, + "LILO_WORKER_DEPLOYMENT": row.model_dump_json(), + "LILO_WORKER_ROLE": role, + } + result = subprocess.run( + [ + sys.executable, + "-c", + """ +import modal +from lilo.providers.modal import deployment_apps +deployment_apps.image_for = lambda backend: modal.Image.debian_slim() +import lilo.providers.modal.deployment_worker_app as worker +assert worker.app.name.startswith("lilo-") +print(worker.app.name) +""", + ], + env=env, + capture_output=True, + text=True, + timeout=30, + ) + assert result.returncode == 0, result.stderr + assert f"lilo-{role}-" in result.stdout + + +def test_spawn_passes_job_configuration_to_saved_trainer(monkeypatch): + import importlib + + app = importlib.import_module("lilo.providers.modal.app") + row = deployment() + calls = [] + + async def spawn(instance_id, config_json): + calls.append((instance_id, config_json)) + return SimpleNamespace(object_id="call-id") + + async def no_error(definition_id): + return None + + definition = SimpleNamespace( + RESOLVED=row, + ENGINE_FUNCTION=SimpleNamespace(spawn=SimpleNamespace(aio=spawn)), + ) + monkeypatch.setattr(app, "deployment_error", no_error) + monkeypatch.setattr(app, "module_for", lambda _: definition) + assert asyncio.run(app._spawn_engine(row.definition_id, "instance")) == "call-id" + assert calls == [("instance", row.model_dump_json())] diff --git a/tests/providers/test_deployment_e2e_helper.py b/tests/providers/test_deployment_e2e_helper.py index 35db6d8..1f50bbe 100644 --- a/tests/providers/test_deployment_e2e_helper.py +++ b/tests/providers/test_deployment_e2e_helper.py @@ -13,7 +13,7 @@ def test_e2e_helper_reads_active_deployed_configuration(monkeypatch, preset): helper = runpy.run_path( str(Path(__file__).parents[2] / "scripts/e2e_engine_definition.py") ) - row = DeploymentRecord.create(load(config_path(preset)), revision="a" * 40, implementation="test") + row = DeploymentRecord.create(load(config_path(preset)), revision="a" * 40) retired = row.model_copy(update={"active": False, "generation": "b" * 64}) registry = SimpleNamespace( get=lambda *args: [retired.model_dump(), row.model_dump()] @@ -21,9 +21,9 @@ def test_e2e_helper_reads_active_deployed_configuration(monkeypatch, preset): monkeypatch.setattr(modal.Dict, "from_name", lambda name: registry) definition, mode = helper["_definition"]("test-frontend", row.spec.name) assert definition.DEFINITION_ID == row.definition_id - assert definition.MAX_CONTEXT_LENGTH == row.spec.model['max_context_length'] - assert definition.GPUS == gpu_count(row.spec.trainer['resources']) - assert mode == row.spec.model['parameterization'] + assert definition.MAX_CONTEXT_LENGTH == row.spec.model["max_context_length"] + assert definition.GPUS == gpu_count(row.spec.trainer["resources"]) + assert mode == row.spec.model["parameterization"] assert definition.MAX_TOKENS_PER_MICROBATCH > 0 with pytest.raises(ValueError, match="one active YAML"): helper["_definition"]("test-frontend", "missing") diff --git a/tests/providers/test_deployment_presets.py b/tests/providers/test_deployment_presets.py index 8587c8e..facaeac 100644 --- a/tests/providers/test_deployment_presets.py +++ b/tests/providers/test_deployment_presets.py @@ -19,7 +19,7 @@ def test_all_packaged_recipes_validate_offline(path): spec = load(path) config = backend_config(spec) - assert config[spec.trainer['backend']]["hf_checkpoint"] == "/assets/pending" + assert config[spec.trainer["backend"]]["hf_checkpoint"] == "/assets/pending" serving_options(spec) @@ -38,7 +38,7 @@ def test_moe_rollout_preserves_attention_data_parallelism(): assert options["tp_size"] == options["dp_size"] == options["ep_size"] == 4 assert options["enable_dp_attention"] is True definition = definition_from_spec( - DeploymentRecord.create(spec, revision="a" * 40, implementation="test"), register_trainer=False + DeploymentRecord.create(spec, revision="a" * 40), register_trainer=False ) assert definition.ROLLOUT_GPUS == 4 assert definition.ROLLOUT_TENSOR_PARALLEL_SIZE == 1 diff --git a/tests/providers/test_lora_pool.py b/tests/providers/test_lora_pool.py index e1dc9f5..8d303df 100644 --- a/tests/providers/test_lora_pool.py +++ b/tests/providers/test_lora_pool.py @@ -44,5 +44,5 @@ def test_pool_revision_comes_from_resolved_generation(): def test_python_definition_cannot_choose_a_pool_revision(): import pytest - with pytest.raises(ValueError, match="YAML deployment id"): + with pytest.raises(ValueError, match="configured deployment id"): LoraPoolSpec("qwen3_5_9b_base_miles_lora_16k") diff --git a/tests/providers/test_modal_app.py b/tests/providers/test_modal_app.py index 5a87921..b70ab67 100644 --- a/tests/providers/test_modal_app.py +++ b/tests/providers/test_modal_app.py @@ -15,7 +15,7 @@ def definition_id(preset): return DeploymentRecord.create( - load(config_path(preset)), revision="a" * 40, implementation="tests" + load(config_path(preset)), revision="a" * 40 ).definition_id diff --git a/tests/test_deployment_cli.py b/tests/test_deployment_cli.py index 4b1535c..7fd8d58 100644 --- a/tests/test_deployment_cli.py +++ b/tests/test_deployment_cli.py @@ -22,7 +22,6 @@ def deployment(): return DeploymentRecord.create( load(config_path("qwen35-9b-lora-16k")), revision="a" * 40, - implementation="test", ) @@ -46,8 +45,10 @@ def run(command, **kwargs): assert "apply_lock" in registry assert "pending" in registry assert "manifest" not in registry - assert "lilo.providers.modal.app" in command - seen.extend(json.loads(kwargs["env"][MANIFEST_ENV])) + if "lilo.providers.modal.app" in command: + seen.extend(json.loads(kwargs["env"][MANIFEST_ENV])) + else: + assert "lilo.providers.modal.deployment_worker_app" in command monkeypatch.setattr(subprocess, "run", run) cli.deploy([row]) @@ -72,12 +73,11 @@ def fail(*args, **kwargs): assert "manifest" not in registry and "apply_lock" not in registry new_spec = deepcopy(row.spec) new_spec.trainer["scaling"]["max_instances"] = 2 - new = DeploymentRecord.create(new_spec, revision="a" * 40, implementation="test") + new = DeploymentRecord.create(new_spec, revision="a" * 40) monkeypatch.setattr(subprocess, "run", lambda *args, **kwargs: None) cli.deploy([new]) assert [(r["generation"], r["active"]) for r in registry["manifest"]] == [ (new.generation, True), - (row.generation, False), ] @@ -130,7 +130,6 @@ def test_compile_pins_revision_at_external_boundary( lookup = Mock(return_value=SimpleNamespace(sha="a" * 40)) monkeypatch.setattr(huggingface_hub.HfApi, "model_info", lookup) monkeypatch.setattr(miles_revision, "resolve_miles_commit", lambda: "b" * 40) - monkeypatch.setattr(cli, "implementation_fingerprint", lambda _: "runtime") (row,) = cli.compile_configs([path]) assert row.spec.model["revision"] == "a" * 40 assert lookup.call_count == lookups @@ -140,24 +139,6 @@ def test_compile_pins_revision_at_external_boundary( cli.compile_configs([path]) -def test_config_edits_do_not_change_runtime_fingerprint(tmp_path, monkeypatch): - import importlib.metadata - - root = tmp_path / "lilo" - (root / "configs").mkdir(parents=True) - runtime = root / "runtime.py" - runtime.write_text("runtime = 1") - config = root / "configs" / "example.py" - config.write_text("gpu = 'H100:4'") - monkeypatch.setattr(cli, "__file__", str(root / "deployment_cli.py")) - monkeypatch.setattr(importlib.metadata, "requires", lambda _: []) - before = cli.implementation_fingerprint("miles-commit") - config.write_text("gpu = 'H200:8'") - assert cli.implementation_fingerprint("miles-commit") == before - runtime.write_text("runtime = 2") - assert cli.implementation_fingerprint("miles-commit") != before - - def test_worker_source_mount_excludes_authoring_configs(): from pathlib import Path from lilo.providers.modal.image_dependencies import ignore_config_source @@ -167,3 +148,110 @@ def test_worker_source_mount_excludes_authoring_configs(): assert ignore_config_source(Path("data.json")) assert not ignore_config_source(Path("deployments.py")) assert not ignore_config_source(Path("backends/miles_config.py")) + + +def test_only_changed_worker_is_deployed(registry, monkeypatch): + row = deployment() + calls = [] + + def run(command, **kwargs): + if "lilo.providers.modal.deployment_worker_app" in command: + calls.append(kwargs["env"]["LILO_WORKER_ROLE"]) + else: + calls.append("frontend") + + monkeypatch.setattr(subprocess, "run", run) + cli.deploy([row]) + assert calls == ["trainer", "inference", "frontend"] + + calls.clear() + cli.deploy([row]) + assert calls == ["frontend"] + + changed = deepcopy(row.spec) + changed.inference["scaling"]["max_replicas"] = 6 + new = DeploymentRecord.create(changed, revision="a" * 40) + calls.clear() + cli.deploy([new]) + assert calls == ["inference", "frontend"] + assert new.trainer_app_name == row.trainer_app_name + assert registry["manifest"][1]["active"] is False + + changed.trainer["runtime_version"] = "new-trainer-code" + newest = DeploymentRecord.create(changed, revision="a" * 40) + calls.clear() + cli.deploy([newest]) + assert calls == ["trainer", "frontend"] + assert newest.inference_app_name == new.inference_app_name + + +def test_backend_update_does_not_redeploy_other_models(registry, monkeypatch): + miles = deployment() + fft = DeploymentRecord.create( + load(config_path("qwen35-4b-fft-64k")), + revision="a" * 40, + ) + calls = [] + + def run(command, **kwargs): + if "LILO_WORKER_DEPLOYMENT" in kwargs["env"]: + calls.append( + ( + kwargs["env"]["LILO_WORKER_ROLE"], + json.loads(kwargs["env"]["LILO_WORKER_DEPLOYMENT"])["spec"]["name"], + ) + ) + + monkeypatch.setattr(subprocess, "run", run) + cli.deploy([miles, fft]) + calls.clear() + spec = deepcopy(miles.spec) + spec.trainer["runtime_version"] = "2" + cli.deploy([DeploymentRecord.create(spec, revision="a" * 40), fft]) + assert calls == [("trainer", miles.spec.name)] + + +def test_retry_preserves_successfully_deployed_workers(registry, monkeypatch): + row = deployment() + calls = [] + + def fail_frontend(command, **kwargs): + calls.append(kwargs["env"].get("LILO_WORKER_ROLE", "frontend")) + if "lilo.providers.modal.app" in command: + raise subprocess.CalledProcessError(1, command) + + monkeypatch.setattr(subprocess, "run", fail_frontend) + with pytest.raises(subprocess.CalledProcessError): + cli.deploy([row]) + assert len(registry["worker_apps"]) == 2 + calls.clear() + monkeypatch.setattr( + subprocess, + "run", + lambda command, **kwargs: calls.append( + kwargs["env"].get("LILO_WORKER_ROLE", "frontend") + ), + ) + cli.deploy([row]) + assert calls == ["frontend"] + + +def test_recover_worker_deployed_before_registry_write(registry, monkeypatch): + row = deployment() + calls = [] + + def lookup(name, **kwargs): + if name == row.trainer_app_name: + return object() + raise modal.exception.NotFoundError("not deployed") + + monkeypatch.setattr(modal.App, "lookup", lookup) + monkeypatch.setattr( + subprocess, + "run", + lambda command, **kwargs: calls.append( + kwargs["env"].get("LILO_WORKER_ROLE", "frontend") + ), + ) + cli.deploy([row]) + assert calls == ["inference", "frontend"] diff --git a/tests/test_deployments.py b/tests/test_deployments.py index b45a948..db97945 100644 --- a/tests/test_deployments.py +++ b/tests/test_deployments.py @@ -33,9 +33,7 @@ def recipe(preset="qwen35-9b-lora-16k", **changes): def resolved(spec=None, **changes): - return DeploymentRecord.create( - spec or recipe(**changes), revision="a" * 40, implementation="test-runtime" - ) + return DeploymentRecord.create(spec or recipe(**changes), revision="a" * 40) def definition(value): @@ -108,12 +106,7 @@ def test_generation_and_asset_paths_include_exact_base(): assert a.generation != resolved(recipe(trainer__resources__gpu="H200:4")).generation b = resolved(recipe(model__id="other/Qwen3.5-9B-Base")) assert a.asset_path != b.asset_path - assert ( - a.asset_path - != DeploymentRecord.create( - a.spec, revision="b" * 40, implementation="test-runtime" - ).asset_path - ) + assert a.asset_path != DeploymentRecord.create(a.spec, revision="b" * 40).asset_path def test_frontend_defaults_and_retained_generations(): @@ -137,10 +130,6 @@ def test_frontend_defaults_and_retained_generations(): routes.select(small.definition_id, "lora").DEFINITION_ID == small.definition_id ) assert routes.capabilities()[0]["max_context_length"] == 65536 - with pytest.raises(ValueError, match="different Lilo/runtime"): - retain_generations( - [small.model_copy(update={"implementation": "old"})], [large] - ) with pytest.raises(ValueError, match="multiple defaults"): validate_frontend( [small.spec, recipe("qwen35-9b-lora-64k", routing__default=True)] @@ -357,16 +346,18 @@ def test_record_creation_copies_without_reparsing(): spec = recipe() original = asdict(spec) - row = DeploymentRecord.create(spec, revision="a" * 40, implementation="test") + row = DeploymentRecord.create(spec, revision="a" * 40) assert asdict(spec) == original assert row.spec.model["revision"] == "a" * 40 - # Keep the existing manifest fields and hash format stable. + # The record hash covers settings and the pinned backend dependency, not source. expected = original | {"model": original["model"] | {"revision": "a" * 40}} expected.pop("routing") assert ( row.generation == hashlib.sha256( - json.dumps(["test", expected], sort_keys=True).encode() + json.dumps( + {"config": expected, "miles_commit": None}, sort_keys=True + ).encode() ).hexdigest() ) row.spec.trainer["config"]["options"]["lora_rank"] = 64 @@ -411,7 +402,7 @@ def test_loading_python_config_does_not_call_backend_readers(monkeypatch): ) spec = load(config_path("qwen35-9b-lora-16k")) original_revision = spec.model["revision"] - record = DeploymentRecord.create(spec, revision="a" * 40, implementation="test") + record = DeploymentRecord.create(spec, revision="a" * 40) assert record.spec.model["revision"] == "a" * 40 assert spec.model["revision"] == original_revision @@ -507,3 +498,28 @@ class Config(BaseConfig): first.trainer["env"]["CUSTOM"] = "value" assert second.inference["scaling"]["max_replicas"] == 8 assert second.trainer["env"] == {} + + +def test_worker_hashes_cover_only_their_settings(): + base = resolved() + inference = resolved(recipe(inference__config__max_running_requests=24)) + assert inference.trainer_hash == base.trainer_hash + assert inference.inference_hash != base.inference_hash + + trainer = resolved(recipe(trainer__config__options__max_tokens_per_gpu=8192)) + assert trainer.trainer_hash != base.trainer_hash + assert trainer.inference_hash == base.inference_hash + + adapter = resolved(recipe(trainer__config__options__lora_rank=64)) + assert adapter.trainer_hash != base.trainer_hash + assert adapter.inference_hash != base.inference_hash + + routing = resolved(recipe(routing__default=False)) + assert routing.trainer_hash == base.trainer_hash + assert routing.inference_hash == base.inference_hash + assert routing.generation == base.generation + + upgraded = resolved(recipe(inference__runtime_version="2")) + assert upgraded.trainer_hash == base.trainer_hash + assert upgraded.inference_hash != base.inference_hash + assert "implementation" not in upgraded.model_dump() From b9a9f94a6ccbdae188297b3c93167a4a4ae5669c Mon Sep 17 00:00:00 2001 From: kailash Date: Wed, 23 Sep 2026 21:45:45 +0000 Subject: [PATCH 17/27] Keep platform settings and worker releases out of model configs --- docs/deployment-configs.md | 35 +++++---- docs/deployment-validation.md | 8 ++ scripts/deploy_models.sh | 2 +- src/lilo/configs/qwen35_4b_fft_64k.py | 2 - src/lilo/configs/qwen35_9b_lora_16k.py | 2 - src/lilo/deployment_cli.py | 71 ++++++++++++++++-- src/lilo/deployments.py | 71 +++++++++++------- src/lilo/providers/modal/app.py | 20 ++--- src/lilo/providers/modal/deployment_apps.py | 40 +++++----- src/lilo/providers/modal/fft_pool.py | 2 +- src/lilo/providers/modal/lora_pool.py | 2 +- tests/providers/test_checkpoint_storage.py | 9 ++- tests/providers/test_deployment_apps.py | 4 +- tests/test_deployment_cli.py | 83 ++++++++++++++++++++- tests/test_deployments.py | 30 ++++++-- 15 files changed, 281 insertions(+), 100 deletions(-) diff --git a/docs/deployment-configs.md b/docs/deployment-configs.md index d7b71f9..6aa3838 100644 --- a/docs/deployment-configs.md +++ b/docs/deployment-configs.md @@ -1,6 +1,6 @@ # Python deployment configs -Each deployment is a Python file exporting a `Config` class that inherits from `BaseConfig`. The class contains the model, trainer, inference and Modal settings. There is no YAML loader or model catalog. +Each deployment is a Python file exporting a `Config` class that inherits from `BaseConfig`. The class contains model, trainer, inference, routing and lifecycle settings. App names, environments and worker-code releases are managed by the deployment command. There is no YAML loader or model catalog. The layout follows the Python recipe approach used by [training-gym](https://github.com/modal-labs/training-gym) and the [multinode training guide](https://github.com/modal-labs/multinode-training-guide/blob/main/nemo-rl/configs/llama3_1_8b_math_2node.py). Only the outer BaseConfig is a dataclass; its sections are ordinary dictionaries. Config files need no decorators or default factories. Importing either project is not required. @@ -39,9 +39,9 @@ Copying happens inside the base class, so edits to a config's nested lists or di Python imports provide reuse. There is no YAML loader or `extends` key. Config files execute as Python when loaded; keep provisioning and training calls outside them. Sibling imports are available while loading a file. -All 14 previous presets are available under [`src/lilo/configs/`](../src/lilo/configs), with the same model, resource and backend settings. The [64K config](../src/lilo/configs/qwen35_9b_lora_64k.py) inherits from the 16K config. The 9B LoRA and 4B FFT configs directly pin the model revisions used in the earlier GPU checks; derived configs inherit them. There is no separate `deployments/` wrapper directory. +All 14 previous presets are available under [`src/lilo/configs/`](../src/lilo/configs), with the same model, resource and backend settings. The [64K config](../src/lilo/configs/qwen35_9b_lora_64k.py) inherits from the 16K config. There is no separate `deployments/` wrapper directory. -`model.revision` is the Hugging Face commit or branch containing the base weights and tokenizer. You can omit it to use `"main"`; the CLI resolves that branch to an exact commit before deployment. A fixed commit makes repeated deployments use the same files even if the repository's main branch changes. This is separate from training steps and published adapter versions. +No revision is required in a config. The CLI resolves the model's Hugging Face `main` branch automatically and saves the exact commit in the deployment record. An explicit `model.revision` remains optional for users who need particular weights; none of the built-in examples specify one. The example defaults therefore no longer pin the weights used in the historical GPU checks. ## Deploy the complete active set @@ -53,7 +53,7 @@ lilo config resolve src/lilo/configs/my_model.py --output /tmp/deployment.json lilo deploy src/lilo/configs/model_a.py src/lilo/configs/model_b.py ``` -`validate` loads the Python classes and checks shared frontend settings. It does not start backend libraries or prove that the model fits in GPU memory. `resolve` additionally pins model revisions and emits the deployment records as JSON; it does not provision compute. +`validate` loads the Python classes and checks routing and shared lifecycle settings. It does not start backend libraries or prove that the model fits in GPU memory. `resolve` additionally pins model revisions and emits the deployment records as JSON; it does not provision compute. For a checked-in list, edit [`scripts/deploy_models.sh`](../scripts/deploy_models.sh). Add a Python config file and its path to the `deployment_files` array, then run: @@ -61,7 +61,13 @@ For a checked-in list, edit [`scripts/deploy_models.sh`](../scripts/deploy_model ./scripts/deploy_models.sh ``` -Supply every configuration that should remain available to new clients. Omitted configurations are retained for existing jobs but removed from new-client selection. All files in the list must agree on shared frontend, region, secrets, storage and lifecycle settings. Pin `LILO_MILES_COMMIT` for repeatable deployments. Credentials remain in Modal secrets; the config contains only secret names. +Supply every configuration that should remain available to new clients. Omitted configurations are retained for existing jobs but removed from new-client selection. All files in the list must agree on shared lifecycle settings. Frontend deployment settings are not part of `BaseConfig`. The command selects the app, environment and region: + +```bash +./scripts/deploy_models.sh --app my-lilo --env dev --region us-west +``` + +These flags are optional; existing provider defaults apply when omitted. Secret and volume names come from provider defaults and are saved as platform metadata in deployment records. Credentials remain in Modal secrets. Pin `LILO_MILES_COMMIT` when a specific Miles build is needed. ## Code path @@ -150,27 +156,26 @@ There is no source fingerprint and no check that a config matches the current Li | Identifier | Inputs | Used for | | --- | --- | --- | -| `generation` | Complete computed config except routing defaults, plus pinned Miles commit | Saved job/checkpoint configuration and routing | -| `trainer_hash` | Deployment name, model, trainer section, shared deployment settings, Miles commit | Independently deployed trainer app name | -| `inference_hash` | Deployment name, model, inference section, shared deployment settings, adapter rank and target modules | Independently deployed inference provisioner app name | +| `generation` | Computed config except routing defaults, platform settings, saved worker releases and Miles commit | Saved job/checkpoint configuration and routing | +| `trainer_hash` | Deployment name, model, trainer section, platform settings, trainer release and Miles commit | Independently deployed trainer app name | +| `inference_hash` | Deployment name, model, inference section, platform settings, inference release, adapter rank and target modules | Independently deployed inference provisioner app name | | Asset hash | Model repository and resolved model commit | Download directory | The hashes use SHA-256 over sorted JSON. Trainer/inference app names use the first 24 hex characters. Definition IDs use the first 16 characters of `generation`. Code is retained by the deployed Modal apps, rather than reconstructed from a source fingerprint. -Both `trainer` and `inference` have a `runtime_version` setting, defaulting to `"1"`. To deploy a code-only trainer update, change `trainer.runtime_version` for the configurations that should use it: +Worker-code versions are not config fields. The CLI keeps the previous release IDs in deployment records. To deploy changed code for selected workers: -```python -overrides = { - "trainer.runtime_version": "2", -} +```bash +./scripts/deploy_models.sh --refresh-trainer qwen35-9b-lora-16k +./scripts/deploy_models.sh --refresh-inference qwen35-9b-lora-16k ``` -An inference code update uses `inference.runtime_version` instead. These are operator-selected release labels, not source hashes or validation requirements. Editing source alone does not update an existing worker app. If shared worker code changes, bump each affected role's version. The frontend itself is redeployed on every apply. +The CLI generates a release ID for each requested update; users do not specify or maintain it. Subsequent ordinary deploys retain that release. Refresh flags can be repeated for multiple config names. Editing source alone does not update existing workers. The frontend itself is redeployed on every apply. Examples: - Change inference concurrency: deploy a new inference provisioner; keep the existing trainer app. -- Change trainer batch settings or trainer runtime version: deploy a new trainer app; keep the inference provisioner. +- Change trainer batch settings or request a trainer code refresh: deploy a new trainer app; keep the inference provisioner. - Change adapter rank or target modules: update both because inference must load the changed adapters. - Change routing defaults: keep both worker apps. - Change one Miles configuration: other Miles configurations and Megatron apps remain deployed as they were. diff --git a/docs/deployment-validation.md b/docs/deployment-validation.md index d6e6cb3..73f0a5a 100644 --- a/docs/deployment-validation.md +++ b/docs/deployment-validation.md @@ -2,6 +2,14 @@ These checks exercise PR #55. The current implementation deploys trainers and inference provisioners independently; the shared frontend references them by app name. Earlier sections record validation of the previous shared-app implementation. +## Model configs contain infrastructure settings only + +Removed the `deployment` section from `BaseConfig` and removed explicit model commits from the built-in examples. Frontend app/environment/region selection belongs to `lilo deploy`; platform metadata and automatically generated worker-release IDs live in saved records. Config authors do not need revision or runtime-version fields. The CLI resolves omitted model revisions from Hugging Face `main`. + +Code-only updates use `--refresh-trainer CONFIG_NAME` or `--refresh-inference CONFIG_NAME`. Subsequent ordinary deploys retain those releases. The deployment script forwards these options. + +Validation: **619 CPU tests passed, 1 skipped**. Added tests cover all 14 examples having no deployment/revision/runtime-version fields, automatic model-commit lookup, command-owned platform settings, and preservation of refreshed workers without changing configs. Existing update-isolation and worker-app construction tests still pass. Ruff, whitespace and script syntax checks passed. No apps were redeployed. + ## Independent worker apps and config hashes Removed the global source fingerprint and its code-upgrade rejection. The config hash identifies saved job settings; separate trainer and inference hashes identify worker apps. Runtime upgrades use explicit per-role `runtime_version` labels. The CLI skips existing worker apps, deploys changed ones before updating the frontend, and records successful workers for retry. Inference provisioners retain their source in an image so idle pools can restart using their original code. diff --git a/scripts/deploy_models.sh b/scripts/deploy_models.sh index 2b74006..2e59ffc 100755 --- a/scripts/deploy_models.sh +++ b/scripts/deploy_models.sh @@ -12,4 +12,4 @@ deployment_files=( src/lilo/configs/qwen35_4b_fft_64k.py ) -lilo deploy "${deployment_files[@]}" +lilo deploy "${deployment_files[@]}" "$@" diff --git a/src/lilo/configs/qwen35_4b_fft_64k.py b/src/lilo/configs/qwen35_4b_fft_64k.py index c0df8ac..f62e54d 100644 --- a/src/lilo/configs/qwen35_4b_fft_64k.py +++ b/src/lilo/configs/qwen35_4b_fft_64k.py @@ -5,12 +5,10 @@ class Config(BaseConfig): name = "qwen35-4b-fft-64k" model = { "id": "Qwen/Qwen3.5-4B", - "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", "parameterization": "full", "max_context_length": 65536, } routing = {"default": True} - deployment = {"frontend": "lilo-yaml"} trainer = { "backend": "megatron", "resources": {"gpu": "H100:4"}, diff --git a/src/lilo/configs/qwen35_9b_lora_16k.py b/src/lilo/configs/qwen35_9b_lora_16k.py index 1c17dc6..6e4a4e4 100644 --- a/src/lilo/configs/qwen35_9b_lora_16k.py +++ b/src/lilo/configs/qwen35_9b_lora_16k.py @@ -5,12 +5,10 @@ class Config(BaseConfig): name = "qwen35-9b-lora-16k" model = { "id": "Qwen/Qwen3.5-9B-Base", - "revision": "68c46c4b3498877f3ef123c856ecfde50c39f404", "parameterization": "lora", "max_context_length": 16384, } routing = {"default": True} - deployment = {"frontend": "lilo-yaml", "mode": "shared"} trainer = { "backend": "miles", "resources": {"gpu": "H100:4", "cpu": 16, "memory_mib": 65536}, diff --git a/src/lilo/deployment_cli.py b/src/lilo/deployment_cli.py index 673d3c8..fa6298d 100644 --- a/src/lilo/deployment_cli.py +++ b/src/lilo/deployment_cli.py @@ -14,13 +14,14 @@ from lilo.deployments import ( DeploymentRecord, + PLATFORM_DEFAULTS, load, config_path, validate_frontend, ) -def compile_configs(paths): +def compile_configs(paths, *, platform=None): specs = [load(path) for path in paths] validate_frontend(specs) from lilo.providers.modal.miles_revision import resolve_miles_commit @@ -32,7 +33,7 @@ def compile_configs(paths): ) records = [] for spec in specs: - revision = spec.model["revision"] + revision = spec.model.get("revision", "main") if not re.fullmatch(r"[0-9a-fA-F]{40,64}", revision): from huggingface_hub import HfApi @@ -44,6 +45,7 @@ def compile_configs(paths): records.append( DeploymentRecord.create( spec, + platform=platform, revision=revision, miles_commit=miles_commit if spec.trainer["backend"] == "miles" @@ -59,7 +61,7 @@ def retain_generations(previous, desired): expected = desired[0] for row in previous: if ( - row.spec.deployment != expected.spec.deployment + row.platform != expected.platform or row.spec.lifecycle != expected.spec.lifecycle ): raise ValueError( @@ -76,11 +78,44 @@ def retain_generations(previous, desired): ] +def select_worker_releases(desired, previous, refresh_trainers, refresh_inference): + """Keep deployed code unless the operator explicitly requests a worker update.""" + names = {row.spec.name for row in desired} + unknown = (set(refresh_trainers) | set(refresh_inference)) - names + if unknown: + raise ValueError(f"unknown configs to refresh: {sorted(unknown)}") + active = {row.spec.name: row for row in previous if row.active} + result = [] + for row in desired: + old = active.get(row.spec.name, row) + trainer = ( + uuid.uuid4().hex + if row.spec.name in refresh_trainers + else old.trainer_release + ) + inference = ( + uuid.uuid4().hex + if row.spec.name in refresh_inference + else old.inference_release + ) + result.append( + DeploymentRecord.create( + row.spec, + revision=row.spec.model["revision"], + miles_commit=row.miles_commit, + platform=row.platform, + trainer_release=trainer, + inference_release=inference, + ) + ) + return result + + def worker_apps_ready(row, deployed): return row.trainer_app_name in deployed and row.inference_app_name in deployed -def deploy(desired): +def deploy(desired, *, refresh_trainers=(), refresh_inference=()): """Serialize operator applies and retain interrupted attempts for safe recovery.""" import modal from lilo.providers.modal.deployment_apps import MANIFEST_ENV @@ -89,7 +124,7 @@ def deploy(desired): raise ValueError( "Python deployment requires Python 3.12 to match the serialized GPU runtime images" ) - settings = desired[0].spec.deployment + settings = desired[0].platform registry = modal.Dict.from_name( f"{settings['frontend']}-yaml-deployments", create_if_missing=True, @@ -122,9 +157,11 @@ def deploy(desired): if worker_apps_ready(DeploymentRecord.model_validate(row), deployed) ] rows = {r["generation"]: r for r in [*rows, *pending]} - manifest = retain_generations( - [DeploymentRecord.model_validate(row) for row in rows.values()], desired + previous = [DeploymentRecord.model_validate(row) for row in rows.values()] + desired = select_worker_releases( + desired, previous, refresh_trainers, refresh_inference ) + manifest = retain_generations(previous, desired) data = [row.model_dump(mode="json") for row in manifest] env = { **os.environ, @@ -210,6 +247,15 @@ def parser(): help="Deploy the complete active Python config set behind one frontend", ) apply.add_argument("files", nargs="+") + apply.add_argument("--app", default=PLATFORM_DEFAULTS["frontend"]) + apply.add_argument("--env") + apply.add_argument("--region", default=PLATFORM_DEFAULTS["modal"]["region"]) + apply.add_argument( + "--refresh-trainer", action="append", default=[], metavar="CONFIG_NAME" + ) + apply.add_argument( + "--refresh-inference", action="append", default=[], metavar="CONFIG_NAME" + ) management = commands.add_parser("deployment").add_subparsers( dest="action", required=True ) @@ -258,7 +304,16 @@ def main(argv=None): else: print(output, end="") elif args.command == "deploy": - deploy(compile_configs(args.files)) + from copy import deepcopy + + platform = deepcopy(PLATFORM_DEFAULTS) + platform["frontend"] = args.app + platform["modal"].update(environment=args.env, region=args.region) + deploy( + compile_configs(args.files, platform=platform), + refresh_trainers=args.refresh_trainer, + refresh_inference=args.refresh_inference, + ) else: import modal diff --git a/src/lilo/deployments.py b/src/lilo/deployments.py index e3651e5..34fdb5a 100644 --- a/src/lilo/deployments.py +++ b/src/lilo/deployments.py @@ -13,17 +13,31 @@ import sys from typing import Any, ClassVar -from pydantic import BaseModel, ConfigDict +from pydantic import BaseModel, ConfigDict, Field +PLATFORM_DEFAULTS = { + "frontend": "lilo-yaml", + "modal": {"environment": None, "region": "us-west"}, + "secrets": { + "api": "lilo-api", + "sampler_proxy": "lilo-proxy", + "huggingface": "huggingface-secret", + }, + "storage": { + "assets": "lilo-model-assets", + "checkpoints": "lilo-checkpoints", + "bulletin": "lilo-snapshot-bulletin", + }, +} + # Shared orchestration defaults. Backend option dictionaries have no schema here. _DEFAULTS = { "api_version": "lilo/v1", - "model": {"revision": "main", "parameterization": "lora"}, + "model": {"parameterization": "lora"}, "routing": {"default": False, "sampling_default": False}, "trainer": { "backend": "miles", - "runtime_version": "1", "resources": {"cpu": 8, "memory_mib": 32768, "timeout_s": 86400}, "scaling": {"min_instances": 0, "max_instances": 1}, "engine": {"max_clients_per_instance": 1, "sampler_persistence_concurrency": 8}, @@ -32,7 +46,6 @@ }, "inference": { "backend": "sglang", - "runtime_version": "1", "resources": {"cpu": 8, "memory_mib": 32768, "timeout_s": 86400}, "scaling": { "min_replicas": 0, @@ -43,21 +56,6 @@ "config": {}, "env": {}, }, - "deployment": { - "frontend": "lilo-yaml", - "mode": "shared", - "modal": {"environment": None, "region": "us-west"}, - "secrets": { - "api": "lilo-api", - "sampler_proxy": "lilo-proxy", - "huggingface": "huggingface-secret", - }, - "storage": { - "assets": "lilo-model-assets", - "checkpoints": "lilo-checkpoints", - "bulletin": "lilo-snapshot-bulletin", - }, - }, "lifecycle": { "session_idle_timeout_s": 300, "pool_idle_timeout_s": 300, @@ -86,7 +84,6 @@ class BaseConfig: inference: dict[str, Any] api_version: str routing: dict[str, Any] - deployment: dict[str, Any] lifecycle: dict[str, Any] overrides: ClassVar[dict[str, Any]] = {} @@ -141,6 +138,11 @@ class DeploymentRecord(BaseModel): model_config = ConfigDict(extra="forbid") spec: BaseConfig + platform: dict[str, Any] = Field( + default_factory=lambda: deepcopy(PLATFORM_DEFAULTS) + ) + trainer_release: str = "initial" + inference_release: str = "initial" miles_commit: str | None = None generation: str active: bool = True @@ -152,6 +154,9 @@ def create( *, revision: str, miles_commit: str | None = None, + platform: dict | None = None, + trainer_release: str = "initial", + inference_release: str = "initial", ) -> DeploymentRecord: """Record an already-resolved revision without reparsing the configuration.""" pinned = deepcopy(spec) @@ -159,22 +164,35 @@ def create( # Changing routing defaults should not restart an existing trainer. identity = asdict(pinned) identity.pop("routing") - generation = settings_hash({"config": identity, "miles_commit": miles_commit}) + platform = deepcopy(PLATFORM_DEFAULTS if platform is None else platform) + generation = settings_hash( + { + "config": identity, + "platform": platform, + "miles_commit": miles_commit, + "trainer_release": trainer_release, + "inference_release": inference_release, + } + ) return cls( spec=pinned, + platform=platform, + trainer_release=trainer_release, + inference_release=inference_release, generation=generation, miles_commit=miles_commit, ) @property def trainer_hash(self) -> str: - """Identify trainer settings and the operator-selected runtime version.""" + """Identify trainer settings and the deployment-managed code release.""" return settings_hash( { "name": self.spec.name, "model": self.spec.model, "trainer": self.spec.trainer, - "deployment": self.spec.deployment, + "release": self.trainer_release, + "platform": self.platform, "miles_commit": self.miles_commit, } ) @@ -194,7 +212,8 @@ def inference_hash(self) -> str: "name": self.spec.name, "model": self.spec.model, "inference": self.spec.inference, - "deployment": self.spec.deployment, + "release": self.inference_release, + "platform": self.platform, "adapter": adapter, } ) @@ -253,9 +272,9 @@ def validate_frontend(specs: list[BaseConfig]) -> None: raise ValueError("duplicate deployment name") first = specs[0] for spec in specs: - if spec.deployment != first.deployment or spec.lifecycle != first.lifecycle: + if spec.lifecycle != first.lifecycle: raise ValueError( - "deployments on one frontend must share deployment and lifecycle settings" + "deployments on one frontend must share lifecycle settings" ) defaults, sampling = set(), set() for spec in specs: diff --git a/src/lilo/providers/modal/app.py b/src/lilo/providers/modal/app.py index 4942cd4..0fe2e03 100644 --- a/src/lilo/providers/modal/app.py +++ b/src/lilo/providers/modal/app.py @@ -56,21 +56,21 @@ ) SETTINGS = frontend_settings() -APP_NAME = SETTINGS.deployment["frontend"] -ROUTING_REGION = SETTINGS.deployment["modal"]["region"] +APP_NAME = SETTINGS.platform["frontend"] +ROUTING_REGION = SETTINGS.platform["modal"]["region"] MODEL_ASSET_ROOT = "/assets" -SESSION_IDLE_TIMEOUT = SETTINGS.lifecycle["session_idle_timeout_s"] -FFT_POOL_IDLE_TIMEOUT = LORA_POOL_IDLE_TIMEOUT = SETTINGS.lifecycle[ +SESSION_IDLE_TIMEOUT = SETTINGS.spec.lifecycle["session_idle_timeout_s"] +FFT_POOL_IDLE_TIMEOUT = LORA_POOL_IDLE_TIMEOUT = SETTINGS.spec.lifecycle[ "pool_idle_timeout_s" ] FFT_POOL_TOUCH_INTERVAL = 60.0 LORA_POOL_CHECK_INTERVAL = 60.0 -SWEEP_PERIOD = modal.Period(seconds=SETTINGS.lifecycle["sweep_interval_s"]) +SWEEP_PERIOD = modal.Period(seconds=SETTINGS.spec.lifecycle["sweep_interval_s"]) CHECKPOINT_READ_LOCK = asyncio.Lock() _pool_touches: dict[str, float] = {} _lora_pool_gateways: dict[str, tuple[float, str]] = {} _lora_pool_checks: dict[str, asyncio.Lock] = {} -CHECKPOINT_VOLUME_NAME = SETTINGS.deployment["storage"]["checkpoints"] +CHECKPOINT_VOLUME_NAME = SETTINGS.platform["storage"]["checkpoints"] checkpoint_volume = modal.Volume.from_name( CHECKPOINT_VOLUME_NAME, create_if_missing=True, version=2 ) @@ -110,13 +110,13 @@ async def _delete_checkpoint(uri: str) -> None: .add_local_python_source("lilo", ignore=ignore_config_source) ) model_assets = modal.Volume.from_name( - SETTINGS.deployment["storage"]["assets"], + SETTINGS.platform["storage"]["assets"], create_if_missing=True, ) -API_SECRET_NAME = SETTINGS.deployment["secrets"]["api"] -HF_SECRET_NAME = SETTINGS.deployment["secrets"]["huggingface"] +API_SECRET_NAME = SETTINGS.platform["secrets"]["api"] +HF_SECRET_NAME = SETTINGS.platform["secrets"]["huggingface"] proxy_secret = modal.Secret.from_name( - SETTINGS.deployment["secrets"]["sampler_proxy"], + SETTINGS.platform["secrets"]["sampler_proxy"], required_keys=["MODAL_PROXY_TOKEN_ID", "MODAL_PROXY_TOKEN_SECRET"], ) diff --git a/src/lilo/providers/modal/deployment_apps.py b/src/lilo/providers/modal/deployment_apps.py index ceb8c6a..0ca5428 100644 --- a/src/lilo/providers/modal/deployment_apps.py +++ b/src/lilo/providers/modal/deployment_apps.py @@ -37,7 +37,7 @@ def frontend_settings(): validate_frontend(active) if len({row.definition_id for row in deployments}) != len(deployments): raise ValueError("duplicate deployment generation") - return active[0] + return deployments[0] def image_for(backend): @@ -52,8 +52,8 @@ def image_for(backend): return image -def volumes_for(spec): - storage = spec.deployment["storage"] +def volumes_for(record): + storage = record.platform["storage"] return { "/assets": modal.Volume.from_name(storage["assets"], create_if_missing=True), "/checkpoints": modal.Volume.from_name( @@ -65,8 +65,8 @@ def volumes_for(spec): } -def secrets_for(spec, *, training=False): - names = spec.deployment["secrets"] +def secrets_for(record, *, training=False): + names = record.platform["secrets"] result = [modal.Secret.from_name(names["api"], required_keys=["TINKER_API_KEY"])] if training: result.append( @@ -97,7 +97,7 @@ def build_trainer_app(resolved: DeploymentRecord, *, image=None): env = { **trainer_deployment_env(), **deployment_env(spec.trainer["env"]), - "LILO_APP_NAME": spec.deployment["frontend"], + "LILO_APP_NAME": resolved.platform["frontend"], } @app.function( @@ -105,7 +105,7 @@ def build_trainer_app(resolved: DeploymentRecord, *, image=None): serialized=True, image=image if image is not None else image_for(spec.trainer["backend"]), gpu=resource["gpu"], - region=spec.deployment["modal"]["region"], + region=resolved.platform["modal"]["region"], cpu=resource["cpu"], memory=resource["memory_mib"], timeout=resource["timeout_s"], @@ -114,8 +114,8 @@ def build_trainer_app(resolved: DeploymentRecord, *, image=None): max_containers=None, min_containers=0, single_use_containers=True, - volumes=volumes_for(spec), - secrets=secrets_for(spec, training=True), + volumes=volumes_for(resolved), + secrets=secrets_for(resolved, training=True), env=env, ) def trainer(instance_id: str, config_json: str): @@ -136,17 +136,17 @@ def run_trainer(resolved, instance_id): settings = backend_config(spec, resolved.asset_path) # Assets are prepared by the frontend before demand is registered. Reload once # on startup to see the committed exact snapshot; never race a trainer download. - volumes_for(spec)["/assets"].reload() + volumes_for(resolved)["/assets"].reload() env = { **deployment_env(spec.trainer["env"]), - "LILO_APP_NAME": spec.deployment["frontend"], + "LILO_APP_NAME": resolved.platform["frontend"], "LILO_BACKEND_CONFIG": json.dumps(settings), "LILO_BASE_MODEL": spec.model["id"], "LILO_BASE_MODEL_REVISION": spec.model["revision"], "LILO_DEFINITION_ID": resolved.definition_id, - "LILO_CHECKPOINT_VOLUME": spec.deployment["storage"]["checkpoints"], + "LILO_CHECKPOINT_VOLUME": resolved.platform["storage"]["checkpoints"], "LILO_BULLETIN_ROOT": "/bulletin", - "LILO_BULLETIN_VOLUME": spec.deployment["storage"]["bulletin"], + "LILO_BULLETIN_VOLUME": resolved.platform["storage"]["bulletin"], "LILO_DEFINITION_REVISION": resolved.generation, } executor = ( @@ -206,7 +206,7 @@ def definition_from_spec(resolved, *, register_trainer=True, image=None): definition.ENGINE_FUNCTION = modal.Function.from_name( resolved.trainer_app_name, "trainer", - environment_name=spec.deployment["modal"]["environment"], + environment_name=resolved.platform["modal"]["environment"], ) return definition @@ -264,8 +264,8 @@ def build_rollout_app(resolved, pool, *, image=None): gpu=resources["gpu"], cpu=resources["cpu"], memory=resources["memory_mib"], - volumes=volumes_for(spec), - secrets=secrets_for(spec), + volumes=volumes_for(resolved), + secrets=secrets_for(resolved), env=deployment_env(spec.inference["env"]), min_containers=scaling["min_replicas"] if minimum is None else minimum, max_containers=scaling["max_replicas"] if maximum is None else maximum, @@ -274,8 +274,8 @@ def build_rollout_app(resolved, pool, *, image=None): startup_timeout=1200, exit_grace_period=300, port=8000, - routing_region=spec.deployment["modal"]["region"], - compute_region=spec.deployment["modal"]["region"], + routing_region=resolved.platform["modal"]["region"], + compute_region=resolved.platform["modal"]["region"], ) class Server: @modal.enter() @@ -306,7 +306,7 @@ def start(self): port=8000, sglang_port=8001, bulletin_root="/bulletin", - bulletin_volume=spec.deployment["storage"]["bulletin"], + bulletin_volume=resolved.platform["storage"]["bulletin"], ) self.sidecar = ( start_lora_sidecar(**kwargs) @@ -340,7 +340,7 @@ def provision_pool(record, pool): provision = modal.Function.from_name( record.inference_app_name, "provision", - environment_name=record.spec.deployment["modal"]["environment"], + environment_name=record.platform["modal"]["environment"], ) return provision.remote(record.model_dump_json(), pool.as_dict()) diff --git a/src/lilo/providers/modal/fft_pool.py b/src/lilo/providers/modal/fft_pool.py index 8c57c86..fda9cf2 100644 --- a/src/lilo/providers/modal/fft_pool.py +++ b/src/lilo/providers/modal/fft_pool.py @@ -150,7 +150,7 @@ def deploy_pool(spec: FFTPoolSpec, *, record=None) -> str: "--name", spec.app_name, ] - environment = record.spec.deployment["modal"]["environment"] + environment = record.platform["modal"]["environment"] if environment: command.extend(["--env", environment]) subprocess.run(command, env=env, check=True) diff --git a/src/lilo/providers/modal/lora_pool.py b/src/lilo/providers/modal/lora_pool.py index 42635f0..5d9f592 100644 --- a/src/lilo/providers/modal/lora_pool.py +++ b/src/lilo/providers/modal/lora_pool.py @@ -83,7 +83,7 @@ def deploy_pool(spec: LoraPoolSpec, *, record=None) -> str: "--name", spec.app_name, ] - environment = record.spec.deployment["modal"]["environment"] + environment = record.platform["modal"]["environment"] if environment: command.extend(["--env", environment]) subprocess.run(command, env={**os.environ, **spec.env(), **recipe_env}, check=True) diff --git a/tests/providers/test_checkpoint_storage.py b/tests/providers/test_checkpoint_storage.py index c02778c..77dfb72 100644 --- a/tests/providers/test_checkpoint_storage.py +++ b/tests/providers/test_checkpoint_storage.py @@ -42,12 +42,15 @@ def test_yaml_definitions_use_configured_checkpoint_storage(): with patch.object( modal.Volume, "from_name", side_effect=lambda name, **kwargs: (name, kwargs) ): - volumes = volumes_for(spec) + volumes = volumes_for(definition.RESOLVED) assert volumes[CHECKPOINT_ROOT] == ( - spec.deployment['storage']['checkpoints'], + definition.RESOLVED.platform["storage"]["checkpoints"], {"create_if_missing": True, "version": 2}, ) - assert volumes["/bulletin"][0] == spec.deployment['storage']['bulletin'] + assert ( + volumes["/bulletin"][0] + == definition.RESOLVED.platform["storage"]["bulletin"] + ) assert backend_config(spec)["checkpoint_dir"] == CHECKPOINT_ROOT diff --git a/tests/providers/test_deployment_apps.py b/tests/providers/test_deployment_apps.py index 4acd304..4fbfefc 100644 --- a/tests/providers/test_deployment_apps.py +++ b/tests/providers/test_deployment_apps.py @@ -54,7 +54,7 @@ def test_trainer_declaration_and_executor_configuration( builders, monkeypatch, preset, backend, clients, nproc ): row = deployment(preset) - row.spec.deployment["storage"]["checkpoints"] = "test-custom-checkpoints" + row.platform["storage"]["checkpoints"] = "test-custom-checkpoints" image = object() app, trainer = deployment_apps.build_trainer_app(row, image=image) declaration, _ = app.functions["trainer"] @@ -359,7 +359,7 @@ def test_provisioner_rejects_wrong_settings_and_uses_saved_record( assert provision(row.model_dump_json(), pool.as_dict()) == "https://pool" assert calls[0][1].inference_hash == row.inference_hash changed = row.model_copy(deep=True) - changed.spec.inference["runtime_version"] = "2" + changed.inference_release = "2" with pytest.raises(ValueError, match="inference settings"): provision(changed.model_dump_json(), pool.as_dict()) diff --git a/tests/test_deployment_cli.py b/tests/test_deployment_cli.py index 7fd8d58..43d4d26 100644 --- a/tests/test_deployment_cli.py +++ b/tests/test_deployment_cli.py @@ -177,10 +177,9 @@ def run(command, **kwargs): assert new.trainer_app_name == row.trainer_app_name assert registry["manifest"][1]["active"] is False - changed.trainer["runtime_version"] = "new-trainer-code" newest = DeploymentRecord.create(changed, revision="a" * 40) calls.clear() - cli.deploy([newest]) + cli.deploy([newest], refresh_trainers=[newest.spec.name]) assert calls == ["trainer", "frontend"] assert newest.inference_app_name == new.inference_app_name @@ -206,8 +205,10 @@ def run(command, **kwargs): cli.deploy([miles, fft]) calls.clear() spec = deepcopy(miles.spec) - spec.trainer["runtime_version"] = "2" - cli.deploy([DeploymentRecord.create(spec, revision="a" * 40), fft]) + cli.deploy( + [DeploymentRecord.create(spec, revision="a" * 40), fft], + refresh_trainers=[miles.spec.name], + ) assert calls == [("trainer", miles.spec.name)] @@ -255,3 +256,77 @@ def lookup(name, **kwargs): ) cli.deploy([row]) assert calls == ["inference", "frontend"] + + +def test_worker_refresh_is_retained_without_editing_config(registry, monkeypatch): + row = deployment() + calls = [] + monkeypatch.setattr( + subprocess, + "run", + lambda command, **kwargs: calls.append( + kwargs["env"].get("LILO_WORKER_ROLE", "frontend") + ), + ) + cli.deploy([row]) + cli.deploy([row], refresh_inference=[row.spec.name]) + active = DeploymentRecord.model_validate(registry["manifest"][0]) + assert active.inference_release != "initial" + assert active.trainer_release == "initial" + assert active.spec.inference == row.spec.inference + calls.clear() + cli.deploy([row]) + assert calls == ["frontend"] + assert registry["manifest"][0]["inference_release"] == active.inference_release + + +def test_deploy_command_owns_platform_settings(monkeypatch): + seen = {} + + def compile(paths, *, platform): + seen["platform"] = platform + return ["record"] + + def deploy(rows, **kwargs): + seen["rows"] = rows + seen.update(kwargs) + + monkeypatch.setattr(cli, "compile_configs", compile) + monkeypatch.setattr(cli, "deploy", deploy) + cli.main( + [ + "deploy", + "model.py", + "--app", + "my-lilo", + "--env", + "dev", + "--region", + "us-east", + "--refresh-trainer", + "my-model", + ] + ) + assert seen["platform"]["frontend"] == "my-lilo" + assert seen["platform"]["modal"] == {"environment": "dev", "region": "us-east"} + assert seen["refresh_trainers"] == ["my-model"] + assert seen["rows"] == ["record"] + + +def test_builtin_config_resolves_revision_automatically(monkeypatch): + from types import SimpleNamespace + import huggingface_hub + from lilo.providers.modal import miles_revision + + calls = [] + monkeypatch.setattr(miles_revision, "resolve_miles_commit", lambda: "b" * 40) + monkeypatch.setattr( + huggingface_hub.HfApi, + "model_info", + lambda self, model, *, revision: calls.append((model, revision)) + or SimpleNamespace(sha="a" * 40), + ) + (row,) = cli.compile_configs([config_path("qwen35-9b-lora-16k")]) + assert calls == [("Qwen/Qwen3.5-9B-Base", "main")] + assert row.spec.model["revision"] == "a" * 40 + assert "revision" not in load(config_path("qwen35-9b-lora-16k")).model diff --git a/tests/test_deployments.py b/tests/test_deployments.py index db97945..5e24283 100644 --- a/tests/test_deployments.py +++ b/tests/test_deployments.py @@ -356,7 +356,14 @@ def test_record_creation_copies_without_reparsing(): row.generation == hashlib.sha256( json.dumps( - {"config": expected, "miles_commit": None}, sort_keys=True + { + "config": expected, + "miles_commit": None, + "platform": row.platform, + "trainer_release": "initial", + "inference_release": "initial", + }, + sort_keys=True, ).encode() ).hexdigest() ) @@ -401,10 +408,10 @@ def test_loading_python_config_does_not_call_backend_readers(monkeypatch): backends, "serving_options", lambda *a: pytest.fail("serving read") ) spec = load(config_path("qwen35-9b-lora-16k")) - original_revision = spec.model["revision"] + assert "revision" not in spec.model record = DeploymentRecord.create(spec, revision="a" * 40) assert record.spec.model["revision"] == "a" * 40 - assert spec.model["revision"] == original_revision + assert "revision" not in spec.model @pytest.mark.parametrize("source", ["Config = {}", "class Config: pass", "value = 1"]) @@ -491,7 +498,7 @@ class Config(BaseConfig): first, second = Config(), Config() assert type(first.model) is dict assert type(first.trainer) is dict - assert first.model["revision"] == "main" + assert "revision" not in first.model assert first.trainer["resources"]["cpu"] == 8 assert first.trainer["config"] == {"future_option": False} first.inference["scaling"]["max_replicas"] = 2 @@ -519,7 +526,20 @@ def test_worker_hashes_cover_only_their_settings(): assert routing.inference_hash == base.inference_hash assert routing.generation == base.generation - upgraded = resolved(recipe(inference__runtime_version="2")) + upgraded = DeploymentRecord.create( + base.spec, revision="a" * 40, inference_release="2" + ) assert upgraded.trainer_hash == base.trainer_hash assert upgraded.inference_hash != base.inference_hash assert "implementation" not in upgraded.model_dump() + + +def test_examples_only_contain_model_infrastructure(): + from pathlib import Path + + for path in Path(config_path("qwen35-9b-lora-16k")).parent.glob("qwen*.py"): + config = load(path) + assert not hasattr(config, "deployment") + assert "revision" not in config.model + assert "runtime_version" not in config.trainer + assert "runtime_version" not in config.inference From 3faaab5f6d8ace028ca41970b71e2649505bb648 Mon Sep 17 00:00:00 2001 From: kailash Date: Wed, 23 Sep 2026 22:08:47 +0000 Subject: [PATCH 18/27] Use backend config fields directly and simplify argparse overrides --- docs/deployment-configs.md | 56 +++++--- docs/deployment-validation.md | 8 ++ src/lilo/argparse_config.py | 51 ++++++++ src/lilo/backend_options.py | 14 -- src/lilo/backends/megatron_deployment.py | 68 ++++------ src/lilo/backends/miles_config.py | 16 ++- src/lilo/backends/miles_deployment.py | 57 +++------ src/lilo/backends/miles_runtime/runtime.py | 6 +- src/lilo/config_validation.py | 9 ++ src/lilo/configs/qwen35_35b_a3b_fft_64k.py | 28 ++-- src/lilo/configs/qwen35_4b_fft_64k.py | 20 ++- src/lilo/configs/qwen35_9b_fft_64k.py | 20 ++- .../configs/qwen35_9b_instruct_lora_16k.py | 21 ++- .../qwen35_9b_instruct_lora_16k_dp2.py | 2 +- src/lilo/configs/qwen35_9b_lora_16k.py | 28 ++-- src/lilo/configs/qwen35_9b_lora_2k.py | 28 ++-- src/lilo/configs/qwen35_9b_lora_64k.py | 4 +- src/lilo/configs/qwen36_27b_fft_64k.py | 24 ++-- src/lilo/configs/qwen38_27b_lora_128k.py | 6 +- src/lilo/configs/qwen38_27b_lora_16k.py | 21 ++- src/lilo/configs/qwen38_27b_lora_64k.py | 4 +- src/lilo/deployments.py | 4 +- src/lilo/inference/native_sglang.py | 4 +- src/lilo/inference/sglang_deployment.py | 6 +- src/lilo/native_options.py | 92 -------------- tests/backends/test_native_megatron_config.py | 10 +- tests/test_deployments.py | 120 ++++++++++-------- 27 files changed, 334 insertions(+), 393 deletions(-) create mode 100644 src/lilo/argparse_config.py delete mode 100644 src/lilo/backend_options.py create mode 100644 src/lilo/config_validation.py delete mode 100644 src/lilo/native_options.py diff --git a/docs/deployment-configs.md b/docs/deployment-configs.md index 6aa3838..01f4e88 100644 --- a/docs/deployment-configs.md +++ b/docs/deployment-configs.md @@ -23,8 +23,8 @@ class Config(ParentConfig): overrides = { "model.max_context_length": 65536, "trainer.resources.gpu": "H200:8", - "trainer.config.options.tensor_model_parallel_size": 8, - "trainer.config.options.max_tokens_per_gpu": 65536, + "trainer.config.tensor_model_parallel_size": 8, + "trainer.config.max_tokens_per_gpu": 65536, "inference.scaling.max_replicas": 6, } ``` @@ -103,37 +103,55 @@ lilo deploy config.py Inference pools are separate apps created on demand by a deployed `provision` function. That function keeps the source used when its inference app was deployed, including when an idle pool needs to be recreated after a frontend upgrade. `build_rollout_app()` constructs a server from the saved configuration. Startup launches SGLang with native options, waits for its health endpoint, then starts the LoRA or FFT sidecar. -## Backend options +## Backend config path -`trainer.config` and `inference.config` remain open dictionaries. Adding an upstream option does not require adding a deployment dataclass field. Backend readers check values used by Lilo's integration, while installed backend libraries check native options at worker startup. +Each backend already has a config class used by its trainer. `trainer.config` uses that class's field names directly. The deployment code supplies the downloaded model path, GPU count and context length; it does not rename user fields. -For Miles, `trainer.config` contains: +| Backend | Config consumer | Additional backend options | +| --- | --- | --- | +| Miles | `MilesBackendConfig` in `backends/miles_config.py` | `cli_options` goes to Miles's argparse parser | +| Megatron | `EngineModelConfig` in `backends/megatron_runtime/common/config.py` | `provider_overrides`, `optimizer_overrides`, `distributed_overrides` go to their respective Megatron constructors | +| SGLang | Its own `ServerArgs` parser | The entire `inference.config` dictionary | + +For Miles: ```python -{ - "model_args": "qwen3.5-9B", - "options": { - "tensor_model_parallel_size": 4, - "multi_lora_n_adapters": 6, +"config": { + "model_type": "qwen3.5-9B", + "tensor_model_parallel_size": 4, + "max_lora_slots": 6, + "max_lora_rank": 32, + "cli_options": { "recompute_granularity": "full", + "recompute_method": "uniform", + "recompute_num_layers": 1, }, } ``` -`model_args` selects a Miles architecture preset; it can be omitted when explicit architecture options are provided. Miles's actual argument parser validates native options. Lilo also reads parallelism, adapter slots/rank and targets for admission and adapter export. +The path is `trainer.config → MilesBackendConfig(**settings) → miles_arguments() → Miles parser`. The existing `miles_arguments()` method translates Lilo's backend settings into Miles flags once. `cli_options` contains additional upstream flags using their argparse destination names. For a model without a Miles architecture preset, set `model_type=""` and supply its architecture options there. + +For Megatron: -For Megatron, the [FFT example](../src/lilo/configs/qwen35_4b_fft_64k.py) separates: +```python +"config": { + "tensor_model_parallel_size": 2, + "context_parallel_size": 2, + "micro_batch_size": 1, + "max_tokens_per_microbatch": 65536, + "provider_overrides": {"recompute_granularity": "full"}, + "optimizer": {"lr": 0.0001, "min_lr": 0.0001}, + "optimizer_overrides": {"adam_eps": 1e-8}, +} +``` -- `runtime`: Lilo's training loop, packing and parallelism settings. -- `provider`: attributes assigned to the actual Megatron Bridge model provider. -- `optimizer`: optimizer settings, including additional native Megatron constructor options. -- `distributed`: additional `DistributedDataParallelConfig` constructor options. +The path is `trainer.config → parse_backend_config() → EngineModelConfig`. The existing reader constructs the nested `OptimizerConfig` from `optimizer`. There is no `runtime` section or deployment-specific field map. `optimizer_overrides` explicitly supplies additional Megatron constructor options; the deployment code does not split optimizer fields by name. Tinker optimizer steps still apply the client's Adam parameters. -These section names are owned by Lilo. Native provider/optimizer/distributed fields do not need a Lilo allowlist. Shared precision and parallelism controls remain protected so Lilo's packing and collectives agree with Megatron. Tinker optimizer steps use Adam parameters supplied by the client. Native optimizer/distributed settings are recorded and checked for exact FFT checkpoint resume. +“Training loop” refers to the code that performs forward/backward passes and optimizer steps. It is not a separate config category. -`inference.config` contains SGLang options directly, such as `tp_size`, `mem_fraction_static` and `max_running_requests`. The real SGLang parser checks them during startup. Miles and SGLang support ordinary scalar, boolean and list arguments; unsupported custom/repeated argparse actions fail explicitly. +[`argparse_config.py`](../src/lilo/argparse_config.py) is used only at the Miles/SGLang process boundary. It appends configured values after preset arguments, leaving type conversion, choices and argument counts to the backend's parser. Boolean flags need special treatment: a `store_true` flag cannot express `False`, so the helper removes that preset flag and sets its default explicitly. Unsupported custom/repeated actions fail with an error. -Backend readers still check managed paths, GPU topology, adapter capacity and communication settings. Python configs do not bypass backend compatibility, adapter export/load requirements or available GPU memory. +[`config_validation.py`](../src/lilo/config_validation.py) only rejects options that would overwrite settings Lilo manages, such as checkpoint paths or parallelism already used by its collectives. Backend-specific modules retain those checks and resource/admission checks. They do not define another configuration schema. Backend support and memory capacity are still checked by the backend when workers start. ## Routing and updates diff --git a/docs/deployment-validation.md b/docs/deployment-validation.md index 73f0a5a..a8aa5d7 100644 --- a/docs/deployment-validation.md +++ b/docs/deployment-validation.md @@ -2,6 +2,14 @@ These checks exercise PR #55. The current implementation deploys trainers and inference provisioners independently; the shared frontend references them by app name. Earlier sections record validation of the previous shared-app implementation. +## Direct backend configuration + +Removed the deployment-specific Miles field-renaming table and Megatron `runtime/provider/optimizer/distributed` schema. Configs now use the existing `MilesBackendConfig` and `EngineModelConfig` field names. Megatron's existing config reader constructs its nested optimizer; extra Megatron constructor settings are explicit `*_overrides` dictionaries instead of being split by field name. + +Replaced `native_options.py` with `argparse_config.py`: configured values become arguments and the actual backend parser validates types/choices. Explicit false boolean flags use defaults after removing the preset flag. Renamed the Miles passthrough dictionary to `cli_options`. Replaced `backend_options.py` with a small `reject_managed_options` check. + +Validation: **619 CPU tests passed, 1 skipped**. All 14 examples produce the same effective trainer and inference settings as before (with the passthrough field renamed). CPU coverage includes Megatron constructor forwarding, Miles/SGLang overrides, boolean false, list/alias overrides, custom type converters and backend rejection of invalid types/choices. Ruff and whitespace checks passed. No apps were redeployed. + ## Model configs contain infrastructure settings only Removed the `deployment` section from `BaseConfig` and removed explicit model commits from the built-in examples. Frontend app/environment/region selection belongs to `lilo deploy`; platform metadata and automatically generated worker-release IDs live in saved records. Config authors do not need revision or runtime-version fields. The CLI resolves omitted model revisions from Hugging Face `main`. diff --git a/src/lilo/argparse_config.py b/src/lilo/argparse_config.py new file mode 100644 index 0000000..fa8a633 --- /dev/null +++ b/src/lilo/argparse_config.py @@ -0,0 +1,51 @@ +"""Translate config dictionaries into arguments for Miles and SGLang. + +Those backends expose argparse parsers. Append configured values after their +preset arguments so argparse itself handles types, choices and required options. +Boolean flags need defaults because a store_true flag cannot express False. +""" + +import argparse + + +def apply_config_overrides(parser, options, argv): + """Mutate argv/defaults before the backend calls parse_args.""" + boolean_actions = ( + argparse._StoreTrueAction, + argparse._StoreFalseAction, + argparse.BooleanOptionalAction, + ) + actions_by_name = {} + for action in parser._actions: + actions_by_name.setdefault(action.dest, []).append(action) + + for name, value in options.items(): + actions = actions_by_name.get(name, []) + if not actions or not all(action.option_strings for action in actions): + raise ValueError(f"unknown backend option: {name}") + if value is None: + raise ValueError(f"backend option {name} cannot be null") + if all(isinstance(action, boolean_actions) for action in actions): + if not isinstance(value, bool): + raise ValueError(f"backend option {name} requires a boolean") + flags = {flag for action in actions for flag in action.option_strings} + argv[:] = [arg for arg in argv if arg.split("=", 1)[0] not in flags] + for action in actions: + action.required = False + parser.set_defaults(**{name: value}) + continue + if len(actions) != 1 or type(actions[0]) is not argparse._StoreAction: + raise ValueError( + f"backend option {name} uses an unsupported argparse action" + ) + action = actions[0] + multiple = action.nargs in ("+", "*") or isinstance(action.nargs, int) + flag = action.option_strings[0] + if multiple: + if not isinstance(value, list): + raise ValueError(f"backend option {name} requires a list") + argv.extend([flag, *(str(item) for item in value)]) + else: + if isinstance(value, (dict, list, bool)): + raise ValueError(f"backend option {name} requires a scalar") + argv.append(f"{flag}={value}") diff --git a/src/lilo/backend_options.py b/src/lilo/backend_options.py deleted file mode 100644 index 831f8f6..0000000 --- a/src/lilo/backend_options.py +++ /dev/null @@ -1,14 +0,0 @@ -"""Helpers shared by backend-owned deployment configuration readers.""" - - -def native_options(values, protected): - if not isinstance(values, dict): - raise ValueError("backend configuration must be a mapping") - result = {} - for key, value in values.items(): - if not isinstance(key, str) or not key.isidentifier(): - raise ValueError(f"native option must use its underscore name: {key}") - if key in protected: - raise ValueError(f"option {key} is managed by Lilo") - result[key] = value - return result diff --git a/src/lilo/backends/megatron_deployment.py b/src/lilo/backends/megatron_deployment.py index 34a13af..532aeef 100644 --- a/src/lilo/backends/megatron_deployment.py +++ b/src/lilo/backends/megatron_deployment.py @@ -1,14 +1,13 @@ -"""Build Lilo loop settings while preserving native Megatron configuration.""" +"""Construct the existing Megatron training config without a second schema.""" - -from dataclasses import asdict, fields +from dataclasses import asdict from lilo.deployments import gpu_count -from lilo.backend_options import native_options -from .megatron_runtime.common.config import EngineModelConfig, OptimizerConfig +from lilo.config_validation import reject_managed_options +from .megatron_config import parse_backend_config # These values also control packing, collectives, and checkpoint metadata in Lilo. -# Configure them once under runtime so the provider and training loop agree. +# Configure them once on EngineModelConfig so packing and the provider agree. PROVIDER_MANAGED = { "tensor_model_parallel_size", "pipeline_model_parallel_size", @@ -44,48 +43,25 @@ def build_config(spec, asset_path): raise ValueError("FFT trainers admit one client per instance") if trainer["engine"]["sampler_persistence_concurrency"] != 1: raise ValueError("Megatron requires sampler_persistence_concurrency: 1") - sections = native_options(trainer["config"], set()) - unknown = sections.keys() - {"runtime", "provider", "optimizer", "distributed"} - if unknown: - raise ValueError(f"unknown Megatron config sections: {sorted(unknown)}") - runtime = native_options( - sections.get("runtime", {}), + settings = trainer["config"] + reject_managed_options(settings, {"hf_checkpoint", "seq_length"}) + reject_managed_options(settings.get("provider_overrides", {}), PROVIDER_MANAGED) + reject_managed_options( + settings.get("optimizer_overrides", {}), OPTIMIZER_MANAGED | {"optimizer"} + ) + reject_managed_options( + settings.get("distributed_overrides", {}), DISTRIBUTED_MANAGED + ) + config, _ = parse_backend_config( { - "hf_checkpoint", - "seq_length", - "max_lora_slots", - "max_lora_rank", - "optimizer", - "provider_overrides", - "optimizer_overrides", - "distributed_overrides", - }, + "megatron": { + **settings, + "hf_checkpoint": asset_path, + "seq_length": spec.model["max_context_length"], + } + } ) - provider = native_options(sections.get("provider", {}), PROVIDER_MANAGED) - optimizer = native_options(sections.get("optimizer", {}), OPTIMIZER_MANAGED) - distributed = native_options(sections.get("distributed", {}), DISTRIBUTED_MANAGED) - if optimizer.get("optimizer", "adam") != "adam": + if config.optimizer.optimizer != "adam": raise ValueError("Tinker optim_step requires an Adam optimizer") - # Keep scheduling and per-request optimizer settings available to Lilo. Every - # other optimizer field goes to the installed Megatron constructor unchanged. - loop_fields = {field.name for field in fields(OptimizerConfig)} - loop_optimizer = { - key: value for key, value in optimizer.items() if key in loop_fields - } - native_optimizer = { - key: value for key, value in optimizer.items() if key not in loop_fields - } - try: - config = EngineModelConfig( - hf_checkpoint=asset_path, - seq_length=spec.model["max_context_length"], - optimizer=OptimizerConfig(**loop_optimizer), - provider_overrides=provider, - optimizer_overrides=native_optimizer, - distributed_overrides=distributed, - **runtime, - ) - except TypeError as exc: - raise ValueError(f"invalid Megatron runtime options: {exc}") from exc config.validate(gpu_count(trainer["resources"])) return {"megatron": asdict(config), "checkpoint_dir": "/checkpoints"} diff --git a/src/lilo/backends/miles_config.py b/src/lilo/backends/miles_config.py index dfedcbb..1708647 100644 --- a/src/lilo/backends/miles_config.py +++ b/src/lilo/backends/miles_config.py @@ -44,7 +44,7 @@ class MilesBackendConfig: ) max_tokens_per_gpu: int = 8192 extra_args: tuple[str, ...] = () - native_options: dict[str, Any] = field(default_factory=dict) + cli_options: dict[str, Any] = field(default_factory=dict) @property def world_size(self) -> int: @@ -78,16 +78,20 @@ def validate(self) -> None: "max_tokens_per_gpu": self.max_tokens_per_gpu, } for name, value in positive.items(): - if value < 1: - raise ValueError(f"{name} must be at least 1") + if type(value) is not int or value < 1: + raise ValueError(f"{name} must be a positive integer") if not self.hf_checkpoint: raise ValueError("hf_checkpoint is required") - if not self.model_type and not self.native_options: + if not self.model_type and not self.cli_options: raise ValueError( "model_type or explicit native architecture options are required" ) - if not self.target_modules: - raise ValueError("target_modules must not be empty") + if ( + not isinstance(self.target_modules, (list, tuple)) + or not self.target_modules + or not all(isinstance(name, str) and name for name in self.target_modules) + ): + raise ValueError("target_modules must be a nonempty list of module names") if ( self.default_lora_alpha <= 0 or not float(self.default_lora_alpha).is_integer() diff --git a/src/lilo/backends/miles_deployment.py b/src/lilo/backends/miles_deployment.py index e6ebb70..e5e1e5d 100644 --- a/src/lilo/backends/miles_deployment.py +++ b/src/lilo/backends/miles_deployment.py @@ -1,13 +1,22 @@ """Miles deployment integration. Native options are validated by Miles at startup.""" - from dataclasses import asdict from lilo.deployments import gpu_count -from lilo.backend_options import native_options +from lilo.config_validation import reject_managed_options from .miles_config import MilesBackendConfig MILES_MANAGED = { + "context_parallel_size", + "expert_model_parallel_size", + "expert_tensor_parallel_size", + "lora_alpha", + "lora_dropout", + "lora_rank", + "max_tokens_per_gpu", + "multi_lora_n_adapters", + "target_modules", + "tensor_model_parallel_size", "hf_checkpoint", "load", "pretrained_checkpoint", @@ -37,48 +46,18 @@ def build_config(spec, asset_path): + """Pass trainer.config directly to MilesBackendConfig.""" trainer = spec.trainer if spec.model["parameterization"] != "lora": raise ValueError("Miles requires lora parameterization") - values = native_options(trainer["config"], set()) - unknown = values.keys() - {"model_args", "options"} - if unknown: - raise ValueError(f"unknown Miles config sections: {sorted(unknown)}") - options = native_options(values.get("options", {}), MILES_MANAGED) - names = { - "tensor_model_parallel_size": "tensor_model_parallel_size", - "context_parallel_size": "context_parallel_size", - "expert_model_parallel_size": "expert_model_parallel_size", - "expert_tensor_parallel_size": "expert_tensor_parallel_size", - "multi_lora_n_adapters": "max_lora_slots", - "lora_rank": "max_lora_rank", - "lora_alpha": "default_lora_alpha", - "lora_dropout": "lora_dropout", - "target_modules": "target_modules", - "max_tokens_per_gpu": "max_tokens_per_gpu", - } - settings = { - dest: options.pop(source) for source, dest in names.items() if source in options - } - for name, value in settings.items(): - if ( - name not in {"target_modules", "default_lora_alpha", "lora_dropout"} - and type(value) is not int - ): - raise ValueError(f"Miles option {name} must be an integer") - if "target_modules" in settings: - value = settings["target_modules"] - value = value.split(",") if isinstance(value, str) else value - if not isinstance(value, list) or not all( - isinstance(item, str) and item for item in value - ): - raise ValueError("target_modules must be a nonempty list of module names") - settings["target_modules"] = tuple(value) + settings = trainer["config"] + reject_managed_options( + settings, {"hf_checkpoint", "actor_num_gpus_per_node", "extra_args"} + ) + reject_managed_options(settings.get("cli_options", {}), MILES_MANAGED) config = MilesBackendConfig( hf_checkpoint=asset_path, - model_type=values.get("model_args") or "", actor_num_gpus_per_node=gpu_count(trainer["resources"]), - native_options=options, extra_args=("--seq-length", str(spec.model["max_context_length"])), **settings, ) @@ -88,5 +67,5 @@ def build_config(spec, asset_path): ): raise ValueError("expert parallel sizes must divide the trainer GPU allocation") if trainer["engine"]["max_clients_per_instance"] > config.max_lora_slots: - raise ValueError("max_clients_per_instance exceeds multi_lora_n_adapters") + raise ValueError("max_clients_per_instance exceeds max_lora_slots") return {"miles": asdict(config), "checkpoint_dir": "/checkpoints"} diff --git a/src/lilo/backends/miles_runtime/runtime.py b/src/lilo/backends/miles_runtime/runtime.py index 7a4d070..f0a902c 100644 --- a/src/lilo/backends/miles_runtime/runtime.py +++ b/src/lilo/backends/miles_runtime/runtime.py @@ -198,11 +198,11 @@ async def _start(self) -> None: else [] ) with _temporary_argv([*architecture, *self.config.miles_arguments()]): - if self.config.native_options: - from lilo.native_options import apply_defaults + if self.config.cli_options: + from lilo.argparse_config import apply_config_overrides def configure(parser): - apply_defaults(parser, self.config.native_options, sys.argv) + apply_config_overrides(parser, self.config.cli_options, sys.argv) return parser args = parse_args(add_custom_arguments=configure, entry="serve") diff --git a/src/lilo/config_validation.py b/src/lilo/config_validation.py new file mode 100644 index 0000000..a9a10ad --- /dev/null +++ b/src/lilo/config_validation.py @@ -0,0 +1,9 @@ +"""Checks for backend options that would override Lilo-managed settings.""" + + +def reject_managed_options(options, managed): + if not isinstance(options, dict): + raise ValueError("backend configuration must be a mapping") + conflicts = options.keys() & managed + if conflicts: + raise ValueError(f"options managed by Lilo: {sorted(conflicts)}") diff --git a/src/lilo/configs/qwen35_35b_a3b_fft_64k.py b/src/lilo/configs/qwen35_35b_a3b_fft_64k.py index 0269890..43a29db 100644 --- a/src/lilo/configs/qwen35_35b_a3b_fft_64k.py +++ b/src/lilo/configs/qwen35_35b_a3b_fft_64k.py @@ -18,21 +18,19 @@ class Config(BaseConfig): "TORCHINDUCTOR_COMPILE_THREADS": "1", }, "config": { - "runtime": { - "tensor_model_parallel_size": 4, - "pipeline_model_parallel_size": 1, - "context_parallel_size": 2, - "expert_model_parallel_size": 8, - "expert_tensor_parallel_size": 1, - "sequence_parallel": True, - "micro_batch_size": 1, - "max_tokens_per_microbatch": 65536, - "bf16": True, - "fp16": False, - "gpu_memory_fraction": 0.9, - "use_distributed_optimizer": True, - }, - "provider": { + "tensor_model_parallel_size": 4, + "pipeline_model_parallel_size": 1, + "context_parallel_size": 2, + "expert_model_parallel_size": 8, + "expert_tensor_parallel_size": 1, + "sequence_parallel": True, + "micro_batch_size": 1, + "max_tokens_per_microbatch": 65536, + "bf16": True, + "fp16": False, + "gpu_memory_fraction": 0.9, + "use_distributed_optimizer": True, + "provider_overrides": { "mtp_num_layers": 0, "recompute_granularity": "selective", "moe_layer_recompute": True, diff --git a/src/lilo/configs/qwen35_4b_fft_64k.py b/src/lilo/configs/qwen35_4b_fft_64k.py index f62e54d..b6a8025 100644 --- a/src/lilo/configs/qwen35_4b_fft_64k.py +++ b/src/lilo/configs/qwen35_4b_fft_64k.py @@ -14,17 +14,15 @@ class Config(BaseConfig): "resources": {"gpu": "H100:4"}, "engine": {"max_clients_per_instance": 1, "sampler_persistence_concurrency": 1}, "config": { - "runtime": { - "tensor_model_parallel_size": 2, - "context_parallel_size": 2, - "sequence_parallel": True, - "micro_batch_size": 1, - "max_tokens_per_microbatch": 65536, - "defer_fp32_logits": True, - "fp32_lm_head": True, - "use_distributed_optimizer": True, - }, - "provider": { + "tensor_model_parallel_size": 2, + "context_parallel_size": 2, + "sequence_parallel": True, + "micro_batch_size": 1, + "max_tokens_per_microbatch": 65536, + "defer_fp32_logits": True, + "fp32_lm_head": True, + "use_distributed_optimizer": True, + "provider_overrides": { "mtp_num_layers": 0, "recompute_granularity": "full", "recompute_method": "uniform", diff --git a/src/lilo/configs/qwen35_9b_fft_64k.py b/src/lilo/configs/qwen35_9b_fft_64k.py index 1ace764..aff3bc8 100644 --- a/src/lilo/configs/qwen35_9b_fft_64k.py +++ b/src/lilo/configs/qwen35_9b_fft_64k.py @@ -18,17 +18,15 @@ class Config(BaseConfig): "TORCHINDUCTOR_COMPILE_THREADS": "1", }, "config": { - "runtime": { - "tensor_model_parallel_size": 2, - "context_parallel_size": 2, - "sequence_parallel": True, - "micro_batch_size": 1, - "max_tokens_per_microbatch": 65536, - "defer_fp32_logits": True, - "fp32_lm_head": True, - "use_distributed_optimizer": True, - }, - "provider": { + "tensor_model_parallel_size": 2, + "context_parallel_size": 2, + "sequence_parallel": True, + "micro_batch_size": 1, + "max_tokens_per_microbatch": 65536, + "defer_fp32_logits": True, + "fp32_lm_head": True, + "use_distributed_optimizer": True, + "provider_overrides": { "mtp_num_layers": 0, "recompute_granularity": "full", "recompute_method": "uniform", diff --git a/src/lilo/configs/qwen35_9b_instruct_lora_16k.py b/src/lilo/configs/qwen35_9b_instruct_lora_16k.py index fdeb869..0138a62 100644 --- a/src/lilo/configs/qwen35_9b_instruct_lora_16k.py +++ b/src/lilo/configs/qwen35_9b_instruct_lora_16k.py @@ -18,22 +18,17 @@ class Config(BaseConfig): "TORCHINDUCTOR_COMPILE_THREADS": "1", }, "config": { - "model_args": "qwen3.5-9B", - "options": { - "tensor_model_parallel_size": 8, - "target_modules": [ - "linear_qkv", - "linear_proj", - "linear_fc1", - "linear_fc2", - ], - "max_tokens_per_gpu": 16384, + "model_type": "qwen3.5-9B", + "tensor_model_parallel_size": 8, + "target_modules": ["linear_qkv", "linear_proj", "linear_fc1", "linear_fc2"], + "max_tokens_per_gpu": 16384, + "max_lora_slots": 6, + "max_lora_rank": 32, + "default_lora_alpha": 32, + "cli_options": { "recompute_granularity": "full", "recompute_method": "uniform", "recompute_num_layers": 1, - "multi_lora_n_adapters": 6, - "lora_rank": 32, - "lora_alpha": 32, }, }, } diff --git a/src/lilo/configs/qwen35_9b_instruct_lora_16k_dp2.py b/src/lilo/configs/qwen35_9b_instruct_lora_16k_dp2.py index 00b460a..961ad8c 100644 --- a/src/lilo/configs/qwen35_9b_instruct_lora_16k_dp2.py +++ b/src/lilo/configs/qwen35_9b_instruct_lora_16k_dp2.py @@ -5,5 +5,5 @@ class Config(ParentConfig): name = "qwen35-9b-instruct-lora-16k-dp2" overrides = { "routing.default": False, - "trainer.config.options.tensor_model_parallel_size": 4, + "trainer.config.tensor_model_parallel_size": 4, } diff --git a/src/lilo/configs/qwen35_9b_lora_16k.py b/src/lilo/configs/qwen35_9b_lora_16k.py index 6e4a4e4..d02f918 100644 --- a/src/lilo/configs/qwen35_9b_lora_16k.py +++ b/src/lilo/configs/qwen35_9b_lora_16k.py @@ -19,20 +19,20 @@ class Config(BaseConfig): "TORCHINDUCTOR_COMPILE_THREADS": "1", }, "config": { - "model_args": "qwen3.5-9B", - "options": { - "tensor_model_parallel_size": 4, - "multi_lora_n_adapters": 6, - "lora_rank": 32, - "lora_alpha": 32, - "target_modules": [ - "linear_qkv", - "linear_proj", - "linear_fc1", - "linear_fc2", - "output_layer", - ], - "max_tokens_per_gpu": 16384, + "model_type": "qwen3.5-9B", + "tensor_model_parallel_size": 4, + "max_lora_slots": 6, + "max_lora_rank": 32, + "default_lora_alpha": 32, + "target_modules": [ + "linear_qkv", + "linear_proj", + "linear_fc1", + "linear_fc2", + "output_layer", + ], + "max_tokens_per_gpu": 16384, + "cli_options": { "recompute_granularity": "full", "recompute_method": "uniform", "recompute_num_layers": 1, diff --git a/src/lilo/configs/qwen35_9b_lora_2k.py b/src/lilo/configs/qwen35_9b_lora_2k.py index eb0de08..2ee6759 100644 --- a/src/lilo/configs/qwen35_9b_lora_2k.py +++ b/src/lilo/configs/qwen35_9b_lora_2k.py @@ -18,23 +18,23 @@ class Config(BaseConfig): "TORCHINDUCTOR_COMPILE_THREADS": "1", }, "config": { - "model_args": "qwen3.5-9B", - "options": { - "tensor_model_parallel_size": 4, - "target_modules": [ - "linear_qkv", - "linear_proj", - "linear_fc1", - "linear_fc2", - "output_layer", - ], - "max_tokens_per_gpu": 2048, + "model_type": "qwen3.5-9B", + "tensor_model_parallel_size": 4, + "target_modules": [ + "linear_qkv", + "linear_proj", + "linear_fc1", + "linear_fc2", + "output_layer", + ], + "max_tokens_per_gpu": 2048, + "max_lora_slots": 4, + "max_lora_rank": 32, + "default_lora_alpha": 32, + "cli_options": { "recompute_granularity": "full", "recompute_method": "uniform", "recompute_num_layers": 1, - "multi_lora_n_adapters": 4, - "lora_rank": 32, - "lora_alpha": 32, }, }, } diff --git a/src/lilo/configs/qwen35_9b_lora_64k.py b/src/lilo/configs/qwen35_9b_lora_64k.py index 7420a6a..d7646cd 100644 --- a/src/lilo/configs/qwen35_9b_lora_64k.py +++ b/src/lilo/configs/qwen35_9b_lora_64k.py @@ -7,6 +7,6 @@ class Config(ParentConfig): "model.max_context_length": 65536, "routing.default": False, "trainer.resources.gpu": "H200:8", - "trainer.config.options.tensor_model_parallel_size": 8, - "trainer.config.options.max_tokens_per_gpu": 65536, + "trainer.config.tensor_model_parallel_size": 8, + "trainer.config.max_tokens_per_gpu": 65536, } diff --git a/src/lilo/configs/qwen36_27b_fft_64k.py b/src/lilo/configs/qwen36_27b_fft_64k.py index 60bb5c4..4dd8fab 100644 --- a/src/lilo/configs/qwen36_27b_fft_64k.py +++ b/src/lilo/configs/qwen36_27b_fft_64k.py @@ -18,19 +18,17 @@ class Config(BaseConfig): "TORCHINDUCTOR_COMPILE_THREADS": "1", }, "config": { - "runtime": { - "tensor_model_parallel_size": 4, - "pipeline_model_parallel_size": 1, - "context_parallel_size": 2, - "sequence_parallel": True, - "micro_batch_size": 1, - "max_tokens_per_microbatch": 65536, - "bf16": True, - "fp16": False, - "gpu_memory_fraction": 0.9, - "use_distributed_optimizer": True, - }, - "provider": { + "tensor_model_parallel_size": 4, + "pipeline_model_parallel_size": 1, + "context_parallel_size": 2, + "sequence_parallel": True, + "micro_batch_size": 1, + "max_tokens_per_microbatch": 65536, + "bf16": True, + "fp16": False, + "gpu_memory_fraction": 0.9, + "use_distributed_optimizer": True, + "provider_overrides": { "mtp_num_layers": 0, "recompute_granularity": "full", "recompute_method": "uniform", diff --git a/src/lilo/configs/qwen38_27b_lora_128k.py b/src/lilo/configs/qwen38_27b_lora_128k.py index f6af875..1549fea 100644 --- a/src/lilo/configs/qwen38_27b_lora_128k.py +++ b/src/lilo/configs/qwen38_27b_lora_128k.py @@ -6,9 +6,9 @@ class Config(ParentConfig): overrides = { "model.max_context_length": 131072, "routing.default": False, - "trainer.config.options.tensor_model_parallel_size": 2, - "trainer.config.options.context_parallel_size": 4, - "trainer.config.options.max_tokens_per_gpu": 32768, + "trainer.config.tensor_model_parallel_size": 2, + "trainer.config.context_parallel_size": 4, + "trainer.config.max_tokens_per_gpu": 32768, "inference.resources.gpu": "H200:2", "inference.scaling.max_replicas": 4, "inference.scaling.target_concurrency": 4, diff --git a/src/lilo/configs/qwen38_27b_lora_16k.py b/src/lilo/configs/qwen38_27b_lora_16k.py index 2cb36a6..6c0432e 100644 --- a/src/lilo/configs/qwen38_27b_lora_16k.py +++ b/src/lilo/configs/qwen38_27b_lora_16k.py @@ -18,22 +18,17 @@ class Config(BaseConfig): "TORCHINDUCTOR_COMPILE_THREADS": "1", }, "config": { - "model_args": "qwen3.8-27B", - "options": { - "tensor_model_parallel_size": 4, - "target_modules": [ - "linear_qkv", - "linear_proj", - "linear_fc1", - "linear_fc2", - ], - "max_tokens_per_gpu": 16384, + "model_type": "qwen3.8-27B", + "tensor_model_parallel_size": 4, + "target_modules": ["linear_qkv", "linear_proj", "linear_fc1", "linear_fc2"], + "max_tokens_per_gpu": 16384, + "max_lora_slots": 6, + "max_lora_rank": 32, + "default_lora_alpha": 32, + "cli_options": { "recompute_granularity": "full", "recompute_method": "uniform", "recompute_num_layers": 1, - "multi_lora_n_adapters": 6, - "lora_rank": 32, - "lora_alpha": 32, }, }, } diff --git a/src/lilo/configs/qwen38_27b_lora_64k.py b/src/lilo/configs/qwen38_27b_lora_64k.py index 2bb796d..c65b30b 100644 --- a/src/lilo/configs/qwen38_27b_lora_64k.py +++ b/src/lilo/configs/qwen38_27b_lora_64k.py @@ -6,8 +6,8 @@ class Config(ParentConfig): overrides = { "model.max_context_length": 65536, "routing.default": False, - "trainer.config.options.context_parallel_size": 2, - "trainer.config.options.max_tokens_per_gpu": 32768, + "trainer.config.context_parallel_size": 2, + "trainer.config.max_tokens_per_gpu": 32768, "inference.scaling.target_concurrency": 8, "inference.config.max_running_requests": 16, } diff --git a/src/lilo/deployments.py b/src/lilo/deployments.py index 34fdb5a..46ba58e 100644 --- a/src/lilo/deployments.py +++ b/src/lilo/deployments.py @@ -202,9 +202,9 @@ def inference_hash(self) -> str: """Identify inference settings, including the adapter shape it must load.""" adapter = {} if self.spec.model["parameterization"] == "lora": - options = self.spec.trainer["config"].get("options", {}) + options = self.spec.trainer["config"] adapter = { - "lora_rank": options.get("lora_rank"), + "max_lora_rank": options.get("max_lora_rank"), "target_modules": options.get("target_modules"), } return settings_hash( diff --git a/src/lilo/inference/native_sglang.py b/src/lilo/inference/native_sglang.py index 12c193a..56ebd77 100644 --- a/src/lilo/inference/native_sglang.py +++ b/src/lilo/inference/native_sglang.py @@ -6,7 +6,7 @@ import os import sys -from lilo.native_options import apply_defaults +from lilo.argparse_config import apply_config_overrides def main(): @@ -19,7 +19,7 @@ def main(): parser = argparse.ArgumentParser() ServerArgs.add_cli_args(parser) argv = ["--model-path", sys.argv[1], "--host", "127.0.0.1", "--port", "8001"] - apply_defaults(parser, json.loads(sys.argv[2]), argv) + apply_config_overrides(parser, json.loads(sys.argv[2]), argv) raw = parser.parse_args(argv) logging.basicConfig(level=getattr(logging, raw.log_level.upper())) args = ServerArgs.from_cli_args(raw) diff --git a/src/lilo/inference/sglang_deployment.py b/src/lilo/inference/sglang_deployment.py index 3242e83..84cc0ea 100644 --- a/src/lilo/inference/sglang_deployment.py +++ b/src/lilo/inference/sglang_deployment.py @@ -1,8 +1,7 @@ """SGLang settings that must agree with Lilo replica orchestration.""" - from lilo.deployments import gpu_count -from lilo.backend_options import native_options +from lilo.config_validation import reject_managed_options SGLANG_MANAGED = { "model_path", @@ -32,7 +31,8 @@ def build_config(spec): - options = native_options(spec.inference["config"], SGLANG_MANAGED) + options = dict(spec.inference["config"]) + reject_managed_options(options, SGLANG_MANAGED) tp = options.get("tp_size", gpu_count(spec.inference["resources"])) ep = options.get("ep_size", 1) if not isinstance(tp, int) or tp != gpu_count(spec.inference["resources"]): diff --git a/src/lilo/native_options.py b/src/lilo/native_options.py deleted file mode 100644 index b29c5d0..0000000 --- a/src/lilo/native_options.py +++ /dev/null @@ -1,92 +0,0 @@ -"""Apply typed config values through a backend's own argparse schema.""" - -from __future__ import annotations - -import argparse - - -def apply_defaults( - parser: argparse.ArgumentParser, options: dict, argv: list[str] -) -> None: - """Override preset arguments after parser construction, before backend validation. - - Values use argparse destination names. Ordinary scalar/list options and boolean - flags are supported. Custom argparse actions fail explicitly instead of being - bypassed by set_defaults. - """ - by_name = {} - for action in parser._actions: - by_name.setdefault(action.dest, []).append(action) - flags, defaults = {}, {} - booleans = ( - argparse._StoreTrueAction, - argparse._StoreFalseAction, - argparse.BooleanOptionalAction, - ) - for key, value in options.items(): - actions = by_name.get(key, []) - if not actions or not all(action.option_strings for action in actions): - raise ValueError(f"unknown backend option: {key}") - if value is None: - raise ValueError(f"backend option {key} cannot be null") - if all(isinstance(action, booleans) for action in actions): - if not isinstance(value, bool): - raise ValueError(f"backend option {key} requires a boolean") - else: - if len(actions) != 1 or type(actions[0]) is not argparse._StoreAction: - raise ValueError( - f"backend option {key} uses an unsupported argparse action" - ) - action = actions[0] - multiple = action.nargs in ("+", "*") or isinstance(action.nargs, int) - if multiple and not isinstance(value, list): - raise ValueError(f"backend option {key} requires a list") - if not multiple and isinstance(value, (dict, list, bool)): - raise ValueError(f"backend option {key} requires a scalar") - values = value if multiple else [value] - if action.nargs == "+" and not values: - raise ValueError(f"backend option {key} requires a nonempty list") - if isinstance(action.nargs, int) and len(values) != action.nargs: - raise ValueError(f"backend option {key} requires {action.nargs} values") - cast = action.type or (lambda x: x) - try: - # argparse type callbacks receive command-line text, including - # custom converters such as SGLang human_readable_int. - values = [cast(str(v)) for v in values] - except (ValueError, TypeError, argparse.ArgumentTypeError) as exc: - raise ValueError( - f"invalid value for backend option {key}: {exc}" - ) from exc - if action.choices is not None and any( - v not in action.choices for v in values - ): - raise ValueError(f"invalid choice for backend option {key}") - value = values if multiple else values[0] - defaults[key] = value - for action in actions: - action.required = False - for flag in action.option_strings: - flags[flag] = action - known = {flag for action in parser._actions for flag in action.option_strings} - kept, i = [], 0 - while i < len(argv): - token = argv[i] - action = flags.get(token.split("=", 1)[0]) - if action is None: - kept.append(token) - i += 1 - continue - i += 1 - if "=" in token or action.nargs == 0: - continue - remaining = action.nargs if isinstance(action.nargs, int) else 1 - while i < len(argv): - if argv[i].split("=", 1)[0] in known or argv[i].startswith("--"): - break - i += 1 - if action.nargs not in ("+", "*"): - remaining -= 1 - if remaining == 0: - break - argv[:] = kept - parser.set_defaults(**defaults) diff --git a/tests/backends/test_native_megatron_config.py b/tests/backends/test_native_megatron_config.py index eecdcd9..d62283e 100644 --- a/tests/backends/test_native_megatron_config.py +++ b/tests/backends/test_native_megatron_config.py @@ -17,11 +17,13 @@ from lilo.backends.megatron_runtime.common import modeling -def test_yaml_native_values_reach_megatron(monkeypatch): +def test_config_overrides_reach_megatron(monkeypatch): data = asdict(load(config_path("qwen35-4b-fft-64k"))) - data["trainer"]["config"]["optimizer"]["native_optimizer_setting"] = False - data["trainer"]["config"]["distributed"] = {"native_ddp_setting": 123} - data["trainer"]["config"]["provider"]["native_provider_setting"] = [1, 2] + data["trainer"]["config"]["optimizer_overrides"] = { + "native_optimizer_setting": False + } + data["trainer"]["config"]["distributed_overrides"] = {"native_ddp_setting": 123} + data["trainer"]["config"]["provider_overrides"]["native_provider_setting"] = [1, 2] config, _ = parse_backend_config( backend_config(TypeAdapter(BaseConfig).validate_python(data)) ) diff --git a/tests/test_deployments.py b/tests/test_deployments.py index 5e24283..e7fbea4 100644 --- a/tests/test_deployments.py +++ b/tests/test_deployments.py @@ -18,7 +18,7 @@ from lilo.control_plane.deployments import DeploymentRoutes from lilo.backends.deployment import backend_config from lilo.providers.modal.deployment_apps import definition_from_spec -from lilo.native_options import apply_defaults +from lilo.argparse_config import apply_config_overrides def recipe(preset="qwen35-9b-lora-16k", **changes): @@ -48,7 +48,7 @@ def test_presets_context_topology_and_backend_options(): config["tensor_model_parallel_size"], config["max_lora_slots"], ) == (4, 4, 6) - assert config["native_options"]["recompute_num_layers"] == 1 + assert config["cli_options"]["recompute_num_layers"] == 1 assert config["extra_args"] == ("--seq-length", "16384") large = recipe("qwen35-9b-lora-64k") assert large.model["max_context_length"] == 65536 @@ -61,8 +61,8 @@ def test_presets_context_topology_and_backend_options(): def test_no_model_catalog_required(): spec = recipe( model__id="my-org/new-model", - trainer__config__model_args=None, - trainer__config__options={ + trainer__config__model_type="", + trainer__config__cli_options={ "num_layers": 12, "hidden_size": 768, "num_attention_heads": 12, @@ -75,11 +75,14 @@ def test_no_model_catalog_required(): @pytest.mark.parametrize( "changes,match", [ - ({"trainer__config__options": {"hf_checkpoint": "other"}}, "managed"), - ({"trainer__config__options": {"pipeline_model_parallel_size": 2}}, "managed"), + ({"trainer__config__cli_options": {"hf_checkpoint": "other"}}, "managed"), + ( + {"trainer__config__cli_options": {"pipeline_model_parallel_size": 2}}, + "managed", + ), ({"inference__config": {"model_path": "other"}}, "managed"), ({"inference__config": {"tp_size": 2}}, "replica GPU"), - ({"trainer__engine__max_clients_per_instance": 7}, "multi_lora_n_adapters"), + ({"trainer__engine__max_clients_per_instance": 7}, "max_lora_slots"), ( { "inference__config": { @@ -169,7 +172,9 @@ def test_native_false_list_aliases_and_scalar_overrides(): parser.add_argument("--tp", "--tensor-parallel-size", dest="tp_size", type=int) parser.add_argument("--unchanged") argv = ["--use-feature", "--layers", "1", "2", "--tp=4", "--unchanged", "keep"] - apply_defaults(parser, {"use_feature": False, "layers": [3], "tp_size": 8}, argv) + apply_config_overrides( + parser, {"use_feature": False, "layers": [3], "tp_size": 8}, argv + ) parsed = parser.parse_args(argv) assert vars(parsed) == { "use_feature": False, @@ -178,11 +183,11 @@ def test_native_false_list_aliases_and_scalar_overrides(): "unchanged": "keep", } with pytest.raises(ValueError, match="unknown backend option"): - apply_defaults(parser, {"typo": 1}, []) + apply_config_overrides(parser, {"typo": 1}, []) with pytest.raises(ValueError, match="boolean"): - apply_defaults(parser, {"use_feature": "false"}, []) + apply_config_overrides(parser, {"use_feature": "false"}, []) with pytest.raises(ValueError, match="list"): - apply_defaults(parser, {"layers": "1,2"}, []) + apply_config_overrides(parser, {"layers": "1,2"}, []) def test_multiple_models_same_http_service_and_old_binding_survives_switch(): @@ -234,7 +239,7 @@ def test_native_boolean_opposite_flags_and_optional_value(): parser.add_argument("--optional", nargs="?") parser.add_argument("--keep", action="store_true") argv = ["--bias", "--optional", "--keep"] - apply_defaults(parser, {"bias": False, "optional": "supplied"}, argv) + apply_config_overrides(parser, {"bias": False, "optional": "supplied"}, argv) assert vars(parser.parse_args(argv)) == { "bias": False, "optional": "supplied", @@ -242,7 +247,7 @@ def test_native_boolean_opposite_flags_and_optional_value(): } parser.add_argument("--custom", action="append") with pytest.raises(ValueError, match="unsupported argparse action"): - apply_defaults(parser, {"custom": [1]}, []) + apply_config_overrides(parser, {"custom": [1]}, []) def test_native_type_callbacks_receive_text(): @@ -254,8 +259,9 @@ def readable_int(value): parser = argparse.ArgumentParser() parser.add_argument("--context-length", type=readable_int) parser.add_argument("--sizes", nargs="+", type=readable_int) - apply_defaults(parser, {"context_length": 65536, "sizes": [32, "2k"]}, []) - args = parser.parse_args([]) + argv = [] + apply_config_overrides(parser, {"context_length": 65536, "sizes": [32, "2k"]}, argv) + args = parser.parse_args(argv) assert args.context_length == 65536 assert args.sizes == [32, 2000] @@ -266,12 +272,14 @@ def test_native_sections_survive_serialization_without_allowlist(): spec = recipe("qwen35-4b-fft-64k") data = asdict(spec) - data["trainer"]["config"]["provider"]["future_provider_option"] = { + data["trainer"]["config"]["provider_overrides"]["future_provider_option"] = { "layers": [1, 4], "enabled": False, } - data["trainer"]["config"]["optimizer"]["future_optimizer_option"] = 0.125 - data["trainer"]["config"]["distributed"] = {"future_ddp_option": False} + data["trainer"]["config"]["optimizer_overrides"] = { + "future_optimizer_option": 0.125 + } + data["trainer"]["config"]["distributed_overrides"] = {"future_ddp_option": False} spec = TypeAdapter(BaseConfig).validate_python(data) settings = backend_config(spec, "/assets/pinned") config, _ = parse_backend_config(json.loads(json.dumps(settings))) @@ -284,22 +292,20 @@ def test_native_sections_survive_serialization_without_allowlist(): assert config.optimizer_overrides == {"future_optimizer_option": 0.125} assert config.distributed_overrides == {"future_ddp_option": False} assert config.optimizer.lr == 0.0001 - assert asdict(spec) == data # Building does not consume or mutate YAML. + assert asdict(spec) == data # Building does not consume or mutate the config. @pytest.mark.parametrize( "section,options,match", [ - ("provider", {"context_parallel_size": 4}, "managed"), - ("optimizer", {"bf16": False}, "managed"), - ("distributed", {"use_distributed_optimizer": False}, "managed"), - ("runtime", {"optimizer_overrides": {}}, "managed"), - ("runtime", {"misspelled_loop_option": 1}, "runtime options"), - ("provider", [], "mapping"), + ("provider_overrides", {"context_parallel_size": 4}, "managed"), + ("optimizer_overrides", {"bf16": False}, "managed"), + ("distributed_overrides", {"use_distributed_optimizer": False}, "managed"), + ("provider_overrides", [], "mapping"), ("optimizer", {"optimizer": "sgd"}, "Adam"), ], ) -def test_megatron_native_options_preserve_integration_contract(section, options, match): +def test_megatron_cli_options_preserve_integration_contract(section, options, match): spec = recipe("qwen35-4b-fft-64k", **{f"trainer__config__{section}": options}) with pytest.raises(ValueError, match=match): backend_config(spec) @@ -315,10 +321,10 @@ def test_new_miles_and_sglang_options_need_no_deployment_schema_change(): from lilo.backends.deployment import serving_options spec = recipe( - trainer__config__options__future_miles_option=[1, 2], + trainer__config__cli_options__future_miles_option=[1, 2], inference__config__future_sglang_option=False, ) - assert backend_config(spec)["miles"]["native_options"]["future_miles_option"] == [ + assert backend_config(spec)["miles"]["cli_options"]["future_miles_option"] == [ 1, 2, ] @@ -367,8 +373,8 @@ def test_record_creation_copies_without_reparsing(): ).encode() ).hexdigest() ) - row.spec.trainer["config"]["options"]["lora_rank"] = 64 - assert spec.trainer["config"]["options"]["lora_rank"] == 32 + row.spec.trainer["config"]["max_lora_rank"] = 64 + assert spec.trainer["config"]["max_lora_rank"] == 32 saved = row.model_dump_json() assert DeploymentRecord.model_validate_json(saved) == row @@ -381,18 +387,18 @@ def test_python_config_inheritance_and_independent_defaults(tmp_path): "from lilo.configs.qwen35_9b_lora_64k import Config as ParentConfig\n" "class Config(ParentConfig):\n" " name = 'custom'\n" - " overrides = {'trainer.config.options.new_backend_option': False}\n" + " overrides = {'trainer.config.cli_options.new_backend_option': False}\n" ) first, second = load(path), load(path) assert is_dataclass(first) assert first.name == "custom" assert first.model["max_context_length"] == 65536 - assert first.trainer["config"]["options"]["new_backend_option"] is False - first.trainer["config"]["options"]["target_modules"].append("extra") - assert "extra" not in second.trainer["config"]["options"]["target_modules"] + assert first.trainer["config"]["cli_options"]["new_backend_option"] is False + first.trainer["config"]["target_modules"].append("extra") + assert "extra" not in second.trainer["config"]["target_modules"] assert ( "extra" - not in load(config_path("qwen35-9b-lora-16k")).trainer["config"]["options"][ + not in load(config_path("qwen35-9b-lora-16k")).trainer["config"][ "target_modules" ] ) @@ -443,36 +449,38 @@ def test_overrides_inherit_replace_and_copy_values(): class Parent(Example): overrides = { - "trainer.config.options.target_modules": ["parent"], - "trainer.config.options.future_option": {"enabled": True}, + "trainer.config.target_modules": ["parent"], + "trainer.config.cli_options.future_option": {"enabled": True}, "inference.config.max_running_requests": 24, } class Child(Parent): name = "child" overrides = { - "trainer.config.options.target_modules": ["child"], - "trainer.config.options.future_option": {"enabled": False}, + "trainer.config.target_modules": ["child"], + "trainer.config.cli_options.future_option": {"enabled": False}, } child = Child() assert child.inference["config"]["max_running_requests"] == 24 - assert child.trainer["config"]["options"]["target_modules"] == ["child"] - assert child.trainer["config"]["options"]["future_option"] == {"enabled": False} - child.trainer["config"]["options"]["target_modules"].append("changed") - child.trainer["config"]["options"]["future_option"]["enabled"] = True - assert Child.overrides["trainer.config.options.target_modules"] == ["child"] - assert Child().trainer["config"]["options"]["future_option"] == {"enabled": False} - assert Parent().trainer["config"]["options"]["target_modules"] == ["parent"] + assert child.trainer["config"]["target_modules"] == ["child"] + assert child.trainer["config"]["cli_options"]["future_option"] == {"enabled": False} + child.trainer["config"]["target_modules"].append("changed") + child.trainer["config"]["cli_options"]["future_option"]["enabled"] = True + assert Child.overrides["trainer.config.target_modules"] == ["child"] + assert Child().trainer["config"]["cli_options"]["future_option"] == { + "enabled": False + } + assert Parent().trainer["config"]["target_modules"] == ["parent"] assert "overrides" not in asdict(child) assert Child(name="keyword").name == "keyword" class Replacement(Parent): - trainer = {"resources": {"gpu": "H200:8"}, "config": {"options": {}}} - overrides = {"trainer.config.options.new_option": 1} + trainer = {"resources": {"gpu": "H200:8"}, "config": {"cli_options": {}}} + overrides = {"trainer.config.cli_options.new_option": 1} # A child's complete field replacement wins over its parent's dotted edits. - assert Replacement().trainer["config"] == {"options": {"new_option": 1}} + assert Replacement().trainer["config"] == {"cli_options": {"new_option": 1}} @pytest.mark.parametrize( @@ -513,11 +521,11 @@ def test_worker_hashes_cover_only_their_settings(): assert inference.trainer_hash == base.trainer_hash assert inference.inference_hash != base.inference_hash - trainer = resolved(recipe(trainer__config__options__max_tokens_per_gpu=8192)) + trainer = resolved(recipe(trainer__config__max_tokens_per_gpu=8192)) assert trainer.trainer_hash != base.trainer_hash assert trainer.inference_hash == base.inference_hash - adapter = resolved(recipe(trainer__config__options__lora_rank=64)) + adapter = resolved(recipe(trainer__config__max_lora_rank=64)) assert adapter.trainer_hash != base.trainer_hash assert adapter.inference_hash != base.inference_hash @@ -543,3 +551,13 @@ def test_examples_only_contain_model_infrastructure(): assert "revision" not in config.model assert "runtime_version" not in config.trainer assert "runtime_version" not in config.inference + + +@pytest.mark.parametrize("value", ["invalid", 7]) +def test_backend_parser_validates_configured_types_and_choices(value): + parser = argparse.ArgumentParser() + parser.add_argument("--count", type=int, choices=[1, 2], required=True) + argv = ["--count", "1"] + apply_config_overrides(parser, {"count": value}, argv) + with pytest.raises(SystemExit): + parser.parse_args(argv) From 9468b53404a52f1af49d1c3d57aaf2692b31b9cc Mon Sep 17 00:00:00 2001 From: kailash Date: Wed, 23 Sep 2026 22:46:57 +0000 Subject: [PATCH 19/27] Compose typed deployment configs and resolve backend settings once --- docs/deployment-configs.md | 276 +++++++---------- docs/deployment-validation.md | 14 + scripts/definition_smoke.py | 48 +-- scripts/e2e_engine_definition.py | 14 +- src/lilo/backends/deployment.py | 52 ++-- src/lilo/backends/megatron_deployment.py | 49 +-- .../megatron_runtime/common/config.py | 5 + .../megatron_runtime/common/modeling.py | 48 +-- .../megatron_runtime/common/settings.py | 57 ++++ .../miles_arguments.py} | 4 +- src/lilo/backends/miles_deployment.py | 16 +- src/lilo/backends/miles_runtime/runtime.py | 2 +- src/lilo/configs/__init__.py | 2 +- src/lilo/configs/qwen35_35b_a3b_fft_64k.py | 47 ++- src/lilo/configs/qwen35_4b_fft_64k.py | 38 ++- src/lilo/configs/qwen35_9b_fft_64k.py | 47 ++- .../configs/qwen35_9b_instruct_lora_16k.py | 53 ++-- .../qwen35_9b_instruct_lora_16k_dp2.py | 17 +- src/lilo/configs/qwen35_9b_lora_16k.py | 46 ++- src/lilo/configs/qwen35_9b_lora_16k_single.py | 12 +- src/lilo/configs/qwen35_9b_lora_2k.py | 53 ++-- src/lilo/configs/qwen35_9b_lora_64k.py | 27 +- src/lilo/configs/qwen36_27b_fft_64k.py | 47 ++- src/lilo/configs/qwen36_35b_a3b_fft_64k.py | 19 +- src/lilo/configs/qwen38_27b_lora_128k.py | 39 ++- src/lilo/configs/qwen38_27b_lora_16k.py | 53 ++-- src/lilo/configs/qwen38_27b_lora_256k.py | 38 +++ src/lilo/configs/qwen38_27b_lora_64k.py | 32 +- src/lilo/configuration.py | 105 +++++++ src/lilo/deployment_cli.py | 50 ++-- src/lilo/deployments.py | 242 ++++++--------- src/lilo/inference/native_sglang.py | 33 -- src/lilo/inference/sglang.py | 30 ++ src/lilo/inference/sglang_deployment.py | 12 +- src/lilo/providers/modal/app.py | 13 +- .../qwen3_8_27b_miles_lora_256k.py | 223 -------------- src/lilo/providers/modal/deployment_apps.py | 283 ++++++++---------- .../providers/modal/deployment_pool_app.py | 6 +- .../providers/modal/deployment_records.py | 47 +++ .../providers/modal/deployment_worker_app.py | 3 +- src/lilo/providers/modal/fft_pool.py | 10 +- src/lilo/providers/modal/lora_pool.py | 8 +- ... => test_megatron_constructor_settings.py} | 24 +- tests/providers/test_deployment_apps.py | 171 +++++++++-- tests/providers/test_deployment_e2e_helper.py | 8 +- tests/providers/test_deployment_presets.py | 12 +- tests/providers/test_sglang_entrypoint.py | 48 +++ tests/test_deployment_cli.py | 19 +- tests/test_deployments.py | 244 +++++++-------- 49 files changed, 1336 insertions(+), 1410 deletions(-) create mode 100644 src/lilo/backends/megatron_runtime/common/settings.py rename src/lilo/{argparse_config.py => backends/miles_arguments.py} (93%) create mode 100644 src/lilo/configs/qwen38_27b_lora_256k.py create mode 100644 src/lilo/configuration.py delete mode 100644 src/lilo/inference/native_sglang.py create mode 100644 src/lilo/inference/sglang.py delete mode 100644 src/lilo/providers/modal/definitions/qwen3_8_27b_miles_lora_256k.py create mode 100644 src/lilo/providers/modal/deployment_records.py rename tests/backends/{test_native_megatron_config.py => test_megatron_constructor_settings.py} (80%) create mode 100644 tests/providers/test_sglang_entrypoint.py diff --git a/docs/deployment-configs.md b/docs/deployment-configs.md index 01f4e88..67f4bdb 100644 --- a/docs/deployment-configs.md +++ b/docs/deployment-configs.md @@ -1,213 +1,137 @@ -# Python deployment configs - -Each deployment is a Python file exporting a `Config` class that inherits from `BaseConfig`. The class contains model, trainer, inference, routing and lifecycle settings. App names, environments and worker-code releases are managed by the deployment command. There is no YAML loader or model catalog. - -The layout follows the Python recipe approach used by [training-gym](https://github.com/modal-labs/training-gym) and the [multinode training guide](https://github.com/modal-labs/multinode-training-guide/blob/main/nemo-rl/configs/llama3_1_8b_math_2node.py). Only the outer BaseConfig is a dataclass; its sections are ordinary dictionaries. Config files need no decorators or default factories. Importing either project is not required. - -## Define a deployment - -Start with an example: - -```bash -lilo config init --preset qwen35-9b-lora-16k > src/lilo/configs/my_model.py -``` - -The generated file imports a packaged config and subclasses it. Customize it using ordinary Python: - -```python -from lilo.configs.qwen35_9b_lora_16k import Config as ParentConfig - - -class Config(ParentConfig): - name = "my-9b-64k" - overrides = { - "model.max_context_length": 65536, - "trainer.resources.gpu": "H200:8", - "trainer.config.tensor_model_parallel_size": 8, - "trainer.config.max_tokens_per_gpu": 65536, - "inference.scaling.max_replicas": 6, - } -``` - -New deployments can inherit directly from `BaseConfig` and declare ordinary class defaults such as `model = {"id": "...", "max_context_length": 16384}` and `trainer = {"resources": {"gpu": "H100:4"}, "config": {...}}`. The [16K example](../src/lilo/configs/qwen35_9b_lora_16k.py) shows the complete structure. No `@dataclass`, `field(default_factory=...)`, or `__post_init__` is needed in config files. - -`BaseConfig` copies defaults for each instance, then applies each parent's overrides before its child's. Dotted paths select a section and traverse its dictionary keys. A value replaces the selected field/key, including whole lists and dictionaries; other keys remain unchanged. Native backend dictionaries can receive new option names. Unknown top-level sections and missing intermediate paths raise an error naming the override. Section keys are open; their consumers read the options they need. Constructor keywords, when supplied, replace top-level fields last. - -Omitted orchestration settings come from one defaults dictionary in `deployments.py`. Backend options have no schema there. Replacing an entire section fills its omitted orchestration defaults; use dotted overrides to retain the parent’s other settings. - -Copying happens inside the base class, so edits to a config's nested lists or dictionaries do not change its parent, another instance, or the class's override dictionary. Only the resulting settings enter the saved deployment record; workers do not apply inheritance again. - -Python imports provide reuse. There is no YAML loader or `extends` key. Config files execute as Python when loaded; keep provisioning and training calls outside them. Sibling imports are available while loading a file. +# Model infrastructure configs + +A config file exports a config object made from the components in [configuration.py](../src/lilo/configuration.py). These are frozen, validated dataclasses. Unknown fields, invalid types, negative capacities and inconsistent scaling limits fail at construction. Backend option dictionaries and environment dictionaries remain open. + +~~~python +from lilo.configuration import Compute, Deployment, Inference, Model, Trainer + +config = Deployment( + name="my-9b", + model=Model(id="Qwen/Qwen3.5-9B-Base", max_context_length=16384), + trainer=Trainer( + compute=Compute(gpu="H100", gpus_per_node=4, cpu=16, memory_mib=65536), + max_clients_per_instance=6, + config={ + "model_type": "qwen3.5-9B", + "tensor_model_parallel_size": 4, + "max_lora_slots": 6, + "max_lora_rank": 32, + }, + ), + inference=Inference( + compute=Compute(gpu="H200"), + max_replicas=8, + startup_timeout_s=1200, + config={"mem_fraction_static": 0.8}, + ), +) +~~~ -All 14 previous presets are available under [`src/lilo/configs/`](../src/lilo/configs), with the same model, resource and backend settings. The [64K config](../src/lilo/configs/qwen35_9b_lora_64k.py) inherits from the 16K config. There is no separate `deployments/` wrapper directory. +See the [9B LoRA example](../src/lilo/configs/qwen35_9b_lora_16k.py) and [4B FFT example](../src/lilo/configs/qwen35_4b_fft_64k.py) for complete configurations. No model revision is required; deployment resolves Hugging Face main to an exact commit and records it. An explicit Model(revision=...) is optional. -No revision is required in a config. The CLI resolves the model's Hugging Face `main` branch automatically and saves the exact commit in the deployment record. An explicit `model.revision` remains optional for users who need particular weights; none of the built-in examples specify one. The example defaults therefore no longer pin the weights used in the historical GPU checks. +## Composition -## Deploy the complete active set +Use standard Python composition and dataclasses.replace: -Use Python 3.12, matching the serialized trainer and inference images: +~~~python +from dataclasses import replace +from lilo.configs.qwen35_9b_lora_16k import config as base -```bash -lilo config validate src/lilo/configs/my_model.py -lilo config resolve src/lilo/configs/my_model.py --output /tmp/deployment.json -lilo deploy src/lilo/configs/model_a.py src/lilo/configs/model_b.py -``` +config = replace( + base, + name="my-9b-more-memory", + trainer=replace( + base.trainer, + compute=replace(base.trainer.compute, memory_mib=98304), + ), +) +~~~ -`validate` loads the Python classes and checks routing and shared lifecycle settings. It does not start backend libraries or prove that the model fits in GPU memory. `resolve` additionally pins model revisions and emits the deployment records as JSON; it does not provision compute. +This retains the other trainer settings. There is no inheritance interpreter, dotted-path override syntax or implicit dictionary merge. Backend dictionaries can be composed explicitly with {**base.trainer.config, "max_tokens_per_gpu": 8192}. Treat configs as values; construct a variant instead of editing an imported object's dictionaries. -For a checked-in list, edit [`scripts/deploy_models.sh`](../scripts/deploy_models.sh). Add a Python config file and its path to the `deployment_files` array, then run: +## Ownership and validation -```bash -./scripts/deploy_models.sh -``` +| Setting | Owner and behavior | +| --- | --- | +| Compute | GPU type, GPUs per node, nodes, CPU and memory. Unknown keys such as memroy_mib are rejected. | +| Trainer | Maximum instances/clients, publication concurrency and function timeout. Trainers start on demand; there is no min_instances field. | +| Inference | Replica scaling and startup_timeout_s, passed to the Modal server and startup health checks. There is no unused timeout_s. Each replica uses one node. | +| trainer.config | Existing MilesBackendConfig or EngineModelConfig fields, plus their explicit extra-option dictionaries. | +| inference.config | SGLang ServerArgs fields. Lilo reserves paths, context, topology and adapter settings that must agree with its own configuration. | -Supply every configuration that should remain available to new clients. Omitted configurations are retained for existing jobs but removed from new-client selection. All files in the list must agree on shared lifecycle settings. Frontend deployment settings are not part of `BaseConfig`. The command selects the app, environment and region: +Compute topology is configured in Compute; Miles receives actor_num_nodes and actor_num_gpus_per_node from it. Setting those again in backend options is rejected. -```bash -./scripts/deploy_models.sh --app my-lilo --env dev --region us-west -``` +Megatron's provider_overrides, optimizer_overrides and distributed_overrides may add backend fields, but may not replace Lilo-owned fields. For example, put the learning rate in optimizer={"lr": ...}; optimizer_overrides={"lr": ...} is rejected. The same settings builders are used during validation and worker construction. The provider is constructed with dataclasses.replace, without an override-by-setattr pass. -These flags are optional; existing provider defaults apply when omitted. Secret and volume names come from provider defaults and are saved as platform metadata in deployment records. Credentials remain in Modal secrets. Pin `LILO_MILES_COMMIT` when a specific Miles build is needed. +Miles has a cli_options dictionary for additional Miles arguments. Its argument conversion is isolated in [miles_arguments.py](../src/lilo/backends/miles_arguments.py), because Miles exposes an argparse interface. SGLang uses ServerArgs(**settings) directly in the worker-only [sglang.py](../src/lilo/inference/sglang.py) entrypoint. -## Code path +Backend libraries validate their own extra options when workers start. Lilo does not maintain another schema for every upstream tuning option. Frontend config imports remain CPU-only. -```text -lilo deploy config.py - deployment_cli.main() - compile_configs() - deployments.load() → execute config.py → Config() - resolve model tag to a Hugging Face commit - DeploymentRecord.create() → copy settings and compute config hash - deploy() - retain existing job configurations - deploy new trainer/inference apps; skip existing apps - save/pass manifest JSON - modal deploy -m lilo.providers.modal.app -``` +## Resolve and launch -[`load()`](../src/lilo/deployments.py) executes the file and instantiates its exported `Config` subclass. The returned dataclass goes directly to the orchestration code. There is no dictionary-to-config conversion on this path. +~~~text +load(config.py) → config: Deployment + → validate typed compute/scaling/model settings + → resolve model commit + → resolve_backend_settings(config, asset_path) + → save DeploymentRecord with trainer_settings and inference_settings + → deploy independent worker apps + → update frontend references +~~~ -`DeploymentRecord` adds the pinned Miles revision, configuration hash and active status. JSON is used only to store records and pass them to other processes. Pydantic reconstructs BaseConfig with its dictionary sections when reading those records; it does not import or run the user's config file in a GPU worker. Records preserve the full computed settings rather than a reference to the original Python file. +The launcher consumes the saved settings. It does not reparse backend configuration or add another set of backend defaults. JSON decoding in a worker reconstructs the saved record, without importing the author's config file. | File | Responsibility | | --- | --- | -| [`deployments.py`](../src/lilo/deployments.py) | BaseConfig defaults and overrides, Python file loader, saved record and shared-frontend checks | -| [`deployment_cli.py`](../src/lilo/deployment_cli.py) | Revision lookup, manifest updates and Modal deployment | -| [`deployment_apps.py`](../src/lilo/providers/modal/deployment_apps.py) | Independent trainer apps, inference provisioners, server classes and process startup | -| [`app.py`](../src/lilo/providers/modal/app.py) | Shared frontend referencing deployed trainer functions | -| [`deployment_worker_app.py`](../src/lilo/providers/modal/deployment_worker_app.py) | Entrypoint for deploying one trainer or inference provisioner | -| [`deployment_pool_app.py`](../src/lilo/providers/modal/deployment_pool_app.py) | Construct an inference pool from its saved record | -| [`backends/deployment.py`](../src/lilo/backends/deployment.py) | Select the backend configuration reader | - -`build_trainer_app(record)` configures resources, secrets, storage and limits. The CLI deploys it as its own app. The frontend looks up its `trainer` function by app name. When that function starts, `run_trainer()` obtains backend settings and passes them as `LILO_BACKEND_CONFIG` to the existing Miles or Megatron executor. +| [configuration.py](../src/lilo/configuration.py) | Typed components and shared orchestration constraints | +| [deployments.py](../src/lilo/deployments.py) | Python object loader, resolved records and config hashes | +| [backends/deployment.py](../src/lilo/backends/deployment.py) | Resolve backend settings before launch | +| [megatron_runtime/common/settings.py](../src/lilo/backends/megatron_runtime/common/settings.py) | Shared Megatron ownership rules and constructor dictionaries | +| [deployment_cli.py](../src/lilo/deployment_cli.py) | Model lookup, worker release selection and deploy ordering | +| [deployment_apps.py](../src/lilo/providers/modal/deployment_apps.py) | Resource declarations and launch using resolved settings | +| [deployment_records.py](../src/lilo/providers/modal/deployment_records.py) | Read saved records and route pool provisioning | +| [deployment_worker_app.py](../src/lilo/providers/modal/deployment_worker_app.py) | Deploy one trainer or inference provisioner | -Inference pools are separate apps created on demand by a deployed `provision` function. That function keeps the source used when its inference app was deployed, including when an idle pool needs to be recreated after a frontend upgrade. `build_rollout_app()` constructs a server from the saved configuration. Startup launches SGLang with native options, waits for its health endpoint, then starts the LoRA or FFT sidecar. +## Multi-node Miles -## Backend config path +[qwen38_27b_lora_256k.py](../src/lilo/configs/qwen38_27b_lora_256k.py) configures two nodes with eight H200s per node, TP2 and CP8. The trainer app uses Modal's clustered launcher and RDMA. Every rank mounts the same volumes; rank 0 starts the engine, and the other ranks join Ray using the launcher merged in #39. The driver receives the Ray address. GPU counts cannot be independently overridden through Miles options. -Each backend already has a config class used by its trainer. `trainer.config` uses that class's field names directly. The deployment code supplies the downloaded model path, GPU count and context length; it does not rename user fields. +This path has CPU construction/topology tests. This refactor has not been redeployed or tested on multiple GPU nodes. -| Backend | Config consumer | Additional backend options | -| --- | --- | --- | -| Miles | `MilesBackendConfig` in `backends/miles_config.py` | `cli_options` goes to Miles's argparse parser | -| Megatron | `EngineModelConfig` in `backends/megatron_runtime/common/config.py` | `provider_overrides`, `optimizer_overrides`, `distributed_overrides` go to their respective Megatron constructors | -| SGLang | Its own `ServerArgs` parser | The entire `inference.config` dictionary | +## Deploy and update -For Miles: +~~~bash +lilo config init --preset qwen35-9b-lora-16k > my_model.py +lilo config validate my_model.py +lilo deploy my_model.py +~~~ -```python -"config": { - "model_type": "qwen3.5-9B", - "tensor_model_parallel_size": 4, - "max_lora_slots": 6, - "max_lora_rank": 32, - "cli_options": { - "recompute_granularity": "full", - "recompute_method": "uniform", - "recompute_num_layers": 1, - }, -} -``` +Validation checks orchestration and integration constraints without loading GPU libraries or provisioning compute. Backend option support and GPU memory capacity still require worker startup. -The path is `trainer.config → MilesBackendConfig(**settings) → miles_arguments() → Miles parser`. The existing `miles_arguments()` method translates Lilo's backend settings into Miles flags once. `cli_options` contains additional upstream flags using their argparse destination names. For a model without a Miles architecture preset, set `model_type=""` and supply its architecture options there. +The checked-in [deploy_models.sh](../scripts/deploy_models.sh) lists the complete active config set. Add a config path there, then run it. The deployment command owns frontend selection and worker-code updates: -For Megatron: - -```python -"config": { - "tensor_model_parallel_size": 2, - "context_parallel_size": 2, - "micro_batch_size": 1, - "max_tokens_per_microbatch": 65536, - "provider_overrides": {"recompute_granularity": "full"}, - "optimizer": {"lr": 0.0001, "min_lr": 0.0001}, - "optimizer_overrides": {"adam_eps": 1e-8}, -} -``` - -The path is `trainer.config → parse_backend_config() → EngineModelConfig`. The existing reader constructs the nested `OptimizerConfig` from `optimizer`. There is no `runtime` section or deployment-specific field map. `optimizer_overrides` explicitly supplies additional Megatron constructor options; the deployment code does not split optimizer fields by name. Tinker optimizer steps still apply the client's Adam parameters. - -“Training loop” refers to the code that performs forward/backward passes and optimizer steps. It is not a separate config category. - -[`argparse_config.py`](../src/lilo/argparse_config.py) is used only at the Miles/SGLang process boundary. It appends configured values after preset arguments, leaving type conversion, choices and argument counts to the backend's parser. Boolean flags need special treatment: a `store_true` flag cannot express `False`, so the helper removes that preset flag and sets its default explicitly. Unsupported custom/repeated actions fail with an error. - -[`config_validation.py`](../src/lilo/config_validation.py) only rejects options that would overwrite settings Lilo manages, such as checkpoint paths or parallelism already used by its collectives. Backend-specific modules retain those checks and resource/admission checks. They do not define another configuration schema. Backend support and memory capacity are still checked by the backend when workers start. - -## Routing and updates - -The frontend URL is shared across models. Clients select the model through `base_model`: - -```python -service = tinker.ServiceClient(base_url=lilo_url, api_key=api_key) -training = service.create_lora_training_client( - base_model="Qwen/Qwen3.5-9B-Base", rank=32, -) -``` - -If multiple configurations serve the same model and training mode, select one with `routing.default=True`. Otherwise client creation reports the ambiguity. `routing.sampling_default` resolves sampling-only selection across LoRA and FFT configurations. Training-derived sampling remains attached to the client's saved configuration. - -Every client stores its selected definition ID. Changing routing defaults affects new clients. Changing compute or backend settings creates a new configuration ID, and old configurations remain registered for existing jobs and checkpoints. The CLI serializes applies and retains interrupted deployments for recovery. - -## Hashes and independent updates - -There is no source fingerprint and no check that a config matches the current Lilo checkout. Hashes identify settings; they do not certify model support or successful training. - -| Identifier | Inputs | Used for | -| --- | --- | --- | -| `generation` | Computed config except routing defaults, platform settings, saved worker releases and Miles commit | Saved job/checkpoint configuration and routing | -| `trainer_hash` | Deployment name, model, trainer section, platform settings, trainer release and Miles commit | Independently deployed trainer app name | -| `inference_hash` | Deployment name, model, inference section, platform settings, inference release, adapter rank and target modules | Independently deployed inference provisioner app name | -| Asset hash | Model repository and resolved model commit | Download directory | - -The hashes use SHA-256 over sorted JSON. Trainer/inference app names use the first 24 hex characters. Definition IDs use the first 16 characters of `generation`. Code is retained by the deployed Modal apps, rather than reconstructed from a source fingerprint. - -Worker-code versions are not config fields. The CLI keeps the previous release IDs in deployment records. To deploy changed code for selected workers: - -```bash +~~~bash +./scripts/deploy_models.sh --app my-lilo --env dev --region us-west ./scripts/deploy_models.sh --refresh-trainer qwen35-9b-lora-16k ./scripts/deploy_models.sh --refresh-inference qwen35-9b-lora-16k -``` - -The CLI generates a release ID for each requested update; users do not specify or maintain it. Subsequent ordinary deploys retain that release. Refresh flags can be repeated for multiple config names. Editing source alone does not update existing workers. The frontend itself is redeployed on every apply. +~~~ -Examples: +These settings are not model-config fields. Secret/volume names come from provider defaults. Credentials remain in Modal secrets. -- Change inference concurrency: deploy a new inference provisioner; keep the existing trainer app. -- Change trainer batch settings or request a trainer code refresh: deploy a new trainer app; keep the inference provisioner. -- Change adapter rank or target modules: update both because inference must load the changed adapters. -- Change routing defaults: keep both worker apps. -- Change one Miles configuration: other Miles configurations and Megatron apps remain deployed as they were. +One frontend serves all models through Tinker's base_model. Routing(default=True) selects among multiple training configurations for one model; sampling_default=True selects a sampling configuration when LoRA/FFT configurations coexist. -The CLI records each successfully deployed worker app before updating the frontend. Retries skip completed apps and recover a worker deployment that succeeded before its registry write. Retained configurations reference old apps; the CLI does not rebuild those apps with new source. Old apps are retained, with GPU trainers scaling to zero when idle. Automatic deletion of unused worker app definitions is not implemented. +## Hashes and update isolation -Trainer capacity is enforced per saved definition by control-plane admission and reconciliation. The reusable Modal trainer function has no additional global container cap, so retained jobs do not prevent new definitions from starting. Old and new definitions can consume their configured capacity simultaneously during an update. +Hashes identify settings, not compatibility with a source checkout: -Pools remain scoped to the full deployment definition (and FFT job/version). A trainer update can therefore cause a new job to obtain a separate pool even when it uses the same inference provisioner. Existing pools are not redeployed. +- generation identifies the saved configuration, resolved backend settings, platform settings and code releases. Routing defaults are excluded. +- trainer_hash includes trainer/model/platform settings, resolved trainer settings, Miles commit and its recorded code release. +- inference_hash includes inference/model/platform settings, resolved inference settings and its recorded code release. -Backend startup checks and optional smoke tests remain separate from these identifiers. Worker/frontend protocol changes still require compatible APIs or an explicit migration; removing the source fingerprint does not guarantee arbitrary old and new versions interoperate. +SHA-256 hashes sorted JSON. Trainer and inference app names use the first 24 hex characters; definition IDs use the first 16 characters of generation. There is no whole-source fingerprint. -This changes the draft's deployment-record format and replaces its earlier shared-app trainers. Existing deployments from that draft need a fresh frontend/registry or an explicit migration; this change does not silently convert running shared-app trainers. +An inference-only change reuses the trainer app. A trainer batch-setting change reuses the inference provisioner. Changes to adapter rank/targets update both. Code-only upgrades use the refresh flags; later ordinary deploys retain those releases. -Existing resource names (`lilo-yaml`, `*-yaml-deployments`, and the `yaml_` definition prefix) are retained for naming continuity. They no longer indicate a YAML ingestion path. PyYAML is not a direct Lilo dependency; other installed libraries may depend on it. +Old jobs keep their recorded worker apps. Deployment retries reuse workers already created successfully. Trainer limits are enforced per saved definition; old and new definitions may use their configured capacity simultaneously while old jobs finish. Pools remain associated with full job configurations, so new jobs may get separate pools even if they share an inference provisioner. -See [validation results](deployment-validation.md) for CPU coverage and the earlier GPU smoke tests. The Python-config migration has not been redeployed. +Old worker app definitions are retained; automatic cleanup is not implemented. Changes to the frontend/worker protocol still need deliberate compatibility handling. Earlier draft manifests need migration or a fresh frontend/registry. See [validation history](deployment-validation.md) for the distinction between current CPU checks and historical GPU runs. diff --git a/docs/deployment-validation.md b/docs/deployment-validation.md index a8aa5d7..4fad77e 100644 --- a/docs/deployment-validation.md +++ b/docs/deployment-validation.md @@ -2,6 +2,20 @@ These checks exercise PR #55. The current implementation deploys trainers and inference provisioners independently; the shared frontend references them by app name. Earlier sections record validation of the previous shared-app implementation. +## Current: typed config composition and one backend resolution + +Configs export a `Deployment` built from validated dataclasses. Ordinary `dataclasses.replace` composes variants; custom inheritance, dotted overrides, and recursive defaults have been removed. Lilo-owned fields reject unknown keys. Backend tuning remains in explicit dictionaries, with duplicates of Lilo-managed settings rejected before constructor calls. + +Deployment records save complete trainer and inference settings. Workers consume those settings without resolving them again. SGLang receives `ServerArgs(**settings)` directly; the argparse adapter is now Miles-only. Megatron provider, optimizer, and distributed constructors receive one merged settings dictionary, without attribute-patching loops. + +Trainer timeout, CPU, memory, inference startup timeout, and replica scaling are wired through. Unsupported trainer minimum instances and generic inference resource timeout fields are rejected. The merged multi-node Miles launcher is integrated with `Compute.nodes`; a two-node Qwen3.8-27B 256K config replaces the legacy catalog definition. + +Validation: **651 CPU tests passed, 1 skipped**. Regression coverage includes misspelled orchestration fields, unsupported settings, optimizer/provider/distributed override collisions, all 15 example configs, saved-record round trips, direct SGLang construction, compute-setting propagation, independent worker updates, and multi-node launcher wiring. No apps were redeployed. GPU backend startup and the new multi-node config have not been live-tested in this revision. + +## Historical validation + +The sections below describe earlier revisions, including APIs that have since been removed. Their live results do not validate the current implementation. + ## Direct backend configuration Removed the deployment-specific Miles field-renaming table and Megatron `runtime/provider/optimizer/distributed` schema. Configs now use the existing `MilesBackendConfig` and `EngineModelConfig` field names. Megatron's existing config reader constructs its nested optimizer; extra Megatron constructor settings are explicit `*_overrides` dictionaries instead of being split by field name. diff --git a/scripts/definition_smoke.py b/scripts/definition_smoke.py index 97d4c42..d9c2725 100644 --- a/scripts/definition_smoke.py +++ b/scripts/definition_smoke.py @@ -15,11 +15,8 @@ --definition-id qwen3_5_9b_miles_lora_16k \ --definition-id qwen3_5_4b_full_64k -Any id from ``lilo.providers.modal.definitions`` works; the client type -(full or LoRA) and the LoRA target flags are read from the definition module so -the request matches what the deployment accepts. The definition id is passed as -``base_model`` so non-cataloged definitions can be targeted directly. -Definitions run sequentially unless ``--parallel`` is set. +IDs come from the deployed manifest selected by --app / --env. Use --list to +see active configurations. No local model catalog or GPU worker imports are needed. """ from __future__ import annotations @@ -34,15 +31,16 @@ from pathlib import Path from typing import Any +import modal import httpx import tinker from tinker import types from lilo.backends.miles_config import lora_target_flags from lilo.client import create_full_training_client -from lilo.providers.modal.app import DEFINITIONS, module_for +from lilo.deployments import DeploymentRecord +from lilo.providers.modal.deployment_apps import definition_from_spec -DEFAULT_DEFINITION = "qwen3_8_27b_miles_lora_64k" TIMEOUT = 60 * 60 PROMPT = "Question: What is two plus two?\nAnswer:" @@ -172,9 +170,11 @@ def _create_training_client( if definition.PARAMETERIZATION == "full": training = create_full_training_client(service, definition_id) return training, {"parameterization": "full"} - train_attn, train_mlp, train_unembed = lora_target_flags(definition.TARGET_MODULES) + train_attn, train_mlp, train_unembed = lora_target_flags( + definition.RESOLVED.trainer_settings["miles"]["target_modules"] + ) if rank is None: - rank = min(16, definition.MAX_LORA_RANK) + rank = min(16, definition.RESOLVED.trainer_settings["miles"]["max_lora_rank"]) training = service.create_lora_training_client( base_model=definition_id, rank=rank, @@ -192,13 +192,14 @@ def _create_training_client( def _run_definition( - definition_id: str, + definition: Any, *, base_url: str, api_key: str, rank: int | None, max_tokens: int, ) -> dict: + definition_id = definition.DEFINITION_ID report: dict[str, Any] = { "definition_id": definition_id, "status": "running", @@ -207,7 +208,6 @@ def _run_definition( } training = None try: - definition = module_for(definition_id) service = tinker.ServiceClient(base_url=base_url, api_key=api_key) started = time.perf_counter() training, spec = _create_training_client(service, definition, rank) @@ -272,14 +272,12 @@ def _run_definition( def main() -> None: parser = argparse.ArgumentParser(description=__doc__.split("\n\n")[0]) - known = [definition.DEFINITION_ID for definition in DEFINITIONS] parser.add_argument( "--definition-id", action="append", dest="definition_ids", - choices=known, metavar="ID", - help=f"engine definition id (repeatable); default {DEFAULT_DEFINITION}", + help="engine definition id (repeatable); defaults to the first active deployment", ) parser.add_argument( "--list", action="store_true", help="print known definitions and exit" @@ -297,12 +295,26 @@ def main() -> None: type=Path, default=Path("scripts/results/definition_smoke.json"), ) + parser.add_argument("--app", default="lilo-yaml") + parser.add_argument("--env") args = parser.parse_args() + rows = modal.Dict.from_name( + f"{args.app}-yaml-deployments", environment_name=args.env + ).get("manifest", []) + definitions = { + row.definition_id: definition_from_spec(row, register_trainer=False) + for data in rows + if (row := DeploymentRecord.model_validate(data)).active + } + if not definitions: + parser.error("no active deployments in the saved manifest") + if set(args.definition_ids or []) - definitions.keys(): + parser.error("unknown deployment ID; use --list to see deployed configurations") if args.list: - for definition in DEFINITIONS: + for definition in definitions.values(): print( f"{definition.DEFINITION_ID:40} {definition.PARAMETERIZATION:5} " - f"{definition.GPUS}x{definition.GPU_TYPE} " + f"{definition.RESOLVED.spec.trainer.compute.nodes} nodes x {definition.RESOLVED.spec.trainer.compute.modal_gpu} " f"ctx={definition.MAX_CONTEXT_LENGTH}" + ("" if definition.CATALOG_VISIBLE else " (not cataloged)") ) @@ -312,11 +324,11 @@ def main() -> None: api_key = os.environ.get("TINKER_API_KEY") if not api_key: parser.error("TINKER_API_KEY is required") - definition_ids = args.definition_ids or [DEFAULT_DEFINITION] + definition_ids = args.definition_ids or [next(iter(definitions))] def run(definition_id: str) -> dict: return _run_definition( - definition_id, + definitions[definition_id], base_url=args.base_url, api_key=api_key, rank=args.rank, diff --git a/scripts/e2e_engine_definition.py b/scripts/e2e_engine_definition.py index e400b66..553480f 100644 --- a/scripts/e2e_engine_definition.py +++ b/scripts/e2e_engine_definition.py @@ -16,7 +16,7 @@ import tinker from tinker import types -from lilo.deployments import DeploymentRecord, gpu_count +from lilo.deployments import DeploymentRecord from lilo.backends.deployment import backend_config TIMEOUT = 3 * 60 * 60 @@ -35,18 +35,18 @@ def _definition(frontend: str, name: str) -> tuple[Any, str]: ) resolved = matches[0] spec = resolved.spec - settings = backend_config(spec)[spec.trainer['backend']] + settings = backend_config(spec)[spec.trainer.backend] definition = SimpleNamespace( DEFINITION_ID=resolved.definition_id, - MODEL_NAME=spec.model['id'], - PARAMETERIZATION=spec.model['parameterization'], - MAX_CONTEXT_LENGTH=spec.model['max_context_length'], + MODEL_NAME=spec.model.id, + PARAMETERIZATION=spec.model.parameterization, + MAX_CONTEXT_LENGTH=spec.model.max_context_length, MAX_TOKENS_PER_MICROBATCH=settings.get( "max_tokens_per_microbatch", settings.get("max_tokens_per_gpu") ), MICRO_BATCH_SIZE=settings.get("micro_batch_size", 1), - GPU_TYPE=spec.trainer['resources']['gpu'].split(":")[0], - GPUS=gpu_count(spec.trainer['resources']), + GPU_TYPE=spec.trainer.compute.modal_gpu.split(":")[0], + GPUS=spec.trainer.compute.gpus_per_node, LORA_RANK=settings.get("max_lora_rank"), ) return definition, definition.PARAMETERIZATION diff --git a/src/lilo/backends/deployment.py b/src/lilo/backends/deployment.py index f2b9a98..2219f39 100644 --- a/src/lilo/backends/deployment.py +++ b/src/lilo/backends/deployment.py @@ -1,28 +1,40 @@ -"""Dispatch backend configuration to its backend-owned integration. +"""Resolve lightweight backend settings once, before creating worker apps.""" -These readers are CPU-only. Native libraries validate their options in workers. -""" +from lilo.backends.megatron_deployment import build_config as megatron_config +from lilo.backends.miles_config import MilesBackendConfig +from lilo.backends.miles_deployment import build_config as miles_config +from lilo.inference.sglang_deployment import build_config as sglang_config -from importlib import import_module - -TRAINERS = { - "miles": "lilo.backends.miles_deployment", - "megatron": "lilo.backends.megatron_deployment", -} -INFERENCE = {"sglang": "lilo.inference.sglang_deployment"} - - -def _reader(registry, backend): - try: - module = registry[backend] - except KeyError: - raise ValueError(f"unknown deployment backend: {backend}") from None - return import_module(module).build_config +TRAINERS = {"miles": miles_config, "megatron": megatron_config} def backend_config(spec, asset_path="/assets/pending"): - return _reader(TRAINERS, spec.trainer["backend"])(spec, asset_path) + return TRAINERS[spec.trainer.backend](spec, asset_path) def serving_options(spec): - return _reader(INFERENCE, spec.inference["backend"])(spec) + return sglang_config(spec) + + +def resolve_backend_settings(spec, asset_path): + trainer = backend_config(spec, asset_path) + inference = { + "context_length": spec.model.max_context_length, + "tp_size": spec.inference.compute.gpus_per_node, + "mem_fraction_static": 0.8, + "max_running_requests": 32, + "weight_loader_disable_mmap": True, + **serving_options(spec), + } + if spec.model.parameterization == "lora": + miles = MilesBackendConfig(**trainer["miles"]) + inference.update( + enable_lora=True, + max_lora_rank=miles.max_lora_rank, + lora_target_modules=list(miles.peft_target_modules), + ) + inference.setdefault("max_loaded_loras", 64) + inference.setdefault("max_loras_per_batch", 8) + else: + inference["enable_cpu_weight_cache"] = True + return trainer, inference diff --git a/src/lilo/backends/megatron_deployment.py b/src/lilo/backends/megatron_deployment.py index 532aeef..5d99550 100644 --- a/src/lilo/backends/megatron_deployment.py +++ b/src/lilo/backends/megatron_deployment.py @@ -2,66 +2,25 @@ from dataclasses import asdict -from lilo.deployments import gpu_count from lilo.config_validation import reject_managed_options -from .megatron_config import parse_backend_config -# These values also control packing, collectives, and checkpoint metadata in Lilo. -# Configure them once on EngineModelConfig so packing and the provider agree. -PROVIDER_MANAGED = { - "tensor_model_parallel_size", - "pipeline_model_parallel_size", - "virtual_pipeline_model_parallel_size", - "context_parallel_size", - "expert_model_parallel_size", - "expert_tensor_parallel_size", - "sequence_parallel", - "variable_seq_lengths", - "params_dtype", - "seq_length", -} -DISTRIBUTED_MANAGED = { - "use_distributed_optimizer", - "overlap_grad_reduce", - "overlap_param_gather", - "align_param_gather", -} -OPTIMIZER_MANAGED = { - "bf16", - "fp16", - "params_dtype", - "use_distributed_optimizer", - "overlap_param_gather", -} +from .megatron_config import parse_backend_config def build_config(spec, asset_path): trainer = spec.trainer - if spec.model["parameterization"] != "full": - raise ValueError("Megatron deployments require full parameterization") - if trainer["engine"]["max_clients_per_instance"] != 1: - raise ValueError("FFT trainers admit one client per instance") - if trainer["engine"]["sampler_persistence_concurrency"] != 1: - raise ValueError("Megatron requires sampler_persistence_concurrency: 1") - settings = trainer["config"] + settings = trainer.config reject_managed_options(settings, {"hf_checkpoint", "seq_length"}) - reject_managed_options(settings.get("provider_overrides", {}), PROVIDER_MANAGED) - reject_managed_options( - settings.get("optimizer_overrides", {}), OPTIMIZER_MANAGED | {"optimizer"} - ) - reject_managed_options( - settings.get("distributed_overrides", {}), DISTRIBUTED_MANAGED - ) config, _ = parse_backend_config( { "megatron": { **settings, "hf_checkpoint": asset_path, - "seq_length": spec.model["max_context_length"], + "seq_length": spec.model.max_context_length, } } ) if config.optimizer.optimizer != "adam": raise ValueError("Tinker optim_step requires an Adam optimizer") - config.validate(gpu_count(trainer["resources"])) + config.validate(trainer.compute.gpus_per_node) return {"megatron": asdict(config), "checkpoint_dir": "/checkpoints"} diff --git a/src/lilo/backends/megatron_runtime/common/config.py b/src/lilo/backends/megatron_runtime/common/config.py index 8ac08cf..42b52d1 100644 --- a/src/lilo/backends/megatron_runtime/common/config.py +++ b/src/lilo/backends/megatron_runtime/common/config.py @@ -3,6 +3,8 @@ import math from dataclasses import dataclass, field +from .settings import provider_settings, optimizer_settings, distributed_settings + @dataclass(frozen=True, slots=True) class OptimizerConfig: @@ -109,6 +111,9 @@ def packed_token_capacity(self) -> int: ) def validate(self, world_size: int) -> None: + provider_settings(self, None) + optimizer_settings(self, None, self.use_distributed_optimizer) + distributed_settings(self, self.use_distributed_optimizer) parallel_sizes = { "tensor_model_parallel_size": self.tensor_model_parallel_size, "pipeline_model_parallel_size": self.pipeline_model_parallel_size, diff --git a/src/lilo/backends/megatron_runtime/common/modeling.py b/src/lilo/backends/megatron_runtime/common/modeling.py index 3f16ab8..9212b07 100644 --- a/src/lilo/backends/megatron_runtime/common/modeling.py +++ b/src/lilo/backends/megatron_runtime/common/modeling.py @@ -1,5 +1,7 @@ from __future__ import annotations +from dataclasses import replace + import torch from megatron.bridge import AutoBridge from megatron.core.distributed import DistributedDataParallelConfig @@ -7,6 +9,7 @@ from megatron.core.transformer.enums import AttnBackend from .config import EngineModelConfig +from .settings import distributed_settings, optimizer_settings, provider_settings def model_provider(config: EngineModelConfig): @@ -15,27 +18,11 @@ def model_provider(config: EngineModelConfig): config.hf_checkpoint, trust_remote_code=True, ) - provider = bridge.to_megatron_provider() - provider.tensor_model_parallel_size = config.tensor_model_parallel_size - provider.pipeline_model_parallel_size = config.pipeline_model_parallel_size - provider.virtual_pipeline_model_parallel_size = ( - config.virtual_pipeline_model_parallel_size + settings = provider_settings(config, dtype) + provider = replace( + bridge.to_megatron_provider(), + **{**settings, "attention_backend": AttnBackend[config.attention_backend]}, ) - provider.context_parallel_size = config.context_parallel_size - provider.expert_model_parallel_size = config.expert_model_parallel_size - provider.expert_tensor_parallel_size = config.expert_tensor_parallel_size - provider.sequence_parallel = config.sequence_parallel - provider.variable_seq_lengths = True - if getattr(provider, "moe_token_dispatcher_type", None) == "allgather": - provider.moe_token_dispatcher_type = "alltoall" - provider.calculate_per_token_loss = config.calculate_per_token_loss - provider.attention_backend = AttnBackend[config.attention_backend] - provider.cross_entropy_loss_fusion = config.cross_entropy_loss_fusion - provider.params_dtype = dtype - for name, value in config.provider_overrides.items(): - if not hasattr(provider, name): - raise ValueError(f"unknown Megatron provider override: {name}") - setattr(provider, name, value) return bridge, provider, dtype @@ -57,11 +44,7 @@ def distributed_model( ): return provider.provide_distributed_model( ddp_config=DistributedDataParallelConfig( - use_distributed_optimizer=distributed_optimizer, - overlap_grad_reduce=config.overlap_grad_reduce, - overlap_param_gather=config.overlap_param_gather, - align_param_gather=config.align_param_gather, - **config.distributed_overrides, + **distributed_settings(config, distributed_optimizer) ), bf16=config.bf16, fp16=config.fp16, @@ -75,18 +58,5 @@ def optimizer_config( distributed_optimizer: bool, ): return MCoreOptimizerConfig( - optimizer=config.optimizer.optimizer, - lr=config.optimizer.lr, - weight_decay=config.optimizer.weight_decay, - adam_beta1=config.optimizer.adam_beta1, - adam_beta2=config.optimizer.adam_beta2, - adam_eps=config.optimizer.adam_eps, - clip_grad=config.optimizer.clip_grad, - loss_scale=config.optimizer.loss_scale, - bf16=config.bf16, - fp16=config.fp16, - params_dtype=dtype, - use_distributed_optimizer=distributed_optimizer, - overlap_param_gather=config.overlap_param_gather, - **config.optimizer_overrides, + **optimizer_settings(config, dtype, distributed_optimizer) ) diff --git a/src/lilo/backends/megatron_runtime/common/settings.py b/src/lilo/backends/megatron_runtime/common/settings.py new file mode 100644 index 0000000..6504a8a --- /dev/null +++ b/src/lilo/backends/megatron_runtime/common/settings.py @@ -0,0 +1,57 @@ +"""One ownership rule: backend extras may add fields, never replace Lilo settings.""" + +from dataclasses import asdict + +from lilo.config_validation import reject_managed_options + + +def provider_settings(config, dtype): + owned = { + "tensor_model_parallel_size": config.tensor_model_parallel_size, + "pipeline_model_parallel_size": config.pipeline_model_parallel_size, + "virtual_pipeline_model_parallel_size": config.virtual_pipeline_model_parallel_size, + "context_parallel_size": config.context_parallel_size, + "expert_model_parallel_size": config.expert_model_parallel_size, + "expert_tensor_parallel_size": config.expert_tensor_parallel_size, + "sequence_parallel": config.sequence_parallel, + "variable_seq_lengths": True, + "calculate_per_token_loss": config.calculate_per_token_loss, + "attention_backend": config.attention_backend, + "cross_entropy_loss_fusion": config.cross_entropy_loss_fusion, + "params_dtype": dtype, + } + reject_managed_options(config.provider_overrides, owned.keys() | {"seq_length"}) + return {**owned, **config.provider_overrides} + + +def optimizer_settings(config, dtype, distributed_optimizer): + owned = { + "optimizer": config.optimizer.optimizer, + "lr": config.optimizer.lr, + "weight_decay": config.optimizer.weight_decay, + "adam_beta1": config.optimizer.adam_beta1, + "adam_beta2": config.optimizer.adam_beta2, + "adam_eps": config.optimizer.adam_eps, + "clip_grad": config.optimizer.clip_grad, + "loss_scale": config.optimizer.loss_scale, + "bf16": config.bf16, + "fp16": config.fp16, + "params_dtype": dtype, + "use_distributed_optimizer": distributed_optimizer, + "overlap_param_gather": config.overlap_param_gather, + } + reject_managed_options( + config.optimizer_overrides, owned.keys() | asdict(config.optimizer).keys() + ) + return {**owned, **config.optimizer_overrides} + + +def distributed_settings(config, distributed_optimizer): + owned = { + "use_distributed_optimizer": distributed_optimizer, + "overlap_grad_reduce": config.overlap_grad_reduce, + "overlap_param_gather": config.overlap_param_gather, + "align_param_gather": config.align_param_gather, + } + reject_managed_options(config.distributed_overrides, owned.keys()) + return {**owned, **config.distributed_overrides} diff --git a/src/lilo/argparse_config.py b/src/lilo/backends/miles_arguments.py similarity index 93% rename from src/lilo/argparse_config.py rename to src/lilo/backends/miles_arguments.py index fa8a633..1a8af52 100644 --- a/src/lilo/argparse_config.py +++ b/src/lilo/backends/miles_arguments.py @@ -1,6 +1,6 @@ -"""Translate config dictionaries into arguments for Miles and SGLang. +"""Translate config dictionaries into arguments for Miles. -Those backends expose argparse parsers. Append configured values after their +Miles exposes an argparse parser. Append configured values after their preset arguments so argparse itself handles types, choices and required options. Boolean flags need defaults because a store_true flag cannot express False. """ diff --git a/src/lilo/backends/miles_deployment.py b/src/lilo/backends/miles_deployment.py index e5e1e5d..0d05a55 100644 --- a/src/lilo/backends/miles_deployment.py +++ b/src/lilo/backends/miles_deployment.py @@ -2,8 +2,8 @@ from dataclasses import asdict -from lilo.deployments import gpu_count from lilo.config_validation import reject_managed_options + from .miles_config import MilesBackendConfig MILES_MANAGED = { @@ -48,17 +48,17 @@ def build_config(spec, asset_path): """Pass trainer.config directly to MilesBackendConfig.""" trainer = spec.trainer - if spec.model["parameterization"] != "lora": - raise ValueError("Miles requires lora parameterization") - settings = trainer["config"] + settings = trainer.config reject_managed_options( - settings, {"hf_checkpoint", "actor_num_gpus_per_node", "extra_args"} + settings, + {"hf_checkpoint", "actor_num_gpus_per_node", "actor_num_nodes", "extra_args"}, ) reject_managed_options(settings.get("cli_options", {}), MILES_MANAGED) config = MilesBackendConfig( hf_checkpoint=asset_path, - actor_num_gpus_per_node=gpu_count(trainer["resources"]), - extra_args=("--seq-length", str(spec.model["max_context_length"])), + actor_num_gpus_per_node=trainer.compute.gpus_per_node, + actor_num_nodes=trainer.compute.nodes, + extra_args=("--seq-length", str(spec.model.max_context_length)), **settings, ) config.validate() @@ -66,6 +66,6 @@ def build_config(spec, asset_path): config.expert_model_parallel_size * config.expert_tensor_parallel_size ): raise ValueError("expert parallel sizes must divide the trainer GPU allocation") - if trainer["engine"]["max_clients_per_instance"] > config.max_lora_slots: + if trainer.max_clients_per_instance > config.max_lora_slots: raise ValueError("max_clients_per_instance exceeds max_lora_slots") return {"miles": asdict(config), "checkpoint_dir": "/checkpoints"} diff --git a/src/lilo/backends/miles_runtime/runtime.py b/src/lilo/backends/miles_runtime/runtime.py index 2c09cbd..be2289d 100644 --- a/src/lilo/backends/miles_runtime/runtime.py +++ b/src/lilo/backends/miles_runtime/runtime.py @@ -14,6 +14,7 @@ from typing import Any from lilo.backends.miles_config import MilesBackendConfig +from lilo.backends.miles_arguments import apply_config_overrides from lilo.errors import BackendFailed @@ -200,7 +201,6 @@ async def _start(self) -> None: ) with _temporary_argv([*architecture, *self.config.miles_arguments()]): if self.config.cli_options: - from lilo.argparse_config import apply_config_overrides def configure(parser): apply_config_overrides(parser, self.config.cli_options, sys.argv) diff --git a/src/lilo/configs/__init__.py b/src/lilo/configs/__init__.py index 4c40641..ad974da 100644 --- a/src/lilo/configs/__init__.py +++ b/src/lilo/configs/__init__.py @@ -1 +1 @@ -"""Example deployment dataclasses; import and subclass any Config to customize it.""" +"""Example infrastructure objects; compose variants with dataclasses.replace.""" diff --git a/src/lilo/configs/qwen35_35b_a3b_fft_64k.py b/src/lilo/configs/qwen35_35b_a3b_fft_64k.py index 43a29db..e5dfc23 100644 --- a/src/lilo/configs/qwen35_35b_a3b_fft_64k.py +++ b/src/lilo/configs/qwen35_35b_a3b_fft_64k.py @@ -1,23 +1,14 @@ -from lilo.deployments import BaseConfig +from lilo.configuration import Compute, Deployment, Inference, Model, Routing, Trainer - -class Config(BaseConfig): - name = "qwen35-35b-a3b-fft-64k" - model = { - "id": "Qwen/Qwen3.5-35B-A3B", - "parameterization": "full", - "max_context_length": 65536, - } - routing = {"default": True} - trainer = { - "backend": "megatron", - "resources": {"gpu": "H200:8"}, - "engine": {"max_clients_per_instance": 1, "sampler_persistence_concurrency": 1}, - "env": { - "PYTORCH_CUDA_ALLOC_CONF": "expandable_segments:True", - "TORCHINDUCTOR_COMPILE_THREADS": "1", - }, - "config": { +config = Deployment( + name="qwen35-35b-a3b-fft-64k", + model=Model( + parameterization="full", id="Qwen/Qwen3.5-35B-A3B", max_context_length=65536 + ), + trainer=Trainer( + compute=Compute(gpu="H200", gpus_per_node=8), + backend="megatron", + config={ "tensor_model_parallel_size": 4, "pipeline_model_parallel_size": 1, "context_parallel_size": 2, @@ -43,11 +34,15 @@ class Config(BaseConfig): }, "optimizer": {"optimizer": "adam", "lr": 0.0001, "min_lr": 0.0001}, }, - } - inference = { - "resources": {"gpu": "H200:4"}, - "scaling": {"min_replicas": 0, "max_replicas": 8, "target_concurrency": 16}, - "config": { + env={ + "PYTORCH_CUDA_ALLOC_CONF": "expandable_segments:True", + "TORCHINDUCTOR_COMPILE_THREADS": "1", + }, + sampler_persistence_concurrency=1, + ), + inference=Inference( + compute=Compute(gpu="H200", gpus_per_node=4), + config={ "tp_size": 4, "ep_size": 4, "mem_fraction_static": 0.9, @@ -57,4 +52,6 @@ class Config(BaseConfig): "dp_size": 4, "enable_dp_attention": True, }, - } + ), + routing=Routing(default=True), +) diff --git a/src/lilo/configs/qwen35_4b_fft_64k.py b/src/lilo/configs/qwen35_4b_fft_64k.py index b6a8025..ce2d827 100644 --- a/src/lilo/configs/qwen35_4b_fft_64k.py +++ b/src/lilo/configs/qwen35_4b_fft_64k.py @@ -1,19 +1,14 @@ -from lilo.deployments import BaseConfig +from lilo.configuration import Compute, Deployment, Inference, Model, Routing, Trainer - -class Config(BaseConfig): - name = "qwen35-4b-fft-64k" - model = { - "id": "Qwen/Qwen3.5-4B", - "parameterization": "full", - "max_context_length": 65536, - } - routing = {"default": True} - trainer = { - "backend": "megatron", - "resources": {"gpu": "H100:4"}, - "engine": {"max_clients_per_instance": 1, "sampler_persistence_concurrency": 1}, - "config": { +config = Deployment( + name="qwen35-4b-fft-64k", + model=Model( + parameterization="full", id="Qwen/Qwen3.5-4B", max_context_length=65536 + ), + trainer=Trainer( + compute=Compute(gpu="H100", gpus_per_node=4), + backend="megatron", + config={ "tensor_model_parallel_size": 2, "context_parallel_size": 2, "sequence_parallel": True, @@ -30,14 +25,17 @@ class Config(BaseConfig): }, "optimizer": {"lr": 0.0001, "min_lr": 0.0001, "loss_scale": 1.0}, }, - } - inference = { - "resources": {"gpu": "H100:1"}, - "config": { + sampler_persistence_concurrency=1, + ), + inference=Inference( + compute=Compute(gpu="H100"), + config={ "tp_size": 1, "mem_fraction_static": 0.85, "max_running_requests": 32, "max_queued_requests": 4, "cpu_weight_cache_max_compile_group_gb": 16, }, - } + ), + routing=Routing(default=True), +) diff --git a/src/lilo/configs/qwen35_9b_fft_64k.py b/src/lilo/configs/qwen35_9b_fft_64k.py index aff3bc8..128dc6d 100644 --- a/src/lilo/configs/qwen35_9b_fft_64k.py +++ b/src/lilo/configs/qwen35_9b_fft_64k.py @@ -1,23 +1,14 @@ -from lilo.deployments import BaseConfig +from lilo.configuration import Compute, Deployment, Inference, Model, Routing, Trainer - -class Config(BaseConfig): - name = "qwen35-9b-fft-64k" - model = { - "id": "Qwen/Qwen3.5-9B", - "parameterization": "full", - "max_context_length": 65536, - } - routing = {"default": True} - trainer = { - "backend": "megatron", - "resources": {"gpu": "H200:4"}, - "engine": {"max_clients_per_instance": 1, "sampler_persistence_concurrency": 1}, - "env": { - "PYTORCH_CUDA_ALLOC_CONF": "expandable_segments:True", - "TORCHINDUCTOR_COMPILE_THREADS": "1", - }, - "config": { +config = Deployment( + name="qwen35-9b-fft-64k", + model=Model( + parameterization="full", id="Qwen/Qwen3.5-9B", max_context_length=65536 + ), + trainer=Trainer( + compute=Compute(gpu="H200", gpus_per_node=4), + backend="megatron", + config={ "tensor_model_parallel_size": 2, "context_parallel_size": 2, "sequence_parallel": True, @@ -34,11 +25,15 @@ class Config(BaseConfig): }, "optimizer": {"lr": 0.0001, "min_lr": 0.0001, "loss_scale": 1.0}, }, - } - inference = { - "resources": {"gpu": "H200:1"}, - "scaling": {"min_replicas": 0, "max_replicas": 8, "target_concurrency": 16}, - "config": { + env={ + "PYTORCH_CUDA_ALLOC_CONF": "expandable_segments:True", + "TORCHINDUCTOR_COMPILE_THREADS": "1", + }, + sampler_persistence_concurrency=1, + ), + inference=Inference( + compute=Compute(gpu="H200"), + config={ "tp_size": 1, "ep_size": 1, "mem_fraction_static": 0.85, @@ -46,4 +41,6 @@ class Config(BaseConfig): "max_queued_requests": 4, "cpu_weight_cache_max_compile_group_gb": 16, }, - } + ), + routing=Routing(default=True), +) diff --git a/src/lilo/configs/qwen35_9b_instruct_lora_16k.py b/src/lilo/configs/qwen35_9b_instruct_lora_16k.py index 0138a62..44406e6 100644 --- a/src/lilo/configs/qwen35_9b_instruct_lora_16k.py +++ b/src/lilo/configs/qwen35_9b_instruct_lora_16k.py @@ -1,23 +1,11 @@ -from lilo.deployments import BaseConfig +from lilo.configuration import Compute, Deployment, Inference, Model, Routing, Trainer - -class Config(BaseConfig): - name = "qwen35-9b-instruct-lora-16k" - model = { - "id": "Qwen/Qwen3.5-9B", - "parameterization": "lora", - "max_context_length": 16384, - } - routing = {"default": True} - trainer = { - "backend": "miles", - "resources": {"gpu": "H100:8"}, - "engine": {"max_clients_per_instance": 6, "sampler_persistence_concurrency": 8}, - "env": { - "PYTORCH_CUDA_ALLOC_CONF": "expandable_segments:True", - "TORCHINDUCTOR_COMPILE_THREADS": "1", - }, - "config": { +config = Deployment( + name="qwen35-9b-instruct-lora-16k", + model=Model(id="Qwen/Qwen3.5-9B", max_context_length=16384), + trainer=Trainer( + compute=Compute(gpu="H100", gpus_per_node=8), + config={ "model_type": "qwen3.5-9B", "tensor_model_parallel_size": 8, "target_modules": ["linear_qkv", "linear_proj", "linear_fc1", "linear_fc2"], @@ -31,11 +19,15 @@ class Config(BaseConfig): "recompute_num_layers": 1, }, }, - } - inference = { - "resources": {"gpu": "H200:1"}, - "scaling": {"min_replicas": 0, "max_replicas": 8, "target_concurrency": 16}, - "config": { + env={ + "PYTORCH_CUDA_ALLOC_CONF": "expandable_segments:True", + "TORCHINDUCTOR_COMPILE_THREADS": "1", + }, + max_clients_per_instance=6, + ), + inference=Inference( + compute=Compute(gpu="H200"), + config={ "tp_size": 1, "ep_size": 1, "mem_fraction_static": 0.8, @@ -43,15 +35,8 @@ class Config(BaseConfig): "max_queued_requests": 8, "max_loaded_loras": 256, "max_loras_per_batch": 8, - "lora_target_modules": [ - "q_proj", - "k_proj", - "v_proj", - "o_proj", - "gate_proj", - "up_proj", - "down_proj", - ], "schedule_policy": "lpm", }, - } + ), + routing=Routing(default=True), +) diff --git a/src/lilo/configs/qwen35_9b_instruct_lora_16k_dp2.py b/src/lilo/configs/qwen35_9b_instruct_lora_16k_dp2.py index 961ad8c..5dd2097 100644 --- a/src/lilo/configs/qwen35_9b_instruct_lora_16k_dp2.py +++ b/src/lilo/configs/qwen35_9b_instruct_lora_16k_dp2.py @@ -1,9 +1,12 @@ -from lilo.configs.qwen35_9b_instruct_lora_16k import Config as ParentConfig +from dataclasses import replace +from lilo.configs.qwen35_9b_instruct_lora_16k import config as base -class Config(ParentConfig): - name = "qwen35-9b-instruct-lora-16k-dp2" - overrides = { - "routing.default": False, - "trainer.config.tensor_model_parallel_size": 4, - } +config = replace( + base, + name="qwen35-9b-instruct-lora-16k-dp2", + trainer=replace( + base.trainer, config={**base.trainer.config, "tensor_model_parallel_size": 4} + ), + routing=replace(base.routing, default=False), +) diff --git a/src/lilo/configs/qwen35_9b_lora_16k.py b/src/lilo/configs/qwen35_9b_lora_16k.py index d02f918..e86279d 100644 --- a/src/lilo/configs/qwen35_9b_lora_16k.py +++ b/src/lilo/configs/qwen35_9b_lora_16k.py @@ -1,24 +1,11 @@ -from lilo.deployments import BaseConfig +from lilo.configuration import Compute, Deployment, Inference, Model, Routing, Trainer - -class Config(BaseConfig): - name = "qwen35-9b-lora-16k" - model = { - "id": "Qwen/Qwen3.5-9B-Base", - "parameterization": "lora", - "max_context_length": 16384, - } - routing = {"default": True} - trainer = { - "backend": "miles", - "resources": {"gpu": "H100:4", "cpu": 16, "memory_mib": 65536}, - "scaling": {"max_instances": 1}, - "engine": {"max_clients_per_instance": 6, "sampler_persistence_concurrency": 8}, - "env": { - "PYTORCH_CUDA_ALLOC_CONF": "expandable_segments:True", - "TORCHINDUCTOR_COMPILE_THREADS": "1", - }, - "config": { +config = Deployment( + name="qwen35-9b-lora-16k", + model=Model(id="Qwen/Qwen3.5-9B-Base", max_context_length=16384), + trainer=Trainer( + compute=Compute(gpu="H100", gpus_per_node=4, cpu=16, memory_mib=65536), + config={ "model_type": "qwen3.5-9B", "tensor_model_parallel_size": 4, "max_lora_slots": 6, @@ -38,12 +25,15 @@ class Config(BaseConfig): "recompute_num_layers": 1, }, }, - } - inference = { - "backend": "sglang", - "resources": {"gpu": "H200:1"}, - "scaling": {"min_replicas": 0, "max_replicas": 8, "target_concurrency": 16}, - "config": { + env={ + "PYTORCH_CUDA_ALLOC_CONF": "expandable_segments:True", + "TORCHINDUCTOR_COMPILE_THREADS": "1", + }, + max_clients_per_instance=6, + ), + inference=Inference( + compute=Compute(gpu="H200"), + config={ "tp_size": 1, "mem_fraction_static": 0.8, "max_running_requests": 32, @@ -51,4 +41,6 @@ class Config(BaseConfig): "max_loaded_loras": 64, "max_loras_per_batch": 8, }, - } + ), + routing=Routing(default=True), +) diff --git a/src/lilo/configs/qwen35_9b_lora_16k_single.py b/src/lilo/configs/qwen35_9b_lora_16k_single.py index 459e2d9..24cb691 100644 --- a/src/lilo/configs/qwen35_9b_lora_16k_single.py +++ b/src/lilo/configs/qwen35_9b_lora_16k_single.py @@ -1,6 +1,10 @@ -from lilo.configs.qwen35_9b_lora_16k import Config as ParentConfig +from dataclasses import replace +from lilo.configs.qwen35_9b_lora_16k import config as base -class Config(ParentConfig): - name = "qwen35-9b-lora-16k-single" - overrides = {"routing.default": False, "trainer.engine.max_clients_per_instance": 1} +config = replace( + base, + name="qwen35-9b-lora-16k-single", + trainer=replace(base.trainer, max_clients_per_instance=1), + routing=replace(base.routing, default=False), +) diff --git a/src/lilo/configs/qwen35_9b_lora_2k.py b/src/lilo/configs/qwen35_9b_lora_2k.py index 2ee6759..60e4a42 100644 --- a/src/lilo/configs/qwen35_9b_lora_2k.py +++ b/src/lilo/configs/qwen35_9b_lora_2k.py @@ -1,23 +1,11 @@ -from lilo.deployments import BaseConfig +from lilo.configuration import Compute, Deployment, Inference, Model, Trainer - -class Config(BaseConfig): - name = "qwen35-9b-lora-2k" - model = { - "id": "Qwen/Qwen3.5-9B-Base", - "parameterization": "lora", - "max_context_length": 2048, - } - routing = {"default": False} - trainer = { - "backend": "miles", - "resources": {"gpu": "H200:4"}, - "engine": {"max_clients_per_instance": 4, "sampler_persistence_concurrency": 8}, - "env": { - "PYTORCH_CUDA_ALLOC_CONF": "expandable_segments:True", - "TORCHINDUCTOR_COMPILE_THREADS": "1", - }, - "config": { +config = Deployment( + name="qwen35-9b-lora-2k", + model=Model(id="Qwen/Qwen3.5-9B-Base", max_context_length=2048), + trainer=Trainer( + compute=Compute(gpu="H200", gpus_per_node=4), + config={ "model_type": "qwen3.5-9B", "tensor_model_parallel_size": 4, "target_modules": [ @@ -37,11 +25,15 @@ class Config(BaseConfig): "recompute_num_layers": 1, }, }, - } - inference = { - "resources": {"gpu": "H200:1"}, - "scaling": {"min_replicas": 0, "max_replicas": 8, "target_concurrency": 16}, - "config": { + env={ + "PYTORCH_CUDA_ALLOC_CONF": "expandable_segments:True", + "TORCHINDUCTOR_COMPILE_THREADS": "1", + }, + max_clients_per_instance=4, + ), + inference=Inference( + compute=Compute(gpu="H200"), + config={ "tp_size": 1, "ep_size": 1, "mem_fraction_static": 0.8, @@ -49,16 +41,7 @@ class Config(BaseConfig): "max_queued_requests": 8, "max_loaded_loras": 32, "max_loras_per_batch": 8, - "lora_target_modules": [ - "q_proj", - "k_proj", - "v_proj", - "o_proj", - "gate_proj", - "up_proj", - "down_proj", - "lm_head", - ], "schedule_policy": "lpm", }, - } + ), +) diff --git a/src/lilo/configs/qwen35_9b_lora_64k.py b/src/lilo/configs/qwen35_9b_lora_64k.py index d7646cd..1f2886f 100644 --- a/src/lilo/configs/qwen35_9b_lora_64k.py +++ b/src/lilo/configs/qwen35_9b_lora_64k.py @@ -1,12 +1,19 @@ -from lilo.configs.qwen35_9b_lora_16k import Config as ParentConfig +from dataclasses import replace +from lilo.configs.qwen35_9b_lora_16k import config as base -class Config(ParentConfig): - name = "qwen35-9b-lora-64k" - overrides = { - "model.max_context_length": 65536, - "routing.default": False, - "trainer.resources.gpu": "H200:8", - "trainer.config.tensor_model_parallel_size": 8, - "trainer.config.max_tokens_per_gpu": 65536, - } +config = replace( + base, + name="qwen35-9b-lora-64k", + model=replace(base.model, max_context_length=65536), + trainer=replace( + base.trainer, + compute=replace(base.trainer.compute, gpu="H200", gpus_per_node=8), + config={ + **base.trainer.config, + "tensor_model_parallel_size": 8, + "max_tokens_per_gpu": 65536, + }, + ), + routing=replace(base.routing, default=False), +) diff --git a/src/lilo/configs/qwen36_27b_fft_64k.py b/src/lilo/configs/qwen36_27b_fft_64k.py index 4dd8fab..90cf2a5 100644 --- a/src/lilo/configs/qwen36_27b_fft_64k.py +++ b/src/lilo/configs/qwen36_27b_fft_64k.py @@ -1,23 +1,14 @@ -from lilo.deployments import BaseConfig +from lilo.configuration import Compute, Deployment, Inference, Model, Routing, Trainer - -class Config(BaseConfig): - name = "qwen36-27b-fft-64k" - model = { - "id": "Qwen/Qwen3.6-27B", - "parameterization": "full", - "max_context_length": 65536, - } - routing = {"default": True} - trainer = { - "backend": "megatron", - "resources": {"gpu": "H200:8"}, - "engine": {"max_clients_per_instance": 1, "sampler_persistence_concurrency": 1}, - "env": { - "PYTORCH_CUDA_ALLOC_CONF": "expandable_segments:True", - "TORCHINDUCTOR_COMPILE_THREADS": "1", - }, - "config": { +config = Deployment( + name="qwen36-27b-fft-64k", + model=Model( + parameterization="full", id="Qwen/Qwen3.6-27B", max_context_length=65536 + ), + trainer=Trainer( + compute=Compute(gpu="H200", gpus_per_node=8), + backend="megatron", + config={ "tensor_model_parallel_size": 4, "pipeline_model_parallel_size": 1, "context_parallel_size": 2, @@ -36,11 +27,15 @@ class Config(BaseConfig): }, "optimizer": {"optimizer": "adam", "lr": 0.0001, "min_lr": 0.0001}, }, - } - inference = { - "resources": {"gpu": "H200:4"}, - "scaling": {"min_replicas": 0, "max_replicas": 8, "target_concurrency": 16}, - "config": { + env={ + "PYTORCH_CUDA_ALLOC_CONF": "expandable_segments:True", + "TORCHINDUCTOR_COMPILE_THREADS": "1", + }, + sampler_persistence_concurrency=1, + ), + inference=Inference( + compute=Compute(gpu="H200", gpus_per_node=4), + config={ "tp_size": 4, "ep_size": 1, "mem_fraction_static": 0.9, @@ -48,4 +43,6 @@ class Config(BaseConfig): "max_queued_requests": 4, "cpu_weight_cache_max_compile_group_gb": 32, }, - } + ), + routing=Routing(default=True), +) diff --git a/src/lilo/configs/qwen36_35b_a3b_fft_64k.py b/src/lilo/configs/qwen36_35b_a3b_fft_64k.py index 8d29075..df04a3f 100644 --- a/src/lilo/configs/qwen36_35b_a3b_fft_64k.py +++ b/src/lilo/configs/qwen36_35b_a3b_fft_64k.py @@ -1,10 +1,13 @@ -from lilo.configs.qwen35_35b_a3b_fft_64k import Config as ParentConfig +from dataclasses import replace +from lilo.configs.qwen35_35b_a3b_fft_64k import config as base -class Config(ParentConfig): - name = "qwen36-35b-a3b-fft-64k" - overrides = { - "model.id": "Qwen/Qwen3.6-35B-A3B", - "inference.config.dp_size": 1, - "inference.config.enable_dp_attention": False, - } +config = replace( + base, + name="qwen36-35b-a3b-fft-64k", + model=replace(base.model, id="Qwen/Qwen3.6-35B-A3B"), + inference=replace( + base.inference, + config={**base.inference.config, "dp_size": 1, "enable_dp_attention": False}, + ), +) diff --git a/src/lilo/configs/qwen38_27b_lora_128k.py b/src/lilo/configs/qwen38_27b_lora_128k.py index 1549fea..c9303c2 100644 --- a/src/lilo/configs/qwen38_27b_lora_128k.py +++ b/src/lilo/configs/qwen38_27b_lora_128k.py @@ -1,17 +1,26 @@ -from lilo.configs.qwen38_27b_lora_16k import Config as ParentConfig +from dataclasses import replace +from lilo.configs.qwen38_27b_lora_16k import config as base -class Config(ParentConfig): - name = "qwen38-27b-lora-128k" - overrides = { - "model.max_context_length": 131072, - "routing.default": False, - "trainer.config.tensor_model_parallel_size": 2, - "trainer.config.context_parallel_size": 4, - "trainer.config.max_tokens_per_gpu": 32768, - "inference.resources.gpu": "H200:2", - "inference.scaling.max_replicas": 4, - "inference.scaling.target_concurrency": 4, - "inference.config.tp_size": 2, - "inference.config.max_running_requests": 8, - } +config = replace( + base, + name="qwen38-27b-lora-128k", + model=replace(base.model, max_context_length=131072), + trainer=replace( + base.trainer, + config={ + **base.trainer.config, + "tensor_model_parallel_size": 2, + "max_tokens_per_gpu": 32768, + "context_parallel_size": 4, + }, + ), + inference=replace( + base.inference, + compute=replace(base.inference.compute, gpus_per_node=2), + max_replicas=4, + target_concurrency=4, + config={**base.inference.config, "tp_size": 2, "max_running_requests": 8}, + ), + routing=replace(base.routing, default=False), +) diff --git a/src/lilo/configs/qwen38_27b_lora_16k.py b/src/lilo/configs/qwen38_27b_lora_16k.py index 6c0432e..2835e7a 100644 --- a/src/lilo/configs/qwen38_27b_lora_16k.py +++ b/src/lilo/configs/qwen38_27b_lora_16k.py @@ -1,23 +1,11 @@ -from lilo.deployments import BaseConfig +from lilo.configuration import Compute, Deployment, Inference, Model, Routing, Trainer - -class Config(BaseConfig): - name = "qwen38-27b-lora-16k" - model = { - "id": "Qwen/Qwen3.8-27B", - "parameterization": "lora", - "max_context_length": 16384, - } - routing = {"default": True} - trainer = { - "backend": "miles", - "resources": {"gpu": "H200:8"}, - "engine": {"max_clients_per_instance": 6, "sampler_persistence_concurrency": 8}, - "env": { - "PYTORCH_CUDA_ALLOC_CONF": "expandable_segments:True", - "TORCHINDUCTOR_COMPILE_THREADS": "1", - }, - "config": { +config = Deployment( + name="qwen38-27b-lora-16k", + model=Model(id="Qwen/Qwen3.8-27B", max_context_length=16384), + trainer=Trainer( + compute=Compute(gpu="H200", gpus_per_node=8), + config={ "model_type": "qwen3.8-27B", "tensor_model_parallel_size": 4, "target_modules": ["linear_qkv", "linear_proj", "linear_fc1", "linear_fc2"], @@ -31,11 +19,15 @@ class Config(BaseConfig): "recompute_num_layers": 1, }, }, - } - inference = { - "resources": {"gpu": "H200:1"}, - "scaling": {"min_replicas": 0, "max_replicas": 8, "target_concurrency": 16}, - "config": { + env={ + "PYTORCH_CUDA_ALLOC_CONF": "expandable_segments:True", + "TORCHINDUCTOR_COMPILE_THREADS": "1", + }, + max_clients_per_instance=6, + ), + inference=Inference( + compute=Compute(gpu="H200"), + config={ "tp_size": 1, "ep_size": 1, "mem_fraction_static": 0.8, @@ -43,15 +35,8 @@ class Config(BaseConfig): "max_queued_requests": 8, "max_loaded_loras": 256, "max_loras_per_batch": 8, - "lora_target_modules": [ - "q_proj", - "k_proj", - "v_proj", - "o_proj", - "gate_proj", - "up_proj", - "down_proj", - ], "schedule_policy": "lpm", }, - } + ), + routing=Routing(default=True), +) diff --git a/src/lilo/configs/qwen38_27b_lora_256k.py b/src/lilo/configs/qwen38_27b_lora_256k.py new file mode 100644 index 0000000..8aab4f7 --- /dev/null +++ b/src/lilo/configs/qwen38_27b_lora_256k.py @@ -0,0 +1,38 @@ +from dataclasses import replace + +from lilo.configs.qwen38_27b_lora_16k import config as base + +config = replace( + base, + name="qwen38-27b-lora-256k", + model=replace(base.model, max_context_length=262144), + routing=replace(base.routing, default=False), + trainer=replace( + base.trainer, + compute=replace(base.trainer.compute, nodes=2), + config={ + **base.trainer.config, + "tensor_model_parallel_size": 2, + "context_parallel_size": 8, + "max_tokens_per_gpu": 32768, + "cli_options": { + **base.trainer.config["cli_options"], + "distributed_timeout_minutes": 120, + }, + }, + ), + inference=replace( + base.inference, + compute=replace(base.inference.compute, gpus_per_node=4), + min_replicas=2, + max_replicas=2, + target_concurrency=2, + config={ + **base.inference.config, + "tp_size": 4, + "max_running_requests": 4, + "max_queued_requests": 8, + "max_loaded_loras": 256, + }, + ), +) diff --git a/src/lilo/configs/qwen38_27b_lora_64k.py b/src/lilo/configs/qwen38_27b_lora_64k.py index c65b30b..3c75488 100644 --- a/src/lilo/configs/qwen38_27b_lora_64k.py +++ b/src/lilo/configs/qwen38_27b_lora_64k.py @@ -1,13 +1,23 @@ -from lilo.configs.qwen38_27b_lora_16k import Config as ParentConfig +from dataclasses import replace +from lilo.configs.qwen38_27b_lora_16k import config as base -class Config(ParentConfig): - name = "qwen38-27b-lora-64k" - overrides = { - "model.max_context_length": 65536, - "routing.default": False, - "trainer.config.context_parallel_size": 2, - "trainer.config.max_tokens_per_gpu": 32768, - "inference.scaling.target_concurrency": 8, - "inference.config.max_running_requests": 16, - } +config = replace( + base, + name="qwen38-27b-lora-64k", + model=replace(base.model, max_context_length=65536), + trainer=replace( + base.trainer, + config={ + **base.trainer.config, + "max_tokens_per_gpu": 32768, + "context_parallel_size": 2, + }, + ), + inference=replace( + base.inference, + target_concurrency=8, + config={**base.inference.config, "max_running_requests": 16}, + ), + routing=replace(base.routing, default=False), +) diff --git a/src/lilo/configuration.py b/src/lilo/configuration.py new file mode 100644 index 0000000..b092970 --- /dev/null +++ b/src/lilo/configuration.py @@ -0,0 +1,105 @@ +"""Validated Python components for model infrastructure. + +Only backend config and environment dictionaries accept arbitrary keys. +Use dataclasses.replace to compose variants without mutating another config. +""" + +from dataclasses import field +from typing import Annotated, Literal + +from pydantic import ConfigDict, Field, StrictBool +from pydantic.dataclasses import dataclass + +PositiveInt = Annotated[int, Field(strict=True, gt=0)] +NonnegativeInt = Annotated[int, Field(strict=True, ge=0)] +NonemptyString = Annotated[str, Field(strict=True, min_length=1)] +CONFIG = ConfigDict(extra="forbid", validate_default=True) + + +@dataclass(config=CONFIG, frozen=True, kw_only=True) +class Compute: + gpu: Annotated[str, Field(pattern=r"^[A-Za-z0-9-]+$")] + gpus_per_node: PositiveInt = 1 + nodes: PositiveInt = 1 + cpu: Annotated[float, Field(gt=0)] = 8 + memory_mib: PositiveInt = 32768 + + @property + def modal_gpu(self) -> str: + return f"{self.gpu}:{self.gpus_per_node}" + + +@dataclass(config=CONFIG, frozen=True, kw_only=True) +class Model: + id: NonemptyString + max_context_length: PositiveInt + parameterization: Literal["lora", "full"] = "lora" + revision: NonemptyString = "main" + + +@dataclass(config=CONFIG, frozen=True, kw_only=True) +class Trainer: + compute: Compute + backend: Literal["miles", "megatron"] = "miles" + max_instances: PositiveInt = 1 + max_clients_per_instance: PositiveInt = 1 + sampler_persistence_concurrency: PositiveInt = 8 + timeout_s: PositiveInt = 86400 + config: dict[str, object] = field(default_factory=dict) + env: dict[str, str] = field(default_factory=dict) + + +@dataclass(config=CONFIG, frozen=True, kw_only=True) +class Inference: + compute: Compute + backend: Literal["sglang"] = "sglang" + min_replicas: NonnegativeInt = 0 + max_replicas: PositiveInt = 8 + target_concurrency: PositiveInt = 16 + scaledown_window_s: PositiveInt = 300 + startup_timeout_s: PositiveInt = 1200 + config: dict[str, object] = field(default_factory=dict) + env: dict[str, str] = field(default_factory=dict) + + +@dataclass(config=CONFIG, frozen=True, kw_only=True) +class Routing: + default: StrictBool = False + sampling_default: StrictBool = False + + +@dataclass(config=CONFIG, frozen=True, kw_only=True) +class Lifecycle: + session_idle_timeout_s: PositiveInt = 300 + pool_idle_timeout_s: PositiveInt = 300 + sweep_interval_s: PositiveInt = 300 + + +@dataclass(config=CONFIG, frozen=True, kw_only=True) +class Deployment: + name: Annotated[str, Field(pattern=r"^[A-Za-z0-9_-]+$")] + model: Model + trainer: Trainer + inference: Inference + routing: Routing = field(default_factory=Routing) + lifecycle: Lifecycle = field(default_factory=Lifecycle) + + def __post_init__(self): + if self.inference.min_replicas > self.inference.max_replicas: + raise ValueError("min_replicas must not exceed max_replicas") + if self.inference.compute.nodes != 1: + raise ValueError("each inference replica uses one node") + if self.trainer.backend == "megatron": + if self.trainer.compute.nodes != 1: + raise ValueError("multi-node training currently requires Miles") + if self.model.parameterization != "full": + raise ValueError("Megatron requires full parameterization") + if self.trainer.max_clients_per_instance != 1: + raise ValueError("FFT trainers admit one client per instance") + if self.trainer.sampler_persistence_concurrency != 1: + raise ValueError("Megatron requires sampler_persistence_concurrency: 1") + elif self.model.parameterization != "lora": + raise ValueError("Miles requires lora parameterization") + for env in (self.trainer.env, self.inference.env): + if any(key.startswith("LILO_") for key in env): + raise ValueError("LILO_ environment variables are managed by Lilo") diff --git a/src/lilo/deployment_cli.py b/src/lilo/deployment_cli.py index fa6298d..a50b12e 100644 --- a/src/lilo/deployment_cli.py +++ b/src/lilo/deployment_cli.py @@ -5,51 +5,52 @@ import argparse import json import os -from pathlib import Path import re import subprocess import sys import uuid +from copy import deepcopy +from pathlib import Path +import modal +from huggingface_hub import HfApi +from lilo.backends.deployment import resolve_backend_settings from lilo.deployments import ( - DeploymentRecord, PLATFORM_DEFAULTS, - load, + DeploymentRecord, config_path, + load, validate_frontend, ) +from lilo.providers.modal.deployment_records import MANIFEST_ENV +from lilo.providers.modal.miles_revision import resolve_miles_commit def compile_configs(paths, *, platform=None): specs = [load(path) for path in paths] validate_frontend(specs) - from lilo.providers.modal.miles_revision import resolve_miles_commit miles_commit = ( resolve_miles_commit() - if any(spec.trainer["backend"] == "miles" for spec in specs) + if any(spec.trainer.backend == "miles" for spec in specs) else None ) records = [] for spec in specs: - revision = spec.model.get("revision", "main") + revision = spec.model.revision if not re.fullmatch(r"[0-9a-fA-F]{40,64}", revision): - from huggingface_hub import HfApi - - revision = HfApi().model_info(spec.model["id"], revision=revision).sha + revision = HfApi().model_info(spec.model.id, revision=revision).sha if not revision or not re.fullmatch(r"[0-9a-fA-F]{40,64}", revision): raise ValueError( - f"Hugging Face did not return a commit for {spec.model['id']}" + f"Hugging Face did not return a commit for {spec.model.id}" ) records.append( DeploymentRecord.create( spec, platform=platform, revision=revision, - miles_commit=miles_commit - if spec.trainer["backend"] == "miles" - else None, + miles_commit=miles_commit if spec.trainer.backend == "miles" else None, ) ) return records @@ -98,16 +99,7 @@ def select_worker_releases(desired, previous, refresh_trainers, refresh_inferenc if row.spec.name in refresh_inference else old.inference_release ) - result.append( - DeploymentRecord.create( - row.spec, - revision=row.spec.model["revision"], - miles_commit=row.miles_commit, - platform=row.platform, - trainer_release=trainer, - inference_release=inference, - ) - ) + result.append(row.with_releases(trainer, inference)) return result @@ -117,8 +109,6 @@ def worker_apps_ready(row, deployed): def deploy(desired, *, refresh_trainers=(), refresh_inference=()): """Serialize operator applies and retain interrupted attempts for safe recovery.""" - import modal - from lilo.providers.modal.deployment_apps import MANIFEST_ENV if sys.version_info[:2] != (3, 12): raise ValueError( @@ -278,13 +268,13 @@ def main(argv=None): if not config_path(args.preset).is_file(): raise ValueError(f"unknown example config: {args.preset}") print( - f"from lilo.configs.{module} import Config as ParentConfig\n\n\n" - "class Config(ParentConfig):\n" - " overrides = {}" + f'from dataclasses import replace\nfrom lilo.configs.{module} import config as base\n\nconfig = replace(base, name="my-model")' ) elif args.action == "validate": specs = [load(path) for path in args.files] validate_frontend(specs) + for spec in specs: + resolve_backend_settings(spec, "/assets/pending") print( f"Validated {len(specs)} deployment(s). Backend integration settings are checked when preparing trainers and pools; native options are checked at worker startup." ) @@ -304,8 +294,6 @@ def main(argv=None): else: print(output, end="") elif args.command == "deploy": - from copy import deepcopy - platform = deepcopy(PLATFORM_DEFAULTS) platform["frontend"] = args.app platform["modal"].update(environment=args.env, region=args.region) @@ -315,8 +303,6 @@ def main(argv=None): refresh_inference=args.refresh_inference, ) else: - import modal - if args.action == "unlock": registry = modal.Dict.from_name( f"{args.frontend}-yaml-deployments", environment_name=args.env diff --git a/src/lilo/deployments.py b/src/lilo/deployments.py index 46ba58e..37e5347 100644 --- a/src/lilo/deployments.py +++ b/src/lilo/deployments.py @@ -2,19 +2,21 @@ from __future__ import annotations -from copy import deepcopy -from dataclasses import asdict, dataclass, fields import hashlib -from importlib.resources import files import json -from pathlib import Path import re import runpy import sys -from typing import Any, ClassVar +from copy import deepcopy +from dataclasses import asdict, replace +from importlib.resources import files +from pathlib import Path +from typing import Any from pydantic import BaseModel, ConfigDict, Field +from lilo.backends.deployment import resolve_backend_settings +from lilo.configuration import Deployment PLATFORM_DEFAULTS = { "frontend": "lilo-yaml", @@ -31,102 +33,40 @@ }, } -# Shared orchestration defaults. Backend option dictionaries have no schema here. -_DEFAULTS = { - "api_version": "lilo/v1", - "model": {"parameterization": "lora"}, - "routing": {"default": False, "sampling_default": False}, - "trainer": { - "backend": "miles", - "resources": {"cpu": 8, "memory_mib": 32768, "timeout_s": 86400}, - "scaling": {"min_instances": 0, "max_instances": 1}, - "engine": {"max_clients_per_instance": 1, "sampler_persistence_concurrency": 8}, - "config": {}, - "env": {}, - }, - "inference": { - "backend": "sglang", - "resources": {"cpu": 8, "memory_mib": 32768, "timeout_s": 86400}, - "scaling": { - "min_replicas": 0, - "max_replicas": 8, - "target_concurrency": 16, - "scaledown_window_s": 300, - }, - "config": {}, - "env": {}, - }, - "lifecycle": { - "session_idle_timeout_s": 300, - "pool_idle_timeout_s": 300, - "sweep_interval_s": 300, - }, -} - - -def _with_defaults(defaults, values): - """Fill omitted orchestration settings; explicit values win.""" - if not isinstance(defaults, dict) or not isinstance(values, dict): - return deepcopy(values) - result = deepcopy(defaults) - for key, value in values.items(): - result[key] = _with_defaults(defaults.get(key), value) - return result - - -@dataclass(kw_only=True, init=False) -class BaseConfig: - """Declare plain section dictionaries and optional inherited overrides.""" - - name: str - model: dict[str, Any] - trainer: dict[str, Any] - inference: dict[str, Any] - api_version: str - routing: dict[str, Any] - lifecycle: dict[str, Any] - overrides: ClassVar[dict[str, Any]] = {} - - def __init__(self, **kwargs): - names = {item.name for item in fields(BaseConfig)} - unknown = kwargs.keys() - names - if unknown: - raise TypeError(f"unknown config fields: {sorted(unknown)}") - values = deepcopy(_DEFAULTS) - for cls in reversed(type(self).__mro__): - for name in names & vars(cls).keys(): - values[name] = _with_defaults(_DEFAULTS.get(name), vars(cls)[name]) - for path, value in vars(cls).get("overrides", {}).items(): - parts = path.split(".") - if parts[0] not in names: - raise ValueError(f"unknown config override: {path}") - target = values - try: - for part in parts[:-1]: - target = target[part] - target[parts[-1]] = deepcopy(value) - except (KeyError, TypeError) as exc: - raise ValueError(f"unknown config override: {path}") from exc - for name, value in kwargs.items(): - values[name] = _with_defaults(_DEFAULTS.get(name), value) - missing = names - values.keys() - if missing: - raise TypeError(f"missing config fields: {sorted(missing)}") - self.__dict__.update(values) - - -def gpu_count(resources): - gpu = resources["gpu"] - if not re.fullmatch(r"[A-Za-z0-9-]+(?::[1-9][0-9]*)?", gpu): - raise ValueError(f"invalid GPU resource: {gpu}") - return int(gpu.split(":")[1]) if ":" in gpu else 1 - def settings_hash(settings: dict) -> str: """Stable identifier for settings, not a claim that they have been validated.""" return hashlib.sha256(json.dumps(settings, sort_keys=True).encode()).hexdigest() +def deployment_generation( + spec, + platform, + miles_commit, + trainer_release, + inference_release, + trainer_settings, + inference_settings, +): + identity = asdict(spec) + identity.pop("routing") + return settings_hash( + { + "config": identity, + "platform": platform, + "miles_commit": miles_commit, + "trainer_release": trainer_release, + "inference_release": inference_release, + "trainer_settings": trainer_settings, + "inference_settings": inference_settings, + } + ) + + +def platform_defaults(): + return deepcopy(PLATFORM_DEFAULTS) + + class DeploymentRecord(BaseModel): """Saved deployment metadata around a Python configuration. @@ -137,20 +77,20 @@ class DeploymentRecord(BaseModel): model_config = ConfigDict(extra="forbid") - spec: BaseConfig - platform: dict[str, Any] = Field( - default_factory=lambda: deepcopy(PLATFORM_DEFAULTS) - ) + spec: Deployment + platform: dict[str, Any] = Field(default_factory=platform_defaults) trainer_release: str = "initial" inference_release: str = "initial" miles_commit: str | None = None + trainer_settings: dict + inference_settings: dict generation: str active: bool = True @classmethod def create( cls, - spec: BaseConfig, + spec: Deployment, *, revision: str, miles_commit: str | None = None, @@ -159,23 +99,25 @@ def create( inference_release: str = "initial", ) -> DeploymentRecord: """Record an already-resolved revision without reparsing the configuration.""" - pinned = deepcopy(spec) - pinned.model["revision"] = revision - # Changing routing defaults should not restart an existing trainer. - identity = asdict(pinned) - identity.pop("routing") + pinned = replace(deepcopy(spec), model=replace(spec.model, revision=revision)) + asset_path = model_asset_path(pinned.model.id, revision) + trainer_settings, inference_settings = resolve_backend_settings( + pinned, asset_path + ) platform = deepcopy(PLATFORM_DEFAULTS if platform is None else platform) - generation = settings_hash( - { - "config": identity, - "platform": platform, - "miles_commit": miles_commit, - "trainer_release": trainer_release, - "inference_release": inference_release, - } + generation = deployment_generation( + pinned, + platform, + miles_commit, + trainer_release, + inference_release, + trainer_settings, + inference_settings, ) return cls( spec=pinned, + trainer_settings=trainer_settings, + inference_settings=inference_settings, platform=platform, trainer_release=trainer_release, inference_release=inference_release, @@ -183,14 +125,33 @@ def create( miles_commit=miles_commit, ) + def with_releases(self, trainer_release, inference_release): + generation = deployment_generation( + self.spec, + self.platform, + self.miles_commit, + trainer_release, + inference_release, + self.trainer_settings, + self.inference_settings, + ) + return self.model_copy( + update={ + "trainer_release": trainer_release, + "inference_release": inference_release, + "generation": generation, + } + ) + @property def trainer_hash(self) -> str: """Identify trainer settings and the deployment-managed code release.""" return settings_hash( { "name": self.spec.name, - "model": self.spec.model, - "trainer": self.spec.trainer, + "model": asdict(self.spec.model), + "trainer": asdict(self.spec.trainer), + "settings": self.trainer_settings, "release": self.trainer_release, "platform": self.platform, "miles_commit": self.miles_commit, @@ -200,21 +161,14 @@ def trainer_hash(self) -> str: @property def inference_hash(self) -> str: """Identify inference settings, including the adapter shape it must load.""" - adapter = {} - if self.spec.model["parameterization"] == "lora": - options = self.spec.trainer["config"] - adapter = { - "max_lora_rank": options.get("max_lora_rank"), - "target_modules": options.get("target_modules"), - } return settings_hash( { "name": self.spec.name, - "model": self.spec.model, - "inference": self.spec.inference, + "model": asdict(self.spec.model), + "inference": asdict(self.spec.inference), + "settings": self.inference_settings, "release": self.inference_release, "platform": self.platform, - "adapter": adapter, } ) @@ -232,10 +186,12 @@ def definition_id(self) -> str: @property def asset_path(self) -> str: - digest = hashlib.sha256( - f"{self.spec.model['id']}@{self.spec.model['revision']}".encode() - ).hexdigest() - return f"/assets/{digest}" + return model_asset_path(self.spec.model.id, self.spec.model.revision) + + +def model_asset_path(model_id, revision): + digest = hashlib.sha256(f"{model_id}@{revision}".encode()).hexdigest() + return f"/assets/{digest}" def config_path(name: str) -> Path: @@ -245,8 +201,8 @@ def config_path(name: str) -> Path: return Path(str(files("lilo").joinpath("configs", name.replace("-", "_") + ".py"))) -def load(path: str | Path) -> BaseConfig: - """Execute a Python config file and instantiate its exported Config class.""" +def load(path: str | Path) -> Deployment: + """Execute a Python config file and read its exported config object.""" path = Path(path).resolve() if path.suffix != ".py": raise ValueError("deployment configs must be Python .py files") @@ -255,17 +211,15 @@ def load(path: str | Path) -> BaseConfig: sys.path.insert(0, str(path.parent)) try: namespace = runpy.run_path(str(path)) - config_class = namespace.get("Config") - if not isinstance(config_class, type) or not issubclass( - config_class, BaseConfig - ): - raise ValueError(f"{path} must export a Config subclass of BaseConfig") - return config_class() + config = namespace.get("config") + if not isinstance(config, Deployment): + raise ValueError(f"{path} must export a Deployment object named config") + return config finally: sys.path[:] = original_path -def validate_frontend(specs: list[BaseConfig]) -> None: +def validate_frontend(specs: list[Deployment]) -> None: if not specs: raise ValueError("at least one deployment is required") if len({s.name for s in specs}) != len(specs): @@ -278,12 +232,12 @@ def validate_frontend(specs: list[BaseConfig]) -> None: ) defaults, sampling = set(), set() for spec in specs: - key = (spec.model["id"], spec.model["parameterization"]) - if spec.routing["default"]: + key = (spec.model.id, spec.model.parameterization) + if spec.routing.default: if key in defaults: raise ValueError(f"multiple defaults for {key}") defaults.add(key) - if spec.routing["sampling_default"]: - if spec.model["id"] in sampling: - raise ValueError(f"multiple sampling defaults for {spec.model['id']}") - sampling.add(spec.model["id"]) + if spec.routing.sampling_default: + if spec.model.id in sampling: + raise ValueError(f"multiple sampling defaults for {spec.model.id}") + sampling.add(spec.model.id) diff --git a/src/lilo/inference/native_sglang.py b/src/lilo/inference/native_sglang.py deleted file mode 100644 index 56ebd77..0000000 --- a/src/lilo/inference/native_sglang.py +++ /dev/null @@ -1,33 +0,0 @@ -"""SGLang entrypoint for resolved YAML serving options.""" - -import argparse -import json -import logging -import os -import sys - -from lilo.argparse_config import apply_config_overrides - - -def main(): - from sglang.srt.server_args import ServerArgs - from sglang.launch_server import run_server - from sglang.srt.utils import kill_process_tree - from sglang.srt.plugins import load_plugins - - load_plugins() - parser = argparse.ArgumentParser() - ServerArgs.add_cli_args(parser) - argv = ["--model-path", sys.argv[1], "--host", "127.0.0.1", "--port", "8001"] - apply_config_overrides(parser, json.loads(sys.argv[2]), argv) - raw = parser.parse_args(argv) - logging.basicConfig(level=getattr(logging, raw.log_level.upper())) - args = ServerArgs.from_cli_args(raw) - try: - run_server(args) - finally: - kill_process_tree(os.getpid(), include_parent=False) - - -if __name__ == "__main__": - main() diff --git a/src/lilo/inference/sglang.py b/src/lilo/inference/sglang.py new file mode 100644 index 0000000..04a5416 --- /dev/null +++ b/src/lilo/inference/sglang.py @@ -0,0 +1,30 @@ +"""SGLang worker entrypoint. The frontend never imports this GPU-only module.""" + +import json +import logging +import os +import sys + +from sglang.launch_server import run_server +from sglang.srt.plugins import load_plugins +from sglang.srt.server_args import ServerArgs +from sglang.srt.utils import kill_process_tree + + +def main(): + load_plugins() + args = ServerArgs( + model_path=sys.argv[1], + host="127.0.0.1", + port=8001, + **json.loads(sys.argv[2]), + ) + logging.basicConfig(level=args.log_level.upper()) + try: + run_server(args) + finally: + kill_process_tree(os.getpid(), include_parent=False) + + +if __name__ == "__main__": + main() diff --git a/src/lilo/inference/sglang_deployment.py b/src/lilo/inference/sglang_deployment.py index 84cc0ea..b0491d4 100644 --- a/src/lilo/inference/sglang_deployment.py +++ b/src/lilo/inference/sglang_deployment.py @@ -1,6 +1,5 @@ """SGLang settings that must agree with Lilo replica orchestration.""" -from lilo.deployments import gpu_count from lilo.config_validation import reject_managed_options SGLANG_MANAGED = { @@ -11,6 +10,7 @@ "context_length", "enable_lora", "max_lora_rank", + "lora_target_modules", "enable_cpu_weight_cache", "api_key", "pp_size", @@ -31,13 +31,13 @@ def build_config(spec): - options = dict(spec.inference["config"]) + options = dict(spec.inference.config) reject_managed_options(options, SGLANG_MANAGED) - tp = options.get("tp_size", gpu_count(spec.inference["resources"])) + tp = options.get("tp_size", spec.inference.compute.gpus_per_node) ep = options.get("ep_size", 1) - if not isinstance(tp, int) or tp != gpu_count(spec.inference["resources"]): + if type(tp) is not int or tp != spec.inference.compute.gpus_per_node: raise ValueError("sglang.tp_size must equal the replica GPU allocation") - if not isinstance(ep, int) or ep < 1 or tp % ep: + if type(ep) is not int or ep < 1 or tp % ep: raise ValueError("sglang.ep_size must divide the replica GPU allocation") dp = options.get("dp_size", 1) dp_attention = options.get("enable_dp_attention", False) @@ -53,7 +53,7 @@ def build_config(spec): "max_running_requests", "max_queued_requests", ): - if key in options and (not isinstance(options[key], int) or options[key] < 1): + if key in options and (type(options[key]) is not int or options[key] < 1): raise ValueError(f"sglang.{key} must be positive") if not 0 < options.get("mem_fraction_static", 0.8) < 1: raise ValueError("sglang.mem_fraction_static must be between zero and one") diff --git a/src/lilo/providers/modal/app.py b/src/lilo/providers/modal/app.py index 0fe2e03..5e6ee50 100644 --- a/src/lilo/providers/modal/app.py +++ b/src/lilo/providers/modal/app.py @@ -48,24 +48,23 @@ stop_pool as stop_lora_pool, ) from .sampling import ModalSamplingTaskPlatform +from .deployment_records import MANIFEST_ENV, manifest_from_env from .deployment_apps import ( - MANIFEST_ENV, definition_from_spec, frontend_settings, - manifest_from_env, ) SETTINGS = frontend_settings() APP_NAME = SETTINGS.platform["frontend"] ROUTING_REGION = SETTINGS.platform["modal"]["region"] MODEL_ASSET_ROOT = "/assets" -SESSION_IDLE_TIMEOUT = SETTINGS.spec.lifecycle["session_idle_timeout_s"] -FFT_POOL_IDLE_TIMEOUT = LORA_POOL_IDLE_TIMEOUT = SETTINGS.spec.lifecycle[ - "pool_idle_timeout_s" -] +SESSION_IDLE_TIMEOUT = SETTINGS.spec.lifecycle.session_idle_timeout_s +FFT_POOL_IDLE_TIMEOUT = LORA_POOL_IDLE_TIMEOUT = ( + SETTINGS.spec.lifecycle.pool_idle_timeout_s +) FFT_POOL_TOUCH_INTERVAL = 60.0 LORA_POOL_CHECK_INTERVAL = 60.0 -SWEEP_PERIOD = modal.Period(seconds=SETTINGS.spec.lifecycle["sweep_interval_s"]) +SWEEP_PERIOD = modal.Period(seconds=SETTINGS.spec.lifecycle.sweep_interval_s) CHECKPOINT_READ_LOCK = asyncio.Lock() _pool_touches: dict[str, float] = {} _lora_pool_gateways: dict[str, tuple[float, str]] = {} diff --git a/src/lilo/providers/modal/definitions/qwen3_8_27b_miles_lora_256k.py b/src/lilo/providers/modal/definitions/qwen3_8_27b_miles_lora_256k.py deleted file mode 100644 index 85bc2ea..0000000 --- a/src/lilo/providers/modal/definitions/qwen3_8_27b_miles_lora_256k.py +++ /dev/null @@ -1,223 +0,0 @@ -from __future__ import annotations - -import os - -import modal -import modal.experimental - -from ..checkpoint_storage import ( - CHECKPOINT_ROOT, - CHECKPOINT_VOLUME_NAME, - checkpoint_volume, -) -from ..deployment import trainer_deployment_env, trainer_max_containers - -MODEL_NAME = "Qwen/Qwen3.8-27B" -HF_CHECKPOINT = "/assets/Qwen3.8-27B" -DEFINITION_ID = "qwen3_8_27b_miles_lora_256k" -PARAMETERIZATION = "lora" -CATALOG_VISIBLE = False -MAX_CONTEXT_LENGTH = 262_144 - -GPU_TYPE = "H200" -GPUS = 8 -TRAINER_NODES = 2 -# 16 = TP2 x CP8 x DP1, the topology raw Miles runs 256k on. TP2 keeps the 27B -# base weights at ~27 GB/GPU of the 141 GB H200, and spending the rest of the -# world size on CP is what shrinks the activation working set: 32k tokens per -# rank instead of 87k under TP8 x CP3. -TENSOR_MODEL_PARALLEL_SIZE = 2 -CONTEXT_PARALLEL_SIZE = 8 -# Divisible by 2 * cp (zigzag chunks) and by tp (sequence parallelism). -_SEQ_ALIGNMENT = 2 * CONTEXT_PARALLEL_SIZE * TENSOR_MODEL_PARALLEL_SIZE -SEQ_LENGTH = -(-MAX_CONTEXT_LENGTH // _SEQ_ALIGNMENT) * _SEQ_ALIGNMENT -MAX_TOKENS_PER_GPU = SEQ_LENGTH // CONTEXT_PARALLEL_SIZE -MAX_LORA_SLOTS = 6 -MAX_LORA_RANK = 32 -DEFAULT_LORA_ALPHA = 32 -TARGET_MODULES = ( - "linear_qkv", - "linear_proj", - "linear_fc1", - "linear_fc2", -) -TRAINER_MODELS_PER_INSTANCE = MAX_LORA_SLOTS -# Megatron's default 10-minute collective timeout is the NCCL watchdog budget -# for a single collective. At 250k tokens the inter-node context-parallel -# all-to-all runs behind a straggler's recompute, so a rank can sit in one -# collective far longer than ten minutes without anything being wrong. -DISTRIBUTED_TIMEOUT_MINUTES = 120 - -ROLLOUT_GPU_TYPE = "H200" -ROLLOUT_GPUS = 4 -ROLLOUT_TENSOR_PARALLEL_SIZE = 4 -ROLLOUT_EXPERT_PARALLEL_SIZE = 1 -ROLLOUT_EXPERT_TENSOR_PARALLEL_SIZE = 1 -ROLLOUT_MEMORY_FRACTION = 0.8 -ROLLOUT_MAX_RUNNING_REQUESTS = 4 -ROLLOUT_MAX_QUEUED_REQUESTS = 8 -ROLLOUT_TARGET_CONCURRENCY = 2 -ROLLOUT_MAX_LOADED_LORAS = 256 -ROLLOUT_MIN_CONTAINERS = 2 -ROLLOUT_MAX_CONTAINERS = 2 -ROLLOUT_MAX_LORAS_PER_BATCH = 8 -ROLLOUT_LORA_TARGET_MODULES = ( - "q_proj", - "k_proj", - "v_proj", - "o_proj", - "gate_proj", - "up_proj", - "down_proj", -) - -BULLETIN_ROOT = "/bulletin" -BULLETIN_VOLUME_NAME = "lilo-snapshot-bulletin" -app = modal.App(f"lilo-{DEFINITION_ID}") - -if modal.is_local(): - from ..miles_image import image -else: - image = modal.Image.debian_slim() - -assets = modal.Volume.from_name("lilo-model-assets", create_if_missing=True) -bulletin = modal.Volume.from_name( - BULLETIN_VOLUME_NAME, - create_if_missing=True, - version=2, -) -TRAINER_VOLUMES = { - "/assets": assets, - BULLETIN_ROOT: bulletin, - CHECKPOINT_ROOT: checkpoint_volume, -} -api_secret = modal.Secret.from_name( - "lilo-api", - required_keys=["TINKER_API_KEY"], -) -proxy_secret = modal.Secret.from_name( - "lilo-proxy", - required_keys=["MODAL_PROXY_TOKEN_ID", "MODAL_PROXY_TOKEN_SECRET"], -) -huggingface_secret = modal.Secret.from_name("huggingface-secret") - - -def ensure_assets() -> None: - from huggingface_hub import snapshot_download - - if not os.path.exists(HF_CHECKPOINT): - snapshot_download(repo_id=MODEL_NAME, local_dir=HF_CHECKPOINT) - assets.commit() - - -@app.function( - image=image, - gpu=f"{GPU_TYPE}:{GPUS}", - volumes=TRAINER_VOLUMES, - secrets=[api_secret, proxy_secret, huggingface_secret], - env=trainer_deployment_env(), - timeout=86_400, - max_containers=trainer_max_containers(), - single_use_containers=True, - experimental_options={"efa_enabled": True}, -) -@modal.experimental.clustered(TRAINER_NODES, rdma=True) -def qwen3_8_27b_miles_lora_256k(instance_id: str) -> None: - from lilo.providers.modal.ray_cluster import start_trainer_cluster - - ray_address = start_trainer_cluster( - TRAINER_NODES, - before_head=ensure_assets, - before_worker_join=assets.reload, - ) - if ray_address is None: - return - run_trainer(instance_id, ray_address=ray_address) - - -def backend_config( - instance_id: str, - *, - deterministic_training: bool = False, -) -> dict: - config = { - "miles": { - "hf_checkpoint": HF_CHECKPOINT, - "model_type": "qwen3.8-27B", - "actor_num_gpus_per_node": GPUS, - "actor_num_nodes": TRAINER_NODES, - "tensor_model_parallel_size": TENSOR_MODEL_PARALLEL_SIZE, - "context_parallel_size": CONTEXT_PARALLEL_SIZE, - "max_lora_slots": MAX_LORA_SLOTS, - "max_lora_rank": MAX_LORA_RANK, - "default_lora_alpha": DEFAULT_LORA_ALPHA, - "target_modules": TARGET_MODULES, - "max_tokens_per_gpu": MAX_TOKENS_PER_GPU, - "extra_args": ( - "--seq-length", - str(SEQ_LENGTH), - "--recompute-granularity", - "full", - "--recompute-method", - "uniform", - "--recompute-num-layers", - "1", - "--distributed-timeout-minutes", - str(DISTRIBUTED_TIMEOUT_MINUTES), - ), - }, - "checkpoint_dir": CHECKPOINT_ROOT, - "capture_dir": f"{CHECKPOINT_ROOT}/.captures/{instance_id}", - } - if deterministic_training: - config["miles"].update( - tp_reduce_precision="float64", deterministic_attention=True - ) - return config - - -def run_trainer( - instance_id: str, - *, - definition_id: str = DEFINITION_ID, - max_models: int = MAX_LORA_SLOTS, - deterministic_training: bool = False, - ray_address: str | None = None, -) -> None: - import json - - from modal.config import config - - from lilo.providers.modal.kv import shared_kv - from lilo.providers.modal.serve import run_engine_with_backend - - ensure_assets() - config_payload = backend_config( - instance_id, - deterministic_training=deterministic_training, - ) - run_engine_with_backend( - shared_kv(), - "lilo.backends.miles_lora:build_executor", - definition_id=definition_id, - revision=config["image_id"], - instance_id=instance_id, - backend_env={ - "LILO_BACKEND_CONFIG": json.dumps(config_payload), - "LILO_BASE_MODEL": MODEL_NAME, - "LILO_DEFINITION_ID": definition_id, - "LILO_CHECKPOINT_VOLUME": CHECKPOINT_VOLUME_NAME, - "LILO_BULLETIN_ROOT": BULLETIN_ROOT, - "LILO_BULLETIN_VOLUME": BULLETIN_VOLUME_NAME, - "LILO_DEFINITION_REVISION": config["image_id"], - "PYTORCH_CUDA_ALLOC_CONF": "expandable_segments:True", - "TORCHINDUCTOR_COMPILE_THREADS": "1", - **({"LILO_RAY_ADDRESS": ray_address} if ray_address else {}), - }, - nproc=1, - max_models=max_models, - sampler_persistence_concurrency=8, - ) - - -ENGINE_FUNCTION = qwen3_8_27b_miles_lora_256k diff --git a/src/lilo/providers/modal/deployment_apps.py b/src/lilo/providers/modal/deployment_apps.py index 0ca5428..db1ffab 100644 --- a/src/lilo/providers/modal/deployment_apps.py +++ b/src/lilo/providers/modal/deployment_apps.py @@ -7,28 +7,41 @@ from __future__ import annotations import json -import os +import subprocess +import sys +from importlib import import_module from types import SimpleNamespace import modal - -from lilo.deployments import DeploymentRecord, gpu_count, validate_frontend -from lilo.backends.deployment import backend_config, serving_options - -MANIFEST_ENV = "LILO_DEPLOYMENT_MANIFEST" -POOL_CONFIG_ENV = "LILO_POOL_DEPLOYMENT" - - -def manifest_from_env(): - data = os.environ.get(MANIFEST_ENV) - if not data: - raise ValueError( - "Missing deployment manifest. Use lilo deploy with your Python config files." - ) - rows = json.loads(data) - if not isinstance(rows, list) or not rows: - raise ValueError("Deployment manifest must be a nonempty list") - return [DeploymentRecord.model_validate(row) for row in rows] +import modal.experimental +from modal.config import config + +from lilo.deployments import DeploymentRecord, validate_frontend +from lilo.inference.serving import ( + start_fft_sidecar, + start_lora_sidecar, + supervise_children, + terminate, + wait_http, +) + +from .deployment import trainer_deployment_env +from .deployment_records import ( + manifest_from_env, +) +from .fft_pool import FFTPoolSpec +from .fft_pool import deploy_pool as deploy_fft +from .image_dependencies import ( + CORE_PACKAGES, + STITCH_PACKAGE, + TINKER_PACKAGE, + ignore_config_source, +) +from .kv import shared_kv +from .lora_pool import LoraPoolSpec +from .lora_pool import deploy_pool as deploy_lora +from .ray_cluster import start_trainer_cluster +from .serve import run_engine_with_backend def frontend_settings(): @@ -43,13 +56,13 @@ def frontend_settings(): def image_for(backend): if not modal.is_local(): return modal.Image.debian_slim() - if backend == "miles": - from .miles_image import image - elif backend == "megatron": - from .megatron_image import image - else: - from .rollout_image import image - return image + # Image definitions are imported only by the selected worker's deploy process. + modules = { + "miles": ".miles_image", + "megatron": ".megatron_image", + "sglang": ".rollout_image", + } + return import_module(modules[backend], __package__).image def volumes_for(record): @@ -91,24 +104,31 @@ def build_trainer_app(resolved: DeploymentRecord, *, image=None): spec = resolved.spec trainer_hash = resolved.trainer_hash app = modal.App(resolved.trainer_app_name) - resource = spec.trainer["resources"] - from .deployment import trainer_deployment_env + resource = spec.trainer.compute env = { **trainer_deployment_env(), - **deployment_env(spec.trainer["env"]), + **deployment_env(spec.trainer.env), "LILO_APP_NAME": resolved.platform["frontend"], } - @app.function( + def trainer(instance_id: str, config_json: str): + record = DeploymentRecord.model_validate_json(config_json) + if record.trainer_hash != trainer_hash: + raise ValueError("trainer settings do not match the deployed app") + run_trainer(record, instance_id) + + if resource.nodes > 1: + trainer = modal.experimental.clustered(resource.nodes, rdma=True)(trainer) + trainer = app.function( name="trainer", serialized=True, - image=image if image is not None else image_for(spec.trainer["backend"]), - gpu=resource["gpu"], + image=image if image is not None else image_for(spec.trainer.backend), + gpu=resource.modal_gpu, region=resolved.platform["modal"]["region"], - cpu=resource["cpu"], - memory=resource["memory_mib"], - timeout=resource["timeout_s"], + cpu=resource.cpu, + memory=resource.memory_mib, + timeout=spec.trainer.timeout_s, # Admission/reconciliation caps each definition. A function-wide cap # would block new definitions behind retained jobs sharing this app. max_containers=None, @@ -117,41 +137,46 @@ def build_trainer_app(resolved: DeploymentRecord, *, image=None): volumes=volumes_for(resolved), secrets=secrets_for(resolved, training=True), env=env, - ) - def trainer(instance_id: str, config_json: str): - record = DeploymentRecord.model_validate_json(config_json) - if record.trainer_hash != trainer_hash: - raise ValueError("trainer settings do not match the deployed app") - run_trainer(record, instance_id) + experimental_options={"efa_enabled": True} if resource.nodes > 1 else {}, + )(trainer) return app, trainer def run_trainer(resolved, instance_id): - from modal.config import config - from .kv import shared_kv - from .serve import run_engine_with_backend - spec = resolved.spec - settings = backend_config(spec, resolved.asset_path) + settings = resolved.trainer_settings # Assets are prepared by the frontend before demand is registered. Reload once # on startup to see the committed exact snapshot; never race a trainer download. - volumes_for(resolved)["/assets"].reload() + assets = volumes_for(resolved)["/assets"] + if spec.trainer.compute.nodes > 1: + ray_address = start_trainer_cluster( + spec.trainer.compute.nodes, + before_head=assets.reload, + before_worker_join=assets.reload, + ) + if ray_address is None: + return + else: + assets.reload() + ray_address = None env = { - **deployment_env(spec.trainer["env"]), + **deployment_env(spec.trainer.env), "LILO_APP_NAME": resolved.platform["frontend"], "LILO_BACKEND_CONFIG": json.dumps(settings), - "LILO_BASE_MODEL": spec.model["id"], - "LILO_BASE_MODEL_REVISION": spec.model["revision"], + "LILO_BASE_MODEL": spec.model.id, + "LILO_BASE_MODEL_REVISION": spec.model.revision, "LILO_DEFINITION_ID": resolved.definition_id, "LILO_CHECKPOINT_VOLUME": resolved.platform["storage"]["checkpoints"], "LILO_BULLETIN_ROOT": "/bulletin", "LILO_BULLETIN_VOLUME": resolved.platform["storage"]["bulletin"], "LILO_DEFINITION_REVISION": resolved.generation, } + if ray_address is not None: + env["LILO_RAY_ADDRESS"] = ray_address executor = ( "lilo.backends.miles_lora:build_executor" - if spec.trainer["backend"] == "miles" + if spec.trainer.backend == "miles" else "lilo.backends.megatron_fft:build_executor" ) @@ -169,38 +194,36 @@ async def failed(error): instance_id=instance_id, backend_env=env, nproc=1 - if spec.trainer["backend"] == "miles" - else gpu_count(spec.trainer["resources"]), - max_models=spec.trainer["engine"]["max_clients_per_instance"], - sampler_persistence_concurrency=spec.trainer["engine"][ - "sampler_persistence_concurrency" - ], + if spec.trainer.backend == "miles" + else spec.trainer.compute.gpus_per_node, + max_models=spec.trainer.max_clients_per_instance, + sampler_persistence_concurrency=spec.trainer.sampler_persistence_concurrency, on_startup_error=failed, ) def definition_from_spec(resolved, *, register_trainer=True, image=None): spec = resolved.spec - native = serving_options(spec) + serving = resolved.inference_settings definition = SimpleNamespace( DEFINITION_ID=resolved.definition_id, - MODEL_NAME=spec.model["id"], - MODEL_REVISION=spec.model["revision"], + MODEL_NAME=spec.model.id, + MODEL_REVISION=spec.model.revision, HF_CHECKPOINT=resolved.asset_path, - PARAMETERIZATION=spec.model["parameterization"], + PARAMETERIZATION=spec.model.parameterization, CATALOG_VISIBLE=resolved.active, - ROUTING_DEFAULT=spec.routing["default"], - SAMPLING_DEFAULT=spec.routing["sampling_default"], + ROUTING_DEFAULT=spec.routing.default, + SAMPLING_DEFAULT=spec.routing.sampling_default, DEPLOYMENT_NAME=spec.name, RESOLVED=resolved, - MAX_CONTEXT_LENGTH=spec.model["max_context_length"], - TRAINER_MODELS_PER_INSTANCE=spec.trainer["engine"]["max_clients_per_instance"], - TRAINER_MAX_CONTAINERS=spec.trainer["scaling"]["max_instances"], - ROLLOUT_GPUS=gpu_count(spec.inference["resources"]), - ROLLOUT_TENSOR_PARALLEL_SIZE=native.get( - "tp_size", gpu_count(spec.inference["resources"]) + MAX_CONTEXT_LENGTH=spec.model.max_context_length, + TRAINER_MODELS_PER_INSTANCE=spec.trainer.max_clients_per_instance, + TRAINER_MAX_CONTAINERS=spec.trainer.max_instances, + ROLLOUT_GPUS=spec.inference.compute.gpus_per_node, + ROLLOUT_TENSOR_PARALLEL_SIZE=serving.get( + "tp_size", spec.inference.compute.gpus_per_node ) - // (native.get("dp_size", 1) if native.get("enable_dp_attention") else 1), + // (serving.get("dp_size", 1) if serving.get("enable_dp_attention") else 1), ) if register_trainer: definition.ENGINE_FUNCTION = modal.Function.from_name( @@ -211,67 +234,39 @@ def definition_from_spec(resolved, *, register_trainer=True, image=None): return definition -def pool_deployment(definition_id): - for row in manifest_from_env(): - if row.definition_id == definition_id: - return row - return None - - -def pool_environment(definition_id): - resolved = pool_deployment(definition_id) - if resolved is None: - raise ValueError(f"missing recorded deployment: {definition_id}") - return {POOL_CONFIG_ENV: resolved.model_dump_json()} - - def build_rollout_app(resolved, pool, *, image=None): """Create one frozen-base LoRA pool or one FFT latest/pinned/base pool.""" spec = resolved.spec - lora = spec.model["parameterization"] == "lora" + lora = spec.model.parameterization == "lora" if pool.definition_id != resolved.definition_id: raise ValueError("pool generation does not match deployment") app = modal.App(pool.app_name) - resources, scaling = spec.inference["resources"], spec.inference["scaling"] - options = { - "context_length": spec.model["max_context_length"], - "tp_size": gpu_count(resources), - "mem_fraction_static": 0.8, - "max_running_requests": 32, - "weight_loader_disable_mmap": True, - **serving_options(spec), - } - if lora: - from lilo.backends.miles_config import MilesBackendConfig - - config = backend_config(spec)["miles"] - targets = MilesBackendConfig(**config).peft_target_modules - options.update(enable_lora=True, max_lora_rank=config["max_lora_rank"]) - options.setdefault("lora_target_modules", list(targets)) - options.setdefault("max_loaded_loras", 64) - options.setdefault("max_loras_per_batch", 8) - else: - options["enable_cpu_weight_cache"] = True - minimum = getattr(pool, "min_containers", None) - maximum = getattr(pool, "max_containers", None) - window = getattr(pool, "scaledown_window", None) + resources, scaling = spec.inference.compute, spec.inference + options = resolved.inference_settings + minimum = scaling.min_replicas + maximum = scaling.max_replicas + window = scaling.scaledown_window_s + if isinstance(pool, FFTPoolSpec): + minimum = minimum if pool.min_containers is None else pool.min_containers + maximum = maximum if pool.max_containers is None else pool.max_containers + window = window if pool.scaledown_window is None else pool.scaledown_window # Serialized class captures the spec; no model-specific module is imported. @app.server( name="Server", serialized=True, image=image if image is not None else image_for("sglang"), - gpu=resources["gpu"], - cpu=resources["cpu"], - memory=resources["memory_mib"], + gpu=resources.modal_gpu, + cpu=resources.cpu, + memory=resources.memory_mib, volumes=volumes_for(resolved), secrets=secrets_for(resolved), - env=deployment_env(spec.inference["env"]), - min_containers=scaling["min_replicas"] if minimum is None else minimum, - max_containers=scaling["max_replicas"] if maximum is None else maximum, - target_concurrency=scaling["target_concurrency"], - scaledown_window=scaling["scaledown_window_s"] if window is None else window, - startup_timeout=1200, + env=deployment_env(spec.inference.env), + min_containers=minimum, + max_containers=maximum, + target_concurrency=scaling.target_concurrency, + scaledown_window=window, + startup_timeout=spec.inference.startup_timeout_s, exit_grace_period=300, port=8000, routing_region=resolved.platform["modal"]["region"], @@ -280,28 +275,25 @@ def build_rollout_app(resolved, pool, *, image=None): class Server: @modal.enter() def start(self): - import subprocess - import sys - from lilo.inference.serving import ( - start_lora_sidecar, - start_fft_sidecar, - supervise_children, - terminate, - wait_http, - ) + self.sidecar = None + self.sglang = None self.sglang = subprocess.Popen( [ sys.executable, "-m", - "lilo.inference.native_sglang", + "lilo.inference.sglang", resolved.asset_path, json.dumps(options), ], start_new_session=True, ) try: - wait_http("http://127.0.0.1:8001/health", self.sglang, 1200) + wait_http( + "http://127.0.0.1:8001/health", + self.sglang, + spec.inference.startup_timeout_s, + ) kwargs = dict( port=8000, sglang_port=8001, @@ -319,40 +311,26 @@ def start(self): ) ) self.supervisor = supervise_children(self.sglang, self.sidecar) - wait_http("http://127.0.0.1:8000/health", self.sidecar, 1200) + wait_http( + "http://127.0.0.1:8000/health", + self.sidecar, + spec.inference.startup_timeout_s, + ) except BaseException: - terminate(getattr(self, "sidecar", None)) + terminate(self.sidecar) terminate(self.sglang) raise @modal.exit() def stop(self): - from lilo.inference.serving import terminate - - terminate(getattr(self, "sidecar", None)) - terminate(getattr(self, "sglang", None)) + terminate(self.sidecar) + terminate(self.sglang) return app, Server -def provision_pool(record, pool): - """Ask the saved inference app to create a pool using its original code.""" - provision = modal.Function.from_name( - record.inference_app_name, - "provision", - environment_name=record.platform["modal"]["environment"], - ) - return provision.remote(record.model_dump_json(), pool.as_dict()) - - def build_inference_app(record, *, image=None): """Freeze pool-building code so idle pools can restart after frontend upgrades.""" - from .image_dependencies import ( - CORE_PACKAGES, - STITCH_PACKAGE, - TINKER_PACKAGE, - ignore_config_source, - ) if image is None: image = ( @@ -368,15 +346,12 @@ def build_inference_app(record, *, image=None): @app.function(name="provision", image=image, serialized=True, timeout=1800) def provision(config_json: str, pool_data: dict): - from .lora_pool import LoraPoolSpec, deploy_pool as deploy_lora - from .fft_pool import FFTPoolSpec, deploy_pool as deploy_fft - saved = DeploymentRecord.model_validate_json(config_json) if saved.inference_hash != inference_hash: raise ValueError("inference settings do not match the deployed app") if pool_data["definition_id"] != saved.definition_id: raise ValueError("pool definition does not match deployment") - if saved.spec.model["parameterization"] == "lora": + if saved.spec.model.parameterization == "lora": return deploy_lora(LoraPoolSpec.from_dict(pool_data), record=saved) return deploy_fft(FFTPoolSpec.from_dict(pool_data), record=saved) diff --git a/src/lilo/providers/modal/deployment_pool_app.py b/src/lilo/providers/modal/deployment_pool_app.py index a4348ec..402ddce 100644 --- a/src/lilo/providers/modal/deployment_pool_app.py +++ b/src/lilo/providers/modal/deployment_pool_app.py @@ -3,12 +3,14 @@ import os from lilo.deployments import DeploymentRecord + +from .deployment_apps import build_rollout_app +from .deployment_records import POOL_CONFIG_ENV from .fft_pool import FFTPoolSpec from .lora_pool import LoraPoolSpec -from .deployment_apps import POOL_CONFIG_ENV, build_rollout_app resolved = DeploymentRecord.model_validate_json(os.environ[POOL_CONFIG_ENV]) -if resolved.spec.model["parameterization"] == "lora": +if resolved.spec.model.parameterization == "lora": pool = LoraPoolSpec(resolved.definition_id, revision=resolved.generation[:16]) else: pool = FFTPoolSpec( diff --git a/src/lilo/providers/modal/deployment_records.py b/src/lilo/providers/modal/deployment_records.py new file mode 100644 index 0000000..586fd8c --- /dev/null +++ b/src/lilo/providers/modal/deployment_records.py @@ -0,0 +1,47 @@ +"""Saved records used to route worker and pool requests.""" + +import json +import os + +import modal + +from lilo.deployments import DeploymentRecord + +MANIFEST_ENV = "LILO_DEPLOYMENT_MANIFEST" +POOL_CONFIG_ENV = "LILO_POOL_DEPLOYMENT" + + +def manifest_from_env(): + data = os.environ.get(MANIFEST_ENV) + if not data: + raise ValueError( + "Missing deployment manifest. Use lilo deploy with your Python config files." + ) + rows = json.loads(data) + if not isinstance(rows, list) or not rows: + raise ValueError("Deployment manifest must be a nonempty list") + return [DeploymentRecord.model_validate(row) for row in rows] + + +def pool_deployment(definition_id): + for row in manifest_from_env(): + if row.definition_id == definition_id: + return row + return None + + +def pool_environment(definition_id): + resolved = pool_deployment(definition_id) + if resolved is None: + raise ValueError(f"missing recorded deployment: {definition_id}") + return {POOL_CONFIG_ENV: resolved.model_dump_json()} + + +def provision_pool(record, pool): + """Ask the saved inference app to create a pool using its original code.""" + provision = modal.Function.from_name( + record.inference_app_name, + "provision", + environment_name=record.platform["modal"]["environment"], + ) + return provision.remote(record.model_dump_json(), pool.as_dict()) diff --git a/src/lilo/providers/modal/deployment_worker_app.py b/src/lilo/providers/modal/deployment_worker_app.py index ed8fac8..89f75c1 100644 --- a/src/lilo/providers/modal/deployment_worker_app.py +++ b/src/lilo/providers/modal/deployment_worker_app.py @@ -3,7 +3,8 @@ import os from lilo.deployments import DeploymentRecord -from .deployment_apps import build_trainer_app, build_inference_app + +from .deployment_apps import build_inference_app, build_trainer_app record = DeploymentRecord.model_validate_json(os.environ["LILO_WORKER_DEPLOYMENT"]) if record.miles_commit: diff --git a/src/lilo/providers/modal/fft_pool.py b/src/lilo/providers/modal/fft_pool.py index fda9cf2..be64006 100644 --- a/src/lilo/providers/modal/fft_pool.py +++ b/src/lilo/providers/modal/fft_pool.py @@ -3,6 +3,9 @@ import hashlib import logging import os +import modal + +from .deployment_records import POOL_CONFIG_ENV, pool_deployment, provision_pool import shutil import subprocess from concurrent.futures import ThreadPoolExecutor @@ -73,8 +76,6 @@ def __init__(self, definition_id: str, model_id: str) -> None: ) def discover_replicas(self) -> list[str]: - import modal - try: return super().discover_replicas() except modal.exception.NotFoundError: @@ -124,13 +125,9 @@ def deploy_pool(spec: FFTPoolSpec, *, record=None) -> str: try: return pool.gateway_url() except Exception as exc: - import modal - if not isinstance(exc, modal.exception.NotFoundError): raise if record is None: - from .deployment_apps import pool_deployment, provision_pool - saved = pool_deployment(spec.definition_id) if saved is None: raise ValueError(f"missing recorded deployment: {spec.definition_id}") @@ -138,7 +135,6 @@ def deploy_pool(spec: FFTPoolSpec, *, record=None) -> str: modal_cli = shutil.which("modal") if modal_cli is None: raise RuntimeError("modal CLI is unavailable") - from .deployment_apps import POOL_CONFIG_ENV recipe_env = {POOL_CONFIG_ENV: record.model_dump_json()} env = {**os.environ, **spec.env(), **recipe_env} diff --git a/src/lilo/providers/modal/lora_pool.py b/src/lilo/providers/modal/lora_pool.py index 5d9f592..034f3ac 100644 --- a/src/lilo/providers/modal/lora_pool.py +++ b/src/lilo/providers/modal/lora_pool.py @@ -2,6 +2,9 @@ import hashlib import os +import modal + +from .deployment_records import POOL_CONFIG_ENV, pool_deployment, provision_pool import shutil import subprocess from dataclasses import asdict, dataclass @@ -58,13 +61,9 @@ def deploy_pool(spec: LoraPoolSpec, *, record=None) -> str: try: return pool.gateway_url() except Exception as exc: - import modal - if not isinstance(exc, modal.exception.NotFoundError): raise if record is None: - from .deployment_apps import pool_deployment, provision_pool - saved = pool_deployment(spec.definition_id) if saved is None: raise ValueError(f"missing recorded deployment: {spec.definition_id}") @@ -72,7 +71,6 @@ def deploy_pool(spec: LoraPoolSpec, *, record=None) -> str: modal_cli = shutil.which("modal") if modal_cli is None: raise RuntimeError("modal CLI is unavailable") - from .deployment_apps import POOL_CONFIG_ENV recipe_env = {POOL_CONFIG_ENV: record.model_dump_json()} command = [ diff --git a/tests/backends/test_native_megatron_config.py b/tests/backends/test_megatron_constructor_settings.py similarity index 80% rename from tests/backends/test_native_megatron_config.py rename to tests/backends/test_megatron_constructor_settings.py index d62283e..bcbd696 100644 --- a/tests/backends/test_native_megatron_config.py +++ b/tests/backends/test_megatron_constructor_settings.py @@ -1,6 +1,6 @@ """CPU checks for native config forwarding into Megatron constructors.""" -from dataclasses import asdict +from dataclasses import asdict, make_dataclass from pydantic import TypeAdapter from types import SimpleNamespace @@ -11,7 +11,7 @@ from lilo.backends.deployment import backend_config from lilo.backends.megatron_config import parse_backend_config -from lilo.deployments import BaseConfig, load, config_path +from lilo.deployments import Deployment, load, config_path with backend_runtime_imports(): from lilo.backends.megatron_runtime.common import modeling @@ -25,18 +25,17 @@ def test_config_overrides_reach_megatron(monkeypatch): data["trainer"]["config"]["distributed_overrides"] = {"native_ddp_setting": 123} data["trainer"]["config"]["provider_overrides"]["native_provider_setting"] = [1, 2] config, _ = parse_backend_config( - backend_config(TypeAdapter(BaseConfig).validate_python(data)) + backend_config(TypeAdapter(Deployment).validate_python(data)) ) # These stand in for an installed upstream version with extra fields. The # deployment reader must not need its own list of those fields. - provider = SimpleNamespace( - native_provider_setting=None, - mtp_num_layers=0, - recompute_granularity=None, - recompute_method=None, - recompute_num_layers=None, - provide_distributed_model=Mock(return_value="model"), - ) + # Model providers in Megatron Bridge are dataclasses. + fields = { + **modeling.provider_settings(config, "bf16"), + "provide_distributed_model": Mock(return_value="model"), + } + Provider = make_dataclass("Provider", [(name, object) for name in fields]) + provider = Provider(**fields) bridge = SimpleNamespace(to_megatron_provider=lambda: provider) monkeypatch.setattr( modeling, @@ -50,6 +49,7 @@ def test_config_overrides_reach_megatron(monkeypatch): monkeypatch.setattr(modeling, "DistributedDataParallelConfig", ddp_constructor) _, actual_provider, _ = modeling.model_provider(config) assert actual_provider.native_provider_setting == [1, 2] + assert actual_provider is not provider assert actual_provider.tensor_model_parallel_size == 2 optimizer = modeling.optimizer_config(config, "bf16", distributed_optimizer=True) assert optimizer["native_optimizer_setting"] is False @@ -62,5 +62,5 @@ def test_config_overrides_reach_megatron(monkeypatch): assert ddp_constructor.call_args.kwargs["use_distributed_optimizer"] is True # An unsupported native field is the installed backend's error at startup. config.provider_overrides["unknown_field"] = True - with pytest.raises(ValueError, match="unknown Megatron provider override"): + with pytest.raises(TypeError, match="unknown_field"): modeling.model_provider(config) diff --git a/tests/providers/test_deployment_apps.py b/tests/providers/test_deployment_apps.py index 4fbfefc..501a2ff 100644 --- a/tests/providers/test_deployment_apps.py +++ b/tests/providers/test_deployment_apps.py @@ -1,3 +1,4 @@ +from dataclasses import replace import asyncio import json from types import SimpleNamespace @@ -6,7 +7,7 @@ import pytest from lilo.deployments import load, config_path, DeploymentRecord -from lilo.providers.modal import deployment_apps +from lilo.providers.modal import deployment_apps, deployment_records from lilo.providers.modal.fft_pool import FFTPoolSpec from lilo.providers.modal.lora_pool import LoraPoolSpec @@ -65,16 +66,15 @@ def test_trainer_declaration_and_executor_configuration( assert declaration["image"] is image calls = [] reloaded = [] - from lilo.providers.modal import serve, kv - monkeypatch.setattr(kv, "shared_kv", lambda: "store") + monkeypatch.setattr(deployment_apps, "shared_kv", lambda: "store") monkeypatch.setattr( deployment_apps, "volumes_for", lambda spec: {"/assets": SimpleNamespace(reload=lambda: reloaded.append(True))}, ) monkeypatch.setattr( - serve, + deployment_apps, "run_engine_with_backend", lambda *args, **kwargs: calls.append((args, kwargs)), ) @@ -86,7 +86,7 @@ def test_trainer_declaration_and_executor_configuration( assert kwargs["backend_env"]["LILO_CHECKPOINT_VOLUME"] == "test-custom-checkpoints" assert kwargs["backend_env"]["LILO_BASE_MODEL_REVISION"] == "a" * 40 config = json.loads(kwargs["backend_env"]["LILO_BACKEND_CONFIG"]) - assert config[row.spec.trainer["backend"]]["hf_checkpoint"] == row.asset_path + assert config[row.spec.trainer.backend]["hf_checkpoint"] == row.asset_path assert config["checkpoint_dir"] == "/checkpoints" assert reloaded == [True] @@ -106,11 +106,10 @@ def test_pool_starts_native_server_and_correct_sidecar(builders, monkeypatch, ki app, server = deployment_apps.build_rollout_app(row, pool, image="test-image") settings, _ = app.servers["Server"] assert app.name == pool.app_name - assert settings["gpu"] == row.spec.inference["resources"]["gpu"] + assert settings["gpu"] == row.spec.inference.compute.modal_gpu assert settings["min_containers"] == 0 assert settings["target_concurrency"] == 16 assert settings["compute_region"] == "us-west" - from lilo.inference import serving import subprocess calls, commands, stops = [], [], [] @@ -118,25 +117,25 @@ def test_pool_starts_native_server_and_correct_sidecar(builders, monkeypatch, ki monkeypatch.setattr( subprocess, "Popen", lambda argv, **kw: (commands.append(argv) or process) ) - monkeypatch.setattr(serving, "wait_http", lambda *args: None) - monkeypatch.setattr(serving, "supervise_children", lambda *args: None) + monkeypatch.setattr(deployment_apps, "wait_http", lambda *args: None) + monkeypatch.setattr(deployment_apps, "supervise_children", lambda *args: None) monkeypatch.setattr( - serving, + deployment_apps, "start_lora_sidecar", lambda **kw: (calls.append(("lora", kw)) or process), ) monkeypatch.setattr( - serving, + deployment_apps, "start_fft_sidecar", lambda **kw: (calls.append(("fft", kw)) or process), ) - monkeypatch.setattr(serving, "terminate", stops.append) + monkeypatch.setattr(deployment_apps, "terminate", stops.append) replica = server() replica.start() - assert commands[0][2] == "lilo.inference.native_sglang" + assert commands[0][2] == "lilo.inference.sglang" assert commands[0][3] == row.asset_path native = json.loads(commands[0][4]) - assert native["context_length"] == row.spec.model["max_context_length"] + assert native["context_length"] == row.spec.model.max_context_length if kind == "lora": assert native["enable_lora"] is True assert native["max_lora_rank"] == 32 @@ -161,15 +160,16 @@ def test_pool_starts_native_server_and_correct_sidecar(builders, monkeypatch, ki def test_pool_subprocess_receives_recorded_generation(monkeypatch): row = deployment() - monkeypatch.setenv(deployment_apps.MANIFEST_ENV, json.dumps([row.model_dump()])) - env = deployment_apps.pool_environment(row.definition_id) + monkeypatch.setenv(deployment_records.MANIFEST_ENV, json.dumps([row.model_dump()])) + env = deployment_records.pool_environment(row.definition_id) assert ( - json.loads(env[deployment_apps.POOL_CONFIG_ENV])["generation"] == row.generation + json.loads(env[deployment_records.POOL_CONFIG_ENV])["generation"] + == row.generation ) with pytest.raises(ValueError, match="missing recorded"): - deployment_apps.pool_environment("yaml_missing_123") + deployment_records.pool_environment("yaml_missing_123") with pytest.raises(ValueError, match="missing recorded"): - deployment_apps.pool_environment("unconfigured-python-definition") + deployment_records.pool_environment("unconfigured-python-definition") def test_startup_failure_is_visible_and_blocks_new_spawns(monkeypatch): @@ -199,14 +199,17 @@ def test_real_modal_app_constructs_from_manifest_without_legacy_catalog(monkeypa import sys row = deployment() - env = {**os.environ, deployment_apps.MANIFEST_ENV: json.dumps([row.model_dump()])} + env = { + **os.environ, + deployment_records.MANIFEST_ENV: json.dumps([row.model_dump()]), + } result = subprocess.run( [ sys.executable, "-c", """ import importlib, sys, modal -from lilo.providers.modal import deployment_apps +from lilo.providers.modal import deployment_apps, deployment_records deployment_apps.image_for = lambda backend: modal.Image.debian_slim() app = importlib.import_module('lilo.providers.modal.app') assert len(app.DEFINITIONS) == 1 @@ -232,12 +235,20 @@ def test_admission_changes_preserve_serialized_trainer(builders): old_bytes = serialize(deployment_apps.build_trainer_app(first, image="test")[1]) changed = first.model_copy(deep=True) changed.active = False - changed.spec.routing["default"] = not first.spec.routing["default"] - changed.spec.routing["sampling_default"] = True + changed.spec = replace( + changed.spec, + routing=replace(changed.spec.routing, default=False, sampling_default=True), + ) new_bytes = serialize(deployment_apps.build_trainer_app(changed, image="test")[1]) assert new_bytes == old_bytes - assert first.active is True and first.spec.routing["default"] is True - changed.spec.trainer["resources"]["gpu"] = "H200:4" + assert first.active is True and first.spec.routing.default is True + changed.spec = replace( + changed.spec, + trainer=replace( + changed.spec.trainer, + compute=replace(changed.spec.trainer.compute, gpu="H200"), + ), + ) assert ( serialize(deployment_apps.build_trainer_app(changed, image="test")[1]) != old_bytes @@ -249,7 +260,7 @@ def test_pool_launch_uses_only_generic_yaml_app(monkeypatch, kind): from lilo.providers.modal import fft_pool, lora_pool row = deployment("qwen35-9b-lora-16k" if kind == "lora" else "qwen35-4b-fft-64k") - monkeypatch.setenv(deployment_apps.MANIFEST_ENV, json.dumps([row.model_dump()])) + monkeypatch.setenv(deployment_records.MANIFEST_ENV, json.dumps([row.model_dump()])) module = lora_pool if kind == "lora" else fft_pool spec = ( LoraPoolSpec(row.definition_id) @@ -281,7 +292,7 @@ def gateway_url(self): command[command.index("-m") + 1] == "lilo.providers.modal.deployment_pool_app" ) assert ( - json.loads(kwargs["env"][deployment_apps.POOL_CONFIG_ENV])["generation"] + json.loads(kwargs["env"][deployment_records.POOL_CONFIG_ENV])["generation"] == row.generation ) @@ -309,7 +320,7 @@ def test_missing_pool_uses_saved_provisioner(monkeypatch, kind): from lilo.providers.modal import fft_pool, lora_pool row = deployment("qwen35-9b-lora-16k" if kind == "lora" else "qwen35-4b-fft-64k") - monkeypatch.setenv(deployment_apps.MANIFEST_ENV, json.dumps([row.model_dump()])) + monkeypatch.setenv(deployment_records.MANIFEST_ENV, json.dumps([row.model_dump()])) module = lora_pool if kind == "lora" else fft_pool spec = ( LoraPoolSpec(row.definition_id) @@ -344,15 +355,13 @@ def gateway_url(self): def test_provisioner_rejects_wrong_settings_and_uses_saved_record( builders, monkeypatch ): - from lilo.providers.modal import lora_pool - row = deployment() app, provision = deployment_apps.build_inference_app(row, image="test") assert app.name == row.inference_app_name calls = [] monkeypatch.setattr( - lora_pool, - "deploy_pool", + deployment_apps, + "deploy_lora", lambda pool, *, record: calls.append((pool, record)) or "https://pool", ) pool = LoraPoolSpec(row.definition_id) @@ -382,7 +391,7 @@ def test_real_worker_entrypoint_constructs_offline(monkeypatch, role): "-c", """ import modal -from lilo.providers.modal import deployment_apps +from lilo.providers.modal import deployment_apps, deployment_records deployment_apps.image_for = lambda backend: modal.Image.debian_slim() import lilo.providers.modal.deployment_worker_app as worker assert worker.app.name.startswith("lilo-") @@ -420,3 +429,99 @@ async def no_error(definition_id): monkeypatch.setattr(app, "module_for", lambda _: definition) assert asyncio.run(app._spawn_engine(row.definition_id, "instance")) == "call-id" assert calls == [("instance", row.model_dump_json())] + + +def test_declared_compute_settings_reach_modal(builders): + base = deployment().spec + spec = replace( + base, + trainer=replace( + base.trainer, + timeout_s=90, + compute=replace(base.trainer.compute, cpu=12, memory_mib=123456), + ), + inference=replace( + base.inference, + startup_timeout_s=90, + min_replicas=1, + max_replicas=3, + compute=replace(base.inference.compute, cpu=6, memory_mib=45000), + ), + ) + row = DeploymentRecord.create(spec, revision="a" * 40) + trainer_app, _ = deployment_apps.build_trainer_app(row, image="test") + trainer, _ = trainer_app.functions["trainer"] + assert (trainer["cpu"], trainer["memory"], trainer["timeout"]) == (12, 123456, 90) + pool_app, _ = deployment_apps.build_rollout_app( + row, LoraPoolSpec(row.definition_id), image="test" + ) + server, _ = pool_app.servers["Server"] + assert (server["cpu"], server["memory"], server["startup_timeout"]) == ( + 6, + 45000, + 90, + ) + assert (server["min_containers"], server["max_containers"]) == (1, 3) + + +def test_multinode_trainer_uses_cluster_launcher(builders, monkeypatch): + row = deployment("qwen38-27b-lora-256k") + clusters = [] + + def clustered(nodes, *, rdma): + clusters.append((nodes, rdma)) + + def decorate(fn): + return fn + + return decorate + + monkeypatch.setattr(modal.experimental, "clustered", clustered) + app, _ = deployment_apps.build_trainer_app(row, image="test") + settings, _ = app.functions["trainer"] + assert clusters == [(2, True)] + assert settings["gpu"] == "H200:8" + assert settings["experimental_options"] == {"efa_enabled": True} + calls = [] + monkeypatch.setattr(deployment_apps, "shared_kv", lambda: "store") + monkeypatch.setattr( + deployment_apps, + "volumes_for", + lambda _: {"/assets": SimpleNamespace(reload=lambda: None)}, + ) + monkeypatch.setattr( + deployment_apps, + "start_trainer_cluster", + lambda nodes, **kwargs: "10.0.0.1:6379", + ) + monkeypatch.setattr( + deployment_apps, "run_engine_with_backend", lambda *a, **kw: calls.append(kw) + ) + deployment_apps.run_trainer(row, "instance") + assert calls[0]["backend_env"]["LILO_RAY_ADDRESS"] == "10.0.0.1:6379" + assert ( + json.loads(calls[0]["backend_env"]["LILO_BACKEND_CONFIG"])["miles"][ + "actor_num_nodes" + ] + == 2 + ) + monkeypatch.setattr(deployment_apps, "start_trainer_cluster", lambda *a, **k: None) + deployment_apps.run_trainer(row, "worker") + assert len(calls) == 1 + + +def test_launchers_do_not_reparse_backend_config(builders, monkeypatch): + import lilo.backends.deployment as backend + + row = deployment() + + def unexpected(*args, **kwargs): + pytest.fail("launcher must use the saved resolved settings") + + monkeypatch.setattr(backend, "backend_config", unexpected) + monkeypatch.setattr(backend, "serving_options", unexpected) + deployment_apps.build_trainer_app(row, image="test") + deployment_apps.build_rollout_app( + row, LoraPoolSpec(row.definition_id), image="test" + ) + deployment_apps.definition_from_spec(row, register_trainer=False) diff --git a/tests/providers/test_deployment_e2e_helper.py b/tests/providers/test_deployment_e2e_helper.py index 1f50bbe..3beda47 100644 --- a/tests/providers/test_deployment_e2e_helper.py +++ b/tests/providers/test_deployment_e2e_helper.py @@ -5,7 +5,7 @@ import modal import pytest -from lilo.deployments import load, config_path, DeploymentRecord, gpu_count +from lilo.deployments import load, config_path, DeploymentRecord @pytest.mark.parametrize("preset", ["qwen35-9b-lora-16k", "qwen35-4b-fft-64k"]) @@ -21,9 +21,9 @@ def test_e2e_helper_reads_active_deployed_configuration(monkeypatch, preset): monkeypatch.setattr(modal.Dict, "from_name", lambda name: registry) definition, mode = helper["_definition"]("test-frontend", row.spec.name) assert definition.DEFINITION_ID == row.definition_id - assert definition.MAX_CONTEXT_LENGTH == row.spec.model["max_context_length"] - assert definition.GPUS == gpu_count(row.spec.trainer["resources"]) - assert mode == row.spec.model["parameterization"] + assert definition.MAX_CONTEXT_LENGTH == row.spec.model.max_context_length + assert definition.GPUS == row.spec.trainer.compute.gpus_per_node + assert mode == row.spec.model.parameterization assert definition.MAX_TOKENS_PER_MICROBATCH > 0 with pytest.raises(ValueError, match="one active YAML"): helper["_definition"]("test-frontend", "missing") diff --git a/tests/providers/test_deployment_presets.py b/tests/providers/test_deployment_presets.py index facaeac..eed56e5 100644 --- a/tests/providers/test_deployment_presets.py +++ b/tests/providers/test_deployment_presets.py @@ -4,10 +4,10 @@ from lilo.deployments import load, config_path, DeploymentRecord from lilo.backends.deployment import backend_config, serving_options +from lilo.providers.modal.deployment_records import manifest_from_env from lilo.providers.modal.deployment_apps import ( definition_from_spec, frontend_settings, - manifest_from_env, ) @@ -19,7 +19,7 @@ def test_all_packaged_recipes_validate_offline(path): spec = load(path) config = backend_config(spec) - assert config[spec.trainer["backend"]]["hf_checkpoint"] == "/assets/pending" + assert config[spec.trainer.backend]["hf_checkpoint"] == "/assets/pending" serving_options(spec) @@ -56,8 +56,8 @@ def test_qwen38_context_parallel_token_budget(context, cp): def test_single_client_recipe_keeps_shared_backend_capacity(): shared = load(config_path("qwen35-9b-lora-16k")) single = load(config_path("qwen35-9b-lora-16k-single")) - assert single.trainer["engine"]["max_clients_per_instance"] == 1 - assert single.trainer["resources"] == shared.trainer["resources"] + assert single.trainer.max_clients_per_instance == 1 + assert single.trainer.compute == shared.trainer.compute assert backend_config(single) == backend_config(shared) @@ -82,12 +82,12 @@ def test_missing_manifest_has_no_python_catalog_fallback(monkeypatch, value): ], ) def test_invalid_attention_parallelism_is_rejected(options): - from lilo.deployments import BaseConfig + from lilo.deployments import Deployment data = asdict(load(config_path("qwen35-35b-a3b-fft-64k"))) data["inference"]["config"].update(options) from lilo.backends.deployment import serving_options - spec = TypeAdapter(BaseConfig).validate_python(data) + spec = TypeAdapter(Deployment).validate_python(data) with pytest.raises(ValueError, match="sglang"): serving_options(spec) diff --git a/tests/providers/test_sglang_entrypoint.py b/tests/providers/test_sglang_entrypoint.py new file mode 100644 index 0000000..f10c770 --- /dev/null +++ b/tests/providers/test_sglang_entrypoint.py @@ -0,0 +1,48 @@ +"""The SGLang worker receives its resolved constructor options without argparse.""" + +import json +import runpy +import sys +from types import ModuleType + +from lilo.deployments import DeploymentRecord, config_path, load + + +def test_sglang_constructor_receives_recorded_settings(monkeypatch): + row = DeploymentRecord.create( + load(config_path("qwen35-9b-lora-16k")), revision="a" * 40 + ) + calls = [] + + class ServerArgs: + def __init__(self, **kwargs): + calls.append(kwargs) + self.log_level = "info" + + modules = { + "sglang": {"__path__": []}, + "sglang.srt": {"__path__": []}, + "sglang.launch_server": {"run_server": lambda args: calls.append("run")}, + "sglang.srt.plugins": {"load_plugins": lambda: calls.append("plugins")}, + "sglang.srt.server_args": {"ServerArgs": ServerArgs}, + "sglang.srt.utils": {"kill_process_tree": lambda *a, **k: calls.append("stop")}, + } + for name, attrs in modules.items(): + module = ModuleType(name) + module.__dict__.update(attrs) + monkeypatch.setitem(sys.modules, name, module) + monkeypatch.setattr( + sys, "argv", ["sglang", row.asset_path, json.dumps(row.inference_settings)] + ) + runpy.run_module("lilo.inference.sglang", run_name="__main__") + assert calls == [ + "plugins", + { + "model_path": row.asset_path, + "host": "127.0.0.1", + "port": 8001, + **row.inference_settings, + }, + "run", + "stop", + ] diff --git a/tests/test_deployment_cli.py b/tests/test_deployment_cli.py index 43d4d26..e3f59e7 100644 --- a/tests/test_deployment_cli.py +++ b/tests/test_deployment_cli.py @@ -1,4 +1,5 @@ from copy import deepcopy +from dataclasses import replace import json import subprocess @@ -7,7 +8,7 @@ from lilo import deployment_cli as cli from lilo.deployments import load, config_path, DeploymentRecord -from lilo.providers.modal.deployment_apps import MANIFEST_ENV +from lilo.providers.modal.deployment_records import MANIFEST_ENV class Registry(dict): @@ -72,7 +73,7 @@ def fail(*args, **kwargs): assert registry["pending"][0]["generation"] == row.generation assert "manifest" not in registry and "apply_lock" not in registry new_spec = deepcopy(row.spec) - new_spec.trainer["scaling"]["max_instances"] = 2 + new_spec = replace(new_spec, trainer=replace(new_spec.trainer, max_instances=2)) new = DeploymentRecord.create(new_spec, revision="a" * 40) monkeypatch.setattr(subprocess, "run", lambda *args, **kwargs: None) cli.deploy([new]) @@ -123,15 +124,15 @@ def test_compile_pins_revision_at_external_boundary( path = tmp_path / "model.py" path.write_text( - "from lilo.configs.qwen35_9b_lora_16k import Config as ParentConfig\n" - "class Config(ParentConfig):\n" - f" overrides = {{'model.revision': {revision!r}}}\n" + "from dataclasses import replace\n" + "from lilo.configs.qwen35_9b_lora_16k import config as base\n" + f"config = replace(base, model=replace(base.model, revision={revision!r}))\n" ) lookup = Mock(return_value=SimpleNamespace(sha="a" * 40)) monkeypatch.setattr(huggingface_hub.HfApi, "model_info", lookup) monkeypatch.setattr(miles_revision, "resolve_miles_commit", lambda: "b" * 40) (row,) = cli.compile_configs([path]) - assert row.spec.model["revision"] == "a" * 40 + assert row.spec.model.revision == "a" * 40 assert lookup.call_count == lookups if lookups: lookup.return_value.sha = None @@ -169,7 +170,7 @@ def run(command, **kwargs): assert calls == ["frontend"] changed = deepcopy(row.spec) - changed.inference["scaling"]["max_replicas"] = 6 + changed = replace(changed, inference=replace(changed.inference, max_replicas=6)) new = DeploymentRecord.create(changed, revision="a" * 40) calls.clear() cli.deploy([new]) @@ -328,5 +329,5 @@ def test_builtin_config_resolves_revision_automatically(monkeypatch): ) (row,) = cli.compile_configs([config_path("qwen35-9b-lora-16k")]) assert calls == [("Qwen/Qwen3.5-9B-Base", "main")] - assert row.spec.model["revision"] == "a" * 40 - assert "revision" not in load(config_path("qwen35-9b-lora-16k")).model + assert row.spec.model.revision == "a" * 40 + assert load(config_path("qwen35-9b-lora-16k")).model.revision == "main" diff --git a/tests/test_deployments.py b/tests/test_deployments.py index e7fbea4..424d38f 100644 --- a/tests/test_deployments.py +++ b/tests/test_deployments.py @@ -1,4 +1,4 @@ -from dataclasses import asdict +from dataclasses import asdict, replace, FrozenInstanceError from pydantic import TypeAdapter import argparse import asyncio @@ -9,7 +9,7 @@ from lilo.deployments import ( DeploymentRecord, - BaseConfig, + Deployment, load, config_path, validate_frontend, @@ -18,7 +18,7 @@ from lilo.control_plane.deployments import DeploymentRoutes from lilo.backends.deployment import backend_config from lilo.providers.modal.deployment_apps import definition_from_spec -from lilo.argparse_config import apply_config_overrides +from lilo.backends.miles_arguments import apply_config_overrides def recipe(preset="qwen35-9b-lora-16k", **changes): @@ -29,7 +29,7 @@ def recipe(preset="qwen35-9b-lora-16k", **changes): for key in keys[:-1]: target = target[key] target[keys[-1]] = value - return TypeAdapter(BaseConfig).validate_python(data) + return TypeAdapter(Deployment).validate_python(data) def resolved(spec=None, **changes): @@ -51,7 +51,7 @@ def test_presets_context_topology_and_backend_options(): assert config["cli_options"]["recompute_num_layers"] == 1 assert config["extra_args"] == ("--seq-length", "16384") large = recipe("qwen35-9b-lora-64k") - assert large.model["max_context_length"] == 65536 + assert large.model.max_context_length == 65536 assert backend_config(large)["miles"]["actor_num_gpus_per_node"] == 8 fft = backend_config(recipe("qwen35-4b-fft-64k"))["megatron"] assert (fft["tensor_model_parallel_size"], fft["context_parallel_size"]) == (2, 2) @@ -82,7 +82,7 @@ def test_no_model_catalog_required(): ), ({"inference__config": {"model_path": "other"}}, "managed"), ({"inference__config": {"tp_size": 2}}, "replica GPU"), - ({"trainer__engine__max_clients_per_instance": 7}, "max_lora_slots"), + ({"trainer__max_clients_per_instance": 7}, "max_lora_slots"), ( { "inference__config": { @@ -106,7 +106,7 @@ def test_invalid_integrations_fail_when_building_backend_settings(changes, match def test_generation_and_asset_paths_include_exact_base(): a = resolved() assert a.generation == resolved(recipe(routing__default=False)).generation - assert a.generation != resolved(recipe(trainer__resources__gpu="H200:4")).generation + assert a.generation != resolved(recipe(trainer__compute__gpu="H200")).generation b = resolved(recipe(model__id="other/Qwen3.5-9B-Base")) assert a.asset_path != b.asset_path assert a.asset_path != DeploymentRecord.create(a.spec, revision="b" * 40).asset_path @@ -117,16 +117,14 @@ def test_frontend_defaults_and_retained_generations(): large = resolved(recipe("qwen35-9b-lora-64k")) routes = DeploymentRoutes([definition(small), definition(large)]) assert ( - routes.select(small.spec.model["id"], "lora").DEFINITION_ID - == small.definition_id + routes.select(small.spec.model.id, "lora").DEFINITION_ID == small.definition_id ) switched = retain_generations( [small, large], [resolved(recipe("qwen35-9b-lora-64k", routing__default=True))] ) routes = DeploymentRoutes(map(definition, switched)) assert ( - routes.select(small.spec.model["id"], "lora").DEFINITION_ID - == large.definition_id + routes.select(small.spec.model.id, "lora").DEFINITION_ID == large.definition_id ) # Saved model/checkpoint records continue using their original definition. assert ( @@ -146,20 +144,22 @@ def test_ambiguous_model_does_not_get_random_configuration(): ] routes = DeploymentRoutes(map(definition, rows)) with pytest.raises(ValueError, match="ambiguous.*16k.*64k"): - routes.select(rows[0].spec.model["id"], "lora") + routes.select(rows[0].spec.model.id, "lora") assert routes.capabilities() == [] def test_sampling_requires_default_across_training_modes(): lora = resolved() - fft = resolved(recipe("qwen35-4b-fft-64k", model__id=lora.spec.model["id"])) + fft = resolved(recipe("qwen35-4b-fft-64k", model__id=lora.spec.model.id)) routes = DeploymentRoutes(map(definition, [lora, fft])) with pytest.raises(ValueError, match="sampling_default"): - routes.sampling(lora.spec.model["id"]) - fft.spec.routing["sampling_default"] = True + routes.sampling(lora.spec.model.id) + fft.spec = replace( + fft.spec, routing=replace(fft.spec.routing, sampling_default=True) + ) assert ( DeploymentRoutes(map(definition, [lora, fft])) - .sampling(lora.spec.model["id"]) + .sampling(lora.spec.model.id) .DEFINITION_ID == fft.definition_id ) @@ -217,7 +217,7 @@ async def run(): json={ "session_id": session, "model_seq_id": seq, - "base_model": row.spec.model["id"], + "base_model": row.spec.model.id, "lora_config": {"rank": 32}, }, ) @@ -280,11 +280,11 @@ def test_native_sections_survive_serialization_without_allowlist(): "future_optimizer_option": 0.125 } data["trainer"]["config"]["distributed_overrides"] = {"future_ddp_option": False} - spec = TypeAdapter(BaseConfig).validate_python(data) + spec = TypeAdapter(Deployment).validate_python(data) settings = backend_config(spec, "/assets/pinned") config, _ = parse_backend_config(json.loads(json.dumps(settings))) assert config.hf_checkpoint == "/assets/pinned" - assert config.seq_length == spec.model["max_context_length"] + assert config.seq_length == spec.model.max_context_length assert config.provider_overrides["future_provider_option"] == { "layers": [1, 4], "enabled": False, @@ -312,9 +312,8 @@ def test_megatron_cli_options_preserve_integration_contract(section, options, ma def test_backend_dispatch_rejects_unknown_backend(): - spec = recipe(trainer__backend="missing") - with pytest.raises(ValueError, match="unknown deployment backend"): - backend_config(spec) + with pytest.raises(ValueError, match="backend"): + recipe(trainer__backend="missing") def test_new_miles_and_sglang_options_need_no_deployment_schema_change(): @@ -332,17 +331,15 @@ def test_new_miles_and_sglang_options_need_no_deployment_schema_change(): def test_fft_capacity_is_checked_by_backend_setup(): - spec = recipe("qwen35-4b-fft-64k", trainer__engine__max_clients_per_instance=2) with pytest.raises(ValueError, match="FFT trainers admit one client"): - backend_config(spec) + recipe("qwen35-4b-fft-64k", trainer__max_clients_per_instance=2) def test_reserved_environment_is_checked_by_modal_setup(): from lilo.providers.modal.deployment_apps import deployment_env - spec = recipe(trainer__env={"LILO_BACKEND_CONFIG": "oops"}) with pytest.raises(ValueError, match="managed"): - deployment_env(spec.trainer["env"]) + recipe(trainer__env={"LILO_BACKEND_CONFIG": "oops"}) assert deployment_env({"MY_SETTING": "value"}) == {"MY_SETTING": "value"} @@ -354,7 +351,7 @@ def test_record_creation_copies_without_reparsing(): original = asdict(spec) row = DeploymentRecord.create(spec, revision="a" * 40) assert asdict(spec) == original - assert row.spec.model["revision"] == "a" * 40 + assert row.spec.model.revision == "a" * 40 # The record hash covers settings and the pinned backend dependency, not source. expected = original | {"model": original["model"] | {"revision": "a" * 40}} expected.pop("routing") @@ -368,63 +365,53 @@ def test_record_creation_copies_without_reparsing(): "platform": row.platform, "trainer_release": "initial", "inference_release": "initial", + "trainer_settings": row.trainer_settings, + "inference_settings": row.inference_settings, }, sort_keys=True, ).encode() ).hexdigest() ) - row.spec.trainer["config"]["max_lora_rank"] = 64 - assert spec.trainer["config"]["max_lora_rank"] == 32 + row.spec.trainer.config["max_lora_rank"] = 64 + assert spec.trainer.config["max_lora_rank"] == 32 saved = row.model_dump_json() - assert DeploymentRecord.model_validate_json(saved) == row - + assert DeploymentRecord.model_validate_json(saved).model_dump( + mode="json" + ) == row.model_dump(mode="json") -def test_python_config_inheritance_and_independent_defaults(tmp_path): - from dataclasses import is_dataclass +def test_python_config_composition(tmp_path): path = tmp_path / "model.py" path.write_text( - "from lilo.configs.qwen35_9b_lora_64k import Config as ParentConfig\n" - "class Config(ParentConfig):\n" - " name = 'custom'\n" - " overrides = {'trainer.config.cli_options.new_backend_option': False}\n" - ) - first, second = load(path), load(path) - assert is_dataclass(first) - assert first.name == "custom" - assert first.model["max_context_length"] == 65536 - assert first.trainer["config"]["cli_options"]["new_backend_option"] is False - first.trainer["config"]["target_modules"].append("extra") - assert "extra" not in second.trainer["config"]["target_modules"] - assert ( - "extra" - not in load(config_path("qwen35-9b-lora-16k")).trainer["config"][ - "target_modules" - ] + "from dataclasses import replace\n" + "from lilo.configs.qwen35_9b_lora_16k import config as base\n" + "config = replace(base, name='custom', trainer=replace(base.trainer, " + "compute=replace(base.trainer.compute, memory_mib=123456)))\n" ) + custom = load(path) + original = recipe() + assert custom.name == "custom" + assert custom.trainer.compute.memory_mib == 123456 + assert custom.trainer.config == original.trainer.config + assert original.trainer.compute.memory_mib == 65536 + with pytest.raises(FrozenInstanceError): + custom.trainer.max_instances = 9 -def test_loading_python_config_does_not_call_backend_readers(monkeypatch): - import lilo.backends.deployment as backends - - monkeypatch.setattr( - backends, "backend_config", lambda *a: pytest.fail("backend read") - ) - monkeypatch.setattr( - backends, "serving_options", lambda *a: pytest.fail("serving read") - ) - spec = load(config_path("qwen35-9b-lora-16k")) - assert "revision" not in spec.model - record = DeploymentRecord.create(spec, revision="a" * 40) - assert record.spec.model["revision"] == "a" * 40 - assert "revision" not in spec.model +def test_worker_record_contains_resolved_settings(monkeypatch): + record = resolved() + assert record.trainer_settings["miles"]["actor_num_gpus_per_node"] == 4 + assert record.inference_settings["max_lora_rank"] == 32 + assert DeploymentRecord.model_validate_json(record.model_dump_json()).model_dump( + mode="json" + ) == record.model_dump(mode="json") @pytest.mark.parametrize("source", ["Config = {}", "class Config: pass", "value = 1"]) def test_config_file_must_export_config_subclass(tmp_path, source): path = tmp_path / "model.py" path.write_text(source) - with pytest.raises(ValueError, match="Config subclass"): + with pytest.raises(ValueError, match="Deployment object"): load(path) @@ -444,75 +431,25 @@ def test_no_yaml_config_ingestion(tmp_path): load(tmp_path / "old.yaml") -def test_overrides_inherit_replace_and_copy_values(): - from lilo.configs.qwen35_9b_lora_16k import Config as Example - - class Parent(Example): - overrides = { - "trainer.config.target_modules": ["parent"], - "trainer.config.cli_options.future_option": {"enabled": True}, - "inference.config.max_running_requests": 24, - } - - class Child(Parent): - name = "child" - overrides = { - "trainer.config.target_modules": ["child"], - "trainer.config.cli_options.future_option": {"enabled": False}, - } - - child = Child() - assert child.inference["config"]["max_running_requests"] == 24 - assert child.trainer["config"]["target_modules"] == ["child"] - assert child.trainer["config"]["cli_options"]["future_option"] == {"enabled": False} - child.trainer["config"]["target_modules"].append("changed") - child.trainer["config"]["cli_options"]["future_option"]["enabled"] = True - assert Child.overrides["trainer.config.target_modules"] == ["child"] - assert Child().trainer["config"]["cli_options"]["future_option"] == { - "enabled": False - } - assert Parent().trainer["config"]["target_modules"] == ["parent"] - assert "overrides" not in asdict(child) - assert Child(name="keyword").name == "keyword" - - class Replacement(Parent): - trainer = {"resources": {"gpu": "H200:8"}, "config": {"cli_options": {}}} - overrides = {"trainer.config.cli_options.new_option": 1} - - # A child's complete field replacement wins over its parent's dotted edits. - assert Replacement().trainer["config"] == {"cli_options": {"new_option": 1}} - - @pytest.mark.parametrize( - "path", ["model.missing.value", "trainer.missing.value", "typo"] + "section,key", + [ + ("compute", "memroy_mib"), + ("trainer", "max_instnaces"), + ("trainer", "min_instances"), + ("inference", "timeout_s"), + ("compute", "timeout_s"), + ], ) -def test_override_typos_fail_with_the_path(path): - from lilo.configs.qwen35_9b_lora_16k import Config as Example - - class Config(Example): - overrides = {path: 1} - - with pytest.raises(ValueError, match=path): - Config() - - -def test_plain_sections_fill_defaults_without_sharing_values(): - class Config(BaseConfig): - name = "plain" - model = {"id": "example/model", "max_context_length": 2048} - trainer = {"resources": {"gpu": "H100:4"}, "config": {"future_option": False}} - inference = {"resources": {"gpu": "H200"}} - - first, second = Config(), Config() - assert type(first.model) is dict - assert type(first.trainer) is dict - assert "revision" not in first.model - assert first.trainer["resources"]["cpu"] == 8 - assert first.trainer["config"] == {"future_option": False} - first.inference["scaling"]["max_replicas"] = 2 - first.trainer["env"]["CUSTOM"] = "value" - assert second.inference["scaling"]["max_replicas"] == 8 - assert second.trainer["env"] == {} +def test_orchestration_typos_and_unused_fields_are_rejected(section, key): + base = recipe() + component = { + "compute": base.trainer.compute, + "trainer": base.trainer, + "inference": base.inference, + }[section] + with pytest.raises(ValueError, match=key): + replace(component, **{key: 9}) def test_worker_hashes_cover_only_their_settings(): @@ -548,9 +485,9 @@ def test_examples_only_contain_model_infrastructure(): for path in Path(config_path("qwen35-9b-lora-16k")).parent.glob("qwen*.py"): config = load(path) assert not hasattr(config, "deployment") - assert "revision" not in config.model - assert "runtime_version" not in config.trainer - assert "runtime_version" not in config.inference + assert config.model.revision == "main" + assert "runtime_version" not in asdict(config.trainer) + assert "runtime_version" not in asdict(config.inference) @pytest.mark.parametrize("value", ["invalid", 7]) @@ -561,3 +498,40 @@ def test_backend_parser_validates_configured_types_and_choices(value): apply_config_overrides(parser, {"count": value}, argv) with pytest.raises(SystemExit): parser.parse_args(argv) + + +@pytest.mark.parametrize( + "section,field", + [ + ("optimizer_overrides", "lr"), + ("optimizer_overrides", "adam_eps"), + ("provider_overrides", "calculate_per_token_loss"), + ("provider_overrides", "attention_backend"), + ("distributed_overrides", "overlap_grad_reduce"), + ], +) +def test_managed_backend_values_fail_before_record_creation(section, field): + base = recipe("qwen35-4b-fft-64k") + config = {**base.trainer.config, section: {field: 1}} + candidate = replace(base, trainer=replace(base.trainer, config=config)) + with pytest.raises(ValueError, match=field): + DeploymentRecord.create(candidate, revision="a" * 40) + + +def test_multinode_ownership_and_topology(): + config = load(config_path("qwen38-27b-lora-256k")) + row = DeploymentRecord.create(config, revision="a" * 40) + miles = row.trainer_settings["miles"] + assert miles["actor_num_nodes"] == 2 + assert miles["actor_num_gpus_per_node"] == 8 + assert miles["tensor_model_parallel_size"] == 2 + assert miles["context_parallel_size"] == 8 + invalid = replace( + config, + trainer=replace( + config.trainer, + config={**config.trainer.config, "actor_num_nodes": 3}, + ), + ) + with pytest.raises(ValueError, match="actor_num_nodes"): + DeploymentRecord.create(invalid, revision="a" * 40) From 17f24ea478c1158d8e52a12edc6c2079835a21cd Mon Sep 17 00:00:00 2001 From: kailash Date: Wed, 23 Sep 2026 22:57:31 +0000 Subject: [PATCH 20/27] Remove backend dispatch layers and duplicated preset definitions --- docs/deployment-validation.md | 2 + src/lilo/backends/deployment.py | 143 +++++++++++++++++- src/lilo/backends/megatron_deployment.py | 26 ---- src/lilo/backends/miles_deployment.py | 71 --------- src/lilo/configs/qwen35_9b_fft_64k.py | 50 ++---- src/lilo/configs/qwen35_9b_lora_2k.py | 51 ++----- src/lilo/inference/sglang_deployment.py | 62 -------- .../providers/modal/deployment_records.py | 7 - tests/providers/test_deployment_apps.py | 14 +- 9 files changed, 170 insertions(+), 256 deletions(-) delete mode 100644 src/lilo/backends/megatron_deployment.py delete mode 100644 src/lilo/backends/miles_deployment.py delete mode 100644 src/lilo/inference/sglang_deployment.py diff --git a/docs/deployment-validation.md b/docs/deployment-validation.md index 4fad77e..bd7ce21 100644 --- a/docs/deployment-validation.md +++ b/docs/deployment-validation.md @@ -12,6 +12,8 @@ Trainer timeout, CPU, memory, inference startup timeout, and replica scaling are Validation: **651 CPU tests passed, 1 skipped**. Regression coverage includes misspelled orchestration fields, unsupported settings, optimizer/provider/distributed override collisions, all 15 example configs, saved-record round trips, direct SGLang construction, compute-setting propagation, independent worker updates, and multi-node launcher wiring. No apps were redeployed. GPU backend startup and the new multi-node config have not been live-tested in this revision. +The follow-up reduction removes three adapter modules and their dispatch wrappers, the unused pool-environment helper, and duplicated preset definitions. The two composed presets were compared field-for-field with their previous values. The CPU suite remains at **651 passed, 1 skipped**. + ## Historical validation The sections below describe earlier revisions, including APIs that have since been removed. Their live results do not validate the current implementation. diff --git a/src/lilo/backends/deployment.py b/src/lilo/backends/deployment.py index 2219f39..1612faf 100644 --- a/src/lilo/backends/deployment.py +++ b/src/lilo/backends/deployment.py @@ -1,19 +1,148 @@ -"""Resolve lightweight backend settings once, before creating worker apps.""" +"""Resolve backend settings before launch; reserve fields owned by Lilo.""" -from lilo.backends.megatron_deployment import build_config as megatron_config +from dataclasses import asdict + +from lilo.backends.megatron_config import parse_backend_config from lilo.backends.miles_config import MilesBackendConfig -from lilo.backends.miles_deployment import build_config as miles_config -from lilo.inference.sglang_deployment import build_config as sglang_config +from lilo.config_validation import reject_managed_options + +MILES_MANAGED = { + "context_parallel_size", + "expert_model_parallel_size", + "expert_tensor_parallel_size", + "lora_alpha", + "lora_dropout", + "lora_rank", + "max_tokens_per_gpu", + "multi_lora_n_adapters", + "target_modules", + "tensor_model_parallel_size", + "hf_checkpoint", + "load", + "pretrained_checkpoint", + "train_backend", + "actor_num_nodes", + "actor_num_gpus_per_node", + "rollout_num_gpus", + "debug_train_only", + "megatron_to_hf_mode", + "seq_length", + "pipeline_model_parallel_size", + "virtual_pipeline_model_parallel_size", + "colocate", + "custom_actor", + "sglang_model_path", + "use_dynamic_global_batch_size", + "delay_split_train_data_by_dp", + "use_dynamic_batch_size", + "optimizer", + "gradient_accumulation_fusion", + "save", + "save_interval", + "ckpt_step", + "rollout_num_gpus_per_engine", + "num_gpus_per_node", +} -TRAINERS = {"miles": miles_config, "megatron": megatron_config} + +SGLANG_MANAGED = { + "model_path", + "model", + "host", + "port", + "context_length", + "enable_lora", + "max_lora_rank", + "lora_target_modules", + "enable_cpu_weight_cache", + "api_key", + "pp_size", + "lora_paths", + "dist_init_addr", + "nnodes", + "node_rank", + "tokenizer_path", + "tokenizer_revision", + "revision", + "grpc_mode", + "smg_grpc_mode", + "encoder_only", + "use_ray", + "disaggregation_mode", + "skip_tokenizer_init", +} def backend_config(spec, asset_path="/assets/pending"): - return TRAINERS[spec.trainer.backend](spec, asset_path) + trainer = spec.trainer + settings = trainer.config + if trainer.backend == "megatron": + reject_managed_options(settings, {"hf_checkpoint", "seq_length"}) + config, _ = parse_backend_config( + { + "megatron": { + **settings, + "hf_checkpoint": asset_path, + "seq_length": spec.model.max_context_length, + } + } + ) + if config.optimizer.optimizer != "adam": + raise ValueError("Tinker optim_step requires an Adam optimizer") + config.validate(trainer.compute.gpus_per_node) + return {"megatron": asdict(config), "checkpoint_dir": "/checkpoints"} + reject_managed_options( + settings, + {"hf_checkpoint", "actor_num_gpus_per_node", "actor_num_nodes", "extra_args"}, + ) + reject_managed_options(settings.get("cli_options", {}), MILES_MANAGED) + config = MilesBackendConfig( + hf_checkpoint=asset_path, + actor_num_gpus_per_node=trainer.compute.gpus_per_node, + actor_num_nodes=trainer.compute.nodes, + extra_args=("--seq-length", str(spec.model.max_context_length)), + **settings, + ) + config.validate() + if config.world_size % ( + config.expert_model_parallel_size * config.expert_tensor_parallel_size + ): + raise ValueError("expert parallel sizes must divide the trainer GPU allocation") + if trainer.max_clients_per_instance > config.max_lora_slots: + raise ValueError("max_clients_per_instance exceeds max_lora_slots") + return {"miles": asdict(config), "checkpoint_dir": "/checkpoints"} def serving_options(spec): - return sglang_config(spec) + options = dict(spec.inference.config) + reject_managed_options(options, SGLANG_MANAGED) + tp = options.get("tp_size", spec.inference.compute.gpus_per_node) + ep = options.get("ep_size", 1) + if type(tp) is not int or tp != spec.inference.compute.gpus_per_node: + raise ValueError("sglang.tp_size must equal the replica GPU allocation") + if type(ep) is not int or ep < 1 or tp % ep: + raise ValueError("sglang.ep_size must divide the replica GPU allocation") + dp = options.get("dp_size", 1) + dp_attention = options.get("enable_dp_attention", False) + if type(dp) is not int or dp < 1 or tp % dp: + raise ValueError("sglang.dp_size must divide the replica GPU allocation") + if not isinstance(dp_attention, bool): + raise ValueError("sglang.enable_dp_attention must be a boolean") + if dp > 1 and not dp_attention: + raise ValueError("sglang.dp_size > 1 requires enable_dp_attention") + for key in ( + "max_loaded_loras", + "max_loras_per_batch", + "max_running_requests", + "max_queued_requests", + ): + if key in options and (type(options[key]) is not int or options[key] < 1): + raise ValueError(f"sglang.{key} must be positive") + if not 0 < options.get("mem_fraction_static", 0.8) < 1: + raise ValueError("sglang.mem_fraction_static must be between zero and one") + if options.get("max_loaded_loras", 64) < options.get("max_loras_per_batch", 8): + raise ValueError("max_loaded_loras must be >= max_loras_per_batch") + return options def resolve_backend_settings(spec, asset_path): diff --git a/src/lilo/backends/megatron_deployment.py b/src/lilo/backends/megatron_deployment.py deleted file mode 100644 index 5d99550..0000000 --- a/src/lilo/backends/megatron_deployment.py +++ /dev/null @@ -1,26 +0,0 @@ -"""Construct the existing Megatron training config without a second schema.""" - -from dataclasses import asdict - -from lilo.config_validation import reject_managed_options - -from .megatron_config import parse_backend_config - - -def build_config(spec, asset_path): - trainer = spec.trainer - settings = trainer.config - reject_managed_options(settings, {"hf_checkpoint", "seq_length"}) - config, _ = parse_backend_config( - { - "megatron": { - **settings, - "hf_checkpoint": asset_path, - "seq_length": spec.model.max_context_length, - } - } - ) - if config.optimizer.optimizer != "adam": - raise ValueError("Tinker optim_step requires an Adam optimizer") - config.validate(trainer.compute.gpus_per_node) - return {"megatron": asdict(config), "checkpoint_dir": "/checkpoints"} diff --git a/src/lilo/backends/miles_deployment.py b/src/lilo/backends/miles_deployment.py deleted file mode 100644 index 0d05a55..0000000 --- a/src/lilo/backends/miles_deployment.py +++ /dev/null @@ -1,71 +0,0 @@ -"""Miles deployment integration. Native options are validated by Miles at startup.""" - -from dataclasses import asdict - -from lilo.config_validation import reject_managed_options - -from .miles_config import MilesBackendConfig - -MILES_MANAGED = { - "context_parallel_size", - "expert_model_parallel_size", - "expert_tensor_parallel_size", - "lora_alpha", - "lora_dropout", - "lora_rank", - "max_tokens_per_gpu", - "multi_lora_n_adapters", - "target_modules", - "tensor_model_parallel_size", - "hf_checkpoint", - "load", - "pretrained_checkpoint", - "train_backend", - "actor_num_nodes", - "actor_num_gpus_per_node", - "rollout_num_gpus", - "debug_train_only", - "megatron_to_hf_mode", - "seq_length", - "pipeline_model_parallel_size", - "virtual_pipeline_model_parallel_size", - "colocate", - "custom_actor", - "sglang_model_path", - "use_dynamic_global_batch_size", - "delay_split_train_data_by_dp", - "use_dynamic_batch_size", - "optimizer", - "gradient_accumulation_fusion", - "save", - "save_interval", - "ckpt_step", - "rollout_num_gpus_per_engine", - "num_gpus_per_node", -} - - -def build_config(spec, asset_path): - """Pass trainer.config directly to MilesBackendConfig.""" - trainer = spec.trainer - settings = trainer.config - reject_managed_options( - settings, - {"hf_checkpoint", "actor_num_gpus_per_node", "actor_num_nodes", "extra_args"}, - ) - reject_managed_options(settings.get("cli_options", {}), MILES_MANAGED) - config = MilesBackendConfig( - hf_checkpoint=asset_path, - actor_num_gpus_per_node=trainer.compute.gpus_per_node, - actor_num_nodes=trainer.compute.nodes, - extra_args=("--seq-length", str(spec.model.max_context_length)), - **settings, - ) - config.validate() - if config.world_size % ( - config.expert_model_parallel_size * config.expert_tensor_parallel_size - ): - raise ValueError("expert parallel sizes must divide the trainer GPU allocation") - if trainer.max_clients_per_instance > config.max_lora_slots: - raise ValueError("max_clients_per_instance exceeds max_lora_slots") - return {"miles": asdict(config), "checkpoint_dir": "/checkpoints"} diff --git a/src/lilo/configs/qwen35_9b_fft_64k.py b/src/lilo/configs/qwen35_9b_fft_64k.py index 128dc6d..1d4a5fe 100644 --- a/src/lilo/configs/qwen35_9b_fft_64k.py +++ b/src/lilo/configs/qwen35_9b_fft_64k.py @@ -1,46 +1,22 @@ -from lilo.configuration import Compute, Deployment, Inference, Model, Routing, Trainer +from dataclasses import replace -config = Deployment( +from lilo.configs.qwen35_4b_fft_64k import config as base + +config = replace( + base, name="qwen35-9b-fft-64k", - model=Model( - parameterization="full", id="Qwen/Qwen3.5-9B", max_context_length=65536 - ), - trainer=Trainer( - compute=Compute(gpu="H200", gpus_per_node=4), - backend="megatron", - config={ - "tensor_model_parallel_size": 2, - "context_parallel_size": 2, - "sequence_parallel": True, - "micro_batch_size": 1, - "max_tokens_per_microbatch": 65536, - "defer_fp32_logits": True, - "fp32_lm_head": True, - "use_distributed_optimizer": True, - "provider_overrides": { - "mtp_num_layers": 0, - "recompute_granularity": "full", - "recompute_method": "uniform", - "recompute_num_layers": 1, - }, - "optimizer": {"lr": 0.0001, "min_lr": 0.0001, "loss_scale": 1.0}, - }, + model=replace(base.model, id="Qwen/Qwen3.5-9B"), + trainer=replace( + base.trainer, + compute=replace(base.trainer.compute, gpu="H200"), env={ "PYTORCH_CUDA_ALLOC_CONF": "expandable_segments:True", "TORCHINDUCTOR_COMPILE_THREADS": "1", }, - sampler_persistence_concurrency=1, ), - inference=Inference( - compute=Compute(gpu="H200"), - config={ - "tp_size": 1, - "ep_size": 1, - "mem_fraction_static": 0.85, - "max_running_requests": 32, - "max_queued_requests": 4, - "cpu_weight_cache_max_compile_group_gb": 16, - }, + inference=replace( + base.inference, + compute=replace(base.inference.compute, gpu="H200"), + config={**base.inference.config, "ep_size": 1}, ), - routing=Routing(default=True), ) diff --git a/src/lilo/configs/qwen35_9b_lora_2k.py b/src/lilo/configs/qwen35_9b_lora_2k.py index 60e4a42..3a6c916 100644 --- a/src/lilo/configs/qwen35_9b_lora_2k.py +++ b/src/lilo/configs/qwen35_9b_lora_2k.py @@ -1,47 +1,26 @@ -from lilo.configuration import Compute, Deployment, Inference, Model, Trainer +from dataclasses import replace -config = Deployment( +from lilo.configuration import Routing +from lilo.configs.qwen35_9b_lora_16k import config as base + +config = replace( + base, name="qwen35-9b-lora-2k", - model=Model(id="Qwen/Qwen3.5-9B-Base", max_context_length=2048), - trainer=Trainer( - compute=Compute(gpu="H200", gpus_per_node=4), - config={ - "model_type": "qwen3.5-9B", - "tensor_model_parallel_size": 4, - "target_modules": [ - "linear_qkv", - "linear_proj", - "linear_fc1", - "linear_fc2", - "output_layer", - ], - "max_tokens_per_gpu": 2048, - "max_lora_slots": 4, - "max_lora_rank": 32, - "default_lora_alpha": 32, - "cli_options": { - "recompute_granularity": "full", - "recompute_method": "uniform", - "recompute_num_layers": 1, - }, - }, - env={ - "PYTORCH_CUDA_ALLOC_CONF": "expandable_segments:True", - "TORCHINDUCTOR_COMPILE_THREADS": "1", - }, + model=replace(base.model, max_context_length=2048), + trainer=replace( + base.trainer, + compute=replace(base.trainer.compute, gpu="H200", cpu=8, memory_mib=32768), max_clients_per_instance=4, + config={**base.trainer.config, "max_tokens_per_gpu": 2048, "max_lora_slots": 4}, ), - inference=Inference( - compute=Compute(gpu="H200"), + inference=replace( + base.inference, config={ - "tp_size": 1, + **base.inference.config, "ep_size": 1, - "mem_fraction_static": 0.8, - "max_running_requests": 32, - "max_queued_requests": 8, "max_loaded_loras": 32, - "max_loras_per_batch": 8, "schedule_policy": "lpm", }, ), + routing=Routing(), ) diff --git a/src/lilo/inference/sglang_deployment.py b/src/lilo/inference/sglang_deployment.py deleted file mode 100644 index b0491d4..0000000 --- a/src/lilo/inference/sglang_deployment.py +++ /dev/null @@ -1,62 +0,0 @@ -"""SGLang settings that must agree with Lilo replica orchestration.""" - -from lilo.config_validation import reject_managed_options - -SGLANG_MANAGED = { - "model_path", - "model", - "host", - "port", - "context_length", - "enable_lora", - "max_lora_rank", - "lora_target_modules", - "enable_cpu_weight_cache", - "api_key", - "pp_size", - "lora_paths", - "dist_init_addr", - "nnodes", - "node_rank", - "tokenizer_path", - "tokenizer_revision", - "revision", - "grpc_mode", - "smg_grpc_mode", - "encoder_only", - "use_ray", - "disaggregation_mode", - "skip_tokenizer_init", -} - - -def build_config(spec): - options = dict(spec.inference.config) - reject_managed_options(options, SGLANG_MANAGED) - tp = options.get("tp_size", spec.inference.compute.gpus_per_node) - ep = options.get("ep_size", 1) - if type(tp) is not int or tp != spec.inference.compute.gpus_per_node: - raise ValueError("sglang.tp_size must equal the replica GPU allocation") - if type(ep) is not int or ep < 1 or tp % ep: - raise ValueError("sglang.ep_size must divide the replica GPU allocation") - dp = options.get("dp_size", 1) - dp_attention = options.get("enable_dp_attention", False) - if type(dp) is not int or dp < 1 or tp % dp: - raise ValueError("sglang.dp_size must divide the replica GPU allocation") - if not isinstance(dp_attention, bool): - raise ValueError("sglang.enable_dp_attention must be a boolean") - if dp > 1 and not dp_attention: - raise ValueError("sglang.dp_size > 1 requires enable_dp_attention") - for key in ( - "max_loaded_loras", - "max_loras_per_batch", - "max_running_requests", - "max_queued_requests", - ): - if key in options and (type(options[key]) is not int or options[key] < 1): - raise ValueError(f"sglang.{key} must be positive") - if not 0 < options.get("mem_fraction_static", 0.8) < 1: - raise ValueError("sglang.mem_fraction_static must be between zero and one") - if options.get("max_loaded_loras", 64) < options.get("max_loras_per_batch", 8): - raise ValueError("max_loaded_loras must be >= max_loras_per_batch") - return options diff --git a/src/lilo/providers/modal/deployment_records.py b/src/lilo/providers/modal/deployment_records.py index 586fd8c..a9fedd7 100644 --- a/src/lilo/providers/modal/deployment_records.py +++ b/src/lilo/providers/modal/deployment_records.py @@ -30,13 +30,6 @@ def pool_deployment(definition_id): return None -def pool_environment(definition_id): - resolved = pool_deployment(definition_id) - if resolved is None: - raise ValueError(f"missing recorded deployment: {definition_id}") - return {POOL_CONFIG_ENV: resolved.model_dump_json()} - - def provision_pool(record, pool): """Ask the saved inference app to create a pool using its original code.""" provision = modal.Function.from_name( diff --git a/tests/providers/test_deployment_apps.py b/tests/providers/test_deployment_apps.py index 501a2ff..be3296b 100644 --- a/tests/providers/test_deployment_apps.py +++ b/tests/providers/test_deployment_apps.py @@ -158,18 +158,12 @@ def test_pool_starts_native_server_and_correct_sidecar(builders, monkeypatch, ki assert stops == [process, process] -def test_pool_subprocess_receives_recorded_generation(monkeypatch): +def test_pool_lookup_uses_recorded_generation(monkeypatch): row = deployment() monkeypatch.setenv(deployment_records.MANIFEST_ENV, json.dumps([row.model_dump()])) - env = deployment_records.pool_environment(row.definition_id) - assert ( - json.loads(env[deployment_records.POOL_CONFIG_ENV])["generation"] - == row.generation - ) - with pytest.raises(ValueError, match="missing recorded"): - deployment_records.pool_environment("yaml_missing_123") - with pytest.raises(ValueError, match="missing recorded"): - deployment_records.pool_environment("unconfigured-python-definition") + saved = deployment_records.pool_deployment(row.definition_id) + assert saved.generation == row.generation + assert deployment_records.pool_deployment("yaml_missing_123") is None def test_startup_failure_is_visible_and_blocks_new_spawns(monkeypatch): From 1a3a8e34af32b9e3f8838186f700ddc16540a9cb Mon Sep 17 00:00:00 2001 From: kailash Date: Thu, 24 Sep 2026 16:10:05 +0000 Subject: [PATCH 21/27] Simplify deployment recipes with BaseConfig inheritance --- README.md | 2 +- docs/deployment-configs.md | 89 +++++++------ docs/deployment-validation.md | 22 ++- docs/design.md | 2 +- scripts/definition_smoke.py | 2 +- scripts/e2e_engine_definition.py | 10 +- src/lilo/backends/deployment.py | 20 +-- src/lilo/configs/__init__.py | 2 +- src/lilo/configs/qwen35_35b_a3b_fft_64k.py | 43 +++--- src/lilo/configs/qwen35_4b_fft_64k.py | 40 +++--- src/lilo/configs/qwen35_9b_fft_64k.py | 34 ++--- .../configs/qwen35_9b_instruct_lora_16k.py | 37 +++--- .../qwen35_9b_instruct_lora_16k_dp2.py | 25 ++-- src/lilo/configs/qwen35_9b_lora_16k.py | 39 +++--- src/lilo/configs/qwen35_9b_lora_16k_single.py | 16 +-- src/lilo/configs/qwen35_9b_lora_2k.py | 45 ++++--- src/lilo/configs/qwen35_9b_lora_64k.py | 29 ++-- src/lilo/configs/qwen36_27b_fft_64k.py | 43 +++--- src/lilo/configs/qwen36_35b_a3b_fft_64k.py | 30 +++-- src/lilo/configs/qwen38_27b_lora_128k.py | 44 +++--- src/lilo/configs/qwen38_27b_lora_16k.py | 37 +++--- src/lilo/configs/qwen38_27b_lora_256k.py | 48 +++---- src/lilo/configs/qwen38_27b_lora_64k.py | 36 ++--- src/lilo/configuration.py | 87 ++++++------ src/lilo/deployment_cli.py | 15 ++- src/lilo/deployments.py | 39 ++++-- src/lilo/providers/modal/app.py | 8 +- src/lilo/providers/modal/deployment_apps.py | 60 ++++----- .../providers/modal/deployment_pool_app.py | 2 +- tests/providers/test_deployment_apps.py | 20 ++- tests/providers/test_deployment_e2e_helper.py | 6 +- tests/providers/test_deployment_presets.py | 6 +- tests/test_deployment_cli.py | 12 +- tests/test_deployments.py | 125 ++++++++++++------ 34 files changed, 595 insertions(+), 480 deletions(-) diff --git a/README.md b/README.md index 986cd59..0fb382a 100644 --- a/README.md +++ b/README.md @@ -48,7 +48,7 @@ training = service.create_lora_training_client( ## Shared deployment quick start -Shared deployments use Python dataclasses. See [Python deployment configs](docs/deployment-configs.md). Keep the active Python config list in [scripts/deploy_models.sh](scripts/deploy_models.sh); run it to deploy the complete list. +Shared deployments use Python recipes inheriting from `BaseConfig`. See [Python deployment configs](docs/deployment-configs.md). Keep the active Python config list in [scripts/deploy_models.sh](scripts/deploy_models.sh); run it to deploy the complete list. Install Lilo into your own Python project, deploy it once to Modal, then call its API from your training scripts. The commands below work in Bash or Zsh. diff --git a/docs/deployment-configs.md b/docs/deployment-configs.md index 67f4bdb..f19a7d2 100644 --- a/docs/deployment-configs.md +++ b/docs/deployment-configs.md @@ -1,65 +1,68 @@ -# Model infrastructure configs +# Python deployment configs -A config file exports a config object made from the components in [configuration.py](../src/lilo/configuration.py). These are frozen, validated dataclasses. Unknown fields, invalid types, negative capacities and inconsistent scaling limits fail at construction. Backend option dictionaries and environment dictionaries remain open. +A recipe subclasses `BaseConfig` and exports `config = Config()`. Model settings are ordinary class attributes; trainer and inference settings are dictionaries. No `Compute`, `Model`, or `Routing` constructors are needed. -~~~python -from lilo.configuration import Compute, Deployment, Inference, Model, Trainer +```python +from lilo.configuration import BaseConfig -config = Deployment( - name="my-9b", - model=Model(id="Qwen/Qwen3.5-9B-Base", max_context_length=16384), - trainer=Trainer( - compute=Compute(gpu="H100", gpus_per_node=4, cpu=16, memory_mib=65536), - max_clients_per_instance=6, - config={ + +class Config(BaseConfig): + name = "my-9b" + model = "Qwen/Qwen3.5-9B-Base" + max_context_length = 16384 + + trainer = { + "gpu": "H100", + "gpus_per_node": 4, + "cpu": 16, + "memory_mib": 65536, + "max_clients_per_instance": 6, + "config": { "model_type": "qwen3.5-9B", "tensor_model_parallel_size": 4, "max_lora_slots": 6, "max_lora_rank": 32, }, - ), - inference=Inference( - compute=Compute(gpu="H200"), - max_replicas=8, - startup_timeout_s=1200, - config={"mem_fraction_static": 0.8}, - ), -) -~~~ + } + inference = {"gpu": "H200", "max_replicas": 8} -See the [9B LoRA example](../src/lilo/configs/qwen35_9b_lora_16k.py) and [4B FFT example](../src/lilo/configs/qwen35_4b_fft_64k.py) for complete configurations. No model revision is required; deployment resolves Hugging Face main to an exact commit and records it. An explicit Model(revision=...) is optional. -## Composition +config = Config() +``` -Use standard Python composition and dataclasses.replace: +See the [9B LoRA recipe](../src/lilo/configs/qwen35_9b_lora_16k.py) and [4B FFT recipe](../src/lilo/configs/qwen35_4b_fft_64k.py) for complete examples. Deployment resolves the model's `main` revision to an exact commit; set `revision` only when you want a different revision. -~~~python -from dataclasses import replace -from lilo.configs.qwen35_9b_lora_16k import config as base +## Variants -config = replace( - base, - name="my-9b-more-memory", - trainer=replace( - base.trainer, - compute=replace(base.trainer.compute, memory_mib=98304), - ), -) -~~~ +Use Python inheritance to change a recipe. Extend a section explicitly when you want to keep its other settings: + +```python +from lilo.configs.qwen35_9b_lora_16k import Config as Parent + + +class Config(Parent): + name = "my-9b-more-memory" + trainer = {**Parent.trainer, "memory_mib": 98304} + + +config = Config() +``` + +Assigning a new dictionary replaces that section. There is no implicit deep merge or dotted override language. Extend backend options with `{**Parent.trainer["config"], "max_tokens_per_gpu": 8192}`. Constructor arguments can also override fields: `Config(name="another-run")`. -This retains the other trainer settings. There is no inheritance interpreter, dotted-path override syntax or implicit dictionary merge. Backend dictionaries can be composed explicitly with {**base.trainer.config, "max_tokens_per_gpu": 8192}. Treat configs as values; construct a variant instead of editing an imported object's dictionaries. +Construction copies the recipe's dictionaries and validates Lilo-owned fields. Unknown fields, invalid types, negative capacities, and inconsistent scaling limits fail before deployment. The validated instance has attribute access (`config.trainer.gpu`); backend options stay dictionaries. Instances do not share mutable options with each other or with their recipe class. ## Ownership and validation | Setting | Owner and behavior | | --- | --- | -| Compute | GPU type, GPUs per node, nodes, CPU and memory. Unknown keys such as memroy_mib are rejected. | -| Trainer | Maximum instances/clients, publication concurrency and function timeout. Trainers start on demand; there is no min_instances field. | -| Inference | Replica scaling and startup_timeout_s, passed to the Modal server and startup health checks. There is no unused timeout_s. Each replica uses one node. | +| trainer / inference resources | GPU type, GPUs per node, CPU and memory directly in each section; `nodes` is trainer-only. Unknown keys such as `memroy_mib` are rejected. | +| trainer | Maximum instances/clients, publication concurrency and function timeout. Trainers start on demand; there is no min_instances field. | +| inference | Replica scaling and startup_timeout_s, passed to the Modal server and startup health checks. There is no unused timeout_s. Each replica uses one node. | | trainer.config | Existing MilesBackendConfig or EngineModelConfig fields, plus their explicit extra-option dictionaries. | | inference.config | SGLang ServerArgs fields. Lilo reserves paths, context, topology and adapter settings that must agree with its own configuration. | -Compute topology is configured in Compute; Miles receives actor_num_nodes and actor_num_gpus_per_node from it. Setting those again in backend options is rejected. +Compute topology is configured directly in `trainer`; Miles receives actor_num_nodes and actor_num_gpus_per_node from it. Setting those again in backend options is rejected. Megatron's provider_overrides, optimizer_overrides and distributed_overrides may add backend fields, but may not replace Lilo-owned fields. For example, put the learning rate in optimizer={"lr": ...}; optimizer_overrides={"lr": ...} is rejected. The same settings builders are used during validation and worker construction. The provider is constructed with dataclasses.replace, without an override-by-setattr pass. @@ -70,7 +73,7 @@ Backend libraries validate their own extra options when workers start. Lilo does ## Resolve and launch ~~~text -load(config.py) → config: Deployment +load(config.py) → config: BaseConfig → validate typed compute/scaling/model settings → resolve model commit → resolve_backend_settings(config, asset_path) @@ -83,7 +86,7 @@ The launcher consumes the saved settings. It does not reparse backend configurat | File | Responsibility | | --- | --- | -| [configuration.py](../src/lilo/configuration.py) | Typed components and shared orchestration constraints | +| [configuration.py](../src/lilo/configuration.py) | BaseConfig and validation of model, trainer, and inference settings | | [deployments.py](../src/lilo/deployments.py) | Python object loader, resolved records and config hashes | | [backends/deployment.py](../src/lilo/backends/deployment.py) | Resolve backend settings before launch | | [megatron_runtime/common/settings.py](../src/lilo/backends/megatron_runtime/common/settings.py) | Shared Megatron ownership rules and constructor dictionaries | @@ -118,7 +121,7 @@ The checked-in [deploy_models.sh](../scripts/deploy_models.sh) lists the complet These settings are not model-config fields. Secret/volume names come from provider defaults. Credentials remain in Modal secrets. -One frontend serves all models through Tinker's base_model. Routing(default=True) selects among multiple training configurations for one model; sampling_default=True selects a sampling configuration when LoRA/FFT configurations coexist. +One frontend serves all models through Tinker's base_model. `default = True` selects among multiple training configurations for one model; `sampling_default = True` selects a sampling configuration when LoRA/FFT configurations coexist. ## Hashes and update isolation diff --git a/docs/deployment-validation.md b/docs/deployment-validation.md index bd7ce21..889473d 100644 --- a/docs/deployment-validation.md +++ b/docs/deployment-validation.md @@ -1,8 +1,22 @@ # Deployment configuration validation -These checks exercise PR #55. The current implementation deploys trainers and inference provisioners independently; the shared frontend references them by app name. Earlier sections record validation of the previous shared-app implementation. +These checks exercise PR #55. The current implementation deploys trainers and inference provisioners independently; the shared frontend references them by app name. Historical sections record validation of previous implementations. -## Current: typed config composition and one backend resolution +## Current: simple BaseConfig recipes + +Recipes use one `BaseConfig` subclass with top-level model settings and plain trainer/inference dictionaries. Removed the `Compute`, `Model`, `Routing`, and `Lifecycle` wrapper classes and the matching nested fields. Compute settings live directly in trainer/inference; routing flags and lifecycle timeouts live at the top level. The supported inference backend remains SGLang, so its redundant backend selector was removed. + +Inheritance follows Python attribute replacement. A section is extended explicitly with dictionary unpacking; there is no deep-merge or dotted-override parser. Construction copies class dictionaries and validates against one schema, and workers continue to consume the saved resolved backend settings. + +Validation: **656 CPU tests passed, 1 skipped** (including **122 focused config/deployment tests**). Coverage includes inherited recipes, constructor overrides, nested option isolation, strict field validation, CLI revision resolution, saved-record round trips, and worker construction. The CLI-generated recipe also loaded and validated successfully. Ruff and whitespace checks passed. The full suite ran outside the socket-restricted sandbox so its local HTTP and SDK tests could run. + +All 15 recipes were compared with their pre-change values: compute, model, routing, lifecycle, and resolved trainer/inference settings are unchanged. No deployments or GPU jobs were modified. Existing draft manifests use the previous schema and require a fresh registry or explicit migration. + +## Historical validation + +These results describe earlier APIs and do not validate the current implementation. + +### Previous: typed config composition and one backend resolution Configs export a `Deployment` built from validated dataclasses. Ordinary `dataclasses.replace` composes variants; custom inheritance, dotted overrides, and recursive defaults have been removed. Lilo-owned fields reject unknown keys. Backend tuning remains in explicit dictionaries, with duplicates of Lilo-managed settings rejected before constructor calls. @@ -14,10 +28,6 @@ Validation: **651 CPU tests passed, 1 skipped**. Regression coverage includes mi The follow-up reduction removes three adapter modules and their dispatch wrappers, the unused pool-environment helper, and duplicated preset definitions. The two composed presets were compared field-for-field with their previous values. The CPU suite remains at **651 passed, 1 skipped**. -## Historical validation - -The sections below describe earlier revisions, including APIs that have since been removed. Their live results do not validate the current implementation. - ## Direct backend configuration Removed the deployment-specific Miles field-renaming table and Megatron `runtime/provider/optimizer/distributed` schema. Configs now use the existing `MilesBackendConfig` and `EngineModelConfig` field names. Megatron's existing config reader constructs its nested optimizer; extra Megatron constructor settings are explicit `*_overrides` dictionaries instead of being split by field name. diff --git a/docs/design.md b/docs/design.md index 169d6da..c3cb213 100644 --- a/docs/design.md +++ b/docs/design.md @@ -132,7 +132,7 @@ A Python deployment dataclass specifies the base model, training mode, context l No central model catalog registration is needed. The generic builders in [`deployment_apps.py`](../src/lilo/providers/modal/deployment_apps.py) construct trainer functions and inference apps from the saved config. See [Python deployment configs](deployment-configs.md) for the configuration schema and app structure. -Set each configuration's trainer limit with `trainer.scaling.max_instances` and its inference limits with `inference.scaling`. Trainer limits are read from the config; `LILO_TRAINER_MAX_CONTAINERS` is no longer used. +Set each configuration's trainer limit with `trainer.max_instances` and its inference limits with `inference.min_replicas` and `inference.max_replicas`. Trainer limits are read from the config; `LILO_TRAINER_MAX_CONTAINERS` is no longer used. To check training, publication, and sampling against a deployed configuration: diff --git a/scripts/definition_smoke.py b/scripts/definition_smoke.py index d9c2725..842c339 100644 --- a/scripts/definition_smoke.py +++ b/scripts/definition_smoke.py @@ -314,7 +314,7 @@ def main() -> None: for definition in definitions.values(): print( f"{definition.DEFINITION_ID:40} {definition.PARAMETERIZATION:5} " - f"{definition.RESOLVED.spec.trainer.compute.nodes} nodes x {definition.RESOLVED.spec.trainer.compute.modal_gpu} " + f"{definition.RESOLVED.spec.trainer.nodes} nodes x {definition.RESOLVED.spec.trainer.gpu}:{definition.RESOLVED.spec.trainer.gpus_per_node} " f"ctx={definition.MAX_CONTEXT_LENGTH}" + ("" if definition.CATALOG_VISIBLE else " (not cataloged)") ) diff --git a/scripts/e2e_engine_definition.py b/scripts/e2e_engine_definition.py index 553480f..bb5cbea 100644 --- a/scripts/e2e_engine_definition.py +++ b/scripts/e2e_engine_definition.py @@ -38,15 +38,15 @@ def _definition(frontend: str, name: str) -> tuple[Any, str]: settings = backend_config(spec)[spec.trainer.backend] definition = SimpleNamespace( DEFINITION_ID=resolved.definition_id, - MODEL_NAME=spec.model.id, - PARAMETERIZATION=spec.model.parameterization, - MAX_CONTEXT_LENGTH=spec.model.max_context_length, + MODEL_NAME=spec.model, + PARAMETERIZATION=spec.parameterization, + MAX_CONTEXT_LENGTH=spec.max_context_length, MAX_TOKENS_PER_MICROBATCH=settings.get( "max_tokens_per_microbatch", settings.get("max_tokens_per_gpu") ), MICRO_BATCH_SIZE=settings.get("micro_batch_size", 1), - GPU_TYPE=spec.trainer.compute.modal_gpu.split(":")[0], - GPUS=spec.trainer.compute.gpus_per_node, + GPU_TYPE=spec.trainer.gpu, + GPUS=spec.trainer.gpus_per_node, LORA_RANK=settings.get("max_lora_rank"), ) return definition, definition.PARAMETERIZATION diff --git a/src/lilo/backends/deployment.py b/src/lilo/backends/deployment.py index 1612faf..b1b7338 100644 --- a/src/lilo/backends/deployment.py +++ b/src/lilo/backends/deployment.py @@ -83,13 +83,13 @@ def backend_config(spec, asset_path="/assets/pending"): "megatron": { **settings, "hf_checkpoint": asset_path, - "seq_length": spec.model.max_context_length, + "seq_length": spec.max_context_length, } } ) if config.optimizer.optimizer != "adam": raise ValueError("Tinker optim_step requires an Adam optimizer") - config.validate(trainer.compute.gpus_per_node) + config.validate(trainer.gpus_per_node) return {"megatron": asdict(config), "checkpoint_dir": "/checkpoints"} reject_managed_options( settings, @@ -98,9 +98,9 @@ def backend_config(spec, asset_path="/assets/pending"): reject_managed_options(settings.get("cli_options", {}), MILES_MANAGED) config = MilesBackendConfig( hf_checkpoint=asset_path, - actor_num_gpus_per_node=trainer.compute.gpus_per_node, - actor_num_nodes=trainer.compute.nodes, - extra_args=("--seq-length", str(spec.model.max_context_length)), + actor_num_gpus_per_node=trainer.gpus_per_node, + actor_num_nodes=trainer.nodes, + extra_args=("--seq-length", str(spec.max_context_length)), **settings, ) config.validate() @@ -116,9 +116,9 @@ def backend_config(spec, asset_path="/assets/pending"): def serving_options(spec): options = dict(spec.inference.config) reject_managed_options(options, SGLANG_MANAGED) - tp = options.get("tp_size", spec.inference.compute.gpus_per_node) + tp = options.get("tp_size", spec.inference.gpus_per_node) ep = options.get("ep_size", 1) - if type(tp) is not int or tp != spec.inference.compute.gpus_per_node: + if type(tp) is not int or tp != spec.inference.gpus_per_node: raise ValueError("sglang.tp_size must equal the replica GPU allocation") if type(ep) is not int or ep < 1 or tp % ep: raise ValueError("sglang.ep_size must divide the replica GPU allocation") @@ -148,14 +148,14 @@ def serving_options(spec): def resolve_backend_settings(spec, asset_path): trainer = backend_config(spec, asset_path) inference = { - "context_length": spec.model.max_context_length, - "tp_size": spec.inference.compute.gpus_per_node, + "context_length": spec.max_context_length, + "tp_size": spec.inference.gpus_per_node, "mem_fraction_static": 0.8, "max_running_requests": 32, "weight_loader_disable_mmap": True, **serving_options(spec), } - if spec.model.parameterization == "lora": + if spec.parameterization == "lora": miles = MilesBackendConfig(**trainer["miles"]) inference.update( enable_lora=True, diff --git a/src/lilo/configs/__init__.py b/src/lilo/configs/__init__.py index ad974da..1fe58e9 100644 --- a/src/lilo/configs/__init__.py +++ b/src/lilo/configs/__init__.py @@ -1 +1 @@ -"""Example infrastructure objects; compose variants with dataclasses.replace.""" +"""Example BaseConfig recipes; compose variants with ordinary Python inheritance.""" diff --git a/src/lilo/configs/qwen35_35b_a3b_fft_64k.py b/src/lilo/configs/qwen35_35b_a3b_fft_64k.py index e5dfc23..aaa0e7d 100644 --- a/src/lilo/configs/qwen35_35b_a3b_fft_64k.py +++ b/src/lilo/configs/qwen35_35b_a3b_fft_64k.py @@ -1,14 +1,16 @@ -from lilo.configuration import Compute, Deployment, Inference, Model, Routing, Trainer +from lilo.configuration import BaseConfig -config = Deployment( - name="qwen35-35b-a3b-fft-64k", - model=Model( - parameterization="full", id="Qwen/Qwen3.5-35B-A3B", max_context_length=65536 - ), - trainer=Trainer( - compute=Compute(gpu="H200", gpus_per_node=8), - backend="megatron", - config={ + +class Config(BaseConfig): + name = "qwen35-35b-a3b-fft-64k" + parameterization = "full" + model = "Qwen/Qwen3.5-35B-A3B" + max_context_length = 65536 + trainer = { + "gpu": "H200", + "gpus_per_node": 8, + "backend": "megatron", + "config": { "tensor_model_parallel_size": 4, "pipeline_model_parallel_size": 1, "context_parallel_size": 2, @@ -34,15 +36,16 @@ }, "optimizer": {"optimizer": "adam", "lr": 0.0001, "min_lr": 0.0001}, }, - env={ + "env": { "PYTORCH_CUDA_ALLOC_CONF": "expandable_segments:True", "TORCHINDUCTOR_COMPILE_THREADS": "1", }, - sampler_persistence_concurrency=1, - ), - inference=Inference( - compute=Compute(gpu="H200", gpus_per_node=4), - config={ + "sampler_persistence_concurrency": 1, + } + inference = { + "gpu": "H200", + "gpus_per_node": 4, + "config": { "tp_size": 4, "ep_size": 4, "mem_fraction_static": 0.9, @@ -52,6 +55,8 @@ "dp_size": 4, "enable_dp_attention": True, }, - ), - routing=Routing(default=True), -) + } + default = True + + +config = Config() diff --git a/src/lilo/configs/qwen35_4b_fft_64k.py b/src/lilo/configs/qwen35_4b_fft_64k.py index ce2d827..8bc58d1 100644 --- a/src/lilo/configs/qwen35_4b_fft_64k.py +++ b/src/lilo/configs/qwen35_4b_fft_64k.py @@ -1,14 +1,16 @@ -from lilo.configuration import Compute, Deployment, Inference, Model, Routing, Trainer +from lilo.configuration import BaseConfig -config = Deployment( - name="qwen35-4b-fft-64k", - model=Model( - parameterization="full", id="Qwen/Qwen3.5-4B", max_context_length=65536 - ), - trainer=Trainer( - compute=Compute(gpu="H100", gpus_per_node=4), - backend="megatron", - config={ + +class Config(BaseConfig): + name = "qwen35-4b-fft-64k" + parameterization = "full" + model = "Qwen/Qwen3.5-4B" + max_context_length = 65536 + trainer = { + "gpu": "H100", + "gpus_per_node": 4, + "backend": "megatron", + "config": { "tensor_model_parallel_size": 2, "context_parallel_size": 2, "sequence_parallel": True, @@ -25,17 +27,19 @@ }, "optimizer": {"lr": 0.0001, "min_lr": 0.0001, "loss_scale": 1.0}, }, - sampler_persistence_concurrency=1, - ), - inference=Inference( - compute=Compute(gpu="H100"), - config={ + "sampler_persistence_concurrency": 1, + } + inference = { + "gpu": "H100", + "config": { "tp_size": 1, "mem_fraction_static": 0.85, "max_running_requests": 32, "max_queued_requests": 4, "cpu_weight_cache_max_compile_group_gb": 16, }, - ), - routing=Routing(default=True), -) + } + default = True + + +config = Config() diff --git a/src/lilo/configs/qwen35_9b_fft_64k.py b/src/lilo/configs/qwen35_9b_fft_64k.py index 1d4a5fe..cc0d414 100644 --- a/src/lilo/configs/qwen35_9b_fft_64k.py +++ b/src/lilo/configs/qwen35_9b_fft_64k.py @@ -1,22 +1,22 @@ -from dataclasses import replace +from lilo.configs.qwen35_4b_fft_64k import Config as Parent -from lilo.configs.qwen35_4b_fft_64k import config as base -config = replace( - base, - name="qwen35-9b-fft-64k", - model=replace(base.model, id="Qwen/Qwen3.5-9B"), - trainer=replace( - base.trainer, - compute=replace(base.trainer.compute, gpu="H200"), - env={ +class Config(Parent): + name = "qwen35-9b-fft-64k" + model = "Qwen/Qwen3.5-9B" + trainer = { + **Parent.trainer, + "gpu": "H200", + "env": { "PYTORCH_CUDA_ALLOC_CONF": "expandable_segments:True", "TORCHINDUCTOR_COMPILE_THREADS": "1", }, - ), - inference=replace( - base.inference, - compute=replace(base.inference.compute, gpu="H200"), - config={**base.inference.config, "ep_size": 1}, - ), -) + } + inference = { + **Parent.inference, + "gpu": "H200", + "config": {**Parent.inference["config"], "ep_size": 1}, + } + + +config = Config() diff --git a/src/lilo/configs/qwen35_9b_instruct_lora_16k.py b/src/lilo/configs/qwen35_9b_instruct_lora_16k.py index 44406e6..8cbeabe 100644 --- a/src/lilo/configs/qwen35_9b_instruct_lora_16k.py +++ b/src/lilo/configs/qwen35_9b_instruct_lora_16k.py @@ -1,11 +1,14 @@ -from lilo.configuration import Compute, Deployment, Inference, Model, Routing, Trainer +from lilo.configuration import BaseConfig -config = Deployment( - name="qwen35-9b-instruct-lora-16k", - model=Model(id="Qwen/Qwen3.5-9B", max_context_length=16384), - trainer=Trainer( - compute=Compute(gpu="H100", gpus_per_node=8), - config={ + +class Config(BaseConfig): + name = "qwen35-9b-instruct-lora-16k" + model = "Qwen/Qwen3.5-9B" + max_context_length = 16384 + trainer = { + "gpu": "H100", + "gpus_per_node": 8, + "config": { "model_type": "qwen3.5-9B", "tensor_model_parallel_size": 8, "target_modules": ["linear_qkv", "linear_proj", "linear_fc1", "linear_fc2"], @@ -19,15 +22,15 @@ "recompute_num_layers": 1, }, }, - env={ + "env": { "PYTORCH_CUDA_ALLOC_CONF": "expandable_segments:True", "TORCHINDUCTOR_COMPILE_THREADS": "1", }, - max_clients_per_instance=6, - ), - inference=Inference( - compute=Compute(gpu="H200"), - config={ + "max_clients_per_instance": 6, + } + inference = { + "gpu": "H200", + "config": { "tp_size": 1, "ep_size": 1, "mem_fraction_static": 0.8, @@ -37,6 +40,8 @@ "max_loras_per_batch": 8, "schedule_policy": "lpm", }, - ), - routing=Routing(default=True), -) + } + default = True + + +config = Config() diff --git a/src/lilo/configs/qwen35_9b_instruct_lora_16k_dp2.py b/src/lilo/configs/qwen35_9b_instruct_lora_16k_dp2.py index 5dd2097..d749920 100644 --- a/src/lilo/configs/qwen35_9b_instruct_lora_16k_dp2.py +++ b/src/lilo/configs/qwen35_9b_instruct_lora_16k_dp2.py @@ -1,12 +1,13 @@ -from dataclasses import replace - -from lilo.configs.qwen35_9b_instruct_lora_16k import config as base - -config = replace( - base, - name="qwen35-9b-instruct-lora-16k-dp2", - trainer=replace( - base.trainer, config={**base.trainer.config, "tensor_model_parallel_size": 4} - ), - routing=replace(base.routing, default=False), -) +from lilo.configs.qwen35_9b_instruct_lora_16k import Config as Parent + + +class Config(Parent): + name = "qwen35-9b-instruct-lora-16k-dp2" + trainer = { + **Parent.trainer, + "config": {**Parent.trainer["config"], "tensor_model_parallel_size": 4}, + } + default = False + + +config = Config() diff --git a/src/lilo/configs/qwen35_9b_lora_16k.py b/src/lilo/configs/qwen35_9b_lora_16k.py index e86279d..c04d24c 100644 --- a/src/lilo/configs/qwen35_9b_lora_16k.py +++ b/src/lilo/configs/qwen35_9b_lora_16k.py @@ -1,11 +1,16 @@ -from lilo.configuration import Compute, Deployment, Inference, Model, Routing, Trainer +from lilo.configuration import BaseConfig -config = Deployment( - name="qwen35-9b-lora-16k", - model=Model(id="Qwen/Qwen3.5-9B-Base", max_context_length=16384), - trainer=Trainer( - compute=Compute(gpu="H100", gpus_per_node=4, cpu=16, memory_mib=65536), - config={ + +class Config(BaseConfig): + name = "qwen35-9b-lora-16k" + model = "Qwen/Qwen3.5-9B-Base" + max_context_length = 16384 + trainer = { + "gpu": "H100", + "gpus_per_node": 4, + "cpu": 16, + "memory_mib": 65536, + "config": { "model_type": "qwen3.5-9B", "tensor_model_parallel_size": 4, "max_lora_slots": 6, @@ -25,15 +30,15 @@ "recompute_num_layers": 1, }, }, - env={ + "env": { "PYTORCH_CUDA_ALLOC_CONF": "expandable_segments:True", "TORCHINDUCTOR_COMPILE_THREADS": "1", }, - max_clients_per_instance=6, - ), - inference=Inference( - compute=Compute(gpu="H200"), - config={ + "max_clients_per_instance": 6, + } + inference = { + "gpu": "H200", + "config": { "tp_size": 1, "mem_fraction_static": 0.8, "max_running_requests": 32, @@ -41,6 +46,8 @@ "max_loaded_loras": 64, "max_loras_per_batch": 8, }, - ), - routing=Routing(default=True), -) + } + default = True + + +config = Config() diff --git a/src/lilo/configs/qwen35_9b_lora_16k_single.py b/src/lilo/configs/qwen35_9b_lora_16k_single.py index 24cb691..0e84a04 100644 --- a/src/lilo/configs/qwen35_9b_lora_16k_single.py +++ b/src/lilo/configs/qwen35_9b_lora_16k_single.py @@ -1,10 +1,10 @@ -from dataclasses import replace +from lilo.configs.qwen35_9b_lora_16k import Config as Parent -from lilo.configs.qwen35_9b_lora_16k import config as base -config = replace( - base, - name="qwen35-9b-lora-16k-single", - trainer=replace(base.trainer, max_clients_per_instance=1), - routing=replace(base.routing, default=False), -) +class Config(Parent): + name = "qwen35-9b-lora-16k-single" + trainer = {**Parent.trainer, "max_clients_per_instance": 1} + default = False + + +config = Config() diff --git a/src/lilo/configs/qwen35_9b_lora_2k.py b/src/lilo/configs/qwen35_9b_lora_2k.py index 3a6c916..6837b08 100644 --- a/src/lilo/configs/qwen35_9b_lora_2k.py +++ b/src/lilo/configs/qwen35_9b_lora_2k.py @@ -1,26 +1,31 @@ -from dataclasses import replace +from lilo.configs.qwen35_9b_lora_16k import Config as Parent -from lilo.configuration import Routing -from lilo.configs.qwen35_9b_lora_16k import config as base -config = replace( - base, - name="qwen35-9b-lora-2k", - model=replace(base.model, max_context_length=2048), - trainer=replace( - base.trainer, - compute=replace(base.trainer.compute, gpu="H200", cpu=8, memory_mib=32768), - max_clients_per_instance=4, - config={**base.trainer.config, "max_tokens_per_gpu": 2048, "max_lora_slots": 4}, - ), - inference=replace( - base.inference, - config={ - **base.inference.config, +class Config(Parent): + name = "qwen35-9b-lora-2k" + max_context_length = 2048 + trainer = { + **Parent.trainer, + "gpu": "H200", + "cpu": 8, + "memory_mib": 32768, + "max_clients_per_instance": 4, + "config": { + **Parent.trainer["config"], + "max_tokens_per_gpu": 2048, + "max_lora_slots": 4, + }, + } + inference = { + **Parent.inference, + "config": { + **Parent.inference["config"], "ep_size": 1, "max_loaded_loras": 32, "schedule_policy": "lpm", }, - ), - routing=Routing(), -) + } + default = False + + +config = Config() diff --git a/src/lilo/configs/qwen35_9b_lora_64k.py b/src/lilo/configs/qwen35_9b_lora_64k.py index 1f2886f..0e343c2 100644 --- a/src/lilo/configs/qwen35_9b_lora_64k.py +++ b/src/lilo/configs/qwen35_9b_lora_64k.py @@ -1,19 +1,20 @@ -from dataclasses import replace +from lilo.configs.qwen35_9b_lora_16k import Config as Parent -from lilo.configs.qwen35_9b_lora_16k import config as base -config = replace( - base, - name="qwen35-9b-lora-64k", - model=replace(base.model, max_context_length=65536), - trainer=replace( - base.trainer, - compute=replace(base.trainer.compute, gpu="H200", gpus_per_node=8), - config={ - **base.trainer.config, +class Config(Parent): + name = "qwen35-9b-lora-64k" + max_context_length = 65536 + trainer = { + **Parent.trainer, + "gpu": "H200", + "gpus_per_node": 8, + "config": { + **Parent.trainer["config"], "tensor_model_parallel_size": 8, "max_tokens_per_gpu": 65536, }, - ), - routing=replace(base.routing, default=False), -) + } + default = False + + +config = Config() diff --git a/src/lilo/configs/qwen36_27b_fft_64k.py b/src/lilo/configs/qwen36_27b_fft_64k.py index 90cf2a5..1a4ef5f 100644 --- a/src/lilo/configs/qwen36_27b_fft_64k.py +++ b/src/lilo/configs/qwen36_27b_fft_64k.py @@ -1,14 +1,16 @@ -from lilo.configuration import Compute, Deployment, Inference, Model, Routing, Trainer +from lilo.configuration import BaseConfig -config = Deployment( - name="qwen36-27b-fft-64k", - model=Model( - parameterization="full", id="Qwen/Qwen3.6-27B", max_context_length=65536 - ), - trainer=Trainer( - compute=Compute(gpu="H200", gpus_per_node=8), - backend="megatron", - config={ + +class Config(BaseConfig): + name = "qwen36-27b-fft-64k" + parameterization = "full" + model = "Qwen/Qwen3.6-27B" + max_context_length = 65536 + trainer = { + "gpu": "H200", + "gpus_per_node": 8, + "backend": "megatron", + "config": { "tensor_model_parallel_size": 4, "pipeline_model_parallel_size": 1, "context_parallel_size": 2, @@ -27,15 +29,16 @@ }, "optimizer": {"optimizer": "adam", "lr": 0.0001, "min_lr": 0.0001}, }, - env={ + "env": { "PYTORCH_CUDA_ALLOC_CONF": "expandable_segments:True", "TORCHINDUCTOR_COMPILE_THREADS": "1", }, - sampler_persistence_concurrency=1, - ), - inference=Inference( - compute=Compute(gpu="H200", gpus_per_node=4), - config={ + "sampler_persistence_concurrency": 1, + } + inference = { + "gpu": "H200", + "gpus_per_node": 4, + "config": { "tp_size": 4, "ep_size": 1, "mem_fraction_static": 0.9, @@ -43,6 +46,8 @@ "max_queued_requests": 4, "cpu_weight_cache_max_compile_group_gb": 32, }, - ), - routing=Routing(default=True), -) + } + default = True + + +config = Config() diff --git a/src/lilo/configs/qwen36_35b_a3b_fft_64k.py b/src/lilo/configs/qwen36_35b_a3b_fft_64k.py index df04a3f..1a82fee 100644 --- a/src/lilo/configs/qwen36_35b_a3b_fft_64k.py +++ b/src/lilo/configs/qwen36_35b_a3b_fft_64k.py @@ -1,13 +1,17 @@ -from dataclasses import replace - -from lilo.configs.qwen35_35b_a3b_fft_64k import config as base - -config = replace( - base, - name="qwen36-35b-a3b-fft-64k", - model=replace(base.model, id="Qwen/Qwen3.6-35B-A3B"), - inference=replace( - base.inference, - config={**base.inference.config, "dp_size": 1, "enable_dp_attention": False}, - ), -) +from lilo.configs.qwen35_35b_a3b_fft_64k import Config as Parent + + +class Config(Parent): + name = "qwen36-35b-a3b-fft-64k" + model = "Qwen/Qwen3.6-35B-A3B" + inference = { + **Parent.inference, + "config": { + **Parent.inference["config"], + "dp_size": 1, + "enable_dp_attention": False, + }, + } + + +config = Config() diff --git a/src/lilo/configs/qwen38_27b_lora_128k.py b/src/lilo/configs/qwen38_27b_lora_128k.py index c9303c2..752b4b1 100644 --- a/src/lilo/configs/qwen38_27b_lora_128k.py +++ b/src/lilo/configs/qwen38_27b_lora_128k.py @@ -1,26 +1,30 @@ -from dataclasses import replace +from lilo.configs.qwen38_27b_lora_16k import Config as Parent -from lilo.configs.qwen38_27b_lora_16k import config as base -config = replace( - base, - name="qwen38-27b-lora-128k", - model=replace(base.model, max_context_length=131072), - trainer=replace( - base.trainer, - config={ - **base.trainer.config, +class Config(Parent): + name = "qwen38-27b-lora-128k" + max_context_length = 131072 + trainer = { + **Parent.trainer, + "config": { + **Parent.trainer["config"], "tensor_model_parallel_size": 2, "max_tokens_per_gpu": 32768, "context_parallel_size": 4, }, - ), - inference=replace( - base.inference, - compute=replace(base.inference.compute, gpus_per_node=2), - max_replicas=4, - target_concurrency=4, - config={**base.inference.config, "tp_size": 2, "max_running_requests": 8}, - ), - routing=replace(base.routing, default=False), -) + } + inference = { + **Parent.inference, + "gpus_per_node": 2, + "max_replicas": 4, + "target_concurrency": 4, + "config": { + **Parent.inference["config"], + "tp_size": 2, + "max_running_requests": 8, + }, + } + default = False + + +config = Config() diff --git a/src/lilo/configs/qwen38_27b_lora_16k.py b/src/lilo/configs/qwen38_27b_lora_16k.py index 2835e7a..f97a257 100644 --- a/src/lilo/configs/qwen38_27b_lora_16k.py +++ b/src/lilo/configs/qwen38_27b_lora_16k.py @@ -1,11 +1,14 @@ -from lilo.configuration import Compute, Deployment, Inference, Model, Routing, Trainer +from lilo.configuration import BaseConfig -config = Deployment( - name="qwen38-27b-lora-16k", - model=Model(id="Qwen/Qwen3.8-27B", max_context_length=16384), - trainer=Trainer( - compute=Compute(gpu="H200", gpus_per_node=8), - config={ + +class Config(BaseConfig): + name = "qwen38-27b-lora-16k" + model = "Qwen/Qwen3.8-27B" + max_context_length = 16384 + trainer = { + "gpu": "H200", + "gpus_per_node": 8, + "config": { "model_type": "qwen3.8-27B", "tensor_model_parallel_size": 4, "target_modules": ["linear_qkv", "linear_proj", "linear_fc1", "linear_fc2"], @@ -19,15 +22,15 @@ "recompute_num_layers": 1, }, }, - env={ + "env": { "PYTORCH_CUDA_ALLOC_CONF": "expandable_segments:True", "TORCHINDUCTOR_COMPILE_THREADS": "1", }, - max_clients_per_instance=6, - ), - inference=Inference( - compute=Compute(gpu="H200"), - config={ + "max_clients_per_instance": 6, + } + inference = { + "gpu": "H200", + "config": { "tp_size": 1, "ep_size": 1, "mem_fraction_static": 0.8, @@ -37,6 +40,8 @@ "max_loras_per_batch": 8, "schedule_policy": "lpm", }, - ), - routing=Routing(default=True), -) + } + default = True + + +config = Config() diff --git a/src/lilo/configs/qwen38_27b_lora_256k.py b/src/lilo/configs/qwen38_27b_lora_256k.py index 8aab4f7..a8961ee 100644 --- a/src/lilo/configs/qwen38_27b_lora_256k.py +++ b/src/lilo/configs/qwen38_27b_lora_256k.py @@ -1,38 +1,38 @@ -from dataclasses import replace +from lilo.configs.qwen38_27b_lora_16k import Config as Parent -from lilo.configs.qwen38_27b_lora_16k import config as base -config = replace( - base, - name="qwen38-27b-lora-256k", - model=replace(base.model, max_context_length=262144), - routing=replace(base.routing, default=False), - trainer=replace( - base.trainer, - compute=replace(base.trainer.compute, nodes=2), - config={ - **base.trainer.config, +class Config(Parent): + name = "qwen38-27b-lora-256k" + max_context_length = 262144 + default = False + trainer = { + **Parent.trainer, + "nodes": 2, + "config": { + **Parent.trainer["config"], "tensor_model_parallel_size": 2, "context_parallel_size": 8, "max_tokens_per_gpu": 32768, "cli_options": { - **base.trainer.config["cli_options"], + **Parent.trainer["config"]["cli_options"], "distributed_timeout_minutes": 120, }, }, - ), - inference=replace( - base.inference, - compute=replace(base.inference.compute, gpus_per_node=4), - min_replicas=2, - max_replicas=2, - target_concurrency=2, - config={ - **base.inference.config, + } + inference = { + **Parent.inference, + "gpus_per_node": 4, + "min_replicas": 2, + "max_replicas": 2, + "target_concurrency": 2, + "config": { + **Parent.inference["config"], "tp_size": 4, "max_running_requests": 4, "max_queued_requests": 8, "max_loaded_loras": 256, }, - ), -) + } + + +config = Config() diff --git a/src/lilo/configs/qwen38_27b_lora_64k.py b/src/lilo/configs/qwen38_27b_lora_64k.py index 3c75488..71cb28b 100644 --- a/src/lilo/configs/qwen38_27b_lora_64k.py +++ b/src/lilo/configs/qwen38_27b_lora_64k.py @@ -1,23 +1,23 @@ -from dataclasses import replace +from lilo.configs.qwen38_27b_lora_16k import Config as Parent -from lilo.configs.qwen38_27b_lora_16k import config as base -config = replace( - base, - name="qwen38-27b-lora-64k", - model=replace(base.model, max_context_length=65536), - trainer=replace( - base.trainer, - config={ - **base.trainer.config, +class Config(Parent): + name = "qwen38-27b-lora-64k" + max_context_length = 65536 + trainer = { + **Parent.trainer, + "config": { + **Parent.trainer["config"], "max_tokens_per_gpu": 32768, "context_parallel_size": 2, }, - ), - inference=replace( - base.inference, - target_concurrency=8, - config={**base.inference.config, "max_running_requests": 16}, - ), - routing=replace(base.routing, default=False), -) + } + inference = { + **Parent.inference, + "target_concurrency": 8, + "config": {**Parent.inference["config"], "max_running_requests": 16}, + } + default = False + + +config = Config() diff --git a/src/lilo/configuration.py b/src/lilo/configuration.py index b092970..50bcf33 100644 --- a/src/lilo/configuration.py +++ b/src/lilo/configuration.py @@ -1,8 +1,4 @@ -"""Validated Python components for model infrastructure. - -Only backend config and environment dictionaries accept arbitrary keys. -Use dataclasses.replace to compose variants without mutating another config. -""" +"""Readable Python recipes with validated model, trainer, and inference settings.""" from dataclasses import field from typing import Annotated, Literal @@ -13,33 +9,17 @@ PositiveInt = Annotated[int, Field(strict=True, gt=0)] NonnegativeInt = Annotated[int, Field(strict=True, ge=0)] NonemptyString = Annotated[str, Field(strict=True, min_length=1)] +GPU = Annotated[str, Field(strict=True, pattern=r"^[A-Za-z0-9-]+$")] CONFIG = ConfigDict(extra="forbid", validate_default=True) @dataclass(config=CONFIG, frozen=True, kw_only=True) -class Compute: - gpu: Annotated[str, Field(pattern=r"^[A-Za-z0-9-]+$")] +class Trainer: + gpu: GPU gpus_per_node: PositiveInt = 1 nodes: PositiveInt = 1 cpu: Annotated[float, Field(gt=0)] = 8 memory_mib: PositiveInt = 32768 - - @property - def modal_gpu(self) -> str: - return f"{self.gpu}:{self.gpus_per_node}" - - -@dataclass(config=CONFIG, frozen=True, kw_only=True) -class Model: - id: NonemptyString - max_context_length: PositiveInt - parameterization: Literal["lora", "full"] = "lora" - revision: NonemptyString = "main" - - -@dataclass(config=CONFIG, frozen=True, kw_only=True) -class Trainer: - compute: Compute backend: Literal["miles", "megatron"] = "miles" max_instances: PositiveInt = 1 max_clients_per_instance: PositiveInt = 1 @@ -51,8 +31,10 @@ class Trainer: @dataclass(config=CONFIG, frozen=True, kw_only=True) class Inference: - compute: Compute - backend: Literal["sglang"] = "sglang" + gpu: GPU + gpus_per_node: PositiveInt = 1 + cpu: Annotated[float, Field(gt=0)] = 8 + memory_mib: PositiveInt = 32768 min_replicas: NonnegativeInt = 0 max_replicas: PositiveInt = 8 target_concurrency: PositiveInt = 16 @@ -63,43 +45,56 @@ class Inference: @dataclass(config=CONFIG, frozen=True, kw_only=True) -class Routing: +class Deployment: + name: Annotated[str, Field(strict=True, pattern=r"^[A-Za-z0-9_-]+$")] + model: NonemptyString + max_context_length: PositiveInt + trainer: Trainer + inference: Inference + parameterization: Literal["lora", "full"] = "lora" + revision: NonemptyString = "main" default: StrictBool = False sampling_default: StrictBool = False - - -@dataclass(config=CONFIG, frozen=True, kw_only=True) -class Lifecycle: session_idle_timeout_s: PositiveInt = 300 pool_idle_timeout_s: PositiveInt = 300 sweep_interval_s: PositiveInt = 300 - -@dataclass(config=CONFIG, frozen=True, kw_only=True) -class Deployment: - name: Annotated[str, Field(pattern=r"^[A-Za-z0-9_-]+$")] - model: Model - trainer: Trainer - inference: Inference - routing: Routing = field(default_factory=Routing) - lifecycle: Lifecycle = field(default_factory=Lifecycle) - def __post_init__(self): if self.inference.min_replicas > self.inference.max_replicas: raise ValueError("min_replicas must not exceed max_replicas") - if self.inference.compute.nodes != 1: - raise ValueError("each inference replica uses one node") if self.trainer.backend == "megatron": - if self.trainer.compute.nodes != 1: + if self.trainer.nodes != 1: raise ValueError("multi-node training currently requires Miles") - if self.model.parameterization != "full": + if self.parameterization != "full": raise ValueError("Megatron requires full parameterization") if self.trainer.max_clients_per_instance != 1: raise ValueError("FFT trainers admit one client per instance") if self.trainer.sampler_persistence_concurrency != 1: raise ValueError("Megatron requires sampler_persistence_concurrency: 1") - elif self.model.parameterization != "lora": + elif self.parameterization != "lora": raise ValueError("Miles requires lora parameterization") for env in (self.trainer.env, self.inference.env): if any(key.startswith("LILO_") for key in env): raise ValueError("LILO_ environment variables are managed by Lilo") + + +class BaseConfig(Deployment): + """Author a recipe with class attributes and ordinary Python inheritance. + + Sections are dictionaries in the recipe and validated values in an instance. + Overriding a section replaces it; use ``{**Parent.trainer, ...}`` to extend it. + """ + + def __init__(self, **overrides): + from copy import deepcopy + + values = {} + for cls in reversed(type(self).__mro__): + if cls in (object, Deployment, BaseConfig): + continue + values.update( + (key, value) + for key, value in vars(cls).items() + if not key.startswith("_") + ) + super().__init__(**deepcopy(values | overrides)) diff --git a/src/lilo/deployment_cli.py b/src/lilo/deployment_cli.py index a50b12e..58b9ad1 100644 --- a/src/lilo/deployment_cli.py +++ b/src/lilo/deployment_cli.py @@ -17,6 +17,7 @@ from lilo.backends.deployment import resolve_backend_settings from lilo.deployments import ( + LIFECYCLE_FIELDS, PLATFORM_DEFAULTS, DeploymentRecord, config_path, @@ -38,12 +39,12 @@ def compile_configs(paths, *, platform=None): ) records = [] for spec in specs: - revision = spec.model.revision + revision = spec.revision if not re.fullmatch(r"[0-9a-fA-F]{40,64}", revision): - revision = HfApi().model_info(spec.model.id, revision=revision).sha + revision = HfApi().model_info(spec.model, revision=revision).sha if not revision or not re.fullmatch(r"[0-9a-fA-F]{40,64}", revision): raise ValueError( - f"Hugging Face did not return a commit for {spec.model.id}" + f"Hugging Face did not return a commit for {spec.model}" ) records.append( DeploymentRecord.create( @@ -61,9 +62,9 @@ def retain_generations(previous, desired): validate_frontend([row.spec for row in desired]) expected = desired[0] for row in previous: - if ( - row.platform != expected.platform - or row.spec.lifecycle != expected.spec.lifecycle + if row.platform != expected.platform or any( + getattr(row.spec, field) != getattr(expected.spec, field) + for field in LIFECYCLE_FIELDS ): raise ValueError( "Cannot change shared storage, secrets, region, or lifecycle while retaining generations; use a separate frontend." @@ -268,7 +269,7 @@ def main(argv=None): if not config_path(args.preset).is_file(): raise ValueError(f"unknown example config: {args.preset}") print( - f'from dataclasses import replace\nfrom lilo.configs.{module} import config as base\n\nconfig = replace(base, name="my-model")' + f'from lilo.configs.{module} import Config as Parent\n\n\nclass Config(Parent):\n name = "my-model"\n\n\nconfig = Config()' ) elif args.action == "validate": specs = [load(path) for path in args.files] diff --git a/src/lilo/deployments.py b/src/lilo/deployments.py index 37e5347..e73baca 100644 --- a/src/lilo/deployments.py +++ b/src/lilo/deployments.py @@ -18,6 +18,8 @@ from lilo.backends.deployment import resolve_backend_settings from lilo.configuration import Deployment +LIFECYCLE_FIELDS = ("session_idle_timeout_s", "pool_idle_timeout_s", "sweep_interval_s") + PLATFORM_DEFAULTS = { "frontend": "lilo-yaml", "modal": {"environment": None, "region": "us-west"}, @@ -49,7 +51,8 @@ def deployment_generation( inference_settings, ): identity = asdict(spec) - identity.pop("routing") + identity.pop("default") + identity.pop("sampling_default") return settings_hash( { "config": identity, @@ -99,8 +102,8 @@ def create( inference_release: str = "initial", ) -> DeploymentRecord: """Record an already-resolved revision without reparsing the configuration.""" - pinned = replace(deepcopy(spec), model=replace(spec.model, revision=revision)) - asset_path = model_asset_path(pinned.model.id, revision) + pinned = replace(deepcopy(spec), revision=revision) + asset_path = model_asset_path(pinned.model, revision) trainer_settings, inference_settings = resolve_backend_settings( pinned, asset_path ) @@ -149,7 +152,10 @@ def trainer_hash(self) -> str: return settings_hash( { "name": self.spec.name, - "model": asdict(self.spec.model), + "model": self.spec.model, + "revision": self.spec.revision, + "parameterization": self.spec.parameterization, + "max_context_length": self.spec.max_context_length, "trainer": asdict(self.spec.trainer), "settings": self.trainer_settings, "release": self.trainer_release, @@ -164,7 +170,10 @@ def inference_hash(self) -> str: return settings_hash( { "name": self.spec.name, - "model": asdict(self.spec.model), + "model": self.spec.model, + "revision": self.spec.revision, + "parameterization": self.spec.parameterization, + "max_context_length": self.spec.max_context_length, "inference": asdict(self.spec.inference), "settings": self.inference_settings, "release": self.inference_release, @@ -186,7 +195,7 @@ def definition_id(self) -> str: @property def asset_path(self) -> str: - return model_asset_path(self.spec.model.id, self.spec.model.revision) + return model_asset_path(self.spec.model, self.spec.revision) def model_asset_path(model_id, revision): @@ -213,7 +222,7 @@ def load(path: str | Path) -> Deployment: namespace = runpy.run_path(str(path)) config = namespace.get("config") if not isinstance(config, Deployment): - raise ValueError(f"{path} must export a Deployment object named config") + raise ValueError(f"{path} must export a BaseConfig instance named config") return config finally: sys.path[:] = original_path @@ -226,18 +235,20 @@ def validate_frontend(specs: list[Deployment]) -> None: raise ValueError("duplicate deployment name") first = specs[0] for spec in specs: - if spec.lifecycle != first.lifecycle: + if any( + getattr(spec, field) != getattr(first, field) for field in LIFECYCLE_FIELDS + ): raise ValueError( "deployments on one frontend must share lifecycle settings" ) defaults, sampling = set(), set() for spec in specs: - key = (spec.model.id, spec.model.parameterization) - if spec.routing.default: + key = (spec.model, spec.parameterization) + if spec.default: if key in defaults: raise ValueError(f"multiple defaults for {key}") defaults.add(key) - if spec.routing.sampling_default: - if spec.model.id in sampling: - raise ValueError(f"multiple sampling defaults for {spec.model.id}") - sampling.add(spec.model.id) + if spec.sampling_default: + if spec.model in sampling: + raise ValueError(f"multiple sampling defaults for {spec.model}") + sampling.add(spec.model) diff --git a/src/lilo/providers/modal/app.py b/src/lilo/providers/modal/app.py index 5e6ee50..07a456b 100644 --- a/src/lilo/providers/modal/app.py +++ b/src/lilo/providers/modal/app.py @@ -58,13 +58,11 @@ APP_NAME = SETTINGS.platform["frontend"] ROUTING_REGION = SETTINGS.platform["modal"]["region"] MODEL_ASSET_ROOT = "/assets" -SESSION_IDLE_TIMEOUT = SETTINGS.spec.lifecycle.session_idle_timeout_s -FFT_POOL_IDLE_TIMEOUT = LORA_POOL_IDLE_TIMEOUT = ( - SETTINGS.spec.lifecycle.pool_idle_timeout_s -) +SESSION_IDLE_TIMEOUT = SETTINGS.spec.session_idle_timeout_s +FFT_POOL_IDLE_TIMEOUT = LORA_POOL_IDLE_TIMEOUT = SETTINGS.spec.pool_idle_timeout_s FFT_POOL_TOUCH_INTERVAL = 60.0 LORA_POOL_CHECK_INTERVAL = 60.0 -SWEEP_PERIOD = modal.Period(seconds=SETTINGS.spec.lifecycle.sweep_interval_s) +SWEEP_PERIOD = modal.Period(seconds=SETTINGS.spec.sweep_interval_s) CHECKPOINT_READ_LOCK = asyncio.Lock() _pool_touches: dict[str, float] = {} _lora_pool_gateways: dict[str, tuple[float, str]] = {} diff --git a/src/lilo/providers/modal/deployment_apps.py b/src/lilo/providers/modal/deployment_apps.py index db1ffab..1775e45 100644 --- a/src/lilo/providers/modal/deployment_apps.py +++ b/src/lilo/providers/modal/deployment_apps.py @@ -104,7 +104,7 @@ def build_trainer_app(resolved: DeploymentRecord, *, image=None): spec = resolved.spec trainer_hash = resolved.trainer_hash app = modal.App(resolved.trainer_app_name) - resource = spec.trainer.compute + resource = spec.trainer env = { **trainer_deployment_env(), @@ -124,7 +124,7 @@ def trainer(instance_id: str, config_json: str): name="trainer", serialized=True, image=image if image is not None else image_for(spec.trainer.backend), - gpu=resource.modal_gpu, + gpu=f"{resource.gpu}:{resource.gpus_per_node}", region=resolved.platform["modal"]["region"], cpu=resource.cpu, memory=resource.memory_mib, @@ -149,9 +149,9 @@ def run_trainer(resolved, instance_id): # Assets are prepared by the frontend before demand is registered. Reload once # on startup to see the committed exact snapshot; never race a trainer download. assets = volumes_for(resolved)["/assets"] - if spec.trainer.compute.nodes > 1: + if spec.trainer.nodes > 1: ray_address = start_trainer_cluster( - spec.trainer.compute.nodes, + spec.trainer.nodes, before_head=assets.reload, before_worker_join=assets.reload, ) @@ -164,8 +164,8 @@ def run_trainer(resolved, instance_id): **deployment_env(spec.trainer.env), "LILO_APP_NAME": resolved.platform["frontend"], "LILO_BACKEND_CONFIG": json.dumps(settings), - "LILO_BASE_MODEL": spec.model.id, - "LILO_BASE_MODEL_REVISION": spec.model.revision, + "LILO_BASE_MODEL": spec.model, + "LILO_BASE_MODEL_REVISION": spec.revision, "LILO_DEFINITION_ID": resolved.definition_id, "LILO_CHECKPOINT_VOLUME": resolved.platform["storage"]["checkpoints"], "LILO_BULLETIN_ROOT": "/bulletin", @@ -193,9 +193,7 @@ async def failed(error): revision=config["image_id"], instance_id=instance_id, backend_env=env, - nproc=1 - if spec.trainer.backend == "miles" - else spec.trainer.compute.gpus_per_node, + nproc=1 if spec.trainer.backend == "miles" else spec.trainer.gpus_per_node, max_models=spec.trainer.max_clients_per_instance, sampler_persistence_concurrency=spec.trainer.sampler_persistence_concurrency, on_startup_error=failed, @@ -207,21 +205,21 @@ def definition_from_spec(resolved, *, register_trainer=True, image=None): serving = resolved.inference_settings definition = SimpleNamespace( DEFINITION_ID=resolved.definition_id, - MODEL_NAME=spec.model.id, - MODEL_REVISION=spec.model.revision, + MODEL_NAME=spec.model, + MODEL_REVISION=spec.revision, HF_CHECKPOINT=resolved.asset_path, - PARAMETERIZATION=spec.model.parameterization, + PARAMETERIZATION=spec.parameterization, CATALOG_VISIBLE=resolved.active, - ROUTING_DEFAULT=spec.routing.default, - SAMPLING_DEFAULT=spec.routing.sampling_default, + ROUTING_DEFAULT=spec.default, + SAMPLING_DEFAULT=spec.sampling_default, DEPLOYMENT_NAME=spec.name, RESOLVED=resolved, - MAX_CONTEXT_LENGTH=spec.model.max_context_length, + MAX_CONTEXT_LENGTH=spec.max_context_length, TRAINER_MODELS_PER_INSTANCE=spec.trainer.max_clients_per_instance, TRAINER_MAX_CONTAINERS=spec.trainer.max_instances, - ROLLOUT_GPUS=spec.inference.compute.gpus_per_node, + ROLLOUT_GPUS=spec.inference.gpus_per_node, ROLLOUT_TENSOR_PARALLEL_SIZE=serving.get( - "tp_size", spec.inference.compute.gpus_per_node + "tp_size", spec.inference.gpus_per_node ) // (serving.get("dp_size", 1) if serving.get("enable_dp_attention") else 1), ) @@ -237,15 +235,15 @@ def definition_from_spec(resolved, *, register_trainer=True, image=None): def build_rollout_app(resolved, pool, *, image=None): """Create one frozen-base LoRA pool or one FFT latest/pinned/base pool.""" spec = resolved.spec - lora = spec.model.parameterization == "lora" + lora = spec.parameterization == "lora" if pool.definition_id != resolved.definition_id: raise ValueError("pool generation does not match deployment") app = modal.App(pool.app_name) - resources, scaling = spec.inference.compute, spec.inference + inference = spec.inference options = resolved.inference_settings - minimum = scaling.min_replicas - maximum = scaling.max_replicas - window = scaling.scaledown_window_s + minimum = inference.min_replicas + maximum = inference.max_replicas + window = inference.scaledown_window_s if isinstance(pool, FFTPoolSpec): minimum = minimum if pool.min_containers is None else pool.min_containers maximum = maximum if pool.max_containers is None else pool.max_containers @@ -256,17 +254,17 @@ def build_rollout_app(resolved, pool, *, image=None): name="Server", serialized=True, image=image if image is not None else image_for("sglang"), - gpu=resources.modal_gpu, - cpu=resources.cpu, - memory=resources.memory_mib, + gpu=f"{inference.gpu}:{inference.gpus_per_node}", + cpu=inference.cpu, + memory=inference.memory_mib, volumes=volumes_for(resolved), secrets=secrets_for(resolved), - env=deployment_env(spec.inference.env), + env=deployment_env(inference.env), min_containers=minimum, max_containers=maximum, - target_concurrency=scaling.target_concurrency, + target_concurrency=inference.target_concurrency, scaledown_window=window, - startup_timeout=spec.inference.startup_timeout_s, + startup_timeout=inference.startup_timeout_s, exit_grace_period=300, port=8000, routing_region=resolved.platform["modal"]["region"], @@ -292,7 +290,7 @@ def start(self): wait_http( "http://127.0.0.1:8001/health", self.sglang, - spec.inference.startup_timeout_s, + inference.startup_timeout_s, ) kwargs = dict( port=8000, @@ -314,7 +312,7 @@ def start(self): wait_http( "http://127.0.0.1:8000/health", self.sidecar, - spec.inference.startup_timeout_s, + inference.startup_timeout_s, ) except BaseException: terminate(self.sidecar) @@ -351,7 +349,7 @@ def provision(config_json: str, pool_data: dict): raise ValueError("inference settings do not match the deployed app") if pool_data["definition_id"] != saved.definition_id: raise ValueError("pool definition does not match deployment") - if saved.spec.model.parameterization == "lora": + if saved.spec.parameterization == "lora": return deploy_lora(LoraPoolSpec.from_dict(pool_data), record=saved) return deploy_fft(FFTPoolSpec.from_dict(pool_data), record=saved) diff --git a/src/lilo/providers/modal/deployment_pool_app.py b/src/lilo/providers/modal/deployment_pool_app.py index 402ddce..7cccfae 100644 --- a/src/lilo/providers/modal/deployment_pool_app.py +++ b/src/lilo/providers/modal/deployment_pool_app.py @@ -10,7 +10,7 @@ from .lora_pool import LoraPoolSpec resolved = DeploymentRecord.model_validate_json(os.environ[POOL_CONFIG_ENV]) -if resolved.spec.model.parameterization == "lora": +if resolved.spec.parameterization == "lora": pool = LoraPoolSpec(resolved.definition_id, revision=resolved.generation[:16]) else: pool = FFTPoolSpec( diff --git a/tests/providers/test_deployment_apps.py b/tests/providers/test_deployment_apps.py index be3296b..242edb9 100644 --- a/tests/providers/test_deployment_apps.py +++ b/tests/providers/test_deployment_apps.py @@ -106,7 +106,10 @@ def test_pool_starts_native_server_and_correct_sidecar(builders, monkeypatch, ki app, server = deployment_apps.build_rollout_app(row, pool, image="test-image") settings, _ = app.servers["Server"] assert app.name == pool.app_name - assert settings["gpu"] == row.spec.inference.compute.modal_gpu + assert ( + settings["gpu"] + == f"{row.spec.inference.gpu}:{row.spec.inference.gpus_per_node}" + ) assert settings["min_containers"] == 0 assert settings["target_concurrency"] == 16 assert settings["compute_region"] == "us-west" @@ -135,7 +138,7 @@ def test_pool_starts_native_server_and_correct_sidecar(builders, monkeypatch, ki assert commands[0][2] == "lilo.inference.sglang" assert commands[0][3] == row.asset_path native = json.loads(commands[0][4]) - assert native["context_length"] == row.spec.model.max_context_length + assert native["context_length"] == row.spec.max_context_length if kind == "lora": assert native["enable_lora"] is True assert native["max_lora_rank"] == 32 @@ -231,16 +234,17 @@ def test_admission_changes_preserve_serialized_trainer(builders): changed.active = False changed.spec = replace( changed.spec, - routing=replace(changed.spec.routing, default=False, sampling_default=True), + default=False, + sampling_default=True, ) new_bytes = serialize(deployment_apps.build_trainer_app(changed, image="test")[1]) assert new_bytes == old_bytes - assert first.active is True and first.spec.routing.default is True + assert first.active is True and first.spec.default is True changed.spec = replace( changed.spec, trainer=replace( changed.spec.trainer, - compute=replace(changed.spec.trainer.compute, gpu="H200"), + gpu="H200", ), ) assert ( @@ -432,14 +436,16 @@ def test_declared_compute_settings_reach_modal(builders): trainer=replace( base.trainer, timeout_s=90, - compute=replace(base.trainer.compute, cpu=12, memory_mib=123456), + cpu=12, + memory_mib=123456, ), inference=replace( base.inference, startup_timeout_s=90, min_replicas=1, max_replicas=3, - compute=replace(base.inference.compute, cpu=6, memory_mib=45000), + cpu=6, + memory_mib=45000, ), ) row = DeploymentRecord.create(spec, revision="a" * 40) diff --git a/tests/providers/test_deployment_e2e_helper.py b/tests/providers/test_deployment_e2e_helper.py index 3beda47..9541207 100644 --- a/tests/providers/test_deployment_e2e_helper.py +++ b/tests/providers/test_deployment_e2e_helper.py @@ -21,9 +21,9 @@ def test_e2e_helper_reads_active_deployed_configuration(monkeypatch, preset): monkeypatch.setattr(modal.Dict, "from_name", lambda name: registry) definition, mode = helper["_definition"]("test-frontend", row.spec.name) assert definition.DEFINITION_ID == row.definition_id - assert definition.MAX_CONTEXT_LENGTH == row.spec.model.max_context_length - assert definition.GPUS == row.spec.trainer.compute.gpus_per_node - assert mode == row.spec.model.parameterization + assert definition.MAX_CONTEXT_LENGTH == row.spec.max_context_length + assert definition.GPUS == row.spec.trainer.gpus_per_node + assert mode == row.spec.parameterization assert definition.MAX_TOKENS_PER_MICROBATCH > 0 with pytest.raises(ValueError, match="one active YAML"): helper["_definition"]("test-frontend", "missing") diff --git a/tests/providers/test_deployment_presets.py b/tests/providers/test_deployment_presets.py index eed56e5..1a5b46e 100644 --- a/tests/providers/test_deployment_presets.py +++ b/tests/providers/test_deployment_presets.py @@ -57,7 +57,11 @@ def test_single_client_recipe_keeps_shared_backend_capacity(): shared = load(config_path("qwen35-9b-lora-16k")) single = load(config_path("qwen35-9b-lora-16k-single")) assert single.trainer.max_clients_per_instance == 1 - assert single.trainer.compute == shared.trainer.compute + assert (single.trainer.gpu, single.trainer.gpus_per_node, single.trainer.nodes) == ( + shared.trainer.gpu, + shared.trainer.gpus_per_node, + shared.trainer.nodes, + ) assert backend_config(single) == backend_config(shared) diff --git a/tests/test_deployment_cli.py b/tests/test_deployment_cli.py index e3f59e7..e7439c4 100644 --- a/tests/test_deployment_cli.py +++ b/tests/test_deployment_cli.py @@ -124,15 +124,15 @@ def test_compile_pins_revision_at_external_boundary( path = tmp_path / "model.py" path.write_text( - "from dataclasses import replace\n" - "from lilo.configs.qwen35_9b_lora_16k import config as base\n" - f"config = replace(base, model=replace(base.model, revision={revision!r}))\n" + "from lilo.configs.qwen35_9b_lora_16k import Config as Parent\n" + f"class Config(Parent):\n revision = {revision!r}\n" + "config = Config()\n" ) lookup = Mock(return_value=SimpleNamespace(sha="a" * 40)) monkeypatch.setattr(huggingface_hub.HfApi, "model_info", lookup) monkeypatch.setattr(miles_revision, "resolve_miles_commit", lambda: "b" * 40) (row,) = cli.compile_configs([path]) - assert row.spec.model.revision == "a" * 40 + assert row.spec.revision == "a" * 40 assert lookup.call_count == lookups if lookups: lookup.return_value.sha = None @@ -329,5 +329,5 @@ def test_builtin_config_resolves_revision_automatically(monkeypatch): ) (row,) = cli.compile_configs([config_path("qwen35-9b-lora-16k")]) assert calls == [("Qwen/Qwen3.5-9B-Base", "main")] - assert row.spec.model.revision == "a" * 40 - assert load(config_path("qwen35-9b-lora-16k")).model.revision == "main" + assert row.spec.revision == "a" * 40 + assert load(config_path("qwen35-9b-lora-16k")).revision == "main" diff --git a/tests/test_deployments.py b/tests/test_deployments.py index 424d38f..1bcae53 100644 --- a/tests/test_deployments.py +++ b/tests/test_deployments.py @@ -51,7 +51,7 @@ def test_presets_context_topology_and_backend_options(): assert config["cli_options"]["recompute_num_layers"] == 1 assert config["extra_args"] == ("--seq-length", "16384") large = recipe("qwen35-9b-lora-64k") - assert large.model.max_context_length == 65536 + assert large.max_context_length == 65536 assert backend_config(large)["miles"]["actor_num_gpus_per_node"] == 8 fft = backend_config(recipe("qwen35-4b-fft-64k"))["megatron"] assert (fft["tensor_model_parallel_size"], fft["context_parallel_size"]) == (2, 2) @@ -60,7 +60,7 @@ def test_presets_context_topology_and_backend_options(): def test_no_model_catalog_required(): spec = recipe( - model__id="my-org/new-model", + model="my-org/new-model", trainer__config__model_type="", trainer__config__cli_options={ "num_layers": 12, @@ -105,9 +105,9 @@ def test_invalid_integrations_fail_when_building_backend_settings(changes, match def test_generation_and_asset_paths_include_exact_base(): a = resolved() - assert a.generation == resolved(recipe(routing__default=False)).generation - assert a.generation != resolved(recipe(trainer__compute__gpu="H200")).generation - b = resolved(recipe(model__id="other/Qwen3.5-9B-Base")) + assert a.generation == resolved(recipe(default=False)).generation + assert a.generation != resolved(recipe(trainer__gpu="H200")).generation + b = resolved(recipe(model="other/Qwen3.5-9B-Base")) assert a.asset_path != b.asset_path assert a.asset_path != DeploymentRecord.create(a.spec, revision="b" * 40).asset_path @@ -116,50 +116,42 @@ def test_frontend_defaults_and_retained_generations(): small = resolved() large = resolved(recipe("qwen35-9b-lora-64k")) routes = DeploymentRoutes([definition(small), definition(large)]) - assert ( - routes.select(small.spec.model.id, "lora").DEFINITION_ID == small.definition_id - ) + assert routes.select(small.spec.model, "lora").DEFINITION_ID == small.definition_id switched = retain_generations( - [small, large], [resolved(recipe("qwen35-9b-lora-64k", routing__default=True))] + [small, large], [resolved(recipe("qwen35-9b-lora-64k", default=True))] ) routes = DeploymentRoutes(map(definition, switched)) - assert ( - routes.select(small.spec.model.id, "lora").DEFINITION_ID == large.definition_id - ) + assert routes.select(small.spec.model, "lora").DEFINITION_ID == large.definition_id # Saved model/checkpoint records continue using their original definition. assert ( routes.select(small.definition_id, "lora").DEFINITION_ID == small.definition_id ) assert routes.capabilities()[0]["max_context_length"] == 65536 with pytest.raises(ValueError, match="multiple defaults"): - validate_frontend( - [small.spec, recipe("qwen35-9b-lora-64k", routing__default=True)] - ) + validate_frontend([small.spec, recipe("qwen35-9b-lora-64k", default=True)]) def test_ambiguous_model_does_not_get_random_configuration(): rows = [ - resolved(recipe(routing__default=False)), + resolved(recipe(default=False)), resolved(recipe("qwen35-9b-lora-64k")), ] routes = DeploymentRoutes(map(definition, rows)) with pytest.raises(ValueError, match="ambiguous.*16k.*64k"): - routes.select(rows[0].spec.model.id, "lora") + routes.select(rows[0].spec.model, "lora") assert routes.capabilities() == [] def test_sampling_requires_default_across_training_modes(): lora = resolved() - fft = resolved(recipe("qwen35-4b-fft-64k", model__id=lora.spec.model.id)) + fft = resolved(recipe("qwen35-4b-fft-64k", model=lora.spec.model)) routes = DeploymentRoutes(map(definition, [lora, fft])) with pytest.raises(ValueError, match="sampling_default"): - routes.sampling(lora.spec.model.id) - fft.spec = replace( - fft.spec, routing=replace(fft.spec.routing, sampling_default=True) - ) + routes.sampling(lora.spec.model) + fft.spec = replace(fft.spec, sampling_default=True) assert ( DeploymentRoutes(map(definition, [lora, fft])) - .sampling(lora.spec.model.id) + .sampling(lora.spec.model) .DEFINITION_ID == fft.definition_id ) @@ -199,7 +191,7 @@ async def run(): store = InMemoryKeyValueStore() plane = ControlPlane(store, SimpleNamespace()) first = resolved() - other = resolved(recipe(name="other", model__id="org/other-model")) + other = resolved(recipe(name="other", model="org/other-model")) app = create_control_plane_app( plane, list(map(definition, [first, other])), api_key="test" ) @@ -217,7 +209,7 @@ async def run(): json={ "session_id": session, "model_seq_id": seq, - "base_model": row.spec.model.id, + "base_model": row.spec.model, "lora_config": {"rank": 32}, }, ) @@ -284,7 +276,7 @@ def test_native_sections_survive_serialization_without_allowlist(): settings = backend_config(spec, "/assets/pinned") config, _ = parse_backend_config(json.loads(json.dumps(settings))) assert config.hf_checkpoint == "/assets/pinned" - assert config.seq_length == spec.model.max_context_length + assert config.seq_length == spec.max_context_length assert config.provider_overrides["future_provider_option"] == { "layers": [1, 4], "enabled": False, @@ -351,10 +343,11 @@ def test_record_creation_copies_without_reparsing(): original = asdict(spec) row = DeploymentRecord.create(spec, revision="a" * 40) assert asdict(spec) == original - assert row.spec.model.revision == "a" * 40 + assert row.spec.revision == "a" * 40 # The record hash covers settings and the pinned backend dependency, not source. - expected = original | {"model": original["model"] | {"revision": "a" * 40}} - expected.pop("routing") + expected = original | {"revision": "a" * 40} + expected.pop("default") + expected.pop("sampling_default") assert ( row.generation == hashlib.sha256( @@ -383,17 +376,18 @@ def test_record_creation_copies_without_reparsing(): def test_python_config_composition(tmp_path): path = tmp_path / "model.py" path.write_text( - "from dataclasses import replace\n" - "from lilo.configs.qwen35_9b_lora_16k import config as base\n" - "config = replace(base, name='custom', trainer=replace(base.trainer, " - "compute=replace(base.trainer.compute, memory_mib=123456)))\n" + "from lilo.configs.qwen35_9b_lora_16k import Config as Parent\n" + "class Config(Parent):\n" + " name = 'custom'\n" + " trainer = {**Parent.trainer, 'memory_mib': 123456}\n" + "config = Config()\n" ) custom = load(path) original = recipe() assert custom.name == "custom" - assert custom.trainer.compute.memory_mib == 123456 + assert custom.trainer.memory_mib == 123456 assert custom.trainer.config == original.trainer.config - assert original.trainer.compute.memory_mib == 65536 + assert original.trainer.memory_mib == 65536 with pytest.raises(FrozenInstanceError): custom.trainer.max_instances = 9 @@ -411,7 +405,7 @@ def test_worker_record_contains_resolved_settings(monkeypatch): def test_config_file_must_export_config_subclass(tmp_path, source): path = tmp_path / "model.py" path.write_text(source) - with pytest.raises(ValueError, match="Deployment object"): + with pytest.raises(ValueError, match="BaseConfig instance"): load(path) @@ -434,17 +428,16 @@ def test_no_yaml_config_ingestion(tmp_path): @pytest.mark.parametrize( "section,key", [ - ("compute", "memroy_mib"), + ("trainer", "memroy_mib"), ("trainer", "max_instnaces"), ("trainer", "min_instances"), ("inference", "timeout_s"), - ("compute", "timeout_s"), + ("inference", "nodes"), ], ) def test_orchestration_typos_and_unused_fields_are_rejected(section, key): base = recipe() component = { - "compute": base.trainer.compute, "trainer": base.trainer, "inference": base.inference, }[section] @@ -466,7 +459,7 @@ def test_worker_hashes_cover_only_their_settings(): assert adapter.trainer_hash != base.trainer_hash assert adapter.inference_hash != base.inference_hash - routing = resolved(recipe(routing__default=False)) + routing = resolved(recipe(default=False)) assert routing.trainer_hash == base.trainer_hash assert routing.inference_hash == base.inference_hash assert routing.generation == base.generation @@ -485,7 +478,7 @@ def test_examples_only_contain_model_infrastructure(): for path in Path(config_path("qwen35-9b-lora-16k")).parent.glob("qwen*.py"): config = load(path) assert not hasattr(config, "deployment") - assert config.model.revision == "main" + assert config.revision == "main" assert "runtime_version" not in asdict(config.trainer) assert "runtime_version" not in asdict(config.inference) @@ -535,3 +528,53 @@ def test_multinode_ownership_and_topology(): ) with pytest.raises(ValueError, match="actor_num_nodes"): DeploymentRecord.create(invalid, revision="a" * 40) + + +def test_config_inheritance_and_constructor_overrides_copy_nested_options(): + from lilo.configs.qwen35_9b_lora_16k import Config as Parent + + class Child(Parent): + name = "child" + max_context_length = 8192 + trainer = { + **Parent.trainer, + "config": {**Parent.trainer["config"], "max_tokens_per_gpu": 8192}, + } + + first, second = Child(), Child(name="second") + first.trainer.config["target_modules"].append("extra") + first.trainer.config["cli_options"]["recompute_num_layers"] = 2 + assert second.name == "second" + assert second.max_context_length == 8192 + assert second.trainer.gpu == "H100" + assert backend_config(second)["miles"]["max_tokens_per_gpu"] == 8192 + assert Parent().max_context_length == 16384 + for config in (second, Parent()): + assert "extra" not in config.trainer.config["target_modules"] + assert config.trainer.config["cli_options"]["recompute_num_layers"] == 1 + + +@pytest.mark.parametrize( + "field,value", + [("max_contex_length", 8192), ("max_context_length", "8192"), ("default", "false")], +) +def test_recipe_class_fields_and_constructor_overrides_are_validated(field, value): + from lilo.configs.qwen35_9b_lora_16k import Config as Parent + + cls = type("Invalid", (Parent,), {field: value}) + with pytest.raises(ValueError, match=field): + cls() + with pytest.raises(ValueError, match=field): + Parent(**{field: value}) + + +def test_recipe_section_replacement_uses_defaults_without_implicit_merge(): + from lilo.configs.qwen35_9b_lora_16k import Config as Parent + + class Child(Parent): + inference = {"gpu": "H100", "config": {"max_running_requests": 4}} + + config = Child() + assert config.inference.gpu == "H100" + assert config.inference.max_replicas == 8 + assert config.inference.config == {"max_running_requests": 4} From dc701552c69ba1a8162a26e7b146e57c1579c182 Mon Sep 17 00:00:00 2001 From: kailash Date: Thu, 24 Sep 2026 16:18:20 +0000 Subject: [PATCH 22/27] Use dotted overrides for inherited deployment recipes --- docs/deployment-configs.md | 12 ++- docs/deployment-validation.md | 6 +- src/lilo/configs/qwen35_9b_fft_64k.py | 14 ++- .../qwen35_9b_instruct_lora_16k_dp2.py | 5 +- src/lilo/configs/qwen35_9b_lora_16k_single.py | 2 +- src/lilo/configs/qwen35_9b_lora_2k.py | 32 +++---- src/lilo/configs/qwen35_9b_lora_64k.py | 16 ++-- src/lilo/configs/qwen36_35b_a3b_fft_64k.py | 10 +-- src/lilo/configs/qwen38_27b_lora_128k.py | 30 +++---- src/lilo/configs/qwen38_27b_lora_256k.py | 41 +++------ src/lilo/configs/qwen38_27b_lora_64k.py | 19 ++-- src/lilo/configuration.py | 43 ++++++--- tests/test_deployments.py | 89 +++++++++++++++++-- 13 files changed, 186 insertions(+), 133 deletions(-) diff --git a/docs/deployment-configs.md b/docs/deployment-configs.md index f19a7d2..1a04f04 100644 --- a/docs/deployment-configs.md +++ b/docs/deployment-configs.md @@ -34,7 +34,7 @@ See the [9B LoRA recipe](../src/lilo/configs/qwen35_9b_lora_16k.py) and [4B FFT ## Variants -Use Python inheritance to change a recipe. Extend a section explicitly when you want to keep its other settings: +Use Python inheritance and an `overrides` dictionary to change only the settings you need: ```python from lilo.configs.qwen35_9b_lora_16k import Config as Parent @@ -42,13 +42,19 @@ from lilo.configs.qwen35_9b_lora_16k import Config as Parent class Config(Parent): name = "my-9b-more-memory" - trainer = {**Parent.trainer, "memory_mib": 98304} + overrides = { + "trainer.memory_mib": 98304, + "trainer.config.max_tokens_per_gpu": 8192, + "inference.gpu": "H200", + } config = Config() ``` -Assigning a new dictionary replaces that section. There is no implicit deep merge or dotted override language. Extend backend options with `{**Parent.trainer["config"], "max_tokens_per_gpu": 8192}`. Constructor arguments can also override fields: `Config(name="another-run")`. +Dotted paths set individual values. Parent settings and overrides apply first, then child settings and overrides; a child does not need to repeat its parent’s `overrides`. Assigning a dictionary or list replaces the value at that path: `"trainer.env": {}` clears inherited environment settings. Assigning a whole section as a class attribute still replaces that section. + +Constructor fields apply last, followed by constructor overrides: `Config(name="another-run", overrides={"trainer.gpu": "H200"})`. These are ordinary Python values, with no expressions or merge directives. Only the resolved settings are saved for workers. Construction copies the recipe's dictionaries and validates Lilo-owned fields. Unknown fields, invalid types, negative capacities, and inconsistent scaling limits fail before deployment. The validated instance has attribute access (`config.trainer.gpu`); backend options stay dictionaries. Instances do not share mutable options with each other or with their recipe class. diff --git a/docs/deployment-validation.md b/docs/deployment-validation.md index 889473d..b68969a 100644 --- a/docs/deployment-validation.md +++ b/docs/deployment-validation.md @@ -6,11 +6,11 @@ These checks exercise PR #55. The current implementation deploys trainers and in Recipes use one `BaseConfig` subclass with top-level model settings and plain trainer/inference dictionaries. Removed the `Compute`, `Model`, `Routing`, and `Lifecycle` wrapper classes and the matching nested fields. Compute settings live directly in trainer/inference; routing flags and lifecycle timeouts live at the top level. The supported inference backend remains SGLang, so its redundant backend selector was removed. -Inheritance follows Python attribute replacement. A section is extended explicitly with dictionary unpacking; there is no deep-merge or dotted-override parser. Construction copies class dictionaries and validates against one schema, and workers continue to consume the saved resolved backend settings. +Variants use dotted `overrides` dictionaries to change inherited settings. Each parent’s settings and overrides apply before its child’s; dictionary and list values replace the value at their path. Constructor fields and overrides apply last. Construction copies mutable settings and validates against the existing schema. Workers continue to consume saved resolved backend settings. -Validation: **656 CPU tests passed, 1 skipped** (including **122 focused config/deployment tests**). Coverage includes inherited recipes, constructor overrides, nested option isolation, strict field validation, CLI revision resolution, saved-record round trips, and worker construction. The CLI-generated recipe also loaded and validated successfully. Ruff and whitespace checks passed. The full suite ran outside the socket-restricted sandbox so its local HTTP and SDK tests could run. +Validation: **666 CPU tests passed, 1 skipped** (including **132 focused config/deployment tests**). Coverage includes multilevel inherited overrides, constructor precedence, dictionary/list replacement, nested option isolation, invalid override paths, strict field validation, CLI revision resolution, saved-record round trips, and worker construction. The CLI-generated recipe also loaded and validated successfully. Ruff and whitespace checks passed. The full suite ran outside the socket-restricted sandbox so its local HTTP and SDK tests could run. -All 15 recipes were compared with their pre-change values: compute, model, routing, lifecycle, and resolved trainer/inference settings are unchanged. No deployments or GPU jobs were modified. Existing draft manifests use the previous schema and require a fresh registry or explicit migration. +All 15 recipes were compared with their pre-change values: compute, model, routing, lifecycle, and resolved trainer/inference settings are unchanged. The conversion to dotted overrides also preserves every deployment generation and trainer/inference hash. No deployments or GPU jobs were modified. Existing draft manifests use the previous schema and require a fresh registry or explicit migration. ## Historical validation diff --git a/src/lilo/configs/qwen35_9b_fft_64k.py b/src/lilo/configs/qwen35_9b_fft_64k.py index cc0d414..a23116c 100644 --- a/src/lilo/configs/qwen35_9b_fft_64k.py +++ b/src/lilo/configs/qwen35_9b_fft_64k.py @@ -4,18 +4,14 @@ class Config(Parent): name = "qwen35-9b-fft-64k" model = "Qwen/Qwen3.5-9B" - trainer = { - **Parent.trainer, - "gpu": "H200", - "env": { + overrides = { + "trainer.gpu": "H200", + "trainer.env": { "PYTORCH_CUDA_ALLOC_CONF": "expandable_segments:True", "TORCHINDUCTOR_COMPILE_THREADS": "1", }, - } - inference = { - **Parent.inference, - "gpu": "H200", - "config": {**Parent.inference["config"], "ep_size": 1}, + "inference.gpu": "H200", + "inference.config.ep_size": 1, } diff --git a/src/lilo/configs/qwen35_9b_instruct_lora_16k_dp2.py b/src/lilo/configs/qwen35_9b_instruct_lora_16k_dp2.py index d749920..5127c4b 100644 --- a/src/lilo/configs/qwen35_9b_instruct_lora_16k_dp2.py +++ b/src/lilo/configs/qwen35_9b_instruct_lora_16k_dp2.py @@ -3,11 +3,8 @@ class Config(Parent): name = "qwen35-9b-instruct-lora-16k-dp2" - trainer = { - **Parent.trainer, - "config": {**Parent.trainer["config"], "tensor_model_parallel_size": 4}, - } default = False + overrides = {"trainer.config.tensor_model_parallel_size": 4} config = Config() diff --git a/src/lilo/configs/qwen35_9b_lora_16k_single.py b/src/lilo/configs/qwen35_9b_lora_16k_single.py index 0e84a04..db0bdf7 100644 --- a/src/lilo/configs/qwen35_9b_lora_16k_single.py +++ b/src/lilo/configs/qwen35_9b_lora_16k_single.py @@ -3,8 +3,8 @@ class Config(Parent): name = "qwen35-9b-lora-16k-single" - trainer = {**Parent.trainer, "max_clients_per_instance": 1} default = False + overrides = {"trainer.max_clients_per_instance": 1} config = Config() diff --git a/src/lilo/configs/qwen35_9b_lora_2k.py b/src/lilo/configs/qwen35_9b_lora_2k.py index 6837b08..8635fea 100644 --- a/src/lilo/configs/qwen35_9b_lora_2k.py +++ b/src/lilo/configs/qwen35_9b_lora_2k.py @@ -4,28 +4,18 @@ class Config(Parent): name = "qwen35-9b-lora-2k" max_context_length = 2048 - trainer = { - **Parent.trainer, - "gpu": "H200", - "cpu": 8, - "memory_mib": 32768, - "max_clients_per_instance": 4, - "config": { - **Parent.trainer["config"], - "max_tokens_per_gpu": 2048, - "max_lora_slots": 4, - }, - } - inference = { - **Parent.inference, - "config": { - **Parent.inference["config"], - "ep_size": 1, - "max_loaded_loras": 32, - "schedule_policy": "lpm", - }, - } default = False + overrides = { + "trainer.gpu": "H200", + "trainer.cpu": 8, + "trainer.memory_mib": 32768, + "trainer.max_clients_per_instance": 4, + "trainer.config.max_tokens_per_gpu": 2048, + "trainer.config.max_lora_slots": 4, + "inference.config.ep_size": 1, + "inference.config.max_loaded_loras": 32, + "inference.config.schedule_policy": "lpm", + } config = Config() diff --git a/src/lilo/configs/qwen35_9b_lora_64k.py b/src/lilo/configs/qwen35_9b_lora_64k.py index 0e343c2..f888f2e 100644 --- a/src/lilo/configs/qwen35_9b_lora_64k.py +++ b/src/lilo/configs/qwen35_9b_lora_64k.py @@ -4,17 +4,13 @@ class Config(Parent): name = "qwen35-9b-lora-64k" max_context_length = 65536 - trainer = { - **Parent.trainer, - "gpu": "H200", - "gpus_per_node": 8, - "config": { - **Parent.trainer["config"], - "tensor_model_parallel_size": 8, - "max_tokens_per_gpu": 65536, - }, - } default = False + overrides = { + "trainer.gpu": "H200", + "trainer.gpus_per_node": 8, + "trainer.config.tensor_model_parallel_size": 8, + "trainer.config.max_tokens_per_gpu": 65536, + } config = Config() diff --git a/src/lilo/configs/qwen36_35b_a3b_fft_64k.py b/src/lilo/configs/qwen36_35b_a3b_fft_64k.py index 1a82fee..cd39cc1 100644 --- a/src/lilo/configs/qwen36_35b_a3b_fft_64k.py +++ b/src/lilo/configs/qwen36_35b_a3b_fft_64k.py @@ -4,13 +4,9 @@ class Config(Parent): name = "qwen36-35b-a3b-fft-64k" model = "Qwen/Qwen3.6-35B-A3B" - inference = { - **Parent.inference, - "config": { - **Parent.inference["config"], - "dp_size": 1, - "enable_dp_attention": False, - }, + overrides = { + "inference.config.dp_size": 1, + "inference.config.enable_dp_attention": False, } diff --git a/src/lilo/configs/qwen38_27b_lora_128k.py b/src/lilo/configs/qwen38_27b_lora_128k.py index 752b4b1..b53040a 100644 --- a/src/lilo/configs/qwen38_27b_lora_128k.py +++ b/src/lilo/configs/qwen38_27b_lora_128k.py @@ -4,27 +4,17 @@ class Config(Parent): name = "qwen38-27b-lora-128k" max_context_length = 131072 - trainer = { - **Parent.trainer, - "config": { - **Parent.trainer["config"], - "tensor_model_parallel_size": 2, - "max_tokens_per_gpu": 32768, - "context_parallel_size": 4, - }, - } - inference = { - **Parent.inference, - "gpus_per_node": 2, - "max_replicas": 4, - "target_concurrency": 4, - "config": { - **Parent.inference["config"], - "tp_size": 2, - "max_running_requests": 8, - }, - } default = False + overrides = { + "trainer.config.tensor_model_parallel_size": 2, + "trainer.config.max_tokens_per_gpu": 32768, + "trainer.config.context_parallel_size": 4, + "inference.gpus_per_node": 2, + "inference.max_replicas": 4, + "inference.target_concurrency": 4, + "inference.config.tp_size": 2, + "inference.config.max_running_requests": 8, + } config = Config() diff --git a/src/lilo/configs/qwen38_27b_lora_256k.py b/src/lilo/configs/qwen38_27b_lora_256k.py index a8961ee..ae45f59 100644 --- a/src/lilo/configs/qwen38_27b_lora_256k.py +++ b/src/lilo/configs/qwen38_27b_lora_256k.py @@ -5,33 +5,20 @@ class Config(Parent): name = "qwen38-27b-lora-256k" max_context_length = 262144 default = False - trainer = { - **Parent.trainer, - "nodes": 2, - "config": { - **Parent.trainer["config"], - "tensor_model_parallel_size": 2, - "context_parallel_size": 8, - "max_tokens_per_gpu": 32768, - "cli_options": { - **Parent.trainer["config"]["cli_options"], - "distributed_timeout_minutes": 120, - }, - }, - } - inference = { - **Parent.inference, - "gpus_per_node": 4, - "min_replicas": 2, - "max_replicas": 2, - "target_concurrency": 2, - "config": { - **Parent.inference["config"], - "tp_size": 4, - "max_running_requests": 4, - "max_queued_requests": 8, - "max_loaded_loras": 256, - }, + overrides = { + "trainer.nodes": 2, + "trainer.config.tensor_model_parallel_size": 2, + "trainer.config.context_parallel_size": 8, + "trainer.config.max_tokens_per_gpu": 32768, + "trainer.config.cli_options.distributed_timeout_minutes": 120, + "inference.gpus_per_node": 4, + "inference.min_replicas": 2, + "inference.max_replicas": 2, + "inference.target_concurrency": 2, + "inference.config.tp_size": 4, + "inference.config.max_running_requests": 4, + "inference.config.max_queued_requests": 8, + "inference.config.max_loaded_loras": 256, } diff --git a/src/lilo/configs/qwen38_27b_lora_64k.py b/src/lilo/configs/qwen38_27b_lora_64k.py index 71cb28b..b036095 100644 --- a/src/lilo/configs/qwen38_27b_lora_64k.py +++ b/src/lilo/configs/qwen38_27b_lora_64k.py @@ -4,20 +4,13 @@ class Config(Parent): name = "qwen38-27b-lora-64k" max_context_length = 65536 - trainer = { - **Parent.trainer, - "config": { - **Parent.trainer["config"], - "max_tokens_per_gpu": 32768, - "context_parallel_size": 2, - }, - } - inference = { - **Parent.inference, - "target_concurrency": 8, - "config": {**Parent.inference["config"], "max_running_requests": 16}, - } default = False + overrides = { + "trainer.config.max_tokens_per_gpu": 32768, + "trainer.config.context_parallel_size": 2, + "inference.target_concurrency": 8, + "inference.config.max_running_requests": 16, + } config = Config() diff --git a/src/lilo/configuration.py b/src/lilo/configuration.py index 50bcf33..5cff6b1 100644 --- a/src/lilo/configuration.py +++ b/src/lilo/configuration.py @@ -1,5 +1,6 @@ """Readable Python recipes with validated model, trainer, and inference settings.""" +from copy import deepcopy from dataclasses import field from typing import Annotated, Literal @@ -78,23 +79,45 @@ def __post_init__(self): raise ValueError("LILO_ environment variables are managed by Lilo") +def _apply_overrides(values, overrides): + """Set dotted dictionary paths; the deployment schema validates the result.""" + if not isinstance(overrides, dict): + raise ValueError("overrides must be a dictionary of dotted paths") + for path, value in overrides.items(): + if not isinstance(path, str) or any(not part for part in path.split(".")): + raise ValueError(f"invalid override path: {path!r}") + target = values + for part in path.split(".")[:-1]: + target = target.setdefault(part, {}) + if not isinstance(target, dict): + raise ValueError(f"override {path!r} traverses a non-dictionary value") + target[path.rsplit(".", 1)[-1]] = deepcopy(value) + + class BaseConfig(Deployment): - """Author a recipe with class attributes and ordinary Python inheritance. + """Declare a recipe, then change inherited settings with dotted overrides. - Sections are dictionaries in the recipe and validated values in an instance. - Overriding a section replaces it; use ``{**Parent.trainer, ...}`` to extend it. + Each class's attributes and overrides apply in parent-to-child order. + Constructor fields and overrides apply last. Dictionary/list values replace + the value at that path, and each instance owns its mutable settings. """ - def __init__(self, **overrides): - from copy import deepcopy - + def __init__(self, **kwargs): values = {} for cls in reversed(type(self).__mro__): if cls in (object, Deployment, BaseConfig): continue values.update( - (key, value) - for key, value in vars(cls).items() - if not key.startswith("_") + deepcopy( + { + key: value + for key, value in vars(cls).items() + if not key.startswith("_") and key != "overrides" + } + ) ) - super().__init__(**deepcopy(values | overrides)) + _apply_overrides(values, vars(cls).get("overrides", {})) + overrides = kwargs.pop("overrides", {}) + values.update(deepcopy(kwargs)) + _apply_overrides(values, overrides) + super().__init__(**values) diff --git a/tests/test_deployments.py b/tests/test_deployments.py index 1bcae53..5a942e5 100644 --- a/tests/test_deployments.py +++ b/tests/test_deployments.py @@ -379,7 +379,7 @@ def test_python_config_composition(tmp_path): "from lilo.configs.qwen35_9b_lora_16k import Config as Parent\n" "class Config(Parent):\n" " name = 'custom'\n" - " trainer = {**Parent.trainer, 'memory_mib': 123456}\n" + " overrides = {'trainer.memory_mib': 123456}\n" "config = Config()\n" ) custom = load(path) @@ -536,10 +536,7 @@ def test_config_inheritance_and_constructor_overrides_copy_nested_options(): class Child(Parent): name = "child" max_context_length = 8192 - trainer = { - **Parent.trainer, - "config": {**Parent.trainer["config"], "max_tokens_per_gpu": 8192}, - } + overrides = {"trainer.config.max_tokens_per_gpu": 8192} first, second = Child(), Child(name="second") first.trainer.config["target_modules"].append("extra") @@ -578,3 +575,85 @@ class Child(Parent): assert config.inference.gpu == "H100" assert config.inference.max_replicas == 8 assert config.inference.config == {"max_running_requests": 4} + + +def test_overrides_compose_across_generations_and_constructor(): + from lilo.configs.qwen35_9b_lora_16k import Config as Parent + + class Child(Parent): + overrides = { + "trainer.gpu": "H200", + "trainer.env.FIRST": "1", + "trainer.config.max_tokens_per_gpu": 8192, + "inference.config.future_option.nested": [1, 2], + } + + class Grandchild(Child): + overrides = { + "trainer.config.max_tokens_per_gpu": 4096, + "trainer.env.SECOND": "2", + "default": False, + } + + config = Grandchild( + name="custom", + overrides={"trainer.config.max_tokens_per_gpu": 2048}, + ) + assert config.name == "custom" + assert config.trainer.gpu == "H200" + assert config.trainer.env["FIRST"] == "1" + assert config.trainer.env["SECOND"] == "2" + assert config.trainer.config["max_tokens_per_gpu"] == 2048 + assert not config.default + assert Grandchild().trainer.config["max_tokens_per_gpu"] == 4096 + assert Child().trainer.config["max_tokens_per_gpu"] == 8192 + config.inference.config["future_option"]["nested"].append(3) + assert Child.overrides["inference.config.future_option.nested"] == [1, 2] + assert Grandchild().inference.config["future_option"]["nested"] == [1, 2] + assert "FIRST" not in Parent().trainer.env + + +def test_override_values_replace_dictionaries_and_lists(): + from lilo.configs.qwen35_9b_lora_16k import Config as Parent + + class Child(Parent): + overrides = {"trainer.env": {"FIRST": "1"}} + + class Grandchild(Child): + overrides = { + "trainer.env": {}, + "trainer.config.target_modules": ["q_proj"], + } + + config = Grandchild() + assert config.trainer.env == {} + assert config.trainer.config["target_modules"] == ["q_proj"] + assert Child().trainer.env == {"FIRST": "1"} + # Constructor fields replace the inherited section, then overrides apply. + config = Grandchild(inference={"gpu": "H100"}, overrides={"inference.gpu": "H200"}) + assert config.inference.gpu == "H200" + assert config.inference.config == {} + assert replace(config, name="copy").trainer == config.trainer + + +@pytest.mark.parametrize( + "overrides,match", + [ + ([], "overrides must be a dictionary"), + ({"": 1}, "invalid override path"), + ({"trainer..gpu": "H200"}, "invalid override path"), + ({1: "H200"}, "invalid override path"), + ({"trainer.gpu.type": "H200"}, "non-dictionary"), + ({"trainer.memroy_mib": 123}, "memroy_mib"), + ({"trianer.gpu": "H200"}, "trianer"), + ({"trainer.gpus_per_node": "4"}, "gpus_per_node"), + ], +) +def test_dotted_overrides_are_validated(overrides, match): + from lilo.configs.qwen35_9b_lora_16k import Config as Parent + + cls = type("Invalid", (Parent,), {"overrides": overrides}) + with pytest.raises(ValueError, match=match): + cls() + with pytest.raises(ValueError, match=match): + Parent(overrides=overrides) From bac50b97e06f2c5815081e11aa6a45767df0b7c4 Mon Sep 17 00:00:00 2001 From: kailash Date: Thu, 24 Sep 2026 16:29:38 +0000 Subject: [PATCH 23/27] Simplify recipe routing to deployment order and explicit IDs --- docs/deployment-configs.md | 4 +- docs/deployment-validation.md | 8 +- scripts/definition_smoke.py | 1 - scripts/deploy_models.sh | 2 +- src/lilo/configs/qwen35_35b_a3b_fft_64k.py | 1 - src/lilo/configs/qwen35_4b_fft_64k.py | 1 - .../configs/qwen35_9b_instruct_lora_16k.py | 1 - .../qwen35_9b_instruct_lora_16k_dp2.py | 1 - src/lilo/configs/qwen35_9b_lora_16k.py | 1 - src/lilo/configs/qwen35_9b_lora_16k_single.py | 1 - src/lilo/configs/qwen35_9b_lora_2k.py | 1 - src/lilo/configs/qwen35_9b_lora_64k.py | 1 - src/lilo/configs/qwen36_27b_fft_64k.py | 1 - src/lilo/configs/qwen38_27b_lora_128k.py | 1 - src/lilo/configs/qwen38_27b_lora_16k.py | 1 - src/lilo/configs/qwen38_27b_lora_256k.py | 1 - src/lilo/configs/qwen38_27b_lora_64k.py | 1 - src/lilo/configuration.py | 4 +- src/lilo/control_plane/deployments.py | 101 +++++------------- src/lilo/control_plane/http.py | 19 ++-- src/lilo/deployments.py | 15 +-- src/lilo/providers/modal/deployment_apps.py | 3 - src/lilo/providers/modal/scoped.py | 1 - tests/control_plane/test_http.py | 16 +-- tests/control_plane/test_sdk_e2e.py | 2 - tests/providers/test_definition_registry.py | 9 -- tests/providers/test_deployment_apps.py | 7 +- tests/test_deployments.py | 70 ++++++------ tests/test_system.py | 1 - 29 files changed, 80 insertions(+), 196 deletions(-) diff --git a/docs/deployment-configs.md b/docs/deployment-configs.md index 1a04f04..9cd4633 100644 --- a/docs/deployment-configs.md +++ b/docs/deployment-configs.md @@ -127,13 +127,13 @@ The checked-in [deploy_models.sh](../scripts/deploy_models.sh) lists the complet These settings are not model-config fields. Secret/volume names come from provider defaults. Credentials remain in Modal secrets. -One frontend serves all models through Tinker's base_model. `default = True` selects among multiple training configurations for one model; `sampling_default = True` selects a sampling configuration when LoRA/FFT configurations coexist. +One frontend serves all models through Tinker’s `base_model`. When recipes share a model, the first matching recipe in the `lilo deploy` argument list is used. Training also matches the requested LoRA/FFT mode; base sampling uses the first recipe regardless of training mode. To select a specific recipe, pass its definition ID from `/api/v1/lilo/deployments` as `base_model`. All supplied definitions are listed, including retained generations. Current recipes precede retained ones; existing jobs keep their saved definition IDs. ## Hashes and update isolation Hashes identify settings, not compatibility with a source checkout: -- generation identifies the saved configuration, resolved backend settings, platform settings and code releases. Routing defaults are excluded. +- generation identifies the saved configuration, resolved backend settings, platform settings and code releases. - trainer_hash includes trainer/model/platform settings, resolved trainer settings, Miles commit and its recorded code release. - inference_hash includes inference/model/platform settings, resolved inference settings and its recorded code release. diff --git a/docs/deployment-validation.md b/docs/deployment-validation.md index b68969a..83fbfd2 100644 --- a/docs/deployment-validation.md +++ b/docs/deployment-validation.md @@ -4,13 +4,15 @@ These checks exercise PR #55. The current implementation deploys trainers and in ## Current: simple BaseConfig recipes -Recipes use one `BaseConfig` subclass with top-level model settings and plain trainer/inference dictionaries. Removed the `Compute`, `Model`, `Routing`, and `Lifecycle` wrapper classes and the matching nested fields. Compute settings live directly in trainer/inference; routing flags and lifecycle timeouts live at the top level. The supported inference backend remains SGLang, so its redundant backend selector was removed. +Recipes use one `BaseConfig` subclass with top-level model settings and plain trainer/inference dictionaries. Removed the `Compute`, `Model`, `Routing`, and `Lifecycle` wrapper classes and the matching nested fields. Compute settings live directly in trainer/inference; lifecycle timeouts live at the top level. The supported inference backend remains SGLang, so its redundant backend selector was removed. Variants use dotted `overrides` dictionaries to change inherited settings. Each parent’s settings and overrides apply before its child’s; dictionary and list values replace the value at their path. Constructor fields and overrides apply last. Construction copies mutable settings and validates against the existing schema. Workers continue to consume saved resolved backend settings. -Validation: **666 CPU tests passed, 1 skipped** (including **132 focused config/deployment tests**). Coverage includes multilevel inherited overrides, constructor precedence, dictionary/list replacement, nested option isolation, invalid override paths, strict field validation, CLI revision resolution, saved-record round trips, and worker construction. The CLI-generated recipe also loaded and validated successfully. Ruff and whitespace checks passed. The full suite ran outside the socket-restricted sandbox so its local HTTP and SDK tests could run. +Validation: **664 CPU tests passed, 1 skipped**. Coverage includes multilevel inherited overrides, constructor precedence, dictionary/list replacement, nested option isolation, invalid override paths, strict field validation, CLI revision resolution, saved-record round trips, and worker construction. The CLI-generated recipe also loaded and validated successfully. Ruff and whitespace checks passed. The full suite ran outside the socket-restricted sandbox so its local HTTP and SDK tests could run. -All 15 recipes were compared with their pre-change values: compute, model, routing, lifecycle, and resolved trainer/inference settings are unchanged. The conversion to dotted overrides also preserves every deployment generation and trainer/inference hash. No deployments or GPU jobs were modified. Existing draft manifests use the previous schema and require a fresh registry or explicit migration. +Routing uses the first matching recipe in deployment order, or an explicit definition ID. Visibility and training/sampling default flags have been removed. Tests cover recipe ordering, LoRA/FFT selection, explicit IDs, listing all definitions, and retained generations. + +All 15 recipes preserve compute, model, lifecycle, resolved trainer/inference settings, and deployment hashes. Routing flags were removed from the recipe schema; retained definitions remain selectable and listed. No deployments or GPU jobs were modified. Existing draft manifests use the previous schema and require a fresh registry or explicit migration. ## Historical validation diff --git a/scripts/definition_smoke.py b/scripts/definition_smoke.py index 842c339..5cc2c9f 100644 --- a/scripts/definition_smoke.py +++ b/scripts/definition_smoke.py @@ -316,7 +316,6 @@ def main() -> None: f"{definition.DEFINITION_ID:40} {definition.PARAMETERIZATION:5} " f"{definition.RESOLVED.spec.trainer.nodes} nodes x {definition.RESOLVED.spec.trainer.gpu}:{definition.RESOLVED.spec.trainer.gpus_per_node} " f"ctx={definition.MAX_CONTEXT_LENGTH}" - + ("" if definition.CATALOG_VISIBLE else " (not cataloged)") ) return if not args.base_url: diff --git a/scripts/deploy_models.sh b/scripts/deploy_models.sh index 2e59ffc..6d06c17 100755 --- a/scripts/deploy_models.sh +++ b/scripts/deploy_models.sh @@ -5,7 +5,7 @@ set -e cd "$(dirname "$0")/.." # Add a model by creating its Python config in src/lilo/configs/ and adding it here. -# Keep every configuration that should be available to new clients in this list. +# Put the preferred recipe first when multiple configs share a model. deployment_files=( src/lilo/configs/qwen35_9b_lora_16k.py src/lilo/configs/qwen35_9b_lora_64k.py diff --git a/src/lilo/configs/qwen35_35b_a3b_fft_64k.py b/src/lilo/configs/qwen35_35b_a3b_fft_64k.py index aaa0e7d..af2e356 100644 --- a/src/lilo/configs/qwen35_35b_a3b_fft_64k.py +++ b/src/lilo/configs/qwen35_35b_a3b_fft_64k.py @@ -56,7 +56,6 @@ class Config(BaseConfig): "enable_dp_attention": True, }, } - default = True config = Config() diff --git a/src/lilo/configs/qwen35_4b_fft_64k.py b/src/lilo/configs/qwen35_4b_fft_64k.py index 8bc58d1..d45b9df 100644 --- a/src/lilo/configs/qwen35_4b_fft_64k.py +++ b/src/lilo/configs/qwen35_4b_fft_64k.py @@ -39,7 +39,6 @@ class Config(BaseConfig): "cpu_weight_cache_max_compile_group_gb": 16, }, } - default = True config = Config() diff --git a/src/lilo/configs/qwen35_9b_instruct_lora_16k.py b/src/lilo/configs/qwen35_9b_instruct_lora_16k.py index 8cbeabe..1aedccf 100644 --- a/src/lilo/configs/qwen35_9b_instruct_lora_16k.py +++ b/src/lilo/configs/qwen35_9b_instruct_lora_16k.py @@ -41,7 +41,6 @@ class Config(BaseConfig): "schedule_policy": "lpm", }, } - default = True config = Config() diff --git a/src/lilo/configs/qwen35_9b_instruct_lora_16k_dp2.py b/src/lilo/configs/qwen35_9b_instruct_lora_16k_dp2.py index 5127c4b..caaeabe 100644 --- a/src/lilo/configs/qwen35_9b_instruct_lora_16k_dp2.py +++ b/src/lilo/configs/qwen35_9b_instruct_lora_16k_dp2.py @@ -3,7 +3,6 @@ class Config(Parent): name = "qwen35-9b-instruct-lora-16k-dp2" - default = False overrides = {"trainer.config.tensor_model_parallel_size": 4} diff --git a/src/lilo/configs/qwen35_9b_lora_16k.py b/src/lilo/configs/qwen35_9b_lora_16k.py index c04d24c..586c3de 100644 --- a/src/lilo/configs/qwen35_9b_lora_16k.py +++ b/src/lilo/configs/qwen35_9b_lora_16k.py @@ -47,7 +47,6 @@ class Config(BaseConfig): "max_loras_per_batch": 8, }, } - default = True config = Config() diff --git a/src/lilo/configs/qwen35_9b_lora_16k_single.py b/src/lilo/configs/qwen35_9b_lora_16k_single.py index db0bdf7..1cf0afb 100644 --- a/src/lilo/configs/qwen35_9b_lora_16k_single.py +++ b/src/lilo/configs/qwen35_9b_lora_16k_single.py @@ -3,7 +3,6 @@ class Config(Parent): name = "qwen35-9b-lora-16k-single" - default = False overrides = {"trainer.max_clients_per_instance": 1} diff --git a/src/lilo/configs/qwen35_9b_lora_2k.py b/src/lilo/configs/qwen35_9b_lora_2k.py index 8635fea..a3b149d 100644 --- a/src/lilo/configs/qwen35_9b_lora_2k.py +++ b/src/lilo/configs/qwen35_9b_lora_2k.py @@ -4,7 +4,6 @@ class Config(Parent): name = "qwen35-9b-lora-2k" max_context_length = 2048 - default = False overrides = { "trainer.gpu": "H200", "trainer.cpu": 8, diff --git a/src/lilo/configs/qwen35_9b_lora_64k.py b/src/lilo/configs/qwen35_9b_lora_64k.py index f888f2e..3132989 100644 --- a/src/lilo/configs/qwen35_9b_lora_64k.py +++ b/src/lilo/configs/qwen35_9b_lora_64k.py @@ -4,7 +4,6 @@ class Config(Parent): name = "qwen35-9b-lora-64k" max_context_length = 65536 - default = False overrides = { "trainer.gpu": "H200", "trainer.gpus_per_node": 8, diff --git a/src/lilo/configs/qwen36_27b_fft_64k.py b/src/lilo/configs/qwen36_27b_fft_64k.py index 1a4ef5f..e185d89 100644 --- a/src/lilo/configs/qwen36_27b_fft_64k.py +++ b/src/lilo/configs/qwen36_27b_fft_64k.py @@ -47,7 +47,6 @@ class Config(BaseConfig): "cpu_weight_cache_max_compile_group_gb": 32, }, } - default = True config = Config() diff --git a/src/lilo/configs/qwen38_27b_lora_128k.py b/src/lilo/configs/qwen38_27b_lora_128k.py index b53040a..df6ff85 100644 --- a/src/lilo/configs/qwen38_27b_lora_128k.py +++ b/src/lilo/configs/qwen38_27b_lora_128k.py @@ -4,7 +4,6 @@ class Config(Parent): name = "qwen38-27b-lora-128k" max_context_length = 131072 - default = False overrides = { "trainer.config.tensor_model_parallel_size": 2, "trainer.config.max_tokens_per_gpu": 32768, diff --git a/src/lilo/configs/qwen38_27b_lora_16k.py b/src/lilo/configs/qwen38_27b_lora_16k.py index f97a257..db54bce 100644 --- a/src/lilo/configs/qwen38_27b_lora_16k.py +++ b/src/lilo/configs/qwen38_27b_lora_16k.py @@ -41,7 +41,6 @@ class Config(BaseConfig): "schedule_policy": "lpm", }, } - default = True config = Config() diff --git a/src/lilo/configs/qwen38_27b_lora_256k.py b/src/lilo/configs/qwen38_27b_lora_256k.py index ae45f59..51d7f0f 100644 --- a/src/lilo/configs/qwen38_27b_lora_256k.py +++ b/src/lilo/configs/qwen38_27b_lora_256k.py @@ -4,7 +4,6 @@ class Config(Parent): name = "qwen38-27b-lora-256k" max_context_length = 262144 - default = False overrides = { "trainer.nodes": 2, "trainer.config.tensor_model_parallel_size": 2, diff --git a/src/lilo/configs/qwen38_27b_lora_64k.py b/src/lilo/configs/qwen38_27b_lora_64k.py index b036095..f6f2216 100644 --- a/src/lilo/configs/qwen38_27b_lora_64k.py +++ b/src/lilo/configs/qwen38_27b_lora_64k.py @@ -4,7 +4,6 @@ class Config(Parent): name = "qwen38-27b-lora-64k" max_context_length = 65536 - default = False overrides = { "trainer.config.max_tokens_per_gpu": 32768, "trainer.config.context_parallel_size": 2, diff --git a/src/lilo/configuration.py b/src/lilo/configuration.py index 5cff6b1..34f31f9 100644 --- a/src/lilo/configuration.py +++ b/src/lilo/configuration.py @@ -4,7 +4,7 @@ from dataclasses import field from typing import Annotated, Literal -from pydantic import ConfigDict, Field, StrictBool +from pydantic import ConfigDict, Field from pydantic.dataclasses import dataclass PositiveInt = Annotated[int, Field(strict=True, gt=0)] @@ -54,8 +54,6 @@ class Deployment: inference: Inference parameterization: Literal["lora", "full"] = "lora" revision: NonemptyString = "main" - default: StrictBool = False - sampling_default: StrictBool = False session_idle_timeout_s: PositiveInt = 300 pool_idle_timeout_s: PositiveInt = 300 sweep_interval_s: PositiveInt = 300 diff --git a/src/lilo/control_plane/deployments.py b/src/lilo/control_plane/deployments.py index 1b9fab2..4170440 100644 --- a/src/lilo/control_plane/deployments.py +++ b/src/lilo/control_plane/deployments.py @@ -1,86 +1,33 @@ -"""Model-name routing for configured deployments and scoped engines.""" +"""Select recipes in deployment order, or explicitly by definition ID.""" class DeploymentRoutes: def __init__(self, definitions): - self.all = tuple(definitions) - self.visible = tuple(d for d in self.all if d.CATALOG_VISIBLE) - for model, mode in {(d.MODEL_NAME, d.PARAMETERIZATION) for d in self.visible}: - defaults = [ - d - for d in self.visible - if d.MODEL_NAME == model - and d.PARAMETERIZATION == mode - and getattr(d, "ROUTING_DEFAULT", False) - ] - if len(defaults) > 1: - raise ValueError(f"multiple defaults for {model} ({mode})") + self.definitions = tuple(definitions) - def select(self, model, mode): - explicit = [ - d - for d in self.all - if d.DEFINITION_ID == model and d.PARAMETERIZATION == mode + def select(self, model, mode=None): + candidates = [ + d for d in self.definitions if mode is None or d.PARAMETERIZATION == mode ] - if explicit: - return explicit[0] - matches = [ - d - for d in self.visible - if d.MODEL_NAME == model and d.PARAMETERIZATION == mode - ] - if len(matches) <= 1: - return next(iter(matches), None) - defaults = [d for d in matches if getattr(d, "ROUTING_DEFAULT", False)] - if len(defaults) == 1: - return defaults[0] - choices = ", ".join( - f"{getattr(d, 'DEPLOYMENT_NAME', d.DEFINITION_ID)} ({d.MAX_CONTEXT_LENGTH} tokens)" - for d in matches - ) - raise ValueError( - f"ambiguous {mode} deployment for {model}; configure routing.default: {choices}" - ) - - def sampling(self, model): - requested = [d for d in self.all if d.DEFINITION_ID == model] - if len(requested) == 1: - return requested[0] - matches = [d for d in self.visible if d.MODEL_NAME == model] - explicit = [d for d in matches if getattr(d, "SAMPLING_DEFAULT", False)] - if len(explicit) == 1: - return explicit[0] - if len(explicit) > 1: - raise ValueError(f"multiple sampling defaults for {model}") - selected = [ - d - for mode in {d.PARAMETERIZATION for d in matches} - if (d := self.select(model, mode)) is not None - ] - if len(selected) > 1: - raise ValueError( - f"ambiguous sampling deployment for {model}; configure routing.sampling_default" - ) - return next(iter(selected), None) + for field in ("DEFINITION_ID", "MODEL_NAME"): + for definition in candidates: + if getattr(definition, field) == model: + return definition + return None def capabilities(self): - result = [] - for model in dict.fromkeys(d.MODEL_NAME for d in self.visible): - try: - selected = [ - self.select(model, mode) - for mode in { - d.PARAMETERIZATION - for d in self.visible - if d.MODEL_NAME == model - } - ] - except ValueError: - continue - result.append( - { - "model_name": model, - "max_context_length": min(d.MAX_CONTEXT_LENGTH for d in selected), - } + selected = {} + for definition in self.definitions: + selected.setdefault( + (definition.MODEL_NAME, definition.PARAMETERIZATION), definition + ) + contexts = {} + for (model, _), definition in selected.items(): + contexts[model] = min( + contexts.get(model, definition.MAX_CONTEXT_LENGTH), + definition.MAX_CONTEXT_LENGTH, ) - return result + return [ + {"model_name": model, "max_context_length": context} + for model, context in contexts.items() + ] diff --git a/src/lilo/control_plane/http.py b/src/lilo/control_plane/http.py index 8450cd5..0c5f7e3 100644 --- a/src/lilo/control_plane/http.py +++ b/src/lilo/control_plane/http.py @@ -160,14 +160,11 @@ def create_control_plane_app( retrieve_window: float = 30.0, checkpoint_volume: str = "lilo-checkpoints", ) -> FastAPI: - all_definitions = tuple(definitions) - definitions = tuple( - definition for definition in all_definitions if definition.CATALOG_VISIBLE - ) + definitions = tuple(definitions) from .deployments import DeploymentRoutes - routes = DeploymentRoutes(all_definitions) + routes = DeploymentRoutes(definitions) def definition_for(model_name, parameterization): selected = routes.select(model_name, parameterization) @@ -177,7 +174,7 @@ def supports_model(model_name: str) -> bool: return any( definition.MODEL_NAME == model_name for definition in definitions ) or any( - definition.DEFINITION_ID == model_name for definition in all_definitions + definition.DEFINITION_ID == model_name for definition in definitions ) async def authorize(request: Request) -> None: @@ -280,8 +277,6 @@ async def list_deployments(): "base_model": d.MODEL_NAME, "parameterization": d.PARAMETERIZATION, "max_context_length": d.MAX_CONTEXT_LENGTH, - "default": getattr(d, "ROUTING_DEFAULT", False), - "sampling_default": getattr(d, "SAMPLING_DEFAULT", False), } for d in definitions ] @@ -360,7 +355,7 @@ async def create_model(body: CreateModelBody) -> dict[str, object]: spec={ "base_model": next( d.MODEL_NAME - for d in all_definitions + for d in definitions if d.DEFINITION_ID == definition_id ), "lora_config": body.lora_config, @@ -383,7 +378,7 @@ async def create_sampling_session( status_code=400, detail="base_model or model_path is required", ) - selected = routes.sampling(body.base_model) if body.base_model else None + selected = routes.select(body.base_model) if body.base_model else None definition_id = selected.DEFINITION_ID if selected else None if body.model_path is None and definition_id is None: raise HTTPException( @@ -396,7 +391,7 @@ async def create_sampling_session( base_model=( next( d.MODEL_NAME - for d in all_definitions + for d in definitions if d.DEFINITION_ID == definition_id ) if definition_id @@ -494,7 +489,7 @@ async def load_weights(body: LoadWeightsBody) -> dict[str, str]: base_model=body.base_model, user_metadata=body.user_metadata, optimizer=body.optimizer, - definition_ids={d.DEFINITION_ID for d in all_definitions}, + definition_ids={d.DEFINITION_ID for d in definitions}, ) return { "request_id": creation.request_id, diff --git a/src/lilo/deployments.py b/src/lilo/deployments.py index e73baca..562f3f0 100644 --- a/src/lilo/deployments.py +++ b/src/lilo/deployments.py @@ -51,8 +51,6 @@ def deployment_generation( inference_settings, ): identity = asdict(spec) - identity.pop("default") - identity.pop("sampling_default") return settings_hash( { "config": identity, @@ -75,7 +73,7 @@ class DeploymentRecord(BaseModel): The CLI resolves the model revision before creating this record. Its hash binds jobs to their original configuration across later deploys; - active controls whether new clients can select it. + active marks recipes in the latest deploy command; retained recipes remain usable. """ model_config = ConfigDict(extra="forbid") @@ -241,14 +239,3 @@ def validate_frontend(specs: list[Deployment]) -> None: raise ValueError( "deployments on one frontend must share lifecycle settings" ) - defaults, sampling = set(), set() - for spec in specs: - key = (spec.model, spec.parameterization) - if spec.default: - if key in defaults: - raise ValueError(f"multiple defaults for {key}") - defaults.add(key) - if spec.sampling_default: - if spec.model in sampling: - raise ValueError(f"multiple sampling defaults for {spec.model}") - sampling.add(spec.model) diff --git a/src/lilo/providers/modal/deployment_apps.py b/src/lilo/providers/modal/deployment_apps.py index 1775e45..c88b876 100644 --- a/src/lilo/providers/modal/deployment_apps.py +++ b/src/lilo/providers/modal/deployment_apps.py @@ -209,9 +209,6 @@ def definition_from_spec(resolved, *, register_trainer=True, image=None): MODEL_REVISION=spec.revision, HF_CHECKPOINT=resolved.asset_path, PARAMETERIZATION=spec.parameterization, - CATALOG_VISIBLE=resolved.active, - ROUTING_DEFAULT=spec.default, - SAMPLING_DEFAULT=spec.sampling_default, DEPLOYMENT_NAME=spec.name, RESOLVED=resolved, MAX_CONTEXT_LENGTH=spec.max_context_length, diff --git a/src/lilo/providers/modal/scoped.py b/src/lilo/providers/modal/scoped.py index 814dcd3..a96ed3b 100644 --- a/src/lilo/providers/modal/scoped.py +++ b/src/lilo/providers/modal/scoped.py @@ -455,7 +455,6 @@ async def spawn_sampling(task): checkpoint_root=storage.root, ) definition = SimpleNamespace( - CATALOG_VISIBLE=True, DEFINITION_ID=engine.name, MODEL_NAME=engine.model, PARAMETERIZATION="full", diff --git a/tests/control_plane/test_http.py b/tests/control_plane/test_http.py index 402c2b4..0a32d73 100644 --- a/tests/control_plane/test_http.py +++ b/tests/control_plane/test_http.py @@ -19,14 +19,12 @@ DEFINITION_ID=DEFINITION, MODEL_NAME=BASE_MODEL, PARAMETERIZATION="lora", - CATALOG_VISIBLE=True, MAX_CONTEXT_LENGTH=16_384, ), SimpleNamespace( DEFINITION_ID=f"{DEFINITION}_full", MODEL_NAME=BASE_MODEL, PARAMETERIZATION="full", - CATALOG_VISIBLE=True, MAX_CONTEXT_LENGTH=65_536, ), ) @@ -283,13 +281,13 @@ async def run() -> None: asyncio.run(run()) -def test_base_sampling_session_uses_explicit_sampling_default() -> None: +def test_base_sampling_session_uses_first_deployment() -> None: async def run() -> None: plane = ControlPlane( InMemoryKeyValueStore(), LocalEnginePlatform(DEFINITION, EchoExecutor), ) - definitions = [SimpleNamespace(**vars(d), SAMPLING_DEFAULT=d.PARAMETERIZATION == "full") for d in DEFINITIONS] + definitions = list(reversed(DEFINITIONS)) app = create_control_plane_app(plane, definitions, api_key=None) client = httpx.AsyncClient( base_url="http://control-plane", @@ -467,13 +465,15 @@ def create(seq: int, **body): asyncio.run(run()) -def test_explicit_hidden_deployment_keeps_canonical_model_name() -> None: +def test_explicit_deployment_keeps_canonical_model_name() -> None: async def run(): - hidden = SimpleNamespace(DEFINITION_ID="isolated", MODEL_NAME=BASE_MODEL, - PARAMETERIZATION="lora", CATALOG_VISIBLE=False, MAX_CONTEXT_LENGTH=16384) + explicit = SimpleNamespace(DEFINITION_ID="isolated", MODEL_NAME=BASE_MODEL, + PARAMETERIZATION="lora", MAX_CONTEXT_LENGTH=16384) plane = ControlPlane(InMemoryKeyValueStore(), LocalEnginePlatform("isolated", EchoExecutor)) - app = create_control_plane_app(plane, (*DEFINITIONS, hidden), retrieve_window=1.0) + app = create_control_plane_app(plane, (*DEFINITIONS, explicit), retrieve_window=1.0) async with httpx.AsyncClient(base_url="http://test", transport=httpx.ASGITransport(app=app)) as client: + listed = (await client.get("/api/v1/lilo/deployments")).json()["deployments"] + assert [row["generation"] for row in listed] == [d.DEFINITION_ID for d in (*DEFINITIONS, explicit)] session = (await client.post("/api/v1/create_session", json={"tags": [], "sdk_version": "0.5.0"})).json()["session_id"] response = await client.post("/api/v1/create_model", json={"session_id": session, "model_seq_id": 0, "base_model": "isolated", "lora_config": {"rank": 16}}) diff --git a/tests/control_plane/test_sdk_e2e.py b/tests/control_plane/test_sdk_e2e.py index 2ec1b37..cb7a726 100644 --- a/tests/control_plane/test_sdk_e2e.py +++ b/tests/control_plane/test_sdk_e2e.py @@ -32,7 +32,6 @@ DEFINITION_ID=DEFINITION, MODEL_NAME=BASE_MODEL, PARAMETERIZATION="lora", - CATALOG_VISIBLE=True, MAX_CONTEXT_LENGTH=MAX_CONTEXT_LENGTH, ), ) @@ -41,7 +40,6 @@ DEFINITION_ID=FULL_DEFINITION, MODEL_NAME=BASE_MODEL, PARAMETERIZATION="full", - CATALOG_VISIBLE=True, MAX_CONTEXT_LENGTH=MAX_CONTEXT_LENGTH, ), ) diff --git a/tests/providers/test_definition_registry.py b/tests/providers/test_definition_registry.py index 74e829e..e9c053f 100644 --- a/tests/providers/test_definition_registry.py +++ b/tests/providers/test_definition_registry.py @@ -12,12 +12,3 @@ def test_definition_registry_resolves_every_definition() -> None: definition.PARAMETERIZATION ) assert definition.ENGINE_FUNCTION is not None - - -def test_public_model_parameterizations_are_unique() -> None: - visible = [ - (definition.MODEL_NAME, definition.PARAMETERIZATION) - for definition in DEFINITIONS - if definition.CATALOG_VISIBLE - ] - assert len(visible) == len(set(visible)) diff --git a/tests/providers/test_deployment_apps.py b/tests/providers/test_deployment_apps.py index 242edb9..3d53bc6 100644 --- a/tests/providers/test_deployment_apps.py +++ b/tests/providers/test_deployment_apps.py @@ -232,14 +232,9 @@ def test_admission_changes_preserve_serialized_trainer(builders): old_bytes = serialize(deployment_apps.build_trainer_app(first, image="test")[1]) changed = first.model_copy(deep=True) changed.active = False - changed.spec = replace( - changed.spec, - default=False, - sampling_default=True, - ) new_bytes = serialize(deployment_apps.build_trainer_app(changed, image="test")[1]) assert new_bytes == old_bytes - assert first.active is True and first.spec.default is True + assert first.active is True changed.spec = replace( changed.spec, trainer=replace( diff --git a/tests/test_deployments.py b/tests/test_deployments.py index 5a942e5..c970f37 100644 --- a/tests/test_deployments.py +++ b/tests/test_deployments.py @@ -105,56 +105,51 @@ def test_invalid_integrations_fail_when_building_backend_settings(changes, match def test_generation_and_asset_paths_include_exact_base(): a = resolved() - assert a.generation == resolved(recipe(default=False)).generation assert a.generation != resolved(recipe(trainer__gpu="H200")).generation b = resolved(recipe(model="other/Qwen3.5-9B-Base")) assert a.asset_path != b.asset_path assert a.asset_path != DeploymentRecord.create(a.spec, revision="b" * 40).asset_path -def test_frontend_defaults_and_retained_generations(): +def test_routing_uses_deployment_order_and_preserves_explicit_generations(): small = resolved() large = resolved(recipe("qwen35-9b-lora-64k")) - routes = DeploymentRoutes([definition(small), definition(large)]) + routes = DeploymentRoutes(map(definition, [small, large])) assert routes.select(small.spec.model, "lora").DEFINITION_ID == small.definition_id - switched = retain_generations( - [small, large], [resolved(recipe("qwen35-9b-lora-64k", default=True))] - ) + assert routes.capabilities()[0]["max_context_length"] == 16384 + validate_frontend([small.spec, large.spec]) + + switched = retain_generations([small, large], [large]) routes = DeploymentRoutes(map(definition, switched)) assert routes.select(small.spec.model, "lora").DEFINITION_ID == large.definition_id - # Saved model/checkpoint records continue using their original definition. assert ( routes.select(small.definition_id, "lora").DEFINITION_ID == small.definition_id ) assert routes.capabilities()[0]["max_context_length"] == 65536 - with pytest.raises(ValueError, match="multiple defaults"): - validate_frontend([small.spec, recipe("qwen35-9b-lora-64k", default=True)]) - -def test_ambiguous_model_does_not_get_random_configuration(): - rows = [ - resolved(recipe(default=False)), - resolved(recipe("qwen35-9b-lora-64k")), - ] - routes = DeploymentRoutes(map(definition, rows)) - with pytest.raises(ValueError, match="ambiguous.*16k.*64k"): - routes.select(rows[0].spec.model, "lora") - assert routes.capabilities() == [] + # A retained-only model is still listed and selectable. + other = resolved(recipe(name="other", model="org/other")) + routes = DeploymentRoutes(map(definition, retain_generations([small], [other]))) + assert routes.select(small.spec.model, "lora").DEFINITION_ID == small.definition_id + assert {row["model_name"] for row in routes.capabilities()} == { + small.spec.model, + other.spec.model, + } -def test_sampling_requires_default_across_training_modes(): +def test_sampling_uses_order_and_training_filters_parameterization(): lora = resolved() fft = resolved(recipe("qwen35-4b-fft-64k", model=lora.spec.model)) - routes = DeploymentRoutes(map(definition, [lora, fft])) - with pytest.raises(ValueError, match="sampling_default"): - routes.sampling(lora.spec.model) - fft.spec = replace(fft.spec, sampling_default=True) - assert ( - DeploymentRoutes(map(definition, [lora, fft])) - .sampling(lora.spec.model) - .DEFINITION_ID - == fft.definition_id - ) + for first, second in ((lora, fft), (fft, lora)): + routes = DeploymentRoutes(map(definition, [first, second])) + assert routes.select(lora.spec.model).DEFINITION_ID == first.definition_id + assert ( + routes.select(lora.spec.model, "lora").DEFINITION_ID == lora.definition_id + ) + assert routes.select(lora.spec.model, "full").DEFINITION_ID == fft.definition_id + assert routes.select(second.definition_id).DEFINITION_ID == second.definition_id + assert routes.select("missing") is None + assert routes.select(lora.definition_id, "full") is None def test_native_false_list_aliases_and_scalar_overrides(): @@ -346,8 +341,6 @@ def test_record_creation_copies_without_reparsing(): assert row.spec.revision == "a" * 40 # The record hash covers settings and the pinned backend dependency, not source. expected = original | {"revision": "a" * 40} - expected.pop("default") - expected.pop("sampling_default") assert ( row.generation == hashlib.sha256( @@ -459,11 +452,6 @@ def test_worker_hashes_cover_only_their_settings(): assert adapter.trainer_hash != base.trainer_hash assert adapter.inference_hash != base.inference_hash - routing = resolved(recipe(default=False)) - assert routing.trainer_hash == base.trainer_hash - assert routing.inference_hash == base.inference_hash - assert routing.generation == base.generation - upgraded = DeploymentRecord.create( base.spec, revision="a" * 40, inference_release="2" ) @@ -553,7 +541,11 @@ class Child(Parent): @pytest.mark.parametrize( "field,value", - [("max_contex_length", 8192), ("max_context_length", "8192"), ("default", "false")], + [ + ("max_contex_length", 8192), + ("max_context_length", "8192"), + ("pool_idle_timeout_s", "300"), + ], ) def test_recipe_class_fields_and_constructor_overrides_are_validated(field, value): from lilo.configs.qwen35_9b_lora_16k import Config as Parent @@ -592,7 +584,6 @@ class Grandchild(Child): overrides = { "trainer.config.max_tokens_per_gpu": 4096, "trainer.env.SECOND": "2", - "default": False, } config = Grandchild( @@ -604,7 +595,6 @@ class Grandchild(Child): assert config.trainer.env["FIRST"] == "1" assert config.trainer.env["SECOND"] == "2" assert config.trainer.config["max_tokens_per_gpu"] == 2048 - assert not config.default assert Grandchild().trainer.config["max_tokens_per_gpu"] == 4096 assert Child().trainer.config["max_tokens_per_gpu"] == 8192 config.inference.config["future_option"]["nested"].append(3) diff --git a/tests/test_system.py b/tests/test_system.py index f81a190..55aa9d9 100644 --- a/tests/test_system.py +++ b/tests/test_system.py @@ -19,7 +19,6 @@ DEFINITION_ID=DEFINITION, MODEL_NAME=BASE_MODEL, PARAMETERIZATION="lora", - CATALOG_VISIBLE=True, ), ) API_KEY = "tml-test" From 0b0da3e03bb93e68ada8fbaa0ba086d3c2164874 Mon Sep 17 00:00:00 2001 From: kailash Date: Thu, 24 Sep 2026 17:03:00 +0000 Subject: [PATCH 24/27] Flatten researcher configs into plain Python attributes --- docs/deployment-configs.md | 71 ++--- docs/deployment-validation.md | 12 +- scripts/definition_smoke.py | 2 +- scripts/e2e_engine_definition.py | 6 +- src/lilo/backends/deployment.py | 57 ++-- src/lilo/configs/qwen35_35b_a3b_fft_64k.py | 90 +++--- src/lilo/configs/qwen35_4b_fft_64k.py | 56 ++-- src/lilo/configs/qwen35_9b_fft_64k.py | 8 +- .../configs/qwen35_9b_instruct_lora_16k.py | 62 ++--- .../qwen35_9b_instruct_lora_16k_dp2.py | 2 +- src/lilo/configs/qwen35_9b_lora_16k.py | 74 +++-- src/lilo/configs/qwen35_9b_lora_16k_single.py | 2 +- src/lilo/configs/qwen35_9b_lora_2k.py | 18 +- src/lilo/configs/qwen35_9b_lora_64k.py | 8 +- src/lilo/configs/qwen36_27b_fft_64k.py | 72 +++-- src/lilo/configs/qwen36_35b_a3b_fft_64k.py | 5 +- src/lilo/configs/qwen38_27b_lora_128k.py | 16 +- src/lilo/configs/qwen38_27b_lora_16k.py | 62 ++--- src/lilo/configs/qwen38_27b_lora_256k.py | 26 +- src/lilo/configs/qwen38_27b_lora_64k.py | 8 +- src/lilo/configuration.py | 144 ++++------ src/lilo/deployment_cli.py | 4 +- src/lilo/deployments.py | 50 +++- src/lilo/providers/modal/deployment_apps.py | 64 +++-- .../test_megatron_constructor_settings.py | 19 +- tests/providers/test_deployment_apps.py | 41 +-- tests/providers/test_deployment_e2e_helper.py | 2 +- tests/providers/test_deployment_presets.py | 34 +-- tests/test_deployment_cli.py | 7 +- tests/test_deployments.py | 263 ++++++------------ 30 files changed, 532 insertions(+), 753 deletions(-) diff --git a/docs/deployment-configs.md b/docs/deployment-configs.md index 9cd4633..c353317 100644 --- a/docs/deployment-configs.md +++ b/docs/deployment-configs.md @@ -1,6 +1,6 @@ # Python deployment configs -A recipe subclasses `BaseConfig` and exports `config = Config()`. Model settings are ordinary class attributes; trainer and inference settings are dictionaries. No `Compute`, `Model`, or `Routing` constructors are needed. +A recipe subclasses `BaseConfig` and exports `config = Config()`. Settings are flat, untyped Python attributes. Backend options are ordinary dictionaries. ```python from lilo.configuration import BaseConfig @@ -10,21 +10,21 @@ class Config(BaseConfig): name = "my-9b" model = "Qwen/Qwen3.5-9B-Base" max_context_length = 16384 - - trainer = { - "gpu": "H100", - "gpus_per_node": 4, - "cpu": 16, - "memory_mib": 65536, - "max_clients_per_instance": 6, - "config": { - "model_type": "qwen3.5-9B", - "tensor_model_parallel_size": 4, - "max_lora_slots": 6, - "max_lora_rank": 32, - }, + backend = "miles" + trainer_gpu = "H100" + trainer_gpus_per_node = 4 + trainer_cpu = 16 + trainer_memory_mib = 65536 + trainer_max_clients_per_instance = 6 + inference_gpu = "H200" + inference_max_replicas = 8 + miles_cfg = { + "model_type": "qwen3.5-9B", + "tensor_model_parallel_size": 4, + "max_lora_slots": 6, + "max_lora_rank": 32, } - inference = {"gpu": "H200", "max_replicas": 8} + sglang_cfg = {"max_running_requests": 32} config = Config() @@ -34,7 +34,7 @@ See the [9B LoRA recipe](../src/lilo/configs/qwen35_9b_lora_16k.py) and [4B FFT ## Variants -Use Python inheritance and an `overrides` dictionary to change only the settings you need: +Override ordinary attributes directly. Use dotted `overrides` to change individual backend options: ```python from lilo.configs.qwen35_9b_lora_16k import Config as Parent @@ -42,45 +42,34 @@ from lilo.configs.qwen35_9b_lora_16k import Config as Parent class Config(Parent): name = "my-9b-more-memory" - overrides = { - "trainer.memory_mib": 98304, - "trainer.config.max_tokens_per_gpu": 8192, - "inference.gpu": "H200", - } + trainer_memory_mib = 98304 + overrides = {"miles_cfg.max_tokens_per_gpu": 8192} config = Config() ``` -Dotted paths set individual values. Parent settings and overrides apply first, then child settings and overrides; a child does not need to repeat its parent’s `overrides`. Assigning a dictionary or list replaces the value at that path: `"trainer.env": {}` clears inherited environment settings. Assigning a whole section as a class attribute still replaces that section. - -Constructor fields apply last, followed by constructor overrides: `Config(name="another-run", overrides={"trainer.gpu": "H200"})`. These are ordinary Python values, with no expressions or merge directives. Only the resolved settings are saved for workers. +Each parent's settings and overrides apply before its child's. Constructor fields and overrides apply last: `Config(trainer_gpu="H200", overrides={"sglang_cfg.max_running_requests": 16})`. Assigning a dictionary or list replaces that value; `trainer_env = {}` clears inherited environment settings. Instances own independent copies of mutable values and can also be edited directly. -Construction copies the recipe's dictionaries and validates Lilo-owned fields. Unknown fields, invalid types, negative capacities, and inconsistent scaling limits fail before deployment. The validated instance has attribute access (`config.trainer.gpu`); backend options stay dictionaries. Instances do not share mutable options with each other or with their recipe class. +## Backend options -## Ownership and validation - -| Setting | Owner and behavior | +| Setting | Consumed by | | --- | --- | -| trainer / inference resources | GPU type, GPUs per node, CPU and memory directly in each section; `nodes` is trainer-only. Unknown keys such as `memroy_mib` are rejected. | -| trainer | Maximum instances/clients, publication concurrency and function timeout. Trainers start on demand; there is no min_instances field. | -| inference | Replica scaling and startup_timeout_s, passed to the Modal server and startup health checks. There is no unused timeout_s. Each replica uses one node. | -| trainer.config | Existing MilesBackendConfig or EngineModelConfig fields, plus their explicit extra-option dictionaries. | -| inference.config | SGLang ServerArgs fields. Lilo reserves paths, context, topology and adapter settings that must agree with its own configuration. | - -Compute topology is configured directly in `trainer`; Miles receives actor_num_nodes and actor_num_gpus_per_node from it. Setting those again in backend options is rejected. +| `trainer_*`, `inference_*` | Modal GPU/CPU/memory allocation, scaling, timeouts and Lilo admission limits | +| `megatron_cfg` | Existing Megatron `EngineModelConfig`; provider, optimizer and distributed options use its native dictionaries | +| `miles_cfg` | Existing `MilesBackendConfig`; `cli_options` supplies additional Miles arguments | +| `sglang_cfg` | SGLang `ServerArgs` | -Megatron's provider_overrides, optimizer_overrides and distributed_overrides may add backend fields, but may not replace Lilo-owned fields. For example, put the learning rate in optimizer={"lr": ...}; optimizer_overrides={"lr": ...} is rejected. The same settings builders are used during validation and worker construction. The provider is constructed with dataclasses.replace, without an override-by-setattr pass. +Modal and backend libraries validate their own options. Lilo checks integration requirements such as trainer slot capacity, supported training modes, and parallelism agreeing with allocated GPUs. It supplies managed model paths, context length and adapter settings; conflicting backend overrides are rejected. -Miles has a cli_options dictionary for additional Miles arguments. Its argument conversion is isolated in [miles_arguments.py](../src/lilo/backends/miles_arguments.py), because Miles exposes an argparse interface. SGLang uses ServerArgs(**settings) directly in the worker-only [sglang.py](../src/lilo/inference/sglang.py) entrypoint. +`BaseConfig` does not enforce field types or reject arbitrary attributes. Extra backend options belong in the corresponding backend dictionary. A misspelled top-level attribute is ordinary Python data and may be unused. -Backend libraries validate their own extra options when workers start. Lilo does not maintain another schema for every upstream tuning option. Frontend config imports remain CPU-only. +Miles argument conversion lives in [miles_arguments.py](../src/lilo/backends/miles_arguments.py). SGLang receives `ServerArgs(**settings)` in its [worker entrypoint](../src/lilo/inference/sglang.py). Backend libraries validate native options when workers start; frontend config imports remain CPU-only. ## Resolve and launch ~~~text load(config.py) → config: BaseConfig - → validate typed compute/scaling/model settings → resolve model commit → resolve_backend_settings(config, asset_path) → save DeploymentRecord with trainer_settings and inference_settings @@ -92,7 +81,7 @@ The launcher consumes the saved settings. It does not reparse backend configurat | File | Responsibility | | --- | --- | -| [configuration.py](../src/lilo/configuration.py) | BaseConfig and validation of model, trainer, and inference settings | +| [configuration.py](../src/lilo/configuration.py) | BaseConfig defaults, inheritance and overrides | | [deployments.py](../src/lilo/deployments.py) | Python object loader, resolved records and config hashes | | [backends/deployment.py](../src/lilo/backends/deployment.py) | Resolve backend settings before launch | | [megatron_runtime/common/settings.py](../src/lilo/backends/megatron_runtime/common/settings.py) | Shared Megatron ownership rules and constructor dictionaries | @@ -115,7 +104,7 @@ lilo config validate my_model.py lilo deploy my_model.py ~~~ -Validation checks orchestration and integration constraints without loading GPU libraries or provisioning compute. Backend option support and GPU memory capacity still require worker startup. +Validation resolves backend settings and checks Lilo integration constraints without loading GPU libraries or provisioning compute. Backend option support and GPU memory capacity still require worker startup. The checked-in [deploy_models.sh](../scripts/deploy_models.sh) lists the complete active config set. Add a config path there, then run it. The deployment command owns frontend selection and worker-code updates: diff --git a/docs/deployment-validation.md b/docs/deployment-validation.md index 83fbfd2..d36de43 100644 --- a/docs/deployment-validation.md +++ b/docs/deployment-validation.md @@ -2,17 +2,15 @@ These checks exercise PR #55. The current implementation deploys trainers and inference provisioners independently; the shared frontend references them by app name. Historical sections record validation of previous implementations. -## Current: simple BaseConfig recipes +## Current: flat BaseConfig recipes -Recipes use one `BaseConfig` subclass with top-level model settings and plain trainer/inference dictionaries. Removed the `Compute`, `Model`, `Routing`, and `Lifecycle` wrapper classes and the matching nested fields. Compute settings live directly in trainer/inference; lifecycle timeouts live at the top level. The supported inference backend remains SGLang, so its redundant backend selector was removed. +Recipes use one plain `BaseConfig` with untyped attributes such as `trainer_gpu` and `inference_max_replicas`. Backend options live in `megatron_cfg`, `miles_cfg`, and `sglang_cfg` dictionaries. The `Trainer`, `Inference`, and `Deployment` config wrappers, constrained type aliases, and Pydantic recipe validation are removed. Each consuming component validates its own options; Lilo retains cross-component integration checks. -Variants use dotted `overrides` dictionaries to change inherited settings. Each parent’s settings and overrides apply before its child’s; dictionary and list values replace the value at their path. Constructor fields and overrides apply last. Construction copies mutable settings and validates against the existing schema. Workers continue to consume saved resolved backend settings. +Variants inherit class attributes and use dotted `overrides` for nested backend options. Mutable settings are copied per instance. Workers consume saved resolved backend settings. Routing uses deployment order or explicit definition IDs, without visibility or default flags. -Validation: **664 CPU tests passed, 1 skipped**. Coverage includes multilevel inherited overrides, constructor precedence, dictionary/list replacement, nested option isolation, invalid override paths, strict field validation, CLI revision resolution, saved-record round trips, and worker construction. The CLI-generated recipe also loaded and validated successfully. Ruff and whitespace checks passed. The full suite ran outside the socket-restricted sandbox so its local HTTP and SDK tests could run. +Validation: **644 CPU tests passed, 1 skipped**, including **111 focused config/deployment tests**. Coverage includes inherited overrides, mutable option isolation, backend passthrough, saved-record JSON round trips, worker resources, update isolation and deployment recovery. The CLI-generated recipe loaded and validated successfully. Ruff and whitespace checks passed. -Routing uses the first matching recipe in deployment order, or an explicit definition ID. Visibility and training/sampling default flags have been removed. Tests cover recipe ordering, LoRA/FFT selection, explicit IDs, listing all definitions, and retained generations. - -All 15 recipes preserve compute, model, lifecycle, resolved trainer/inference settings, and deployment hashes. Routing flags were removed from the recipe schema; retained definitions remain selectable and listed. No deployments or GPU jobs were modified. Existing draft manifests use the previous schema and require a fresh registry or explicit migration. +All 15 recipes preserve compute, model, lifecycle and resolved trainer/inference settings. The flat saved config layout changes deployment hashes. Existing draft manifests need regeneration or explicit migration. No deployments or GPU jobs were modified; GPU backend startup has not been rerun for this revision. ## Historical validation diff --git a/scripts/definition_smoke.py b/scripts/definition_smoke.py index 5cc2c9f..f342874 100644 --- a/scripts/definition_smoke.py +++ b/scripts/definition_smoke.py @@ -314,7 +314,7 @@ def main() -> None: for definition in definitions.values(): print( f"{definition.DEFINITION_ID:40} {definition.PARAMETERIZATION:5} " - f"{definition.RESOLVED.spec.trainer.nodes} nodes x {definition.RESOLVED.spec.trainer.gpu}:{definition.RESOLVED.spec.trainer.gpus_per_node} " + f"{definition.RESOLVED.spec.trainer_nodes} nodes x {definition.RESOLVED.spec.trainer_gpu}:{definition.RESOLVED.spec.trainer_gpus_per_node} " f"ctx={definition.MAX_CONTEXT_LENGTH}" ) return diff --git a/scripts/e2e_engine_definition.py b/scripts/e2e_engine_definition.py index bb5cbea..dae7f0a 100644 --- a/scripts/e2e_engine_definition.py +++ b/scripts/e2e_engine_definition.py @@ -35,7 +35,7 @@ def _definition(frontend: str, name: str) -> tuple[Any, str]: ) resolved = matches[0] spec = resolved.spec - settings = backend_config(spec)[spec.trainer.backend] + settings = backend_config(spec)[spec.backend] definition = SimpleNamespace( DEFINITION_ID=resolved.definition_id, MODEL_NAME=spec.model, @@ -45,8 +45,8 @@ def _definition(frontend: str, name: str) -> tuple[Any, str]: "max_tokens_per_microbatch", settings.get("max_tokens_per_gpu") ), MICRO_BATCH_SIZE=settings.get("micro_batch_size", 1), - GPU_TYPE=spec.trainer.gpu, - GPUS=spec.trainer.gpus_per_node, + GPU_TYPE=spec.trainer_gpu, + GPUS=spec.trainer_gpus_per_node, LORA_RANK=settings.get("max_lora_rank"), ) return definition, definition.PARAMETERIZATION diff --git a/src/lilo/backends/deployment.py b/src/lilo/backends/deployment.py index b1b7338..3af8479 100644 --- a/src/lilo/backends/deployment.py +++ b/src/lilo/backends/deployment.py @@ -74,9 +74,16 @@ def backend_config(spec, asset_path="/assets/pending"): - trainer = spec.trainer - settings = trainer.config - if trainer.backend == "megatron": + if spec.backend == "megatron": + if spec.trainer_nodes != 1: + raise ValueError("multi-node training currently requires Miles") + if spec.parameterization != "full": + raise ValueError("Megatron requires full parameterization") + if spec.trainer_max_clients_per_instance != 1: + raise ValueError("FFT trainers admit one client per instance") + if spec.sampler_persistence_concurrency != 1: + raise ValueError("Megatron requires sampler_persistence_concurrency: 1") + settings = spec.megatron_cfg reject_managed_options(settings, {"hf_checkpoint", "seq_length"}) config, _ = parse_backend_config( { @@ -89,8 +96,13 @@ def backend_config(spec, asset_path="/assets/pending"): ) if config.optimizer.optimizer != "adam": raise ValueError("Tinker optim_step requires an Adam optimizer") - config.validate(trainer.gpus_per_node) + config.validate(spec.trainer_gpus_per_node) return {"megatron": asdict(config), "checkpoint_dir": "/checkpoints"} + if spec.backend != "miles": + raise ValueError(f"unknown backend: {spec.backend}") + if spec.parameterization != "lora": + raise ValueError("Miles requires lora parameterization") + settings = spec.miles_cfg reject_managed_options( settings, {"hf_checkpoint", "actor_num_gpus_per_node", "actor_num_nodes", "extra_args"}, @@ -98,8 +110,8 @@ def backend_config(spec, asset_path="/assets/pending"): reject_managed_options(settings.get("cli_options", {}), MILES_MANAGED) config = MilesBackendConfig( hf_checkpoint=asset_path, - actor_num_gpus_per_node=trainer.gpus_per_node, - actor_num_nodes=trainer.nodes, + actor_num_gpus_per_node=spec.trainer_gpus_per_node, + actor_num_nodes=spec.trainer_nodes, extra_args=("--seq-length", str(spec.max_context_length)), **settings, ) @@ -108,40 +120,17 @@ def backend_config(spec, asset_path="/assets/pending"): config.expert_model_parallel_size * config.expert_tensor_parallel_size ): raise ValueError("expert parallel sizes must divide the trainer GPU allocation") - if trainer.max_clients_per_instance > config.max_lora_slots: + if spec.trainer_max_clients_per_instance > config.max_lora_slots: raise ValueError("max_clients_per_instance exceeds max_lora_slots") return {"miles": asdict(config), "checkpoint_dir": "/checkpoints"} def serving_options(spec): - options = dict(spec.inference.config) + options = dict(spec.sglang_cfg) reject_managed_options(options, SGLANG_MANAGED) - tp = options.get("tp_size", spec.inference.gpus_per_node) - ep = options.get("ep_size", 1) - if type(tp) is not int or tp != spec.inference.gpus_per_node: + tp = options.get("tp_size", spec.inference_gpus_per_node) + if tp != spec.inference_gpus_per_node: raise ValueError("sglang.tp_size must equal the replica GPU allocation") - if type(ep) is not int or ep < 1 or tp % ep: - raise ValueError("sglang.ep_size must divide the replica GPU allocation") - dp = options.get("dp_size", 1) - dp_attention = options.get("enable_dp_attention", False) - if type(dp) is not int or dp < 1 or tp % dp: - raise ValueError("sglang.dp_size must divide the replica GPU allocation") - if not isinstance(dp_attention, bool): - raise ValueError("sglang.enable_dp_attention must be a boolean") - if dp > 1 and not dp_attention: - raise ValueError("sglang.dp_size > 1 requires enable_dp_attention") - for key in ( - "max_loaded_loras", - "max_loras_per_batch", - "max_running_requests", - "max_queued_requests", - ): - if key in options and (type(options[key]) is not int or options[key] < 1): - raise ValueError(f"sglang.{key} must be positive") - if not 0 < options.get("mem_fraction_static", 0.8) < 1: - raise ValueError("sglang.mem_fraction_static must be between zero and one") - if options.get("max_loaded_loras", 64) < options.get("max_loras_per_batch", 8): - raise ValueError("max_loaded_loras must be >= max_loras_per_batch") return options @@ -149,7 +138,7 @@ def resolve_backend_settings(spec, asset_path): trainer = backend_config(spec, asset_path) inference = { "context_length": spec.max_context_length, - "tp_size": spec.inference.gpus_per_node, + "tp_size": spec.inference_gpus_per_node, "mem_fraction_static": 0.8, "max_running_requests": 32, "weight_loader_disable_mmap": True, diff --git a/src/lilo/configs/qwen35_35b_a3b_fft_64k.py b/src/lilo/configs/qwen35_35b_a3b_fft_64k.py index af2e356..d659372 100644 --- a/src/lilo/configs/qwen35_35b_a3b_fft_64k.py +++ b/src/lilo/configs/qwen35_35b_a3b_fft_64k.py @@ -6,55 +6,51 @@ class Config(BaseConfig): parameterization = "full" model = "Qwen/Qwen3.5-35B-A3B" max_context_length = 65536 - trainer = { - "gpu": "H200", - "gpus_per_node": 8, - "backend": "megatron", - "config": { - "tensor_model_parallel_size": 4, - "pipeline_model_parallel_size": 1, - "context_parallel_size": 2, - "expert_model_parallel_size": 8, - "expert_tensor_parallel_size": 1, - "sequence_parallel": True, - "micro_batch_size": 1, - "max_tokens_per_microbatch": 65536, - "bf16": True, - "fp16": False, - "gpu_memory_fraction": 0.9, - "use_distributed_optimizer": True, - "provider_overrides": { - "mtp_num_layers": 0, - "recompute_granularity": "selective", - "moe_layer_recompute": True, - "moe_token_dispatcher_type": "alltoall", - "moe_router_fusion": True, - "moe_permute_fusion": True, - "moe_grouped_gemm": True, - "moe_shared_expert_overlap": False, - "moe_aux_loss_coeff": 0.0, - }, - "optimizer": {"optimizer": "adam", "lr": 0.0001, "min_lr": 0.0001}, + trainer_gpu = "H200" + trainer_gpus_per_node = 8 + backend = "megatron" + megatron_cfg = { + "tensor_model_parallel_size": 4, + "pipeline_model_parallel_size": 1, + "context_parallel_size": 2, + "expert_model_parallel_size": 8, + "expert_tensor_parallel_size": 1, + "sequence_parallel": True, + "micro_batch_size": 1, + "max_tokens_per_microbatch": 65536, + "bf16": True, + "fp16": False, + "gpu_memory_fraction": 0.9, + "use_distributed_optimizer": True, + "provider_overrides": { + "mtp_num_layers": 0, + "recompute_granularity": "selective", + "moe_layer_recompute": True, + "moe_token_dispatcher_type": "alltoall", + "moe_router_fusion": True, + "moe_permute_fusion": True, + "moe_grouped_gemm": True, + "moe_shared_expert_overlap": False, + "moe_aux_loss_coeff": 0.0, }, - "env": { - "PYTORCH_CUDA_ALLOC_CONF": "expandable_segments:True", - "TORCHINDUCTOR_COMPILE_THREADS": "1", - }, - "sampler_persistence_concurrency": 1, + "optimizer": {"optimizer": "adam", "lr": 0.0001, "min_lr": 0.0001}, } - inference = { - "gpu": "H200", - "gpus_per_node": 4, - "config": { - "tp_size": 4, - "ep_size": 4, - "mem_fraction_static": 0.9, - "max_running_requests": 32, - "max_queued_requests": 4, - "cpu_weight_cache_max_compile_group_gb": 32, - "dp_size": 4, - "enable_dp_attention": True, - }, + trainer_env = { + "PYTORCH_CUDA_ALLOC_CONF": "expandable_segments:True", + "TORCHINDUCTOR_COMPILE_THREADS": "1", + } + sampler_persistence_concurrency = 1 + inference_gpu = "H200" + inference_gpus_per_node = 4 + sglang_cfg = { + "tp_size": 4, + "ep_size": 4, + "mem_fraction_static": 0.9, + "max_running_requests": 32, + "max_queued_requests": 4, + "cpu_weight_cache_max_compile_group_gb": 32, + "dp_size": 4, + "enable_dp_attention": True, } diff --git a/src/lilo/configs/qwen35_4b_fft_64k.py b/src/lilo/configs/qwen35_4b_fft_64k.py index d45b9df..aaab108 100644 --- a/src/lilo/configs/qwen35_4b_fft_64k.py +++ b/src/lilo/configs/qwen35_4b_fft_64k.py @@ -6,38 +6,34 @@ class Config(BaseConfig): parameterization = "full" model = "Qwen/Qwen3.5-4B" max_context_length = 65536 - trainer = { - "gpu": "H100", - "gpus_per_node": 4, - "backend": "megatron", - "config": { - "tensor_model_parallel_size": 2, - "context_parallel_size": 2, - "sequence_parallel": True, - "micro_batch_size": 1, - "max_tokens_per_microbatch": 65536, - "defer_fp32_logits": True, - "fp32_lm_head": True, - "use_distributed_optimizer": True, - "provider_overrides": { - "mtp_num_layers": 0, - "recompute_granularity": "full", - "recompute_method": "uniform", - "recompute_num_layers": 1, - }, - "optimizer": {"lr": 0.0001, "min_lr": 0.0001, "loss_scale": 1.0}, + trainer_gpu = "H100" + trainer_gpus_per_node = 4 + backend = "megatron" + megatron_cfg = { + "tensor_model_parallel_size": 2, + "context_parallel_size": 2, + "sequence_parallel": True, + "micro_batch_size": 1, + "max_tokens_per_microbatch": 65536, + "defer_fp32_logits": True, + "fp32_lm_head": True, + "use_distributed_optimizer": True, + "provider_overrides": { + "mtp_num_layers": 0, + "recompute_granularity": "full", + "recompute_method": "uniform", + "recompute_num_layers": 1, }, - "sampler_persistence_concurrency": 1, + "optimizer": {"lr": 0.0001, "min_lr": 0.0001, "loss_scale": 1.0}, } - inference = { - "gpu": "H100", - "config": { - "tp_size": 1, - "mem_fraction_static": 0.85, - "max_running_requests": 32, - "max_queued_requests": 4, - "cpu_weight_cache_max_compile_group_gb": 16, - }, + sampler_persistence_concurrency = 1 + inference_gpu = "H100" + sglang_cfg = { + "tp_size": 1, + "mem_fraction_static": 0.85, + "max_running_requests": 32, + "max_queued_requests": 4, + "cpu_weight_cache_max_compile_group_gb": 16, } diff --git a/src/lilo/configs/qwen35_9b_fft_64k.py b/src/lilo/configs/qwen35_9b_fft_64k.py index a23116c..b8682ed 100644 --- a/src/lilo/configs/qwen35_9b_fft_64k.py +++ b/src/lilo/configs/qwen35_9b_fft_64k.py @@ -5,13 +5,13 @@ class Config(Parent): name = "qwen35-9b-fft-64k" model = "Qwen/Qwen3.5-9B" overrides = { - "trainer.gpu": "H200", - "trainer.env": { + "trainer_gpu": "H200", + "trainer_env": { "PYTORCH_CUDA_ALLOC_CONF": "expandable_segments:True", "TORCHINDUCTOR_COMPILE_THREADS": "1", }, - "inference.gpu": "H200", - "inference.config.ep_size": 1, + "inference_gpu": "H200", + "sglang_cfg.ep_size": 1, } diff --git a/src/lilo/configs/qwen35_9b_instruct_lora_16k.py b/src/lilo/configs/qwen35_9b_instruct_lora_16k.py index 1aedccf..a722cfd 100644 --- a/src/lilo/configs/qwen35_9b_instruct_lora_16k.py +++ b/src/lilo/configs/qwen35_9b_instruct_lora_16k.py @@ -5,41 +5,37 @@ class Config(BaseConfig): name = "qwen35-9b-instruct-lora-16k" model = "Qwen/Qwen3.5-9B" max_context_length = 16384 - trainer = { - "gpu": "H100", - "gpus_per_node": 8, - "config": { - "model_type": "qwen3.5-9B", - "tensor_model_parallel_size": 8, - "target_modules": ["linear_qkv", "linear_proj", "linear_fc1", "linear_fc2"], - "max_tokens_per_gpu": 16384, - "max_lora_slots": 6, - "max_lora_rank": 32, - "default_lora_alpha": 32, - "cli_options": { - "recompute_granularity": "full", - "recompute_method": "uniform", - "recompute_num_layers": 1, - }, + trainer_gpu = "H100" + trainer_gpus_per_node = 8 + miles_cfg = { + "model_type": "qwen3.5-9B", + "tensor_model_parallel_size": 8, + "target_modules": ["linear_qkv", "linear_proj", "linear_fc1", "linear_fc2"], + "max_tokens_per_gpu": 16384, + "max_lora_slots": 6, + "max_lora_rank": 32, + "default_lora_alpha": 32, + "cli_options": { + "recompute_granularity": "full", + "recompute_method": "uniform", + "recompute_num_layers": 1, }, - "env": { - "PYTORCH_CUDA_ALLOC_CONF": "expandable_segments:True", - "TORCHINDUCTOR_COMPILE_THREADS": "1", - }, - "max_clients_per_instance": 6, } - inference = { - "gpu": "H200", - "config": { - "tp_size": 1, - "ep_size": 1, - "mem_fraction_static": 0.8, - "max_running_requests": 32, - "max_queued_requests": 8, - "max_loaded_loras": 256, - "max_loras_per_batch": 8, - "schedule_policy": "lpm", - }, + trainer_env = { + "PYTORCH_CUDA_ALLOC_CONF": "expandable_segments:True", + "TORCHINDUCTOR_COMPILE_THREADS": "1", + } + trainer_max_clients_per_instance = 6 + inference_gpu = "H200" + sglang_cfg = { + "tp_size": 1, + "ep_size": 1, + "mem_fraction_static": 0.8, + "max_running_requests": 32, + "max_queued_requests": 8, + "max_loaded_loras": 256, + "max_loras_per_batch": 8, + "schedule_policy": "lpm", } diff --git a/src/lilo/configs/qwen35_9b_instruct_lora_16k_dp2.py b/src/lilo/configs/qwen35_9b_instruct_lora_16k_dp2.py index caaeabe..5f7fbfe 100644 --- a/src/lilo/configs/qwen35_9b_instruct_lora_16k_dp2.py +++ b/src/lilo/configs/qwen35_9b_instruct_lora_16k_dp2.py @@ -3,7 +3,7 @@ class Config(Parent): name = "qwen35-9b-instruct-lora-16k-dp2" - overrides = {"trainer.config.tensor_model_parallel_size": 4} + overrides = {"miles_cfg.tensor_model_parallel_size": 4} config = Config() diff --git a/src/lilo/configs/qwen35_9b_lora_16k.py b/src/lilo/configs/qwen35_9b_lora_16k.py index 586c3de..4851bb5 100644 --- a/src/lilo/configs/qwen35_9b_lora_16k.py +++ b/src/lilo/configs/qwen35_9b_lora_16k.py @@ -5,47 +5,43 @@ class Config(BaseConfig): name = "qwen35-9b-lora-16k" model = "Qwen/Qwen3.5-9B-Base" max_context_length = 16384 - trainer = { - "gpu": "H100", - "gpus_per_node": 4, - "cpu": 16, - "memory_mib": 65536, - "config": { - "model_type": "qwen3.5-9B", - "tensor_model_parallel_size": 4, - "max_lora_slots": 6, - "max_lora_rank": 32, - "default_lora_alpha": 32, - "target_modules": [ - "linear_qkv", - "linear_proj", - "linear_fc1", - "linear_fc2", - "output_layer", - ], - "max_tokens_per_gpu": 16384, - "cli_options": { - "recompute_granularity": "full", - "recompute_method": "uniform", - "recompute_num_layers": 1, - }, + trainer_gpu = "H100" + trainer_gpus_per_node = 4 + trainer_cpu = 16 + trainer_memory_mib = 65536 + miles_cfg = { + "model_type": "qwen3.5-9B", + "tensor_model_parallel_size": 4, + "max_lora_slots": 6, + "max_lora_rank": 32, + "default_lora_alpha": 32, + "target_modules": [ + "linear_qkv", + "linear_proj", + "linear_fc1", + "linear_fc2", + "output_layer", + ], + "max_tokens_per_gpu": 16384, + "cli_options": { + "recompute_granularity": "full", + "recompute_method": "uniform", + "recompute_num_layers": 1, }, - "env": { - "PYTORCH_CUDA_ALLOC_CONF": "expandable_segments:True", - "TORCHINDUCTOR_COMPILE_THREADS": "1", - }, - "max_clients_per_instance": 6, } - inference = { - "gpu": "H200", - "config": { - "tp_size": 1, - "mem_fraction_static": 0.8, - "max_running_requests": 32, - "max_queued_requests": 8, - "max_loaded_loras": 64, - "max_loras_per_batch": 8, - }, + trainer_env = { + "PYTORCH_CUDA_ALLOC_CONF": "expandable_segments:True", + "TORCHINDUCTOR_COMPILE_THREADS": "1", + } + trainer_max_clients_per_instance = 6 + inference_gpu = "H200" + sglang_cfg = { + "tp_size": 1, + "mem_fraction_static": 0.8, + "max_running_requests": 32, + "max_queued_requests": 8, + "max_loaded_loras": 64, + "max_loras_per_batch": 8, } diff --git a/src/lilo/configs/qwen35_9b_lora_16k_single.py b/src/lilo/configs/qwen35_9b_lora_16k_single.py index 1cf0afb..7a596e9 100644 --- a/src/lilo/configs/qwen35_9b_lora_16k_single.py +++ b/src/lilo/configs/qwen35_9b_lora_16k_single.py @@ -3,7 +3,7 @@ class Config(Parent): name = "qwen35-9b-lora-16k-single" - overrides = {"trainer.max_clients_per_instance": 1} + overrides = {"trainer_max_clients_per_instance": 1} config = Config() diff --git a/src/lilo/configs/qwen35_9b_lora_2k.py b/src/lilo/configs/qwen35_9b_lora_2k.py index a3b149d..d900ca4 100644 --- a/src/lilo/configs/qwen35_9b_lora_2k.py +++ b/src/lilo/configs/qwen35_9b_lora_2k.py @@ -5,15 +5,15 @@ class Config(Parent): name = "qwen35-9b-lora-2k" max_context_length = 2048 overrides = { - "trainer.gpu": "H200", - "trainer.cpu": 8, - "trainer.memory_mib": 32768, - "trainer.max_clients_per_instance": 4, - "trainer.config.max_tokens_per_gpu": 2048, - "trainer.config.max_lora_slots": 4, - "inference.config.ep_size": 1, - "inference.config.max_loaded_loras": 32, - "inference.config.schedule_policy": "lpm", + "trainer_gpu": "H200", + "trainer_cpu": 8, + "trainer_memory_mib": 32768, + "trainer_max_clients_per_instance": 4, + "miles_cfg.max_tokens_per_gpu": 2048, + "miles_cfg.max_lora_slots": 4, + "sglang_cfg.ep_size": 1, + "sglang_cfg.max_loaded_loras": 32, + "sglang_cfg.schedule_policy": "lpm", } diff --git a/src/lilo/configs/qwen35_9b_lora_64k.py b/src/lilo/configs/qwen35_9b_lora_64k.py index 3132989..a0354d3 100644 --- a/src/lilo/configs/qwen35_9b_lora_64k.py +++ b/src/lilo/configs/qwen35_9b_lora_64k.py @@ -5,10 +5,10 @@ class Config(Parent): name = "qwen35-9b-lora-64k" max_context_length = 65536 overrides = { - "trainer.gpu": "H200", - "trainer.gpus_per_node": 8, - "trainer.config.tensor_model_parallel_size": 8, - "trainer.config.max_tokens_per_gpu": 65536, + "trainer_gpu": "H200", + "trainer_gpus_per_node": 8, + "miles_cfg.tensor_model_parallel_size": 8, + "miles_cfg.max_tokens_per_gpu": 65536, } diff --git a/src/lilo/configs/qwen36_27b_fft_64k.py b/src/lilo/configs/qwen36_27b_fft_64k.py index e185d89..0af1b67 100644 --- a/src/lilo/configs/qwen36_27b_fft_64k.py +++ b/src/lilo/configs/qwen36_27b_fft_64k.py @@ -6,46 +6,42 @@ class Config(BaseConfig): parameterization = "full" model = "Qwen/Qwen3.6-27B" max_context_length = 65536 - trainer = { - "gpu": "H200", - "gpus_per_node": 8, - "backend": "megatron", - "config": { - "tensor_model_parallel_size": 4, - "pipeline_model_parallel_size": 1, - "context_parallel_size": 2, - "sequence_parallel": True, - "micro_batch_size": 1, - "max_tokens_per_microbatch": 65536, - "bf16": True, - "fp16": False, - "gpu_memory_fraction": 0.9, - "use_distributed_optimizer": True, - "provider_overrides": { - "mtp_num_layers": 0, - "recompute_granularity": "full", - "recompute_method": "uniform", - "recompute_num_layers": 1, - }, - "optimizer": {"optimizer": "adam", "lr": 0.0001, "min_lr": 0.0001}, + trainer_gpu = "H200" + trainer_gpus_per_node = 8 + backend = "megatron" + megatron_cfg = { + "tensor_model_parallel_size": 4, + "pipeline_model_parallel_size": 1, + "context_parallel_size": 2, + "sequence_parallel": True, + "micro_batch_size": 1, + "max_tokens_per_microbatch": 65536, + "bf16": True, + "fp16": False, + "gpu_memory_fraction": 0.9, + "use_distributed_optimizer": True, + "provider_overrides": { + "mtp_num_layers": 0, + "recompute_granularity": "full", + "recompute_method": "uniform", + "recompute_num_layers": 1, }, - "env": { - "PYTORCH_CUDA_ALLOC_CONF": "expandable_segments:True", - "TORCHINDUCTOR_COMPILE_THREADS": "1", - }, - "sampler_persistence_concurrency": 1, + "optimizer": {"optimizer": "adam", "lr": 0.0001, "min_lr": 0.0001}, } - inference = { - "gpu": "H200", - "gpus_per_node": 4, - "config": { - "tp_size": 4, - "ep_size": 1, - "mem_fraction_static": 0.9, - "max_running_requests": 32, - "max_queued_requests": 4, - "cpu_weight_cache_max_compile_group_gb": 32, - }, + trainer_env = { + "PYTORCH_CUDA_ALLOC_CONF": "expandable_segments:True", + "TORCHINDUCTOR_COMPILE_THREADS": "1", + } + sampler_persistence_concurrency = 1 + inference_gpu = "H200" + inference_gpus_per_node = 4 + sglang_cfg = { + "tp_size": 4, + "ep_size": 1, + "mem_fraction_static": 0.9, + "max_running_requests": 32, + "max_queued_requests": 4, + "cpu_weight_cache_max_compile_group_gb": 32, } diff --git a/src/lilo/configs/qwen36_35b_a3b_fft_64k.py b/src/lilo/configs/qwen36_35b_a3b_fft_64k.py index cd39cc1..25a854a 100644 --- a/src/lilo/configs/qwen36_35b_a3b_fft_64k.py +++ b/src/lilo/configs/qwen36_35b_a3b_fft_64k.py @@ -4,10 +4,7 @@ class Config(Parent): name = "qwen36-35b-a3b-fft-64k" model = "Qwen/Qwen3.6-35B-A3B" - overrides = { - "inference.config.dp_size": 1, - "inference.config.enable_dp_attention": False, - } + overrides = {"sglang_cfg.dp_size": 1, "sglang_cfg.enable_dp_attention": False} config = Config() diff --git a/src/lilo/configs/qwen38_27b_lora_128k.py b/src/lilo/configs/qwen38_27b_lora_128k.py index df6ff85..b18b718 100644 --- a/src/lilo/configs/qwen38_27b_lora_128k.py +++ b/src/lilo/configs/qwen38_27b_lora_128k.py @@ -5,14 +5,14 @@ class Config(Parent): name = "qwen38-27b-lora-128k" max_context_length = 131072 overrides = { - "trainer.config.tensor_model_parallel_size": 2, - "trainer.config.max_tokens_per_gpu": 32768, - "trainer.config.context_parallel_size": 4, - "inference.gpus_per_node": 2, - "inference.max_replicas": 4, - "inference.target_concurrency": 4, - "inference.config.tp_size": 2, - "inference.config.max_running_requests": 8, + "miles_cfg.tensor_model_parallel_size": 2, + "miles_cfg.max_tokens_per_gpu": 32768, + "miles_cfg.context_parallel_size": 4, + "inference_gpus_per_node": 2, + "inference_max_replicas": 4, + "inference_target_concurrency": 4, + "sglang_cfg.tp_size": 2, + "sglang_cfg.max_running_requests": 8, } diff --git a/src/lilo/configs/qwen38_27b_lora_16k.py b/src/lilo/configs/qwen38_27b_lora_16k.py index db54bce..b77bb6b 100644 --- a/src/lilo/configs/qwen38_27b_lora_16k.py +++ b/src/lilo/configs/qwen38_27b_lora_16k.py @@ -5,41 +5,37 @@ class Config(BaseConfig): name = "qwen38-27b-lora-16k" model = "Qwen/Qwen3.8-27B" max_context_length = 16384 - trainer = { - "gpu": "H200", - "gpus_per_node": 8, - "config": { - "model_type": "qwen3.8-27B", - "tensor_model_parallel_size": 4, - "target_modules": ["linear_qkv", "linear_proj", "linear_fc1", "linear_fc2"], - "max_tokens_per_gpu": 16384, - "max_lora_slots": 6, - "max_lora_rank": 32, - "default_lora_alpha": 32, - "cli_options": { - "recompute_granularity": "full", - "recompute_method": "uniform", - "recompute_num_layers": 1, - }, + trainer_gpu = "H200" + trainer_gpus_per_node = 8 + miles_cfg = { + "model_type": "qwen3.8-27B", + "tensor_model_parallel_size": 4, + "target_modules": ["linear_qkv", "linear_proj", "linear_fc1", "linear_fc2"], + "max_tokens_per_gpu": 16384, + "max_lora_slots": 6, + "max_lora_rank": 32, + "default_lora_alpha": 32, + "cli_options": { + "recompute_granularity": "full", + "recompute_method": "uniform", + "recompute_num_layers": 1, }, - "env": { - "PYTORCH_CUDA_ALLOC_CONF": "expandable_segments:True", - "TORCHINDUCTOR_COMPILE_THREADS": "1", - }, - "max_clients_per_instance": 6, } - inference = { - "gpu": "H200", - "config": { - "tp_size": 1, - "ep_size": 1, - "mem_fraction_static": 0.8, - "max_running_requests": 32, - "max_queued_requests": 8, - "max_loaded_loras": 256, - "max_loras_per_batch": 8, - "schedule_policy": "lpm", - }, + trainer_env = { + "PYTORCH_CUDA_ALLOC_CONF": "expandable_segments:True", + "TORCHINDUCTOR_COMPILE_THREADS": "1", + } + trainer_max_clients_per_instance = 6 + inference_gpu = "H200" + sglang_cfg = { + "tp_size": 1, + "ep_size": 1, + "mem_fraction_static": 0.8, + "max_running_requests": 32, + "max_queued_requests": 8, + "max_loaded_loras": 256, + "max_loras_per_batch": 8, + "schedule_policy": "lpm", } diff --git a/src/lilo/configs/qwen38_27b_lora_256k.py b/src/lilo/configs/qwen38_27b_lora_256k.py index 51d7f0f..0ab78e9 100644 --- a/src/lilo/configs/qwen38_27b_lora_256k.py +++ b/src/lilo/configs/qwen38_27b_lora_256k.py @@ -5,19 +5,19 @@ class Config(Parent): name = "qwen38-27b-lora-256k" max_context_length = 262144 overrides = { - "trainer.nodes": 2, - "trainer.config.tensor_model_parallel_size": 2, - "trainer.config.context_parallel_size": 8, - "trainer.config.max_tokens_per_gpu": 32768, - "trainer.config.cli_options.distributed_timeout_minutes": 120, - "inference.gpus_per_node": 4, - "inference.min_replicas": 2, - "inference.max_replicas": 2, - "inference.target_concurrency": 2, - "inference.config.tp_size": 4, - "inference.config.max_running_requests": 4, - "inference.config.max_queued_requests": 8, - "inference.config.max_loaded_loras": 256, + "trainer_nodes": 2, + "miles_cfg.tensor_model_parallel_size": 2, + "miles_cfg.context_parallel_size": 8, + "miles_cfg.max_tokens_per_gpu": 32768, + "miles_cfg.cli_options.distributed_timeout_minutes": 120, + "inference_gpus_per_node": 4, + "inference_min_replicas": 2, + "inference_max_replicas": 2, + "inference_target_concurrency": 2, + "sglang_cfg.tp_size": 4, + "sglang_cfg.max_running_requests": 4, + "sglang_cfg.max_queued_requests": 8, + "sglang_cfg.max_loaded_loras": 256, } diff --git a/src/lilo/configs/qwen38_27b_lora_64k.py b/src/lilo/configs/qwen38_27b_lora_64k.py index f6f2216..d864a3e 100644 --- a/src/lilo/configs/qwen38_27b_lora_64k.py +++ b/src/lilo/configs/qwen38_27b_lora_64k.py @@ -5,10 +5,10 @@ class Config(Parent): name = "qwen38-27b-lora-64k" max_context_length = 65536 overrides = { - "trainer.config.max_tokens_per_gpu": 32768, - "trainer.config.context_parallel_size": 2, - "inference.target_concurrency": 8, - "inference.config.max_running_requests": 16, + "miles_cfg.max_tokens_per_gpu": 32768, + "miles_cfg.context_parallel_size": 2, + "inference_target_concurrency": 8, + "sglang_cfg.max_running_requests": 16, } diff --git a/src/lilo/configuration.py b/src/lilo/configuration.py index 34f31f9..4a2b085 100644 --- a/src/lilo/configuration.py +++ b/src/lilo/configuration.py @@ -1,116 +1,66 @@ -"""Readable Python recipes with validated model, trainer, and inference settings.""" +"""Flat Python recipes. Modal and the backend validate their own options.""" from copy import deepcopy -from dataclasses import field -from typing import Annotated, Literal - -from pydantic import ConfigDict, Field -from pydantic.dataclasses import dataclass - -PositiveInt = Annotated[int, Field(strict=True, gt=0)] -NonnegativeInt = Annotated[int, Field(strict=True, ge=0)] -NonemptyString = Annotated[str, Field(strict=True, min_length=1)] -GPU = Annotated[str, Field(strict=True, pattern=r"^[A-Za-z0-9-]+$")] -CONFIG = ConfigDict(extra="forbid", validate_default=True) - - -@dataclass(config=CONFIG, frozen=True, kw_only=True) -class Trainer: - gpu: GPU - gpus_per_node: PositiveInt = 1 - nodes: PositiveInt = 1 - cpu: Annotated[float, Field(gt=0)] = 8 - memory_mib: PositiveInt = 32768 - backend: Literal["miles", "megatron"] = "miles" - max_instances: PositiveInt = 1 - max_clients_per_instance: PositiveInt = 1 - sampler_persistence_concurrency: PositiveInt = 8 - timeout_s: PositiveInt = 86400 - config: dict[str, object] = field(default_factory=dict) - env: dict[str, str] = field(default_factory=dict) - - -@dataclass(config=CONFIG, frozen=True, kw_only=True) -class Inference: - gpu: GPU - gpus_per_node: PositiveInt = 1 - cpu: Annotated[float, Field(gt=0)] = 8 - memory_mib: PositiveInt = 32768 - min_replicas: NonnegativeInt = 0 - max_replicas: PositiveInt = 8 - target_concurrency: PositiveInt = 16 - scaledown_window_s: PositiveInt = 300 - startup_timeout_s: PositiveInt = 1200 - config: dict[str, object] = field(default_factory=dict) - env: dict[str, str] = field(default_factory=dict) - - -@dataclass(config=CONFIG, frozen=True, kw_only=True) -class Deployment: - name: Annotated[str, Field(strict=True, pattern=r"^[A-Za-z0-9_-]+$")] - model: NonemptyString - max_context_length: PositiveInt - trainer: Trainer - inference: Inference - parameterization: Literal["lora", "full"] = "lora" - revision: NonemptyString = "main" - session_idle_timeout_s: PositiveInt = 300 - pool_idle_timeout_s: PositiveInt = 300 - sweep_interval_s: PositiveInt = 300 - - def __post_init__(self): - if self.inference.min_replicas > self.inference.max_replicas: - raise ValueError("min_replicas must not exceed max_replicas") - if self.trainer.backend == "megatron": - if self.trainer.nodes != 1: - raise ValueError("multi-node training currently requires Miles") - if self.parameterization != "full": - raise ValueError("Megatron requires full parameterization") - if self.trainer.max_clients_per_instance != 1: - raise ValueError("FFT trainers admit one client per instance") - if self.trainer.sampler_persistence_concurrency != 1: - raise ValueError("Megatron requires sampler_persistence_concurrency: 1") - elif self.parameterization != "lora": - raise ValueError("Miles requires lora parameterization") - for env in (self.trainer.env, self.inference.env): - if any(key.startswith("LILO_") for key in env): - raise ValueError("LILO_ environment variables are managed by Lilo") def _apply_overrides(values, overrides): - """Set dotted dictionary paths; the deployment schema validates the result.""" - if not isinstance(overrides, dict): - raise ValueError("overrides must be a dictionary of dotted paths") for path, value in overrides.items(): - if not isinstance(path, str) or any(not part for part in path.split(".")): - raise ValueError(f"invalid override path: {path!r}") + parts = path.split(".") target = values - for part in path.split(".")[:-1]: + for part in parts[:-1]: target = target.setdefault(part, {}) - if not isinstance(target, dict): - raise ValueError(f"override {path!r} traverses a non-dictionary value") - target[path.rsplit(".", 1)[-1]] = deepcopy(value) - - -class BaseConfig(Deployment): - """Declare a recipe, then change inherited settings with dotted overrides. - - Each class's attributes and overrides apply in parent-to-child order. - Constructor fields and overrides apply last. Dictionary/list values replace - the value at that path, and each instance owns its mutable settings. - """ + target[parts[-1]] = deepcopy(value) + + +class BaseConfig: + name = "" + model = "" + max_context_length = 16384 + parameterization = "lora" + revision = "main" + backend = "miles" + + trainer_gpu = "H100" + trainer_gpus_per_node = 1 + trainer_nodes = 1 + trainer_cpu = 8 + trainer_memory_mib = 32768 + trainer_max_instances = 1 + trainer_max_clients_per_instance = 1 + trainer_timeout_s = 86400 + trainer_env = {} + sampler_persistence_concurrency = 8 + + inference_gpu = "H100" + inference_gpus_per_node = 1 + inference_cpu = 8 + inference_memory_mib = 32768 + inference_min_replicas = 0 + inference_max_replicas = 8 + inference_target_concurrency = 16 + inference_scaledown_window_s = 300 + inference_startup_timeout_s = 1200 + inference_env = {} + + miles_cfg = {} + megatron_cfg = {} + sglang_cfg = {} + session_idle_timeout_s = 300 + pool_idle_timeout_s = 300 + sweep_interval_s = 300 def __init__(self, **kwargs): + # Copy each ancestor's values before applying its overrides. values = {} for cls in reversed(type(self).__mro__): - if cls in (object, Deployment, BaseConfig): - continue values.update( deepcopy( { key: value for key, value in vars(cls).items() - if not key.startswith("_") and key != "overrides" + if not key.startswith("_") + and key != "overrides" + and not callable(value) } ) ) @@ -118,4 +68,4 @@ def __init__(self, **kwargs): overrides = kwargs.pop("overrides", {}) values.update(deepcopy(kwargs)) _apply_overrides(values, overrides) - super().__init__(**values) + self.__dict__.update(values) diff --git a/src/lilo/deployment_cli.py b/src/lilo/deployment_cli.py index 58b9ad1..d2097f3 100644 --- a/src/lilo/deployment_cli.py +++ b/src/lilo/deployment_cli.py @@ -34,7 +34,7 @@ def compile_configs(paths, *, platform=None): miles_commit = ( resolve_miles_commit() - if any(spec.trainer.backend == "miles" for spec in specs) + if any(spec.backend == "miles" for spec in specs) else None ) records = [] @@ -51,7 +51,7 @@ def compile_configs(paths, *, platform=None): spec, platform=platform, revision=revision, - miles_commit=miles_commit if spec.trainer.backend == "miles" else None, + miles_commit=miles_commit if spec.backend == "miles" else None, ) ) return records diff --git a/src/lilo/deployments.py b/src/lilo/deployments.py index 562f3f0..6de312f 100644 --- a/src/lilo/deployments.py +++ b/src/lilo/deployments.py @@ -8,15 +8,14 @@ import runpy import sys from copy import deepcopy -from dataclasses import asdict, replace from importlib.resources import files from pathlib import Path from typing import Any -from pydantic import BaseModel, ConfigDict, Field +from pydantic import BaseModel, ConfigDict, Field, field_serializer, field_validator from lilo.backends.deployment import resolve_backend_settings -from lilo.configuration import Deployment +from lilo.configuration import BaseConfig LIFECYCLE_FIELDS = ("session_idle_timeout_s", "pool_idle_timeout_s", "sweep_interval_s") @@ -50,7 +49,7 @@ def deployment_generation( trainer_settings, inference_settings, ): - identity = asdict(spec) + identity = vars(spec) return settings_hash( { "config": identity, @@ -76,9 +75,9 @@ class DeploymentRecord(BaseModel): active marks recipes in the latest deploy command; retained recipes remain usable. """ - model_config = ConfigDict(extra="forbid") + model_config = ConfigDict(extra="forbid", arbitrary_types_allowed=True) - spec: Deployment + spec: BaseConfig platform: dict[str, Any] = Field(default_factory=platform_defaults) trainer_release: str = "initial" inference_release: str = "initial" @@ -88,10 +87,19 @@ class DeploymentRecord(BaseModel): generation: str active: bool = True + @field_validator("spec", mode="before") + @classmethod + def restore_config(cls, value): + return BaseConfig(**value) if isinstance(value, dict) else value + + @field_serializer("spec") + def serialize_config(self, config): + return vars(config) + @classmethod def create( cls, - spec: Deployment, + spec: BaseConfig, *, revision: str, miles_commit: str | None = None, @@ -100,7 +108,8 @@ def create( inference_release: str = "initial", ) -> DeploymentRecord: """Record an already-resolved revision without reparsing the configuration.""" - pinned = replace(deepcopy(spec), revision=revision) + pinned = deepcopy(spec) + pinned.revision = revision asset_path = model_asset_path(pinned.model, revision) trainer_settings, inference_settings = resolve_backend_settings( pinned, asset_path @@ -154,7 +163,18 @@ def trainer_hash(self) -> str: "revision": self.spec.revision, "parameterization": self.spec.parameterization, "max_context_length": self.spec.max_context_length, - "trainer": asdict(self.spec.trainer), + "trainer": { + key: value + for key, value in vars(self.spec).items() + if key.startswith("trainer_") + or key + in ( + "backend", + "miles_cfg", + "megatron_cfg", + "sampler_persistence_concurrency", + ) + }, "settings": self.trainer_settings, "release": self.trainer_release, "platform": self.platform, @@ -172,7 +192,11 @@ def inference_hash(self) -> str: "revision": self.spec.revision, "parameterization": self.spec.parameterization, "max_context_length": self.spec.max_context_length, - "inference": asdict(self.spec.inference), + "inference": { + key: value + for key, value in vars(self.spec).items() + if key.startswith("inference_") or key == "sglang_cfg" + }, "settings": self.inference_settings, "release": self.inference_release, "platform": self.platform, @@ -208,7 +232,7 @@ def config_path(name: str) -> Path: return Path(str(files("lilo").joinpath("configs", name.replace("-", "_") + ".py"))) -def load(path: str | Path) -> Deployment: +def load(path: str | Path) -> BaseConfig: """Execute a Python config file and read its exported config object.""" path = Path(path).resolve() if path.suffix != ".py": @@ -219,14 +243,14 @@ def load(path: str | Path) -> Deployment: try: namespace = runpy.run_path(str(path)) config = namespace.get("config") - if not isinstance(config, Deployment): + if not isinstance(config, BaseConfig): raise ValueError(f"{path} must export a BaseConfig instance named config") return config finally: sys.path[:] = original_path -def validate_frontend(specs: list[Deployment]) -> None: +def validate_frontend(specs: list[BaseConfig]) -> None: if not specs: raise ValueError("at least one deployment is required") if len({s.name for s in specs}) != len(specs): diff --git a/src/lilo/providers/modal/deployment_apps.py b/src/lilo/providers/modal/deployment_apps.py index c88b876..c0eadef 100644 --- a/src/lilo/providers/modal/deployment_apps.py +++ b/src/lilo/providers/modal/deployment_apps.py @@ -104,11 +104,10 @@ def build_trainer_app(resolved: DeploymentRecord, *, image=None): spec = resolved.spec trainer_hash = resolved.trainer_hash app = modal.App(resolved.trainer_app_name) - resource = spec.trainer env = { **trainer_deployment_env(), - **deployment_env(spec.trainer.env), + **deployment_env(spec.trainer_env), "LILO_APP_NAME": resolved.platform["frontend"], } @@ -118,17 +117,17 @@ def trainer(instance_id: str, config_json: str): raise ValueError("trainer settings do not match the deployed app") run_trainer(record, instance_id) - if resource.nodes > 1: - trainer = modal.experimental.clustered(resource.nodes, rdma=True)(trainer) + if spec.trainer_nodes > 1: + trainer = modal.experimental.clustered(spec.trainer_nodes, rdma=True)(trainer) trainer = app.function( name="trainer", serialized=True, - image=image if image is not None else image_for(spec.trainer.backend), - gpu=f"{resource.gpu}:{resource.gpus_per_node}", + image=image if image is not None else image_for(spec.backend), + gpu=f"{spec.trainer_gpu}:{spec.trainer_gpus_per_node}", region=resolved.platform["modal"]["region"], - cpu=resource.cpu, - memory=resource.memory_mib, - timeout=spec.trainer.timeout_s, + cpu=spec.trainer_cpu, + memory=spec.trainer_memory_mib, + timeout=spec.trainer_timeout_s, # Admission/reconciliation caps each definition. A function-wide cap # would block new definitions behind retained jobs sharing this app. max_containers=None, @@ -137,7 +136,7 @@ def trainer(instance_id: str, config_json: str): volumes=volumes_for(resolved), secrets=secrets_for(resolved, training=True), env=env, - experimental_options={"efa_enabled": True} if resource.nodes > 1 else {}, + experimental_options={"efa_enabled": True} if spec.trainer_nodes > 1 else {}, )(trainer) return app, trainer @@ -149,9 +148,9 @@ def run_trainer(resolved, instance_id): # Assets are prepared by the frontend before demand is registered. Reload once # on startup to see the committed exact snapshot; never race a trainer download. assets = volumes_for(resolved)["/assets"] - if spec.trainer.nodes > 1: + if spec.trainer_nodes > 1: ray_address = start_trainer_cluster( - spec.trainer.nodes, + spec.trainer_nodes, before_head=assets.reload, before_worker_join=assets.reload, ) @@ -161,7 +160,7 @@ def run_trainer(resolved, instance_id): assets.reload() ray_address = None env = { - **deployment_env(spec.trainer.env), + **deployment_env(spec.trainer_env), "LILO_APP_NAME": resolved.platform["frontend"], "LILO_BACKEND_CONFIG": json.dumps(settings), "LILO_BASE_MODEL": spec.model, @@ -176,7 +175,7 @@ def run_trainer(resolved, instance_id): env["LILO_RAY_ADDRESS"] = ray_address executor = ( "lilo.backends.miles_lora:build_executor" - if spec.trainer.backend == "miles" + if spec.backend == "miles" else "lilo.backends.megatron_fft:build_executor" ) @@ -193,9 +192,9 @@ async def failed(error): revision=config["image_id"], instance_id=instance_id, backend_env=env, - nproc=1 if spec.trainer.backend == "miles" else spec.trainer.gpus_per_node, - max_models=spec.trainer.max_clients_per_instance, - sampler_persistence_concurrency=spec.trainer.sampler_persistence_concurrency, + nproc=1 if spec.backend == "miles" else spec.trainer_gpus_per_node, + max_models=spec.trainer_max_clients_per_instance, + sampler_persistence_concurrency=spec.sampler_persistence_concurrency, on_startup_error=failed, ) @@ -212,11 +211,11 @@ def definition_from_spec(resolved, *, register_trainer=True, image=None): DEPLOYMENT_NAME=spec.name, RESOLVED=resolved, MAX_CONTEXT_LENGTH=spec.max_context_length, - TRAINER_MODELS_PER_INSTANCE=spec.trainer.max_clients_per_instance, - TRAINER_MAX_CONTAINERS=spec.trainer.max_instances, - ROLLOUT_GPUS=spec.inference.gpus_per_node, + TRAINER_MODELS_PER_INSTANCE=spec.trainer_max_clients_per_instance, + TRAINER_MAX_CONTAINERS=spec.trainer_max_instances, + ROLLOUT_GPUS=spec.inference_gpus_per_node, ROLLOUT_TENSOR_PARALLEL_SIZE=serving.get( - "tp_size", spec.inference.gpus_per_node + "tp_size", spec.inference_gpus_per_node ) // (serving.get("dp_size", 1) if serving.get("enable_dp_attention") else 1), ) @@ -236,11 +235,10 @@ def build_rollout_app(resolved, pool, *, image=None): if pool.definition_id != resolved.definition_id: raise ValueError("pool generation does not match deployment") app = modal.App(pool.app_name) - inference = spec.inference options = resolved.inference_settings - minimum = inference.min_replicas - maximum = inference.max_replicas - window = inference.scaledown_window_s + minimum = spec.inference_min_replicas + maximum = spec.inference_max_replicas + window = spec.inference_scaledown_window_s if isinstance(pool, FFTPoolSpec): minimum = minimum if pool.min_containers is None else pool.min_containers maximum = maximum if pool.max_containers is None else pool.max_containers @@ -251,17 +249,17 @@ def build_rollout_app(resolved, pool, *, image=None): name="Server", serialized=True, image=image if image is not None else image_for("sglang"), - gpu=f"{inference.gpu}:{inference.gpus_per_node}", - cpu=inference.cpu, - memory=inference.memory_mib, + gpu=f"{spec.inference_gpu}:{spec.inference_gpus_per_node}", + cpu=spec.inference_cpu, + memory=spec.inference_memory_mib, volumes=volumes_for(resolved), secrets=secrets_for(resolved), - env=deployment_env(inference.env), + env=deployment_env(spec.inference_env), min_containers=minimum, max_containers=maximum, - target_concurrency=inference.target_concurrency, + target_concurrency=spec.inference_target_concurrency, scaledown_window=window, - startup_timeout=inference.startup_timeout_s, + startup_timeout=spec.inference_startup_timeout_s, exit_grace_period=300, port=8000, routing_region=resolved.platform["modal"]["region"], @@ -287,7 +285,7 @@ def start(self): wait_http( "http://127.0.0.1:8001/health", self.sglang, - inference.startup_timeout_s, + spec.inference_startup_timeout_s, ) kwargs = dict( port=8000, @@ -309,7 +307,7 @@ def start(self): wait_http( "http://127.0.0.1:8000/health", self.sidecar, - inference.startup_timeout_s, + spec.inference_startup_timeout_s, ) except BaseException: terminate(self.sidecar) diff --git a/tests/backends/test_megatron_constructor_settings.py b/tests/backends/test_megatron_constructor_settings.py index bcbd696..bb1ad5f 100644 --- a/tests/backends/test_megatron_constructor_settings.py +++ b/tests/backends/test_megatron_constructor_settings.py @@ -1,7 +1,6 @@ """CPU checks for native config forwarding into Megatron constructors.""" -from dataclasses import asdict, make_dataclass -from pydantic import TypeAdapter +from dataclasses import make_dataclass from types import SimpleNamespace from unittest.mock import Mock @@ -11,22 +10,18 @@ from lilo.backends.deployment import backend_config from lilo.backends.megatron_config import parse_backend_config -from lilo.deployments import Deployment, load, config_path +from lilo.deployments import load, config_path with backend_runtime_imports(): from lilo.backends.megatron_runtime.common import modeling def test_config_overrides_reach_megatron(monkeypatch): - data = asdict(load(config_path("qwen35-4b-fft-64k"))) - data["trainer"]["config"]["optimizer_overrides"] = { - "native_optimizer_setting": False - } - data["trainer"]["config"]["distributed_overrides"] = {"native_ddp_setting": 123} - data["trainer"]["config"]["provider_overrides"]["native_provider_setting"] = [1, 2] - config, _ = parse_backend_config( - backend_config(TypeAdapter(Deployment).validate_python(data)) - ) + spec = load(config_path("qwen35-4b-fft-64k")) + spec.megatron_cfg["optimizer_overrides"] = {"native_optimizer_setting": False} + spec.megatron_cfg["distributed_overrides"] = {"native_ddp_setting": 123} + spec.megatron_cfg["provider_overrides"]["native_provider_setting"] = [1, 2] + config, _ = parse_backend_config(backend_config(spec)) # These stand in for an installed upstream version with extra fields. The # deployment reader must not need its own list of those fields. # Model providers in Megatron Bridge are dataclasses. diff --git a/tests/providers/test_deployment_apps.py b/tests/providers/test_deployment_apps.py index 3d53bc6..a73d9ce 100644 --- a/tests/providers/test_deployment_apps.py +++ b/tests/providers/test_deployment_apps.py @@ -1,4 +1,4 @@ -from dataclasses import replace +from copy import deepcopy import asyncio import json from types import SimpleNamespace @@ -86,7 +86,7 @@ def test_trainer_declaration_and_executor_configuration( assert kwargs["backend_env"]["LILO_CHECKPOINT_VOLUME"] == "test-custom-checkpoints" assert kwargs["backend_env"]["LILO_BASE_MODEL_REVISION"] == "a" * 40 config = json.loads(kwargs["backend_env"]["LILO_BACKEND_CONFIG"]) - assert config[row.spec.trainer.backend]["hf_checkpoint"] == row.asset_path + assert config[row.spec.backend]["hf_checkpoint"] == row.asset_path assert config["checkpoint_dir"] == "/checkpoints" assert reloaded == [True] @@ -108,7 +108,7 @@ def test_pool_starts_native_server_and_correct_sidecar(builders, monkeypatch, ki assert app.name == pool.app_name assert ( settings["gpu"] - == f"{row.spec.inference.gpu}:{row.spec.inference.gpus_per_node}" + == f"{row.spec.inference_gpu}:{row.spec.inference_gpus_per_node}" ) assert settings["min_containers"] == 0 assert settings["target_concurrency"] == 16 @@ -235,13 +235,7 @@ def test_admission_changes_preserve_serialized_trainer(builders): new_bytes = serialize(deployment_apps.build_trainer_app(changed, image="test")[1]) assert new_bytes == old_bytes assert first.active is True - changed.spec = replace( - changed.spec, - trainer=replace( - changed.spec.trainer, - gpu="H200", - ), - ) + changed.spec.trainer_gpu = "H200" assert ( serialize(deployment_apps.build_trainer_app(changed, image="test")[1]) != old_bytes @@ -425,24 +419,15 @@ async def no_error(definition_id): def test_declared_compute_settings_reach_modal(builders): - base = deployment().spec - spec = replace( - base, - trainer=replace( - base.trainer, - timeout_s=90, - cpu=12, - memory_mib=123456, - ), - inference=replace( - base.inference, - startup_timeout_s=90, - min_replicas=1, - max_replicas=3, - cpu=6, - memory_mib=45000, - ), - ) + spec = deepcopy(deployment().spec) + spec.trainer_timeout_s = 90 + spec.trainer_cpu = 12 + spec.trainer_memory_mib = 123456 + spec.inference_startup_timeout_s = 90 + spec.inference_min_replicas = 1 + spec.inference_max_replicas = 3 + spec.inference_cpu = 6 + spec.inference_memory_mib = 45000 row = DeploymentRecord.create(spec, revision="a" * 40) trainer_app, _ = deployment_apps.build_trainer_app(row, image="test") trainer, _ = trainer_app.functions["trainer"] diff --git a/tests/providers/test_deployment_e2e_helper.py b/tests/providers/test_deployment_e2e_helper.py index 9541207..2e59e5c 100644 --- a/tests/providers/test_deployment_e2e_helper.py +++ b/tests/providers/test_deployment_e2e_helper.py @@ -22,7 +22,7 @@ def test_e2e_helper_reads_active_deployed_configuration(monkeypatch, preset): definition, mode = helper["_definition"]("test-frontend", row.spec.name) assert definition.DEFINITION_ID == row.definition_id assert definition.MAX_CONTEXT_LENGTH == row.spec.max_context_length - assert definition.GPUS == row.spec.trainer.gpus_per_node + assert definition.GPUS == row.spec.trainer_gpus_per_node assert mode == row.spec.parameterization assert definition.MAX_TOKENS_PER_MICROBATCH > 0 with pytest.raises(ValueError, match="one active YAML"): diff --git a/tests/providers/test_deployment_presets.py b/tests/providers/test_deployment_presets.py index 1a5b46e..cc9a51f 100644 --- a/tests/providers/test_deployment_presets.py +++ b/tests/providers/test_deployment_presets.py @@ -1,5 +1,3 @@ -from dataclasses import asdict -from pydantic import TypeAdapter import pytest from lilo.deployments import load, config_path, DeploymentRecord @@ -19,7 +17,7 @@ def test_all_packaged_recipes_validate_offline(path): spec = load(path) config = backend_config(spec) - assert config[spec.trainer.backend]["hf_checkpoint"] == "/assets/pending" + assert config[spec.backend]["hf_checkpoint"] == "/assets/pending" serving_options(spec) @@ -56,11 +54,11 @@ def test_qwen38_context_parallel_token_budget(context, cp): def test_single_client_recipe_keeps_shared_backend_capacity(): shared = load(config_path("qwen35-9b-lora-16k")) single = load(config_path("qwen35-9b-lora-16k-single")) - assert single.trainer.max_clients_per_instance == 1 - assert (single.trainer.gpu, single.trainer.gpus_per_node, single.trainer.nodes) == ( - shared.trainer.gpu, - shared.trainer.gpus_per_node, - shared.trainer.nodes, + assert single.trainer_max_clients_per_instance == 1 + assert (single.trainer_gpu, single.trainer_gpus_per_node, single.trainer_nodes) == ( + shared.trainer_gpu, + shared.trainer_gpus_per_node, + shared.trainer_nodes, ) assert backend_config(single) == backend_config(shared) @@ -75,23 +73,3 @@ def test_missing_manifest_has_no_python_catalog_fallback(monkeypatch, value): manifest_from_env() with pytest.raises(ValueError, match="manifest"): frontend_settings() - - -@pytest.mark.parametrize( - "options", - [ - {"dp_size": 3, "enable_dp_attention": True}, - {"dp_size": 2, "enable_dp_attention": False}, - {"dp_size": 2, "enable_dp_attention": "true"}, - ], -) -def test_invalid_attention_parallelism_is_rejected(options): - from lilo.deployments import Deployment - - data = asdict(load(config_path("qwen35-35b-a3b-fft-64k"))) - data["inference"]["config"].update(options) - from lilo.backends.deployment import serving_options - - spec = TypeAdapter(Deployment).validate_python(data) - with pytest.raises(ValueError, match="sglang"): - serving_options(spec) diff --git a/tests/test_deployment_cli.py b/tests/test_deployment_cli.py index e7439c4..c9949b3 100644 --- a/tests/test_deployment_cli.py +++ b/tests/test_deployment_cli.py @@ -1,5 +1,4 @@ from copy import deepcopy -from dataclasses import replace import json import subprocess @@ -73,7 +72,7 @@ def fail(*args, **kwargs): assert registry["pending"][0]["generation"] == row.generation assert "manifest" not in registry and "apply_lock" not in registry new_spec = deepcopy(row.spec) - new_spec = replace(new_spec, trainer=replace(new_spec.trainer, max_instances=2)) + new_spec.trainer_max_instances = 2 new = DeploymentRecord.create(new_spec, revision="a" * 40) monkeypatch.setattr(subprocess, "run", lambda *args, **kwargs: None) cli.deploy([new]) @@ -170,7 +169,7 @@ def run(command, **kwargs): assert calls == ["frontend"] changed = deepcopy(row.spec) - changed = replace(changed, inference=replace(changed.inference, max_replicas=6)) + changed.inference_max_replicas = 6 new = DeploymentRecord.create(changed, revision="a" * 40) calls.clear() cli.deploy([new]) @@ -274,7 +273,7 @@ def test_worker_refresh_is_retained_without_editing_config(registry, monkeypatch active = DeploymentRecord.model_validate(registry["manifest"][0]) assert active.inference_release != "initial" assert active.trainer_release == "initial" - assert active.spec.inference == row.spec.inference + assert vars(active.spec) == vars(row.spec) calls.clear() cli.deploy([row]) assert calls == ["frontend"] diff --git a/tests/test_deployments.py b/tests/test_deployments.py index c970f37..5020e1e 100644 --- a/tests/test_deployments.py +++ b/tests/test_deployments.py @@ -1,5 +1,5 @@ -from dataclasses import asdict, replace, FrozenInstanceError -from pydantic import TypeAdapter +from lilo.configuration import BaseConfig +from copy import deepcopy import argparse import asyncio from types import SimpleNamespace @@ -9,7 +9,6 @@ from lilo.deployments import ( DeploymentRecord, - Deployment, load, config_path, validate_frontend, @@ -22,14 +21,14 @@ def recipe(preset="qwen35-9b-lora-16k", **changes): - data = asdict(load(config_path(preset))) + data = vars(load(config_path(preset))) for path, value in changes.items(): keys = path.split("__") target = data for key in keys[:-1]: target = target[key] target[keys[-1]] = value - return TypeAdapter(Deployment).validate_python(data) + return BaseConfig(**data) def resolved(spec=None, **changes): @@ -61,8 +60,8 @@ def test_presets_context_topology_and_backend_options(): def test_no_model_catalog_required(): spec = recipe( model="my-org/new-model", - trainer__config__model_type="", - trainer__config__cli_options={ + miles_cfg__model_type="", + miles_cfg__cli_options={ "num_layers": 12, "hidden_size": 768, "num_attention_heads": 12, @@ -75,23 +74,14 @@ def test_no_model_catalog_required(): @pytest.mark.parametrize( "changes,match", [ - ({"trainer__config__cli_options": {"hf_checkpoint": "other"}}, "managed"), + ({"miles_cfg__cli_options": {"hf_checkpoint": "other"}}, "managed"), ( - {"trainer__config__cli_options": {"pipeline_model_parallel_size": 2}}, + {"miles_cfg__cli_options": {"pipeline_model_parallel_size": 2}}, "managed", ), - ({"inference__config": {"model_path": "other"}}, "managed"), - ({"inference__config": {"tp_size": 2}}, "replica GPU"), - ({"trainer__max_clients_per_instance": 7}, "max_lora_slots"), - ( - { - "inference__config": { - "max_loaded_loras": 2, - "max_loras_per_batch": 8, - } - }, - "max_loaded_loras", - ), + ({"sglang_cfg": {"model_path": "other"}}, "managed"), + ({"sglang_cfg": {"tp_size": 2}}, "replica GPU"), + ({"trainer_max_clients_per_instance": 7}, "max_lora_slots"), ], ) def test_invalid_integrations_fail_when_building_backend_settings(changes, match): @@ -105,7 +95,7 @@ def test_invalid_integrations_fail_when_building_backend_settings(changes, match def test_generation_and_asset_paths_include_exact_base(): a = resolved() - assert a.generation != resolved(recipe(trainer__gpu="H200")).generation + assert a.generation != resolved(recipe(trainer_gpu="H200")).generation b = resolved(recipe(model="other/Qwen3.5-9B-Base")) assert a.asset_path != b.asset_path assert a.asset_path != DeploymentRecord.create(a.spec, revision="b" * 40).asset_path @@ -258,16 +248,14 @@ def test_native_sections_survive_serialization_without_allowlist(): from lilo.backends.megatron_config import parse_backend_config spec = recipe("qwen35-4b-fft-64k") - data = asdict(spec) - data["trainer"]["config"]["provider_overrides"]["future_provider_option"] = { + data = vars(spec) + data["megatron_cfg"]["provider_overrides"]["future_provider_option"] = { "layers": [1, 4], "enabled": False, } - data["trainer"]["config"]["optimizer_overrides"] = { - "future_optimizer_option": 0.125 - } - data["trainer"]["config"]["distributed_overrides"] = {"future_ddp_option": False} - spec = TypeAdapter(Deployment).validate_python(data) + data["megatron_cfg"]["optimizer_overrides"] = {"future_optimizer_option": 0.125} + data["megatron_cfg"]["distributed_overrides"] = {"future_ddp_option": False} + spec = BaseConfig(**data) settings = backend_config(spec, "/assets/pinned") config, _ = parse_backend_config(json.loads(json.dumps(settings))) assert config.hf_checkpoint == "/assets/pinned" @@ -279,7 +267,7 @@ def test_native_sections_survive_serialization_without_allowlist(): assert config.optimizer_overrides == {"future_optimizer_option": 0.125} assert config.distributed_overrides == {"future_ddp_option": False} assert config.optimizer.lr == 0.0001 - assert asdict(spec) == data # Building does not consume or mutate the config. + assert vars(spec) == data # Building does not consume or mutate the config. @pytest.mark.parametrize( @@ -293,22 +281,22 @@ def test_native_sections_survive_serialization_without_allowlist(): ], ) def test_megatron_cli_options_preserve_integration_contract(section, options, match): - spec = recipe("qwen35-4b-fft-64k", **{f"trainer__config__{section}": options}) + spec = recipe("qwen35-4b-fft-64k", **{f"megatron_cfg__{section}": options}) with pytest.raises(ValueError, match=match): backend_config(spec) def test_backend_dispatch_rejects_unknown_backend(): with pytest.raises(ValueError, match="backend"): - recipe(trainer__backend="missing") + backend_config(recipe(backend="missing")) def test_new_miles_and_sglang_options_need_no_deployment_schema_change(): from lilo.backends.deployment import serving_options spec = recipe( - trainer__config__cli_options__future_miles_option=[1, 2], - inference__config__future_sglang_option=False, + miles_cfg__cli_options__future_miles_option=[1, 2], + sglang_cfg__future_sglang_option=False, ) assert backend_config(spec)["miles"]["cli_options"]["future_miles_option"] == [ 1, @@ -319,47 +307,25 @@ def test_new_miles_and_sglang_options_need_no_deployment_schema_change(): def test_fft_capacity_is_checked_by_backend_setup(): with pytest.raises(ValueError, match="FFT trainers admit one client"): - recipe("qwen35-4b-fft-64k", trainer__max_clients_per_instance=2) + backend_config(recipe("qwen35-4b-fft-64k", trainer_max_clients_per_instance=2)) def test_reserved_environment_is_checked_by_modal_setup(): from lilo.providers.modal.deployment_apps import deployment_env with pytest.raises(ValueError, match="managed"): - recipe(trainer__env={"LILO_BACKEND_CONFIG": "oops"}) + deployment_env({"LILO_BACKEND_CONFIG": "oops"}) assert deployment_env({"MY_SETTING": "value"}) == {"MY_SETTING": "value"} def test_record_creation_copies_without_reparsing(): - import hashlib - import json - spec = recipe() - original = asdict(spec) + original = vars(spec) row = DeploymentRecord.create(spec, revision="a" * 40) - assert asdict(spec) == original + assert vars(spec) == original assert row.spec.revision == "a" * 40 - # The record hash covers settings and the pinned backend dependency, not source. - expected = original | {"revision": "a" * 40} - assert ( - row.generation - == hashlib.sha256( - json.dumps( - { - "config": expected, - "miles_commit": None, - "platform": row.platform, - "trainer_release": "initial", - "inference_release": "initial", - "trainer_settings": row.trainer_settings, - "inference_settings": row.inference_settings, - }, - sort_keys=True, - ).encode() - ).hexdigest() - ) - row.spec.trainer.config["max_lora_rank"] = 64 - assert spec.trainer.config["max_lora_rank"] == 32 + row.spec.miles_cfg["max_lora_rank"] = 64 + assert spec.miles_cfg["max_lora_rank"] == 32 saved = row.model_dump_json() assert DeploymentRecord.model_validate_json(saved).model_dump( mode="json" @@ -372,17 +338,18 @@ def test_python_config_composition(tmp_path): "from lilo.configs.qwen35_9b_lora_16k import Config as Parent\n" "class Config(Parent):\n" " name = 'custom'\n" - " overrides = {'trainer.memory_mib': 123456}\n" + " overrides = {'trainer_memory_mib': 123456}\n" "config = Config()\n" ) custom = load(path) original = recipe() assert custom.name == "custom" - assert custom.trainer.memory_mib == 123456 - assert custom.trainer.config == original.trainer.config - assert original.trainer.memory_mib == 65536 - with pytest.raises(FrozenInstanceError): - custom.trainer.max_instances = 9 + assert custom.trainer_memory_mib == 123456 + assert custom.miles_cfg == original.miles_cfg + assert original.trainer_memory_mib == 65536 + custom.trainer_max_instances = 9 + assert custom.trainer_max_instances == 9 + assert original.trainer_max_instances == 1 def test_worker_record_contains_resolved_settings(monkeypatch): @@ -418,37 +385,17 @@ def test_no_yaml_config_ingestion(tmp_path): load(tmp_path / "old.yaml") -@pytest.mark.parametrize( - "section,key", - [ - ("trainer", "memroy_mib"), - ("trainer", "max_instnaces"), - ("trainer", "min_instances"), - ("inference", "timeout_s"), - ("inference", "nodes"), - ], -) -def test_orchestration_typos_and_unused_fields_are_rejected(section, key): - base = recipe() - component = { - "trainer": base.trainer, - "inference": base.inference, - }[section] - with pytest.raises(ValueError, match=key): - replace(component, **{key: 9}) - - def test_worker_hashes_cover_only_their_settings(): base = resolved() - inference = resolved(recipe(inference__config__max_running_requests=24)) + inference = resolved(recipe(sglang_cfg__max_running_requests=24)) assert inference.trainer_hash == base.trainer_hash assert inference.inference_hash != base.inference_hash - trainer = resolved(recipe(trainer__config__max_tokens_per_gpu=8192)) + trainer = resolved(recipe(miles_cfg__max_tokens_per_gpu=8192)) assert trainer.trainer_hash != base.trainer_hash assert trainer.inference_hash == base.inference_hash - adapter = resolved(recipe(trainer__config__max_lora_rank=64)) + adapter = resolved(recipe(miles_cfg__max_lora_rank=64)) assert adapter.trainer_hash != base.trainer_hash assert adapter.inference_hash != base.inference_hash @@ -467,8 +414,6 @@ def test_examples_only_contain_model_infrastructure(): config = load(path) assert not hasattr(config, "deployment") assert config.revision == "main" - assert "runtime_version" not in asdict(config.trainer) - assert "runtime_version" not in asdict(config.inference) @pytest.mark.parametrize("value", ["invalid", 7]) @@ -493,8 +438,8 @@ def test_backend_parser_validates_configured_types_and_choices(value): ) def test_managed_backend_values_fail_before_record_creation(section, field): base = recipe("qwen35-4b-fft-64k") - config = {**base.trainer.config, section: {field: 1}} - candidate = replace(base, trainer=replace(base.trainer, config=config)) + candidate = deepcopy(base) + candidate.megatron_cfg[section] = {field: 1} with pytest.raises(ValueError, match=field): DeploymentRecord.create(candidate, revision="a" * 40) @@ -507,13 +452,8 @@ def test_multinode_ownership_and_topology(): assert miles["actor_num_gpus_per_node"] == 8 assert miles["tensor_model_parallel_size"] == 2 assert miles["context_parallel_size"] == 8 - invalid = replace( - config, - trainer=replace( - config.trainer, - config={**config.trainer.config, "actor_num_nodes": 3}, - ), - ) + invalid = deepcopy(config) + invalid.miles_cfg["actor_num_nodes"] = 3 with pytest.raises(ValueError, match="actor_num_nodes"): DeploymentRecord.create(invalid, revision="a" * 40) @@ -524,49 +464,32 @@ def test_config_inheritance_and_constructor_overrides_copy_nested_options(): class Child(Parent): name = "child" max_context_length = 8192 - overrides = {"trainer.config.max_tokens_per_gpu": 8192} + overrides = {"miles_cfg.max_tokens_per_gpu": 8192} first, second = Child(), Child(name="second") - first.trainer.config["target_modules"].append("extra") - first.trainer.config["cli_options"]["recompute_num_layers"] = 2 + first.miles_cfg["target_modules"].append("extra") + first.miles_cfg["cli_options"]["recompute_num_layers"] = 2 assert second.name == "second" assert second.max_context_length == 8192 - assert second.trainer.gpu == "H100" + assert second.trainer_gpu == "H100" assert backend_config(second)["miles"]["max_tokens_per_gpu"] == 8192 assert Parent().max_context_length == 16384 for config in (second, Parent()): - assert "extra" not in config.trainer.config["target_modules"] - assert config.trainer.config["cli_options"]["recompute_num_layers"] == 1 + assert "extra" not in config.miles_cfg["target_modules"] + assert config.miles_cfg["cli_options"]["recompute_num_layers"] == 1 -@pytest.mark.parametrize( - "field,value", - [ - ("max_contex_length", 8192), - ("max_context_length", "8192"), - ("pool_idle_timeout_s", "300"), - ], -) -def test_recipe_class_fields_and_constructor_overrides_are_validated(field, value): - from lilo.configs.qwen35_9b_lora_16k import Config as Parent - - cls = type("Invalid", (Parent,), {field: value}) - with pytest.raises(ValueError, match=field): - cls() - with pytest.raises(ValueError, match=field): - Parent(**{field: value}) - - -def test_recipe_section_replacement_uses_defaults_without_implicit_merge(): +def test_backend_dictionary_assignment_replaces_inherited_values(): from lilo.configs.qwen35_9b_lora_16k import Config as Parent class Child(Parent): - inference = {"gpu": "H100", "config": {"max_running_requests": 4}} + inference_gpu = "H100" + sglang_cfg = {"max_running_requests": 4} config = Child() - assert config.inference.gpu == "H100" - assert config.inference.max_replicas == 8 - assert config.inference.config == {"max_running_requests": 4} + assert config.inference_gpu == "H100" + assert config.inference_max_replicas == 8 + assert config.sglang_cfg == {"max_running_requests": 4} def test_overrides_compose_across_generations_and_constructor(): @@ -574,76 +497,54 @@ def test_overrides_compose_across_generations_and_constructor(): class Child(Parent): overrides = { - "trainer.gpu": "H200", - "trainer.env.FIRST": "1", - "trainer.config.max_tokens_per_gpu": 8192, - "inference.config.future_option.nested": [1, 2], + "trainer_gpu": "H200", + "trainer_env.FIRST": "1", + "miles_cfg.max_tokens_per_gpu": 8192, + "sglang_cfg.future_option.nested": [1, 2], } class Grandchild(Child): overrides = { - "trainer.config.max_tokens_per_gpu": 4096, - "trainer.env.SECOND": "2", + "miles_cfg.max_tokens_per_gpu": 4096, + "trainer_env.SECOND": "2", } config = Grandchild( name="custom", - overrides={"trainer.config.max_tokens_per_gpu": 2048}, + overrides={"miles_cfg.max_tokens_per_gpu": 2048}, ) assert config.name == "custom" - assert config.trainer.gpu == "H200" - assert config.trainer.env["FIRST"] == "1" - assert config.trainer.env["SECOND"] == "2" - assert config.trainer.config["max_tokens_per_gpu"] == 2048 - assert Grandchild().trainer.config["max_tokens_per_gpu"] == 4096 - assert Child().trainer.config["max_tokens_per_gpu"] == 8192 - config.inference.config["future_option"]["nested"].append(3) - assert Child.overrides["inference.config.future_option.nested"] == [1, 2] - assert Grandchild().inference.config["future_option"]["nested"] == [1, 2] - assert "FIRST" not in Parent().trainer.env + assert config.trainer_gpu == "H200" + assert config.trainer_env["FIRST"] == "1" + assert config.trainer_env["SECOND"] == "2" + assert config.miles_cfg["max_tokens_per_gpu"] == 2048 + assert Grandchild().miles_cfg["max_tokens_per_gpu"] == 4096 + assert Child().miles_cfg["max_tokens_per_gpu"] == 8192 + config.sglang_cfg["future_option"]["nested"].append(3) + assert Child.overrides["sglang_cfg.future_option.nested"] == [1, 2] + assert Grandchild().sglang_cfg["future_option"]["nested"] == [1, 2] + assert "FIRST" not in Parent().trainer_env def test_override_values_replace_dictionaries_and_lists(): from lilo.configs.qwen35_9b_lora_16k import Config as Parent class Child(Parent): - overrides = {"trainer.env": {"FIRST": "1"}} + overrides = {"trainer_env": {"FIRST": "1"}} class Grandchild(Child): overrides = { - "trainer.env": {}, - "trainer.config.target_modules": ["q_proj"], + "trainer_env": {}, + "miles_cfg.target_modules": ["q_proj"], } config = Grandchild() - assert config.trainer.env == {} - assert config.trainer.config["target_modules"] == ["q_proj"] - assert Child().trainer.env == {"FIRST": "1"} + assert config.trainer_env == {} + assert config.miles_cfg["target_modules"] == ["q_proj"] + assert Child().trainer_env == {"FIRST": "1"} # Constructor fields replace the inherited section, then overrides apply. - config = Grandchild(inference={"gpu": "H100"}, overrides={"inference.gpu": "H200"}) - assert config.inference.gpu == "H200" - assert config.inference.config == {} - assert replace(config, name="copy").trainer == config.trainer - - -@pytest.mark.parametrize( - "overrides,match", - [ - ([], "overrides must be a dictionary"), - ({"": 1}, "invalid override path"), - ({"trainer..gpu": "H200"}, "invalid override path"), - ({1: "H200"}, "invalid override path"), - ({"trainer.gpu.type": "H200"}, "non-dictionary"), - ({"trainer.memroy_mib": 123}, "memroy_mib"), - ({"trianer.gpu": "H200"}, "trianer"), - ({"trainer.gpus_per_node": "4"}, "gpus_per_node"), - ], -) -def test_dotted_overrides_are_validated(overrides, match): - from lilo.configs.qwen35_9b_lora_16k import Config as Parent - - cls = type("Invalid", (Parent,), {"overrides": overrides}) - with pytest.raises(ValueError, match=match): - cls() - with pytest.raises(ValueError, match=match): - Parent(overrides=overrides) + config = Grandchild( + sglang_cfg={}, inference_gpu="H100", overrides={"inference_gpu": "H200"} + ) + assert config.inference_gpu == "H200" + assert config.sglang_cfg == {} From 7915b44d4d82acd2e175c37aeec635da06fadfb1 Mon Sep 17 00:00:00 2001 From: kailash Date: Thu, 24 Sep 2026 17:09:47 +0000 Subject: [PATCH 25/27] Use deployed Modal apps instead of a separate deployment registry --- README.md | 2 +- docs/deployment-configs.md | 6 +- docs/deployment-validation.md | 4 +- scripts/definition_smoke.py | 8 +- scripts/deployment_smoke.py | 5 +- scripts/e2e_engine_definition.py | 5 +- src/lilo/deployment_cli.py | 157 +++------- src/lilo/deployments.py | 4 +- src/lilo/providers/modal/app.py | 7 + .../providers/modal/deployment_records.py | 12 + src/lilo/providers/modal/lora_pool.py | 2 +- tests/providers/test_checkpoint_storage.py | 2 +- tests/providers/test_deployment_apps.py | 7 +- tests/providers/test_deployment_e2e_helper.py | 18 +- tests/providers/test_lora_pool.py | 6 +- tests/test_deployment_cli.py | 294 +++++++----------- 16 files changed, 203 insertions(+), 336 deletions(-) diff --git a/README.md b/README.md index 0fb382a..33658e0 100644 --- a/README.md +++ b/README.md @@ -156,7 +156,7 @@ has finished in the Modal dashboard or list apps with: uv run modal app list ``` -To tear down the deployment, stop its `lilo-fft-...` sampler apps, then the frontend named in your config (`lilo-yaml` by default), +To tear down the deployment, stop its `lilo-fft-...` sampler apps, then the frontend selected with `--app` (`lilo` by default), using `uv run modal app stop `. Stopping the frontend does not stop sampler apps. ## Next steps diff --git a/docs/deployment-configs.md b/docs/deployment-configs.md index c353317..313e0d2 100644 --- a/docs/deployment-configs.md +++ b/docs/deployment-configs.md @@ -130,6 +130,8 @@ SHA-256 hashes sorted JSON. Trainer and inference app names use the first 24 hex An inference-only change reuses the trainer app. A trainer batch-setting change reuses the inference provisioner. Changes to adapter rank/targets update both. Code-only upgrades use the refresh flags; later ordinary deploys retain those releases. -Old jobs keep their recorded worker apps. Deployment retries reuse workers already created successfully. Trainer limits are enforced per saved definition; old and new definitions may use their configured capacity simultaneously while old jobs finish. Pools remain associated with full job configurations, so new jobs may get separate pools even if they share an inference provisioner. +The frontend carries its resolved configurations and exposes them through the Modal `deployment_manifest` function. The CLI reads that function to retain existing definitions and worker releases, and checks worker existence directly with Modal. There is no separate deployment Dict, pending journal, or deployment lock. Run deploy commands sequentially. -Old worker app definitions are retained; automatic cleanup is not implemented. Changes to the frontend/worker protocol still need deliberate compatibility handling. Earlier draft manifests need migration or a fresh frontend/registry. See [validation history](deployment-validation.md) for the distinction between current CPU checks and historical GPU runs. +Old jobs keep their recorded worker apps. Deployment retries discover and reuse workers already created successfully. Trainer limits are enforced per saved definition; old and new definitions may use their configured capacity simultaneously while old jobs finish. Pools remain associated with full job configurations, so new jobs may get separate pools even if they share an inference provisioner. + +Old worker app definitions are retained; automatic cleanup is not implemented. Changes to the frontend/worker protocol still need deliberate compatibility handling. Earlier draft config formats require migration or a fresh frontend. See [validation history](deployment-validation.md) for the distinction between current CPU checks and historical GPU runs. diff --git a/docs/deployment-validation.md b/docs/deployment-validation.md index d36de43..bb03f14 100644 --- a/docs/deployment-validation.md +++ b/docs/deployment-validation.md @@ -8,7 +8,9 @@ Recipes use one plain `BaseConfig` with untyped attributes such as `trainer_gpu` Variants inherit class attributes and use dotted `overrides` for nested backend options. Mutable settings are copied per instance. Workers consume saved resolved backend settings. Routing uses deployment order or explicit definition IDs, without visibility or default flags. -Validation: **644 CPU tests passed, 1 skipped**, including **111 focused config/deployment tests**. Coverage includes inherited overrides, mutable option isolation, backend passthrough, saved-record JSON round trips, worker resources, update isolation and deployment recovery. The CLI-generated recipe loaded and validated successfully. Ruff and whitespace checks passed. +Validation: **642 CPU tests passed, 1 skipped**, including **41 focused deployment/provider tests**. Coverage includes inherited overrides, mutable option isolation, backend passthrough, saved-record JSON round trips, worker resources, update isolation and deployment recovery. The CLI-generated recipe loaded and validated successfully. Ruff and whitespace checks passed. + +Deployment uses Modal app lookups and the frontend’s `deployment_manifest` function. The separate Dict registry, pending records, deployment lock, unlock command and registry-less-frontend rejection are removed. Tests cover worker reuse, retries after partial deployment, missing worker recreation, retained configurations and code refresh. The default frontend is `lilo`; definition IDs use `deployment_`. All 15 recipes preserve compute, model, lifecycle and resolved trainer/inference settings. The flat saved config layout changes deployment hashes. Existing draft manifests need regeneration or explicit migration. No deployments or GPU jobs were modified; GPU backend startup has not been rerun for this revision. diff --git a/scripts/definition_smoke.py b/scripts/definition_smoke.py index f342874..4a1242c 100644 --- a/scripts/definition_smoke.py +++ b/scripts/definition_smoke.py @@ -31,7 +31,6 @@ from pathlib import Path from typing import Any -import modal import httpx import tinker from tinker import types @@ -39,6 +38,7 @@ from lilo.backends.miles_config import lora_target_flags from lilo.client import create_full_training_client from lilo.deployments import DeploymentRecord +from lilo.providers.modal.deployment_records import deployed_manifest from lilo.providers.modal.deployment_apps import definition_from_spec TIMEOUT = 60 * 60 @@ -295,12 +295,10 @@ def main() -> None: type=Path, default=Path("scripts/results/definition_smoke.json"), ) - parser.add_argument("--app", default="lilo-yaml") + parser.add_argument("--app", default="lilo") parser.add_argument("--env") args = parser.parse_args() - rows = modal.Dict.from_name( - f"{args.app}-yaml-deployments", environment_name=args.env - ).get("manifest", []) + rows = deployed_manifest(args.app, args.env) definitions = { row.definition_id: definition_from_spec(row, register_trainer=False) for data in rows diff --git a/scripts/deployment_smoke.py b/scripts/deployment_smoke.py index ef94c55..dca378a 100644 --- a/scripts/deployment_smoke.py +++ b/scripts/deployment_smoke.py @@ -18,6 +18,7 @@ import tinker from tinker import types +from lilo.providers.modal.deployment_records import deployed_manifest from lilo.client import create_full_training_client from lilo.providers.modal.kv import app_store_name @@ -100,9 +101,9 @@ def run(args): server = modal.Function.from_name(args.frontend, "server") url = server.get_web_url() headers = {"X-API-Key": os.environ["TINKER_API_KEY"]} - rows = modal.Dict.from_name(f"{args.frontend}-yaml-deployments").get("manifest") + rows = deployed_manifest(args.frontend) row = next(r for r in rows if r["active"] and r["spec"]["name"] == args.name) - definition_id = f"yaml_{args.name}_{row['generation'][:16]}" + definition_id = f"deployment_{args.name}_{row['generation'][:16]}" report = { "frontend": args.frontend, "app_id": app.app_id, diff --git a/scripts/e2e_engine_definition.py b/scripts/e2e_engine_definition.py index dae7f0a..e39985c 100644 --- a/scripts/e2e_engine_definition.py +++ b/scripts/e2e_engine_definition.py @@ -17,13 +17,14 @@ from tinker import types from lilo.deployments import DeploymentRecord +from lilo.providers.modal.deployment_records import deployed_manifest from lilo.backends.deployment import backend_config TIMEOUT = 3 * 60 * 60 def _definition(frontend: str, name: str) -> tuple[Any, str]: - rows = modal.Dict.from_name(f"{frontend}-yaml-deployments").get("manifest", []) + rows = deployed_manifest(frontend) matches = [ DeploymentRecord.model_validate(row) for row in rows @@ -31,7 +32,7 @@ def _definition(frontend: str, name: str) -> tuple[Any, str]: ] if len(matches) != 1: raise ValueError( - f"expected one active YAML configuration named {name} in {frontend}" + f"expected one active configuration named {name} in {frontend}" ) resolved = matches[0] spec = resolved.spec diff --git a/src/lilo/deployment_cli.py b/src/lilo/deployment_cli.py index d2097f3..9ca6abf 100644 --- a/src/lilo/deployment_cli.py +++ b/src/lilo/deployment_cli.py @@ -24,7 +24,7 @@ load, validate_frontend, ) -from lilo.providers.modal.deployment_records import MANIFEST_ENV +from lilo.providers.modal.deployment_records import MANIFEST_ENV, deployed_manifest from lilo.providers.modal.miles_revision import resolve_miles_commit @@ -104,96 +104,34 @@ def select_worker_releases(desired, previous, refresh_trainers, refresh_inferenc return result -def worker_apps_ready(row, deployed): - return row.trainer_app_name in deployed and row.inference_app_name in deployed - - def deploy(desired, *, refresh_trainers=(), refresh_inference=()): - """Serialize operator applies and retain interrupted attempts for safe recovery.""" - + """Deploy missing workers, then publish the frontend's configuration.""" if sys.version_info[:2] != (3, 12): raise ValueError( "Python deployment requires Python 3.12 to match the serialized GPU runtime images" ) settings = desired[0].platform - registry = modal.Dict.from_name( - f"{settings['frontend']}-yaml-deployments", - create_if_missing=True, - environment_name=settings["modal"]["environment"], + environment = settings["modal"]["environment"] + previous = [ + DeploymentRecord.model_validate(row) + for row in deployed_manifest(settings["frontend"], environment) + ] + desired = select_worker_releases( + desired, previous, refresh_trainers, refresh_inference ) - owner = uuid.uuid4().hex - if not registry.put("apply_lock", owner, skip_if_exists=True): - raise ValueError( - f"An apply owns {settings['frontend']}. If it was interrupted, confirm it has stopped before running lilo deployment unlock --frontend {settings['frontend']}." - ) - try: - rows = registry.get("manifest", []) - if not rows and not registry.get("pending", []): + manifest = retain_generations(previous, desired) + command = [sys.executable, "-m", "modal", "deploy"] + if environment: + command += ["--env", environment] + + for row in desired: + for role, app_name in ( + ("trainer", row.trainer_app_name), + ("inference", row.inference_app_name), + ): try: - modal.App.lookup( - settings["frontend"], - environment_name=settings["modal"]["environment"], - ) + modal.App.lookup(app_name, environment_name=environment) except modal.exception.NotFoundError: - pass - else: - raise ValueError( - "The frontend already exists without a deployment registry. Choose a new frontend name; an app with no deployment registry cannot be safely updated." - ) - # Only complete worker pairs could have been exposed by a pending frontend. - deployed = set(registry.get("worker_apps", [])) - pending = [ - row - for row in registry.get("pending", []) - if worker_apps_ready(DeploymentRecord.model_validate(row), deployed) - ] - rows = {r["generation"]: r for r in [*rows, *pending]} - previous = [DeploymentRecord.model_validate(row) for row in rows.values()] - desired = select_worker_releases( - desired, previous, refresh_trainers, refresh_inference - ) - manifest = retain_generations(previous, desired) - data = [row.model_dump(mode="json") for row in manifest] - env = { - **os.environ, - MANIFEST_ENV: json.dumps(data), - "LILO_APP_NAME": settings["frontend"], - } - command = [ - sys.executable, - "-m", - "modal", - "deploy", - "-m", - "lilo.providers.modal.app", - ] - if settings["modal"]["environment"]: - command += ["--env", settings["modal"]["environment"]] - registry.put("pending", data) - # Deploy each worker app once. Retained apps keep their original code. - deployed = set(registry.get("worker_apps", [])) - for row in manifest: - for role, app_name in ( - ("trainer", row.trainer_app_name), - ("inference", row.inference_app_name), - ): - if app_name in deployed: - continue - if not row.active: - raise ValueError( - f"Retained worker {app_name} is missing; restore its original deployment." - ) - # Recover a crash after Modal succeeded but before the registry write. - try: - modal.App.lookup( - app_name, environment_name=settings["modal"]["environment"] - ) - except modal.exception.NotFoundError: - pass - else: - deployed.add(app_name) - registry.put("worker_apps", sorted(deployed)) - continue worker_env = { **os.environ, "LILO_WORKER_DEPLOYMENT": row.model_dump_json(), @@ -201,25 +139,20 @@ def deploy(desired, *, refresh_trainers=(), refresh_inference=()): } if row.miles_commit: worker_env["LILO_MILES_COMMIT"] = row.miles_commit - worker_command = [ - sys.executable, - "-m", - "modal", - "deploy", - "-m", - "lilo.providers.modal.deployment_worker_app", - ] - if settings["modal"]["environment"]: - worker_command += ["--env", settings["modal"]["environment"]] - subprocess.run(worker_command, check=True, env=worker_env) - deployed.add(app_name) - registry.put("worker_apps", sorted(deployed)) - subprocess.run(command, check=True, env=env) - registry.put("manifest", data) - registry.pop("pending", None) - finally: - if registry.get("apply_lock") == owner: - registry.pop("apply_lock", None) + subprocess.run( + [*command, "-m", "lilo.providers.modal.deployment_worker_app"], + check=True, + env=worker_env, + ) + subprocess.run( + [*command, "-m", "lilo.providers.modal.app"], + check=True, + env={ + **os.environ, + MANIFEST_ENV: json.dumps([row.model_dump(mode="json") for row in manifest]), + "LILO_APP_NAME": settings["frontend"], + }, + ) def parser(): @@ -250,12 +183,10 @@ def parser(): management = commands.add_parser("deployment").add_subparsers( dest="action", required=True ) - for name in ("unlock", "retry"): - cmd = management.add_parser(name) - cmd.add_argument("--frontend", required=True) - cmd.add_argument("--env") - if name == "retry": - cmd.add_argument("definition_id") + retry = management.add_parser("retry") + retry.add_argument("--frontend", required=True) + retry.add_argument("--env") + retry.add_argument("definition_id") return result @@ -304,15 +235,9 @@ def main(argv=None): refresh_inference=args.refresh_inference, ) else: - if args.action == "unlock": - registry = modal.Dict.from_name( - f"{args.frontend}-yaml-deployments", environment_name=args.env - ) - registry.pop("apply_lock", None) - else: - modal.Function.from_name( - args.frontend, "clear_deployment_failure", environment_name=args.env - ).remote(args.definition_id) + modal.Function.from_name( + args.frontend, "clear_deployment_failure", environment_name=args.env + ).remote(args.definition_id) except (ValueError, OSError, subprocess.CalledProcessError) as exc: cli.exit(1, f"{exc}\n") diff --git a/src/lilo/deployments.py b/src/lilo/deployments.py index 6de312f..a3c5b55 100644 --- a/src/lilo/deployments.py +++ b/src/lilo/deployments.py @@ -20,7 +20,7 @@ LIFECYCLE_FIELDS = ("session_idle_timeout_s", "pool_idle_timeout_s", "sweep_interval_s") PLATFORM_DEFAULTS = { - "frontend": "lilo-yaml", + "frontend": "lilo", "modal": {"environment": None, "region": "us-west"}, "secrets": { "api": "lilo-api", @@ -213,7 +213,7 @@ def inference_app_name(self) -> str: @property def definition_id(self) -> str: - return f"yaml_{self.spec.name}_{self.generation[:16]}" + return f"deployment_{self.spec.name}_{self.generation[:16]}" @property def asset_path(self) -> str: diff --git a/src/lilo/providers/modal/app.py b/src/lilo/providers/modal/app.py index 07a456b..bd96cd5 100644 --- a/src/lilo/providers/modal/app.py +++ b/src/lilo/providers/modal/app.py @@ -106,6 +106,13 @@ async def _delete_checkpoint(uri: str) -> None: .env(TRAINER_DEPLOYMENT_ENV) .add_local_python_source("lilo", ignore=ignore_config_source) ) + + +@app.function(image=image) +def deployment_manifest(): + return [row.model_dump(mode="json") for row in manifest_from_env()] + + model_assets = modal.Volume.from_name( SETTINGS.platform["storage"]["assets"], create_if_missing=True, diff --git a/src/lilo/providers/modal/deployment_records.py b/src/lilo/providers/modal/deployment_records.py index a9fedd7..c44d1fc 100644 --- a/src/lilo/providers/modal/deployment_records.py +++ b/src/lilo/providers/modal/deployment_records.py @@ -23,6 +23,18 @@ def manifest_from_env(): return [DeploymentRecord.model_validate(row) for row in rows] +def deployed_manifest(frontend, environment=None): + """Read the configuration carried by the currently deployed frontend.""" + function = modal.Function.from_name( + frontend, "deployment_manifest", environment_name=environment + ) + try: + function.hydrate() + except modal.exception.NotFoundError: + return [] + return function.remote() + + def pool_deployment(definition_id): for row in manifest_from_env(): if row.definition_id == definition_id: diff --git a/src/lilo/providers/modal/lora_pool.py b/src/lilo/providers/modal/lora_pool.py index 034f3ac..299649c 100644 --- a/src/lilo/providers/modal/lora_pool.py +++ b/src/lilo/providers/modal/lora_pool.py @@ -106,6 +106,6 @@ def stop_pool(spec: LoraPoolSpec) -> None: def _definition_revision(definition_id: str) -> str: - if not definition_id.startswith("yaml_"): + if not definition_id.startswith("deployment_"): raise ValueError(f"expected a configured deployment id: {definition_id}") return definition_id.rsplit("_", 1)[-1] diff --git a/tests/providers/test_checkpoint_storage.py b/tests/providers/test_checkpoint_storage.py index 77dfb72..1e290a0 100644 --- a/tests/providers/test_checkpoint_storage.py +++ b/tests/providers/test_checkpoint_storage.py @@ -33,7 +33,7 @@ def test_checkpoint_storage_creates_one_v2_volume_without_live_lookup() -> None: assert storage["CHECKPOINT_ROOT"] == "/checkpoints" -def test_yaml_definitions_use_configured_checkpoint_storage(): +def test_deployments_use_configured_checkpoint_storage(): from lilo.providers.modal.deployment_apps import volumes_for from lilo.backends.deployment import backend_config diff --git a/tests/providers/test_deployment_apps.py b/tests/providers/test_deployment_apps.py index a73d9ce..55dfa9f 100644 --- a/tests/providers/test_deployment_apps.py +++ b/tests/providers/test_deployment_apps.py @@ -166,7 +166,7 @@ def test_pool_lookup_uses_recorded_generation(monkeypatch): monkeypatch.setenv(deployment_records.MANIFEST_ENV, json.dumps([row.model_dump()])) saved = deployment_records.pool_deployment(row.definition_id) assert saved.generation == row.generation - assert deployment_records.pool_deployment("yaml_missing_123") is None + assert deployment_records.pool_deployment("deployment_missing_123") is None def test_startup_failure_is_visible_and_blocks_new_spawns(monkeypatch): @@ -210,7 +210,8 @@ def test_real_modal_app_constructs_from_manifest_without_legacy_catalog(monkeypa deployment_apps.image_for = lambda backend: modal.Image.debian_slim() app = importlib.import_module('lilo.providers.modal.app') assert len(app.DEFINITIONS) == 1 -assert app.APP_NAME == 'lilo-yaml' +assert app.APP_NAME == 'lilo' +assert app.deployment_manifest.local() == [d.model_dump(mode='json') for d in app.manifest_from_env()] assert app.DEFINITIONS[0].ENGINE_FUNCTION is not None assert not any(name.startswith('lilo.providers.modal.definitions.') for name in sys.modules) print('constructed') @@ -243,7 +244,7 @@ def test_admission_changes_preserve_serialized_trainer(builders): @pytest.mark.parametrize("kind", ["lora", "full"]) -def test_pool_launch_uses_only_generic_yaml_app(monkeypatch, kind): +def test_pool_launch_uses_only_generic_deployment_app(monkeypatch, kind): from lilo.providers.modal import fft_pool, lora_pool row = deployment("qwen35-9b-lora-16k" if kind == "lora" else "qwen35-4b-fft-64k") diff --git a/tests/providers/test_deployment_e2e_helper.py b/tests/providers/test_deployment_e2e_helper.py index 2e59e5c..2b84993 100644 --- a/tests/providers/test_deployment_e2e_helper.py +++ b/tests/providers/test_deployment_e2e_helper.py @@ -1,29 +1,29 @@ from pathlib import Path import runpy -from types import SimpleNamespace -import modal import pytest from lilo.deployments import load, config_path, DeploymentRecord +from lilo.providers.modal import deployment_records @pytest.mark.parametrize("preset", ["qwen35-9b-lora-16k", "qwen35-4b-fft-64k"]) def test_e2e_helper_reads_active_deployed_configuration(monkeypatch, preset): - helper = runpy.run_path( - str(Path(__file__).parents[2] / "scripts/e2e_engine_definition.py") - ) row = DeploymentRecord.create(load(config_path(preset)), revision="a" * 40) retired = row.model_copy(update={"active": False, "generation": "b" * 64}) - registry = SimpleNamespace( - get=lambda *args: [retired.model_dump(), row.model_dump()] + monkeypatch.setattr( + deployment_records, + "deployed_manifest", + lambda frontend: [retired.model_dump(), row.model_dump()], + ) + helper = runpy.run_path( + str(Path(__file__).parents[2] / "scripts/e2e_engine_definition.py") ) - monkeypatch.setattr(modal.Dict, "from_name", lambda name: registry) definition, mode = helper["_definition"]("test-frontend", row.spec.name) assert definition.DEFINITION_ID == row.definition_id assert definition.MAX_CONTEXT_LENGTH == row.spec.max_context_length assert definition.GPUS == row.spec.trainer_gpus_per_node assert mode == row.spec.parameterization assert definition.MAX_TOKENS_PER_MICROBATCH > 0 - with pytest.raises(ValueError, match="one active YAML"): + with pytest.raises(ValueError, match="one active configuration"): helper["_definition"]("test-frontend", "missing") diff --git a/tests/providers/test_lora_pool.py b/tests/providers/test_lora_pool.py index 8d303df..099800a 100644 --- a/tests/providers/test_lora_pool.py +++ b/tests/providers/test_lora_pool.py @@ -2,7 +2,7 @@ def test_lora_pool_is_shared_by_every_adapter_for_definition() -> None: - first = LoraPoolSpec("yaml_example_0123456789abcdef") + first = LoraPoolSpec("deployment_example_0123456789abcdef") second = LoraPoolSpec.from_dict(first.as_dict()) assert first == second @@ -35,8 +35,8 @@ def test_stop_already_stopped_lora_pool_succeeds_but_real_failure_propagates( def test_pool_revision_comes_from_resolved_generation(): - first = LoraPoolSpec("yaml_example_0123456789abcdef") - changed = LoraPoolSpec("yaml_example_fedcba9876543210") + first = LoraPoolSpec("deployment_example_0123456789abcdef") + changed = LoraPoolSpec("deployment_example_fedcba9876543210") assert first.revision == "0123456789abcdef" assert first.app_name != changed.app_name diff --git a/tests/test_deployment_cli.py b/tests/test_deployment_cli.py index c9949b3..e254c06 100644 --- a/tests/test_deployment_cli.py +++ b/tests/test_deployment_cli.py @@ -1,93 +1,141 @@ from copy import deepcopy import json import subprocess +from types import SimpleNamespace +from unittest.mock import Mock import modal import pytest from lilo import deployment_cli as cli from lilo.deployments import load, config_path, DeploymentRecord -from lilo.providers.modal.deployment_records import MANIFEST_ENV - - -class Registry(dict): - def put(self, key, value, skip_if_exists=False): - if skip_if_exists and key in self: - return False - self[key] = value - return True +from lilo.providers.modal.deployment_records import MANIFEST_ENV, deployed_manifest def deployment(): return DeploymentRecord.create( - load(config_path("qwen35-9b-lora-16k")), - revision="a" * 40, + load(config_path("qwen35-9b-lora-16k")), revision="a" * 40 ) @pytest.fixture -def registry(monkeypatch): - value = Registry() - monkeypatch.setattr(modal.Dict, "from_name", lambda *args, **kwargs: value) - - def missing(*args, **kwargs): - raise modal.exception.NotFoundError("not deployed") +def deployed(monkeypatch): + state = SimpleNamespace(apps=set(), manifest=[], calls=[], fail=None) - monkeypatch.setattr(modal.App, "lookup", missing) - return value + def read_manifest(frontend, environment): + return state.manifest + def lookup(name, **kwargs): + if name not in state.apps: + raise modal.exception.NotFoundError("not deployed") -def test_deploy_is_serialized_and_commits_only_after_success(registry, monkeypatch): - row = deployment() - seen = [] - - def run(command, **kwargs): - assert "apply_lock" in registry - assert "pending" in registry - assert "manifest" not in registry - if "lilo.providers.modal.app" in command: - seen.extend(json.loads(kwargs["env"][MANIFEST_ENV])) + def run(command, *, env, check): + role = env.get("LILO_WORKER_ROLE", "frontend") + state.calls.append(role) + if state.fail == role: + raise subprocess.CalledProcessError(1, command) + if role == "frontend": + state.manifest = json.loads(env[MANIFEST_ENV]) else: - assert "lilo.providers.modal.deployment_worker_app" in command + row = DeploymentRecord.model_validate_json(env["LILO_WORKER_DEPLOYMENT"]) + state.apps.add( + row.trainer_app_name if role == "trainer" else row.inference_app_name + ) + monkeypatch.setattr(cli, "deployed_manifest", read_manifest) + monkeypatch.setattr(modal.App, "lookup", lookup) monkeypatch.setattr(subprocess, "run", run) - cli.deploy([row]) - assert seen == registry["manifest"] - assert "pending" not in registry and "apply_lock" not in registry - registry["apply_lock"] = "other-operator" - with pytest.raises(ValueError, match="An apply owns"): - cli.deploy([row]) - assert registry["apply_lock"] == "other-operator" + monkeypatch.setattr( + modal.Dict, + "from_name", + lambda *a, **k: pytest.fail("deployment must not use a Dict registry"), + ) + return state -def test_failed_apply_keeps_pending_generations_for_next_attempt(registry, monkeypatch): +def test_deploy_reuses_modal_apps_and_retains_previous_config(deployed): row = deployment() + cli.deploy([row]) + assert deployed.calls == ["trainer", "inference", "frontend"] + deployed.calls.clear() + cli.deploy([row]) + assert deployed.calls == ["frontend"] + + spec = deepcopy(row.spec) + spec.inference_max_replicas = 6 + changed = DeploymentRecord.create(spec, revision="a" * 40) + deployed.calls.clear() + cli.deploy([changed]) + assert deployed.calls == ["inference", "frontend"] + assert [(r["generation"], r["active"]) for r in deployed.manifest] == [ + (changed.generation, True), + (row.generation, False), + ] + + # Modal's actual state wins over a frontend record that mentions an old app. + deployed.apps.remove(changed.inference_app_name) + deployed.calls.clear() + cli.deploy([changed]) + assert deployed.calls == ["inference", "frontend"] - def fail(*args, **kwargs): - raise subprocess.CalledProcessError(1, args[0]) - monkeypatch.setattr(subprocess, "run", fail) +@pytest.mark.parametrize("failure", ["inference", "frontend"]) +def test_retry_discovers_completed_workers_without_pending_records(deployed, failure): + row = deployment() + deployed.fail = failure with pytest.raises(subprocess.CalledProcessError): cli.deploy([row]) - assert registry["pending"][0]["generation"] == row.generation - assert "manifest" not in registry and "apply_lock" not in registry - new_spec = deepcopy(row.spec) - new_spec.trainer_max_instances = 2 - new = DeploymentRecord.create(new_spec, revision="a" * 40) - monkeypatch.setattr(subprocess, "run", lambda *args, **kwargs: None) - cli.deploy([new]) - assert [(r["generation"], r["active"]) for r in registry["manifest"]] == [ - (new.generation, True), - ] + assert deployed.manifest == [] + assert row.trainer_app_name in deployed.apps + deployed.calls.clear() + deployed.fail = None + cli.deploy([row]) + assert deployed.calls == ( + ["inference", "frontend"] if failure == "inference" else ["frontend"] + ) -def test_refuse_overwriting_legacy_frontend(registry, monkeypatch): - monkeypatch.setattr(modal.App, "lookup", lambda *args, **kwargs: object()) - with pytest.raises( - ValueError, match="already exists without a deployment registry" - ): - cli.deploy([deployment()]) - assert "pending" not in registry +def test_refresh_keeps_other_workers_and_is_remembered_by_frontend(deployed): + miles = deployment() + fft = DeploymentRecord.create( + load(config_path("qwen35-4b-fft-64k")), revision="a" * 40 + ) + cli.deploy([miles, fft]) + deployed.calls.clear() + cli.deploy([miles, fft], refresh_trainers=[miles.spec.name]) + assert deployed.calls == ["trainer", "frontend"] + updated = deployed.manifest[0] + assert updated["trainer_release"] != "initial" + assert updated["inference_release"] == "initial" + deployed.calls.clear() + cli.deploy([miles, fft]) + assert deployed.calls == ["frontend"] + assert deployed.manifest[0] == updated + + +def test_existing_frontend_without_deployment_metadata_can_be_updated(deployed): + row = deployment() + deployed.apps.add(row.platform["frontend"]) + cli.deploy([row]) + assert deployed.manifest[0]["generation"] == row.generation + + +def test_read_manifest_from_deployed_function(monkeypatch): + function = SimpleNamespace( + hydrate=Mock(), remote=Mock(return_value=[{"configuration": "saved"}]) + ) + lookup = Mock(return_value=function) + monkeypatch.setattr(modal.Function, "from_name", lookup) + assert deployed_manifest("my-app", "dev") == [{"configuration": "saved"}] + lookup.assert_called_once_with( + "my-app", "deployment_manifest", environment_name="dev" + ) + function.hydrate.side_effect = modal.exception.NotFoundError("no function") + assert deployed_manifest("my-app", "dev") == [] + function.hydrate.side_effect = None + function.remote.side_effect = RuntimeError("frontend failed") + with pytest.raises(RuntimeError, match="frontend failed"): + deployed_manifest("my-app", "dev") def test_validate_never_resolves_or_deploys(monkeypatch, capsys): @@ -150,136 +198,6 @@ def test_worker_source_mount_excludes_authoring_configs(): assert not ignore_config_source(Path("backends/miles_config.py")) -def test_only_changed_worker_is_deployed(registry, monkeypatch): - row = deployment() - calls = [] - - def run(command, **kwargs): - if "lilo.providers.modal.deployment_worker_app" in command: - calls.append(kwargs["env"]["LILO_WORKER_ROLE"]) - else: - calls.append("frontend") - - monkeypatch.setattr(subprocess, "run", run) - cli.deploy([row]) - assert calls == ["trainer", "inference", "frontend"] - - calls.clear() - cli.deploy([row]) - assert calls == ["frontend"] - - changed = deepcopy(row.spec) - changed.inference_max_replicas = 6 - new = DeploymentRecord.create(changed, revision="a" * 40) - calls.clear() - cli.deploy([new]) - assert calls == ["inference", "frontend"] - assert new.trainer_app_name == row.trainer_app_name - assert registry["manifest"][1]["active"] is False - - newest = DeploymentRecord.create(changed, revision="a" * 40) - calls.clear() - cli.deploy([newest], refresh_trainers=[newest.spec.name]) - assert calls == ["trainer", "frontend"] - assert newest.inference_app_name == new.inference_app_name - - -def test_backend_update_does_not_redeploy_other_models(registry, monkeypatch): - miles = deployment() - fft = DeploymentRecord.create( - load(config_path("qwen35-4b-fft-64k")), - revision="a" * 40, - ) - calls = [] - - def run(command, **kwargs): - if "LILO_WORKER_DEPLOYMENT" in kwargs["env"]: - calls.append( - ( - kwargs["env"]["LILO_WORKER_ROLE"], - json.loads(kwargs["env"]["LILO_WORKER_DEPLOYMENT"])["spec"]["name"], - ) - ) - - monkeypatch.setattr(subprocess, "run", run) - cli.deploy([miles, fft]) - calls.clear() - spec = deepcopy(miles.spec) - cli.deploy( - [DeploymentRecord.create(spec, revision="a" * 40), fft], - refresh_trainers=[miles.spec.name], - ) - assert calls == [("trainer", miles.spec.name)] - - -def test_retry_preserves_successfully_deployed_workers(registry, monkeypatch): - row = deployment() - calls = [] - - def fail_frontend(command, **kwargs): - calls.append(kwargs["env"].get("LILO_WORKER_ROLE", "frontend")) - if "lilo.providers.modal.app" in command: - raise subprocess.CalledProcessError(1, command) - - monkeypatch.setattr(subprocess, "run", fail_frontend) - with pytest.raises(subprocess.CalledProcessError): - cli.deploy([row]) - assert len(registry["worker_apps"]) == 2 - calls.clear() - monkeypatch.setattr( - subprocess, - "run", - lambda command, **kwargs: calls.append( - kwargs["env"].get("LILO_WORKER_ROLE", "frontend") - ), - ) - cli.deploy([row]) - assert calls == ["frontend"] - - -def test_recover_worker_deployed_before_registry_write(registry, monkeypatch): - row = deployment() - calls = [] - - def lookup(name, **kwargs): - if name == row.trainer_app_name: - return object() - raise modal.exception.NotFoundError("not deployed") - - monkeypatch.setattr(modal.App, "lookup", lookup) - monkeypatch.setattr( - subprocess, - "run", - lambda command, **kwargs: calls.append( - kwargs["env"].get("LILO_WORKER_ROLE", "frontend") - ), - ) - cli.deploy([row]) - assert calls == ["inference", "frontend"] - - -def test_worker_refresh_is_retained_without_editing_config(registry, monkeypatch): - row = deployment() - calls = [] - monkeypatch.setattr( - subprocess, - "run", - lambda command, **kwargs: calls.append( - kwargs["env"].get("LILO_WORKER_ROLE", "frontend") - ), - ) - cli.deploy([row]) - cli.deploy([row], refresh_inference=[row.spec.name]) - active = DeploymentRecord.model_validate(registry["manifest"][0]) - assert active.inference_release != "initial" - assert active.trainer_release == "initial" - assert vars(active.spec) == vars(row.spec) - calls.clear() - cli.deploy([row]) - assert calls == ["frontend"] - assert registry["manifest"][0]["inference_release"] == active.inference_release - - def test_deploy_command_owns_platform_settings(monkeypatch): seen = {} From d571aff6c3849d56d4b367d774ed58faac0948be Mon Sep 17 00:00:00 2001 From: kailash Date: Mon, 28 Sep 2026 00:04:56 +0000 Subject: [PATCH 26/27] Simplify deployment configs and use named worker apps --- README.md | 4 +- docs/deployment-configs.md | 51 +- docs/deployment-validation.md | 18 +- scripts/definition_smoke.py | 31 +- scripts/deployment_smoke.py | 12 +- scripts/e2e_engine_definition.py | 16 +- .../megatron_runtime/fft/checkpoint.py | 2 - src/lilo/backends/miles_config.py | 2 - src/lilo/backends/miles_lora.py | 7 - src/lilo/configuration.py | 1 - src/lilo/control_plane/deployments.py | 20 +- src/lilo/control_plane/http.py | 28 +- src/lilo/deployment_cli.py | 222 +-- src/lilo/deployments.py | 303 ++- src/lilo/providers/modal/app.py | 66 +- src/lilo/providers/modal/deployment_apps.py | 220 +-- .../providers/modal/deployment_configs.py | 58 + .../providers/modal/deployment_pool_app.py | 14 +- .../providers/modal/deployment_records.py | 52 - .../providers/modal/deployment_worker_app.py | 17 - src/lilo/providers/modal/fft_pool.py | 33 +- src/lilo/providers/modal/lora_pool.py | 49 +- src/lilo/providers/modal/miles_image.py | 6 +- src/lilo/providers/modal/miles_revision.py | 34 - src/lilo/providers/modal/scoped.py | 9 +- tests/backends/test_megatron_fft.py | 18 +- tests/backends/test_miles.py | 19 - tests/conftest.py | 3 +- tests/control_plane/test_http.py | 31 +- tests/control_plane/test_sdk_e2e.py | 18 +- tests/providers/conftest.py | 10 +- tests/providers/test_checkpoint_storage.py | 15 +- tests/providers/test_definition_registry.py | 8 +- tests/providers/test_deployment_apps.py | 124 +- tests/providers/test_deployment_e2e_helper.py | 23 +- tests/providers/test_deployment_presets.py | 29 +- tests/providers/test_lora_pool.py | 21 +- tests/providers/test_miles_revision.py | 55 - tests/providers/test_modal_app.py | 41 +- tests/providers/test_sglang_entrypoint.py | 6 +- tests/test_deployment_cli.py | 180 +- tests/test_deployments.py | 108 +- tests/test_system.py | 8 +- uv.lock | 1622 ++++++++--------- 44 files changed, 1635 insertions(+), 1979 deletions(-) create mode 100644 src/lilo/providers/modal/deployment_configs.py delete mode 100644 src/lilo/providers/modal/deployment_records.py delete mode 100644 src/lilo/providers/modal/deployment_worker_app.py delete mode 100644 src/lilo/providers/modal/miles_revision.py delete mode 100644 tests/providers/test_miles_revision.py diff --git a/README.md b/README.md index 33658e0..409ba3b 100644 --- a/README.md +++ b/README.md @@ -139,9 +139,9 @@ uv run lilo config validate deployment.py uv run lilo deploy deployment.py ``` -This deploys the shared app and prints its `server` URL. Add more Python config files to the same command to serve more configurations. Always supply the complete active set. Pin model revisions and `LILO_MILES_COMMIT` for repeatable applies; see [Python deployment configs](docs/deployment-configs.md). +This deploys the shared app and prints its `server` URL. Add more Python config files to the same command to serve more recipes. Always supply the complete current set. The Miles commit is pinned in `miles_image.py`; see [Python deployment configs](docs/deployment-configs.md). -From a repository checkout, maintain the list in `scripts/deploy_models.sh` and run that script. `lilo deploy` supplies the saved configuration to Modal; importing the shared app directly without a manifest is no longer a deployment entrypoint. +From a repository checkout, maintain the list in `scripts/deploy_models.sh` and run that script. `lilo deploy` supplies the current configs and frontend platform settings to Modal. Deploying the server doesn't allocate any GPUs; rather, this allocation for both the training and sampling sides are done on demand. See [cold starts and capacity configuration](docs/full-fine-tunes.md#performance-and-behavior-considerations) before running a larger workload. diff --git a/docs/deployment-configs.md b/docs/deployment-configs.md index 313e0d2..ec044e2 100644 --- a/docs/deployment-configs.md +++ b/docs/deployment-configs.md @@ -30,7 +30,7 @@ class Config(BaseConfig): config = Config() ``` -See the [9B LoRA recipe](../src/lilo/configs/qwen35_9b_lora_16k.py) and [4B FFT recipe](../src/lilo/configs/qwen35_4b_fft_64k.py) for complete examples. Deployment resolves the model's `main` revision to an exact commit; set `revision` only when you want a different revision. +See the [9B LoRA recipe](../src/lilo/configs/qwen35_9b_lora_16k.py) and [4B FFT recipe](../src/lilo/configs/qwen35_4b_fft_64k.py) for complete examples. Weights are downloaded to `/assets/`. ## Variants @@ -66,29 +66,26 @@ Modal and backend libraries validate their own options. Lilo checks integration Miles argument conversion lives in [miles_arguments.py](../src/lilo/backends/miles_arguments.py). SGLang receives `ServerArgs(**settings)` in its [worker entrypoint](../src/lilo/inference/sglang.py). Backend libraries validate native options when workers start; frontend config imports remain CPU-only. -## Resolve and launch +## Prepare and launch ~~~text -load(config.py) → config: BaseConfig - → resolve model commit - → resolve_backend_settings(config, asset_path) - → save DeploymentRecord with trainer_settings and inference_settings - → deploy independent worker apps - → update frontend references +load(config.py) → recipe: BaseConfig + → DeploymentConfig.create(recipe) + → deploy named trainer and inference apps if they are missing or changed + → deploy the frontend with the current config set ~~~ -The launcher consumes the saved settings. It does not reparse backend configuration or add another set of backend defaults. JSON decoding in a worker reconstructs the saved record, without importing the author's config file. +The launcher consumes the saved settings. It does not reparse backend configuration or add another set of backend defaults. JSON decoding in a trainer reconstructs the saved config, without importing the author's config file. | File | Responsibility | | --- | --- | | [configuration.py](../src/lilo/configuration.py) | BaseConfig defaults, inheritance and overrides | -| [deployments.py](../src/lilo/deployments.py) | Python object loader, resolved records and config hashes | -| [backends/deployment.py](../src/lilo/backends/deployment.py) | Resolve backend settings before launch | +| [deployments.py](../src/lilo/deployments.py) | Python recipe loader and DeploymentConfig | +| [backends/deployment.py](../src/lilo/backends/deployment.py) | Backend settings used when creating a DeploymentConfig | | [megatron_runtime/common/settings.py](../src/lilo/backends/megatron_runtime/common/settings.py) | Shared Megatron ownership rules and constructor dictionaries | -| [deployment_cli.py](../src/lilo/deployment_cli.py) | Model lookup, worker release selection and deploy ordering | -| [deployment_apps.py](../src/lilo/providers/modal/deployment_apps.py) | Resource declarations and launch using resolved settings | -| [deployment_records.py](../src/lilo/providers/modal/deployment_records.py) | Read saved records and route pool provisioning | -| [deployment_worker_app.py](../src/lilo/providers/modal/deployment_worker_app.py) | Deploy one trainer or inference provisioner | +| [deployment_apps.py](../src/lilo/providers/modal/deployment_apps.py) | Trainer, inference, and pool app builders | +| [deployment_configs.py](../src/lilo/providers/modal/deployment_configs.py) | Read saved configs and frontend platform settings | +| [deployment_cli.py](../src/lilo/deployment_cli.py) | Deploy trainer and inference apps by calling the builders, then deploy the frontend | ## Multi-node Miles @@ -104,9 +101,9 @@ lilo config validate my_model.py lilo deploy my_model.py ~~~ -Validation resolves backend settings and checks Lilo integration constraints without loading GPU libraries or provisioning compute. Backend option support and GPU memory capacity still require worker startup. +Validation loads the recipes and checks that they can share one frontend. Backend integration settings are checked when creating a DeploymentConfig at deploy time. Native backend options are checked at engine startup. -The checked-in [deploy_models.sh](../scripts/deploy_models.sh) lists the complete active config set. Add a config path there, then run it. The deployment command owns frontend selection and worker-code updates: +The checked-in [deploy_models.sh](../scripts/deploy_models.sh) lists the complete active config set. Add a config path there, then run it. The deployment command owns frontend selection and trainer/inference updates: ~~~bash ./scripts/deploy_models.sh --app my-lilo --env dev --region us-west @@ -116,22 +113,16 @@ The checked-in [deploy_models.sh](../scripts/deploy_models.sh) lists the complet These settings are not model-config fields. Secret/volume names come from provider defaults. Credentials remain in Modal secrets. -One frontend serves all models through Tinker’s `base_model`. When recipes share a model, the first matching recipe in the `lilo deploy` argument list is used. Training also matches the requested LoRA/FFT mode; base sampling uses the first recipe regardless of training mode. To select a specific recipe, pass its definition ID from `/api/v1/lilo/deployments` as `base_model`. All supplied definitions are listed, including retained generations. Current recipes precede retained ones; existing jobs keep their saved definition IDs. +One frontend serves all models through Tinker’s `base_model`. When recipes share a model, the first matching recipe in the `lilo deploy` argument list is used. Training also matches the requested LoRA/FFT mode; base sampling uses the first recipe regardless of training mode. To select a specific recipe, pass its name from `/api/v1/lilo/deployments` as `base_model`. The frontend lists the current config set only. -## Hashes and update isolation +## Update isolation -Hashes identify settings, not compatibility with a source checkout: +Trainer apps are named `lilo-trainer-{name}` and inference apps `lilo-inference-{name}`. An inference-only change reuses the trainer app. A trainer-only change reuses the inference app. Source-only upgrades use `--refresh-trainer` / `--refresh-inference`. -- generation identifies the saved configuration, resolved backend settings, platform settings and code releases. -- trainer_hash includes trainer/model/platform settings, resolved trainer settings, Miles commit and its recorded code release. -- inference_hash includes inference/model/platform settings, resolved inference settings and its recorded code release. +The frontend carries the current configs and exposes them through `deployed_configs`. The CLI reads that function to compare settings, and checks trainer and inference app existence directly with Modal. There is no retained history of earlier configs. Run deploy commands sequentially. -SHA-256 hashes sorted JSON. Trainer and inference app names use the first 24 hex characters; definition IDs use the first 16 characters of generation. There is no whole-source fingerprint. +`--app`, `--env`, `--region`, secrets, and storage names are frontend/platform settings. They are not copied onto each recipe. -An inference-only change reuses the trainer app. A trainer batch-setting change reuses the inference provisioner. Changes to adapter rank/targets update both. Code-only upgrades use the refresh flags; later ordinary deploys retain those releases. +LoRA pools are named `lilo-lora-{name}`. FFT pools are named `lilo-fft-{name}-{model_id}-latest` or `lilo-fft-{name}-{model_id}-v{version}`. -The frontend carries its resolved configurations and exposes them through the Modal `deployment_manifest` function. The CLI reads that function to retain existing definitions and worker releases, and checks worker existence directly with Modal. There is no separate deployment Dict, pending journal, or deployment lock. Run deploy commands sequentially. - -Old jobs keep their recorded worker apps. Deployment retries discover and reuse workers already created successfully. Trainer limits are enforced per saved definition; old and new definitions may use their configured capacity simultaneously while old jobs finish. Pools remain associated with full job configurations, so new jobs may get separate pools even if they share an inference provisioner. - -Old worker app definitions are retained; automatic cleanup is not implemented. Changes to the frontend/worker protocol still need deliberate compatibility handling. Earlier draft config formats require migration or a fresh frontend. See [validation history](deployment-validation.md) for the distinction between current CPU checks and historical GPU runs. +See [validation history](deployment-validation.md) for the distinction between current CPU checks and historical GPU runs. diff --git a/docs/deployment-validation.md b/docs/deployment-validation.md index bb03f14..18338dd 100644 --- a/docs/deployment-validation.md +++ b/docs/deployment-validation.md @@ -2,17 +2,15 @@ These checks exercise PR #55. The current implementation deploys trainers and inference provisioners independently; the shared frontend references them by app name. Historical sections record validation of previous implementations. -## Current: flat BaseConfig recipes +## Current: recipe plus backend settings -Recipes use one plain `BaseConfig` with untyped attributes such as `trainer_gpu` and `inference_max_replicas`. Backend options live in `megatron_cfg`, `miles_cfg`, and `sglang_cfg` dictionaries. The `Trainer`, `Inference`, and `Deployment` config wrappers, constrained type aliases, and Pydantic recipe validation are removed. Each consuming component validates its own options; Lilo retains cross-component integration checks. +A recipe is a `BaseConfig` instance. `DeploymentConfig` is that recipe plus trainer and inference settings. Definition identity is the recipe `name`. Trainer and inference apps are `lilo-trainer-{name}` and `lilo-inference-{name}`. LoRA pools are `lilo-lora-{name}`; FFT pools are `lilo-fft-{name}-{model_id}-latest` or `lilo-fft-{name}-{model_id}-v{version}`. -Variants inherit class attributes and use dotted `overrides` for nested backend options. Mutable settings are copied per instance. Workers consume saved resolved backend settings. Routing uses deployment order or explicit definition IDs, without visibility or default flags. +`--app`, `--env`, `--region`, secrets, and storage live on the frontend, not on each recipe. The frontend is parameterized by `LILO_DEPLOYMENT_CONFIGS` and `LILO_PLATFORM`. `lilo config validate` loads recipes and checks shared lifecycle fields; backend settings are built once when creating a `DeploymentConfig` at deploy time. -Validation: **642 CPU tests passed, 1 skipped**, including **41 focused deployment/provider tests**. Coverage includes inherited overrides, mutable option isolation, backend passthrough, saved-record JSON round trips, worker resources, update isolation and deployment recovery. The CLI-generated recipe loaded and validated successfully. Ruff and whitespace checks passed. +Isolation is field equality plus named-app lookup. An unchanged trainer or inference app is skipped. There is no config hash, worker hash, retained release, or inactive-generation list. After a recipe changes, the new config is the current one. -Deployment uses Modal app lookups and the frontend’s `deployment_manifest` function. The separate Dict registry, pending records, deployment lock, unlock command and registry-less-frontend rejection are removed. Tests cover worker reuse, retries after partial deployment, missing worker recreation, retained configurations and code refresh. The default frontend is `lilo`; definition IDs use `deployment_`. - -All 15 recipes preserve compute, model, lifecycle and resolved trainer/inference settings. The flat saved config layout changes deployment hashes. Existing draft manifests need regeneration or explicit migration. No deployments or GPU jobs were modified; GPU backend startup has not been rerun for this revision. +Validation: **271 focused deployment/control-plane tests passed**. Coverage includes inherited overrides, mutable option isolation, backend passthrough, JSON round trips, named-app reuse, retries after partial deployment, missing-app recreation, and code refresh. No GPU jobs were redeployed for this revision. ## Historical validation @@ -81,11 +79,11 @@ Validation: **597 CPU tests passed, 1 skipped**; Ruff passed for changed Python This migration has not been redeployed or GPU-tested. The GPU results below describe the earlier YAML-based source at `13d2a31`, before the subsequent CPU-only refactors. Earlier cleanup counts below are historical. -## Deployment record simplification +## Deployment config simplification -The saved manifest wrapper is now named `DeploymentRecord`. Its factory copies the parsed specification and computes the configuration hash without a dictionary-to-model round trip. Model tag resolution and validation of the Hugging Face result stay in the CLI. The standalone `resolve()` function was removed. Serialized record fields and hash format are unchanged. +The assembled wrapper is now named `DeploymentConfig`. Its factory copies the parsed specification and computes the configuration hash without a dictionary-to-model round trip. Model tag resolution and validation of the Hugging Face result stay in the CLI. The standalone `resolve()` function was removed. Serialized config fields and hash format are unchanged. -Validation: **595 CPU tests passed, 1 skipped**; Ruff passed for changed Python files and whitespace checks passed. Added coverage verifies model tag pinning, no lookup for an existing commit, record isolation from subsequent mutations, JSON round trips, and unchanged configuration hashes. No apps were redeployed. +Validation: **595 CPU tests passed, 1 skipped**; Ruff passed for changed Python files and whitespace checks passed. Added coverage verifies model tag pinning, no lookup for an existing commit, config isolation from subsequent mutations, JSON round trips, and unchanged configuration hashes. No apps were redeployed. ## YAML loader simplification diff --git a/scripts/definition_smoke.py b/scripts/definition_smoke.py index 4a1242c..5dd5caf 100644 --- a/scripts/definition_smoke.py +++ b/scripts/definition_smoke.py @@ -15,7 +15,7 @@ --definition-id qwen3_5_9b_miles_lora_16k \ --definition-id qwen3_5_4b_full_64k -IDs come from the deployed manifest selected by --app / --env. Use --list to +IDs come from the deployed records selected by --app / --env. Use --list to see active configurations. No local model catalog or GPU worker imports are needed. """ @@ -37,9 +37,8 @@ from lilo.backends.miles_config import lora_target_flags from lilo.client import create_full_training_client -from lilo.deployments import DeploymentRecord -from lilo.providers.modal.deployment_records import deployed_manifest -from lilo.providers.modal.deployment_apps import definition_from_spec +from lilo.deployments import DeploymentConfig +from lilo.providers.modal.deployment_configs import deployed_configs TIMEOUT = 60 * 60 PROMPT = "Question: What is two plus two?\nAnswer:" @@ -166,15 +165,15 @@ def _unload(base_url: str, api_key: str, model_id: str) -> None: def _create_training_client( service: tinker.ServiceClient, definition: Any, rank: int | None ) -> tuple[tinker.TrainingClient, dict[str, Any]]: - definition_id = definition.DEFINITION_ID - if definition.PARAMETERIZATION == "full": + definition_id = definition.definition_id + if definition.parameterization == "full": training = create_full_training_client(service, definition_id) return training, {"parameterization": "full"} train_attn, train_mlp, train_unembed = lora_target_flags( - definition.RESOLVED.trainer_settings["miles"]["target_modules"] + definition.trainer_settings["miles"]["target_modules"] ) if rank is None: - rank = min(16, definition.RESOLVED.trainer_settings["miles"]["max_lora_rank"]) + rank = min(16, definition.trainer_settings["miles"]["max_lora_rank"]) training = service.create_lora_training_client( base_model=definition_id, rank=rank, @@ -199,7 +198,7 @@ def _run_definition( rank: int | None, max_tokens: int, ) -> dict: - definition_id = definition.DEFINITION_ID + definition_id = definition.definition_id report: dict[str, Any] = { "definition_id": definition_id, "status": "running", @@ -298,22 +297,22 @@ def main() -> None: parser.add_argument("--app", default="lilo") parser.add_argument("--env") args = parser.parse_args() - rows = deployed_manifest(args.app, args.env) + rows = deployed_configs(args.app, args.env) definitions = { - row.definition_id: definition_from_spec(row, register_trainer=False) + row.definition_id: row for data in rows - if (row := DeploymentRecord.model_validate(data)).active + if (row := DeploymentConfig.model_validate(data)) } if not definitions: - parser.error("no active deployments in the saved manifest") + parser.error("no deployments in the saved configs") if set(args.definition_ids or []) - definitions.keys(): parser.error("unknown deployment ID; use --list to see deployed configurations") if args.list: for definition in definitions.values(): print( - f"{definition.DEFINITION_ID:40} {definition.PARAMETERIZATION:5} " - f"{definition.RESOLVED.spec.trainer_nodes} nodes x {definition.RESOLVED.spec.trainer_gpu}:{definition.RESOLVED.spec.trainer_gpus_per_node} " - f"ctx={definition.MAX_CONTEXT_LENGTH}" + f"{definition.definition_id:40} {definition.parameterization:5} " + f"{definition.recipe.trainer_nodes} nodes x {definition.recipe.trainer_gpu}:{definition.recipe.trainer_gpus_per_node} " + f"ctx={definition.max_context_length}" ) return if not args.base_url: diff --git a/scripts/deployment_smoke.py b/scripts/deployment_smoke.py index dca378a..fa4fa11 100644 --- a/scripts/deployment_smoke.py +++ b/scripts/deployment_smoke.py @@ -18,7 +18,7 @@ import tinker from tinker import types -from lilo.providers.modal.deployment_records import deployed_manifest +from lilo.providers.modal.deployment_configs import deployed_configs from lilo.client import create_full_training_client from lilo.providers.modal.kv import app_store_name @@ -101,15 +101,15 @@ def run(args): server = modal.Function.from_name(args.frontend, "server") url = server.get_web_url() headers = {"X-API-Key": os.environ["TINKER_API_KEY"]} - rows = deployed_manifest(args.frontend) - row = next(r for r in rows if r["active"] and r["spec"]["name"] == args.name) - definition_id = f"deployment_{args.name}_{row['generation'][:16]}" + rows = deployed_configs(args.frontend) + row = next(r for r in rows if r["recipe"]["name"] == args.name) + definition_id = args.name report = { "frontend": args.frontend, "app_id": app.app_id, "name": args.name, - "generation": row["generation"], - "spec": row["spec"], + "definition_id": definition_id, + "recipe": row["recipe"], "status": "creating", "steps": [], } diff --git a/scripts/e2e_engine_definition.py b/scripts/e2e_engine_definition.py index e39985c..b8d0ba0 100644 --- a/scripts/e2e_engine_definition.py +++ b/scripts/e2e_engine_definition.py @@ -16,26 +16,24 @@ import tinker from tinker import types -from lilo.deployments import DeploymentRecord -from lilo.providers.modal.deployment_records import deployed_manifest +from lilo.deployments import DeploymentConfig +from lilo.providers.modal.deployment_configs import deployed_configs from lilo.backends.deployment import backend_config TIMEOUT = 3 * 60 * 60 def _definition(frontend: str, name: str) -> tuple[Any, str]: - rows = deployed_manifest(frontend) + rows = deployed_configs(frontend) matches = [ - DeploymentRecord.model_validate(row) + DeploymentConfig.model_validate(row) for row in rows - if row["active"] and row["spec"]["name"] == name + if row["recipe"]["name"] == name ] if len(matches) != 1: - raise ValueError( - f"expected one active configuration named {name} in {frontend}" - ) + raise ValueError(f"expected one configuration named {name} in {frontend}") resolved = matches[0] - spec = resolved.spec + spec = resolved.recipe settings = backend_config(spec)[spec.backend] definition = SimpleNamespace( DEFINITION_ID=resolved.definition_id, diff --git a/src/lilo/backends/megatron_runtime/fft/checkpoint.py b/src/lilo/backends/megatron_runtime/fft/checkpoint.py index 68e0ea9..eadf0b9 100644 --- a/src/lilo/backends/megatron_runtime/fft/checkpoint.py +++ b/src/lilo/backends/megatron_runtime/fft/checkpoint.py @@ -91,7 +91,6 @@ def create_fft_checkpoint_metadata( version=CHECKPOINT_FORMAT_VERSION, checkpoint_id=checkpoint_id, base_model=base_model, - base_model_revision=os.environ.get("LILO_BASE_MODEL_REVISION"), world_size=world_size, tensor_model_parallel_size=config.tensor_model_parallel_size, pipeline_model_parallel_size=config.pipeline_model_parallel_size, @@ -293,7 +292,6 @@ def load_fft_training_checkpoint( "format", "version", "base_model", - "base_model_revision", "world_size", "tensor_model_parallel_size", "pipeline_model_parallel_size", diff --git a/src/lilo/backends/miles_config.py b/src/lilo/backends/miles_config.py index 54ae7ad..68e41de 100644 --- a/src/lilo/backends/miles_config.py +++ b/src/lilo/backends/miles_config.py @@ -4,8 +4,6 @@ from pathlib import Path from typing import Any -MILES_REF = "main" - _PEFT_TARGETS = { "linear_qkv": ("q_proj", "k_proj", "v_proj"), "linear_q": ("q_proj",), diff --git a/src/lilo/backends/miles_lora.py b/src/lilo/backends/miles_lora.py index 390d920..6a59f83 100644 --- a/src/lilo/backends/miles_lora.py +++ b/src/lilo/backends/miles_lora.py @@ -284,7 +284,6 @@ def capture_checkpoint( "backend": "miles", "miles_revision": self.runtime.revision, "base_model": self.base_model, - "base_model_revision": os.environ.get("LILO_BASE_MODEL_REVISION"), "engine_definition_id": os.environ.get("LILO_DEFINITION_ID"), "parameterization": {"type": "lora"}, "lora_config": { @@ -593,12 +592,6 @@ def _validate_checkpoint( raise ValueError("unsupported Miles checkpoint format") if metadata.get("base_model") != self.base_model: raise ValueError("checkpoint base model does not match the deployment") - if metadata.get("base_model_revision") != os.environ.get( - "LILO_BASE_MODEL_REVISION" - ): - raise ValueError( - "checkpoint base model revision does not match the deployment" - ) lora = metadata.get("lora_config") or {} if ( int(lora.get("rank", 0)) != state.rank diff --git a/src/lilo/configuration.py b/src/lilo/configuration.py index 4a2b085..32fb097 100644 --- a/src/lilo/configuration.py +++ b/src/lilo/configuration.py @@ -17,7 +17,6 @@ class BaseConfig: model = "" max_context_length = 16384 parameterization = "lora" - revision = "main" backend = "miles" trainer_gpu = "H100" diff --git a/src/lilo/control_plane/deployments.py b/src/lilo/control_plane/deployments.py index 4170440..430ff45 100644 --- a/src/lilo/control_plane/deployments.py +++ b/src/lilo/control_plane/deployments.py @@ -7,25 +7,29 @@ def __init__(self, definitions): def select(self, model, mode=None): candidates = [ - d for d in self.definitions if mode is None or d.PARAMETERIZATION == mode + d + for d in self.definitions + if mode is None or d.parameterization == mode ] - for field in ("DEFINITION_ID", "MODEL_NAME"): - for definition in candidates: - if getattr(definition, field) == model: - return definition + for definition in candidates: + if definition.definition_id == model: + return definition + for definition in candidates: + if definition.model == model: + return definition return None def capabilities(self): selected = {} for definition in self.definitions: selected.setdefault( - (definition.MODEL_NAME, definition.PARAMETERIZATION), definition + (definition.model, definition.parameterization), definition ) contexts = {} for (model, _), definition in selected.items(): contexts[model] = min( - contexts.get(model, definition.MAX_CONTEXT_LENGTH), - definition.MAX_CONTEXT_LENGTH, + contexts.get(model, definition.max_context_length), + definition.max_context_length, ) return [ {"model_name": model, "max_context_length": context} diff --git a/src/lilo/control_plane/http.py b/src/lilo/control_plane/http.py index 0c5f7e3..4501c7b 100644 --- a/src/lilo/control_plane/http.py +++ b/src/lilo/control_plane/http.py @@ -168,13 +168,13 @@ def create_control_plane_app( def definition_for(model_name, parameterization): selected = routes.select(model_name, parameterization) - return selected.DEFINITION_ID if selected else None + return selected.definition_id if selected else None def supports_model(model_name: str) -> bool: return any( - definition.MODEL_NAME == model_name for definition in definitions + definition.model == model_name for definition in definitions ) or any( - definition.DEFINITION_ID == model_name for definition in definitions + definition.definition_id == model_name for definition in definitions ) async def authorize(request: Request) -> None: @@ -272,11 +272,11 @@ async def list_deployments(): return { "deployments": [ { - "name": getattr(d, "DEPLOYMENT_NAME", d.DEFINITION_ID), - "generation": d.DEFINITION_ID, - "base_model": d.MODEL_NAME, - "parameterization": d.PARAMETERIZATION, - "max_context_length": d.MAX_CONTEXT_LENGTH, + "name": d.name, + "definition_id": d.definition_id, + "base_model": d.model, + "parameterization": d.parameterization, + "max_context_length": d.max_context_length, } for d in definitions ] @@ -354,9 +354,9 @@ async def create_model(body: CreateModelBody) -> dict[str, object]: definition_id=definition_id, spec={ "base_model": next( - d.MODEL_NAME + d.model for d in definitions - if d.DEFINITION_ID == definition_id + if d.definition_id == definition_id ), "lora_config": body.lora_config, "parameterization": {"type": parameterization}, @@ -379,7 +379,7 @@ async def create_sampling_session( detail="base_model or model_path is required", ) selected = routes.select(body.base_model) if body.base_model else None - definition_id = selected.DEFINITION_ID if selected else None + definition_id = selected.definition_id if selected else None if body.model_path is None and definition_id is None: raise HTTPException( status_code=400, @@ -390,9 +390,9 @@ async def create_sampling_session( sampling_session_seq_id=body.sampling_session_seq_id, base_model=( next( - d.MODEL_NAME + d.model for d in definitions - if d.DEFINITION_ID == definition_id + if d.definition_id == definition_id ) if definition_id else body.base_model @@ -489,7 +489,7 @@ async def load_weights(body: LoadWeightsBody) -> dict[str, str]: base_model=body.base_model, user_metadata=body.user_metadata, optimizer=body.optimizer, - definition_ids={d.DEFINITION_ID for d in definitions}, + definition_ids={d.definition_id for d in definitions}, ) return { "request_id": creation.request_id, diff --git a/src/lilo/deployment_cli.py b/src/lilo/deployment_cli.py index 9ca6abf..aaccd5d 100644 --- a/src/lilo/deployment_cli.py +++ b/src/lilo/deployment_cli.py @@ -5,152 +5,113 @@ import argparse import json import os -import re import subprocess import sys -import uuid -from copy import deepcopy -from pathlib import Path import modal -from huggingface_hub import HfApi -from lilo.backends.deployment import resolve_backend_settings from lilo.deployments import ( - LIFECYCLE_FIELDS, PLATFORM_DEFAULTS, - DeploymentRecord, + DeploymentConfig, config_path, load, + platform_defaults, validate_frontend, ) -from lilo.providers.modal.deployment_records import MANIFEST_ENV, deployed_manifest -from lilo.providers.modal.miles_revision import resolve_miles_commit +from lilo.providers.modal.deployment_apps import ( + build_inference_app, + build_trainer_app, +) +from lilo.providers.modal.deployment_configs import ( + CONFIGS_ENV, + PLATFORM_ENV, + deployed_configs, +) -def compile_configs(paths, *, platform=None): - specs = [load(path) for path in paths] - validate_frontend(specs) +def compile_configs(paths): + recipes = [load(path) for path in paths] + validate_frontend(recipes) + return [DeploymentConfig.create(recipe) for recipe in recipes] - miles_commit = ( - resolve_miles_commit() - if any(spec.backend == "miles" for spec in specs) - else None - ) - records = [] - for spec in specs: - revision = spec.revision - if not re.fullmatch(r"[0-9a-fA-F]{40,64}", revision): - revision = HfApi().model_info(spec.model, revision=revision).sha - if not revision or not re.fullmatch(r"[0-9a-fA-F]{40,64}", revision): - raise ValueError( - f"Hugging Face did not return a commit for {spec.model}" - ) - records.append( - DeploymentRecord.create( - spec, - platform=platform, - revision=revision, - miles_commit=miles_commit if spec.backend == "miles" else None, - ) - ) - return records +def validate_refresh(configs, refresh_trainers, refresh_inference): + names = {deployment.name for deployment in configs} + unknown = (refresh_trainers | refresh_inference) - names + if unknown: + raise ValueError(f"unknown configs to refresh: {sorted(unknown)}") -def retain_generations(previous, desired): - """The command supplies the complete active set; old generations remain usable.""" - validate_frontend([row.spec for row in desired]) - expected = desired[0] - for row in previous: - if row.platform != expected.platform or any( - getattr(row.spec, field) != getattr(expected.spec, field) - for field in LIFECYCLE_FIELDS - ): - raise ValueError( - "Cannot change shared storage, secrets, region, or lifecycle while retaining generations; use a separate frontend." - ) - ids = {row.definition_id for row in desired} - return [ - *desired, - *( - row.model_copy(update={"active": False}) - for row in previous - if row.definition_id not in ids - ), - ] +def current_config(configs, name): + return next((deployment for deployment in configs if deployment.name == name), None) -def select_worker_releases(desired, previous, refresh_trainers, refresh_inference): - """Keep deployed code unless the operator explicitly requests a worker update.""" - names = {row.spec.name for row in desired} - unknown = (set(refresh_trainers) | set(refresh_inference)) - names - if unknown: - raise ValueError(f"unknown configs to refresh: {sorted(unknown)}") - active = {row.spec.name: row for row in previous if row.active} - result = [] - for row in desired: - old = active.get(row.spec.name, row) - trainer = ( - uuid.uuid4().hex - if row.spec.name in refresh_trainers - else old.trainer_release - ) - inference = ( - uuid.uuid4().hex - if row.spec.name in refresh_inference - else old.inference_release - ) - result.append(row.with_releases(trainer, inference)) - return result +def app_is_missing(app_name, environment): + try: + modal.App.lookup(app_name, environment_name=environment) + except modal.exception.NotFoundError: + return True + return False + + +def apps_to_publish(configs, current_configs, environment, trainers, inference): + """Add names whose apps are missing or whose trainer/inference settings changed.""" + for deployment in configs: + current = current_config(current_configs, deployment.name) + if app_is_missing(deployment.trainer_app_name, environment) or ( + current is not None and not deployment.same_trainer(current) + ): + trainers.add(deployment.name) + if app_is_missing(deployment.inference_app_name, environment) or ( + current is not None and not deployment.same_inference(current) + ): + inference.add(deployment.name) + return trainers, inference -def deploy(desired, *, refresh_trainers=(), refresh_inference=()): - """Deploy missing workers, then publish the frontend's configuration.""" + +def deploy(configs, platform=None, *, refresh_trainers=(), refresh_inference=()): + """Deploy missing or changed trainer and inference apps, then the frontend.""" if sys.version_info[:2] != (3, 12): raise ValueError( "Python deployment requires Python 3.12 to match the serialized GPU runtime images" ) - settings = desired[0].platform - environment = settings["modal"]["environment"] - previous = [ - DeploymentRecord.model_validate(row) - for row in deployed_manifest(settings["frontend"], environment) + trainers = set(refresh_trainers) + inference = set(refresh_inference) + validate_refresh(configs, trainers, inference) + if platform is None: + platform = platform_defaults() + environment = platform["modal"]["environment"] + current_configs = [ + DeploymentConfig.model_validate(row) + for row in deployed_configs(platform["frontend"], environment) ] - desired = select_worker_releases( - desired, previous, refresh_trainers, refresh_inference + trainers, inference = apps_to_publish( + configs, current_configs, environment, trainers, inference ) - manifest = retain_generations(previous, desired) + + with modal.enable_output(): + for deployment in configs: + if deployment.name in trainers: + build_trainer_app(deployment, platform)[0].deploy( + environment_name=environment + ) + if deployment.name in inference: + build_inference_app(deployment, platform)[0].deploy( + environment_name=environment + ) command = [sys.executable, "-m", "modal", "deploy"] if environment: command += ["--env", environment] - - for row in desired: - for role, app_name in ( - ("trainer", row.trainer_app_name), - ("inference", row.inference_app_name), - ): - try: - modal.App.lookup(app_name, environment_name=environment) - except modal.exception.NotFoundError: - worker_env = { - **os.environ, - "LILO_WORKER_DEPLOYMENT": row.model_dump_json(), - "LILO_WORKER_ROLE": role, - } - if row.miles_commit: - worker_env["LILO_MILES_COMMIT"] = row.miles_commit - subprocess.run( - [*command, "-m", "lilo.providers.modal.deployment_worker_app"], - check=True, - env=worker_env, - ) subprocess.run( [*command, "-m", "lilo.providers.modal.app"], check=True, env={ **os.environ, - MANIFEST_ENV: json.dumps([row.model_dump(mode="json") for row in manifest]), - "LILO_APP_NAME": settings["frontend"], + CONFIGS_ENV: json.dumps( + [deployment.model_dump(mode="json") for deployment in configs] + ), + PLATFORM_ENV: json.dumps(platform), + "LILO_APP_NAME": platform["frontend"], }, ) @@ -161,11 +122,8 @@ def parser(): config = commands.add_parser("config").add_subparsers(dest="action", required=True) init = config.add_parser("init") init.add_argument("--preset", required=True) - for name in ("validate", "resolve"): - cmd = config.add_parser(name) - cmd.add_argument("files", nargs="+") - if name == "resolve": - cmd.add_argument("--output") + validate = config.add_parser("validate") + validate.add_argument("files", nargs="+") apply = commands.add_parser( "deploy", help="Deploy the complete active Python config set behind one frontend", @@ -202,35 +160,19 @@ def main(argv=None): print( f'from lilo.configs.{module} import Config as Parent\n\n\nclass Config(Parent):\n name = "my-model"\n\n\nconfig = Config()' ) - elif args.action == "validate": - specs = [load(path) for path in args.files] - validate_frontend(specs) - for spec in specs: - resolve_backend_settings(spec, "/assets/pending") - print( - f"Validated {len(specs)} deployment(s). Backend integration settings are checked when preparing trainers and pools; native options are checked at worker startup." - ) else: - output = ( - json.dumps( - [ - row.model_dump(mode="json") - for row in compile_configs(args.files) - ], - indent=2, - ) - + "\n" + recipes = [load(path) for path in args.files] + validate_frontend(recipes) + print( + f"Validated {len(recipes)} deployment(s). Backend integration settings are checked when preparing trainers and pools; native options are checked at engine startup." ) - if args.output: - Path(args.output).write_text(output) - else: - print(output, end="") elif args.command == "deploy": - platform = deepcopy(PLATFORM_DEFAULTS) + platform = platform_defaults() platform["frontend"] = args.app platform["modal"].update(environment=args.env, region=args.region) deploy( - compile_configs(args.files, platform=platform), + compile_configs(args.files), + platform, refresh_trainers=args.refresh_trainer, refresh_inference=args.refresh_inference, ) diff --git a/src/lilo/deployments.py b/src/lilo/deployments.py index a3c5b55..c6c054c 100644 --- a/src/lilo/deployments.py +++ b/src/lilo/deployments.py @@ -1,8 +1,7 @@ -"""Python deployment configs and the records saved by the deployment CLI.""" +"""Python deployment configs: the recipe plus backend settings used to deploy it.""" from __future__ import annotations -import hashlib import json import re import runpy @@ -10,15 +9,11 @@ from copy import deepcopy from importlib.resources import files from pathlib import Path -from typing import Any - -from pydantic import BaseModel, ConfigDict, Field, field_serializer, field_validator +from pydantic import BaseModel, ConfigDict, field_serializer, field_validator from lilo.backends.deployment import resolve_backend_settings from lilo.configuration import BaseConfig -LIFECYCLE_FIELDS = ("session_idle_timeout_s", "pool_idle_timeout_s", "sweep_interval_s") - PLATFORM_DEFAULTS = { "frontend": "lilo", "modal": {"environment": None, "region": "us-west"}, @@ -35,194 +30,169 @@ } -def settings_hash(settings: dict) -> str: - """Stable identifier for settings, not a claim that they have been validated.""" - return hashlib.sha256(json.dumps(settings, sort_keys=True).encode()).hexdigest() - - -def deployment_generation( - spec, - platform, - miles_commit, - trainer_release, - inference_release, - trainer_settings, - inference_settings, -): - identity = vars(spec) - return settings_hash( - { - "config": identity, - "platform": platform, - "miles_commit": miles_commit, - "trainer_release": trainer_release, - "inference_release": inference_release, - "trainer_settings": trainer_settings, - "inference_settings": inference_settings, - } - ) - - def platform_defaults(): return deepcopy(PLATFORM_DEFAULTS) -class DeploymentRecord(BaseModel): - """Saved deployment metadata around a Python configuration. +def _jsonable(value): + return json.loads(json.dumps(value)) - The CLI resolves the model revision before creating this record. Its hash - binds jobs to their original configuration across later deploys; - active marks recipes in the latest deploy command; retained recipes remain usable. - """ + +class DeploymentConfig(BaseModel): + """Recipe plus the backend settings used to deploy one model.""" model_config = ConfigDict(extra="forbid", arbitrary_types_allowed=True) - spec: BaseConfig - platform: dict[str, Any] = Field(default_factory=platform_defaults) - trainer_release: str = "initial" - inference_release: str = "initial" - miles_commit: str | None = None + recipe: BaseConfig trainer_settings: dict inference_settings: dict - generation: str - active: bool = True - @field_validator("spec", mode="before") + @field_validator("recipe", mode="before") @classmethod - def restore_config(cls, value): + def restore_recipe(cls, value): return BaseConfig(**value) if isinstance(value, dict) else value - @field_serializer("spec") - def serialize_config(self, config): - return vars(config) + @field_serializer("recipe") + def serialize_recipe(self, recipe): + return vars(recipe) @classmethod - def create( - cls, - spec: BaseConfig, - *, - revision: str, - miles_commit: str | None = None, - platform: dict | None = None, - trainer_release: str = "initial", - inference_release: str = "initial", - ) -> DeploymentRecord: - """Record an already-resolved revision without reparsing the configuration.""" - pinned = deepcopy(spec) - pinned.revision = revision - asset_path = model_asset_path(pinned.model, revision) + def create(cls, recipe: BaseConfig) -> DeploymentConfig: + """Copy the recipe and attach backend settings.""" + pinned = BaseConfig(**_jsonable(vars(recipe))) trainer_settings, inference_settings = resolve_backend_settings( - pinned, asset_path - ) - platform = deepcopy(PLATFORM_DEFAULTS if platform is None else platform) - generation = deployment_generation( - pinned, - platform, - miles_commit, - trainer_release, - inference_release, - trainer_settings, - inference_settings, + pinned, model_asset_path(pinned.model) ) return cls( - spec=pinned, - trainer_settings=trainer_settings, - inference_settings=inference_settings, - platform=platform, - trainer_release=trainer_release, - inference_release=inference_release, - generation=generation, - miles_commit=miles_commit, + recipe=pinned, + trainer_settings=_jsonable(trainer_settings), + inference_settings=_jsonable(inference_settings), ) - def with_releases(self, trainer_release, inference_release): - generation = deployment_generation( - self.spec, - self.platform, - self.miles_commit, - trainer_release, - inference_release, - self.trainer_settings, - self.inference_settings, - ) - return self.model_copy( - update={ - "trainer_release": trainer_release, - "inference_release": inference_release, - "generation": generation, - } - ) + @property + def name(self) -> str: + return self.recipe.name @property - def trainer_hash(self) -> str: - """Identify trainer settings and the deployment-managed code release.""" - return settings_hash( - { - "name": self.spec.name, - "model": self.spec.model, - "revision": self.spec.revision, - "parameterization": self.spec.parameterization, - "max_context_length": self.spec.max_context_length, - "trainer": { - key: value - for key, value in vars(self.spec).items() - if key.startswith("trainer_") - or key - in ( - "backend", - "miles_cfg", - "megatron_cfg", - "sampler_persistence_concurrency", - ) - }, - "settings": self.trainer_settings, - "release": self.trainer_release, - "platform": self.platform, - "miles_commit": self.miles_commit, - } - ) + def model(self) -> str: + return self.recipe.model @property - def inference_hash(self) -> str: - """Identify inference settings, including the adapter shape it must load.""" - return settings_hash( - { - "name": self.spec.name, - "model": self.spec.model, - "revision": self.spec.revision, - "parameterization": self.spec.parameterization, - "max_context_length": self.spec.max_context_length, - "inference": { - key: value - for key, value in vars(self.spec).items() - if key.startswith("inference_") or key == "sglang_cfg" - }, - "settings": self.inference_settings, - "release": self.inference_release, - "platform": self.platform, - } - ) + def parameterization(self) -> str: + return self.recipe.parameterization + + @property + def max_context_length(self) -> int: + return self.recipe.max_context_length + + @property + def definition_id(self) -> str: + return self.recipe.name @property def trainer_app_name(self) -> str: - return f"lilo-trainer-{self.trainer_hash[:24]}" + return f"lilo-trainer-{self.recipe.name}" @property def inference_app_name(self) -> str: - return f"lilo-inference-{self.inference_hash[:24]}" + return f"lilo-inference-{self.recipe.name}" @property - def definition_id(self) -> str: - return f"deployment_{self.spec.name}_{self.generation[:16]}" + def asset_path(self) -> str: + return model_asset_path(self.recipe.model) @property - def asset_path(self) -> str: - return model_asset_path(self.spec.model, self.spec.revision) + def rollout_tensor_parallel_size(self) -> int: + serving = self.inference_settings + return serving.get("tp_size", self.recipe.inference_gpus_per_node) // ( + serving.get("dp_size", 1) if serving.get("enable_dp_attention") else 1 + ) + + def same_trainer(self, other) -> bool: + recipe = self.recipe + current = other.recipe + return ( + recipe.name, + recipe.model, + recipe.parameterization, + recipe.max_context_length, + recipe.backend, + recipe.miles_cfg, + recipe.megatron_cfg, + recipe.sampler_persistence_concurrency, + recipe.trainer_gpu, + recipe.trainer_gpus_per_node, + recipe.trainer_nodes, + recipe.trainer_cpu, + recipe.trainer_memory_mib, + recipe.trainer_max_instances, + recipe.trainer_max_clients_per_instance, + recipe.trainer_timeout_s, + recipe.trainer_env, + self.trainer_settings, + ) == ( + current.name, + current.model, + current.parameterization, + current.max_context_length, + current.backend, + current.miles_cfg, + current.megatron_cfg, + current.sampler_persistence_concurrency, + current.trainer_gpu, + current.trainer_gpus_per_node, + current.trainer_nodes, + current.trainer_cpu, + current.trainer_memory_mib, + current.trainer_max_instances, + current.trainer_max_clients_per_instance, + current.trainer_timeout_s, + current.trainer_env, + other.trainer_settings, + ) + + def same_inference(self, other) -> bool: + recipe = self.recipe + current = other.recipe + return ( + recipe.name, + recipe.model, + recipe.parameterization, + recipe.max_context_length, + recipe.sglang_cfg, + recipe.inference_gpu, + recipe.inference_gpus_per_node, + recipe.inference_cpu, + recipe.inference_memory_mib, + recipe.inference_min_replicas, + recipe.inference_max_replicas, + recipe.inference_target_concurrency, + recipe.inference_scaledown_window_s, + recipe.inference_startup_timeout_s, + recipe.inference_env, + self.inference_settings, + ) == ( + current.name, + current.model, + current.parameterization, + current.max_context_length, + current.sglang_cfg, + current.inference_gpu, + current.inference_gpus_per_node, + current.inference_cpu, + current.inference_memory_mib, + current.inference_min_replicas, + current.inference_max_replicas, + current.inference_target_concurrency, + current.inference_scaledown_window_s, + current.inference_startup_timeout_s, + current.inference_env, + other.inference_settings, + ) -def model_asset_path(model_id, revision): - digest = hashlib.sha256(f"{model_id}@{revision}".encode()).hexdigest() - return f"/assets/{digest}" +def model_asset_path(model_id): + return f"/assets/{model_id}" def config_path(name: str) -> Path: @@ -237,7 +207,6 @@ def load(path: str | Path) -> BaseConfig: path = Path(path).resolve() if path.suffix != ".py": raise ValueError("deployment configs must be Python .py files") - # Let a config import sibling modules using normal Python imports. original_path = sys.path[:] sys.path.insert(0, str(path.parent)) try: @@ -250,15 +219,17 @@ def load(path: str | Path) -> BaseConfig: sys.path[:] = original_path -def validate_frontend(specs: list[BaseConfig]) -> None: - if not specs: +def validate_frontend(recipes: list[BaseConfig]) -> None: + if not recipes: raise ValueError("at least one deployment is required") - if len({s.name for s in specs}) != len(specs): + if len({recipe.name for recipe in recipes}) != len(recipes): raise ValueError("duplicate deployment name") - first = specs[0] - for spec in specs: - if any( - getattr(spec, field) != getattr(first, field) for field in LIFECYCLE_FIELDS + first = recipes[0] + for recipe in recipes: + if ( + recipe.session_idle_timeout_s != first.session_idle_timeout_s + or recipe.pool_idle_timeout_s != first.pool_idle_timeout_s + or recipe.sweep_interval_s != first.sweep_interval_s ): raise ValueError( "deployments on one frontend must share lifecycle settings" diff --git a/src/lilo/providers/modal/app.py b/src/lilo/providers/modal/app.py index bd96cd5..b6440e6 100644 --- a/src/lilo/providers/modal/app.py +++ b/src/lilo/providers/modal/app.py @@ -48,33 +48,40 @@ stop_pool as stop_lora_pool, ) from .sampling import ModalSamplingTaskPlatform -from .deployment_records import MANIFEST_ENV, manifest_from_env +from .deployment_configs import ( + CONFIGS_ENV, + PLATFORM_ENV, + configs_from_env, + platform_from_env, +) from .deployment_apps import ( - definition_from_spec, frontend_settings, + trainer_function, ) SETTINGS = frontend_settings() -APP_NAME = SETTINGS.platform["frontend"] -ROUTING_REGION = SETTINGS.platform["modal"]["region"] +PLATFORM = platform_from_env() +APP_NAME = PLATFORM["frontend"] +ROUTING_REGION = PLATFORM["modal"]["region"] MODEL_ASSET_ROOT = "/assets" -SESSION_IDLE_TIMEOUT = SETTINGS.spec.session_idle_timeout_s -FFT_POOL_IDLE_TIMEOUT = LORA_POOL_IDLE_TIMEOUT = SETTINGS.spec.pool_idle_timeout_s +SESSION_IDLE_TIMEOUT = SETTINGS.recipe.session_idle_timeout_s +FFT_POOL_IDLE_TIMEOUT = LORA_POOL_IDLE_TIMEOUT = SETTINGS.recipe.pool_idle_timeout_s FFT_POOL_TOUCH_INTERVAL = 60.0 LORA_POOL_CHECK_INTERVAL = 60.0 -SWEEP_PERIOD = modal.Period(seconds=SETTINGS.spec.sweep_interval_s) +SWEEP_PERIOD = modal.Period(seconds=SETTINGS.recipe.sweep_interval_s) CHECKPOINT_READ_LOCK = asyncio.Lock() _pool_touches: dict[str, float] = {} _lora_pool_gateways: dict[str, tuple[float, str]] = {} _lora_pool_checks: dict[str, asyncio.Lock] = {} -CHECKPOINT_VOLUME_NAME = SETTINGS.platform["storage"]["checkpoints"] +CHECKPOINT_VOLUME_NAME = PLATFORM["storage"]["checkpoints"] checkpoint_volume = modal.Volume.from_name( CHECKPOINT_VOLUME_NAME, create_if_missing=True, version=2 ) -DEFINITIONS = tuple(definition_from_spec(row) for row in manifest_from_env()) +DEFINITIONS = tuple(configs_from_env()) TRAINER_DEPLOYMENT_ENV = { **trainer_deployment_env(), - MANIFEST_ENV: os.environ[MANIFEST_ENV], + CONFIGS_ENV: os.environ[CONFIGS_ENV], + PLATFORM_ENV: os.environ.get(PLATFORM_ENV, ""), "LILO_APP_NAME": APP_NAME, } app = modal.App(APP_NAME) @@ -109,18 +116,18 @@ async def _delete_checkpoint(uri: str) -> None: @app.function(image=image) -def deployment_manifest(): - return [row.model_dump(mode="json") for row in manifest_from_env()] +def deployed_configs(): + return [row.model_dump(mode="json") for row in configs_from_env()] model_assets = modal.Volume.from_name( - SETTINGS.platform["storage"]["assets"], + PLATFORM["storage"]["assets"], create_if_missing=True, ) -API_SECRET_NAME = SETTINGS.platform["secrets"]["api"] -HF_SECRET_NAME = SETTINGS.platform["secrets"]["huggingface"] +API_SECRET_NAME = PLATFORM["secrets"]["api"] +HF_SECRET_NAME = PLATFORM["secrets"]["huggingface"] proxy_secret = modal.Secret.from_name( - SETTINGS.platform["secrets"]["sampler_proxy"], + PLATFORM["secrets"]["sampler_proxy"], required_keys=["MODAL_PROXY_TOKEN_ID", "MODAL_PROXY_TOKEN_SECRET"], ) @@ -137,16 +144,15 @@ def prepare_model_assets(definition_id: str) -> None: from huggingface_hub import snapshot_download definition = module_for(definition_id) - checkpoint = os.path.normpath(definition.HF_CHECKPOINT) + checkpoint = os.path.normpath(definition.asset_path) if ( os.path.commonpath((MODEL_ASSET_ROOT, checkpoint)) != MODEL_ASSET_ROOT or checkpoint == MODEL_ASSET_ROOT ): raise ValueError(f"invalid model asset path: {checkpoint}") snapshot_download( - repo_id=definition.MODEL_NAME, + repo_id=definition.model, local_dir=checkpoint, - revision=definition.MODEL_REVISION, ) model_assets.commit() @@ -248,8 +254,8 @@ async def _execute_sample(task: dict, stats: dict) -> dict: if parameterization not in {"full", "lora"}: raise ValueError(f"unsupported sampling definition: {definition_id}") definition = module_for(definition_id) - rollout_world_size = definition.ROLLOUT_GPUS - rollout_tensor_parallel_size = definition.ROLLOUT_TENSOR_PARALLEL_SIZE + rollout_world_size = definition.recipe.inference_gpus_per_node + rollout_tensor_parallel_size = definition.rollout_tensor_parallel_size if rollout_world_size % rollout_tensor_parallel_size: raise ValueError("rollout GPU count must be divisible by tensor parallel size") rollout_data_parallel_size = rollout_world_size // rollout_tensor_parallel_size @@ -267,7 +273,7 @@ async def keep_pool_ready() -> None: gateway, data_parallel_size=rollout_data_parallel_size, headers=proxy_auth_headers(), - context_length=definition.MAX_CONTEXT_LENGTH, + context_length=definition.max_context_length, on_wait=keep_pool_ready, stats=stats, ) @@ -288,7 +294,7 @@ async def keep_pool_ready() -> None: data_parallel_size=rollout_data_parallel_size, headers=proxy_auth_headers(), on_wait=lambda: _touch_fft_pool(spec), - context_length=definition.MAX_CONTEXT_LENGTH, + context_length=definition.max_context_length, stats=stats, ) @@ -312,14 +318,14 @@ async def _model_record(kv, model_id: str): def module_for(definition_id: str): for definition in DEFINITIONS: - if definition.DEFINITION_ID == definition_id: + if definition.definition_id == definition_id: return definition raise KeyError(definition_id) def parameterization_for(definition_id: str) -> Parameterization | None: try: - return module_for(definition_id).PARAMETERIZATION + return module_for(definition_id).parameterization except KeyError: return None @@ -353,7 +359,7 @@ async def run(definition_id: str, token: str) -> None: await complete_reconcile(definition_id, token) return module = module_for(definition_id) - maximum_instances = module.TRAINER_MAX_CONTAINERS + maximum_instances = module.recipe.trainer_max_instances try: await reconcile_trainers( shared_kv(), @@ -361,7 +367,7 @@ async def run(definition_id: str, token: str) -> None: definition_id, revision=None, maximum_instances=maximum_instances, - models_per_instance=module.TRAINER_MODELS_PER_INSTANCE, + models_per_instance=module.recipe.trainer_max_clients_per_instance, scale_up=trainer_autoscaling(definition_id), ) except Exception: @@ -400,8 +406,8 @@ async def _spawn_engine(definition_id: str, instance_id: str) -> str: if error := await deployment_error(definition_id): raise ValueError(error) definition = module_for(definition_id) - engine = definition.ENGINE_FUNCTION - call = await engine.spawn.aio(instance_id, definition.RESOLVED.model_dump_json()) + engine = trainer_function(definition, PLATFORM) + call = await engine.spawn.aio(instance_id, definition.model_dump_json()) return call.object_id @@ -479,7 +485,7 @@ async def kick_trainers(definition_id: str) -> bool: await kick_trainer_reconciler(definition_id) if not trainer_autoscaling(definition_id): return False - maximum = module_for(definition_id).TRAINER_MAX_CONTAINERS + maximum = module_for(definition_id).recipe.trainer_max_instances instances = [ instance for instance in await engines.list_instances() diff --git a/src/lilo/providers/modal/deployment_apps.py b/src/lilo/providers/modal/deployment_apps.py index c0eadef..48bdafd 100644 --- a/src/lilo/providers/modal/deployment_apps.py +++ b/src/lilo/providers/modal/deployment_apps.py @@ -10,13 +10,12 @@ import subprocess import sys from importlib import import_module -from types import SimpleNamespace import modal import modal.experimental from modal.config import config -from lilo.deployments import DeploymentRecord, validate_frontend +from lilo.deployments import DeploymentConfig, platform_defaults, validate_frontend from lilo.inference.serving import ( start_fft_sidecar, start_lora_sidecar, @@ -26,9 +25,7 @@ ) from .deployment import trainer_deployment_env -from .deployment_records import ( - manifest_from_env, -) +from .deployment_configs import configs_from_env, platform_from_env from .fft_pool import FFTPoolSpec from .fft_pool import deploy_pool as deploy_fft from .image_dependencies import ( @@ -45,18 +42,14 @@ def frontend_settings(): - deployments = manifest_from_env() - active = [row.spec for row in deployments if row.active] - validate_frontend(active) - if len({row.definition_id for row in deployments}) != len(deployments): - raise ValueError("duplicate deployment generation") + deployments = configs_from_env() + validate_frontend([deployment.recipe for deployment in deployments]) return deployments[0] def image_for(backend): if not modal.is_local(): return modal.Image.debian_slim() - # Image definitions are imported only by the selected worker's deploy process. modules = { "miles": ".miles_image", "megatron": ".megatron_image", @@ -65,8 +58,8 @@ def image_for(backend): return import_module(modules[backend], __package__).image -def volumes_for(record): - storage = record.platform["storage"] +def volumes_for(platform): + storage = platform["storage"] return { "/assets": modal.Volume.from_name(storage["assets"], create_if_missing=True), "/checkpoints": modal.Volume.from_name( @@ -78,8 +71,8 @@ def volumes_for(record): } -def secrets_for(record, *, training=False): - names = record.platform["secrets"] +def secrets_for(platform, *, training=False): + names = platform["secrets"] result = [modal.Secret.from_name(names["api"], required_keys=["TINKER_API_KEY"])] if training: result.append( @@ -100,57 +93,65 @@ def deployment_env(values): return dict(values) -def build_trainer_app(resolved: DeploymentRecord, *, image=None): - spec = resolved.spec - trainer_hash = resolved.trainer_hash - app = modal.App(resolved.trainer_app_name) +def trainer_function(deployment, platform): + return modal.Function.from_name( + deployment.trainer_app_name, + "trainer", + environment_name=platform["modal"]["environment"], + ) + + +def build_trainer_app(deployment: DeploymentConfig, platform=None, *, image=None): + if platform is None: + platform = platform_defaults() + recipe = deployment.recipe + trainer_name = recipe.name + app = modal.App(deployment.trainer_app_name) env = { **trainer_deployment_env(), - **deployment_env(spec.trainer_env), - "LILO_APP_NAME": resolved.platform["frontend"], + **deployment_env(recipe.trainer_env), + "LILO_APP_NAME": platform["frontend"], } def trainer(instance_id: str, config_json: str): - record = DeploymentRecord.model_validate_json(config_json) - if record.trainer_hash != trainer_hash: + saved = DeploymentConfig.model_validate_json(config_json) + if saved.name != trainer_name: raise ValueError("trainer settings do not match the deployed app") - run_trainer(record, instance_id) + run_trainer(saved, instance_id, platform) - if spec.trainer_nodes > 1: - trainer = modal.experimental.clustered(spec.trainer_nodes, rdma=True)(trainer) + if recipe.trainer_nodes > 1: + trainer = modal.experimental.clustered(recipe.trainer_nodes, rdma=True)(trainer) trainer = app.function( name="trainer", serialized=True, - image=image if image is not None else image_for(spec.backend), - gpu=f"{spec.trainer_gpu}:{spec.trainer_gpus_per_node}", - region=resolved.platform["modal"]["region"], - cpu=spec.trainer_cpu, - memory=spec.trainer_memory_mib, - timeout=spec.trainer_timeout_s, - # Admission/reconciliation caps each definition. A function-wide cap - # would block new definitions behind retained jobs sharing this app. + image=image if image is not None else image_for(recipe.backend), + gpu=f"{recipe.trainer_gpu}:{recipe.trainer_gpus_per_node}", + region=platform["modal"]["region"], + cpu=recipe.trainer_cpu, + memory=recipe.trainer_memory_mib, + timeout=recipe.trainer_timeout_s, max_containers=None, min_containers=0, single_use_containers=True, - volumes=volumes_for(resolved), - secrets=secrets_for(resolved, training=True), + volumes=volumes_for(platform), + secrets=secrets_for(platform, training=True), env=env, - experimental_options={"efa_enabled": True} if spec.trainer_nodes > 1 else {}, + experimental_options={"efa_enabled": True} if recipe.trainer_nodes > 1 else {}, )(trainer) return app, trainer -def run_trainer(resolved, instance_id): - spec = resolved.spec - settings = resolved.trainer_settings - # Assets are prepared by the frontend before demand is registered. Reload once - # on startup to see the committed exact snapshot; never race a trainer download. - assets = volumes_for(resolved)["/assets"] - if spec.trainer_nodes > 1: +def run_trainer(deployment, instance_id, platform=None): + if platform is None: + platform = platform_defaults() + recipe = deployment.recipe + settings = deployment.trainer_settings + assets = volumes_for(platform)["/assets"] + if recipe.trainer_nodes > 1: ray_address = start_trainer_cluster( - spec.trainer_nodes, + recipe.trainer_nodes, before_head=assets.reload, before_worker_join=assets.reload, ) @@ -160,110 +161,81 @@ def run_trainer(resolved, instance_id): assets.reload() ray_address = None env = { - **deployment_env(spec.trainer_env), - "LILO_APP_NAME": resolved.platform["frontend"], + **deployment_env(recipe.trainer_env), + "LILO_APP_NAME": platform["frontend"], "LILO_BACKEND_CONFIG": json.dumps(settings), - "LILO_BASE_MODEL": spec.model, - "LILO_BASE_MODEL_REVISION": spec.revision, - "LILO_DEFINITION_ID": resolved.definition_id, - "LILO_CHECKPOINT_VOLUME": resolved.platform["storage"]["checkpoints"], + "LILO_BASE_MODEL": recipe.model, + "LILO_DEFINITION_ID": deployment.definition_id, + "LILO_CHECKPOINT_VOLUME": platform["storage"]["checkpoints"], "LILO_BULLETIN_ROOT": "/bulletin", - "LILO_BULLETIN_VOLUME": resolved.platform["storage"]["bulletin"], - "LILO_DEFINITION_REVISION": resolved.generation, + "LILO_BULLETIN_VOLUME": platform["storage"]["bulletin"], + "LILO_DEFINITION_REVISION": deployment.definition_id, } if ray_address is not None: env["LILO_RAY_ADDRESS"] = ray_address executor = ( "lilo.backends.miles_lora:build_executor" - if spec.backend == "miles" + if recipe.backend == "miles" else "lilo.backends.megatron_fft:build_executor" ) async def failed(error): await shared_kv().put( - f"deployment_failure:{resolved.definition_id}", + f"deployment_failure:{deployment.definition_id}", {"error": str(error), "instance_id": instance_id}, ) run_engine_with_backend( shared_kv(), executor, - definition_id=resolved.definition_id, + definition_id=deployment.definition_id, revision=config["image_id"], instance_id=instance_id, backend_env=env, - nproc=1 if spec.backend == "miles" else spec.trainer_gpus_per_node, - max_models=spec.trainer_max_clients_per_instance, - sampler_persistence_concurrency=spec.sampler_persistence_concurrency, + nproc=1 if recipe.backend == "miles" else recipe.trainer_gpus_per_node, + max_models=recipe.trainer_max_clients_per_instance, + sampler_persistence_concurrency=recipe.sampler_persistence_concurrency, on_startup_error=failed, ) -def definition_from_spec(resolved, *, register_trainer=True, image=None): - spec = resolved.spec - serving = resolved.inference_settings - definition = SimpleNamespace( - DEFINITION_ID=resolved.definition_id, - MODEL_NAME=spec.model, - MODEL_REVISION=spec.revision, - HF_CHECKPOINT=resolved.asset_path, - PARAMETERIZATION=spec.parameterization, - DEPLOYMENT_NAME=spec.name, - RESOLVED=resolved, - MAX_CONTEXT_LENGTH=spec.max_context_length, - TRAINER_MODELS_PER_INSTANCE=spec.trainer_max_clients_per_instance, - TRAINER_MAX_CONTAINERS=spec.trainer_max_instances, - ROLLOUT_GPUS=spec.inference_gpus_per_node, - ROLLOUT_TENSOR_PARALLEL_SIZE=serving.get( - "tp_size", spec.inference_gpus_per_node - ) - // (serving.get("dp_size", 1) if serving.get("enable_dp_attention") else 1), - ) - if register_trainer: - definition.ENGINE_FUNCTION = modal.Function.from_name( - resolved.trainer_app_name, - "trainer", - environment_name=resolved.platform["modal"]["environment"], - ) - return definition - - -def build_rollout_app(resolved, pool, *, image=None): +def build_rollout_app(deployment, pool, platform=None, *, image=None): """Create one frozen-base LoRA pool or one FFT latest/pinned/base pool.""" - spec = resolved.spec - lora = spec.parameterization == "lora" - if pool.definition_id != resolved.definition_id: - raise ValueError("pool generation does not match deployment") + if platform is None: + platform = platform_from_env() + recipe = deployment.recipe + lora = recipe.parameterization == "lora" + if pool.definition_id != deployment.definition_id: + raise ValueError("pool does not match deployment") app = modal.App(pool.app_name) - options = resolved.inference_settings - minimum = spec.inference_min_replicas - maximum = spec.inference_max_replicas - window = spec.inference_scaledown_window_s + options = deployment.inference_settings + minimum = recipe.inference_min_replicas + maximum = recipe.inference_max_replicas + window = recipe.inference_scaledown_window_s if isinstance(pool, FFTPoolSpec): minimum = minimum if pool.min_containers is None else pool.min_containers maximum = maximum if pool.max_containers is None else pool.max_containers window = window if pool.scaledown_window is None else pool.scaledown_window - # Serialized class captures the spec; no model-specific module is imported. @app.server( name="Server", serialized=True, image=image if image is not None else image_for("sglang"), - gpu=f"{spec.inference_gpu}:{spec.inference_gpus_per_node}", - cpu=spec.inference_cpu, - memory=spec.inference_memory_mib, - volumes=volumes_for(resolved), - secrets=secrets_for(resolved), - env=deployment_env(spec.inference_env), + gpu=f"{recipe.inference_gpu}:{recipe.inference_gpus_per_node}", + cpu=recipe.inference_cpu, + memory=recipe.inference_memory_mib, + volumes=volumes_for(platform), + secrets=secrets_for(platform), + env=deployment_env(recipe.inference_env), min_containers=minimum, max_containers=maximum, - target_concurrency=spec.inference_target_concurrency, + target_concurrency=recipe.inference_target_concurrency, scaledown_window=window, - startup_timeout=spec.inference_startup_timeout_s, + startup_timeout=recipe.inference_startup_timeout_s, exit_grace_period=300, port=8000, - routing_region=resolved.platform["modal"]["region"], - compute_region=resolved.platform["modal"]["region"], + routing_region=platform["modal"]["region"], + compute_region=platform["modal"]["region"], ) class Server: @modal.enter() @@ -276,7 +248,7 @@ def start(self): sys.executable, "-m", "lilo.inference.sglang", - resolved.asset_path, + deployment.asset_path, json.dumps(options), ], start_new_session=True, @@ -285,20 +257,20 @@ def start(self): wait_http( "http://127.0.0.1:8001/health", self.sglang, - spec.inference_startup_timeout_s, + recipe.inference_startup_timeout_s, ) kwargs = dict( port=8000, sglang_port=8001, bulletin_root="/bulletin", - bulletin_volume=resolved.platform["storage"]["bulletin"], + bulletin_volume=platform["storage"]["bulletin"], ) self.sidecar = ( start_lora_sidecar(**kwargs) if lora else start_fft_sidecar( **kwargs, - model_path=resolved.asset_path, + model_path=deployment.asset_path, run_id=pool.model_id, pinned_version=None if pool.latest else pool.version, ) @@ -307,7 +279,7 @@ def start(self): wait_http( "http://127.0.0.1:8000/health", self.sidecar, - spec.inference_startup_timeout_s, + recipe.inference_startup_timeout_s, ) except BaseException: terminate(self.sidecar) @@ -322,7 +294,9 @@ def stop(self): return app, Server -def build_inference_app(record, *, image=None): +def build_inference_app(deployment, platform=None, *, image=None): + if platform is None: + platform = platform_defaults() """Freeze pool-building code so idle pools can restart after frontend upgrades.""" if image is None: @@ -334,18 +308,22 @@ def build_inference_app(record, *, image=None): ) .add_local_python_source("lilo", copy=True, ignore=ignore_config_source) ) - app = modal.App(record.inference_app_name) - inference_hash = record.inference_hash + app = modal.App(deployment.inference_app_name) + inference_name = deployment.name @app.function(name="provision", image=image, serialized=True, timeout=1800) def provision(config_json: str, pool_data: dict): - saved = DeploymentRecord.model_validate_json(config_json) - if saved.inference_hash != inference_hash: + saved = DeploymentConfig.model_validate_json(config_json) + if saved.name != inference_name: raise ValueError("inference settings do not match the deployed app") if pool_data["definition_id"] != saved.definition_id: raise ValueError("pool definition does not match deployment") - if saved.spec.parameterization == "lora": - return deploy_lora(LoraPoolSpec.from_dict(pool_data), record=saved) - return deploy_fft(FFTPoolSpec.from_dict(pool_data), record=saved) + if saved.parameterization == "lora": + return deploy_lora( + LoraPoolSpec.from_dict(pool_data), config=saved, platform=platform + ) + return deploy_fft( + FFTPoolSpec.from_dict(pool_data), config=saved, platform=platform + ) return app, provision diff --git a/src/lilo/providers/modal/deployment_configs.py b/src/lilo/providers/modal/deployment_configs.py new file mode 100644 index 0000000..7bc595f --- /dev/null +++ b/src/lilo/providers/modal/deployment_configs.py @@ -0,0 +1,58 @@ +"""Deployment configs and frontend-only platform settings.""" + +import json +import os + +import modal + +from lilo.deployments import DeploymentConfig, platform_defaults + +CONFIGS_ENV = "LILO_DEPLOYMENT_CONFIGS" +PLATFORM_ENV = "LILO_PLATFORM" +POOL_CONFIG_ENV = "LILO_POOL_DEPLOYMENT" + + +def configs_from_env(): + data = os.environ.get(CONFIGS_ENV) + if not data: + raise ValueError( + "Missing deployment configs. Use lilo deploy with your Python config files." + ) + rows = json.loads(data) + if not isinstance(rows, list) or not rows: + raise ValueError("Deployment configs must be a nonempty list") + return [DeploymentConfig.model_validate(row) for row in rows] + + +def platform_from_env(): + data = os.environ.get(PLATFORM_ENV) + if not data: + return platform_defaults() + return json.loads(data) + + +def deployed_configs(frontend, environment=None): + """Read the configuration carried by the currently deployed frontend.""" + try: + return modal.Function.from_name( + frontend, "deployed_configs", environment_name=environment + ).remote() + except modal.exception.NotFoundError: + return [] + + +def pool_config(definition_id): + for deployment in configs_from_env(): + if deployment.definition_id == definition_id: + return deployment + return None + + +def provision_pool(deployment, pool, platform): + """Ask the inference app to create a pool using its original code.""" + provision = modal.Function.from_name( + deployment.inference_app_name, + "provision", + environment_name=platform["modal"]["environment"], + ) + return provision.remote(deployment.model_dump_json(), pool.as_dict()) diff --git a/src/lilo/providers/modal/deployment_pool_app.py b/src/lilo/providers/modal/deployment_pool_app.py index 7cccfae..787661e 100644 --- a/src/lilo/providers/modal/deployment_pool_app.py +++ b/src/lilo/providers/modal/deployment_pool_app.py @@ -2,19 +2,19 @@ import os -from lilo.deployments import DeploymentRecord +from lilo.deployments import DeploymentConfig from .deployment_apps import build_rollout_app -from .deployment_records import POOL_CONFIG_ENV +from .deployment_configs import POOL_CONFIG_ENV, platform_from_env from .fft_pool import FFTPoolSpec from .lora_pool import LoraPoolSpec -resolved = DeploymentRecord.model_validate_json(os.environ[POOL_CONFIG_ENV]) -if resolved.spec.parameterization == "lora": - pool = LoraPoolSpec(resolved.definition_id, revision=resolved.generation[:16]) +deployment = DeploymentConfig.model_validate_json(os.environ[POOL_CONFIG_ENV]) +if deployment.parameterization == "lora": + pool = LoraPoolSpec(deployment.definition_id) else: pool = FFTPoolSpec( - definition_id=resolved.definition_id, + definition_id=deployment.definition_id, model_id=os.environ["LILO_FFT_POOL_MODEL_ID"], latest=os.environ["LILO_FFT_POOL_LATEST"] == "1", version=int(os.environ["LILO_FFT_POOL_VERSION"]), @@ -24,4 +24,4 @@ if f"LILO_FFT_POOL_{key.upper()}" in os.environ }, ) -app, Server = build_rollout_app(resolved, pool) +app, Server = build_rollout_app(deployment, pool, platform_from_env()) diff --git a/src/lilo/providers/modal/deployment_records.py b/src/lilo/providers/modal/deployment_records.py deleted file mode 100644 index c44d1fc..0000000 --- a/src/lilo/providers/modal/deployment_records.py +++ /dev/null @@ -1,52 +0,0 @@ -"""Saved records used to route worker and pool requests.""" - -import json -import os - -import modal - -from lilo.deployments import DeploymentRecord - -MANIFEST_ENV = "LILO_DEPLOYMENT_MANIFEST" -POOL_CONFIG_ENV = "LILO_POOL_DEPLOYMENT" - - -def manifest_from_env(): - data = os.environ.get(MANIFEST_ENV) - if not data: - raise ValueError( - "Missing deployment manifest. Use lilo deploy with your Python config files." - ) - rows = json.loads(data) - if not isinstance(rows, list) or not rows: - raise ValueError("Deployment manifest must be a nonempty list") - return [DeploymentRecord.model_validate(row) for row in rows] - - -def deployed_manifest(frontend, environment=None): - """Read the configuration carried by the currently deployed frontend.""" - function = modal.Function.from_name( - frontend, "deployment_manifest", environment_name=environment - ) - try: - function.hydrate() - except modal.exception.NotFoundError: - return [] - return function.remote() - - -def pool_deployment(definition_id): - for row in manifest_from_env(): - if row.definition_id == definition_id: - return row - return None - - -def provision_pool(record, pool): - """Ask the saved inference app to create a pool using its original code.""" - provision = modal.Function.from_name( - record.inference_app_name, - "provision", - environment_name=record.platform["modal"]["environment"], - ) - return provision.remote(record.model_dump_json(), pool.as_dict()) diff --git a/src/lilo/providers/modal/deployment_worker_app.py b/src/lilo/providers/modal/deployment_worker_app.py deleted file mode 100644 index 89f75c1..0000000 --- a/src/lilo/providers/modal/deployment_worker_app.py +++ /dev/null @@ -1,17 +0,0 @@ -"""Deploy one trainer or inference provisioner; the frontend references its name.""" - -import os - -from lilo.deployments import DeploymentRecord - -from .deployment_apps import build_inference_app, build_trainer_app - -record = DeploymentRecord.model_validate_json(os.environ["LILO_WORKER_DEPLOYMENT"]) -if record.miles_commit: - os.environ["LILO_MILES_COMMIT"] = record.miles_commit -if os.environ["LILO_WORKER_ROLE"] == "trainer": - app, trainer = build_trainer_app(record) -elif os.environ["LILO_WORKER_ROLE"] == "inference": - app, provision = build_inference_app(record) -else: - raise ValueError("unknown worker role") diff --git a/src/lilo/providers/modal/fft_pool.py b/src/lilo/providers/modal/fft_pool.py index be64006..6f34a44 100644 --- a/src/lilo/providers/modal/fft_pool.py +++ b/src/lilo/providers/modal/fft_pool.py @@ -1,11 +1,17 @@ from __future__ import annotations -import hashlib +import json import logging import os import modal -from .deployment_records import POOL_CONFIG_ENV, pool_deployment, provision_pool +from .deployment_configs import ( + PLATFORM_ENV, + POOL_CONFIG_ENV, + platform_from_env, + pool_config, + provision_pool, +) import shutil import subprocess from concurrent.futures import ThreadPoolExecutor @@ -46,11 +52,8 @@ def from_dict(cls, value: dict[str, Any]) -> FFTPoolSpec: @property def app_name(self) -> str: - digest = hashlib.sha256( - f"{self.definition_id}\0{self.model_id}".encode() - ).hexdigest()[:16] suffix = "latest" if self.latest else f"v{self.version}" - return f"lilo-fft-{digest}-{suffix}" + return f"lilo-fft-{self.definition_id}-{self.model_id}-{suffix}" def as_dict(self) -> dict[str, Any]: return asdict(self) @@ -120,23 +123,27 @@ async def pool_gateway(spec: FFTPoolSpec) -> str: return await ModalFlashPool(spec.app_name, "Server").gateway_url_async() -def deploy_pool(spec: FFTPoolSpec, *, record=None) -> str: +def deploy_pool(spec: FFTPoolSpec, *, config=None, platform=None) -> str: pool = ModalFlashPool(spec.app_name, "Server") try: return pool.gateway_url() except Exception as exc: if not isinstance(exc, modal.exception.NotFoundError): raise - if record is None: - saved = pool_deployment(spec.definition_id) + if config is None: + saved = pool_config(spec.definition_id) if saved is None: - raise ValueError(f"missing recorded deployment: {spec.definition_id}") - return provision_pool(saved, spec) + raise ValueError(f"missing deployment config: {spec.definition_id}") + return provision_pool(saved, spec, platform or platform_from_env()) modal_cli = shutil.which("modal") if modal_cli is None: raise RuntimeError("modal CLI is unavailable") - recipe_env = {POOL_CONFIG_ENV: record.model_dump_json()} + platform = platform or platform_from_env() + recipe_env = { + POOL_CONFIG_ENV: config.model_dump_json(), + PLATFORM_ENV: json.dumps(platform), + } env = {**os.environ, **spec.env(), **recipe_env} command = [ modal_cli, @@ -146,7 +153,7 @@ def deploy_pool(spec: FFTPoolSpec, *, record=None) -> str: "--name", spec.app_name, ] - environment = record.platform["modal"]["environment"] + environment = platform["modal"]["environment"] if environment: command.extend(["--env", environment]) subprocess.run(command, env=env, check=True) diff --git a/src/lilo/providers/modal/lora_pool.py b/src/lilo/providers/modal/lora_pool.py index 299649c..baf9fcb 100644 --- a/src/lilo/providers/modal/lora_pool.py +++ b/src/lilo/providers/modal/lora_pool.py @@ -1,10 +1,16 @@ from __future__ import annotations -import hashlib +import json import os import modal -from .deployment_records import POOL_CONFIG_ENV, pool_deployment, provision_pool +from .deployment_configs import ( + PLATFORM_ENV, + POOL_CONFIG_ENV, + platform_from_env, + pool_config, + provision_pool, +) import shutil import subprocess from dataclasses import asdict, dataclass @@ -16,31 +22,18 @@ @dataclass(frozen=True) class LoraPoolSpec: definition_id: str - revision: str = "" def __post_init__(self) -> None: if Path(self.definition_id).name != self.definition_id: raise ValueError(f"invalid definition id: {self.definition_id!r}") - if not self.revision: - object.__setattr__( - self, - "revision", - _definition_revision(self.definition_id), - ) @classmethod def from_dict(cls, value: dict) -> LoraPoolSpec: - return cls( - definition_id=str(value["definition_id"]), - revision=str(value.get("revision", "")), - ) + return cls(definition_id=str(value["definition_id"])) @property def app_name(self) -> str: - digest = hashlib.sha256( - f"{self.definition_id}\0{self.revision}".encode() - ).hexdigest()[:16] - return f"lilo-lora-{digest}" + return f"lilo-lora-{self.definition_id}" def as_dict(self) -> dict[str, str]: return asdict(self) @@ -56,23 +49,27 @@ async def pool_gateway(spec: LoraPoolSpec) -> str: return await ModalFlashPool(spec.app_name, "Server").gateway_url_async() -def deploy_pool(spec: LoraPoolSpec, *, record=None) -> str: +def deploy_pool(spec: LoraPoolSpec, *, config=None, platform=None) -> str: pool = ModalFlashPool(spec.app_name, "Server") try: return pool.gateway_url() except Exception as exc: if not isinstance(exc, modal.exception.NotFoundError): raise - if record is None: - saved = pool_deployment(spec.definition_id) + if config is None: + saved = pool_config(spec.definition_id) if saved is None: - raise ValueError(f"missing recorded deployment: {spec.definition_id}") - return provision_pool(saved, spec) + raise ValueError(f"missing deployment config: {spec.definition_id}") + return provision_pool(saved, spec, platform or platform_from_env()) modal_cli = shutil.which("modal") if modal_cli is None: raise RuntimeError("modal CLI is unavailable") - recipe_env = {POOL_CONFIG_ENV: record.model_dump_json()} + platform = platform or platform_from_env() + recipe_env = { + POOL_CONFIG_ENV: config.model_dump_json(), + PLATFORM_ENV: json.dumps(platform), + } command = [ modal_cli, "deploy", @@ -81,7 +78,7 @@ def deploy_pool(spec: LoraPoolSpec, *, record=None) -> str: "--name", spec.app_name, ] - environment = record.platform["modal"]["environment"] + environment = platform["modal"]["environment"] if environment: command.extend(["--env", environment]) subprocess.run(command, env={**os.environ, **spec.env(), **recipe_env}, check=True) @@ -105,7 +102,3 @@ def stop_pool(spec: LoraPoolSpec) -> None: result.check_returncode() -def _definition_revision(definition_id: str) -> str: - if not definition_id.startswith("deployment_"): - raise ValueError(f"expected a configured deployment id: {definition_id}") - return definition_id.rsplit("_", 1)[-1] diff --git a/src/lilo/providers/modal/miles_image.py b/src/lilo/providers/modal/miles_image.py index fd66aba..0a6802f 100644 --- a/src/lilo/providers/modal/miles_image.py +++ b/src/lilo/providers/modal/miles_image.py @@ -2,8 +2,6 @@ from .image_dependencies import ignore_config_source -from .miles_revision import MILES_REPOSITORY, resolve_miles_commit - from .image_dependencies import ( CORE_PACKAGES, MEGATRON_RUNTIME_CHECK, @@ -12,7 +10,9 @@ ) BASE_IMAGE = "radixark/miles:v0.1.0" -MILES_COMMIT = resolve_miles_commit() +MILES_REPOSITORY = "https://github.com/radixark/miles.git" +# Update this when Miles main should be picked up, then refresh the trainer app. +MILES_COMMIT = "5510af675238be8271c0a24740f70d5116f6d32b" MILES_PATH = "/root/miles" MEGATRON_REPOSITORY = "https://github.com/radixark/Megatron-LM.git" MEGATRON_REVISION = "8c1e05747eb612b382df2632783df5c83a853646" diff --git a/src/lilo/providers/modal/miles_revision.py b/src/lilo/providers/modal/miles_revision.py deleted file mode 100644 index 7b24e91..0000000 --- a/src/lilo/providers/modal/miles_revision.py +++ /dev/null @@ -1,34 +0,0 @@ -"""Resolve a moving Miles ref before constructing the cached Modal image.""" - -import os -import re -import subprocess -from functools import lru_cache - -from lilo.backends.miles_config import MILES_REF - -MILES_REPOSITORY = "https://github.com/radixark/miles.git" - - -def validate_commit(value: str) -> str: - if re.fullmatch(r"[0-9a-f]{40}", value) is None: - raise ValueError("LILO_MILES_COMMIT must be a full lowercase Git commit SHA") - return value - - -@lru_cache(maxsize=1) -def resolve_miles_commit() -> str: - """Resolve once per deployment process; allow an exact reproducibility override.""" - override = os.environ.get("LILO_MILES_COMMIT") - if override is not None: - return validate_commit(override) - ref = f"refs/heads/{MILES_REF}" - result = subprocess.run( - ["git", "ls-remote", "--exit-code", MILES_REPOSITORY, ref], - check=True, capture_output=True, text=True, timeout=30, - ) - entries = [line.split() for line in result.stdout.splitlines()] - commits = [sha for sha, name in entries if name == ref] - if len(commits) != 1: - raise RuntimeError(f"Expected exactly one Miles {ref} commit") - return validate_commit(commits[0]) diff --git a/src/lilo/providers/modal/scoped.py b/src/lilo/providers/modal/scoped.py index a96ed3b..0393de8 100644 --- a/src/lilo/providers/modal/scoped.py +++ b/src/lilo/providers/modal/scoped.py @@ -455,10 +455,11 @@ async def spawn_sampling(task): checkpoint_root=storage.root, ) definition = SimpleNamespace( - DEFINITION_ID=engine.name, - MODEL_NAME=engine.model, - PARAMETERIZATION="full", - MAX_CONTEXT_LENGTH=engine.training.seq_length, + definition_id=engine.name, + name=engine.name, + model=engine.model, + parameterization="full", + max_context_length=engine.training.seq_length, ) return create_control_plane_app( plane, diff --git a/tests/backends/test_megatron_fft.py b/tests/backends/test_megatron_fft.py index 7abb965..3e12ef7 100644 --- a/tests/backends/test_megatron_fft.py +++ b/tests/backends/test_megatron_fft.py @@ -916,11 +916,8 @@ def _impl(self, **kwargs): assert not hasattr(model.decoder, "_forward_impl") -def test_native_resume_rejects_different_or_unknown_base_revision( - tmp_path, monkeypatch -): +def test_checkpoint_metadata_omits_base_revision(): config = EngineModelConfig(hf_checkpoint="/model") - monkeypatch.setenv("LILO_BASE_MODEL_REVISION", "a" * 40) metadata = fft_checkpoint.create_fft_checkpoint_metadata( config, checkpoint_id="snapshot", @@ -928,14 +925,5 @@ def test_native_resume_rejects_different_or_unknown_base_revision( include_optimizer=True, world_size=1, ) - assert metadata.base_model_revision == "a" * 40 - for revision in ("b" * 40, None): - saved = metadata.to_dict() - saved["base_model_revision"] = revision - (tmp_path / fft_checkpoint.CHECKPOINT_METADATA_FILENAME).write_text( - json.dumps(saved) - ) - with pytest.raises(ValueError, match="base_model_revision"): - fft_checkpoint.load_fft_training_checkpoint( - str(tmp_path), config, base_model=BASE_MODEL, world_size=1 - ) + assert metadata.base_model_revision is None + assert "base_model_revision" not in metadata.to_dict() diff --git a/tests/backends/test_miles.py b/tests/backends/test_miles.py index e977a81..f314a3b 100644 --- a/tests/backends/test_miles.py +++ b/tests/backends/test_miles.py @@ -793,25 +793,6 @@ def test_optimizer_worker_failure_is_fatal(tmp_path, outcome): backend.optim_step(("a",), AdamParams(learning_rate=1e-4)) -def test_checkpoint_rejects_different_or_unknown_pinned_base_revision( - tmp_path, monkeypatch -): - backend = _backend(tmp_path) - backend.accept_model("model-a", _spec()) - monkeypatch.setenv("LILO_BASE_MODEL_REVISION", "a" * 40) - backend.capture_checkpoint( - "model-a", "capture", destination="step-1", include_optimizer=False - ) - uri = Path(backend.persist_checkpoint("capture", "step-1")) - metadata = json.loads((uri / "metadata.json").read_text()) - assert metadata["base_model_revision"] == "a" * 40 - backend._validate_checkpoint(metadata, backend.jobs["model-a"], False) - for revision in ("b" * 40, None): - metadata["base_model_revision"] = revision - with pytest.raises(ValueError, match="base model revision"): - backend._validate_checkpoint(metadata, backend.jobs["model-a"], False) - - class _FakeCheckpointVolume: """A volume whose committed state holds shards this container never wrote.""" diff --git a/tests/conftest.py b/tests/conftest.py index 877cfd8..34fe8c8 100644 --- a/tests/conftest.py +++ b/tests/conftest.py @@ -1,5 +1,4 @@ -"""Keep image-definition imports offline during CPU tests.""" +"""The Miles runtime reads this commit from the environment.""" import os -# No image is built by this suite. Resolver tests explicitly clear this override. os.environ.setdefault("LILO_MILES_COMMIT", "a" * 40) diff --git a/tests/control_plane/test_http.py b/tests/control_plane/test_http.py index 0a32d73..1e84428 100644 --- a/tests/control_plane/test_http.py +++ b/tests/control_plane/test_http.py @@ -16,16 +16,18 @@ DEFINITION = "qwen3_8b" DEFINITIONS = ( SimpleNamespace( - DEFINITION_ID=DEFINITION, - MODEL_NAME=BASE_MODEL, - PARAMETERIZATION="lora", - MAX_CONTEXT_LENGTH=16_384, + definition_id=DEFINITION, + name=DEFINITION, + model=BASE_MODEL, + parameterization="lora", + max_context_length=16_384, ), SimpleNamespace( - DEFINITION_ID=f"{DEFINITION}_full", - MODEL_NAME=BASE_MODEL, - PARAMETERIZATION="full", - MAX_CONTEXT_LENGTH=65_536, + definition_id=f"{DEFINITION}_full", + name=f"{DEFINITION}_full", + model=BASE_MODEL, + parameterization="full", + max_context_length=65_536, ), ) @@ -467,13 +469,20 @@ def create(seq: int, **body): def test_explicit_deployment_keeps_canonical_model_name() -> None: async def run(): - explicit = SimpleNamespace(DEFINITION_ID="isolated", MODEL_NAME=BASE_MODEL, - PARAMETERIZATION="lora", MAX_CONTEXT_LENGTH=16384) + explicit = SimpleNamespace( + definition_id="isolated", + name="isolated", + model=BASE_MODEL, + parameterization="lora", + max_context_length=16384, + ) plane = ControlPlane(InMemoryKeyValueStore(), LocalEnginePlatform("isolated", EchoExecutor)) app = create_control_plane_app(plane, (*DEFINITIONS, explicit), retrieve_window=1.0) async with httpx.AsyncClient(base_url="http://test", transport=httpx.ASGITransport(app=app)) as client: listed = (await client.get("/api/v1/lilo/deployments")).json()["deployments"] - assert [row["generation"] for row in listed] == [d.DEFINITION_ID for d in (*DEFINITIONS, explicit)] + assert [row["definition_id"] for row in listed] == [ + d.definition_id for d in (*DEFINITIONS, explicit) + ] session = (await client.post("/api/v1/create_session", json={"tags": [], "sdk_version": "0.5.0"})).json()["session_id"] response = await client.post("/api/v1/create_model", json={"session_id": session, "model_seq_id": 0, "base_model": "isolated", "lora_config": {"rank": 16}}) diff --git a/tests/control_plane/test_sdk_e2e.py b/tests/control_plane/test_sdk_e2e.py index cb7a726..6d69796 100644 --- a/tests/control_plane/test_sdk_e2e.py +++ b/tests/control_plane/test_sdk_e2e.py @@ -29,18 +29,20 @@ MAX_CONTEXT_LENGTH = 32_768 DEFINITIONS = ( SimpleNamespace( - DEFINITION_ID=DEFINITION, - MODEL_NAME=BASE_MODEL, - PARAMETERIZATION="lora", - MAX_CONTEXT_LENGTH=MAX_CONTEXT_LENGTH, + definition_id=DEFINITION, + name=DEFINITION, + model=BASE_MODEL, + parameterization="lora", + max_context_length=MAX_CONTEXT_LENGTH, ), ) FULL_DEFINITIONS = ( SimpleNamespace( - DEFINITION_ID=FULL_DEFINITION, - MODEL_NAME=BASE_MODEL, - PARAMETERIZATION="full", - MAX_CONTEXT_LENGTH=MAX_CONTEXT_LENGTH, + definition_id=FULL_DEFINITION, + name=FULL_DEFINITION, + model=BASE_MODEL, + parameterization="full", + max_context_length=MAX_CONTEXT_LENGTH, ), ) diff --git a/tests/providers/conftest.py b/tests/providers/conftest.py index 5400f8a..316bcf8 100644 --- a/tests/providers/conftest.py +++ b/tests/providers/conftest.py @@ -1,17 +1,15 @@ -"""Shared provider tests construct the app from an explicit offline manifest.""" +"""Shared provider tests construct the app from explicit offline configs.""" import json import os -from lilo.deployments import load, config_path, DeploymentRecord +from lilo.deployments import load, config_path, DeploymentConfig os.environ.setdefault( - "LILO_DEPLOYMENT_MANIFEST", + "LILO_DEPLOYMENT_CONFIGS", json.dumps( [ - DeploymentRecord.create( - load(config_path(name)), revision="a" * 40 - ).model_dump(mode="json") + DeploymentConfig.create(load(config_path(name))).model_dump(mode="json") for name in ("qwen35-9b-fft-64k", "qwen35-9b-lora-16k") ] ), diff --git a/tests/providers/test_checkpoint_storage.py b/tests/providers/test_checkpoint_storage.py index 1e290a0..2bc56dd 100644 --- a/tests/providers/test_checkpoint_storage.py +++ b/tests/providers/test_checkpoint_storage.py @@ -10,7 +10,7 @@ ) FULL_DEFINITIONS = tuple( - definition for definition in DEFINITIONS if definition.PARAMETERIZATION == "full" + definition for definition in DEFINITIONS if definition.parameterization == "full" ) @@ -35,23 +35,20 @@ def test_checkpoint_storage_creates_one_v2_volume_without_live_lookup() -> None: def test_deployments_use_configured_checkpoint_storage(): from lilo.providers.modal.deployment_apps import volumes_for + from lilo.providers.modal.app import PLATFORM from lilo.backends.deployment import backend_config for definition in DEFINITIONS: - spec = definition.RESOLVED.spec with patch.object( modal.Volume, "from_name", side_effect=lambda name, **kwargs: (name, kwargs) ): - volumes = volumes_for(definition.RESOLVED) + volumes = volumes_for(PLATFORM) assert volumes[CHECKPOINT_ROOT] == ( - definition.RESOLVED.platform["storage"]["checkpoints"], + PLATFORM["storage"]["checkpoints"], {"create_if_missing": True, "version": 2}, ) - assert ( - volumes["/bulletin"][0] - == definition.RESOLVED.platform["storage"]["bulletin"] - ) - assert backend_config(spec)["checkpoint_dir"] == CHECKPOINT_ROOT + assert volumes["/bulletin"][0] == PLATFORM["storage"]["bulletin"] + assert backend_config(definition.recipe)["checkpoint_dir"] == CHECKPOINT_ROOT def test_same_checkpoint_name_isolated_by_model(tmp_path, monkeypatch) -> None: diff --git a/tests/providers/test_definition_registry.py b/tests/providers/test_definition_registry.py index e9c053f..cf48119 100644 --- a/tests/providers/test_definition_registry.py +++ b/tests/providers/test_definition_registry.py @@ -7,8 +7,8 @@ def test_definition_registry_resolves_every_definition() -> None: for definition in DEFINITIONS: - assert module_for(definition.DEFINITION_ID) is definition - assert parameterization_for(definition.DEFINITION_ID) == ( - definition.PARAMETERIZATION + assert module_for(definition.definition_id) is definition + assert parameterization_for(definition.definition_id) == ( + definition.parameterization ) - assert definition.ENGINE_FUNCTION is not None + assert definition.trainer_app_name diff --git a/tests/providers/test_deployment_apps.py b/tests/providers/test_deployment_apps.py index 55dfa9f..5f73073 100644 --- a/tests/providers/test_deployment_apps.py +++ b/tests/providers/test_deployment_apps.py @@ -6,14 +6,14 @@ import modal import pytest -from lilo.deployments import load, config_path, DeploymentRecord -from lilo.providers.modal import deployment_apps, deployment_records +from lilo.deployments import load, config_path, DeploymentConfig, platform_defaults +from lilo.providers.modal import deployment_apps, deployment_configs from lilo.providers.modal.fft_pool import FFTPoolSpec from lilo.providers.modal.lora_pool import LoraPoolSpec def deployment(preset="qwen35-9b-lora-16k"): - return DeploymentRecord.create(load(config_path(preset)), revision="a" * 40) + return DeploymentConfig.create(load(config_path(preset))) class App: @@ -55,9 +55,10 @@ def test_trainer_declaration_and_executor_configuration( builders, monkeypatch, preset, backend, clients, nproc ): row = deployment(preset) - row.platform["storage"]["checkpoints"] = "test-custom-checkpoints" + platform = platform_defaults() + platform["storage"]["checkpoints"] = "test-custom-checkpoints" image = object() - app, trainer = deployment_apps.build_trainer_app(row, image=image) + app, trainer = deployment_apps.build_trainer_app(row, platform, image=image) declaration, _ = app.functions["trainer"] assert declaration["gpu"] == "H100:4" assert declaration["region"] == "us-west" @@ -84,9 +85,9 @@ def test_trainer_declaration_and_executor_configuration( assert kwargs["max_models"] == clients assert kwargs["nproc"] == nproc assert kwargs["backend_env"]["LILO_CHECKPOINT_VOLUME"] == "test-custom-checkpoints" - assert kwargs["backend_env"]["LILO_BASE_MODEL_REVISION"] == "a" * 40 + assert kwargs["backend_env"]["LILO_BASE_MODEL"] == row.model config = json.loads(kwargs["backend_env"]["LILO_BACKEND_CONFIG"]) - assert config[row.spec.backend]["hf_checkpoint"] == row.asset_path + assert config[row.recipe.backend]["hf_checkpoint"] == row.asset_path assert config["checkpoint_dir"] == "/checkpoints" assert reloaded == [True] @@ -108,7 +109,7 @@ def test_pool_starts_native_server_and_correct_sidecar(builders, monkeypatch, ki assert app.name == pool.app_name assert ( settings["gpu"] - == f"{row.spec.inference_gpu}:{row.spec.inference_gpus_per_node}" + == f"{row.recipe.inference_gpu}:{row.recipe.inference_gpus_per_node}" ) assert settings["min_containers"] == 0 assert settings["target_concurrency"] == 16 @@ -138,7 +139,7 @@ def test_pool_starts_native_server_and_correct_sidecar(builders, monkeypatch, ki assert commands[0][2] == "lilo.inference.sglang" assert commands[0][3] == row.asset_path native = json.loads(commands[0][4]) - assert native["context_length"] == row.spec.max_context_length + assert native["context_length"] == row.max_context_length if kind == "lora": assert native["enable_lora"] is True assert native["max_lora_rank"] == 32 @@ -161,12 +162,12 @@ def test_pool_starts_native_server_and_correct_sidecar(builders, monkeypatch, ki assert stops == [process, process] -def test_pool_lookup_uses_recorded_generation(monkeypatch): +def test_pool_lookup_uses_definition_name(monkeypatch): row = deployment() - monkeypatch.setenv(deployment_records.MANIFEST_ENV, json.dumps([row.model_dump()])) - saved = deployment_records.pool_deployment(row.definition_id) - assert saved.generation == row.generation - assert deployment_records.pool_deployment("deployment_missing_123") is None + monkeypatch.setenv(deployment_configs.CONFIGS_ENV, json.dumps([row.model_dump()])) + saved = deployment_configs.pool_config(row.definition_id) + assert saved.definition_id == row.definition_id + assert deployment_configs.pool_config("missing") is None def test_startup_failure_is_visible_and_blocks_new_spawns(monkeypatch): @@ -190,7 +191,7 @@ async def run(): asyncio.run(run()) -def test_real_modal_app_constructs_from_manifest_without_legacy_catalog(monkeypatch): +def test_real_modal_app_constructs_from_configs_without_legacy_catalog(monkeypatch): import os import subprocess import sys @@ -198,7 +199,7 @@ def test_real_modal_app_constructs_from_manifest_without_legacy_catalog(monkeypa row = deployment() env = { **os.environ, - deployment_records.MANIFEST_ENV: json.dumps([row.model_dump()]), + deployment_configs.CONFIGS_ENV: json.dumps([row.model_dump()]), } result = subprocess.run( [ @@ -206,13 +207,13 @@ def test_real_modal_app_constructs_from_manifest_without_legacy_catalog(monkeypa "-c", """ import importlib, sys, modal -from lilo.providers.modal import deployment_apps, deployment_records +from lilo.providers.modal import deployment_apps, deployment_configs deployment_apps.image_for = lambda backend: modal.Image.debian_slim() app = importlib.import_module('lilo.providers.modal.app') assert len(app.DEFINITIONS) == 1 assert app.APP_NAME == 'lilo' -assert app.deployment_manifest.local() == [d.model_dump(mode='json') for d in app.manifest_from_env()] -assert app.DEFINITIONS[0].ENGINE_FUNCTION is not None +assert app.deployed_configs.local() == [d.model_dump(mode='json') for d in app.configs_from_env()] +assert app.DEFINITIONS[0].trainer_app_name assert not any(name.startswith('lilo.providers.modal.definitions.') for name in sys.modules) print('constructed') """, @@ -232,11 +233,15 @@ def test_admission_changes_preserve_serialized_trainer(builders): first = deployment() old_bytes = serialize(deployment_apps.build_trainer_app(first, image="test")[1]) changed = first.model_copy(deep=True) - changed.active = False + changed.recipe.trainer_timeout_s = 1 new_bytes = serialize(deployment_apps.build_trainer_app(changed, image="test")[1]) assert new_bytes == old_bytes - assert first.active is True - changed.spec.trainer_gpu = "H200" + changed.recipe.trainer_gpu = "H200" + assert ( + serialize(deployment_apps.build_trainer_app(changed, image="test")[1]) + == old_bytes + ) + changed.recipe.name = "other-model" assert ( serialize(deployment_apps.build_trainer_app(changed, image="test")[1]) != old_bytes @@ -248,7 +253,7 @@ def test_pool_launch_uses_only_generic_deployment_app(monkeypatch, kind): from lilo.providers.modal import fft_pool, lora_pool row = deployment("qwen35-9b-lora-16k" if kind == "lora" else "qwen35-4b-fft-64k") - monkeypatch.setenv(deployment_records.MANIFEST_ENV, json.dumps([row.model_dump()])) + monkeypatch.setenv(deployment_configs.CONFIGS_ENV, json.dumps([row.model_dump()])) module = lora_pool if kind == "lora" else fft_pool spec = ( LoraPoolSpec(row.definition_id) @@ -274,14 +279,14 @@ def gateway_url(self): "run", lambda command, **kwargs: calls.append((command, kwargs)), ) - assert module.deploy_pool(spec, record=row) == "https://pool" + assert module.deploy_pool(spec, config=row) == "https://pool" command, kwargs = calls[0] assert ( command[command.index("-m") + 1] == "lilo.providers.modal.deployment_pool_app" ) assert ( - json.loads(kwargs["env"][deployment_records.POOL_CONFIG_ENV])["generation"] - == row.generation + json.loads(kwargs["env"][deployment_configs.POOL_CONFIG_ENV])["recipe"]["name"] + == row.name ) @@ -298,8 +303,8 @@ def test_frontend_uses_deployed_trainer_without_building_it(monkeypatch): "from_name", lambda *a, **k: calls.append((a, k)) or "remote-trainer", ) - definition = deployment_apps.definition_from_spec(row) - assert definition.ENGINE_FUNCTION == "remote-trainer" + function = deployment_apps.trainer_function(row, platform_defaults()) + assert function == "remote-trainer" assert calls[0][0] == (row.trainer_app_name, "trainer") @@ -308,7 +313,7 @@ def test_missing_pool_uses_saved_provisioner(monkeypatch, kind): from lilo.providers.modal import fft_pool, lora_pool row = deployment("qwen35-9b-lora-16k" if kind == "lora" else "qwen35-4b-fft-64k") - monkeypatch.setenv(deployment_records.MANIFEST_ENV, json.dumps([row.model_dump()])) + monkeypatch.setenv(deployment_configs.CONFIGS_ENV, json.dumps([row.model_dump()])) module = lora_pool if kind == "lora" else fft_pool spec = ( LoraPoolSpec(row.definition_id) @@ -340,7 +345,7 @@ def gateway_url(self): assert calls == [(row.inference_app_name, "provision")] -def test_provisioner_rejects_wrong_settings_and_uses_saved_record( +def test_provisioner_rejects_wrong_settings_and_uses_saved_config( builders, monkeypatch ): row = deployment() @@ -350,49 +355,28 @@ def test_provisioner_rejects_wrong_settings_and_uses_saved_record( monkeypatch.setattr( deployment_apps, "deploy_lora", - lambda pool, *, record: calls.append((pool, record)) or "https://pool", + lambda pool, *, config, platform=None: calls.append((pool, config)) + or "https://pool", ) pool = LoraPoolSpec(row.definition_id) assert provision(row.model_dump_json(), pool.as_dict()) == "https://pool" - assert calls[0][1].inference_hash == row.inference_hash + assert calls[0][1].name == row.name changed = row.model_copy(deep=True) - changed.inference_release = "2" + changed.recipe.name = "other-model" with pytest.raises(ValueError, match="inference settings"): provision(changed.model_dump_json(), pool.as_dict()) @pytest.mark.parametrize("role", ["trainer", "inference"]) -def test_real_worker_entrypoint_constructs_offline(monkeypatch, role): - import os - import subprocess - import sys - +def test_trainer_and_inference_apps_construct_offline(builders, role): row = deployment() - env = { - **os.environ, - "LILO_WORKER_DEPLOYMENT": row.model_dump_json(), - "LILO_WORKER_ROLE": role, - } - result = subprocess.run( - [ - sys.executable, - "-c", - """ -import modal -from lilo.providers.modal import deployment_apps, deployment_records -deployment_apps.image_for = lambda backend: modal.Image.debian_slim() -import lilo.providers.modal.deployment_worker_app as worker -assert worker.app.name.startswith("lilo-") -print(worker.app.name) -""", - ], - env=env, - capture_output=True, - text=True, - timeout=30, + if role == "trainer": + app, _ = deployment_apps.build_trainer_app(row, image="test") + else: + app, _ = deployment_apps.build_inference_app(row, image="test") + assert app.name == ( + row.trainer_app_name if role == "trainer" else row.inference_app_name ) - assert result.returncode == 0, result.stderr - assert f"lilo-{role}-" in result.stdout def test_spawn_passes_job_configuration_to_saved_trainer(monkeypatch): @@ -409,18 +393,19 @@ async def spawn(instance_id, config_json): async def no_error(definition_id): return None - definition = SimpleNamespace( - RESOLVED=row, - ENGINE_FUNCTION=SimpleNamespace(spawn=SimpleNamespace(aio=spawn)), - ) monkeypatch.setattr(app, "deployment_error", no_error) - monkeypatch.setattr(app, "module_for", lambda _: definition) + monkeypatch.setattr(app, "module_for", lambda _: row) + monkeypatch.setattr( + app, + "trainer_function", + lambda deployment, platform: SimpleNamespace(spawn=SimpleNamespace(aio=spawn)), + ) assert asyncio.run(app._spawn_engine(row.definition_id, "instance")) == "call-id" assert calls == [("instance", row.model_dump_json())] def test_declared_compute_settings_reach_modal(builders): - spec = deepcopy(deployment().spec) + spec = deepcopy(deployment().recipe) spec.trainer_timeout_s = 90 spec.trainer_cpu = 12 spec.trainer_memory_mib = 123456 @@ -429,7 +414,7 @@ def test_declared_compute_settings_reach_modal(builders): spec.inference_max_replicas = 3 spec.inference_cpu = 6 spec.inference_memory_mib = 45000 - row = DeploymentRecord.create(spec, revision="a" * 40) + row = DeploymentConfig.create(spec) trainer_app, _ = deployment_apps.build_trainer_app(row, image="test") trainer, _ = trainer_app.functions["trainer"] assert (trainer["cpu"], trainer["memory"], trainer["timeout"]) == (12, 123456, 90) @@ -505,4 +490,3 @@ def unexpected(*args, **kwargs): deployment_apps.build_rollout_app( row, LoraPoolSpec(row.definition_id), image="test" ) - deployment_apps.definition_from_spec(row, register_trainer=False) diff --git a/tests/providers/test_deployment_e2e_helper.py b/tests/providers/test_deployment_e2e_helper.py index 2b84993..b7ead8e 100644 --- a/tests/providers/test_deployment_e2e_helper.py +++ b/tests/providers/test_deployment_e2e_helper.py @@ -3,27 +3,26 @@ import pytest -from lilo.deployments import load, config_path, DeploymentRecord -from lilo.providers.modal import deployment_records +from lilo.deployments import load, config_path, DeploymentConfig +from lilo.providers.modal import deployment_configs @pytest.mark.parametrize("preset", ["qwen35-9b-lora-16k", "qwen35-4b-fft-64k"]) def test_e2e_helper_reads_active_deployed_configuration(monkeypatch, preset): - row = DeploymentRecord.create(load(config_path(preset)), revision="a" * 40) - retired = row.model_copy(update={"active": False, "generation": "b" * 64}) + row = DeploymentConfig.create(load(config_path(preset))) monkeypatch.setattr( - deployment_records, - "deployed_manifest", - lambda frontend: [retired.model_dump(), row.model_dump()], + deployment_configs, + "deployed_configs", + lambda frontend: [row.model_dump()], ) helper = runpy.run_path( str(Path(__file__).parents[2] / "scripts/e2e_engine_definition.py") ) - definition, mode = helper["_definition"]("test-frontend", row.spec.name) + definition, mode = helper["_definition"]("test-frontend", row.name) assert definition.DEFINITION_ID == row.definition_id - assert definition.MAX_CONTEXT_LENGTH == row.spec.max_context_length - assert definition.GPUS == row.spec.trainer_gpus_per_node - assert mode == row.spec.parameterization + assert definition.MAX_CONTEXT_LENGTH == row.max_context_length + assert definition.GPUS == row.recipe.trainer_gpus_per_node + assert mode == row.parameterization assert definition.MAX_TOKENS_PER_MICROBATCH > 0 - with pytest.raises(ValueError, match="one active configuration"): + with pytest.raises(ValueError, match="one configuration"): helper["_definition"]("test-frontend", "missing") diff --git a/tests/providers/test_deployment_presets.py b/tests/providers/test_deployment_presets.py index cc9a51f..c33fabc 100644 --- a/tests/providers/test_deployment_presets.py +++ b/tests/providers/test_deployment_presets.py @@ -1,12 +1,9 @@ import pytest -from lilo.deployments import load, config_path, DeploymentRecord +from lilo.deployments import load, config_path, DeploymentConfig from lilo.backends.deployment import backend_config, serving_options -from lilo.providers.modal.deployment_records import manifest_from_env -from lilo.providers.modal.deployment_apps import ( - definition_from_spec, - frontend_settings, -) +from lilo.providers.modal.deployment_configs import configs_from_env +from lilo.providers.modal.deployment_apps import frontend_settings @pytest.mark.parametrize( @@ -35,11 +32,9 @@ def test_moe_rollout_preserves_attention_data_parallelism(): options = serving_options(spec) assert options["tp_size"] == options["dp_size"] == options["ep_size"] == 4 assert options["enable_dp_attention"] is True - definition = definition_from_spec( - DeploymentRecord.create(spec, revision="a" * 40), register_trainer=False - ) - assert definition.ROLLOUT_GPUS == 4 - assert definition.ROLLOUT_TENSOR_PARALLEL_SIZE == 1 + definition = DeploymentConfig.create(spec) + assert definition.recipe.inference_gpus_per_node == 4 + assert definition.rollout_tensor_parallel_size == 1 @pytest.mark.parametrize("context,cp", [(16384, 1), (65536, 2), (131072, 4)]) @@ -64,12 +59,12 @@ def test_single_client_recipe_keeps_shared_backend_capacity(): @pytest.mark.parametrize("value", [None, "", "[]", "{}"]) -def test_missing_manifest_has_no_python_catalog_fallback(monkeypatch, value): +def test_missing_configs_have_no_python_catalog_fallback(monkeypatch, value): if value is None: - monkeypatch.delenv("LILO_DEPLOYMENT_MANIFEST", raising=False) + monkeypatch.delenv("LILO_DEPLOYMENT_CONFIGS", raising=False) else: - monkeypatch.setenv("LILO_DEPLOYMENT_MANIFEST", value) - with pytest.raises(ValueError, match="manifest"): - manifest_from_env() - with pytest.raises(ValueError, match="manifest"): + monkeypatch.setenv("LILO_DEPLOYMENT_CONFIGS", value) + with pytest.raises(ValueError, match="configs"): + configs_from_env() + with pytest.raises(ValueError, match="configs"): frontend_settings() diff --git a/tests/providers/test_lora_pool.py b/tests/providers/test_lora_pool.py index 099800a..98b317c 100644 --- a/tests/providers/test_lora_pool.py +++ b/tests/providers/test_lora_pool.py @@ -2,13 +2,12 @@ def test_lora_pool_is_shared_by_every_adapter_for_definition() -> None: - first = LoraPoolSpec("deployment_example_0123456789abcdef") + first = LoraPoolSpec("qwen35-9b-lora-16k") second = LoraPoolSpec.from_dict(first.as_dict()) assert first == second assert first.app_name == second.app_name - assert first.app_name != LoraPoolSpec(first.definition_id, "old").app_name - assert first.app_name.startswith("lilo-lora-") + assert first.app_name == "lilo-lora-qwen35-9b-lora-16k" assert first.env() == { "LILO_LORA_POOL_APP_NAME": first.app_name, "LILO_LORA_POOL_DEFINITION_ID": first.definition_id, @@ -27,22 +26,8 @@ def test_stop_already_stopped_lora_pool_succeeds_but_real_failure_propagates( [], 1, "", "App is already stopped. (Stopped yesterday).\n" ) monkeypatch.setattr(lora_pool.subprocess, "run", lambda *args, **kwargs: result) - spec = LoraPoolSpec("definition", "revision") + spec = LoraPoolSpec("qwen35-9b-lora-16k") lora_pool.stop_pool(spec) result.stderr = "authentication failed" with pytest.raises(subprocess.CalledProcessError): lora_pool.stop_pool(spec) - - -def test_pool_revision_comes_from_resolved_generation(): - first = LoraPoolSpec("deployment_example_0123456789abcdef") - changed = LoraPoolSpec("deployment_example_fedcba9876543210") - assert first.revision == "0123456789abcdef" - assert first.app_name != changed.app_name - - -def test_python_definition_cannot_choose_a_pool_revision(): - import pytest - - with pytest.raises(ValueError, match="configured deployment id"): - LoraPoolSpec("qwen3_5_9b_base_miles_lora_16k") diff --git a/tests/providers/test_miles_revision.py b/tests/providers/test_miles_revision.py deleted file mode 100644 index 9e71524..0000000 --- a/tests/providers/test_miles_revision.py +++ /dev/null @@ -1,55 +0,0 @@ -import subprocess -from types import SimpleNamespace - -import pytest - -from lilo.providers.modal.miles_revision import resolve_miles_commit - - -@pytest.fixture(autouse=True) -def fresh_resolution(monkeypatch): - monkeypatch.delenv("LILO_MILES_COMMIT", raising=False) - resolve_miles_commit.cache_clear() - yield - resolve_miles_commit.cache_clear() - - -def test_main_resolves_once_and_next_deployment_can_advance(monkeypatch): - calls = [] - head = ["b" * 40] - - def run(argv, **kwargs): - calls.append(argv) - assert argv[-1] == "refs/heads/main" - assert kwargs["check"] and kwargs["timeout"] == 30 - return SimpleNamespace(stdout=f"{head[0]}\trefs/heads/main\n") - - monkeypatch.setattr(subprocess, "run", run) - assert resolve_miles_commit() == "b" * 40 - head[0] = "c" * 40 - assert resolve_miles_commit() == "b" * 40 - assert len(calls) == 1 - resolve_miles_commit.cache_clear() - assert resolve_miles_commit() == "c" * 40 - - -def test_exact_override_does_not_lookup_main(monkeypatch): - monkeypatch.setenv("LILO_MILES_COMMIT", "d" * 40) - def unexpected(*args, **kwargs): - pytest.fail("Pinned reproduction should not query main") - monkeypatch.setattr(subprocess, "run", unexpected) - assert resolve_miles_commit() == "d" * 40 - resolve_miles_commit.cache_clear() - monkeypatch.setenv("LILO_MILES_COMMIT", "main") - with pytest.raises(ValueError, match="full lowercase Git commit"): - resolve_miles_commit() - - -def test_lookup_failure_is_not_cached(monkeypatch): - def fail(*args, **kwargs): - raise subprocess.CalledProcessError(2, args[0]) - monkeypatch.setattr(subprocess, "run", fail) - with pytest.raises(subprocess.CalledProcessError): - resolve_miles_commit() - monkeypatch.setattr(subprocess, "run", lambda *a, **k: SimpleNamespace(stdout=f'{"e" * 40}\trefs/heads/main\n')) - assert resolve_miles_commit() == "e" * 40 diff --git a/tests/providers/test_modal_app.py b/tests/providers/test_modal_app.py index b70ab67..4e3d724 100644 --- a/tests/providers/test_modal_app.py +++ b/tests/providers/test_modal_app.py @@ -10,13 +10,11 @@ from lilo.providers.modal.fft_pool import FFTPoolSpec from lilo.providers.modal.lora_pool import LoraPoolSpec -from lilo.deployments import load, config_path, DeploymentRecord +from lilo.deployments import load, config_path, DeploymentConfig def definition_id(preset): - return DeploymentRecord.create( - load(config_path(preset)), revision="a" * 40 - ).definition_id + return DeploymentConfig.create(load(config_path(preset))).definition_id FULL_DEFINITION = definition_id("qwen35-9b-fft-64k") @@ -30,9 +28,9 @@ def reset_lora_pool_cache(monkeypatch): monkeypatch.setattr(modal_app, "_lora_pool_checks", {}) -def test_definitions_come_only_from_the_configured_manifest() -> None: +def test_definitions_come_only_from_the_configured_records() -> None: modal_app = importlib.import_module("lilo.providers.modal.app") - assert {definition.DEFINITION_ID for definition in modal_app.DEFINITIONS} == { + assert {definition.definition_id for definition in modal_app.DEFINITIONS} == { FULL_DEFINITION, LORA_DEFINITION, } @@ -70,9 +68,7 @@ async def kick(_definition_id: str) -> None: monkeypatch.setattr(modal_app, "ModalSessionKeyValueStores", SimpleNamespace) monkeypatch.setattr(modal_app, "ModalEnginePlatform", lambda *args: engines) monkeypatch.setattr(modal_app, "kick_trainer_reconciler", kick) - monkeypatch.setattr( - modal_app.module_for(LORA_DEFINITION), "TRAINER_MAX_CONTAINERS", 1 - ) + modal_app.module_for(LORA_DEFINITION).recipe.trainer_max_instances = 1 plane = modal_app._plane() assert asyncio.run(plane.reconcile_trainers(LORA_DEFINITION)) is available @@ -186,8 +182,8 @@ def test_prepare_model_assets_validates_snapshot_before_commit(monkeypatch) -> N modal_app = importlib.import_module("lilo.providers.modal.app") events = [] - def download(*, repo_id: str, local_dir: str, revision: str) -> None: - events.append(("download", repo_id, local_dir, revision)) + def download(*, repo_id: str, local_dir: str) -> None: + events.append(("download", repo_id, local_dir)) monkeypatch.setattr("huggingface_hub.snapshot_download", download) monkeypatch.setattr( @@ -200,12 +196,7 @@ def download(*, repo_id: str, local_dir: str, revision: str) -> None: definition = modal_app.module_for(FULL_DEFINITION) assert events == [ - ( - "download", - definition.MODEL_NAME, - definition.HF_CHECKPOINT, - definition.MODEL_REVISION, - ), + ("download", definition.model, definition.asset_path), ("commit",), ] @@ -499,7 +490,7 @@ def test_cleanup_stops_superseded_lora_pool(monkeypatch) -> None: modal_app = importlib.import_module("lilo.providers.modal.app") registry = InMemoryKeyValueStore() current = LoraPoolSpec(LORA_DEFINITION) - superseded = LoraPoolSpec(LORA_DEFINITION, revision="superseded") + superseded = LoraPoolSpec("retired-lora") stopped = [] model = ModelRecord( @@ -864,7 +855,7 @@ async def run(): assert lookups == [LoraPoolSpec(LORA_DEFINITION)] -def test_lora_readiness_cache_is_revision_specific(monkeypatch): +def test_lora_readiness_cache_is_per_definition(monkeypatch): modal_app = importlib.import_module("lilo.providers.modal.app") lookups = [] @@ -876,27 +867,25 @@ async def lookup(spec): monkeypatch.setattr(modal_app, "shared_kv", InMemoryKeyValueStore) async def run(): - first = LoraPoolSpec(LORA_DEFINITION, "old") - second = LoraPoolSpec(LORA_DEFINITION, "new") + first = LoraPoolSpec(LORA_DEFINITION) + second = LoraPoolSpec("other-lora") assert await modal_app._ready_lora_pool( first ) != await modal_app._ready_lora_pool(second) await modal_app._ready_lora_pool(first) asyncio.run(run()) - assert [spec.revision for spec in lookups] == ["old", "new"] + assert [spec.definition_id for spec in lookups] == [LORA_DEFINITION, "other-lora"] def test_lora_cleanup_continues_after_failure_and_retries(monkeypatch, caplog): modal_app = importlib.import_module("lilo.providers.modal.app") registry = InMemoryKeyValueStore() - specs = [ - LoraPoolSpec(LORA_DEFINITION, revision) for revision in ("failed", "healthy") - ] + specs = [LoraPoolSpec("failed-lora"), LoraPoolSpec("healthy-lora")] fail = True def stop(spec): - if spec.revision == "failed" and fail: + if spec.definition_id == "failed-lora" and fail: raise RuntimeError("stop unavailable") monkeypatch.setattr(modal_app, "shared_kv", lambda: registry) diff --git a/tests/providers/test_sglang_entrypoint.py b/tests/providers/test_sglang_entrypoint.py index f10c770..ef83c72 100644 --- a/tests/providers/test_sglang_entrypoint.py +++ b/tests/providers/test_sglang_entrypoint.py @@ -5,13 +5,11 @@ import sys from types import ModuleType -from lilo.deployments import DeploymentRecord, config_path, load +from lilo.deployments import DeploymentConfig, config_path, load def test_sglang_constructor_receives_recorded_settings(monkeypatch): - row = DeploymentRecord.create( - load(config_path("qwen35-9b-lora-16k")), revision="a" * 40 - ) + row = DeploymentConfig.create(load(config_path("qwen35-9b-lora-16k"))) calls = [] class ServerArgs: diff --git a/tests/test_deployment_cli.py b/tests/test_deployment_cli.py index e254c06..2c58aa6 100644 --- a/tests/test_deployment_cli.py +++ b/tests/test_deployment_cli.py @@ -1,3 +1,4 @@ +from contextlib import nullcontext from copy import deepcopy import json import subprocess @@ -8,42 +9,59 @@ import pytest from lilo import deployment_cli as cli -from lilo.deployments import load, config_path, DeploymentRecord -from lilo.providers.modal.deployment_records import MANIFEST_ENV, deployed_manifest +from lilo.deployments import load, config_path, DeploymentConfig +from lilo.providers.modal.deployment_configs import CONFIGS_ENV, deployed_configs def deployment(): - return DeploymentRecord.create( - load(config_path("qwen35-9b-lora-16k")), revision="a" * 40 - ) + return DeploymentConfig.create(load(config_path("qwen35-9b-lora-16k"))) @pytest.fixture def deployed(monkeypatch): - state = SimpleNamespace(apps=set(), manifest=[], calls=[], fail=None) + monkeypatch.setattr(cli.sys, "version_info", (3, 12, 0)) + state = SimpleNamespace(apps=set(), configs=[], calls=[], fail=None) + + class FakeApp: + def __init__(self, name, role): + self.name = name + self.role = role + + def deploy(self, environment_name=None): + state.calls.append(self.role) + if state.fail == self.role: + raise RuntimeError(f"{self.role} deploy failed") + state.apps.add(self.name) - def read_manifest(frontend, environment): - return state.manifest + def read_configs(frontend, environment): + return state.configs def lookup(name, **kwargs): if name not in state.apps: raise modal.exception.NotFoundError("not deployed") def run(command, *, env, check): - role = env.get("LILO_WORKER_ROLE", "frontend") - state.calls.append(role) - if state.fail == role: + state.calls.append("frontend") + if state.fail == "frontend": raise subprocess.CalledProcessError(1, command) - if role == "frontend": - state.manifest = json.loads(env[MANIFEST_ENV]) - else: - row = DeploymentRecord.model_validate_json(env["LILO_WORKER_DEPLOYMENT"]) - state.apps.add( - row.trainer_app_name if role == "trainer" else row.inference_app_name - ) - - monkeypatch.setattr(cli, "deployed_manifest", read_manifest) + state.configs = json.loads(env[CONFIGS_ENV]) + + monkeypatch.setattr(cli, "deployed_configs", read_configs) monkeypatch.setattr(modal.App, "lookup", lookup) + monkeypatch.setattr( + cli, + "build_trainer_app", + lambda row, platform=None, **k: (FakeApp(row.trainer_app_name, "trainer"), None), + ) + monkeypatch.setattr( + cli, + "build_inference_app", + lambda row, platform=None, **k: ( + FakeApp(row.inference_app_name, "inference"), + None, + ), + ) + monkeypatch.setattr(cli.modal, "enable_output", lambda: nullcontext()) monkeypatch.setattr(subprocess, "run", run) monkeypatch.setattr( modal.Dict, @@ -53,7 +71,7 @@ def run(command, *, env, check): return state -def test_deploy_reuses_modal_apps_and_retains_previous_config(deployed): +def test_deploy_reuses_unchanged_apps(deployed): row = deployment() cli.deploy([row]) assert deployed.calls == ["trainer", "inference", "frontend"] @@ -61,31 +79,33 @@ def test_deploy_reuses_modal_apps_and_retains_previous_config(deployed): cli.deploy([row]) assert deployed.calls == ["frontend"] - spec = deepcopy(row.spec) + spec = deepcopy(row.recipe) spec.inference_max_replicas = 6 - changed = DeploymentRecord.create(spec, revision="a" * 40) + changed = DeploymentConfig.create(spec) deployed.calls.clear() cli.deploy([changed]) assert deployed.calls == ["inference", "frontend"] - assert [(r["generation"], r["active"]) for r in deployed.manifest] == [ - (changed.generation, True), - (row.generation, False), - ] + assert [r["recipe"]["name"] for r in deployed.configs] == [changed.name] - # Modal's actual state wins over a frontend record that mentions an old app. + # Modal's actual state wins over a frontend config that mentions an old app. deployed.apps.remove(changed.inference_app_name) deployed.calls.clear() cli.deploy([changed]) assert deployed.calls == ["inference", "frontend"] -@pytest.mark.parametrize("failure", ["inference", "frontend"]) -def test_retry_discovers_completed_workers_without_pending_records(deployed, failure): +@pytest.mark.parametrize( + "failure,error", + [("inference", RuntimeError), ("frontend", subprocess.CalledProcessError)], +) +def test_retry_discovers_completed_apps_without_pending_configs( + deployed, failure, error +): row = deployment() deployed.fail = failure - with pytest.raises(subprocess.CalledProcessError): + with pytest.raises(error): cli.deploy([row]) - assert deployed.manifest == [] + assert deployed.configs == [] assert row.trainer_app_name in deployed.apps deployed.calls.clear() deployed.fail = None @@ -95,52 +115,43 @@ def test_retry_discovers_completed_workers_without_pending_records(deployed, fai ) -def test_refresh_keeps_other_workers_and_is_remembered_by_frontend(deployed): +def test_refresh_redeploys_only_that_trainer(deployed): miles = deployment() - fft = DeploymentRecord.create( - load(config_path("qwen35-4b-fft-64k")), revision="a" * 40 - ) + fft = DeploymentConfig.create(load(config_path("qwen35-4b-fft-64k"))) cli.deploy([miles, fft]) deployed.calls.clear() - cli.deploy([miles, fft], refresh_trainers=[miles.spec.name]) + cli.deploy([miles, fft], refresh_trainers=[miles.name]) assert deployed.calls == ["trainer", "frontend"] - updated = deployed.manifest[0] - assert updated["trainer_release"] != "initial" - assert updated["inference_release"] == "initial" deployed.calls.clear() cli.deploy([miles, fft]) assert deployed.calls == ["frontend"] - assert deployed.manifest[0] == updated def test_existing_frontend_without_deployment_metadata_can_be_updated(deployed): row = deployment() - deployed.apps.add(row.platform["frontend"]) + deployed.apps.add("lilo") cli.deploy([row]) - assert deployed.manifest[0]["generation"] == row.generation + assert deployed.configs[0]["recipe"]["name"] == row.name -def test_read_manifest_from_deployed_function(monkeypatch): - function = SimpleNamespace( - hydrate=Mock(), remote=Mock(return_value=[{"configuration": "saved"}]) - ) +def test_read_configs_from_deployed_function(monkeypatch): + function = SimpleNamespace(remote=Mock(return_value=[{"configuration": "saved"}])) lookup = Mock(return_value=function) monkeypatch.setattr(modal.Function, "from_name", lookup) - assert deployed_manifest("my-app", "dev") == [{"configuration": "saved"}] + assert deployed_configs("my-app", "dev") == [{"configuration": "saved"}] lookup.assert_called_once_with( - "my-app", "deployment_manifest", environment_name="dev" + "my-app", "deployed_configs", environment_name="dev" ) - function.hydrate.side_effect = modal.exception.NotFoundError("no function") - assert deployed_manifest("my-app", "dev") == [] - function.hydrate.side_effect = None + function.remote.side_effect = modal.exception.NotFoundError("no function") + assert deployed_configs("my-app", "dev") == [] function.remote.side_effect = RuntimeError("frontend failed") with pytest.raises(RuntimeError, match="frontend failed"): - deployed_manifest("my-app", "dev") + deployed_configs("my-app", "dev") def test_validate_never_resolves_or_deploys(monkeypatch, capsys): monkeypatch.setattr( - cli, "compile_configs", lambda *args: pytest.fail("unexpected resolution") + cli, "compile_configs", lambda *args, **kwargs: pytest.fail("unexpected compile") ) monkeypatch.setattr( cli, "deploy", lambda *args: pytest.fail("unexpected deployment") @@ -160,34 +171,7 @@ def test_deploy_rejects_python_mismatch_before_remote_changes(monkeypatch): cli.deploy([deployment()]) -@pytest.mark.parametrize("revision,lookups", [("main", 1), ("a" * 40, 0)]) -def test_compile_pins_revision_at_external_boundary( - tmp_path, monkeypatch, revision, lookups -): - from types import SimpleNamespace - from unittest.mock import Mock - import huggingface_hub - from lilo.providers.modal import miles_revision - - path = tmp_path / "model.py" - path.write_text( - "from lilo.configs.qwen35_9b_lora_16k import Config as Parent\n" - f"class Config(Parent):\n revision = {revision!r}\n" - "config = Config()\n" - ) - lookup = Mock(return_value=SimpleNamespace(sha="a" * 40)) - monkeypatch.setattr(huggingface_hub.HfApi, "model_info", lookup) - monkeypatch.setattr(miles_revision, "resolve_miles_commit", lambda: "b" * 40) - (row,) = cli.compile_configs([path]) - assert row.spec.revision == "a" * 40 - assert lookup.call_count == lookups - if lookups: - lookup.return_value.sha = None - with pytest.raises(ValueError, match="did not return a commit"): - cli.compile_configs([path]) - - -def test_worker_source_mount_excludes_authoring_configs(): +def test_source_mount_excludes_authoring_configs(): from pathlib import Path from lilo.providers.modal.image_dependencies import ignore_config_source @@ -201,12 +185,13 @@ def test_worker_source_mount_excludes_authoring_configs(): def test_deploy_command_owns_platform_settings(monkeypatch): seen = {} - def compile(paths, *, platform): - seen["platform"] = platform - return ["record"] + def compile(paths): + seen["compiled"] = paths + return ["config"] - def deploy(rows, **kwargs): - seen["rows"] = rows + def deploy(configs, platform, **kwargs): + seen["configs"] = configs + seen["platform"] = platform seen.update(kwargs) monkeypatch.setattr(cli, "compile_configs", compile) @@ -228,23 +213,10 @@ def deploy(rows, **kwargs): assert seen["platform"]["frontend"] == "my-lilo" assert seen["platform"]["modal"] == {"environment": "dev", "region": "us-east"} assert seen["refresh_trainers"] == ["my-model"] - assert seen["rows"] == ["record"] - + assert seen["configs"] == ["config"] -def test_builtin_config_resolves_revision_automatically(monkeypatch): - from types import SimpleNamespace - import huggingface_hub - from lilo.providers.modal import miles_revision - calls = [] - monkeypatch.setattr(miles_revision, "resolve_miles_commit", lambda: "b" * 40) - monkeypatch.setattr( - huggingface_hub.HfApi, - "model_info", - lambda self, model, *, revision: calls.append((model, revision)) - or SimpleNamespace(sha="a" * 40), - ) +def test_builtin_config_keeps_the_model_path(): (row,) = cli.compile_configs([config_path("qwen35-9b-lora-16k")]) - assert calls == [("Qwen/Qwen3.5-9B-Base", "main")] - assert row.spec.revision == "a" * 40 - assert load(config_path("qwen35-9b-lora-16k")).revision == "main" + assert row.asset_path == "/assets/Qwen/Qwen3.5-9B-Base" + assert not hasattr(load(config_path("qwen35-9b-lora-16k")), "revision") diff --git a/tests/test_deployments.py b/tests/test_deployments.py index 5020e1e..4fa8b2d 100644 --- a/tests/test_deployments.py +++ b/tests/test_deployments.py @@ -8,15 +8,13 @@ import pytest from lilo.deployments import ( - DeploymentRecord, + DeploymentConfig, load, config_path, validate_frontend, ) -from lilo.deployment_cli import retain_generations from lilo.control_plane.deployments import DeploymentRoutes from lilo.backends.deployment import backend_config -from lilo.providers.modal.deployment_apps import definition_from_spec from lilo.backends.miles_arguments import apply_config_overrides @@ -32,11 +30,7 @@ def recipe(preset="qwen35-9b-lora-16k", **changes): def resolved(spec=None, **changes): - return DeploymentRecord.create(spec or recipe(**changes), revision="a" * 40) - - -def definition(value): - return definition_from_spec(value, register_trainer=False) + return DeploymentConfig.create(spec or recipe(**changes)) def test_presets_context_topology_and_backend_options(): @@ -67,7 +61,7 @@ def test_no_model_catalog_required(): "num_attention_heads": 12, }, ) - assert definition(resolved(spec)).MODEL_NAME == "my-org/new-model" + assert resolved(spec).model == "my-org/new-model" assert backend_config(spec)["miles"]["model_type"] == "" @@ -93,51 +87,40 @@ def test_invalid_integrations_fail_when_building_backend_settings(changes, match serving_options(spec) -def test_generation_and_asset_paths_include_exact_base(): +def test_asset_paths_follow_the_model(): a = resolved() - assert a.generation != resolved(recipe(trainer_gpu="H200")).generation + assert a.asset_path == f"/assets/{a.recipe.model}" b = resolved(recipe(model="other/Qwen3.5-9B-Base")) assert a.asset_path != b.asset_path - assert a.asset_path != DeploymentRecord.create(a.spec, revision="b" * 40).asset_path -def test_routing_uses_deployment_order_and_preserves_explicit_generations(): +def test_routing_uses_deployment_order(): small = resolved() large = resolved(recipe("qwen35-9b-lora-64k")) - routes = DeploymentRoutes(map(definition, [small, large])) - assert routes.select(small.spec.model, "lora").DEFINITION_ID == small.definition_id + routes = DeploymentRoutes([small, large]) + assert routes.select(small.model, "lora").definition_id == small.definition_id assert routes.capabilities()[0]["max_context_length"] == 16384 - validate_frontend([small.spec, large.spec]) + validate_frontend([small.recipe, large.recipe]) - switched = retain_generations([small, large], [large]) - routes = DeploymentRoutes(map(definition, switched)) - assert routes.select(small.spec.model, "lora").DEFINITION_ID == large.definition_id - assert ( - routes.select(small.definition_id, "lora").DEFINITION_ID == small.definition_id - ) + routes = DeploymentRoutes([large]) + assert routes.select(small.model, "lora").definition_id == large.definition_id assert routes.capabilities()[0]["max_context_length"] == 65536 - # A retained-only model is still listed and selectable. other = resolved(recipe(name="other", model="org/other")) - routes = DeploymentRoutes(map(definition, retain_generations([small], [other]))) - assert routes.select(small.spec.model, "lora").DEFINITION_ID == small.definition_id - assert {row["model_name"] for row in routes.capabilities()} == { - small.spec.model, - other.spec.model, - } + routes = DeploymentRoutes([other]) + assert routes.select(other.model, "lora").definition_id == other.definition_id + assert {row["model_name"] for row in routes.capabilities()} == {other.model} def test_sampling_uses_order_and_training_filters_parameterization(): lora = resolved() - fft = resolved(recipe("qwen35-4b-fft-64k", model=lora.spec.model)) + fft = resolved(recipe("qwen35-4b-fft-64k", model=lora.model)) for first, second in ((lora, fft), (fft, lora)): - routes = DeploymentRoutes(map(definition, [first, second])) - assert routes.select(lora.spec.model).DEFINITION_ID == first.definition_id - assert ( - routes.select(lora.spec.model, "lora").DEFINITION_ID == lora.definition_id - ) - assert routes.select(lora.spec.model, "full").DEFINITION_ID == fft.definition_id - assert routes.select(second.definition_id).DEFINITION_ID == second.definition_id + routes = DeploymentRoutes([first, second]) + assert routes.select(lora.model).definition_id == first.definition_id + assert routes.select(lora.model, "lora").definition_id == lora.definition_id + assert routes.select(lora.model, "full").definition_id == fft.definition_id + assert routes.select(second.definition_id).definition_id == second.definition_id assert routes.select("missing") is None assert routes.select(lora.definition_id, "full") is None @@ -178,7 +161,7 @@ async def run(): first = resolved() other = resolved(recipe(name="other", model="org/other-model")) app = create_control_plane_app( - plane, list(map(definition, [first, other])), api_key="test" + plane, [first, other], api_key="test" ) async with httpx.AsyncClient( transport=httpx.ASGITransport(app=app), @@ -194,7 +177,7 @@ async def run(): json={ "session_id": session, "model_seq_id": seq, - "base_model": row.spec.model, + "base_model": row.model, "lora_config": {"rank": 32}, }, ) @@ -321,13 +304,13 @@ def test_reserved_environment_is_checked_by_modal_setup(): def test_record_creation_copies_without_reparsing(): spec = recipe() original = vars(spec) - row = DeploymentRecord.create(spec, revision="a" * 40) + row = DeploymentConfig.create(spec) assert vars(spec) == original - assert row.spec.revision == "a" * 40 - row.spec.miles_cfg["max_lora_rank"] = 64 + assert row.asset_path == f"/assets/{spec.model}" + row.recipe.miles_cfg["max_lora_rank"] = 64 assert spec.miles_cfg["max_lora_rank"] == 32 saved = row.model_dump_json() - assert DeploymentRecord.model_validate_json(saved).model_dump( + assert DeploymentConfig.model_validate_json(saved).model_dump( mode="json" ) == row.model_dump(mode="json") @@ -356,7 +339,7 @@ def test_worker_record_contains_resolved_settings(monkeypatch): record = resolved() assert record.trainer_settings["miles"]["actor_num_gpus_per_node"] == 4 assert record.inference_settings["max_lora_rank"] == 32 - assert DeploymentRecord.model_validate_json(record.model_dump_json()).model_dump( + assert DeploymentConfig.model_validate_json(record.model_dump_json()).model_dump( mode="json" ) == record.model_dump(mode="json") @@ -385,26 +368,25 @@ def test_no_yaml_config_ingestion(tmp_path): load(tmp_path / "old.yaml") -def test_worker_hashes_cover_only_their_settings(): +def test_recipe_app_names_follow_the_recipe_name(): base = resolved() + assert base.trainer_app_name == f"lilo-trainer-{base.name}" + assert base.inference_app_name == f"lilo-inference-{base.name}" + inference = resolved(recipe(sglang_cfg__max_running_requests=24)) - assert inference.trainer_hash == base.trainer_hash - assert inference.inference_hash != base.inference_hash + assert inference.trainer_app_name == base.trainer_app_name + assert inference.same_trainer(base) + assert not inference.same_inference(base) trainer = resolved(recipe(miles_cfg__max_tokens_per_gpu=8192)) - assert trainer.trainer_hash != base.trainer_hash - assert trainer.inference_hash == base.inference_hash + assert trainer.trainer_app_name == base.trainer_app_name + assert not trainer.same_trainer(base) + assert trainer.same_inference(base) - adapter = resolved(recipe(miles_cfg__max_lora_rank=64)) - assert adapter.trainer_hash != base.trainer_hash - assert adapter.inference_hash != base.inference_hash - - upgraded = DeploymentRecord.create( - base.spec, revision="a" * 40, inference_release="2" - ) - assert upgraded.trainer_hash == base.trainer_hash - assert upgraded.inference_hash != base.inference_hash - assert "implementation" not in upgraded.model_dump() + other = resolved(recipe("qwen35-9b-lora-64k")) + assert other.trainer_app_name != base.trainer_app_name + assert other.inference_app_name != base.inference_app_name + assert "implementation" not in other.model_dump() def test_examples_only_contain_model_infrastructure(): @@ -413,7 +395,7 @@ def test_examples_only_contain_model_infrastructure(): for path in Path(config_path("qwen35-9b-lora-16k")).parent.glob("qwen*.py"): config = load(path) assert not hasattr(config, "deployment") - assert config.revision == "main" + assert not hasattr(config, "revision") @pytest.mark.parametrize("value", ["invalid", 7]) @@ -441,12 +423,12 @@ def test_managed_backend_values_fail_before_record_creation(section, field): candidate = deepcopy(base) candidate.megatron_cfg[section] = {field: 1} with pytest.raises(ValueError, match=field): - DeploymentRecord.create(candidate, revision="a" * 40) + DeploymentConfig.create(candidate) def test_multinode_ownership_and_topology(): config = load(config_path("qwen38-27b-lora-256k")) - row = DeploymentRecord.create(config, revision="a" * 40) + row = DeploymentConfig.create(config) miles = row.trainer_settings["miles"] assert miles["actor_num_nodes"] == 2 assert miles["actor_num_gpus_per_node"] == 8 @@ -455,7 +437,7 @@ def test_multinode_ownership_and_topology(): invalid = deepcopy(config) invalid.miles_cfg["actor_num_nodes"] = 3 with pytest.raises(ValueError, match="actor_num_nodes"): - DeploymentRecord.create(invalid, revision="a" * 40) + DeploymentConfig.create(invalid) def test_config_inheritance_and_constructor_overrides_copy_nested_options(): diff --git a/tests/test_system.py b/tests/test_system.py index 55aa9d9..d6e1bb0 100644 --- a/tests/test_system.py +++ b/tests/test_system.py @@ -16,9 +16,11 @@ DEFINITION = "qwen3_8b" DEFINITIONS = ( SimpleNamespace( - DEFINITION_ID=DEFINITION, - MODEL_NAME=BASE_MODEL, - PARAMETERIZATION="lora", + definition_id=DEFINITION, + name=DEFINITION, + model=BASE_MODEL, + parameterization="lora", + max_context_length=16384, ), ) API_KEY = "tml-test" diff --git a/uv.lock b/uv.lock index 3d76c80..83bac41 100644 --- a/uv.lock +++ b/uv.lock @@ -9,16 +9,16 @@ resolution-markers = [ [[package]] name = "aiohappyeyeballs" version = "2.7.1" -source = { registry = "https://pypi.org/simple" } -sdist = { url = "https://files.pythonhosted.org/packages/ce/f4/eec0465c2f67b2664688d0240b3212d5196fd89e741df67ddb81f8d35658/aiohappyeyeballs-2.7.1.tar.gz", hash = "sha256:065665c041c42a5938ed220bdcd7230f22527fbec085e1853d2402c8a3615d9d", size = 24757, upload-time = "2026-07-01T17:11:55.501Z" } +source = { registry = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/" } +sdist = { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/aiohappyeyeballs/2.7.1/aiohappyeyeballs-2.7.1.tar.gz", hash = "sha256:065665c041c42a5938ed220bdcd7230f22527fbec085e1853d2402c8a3615d9d", size = 24757, upload-time = "2026-07-01T17:11:55.501Z" } wheels = [ - { url = "https://files.pythonhosted.org/packages/71/43/1947f06babed6b3f1d7f38b0c767f52df66bfb2bc10b468c4a7de9eceff2/aiohappyeyeballs-2.7.1-py3-none-any.whl", hash = "sha256:9243213661e29250eb41368e5daa826fc017156c3b8a11440826b2e3ed376472", size = 15038, upload-time = "2026-07-01T17:11:54.055Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/aiohappyeyeballs/2.7.1/aiohappyeyeballs-2.7.1-py3-none-any.whl", hash = "sha256:9243213661e29250eb41368e5daa826fc017156c3b8a11440826b2e3ed376472", size = 15038, upload-time = "2026-07-01T17:11:54.055Z" }, ] [[package]] name = "aiohttp" version = "3.14.3" -source = { registry = "https://pypi.org/simple" } +source = { registry = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/" } dependencies = [ { name = "aiohappyeyeballs" }, { name = "aiosignal" }, @@ -29,223 +29,223 @@ dependencies = [ { name = "typing-extensions" }, { name = "yarl" }, ] -sdist = { url = "https://files.pythonhosted.org/packages/58/d9/22ce5786ac0c1653ae8b6c23bded02c1686d11f0dbb45b31ce128e0df985/aiohttp-3.14.3.tar.gz", hash = "sha256:9491196535a88924a60afd5b5f434b5b203b6cc616250878dbdb223a8f7844bc", size = 7971213, upload-time = "2026-07-23T01:57:27.037Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/f8/5c/b3e4ff8ad43a8afef9602c5e90285936da1beaea8b029016b793891f03c3/aiohttp-3.14.3-cp311-cp311-macosx_10_9_universal2.whl", hash = "sha256:e568e14940c09955aa51f4e645b6daa18a581c5dcfcd73744dcc86a856e3ced3", size = 764250, upload-time = "2026-07-23T01:52:48.525Z" }, - { url = "https://files.pythonhosted.org/packages/0e/da/f1b384465e51449d844056b75070461da03a9a23e6c1747003695bf4172a/aiohttp-3.14.3-cp311-cp311-macosx_10_9_x86_64.whl", hash = "sha256:54cfcdee2770dac994417cbb0ee1f3eb0e7cb6b30c79bf44f2c02ff79ec5124a", size = 516281, upload-time = "2026-07-23T01:52:51.047Z" }, - { url = "https://files.pythonhosted.org/packages/b9/3f/01264f820ee2e3712a827892b1cd6ff80f3300c1fcbffbb45714a915d47a/aiohttp-3.14.3-cp311-cp311-macosx_11_0_arm64.whl", hash = "sha256:21c016079415ed3fd676963e9793700a566d85dbbd6bfc564b9b2d209147dcc8", size = 514742, upload-time = "2026-07-23T01:52:53.779Z" }, - { url = "https://files.pythonhosted.org/packages/9e/8d/a71c6f2db52ac1ed142b133f7feddaa6b70539c3f4de24d7e226c95b794c/aiohttp-3.14.3-cp311-cp311-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:d6088ec9894113802bddb3c09e974929aed2c7b3a8c456219b8aab4481f1a239", size = 1780613, upload-time = "2026-07-23T01:52:56.948Z" }, - { url = "https://files.pythonhosted.org/packages/a5/11/3dd9b3fb3a170f6ec9011b5291d876a6fab4086714c9e158600edf01b4fd/aiohttp-3.14.3-cp311-cp311-manylinux2014_armv7l.manylinux_2_17_armv7l.manylinux_2_31_armv7l.whl", hash = "sha256:16ea7e24c309fb7c0bbd505d149abe4fe4dccfb8db911db7dbec0921bc889a6f", size = 1737688, upload-time = "2026-07-23T01:52:59.294Z" }, - { url = "https://files.pythonhosted.org/packages/6d/3e/834c26918be7d88068822b40e0db30fca50b5f4fe79104aa16a93f1d74e6/aiohttp-3.14.3-cp311-cp311-manylinux2014_ppc64le.manylinux_2_17_ppc64le.manylinux_2_28_ppc64le.whl", hash = "sha256:56f355e79f71aef2a85c80305cc915f894b170dba76de5fe84f6351939b83c06", size = 1845742, upload-time = "2026-07-23T01:53:01.641Z" }, - { url = "https://files.pythonhosted.org/packages/cc/c9/49ab8572df7d66bc13d11e31f781292badb04180dd87ba98733066c6aed7/aiohttp-3.14.3-cp311-cp311-manylinux2014_s390x.manylinux_2_17_s390x.manylinux_2_28_s390x.whl", hash = "sha256:18c441d0a8fca6de8d1f546849b9f0ab20d435993e2c5b59562b2fae6be2f929", size = 1928412, upload-time = "2026-07-23T01:53:04.018Z" }, - { url = "https://files.pythonhosted.org/packages/a5/b9/2b8f0c0ce09c87a1daf80fd483431b56b1435d3f62789bc86f572e1245de/aiohttp-3.14.3-cp311-cp311-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:53e7b4ce82b54a8bcc71b3b67a5cbd177ca1d7f592cbc92cd38b7349f73482db", size = 1786220, upload-time = "2026-07-23T01:53:06.481Z" }, - { url = "https://files.pythonhosted.org/packages/85/00/9c45f81de11710460edfa1dc81317b6e882703b160926c879a9d20da9fcc/aiohttp-3.14.3-cp311-cp311-manylinux_2_31_riscv64.manylinux_2_39_riscv64.whl", hash = "sha256:f55119f7bf25f49ed210f6096090715da24f2943c62102448915fde3c62877ce", size = 1637231, upload-time = "2026-07-23T01:53:10.258Z" }, - { url = "https://files.pythonhosted.org/packages/19/ce/967d628e910756f3539c6107cb7844a1b69440dcb3029a5ee7871b09ab63/aiohttp-3.14.3-cp311-cp311-musllinux_1_2_aarch64.whl", hash = "sha256:9aa6e61fdf20105c4144e755bd586008ff450791d67b1c8146fdc15959c4d51c", size = 1753161, upload-time = "2026-07-23T01:53:13.817Z" }, - { url = "https://files.pythonhosted.org/packages/11/b2/0c3d4114f0aee4f580f5b3b4eb71b24d7a23b834ea506a4dfebe76513f35/aiohttp-3.14.3-cp311-cp311-musllinux_1_2_armv7l.whl", hash = "sha256:ccd4893707b3e2a13e39c90d43cf80edf2e4d0457935bcc103bf2346214c3f15", size = 1756356, upload-time = "2026-07-23T01:53:16.211Z" }, - { url = "https://files.pythonhosted.org/packages/63/5d/99e7d91c82f1399d1ae2a854e080bd1493fbc31e5e959dbc4ec33dac3bec/aiohttp-3.14.3-cp311-cp311-musllinux_1_2_ppc64le.whl", hash = "sha256:b2466434105a4e03113c36ec775cc2ebe6676b62eae326fa670bb607ef788c1c", size = 1819846, upload-time = "2026-07-23T01:53:18.289Z" }, - { url = "https://files.pythonhosted.org/packages/ad/05/d5e1cb6480eeffd3f901d40a2c5e2d1e7effdc797837da3b490272699f13/aiohttp-3.14.3-cp311-cp311-musllinux_1_2_riscv64.whl", hash = "sha256:ba59d59aba08ac02fc03b0c8983ccd5ee39a199d0552ce9e6d2b4845b34d59ae", size = 1628531, upload-time = "2026-07-23T01:53:23.86Z" }, - { url = "https://files.pythonhosted.org/packages/c9/90/b934682bcaefae18a9e04f3dff5b68522ba810906358ae5029b68110ea3b/aiohttp-3.14.3-cp311-cp311-musllinux_1_2_s390x.whl", hash = "sha256:ed099d105449c4f9e84f24af203cd131349d4761d8813fa7e02c32e7128cd910", size = 1832712, upload-time = "2026-07-23T01:53:27.551Z" }, - { url = "https://files.pythonhosted.org/packages/21/df/6061679faaf81fac746e7307c7adb71e858071a5d34c27583afefc64f543/aiohttp-3.14.3-cp311-cp311-musllinux_1_2_x86_64.whl", hash = "sha256:152516815ef926786a0b6ae2b8f1fd2e0c71582dee0b435636865316fd4891b7", size = 1775014, upload-time = "2026-07-23T01:53:30.223Z" }, - { url = "https://files.pythonhosted.org/packages/8a/1d/f854878bbc69b88faefe924b619a34a6f59ec05fd387c77690667eaa75eb/aiohttp-3.14.3-cp311-cp311-win32.whl", hash = "sha256:a4af35c443e0b1a1bd6a8af3f3485d7fda15c142751a00f3ff8090f0b93346fa", size = 456006, upload-time = "2026-07-23T01:53:34.97Z" }, - { url = "https://files.pythonhosted.org/packages/73/0c/2af9d1674baccd1dbd47282a93d660a22e57ef6167c856deb24b4214fbab/aiohttp-3.14.3-cp311-cp311-win_amd64.whl", hash = "sha256:e1e74298bab6ee0d6e749ed4fd1901c7e604bdda32c03d787a2cc71c46d0433d", size = 481069, upload-time = "2026-07-23T01:53:39.673Z" }, - { url = "https://files.pythonhosted.org/packages/8e/76/88401ff3fc95e85c5fc38d588f36f55e61ecb64343b2bc8d69326f453cc0/aiohttp-3.14.3-cp311-cp311-win_arm64.whl", hash = "sha256:03cd2bde3d7f085b64e549c985f4bb928cad7e8ecf5323bfca320db548d81b39", size = 453021, upload-time = "2026-07-23T01:53:43.749Z" }, - { url = "https://files.pythonhosted.org/packages/18/d4/eb96299230e20acf2efae207cb8d69051f1f68e357e5ea5e479bf6fb097a/aiohttp-3.14.3-cp312-cp312-macosx_10_13_universal2.whl", hash = "sha256:39aded8c7f3b935b54aab1d8d73c70ec0ee2d3ec3b943e0e86611bc150ba47f5", size = 754690, upload-time = "2026-07-23T01:53:47.332Z" }, - { url = "https://files.pythonhosted.org/packages/88/11/e7a70a209eb9a067c0d3212b518a0134e3484f5178c7533878b6b514d469/aiohttp-3.14.3-cp312-cp312-macosx_10_13_x86_64.whl", hash = "sha256:5bcb6ff3fdab1258a192679ff1a05d44f59626430aa05cd1a9d2447423599228", size = 509484, upload-time = "2026-07-23T01:53:51.159Z" }, - { url = "https://files.pythonhosted.org/packages/30/07/4bbc222cc8dbe31d4c3e8a5baad2286e4d42026ac0c570027b89afce6344/aiohttp-3.14.3-cp312-cp312-macosx_11_0_arm64.whl", hash = "sha256:617105e2c3018ee38d0c8ce5ee3c84f621a6d8b9f723202aacaff28449ca91ee", size = 511949, upload-time = "2026-07-23T01:53:55.083Z" }, - { url = "https://files.pythonhosted.org/packages/54/b9/42e74c46b7b7c794b995bbc1f573fb48950c38b19d8600c62a6804ee2d67/aiohttp-3.14.3-cp312-cp312-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:f631fe87a6f30df5fbe6d79640b25e4cffb38c31c7fb6f10871517b84b0f8c1a", size = 1765282, upload-time = "2026-07-23T01:53:59.662Z" }, - { url = "https://files.pythonhosted.org/packages/6b/ed/62bc4d74363ad346d518e0720363a949f63e2e23439a79eb5813d4d29bb3/aiohttp-3.14.3-cp312-cp312-manylinux2014_armv7l.manylinux_2_17_armv7l.manylinux_2_31_armv7l.whl", hash = "sha256:a94dbaae5ae27bd849c93570669bff91e0510f33a80805738e3de72a7be0447b", size = 1741511, upload-time = "2026-07-23T01:54:04.063Z" }, - { url = "https://files.pythonhosted.org/packages/d0/9f/181e8a8bc79e47d13c7fc4540bd7a3b729d9505609c61f392a8dd2fbfe55/aiohttp-3.14.3-cp312-cp312-manylinux2014_ppc64le.manylinux_2_17_ppc64le.manylinux_2_28_ppc64le.whl", hash = "sha256:8f2f1c4c032c7cedd7d8da6f54c97b70266c6570c3108d3fdffee7188bb70529", size = 1810680, upload-time = "2026-07-23T01:54:09.882Z" }, - { url = "https://files.pythonhosted.org/packages/5c/9a/dec94d6ad694552fe3424e3f1928d7a606a5d9d9433a04e7ecdd9d38ae7f/aiohttp-3.14.3-cp312-cp312-manylinux2014_s390x.manylinux_2_17_s390x.manylinux_2_28_s390x.whl", hash = "sha256:ea05e1f97ceea523942d9b2a7d7c0359d781d683d6b043f5943a602b14da4787", size = 1905646, upload-time = "2026-07-23T01:54:13.475Z" }, - { url = "https://files.pythonhosted.org/packages/52/b7/7cd31f29d6055bd711ae6e669367fba6f5ae9de463910a793e30556a8db7/aiohttp-3.14.3-cp312-cp312-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:543906c127fb1d929b95076db19b83fa2d46751006ff1e23b093aa5ac4d8db42", size = 1792122, upload-time = "2026-07-23T01:54:15.752Z" }, - { url = "https://files.pythonhosted.org/packages/66/73/10b1ef93afa61f4963c746257b70ced619cf31a4798671de5fdb2608501d/aiohttp-3.14.3-cp312-cp312-manylinux_2_31_riscv64.manylinux_2_39_riscv64.whl", hash = "sha256:0a5ff2dfbb9ce645fa5b8ef3e02c6c0b9cc3f6030ff863d0c51fffc50cb5541b", size = 1591127, upload-time = "2026-07-23T01:54:19.489Z" }, - { url = "https://files.pythonhosted.org/packages/49/ed/3b203fa6de1b338c14acdc06bf6ca9b043b7944f005966958c2ced932cde/aiohttp-3.14.3-cp312-cp312-musllinux_1_2_aarch64.whl", hash = "sha256:041badb8f84396357c4d3ad26de6afd7a32b112f43d3c63045c0c8278cfd2043", size = 1725210, upload-time = "2026-07-23T01:54:24.129Z" }, - { url = "https://files.pythonhosted.org/packages/28/b7/1c2aab8c706436dcc28598452488ac9cd7c409da815237c28c27d58993e6/aiohttp-3.14.3-cp312-cp312-musllinux_1_2_armv7l.whl", hash = "sha256:530125ee1163c4219af35dc3aa1206e541e7b31b6efc1a3f93b70a136f65d427", size = 1764848, upload-time = "2026-07-23T01:54:27.973Z" }, - { url = "https://files.pythonhosted.org/packages/54/50/94c28f08b131c4bf10984ea2c7a536c9920608bb2d6e7f95642c30cc87b7/aiohttp-3.14.3-cp312-cp312-musllinux_1_2_ppc64le.whl", hash = "sha256:c8653fd547c93a61aadc612007790f5555cdd18946fa48cf45e26d8ea4ea473d", size = 1777102, upload-time = "2026-07-23T01:54:31.775Z" }, - { url = "https://files.pythonhosted.org/packages/13/d4/e7d09ba7d345fb2d74440fd2fa033c5e079fac05552927705986f41a364f/aiohttp-3.14.3-cp312-cp312-musllinux_1_2_riscv64.whl", hash = "sha256:89176250f686cb9853c0fb7ead90e639e915b84a6f43eedc2a4e7ec21f1037f0", size = 1580205, upload-time = "2026-07-23T01:54:34.518Z" }, - { url = "https://files.pythonhosted.org/packages/a3/84/072a91d68e1e1eb587985b54baab94221277f877e8ef274fc213a0ceae28/aiohttp-3.14.3-cp312-cp312-musllinux_1_2_s390x.whl", hash = "sha256:3a26434dafe408229ff3403458ca58de24fb51936504decac49ce6755f77e59d", size = 1797219, upload-time = "2026-07-23T01:54:36.995Z" }, - { url = "https://files.pythonhosted.org/packages/e0/eb/aad34e897e668424d6e995da5dff8a4a09af93363d3392488772957a63aa/aiohttp-3.14.3-cp312-cp312-musllinux_1_2_x86_64.whl", hash = "sha256:d1558173930a5a8d3069cee5c92fc91c87c4dbcb099debbb3622053717145a19", size = 1768629, upload-time = "2026-07-23T01:54:40.103Z" }, - { url = "https://files.pythonhosted.org/packages/b6/2b/6bb88ddba0fecd9122aa3ebcad25996cf6c083a4a7040dbb3a4f97972af6/aiohttp-3.14.3-cp312-cp312-win32.whl", hash = "sha256:16100ad3ab8d649fdfbee87602d9d2dcdca9df0b9eda8a1b5fdc0d41f96da559", size = 451481, upload-time = "2026-07-23T01:54:42.547Z" }, - { url = "https://files.pythonhosted.org/packages/76/9b/f2f8f108da17ecef2cc3efc424e8b7ad3782b1a8360f7b8eae8ced84f6ea/aiohttp-3.14.3-cp312-cp312-win_amd64.whl", hash = "sha256:33a2d7c28d33797a2e99923dffa63f83d908a19b6bf26cfe80fa790aa5e1a75a", size = 476845, upload-time = "2026-07-23T01:54:44.853Z" }, - { url = "https://files.pythonhosted.org/packages/3e/44/28dac80a8941b604f4da10ce21097614ca1bf905ce93dca28d8d7de9c1e7/aiohttp-3.14.3-cp312-cp312-win_arm64.whl", hash = "sha256:362a3fd481769cac1a824514bcd86fda51c65e8fe6e051099e008fddde6db17c", size = 448050, upload-time = "2026-07-23T01:54:47.087Z" }, +sdist = { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/aiohttp/3.14.3/aiohttp-3.14.3.tar.gz", hash = "sha256:9491196535a88924a60afd5b5f434b5b203b6cc616250878dbdb223a8f7844bc", size = 7971213, upload-time = "2026-07-23T01:57:27.037Z" } +wheels = [ + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/aiohttp/3.14.3/aiohttp-3.14.3-cp311-cp311-macosx_10_9_universal2.whl", hash = "sha256:e568e14940c09955aa51f4e645b6daa18a581c5dcfcd73744dcc86a856e3ced3", size = 764250, upload-time = "2026-07-23T01:52:48.525Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/aiohttp/3.14.3/aiohttp-3.14.3-cp311-cp311-macosx_10_9_x86_64.whl", hash = "sha256:54cfcdee2770dac994417cbb0ee1f3eb0e7cb6b30c79bf44f2c02ff79ec5124a", size = 516281, upload-time = "2026-07-23T01:52:51.047Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/aiohttp/3.14.3/aiohttp-3.14.3-cp311-cp311-macosx_11_0_arm64.whl", hash = "sha256:21c016079415ed3fd676963e9793700a566d85dbbd6bfc564b9b2d209147dcc8", size = 514742, upload-time = "2026-07-23T01:52:53.779Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/aiohttp/3.14.3/aiohttp-3.14.3-cp311-cp311-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:d6088ec9894113802bddb3c09e974929aed2c7b3a8c456219b8aab4481f1a239", size = 1780613, upload-time = "2026-07-23T01:52:56.948Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/aiohttp/3.14.3/aiohttp-3.14.3-cp311-cp311-manylinux2014_armv7l.manylinux_2_17_armv7l.manylinux_2_31_armv7l.whl", hash = "sha256:16ea7e24c309fb7c0bbd505d149abe4fe4dccfb8db911db7dbec0921bc889a6f", size = 1737688, upload-time = "2026-07-23T01:52:59.294Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/aiohttp/3.14.3/aiohttp-3.14.3-cp311-cp311-manylinux2014_ppc64le.manylinux_2_17_ppc64le.manylinux_2_28_ppc64le.whl", hash = "sha256:56f355e79f71aef2a85c80305cc915f894b170dba76de5fe84f6351939b83c06", size = 1845742, upload-time = "2026-07-23T01:53:01.641Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/aiohttp/3.14.3/aiohttp-3.14.3-cp311-cp311-manylinux2014_s390x.manylinux_2_17_s390x.manylinux_2_28_s390x.whl", hash = "sha256:18c441d0a8fca6de8d1f546849b9f0ab20d435993e2c5b59562b2fae6be2f929", size = 1928412, upload-time = "2026-07-23T01:53:04.018Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/aiohttp/3.14.3/aiohttp-3.14.3-cp311-cp311-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:53e7b4ce82b54a8bcc71b3b67a5cbd177ca1d7f592cbc92cd38b7349f73482db", size = 1786220, upload-time = "2026-07-23T01:53:06.481Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/aiohttp/3.14.3/aiohttp-3.14.3-cp311-cp311-manylinux_2_31_riscv64.manylinux_2_39_riscv64.whl", hash = "sha256:f55119f7bf25f49ed210f6096090715da24f2943c62102448915fde3c62877ce", size = 1637231, upload-time = "2026-07-23T01:53:10.258Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/aiohttp/3.14.3/aiohttp-3.14.3-cp311-cp311-musllinux_1_2_aarch64.whl", hash = "sha256:9aa6e61fdf20105c4144e755bd586008ff450791d67b1c8146fdc15959c4d51c", size = 1753161, upload-time = "2026-07-23T01:53:13.817Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/aiohttp/3.14.3/aiohttp-3.14.3-cp311-cp311-musllinux_1_2_armv7l.whl", hash = "sha256:ccd4893707b3e2a13e39c90d43cf80edf2e4d0457935bcc103bf2346214c3f15", size = 1756356, upload-time = "2026-07-23T01:53:16.211Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/aiohttp/3.14.3/aiohttp-3.14.3-cp311-cp311-musllinux_1_2_ppc64le.whl", hash = "sha256:b2466434105a4e03113c36ec775cc2ebe6676b62eae326fa670bb607ef788c1c", size = 1819846, upload-time = "2026-07-23T01:53:18.289Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/aiohttp/3.14.3/aiohttp-3.14.3-cp311-cp311-musllinux_1_2_riscv64.whl", hash = "sha256:ba59d59aba08ac02fc03b0c8983ccd5ee39a199d0552ce9e6d2b4845b34d59ae", size = 1628531, upload-time = "2026-07-23T01:53:23.86Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/aiohttp/3.14.3/aiohttp-3.14.3-cp311-cp311-musllinux_1_2_s390x.whl", hash = "sha256:ed099d105449c4f9e84f24af203cd131349d4761d8813fa7e02c32e7128cd910", size = 1832712, upload-time = "2026-07-23T01:53:27.551Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/aiohttp/3.14.3/aiohttp-3.14.3-cp311-cp311-musllinux_1_2_x86_64.whl", hash = "sha256:152516815ef926786a0b6ae2b8f1fd2e0c71582dee0b435636865316fd4891b7", size = 1775014, upload-time = "2026-07-23T01:53:30.223Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/aiohttp/3.14.3/aiohttp-3.14.3-cp311-cp311-win32.whl", hash = "sha256:a4af35c443e0b1a1bd6a8af3f3485d7fda15c142751a00f3ff8090f0b93346fa", size = 456006, upload-time = "2026-07-23T01:53:34.97Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/aiohttp/3.14.3/aiohttp-3.14.3-cp311-cp311-win_amd64.whl", hash = "sha256:e1e74298bab6ee0d6e749ed4fd1901c7e604bdda32c03d787a2cc71c46d0433d", size = 481069, upload-time = "2026-07-23T01:53:39.673Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/aiohttp/3.14.3/aiohttp-3.14.3-cp311-cp311-win_arm64.whl", hash = "sha256:03cd2bde3d7f085b64e549c985f4bb928cad7e8ecf5323bfca320db548d81b39", size = 453021, upload-time = "2026-07-23T01:53:43.749Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/aiohttp/3.14.3/aiohttp-3.14.3-cp312-cp312-macosx_10_13_universal2.whl", hash = "sha256:39aded8c7f3b935b54aab1d8d73c70ec0ee2d3ec3b943e0e86611bc150ba47f5", size = 754690, upload-time = "2026-07-23T01:53:47.332Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/aiohttp/3.14.3/aiohttp-3.14.3-cp312-cp312-macosx_10_13_x86_64.whl", hash = "sha256:5bcb6ff3fdab1258a192679ff1a05d44f59626430aa05cd1a9d2447423599228", size = 509484, upload-time = "2026-07-23T01:53:51.159Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/aiohttp/3.14.3/aiohttp-3.14.3-cp312-cp312-macosx_11_0_arm64.whl", hash = "sha256:617105e2c3018ee38d0c8ce5ee3c84f621a6d8b9f723202aacaff28449ca91ee", size = 511949, upload-time = "2026-07-23T01:53:55.083Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/aiohttp/3.14.3/aiohttp-3.14.3-cp312-cp312-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:f631fe87a6f30df5fbe6d79640b25e4cffb38c31c7fb6f10871517b84b0f8c1a", size = 1765282, upload-time = "2026-07-23T01:53:59.662Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/aiohttp/3.14.3/aiohttp-3.14.3-cp312-cp312-manylinux2014_armv7l.manylinux_2_17_armv7l.manylinux_2_31_armv7l.whl", hash = "sha256:a94dbaae5ae27bd849c93570669bff91e0510f33a80805738e3de72a7be0447b", size = 1741511, upload-time = "2026-07-23T01:54:04.063Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/aiohttp/3.14.3/aiohttp-3.14.3-cp312-cp312-manylinux2014_ppc64le.manylinux_2_17_ppc64le.manylinux_2_28_ppc64le.whl", hash = "sha256:8f2f1c4c032c7cedd7d8da6f54c97b70266c6570c3108d3fdffee7188bb70529", size = 1810680, upload-time = "2026-07-23T01:54:09.882Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/aiohttp/3.14.3/aiohttp-3.14.3-cp312-cp312-manylinux2014_s390x.manylinux_2_17_s390x.manylinux_2_28_s390x.whl", hash = "sha256:ea05e1f97ceea523942d9b2a7d7c0359d781d683d6b043f5943a602b14da4787", size = 1905646, upload-time = "2026-07-23T01:54:13.475Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/aiohttp/3.14.3/aiohttp-3.14.3-cp312-cp312-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:543906c127fb1d929b95076db19b83fa2d46751006ff1e23b093aa5ac4d8db42", size = 1792122, upload-time = "2026-07-23T01:54:15.752Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/aiohttp/3.14.3/aiohttp-3.14.3-cp312-cp312-manylinux_2_31_riscv64.manylinux_2_39_riscv64.whl", hash = "sha256:0a5ff2dfbb9ce645fa5b8ef3e02c6c0b9cc3f6030ff863d0c51fffc50cb5541b", size = 1591127, upload-time = "2026-07-23T01:54:19.489Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/aiohttp/3.14.3/aiohttp-3.14.3-cp312-cp312-musllinux_1_2_aarch64.whl", hash = "sha256:041badb8f84396357c4d3ad26de6afd7a32b112f43d3c63045c0c8278cfd2043", size = 1725210, upload-time = "2026-07-23T01:54:24.129Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/aiohttp/3.14.3/aiohttp-3.14.3-cp312-cp312-musllinux_1_2_armv7l.whl", hash = "sha256:530125ee1163c4219af35dc3aa1206e541e7b31b6efc1a3f93b70a136f65d427", size = 1764848, upload-time = "2026-07-23T01:54:27.973Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/aiohttp/3.14.3/aiohttp-3.14.3-cp312-cp312-musllinux_1_2_ppc64le.whl", hash = "sha256:c8653fd547c93a61aadc612007790f5555cdd18946fa48cf45e26d8ea4ea473d", size = 1777102, upload-time = "2026-07-23T01:54:31.775Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/aiohttp/3.14.3/aiohttp-3.14.3-cp312-cp312-musllinux_1_2_riscv64.whl", hash = "sha256:89176250f686cb9853c0fb7ead90e639e915b84a6f43eedc2a4e7ec21f1037f0", size = 1580205, upload-time = "2026-07-23T01:54:34.518Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/aiohttp/3.14.3/aiohttp-3.14.3-cp312-cp312-musllinux_1_2_s390x.whl", hash = "sha256:3a26434dafe408229ff3403458ca58de24fb51936504decac49ce6755f77e59d", size = 1797219, upload-time = "2026-07-23T01:54:36.995Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/aiohttp/3.14.3/aiohttp-3.14.3-cp312-cp312-musllinux_1_2_x86_64.whl", hash = "sha256:d1558173930a5a8d3069cee5c92fc91c87c4dbcb099debbb3622053717145a19", size = 1768629, upload-time = "2026-07-23T01:54:40.103Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/aiohttp/3.14.3/aiohttp-3.14.3-cp312-cp312-win32.whl", hash = "sha256:16100ad3ab8d649fdfbee87602d9d2dcdca9df0b9eda8a1b5fdc0d41f96da559", size = 451481, upload-time = "2026-07-23T01:54:42.547Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/aiohttp/3.14.3/aiohttp-3.14.3-cp312-cp312-win_amd64.whl", hash = "sha256:33a2d7c28d33797a2e99923dffa63f83d908a19b6bf26cfe80fa790aa5e1a75a", size = 476845, upload-time = "2026-07-23T01:54:44.853Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/aiohttp/3.14.3/aiohttp-3.14.3-cp312-cp312-win_arm64.whl", hash = "sha256:362a3fd481769cac1a824514bcd86fda51c65e8fe6e051099e008fddde6db17c", size = 448050, upload-time = "2026-07-23T01:54:47.087Z" }, ] [[package]] name = "aiosignal" version = "1.4.0" -source = { registry = "https://pypi.org/simple" } +source = { registry = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/" } dependencies = [ { name = "frozenlist" }, { name = "typing-extensions" }, ] -sdist = { url = "https://files.pythonhosted.org/packages/61/62/06741b579156360248d1ec624842ad0edf697050bbaf7c3e46394e106ad1/aiosignal-1.4.0.tar.gz", hash = "sha256:f47eecd9468083c2029cc99945502cb7708b082c232f9aca65da147157b251c7", size = 25007, upload-time = "2025-07-03T22:54:43.528Z" } +sdist = { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/aiosignal/1.4.0/aiosignal-1.4.0.tar.gz", hash = "sha256:f47eecd9468083c2029cc99945502cb7708b082c232f9aca65da147157b251c7", size = 25007, upload-time = "2025-07-03T22:54:43.528Z" } wheels = [ - { url = "https://files.pythonhosted.org/packages/fb/76/641ae371508676492379f16e2fa48f4e2c11741bd63c48be4b12a6b09cba/aiosignal-1.4.0-py3-none-any.whl", hash = "sha256:053243f8b92b990551949e63930a839ff0cf0b0ebbe0597b0f3fb19e1a0fe82e", size = 7490, upload-time = "2025-07-03T22:54:42.156Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/aiosignal/1.4.0/aiosignal-1.4.0-py3-none-any.whl", hash = "sha256:053243f8b92b990551949e63930a839ff0cf0b0ebbe0597b0f3fb19e1a0fe82e", size = 7490, upload-time = "2025-07-03T22:54:42.156Z" }, ] [[package]] name = "annotated-doc" version = "0.0.5" -source = { registry = "https://pypi.org/simple" } -sdist = { url = "https://files.pythonhosted.org/packages/5a/8e/38aa427ed5402449e226975b649c5dc73ccadfefeb95e6aecb8f8ea4b6b6/annotated_doc-0.0.5.tar.gz", hash = "sha256:c7e58ce09192557605d8bbd92836d7e1d520ac9580096042c0bfd197efacf1bb", size = 10758, upload-time = "2026-07-28T13:50:58.129Z" } +source = { registry = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/" } +sdist = { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/annotated-doc/0.0.5/annotated_doc-0.0.5.tar.gz", hash = "sha256:c7e58ce09192557605d8bbd92836d7e1d520ac9580096042c0bfd197efacf1bb", size = 10758, upload-time = "2026-07-28T13:50:58.129Z" } wheels = [ - { url = "https://files.pythonhosted.org/packages/3e/30/e900b21425a860e195f32e37657aa1f7c7f2b1bfb26f03ca209b90933c06/annotated_doc-0.0.5-py3-none-any.whl", hash = "sha256:117bac03a25ede5df5440e855b32d556049ca169ead221505badf432fed4b101", size = 5302, upload-time = "2026-07-28T13:50:57.239Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/annotated-doc/0.0.5/annotated_doc-0.0.5-py3-none-any.whl", hash = "sha256:117bac03a25ede5df5440e855b32d556049ca169ead221505badf432fed4b101", size = 5302, upload-time = "2026-07-28T13:50:57.239Z" }, ] [[package]] name = "annotated-types" version = "0.8.0" -source = { registry = "https://pypi.org/simple" } -sdist = { url = "https://files.pythonhosted.org/packages/5f/56/a8120250d128bed162cd73c76d45f6ef9991f3e068f62a8ee060afa3104a/annotated_types-0.8.0.tar.gz", hash = "sha256:13b2beaad985e05e2d6407ee4c4f35590b11f8d693a258a561055cac8f64cab7", size = 15893, upload-time = "2026-07-23T20:16:13.995Z" } +source = { registry = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/" } +sdist = { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/annotated-types/0.8.0/annotated_types-0.8.0.tar.gz", hash = "sha256:13b2beaad985e05e2d6407ee4c4f35590b11f8d693a258a561055cac8f64cab7", size = 15893, upload-time = "2026-07-23T20:16:13.995Z" } wheels = [ - { url = "https://files.pythonhosted.org/packages/99/91/8acff4f5e50511b911bbccb72b8628a49c68ce14148cd9f6431094859a90/annotated_types-0.8.0-py3-none-any.whl", hash = "sha256:f072f4d804ea359e4eaf198b1af7a8b0943881a87f31bb764f8bf219bb9419e0", size = 13427, upload-time = "2026-07-23T20:16:12.938Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/annotated-types/0.8.0/annotated_types-0.8.0-py3-none-any.whl", hash = "sha256:f072f4d804ea359e4eaf198b1af7a8b0943881a87f31bb764f8bf219bb9419e0", size = 13427, upload-time = "2026-07-23T20:16:12.938Z" }, ] [[package]] name = "anyio" version = "4.14.2" -source = { registry = "https://pypi.org/simple" } +source = { registry = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/" } dependencies = [ { name = "idna" }, { name = "typing-extensions" }, ] -sdist = { url = "https://files.pythonhosted.org/packages/61/cc/a381afa6efea9f496eff839d4a6a1aed3bfafc7b3ab4b0d1b243a12573dd/anyio-4.14.2.tar.gz", hash = "sha256:cfa139f3ed1a23ee8f88a145ddb5ac7605b8bbfd8592baacd7ce3d8bb4313c7f", size = 260176, upload-time = "2026-07-12T20:29:07.082Z" } +sdist = { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/anyio/4.14.2/anyio-4.14.2.tar.gz", hash = "sha256:cfa139f3ed1a23ee8f88a145ddb5ac7605b8bbfd8592baacd7ce3d8bb4313c7f", size = 260176, upload-time = "2026-07-12T20:29:07.082Z" } wheels = [ - { url = "https://files.pythonhosted.org/packages/da/35/f2287558c17e29fafc8ef3daf819bb9834061cfa43bff8014f7df7f63bdc/anyio-4.14.2-py3-none-any.whl", hash = "sha256:9f505dda5ac9f0c8309b5e8bd445a8c2bf7246f3ce950121e45ea15bc41d1494", size = 125813, upload-time = "2026-07-12T20:29:05.763Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/anyio/4.14.2/anyio-4.14.2-py3-none-any.whl", hash = "sha256:9f505dda5ac9f0c8309b5e8bd445a8c2bf7246f3ce950121e45ea15bc41d1494", size = 125813, upload-time = "2026-07-12T20:29:05.763Z" }, ] [[package]] name = "attrs" version = "26.1.0" -source = { registry = "https://pypi.org/simple" } -sdist = { url = "https://files.pythonhosted.org/packages/9a/8e/82a0fe20a541c03148528be8cac2408564a6c9a0cc7e9171802bc1d26985/attrs-26.1.0.tar.gz", hash = "sha256:d03ceb89cb322a8fd706d4fb91940737b6642aa36998fe130a9bc96c985eff32", size = 952055, upload-time = "2026-03-19T14:22:25.026Z" } +source = { registry = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/" } +sdist = { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/attrs/26.1.0/attrs-26.1.0.tar.gz", hash = "sha256:d03ceb89cb322a8fd706d4fb91940737b6642aa36998fe130a9bc96c985eff32", size = 952055, upload-time = "2026-03-19T14:22:25.026Z" } wheels = [ - { url = "https://files.pythonhosted.org/packages/64/b4/17d4b0b2a2dc85a6df63d1157e028ed19f90d4cd97c36717afef2bc2f395/attrs-26.1.0-py3-none-any.whl", hash = "sha256:c647aa4a12dfbad9333ca4e71fe62ddc36f4e63b2d260a37a8b83d2f043ac309", size = 67548, upload-time = "2026-03-19T14:22:23.645Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/attrs/26.1.0/attrs-26.1.0-py3-none-any.whl", hash = "sha256:c647aa4a12dfbad9333ca4e71fe62ddc36f4e63b2d260a37a8b83d2f043ac309", size = 67548, upload-time = "2026-03-19T14:22:23.645Z" }, ] [[package]] name = "cbor2" version = "6.1.4" -source = { registry = "https://pypi.org/simple" } -sdist = { url = "https://files.pythonhosted.org/packages/c6/14/b02446bacfe44351b1689c04937ade007588f44570431880a6937e525e6c/cbor2-6.1.4.tar.gz", hash = "sha256:01ecc79a28f33d17331943ce508fc1e21f4b06553c73f874f4c77120d72b2ef9", size = 90840, upload-time = "2026-08-01T20:41:39.797Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/a7/84/1e363301c06f509963d134f5479e82b3ade87fb1495ddacf9bf7ff24ac42/cbor2-6.1.4-cp311-cp311-macosx_11_0_arm64.whl", hash = "sha256:8156fdeb73c3ff6c8cf67ad414fb5c887cd708ff0af6d61f62629f41cb4c17b2", size = 414947, upload-time = "2026-08-01T20:40:37.405Z" }, - { url = "https://files.pythonhosted.org/packages/8d/96/d8e1ed3e79ea20a3423a96b5c89ce794fa02cb428e4429e601f8ebcbac7c/cbor2-6.1.4-cp311-cp311-manylinux_2_28_aarch64.whl", hash = "sha256:e1fe2d62c50df290576280b18247ec63486f78be73e285bae269c2456c6ddff0", size = 457343, upload-time = "2026-08-01T20:40:38.868Z" }, - { url = "https://files.pythonhosted.org/packages/d5/0c/5796c2ed2dcd0696fc4abedf0ea0dfd5361b3f022a311481f977fa51b2b8/cbor2-6.1.4-cp311-cp311-manylinux_2_28_x86_64.whl", hash = "sha256:c204a75f91f8cd9ed0881f6b88ec395c59aeac9fcf4d08155e7f899db2a1c46e", size = 464314, upload-time = "2026-08-01T20:40:40.63Z" }, - { url = "https://files.pythonhosted.org/packages/b1/88/de524c6c2c91b740e5df6e6955a113fb616e979b26fd2e6a0693082d36e0/cbor2-6.1.4-cp311-cp311-musllinux_1_2_aarch64.whl", hash = "sha256:28fa5db05a7eae8fd80709959988d8a7f12838c6d4e5c58ec951414058641195", size = 523053, upload-time = "2026-08-01T20:40:42.602Z" }, - { url = "https://files.pythonhosted.org/packages/84/07/cb5fd92834633508d680a5b5695aeaf99d33ca0bdc5b844550d538f335b0/cbor2-6.1.4-cp311-cp311-musllinux_1_2_x86_64.whl", hash = "sha256:316e217a496640418d3137483279d0e70053b000cdd4b52a4dbf20ea478bc40a", size = 532177, upload-time = "2026-08-01T20:40:44.058Z" }, - { url = "https://files.pythonhosted.org/packages/c9/19/be98721365edfe6fc23e6bcd1385afa0e960b247c5f0b50bb67f5d05e2d9/cbor2-6.1.4-cp311-cp311-win32.whl", hash = "sha256:4903f24e0f9087275a0b6606c8b0aa586277001d51e4844fcdbc5b7211330aa8", size = 281660, upload-time = "2026-08-01T20:40:45.761Z" }, - { url = "https://files.pythonhosted.org/packages/16/23/d54f679d4b155918f5a0879dab78203ce4fd514d311b7cfeba27dafe480b/cbor2-6.1.4-cp311-cp311-win_amd64.whl", hash = "sha256:5b99305d4013867e059f147752b95f728680682ab03d75a3f4dcfbb270d8dfe9", size = 303207, upload-time = "2026-08-01T20:40:47.293Z" }, - { url = "https://files.pythonhosted.org/packages/53/3c/b3839d6213c88b249ba860525df05ff18b27bdc28ebc09cb1547790f001a/cbor2-6.1.4-cp311-cp311-win_arm64.whl", hash = "sha256:bd20ecc5c8ece24db952e48a91c8c47319eaa6358af707c85ac2bb388a79abc8", size = 296123, upload-time = "2026-08-01T20:40:48.808Z" }, - { url = "https://files.pythonhosted.org/packages/2e/76/fb64293c19cafb860060310c57b768fd9cfb7cf592449660b756538cc116/cbor2-6.1.4-cp312-cp312-macosx_11_0_arm64.whl", hash = "sha256:1fc15061553e4494dc10883237501e3402c645fe509248dd698e1faf2460d68b", size = 404608, upload-time = "2026-08-01T20:40:50.219Z" }, - { url = "https://files.pythonhosted.org/packages/96/ac/f58b3bafce7c86ada2ad8eaf189453136d2cf5bae526ea0540e1b9bc9d06/cbor2-6.1.4-cp312-cp312-manylinux_2_28_aarch64.whl", hash = "sha256:d9ada5a6ccfbb8ea7a3aa2aeb028421b52d8e0cd9323f0a2aeaa9c09d25fbce2", size = 449851, upload-time = "2026-08-01T20:40:51.725Z" }, - { url = "https://files.pythonhosted.org/packages/f0/a5/10c6c126d59b07f2bd005094dd12a20afa46146f7e2673ed6f61a57641a7/cbor2-6.1.4-cp312-cp312-manylinux_2_28_x86_64.whl", hash = "sha256:310f3dfb296ba48fe9b63c5cf26e691e3548a1eae6901d2f0c18e941d151f220", size = 461193, upload-time = "2026-08-01T20:40:53.446Z" }, - { url = "https://files.pythonhosted.org/packages/15/e4/4445e6237088d1cca3b8536daeb90d6b4e23776de5609c9fa46773874757/cbor2-6.1.4-cp312-cp312-musllinux_1_2_aarch64.whl", hash = "sha256:5e6c76004d674ad1c620660cb0bc5a8a0b72a5d8c7b70926d8e09e6d7e87332f", size = 516937, upload-time = "2026-08-01T20:40:54.952Z" }, - { url = "https://files.pythonhosted.org/packages/8c/87/9c0959510f7a402e5995c81ccfd82cb9f314140dc0cce88c12836e5b93f1/cbor2-6.1.4-cp312-cp312-musllinux_1_2_x86_64.whl", hash = "sha256:32a4663425fbca4a4a7aa918eb5789d844c406439e58424cf34511f79f559242", size = 529229, upload-time = "2026-08-01T20:40:56.365Z" }, - { url = "https://files.pythonhosted.org/packages/91/8e/6811e4ee84203ac657f6f461a37c7c9ba0287bde80eb83c7971e9b3fe156/cbor2-6.1.4-cp312-cp312-win32.whl", hash = "sha256:2310f07db3f9ba26f2a623774ff9f3dc7185af54f732ea119785a6b1bf7e1e7e", size = 278810, upload-time = "2026-08-01T20:40:57.76Z" }, - { url = "https://files.pythonhosted.org/packages/da/27/87440788fc0d9513534c3c699238e2a9ca6010f8cb72e9c203b7af20a9f6/cbor2-6.1.4-cp312-cp312-win_amd64.whl", hash = "sha256:cc8cd300e236e9797b2e1ce306109dc481fcccf78bfa2682bf36d99e6eab1ec6", size = 299971, upload-time = "2026-08-01T20:40:59.256Z" }, - { url = "https://files.pythonhosted.org/packages/23/f9/77981e6e63092de19d7306a09a12b0eb3fd2907dc22c10dd5d389eb27faf/cbor2-6.1.4-cp312-cp312-win_arm64.whl", hash = "sha256:553a46bda7d09552631a714e22b91e6ff2c867ecd91511596ce290d8879b8d5b", size = 290662, upload-time = "2026-08-01T20:41:00.89Z" }, +source = { registry = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/" } +sdist = { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/cbor2/6.1.4/cbor2-6.1.4.tar.gz", hash = "sha256:01ecc79a28f33d17331943ce508fc1e21f4b06553c73f874f4c77120d72b2ef9", size = 90840, upload-time = "2026-08-01T20:41:39.797Z" } +wheels = [ + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/cbor2/6.1.4/cbor2-6.1.4-cp311-cp311-macosx_11_0_arm64.whl", hash = "sha256:8156fdeb73c3ff6c8cf67ad414fb5c887cd708ff0af6d61f62629f41cb4c17b2", size = 414947, upload-time = "2026-08-01T20:40:37.405Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/cbor2/6.1.4/cbor2-6.1.4-cp311-cp311-manylinux_2_28_aarch64.whl", hash = "sha256:e1fe2d62c50df290576280b18247ec63486f78be73e285bae269c2456c6ddff0", size = 457343, upload-time = "2026-08-01T20:40:38.868Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/cbor2/6.1.4/cbor2-6.1.4-cp311-cp311-manylinux_2_28_x86_64.whl", hash = "sha256:c204a75f91f8cd9ed0881f6b88ec395c59aeac9fcf4d08155e7f899db2a1c46e", size = 464314, upload-time = "2026-08-01T20:40:40.63Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/cbor2/6.1.4/cbor2-6.1.4-cp311-cp311-musllinux_1_2_aarch64.whl", hash = "sha256:28fa5db05a7eae8fd80709959988d8a7f12838c6d4e5c58ec951414058641195", size = 523053, upload-time = "2026-08-01T20:40:42.602Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/cbor2/6.1.4/cbor2-6.1.4-cp311-cp311-musllinux_1_2_x86_64.whl", hash = "sha256:316e217a496640418d3137483279d0e70053b000cdd4b52a4dbf20ea478bc40a", size = 532177, upload-time = "2026-08-01T20:40:44.058Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/cbor2/6.1.4/cbor2-6.1.4-cp311-cp311-win32.whl", hash = "sha256:4903f24e0f9087275a0b6606c8b0aa586277001d51e4844fcdbc5b7211330aa8", size = 281660, upload-time = "2026-08-01T20:40:45.761Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/cbor2/6.1.4/cbor2-6.1.4-cp311-cp311-win_amd64.whl", hash = "sha256:5b99305d4013867e059f147752b95f728680682ab03d75a3f4dcfbb270d8dfe9", size = 303207, upload-time = "2026-08-01T20:40:47.293Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/cbor2/6.1.4/cbor2-6.1.4-cp311-cp311-win_arm64.whl", hash = "sha256:bd20ecc5c8ece24db952e48a91c8c47319eaa6358af707c85ac2bb388a79abc8", size = 296123, upload-time = "2026-08-01T20:40:48.808Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/cbor2/6.1.4/cbor2-6.1.4-cp312-cp312-macosx_11_0_arm64.whl", hash = "sha256:1fc15061553e4494dc10883237501e3402c645fe509248dd698e1faf2460d68b", size = 404608, upload-time = "2026-08-01T20:40:50.219Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/cbor2/6.1.4/cbor2-6.1.4-cp312-cp312-manylinux_2_28_aarch64.whl", hash = "sha256:d9ada5a6ccfbb8ea7a3aa2aeb028421b52d8e0cd9323f0a2aeaa9c09d25fbce2", size = 449851, upload-time = "2026-08-01T20:40:51.725Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/cbor2/6.1.4/cbor2-6.1.4-cp312-cp312-manylinux_2_28_x86_64.whl", hash = "sha256:310f3dfb296ba48fe9b63c5cf26e691e3548a1eae6901d2f0c18e941d151f220", size = 461193, upload-time = "2026-08-01T20:40:53.446Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/cbor2/6.1.4/cbor2-6.1.4-cp312-cp312-musllinux_1_2_aarch64.whl", hash = "sha256:5e6c76004d674ad1c620660cb0bc5a8a0b72a5d8c7b70926d8e09e6d7e87332f", size = 516937, upload-time = "2026-08-01T20:40:54.952Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/cbor2/6.1.4/cbor2-6.1.4-cp312-cp312-musllinux_1_2_x86_64.whl", hash = "sha256:32a4663425fbca4a4a7aa918eb5789d844c406439e58424cf34511f79f559242", size = 529229, upload-time = "2026-08-01T20:40:56.365Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/cbor2/6.1.4/cbor2-6.1.4-cp312-cp312-win32.whl", hash = "sha256:2310f07db3f9ba26f2a623774ff9f3dc7185af54f732ea119785a6b1bf7e1e7e", size = 278810, upload-time = "2026-08-01T20:40:57.76Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/cbor2/6.1.4/cbor2-6.1.4-cp312-cp312-win_amd64.whl", hash = "sha256:cc8cd300e236e9797b2e1ce306109dc481fcccf78bfa2682bf36d99e6eab1ec6", size = 299971, upload-time = "2026-08-01T20:40:59.256Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/cbor2/6.1.4/cbor2-6.1.4-cp312-cp312-win_arm64.whl", hash = "sha256:553a46bda7d09552631a714e22b91e6ff2c867ecd91511596ce290d8879b8d5b", size = 290662, upload-time = "2026-08-01T20:41:00.89Z" }, ] [[package]] name = "certifi" version = "2026.7.22" -source = { registry = "https://pypi.org/simple" } -sdist = { url = "https://files.pythonhosted.org/packages/a3/c2/24167ea9858356b47a87a50d39908bfdb72ceeefe0041586e704e5376b3a/certifi-2026.7.22.tar.gz", hash = "sha256:741e2c3b351ddf169a738da9f2c048608ff7f2c5cc02f1ebc6b118bb090d5d55", size = 138112, upload-time = "2026-07-22T03:35:12.644Z" } +source = { registry = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/" } +sdist = { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/certifi/2026.7.22/certifi-2026.7.22.tar.gz", hash = "sha256:741e2c3b351ddf169a738da9f2c048608ff7f2c5cc02f1ebc6b118bb090d5d55", size = 138112, upload-time = "2026-07-22T03:35:12.644Z" } wheels = [ - { url = "https://files.pythonhosted.org/packages/0b/a7/71ac2cff56fec219ed242bb11b8efb69fcc4bec75db06fb7bfe35de520e6/certifi-2026.7.22-py3-none-any.whl", hash = "sha256:62f22742b58a1a33014a2b6b706588a8d7e2a88ae7bd1a6ebe8c992928483775", size = 136983, upload-time = "2026-07-22T03:35:11.276Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/certifi/2026.7.22/certifi-2026.7.22-py3-none-any.whl", hash = "sha256:62f22742b58a1a33014a2b6b706588a8d7e2a88ae7bd1a6ebe8c992928483775", size = 136983, upload-time = "2026-07-22T03:35:11.276Z" }, ] [[package]] name = "charset-normalizer" version = "3.5.1" -source = { registry = "https://pypi.org/simple" } -sdist = { url = "https://files.pythonhosted.org/packages/e5/3f/143b048436775b0f76ac3eec145c019e8173ccc2885c8f20319b996d5e83/charset_normalizer-3.5.1.tar.gz", hash = "sha256:6117b84ea48435e5356dc737f5121485c30920ba43375fa7b434fd753df0eac3", size = 171764, upload-time = "2026-08-15T08:20:44.807Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/6a/b6/034f6802e9c3f6418966cfabb7db8c9252cc2429c5098f41cc43af804149/charset_normalizer-3.5.1-cp311-cp311-macosx_10_9_universal2.whl", hash = "sha256:eda059b6bc8bc0812d626fd91a7ce01bf583df0a61296eff390fd94141a34e30", size = 363585, upload-time = "2026-08-15T08:16:46.646Z" }, - { url = "https://files.pythonhosted.org/packages/d5/fa/6a7e2a7c4b5451912b8c417732df79574354443592a88d616de03da66ae5/charset_normalizer-3.5.1-cp311-cp311-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:aa2bb0b37202dca27175591f761108b5d34096ade1191ffe4808bdf6b1571488", size = 251189, upload-time = "2026-08-15T08:16:48.287Z" }, - { url = "https://files.pythonhosted.org/packages/a4/c8/ab42b07cfd82e919f427fcfaa7c41abae8242833ad1aad66d42bae40b669/charset_normalizer-3.5.1-cp311-cp311-manylinux2014_armv7l.manylinux_2_17_armv7l.manylinux_2_31_armv7l.whl", hash = "sha256:0b2b1b3fa5670c127b246df1d0c059defd41f689a868a3b9d79df9b1cac42d22", size = 239724, upload-time = "2026-08-15T08:16:49.67Z" }, - { url = "https://files.pythonhosted.org/packages/e7/80/b9348b5d3041209f98b4cdad7655766369233f1d533f4f4f7558e9717bec/charset_normalizer-3.5.1-cp311-cp311-manylinux2014_ppc64le.manylinux_2_17_ppc64le.manylinux_2_28_ppc64le.whl", hash = "sha256:6e5e4d73d588ca5ed09df1b7dcd1b203d1df3c542e3f50d126c947d432b10731", size = 280078, upload-time = "2026-08-15T08:16:51.228Z" }, - { url = "https://files.pythonhosted.org/packages/82/38/083a24028304bc85bb9e376fed801178423dcbb67495f73b6ea0624e1894/charset_normalizer-3.5.1-cp311-cp311-manylinux2014_s390x.manylinux_2_17_s390x.manylinux_2_28_s390x.whl", hash = "sha256:b54e7e13267d49ffbfe68e25b3cbd774dab38fa37238f71265e91b36146eb21c", size = 276650, upload-time = "2026-08-15T08:16:52.625Z" }, - { url = "https://files.pythonhosted.org/packages/0d/35/731ac04aa0a097fc1c97f0994c375bdb230c6c96619db794208fe664e9ce/charset_normalizer-3.5.1-cp311-cp311-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:c7b742bf31c88566b4bb6335a7f393bb322e580b6bb98df7bd0c25e6e3519ce8", size = 262325, upload-time = "2026-08-15T08:16:54.085Z" }, - { url = "https://files.pythonhosted.org/packages/f5/28/c2028e7021fb89c6e56868ed0e387b8e9aa811abdd2ab3208d6578d2c930/charset_normalizer-3.5.1-cp311-cp311-manylinux_2_31_riscv64.manylinux_2_39_riscv64.whl", hash = "sha256:6ba32c4d2abf1d2fe7cf27d280f4cca5664233b0f885549c7761719eb977f486", size = 261140, upload-time = "2026-08-15T08:16:55.604Z" }, - { url = "https://files.pythonhosted.org/packages/28/f0/0c0ceec6d98b7daa62e361e418135d59685811d79ba11529aad5cdf15e84/charset_normalizer-3.5.1-cp311-cp311-musllinux_1_2_aarch64.whl", hash = "sha256:0722590aabf9dc6a6c0343d523c05458fa2b5047dbe6302fd526bb570600753f", size = 252791, upload-time = "2026-08-15T08:16:57.103Z" }, - { url = "https://files.pythonhosted.org/packages/f0/3e/48f4cd187b1c33189d86039e9cbe4f92c05454175504b44ff81806d4d1bf/charset_normalizer-3.5.1-cp311-cp311-musllinux_1_2_armv7l.whl", hash = "sha256:aa1099b956fb795e686d073568f6dc002a0bb89765ea6d5b055dd7d9bf1b116c", size = 240730, upload-time = "2026-08-15T08:16:58.418Z" }, - { url = "https://files.pythonhosted.org/packages/42/85/f9e22af69af67c54cce42be9455d9c81294f918b4ccc454db01f66efcac2/charset_normalizer-3.5.1-cp311-cp311-musllinux_1_2_ppc64le.whl", hash = "sha256:bd6c173f04743d483881bffa1478d5a4624475b8cd1d2194956a75548e191c18", size = 280791, upload-time = "2026-08-15T08:16:59.918Z" }, - { url = "https://files.pythonhosted.org/packages/fd/4c/9044135f42127630b6fa742feb51256353f6ab87a78f2fdd1de3de955a7f/charset_normalizer-3.5.1-cp311-cp311-musllinux_1_2_riscv64.whl", hash = "sha256:f298e218441525d3794428b4c8b8fb8662c6d3ea79925d4807ee6b9a96a3bca5", size = 259598, upload-time = "2026-08-15T08:17:01.421Z" }, - { url = "https://files.pythonhosted.org/packages/ba/ed/1dd7cfebb4e75812934c49ca3b79757d11948053f7937ab7070c151f3c55/charset_normalizer-3.5.1-cp311-cp311-musllinux_1_2_s390x.whl", hash = "sha256:6e2912d4babbc65196ac13c2f53468dc57fb8b9c25ef913e8c59ddf7c6dc0e1b", size = 278217, upload-time = "2026-08-15T08:17:02.782Z" }, - { url = "https://files.pythonhosted.org/packages/bf/eb/239c84503cc9e3ba6eb34686a24bc66e84f3924efdd7e38e751a19f6bc10/charset_normalizer-3.5.1-cp311-cp311-musllinux_1_2_x86_64.whl", hash = "sha256:3d27167433c0d5f18dc850f07d0b3816221984fecdc405d6c157a6f0b8f8e9e6", size = 263417, upload-time = "2026-08-15T08:17:04.216Z" }, - { url = "https://files.pythonhosted.org/packages/37/ab/4e4510e1e288478e2c8333131d1c1382382ba8cd2165053c79e39d1da961/charset_normalizer-3.5.1-cp311-cp311-win32.whl", hash = "sha256:ac00177c4831ffa650f8609e4bdddd5fe09c03b1c0c47acece7e6ea20421598b", size = 181774, upload-time = "2026-08-15T08:17:05.58Z" }, - { url = "https://files.pythonhosted.org/packages/e3/57/32f0ccea59e8612057c61d6fd22ef2cb63cca93c9fe594094919696ac170/charset_normalizer-3.5.1-cp311-cp311-win_amd64.whl", hash = "sha256:f9b1e28d0e8dbfa858abdba91d6b547beaf2df1a59bec6da6faae7b96a4991a9", size = 206653, upload-time = "2026-08-15T08:17:07.075Z" }, - { url = "https://files.pythonhosted.org/packages/17/d4/b65c433fc521e58b5f54293982a5e51c05cb5f2dd3f1c7a6acb65b75324e/charset_normalizer-3.5.1-cp311-cp311-win_arm64.whl", hash = "sha256:ae31a1a1db2ee6cc2942fccaf695c934bc7f3db9f2133a3fef1f367cf1a4ab10", size = 185630, upload-time = "2026-08-15T08:17:08.502Z" }, - { url = "https://files.pythonhosted.org/packages/30/27/78873dc8b6a56357517b74b6bb9568b80450e7bb4f6ef7e3fa9d22aa0bd7/charset_normalizer-3.5.1-cp312-cp312-macosx_10_13_universal2.whl", hash = "sha256:5b6d1386bf0096d26d3a863dc0a487a5b4eb9aa93cf5ba69683d29dde6b9d60f", size = 344456, upload-time = "2026-08-15T08:17:10.072Z" }, - { url = "https://files.pythonhosted.org/packages/9a/4c/be49ada26b1f0232d57aa89bbebf997a5cc2332a5616b6eca26ff680044d/charset_normalizer-3.5.1-cp312-cp312-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:4582c27e8c889d64811987b5967fbd3ae0c823fe1fd933b543d55ac20bb475fa", size = 238530, upload-time = "2026-08-15T08:17:11.563Z" }, - { url = "https://files.pythonhosted.org/packages/76/84/6f1290fa07ae6978d3960caa3eb1b8019bf9284ab7c2297b00c099ef4250/charset_normalizer-3.5.1-cp312-cp312-manylinux2014_armv7l.manylinux_2_17_armv7l.manylinux_2_31_armv7l.whl", hash = "sha256:1d1c7a53a6c2103925cdd6d7229f8c567379f211c869793df679f2e9f738c369", size = 230200, upload-time = "2026-08-15T08:17:12.919Z" }, - { url = "https://files.pythonhosted.org/packages/e7/a0/47b18adeed31c8f16ba9700f32c1b18594cfa09f47eb672a488c273c22bf/charset_normalizer-3.5.1-cp312-cp312-manylinux2014_ppc64le.manylinux_2_17_ppc64le.manylinux_2_28_ppc64le.whl", hash = "sha256:e6621fb2a4988d6e53eedc455e5903e2679f3967b8acb3d639f1b63c14a2e893", size = 262222, upload-time = "2026-08-15T08:17:14.571Z" }, - { url = "https://files.pythonhosted.org/packages/38/fe/341861ac118dae06f3ec0eb487488af52128f2ef2faf0b11003944d22259/charset_normalizer-3.5.1-cp312-cp312-manylinux2014_s390x.manylinux_2_17_s390x.manylinux_2_28_s390x.whl", hash = "sha256:7c0c10730342b0c9b35dd1d619beb8214e520bd96a1f870f452680b238aab3e0", size = 258951, upload-time = "2026-08-15T08:17:16.158Z" }, - { url = "https://files.pythonhosted.org/packages/6f/89/bb5108dc6c3651dca963f2b0a3ba19bbcb370c94e1b6d3e0e844a58e6dca/charset_normalizer-3.5.1-cp312-cp312-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:b9af956078716df40d985fb0dfeb2c2120c5ca92ba4ff4b388acfd01cdc14d08", size = 248801, upload-time = "2026-08-15T08:17:17.683Z" }, - { url = "https://files.pythonhosted.org/packages/b1/ba/ef83ae3aca816393decfa3530976f38a79812d707b80b580ac33b83f9877/charset_normalizer-3.5.1-cp312-cp312-manylinux_2_31_riscv64.manylinux_2_39_riscv64.whl", hash = "sha256:f9f8405c2c758532c74fed975dbee57be1f31a6e865c031870c79a6ed3212ada", size = 244070, upload-time = "2026-08-15T08:17:19.191Z" }, - { url = "https://files.pythonhosted.org/packages/f6/0b/c5292a2462d69b7378ea89793bbb5b2b6fcf6f7dd6d1667f9619094ad553/charset_normalizer-3.5.1-cp312-cp312-musllinux_1_2_aarch64.whl", hash = "sha256:96fef3e886d6a9874b14f27fc193fbdc69d5d8035783d86aa4e1cea594e695f9", size = 240110, upload-time = "2026-08-15T08:17:20.547Z" }, - { url = "https://files.pythonhosted.org/packages/46/22/111e5be3b740d5c2a5bfcedb3d237b6591e5c2e82ae9d6ffcb121fe0909c/charset_normalizer-3.5.1-cp312-cp312-musllinux_1_2_armv7l.whl", hash = "sha256:5d8531a6569d025f68e2321e7638fb7978f23db58e5f69f56913837aae03816e", size = 232836, upload-time = "2026-08-15T08:17:21.895Z" }, - { url = "https://files.pythonhosted.org/packages/f9/d2/d2aad6fe0dbb44b194bf3becb60f5a0ac48446ade999a47fe7bb41eb09a7/charset_normalizer-3.5.1-cp312-cp312-musllinux_1_2_ppc64le.whl", hash = "sha256:aae2ee51122d3ae968a3837d97dc24a0aeebb0dea23694422cd172bd30017cd6", size = 262712, upload-time = "2026-08-15T08:17:23.727Z" }, - { url = "https://files.pythonhosted.org/packages/35/5a/337e4663a5eae6de99db940ee8066d4145caafb61327db62deda15313cce/charset_normalizer-3.5.1-cp312-cp312-musllinux_1_2_riscv64.whl", hash = "sha256:7235dc28fc6dd9d832ac7c7bce95367dedb85929f17368a0c2bee1e080b9acbf", size = 242977, upload-time = "2026-08-15T08:17:25.157Z" }, - { url = "https://files.pythonhosted.org/packages/ca/85/f82f8a92e31c7519410e2e1afdc630f28ec47490ce2c09a11c1a43cbb459/charset_normalizer-3.5.1-cp312-cp312-musllinux_1_2_s390x.whl", hash = "sha256:4abdc5f9ad448c1ecbfae2974b820535d6bc6e7eef63babbab3d81cf46968c71", size = 260207, upload-time = "2026-08-15T08:17:26.602Z" }, - { url = "https://files.pythonhosted.org/packages/b7/52/643d11ffd60e9ac2fd1fb87e167a19285b9eefeff4a40e63c87cbfbeab36/charset_normalizer-3.5.1-cp312-cp312-musllinux_1_2_x86_64.whl", hash = "sha256:ba501e667c17d8411f98e67a022d9604ef179aff0e459b7e292c796837c13573", size = 250562, upload-time = "2026-08-15T08:17:27.971Z" }, - { url = "https://files.pythonhosted.org/packages/62/16/46556278c2168d12df9da7fede5dc6fc70e60301b26a82bbeec238c9cfe3/charset_normalizer-3.5.1-cp312-cp312-win32.whl", hash = "sha256:cfa1c0cc3a8f9f53f1243a5a99ac36fd003880199383b37672e86ddda9cb07e2", size = 178507, upload-time = "2026-08-15T08:17:29.277Z" }, - { url = "https://files.pythonhosted.org/packages/9d/7a/4c6c298171e6b3e745633180ff59350fc0ca0db1ffd28df1e369e0579f71/charset_normalizer-3.5.1-cp312-cp312-win_amd64.whl", hash = "sha256:3617ac3cfd8b9888f145ad89dd6e692285834b0201c6074a5eeaad3fd4d668c2", size = 200551, upload-time = "2026-08-15T08:17:30.668Z" }, - { url = "https://files.pythonhosted.org/packages/cd/d7/eb95a042f0dd22e304b0b6472b154f3546a1a039a9ee89ccb2a7f61591fc/charset_normalizer-3.5.1-cp312-cp312-win_arm64.whl", hash = "sha256:88e85ab89cb822c1e635f51d6d32e488f94e002e70e2f492bdb8b945543f345a", size = 180700, upload-time = "2026-08-15T08:17:32.028Z" }, - { url = "https://files.pythonhosted.org/packages/5b/97/fb4e82231aba271ffd775a1b4993b0defc4e3059f286ae41d9433409fe85/charset_normalizer-3.5.1-cp37-abi3-macosx_10_9_universal2.whl", hash = "sha256:41876ee62a3dddf48ff1121ad8f0798032aa03f2fd35f21f34a4cab14f18d8d2", size = 331467, upload-time = "2026-08-15T08:19:50.959Z" }, - { url = "https://files.pythonhosted.org/packages/9f/2f/fe3f187327aac18e2d54e9d2b08e15d27bf9b642d9e51c219f130fc34d1a/charset_normalizer-3.5.1-cp37-abi3-manylinux1_x86_64.manylinux_2_28_x86_64.manylinux_2_5_x86_64.whl", hash = "sha256:a6dac12ff6b846103483683f60c5f8fee205121adc58ffd87e90a90a3af69e99", size = 253057, upload-time = "2026-08-15T08:19:52.654Z" }, - { url = "https://files.pythonhosted.org/packages/d7/c7/9e48cee5c161fe24da823b61bf381921d77cb994a0a4de148e95018c1984/charset_normalizer-3.5.1-cp37-abi3-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:cee5dd7c6fb5dd52a0fe2a740f9bc6e3593f5f8b1788bde49de02086f30182b2", size = 240930, upload-time = "2026-08-15T08:19:54.163Z" }, - { url = "https://files.pythonhosted.org/packages/49/e0/716601f3cc69be7b198951150c75ead1ece33c3c8036ff6ffa46029659a0/charset_normalizer-3.5.1-cp37-abi3-manylinux2014_armv7l.manylinux_2_17_armv7l.manylinux_2_31_armv7l.whl", hash = "sha256:343fb4f2821043bd87095f7b08a1a181febc8e36ac64212143bbfd0a0e1bc235", size = 230822, upload-time = "2026-08-15T08:19:55.807Z" }, - { url = "https://files.pythonhosted.org/packages/d3/05/71bfc5caa0abcc45aea1f6a4d50ac68e59605ddc7666fe8494f4cd229665/charset_normalizer-3.5.1-cp37-abi3-manylinux2014_ppc64le.manylinux_2_17_ppc64le.manylinux_2_28_ppc64le.whl", hash = "sha256:ae4a097991662cd4fff0ddc74e0fe7874f82e00042fa0ea00855645ed0c79598", size = 260037, upload-time = "2026-08-15T08:19:57.312Z" }, - { url = "https://files.pythonhosted.org/packages/c3/92/de7e32ed05341e7a9c4c877c318418197b7f2d66a3b68d561bf2ac57ca3e/charset_normalizer-3.5.1-cp37-abi3-manylinux2014_s390x.manylinux_2_17_s390x.manylinux_2_28_s390x.whl", hash = "sha256:4b599739b93b2cbeded49645ae3c8d1405c29ddfbceac1545c87a3f9580a9e96", size = 255097, upload-time = "2026-08-15T08:19:59.056Z" }, - { url = "https://files.pythonhosted.org/packages/f5/7b/ade0a122600319dfa0b1000ab0f9731c94a817904cf3c5de408c73a4ede7/charset_normalizer-3.5.1-cp37-abi3-manylinux_2_31_riscv64.manylinux_2_39_riscv64.whl", hash = "sha256:b39b69b347e5e47a3b5b8cfc005c68c1ba347474e3960236c4944a8ecd174962", size = 250166, upload-time = "2026-08-15T08:20:00.612Z" }, - { url = "https://files.pythonhosted.org/packages/75/9c/019fbb9f4834491a160951349b1a3714439376f66e5f7cf18b4f18f0c7aa/charset_normalizer-3.5.1-cp37-abi3-musllinux_1_2_aarch64.whl", hash = "sha256:a2028475ba855475b8b4d3cfeb4994269c967aea8b9892dfba907f4263a863a3", size = 241821, upload-time = "2026-08-15T08:20:02.321Z" }, - { url = "https://files.pythonhosted.org/packages/2b/b8/11d4840bfc99330cc7fbcc2681ee5a044553a6e77655508d8f9b2bff7b34/charset_normalizer-3.5.1-cp37-abi3-musllinux_1_2_armv7l.whl", hash = "sha256:36047af20e17097c3bb9476c2b7655f2f7aa51322c0ba58c07695bedf755a950", size = 232529, upload-time = "2026-08-15T08:20:04.008Z" }, - { url = "https://files.pythonhosted.org/packages/18/96/2b3a21492d9f65171ac75d872f5018260013d00bfa0ff70ec9f179148cbd/charset_normalizer-3.5.1-cp37-abi3-musllinux_1_2_ppc64le.whl", hash = "sha256:4c4fb141a727957c93edfe5c32a26ceb6b5f6461d67146e2d39f51e16170bea8", size = 260348, upload-time = "2026-08-15T08:20:05.877Z" }, - { url = "https://files.pythonhosted.org/packages/d6/aa/a69a2028e8bd052476c245460ab19d7de595de084dd968f2d75cd50c3e25/charset_normalizer-3.5.1-cp37-abi3-musllinux_1_2_riscv64.whl", hash = "sha256:2f293479cce755c75f1697e87c409b7ae4c555c7dfecb6e988ad13abba943031", size = 247234, upload-time = "2026-08-15T08:20:07.487Z" }, - { url = "https://files.pythonhosted.org/packages/35/8a/3d130aeabcaf3d2466af76b7b141c08d9e89c9016ab4b7cdd0f7dc2d1c62/charset_normalizer-3.5.1-cp37-abi3-musllinux_1_2_s390x.whl", hash = "sha256:3588e376b3ea2eea84976f67273d679f229e24c66dce7b82ae45aef04ff6e072", size = 256917, upload-time = "2026-08-15T08:20:09.142Z" }, - { url = "https://files.pythonhosted.org/packages/80/c2/a7379b840292d0c1ab9fbd17d1f3967aa81794dc95bc74be8999d7fedcf7/charset_normalizer-3.5.1-cp37-abi3-musllinux_1_2_x86_64.whl", hash = "sha256:e199fb99720074809a7720f1c0b4d919eea8b87e88713e0f8f602f7bef543d9d", size = 254846, upload-time = "2026-08-15T08:20:10.727Z" }, - { url = "https://files.pythonhosted.org/packages/01/65/d43b714731bb2f40d4053dfa00ecfc1c5a301f8e3316c5db3a09af59fe94/charset_normalizer-3.5.1-cp37-abi3-win32.whl", hash = "sha256:dd732602a7009217f658d5863d12d79d373a4de0eebc111094bcdd3bb8e0a6cc", size = 174216, upload-time = "2026-08-15T08:20:12.334Z" }, - { url = "https://files.pythonhosted.org/packages/35/4f/b911ed898b26a09789eba9c9200c999aff6c61b4bafaf4838e56d1a1e1a3/charset_normalizer-3.5.1-cp37-abi3-win_amd64.whl", hash = "sha256:70055ff39b97c99e7ae40ea3e393fb62aa2e44dbd9b29f8d14f42fb0025c3959", size = 199764, upload-time = "2026-08-15T08:20:13.908Z" }, - { url = "https://files.pythonhosted.org/packages/f0/a7/920baf467bfd9bf689f3b318340f37aee4572a71f162bd8db51da55ba4fa/charset_normalizer-3.5.1-cp37-abi3-win_arm64.whl", hash = "sha256:87e4f41d375c0b9be2fb5251aee4b8a689169e134535aed81bf085c3b647451e", size = 287318, upload-time = "2026-08-15T08:20:15.551Z" }, - { url = "https://files.pythonhosted.org/packages/cc/61/d01fc49b8dea277640b55a9e15960dbca9fdc8c9fde18e572d39c59f4019/charset_normalizer-3.5.1-py3-none-any.whl", hash = "sha256:6df0ec430f9a831772c23ca5a224cba36517a58a84bb32c32bb59a9fa67c47f6", size = 68658, upload-time = "2026-08-15T08:20:43.306Z" }, +source = { registry = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/" } +sdist = { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/charset-normalizer/3.5.1/charset_normalizer-3.5.1.tar.gz", hash = "sha256:6117b84ea48435e5356dc737f5121485c30920ba43375fa7b434fd753df0eac3", size = 171764, upload-time = "2026-08-15T08:20:44.807Z" } +wheels = [ + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/charset-normalizer/3.5.1/charset_normalizer-3.5.1-cp311-cp311-macosx_10_9_universal2.whl", hash = "sha256:eda059b6bc8bc0812d626fd91a7ce01bf583df0a61296eff390fd94141a34e30", size = 363585, upload-time = "2026-08-15T08:16:46.646Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/charset-normalizer/3.5.1/charset_normalizer-3.5.1-cp311-cp311-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:aa2bb0b37202dca27175591f761108b5d34096ade1191ffe4808bdf6b1571488", size = 251189, upload-time = "2026-08-15T08:16:48.287Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/charset-normalizer/3.5.1/charset_normalizer-3.5.1-cp311-cp311-manylinux2014_armv7l.manylinux_2_17_armv7l.manylinux_2_31_armv7l.whl", hash = "sha256:0b2b1b3fa5670c127b246df1d0c059defd41f689a868a3b9d79df9b1cac42d22", size = 239724, upload-time = "2026-08-15T08:16:49.67Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/charset-normalizer/3.5.1/charset_normalizer-3.5.1-cp311-cp311-manylinux2014_ppc64le.manylinux_2_17_ppc64le.manylinux_2_28_ppc64le.whl", hash = "sha256:6e5e4d73d588ca5ed09df1b7dcd1b203d1df3c542e3f50d126c947d432b10731", size = 280078, upload-time = "2026-08-15T08:16:51.228Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/charset-normalizer/3.5.1/charset_normalizer-3.5.1-cp311-cp311-manylinux2014_s390x.manylinux_2_17_s390x.manylinux_2_28_s390x.whl", hash = "sha256:b54e7e13267d49ffbfe68e25b3cbd774dab38fa37238f71265e91b36146eb21c", size = 276650, upload-time = "2026-08-15T08:16:52.625Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/charset-normalizer/3.5.1/charset_normalizer-3.5.1-cp311-cp311-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:c7b742bf31c88566b4bb6335a7f393bb322e580b6bb98df7bd0c25e6e3519ce8", size = 262325, upload-time = "2026-08-15T08:16:54.085Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/charset-normalizer/3.5.1/charset_normalizer-3.5.1-cp311-cp311-manylinux_2_31_riscv64.manylinux_2_39_riscv64.whl", hash = "sha256:6ba32c4d2abf1d2fe7cf27d280f4cca5664233b0f885549c7761719eb977f486", size = 261140, upload-time = "2026-08-15T08:16:55.604Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/charset-normalizer/3.5.1/charset_normalizer-3.5.1-cp311-cp311-musllinux_1_2_aarch64.whl", hash = "sha256:0722590aabf9dc6a6c0343d523c05458fa2b5047dbe6302fd526bb570600753f", size = 252791, upload-time = "2026-08-15T08:16:57.103Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/charset-normalizer/3.5.1/charset_normalizer-3.5.1-cp311-cp311-musllinux_1_2_armv7l.whl", hash = "sha256:aa1099b956fb795e686d073568f6dc002a0bb89765ea6d5b055dd7d9bf1b116c", size = 240730, upload-time = "2026-08-15T08:16:58.418Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/charset-normalizer/3.5.1/charset_normalizer-3.5.1-cp311-cp311-musllinux_1_2_ppc64le.whl", hash = "sha256:bd6c173f04743d483881bffa1478d5a4624475b8cd1d2194956a75548e191c18", size = 280791, upload-time = "2026-08-15T08:16:59.918Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/charset-normalizer/3.5.1/charset_normalizer-3.5.1-cp311-cp311-musllinux_1_2_riscv64.whl", hash = "sha256:f298e218441525d3794428b4c8b8fb8662c6d3ea79925d4807ee6b9a96a3bca5", size = 259598, upload-time = "2026-08-15T08:17:01.421Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/charset-normalizer/3.5.1/charset_normalizer-3.5.1-cp311-cp311-musllinux_1_2_s390x.whl", hash = "sha256:6e2912d4babbc65196ac13c2f53468dc57fb8b9c25ef913e8c59ddf7c6dc0e1b", size = 278217, upload-time = "2026-08-15T08:17:02.782Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/charset-normalizer/3.5.1/charset_normalizer-3.5.1-cp311-cp311-musllinux_1_2_x86_64.whl", hash = "sha256:3d27167433c0d5f18dc850f07d0b3816221984fecdc405d6c157a6f0b8f8e9e6", size = 263417, upload-time = "2026-08-15T08:17:04.216Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/charset-normalizer/3.5.1/charset_normalizer-3.5.1-cp311-cp311-win32.whl", hash = "sha256:ac00177c4831ffa650f8609e4bdddd5fe09c03b1c0c47acece7e6ea20421598b", size = 181774, upload-time = "2026-08-15T08:17:05.58Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/charset-normalizer/3.5.1/charset_normalizer-3.5.1-cp311-cp311-win_amd64.whl", hash = "sha256:f9b1e28d0e8dbfa858abdba91d6b547beaf2df1a59bec6da6faae7b96a4991a9", size = 206653, upload-time = "2026-08-15T08:17:07.075Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/charset-normalizer/3.5.1/charset_normalizer-3.5.1-cp311-cp311-win_arm64.whl", hash = "sha256:ae31a1a1db2ee6cc2942fccaf695c934bc7f3db9f2133a3fef1f367cf1a4ab10", size = 185630, upload-time = "2026-08-15T08:17:08.502Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/charset-normalizer/3.5.1/charset_normalizer-3.5.1-cp312-cp312-macosx_10_13_universal2.whl", hash = "sha256:5b6d1386bf0096d26d3a863dc0a487a5b4eb9aa93cf5ba69683d29dde6b9d60f", size = 344456, upload-time = "2026-08-15T08:17:10.072Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/charset-normalizer/3.5.1/charset_normalizer-3.5.1-cp312-cp312-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:4582c27e8c889d64811987b5967fbd3ae0c823fe1fd933b543d55ac20bb475fa", size = 238530, upload-time = "2026-08-15T08:17:11.563Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/charset-normalizer/3.5.1/charset_normalizer-3.5.1-cp312-cp312-manylinux2014_armv7l.manylinux_2_17_armv7l.manylinux_2_31_armv7l.whl", hash = "sha256:1d1c7a53a6c2103925cdd6d7229f8c567379f211c869793df679f2e9f738c369", size = 230200, upload-time = "2026-08-15T08:17:12.919Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/charset-normalizer/3.5.1/charset_normalizer-3.5.1-cp312-cp312-manylinux2014_ppc64le.manylinux_2_17_ppc64le.manylinux_2_28_ppc64le.whl", hash = "sha256:e6621fb2a4988d6e53eedc455e5903e2679f3967b8acb3d639f1b63c14a2e893", size = 262222, upload-time = "2026-08-15T08:17:14.571Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/charset-normalizer/3.5.1/charset_normalizer-3.5.1-cp312-cp312-manylinux2014_s390x.manylinux_2_17_s390x.manylinux_2_28_s390x.whl", hash = "sha256:7c0c10730342b0c9b35dd1d619beb8214e520bd96a1f870f452680b238aab3e0", size = 258951, upload-time = "2026-08-15T08:17:16.158Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/charset-normalizer/3.5.1/charset_normalizer-3.5.1-cp312-cp312-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:b9af956078716df40d985fb0dfeb2c2120c5ca92ba4ff4b388acfd01cdc14d08", size = 248801, upload-time = "2026-08-15T08:17:17.683Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/charset-normalizer/3.5.1/charset_normalizer-3.5.1-cp312-cp312-manylinux_2_31_riscv64.manylinux_2_39_riscv64.whl", hash = "sha256:f9f8405c2c758532c74fed975dbee57be1f31a6e865c031870c79a6ed3212ada", size = 244070, upload-time = "2026-08-15T08:17:19.191Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/charset-normalizer/3.5.1/charset_normalizer-3.5.1-cp312-cp312-musllinux_1_2_aarch64.whl", hash = "sha256:96fef3e886d6a9874b14f27fc193fbdc69d5d8035783d86aa4e1cea594e695f9", size = 240110, upload-time = "2026-08-15T08:17:20.547Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/charset-normalizer/3.5.1/charset_normalizer-3.5.1-cp312-cp312-musllinux_1_2_armv7l.whl", hash = "sha256:5d8531a6569d025f68e2321e7638fb7978f23db58e5f69f56913837aae03816e", size = 232836, upload-time = "2026-08-15T08:17:21.895Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/charset-normalizer/3.5.1/charset_normalizer-3.5.1-cp312-cp312-musllinux_1_2_ppc64le.whl", hash = "sha256:aae2ee51122d3ae968a3837d97dc24a0aeebb0dea23694422cd172bd30017cd6", size = 262712, upload-time = "2026-08-15T08:17:23.727Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/charset-normalizer/3.5.1/charset_normalizer-3.5.1-cp312-cp312-musllinux_1_2_riscv64.whl", hash = "sha256:7235dc28fc6dd9d832ac7c7bce95367dedb85929f17368a0c2bee1e080b9acbf", size = 242977, upload-time = "2026-08-15T08:17:25.157Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/charset-normalizer/3.5.1/charset_normalizer-3.5.1-cp312-cp312-musllinux_1_2_s390x.whl", hash = "sha256:4abdc5f9ad448c1ecbfae2974b820535d6bc6e7eef63babbab3d81cf46968c71", size = 260207, upload-time = "2026-08-15T08:17:26.602Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/charset-normalizer/3.5.1/charset_normalizer-3.5.1-cp312-cp312-musllinux_1_2_x86_64.whl", hash = "sha256:ba501e667c17d8411f98e67a022d9604ef179aff0e459b7e292c796837c13573", size = 250562, upload-time = "2026-08-15T08:17:27.971Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/charset-normalizer/3.5.1/charset_normalizer-3.5.1-cp312-cp312-win32.whl", hash = "sha256:cfa1c0cc3a8f9f53f1243a5a99ac36fd003880199383b37672e86ddda9cb07e2", size = 178507, upload-time = "2026-08-15T08:17:29.277Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/charset-normalizer/3.5.1/charset_normalizer-3.5.1-cp312-cp312-win_amd64.whl", hash = "sha256:3617ac3cfd8b9888f145ad89dd6e692285834b0201c6074a5eeaad3fd4d668c2", size = 200551, upload-time = "2026-08-15T08:17:30.668Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/charset-normalizer/3.5.1/charset_normalizer-3.5.1-cp312-cp312-win_arm64.whl", hash = "sha256:88e85ab89cb822c1e635f51d6d32e488f94e002e70e2f492bdb8b945543f345a", size = 180700, upload-time = "2026-08-15T08:17:32.028Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/charset-normalizer/3.5.1/charset_normalizer-3.5.1-cp37-abi3-macosx_10_9_universal2.whl", hash = "sha256:41876ee62a3dddf48ff1121ad8f0798032aa03f2fd35f21f34a4cab14f18d8d2", size = 331467, upload-time = "2026-08-15T08:19:50.959Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/charset-normalizer/3.5.1/charset_normalizer-3.5.1-cp37-abi3-manylinux1_x86_64.manylinux_2_28_x86_64.manylinux_2_5_x86_64.whl", hash = "sha256:a6dac12ff6b846103483683f60c5f8fee205121adc58ffd87e90a90a3af69e99", size = 253057, upload-time = "2026-08-15T08:19:52.654Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/charset-normalizer/3.5.1/charset_normalizer-3.5.1-cp37-abi3-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:cee5dd7c6fb5dd52a0fe2a740f9bc6e3593f5f8b1788bde49de02086f30182b2", size = 240930, upload-time = "2026-08-15T08:19:54.163Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/charset-normalizer/3.5.1/charset_normalizer-3.5.1-cp37-abi3-manylinux2014_armv7l.manylinux_2_17_armv7l.manylinux_2_31_armv7l.whl", hash = "sha256:343fb4f2821043bd87095f7b08a1a181febc8e36ac64212143bbfd0a0e1bc235", size = 230822, upload-time = "2026-08-15T08:19:55.807Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/charset-normalizer/3.5.1/charset_normalizer-3.5.1-cp37-abi3-manylinux2014_ppc64le.manylinux_2_17_ppc64le.manylinux_2_28_ppc64le.whl", hash = "sha256:ae4a097991662cd4fff0ddc74e0fe7874f82e00042fa0ea00855645ed0c79598", size = 260037, upload-time = "2026-08-15T08:19:57.312Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/charset-normalizer/3.5.1/charset_normalizer-3.5.1-cp37-abi3-manylinux2014_s390x.manylinux_2_17_s390x.manylinux_2_28_s390x.whl", hash = "sha256:4b599739b93b2cbeded49645ae3c8d1405c29ddfbceac1545c87a3f9580a9e96", size = 255097, upload-time = "2026-08-15T08:19:59.056Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/charset-normalizer/3.5.1/charset_normalizer-3.5.1-cp37-abi3-manylinux_2_31_riscv64.manylinux_2_39_riscv64.whl", hash = "sha256:b39b69b347e5e47a3b5b8cfc005c68c1ba347474e3960236c4944a8ecd174962", size = 250166, upload-time = "2026-08-15T08:20:00.612Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/charset-normalizer/3.5.1/charset_normalizer-3.5.1-cp37-abi3-musllinux_1_2_aarch64.whl", hash = "sha256:a2028475ba855475b8b4d3cfeb4994269c967aea8b9892dfba907f4263a863a3", size = 241821, upload-time = "2026-08-15T08:20:02.321Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/charset-normalizer/3.5.1/charset_normalizer-3.5.1-cp37-abi3-musllinux_1_2_armv7l.whl", hash = "sha256:36047af20e17097c3bb9476c2b7655f2f7aa51322c0ba58c07695bedf755a950", size = 232529, upload-time = "2026-08-15T08:20:04.008Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/charset-normalizer/3.5.1/charset_normalizer-3.5.1-cp37-abi3-musllinux_1_2_ppc64le.whl", hash = "sha256:4c4fb141a727957c93edfe5c32a26ceb6b5f6461d67146e2d39f51e16170bea8", size = 260348, upload-time = "2026-08-15T08:20:05.877Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/charset-normalizer/3.5.1/charset_normalizer-3.5.1-cp37-abi3-musllinux_1_2_riscv64.whl", hash = "sha256:2f293479cce755c75f1697e87c409b7ae4c555c7dfecb6e988ad13abba943031", size = 247234, upload-time = "2026-08-15T08:20:07.487Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/charset-normalizer/3.5.1/charset_normalizer-3.5.1-cp37-abi3-musllinux_1_2_s390x.whl", hash = "sha256:3588e376b3ea2eea84976f67273d679f229e24c66dce7b82ae45aef04ff6e072", size = 256917, upload-time = "2026-08-15T08:20:09.142Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/charset-normalizer/3.5.1/charset_normalizer-3.5.1-cp37-abi3-musllinux_1_2_x86_64.whl", hash = "sha256:e199fb99720074809a7720f1c0b4d919eea8b87e88713e0f8f602f7bef543d9d", size = 254846, upload-time = "2026-08-15T08:20:10.727Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/charset-normalizer/3.5.1/charset_normalizer-3.5.1-cp37-abi3-win32.whl", hash = "sha256:dd732602a7009217f658d5863d12d79d373a4de0eebc111094bcdd3bb8e0a6cc", size = 174216, upload-time = "2026-08-15T08:20:12.334Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/charset-normalizer/3.5.1/charset_normalizer-3.5.1-cp37-abi3-win_amd64.whl", hash = "sha256:70055ff39b97c99e7ae40ea3e393fb62aa2e44dbd9b29f8d14f42fb0025c3959", size = 199764, upload-time = "2026-08-15T08:20:13.908Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/charset-normalizer/3.5.1/charset_normalizer-3.5.1-cp37-abi3-win_arm64.whl", hash = "sha256:87e4f41d375c0b9be2fb5251aee4b8a689169e134535aed81bf085c3b647451e", size = 287318, upload-time = "2026-08-15T08:20:15.551Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/charset-normalizer/3.5.1/charset_normalizer-3.5.1-py3-none-any.whl", hash = "sha256:6df0ec430f9a831772c23ca5a224cba36517a58a84bb32c32bb59a9fa67c47f6", size = 68658, upload-time = "2026-08-15T08:20:43.306Z" }, ] [[package]] name = "click" version = "8.4.2" -source = { registry = "https://pypi.org/simple" } +source = { registry = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/" } dependencies = [ { name = "colorama", marker = "sys_platform == 'win32'" }, ] -sdist = { url = "https://files.pythonhosted.org/packages/76/d4/81420972a676e8ffea40450d8c8c92943e7218a78fe9b64359836cc9876b/click-8.4.2.tar.gz", hash = "sha256:9a6cea6e60b17ebe0a44c5cc636d94f09bd66142c1cd7d8b4cd731c4917a15f6", size = 338000, upload-time = "2026-06-24T17:45:15.148Z" } +sdist = { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/click/8.4.2/click-8.4.2.tar.gz", hash = "sha256:9a6cea6e60b17ebe0a44c5cc636d94f09bd66142c1cd7d8b4cd731c4917a15f6", size = 338000, upload-time = "2026-06-24T17:45:15.148Z" } wheels = [ - { url = "https://files.pythonhosted.org/packages/fb/e2/79c688af8b210d232694e31e59da9f6ec747bae31c3f5946e4e9b98860d5/click-8.4.2-py3-none-any.whl", hash = "sha256:e6f9f66136c816745b9d65817da91d61d957fb16e02e4dcd0552553c5a197b76", size = 119243, upload-time = "2026-06-24T17:45:13.73Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/click/8.4.2/click-8.4.2-py3-none-any.whl", hash = "sha256:e6f9f66136c816745b9d65817da91d61d957fb16e02e4dcd0552553c5a197b76", size = 119243, upload-time = "2026-06-24T17:45:13.73Z" }, ] [[package]] name = "colorama" version = "0.4.6" -source = { registry = "https://pypi.org/simple" } -sdist = { url = "https://files.pythonhosted.org/packages/d8/53/6f443c9a4a8358a93a6792e2acffb9d9d5cb0a5cfd8802644b7b1c9a02e4/colorama-0.4.6.tar.gz", hash = "sha256:08695f5cb7ed6e0531a20572697297273c47b8cae5a63ffc6d6ed5c201be6e44", size = 27697, upload-time = "2022-10-25T02:36:22.414Z" } +source = { registry = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/" } +sdist = { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/colorama/0.4.6/colorama-0.4.6.tar.gz", hash = "sha256:08695f5cb7ed6e0531a20572697297273c47b8cae5a63ffc6d6ed5c201be6e44", size = 27697, upload-time = "2022-10-25T02:36:22.414Z" } wheels = [ - { url = "https://files.pythonhosted.org/packages/d1/d6/3965ed04c63042e047cb6a3e6ed1a63a35087b6a609aa3a15ed8ac56c221/colorama-0.4.6-py2.py3-none-any.whl", hash = "sha256:4f1d9991f5acc0ca119f9d443620b77f9d6b33703e51011c16baf57afb285fc6", size = 25335, upload-time = "2022-10-25T02:36:20.889Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/colorama/0.4.6/colorama-0.4.6-py2.py3-none-any.whl", hash = "sha256:4f1d9991f5acc0ca119f9d443620b77f9d6b33703e51011c16baf57afb285fc6", size = 25335, upload-time = "2022-10-25T02:36:20.889Z" }, ] [[package]] name = "distro" version = "1.9.0" -source = { registry = "https://pypi.org/simple" } -sdist = { url = "https://files.pythonhosted.org/packages/fc/f8/98eea607f65de6527f8a2e8885fc8015d3e6f5775df186e443e0964a11c3/distro-1.9.0.tar.gz", hash = "sha256:2fa77c6fd8940f116ee1d6b94a2f90b13b5ea8d019b98bc8bafdcabcdd9bdbed", size = 60722, upload-time = "2023-12-24T09:54:32.31Z" } +source = { registry = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/" } +sdist = { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/distro/1.9.0/distro-1.9.0.tar.gz", hash = "sha256:2fa77c6fd8940f116ee1d6b94a2f90b13b5ea8d019b98bc8bafdcabcdd9bdbed", size = 60722, upload-time = "2023-12-24T09:54:32.31Z" } wheels = [ - { url = "https://files.pythonhosted.org/packages/12/b3/231ffd4ab1fc9d679809f356cebee130ac7daa00d6d6f3206dd4fd137e9e/distro-1.9.0-py3-none-any.whl", hash = "sha256:7bffd925d65168f85027d8da9af6bddab658135b840670a223589bc0c8ef02b2", size = 20277, upload-time = "2023-12-24T09:54:30.421Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/distro/1.9.0/distro-1.9.0-py3-none-any.whl", hash = "sha256:7bffd925d65168f85027d8da9af6bddab658135b840670a223589bc0c8ef02b2", size = 20277, upload-time = "2023-12-24T09:54:30.421Z" }, ] [[package]] name = "fastapi" version = "0.141.1" -source = { registry = "https://pypi.org/simple" } +source = { registry = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/" } dependencies = [ { name = "annotated-doc" }, { name = "pydantic" }, @@ -253,168 +253,168 @@ dependencies = [ { name = "typing-extensions" }, { name = "typing-inspection" }, ] -sdist = { url = "https://files.pythonhosted.org/packages/8a/02/91e3416a8fdd715abb903a952a6bec7cdd8d14eed55d415fc8595524c319/fastapi-0.141.1.tar.gz", hash = "sha256:e8822fc40db1e1858054d7a949a888695bc9bdce70139178e33bd2871a453ca1", size = 425799, upload-time = "2026-07-29T17:18:05.568Z" } +sdist = { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/fastapi/0.141.1/fastapi-0.141.1.tar.gz", hash = "sha256:e8822fc40db1e1858054d7a949a888695bc9bdce70139178e33bd2871a453ca1", size = 425799, upload-time = "2026-07-29T17:18:05.568Z" } wheels = [ - { url = "https://files.pythonhosted.org/packages/cb/03/10388a42375ee7e4ac9b94eb2c5c569c8b5795e377e701c9ac3ad63de890/fastapi-0.141.1-py3-none-any.whl", hash = "sha256:bfb91aa2d334c61cb35ba9a116fc123b3d3df31640b801cf57a7a78ec3f603b3", size = 131954, upload-time = "2026-07-29T17:18:04.364Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/fastapi/0.141.1/fastapi-0.141.1-py3-none-any.whl", hash = "sha256:bfb91aa2d334c61cb35ba9a116fc123b3d3df31640b801cf57a7a78ec3f603b3", size = 131954, upload-time = "2026-07-29T17:18:04.364Z" }, ] [[package]] name = "filelock" version = "3.32.4" -source = { registry = "https://pypi.org/simple" } -sdist = { url = "https://files.pythonhosted.org/packages/6d/30/03b03951873a1a0ffc7e8ca0e10c15597b59e8d0e39260704cd2ea087bc4/filelock-3.32.4.tar.gz", hash = "sha256:2bde2e4cf732e0153406d8a7bc80620ecf5e621fe0d25e41143c4e3b4733ff30", size = 222126, upload-time = "2026-08-23T17:37:55.363Z" } +source = { registry = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/" } +sdist = { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/filelock/3.32.4/filelock-3.32.4.tar.gz", hash = "sha256:2bde2e4cf732e0153406d8a7bc80620ecf5e621fe0d25e41143c4e3b4733ff30", size = 222126, upload-time = "2026-08-23T17:37:55.363Z" } wheels = [ - { url = "https://files.pythonhosted.org/packages/01/a4/9b63d595d748e3aff8812b65eacc1a2c4bd90b7c2012e08e72373b4835eb/filelock-3.32.4-py3-none-any.whl", hash = "sha256:22e58ca3b1ae3b98993b762d7338367ae64fe50252bf78d59da3bfebcdf1cedd", size = 99864, upload-time = "2026-08-23T17:37:53.913Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/filelock/3.32.4/filelock-3.32.4-py3-none-any.whl", hash = "sha256:22e58ca3b1ae3b98993b762d7338367ae64fe50252bf78d59da3bfebcdf1cedd", size = 99864, upload-time = "2026-08-23T17:37:53.913Z" }, ] [[package]] name = "frozenlist" version = "1.8.0" -source = { registry = "https://pypi.org/simple" } -sdist = { url = "https://files.pythonhosted.org/packages/2d/f5/c831fac6cc817d26fd54c7eaccd04ef7e0288806943f7cc5bbf69f3ac1f0/frozenlist-1.8.0.tar.gz", hash = "sha256:3ede829ed8d842f6cd48fc7081d7a41001a56f1f38603f9d49bf3020d59a31ad", size = 45875, upload-time = "2025-10-06T05:38:17.865Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/bc/03/077f869d540370db12165c0aa51640a873fb661d8b315d1d4d67b284d7ac/frozenlist-1.8.0-cp311-cp311-macosx_10_9_universal2.whl", hash = "sha256:09474e9831bc2b2199fad6da3c14c7b0fbdd377cce9d3d77131be28906cb7d84", size = 86912, upload-time = "2025-10-06T05:35:45.98Z" }, - { url = "https://files.pythonhosted.org/packages/df/b5/7610b6bd13e4ae77b96ba85abea1c8cb249683217ef09ac9e0ae93f25a91/frozenlist-1.8.0-cp311-cp311-macosx_10_9_x86_64.whl", hash = "sha256:17c883ab0ab67200b5f964d2b9ed6b00971917d5d8a92df149dc2c9779208ee9", size = 50046, upload-time = "2025-10-06T05:35:47.009Z" }, - { url = "https://files.pythonhosted.org/packages/6e/ef/0e8f1fe32f8a53dd26bdd1f9347efe0778b0fddf62789ea683f4cc7d787d/frozenlist-1.8.0-cp311-cp311-macosx_11_0_arm64.whl", hash = "sha256:fa47e444b8ba08fffd1c18e8cdb9a75db1b6a27f17507522834ad13ed5922b93", size = 50119, upload-time = "2025-10-06T05:35:48.38Z" }, - { url = "https://files.pythonhosted.org/packages/11/b1/71a477adc7c36e5fb628245dfbdea2166feae310757dea848d02bd0689fd/frozenlist-1.8.0-cp311-cp311-manylinux1_x86_64.manylinux_2_28_x86_64.manylinux_2_5_x86_64.whl", hash = "sha256:2552f44204b744fba866e573be4c1f9048d6a324dfe14475103fd51613eb1d1f", size = 231067, upload-time = "2025-10-06T05:35:49.97Z" }, - { url = "https://files.pythonhosted.org/packages/45/7e/afe40eca3a2dc19b9904c0f5d7edfe82b5304cb831391edec0ac04af94c2/frozenlist-1.8.0-cp311-cp311-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:957e7c38f250991e48a9a73e6423db1bb9dd14e722a10f6b8bb8e16a0f55f695", size = 233160, upload-time = "2025-10-06T05:35:51.729Z" }, - { url = "https://files.pythonhosted.org/packages/a6/aa/7416eac95603ce428679d273255ffc7c998d4132cfae200103f164b108aa/frozenlist-1.8.0-cp311-cp311-manylinux2014_armv7l.manylinux_2_17_armv7l.manylinux_2_31_armv7l.whl", hash = "sha256:8585e3bb2cdea02fc88ffa245069c36555557ad3609e83be0ec71f54fd4abb52", size = 228544, upload-time = "2025-10-06T05:35:53.246Z" }, - { url = "https://files.pythonhosted.org/packages/8b/3d/2a2d1f683d55ac7e3875e4263d28410063e738384d3adc294f5ff3d7105e/frozenlist-1.8.0-cp311-cp311-manylinux2014_ppc64le.manylinux_2_17_ppc64le.manylinux_2_28_ppc64le.whl", hash = "sha256:edee74874ce20a373d62dc28b0b18b93f645633c2943fd90ee9d898550770581", size = 243797, upload-time = "2025-10-06T05:35:54.497Z" }, - { url = "https://files.pythonhosted.org/packages/78/1e/2d5565b589e580c296d3bb54da08d206e797d941a83a6fdea42af23be79c/frozenlist-1.8.0-cp311-cp311-manylinux2014_s390x.manylinux_2_17_s390x.manylinux_2_28_s390x.whl", hash = "sha256:c9a63152fe95756b85f31186bddf42e4c02c6321207fd6601a1c89ebac4fe567", size = 247923, upload-time = "2025-10-06T05:35:55.861Z" }, - { url = "https://files.pythonhosted.org/packages/aa/c3/65872fcf1d326a7f101ad4d86285c403c87be7d832b7470b77f6d2ed5ddc/frozenlist-1.8.0-cp311-cp311-musllinux_1_2_aarch64.whl", hash = "sha256:b6db2185db9be0a04fecf2f241c70b63b1a242e2805be291855078f2b404dd6b", size = 230886, upload-time = "2025-10-06T05:35:57.399Z" }, - { url = "https://files.pythonhosted.org/packages/a0/76/ac9ced601d62f6956f03cc794f9e04c81719509f85255abf96e2510f4265/frozenlist-1.8.0-cp311-cp311-musllinux_1_2_armv7l.whl", hash = "sha256:f4be2e3d8bc8aabd566f8d5b8ba7ecc09249d74ba3c9ed52e54dc23a293f0b92", size = 245731, upload-time = "2025-10-06T05:35:58.563Z" }, - { url = "https://files.pythonhosted.org/packages/b9/49/ecccb5f2598daf0b4a1415497eba4c33c1e8ce07495eb07d2860c731b8d5/frozenlist-1.8.0-cp311-cp311-musllinux_1_2_ppc64le.whl", hash = "sha256:c8d1634419f39ea6f5c427ea2f90ca85126b54b50837f31497f3bf38266e853d", size = 241544, upload-time = "2025-10-06T05:35:59.719Z" }, - { url = "https://files.pythonhosted.org/packages/53/4b/ddf24113323c0bbcc54cb38c8b8916f1da7165e07b8e24a717b4a12cbf10/frozenlist-1.8.0-cp311-cp311-musllinux_1_2_s390x.whl", hash = "sha256:1a7fa382a4a223773ed64242dbe1c9c326ec09457e6b8428efb4118c685c3dfd", size = 241806, upload-time = "2025-10-06T05:36:00.959Z" }, - { url = "https://files.pythonhosted.org/packages/a7/fb/9b9a084d73c67175484ba2789a59f8eebebd0827d186a8102005ce41e1ba/frozenlist-1.8.0-cp311-cp311-musllinux_1_2_x86_64.whl", hash = "sha256:11847b53d722050808926e785df837353bd4d75f1d494377e59b23594d834967", size = 229382, upload-time = "2025-10-06T05:36:02.22Z" }, - { url = "https://files.pythonhosted.org/packages/95/a3/c8fb25aac55bf5e12dae5c5aa6a98f85d436c1dc658f21c3ac73f9fa95e5/frozenlist-1.8.0-cp311-cp311-win32.whl", hash = "sha256:27c6e8077956cf73eadd514be8fb04d77fc946a7fe9f7fe167648b0b9085cc25", size = 39647, upload-time = "2025-10-06T05:36:03.409Z" }, - { url = "https://files.pythonhosted.org/packages/0a/f5/603d0d6a02cfd4c8f2a095a54672b3cf967ad688a60fb9faf04fc4887f65/frozenlist-1.8.0-cp311-cp311-win_amd64.whl", hash = "sha256:ac913f8403b36a2c8610bbfd25b8013488533e71e62b4b4adce9c86c8cea905b", size = 44064, upload-time = "2025-10-06T05:36:04.368Z" }, - { url = "https://files.pythonhosted.org/packages/5d/16/c2c9ab44e181f043a86f9a8f84d5124b62dbcb3a02c0977ec72b9ac1d3e0/frozenlist-1.8.0-cp311-cp311-win_arm64.whl", hash = "sha256:d4d3214a0f8394edfa3e303136d0575eece0745ff2b47bd2cb2e66dd92d4351a", size = 39937, upload-time = "2025-10-06T05:36:05.669Z" }, - { url = "https://files.pythonhosted.org/packages/69/29/948b9aa87e75820a38650af445d2ef2b6b8a6fab1a23b6bb9e4ef0be2d59/frozenlist-1.8.0-cp312-cp312-macosx_10_13_universal2.whl", hash = "sha256:78f7b9e5d6f2fdb88cdde9440dc147259b62b9d3b019924def9f6478be254ac1", size = 87782, upload-time = "2025-10-06T05:36:06.649Z" }, - { url = "https://files.pythonhosted.org/packages/64/80/4f6e318ee2a7c0750ed724fa33a4bdf1eacdc5a39a7a24e818a773cd91af/frozenlist-1.8.0-cp312-cp312-macosx_10_13_x86_64.whl", hash = "sha256:229bf37d2e4acdaf808fd3f06e854a4a7a3661e871b10dc1f8f1896a3b05f18b", size = 50594, upload-time = "2025-10-06T05:36:07.69Z" }, - { url = "https://files.pythonhosted.org/packages/2b/94/5c8a2b50a496b11dd519f4a24cb5496cf125681dd99e94c604ccdea9419a/frozenlist-1.8.0-cp312-cp312-macosx_11_0_arm64.whl", hash = "sha256:f833670942247a14eafbb675458b4e61c82e002a148f49e68257b79296e865c4", size = 50448, upload-time = "2025-10-06T05:36:08.78Z" }, - { url = "https://files.pythonhosted.org/packages/6a/bd/d91c5e39f490a49df14320f4e8c80161cfcce09f1e2cde1edd16a551abb3/frozenlist-1.8.0-cp312-cp312-manylinux1_x86_64.manylinux_2_28_x86_64.manylinux_2_5_x86_64.whl", hash = "sha256:494a5952b1c597ba44e0e78113a7266e656b9794eec897b19ead706bd7074383", size = 242411, upload-time = "2025-10-06T05:36:09.801Z" }, - { url = "https://files.pythonhosted.org/packages/8f/83/f61505a05109ef3293dfb1ff594d13d64a2324ac3482be2cedc2be818256/frozenlist-1.8.0-cp312-cp312-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:96f423a119f4777a4a056b66ce11527366a8bb92f54e541ade21f2374433f6d4", size = 243014, upload-time = "2025-10-06T05:36:11.394Z" }, - { url = "https://files.pythonhosted.org/packages/d8/cb/cb6c7b0f7d4023ddda30cf56b8b17494eb3a79e3fda666bf735f63118b35/frozenlist-1.8.0-cp312-cp312-manylinux2014_armv7l.manylinux_2_17_armv7l.manylinux_2_31_armv7l.whl", hash = "sha256:3462dd9475af2025c31cc61be6652dfa25cbfb56cbbf52f4ccfe029f38decaf8", size = 234909, upload-time = "2025-10-06T05:36:12.598Z" }, - { url = "https://files.pythonhosted.org/packages/31/c5/cd7a1f3b8b34af009fb17d4123c5a778b44ae2804e3ad6b86204255f9ec5/frozenlist-1.8.0-cp312-cp312-manylinux2014_ppc64le.manylinux_2_17_ppc64le.manylinux_2_28_ppc64le.whl", hash = "sha256:c4c800524c9cd9bac5166cd6f55285957fcfc907db323e193f2afcd4d9abd69b", size = 250049, upload-time = "2025-10-06T05:36:14.065Z" }, - { url = "https://files.pythonhosted.org/packages/c0/01/2f95d3b416c584a1e7f0e1d6d31998c4a795f7544069ee2e0962a4b60740/frozenlist-1.8.0-cp312-cp312-manylinux2014_s390x.manylinux_2_17_s390x.manylinux_2_28_s390x.whl", hash = "sha256:d6a5df73acd3399d893dafc71663ad22534b5aa4f94e8a2fabfe856c3c1b6a52", size = 256485, upload-time = "2025-10-06T05:36:15.39Z" }, - { url = "https://files.pythonhosted.org/packages/ce/03/024bf7720b3abaebcff6d0793d73c154237b85bdf67b7ed55e5e9596dc9a/frozenlist-1.8.0-cp312-cp312-musllinux_1_2_aarch64.whl", hash = "sha256:405e8fe955c2280ce66428b3ca55e12b3c4e9c336fb2103a4937e891c69a4a29", size = 237619, upload-time = "2025-10-06T05:36:16.558Z" }, - { url = "https://files.pythonhosted.org/packages/69/fa/f8abdfe7d76b731f5d8bd217827cf6764d4f1d9763407e42717b4bed50a0/frozenlist-1.8.0-cp312-cp312-musllinux_1_2_armv7l.whl", hash = "sha256:908bd3f6439f2fef9e85031b59fd4f1297af54415fb60e4254a95f75b3cab3f3", size = 250320, upload-time = "2025-10-06T05:36:17.821Z" }, - { url = "https://files.pythonhosted.org/packages/f5/3c/b051329f718b463b22613e269ad72138cc256c540f78a6de89452803a47d/frozenlist-1.8.0-cp312-cp312-musllinux_1_2_ppc64le.whl", hash = "sha256:294e487f9ec720bd8ffcebc99d575f7eff3568a08a253d1ee1a0378754b74143", size = 246820, upload-time = "2025-10-06T05:36:19.046Z" }, - { url = "https://files.pythonhosted.org/packages/0f/ae/58282e8f98e444b3f4dd42448ff36fa38bef29e40d40f330b22e7108f565/frozenlist-1.8.0-cp312-cp312-musllinux_1_2_s390x.whl", hash = "sha256:74c51543498289c0c43656701be6b077f4b265868fa7f8a8859c197006efb608", size = 250518, upload-time = "2025-10-06T05:36:20.763Z" }, - { url = "https://files.pythonhosted.org/packages/8f/96/007e5944694d66123183845a106547a15944fbbb7154788cbf7272789536/frozenlist-1.8.0-cp312-cp312-musllinux_1_2_x86_64.whl", hash = "sha256:776f352e8329135506a1d6bf16ac3f87bc25b28e765949282dcc627af36123aa", size = 239096, upload-time = "2025-10-06T05:36:22.129Z" }, - { url = "https://files.pythonhosted.org/packages/66/bb/852b9d6db2fa40be96f29c0d1205c306288f0684df8fd26ca1951d461a56/frozenlist-1.8.0-cp312-cp312-win32.whl", hash = "sha256:433403ae80709741ce34038da08511d4a77062aa924baf411ef73d1146e74faf", size = 39985, upload-time = "2025-10-06T05:36:23.661Z" }, - { url = "https://files.pythonhosted.org/packages/b8/af/38e51a553dd66eb064cdf193841f16f077585d4d28394c2fa6235cb41765/frozenlist-1.8.0-cp312-cp312-win_amd64.whl", hash = "sha256:34187385b08f866104f0c0617404c8eb08165ab1272e884abc89c112e9c00746", size = 44591, upload-time = "2025-10-06T05:36:24.958Z" }, - { url = "https://files.pythonhosted.org/packages/a7/06/1dc65480ab147339fecc70797e9c2f69d9cea9cf38934ce08df070fdb9cb/frozenlist-1.8.0-cp312-cp312-win_arm64.whl", hash = "sha256:fe3c58d2f5db5fbd18c2987cba06d51b0529f52bc3a6cdc33d3f4eab725104bd", size = 40102, upload-time = "2025-10-06T05:36:26.333Z" }, - { url = "https://files.pythonhosted.org/packages/9a/9a/e35b4a917281c0b8419d4207f4334c8e8c5dbf4f3f5f9ada73958d937dcc/frozenlist-1.8.0-py3-none-any.whl", hash = "sha256:0c18a16eab41e82c295618a77502e17b195883241c563b00f0aa5106fc4eaa0d", size = 13409, upload-time = "2025-10-06T05:38:16.721Z" }, +source = { registry = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/" } +sdist = { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/frozenlist/1.8.0/frozenlist-1.8.0.tar.gz", hash = "sha256:3ede829ed8d842f6cd48fc7081d7a41001a56f1f38603f9d49bf3020d59a31ad", size = 45875, upload-time = "2025-10-06T05:38:17.865Z" } +wheels = [ + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/frozenlist/1.8.0/frozenlist-1.8.0-cp311-cp311-macosx_10_9_universal2.whl", hash = "sha256:09474e9831bc2b2199fad6da3c14c7b0fbdd377cce9d3d77131be28906cb7d84", size = 86912, upload-time = "2025-10-06T05:35:45.98Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/frozenlist/1.8.0/frozenlist-1.8.0-cp311-cp311-macosx_10_9_x86_64.whl", hash = "sha256:17c883ab0ab67200b5f964d2b9ed6b00971917d5d8a92df149dc2c9779208ee9", size = 50046, upload-time = "2025-10-06T05:35:47.009Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/frozenlist/1.8.0/frozenlist-1.8.0-cp311-cp311-macosx_11_0_arm64.whl", hash = "sha256:fa47e444b8ba08fffd1c18e8cdb9a75db1b6a27f17507522834ad13ed5922b93", size = 50119, upload-time = "2025-10-06T05:35:48.38Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/frozenlist/1.8.0/frozenlist-1.8.0-cp311-cp311-manylinux1_x86_64.manylinux_2_28_x86_64.manylinux_2_5_x86_64.whl", hash = "sha256:2552f44204b744fba866e573be4c1f9048d6a324dfe14475103fd51613eb1d1f", size = 231067, upload-time = "2025-10-06T05:35:49.97Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/frozenlist/1.8.0/frozenlist-1.8.0-cp311-cp311-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:957e7c38f250991e48a9a73e6423db1bb9dd14e722a10f6b8bb8e16a0f55f695", size = 233160, upload-time = "2025-10-06T05:35:51.729Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/frozenlist/1.8.0/frozenlist-1.8.0-cp311-cp311-manylinux2014_armv7l.manylinux_2_17_armv7l.manylinux_2_31_armv7l.whl", hash = "sha256:8585e3bb2cdea02fc88ffa245069c36555557ad3609e83be0ec71f54fd4abb52", size = 228544, upload-time = "2025-10-06T05:35:53.246Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/frozenlist/1.8.0/frozenlist-1.8.0-cp311-cp311-manylinux2014_ppc64le.manylinux_2_17_ppc64le.manylinux_2_28_ppc64le.whl", hash = "sha256:edee74874ce20a373d62dc28b0b18b93f645633c2943fd90ee9d898550770581", size = 243797, upload-time = "2025-10-06T05:35:54.497Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/frozenlist/1.8.0/frozenlist-1.8.0-cp311-cp311-manylinux2014_s390x.manylinux_2_17_s390x.manylinux_2_28_s390x.whl", hash = "sha256:c9a63152fe95756b85f31186bddf42e4c02c6321207fd6601a1c89ebac4fe567", size = 247923, upload-time = "2025-10-06T05:35:55.861Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/frozenlist/1.8.0/frozenlist-1.8.0-cp311-cp311-musllinux_1_2_aarch64.whl", hash = "sha256:b6db2185db9be0a04fecf2f241c70b63b1a242e2805be291855078f2b404dd6b", size = 230886, upload-time = "2025-10-06T05:35:57.399Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/frozenlist/1.8.0/frozenlist-1.8.0-cp311-cp311-musllinux_1_2_armv7l.whl", hash = "sha256:f4be2e3d8bc8aabd566f8d5b8ba7ecc09249d74ba3c9ed52e54dc23a293f0b92", size = 245731, upload-time = "2025-10-06T05:35:58.563Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/frozenlist/1.8.0/frozenlist-1.8.0-cp311-cp311-musllinux_1_2_ppc64le.whl", hash = "sha256:c8d1634419f39ea6f5c427ea2f90ca85126b54b50837f31497f3bf38266e853d", size = 241544, upload-time = "2025-10-06T05:35:59.719Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/frozenlist/1.8.0/frozenlist-1.8.0-cp311-cp311-musllinux_1_2_s390x.whl", hash = "sha256:1a7fa382a4a223773ed64242dbe1c9c326ec09457e6b8428efb4118c685c3dfd", size = 241806, upload-time = "2025-10-06T05:36:00.959Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/frozenlist/1.8.0/frozenlist-1.8.0-cp311-cp311-musllinux_1_2_x86_64.whl", hash = "sha256:11847b53d722050808926e785df837353bd4d75f1d494377e59b23594d834967", size = 229382, upload-time = "2025-10-06T05:36:02.22Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/frozenlist/1.8.0/frozenlist-1.8.0-cp311-cp311-win32.whl", hash = "sha256:27c6e8077956cf73eadd514be8fb04d77fc946a7fe9f7fe167648b0b9085cc25", size = 39647, upload-time = "2025-10-06T05:36:03.409Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/frozenlist/1.8.0/frozenlist-1.8.0-cp311-cp311-win_amd64.whl", hash = "sha256:ac913f8403b36a2c8610bbfd25b8013488533e71e62b4b4adce9c86c8cea905b", size = 44064, upload-time = "2025-10-06T05:36:04.368Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/frozenlist/1.8.0/frozenlist-1.8.0-cp311-cp311-win_arm64.whl", hash = "sha256:d4d3214a0f8394edfa3e303136d0575eece0745ff2b47bd2cb2e66dd92d4351a", size = 39937, upload-time = "2025-10-06T05:36:05.669Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/frozenlist/1.8.0/frozenlist-1.8.0-cp312-cp312-macosx_10_13_universal2.whl", hash = "sha256:78f7b9e5d6f2fdb88cdde9440dc147259b62b9d3b019924def9f6478be254ac1", size = 87782, upload-time = "2025-10-06T05:36:06.649Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/frozenlist/1.8.0/frozenlist-1.8.0-cp312-cp312-macosx_10_13_x86_64.whl", hash = "sha256:229bf37d2e4acdaf808fd3f06e854a4a7a3661e871b10dc1f8f1896a3b05f18b", size = 50594, upload-time = "2025-10-06T05:36:07.69Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/frozenlist/1.8.0/frozenlist-1.8.0-cp312-cp312-macosx_11_0_arm64.whl", hash = "sha256:f833670942247a14eafbb675458b4e61c82e002a148f49e68257b79296e865c4", size = 50448, upload-time = "2025-10-06T05:36:08.78Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/frozenlist/1.8.0/frozenlist-1.8.0-cp312-cp312-manylinux1_x86_64.manylinux_2_28_x86_64.manylinux_2_5_x86_64.whl", hash = "sha256:494a5952b1c597ba44e0e78113a7266e656b9794eec897b19ead706bd7074383", size = 242411, upload-time = "2025-10-06T05:36:09.801Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/frozenlist/1.8.0/frozenlist-1.8.0-cp312-cp312-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:96f423a119f4777a4a056b66ce11527366a8bb92f54e541ade21f2374433f6d4", size = 243014, upload-time = "2025-10-06T05:36:11.394Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/frozenlist/1.8.0/frozenlist-1.8.0-cp312-cp312-manylinux2014_armv7l.manylinux_2_17_armv7l.manylinux_2_31_armv7l.whl", hash = "sha256:3462dd9475af2025c31cc61be6652dfa25cbfb56cbbf52f4ccfe029f38decaf8", size = 234909, upload-time = "2025-10-06T05:36:12.598Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/frozenlist/1.8.0/frozenlist-1.8.0-cp312-cp312-manylinux2014_ppc64le.manylinux_2_17_ppc64le.manylinux_2_28_ppc64le.whl", hash = "sha256:c4c800524c9cd9bac5166cd6f55285957fcfc907db323e193f2afcd4d9abd69b", size = 250049, upload-time = "2025-10-06T05:36:14.065Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/frozenlist/1.8.0/frozenlist-1.8.0-cp312-cp312-manylinux2014_s390x.manylinux_2_17_s390x.manylinux_2_28_s390x.whl", hash = "sha256:d6a5df73acd3399d893dafc71663ad22534b5aa4f94e8a2fabfe856c3c1b6a52", size = 256485, upload-time = "2025-10-06T05:36:15.39Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/frozenlist/1.8.0/frozenlist-1.8.0-cp312-cp312-musllinux_1_2_aarch64.whl", hash = "sha256:405e8fe955c2280ce66428b3ca55e12b3c4e9c336fb2103a4937e891c69a4a29", size = 237619, upload-time = "2025-10-06T05:36:16.558Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/frozenlist/1.8.0/frozenlist-1.8.0-cp312-cp312-musllinux_1_2_armv7l.whl", hash = "sha256:908bd3f6439f2fef9e85031b59fd4f1297af54415fb60e4254a95f75b3cab3f3", size = 250320, upload-time = "2025-10-06T05:36:17.821Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/frozenlist/1.8.0/frozenlist-1.8.0-cp312-cp312-musllinux_1_2_ppc64le.whl", hash = "sha256:294e487f9ec720bd8ffcebc99d575f7eff3568a08a253d1ee1a0378754b74143", size = 246820, upload-time = "2025-10-06T05:36:19.046Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/frozenlist/1.8.0/frozenlist-1.8.0-cp312-cp312-musllinux_1_2_s390x.whl", hash = "sha256:74c51543498289c0c43656701be6b077f4b265868fa7f8a8859c197006efb608", size = 250518, upload-time = "2025-10-06T05:36:20.763Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/frozenlist/1.8.0/frozenlist-1.8.0-cp312-cp312-musllinux_1_2_x86_64.whl", hash = "sha256:776f352e8329135506a1d6bf16ac3f87bc25b28e765949282dcc627af36123aa", size = 239096, upload-time = "2025-10-06T05:36:22.129Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/frozenlist/1.8.0/frozenlist-1.8.0-cp312-cp312-win32.whl", hash = "sha256:433403ae80709741ce34038da08511d4a77062aa924baf411ef73d1146e74faf", size = 39985, upload-time = "2025-10-06T05:36:23.661Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/frozenlist/1.8.0/frozenlist-1.8.0-cp312-cp312-win_amd64.whl", hash = "sha256:34187385b08f866104f0c0617404c8eb08165ab1272e884abc89c112e9c00746", size = 44591, upload-time = "2025-10-06T05:36:24.958Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/frozenlist/1.8.0/frozenlist-1.8.0-cp312-cp312-win_arm64.whl", hash = "sha256:fe3c58d2f5db5fbd18c2987cba06d51b0529f52bc3a6cdc33d3f4eab725104bd", size = 40102, upload-time = "2025-10-06T05:36:26.333Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/frozenlist/1.8.0/frozenlist-1.8.0-py3-none-any.whl", hash = "sha256:0c18a16eab41e82c295618a77502e17b195883241c563b00f0aa5106fc4eaa0d", size = 13409, upload-time = "2025-10-06T05:38:16.721Z" }, ] [[package]] name = "fsspec" version = "2026.7.0" -source = { registry = "https://pypi.org/simple" } -sdist = { url = "https://files.pythonhosted.org/packages/00/78/f34251dadb8f3921264a1d9b8946f5e542014ee2614b285261b4e40e6775/fsspec-2026.7.0.tar.gz", hash = "sha256:c803c40f4cf860b49dea58ee3e1c33cb9c790520e233537e1340049f89b82a88", size = 317040, upload-time = "2026-07-28T16:34:51.052Z" } +source = { registry = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/" } +sdist = { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/fsspec/2026.7.0/fsspec-2026.7.0.tar.gz", hash = "sha256:c803c40f4cf860b49dea58ee3e1c33cb9c790520e233537e1340049f89b82a88", size = 317040, upload-time = "2026-07-28T16:34:51.052Z" } wheels = [ - { url = "https://files.pythonhosted.org/packages/fd/3c/6a2bf344106328fd04963664a60b9bb6496fc25df8e962fcdc1367285fb9/fsspec-2026.7.0-py3-none-any.whl", hash = "sha256:b57ddbafedfaef7018c1ecab32aa200a9d7ca26b77965f64e48b70061249d279", size = 206583, upload-time = "2026-07-28T16:34:49.538Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/fsspec/2026.7.0/fsspec-2026.7.0-py3-none-any.whl", hash = "sha256:b57ddbafedfaef7018c1ecab32aa200a9d7ca26b77965f64e48b70061249d279", size = 206583, upload-time = "2026-07-28T16:34:49.538Z" }, ] [[package]] name = "googleapis-common-protos" version = "1.75.3" -source = { registry = "https://pypi.org/simple" } +source = { registry = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/" } dependencies = [ { name = "protobuf" }, ] -sdist = { url = "https://files.pythonhosted.org/packages/8a/c5/4353a188e2c335aee33269e8b654af228278cca8e5f0b4b5f11e5d0e9adb/googleapis_common_protos-1.75.3.tar.gz", hash = "sha256:57c435ac2c68b108999b6db075d9053e4d7a936ba57b4a3d45667b1346f1738a", size = 153905, upload-time = "2026-09-03T22:31:21.869Z" } +sdist = { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/googleapis-common-protos/1.75.3/googleapis_common_protos-1.75.3.tar.gz", hash = "sha256:57c435ac2c68b108999b6db075d9053e4d7a936ba57b4a3d45667b1346f1738a", size = 153905, upload-time = "2026-09-03T22:31:21.869Z" } wheels = [ - { url = "https://files.pythonhosted.org/packages/1a/7a/7d79170c6ce6f12e109df2b3879d6b934010cf4f99aea8de8b7e5408c174/googleapis_common_protos-1.75.3-py3-none-any.whl", hash = "sha256:a018d2bf098ca9fb6faa08d5bb780e2a2c2f73c566f069761331386c9596d3f2", size = 306984, upload-time = "2026-09-03T22:30:45.133Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/googleapis-common-protos/1.75.3/googleapis_common_protos-1.75.3-py3-none-any.whl", hash = "sha256:a018d2bf098ca9fb6faa08d5bb780e2a2c2f73c566f069761331386c9596d3f2", size = 306984, upload-time = "2026-09-03T22:30:45.133Z" }, ] [[package]] name = "grpclib" version = "0.4.9" -source = { registry = "https://pypi.org/simple" } +source = { registry = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/" } dependencies = [ { name = "h2" }, { name = "multidict" }, ] -sdist = { url = "https://files.pythonhosted.org/packages/5b/28/5a2c299ec82a876a252c5919aa895a6f1d1d35c96417c5ce4a4660dc3a80/grpclib-0.4.9.tar.gz", hash = "sha256:cc589c330fa81004c6400a52a566407574498cb5b055fa927013361e21466c46", size = 84798, upload-time = "2025-12-14T22:23:14.349Z" } +sdist = { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/grpclib/0.4.9/grpclib-0.4.9.tar.gz", hash = "sha256:cc589c330fa81004c6400a52a566407574498cb5b055fa927013361e21466c46", size = 84798, upload-time = "2025-12-14T22:23:14.349Z" } wheels = [ - { url = "https://files.pythonhosted.org/packages/5c/90/b0cbbd9efcc82816c58f31a34963071aa19fb792a212a5d9caf8e0fc3097/grpclib-0.4.9-py3-none-any.whl", hash = "sha256:7762ec1c8ed94dfad597475152dd35cbd11aecaaca2f243e29702435ca24cf0e", size = 77063, upload-time = "2025-12-14T22:23:13.224Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/grpclib/0.4.9/grpclib-0.4.9-py3-none-any.whl", hash = "sha256:7762ec1c8ed94dfad597475152dd35cbd11aecaaca2f243e29702435ca24cf0e", size = 77063, upload-time = "2025-12-14T22:23:13.224Z" }, ] [[package]] name = "h11" version = "0.16.0" -source = { registry = "https://pypi.org/simple" } -sdist = { url = "https://files.pythonhosted.org/packages/01/ee/02a2c011bdab74c6fb3c75474d40b3052059d95df7e73351460c8588d963/h11-0.16.0.tar.gz", hash = "sha256:4e35b956cf45792e4caa5885e69fba00bdbc6ffafbfa020300e549b208ee5ff1", size = 101250, upload-time = "2025-04-24T03:35:25.427Z" } +source = { registry = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/" } +sdist = { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/h11/0.16.0/h11-0.16.0.tar.gz", hash = "sha256:4e35b956cf45792e4caa5885e69fba00bdbc6ffafbfa020300e549b208ee5ff1", size = 101250, upload-time = "2025-04-24T03:35:25.427Z" } wheels = [ - { url = "https://files.pythonhosted.org/packages/04/4b/29cac41a4d98d144bf5f6d33995617b185d14b22401f75ca86f384e87ff1/h11-0.16.0-py3-none-any.whl", hash = "sha256:63cf8bbe7522de3bf65932fda1d9c2772064ffb3dae62d55932da54b31cb6c86", size = 37515, upload-time = "2025-04-24T03:35:24.344Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/h11/0.16.0/h11-0.16.0-py3-none-any.whl", hash = "sha256:63cf8bbe7522de3bf65932fda1d9c2772064ffb3dae62d55932da54b31cb6c86", size = 37515, upload-time = "2025-04-24T03:35:24.344Z" }, ] [[package]] name = "h2" version = "4.4.1" -source = { registry = "https://pypi.org/simple" } +source = { registry = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/" } dependencies = [ { name = "hpack" }, { name = "hyperframe" }, ] -sdist = { url = "https://files.pythonhosted.org/packages/e7/85/7c366e69d84c17bb778fe41419e1fbcce3033d5b7ce29bbffff0a98b859f/h2-4.4.1.tar.gz", hash = "sha256:4e866ffb1a869ae14dd9b5e6beb5c24a13da0495ad72b65925ded182521c1516", size = 2157281, upload-time = "2026-08-03T11:45:09.509Z" } +sdist = { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/h2/4.4.1/h2-4.4.1.tar.gz", hash = "sha256:4e866ffb1a869ae14dd9b5e6beb5c24a13da0495ad72b65925ded182521c1516", size = 2157281, upload-time = "2026-08-03T11:45:09.509Z" } wheels = [ - { url = "https://files.pythonhosted.org/packages/7e/22/e85faf23bd72a92d1921e37d674ca56eb298a3c8be31fdecef0ff2b3aaac/h2-4.4.1-py3-none-any.whl", hash = "sha256:0e25f1462b23c9cb82d9eb02e28bc706dac2a68cb457c6a0d74d63c8a2a5d0e6", size = 62636, upload-time = "2026-08-03T11:44:59.164Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/h2/4.4.1/h2-4.4.1-py3-none-any.whl", hash = "sha256:0e25f1462b23c9cb82d9eb02e28bc706dac2a68cb457c6a0d74d63c8a2a5d0e6", size = 62636, upload-time = "2026-08-03T11:44:59.164Z" }, ] [[package]] name = "hf-xet" version = "1.6.0" -source = { registry = "https://pypi.org/simple" } -sdist = { url = "https://files.pythonhosted.org/packages/1b/ab/522a2ab67f27971a9d48ca666d4fca85ef7d5282d142e31fd087e27b1bbe/hf_xet-1.6.0.tar.gz", hash = "sha256:2e58454a340b3556dfa4972d5451aff4fba8dd42a236600ba1a1d2b1514f0fef", size = 920527, upload-time = "2026-08-03T22:33:13.243Z" } +source = { registry = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/" } +sdist = { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/hf-xet/1.6.0/hf_xet-1.6.0.tar.gz", hash = "sha256:2e58454a340b3556dfa4972d5451aff4fba8dd42a236600ba1a1d2b1514f0fef", size = 920527, upload-time = "2026-08-03T22:33:13.243Z" } wheels = [ - { url = "https://files.pythonhosted.org/packages/a2/50/7afa2c9c787405864fc47a0d1bbc02c62e9101947ed43c1f43899fc7d91d/hf_xet-1.6.0-cp38-abi3-macosx_10_12_x86_64.whl", hash = "sha256:633dc0cd71d32da58ab8c03ad38e2fac452c15c2b0a2866ebf6ededfe0a5061d", size = 4071729, upload-time = "2026-08-03T22:33:00.721Z" }, - { url = "https://files.pythonhosted.org/packages/4b/69/55b8dcf636142ae660fec1869fcac14c4da2e8412e14d6eee1523be77e9f/hf_xet-1.6.0-cp38-abi3-macosx_11_0_arm64.whl", hash = "sha256:f0906082d9932ae0c0057fa194041c22b4e2cdb46b2592ef3b91f020d62a081a", size = 3876287, upload-time = "2026-08-03T22:33:02.251Z" }, - { url = "https://files.pythonhosted.org/packages/67/4e/a28359bf1c1ecf11eba22123168c138698f7cb576ac678f5a2e16cd5da08/hf_xet-1.6.0-cp38-abi3-manylinux2014_x86_64.manylinux_2_17_x86_64.whl", hash = "sha256:d62671bb130879cef0ee4c9ebe47a14af6c66ec53e6d84dc15936e5ffdfac82f", size = 4464663, upload-time = "2026-08-03T22:33:03.802Z" }, - { url = "https://files.pythonhosted.org/packages/9a/69/1f0cbc2fb22ae6082d094f743d1b8945a3f36f6089cb95f42b7ee348cda7/hf_xet-1.6.0-cp38-abi3-manylinux_2_28_aarch64.whl", hash = "sha256:0e6e21fa3cdfcdcd76748564bf593870a5e013f47d97cf10aed63aa222cff5b7", size = 4262538, upload-time = "2026-08-03T22:33:05.287Z" }, - { url = "https://files.pythonhosted.org/packages/d1/3a/4f4f2301ade26e404462d3336fa11f7958d914cabbabdd6e03c3c5d5658c/hf_xet-1.6.0-cp38-abi3-musllinux_1_2_aarch64.whl", hash = "sha256:4fc74352a17015bd0ee90038bc9efe38db894cde45f268b6712b04fce8cd0acb", size = 4460520, upload-time = "2026-08-03T22:33:06.81Z" }, - { url = "https://files.pythonhosted.org/packages/ab/5f/311725e2a905534dfee2dcb5b08414f249147f1f12252bfc2bd24caa075c/hf_xet-1.6.0-cp38-abi3-musllinux_1_2_x86_64.whl", hash = "sha256:8fb4f71cba6129110c3374a33f919001ff130488fc23553698e34cc1c2a1198c", size = 4675937, upload-time = "2026-08-03T22:33:08.616Z" }, - { url = "https://files.pythonhosted.org/packages/98/b7/8c59a66d15205024662f1d66968136f13893f96df1ddc5087e2e281fc95f/hf_xet-1.6.0-cp38-abi3-win_amd64.whl", hash = "sha256:fb4fadde1b2b70bf4c0c14a6dccbe7194b1c28947fefd5bbe3fed9d940676c3b", size = 4033128, upload-time = "2026-08-03T22:33:10.171Z" }, - { url = "https://files.pythonhosted.org/packages/73/63/ca511b6f802f28cf3489b280fe77475bcca8de85e81a6299d7916b5b5555/hf_xet-1.6.0-cp38-abi3-win_arm64.whl", hash = "sha256:3dc3e35441ba395006af5aaacc40ef2e603c51ef46c3530b9156185f00935ea3", size = 3859359, upload-time = "2026-08-03T22:33:11.725Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/hf-xet/1.6.0/hf_xet-1.6.0-cp38-abi3-macosx_10_12_x86_64.whl", hash = "sha256:633dc0cd71d32da58ab8c03ad38e2fac452c15c2b0a2866ebf6ededfe0a5061d", size = 4071729, upload-time = "2026-08-03T22:33:00.721Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/hf-xet/1.6.0/hf_xet-1.6.0-cp38-abi3-macosx_11_0_arm64.whl", hash = "sha256:f0906082d9932ae0c0057fa194041c22b4e2cdb46b2592ef3b91f020d62a081a", size = 3876287, upload-time = "2026-08-03T22:33:02.251Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/hf-xet/1.6.0/hf_xet-1.6.0-cp38-abi3-manylinux2014_x86_64.manylinux_2_17_x86_64.whl", hash = "sha256:d62671bb130879cef0ee4c9ebe47a14af6c66ec53e6d84dc15936e5ffdfac82f", size = 4464663, upload-time = "2026-08-03T22:33:03.802Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/hf-xet/1.6.0/hf_xet-1.6.0-cp38-abi3-manylinux_2_28_aarch64.whl", hash = "sha256:0e6e21fa3cdfcdcd76748564bf593870a5e013f47d97cf10aed63aa222cff5b7", size = 4262538, upload-time = "2026-08-03T22:33:05.287Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/hf-xet/1.6.0/hf_xet-1.6.0-cp38-abi3-musllinux_1_2_aarch64.whl", hash = "sha256:4fc74352a17015bd0ee90038bc9efe38db894cde45f268b6712b04fce8cd0acb", size = 4460520, upload-time = "2026-08-03T22:33:06.81Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/hf-xet/1.6.0/hf_xet-1.6.0-cp38-abi3-musllinux_1_2_x86_64.whl", hash = "sha256:8fb4f71cba6129110c3374a33f919001ff130488fc23553698e34cc1c2a1198c", size = 4675937, upload-time = "2026-08-03T22:33:08.616Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/hf-xet/1.6.0/hf_xet-1.6.0-cp38-abi3-win_amd64.whl", hash = "sha256:fb4fadde1b2b70bf4c0c14a6dccbe7194b1c28947fefd5bbe3fed9d940676c3b", size = 4033128, upload-time = "2026-08-03T22:33:10.171Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/hf-xet/1.6.0/hf_xet-1.6.0-cp38-abi3-win_arm64.whl", hash = "sha256:3dc3e35441ba395006af5aaacc40ef2e603c51ef46c3530b9156185f00935ea3", size = 3859359, upload-time = "2026-08-03T22:33:11.725Z" }, ] [[package]] name = "hpack" version = "4.2.0" -source = { registry = "https://pypi.org/simple" } -sdist = { url = "https://files.pythonhosted.org/packages/26/5b/fcabf6028144a8723726318b07a32c2f3314acdff6265743cf08a344b18e/hpack-4.2.0.tar.gz", hash = "sha256:0895cfa3b5531fc65fe439c05eb65144f123bf7a394fcaa56aa423548d8e45c0", size = 51300, upload-time = "2026-06-23T18:34:46.667Z" } +source = { registry = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/" } +sdist = { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/hpack/4.2.0/hpack-4.2.0.tar.gz", hash = "sha256:0895cfa3b5531fc65fe439c05eb65144f123bf7a394fcaa56aa423548d8e45c0", size = 51300, upload-time = "2026-06-23T18:34:46.667Z" } wheels = [ - { url = "https://files.pythonhosted.org/packages/71/b4/4a9fcfb2aef6ba44d9073ecd301443aa00b3dac95de5619f2a7de7ec8a91/hpack-4.2.0-py3-none-any.whl", hash = "sha256:858ac0b02280fa582b5080d68db0899c62a80375e0e5413a74970c5e518b6986", size = 34246, upload-time = "2026-06-23T18:34:45.472Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/hpack/4.2.0/hpack-4.2.0-py3-none-any.whl", hash = "sha256:858ac0b02280fa582b5080d68db0899c62a80375e0e5413a74970c5e518b6986", size = 34246, upload-time = "2026-06-23T18:34:45.472Z" }, ] [[package]] name = "httpcore" version = "1.0.9" -source = { registry = "https://pypi.org/simple" } +source = { registry = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/" } dependencies = [ { name = "certifi" }, { name = "h11" }, ] -sdist = { url = "https://files.pythonhosted.org/packages/06/94/82699a10bca87a5556c9c59b5963f2d039dbd239f25bc2a63907a05a14cb/httpcore-1.0.9.tar.gz", hash = "sha256:6e34463af53fd2ab5d807f399a9b45ea31c3dfa2276f15a2c3f00afff6e176e8", size = 85484, upload-time = "2025-04-24T22:06:22.219Z" } +sdist = { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/httpcore/1.0.9/httpcore-1.0.9.tar.gz", hash = "sha256:6e34463af53fd2ab5d807f399a9b45ea31c3dfa2276f15a2c3f00afff6e176e8", size = 85484, upload-time = "2025-04-24T22:06:22.219Z" } wheels = [ - { url = "https://files.pythonhosted.org/packages/7e/f5/f66802a942d491edb555dd61e3a9961140fd64c90bce1eafd741609d334d/httpcore-1.0.9-py3-none-any.whl", hash = "sha256:2d400746a40668fc9dec9810239072b40b4484b640a8c38fd654a024c7a1bf55", size = 78784, upload-time = "2025-04-24T22:06:20.566Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/httpcore/1.0.9/httpcore-1.0.9-py3-none-any.whl", hash = "sha256:2d400746a40668fc9dec9810239072b40b4484b640a8c38fd654a024c7a1bf55", size = 78784, upload-time = "2025-04-24T22:06:20.566Z" }, ] [[package]] name = "httpx" version = "0.28.1" -source = { registry = "https://pypi.org/simple" } +source = { registry = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/" } dependencies = [ { name = "anyio" }, { name = "certifi" }, { name = "httpcore" }, { name = "idna" }, ] -sdist = { url = "https://files.pythonhosted.org/packages/b1/df/48c586a5fe32a0f01324ee087459e112ebb7224f646c0b5023f5e79e9956/httpx-0.28.1.tar.gz", hash = "sha256:75e98c5f16b0f35b567856f597f06ff2270a374470a5c2392242528e3e3e42fc", size = 141406, upload-time = "2024-12-06T15:37:23.222Z" } +sdist = { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/httpx/0.28.1/httpx-0.28.1.tar.gz", hash = "sha256:75e98c5f16b0f35b567856f597f06ff2270a374470a5c2392242528e3e3e42fc", size = 141406, upload-time = "2024-12-06T15:37:23.222Z" } wheels = [ - { url = "https://files.pythonhosted.org/packages/2a/39/e50c7c3a983047577ee07d2a9e53faf5a69493943ec3f6a384bdc792deb2/httpx-0.28.1-py3-none-any.whl", hash = "sha256:d909fcccc110f8c7faf814ca82a9a4d816bc5a6dbfea25d6591d6985b8ba59ad", size = 73517, upload-time = "2024-12-06T15:37:21.509Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/httpx/0.28.1/httpx-0.28.1-py3-none-any.whl", hash = "sha256:d909fcccc110f8c7faf814ca82a9a4d816bc5a6dbfea25d6591d6985b8ba59ad", size = 73517, upload-time = "2024-12-06T15:37:21.509Z" }, ] [package.optional-dependencies] @@ -425,7 +425,7 @@ http2 = [ [[package]] name = "huggingface-hub" version = "1.28.0" -source = { registry = "https://pypi.org/simple" } +source = { registry = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/" } dependencies = [ { name = "click" }, { name = "filelock" }, @@ -437,36 +437,36 @@ dependencies = [ { name = "tqdm" }, { name = "typing-extensions" }, ] -sdist = { url = "https://files.pythonhosted.org/packages/c6/ae/222a91937ebee7f62c0ca8f5ee0afd97577caf24c0abb927d1f5c7e9f6d2/huggingface_hub-1.28.0.tar.gz", hash = "sha256:46a2e950c09234de54093d587d1675382f0d08dbd600d9fb599b5932f5b2c6cb", size = 959609, upload-time = "2026-08-18T12:27:15.101Z" } +sdist = { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/huggingface-hub/1.28.0/huggingface_hub-1.28.0.tar.gz", hash = "sha256:46a2e950c09234de54093d587d1675382f0d08dbd600d9fb599b5932f5b2c6cb", size = 959609, upload-time = "2026-08-18T12:27:15.101Z" } wheels = [ - { url = "https://files.pythonhosted.org/packages/51/0e/eafef18f1a75e125e68395db21131db0cf868a128ecd2fce69b4df6c584b/huggingface_hub-1.28.0-py3-none-any.whl", hash = "sha256:58a8bacb03072edfc38067065e9dc24bbb34805410fcd36a1632de0b329660bb", size = 793202, upload-time = "2026-08-18T12:27:12.719Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/huggingface-hub/1.28.0/huggingface_hub-1.28.0-py3-none-any.whl", hash = "sha256:58a8bacb03072edfc38067065e9dc24bbb34805410fcd36a1632de0b329660bb", size = 793202, upload-time = "2026-08-18T12:27:12.719Z" }, ] [[package]] name = "hyperframe" version = "6.1.0" -source = { registry = "https://pypi.org/simple" } -sdist = { url = "https://files.pythonhosted.org/packages/02/e7/94f8232d4a74cc99514c13a9f995811485a6903d48e5d952771ef6322e30/hyperframe-6.1.0.tar.gz", hash = "sha256:f630908a00854a7adeabd6382b43923a4c4cd4b821fcb527e6ab9e15382a3b08", size = 26566, upload-time = "2025-01-22T21:41:49.302Z" } +source = { registry = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/" } +sdist = { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/hyperframe/6.1.0/hyperframe-6.1.0.tar.gz", hash = "sha256:f630908a00854a7adeabd6382b43923a4c4cd4b821fcb527e6ab9e15382a3b08", size = 26566, upload-time = "2025-01-22T21:41:49.302Z" } wheels = [ - { url = "https://files.pythonhosted.org/packages/48/30/47d0bf6072f7252e6521f3447ccfa40b421b6824517f82854703d0f5a98b/hyperframe-6.1.0-py3-none-any.whl", hash = "sha256:b03380493a519fce58ea5af42e4a42317bf9bd425596f7a0835ffce80f1a42e5", size = 13007, upload-time = "2025-01-22T21:41:47.295Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/hyperframe/6.1.0/hyperframe-6.1.0-py3-none-any.whl", hash = "sha256:b03380493a519fce58ea5af42e4a42317bf9bd425596f7a0835ffce80f1a42e5", size = 13007, upload-time = "2025-01-22T21:41:47.295Z" }, ] [[package]] name = "idna" version = "3.19" -source = { registry = "https://pypi.org/simple" } -sdist = { url = "https://files.pythonhosted.org/packages/5f/f7/abb373e5757eaec4b922b92f97ec8d6d7e057cf06778247604fbc4e7c3f3/idna-3.19.tar.gz", hash = "sha256:5e0811a4383b21dc5838069f801c4fb62113b7447663d2530d2bd6e77b49bf15", size = 215237, upload-time = "2026-08-18T05:14:24.27Z" } +source = { registry = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/" } +sdist = { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/idna/3.19/idna-3.19.tar.gz", hash = "sha256:5e0811a4383b21dc5838069f801c4fb62113b7447663d2530d2bd6e77b49bf15", size = 215237, upload-time = "2026-08-18T05:14:24.27Z" } wheels = [ - { url = "https://files.pythonhosted.org/packages/57/b0/0e52c878c53f245edd3a11020f20979b3f490f245af532c7cae3027754b5/idna-3.19-py3-none-any.whl", hash = "sha256:815e7be7a7806d54abb586dc943addc79e8b2ee16915059658cbeff4b1b43bf4", size = 68550, upload-time = "2026-08-18T05:14:22.343Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/idna/3.19/idna-3.19-py3-none-any.whl", hash = "sha256:815e7be7a7806d54abb586dc943addc79e8b2ee16915059658cbeff4b1b43bf4", size = 68550, upload-time = "2026-08-18T05:14:22.343Z" }, ] [[package]] name = "iniconfig" version = "2.3.0" -source = { registry = "https://pypi.org/simple" } -sdist = { url = "https://files.pythonhosted.org/packages/72/34/14ca021ce8e5dfedc35312d08ba8bf51fdd999c576889fc2c24cb97f4f10/iniconfig-2.3.0.tar.gz", hash = "sha256:c76315c77db068650d49c5b56314774a7804df16fee4402c1f19d6d15d8c4730", size = 20503, upload-time = "2025-10-18T21:55:43.219Z" } +source = { registry = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/" } +sdist = { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/iniconfig/2.3.0/iniconfig-2.3.0.tar.gz", hash = "sha256:c76315c77db068650d49c5b56314774a7804df16fee4402c1f19d6d15d8c4730", size = 20503, upload-time = "2025-10-18T21:55:43.219Z" } wheels = [ - { url = "https://files.pythonhosted.org/packages/cb/b1/3846dd7f199d53cb17f49cba7e651e9ce294d8497c8c150530ed11865bb8/iniconfig-2.3.0-py3-none-any.whl", hash = "sha256:f631c04d2c48c52b84d0d0549c99ff3859c98df65b3101406327ecc7d53fbf12", size = 7484, upload-time = "2025-10-18T21:55:41.639Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/iniconfig/2.3.0/iniconfig-2.3.0-py3-none-any.whl", hash = "sha256:f631c04d2c48c52b84d0d0549c99ff3859c98df65b3101406327ecc7d53fbf12", size = 7484, upload-time = "2025-10-18T21:55:41.639Z" }, ] [[package]] @@ -517,28 +517,28 @@ dev = [{ name = "pytest", specifier = ">=9.1.1" }] [[package]] name = "markdown-it-py" version = "4.2.0" -source = { registry = "https://pypi.org/simple" } +source = { registry = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/" } dependencies = [ { name = "mdurl" }, ] -sdist = { url = "https://files.pythonhosted.org/packages/06/ff/7841249c247aa650a76b9ee4bbaeae59370dc8bfd2f6c01f3630c35eb134/markdown_it_py-4.2.0.tar.gz", hash = "sha256:04a21681d6fbb623de53f6f364d352309d4094dd4194040a10fd51833e418d49", size = 82454, upload-time = "2026-05-07T12:08:28.36Z" } +sdist = { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/markdown-it-py/4.2.0/markdown_it_py-4.2.0.tar.gz", hash = "sha256:04a21681d6fbb623de53f6f364d352309d4094dd4194040a10fd51833e418d49", size = 82454, upload-time = "2026-05-07T12:08:28.36Z" } wheels = [ - { url = "https://files.pythonhosted.org/packages/b3/81/4da04ced5a082363ecfa159c010d200ecbd959ae410c10c0264a38cac0f5/markdown_it_py-4.2.0-py3-none-any.whl", hash = "sha256:9f7ebbcd14fe59494226453aed97c1070d83f8d24b6fc3a3bcf9a38092641c4a", size = 91687, upload-time = "2026-05-07T12:08:27.182Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/markdown-it-py/4.2.0/markdown_it_py-4.2.0-py3-none-any.whl", hash = "sha256:9f7ebbcd14fe59494226453aed97c1070d83f8d24b6fc3a3bcf9a38092641c4a", size = 91687, upload-time = "2026-05-07T12:08:27.182Z" }, ] [[package]] name = "mdurl" version = "0.1.2" -source = { registry = "https://pypi.org/simple" } -sdist = { url = "https://files.pythonhosted.org/packages/d6/54/cfe61301667036ec958cb99bd3efefba235e65cdeb9c84d24a8293ba1d90/mdurl-0.1.2.tar.gz", hash = "sha256:bb413d29f5eea38f31dd4754dd7377d4465116fb207585f97bf925588687c1ba", size = 8729, upload-time = "2022-08-14T12:40:10.846Z" } +source = { registry = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/" } +sdist = { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/mdurl/0.1.2/mdurl-0.1.2.tar.gz", hash = "sha256:bb413d29f5eea38f31dd4754dd7377d4465116fb207585f97bf925588687c1ba", size = 8729, upload-time = "2022-08-14T12:40:10.846Z" } wheels = [ - { url = "https://files.pythonhosted.org/packages/b3/38/89ba8ad64ae25be8de66a6d463314cf1eb366222074cfda9ee839c56a4b4/mdurl-0.1.2-py3-none-any.whl", hash = "sha256:84008a41e51615a49fc9966191ff91509e3c40b939176e643fd50a5c2196b8f8", size = 9979, upload-time = "2022-08-14T12:40:09.779Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/mdurl/0.1.2/mdurl-0.1.2-py3-none-any.whl", hash = "sha256:84008a41e51615a49fc9966191ff91509e3c40b939176e643fd50a5c2196b8f8", size = 9979, upload-time = "2022-08-14T12:40:09.779Z" }, ] [[package]] name = "modal" version = "1.5.5" -source = { registry = "https://pypi.org/simple" } +source = { registry = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/" } dependencies = [ { name = "aiohttp" }, { name = "cbor2" }, @@ -554,146 +554,146 @@ dependencies = [ { name = "typing-extensions" }, { name = "watchfiles" }, ] -sdist = { url = "https://files.pythonhosted.org/packages/b9/a5/9e322043716e511b7f6c9120804f33b92e21a818bf4709b2d98b3c76d665/modal-1.5.5.tar.gz", hash = "sha256:30df363ed1898cc3d91a09ff3f95c38ab043f6b6294011b01085312c6a0ac777", size = 870356, upload-time = "2026-08-28T19:51:34.881Z" } +sdist = { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/modal/1.5.5/modal-1.5.5.tar.gz", hash = "sha256:30df363ed1898cc3d91a09ff3f95c38ab043f6b6294011b01085312c6a0ac777", size = 870356, upload-time = "2026-08-28T19:51:34.881Z" } wheels = [ - { url = "https://files.pythonhosted.org/packages/04/19/b3dca8baec119126058b12ca620e44022e4d08a091956668bd04180c89a7/modal-1.5.5-py3-none-any.whl", hash = "sha256:8d10d3ee09818aaba1973b73ce2521ab8961b63a29b5b52e3ff0d25e7a74808e", size = 985163, upload-time = "2026-08-28T19:51:32.404Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/modal/1.5.5/modal-1.5.5-py3-none-any.whl", hash = "sha256:8d10d3ee09818aaba1973b73ce2521ab8961b63a29b5b52e3ff0d25e7a74808e", size = 985163, upload-time = "2026-08-28T19:51:32.404Z" }, ] [[package]] name = "multidict" version = "6.7.1" -source = { registry = "https://pypi.org/simple" } -sdist = { url = "https://files.pythonhosted.org/packages/1a/c2/c2d94cbe6ac1753f3fc980da97b3d930efe1da3af3c9f5125354436c073d/multidict-6.7.1.tar.gz", hash = "sha256:ec6652a1bee61c53a3e5776b6049172c53b6aaba34f18c9ad04f82712bac623d", size = 102010, upload-time = "2026-01-26T02:46:45.979Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/ce/f1/a90635c4f88fb913fbf4ce660b83b7445b7a02615bda034b2f8eb38fd597/multidict-6.7.1-cp311-cp311-macosx_10_9_universal2.whl", hash = "sha256:7ff981b266af91d7b4b3793ca3382e53229088d193a85dfad6f5f4c27fc73e5d", size = 76626, upload-time = "2026-01-26T02:43:26.485Z" }, - { url = "https://files.pythonhosted.org/packages/a6/9b/267e64eaf6fc637a15b35f5de31a566634a2740f97d8d094a69d34f524a4/multidict-6.7.1-cp311-cp311-macosx_10_9_x86_64.whl", hash = "sha256:844c5bca0b5444adb44a623fb0a1310c2f4cd41f402126bb269cd44c9b3f3e1e", size = 44706, upload-time = "2026-01-26T02:43:27.607Z" }, - { url = "https://files.pythonhosted.org/packages/dd/a4/d45caf2b97b035c57267791ecfaafbd59c68212004b3842830954bb4b02e/multidict-6.7.1-cp311-cp311-macosx_11_0_arm64.whl", hash = "sha256:f2a0a924d4c2e9afcd7ec64f9de35fcd96915149b2216e1cb2c10a56df483855", size = 44356, upload-time = "2026-01-26T02:43:28.661Z" }, - { url = "https://files.pythonhosted.org/packages/fd/d2/0a36c8473f0cbaeadd5db6c8b72d15bbceeec275807772bfcd059bef487d/multidict-6.7.1-cp311-cp311-manylinux1_i686.manylinux_2_28_i686.manylinux_2_5_i686.whl", hash = "sha256:8be1802715a8e892c784c0197c2ace276ea52702a0ede98b6310c8f255a5afb3", size = 244355, upload-time = "2026-01-26T02:43:31.165Z" }, - { url = "https://files.pythonhosted.org/packages/5d/16/8c65be997fd7dd311b7d39c7b6e71a0cb449bad093761481eccbbe4b42a2/multidict-6.7.1-cp311-cp311-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:2e2d2ed645ea29f31c4c7ea1552fcfd7cb7ba656e1eafd4134a6620c9f5fdd9e", size = 246433, upload-time = "2026-01-26T02:43:32.581Z" }, - { url = "https://files.pythonhosted.org/packages/01/fb/4dbd7e848d2799c6a026ec88ad39cf2b8416aa167fcc903baa55ecaa045c/multidict-6.7.1-cp311-cp311-manylinux2014_armv7l.manylinux_2_17_armv7l.manylinux_2_31_armv7l.whl", hash = "sha256:95922cee9a778659e91db6497596435777bd25ed116701a4c034f8e46544955a", size = 225376, upload-time = "2026-01-26T02:43:34.417Z" }, - { url = "https://files.pythonhosted.org/packages/b6/8a/4a3a6341eac3830f6053062f8fbc9a9e54407c80755b3f05bc427295c2d0/multidict-6.7.1-cp311-cp311-manylinux2014_ppc64le.manylinux_2_17_ppc64le.manylinux_2_28_ppc64le.whl", hash = "sha256:6b83cabdc375ffaaa15edd97eb7c0c672ad788e2687004990074d7d6c9b140c8", size = 257365, upload-time = "2026-01-26T02:43:35.741Z" }, - { url = "https://files.pythonhosted.org/packages/f7/a2/dd575a69c1aa206e12d27d0770cdf9b92434b48a9ef0cd0d1afdecaa93c4/multidict-6.7.1-cp311-cp311-manylinux2014_s390x.manylinux_2_17_s390x.manylinux_2_28_s390x.whl", hash = "sha256:38fb49540705369bab8484db0689d86c0a33a0a9f2c1b197f506b71b4b6c19b0", size = 254747, upload-time = "2026-01-26T02:43:36.976Z" }, - { url = "https://files.pythonhosted.org/packages/5a/56/21b27c560c13822ed93133f08aa6372c53a8e067f11fbed37b4adcdac922/multidict-6.7.1-cp311-cp311-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:439cbebd499f92e9aa6793016a8acaa161dfa749ae86d20960189f5398a19144", size = 246293, upload-time = "2026-01-26T02:43:38.258Z" }, - { url = "https://files.pythonhosted.org/packages/5a/a4/23466059dc3854763423d0ad6c0f3683a379d97673b1b89ec33826e46728/multidict-6.7.1-cp311-cp311-musllinux_1_2_aarch64.whl", hash = "sha256:6d3bc717b6fe763b8be3f2bee2701d3c8eb1b2a8ae9f60910f1b2860c82b6c49", size = 242962, upload-time = "2026-01-26T02:43:40.034Z" }, - { url = "https://files.pythonhosted.org/packages/1f/67/51dd754a3524d685958001e8fa20a0f5f90a6a856e0a9dcabff69be3dbb7/multidict-6.7.1-cp311-cp311-musllinux_1_2_armv7l.whl", hash = "sha256:619e5a1ac57986dbfec9f0b301d865dddf763696435e2962f6d9cf2fdff2bb71", size = 237360, upload-time = "2026-01-26T02:43:41.752Z" }, - { url = "https://files.pythonhosted.org/packages/64/3f/036dfc8c174934d4b55d86ff4f978e558b0e585cef70cfc1ad01adc6bf18/multidict-6.7.1-cp311-cp311-musllinux_1_2_i686.whl", hash = "sha256:0b38ebffd9be37c1170d33bc0f36f4f262e0a09bc1aac1c34c7aa51a7293f0b3", size = 245940, upload-time = "2026-01-26T02:43:43.042Z" }, - { url = "https://files.pythonhosted.org/packages/3d/20/6214d3c105928ebc353a1c644a6ef1408bc5794fcb4f170bb524a3c16311/multidict-6.7.1-cp311-cp311-musllinux_1_2_ppc64le.whl", hash = "sha256:10ae39c9cfe6adedcdb764f5e8411d4a92b055e35573a2eaa88d3323289ef93c", size = 253502, upload-time = "2026-01-26T02:43:44.371Z" }, - { url = "https://files.pythonhosted.org/packages/b1/e2/c653bc4ae1be70a0f836b82172d643fcf1dade042ba2676ab08ec08bff0f/multidict-6.7.1-cp311-cp311-musllinux_1_2_s390x.whl", hash = "sha256:25167cc263257660290fba06b9318d2026e3c910be240a146e1f66dd114af2b0", size = 247065, upload-time = "2026-01-26T02:43:45.745Z" }, - { url = "https://files.pythonhosted.org/packages/c8/11/a854b4154cd3bd8b1fd375e8a8ca9d73be37610c361543d56f764109509b/multidict-6.7.1-cp311-cp311-musllinux_1_2_x86_64.whl", hash = "sha256:128441d052254f42989ef98b7b6a6ecb1e6f708aa962c7984235316db59f50fa", size = 241870, upload-time = "2026-01-26T02:43:47.054Z" }, - { url = "https://files.pythonhosted.org/packages/13/bf/9676c0392309b5fdae322333d22a829715b570edb9baa8016a517b55b558/multidict-6.7.1-cp311-cp311-win32.whl", hash = "sha256:d62b7f64ffde3b99d06b707a280db04fb3855b55f5a06df387236051d0668f4a", size = 41302, upload-time = "2026-01-26T02:43:48.753Z" }, - { url = "https://files.pythonhosted.org/packages/c9/68/f16a3a8ba6f7b6dc92a1f19669c0810bd2c43fc5a02da13b1cbf8e253845/multidict-6.7.1-cp311-cp311-win_amd64.whl", hash = "sha256:bdbf9f3b332abd0cdb306e7c2113818ab1e922dc84b8f8fd06ec89ed2a19ab8b", size = 45981, upload-time = "2026-01-26T02:43:49.921Z" }, - { url = "https://files.pythonhosted.org/packages/ac/ad/9dd5305253fa00cd3c7555dbef69d5bf4133debc53b87ab8d6a44d411665/multidict-6.7.1-cp311-cp311-win_arm64.whl", hash = "sha256:b8c990b037d2fff2f4e33d3f21b9b531c5745b33a49a7d6dbe7a177266af44f6", size = 43159, upload-time = "2026-01-26T02:43:51.635Z" }, - { url = "https://files.pythonhosted.org/packages/8d/9c/f20e0e2cf80e4b2e4b1c365bf5fe104ee633c751a724246262db8f1a0b13/multidict-6.7.1-cp312-cp312-macosx_10_13_universal2.whl", hash = "sha256:a90f75c956e32891a4eda3639ce6dd86e87105271f43d43442a3aedf3cddf172", size = 76893, upload-time = "2026-01-26T02:43:52.754Z" }, - { url = "https://files.pythonhosted.org/packages/fe/cf/18ef143a81610136d3da8193da9d80bfe1cb548a1e2d1c775f26b23d024a/multidict-6.7.1-cp312-cp312-macosx_10_13_x86_64.whl", hash = "sha256:3fccb473e87eaa1382689053e4a4618e7ba7b9b9b8d6adf2027ee474597128cd", size = 45456, upload-time = "2026-01-26T02:43:53.893Z" }, - { url = "https://files.pythonhosted.org/packages/a9/65/1caac9d4cd32e8433908683446eebc953e82d22b03d10d41a5f0fefe991b/multidict-6.7.1-cp312-cp312-macosx_11_0_arm64.whl", hash = "sha256:b0fa96985700739c4c7853a43c0b3e169360d6855780021bfc6d0f1ce7c123e7", size = 43872, upload-time = "2026-01-26T02:43:55.041Z" }, - { url = "https://files.pythonhosted.org/packages/cf/3b/d6bd75dc4f3ff7c73766e04e705b00ed6dbbaccf670d9e05a12b006f5a21/multidict-6.7.1-cp312-cp312-manylinux1_i686.manylinux_2_28_i686.manylinux_2_5_i686.whl", hash = "sha256:cb2a55f408c3043e42b40cc8eecd575afa27b7e0b956dfb190de0f8499a57a53", size = 251018, upload-time = "2026-01-26T02:43:56.198Z" }, - { url = "https://files.pythonhosted.org/packages/fd/80/c959c5933adedb9ac15152e4067c702a808ea183a8b64cf8f31af8ad3155/multidict-6.7.1-cp312-cp312-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:eb0ce7b2a32d09892b3dd6cc44877a0d02a33241fafca5f25c8b6b62374f8b75", size = 258883, upload-time = "2026-01-26T02:43:57.499Z" }, - { url = "https://files.pythonhosted.org/packages/86/85/7ed40adafea3d4f1c8b916e3b5cc3a8e07dfcdcb9cd72800f4ed3ca1b387/multidict-6.7.1-cp312-cp312-manylinux2014_armv7l.manylinux_2_17_armv7l.manylinux_2_31_armv7l.whl", hash = "sha256:c3a32d23520ee37bf327d1e1a656fec76a2edd5c038bf43eddfa0572ec49c60b", size = 242413, upload-time = "2026-01-26T02:43:58.755Z" }, - { url = "https://files.pythonhosted.org/packages/d2/57/b8565ff533e48595503c785f8361ff9a4fde4d67de25c207cd0ba3befd03/multidict-6.7.1-cp312-cp312-manylinux2014_ppc64le.manylinux_2_17_ppc64le.manylinux_2_28_ppc64le.whl", hash = "sha256:9c90fed18bffc0189ba814749fdcc102b536e83a9f738a9003e569acd540a733", size = 268404, upload-time = "2026-01-26T02:44:00.216Z" }, - { url = "https://files.pythonhosted.org/packages/e0/50/9810c5c29350f7258180dfdcb2e52783a0632862eb334c4896ac717cebcb/multidict-6.7.1-cp312-cp312-manylinux2014_s390x.manylinux_2_17_s390x.manylinux_2_28_s390x.whl", hash = "sha256:da62917e6076f512daccfbbde27f46fed1c98fee202f0559adec8ee0de67f71a", size = 269456, upload-time = "2026-01-26T02:44:02.202Z" }, - { url = "https://files.pythonhosted.org/packages/f3/8d/5e5be3ced1d12966fefb5c4ea3b2a5b480afcea36406559442c6e31d4a48/multidict-6.7.1-cp312-cp312-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:bfde23ef6ed9db7eaee6c37dcec08524cb43903c60b285b172b6c094711b3961", size = 256322, upload-time = "2026-01-26T02:44:03.56Z" }, - { url = "https://files.pythonhosted.org/packages/31/6e/d8a26d81ac166a5592782d208dd90dfdc0a7a218adaa52b45a672b46c122/multidict-6.7.1-cp312-cp312-musllinux_1_2_aarch64.whl", hash = "sha256:3758692429e4e32f1ba0df23219cd0b4fc0a52f476726fff9337d1a57676a582", size = 253955, upload-time = "2026-01-26T02:44:04.845Z" }, - { url = "https://files.pythonhosted.org/packages/59/4c/7c672c8aad41534ba619bcd4ade7a0dc87ed6b8b5c06149b85d3dd03f0cd/multidict-6.7.1-cp312-cp312-musllinux_1_2_armv7l.whl", hash = "sha256:398c1478926eca669f2fd6a5856b6de9c0acf23a2cb59a14c0ba5844fa38077e", size = 251254, upload-time = "2026-01-26T02:44:06.133Z" }, - { url = "https://files.pythonhosted.org/packages/7b/bd/84c24de512cbafbdbc39439f74e967f19570ce7924e3007174a29c348916/multidict-6.7.1-cp312-cp312-musllinux_1_2_i686.whl", hash = "sha256:c102791b1c4f3ab36ce4101154549105a53dc828f016356b3e3bcae2e3a039d3", size = 252059, upload-time = "2026-01-26T02:44:07.518Z" }, - { url = "https://files.pythonhosted.org/packages/fa/ba/f5449385510825b73d01c2d4087bf6d2fccc20a2d42ac34df93191d3dd03/multidict-6.7.1-cp312-cp312-musllinux_1_2_ppc64le.whl", hash = "sha256:a088b62bd733e2ad12c50dad01b7d0166c30287c166e137433d3b410add807a6", size = 263588, upload-time = "2026-01-26T02:44:09.382Z" }, - { url = "https://files.pythonhosted.org/packages/d7/11/afc7c677f68f75c84a69fe37184f0f82fce13ce4b92f49f3db280b7e92b3/multidict-6.7.1-cp312-cp312-musllinux_1_2_s390x.whl", hash = "sha256:3d51ff4785d58d3f6c91bdbffcb5e1f7ddfda557727043aa20d20ec4f65e324a", size = 259642, upload-time = "2026-01-26T02:44:10.73Z" }, - { url = "https://files.pythonhosted.org/packages/2b/17/ebb9644da78c4ab36403739e0e6e0e30ebb135b9caf3440825001a0bddcb/multidict-6.7.1-cp312-cp312-musllinux_1_2_x86_64.whl", hash = "sha256:fc5907494fccf3e7d3f94f95c91d6336b092b5fc83811720fae5e2765890dfba", size = 251377, upload-time = "2026-01-26T02:44:12.042Z" }, - { url = "https://files.pythonhosted.org/packages/ca/a4/840f5b97339e27846c46307f2530a2805d9d537d8b8bd416af031cad7fa0/multidict-6.7.1-cp312-cp312-win32.whl", hash = "sha256:28ca5ce2fd9716631133d0e9a9b9a745ad7f60bac2bccafb56aa380fc0b6c511", size = 41887, upload-time = "2026-01-26T02:44:14.245Z" }, - { url = "https://files.pythonhosted.org/packages/80/31/0b2517913687895f5904325c2069d6a3b78f66cc641a86a2baf75a05dcbb/multidict-6.7.1-cp312-cp312-win_amd64.whl", hash = "sha256:fcee94dfbd638784645b066074b338bc9cc155d4b4bffa4adce1615c5a426c19", size = 46053, upload-time = "2026-01-26T02:44:15.371Z" }, - { url = "https://files.pythonhosted.org/packages/0c/5b/aba28e4ee4006ae4c7df8d327d31025d760ffa992ea23812a601d226e682/multidict-6.7.1-cp312-cp312-win_arm64.whl", hash = "sha256:ba0a9fb644d0c1a2194cf7ffb043bd852cea63a57f66fbd33959f7dae18517bf", size = 43307, upload-time = "2026-01-26T02:44:16.852Z" }, - { url = "https://files.pythonhosted.org/packages/81/08/7036c080d7117f28a4af526d794aab6a84463126db031b007717c1a6676e/multidict-6.7.1-py3-none-any.whl", hash = "sha256:55d97cc6dae627efa6a6e548885712d4864b81110ac76fa4e534c03819fa4a56", size = 12319, upload-time = "2026-01-26T02:46:44.004Z" }, +source = { registry = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/" } +sdist = { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/multidict/6.7.1/multidict-6.7.1.tar.gz", hash = "sha256:ec6652a1bee61c53a3e5776b6049172c53b6aaba34f18c9ad04f82712bac623d", size = 102010, upload-time = "2026-01-26T02:46:45.979Z" } +wheels = [ + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/multidict/6.7.1/multidict-6.7.1-cp311-cp311-macosx_10_9_universal2.whl", hash = "sha256:7ff981b266af91d7b4b3793ca3382e53229088d193a85dfad6f5f4c27fc73e5d", size = 76626, upload-time = "2026-01-26T02:43:26.485Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/multidict/6.7.1/multidict-6.7.1-cp311-cp311-macosx_10_9_x86_64.whl", hash = "sha256:844c5bca0b5444adb44a623fb0a1310c2f4cd41f402126bb269cd44c9b3f3e1e", size = 44706, upload-time = "2026-01-26T02:43:27.607Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/multidict/6.7.1/multidict-6.7.1-cp311-cp311-macosx_11_0_arm64.whl", hash = "sha256:f2a0a924d4c2e9afcd7ec64f9de35fcd96915149b2216e1cb2c10a56df483855", size = 44356, upload-time = "2026-01-26T02:43:28.661Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/multidict/6.7.1/multidict-6.7.1-cp311-cp311-manylinux1_i686.manylinux_2_28_i686.manylinux_2_5_i686.whl", hash = "sha256:8be1802715a8e892c784c0197c2ace276ea52702a0ede98b6310c8f255a5afb3", size = 244355, upload-time = "2026-01-26T02:43:31.165Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/multidict/6.7.1/multidict-6.7.1-cp311-cp311-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:2e2d2ed645ea29f31c4c7ea1552fcfd7cb7ba656e1eafd4134a6620c9f5fdd9e", size = 246433, upload-time = "2026-01-26T02:43:32.581Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/multidict/6.7.1/multidict-6.7.1-cp311-cp311-manylinux2014_armv7l.manylinux_2_17_armv7l.manylinux_2_31_armv7l.whl", hash = "sha256:95922cee9a778659e91db6497596435777bd25ed116701a4c034f8e46544955a", size = 225376, upload-time = "2026-01-26T02:43:34.417Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/multidict/6.7.1/multidict-6.7.1-cp311-cp311-manylinux2014_ppc64le.manylinux_2_17_ppc64le.manylinux_2_28_ppc64le.whl", hash = "sha256:6b83cabdc375ffaaa15edd97eb7c0c672ad788e2687004990074d7d6c9b140c8", size = 257365, upload-time = "2026-01-26T02:43:35.741Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/multidict/6.7.1/multidict-6.7.1-cp311-cp311-manylinux2014_s390x.manylinux_2_17_s390x.manylinux_2_28_s390x.whl", hash = "sha256:38fb49540705369bab8484db0689d86c0a33a0a9f2c1b197f506b71b4b6c19b0", size = 254747, upload-time = "2026-01-26T02:43:36.976Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/multidict/6.7.1/multidict-6.7.1-cp311-cp311-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:439cbebd499f92e9aa6793016a8acaa161dfa749ae86d20960189f5398a19144", size = 246293, upload-time = "2026-01-26T02:43:38.258Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/multidict/6.7.1/multidict-6.7.1-cp311-cp311-musllinux_1_2_aarch64.whl", hash = "sha256:6d3bc717b6fe763b8be3f2bee2701d3c8eb1b2a8ae9f60910f1b2860c82b6c49", size = 242962, upload-time = "2026-01-26T02:43:40.034Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/multidict/6.7.1/multidict-6.7.1-cp311-cp311-musllinux_1_2_armv7l.whl", hash = "sha256:619e5a1ac57986dbfec9f0b301d865dddf763696435e2962f6d9cf2fdff2bb71", size = 237360, upload-time = "2026-01-26T02:43:41.752Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/multidict/6.7.1/multidict-6.7.1-cp311-cp311-musllinux_1_2_i686.whl", hash = "sha256:0b38ebffd9be37c1170d33bc0f36f4f262e0a09bc1aac1c34c7aa51a7293f0b3", size = 245940, upload-time = "2026-01-26T02:43:43.042Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/multidict/6.7.1/multidict-6.7.1-cp311-cp311-musllinux_1_2_ppc64le.whl", hash = "sha256:10ae39c9cfe6adedcdb764f5e8411d4a92b055e35573a2eaa88d3323289ef93c", size = 253502, upload-time = "2026-01-26T02:43:44.371Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/multidict/6.7.1/multidict-6.7.1-cp311-cp311-musllinux_1_2_s390x.whl", hash = "sha256:25167cc263257660290fba06b9318d2026e3c910be240a146e1f66dd114af2b0", size = 247065, upload-time = "2026-01-26T02:43:45.745Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/multidict/6.7.1/multidict-6.7.1-cp311-cp311-musllinux_1_2_x86_64.whl", hash = "sha256:128441d052254f42989ef98b7b6a6ecb1e6f708aa962c7984235316db59f50fa", size = 241870, upload-time = "2026-01-26T02:43:47.054Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/multidict/6.7.1/multidict-6.7.1-cp311-cp311-win32.whl", hash = "sha256:d62b7f64ffde3b99d06b707a280db04fb3855b55f5a06df387236051d0668f4a", size = 41302, upload-time = "2026-01-26T02:43:48.753Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/multidict/6.7.1/multidict-6.7.1-cp311-cp311-win_amd64.whl", hash = "sha256:bdbf9f3b332abd0cdb306e7c2113818ab1e922dc84b8f8fd06ec89ed2a19ab8b", size = 45981, upload-time = "2026-01-26T02:43:49.921Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/multidict/6.7.1/multidict-6.7.1-cp311-cp311-win_arm64.whl", hash = "sha256:b8c990b037d2fff2f4e33d3f21b9b531c5745b33a49a7d6dbe7a177266af44f6", size = 43159, upload-time = "2026-01-26T02:43:51.635Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/multidict/6.7.1/multidict-6.7.1-cp312-cp312-macosx_10_13_universal2.whl", hash = "sha256:a90f75c956e32891a4eda3639ce6dd86e87105271f43d43442a3aedf3cddf172", size = 76893, upload-time = "2026-01-26T02:43:52.754Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/multidict/6.7.1/multidict-6.7.1-cp312-cp312-macosx_10_13_x86_64.whl", hash = "sha256:3fccb473e87eaa1382689053e4a4618e7ba7b9b9b8d6adf2027ee474597128cd", size = 45456, upload-time = "2026-01-26T02:43:53.893Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/multidict/6.7.1/multidict-6.7.1-cp312-cp312-macosx_11_0_arm64.whl", hash = "sha256:b0fa96985700739c4c7853a43c0b3e169360d6855780021bfc6d0f1ce7c123e7", size = 43872, upload-time = "2026-01-26T02:43:55.041Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/multidict/6.7.1/multidict-6.7.1-cp312-cp312-manylinux1_i686.manylinux_2_28_i686.manylinux_2_5_i686.whl", hash = "sha256:cb2a55f408c3043e42b40cc8eecd575afa27b7e0b956dfb190de0f8499a57a53", size = 251018, upload-time = "2026-01-26T02:43:56.198Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/multidict/6.7.1/multidict-6.7.1-cp312-cp312-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:eb0ce7b2a32d09892b3dd6cc44877a0d02a33241fafca5f25c8b6b62374f8b75", size = 258883, upload-time = "2026-01-26T02:43:57.499Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/multidict/6.7.1/multidict-6.7.1-cp312-cp312-manylinux2014_armv7l.manylinux_2_17_armv7l.manylinux_2_31_armv7l.whl", hash = "sha256:c3a32d23520ee37bf327d1e1a656fec76a2edd5c038bf43eddfa0572ec49c60b", size = 242413, upload-time = "2026-01-26T02:43:58.755Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/multidict/6.7.1/multidict-6.7.1-cp312-cp312-manylinux2014_ppc64le.manylinux_2_17_ppc64le.manylinux_2_28_ppc64le.whl", hash = "sha256:9c90fed18bffc0189ba814749fdcc102b536e83a9f738a9003e569acd540a733", size = 268404, upload-time = "2026-01-26T02:44:00.216Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/multidict/6.7.1/multidict-6.7.1-cp312-cp312-manylinux2014_s390x.manylinux_2_17_s390x.manylinux_2_28_s390x.whl", hash = "sha256:da62917e6076f512daccfbbde27f46fed1c98fee202f0559adec8ee0de67f71a", size = 269456, upload-time = "2026-01-26T02:44:02.202Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/multidict/6.7.1/multidict-6.7.1-cp312-cp312-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:bfde23ef6ed9db7eaee6c37dcec08524cb43903c60b285b172b6c094711b3961", size = 256322, upload-time = "2026-01-26T02:44:03.56Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/multidict/6.7.1/multidict-6.7.1-cp312-cp312-musllinux_1_2_aarch64.whl", hash = "sha256:3758692429e4e32f1ba0df23219cd0b4fc0a52f476726fff9337d1a57676a582", size = 253955, upload-time = "2026-01-26T02:44:04.845Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/multidict/6.7.1/multidict-6.7.1-cp312-cp312-musllinux_1_2_armv7l.whl", hash = "sha256:398c1478926eca669f2fd6a5856b6de9c0acf23a2cb59a14c0ba5844fa38077e", size = 251254, upload-time = "2026-01-26T02:44:06.133Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/multidict/6.7.1/multidict-6.7.1-cp312-cp312-musllinux_1_2_i686.whl", hash = "sha256:c102791b1c4f3ab36ce4101154549105a53dc828f016356b3e3bcae2e3a039d3", size = 252059, upload-time = "2026-01-26T02:44:07.518Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/multidict/6.7.1/multidict-6.7.1-cp312-cp312-musllinux_1_2_ppc64le.whl", hash = "sha256:a088b62bd733e2ad12c50dad01b7d0166c30287c166e137433d3b410add807a6", size = 263588, upload-time = "2026-01-26T02:44:09.382Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/multidict/6.7.1/multidict-6.7.1-cp312-cp312-musllinux_1_2_s390x.whl", hash = "sha256:3d51ff4785d58d3f6c91bdbffcb5e1f7ddfda557727043aa20d20ec4f65e324a", size = 259642, upload-time = "2026-01-26T02:44:10.73Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/multidict/6.7.1/multidict-6.7.1-cp312-cp312-musllinux_1_2_x86_64.whl", hash = "sha256:fc5907494fccf3e7d3f94f95c91d6336b092b5fc83811720fae5e2765890dfba", size = 251377, upload-time = "2026-01-26T02:44:12.042Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/multidict/6.7.1/multidict-6.7.1-cp312-cp312-win32.whl", hash = "sha256:28ca5ce2fd9716631133d0e9a9b9a745ad7f60bac2bccafb56aa380fc0b6c511", size = 41887, upload-time = "2026-01-26T02:44:14.245Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/multidict/6.7.1/multidict-6.7.1-cp312-cp312-win_amd64.whl", hash = "sha256:fcee94dfbd638784645b066074b338bc9cc155d4b4bffa4adce1615c5a426c19", size = 46053, upload-time = "2026-01-26T02:44:15.371Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/multidict/6.7.1/multidict-6.7.1-cp312-cp312-win_arm64.whl", hash = "sha256:ba0a9fb644d0c1a2194cf7ffb043bd852cea63a57f66fbd33959f7dae18517bf", size = 43307, upload-time = "2026-01-26T02:44:16.852Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/multidict/6.7.1/multidict-6.7.1-py3-none-any.whl", hash = "sha256:55d97cc6dae627efa6a6e548885712d4864b81110ac76fa4e534c03819fa4a56", size = 12319, upload-time = "2026-01-26T02:46:44.004Z" }, ] [[package]] name = "numpy" version = "2.4.6" -source = { registry = "https://pypi.org/simple" } +source = { registry = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/" } resolution-markers = [ "python_full_version < '3.12'", ] -sdist = { url = "https://files.pythonhosted.org/packages/d0/ad/fed0499ce6a338d2a03ebae59cd15093910c8875328855781952abf6c2fe/numpy-2.4.6.tar.gz", hash = "sha256:f3a3570c4a2a16746ac2c31a7c7c7b0c186b95ce902e33db6f28094ed7387dda", size = 20735807, upload-time = "2026-05-18T23:37:14.07Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/b3/49/ec46835a70be8fa6446c495126ac84fdb28cb2558e1620ffb87a10c8b64c/numpy-2.4.6-cp311-cp311-macosx_10_9_x86_64.whl", hash = "sha256:0280e0356c0829a18d9de1cb7eee50ec22ca639878d7240307ca0943d73cd2c4", size = 16969194, upload-time = "2026-05-18T23:33:13.503Z" }, - { url = "https://files.pythonhosted.org/packages/0e/0d/f5957185c0ee2f3e12f78715aa9e3b353fd83633316c8532b38faa37e3f6/numpy-2.4.6-cp311-cp311-macosx_11_0_arm64.whl", hash = "sha256:110f8b71aacb688ec69062bb7f6938a0f8acb01b7c1c4beb453c65b6d234584d", size = 14964111, upload-time = "2026-05-18T23:33:17.795Z" }, - { url = "https://files.pythonhosted.org/packages/ad/40/40a40ee0ddf7ceb782c49af278894b686e586d65d8c1889c8b5da01a3d7d/numpy-2.4.6-cp311-cp311-macosx_14_0_arm64.whl", hash = "sha256:4cfe66903cc32a9921a6733d96b19bb6abf310397581bbad89c228f5abaf0ee8", size = 5469159, upload-time = "2026-05-18T23:33:20.654Z" }, - { url = "https://files.pythonhosted.org/packages/63/13/f9a8046535cb21deae82f8d03de9617e08882d274fad2539630761888228/numpy-2.4.6-cp311-cp311-macosx_14_0_x86_64.whl", hash = "sha256:8155154c7c691289fe18f510b5d4657c68c67989f293f0535a91360392ff6538", size = 6798936, upload-time = "2026-05-18T23:33:22.987Z" }, - { url = "https://files.pythonhosted.org/packages/33/a8/6fa8c1a345a8c85dbb21932c447bee07c30a2c2a3f31e369c0a84b300147/numpy-2.4.6-cp311-cp311-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:0ab0a9c4ffb1a6d95ef519fe4247dba8eb6b18ad93999f76b7f657039acabd47", size = 15966692, upload-time = "2026-05-18T23:33:26.62Z" }, - { url = "https://files.pythonhosted.org/packages/02/03/74fe2a4cb3817d94d86402f2506554130a2f01414e299b5a843e5a8a957f/numpy-2.4.6-cp311-cp311-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:89cd468399cfd2504718f0ba50e410dca55a170b61a02ad92bb18c8a65186e93", size = 16918164, upload-time = "2026-05-18T23:33:29.955Z" }, - { url = "https://files.pythonhosted.org/packages/c5/80/3615be3313f7e7696609bc194b9f0101da809df79e859bdb84e0cd043f46/numpy-2.4.6-cp311-cp311-musllinux_1_2_aarch64.whl", hash = "sha256:c2d37ab77531417474168eb79d6d80b14f821a966818505d03013d0833edb7a8", size = 17322877, upload-time = "2026-05-18T23:33:34.724Z" }, - { url = "https://files.pythonhosted.org/packages/ca/ac/a691e0fe2675e370d0e08ff905adc49a1c8830e8cae03efe4477e92cd55d/numpy-2.4.6-cp311-cp311-musllinux_1_2_x86_64.whl", hash = "sha256:f407cb6b8e9d6d8c626bc73c945db1706035af8fd632295547bf1c9e46d092d6", size = 18651487, upload-time = "2026-05-18T23:33:38.217Z" }, - { url = "https://files.pythonhosted.org/packages/15/a7/9bc1cd626d7bf6869bfedf27b91b6ab5dd607758bf8e959d6fa80c6a59cb/numpy-2.4.6-cp311-cp311-win32.whl", hash = "sha256:ddea102b48f9e339f3948bf22040944184627a30fdf7f858667673b9c5f033c8", size = 6233945, upload-time = "2026-05-18T23:33:41.331Z" }, - { url = "https://files.pythonhosted.org/packages/c5/31/7fc6239c12bce7e931463251cca4426c465e1876ba3cc785402ef4dd8f4e/numpy-2.4.6-cp311-cp311-win_amd64.whl", hash = "sha256:1e254a00cdf42b1e4d5b3d68d33af63268d41340d8885df2ab6470f2e1500147", size = 12608406, upload-time = "2026-05-18T23:33:44.131Z" }, - { url = "https://files.pythonhosted.org/packages/27/83/140f85a466595a16382996a1bf06b2b54bcd597488921b0c9daaeeda72af/numpy-2.4.6-cp311-cp311-win_arm64.whl", hash = "sha256:ed9749eef4cbd126da3dc1d6bcb3a57f5eb7ac6a6484146bdbf743f552dfc577", size = 10479528, upload-time = "2026-05-18T23:33:50.725Z" }, - { url = "https://files.pythonhosted.org/packages/95/2a/3d7b5ac8aac24feaf9ad7ed58f45b0bbc06d37e4338ae84c9f2298b570f9/numpy-2.4.6-cp312-cp312-macosx_10_13_x86_64.whl", hash = "sha256:001fbb8e08d942dd57599e781f2472269ee7f2755fae407b4f67b2f0b17da3f1", size = 16689119, upload-time = "2026-05-18T23:33:54.065Z" }, - { url = "https://files.pythonhosted.org/packages/ea/12/92c4c131527599e8288d6918e888d88726f84d805d784b771f32408aeaef/numpy-2.4.6-cp312-cp312-macosx_11_0_arm64.whl", hash = "sha256:ebfb099f8dcf083deef3ac1ca4c1503f387cf76296fcb3816b66f5ecb5f54fdb", size = 14699246, upload-time = "2026-05-18T23:33:57.621Z" }, - { url = "https://files.pythonhosted.org/packages/ad/fe/c0a6b7b2ca128a8fb228575147073b660656734b8ebe4d76c8fd748dcc79/numpy-2.4.6-cp312-cp312-macosx_14_0_arm64.whl", hash = "sha256:3213d622a0283a39a93d188f3cf72b26862df52fbb4ca3697f51705016523d41", size = 5204410, upload-time = "2026-05-18T23:34:00.302Z" }, - { url = "https://files.pythonhosted.org/packages/f3/d4/9770d14ba719432bb90a421bfd443872ed0f70f7264b64bec12ea363d5fd/numpy-2.4.6-cp312-cp312-macosx_14_0_x86_64.whl", hash = "sha256:357cc07a6d7b0b182ff02249616a03742827ebb1277546b5c7cd7f7620a45698", size = 6551240, upload-time = "2026-05-18T23:34:02.852Z" }, - { url = "https://files.pythonhosted.org/packages/c9/c6/50a46a6205feba2343f1d6d17438107c5dc491ed1c736e6ea68689fd906b/numpy-2.4.6-cp312-cp312-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:5f9fb9157b4ce2971008323afe46053787b526ef624fea915b261468a8421a0f", size = 15671012, upload-time = "2026-05-18T23:34:05.485Z" }, - { url = "https://files.pythonhosted.org/packages/99/60/14115e6364fa676c5397c2ad3004e527e9aa487abf5d0706ec81bbd08529/numpy-2.4.6-cp312-cp312-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:90f9849678c75fe7afa2d348ac842c168b0a4d3d61919687216dfc547976d853", size = 16645538, upload-time = "2026-05-18T23:34:09.265Z" }, - { url = "https://files.pythonhosted.org/packages/ae/c5/693cbe59e57db94d2231fa519ca3978dc9e19da5a8f088588f5c6e947ff2/numpy-2.4.6-cp312-cp312-musllinux_1_2_aarch64.whl", hash = "sha256:c1a2af6c6ef86344a6b0db6b97834208bf598db514f2b155042439b62605601a", size = 17020706, upload-time = "2026-05-18T23:34:13.053Z" }, - { url = "https://files.pythonhosted.org/packages/ef/fc/85b7c4eff9b4966ade25c2273cf7e7012e92366c032058653934b37de044/numpy-2.4.6-cp312-cp312-musllinux_1_2_x86_64.whl", hash = "sha256:e5805d5a22fd19c8ccff10a9561f9df94436b0545619ea579db2d3c35294bce2", size = 18368541, upload-time = "2026-05-18T23:34:17.024Z" }, - { url = "https://files.pythonhosted.org/packages/f6/81/e1b27545deedce7f4a0b348618c6b62d74e36a4dc9ccd42f3eb2f85eee32/numpy-2.4.6-cp312-cp312-win32.whl", hash = "sha256:e3eeb0aabd6bd5ce64faae67e9935203a6991b4bc2a485a767fbafb2c5125f45", size = 5962825, upload-time = "2026-05-18T23:34:20.3Z" }, - { url = "https://files.pythonhosted.org/packages/ab/ca/feab00bd44aa5fe1ad2c18f08b4d3bb92e26484b0b1d1443897809ed528c/numpy-2.4.6-cp312-cp312-win_amd64.whl", hash = "sha256:d8e8286dd7cea7895157318d1b91cdacac64c479f3cbc8dce548331728484751", size = 12321687, upload-time = "2026-05-18T23:34:23.095Z" }, - { url = "https://files.pythonhosted.org/packages/63/cf/5a6d34850a39d1093558564f77ee8e8e0bee5061151b8f05a55711001ec7/numpy-2.4.6-cp312-cp312-win_arm64.whl", hash = "sha256:4081eb135ac24158bd51cdfbef16f1c64df7063b1143f24731387137c092bec8", size = 10221482, upload-time = "2026-05-18T23:34:25.876Z" }, - { url = "https://files.pythonhosted.org/packages/de/12/b422cc84439adc0d00de605bf4a308890ae5c26f2c71fbd73e5d08fbb0dd/numpy-2.4.6-pp311-pypy311_pp73-macosx_10_15_x86_64.whl", hash = "sha256:55cced7c52e981362f708ad635198e97a752dfba412cc03c23bbf3bd8d5cd662", size = 16847511, upload-time = "2026-05-18T23:36:50.673Z" }, - { url = "https://files.pythonhosted.org/packages/44/53/f481bef68011740f8849418d82db07230e825013f31f4eef5ba5b805316a/numpy-2.4.6-pp311-pypy311_pp73-macosx_11_0_arm64.whl", hash = "sha256:d6da64deb6b8ed903e7560180a92f2d804ee1ba5eeb849ac2748b8c1aba1f6d7", size = 14889064, upload-time = "2026-05-18T23:36:53.879Z" }, - { url = "https://files.pythonhosted.org/packages/7f/57/42ed575c10ced8af951d426bc4e1f8aff16fd851db33f067036215a7f860/numpy-2.4.6-pp311-pypy311_pp73-macosx_14_0_arm64.whl", hash = "sha256:68a5124b13fa6cc2086764a20005d30bc0548146f7f5322f02fce212ca14317f", size = 5394157, upload-time = "2026-05-18T23:36:57.194Z" }, - { url = "https://files.pythonhosted.org/packages/6a/ef/f66cc724fcc36c1e364c67f51ae9146090b8b584f27d58b97fdae3edd737/numpy-2.4.6-pp311-pypy311_pp73-macosx_14_0_x86_64.whl", hash = "sha256:948424b06129ce883307e8cff868c31396d8dc7630a59c61d70d98dbe70f222c", size = 6708728, upload-time = "2026-05-18T23:36:59.575Z" }, - { url = "https://files.pythonhosted.org/packages/1a/9c/c531f2293b91265d8b48e9b329f54fdd7ffae73cb4134ea10cca4237e9cc/numpy-2.4.6-pp311-pypy311_pp73-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:5dbbdb29840ca3d91ee0fece42fc29278886d908280bfec0a5846c6f901a3eb0", size = 15798374, upload-time = "2026-05-18T23:37:02.674Z" }, - { url = "https://files.pythonhosted.org/packages/1a/b0/413077f6b1153ed3cba361401c6783bbad6114804a000cc22eb71c13e190/numpy-2.4.6-pp311-pypy311_pp73-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:8ad03c0965fb3c692200e74d458ca28c1dbb4ce96f9a479a8aa041ad5fabca02", size = 16747286, upload-time = "2026-05-18T23:37:06.327Z" }, - { url = "https://files.pythonhosted.org/packages/15/ce/e5ec180bc41812edcd8daeb8639d205622c0e8c02259d8ab25a0201b3c2a/numpy-2.4.6-pp311-pypy311_pp73-win_amd64.whl", hash = "sha256:2803abfebfc990042cd494d8ce2d5f82e9d847af6d35ec486923aa19dbad5e73", size = 12504263, upload-time = "2026-05-18T23:37:09.715Z" }, +sdist = { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/numpy/2.4.6/numpy-2.4.6.tar.gz", hash = "sha256:f3a3570c4a2a16746ac2c31a7c7c7b0c186b95ce902e33db6f28094ed7387dda", size = 20735807, upload-time = "2026-05-18T23:37:14.07Z" } +wheels = [ + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/numpy/2.4.6/numpy-2.4.6-cp311-cp311-macosx_10_9_x86_64.whl", hash = "sha256:0280e0356c0829a18d9de1cb7eee50ec22ca639878d7240307ca0943d73cd2c4", size = 16969194, upload-time = "2026-05-18T23:33:13.503Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/numpy/2.4.6/numpy-2.4.6-cp311-cp311-macosx_11_0_arm64.whl", hash = "sha256:110f8b71aacb688ec69062bb7f6938a0f8acb01b7c1c4beb453c65b6d234584d", size = 14964111, upload-time = "2026-05-18T23:33:17.795Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/numpy/2.4.6/numpy-2.4.6-cp311-cp311-macosx_14_0_arm64.whl", hash = "sha256:4cfe66903cc32a9921a6733d96b19bb6abf310397581bbad89c228f5abaf0ee8", size = 5469159, upload-time = "2026-05-18T23:33:20.654Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/numpy/2.4.6/numpy-2.4.6-cp311-cp311-macosx_14_0_x86_64.whl", hash = "sha256:8155154c7c691289fe18f510b5d4657c68c67989f293f0535a91360392ff6538", size = 6798936, upload-time = "2026-05-18T23:33:22.987Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/numpy/2.4.6/numpy-2.4.6-cp311-cp311-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:0ab0a9c4ffb1a6d95ef519fe4247dba8eb6b18ad93999f76b7f657039acabd47", size = 15966692, upload-time = "2026-05-18T23:33:26.62Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/numpy/2.4.6/numpy-2.4.6-cp311-cp311-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:89cd468399cfd2504718f0ba50e410dca55a170b61a02ad92bb18c8a65186e93", size = 16918164, upload-time = "2026-05-18T23:33:29.955Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/numpy/2.4.6/numpy-2.4.6-cp311-cp311-musllinux_1_2_aarch64.whl", hash = "sha256:c2d37ab77531417474168eb79d6d80b14f821a966818505d03013d0833edb7a8", size = 17322877, upload-time = "2026-05-18T23:33:34.724Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/numpy/2.4.6/numpy-2.4.6-cp311-cp311-musllinux_1_2_x86_64.whl", hash = "sha256:f407cb6b8e9d6d8c626bc73c945db1706035af8fd632295547bf1c9e46d092d6", size = 18651487, upload-time = "2026-05-18T23:33:38.217Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/numpy/2.4.6/numpy-2.4.6-cp311-cp311-win32.whl", hash = "sha256:ddea102b48f9e339f3948bf22040944184627a30fdf7f858667673b9c5f033c8", size = 6233945, upload-time = "2026-05-18T23:33:41.331Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/numpy/2.4.6/numpy-2.4.6-cp311-cp311-win_amd64.whl", hash = "sha256:1e254a00cdf42b1e4d5b3d68d33af63268d41340d8885df2ab6470f2e1500147", size = 12608406, upload-time = "2026-05-18T23:33:44.131Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/numpy/2.4.6/numpy-2.4.6-cp311-cp311-win_arm64.whl", hash = "sha256:ed9749eef4cbd126da3dc1d6bcb3a57f5eb7ac6a6484146bdbf743f552dfc577", size = 10479528, upload-time = "2026-05-18T23:33:50.725Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/numpy/2.4.6/numpy-2.4.6-cp312-cp312-macosx_10_13_x86_64.whl", hash = "sha256:001fbb8e08d942dd57599e781f2472269ee7f2755fae407b4f67b2f0b17da3f1", size = 16689119, upload-time = "2026-05-18T23:33:54.065Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/numpy/2.4.6/numpy-2.4.6-cp312-cp312-macosx_11_0_arm64.whl", hash = "sha256:ebfb099f8dcf083deef3ac1ca4c1503f387cf76296fcb3816b66f5ecb5f54fdb", size = 14699246, upload-time = "2026-05-18T23:33:57.621Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/numpy/2.4.6/numpy-2.4.6-cp312-cp312-macosx_14_0_arm64.whl", hash = "sha256:3213d622a0283a39a93d188f3cf72b26862df52fbb4ca3697f51705016523d41", size = 5204410, upload-time = "2026-05-18T23:34:00.302Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/numpy/2.4.6/numpy-2.4.6-cp312-cp312-macosx_14_0_x86_64.whl", hash = "sha256:357cc07a6d7b0b182ff02249616a03742827ebb1277546b5c7cd7f7620a45698", size = 6551240, upload-time = "2026-05-18T23:34:02.852Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/numpy/2.4.6/numpy-2.4.6-cp312-cp312-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:5f9fb9157b4ce2971008323afe46053787b526ef624fea915b261468a8421a0f", size = 15671012, upload-time = "2026-05-18T23:34:05.485Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/numpy/2.4.6/numpy-2.4.6-cp312-cp312-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:90f9849678c75fe7afa2d348ac842c168b0a4d3d61919687216dfc547976d853", size = 16645538, upload-time = "2026-05-18T23:34:09.265Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/numpy/2.4.6/numpy-2.4.6-cp312-cp312-musllinux_1_2_aarch64.whl", hash = "sha256:c1a2af6c6ef86344a6b0db6b97834208bf598db514f2b155042439b62605601a", size = 17020706, upload-time = "2026-05-18T23:34:13.053Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/numpy/2.4.6/numpy-2.4.6-cp312-cp312-musllinux_1_2_x86_64.whl", hash = "sha256:e5805d5a22fd19c8ccff10a9561f9df94436b0545619ea579db2d3c35294bce2", size = 18368541, upload-time = "2026-05-18T23:34:17.024Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/numpy/2.4.6/numpy-2.4.6-cp312-cp312-win32.whl", hash = "sha256:e3eeb0aabd6bd5ce64faae67e9935203a6991b4bc2a485a767fbafb2c5125f45", size = 5962825, upload-time = "2026-05-18T23:34:20.3Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/numpy/2.4.6/numpy-2.4.6-cp312-cp312-win_amd64.whl", hash = "sha256:d8e8286dd7cea7895157318d1b91cdacac64c479f3cbc8dce548331728484751", size = 12321687, upload-time = "2026-05-18T23:34:23.095Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/numpy/2.4.6/numpy-2.4.6-cp312-cp312-win_arm64.whl", hash = "sha256:4081eb135ac24158bd51cdfbef16f1c64df7063b1143f24731387137c092bec8", size = 10221482, upload-time = "2026-05-18T23:34:25.876Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/numpy/2.4.6/numpy-2.4.6-pp311-pypy311_pp73-macosx_10_15_x86_64.whl", hash = "sha256:55cced7c52e981362f708ad635198e97a752dfba412cc03c23bbf3bd8d5cd662", size = 16847511, upload-time = "2026-05-18T23:36:50.673Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/numpy/2.4.6/numpy-2.4.6-pp311-pypy311_pp73-macosx_11_0_arm64.whl", hash = "sha256:d6da64deb6b8ed903e7560180a92f2d804ee1ba5eeb849ac2748b8c1aba1f6d7", size = 14889064, upload-time = "2026-05-18T23:36:53.879Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/numpy/2.4.6/numpy-2.4.6-pp311-pypy311_pp73-macosx_14_0_arm64.whl", hash = "sha256:68a5124b13fa6cc2086764a20005d30bc0548146f7f5322f02fce212ca14317f", size = 5394157, upload-time = "2026-05-18T23:36:57.194Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/numpy/2.4.6/numpy-2.4.6-pp311-pypy311_pp73-macosx_14_0_x86_64.whl", hash = "sha256:948424b06129ce883307e8cff868c31396d8dc7630a59c61d70d98dbe70f222c", size = 6708728, upload-time = "2026-05-18T23:36:59.575Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/numpy/2.4.6/numpy-2.4.6-pp311-pypy311_pp73-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:5dbbdb29840ca3d91ee0fece42fc29278886d908280bfec0a5846c6f901a3eb0", size = 15798374, upload-time = "2026-05-18T23:37:02.674Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/numpy/2.4.6/numpy-2.4.6-pp311-pypy311_pp73-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:8ad03c0965fb3c692200e74d458ca28c1dbb4ce96f9a479a8aa041ad5fabca02", size = 16747286, upload-time = "2026-05-18T23:37:06.327Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/numpy/2.4.6/numpy-2.4.6-pp311-pypy311_pp73-win_amd64.whl", hash = "sha256:2803abfebfc990042cd494d8ce2d5f82e9d847af6d35ec486923aa19dbad5e73", size = 12504263, upload-time = "2026-05-18T23:37:09.715Z" }, ] [[package]] name = "numpy" version = "2.5.2" -source = { registry = "https://pypi.org/simple" } +source = { registry = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/" } resolution-markers = [ "python_full_version >= '3.12'", ] -sdist = { url = "https://files.pythonhosted.org/packages/9a/80/db0b4559e57ec36362bedbb05530a87fafbcb6067708c946967a41d449e7/numpy-2.5.2.tar.gz", hash = "sha256:d482d171c406ae88c5b19cad3b6a1c4c5209f886ab74bc44c2c865c23f52d860", size = 20773161, upload-time = "2026-08-09T13:48:27.962Z" } +sdist = { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/numpy/2.5.2/numpy-2.5.2.tar.gz", hash = "sha256:d482d171c406ae88c5b19cad3b6a1c4c5209f886ab74bc44c2c865c23f52d860", size = 20773161, upload-time = "2026-08-09T13:48:27.962Z" } wheels = [ - { url = "https://files.pythonhosted.org/packages/69/72/dccb0aaf40972777283303919f613964227266d0c13adebb79ac124f1c3e/numpy-2.5.2-cp312-cp312-macosx_10_13_x86_64.whl", hash = "sha256:14e373cfc6387177e8409dac3c7159be8eb05cd77096cd7c950268b86f62831c", size = 16891693, upload-time = "2026-08-09T13:44:51.702Z" }, - { url = "https://files.pythonhosted.org/packages/60/2e/b5aee50a1f74ac815cf8331812cb8251e29024025de462e0c047641c614c/numpy-2.5.2-cp312-cp312-macosx_11_0_arm64.whl", hash = "sha256:4bbd96c833ecc8cc069ce518078fc8c60cb9cbfb0fea5b7a803ad65035596d03", size = 11903109, upload-time = "2026-08-09T13:44:55.501Z" }, - { url = "https://files.pythonhosted.org/packages/f3/f4/29e78102a80601cf034d4e9767022cffeca2c3b4c926e1754572ca95593d/numpy-2.5.2-cp312-cp312-macosx_14_0_arm64.whl", hash = "sha256:6e8172ddfcf5cf74b811d372b570b83c60bd2de87a6fbfbebdadb4a9bd9c6cbb", size = 5350202, upload-time = "2026-08-09T13:44:58.401Z" }, - { url = "https://files.pythonhosted.org/packages/11/4b/dcd3b7eadaf4035d2c7a4289d232523a6964f602598ef7674e4bd7291f93/numpy-2.5.2-cp312-cp312-macosx_14_0_x86_64.whl", hash = "sha256:65f188481f1669e26f62b701e8205d19e460fa4a9b52a1414ba382330e4a3414", size = 6687736, upload-time = "2026-08-09T13:45:00.813Z" }, - { url = "https://files.pythonhosted.org/packages/e5/21/4947e0e9d6c9fc2e2ff15b8949049ee44f63adb9cacc729ab8793f97e712/numpy-2.5.2-cp312-cp312-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:8ee9c4eeb8454b3660a8b53493563c3e121c2fc94fbd72b848ef814ed7b676a9", size = 15612696, upload-time = "2026-08-09T13:45:04.151Z" }, - { url = "https://files.pythonhosted.org/packages/3a/5f/62d28cf019460c7f1394105b4d49d9911a9c444cb77ab0bd95a204c5a6de/numpy-2.5.2-cp312-cp312-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:3cdec01fa790a186d430433fdd4d4ffb70eed6f0eeb4bf05c8dbe2dce0a9bcb8", size = 16722264, upload-time = "2026-08-09T13:45:07.714Z" }, - { url = "https://files.pythonhosted.org/packages/14/25/3f0be4c1b9fdf5dd5e708a6806978564d7c46a055c000496309ff2a2f8af/numpy-2.5.2-cp312-cp312-musllinux_1_2_aarch64.whl", hash = "sha256:7999d4ddb0c4025018373fd787510d46e04c769467af22869707b3c1cfd459ab", size = 16974396, upload-time = "2026-08-09T13:45:11.316Z" }, - { url = "https://files.pythonhosted.org/packages/22/72/6262cbdeeb45da9d971e40715f579d791603ba8ec0b5e2db1ac55454421d/numpy-2.5.2-cp312-cp312-musllinux_1_2_x86_64.whl", hash = "sha256:c1f017dc0875c9209d219f97feceb7d54c2661bb243deb4114478e1295808af7", size = 18476044, upload-time = "2026-08-09T13:45:14.869Z" }, - { url = "https://files.pythonhosted.org/packages/36/33/29208b8b075bde62d26a81d14b358c42b0f69b6cabd98d4ff97f37f22b05/numpy-2.5.2-cp312-cp312-win32.whl", hash = "sha256:d6a48072864e3324e194a8fbb3c657bcc5b5c869dbc64c9537b1d5c862572c0a", size = 6072817, upload-time = "2026-08-09T13:45:17.867Z" }, - { url = "https://files.pythonhosted.org/packages/7f/b9/87fea2769fe1c47c1b5b01d8310772c9d1a85d485de7cf386ef7a3332b02/numpy-2.5.2-cp312-cp312-win_amd64.whl", hash = "sha256:28ac63476ec7651484215ee7fa15a1f78b57c14621f01e392afe17b9a1390ce4", size = 12464674, upload-time = "2026-08-09T13:45:20.734Z" }, - { url = "https://files.pythonhosted.org/packages/14/52/032b97e00461ab0809bbe4c588b035620e5a14b8cdee47ecddefc7b17d33/numpy-2.5.2-cp312-cp312-win_arm64.whl", hash = "sha256:27650bb0e7140fa3d37b9923b4803645e0b125d190f326eecfd3f4dad8e8ade1", size = 10397131, upload-time = "2026-08-09T13:45:23.73Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/numpy/2.5.2/numpy-2.5.2-cp312-cp312-macosx_10_13_x86_64.whl", hash = "sha256:14e373cfc6387177e8409dac3c7159be8eb05cd77096cd7c950268b86f62831c", size = 16891693, upload-time = "2026-08-09T13:44:51.702Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/numpy/2.5.2/numpy-2.5.2-cp312-cp312-macosx_11_0_arm64.whl", hash = "sha256:4bbd96c833ecc8cc069ce518078fc8c60cb9cbfb0fea5b7a803ad65035596d03", size = 11903109, upload-time = "2026-08-09T13:44:55.501Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/numpy/2.5.2/numpy-2.5.2-cp312-cp312-macosx_14_0_arm64.whl", hash = "sha256:6e8172ddfcf5cf74b811d372b570b83c60bd2de87a6fbfbebdadb4a9bd9c6cbb", size = 5350202, upload-time = "2026-08-09T13:44:58.401Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/numpy/2.5.2/numpy-2.5.2-cp312-cp312-macosx_14_0_x86_64.whl", hash = "sha256:65f188481f1669e26f62b701e8205d19e460fa4a9b52a1414ba382330e4a3414", size = 6687736, upload-time = "2026-08-09T13:45:00.813Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/numpy/2.5.2/numpy-2.5.2-cp312-cp312-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:8ee9c4eeb8454b3660a8b53493563c3e121c2fc94fbd72b848ef814ed7b676a9", size = 15612696, upload-time = "2026-08-09T13:45:04.151Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/numpy/2.5.2/numpy-2.5.2-cp312-cp312-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:3cdec01fa790a186d430433fdd4d4ffb70eed6f0eeb4bf05c8dbe2dce0a9bcb8", size = 16722264, upload-time = "2026-08-09T13:45:07.714Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/numpy/2.5.2/numpy-2.5.2-cp312-cp312-musllinux_1_2_aarch64.whl", hash = "sha256:7999d4ddb0c4025018373fd787510d46e04c769467af22869707b3c1cfd459ab", size = 16974396, upload-time = "2026-08-09T13:45:11.316Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/numpy/2.5.2/numpy-2.5.2-cp312-cp312-musllinux_1_2_x86_64.whl", hash = "sha256:c1f017dc0875c9209d219f97feceb7d54c2661bb243deb4114478e1295808af7", size = 18476044, upload-time = "2026-08-09T13:45:14.869Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/numpy/2.5.2/numpy-2.5.2-cp312-cp312-win32.whl", hash = "sha256:d6a48072864e3324e194a8fbb3c657bcc5b5c869dbc64c9537b1d5c862572c0a", size = 6072817, upload-time = "2026-08-09T13:45:17.867Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/numpy/2.5.2/numpy-2.5.2-cp312-cp312-win_amd64.whl", hash = "sha256:28ac63476ec7651484215ee7fa15a1f78b57c14621f01e392afe17b9a1390ce4", size = 12464674, upload-time = "2026-08-09T13:45:20.734Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/numpy/2.5.2/numpy-2.5.2-cp312-cp312-win_arm64.whl", hash = "sha256:27650bb0e7140fa3d37b9923b4803645e0b125d190f326eecfd3f4dad8e8ade1", size = 10397131, upload-time = "2026-08-09T13:45:23.73Z" }, ] [[package]] name = "opentelemetry-api" version = "1.44.0" -source = { registry = "https://pypi.org/simple" } +source = { registry = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/" } dependencies = [ { name = "typing-extensions" }, ] -sdist = { url = "https://files.pythonhosted.org/packages/ee/8b/aa9e2d8b8dfa7c946f7dec5d1f8f6ba8eca062f43509a06bdb5ce93d26c0/opentelemetry_api-1.44.0.tar.gz", hash = "sha256:67647e5e9566edcf421166fdf022b3537f818635daa852b289e34604dc6fb33a", size = 72406, upload-time = "2026-07-16T15:25:32.678Z" } +sdist = { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/opentelemetry-api/1.44.0/opentelemetry_api-1.44.0.tar.gz", hash = "sha256:67647e5e9566edcf421166fdf022b3537f818635daa852b289e34604dc6fb33a", size = 72406, upload-time = "2026-07-16T15:25:32.678Z" } wheels = [ - { url = "https://files.pythonhosted.org/packages/ca/6f/a04e900f465ff3221ccc395522503e2d10e79fa21f2723c8e177aae1e0d1/opentelemetry_api-1.44.0-py3-none-any.whl", hash = "sha256:94b98c893a91b88657eaac1e3ba89618cdb85be6918196705354f34728b2cdef", size = 60018, upload-time = "2026-07-16T15:25:11.657Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/opentelemetry-api/1.44.0/opentelemetry_api-1.44.0-py3-none-any.whl", hash = "sha256:94b98c893a91b88657eaac1e3ba89618cdb85be6918196705354f34728b2cdef", size = 60018, upload-time = "2026-07-16T15:25:11.657Z" }, ] [[package]] name = "opentelemetry-exporter-otlp-proto-common" version = "1.44.0" -source = { registry = "https://pypi.org/simple" } +source = { registry = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/" } dependencies = [ { name = "opentelemetry-proto" }, ] -sdist = { url = "https://files.pythonhosted.org/packages/61/09/4d717852c1cf3f854b76c7110a5d00883bc3c99288b9b0dbcbeb9e306eb6/opentelemetry_exporter_otlp_proto_common-1.44.0.tar.gz", hash = "sha256:dc87a5a5bc58f149a56d1547e4691588fa12994cdc3bc039a694ccb3375862ac", size = 20202, upload-time = "2026-07-16T15:25:37.658Z" } +sdist = { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/opentelemetry-exporter-otlp-proto-common/1.44.0/opentelemetry_exporter_otlp_proto_common-1.44.0.tar.gz", hash = "sha256:dc87a5a5bc58f149a56d1547e4691588fa12994cdc3bc039a694ccb3375862ac", size = 20202, upload-time = "2026-07-16T15:25:37.658Z" } wheels = [ - { url = "https://files.pythonhosted.org/packages/5e/71/65fd9d54c10b860f87c045ccee1264cab7011268895d3528818a29c1172a/opentelemetry_exporter_otlp_proto_common-1.44.0-py3-none-any.whl", hash = "sha256:9a9fe61bba73d802904bc989f1d6b4a7b1ee40f06c40e98d6f85af65aaebb694", size = 17045, upload-time = "2026-07-16T15:25:18.201Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/opentelemetry-exporter-otlp-proto-common/1.44.0/opentelemetry_exporter_otlp_proto_common-1.44.0-py3-none-any.whl", hash = "sha256:9a9fe61bba73d802904bc989f1d6b4a7b1ee40f06c40e98d6f85af65aaebb694", size = 17045, upload-time = "2026-07-16T15:25:18.201Z" }, ] [[package]] name = "opentelemetry-exporter-otlp-proto-http" version = "1.44.0" -source = { registry = "https://pypi.org/simple" } +source = { registry = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/" } dependencies = [ { name = "googleapis-common-protos" }, { name = "opentelemetry-api" }, @@ -703,273 +703,273 @@ dependencies = [ { name = "requests" }, { name = "typing-extensions" }, ] -sdist = { url = "https://files.pythonhosted.org/packages/1a/87/95e2a5aaa795b4e2260d74e16df2d5541deb2ea9de010bcd615f4dee2654/opentelemetry_exporter_otlp_proto_http-1.44.0.tar.gz", hash = "sha256:c633d7270ad6b57cd4cfbe8b0007a9e2e7c0cb50bd6c50fe2a7b245f721a09d8", size = 25806, upload-time = "2026-07-16T15:25:39.162Z" } +sdist = { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/opentelemetry-exporter-otlp-proto-http/1.44.0/opentelemetry_exporter_otlp_proto_http-1.44.0.tar.gz", hash = "sha256:c633d7270ad6b57cd4cfbe8b0007a9e2e7c0cb50bd6c50fe2a7b245f721a09d8", size = 25806, upload-time = "2026-07-16T15:25:39.162Z" } wheels = [ - { url = "https://files.pythonhosted.org/packages/cd/d0/fdeb1a98d8d3a6205f5f297c51b4a9bfe65126ab60339669bbe3dd54c2e2/opentelemetry_exporter_otlp_proto_http-1.44.0-py3-none-any.whl", hash = "sha256:838592fce774c1c8bb7b9a0a7facbfa82e17be5a8a4e94cef10cb84ae026bae3", size = 21850, upload-time = "2026-07-16T15:25:20.006Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/opentelemetry-exporter-otlp-proto-http/1.44.0/opentelemetry_exporter_otlp_proto_http-1.44.0-py3-none-any.whl", hash = "sha256:838592fce774c1c8bb7b9a0a7facbfa82e17be5a8a4e94cef10cb84ae026bae3", size = 21850, upload-time = "2026-07-16T15:25:20.006Z" }, ] [[package]] name = "opentelemetry-proto" version = "1.44.0" -source = { registry = "https://pypi.org/simple" } +source = { registry = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/" } dependencies = [ { name = "protobuf" }, ] -sdist = { url = "https://files.pythonhosted.org/packages/64/01/40ac4ae9a149263cc52c2cee200ddd80cb6d8db1a4610abf8eabce0fe771/opentelemetry_proto-1.44.0.tar.gz", hash = "sha256:c547a79c2f8c0c515d31509154682e5921c7cfd5ca67b70e1f9266e2c3e103f3", size = 46488, upload-time = "2026-07-16T15:25:45.34Z" } +sdist = { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/opentelemetry-proto/1.44.0/opentelemetry_proto-1.44.0.tar.gz", hash = "sha256:c547a79c2f8c0c515d31509154682e5921c7cfd5ca67b70e1f9266e2c3e103f3", size = 46488, upload-time = "2026-07-16T15:25:45.34Z" } wheels = [ - { url = "https://files.pythonhosted.org/packages/d1/7c/8be563d68e93bbefa5c8affb82ddcff91b3ad858ce49957ba7b16fd3e0ab/opentelemetry_proto-1.44.0-py3-none-any.whl", hash = "sha256:898b155a0e1557afd867478fb6158e8122a46329ca0bb8dc53cc55e98f017f56", size = 72483, upload-time = "2026-07-16T15:25:28.429Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/opentelemetry-proto/1.44.0/opentelemetry_proto-1.44.0-py3-none-any.whl", hash = "sha256:898b155a0e1557afd867478fb6158e8122a46329ca0bb8dc53cc55e98f017f56", size = 72483, upload-time = "2026-07-16T15:25:28.429Z" }, ] [[package]] name = "opentelemetry-sdk" version = "1.44.0" -source = { registry = "https://pypi.org/simple" } +source = { registry = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/" } dependencies = [ { name = "opentelemetry-api" }, { name = "opentelemetry-semantic-conventions" }, { name = "typing-extensions" }, ] -sdist = { url = "https://files.pythonhosted.org/packages/5d/77/a6592cbc7c8d9bcc9d6757a9df45e04a7c585e3e6e7a13456da522b21109/opentelemetry_sdk-1.44.0.tar.gz", hash = "sha256:cebe7f65dc12f26ead75c6064de12fd2a9052e5060c0272d402cfa203aae123b", size = 208624, upload-time = "2026-07-16T15:25:46.078Z" } +sdist = { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/opentelemetry-sdk/1.44.0/opentelemetry_sdk-1.44.0.tar.gz", hash = "sha256:cebe7f65dc12f26ead75c6064de12fd2a9052e5060c0272d402cfa203aae123b", size = 208624, upload-time = "2026-07-16T15:25:46.078Z" } wheels = [ - { url = "https://files.pythonhosted.org/packages/e7/23/ff077e61886ee020a17ce9c8b6fa11c601c8d8345b09ea24f605445df62a/opentelemetry_sdk-1.44.0-py3-none-any.whl", hash = "sha256:df081c4c6bcfdb1211e3e86140376792643128a25f8d72d1d27675936e7e96ad", size = 137221, upload-time = "2026-07-16T15:25:29.534Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/opentelemetry-sdk/1.44.0/opentelemetry_sdk-1.44.0-py3-none-any.whl", hash = "sha256:df081c4c6bcfdb1211e3e86140376792643128a25f8d72d1d27675936e7e96ad", size = 137221, upload-time = "2026-07-16T15:25:29.534Z" }, ] [[package]] name = "opentelemetry-semantic-conventions" version = "0.65b0" -source = { registry = "https://pypi.org/simple" } +source = { registry = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/" } dependencies = [ { name = "opentelemetry-api" }, { name = "typing-extensions" }, ] -sdist = { url = "https://files.pythonhosted.org/packages/8f/73/0cbdebcb4cf545fdd328da14f5137e37d0770c3f26185e478b0d15d94f50/opentelemetry_semantic_conventions-0.65b0.tar.gz", hash = "sha256:f9b2b81e9d5b64f11bc952075e7e9c7fb0aab075c7fd1c46d597f1b919852d60", size = 148774, upload-time = "2026-07-16T15:25:46.902Z" } +sdist = { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/opentelemetry-semantic-conventions/0.65b0/opentelemetry_semantic_conventions-0.65b0.tar.gz", hash = "sha256:f9b2b81e9d5b64f11bc952075e7e9c7fb0aab075c7fd1c46d597f1b919852d60", size = 148774, upload-time = "2026-07-16T15:25:46.902Z" } wheels = [ - { url = "https://files.pythonhosted.org/packages/a6/0e/49df70d9b81fb5cbae4bbf2a49d865b09bcbcbc4eb53f5851b1027738d78/opentelemetry_semantic_conventions-0.65b0-py3-none-any.whl", hash = "sha256:1cacde7b0ad306f84c5ef08c3dbe1bbaf20165bba6f8bff43b670e555a086bcb", size = 204645, upload-time = "2026-07-16T15:25:30.688Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/opentelemetry-semantic-conventions/0.65b0/opentelemetry_semantic_conventions-0.65b0-py3-none-any.whl", hash = "sha256:1cacde7b0ad306f84c5ef08c3dbe1bbaf20165bba6f8bff43b670e555a086bcb", size = 204645, upload-time = "2026-07-16T15:25:30.688Z" }, ] [[package]] name = "orjson" version = "3.12.0" -source = { registry = "https://pypi.org/simple" } -sdist = { url = "https://files.pythonhosted.org/packages/0f/f3/742fb1f62b825f2c010697eaf4e828004bc2a81e7e806666989c132c7c42/orjson-3.12.0.tar.gz", hash = "sha256:d14203fb1aae2ad9b3d52f8a0e82aeb10197ef1c9bc61da7f358bd70b00123d5", size = 4142915, upload-time = "2026-08-14T16:13:30.607Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/75/1a/a7075a8e8b0d3f5097d17ac3099017104b6b7b42012041147995d5b2da05/orjson-3.12.0-cp311-cp311-macosx_10_15_x86_64.macosx_11_0_arm64.macosx_10_15_universal2.whl", hash = "sha256:a94f0f0c6fcbb2b5bd9734c57a489c7584a732bbdf04a39e8c83b861e9d03e92", size = 223409, upload-time = "2026-08-14T16:12:12.654Z" }, - { url = "https://files.pythonhosted.org/packages/05/34/c2eb3b2900e5597db7841a4c6416ac2d90081bd956b02d4dd1833fa2b96b/orjson-3.12.0-cp311-cp311-macosx_15_0_arm64.whl", hash = "sha256:a696529ec96a90d9a5f9570207efe403c8b08f8e4aa2783ee3403511e2fdfa10", size = 124015, upload-time = "2026-08-14T16:12:14.025Z" }, - { url = "https://files.pythonhosted.org/packages/1c/df/b49081766a75b6a37b3d33bdc0a39e492abab8441dd25e3e1998e7b83fcb/orjson-3.12.0-cp311-cp311-manylinux2014_armv7l.manylinux_2_17_armv7l.whl", hash = "sha256:e4ac5059baab4b3acbd99485de019ff8cda0fdf34b61fa74f7197a53db78bfe8", size = 113471, upload-time = "2026-08-14T16:12:15.81Z" }, - { url = "https://files.pythonhosted.org/packages/48/d4/58ea28eeef95c2a27358ed927380a621162cf20bd740bbccf9c3f09a200a/orjson-3.12.0-cp311-cp311-manylinux2014_i686.manylinux_2_17_i686.whl", hash = "sha256:8e29957429c35bbb5a185a119c523aa2428b7bbf1a293724c7b9375ed8f892a3", size = 129998, upload-time = "2026-08-14T16:12:17.503Z" }, - { url = "https://files.pythonhosted.org/packages/e2/f4/1e82aa2efc9916422d804697876ce433c907a1abd7c7e5c6d3d48565e5f9/orjson-3.12.0-cp311-cp311-manylinux_2_17_aarch64.manylinux2014_aarch64.whl", hash = "sha256:dce0166feb0a737ab84f598c9a338cbc0b764a036617aa686194f53c7eba0c3e", size = 130891, upload-time = "2026-08-14T16:12:18.762Z" }, - { url = "https://files.pythonhosted.org/packages/5b/e1/15169e9d22b59a406264f99d6db387c0b0b12b6357a8a0169917c2a713eb/orjson-3.12.0-cp311-cp311-manylinux_2_17_x86_64.manylinux2014_x86_64.whl", hash = "sha256:9caf3d09f47c3c70c4451ada20ef9bc4a4cdffa26f49862cf0a253b329aae2d5", size = 131285, upload-time = "2026-08-14T16:12:20.251Z" }, - { url = "https://files.pythonhosted.org/packages/a4/3a/763dbd426290d044ec3e615a05e70adb6d8b6f95bf17dc355c0081a5e8b6/orjson-3.12.0-cp311-cp311-musllinux_1_2_aarch64.whl", hash = "sha256:b9dca132b1fda5565088e65a6b6e742285e0aeceb6fae549fa8863e16c7d3998", size = 135707, upload-time = "2026-08-14T16:12:21.652Z" }, - { url = "https://files.pythonhosted.org/packages/04/d1/3b2038ed168d22e14182ed715d6963f9c073a83a2ba43cfe918a4fc43c64/orjson-3.12.0-cp311-cp311-musllinux_1_2_x86_64.whl", hash = "sha256:a791f793b287bbc135b8e87c34e35c8bfc693e2a8a620fab1ae682b925f9a32e", size = 127669, upload-time = "2026-08-14T16:12:22.926Z" }, - { url = "https://files.pythonhosted.org/packages/88/ae/b84b3d3e65f5629ada0edcb1d2bccc55d7c5f89d8b981537ecdc3d6f31ec/orjson-3.12.0-cp311-cp311-win32.whl", hash = "sha256:31ed278a36304390adc3eec5d7f6fd593a7c3e99e5a06cd07866396c4b1b4710", size = 128043, upload-time = "2026-08-14T16:12:24.367Z" }, - { url = "https://files.pythonhosted.org/packages/35/24/2ed0e6f51ea3d0af45d807233a851175af75bec83ef5fd0d6a2601904ec0/orjson-3.12.0-cp311-cp311-win_amd64.whl", hash = "sha256:fb2539159dfe8d371914f354360fa50e4a577cc89222a3828b9650a5e5040252", size = 122084, upload-time = "2026-08-14T16:12:25.813Z" }, - { url = "https://files.pythonhosted.org/packages/21/dd/95d25fcfbc9471799ef6bb01c552d64ee5cde93ee40ba2f423dd3442c708/orjson-3.12.0-cp311-cp311-win_arm64.whl", hash = "sha256:61318b6de893c7a9d9f3e5ecbadccbfc26a7eb417ccc7bbf0771de3b4d72f868", size = 127035, upload-time = "2026-08-14T16:12:27.201Z" }, - { url = "https://files.pythonhosted.org/packages/be/4a/295da39c651c2faac8bd351a2a346f0fdedd9d50b847ee9dfc27d2207ef6/orjson-3.12.0-cp312-cp312-macosx_10_15_x86_64.macosx_11_0_arm64.macosx_10_15_universal2.whl", hash = "sha256:aa3e43a6846e91d7bde3d5a9c66090fcd8744f569a9b6cffc5e1ca38f6a461c0", size = 223427, upload-time = "2026-08-14T16:12:28.525Z" }, - { url = "https://files.pythonhosted.org/packages/29/98/758cf90fbeaaafb7f8141bfac75a432099959f3a2f5db93a412e876415d8/orjson-3.12.0-cp312-cp312-macosx_15_0_arm64.whl", hash = "sha256:11edb4660a6680abee9788a3a9072208a2c96538cc1322bd79542065229d8e54", size = 123725, upload-time = "2026-08-14T16:12:30.013Z" }, - { url = "https://files.pythonhosted.org/packages/32/b5/5b934d251f8651f7e41df180ad0c57a6e1cabe15c7bd331638413a50ebc9/orjson-3.12.0-cp312-cp312-manylinux2014_armv7l.manylinux_2_17_armv7l.whl", hash = "sha256:2d3a9da945a4d96ae758fdaaca56742e6b73b6fd554c5d8876f252a6dad70b83", size = 113375, upload-time = "2026-08-14T16:12:31.209Z" }, - { url = "https://files.pythonhosted.org/packages/cd/d2/37efb5b12a176ce3ced29f4144f20da57d02757f78ce549637dc1b4e1fc8/orjson-3.12.0-cp312-cp312-manylinux2014_i686.manylinux_2_17_i686.whl", hash = "sha256:92ffc09e07233a6ab6d4e067f7841edcbcc134cb4812155cf171ea5255a421d7", size = 129983, upload-time = "2026-08-14T16:12:32.721Z" }, - { url = "https://files.pythonhosted.org/packages/50/22/0644b87c73f13e0092df8f35a1fe280d991e5e90072087411e0dd7e44e0c/orjson-3.12.0-cp312-cp312-manylinux_2_17_aarch64.manylinux2014_aarch64.whl", hash = "sha256:bf44e374aadde77b1f6109f1030be51433eb61984379852766b6f4e187db7b1e", size = 130629, upload-time = "2026-08-14T16:12:34.084Z" }, - { url = "https://files.pythonhosted.org/packages/8c/57/80b986ebfecd9c6a177ddf1c2319717f0cd8feffb2b78946595a18a2fc88/orjson-3.12.0-cp312-cp312-manylinux_2_17_x86_64.manylinux2014_x86_64.whl", hash = "sha256:1192a7021b6d071aaf909864f6e924d6a2675ca360485b972b8401749311750b", size = 131245, upload-time = "2026-08-14T16:12:35.713Z" }, - { url = "https://files.pythonhosted.org/packages/80/3d/75c5ac5a69161f44492a68fbdde66f4cc4ce48cd5e1fb05918e46f0c8848/orjson-3.12.0-cp312-cp312-musllinux_1_2_aarch64.whl", hash = "sha256:53c0c474a9d9aff9aebfc0c88de1f28f843d940e6e3a80729abdf6a20274356f", size = 135397, upload-time = "2026-08-14T16:12:37.128Z" }, - { url = "https://files.pythonhosted.org/packages/71/93/4d71f2df314a97ff0d27a4559bf5888fc8406e3c6dec90e92291e3511215/orjson-3.12.0-cp312-cp312-musllinux_1_2_x86_64.whl", hash = "sha256:532ff8cd4bd59a327a953a7dcde922c7fc25b85e29721bb8633265430d3a3873", size = 127693, upload-time = "2026-08-14T16:12:38.627Z" }, - { url = "https://files.pythonhosted.org/packages/bc/1d/0dbc6be5adfd1730491072fb60beb6bcdf5d7b2596ee41b7fc2e298bfc09/orjson-3.12.0-cp312-cp312-win32.whl", hash = "sha256:a6cf4b18e7de173f209f2084ffbd736dd72389a396326ee80a7022168be232e5", size = 128000, upload-time = "2026-08-14T16:12:39.954Z" }, - { url = "https://files.pythonhosted.org/packages/2d/c9/97b1ce0112ebf5e949c775ed5b1755e562233179f3584579673cc24d6378/orjson-3.12.0-cp312-cp312-win_amd64.whl", hash = "sha256:010811c1b69773450a01cef97727a67b223242f350b77d4ca000e59a9ef2155a", size = 122106, upload-time = "2026-08-14T16:12:41.324Z" }, - { url = "https://files.pythonhosted.org/packages/a8/6a/facd8b312e4a0d3a7fa978c7e15821f74a336adf1d65529faec33b48e18b/orjson-3.12.0-cp312-cp312-win_arm64.whl", hash = "sha256:ad29eece0c601737f2a60edc2752a84e7a0785df3efb62e3012834700a5afe0d", size = 126869, upload-time = "2026-08-14T16:12:42.651Z" }, +source = { registry = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/" } +sdist = { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/orjson/3.12.0/orjson-3.12.0.tar.gz", hash = "sha256:d14203fb1aae2ad9b3d52f8a0e82aeb10197ef1c9bc61da7f358bd70b00123d5", size = 4142915, upload-time = "2026-08-14T16:13:30.607Z" } +wheels = [ + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/orjson/3.12.0/orjson-3.12.0-cp311-cp311-macosx_10_15_x86_64.macosx_11_0_arm64.macosx_10_15_universal2.whl", hash = "sha256:a94f0f0c6fcbb2b5bd9734c57a489c7584a732bbdf04a39e8c83b861e9d03e92", size = 223409, upload-time = "2026-08-14T16:12:12.654Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/orjson/3.12.0/orjson-3.12.0-cp311-cp311-macosx_15_0_arm64.whl", hash = "sha256:a696529ec96a90d9a5f9570207efe403c8b08f8e4aa2783ee3403511e2fdfa10", size = 124015, upload-time = "2026-08-14T16:12:14.025Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/orjson/3.12.0/orjson-3.12.0-cp311-cp311-manylinux2014_armv7l.manylinux_2_17_armv7l.whl", hash = "sha256:e4ac5059baab4b3acbd99485de019ff8cda0fdf34b61fa74f7197a53db78bfe8", size = 113471, upload-time = "2026-08-14T16:12:15.81Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/orjson/3.12.0/orjson-3.12.0-cp311-cp311-manylinux2014_i686.manylinux_2_17_i686.whl", hash = "sha256:8e29957429c35bbb5a185a119c523aa2428b7bbf1a293724c7b9375ed8f892a3", size = 129998, upload-time = "2026-08-14T16:12:17.503Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/orjson/3.12.0/orjson-3.12.0-cp311-cp311-manylinux_2_17_aarch64.manylinux2014_aarch64.whl", hash = "sha256:dce0166feb0a737ab84f598c9a338cbc0b764a036617aa686194f53c7eba0c3e", size = 130891, upload-time = "2026-08-14T16:12:18.762Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/orjson/3.12.0/orjson-3.12.0-cp311-cp311-manylinux_2_17_x86_64.manylinux2014_x86_64.whl", hash = "sha256:9caf3d09f47c3c70c4451ada20ef9bc4a4cdffa26f49862cf0a253b329aae2d5", size = 131285, upload-time = "2026-08-14T16:12:20.251Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/orjson/3.12.0/orjson-3.12.0-cp311-cp311-musllinux_1_2_aarch64.whl", hash = "sha256:b9dca132b1fda5565088e65a6b6e742285e0aeceb6fae549fa8863e16c7d3998", size = 135707, upload-time = "2026-08-14T16:12:21.652Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/orjson/3.12.0/orjson-3.12.0-cp311-cp311-musllinux_1_2_x86_64.whl", hash = "sha256:a791f793b287bbc135b8e87c34e35c8bfc693e2a8a620fab1ae682b925f9a32e", size = 127669, upload-time = "2026-08-14T16:12:22.926Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/orjson/3.12.0/orjson-3.12.0-cp311-cp311-win32.whl", hash = "sha256:31ed278a36304390adc3eec5d7f6fd593a7c3e99e5a06cd07866396c4b1b4710", size = 128043, upload-time = "2026-08-14T16:12:24.367Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/orjson/3.12.0/orjson-3.12.0-cp311-cp311-win_amd64.whl", hash = "sha256:fb2539159dfe8d371914f354360fa50e4a577cc89222a3828b9650a5e5040252", size = 122084, upload-time = "2026-08-14T16:12:25.813Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/orjson/3.12.0/orjson-3.12.0-cp311-cp311-win_arm64.whl", hash = "sha256:61318b6de893c7a9d9f3e5ecbadccbfc26a7eb417ccc7bbf0771de3b4d72f868", size = 127035, upload-time = "2026-08-14T16:12:27.201Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/orjson/3.12.0/orjson-3.12.0-cp312-cp312-macosx_10_15_x86_64.macosx_11_0_arm64.macosx_10_15_universal2.whl", hash = "sha256:aa3e43a6846e91d7bde3d5a9c66090fcd8744f569a9b6cffc5e1ca38f6a461c0", size = 223427, upload-time = "2026-08-14T16:12:28.525Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/orjson/3.12.0/orjson-3.12.0-cp312-cp312-macosx_15_0_arm64.whl", hash = "sha256:11edb4660a6680abee9788a3a9072208a2c96538cc1322bd79542065229d8e54", size = 123725, upload-time = "2026-08-14T16:12:30.013Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/orjson/3.12.0/orjson-3.12.0-cp312-cp312-manylinux2014_armv7l.manylinux_2_17_armv7l.whl", hash = "sha256:2d3a9da945a4d96ae758fdaaca56742e6b73b6fd554c5d8876f252a6dad70b83", size = 113375, upload-time = "2026-08-14T16:12:31.209Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/orjson/3.12.0/orjson-3.12.0-cp312-cp312-manylinux2014_i686.manylinux_2_17_i686.whl", hash = "sha256:92ffc09e07233a6ab6d4e067f7841edcbcc134cb4812155cf171ea5255a421d7", size = 129983, upload-time = "2026-08-14T16:12:32.721Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/orjson/3.12.0/orjson-3.12.0-cp312-cp312-manylinux_2_17_aarch64.manylinux2014_aarch64.whl", hash = "sha256:bf44e374aadde77b1f6109f1030be51433eb61984379852766b6f4e187db7b1e", size = 130629, upload-time = "2026-08-14T16:12:34.084Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/orjson/3.12.0/orjson-3.12.0-cp312-cp312-manylinux_2_17_x86_64.manylinux2014_x86_64.whl", hash = "sha256:1192a7021b6d071aaf909864f6e924d6a2675ca360485b972b8401749311750b", size = 131245, upload-time = "2026-08-14T16:12:35.713Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/orjson/3.12.0/orjson-3.12.0-cp312-cp312-musllinux_1_2_aarch64.whl", hash = "sha256:53c0c474a9d9aff9aebfc0c88de1f28f843d940e6e3a80729abdf6a20274356f", size = 135397, upload-time = "2026-08-14T16:12:37.128Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/orjson/3.12.0/orjson-3.12.0-cp312-cp312-musllinux_1_2_x86_64.whl", hash = "sha256:532ff8cd4bd59a327a953a7dcde922c7fc25b85e29721bb8633265430d3a3873", size = 127693, upload-time = "2026-08-14T16:12:38.627Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/orjson/3.12.0/orjson-3.12.0-cp312-cp312-win32.whl", hash = "sha256:a6cf4b18e7de173f209f2084ffbd736dd72389a396326ee80a7022168be232e5", size = 128000, upload-time = "2026-08-14T16:12:39.954Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/orjson/3.12.0/orjson-3.12.0-cp312-cp312-win_amd64.whl", hash = "sha256:010811c1b69773450a01cef97727a67b223242f350b77d4ca000e59a9ef2155a", size = 122106, upload-time = "2026-08-14T16:12:41.324Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/orjson/3.12.0/orjson-3.12.0-cp312-cp312-win_arm64.whl", hash = "sha256:ad29eece0c601737f2a60edc2752a84e7a0785df3efb62e3012834700a5afe0d", size = 126869, upload-time = "2026-08-14T16:12:42.651Z" }, ] [[package]] name = "packaging" version = "26.3" -source = { registry = "https://pypi.org/simple" } -sdist = { url = "https://files.pythonhosted.org/packages/7d/fa/3944b40b07da9ce895c0e6303a5ab7d53da063554f534556b134a54d6093/packaging-26.3.tar.gz", hash = "sha256:94edc256424af38762eb31306eed28beb9f0efc50a8837492c9d6fd6004aed79", size = 313412, upload-time = "2026-08-04T18:15:28.737Z" } +source = { registry = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/" } +sdist = { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/packaging/26.3/packaging-26.3.tar.gz", hash = "sha256:94edc256424af38762eb31306eed28beb9f0efc50a8837492c9d6fd6004aed79", size = 313412, upload-time = "2026-08-04T18:15:28.737Z" } wheels = [ - { url = "https://files.pythonhosted.org/packages/63/34/ba1c580383c9eada3711951fef0795c80b829a078d72188184bcab9dd527/packaging-26.3-py3-none-any.whl", hash = "sha256:d7193f7c8e4e93f444fde0262bf90af30e16fa0ad0ad44cb553c87339b23cd1c", size = 129956, upload-time = "2026-08-04T18:15:27.159Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/packaging/26.3/packaging-26.3-py3-none-any.whl", hash = "sha256:d7193f7c8e4e93f444fde0262bf90af30e16fa0ad0ad44cb553c87339b23cd1c", size = 129956, upload-time = "2026-08-04T18:15:27.159Z" }, ] [[package]] name = "pluggy" version = "1.6.0" -source = { registry = "https://pypi.org/simple" } -sdist = { url = "https://files.pythonhosted.org/packages/f9/e2/3e91f31a7d2b083fe6ef3fa267035b518369d9511ffab804f839851d2779/pluggy-1.6.0.tar.gz", hash = "sha256:7dcc130b76258d33b90f61b658791dede3486c3e6bfb003ee5c9bfb396dd22f3", size = 69412, upload-time = "2025-05-15T12:30:07.975Z" } +source = { registry = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/" } +sdist = { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pluggy/1.6.0/pluggy-1.6.0.tar.gz", hash = "sha256:7dcc130b76258d33b90f61b658791dede3486c3e6bfb003ee5c9bfb396dd22f3", size = 69412, upload-time = "2025-05-15T12:30:07.975Z" } wheels = [ - { url = "https://files.pythonhosted.org/packages/54/20/4d324d65cc6d9205fabedc306948156824eb9f0ee1633355a8f7ec5c66bf/pluggy-1.6.0-py3-none-any.whl", hash = "sha256:e920276dd6813095e9377c0bc5566d94c932c33b27a3e3945d8389c374dd4746", size = 20538, upload-time = "2025-05-15T12:30:06.134Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pluggy/1.6.0/pluggy-1.6.0-py3-none-any.whl", hash = "sha256:e920276dd6813095e9377c0bc5566d94c932c33b27a3e3945d8389c374dd4746", size = 20538, upload-time = "2025-05-15T12:30:06.134Z" }, ] [[package]] name = "propcache" version = "0.5.2" -source = { registry = "https://pypi.org/simple" } -sdist = { url = "https://files.pythonhosted.org/packages/ec/44/c87281c333769159c50594f22610f77398a47ccbfbbf23074e744e86f87c/propcache-0.5.2.tar.gz", hash = "sha256:01c4fc7480cd0598bb4b57022df55b9ca296da7fc5a8760bd8451a7e63a7d427", size = 50208, upload-time = "2026-05-08T21:02:12.199Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/e7/f1/8a8cc1c2c7e7934ab77e0163414f736fadbc0f5e8dd9673b952355ac175b/propcache-0.5.2-cp311-cp311-macosx_10_9_universal2.whl", hash = "sha256:74b70780220e2dd89175ca24b81b68b67c83db499ae611e7f2313cb329801c78", size = 90744, upload-time = "2026-05-08T20:59:45.799Z" }, - { url = "https://files.pythonhosted.org/packages/c2/f4/651b1225e976bd1a2ba5cfba0c29d096581c2636b437e3a9a7ab6276270a/propcache-0.5.2-cp311-cp311-macosx_10_9_x86_64.whl", hash = "sha256:a4840ab0ae0216d952f4b53dc6d0b992bfc2bedbfe360bdd9b548bc184c08959", size = 52033, upload-time = "2026-05-08T20:59:47.408Z" }, - { url = "https://files.pythonhosted.org/packages/15/a8/8ede85d6aa1f79fc7dc2f8fd2c8d65920b8272c3892903c8a1affde48cfb/propcache-0.5.2-cp311-cp311-macosx_11_0_arm64.whl", hash = "sha256:c6844ba6364fb12f403928a82cfd295ab103a2b315c77c747b2dbe4a41894ea7", size = 52754, upload-time = "2026-05-08T20:59:49.202Z" }, - { url = "https://files.pythonhosted.org/packages/7d/fe/b3551b41bbc2f5b5bb088fc6920567cd43101253e68fbaa261339eb96fe1/propcache-0.5.2-cp311-cp311-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:2293949b855ce597f2826452d17c2d545fb5622379c4ea6fdf525e9b8e8a2511", size = 57573, upload-time = "2026-05-08T20:59:50.778Z" }, - { url = "https://files.pythonhosted.org/packages/83/27/ab851ebd1b7172e3e161f5f8d39e315d54a91bea246f01f4d872d3376aef/propcache-0.5.2-cp311-cp311-manylinux2014_ppc64le.manylinux_2_17_ppc64le.manylinux_2_28_ppc64le.whl", hash = "sha256:0fd59b5af35f74da48d905dcbad55449ba13be91823cb05a9bd590bbf5b61660", size = 60645, upload-time = "2026-05-08T20:59:52.227Z" }, - { url = "https://files.pythonhosted.org/packages/95/7d/466b3d18022e9897cbda9c735c493c5bd747d7a4c6f5ea1480b4cec434b6/propcache-0.5.2-cp311-cp311-manylinux2014_s390x.manylinux_2_17_s390x.manylinux_2_28_s390x.whl", hash = "sha256:29f9309a2e42b0d273be006fdb4be2d6c39a47f6f57d8fb1cf9f81481df81b66", size = 61563, upload-time = "2026-05-08T20:59:53.866Z" }, - { url = "https://files.pythonhosted.org/packages/27/1b/16ab7f2cf2041da2f60d156ba64c2484eadf9168075b4ff43c3ef60045af/propcache-0.5.2-cp311-cp311-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:5aaa2b923c1944ac8febd6609cb373540a5563e7cbcb0fd770f75dace2eb817b", size = 58888, upload-time = "2026-05-08T20:59:55.457Z" }, - { url = "https://files.pythonhosted.org/packages/0a/67/bb777ffd907633563bf35fd859c4ce97b0512c32f4633cf5d1eb7c33512b/propcache-0.5.2-cp311-cp311-manylinux_2_31_riscv64.manylinux_2_39_riscv64.whl", hash = "sha256:66ea454f095ddf5b6b14f56c064c0941c4788be11e18d2464cf643bf7203ff67", size = 59253, upload-time = "2026-05-08T20:59:57.075Z" }, - { url = "https://files.pythonhosted.org/packages/b9/42/64f8d90b73fd9cdc1499b48057ff6d9cd2a98a25734c9bb62ecf07e87061/propcache-0.5.2-cp311-cp311-musllinux_1_2_aarch64.whl", hash = "sha256:95f1e3f4760d404b13c9976c0229b2b49a3c8e2c62a9ce92efdd2b11ada75e3f", size = 57558, upload-time = "2026-05-08T20:59:58.602Z" }, - { url = "https://files.pythonhosted.org/packages/eb/02/dba5bc03c9041f2092ea55a449caf5dfe68352c6654511b29ba0654ddb69/propcache-0.5.2-cp311-cp311-musllinux_1_2_armv7l.whl", hash = "sha256:85341b12b9d55bad0bded24cac341bb34289469e03a11f3f583ea1cc1db0326c", size = 55007, upload-time = "2026-05-08T20:59:59.837Z" }, - { url = "https://files.pythonhosted.org/packages/14/c0/43f649c7aa2a77a3b100d84e9dea3a483120ecb608bfe36ce49eaff517fe/propcache-0.5.2-cp311-cp311-musllinux_1_2_ppc64le.whl", hash = "sha256:26a4dca084132874e639895c3135dfad5eb20bae209f62d1aeb31b03e601c3c0", size = 60355, upload-time = "2026-05-08T21:00:01.144Z" }, - { url = "https://files.pythonhosted.org/packages/83/c0/435dafd27f1cb4a495381dae60e25883ccfe4020bb72818e8184c1678092/propcache-0.5.2-cp311-cp311-musllinux_1_2_riscv64.whl", hash = "sha256:3b199b9b2b3d6a7edf3183ba8a9a137a22b97f7df525feb5ae1eccf026d2a9c6", size = 59057, upload-time = "2026-05-08T21:00:02.401Z" }, - { url = "https://files.pythonhosted.org/packages/53/ae/6e292df9135d659944e96cb3389258e4a663e5b2b5f6c217ef0ddc8d2f73/propcache-0.5.2-cp311-cp311-musllinux_1_2_s390x.whl", hash = "sha256:e59bc9e66329185b93dab73f210f1a37f81cb40f321501db8017c9aea15dba27", size = 61938, upload-time = "2026-05-08T21:00:03.638Z" }, - { url = "https://files.pythonhosted.org/packages/0b/42/314ebc50d8159055411fd6b0bda322ff510e4b1f7d2e4927940ad0f6af20/propcache-0.5.2-cp311-cp311-musllinux_1_2_x86_64.whl", hash = "sha256:552ffadf6ad409844bc5919c42a0a83d88314cedddaea0e41e80a8b8fffe881f", size = 59731, upload-time = "2026-05-08T21:00:04.881Z" }, - { url = "https://files.pythonhosted.org/packages/b8/9b/2da6dee38871c3c8772fabc2758325a5c9077d6d18c597737dc04dd884cd/propcache-0.5.2-cp311-cp311-win32.whl", hash = "sha256:cd416c1de191973c52ff1a12a57446bfc7642797b282d7caf2162d7d1b8aa9a0", size = 38966, upload-time = "2026-05-08T21:00:06.511Z" }, - { url = "https://files.pythonhosted.org/packages/42/4e/f17363fb58c0afe05b067361cb6d86ed2d29de6506779a27547c4d183075/propcache-0.5.2-cp311-cp311-win_amd64.whl", hash = "sha256:44e488ef40dbb452700b2b1f8188934121f6648f52c295055662d2191959ff82", size = 42135, upload-time = "2026-05-08T21:00:08.088Z" }, - { url = "https://files.pythonhosted.org/packages/c6/eb/6af6685077d22e8b33358d3c548e3282706a0b3cd85044ffba4e5dd08e3b/propcache-0.5.2-cp311-cp311-win_arm64.whl", hash = "sha256:54adaa85a22078d1e306304a40984dc5be99d599bf3dc0a24dc98f7daeab89ab", size = 38381, upload-time = "2026-05-08T21:00:09.692Z" }, - { url = "https://files.pythonhosted.org/packages/4a/cb/e27bc2b2737a0bb49962b275efa051e8f1c35a936df7d5139b6b658b7dc9/propcache-0.5.2-cp312-cp312-macosx_10_13_universal2.whl", hash = "sha256:806719138ecd720339a12410fb9614ac9b2b2d3a5fdf8235d56981c36f4039ba", size = 95887, upload-time = "2026-05-08T21:00:11.277Z" }, - { url = "https://files.pythonhosted.org/packages/e6/13/b8ae04c59392f8d11c6cd9fb4011d1dc7c86b81225c770280300e259ffe1/propcache-0.5.2-cp312-cp312-macosx_10_13_x86_64.whl", hash = "sha256:db2b80ea58eab4f86b2beec3cc8b39e8ff9276ac20e96b7cce43c8ae84cd6b5a", size = 54654, upload-time = "2026-05-08T21:00:12.604Z" }, - { url = "https://files.pythonhosted.org/packages/2c/7d/49777a3e20b55863d4794384a38acd460c04157b0a00f8602b0d508b8431/propcache-0.5.2-cp312-cp312-macosx_11_0_arm64.whl", hash = "sha256:e5cbfac9f61484f7e9f3597775500cd3ebe8274e9b050c38f9525c77c97520bf", size = 55190, upload-time = "2026-05-08T21:00:13.935Z" }, - { url = "https://files.pythonhosted.org/packages/44/c7/085d0cd63062e84044e3f05797749c3f8e3938ff3aeb0eb2f69d43fafc91/propcache-0.5.2-cp312-cp312-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:5dbc581d2814337da56222fab8dc5f161cd798a434e49bac27930aaef798e144", size = 59995, upload-time = "2026-05-08T21:00:15.526Z" }, - { url = "https://files.pythonhosted.org/packages/9c/42/32cf8e3009e92b2645cf1e944f701e8ea4e924dffde1ee26db860bcbf7e4/propcache-0.5.2-cp312-cp312-manylinux2014_ppc64le.manylinux_2_17_ppc64le.manylinux_2_28_ppc64le.whl", hash = "sha256:857187f381f88c8e2fa2fe56ab94879d011b883d5a2ee5a1b60a8cd2a06846d9", size = 63422, upload-time = "2026-05-08T21:00:16.824Z" }, - { url = "https://files.pythonhosted.org/packages/9e/1b/f112433f99fc979431b87a39ef169e3f8df070d99a72792c56d6937ac48b/propcache-0.5.2-cp312-cp312-manylinux2014_s390x.manylinux_2_17_s390x.manylinux_2_28_s390x.whl", hash = "sha256:178b4a2cdaac1818e2bf1c5a99b94383fa73ea5382e032a48dec07dc5668dc42", size = 64342, upload-time = "2026-05-08T21:00:18.362Z" }, - { url = "https://files.pythonhosted.org/packages/14/15/5574111ae50dd6e879456888c0eadd4c5a869959775854e18e18a6b345f3/propcache-0.5.2-cp312-cp312-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:6f328175a2cde1f0ff2c4ed8ce968b9dcfb55f3a7153f39e2957ed994da13476", size = 61639, upload-time = "2026-05-08T21:00:19.692Z" }, - { url = "https://files.pythonhosted.org/packages/cc/da/4d775080b1490c0ae604acda868bd71aabe3a89ed16f2aa4339eb8a283e7/propcache-0.5.2-cp312-cp312-manylinux_2_31_riscv64.manylinux_2_39_riscv64.whl", hash = "sha256:5671d09a36b06d0fd4a3da0fccbcae360e9b1570924171a15e9e0997f0249fba", size = 61588, upload-time = "2026-05-08T21:00:21.155Z" }, - { url = "https://files.pythonhosted.org/packages/04/ac/f076982cbe2195ee9cf32de5a1e46951d9fb399fc207f390562dd0fd8fb2/propcache-0.5.2-cp312-cp312-musllinux_1_2_aarch64.whl", hash = "sha256:80168e2ebe4d3ec6599d10ad8f520304ae1cad9b6c5a95372aef1b66b7bfb53a", size = 60029, upload-time = "2026-05-08T21:00:22.713Z" }, - { url = "https://files.pythonhosted.org/packages/70/60/189be62e0dd898dce3b331e1b8c7a543cd3a405ac0c81fe8ee8a9d5d77e1/propcache-0.5.2-cp312-cp312-musllinux_1_2_armv7l.whl", hash = "sha256:45f11346f884bc47444f6e6647131055844134c3175b629f84952e2b5cd62b64", size = 56774, upload-time = "2026-05-08T21:00:24.001Z" }, - { url = "https://files.pythonhosted.org/packages/ea/9e/93377b9c7939c1ffae98f878dee955efadfd638078bc86dbc21f9d52f651/propcache-0.5.2-cp312-cp312-musllinux_1_2_ppc64le.whl", hash = "sha256:8e778ebd44ef4f66ed60a0416b06b489687db264a9c0b3620362f26489492913", size = 63532, upload-time = "2026-05-08T21:00:25.545Z" }, - { url = "https://files.pythonhosted.org/packages/14/f9/590ef6cfb9b8028d516d287812ece32bb0bc5f11fbb9c8bf6b2e6313fec8/propcache-0.5.2-cp312-cp312-musllinux_1_2_riscv64.whl", hash = "sha256:c0cb9ed24c8964e172768d455a38254c2dd8a552905729ce006cad3d3dda59b1", size = 61592, upload-time = "2026-05-08T21:00:27.186Z" }, - { url = "https://files.pythonhosted.org/packages/b4/5e/70958b3034c297a630bba2f17ca7abc2d5f39a803ad7e370ab79d1ecd022/propcache-0.5.2-cp312-cp312-musllinux_1_2_s390x.whl", hash = "sha256:1d1ad32d9d4355e2be65574fd0bfd3677e7066b009cd5b9b2dee8aa6a6393b33", size = 64788, upload-time = "2026-05-08T21:00:28.8Z" }, - { url = "https://files.pythonhosted.org/packages/12/fd/77fe5936d8c3086ca9048f7f415f122ed82e53884a9ec193646b42deef06/propcache-0.5.2-cp312-cp312-musllinux_1_2_x86_64.whl", hash = "sha256:c80f4ba3e8f00189165999a742ee526ebeccedf6c3f7beb0c7df821e9772435a", size = 62514, upload-time = "2026-05-08T21:00:30.098Z" }, - { url = "https://files.pythonhosted.org/packages/cf/74/66bd798b5b3be70aa1b391f5cc9d6a0a5532d7fd3b19ec0b213e72e6ad9d/propcache-0.5.2-cp312-cp312-win32.whl", hash = "sha256:8c7972d8f193740d9175f0998ab38717e6cd322d5935c5b0fef8c0d323fd9031", size = 39018, upload-time = "2026-05-08T21:00:31.622Z" }, - { url = "https://files.pythonhosted.org/packages/61/7c/5c0d34aa3024694d6dcb9271cdbdd08c4e47c1c0ad95ec7e7bc74cdea145/propcache-0.5.2-cp312-cp312-win_amd64.whl", hash = "sha256:d9ee8826a7d47863a08ac44e1a5f611a462eefc3a194b492da242128bec75b42", size = 42322, upload-time = "2026-05-08T21:00:32.918Z" }, - { url = "https://files.pythonhosted.org/packages/4d/91/875812f1a3feb20ceba818ef39fbe4d92f1081e04ac815c822496d0d038b/propcache-0.5.2-cp312-cp312-win_arm64.whl", hash = "sha256:2800a4a8ead6b28cccd1ec54b59346f0def7922ee1c7598e8499c733cfbb7c84", size = 38172, upload-time = "2026-05-08T21:00:35.124Z" }, - { url = "https://files.pythonhosted.org/packages/3a/ed/1cdcab6ba3d6ab7feca11fc14f0eeea80755bb53ef4e892079f31b10a25f/propcache-0.5.2-py3-none-any.whl", hash = "sha256:be1ddfcbb376e3de5d2e2db1d58d6d67463e6b4f9f040c000de8e300295465fe", size = 14036, upload-time = "2026-05-08T21:02:10.673Z" }, +source = { registry = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/" } +sdist = { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/propcache/0.5.2/propcache-0.5.2.tar.gz", hash = "sha256:01c4fc7480cd0598bb4b57022df55b9ca296da7fc5a8760bd8451a7e63a7d427", size = 50208, upload-time = "2026-05-08T21:02:12.199Z" } +wheels = [ + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/propcache/0.5.2/propcache-0.5.2-cp311-cp311-macosx_10_9_universal2.whl", hash = "sha256:74b70780220e2dd89175ca24b81b68b67c83db499ae611e7f2313cb329801c78", size = 90744, upload-time = "2026-05-08T20:59:45.799Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/propcache/0.5.2/propcache-0.5.2-cp311-cp311-macosx_10_9_x86_64.whl", hash = "sha256:a4840ab0ae0216d952f4b53dc6d0b992bfc2bedbfe360bdd9b548bc184c08959", size = 52033, upload-time = "2026-05-08T20:59:47.408Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/propcache/0.5.2/propcache-0.5.2-cp311-cp311-macosx_11_0_arm64.whl", hash = "sha256:c6844ba6364fb12f403928a82cfd295ab103a2b315c77c747b2dbe4a41894ea7", size = 52754, upload-time = "2026-05-08T20:59:49.202Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/propcache/0.5.2/propcache-0.5.2-cp311-cp311-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:2293949b855ce597f2826452d17c2d545fb5622379c4ea6fdf525e9b8e8a2511", size = 57573, upload-time = "2026-05-08T20:59:50.778Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/propcache/0.5.2/propcache-0.5.2-cp311-cp311-manylinux2014_ppc64le.manylinux_2_17_ppc64le.manylinux_2_28_ppc64le.whl", hash = "sha256:0fd59b5af35f74da48d905dcbad55449ba13be91823cb05a9bd590bbf5b61660", size = 60645, upload-time = "2026-05-08T20:59:52.227Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/propcache/0.5.2/propcache-0.5.2-cp311-cp311-manylinux2014_s390x.manylinux_2_17_s390x.manylinux_2_28_s390x.whl", hash = "sha256:29f9309a2e42b0d273be006fdb4be2d6c39a47f6f57d8fb1cf9f81481df81b66", size = 61563, upload-time = "2026-05-08T20:59:53.866Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/propcache/0.5.2/propcache-0.5.2-cp311-cp311-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:5aaa2b923c1944ac8febd6609cb373540a5563e7cbcb0fd770f75dace2eb817b", size = 58888, upload-time = "2026-05-08T20:59:55.457Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/propcache/0.5.2/propcache-0.5.2-cp311-cp311-manylinux_2_31_riscv64.manylinux_2_39_riscv64.whl", hash = "sha256:66ea454f095ddf5b6b14f56c064c0941c4788be11e18d2464cf643bf7203ff67", size = 59253, upload-time = "2026-05-08T20:59:57.075Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/propcache/0.5.2/propcache-0.5.2-cp311-cp311-musllinux_1_2_aarch64.whl", hash = "sha256:95f1e3f4760d404b13c9976c0229b2b49a3c8e2c62a9ce92efdd2b11ada75e3f", size = 57558, upload-time = "2026-05-08T20:59:58.602Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/propcache/0.5.2/propcache-0.5.2-cp311-cp311-musllinux_1_2_armv7l.whl", hash = "sha256:85341b12b9d55bad0bded24cac341bb34289469e03a11f3f583ea1cc1db0326c", size = 55007, upload-time = "2026-05-08T20:59:59.837Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/propcache/0.5.2/propcache-0.5.2-cp311-cp311-musllinux_1_2_ppc64le.whl", hash = "sha256:26a4dca084132874e639895c3135dfad5eb20bae209f62d1aeb31b03e601c3c0", size = 60355, upload-time = "2026-05-08T21:00:01.144Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/propcache/0.5.2/propcache-0.5.2-cp311-cp311-musllinux_1_2_riscv64.whl", hash = "sha256:3b199b9b2b3d6a7edf3183ba8a9a137a22b97f7df525feb5ae1eccf026d2a9c6", size = 59057, upload-time = "2026-05-08T21:00:02.401Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/propcache/0.5.2/propcache-0.5.2-cp311-cp311-musllinux_1_2_s390x.whl", hash = "sha256:e59bc9e66329185b93dab73f210f1a37f81cb40f321501db8017c9aea15dba27", size = 61938, upload-time = "2026-05-08T21:00:03.638Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/propcache/0.5.2/propcache-0.5.2-cp311-cp311-musllinux_1_2_x86_64.whl", hash = "sha256:552ffadf6ad409844bc5919c42a0a83d88314cedddaea0e41e80a8b8fffe881f", size = 59731, upload-time = "2026-05-08T21:00:04.881Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/propcache/0.5.2/propcache-0.5.2-cp311-cp311-win32.whl", hash = "sha256:cd416c1de191973c52ff1a12a57446bfc7642797b282d7caf2162d7d1b8aa9a0", size = 38966, upload-time = "2026-05-08T21:00:06.511Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/propcache/0.5.2/propcache-0.5.2-cp311-cp311-win_amd64.whl", hash = "sha256:44e488ef40dbb452700b2b1f8188934121f6648f52c295055662d2191959ff82", size = 42135, upload-time = "2026-05-08T21:00:08.088Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/propcache/0.5.2/propcache-0.5.2-cp311-cp311-win_arm64.whl", hash = "sha256:54adaa85a22078d1e306304a40984dc5be99d599bf3dc0a24dc98f7daeab89ab", size = 38381, upload-time = "2026-05-08T21:00:09.692Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/propcache/0.5.2/propcache-0.5.2-cp312-cp312-macosx_10_13_universal2.whl", hash = "sha256:806719138ecd720339a12410fb9614ac9b2b2d3a5fdf8235d56981c36f4039ba", size = 95887, upload-time = "2026-05-08T21:00:11.277Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/propcache/0.5.2/propcache-0.5.2-cp312-cp312-macosx_10_13_x86_64.whl", hash = "sha256:db2b80ea58eab4f86b2beec3cc8b39e8ff9276ac20e96b7cce43c8ae84cd6b5a", size = 54654, upload-time = "2026-05-08T21:00:12.604Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/propcache/0.5.2/propcache-0.5.2-cp312-cp312-macosx_11_0_arm64.whl", hash = "sha256:e5cbfac9f61484f7e9f3597775500cd3ebe8274e9b050c38f9525c77c97520bf", size = 55190, upload-time = "2026-05-08T21:00:13.935Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/propcache/0.5.2/propcache-0.5.2-cp312-cp312-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:5dbc581d2814337da56222fab8dc5f161cd798a434e49bac27930aaef798e144", size = 59995, upload-time = "2026-05-08T21:00:15.526Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/propcache/0.5.2/propcache-0.5.2-cp312-cp312-manylinux2014_ppc64le.manylinux_2_17_ppc64le.manylinux_2_28_ppc64le.whl", hash = "sha256:857187f381f88c8e2fa2fe56ab94879d011b883d5a2ee5a1b60a8cd2a06846d9", size = 63422, upload-time = "2026-05-08T21:00:16.824Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/propcache/0.5.2/propcache-0.5.2-cp312-cp312-manylinux2014_s390x.manylinux_2_17_s390x.manylinux_2_28_s390x.whl", hash = "sha256:178b4a2cdaac1818e2bf1c5a99b94383fa73ea5382e032a48dec07dc5668dc42", size = 64342, upload-time = "2026-05-08T21:00:18.362Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/propcache/0.5.2/propcache-0.5.2-cp312-cp312-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:6f328175a2cde1f0ff2c4ed8ce968b9dcfb55f3a7153f39e2957ed994da13476", size = 61639, upload-time = "2026-05-08T21:00:19.692Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/propcache/0.5.2/propcache-0.5.2-cp312-cp312-manylinux_2_31_riscv64.manylinux_2_39_riscv64.whl", hash = "sha256:5671d09a36b06d0fd4a3da0fccbcae360e9b1570924171a15e9e0997f0249fba", size = 61588, upload-time = "2026-05-08T21:00:21.155Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/propcache/0.5.2/propcache-0.5.2-cp312-cp312-musllinux_1_2_aarch64.whl", hash = "sha256:80168e2ebe4d3ec6599d10ad8f520304ae1cad9b6c5a95372aef1b66b7bfb53a", size = 60029, upload-time = "2026-05-08T21:00:22.713Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/propcache/0.5.2/propcache-0.5.2-cp312-cp312-musllinux_1_2_armv7l.whl", hash = "sha256:45f11346f884bc47444f6e6647131055844134c3175b629f84952e2b5cd62b64", size = 56774, upload-time = "2026-05-08T21:00:24.001Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/propcache/0.5.2/propcache-0.5.2-cp312-cp312-musllinux_1_2_ppc64le.whl", hash = "sha256:8e778ebd44ef4f66ed60a0416b06b489687db264a9c0b3620362f26489492913", size = 63532, upload-time = "2026-05-08T21:00:25.545Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/propcache/0.5.2/propcache-0.5.2-cp312-cp312-musllinux_1_2_riscv64.whl", hash = "sha256:c0cb9ed24c8964e172768d455a38254c2dd8a552905729ce006cad3d3dda59b1", size = 61592, upload-time = "2026-05-08T21:00:27.186Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/propcache/0.5.2/propcache-0.5.2-cp312-cp312-musllinux_1_2_s390x.whl", hash = "sha256:1d1ad32d9d4355e2be65574fd0bfd3677e7066b009cd5b9b2dee8aa6a6393b33", size = 64788, upload-time = "2026-05-08T21:00:28.8Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/propcache/0.5.2/propcache-0.5.2-cp312-cp312-musllinux_1_2_x86_64.whl", hash = "sha256:c80f4ba3e8f00189165999a742ee526ebeccedf6c3f7beb0c7df821e9772435a", size = 62514, upload-time = "2026-05-08T21:00:30.098Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/propcache/0.5.2/propcache-0.5.2-cp312-cp312-win32.whl", hash = "sha256:8c7972d8f193740d9175f0998ab38717e6cd322d5935c5b0fef8c0d323fd9031", size = 39018, upload-time = "2026-05-08T21:00:31.622Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/propcache/0.5.2/propcache-0.5.2-cp312-cp312-win_amd64.whl", hash = "sha256:d9ee8826a7d47863a08ac44e1a5f611a462eefc3a194b492da242128bec75b42", size = 42322, upload-time = "2026-05-08T21:00:32.918Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/propcache/0.5.2/propcache-0.5.2-cp312-cp312-win_arm64.whl", hash = "sha256:2800a4a8ead6b28cccd1ec54b59346f0def7922ee1c7598e8499c733cfbb7c84", size = 38172, upload-time = "2026-05-08T21:00:35.124Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/propcache/0.5.2/propcache-0.5.2-py3-none-any.whl", hash = "sha256:be1ddfcbb376e3de5d2e2db1d58d6d67463e6b4f9f040c000de8e300295465fe", size = 14036, upload-time = "2026-05-08T21:02:10.673Z" }, ] [[package]] name = "protobuf" version = "6.33.6" -source = { registry = "https://pypi.org/simple" } -sdist = { url = "https://files.pythonhosted.org/packages/66/70/e908e9c5e52ef7c3a6c7902c9dfbb34c7e29c25d2f81ade3856445fd5c94/protobuf-6.33.6.tar.gz", hash = "sha256:a6768d25248312c297558af96a9f9c929e8c4cee0659cb07e780731095f38135", size = 444531, upload-time = "2026-03-18T19:05:00.988Z" } +source = { registry = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/" } +sdist = { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/protobuf/6.33.6/protobuf-6.33.6.tar.gz", hash = "sha256:a6768d25248312c297558af96a9f9c929e8c4cee0659cb07e780731095f38135", size = 444531, upload-time = "2026-03-18T19:05:00.988Z" } wheels = [ - { url = "https://files.pythonhosted.org/packages/fc/9f/2f509339e89cfa6f6a4c4ff50438db9ca488dec341f7e454adad60150b00/protobuf-6.33.6-cp310-abi3-win32.whl", hash = "sha256:7d29d9b65f8afef196f8334e80d6bc1d5d4adedb449971fefd3723824e6e77d3", size = 425739, upload-time = "2026-03-18T19:04:48.373Z" }, - { url = "https://files.pythonhosted.org/packages/76/5d/683efcd4798e0030c1bab27374fd13a89f7c2515fb1f3123efdfaa5eab57/protobuf-6.33.6-cp310-abi3-win_amd64.whl", hash = "sha256:0cd27b587afca21b7cfa59a74dcbd48a50f0a6400cfb59391340ad729d91d326", size = 437089, upload-time = "2026-03-18T19:04:50.381Z" }, - { url = "https://files.pythonhosted.org/packages/5c/01/a3c3ed5cd186f39e7880f8303cc51385a198a81469d53d0fdecf1f64d929/protobuf-6.33.6-cp39-abi3-macosx_10_9_universal2.whl", hash = "sha256:9720e6961b251bde64edfdab7d500725a2af5280f3f4c87e57c0208376aa8c3a", size = 427737, upload-time = "2026-03-18T19:04:51.866Z" }, - { url = "https://files.pythonhosted.org/packages/ee/90/b3c01fdec7d2f627b3a6884243ba328c1217ed2d978def5c12dc50d328a3/protobuf-6.33.6-cp39-abi3-manylinux2014_aarch64.whl", hash = "sha256:e2afbae9b8e1825e3529f88d514754e094278bb95eadc0e199751cdd9a2e82a2", size = 324610, upload-time = "2026-03-18T19:04:53.096Z" }, - { url = "https://files.pythonhosted.org/packages/9b/ca/25afc144934014700c52e05103c2421997482d561f3101ff352e1292fb81/protobuf-6.33.6-cp39-abi3-manylinux2014_s390x.whl", hash = "sha256:c96c37eec15086b79762ed265d59ab204dabc53056e3443e702d2681f4b39ce3", size = 339381, upload-time = "2026-03-18T19:04:54.616Z" }, - { url = "https://files.pythonhosted.org/packages/16/92/d1e32e3e0d894fe00b15ce28ad4944ab692713f2e7f0a99787405e43533a/protobuf-6.33.6-cp39-abi3-manylinux2014_x86_64.whl", hash = "sha256:e9db7e292e0ab79dd108d7f1a94fe31601ce1ee3f7b79e0692043423020b0593", size = 323436, upload-time = "2026-03-18T19:04:55.768Z" }, - { url = "https://files.pythonhosted.org/packages/c4/72/02445137af02769918a93807b2b7890047c32bfb9f90371cbc12688819eb/protobuf-6.33.6-py3-none-any.whl", hash = "sha256:77179e006c476e69bf8e8ce866640091ec42e1beb80b213c3900006ecfba6901", size = 170656, upload-time = "2026-03-18T19:04:59.826Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/protobuf/6.33.6/protobuf-6.33.6-cp310-abi3-win32.whl", hash = "sha256:7d29d9b65f8afef196f8334e80d6bc1d5d4adedb449971fefd3723824e6e77d3", size = 425739, upload-time = "2026-03-18T19:04:48.373Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/protobuf/6.33.6/protobuf-6.33.6-cp310-abi3-win_amd64.whl", hash = "sha256:0cd27b587afca21b7cfa59a74dcbd48a50f0a6400cfb59391340ad729d91d326", size = 437089, upload-time = "2026-03-18T19:04:50.381Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/protobuf/6.33.6/protobuf-6.33.6-cp39-abi3-macosx_10_9_universal2.whl", hash = "sha256:9720e6961b251bde64edfdab7d500725a2af5280f3f4c87e57c0208376aa8c3a", size = 427737, upload-time = "2026-03-18T19:04:51.866Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/protobuf/6.33.6/protobuf-6.33.6-cp39-abi3-manylinux2014_aarch64.whl", hash = "sha256:e2afbae9b8e1825e3529f88d514754e094278bb95eadc0e199751cdd9a2e82a2", size = 324610, upload-time = "2026-03-18T19:04:53.096Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/protobuf/6.33.6/protobuf-6.33.6-cp39-abi3-manylinux2014_s390x.whl", hash = "sha256:c96c37eec15086b79762ed265d59ab204dabc53056e3443e702d2681f4b39ce3", size = 339381, upload-time = "2026-03-18T19:04:54.616Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/protobuf/6.33.6/protobuf-6.33.6-cp39-abi3-manylinux2014_x86_64.whl", hash = "sha256:e9db7e292e0ab79dd108d7f1a94fe31601ce1ee3f7b79e0692043423020b0593", size = 323436, upload-time = "2026-03-18T19:04:55.768Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/protobuf/6.33.6/protobuf-6.33.6-py3-none-any.whl", hash = "sha256:77179e006c476e69bf8e8ce866640091ec42e1beb80b213c3900006ecfba6901", size = 170656, upload-time = "2026-03-18T19:04:59.826Z" }, ] [[package]] name = "pydantic" version = "2.13.4" -source = { registry = "https://pypi.org/simple" } +source = { registry = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/" } dependencies = [ { name = "annotated-types" }, { name = "pydantic-core" }, { name = "typing-extensions" }, { name = "typing-inspection" }, ] -sdist = { url = "https://files.pythonhosted.org/packages/18/a5/b60d21ac674192f8ab0ba4e9fd860690f9b4a6e51ca5df118733b487d8d6/pydantic-2.13.4.tar.gz", hash = "sha256:c40756b57adaa8b1efeeced5c196f3f3b7c435f90e84ea7f443901bec8099ef6", size = 844775, upload-time = "2026-05-06T13:43:05.343Z" } +sdist = { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pydantic/2.13.4/pydantic-2.13.4.tar.gz", hash = "sha256:c40756b57adaa8b1efeeced5c196f3f3b7c435f90e84ea7f443901bec8099ef6", size = 844775, upload-time = "2026-05-06T13:43:05.343Z" } wheels = [ - { url = "https://files.pythonhosted.org/packages/fd/7b/122376b1fd3c62c1ed9dc80c931ace4844b3c55407b6fb2d199377c9736f/pydantic-2.13.4-py3-none-any.whl", hash = "sha256:45a282cde31d808236fd7ea9d919b128653c8b38b393d1c4ab335c62924d9aba", size = 472262, upload-time = "2026-05-06T13:43:02.641Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pydantic/2.13.4/pydantic-2.13.4-py3-none-any.whl", hash = "sha256:45a282cde31d808236fd7ea9d919b128653c8b38b393d1c4ab335c62924d9aba", size = 472262, upload-time = "2026-05-06T13:43:02.641Z" }, ] [[package]] name = "pydantic-core" version = "2.46.4" -source = { registry = "https://pypi.org/simple" } +source = { registry = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/" } dependencies = [ { name = "typing-extensions" }, ] -sdist = { url = "https://files.pythonhosted.org/packages/9d/56/921726b776ace8d8f5db44c4ef961006580d91dc52b803c489fafd1aa249/pydantic_core-2.46.4.tar.gz", hash = "sha256:62f875393d7f270851f20523dd2e29f082bcc82292d66db2b64ea71f64b6e1c1", size = 471464, upload-time = "2026-05-06T13:37:06.98Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/5c/fa/6d7708d2cfc1a832acb6aeb0cd16e801902df8a0f583bb3b4b527fde022e/pydantic_core-2.46.4-cp311-cp311-macosx_10_12_x86_64.whl", hash = "sha256:0e96592440881c74a213e5ad528e2b24d3d4f940de2766bed9010ab1d9e51594", size = 2111872, upload-time = "2026-05-06T13:40:27.596Z" }, - { url = "https://files.pythonhosted.org/packages/ae/6f/aa064a3e74b5745afbdf250594f38e7ead05e2d651bcb35994b9417a0d4d/pydantic_core-2.46.4-cp311-cp311-macosx_11_0_arm64.whl", hash = "sha256:e0d65b8c354be7fb5f720c3caa8bc940bc2d20ce749c8e06135f07f8ed95dd7c", size = 1948255, upload-time = "2026-05-06T13:39:12.574Z" }, - { url = "https://files.pythonhosted.org/packages/43/3a/41114a9f7569b84b4d84e7a018c57c56347dac30c0d4a872946ec4e36c46/pydantic_core-2.46.4-cp311-cp311-manylinux_2_17_aarch64.manylinux2014_aarch64.whl", hash = "sha256:7bfb192b3f4b9e8a89b6277b6ce787564f62cfd272055f6e685726b111dc7826", size = 1972827, upload-time = "2026-05-06T13:38:19.841Z" }, - { url = "https://files.pythonhosted.org/packages/ef/25/1ab42e8048fe551934d9884e8d64daa7e990ad386f310a15981aeb6a5b08/pydantic_core-2.46.4-cp311-cp311-manylinux_2_17_armv7l.manylinux2014_armv7l.whl", hash = "sha256:9037063db01f09b09e237c282b6792bd4da634b5402c4e7f0c61effed7701a04", size = 2041051, upload-time = "2026-05-06T13:38:10.447Z" }, - { url = "https://files.pythonhosted.org/packages/94/c2/1a934597ddf08da410385b3b7aae91956a5a76c635effef456074fad7e88/pydantic_core-2.46.4-cp311-cp311-manylinux_2_17_ppc64le.manylinux2014_ppc64le.whl", hash = "sha256:fc010ab034c8c7452522748bf937df58020d256ccae0874463d1f4d01758af8e", size = 2221314, upload-time = "2026-05-06T13:40:13.089Z" }, - { url = "https://files.pythonhosted.org/packages/02/6d/9e8ad178c9c4df27ad3c8f25d1fe2a7ab0d2ba0559fad4aee5d3d1f16771/pydantic_core-2.46.4-cp311-cp311-manylinux_2_17_s390x.manylinux2014_s390x.whl", hash = "sha256:8c5dac79fa1614d1e06ca695109c6105923bd9c7d1d6c918d4e637b7e6b32fd3", size = 2285146, upload-time = "2026-05-06T13:38:59.224Z" }, - { url = "https://files.pythonhosted.org/packages/80/50/540cd3aeefc041beb111125c4bff779831a2111fc6b15a9138cda277d32c/pydantic_core-2.46.4-cp311-cp311-manylinux_2_17_x86_64.manylinux2014_x86_64.whl", hash = "sha256:f9fa868638bf362d3d138ea55829cefb3d5f4b0d7f142234382a15e2485dbec4", size = 2089685, upload-time = "2026-05-06T13:38:17.762Z" }, - { url = "https://files.pythonhosted.org/packages/6b/a4/b440ad35f05f6a38f89fa0f149accb3f0e02be94ca5e15f3c449a61b4bc9/pydantic_core-2.46.4-cp311-cp311-manylinux_2_31_riscv64.whl", hash = "sha256:17299feefe090f2caa5b8e37222bb5f663e4935a8bfa6931d4102e5df1a9f398", size = 2115420, upload-time = "2026-05-06T13:37:58.195Z" }, - { url = "https://files.pythonhosted.org/packages/99/61/de4f55db8dfd57bfdfa9a12ec90fe1b57c4f41062f7ca86f08586b3e0ac0/pydantic_core-2.46.4-cp311-cp311-manylinux_2_5_i686.manylinux1_i686.whl", hash = "sha256:4c63ebc82684aa89d9a3bcbd13d515b3be44250dc68dd3bd81526c1cb31286c3", size = 2165122, upload-time = "2026-05-06T13:37:01.167Z" }, - { url = "https://files.pythonhosted.org/packages/f7/52/7c529d7bdb2d1068bd52f51fe32572c8301f9a4febf1948f10639f1436f5/pydantic_core-2.46.4-cp311-cp311-musllinux_1_1_aarch64.whl", hash = "sha256:aaa2a54443eff1950ba5ddc6b6ccda0d9c84a364276a62f969bdf2a390650848", size = 2182573, upload-time = "2026-05-06T13:38:45.04Z" }, - { url = "https://files.pythonhosted.org/packages/37/b3/7c40325848ba78247f2812dcf9c7274e38cd801820ca6dd9fe63bcfb0eb4/pydantic_core-2.46.4-cp311-cp311-musllinux_1_1_armv7l.whl", hash = "sha256:18e5ceec2ab67e6d5f1a9085e5a24c9c4e2ac4545730bfe668680bca05e555f3", size = 2317139, upload-time = "2026-05-06T13:37:15.539Z" }, - { url = "https://files.pythonhosted.org/packages/d9/37/f913f81a657c865b75da6c0dbed79876073c2a43b5bd9edbe8da785e4d49/pydantic_core-2.46.4-cp311-cp311-musllinux_1_1_x86_64.whl", hash = "sha256:a0f62d0a58f4e7da165457e995725421e0064f2255d8eccebc49f41bbc23b109", size = 2360433, upload-time = "2026-05-06T13:37:30.099Z" }, - { url = "https://files.pythonhosted.org/packages/c4/67/6acaa1be2567f9256b056d8477158cac7240813956ce86e49deae8e173b4/pydantic_core-2.46.4-cp311-cp311-win32.whl", hash = "sha256:041bde0a48fd37cf71cab1c9d56d3e8625a3793fef1f7dd232b3ff37e978ecda", size = 1985513, upload-time = "2026-05-06T13:38:15.669Z" }, - { url = "https://files.pythonhosted.org/packages/aa/e6/c505f83dfeda9a2e5c995cfd872949e4d05e12f7feb3dca72f633daefa94/pydantic_core-2.46.4-cp311-cp311-win_amd64.whl", hash = "sha256:6f2eeda33a839975441c86a4119e1383c50b47faf0cbb5176985565c6bb02c33", size = 2071114, upload-time = "2026-05-06T13:40:35.416Z" }, - { url = "https://files.pythonhosted.org/packages/0f/da/7a263a96d965d9d0df5e8de8a475f33495451117035b09acb110288c381f/pydantic_core-2.46.4-cp311-cp311-win_arm64.whl", hash = "sha256:14f4c5d6db102bd796a627bbb3a17b4cf4574b9ae861d8b7c9a9661c6dd3362d", size = 2044298, upload-time = "2026-05-06T13:38:29.754Z" }, - { url = "https://files.pythonhosted.org/packages/ce/8c/af022f0af448d7747c5154288d46b5f2bc5f17366eaa0e23e9aa04d59f3b/pydantic_core-2.46.4-cp312-cp312-macosx_10_12_x86_64.whl", hash = "sha256:3245406455a5d98187ec35530fd772b1d799b26667980872c8d4614991e2c4a2", size = 2106158, upload-time = "2026-05-06T13:38:57.215Z" }, - { url = "https://files.pythonhosted.org/packages/19/95/6195171e385007300f0f5574592e467c568becce2d937a0b6804f218bc49/pydantic_core-2.46.4-cp312-cp312-macosx_11_0_arm64.whl", hash = "sha256:962ccbab7b642487b1d8b7df90ef677e03134cf1fd8880bf698649b22a69371f", size = 1951724, upload-time = "2026-05-06T13:37:02.697Z" }, - { url = "https://files.pythonhosted.org/packages/8e/bc/f47d1ff9cbb1620e1b5b697eef06010035735f07820180e74178226b27b3/pydantic_core-2.46.4-cp312-cp312-manylinux_2_17_aarch64.manylinux2014_aarch64.whl", hash = "sha256:8233f2947cf85404441fd7e0085f53b10c93e0ee78611099b5c7237e36aacbf7", size = 1975742, upload-time = "2026-05-06T13:37:09.448Z" }, - { url = "https://files.pythonhosted.org/packages/5b/11/9b9a5b0306345664a2da6410877af6e8082481b5884b3ddd78d47c6013ce/pydantic_core-2.46.4-cp312-cp312-manylinux_2_17_armv7l.manylinux2014_armv7l.whl", hash = "sha256:3a233125ac121aa3ffba9a2b59edfc4a985a76092dc8279586ab4b71390875e7", size = 2052418, upload-time = "2026-05-06T13:37:38.234Z" }, - { url = "https://files.pythonhosted.org/packages/f1/b7/a65fec226f5d78fc39f4a13c4cc0c768c22b113438f60c14adc9d2865038/pydantic_core-2.46.4-cp312-cp312-manylinux_2_17_ppc64le.manylinux2014_ppc64le.whl", hash = "sha256:5b712b53160b79a5850310b912a5ef8e57e56947c8ad690c227f5c9d7e561712", size = 2232274, upload-time = "2026-05-06T13:38:27.753Z" }, - { url = "https://files.pythonhosted.org/packages/68/f0/92039db98b907ef49269a8271f67db9cb78ae2fc68062ef7e4e77adb5f61/pydantic_core-2.46.4-cp312-cp312-manylinux_2_17_s390x.manylinux2014_s390x.whl", hash = "sha256:9401557acd873c3a7f3eb9383edef8ac4968f9510e340f4808d427e75667e7b4", size = 2309940, upload-time = "2026-05-06T13:38:05.353Z" }, - { url = "https://files.pythonhosted.org/packages/5f/97/2aab507d3d00ca626e8e57c1eac6a79e4e5fbcc63eb99733ff55d1717f65/pydantic_core-2.46.4-cp312-cp312-manylinux_2_17_x86_64.manylinux2014_x86_64.whl", hash = "sha256:926c9541b14b12b1681dca8a0b75feb510b06c6341b70a8e500c2fdcff837cce", size = 2094516, upload-time = "2026-05-06T13:39:10.577Z" }, - { url = "https://files.pythonhosted.org/packages/22/37/a8aca44d40d737dde2bc05b3c6c07dff0de07ce6f82e9f3167aeaf4d5dea/pydantic_core-2.46.4-cp312-cp312-manylinux_2_31_riscv64.whl", hash = "sha256:56cb4851bcaf3d117eddcef4fe66afd750a50274b0da8e22be256d10e5611987", size = 2136854, upload-time = "2026-05-06T13:40:22.59Z" }, - { url = "https://files.pythonhosted.org/packages/24/99/fcef1b79238c06a8cbec70819ac722ba76e02bc8ada9b0fd66eba40da01b/pydantic_core-2.46.4-cp312-cp312-manylinux_2_5_i686.manylinux1_i686.whl", hash = "sha256:c68fcd102d71ea85c5b2dfac3f4f8476eff42a9e078fd5faefff6d145063536b", size = 2180306, upload-time = "2026-05-06T13:40:10.666Z" }, - { url = "https://files.pythonhosted.org/packages/ae/6c/fc44000918855b42779d007ae63b0532794739027b2f417321cddbc44f6a/pydantic_core-2.46.4-cp312-cp312-musllinux_1_1_aarch64.whl", hash = "sha256:b2f69dec1725e79a012d920df1707de5caf7ed5e08f3be4435e25803efc47458", size = 2190044, upload-time = "2026-05-06T13:40:43.231Z" }, - { url = "https://files.pythonhosted.org/packages/6b/65/d9cadc9f1920d7a127ad2edba16c1db7916e59719285cd6c94600b0080ba/pydantic_core-2.46.4-cp312-cp312-musllinux_1_1_armv7l.whl", hash = "sha256:8d0820e8192167f80d88d64038e609c31452eeca865b4e1d9950a27a4609b00b", size = 2329133, upload-time = "2026-05-06T13:39:57.365Z" }, - { url = "https://files.pythonhosted.org/packages/d0/cf/c873d91679f3a30bcf5e7ac280ce5573483e72295307685120d0d5ad3416/pydantic_core-2.46.4-cp312-cp312-musllinux_1_1_x86_64.whl", hash = "sha256:fbdb89b3e1c94a30cc5edfce477c6e6a5dc4d8f84665b455c27582f211a1c72c", size = 2374464, upload-time = "2026-05-06T13:38:06.976Z" }, - { url = "https://files.pythonhosted.org/packages/47/bd/6f2fc8188f31bf10590f1e98e7b306336161fac930a8c514cd7bd828c7dc/pydantic_core-2.46.4-cp312-cp312-win32.whl", hash = "sha256:9aa768456404a8bf48a4406685ac2bec8e72b62c69313734fa3b73cf33b3a894", size = 1974823, upload-time = "2026-05-06T13:40:47.985Z" }, - { url = "https://files.pythonhosted.org/packages/40/8c/985c1d41ea1107c2534abd9870e4ed5c8e7669b5c308297835c001e7a1c4/pydantic_core-2.46.4-cp312-cp312-win_amd64.whl", hash = "sha256:e9c26f834c65f5752f3f06cb08cb86a913ceb7274d0db6e267808a708b46bc89", size = 2072919, upload-time = "2026-05-06T13:39:21.153Z" }, - { url = "https://files.pythonhosted.org/packages/c4/ba/f463d006e0c47373ca7ec5e1a261c59dc01ef4d62b2657af925fb0deee3a/pydantic_core-2.46.4-cp312-cp312-win_arm64.whl", hash = "sha256:4fc73cb559bdb54b1134a706a2802a4cddd27a0633f5abb7e53056268751ac6a", size = 2027604, upload-time = "2026-05-06T13:39:03.753Z" }, - { url = "https://files.pythonhosted.org/packages/ee/a4/73995fd4ebbb46ba0ee51e6fa049b8f02c40daebb762208feda8a6b7894d/pydantic_core-2.46.4-graalpy311-graalpy242_311_native-macosx_10_12_x86_64.whl", hash = "sha256:14d4edf427bdcf950a8a02d7cb44a08614388dd6e1bdcbf4f67504fa7887da9c", size = 2111589, upload-time = "2026-05-06T13:37:10.817Z" }, - { url = "https://files.pythonhosted.org/packages/fb/7f/f37d3a5e8bfcc2e403f5c57a730f2d815693fb42119e8ea48b3789335af1/pydantic_core-2.46.4-graalpy311-graalpy242_311_native-macosx_11_0_arm64.whl", hash = "sha256:0ce40cd7b21210e99342afafbd4d0f76d784eb5b1d60f3bdc566be4983c6c73b", size = 1944552, upload-time = "2026-05-06T13:36:56.717Z" }, - { url = "https://files.pythonhosted.org/packages/15/3c/d7eb777b3ff43e8433a4efb39a17aa8fd98a4ee8561a24a67ef5db07b2d6/pydantic_core-2.46.4-graalpy311-graalpy242_311_native-manylinux_2_17_aarch64.manylinux2014_aarch64.whl", hash = "sha256:90884113d8b48f760e9587002789ddd741e76ab9f89518cd1e43b1f1a52ec44b", size = 1982984, upload-time = "2026-05-06T13:39:06.207Z" }, - { url = "https://files.pythonhosted.org/packages/63/87/70b9f40170a81afd55ca26c9b2acb25c20d64bcfbf888fafecb3ba077d4c/pydantic_core-2.46.4-graalpy311-graalpy242_311_native-manylinux_2_17_x86_64.manylinux2014_x86_64.whl", hash = "sha256:66ce7632c22d837c95301830e111ad0128a32b8207533b60896a96c4915192ea", size = 2138417, upload-time = "2026-05-06T13:39:45.476Z" }, - { url = "https://files.pythonhosted.org/packages/9d/1d/8987ad40f65ae1432753072f214fb5c74fe47ffbd0698bb9cbbb585664f8/pydantic_core-2.46.4-graalpy312-graalpy250_312_native-macosx_10_12_x86_64.whl", hash = "sha256:1d8ba486450b14f3b1d63bc521d410ec7565e52f887b9fb671791886436a42f7", size = 2095527, upload-time = "2026-05-06T13:39:52.283Z" }, - { url = "https://files.pythonhosted.org/packages/64/d3/84c282a7eee1d3ac4c0377546ef5a1ea436ce26840d9ac3b7ed54a377507/pydantic_core-2.46.4-graalpy312-graalpy250_312_native-macosx_11_0_arm64.whl", hash = "sha256:3009f12e4e90b7f88b4f9adb1b0c4a3d58fe7820f3238c190047209d148026df", size = 1936024, upload-time = "2026-05-06T13:40:15.671Z" }, - { url = "https://files.pythonhosted.org/packages/d7/ca/eac61596cdeb4d7e174d3dc0bd8a6238f14f75f97a24e7b7db4c7e7340a0/pydantic_core-2.46.4-graalpy312-graalpy250_312_native-manylinux_2_17_aarch64.manylinux2014_aarch64.whl", hash = "sha256:ad785e92e6dc634c21555edc8bd6b64957ab844541bcb96a1366c202951ae526", size = 1990696, upload-time = "2026-05-06T13:38:34.717Z" }, - { url = "https://files.pythonhosted.org/packages/fa/c3/7c8b240552251faf6b3a957db200fcfbbcec36763c050428b601e0c9b83b/pydantic_core-2.46.4-graalpy312-graalpy250_312_native-manylinux_2_17_x86_64.manylinux2014_x86_64.whl", hash = "sha256:00c603d540afdd6b80eb39f078f33ebd46211f02f33e34a32d9f053bba711de0", size = 2147590, upload-time = "2026-05-06T13:39:29.883Z" }, - { url = "https://files.pythonhosted.org/packages/11/cb/428de0385b6c8d44b716feba566abfacfbd23ee3c4439faa789a1456242f/pydantic_core-2.46.4-pp311-pypy311_pp73-macosx_10_12_x86_64.whl", hash = "sha256:0c563b08bca408dc7f65f700633d8442fffb2421fc47b8101377e9fd65051ff0", size = 2112782, upload-time = "2026-05-06T13:37:04.016Z" }, - { url = "https://files.pythonhosted.org/packages/0b/b5/6a17bdadd0fc1f170adfd05a20d37c832f52b117b4d9131da1f41bb097ce/pydantic_core-2.46.4-pp311-pypy311_pp73-macosx_11_0_arm64.whl", hash = "sha256:db06ffe51636ffe9ca531fe9023dd64bdd794be8754cb5df57c5498ae5b518a7", size = 1952146, upload-time = "2026-05-06T13:39:43.092Z" }, - { url = "https://files.pythonhosted.org/packages/2a/dc/03734d80e362cd43ef65428e9de77c730ce7f2f11c60d2b1e1b39f0fbf99/pydantic_core-2.46.4-pp311-pypy311_pp73-manylinux_2_17_x86_64.manylinux2014_x86_64.whl", hash = "sha256:133878133d271ade3d41d1bfb2a45ec38dbdbda40bc065921c6b04e4630127e2", size = 2134492, upload-time = "2026-05-06T13:36:58.124Z" }, - { url = "https://files.pythonhosted.org/packages/de/df/5e5ffc085ed07cc22d298134d3d911c63e91f6a0eb91fe646750a3209910/pydantic_core-2.46.4-pp311-pypy311_pp73-manylinux_2_5_i686.manylinux1_i686.whl", hash = "sha256:9bc519fbf2b7578398853d815009ae5e4d4603d12f4e3f91da8c06852d3da3e9", size = 2156604, upload-time = "2026-05-06T13:37:49.88Z" }, - { url = "https://files.pythonhosted.org/packages/81/44/6e112a4253e56f5705467cbab7ab5e91ee7398ba3d56d358635958893d3e/pydantic_core-2.46.4-pp311-pypy311_pp73-musllinux_1_1_aarch64.whl", hash = "sha256:c7a7bd4e39e8e4c12c39cd480356842b6a8a06e41b23a55a5e3e191718838ddf", size = 2183828, upload-time = "2026-05-06T13:37:43.053Z" }, - { url = "https://files.pythonhosted.org/packages/ac/ad/5565071e937d8e752842ac241463944c9eb14c87e2d269f2658a5bd05e98/pydantic_core-2.46.4-pp311-pypy311_pp73-musllinux_1_1_armv7l.whl", hash = "sha256:d396ec2b979760aaf3218e76c24e65bd0aca24983298653b3a9d7a45f9e47b30", size = 2310000, upload-time = "2026-05-06T13:37:56.694Z" }, - { url = "https://files.pythonhosted.org/packages/4f/c3/66883a5cec183e7fba4d024b4cbbe61851a63750ef606b0afecc46d1f2bf/pydantic_core-2.46.4-pp311-pypy311_pp73-musllinux_1_1_x86_64.whl", hash = "sha256:86e1a4418c6cd97d60c95c71164158eaf7324fae7b0923264016baa993eba6fc", size = 2361286, upload-time = "2026-05-06T13:40:05.667Z" }, - { url = "https://files.pythonhosted.org/packages/4b/2d/69abac8f838090bbecd5df894befb2c2619e7996a98ddb949db9f3b93225/pydantic_core-2.46.4-pp311-pypy311_pp73-win_amd64.whl", hash = "sha256:d51026d73fcfd93610abc7b27789c26b313920fcfb20e27462d74a7f8b06e983", size = 2193071, upload-time = "2026-05-06T13:38:08.682Z" }, +sdist = { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pydantic-core/2.46.4/pydantic_core-2.46.4.tar.gz", hash = "sha256:62f875393d7f270851f20523dd2e29f082bcc82292d66db2b64ea71f64b6e1c1", size = 471464, upload-time = "2026-05-06T13:37:06.98Z" } +wheels = [ + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pydantic-core/2.46.4/pydantic_core-2.46.4-cp311-cp311-macosx_10_12_x86_64.whl", hash = "sha256:0e96592440881c74a213e5ad528e2b24d3d4f940de2766bed9010ab1d9e51594", size = 2111872, upload-time = "2026-05-06T13:40:27.596Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pydantic-core/2.46.4/pydantic_core-2.46.4-cp311-cp311-macosx_11_0_arm64.whl", hash = "sha256:e0d65b8c354be7fb5f720c3caa8bc940bc2d20ce749c8e06135f07f8ed95dd7c", size = 1948255, upload-time = "2026-05-06T13:39:12.574Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pydantic-core/2.46.4/pydantic_core-2.46.4-cp311-cp311-manylinux_2_17_aarch64.manylinux2014_aarch64.whl", hash = "sha256:7bfb192b3f4b9e8a89b6277b6ce787564f62cfd272055f6e685726b111dc7826", size = 1972827, upload-time = "2026-05-06T13:38:19.841Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pydantic-core/2.46.4/pydantic_core-2.46.4-cp311-cp311-manylinux_2_17_armv7l.manylinux2014_armv7l.whl", hash = "sha256:9037063db01f09b09e237c282b6792bd4da634b5402c4e7f0c61effed7701a04", size = 2041051, upload-time = "2026-05-06T13:38:10.447Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pydantic-core/2.46.4/pydantic_core-2.46.4-cp311-cp311-manylinux_2_17_ppc64le.manylinux2014_ppc64le.whl", hash = "sha256:fc010ab034c8c7452522748bf937df58020d256ccae0874463d1f4d01758af8e", size = 2221314, upload-time = "2026-05-06T13:40:13.089Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pydantic-core/2.46.4/pydantic_core-2.46.4-cp311-cp311-manylinux_2_17_s390x.manylinux2014_s390x.whl", hash = "sha256:8c5dac79fa1614d1e06ca695109c6105923bd9c7d1d6c918d4e637b7e6b32fd3", size = 2285146, upload-time = "2026-05-06T13:38:59.224Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pydantic-core/2.46.4/pydantic_core-2.46.4-cp311-cp311-manylinux_2_17_x86_64.manylinux2014_x86_64.whl", hash = "sha256:f9fa868638bf362d3d138ea55829cefb3d5f4b0d7f142234382a15e2485dbec4", size = 2089685, upload-time = "2026-05-06T13:38:17.762Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pydantic-core/2.46.4/pydantic_core-2.46.4-cp311-cp311-manylinux_2_31_riscv64.whl", hash = "sha256:17299feefe090f2caa5b8e37222bb5f663e4935a8bfa6931d4102e5df1a9f398", size = 2115420, upload-time = "2026-05-06T13:37:58.195Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pydantic-core/2.46.4/pydantic_core-2.46.4-cp311-cp311-manylinux_2_5_i686.manylinux1_i686.whl", hash = "sha256:4c63ebc82684aa89d9a3bcbd13d515b3be44250dc68dd3bd81526c1cb31286c3", size = 2165122, upload-time = "2026-05-06T13:37:01.167Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pydantic-core/2.46.4/pydantic_core-2.46.4-cp311-cp311-musllinux_1_1_aarch64.whl", hash = "sha256:aaa2a54443eff1950ba5ddc6b6ccda0d9c84a364276a62f969bdf2a390650848", size = 2182573, upload-time = "2026-05-06T13:38:45.04Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pydantic-core/2.46.4/pydantic_core-2.46.4-cp311-cp311-musllinux_1_1_armv7l.whl", hash = "sha256:18e5ceec2ab67e6d5f1a9085e5a24c9c4e2ac4545730bfe668680bca05e555f3", size = 2317139, upload-time = "2026-05-06T13:37:15.539Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pydantic-core/2.46.4/pydantic_core-2.46.4-cp311-cp311-musllinux_1_1_x86_64.whl", hash = "sha256:a0f62d0a58f4e7da165457e995725421e0064f2255d8eccebc49f41bbc23b109", size = 2360433, upload-time = "2026-05-06T13:37:30.099Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pydantic-core/2.46.4/pydantic_core-2.46.4-cp311-cp311-win32.whl", hash = "sha256:041bde0a48fd37cf71cab1c9d56d3e8625a3793fef1f7dd232b3ff37e978ecda", size = 1985513, upload-time = "2026-05-06T13:38:15.669Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pydantic-core/2.46.4/pydantic_core-2.46.4-cp311-cp311-win_amd64.whl", hash = "sha256:6f2eeda33a839975441c86a4119e1383c50b47faf0cbb5176985565c6bb02c33", size = 2071114, upload-time = "2026-05-06T13:40:35.416Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pydantic-core/2.46.4/pydantic_core-2.46.4-cp311-cp311-win_arm64.whl", hash = "sha256:14f4c5d6db102bd796a627bbb3a17b4cf4574b9ae861d8b7c9a9661c6dd3362d", size = 2044298, upload-time = "2026-05-06T13:38:29.754Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pydantic-core/2.46.4/pydantic_core-2.46.4-cp312-cp312-macosx_10_12_x86_64.whl", hash = "sha256:3245406455a5d98187ec35530fd772b1d799b26667980872c8d4614991e2c4a2", size = 2106158, upload-time = "2026-05-06T13:38:57.215Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pydantic-core/2.46.4/pydantic_core-2.46.4-cp312-cp312-macosx_11_0_arm64.whl", hash = "sha256:962ccbab7b642487b1d8b7df90ef677e03134cf1fd8880bf698649b22a69371f", size = 1951724, upload-time = "2026-05-06T13:37:02.697Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pydantic-core/2.46.4/pydantic_core-2.46.4-cp312-cp312-manylinux_2_17_aarch64.manylinux2014_aarch64.whl", hash = "sha256:8233f2947cf85404441fd7e0085f53b10c93e0ee78611099b5c7237e36aacbf7", size = 1975742, upload-time = "2026-05-06T13:37:09.448Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pydantic-core/2.46.4/pydantic_core-2.46.4-cp312-cp312-manylinux_2_17_armv7l.manylinux2014_armv7l.whl", hash = "sha256:3a233125ac121aa3ffba9a2b59edfc4a985a76092dc8279586ab4b71390875e7", size = 2052418, upload-time = "2026-05-06T13:37:38.234Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pydantic-core/2.46.4/pydantic_core-2.46.4-cp312-cp312-manylinux_2_17_ppc64le.manylinux2014_ppc64le.whl", hash = "sha256:5b712b53160b79a5850310b912a5ef8e57e56947c8ad690c227f5c9d7e561712", size = 2232274, upload-time = "2026-05-06T13:38:27.753Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pydantic-core/2.46.4/pydantic_core-2.46.4-cp312-cp312-manylinux_2_17_s390x.manylinux2014_s390x.whl", hash = "sha256:9401557acd873c3a7f3eb9383edef8ac4968f9510e340f4808d427e75667e7b4", size = 2309940, upload-time = "2026-05-06T13:38:05.353Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pydantic-core/2.46.4/pydantic_core-2.46.4-cp312-cp312-manylinux_2_17_x86_64.manylinux2014_x86_64.whl", hash = "sha256:926c9541b14b12b1681dca8a0b75feb510b06c6341b70a8e500c2fdcff837cce", size = 2094516, upload-time = "2026-05-06T13:39:10.577Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pydantic-core/2.46.4/pydantic_core-2.46.4-cp312-cp312-manylinux_2_31_riscv64.whl", hash = "sha256:56cb4851bcaf3d117eddcef4fe66afd750a50274b0da8e22be256d10e5611987", size = 2136854, upload-time = "2026-05-06T13:40:22.59Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pydantic-core/2.46.4/pydantic_core-2.46.4-cp312-cp312-manylinux_2_5_i686.manylinux1_i686.whl", hash = "sha256:c68fcd102d71ea85c5b2dfac3f4f8476eff42a9e078fd5faefff6d145063536b", size = 2180306, upload-time = "2026-05-06T13:40:10.666Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pydantic-core/2.46.4/pydantic_core-2.46.4-cp312-cp312-musllinux_1_1_aarch64.whl", hash = "sha256:b2f69dec1725e79a012d920df1707de5caf7ed5e08f3be4435e25803efc47458", size = 2190044, upload-time = "2026-05-06T13:40:43.231Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pydantic-core/2.46.4/pydantic_core-2.46.4-cp312-cp312-musllinux_1_1_armv7l.whl", hash = "sha256:8d0820e8192167f80d88d64038e609c31452eeca865b4e1d9950a27a4609b00b", size = 2329133, upload-time = "2026-05-06T13:39:57.365Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pydantic-core/2.46.4/pydantic_core-2.46.4-cp312-cp312-musllinux_1_1_x86_64.whl", hash = "sha256:fbdb89b3e1c94a30cc5edfce477c6e6a5dc4d8f84665b455c27582f211a1c72c", size = 2374464, upload-time = "2026-05-06T13:38:06.976Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pydantic-core/2.46.4/pydantic_core-2.46.4-cp312-cp312-win32.whl", hash = "sha256:9aa768456404a8bf48a4406685ac2bec8e72b62c69313734fa3b73cf33b3a894", size = 1974823, upload-time = "2026-05-06T13:40:47.985Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pydantic-core/2.46.4/pydantic_core-2.46.4-cp312-cp312-win_amd64.whl", hash = "sha256:e9c26f834c65f5752f3f06cb08cb86a913ceb7274d0db6e267808a708b46bc89", size = 2072919, upload-time = "2026-05-06T13:39:21.153Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pydantic-core/2.46.4/pydantic_core-2.46.4-cp312-cp312-win_arm64.whl", hash = "sha256:4fc73cb559bdb54b1134a706a2802a4cddd27a0633f5abb7e53056268751ac6a", size = 2027604, upload-time = "2026-05-06T13:39:03.753Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pydantic-core/2.46.4/pydantic_core-2.46.4-graalpy311-graalpy242_311_native-macosx_10_12_x86_64.whl", hash = "sha256:14d4edf427bdcf950a8a02d7cb44a08614388dd6e1bdcbf4f67504fa7887da9c", size = 2111589, upload-time = "2026-05-06T13:37:10.817Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pydantic-core/2.46.4/pydantic_core-2.46.4-graalpy311-graalpy242_311_native-macosx_11_0_arm64.whl", hash = "sha256:0ce40cd7b21210e99342afafbd4d0f76d784eb5b1d60f3bdc566be4983c6c73b", size = 1944552, upload-time = "2026-05-06T13:36:56.717Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pydantic-core/2.46.4/pydantic_core-2.46.4-graalpy311-graalpy242_311_native-manylinux_2_17_aarch64.manylinux2014_aarch64.whl", hash = "sha256:90884113d8b48f760e9587002789ddd741e76ab9f89518cd1e43b1f1a52ec44b", size = 1982984, upload-time = "2026-05-06T13:39:06.207Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pydantic-core/2.46.4/pydantic_core-2.46.4-graalpy311-graalpy242_311_native-manylinux_2_17_x86_64.manylinux2014_x86_64.whl", hash = "sha256:66ce7632c22d837c95301830e111ad0128a32b8207533b60896a96c4915192ea", size = 2138417, upload-time = "2026-05-06T13:39:45.476Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pydantic-core/2.46.4/pydantic_core-2.46.4-graalpy312-graalpy250_312_native-macosx_10_12_x86_64.whl", hash = "sha256:1d8ba486450b14f3b1d63bc521d410ec7565e52f887b9fb671791886436a42f7", size = 2095527, upload-time = "2026-05-06T13:39:52.283Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pydantic-core/2.46.4/pydantic_core-2.46.4-graalpy312-graalpy250_312_native-macosx_11_0_arm64.whl", hash = "sha256:3009f12e4e90b7f88b4f9adb1b0c4a3d58fe7820f3238c190047209d148026df", size = 1936024, upload-time = "2026-05-06T13:40:15.671Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pydantic-core/2.46.4/pydantic_core-2.46.4-graalpy312-graalpy250_312_native-manylinux_2_17_aarch64.manylinux2014_aarch64.whl", hash = "sha256:ad785e92e6dc634c21555edc8bd6b64957ab844541bcb96a1366c202951ae526", size = 1990696, upload-time = "2026-05-06T13:38:34.717Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pydantic-core/2.46.4/pydantic_core-2.46.4-graalpy312-graalpy250_312_native-manylinux_2_17_x86_64.manylinux2014_x86_64.whl", hash = "sha256:00c603d540afdd6b80eb39f078f33ebd46211f02f33e34a32d9f053bba711de0", size = 2147590, upload-time = "2026-05-06T13:39:29.883Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pydantic-core/2.46.4/pydantic_core-2.46.4-pp311-pypy311_pp73-macosx_10_12_x86_64.whl", hash = "sha256:0c563b08bca408dc7f65f700633d8442fffb2421fc47b8101377e9fd65051ff0", size = 2112782, upload-time = "2026-05-06T13:37:04.016Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pydantic-core/2.46.4/pydantic_core-2.46.4-pp311-pypy311_pp73-macosx_11_0_arm64.whl", hash = "sha256:db06ffe51636ffe9ca531fe9023dd64bdd794be8754cb5df57c5498ae5b518a7", size = 1952146, upload-time = "2026-05-06T13:39:43.092Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pydantic-core/2.46.4/pydantic_core-2.46.4-pp311-pypy311_pp73-manylinux_2_17_x86_64.manylinux2014_x86_64.whl", hash = "sha256:133878133d271ade3d41d1bfb2a45ec38dbdbda40bc065921c6b04e4630127e2", size = 2134492, upload-time = "2026-05-06T13:36:58.124Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pydantic-core/2.46.4/pydantic_core-2.46.4-pp311-pypy311_pp73-manylinux_2_5_i686.manylinux1_i686.whl", hash = "sha256:9bc519fbf2b7578398853d815009ae5e4d4603d12f4e3f91da8c06852d3da3e9", size = 2156604, upload-time = "2026-05-06T13:37:49.88Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pydantic-core/2.46.4/pydantic_core-2.46.4-pp311-pypy311_pp73-musllinux_1_1_aarch64.whl", hash = "sha256:c7a7bd4e39e8e4c12c39cd480356842b6a8a06e41b23a55a5e3e191718838ddf", size = 2183828, upload-time = "2026-05-06T13:37:43.053Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pydantic-core/2.46.4/pydantic_core-2.46.4-pp311-pypy311_pp73-musllinux_1_1_armv7l.whl", hash = "sha256:d396ec2b979760aaf3218e76c24e65bd0aca24983298653b3a9d7a45f9e47b30", size = 2310000, upload-time = "2026-05-06T13:37:56.694Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pydantic-core/2.46.4/pydantic_core-2.46.4-pp311-pypy311_pp73-musllinux_1_1_x86_64.whl", hash = "sha256:86e1a4418c6cd97d60c95c71164158eaf7324fae7b0923264016baa993eba6fc", size = 2361286, upload-time = "2026-05-06T13:40:05.667Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pydantic-core/2.46.4/pydantic_core-2.46.4-pp311-pypy311_pp73-win_amd64.whl", hash = "sha256:d51026d73fcfd93610abc7b27789c26b313920fcfb20e27462d74a7f8b06e983", size = 2193071, upload-time = "2026-05-06T13:38:08.682Z" }, ] [[package]] name = "pygments" version = "2.21.0" -source = { registry = "https://pypi.org/simple" } -sdist = { url = "https://files.pythonhosted.org/packages/49/2e/ced460408999b33da6b31b0021b0f37d329e202d4169aeb164493778f25b/pygments-2.21.0.tar.gz", hash = "sha256:610ca751c9bc2492b38eb9a38a7fbc93edbbb2d7182edaf34e66ae493dee5c8c", size = 5005329, upload-time = "2026-08-17T08:02:48.824Z" } +source = { registry = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/" } +sdist = { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pygments/2.21.0/pygments-2.21.0.tar.gz", hash = "sha256:610ca751c9bc2492b38eb9a38a7fbc93edbbb2d7182edaf34e66ae493dee5c8c", size = 5005329, upload-time = "2026-08-17T08:02:48.824Z" } wheels = [ - { url = "https://files.pythonhosted.org/packages/71/46/17f022dd3e953bf20a04a028a21ec746d942f8d2af30fa0f124fa0e6a684/pygments-2.21.0-py3-none-any.whl", hash = "sha256:2363c69b61c4a97c838da3b130dcd6468f4848992b21a82f2a63ec34377137d9", size = 1250147, upload-time = "2026-08-17T08:02:44.912Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pygments/2.21.0/pygments-2.21.0-py3-none-any.whl", hash = "sha256:2363c69b61c4a97c838da3b130dcd6468f4848992b21a82f2a63ec34377137d9", size = 1250147, upload-time = "2026-08-17T08:02:44.912Z" }, ] [[package]] name = "pyqwest" version = "0.10.0" -source = { registry = "https://pypi.org/simple" } +source = { registry = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/" } dependencies = [ { name = "opentelemetry-api" }, ] -sdist = { url = "https://files.pythonhosted.org/packages/07/ee/0ff9facfa9e7a4f6df2a770d4eaf1ad0f74165da7e8c28e888461f07604c/pyqwest-0.10.0.tar.gz", hash = "sha256:6c1a693be17d57d2c2eca4085e32c2809c53090c16719a907c90ebcf1f40dc01", size = 482248, upload-time = "2026-08-21T06:09:20.656Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/43/ee/b1a28f57c689606cfd065d8a553841150f7daaa91d20e58dcc2c5ea191f8/pyqwest-0.10.0-cp310-abi3-macosx_10_12_x86_64.whl", hash = "sha256:aa492d5777dd145a60795ed95d9d4707a3cd1091fdcdfc93a82ac7fdc43ebacd", size = 5261059, upload-time = "2026-08-21T06:08:04.999Z" }, - { url = "https://files.pythonhosted.org/packages/dc/13/9c5046cfd6ef705bde0b620ba8a794335bcabc0839342a2a647f2427b27e/pyqwest-0.10.0-cp310-abi3-macosx_11_0_arm64.whl", hash = "sha256:59f3f16628e518c674102e7b5fcff2101bba6abb4f6737ec5fade9b9278e6a53", size = 5134207, upload-time = "2026-08-21T06:08:06.955Z" }, - { url = "https://files.pythonhosted.org/packages/93/7d/50021dd88d82d6966ab1c27593ceaee9d1ed62fbe597c40e8dc187cfa5fd/pyqwest-0.10.0-cp310-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl", hash = "sha256:d6e7db305a8318b1f3218053e87501f8f245ca8bd63e948e0282d04bf0883470", size = 5640730, upload-time = "2026-08-21T06:08:09.135Z" }, - { url = "https://files.pythonhosted.org/packages/ff/3f/5bf6c32e9e701837a8c47ce6e3ad38978cfec8eb7bc6596181e5f9e1eaeb/pyqwest-0.10.0-cp310-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl", hash = "sha256:a5c757cfac5f53c8671dcb4850d5fc4c4339ea3e90636331c9318f8e3ddabc06", size = 5561462, upload-time = "2026-08-21T06:08:10.836Z" }, - { url = "https://files.pythonhosted.org/packages/5f/61/ca9ba5721461b7ce5cfaac373ab3a1723ddcc434af7430f8bd628da6e623/pyqwest-0.10.0-cp310-abi3-musllinux_1_2_aarch64.whl", hash = "sha256:234b3f71e3f314d997c203d8cf829b7117edd041153f9c277d0060ab90134148", size = 5801847, upload-time = "2026-08-21T06:08:12.502Z" }, - { url = "https://files.pythonhosted.org/packages/d2/44/95593919b996a417093f598d887822b9b899e8d025588c9bfaf8c60dd812/pyqwest-0.10.0-cp310-abi3-musllinux_1_2_x86_64.whl", hash = "sha256:5637256a0dac0ef57e0eaa02b032014965e4a4c995e1deca1b1b97e6d1765f78", size = 5978692, upload-time = "2026-08-21T06:08:14.383Z" }, - { url = "https://files.pythonhosted.org/packages/a5/b9/b2821ce5188457168ebb25d5ff65b1ca1bf27bc6b4a33df4bcc2357e625c/pyqwest-0.10.0-cp310-abi3-win_amd64.whl", hash = "sha256:7ea761937acf3a00d1a7e70e982949d18946e5471d1419266ab3a78bbfa19759", size = 4876627, upload-time = "2026-08-21T06:08:16.084Z" }, - { url = "https://files.pythonhosted.org/packages/86/b4/16ccef1c203fa258ce46a86aefc1a79c13b5f0b8d49627347d90eef25efd/pyqwest-0.10.0-cp312-cp312-macosx_10_12_x86_64.whl", hash = "sha256:a21f1f15252a8303623b4f17b9c6de595ace11b3ade07f2adb6d07121e8191aa", size = 5274815, upload-time = "2026-08-21T06:08:17.777Z" }, - { url = "https://files.pythonhosted.org/packages/5f/61/6a87f84f571441ea43279587d4bfcad4543505918ae2b83a1ebdcfa98be5/pyqwest-0.10.0-cp312-cp312-macosx_11_0_arm64.whl", hash = "sha256:eb472c6e5d6833ebfec79db310e426eb17b01ac64c0e2c251bd9192c0d2ee0c5", size = 5123656, upload-time = "2026-08-21T06:08:19.501Z" }, - { url = "https://files.pythonhosted.org/packages/8b/f2/ab69e581cf9b798b0e169f7b27fd3f8b6f9f1631bd4d3b6e22e5abaf8d8b/pyqwest-0.10.0-cp312-cp312-manylinux_2_17_aarch64.manylinux2014_aarch64.whl", hash = "sha256:83578e24cccd5e0dc04d60a0af7bfb43325b5f22d03ff74ff79ed0ecf553b50d", size = 5641253, upload-time = "2026-08-21T06:08:21.627Z" }, - { url = "https://files.pythonhosted.org/packages/9f/dd/f1a62eebf8321ace506bd94551a01431f6a45b882225455eae3ea6e8c6d1/pyqwest-0.10.0-cp312-cp312-manylinux_2_17_x86_64.manylinux2014_x86_64.whl", hash = "sha256:6bb511c434f79c641efb5573e5795e56dc972252f4b96e52a9636d4ece5231a4", size = 5567341, upload-time = "2026-08-21T06:08:23.346Z" }, - { url = "https://files.pythonhosted.org/packages/4c/48/51d767691973e046887f5e6d96e32142fe296823b163ccba732233a6ef72/pyqwest-0.10.0-cp312-cp312-musllinux_1_2_aarch64.whl", hash = "sha256:1aaccd8a9db9430b2aedb5bad8ead80742cbc056b85c229516c70dc80539f906", size = 5803874, upload-time = "2026-08-21T06:08:25.096Z" }, - { url = "https://files.pythonhosted.org/packages/58/0a/d2834ccc6e59ad110718895cc65ff2a68aa6e010f0ba8fbe42a57ea33c21/pyqwest-0.10.0-cp312-cp312-musllinux_1_2_x86_64.whl", hash = "sha256:73d9eb438ab4a957a1ce0619d3af8c1c1126bfb9181033b123d792fcf4224531", size = 5981518, upload-time = "2026-08-21T06:08:27.594Z" }, - { url = "https://files.pythonhosted.org/packages/75/f3/274b4c268e9a55fbdb1b3637ac50b5bf42cd3a85d1cfbdc15c602a7b0d9c/pyqwest-0.10.0-cp312-cp312-win_amd64.whl", hash = "sha256:317a74d633abe3bc5bccabf479e069c515dab9e6a755274b0ccb1d8a5bbfede3", size = 4870638, upload-time = "2026-08-21T06:08:29.421Z" }, - { url = "https://files.pythonhosted.org/packages/7a/6f/29f605665f33aab894db8daf11b3b64bdca015fb11135815c53a150191a6/pyqwest-0.10.0-pp311-pypy311_pp73-macosx_10_12_x86_64.whl", hash = "sha256:cfcc7ba0229baa17831582befb046ace167b368140dae022d0b89b8d586ba12c", size = 5264903, upload-time = "2026-08-21T06:09:09.183Z" }, - { url = "https://files.pythonhosted.org/packages/c8/c8/4ce4b40f21397a6482da910fae18b32c28dcdaddf3498aef2fe38597a9e0/pyqwest-0.10.0-pp311-pypy311_pp73-macosx_11_0_arm64.whl", hash = "sha256:eff9ccf427604d34c635954def07b6113d4754f968073eef2df40bdb80b05bf5", size = 5142291, upload-time = "2026-08-21T06:09:10.943Z" }, - { url = "https://files.pythonhosted.org/packages/41/34/7205db9cddc5a2286da459ca013af629e20f09498751c60093b952379add/pyqwest-0.10.0-pp311-pypy311_pp73-manylinux_2_17_aarch64.manylinux2014_aarch64.whl", hash = "sha256:78662158093f9d5c742368f4dd9956aa595f44c6ac0860777c98853ddd5e1610", size = 5648118, upload-time = "2026-08-21T06:09:12.616Z" }, - { url = "https://files.pythonhosted.org/packages/1e/94/75c6cd5d01cf68c99c0cfd66b165dce3fd9d9fce97cd4973117cab68355d/pyqwest-0.10.0-pp311-pypy311_pp73-manylinux_2_17_x86_64.manylinux2014_x86_64.whl", hash = "sha256:0879a3f0b37876372aa328b9fee956165db174c1ab72464146fe515f49399ddb", size = 5565620, upload-time = "2026-08-21T06:09:14.223Z" }, - { url = "https://files.pythonhosted.org/packages/7a/a6/20a910d2f096908cf25ec47fa84b35b1bbb1af63dc9713f53f38ed97802e/pyqwest-0.10.0-pp311-pypy311_pp73-musllinux_1_2_aarch64.whl", hash = "sha256:850b8de6ade09a60bdb2f969871a177c2c304b594b2034ba3f5962c7bea75551", size = 5809262, upload-time = "2026-08-21T06:09:15.897Z" }, - { url = "https://files.pythonhosted.org/packages/e4/f7/f92fe4004d93f08c5118e750feefdbc72a73438b7fd7083300728e999911/pyqwest-0.10.0-pp311-pypy311_pp73-musllinux_1_2_x86_64.whl", hash = "sha256:399802647ea646c6ac9b5460e541b7c209b7a13563c667b2690ded2060185f2e", size = 5985550, upload-time = "2026-08-21T06:09:17.579Z" }, - { url = "https://files.pythonhosted.org/packages/47/3e/2c896e54dbe3f1ba6e3bd10e9d412f422a2e33827ff591d59f67b13e0fa5/pyqwest-0.10.0-pp311-pypy311_pp73-win_amd64.whl", hash = "sha256:c26f3de1feb5d066d7a66802a47407a93ba043696064ad80beda4a0a4bf10056", size = 4873529, upload-time = "2026-08-21T06:09:19.161Z" }, +sdist = { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pyqwest/0.10.0/pyqwest-0.10.0.tar.gz", hash = "sha256:6c1a693be17d57d2c2eca4085e32c2809c53090c16719a907c90ebcf1f40dc01", size = 482248, upload-time = "2026-08-21T06:09:20.656Z" } +wheels = [ + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pyqwest/0.10.0/pyqwest-0.10.0-cp310-abi3-macosx_10_12_x86_64.whl", hash = "sha256:aa492d5777dd145a60795ed95d9d4707a3cd1091fdcdfc93a82ac7fdc43ebacd", size = 5261059, upload-time = "2026-08-21T06:08:04.999Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pyqwest/0.10.0/pyqwest-0.10.0-cp310-abi3-macosx_11_0_arm64.whl", hash = "sha256:59f3f16628e518c674102e7b5fcff2101bba6abb4f6737ec5fade9b9278e6a53", size = 5134207, upload-time = "2026-08-21T06:08:06.955Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pyqwest/0.10.0/pyqwest-0.10.0-cp310-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl", hash = "sha256:d6e7db305a8318b1f3218053e87501f8f245ca8bd63e948e0282d04bf0883470", size = 5640730, upload-time = "2026-08-21T06:08:09.135Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pyqwest/0.10.0/pyqwest-0.10.0-cp310-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl", hash = "sha256:a5c757cfac5f53c8671dcb4850d5fc4c4339ea3e90636331c9318f8e3ddabc06", size = 5561462, upload-time = "2026-08-21T06:08:10.836Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pyqwest/0.10.0/pyqwest-0.10.0-cp310-abi3-musllinux_1_2_aarch64.whl", hash = "sha256:234b3f71e3f314d997c203d8cf829b7117edd041153f9c277d0060ab90134148", size = 5801847, upload-time = "2026-08-21T06:08:12.502Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pyqwest/0.10.0/pyqwest-0.10.0-cp310-abi3-musllinux_1_2_x86_64.whl", hash = "sha256:5637256a0dac0ef57e0eaa02b032014965e4a4c995e1deca1b1b97e6d1765f78", size = 5978692, upload-time = "2026-08-21T06:08:14.383Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pyqwest/0.10.0/pyqwest-0.10.0-cp310-abi3-win_amd64.whl", hash = "sha256:7ea761937acf3a00d1a7e70e982949d18946e5471d1419266ab3a78bbfa19759", size = 4876627, upload-time = "2026-08-21T06:08:16.084Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pyqwest/0.10.0/pyqwest-0.10.0-cp312-cp312-macosx_10_12_x86_64.whl", hash = "sha256:a21f1f15252a8303623b4f17b9c6de595ace11b3ade07f2adb6d07121e8191aa", size = 5274815, upload-time = "2026-08-21T06:08:17.777Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pyqwest/0.10.0/pyqwest-0.10.0-cp312-cp312-macosx_11_0_arm64.whl", hash = "sha256:eb472c6e5d6833ebfec79db310e426eb17b01ac64c0e2c251bd9192c0d2ee0c5", size = 5123656, upload-time = "2026-08-21T06:08:19.501Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pyqwest/0.10.0/pyqwest-0.10.0-cp312-cp312-manylinux_2_17_aarch64.manylinux2014_aarch64.whl", hash = "sha256:83578e24cccd5e0dc04d60a0af7bfb43325b5f22d03ff74ff79ed0ecf553b50d", size = 5641253, upload-time = "2026-08-21T06:08:21.627Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pyqwest/0.10.0/pyqwest-0.10.0-cp312-cp312-manylinux_2_17_x86_64.manylinux2014_x86_64.whl", hash = "sha256:6bb511c434f79c641efb5573e5795e56dc972252f4b96e52a9636d4ece5231a4", size = 5567341, upload-time = "2026-08-21T06:08:23.346Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pyqwest/0.10.0/pyqwest-0.10.0-cp312-cp312-musllinux_1_2_aarch64.whl", hash = "sha256:1aaccd8a9db9430b2aedb5bad8ead80742cbc056b85c229516c70dc80539f906", size = 5803874, upload-time = "2026-08-21T06:08:25.096Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pyqwest/0.10.0/pyqwest-0.10.0-cp312-cp312-musllinux_1_2_x86_64.whl", hash = "sha256:73d9eb438ab4a957a1ce0619d3af8c1c1126bfb9181033b123d792fcf4224531", size = 5981518, upload-time = "2026-08-21T06:08:27.594Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pyqwest/0.10.0/pyqwest-0.10.0-cp312-cp312-win_amd64.whl", hash = "sha256:317a74d633abe3bc5bccabf479e069c515dab9e6a755274b0ccb1d8a5bbfede3", size = 4870638, upload-time = "2026-08-21T06:08:29.421Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pyqwest/0.10.0/pyqwest-0.10.0-pp311-pypy311_pp73-macosx_10_12_x86_64.whl", hash = "sha256:cfcc7ba0229baa17831582befb046ace167b368140dae022d0b89b8d586ba12c", size = 5264903, upload-time = "2026-08-21T06:09:09.183Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pyqwest/0.10.0/pyqwest-0.10.0-pp311-pypy311_pp73-macosx_11_0_arm64.whl", hash = "sha256:eff9ccf427604d34c635954def07b6113d4754f968073eef2df40bdb80b05bf5", size = 5142291, upload-time = "2026-08-21T06:09:10.943Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pyqwest/0.10.0/pyqwest-0.10.0-pp311-pypy311_pp73-manylinux_2_17_aarch64.manylinux2014_aarch64.whl", hash = "sha256:78662158093f9d5c742368f4dd9956aa595f44c6ac0860777c98853ddd5e1610", size = 5648118, upload-time = "2026-08-21T06:09:12.616Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pyqwest/0.10.0/pyqwest-0.10.0-pp311-pypy311_pp73-manylinux_2_17_x86_64.manylinux2014_x86_64.whl", hash = "sha256:0879a3f0b37876372aa328b9fee956165db174c1ab72464146fe515f49399ddb", size = 5565620, upload-time = "2026-08-21T06:09:14.223Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pyqwest/0.10.0/pyqwest-0.10.0-pp311-pypy311_pp73-musllinux_1_2_aarch64.whl", hash = "sha256:850b8de6ade09a60bdb2f969871a177c2c304b594b2034ba3f5962c7bea75551", size = 5809262, upload-time = "2026-08-21T06:09:15.897Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pyqwest/0.10.0/pyqwest-0.10.0-pp311-pypy311_pp73-musllinux_1_2_x86_64.whl", hash = "sha256:399802647ea646c6ac9b5460e541b7c209b7a13563c667b2690ded2060185f2e", size = 5985550, upload-time = "2026-08-21T06:09:17.579Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pyqwest/0.10.0/pyqwest-0.10.0-pp311-pypy311_pp73-win_amd64.whl", hash = "sha256:c26f3de1feb5d066d7a66802a47407a93ba043696064ad80beda4a0a4bf10056", size = 4873529, upload-time = "2026-08-21T06:09:19.161Z" }, ] [[package]] name = "pytest" version = "9.1.1" -source = { registry = "https://pypi.org/simple" } +source = { registry = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/" } dependencies = [ { name = "colorama", marker = "sys_platform == 'win32'" }, { name = "iniconfig" }, @@ -977,159 +977,159 @@ dependencies = [ { name = "pluggy" }, { name = "pygments" }, ] -sdist = { url = "https://files.pythonhosted.org/packages/e4/47/b9efed96c114afcfa3c9d3fe98a76a1d14c74a9e266d397cf6eb64be5e01/pytest-9.1.1.tar.gz", hash = "sha256:1088fbde8f2b49d95a549a195707afa7a76a3ce9bcadc26b6d71f0ffda5fe313", size = 1636369, upload-time = "2026-06-19T10:58:32.857Z" } +sdist = { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pytest/9.1.1/pytest-9.1.1.tar.gz", hash = "sha256:1088fbde8f2b49d95a549a195707afa7a76a3ce9bcadc26b6d71f0ffda5fe313", size = 1636369, upload-time = "2026-06-19T10:58:32.857Z" } wheels = [ - { url = "https://files.pythonhosted.org/packages/24/25/1de2678b631f5a49215c6c96fff41ba892b0a34df68d6d80292b1b48aa7f/pytest-9.1.1-py3-none-any.whl", hash = "sha256:37a86b45efb9a47a61a36449063e8e18d0cab3161329fc099eb21783169c4f0c", size = 386536, upload-time = "2026-06-19T10:58:31.347Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pytest/9.1.1/pytest-9.1.1-py3-none-any.whl", hash = "sha256:37a86b45efb9a47a61a36449063e8e18d0cab3161329fc099eb21783169c4f0c", size = 386536, upload-time = "2026-06-19T10:58:31.347Z" }, ] [[package]] name = "pyyaml" version = "6.0.3" -source = { registry = "https://pypi.org/simple" } -sdist = { url = "https://files.pythonhosted.org/packages/05/8e/961c0007c59b8dd7729d542c61a4d537767a59645b82a0b521206e1e25c2/pyyaml-6.0.3.tar.gz", hash = "sha256:d76623373421df22fb4cf8817020cbb7ef15c725b9d5e45f17e189bfc384190f", size = 130960, upload-time = "2025-09-25T21:33:16.546Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/6d/16/a95b6757765b7b031c9374925bb718d55e0a9ba8a1b6a12d25962ea44347/pyyaml-6.0.3-cp311-cp311-macosx_10_13_x86_64.whl", hash = "sha256:44edc647873928551a01e7a563d7452ccdebee747728c1080d881d68af7b997e", size = 185826, upload-time = "2025-09-25T21:31:58.655Z" }, - { url = "https://files.pythonhosted.org/packages/16/19/13de8e4377ed53079ee996e1ab0a9c33ec2faf808a4647b7b4c0d46dd239/pyyaml-6.0.3-cp311-cp311-macosx_11_0_arm64.whl", hash = "sha256:652cb6edd41e718550aad172851962662ff2681490a8a711af6a4d288dd96824", size = 175577, upload-time = "2025-09-25T21:32:00.088Z" }, - { url = "https://files.pythonhosted.org/packages/0c/62/d2eb46264d4b157dae1275b573017abec435397aa59cbcdab6fc978a8af4/pyyaml-6.0.3-cp311-cp311-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:10892704fc220243f5305762e276552a0395f7beb4dbf9b14ec8fd43b57f126c", size = 775556, upload-time = "2025-09-25T21:32:01.31Z" }, - { url = "https://files.pythonhosted.org/packages/10/cb/16c3f2cf3266edd25aaa00d6c4350381c8b012ed6f5276675b9eba8d9ff4/pyyaml-6.0.3-cp311-cp311-manylinux2014_s390x.manylinux_2_17_s390x.manylinux_2_28_s390x.whl", hash = "sha256:850774a7879607d3a6f50d36d04f00ee69e7fc816450e5f7e58d7f17f1ae5c00", size = 882114, upload-time = "2025-09-25T21:32:03.376Z" }, - { url = "https://files.pythonhosted.org/packages/71/60/917329f640924b18ff085ab889a11c763e0b573da888e8404ff486657602/pyyaml-6.0.3-cp311-cp311-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:b8bb0864c5a28024fac8a632c443c87c5aa6f215c0b126c449ae1a150412f31d", size = 806638, upload-time = "2025-09-25T21:32:04.553Z" }, - { url = "https://files.pythonhosted.org/packages/dd/6f/529b0f316a9fd167281a6c3826b5583e6192dba792dd55e3203d3f8e655a/pyyaml-6.0.3-cp311-cp311-musllinux_1_2_aarch64.whl", hash = "sha256:1d37d57ad971609cf3c53ba6a7e365e40660e3be0e5175fa9f2365a379d6095a", size = 767463, upload-time = "2025-09-25T21:32:06.152Z" }, - { url = "https://files.pythonhosted.org/packages/f2/6a/b627b4e0c1dd03718543519ffb2f1deea4a1e6d42fbab8021936a4d22589/pyyaml-6.0.3-cp311-cp311-musllinux_1_2_x86_64.whl", hash = "sha256:37503bfbfc9d2c40b344d06b2199cf0e96e97957ab1c1b546fd4f87e53e5d3e4", size = 794986, upload-time = "2025-09-25T21:32:07.367Z" }, - { url = "https://files.pythonhosted.org/packages/45/91/47a6e1c42d9ee337c4839208f30d9f09caa9f720ec7582917b264defc875/pyyaml-6.0.3-cp311-cp311-win32.whl", hash = "sha256:8098f252adfa6c80ab48096053f512f2321f0b998f98150cea9bd23d83e1467b", size = 142543, upload-time = "2025-09-25T21:32:08.95Z" }, - { url = "https://files.pythonhosted.org/packages/da/e3/ea007450a105ae919a72393cb06f122f288ef60bba2dc64b26e2646fa315/pyyaml-6.0.3-cp311-cp311-win_amd64.whl", hash = "sha256:9f3bfb4965eb874431221a3ff3fdcddc7e74e3b07799e0e84ca4a0f867d449bf", size = 158763, upload-time = "2025-09-25T21:32:09.96Z" }, - { url = "https://files.pythonhosted.org/packages/d1/33/422b98d2195232ca1826284a76852ad5a86fe23e31b009c9886b2d0fb8b2/pyyaml-6.0.3-cp312-cp312-macosx_10_13_x86_64.whl", hash = "sha256:7f047e29dcae44602496db43be01ad42fc6f1cc0d8cd6c83d342306c32270196", size = 182063, upload-time = "2025-09-25T21:32:11.445Z" }, - { url = "https://files.pythonhosted.org/packages/89/a0/6cf41a19a1f2f3feab0e9c0b74134aa2ce6849093d5517a0c550fe37a648/pyyaml-6.0.3-cp312-cp312-macosx_11_0_arm64.whl", hash = "sha256:fc09d0aa354569bc501d4e787133afc08552722d3ab34836a80547331bb5d4a0", size = 173973, upload-time = "2025-09-25T21:32:12.492Z" }, - { url = "https://files.pythonhosted.org/packages/ed/23/7a778b6bd0b9a8039df8b1b1d80e2e2ad78aa04171592c8a5c43a56a6af4/pyyaml-6.0.3-cp312-cp312-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:9149cad251584d5fb4981be1ecde53a1ca46c891a79788c0df828d2f166bda28", size = 775116, upload-time = "2025-09-25T21:32:13.652Z" }, - { url = "https://files.pythonhosted.org/packages/65/30/d7353c338e12baef4ecc1b09e877c1970bd3382789c159b4f89d6a70dc09/pyyaml-6.0.3-cp312-cp312-manylinux2014_s390x.manylinux_2_17_s390x.manylinux_2_28_s390x.whl", hash = "sha256:5fdec68f91a0c6739b380c83b951e2c72ac0197ace422360e6d5a959d8d97b2c", size = 844011, upload-time = "2025-09-25T21:32:15.21Z" }, - { url = "https://files.pythonhosted.org/packages/8b/9d/b3589d3877982d4f2329302ef98a8026e7f4443c765c46cfecc8858c6b4b/pyyaml-6.0.3-cp312-cp312-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:ba1cc08a7ccde2d2ec775841541641e4548226580ab850948cbfda66a1befcdc", size = 807870, upload-time = "2025-09-25T21:32:16.431Z" }, - { url = "https://files.pythonhosted.org/packages/05/c0/b3be26a015601b822b97d9149ff8cb5ead58c66f981e04fedf4e762f4bd4/pyyaml-6.0.3-cp312-cp312-musllinux_1_2_aarch64.whl", hash = "sha256:8dc52c23056b9ddd46818a57b78404882310fb473d63f17b07d5c40421e47f8e", size = 761089, upload-time = "2025-09-25T21:32:17.56Z" }, - { url = "https://files.pythonhosted.org/packages/be/8e/98435a21d1d4b46590d5459a22d88128103f8da4c2d4cb8f14f2a96504e1/pyyaml-6.0.3-cp312-cp312-musllinux_1_2_x86_64.whl", hash = "sha256:41715c910c881bc081f1e8872880d3c650acf13dfa8214bad49ed4cede7c34ea", size = 790181, upload-time = "2025-09-25T21:32:18.834Z" }, - { url = "https://files.pythonhosted.org/packages/74/93/7baea19427dcfbe1e5a372d81473250b379f04b1bd3c4c5ff825e2327202/pyyaml-6.0.3-cp312-cp312-win32.whl", hash = "sha256:96b533f0e99f6579b3d4d4995707cf36df9100d67e0c8303a0c55b27b5f99bc5", size = 137658, upload-time = "2025-09-25T21:32:20.209Z" }, - { url = "https://files.pythonhosted.org/packages/86/bf/899e81e4cce32febab4fb42bb97dcdf66bc135272882d1987881a4b519e9/pyyaml-6.0.3-cp312-cp312-win_amd64.whl", hash = "sha256:5fcd34e47f6e0b794d17de1b4ff496c00986e1c83f7ab2fb8fcfe9616ff7477b", size = 154003, upload-time = "2025-09-25T21:32:21.167Z" }, - { url = "https://files.pythonhosted.org/packages/1a/08/67bd04656199bbb51dbed1439b7f27601dfb576fb864099c7ef0c3e55531/pyyaml-6.0.3-cp312-cp312-win_arm64.whl", hash = "sha256:64386e5e707d03a7e172c0701abfb7e10f0fb753ee1d773128192742712a98fd", size = 140344, upload-time = "2025-09-25T21:32:22.617Z" }, +source = { registry = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/" } +sdist = { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pyyaml/6.0.3/pyyaml-6.0.3.tar.gz", hash = "sha256:d76623373421df22fb4cf8817020cbb7ef15c725b9d5e45f17e189bfc384190f", size = 130960, upload-time = "2025-09-25T21:33:16.546Z" } +wheels = [ + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pyyaml/6.0.3/pyyaml-6.0.3-cp311-cp311-macosx_10_13_x86_64.whl", hash = "sha256:44edc647873928551a01e7a563d7452ccdebee747728c1080d881d68af7b997e", size = 185826, upload-time = "2025-09-25T21:31:58.655Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pyyaml/6.0.3/pyyaml-6.0.3-cp311-cp311-macosx_11_0_arm64.whl", hash = "sha256:652cb6edd41e718550aad172851962662ff2681490a8a711af6a4d288dd96824", size = 175577, upload-time = "2025-09-25T21:32:00.088Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pyyaml/6.0.3/pyyaml-6.0.3-cp311-cp311-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:10892704fc220243f5305762e276552a0395f7beb4dbf9b14ec8fd43b57f126c", size = 775556, upload-time = "2025-09-25T21:32:01.31Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pyyaml/6.0.3/pyyaml-6.0.3-cp311-cp311-manylinux2014_s390x.manylinux_2_17_s390x.manylinux_2_28_s390x.whl", hash = "sha256:850774a7879607d3a6f50d36d04f00ee69e7fc816450e5f7e58d7f17f1ae5c00", size = 882114, upload-time = "2025-09-25T21:32:03.376Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pyyaml/6.0.3/pyyaml-6.0.3-cp311-cp311-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:b8bb0864c5a28024fac8a632c443c87c5aa6f215c0b126c449ae1a150412f31d", size = 806638, upload-time = "2025-09-25T21:32:04.553Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pyyaml/6.0.3/pyyaml-6.0.3-cp311-cp311-musllinux_1_2_aarch64.whl", hash = "sha256:1d37d57ad971609cf3c53ba6a7e365e40660e3be0e5175fa9f2365a379d6095a", size = 767463, upload-time = "2025-09-25T21:32:06.152Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pyyaml/6.0.3/pyyaml-6.0.3-cp311-cp311-musllinux_1_2_x86_64.whl", hash = "sha256:37503bfbfc9d2c40b344d06b2199cf0e96e97957ab1c1b546fd4f87e53e5d3e4", size = 794986, upload-time = "2025-09-25T21:32:07.367Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pyyaml/6.0.3/pyyaml-6.0.3-cp311-cp311-win32.whl", hash = "sha256:8098f252adfa6c80ab48096053f512f2321f0b998f98150cea9bd23d83e1467b", size = 142543, upload-time = "2025-09-25T21:32:08.95Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pyyaml/6.0.3/pyyaml-6.0.3-cp311-cp311-win_amd64.whl", hash = "sha256:9f3bfb4965eb874431221a3ff3fdcddc7e74e3b07799e0e84ca4a0f867d449bf", size = 158763, upload-time = "2025-09-25T21:32:09.96Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pyyaml/6.0.3/pyyaml-6.0.3-cp312-cp312-macosx_10_13_x86_64.whl", hash = "sha256:7f047e29dcae44602496db43be01ad42fc6f1cc0d8cd6c83d342306c32270196", size = 182063, upload-time = "2025-09-25T21:32:11.445Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pyyaml/6.0.3/pyyaml-6.0.3-cp312-cp312-macosx_11_0_arm64.whl", hash = "sha256:fc09d0aa354569bc501d4e787133afc08552722d3ab34836a80547331bb5d4a0", size = 173973, upload-time = "2025-09-25T21:32:12.492Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pyyaml/6.0.3/pyyaml-6.0.3-cp312-cp312-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:9149cad251584d5fb4981be1ecde53a1ca46c891a79788c0df828d2f166bda28", size = 775116, upload-time = "2025-09-25T21:32:13.652Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pyyaml/6.0.3/pyyaml-6.0.3-cp312-cp312-manylinux2014_s390x.manylinux_2_17_s390x.manylinux_2_28_s390x.whl", hash = "sha256:5fdec68f91a0c6739b380c83b951e2c72ac0197ace422360e6d5a959d8d97b2c", size = 844011, upload-time = "2025-09-25T21:32:15.21Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pyyaml/6.0.3/pyyaml-6.0.3-cp312-cp312-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:ba1cc08a7ccde2d2ec775841541641e4548226580ab850948cbfda66a1befcdc", size = 807870, upload-time = "2025-09-25T21:32:16.431Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pyyaml/6.0.3/pyyaml-6.0.3-cp312-cp312-musllinux_1_2_aarch64.whl", hash = "sha256:8dc52c23056b9ddd46818a57b78404882310fb473d63f17b07d5c40421e47f8e", size = 761089, upload-time = "2025-09-25T21:32:17.56Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pyyaml/6.0.3/pyyaml-6.0.3-cp312-cp312-musllinux_1_2_x86_64.whl", hash = "sha256:41715c910c881bc081f1e8872880d3c650acf13dfa8214bad49ed4cede7c34ea", size = 790181, upload-time = "2025-09-25T21:32:18.834Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pyyaml/6.0.3/pyyaml-6.0.3-cp312-cp312-win32.whl", hash = "sha256:96b533f0e99f6579b3d4d4995707cf36df9100d67e0c8303a0c55b27b5f99bc5", size = 137658, upload-time = "2025-09-25T21:32:20.209Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pyyaml/6.0.3/pyyaml-6.0.3-cp312-cp312-win_amd64.whl", hash = "sha256:5fcd34e47f6e0b794d17de1b4ff496c00986e1c83f7ab2fb8fcfe9616ff7477b", size = 154003, upload-time = "2025-09-25T21:32:21.167Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/pyyaml/6.0.3/pyyaml-6.0.3-cp312-cp312-win_arm64.whl", hash = "sha256:64386e5e707d03a7e172c0701abfb7e10f0fb753ee1d773128192742712a98fd", size = 140344, upload-time = "2025-09-25T21:32:22.617Z" }, ] [[package]] name = "regex" version = "2026.7.19" -source = { registry = "https://pypi.org/simple" } -sdist = { url = "https://files.pythonhosted.org/packages/20/98/04b13f1ddfb63158025291c02e03eb42fbb7acb51d091d541050eb4e35e8/regex-2026.7.19.tar.gz", hash = "sha256:7e77b324909c1617cbb4c668677e2c6ae13f44d7c1de0d4f15f2e3c10f3315b5", size = 416440, upload-time = "2026-07-19T00:19:48.923Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/05/e5/cef4de2bac939280b68d32adc659478845238a8274f2f79c465063f590ad/regex-2026.7.19-cp311-cp311-macosx_10_9_universal2.whl", hash = "sha256:ac777001cdfc28b72477d93c8564bb7583081ea8fb45cdca3d568e0a4f87183c", size = 494012, upload-time = "2026-07-19T00:16:39.927Z" }, - { url = "https://files.pythonhosted.org/packages/ff/87/e86f51eb117457bb7803132ffe5cb6e2841e2b5bea4cc85d397f3c6e257d/regex-2026.7.19-cp311-cp311-macosx_10_9_x86_64.whl", hash = "sha256:59787bd5f8c70aa339084e961d2996b53fbdeab4d5393bba5c1fe1fc32e02bae", size = 295281, upload-time = "2026-07-19T00:16:41.433Z" }, - { url = "https://files.pythonhosted.org/packages/41/2e/2360c41d8080a3d9ec7e5c90fad6eab3b50192869d10e9a5609e48c8177b/regex-2026.7.19-cp311-cp311-macosx_11_0_arm64.whl", hash = "sha256:90c633e7e8d6bf4e992b8b36ce69e018f834b641dd6de8cea6d78c06ffa119c5", size = 290615, upload-time = "2026-07-19T00:16:43.058Z" }, - { url = "https://files.pythonhosted.org/packages/cf/69/b65ba4344efbc771b28fe5dde84cbbb6c8f9551165952fe78def5b9dde6a/regex-2026.7.19-cp311-cp311-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:87ccab0db8d5f4fbb0272642113c1adb2ffc698c16d3a0944580222331fa7a20", size = 791804, upload-time = "2026-07-19T00:16:44.662Z" }, - { url = "https://files.pythonhosted.org/packages/81/b6/a40dfa0dc6224b36f620c00296eacc830489cbf8c2837b6750dfe6170375/regex-2026.7.19-cp311-cp311-manylinux2014_ppc64le.manylinux_2_17_ppc64le.manylinux_2_28_ppc64le.whl", hash = "sha256:9e50d748a32da622f256e8d505867f5d3c43a837c6a9f0efb149655fadd1042a", size = 861723, upload-time = "2026-07-19T00:16:46.412Z" }, - { url = "https://files.pythonhosted.org/packages/e3/02/735991dee71abd83196a7962f7ed8bf5aa05720ff06e2d3ff896a85e2bbb/regex-2026.7.19-cp311-cp311-manylinux2014_s390x.manylinux_2_17_s390x.manylinux_2_28_s390x.whl", hash = "sha256:bf1516fe58fc104f39b2d1dbe2d5e27d0cd45c4be2e42ba6ee0cc763701ec3c7", size = 905932, upload-time = "2026-07-19T00:16:47.956Z" }, - { url = "https://files.pythonhosted.org/packages/45/6c/e7098d8b846ccdbf431d8c081b61e496526a27a28094ed09e0dce21b3f54/regex-2026.7.19-cp311-cp311-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:09f3e5287f94f17b709dc9a9e70865855feee835c861613be144218ce4ca82cc", size = 801407, upload-time = "2026-07-19T00:16:49.43Z" }, - { url = "https://files.pythonhosted.org/packages/8a/18/34b69274e2649bcc7d9b089c2b2983fb2632d8ecf667e359593be9072e79/regex-2026.7.19-cp311-cp311-manylinux_2_31_riscv64.manylinux_2_39_riscv64.whl", hash = "sha256:6383cd2ed53a646c659ba1fe65727db76437fdaa069e697a0b44a51d5843d864", size = 774448, upload-time = "2026-07-19T00:16:51.352Z" }, - { url = "https://files.pythonhosted.org/packages/bb/e6/0a72247d025585fd3800b98e040b84d562a88af6303347100484849f4f01/regex-2026.7.19-cp311-cp311-musllinux_1_2_aarch64.whl", hash = "sha256:09d3007fc76249a83cdd33de160d50e6cb77f54e09d8fa9e7148e10607ce24af", size = 783297, upload-time = "2026-07-19T00:16:53.071Z" }, - { url = "https://files.pythonhosted.org/packages/b1/aa/c4f65ae7dd02a36b323a70c4cff326e1f3442361aaebc9311100a130d54f/regex-2026.7.19-cp311-cp311-musllinux_1_2_ppc64le.whl", hash = "sha256:6f8c6e7a1cfa3dc9d0ee2de0e65e834537fa29992cc3976ffec914afc35c5dd5", size = 854736, upload-time = "2026-07-19T00:16:54.607Z" }, - { url = "https://files.pythonhosted.org/packages/62/c3/668082bcc817b9e694189b84997aeba7385b7779faa6711788679c482e35/regex-2026.7.19-cp311-cp311-musllinux_1_2_riscv64.whl", hash = "sha256:b2ea4a3e8357be8849e833beeae757ac3c7a6b3fc055c03c808a53c91ad30d82", size = 763298, upload-time = "2026-07-19T00:16:56.289Z" }, - { url = "https://files.pythonhosted.org/packages/4b/fb/2d07ad555e7af88aa5f867fdafa47a8d945ee237c20af3ebceb46a820835/regex-2026.7.19-cp311-cp311-musllinux_1_2_s390x.whl", hash = "sha256:80115dd39481fd3a4b4080220799dbcacb921a844de4b827264ececacbe17c78", size = 844430, upload-time = "2026-07-19T00:16:57.933Z" }, - { url = "https://files.pythonhosted.org/packages/51/15/c82a471fe3dce56f03745635b43aa456c40dc0db089e07ef148b331507d1/regex-2026.7.19-cp311-cp311-musllinux_1_2_x86_64.whl", hash = "sha256:d6ce43a0269d68cee79a7d1ade7def53c20f8f2a047b92d7b5d5bcc73ae88327", size = 789683, upload-time = "2026-07-19T00:16:59.583Z" }, - { url = "https://files.pythonhosted.org/packages/b5/f4/7532a2c59d56f5398902c20de60f0c9a5d1cd364e42a051b48e1b210be7b/regex-2026.7.19-cp311-cp311-win32.whl", hash = "sha256:9be2a6647740dd3cca6acb24e87f03d7632cd280dbce9bbe40c26353a215a45d", size = 266778, upload-time = "2026-07-19T00:17:01.032Z" }, - { url = "https://files.pythonhosted.org/packages/83/2b/cf1bc631db154eb95520d9d5dbc2371ff77a0f014bbf7d748fed8496aa63/regex-2026.7.19-cp311-cp311-win_amd64.whl", hash = "sha256:8d3469c91dd92ee41b7c95280edbd975ef1ba9195086686623a1c6e8935ce965", size = 277983, upload-time = "2026-07-19T00:17:02.571Z" }, - { url = "https://files.pythonhosted.org/packages/8d/bd/56ceaf170e875d5a6761bf2bfd0d040f1cacc896850d5e40cb29b11bbd06/regex-2026.7.19-cp311-cp311-win_arm64.whl", hash = "sha256:36aacfb15faaff3ced55afbf35ec72f50d4aee22082c4f7fe0573a33e2fca92e", size = 276961, upload-time = "2026-07-19T00:17:04.135Z" }, - { url = "https://files.pythonhosted.org/packages/3b/b9/d11d7e501ac8fd7d617684423ebb9561e0b998481c1e4cbc0cb212c5d74a/regex-2026.7.19-cp312-cp312-macosx_10_13_universal2.whl", hash = "sha256:2cc3460cedf7579948486eab03bc9ad7089df4d7281c0f47f4afe03e8d13f02d", size = 496778, upload-time = "2026-07-19T00:17:05.677Z" }, - { url = "https://files.pythonhosted.org/packages/3f/a9/a5ab6f312f24318019170dc485d5421fe4f89e43a98640da50d95a8a7041/regex-2026.7.19-cp312-cp312-macosx_10_13_x86_64.whl", hash = "sha256:0e9554c8785eac5cffe6300f69a91f58ba72bc88a5f8d661235ad7c6aa5b8ccd", size = 297122, upload-time = "2026-07-19T00:17:07.59Z" }, - { url = "https://files.pythonhosted.org/packages/b3/63/4cab4d7f2d384a144d420b763d97674cb70619c878ea6fcd7640d0e62143/regex-2026.7.19-cp312-cp312-macosx_11_0_arm64.whl", hash = "sha256:d7da47a0f248977f08e2cb659ff3c17ddc13a4d39b3a7baa0a81bf5b415430f6", size = 292009, upload-time = "2026-07-19T00:17:09.648Z" }, - { url = "https://files.pythonhosted.org/packages/22/85/102a81b218298957d4ea7d2f084fae537a71add9d6ff93c8e67284c5f45e/regex-2026.7.19-cp312-cp312-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:93db40c8de0815baab96a06e08a984bac71f989d13bab789e382158c5d426797", size = 796708, upload-time = "2026-07-19T00:17:11.542Z" }, - { url = "https://files.pythonhosted.org/packages/78/b5/dc136af5629938a037cd2b304c12240e132ec92f38be8ff9cc89af2a1f2d/regex-2026.7.19-cp312-cp312-manylinux2014_ppc64le.manylinux_2_17_ppc64le.manylinux_2_28_ppc64le.whl", hash = "sha256:66bd62c59a5427746e8c44becae1d9b99d22fb13f30f492083dfb9ad7c45cc18", size = 865651, upload-time = "2026-07-19T00:17:13.312Z" }, - { url = "https://files.pythonhosted.org/packages/e0/75/67402ae3cd9c8c988a4c805d15ee3eef015e7ca4cb112cf3e640fc1f4153/regex-2026.7.19-cp312-cp312-manylinux2014_s390x.manylinux_2_17_s390x.manylinux_2_28_s390x.whl", hash = "sha256:1649eb39fcc9ea80c4d2f110fde2b8ab2aef3877b98f02ab9b14e961f418c511", size = 911756, upload-time = "2026-07-19T00:17:15.015Z" }, - { url = "https://files.pythonhosted.org/packages/2a/8e/096d00c7c480ef2ff4265349b14e2261d4ab787ba1f74e2e80d1c58079c3/regex-2026.7.19-cp312-cp312-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:9dce8ec9695f531a1b8a6f314fd4b393adcccf2ea861db480cdf97a301d01a68", size = 801798, upload-time = "2026-07-19T00:17:17.208Z" }, - { url = "https://files.pythonhosted.org/packages/f0/41/e7ecac6edb5722417f85cc67eaf386322fbe8acf6918ec2fdc37c20dd9d0/regex-2026.7.19-cp312-cp312-manylinux_2_31_riscv64.manylinux_2_39_riscv64.whl", hash = "sha256:3080a7fd38ef049bd489e01c970c97dd84ff446a885b0f1f6b26d9b1ad13ce11", size = 776933, upload-time = "2026-07-19T00:17:19.347Z" }, - { url = "https://files.pythonhosted.org/packages/6f/69/03c9b3f058d66403e0ca2c938696e81d51cd4c6d47ec5265f02f96948d9a/regex-2026.7.19-cp312-cp312-musllinux_1_2_aarch64.whl", hash = "sha256:1d793a7988e04fcb1e2e135567443d82173225d657419ec09414a9b5a145b986", size = 784338, upload-time = "2026-07-19T00:17:21.057Z" }, - { url = "https://files.pythonhosted.org/packages/f6/f7/b38ab3d43f284afbb618fcd15d0e77eb786ae461ce1f6bc7494619ddc0f2/regex-2026.7.19-cp312-cp312-musllinux_1_2_ppc64le.whl", hash = "sha256:e8b0abe7d870f53ca5143895fef7d1041a0c831a140d3dc2c760dd7ba25d4a8b", size = 860452, upload-time = "2026-07-19T00:17:23.119Z" }, - { url = "https://files.pythonhosted.org/packages/15/5c/ff60ef0571121714f3cf9920bc183071e384a10b556d042e0fdb06cc07a5/regex-2026.7.19-cp312-cp312-musllinux_1_2_riscv64.whl", hash = "sha256:4e5413bd5f13d3a4e3539ca98f70f75e7fca92518dd7f117f030ebedd10b60cb", size = 765958, upload-time = "2026-07-19T00:17:24.81Z" }, - { url = "https://files.pythonhosted.org/packages/aa/0f/bd34021162c0ab47f9a315bd56cd5642e920c8e5668a75ef6c6a6fca590d/regex-2026.7.19-cp312-cp312-musllinux_1_2_s390x.whl", hash = "sha256:73b133a9e6fb512858e7f065e96f1180aa46646bc74a83aea62f1d314f3dd035", size = 851765, upload-time = "2026-07-19T00:17:26.993Z" }, - { url = "https://files.pythonhosted.org/packages/2a/20/a2ca43edade0595cccfdc98636739f536d9e26898e7dbddc2b9e98898953/regex-2026.7.19-cp312-cp312-musllinux_1_2_x86_64.whl", hash = "sha256:dbe6493fbd27321b1d1f2dd4f5c7e5bd4d8b1d7cab7f32fd67db3d0b2ed8248a", size = 789714, upload-time = "2026-07-19T00:17:28.699Z" }, - { url = "https://files.pythonhosted.org/packages/5d/47/e02db4015d424fc83c00ea0ac8c5e5ec14397943de9abf909d5ce3a25931/regex-2026.7.19-cp312-cp312-win32.whl", hash = "sha256:ddd67571c10869f65a5d7dde536d1e066e306cc90de57d7de4d5f34802428bb5", size = 267157, upload-time = "2026-07-19T00:17:31.051Z" }, - { url = "https://files.pythonhosted.org/packages/08/8e/c780c131f79b42ed22d1bd7da4096c2c35f813e835acd02ef0f018bd892c/regex-2026.7.19-cp312-cp312-win_amd64.whl", hash = "sha256:e30d40268a28d54ce0437031750497004c22602b8e3ab891f759b795a003b312", size = 277777, upload-time = "2026-07-19T00:17:32.848Z" }, - { url = "https://files.pythonhosted.org/packages/3e/4c/e4d7e086449bdf379d89774bf1f89dc4a41943f3c5a6125a03905b34b5fb/regex-2026.7.19-cp312-cp312-win_arm64.whl", hash = "sha256:de9208bb427130c82a5dbfd104f92c8876fc9559278c880b3002755bbbe9c83d", size = 277136, upload-time = "2026-07-19T00:17:34.803Z" }, +source = { registry = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/" } +sdist = { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/regex/2026.7.19/regex-2026.7.19.tar.gz", hash = "sha256:7e77b324909c1617cbb4c668677e2c6ae13f44d7c1de0d4f15f2e3c10f3315b5", size = 416440, upload-time = "2026-07-19T00:19:48.923Z" } +wheels = [ + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/regex/2026.7.19/regex-2026.7.19-cp311-cp311-macosx_10_9_universal2.whl", hash = "sha256:ac777001cdfc28b72477d93c8564bb7583081ea8fb45cdca3d568e0a4f87183c", size = 494012, upload-time = "2026-07-19T00:16:39.927Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/regex/2026.7.19/regex-2026.7.19-cp311-cp311-macosx_10_9_x86_64.whl", hash = "sha256:59787bd5f8c70aa339084e961d2996b53fbdeab4d5393bba5c1fe1fc32e02bae", size = 295281, upload-time = "2026-07-19T00:16:41.433Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/regex/2026.7.19/regex-2026.7.19-cp311-cp311-macosx_11_0_arm64.whl", hash = "sha256:90c633e7e8d6bf4e992b8b36ce69e018f834b641dd6de8cea6d78c06ffa119c5", size = 290615, upload-time = "2026-07-19T00:16:43.058Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/regex/2026.7.19/regex-2026.7.19-cp311-cp311-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:87ccab0db8d5f4fbb0272642113c1adb2ffc698c16d3a0944580222331fa7a20", size = 791804, upload-time = "2026-07-19T00:16:44.662Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/regex/2026.7.19/regex-2026.7.19-cp311-cp311-manylinux2014_ppc64le.manylinux_2_17_ppc64le.manylinux_2_28_ppc64le.whl", hash = "sha256:9e50d748a32da622f256e8d505867f5d3c43a837c6a9f0efb149655fadd1042a", size = 861723, upload-time = "2026-07-19T00:16:46.412Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/regex/2026.7.19/regex-2026.7.19-cp311-cp311-manylinux2014_s390x.manylinux_2_17_s390x.manylinux_2_28_s390x.whl", hash = "sha256:bf1516fe58fc104f39b2d1dbe2d5e27d0cd45c4be2e42ba6ee0cc763701ec3c7", size = 905932, upload-time = "2026-07-19T00:16:47.956Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/regex/2026.7.19/regex-2026.7.19-cp311-cp311-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:09f3e5287f94f17b709dc9a9e70865855feee835c861613be144218ce4ca82cc", size = 801407, upload-time = "2026-07-19T00:16:49.43Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/regex/2026.7.19/regex-2026.7.19-cp311-cp311-manylinux_2_31_riscv64.manylinux_2_39_riscv64.whl", hash = "sha256:6383cd2ed53a646c659ba1fe65727db76437fdaa069e697a0b44a51d5843d864", size = 774448, upload-time = "2026-07-19T00:16:51.352Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/regex/2026.7.19/regex-2026.7.19-cp311-cp311-musllinux_1_2_aarch64.whl", hash = "sha256:09d3007fc76249a83cdd33de160d50e6cb77f54e09d8fa9e7148e10607ce24af", size = 783297, upload-time = "2026-07-19T00:16:53.071Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/regex/2026.7.19/regex-2026.7.19-cp311-cp311-musllinux_1_2_ppc64le.whl", hash = "sha256:6f8c6e7a1cfa3dc9d0ee2de0e65e834537fa29992cc3976ffec914afc35c5dd5", size = 854736, upload-time = "2026-07-19T00:16:54.607Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/regex/2026.7.19/regex-2026.7.19-cp311-cp311-musllinux_1_2_riscv64.whl", hash = "sha256:b2ea4a3e8357be8849e833beeae757ac3c7a6b3fc055c03c808a53c91ad30d82", size = 763298, upload-time = "2026-07-19T00:16:56.289Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/regex/2026.7.19/regex-2026.7.19-cp311-cp311-musllinux_1_2_s390x.whl", hash = "sha256:80115dd39481fd3a4b4080220799dbcacb921a844de4b827264ececacbe17c78", size = 844430, upload-time = "2026-07-19T00:16:57.933Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/regex/2026.7.19/regex-2026.7.19-cp311-cp311-musllinux_1_2_x86_64.whl", hash = "sha256:d6ce43a0269d68cee79a7d1ade7def53c20f8f2a047b92d7b5d5bcc73ae88327", size = 789683, upload-time = "2026-07-19T00:16:59.583Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/regex/2026.7.19/regex-2026.7.19-cp311-cp311-win32.whl", hash = "sha256:9be2a6647740dd3cca6acb24e87f03d7632cd280dbce9bbe40c26353a215a45d", size = 266778, upload-time = "2026-07-19T00:17:01.032Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/regex/2026.7.19/regex-2026.7.19-cp311-cp311-win_amd64.whl", hash = "sha256:8d3469c91dd92ee41b7c95280edbd975ef1ba9195086686623a1c6e8935ce965", size = 277983, upload-time = "2026-07-19T00:17:02.571Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/regex/2026.7.19/regex-2026.7.19-cp311-cp311-win_arm64.whl", hash = "sha256:36aacfb15faaff3ced55afbf35ec72f50d4aee22082c4f7fe0573a33e2fca92e", size = 276961, upload-time = "2026-07-19T00:17:04.135Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/regex/2026.7.19/regex-2026.7.19-cp312-cp312-macosx_10_13_universal2.whl", hash = "sha256:2cc3460cedf7579948486eab03bc9ad7089df4d7281c0f47f4afe03e8d13f02d", size = 496778, upload-time = "2026-07-19T00:17:05.677Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/regex/2026.7.19/regex-2026.7.19-cp312-cp312-macosx_10_13_x86_64.whl", hash = "sha256:0e9554c8785eac5cffe6300f69a91f58ba72bc88a5f8d661235ad7c6aa5b8ccd", size = 297122, upload-time = "2026-07-19T00:17:07.59Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/regex/2026.7.19/regex-2026.7.19-cp312-cp312-macosx_11_0_arm64.whl", hash = "sha256:d7da47a0f248977f08e2cb659ff3c17ddc13a4d39b3a7baa0a81bf5b415430f6", size = 292009, upload-time = "2026-07-19T00:17:09.648Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/regex/2026.7.19/regex-2026.7.19-cp312-cp312-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:93db40c8de0815baab96a06e08a984bac71f989d13bab789e382158c5d426797", size = 796708, upload-time = "2026-07-19T00:17:11.542Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/regex/2026.7.19/regex-2026.7.19-cp312-cp312-manylinux2014_ppc64le.manylinux_2_17_ppc64le.manylinux_2_28_ppc64le.whl", hash = "sha256:66bd62c59a5427746e8c44becae1d9b99d22fb13f30f492083dfb9ad7c45cc18", size = 865651, upload-time = "2026-07-19T00:17:13.312Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/regex/2026.7.19/regex-2026.7.19-cp312-cp312-manylinux2014_s390x.manylinux_2_17_s390x.manylinux_2_28_s390x.whl", hash = "sha256:1649eb39fcc9ea80c4d2f110fde2b8ab2aef3877b98f02ab9b14e961f418c511", size = 911756, upload-time = "2026-07-19T00:17:15.015Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/regex/2026.7.19/regex-2026.7.19-cp312-cp312-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:9dce8ec9695f531a1b8a6f314fd4b393adcccf2ea861db480cdf97a301d01a68", size = 801798, upload-time = "2026-07-19T00:17:17.208Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/regex/2026.7.19/regex-2026.7.19-cp312-cp312-manylinux_2_31_riscv64.manylinux_2_39_riscv64.whl", hash = "sha256:3080a7fd38ef049bd489e01c970c97dd84ff446a885b0f1f6b26d9b1ad13ce11", size = 776933, upload-time = "2026-07-19T00:17:19.347Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/regex/2026.7.19/regex-2026.7.19-cp312-cp312-musllinux_1_2_aarch64.whl", hash = "sha256:1d793a7988e04fcb1e2e135567443d82173225d657419ec09414a9b5a145b986", size = 784338, upload-time = "2026-07-19T00:17:21.057Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/regex/2026.7.19/regex-2026.7.19-cp312-cp312-musllinux_1_2_ppc64le.whl", hash = "sha256:e8b0abe7d870f53ca5143895fef7d1041a0c831a140d3dc2c760dd7ba25d4a8b", size = 860452, upload-time = "2026-07-19T00:17:23.119Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/regex/2026.7.19/regex-2026.7.19-cp312-cp312-musllinux_1_2_riscv64.whl", hash = "sha256:4e5413bd5f13d3a4e3539ca98f70f75e7fca92518dd7f117f030ebedd10b60cb", size = 765958, upload-time = "2026-07-19T00:17:24.81Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/regex/2026.7.19/regex-2026.7.19-cp312-cp312-musllinux_1_2_s390x.whl", hash = "sha256:73b133a9e6fb512858e7f065e96f1180aa46646bc74a83aea62f1d314f3dd035", size = 851765, upload-time = "2026-07-19T00:17:26.993Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/regex/2026.7.19/regex-2026.7.19-cp312-cp312-musllinux_1_2_x86_64.whl", hash = "sha256:dbe6493fbd27321b1d1f2dd4f5c7e5bd4d8b1d7cab7f32fd67db3d0b2ed8248a", size = 789714, upload-time = "2026-07-19T00:17:28.699Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/regex/2026.7.19/regex-2026.7.19-cp312-cp312-win32.whl", hash = "sha256:ddd67571c10869f65a5d7dde536d1e066e306cc90de57d7de4d5f34802428bb5", size = 267157, upload-time = "2026-07-19T00:17:31.051Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/regex/2026.7.19/regex-2026.7.19-cp312-cp312-win_amd64.whl", hash = "sha256:e30d40268a28d54ce0437031750497004c22602b8e3ab891f759b795a003b312", size = 277777, upload-time = "2026-07-19T00:17:32.848Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/regex/2026.7.19/regex-2026.7.19-cp312-cp312-win_arm64.whl", hash = "sha256:de9208bb427130c82a5dbfd104f92c8876fc9559278c880b3002755bbbe9c83d", size = 277136, upload-time = "2026-07-19T00:17:34.803Z" }, ] [[package]] name = "requests" version = "2.34.2" -source = { registry = "https://pypi.org/simple" } +source = { registry = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/" } dependencies = [ { name = "certifi" }, { name = "charset-normalizer" }, { name = "idna" }, { name = "urllib3" }, ] -sdist = { url = "https://files.pythonhosted.org/packages/ac/c3/e2a2b89f2d3e2179abd6d00ebd70bff6273f37fb3e0cc209f48b39d00cbf/requests-2.34.2.tar.gz", hash = "sha256:f288924cae4e29463698d6d60bc6a4da69c89185ad1e0bcc4104f584e960b9ed", size = 142856, upload-time = "2026-05-14T19:25:27.735Z" } +sdist = { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/requests/2.34.2/requests-2.34.2.tar.gz", hash = "sha256:f288924cae4e29463698d6d60bc6a4da69c89185ad1e0bcc4104f584e960b9ed", size = 142856, upload-time = "2026-05-14T19:25:27.735Z" } wheels = [ - { url = "https://files.pythonhosted.org/packages/a0/f4/c67b0b3f1b9245e8d266f0f112c500d50e5b4e83cb6f3b71b6528104182a/requests-2.34.2-py3-none-any.whl", hash = "sha256:2a0d60c172f83ac6ab31e4554906c0f3b3588d37b5cb939b1c061f4907e278e0", size = 73075, upload-time = "2026-05-14T19:25:26.443Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/requests/2.34.2/requests-2.34.2-py3-none-any.whl", hash = "sha256:2a0d60c172f83ac6ab31e4554906c0f3b3588d37b5cb939b1c061f4907e278e0", size = 73075, upload-time = "2026-05-14T19:25:26.443Z" }, ] [[package]] name = "rich" version = "15.0.0" -source = { registry = "https://pypi.org/simple" } +source = { registry = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/" } dependencies = [ { name = "markdown-it-py" }, { name = "pygments" }, ] -sdist = { url = "https://files.pythonhosted.org/packages/c0/8f/0722ca900cc807c13a6a0c696dacf35430f72e0ec571c4275d2371fca3e9/rich-15.0.0.tar.gz", hash = "sha256:edd07a4824c6b40189fb7ac9bc4c52536e9780fbbfbddf6f1e2502c31b068c36", size = 230680, upload-time = "2026-04-12T08:24:00.75Z" } +sdist = { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/rich/15.0.0/rich-15.0.0.tar.gz", hash = "sha256:edd07a4824c6b40189fb7ac9bc4c52536e9780fbbfbddf6f1e2502c31b068c36", size = 230680, upload-time = "2026-04-12T08:24:00.75Z" } wheels = [ - { url = "https://files.pythonhosted.org/packages/82/3b/64d4899d73f91ba49a8c18a8ff3f0ea8f1c1d75481760df8c68ef5235bf5/rich-15.0.0-py3-none-any.whl", hash = "sha256:33bd4ef74232fb73fe9279a257718407f169c09b78a87ad3d296f548e27de0bb", size = 310654, upload-time = "2026-04-12T08:24:02.83Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/rich/15.0.0/rich-15.0.0-py3-none-any.whl", hash = "sha256:33bd4ef74232fb73fe9279a257718407f169c09b78a87ad3d296f548e27de0bb", size = 310654, upload-time = "2026-04-12T08:24:02.83Z" }, ] [[package]] name = "safetensors" version = "0.8.0" -source = { registry = "https://pypi.org/simple" } -sdist = { url = "https://files.pythonhosted.org/packages/45/06/f955dbbb1859e3bd23c8ac6141af5106e7ad5fedec4a3a6e3d60f94b7001/safetensors-0.8.0.tar.gz", hash = "sha256:fabaf3e0f18a6618d9b36560682562157f77c2b71fcffc7b432be2baed9d753d", size = 325846, upload-time = "2026-06-09T07:52:25.563Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/39/a0/f718cda65b05407d228f97602cf60dca269c979867aa5beb25410de26cd3/safetensors-0.8.0-cp310-abi3-macosx_10_12_x86_64.whl", hash = "sha256:c554f85858e05226d3c2828e32395e677434685d6d94594a41643361c5e837f0", size = 473568, upload-time = "2026-06-09T07:52:18.829Z" }, - { url = "https://files.pythonhosted.org/packages/f5/b1/fa7c600e7dceae12e9606c7578cbc9ff1e1ed55844883ee5c92205e86226/safetensors-0.8.0-cp310-abi3-macosx_11_0_arm64.whl", hash = "sha256:c80201d22cbf405b80647a60ada77bba06c8fba2da2743ba1e89cdcc39a81f25", size = 484562, upload-time = "2026-06-09T07:52:17.518Z" }, - { url = "https://files.pythonhosted.org/packages/09/7d/65a7de0af421317bb36a067241e4235fff194eed60b961ed6d3f59a3fc60/safetensors-0.8.0-cp310-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl", hash = "sha256:7a46e5ff292c356d6991e60942ba7f79817682d3a2cef0702136448cb9c4d235", size = 502844, upload-time = "2026-06-09T07:52:07.624Z" }, - { url = "https://files.pythonhosted.org/packages/91/4f/3175c9d75634e0e0dda0082794193521035edd7c70a6f212bf33ca06ddf4/safetensors-0.8.0-cp310-abi3-manylinux_2_17_armv7l.manylinux2014_armv7l.whl", hash = "sha256:4124502b78f03534117c848f87a39b8f31e577b15eff423bf8bfb95f2a8c30d0", size = 511823, upload-time = "2026-06-09T07:52:09.565Z" }, - { url = "https://files.pythonhosted.org/packages/20/87/846c289e7aa2299eff406335717cf43ce8777194ece8aad75772e0411615/safetensors-0.8.0-cp310-abi3-manylinux_2_17_ppc64le.manylinux2014_ppc64le.whl", hash = "sha256:7bc0a787ba8a35be368ee3574edfa2b1ad389eebd0a72e482ae275490e3f6c98", size = 633461, upload-time = "2026-06-09T07:52:11.128Z" }, - { url = "https://files.pythonhosted.org/packages/76/22/8d64d9df2c45d5ded401df889d0ad90882804ca172d79ec4f0df8f727fe0/safetensors-0.8.0-cp310-abi3-manylinux_2_17_s390x.manylinux2014_s390x.whl", hash = "sha256:040070828e36dc8e122178bbbd5830ff9e97920affb84cbe0f46442497bed358", size = 545148, upload-time = "2026-06-09T07:52:13.603Z" }, - { url = "https://files.pythonhosted.org/packages/28/50/f203ff3a3ddfe19308efc83c5a3a29ed02bf786732ec35e68bf9162f3365/safetensors-0.8.0-cp310-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl", hash = "sha256:fd6f3f93c9a0a7cc2788ee63fb763353d4bd2e89b0751bc78fcf7dda00bea774", size = 516040, upload-time = "2026-06-09T07:52:16.29Z" }, - { url = "https://files.pythonhosted.org/packages/46/fb/cdaed17ceb2948784fd9c36b6fd3e951b608547cea81a48e8ee6f8cfdfcb/safetensors-0.8.0-cp310-abi3-manylinux_2_31_riscv64.whl", hash = "sha256:fcdd41ec4628fee5799f807c73c353629130fbd942aa23d83c623dd6c9d52d78", size = 513832, upload-time = "2026-06-09T07:52:12.37Z" }, - { url = "https://files.pythonhosted.org/packages/0d/49/1e15de264dcc3b77943d2d0c56a95809956883b1c2d6d585c792523f180b/safetensors-0.8.0-cp310-abi3-manylinux_2_5_i686.manylinux1_i686.whl", hash = "sha256:8e9f537aa183a38ace122d27303dcd986b26bd2a7591f9181d7f0c396f4677ca", size = 559930, upload-time = "2026-06-09T07:52:14.743Z" }, - { url = "https://files.pythonhosted.org/packages/2a/43/bf38443278eab4b1be1fce2931e2b012ad9cb7df52ada751d0aab8f7659a/safetensors-0.8.0-cp310-abi3-musllinux_1_2_aarch64.whl", hash = "sha256:87eec7ffed2b809f05a398a8becb7d013f19f7837cd15d9748580d6cf30dbaf4", size = 678670, upload-time = "2026-06-09T07:52:20.032Z" }, - { url = "https://files.pythonhosted.org/packages/72/e3/68cd3fa5b48488e84add63e04cb12f3bc28ae4638c06d4508c6e88823d0e/safetensors-0.8.0-cp310-abi3-musllinux_1_2_armv7l.whl", hash = "sha256:4a95ae2b05d7726d751da4ebf626a2ca782b706e101bd894c95bc2450b1cffcc", size = 786679, upload-time = "2026-06-09T07:52:21.322Z" }, - { url = "https://files.pythonhosted.org/packages/29/4b/1c19c509d56e01f4fbb3d0a2e597450f6cc04d1d56cf52defb0a62dfd715/safetensors-0.8.0-cp310-abi3-musllinux_1_2_i686.whl", hash = "sha256:3ae091f16662658bdc019a4ff6cb4c085bb7d725eb5978b183ffd265863b6d2d", size = 765683, upload-time = "2026-06-09T07:52:22.594Z" }, - { url = "https://files.pythonhosted.org/packages/27/43/41c1621732edd934d868a00d1b891584c892a7b62a9aab82ea5a0a5623ee/safetensors-0.8.0-cp310-abi3-musllinux_1_2_x86_64.whl", hash = "sha256:8e080062fcde23be189565e1c3305d16751a218ecf9412c8601e64204eb6f846", size = 722361, upload-time = "2026-06-09T07:52:23.924Z" }, - { url = "https://files.pythonhosted.org/packages/8e/3f/73ccf82579412b4a71c4ca673f10b5f1f888d7cf5af7fe24f27d30307be4/safetensors-0.8.0-cp310-abi3-win32.whl", hash = "sha256:2ddf52eac562eda224f99acfa7889d02968c1fd59a5b011ae7d8137c37e9c02d", size = 342401, upload-time = "2026-06-09T07:52:28.895Z" }, - { url = "https://files.pythonhosted.org/packages/1b/6d/3fba214c1e5e0f69991677ec3bc17023f0421776975e1de0c682dca475e2/safetensors-0.8.0-cp310-abi3-win_amd64.whl", hash = "sha256:096ec1a98435df7beb08853bb5aa9081a84f23d0adc67ed1a0a10550f608373f", size = 355540, upload-time = "2026-06-09T07:52:27.832Z" }, - { url = "https://files.pythonhosted.org/packages/8d/fc/7eedc3510d97878876e32774eebbeb61c43f148a96e915c84229a3e967aa/safetensors-0.8.0-cp310-abi3-win_arm64.whl", hash = "sha256:f7838e5135a406ad3e02efdcb8cf2e5397d368b0154537c4fec682dbc544d452", size = 340500, upload-time = "2026-06-09T07:52:26.745Z" }, +source = { registry = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/" } +sdist = { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/safetensors/0.8.0/safetensors-0.8.0.tar.gz", hash = "sha256:fabaf3e0f18a6618d9b36560682562157f77c2b71fcffc7b432be2baed9d753d", size = 325846, upload-time = "2026-06-09T07:52:25.563Z" } +wheels = [ + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/safetensors/0.8.0/safetensors-0.8.0-cp310-abi3-macosx_10_12_x86_64.whl", hash = "sha256:c554f85858e05226d3c2828e32395e677434685d6d94594a41643361c5e837f0", size = 473568, upload-time = "2026-06-09T07:52:18.829Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/safetensors/0.8.0/safetensors-0.8.0-cp310-abi3-macosx_11_0_arm64.whl", hash = "sha256:c80201d22cbf405b80647a60ada77bba06c8fba2da2743ba1e89cdcc39a81f25", size = 484562, upload-time = "2026-06-09T07:52:17.518Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/safetensors/0.8.0/safetensors-0.8.0-cp310-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl", hash = "sha256:7a46e5ff292c356d6991e60942ba7f79817682d3a2cef0702136448cb9c4d235", size = 502844, upload-time = "2026-06-09T07:52:07.624Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/safetensors/0.8.0/safetensors-0.8.0-cp310-abi3-manylinux_2_17_armv7l.manylinux2014_armv7l.whl", hash = "sha256:4124502b78f03534117c848f87a39b8f31e577b15eff423bf8bfb95f2a8c30d0", size = 511823, upload-time = "2026-06-09T07:52:09.565Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/safetensors/0.8.0/safetensors-0.8.0-cp310-abi3-manylinux_2_17_ppc64le.manylinux2014_ppc64le.whl", hash = "sha256:7bc0a787ba8a35be368ee3574edfa2b1ad389eebd0a72e482ae275490e3f6c98", size = 633461, upload-time = "2026-06-09T07:52:11.128Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/safetensors/0.8.0/safetensors-0.8.0-cp310-abi3-manylinux_2_17_s390x.manylinux2014_s390x.whl", hash = "sha256:040070828e36dc8e122178bbbd5830ff9e97920affb84cbe0f46442497bed358", size = 545148, upload-time = "2026-06-09T07:52:13.603Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/safetensors/0.8.0/safetensors-0.8.0-cp310-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl", hash = "sha256:fd6f3f93c9a0a7cc2788ee63fb763353d4bd2e89b0751bc78fcf7dda00bea774", size = 516040, upload-time = "2026-06-09T07:52:16.29Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/safetensors/0.8.0/safetensors-0.8.0-cp310-abi3-manylinux_2_31_riscv64.whl", hash = "sha256:fcdd41ec4628fee5799f807c73c353629130fbd942aa23d83c623dd6c9d52d78", size = 513832, upload-time = "2026-06-09T07:52:12.37Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/safetensors/0.8.0/safetensors-0.8.0-cp310-abi3-manylinux_2_5_i686.manylinux1_i686.whl", hash = "sha256:8e9f537aa183a38ace122d27303dcd986b26bd2a7591f9181d7f0c396f4677ca", size = 559930, upload-time = "2026-06-09T07:52:14.743Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/safetensors/0.8.0/safetensors-0.8.0-cp310-abi3-musllinux_1_2_aarch64.whl", hash = "sha256:87eec7ffed2b809f05a398a8becb7d013f19f7837cd15d9748580d6cf30dbaf4", size = 678670, upload-time = "2026-06-09T07:52:20.032Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/safetensors/0.8.0/safetensors-0.8.0-cp310-abi3-musllinux_1_2_armv7l.whl", hash = "sha256:4a95ae2b05d7726d751da4ebf626a2ca782b706e101bd894c95bc2450b1cffcc", size = 786679, upload-time = "2026-06-09T07:52:21.322Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/safetensors/0.8.0/safetensors-0.8.0-cp310-abi3-musllinux_1_2_i686.whl", hash = "sha256:3ae091f16662658bdc019a4ff6cb4c085bb7d725eb5978b183ffd265863b6d2d", size = 765683, upload-time = "2026-06-09T07:52:22.594Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/safetensors/0.8.0/safetensors-0.8.0-cp310-abi3-musllinux_1_2_x86_64.whl", hash = "sha256:8e080062fcde23be189565e1c3305d16751a218ecf9412c8601e64204eb6f846", size = 722361, upload-time = "2026-06-09T07:52:23.924Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/safetensors/0.8.0/safetensors-0.8.0-cp310-abi3-win32.whl", hash = "sha256:2ddf52eac562eda224f99acfa7889d02968c1fd59a5b011ae7d8137c37e9c02d", size = 342401, upload-time = "2026-06-09T07:52:28.895Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/safetensors/0.8.0/safetensors-0.8.0-cp310-abi3-win_amd64.whl", hash = "sha256:096ec1a98435df7beb08853bb5aa9081a84f23d0adc67ed1a0a10550f608373f", size = 355540, upload-time = "2026-06-09T07:52:27.832Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/safetensors/0.8.0/safetensors-0.8.0-cp310-abi3-win_arm64.whl", hash = "sha256:f7838e5135a406ad3e02efdcb8cf2e5397d368b0154537c4fec682dbc544d452", size = 340500, upload-time = "2026-06-09T07:52:26.745Z" }, ] [[package]] name = "shellingham" version = "1.5.4" -source = { registry = "https://pypi.org/simple" } -sdist = { url = "https://files.pythonhosted.org/packages/58/15/8b3609fd3830ef7b27b655beb4b4e9c62313a4e8da8c676e142cc210d58e/shellingham-1.5.4.tar.gz", hash = "sha256:8dbca0739d487e5bd35ab3ca4b36e11c4078f3a234bfce294b0a0291363404de", size = 10310, upload-time = "2023-10-24T04:13:40.426Z" } +source = { registry = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/" } +sdist = { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/shellingham/1.5.4/shellingham-1.5.4.tar.gz", hash = "sha256:8dbca0739d487e5bd35ab3ca4b36e11c4078f3a234bfce294b0a0291363404de", size = 10310, upload-time = "2023-10-24T04:13:40.426Z" } wheels = [ - { url = "https://files.pythonhosted.org/packages/e0/f9/0595336914c5619e5f28a1fb793285925a8cd4b432c9da0a987836c7f822/shellingham-1.5.4-py2.py3-none-any.whl", hash = "sha256:7ecfff8f2fd72616f7481040475a65b2bf8af90a56c89140852d1120324e8686", size = 9755, upload-time = "2023-10-24T04:13:38.866Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/shellingham/1.5.4/shellingham-1.5.4-py2.py3-none-any.whl", hash = "sha256:7ecfff8f2fd72616f7481040475a65b2bf8af90a56c89140852d1120324e8686", size = 9755, upload-time = "2023-10-24T04:13:38.866Z" }, ] [[package]] name = "sniffio" version = "1.3.1" -source = { registry = "https://pypi.org/simple" } -sdist = { url = "https://files.pythonhosted.org/packages/a2/87/a6771e1546d97e7e041b6ae58d80074f81b7d5121207425c964ddf5cfdbd/sniffio-1.3.1.tar.gz", hash = "sha256:f4324edc670a0f49750a81b895f35c3adb843cca46f0530f79fc1babb23789dc", size = 20372, upload-time = "2024-02-25T23:20:04.057Z" } +source = { registry = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/" } +sdist = { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/sniffio/1.3.1/sniffio-1.3.1.tar.gz", hash = "sha256:f4324edc670a0f49750a81b895f35c3adb843cca46f0530f79fc1babb23789dc", size = 20372, upload-time = "2024-02-25T23:20:04.057Z" } wheels = [ - { url = "https://files.pythonhosted.org/packages/e9/44/75a9c9421471a6c4805dbf2356f7c181a29c1879239abab1ea2cc8f38b40/sniffio-1.3.1-py3-none-any.whl", hash = "sha256:2f6da418d1f1e0fddd844478f41680e794e6051915791a034ff65e5f100525a2", size = 10235, upload-time = "2024-02-25T23:20:01.196Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/sniffio/1.3.1/sniffio-1.3.1-py3-none-any.whl", hash = "sha256:2f6da418d1f1e0fddd844478f41680e794e6051915791a034ff65e5f100525a2", size = 10235, upload-time = "2024-02-25T23:20:01.196Z" }, ] [[package]] name = "starlette" version = "1.6.0" -source = { registry = "https://pypi.org/simple" } +source = { registry = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/" } dependencies = [ { name = "anyio" }, { name = "typing-extensions" }, ] -sdist = { url = "https://files.pythonhosted.org/packages/b5/b4/205b0d5241d934e8add0c38aa924c4f9fb7330834ff11e5444db964ec3f9/starlette-1.6.0.tar.gz", hash = "sha256:d4e3ac5e546444960c710297a3c9fc3f7ebae1b7e963f3d36173b49da535be9b", size = 2716969, upload-time = "2026-08-08T18:27:57.512Z" } +sdist = { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/starlette/1.6.0/starlette-1.6.0.tar.gz", hash = "sha256:d4e3ac5e546444960c710297a3c9fc3f7ebae1b7e963f3d36173b49da535be9b", size = 2716969, upload-time = "2026-08-08T18:27:57.512Z" } wheels = [ - { url = "https://files.pythonhosted.org/packages/c8/cb/6a6a47d5b464bd08695d254f3da6e7986cc70c9fa5d778eda57538edfe56/starlette-1.6.0-py3-none-any.whl", hash = "sha256:a86dd39d14bb45f85a3d18525215a9ef0cfd1f192ac793220e72598c90335f0c", size = 75969, upload-time = "2026-08-08T18:27:56.196Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/starlette/1.6.0/starlette-1.6.0-py3-none-any.whl", hash = "sha256:a86dd39d14bb45f85a3d18525215a9ef0cfd1f192ac793220e72598c90335f0c", size = 75969, upload-time = "2026-08-08T18:27:56.196Z" }, ] [[package]] @@ -1140,26 +1140,26 @@ source = { git = "https://github.com/modal-projects/stitch.git?rev=375a9396a7b05 [[package]] name = "synchronicity" version = "0.12.5" -source = { registry = "https://pypi.org/simple" } +source = { registry = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/" } dependencies = [ { name = "typing-extensions" }, ] -sdist = { url = "https://files.pythonhosted.org/packages/5d/1c/f51dc54bbd302991026a53f9790735540e0e9e1184e9d5939f02446aa5bc/synchronicity-0.12.5.tar.gz", hash = "sha256:94d96b1d85698e3056b96a793b8c0949af6584e4a7d877fabdeb5385efe230aa", size = 60745, upload-time = "2026-06-18T21:06:23.545Z" } +sdist = { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/synchronicity/0.12.5/synchronicity-0.12.5.tar.gz", hash = "sha256:94d96b1d85698e3056b96a793b8c0949af6584e4a7d877fabdeb5385efe230aa", size = 60745, upload-time = "2026-06-18T21:06:23.545Z" } wheels = [ - { url = "https://files.pythonhosted.org/packages/06/74/ad9b99520f70c0bc3318e582e359d360cfc0f7afd7bf368a7f24013cece7/synchronicity-0.12.5-py3-none-any.whl", hash = "sha256:fdbbb10d437bc08a6b0f814fc66fddd1b58ffed314533d42f1ab555801e781af", size = 41107, upload-time = "2026-06-18T21:06:22.505Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/synchronicity/0.12.5/synchronicity-0.12.5-py3-none-any.whl", hash = "sha256:fdbbb10d437bc08a6b0f814fc66fddd1b58ffed314533d42f1ab555801e781af", size = 41107, upload-time = "2026-06-18T21:06:22.505Z" }, ] [[package]] name = "tinker" version = "0.24.1" -source = { registry = "https://pypi.org/simple" } +source = { registry = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/" } dependencies = [ { name = "anyio" }, { name = "click" }, { name = "distro" }, { name = "httpx", extra = ["http2"] }, - { name = "numpy", version = "2.4.6", source = { registry = "https://pypi.org/simple" }, marker = "python_full_version < '3.12'" }, - { name = "numpy", version = "2.5.2", source = { registry = "https://pypi.org/simple" }, marker = "python_full_version >= '3.12'" }, + { name = "numpy", version = "2.4.6", source = { registry = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/" }, marker = "python_full_version < '3.12'" }, + { name = "numpy", version = "2.5.2", source = { registry = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/" }, marker = "python_full_version >= '3.12'" }, { name = "orjson" }, { name = "protobuf" }, { name = "pydantic" }, @@ -1170,66 +1170,66 @@ dependencies = [ { name = "typing-extensions" }, { name = "zstandard" }, ] -sdist = { url = "https://files.pythonhosted.org/packages/29/45/b1a30a92b17b237e3661eeab3812546f11fd1ff58c3b05915b07cb09fbc7/tinker-0.24.1.tar.gz", hash = "sha256:1cc8e052155b624d13c6e3ae5e17b3c1b91c36f31722e5d75b38bdd3621a11c8", size = 260945, upload-time = "2026-08-06T21:07:18.945Z" } +sdist = { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/tinker/0.24.1/tinker-0.24.1.tar.gz", hash = "sha256:1cc8e052155b624d13c6e3ae5e17b3c1b91c36f31722e5d75b38bdd3621a11c8", size = 260945, upload-time = "2026-08-06T21:07:18.945Z" } wheels = [ - { url = "https://files.pythonhosted.org/packages/bc/64/6bc22116fd1f7afd57153ce88849db711c85d1244585eca8abe5e6b772f5/tinker-0.24.1-py3-none-any.whl", hash = "sha256:67a6564b2563f970e36f870292eabbbe698b539e84c04ee96cec0ac6af2206a5", size = 250296, upload-time = "2026-08-06T21:07:17.512Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/tinker/0.24.1/tinker-0.24.1-py3-none-any.whl", hash = "sha256:67a6564b2563f970e36f870292eabbbe698b539e84c04ee96cec0ac6af2206a5", size = 250296, upload-time = "2026-08-06T21:07:17.512Z" }, ] [[package]] name = "tokenizers" version = "0.22.2" -source = { registry = "https://pypi.org/simple" } +source = { registry = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/" } dependencies = [ { name = "huggingface-hub" }, ] -sdist = { url = "https://files.pythonhosted.org/packages/73/6f/f80cfef4a312e1fb34baf7d85c72d4411afde10978d4657f8cdd811d3ccc/tokenizers-0.22.2.tar.gz", hash = "sha256:473b83b915e547aa366d1eee11806deaf419e17be16310ac0a14077f1e28f917", size = 372115, upload-time = "2026-01-05T10:45:15.988Z" } +sdist = { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/tokenizers/0.22.2/tokenizers-0.22.2.tar.gz", hash = "sha256:473b83b915e547aa366d1eee11806deaf419e17be16310ac0a14077f1e28f917", size = 372115, upload-time = "2026-01-05T10:45:15.988Z" } wheels = [ - { url = "https://files.pythonhosted.org/packages/92/97/5dbfabf04c7e348e655e907ed27913e03db0923abb5dfdd120d7b25630e1/tokenizers-0.22.2-cp39-abi3-macosx_10_12_x86_64.whl", hash = "sha256:544dd704ae7238755d790de45ba8da072e9af3eea688f698b137915ae959281c", size = 3100275, upload-time = "2026-01-05T10:41:02.158Z" }, - { url = "https://files.pythonhosted.org/packages/2e/47/174dca0502ef88b28f1c9e06b73ce33500eedfac7a7692108aec220464e7/tokenizers-0.22.2-cp39-abi3-macosx_11_0_arm64.whl", hash = "sha256:1e418a55456beedca4621dbab65a318981467a2b188e982a23e117f115ce5001", size = 2981472, upload-time = "2026-01-05T10:41:00.276Z" }, - { url = "https://files.pythonhosted.org/packages/d6/84/7990e799f1309a8b87af6b948f31edaa12a3ed22d11b352eaf4f4b2e5753/tokenizers-0.22.2-cp39-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl", hash = "sha256:2249487018adec45d6e3554c71d46eb39fa8ea67156c640f7513eb26f318cec7", size = 3290736, upload-time = "2026-01-05T10:40:32.165Z" }, - { url = "https://files.pythonhosted.org/packages/78/59/09d0d9ba94dcd5f4f1368d4858d24546b4bdc0231c2354aa31d6199f0399/tokenizers-0.22.2-cp39-abi3-manylinux_2_17_armv7l.manylinux2014_armv7l.whl", hash = "sha256:25b85325d0815e86e0bac263506dd114578953b7b53d7de09a6485e4a160a7dd", size = 3168835, upload-time = "2026-01-05T10:40:38.847Z" }, - { url = "https://files.pythonhosted.org/packages/47/50/b3ebb4243e7160bda8d34b731e54dd8ab8b133e50775872e7a434e524c28/tokenizers-0.22.2-cp39-abi3-manylinux_2_17_i686.manylinux2014_i686.whl", hash = "sha256:bfb88f22a209ff7b40a576d5324bf8286b519d7358663db21d6246fb17eea2d5", size = 3521673, upload-time = "2026-01-05T10:40:56.614Z" }, - { url = "https://files.pythonhosted.org/packages/e0/fa/89f4cb9e08df770b57adb96f8cbb7e22695a4cb6c2bd5f0c4f0ebcf33b66/tokenizers-0.22.2-cp39-abi3-manylinux_2_17_ppc64le.manylinux2014_ppc64le.whl", hash = "sha256:1c774b1276f71e1ef716e5486f21e76333464f47bece56bbd554485982a9e03e", size = 3724818, upload-time = "2026-01-05T10:40:44.507Z" }, - { url = "https://files.pythonhosted.org/packages/64/04/ca2363f0bfbe3b3d36e95bf67e56a4c88c8e3362b658e616d1ac185d47f2/tokenizers-0.22.2-cp39-abi3-manylinux_2_17_s390x.manylinux2014_s390x.whl", hash = "sha256:df6c4265b289083bf710dff49bc51ef252f9d5be33a45ee2bed151114a56207b", size = 3379195, upload-time = "2026-01-05T10:40:51.139Z" }, - { url = "https://files.pythonhosted.org/packages/2e/76/932be4b50ef6ccedf9d3c6639b056a967a86258c6d9200643f01269211ca/tokenizers-0.22.2-cp39-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl", hash = "sha256:369cc9fc8cc10cb24143873a0d95438bb8ee257bb80c71989e3ee290e8d72c67", size = 3274982, upload-time = "2026-01-05T10:40:58.331Z" }, - { url = "https://files.pythonhosted.org/packages/1d/28/5f9f5a4cc211b69e89420980e483831bcc29dade307955cc9dc858a40f01/tokenizers-0.22.2-cp39-abi3-musllinux_1_2_aarch64.whl", hash = "sha256:29c30b83d8dcd061078b05ae0cb94d3c710555fbb44861139f9f83dcca3dc3e4", size = 9478245, upload-time = "2026-01-05T10:41:04.053Z" }, - { url = "https://files.pythonhosted.org/packages/6c/fb/66e2da4704d6aadebf8cb39f1d6d1957df667ab24cff2326b77cda0dcb85/tokenizers-0.22.2-cp39-abi3-musllinux_1_2_armv7l.whl", hash = "sha256:37ae80a28c1d3265bb1f22464c856bd23c02a05bb211e56d0c5301a435be6c1a", size = 9560069, upload-time = "2026-01-05T10:45:10.673Z" }, - { url = "https://files.pythonhosted.org/packages/16/04/fed398b05caa87ce9b1a1bb5166645e38196081b225059a6edaff6440fac/tokenizers-0.22.2-cp39-abi3-musllinux_1_2_i686.whl", hash = "sha256:791135ee325f2336f498590eb2f11dc5c295232f288e75c99a36c5dbce63088a", size = 9899263, upload-time = "2026-01-05T10:45:12.559Z" }, - { url = "https://files.pythonhosted.org/packages/05/a1/d62dfe7376beaaf1394917e0f8e93ee5f67fea8fcf4107501db35996586b/tokenizers-0.22.2-cp39-abi3-musllinux_1_2_x86_64.whl", hash = "sha256:38337540fbbddff8e999d59970f3c6f35a82de10053206a7562f1ea02d046fa5", size = 10033429, upload-time = "2026-01-05T10:45:14.333Z" }, - { url = "https://files.pythonhosted.org/packages/fd/18/a545c4ea42af3df6effd7d13d250ba77a0a86fb20393143bbb9a92e434d4/tokenizers-0.22.2-cp39-abi3-win32.whl", hash = "sha256:a6bf3f88c554a2b653af81f3204491c818ae2ac6fbc09e76ef4773351292bc92", size = 2502363, upload-time = "2026-01-05T10:45:20.593Z" }, - { url = "https://files.pythonhosted.org/packages/65/71/0670843133a43d43070abeb1949abfdef12a86d490bea9cd9e18e37c5ff7/tokenizers-0.22.2-cp39-abi3-win_amd64.whl", hash = "sha256:c9ea31edff2968b44a88f97d784c2f16dc0729b8b143ed004699ebca91f05c48", size = 2747786, upload-time = "2026-01-05T10:45:18.411Z" }, - { url = "https://files.pythonhosted.org/packages/72/f4/0de46cfa12cdcbcd464cc59fde36912af405696f687e53a091fb432f694c/tokenizers-0.22.2-cp39-abi3-win_arm64.whl", hash = "sha256:9ce725d22864a1e965217204946f830c37876eee3b2ba6fc6255e8e903d5fcbc", size = 2612133, upload-time = "2026-01-05T10:45:17.232Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/tokenizers/0.22.2/tokenizers-0.22.2-cp39-abi3-macosx_10_12_x86_64.whl", hash = "sha256:544dd704ae7238755d790de45ba8da072e9af3eea688f698b137915ae959281c", size = 3100275, upload-time = "2026-01-05T10:41:02.158Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/tokenizers/0.22.2/tokenizers-0.22.2-cp39-abi3-macosx_11_0_arm64.whl", hash = "sha256:1e418a55456beedca4621dbab65a318981467a2b188e982a23e117f115ce5001", size = 2981472, upload-time = "2026-01-05T10:41:00.276Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/tokenizers/0.22.2/tokenizers-0.22.2-cp39-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl", hash = "sha256:2249487018adec45d6e3554c71d46eb39fa8ea67156c640f7513eb26f318cec7", size = 3290736, upload-time = "2026-01-05T10:40:32.165Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/tokenizers/0.22.2/tokenizers-0.22.2-cp39-abi3-manylinux_2_17_armv7l.manylinux2014_armv7l.whl", hash = "sha256:25b85325d0815e86e0bac263506dd114578953b7b53d7de09a6485e4a160a7dd", size = 3168835, upload-time = "2026-01-05T10:40:38.847Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/tokenizers/0.22.2/tokenizers-0.22.2-cp39-abi3-manylinux_2_17_i686.manylinux2014_i686.whl", hash = "sha256:bfb88f22a209ff7b40a576d5324bf8286b519d7358663db21d6246fb17eea2d5", size = 3521673, upload-time = "2026-01-05T10:40:56.614Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/tokenizers/0.22.2/tokenizers-0.22.2-cp39-abi3-manylinux_2_17_ppc64le.manylinux2014_ppc64le.whl", hash = "sha256:1c774b1276f71e1ef716e5486f21e76333464f47bece56bbd554485982a9e03e", size = 3724818, upload-time = "2026-01-05T10:40:44.507Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/tokenizers/0.22.2/tokenizers-0.22.2-cp39-abi3-manylinux_2_17_s390x.manylinux2014_s390x.whl", hash = "sha256:df6c4265b289083bf710dff49bc51ef252f9d5be33a45ee2bed151114a56207b", size = 3379195, upload-time = "2026-01-05T10:40:51.139Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/tokenizers/0.22.2/tokenizers-0.22.2-cp39-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl", hash = "sha256:369cc9fc8cc10cb24143873a0d95438bb8ee257bb80c71989e3ee290e8d72c67", size = 3274982, upload-time = "2026-01-05T10:40:58.331Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/tokenizers/0.22.2/tokenizers-0.22.2-cp39-abi3-musllinux_1_2_aarch64.whl", hash = "sha256:29c30b83d8dcd061078b05ae0cb94d3c710555fbb44861139f9f83dcca3dc3e4", size = 9478245, upload-time = "2026-01-05T10:41:04.053Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/tokenizers/0.22.2/tokenizers-0.22.2-cp39-abi3-musllinux_1_2_armv7l.whl", hash = "sha256:37ae80a28c1d3265bb1f22464c856bd23c02a05bb211e56d0c5301a435be6c1a", size = 9560069, upload-time = "2026-01-05T10:45:10.673Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/tokenizers/0.22.2/tokenizers-0.22.2-cp39-abi3-musllinux_1_2_i686.whl", hash = "sha256:791135ee325f2336f498590eb2f11dc5c295232f288e75c99a36c5dbce63088a", size = 9899263, upload-time = "2026-01-05T10:45:12.559Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/tokenizers/0.22.2/tokenizers-0.22.2-cp39-abi3-musllinux_1_2_x86_64.whl", hash = "sha256:38337540fbbddff8e999d59970f3c6f35a82de10053206a7562f1ea02d046fa5", size = 10033429, upload-time = "2026-01-05T10:45:14.333Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/tokenizers/0.22.2/tokenizers-0.22.2-cp39-abi3-win32.whl", hash = "sha256:a6bf3f88c554a2b653af81f3204491c818ae2ac6fbc09e76ef4773351292bc92", size = 2502363, upload-time = "2026-01-05T10:45:20.593Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/tokenizers/0.22.2/tokenizers-0.22.2-cp39-abi3-win_amd64.whl", hash = "sha256:c9ea31edff2968b44a88f97d784c2f16dc0729b8b143ed004699ebca91f05c48", size = 2747786, upload-time = "2026-01-05T10:45:18.411Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/tokenizers/0.22.2/tokenizers-0.22.2-cp39-abi3-win_arm64.whl", hash = "sha256:9ce725d22864a1e965217204946f830c37876eee3b2ba6fc6255e8e903d5fcbc", size = 2612133, upload-time = "2026-01-05T10:45:17.232Z" }, ] [[package]] name = "toml" version = "0.10.2" -source = { registry = "https://pypi.org/simple" } -sdist = { url = "https://files.pythonhosted.org/packages/be/ba/1f744cdc819428fc6b5084ec34d9b30660f6f9daaf70eead706e3203ec3c/toml-0.10.2.tar.gz", hash = "sha256:b3bda1d108d5dd99f4a20d24d9c348e91c4db7ab1b749200bded2f839ccbe68f", size = 22253, upload-time = "2020-11-01T01:40:22.204Z" } +source = { registry = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/" } +sdist = { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/toml/0.10.2/toml-0.10.2.tar.gz", hash = "sha256:b3bda1d108d5dd99f4a20d24d9c348e91c4db7ab1b749200bded2f839ccbe68f", size = 22253, upload-time = "2020-11-01T01:40:22.204Z" } wheels = [ - { url = "https://files.pythonhosted.org/packages/44/6f/7120676b6d73228c96e17f1f794d8ab046fc910d781c8d151120c3f1569e/toml-0.10.2-py2.py3-none-any.whl", hash = "sha256:806143ae5bfb6a3c6e736a764057db0e6a0e05e338b5630894a5f779cabb4f9b", size = 16588, upload-time = "2020-11-01T01:40:20.672Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/toml/0.10.2/toml-0.10.2-py2.py3-none-any.whl", hash = "sha256:806143ae5bfb6a3c6e736a764057db0e6a0e05e338b5630894a5f779cabb4f9b", size = 16588, upload-time = "2020-11-01T01:40:20.672Z" }, ] [[package]] name = "tqdm" version = "4.70.0" -source = { registry = "https://pypi.org/simple" } +source = { registry = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/" } dependencies = [ { name = "colorama", marker = "sys_platform == 'win32'" }, ] -sdist = { url = "https://files.pythonhosted.org/packages/21/3b/6c24bec5be5e743ffd99576daa5cc077722fc7d5bbc00bd133fa0c698dc6/tqdm-4.70.0.tar.gz", hash = "sha256:55b0b0dbd97462d06ebee91e4dac24ed4d4702be82b24f07e6c1d27e08cea220", size = 795438, upload-time = "2026-07-27T11:33:15.271Z" } +sdist = { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/tqdm/4.70.0/tqdm-4.70.0.tar.gz", hash = "sha256:55b0b0dbd97462d06ebee91e4dac24ed4d4702be82b24f07e6c1d27e08cea220", size = 795438, upload-time = "2026-07-27T11:33:15.271Z" } wheels = [ - { url = "https://files.pythonhosted.org/packages/f9/1c/01bfd571a64e7f270e6bab5e33777debe0edc56759233ce84f27dec92d14/tqdm-4.70.0-py3-none-any.whl", hash = "sha256:7f585706bfddbdebf89daac705b2dfcc16890130727d3197ca62c732b4310953", size = 80184, upload-time = "2026-07-27T11:33:13.167Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/tqdm/4.70.0/tqdm-4.70.0-py3-none-any.whl", hash = "sha256:7f585706bfddbdebf89daac705b2dfcc16890130727d3197ca62c732b4310953", size = 80184, upload-time = "2026-07-27T11:33:13.167Z" }, ] [[package]] name = "transformers" version = "5.15.1" -source = { registry = "https://pypi.org/simple" } +source = { registry = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/" } dependencies = [ { name = "huggingface-hub" }, - { name = "numpy", version = "2.4.6", source = { registry = "https://pypi.org/simple" }, marker = "python_full_version < '3.12'" }, - { name = "numpy", version = "2.5.2", source = { registry = "https://pypi.org/simple" }, marker = "python_full_version >= '3.12'" }, + { name = "numpy", version = "2.4.6", source = { registry = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/" }, marker = "python_full_version < '3.12'" }, + { name = "numpy", version = "2.5.2", source = { registry = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/" }, marker = "python_full_version >= '3.12'" }, { name = "packaging" }, { name = "pyyaml" }, { name = "regex" }, @@ -1238,274 +1238,274 @@ dependencies = [ { name = "tqdm" }, { name = "typer" }, ] -sdist = { url = "https://files.pythonhosted.org/packages/2a/92/c50c61da7046bbb59a4d011291aeadcfb4d7980ab36fdb31e93823a3fb93/transformers-5.15.1.tar.gz", hash = "sha256:27c996bd9075ddc82d40f8590dfdc81ea45f611bfca477e0db5d7fd257a482f7", size = 9378434, upload-time = "2026-08-19T11:28:20.33Z" } +sdist = { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/transformers/5.15.1/transformers-5.15.1.tar.gz", hash = "sha256:27c996bd9075ddc82d40f8590dfdc81ea45f611bfca477e0db5d7fd257a482f7", size = 9378434, upload-time = "2026-08-19T11:28:20.33Z" } wheels = [ - { url = "https://files.pythonhosted.org/packages/41/c4/a12e1d9b387fb0c40a57116db82b457e8c771cb419163cda29204d74a595/transformers-5.15.1-py3-none-any.whl", hash = "sha256:b7cdf238ff583e3a58dbc7fa34da1aaf091ce063141f65a30538160bd5afe93f", size = 11749582, upload-time = "2026-08-19T11:28:16.726Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/transformers/5.15.1/transformers-5.15.1-py3-none-any.whl", hash = "sha256:b7cdf238ff583e3a58dbc7fa34da1aaf091ce063141f65a30538160bd5afe93f", size = 11749582, upload-time = "2026-08-19T11:28:16.726Z" }, ] [[package]] name = "typer" version = "0.27.1" -source = { registry = "https://pypi.org/simple" } +source = { registry = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/" } dependencies = [ { name = "annotated-doc" }, { name = "colorama", marker = "sys_platform == 'win32'" }, { name = "rich" }, { name = "shellingham" }, ] -sdist = { url = "https://files.pythonhosted.org/packages/ae/40/4a3db7990d1f62a53182aa96eaef57aeb2886a27f90a195bc66713565d31/typer-0.27.1.tar.gz", hash = "sha256:a79bef8469a79c45498e7b814ecf8d603cc7644e9acbd9e19cac0334240b18df", size = 203994, upload-time = "2026-08-03T14:41:03.438Z" } +sdist = { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/typer/0.27.1/typer-0.27.1.tar.gz", hash = "sha256:a79bef8469a79c45498e7b814ecf8d603cc7644e9acbd9e19cac0334240b18df", size = 203994, upload-time = "2026-08-03T14:41:03.438Z" } wheels = [ - { url = "https://files.pythonhosted.org/packages/43/89/9518bc0c3929bee36b3a4a8e3daddd6e03f92f9961c66d4983b837160543/typer-0.27.1-py3-none-any.whl", hash = "sha256:53150287edd11baeb4e4722c8e394fcdf8181c0ae89485cba8d25c778d5edd56", size = 122874, upload-time = "2026-08-03T14:41:04.391Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/typer/0.27.1/typer-0.27.1-py3-none-any.whl", hash = "sha256:53150287edd11baeb4e4722c8e394fcdf8181c0ae89485cba8d25c778d5edd56", size = 122874, upload-time = "2026-08-03T14:41:04.391Z" }, ] [[package]] name = "types-certifi" version = "2021.10.8.3" -source = { registry = "https://pypi.org/simple" } -sdist = { url = "https://files.pythonhosted.org/packages/52/68/943c3aeaf14624712a0357c4a67814dba5cea36d194f5c764dad7959a00c/types-certifi-2021.10.8.3.tar.gz", hash = "sha256:72cf7798d165bc0b76e1c10dd1ea3097c7063c42c21d664523b928e88b554a4f", size = 2095, upload-time = "2022-06-09T15:19:05.244Z" } +source = { registry = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/" } +sdist = { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/types-certifi/2021.10.8.3/types-certifi-2021.10.8.3.tar.gz", hash = "sha256:72cf7798d165bc0b76e1c10dd1ea3097c7063c42c21d664523b928e88b554a4f", size = 2095, upload-time = "2022-06-09T15:19:05.244Z" } wheels = [ - { url = "https://files.pythonhosted.org/packages/b5/63/2463d89481e811f007b0e1cd0a91e52e141b47f9de724d20db7b861dcfec/types_certifi-2021.10.8.3-py3-none-any.whl", hash = "sha256:b2d1e325e69f71f7c78e5943d410e650b4707bb0ef32e4ddf3da37f54176e88a", size = 2136, upload-time = "2022-06-09T15:19:03.127Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/types-certifi/2021.10.8.3/types_certifi-2021.10.8.3-py3-none-any.whl", hash = "sha256:b2d1e325e69f71f7c78e5943d410e650b4707bb0ef32e4ddf3da37f54176e88a", size = 2136, upload-time = "2022-06-09T15:19:03.127Z" }, ] [[package]] name = "types-toml" version = "0.10.8.20260518" -source = { registry = "https://pypi.org/simple" } -sdist = { url = "https://files.pythonhosted.org/packages/4b/11/6ece999e91f2ccb848ab4420f3f4816e78ac0541f739e6864affdaaa5737/types_toml-0.10.8.20260518.tar.gz", hash = "sha256:80e10facd24fdeda9d5c672187d72be3ac284843788d67f5aae59e3e016db6fe", size = 9419, upload-time = "2026-05-18T06:02:16.719Z" } +source = { registry = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/" } +sdist = { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/types-toml/0.10.8.20260518/types_toml-0.10.8.20260518.tar.gz", hash = "sha256:80e10facd24fdeda9d5c672187d72be3ac284843788d67f5aae59e3e016db6fe", size = 9419, upload-time = "2026-05-18T06:02:16.719Z" } wheels = [ - { url = "https://files.pythonhosted.org/packages/91/25/489751806bf5c95e4007f8e17409199c54d31e49ffbea07c5729b1286c8e/types_toml-0.10.8.20260518-py3-none-any.whl", hash = "sha256:0e564ab05f6fde62a315b3b5a9b6624fda569399795d30a37e64705a70459303", size = 9669, upload-time = "2026-05-18T06:02:15.86Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/types-toml/0.10.8.20260518/types_toml-0.10.8.20260518-py3-none-any.whl", hash = "sha256:0e564ab05f6fde62a315b3b5a9b6624fda569399795d30a37e64705a70459303", size = 9669, upload-time = "2026-05-18T06:02:15.86Z" }, ] [[package]] name = "typing-extensions" version = "4.16.0" -source = { registry = "https://pypi.org/simple" } -sdist = { url = "https://files.pythonhosted.org/packages/f6/cc/6253133b5bb138fc3306cebfbda2c520f545d36b5be2c7255cc528bb45d6/typing_extensions-4.16.0.tar.gz", hash = "sha256:dc983d19a509c94dba722ee6abd33940f7c05a89e243c47e907eb4db6f1a43e5", size = 113555, upload-time = "2026-07-02T08:40:05.92Z" } +source = { registry = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/" } +sdist = { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/typing-extensions/4.16.0/typing_extensions-4.16.0.tar.gz", hash = "sha256:dc983d19a509c94dba722ee6abd33940f7c05a89e243c47e907eb4db6f1a43e5", size = 113555, upload-time = "2026-07-02T08:40:05.92Z" } wheels = [ - { url = "https://files.pythonhosted.org/packages/49/d3/b8441a820a491ddfc024b0b0cf0393375b75ea13866d9c66727e54c2fc80/typing_extensions-4.16.0-py3-none-any.whl", hash = "sha256:481caa481374e813c1b176ada14e97f1f67a4539ce9cfeb3f350d78d6370c2e8", size = 45571, upload-time = "2026-07-02T08:40:04.659Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/typing-extensions/4.16.0/typing_extensions-4.16.0-py3-none-any.whl", hash = "sha256:481caa481374e813c1b176ada14e97f1f67a4539ce9cfeb3f350d78d6370c2e8", size = 45571, upload-time = "2026-07-02T08:40:04.659Z" }, ] [[package]] name = "typing-inspection" version = "0.4.4" -source = { registry = "https://pypi.org/simple" } +source = { registry = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/" } dependencies = [ { name = "typing-extensions" }, ] -sdist = { url = "https://files.pythonhosted.org/packages/a3/26/b09b8010994eccc3c09092e6b34058f36a460eea2d4c3e8b910c695975a0/typing_inspection-0.4.4.tar.gz", hash = "sha256:547274fa6b0a561ccf549cc9524b999a578e737d015d8709d021f9d0d13bea47", size = 76928, upload-time = "2026-08-12T12:37:25.997Z" } +sdist = { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/typing-inspection/0.4.4/typing_inspection-0.4.4.tar.gz", hash = "sha256:547274fa6b0a561ccf549cc9524b999a578e737d015d8709d021f9d0d13bea47", size = 76928, upload-time = "2026-08-12T12:37:25.997Z" } wheels = [ - { url = "https://files.pythonhosted.org/packages/67/81/4add07e5172b7ac40d8ed5ff580409a7801a4fe26d529bdd915401dabfbe/typing_inspection-0.4.4-py3-none-any.whl", hash = "sha256:65b8397ba37ccbce054456aaccddfc91e6e3083c92824df348d96ca832f3f147", size = 14750, upload-time = "2026-08-12T12:37:24.648Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/typing-inspection/0.4.4/typing_inspection-0.4.4-py3-none-any.whl", hash = "sha256:65b8397ba37ccbce054456aaccddfc91e6e3083c92824df348d96ca832f3f147", size = 14750, upload-time = "2026-08-12T12:37:24.648Z" }, ] [[package]] name = "urllib3" version = "2.7.0" -source = { registry = "https://pypi.org/simple" } -sdist = { url = "https://files.pythonhosted.org/packages/53/0c/06f8b233b8fd13b9e5ee11424ef85419ba0d8ba0b3138bf360be2ff56953/urllib3-2.7.0.tar.gz", hash = "sha256:231e0ec3b63ceb14667c67be60f2f2c40a518cb38b03af60abc813da26505f4c", size = 433602, upload-time = "2026-05-07T16:13:18.596Z" } +source = { registry = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/" } +sdist = { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/urllib3/2.7.0/urllib3-2.7.0.tar.gz", hash = "sha256:231e0ec3b63ceb14667c67be60f2f2c40a518cb38b03af60abc813da26505f4c", size = 433602, upload-time = "2026-05-07T16:13:18.596Z" } wheels = [ - { url = "https://files.pythonhosted.org/packages/7f/3e/5db95bcf282c52709639744ca2a8b149baccf648e39c8cc87553df9eae0c/urllib3-2.7.0-py3-none-any.whl", hash = "sha256:9fb4c81ebbb1ce9531cce37674bbc6f1360472bc18ca9a553ede278ef7276897", size = 131087, upload-time = "2026-05-07T16:13:17.151Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/urllib3/2.7.0/urllib3-2.7.0-py3-none-any.whl", hash = "sha256:9fb4c81ebbb1ce9531cce37674bbc6f1360472bc18ca9a553ede278ef7276897", size = 131087, upload-time = "2026-05-07T16:13:17.151Z" }, ] [[package]] name = "uvicorn" version = "0.52.4" -source = { registry = "https://pypi.org/simple" } +source = { registry = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/" } dependencies = [ { name = "click" }, { name = "h11" }, ] -sdist = { url = "https://files.pythonhosted.org/packages/f2/0f/3f86e61397dd33bf2ccf28188c40db6a740658aeebbbf6e7dbc101a1f487/uvicorn-0.52.4.tar.gz", hash = "sha256:73acfee47a0b133c5de13d219492d62d8a31e935f4fe6e41a232451a15379f86", size = 100627, upload-time = "2026-08-19T06:27:41.821Z" } +sdist = { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/uvicorn/0.52.4/uvicorn-0.52.4.tar.gz", hash = "sha256:73acfee47a0b133c5de13d219492d62d8a31e935f4fe6e41a232451a15379f86", size = 100627, upload-time = "2026-08-19T06:27:41.821Z" } wheels = [ - { url = "https://files.pythonhosted.org/packages/f1/79/4a20b54ab0491485ccd8c077db2d39187c7f12b3e15485d38a7be37c81b4/uvicorn-0.52.4-py3-none-any.whl", hash = "sha256:f86e41a149d7d05a9969337e3946a9c171c06a5d42680896daaba624aeac8da1", size = 79871, upload-time = "2026-08-19T06:27:40.36Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/uvicorn/0.52.4/uvicorn-0.52.4-py3-none-any.whl", hash = "sha256:f86e41a149d7d05a9969337e3946a9c171c06a5d42680896daaba624aeac8da1", size = 79871, upload-time = "2026-08-19T06:27:40.36Z" }, ] [[package]] name = "watchfiles" version = "1.2.0" -source = { registry = "https://pypi.org/simple" } +source = { registry = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/" } dependencies = [ { name = "anyio" }, ] -sdist = { url = "https://files.pythonhosted.org/packages/cd/41/5e1a4bb12aac5f1493fa1bdc11154eca3b258ca4eba65d39c473fe19d8e9/watchfiles-1.2.0.tar.gz", hash = "sha256:c995fba777f1ea992f090f9236e9284cf7a5d1a0130dd5a3d82c598cacd76838", size = 108252, upload-time = "2026-05-18T04:32:04.251Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/fc/3d/8024c801df84d1587740d0359e7fdd80afeae3d159011f3d5376dd82f18e/watchfiles-1.2.0-cp311-cp311-macosx_10_12_x86_64.whl", hash = "sha256:704fd259e332e01f9b9c178f4bce9e49027e5587cc2600eeeaf8e76e1c846201", size = 400242, upload-time = "2026-05-18T04:31:19.014Z" }, - { url = "https://files.pythonhosted.org/packages/87/5b/f4dfd45323e949984a3a7f9dc31d1cbb049921e7d98253488dda72ccdaa9/watchfiles-1.2.0-cp311-cp311-macosx_11_0_arm64.whl", hash = "sha256:6543cf55d170003296d185c0af981f3e1311564907e1f4e08671fc7693a890a5", size = 394562, upload-time = "2026-05-18T04:30:08.46Z" }, - { url = "https://files.pythonhosted.org/packages/98/d8/19483ef075d601c409bce8bcbb5c0f81a10876fff870400568f08ce484a1/watchfiles-1.2.0-cp311-cp311-manylinux_2_17_aarch64.manylinux2014_aarch64.whl", hash = "sha256:89d8c2394a065ca86f5d2910ff263ae67c127e1376ccc4f9fc35c71db879f80a", size = 456611, upload-time = "2026-05-18T04:30:45.723Z" }, - { url = "https://files.pythonhosted.org/packages/b1/6a/cc81fbe7ee42f2f22e661a6e12def7807e01b14b2f39e0ff83fd373fd307/watchfiles-1.2.0-cp311-cp311-manylinux_2_17_armv7l.manylinux2014_armv7l.whl", hash = "sha256:772b80df316480d894a0e3165fdd19cf77f5d17f9a787f94029465ad0e3529d1", size = 461379, upload-time = "2026-05-18T04:31:29.292Z" }, - { url = "https://files.pythonhosted.org/packages/b1/57/7e669002082c0a0f4fb5113bb70125f7110124b846b0a11bc5ae8e90eac1/watchfiles-1.2.0-cp311-cp311-manylinux_2_17_i686.manylinux2014_i686.whl", hash = "sha256:d158cd89df6053823533e06fb1d73c549133bff5f0396170c0e53d9559340717", size = 493556, upload-time = "2026-05-18T04:30:05.44Z" }, - { url = "https://files.pythonhosted.org/packages/45/7d/f60a2b19807b21fe8281f3a8da4f59eef0d5f96825ac4680ba2d4f2ebf91/watchfiles-1.2.0-cp311-cp311-manylinux_2_17_ppc64le.manylinux2014_ppc64le.whl", hash = "sha256:d516b3283a758e087841aedb8031549fb41ced08f3db10aa6d2bf32dc042525b", size = 575255, upload-time = "2026-05-18T04:30:40.568Z" }, - { url = "https://files.pythonhosted.org/packages/bd/49/77f5b5e6efbcd57482f74948ebb1b97e5c0046d6b61475042d830c84b3ff/watchfiles-1.2.0-cp311-cp311-manylinux_2_17_s390x.manylinux2014_s390x.whl", hash = "sha256:53b2290c92e0506d102cd448fbc610d87079553f86caa39d67440856a8b8bba5", size = 467052, upload-time = "2026-05-18T04:31:17.942Z" }, - { url = "https://files.pythonhosted.org/packages/ee/5a/73e2959af1b97fd5d556f9a8bdba017be23ceeef731869d5eaa0a753d5a3/watchfiles-1.2.0-cp311-cp311-manylinux_2_17_x86_64.manylinux2014_x86_64.whl", hash = "sha256:a711b51aec4370d0dcda5b6c09463206f133a5759341d7744b953a7b62e1100e", size = 456858, upload-time = "2026-05-18T04:30:30.182Z" }, - { url = "https://files.pythonhosted.org/packages/50/57/1bc8c27fad7e6c19bddee15d276dbb6ab72480ec01c127afff1673aee417/watchfiles-1.2.0-cp311-cp311-manylinux_2_31_riscv64.whl", hash = "sha256:e2ca07fa7d89195ec0865d3d285666286740bfa83d83e5cee204043a31ecc165", size = 467579, upload-time = "2026-05-18T04:32:15.897Z" }, - { url = "https://files.pythonhosted.org/packages/09/6c/3c2e44edba3553c5e3c3b8c8a2a6dee6b9e12ae2cf4bd2378bebf9dc3038/watchfiles-1.2.0-cp311-cp311-musllinux_1_1_aarch64.whl", hash = "sha256:e0618518f282c4ebff60f5e5b1247b6d91bb8b9f4476947563a1e74acc66f3c6", size = 633253, upload-time = "2026-05-18T04:31:37.123Z" }, - { url = "https://files.pythonhosted.org/packages/30/c2/d8c84a882ab39bbefcc4915ab3e91830b7a7e990c5570b0b69075aba3faf/watchfiles-1.2.0-cp311-cp311-musllinux_1_1_x86_64.whl", hash = "sha256:0d191c054d0715c3c95c99df9b8dbf6fd096d8c1e021e8f212e1bd8bc444ccb5", size = 660713, upload-time = "2026-05-18T04:31:24.62Z" }, - { url = "https://files.pythonhosted.org/packages/a9/07/f97736a5fc605364fe67b25e9fa4a6965dfd4840d50c406ada507e9d735f/watchfiles-1.2.0-cp311-cp311-win32.whl", hash = "sha256:9342472aff9b093c5acd4f6d8f70ae0937964ab56542502bcf5579782da69ae8", size = 277222, upload-time = "2026-05-18T04:31:21.131Z" }, - { url = "https://files.pythonhosted.org/packages/cf/99/2b04981977fc2608afd60360d928c6aecf6b950292ca221d98f4005f6694/watchfiles-1.2.0-cp311-cp311-win_amd64.whl", hash = "sha256:dbd6c97045dad81227c8d040173da044c1de08de64a5ea8b555da4aee1d5fa22", size = 290274, upload-time = "2026-05-18T04:31:45.966Z" }, - { url = "https://files.pythonhosted.org/packages/3c/74/f7f58a7075ee9cf612b0cfcddb78b8cd8234f0742d6f0075cf0da2dde1c6/watchfiles-1.2.0-cp311-cp311-win_arm64.whl", hash = "sha256:57a2d9fa4fb4c2ecae57b13dfff2c7ab53e21a2ba674fe9f05506680fcdcc0d7", size = 283460, upload-time = "2026-05-18T04:31:39.126Z" }, - { url = "https://files.pythonhosted.org/packages/b8/2f/e42c992d2afda3108ea1c02acecc991b9f31d05c14adc2a7cee9ee211fc4/watchfiles-1.2.0-cp312-cp312-macosx_10_12_x86_64.whl", hash = "sha256:bc13eb17538be00c874699dc0abe4ee2bc8d50bb1166a6b9e175ef3fd7eb8f26", size = 400115, upload-time = "2026-05-18T04:32:02.06Z" }, - { url = "https://files.pythonhosted.org/packages/5f/8f/6af2ea19065c91d8b0ea3516fdfc8c0d349f407e8e9fbf4e5a17360de8ad/watchfiles-1.2.0-cp312-cp312-macosx_11_0_arm64.whl", hash = "sha256:2d95ddc1eb6914154253d239089900813f6a767e174b8e6a50e7fdacb7e4236c", size = 393659, upload-time = "2026-05-18T04:30:50.951Z" }, - { url = "https://files.pythonhosted.org/packages/13/01/b32a967c56fb3e3e5be3db52c3d3b87fa4513aa367d8ed1ad96d42952e5f/watchfiles-1.2.0-cp312-cp312-manylinux_2_17_aarch64.manylinux2014_aarch64.whl", hash = "sha256:8f70d8b291ef6e88d19b1f297a6905ddb978888d9272b0d05e6f53309856bcfc", size = 453207, upload-time = "2026-05-18T04:31:04.231Z" }, - { url = "https://files.pythonhosted.org/packages/04/98/97557a812180338cb1abd32e1cffcc4588f59b5f23e0cb006b2ba95ba64a/watchfiles-1.2.0-cp312-cp312-manylinux_2_17_armv7l.manylinux2014_armv7l.whl", hash = "sha256:56d8641cf834c2836922899105bd3ce3d0dfc69291d52edf0b4d0436829b34c0", size = 459273, upload-time = "2026-05-18T04:31:50.377Z" }, - { url = "https://files.pythonhosted.org/packages/e8/a8/b4b08dcb7653b8087c6586f7ce649505900e866bbcfe40dc9587af02e686/watchfiles-1.2.0-cp312-cp312-manylinux_2_17_i686.manylinux2014_i686.whl", hash = "sha256:2581a94056e55d7d0a31a823ea92bf73749c489ca2285bfdc0fbe6b2bb49d50c", size = 489927, upload-time = "2026-05-18T04:31:42.485Z" }, - { url = "https://files.pythonhosted.org/packages/50/94/3dceea03545d2e5ddfd839f0ddd5e1cecbf1697b5a428d5ba11cef6af95d/watchfiles-1.2.0-cp312-cp312-manylinux_2_17_ppc64le.manylinux2014_ppc64le.whl", hash = "sha256:41bc1199f7523b3f82843c88cbb979180c949caef0342cf90968f178e5d49b01", size = 570476, upload-time = "2026-05-18T04:31:03.071Z" }, - { url = "https://files.pythonhosted.org/packages/cc/f2/d39a5450c3532092b91f81d274360e613c2371bc874a89c7a1a3c5e8d138/watchfiles-1.2.0-cp312-cp312-manylinux_2_17_s390x.manylinux2014_s390x.whl", hash = "sha256:7571e4464cb6e434958f867f7f730b8ab0b75e3f8e5eac0499168486ab3c33a8", size = 465650, upload-time = "2026-05-18T04:30:12.701Z" }, - { url = "https://files.pythonhosted.org/packages/22/24/ed72f68cbc1333ca9b9f2200aa048bb6658ae41709bc1caad4310f4bdffd/watchfiles-1.2.0-cp312-cp312-manylinux_2_17_x86_64.manylinux2014_x86_64.whl", hash = "sha256:e53a384f76b631c3ae5334ce6a52f0baa3a911eb94a4eac7f160079868b716d5", size = 456398, upload-time = "2026-05-18T04:30:13.784Z" }, - { url = "https://files.pythonhosted.org/packages/0d/64/982ef4a4e5bab5b6e5b6becc8cd5e732f6130a78b855f0abec6439a9a135/watchfiles-1.2.0-cp312-cp312-manylinux_2_31_riscv64.whl", hash = "sha256:d20029a60a71a052a24c4db7673bc4de39ab89adbaccbfb5d67987c5d73f424d", size = 465140, upload-time = "2026-05-18T04:31:52.111Z" }, - { url = "https://files.pythonhosted.org/packages/a0/0c/95282abf4ed680b6096010bcfc30c5fa7a041fc5aa5a2ad17a2cc6c75bba/watchfiles-1.2.0-cp312-cp312-musllinux_1_1_aarch64.whl", hash = "sha256:2cb93af48550faf1cea04c303107c8b75833de7013e57ce27d3b8d21d8d0f58c", size = 630259, upload-time = "2026-05-18T04:31:25.676Z" }, - { url = "https://files.pythonhosted.org/packages/30/45/607c1de1530c4bdcf2cf1d1ecc2505ddba5d96bd43ba9f2b0e79876f850f/watchfiles-1.2.0-cp312-cp312-musllinux_1_1_x86_64.whl", hash = "sha256:2995c176de7692b86a2e4c58d9ec718f753150a979cb4a754e2b4ffa38e70906", size = 659859, upload-time = "2026-05-18T04:30:24.333Z" }, - { url = "https://files.pythonhosted.org/packages/fa/08/d9e2e0f9e8e6791d33aefc694ad7eefa7f901f63caff84a81ded38692f9c/watchfiles-1.2.0-cp312-cp312-win32.whl", hash = "sha256:7a2cffd17d27d2ecbb310c2b1d8174f222a5495b1a721894afa88ec11e25b898", size = 275480, upload-time = "2026-05-18T04:30:31.307Z" }, - { url = "https://files.pythonhosted.org/packages/1c/e6/9d42569c0102645cc8cea5d8c7d8a1e9d4ada2cb7f05f75e554b8aa2202a/watchfiles-1.2.0-cp312-cp312-win_amd64.whl", hash = "sha256:f155b3a1b2a5fc89cdc70d47ee5d54e3b75e88efa34982028a35daef9ba00379", size = 288718, upload-time = "2026-05-18T04:32:10.745Z" }, - { url = "https://files.pythonhosted.org/packages/0a/26/88e0dc6ee3898169d7fa22bb6a69cabf2502d2ee25cb8c876d1262d204f8/watchfiles-1.2.0-cp312-cp312-win_arm64.whl", hash = "sha256:8fa585ede612ee9f9e91b18bebf9ba11b9ae29a4e3a0d0cf6fca3e382133f0d5", size = 281026, upload-time = "2026-05-18T04:30:22.23Z" }, - { url = "https://files.pythonhosted.org/packages/23/f4/7513ef1e85fc4c6331b59479d6d72661fc391fbe543678052ac72c8b6c19/watchfiles-1.2.0-pp311-pypy311_pp73-macosx_10_12_x86_64.whl", hash = "sha256:4674d49eb94706dfe666c069fc0a1b646ffcf920473492e209f6d5f60d3f0cc2", size = 403050, upload-time = "2026-05-18T04:30:36.753Z" }, - { url = "https://files.pythonhosted.org/packages/27/0b/a54103cfd732bb703c7a749222011a0483ef3705948dae3b203158601119/watchfiles-1.2.0-pp311-pypy311_pp73-macosx_11_0_arm64.whl", hash = "sha256:094b9b70103d4e963499bdea001ee3c2697b144cd9ae6218a62c0f89ec9e31db", size = 396629, upload-time = "2026-05-18T04:32:03.268Z" }, - { url = "https://files.pythonhosted.org/packages/5e/2c/73f31a3b893886206c3f54d73e8ad8dee58cdb2f69ad2622e0a8a9e07f4e/watchfiles-1.2.0-pp311-pypy311_pp73-manylinux_2_17_aarch64.manylinux2014_aarch64.whl", hash = "sha256:b0ef001f8c25ad0fa9529f914c1600647ecd0f542d11c19b7894768c67b6acb7", size = 457318, upload-time = "2026-05-18T04:31:01.932Z" }, - { url = "https://files.pythonhosted.org/packages/e9/f9/45d021e4a5cc7b9dd567f7cbb06d3b75f751a690063fb6cc7ec60f4e46b7/watchfiles-1.2.0-pp311-pypy311_pp73-manylinux_2_17_x86_64.manylinux2014_x86_64.whl", hash = "sha256:a88fc94e647bc4eec523f1caa540258eb71d14278b9daf72fa1e2658a98df0f0", size = 457771, upload-time = "2026-05-18T04:30:56.331Z" }, +sdist = { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/watchfiles/1.2.0/watchfiles-1.2.0.tar.gz", hash = "sha256:c995fba777f1ea992f090f9236e9284cf7a5d1a0130dd5a3d82c598cacd76838", size = 108252, upload-time = "2026-05-18T04:32:04.251Z" } +wheels = [ + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/watchfiles/1.2.0/watchfiles-1.2.0-cp311-cp311-macosx_10_12_x86_64.whl", hash = "sha256:704fd259e332e01f9b9c178f4bce9e49027e5587cc2600eeeaf8e76e1c846201", size = 400242, upload-time = "2026-05-18T04:31:19.014Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/watchfiles/1.2.0/watchfiles-1.2.0-cp311-cp311-macosx_11_0_arm64.whl", hash = "sha256:6543cf55d170003296d185c0af981f3e1311564907e1f4e08671fc7693a890a5", size = 394562, upload-time = "2026-05-18T04:30:08.46Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/watchfiles/1.2.0/watchfiles-1.2.0-cp311-cp311-manylinux_2_17_aarch64.manylinux2014_aarch64.whl", hash = "sha256:89d8c2394a065ca86f5d2910ff263ae67c127e1376ccc4f9fc35c71db879f80a", size = 456611, upload-time = "2026-05-18T04:30:45.723Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/watchfiles/1.2.0/watchfiles-1.2.0-cp311-cp311-manylinux_2_17_armv7l.manylinux2014_armv7l.whl", hash = "sha256:772b80df316480d894a0e3165fdd19cf77f5d17f9a787f94029465ad0e3529d1", size = 461379, upload-time = "2026-05-18T04:31:29.292Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/watchfiles/1.2.0/watchfiles-1.2.0-cp311-cp311-manylinux_2_17_i686.manylinux2014_i686.whl", hash = "sha256:d158cd89df6053823533e06fb1d73c549133bff5f0396170c0e53d9559340717", size = 493556, upload-time = "2026-05-18T04:30:05.44Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/watchfiles/1.2.0/watchfiles-1.2.0-cp311-cp311-manylinux_2_17_ppc64le.manylinux2014_ppc64le.whl", hash = "sha256:d516b3283a758e087841aedb8031549fb41ced08f3db10aa6d2bf32dc042525b", size = 575255, upload-time = "2026-05-18T04:30:40.568Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/watchfiles/1.2.0/watchfiles-1.2.0-cp311-cp311-manylinux_2_17_s390x.manylinux2014_s390x.whl", hash = "sha256:53b2290c92e0506d102cd448fbc610d87079553f86caa39d67440856a8b8bba5", size = 467052, upload-time = "2026-05-18T04:31:17.942Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/watchfiles/1.2.0/watchfiles-1.2.0-cp311-cp311-manylinux_2_17_x86_64.manylinux2014_x86_64.whl", hash = "sha256:a711b51aec4370d0dcda5b6c09463206f133a5759341d7744b953a7b62e1100e", size = 456858, upload-time = "2026-05-18T04:30:30.182Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/watchfiles/1.2.0/watchfiles-1.2.0-cp311-cp311-manylinux_2_31_riscv64.whl", hash = "sha256:e2ca07fa7d89195ec0865d3d285666286740bfa83d83e5cee204043a31ecc165", size = 467579, upload-time = "2026-05-18T04:32:15.897Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/watchfiles/1.2.0/watchfiles-1.2.0-cp311-cp311-musllinux_1_1_aarch64.whl", hash = "sha256:e0618518f282c4ebff60f5e5b1247b6d91bb8b9f4476947563a1e74acc66f3c6", size = 633253, upload-time = "2026-05-18T04:31:37.123Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/watchfiles/1.2.0/watchfiles-1.2.0-cp311-cp311-musllinux_1_1_x86_64.whl", hash = "sha256:0d191c054d0715c3c95c99df9b8dbf6fd096d8c1e021e8f212e1bd8bc444ccb5", size = 660713, upload-time = "2026-05-18T04:31:24.62Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/watchfiles/1.2.0/watchfiles-1.2.0-cp311-cp311-win32.whl", hash = "sha256:9342472aff9b093c5acd4f6d8f70ae0937964ab56542502bcf5579782da69ae8", size = 277222, upload-time = "2026-05-18T04:31:21.131Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/watchfiles/1.2.0/watchfiles-1.2.0-cp311-cp311-win_amd64.whl", hash = "sha256:dbd6c97045dad81227c8d040173da044c1de08de64a5ea8b555da4aee1d5fa22", size = 290274, upload-time = "2026-05-18T04:31:45.966Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/watchfiles/1.2.0/watchfiles-1.2.0-cp311-cp311-win_arm64.whl", hash = "sha256:57a2d9fa4fb4c2ecae57b13dfff2c7ab53e21a2ba674fe9f05506680fcdcc0d7", size = 283460, upload-time = "2026-05-18T04:31:39.126Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/watchfiles/1.2.0/watchfiles-1.2.0-cp312-cp312-macosx_10_12_x86_64.whl", hash = "sha256:bc13eb17538be00c874699dc0abe4ee2bc8d50bb1166a6b9e175ef3fd7eb8f26", size = 400115, upload-time = "2026-05-18T04:32:02.06Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/watchfiles/1.2.0/watchfiles-1.2.0-cp312-cp312-macosx_11_0_arm64.whl", hash = "sha256:2d95ddc1eb6914154253d239089900813f6a767e174b8e6a50e7fdacb7e4236c", size = 393659, upload-time = "2026-05-18T04:30:50.951Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/watchfiles/1.2.0/watchfiles-1.2.0-cp312-cp312-manylinux_2_17_aarch64.manylinux2014_aarch64.whl", hash = "sha256:8f70d8b291ef6e88d19b1f297a6905ddb978888d9272b0d05e6f53309856bcfc", size = 453207, upload-time = "2026-05-18T04:31:04.231Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/watchfiles/1.2.0/watchfiles-1.2.0-cp312-cp312-manylinux_2_17_armv7l.manylinux2014_armv7l.whl", hash = "sha256:56d8641cf834c2836922899105bd3ce3d0dfc69291d52edf0b4d0436829b34c0", size = 459273, upload-time = "2026-05-18T04:31:50.377Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/watchfiles/1.2.0/watchfiles-1.2.0-cp312-cp312-manylinux_2_17_i686.manylinux2014_i686.whl", hash = "sha256:2581a94056e55d7d0a31a823ea92bf73749c489ca2285bfdc0fbe6b2bb49d50c", size = 489927, upload-time = "2026-05-18T04:31:42.485Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/watchfiles/1.2.0/watchfiles-1.2.0-cp312-cp312-manylinux_2_17_ppc64le.manylinux2014_ppc64le.whl", hash = "sha256:41bc1199f7523b3f82843c88cbb979180c949caef0342cf90968f178e5d49b01", size = 570476, upload-time = "2026-05-18T04:31:03.071Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/watchfiles/1.2.0/watchfiles-1.2.0-cp312-cp312-manylinux_2_17_s390x.manylinux2014_s390x.whl", hash = "sha256:7571e4464cb6e434958f867f7f730b8ab0b75e3f8e5eac0499168486ab3c33a8", size = 465650, upload-time = "2026-05-18T04:30:12.701Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/watchfiles/1.2.0/watchfiles-1.2.0-cp312-cp312-manylinux_2_17_x86_64.manylinux2014_x86_64.whl", hash = "sha256:e53a384f76b631c3ae5334ce6a52f0baa3a911eb94a4eac7f160079868b716d5", size = 456398, upload-time = "2026-05-18T04:30:13.784Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/watchfiles/1.2.0/watchfiles-1.2.0-cp312-cp312-manylinux_2_31_riscv64.whl", hash = "sha256:d20029a60a71a052a24c4db7673bc4de39ab89adbaccbfb5d67987c5d73f424d", size = 465140, upload-time = "2026-05-18T04:31:52.111Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/watchfiles/1.2.0/watchfiles-1.2.0-cp312-cp312-musllinux_1_1_aarch64.whl", hash = "sha256:2cb93af48550faf1cea04c303107c8b75833de7013e57ce27d3b8d21d8d0f58c", size = 630259, upload-time = "2026-05-18T04:31:25.676Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/watchfiles/1.2.0/watchfiles-1.2.0-cp312-cp312-musllinux_1_1_x86_64.whl", hash = "sha256:2995c176de7692b86a2e4c58d9ec718f753150a979cb4a754e2b4ffa38e70906", size = 659859, upload-time = "2026-05-18T04:30:24.333Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/watchfiles/1.2.0/watchfiles-1.2.0-cp312-cp312-win32.whl", hash = "sha256:7a2cffd17d27d2ecbb310c2b1d8174f222a5495b1a721894afa88ec11e25b898", size = 275480, upload-time = "2026-05-18T04:30:31.307Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/watchfiles/1.2.0/watchfiles-1.2.0-cp312-cp312-win_amd64.whl", hash = "sha256:f155b3a1b2a5fc89cdc70d47ee5d54e3b75e88efa34982028a35daef9ba00379", size = 288718, upload-time = "2026-05-18T04:32:10.745Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/watchfiles/1.2.0/watchfiles-1.2.0-cp312-cp312-win_arm64.whl", hash = "sha256:8fa585ede612ee9f9e91b18bebf9ba11b9ae29a4e3a0d0cf6fca3e382133f0d5", size = 281026, upload-time = "2026-05-18T04:30:22.23Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/watchfiles/1.2.0/watchfiles-1.2.0-pp311-pypy311_pp73-macosx_10_12_x86_64.whl", hash = "sha256:4674d49eb94706dfe666c069fc0a1b646ffcf920473492e209f6d5f60d3f0cc2", size = 403050, upload-time = "2026-05-18T04:30:36.753Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/watchfiles/1.2.0/watchfiles-1.2.0-pp311-pypy311_pp73-macosx_11_0_arm64.whl", hash = "sha256:094b9b70103d4e963499bdea001ee3c2697b144cd9ae6218a62c0f89ec9e31db", size = 396629, upload-time = "2026-05-18T04:32:03.268Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/watchfiles/1.2.0/watchfiles-1.2.0-pp311-pypy311_pp73-manylinux_2_17_aarch64.manylinux2014_aarch64.whl", hash = "sha256:b0ef001f8c25ad0fa9529f914c1600647ecd0f542d11c19b7894768c67b6acb7", size = 457318, upload-time = "2026-05-18T04:31:01.932Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/watchfiles/1.2.0/watchfiles-1.2.0-pp311-pypy311_pp73-manylinux_2_17_x86_64.manylinux2014_x86_64.whl", hash = "sha256:a88fc94e647bc4eec523f1caa540258eb71d14278b9daf72fa1e2658a98df0f0", size = 457771, upload-time = "2026-05-18T04:30:56.331Z" }, ] [[package]] name = "xxhash" version = "4.0.1" -source = { registry = "https://pypi.org/simple" } -sdist = { url = "https://files.pythonhosted.org/packages/f6/a5/1386f35da1475fcaeef42581deae73417c6d2a6a0b2d2e8914de18844dcd/xxhash-4.0.1.tar.gz", hash = "sha256:d55bf4ef10eb09b8b6866790e083d26d087d84caa3cc0946ba87c3ca7ecaf7b7", size = 101513, upload-time = "2026-08-17T08:24:08.557Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/0e/58/bc81e25cceab76ce4b400441e3a43312bf3887fedbb2e5f80cc5a7dd7f75/xxhash-4.0.1-cp311-cp311-macosx_10_9_x86_64.whl", hash = "sha256:8b4477edc03091f51f5309406d230851c23cf4822029e3bf40b8df53093fff1c", size = 38473, upload-time = "2026-08-17T08:20:34.303Z" }, - { url = "https://files.pythonhosted.org/packages/64/8d/d95e810c9a2930906f1fbd0e38f77554abb8e042a190aa9cbb24da39e9a2/xxhash-4.0.1-cp311-cp311-macosx_11_0_arm64.whl", hash = "sha256:04f9a24de11a6647666d5302fd73d6a5224ce50ddc965fb0bb44cee736e6bd7c", size = 36228, upload-time = "2026-08-17T08:20:36.641Z" }, - { url = "https://files.pythonhosted.org/packages/72/be/ebcded2a32ba664a17a1737a91b6baa298bcda6942acdad81af81660b74c/xxhash-4.0.1-cp311-cp311-manylinux1_i686.manylinux_2_28_i686.manylinux_2_5_i686.whl", hash = "sha256:6a8c5ce76b94ba49f3be8a8f2611abc6564210702c72ac9e237ca2bebfd17794", size = 253292, upload-time = "2026-08-17T08:21:27.206Z" }, - { url = "https://files.pythonhosted.org/packages/52/2a/72d31d787d988d1130253bc34d249c553f02c9387e7dda789b6d5aa7963b/xxhash-4.0.1-cp311-cp311-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:b4c8842fb19d78b5e8c2a52baf4c8357658cc56c62bc822b86ce0f942f28e286", size = 276545, upload-time = "2026-08-17T08:20:26.726Z" }, - { url = "https://files.pythonhosted.org/packages/39/ae/048f3b1f283a340bef3697fe5a9d4f10de1696a379af9af6b2352c24d82a/xxhash-4.0.1-cp311-cp311-manylinux2014_armv7l.manylinux_2_17_armv7l.manylinux_2_31_armv7l.whl", hash = "sha256:a43418e1a90b4809a9caf64aeb8b0696e3e1f300a323acc1e6ee2f93ae319fcf", size = 296295, upload-time = "2026-08-17T08:20:15.759Z" }, - { url = "https://files.pythonhosted.org/packages/8e/43/e9a593c81445c8e8d669402edb8264144624e7e5e2406df7d33e3c164d90/xxhash-4.0.1-cp311-cp311-manylinux2014_ppc64le.manylinux_2_17_ppc64le.manylinux_2_28_ppc64le.whl", hash = "sha256:b3662719007e059abde7eddacf8517142ba076ddc7b30c807260e57d28c3c191", size = 279966, upload-time = "2026-08-17T08:21:22.217Z" }, - { url = "https://files.pythonhosted.org/packages/f9/4c/29da2b166955a300521c0abd498ad4481916c3743949b106192b781f58fc/xxhash-4.0.1-cp311-cp311-manylinux2014_s390x.manylinux_2_17_s390x.manylinux_2_28_s390x.whl", hash = "sha256:06d7fbd609503c3be5e65cdb6bb2f040d6a98574404e2e1d5c60815c97fff4aa", size = 509850, upload-time = "2026-08-17T08:20:32.731Z" }, - { url = "https://files.pythonhosted.org/packages/40/ff/e39e1900179ce7f0f23e7000d9fc672471a8d05b6b2dc657cb496c8975d0/xxhash-4.0.1-cp311-cp311-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:101aa300de6ceef3d9c77569706330d8921fc45dd82bceed2084f1e9f2557a24", size = 261005, upload-time = "2026-08-17T08:20:40.735Z" }, - { url = "https://files.pythonhosted.org/packages/89/79/76ee26720d13219458f4b2ec0b22f539fe2bcb1f83dc22a24be4cda4e285/xxhash-4.0.1-cp311-cp311-manylinux_2_31_riscv64.manylinux_2_39_riscv64.whl", hash = "sha256:e4296fcc790876a8b0f297edc83d3b088457b774d8f67b4636807f8a2ec69a79", size = 339620, upload-time = "2026-08-17T08:20:50.371Z" }, - { url = "https://files.pythonhosted.org/packages/2f/c7/56b53252d64ecb2f3a0d283b41033561b1bec9c268095150153b694c2e52/xxhash-4.0.1-cp311-cp311-musllinux_1_2_aarch64.whl", hash = "sha256:57d7fa8f23908d173001c21a9e82bfc6ad997d1b6c270fb121812b7ed158891c", size = 272558, upload-time = "2026-08-17T08:21:33.883Z" }, - { url = "https://files.pythonhosted.org/packages/fc/76/83101a2f2ad3eb6b4b5d571e0aa394a576e3e23121707d507d80197ac385/xxhash-4.0.1-cp311-cp311-musllinux_1_2_armv7l.whl", hash = "sha256:85e402dab0f9acd3604539747c6fcc57dc188a18af6ab07eb8189351cd32466c", size = 300417, upload-time = "2026-08-17T08:20:36.183Z" }, - { url = "https://files.pythonhosted.org/packages/34/5d/5be1166ec4fc4bc896e508dc51189a2dad200b373269a430e214fb152693/xxhash-4.0.1-cp311-cp311-musllinux_1_2_i686.whl", hash = "sha256:fb59a0dd61fb2ad481c03fda399d78ce57dab6bb62c2c8fdb446a7ba4754b89a", size = 259286, upload-time = "2026-08-17T08:20:21.51Z" }, - { url = "https://files.pythonhosted.org/packages/77/6c/42f6201f13278787c58a6b3a1cd597b47e8665a7d4f68a3440d001c05a75/xxhash-4.0.1-cp311-cp311-musllinux_1_2_ppc64le.whl", hash = "sha256:0b20a06454b34f1531fc677c54efe2ecdec691ef9224f7fa919bf2c1363f7ff1", size = 278230, upload-time = "2026-08-17T08:21:47.62Z" }, - { url = "https://files.pythonhosted.org/packages/7f/37/3cb7f149a5a469c628604c33bec6d79764698b315700a4bbf1c6fb15622b/xxhash-4.0.1-cp311-cp311-musllinux_1_2_riscv64.whl", hash = "sha256:f7db035447a0ac8959aa230c5d36545ecf9f547413eb1711c0ca6f0ba1418925", size = 329846, upload-time = "2026-08-17T08:35:10.922Z" }, - { url = "https://files.pythonhosted.org/packages/8c/ee/934eb1e11f4d0e95f3828825f8da4b6d59e338235d63edc3d929769262d2/xxhash-4.0.1-cp311-cp311-musllinux_1_2_s390x.whl", hash = "sha256:34ed93e20bfd98d722b902121643791eeb4b1641871e2dc63d0d4c2d93f187df", size = 477285, upload-time = "2026-08-17T08:20:46.297Z" }, - { url = "https://files.pythonhosted.org/packages/e2/0b/77835dcd7aec970db74f5724c94207ada83b3e2f5e1474ab44b9c8fe5ea5/xxhash-4.0.1-cp311-cp311-musllinux_1_2_x86_64.whl", hash = "sha256:f6247f5e23ee94f2557ac9dab738a336f607c6ff476fcf66ca70c3aef5eee15a", size = 257552, upload-time = "2026-08-17T08:20:58.211Z" }, - { url = "https://files.pythonhosted.org/packages/da/6c/a3ce7a1a9c4ec9ad8babba591a996a71c474ff9630c5fb04c8c9d8b995ca/xxhash-4.0.1-cp311-cp311-win32.whl", hash = "sha256:348c8f288dc961d6bbd1985c8152a3ed7a85c95df00e82320f0c5215d922a399", size = 34630, upload-time = "2026-08-17T08:21:53.323Z" }, - { url = "https://files.pythonhosted.org/packages/d9/58/60d2170e8cda0891aab25dcaa74797b420c3c6cde8a5bc8372f17b30c0cf/xxhash-4.0.1-cp311-cp311-win_amd64.whl", hash = "sha256:ac0f291ab6485bd71f33941f9b92771318332a05d505460b41e893a549caadc0", size = 36992, upload-time = "2026-08-17T08:20:41.973Z" }, - { url = "https://files.pythonhosted.org/packages/ac/de/b229a39f9bbbe30cbcf9afaaa8993cb65286715c90a3b058c3769196ae02/xxhash-4.0.1-cp311-cp311-win_arm64.whl", hash = "sha256:72f34834518157a75e7090f328ee7a16c70c804cfc7c694fa069cc888e9fc03e", size = 33289, upload-time = "2026-08-17T08:20:32.176Z" }, - { url = "https://files.pythonhosted.org/packages/26/6c/dc7cffeadd06336cd934947187cd38abb263103bbc552ca0f55fe4ff595a/xxhash-4.0.1-cp312-cp312-macosx_10_13_x86_64.whl", hash = "sha256:1ee523f51718e41753f04f7102bb4dc55a18d2ea5cbaceef8ec7ca08571bd428", size = 38444, upload-time = "2026-08-17T08:21:54.332Z" }, - { url = "https://files.pythonhosted.org/packages/75/c9/cf736f6db8c3273af18925061572db0d4357818a9ce425f4b5fb0021918e/xxhash-4.0.1-cp312-cp312-macosx_11_0_arm64.whl", hash = "sha256:515a822c73abbf6a0b7c70976d9662be342835c9d78b8dc7c023411f39c35dbc", size = 36195, upload-time = "2026-08-17T08:35:13.004Z" }, - { url = "https://files.pythonhosted.org/packages/da/a2/ca1929354b6851529d0148f7f335b5e2b0281f83bab3e19f0896dc579796/xxhash-4.0.1-cp312-cp312-manylinux1_i686.manylinux_2_28_i686.manylinux_2_5_i686.whl", hash = "sha256:f5d031f35962e5483a613214e61f09fe24ab523062c3646d592dc16c4a217451", size = 253113, upload-time = "2026-08-17T08:20:52.152Z" }, - { url = "https://files.pythonhosted.org/packages/de/bb/542005206af59518bc8d78a210f1e0172217bc53beb32f64a5b632e72b6b/xxhash-4.0.1-cp312-cp312-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:da0264844a09b538c894e5eff25313d941deb4dedec2131b98418a71a3c9944e", size = 276525, upload-time = "2026-08-17T08:21:01.886Z" }, - { url = "https://files.pythonhosted.org/packages/1b/df/607cff25dcb0f1d35c3b04493f6ad8471edb03fd4eacbdcc5ceddef1f3e9/xxhash-4.0.1-cp312-cp312-manylinux2014_armv7l.manylinux_2_17_armv7l.manylinux_2_31_armv7l.whl", hash = "sha256:1642907941ee4b75aacc3db688af52ea02ca2305ab22af7ee686ed726b332684", size = 297703, upload-time = "2026-08-17T08:21:57.958Z" }, - { url = "https://files.pythonhosted.org/packages/15/ba/9d2275eea0b9d9c6b02921be23f7588356c60df95c763b25f0e045894d43/xxhash-4.0.1-cp312-cp312-manylinux2014_ppc64le.manylinux_2_17_ppc64le.manylinux_2_28_ppc64le.whl", hash = "sha256:4af350bc3f329970c0e3a59af84a8a30998bf8a9167eb50cd48e59baaa1d7bec", size = 280252, upload-time = "2026-08-17T08:20:47.299Z" }, - { url = "https://files.pythonhosted.org/packages/1d/aa/2299d9f6369e550aef2abb64945e39daa34412725aa46a20d99b74d76f67/xxhash-4.0.1-cp312-cp312-manylinux2014_s390x.manylinux_2_17_s390x.manylinux_2_28_s390x.whl", hash = "sha256:8ba782ca3bf1e81492611152b9a0d5264971339e95e34d69de0ac2c926be496d", size = 511041, upload-time = "2026-08-17T08:20:36.771Z" }, - { url = "https://files.pythonhosted.org/packages/83/97/31bd8b8279e6935a0719f6910ced15e9d5a2cd554b253f6027ce1b5a1c2c/xxhash-4.0.1-cp312-cp312-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:237b8f63a2a0fcfb1ffc06e21dad23add44e6d354b2b014364a1d41e419a4dee", size = 261812, upload-time = "2026-08-17T08:22:00.469Z" }, - { url = "https://files.pythonhosted.org/packages/2d/c1/d180a2da23c105d8e0b02d54f9f5841013fc81c233010ec781e31f1aee4c/xxhash-4.0.1-cp312-cp312-manylinux_2_31_riscv64.manylinux_2_39_riscv64.whl", hash = "sha256:81507a68ba84c55241fb61cce1469f473a5da4205fc8ef6f698e5948eea8dd88", size = 339878, upload-time = "2026-08-17T08:35:17.626Z" }, - { url = "https://files.pythonhosted.org/packages/a8/3d/f584cd3172fe934f0f5a0a3917d0d7ce781f74d794fd43bb72be71c3ef6f/xxhash-4.0.1-cp312-cp312-musllinux_1_2_aarch64.whl", hash = "sha256:5f1ea31d61bcd2cd2f3ec4ca80a64187bbd7948f490b63cf0dcbc6e717b4c1e9", size = 272871, upload-time = "2026-08-17T08:20:56.067Z" }, - { url = "https://files.pythonhosted.org/packages/34/50/2c7956b2b551682e00b9aebce9ceb0a991a131d65f9850c09f5f9760be2e/xxhash-4.0.1-cp312-cp312-musllinux_1_2_armv7l.whl", hash = "sha256:06713a5aaf1d0905c5579416c020c02e42b3ceb931e86c7d3b7fb85403dee3f3", size = 301440, upload-time = "2026-08-17T08:21:35.911Z" }, - { url = "https://files.pythonhosted.org/packages/eb/a2/0739f6482184a8026f4b022718f5f815d352059312e80696825433f0a8e7/xxhash-4.0.1-cp312-cp312-musllinux_1_2_i686.whl", hash = "sha256:e8cda075b10bb3917b002c74a04f9e02b7d13b5bf732571404d51c52b11c7329", size = 260157, upload-time = "2026-08-17T08:22:01.416Z" }, - { url = "https://files.pythonhosted.org/packages/a1/25/b31a7bcf1d7d116842812e54f9b944843b4236ea4fa85634e8259f342212/xxhash-4.0.1-cp312-cp312-musllinux_1_2_ppc64le.whl", hash = "sha256:c10b9206753b64aa791b35b201485477525b26fdec5bf86e8364c388a03e2592", size = 278233, upload-time = "2026-08-17T08:21:15.674Z" }, - { url = "https://files.pythonhosted.org/packages/db/e8/5293bae090fc6119dbc5fcf5c4cc0e1536394b52d73b7904d033836c73db/xxhash-4.0.1-cp312-cp312-musllinux_1_2_riscv64.whl", hash = "sha256:f3e1a44af01b6692de0ec6caba5f0bf93ceb36896e02b7fc00952c6ea7ef39e1", size = 330270, upload-time = "2026-08-17T08:20:51.128Z" }, - { url = "https://files.pythonhosted.org/packages/72/9e/e2ab12d40921f3f34c9317637d65e011aeababf8288356ea8d527de2c1d0/xxhash-4.0.1-cp312-cp312-musllinux_1_2_s390x.whl", hash = "sha256:c6fc415b5568bd9accc7187f1729a99707330c0a67a8b9f93c1149ed573ed75d", size = 478555, upload-time = "2026-08-17T08:22:04.183Z" }, - { url = "https://files.pythonhosted.org/packages/6d/32/c6148d39a49efa95f39b4cf0d41ef35a487f3b30f6fb1fc8fe8d8eab577e/xxhash-4.0.1-cp312-cp312-musllinux_1_2_x86_64.whl", hash = "sha256:96d8de55029d42251945531f6aa7590c32b48163c66a43bf29d8657d7446a377", size = 258174, upload-time = "2026-08-17T08:35:21.18Z" }, - { url = "https://files.pythonhosted.org/packages/8f/fb/0b04b68d6c5bc71c7a2c344f1287327b67e607f28fbcfd937697caca64b6/xxhash-4.0.1-cp312-cp312-pyemscripten_2024_0_wasm32.whl", hash = "sha256:0163b5d259de23ae9e07b7eabf435ce4704f6f205589a2b154e6af4be985ce1b", size = 20767, upload-time = "2026-08-17T08:21:00.806Z" }, - { url = "https://files.pythonhosted.org/packages/a6/be/476092aba34d1fcd313e1613a3bb3bc692f253d167b54bc90049043b5034/xxhash-4.0.1-cp312-cp312-win32.whl", hash = "sha256:1216f7ba5683f17a89eb7dcb4bc50a0b743dfe1902278d7b3d0786f538118433", size = 34669, upload-time = "2026-08-17T08:21:49.486Z" }, - { url = "https://files.pythonhosted.org/packages/aa/02/f9413d94fae43cec6d1a74c4f12156c6f4a7f5fd50e1d34defebdee3dec9/xxhash-4.0.1-cp312-cp312-win_amd64.whl", hash = "sha256:5c2d525a3afabcd8e3549d85fc7e111fde6bc302d06a1893fe73adb79823415e", size = 37073, upload-time = "2026-08-17T08:22:04.886Z" }, - { url = "https://files.pythonhosted.org/packages/c1/83/6fe93c1b95acf962bc61a246df09dc2dcce895ccfc1080c9f48d0b652b92/xxhash-4.0.1-cp312-cp312-win_arm64.whl", hash = "sha256:86b2b12bec60c678ed8f5cca0258ad93a8928ebddb6ca7732f0875afe1451d1a", size = 33299, upload-time = "2026-08-17T08:35:12.708Z" }, - { url = "https://files.pythonhosted.org/packages/86/79/9127ff42a887a348dc4ce3211cf1a962836887adee6f57078132bfba78b4/xxhash-4.0.1-graalpy312-graalpy250_312_native-macosx_10_13_x86_64.whl", hash = "sha256:ff48915bf1871a1f19f74c11834c6329443d306cedc0c05fe7fe617810422a80", size = 31836, upload-time = "2026-08-17T08:36:28.261Z" }, - { url = "https://files.pythonhosted.org/packages/0a/e6/f238693bfdd642adb59c99683964d46d9947fe721ff44d3bd850ae675407/xxhash-4.0.1-graalpy312-graalpy250_312_native-macosx_11_0_arm64.whl", hash = "sha256:4a76345f5aceb4ec404918edf9c7f2b5507db864dc0d7455982009ac0890b57b", size = 34453, upload-time = "2026-08-17T08:23:49.795Z" }, - { url = "https://files.pythonhosted.org/packages/40/4b/796ace33cdfb75c91ba6d11615c3bd436355b9f3103e05865bbee9abce57/xxhash-4.0.1-graalpy312-graalpy250_312_native-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:31d86f9e81f3e84e00131ac7c54caf5119ae4ddd82c09c31cff597c813ce1ee2", size = 38488, upload-time = "2026-08-17T08:23:59.901Z" }, - { url = "https://files.pythonhosted.org/packages/ad/23/2d549e5d5d7759eaf9ac2d2d2ab81ff60f1bb2b52cdaae8e5ec5c6524354/xxhash-4.0.1-graalpy312-graalpy250_312_native-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:deca2a30d983d240b8375ec2ee0a4288e72042827fc61df2f7671f8467e4cb2f", size = 38206, upload-time = "2026-08-17T08:36:32.193Z" }, - { url = "https://files.pythonhosted.org/packages/79/98/1ee576b27f78e6107ee4ea8ac03e8a52888dff256e57d560f8282c195563/xxhash-4.0.1-graalpy312-graalpy250_312_native-win_amd64.whl", hash = "sha256:7c343ee174d417a44d0c3355602c0cbbfa52a04d1bbbf1723378c7d2c8f60626", size = 37127, upload-time = "2026-08-17T08:23:42.705Z" }, - { url = "https://files.pythonhosted.org/packages/ea/4f/e0648288a17d0d1084ca4f7bef206097831988fc86af74aa1dff8f1fbd68/xxhash-4.0.1-pp311-pypy311_pp73-macosx_10_15_x86_64.whl", hash = "sha256:554f87034635bcec47c5d72447bf3db7e02da1bf493a0ada010db28a76f891c6", size = 36333, upload-time = "2026-08-17T08:23:52.913Z" }, - { url = "https://files.pythonhosted.org/packages/22/15/34b7f72e9b5a8bfd7e6178de9e1e342bc3de9f07111a5ae26c00506d9edf/xxhash-4.0.1-pp311-pypy311_pp73-macosx_11_0_arm64.whl", hash = "sha256:3c2445edafc300cc40feb6a25a8356a971c30cd0bf47b5349c2ad74c508343b1", size = 33519, upload-time = "2026-08-17T08:36:03.04Z" }, - { url = "https://files.pythonhosted.org/packages/06/4b/e0af324ccf701bf84a7a060bef11c915d45d9e3c5b9caf5b94d62ecb040b/xxhash-4.0.1-pp311-pypy311_pp73-manylinux1_i686.manylinux_2_28_i686.manylinux_2_5_i686.whl", hash = "sha256:bdd16718b63aa3ebd68aabb79021a40e47c81374852d41a306b9453141bbcbee", size = 47995, upload-time = "2026-08-17T08:24:14.335Z" }, - { url = "https://files.pythonhosted.org/packages/59/46/cc7130e6ca6b41ab72eb6a03177b933c7c74145545c99a8610ecc208c449/xxhash-4.0.1-pp311-pypy311_pp73-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:8b99ebaf9e816ac5069423b1367ee7e8078fbcebcf62545506bb0608d2f4f468", size = 42725, upload-time = "2026-08-17T08:36:34.144Z" }, - { url = "https://files.pythonhosted.org/packages/1c/24/4f26ff9a7dd0998f6d1036bdddef7ce3e78972a74f7fffa7967e7bc3b7e4/xxhash-4.0.1-pp311-pypy311_pp73-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:f484ed57bb3e4142f9d6439568658c38be5f94b702ba00a1ff32c69783b6c66d", size = 39532, upload-time = "2026-08-17T08:23:57.556Z" }, - { url = "https://files.pythonhosted.org/packages/7d/f3/1ac078fc8fceadcf066469acecacb35d2821cbfaf7d6fc5ac2107c7a314d/xxhash-4.0.1-pp311-pypy311_pp73-win_amd64.whl", hash = "sha256:fac4832b638000106207bc44e44b9616a6a416aaee56c62b01d61f3705e49f58", size = 37252, upload-time = "2026-08-17T08:24:06.652Z" }, +source = { registry = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/" } +sdist = { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/xxhash/4.0.1/xxhash-4.0.1.tar.gz", hash = "sha256:d55bf4ef10eb09b8b6866790e083d26d087d84caa3cc0946ba87c3ca7ecaf7b7", size = 101513, upload-time = "2026-08-17T08:24:08.557Z" } +wheels = [ + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/xxhash/4.0.1/xxhash-4.0.1-cp311-cp311-macosx_10_9_x86_64.whl", hash = "sha256:8b4477edc03091f51f5309406d230851c23cf4822029e3bf40b8df53093fff1c", size = 38473, upload-time = "2026-08-17T08:20:34.303Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/xxhash/4.0.1/xxhash-4.0.1-cp311-cp311-macosx_11_0_arm64.whl", hash = "sha256:04f9a24de11a6647666d5302fd73d6a5224ce50ddc965fb0bb44cee736e6bd7c", size = 36228, upload-time = "2026-08-17T08:20:36.641Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/xxhash/4.0.1/xxhash-4.0.1-cp311-cp311-manylinux1_i686.manylinux_2_28_i686.manylinux_2_5_i686.whl", hash = "sha256:6a8c5ce76b94ba49f3be8a8f2611abc6564210702c72ac9e237ca2bebfd17794", size = 253292, upload-time = "2026-08-17T08:21:27.206Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/xxhash/4.0.1/xxhash-4.0.1-cp311-cp311-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:b4c8842fb19d78b5e8c2a52baf4c8357658cc56c62bc822b86ce0f942f28e286", size = 276545, upload-time = "2026-08-17T08:20:26.726Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/xxhash/4.0.1/xxhash-4.0.1-cp311-cp311-manylinux2014_armv7l.manylinux_2_17_armv7l.manylinux_2_31_armv7l.whl", hash = "sha256:a43418e1a90b4809a9caf64aeb8b0696e3e1f300a323acc1e6ee2f93ae319fcf", size = 296295, upload-time = "2026-08-17T08:20:15.759Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/xxhash/4.0.1/xxhash-4.0.1-cp311-cp311-manylinux2014_ppc64le.manylinux_2_17_ppc64le.manylinux_2_28_ppc64le.whl", hash = "sha256:b3662719007e059abde7eddacf8517142ba076ddc7b30c807260e57d28c3c191", size = 279966, upload-time = "2026-08-17T08:21:22.217Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/xxhash/4.0.1/xxhash-4.0.1-cp311-cp311-manylinux2014_s390x.manylinux_2_17_s390x.manylinux_2_28_s390x.whl", hash = "sha256:06d7fbd609503c3be5e65cdb6bb2f040d6a98574404e2e1d5c60815c97fff4aa", size = 509850, upload-time = "2026-08-17T08:20:32.731Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/xxhash/4.0.1/xxhash-4.0.1-cp311-cp311-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:101aa300de6ceef3d9c77569706330d8921fc45dd82bceed2084f1e9f2557a24", size = 261005, upload-time = "2026-08-17T08:20:40.735Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/xxhash/4.0.1/xxhash-4.0.1-cp311-cp311-manylinux_2_31_riscv64.manylinux_2_39_riscv64.whl", hash = "sha256:e4296fcc790876a8b0f297edc83d3b088457b774d8f67b4636807f8a2ec69a79", size = 339620, upload-time = "2026-08-17T08:20:50.371Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/xxhash/4.0.1/xxhash-4.0.1-cp311-cp311-musllinux_1_2_aarch64.whl", hash = "sha256:57d7fa8f23908d173001c21a9e82bfc6ad997d1b6c270fb121812b7ed158891c", size = 272558, upload-time = "2026-08-17T08:21:33.883Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/xxhash/4.0.1/xxhash-4.0.1-cp311-cp311-musllinux_1_2_armv7l.whl", hash = "sha256:85e402dab0f9acd3604539747c6fcc57dc188a18af6ab07eb8189351cd32466c", size = 300417, upload-time = "2026-08-17T08:20:36.183Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/xxhash/4.0.1/xxhash-4.0.1-cp311-cp311-musllinux_1_2_i686.whl", hash = "sha256:fb59a0dd61fb2ad481c03fda399d78ce57dab6bb62c2c8fdb446a7ba4754b89a", size = 259286, upload-time = "2026-08-17T08:20:21.51Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/xxhash/4.0.1/xxhash-4.0.1-cp311-cp311-musllinux_1_2_ppc64le.whl", hash = "sha256:0b20a06454b34f1531fc677c54efe2ecdec691ef9224f7fa919bf2c1363f7ff1", size = 278230, upload-time = "2026-08-17T08:21:47.62Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/xxhash/4.0.1/xxhash-4.0.1-cp311-cp311-musllinux_1_2_riscv64.whl", hash = "sha256:f7db035447a0ac8959aa230c5d36545ecf9f547413eb1711c0ca6f0ba1418925", size = 329846, upload-time = "2026-08-17T08:35:10.922Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/xxhash/4.0.1/xxhash-4.0.1-cp311-cp311-musllinux_1_2_s390x.whl", hash = "sha256:34ed93e20bfd98d722b902121643791eeb4b1641871e2dc63d0d4c2d93f187df", size = 477285, upload-time = "2026-08-17T08:20:46.297Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/xxhash/4.0.1/xxhash-4.0.1-cp311-cp311-musllinux_1_2_x86_64.whl", hash = "sha256:f6247f5e23ee94f2557ac9dab738a336f607c6ff476fcf66ca70c3aef5eee15a", size = 257552, upload-time = "2026-08-17T08:20:58.211Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/xxhash/4.0.1/xxhash-4.0.1-cp311-cp311-win32.whl", hash = "sha256:348c8f288dc961d6bbd1985c8152a3ed7a85c95df00e82320f0c5215d922a399", size = 34630, upload-time = "2026-08-17T08:21:53.323Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/xxhash/4.0.1/xxhash-4.0.1-cp311-cp311-win_amd64.whl", hash = "sha256:ac0f291ab6485bd71f33941f9b92771318332a05d505460b41e893a549caadc0", size = 36992, upload-time = "2026-08-17T08:20:41.973Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/xxhash/4.0.1/xxhash-4.0.1-cp311-cp311-win_arm64.whl", hash = "sha256:72f34834518157a75e7090f328ee7a16c70c804cfc7c694fa069cc888e9fc03e", size = 33289, upload-time = "2026-08-17T08:20:32.176Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/xxhash/4.0.1/xxhash-4.0.1-cp312-cp312-macosx_10_13_x86_64.whl", hash = "sha256:1ee523f51718e41753f04f7102bb4dc55a18d2ea5cbaceef8ec7ca08571bd428", size = 38444, upload-time = "2026-08-17T08:21:54.332Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/xxhash/4.0.1/xxhash-4.0.1-cp312-cp312-macosx_11_0_arm64.whl", hash = "sha256:515a822c73abbf6a0b7c70976d9662be342835c9d78b8dc7c023411f39c35dbc", size = 36195, upload-time = "2026-08-17T08:35:13.004Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/xxhash/4.0.1/xxhash-4.0.1-cp312-cp312-manylinux1_i686.manylinux_2_28_i686.manylinux_2_5_i686.whl", hash = "sha256:f5d031f35962e5483a613214e61f09fe24ab523062c3646d592dc16c4a217451", size = 253113, upload-time = "2026-08-17T08:20:52.152Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/xxhash/4.0.1/xxhash-4.0.1-cp312-cp312-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:da0264844a09b538c894e5eff25313d941deb4dedec2131b98418a71a3c9944e", size = 276525, upload-time = "2026-08-17T08:21:01.886Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/xxhash/4.0.1/xxhash-4.0.1-cp312-cp312-manylinux2014_armv7l.manylinux_2_17_armv7l.manylinux_2_31_armv7l.whl", hash = "sha256:1642907941ee4b75aacc3db688af52ea02ca2305ab22af7ee686ed726b332684", size = 297703, upload-time = "2026-08-17T08:21:57.958Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/xxhash/4.0.1/xxhash-4.0.1-cp312-cp312-manylinux2014_ppc64le.manylinux_2_17_ppc64le.manylinux_2_28_ppc64le.whl", hash = "sha256:4af350bc3f329970c0e3a59af84a8a30998bf8a9167eb50cd48e59baaa1d7bec", size = 280252, upload-time = "2026-08-17T08:20:47.299Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/xxhash/4.0.1/xxhash-4.0.1-cp312-cp312-manylinux2014_s390x.manylinux_2_17_s390x.manylinux_2_28_s390x.whl", hash = "sha256:8ba782ca3bf1e81492611152b9a0d5264971339e95e34d69de0ac2c926be496d", size = 511041, upload-time = "2026-08-17T08:20:36.771Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/xxhash/4.0.1/xxhash-4.0.1-cp312-cp312-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:237b8f63a2a0fcfb1ffc06e21dad23add44e6d354b2b014364a1d41e419a4dee", size = 261812, upload-time = "2026-08-17T08:22:00.469Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/xxhash/4.0.1/xxhash-4.0.1-cp312-cp312-manylinux_2_31_riscv64.manylinux_2_39_riscv64.whl", hash = "sha256:81507a68ba84c55241fb61cce1469f473a5da4205fc8ef6f698e5948eea8dd88", size = 339878, upload-time = "2026-08-17T08:35:17.626Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/xxhash/4.0.1/xxhash-4.0.1-cp312-cp312-musllinux_1_2_aarch64.whl", hash = "sha256:5f1ea31d61bcd2cd2f3ec4ca80a64187bbd7948f490b63cf0dcbc6e717b4c1e9", size = 272871, upload-time = "2026-08-17T08:20:56.067Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/xxhash/4.0.1/xxhash-4.0.1-cp312-cp312-musllinux_1_2_armv7l.whl", hash = "sha256:06713a5aaf1d0905c5579416c020c02e42b3ceb931e86c7d3b7fb85403dee3f3", size = 301440, upload-time = "2026-08-17T08:21:35.911Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/xxhash/4.0.1/xxhash-4.0.1-cp312-cp312-musllinux_1_2_i686.whl", hash = "sha256:e8cda075b10bb3917b002c74a04f9e02b7d13b5bf732571404d51c52b11c7329", size = 260157, upload-time = "2026-08-17T08:22:01.416Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/xxhash/4.0.1/xxhash-4.0.1-cp312-cp312-musllinux_1_2_ppc64le.whl", hash = "sha256:c10b9206753b64aa791b35b201485477525b26fdec5bf86e8364c388a03e2592", size = 278233, upload-time = "2026-08-17T08:21:15.674Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/xxhash/4.0.1/xxhash-4.0.1-cp312-cp312-musllinux_1_2_riscv64.whl", hash = "sha256:f3e1a44af01b6692de0ec6caba5f0bf93ceb36896e02b7fc00952c6ea7ef39e1", size = 330270, upload-time = "2026-08-17T08:20:51.128Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/xxhash/4.0.1/xxhash-4.0.1-cp312-cp312-musllinux_1_2_s390x.whl", hash = "sha256:c6fc415b5568bd9accc7187f1729a99707330c0a67a8b9f93c1149ed573ed75d", size = 478555, upload-time = "2026-08-17T08:22:04.183Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/xxhash/4.0.1/xxhash-4.0.1-cp312-cp312-musllinux_1_2_x86_64.whl", hash = "sha256:96d8de55029d42251945531f6aa7590c32b48163c66a43bf29d8657d7446a377", size = 258174, upload-time = "2026-08-17T08:35:21.18Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/xxhash/4.0.1/xxhash-4.0.1-cp312-cp312-pyemscripten_2024_0_wasm32.whl", hash = "sha256:0163b5d259de23ae9e07b7eabf435ce4704f6f205589a2b154e6af4be985ce1b", size = 20767, upload-time = "2026-08-17T08:21:00.806Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/xxhash/4.0.1/xxhash-4.0.1-cp312-cp312-win32.whl", hash = "sha256:1216f7ba5683f17a89eb7dcb4bc50a0b743dfe1902278d7b3d0786f538118433", size = 34669, upload-time = "2026-08-17T08:21:49.486Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/xxhash/4.0.1/xxhash-4.0.1-cp312-cp312-win_amd64.whl", hash = "sha256:5c2d525a3afabcd8e3549d85fc7e111fde6bc302d06a1893fe73adb79823415e", size = 37073, upload-time = "2026-08-17T08:22:04.886Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/xxhash/4.0.1/xxhash-4.0.1-cp312-cp312-win_arm64.whl", hash = "sha256:86b2b12bec60c678ed8f5cca0258ad93a8928ebddb6ca7732f0875afe1451d1a", size = 33299, upload-time = "2026-08-17T08:35:12.708Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/xxhash/4.0.1/xxhash-4.0.1-graalpy312-graalpy250_312_native-macosx_10_13_x86_64.whl", hash = "sha256:ff48915bf1871a1f19f74c11834c6329443d306cedc0c05fe7fe617810422a80", size = 31836, upload-time = "2026-08-17T08:36:28.261Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/xxhash/4.0.1/xxhash-4.0.1-graalpy312-graalpy250_312_native-macosx_11_0_arm64.whl", hash = "sha256:4a76345f5aceb4ec404918edf9c7f2b5507db864dc0d7455982009ac0890b57b", size = 34453, upload-time = "2026-08-17T08:23:49.795Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/xxhash/4.0.1/xxhash-4.0.1-graalpy312-graalpy250_312_native-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:31d86f9e81f3e84e00131ac7c54caf5119ae4ddd82c09c31cff597c813ce1ee2", size = 38488, upload-time = "2026-08-17T08:23:59.901Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/xxhash/4.0.1/xxhash-4.0.1-graalpy312-graalpy250_312_native-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:deca2a30d983d240b8375ec2ee0a4288e72042827fc61df2f7671f8467e4cb2f", size = 38206, upload-time = "2026-08-17T08:36:32.193Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/xxhash/4.0.1/xxhash-4.0.1-graalpy312-graalpy250_312_native-win_amd64.whl", hash = "sha256:7c343ee174d417a44d0c3355602c0cbbfa52a04d1bbbf1723378c7d2c8f60626", size = 37127, upload-time = "2026-08-17T08:23:42.705Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/xxhash/4.0.1/xxhash-4.0.1-pp311-pypy311_pp73-macosx_10_15_x86_64.whl", hash = "sha256:554f87034635bcec47c5d72447bf3db7e02da1bf493a0ada010db28a76f891c6", size = 36333, upload-time = "2026-08-17T08:23:52.913Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/xxhash/4.0.1/xxhash-4.0.1-pp311-pypy311_pp73-macosx_11_0_arm64.whl", hash = "sha256:3c2445edafc300cc40feb6a25a8356a971c30cd0bf47b5349c2ad74c508343b1", size = 33519, upload-time = "2026-08-17T08:36:03.04Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/xxhash/4.0.1/xxhash-4.0.1-pp311-pypy311_pp73-manylinux1_i686.manylinux_2_28_i686.manylinux_2_5_i686.whl", hash = "sha256:bdd16718b63aa3ebd68aabb79021a40e47c81374852d41a306b9453141bbcbee", size = 47995, upload-time = "2026-08-17T08:24:14.335Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/xxhash/4.0.1/xxhash-4.0.1-pp311-pypy311_pp73-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:8b99ebaf9e816ac5069423b1367ee7e8078fbcebcf62545506bb0608d2f4f468", size = 42725, upload-time = "2026-08-17T08:36:34.144Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/xxhash/4.0.1/xxhash-4.0.1-pp311-pypy311_pp73-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:f484ed57bb3e4142f9d6439568658c38be5f94b702ba00a1ff32c69783b6c66d", size = 39532, upload-time = "2026-08-17T08:23:57.556Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/xxhash/4.0.1/xxhash-4.0.1-pp311-pypy311_pp73-win_amd64.whl", hash = "sha256:fac4832b638000106207bc44e44b9616a6a416aaee56c62b01d61f3705e49f58", size = 37252, upload-time = "2026-08-17T08:24:06.652Z" }, ] [[package]] name = "yarl" version = "1.24.5" -source = { registry = "https://pypi.org/simple" } +source = { registry = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/" } dependencies = [ { name = "idna" }, { name = "multidict" }, { name = "propcache" }, ] -sdist = { url = "https://files.pythonhosted.org/packages/31/33/ebe9e3d1f86c7a0b51094c0a146392045ca1631d2664889539dec8088a33/yarl-1.24.5.tar.gz", hash = "sha256:e81b83143bee16329c23db3c1b2d82b29892fcbcb849186d2f6e98a5abe9a57f", size = 228679, upload-time = "2026-07-20T02:07:45.435Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/fe/db/3cb5df059756a45761cc3dee8fd25ec82b83a6585ea3542b969fda850f99/yarl-1.24.5-cp311-cp311-macosx_10_9_universal2.whl", hash = "sha256:2c1fe720934a16ea8e7146175cba2126f87f54912c8c5435e7f7c7a51ef808d3", size = 135043, upload-time = "2026-07-20T02:04:52.39Z" }, - { url = "https://files.pythonhosted.org/packages/44/f8/767d6bd5a03db63bc467df2fb56d6fafeae9667d74aea92cd6af399f828b/yarl-1.24.5-cp311-cp311-macosx_10_9_x86_64.whl", hash = "sha256:c687ed078e145f5fd53a14854beff320e1d2ab76df03e2009c98f39a0f68f39a", size = 96942, upload-time = "2026-07-20T02:04:54.26Z" }, - { url = "https://files.pythonhosted.org/packages/ce/97/10b939c44d7b28d1dbc389cfc7012306d1ea8dba01eaef44b39fffaee52a/yarl-1.24.5-cp311-cp311-macosx_11_0_arm64.whl", hash = "sha256:709f1efed56c4a145793c046cd4939f9959bcd818979a787b77d8e09c57a0840", size = 97046, upload-time = "2026-07-20T02:04:56.638Z" }, - { url = "https://files.pythonhosted.org/packages/5b/7a/b410dbe39b6255c55fb2a2bcee96eb844d0789235ddc381a889a90dc72d6/yarl-1.24.5-cp311-cp311-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:874019bd513008b009f58657134e5d0c5e030b3559bd0553976837adf52fe966", size = 110512, upload-time = "2026-07-20T02:04:58.955Z" }, - { url = "https://files.pythonhosted.org/packages/83/c7/da591971f78a5617e1f21f5699858ebccd836fe181a6493788ffc91ba69b/yarl-1.24.5-cp311-cp311-manylinux2014_armv7l.manylinux_2_17_armv7l.manylinux_2_31_armv7l.whl", hash = "sha256:a4582acf7ef76482f6f511ebaf1946dae7f2e85ec4728b81a678c01df63bd723", size = 102454, upload-time = "2026-07-20T02:05:00.623Z" }, - { url = "https://files.pythonhosted.org/packages/c4/8e/73b0ed4de47289a78a96045d76d1cfe5e41848bf0da59ce25b2ec87ee05d/yarl-1.24.5-cp311-cp311-manylinux2014_ppc64le.manylinux_2_17_ppc64le.manylinux_2_28_ppc64le.whl", hash = "sha256:2cabe6546e41dabe439999a23fcb5246e0c3b595b4315b96ef755252be90caeb", size = 117617, upload-time = "2026-07-20T02:05:02.325Z" }, - { url = "https://files.pythonhosted.org/packages/cf/14/b744747bc4f57a8d55bd744df463457524583e1e9f7538b5ace0346ab92e/yarl-1.24.5-cp311-cp311-manylinux2014_s390x.manylinux_2_17_s390x.manylinux_2_28_s390x.whl", hash = "sha256:17f57620f5475b3c69109376cc87e42a7af5db13c9398e4292772a706ff10780", size = 116135, upload-time = "2026-07-20T02:05:04.05Z" }, - { url = "https://files.pythonhosted.org/packages/66/ca/95aa4d0e5b7ea4f20e4d577c42d001ed9df207569fdb063cc5ed4ebb496b/yarl-1.24.5-cp311-cp311-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:570fec8fbd22b032733625f03f10b7ff023bc399213db15e72a7acaef28c2f4e", size = 111935, upload-time = "2026-07-20T02:05:05.738Z" }, - { url = "https://files.pythonhosted.org/packages/72/0d/d2ad8d6b147832d177a4e720ba1962fe686eb0913b74503b3eca094b8bba/yarl-1.24.5-cp311-cp311-manylinux_2_31_riscv64.manylinux_2_39_riscv64.whl", hash = "sha256:5fede79c6f73ff2c3ef822864cb1ada23196e62756df53bc6231d351a49516a2", size = 110010, upload-time = "2026-07-20T02:05:07.471Z" }, - { url = "https://files.pythonhosted.org/packages/50/18/eb335e4120903903f4865041355ae46256a2406eb2865bc24827f4f27b61/yarl-1.24.5-cp311-cp311-musllinux_1_2_aarch64.whl", hash = "sha256:8ccf9aca873b767977c73df497a85dbedee4ee086ae9ae49dc461333b9b79f58", size = 110058, upload-time = "2026-07-20T02:05:09.246Z" }, - { url = "https://files.pythonhosted.org/packages/44/70/97353add32c62ad6f206d948ac5a5ee84398225e534dc6ed6433d1b335b6/yarl-1.24.5-cp311-cp311-musllinux_1_2_armv7l.whl", hash = "sha256:ad5d8201d310b031e6cd839d9bac2d4e5a01533ce5d3d5b50b7de1ef3af1de61", size = 103308, upload-time = "2026-07-20T02:05:11.31Z" }, - { url = "https://files.pythonhosted.org/packages/68/39/5e7398d4b6f6b3c9062823ebc60802df5b272e3fe9e788f9734c6ee46c85/yarl-1.24.5-cp311-cp311-musllinux_1_2_ppc64le.whl", hash = "sha256:841f0852f48fefea3b12c9dfec00704dfa3aef5215d0e3ce564bb3d7cd8d57c6", size = 116898, upload-time = "2026-07-20T02:05:13.099Z" }, - { url = "https://files.pythonhosted.org/packages/e4/c9/09e52f2239e8b96357eccca05915382e4ba5405ebfb623b6036040d99654/yarl-1.24.5-cp311-cp311-musllinux_1_2_riscv64.whl", hash = "sha256:9baafc71b04f8f4bb0703b21d6fc9f0c30b346c636a532ff16ec8491a5ea4b1f", size = 109400, upload-time = "2026-07-20T02:05:14.821Z" }, - { url = "https://files.pythonhosted.org/packages/4b/6a/e94133d4c2d1a14d2384310bf3e79d9cf32c9d1eae1c6f034fb80d098fa1/yarl-1.24.5-cp311-cp311-musllinux_1_2_s390x.whl", hash = "sha256:d897129df1a22b12aeed2c2c98df0785a2e8e6e0bde87b389491d0025c187077", size = 115934, upload-time = "2026-07-20T02:05:17.78Z" }, - { url = "https://files.pythonhosted.org/packages/4e/3c/34955ed967b976fc38edcbb6d538dee79dbda4cb7fc7f72a0907a7c78e0f/yarl-1.24.5-cp311-cp311-musllinux_1_2_x86_64.whl", hash = "sha256:dd625535328fd9882374356269227670189adfcc6a2d90284f323c05862eecbd", size = 112178, upload-time = "2026-07-20T02:05:19.675Z" }, - { url = "https://files.pythonhosted.org/packages/f5/46/d7bd3a8859d47dcfaffd7127af7076032a7da278a9a02e17b5f37bfb6712/yarl-1.24.5-cp311-cp311-win_amd64.whl", hash = "sha256:f4239bbec5a3577ddb49e4b50aeb32d8e5792098262ae2f63723f916a29b1a25", size = 97544, upload-time = "2026-07-20T02:05:21.523Z" }, - { url = "https://files.pythonhosted.org/packages/01/69/c1bfd21e32c638974ea2c542a0b8c53ef1fa9eff336020f5d014f9503ff2/yarl-1.24.5-cp311-cp311-win_arm64.whl", hash = "sha256:3ac6aff147deb9c09461b2d4bbdf6256831198f5d8a23f5d37138213090b6d8a", size = 93359, upload-time = "2026-07-20T02:05:23.493Z" }, - { url = "https://files.pythonhosted.org/packages/1b/84/71d051c850b5af41d168c679d9eb67eb7c55283ac4ee131673edf134bc4e/yarl-1.24.5-cp312-cp312-macosx_10_13_universal2.whl", hash = "sha256:d693396e5aea78db03decd60aec9ece16c9b40ba00a587f089615ff4e718a81d", size = 136035, upload-time = "2026-07-20T02:05:25.489Z" }, - { url = "https://files.pythonhosted.org/packages/03/4d/8ad27f9a1b7e69313cca5d695b925b48efe51208d3490e0844bae97cabc0/yarl-1.24.5-cp312-cp312-macosx_10_13_x86_64.whl", hash = "sha256:3363fcc96e665878946ad7a106b9a13eac0541766a690ef287c0232ac768b6ec", size = 97642, upload-time = "2026-07-20T02:05:27.429Z" }, - { url = "https://files.pythonhosted.org/packages/ea/b4/05b4131c407006cd1e410e9c6539f16a0945724677e5364447313c15ea3e/yarl-1.24.5-cp312-cp312-macosx_11_0_arm64.whl", hash = "sha256:9d399bdcfb4a0f659b9b3788bbc89babe63d9a6a65aacdf4d4e7065ff2e6316c", size = 97323, upload-time = "2026-07-20T02:05:29.441Z" }, - { url = "https://files.pythonhosted.org/packages/20/16/e618c875c73e0e39611f20a581b3d5e8d59b8857bf001bee3263044c6deb/yarl-1.24.5-cp312-cp312-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:90333fd89b43c0d08ac85f3f1447593fc2c66de18c3d6378d7125ea118dc7a54", size = 107741, upload-time = "2026-07-20T02:05:31.367Z" }, - { url = "https://files.pythonhosted.org/packages/d9/9a/c4defeaf3ed33fcb346aacf9c6e971a8d4e2bde04a0310e79abb208e7965/yarl-1.24.5-cp312-cp312-manylinux2014_armv7l.manylinux_2_17_armv7l.manylinux_2_31_armv7l.whl", hash = "sha256:665b0a2c463cc9423dd647e0bfd9f4ccc9b50f768c55304d5e9f80b177c1de12", size = 103570, upload-time = "2026-07-20T02:05:33.303Z" }, - { url = "https://files.pythonhosted.org/packages/5f/e7/0e0e0de5865ebd5914537ef486f36c727a59865c3ac0cf5ff1b32aececbf/yarl-1.24.5-cp312-cp312-manylinux2014_ppc64le.manylinux_2_17_ppc64le.manylinux_2_28_ppc64le.whl", hash = "sha256:e006d3a974c4ee19512e5f058abedb6eef36a5e553c14812bdeba1758d812e6d", size = 115815, upload-time = "2026-07-20T02:05:35.292Z" }, - { url = "https://files.pythonhosted.org/packages/2b/27/ca56b700cb170aba25a3893b75355b213935657dc5714d2383354a270e62/yarl-1.24.5-cp312-cp312-manylinux2014_s390x.manylinux_2_17_s390x.manylinux_2_28_s390x.whl", hash = "sha256:e7d42c531243450ef0d4d9c172e7ed6ef052640f195629065041b5add4e058d1", size = 116025, upload-time = "2026-07-20T02:05:37.503Z" }, - { url = "https://files.pythonhosted.org/packages/d6/d0/d56c859b8222116f5d68459199f48359e0bf121b6f65a69bf329b3602ba0/yarl-1.24.5-cp312-cp312-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:f08c7513ecef5aad65687bfdf6bc601ae9fccd04a42904501f8f7141abad9eb9", size = 109835, upload-time = "2026-07-20T02:05:39.506Z" }, - { url = "https://files.pythonhosted.org/packages/70/a2/3a35557e4d1a79425040eba202ccaf08bdc8717680fc77e2498a1ad2e0a5/yarl-1.24.5-cp312-cp312-manylinux_2_31_riscv64.manylinux_2_39_riscv64.whl", hash = "sha256:6c95b17fe34ed802f17e205112e6e10db92275c34fee290aa9bdc55a9c724027", size = 108884, upload-time = "2026-07-20T02:05:41.584Z" }, - { url = "https://files.pythonhosted.org/packages/e4/35/ef4c26356b7913c68983bac2d72a4212b3347af551cb8d250b99b5ed7b7f/yarl-1.24.5-cp312-cp312-musllinux_1_2_aarch64.whl", hash = "sha256:56b149b22de33b23b0c6077ab9518c6dcb538ad462e1830e68d06591ccf6e38b", size = 107308, upload-time = "2026-07-20T02:05:43.697Z" }, - { url = "https://files.pythonhosted.org/packages/d5/91/ff0dc66c2ccf3e0153ab97ff61eabab4400e6a5264af427ab30cd69f1857/yarl-1.24.5-cp312-cp312-musllinux_1_2_armv7l.whl", hash = "sha256:a8fe66b8f300da93798025a785a5b90b42f3810dc2b72283ff84a41aaaebc293", size = 103646, upload-time = "2026-07-20T02:05:45.895Z" }, - { url = "https://files.pythonhosted.org/packages/74/f0/33b9271c7f881766359d58266fa0811d2e5210ed860e28da7dc6d7786344/yarl-1.24.5-cp312-cp312-musllinux_1_2_ppc64le.whl", hash = "sha256:377fe3732edbaf78ee74efdf2c9f49f6e99f20e7f9d2649fda3eb4badd77d76e", size = 115305, upload-time = "2026-07-20T02:05:47.832Z" }, - { url = "https://files.pythonhosted.org/packages/ef/65/fd79fb1868c4a80db8661091de525bf430f63c3bea1b20e8b6a84fc7d359/yarl-1.24.5-cp312-cp312-musllinux_1_2_riscv64.whl", hash = "sha256:e8ffa78582120024f476a611d7befc123cee59e47e8309d470cf667d806e613b", size = 108404, upload-time = "2026-07-20T02:05:49.604Z" }, - { url = "https://files.pythonhosted.org/packages/ff/ba/dbabe6b262f17a816c70cfc09558dbf03ece3ec76684d02f911a3d3a189c/yarl-1.24.5-cp312-cp312-musllinux_1_2_s390x.whl", hash = "sha256:daba5e594f06114e37db186efd2dd916609071e59daca901a0a2e71f02b142ce", size = 115940, upload-time = "2026-07-20T02:05:51.741Z" }, - { url = "https://files.pythonhosted.org/packages/a5/43/fab2d1dad9d340a268cdde63756a123d069723efff6a372d123fa74a9517/yarl-1.24.5-cp312-cp312-musllinux_1_2_x86_64.whl", hash = "sha256:65be18ec59496c13908f02a2472751d9ef840b4f3fb5726f129306bf6a2a7bba", size = 110006, upload-time = "2026-07-20T02:05:53.554Z" }, - { url = "https://files.pythonhosted.org/packages/c4/27/41eb51bbd1b8d89546b83897cfb0164f1e109304fd408dbb151b639eec0f/yarl-1.24.5-cp312-cp312-win_amd64.whl", hash = "sha256:a929d878fec099030c292803b31e5d5540a7b6a31e6a3cc76cb4685fc2a2f51b", size = 97618, upload-time = "2026-07-20T02:05:55.57Z" }, - { url = "https://files.pythonhosted.org/packages/3c/25/b2553764b3d65db711d8f45416351ec4f420847558eb669edcbcaadf5780/yarl-1.24.5-cp312-cp312-win_arm64.whl", hash = "sha256:7ce27823052e2013b597e0c738b13e7e36b8ccb9400df8959417b052ab0fd92c", size = 93018, upload-time = "2026-07-20T02:05:57.554Z" }, - { url = "https://files.pythonhosted.org/packages/61/02/962c1cbfc401a30c1d034dc67ff395f64b52302c6d62de556c1fca99acc0/yarl-1.24.5-py3-none-any.whl", hash = "sha256:a33700d13d9b7d84fd10947b09ff69fb9a792e519c8cb9764a3ca70baa6c23a7", size = 58612, upload-time = "2026-07-20T02:07:43.461Z" }, +sdist = { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/yarl/1.24.5/yarl-1.24.5.tar.gz", hash = "sha256:e81b83143bee16329c23db3c1b2d82b29892fcbcb849186d2f6e98a5abe9a57f", size = 228679, upload-time = "2026-07-20T02:07:45.435Z" } +wheels = [ + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/yarl/1.24.5/yarl-1.24.5-cp311-cp311-macosx_10_9_universal2.whl", hash = "sha256:2c1fe720934a16ea8e7146175cba2126f87f54912c8c5435e7f7c7a51ef808d3", size = 135043, upload-time = "2026-07-20T02:04:52.39Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/yarl/1.24.5/yarl-1.24.5-cp311-cp311-macosx_10_9_x86_64.whl", hash = "sha256:c687ed078e145f5fd53a14854beff320e1d2ab76df03e2009c98f39a0f68f39a", size = 96942, upload-time = "2026-07-20T02:04:54.26Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/yarl/1.24.5/yarl-1.24.5-cp311-cp311-macosx_11_0_arm64.whl", hash = "sha256:709f1efed56c4a145793c046cd4939f9959bcd818979a787b77d8e09c57a0840", size = 97046, upload-time = "2026-07-20T02:04:56.638Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/yarl/1.24.5/yarl-1.24.5-cp311-cp311-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:874019bd513008b009f58657134e5d0c5e030b3559bd0553976837adf52fe966", size = 110512, upload-time = "2026-07-20T02:04:58.955Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/yarl/1.24.5/yarl-1.24.5-cp311-cp311-manylinux2014_armv7l.manylinux_2_17_armv7l.manylinux_2_31_armv7l.whl", hash = "sha256:a4582acf7ef76482f6f511ebaf1946dae7f2e85ec4728b81a678c01df63bd723", size = 102454, upload-time = "2026-07-20T02:05:00.623Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/yarl/1.24.5/yarl-1.24.5-cp311-cp311-manylinux2014_ppc64le.manylinux_2_17_ppc64le.manylinux_2_28_ppc64le.whl", hash = "sha256:2cabe6546e41dabe439999a23fcb5246e0c3b595b4315b96ef755252be90caeb", size = 117617, upload-time = "2026-07-20T02:05:02.325Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/yarl/1.24.5/yarl-1.24.5-cp311-cp311-manylinux2014_s390x.manylinux_2_17_s390x.manylinux_2_28_s390x.whl", hash = "sha256:17f57620f5475b3c69109376cc87e42a7af5db13c9398e4292772a706ff10780", size = 116135, upload-time = "2026-07-20T02:05:04.05Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/yarl/1.24.5/yarl-1.24.5-cp311-cp311-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:570fec8fbd22b032733625f03f10b7ff023bc399213db15e72a7acaef28c2f4e", size = 111935, upload-time = "2026-07-20T02:05:05.738Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/yarl/1.24.5/yarl-1.24.5-cp311-cp311-manylinux_2_31_riscv64.manylinux_2_39_riscv64.whl", hash = "sha256:5fede79c6f73ff2c3ef822864cb1ada23196e62756df53bc6231d351a49516a2", size = 110010, upload-time = "2026-07-20T02:05:07.471Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/yarl/1.24.5/yarl-1.24.5-cp311-cp311-musllinux_1_2_aarch64.whl", hash = "sha256:8ccf9aca873b767977c73df497a85dbedee4ee086ae9ae49dc461333b9b79f58", size = 110058, upload-time = "2026-07-20T02:05:09.246Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/yarl/1.24.5/yarl-1.24.5-cp311-cp311-musllinux_1_2_armv7l.whl", hash = "sha256:ad5d8201d310b031e6cd839d9bac2d4e5a01533ce5d3d5b50b7de1ef3af1de61", size = 103308, upload-time = "2026-07-20T02:05:11.31Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/yarl/1.24.5/yarl-1.24.5-cp311-cp311-musllinux_1_2_ppc64le.whl", hash = "sha256:841f0852f48fefea3b12c9dfec00704dfa3aef5215d0e3ce564bb3d7cd8d57c6", size = 116898, upload-time = "2026-07-20T02:05:13.099Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/yarl/1.24.5/yarl-1.24.5-cp311-cp311-musllinux_1_2_riscv64.whl", hash = "sha256:9baafc71b04f8f4bb0703b21d6fc9f0c30b346c636a532ff16ec8491a5ea4b1f", size = 109400, upload-time = "2026-07-20T02:05:14.821Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/yarl/1.24.5/yarl-1.24.5-cp311-cp311-musllinux_1_2_s390x.whl", hash = "sha256:d897129df1a22b12aeed2c2c98df0785a2e8e6e0bde87b389491d0025c187077", size = 115934, upload-time = "2026-07-20T02:05:17.78Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/yarl/1.24.5/yarl-1.24.5-cp311-cp311-musllinux_1_2_x86_64.whl", hash = "sha256:dd625535328fd9882374356269227670189adfcc6a2d90284f323c05862eecbd", size = 112178, upload-time = "2026-07-20T02:05:19.675Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/yarl/1.24.5/yarl-1.24.5-cp311-cp311-win_amd64.whl", hash = "sha256:f4239bbec5a3577ddb49e4b50aeb32d8e5792098262ae2f63723f916a29b1a25", size = 97544, upload-time = "2026-07-20T02:05:21.523Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/yarl/1.24.5/yarl-1.24.5-cp311-cp311-win_arm64.whl", hash = "sha256:3ac6aff147deb9c09461b2d4bbdf6256831198f5d8a23f5d37138213090b6d8a", size = 93359, upload-time = "2026-07-20T02:05:23.493Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/yarl/1.24.5/yarl-1.24.5-cp312-cp312-macosx_10_13_universal2.whl", hash = "sha256:d693396e5aea78db03decd60aec9ece16c9b40ba00a587f089615ff4e718a81d", size = 136035, upload-time = "2026-07-20T02:05:25.489Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/yarl/1.24.5/yarl-1.24.5-cp312-cp312-macosx_10_13_x86_64.whl", hash = "sha256:3363fcc96e665878946ad7a106b9a13eac0541766a690ef287c0232ac768b6ec", size = 97642, upload-time = "2026-07-20T02:05:27.429Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/yarl/1.24.5/yarl-1.24.5-cp312-cp312-macosx_11_0_arm64.whl", hash = "sha256:9d399bdcfb4a0f659b9b3788bbc89babe63d9a6a65aacdf4d4e7065ff2e6316c", size = 97323, upload-time = "2026-07-20T02:05:29.441Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/yarl/1.24.5/yarl-1.24.5-cp312-cp312-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:90333fd89b43c0d08ac85f3f1447593fc2c66de18c3d6378d7125ea118dc7a54", size = 107741, upload-time = "2026-07-20T02:05:31.367Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/yarl/1.24.5/yarl-1.24.5-cp312-cp312-manylinux2014_armv7l.manylinux_2_17_armv7l.manylinux_2_31_armv7l.whl", hash = "sha256:665b0a2c463cc9423dd647e0bfd9f4ccc9b50f768c55304d5e9f80b177c1de12", size = 103570, upload-time = "2026-07-20T02:05:33.303Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/yarl/1.24.5/yarl-1.24.5-cp312-cp312-manylinux2014_ppc64le.manylinux_2_17_ppc64le.manylinux_2_28_ppc64le.whl", hash = "sha256:e006d3a974c4ee19512e5f058abedb6eef36a5e553c14812bdeba1758d812e6d", size = 115815, upload-time = "2026-07-20T02:05:35.292Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/yarl/1.24.5/yarl-1.24.5-cp312-cp312-manylinux2014_s390x.manylinux_2_17_s390x.manylinux_2_28_s390x.whl", hash = "sha256:e7d42c531243450ef0d4d9c172e7ed6ef052640f195629065041b5add4e058d1", size = 116025, upload-time = "2026-07-20T02:05:37.503Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/yarl/1.24.5/yarl-1.24.5-cp312-cp312-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:f08c7513ecef5aad65687bfdf6bc601ae9fccd04a42904501f8f7141abad9eb9", size = 109835, upload-time = "2026-07-20T02:05:39.506Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/yarl/1.24.5/yarl-1.24.5-cp312-cp312-manylinux_2_31_riscv64.manylinux_2_39_riscv64.whl", hash = "sha256:6c95b17fe34ed802f17e205112e6e10db92275c34fee290aa9bdc55a9c724027", size = 108884, upload-time = "2026-07-20T02:05:41.584Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/yarl/1.24.5/yarl-1.24.5-cp312-cp312-musllinux_1_2_aarch64.whl", hash = "sha256:56b149b22de33b23b0c6077ab9518c6dcb538ad462e1830e68d06591ccf6e38b", size = 107308, upload-time = "2026-07-20T02:05:43.697Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/yarl/1.24.5/yarl-1.24.5-cp312-cp312-musllinux_1_2_armv7l.whl", hash = "sha256:a8fe66b8f300da93798025a785a5b90b42f3810dc2b72283ff84a41aaaebc293", size = 103646, upload-time = "2026-07-20T02:05:45.895Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/yarl/1.24.5/yarl-1.24.5-cp312-cp312-musllinux_1_2_ppc64le.whl", hash = "sha256:377fe3732edbaf78ee74efdf2c9f49f6e99f20e7f9d2649fda3eb4badd77d76e", size = 115305, upload-time = "2026-07-20T02:05:47.832Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/yarl/1.24.5/yarl-1.24.5-cp312-cp312-musllinux_1_2_riscv64.whl", hash = "sha256:e8ffa78582120024f476a611d7befc123cee59e47e8309d470cf667d806e613b", size = 108404, upload-time = "2026-07-20T02:05:49.604Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/yarl/1.24.5/yarl-1.24.5-cp312-cp312-musllinux_1_2_s390x.whl", hash = "sha256:daba5e594f06114e37db186efd2dd916609071e59daca901a0a2e71f02b142ce", size = 115940, upload-time = "2026-07-20T02:05:51.741Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/yarl/1.24.5/yarl-1.24.5-cp312-cp312-musllinux_1_2_x86_64.whl", hash = "sha256:65be18ec59496c13908f02a2472751d9ef840b4f3fb5726f129306bf6a2a7bba", size = 110006, upload-time = "2026-07-20T02:05:53.554Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/yarl/1.24.5/yarl-1.24.5-cp312-cp312-win_amd64.whl", hash = "sha256:a929d878fec099030c292803b31e5d5540a7b6a31e6a3cc76cb4685fc2a2f51b", size = 97618, upload-time = "2026-07-20T02:05:55.57Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/yarl/1.24.5/yarl-1.24.5-cp312-cp312-win_arm64.whl", hash = "sha256:7ce27823052e2013b597e0c738b13e7e36b8ccb9400df8959417b052ab0fd92c", size = 93018, upload-time = "2026-07-20T02:05:57.554Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/yarl/1.24.5/yarl-1.24.5-py3-none-any.whl", hash = "sha256:a33700d13d9b7d84fd10947b09ff69fb9a792e519c8cb9764a3ca70baa6c23a7", size = 58612, upload-time = "2026-07-20T02:07:43.461Z" }, ] [[package]] name = "zstandard" version = "0.25.0" -source = { registry = "https://pypi.org/simple" } -sdist = { url = "https://files.pythonhosted.org/packages/fd/aa/3e0508d5a5dd96529cdc5a97011299056e14c6505b678fd58938792794b1/zstandard-0.25.0.tar.gz", hash = "sha256:7713e1179d162cf5c7906da876ec2ccb9c3a9dcbdffef0cc7f70c3667a205f0b", size = 711513, upload-time = "2025-09-14T22:15:54.002Z" } -wheels = [ - { url = "https://files.pythonhosted.org/packages/2a/83/c3ca27c363d104980f1c9cee1101cc8ba724ac8c28a033ede6aab89585b1/zstandard-0.25.0-cp311-cp311-macosx_10_9_x86_64.whl", hash = "sha256:933b65d7680ea337180733cf9e87293cc5500cc0eb3fc8769f4d3c88d724ec5c", size = 795254, upload-time = "2025-09-14T22:16:26.137Z" }, - { url = "https://files.pythonhosted.org/packages/ac/4d/e66465c5411a7cf4866aeadc7d108081d8ceba9bc7abe6b14aa21c671ec3/zstandard-0.25.0-cp311-cp311-macosx_11_0_arm64.whl", hash = "sha256:a3f79487c687b1fc69f19e487cd949bf3aae653d181dfb5fde3bf6d18894706f", size = 640559, upload-time = "2025-09-14T22:16:27.973Z" }, - { url = "https://files.pythonhosted.org/packages/12/56/354fe655905f290d3b147b33fe946b0f27e791e4b50a5f004c802cb3eb7b/zstandard-0.25.0-cp311-cp311-manylinux2010_i686.manylinux2014_i686.manylinux_2_12_i686.manylinux_2_17_i686.whl", hash = "sha256:0bbc9a0c65ce0eea3c34a691e3c4b6889f5f3909ba4822ab385fab9057099431", size = 5348020, upload-time = "2025-09-14T22:16:29.523Z" }, - { url = "https://files.pythonhosted.org/packages/3b/13/2b7ed68bd85e69a2069bcc72141d378f22cae5a0f3b353a2c8f50ef30c1b/zstandard-0.25.0-cp311-cp311-manylinux2014_aarch64.manylinux_2_17_aarch64.whl", hash = "sha256:01582723b3ccd6939ab7b3a78622c573799d5d8737b534b86d0e06ac18dbde4a", size = 5058126, upload-time = "2025-09-14T22:16:31.811Z" }, - { url = "https://files.pythonhosted.org/packages/c9/dd/fdaf0674f4b10d92cb120ccff58bbb6626bf8368f00ebfd2a41ba4a0dc99/zstandard-0.25.0-cp311-cp311-manylinux2014_ppc64le.manylinux_2_17_ppc64le.whl", hash = "sha256:5f1ad7bf88535edcf30038f6919abe087f606f62c00a87d7e33e7fc57cb69fcc", size = 5405390, upload-time = "2025-09-14T22:16:33.486Z" }, - { url = "https://files.pythonhosted.org/packages/0f/67/354d1555575bc2490435f90d67ca4dd65238ff2f119f30f72d5cde09c2ad/zstandard-0.25.0-cp311-cp311-manylinux2014_s390x.manylinux_2_17_s390x.whl", hash = "sha256:06acb75eebeedb77b69048031282737717a63e71e4ae3f77cc0c3b9508320df6", size = 5452914, upload-time = "2025-09-14T22:16:35.277Z" }, - { url = "https://files.pythonhosted.org/packages/bb/1f/e9cfd801a3f9190bf3e759c422bbfd2247db9d7f3d54a56ecde70137791a/zstandard-0.25.0-cp311-cp311-manylinux2014_x86_64.manylinux_2_17_x86_64.whl", hash = "sha256:9300d02ea7c6506f00e627e287e0492a5eb0371ec1670ae852fefffa6164b072", size = 5559635, upload-time = "2025-09-14T22:16:37.141Z" }, - { url = "https://files.pythonhosted.org/packages/21/88/5ba550f797ca953a52d708c8e4f380959e7e3280af029e38fbf47b55916e/zstandard-0.25.0-cp311-cp311-musllinux_1_1_aarch64.whl", hash = "sha256:bfd06b1c5584b657a2892a6014c2f4c20e0db0208c159148fa78c65f7e0b0277", size = 5048277, upload-time = "2025-09-14T22:16:38.807Z" }, - { url = "https://files.pythonhosted.org/packages/46/c0/ca3e533b4fa03112facbe7fbe7779cb1ebec215688e5df576fe5429172e0/zstandard-0.25.0-cp311-cp311-musllinux_1_1_x86_64.whl", hash = "sha256:f373da2c1757bb7f1acaf09369cdc1d51d84131e50d5fa9863982fd626466313", size = 5574377, upload-time = "2025-09-14T22:16:40.523Z" }, - { url = "https://files.pythonhosted.org/packages/12/9b/3fb626390113f272abd0799fd677ea33d5fc3ec185e62e6be534493c4b60/zstandard-0.25.0-cp311-cp311-musllinux_1_2_aarch64.whl", hash = "sha256:6c0e5a65158a7946e7a7affa6418878ef97ab66636f13353b8502d7ea03c8097", size = 4961493, upload-time = "2025-09-14T22:16:43.3Z" }, - { url = "https://files.pythonhosted.org/packages/cb/d3/23094a6b6a4b1343b27ae68249daa17ae0651fcfec9ed4de09d14b940285/zstandard-0.25.0-cp311-cp311-musllinux_1_2_i686.whl", hash = "sha256:c8e167d5adf59476fa3e37bee730890e389410c354771a62e3c076c86f9f7778", size = 5269018, upload-time = "2025-09-14T22:16:45.292Z" }, - { url = "https://files.pythonhosted.org/packages/8c/a7/bb5a0c1c0f3f4b5e9d5b55198e39de91e04ba7c205cc46fcb0f95f0383c1/zstandard-0.25.0-cp311-cp311-musllinux_1_2_ppc64le.whl", hash = "sha256:98750a309eb2f020da61e727de7d7ba3c57c97cf6213f6f6277bb7fb42a8e065", size = 5443672, upload-time = "2025-09-14T22:16:47.076Z" }, - { url = "https://files.pythonhosted.org/packages/27/22/503347aa08d073993f25109c36c8d9f029c7d5949198050962cb568dfa5e/zstandard-0.25.0-cp311-cp311-musllinux_1_2_s390x.whl", hash = "sha256:22a086cff1b6ceca18a8dd6096ec631e430e93a8e70a9ca5efa7561a00f826fa", size = 5822753, upload-time = "2025-09-14T22:16:49.316Z" }, - { url = "https://files.pythonhosted.org/packages/e2/be/94267dc6ee64f0f8ba2b2ae7c7a2df934a816baaa7291db9e1aa77394c3c/zstandard-0.25.0-cp311-cp311-musllinux_1_2_x86_64.whl", hash = "sha256:72d35d7aa0bba323965da807a462b0966c91608ef3a48ba761678cb20ce5d8b7", size = 5366047, upload-time = "2025-09-14T22:16:51.328Z" }, - { url = "https://files.pythonhosted.org/packages/7b/a3/732893eab0a3a7aecff8b99052fecf9f605cf0fb5fb6d0290e36beee47a4/zstandard-0.25.0-cp311-cp311-win32.whl", hash = "sha256:f5aeea11ded7320a84dcdd62a3d95b5186834224a9e55b92ccae35d21a8b63d4", size = 436484, upload-time = "2025-09-14T22:16:55.005Z" }, - { url = "https://files.pythonhosted.org/packages/43/a3/c6155f5c1cce691cb80dfd38627046e50af3ee9ddc5d0b45b9b063bfb8c9/zstandard-0.25.0-cp311-cp311-win_amd64.whl", hash = "sha256:daab68faadb847063d0c56f361a289c4f268706b598afbf9ad113cbe5c38b6b2", size = 506183, upload-time = "2025-09-14T22:16:52.753Z" }, - { url = "https://files.pythonhosted.org/packages/8c/3e/8945ab86a0820cc0e0cdbf38086a92868a9172020fdab8a03ac19662b0e5/zstandard-0.25.0-cp311-cp311-win_arm64.whl", hash = "sha256:22a06c5df3751bb7dc67406f5374734ccee8ed37fc5981bf1ad7041831fa1137", size = 462533, upload-time = "2025-09-14T22:16:53.878Z" }, - { url = "https://files.pythonhosted.org/packages/82/fc/f26eb6ef91ae723a03e16eddb198abcfce2bc5a42e224d44cc8b6765e57e/zstandard-0.25.0-cp312-cp312-macosx_10_13_x86_64.whl", hash = "sha256:7b3c3a3ab9daa3eed242d6ecceead93aebbb8f5f84318d82cee643e019c4b73b", size = 795738, upload-time = "2025-09-14T22:16:56.237Z" }, - { url = "https://files.pythonhosted.org/packages/aa/1c/d920d64b22f8dd028a8b90e2d756e431a5d86194caa78e3819c7bf53b4b3/zstandard-0.25.0-cp312-cp312-macosx_11_0_arm64.whl", hash = "sha256:913cbd31a400febff93b564a23e17c3ed2d56c064006f54efec210d586171c00", size = 640436, upload-time = "2025-09-14T22:16:57.774Z" }, - { url = "https://files.pythonhosted.org/packages/53/6c/288c3f0bd9fcfe9ca41e2c2fbfd17b2097f6af57b62a81161941f09afa76/zstandard-0.25.0-cp312-cp312-manylinux2010_i686.manylinux2014_i686.manylinux_2_12_i686.manylinux_2_17_i686.whl", hash = "sha256:011d388c76b11a0c165374ce660ce2c8efa8e5d87f34996aa80f9c0816698b64", size = 5343019, upload-time = "2025-09-14T22:16:59.302Z" }, - { url = "https://files.pythonhosted.org/packages/1e/15/efef5a2f204a64bdb5571e6161d49f7ef0fffdbca953a615efbec045f60f/zstandard-0.25.0-cp312-cp312-manylinux2014_aarch64.manylinux_2_17_aarch64.whl", hash = "sha256:6dffecc361d079bb48d7caef5d673c88c8988d3d33fb74ab95b7ee6da42652ea", size = 5063012, upload-time = "2025-09-14T22:17:01.156Z" }, - { url = "https://files.pythonhosted.org/packages/b7/37/a6ce629ffdb43959e92e87ebdaeebb5ac81c944b6a75c9c47e300f85abdf/zstandard-0.25.0-cp312-cp312-manylinux2014_ppc64le.manylinux_2_17_ppc64le.whl", hash = "sha256:7149623bba7fdf7e7f24312953bcf73cae103db8cae49f8154dd1eadc8a29ecb", size = 5394148, upload-time = "2025-09-14T22:17:03.091Z" }, - { url = "https://files.pythonhosted.org/packages/e3/79/2bf870b3abeb5c070fe2d670a5a8d1057a8270f125ef7676d29ea900f496/zstandard-0.25.0-cp312-cp312-manylinux2014_s390x.manylinux_2_17_s390x.whl", hash = "sha256:6a573a35693e03cf1d67799fd01b50ff578515a8aeadd4595d2a7fa9f3ec002a", size = 5451652, upload-time = "2025-09-14T22:17:04.979Z" }, - { url = "https://files.pythonhosted.org/packages/53/60/7be26e610767316c028a2cbedb9a3beabdbe33e2182c373f71a1c0b88f36/zstandard-0.25.0-cp312-cp312-manylinux2014_x86_64.manylinux_2_17_x86_64.whl", hash = "sha256:5a56ba0db2d244117ed744dfa8f6f5b366e14148e00de44723413b2f3938a902", size = 5546993, upload-time = "2025-09-14T22:17:06.781Z" }, - { url = "https://files.pythonhosted.org/packages/85/c7/3483ad9ff0662623f3648479b0380d2de5510abf00990468c286c6b04017/zstandard-0.25.0-cp312-cp312-musllinux_1_1_aarch64.whl", hash = "sha256:10ef2a79ab8e2974e2075fb984e5b9806c64134810fac21576f0668e7ea19f8f", size = 5046806, upload-time = "2025-09-14T22:17:08.415Z" }, - { url = "https://files.pythonhosted.org/packages/08/b3/206883dd25b8d1591a1caa44b54c2aad84badccf2f1de9e2d60a446f9a25/zstandard-0.25.0-cp312-cp312-musllinux_1_1_x86_64.whl", hash = "sha256:aaf21ba8fb76d102b696781bddaa0954b782536446083ae3fdaa6f16b25a1c4b", size = 5576659, upload-time = "2025-09-14T22:17:10.164Z" }, - { url = "https://files.pythonhosted.org/packages/9d/31/76c0779101453e6c117b0ff22565865c54f48f8bd807df2b00c2c404b8e0/zstandard-0.25.0-cp312-cp312-musllinux_1_2_aarch64.whl", hash = "sha256:1869da9571d5e94a85a5e8d57e4e8807b175c9e4a6294e3b66fa4efb074d90f6", size = 4953933, upload-time = "2025-09-14T22:17:11.857Z" }, - { url = "https://files.pythonhosted.org/packages/18/e1/97680c664a1bf9a247a280a053d98e251424af51f1b196c6d52f117c9720/zstandard-0.25.0-cp312-cp312-musllinux_1_2_i686.whl", hash = "sha256:809c5bcb2c67cd0ed81e9229d227d4ca28f82d0f778fc5fea624a9def3963f91", size = 5268008, upload-time = "2025-09-14T22:17:13.627Z" }, - { url = "https://files.pythonhosted.org/packages/1e/73/316e4010de585ac798e154e88fd81bb16afc5c5cb1a72eeb16dd37e8024a/zstandard-0.25.0-cp312-cp312-musllinux_1_2_ppc64le.whl", hash = "sha256:f27662e4f7dbf9f9c12391cb37b4c4c3cb90ffbd3b1fb9284dadbbb8935fa708", size = 5433517, upload-time = "2025-09-14T22:17:16.103Z" }, - { url = "https://files.pythonhosted.org/packages/5b/60/dd0f8cfa8129c5a0ce3ea6b7f70be5b33d2618013a161e1ff26c2b39787c/zstandard-0.25.0-cp312-cp312-musllinux_1_2_s390x.whl", hash = "sha256:99c0c846e6e61718715a3c9437ccc625de26593fea60189567f0118dc9db7512", size = 5814292, upload-time = "2025-09-14T22:17:17.827Z" }, - { url = "https://files.pythonhosted.org/packages/fc/5f/75aafd4b9d11b5407b641b8e41a57864097663699f23e9ad4dbb91dc6bfe/zstandard-0.25.0-cp312-cp312-musllinux_1_2_x86_64.whl", hash = "sha256:474d2596a2dbc241a556e965fb76002c1ce655445e4e3bf38e5477d413165ffa", size = 5360237, upload-time = "2025-09-14T22:17:19.954Z" }, - { url = "https://files.pythonhosted.org/packages/ff/8d/0309daffea4fcac7981021dbf21cdb2e3427a9e76bafbcdbdf5392ff99a4/zstandard-0.25.0-cp312-cp312-win32.whl", hash = "sha256:23ebc8f17a03133b4426bcc04aabd68f8236eb78c3760f12783385171b0fd8bd", size = 436922, upload-time = "2025-09-14T22:17:24.398Z" }, - { url = "https://files.pythonhosted.org/packages/79/3b/fa54d9015f945330510cb5d0b0501e8253c127cca7ebe8ba46a965df18c5/zstandard-0.25.0-cp312-cp312-win_amd64.whl", hash = "sha256:ffef5a74088f1e09947aecf91011136665152e0b4b359c42be3373897fb39b01", size = 506276, upload-time = "2025-09-14T22:17:21.429Z" }, - { url = "https://files.pythonhosted.org/packages/ea/6b/8b51697e5319b1f9ac71087b0af9a40d8a6288ff8025c36486e0c12abcc4/zstandard-0.25.0-cp312-cp312-win_arm64.whl", hash = "sha256:181eb40e0b6a29b3cd2849f825e0fa34397f649170673d385f3598ae17cca2e9", size = 462679, upload-time = "2025-09-14T22:17:23.147Z" }, +source = { registry = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/" } +sdist = { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/zstandard/0.25.0/zstandard-0.25.0.tar.gz", hash = "sha256:7713e1179d162cf5c7906da876ec2ccb9c3a9dcbdffef0cc7f70c3667a205f0b", size = 711513, upload-time = "2025-09-14T22:15:54.002Z" } +wheels = [ + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/zstandard/0.25.0/zstandard-0.25.0-cp311-cp311-macosx_10_9_x86_64.whl", hash = "sha256:933b65d7680ea337180733cf9e87293cc5500cc0eb3fc8769f4d3c88d724ec5c", size = 795254, upload-time = "2025-09-14T22:16:26.137Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/zstandard/0.25.0/zstandard-0.25.0-cp311-cp311-macosx_11_0_arm64.whl", hash = "sha256:a3f79487c687b1fc69f19e487cd949bf3aae653d181dfb5fde3bf6d18894706f", size = 640559, upload-time = "2025-09-14T22:16:27.973Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/zstandard/0.25.0/zstandard-0.25.0-cp311-cp311-manylinux2010_i686.manylinux2014_i686.manylinux_2_12_i686.manylinux_2_17_i686.whl", hash = "sha256:0bbc9a0c65ce0eea3c34a691e3c4b6889f5f3909ba4822ab385fab9057099431", size = 5348020, upload-time = "2025-09-14T22:16:29.523Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/zstandard/0.25.0/zstandard-0.25.0-cp311-cp311-manylinux2014_aarch64.manylinux_2_17_aarch64.whl", hash = "sha256:01582723b3ccd6939ab7b3a78622c573799d5d8737b534b86d0e06ac18dbde4a", size = 5058126, upload-time = "2025-09-14T22:16:31.811Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/zstandard/0.25.0/zstandard-0.25.0-cp311-cp311-manylinux2014_ppc64le.manylinux_2_17_ppc64le.whl", hash = "sha256:5f1ad7bf88535edcf30038f6919abe087f606f62c00a87d7e33e7fc57cb69fcc", size = 5405390, upload-time = "2025-09-14T22:16:33.486Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/zstandard/0.25.0/zstandard-0.25.0-cp311-cp311-manylinux2014_s390x.manylinux_2_17_s390x.whl", hash = "sha256:06acb75eebeedb77b69048031282737717a63e71e4ae3f77cc0c3b9508320df6", size = 5452914, upload-time = "2025-09-14T22:16:35.277Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/zstandard/0.25.0/zstandard-0.25.0-cp311-cp311-manylinux2014_x86_64.manylinux_2_17_x86_64.whl", hash = "sha256:9300d02ea7c6506f00e627e287e0492a5eb0371ec1670ae852fefffa6164b072", size = 5559635, upload-time = "2025-09-14T22:16:37.141Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/zstandard/0.25.0/zstandard-0.25.0-cp311-cp311-musllinux_1_1_aarch64.whl", hash = "sha256:bfd06b1c5584b657a2892a6014c2f4c20e0db0208c159148fa78c65f7e0b0277", size = 5048277, upload-time = "2025-09-14T22:16:38.807Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/zstandard/0.25.0/zstandard-0.25.0-cp311-cp311-musllinux_1_1_x86_64.whl", hash = "sha256:f373da2c1757bb7f1acaf09369cdc1d51d84131e50d5fa9863982fd626466313", size = 5574377, upload-time = "2025-09-14T22:16:40.523Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/zstandard/0.25.0/zstandard-0.25.0-cp311-cp311-musllinux_1_2_aarch64.whl", hash = "sha256:6c0e5a65158a7946e7a7affa6418878ef97ab66636f13353b8502d7ea03c8097", size = 4961493, upload-time = "2025-09-14T22:16:43.3Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/zstandard/0.25.0/zstandard-0.25.0-cp311-cp311-musllinux_1_2_i686.whl", hash = "sha256:c8e167d5adf59476fa3e37bee730890e389410c354771a62e3c076c86f9f7778", size = 5269018, upload-time = "2025-09-14T22:16:45.292Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/zstandard/0.25.0/zstandard-0.25.0-cp311-cp311-musllinux_1_2_ppc64le.whl", hash = "sha256:98750a309eb2f020da61e727de7d7ba3c57c97cf6213f6f6277bb7fb42a8e065", size = 5443672, upload-time = "2025-09-14T22:16:47.076Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/zstandard/0.25.0/zstandard-0.25.0-cp311-cp311-musllinux_1_2_s390x.whl", hash = "sha256:22a086cff1b6ceca18a8dd6096ec631e430e93a8e70a9ca5efa7561a00f826fa", size = 5822753, upload-time = "2025-09-14T22:16:49.316Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/zstandard/0.25.0/zstandard-0.25.0-cp311-cp311-musllinux_1_2_x86_64.whl", hash = "sha256:72d35d7aa0bba323965da807a462b0966c91608ef3a48ba761678cb20ce5d8b7", size = 5366047, upload-time = "2025-09-14T22:16:51.328Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/zstandard/0.25.0/zstandard-0.25.0-cp311-cp311-win32.whl", hash = "sha256:f5aeea11ded7320a84dcdd62a3d95b5186834224a9e55b92ccae35d21a8b63d4", size = 436484, upload-time = "2025-09-14T22:16:55.005Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/zstandard/0.25.0/zstandard-0.25.0-cp311-cp311-win_amd64.whl", hash = "sha256:daab68faadb847063d0c56f361a289c4f268706b598afbf9ad113cbe5c38b6b2", size = 506183, upload-time = "2025-09-14T22:16:52.753Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/zstandard/0.25.0/zstandard-0.25.0-cp311-cp311-win_arm64.whl", hash = "sha256:22a06c5df3751bb7dc67406f5374734ccee8ed37fc5981bf1ad7041831fa1137", size = 462533, upload-time = "2025-09-14T22:16:53.878Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/zstandard/0.25.0/zstandard-0.25.0-cp312-cp312-macosx_10_13_x86_64.whl", hash = "sha256:7b3c3a3ab9daa3eed242d6ecceead93aebbb8f5f84318d82cee643e019c4b73b", size = 795738, upload-time = "2025-09-14T22:16:56.237Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/zstandard/0.25.0/zstandard-0.25.0-cp312-cp312-macosx_11_0_arm64.whl", hash = "sha256:913cbd31a400febff93b564a23e17c3ed2d56c064006f54efec210d586171c00", size = 640436, upload-time = "2025-09-14T22:16:57.774Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/zstandard/0.25.0/zstandard-0.25.0-cp312-cp312-manylinux2010_i686.manylinux2014_i686.manylinux_2_12_i686.manylinux_2_17_i686.whl", hash = "sha256:011d388c76b11a0c165374ce660ce2c8efa8e5d87f34996aa80f9c0816698b64", size = 5343019, upload-time = "2025-09-14T22:16:59.302Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/zstandard/0.25.0/zstandard-0.25.0-cp312-cp312-manylinux2014_aarch64.manylinux_2_17_aarch64.whl", hash = "sha256:6dffecc361d079bb48d7caef5d673c88c8988d3d33fb74ab95b7ee6da42652ea", size = 5063012, upload-time = "2025-09-14T22:17:01.156Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/zstandard/0.25.0/zstandard-0.25.0-cp312-cp312-manylinux2014_ppc64le.manylinux_2_17_ppc64le.whl", hash = "sha256:7149623bba7fdf7e7f24312953bcf73cae103db8cae49f8154dd1eadc8a29ecb", size = 5394148, upload-time = "2025-09-14T22:17:03.091Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/zstandard/0.25.0/zstandard-0.25.0-cp312-cp312-manylinux2014_s390x.manylinux_2_17_s390x.whl", hash = "sha256:6a573a35693e03cf1d67799fd01b50ff578515a8aeadd4595d2a7fa9f3ec002a", size = 5451652, upload-time = "2025-09-14T22:17:04.979Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/zstandard/0.25.0/zstandard-0.25.0-cp312-cp312-manylinux2014_x86_64.manylinux_2_17_x86_64.whl", hash = "sha256:5a56ba0db2d244117ed744dfa8f6f5b366e14148e00de44723413b2f3938a902", size = 5546993, upload-time = "2025-09-14T22:17:06.781Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/zstandard/0.25.0/zstandard-0.25.0-cp312-cp312-musllinux_1_1_aarch64.whl", hash = "sha256:10ef2a79ab8e2974e2075fb984e5b9806c64134810fac21576f0668e7ea19f8f", size = 5046806, upload-time = "2025-09-14T22:17:08.415Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/zstandard/0.25.0/zstandard-0.25.0-cp312-cp312-musllinux_1_1_x86_64.whl", hash = "sha256:aaf21ba8fb76d102b696781bddaa0954b782536446083ae3fdaa6f16b25a1c4b", size = 5576659, upload-time = "2025-09-14T22:17:10.164Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/zstandard/0.25.0/zstandard-0.25.0-cp312-cp312-musllinux_1_2_aarch64.whl", hash = "sha256:1869da9571d5e94a85a5e8d57e4e8807b175c9e4a6294e3b66fa4efb074d90f6", size = 4953933, upload-time = "2025-09-14T22:17:11.857Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/zstandard/0.25.0/zstandard-0.25.0-cp312-cp312-musllinux_1_2_i686.whl", hash = "sha256:809c5bcb2c67cd0ed81e9229d227d4ca28f82d0f778fc5fea624a9def3963f91", size = 5268008, upload-time = "2025-09-14T22:17:13.627Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/zstandard/0.25.0/zstandard-0.25.0-cp312-cp312-musllinux_1_2_ppc64le.whl", hash = "sha256:f27662e4f7dbf9f9c12391cb37b4c4c3cb90ffbd3b1fb9284dadbbb8935fa708", size = 5433517, upload-time = "2025-09-14T22:17:16.103Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/zstandard/0.25.0/zstandard-0.25.0-cp312-cp312-musllinux_1_2_s390x.whl", hash = "sha256:99c0c846e6e61718715a3c9437ccc625de26593fea60189567f0118dc9db7512", size = 5814292, upload-time = "2025-09-14T22:17:17.827Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/zstandard/0.25.0/zstandard-0.25.0-cp312-cp312-musllinux_1_2_x86_64.whl", hash = "sha256:474d2596a2dbc241a556e965fb76002c1ce655445e4e3bf38e5477d413165ffa", size = 5360237, upload-time = "2025-09-14T22:17:19.954Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/zstandard/0.25.0/zstandard-0.25.0-cp312-cp312-win32.whl", hash = "sha256:23ebc8f17a03133b4426bcc04aabd68f8236eb78c3760f12783385171b0fd8bd", size = 436922, upload-time = "2025-09-14T22:17:24.398Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/zstandard/0.25.0/zstandard-0.25.0-cp312-cp312-win_amd64.whl", hash = "sha256:ffef5a74088f1e09947aecf91011136665152e0b4b359c42be3373897fb39b01", size = 506276, upload-time = "2025-09-14T22:17:21.429Z" }, + { url = "https://modal-labs-459781239556.d.codeartifact.us-east-1.amazonaws.com/pypi/modal-pypi/simple/zstandard/0.25.0/zstandard-0.25.0-cp312-cp312-win_arm64.whl", hash = "sha256:181eb40e0b6a29b3cd2849f825e0fa34397f649170673d385f3598ae17cca2e9", size = 462679, upload-time = "2025-09-14T22:17:23.147Z" }, ] From f9c5e6d71f7d521ea602165fc65d8d196b6c1d71 Mon Sep 17 00:00:00 2001 From: kailash Date: Mon, 28 Sep 2026 01:16:06 +0000 Subject: [PATCH 27/27] Record current CPU and GPU deployment validation --- docs/deployment-validation.md | 12 ++++++++++-- 1 file changed, 10 insertions(+), 2 deletions(-) diff --git a/docs/deployment-validation.md b/docs/deployment-validation.md index c88e71d..cccf927 100644 --- a/docs/deployment-validation.md +++ b/docs/deployment-validation.md @@ -12,9 +12,17 @@ Regression coverage includes inherited overrides, independent mutable settings, ## Live GPU checks -The isolated frontend is `lilo-pr55-validation-20260928` in `modal-labs/lilo-deploy`, region `us-west`. It uses uniquely named variants of the 9B LoRA/16K and 4B FFT/64K recipes. Each trainer requests 4 H100s; each inference replica requests one H200 for LoRA or one H100 for FFT. Test pools use a maximum of one replica and a 60-second scale-down window. +The isolated frontend is `lilo-pr55-validation-20260928` in `modal-labs/lilo-deploy`, region `us-west`. It uses uniquely named variants of the 9B LoRA/16K and 4B FFT/64K recipes. Each trainer requests 4 H100s. Both test inference pools use one H100 replica kept warm during validation. LoRA inference was changed from the preset’s H200 request after Modal reported insufficient H200 capacity in `us-west`; the production preset still requests H200. -Live training, publication, sampling, and inference-update continuity results are recorded below when complete. +| Check | Result | +| --- | --- | +| 9B LoRA client A | 3 complete training/publication/sampling steps passed | +| 9B LoRA client B, sharing client A’s trainer | 2 initial steps plus 1 after the inference update passed | +| 4B full-parameter training | 3 complete training/publication/sampling steps passed | +| Inference-only update | SGLang `max_running_requests` changed from 32 to 24; the running worker and frontend lookup both reported 24 | +| Trainer continuity | Client B retained the same trainer instance and boot ID through the inference updates | + +Every step returned finite training and sampling log probabilities. The frontend lookup initially returned the prior configuration during the deployment transition; it was checked again, along with the worker’s actual settings, before releasing the continuation. All completed clients unloaded without reported cleanup errors. The isolated validation apps and pools were stopped afterward. These checks use 512 training tokens and a 16-token generation cap. They exercise the configured processes and interfaces, not maximum-context capacity, convergence, autoscaling under load, or throughput. Other model presets and multi-node topology have CPU coverage only.