Skip to content

Latest commit

 

History

History
133 lines (112 loc) · 6.46 KB

File metadata and controls

133 lines (112 loc) · 6.46 KB

Roadmap

Where lm15-python is, and what is planned before and after the stable 1.0 release. Dates are intentions, not promises; everything here follows the same discipline as the code — see How lm15 is specified.

Where we are (September 2026)

  • 1.2.0 is the current release (2026-09-30): the Claude Code release as a setting, and Claude's own output ceiling as the default max_tokens, on top of 1.1.0 (2026-09-26: four open-model hosts and a Google Cloud pass) and 1.0.1 (2026-09-25), the first stable one (1.0.0 was used by a June upload that was removed; PyPI never reuses a number). The chat core — canonical types, serde, errors, request building, response parsing, streaming — is checked against the pinned language-neutral contract, and frozen for 1.x.
  • Non-chat endpoints, stored-cache resources, live sessions and Chat Completions ingest ship as provisional. Contract tests cover recorded behavior; they do not promise that every provider/account works live. The stability boundary below applies throughout 1.x.
  • The model-string router is available. Tool derivation from functions was removed before 1.0 (2026-09-23): a FunctionTool is written out, in every language. The contract's API-family playbook governs the shared public names; an older router-portability proposal is not the authority for the current implementation.
  • Python, Rust and TypeScript pass the shared corpus at their recorded pins. That is evidence of agreement on the recorded cases, not a release approval or a claim about untested runtime behavior.

What ships in 1.0, and what is stable

Already decided: provisional features ship in 1.0, clearly labeled. This follows the ratified contract's spec/SCOPE.md, not a new release-policy decision.

  • Frozen: the canonical chat types, serialization, errors, request and response mapping, streaming, credential resolution and model listing. Removing or changing frozen behavior requires a major release and maintainer ratification. Additions follow the contract's change process.
  • Provisional: non-chat endpoints (files, batches, image/speech/video generation), stored-cache resources, live sessions and Chat Completions ingest. These are included, but may change incompatibly during 1.x with a contract changes/ entry. Additive changes are preferred, not guaranteed. Pin an exact package version and review change entries before upgrading applications that use them.
  • Out of scope: embeddings and canonical provider-executed computer use. The contract describes the limits of provider-specific passthrough; it is not a stability promise for those features.

The stability of a feature is separate from its test coverage. In particular, typed media inside chat belongs to the frozen chat model; the standalone media-generation endpoints are provisional.

After 1.0

1.0.1 was released on 2026-09-25, ahead of the documentation plan it was waiting for, by the maintainer's decision. Still to do, all additive:

  1. Complete documentation site — guides, API reference, specification pages; several guides and the reference are still marked unfinished.
  2. User-experience review pass — read the docs as a new user would. Changes to the frozen chat core must be additive; provisional surfaces may still move.
  3. Scope labels — keep the provisional notices on the relevant guides and release notes.
  4. Release engineering, in place — tag-driven publishing via PyPI trusted publishing (OIDC); CI on Python 3.10–3.14 × Linux, macOS and Windows; a type-checking gate (typecheck/README.md).

Provider coverage

Today, with identical canonical behavior and live-receipt fixtures:

  • OpenAI (Responses API) and OpenAI Codex, including GA Realtime (live) sessions, Sora video, images, and speech
  • Anthropic and Claude Code
  • Google Gemini, including Live (WebSocket) sessions, Veo video, images, and speech
  • xAI (Grok), including subscription OAuth, images, and grok-imagine video
  • Any Chat Completions–compatible server through one dialect adapter with typed compatibility policies — Groq, OpenRouter, DeepSeek, vLLM, SGLang, Ollama

Azure, Bedrock and Vertex already have contract coverage; see cloud hosts for their distinct credentials and settings. Provider/model availability and live verification remain separate from passing recorded cases. Additional providers require evidence-backed fixtures before implementation.

Layers above the foundation

lm15 is deliberately low-level: no automatic tool loop, no retries, no cost ledger, no policy routing (the shipped router is a lookup table — no fallbacks, no ranking). Several companion pieces are under consideration once the foundation's user experience is validated — each as a separate package built on the frozen core, none of them contract-governed:

  • An ergonomic layer — a concise call()-style interface, automatic tool loops, retry/fallback patterns, for people who want three lines and sensible defaults.
  • A model catalog — maintained pricing, context-window, and capability metadata via the entry-point protocol already specified in model-hydration, enabling cost estimation and routing.
  • Recipes — cookbook pages for everything the core deliberately omits (retries, fallback, budget caps, proxying), so each "lm15 doesn't do X" has a one-page answer.

Multi-language

  • Keep Python, Rust and TypeScript aligned with their contract pins and runtime tests. Other ports must meet the same gates before claiming parity; the corpus grows, so use the current reports rather than a historical fixed check count.
  • Publish them (crates.io, Go module, npm) once they pass the full corpus — never before.
  • The promise stays the same in every language: byte-identical wire requests, identical canonical parses, one spec.

Ecosystem and community

  • Integration examples: a FastAPI service, an agent loop, notebooks, and migration guides from other clients.
  • A fixture-first "add a provider" contributor path (see CONTRIBUTING).
  • Benchmarks stay machine-generated and re-run on a schedule — numbers in the README and on this site are never hand-edited.