Skip to content

Epic: Make time-to-answer a product-level performance contract #333

Description

@Joncallim

Goal

Make “total speed” a product property, not an impression.

DockerMap should publish the first useful Docker answer as soon as authoritative Docker evidence exists, keep optional/enrichment work off that critical path, avoid refetch/recompute work when only unrelated evidence changes, and mechanically prevent meaningful regressions.

Product outcome

For the target single-host Compose operator, DockerMap should feel immediate in four different senses:

  • startup speed — the UI/API becomes reachable before expensive collection completes;
  • evidence speed — Docker facts publish independently of slower Compose/host enrichment;
  • reaction speed — a daemon revision reaches the browser with minimal avoidable latency;
  • cognitive speed — common questions resolve with little navigation or recomputation.

This epic is about measured time-to-answer, not synthetic benchmark chasing.

Current-main findings motivating this epic

  • Docker/host provider collection is already well decoupled and bounded.
  • Compose runtime correlation still runs inside the five-second Docker snapshot publication budget.
  • The daemon waits for the initial refresh_cache() before binding its HTTP listener.
  • Independent Docker inventory reads are currently sequential.
  • Browser model refreshes refetch Snapshot + Runtime Map + Findings together when the broad model revision changes, even when only one facet changed.
  • Node SSE polls daemon health on a fixed interval, adding avoidable revision-notification latency and per-stream daemon work.
  • Home still renders the legacy force-layout Service Map; its layout is expensive relative to a summary surface and can rerun on unrelated model revisions.
  • Atlas has controlled performance evidence, but DockerMap as a whole has no end-to-end performance authority.

Sequence

Blocked only by the small authority cleanup in #332 for trustworthy baselines/docs.

Recommended child order:

  1. establish the end-to-end performance harness and baseline;
  2. make first Docker publication independent and fast;
  3. add revision-scoped selective refresh / lower notification latency;
  4. remove avoidable frontend work and bundle cost;
  5. rerun certification and freeze budgets.

Children may overlap only where they do not invalidate the controlled baseline environment.

Non-negotiable constraints

  • Do not weaken modelRevision/source/provenance coherence for speed.
  • Do not make stale provider evidence look fresh.
  • No raw Docker socket fallback.
  • No new runtime telemetry solely to benchmark production users.
  • Optional provider/Compose work may enrich the current model later; it may not delay publication of already-authoritative Docker facts.
  • Performance tests must use controlled fixtures/runners; ordinary CI wall-clock timing is not a reliable gate.
  • No feature removal merely to improve a number without a product decision.

Acceptance criteria

  • A controlled performance contract covers process start → listener, listener → first useful Docker publication, Docker collection, daemon publication → browser awareness, browser update → Review render, buildModel, Findings derivation, Home render, and command/search response.
  • First Docker publication no longer waits for Compose filesystem projection.
  • Daemon HTTP availability no longer waits for the first collection pass.
  • Independent Docker inventory reads are concurrent or there is measured evidence that concurrency is unsafe/not useful.
  • Browser refresh work is scoped to changed semantic facets while preserving atomic/coherent presentation.
  • Per-client SSE streams do not each create unnecessary daemon polling work; revision notification latency is measured and bounded.
  • Provider-only changes do not cause avoidable legacy topology-layout recomputation.
  • Production frontend has a reviewed bundle/code-loading budget.
  • 25/100/250-container reference fixtures remain responsive and truthful.
  • Full security, contract, live-Docker and production-image gates remain green.

Non-goals

  • Multi-host.
  • Monitoring-grade CPU/memory metrics.
  • Kubernetes.
  • Replacing the Atlas/Service Map product decision.
  • Relaxing security/redaction/coherence boundaries.
  • Premature micro-optimization of rule loops without measured evidence.

Closure evidence

Record:

  • baseline and final controlled measurements;
  • exact environment/source revisions;
  • architecture changes that moved work off the critical path;
  • resource/revision refresh matrix;
  • bundle evidence;
  • 25/100/250 fixture results;
  • live-host qualitative “time to first useful answer” exercise;
  • exact tested SHA and full gates.

Risk: MEDIUM · read_only_product_behavior: true · security_sensitive: true (transport/coherence boundaries) · data_migration: false · routing_mode: stable

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions