Skip to content

Bill allocated linear memory - #3727

Merged
kmatasfp merged 22 commits into
mainfrom
kaurmatas/gol-113-memory-billing-from-allocated-linear-memory
Aug 7, 2026
Merged

kmatasfp merged 22 commits into
mainfrom
kaurmatas/gol-113-memory-billing-from-allocated-linear-memory

Conversation

@kmatasfp

@kmatasfp kmatasfp commented Aug 6, 2026 •

Copy link
Copy Markdown
Contributor

Implements GOL-113.

Summary

  • add memory plan defaults, ceilings, account overrides, usage APIs, CLI output, registry accumulation, and protobuf persistence
  • meter allocated Wasmtime linear-memory byte-time while a worker owns its concurrent-agent permit, including replay and host waits while excluding reclaimable cached time
  • use one canonical per-worker aggregate for multi-memory limits, admission grants, growth oplog entries, reconstructed status, billing, and metrics
  • add measured-headroom admission with ratio-aware cgroup-v2 probing, coalesced growth persistence, account-scoped fractional remainders, and successful-delivery observability
  • explicitly disable WebAssembly threads/shared memories and reject shared-memory modules

Behavior

Memory meters integrate each worker's allocated linear-memory bytes over the period in which it owns a concurrent-agent permit. This includes guest execution, replay, host calls, host I/O waits, and scheduler waits that occur while the permit remains held, as well as non-durable work that keeps the worker non-reclaimable. Waiting to acquire a permit is not billed. Loaded-idle, warm-runnable, and unloaded workers accrue nothing after permit release. Long-running occupancy is settled incrementally during regular resource-usage batches.

One canonical per-worker tracker sums every unique unshared Wasmtime linear-memory backing, including non-exported memories while deduplicating aliases. Initial allocation, committed memory.grow, aggregate per-worker limits, executor admission grants, oplog/status reconstruction, billing samples, and metrics all use this same total. Growth is charged prospectively: elapsed time is settled at the old size before the committed delta updates the tracker.

Meters retain byte-nanosecond precision locally and transfer final fractional remainders to account scope, allowing usage from short-lived workers to combine into complete GB-seconds. Usage batches keep the existing non-retry behavior after ambiguous transport failures to avoid double charging, and the Prometheus counter advances only for whole GB-seconds successfully delivered to the registry.

The effective maximum memory per worker and monthly memory allowance resolve from plan defaults, ceilings, and active account overrides. Override changes honor configurability, optional expiry, lower/default and upper/ceiling validation. Plan changes clamp overrides above a new ceiling or remove them when the dimension is no longer user-configurable. The per-worker maximum remains a safety cap; billing always uses allocated memory rather than the limit.

The registry and CLI expose current and historical memory usage plus effective per-worker limits and monthly allowances. Protobuf and oplog additions preserve persisted-worker compatibility, and the Wasmtime fork supplies unique-memory enumeration and committed-growth notification. WebAssembly threads and shared memories are explicitly unsupported and rejected; arbitrary components with multiple unshared memories remain supported.

Wasmtime

@kmatasfp
kmatasfp requested a review from a team August 6, 2026 02:19
@netlify

netlify Bot commented Aug 6, 2026 •

Copy link
Copy Markdown

✅ Deploy Preview for golemcloud canceled.

Name Link
🔨 Latest commit da5c353
🔍 Latest deploy log https://app.netlify.com/projects/golemcloud/deploys/6a7618dcc571e70008806227

@github-actions

github-actions Bot commented Aug 6, 2026 •

Copy link
Copy Markdown

📖 Docs preview: https://docs-dlpy2nia3-golem-cloud.vercel.app

Built from commit da5c353b70c21f8df711c6feccc311242ba27728 by docs.yaml.

@kmatasfp

kmatasfp commented Aug 6, 2026 •

Copy link
Copy Markdown
Contributor Author

Updated verification now covers the permit-ownership billing window end to end. The memory-billing integration test verifies that a 512 MiB allocation retained across a three-second std::thread::sleep accrues memory usage, an interrupted 512 MiB workload accrues during replay, and five seconds of loaded-idle time after permit release accrues nothing.

A durable std::thread::sleep remains inside the running guest invocation and retains the ConcurrentAgentPermit. The invocation loop pauses the memory meter and releases the permit only when the worker returns to idle. This distinguishes permit-owning host waits from waiting to acquire a permit, which is not billed.

@kmatasfp
kmatasfp merged commit e2da057 into main Aug 7, 2026
54 checks passed
@kmatasfp
kmatasfp deleted the kaurmatas/gol-113-memory-billing-from-allocated-linear-memory branch August 7, 2026 18:16
@github-actions github-actions Bot locked and limited conversation to collaborators Aug 7, 2026
Sign up for free to subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant