You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Commit df8f49a
Browse filesBrowse the repository at this point in the historyBrowse files
docs: cover reused-process runtimes (AWS Lambda) in deploy.md
StreamableHTTPSessionManager.run() can only run once per instance.
docs/troubleshooting.md already explains this for a Mount swallowing
a lifespan and for several long-running workers. It doesn't cover a
serverless runtime that reuses one warm process across separate
invocations, which hits the same error the moment a container is
reused, the normal case in production. Adds that as a third cause in
deploy.md, with the per-invocation fix and a SnapStart-specific note,
and links it from troubleshooting.md.
Fixes#3590
Copy file name to clipboardExpand all lines: docs/run/deploy.md
+24Lines changed: 24 additions & 0 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -169,6 +169,29 @@ Nothing about the fan-out cares which server object a stream is attached to. Two
169
169
* The bus carries four small typed events, never JSON-RPC. Acknowledgment, filtering, and stream lifecycle stay in the SDK, so your bus cannot break the protocol; it can only move events between processes.
170
170
* Streams are **not** resumable and events are **not** replayed. Losing a replica drops its streams; the clients re-listen and re-fetch. There is no event store to share and nothing else to configure. This is the one place where scaling out is genuinely just more of the same.
171
171
172
+
## Reused-process runtimes (Lambda, and similar)
173
+
174
+
A different problem from having many workers: having one process that gets reused, sequentially, across calls that don't share a lifespan. AWS Lambda is the common case. A container that handled one invocation can be frozen and woken up for the next one, often minutes or hours later, with the same Python process and the same objects still in memory.
175
+
176
+
Build `mcp.streamable_http_app()` once, at import time, the way you would for a normal server, and the first invocation works fine. The second invocation against that same warm container fails on every request:
177
+
178
+
```text
179
+
RuntimeError: StreamableHTTPSessionManager .run() can only be called once per instance. Create a new instance if you need to run again.
180
+
```
181
+
182
+
The manager's lifespan already ran and finished at the end of the first invocation. **[Troubleshooting](../troubleshooting.md)** covers this same error for two other causes; a reused process is the third, and it's easy to miss because a cold container only ever sees one request during local testing.
183
+
184
+
The fix is to build the app inside the handler, per invocation, instead of once at import time:
185
+
186
+
```python
187
+
defhandler(event, context):
188
+
app = mcp.streamable_http_app() # fresh instance, every invocation
189
+
...
190
+
```
191
+
192
+
!!! warning "SnapStart"
193
+
SnapStart's snapshot is taken before any request arrives and deliberately excludes live network connections and running event loops. An app built once at import time and cached across invocations is exactly what that model assumes you won't do. Build fresh per invocation here too.
194
+
172
195
## What the SDK does not give you
173
196
174
197
An `MCPServer` is a protocol implementation, not an application server. The deployment knobs you go looking for next are missing on purpose:
@@ -186,6 +209,7 @@ An `MCPServer` is a protocol implementation, not an application server. The depl
186
209
* The default `requestState` key is `os.urandom(32)`, minted per process. A multi-round-trip retry that reaches a different worker fails with `-32602`*"Invalid or expired requestState"*.
187
210
* The fix is `RequestStateSecurity(keys=[...])`**and** the same server name on every instance. The name is the token's default audience claim. Same keys, same name.
188
211
* Change notifications cross replicas through one shared `SubscriptionBus`. The SDK's only implementation is in-process; the two-method `Protocol` over your own pub/sub is yours to write.
212
+
* A process reused across invocations (Lambda, and similar) hits the same single-use-manager error as a `Mount` swallowing a lifespan or several workers. Build the app fresh inside the handler, per invocation, not once at import time.
189
213
* There is no `workers=`, no health route, no production settings object. Bring your own ASGI server.
190
214
191
215
The other thing a real hostname needs in front of it is a token: **[Authorization](authorization.md)**.
**[Add to an existing app](run/asgi.md)** is the page for this, including several servers in one app and FastAPI. Two neighbouring strings from the same class:
249
249
250
-
*`StreamableHTTPSessionManager .run() can only be called once per instance. Create a new instance if you need to run again.` The manager is single-use; entering the same app's lifespan twice hits it.
250
+
*`StreamableHTTPSessionManager .run() can only be called once per instance. Create a new instance if you need to run again.` The manager is single-use; entering the same app's lifespan twice hits it. A process reused sequentially across separate invocations (AWS Lambda, and similar) hits this too, once a warm container serves its second request. **[Reused-process runtimes](run/deploy.md#reused-process-runtimes-lambda-and-similar)** is that case specifically.
251
251
*`mcp.session_manager` only exists **after**`streamable_http_app()` has been called, so build the routes first and touch the manager only inside the lifespan.
0 commit comments