docs: add a middleware tutorial that refuses tool calls - #3613
Conversation
The Middleware page listed "Refuse" as something a middleware can do but had no example of its own for it. Add one: a middleware factory that caps how many tool calls run at once and answers the rest with a JSON-RPC error, leaving every other method alone. - docs_src/middleware/tutorial002.py: `max_concurrent_tool_calls(limit)` on the low-level `Server`, with an application-defined `SERVER_BUSY` error code. - docs/advanced/middleware.md: a new "A concurrency cap" section walks through it, and the "Refuse" bullet now points at it as well as at the subscriptions example. - tests/docs_src/test_middleware.py: the cap refuses the call over the limit while the earlier ones are still running, other requests are still answered, and a call that finishes, raises or is cancelled gives its slot back. Fixes #3272
📚 Documentation preview
|
There was a problem hiding this comment.
3 issues found across 3 files
Prompt for AI agents (unresolved issues)
Check if these issues are valid — if so, understand the root cause of each and fix them. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. If appropriate, use sub-agents to investigate and fix each issue separately.
<file name="docs/advanced/middleware.md">
<violation number="1" location="docs/advanced/middleware.md:87">
P3: The Refuse bullet contains the fragment `The concurrency cap above.`, so the new cross-reference reads as incomplete. Change it to `This is the concurrency cap above.`.</violation>
</file>
<file name="docs_src/middleware/tutorial002.py">
<violation number="1" location="docs_src/middleware/tutorial002.py:36">
P3: `on_call_tool` raises an unhandled `KeyError: 'query'` when a client omits the argument. `CallToolRequestParams.arguments` is optional, and the SDK only validates the request shape, never the tool's declared `input_schema`, so a call like `search_books` with no `arguments` gets past validation and becomes a generic `-32603` internal error that leaks the handler internals — the opposite of the clean, explained flows this tutorial teaches. Use `.get("query", "")`, or validate and raise a clean `MCPError(INVALID_PARAMS, ...)` since this page is about error handling.</violation>
<violation number="2" location="docs_src/middleware/tutorial002.py:37">
P2: This sample handler never awaits any work, so it releases `running` before another request can enter and the included server never reaches its limit of four. Make the example hold or await realistic work, or explain that the cap only becomes observable when tool handlers overlap.</violation>
</file>
Reply with feedback, questions, or to request a fix.
Re-trigger cubic
|
|
||
| async def on_call_tool(ctx: ServerRequestContext, params: CallToolRequestParams) -> CallToolResult: | ||
| query = (params.arguments or {})["query"] | ||
| return CallToolResult(content=[TextContent(type="text", text=f"Found 3 books matching {query!r}.")]) |
There was a problem hiding this comment.
P2: This sample handler never awaits any work, so it releases running before another request can enter and the included server never reaches its limit of four. Make the example hold or await realistic work, or explain that the cap only becomes observable when tool handlers overlap.
Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. At docs_src/middleware/tutorial002.py, line 37:
<comment>This sample handler never awaits any work, so it releases `running` before another request can enter and the included server never reaches its limit of four. Make the example hold or await realistic work, or explain that the cap only becomes observable when tool handlers overlap.</comment>
<file context>
@@ -0,0 +1,59 @@
+
+async def on_call_tool(ctx: ServerRequestContext, params: CallToolRequestParams) -> CallToolResult:
+ query = (params.arguments or {})["query"]
+ return CallToolResult(content=[TextContent(type="text", text=f"Found 3 books matching {query!r}.")])
+
+
</file context>
| answered with a JSON-RPC error. The connection stays up; the next message goes through. This is | ||
| how a server gates `subscriptions/listen` per caller: | ||
| answered with a JSON-RPC error. The connection stays up; the next message goes through. The | ||
| concurrency cap above. It is also how a server gates `subscriptions/listen` per caller: |
There was a problem hiding this comment.
P3: The Refuse bullet contains the fragment The concurrency cap above., so the new cross-reference reads as incomplete. Change it to This is the concurrency cap above..
Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. At docs/advanced/middleware.md, line 87:
<comment>The Refuse bullet contains the fragment `The concurrency cap above.`, so the new cross-reference reads as incomplete. Change it to `This is the concurrency cap above.`.</comment>
<file context>
@@ -56,14 +56,35 @@ That is the point. Middleware wraps **every** inbound message:
- answered with a JSON-RPC error. The connection stays up; the next message goes through. This is
- how a server gates `subscriptions/listen` per caller:
+ answered with a JSON-RPC error. The connection stays up; the next message goes through. The
+ concurrency cap above. It is also how a server gates `subscriptions/listen` per caller:
**[Deciding who may watch](../handlers/subscriptions.md#deciding-who-may-watch)** on the
Subscriptions page walks through it.
</file context>
|
|
||
|
|
||
| async def on_call_tool(ctx: ServerRequestContext, params: CallToolRequestParams) -> CallToolResult: | ||
| query = (params.arguments or {})["query"] |
There was a problem hiding this comment.
P3: on_call_tool raises an unhandled KeyError: 'query' when a client omits the argument. CallToolRequestParams.arguments is optional, and the SDK only validates the request shape, never the tool's declared input_schema, so a call like search_books with no arguments gets past validation and becomes a generic -32603 internal error that leaks the handler internals — the opposite of the clean, explained flows this tutorial teaches. Use .get("query", ""), or validate and raise a clean MCPError(INVALID_PARAMS, ...) since this page is about error handling.
Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. At docs_src/middleware/tutorial002.py, line 36:
<comment>`on_call_tool` raises an unhandled `KeyError: 'query'` when a client omits the argument. `CallToolRequestParams.arguments` is optional, and the SDK only validates the request shape, never the tool's declared `input_schema`, so a call like `search_books` with no `arguments` gets past validation and becomes a generic `-32603` internal error that leaks the handler internals — the opposite of the clean, explained flows this tutorial teaches. Use `.get("query", "")`, or validate and raise a clean `MCPError(INVALID_PARAMS, ...)` since this page is about error handling.</comment>
<file context>
@@ -0,0 +1,59 @@
+
+
+async def on_call_tool(ctx: ServerRequestContext, params: CallToolRequestParams) -> CallToolResult:
+ query = (params.arguments or {})["query"]
+ return CallToolResult(content=[TextContent(type="text", text=f"Found 3 books matching {query!r}.")])
+
</file context>
| query = (params.arguments or {})["query"] | |
| query = (params.arguments or {}).get("query", "") |
There was a problem hiding this comment.
Nothing blocking. The comments below are optional suggestions. There is no need to push a fix for them before merging.
Beyond the inline finding, I also checked that the hl_lines="15-16 40-55 59" ranges in docs/advanced/middleware.md line up with the actual lines of docs_src/middleware/tutorial002.py (they do), and that the running >= limit check and running += 1 have no await between them, so the cap cannot be overshot by interleaved requests.
Extended reasoning...
Docs-only change: a new middleware example (tutorial002.py), a new section on the Middleware page embedding it, and seven anyio tests; nothing under src/ and no security-sensitive surface. The one inline finding is a test-hygiene point about an unreachable handler registration, not a correctness bug in the example; the untouched translated pages are a known, routine lag for English-page edits.
| result = await client.call_tool("search_books", {"query": "dune"}) | ||
| assert result.is_error | ||
| assert result.content == [TextContent(type="text", text="Server busy. Try again shortly.")] | ||
|
|
There was a problem hiding this comment.
🟡 (optional) A reviewer following the repo's test rules finds a registered handler that can never run in test_answering_with_an_error_result_returns_it_to_the_caller_instead_of_raising. tests/docs_src/test_middleware.py:282 registers on_call_tool=tutorial002.on_call_tool, but the busy middleware at :281-284 short-circuits every tools/call, so the handler body is dead in this test. Fix: register a handler whose body is raise NotImplementedError (or omit on_call_tool) in every test where the middleware answers before call_next, so the test cannot silently start depending on the real handler.
Why this was flagged
In tests/docs_src/test_middleware.py:279-293 the busy middleware returns a CallToolResult(is_error=True) for tools/call at :282-284 and never calls call_next for that method, so the tutorial002.on_call_tool registered at :282 is never invoked. The repository's test-quality skill at .claude/skills/test-quality/SKILL.md:106 states registered-but-never-invoked handler bodies are raise NotImplementedError so they cannot silently become load-bearing. The dismissal argued the handler is the documented tutorial handler; the rule makes no such exception, and if a later edit made busy fall through, the assertions at :291-292 would fail on content rather than on an explicit NotImplementedError that names the cause. The sibling tests at :151, :189, :205 and :239 likewise register on_list_tools=tutorial002.on_list_tools without ever calling list_tools. Remedy: use a raise NotImplementedError stub for handlers a test never reaches.
Verification: nit. Triggering condition: any run of test_answering_with_an_error_result_returns_it_to_the_caller_instead_of_raising (tests/docs_src/test_middleware.py:279-293). The busy middleware at :281-284 returns a CallToolResult(is_error=True) for every ctx.method == "tools/call", so tutorial002.on_call_tool, registered at :288, is dead in this test. Nothing fails at runtime.
Fixes #3272.
The Middleware page lists "Refuse" among the things a middleware can do, but the only code for it was on the Subscriptions page. This adds a worked example to the Middleware page itself: a middleware that caps how many tool calls run at once and refuses the rest.
It covers the gap #3272 pointed out with a small deterministic example in the middleware docs, rather than the separate model-backed example server the issue proposed.
What it adds
docs_src/middleware/tutorial002.pymax_concurrent_tool_calls(limit), a factory that returns the middleware, on the low-levelServerlike the page's first example.tools/callis counted, so discovery and listing keep answering while tool calls are refused.MCPErrorwith the server's ownSERVER_BUSYcode. MCP has no standard code for "busy", and the spec leaves codes outside the JSON-RPC reserved range to applications.docs/advanced/middleware.mdtests/docs_src/test_middleware.pytools/listis answered with the cap reached.is_error=Trueresult instead of raising reaches the client as a tool result.Notes
src/and there is no new API.AI Disclaimer