feat(indexing): index-url — index a document set from an allow-listed http(s) URL in both transports (#208) - #210
Draft
adityamparikh wants to merge 3 commits into
Draft
adityamparikh wants to merge 3 commits into
adityamparikh wants to merge 3 commits into
Conversation
adityamparikh
added a commit
to adityamparikh/solr-mcp
that referenced
this pull request
Sep 16, 2026
…hem, and harden the policy (apache#208) Review findings from three independent passes over PR apache#210. Spring's response close() drains an unread body to keep the connection reusable, so a refused over-cap or non-2xx response was downloaded in full after the decision to refuse it. UrlFetcher now closes the body stream before returning on every non-body path, which makes the JDK drop the connection; a 2 MB body streamed at 10 KB/s must now fail inside a 5-second preemptive timeout. The policy refused only link-local addresses and the AWS IPv6 metadata literal while the docs claimed cloud-metadata addresses in general; Alibaba Cloud (100.100.100.200) and Azure WireServer (168.63.129.16) are now refused too, and every doc site says the known cloud-metadata addresses. A port above 65535 is rejected as an invalid URL instead of leaking HttpURLConnection's message; a non-numeric Content-Length counts as unknown; syntactic failures on a redirect hop report the redirect message rather than telling the caller to fix a URL they never supplied. Tests: every redirect status, a 3xx without Location, a quoted charset, SolrException in the Solr row, csv and markdown media types, a blank explicit format, an exact IPv6 allow-list entry, exact messages in the integration test, and the field-name summary that pins the indexPayload refactor. THREAT_MODEL §9 records the per-read timeout, header size and DNS-cache facts; FAQ and AGENTS.md tool counts and service lists updated. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: Aditya Parikh <aditya.m.parikh@gmail.com>
This was referenced Sep 16, 2026
… http(s) URL in both transports (apache#208) The inline indexing tools take the payload as a tool argument, so the model writes out every byte; index-url takes a URL instead. The server fetches it from a host on an operator allow-list (GitHub raw content by default), refuses link-local and cloud-metadata addresses on every redirect hop, sends no credentials or caller headers, and caps the body (10 MB), the per-read and total time, and the number of concurrent fetches. The fetched body is indexed exactly as the inline tool for its format would index it: CSV and XML go to Solr's own update handlers through index-csv-documents and index-xml-documents (so XML must be a Solr <add> block), JSON and Markdown are parsed by the server. A Solr 400 on CSV or XML is reported as rejected content with Solr's reason. Also routes the index-data prompt and README to index-url, narrows the threat model's "no tool argument picks a target URL" property to the Solr backend, and records the design in docs/superpowers/specs/2026-09-15-url-ingestion-design.md. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Signed-off-by: Aditya Parikh <aditya.m.parikh@gmail.com>
adityamparikh
force-pushed
the
feat/index-url
branch
from
September 24, 2026 16:48
973af04 to
0b9e41c
Compare
Records the agreed approach: index-url streams JSON/CSV/XML into Solr's update handlers instead of holding the body in memory, which removes SOLR_INDEX_URL_MAX_BYTES; Markdown keeps a fixed internal cap. Verified against SolrJ 10 source that ContentStreamUpdateRequest bodies are streamed, and that a failed content writer is only logged, so the commit must be sent separately and only after the server has seen the whole transfer. For files on the user's machine: Solr's own tooling now, MCP file upload (SEP-2631) later in both transports. Form-mode elicitation cannot carry a file; URL-mode elicitation is recorded as considered and not chosen (no STDIO equivalent, stateless HTTP mode, new authenticated web surface). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Signed-off-by: Aditya Parikh <aditya.m.parikh@gmail.com>
JSON, CSV and XML bodies now stream from the fetched URL straight into Solr's update handlers (JSON via /update/json/docs) through a fixed buffer, so a document of any size costs the server the same memory. SOLR_INDEX_URL_MAX_BYTES and the "too large" error are removed; Markdown, which Solr cannot parse, is read whole with no limit for now. SolrJ only logs a failed read of a streamed body and ends the upload early, which Solr can accept as a shorter valid document. TransferStream records the failure, the total deadline and whether the body reached its end and declared length, and the commit is a separate request sent only when it did; otherwise the tool reports a partial transfer and commits nothing. The XML <add>-only check reads just up to the root element through a marked buffer and rewinds. Verified on Solr 8.11, 9.9 and 10 that /update/json/docs flattens nested objects to dotted field names on one document. Implements phase A of docs/superpowers/specs/2026-09-24-streaming-and-local-file-ingestion-design.md. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Signed-off-by: Aditya Parikh <aditya.m.parikh@gmail.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Implements #208: an
index-urltool that indexes a JSON, CSV, Solr update XML or Markdown document set from an http(s) URL into a Solr collection without the payload passing through the model, registered in both the STDIO and HTTP transports.Design record with every decision and its reason:
docs/superpowers/specs/2026-09-15-url-ingestion-design.md(in this PR).Why
The four inline indexing tools take the payload as a tool-call argument, so the model emits every byte. Measured: 61 records took over two minutes, of which Solr took under a second. The server's differentiator is the onboarding loop (look at data, design schema, index, verify, search) and it broke at the indexing step, with no workaround in chat-only clients such as Claude Desktop. A URL argument is ~20 tokens regardless of payload size.
What
SOLR_INDEX_URL_ALLOWED_HOSTStakes exact hosts,*.suffixpatterns or*. Default:raw.githubusercontent.com,*.githubusercontent.com,github.com, so the tutorial URL works with no configuration and an HTTP deployment is not an SSRF primitive on day one.*—169.254.0.0/16,fe80::/10,fd00:ec2::254(AWS),100.100.100.200(Alibaba Cloud),168.63.129.16(Azure).user:pass@is rejected. OnlyAcceptandUser-Agentare sent.https→httprefused.format> requested URL extension (query string ignored) > final URL extension >Content-Type;text/plainandtext/htmlnever resolve.ContentStreamUpdateRequestthrough a fixed buffer, never held in memory: CSV to/update(asindex-csv-documents), XML to/updateafter the same<add>-only check asindex-xml-documents(reading only up to the root element, then rewinding), JSON to/update/json/docs, which turns each object into a document and flattens nested objects to dotted field names (studio.name; verified identical on Solr 8.11, 9.9 and 10). Markdown, which Solr cannot parse, is read whole and parsed by the server, with no limit for now. The reply for streamed formats is Solr's acceptance and the byte count (Solr reports no document count); a Solr 400 comes back asSolr rejected the URL content as <format>: <reason>.TransferStreamrecords the read failure, the total deadline, and whether the body reached its end and its declaredContent-Length; the commit is a separate request sent only if it did. Otherwise the tool says the transfer failed after N bytes and was not committed, and that documents Solr had already read may still appear (they become durable at Solr's nextautoCommit), so the call should be re-run.RestClientoverHttpURLConnection(already on the classpath); redirect following is disabled at the connection so the policy sees each hop.index-dataprompt and the server instructions route any http(s) URL toindex-url, files on the user's machine tobin/solr postorcurlagainst/update(run where the file is, which also covers Claude Desktop local files), and pasted data to the inline tools.Limits
Bound in
UrlIndexingProperties(solr.index-url.*), validated at startup rather than on the first tool call:allowed-hostsSOLR_INDEX_URL_ALLOWED_HOSTSconnect-timeoutSOLR_INDEX_URL_CONNECT_TIMEOUTread-timeoutSOLR_INDEX_URL_READ_TIMEOUTtotal-timeoutSOLR_INDEX_URL_TOTAL_TIMEOUTmax-concurrent-fetchesSOLR_INDEX_URL_MAX_CONCURRENT_FETCHESThreat model
THREAT_MODEL.md§8.5 said the client "cannot inject a target URL" at critical severity; this tool exercises that deliberately behind an allow-list. The PR narrows §8.5 to the Solr backend and its credentials, adds the allow-listed fetch as a §9 bounded property, adds the*-on-an-internal-network misuse to §11, theindex-urlnon-finding to §11a, a dated entry to §12, qualifies §13, adds the new knobs to §5a, corrects the tool counts in §1/§5a, and updatesdocs/security/stdio.mdandhttp.md.Three residual properties are recorded in §9 rather than fixed: the address check runs on the resolved addresses before the connection is made, so a DNS answer that changes in between (rebinding) can bypass it — the JDK's positive DNS cache narrows the window, and with the default allow-list it requires control of a GitHub host's DNS; response headers are not size-capped; and a Markdown body is read whole with no size limit, bounded only by the concurrency limit and the timeouts.
The default allow-list is the author's choice and needs maintainer confirmation.
Tests
New unit,
HttpServer-backed fetcher and Testcontainers integration tests cover the policy, format resolution, fetching and an end-to-end index of every format (including a plain XML file being refused, as withindex-xml-documents); both MCP client transports assertindex-urland its hints, a URL round trip and a refused metadata address../gradlew build: 530 tests, 0 skipped, 0 failures (the integration test streams an ~11 MB CSV, cuts a transfer off partway, and checks nested JSON, and was also run against Solr 8.11 and 10)./gradlew nativeTest -Pnative: 364 passed, 0 failed (163 skipped are the@DisabledInNativeImageMockito tests)Relationship to #205 and #202
This builds on both: CSV and XML take the same handlers and
<add>check as #205's inline tools, and Markdown uses the same creator asindex-markdown-documents. The design is recorded indocs/superpowers/specs/2026-09-15-url-ingestion-design.mdand, for streaming and files on the user's machine (why not elicitation; MCP file upload via SEP-2631 later),docs/superpowers/specs/2026-09-24-streaming-and-local-file-ingestion-design.md.Open points for maintainers
THREAT_MODEL.md§9.*allow-list should still refuse loopback and private (RFC 1918) addresses; today only link-local and the known cloud-metadata addresses are refused under*.🤖 Generated with Claude Code