Skip to content

feat(indexing): index-url — stream documents into Solr from an http(s) URL in both transports #208

Description

@adityamparikh

Summary

Add an index-url MCP tool that indexes a JSON, CSV, XML or Markdown document set from an http(s) URL into a Solr collection, so the payload never passes through the model's context. Registered in both the STDIO and HTTP transports. Hosts are allow-listed (GitHub raw content by default), the body is capped (10 MB by default) and parsed by the existing document creators, and the model's guidance routes anything larger to Solr's own bulk tooling. This issue records the design and the reasons; the full implementation spec (tool contract, fetch algorithm, exact error messages, test list, threat-model edits) is at https://github.com/adityamparikh/solr-mcp/blob/feat/index-url/docs/superpowers/specs/2026-09-15-url-ingestion-design.md and ships in the PR.

Revised 2026-09-15. The first version of this issue proposed streaming the body through server-side parsers, keeping a STDIO-only index-file, and allowing any host by default. That was built and is green on wip/index-url-streaming, but it duplicated Solr's own parsers, collided with #205, and reversed the #194 position. See "What changed and why" below.

Problem

All four indexing tools take the payload as a tool-call argument, so the model has to emit every byte. Measured when #197 was closed: indexing the 61-record shows.json sample took over two minutes; Solr and the server accounted for under one second. The rest was the model emitting ~9,500 tokens of escaped JSON. Anything larger than one response's output budget has to be split across calls, each of which commits.

This matters because the server's differentiator is the onboarding loop: look at a dataset, design the schema, add fields, index, verify, search, iterate, all in one conversation. The loop breaks at the indexing step, and in a chat-only client such as Claude Desktop there is no workaround because the client cannot run a shell command. A URL argument is ~20 tokens regardless of payload size.

Why a URL and not a file upload

MCP has no client-to-server file transfer. SEP-2356 was closed in favour of SEP-2631 ("File Objects and Transfer"), still an open draft. A file attached in Claude Desktop is extracted into the model's context; the model gets content and a filename, never a path or handle it can pass to a server. A URL is the only non-inline source both transports can reach identically today.

Why a size cap and not streaming

Once the bytes bypass the model, the bottleneck is gone: a 10 MB body parsed in memory by the existing creators is instant, and 10 MB is roughly 15,000 shows-sized records, far beyond what a conversation onboards. Datasets larger than that belong to bin/solr post or Solr's /update handler, which do bulk ingestion better than this server can, with no model in the loop and no new attack surface, and which run where the file is. So the index-data prompt tells the model: URL under the cap → index-url; larger → give the user the exact bin/solr post or curl command; small pasted or attached data → the inline tools as today.

Proposed design

One new @Service, UrlIndexingService, with no @Profile gate:

index-url(collection: String, url: String, format: String?) -> summary string

Fetch behaviour, in order:

  1. Allow-list, checked on the requested URL and on every redirect hop. Property solr.index-url.allowed-hosts / env SOLR_INDEX_URL_ALLOWED_HOSTS, a comma-separated list of exact hosts, *.suffix patterns, or * for any host. Default raw.githubusercontent.com,*.githubusercontent.com,github.com. An empty list allows nothing.
  2. Link-local and cloud-metadata addresses (169.254.0.0/16, fe80::/10, fd00:ec2::254) are refused on every hop even with *. Loopback and RFC1918 are governed by the allow-list like any other host.
  3. No credentials and no caller-supplied headers, ever. Only Accept and User-Agent: solr-mcp are sent. A URL with embedded user:pass@ is rejected.
  4. Redirects: up to five are followed, the sixth redirect response is an error, and an https → http redirect is refused.
  5. Non-2xx is an error before any body is read. A GitHub 404 body is the text 404: Not Found; the CSV parser would index it as a document.
  6. Size cap solr.index-url.max-bytes (default 10MB): rejected from Content-Length before reading when declared, and while reading otherwise. Over the cap, the error tells the model to use Solr's bulk tooling.
  7. Format: explicit format > requested URL's path extension (query string ignored) > final URL's extension > Content-Type > error. text/plain and text/html never resolve; the error suggests format=.
  8. Body read fully into memory, decoded with the declared charset (UTF-8 default), and handed to the existing createSchemalessDocumentsFrom* parsers and indexDocuments. Because parsing completes before indexing starts, every fetch or parse failure truthfully reports "nothing was indexed".

HTTP client: Spring's RestClient over SimpleClientHttpRequestFactory, already on the classpath, with connect-timeout 10s and a per-read read-timeout 30s that covers the body (the JDK HttpClient's timeout covers only the headers, which is what forced a hand-written watchdog in the first version). Redirect following is disabled at the connection so the policy runs per hop.

Decisions and why

  1. Allow-list on by default. The tutorial URL works with no configuration; an HTTP deployment on a corporate network is safe without the operator remembering a knob, matching the project's HTTP-security-on-by-default posture; operators hosting data elsewhere add one host. Default contents are the author's choice; confirm in the PR.
  2. Metadata addresses always refused. Removes the single worst outcome of * with no configuration and no legitimate loss.
  3. In-memory parse with a cap, existing parsers. Reuses tested code, keeps field-name sanitising and nested-XML flattening, adds no parser code, and does not conflict with refactor(indexing): forward CSV and XML to Solr's own update handlers #205. The 10 MB default is the author's choice; confirm in the PR.
  4. Larger datasets go to Solr directly, by guidance. Better tool for the job, and it also covers local files under both transports.
  5. No index-file. feat(indexing): stream local JSON, CSV, XML and Markdown files #194 closed it for giving the transports different surfaces; bin/solr post covers the local-file case it was for. This design reverses nothing recorded on feat(indexing): stream local JSON, CSV, XML and Markdown files #194.
  6. One tool with optional format. feat(indexing)!: fold per-format inline tools into index-documents #197 kept per-format inline tools for token shape; index-url carries no payload, so there is no shape to optimise.
  7. RestClient, not a hand-rolled client. Real read timeout, Micrometer observations, no new dependency.

Locally selected files in HTTP mode

A file selected in Claude Desktop while connected to an HTTP-mode server keeps using the existing inline path: the client places the content in the model's context and the model calls the per-format inline tool. That is unchanged. For a file too large to paste, the model gives the user the bin/solr post command to run where the file is. index-url is not a substitute: it resolves the URL from the server's network position, so http://localhost:8000/file.json served on the caller's machine hits the server's own loopback, and with the default allow-list it is refused anyway.

Threat model impact

THREAT_MODEL.md §8 property 5 says the AI client "cannot repoint the server or inject a target URL", with violation at critical severity, and §12 lists "backend-target from a tool argument" as a model-changing condition. This tool exercises that condition deliberately, behind an allow-list. The PR narrows §8.5 to the Solr backend and its credentials, adds the allow-listed fetch as a §9 bounded property (including the DNS-rebinding window between the address check and the connect), adds the *-on-an-internal-network misuse to §11, adds the index-url non-finding to §11a (KNOWN-NON-FINDING with the default list, OUT-OF-MODEL: trusted-input with *), records the decision in §12 with the date, qualifies the §13 VALID row, adds the three security-relevant knobs to §5a, updates the tool counts in §1 and §5a, and updates docs/security/stdio.md and http.md, which §8.5 cites.

Tests

  • Pure unit tests, all running natively: UrlTargetPolicyTest (allow-list matching rules, metadata refusal even with *, scheme and credential rejection), IndexFormatsTest, UrlFetcherTest (a JDK HttpServer on loopback: media type and charset, 404 before any body, absolute and relative redirects, five followed and the sixth refused, redirect loop, https→http downgrade, policy re-run on a redirect to a metadata or non-listed host, unresolvable host, cap by Content-Length and by reading, read timeout on a stalling handler, and that only Accept and User-Agent are sent).
  • UrlIndexingServiceTest (Mockito, the only class skipped natively): the error-mapping table row by row and format precedence.
  • IndexingServiceTest passes unmodified after the four inline tools are refactored to share one indexPayload method with index-url.
  • UrlIndexingIntegrationTest (Testcontainers Solr): all four formats, query string, media type, redirect; text/plain, HTML, 404 and an over-cap body each fail and leave the collection empty.
  • Both MCP client transports assert index-url and its hints through the shared base, plus one round trip each, plus refused-host and non-listed-host calls that come back as MCP tool errors.
  • Zero skipped tests in ./gradlew build other than the OTLP class already @Disabled on main; under nativeTest -Pnative the skipped count rises by exactly the Mockito class.

Delivery

One PR from feat/index-url against main, opened as a draft first so the threat-model changes can be reviewed early. No conflict with #205 (no parser code added; only the four inline tool methods are touched, to extract indexPayload). Textual overlap only with #196, #202, #203 and #207.

What changed and why (2026-09-15)

The first version streamed the body through server-side parsers with a hand-written idle watchdog and byte-cap stream, kept a STDIO-only index-file, and allowed any host by default. It is built, green under build and nativeTest -Pnative, and preserved at wip/index-url-streaming (6c9b6f4). It was dropped because (1) it duplicated parsing Solr's /update handlers already do, and the spine lived in the CSV and XML creators that #205 deletes; (2) index-file reversed the #194 position for a case bin/solr post serves better; (3) an open-by-default fetch made every HTTP deployment an SSRF primitive, and its mitigations were defending a design a default allow-list makes unnecessary.

Follow-up (not in this PR)

After #205 merges, index-url can keep its exact contract and, for JSON, CSV and XML, stream the fetched body straight into Solr's /update handlers through a ContentStreamUpdateRequest, removing the size cap for those formats. JSON would lose the server's field-name sanitising, which #205 already accepted for CSV and XML. To be filed as its own issue when #205 merges.

Related

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions