Skip to content

feat(indexing): index-url — index a document set from an allow-listed http(s) URL in both transports (#208) - #210

Draft
adityamparikh wants to merge 3 commits into
apache:mainfrom
adityamparikh:feat/index-url
Draft

adityamparikh wants to merge 3 commits into
apache:mainfrom
adityamparikh:feat/index-url

Conversation

@adityamparikh

@adityamparikh adityamparikh commented Sep 16, 2026 •

Copy link
Copy Markdown
Contributor

Implements #208: an index-url tool that indexes a JSON, CSV, Solr update XML or Markdown document set from an http(s) URL into a Solr collection without the payload passing through the model, registered in both the STDIO and HTTP transports.

Design record with every decision and its reason: docs/superpowers/specs/2026-09-15-url-ingestion-design.md (in this PR).

Why

The four inline indexing tools take the payload as a tool-call argument, so the model emits every byte. Measured: 61 records took over two minutes, of which Solr took under a second. The server's differentiator is the onboarding loop (look at data, design schema, index, verify, search) and it broke at the indexing step, with no workaround in chat-only clients such as Claude Desktop. A URL argument is ~20 tokens regardless of payload size.

What

index-url(collection, url, format?)
  • Allow-listed hosts, checked on the requested URL and on every redirect hop. SOLR_INDEX_URL_ALLOWED_HOSTS takes exact hosts, *.suffix patterns or *. Default: raw.githubusercontent.com,*.githubusercontent.com,github.com, so the tutorial URL works with no configuration and an HTTP deployment is not an SSRF primitive on day one.
  • Link-local and the known cloud-metadata addresses refused on every hop even with * — 169.254.0.0/16, fe80::/10, fd00:ec2::254 (AWS), 100.100.100.200 (Alibaba Cloud), 168.63.129.16 (Azure).
  • No credentials or caller headers, ever; embedded user:pass@ is rejected. Only Accept and User-Agent are sent.
  • Redirects: five followed, the sixth is an error, https→http refused.
  • Non-2xx is an error before any body is read (a GitHub 404 body would otherwise index as a CSV row). Every refused or abandoned response is closed without reading its body, so the JDK drops the connection instead of draining the response for keep-alive.
  • Format: explicit format > requested URL extension (query string ignored) > final URL extension > Content-Type; text/plain and text/html never resolve.
  • No size limit: JSON, CSV and XML stream into Solr. The fetched body is copied straight into a ContentStreamUpdateRequest through a fixed buffer, never held in memory: CSV to /update (as index-csv-documents), XML to /update after the same <add>-only check as index-xml-documents (reading only up to the root element, then rewinding), JSON to /update/json/docs, which turns each object into a document and flattens nested objects to dotted field names (studio.name; verified identical on Solr 8.11, 9.9 and 10). Markdown, which Solr cannot parse, is read whole and parsed by the server, with no limit for now. The reply for streamed formats is Solr's acceptance and the byte count (Solr reports no document count); a Solr 400 comes back as Solr rejected the URL content as <format>: <reason>.
  • Commit only after the whole body arrived. SolrJ only logs a failed read of a streamed body and ends the upload early, which Solr can accept as a shorter, valid document. TransferStream records the read failure, the total deadline, and whether the body reached its end and its declared Content-Length; the commit is a separate request sent only if it did. Otherwise the tool says the transfer failed after N bytes and was not committed, and that documents Solr had already read may still appear (they become durable at Solr's next autoCommit), so the call should be re-run.
  • HTTP client: Spring RestClient over HttpURLConnection (already on the classpath); redirect following is disabled at the connection so the policy sees each hop.
  • Guidance: the index-data prompt and the server instructions route any http(s) URL to index-url, files on the user's machine to bin/solr post or curl against /update (run where the file is, which also covers Claude Desktop local files), and pasted data to the inline tools.

Limits

Bound in UrlIndexingProperties (solr.index-url.*), validated at startup rather than on the first tool call:

Property Env Default
allowed-hosts SOLR_INDEX_URL_ALLOWED_HOSTS GitHub raw-content hosts
connect-timeout SOLR_INDEX_URL_CONNECT_TIMEOUT 10s
read-timeout SOLR_INDEX_URL_READ_TIMEOUT 30s, per socket read
total-timeout SOLR_INDEX_URL_TOTAL_TIMEOUT 5m for a whole fetch, redirects included, so a host that drips bytes cannot outlast the per-read timeout
max-concurrent-fetches SOLR_INDEX_URL_MAX_CONCURRENT_FETCHES 4; further calls fail immediately, bounding total memory and outbound connections

Threat model

THREAT_MODEL.md §8.5 said the client "cannot inject a target URL" at critical severity; this tool exercises that deliberately behind an allow-list. The PR narrows §8.5 to the Solr backend and its credentials, adds the allow-listed fetch as a §9 bounded property, adds the *-on-an-internal-network misuse to §11, the index-url non-finding to §11a, a dated entry to §12, qualifies §13, adds the new knobs to §5a, corrects the tool counts in §1/§5a, and updates docs/security/stdio.md and http.md.

Three residual properties are recorded in §9 rather than fixed: the address check runs on the resolved addresses before the connection is made, so a DNS answer that changes in between (rebinding) can bypass it — the JDK's positive DNS cache narrows the window, and with the default allow-list it requires control of a GitHub host's DNS; response headers are not size-capped; and a Markdown body is read whole with no size limit, bounded only by the concurrency limit and the timeouts.

The default allow-list is the author's choice and needs maintainer confirmation.

Tests

New unit, HttpServer-backed fetcher and Testcontainers integration tests cover the policy, format resolution, fetching and an end-to-end index of every format (including a plain XML file being refused, as with index-xml-documents); both MCP client transports assert index-url and its hints, a URL round trip and a refused metadata address.

  • ./gradlew build: 530 tests, 0 skipped, 0 failures (the integration test streams an ~11 MB CSV, cuts a transfer off partway, and checks nested JSON, and was also run against Solr 8.11 and 10)
  • ./gradlew nativeTest -Pnative: 364 passed, 0 failed (163 skipped are the @DisabledInNativeImage Mockito tests)

Relationship to #205 and #202

This builds on both: CSV and XML take the same handlers and <add> check as #205's inline tools, and Markdown uses the same creator as index-markdown-documents. The design is recorded in docs/superpowers/specs/2026-09-15-url-ingestion-design.md and, for streaming and files on the user's machine (why not elicitation; MCP file upload via SEP-2631 later), docs/superpowers/specs/2026-09-24-streaming-and-local-file-ingestion-design.md.

Open points for maintainers

  • The default allow-list contents.
  • Whether Markdown should regain a size limit, and whether the concurrency limit (4) and total timeout (5 m) fit now that the total timeout is the practical bound on file size.
  • Whether to connect to the validated IP rather than re-resolving the host, which would close the DNS-rebinding window recorded in THREAT_MODEL.md §9.
  • Whether an operator-configured * allow-list should still refuse loopback and private (RFC 1918) addresses; today only link-local and the known cloud-metadata addresses are refused under *.

🤖 Generated with Claude Code

adityamparikh added a commit to adityamparikh/solr-mcp that referenced this pull request Sep 16, 2026
…hem, and harden the policy (apache#208)

Review findings from three independent passes over PR apache#210.

Spring's response close() drains an unread body to keep the connection
reusable, so a refused over-cap or non-2xx response was downloaded in
full after the decision to refuse it. UrlFetcher now closes the body
stream before returning on every non-body path, which makes the JDK drop
the connection; a 2 MB body streamed at 10 KB/s must now fail inside a
5-second preemptive timeout.

The policy refused only link-local addresses and the AWS IPv6 metadata
literal while the docs claimed cloud-metadata addresses in general;
Alibaba Cloud (100.100.100.200) and Azure WireServer (168.63.129.16) are
now refused too, and every doc site says the known cloud-metadata
addresses. A port above 65535 is rejected as an invalid URL instead of
leaking HttpURLConnection's message; a non-numeric Content-Length counts
as unknown; syntactic failures on a redirect hop report the redirect
message rather than telling the caller to fix a URL they never supplied.

Tests: every redirect status, a 3xx without Location, a quoted charset,
SolrException in the Solr row, csv and markdown media types, a blank
explicit format, an exact IPv6 allow-list entry, exact messages in the
integration test, and the field-name summary that pins the indexPayload
refactor. THREAT_MODEL §9 records the per-read timeout, header size and
DNS-cache facts; FAQ and AGENTS.md tool counts and service lists updated.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Aditya Parikh <aditya.m.parikh@gmail.com>
… http(s) URL in both transports (apache#208)

The inline indexing tools take the payload as a tool argument, so the model
writes out every byte; index-url takes a URL instead. The server fetches it
from a host on an operator allow-list (GitHub raw content by default),
refuses link-local and cloud-metadata addresses on every redirect hop,
sends no credentials or caller headers, and caps the body (10 MB), the
per-read and total time, and the number of concurrent fetches.

The fetched body is indexed exactly as the inline tool for its format would
index it: CSV and XML go to Solr's own update handlers through
index-csv-documents and index-xml-documents (so XML must be a Solr <add>
block), JSON and Markdown are parsed by the server. A Solr 400 on CSV or XML
is reported as rejected content with Solr's reason.

Also routes the index-data prompt and README to index-url, narrows the
threat model's "no tool argument picks a target URL" property to the Solr
backend, and records the design in
docs/superpowers/specs/2026-09-15-url-ingestion-design.md.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: Aditya Parikh <aditya.m.parikh@gmail.com>
adityamparikh and others added 2 commits September 24, 2026 22:22
Records the agreed approach: index-url streams JSON/CSV/XML into Solr's
update handlers instead of holding the body in memory, which removes
SOLR_INDEX_URL_MAX_BYTES; Markdown keeps a fixed internal cap. Verified
against SolrJ 10 source that ContentStreamUpdateRequest bodies are
streamed, and that a failed content writer is only logged, so the commit
must be sent separately and only after the server has seen the whole
transfer.

For files on the user's machine: Solr's own tooling now, MCP file upload
(SEP-2631) later in both transports. Form-mode elicitation cannot carry a
file; URL-mode elicitation is recorded as considered and not chosen (no
STDIO equivalent, stateless HTTP mode, new authenticated web surface).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: Aditya Parikh <aditya.m.parikh@gmail.com>
JSON, CSV and XML bodies now stream from the fetched URL straight into
Solr's update handlers (JSON via /update/json/docs) through a fixed
buffer, so a document of any size costs the server the same memory.
SOLR_INDEX_URL_MAX_BYTES and the "too large" error are removed; Markdown,
which Solr cannot parse, is read whole with no limit for now.

SolrJ only logs a failed read of a streamed body and ends the upload
early, which Solr can accept as a shorter valid document. TransferStream
records the failure, the total deadline and whether the body reached its
end and declared length, and the commit is a separate request sent only
when it did; otherwise the tool reports a partial transfer and commits
nothing. The XML <add>-only check reads just up to the root element
through a marked buffer and rewinds.

Verified on Solr 8.11, 9.9 and 10 that /update/json/docs flattens nested
objects to dotted field names on one document. Implements phase A of
docs/superpowers/specs/2026-09-24-streaming-and-local-file-ingestion-design.md.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: Aditya Parikh <aditya.m.parikh@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant