Skip to content

feat: site owner configures the model source; no download address in the plugin (#228) - #239

Merged
pluginslab merged 2 commits into
mainfrom
feat/228-owner-configured-model-source
Sep 29, 2026
Merged

pluginslab merged 2 commits into
mainfrom
feat/228-owner-configured-model-source

Conversation

@pluginslab

Copy link
Copy Markdown
Owner

@ivdimova I'd like your opinion on the approach before we submit, not just the code.

Why

WP.org has flagged "calling files remotely" in four rounds, most recently in R agentic-admin/28May26/T5 27Sep26/4.3 (27 Sep). Every time they quote the same WebLLM lines: its built-in GitHub address for model libraries, plus a list of about 140 models with Hugging Face URLs. We've argued it's a service (in the readme and in replies), and they re-sent the same template each time.

The model files can't ship in the plugin: the default model is 1.2 GB, and the compiled model libraries (.wasm) are hosted only on GitHub.

The approach

Treat the model download like the remote LLM endpoint, which the site owner types in and which WP.org has never flagged in five rounds:

  • The plugin contains no model download address. A build step (tools/strip-webllm-prebuilt-loader.js) empties WebLLM's built-in list and GitHub prefix. The build fails if WebLLM changes shape. index.js and sw.js went from 278 huggingface.co / 2 raw.githubusercontent.com addresses to 0 / 0.
  • New option agentic_admin_model_source (weights URL + library URL), with no default. It's saved through core's /wp/v2/settings, so only admins can change it (tested: logged-out → 401, javascript: rejected). It's removed on uninstall.
  • Settings → "Model source" card. Until it's filled in, Load Model is disabled with a notice. Remote and Connector are unaffected.
  • The readme lists the two MLC-AI addresses to paste under External services, and notes you can host the files yourself.

The trade-off I'd like your view on

Pro: it answers the question the reviewer asked back in June: are the assets "configurable by the site owner, user-supplied, or tied to third-party providers?" The answer becomes "configured by the owner". It also removes the exact strings their tool matches.

Con: a new user has to paste two addresses before the local engine works, which is worse first-run UX. Marcel asked about pre-filling them. I advised against it, because a default would put the addresses back in the plugin (PHP instead of JS) and be the same "plugin picked the host" situation. The plan is to ask the reviewer directly in our reply whether a pre-filled, owner-editable default would be acceptable, and add it later only with their OK.

Is this the right call, or would you handle it differently (e.g. pre-fill now, or keep arguing the service exception)?

Testing

  • Unit tests: 81 passed, 3 skipped. There are 6 new ones covering the config builder, and the build step failing if WebLLM changes.
  • Lint: clean. PHP shows 3 warnings that were already there.
  • Build: 0 model addresses in either bundle.
  • WordPress Playground (WP 7.1.2, PHP 8.3, WP_DEBUG on): the option saves and sanitizes through /wp/v2/settings, and the empty state shows correctly.
  • Not yet verified: a full model download in a real browser through the new settings. Marcel is running that in Playground now.

The #238 fixes (Tested up to, escaping) are merged but not yet uploaded. The plan is one upload with this PR included, if we go ahead.

🤖 Generated with Claude Code

…the plugin (#228)

WP.org has flagged WebLLM's built-in model download addresses in four
review rounds (latest R agentic-admin/28May26/T5 27Sep26/4.3). The model
files (1.2 GB+) and their compiled libraries can't ship in the plugin, so
the download source becomes a site-owner setting, the same way the remote
LLM endpoint already is.

- tools/strip-webllm-prebuilt-loader.js: webpack loader that empties
  WebLLM's prebuiltAppConfig.model_list and modelLibURLPrefix. The build
  fails if the expected code is missing, so an upgrade can't bring the
  addresses back. Bundles now contain no huggingface.co or
  raw.githubusercontent.com address.
- Option agentic_admin_model_source (weights_url, library_url), registered
  with show_in_rest so it saves through /wp/v2/settings (manage_options).
  No default. Sanitized with esc_url_raw (http/https) + trailing slash.
  Removed on uninstall.
- model-source.js builds WebLLM's appConfig from the owner's addresses;
  model-loader passes it to CreateMLCEngine, CreateServiceWorkerMLCEngine
  and hasModelInCache. Loading is refused until a source is set.
- Settings gets a "Model source" card; Load Model is disabled with a notice
  until a source is set. Remote and Connector engines are unaffected.
- readme lists the MLC-AI addresses to paste, under External services.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@pluginslab
pluginslab requested a review from ivdimova September 27, 2026 21:29
@ivdimova

Copy link
Copy Markdown
Collaborator

Agree with the direction. After four rounds of the same template, changing the design is better than a fifth reply arguing the service exception. It also answers their June question ("configurable by the site owner?") with a clear yes. And no pre-fill for now: a default in PHP is the same "plugin picked the host" situation, so asking them first is right.

I checked the code against WebLLM 0.2.82. The four model records match WebLLM's own records exactly, appConfig does reach the service worker engine (through setAppConfig), and the model-source tests pass locally. A few things I'd fix or decide before we upload:

1. The strongest objection they could still raise. The .wasm model libraries are executable code, and the readme still tells people to paste the GitHub address. A reviewer could say we moved the URL from the bundle into the readme. In the email reply I'd say it plainly: the plugin makes no request until the owner enters a source, the owner can host the files anywhere, and MLC-AI is one example. Then ask them directly if a pre-filled, editable default is acceptable. Better to say this ourselves than have them find it.

2. Self-hosting needs a specific folder layout, and the readme doesn't say so. WebLLM's cleanModelUrl() adds resolve/main/ to any model URL without /resolve/…/ in it. So a self-hosted weights URL only works if each model sits at <weights_url>/<model_id>/resolve/main/. Since the readme tells owners they can host the files themselves, it should give that layout, or we should build the /resolve/main/ part into the URL ourselves.

3. The library version is now stored per site. The readme address ends in v0_2_80, which matches the modelVersion in the WebLLM we bundle today. When a later WebLLM upgrade bumps that version, existing sites keep the old library URL in their settings and can end up with a runtime and libraries that don't match. Options: store only the base URL and add modelVersion in buildAppConfig(), or at least add a note in the readme/UI and a check in the release process.

4. Existing users lose the local engine when they update. With no source set, isModelCached() returns false, so auto-load stops and Load Model is disabled, even for people whose model is already in the browser cache. Pasting the same MLC-AI addresses brings the cache back (same URLs), but people need to be told. I'd add an Upgrade Notice in the readme and link the warning notice straight to Settings → Model source.

Smaller things:

  • sanitize_model_source() accepts http://. On an https admin page the browser blocks those downloads as mixed content, so I'd accept https only.
  • The Load Model button reads isModelSourceConfigured() from module state at render time. After you save in Settings, it only updates when ModelStatus re-renders. It's worth checking the tab switch in the Playground run.
  • The loader's bracket scan skips // comments but not /* */ block comments. It works on 0.2.82 (the test proves it), and the final check catches any leftover address, so this is fine as long as that check stays.

I haven't done the full in-browser model download; that's Marcel's Playground run. Once that passes and 2–4 are handled or answered, 👍 from me.

@ivdimova

Copy link
Copy Markdown
Collaborator

One more I missed: readme.txt screenshot caption 2 still says "The Qwen 3 1.7B weights (~1.2 GB) are fetched from the MLC-AI / HuggingFace CDN". With this PR nothing downloads until the owner sets a source, and naming a fixed host is exactly what we are trying to remove. Suggest something like "…are fetched from the configured model source". Caption 4 is already updated.

- Library URL is a base address; buildAppConfig() appends WebLLM's
  modelVersion, so a WebLLM upgrade picks matching libraries without
  owners editing their settings.
- Readme and Settings help document the self-hosting layout
  (<weights>/<model>/resolve/main/, <library>/<version>/).
- Model source accepts https only (http is blocked as mixed content).
- Load Model and the Settings warning update live after saving, via a
  model-source subscription.
- The "no model source" notice links to #model-source, which opens
  Settings and scrolls to the card.
- Upgrade Notice for existing local-engine users; screenshot caption 2
  no longer names a fixed host.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@pluginslab

Copy link
Copy Markdown
Owner Author

Thanks @ivdimova, all handled in 6efa590:

  1. Reply wording: rewritten. It now says the plugin makes no request until the owner enters a source, the owner can host the files anywhere, and MLC-AI is one option. It also asks directly about a pre-filled default.
  2. Self-hosting layout: documented in the readme and in the Settings help text: <weights URL>/<model ID>/resolve/main/ and <library URL>/<WebLLM version>/.
  3. Version per site: the library URL is now a base address, and buildAppConfig() adds webllm.modelVersion, so a WebLLM upgrade picks matching libraries without owners editing anything. The readme address no longer ends in v0_2_80.
  4. Existing users: added an Upgrade Notice to the readme. The "no model source" notice links to #model-source, which opens Settings and scrolls to the card. Their cache is reused, because the MLC-AI addresses produce the same URLs as before.

Smaller ones:

  • https only.
  • Load Model and the Settings warning now update live through a model-source subscription.
  • Caption 2 now says "configured model source".
  • Kept the loader's final address check.

82 unit tests pass (added one for the subscription), lint is clean, and both bundles contain 0 model addresses. Marcel's in-browser download test is next.

@pluginslab

Copy link
Copy Markdown
Owner Author

In-browser test passed (Marcel, WordPress Playground with this build): with no source set, Load Model is disabled and the notice links to Settings → Model source. After pasting the MLC-AI addresses and saving, Load Model enables without a reload, and the model downloads and runs. That covers the last open item from your review, so merging.

@pluginslab
pluginslab merged commit 83bb984 into main Sep 29, 2026
4 checks passed
@pluginslab
pluginslab deleted the feat/228-owner-configured-model-source branch September 29, 2026 16:10
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants