Skip to content

feat(vectorisers): Add FastEmbedVectoriser implementation of VectoriserBase and add fast embed dependency group - #224

Open
Tom-Owen-ONS wants to merge 9 commits into
mainfrom
222-make-huggingface_light-vectoriser-dependency-group
Open

Tom-Owen-ONS wants to merge 9 commits into
mainfrom
222-make-huggingface_light-vectoriser-dependency-group

Conversation

@Tom-Owen-ONS

Copy link
Copy Markdown
Contributor

✨ Summary

Add a new FastEmbedVectoriser to ClassifAI as a lightweight local embedding backend that avoids torch and transformers at runtime, alongside the new fastembed optional dependency group, tests, and documentation updates.

I have run a manual check in general_workflow_demo.ipynb with FastEmbedVectoriser and it runs end to end.

I haven't done a performance benchmark, however happy to do so with guidance on how this is done for classifai.

I have not included the model caching or normalisation logic in survey-assist-embed-core to remain consistent with other implementations of VectoriserBase.

📜 Changes Introduced

  • Added FastEmbedVectoriser as a new VectoriserBase implementation using fastembed.TextEmbedding with ClassifAI-style error handling.
  • Added the fastembed optional dependency group and included it in the all extra.
  • Updated vectoriser exports and overview documentation so FastEmbed appears in the public vectoriser surface and generated API docs.
  • Updated installation and usage documentation to explain the lighter classifai[fastembed] install path.
  • Added unit tests covering single-string input, batch input, load failures, transform failures, and empty input handling.
  • Feature implementation (feat:) / bug fix (fix:) / refactoring (chore:) / documentation (docs:) / testing (test:)
  • Updates to tests and/or documentation
  • Terraform changes (if applicable)

✅ Checklist

  • Code passes linting with Ruff
  • Security checks pass using Bandit
  • API and Unit tests are written and pass using pytest
  • Terraform files (if applicable) follow best practices and have been validated (terraform fmt & terraform validate)
  • DocStrings follow Google-style and are added as per Pylint recommendations
  • Documentation has been updated if needed

🔍 How to Test

  1. Sync the repo environment, including optional extras:
    uv sync --all-extras
  2. Run the test suite:
    uv run pytest -q
  3. Smoke test the new backend directly:
    uv run --extra fastembed python -c "from classifai.vectorisers import FastEmbedVectoriser; vectoriser = FastEmbedVectoriser(model_name='sentence-transformers/all-MiniLM-L6-v2'); print(vectoriser.transform('hello world').shape)"
  4. Open DEMO/general_workflow_demo.ipynb and instantiate FastEmbedVectoriser there. This has been manually validated on this branch.
  5. Review documentation updates.

@Tom-Owen-ONS
Tom-Owen-ONS requested a review from jamie-ons August 21, 2026 15:15
@Tom-Owen-ONS
Tom-Owen-ONS requested a review from a team as a code owner August 21, 2026 15:15
@Tom-Owen-ONS Tom-Owen-ONS linked an issue Aug 21, 2026 that may be closed by this pull request
@github-actions github-actions Bot added the enhancement New feature or request label Aug 21, 2026
@jamie-ons

Copy link
Copy Markdown
Contributor

Have made the relevant changes to merge it in. I am now happy with it being merged in, needs to have one more person review it though.

@jamie-ons jamie-ons self-assigned this Sep 9, 2026
jamie-ons
jamie-ons previously approved these changes Sep 9, 2026

@jamie-ons jamie-ons left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This looks great. I have updated it with some changes to make it work well with the rest of the package. I have also added a explicit model path parameter as FastEmbed requires both:

  • a model name
  • a model path

To load models locally and I thought having it in kwargs was too confuesing

frayle-ons
frayle-ons previously approved these changes Sep 14, 2026

@frayle-ons frayle-ons left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I really like this - its fast, and I've tested loading different models and then also using the FastEmbedVectoriser in the general_worflow.ipynb notebook as the tests suggest and it all works well.

Happy to approve as is but I added a few comments for potential changes.

One additional thought is - in the similar HuggingFaceVectoriser class we have a revision argument, but we don't seem to have that in this new FastEmbedVectoriser class. Should we/could we add that?

def __init__(
self,
model_name: str,
specific_model_path: str | None = None,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Consider removing this parameter and then user can pass it as part of the kwarg to the constructor?

Looking at how this is handled in the HuggingFaceVectoriser class (HuggingFace also has its own way of local cache checking for models under the hood), we don't have a special parameter in that constructor.

So mirroring that might be good for consistency

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@jamie-ons added this so not sure of the rationale, don't mind either way.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Will confirm when @jamie-ons returns from leave, but I think this was to allow someone to specify whether to use ONNX format weights on a model card if there's multiple options available

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The HuggingFaceVectoriser allows a user to put in a path to a local model as the model_name, however in fast embed this is not allowed.

I therefore decided to add in this specific model path as a key use of the FastEmbedVectoriser is to run models in enviroments where resources may be lower.

This could often mean that ability to download large packages (torch) or files (model weights) may be restricted. I also don't think the FastEmbed documentation is that clear about how to run models from a local download and so thought adding the argument would save them time in researching how to do it.

I will update the docstrings to make this clearer.

Comment thread CHANGELOG.md

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do we update the changelog with every PR on the package? I thought we just did one changelog update per release, and looked back at the merged PRs when writing it?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Soz haven't had to think about CHANGELOGs in a while as that's done for me on scanner ;)

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Will have a look later today

@jamie-ons
jamie-ons dismissed stale reviews from frayle-ons and themself via 3ce73b2 September 24, 2026 11:18

@frayle-ons frayle-ons left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good to me! But unit tests?

@jamie-ons

Copy link
Copy Markdown
Contributor

Looks good to me! But unit tests?

Sure il add some

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancement New feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Make huggingface_light Vectoriser & dependency group

4 participants