diff --git a/DEMO/README.md b/DEMO/README.md index 5f6474a..8eef653 100644 --- a/DEMO/README.md +++ b/DEMO/README.md @@ -86,6 +86,20 @@ It covers: **Note:** The Evaluation module is currently in development and its API is subject to change in future releases. +### 6. Working with datasets in multiple languages with ClassifAI and Multilingual Vectoriser Models : `multilingual_datasets_and_vectorisers.ipynb` + +This notebook demonstrates how to work with datasets containing text in multiple languages using ClassifAI and multilingual Vectoriser models. + +It covers: + +* How multilingual Vectorisers can encode text from different languages into a shared embedding space, allowing semantically similar text to be represented similarly, regardless of language. + +* Building a `VectorStore` from a multilingual dataset and searching it using queries in different languages. + +* Visualising the multilingual embedding space to show how semantically similar text in different languages is positioned close together. + +* Performing searches in one language and retrieving relevant results in multiple languages, showcasing the multilingual capabilities of the VectorStore. + --- ## Installation of classifai diff --git a/DEMO/custom_vectoriser.ipynb b/DEMO/custom_vectoriser.ipynb index 3b40eab..535c9ed 100644 --- a/DEMO/custom_vectoriser.ipynb +++ b/DEMO/custom_vectoriser.ipynb @@ -58,7 +58,7 @@ "cell_type": "markdown", "metadata": {}, "source": [ - "![Server_Image](files/vectoriser.png)\n", + "![Server_Image](./files/vectoriser.png)\n", "\n", "As seen above, a Vectorisers' sole responsibility is to convert text to a vector representation. Each Vectoriser class must implement a transform() method that will:\n", "\n", diff --git a/DEMO/data/fake_multilingual_soc_dataset.csv b/DEMO/data/fake_multilingual_soc_dataset.csv new file mode 100644 index 0000000..09b84fd --- /dev/null +++ b/DEMO/data/fake_multilingual_soc_dataset.csv @@ -0,0 +1,84 @@ +label,text +101,"Fruit farmer: Grows and harvests fruits such as apples, oranges, and berries." +101,"Maraîcher : cultive et récolte des légumes comme les carottes, les pommes de terre et la laitue." +101,"Cultivateur de vergers : s'occupe de la récolte des pommes et des poires en saison." +102,"Dairy farmer: Manages cows for milk production and processes dairy products." +102,"Éleveur de moutons : élève des moutons pour la laine, la viande et d'autres produits." +103,"Construction laborer: Performs physical tasks on construction sites, such as digging and carrying materials." +103,"Maçon : pose des briques, des blocs de béton et des pierres pour construire des murs et des structures." +103,"Ouvrier du bâtiment : travaille sur différents chantiers, transport de matériaux et travaux de terrassement." +104,"Carpenter: Constructs, installs, and repairs wooden frameworks and structures." +104,"Charpentier : construit, installe et répare des structures et charpentes en bois." +105,"Electrician: Installs, maintains, and repairs electrical systems in buildings and equipment." +106,"Plumber: Installs and repairs water, gas, and drainage systems in homes and businesses." +106,"Plombier : installe et répare des systèmes d'eau, de gaz et de drainage dans les foyers et les entreprises." +107,"Software developer: Designs, writes, and tests computer programs and applications." +107,"Développeur web : conçoit et maintient des sites web et des applications web." +107,"Ingénieur logiciel : participe à la conception et aux tests de nouveaux systèmes informatiques." +108,"Data analyst: Analyzes data to provide insights and support decision-making." +109,"Accountant: Prepares and examines financial records, ensuring accuracy and compliance with regulations." +109,"Auditeur : examine les états financiers et les registres pour assurer la conformité et détecter les fraudes." +110,"Teacher: Educates students in schools, colleges, or universities." +110,"Enseignant : éduque les élèves dans les écoles, les collèges ou les universités." +111,"Nurse: Provides medical care and support to patients in hospitals, clinics, or homes." +111,"Ambulancier : intervient dans des situations médicales d'urgence et fournit des soins préhospitaliers." +111,"Aide-soignant : accompagne les patients au quotidien dans les établissements de santé." +112,"Chef: Prepares and cooks meals in restaurants, hotels, or other food establishments." +112,"Serveur : sert de la nourriture et des boissons aux clients dans les restaurants et les cafés." +113,"Graphic designer: Creates visual concepts for advertisements, websites, and branding." +113,"Illustrateur : produit des œuvres pour des livres, des magazines et d'autres médias." +114,"Mechanic: Repairs and maintains vehicles and machinery." +114,"Technicien automobile : diagnostique et répare les problèmes des voitures et des camions." +115,"Photographer: Captures images for events, advertising, or artistic purposes." +115,"Vidéaste : enregistre et monte du contenu vidéo pour divers projets." +116,"Barista: Prepares and serves coffee and other beverages in cafes and coffee shops." +116,"Gérant de café : supervise les opérations quotidiennes d'un café, y compris la gestion du personnel et des stocks." +117,"Fitness trainer: Designs and leads exercise programs for individuals or groups." +117,"Professeur de yoga : enseigne des cours de yoga pour améliorer la flexibilité, la force et la relaxation." +118,"Librarian: Manages library resources and assists patrons with research." +118,"Archiviste : préserve et organise des documents et registres historiques." +119,"Journalist: Researches and writes news articles for print or online media." +119,"Rédacteur : révise et corrige le contenu écrit en vue de sa publication." +119,"Rédacteur technique : produit des manuels et de la documentation pour des produits techniques." +120,"Scientist: Conducts research and experiments to advance knowledge in a specific field." +120,"Technicien de laboratoire : assiste dans les expériences scientifiques et entretient le matériel de laboratoire." +121,"Police officer: Enforces laws and maintains public safety." +121,"Détective : enquête sur des crimes et rassemble des preuves pour des affaires judiciaires." +122,"Firefighter: Responds to emergencies and extinguishes fires." +122,"Inspecteur incendie : inspecte les bâtiments pour assurer la conformité aux réglementations de sécurité incendie." +123,"Pilot: Operates aircraft to transport passengers or cargo." +123,"Hôtesse de l'air : assure la sécurité et le confort des passagers pendant les vols." +124,"Actor: Performs roles in theater, film, or television productions." +124,"Réalisateur : supervise les aspects créatifs d'un film, d'une pièce de théâtre ou d'une émission de télévision." +125,"Musician: Performs music as a soloist or part of a band or orchestra." +125,"Compositeur : crée de la musique originale pour des spectacles ou des enregistrements." +126,"Athlete: Competes in sports at a professional or amateur level." +126,"Entraîneur sportif : forme et encadre des athlètes pour améliorer leurs performances." +127,"Fashion designer: Creates clothing and accessories for production or custom orders." +127,"Tailleur : ajuste et répare des vêtements pour les clients." +128,"Real estate agent: Assists clients in buying, selling, or renting properties." +128,"Gestionnaire immobilier : supervise l'entretien et la location de biens immobiliers." +129,"Event planner: Organizes and coordinates events such as weddings and conferences." +129,"Coordinateur de mariage : gère tous les aspects de la planification et de l'exécution d'un mariage." +130,"Veterinarian: Provides medical care for animals." +130,"Technicien vétérinaire : assiste les vétérinaires dans le traitement et les soins aux animaux." +131,"Social worker: Supports individuals and families in overcoming challenges." +131,"Conseiller conjugal : accompagne les couples et les familles en difficulté." +132,"Entrepreneur: Starts and manages their own business ventures." +132,"Consultant en entreprise : conseille les entreprises sur des stratégies pour améliorer leurs opérations." +133,"Logistics manager: Oversees the transportation and storage of goods." +133,"Analyste de la chaîne d'approvisionnement : optimise le flux des biens et des matériaux dans une chaîne d'approvisionnement." +134,"Software tester: Tests software applications for bugs and usability issues." +134,"Ingénieur DevOps : gère les processus de développement logiciel et d'exploitation informatique." +135,"Research assistant: Supports researchers in conducting studies and experiments." +135,"Statisticien : analyse des données pour identifier des tendances et faire des prévisions." +136,"HR manager: Oversees recruitment, training, and employee relations." +136,"Chargé de recrutement : trouve et embauche des candidats pour des postes vacants." +137,"Chef de cuisine: Leads the kitchen team in a restaurant." +137,"Second de cuisine : assiste le chef principal dans la gestion des opérations en cuisine." +138,"Bartender: Mixes and serves drinks to customers." +138,"Sommelier : recommande et sert des vins dans les restaurants." +139,"Tour guide: Leads groups on tours and provides information about destinations." +139,"Agent de voyages : planifie et réserve des arrangements de voyage pour les clients." +140,"Game designer: Creates concepts and mechanics for video games." +140,"Concepteur de niveaux : conçoit des niveaux et des environnements pour des jeux vidéo." \ No newline at end of file diff --git a/DEMO/evaluation_workflow_demo.ipynb b/DEMO/evaluation_workflow_demo.ipynb index 89243e0..4b47c7d 100644 --- a/DEMO/evaluation_workflow_demo.ipynb +++ b/DEMO/evaluation_workflow_demo.ipynb @@ -29,7 +29,7 @@ "\n", "Currently the evaluation module only evaluates single label predictions meaning that, while ClassifAI is designed to return a ranked list of several semantically similar candidate entries to a provided query sample, only the top result will be considered when comparing the VectorStore result to a ground truth label provided by a user.\n", "\n", - "![top_1_eval_image](files/eval_top_1_diagram.png)\n", + "![top_1_eval_image](./files/eval_top_1_diagram.png)\n", "\n", "The Evaluation module is currently in development, and in the future its feature set may be extended to include a broader range of evaluation tasks such as multi-class multi-label classification, where potentially multiple labels for a ground truth sample can be compared and evaluated against multiple ranked candidate predictions of the VectorStore." ] @@ -353,7 +353,7 @@ ], "metadata": { "kernelspec": { - "display_name": "Python 3 (ipykernel)", + "display_name": "classifai", "language": "python", "name": "python3" }, @@ -367,7 +367,7 @@ "name": "python", "nbconvert_exporter": "python", "pygments_lexer": "ipython3", - "version": "3.12.4" + "version": "3.13.7" } }, "nbformat": 4, diff --git a/DEMO/files/vectoriser_multilingual.png b/DEMO/files/vectoriser_multilingual.png new file mode 100644 index 0000000..c1f2d11 Binary files /dev/null and b/DEMO/files/vectoriser_multilingual.png differ diff --git a/DEMO/files/vectorstore_2d_vis_multilingual.png b/DEMO/files/vectorstore_2d_vis_multilingual.png new file mode 100644 index 0000000..587c88d Binary files /dev/null and b/DEMO/files/vectorstore_2d_vis_multilingual.png differ diff --git a/DEMO/general_workflow_demo.ipynb b/DEMO/general_workflow_demo.ipynb index 135cb53..d0bf104 100644 --- a/DEMO/general_workflow_demo.ipynb +++ b/DEMO/general_workflow_demo.ipynb @@ -39,7 +39,7 @@ "source": [ "## Vectorising\n", "\n", - "![Vectoriser_image](files/vectoriser.png)\n", + "![Vectoriser_image](./files/vectoriser.png)\n", "\n", "#### We provide several vectoriser classes that you can use to convert text to embeddings/vectors;\n", "```python\n", @@ -114,7 +114,7 @@ "\n", "By default, the vector database is persisted to a local directory named after the input filename. You can use the `output_dir` argument to change the location of the persisted vector database when creating the VectorStore. If the directory already exists, it will exit with a warning - you can pass the `overwrite=True` argument to permit it to overwrite an existing directory. If you don't want the vector database to be persisted at all, you can pass the `skip_save=True` argument - note that this takes precedence over `output_dir` and `overwrite`.\n", "If you want to minimise verbose output, such as INFO-level logs and progress bars, you can set the `quiet_mode` parameter to `True`.\n", - "![VectorStore_image](files/VectorStore.png)\n" + "![VectorStore_image](./files/VectorStore.png)\n" ] }, { @@ -384,7 +384,7 @@ "\n", "#### *Now, how do I host it so others can use it?*\n", "\n", - "![Server_Image](files/servers.png)" + "![Server_Image](./files/servers.png)" ] }, { diff --git a/DEMO/multilingual_datasets_and_vectorisers.ipynb b/DEMO/multilingual_datasets_and_vectorisers.ipynb new file mode 100644 index 0000000..358f987 --- /dev/null +++ b/DEMO/multilingual_datasets_and_vectorisers.ipynb @@ -0,0 +1,397 @@ +{ + "cells": [ + { + "cell_type": "markdown", + "metadata": {}, + "source": [ + "# Working with datasets in multiple languages with ClassifAI and Multilingual Vectoriser Models" + ] + }, + { + "cell_type": "markdown", + "metadata": {}, + "source": [ + "ClassifAI is a tool to help in the creation and serving of searchable vector databases, for text classification tasks. \n", + "\n", + "It has three core components:\n", + "\n", + "1. Vectorisers - Models for converting text to vectors\n", + "2. Indexers - Classes for building VectorStores from text datasets, which you can search\n", + "3. Servers - Allow you to deploy VectorStores with a Rest-API interface\n" + ] + }, + { + "cell_type": "markdown", + "metadata": {}, + "source": [ + "This notebook shows how to work with text data written in more than one language.\n", + "\n", + "If your dataset contains, for example, both English and French text, you can build a ClassifAI `VectorStore` that can be searched in **either** language — and get relevant results back in **both** languages, regardless of which language you searched in.\n", + "\n", + "This works because of **multilingual encoder models**. These are Vectoriser models that convert text from many different languages into a shared embedding space, so that text with the same meaning ends up with similar embeddings, no matter what language it's written in.\n", + "\n", + "This notebook uses a multilingual model from HuggingFace, along with a dataset containing multiple languages, to demonstrate this in practice." + ] + }, + { + "cell_type": "markdown", + "metadata": {}, + "source": [ + "## In this Notebook...\n", + "\n", + "We will show:\n", + "\n", + "* The core workings of the Vectoriser Class and its responsibilities, in a multilingual setting.\n", + "* How to use multilingual encoding models from HuggingFace within ClassifAI using ClassifAI's `HuggingFaceVectoriser` class\n", + "* Examples of building a `VectorStore` knowledgebase containing English and non-English text, and examples of searching the knowledgebase with English and non-English queries." + ] + }, + { + "cell_type": "markdown", + "metadata": {}, + "source": [ + "## Vectorisers and Multilingual Encoding" + ] + }, + { + "cell_type": "markdown", + "metadata": {}, + "source": [ + "![English_Vectoriser_Image](./files/vectoriser.png)\n", + "\n", + "\n", + "As seen above, a Vectorisers' sole responsibility is to convert text to a vector representation. Each Vectoriser class must implement a `transform()` method that will:\n", + "\n", + "1. accept a string or list of N strings as an argument,\n", + "2. return a numpy array of dimension [N,D] where N matches the number of input strings, and D is the embedding dimension.\n", + "\n", + "By enforcing this, the Indexers and Servers modules can reliably work with any Vectoriser object to perform the various search/classification functions required by ClassifAI.\n", + "\n", + "All a developer has to consider when building their own Vectoriser is the logic of this `transform()` method." + ] + }, + { + "cell_type": "markdown", + "metadata": {}, + "source": [ + "![Multilingual_Vectoriser_Image](./files/vectoriser_multilingual.png)\n", + "\n", + "Some Vectoriser (embedding) models are trained to understand many languages at once. These are called **multilingual encoder models**.\n", + "\n", + "As shown above, sentences in different languages go into the model, and come out as embeddings in a **shared embedding space**. If two sentences mean the same thing — even if they're written in different languages — their embeddings will be very similar.\n", + "\n" + ] + }, + { + "cell_type": "markdown", + "metadata": {}, + "source": [ + "In the context of ClassifAI, this means we can build a `VectorStore` from a CSV file containing text in multiple languages, and the resulting embeddings will still reflect the *meaning* of the text — not just its language.\n", + "\n", + "In the diagram below, each dot represents a piece of text, coloured by language (black = English, blue = French, green = Italian). Dots that are **close together** represent sentences that mean similar things, even though they're written in different languages:\n", + "\n", + "![Multilingual_Vectoriser_Image](./files/vectorstore_2d_vis_multilingual.png)" + ] + }, + { + "cell_type": "markdown", + "metadata": {}, + "source": [ + "The `VectorStore` can be searched as normal, using the same functionality as any other ClassifAI use case — but because its using a multilingual Vectoriser, the search query can be written in **any language** the model supports, and it will still return relevant results regardless of the language of the original text." + ] + }, + { + "cell_type": "markdown", + "metadata": {}, + "source": [ + "## Example Implementation" + ] + }, + { + "cell_type": "markdown", + "metadata": { + "vscode": { + "languageId": "plaintext" + } + }, + "source": [ + "### Vectoriser" + ] + }, + { + "cell_type": "markdown", + "metadata": {}, + "source": [ + "This implementation requires an appropriate embedding model that works in a multilingual fashion. For this demo we've chosen to use `Granite-Embedding-97M-Multilingual-R2` - an embedding model provided by IBM and available on HuggingFace that supports over 200 languages. For more information check out the model (and find other multilingual embedding models) on HuggingFace at: https://huggingface.co/ibm-granite/granite-embedding-97m-multilingual-r2\n", + "\n", + "\n", + "With ClassifAI, HuggingFace embedding models can be loaded with the Vectorisers module's `HuggingFaceVectoriser` class, the exact same way a monolingual embedding model would be loaded." + ] + }, + { + "cell_type": "code", + "execution_count": null, + "metadata": {}, + "outputs": [], + "source": [ + "from classifai.vectorisers import HuggingFaceVectoriser\n", + "\n", + "multilingual_vectoriser = HuggingFaceVectoriser(model_name=\"ibm-granite/granite-embedding-97m-multilingual-r2\")" + ] + }, + { + "cell_type": "markdown", + "metadata": {}, + "source": [ + "The vectoriser's `transform()` method can be called that will convert text to embedding representation." + ] + }, + { + "cell_type": "code", + "execution_count": null, + "metadata": {}, + "outputs": [], + "source": [ + "embedding_from_english = multilingual_vectoriser.transform(\"ambulance driver\")\n", + "\n", + "embedding_from_french = multilingual_vectoriser.transform(\"conducteur d'ambulance\")\n", + "\n", + "embedding_from_italian = multilingual_vectoriser.transform(\"autista di ambulanza\")" + ] + }, + { + "cell_type": "code", + "execution_count": null, + "metadata": {}, + "outputs": [], + "source": [ + "print(embedding_from_english.shape)\n", + "print(embedding_from_french.shape)\n", + "print(embedding_from_italian)" + ] + }, + { + "cell_type": "markdown", + "metadata": {}, + "source": [ + "### Dataset" + ] + }, + { + "cell_type": "markdown", + "metadata": {}, + "source": [ + "For this example notebook, Generative AI was used to make an example dataset that contains fake SOC data, with text written in English and French. This dataset can be used, with the Granite-embedding model to build a ClassifAI VectorStore. Then this notebook will try out searching the `VectorStore` in a variety of languages (including languages other than English and French).\n", + "\n", + "The below cell loads the CSV file into a Pandas dataframe, and displays the top 5 entries to showcase the content. It contains profession names with a short description of the work with a corresponding 'fake' SOC label. \n" + ] + }, + { + "cell_type": "code", + "execution_count": null, + "metadata": {}, + "outputs": [], + "source": [ + "import pandas as pd\n", + "\n", + "multilingual_dataset = pd.read_csv(\"./data/fake_multilingual_soc_dataset.csv\")\n", + "multilingual_dataset.head()" + ] + }, + { + "cell_type": "markdown", + "metadata": {}, + "source": [ + "### Building the VectorStore\n", + "\n", + "The process for creating a `VectorStore` is identical to the case of making a `VectorStore` for a CSV file of text all in the same language." + ] + }, + { + "cell_type": "code", + "execution_count": null, + "metadata": {}, + "outputs": [], + "source": [ + "from classifai.indexers import VectorStore\n", + "\n", + "# this is standard code to construct a VectorStore from a CSV file of data and using an instantiated Vectoriser model.\n", + "demo_vectorstore = VectorStore(\n", + " file_name=\"./data/fake_multilingual_soc_dataset.csv\",\n", + " data_type=\"csv\",\n", + " vectoriser=multilingual_vectoriser,\n", + " skip_save=True,\n", + ")" + ] + }, + { + "cell_type": "markdown", + "metadata": {}, + "source": [ + "### Searching with English queries\n", + "\n", + "With the `VectorStore` object instantiated, now call the `search()` method - it can be seen that the below search query passed in English returns relevant results in multiple languages." + ] + }, + { + "cell_type": "code", + "execution_count": null, + "metadata": {}, + "outputs": [], + "source": [ + "from classifai.indexers.dataclasses import VectorStoreSearchInput\n", + "\n", + "# creating a VectorStoreSearchInput object to pass to the search method\n", + "english_search_input = VectorStoreSearchInput({\"id\": [1], \"query\": [\"medical doctor\"]})\n", + "\n", + "# calling the search method and displaying the results\n", + "demo_vectorstore.search(english_search_input, n_results=5)" + ] + }, + { + "cell_type": "markdown", + "metadata": {}, + "source": [ + "The resulting output of the cell above shows that while there isn't a medical doctor profession listed in our dataset, the top few results are all medical related results from both French and English, despite the original query being only in English. " + ] + }, + { + "cell_type": "markdown", + "metadata": {}, + "source": [ + "### Searching with non-English Queries\n", + "\n", + "In the cell below, an example query is written in French.\n", + "\n", + "The English translation of the below query is \"A person who repairs cars and trucks\"" + ] + }, + { + "cell_type": "code", + "execution_count": null, + "metadata": {}, + "outputs": [], + "source": [ + "# creating a VectorStoreSearchInput object to pass to the search method, this time with french\n", + "french_search_input = VectorStoreSearchInput(\n", + " {\"id\": [1], \"query\": [\"Une personne qui répare des voitures et des camions\"]}\n", + ")\n", + "\n", + "# calling the search method and displaying the results\n", + "demo_vectorstore.search(french_search_input, n_results=5)" + ] + }, + { + "cell_type": "markdown", + "metadata": {}, + "source": [ + "Among the top results of the previous cell's output is the 'Mechanic\" entry from the dataset which shows that the VectorStore is accepting a French query and returning an English result. The other top results are in French but are also seemingly relevant." + ] + }, + { + "cell_type": "markdown", + "metadata": {}, + "source": [ + "The VectorStore can also be searched in other languages that are not English or French, because as mentioned earlier the IBM Granite embedding model supports many languages.\n", + "\n", + "Below the query \"Una persona che progetta e testa programmi informatici e applicazioni\" is passed to the VectorStore `search()` method - the English translation is: \"A person who designs and tests computer programs and applications\"" + ] + }, + { + "cell_type": "code", + "execution_count": null, + "metadata": {}, + "outputs": [], + "source": [ + "# creating a VectorStoreSearchInput object to pass to the search method, this time with italian\n", + "italian_search_input = VectorStoreSearchInput(\n", + " {\"id\": [1], \"query\": [\"Una persona che progetta e testa programmi informatici e applicazioni\"]}\n", + ")\n", + "\n", + "# calling the search method and displaying the results\n", + "demo_vectorstore.search(italian_search_input, n_results=5)" + ] + }, + { + "cell_type": "markdown", + "metadata": {}, + "source": [ + "The above output shows several highly ranked results that are relevant from both French and English sources including \"Ingénieur logiciel\" (Software Engineer), \"Software Developer\" (which is an English result), and \"Développeur web\" (the French for Web Developer)" + ] + }, + { + "cell_type": "markdown", + "metadata": {}, + "source": [ + "### Querying in many languages" + ] + }, + { + "cell_type": "markdown", + "metadata": {}, + "source": [ + "Finally, its also possible to pass multiple queries to the search method at one time where the queries are in different languages. In the next cell, the same query is passed in English, French and Italian.\n", + "\n", + "\n", + "The results for each query should be similar, as the queries are translations of one another." + ] + }, + { + "cell_type": "code", + "execution_count": null, + "metadata": {}, + "outputs": [], + "source": [ + "# three queries that all mean the same thing in different languages\n", + "query_en = \"A craftsman who builds and repairs wooden furniture and structures\"\n", + "query_fr = \"Un artisan qui construit et répare des meubles et des structures en bois\"\n", + "query_it = \"Un artigiano che costruisce e ripara mobili e strutture in legno\"\n", + "\n", + "# creating a input object for the search method.\n", + "search_input_multiple = VectorStoreSearchInput({\"id\": [1, 2, 3], \"query\": [query_en, query_fr, query_it]})\n", + "\n", + "# searching, retrieving the top 3 candidates for each query.\n", + "demo_vectorstore.search(search_input_multiple, n_results=3)" + ] + }, + { + "cell_type": "markdown", + "metadata": {}, + "source": [ + "While there is some difference in the top 3 results for each query, there is a large amount of overlap and generally the results are highly relevant throughout." + ] + }, + { + "cell_type": "markdown", + "metadata": {}, + "source": [ + "## Thats it!\n", + "\n", + "* When choosing an embedding model from HuggingFace or from another service such as GCP, be sure to check which languages the embedding model supports. \n", + "* Generally, monolingual models perform stronger on tasks for a specific single language than a multilingual model would. \n", + "* Because these models are accessible through ClassifAI's `HuggingFaceVectoriser` class, these models are directly compatible with the other modules of the Package including the Servers module and Evaluation module," + ] + } + ], + "metadata": { + "kernelspec": { + "display_name": "classifai", + "language": "python", + "name": "python3" + }, + "language_info": { + "codemirror_mode": { + "name": "ipython", + "version": 3 + }, + "file_extension": ".py", + "mimetype": "text/x-python", + "name": "python", + "nbconvert_exporter": "python", + "pygments_lexer": "ipython3", + "version": "3.13.7" + } + }, + "nbformat": 4, + "nbformat_minor": 2 +}