diff --git a/docs/assets/images/guides/mountable_secrets/account-mountable-secrets.png b/docs/assets/images/guides/mountable_secrets/account-mountable-secrets.png new file mode 100644 index 0000000000..9d1583fcdb Binary files /dev/null and b/docs/assets/images/guides/mountable_secrets/account-mountable-secrets.png differ diff --git a/docs/assets/images/guides/trino/catalog-sharing-page.png b/docs/assets/images/guides/trino/catalog-sharing-page.png new file mode 100644 index 0000000000..d47155aa0f Binary files /dev/null and b/docs/assets/images/guides/trino/catalog-sharing-page.png differ diff --git a/docs/assets/images/guides/trino/catalogs-list.png b/docs/assets/images/guides/trino/catalogs-list.png index 3a4ca3c3e6..87fb18a0e1 100644 Binary files a/docs/assets/images/guides/trino/catalogs-list.png and b/docs/assets/images/guides/trino/catalogs-list.png differ diff --git a/docs/assets/images/guides/trino/create-private-catalog.png b/docs/assets/images/guides/trino/create-private-catalog.png new file mode 100644 index 0000000000..b4a571b244 Binary files /dev/null and b/docs/assets/images/guides/trino/create-private-catalog.png differ diff --git a/docs/assets/images/guides/trino/fg-share-subset.png b/docs/assets/images/guides/trino/fg-share-subset.png new file mode 100644 index 0000000000..fe6e3603a4 Binary files /dev/null and b/docs/assets/images/guides/trino/fg-share-subset.png differ diff --git a/docs/assets/images/guides/trino/fg-share-whole.png b/docs/assets/images/guides/trino/fg-share-whole.png new file mode 100644 index 0000000000..5228e76ff1 Binary files /dev/null and b/docs/assets/images/guides/trino/fg-share-whole.png differ diff --git a/docs/assets/images/guides/trino/query-masked-column.png b/docs/assets/images/guides/trino/query-masked-column.png new file mode 100644 index 0000000000..60cd84fe3a Binary files /dev/null and b/docs/assets/images/guides/trino/query-masked-column.png differ diff --git a/docs/assets/images/guides/trino/query-shared-features.png b/docs/assets/images/guides/trino/query-shared-features.png new file mode 100644 index 0000000000..68c5ab5f01 Binary files /dev/null and b/docs/assets/images/guides/trino/query-shared-features.png differ diff --git a/docs/assets/images/guides/trino/query-unshared-feature-denied.png b/docs/assets/images/guides/trino/query-unshared-feature-denied.png new file mode 100644 index 0000000000..87a98a4aae Binary files /dev/null and b/docs/assets/images/guides/trino/query-unshared-feature-denied.png differ diff --git a/docs/assets/images/guides/trino/received-shares.png b/docs/assets/images/guides/trino/received-shares.png new file mode 100644 index 0000000000..17aaf10882 Binary files /dev/null and b/docs/assets/images/guides/trino/received-shares.png differ diff --git a/docs/assets/images/guides/trino/share-dialog-table.png b/docs/assets/images/guides/trino/share-dialog-table.png new file mode 100644 index 0000000000..fd411dc23c Binary files /dev/null and b/docs/assets/images/guides/trino/share-dialog-table.png differ diff --git a/docs/assets/images/guides/trino/share-edit-columns.png b/docs/assets/images/guides/trino/share-edit-columns.png new file mode 100644 index 0000000000..5ae3d36326 Binary files /dev/null and b/docs/assets/images/guides/trino/share-edit-columns.png differ diff --git a/docs/setup_installation/admin/superset.md b/docs/setup_installation/admin/superset.md index ccc6857e9b..3139adfc32 100644 --- a/docs/setup_installation/admin/superset.md +++ b/docs/setup_installation/admin/superset.md @@ -204,9 +204,11 @@ In **Cluster Settings**, choose **Configuration** under _Infrastructure_ in the #### trino_default_catalog - **Description**: Default catalog to use for the Offline Feature Store Connection -- **Default**: `hive` +- **Default**: `delta` - **Values**: `hive`, `delta`, `iceberg`, and `hudi`. +The value is applied when a member's connection is created, so connections created before a change keep the catalog they were created with. + Trino connections are created with multi-catalog enabled (Superset's **Allow changing catalogs** option), so users are not limited to this default. They select the catalog matching their table format from the **Catalog** dropdown in SQL Lab or when adding a dataset. Querying a table whose format does not match the selected catalog raises a Trino `UNSUPPORTED_TABLE_TYPE` error; see the [Trino table type error][trino-table-type-error] troubleshooting in the Superset user guide. diff --git a/docs/setup_installation/admin/trino.md b/docs/setup_installation/admin/trino.md index e71109eeab..cef3df9569 100644 --- a/docs/setup_installation/admin/trino.md +++ b/docs/setup_installation/admin/trino.md @@ -224,6 +224,136 @@ A catalog whose `${HOPSWORKS_SECRET:}` reference no longer resolves cannot The repair reports it, leaves any file it already has in place, because that copy resolved when it was approved and still works, and carries on with every other catalog. Its owner has to repoint the reference at an existing secret. +## Access control and sharing + +The query engine decides who can read what with Trino's file-based access control, from a rules file published into the Trino files store as `access-control/rules.json`. +Hopsworks owns that file and rebuilds it whenever a share changes, and on a schedule every five minutes by default. + +The file is composed from two parts: + +- The base policy, from the Helm value `trino.accessControl.rules`, which the chart renders into the ConfigMap `hopsworks-trino-access-control-base`. + It grants each project its own catalogs and feature store, and each user their private catalogs, written only from projects where the user is a Data Owner. + Administrators see every catalog but read only `system`, `tpch` and `tpcds`, because the query engine's administrator is also the identity Hopsworks itself uses, and a view recorded as owned by it would otherwise read any project's data. + The administrators' SQL console therefore cannot read a project's tables; query them as a member of the project. +- One set of rules per share, for [catalog shares and feature group shares][sharing-catalogs-and-feature-groups]. + A share names the receiving project's existing `__data_owner` and `__data_scientist` groups, so sharing never changes the group file. + +Change the base policy through the Helm value and an upgrade. +An edit to the published `rules.json` is overwritten by the next rebuild, within minutes. + +Every rebuilt file is validated before it is published, and a file that fails validation is not published: the shares that caused it are marked **Failed** with the reason, and the file in place stays as it was. +After publishing, Hopsworks checks that the query engine still answers once it has re-read the file, and restores the last file that worked if it does not, because Trino refuses every query while its rules file is unreadable. + +### The shared feature store catalogs + +The chart ships two kinds of catalog over the feature store: + +- `delta`, `hudi`, `iceberg` and `hive` impersonate the querying user, so HopsFS permissions apply on top of the access-control rules. + They serve a project's own feature groups, and feature groups or feature stores shared whole, which HopsFS grants the receiving project. +- `delta_shared`, `hudi_shared` and `iceberg_shared` do not impersonate. + They read HopsFS as the `trino` user, which is a HopsFS superuser, because a feature group shared with a subset of its features grants the receiving project no HopsFS access. + +For the second kind the access-control rules are the only gate. +The base policy grants nobody access to them, not even administrators, and Hopsworks adds a rule per subset share that allows the receiving project the shared features of that one table and denies the rest. +They are read-only at the connector as well, so no rule can let a query write through them. +Do not add rules for these catalogs to the base policy: any rule that reaches one of them reads every project's feature store. +Administrators are denied them because Trino runs a view as the user recorded as its owner, so a view recorded as owned by an administrator would reach them too. + +### Reading the rules file + +The rules the query engine enforces can be read under **Cluster Settings** → **Query Engine** → **Files**, as `access-control/rules.json`. +Beside it, `access-control/rules.json.last-good` is the last file the query engine loaded. +They differ from a publish until Hopsworks confirms the query engine loaded the new file, a few seconds later. +If they stay different, the new file is not confirmed yet, for example because the query engine was unreachable, and the next reconcile checks again. +A file the query engine refused does not stay: the last good file goes back in its place, and the shares the refused file added are marked **Failed**. +The groups the rules name are in `auth/group.db`. + +#### Who the rules match + +A query runs as a principal named `__`, for example `seeda__seed1000` for user `seed1000` in project `seeda`. +Its groups are the member's role in that project, `__data_owner` or `__data_scientist`, and `__shared_featurestore` for each project `` whose feature store is shared with that project. +Group `admin` has one member, the query engine administrator, which is the identity Hopsworks itself uses. +Project names and usernames cannot contain `__`, so a pattern such as `.*__(.*)` splits a principal unambiguously, and `$1` in a later field stands for what the pattern captured. + +Each section of the file (`catalogs`, `schemas`, `tables`, `functions`, `queries`) is checked on its own. +In a section, the first rule whose user, group and object all match decides, and a request no rule matches is denied. +The order of the rules is therefore the policy: a broader rule placed first would answer before a narrower one. + +#### The order of the rules + +Every section keeps the same order, and the rules Hopsworks adds for shares go in one place in it: + +1. The administrator rules. +2. A deny for each private catalog whose owner's account was deleted, until the catalog is removed. + It comes before the private-owner rules because a later account with the same username would match them. +3. The private-owner rules. + They come before the shares so that sharing a private catalog with a project the owner belongs to never narrows the owner's own access. +4. The share rules. +5. The rest of the base policy: every project's own catalogs and feature store. + +#### The base policy + +The `catalogs` section of the base policy, in order: + +| Rule | Effect | +| --- | --- | +| `group: admin`, `allow: none` on `iceberg_shared`, `delta_shared` and `hudi_shared` | The administrator never sees the shared feature store catalogs. | +| `group: admin`, `catalog: .*`, `allow: read-only` | The administrator sees every other catalog, without writing to any. | +| `user: .*__(.*)`, `group: .*__data_owner`, `catalog: _$1__.*`, `allow: all` | The owner of a private catalog reads and writes it from a project where they are a Data Owner. | +| `user: .*__(.*)`, `catalog: _$1__.*`, `allow: read-only` | The owner reads it from any other project. | +| `catalog: tpch` and `tpcds`, `allow: read-only` | Everyone reads the sample catalogs. | +| `catalog: iceberg`, `delta`, `hive`, `hudi`, `allow: all` | Everyone reaches the feature store catalogs; the table rules decide what they read. | +| `group: (.*)__data_owner`, `catalog: $1__.*`, `allow: all` | A project's Data Owners read and write its catalogs. | +| `group: (.*)__data_scientist`, `catalog: $1__.*`, `allow: read-only` | Its Data Scientists read them. | +| `catalog: system`, `allow: read-only` | Everyone reads the `system` catalog. | + +The `tables` section follows the same pattern: the administrator reads only `system`, `tpch` and `tpcds`, the private-owner rules mirror the catalog ones, and each project reaches the schema `_featurestore` in the feature store catalogs, all of it for its Data Owners, reading for its Data Scientists and for projects its feature store is shared with. +The `schemas` section gives schema ownership, which is what creating and dropping schemas needs, to Data Owners only. +The `functions` section lets everyone run builtin functions, and the Data Owners of a project run the `system` functions of their project's catalogs, such as `system.query` on a JDBC catalog. + +#### The rules a share adds + +Hopsworks writes the names in a share rule as literals between `\Q` and `\E`, so a name containing regular expression syntax matches only itself. +A share names the receiving project's two role groups in one pattern, `\Q\E__data_(?:owner|scientist)`, so searching the file for `\Qseedc\E__data_` finds every rule a share to `seedc` added. + +A share of catalog `seeda__postgresql` with `seedc`, covering table `public.customers` with column `created` unchecked and column `name` masked, adds these rules: + +```json +{"catalogs": [ + {"group": "\\Qseedc\\E__data_(?:owner|scientist)", "catalog": "\\Qseeda__postgresql\\E", "allow": "read-only"} +], +"tables": [ + {"group": "\\Qseedc\\E__data_(?:owner|scientist)", "catalog": "\\Qseeda__postgresql\\E", + "schema": "\\Qpublic\\E", "table": "\\Qcustomers\\E", "privileges": ["SELECT"], + "columns": [{"name": "created", "allow": false}, {"name": "$path", "allow": false}, + {"name": "name", "mask": "'***'"}]}, + {"group": "\\Qseedc\\E__data_(?:owner|scientist)", "catalog": "\\Qseeda__postgresql\\E", + "schema": "\\Qpublic\\E", "table": "\\Qcustomers\\E\\$.*", "privileges": []} +], +"functions": [ + {"group": "\\Qseedc\\E__data_(?:owner|scientist)", "catalog": "\\Qseeda__postgresql\\E", "privileges": []} +]} +``` + +The list of denied columns in the rule is shortened here. + +- The catalog rule makes the catalog visible to the receiving project, read-only. +- The table rule grants `SELECT` on the table and lists the columns it denies: the unchecked ones, and the connector's hidden columns, such as `$path`. + A table shared whole has a table rule without `columns`, a schema shared whole has `table: .*`, and a catalog shared whole has `schema: .*` too. +- The rule after it, with no privileges, denies the table's metadata tables, such as `customers$partitions`, which Trino checks by their own name. +- The function rule denies the receiving project the catalog's functions, which the base policy would otherwise give it. + A share of a private catalog also adds, before that deny, a rule letting the owner keep running the catalog's `system` functions. + +A feature group shared whole adds one table rule on `hive|iceberg|delta|hudi` for its table in the owner's feature store. +A feature group shared with a subset of its features adds a catalog rule on the shared catalog of its format, such as `delta_shared`, and a table rule there that denies every unshared feature and hidden column, followed by the metadata table deny. + +#### Debugging a share + +- The share is **Active** but a query is refused: find the share's rules by the receiving project's group, then look for a rule above them that matches the same principal and object first. +- The share's rules are not in the file: the share is still **Applying**, or it is **Failed** and its status says why. +- A column that should be hidden is readable: it is missing from the rule's `columns`, typically a column added after the share was saved; saving the share again adds it as unshared. +- `rules.json` and `rules.json.last-good` differ for minutes: the newest file is not confirmed, so check that the query engine is reachable; the reconcile verifies it again and restores the last good file if the query engine refuses it. + ## Credential files a project supplies A connector that authenticates with a file, such as an Oracle wallet or a Java keystore, cannot be served by a catalog property alone. @@ -311,7 +441,8 @@ Trino behavior can be customized through cluster configuration variables. To mod **Available Variables:** - **trino_enabled**: Enable or disable Trino cluster-wide (default: `false`) -- **trino_default_catalog**: Default catalog used for Superset queries (default: `hive`) +- **trino_default_catalog**: Default catalog of the Superset database connections created for new project members (default: `delta`). + Connections created before a change keep the catalog they were created with. - **trino_test_coordinator_enabled**: Enable the optional test coordinator that backs the "Test connection" action for user-created catalogs (default: `true`) - **trino_catalog_reconcile_enabled**: Rebuild the user-catalog Secrets from the database on a schedule, for a cluster that has lost them (default: `false`, see [Recovering catalog files lost from the mount][recovering-catalog-files-lost-from-the-mount]) - **trino_catalog_max_per_project**: Catalogs a *newly created* project may create (default: `10`). diff --git a/docs/user_guides/projects/mountable_secrets/mountable_secrets.md b/docs/user_guides/projects/mountable_secrets/mountable_secrets.md index 61d39f753e..1ac7adb934 100644 --- a/docs/user_guides/projects/mountable_secrets/mountable_secrets.md +++ b/docs/user_guides/projects/mountable_secrets/mountable_secrets.md @@ -12,8 +12,9 @@ A **mountable secret** is a named bundle of files that belongs to your project. You upload the files once, then refer to the bundle by name from a catalog property, and Hopsworks substitutes the real location when the catalog is written for Trino. The files are stored where project members cannot read or write them directly, and a catalog can only ever reach its own project's bundles. -Only a project Data Owner can list, create or delete mountable secrets. +Only a project Data Owner can list, create or delete a project's mountable secrets. Through the API the same endpoints need an API key with the `MOUNTABLE_SECRET` scope. +Private catalogs use mountable secrets that belong to your account instead, described in [Mountable secrets for private catalogs][mountable-secrets-for-private-catalogs]. ## Creating a bundle @@ -136,6 +137,27 @@ connection-password=${HOPSWORKS_SECRET:oracle_password} Take the host, port and `service_name` from the alias's entry in the wallet's `tnsnames.ora`, and drop `retry_delay`, which means nothing once `retry_count` is zero. A catalog can keep this form, and doing so records which consumer group it connects to instead of leaving it to an alias name. +## Mountable secrets for private catalogs + +A [private catalog][private-catalogs] follows its owner into every project they are a member of, so it cannot reference a project's bundles: it would carry them into the owner's other projects. +It references bundles that belong to your account instead. + +Open **Account Settings**, then **Secrets**, and use the **Mountable secrets** section below your secrets. +Creating, listing and deleting work as for a project's bundles, and the same limits apply, counted per account rather than per project. + +
+ Mountable secrets on the account Secrets page +
Bundles that belong to your account, for use by your private catalogs from any project
+
+ +A private catalog references your bundles with the same `${HOPSWORKS_MOUNT:}` forms, and a project catalog cannot reference them. +Names, sizes, hashes and timestamps of your bundles are visible to you alone. +The same caution applies as for a project's bundles: where the cluster runs a Trino test coordinator, every mountable secret on it is readable from that coordinator. + +Through the API, your account's bundles are at `/users/mountable-secrets`, with the same operations as a project's and an API key with the `MOUNTABLE_SECRET` scope. + +When your account is deleted, your bundles are deleted with it. + ## When the feature is unavailable An administrator can turn the store off for a whole cluster. diff --git a/docs/user_guides/projects/superset/superset.md b/docs/user_guides/projects/superset/superset.md index a590167b0c..af1cb1f271 100644 --- a/docs/user_guides/projects/superset/superset.md +++ b/docs/user_guides/projects/superset/superset.md @@ -84,7 +84,7 @@ SQL Lab is an interactive SQL query interface for exploring your feature data. For the Trino connection, SQL Lab shows a **Catalog** dropdown between **Database** and **Schema**. -The dropdown appears because the connection has *Allow changing catalogs* enabled, and it lists every Trino catalog (`hive`, `delta`, `iceberg`, `hudi`). +The dropdown appears because the connection has *Allow changing catalogs* enabled, and it lists every Trino catalog you can read (`hive`, `delta`, `iceberg`, `hudi`, your project's catalogs, and catalogs shared with your project). Selecting the catalog that matches your table format lets you query the table without prefixing the catalog in the SQL.
diff --git a/docs/user_guides/projects/trino/catalogs.md b/docs/user_guides/projects/trino/catalogs.md index 9af0e019f0..267373be30 100644 --- a/docs/user_guides/projects/trino/catalogs.md +++ b/docs/user_guides/projects/trino/catalogs.md @@ -8,13 +8,22 @@ A Trino catalog makes an external data source queryable from the query engine. Each catalog names a Trino connector and the properties that connector needs to reach the source, such as a connection URL and credentials. Once a catalog is live, its databases and tables can be queried from the SQL runner alongside your feature groups. -Navigate to **Query Engine** → **Catalogs** in your project to see the project's catalogs together with the cluster's shared default catalogs. -A catalog you create is named `__` and is queryable only inside your own project. -Only a project Data Owner can create, edit or delete one. +Navigate to **Query Engine** → **Catalogs** in your project to see the project's catalogs and your own private catalogs. +The cluster's shared default catalogs are listed once you add **Default** to the **Type** filter. +A catalog belongs either to the project or to you: + +- A **project catalog** is named `__` and is queryable inside the project. + The project's Data Owners create, edit, delete and share it. +- A **private catalog** is named `___` and follows you into every project you are a member of. + Only you edit, delete or share it. + See [Private catalogs][private-catalogs]. + +Only a Data Owner of the current project can create a catalog of either kind. +A catalog is queryable outside its own project, or outside your projects for a private one, only once it is shared, as described in [Sharing Catalogs and Feature Groups][sharing-catalogs-and-feature-groups].
Catalogs list -
The project's catalogs alongside the cluster's shared default catalogs
+
The project's catalogs, a private catalog, and the cluster's shared default catalogs
A catalog change is recorded immediately, but it reaches the query engine only when the engine restarts, because Trino reads catalogs at startup. @@ -93,6 +102,32 @@ If a catalog needs to outlive your account, have someone recreate it under their Create the secret by typing or pasting the value as text when you intend to reference it from a catalog. For a credential that is naturally a file, use a mountable secret instead. +## Private catalogs + +Choose **Private** as the owner when creating a catalog to make it yours rather than the project's. +The name prefix becomes `___`, for example `_meb10000__sales`, and the catalog is listed in every project you are a member of. + +
+ Creating a private catalog +
A private catalog takes your username as its prefix and can reference only your own secrets
+
+ +A private catalog differs from a project catalog in what it can reach and who controls it: + +- It can reference only your own Hopsworks secrets and your own mountable secrets, which you manage under **Account Settings** → **Secrets**. + A project's mountable secrets are not available to it, because the catalog would carry them into every other project you are a member of. + See [Mountable secrets for private catalogs][mountable-secrets-for-private-catalogs]. +- It cannot be created from a data source, because a data source's credential belongs to the project. + Enter the connection details yourself instead. +- Only you can edit, delete or share it, from any of your projects and whatever your role there, so being made a Data Scientist in a project does not lock you out of your own catalogs. + Other members of your projects cannot query it unless you share it with their project. +- You write to it only from a project where you are a Data Owner, and read it from any other. + In a project where you are a Data Scientist you can only read that project's data, and a private catalog writable from there would let you copy the data into a catalog you then read from your other projects. +- The number of private catalogs you can own has the same limit as a project's catalogs, ten by default. + +When your account is deleted, your private catalogs are marked for removal and stop being queryable at once, and their shares are removed with them. +They stay denied to everyone, including a later account with the same username, until they are removed like any deleted catalog, as described in [When the catalog goes live][when-the-catalog-goes-live]. + ## Testing the connection **Create** stays disabled until **Test connection** succeeds, so a catalog that cannot reach its source is caught now rather than after a restart. @@ -171,9 +206,24 @@ Testing the connection before saving catches most of these earlier. ## Who can query a catalog -Access to a user-created catalog is granted at the catalog level per project: a project's Data Owners can read and write, and Data Scientists can read. -There is no per-schema or per-table configuration for these catalogs. -To limit what a catalog exposes, scope the database user in the connection credentials at the source, since the query engine reads the external system as that user and can only ever see what those credentials allow. +Inside the project that owns a project catalog, its Data Owners can read and write, and its Data Scientists can read. +A private catalog can be read by you from any of your projects, and written only from a project where you are a Data Owner. + +Other projects can read a catalog only through a share, which grants read access to the whole catalog, one schema, one table, or some columns of a table, optionally with masked values. +See [Sharing Catalogs and Feature Groups][sharing-catalogs-and-feature-groups]. + +The query engine reads the external system as the database user in the connection credentials, so no share can expose more than those credentials allow. +Scoping that database user at the source remains the strongest limit on what a catalog can reach. + +A Data Owner of the project can also run a JDBC catalog's `system.query` table function, which passes a query to the source database as that database user. +The query engine sends it as a subquery, so the database rejects a statement that changes data, but a database function that changes data as a side effect still runs. +Give the database user only the privileges the catalog's Data Owners should have. +The receiving project of a share cannot run the catalog's functions at all. + +Read access to an Iceberg or Delta Lake table also allows its table procedures, `ALTER TABLE ... EXECUTE` with `optimize`, `expire_snapshots`, `remove_orphan_files` or `rollback_to_snapshot`, because the query engine does not check them against the access rules. +A Data Scientist of the project can therefore rewrite, expire or roll back the tables of the project's Iceberg and Delta Lake catalogs, and so can you on your private catalog from a project where you are a Data Scientist, although neither can write rows. +Set the connector's `iceberg.security` or `delta.security` property to `read_only` on a catalog that must not be changed this way; the query engine then refuses table procedures to everyone and reading is unchanged. +`CALL` procedures, such as `system.unregister_table`, are refused to everyone. ## Creating a catalog from the Python client diff --git a/docs/user_guides/projects/trino/sharing.md b/docs/user_guides/projects/trino/sharing.md new file mode 100644 index 0000000000..304d5ad066 --- /dev/null +++ b/docs/user_guides/projects/trino/sharing.md @@ -0,0 +1,208 @@ +--- +description: Give another project read access to a Trino catalog or a feature group through the query engine, the whole catalog or chosen schemas, tables and columns, with masked columns. +--- + +# Sharing Catalogs and Feature Groups + +The query engine enforces who can read what, so data one project owns is not visible to another project until it is shared. +Two kinds of share reach the query engine: + +- A **catalog share** gives another project read access to one of your Trino catalogs: the whole catalog, or chosen schemas and tables, down to some columns of a table, optionally with masked values. +- A **feature group share**, made from the feature store, also makes the shared feature group queryable through the query engine in the receiving project. + +A share always grants read access, and always to the receiving project's Data Owners and Data Scientists. +It grants tables only: the receiving project cannot run the catalog's functions, such as `system.query` of a JDBC catalog, which would run any SQL on the source database as the catalog's database user. +Changes take effect within seconds, without restarting the query engine. + +## Sharing a catalog + +A project catalog is shared by a Data Owner of the project, and a private catalog by its owner, from any of their projects. +A catalog can be shared once it is **Approved** and while it is not being deleted, because a catalog the query engine has not loaded has nothing to share yet. +It cannot be shared with the project that owns it. +While a catalog is shared, its connector cannot change, because a narrowed share denies the hidden columns of the connector the query engine runs, and the engine switches connector only at its next restart. +Revoke its shares first; its other properties can be edited as usual. + +Click the share icon on the catalog's row in **Query Engine** → **Catalogs** to open its sharing page. +The page lists every project the catalog is shared with, what it shares with each, and whether each share is live. + +
+ Sharing page of a catalog +
A catalog shared whole with one project, and one table with two of its three columns, one of them masked, with another
+
+ +Click **Share** and choose the project. +A project holds one share of a catalog, which covers everything it receives from the catalog, so a project the catalog is shared with already is not offered: edit its share instead. +Then check what to share in the tree of the catalog: + +- Check the catalog to share every schema and table in it, including ones created later. +- Check a schema to share every table in it, including tables created in it later. +- Expand a schema and check some of its tables to share only those tables. +- Expand a table and uncheck some of its columns to share only the checked columns. A checked column can carry a mask. + +Unchecking something inside a checked schema or catalog keeps the rest of it: uncheck `sales.salaries` in a checked schema `sales`, and every other table of `sales` stays shared. +A partly checked box shares only what is checked under it, so a table created later in a partly checked schema is not shared. + +The panel beside the tree shows what the receiving project will see. +Click a table name to see its first rows as they will read them, with unchecked columns left out and masks applied. +The sample is read as you, before anything is saved. + +
+ Sharing one table with some columns +
Sharing one table, with one column left out and one column masked
+
+ +The schemas, tables and columns offered are the ones you can see in the catalog yourself, read through the catalog's own connection. + +### Sharing some columns of a table + +Expand a table in the tree to choose its columns. +An unchecked column cannot be read by the receiving project, and a query that selects it, or selects `*`, is refused. +The table's other ways of revealing a column are closed too: the connector's hidden columns, such as `$path` and `$partition`, or Elasticsearch's `_source`, which holds the whole document, and the table's metadata tables, such as `$partitions`, are denied on a narrowed table. + +Tables of a Kafka, Redis, MongoDB, Cassandra or Thrift catalog cannot be narrowed to some columns, because those connectors can hide columns that are defined outside the query engine and cannot all be denied. +Share such a table whole, or leave it out. + +The receiving project reads only the columns you checked. +Hopsworks reads the table's columns each time it updates the rules, and at least every five minutes by default, and denies every column you did not check, so a column added to the table later, or renamed at the source, is not shared. +Between the change at the source and the next update, the new column is readable. +A renamed column loses its mask with its old name and is denied under the new one. +The sharing page marks such a share with the number of new columns and names them; edit the share and check them to share them. +While the query engine is restarting, the rules keep the columns recorded when the share was saved until it is back, and a narrowed table whose columns cannot be read is left out of the share until they can. + +Iceberg and Delta Lake tables can also be read as of an earlier version, which has the columns the table had then. +A column dropped or renamed before the table was shared is readable that way, under its old name, until those versions expire. + +
+ Editing the columns of a share +
Editing a share: a table with two columns shared, one of them masked, and one left out
+
+ +### Masking a column + +A shared column can carry a mask, which replaces the value the receiving project reads. +A mask is one SQL expression over the row, for example `'***'` or `regexp_replace(email, '.+@', '***@')`, and it must return the column's own type. + +A mask runs as the person querying, not as you, so it can use only what they can read: the checked columns of the same table. +A mask that refers to an unchecked column is refused when you save the share. +A mask that reads another table is accepted, because it is checked as you, but it fails at query time for anyone in the receiving project who cannot read that table. +It cannot contain `;` or a comment, and it is limited to 2000 characters. +The mask is checked against the table when the share is saved, so an expression the query engine cannot evaluate is reported then rather than when someone queries the table. + +
+ Querying a masked column +
The receiving project reads the masked column as ***
+
+ +You always read your own catalog unmasked, and a share never narrows your own access: in a project a private catalog is shared with, its owner keeps the access described in [Private catalogs][private-catalogs]. + +### Editing a share + +Click the edit icon on a share to change what it covers: add or remove schemas and tables, narrow tables to some columns, or go from chosen schemas and tables to the whole catalog. +Saving reads the narrowed tables' columns again, and the change is live within seconds. + +### The status of a share + +A share is live once the query engine has loaded the rules that include it, which takes a few seconds. +The sharing page shows where each share stands and refreshes itself while a change is being applied. + +| Status | Meaning | +| --- | --- | +| Applying | Saved, and being made live. | +| Active | Live: the receiving project can read what it covers. | +| Revoking | Being removed. The receiving project loses access when this finishes. | +| Failed | Could not be made live. The status carries the reason; edit the share and save it to retry. | + +### Objects that no longer exist + +A share names schemas and tables, and an object it names may be dropped at the source later. +The share keeps it, and the sharing page lists it as **Not found**. +If an object with the same name is created again, the share applies to it. +To remove an object that is gone for good, click **Remove them from the share**, or edit the share and uncheck the object, which the tree shows as not found at the source. + +### Revoking a share + +Click the delete icon on a share to revoke it. +The receiving project loses access to everything the share covered as soon as the revoke is live, within seconds, and the share disappears from the page once it is. +The catalog cannot be shared with the same project again until the revoke has finished. + +### What a share does not restrict + +A share grants reading, but the query engine does not check table procedures against it. +Anyone in the receiving project can run `ALTER TABLE ... EXECUTE` on a shared table of an Iceberg or Delta Lake catalog, for example `optimize`, `expire_snapshots` or `rollback_to_snapshot`, and change the table with the catalog's own credentials. +`rollback_to_snapshot` returns the table to an earlier version and drops every write made since. +A table with a masked column is the exception: the query engine refuses table procedures on it. +Share such a catalog only with projects you trust with its tables, or set the connector's own `iceberg.security` or `delta.security` property to `read_only` on the catalog, which refuses table procedures to everyone, you included, and leaves reading unchanged. + +Deleting a catalog revokes all of its shares at once, and so does deleting the receiving project. + +### Shares your project received + +The **Catalogs** tab lists, under **Shared with this project**, every catalog share your project received, who it comes from, and what it covers. +It shows how many columns of a table are masked, but not the mask expressions, which can hold values such as a salt. +A shared catalog is queried by its own name, like any other catalog, with the SQL runner or any Trino client. + +
+ Shares received by a project +
A project catalog shared whole, and one table of another user's private catalog
+
+ +## Sharing feature groups through the query engine + +A feature group shared with another project, as described in [Sharing a feature group with selected features][sharing-a-feature-group-with-selected-features], is also queryable by that project through the query engine. +Where it is read from depends on whether the whole feature group was shared or a subset of its features. + +| Shared | Catalog | Readable | +| --- | --- | --- | +| The whole feature group, or the whole feature store | The catalog named after the feature group's format: `delta`, `hudi`, `iceberg` or `hive` | Every feature | +| A subset of the features | The shared catalog of its format: `delta_shared`, `hudi_shared` or `iceberg_shared` | The shared features only | + +The schema is the owning project's feature store, `_featurestore`, and the table is `_`. + +The share dialog says which catalog the other project will read from, or that the share is not queryable through the query engine. + +
+ Sharing a subset of a feature group +
A subset of the features is read through delta_shared
+
+ +### A subset of the features + +A subset share is read through the shared catalog of the feature group's format. +The primary key and the event time are always shared, and the receiving project reads the other shared features alongside them. + +```sql +SELECT datetime, cc_num, category, amount +FROM delta_shared.fraud_featurestore.transactions_1; +``` + +
+ Querying a subset share +
Only the shared features are listed, and the preview selects them by name
+
+ +Everything else in the table is denied: the unshared features, the connector's hidden columns such as `$path`, and the table's metadata tables such as `transactions_1$history`. +A query that selects an unshared feature, or selects `*`, is refused. + +
+ Querying an unshared feature +
Selecting a feature that was not shared is refused
+
+ +The shared catalogs are visible only to projects that received a subset share, and a project sees only the feature groups shared with it there. +A subset share is not readable through the catalog named after the format, and a whole share is not readable through the shared catalog. + +Adding features to a shared feature group does not widen the share: a new feature stays unreadable by the receiving project until you share it. +The access rules are updated before the request that adds the features returns. +A column that reaches the table some other way, such as a Delta write that merges a new column into the schema, is denied too, once Hopsworks next reads the table's columns: at least every five minutes by default, and on every share change. +If the table's columns cannot be read, the share is left out of the rules until they can. +Unsharing the feature group, or deleting it, removes access within seconds. + +### Feature groups that are not queryable + +Two kinds of feature group cannot be read through the query engine by the receiving project: + +- A feature group on an external data source, whether shared whole or in part, because its data does not live in the feature store's tables. +- A subset of a feature group stored as a plain Hive table, without Delta, Hudi or Iceberg, because there is no shared catalog for that format. + Such a feature group shared whole is read through the `hive` catalog. + +The share dialog says so for these feature groups, and the feature group remains shared and readable through the feature store APIs. diff --git a/mkdocs.yml b/mkdocs.yml index 32a93acf49..9757cd539c 100644 --- a/mkdocs.yml +++ b/mkdocs.yml @@ -228,6 +228,7 @@ nav: - Query Engine (Trino): - Query Engine: user_guides/projects/trino/query_engine.md - Trino Catalogs: user_guides/projects/trino/catalogs.md + - Sharing: user_guides/projects/trino/sharing.md - Superset: user_guides/projects/superset/superset.md - MLOps: - user_guides/mlops/index.md