From 35bf1e559e0a9a188c697d62df52463b5d33e205 Mon Sep 17 00:00:00 2001 From: Ivan Despot <66276597+g-despot@users.noreply.github.com> Date: Fri, 25 Sep 2026 09:48:40 +0200 Subject: [PATCH 1/3] docs: Enterprise Edition landing page, badge, and Shard Self-Recovery (v1.40) --- .../feature-notes/enterprise-edition.mdx | 3 + docs/deploy/configuration/env-vars/index.md | 7 + docs/deploy/configuration/replication.md | 1 + docs/deploy/configuration/self-recovery.mdx | 220 ++++++++++++++++++ docs/deploy/enterprise.mdx | 143 ++++++++++++ sidebars.js | 23 ++ src/components/EnterpriseBadge/index.jsx | 57 +++++ .../EnterpriseBadge/styles.module.scss | 84 +++++++ src/theme/DocSidebarItem/Category/index.tsx | 11 +- src/theme/DocSidebarItem/Link/index.tsx | 5 +- src/theme/MDXComponents.js | 2 + static/img/building-icon.svg | 1 + 12 files changed, 552 insertions(+), 5 deletions(-) create mode 100644 _includes/feature-notes/enterprise-edition.mdx create mode 100644 docs/deploy/configuration/self-recovery.mdx create mode 100644 docs/deploy/enterprise.mdx create mode 100644 src/components/EnterpriseBadge/index.jsx create mode 100644 src/components/EnterpriseBadge/styles.module.scss create mode 100644 static/img/building-icon.svg diff --git a/_includes/feature-notes/enterprise-edition.mdx b/_includes/feature-notes/enterprise-edition.mdx new file mode 100644 index 000000000..00f98fc08 --- /dev/null +++ b/_includes/feature-notes/enterprise-edition.mdx @@ -0,0 +1,3 @@ +:::info Enterprise Edition +This feature is part of the Weaviate Enterprise Edition and requires a license key. See [Weaviate Enterprise Edition](/deploy/enterprise) to learn how to activate it. +::: diff --git a/docs/deploy/configuration/env-vars/index.md b/docs/deploy/configuration/env-vars/index.md index eb712ce9b..d236868ac 100644 --- a/docs/deploy/configuration/env-vars/index.md +++ b/docs/deploy/configuration/env-vars/index.md @@ -72,6 +72,8 @@ import APITable from '@site/src/components/APITable'; | `HNSW_GEO_INDEX_EF` | Balance geo index search speed and recall. This value controls the search depth for geo-based queries. Default: `800`
Added in `v1.31.22` | `string - number` | `1000` | | `LAZY_LOAD_SHARD_COUNT_THRESHOLD` | Number of shards (tenants) in a collection before lazy shard loading activates. Set to `0` to force lazy loading for all collections. Default: `1000`. See [dynamic lazy shard loading](/weaviate/concepts/storage#dynamic-lazy-shard-loading).
Added in `v1.36.6` | `string - number` | `1000` | | `LAZY_LOAD_SHARD_SIZE_THRESHOLD_GB` | Total shard size (in GB) for a collection before lazy shard loading activates. Default: `100`. See [dynamic lazy shard loading](/weaviate/concepts/storage#dynamic-lazy-shard-loading).
Added in `v1.36.6` | `string - number` | `100` | +| `LICENSE_KEY` | The Weaviate license key that activates the [Enterprise Edition](/deploy/enterprise). Mutually exclusive with `LICENSE_KEY_FILE`: if you set both, Weaviate fails to start.
Added in `v1.40` {/* TODO(ivan): parser shipped inert in v1.39.6 — confirm whether to mark v1.39.6 or v1.40 (applies to LICENSE_KEY and LICENSE_KEY_FILE) */} | `string` | `` | +| `LICENSE_KEY_FILE` | Path to a file that contains the Weaviate license key that activates the [Enterprise Edition](/deploy/enterprise). Mutually exclusive with `LICENSE_KEY`. If Weaviate can't read the file, it fails to start.
Added in `v1.40` | `string - file path` | `/etc/weaviate/license.key` | | `LIMIT_RESOURCES` | If `true`, Weaviate will automatically attempt to auto-detect and limit the amount of resources (memory & threads) it uses to (0.8 * total memory) and (number of cores-1). It will override any `GOMEMLIMIT` values, however it will respect `GOMAXPROCS` values. | `boolean` | `false` | | `LOG_FORMAT` | Set the Weaviate logging format

`json` (default): Outputs log data in JSON. e.g. `{"action":"startup","level":"debug","msg":"finished initializing modules","time":"2023-04-12T05:07:43Z"}`
`text`: Outputs log data to a string. e.g. `time="2023-04-12T04:54:23Z" level=debug msg="finished initializing modules" action=startup` | `string` | | | `LOG_LEVEL` | Sets the Weaviate logging level. Default: `info`

`panic`: Panic entries only.
`fatal`: Fatal entries only.
`error`: Error entries only.
`warning`: Warning entries only.
`info`: General operational entries.
`debug`: Very verbose logging.
`trace`: Even finer-grained informational events than `debug`. | `string` | | @@ -284,8 +286,13 @@ For more information on authentication and authorization, see the [Authenticatio | `REPLICA_MOVEMENT_ENABLED` | Enable replica movement and replication operations. When enabled, the replication engine starts and REST API endpoints for replica operations become available. Default: `false`
Added in `v1.32` | `boolean` | `true` | | `REPLICA_MOVEMENT_MINIMUM_ASYNC_WAIT` | How long replica movement waits after file copy but before finalizing the move in order for in progress writes to finish. Default: `60` seconds
Added in `v1.32` | `string - number` | `90` | | `REPLICATED_INDICES_REQUEST_QUEUE_ENABLED` | **Removed in `v1.37.10`**, and in `v1.36.18` on the `v1.36` patch line. Previously enabled a request queue buffer for replicated indices in multi-node clusters, and could be modified at runtime. Default was `false`. The feature was removed; there is no replacement.
Added in `v1.30.19` | `boolean` | `true` | +| `REPLICATION_ENGINE_FILE_COPY_CHUNK_SIZE` | Chunk size in bytes for the file copies of replication operations, including [Shard Self-Recovery](/deploy/configuration/self-recovery.mdx). The source node reads this value. Default: `1048576`
Added in `v1.34.2` | `string - number` | `4194304` | +| `REPLICATION_ENGINE_FILE_COPY_WORKERS` | The number of workers that copy files for a replication operation, including [Shard Self-Recovery](/deploy/configuration/self-recovery.mdx). Default: `10`
Added in `v1.32` | `string - number` | `5` | | `REPLICATION_ENGINE_MAX_WORKERS` | The number of workers to process replica movements in parallel. Default: `10`
Added in `v1.32` | `string - number` | `5` | | `REPLICATION_MINIMUM_FACTOR` | The minimum replication factor for all collections in the cluster. | `string - number` | `3` | +| `SELF_RECOVERY_BARRIER_TIMEOUT` | How long a node that starts without Raft state waits without catch-up progress before it loads its shards. Must be a positive duration, or Weaviate fails to start. See [Shard Self-Recovery](/deploy/configuration/self-recovery.mdx). Default: `3m`
Added in `v1.40` | `string - duration` | `5m` | +| `SELF_RECOVERY_CONCURRENCY` | The maximum number of shard recoveries that run at the same time on a node. Must be between `1` and `32`, or Weaviate fails to start. See [Shard Self-Recovery](/deploy/configuration/self-recovery.mdx). Default: `10`
Added in `v1.40` | `string - number` | `4` | +| `SELF_RECOVERY_ENABLED` | Enable [Shard Self-Recovery](/deploy/configuration/self-recovery.mdx), which restores missing shard directories from healthy replicas. An [Enterprise Edition](/deploy/enterprise) feature: it also requires a valid license key and `REPLICA_MOVEMENT_ENABLED=true`. If `REPLICA_MOVEMENT_ENABLED` is not `true`, Weaviate disables self-recovery and logs a warning. Default: `false`
Added in `v1.40` | `boolean` | `true` | ```mdx-code-block diff --git a/docs/deploy/configuration/replication.md b/docs/deploy/configuration/replication.md index 2d8c81e88..6a8ee9bc5 100644 --- a/docs/deploy/configuration/replication.md +++ b/docs/deploy/configuration/replication.md @@ -143,6 +143,7 @@ Beyond setting the initial replication factor, you can actively manage the place - [Concepts: Replication Architecture](/weaviate/concepts/replication-architecture/index.md) - [Configuring Async Replication](./async-rep.md) +- [Shard Self-Recovery](./self-recovery.mdx) ## Questions and feedback diff --git a/docs/deploy/configuration/self-recovery.mdx b/docs/deploy/configuration/self-recovery.mdx new file mode 100644 index 000000000..e3574ec68 --- /dev/null +++ b/docs/deploy/configuration/self-recovery.mdx @@ -0,0 +1,220 @@ +--- +title: Shard Self-Recovery +description: "Restore missing shard data from healthy replicas when a Weaviate node starts, and monitor and control the recovery." +image: og/docs/configuration.jpg +# tags: ['configuration', 'replication', 'enterprise'] +--- + +{/* DRAFT-HOLD(ivan): PR #11768 unmerged — do not publish before it lands in stable/v1.40 */} + +import EnterpriseEdition from '/_includes/feature-notes/enterprise-edition.mdx'; + +:::info Added in `v1.40` +::: + + + +Shard Self-Recovery restores a shard whose data directory is missing on a node, for example after a disk replacement or a lost volume. Instead of starting the shard empty, the node copies the shard's files from a healthy replica on another node. Self-recovery runs automatically and needs no manual intervention in the common case. + +## Requirements + +To enable Shard Self-Recovery, all of these must be true on every node in the cluster: + +- `SELF_RECOVERY_ENABLED` is `true`. +- Weaviate runs with a valid Enterprise Edition license key. See [Activate the Enterprise Edition](/deploy/enterprise#activate-the-enterprise-edition). +- `REPLICA_MOVEMENT_ENABLED` is `true`. Self-recovery uses the replication engine, which only runs when replica movement is enabled. +- Every node runs `v1.40` or higher. See [Mixed-version clusters](#mixed-version-clusters). + +Without a valid license key, self-recovery does not activate, even when `SELF_RECOVERY_ENABLED` is `true`. + +{/* TODO(ivan): spec — confirm exact no-key failure mode with core */} + +If `SELF_RECOVERY_ENABLED` is `true` but `REPLICA_MOVEMENT_ENABLED` is not, Weaviate starts normally, logs a warning, and disables self-recovery. This is not a startup error, so check the logs if recovery doesn't happen. + +Self-recovery only helps shards that have another replica, so it applies to collections with a [replication factor](./replication.md) greater than `1`. + +## How it works + +### When recovery starts + +A node checks for missing shard directories at two points: + +- **At startup.** While the node loads its shards, each `HOT` shard whose directory is missing is recovered. +- **When a tenant is activated.** For a multi-tenant collection, a `COLD` tenant whose shard directory is missing on this node is recovered when the tenant is activated. This only happens when the collection's schema has more than one replica. + +A node that starts with no Raft state at all, for example with an empty data volume (a *wiped* node), first catches up with the cluster's schema and only then loads its shards. If the catch-up stops making progress for `SELF_RECOVERY_BARRIER_TIMEOUT` (default `3m`), the node logs a warning and loads its shards anyway. In that case some shards can start empty. + +### Choosing a source replica + +The node probes all other replicas of the shard in parallel, in random order. The first replica that reports having data for the shard becomes the source for the copy. + +The node creates an empty shard, without copying any data, only in these cases: + +- The shard has no other replica. +- No replica was unreachable, and at least one replica reported that it has no data for the shard. + +If any replica is unreachable, the node doesn't fall back to an empty shard. It retries, because the unreachable replica might hold the data. A replica that is itself recovering the same shard counts as having no data. + +### Copying the data + +The node copies the source replica's files into a staging directory named `.recovering/` next to the shard's normal location. The copy is resumable per file: a file whose CRC32 checksum already matches the source is skipped. When the copy completes, the node renames the staging directory to the shard directory in a single atomic step, and loads the shard. + +The copy uses the replication engine's settings. See [Configuration](#configuration). + +### Retries and giving up + +A recovery makes up to 10 attempts. Probe errors, unreachable replicas, and failures to register or track the copy operation all count as attempts. Between attempts, the node waits 5 seconds, doubling the wait each time up to a limit of 5 minutes. All attempts together take about 20 minutes. + +If all attempts fail, the node gives up and logs an error. **The shard stays in the `RECOVERING` state.** It doesn't become empty, and it doesn't serve data. To resolve it, do one of the following: + +- Retry the recovery with the [`restart` endpoint](#restart-a-recovery). +- Accept an empty shard with the [`accept-empty` endpoint](#accept-an-empty-shard). +- Restart the node. On the next startup, the node submits the recovery again automatically. + +## Configuration + +Set these [environment variables](/deploy/configuration/env-vars/index.md) on every node. + +| Variable | Default | Description | +| :-- | :-- | :-- | +| `SELF_RECOVERY_ENABLED` | `false` | Enable Shard Self-Recovery. Also requires a valid license key and `REPLICA_MOVEMENT_ENABLED=true`. | +| `SELF_RECOVERY_CONCURRENCY` | `10` | The maximum number of shard recoveries that run at the same time on a node. Must be between `1` and `32`. Any other value makes Weaviate fail to start. | +| `SELF_RECOVERY_BARRIER_TIMEOUT` | `3m` | How long a wiped node waits without catch-up progress before it loads its shards. A duration such as `5m`. Must be positive, or Weaviate fails to start. | +| `REPLICA_MOVEMENT_ENABLED` | `false` | Must be `true`. If it's `false`, Weaviate disables self-recovery and logs a warning. | +| `REPLICATION_ENGINE_FILE_COPY_WORKERS` | `10` | The number of workers that copy files for a replication operation, including a recovery. | +| `REPLICATION_ENGINE_MAX_WORKERS` | `10` | The number of replication operations, including recoveries, that the replication engine processes in parallel. | +| `REPLICATION_ENGINE_FILE_COPY_CHUNK_SIZE` | `1048576` | The chunk size in bytes for file copies. The source node reads this value, so set it on the nodes that serve the data. | + +## The `RECOVERING` state + +While a shard is being recovered, the [nodes endpoint](/deploy/configuration/status.md#cluster-node-data) (`GET /v1/nodes`) reports it with `vectorIndexingStatus: "RECOVERING"` and `loaded: false`. Checking the status doesn't make the node load the shard. + +Requests for a recovering shard are either served by another replica or fail with a `422` error that you can retry. + +While a registered recovery operation for the shard is in progress, the cluster excludes the recovering replica from read and write routing, so other replicas serve the requests. There are short windows where this exclusion isn't in place yet, or no longer is, for example before the recovery operation is registered. A request that reaches the recovering replica during such a window receives a `422` error that you can retry. Configure your clients to retry these requests. + +## Observability + +### Logs + +Self-recovery logs structured entries. Each entry has an `event` field with one of these values: + +| `event` | Meaning | +| :-- | :-- | +| `self_recovery.started` | A recovery started for a shard. | +| `self_recovery.peer_probe` | The result of probing a replica, for example an unreachable replica or a replica that doesn't support self-recovery. | +| `self_recovery.op_registered` | The copy operation was registered. The `op_uuid` field holds the operation ID, and `source_node` holds the source replica. | +| `self_recovery.completed` | The recovery finished. | +| `self_recovery.empty_fallback` | The node created an empty shard because no replica had data. | +| `self_recovery.restart` | An operator restarted the recovery through the debug endpoint. | +| `self_recovery.accept_empty` | An operator accepted an empty shard through the debug endpoint. | +| `self_recovery.skipped_maintenance_mode` | The node skipped a recovery because it's in maintenance mode. | + +There is no dedicated failure event. Failures, retries, and give-ups are logged at the `warning` or `error` level. To find them, filter the self-recovery entries by level. + +### Metrics + +When [Prometheus monitoring](/deploy/configuration/monitoring.md) is enabled, Weaviate exposes these metrics: + +| Metric | Labels | Type | Description | +| :-- | :-- | :-- | :-- | +| `weaviate_self_recovery_in_progress` | | Gauge | The number of recoveries in progress on this node. | +| `weaviate_self_recovery_started_total` | `source_node` | Counter | Recoveries started, by source replica. | +| `weaviate_self_recovery_completed_total` | `result` | Counter | Recoveries finished, by result. | +| `weaviate_self_recovery_duration_seconds` | `result` | Histogram | The total duration of a recovery, by result. | +| `weaviate_self_recovery_no_data_empty_total` | | Counter | Empty shards created on a node that started with its Raft state. A shard directory disappeared and no replica had data. Alert on this metric. | +| `weaviate_self_recovery_no_data_during_bootstrap_total` | | Counter | Empty shards created on a wiped node. This usually means a collection or tenant was created while the node was away, not data loss. | +| `weaviate_self_recovery_unreachable_peer_total` | `peer` | Counter | Probes that couldn't reach a replica, by replica. | +| `weaviate_self_recovery_giveup_total` | | Counter | Recoveries that used up all attempts. The shard stays `RECOVERING`. | +| `weaviate_self_recovery_accept_empty_total` | | Counter | Calls to the `accept-empty` endpoint. | + +The `result` label has one of these values: `success`, `failure`, `empty_fallback`, `cancelled`, or `skipped`. + +## Operator controls + +### Cancel a recovery operation + +A registered recovery is a replication operation. To cancel it, use the standard cancel endpoint, `POST /v1/replication/replicate/{id}/cancel`. The operation ID is the `op_uuid` field of the `self_recovery.op_registered` log entry. See [Cancel a replication operation](./replica-movement.mdx#cancel-a-replication-operation). + +### Debug endpoints + +Two further endpoints are served on the debug port (`GO_PROFILING_PORT`, default `6060`) of the node that holds the shard. They exist only when self-recovery is enabled. The debug listener also requires [`DEBUG_ENDPOINTS_ENABLED`](/deploy/configuration/env-vars/index.md#DEBUG_ENDPOINTS_ENABLED) to be `true`. Otherwise, it answers every request with a `404` status. + +:::caution The debug port is unauthenticated +The debug port serves requests without authentication. Don't expose it outside your cluster's private network. +::: + +Both endpoints take `POST` requests with the query parameters `collection` and `shard`. They return these status codes: + +| Status | Meaning | +| :-- | :-- | +| `202` | The request was accepted. | +| `400` | The `collection` or `shard` parameter is missing or invalid. | +| `404` | The shard isn't in the schema. | +| `409` | `restart` only: the shard's directory already exists on the node. | +| `503` | Self-recovery isn't available on this node. | + +#### Restart a recovery + +Start the recovery of a shard again from the beginning, for example after it gave up: + +```bash +curl -X POST "http://localhost:6060/debug/self-recovery/restart?collection=Article&shard=4DHWE6iYBU7X" +``` + +#### Accept an empty shard + +Accept that no replica has the shard's data, and create the shard empty: + +```bash +curl -X POST "http://localhost:6060/debug/self-recovery/accept-empty?collection=Article&shard=4DHWE6iYBU7X" +``` + +`accept-empty` doesn't cancel a copy operation that is already registered in the cluster. To stop one, [cancel it](#cancel-a-recovery-operation). + +## Operations blocked during replication + +While a replication operation is active, Weaviate rejects changes that would break it with a `replica movement in progress` error. This applies to every replication operation, including self-recovery and [replica movement](./replica-movement.mdx). Retry the change after the operation completes. + +These changes are blocked while a replication operation is active on the collection: + +- Structural changes to a vector index. +- Adding or removing a named vector. +- Disabling a property's `indexFilterable`, `indexSearchable`, or `indexRangeFilters` index. + +Adding a property is not blocked. + +These tenant status changes are blocked while a replication operation is active on the tenant: + +- Changing the tenant to `COLD`, including a deactivation. +- Changing the tenant to `FROZEN`. +- Changing the tenant from `FROZEN` to another status. + +## Limitations + +### Collections and tenants created while a node was down + +If a collection or tenant is created while a node is down, and the node restarts with its Raft state intact, the node creates that shard empty. If [async replication](./async-rep.md) is enabled, it fills the shard from the other replicas over time. Otherwise, the replicas stay out of sync. + +A wiped node doesn't have this gap, because it catches up with the cluster's schema before it loads its shards. + +### Mixed-version clusters + +Upgrade every node to `v1.40` or higher before you enable self-recovery. A replica on an older version doesn't support the recovery probe. The recovering node logs a warning for that replica and keeps retrying. + +### Maintenance mode + +A node in maintenance mode doesn't start new recoveries. It logs a `self_recovery.skipped_maintenance_mode` entry instead. + +## Further resources + +- [Replication](./replication.md) +- [Replica movement](./replica-movement.mdx) +- [Async replication](./async-rep.md) +- [Weaviate Enterprise Edition](/deploy/enterprise) + +## Questions and feedback + +import DocsFeedback from '/_includes/docs-feedback.mdx'; + + diff --git a/docs/deploy/enterprise.mdx b/docs/deploy/enterprise.mdx new file mode 100644 index 000000000..8f474ba0a --- /dev/null +++ b/docs/deploy/enterprise.mdx @@ -0,0 +1,143 @@ +--- +title: Weaviate Enterprise Edition +sidebar_label: Enterprise Edition +description: "How the Weaviate Community Edition and Enterprise Edition relate, how a license key activates the Enterprise Edition, and where to get help." +sidebar_position: 5 +image: og/docs/installation.jpg +# tags: ['installation', 'enterprise', 'license'] +--- + +Weaviate is open source, with an optional Enterprise Edition that adds advanced features under a commercial license. Self-hosted Weaviate is available in two editions: the **Community Edition (CE)** and the **Enterprise Edition (EE)**. Both editions ship in the same Docker image and binary. A license key selects the edition when Weaviate starts. + +## Editions + +### Community Edition + +The Community Edition is Weaviate without a license key. Its code is everything outside the [`wl/` directory](https://github.com/weaviate/weaviate/tree/main/wl) of the [`weaviate/weaviate`](https://github.com/weaviate/weaviate) repository. That code is licensed under the BSD-3-Clause License, as defined in [`LICENSE-BSD`](https://github.com/weaviate/weaviate/blob/main/LICENSE-BSD). + +### Enterprise Edition + +The Enterprise Edition is Weaviate started with a Weaviate license key. It adds the code in the [`wl/` directory](https://github.com/weaviate/weaviate/tree/main/wl) of the same repository. That code is licensed under the Weaviate License, as defined in [`wl/LICENSE-WEAVIATE`](https://github.com/weaviate/weaviate/blob/main/wl/LICENSE-WEAVIATE). + +The repository's top-level [`LICENSE`](https://github.com/weaviate/weaviate/blob/main/LICENSE) file describes which license applies to which code. + +## How it works + +- **One image, one binary.** The Community Edition and the Enterprise Edition use the same Docker image and the same binary. +- **No key means Community Edition.** Without a license key, Weaviate runs as the Community Edition. +- **A key enables the Enterprise Edition.** A valid license key enables Enterprise Edition features at runtime. +- **Upgrading needs no new download.** To move from the Community Edition to the Enterprise Edition, obtain a license key and restart Weaviate with it. You don't need a new image or a new download. + +## Activate the Enterprise Edition + +:::info Added in `v1.40` +::: + +Supply the license key to Weaviate through one of these [environment variables](/deploy/configuration/env-vars/index.md): + +| Variable | Value | +| :-- | :-- | +| `LICENSE_KEY` | The license key itself. | +| `LICENSE_KEY_FILE` | The path to a file that contains the license key. | + +Weaviate reads the key when it starts: + +- `LICENSE_KEY` and `LICENSE_KEY_FILE` are mutually exclusive. If you set both, Weaviate fails to start. +- If `LICENSE_KEY_FILE` points to a file that Weaviate can't read, Weaviate fails to start. +- If the key is empty or malformed, Weaviate starts and logs a warning. Weaviate never writes the key value to its logs. +- If neither variable is set, Weaviate starts normally as the Community Edition. + +## Get a license + +To get an Enterprise Edition license key, contact Weaviate. + +{/* TODO(ivan): license request link */} + +{/* TODO(ivan): legal page pending. License terms link target: https://weaviate.io/service/enterprise-edition */} + +## Enterprise Edition features + +This table lists Enterprise Edition features and the Weaviate version that added each one. + +| Feature | Available since | Docs | +| :-- | :-- | :-- | +| Shard Self-Recovery | `v1.40` | [Shard Self-Recovery](/deploy/configuration/self-recovery) | + +## Frequently asked questions + +{/* TODO(ivan): sync wording with the final legal licensing FAQ before publish */} + +#### Q: Is Weaviate open source? + +
+ Answer + +Weaviate is open source, with an optional Enterprise Edition available under a commercial license. The Community Edition, which is all code outside the [`wl/` directory](https://github.com/weaviate/weaviate/tree/main/wl) of the [`weaviate/weaviate`](https://github.com/weaviate/weaviate) repository, is open source under the BSD-3-Clause License. The Enterprise Edition code in the `wl/` directory is source-visible, but it's licensed under the Weaviate License and requires a commercial agreement. + +
+ +#### Q: What can I do with the Community Edition? + +
+ Answer + +Anything the BSD-3-Clause License allows. You can use, modify, and redistribute the Community Edition, including commercially and in production, at no charge. The license grants no trademark rights, so you may not call a modified version "Weaviate". + +{/* TODO(ivan): link trademark policy when it exists */} + +
+ +#### Q: The Enterprise Edition code is visible on GitHub. Can I use it? + +
+ Answer + +Not without a commercial agreement. The Enterprise Edition code lives in the same public repository for transparency and auditability, but seeing the source doesn't grant a license to use it. Removing or bypassing the license key check doesn't grant any rights either. + +
+ +#### Q: Which license applies to the Docker image? + +
+ Answer + +One image contains both editions. Without a license key, Weaviate runs as the Community Edition, and everything it does is covered by the BSD-3-Clause License. Enterprise Edition features activate only with a license key, and the enterprise agreement governs them. + +{/* TODO(ivan): legal page pending — link weaviate.io/service/enterprise-edition */} + +
+ +#### Q: Do I need a new image or download to upgrade to the Enterprise Edition? + +
+ Answer + +No. Obtain a license key and restart Weaviate with it. See [Activate the Enterprise Edition](#activate-the-enterprise-edition). + +
+ +#### Q: Do Weaviate Cloud users need a license key? + +
+ Answer + +No. License keys apply to self-hosted deployments only. + +
+ +## Support and feature requests + +- **Enterprise Edition:** Submit support requests, including license and activation issues, and feature requests through the [Weaviate support portal](https://support.weaviate.io/). +- **Community Edition:** Report bugs and request features by opening an issue in the [`weaviate/weaviate` GitHub repository](https://github.com/weaviate/weaviate/issues). + +## Further resources + +- [Environment variables](/deploy/configuration/env-vars/index.md) +- [Docker installation](/deploy/installation-guides/docker-installation.md) +- [Kubernetes installation](/deploy/installation-guides/k8s-installation.md) + +## Questions and feedback + +import DocsFeedback from '/_includes/docs-feedback.mdx'; + + diff --git a/sidebars.js b/sidebars.js index 5213f5120..c27aeea0b 100644 --- a/sidebars.js +++ b/sidebars.js @@ -1197,6 +1197,12 @@ const sidebars = { id: "deploy/configuration/replica-movement", className: "sidebar-item", }, + { + type: "doc", + id: "deploy/configuration/self-recovery", + className: "sidebar-item", + customProps: { enterpriseOnly: true }, + }, { type: "doc", id: "deploy/configuration/horizontal-scaling", @@ -1243,6 +1249,23 @@ const sidebars = { type: "html", value: "", }, + { + type: "category", + label: "Licensing", + className: "sidebar-main-category", + collapsible: false, + items: [ + { + type: "doc", + id: "deploy/enterprise", + className: "sidebar-item", + }, + ], + }, + { + type: "html", + value: "", + }, { type: "category", label: "Kubernetes", diff --git a/src/components/EnterpriseBadge/index.jsx b/src/components/EnterpriseBadge/index.jsx new file mode 100644 index 000000000..84729ebee --- /dev/null +++ b/src/components/EnterpriseBadge/index.jsx @@ -0,0 +1,57 @@ +// src/components/EnterpriseBadge/index.jsx +import React from "react"; +import Link from "@docusaurus/Link"; +import styles from "./styles.module.scss"; + +const EnterpriseBadge = ({ + text = "Enterprise Edition", + compactText = "Enterprise", + href = "/deploy/enterprise", + compact = false, + iconOnly = false, + className = "" +}) => { + // data-copy-exclude marks this badge as UI chrome so the "Copy page" markdown + // export (src/components/ContextualMenu) strips it out via [data-copy-exclude]. + const badge = ( + + + {!iconOnly && {compact ? compactText : text}} + + ); + + // The icon-only variant renders inside sidebar links, where a nested link is invalid. + if (iconOnly || !href) { + return badge; + } + + return ( + + {badge} + + ); +}; + +export default EnterpriseBadge; + +// ============================================ +// Example usage in MDX files: +/* +// Globally available through src/theme/MDXComponents.js, no import needed. + +// Simple usage - shows "Enterprise Edition" and links to /deploy/enterprise + + +// Custom text + + +// Compact variant for inline use + + +// Without a link + +*/ diff --git a/src/components/EnterpriseBadge/styles.module.scss b/src/components/EnterpriseBadge/styles.module.scss new file mode 100644 index 000000000..fac8b210f --- /dev/null +++ b/src/components/EnterpriseBadge/styles.module.scss @@ -0,0 +1,84 @@ +/* src/components/EnterpriseBadge/styles.module.scss */ + +.enterpriseBadgeLink { + display: inline-flex; + text-decoration: none; + + &:hover { + text-decoration: none; + } +} + +.enterpriseBadge { + display: inline-flex; + align-items: center; + gap: 0.35rem; + /* White badge with violet text and border */ + background: white; + color: #6d28d9; + padding: 0.25rem 0.6rem; + border-radius: 4px; + font-weight: 600; + font-size: 0.75rem; + line-height: 1.3; + letter-spacing: 0.01em; + margin-bottom: 1rem; + margin-top: 0.5rem; + border: 1px solid #6d28d9; + box-shadow: 0 1px 2px rgba(0, 0, 0, 0.05); + + .enterpriseIcon { + width: 0.8rem; + height: 0.8rem; + filter: invert(17%) sepia(86%) saturate(4165%) hue-rotate(262deg) brightness(84%) contrast(98%); + // This filter converts black SVG to approximately #6d28d9 violet + } +} + +.enterpriseBadgeLink:hover .enterpriseBadge { + box-shadow: 0 1px 3px rgba(109, 40, 217, 0.25); +} + +/* Compact variant for inline use */ +.enterpriseBadge.compact { + padding: 0.2rem 0.5rem; + font-size: 0.7rem; + margin-bottom: 0; + margin-top: 0; + display: inline-flex; + vertical-align: middle; + letter-spacing: 0.02em; + + .enterpriseIcon { + width: 0.7rem; + height: 0.7rem; + } +} + +/* Icon only variant for sidebar use */ +.enterpriseBadge.iconOnly { + padding: 0.15rem 0.35rem; + margin-bottom: 0; + margin-top: 0; + display: inline-flex; + vertical-align: middle; + gap: 0; + + .enterpriseIcon { + width: 0.75rem; + height: 0.75rem; + } +} + +/* Dark theme support */ +[data-theme="dark"] .enterpriseBadge { + background: #1a1a1a; + color: #a78bfa; + border-color: #a78bfa; + box-shadow: 0 1px 3px rgba(0, 0, 0, 0.25); + + .enterpriseIcon { + filter: invert(62%) sepia(38%) saturate(1531%) hue-rotate(209deg) brightness(101%) contrast(97%); + // This filter converts black SVG to approximately #a78bfa light violet + } +} diff --git a/src/theme/DocSidebarItem/Category/index.tsx b/src/theme/DocSidebarItem/Category/index.tsx index ad563c772..1620222d6 100644 --- a/src/theme/DocSidebarItem/Category/index.tsx +++ b/src/theme/DocSidebarItem/Category/index.tsx @@ -4,18 +4,20 @@ import type CategoryType from '@theme/DocSidebarItem/Category'; import type {WrapperProps} from '@docusaurus/types'; import CloudOnlyBadge from '@site/src/components/CloudOnlyBadge'; import AcademyBadge from '@site/src/components/AcademyBadge'; +import EnterpriseBadge from '@site/src/components/EnterpriseBadge'; type Props = WrapperProps; export default function CategoryWrapper(props: Props): ReactNode { - // Check if this sidebar item has cloudOnly or academyOnly customProps + // Check if this sidebar item has cloudOnly, academyOnly or enterpriseOnly customProps const cloudOnly = props.item?.customProps?.cloudOnly; const academyOnly = props.item?.customProps?.academyOnly; + const enterpriseOnly = props.item?.customProps?.enterpriseOnly; const wrapperRef = useRef(null); const [badgeStyle, setBadgeStyle] = useState({}); useEffect(() => { - if ((cloudOnly || academyOnly) && wrapperRef.current) { + if ((cloudOnly || academyOnly || enterpriseOnly) && wrapperRef.current) { // Find the clickable label (menu__link) inside the category const menuLink = wrapperRef.current.querySelector('.menu__link'); if (menuLink) { @@ -32,9 +34,9 @@ export default function CategoryWrapper(props: Props): ReactNode { }); } } - }, [cloudOnly, academyOnly]); + }, [cloudOnly, academyOnly, enterpriseOnly]); - if (!cloudOnly && !academyOnly) { + if (!cloudOnly && !academyOnly && !enterpriseOnly) { // If no badge, just render the original Category without wrapper return ; } @@ -46,6 +48,7 @@ export default function CategoryWrapper(props: Props): ReactNode {
{cloudOnly && } {academyOnly && } + {enterpriseOnly && }
)} diff --git a/src/theme/DocSidebarItem/Link/index.tsx b/src/theme/DocSidebarItem/Link/index.tsx index bf1811c23..e5ff78118 100644 --- a/src/theme/DocSidebarItem/Link/index.tsx +++ b/src/theme/DocSidebarItem/Link/index.tsx @@ -8,6 +8,7 @@ import type LinkType from '@theme/DocSidebarItem/Link'; import type {WrapperProps} from '@docusaurus/types'; import CloudOnlyBadge from '@site/src/components/CloudOnlyBadge'; import AcademyBadge from '@site/src/components/AcademyBadge'; +import EnterpriseBadge from '@site/src/components/EnterpriseBadge'; import styles from './styles.module.scss'; type Props = WrapperProps; @@ -37,6 +38,7 @@ export default function LinkWrapper(props: Props): ReactNode { const openInNewTab = item?.customProps?.openInNewTab; const cloudOnly = item?.customProps?.cloudOnly; const academyOnly = item?.customProps?.academyOnly; + const enterpriseOnly = item?.customProps?.enterpriseOnly; // Render a custom link that opens in a new tab if (openInNewTab && 'href' in item) { @@ -66,7 +68,7 @@ export default function LinkWrapper(props: Props): ReactNode { ); } - if (!cloudOnly && !academyOnly) { + if (!cloudOnly && !academyOnly && !enterpriseOnly) { // If no badge, just render the original Link without wrapper return ; } @@ -80,6 +82,7 @@ export default function LinkWrapper(props: Props): ReactNode {
{cloudOnly && } {academyOnly && } + {enterpriseOnly && }
); diff --git a/src/theme/MDXComponents.js b/src/theme/MDXComponents.js index dc3e7b710..94da2a209 100644 --- a/src/theme/MDXComponents.js +++ b/src/theme/MDXComponents.js @@ -4,6 +4,7 @@ import DocsImage from "../components/DocsImage"; import SkipValidationLink from "../components/SkipValidationLink"; import CloudOnlyBadge from "../components/CloudOnlyBadge"; import AcademyBadge from "../components/AcademyBadge"; +import EnterpriseBadge from "../components/EnterpriseBadge"; import PromptStarter from "../components/PromptStarter"; export default { @@ -13,5 +14,6 @@ export default { SkipValidationLink, CloudOnlyBadge, AcademyBadge, + EnterpriseBadge, PromptStarter, }; diff --git a/static/img/building-icon.svg b/static/img/building-icon.svg new file mode 100644 index 000000000..6b66c82a5 --- /dev/null +++ b/static/img/building-icon.svg @@ -0,0 +1 @@ + From 7086f746bea3001ecfccfb5d43a5995918b41da7 Mon Sep 17 00:00:00 2001 From: Ivan Despot <66276597+g-despot@users.noreply.github.com> Date: Fri, 25 Sep 2026 12:58:09 +0200 Subject: [PATCH 2/3] docs: split self-recovery into concepts and configuration, move metrics to monitoring --- docs/deploy/configuration/monitoring.md | 22 ++++ docs/deploy/configuration/self-recovery.mdx | 101 ++++-------------- .../replication-architecture/consistency.md | 50 +++++++++ 3 files changed, 91 insertions(+), 82 deletions(-) diff --git a/docs/deploy/configuration/monitoring.md b/docs/deploy/configuration/monitoring.md index 383adff38..c0f343d83 100644 --- a/docs/deploy/configuration/monitoring.md +++ b/docs/deploy/configuration/monitoring.md @@ -492,6 +492,28 @@ These metrics track the replication coordinator's read and write operations acro | `replication_coordinator_reads_duration_seconds` | Duration in seconds of read operations from replicas | None | `Histogram` | | `replication_read_repair_duration_seconds` | Duration in seconds of read repair operations | None | `Histogram` | +#### Shard self-recovery + +{/* DRAFT-HOLD(ivan): PR #11768 unmerged — do not publish before it lands in stable/v1.40 */} + +Added in `v1.40`. These metrics track [Shard Self-Recovery](/deploy/configuration/self-recovery), an [Enterprise Edition](/deploy/enterprise) feature that restores missing shard data from healthy replicas. + +| Metric | Description | Labels | Type | +| ------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------- | ------------- | ----------- | +| `weaviate_self_recovery_in_progress` | The number of recoveries in progress on this node. | None | `Gauge` | +| `weaviate_self_recovery_started_total` | Recoveries started, by source replica. | `source_node` | `Counter` | +| `weaviate_self_recovery_completed_total` | Recoveries finished, by result. | `result` | `Counter` | +| `weaviate_self_recovery_duration_seconds` | The total duration of a recovery, by result. | `result` | `Histogram` | +| `weaviate_self_recovery_no_data_empty_total` | Empty shards created on a node that started with its Raft state. A shard directory disappeared and no replica had data. Alert on this metric. | None | `Counter` | +| `weaviate_self_recovery_no_data_during_bootstrap_total` | Empty shards created on a wiped node. This usually means a collection or tenant was created while the node was away, not data loss. | None | `Counter` | +| `weaviate_self_recovery_unreachable_peer_total` | Probes that couldn't reach a replica, by replica. | `peer` | `Counter` | +| `weaviate_self_recovery_giveup_total` | Recoveries that used up all attempts. The shard stays `RECOVERING`. | None | `Counter` | +| `weaviate_self_recovery_accept_empty_total` | Calls to the `accept-empty` endpoint. | None | `Counter` | + +Label values: + +- **`result`**: `success` · `failure` · `empty_fallback` · `cancelled` · `skipped`. + ### MCP server Added in `v1.38`. These metrics track tool traffic, latency, auth failures, and the live state of the runtime write-access flag for the built-in [Weaviate MCP server](/weaviate/configuration/mcp-server.mdx). diff --git a/docs/deploy/configuration/self-recovery.mdx b/docs/deploy/configuration/self-recovery.mdx index e3574ec68..242343507 100644 --- a/docs/deploy/configuration/self-recovery.mdx +++ b/docs/deploy/configuration/self-recovery.mdx @@ -14,7 +14,9 @@ import EnterpriseEdition from '/_includes/feature-notes/enterprise-edition.mdx'; -Shard Self-Recovery restores a shard whose data directory is missing on a node, for example after a disk replacement or a lost volume. Instead of starting the shard empty, the node copies the shard's files from a healthy replica on another node. Self-recovery runs automatically and needs no manual intervention in the common case. +Shard Self-Recovery restores a shard whose data directory is missing on a node, for example after a disk replacement or a lost volume. Instead of starting the shard empty, the node copies the shard's files from a healthy replica on another node. Self-recovery runs automatically when the node starts, or when a missing tenant is activated, and needs no manual intervention in the common case. + +To learn how a recovery works, see [Concepts: Shard self-recovery](/weaviate/concepts/replication-architecture/consistency.md#shard-self-recovery). ## Requirements @@ -33,44 +35,6 @@ If `SELF_RECOVERY_ENABLED` is `true` but `REPLICA_MOVEMENT_ENABLED` is not, Weav Self-recovery only helps shards that have another replica, so it applies to collections with a [replication factor](./replication.md) greater than `1`. -## How it works - -### When recovery starts - -A node checks for missing shard directories at two points: - -- **At startup.** While the node loads its shards, each `HOT` shard whose directory is missing is recovered. -- **When a tenant is activated.** For a multi-tenant collection, a `COLD` tenant whose shard directory is missing on this node is recovered when the tenant is activated. This only happens when the collection's schema has more than one replica. - -A node that starts with no Raft state at all, for example with an empty data volume (a *wiped* node), first catches up with the cluster's schema and only then loads its shards. If the catch-up stops making progress for `SELF_RECOVERY_BARRIER_TIMEOUT` (default `3m`), the node logs a warning and loads its shards anyway. In that case some shards can start empty. - -### Choosing a source replica - -The node probes all other replicas of the shard in parallel, in random order. The first replica that reports having data for the shard becomes the source for the copy. - -The node creates an empty shard, without copying any data, only in these cases: - -- The shard has no other replica. -- No replica was unreachable, and at least one replica reported that it has no data for the shard. - -If any replica is unreachable, the node doesn't fall back to an empty shard. It retries, because the unreachable replica might hold the data. A replica that is itself recovering the same shard counts as having no data. - -### Copying the data - -The node copies the source replica's files into a staging directory named `.recovering/` next to the shard's normal location. The copy is resumable per file: a file whose CRC32 checksum already matches the source is skipped. When the copy completes, the node renames the staging directory to the shard directory in a single atomic step, and loads the shard. - -The copy uses the replication engine's settings. See [Configuration](#configuration). - -### Retries and giving up - -A recovery makes up to 10 attempts. Probe errors, unreachable replicas, and failures to register or track the copy operation all count as attempts. Between attempts, the node waits 5 seconds, doubling the wait each time up to a limit of 5 minutes. All attempts together take about 20 minutes. - -If all attempts fail, the node gives up and logs an error. **The shard stays in the `RECOVERING` state.** It doesn't become empty, and it doesn't serve data. To resolve it, do one of the following: - -- Retry the recovery with the [`restart` endpoint](#restart-a-recovery). -- Accept an empty shard with the [`accept-empty` endpoint](#accept-an-empty-shard). -- Restart the node. On the next startup, the node submits the recovery again automatically. - ## Configuration Set these [environment variables](/deploy/configuration/env-vars/index.md) on every node. @@ -79,62 +43,33 @@ Set these [environment variables](/deploy/configuration/env-vars/index.md) on ev | :-- | :-- | :-- | | `SELF_RECOVERY_ENABLED` | `false` | Enable Shard Self-Recovery. Also requires a valid license key and `REPLICA_MOVEMENT_ENABLED=true`. | | `SELF_RECOVERY_CONCURRENCY` | `10` | The maximum number of shard recoveries that run at the same time on a node. Must be between `1` and `32`. Any other value makes Weaviate fail to start. | -| `SELF_RECOVERY_BARRIER_TIMEOUT` | `3m` | How long a wiped node waits without catch-up progress before it loads its shards. A duration such as `5m`. Must be positive, or Weaviate fails to start. | +| `SELF_RECOVERY_BARRIER_TIMEOUT` | `3m` | How long a node that starts with an empty data volume waits for the cluster's schema without progress before it loads its shards. A duration such as `5m`. Must be positive, or Weaviate fails to start. | | `REPLICA_MOVEMENT_ENABLED` | `false` | Must be `true`. If it's `false`, Weaviate disables self-recovery and logs a warning. | | `REPLICATION_ENGINE_FILE_COPY_WORKERS` | `10` | The number of workers that copy files for a replication operation, including a recovery. | | `REPLICATION_ENGINE_MAX_WORKERS` | `10` | The number of replication operations, including recoveries, that the replication engine processes in parallel. | | `REPLICATION_ENGINE_FILE_COPY_CHUNK_SIZE` | `1048576` | The chunk size in bytes for file copies. The source node reads this value, so set it on the nodes that serve the data. | -## The `RECOVERING` state +## Monitor a recovery While a shard is being recovered, the [nodes endpoint](/deploy/configuration/status.md#cluster-node-data) (`GET /v1/nodes`) reports it with `vectorIndexingStatus: "RECOVERING"` and `loaded: false`. Checking the status doesn't make the node load the shard. -Requests for a recovering shard are either served by another replica or fail with a `422` error that you can retry. +Requests to a recovering shard are retried or served by another replica. A request can occasionally fail with a `422` error, so configure your clients to retry these requests. -While a registered recovery operation for the shard is in progress, the cluster excludes the recovering replica from read and write routing, so other replicas serve the requests. There are short windows where this exclusion isn't in place yet, or no longer is, for example before the recovery operation is registered. A request that reaches the recovering replica during such a window receives a `422` error that you can retry. Configure your clients to retry these requests. +A recovery is a replication operation. To see the recovery operations of a node, [list the replication operations](./replica-movement.mdx#list-replication-operations) and filter by the recovering node as the target node, for example with `GET /v1/replication/replicate/list?targetNode=`. -## Observability +For Prometheus metrics that track recoveries, including in-progress, completed, and given-up recoveries, see [Monitoring: Shard self-recovery](./monitoring.md#shard-self-recovery). -### Logs - -Self-recovery logs structured entries. Each entry has an `event` field with one of these values: +## Operator controls -| `event` | Meaning | -| :-- | :-- | -| `self_recovery.started` | A recovery started for a shard. | -| `self_recovery.peer_probe` | The result of probing a replica, for example an unreachable replica or a replica that doesn't support self-recovery. | -| `self_recovery.op_registered` | The copy operation was registered. The `op_uuid` field holds the operation ID, and `source_node` holds the source replica. | -| `self_recovery.completed` | The recovery finished. | -| `self_recovery.empty_fallback` | The node created an empty shard because no replica had data. | -| `self_recovery.restart` | An operator restarted the recovery through the debug endpoint. | -| `self_recovery.accept_empty` | An operator accepted an empty shard through the debug endpoint. | -| `self_recovery.skipped_maintenance_mode` | The node skipped a recovery because it's in maintenance mode. | - -There is no dedicated failure event. Failures, retries, and give-ups are logged at the `warning` or `error` level. To find them, filter the self-recovery entries by level. - -### Metrics - -When [Prometheus monitoring](/deploy/configuration/monitoring.md) is enabled, Weaviate exposes these metrics: - -| Metric | Labels | Type | Description | -| :-- | :-- | :-- | :-- | -| `weaviate_self_recovery_in_progress` | | Gauge | The number of recoveries in progress on this node. | -| `weaviate_self_recovery_started_total` | `source_node` | Counter | Recoveries started, by source replica. | -| `weaviate_self_recovery_completed_total` | `result` | Counter | Recoveries finished, by result. | -| `weaviate_self_recovery_duration_seconds` | `result` | Histogram | The total duration of a recovery, by result. | -| `weaviate_self_recovery_no_data_empty_total` | | Counter | Empty shards created on a node that started with its Raft state. A shard directory disappeared and no replica had data. Alert on this metric. | -| `weaviate_self_recovery_no_data_during_bootstrap_total` | | Counter | Empty shards created on a wiped node. This usually means a collection or tenant was created while the node was away, not data loss. | -| `weaviate_self_recovery_unreachable_peer_total` | `peer` | Counter | Probes that couldn't reach a replica, by replica. | -| `weaviate_self_recovery_giveup_total` | | Counter | Recoveries that used up all attempts. The shard stays `RECOVERING`. | -| `weaviate_self_recovery_accept_empty_total` | | Counter | Calls to the `accept-empty` endpoint. | - -The `result` label has one of these values: `success`, `failure`, `empty_fallback`, `cancelled`, or `skipped`. +If a recovery can't complete after about 20 minutes of retries, it gives up. **The shard stays in the `RECOVERING` state.** It doesn't become empty, and it doesn't serve data. To resolve it, do one of the following: -## Operator controls +- Retry the recovery with the [`restart` endpoint](#restart-a-recovery). +- Accept an empty shard with the [`accept-empty` endpoint](#accept-an-empty-shard). +- Restart the node. On the next startup, the node submits the recovery again automatically. ### Cancel a recovery operation -A registered recovery is a replication operation. To cancel it, use the standard cancel endpoint, `POST /v1/replication/replicate/{id}/cancel`. The operation ID is the `op_uuid` field of the `self_recovery.op_registered` log entry. See [Cancel a replication operation](./replica-movement.mdx#cancel-a-replication-operation). +To cancel a registered recovery, use the standard cancel endpoint, `POST /v1/replication/replicate/{id}/cancel`. To find the operation ID, [list the replication operations](./replica-movement.mdx#list-replication-operations) that target the recovering node. See [Cancel a replication operation](./replica-movement.mdx#cancel-a-replication-operation). ### Debug endpoints @@ -196,18 +131,20 @@ These tenant status changes are blocked while a replication operation is active If a collection or tenant is created while a node is down, and the node restarts with its Raft state intact, the node creates that shard empty. If [async replication](./async-rep.md) is enabled, it fills the shard from the other replicas over time. Otherwise, the replicas stay out of sync. -A wiped node doesn't have this gap, because it catches up with the cluster's schema before it loads its shards. +A node that starts with an empty data volume doesn't have this gap, because it catches up with the cluster's schema before it loads its shards. ### Mixed-version clusters -Upgrade every node to `v1.40` or higher before you enable self-recovery. A replica on an older version doesn't support the recovery probe. The recovering node logs a warning for that replica and keeps retrying. +Upgrade every node to `v1.40` or higher before you enable self-recovery. A replica on an older version doesn't support self-recovery, so a node that needs that replica's data keeps retrying the recovery. ### Maintenance mode -A node in maintenance mode doesn't start new recoveries. It logs a `self_recovery.skipped_maintenance_mode` entry instead. +A node in maintenance mode doesn't start new recoveries. ## Further resources +- [Concepts: Shard self-recovery](/weaviate/concepts/replication-architecture/consistency.md#shard-self-recovery) +- [Monitoring: Shard self-recovery metrics](./monitoring.md#shard-self-recovery) - [Replication](./replication.md) - [Replica movement](./replica-movement.mdx) - [Async replication](./async-rep.md) diff --git a/docs/weaviate/concepts/replication-architecture/consistency.md b/docs/weaviate/concepts/replication-architecture/consistency.md index b52042b1c..c0955578c 100644 --- a/docs/weaviate/concepts/replication-architecture/consistency.md +++ b/docs/weaviate/concepts/replication-architecture/consistency.md @@ -374,7 +374,57 @@ When a shard replica is copied, the increased replication factor may become an e ::: +## Shard self-recovery + +{/* DRAFT-HOLD(ivan): PR #11768 unmerged — do not publish before it lands in stable/v1.40 */} + +import EnterpriseEdition from '/_includes/feature-notes/enterprise-edition.mdx'; + +:::info Added in `v1.40` +::: + + + +Shard self-recovery restores a shard replica whose data directory is missing on a node, for example after a disk replacement or a lost volume. Instead of starting the shard empty, the node copies the shard's files from a healthy replica on another node. A recovery is a replication operation, so it uses the same engine as [replica movement](#replica-movement). + +This section describes how a recovery works. To enable, monitor, and control self-recovery, see [Configuration: Shard Self-Recovery](/deploy/configuration/self-recovery). + +### When recovery starts + +A node checks for missing shard directories at two points: + +- **At startup.** While the node loads its shards, each `HOT` shard whose directory is missing is recovered. +- **When a tenant is activated.** For a multi-tenant collection, a `COLD` tenant whose shard directory is missing on this node is recovered when the tenant is activated. This only happens when the collection's schema has more than one replica. + +A node that starts with no Raft state at all, for example with an empty data volume (a *wiped* node), first catches up with the cluster's schema and only then loads its shards. If the catch-up stops making progress for `SELF_RECOVERY_BARRIER_TIMEOUT` (default `3m`), the node logs a warning and loads its shards anyway. In that case some shards can start empty. + +### Choosing a source replica + +The node probes all other replicas of the shard in parallel, in random order. The first replica that reports having data for the shard becomes the source for the copy. + +The node creates an empty shard, without copying any data, only in these cases: + +- The shard has no other replica. +- No replica was unreachable, and at least one replica reported that it has no data for the shard. + +If any replica is unreachable, the node doesn't fall back to an empty shard. It retries, because the unreachable replica might hold the data. A replica that is itself recovering the same shard counts as having no data. + +### Copying the data + +The node copies the source replica's files into a staging directory named `.recovering/` next to the shard's normal location. The copy is resumable per file: a file whose CRC32 checksum already matches the source is skipped. When the copy completes, the node renames the staging directory to the shard directory in a single atomic step, and loads the shard. + +### Request routing during recovery + +While a registered recovery operation for the shard is in progress, the cluster excludes the recovering replica from read and write routing, so other replicas serve the requests. There are short windows where this exclusion isn't in place yet, or no longer is, for example before the recovery operation is registered. A request that reaches the recovering replica during such a window receives a `422` error that can be retried. + +### Retries and giving up + +A recovery makes up to 10 attempts. Probe errors, unreachable replicas, and failures to register or track the copy operation all count as attempts. Between attempts, the node waits 5 seconds, doubling the wait each time up to a limit of 5 minutes. All attempts together take about 20 minutes. + +If all attempts fail, the node gives up. The shard stays in the `RECOVERING` state: it doesn't become empty, and it doesn't serve data. On the next startup, the node submits the recovery again. An operator can also [restart the recovery or accept an empty shard](/deploy/configuration/self-recovery#operator-controls). + ## Related pages +- [Configuration: Shard Self-Recovery](/deploy/configuration/self-recovery) - [API References | GraphQL | Get | Consistency Levels](../../api/graphql/get.md#consistency-levels) - API References | REST | Objects From d324fa1985ec85743f41eaee641f6c261bb941c6 Mon Sep 17 00:00:00 2001 From: Ivan Despot <66276597+g-despot@users.noreply.github.com> Date: Fri, 25 Sep 2026 13:39:02 +0200 Subject: [PATCH 3/3] docs: correct self-recovery and license-key details against tested behavior --- docs/deploy/configuration/monitoring.md | 2 +- docs/deploy/configuration/self-recovery.mdx | 12 +++++++++--- docs/deploy/enterprise.mdx | 3 ++- .../concepts/replication-architecture/consistency.md | 4 +++- 4 files changed, 15 insertions(+), 6 deletions(-) diff --git a/docs/deploy/configuration/monitoring.md b/docs/deploy/configuration/monitoring.md index c0f343d83..91d9d429c 100644 --- a/docs/deploy/configuration/monitoring.md +++ b/docs/deploy/configuration/monitoring.md @@ -505,7 +505,7 @@ Added in `v1.40`. These metrics track [Shard Self-Recovery](/deploy/configuratio | `weaviate_self_recovery_completed_total` | Recoveries finished, by result. | `result` | `Counter` | | `weaviate_self_recovery_duration_seconds` | The total duration of a recovery, by result. | `result` | `Histogram` | | `weaviate_self_recovery_no_data_empty_total` | Empty shards created on a node that started with its Raft state. A shard directory disappeared and no replica had data. Alert on this metric. | None | `Counter` | -| `weaviate_self_recovery_no_data_during_bootstrap_total` | Empty shards created on a wiped node. This usually means a collection or tenant was created while the node was away, not data loss. | None | `Counter` | +| `weaviate_self_recovery_no_data_during_bootstrap_total` | Shards that a rejoining node recreated empty because no other replica holds their data. For a shard with a replication factor of `1`, this means the shard's data was lost with the volume. | None | `Counter` | | `weaviate_self_recovery_unreachable_peer_total` | Probes that couldn't reach a replica, by replica. | `peer` | `Counter` | | `weaviate_self_recovery_giveup_total` | Recoveries that used up all attempts. The shard stays `RECOVERING`. | None | `Counter` | | `weaviate_self_recovery_accept_empty_total` | Calls to the `accept-empty` endpoint. | None | `Counter` | diff --git a/docs/deploy/configuration/self-recovery.mdx b/docs/deploy/configuration/self-recovery.mdx index 242343507..c5d802142 100644 --- a/docs/deploy/configuration/self-recovery.mdx +++ b/docs/deploy/configuration/self-recovery.mdx @@ -53,7 +53,9 @@ Set these [environment variables](/deploy/configuration/env-vars/index.md) on ev While a shard is being recovered, the [nodes endpoint](/deploy/configuration/status.md#cluster-node-data) (`GET /v1/nodes`) reports it with `vectorIndexingStatus: "RECOVERING"` and `loaded: false`. Checking the status doesn't make the node load the shard. -Requests to a recovering shard are retried or served by another replica. A request can occasionally fail with a `422` error, so configure your clients to retry these requests. +While a shard recovers, the healthy replicas serve searches and reads, such as fetching an object by ID, without errors. Operations that must consult every replica, such as aggregations, can be delayed until the recovery finishes, and then succeed. + +{/* TODO(ivan): client-facing error contract for direct hits on a recovering shard unconfirmed — internal 503/500 observed, 422 exists only in the local-access code path; confirm with core */} A recovery is a replication operation. To see the recovery operations of a node, [list the replication operations](./replica-movement.mdx#list-replication-operations) and filter by the recovering node as the target node, for example with `GET /v1/replication/replicate/list?targetNode=`. @@ -85,9 +87,9 @@ Both endpoints take `POST` requests with the query parameters `collection` and ` | :-- | :-- | | `202` | The request was accepted. | | `400` | The `collection` or `shard` parameter is missing or invalid. | -| `404` | The shard isn't in the schema. | +| `404` | The collection or shard isn't in the schema, or self-recovery is disabled on this node, so the endpoints aren't registered. | +| `405` | The request uses a method other than `POST`. | | `409` | `restart` only: the shard's directory already exists on the node. | -| `503` | Self-recovery isn't available on this node. | #### Restart a recovery @@ -141,6 +143,10 @@ Upgrade every node to `v1.40` or higher before you enable self-recovery. A repli A node in maintenance mode doesn't start new recoveries. +### Replica nodes that are down during a recovery + +If another node that holds a replica of the shard is down while the shard recovers, the recovery operation stays in the `INTEGRATING` state, and the `weaviate_self_recovery_in_progress` metric stays at `1`, until that node returns. This happens even though the recovered shard is already loaded and serving requests. The operation completes after the node is back. + ## Further resources - [Concepts: Shard self-recovery](/weaviate/concepts/replication-architecture/consistency.md#shard-self-recovery) diff --git a/docs/deploy/enterprise.mdx b/docs/deploy/enterprise.mdx index 8f474ba0a..cda82c928 100644 --- a/docs/deploy/enterprise.mdx +++ b/docs/deploy/enterprise.mdx @@ -44,7 +44,8 @@ Weaviate reads the key when it starts: - `LICENSE_KEY` and `LICENSE_KEY_FILE` are mutually exclusive. If you set both, Weaviate fails to start. - If `LICENSE_KEY_FILE` points to a file that Weaviate can't read, Weaviate fails to start. -- If the key is empty or malformed, Weaviate starts and logs a warning. Weaviate never writes the key value to its logs. +- If the key is malformed, or if `LICENSE_KEY_FILE` points to an empty file, Weaviate starts and logs a warning. Weaviate never writes the key value to its logs. +- An empty `LICENSE_KEY` is treated as unset. - If neither variable is set, Weaviate starts normally as the Community Edition. ## Get a license diff --git a/docs/weaviate/concepts/replication-architecture/consistency.md b/docs/weaviate/concepts/replication-architecture/consistency.md index c0955578c..fdba49351 100644 --- a/docs/weaviate/concepts/replication-architecture/consistency.md +++ b/docs/weaviate/concepts/replication-architecture/consistency.md @@ -415,7 +415,9 @@ The node copies the source replica's files into a staging directory named `