Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 3 additions & 0 deletions _includes/feature-notes/enterprise-edition.mdx
Original file line number Diff line number Diff line change
@@ -0,0 +1,3 @@
:::info Enterprise Edition
This feature is part of the Weaviate Enterprise Edition and requires a license key. See [Weaviate Enterprise Edition](/deploy/enterprise) to learn how to activate it.
:::
7 changes: 7 additions & 0 deletions docs/deploy/configuration/env-vars/index.md
Original file line number Diff line number Diff line change
Expand Up @@ -72,6 +72,8 @@ import APITable from '@site/src/components/APITable';
| `HNSW_GEO_INDEX_EF` | Balance geo index search speed and recall. This value controls the search depth for geo-based queries. Default: `800`<br/>Added in `v1.31.22` | `string - number` | `1000` |
| `LAZY_LOAD_SHARD_COUNT_THRESHOLD` | Number of shards (tenants) in a collection before lazy shard loading activates. Set to `0` to force lazy loading for all collections. Default: `1000`. See [dynamic lazy shard loading](/weaviate/concepts/storage#dynamic-lazy-shard-loading). <br/>Added in `v1.36.6` | `string - number` | `1000` |
| `LAZY_LOAD_SHARD_SIZE_THRESHOLD_GB` | Total shard size (in GB) for a collection before lazy shard loading activates. Default: `100`. See [dynamic lazy shard loading](/weaviate/concepts/storage#dynamic-lazy-shard-loading). <br/>Added in `v1.36.6` | `string - number` | `100` |
| `LICENSE_KEY` | The Weaviate license key that activates the [Enterprise Edition](/deploy/enterprise). Mutually exclusive with `LICENSE_KEY_FILE`: if you set both, Weaviate fails to start.<br/>Added in `v1.40` {/* TODO(ivan): parser shipped inert in v1.39.6 — confirm whether to mark v1.39.6 or v1.40 (applies to LICENSE_KEY and LICENSE_KEY_FILE) */} | `string` | `<your-license-key>` |
| `LICENSE_KEY_FILE` | Path to a file that contains the Weaviate license key that activates the [Enterprise Edition](/deploy/enterprise). Mutually exclusive with `LICENSE_KEY`. If Weaviate can't read the file, it fails to start.<br/>Added in `v1.40` | `string - file path` | `/etc/weaviate/license.key` |
| `LIMIT_RESOURCES` | If `true`, Weaviate will automatically attempt to auto-detect and limit the amount of resources (memory & threads) it uses to (0.8 * total memory) and (number of cores-1). It will override any `GOMEMLIMIT` values, however it will respect `GOMAXPROCS` values. | `boolean` | `false` |
| `LOG_FORMAT` | Set the Weaviate logging format <br/><br/>`json` (default): Outputs log data in JSON. e.g. `{"action":"startup","level":"debug","msg":"finished initializing modules","time":"2023-04-12T05:07:43Z"}` <br/>`text`: Outputs log data to a string. e.g. `time="2023-04-12T04:54:23Z" level=debug msg="finished initializing modules" action=startup` | `string` | |
| `LOG_LEVEL` | Sets the Weaviate logging level. Default: `info`<br/><br/>`panic`: Panic entries only. <br/>`fatal`: Fatal entries only. <br/> `error`: Error entries only. <br/>`warning`: Warning entries only. <br/>`info`: General operational entries. <br/> `debug`: Very verbose logging. <br/>`trace`: Even finer-grained informational events than `debug`. | `string` | |
Expand Down Expand Up @@ -284,8 +286,13 @@ For more information on authentication and authorization, see the [Authenticatio
| `REPLICA_MOVEMENT_ENABLED` | Enable replica movement and replication operations. When enabled, the replication engine starts and REST API endpoints for replica operations become available. Default: `false` <br/>Added in `v1.32` | `boolean` | `true` |
| `REPLICA_MOVEMENT_MINIMUM_ASYNC_WAIT` | How long replica movement waits after file copy but before finalizing the move in order for in progress writes to finish. Default: `60` seconds <br/>Added in `v1.32` | `string - number` | `90` |
| `REPLICATED_INDICES_REQUEST_QUEUE_ENABLED` | **Removed in `v1.37.10`**, and in `v1.36.18` on the `v1.36` patch line. Previously enabled a request queue buffer for replicated indices in multi-node clusters, and could be modified at runtime. Default was `false`. The feature was removed; there is no replacement. <br/>Added in `v1.30.19` | `boolean` | `true` |
| `REPLICATION_ENGINE_FILE_COPY_CHUNK_SIZE` | Chunk size in bytes for the file copies of replication operations, including [Shard Self-Recovery](/deploy/configuration/self-recovery.mdx). The source node reads this value. Default: `1048576`<br/>Added in `v1.34.2` | `string - number` | `4194304` |
| `REPLICATION_ENGINE_FILE_COPY_WORKERS` | The number of workers that copy files for a replication operation, including [Shard Self-Recovery](/deploy/configuration/self-recovery.mdx). Default: `10`<br/>Added in `v1.32` | `string - number` | `5` |
| `REPLICATION_ENGINE_MAX_WORKERS` | The number of workers to process replica movements in parallel. Default: `10` <br/>Added in `v1.32` | `string - number` | `5` |
| `REPLICATION_MINIMUM_FACTOR` | The minimum replication factor for all collections in the cluster. | `string - number` | `3` |
| `SELF_RECOVERY_BARRIER_TIMEOUT` | How long a node that starts without Raft state waits without catch-up progress before it loads its shards. Must be a positive duration, or Weaviate fails to start. See [Shard Self-Recovery](/deploy/configuration/self-recovery.mdx). Default: `3m`<br/>Added in `v1.40` | `string - duration` | `5m` |
| `SELF_RECOVERY_CONCURRENCY` | The maximum number of shard recoveries that run at the same time on a node. Must be between `1` and `32`, or Weaviate fails to start. See [Shard Self-Recovery](/deploy/configuration/self-recovery.mdx). Default: `10`<br/>Added in `v1.40` | `string - number` | `4` |
| `SELF_RECOVERY_ENABLED` | Enable [Shard Self-Recovery](/deploy/configuration/self-recovery.mdx), which restores missing shard directories from healthy replicas. An [Enterprise Edition](/deploy/enterprise) feature: it also requires a valid license key and `REPLICA_MOVEMENT_ENABLED=true`. If `REPLICA_MOVEMENT_ENABLED` is not `true`, Weaviate disables self-recovery and logs a warning. Default: `false`<br/>Added in `v1.40` | `boolean` | `true` |

```mdx-code-block
</APITable>
Expand Down
22 changes: 22 additions & 0 deletions docs/deploy/configuration/monitoring.md
Original file line number Diff line number Diff line change
Expand Up @@ -492,6 +492,28 @@ These metrics track the replication coordinator's read and write operations acro
| `replication_coordinator_reads_duration_seconds` | Duration in seconds of read operations from replicas | None | `Histogram` |
| `replication_read_repair_duration_seconds` | Duration in seconds of read repair operations | None | `Histogram` |

#### Shard self-recovery

{/* DRAFT-HOLD(ivan): PR #11768 unmerged — do not publish before it lands in stable/v1.40 */}

Added in `v1.40`. These metrics track [Shard Self-Recovery](/deploy/configuration/self-recovery), an [Enterprise Edition](/deploy/enterprise) feature that restores missing shard data from healthy replicas.

| Metric | Description | Labels | Type |
| ------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------- | ------------- | ----------- |
| `weaviate_self_recovery_in_progress` | The number of recoveries in progress on this node. | None | `Gauge` |
| `weaviate_self_recovery_started_total` | Recoveries started, by source replica. | `source_node` | `Counter` |
| `weaviate_self_recovery_completed_total` | Recoveries finished, by result. | `result` | `Counter` |
| `weaviate_self_recovery_duration_seconds` | The total duration of a recovery, by result. | `result` | `Histogram` |
| `weaviate_self_recovery_no_data_empty_total` | Empty shards created on a node that started with its Raft state. A shard directory disappeared and no replica had data. Alert on this metric. | None | `Counter` |
| `weaviate_self_recovery_no_data_during_bootstrap_total` | Shards that a rejoining node recreated empty because no other replica holds their data. For a shard with a replication factor of `1`, this means the shard's data was lost with the volume. | None | `Counter` |
| `weaviate_self_recovery_unreachable_peer_total` | Probes that couldn't reach a replica, by replica. | `peer` | `Counter` |
| `weaviate_self_recovery_giveup_total` | Recoveries that used up all attempts. The shard stays `RECOVERING`. | None | `Counter` |
| `weaviate_self_recovery_accept_empty_total` | Calls to the `accept-empty` endpoint. | None | `Counter` |

Label values:

- **`result`**: `success` · `failure` · `empty_fallback` · `cancelled` · `skipped`.

### MCP server

Added in `v1.38`. These metrics track tool traffic, latency, auth failures, and the live state of the runtime write-access flag for the built-in [Weaviate MCP server](/weaviate/configuration/mcp-server.mdx).
Expand Down
1 change: 1 addition & 0 deletions docs/deploy/configuration/replication.md
Original file line number Diff line number Diff line change
Expand Up @@ -143,6 +143,7 @@ Beyond setting the initial replication factor, you can actively manage the place

- [Concepts: Replication Architecture](/weaviate/concepts/replication-architecture/index.md)
- [Configuring Async Replication](./async-rep.md)
- [Shard Self-Recovery](./self-recovery.mdx)

## Questions and feedback

Expand Down
163 changes: 163 additions & 0 deletions docs/deploy/configuration/self-recovery.mdx
Original file line number Diff line number Diff line change
@@ -0,0 +1,163 @@
---
title: Shard Self-Recovery
description: "Restore missing shard data from healthy replicas when a Weaviate node starts, and monitor and control the recovery."
image: og/docs/configuration.jpg
# tags: ['configuration', 'replication', 'enterprise']
---

{/* DRAFT-HOLD(ivan): PR #11768 unmerged — do not publish before it lands in stable/v1.40 */}

import EnterpriseEdition from '/_includes/feature-notes/enterprise-edition.mdx';

:::info Added in `v1.40`
:::

<EnterpriseEdition/>

Shard Self-Recovery restores a shard whose data directory is missing on a node, for example after a disk replacement or a lost volume. Instead of starting the shard empty, the node copies the shard's files from a healthy replica on another node. Self-recovery runs automatically when the node starts, or when a missing tenant is activated, and needs no manual intervention in the common case.

To learn how a recovery works, see [Concepts: Shard self-recovery](/weaviate/concepts/replication-architecture/consistency.md#shard-self-recovery).

## Requirements

To enable Shard Self-Recovery, all of these must be true on every node in the cluster:

- `SELF_RECOVERY_ENABLED` is `true`.
- Weaviate runs with a valid Enterprise Edition license key. See [Activate the Enterprise Edition](/deploy/enterprise#activate-the-enterprise-edition).
- `REPLICA_MOVEMENT_ENABLED` is `true`. Self-recovery uses the replication engine, which only runs when replica movement is enabled.
- Every node runs `v1.40` or higher. See [Mixed-version clusters](#mixed-version-clusters).

Without a valid license key, self-recovery does not activate, even when `SELF_RECOVERY_ENABLED` is `true`.

{/* TODO(ivan): spec — confirm exact no-key failure mode with core */}

If `SELF_RECOVERY_ENABLED` is `true` but `REPLICA_MOVEMENT_ENABLED` is not, Weaviate starts normally, logs a warning, and disables self-recovery. This is not a startup error, so check the logs if recovery doesn't happen.

Self-recovery only helps shards that have another replica, so it applies to collections with a [replication factor](./replication.md) greater than `1`.

## Configuration

Set these [environment variables](/deploy/configuration/env-vars/index.md) on every node.

| Variable | Default | Description |
| :-- | :-- | :-- |
| `SELF_RECOVERY_ENABLED` | `false` | Enable Shard Self-Recovery. Also requires a valid license key and `REPLICA_MOVEMENT_ENABLED=true`. |
| `SELF_RECOVERY_CONCURRENCY` | `10` | The maximum number of shard recoveries that run at the same time on a node. Must be between `1` and `32`. Any other value makes Weaviate fail to start. |
| `SELF_RECOVERY_BARRIER_TIMEOUT` | `3m` | How long a node that starts with an empty data volume waits for the cluster's schema without progress before it loads its shards. A duration such as `5m`. Must be positive, or Weaviate fails to start. |
| `REPLICA_MOVEMENT_ENABLED` | `false` | Must be `true`. If it's `false`, Weaviate disables self-recovery and logs a warning. |
| `REPLICATION_ENGINE_FILE_COPY_WORKERS` | `10` | The number of workers that copy files for a replication operation, including a recovery. |
| `REPLICATION_ENGINE_MAX_WORKERS` | `10` | The number of replication operations, including recoveries, that the replication engine processes in parallel. |
| `REPLICATION_ENGINE_FILE_COPY_CHUNK_SIZE` | `1048576` | The chunk size in bytes for file copies. The source node reads this value, so set it on the nodes that serve the data. |

## Monitor a recovery

While a shard is being recovered, the [nodes endpoint](/deploy/configuration/status.md#cluster-node-data) (`GET /v1/nodes`) reports it with `vectorIndexingStatus: "RECOVERING"` and `loaded: false`. Checking the status doesn't make the node load the shard.

While a shard recovers, the healthy replicas serve searches and reads, such as fetching an object by ID, without errors. Operations that must consult every replica, such as aggregations, can be delayed until the recovery finishes, and then succeed.

{/* TODO(ivan): client-facing error contract for direct hits on a recovering shard unconfirmed — internal 503/500 observed, 422 exists only in the local-access code path; confirm with core */}

A recovery is a replication operation. To see the recovery operations of a node, [list the replication operations](./replica-movement.mdx#list-replication-operations) and filter by the recovering node as the target node, for example with `GET /v1/replication/replicate/list?targetNode=<node>`.

For Prometheus metrics that track recoveries, including in-progress, completed, and given-up recoveries, see [Monitoring: Shard self-recovery](./monitoring.md#shard-self-recovery).

## Operator controls

If a recovery can't complete after about 20 minutes of retries, it gives up. **The shard stays in the `RECOVERING` state.** It doesn't become empty, and it doesn't serve data. To resolve it, do one of the following:

- Retry the recovery with the [`restart` endpoint](#restart-a-recovery).
- Accept an empty shard with the [`accept-empty` endpoint](#accept-an-empty-shard).
- Restart the node. On the next startup, the node submits the recovery again automatically.

### Cancel a recovery operation

To cancel a registered recovery, use the standard cancel endpoint, `POST /v1/replication/replicate/{id}/cancel`. To find the operation ID, [list the replication operations](./replica-movement.mdx#list-replication-operations) that target the recovering node. See [Cancel a replication operation](./replica-movement.mdx#cancel-a-replication-operation).

### Debug endpoints

Two further endpoints are served on the debug port (`GO_PROFILING_PORT`, default `6060`) of the node that holds the shard. They exist only when self-recovery is enabled. The debug listener also requires [`DEBUG_ENDPOINTS_ENABLED`](/deploy/configuration/env-vars/index.md#DEBUG_ENDPOINTS_ENABLED) to be `true`. Otherwise, it answers every request with a `404` status.

:::caution The debug port is unauthenticated
The debug port serves requests without authentication. Don't expose it outside your cluster's private network.
:::

Both endpoints take `POST` requests with the query parameters `collection` and `shard`. They return these status codes:

| Status | Meaning |
| :-- | :-- |
| `202` | The request was accepted. |
| `400` | The `collection` or `shard` parameter is missing or invalid. |
| `404` | The collection or shard isn't in the schema, or self-recovery is disabled on this node, so the endpoints aren't registered. |
| `405` | The request uses a method other than `POST`. |
| `409` | `restart` only: the shard's directory already exists on the node. |

#### Restart a recovery

Start the recovery of a shard again from the beginning, for example after it gave up:

```bash
curl -X POST "http://localhost:6060/debug/self-recovery/restart?collection=Article&shard=4DHWE6iYBU7X"
```

#### Accept an empty shard

Accept that no replica has the shard's data, and create the shard empty:

```bash
curl -X POST "http://localhost:6060/debug/self-recovery/accept-empty?collection=Article&shard=4DHWE6iYBU7X"
```

`accept-empty` doesn't cancel a copy operation that is already registered in the cluster. To stop one, [cancel it](#cancel-a-recovery-operation).

## Operations blocked during replication

While a replication operation is active, Weaviate rejects changes that would break it with a `replica movement in progress` error. This applies to every replication operation, including self-recovery and [replica movement](./replica-movement.mdx). Retry the change after the operation completes.

These changes are blocked while a replication operation is active on the collection:

- Structural changes to a vector index.
- Adding or removing a named vector.
- Disabling a property's `indexFilterable`, `indexSearchable`, or `indexRangeFilters` index.

Adding a property is not blocked.

These tenant status changes are blocked while a replication operation is active on the tenant:

- Changing the tenant to `COLD`, including a deactivation.
- Changing the tenant to `FROZEN`.
- Changing the tenant from `FROZEN` to another status.

## Limitations

### Collections and tenants created while a node was down

If a collection or tenant is created while a node is down, and the node restarts with its Raft state intact, the node creates that shard empty. If [async replication](./async-rep.md) is enabled, it fills the shard from the other replicas over time. Otherwise, the replicas stay out of sync.

A node that starts with an empty data volume doesn't have this gap, because it catches up with the cluster's schema before it loads its shards.

### Mixed-version clusters

Upgrade every node to `v1.40` or higher before you enable self-recovery. A replica on an older version doesn't support self-recovery, so a node that needs that replica's data keeps retrying the recovery.

### Maintenance mode

A node in maintenance mode doesn't start new recoveries.

### Replica nodes that are down during a recovery

If another node that holds a replica of the shard is down while the shard recovers, the recovery operation stays in the `INTEGRATING` state, and the `weaviate_self_recovery_in_progress` metric stays at `1`, until that node returns. This happens even though the recovered shard is already loaded and serving requests. The operation completes after the node is back.

## Further resources

- [Concepts: Shard self-recovery](/weaviate/concepts/replication-architecture/consistency.md#shard-self-recovery)
- [Monitoring: Shard self-recovery metrics](./monitoring.md#shard-self-recovery)
- [Replication](./replication.md)
- [Replica movement](./replica-movement.mdx)
- [Async replication](./async-rep.md)
- [Weaviate Enterprise Edition](/deploy/enterprise)

## Questions and feedback

import DocsFeedback from '/_includes/docs-feedback.mdx';

<DocsFeedback/>
Loading
Loading