Skip to content

feat: implement partition-based verification filtering (PE-8163) - #448

Closed
djwhitt wants to merge 1 commit into
developfrom
PE-8163-verification-partitioning
Closed

djwhitt wants to merge 1 commit into
developfrom
PE-8163-verification-partitioning

Conversation

@djwhitt

@djwhitt djwhitt commented Jul 17, 2025 •

Copy link
Copy Markdown
Collaborator

Summary

This PR implements partition-based verification filtering to reduce verification workload while ensuring network-wide coverage of all data.

Changes

Configuration

  • Add VERIFICATION_PARTITION_COUNT (default: 64) - number of partitions to divide ID space
  • Add VERIFICATION_PARTITION_THRESHOLD (default: 70) - priority threshold for partition filtering

Partition Assignment

  • Each node is assigned to 1 of 64 partitions at startup
  • Assignment based on:
    • AR_IO_WALLET address (if configured) - deterministic
    • Random bytes (if no wallet) - random distribution
  • Partition assignment is logged at startup

Verification Filtering

  • IDs with priority < 70 are filtered by partition
  • High-priority data (priority ≥ 70) bypasses filtering and is verified by all nodes
  • Filtering happens in the data verification worker after retrieving verifiable IDs
  • Skipped IDs have their retry count incremented to ensure eventual verification

Implementation Details

  • Created src/lib/verification-partition.ts with partition calculation functions
  • Updated data verification worker to apply partition filtering
  • No database changes required - filtering happens at application level

Benefits

  • Reduces verification workload by ~98.4% (63/64) for low-priority data
  • Ensures high-priority ArNS data is still verified by all nodes
  • Deterministic partitioning for nodes with wallets
  • Fair distribution through retry count increments

Testing

  • Verify partition assignment at startup
  • Confirm high-priority data bypasses filtering
  • Check retry count increments for skipped IDs
  • Test with different partition counts
  • Verify deterministic behavior with wallets

Related

🤖 Generated with Claude Code

- Add VERIFICATION_PARTITION_COUNT (default: 64) and VERIFICATION_PARTITION_THRESHOLD (default: 70) config
- Create verification-partition.ts with partition calculation functions
- Assign each node to a partition based on wallet address or random seed
- Filter verification workload by partition for low-priority data
- High-priority data (>= 70) bypasses partition filtering
- Skipped IDs increment retry count to ensure eventual verification
- Log partition assignment and verification statistics

This reduces verification workload by ~98.4% for low-priority data while
ensuring high-priority ArNS data is still verified by all nodes.

🤖 Generated with [Claude Code](https://claude.ai/code)

Co-Authored-By: Claude <noreply@anthropic.com>
@vilenarios

Copy link
Copy Markdown
Contributor

Thanks for this, and sorry it sat so long. Coming right after #447 lowered the minimum priority so every node verifies all ArNS data, splitting ordinary ArNS verification into partitions is a sensible way to control that cost. Peers only take data from each other when it is marked verified or trusted, so the network as a whole would still hold verified copies of everything. I'm closing it for these reasons:

  • Each node's own X-AR-IO-Verified would stay false for about 63/64 of its ordinary ArNS data, since the result is recorded only in that node's data.db.
  • Skipped items increment verification_retry_count, so after MAX_VERIFICATION_RETRIES passes they drop out of the queue permanently, and they are recorded as failures rather than as skipped. A gateway without a wallet also picks a new random partition on every restart.
  • rootTxId ?? dataId would start verifying data items whose root transaction is unknown, which develop deliberately skips.

While measuring what this would save on a production gateway, I found that most of the cost was in the SQL, not in the number of items verified. Marking a root verified (WHERE id = @id OR root_transaction_id = @id) scanned the whole contiguous_data_ids table, because root_transaction_id has no index in data.db. getVerifiableDataIds also walked every unverified row whenever fewer than 1000 qualified. I'm fixing both directly in a separate PR, which removes that cost for every gateway without giving up any trust headers.

🤖 Generated with Claude Code

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants