Skip to content
context-dot-devPublic

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

0 watching

Forks

Latest commit

 

History

10 Commits

Folders and files

Repository files navigation

Context.dev

sgrep

Search code and documents with ripgrep and local static embeddings.

Benchmarks · Releases · Context.dev

sgrep searches fresh source with independent lexical and semantic retrieval, including passages that share none of your question's words. It can also rank the output of your own ripgrep command. Everything runs locally; a persistent index reuses chunks, lexical statistics, and embeddings after checking the current file list and metadata.

Benchmarks

Indexed discovery measured 62 ms on a 2,545-file Context snapshot; piped reranking measured 43 ms, including ripgrep. These are warm local measurements; cold costs, exact output parity, and before/after results are reported separately.

Code-search benchmark comparing function retrieval and median latency for sgrep, CK, Jevgrep, and ripgrep.

The hybrid-fusion baseline, before independent semantic discovery, improved complete-function retrieval from 507/600 to 533/600 on the full local RepoQA retrieval suite, compared with current main. Median CLI latency on the separate replay panel was 81 → 69 ms. There were 32 function-retrieval gains and six losses; see the before/after results and limitations. The discovery and stdin changes have a separate 60-query comparison.

The chart above is the earlier published v0.2.0 comparison: sgrep matched CK's 76.7% function-retrieval rate at 71 ms median latency on 60 queries. Jevgrep was best at ranking the correct file first, while CK returned complete functions more often. These source-snapshot results do not establish large-monorepo or general-document performance.

See the full results and methodology, including the exact metric, per-query scores, setup costs, and limitations. The earlier Rust rewrite reduced median latency from 397 ms to 74 ms while preserving all 100 Python-query rankings.

Install

brew install ripgrep # or use your package manager
cargo install --git https://github.com/mrmps/sgrep --locked

Building requires Rust 1.90 or later. You can also download a macOS Apple Silicon binary. The executable needs ripgrep on your PATH, but no Python runtime or API key.

If you installed the earlier Python package, remove it with uv tool uninstall sgrep or pip uninstall sgrep first.

Usage

sgrep "when does a failed scrape consume credits" ./src

rg --json -C 3 'creditCost|shouldBill' ./src \
  | sgrep "when does a failed scrape consume credits" --stdin

--stdin ranks only the supplied match/context blocks. It does not scan the repository or add discovery results. Run both commands from the same directory; an optional path limits which input files are accepted. Empty input returns [] with --json and exit code 1. Use -n to choose the number of passages and --json for structured output.

Discovery returns approximate matches even when no query words occur. A result is not proof that the code implements the requested behavior; inspect the source. Use ripgrep when you need exact-match absence checks.

You can add matches from an independently chosen ripgrep command. Run both commands from the same directory; emitted paths must fall within the sgrep search scope. Ripgrep still controls its regexes, globs, case handling and context:

rg --json -i -C 12 -e 'single.?flight|in.?flight|coalesc' ./src \
  | sgrep "where do identical concurrent requests share work" ./src --rg-json -

--rg-json FILE also accepts saved output. These context blocks join the independently retrieved lexical and semantic candidates, even if they contain none of the question's words. Duplicate locations are scored once. Supplemental matches can explicitly include ignored or hidden files; the default BM25 scan still respects ignore rules. Stale source or paths outside the requested scope produce an error. Ripgrep JSON paths are resolved from the current working directory, including when reading saved output.

Use --json --explain to inspect raw BM25 and semantic scores, their normalized values, and the final fused score. Existing JSON output is unchanged without --explain.

The default --model auto uses general-text embeddings when every eligible passage is a prose file (.md, .mdx, .txt, .rst, .adoc, .org, or .text, case-insensitive). Code, mixed candidates, and other file types use code embeddings. You can choose explicitly when a file extension does not reflect its contents:

sgrep "how do refunds work" . --model text
sgrep "retry with exponential backoff" . --model code

Each model downloads about 32–34 MB on first use and works offline afterward. Code search also downloads and caches the required Tree-sitter grammars on first use. Your source and queries stay local. Existing Hugging Face caches are reused, and HF_HUB_OFFLINE=1 prevents downloads.

Directory searches follow ripgrep's ignore rules and skip hidden and binary files. Results include relative paths, one-based line ranges, and source. Exit codes are 0 for results, 1 for no matches, and 2 for errors.

Cache controls

The search index holds source passages, lexical postings, and vectors. Each search discovers eligible files with ripgrep and hashes their content on every platform. Changes to content, paths, ignore rules, model identity, or index format invalidate it; timestamps alone are never trusted. Source is checked again before publishing or returning indexed results. Concurrent edits are not an atomic repository snapshot, so freeze inputs for reproducible evals.

# Rebuild and replace this scope's index, including offline.
sgrep "retry with exponential backoff" . --refresh-cache

# Search without reading or writing the index.
sgrep "retry with exponential backoff" . --no-cache

# Store indexes, models and parsers in one isolated directory.
sgrep "retry with exponential backoff" . --cache-dir /path/to/sgrep-cache

# Rebuild separate disposable indexes before timing each executable.
python3 tests/benchmark.py BEFORE AFTER queries.json results --refresh-cache

The index cache keeps at most 256 MiB and 128 entries, evicts the least recently used entries, and removes entries unused for 30 days on the next cache-enabled search. Abandoned sgrep temporary writes and legacy embedding caches count toward cleanup. Entries larger than the budget are computed without being saved. SGREP_CACHE_MAX_BYTES changes the byte limit; 0 disables index caching. Atomic replacement can temporarily require one extra entry's disk space. Cleanup preserves unrelated files and symlink targets. Cache I/O failures fall back to uncached search. Piped reranking bypasses the index.

Index directory precedence is SGREP_CACHE_DIR, XDG_CACHE_HOME/sgrep, then ~/.cache/sgrep. Models use HF_HUB_CACHE, HUGGINGFACE_HUB_CACHE, then HF_HOME/hub (default ~/.cache/huggingface/hub). Parsers use TREE_SITTER_LANGUAGE_PACK_CACHE_DIR or the platform cache directory, under tree-sitter-language-pack/v<version>. --cache-dir overrides these with index/, hub/, and the versioned parser directory. Index files have private permissions on Unix. Source copies persist until eviction or manual deletion.

Pinned model weights and tokenizers are checked against expected sizes and SHA-256 digests before use, repaired online, or rejected offline. HF_HUB_OFFLINE=1, true, yes, and on (case-insensitive) prevent model and parser downloads. Shared models (about 64 MB together), parser libraries, and parser bundles are outside the index budget and are not automatically deleted. Parser downloads verify archive checksums; installed libraries are not rehashed on each search.

Use --fresh-assets to download models/parsers into disposable caches and bypass the persistent index. This requires network access, preserves pinned revisions, and cleans up on normal success or error. Forced termination can leave temporary directories. Neither refresh option clears operating-system filesystem caches. Benchmarks record preparation separately from timed queries and clean up disposable caches; the output directory intentionally retains result evidence.

How it works

Ripgrep enumerates eligible files, and Rust reads and chunks them in parallel. BM25 uses statistics from the whole live corpus rather than files selected by natural-language words. Code uses a Rust port of Chonkie 1.7.0 CodeChunker with the exact boundary algorithm from Tree-sitter language-pack 1.21.0, the same pinned grammars, character tokenizer, and 2,048-character size estimate. It skips metadata that CodeChunker discards and avoids parsing files that already fit in one chunk. Supported extensions cover Python, C/C++, Go, Java, Rust, JavaScript/JSX, and TypeScript/TSX. Other files use Rust Chonkie RecursiveChunker with a 2,048-character target. Chunk boundaries expand to whole source lines, so a long line can exceed that target. Discovery takes the union of the top 200 BM25 passages and top 200 semantic passages, then adds any --rg-json blocks. Semantic retrieval considers every eligible passage independently of lexical scores. Each embedding includes the relative path and complete chunk text; tokenizer padding and truncation are disabled. --stdin skips discovery and scores only supplied blocks.

Relative score fusion scales each candidate's BM25 and semantic scores to 0–1 within that candidate pool, then combines 75% semantic and 25% BM25. Single identifiers, and code-like identifiers within longer questions (camelCase, underscores, dollar signs, or digits), also retrieve exact, case-sensitive whole-token occurrences from source and rank them before approximate matches, even outside the two retrieval shortlists. --explain reports this priority as exact_identifier; fused remains the numeric hybrid score. Piped mode applies the same priority only to supplied passages. A flat score distribution contributes zero; ties favor the higher raw BM25 score. Scores are not global confidence values, and changing the candidate pool can change normalization. Results wholly contained in an earlier result are omitted; partial overlaps retain their unique source.

The index preserves the same chunks and scores as uncached search. A warm query scores matching lexical postings and materializes selected source passages; semantic retrieval still scans every vector. Queries containing identifiers also scan cached source records for exact matches. Embedding loads decode only token rows used by the current batch. Piped reranking embeds the query and supplied passages in one batch.

The embeddings come from Minish Lab, the team behind Model2Vec: Potion Code 16M v2 for code and Potion Base 8M for general English text. Both models are MIT-licensed and pinned to specific revisions.

Alternatives

Project Approach
ripgrep Fast literal and regular-expression search when you know what to match.
CK Local semantic and hybrid code search with a persistent index.
Jevgrep Model-guided repository search using Jev through TypeSafe.

Development

cargo build --release --locked
PATH="$PWD/target/release:$PATH" python3 tests/e2e.py > e2e-results.json
PATH="$PWD/target/release:$PATH" python3 tests/codechunker.py > chunker-results.json
PATH="$PWD/target/release:$PATH" python3 tests/hybrid.py > hybrid-results.json
PATH="$PWD/target/release:$PATH" python3 tests/discovery.py > discovery-results.json
SGREP_BASELINE=/path/to/before SGREP_CANDIDATE="$PWD/target/release/sgrep" python3 tests/index.py > index-results.json
python3 tests/cache_lifecycle.py target/release/sgrep target/cache-lifecycle-results.json
python3 tests/cache.py target/release/sgrep target/asset-cache-results.json --online

The end-to-end checks use the real models and write a JSON receipt. To compare two executables on your own queries, run python3 tests/benchmark.py BEFORE AFTER queries.json results. Each query is an object with id, query, and an absolute root path.

MIT · A Context.dev project, built on Minish Lab's static embeddings.

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages