slack: add slack ingest pipeline - #1
Merged
Merged
Conversation
Make Slack conversations and shared text files searchable locally so channel context can survive API pagination, edits, and live updates without manual exports.
A message without replies already arrives complete in the channel scan, so fetching it again spent one of the 50 Slack calls per minute on nothing: on the indexed channel that is 236 calls down to 10. Carrying the root text in ThreadRef also re-runs a thread when its message is edited, which the previous latest_reply comparison missed.
Every retrieval change from here — a different embedding model, grouping short messages, a reranker — is a guess until the same questions are re-scored against it, so the recall@k/MRR scorer lands before any of them. The labels quote an internal channel, so the question set itself stays untracked and only its format and tags are committed. Search moves into its own module so the CLI and the scorer rank identically, including collapsing a source's chunks into a single result.
The 90-day window was hiding more than half the channel: the full history is 443 messages back to 2025-08, and re-scoring the same questions over it drops MRR from 0.78 to 0.74, which is the honest number. An unset or zero SLACK_INDEX_LOOKBACK_DAYS now means no cutoff, and Slack reads a missing `oldest` the same way. The scorer also learns about questions the channel cannot answer: an empty `expected` skips recall and reports the top hit's score instead, because an index that answers confidently about something never discussed is its own kind of failure.
A message that is only a date is an answer, not a document: on its own it is unretrievable, and a five-character vector sits near every query. Messages within 15 minutes of each other now form one document, while a threaded message stays its own, which takes 443 documents down to 166 and the median document from 29 to 126 characters. Scored over the same questions, MRR goes 0.74 to 0.79 and the questions whose answer needs its neighbour go 0.48 to 0.73, with nothing regressing. Rows carry the timestamps they cover so a label written against a single message still matches the conversation that swallowed it, which keeps the question set independent of how the grouping is tuned.
A searcher's wording rarely appears in a chat log: the question is asked once, in passing, and the answer is a fragment. claude-haiku-4-5 now writes a question/summary/resolution/keywords header for each conversation, which lifts MRR 0.79 to 0.81 and recall@1 0.69 to 0.73, with decisions at 0.96 and dates at 0.73. The header is prepended, not substituted as Cerebras does, because their raw transcripts stay searchable in a full-text index and ours have nowhere else to live: a summary drops the catalog numbers and strain names that lexical questions turn on, and those questions held exactly at 0.82 across the change. The cost lands on questions whose answer is two lines long, which now carry a four-line header and score 0.73 to 0.65 — the reranker is the next thing to try against that. The distiller is keyed by model id, so a different model re-distills everything while a new client for the same model does not.
The prompt and the output schema decide what every summary says, but they are module constants that the logic fingerprint cannot see, so editing a prompt would have left the channel indexed by summaries the old prompt wrote. Declaring them as deps fixes that: a one-space edit to the system prompt now changes the fingerprint and re-distills. Sampling was also unpinned, so an unchanged conversation came back with a differently worded summary on every re-run; temperature 0 rides along in the request body because parse() takes no sampling arguments. Re-distilling the same conversations with only the sampling changed moves overall MRR 0.81 to 0.80 and the split slice 0.65 to 0.75, which puts a number on how much of a per-tag reading is noise: the distillation gain claimed in the previous commit is inside that band, while recall@10 rising to 0.99 is not.
The retriever compares a query and a document through one vector each, which is why the answer kept landing in the pool but not at the top: at this point every remaining failure was an ordering failure. A cross-encoder reads the pair together and rescores the 20 sources the retriever shortlists, taking MRR 0.80 to 0.88, recall@1 0.70 to 0.79, and leaving no question whose answer is missing from the results. Dates gain the most — several meetings look alike to an embedding and not to a model reading the question beside them. It also gives the index a way to say nothing fits. Cosine put the five questions the channel cannot answer inside the range of the real ones; the cross-encoder scores them at most 0.002 against a median of 0.690, so a threshold finally means something. Reranking is query-time, so the eval can run with and without it against one index, and the search picks the GPU when there is one: 2.2s per query on mps against 5s on the CPU.
Sweeping the pool over one model load shows the size was chosen too generously: ten and twenty candidates score the same MRR, ten runs in 0.98s per question against 2.33s, and past twenty the extra candidates only give the cross-encoder more to confuse itself with — recall@3 slips from 0.96 to 0.94 at thirty and MRR to 0.87 at fifty. The sweep also settles what to do about the late-interaction embedding model, which was next on the list: a better retriever can only help by lifting the answer into the shortlist, and the current one already does that for 99% of the questions while the reranker needs only ten of them. The headroom is about one question, against the cost of multi-vector storage and MaxSim scoring. The error that is left sits between recall@1 at 0.79 and recall@3 at 0.96 — an ordering problem inside the top three, not a retrieval one.
A dotenv sits in the working tree as plaintext and is invisible to the repository, so the only record of which credentials the pipeline needs was a .env.example nobody validates — and losing the file loses the values with it. secrets.yaml is committed encrypted to the age key this machine already holds for the infrastructure repo, so the set of required credentials is reviewable and the values survive a clean checkout. Nothing reads that file at run time. The dev shell decrypts it into environment variables and a deployment will hand the same variables to the unit, so the application keeps reading only the environment. An empty variable no longer counts as a value: the cocoindex CLI auto-loads the first .env it finds upwards, and a leftover scaffold full of placeholders was blanking out credentials that were set elsewhere.
The question set is the expensive artefact in this repository — every retrieval decision so far was settled by re-scoring the same hand-labelled questions — and it lived only in an ignored directory, where it was deleted once already. It cannot be committed in the clear because the labels quote a private channel, so it is committed through sops to the same age key as the credentials, and the scorer now says which command writes the plaintext it reads.
mulatta
enabled auto-merge
September 23, 2026 07:32
|
❌ Removed from merge queue: Check failed: buildbot/nix-build |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
No description provided.