Skip to content

Latest commit

 

History

4 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

JobOps Agent Harness

JobOps architecture

AI-intern job application harness: agent loop + layered memory (skills, semantic RAG, episodic SQLite) + LLMOps eval/observe/diagnose gate with auto-retry.

① Cron fires 09:00 IST (dry-run only) → ② input guardrails strip secrets/injections → ③ working memory assembles context: skills always injected + TF-IDF RAG top-k over facts + SQLite recency episodes/chat → ④ agent loop calls apify/gmail tools until final answer → ⑤ output guardrails verify sends → ⑥ LLMOps judges every run; failures are diagnosed, hint-patched, and re-traced automatically before anything reaches an inbox.

Layout

core/agent.py — agent loop (LLM <-> tools), guardrails in/out
core/llm.py — LLM client (OpenAI-compatible; set OPENAI_API_KEY)
core/memory/ — working memory assembler
    working_memory.py — assembles system prompt + history + skills + RAG + episodic context
    rag.py — HybridRAG: Pinecone vector search first, TF-IDF fallback over data/facts.md
    pinecone_store.py — Pinecone vector store (env-driven, optional dep; offline-safe)
    episodic.py — SQLite episodic memory (events, past runs, chat history) w/ recency SQL
    skills.py — procedural memory loader from skills/*.md
tools/registry.py — apify (job search) + gmail (send via himalaya) tools
core/llmops.py — eval (LLM-judge rubric) -> observe (metrics JSONL) -> diagnose -> gate -> re-run
run.py — CLI entry: python3 run.py --prompt "..." [--dry-run] [--max-retries N] or python3 run.py --chat
chat.py — interactive chat + harness REPL (persistent history, LLMOps gate per turn, /help)
scripts/reindex_pinecone.py — upsert data/facts.md chunks into Pinecone
cron.sh / cron job — daily scheduled run (dry-run by default)
data/jobs.db — episodic store; data/metrics.jsonl — observations

Memory layers Layer Type Store Retrieval skills procedural skills/*.md always injected semantic RAG facts Pinecone vectors (fallback: TF-IDF over data/facts.md) vector top-k vs prompt episodic events/past chats SQLite SQL ORDER BY recency LIMIT n Chat + harness

python3 chat.py (or python3 run.py --chat) starts a persistent REPL over the same harness: guardrails + skills + HybridRAG + SQLite history + tools + LLMOps gate on every turn. Commands: /dry, /send, /retries N, /stats, /history [N], /reindex, /quit. Single-shot CLI (--prompt) is unchanged. Pinecone vector search

Set PINECONE_API_KEY (+ embedding key OPENAI_API_KEY) and install pip install -r requirements.txt, then python3 scripts/reindex_pinecone.py to index data/facts.md. Without keys/packages the harness logs RAG backend: tfidf (fallback: ...) and runs exactly as before — Pinecone is purely additive. See .env.example. LLMOps gate

Each run: judge scores reply on rubric (relevance, tool-use sanity, no hallucinated jobs, safety). Pass >= threshold -> deliver. Fail -> diagnose (judge returns failure reasons) -> patch hint appended to retry prompt -> re-run up to --max-retries. All runs logged to data/metrics.jsonl. Safety

--dry-run (default for cron): gmail tool returns a preview instead of sending. Real sends need explicit flag.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages