crates.io · Docs · API · GitHub · DeepWiki · Changelog
The agent loop for Rust. Stream from any of 7 LLM protocols, run tools, loop until done.
git clone https://github.com/yologdev/yoagent && cd yoagent
ollama serve &
ollama pull llama3.1:8b # or pass --model <any pulled model>
cargo run --example cli -- --provider ollamaThat's a working coding agent in your terminal — file read/write/edit, shell, ripgrep search, streaming output, skills. No signup, no key, nothing to configure.
yoagent cli — mini coding agent
Type /quit to exit, /clear to reset
model: llama3.1:8b
cwd: /home/user/my-project
> find all TODO comments in src/
▶ search 'TODO' ✓
Found 3 TODOs:
src/main.rs:42: // TODO: handle edge case
src/lib.rs:15: // TODO: add tests
src/utils.rs:8: // TODO: optimize this
tokens: 1250 in / 89 out
Point it at a hosted model instead by swapping the flag:
ANTHROPIC_API_KEY=sk-... cargo run --example cli
GROQ_API_KEY=... cargo run --example cli -- --provider groq --model openai/gpt-oss-120b
cargo run --example cli -- --api-url http://localhost:1234/v1 --model my-model # LM Studio, llama.cpp, vLLM[dependencies]
yoagent = "0.25"
tokio = { version = "1", features = ["full"] }Building for wasm32? Use default-features = false (see
WebAssembly & Cloudflare Workers).
On a native target keep the default native feature: without it reqwest has no TLS, and every
HTTPS provider call fails at runtime.
An agent that actually uses a tool — the thing the crate exists for:
use yoagent::provider::ModelConfig;
use yoagent::{tools, Agent, AgentEvent, StreamDelta};
#[tokio::main]
async fn main() {
// The provider is selected from the config's protocol and the key is read
// from ANTHROPIC_API_KEY. Call `.with_api_key(k)` to pass one explicitly.
let mut agent = Agent::from_config(ModelConfig::claude_sonnet_5())
.with_system_prompt("You are a coding assistant.")
.with_tools(tools::default_tools());
let mut events = agent.prompt("Find every TODO in src/ and summarise them").await;
while let Some(event) = events.recv().await {
match event {
AgentEvent::MessageUpdate { delta: StreamDelta::Text { delta }, .. } => print!("{delta}"),
AgentEvent::ToolExecutionStart { tool_name, .. } => println!("\n▶ {tool_name}"),
AgentEvent::AgentEnd { .. } => break,
_ => {}
}
}
agent.finish().await;
}Swap the model by swapping the config — the provider follows, and the key is read from that provider's conventional env var:
Agent::from_config(ModelConfig::groq("openai/gpt-oss-120b", "GPT-OSS 120B")); // GROQ_API_KEY
Agent::from_config(ModelConfig::google("gemini-3.8-flash", "Gemini 3.8 Flash")); // GEMINI_API_KEY
Agent::from_config(ModelConfig::ollama("http://localhost:11434/v1", "llama3.1:8b")); // no keyyoagent is deliberately narrow. It is the loop, tool execution, and the machinery you need to run that loop in production. It ships no vector stores, embedding pipelines, or task-graph layer — if your problem is retrieval or orchestration, one of these is the better fit:
| If you need | Look at |
|---|---|
| RAG pipelines, vector stores, embeddings, transcription and image generation | rig — "Build modular and scalable LLM Applications in Rust" |
| Typed task graphs and streaming RAG indexing alongside agents | swiftide — "Composable LLM agents and harness, typed task graphs, and streaming RAG pipelines in Rust" |
| A tool-calling loop you host, gate, steer, branch, and record | yoagent |
What that focus bought:
- The loop is a free function.
agent_loop()is stateless and takes everything it needs as arguments.Agentis an optional wrapper that adds history and queues. You can drive the loop yourself without adopting our state model. - 7 native wire protocols, not one OpenAI-compat shim with adapters bolted on. Anthropic Messages, OpenAI Completions, OpenAI Responses, Azure, Gemini, Vertex, and Bedrock each have a real implementation, so provider-specific features (thinking budgets, prompt-cache breakpoints, reasoning deltas) survive instead of being flattened away.
- One plug-in contract for the whole run. An
Extensioncan add tools, check input, allow, modify or deny each tool call, redact results, and check the final answer, with state that starts fresh each run. Install it as host policy and it governs every sub-agent too. - Steer a run that's already going. Inject guidance mid-flight; it's picked up between tool batches without restarting the turn.
- History is a tree, not a list.
Sessionforks, checkpoints, and seeks. Edit an earlier turn and re-run it without destroying the original branch. - Runs are recordable. With
features = ["gasp"], a run becomes an append-only semantic event log in a git repo — restore is clone + replay. Conformance-checked in CI. - The whole loop is testable offline.
MockProviderscripts multi-turn tool-calling conversations and honours cancellation, so abort and steering paths are testable with no network or key.
yoyo-evolve — a coding
agent that evolves its own source in public. It began as 200 lines of Rust; every commit since has
been agent-written and gated on tests. It runs on this loop with the
openapi feature enabled.
Also built on yoagent:
| Project | What it is |
|---|---|
rab |
A lightweight, extensible Rust coding agent |
greatsage |
"Rimuru's Unique Skill, you know the one" |
yoclaw |
OpenClaw reborn in Rust — a single-binary agent that remembers you |
Built something on yoagent? Open a PR and add it here — we'd like to see it.
Each line links to its chapter in the book.
- The loop — a full event stream, parallel / sequential / batched tools, steering and follow-ups, execution limits, retry with backoff and jitter, and the original hooks (
ToolMiddleware, input filters,TurnHook, lifecycle callbacks). Agent loop · Events · Retry · Callbacks & hooks - Extensions — one plug-in contract for the whole run: add tools, check input, gate and rewrite tool calls, redact results, verify the final answer, enforce a dollar
Budget, audit events, and cover sub-agents with host policy. Theyoagent-rutisbridge (not yet on crates.io) installs rutis plugins, in Rust, TypeScript or Python, as one extension. Extensions - Providers — 7 native protocols (Anthropic, OpenAI Completions and Responses, Azure, Gemini, Vertex, Bedrock) reaching 20+ providers, with thinking controls, prompt-cache hints and centralised context-overflow detection. Providers · Prompt caching
- Tools — built-in
bash, file read/write/edit,list_filesandsearch(native), custom tools via one trait, MCP over stdio or HTTP, OpenAPI specs, and per-runToolSources. Tools · MCP · OpenAPI - Sub-agents and shared state — delegate to child loops with their own model and tools; pass large artifacts by reference. Sub-agents
- Context — usage-calibrated tracking, tiered compaction, optional
LlmCompaction, loop detection. Context management - Sessions, skills, structured outputs — branching session trees with JSONL persistence, AgentSkills
SKILL.mdloading, typedprompt_structured::<T>(). Session trees · Skills · Structured outputs - Decision models (feature
decision) — typed yes/no, one-of-N and score judgments in a few hundred ms (TypeSafe Jev, Cloudflare Clef, OpenAI's Decisions API, any logprobs server); a tool gate and an input guard built on them. Decision models - Cost and telemetry — opt-in per-model pricing (
prices::enable_bundled()offline, orenable_livefrom models.dev; nothing is priced by default),SessionStatson every run including sub-agents,tracingspans with tokens and cost. Pricing · Telemetry - Recording and persistence — serde on every core type; record runs into a GASP repo (feature
gasp). Persistence · GASP - WebAssembly —
--no-default-featuresbuilds forwasm32-unknown-unknown, e.g. Cloudflare Workers. WebAssembly & Workers
Seventeen of the runnable examples in examples/ are below; ten need no API key at all (eleven counting cli with a local model). The rest are live-provider harnesses and offline evaluation sweeps.
| Example | What it shows | Key needed |
|---|---|---|
cli |
A ~400-line coding agent — all tools, skills, streaming, colored output. Like a baby Claude Code | optional¹ |
rlm |
An LLM that explores a codebase on its own by spawning sub-agents | yes |
code_review |
Three sub-agents reviewing a diff in parallel, results merged | yes |
shared_state |
Passing a large artifact between sub-agents by reference | yes |
sub_agent |
Delegation basics with a per-sub-agent model | yes |
basic |
The smallest possible agent | yes |
callbacks |
Lifecycle hooks and a custom tool | no |
persistence |
Save and restore a session | no |
telemetry |
tracing spans with token and cost fields |
no |
gasp_emit |
Recording a run into a GASP repo | no |
decision |
Decision-model questions in one line, and attaching a model to an agent (feature decision) |
yes |
extension_policy, _redact, _verifier, _budget, _tree, _audit |
Extensions: a tool policy, redaction, a verifier, budgets, policy over sub-agents, an audit log (guide) | no² |
¹ --provider ollama or --api-url needs no key; hosted providers read their conventional env var.
² Scripted offline by default; -- --live uses DEEPSEEK_API_KEY or ANTHROPIC_API_KEY.
MockProvider scripts a whole multi-turn tool-calling conversation with no network, and honours
cancellation, so abort and steering paths are testable too. See
Testing Your Agent; how the crate itself is tested and what CI runs is
in CONTRIBUTING.
- The book — concepts, guides, a page per provider, and the architecture and module map (source)
- API reference — built with all features enabled
- CHANGELOG — every release
- CONTRIBUTING — how to build, test, and send a PR
MSRV is 1.86, enforced in CI. Raising it is a minor-version change.
MIT — see LICENSE.