Curated datasets, MCP tool server packages, and supporting data files for use
with ThinkingBox. See the
framework's README for
installation, configuration, and the tb CLI overview.
ThinkingBox-Bench is the primary evaluation release in this repository. Version 1.0 contains 507 executable tool-agent-user tasks across retail and e-commerce, travel and hospitality, auto insurance, neobank support, and consulting IT/HR support.
Each task runs in an isolated, stateful tool environment and is evaluated with executable checks over the final backend state and side effects. Some tasks also check required properties of the final response.
Browse the 507 tasks, shared scenarios, and agent configuration in the ThinkingBox-Bench dataset on Hugging Face. The Hugging Face dataset is a viewer-friendly representation; this GitHub repository remains the executable source.
If your goal is to run the published benchmark, follow the complete ThinkingBox-Bench v1.0 installation and run instructions directly. The remaining setup and examples in this README are intended for customized ThinkingBox development, individual scenarios, and smoke testing.
The rest of this repository also contains individual datasets and development fixtures that are not part of ThinkingBox-Bench. The benchmark's canonical task set is defined by the test list linked from its release documentation.
dataset/— scenarios, agents, and test cases.servers/— MCP tool server packages (thinkingbox_tools,tb_business_ops_servers_202606) and the masterservers.yamlconsumed bytb mcp-start.support/— large data files used by some tools (embeddings, knowledge bases). SetTHINKINGBOX_DATA=<path-to-this-repo>so tools can locate them.releases/— supported benchmark releases and their canonical test lists.
The commands below assume both repos are cloned side-by-side and you are
running them from the thinkingbox/ directory, so that uv run picks up
the framework's project and ../thinkingbox-data resolves to this repo:
parent/
├── thinkingbox/ # framework (CLI, Session Proxy, agent loop) ← cwd
└── thinkingbox-data/ # this repo (datasets, server packages, support files)
Install the framework first (see the thinkingbox
README). For the smoke tests
below, install the thinkingbox_tools package into the same environment:
uv pip install --config-settings editable-mode=compat -e ../thinkingbox-data/servers/thinkingbox_toolsBefore running larger scenarios, sanity-check that the framework, the
thinkingbox_tools server package, and your LLM config are wired up
correctly. The four test cases below need only the in-process MCP servers
bundled here — no Typesense, no embeddings, no support/ data files.
In one terminal (from thinkingbox/), start the Session Proxy:
THINKINGBOX_DATA="../thinkingbox-data" \
uv run tb mcp-start --servers ../thinkingbox-data/servers/servers.yamlIn another terminal (also from thinkingbox/), try running a single test
case — output is a single YAML file:
# Banking: agent looks up an account balance
uv run tb infer -c config/config_o4mini.yaml \
--dataset ../thinkingbox-data/dataset --agent think \
--name banking.py:test_get_balance_savings \
--output output.yaml
# MCS defaults: agent answers a store-info question
uv run tb infer -c config/config_o4mini.yaml \
--dataset ../thinkingbox-data/dataset --agent think \
--name mcs_defaults.py:test_mcs_defaults_easy \
--output output.yamlPretty-print the conversation and verify the assertions passed:
uv run tb pp output.yamlOr run a whole test file with multiple repetitions — output is a JSONL with one row per (test, repetition):
# Banking + email: full file (1 test) × 5 repetitions
uv run tb infer -c config/config_o4mini.yaml \
--dataset ../thinkingbox-data/dataset --agent think \
--inputs ../thinkingbox-data/dataset/test_case/banking_email.py \
--repeat 5 --batch-size 5 --output output.jsonl
# Email org: full file (3 tests) × 5 repetitions
uv run tb infer -c config/config_o4mini.yaml \
--dataset ../thinkingbox-data/dataset --agent think \
--inputs ../thinkingbox-data/dataset/test_case/email_system_org.py \
--repeat 5 --batch-size 5 --output output.jsonlAggregate the JSONL into a summary table (pass-rate per test, etc.):
uv run tb agg output.jsonlIf tb pp shows a successful conversation or tb agg reports passing
assertions, the framework, server packages, and LLM endpoint are all wired
up.
Use the canonical ThinkingBox-Bench v1.0 installation and run instructions.
After decoding once, re-run just the test assertions (no LLM calls):
# write to a new file
uv run tb run-test -c config/config_o4mini.yaml \
--dataset ../thinkingbox-data/dataset \
--resultfile output.yaml \
--name banking.py:test_transfer_and_balance \
--output test_result.yaml
# OR update output.yaml in place
uv run tb run-test -c config/config_o4mini.yaml \
--dataset ../thinkingbox-data/dataset \
--resultfile output.yaml --updateChat with a scenario:
uv run tb tui -c config/config_o4mini.yaml \
--dataset ../thinkingbox-data/dataset --agent think \
--scenario retail_banking --query "What's my checking account balance?"Chat with a specific test case (loads the test's user_context so the
simulated user can answer follow-ups):
uv run tb tui -c config/config_o4mini.yaml \
--dataset ../thinkingbox-data/dataset --agent think \
--name banking.py:test_transfer_and_balance --query ""For interactive-mode hotkeys (ESC+ENTER to submit) and slash commands, see Interactive TUI in the framework README.
This repository does not vendor third-party source code. The MCP server packages under servers/ declare their dependencies in their respective pyproject.toml files and install them from public package indexes (PyPI). Each dependency retains its own license.
This project may contain trademarks or logos for projects, products, or services. Authorized use of Microsoft trademarks or logos is subject to and must follow Microsoft's Trademark & Brand Guidelines. Use of Microsoft trademarks or logos in modified versions of this project must not cause confusion or imply Microsoft sponsorship. Any use of third-party trademarks or logos are subject to those third-party's policies.