ARISE: A Repository-level Graph Representation and Toolset for Agentic Program Repair and Fault Localization
Paper: ARISE: A Repository-level Graph Representation and Toolset for Agentic Program Repair and Fault Localization — Shahd Seddik, Fahd Seddik, Amirrezza Esmaeili, Mahdieh Sadat, Fatemeh Fard (University of British Columbia). arXiv:2605.03117
Artifact: https://github.com/FARD-Lab/ARISE
ARISE is a framework-agnostic Python toolset that builds a multi-granularity program graph from a repository and exposes it through a three-tier tool API that any LLM-based agent can query. The central primitive is data-flow slicing — a queryable agent tool that traces, in a single call, which statements define or consume a variable of interest.
On SWE-bench Lite (300 real GitHub issues, 11 Python repositories) with Qwen2.5-Coder-32B-Instruct, mounting ARISE on SWE-agent resolves 22.0% of issues (66/300), a +4.7 pp gain over the unmodified SWE-agent baseline under the identical backbone and host. The gain is mechanistically explained by sharper localization:
| Metric | SWE-agent baseline | ARISE-Full | Δ |
|---|---|---|---|
| Pass@1 (repair) | 17.3% | 22.0% | +4.7 pp |
| Function Recall@1 (localization) | 43.0 | 60.0 | +17.0 pts |
| Line Recall@1 (localization) | 26.0 | 41.0 | +15.0 pts |
Data-flow slicing is the single largest contributor (+7.0 Function Recall@1, +2.0 pp Pass@1); controlled ablations attribute this to the data-flow graph itself, not the additional tool schema entry.
ARISE/
├── src/arise/
│ ├── graph/ Core graph data structure (nodes, edges, traversal)
│ ├── analysis/ Two-pass graph builder (structural AST + data-flow)
│ ├── index/ TF-IDF entity index for text search
│ ├── retrieval/ RetrievalSession: one-time build + disk cache
│ ├── tools/ Three-tier tool API (Tier 1–3 + explain_slice)
│ ├── eval/ FL and APR metric functions + gold label extraction
│ ├── swe_agent_bundle/ SWE-agent integration bundle — ARISE-Full
│ ├── swe_agent_bundle_tier1/ SWE-agent integration bundle — ARISE-Structural
│ ├── swe_agent_bundle_tier2/ SWE-agent integration bundle — ARISE-Slicing
│ ├── swe_agent_bundle_coarse/ SWE-agent integration bundle — ARISE-Coarse
│ └── swe_agent_bundle_explain/ SWE-agent integration bundle — ARISE-ExplainSlice
├── configs/ SWE-agent experiment configs and agent system prompts
│ ├── arise.yaml ARISE-Full tool bundle overlay
│ ├── arise_tier1.yaml ARISE-Structural (Tier 1 only)
│ ├── arise_tier2.yaml ARISE-Slicing (Tier 1+2)
│ ├── arise_coarse.yaml ARISE-Coarse (Tier 1+2 schema, no data-flow graph)
│ ├── arise_explain.yaml ARISE-ExplainSlice (Tier 1+2 + NL explanation)
│ ├── fl.yaml FL task system prompt (ARISE-aware)
│ ├── apr.yaml APR task system prompt (ARISE-aware)
│ ├── fl_baseline.yaml FL task system prompt (baseline, shell-only)
│ ├── apr_baseline.yaml APR task system prompt (baseline, shell-only)
│ └── swe_bench_lite.yaml SWE-bench Lite benchmark instance config
├── evaluation/ Evaluation scripts
│ ├── parse_preds.py Parse LOCATIONS blocks from agent trajectories
│ ├── run_eval.py Compute all FL metrics (Recall@k, MRR, F1, IoU, …)
│ └── run_apr_eval.py Compute APR Pass@1 and Spearman correlation
├── tests/ Unit test suite
└── main.py Graph inspection CLI (quick smoke test)
- Python ≥ 3.10
networkx≥ 3.2 (graph data structure)datasets≥ 2.0 — only for evaluation scripts (pip install -e ".[eval]")openai— only forexplain_slice(optional; communicates with a local vLLM endpoint)
With uv (recommended):
git clone https://github.com/FARD-Lab/ARISE.git
cd arise
uv sync --extra eval # core + evaluation scripts
uv sync --extra eval --group dev # also add mypy, pytest, ruff (for development)With pip (requires an active virtual environment):
git clone https://github.com/FARD-Lab/ARISE.git
cd arise
python3 -m venv .venv && source .venv/bin/activate
pip install -e ".[eval]" # core + evaluation scriptsAll tools are accessed through a RetrievalSession, which builds the graph and indexes once and caches the result to disk.
from arise import RetrievalSession
# Build once — subsequent calls load from cache (keyed by git commit hash)
session = RetrievalSession.build("/path/to/python-repo")from arise import search_entities, get_code_span, traverse_relations
from arise.graph import EdgeType
# TF-IDF search over all entities (modules, classes, functions, methods)
results = search_entities(session.graph, session.entity_index, "parse arguments", top_k=5)
for r in results:
print(r.name, r.file_path, r.start_line, r.score)
# Read source for a specific line range
span = get_code_span(session.root_dir, "src/mypackage/cli.py", start_line=42, end_line=60)
print(span.text)
# BFS traversal: find everything that calls a given function (by its node ID)
fragment = traverse_relations(
session.graph,
seed_id=results[0].id,
relation_types=[EdgeType.CALLED_BY],
max_hops=2,
)
for node in fragment.nodes:
print(node.type, node.name)from arise import get_dataflow_slice
# Backward slice: find all statements that define 'user_id' reaching line 87
steps = get_dataflow_slice(
session.graph,
session.stmt_index,
file_path="src/mypackage/auth.py",
line=87,
variable="user_id",
direction="backward", # "forward" or "both" also supported
)
for step in steps:
print(f" [{step.role}] {step.file_path}:{step.start_line} var={step.variable}")If the seed line has no statement node (e.g., it is a comment or decorator), the function returns an explanatory string and the agent falls back to get_code_span.
from arise import rank_suspect_regions, build_context_bundle
issue_text = "TypeError: 'NoneType' object is not subscriptable in process_result"
# Rank suspicious functions by relevance, call-graph proximity, and slice membership
suspects = rank_suspect_regions(
session.graph,
session.entity_index,
session.scope_index,
issue_text=issue_text,
top_k=10,
)
for s in suspects:
print(f" {s.score:.3f} {s.name} {s.file_path}:{s.start_line}")
# Greedily assemble relevant code spans under a token budget
bundle = build_context_bundle(
session.graph,
session.entity_index,
session.root_dir,
seed_ids=[s.node_id for s in suspects[:3]],
issue_text=issue_text,
token_budget=8000,
)
print(f"Packed {len(bundle.spans)} spans ({bundle.total_tokens} tokens)")ARISE comprises three phases, of which Phases 1 and 2 are the portable toolset — Phase 3 (the agent loop) is supplied by the host framework.
ARISE represents a Python repository as a directed, typed property graph G = (V, E) with node types {DIRECTORY, MODULE, CLASS, FUNCTION, METHOD, STATEMENT} and edge types {CONTAINS, IMPORTS, IMPORTED_BY, CALLS, CALLED_BY, INHERITS, DATAFLOW_DEF_USE, DATAFLOW_USE_DEF}. Construction proceeds in two passes:
Structural pass — parses every .py file with Python's ast module to emit directory, module, class, function, and method nodes with structural edges (CONTAINS, IMPORTS/IMPORTED_BY, CALLS/CALLED_BY, INHERITS). Call resolution is name-based and deliberately conservative: only unambiguous direct and module.function() qualified calls are emitted; dynamic dispatch is dropped to avoid misleading the agent with spurious edges.
Program graph pass — augments the structural graph with intra-procedural data-flow. For each function and method, one STATEMENT node is emitted per top-level AST statement. An intra-procedural reaching-definition scan then identifies def-use pairs and emits DATAFLOW_DEF_USE / DATAFLOW_USE_DEF edges connecting defining statements to using statements, tagged with the flowing variable names. The graph is built once and cached by commit hash.
ARISE exposes the graph through a layered tool API. The three tiers are independently ablatable — a host can mount any subset:
| Tier | Tools | Purpose |
|---|---|---|
| 1 — Structural Retrieval | search_entities, get_entity_info, traverse_relations, get_enclosing_scopes, get_code_span |
Navigate the repository structure |
| 2 — Data-Flow Slicing | get_dataflow_slice, explain_slice |
Trace def-use chains within functions |
| 3 — Context Bundling | build_context_bundle, rank_suspect_regions |
Rank suspects and pack code under a token budget |
All tools share a RetrievalSession that is built once per repository.
ARISE requires only that the host framework (a) inject the ARISE tool schema into the model context and (b) dispatch the resulting tool calls across its turns. The configs/ directory in this repository provides ready-made SWE-agent YAML overlays and system prompts for all experimental conditions. Any other tool-call-capable framework can wrap the same Python API directly.
Any framework that can make Python function calls can use ARISE by wrapping the tool functions as tool definitions and calling them via a RetrievalSession:
from arise import RetrievalSession, search_entities, get_dataflow_slice
session = RetrievalSession.build(repo_path)
def arise_search(query: str, top_k: int = 20):
return [r.to_dict() for r in search_entities(session.graph, session.entity_index, query, top_k=top_k)]
def arise_get_dataflow_slice(file_path: str, line: int, variable: str, direction: str = "backward"):
result = get_dataflow_slice(session.graph, session.stmt_index, file_path, line, variable, direction)
if isinstance(result, str):
return {"message": result}
return [step.to_dict() for step in result]The swe_agent_bundle*/ directories inside the installed package are ready-made SWE-agent tool bundles (a config.yaml with JSON-schema tool definitions and a bin/ directory with command-line wrappers). To use them with SWE-agent:
Step 1: Build the ARISE wheel (for offline install inside SWE-bench containers)
cd arise/
uv build # produces dist/arise-*.whlStep 2: Sync bundles into SWE-agent
Run this from inside your SWE-agent directory after installing ARISE:
import arise, shutil, pathlib
pkg = pathlib.Path(arise.__file__).parent
bundles = [
("swe_agent_bundle", "tools/arise"),
("swe_agent_bundle_tier1", "tools/arise-tier1"),
("swe_agent_bundle_tier2", "tools/arise-tier2"),
("swe_agent_bundle_coarse", "tools/arise-coarse"),
("swe_agent_bundle_explain", "tools/arise-explain"),
]
for src_name, dst_name in bundles:
shutil.copytree(pkg / src_name, pathlib.Path(dst_name), dirs_exist_ok=True)
# Copy wheel for offline install inside containers
import glob
whl = sorted(glob.glob("../arise/dist/arise-*.whl"))[-1]
for dst in ["tools/arise", "tools/arise-tier1", "tools/arise-tier2",
"tools/arise-coarse", "tools/arise-explain"]:
shutil.copy(whl, dst)All experiments use SWE-agent as the host framework and SWE-bench Lite as the benchmark. The backbone model is Qwen2.5-Coder-32B-Instruct (AWQ-quantized checkpoint Qwen/Qwen2.5-Coder-32B-Instruct-AWQ) served locally via vLLM.
| Condition (paper name) | Config files to stack |
|---|---|
| ARISE-Full | arise.yaml + fl.yaml or apr.yaml |
| ARISE-Structural | arise_tier1.yaml + fl.yaml or apr.yaml |
| ARISE-Slicing | arise_tier2.yaml + fl.yaml or apr.yaml |
| ARISE-Coarse | arise_coarse.yaml + fl.yaml or apr.yaml (set ARISE_COARSE_GRAPH=1) |
| ARISE-ExplainSlice | arise_explain.yaml + fl.yaml or apr.yaml |
| Baseline (shell-only) | fl_baseline.yaml or apr_baseline.yaml (no ARISE config) |
1. Install ARISE and build the wheel
git clone https://github.com/FARD-Lab/ARISE.git arise
cd arise
uv sync --extra eval
uv build # produces dist/arise-*.whl2. Install SWE-agent
git clone https://github.com/princeton-nlp/SWE-agent
cd SWE-agent
pip install -e .3. Sync ARISE tool bundles into SWE-agent (run from the SWE-agent directory)
python - <<'EOF'
import arise, shutil, pathlib, glob
pkg = pathlib.Path(arise.__file__).parent
bundles = [
("swe_agent_bundle", "tools/arise"),
("swe_agent_bundle_tier1", "tools/arise-tier1"),
("swe_agent_bundle_tier2", "tools/arise-tier2"),
("swe_agent_bundle_coarse", "tools/arise-coarse"),
("swe_agent_bundle_explain", "tools/arise-explain"),
]
for src_name, dst_name in bundles:
shutil.copytree(pkg / src_name, pathlib.Path(dst_name), dirs_exist_ok=True)
whl = sorted(glob.glob("../arise/dist/arise-*.whl"))[-1]
for dst in ["tools/arise", "tools/arise-tier1", "tools/arise-tier2",
"tools/arise-coarse", "tools/arise-explain"]:
shutil.copy(whl, dst)
print("Done.")
EOF4. Start the vLLM endpoint (requires a GPU with ≥ 48 GB VRAM for 32B model)
VLLM_USE_FLASHINFER_SAMPLER=0 \
python -m vllm.entrypoints.openai.api_server \
--model Qwen/Qwen2.5-Coder-32B-Instruct-AWQ \
--quantization awq_marlin \
--tensor-parallel-size 1 \
--port 80025. Run the FL task — ARISE-Full condition
# From inside SWE-agent/
OPENAI_API_BASE=http://localhost:8002/v1 \
OPENAI_API_KEY=dummy \
python -m sweagent.run.run_batch \
--config config/default.yaml \
--config ../arise/configs/arise.yaml \
--config ../arise/configs/fl.yaml \
--config ../arise/configs/swe_bench_lite.yaml \
--output_dir outputs/qwen25coder/condfull-fl6. Evaluate FL metrics
# From inside arise/
python evaluation/run_eval.py \
--traj-dir ../SWE-agent/outputs/qwen25coder/condfull-fl \
--output results_fl.json7. Run the APR task — ARISE-Full condition
OPENAI_API_BASE=http://localhost:8002/v1 \
OPENAI_API_KEY=dummy \
python -m sweagent.run.run_batch \
--config config/default.yaml \
--config ../arise/configs/arise.yaml \
--config ../arise/configs/apr.yaml \
--config ../arise/configs/swe_bench_lite.yaml \
--instances.evaluate=true \
--output_dir outputs/qwen25coder/condfull-apr8. Evaluate APR metrics
python evaluation/run_apr_eval.py \
--run-dir ../SWE-agent/outputs/qwen25coder/condfull-apr \
--output-pass-at-1 ../SWE-agent/outputs/qwen25coder/condfull-apr/pass_at1.json \
--output results_apr.json9. Spearman correlation (FL ↔ APR)
python evaluation/run_eval.py \
--traj-dir ../SWE-agent/outputs/qwen25coder/condfull-fl \
--pass-at-1 ../SWE-agent/outputs/qwen25coder/condfull-apr/pass_at1.json \
--output results_fl_spearman.jsonFor the ARISE-Coarse condition, set ARISE_COARSE_GRAPH=1 in the environment before running SWE-agent. All other conditions follow the same pattern — substitute the appropriate arise*.yaml overlay from configs/.
uv run pytest tests/ -q
# or
pytest tests/ -qAll metric functions have unit tests in tests/eval/. Graph construction and tool behavior are tested against the tests/fixtures/simple_repo/ fixture.
| Metric | Task | Definition |
|---|---|---|
| File Recall@k | FL | 1 if any gold file appears in top-k predicted files |
| File MRR | FL | Mean reciprocal rank of first correct file |
| Function Recall@k | FL | 1 if any gold (file, function) pair appears in top-k |
| Function F1@k | FL | Harmonic mean of precision@k and recall@k over (file, function) pairs |
| Function MRR | FL | Mean reciprocal rank of first correct (file, function) pair |
| Line Recall@k | FL | 1 if any gold (file, line) pair appears in top-k |
| Line IoU | FL | Mean intersection-over-union of predicted and gold (file, line) sets |
| Coverage@budget | FL | Fraction of gold lines in build_context_bundle output under the token budget |
| Pass@1 | APR | Fraction of instances where the first generated patch passes all tests |
| Spearman ρ | FL+APR | Spearman rank correlation between Function Recall@1 and Pass@1 |
- Intra-procedural scope only. Data-flow slicing stops at function boundaries; def-use chains that cross a call edge require multiple tool calls. This accounts for approximately 45% of localization failures in our error analysis.
- Python-specific. The graph builder uses Python's
astmodule. Supporting other languages requires a language-specific frontend. - Simplified call resolution. Call edges are resolved name-based; dynamic dispatch through attribute access or
getattris not tracked. - Token estimation.
build_context_bundleuses ~4 chars/token as a rough heuristic; actual token counts depend on the tokenizer.