Skip to content

Repository files navigation

ARISE: A Repository-level Graph Representation and Toolset for Agentic Program Repair and Fault Localization

Paper: ARISE: A Repository-level Graph Representation and Toolset for Agentic Program Repair and Fault Localization — Shahd Seddik, Fahd Seddik, Amirrezza Esmaeili, Mahdieh Sadat, Fatemeh Fard (University of British Columbia). arXiv:2605.03117
Artifact: https://github.com/FARD-Lab/ARISE


ARISE is a framework-agnostic Python toolset that builds a multi-granularity program graph from a repository and exposes it through a three-tier tool API that any LLM-based agent can query. The central primitive is data-flow slicing — a queryable agent tool that traces, in a single call, which statements define or consume a variable of interest.

On SWE-bench Lite (300 real GitHub issues, 11 Python repositories) with Qwen2.5-Coder-32B-Instruct, mounting ARISE on SWE-agent resolves 22.0% of issues (66/300), a +4.7 pp gain over the unmodified SWE-agent baseline under the identical backbone and host. The gain is mechanistically explained by sharper localization:

Metric SWE-agent baseline ARISE-Full Δ
Pass@1 (repair) 17.3% 22.0% +4.7 pp
Function Recall@1 (localization) 43.0 60.0 +17.0 pts
Line Recall@1 (localization) 26.0 41.0 +15.0 pts

Data-flow slicing is the single largest contributor (+7.0 Function Recall@1, +2.0 pp Pass@1); controlled ablations attribute this to the data-flow graph itself, not the additional tool schema entry.


Contents

ARISE/
├── src/arise/
│   ├── graph/              Core graph data structure (nodes, edges, traversal)
│   ├── analysis/           Two-pass graph builder (structural AST + data-flow)
│   ├── index/              TF-IDF entity index for text search
│   ├── retrieval/          RetrievalSession: one-time build + disk cache
│   ├── tools/              Three-tier tool API (Tier 1–3 + explain_slice)
│   ├── eval/               FL and APR metric functions + gold label extraction
│   ├── swe_agent_bundle/        SWE-agent integration bundle — ARISE-Full
│   ├── swe_agent_bundle_tier1/  SWE-agent integration bundle — ARISE-Structural
│   ├── swe_agent_bundle_tier2/  SWE-agent integration bundle — ARISE-Slicing
│   ├── swe_agent_bundle_coarse/ SWE-agent integration bundle — ARISE-Coarse
│   └── swe_agent_bundle_explain/ SWE-agent integration bundle — ARISE-ExplainSlice
├── configs/                SWE-agent experiment configs and agent system prompts
│   ├── arise.yaml               ARISE-Full tool bundle overlay
│   ├── arise_tier1.yaml         ARISE-Structural (Tier 1 only)
│   ├── arise_tier2.yaml         ARISE-Slicing (Tier 1+2)
│   ├── arise_coarse.yaml        ARISE-Coarse (Tier 1+2 schema, no data-flow graph)
│   ├── arise_explain.yaml       ARISE-ExplainSlice (Tier 1+2 + NL explanation)
│   ├── fl.yaml                  FL task system prompt (ARISE-aware)
│   ├── apr.yaml                 APR task system prompt (ARISE-aware)
│   ├── fl_baseline.yaml         FL task system prompt (baseline, shell-only)
│   ├── apr_baseline.yaml        APR task system prompt (baseline, shell-only)
│   └── swe_bench_lite.yaml      SWE-bench Lite benchmark instance config
├── evaluation/             Evaluation scripts
│   ├── parse_preds.py           Parse LOCATIONS blocks from agent trajectories
│   ├── run_eval.py              Compute all FL metrics (Recall@k, MRR, F1, IoU, …)
│   └── run_apr_eval.py          Compute APR Pass@1 and Spearman correlation
├── tests/                  Unit test suite
└── main.py                 Graph inspection CLI (quick smoke test)

Requirements

  • Python ≥ 3.10
  • networkx ≥ 3.2 (graph data structure)
  • datasets ≥ 2.0 — only for evaluation scripts (pip install -e ".[eval]")
  • openai — only for explain_slice (optional; communicates with a local vLLM endpoint)

Installation

With uv (recommended):

git clone https://github.com/FARD-Lab/ARISE.git
cd arise
uv sync --extra eval               # core + evaluation scripts
uv sync --extra eval --group dev   # also add mypy, pytest, ruff (for development)

With pip (requires an active virtual environment):

git clone https://github.com/FARD-Lab/ARISE.git
cd arise
python3 -m venv .venv && source .venv/bin/activate
pip install -e ".[eval]"           # core + evaluation scripts

Quick Start

All tools are accessed through a RetrievalSession, which builds the graph and indexes once and caches the result to disk.

from arise import RetrievalSession

# Build once — subsequent calls load from cache (keyed by git commit hash)
session = RetrievalSession.build("/path/to/python-repo")

Tier 1 — Structural Retrieval

from arise import search_entities, get_code_span, traverse_relations
from arise.graph import EdgeType

# TF-IDF search over all entities (modules, classes, functions, methods)
results = search_entities(session.graph, session.entity_index, "parse arguments", top_k=5)
for r in results:
    print(r.name, r.file_path, r.start_line, r.score)

# Read source for a specific line range
span = get_code_span(session.root_dir, "src/mypackage/cli.py", start_line=42, end_line=60)
print(span.text)

# BFS traversal: find everything that calls a given function (by its node ID)
fragment = traverse_relations(
    session.graph,
    seed_id=results[0].id,
    relation_types=[EdgeType.CALLED_BY],
    max_hops=2,
)
for node in fragment.nodes:
    print(node.type, node.name)

Tier 2 — Data-Flow Slicing

from arise import get_dataflow_slice

# Backward slice: find all statements that define 'user_id' reaching line 87
steps = get_dataflow_slice(
    session.graph,
    session.stmt_index,
    file_path="src/mypackage/auth.py",
    line=87,
    variable="user_id",
    direction="backward",   # "forward" or "both" also supported
)
for step in steps:
    print(f"  [{step.role}] {step.file_path}:{step.start_line}  var={step.variable}")

If the seed line has no statement node (e.g., it is a comment or decorator), the function returns an explanatory string and the agent falls back to get_code_span.

Tier 3 — Ranking and Context Assembly

from arise import rank_suspect_regions, build_context_bundle

issue_text = "TypeError: 'NoneType' object is not subscriptable in process_result"

# Rank suspicious functions by relevance, call-graph proximity, and slice membership
suspects = rank_suspect_regions(
    session.graph,
    session.entity_index,
    session.scope_index,
    issue_text=issue_text,
    top_k=10,
)
for s in suspects:
    print(f"  {s.score:.3f}  {s.name}  {s.file_path}:{s.start_line}")

# Greedily assemble relevant code spans under a token budget
bundle = build_context_bundle(
    session.graph,
    session.entity_index,
    session.root_dir,
    seed_ids=[s.node_id for s in suspects[:3]],
    issue_text=issue_text,
    token_budget=8000,
)
print(f"Packed {len(bundle.spans)} spans ({bundle.total_tokens} tokens)")

Architecture

ARISE comprises three phases, of which Phases 1 and 2 are the portable toolset — Phase 3 (the agent loop) is supplied by the host framework.

Phase 1 — Multi-Granularity Program Graph

ARISE represents a Python repository as a directed, typed property graph G = (V, E) with node types {DIRECTORY, MODULE, CLASS, FUNCTION, METHOD, STATEMENT} and edge types {CONTAINS, IMPORTS, IMPORTED_BY, CALLS, CALLED_BY, INHERITS, DATAFLOW_DEF_USE, DATAFLOW_USE_DEF}. Construction proceeds in two passes:

Structural pass — parses every .py file with Python's ast module to emit directory, module, class, function, and method nodes with structural edges (CONTAINS, IMPORTS/IMPORTED_BY, CALLS/CALLED_BY, INHERITS). Call resolution is name-based and deliberately conservative: only unambiguous direct and module.function() qualified calls are emitted; dynamic dispatch is dropped to avoid misleading the agent with spurious edges.

Program graph pass — augments the structural graph with intra-procedural data-flow. For each function and method, one STATEMENT node is emitted per top-level AST statement. An intra-procedural reaching-definition scan then identifies def-use pairs and emits DATAFLOW_DEF_USE / DATAFLOW_USE_DEF edges connecting defining statements to using statements, tagged with the flowing variable names. The graph is built once and cached by commit hash.

Phase 2 — Three-Tier Tool API

ARISE exposes the graph through a layered tool API. The three tiers are independently ablatable — a host can mount any subset:

Tier Tools Purpose
1 — Structural Retrieval search_entities, get_entity_info, traverse_relations, get_enclosing_scopes, get_code_span Navigate the repository structure
2 — Data-Flow Slicing get_dataflow_slice, explain_slice Trace def-use chains within functions
3 — Context Bundling build_context_bundle, rank_suspect_regions Rank suspects and pack code under a token budget

All tools share a RetrievalSession that is built once per repository.

Phase 3 — Agent Loop (host-supplied)

ARISE requires only that the host framework (a) inject the ARISE tool schema into the model context and (b) dispatch the resulting tool calls across its turns. The configs/ directory in this repository provides ready-made SWE-agent YAML overlays and system prompts for all experimental conditions. Any other tool-call-capable framework can wrap the same Python API directly.


Mounting on a Host Framework

Using the Python API directly

Any framework that can make Python function calls can use ARISE by wrapping the tool functions as tool definitions and calling them via a RetrievalSession:

from arise import RetrievalSession, search_entities, get_dataflow_slice

session = RetrievalSession.build(repo_path)

def arise_search(query: str, top_k: int = 20):
    return [r.to_dict() for r in search_entities(session.graph, session.entity_index, query, top_k=top_k)]

def arise_get_dataflow_slice(file_path: str, line: int, variable: str, direction: str = "backward"):
    result = get_dataflow_slice(session.graph, session.stmt_index, file_path, line, variable, direction)
    if isinstance(result, str):
        return {"message": result}
    return [step.to_dict() for step in result]

Using the SWE-agent bundles

The swe_agent_bundle*/ directories inside the installed package are ready-made SWE-agent tool bundles (a config.yaml with JSON-schema tool definitions and a bin/ directory with command-line wrappers). To use them with SWE-agent:

Step 1: Build the ARISE wheel (for offline install inside SWE-bench containers)

cd arise/
uv build                   # produces dist/arise-*.whl

Step 2: Sync bundles into SWE-agent

Run this from inside your SWE-agent directory after installing ARISE:

import arise, shutil, pathlib

pkg = pathlib.Path(arise.__file__).parent
bundles = [
    ("swe_agent_bundle",         "tools/arise"),
    ("swe_agent_bundle_tier1",   "tools/arise-tier1"),
    ("swe_agent_bundle_tier2",   "tools/arise-tier2"),
    ("swe_agent_bundle_coarse",  "tools/arise-coarse"),
    ("swe_agent_bundle_explain", "tools/arise-explain"),
]
for src_name, dst_name in bundles:
    shutil.copytree(pkg / src_name, pathlib.Path(dst_name), dirs_exist_ok=True)

# Copy wheel for offline install inside containers
import glob
whl = sorted(glob.glob("../arise/dist/arise-*.whl"))[-1]
for dst in ["tools/arise", "tools/arise-tier1", "tools/arise-tier2",
            "tools/arise-coarse", "tools/arise-explain"]:
    shutil.copy(whl, dst)

Reproducing the Paper's Experiments

All experiments use SWE-agent as the host framework and SWE-bench Lite as the benchmark. The backbone model is Qwen2.5-Coder-32B-Instruct (AWQ-quantized checkpoint Qwen/Qwen2.5-Coder-32B-Instruct-AWQ) served locally via vLLM.

Experimental Conditions

Condition (paper name) Config files to stack
ARISE-Full arise.yaml + fl.yaml or apr.yaml
ARISE-Structural arise_tier1.yaml + fl.yaml or apr.yaml
ARISE-Slicing arise_tier2.yaml + fl.yaml or apr.yaml
ARISE-Coarse arise_coarse.yaml + fl.yaml or apr.yaml (set ARISE_COARSE_GRAPH=1)
ARISE-ExplainSlice arise_explain.yaml + fl.yaml or apr.yaml
Baseline (shell-only) fl_baseline.yaml or apr_baseline.yaml (no ARISE config)

Step-by-Step

1. Install ARISE and build the wheel

git clone https://github.com/FARD-Lab/ARISE.git arise
cd arise
uv sync --extra eval
uv build                 # produces dist/arise-*.whl

2. Install SWE-agent

git clone https://github.com/princeton-nlp/SWE-agent
cd SWE-agent
pip install -e .

3. Sync ARISE tool bundles into SWE-agent (run from the SWE-agent directory)

python - <<'EOF'
import arise, shutil, pathlib, glob
pkg = pathlib.Path(arise.__file__).parent
bundles = [
    ("swe_agent_bundle",         "tools/arise"),
    ("swe_agent_bundle_tier1",   "tools/arise-tier1"),
    ("swe_agent_bundle_tier2",   "tools/arise-tier2"),
    ("swe_agent_bundle_coarse",  "tools/arise-coarse"),
    ("swe_agent_bundle_explain", "tools/arise-explain"),
]
for src_name, dst_name in bundles:
    shutil.copytree(pkg / src_name, pathlib.Path(dst_name), dirs_exist_ok=True)
whl = sorted(glob.glob("../arise/dist/arise-*.whl"))[-1]
for dst in ["tools/arise", "tools/arise-tier1", "tools/arise-tier2",
            "tools/arise-coarse", "tools/arise-explain"]:
    shutil.copy(whl, dst)
print("Done.")
EOF

4. Start the vLLM endpoint (requires a GPU with ≥ 48 GB VRAM for 32B model)

VLLM_USE_FLASHINFER_SAMPLER=0 \
python -m vllm.entrypoints.openai.api_server \
    --model Qwen/Qwen2.5-Coder-32B-Instruct-AWQ \
    --quantization awq_marlin \
    --tensor-parallel-size 1 \
    --port 8002

5. Run the FL task — ARISE-Full condition

# From inside SWE-agent/
OPENAI_API_BASE=http://localhost:8002/v1 \
OPENAI_API_KEY=dummy \
python -m sweagent.run.run_batch \
    --config config/default.yaml \
    --config ../arise/configs/arise.yaml \
    --config ../arise/configs/fl.yaml \
    --config ../arise/configs/swe_bench_lite.yaml \
    --output_dir outputs/qwen25coder/condfull-fl

6. Evaluate FL metrics

# From inside arise/
python evaluation/run_eval.py \
    --traj-dir ../SWE-agent/outputs/qwen25coder/condfull-fl \
    --output results_fl.json

7. Run the APR task — ARISE-Full condition

OPENAI_API_BASE=http://localhost:8002/v1 \
OPENAI_API_KEY=dummy \
python -m sweagent.run.run_batch \
    --config config/default.yaml \
    --config ../arise/configs/arise.yaml \
    --config ../arise/configs/apr.yaml \
    --config ../arise/configs/swe_bench_lite.yaml \
    --instances.evaluate=true \
    --output_dir outputs/qwen25coder/condfull-apr

8. Evaluate APR metrics

python evaluation/run_apr_eval.py \
    --run-dir ../SWE-agent/outputs/qwen25coder/condfull-apr \
    --output-pass-at-1 ../SWE-agent/outputs/qwen25coder/condfull-apr/pass_at1.json \
    --output results_apr.json

9. Spearman correlation (FL ↔ APR)

python evaluation/run_eval.py \
    --traj-dir ../SWE-agent/outputs/qwen25coder/condfull-fl \
    --pass-at-1 ../SWE-agent/outputs/qwen25coder/condfull-apr/pass_at1.json \
    --output results_fl_spearman.json

For the ARISE-Coarse condition, set ARISE_COARSE_GRAPH=1 in the environment before running SWE-agent. All other conditions follow the same pattern — substitute the appropriate arise*.yaml overlay from configs/.


Running Tests

uv run pytest tests/ -q
# or
pytest tests/ -q

All metric functions have unit tests in tests/eval/. Graph construction and tool behavior are tested against the tests/fixtures/simple_repo/ fixture.


Evaluation Metrics Reference

Metric Task Definition
File Recall@k FL 1 if any gold file appears in top-k predicted files
File MRR FL Mean reciprocal rank of first correct file
Function Recall@k FL 1 if any gold (file, function) pair appears in top-k
Function F1@k FL Harmonic mean of precision@k and recall@k over (file, function) pairs
Function MRR FL Mean reciprocal rank of first correct (file, function) pair
Line Recall@k FL 1 if any gold (file, line) pair appears in top-k
Line IoU FL Mean intersection-over-union of predicted and gold (file, line) sets
Coverage@budget FL Fraction of gold lines in build_context_bundle output under the token budget
Pass@1 APR Fraction of instances where the first generated patch passes all tests
Spearman ρ FL+APR Spearman rank correlation between Function Recall@1 and Pass@1

Limitations

  • Intra-procedural scope only. Data-flow slicing stops at function boundaries; def-use chains that cross a call edge require multiple tool calls. This accounts for approximately 45% of localization failures in our error analysis.
  • Python-specific. The graph builder uses Python's ast module. Supporting other languages requires a language-specific frontend.
  • Simplified call resolution. Call edges are resolved name-based; dynamic dispatch through attribute access or getattr is not tracked.
  • Token estimation. build_context_bundle uses ~4 chars/token as a rough heuristic; actual token counts depend on the tokenizer.

About

Enhancing Agentic Software Engineering with Repository-level Code Graph

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages