-
Notifications
You must be signed in to change notification settings - Fork 2
TreeScan Overview
TreeScan is a statistical surveillance method used to detect unusual patterns in healthcare data.
In this project, TreeScan analyzes emergency department (ED) visit data and searches for clusters of diagnoses that occur more frequently than expected within a recent time window.
Rather than monitoring a small set of predefined syndromes, TreeScan evaluates all diagnosis groups simultaneously using a hierarchical tree of ICD-10 codes. This allows the method to detect unexpected patterns that may not fit into traditional surveillance categories.
TreeScan therefore supports syndrome-agnostic surveillance, meaning it can detect signals even when the underlying syndrome has not been predefined.
TreeScan searches for combinations of:
- diagnosis groups
- time periods
where the number of observed cases is unusually high compared with the expected baseline.
For example, TreeScan might identify:
- a sudden increase in respiratory diagnoses over several days
- an unusual cluster of gastrointestinal diagnoses at a specific point in time
- a spike in a rare diagnosis category that is not normally monitored
These unusual patterns are called signals.
A TreeScan signal indicates that the observed number of cases in a particular diagnosis group and time window is unlikely under the recent baseline distribution of diagnoses.
Importantly, a signal does not mean that an outbreak or causal event has been confirmed.
Signals can arise for many reasons, including:
- real public health events
- changes in clinical coding practices
- data reporting artifacts
- random statistical fluctuation
TreeScan therefore functions as a screening tool that identifies patterns requiring further review.
Traditional syndromic surveillance systems typically monitor a fixed set of syndromes, such as respiratory illness or gastrointestinal illness.
While useful, this approach can miss unusual patterns that fall outside predefined categories.
TreeScan addresses this limitation by:
- scanning across all diagnosis codes
- evaluating many potential clusters simultaneously
- identifying patterns that may not have been anticipated
This makes TreeScan particularly useful for detecting novel or emerging health events.
A full pipeline for running TreeScan-based analyses using R and the TreeScan software.
This project provides a structured workflow to prepare data, run TreeScan, and process results using an R-based pipeline.
- Click the green Code button on GitHub
- Select Download ZIP
- Extract the ZIP file
- Locate the
treescan_projectsubfolder - Move
treescan_projectto your desired working directory
Download and install RStudio:
https://posit.co/download/rstudio-desktop/
Download TreeScan from:
https://www.treescan.org/download_treescan.html
- You must create an account before downloading
- Choose version based on your environment:
- Windows → if running locally
- Linux → if running on a server
- Select the NON-graphical version
- The standard (graphical) version may cause IT/access issues
After downloading:
Move the TreeScan files into the correct subfolder inside treescan_project:
| Environment | Folder |
|---|---|
| Windows | TS_windows/ |
| Linux | TS_linux/ |
- Launch RStudio
- In the bottom-right file explorer:
- Navigate to:
treescan_project/code/ - Open:
run_full_pipeline.R
- Navigate to:
Before running, update the following:
Update line 4 to match your local path:
setwd("~/TreeScan-implementation/treescan_project")Replace with wherever you saved treescan_project.
Modify these variables depending on your setup:
server <- FALSE # Set to TRUE if running on a server
first_time <- TRUE # Set to FALSE after first run- Run the script in RStudio
The pipeline will:
- Execute TreeScan
- Process outputs
- Complete the full analysis workflow
- Ensure the correct TreeScan version is placed in the matching folder (
TS_windowsorTS_linux) - Using the non-graphical version is required
- Incorrect working directory paths will cause errors