Analysis code and non-confidential supporting material for:
Gabapentinoid use, off-label indications, and CNS-active co-prescribing in Catalonia: a population-based drug utilization study
This repository documents the complete analysis workflow used for the manuscript. It intentionally contains no individual-level, pseudonymised, or linkable CatSalut records. The repository supports code inspection and method verification; an end-to-end numerical rerun additionally requires authorised access to the restricted source data described in DATA_AVAILABILITY.md.
| Review goal | Start here | Restricted data needed? |
|---|---|---|
| Orient yourself to the repository | Reviewer guide | No |
| Trace a manuscript table or figure to its code | Manuscript output map | No |
| Inspect diagnostic and regulatory classification | Diagnostic mapping, diagnostic recategorisation, and evidence rules | No |
| Inspect the public manuscript supplement | Supplementary Data File 1 | No |
| Understand the execution order | Analysis scripts guide and reproducibility guide | No |
| Perform the full numerical rerun | Reproducibility guide and documented input schemas under data/raw/ |
Yes |
| Understand the data-sharing boundary | Data and code availability | No |
Reviewers can inspect the computational logic, classification rules, public mappings, software environment, and provenance of every manuscript output without the restricted data. Numerical reproduction of results derived from person-level records must take place in an authorised secure environment.
- R and R Markdown analysis code for historical utilization, regulatory/evidence classification, and CNS-active co-medication analyses.
- The prespecified diagnostic and regulatory mapping used by the code.
- The public Supplementary Data File 1 supplied with the manuscript.
- Input schemas, execution order, output provenance, and the software environment recorded for the final analysis.
- A privacy guard used locally and in GitHub Actions.
Dummy.ID-level records or any other patient-level table.- Raw prescription, dispensation, diagnosis, or co-medication files.
- Derived RDS objects, Excel checkpoints, patient-level audit extracts, or rendered review HTML files.
- The original anonymised ZIP. Pseudonymised or anonymised records remain governed data and are not made public here.
R/ Shared analysis and validation functions
scripts/ Ordered analysis scripts and R Markdown notebooks
data/reference/ Public, non-patient diagnostic/regulatory mapping
data/public/ Public manuscript supplement
data/raw/ Empty documented locations for restricted inputs
environment/ Recorded R and package versions
docs/ Reviewer, reproducibility, and output documentation
tools/ Public-release privacy and code-parse checks
outputs/ Empty; generated files are deliberately ignored
The scripts/ and R/ directories each have a short index explaining the role and order of their files.
- R 4.5.2 was used for the recorded analysis.
- Pandoc is required to render the R Markdown notebooks.
- The exact package versions are listed in
environment/packages.csv, and the full recorded session is inenvironment/session-info.txt.
Install the required R packages from the repository root:
Rscript R/install_packages.R- Obtain the governed input files through the applicable CatSalut approval and secure-analysis process.
- Place them at the exact paths documented under
data/raw/; do not rename identifiers or worksheet headers. - Run the pipeline from the repository root:
Rscript run_all.RThe runner executes all required steps in dependency order and stops on failure so that stale and newly generated results cannot be mixed. On Windows/OneDrive it uses a short temporary rendering path and then synchronises the generated outputs back to the repository working copy.
See the reproducibility guide for the detailed workflow and the manuscript output map for the mapping from manuscript figures and tables to code.
Before committing or publishing, run:
python tools/check_public_release.py
Rscript tools/parse_all_code.RThe first check rejects raw or processed data, checkpoints, patient-level filenames, common healthcare-data formats, rendered HTML, unexpected spreadsheets, and local absolute paths. The two approved public workbooks are pinned by SHA-256 hash. The second check parses all R scripts and R Markdown code chunks without requiring the restricted inputs.
The repository is designed to remain consistent with the manuscript statement that individual-level CatSalut data are subject to data-protection and governance restrictions and are not publicly available. The diagnostic code mapping is supplied as Supplementary Data File 1. See DATA_AVAILABILITY.md for the operational interpretation.
The software code, scripts, workflows, and executable analysis logic are licensed under the MIT License. Documentation, the public supplement, the diagnostic/regulatory mapping, and other non-software public materials are licensed under CC BY 4.0, except where otherwise stated.
Citation metadata are provided in CITATION.cff. The software version and article DOI should be updated when the final article citation is available.
The remaining publication and versioning tasks are recorded in RELEASE_CHECKLIST.md.