Skip to content
#

data-contamination

Here are 26 public repositories matching this topic...

Python .pyc decompiler (3.0–3.14) with a contamination-aware benchmark harness. Rule-only pass + one Codex call per module; evaluated on fuzz-synthetic (LLM-naïve) and *-obf (anonymised) corpora to put a number on the memorisation share. Three independent PyPI packages: pychd, pychd-pyfuzz, pychd-pyobf.

  • Updated May 27, 2026
  • Python

Do quality filters pull benchmark questions into pretraining corpora? Injects MMLU, GSM8K, GPQA, ARC, HellaSwag, PIQA and TruthfulQA items into a corpus and ranks them with six quality classifiers, including DCLM, FineWeb-Edu and Gaperon.

  • Updated Aug 19, 2026
  • Python

Leakage-free real-time evaluation of open-weights LLMs for US CPI inflation forecasting. Introduces the memorization premium (seen vs. unseen forecast-error gap) and a three-role decomposition (direct forecaster, FOMC-text extractor, combiner). Reproduces every number in the IJF manuscript's Table 3 from the committed checkpoint.

  • Updated Aug 9, 2026
  • Python

A contamination-resistant research platform for testing whether public market data carries tradeable information. Ten pre-registered experiments on an explore -> holdout -> live-forward design; none survived.

  • Updated Aug 13, 2026
  • Python

Improve this page

Add a description, image, and links to the data-contamination topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the data-contamination topic, visit your repo's landing page and select "manage topics."

Learn more