Skip to content

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

1 watching

Forks

Repository files navigation

ATM-Bench — Project Page & Live Leaderboard

Static site for ATM-Bench: According to Me — Long-Term Personalized Referential Memory QA.

Project Page Live Leaderboard arXiv Code Dataset

This repository hosts the landing page and the community leaderboard for ATM-Bench (Mei, Chen, Yang, Hou, Li, Byrne — University of Cambridge). It is a hand-written static site — no build step, no dependencies, no generator. Every page is self-contained and can be opened directly as a file.

The benchmark code, harness, and canonical results live in the companion repo JingbiaoMei/ATM-Bench. That repo's README is the source of truth for scores; this site visualizes them.

Pages

Page What it is
index.html Paper landing page — hero + demo video, the four "memory puzzle" cards, benchmark comparison, method figure, and BibTeX. Bilingual (EN / 中文).
leaderboard.html Live leaderboard with three tabs (ATM-Bench, ATM-Bench-Hard, NIAH-100), sortable columns, and system-type filters.

Repository structure

atmbench.github.io/
├── index.html             # landing page (Academic Project Page / Nerfies template)
├── leaderboard.html       # leaderboard — inline CSS + data + vanilla JS, no build
├── .nojekyll              # serve files as-is (skip Jekyll processing)
├── paper/                 # drop the camera-ready PDF here, link it from index.html
└── static/
    ├── css/               # Bulma, Font Awesome, Academicons + page-local index.css
    ├── js/                # Bulma carousel/slider + index.js
    ├── images/            # teaser (ATM-Bench-Demo.png), method (ATM-Method.png), favicon
    ├── videos/            # demo videos (EN / CN)
    └── pdfs/              # supplementary PDFs

Local preview

No toolchain required — just serve the folder:

cd atmbench.github.io
python3 -m http.server 8765
# open http://localhost:8765/  and  http://localhost:8765/leaderboard.html

Opening the file:// URL also works; only Google Fonts and Analytics fail silently under file:// and neither affects layout.

Editing the site

  • Text / authors / links / BibTeX → index.html.
  • Teaser image & social preview → static/images/, wired through the og:image / twitter:image meta tags.
  • Paper PDF → place it in paper/ and update the link in index.html.
  • Styling → reuse the CSS custom properties in the :root block (warm beige page with rust / blue / green / gold / plum accents). Add new tokens only when a genuinely new color is needed.
  • Bilingual copy → the site is EN/中文. Any user-visible string uses a data-i18n (or data-i18n-html) attribute and a matching key in the translations object; add both languages when you add text. Browser language is auto-detected on load.

Updating the leaderboard

All leaderboard data is a plain JavaScript array (TRACKS) near the bottom of leaderboard.html — adding a result is a one-line append to the relevant rows list.

Each row's type is one of Oracle, Agent, Memory, RAG, NIAH, which drives its color pill. The first four also appear as filter chips on the ATM-Bench / ATM-Bench-Hard boards; the NIAH board has no filter chips.

// ATM-Bench-Hard, append to that track's rows:
{ type: 'Agent', harness: 'Claude Code', model: 'Claude Opus 4.8',
  qs: 41.60, recall: null, total_tokens: 4.42,
  link: 'https://github.com/anthropics/claude-code' },

Conventions:

  • Use verbatim numbers from the source (paper table or the ATM-Bench README). Unknown fields are null → the renderer prints - and excludes them from sorting and "best-in-column" highlights.
  • QS (Question Score) and Recall@10 are percentages (two decimals, e.g. 41.60). Total tokens are stored in millions (4.42 → 4.42M). Cost is stored in USD and displayed with two decimals. The Hard table shows these optional fields; Main costs stay in run reports until multiple submissions report them. Preserve exact cost accounting in the linked run report. Cost/token best highlights require at least two reported comparable values in the current filter. Index time is in hours and exists only on the ATM-Bench (first) track.
  • Set link to the system's canonical repo/paper to make the harness name clickable.
  • Put the answer model in the row label. Use short cell footnotes for judge, input, and attribution disclosures; link longer evidence instead of duplicating it in legends.
  • The headline view filters Oracle off by default (it is a no-retrieval upper bound); click the Oracle chip to include it.

Validation

Run node tools/verify_leaderboard.mjs before submitting. It checks numeric bounds, row coverage, best-value eligibility, visible provenance disclosures, and the reported score summary scores against the public per-question score summary.

Deployment

Served by GitHub Pages from the main branch of the org user-site repo atmbench/atmbench.github.io → https://atmbench.github.io. Pushing to main publishes. .nojekyll keeps Pages from running Jekyll, so the static files are served exactly as committed.

Note: *.github.io subdomains (e.g. leaderboard.atmbench.github.io) are not supported by GitHub Pages without a custom domain — the leaderboard lives at atmbench.github.io/leaderboard.html.

Citation

@article{mei2026atm,
  title={According to Me: Long-Term Personalized Referential Memory QA},
  author={Mei, Jingbiao and Chen, Jinghong and Yang, Guangyu and Hou, Xinyu and Li, Margaret and Byrne, Bill},
  journal={arXiv preprint arXiv:2603.01990},
  year={2026},
  url={https://arxiv.org/abs/2603.01990},
  doi={10.48550/arXiv.2603.01990}
}

Links

Credits

The landing page builds on the Academic Project Page Template (itself based on Nerfies), released under CC BY-SA 4.0.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages