Skip to content

fix(curation,dataset): query check_runs_latest in coverage and dataset SQL (#676) - #682

Closed
shobhitagnihotri69 wants to merge 1 commit into
Hebbian-Robotics:mainfrom
shobhitagnihotri69:fix/676-check-runs-latest-coverage-dataset
Closed

shobhitagnihotri69 wants to merge 1 commit into
Hebbian-Robotics:mainfrom
shobhitagnihotri69:fix/676-check-runs-latest-coverage-dataset

Conversation

@shobhitagnihotri69

@shobhitagnihotri69 shobhitagnihotri69 commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Fixes #676

As identified in #676 and confirmed by maintainer guidance in #676 (comment):

  • default_dataset_sql in src/hflow/dataset.py:231 and _collect_coverage in src/hflow/curation.py:752 were querying the append-only check_runs table instead of check_runs_latest.
  • As a result, when a check or enrichment succeeded on an initial run and subsequently encountered an error on replay, the stale settled record allowed the episode to remain in default_dataset_sql and falsely retain its coverage.

Changes

  1. src/hflow/curation.py: Switched query to FROM check_runs_latest in _collect_coverage.
  2. src/hflow/dataset.py: Switched query to FROM check_runs_latest in default_dataset_sql.
  3. Documentation:
    • Updated docs/CATALOG.md (dataset policy and coverage sections).
    • Updated default_dataset_sql docstrings in src/hflow/dataset.py to clarify latest-run rules.
  4. Regression Tests:
    • tests/test_dataset.py: Added test_a_check_that_errors_on_replay_excludes_its_episode confirming that a noncritical check replay error excludes the episode from default_dataset_sql while episode status remains ok.
    • tests/test_catalog_curation.py: Added test_a_check_that_errors_on_replay_loses_coverage on a 2-episode corpus, demonstrating that when a check errors on replay for one episode, its coverage decreases from 2/2 (100%) to 1/2 (50%) while cleanly preserving the total coverage denominator.

All test suites and ruff formatting/lint checks pass cleanly.

…t SQL (Hebbian-Robotics#676)

- default_dataset_sql and _collect_coverage now query check_runs_latest
  instead of append-only check_runs.
- A noncritical check that passes and later errors on replay is excluded
  from the default dataset and withdraws its coverage while retaining the
  coverage denominator.
- Adds regression tests in test_catalog_curation.py and test_dataset.py.
- Updates dataset policy and coverage documentation in docs/CATALOG.md
  and dataset.py docstrings.

Closes Hebbian-Robotics#676
@greptile-apps

greptile-apps Bot commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 4/5

[Critical risk] Changes how dataset selection filters check results.

This PR should not merge until the coverage report keeps checks that fall to zero coverage.

Findings

  1. P1 Failed check disappears from coverage ▶
Summary

The default dataset query and curation coverage now use each check’s latest run, so an error on replay no longer leaves an earlier success counted. The docs and regression tests describe and exercise these rules.

  • Dataset eligibility follows the latest settled result for each step.
  • Coverage drops replay errors from the check count while keeping the episode total.
Diagram
%%{init: {'theme': 'neutral'}}%%
flowchart LR
  A[Check run history] --> B[Latest run per episode and check]
  B --> C[Default dataset selection]
  B --> D[Coverage count]
  D --> E[Coverage report]
Loading

Reviews (1) · Last reviewed commit: "fix(curation,dataset): query check_runs_..."

Comment thread src/hflow/curation.py
f"""
SELECT check_name, count(DISTINCT episode_id) AS episodes_ran
FROM check_runs WHERE status IN ({status_list})
FROM check_runs_latest WHERE status IN ({status_list})

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Failed check disappears from coverage

When a check errors on replay for every episode, check_runs_latest has no result that counts as “ran” for that check. This query drops the check entirely, so the report says “no check runs recorded” instead of showing 0/N. Keep checks with zero coverage in the report so teams can see the loss of evidence.

Knowledge Base Used: Datasets, catalog, and curation

@kstonekuan

Copy link
Copy Markdown
Contributor

Thanks for this. #678 was opened first for #676 and has now merged, so I'm closing this one.

@kstonekuan kstonekuan closed this Oct 4, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

fix(curation,dataset): default_dataset_sql and _collect_coverage query append-only check_runs instead of check_runs_latest, concealing errored checks

2 participants