Skip to content

feat: ingest events from PredictHQ and AskNews onto a shared normalized form - #41

Closed
Alexcchip wants to merge 24 commits into
mainfrom
external-factors-events
Closed

Alexcchip wants to merge 24 commits into
mainfrom
external-factors-events

Conversation

@Alexcchip

Copy link
Copy Markdown
Collaborator

Summary

Adds event ingestion from two external sources — PredictHQ (structured listings) and AskNews (news articles) — behind one shared NormalizedEvent contract that future event sources should also target. Events merge onto canonical rows and feed the pricing model as per-night expected attendance.

Type of change

  • feat — new feature
  • docs — documentation only
  • test — adding or updating tests
  • chore / build / ci — tooling, deps, or pipeline changes

Changes

  • ingest/normalize.py — NormalizedEvent, dedupe_key (normalized title + city + local start date), and the shared enums. Validates on construction.
  • ingest/predicthq.py — structured listings near a point, paged by updated.gte. Requests deleted state so upstream cancellations propagate.
  • ingest/asknews.py + ingest/textdates.py — infers events from articles: per-city scoped search, prose date resolution, and a named reject reason for everything dropped.
  • ingest/repository.py — upsert_event, source-agnostic. Resolves by (source, source_ref) before dedupe_key so a revised start date updates rather than orphans. PredictHQ outranks AskNews on contested fields.
  • ingest/run.py — ingest CLI, with --dry-run to fetch and print without writing.
  • services/features/events.py — stored events to the per-night expected_attendance feature.
  • docs/event-extraction.md — architecture, file inventory, and how to test it.
  • Deps: adds httpx, promotes scikit-learn to a runtime dependency.

Database migrations

  • Includes an Alembic migration (just migrate-create "...")
  • Migration runs cleanly (just migrate) and rolls back (just migrate-down)

Two migrations: canonical events + event_observations, and status on observations. Verified down and back up with data in the table.

Checklist

  • Commits follow Conventional Commits
  • Lint passes (just lint) and code is formatted (just format)
  • Pre-commit hooks run clean
  • Updated docs / .env.example
  • Tested locally (just dev) and verified the API at /docs — N/A, no API surface in this PR; ingestion runs via CLI and scripts

How to test

Full instructions in docs/event-extraction.md → "How to test it".

Quickest check — no credentials, no database, no API credits:

cd backend && uv run python scripts/run_asknews.py --dry-run

Then the suite: just test (132 passing; test_repository.py and test_event_features.py need Postgres via just db-up).

Only --live spends AskNews credits — use --runs 1 if you try it.

Notes for reviewers

Please sanity-check NormalizedEvent. It is the contract every future event source has to fit, and the expensive thing to change later. Specifically: dedupe_key excludes category on purpose (sources disagree, and that must not split one event in two), and coarse dates are kept at MONTH/QUARTER precision rather than rounded away.

Known gaps, not blockers:

  • AskNews extraction has a precision problem. A live run kept 47 events, but the set includes historical entities ("Irish Civil War", "midterms"), wrong-city assignments (Reading Festival → Dublin) and one festival split across four rows. entities.Event is not a scheduled-event list.
  • confidence measures extraction certainty, not event-ness — a live run scored "midterms" at 0.95. Not safe as a quality filter. AskNews events carry no attendance so they contribute nothing to pricing today, but nightly_attendance multiplies by confidence, so this must be fixed before they ever do.
  • Two entry points remain: python -m foresight.ingest.run (PredictHQ) and scripts/run_asknews.py (AskNews). Both now support --dry-run; folding AskNews in as --source asknews would collapse them.
  • Rejected articles are counted, not persisted. No quarantine table yet.

Unverified: --dry-run on the PredictHQ path has no live token behind it on my side — it is covered by a unit test and confirmed to touch no database, but nobody has run it against real PredictHQ data.

Heads-up for anyone running this locally: if you already have Postgres on your machine it binds 127.0.0.1:5432 and wins over Docker, so just db-up succeeds while the app talks to a different database.

Alexcchip and others added 24 commits October 6, 2026 12:57
…as_row

The rebase of 9b83901 onto the src/ -> foresight/ refactor left ingest/
under backend/src/, so foresight.ingest was unimportable and broke model
and Alembic imports. as_row was misspelled mas_row; the tests and the
documented contract both use as_row.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Settings now also reads the repo-root .env, which docker-compose already uses, so commands run from backend/ pick up the API credentials.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Resolves by (source, source_ref) before dedupe_key so a start-date revision updates the existing row instead of orphaning it. Structured sources outrank text-extracted ones, and a revision that lands on another row's key merges the two.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Geo-filters around the hotel, pages through results, and fetches incrementally on updated.gte. Requests deleted states explicitly, since the API returns only active events by default.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
python -m foresight.ingest.run --source predicthq --since 7d --report prints fetched/new/updated/unchanged/rejected plus market totals. Records failing validation are logged and counted as rejected rather than aborting the run.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
PredictHQ often lists one event several times, typically one copy active and one deleted. The first run correctly ignored the deleted copy, but on every later run its sighting already pointed at the event, so it counted as the event's own record and its deletion was applied (20 live Dublin events). Live copies that differed slightly also overwrote each other on every run, so an identical rerun reported 238 updated instead of 0.

Each observation now stores the status it reported, and every sighting of an event is ranked the same way: source precedence, then live over withdrawn, then latest upstream revision, with source_ref as a tie-break. Only the top-ranked sighting sets the row, so reruns settle on one winner. Migration backfills status for existing PredictHQ sightings from their payload.

Repository tests now clear both tables inside their rolled-back transaction, so locally ingested events no longer break their counts.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
services/features/events.py turns the events table into the expected_attendance feature the models train on: same-night events add up, multi-day events are spread over their nights, withdrawn events are dropped, provisional ones count by confidence, and an optional radius drops far venues.

scripts/forecast_events.py prices upcoming nights with and without the real events. generate_mock_data.generate can take attendance as input so the mock competitor, fare and flight placeholders respond to real events, as they do in the training data.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Five PredictHQ ingest runs (Dublin and Boston, including same-query reruns), the pre-fix runs that exposed the duplicate-listing bug, a timed-out run, the model forecast, and a findings page summarising them.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…t CLI

docker compose up already ran the migrate service, but `just db-up` started only the database, and the ingest CLI run from the host never ran migrations at all. Pulling the new observation status column then broke the ingest with a missing-column error until someone ran `just migrate` by hand.

`just db-up` now starts the migrate service with the database, and the ingest CLI upgrades to head before fetching (a no-op when current). Alembic gets no ini file there, so it leaves the run report's logging alone. The findings page drops the now-unneeded manual migration step.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…as_row

backend/src/ is outside the hatch package and the ruff src root, so
foresight.ingest was not importable and foresight/database/models/external.py
could not load. Same move Heather made locally, so the two land identically.
env_file also reads ../.env so commands run from backend/ pick up the
repo-root file.
Text sources write dates as people do. Resolves day ranges (including
ones crossing a month or a new year), month-year and season/quarter
spans, reporting the precision it achieved rather than rounding a vague
date into a confident one. Relative forms are refused outright.
Articles in, provisional events out, with every drop recorded on
IngestReport as a named reason. collect_events is the single call a
scheduled worker makes; hours_back is the scheduling knob.
Replays recorded batches by default so the pipeline is demonstrable
without credentials; --live goes through the same collect_events call.
Brings in the Windows justfile fix. pyproject.toml conflicted only because both sides appended a runtime dependency; keeps both httpx and scikit-learn.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
An unscoped run rejected 188 of 198 articles on city alone -- paid-for noise. One search per city, filtered to that city's country, cuts that to 60 and raises the keep rate from 2% to 38% on three calls instead of four.

Keep rate is not quality. The kept set still contains historical entities ("Irish Civil War", "midterms"), wrong-city assignments (Reading Festival -> Dublin), and one Tribeca Lisboa run split across four rows. entities.Event is not a scheduled-event list, and confidence scores extraction certainty rather than event-ness, so it cannot yet be used as a quality filter. Precision work still outstanding.
…ternal-factors-events

# Conflicts:
#	.env.example
#	backend/foresight/config.py
#	backend/pyproject.toml
Replaces the in-memory EventStore stand-in, which existed only because repository.py was not yet on this branch. One transaction per run; new/updated/unchanged now come from repository.Outcome rather than a local dict.
CityTarget.slug was never called. horizon_days and the plan_searches queries override were knobs nothing turned outside tests. No behaviour change.
Fetch and print without writing, so PredictHQ output can be inspected the way scripts/run_asknews.py --dry-run already allows. With --lat/--lon/--city it touches no database and skips the automatic migration; --hotel-id still needs one to resolve the market.
@Alexcchip Alexcchip linked an issue Oct 7, 2026 that may be closed by this pull request
@Alexcchip Alexcchip closed this Oct 7, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

External factor integration: Events

2 participants