Skip to content

fix(worker): keep a run log from making its document impossible to update - #113

Draft
BatLeDev wants to merge 1 commit into
masterfrom
fix-run-log-size-guard
Draft

BatLeDev wants to merge 1 commit into
masterfrom
fix-run-log-size-guard

Conversation

@BatLeDev

@BatLeDev BatLeDev commented Sep 15, 2026 •

Copy link
Copy Markdown
Member

The whole log of a run is stored in its mongo document with an unbounded $push. A plugin logging a ~5KB message on each of 3 000+ pages filled the 16MB document: the task failed, and since finish() also has to update the document, the run could never be marked finished or killed and the kill loop retried it every 20s forever (Resulting document after update is larger than 16777216).

  • task.ts: msg/extra of a log entry are truncated beyond worker.task.maxLogEntryLength (10 000 chars, the cap already hard-coded for axios errors, now shared). When a $push is refused because the document is full, the log is cut to its tail, an explicit error entry is written and the task fails.
  • runs.ts: finish() applies the same truncation and retries when the document is already too large, so runs that predate this guard (the one currently stuck in production) get finished on the next kill loop iteration.
  • runs-operations.ts: pure helpers isDocumentTooLargeError and truncateLogValue, unit tested.
  • New config worker.task.maxLogEntryLength / WORKER_TASK_MAX_LOG_ENTRY_LENGTH.
  • log-flood.api.spec.ts with a processing-log-flood fixture plugin fills the document in ~25s and checks the run ends in error, every entry is truncated, the guard message is present and finish() could still write its entry.

Why: a run must never become impossible to finish or kill because of its own log; a plugin that floods its log should fail with a clear message instead of wedging the worker.

Heads-up: a non-string extra whose serialization exceeds the cap is now stored as a truncated string instead of an object (the UI renders both). The mongo error is matched on code 17419 or /larger than \d+/ in the message, since the driver wraps it in a Plan executor error message. On deploy, finish() will close the run currently stuck in production, keeping only its last 100 log entries.

The plugin side is fixed separately in processing-import-api (one progress entry instead of a log per page, detection of an API that ignores its pagination parameter).

@github-actions github-actions Bot added the fix label Sep 15, 2026
…date

The whole log of a run is stored in its mongo document with an unbounded
$push. A plugin logging a long message on every page of a paginated
import filled the 16MB document: the task failed, and since finish()
also needs to update the document, the run could never be marked
finished or killed and the kill loop retried it forever.

- truncate msg/extra of a log entry beyond worker.task.maxLogEntryLength
  (10 000 chars, the cap already applied to axios errors)
- when a $push is refused because the document is full, keep the tail
  of the log, write an explicit error entry and fail the task
- in finish(), apply the same truncation and retry when the document is
  already too large (runs predating this guard)
@BatLeDev
BatLeDev force-pushed the fix-run-log-size-guard branch from abe46de to 05ae1af Compare September 15, 2026 09:45
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant