Add automated retry for network errors in install.sh and CI - #374
Merged
Merged
Conversation
FPM git fetches are not resilient to transient network errors, they drop dead after a single failure with no retry. Replace that with scripted git fetches with automatic retry.
Member
Author
|
Once merged, I plan to clone the new CI bits to our other projects using the same CI framework. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
The Caffeine CI Action includes fetching many prerequisites over the network, from both GitHub itself and also from Homebrew and the Ubuntu package servers.
install.shalso clones the assert and julienne dependencies from GitHub. Most of these network operations are currently vulnerable to transient network faults, which periodically arise and usually fail the entire job, creating meaningless failure noise and wasting maintainer time and runner cycles.This PR adds automated retry for most network operations that take place in both Caffeine CI Actions (e.g. to download build prerequisites), and within install.sh (when fetching assert and julienne, or running Homebrew), thereby improving CI/build resilience to transient network problems.
Examples
Before this PR
Here are some recent examples of job failures caused by network transients on otherwise correct code:
Homebrew failure in CI step:
https://github.com/BerkeleyLab/caffeine/actions/runs/29650246717/job/88095270587#step:6:27

Homebrew failure in install.sh:
https://github.com/BerkeleyLab/caffeine/actions/runs/29650246717/job/88095270698#step:16:133
Git fetch failures within
fpm build:https://github.com/BerkeleyLab/caffeine/actions/runs/31532814604/job/93916730519#step:16:1886

And many others:
https://github.com/BerkeleyLab/caffeine/actions/runs/34671967323/job/103494898136#step:21:2559
https://github.com/BerkeleyLab/caffeine/actions/runs/34799622918/job/103839489637#step:21:2530
https://github.com/BerkeleyLab/caffeine/actions/runs/31532814604/job/93916730698#step:16:2745
https://github.com/BerkeleyLab/caffeine/actions/runs/31630099427/job/94226514338#step:16:2751
After this PR
Example retries on a transient
git clonefailure, now handled (with a warning) byinstall.sh:https://github.com/bonachea/caffeine/actions/runs/36297846900/job/108559968903#step:22:2009

https://github.com/bonachea/caffeine/actions/runs/36227503338/job/108364264635#step:30:372

Example retry on a transient

brew updatefailure, in scripted CI step:https://github.com/bonachea/caffeine/actions/runs/36213574257/job/108527369168#step:8:201
Related but orthogonal
GitHub's job setup step already includes some retry when cloning dependent actions, example:
https://github.com/bonachea/caffeine/actions/runs/36297838177/job/108559969245#step:1:33