feat: add node_readiness_rules with enforcement_mode and dry_run labels - #448
Conversation
✅ Deploy Preview for node-readiness-controller canceled.
|
|
Hi @rawadhossain. Thanks for your PR. I'm waiting for a kubernetes-sigs member to verify that this patch is reasonable to test. If it is, they should reply with Tip We noticed you've done this a few times! Consider joining the org to skip this step and gain Once the patch is verified, the new status will be reflected by the I understand the commands that are listed here. DetailsInstructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the kubernetes-sigs/prow repository. |
|
/cc @ajaysundark |
|
/cc @AvineshTripathi |
|
/ok-to-test |
|
/retest |
AvineshTripathi
left a comment
There was a problem hiding this comment.
Since we're already calling ListRules in the collector, we can compute the enforcement_mode/dry_run breakdown there instead of maintaining a separate in-memory sync. The local ruleCache can lag behind actual cluster state and can be unreliable.
cc @ajaysundark wdyt?
58c5646 to
3ac4f69
Compare
|
@AvineshTripathi Good point. I tested both ways and found @ajaysundark since this metric wasn't part of the collector scope, should we switch or keep it separate for now? |
Signed-off-by: Rawad Hossain <rawad.hossain00@gmail.com>
3ac4f69 to
81c256d
Compare
I would prefer testing this delay on the scale before doing the scale. Can we run the scan? |
|
Yes we can and thanks it helped. Ran the scale test. Used the ruleCache staleness (current one):
Single changes are fine, but bigger difference shows up when a large number of rules change together, delay grows quite long. Scrape-time approach:
Since the collector already does that fetch, adding this one on barely adds anything. Based on these, would it make more sense to go with collector instead, specially for scale? Staleness gets pretty noticeable with bulk changes, also extra cost in the collector is very small. @AvineshTripathi @ajaysundark wdyt? I can make the changes if we decide to go with collector approach. |
|
@rawadhossain I'll go with the collector approach! that can be a standard we can go for |
|
@AvineshTripathi I also think so collector approach is the one to go for. Made the changes. Tested and everything's working as expected. cc @ajaysundark |
|
/lgtm havn't tested in local but looks good! |
a89e0b0 to
56d8ec6
Compare
56d8ec6 to
4dc8a70
Compare
|
/lgtm |
|
[APPROVALNOTIFIER] This PR is APPROVED This pull-request has been approved by: ajaysundark, rawadhossain The full list of commands accepted by this bot can be found here. The pull request process is described here DetailsNeeds approval from an approver in each of these files:
Approvers can indicate their approval by writing |
Description
node_readiness_rules{enforcement_mode, dry_run}while keepingnode_readiness_rules_totalunchanged for compatibility.node_readiness_rulesfrom the scrape-time ReadinessCollector. the original ruleCache push approach showed staleness scaling with rule count under bulk changesmonitoring.mdanddocs/TEST_README.md.from the design doc:
Related to #446
Type of Change
/kind feature
Testing
make test,make lint,go vet,gofmt -lall pass.Checklist
make testpassesmake lintpasses