fix(autoscaler,heartbeat): react to runner pods that cannot be scheduled - #28
Merged
Merged
Conversation
…hedulable runner pods fail_safe built its status from every target's maxRunners, but a measurement failure in the first target's loop iteration fired it before the later targets were read, so it died on an unbound MAX index and never lowered anything. Read each target's running count, maxRunners and pods first; a failure there now exits non-zero, and memory measurement failures fail safe with the full state. The memory budget sums every node's capacity, so it cannot see a node that is tainted for disk pressure or a request that fits no single node. A runner pod stuck Unschedulable now lowers the combined maxRunners to the running count and blocks raising it.
… be scheduled Ready nodes and a running controller did not mean a runner could be placed: with every schedulable node full and another tainted for disk pressure, jobs queued behind the runners already busy while runner-fallback-action kept routing to the fleet. The heartbeat now stops refreshing the gist once a runner pod has been unschedulable for HEARTBEAT_UNSCHEDULABLE_AFTER_SECONDS (default one heartbeat interval), so jobs fall back to GitHub-hosted runners until the pods schedule. Grant the heartbeat's service account read access to pods, and document both this and the autoscaler's unschedulable handling.
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
|
🎉 This PR is included in version 1.8.0 🎉 The release is available on:
Your semantic-release bot 📦🚀 |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
A node tainted for disk pressure left runner pods Pending for days while the heartbeat stayed green and the autoscaler kept targeting capacity the cluster could not place, so jobs queued behind the two runners that fit.
The autoscaler's fail_safe died on an unbound MAX index when a measurement failed during the first target's loop iteration, so it never lowered anything. State for every target is now read before any measurement. A runner pod stuck Unschedulable now lowers maxRunners to the running count and blocks raising it. The heartbeat stops refreshing the gist once a runner pod has been unschedulable for a heartbeat interval, so runner-fallback-action falls back to GitHub-hosted runners. The heartbeat's service account gets read access to pods for that.
Tested the three autoscaler paths (healthy, unschedulable, failing top pod) and the heartbeat against a fake kubectl; the original script reproduces the unbound variable crash on the same input. Not run against a live cluster.