Skip to content

fix(autoscaler,heartbeat): react to runner pods that cannot be scheduled - #28

Merged
Mearman merged 2 commits into
mainfrom
fix/runner-capacity-signals
Oct 3, 2026
Merged

Mearman merged 2 commits into
mainfrom
fix/runner-capacity-signals

Conversation

@Mearman

@Mearman Mearman commented Oct 3, 2026

Copy link
Copy Markdown
Member

A node tainted for disk pressure left runner pods Pending for days while the heartbeat stayed green and the autoscaler kept targeting capacity the cluster could not place, so jobs queued behind the two runners that fit.

The autoscaler's fail_safe died on an unbound MAX index when a measurement failed during the first target's loop iteration, so it never lowered anything. State for every target is now read before any measurement. A runner pod stuck Unschedulable now lowers maxRunners to the running count and blocks raising it. The heartbeat stops refreshing the gist once a runner pod has been unschedulable for a heartbeat interval, so runner-fallback-action falls back to GitHub-hosted runners. The heartbeat's service account gets read access to pods for that.

Tested the three autoscaler paths (healthy, unschedulable, failing top pod) and the heartbeat against a fake kubectl; the original script reproduces the unbound variable crash on the same input. Not run against a live cluster.

…hedulable runner pods

fail_safe built its status from every target's maxRunners, but a measurement
failure in the first target's loop iteration fired it before the later
targets were read, so it died on an unbound MAX index and never lowered
anything. Read each target's running count, maxRunners and pods first; a
failure there now exits non-zero, and memory measurement failures fail safe
with the full state.

The memory budget sums every node's capacity, so it cannot see a node that is
tainted for disk pressure or a request that fits no single node. A runner pod
stuck Unschedulable now lowers the combined maxRunners to the running count
and blocks raising it.
… be scheduled

Ready nodes and a running controller did not mean a runner could be placed:
with every schedulable node full and another tainted for disk pressure, jobs
queued behind the runners already busy while runner-fallback-action kept
routing to the fleet. The heartbeat now stops refreshing the gist once a
runner pod has been unschedulable for HEARTBEAT_UNSCHEDULABLE_AFTER_SECONDS
(default one heartbeat interval), so jobs fall back to GitHub-hosted runners
until the pods schedule. Grant the heartbeat's service account read access to
pods, and document both this and the autoscaler's unschedulable handling.
@Mearman
Mearman marked this pull request as ready for review October 3, 2026 07:27
@Mearman
Mearman merged commit f875714 into main Oct 3, 2026
12 checks passed
@Mearman
Mearman deleted the fix/runner-capacity-signals branch October 3, 2026 07:27
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Oct 3, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
🔒 Security Review ✅ Completed 2026-10-03T07:33:12.513082Z 905b416 Draft marked ready
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@github-actions

github-actions Bot commented Oct 3, 2026

Copy link
Copy Markdown

🎉 This PR is included in version 1.8.0 🎉

The release is available on:

Your semantic-release bot 📦🚀

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant