Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
5 changes: 3 additions & 2 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -105,7 +105,7 @@ Provisioning the join key: create a reusable, pre-authorized Tailscale auth key

### Heartbeat

ARC's autoscaling means there is normally no pre-existing "online runner" to check the way the previous design's fallback logic did (querying `/orgs/{org}/actions/runners` for an online match); with `minRunners: 0`, pods exist only while a job is running. To still support falling back to `ubuntu-latest` if the whole fleet goes down, `heartbeat` (an in-cluster Deployment - see `roles/github_runner_arc/tasks/install_platform.yml`) checks every `HEARTBEAT_INTERVAL_SECONDS` (default 180s) that a k3s node and the ARC controller are healthy, and if so, refreshes a single secret gist with a unix timestamp set `HEARTBEAT_WINDOW_SECONDS` (default 600s) into the future. [`exadev/runner-fallback-action`](https://github.com/ExaDev/runner-fallback-action) reads that gist over the unauthenticated GitHub API and routes to `exadev-runners` while the timestamp is still fresh, or to `ubuntu-latest` otherwise. See that repo's `docs/spec.md` for the algorithm.
ARC's autoscaling means there is normally no pre-existing "online runner" to check the way the previous design's fallback logic did (querying `/orgs/{org}/actions/runners` for an online match); with `minRunners: 0`, pods exist only while a job is running. To still support falling back to `ubuntu-latest` if the whole fleet goes down, `heartbeat` (an in-cluster Deployment - see `roles/github_runner_arc/tasks/install_platform.yml`) checks every `HEARTBEAT_INTERVAL_SECONDS` (default 180s) that a k3s node and the ARC controller are healthy and that no runner pod has been unschedulable for `HEARTBEAT_UNSCHEDULABLE_AFTER_SECONDS` (default: one interval), since a fleet with nowhere to place a runner only queues the work sent to it, and if so, refreshes a single secret gist with a unix timestamp set `HEARTBEAT_WINDOW_SECONDS` (default 600s) into the future. [`exadev/runner-fallback-action`](https://github.com/ExaDev/runner-fallback-action) reads that gist over the unauthenticated GitHub API and routes to `exadev-runners` while the timestamp is still fresh, or to `ubuntu-latest` otherwise. See that repo's `docs/spec.md` for the algorithm.

Runs as a genuinely unpinned Kubernetes Deployment (`replicas: 1`, no `nodeSelector`), the same pattern ARC's own controller pods already use, not a host-pinned Docker Compose service — Kubernetes' own scheduler places and reschedules it on any server host with no config to move if one goes down. `kubectl` needs no `KUBECONFIG`: running as a pod with a mounted ServiceAccount token, in-cluster config is auto-detected. RBAC is scoped to exactly what it reads: cluster-wide `nodes` get/list, `deployments` get/list in the `actions-runner-controller` namespace, and `get` on the autoscaler's own status ConfigMap by name (see Autoscaler below) — no wildcards.

Expand All @@ -128,7 +128,8 @@ Each poll (`AUTOSCALER_POLL_SECONDS`, default 45s):
- Reads cluster-wide memory-availability pressure via `kubectl top nodes`, taking the **worst-case (minimum)** available-memory percentage across all nodes, not an average — a pool average can look healthy while the specific node a new runner pod would actually land on is not, consistent with the algorithm's own bias toward lowering eagerly (lowering never disrupts in-flight jobs). This is a genuine improvement over the pod's own earlier host-pinned design, which could only ever see one fixed machine's pressure via `/proc/meminfo`, not the whole pool's. **No swap-pressure signal**: metrics-server's API has no swap field at all, and the only alternative (the kubelet's own Summary API) needs a materially broader RBAC grant for a signal most kubelet versions don't even surface unless an off-by-default feature gate is on — dropping swap detection is a deliberate, documented trade-off once the pod is genuinely unpinned to any node, not an oversight; memory-availability pressure remains the dominant signal (the incident the values file's own comments were tuned from).
- Raises the combined `maxRunners` by exactly +1 per cycle (never straight to a computed target) once headroom has covered a full pod's hard limit for `AUTOSCALER_RAISE_CONFIRM_POLLS` (default 2) consecutive polls, capped at `AUTOSCALER_MAX_CEILING` (`github_runner_arc_autoscaler_max_ceiling`, or, when exactly one pooled profile sets `sizing`, that profile's own derived ceiling): the bin-packing-safe ceiling across the pool at the profile's pod size, not an arbitrary cap.
- Lowers the combined `maxRunners` immediately, with no delay or averaging, the moment headroom drops below a pod's hard limit or real pressure is detected — lowering never disrupts in-flight jobs, since ARC only gates new claims. Never patches any pooled profile below its own currently-running count.
- Fails safe on any measurement error (`kubectl` unreachable, `kubectl top nodes`/`kubectl top pod` failing): lowers the combined total toward `AUTOSCALER_FLOOR` (`github_runner_arc_autoscaler_floor`) immediately rather than skipping the cycle silently. This combined floor need not equal any one profile's own values-file `maxRunners`, since each profile's Helm upgrade reverts to its own static floor independently of how the pool's combined floor is set.
- Treats any unschedulable runner pod (`PodScheduled=False`, reason `Unschedulable`) like pressure: the memory budget sums every node's capacity, so it cannot see a node that is tainted for disk pressure or a request that fits no single node. It lowers the combined `maxRunners` to the running count and never raises while one exists.
- Fails safe on any memory measurement error (`kubectl top nodes`/`kubectl top pod` failing or unparseable): lowers the combined total toward `AUTOSCALER_FLOOR` (`github_runner_arc_autoscaler_floor`) immediately rather than skipping the cycle silently. A failure reading the scale sets' own state or their pods exits non-zero instead, since fail-safe needs that state to lower anything and the API server is then unreachable anyway. This combined floor need not equal any one profile's own values-file `maxRunners`, since each profile's Helm upgrade reverts to its own static floor independently of how the pool's combined floor is set.

Ships with `AUTOSCALER_DRY_RUN=true` by default: it computes and logs the target and writes its own status, but never patches. Generate real concurrent traffic to watch it react: `gh workflow run test-autoscaler.yml`, then watch `kubectl logs -n github-runner-platform deploy/autoscaler -f` and `kubectl get configmap autoscaler-status -n github-runner-platform -o jsonpath='{.data.status\.json}'`. Only set `AUTOSCALER_DRY_RUN=false` after watching real dry-run output across genuine CI traffic. Status and its raise-confirm counter live in a shared Kubernetes ConfigMap (`autoscaler-status`, pre-created empty so neither pod's own ServiceAccount ever needs `create` RBAC on it, only `get`/`patch`) rather than a bind-mounted directory, since the two pods can land on different nodes now — `scripts/heartbeat.sh` reads that same ConfigMap back out and republishes it as a second file in the heartbeat gist, reusing its existing gist-write credential rather than giving the autoscaler its own.

Expand Down
29 changes: 29 additions & 0 deletions roles/github_runner_arc/tasks/install_platform.yml
Original file line number Diff line number Diff line change
Expand Up @@ -50,6 +50,35 @@
name: github-runner-heartbeat-nodes-reader
apiGroup: rbac.authorization.k8s.io

# heartbeat: list runner pods in every scale-set namespace, to spot ones the scheduler cannot place (cluster-wide, since each profile has its own namespace).
- name: "Platform: heartbeat RBAC - read runner pods"
kubernetes.core.k8s:
definition:
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
name: github-runner-heartbeat-pods-reader
rules:
- apiGroups: [""]
resources: ["pods"]
verbs: ["get", "list"]

- name: "Platform: heartbeat RBAC - bind pods-reader"
kubernetes.core.k8s:
definition:
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRoleBinding
metadata:
name: github-runner-heartbeat-pods-reader
subjects:
- kind: ServiceAccount
name: heartbeat
namespace: "{{ github_runner_arc_platform_namespace }}"
roleRef:
kind: ClusterRole
name: github-runner-heartbeat-pods-reader
apiGroup: rbac.authorization.k8s.io

# heartbeat: read the ARC controller's own Deployment readiness, in its own namespace.
- name: "Platform: heartbeat RBAC - read the ARC controller Deployment"
kubernetes.core.k8s:
Expand Down
41 changes: 33 additions & 8 deletions scripts/autoscaler.sh
Original file line number Diff line number Diff line change
Expand Up @@ -135,6 +135,12 @@ raise_by_one() {
NEW_MAX[best]=$(( NEW_MAX[best] + 1 ))
}

# For a failure before every target's current state is known: fail_safe patches from that state, so with none or part of it, the only honest outcome is to stop loudly (the API server is unreachable, so nothing could be patched anyway).
die() {
echo "ERROR: $1" >&2
exit 1
}

fail_safe() {
echo "FAIL-SAFE: $1 — lowering the pooled total toward the combined floor (${FLOOR})" >&2
configmap_set "raise-confirm-count" "0"
Expand All @@ -146,24 +152,41 @@ fail_safe() {
# ---- Gather signals ---------------------------------------------------------

declare -a R MAX
USAGE_MIB=0
WH_MIB=0 # the largest pooled target's own pod memory limit, used as the headroom threshold: raising or lowering must leave room for whichever pooled target's next pod would be the largest.
UNSCHEDULABLE_TOTAL=0

# Phase 1: every target's own running count, maxRunners and unschedulable runner pods. fail_safe needs all of these to lower anything, so a failure here dies instead.
for i in "${!TARGETS[@]}"; do
namespace=${NAMESPACES[$i]}
release=${RELEASES[$i]}
target=${TARGETS[$i]}

# R: the authoritative currently-running count, read from the AutoscalingRunnerSet CRD's own status (not a pod-label guess). The role gives every scale-set profile its own namespace, so each target namespace holds exactly this one scale set and no scale-set-specific label selector is needed for anything below.
r="$(kubectl get autoscalingrunnerset "$release" -n "$namespace" -o jsonpath='{.status.currentRunners}' 2>/dev/null)" \
|| fail_safe "could not read ${target}'s status.currentRunners"
case "$r" in ''|*[!0-9]*) fail_safe "${target}'s status.currentRunners was not a plain integer ('$r')" ;; esac
|| die "could not read ${target}'s status.currentRunners"
case "$r" in ''|*[!0-9]*) die "${target}'s status.currentRunners was not a plain integer ('$r')" ;; esac
R[i]=$r

max="$(kubectl get autoscalingrunnerset "$release" -n "$namespace" -o jsonpath='{.spec.maxRunners}' 2>/dev/null)" \
|| fail_safe "could not read ${target}'s current spec.maxRunners"
case "$max" in ''|*[!0-9]*) fail_safe "${target}'s spec.maxRunners was not a plain integer ('$max')" ;; esac
|| die "could not read ${target}'s current spec.maxRunners"
case "$max" in ''|*[!0-9]*) die "${target}'s spec.maxRunners was not a plain integer ('$max')" ;; esac
MAX[i]=$max

# Runner pods the scheduler has found no node for. The memory budget below sums every node's capacity, so it cannot see that a node is cordoned or tainted (disk pressure, for one) or that a pod's request fits no single node; a pod stuck Unschedulable is the direct evidence that the budgeted capacity is not real.
unschedulable="$(kubectl get pods -n "$namespace" -l actions-ephemeral-runner=True -o json 2>/dev/null \
| jq '[.items[] | select(any(.status.conditions[]?; .type == "PodScheduled" and .status == "False" and .reason == "Unschedulable"))] | length')" \
|| die "could not list ${target}'s unschedulable runner pods"
case "$unschedulable" in ''|*[!0-9]*) die "${target}'s unschedulable runner pod count was not a plain integer ('$unschedulable')" ;; esac
UNSCHEDULABLE_TOTAL=$(( UNSCHEDULABLE_TOTAL + unschedulable ))
done

# Phase 2: memory measurements. Any failure fails safe, which can now lower every target because phase 1 read them all.
USAGE_MIB=0
WH_MIB=0 # the largest pooled target's own pod memory limit, used as the headroom threshold: raising or lowering must leave room for whichever pooled target's next pod would be the largest.
for i in "${!TARGETS[@]}"; do
namespace=${NAMESPACES[$i]}
release=${RELEASES[$i]}
target=${TARGETS[$i]}

# Wh: the pod's own hard memory limit, read from the live deployed spec, not re-parsed from values/*.yaml, so this always matches what is actually running even after a manual --set override.
wh_raw="$(kubectl get autoscalingrunnerset "$release" -n "$namespace" \
-o jsonpath='{.spec.template.spec.containers[0].resources.limits.memory}' 2>/dev/null)" \
Expand Down Expand Up @@ -212,17 +235,19 @@ fi
USABLE_BUDGET_MIB=$(( USABLE_BUDGET_GI * 1024 ))
HEADROOM_MIB=$(( USABLE_BUDGET_MIB - USAGE_MIB ))

echo "R_total=${R_TOTAL} maxRunners_total=${MAX_TOTAL} Wh=${WH_MIB}MiB usage=${USAGE_MIB}MiB headroom=${HEADROOM_MIB}MiB mem_available=${mem_available_pct}% pressure=${pressure} targets=${TARGETS[*]}"
echo "R_total=${R_TOTAL} maxRunners_total=${MAX_TOTAL} Wh=${WH_MIB}MiB usage=${USAGE_MIB}MiB headroom=${HEADROOM_MIB}MiB mem_available=${mem_available_pct}% pressure=${pressure} unschedulable=${UNSCHEDULABLE_TOTAL} targets=${TARGETS[*]}"

if [ "$pressure" = "true" ] || [ "$HEADROOM_MIB" -lt "$WH_MIB" ]; then
if [ "$pressure" = "true" ] || [ "$HEADROOM_MIB" -lt "$WH_MIB" ] || [ "$UNSCHEDULABLE_TOTAL" -gt 0 ]; then
# Lower immediately, no delay or averaging — lowering never disrupts in-flight jobs (ARC only gates new claims), so there is no cost to being trigger-happy in this direction. Can drop below FLOOR (even to 0, spread across targets by lower_to) if pressure is severe enough.
configmap_set "raise-confirm-count" "0"
target_total=$FLOOR
[ "$pressure" = "true" ] && target_total=1
[ "$UNSCHEDULABLE_TOTAL" -gt 0 ] && target_total=0 # clamped to R_TOTAL below: stop asking for runners no node can host, never orphan a running job
[ "$target_total" -lt "$R_TOTAL" ] && target_total=$R_TOTAL
if [ "$target_total" -lt "$MAX_TOTAL" ]; then
reason="lower: headroom=${HEADROOM_MIB}MiB < Wh=${WH_MIB}MiB"
[ "$pressure" = "true" ] && reason="lower: host pressure (mem_available=${mem_available_pct}%)"
[ "$UNSCHEDULABLE_TOTAL" -gt 0 ] && reason="lower: ${UNSCHEDULABLE_TOTAL} runner pod(s) unschedulable, so the budgeted capacity is not all schedulable"
lower_to "$target_total"
apply_new_max "$reason" "$HEADROOM_MIB"
else
Expand Down
14 changes: 14 additions & 0 deletions scripts/heartbeat.sh
Original file line number Diff line number Diff line change
Expand Up @@ -8,6 +8,7 @@
# - HEARTBEAT_STATE_NAMESPACE: namespace holding the autoscaler's own status ConfigMap (see scripts/autoscaler.sh)
# - AUTOSCALER_STATUS_CONFIGMAP: name of that ConfigMap
# - HEARTBEAT_CONTROLLER_NAMESPACE: namespace of the ARC controller's Deployment, when it is not the default actions-runner-controller
# - HEARTBEAT_UNSCHEDULABLE_AFTER_SECONDS: how long a runner pod may stay unschedulable before the heartbeat stops refreshing (default: one heartbeat interval)
set -euo pipefail

HEARTBEAT_GH_TOKEN="${HEARTBEAT_GH_TOKEN:?HEARTBEAT_GH_TOKEN must be set (a PAT with gist scope)}"
Expand All @@ -19,6 +20,8 @@ AUTOSCALER_STATUS_GIST_FILE="${AUTOSCALER_STATUS_GIST_FILE:-autoscaler-status.js
# How far in the future to set the timestamp: comfortably longer than the loop interval, so a single missed/slow tick doesn't look like an outage, but short enough that a real outage is detected promptly.
HEARTBEAT_WINDOW_SECONDS="${HEARTBEAT_WINDOW_SECONDS:-600}"
HEARTBEAT_CONTROLLER_NAMESPACE="${HEARTBEAT_CONTROLLER_NAMESPACE:-actions-runner-controller}"
# How long a runner pod may sit unschedulable before the fleet counts as unable to take work: one heartbeat interval (the loop's own, see heartbeat/loop.sh), so the condition has been seen on at least two consecutive ticks and a pod that is merely being placed never trips it.
HEARTBEAT_UNSCHEDULABLE_AFTER_SECONDS="${HEARTBEAT_UNSCHEDULABLE_AFTER_SECONDS:-${HEARTBEAT_INTERVAL_SECONDS:-180}}"

healthy=true

Expand All @@ -31,6 +34,17 @@ if ! kubectl get deployment -n "$HEARTBEAT_CONTROLLER_NAMESPACE" -l app.kubernet
healthy=false
fi

# A fleet with Ready nodes and a running controller can still have nowhere to put a runner (every schedulable node full, the rest cordoned or tainted for disk pressure, say). Jobs sent to it then queue behind the runners already busy, so stop vouching for it and let runner-fallback-action route to GitHub-hosted runners until the pods schedule again.
now="$(date +%s)"
if ! unschedulable="$(kubectl get pods -A -l actions-ephemeral-runner=True -o json 2>/dev/null \
| jq --argjson now "$now" --argjson after "$HEARTBEAT_UNSCHEDULABLE_AFTER_SECONDS" \
'[.items[] | select(any(.status.conditions[]?; .type == "PodScheduled" and .status == "False" and .reason == "Unschedulable" and ($now - (.lastTransitionTime | fromdateiso8601)) >= $after))] | length')"; then
healthy=false
elif [ "$unschedulable" -gt 0 ]; then
echo "${unschedulable} runner pod(s) unschedulable for at least ${HEARTBEAT_UNSCHEDULABLE_AFTER_SECONDS}s" >&2
healthy=false
fi

if [ "$healthy" != "true" ]; then
echo "Unhealthy - not refreshing the heartbeat gist" >&2
exit 1
Expand Down
Loading