Skip to content

Lower the runner pod memory request once job peaks are measured #46

Description

@Mearman

Runner pods are sized at 2 CPU and 7 Gi regardless of the job, which limits how many run at once on nodes with about 8 Gi allocatable, and observed use is a fraction of that. The measured peaks are now collected (the autoscaler's job-peak key) and summarised in the heartbeat gist's job-peak-summary.json: job count, median, 95th percentile and maximum in MiB next to the pod's memory limit.

Deferred because it needs real CI traffic: the measurement only began recently and the figures so far come from a handful of jobs. Revisit when the summary's job count covers a representative spread of the repositories that use the fleet (a week of ordinary traffic is a reasonable minimum) and the 95th percentile and maximum are stable between readings.

Then decide from the data, not a formula: if the maximum sits comfortably below the limit, lower the request in the fleet's shared runner values file and watch for out-of-memory kills; if a few jobs sit near the limit while most sit far below, that is the case for size tiers (see issue 40), not for a lower default. A lower request also raises how many runners fit, so the autoscaler's ceiling and the static maxRunners in the values file need revisiting with it.

Related: #40

A one-time scheduled run (https://claude.ai/code/routines/trig_01Q4AQebKvmVHtCMSbyUVXiq) posts the gist figures on this issue on 2026-10-10 at 16:00 UTC. It posts figures only.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions