Skip to content

Record per-job peak memory and route jobs to learned size tiers #40

Description

@Mearman

Runner pods are sized by one request (2 CPU, 7 Gi) regardless of the job, and slot count is limited by that request, not by what jobs use. Observed use across running pods was a fraction of the request, so many more jobs would fit if each pod were sized to its job. Nothing records what a job actually needs, so the request cannot be lowered safely, and a smaller default would kill the jobs that do need it.

What is known. ARC creates the runner pod before a job is assigned to it, so a pod cannot be resized for the job it receives. Sizing therefore means several scale sets (small, default, large) with the job choosing one through its runs-on label; the role's scale-set profiles already take a sizing and a runs_on_label. The controller already knows which job each pod got: EphemeralRunner status carries jobRepositoryName, jobWorkflowRef, jobDisplayName and workflowRunId.

What is needed.

  1. Measure a job's peak memory when it ends. Polling kubectl top (the autoscaler's current signal) misses spikes, so the better source is the cgroup's memory.peak read from a job-completed hook (ACTIONS_RUNNER_HOOK_JOB_COMPLETED) inside the runner pod. Unverified: that the runner image's cgroup exposes memory.peak and the hook runs with the permissions to read it.
  2. Store it keyed by repository, workflow and job name. The runner pod must not hold a gist or cluster write credential, so the hook should emit the figure somewhere the platform already reads (the pod's log or the EphemeralRunner), and a platform component should aggregate it. That storage choice is the open design question.
  3. Route a job to a tier. runner-fallback-action already picks the runner label per job, so it is the natural place to read the learned size and choose the label.
  4. Handle under-sizing: a job killed for memory is retried one tier up and that tier is recorded, so a wrong guess costs one rerun.

Start with step 1 alone, and read the data for a representative period before building steps 2 to 4. Lowering the default request is a separate, smaller change that this data would justify.

Related: the pool is currently bounded by memory requests, see #30 for why replicated storage was deferred.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions