Runner pods are sized by one request (2 CPU, 7 Gi) regardless of the job, and slot count is limited by that request, not by what jobs use. Observed use across running pods was a fraction of the request, so many more jobs would fit if each pod were sized to its job. Nothing records what a job actually needs, so the request cannot be lowered safely, and a smaller default would kill the jobs that do need it.
What is known. ARC creates the runner pod before a job is assigned to it, so a pod cannot be resized for the job it receives. Sizing therefore means several scale sets (small, default, large) with the job choosing one through its runs-on label; the role's scale-set profiles already take a sizing and a runs_on_label. The controller already knows which job each pod got: EphemeralRunner status carries jobRepositoryName, jobWorkflowRef, jobDisplayName and workflowRunId.
What is needed.
- Measure a job's peak memory when it ends. Polling kubectl top (the autoscaler's current signal) misses spikes, so the better source is the cgroup's memory.peak read from a job-completed hook (ACTIONS_RUNNER_HOOK_JOB_COMPLETED) inside the runner pod. Unverified: that the runner image's cgroup exposes memory.peak and the hook runs with the permissions to read it.
- Store it keyed by repository, workflow and job name. The runner pod must not hold a gist or cluster write credential, so the hook should emit the figure somewhere the platform already reads (the pod's log or the EphemeralRunner), and a platform component should aggregate it. That storage choice is the open design question.
- Route a job to a tier. runner-fallback-action already picks the runner label per job, so it is the natural place to read the learned size and choose the label.
- Handle under-sizing: a job killed for memory is retried one tier up and that tier is recorded, so a wrong guess costs one rerun.
Start with step 1 alone, and read the data for a representative period before building steps 2 to 4. Lowering the default request is a separate, smaller change that this data would justify.
Related: the pool is currently bounded by memory requests, see #30 for why replicated storage was deferred.
Runner pods are sized by one request (2 CPU, 7 Gi) regardless of the job, and slot count is limited by that request, not by what jobs use. Observed use across running pods was a fraction of the request, so many more jobs would fit if each pod were sized to its job. Nothing records what a job actually needs, so the request cannot be lowered safely, and a smaller default would kill the jobs that do need it.
What is known. ARC creates the runner pod before a job is assigned to it, so a pod cannot be resized for the job it receives. Sizing therefore means several scale sets (small, default, large) with the job choosing one through its runs-on label; the role's scale-set profiles already take a sizing and a runs_on_label. The controller already knows which job each pod got: EphemeralRunner status carries jobRepositoryName, jobWorkflowRef, jobDisplayName and workflowRunId.
What is needed.
Start with step 1 alone, and read the data for a representative period before building steps 2 to 4. Lowering the default request is a separate, smaller change that this data would justify.
Related: the pool is currently bounded by memory requests, see #30 for why replicated storage was deferred.