Pooling the nodes' disks into replicated storage (Longhorn, Rook/Ceph, OpenEBS replicated volumes or similar) came up after the m3 node sat tainted for disk pressure for about 3.5 days and left runner pods Pending. We decided against it for now. This issue records why and what would change that.
The disk that filled was the Docker VM's disk on the M3, shared with unrelated Docker builds outside the cluster, so the kubelet had nothing to evict ("no pods active to evict"). Pooled storage would not have seen that disk. Apart from etcd, which is already replicated across the three server nodes, nothing in the cluster is stateful, and runner pods are ephemeral. The nodes reach each other over Tailscale and the node logs show peer contacts via DERP relays, which would make synchronous replication slow. The node disks are small (roughly 30, 60 and 118 GB, shared with other workloads), so replication would cost a large share of them.
Revisit if a workload that needs persistent volumes lands in the cluster (an in-cluster cache or registry mirror, say), if the node count grows past three, or if the nodes get dedicated disks and direct tailnet paths. Before choosing a system, measure throughput and latency between the nodes on direct and relayed paths, and per-node free disk once the node disks are isolated from other Docker workloads.
Until then, ephemeral-storage requests and limits on runner and platform pods, plus an off-cluster shared cache (Turbo, pnpm and buildkit against an S3-compatible store), cover the same risks. The disk pressure fixes are in #28.
Pooling the nodes' disks into replicated storage (Longhorn, Rook/Ceph, OpenEBS replicated volumes or similar) came up after the m3 node sat tainted for disk pressure for about 3.5 days and left runner pods Pending. We decided against it for now. This issue records why and what would change that.
The disk that filled was the Docker VM's disk on the M3, shared with unrelated Docker builds outside the cluster, so the kubelet had nothing to evict ("no pods active to evict"). Pooled storage would not have seen that disk. Apart from etcd, which is already replicated across the three server nodes, nothing in the cluster is stateful, and runner pods are ephemeral. The nodes reach each other over Tailscale and the node logs show peer contacts via DERP relays, which would make synchronous replication slow. The node disks are small (roughly 30, 60 and 118 GB, shared with other workloads), so replication would cost a large share of them.
Revisit if a workload that needs persistent volumes lands in the cluster (an in-cluster cache or registry mirror, say), if the node count grows past three, or if the nodes get dedicated disks and direct tailnet paths. Before choosing a system, measure throughput and latency between the nodes on direct and relayed paths, and per-node free disk once the node disks are isolated from other Docker workloads.
Until then, ephemeral-storage requests and limits on runner and platform pods, plus an off-cluster shared cache (Turbo, pnpm and buildkit against an S3-compatible store), cover the same risks. The disk pressure fixes are in #28.