Skip to content

Repository files navigation

forkrun — NUMA-Aware Contention-Free Streaming Parallelization

License: MIT

forkrun is a self-tuning, drop-in replacement for GNU Parallel and xargs -P that accelerates shell-based data preparation by 50×–400× for typical shell builtins (up to ~3300× for external-binary no-op microbenchmarks) on modern CPUs and scales linearly on NUMA architectures.

forkrun achieves:

  • 87–99% CPU utilization across all cores depending on mode and input size (vs ~6% for GNU Parallel) — ~95–99% for sustained default/external modes, ~90% aggregate across 396 mixed benchmarks, lower for sub-second or byte-mode jobs by design
  • Born-local NUMA placement: file ingest measures a median of 0.0% cross-socket chunks (tail up to ~20% on small-chunk and byte-mode runs). Under fast-draining pipe input, 2–13% of chunks may be stolen — by design (an idle node costs more than a remote chunk). Real multi-socket topologies raise the steal threshold with distance (1 + distance/10), so these figures — measured on numa=fake=4, where all distances are 10 — are a worst case. (The end-of-stream drain collapses the threshold to 1 regardless of distance; this is bounded to EOF.)
  • Automatic recovery and retry when a worker unexpectedly dies processing a batch (v3.1.0+)

forkrun is built for high-frequency, low-latency workloads on deep NUMA hardware — a regime where existing tools leave most cores idle due to IPC overhead and cross-socket data migration.


🚀 Quick Start (Installation & Usage)

forkrun is distributed as a single bash file with an embedded, self-extracting compiled C extension. There are no external dependencies (no Perl, no Python).

Download and source it directly:

# Option 1: download and source
wget https://raw.githubusercontent.com/jkool702/forkrun/main/frun.bash
source ./frun.bash

# Option 2: source curl stream
source <(curl -sL https://raw.githubusercontent.com/jkool702/forkrun/main/frun.bash)

(Note: Sourcing the script sets up the required C loadable builtins in your shell environment).

Once sourced, frun acts as a drop-in parallelizer:

frun my_bash_func < inputs.txt             # parallelize custom bash functions natively!
cat file_list | frun -k sed 's/old/new/'   # pipe-based input, ordered output
frun -k -s sort < records.tsv              # stdin-passthrough, ordered output
frun -s -I bash -c 'gzip -c >{ID}.gz' < raw_logs   # stdin-passthrough, unique output names

Auditable Builds: the embedded C extension is compiled and injected by a public GitHub Actions workflow; the git history of the base64 blob traces every byte to a specific CI run of forkrun_ring.c. (Reproducible builds with published checksums are on the roadmap and would upgrade this to cryptographic attestation.)


⚡ Benchmarks (14-core/28-thread i9-7940x, 100M+ lines)

Workload forkrun GNU Parallel Speedup Notes
Max batch external binary (-l 1:-1 /bin/true) 191.4 M lines/s ~58 k lines/s ~3300× Zero-copy vfork fast-path
Default external binary (/bin/true) 86.9 M lines/s ~58 k lines/s ~1500× Bypasses Bash AST entirely
Bash Builtin (:, fully-quoted args) 25.0 M lines/s ~58 k lines/s ~430× forkrun standard array mode
Ordered output (-k, external binary) 86.9 M lines/s 57 k lines/s ~1520× ordering has zero measurable overhead
External printf '%s\n' (I/O heavy) 52.6 M lines/s ~58 k lines/s ~900× formatting + output
-s stdin passthrough (no-op) 1.04 B lines/s 6.05 M lines/s (--pipe) ~172× streaming / splice()
-b 512k byte batches (no-op) 2.51 B lines/s 6.02 M lines/s (--pipe) ~417× kernel-limited

Note on benchmark basis: headline throughputs above are conservative 100M-line measurements. Top modes (-s, -b, external-binary) are limited by a ~30 ms fixed pipeline bring-up cost; ≥1B-line runs remove this fixed cost and show 30–50% higher peak rates. The 50×–400× range quoted in the intro is the typical shell-builtin range; microbenchmark extremes (/bin/true, -l 1:-1) reach ~1500–3300× due to GNU Parallel's per-item Perl fork overhead.

Average CPU utilization across 396 benchmarks (mix-dependent)

  • forkrun: ~90% aggregate (27.1 / 28 cores in steady-state default mode = 97%; 27.6/28 = 98.6% for default-mode sustained runs at ≥1B-line scale (100M-scale measures 24.5–25.5/28 for default -X); -U unsafe runs hit 27.1+/28; -b 512k on 100 MB intentionally ~2.6/28) — No centralized dispatcher; all cores do actual work when work exists.
  • GNU Parallel: 9.6% total (2.68 / 28 cores), 6% useful work (1.68 / 28) — 1 full core used strictly for dispatching work; 1.68 cores doing actual work.

Python Frontend (forkrun 0.16.0 — python/)

10M lines (large), median of 5, same i9-7940X class hardware. Method: python/benchmarks/ (run_all.py --scale large). CPU% = attributable process-tree CPU (self + reaped children) over wall × cores — not system-wide. Full record: python/benchmarks/results/large.md + large.csv.

Workload forkrun Python Baseline Speedup CPU% Notes
Python no-op (map) 182 M lines/s — — 6% claim/ack via ctypes; overhead-bound at 10M, not compute-bound
Python transform (upper) 63 M lines/s 9.8 M/s serial 6.4× 7% bytes(batch.data).upper(); parent collect is the serial bottleneck
Python compute (sum) 53 M lines/s — — 21% sum(memoryview)
C plugin callback 54 M lines/s — — 11% ctypes → C function
Spawn external (cat) 15.1 M lines/s — — 12% subprocess amortized by batching
JSONL ingestion 3.6 M records/s — — 27% json.loads per record; highest CPU (payload-bound)
Filter + transform 31 M lines/s — — — grep-like + upper
Aggregation (sum) 50 M lines/s — — — int parse + sum

Unmeasured cells show "—" (pool baselines are small-scale-only by design; per-row CPU was sampled on headline rows). No-op/upper/sum hold or improve from 1M (102→182M), i.e. fixed bring-up amortizes.

Scenario Input Output Peak RSS Notes
No output (discard) 1→8MB 0 +0MB perfectly flat, both scales
map (collect-all) 1→8MB 1→8MB output-sized v0.5 design
stream, slow consumer 10MB 50MB ~40–175MB peak bounded by window, not stream; spread across runs under investigation (see large.md)
stream, 5× amplification 2MB 10MB bounded backpressure active
Metric Value Notes
stream() vs map() 1.8× faster drain overlaps produce (no collect); 1.4× at 1M
ordered vs unordered 1.34× reassembly cost grows with batch count (1.01× at 1M)
First yield latency ~7ms on a 0.5s job

New to the Python frontend? Start at python/docs/QUICKSTART.md — full guides (API, modes, fault tolerance, NUMA, streaming, examples, plugins, performance, troubleshooting, migration) live in python/docs/.

What These Benchmarks Do NOT Measure

  • NUMA multi-node scaling — single-node only; NUMA is Stage 5 P5.
  • TB-scale streaming — v0 materializes input; streaming ingest is v1.
  • aarch64 — x86_64 only; ARM legs are run manually on hardware.
  • GPU workloads — workers are CPU-only by design; GPU work belongs in the parent.
  • Sub-100k-line jobs — fixed ~30ms bring-up dominates; use serial Python.
  • Spawn vs bash -X — Python subprocess (~1–5ms/batch) vs posix_spawnp (~10µs); bash wins for external binaries by design.
  • Cold I/O — inputs are page-cached/tmpfs; cold disk adds a floor both sides.
  • Pool baselines at scale — multiprocessing.Pool rows are small-scale only (per-line pickling would blow the budget proving nothing new).

🧠 How It Works: The Physics of forkrun

Traditional tools like GNU Parallel use heavy regex parsing and IPC dispatch loops that bottleneck multi-socket servers. forkrun operates completely differently. The pipeline has four stages, each designed to preserve physical locality:

  1. Ingest (Born-Local NUMA): Data is splice()'d from stdin into a shared memfd. This is PFS-friendly (avoids Lustre/NFS metadata storms). On multi-socket systems, set_mempolicy(MPOL_BIND) places each chunk's pages on a target NUMA node before any worker touches them. This placement is driven by real-time backpressure from the per-node indexers, making NUMA distribution completely self-load-balancing.
  2. Index: Per-node indexers (pinned to their socket) find record boundaries using AVX2/NEON SIMD scanning at memory bandwidth. They dynamically batch based on runtime conditions, then publish offset markers into a per-node lock-free ring buffer.
  3. Claim (contention-free (no userspace locks or CAS retry loops on the fast path — two amortized atomic RMWs per batch (read_idx + total_lines_consumed), sharded per NUMA node)): Workers claim batches via a single atomic_fetch_add — no CAS retry loops, no locks, no contention. If a worker process crashes, its transaction is safely rolled back and deposited into an escrow pipe for idle workers to steal.
  4. Reclaim: A background fallow thread punches holes behind completed work via fallocate(PUNCH_HOLE), bounding memory usage without breaking the offset coordinate system.

Adaptive tuning is fully automatic. A Pre-Flight AVX2/NEON SIMD popcount computes the globally optimal batch size during fork latency, instantly entering PID steady-state. If a worker spawns before the scan completes, a geometric fallback converges in O(log L) steps. Either way the worker fast-path is a single atomic_fetch_add with no user -n or -j configuration required.


🛠 Requirements & Dependencies

forkrun is designed to run anywhere with zero friction:

  • Required: Bash ≥ 4.4 (mapfile -d needs 4.4; Bash 5.1+ highly recommended for array performance), Linux Kernel ≥ 3.17 (for memfd), GNU coreutils (sed -z, base64 -w 0, truncate --size= are GNU-only — no busybox support). Kernels ≥ 4.5 additionally enable the copy_file_range fast path; older kernels automatically fall back to sendfile/read-write with no functional difference.

Supported bash versions (v3.6.0 verification status — bootstrap = source + load + ring_version; suites = 96 + 264):

Two-axis reality (read both): the bash axis below was verified on new-glibc iron. Independently, the shipped x86-64 loadables carry no GLIBC_ABI_GNU2_TLS requirement (D-TLS: rebuilt with -mtls-dialect=gnu, dropping the gcc-16 TLSDESC codegen; verified zero references, max GLIBC_2.38 across all shipped blobs). So glibc ≥2.38 loads (Ubuntu 24.04's 2.39 OK); Debian 12 (2.36) and RHEL ≤9 remain below the floor. Non-x86 blobs never carried the requirement. Pure-shell parsing works everywhere regardless, but that is not a usable state without the engine.

Bash Bootstrap + smoke ×10 Full suites Status (on new glibc)
4.4 (RHEL 8) ✅ 10/10 incl. round-trips (D-SEGFIX, extracted binary) — (suites run on 5.2/5.3 only) ✅ usable; RHEL8 glibc (2.28) still below the 2.38 floor
5.0 ✅ 10/10 incl. round-trips (D-SEGFIX, extracted binary) — (suites run on 5.2/5.3 only) ✅ usable
5.1 ✅ 10/10 incl. round-trips (D-SEGFIX, source-built 5.1.0 + extracted 5.1.16, byte-exact) — (suites run on 5.2/5.3 only) ✅ usable
5.2 (Ubuntu 24.04, Debian 12) ✅ 10/10 incl. round-trips ✅ 92 + 264, zero failures fully verified
5.3 ✅ 10/10 incl. round-trips ✅ 92 + 264, zero failures fully verified

Effective floor (W-REL5-D): bash ≥4.4. D-SEGFIX landed (worker segfault root-caused to the engine's ARRAY-struct walk + fixed via pair-list flattening; round-trips byte-exact on 4.4/5.0/5.1) and ships in the v3.6.0 blob cycle. glibc ≥2.38 (D-TLS rebuild, this cycle).


🏛️ Legacy Version (v2)

With the release of v3.0.0, forkrun has transitioned to a high-performance C-ring architecture (frun.bash). The older v2, pure-Bash coproc-based version (forkrun.bash) remains available in the legacy/ directory. While v3 (frun.bash) is highly recommended for all modern workloads, v2 (forkrun.bash) remains as an alternate fully-functional high-performance bash stream parallelizer. forkrun v1 is not recommended for use.


🛣 Roadmap

forkrun features robust intra-node fault tolerance and preemption recovery (automatically trapping worker failures and Slurm signals to generate exactly-once checkpoints).

Priorities for the development roadmap include:

  • Cluster-level multi-node resume support across distributed compute fabrics.
  • Deeper integration with facility workload managers for dynamic resource elasticity.

(If forkrun is saving your institution compute-hours, please consider sponsoring its development to accelerate these features!)

About

NUMA-Aware Contention-Free Dynamically-Auto-Tuning Bash-Native Streaming Parallelization Engine

Topics

Resources

Security policy

Stars

362 stars

Watchers

2 watching

Forks

Releases

Sponsor this project

Used by

Contributors

Languages