When your CI runner chokes the host, and `nice` just shrugs
By Saket Jain Published Linux/Unix
When your CI runner chokes the host, and `nice` just shrugs
Technical Briefing | 10/6/2026
CI/CD runners are great until one of ’em decides to be a jerk. You’ve got a perfectly good build server, running a handful of jobs, and then randomly, some builds start taking twice as long, others fail with obscure timeouts, and the whole system just feels slow. You hop on, run `top`, maybe see some high CPU, but it’s transient. The obvious fix, tossing a `nice -n 19` in front of your build command, feels good but often doesn’t actually solve the problem. Been there. This bit me hard on a shared Gitlab Runner host once when a new project’s build started chewing through everything.
`nice` and `ionice` are polite suggestions, not hard boundaries
Look, `nice` has its place. It tells the kernel, “Hey, this process is less important than others, so if you’re feeling generous, give more CPU cycles to the others first.” And `ionice` does something similar for block I/O. The keyword there is “if you’re feeling generous.” When a single process demands all the CPU cores or slams the disk with writes, `nice` might make other processes a bit more responsive, but it won’t stop the greedy process from taking 100% of a core when nothing else is running, or starving others entirely if it’s the only one active in its priority bracket. It’s cooperative multitasking at the process level, which is why it often falls short when you’ve got real resource hogs.
Cgroups: The kernel’s enforcers
If you want actual resource limits – “this process gets no more than X% CPU, or Y MB/s disk write” – then you’ve gotta talk to cgroups. Control groups are a kernel mechanism for organizing processes into hierarchical groups and allocating system resources among them. You define these groups, set hard limits on things like CPU shares, absolute CPU time, memory, or I/O, and then assign processes to them. This is how containers like Docker or Podman get their resource limits, but you can absolutely use them directly for any process, including your CI/CD jobs. It’s how you tell the kernel, “I mean it.”
systemd-run --scope -p CPUQuota=25% -p IOReadBandwidth=/dev/sda:20M -p IOWriteBandwidth=/dev/sda:10M /bin/bash -c "dd if=/dev/zero of=/tmp/testfile bs=1M count=1000 status=progress && stress-ng --cpu 4 --timeout 60s"
That `systemd-run` command is pretty handy for testing or for simple wrapper scripts. It creates a temporary cgroup for that `/bin/bash` command and everything it spawns. We’re telling `systemd` to limit its CPU usage to 25% of one core (or 25% across all cores if it could use more), cap its read bandwidth from `/dev/sda` at 20 megabytes per second, and write bandwidth at 10 megabytes per second. The `dd` and `stress-ng` commands simulate a typical greedy CI job. You could swap that out for `ansible-playbook` or whatever your build script is. This gives you actual, enforced throttling, not just a gentle nudge. Remember to adjust device paths (`/dev/sda`) to match your actual block devices; `lsblk -f` is your friend here.
- Identify your bottleneck device: For I/O limits, make sure you’re targeting the correct block device (e.g., `/dev/sda`, `/dev/nvme0n1`). Check `lsblk -f` or `df -hT` to find out which is which.
- Balance limits: Don’t under-resource your critical jobs, but don’t over-resource your non-critical ones either. Start with generous limits and tighten them based on observed performance and resource contention. This usually means a bit of trial and error.
- Wrapper scripts: For complex CI pipelines, you’ll likely want a simple shell wrapper that calls `systemd-run` around your actual build command. Or, configure your CI runner’s executor to use cgroups if it supports it directly (like some Kubernetes executors do, or Docker’s resource limits which are just cgroups).
This isn’t about perfectly optimizing every single build, it’s about stability. Preventing one runaway job from bringing your entire runner host to its knees keeps your other builds green and your deploy times consistent. You won’t use these specific `systemd-run` flags every day, but knowing `nice` isn’t the final word on resource management can save you a world of hurt when that inexplicable CI slowdown inevitably hits. And trust me, it will hit.
