Stop fighting cloud-init logs that hide why your network failed to come up
Technical Briefing | 10/2/2026
You launch a new node, point your bootc image at it, and wait for the signal. It never comes. You log into the serial console, expecting a smoking gun, but cloud-init has already rotated the logs or, worse, swallowed the failure state into a generic success exit code because it reached the end of its stage. I’ve spent too many hours debugging why a late-running network configuration script didn’t trigger, only to find the culprit was a shell injection error that didn’t even show up in /var/log/cloud-init.log.
Why cloud-init logs are not where you think they are
Most people assume cloud-init.log contains the truth. It is a lie. It’s just a summary. When you are working with immutable images where you can’t just install strace on the fly, you need to be looking at the raw output captured during the datasource fetch and the subsequent stages. If you are using bootc, the lifecycle is even more aggressive about cleaning up transient state. You need to pull the truth out of the individual stage logs before the boot sequence prunes them.
find /var/lib/cloud/instance/ -name 'output*' -exec cat {} +
- Check /var/lib/cloud/instance/scripts for raw bash blobs that actually ran
- Verify /var/lib/cloud/instance/obj.pkl for internal state mismatches
- Look for .err files that cloud-init creates only when specific handlers crash
- Search journald specifically for the cloud-init service unit rather than relying on file tails
If you are iterating on custom images, stop trying to fix the config files on a running system. If it fails, bake the fix, re-image, and push. But before you rebuild, inspect the instance-data.json file generated in the instance directory. It holds the actual metadata the node ingested, which is usually where the mismatch lives. Once you know which module choked, you can inject a drop-in override for that specific stage in your bootc build, keeping your rootfs clean and reproducible.
Next time a node hangs during provisioning, don’t waste time checking global system logs. Go straight to the lib cloud directory, dump those output files, and see exactly which command returned non-zero before the systemd target even thought about starting your application.
