Note

Liveness probes versus model loading

An inference container that loads a large model at startup will be killed by an impatient liveness probe, repeatedly, forever.

The symptom is a pod in a restart loop with no useful logs, because it never got far enough to log anything.

The cause is that the liveness probe started before the model finished loading, declared the container unhealthy, and killed it. On restart the same thing happens.

Two fixes, and they are not interchangeable. A startup probe is the correct tool: it gives the container a generous window to become ready, and only after it succeeds does the liveness probe begin. Inflating the liveness probe's initial delay instead works, and it permanently degrades your detection time for genuine failures after startup.

Use a startup probe. Keep liveness tight.