Lab / lab

How far a single-node K3s cluster goes

Running the full inference stack on one node to find where it actually falls over.

01Hypothesis

A single-node K3s cluster can carry a realistic inference workload plus its monitoring stack, and the binding constraint will be memory rather than CPU.

02Method

Deployed two FastAPI inference services, Prometheus and Grafana on a single node. Increased concurrent request load until something failed, then recorded what failed first.

03Findings

Memory, as expected, and specifically Prometheus retention rather than the inference services. The scrape volume from two services with model-level metrics filled the configured retention faster than anticipated, and the eviction took the monitoring down before the workload it was monitoring.

Which is a genuinely bad failure ordering: you lose observability at exactly the moment the cluster is under stress. Fixed by giving Prometheus a resource request that reserves what it needs and reducing retention, on the reasoning that short-retention working monitoring beats long-retention absent monitoring.

Lesson generalised: give the monitoring stack a guaranteed allocation, not a best-effort one.

04Record