ProductionMLOps platform2025

Inference platform on K3s

Containerised ML training and inference workflows served from K3s, with monitoring that treats a model service as an ordinary production service.

Why it existsModel services fail in ways that HTTP 200 does not capture. The platform had to make that visible.

Overview

Fellowship work at Fusemachines building containerised ML pipelines and serving inference from K3s, with the monitoring and CI/CD needed to make model deployments repeatable.

The framing that shaped the work: an inference service is a production service with an additional failure mode. It can be healthy, fast, and wrong. Ordinary service monitoring catches the first two.

Problem

Model work tends to produce environment drift. Training happens on one machine with one set of library versions, serving happens on another, and the discrepancy surfaces as a prediction difference that nobody can attribute.

Alongside that, deployment was manual enough to be inconsistent, and there was no reliable answer to which model version an endpoint was serving. Reliability signals were thin: request-level metrics existed, model-level ones did not.

Approach

Containerise the whole path, put serving on K3s, and instrument the model service as a first-class production citizen.

Training and inference share a base image, so library versions are identical by construction rather than by convention. Inference is a FastAPI service in a container, deployed to K3s as a versioned deployment with the model version recorded as a label and exposed as a metric.

Monitoring covers both layers. Request rate, latency and error rate for the service. Prediction distribution, input feature ranges and inference duration for the model. A shift in prediction distribution without a corresponding change in traffic is the signal that catches silently wrong models.

Architecture

Serving

FastAPI inference services containerised and deployed to K3s. K3s rather than full Kubernetes because the footprint suited the environment and the operational surface was smaller for a small team.

Images

A shared base image pinning the Python runtime and the ML library versions. Training and inference images extend it. This is the mechanism that eliminated train-serve environment drift.

Pipelines

Containerised training pipelines invoked as jobs, writing model artefacts to object storage with a version identifier. Experiment tracking records parameters, metrics and the artefact location per run.

Delivery

GitHub Actions builds and publishes images on merge, then updates the K3s deployment to the new image digest. Model version travels as a deployment label and a service metric.

Observability

Prometheus scrapes both service metrics and model metrics from the inference services. Grafana dashboards pair them, so latency and prediction distribution are visible on one screen. Logs are structured and carry the model version on every inference line.

Technology

K3s for the cluster. Docker for images. FastAPI for inference services. Prometheus and Grafana for metrics and dashboards. GitHub Actions for CI/CD. Object storage for model artefacts. Helm for packaging the deployment manifests so an environment difference is a values file.

Implementation

Model version is exposed as a Prometheus label on inference metrics. This one detail did more for debuggability than the rest of the instrumentation combined, because it makes "did this change with the deployment" a question you can answer on a graph instead of by correlating timestamps.

Experiment tracking captures parameters, metrics and the artefact URI for every training run, including failed ones. Failed runs are the ones you want later, and they are the ones that get discarded by default.

Structured JSON logging with the model version and a request identifier on every line. When a prediction is disputed, the path from a user report to the exact inference is a log query rather than an investigation.

Rollback is a deployment rollback to a previous image digest, and because the model artefact is baked or pinned per image, rolling back the service rolls back the model with it. No separate model rollback procedure to get wrong under pressure.

Challenges

Silently wrong models. The failure mode that motivated the model-level metrics. A service returning 200s at normal latency with a prediction distribution that has drifted is invisible to standard monitoring. Pairing the two metric layers on one dashboard is what made it visible.

K3s is small, not simple. The reduced footprint removed operational surface but not Kubernetes concepts. Resource requests and limits still needed tuning, and an inference container that occasionally allocated a large batch was the source of the only eviction problems worth remembering.

Image size versus build time. ML images are large. Layer ordering matters more than usual: pinning dependencies in an early layer that changes rarely turned most builds into a small final layer instead of a full rebuild.

Decisions

K3s over managed Kubernetes. For the size of the workload and the team, the operational simplicity was worth more than managed control-plane features that were not going to be used.

Shared base image for training and inference. The single most effective change against train-serve skew, and it costs one Dockerfile.

Model version as a metric label. Small instrumentation decision, disproportionate payoff in incident response.

Rollback through the image, not through a model switch. One rollback procedure instead of two, and no possibility of the service and model versions disagreeing.

Results

Inference services run on K3s with consistent environments between training and serving, deployments are driven by CI on merge, and model versions are traceable from a metric or a log line to the training run that produced them.

The monitoring work is what I would point at. Pairing service metrics with prediction-distribution metrics on a single dashboard changed model incidents from a research exercise into an ordinary operational one.

Stack

  1. Serving

    • FastAPI inference services
    • K3s deployments
    • Helm-packaged manifests
    • Digest-pinned images
  2. Environment

    • Shared base image
    • Pinned runtime and ML libraries
    • Separate training and inference layers
  3. Lifecycle

    • Containerised training jobs
    • Experiment tracking incl. failed runs
    • Versioned artefacts in object storage
  4. Observability

    • Prometheus service metrics
    • Prediction-distribution metrics
    • Grafana paired dashboards
    • Structured logs with model version

Measurements

Cluster
K3sChosen for operational surface, not scale
Serving framework
FastAPIContainerised, versioned per deployment
Rollback unit
Image digestService and model roll back together