Machine Learning

A model can be healthy, fast and wrong

Standard service monitoring tells you an inference endpoint is up and responsive. It cannot tell you the predictions stopped making sense three days ago.

An inference service is a production service with one extra failure mode, and it is the failure mode that none of your existing monitoring covers.

Consider the dashboard for an ordinary web service. Request rate, error rate, latency percentiles, saturation. If those four are healthy, the service is healthy, and that inference is usually sound. Now put a model behind the same endpoint. Request rate is normal. Error rate is zero. Latency is at its usual p99. Every one of those signals is green, and the model has been returning nonsense since a data pipeline changed on Tuesday.

Nothing in the standard toolkit is looking at the one thing that broke.

The missing layer

The gap is that service metrics describe the transport and say nothing about the payload. A 200 response with a well-formed JSON body containing a confident, wrong number is indistinguishable from a correct one at the HTTP layer.

So the inference services at Fusemachines exported a second layer of metrics alongside the usual ones:

  • Prediction distribution. For a classifier, the rate of each predicted class. For a regressor, a histogram of outputs.
  • Input feature ranges. Enough summary statistics to notice when the shape of incoming data moves.
  • Inference duration separated from request duration, so model slowness and serving slowness are distinguishable.
  • Volume of inputs rejected by validation, which is often the earliest signal of an upstream change.

None of this is exotic. It is a handful of Prometheus counters and histograms in the same service that already exports request metrics.

Pair them on one dashboard

Exporting model metrics is necessary and not sufficient. The thing that made them useful was putting them on the same Grafana dashboard as the service metrics, on a shared time axis.

A prediction distribution that shifts is ambiguous on its own. Traffic mix changes, and the model responds to it correctly. But a prediction distribution that shifts while request rate, client mix and input volume all hold steady is a much sharper signal. Something changed on the inside.

You cannot see that correlation on two dashboards owned by two teams. You see it immediately on one row of one dashboard.

Model version as a metric label

The smallest change in this whole piece of work, and the one with the highest return: every model metric carries the model version as a label.

This turns the most common incident question from an investigation into a glance. "Did this start when we deployed?" stops requiring you to line up a deployment timestamp against a graph and start correlating by eye. You split the series by version and look. Either the new version is behaving differently or it is not.

It costs one label. I have not yet found a reason to leave it out.

Structured logs carry the version too

Metrics tell you something changed across a population. They cannot tell you what happened to one specific prediction that a user is complaining about.

Every inference log line was JSON and carried the model version and a request identifier. The path from a user report to the exact inference that produced it became a log query. Without that, the same investigation involves guessing at timestamps and hoping the volume is low enough to eyeball.

Roll the model back with the service

Two rollback procedures is one too many, and the second one always gets performed incorrectly at two in the morning.

Because the model artefact is pinned per image, rolling the deployment back to a previous image digest rolls the model back with it. There is no separate model-switching step and no possibility of the service version and model version disagreeing. One procedure, and it is the same procedure the team already knew.

The honest limitation

Distribution monitoring detects change. It does not detect being wrong.

A model that was already miscalibrated at deployment has a stable, wrong prediction distribution, and everything described here will report it as healthy forever. Catching that requires evaluation against labelled data, which is a different discipline with different infrastructure. What this buys you is the ability to notice degradation from a known baseline, which is the common case and the one that used to go unnoticed for days.

Continue reading