Every metric an inference service exports carries the model version as a label.
The payoff is that "did this change when we deployed" becomes a series split rather than an investigation. You stop correlating a deployment timestamp against a graph by eye and start comparing two series directly.
Cost: one label. There is a cardinality argument against adding labels casually, and it does not apply here, because the number of model versions serving traffic at once is small and bounded by your own deployment practice.