Model Registry & Serving Guides#
A model goes from training into the registry, out through a deployment, and stays under monitoring. The guides below follow that path.
-
Start here
Register a trained model with its metrics, then deploy it in one call.
Register
- Model registry Save a model with metrics, a schema and an input example, one version per save.
- Frameworks TensorFlow, PyTorch, scikit-learn, LLM and plain Python models.
- Import from Hugging Face Bring a Hub model into the registry without training it here.
- Evaluation images Attach plots and confusion matrices to a model version.
Serve
- Deployment creation Deploy a registered model and check its state.
- Predictor and transformer Custom inference code and pre/post-processing on KServe.
- Logging and batching Log requests to a feature group, batch them for throughput.
- Resources and autoscaling CPU, memory, GPU and replica bounds, with scheduling constraints.
- Reach the endpoint API protocol, REST access from outside, troubleshooting.
Observe
- Model monitoring Compare logged inference data against the training dataset on a schedule.
- Provenance Trace a model back to its training data and features.
- Vector database Similarity search over embeddings stored in the feature store.
- Agents Agent tasks as jobs and served interactive agents.