hsml.deployment #
Deployment #
NOT_FOUND_ERROR_CODE class-attribute instance-attribute #
Metadata object representing a deployment in Model Serving.
api_protocol property writable #
API protocol enabled in the deployment (e.g., HTTP or GRPC).
artifact_files_path property #
Path of the artifact files deployed by the predictor.
artifact_path property #
Path of the model artifact deployed by the predictor.
Deprecated
Artifact versions are deprecated in favor of deployment versions.
artifact_version property writable #
Artifact version deployed by the predictor.
Deprecated
Artifact versions are deprecated in favor of deployment versions.
config_file property writable #
Model server configuration file passed to the model deployment.
It can be accessed via CONFIG_FILE_PATH environment variable from a predictor or transformer script. For LLM deployments without a predictor script, this file is used to configure the vLLM engine.
created_at property #
Created at date of the predictor.
creator property #
Creator of the predictor.
description property writable #
Description of the deployment.
env_vars property writable #
Environment variables of the predictor.
environment property writable #
Name of the inference environment the predictor runs in.
Deprecated
Use deployment.predictor.environment, or deployment.transformer.environment for the transformer's own environment.
Changed in 5.2
Setting this on a deployment that was read back moves the predictor only. The transformer keeps the environment it was read with, where before 5.2 both components moved together. Set deployment.transformer.environment as well to move both.
feature_logging property writable #
Feature logging configuration attached to this deployment.
Edit its fields and call save(); a running deployment applies them after restart().
feature_view_name property #
feature_view_name: str | None
Name of the feature view served by a feature view deployment.
feature_view_version property #
feature_view_version: int | None
Version of the feature view served by a feature view deployment.
has_feature_view property #
has_feature_view: bool
Whether this deployment serves a feature view without a model.
has_model property #
Whether the deployment has a model associated.
id property #
Id of the deployment.
inference_batcher property writable #
Configuration of the inference batcher attached to this predictor.
inference_logger property writable #
Configuration of the inference logger attached to this predictor.
knative_mode property writable #
Whether the deployment runs in KServe Knative mode.
True selects Knative mode, False selects Standard mode and None lets the backend decide on creation or keeps the stored mode on an update. See Predictor.knative_mode for the full mode semantics.
Adds Knative mode selection, ~=5.1.0
Deployments can now select between KServe Knative and Standard mode.
missing_mandatory_tags property #
Mandatory tags configured for deployments that this deployment is missing.
Populated from the backend response. Empty when all mandatory deployment tags are set.
model_name property writable #
Name of the model deployed by the predictor.
model_path property writable #
Model path deployed by the predictor.
model_registry_id property writable #
Model Registry Id of the deployment.
model_server property writable #
Model server ran by the predictor.
model_version property writable #
Model version deployed by the predictor.
name property writable #
Name of the deployment.
predictor property writable #
Predictor used in the deployment.
project_name property writable #
Name of the project the deployment belongs to.
project_namespace property writable #
Name of the Kubernetes namespace the project is in.
requested_instances property #
Total number of requested instances in the deployment.
resources property writable #
Resource configuration for the predictor.
scaling_configuration property writable #
Scaling configuration for the deployment.
schema property writable #
Deployment schema, or None; see Predictor.schema.
script_file property writable #
Script file used by the predictor.
serving_tool property writable #
Serving tool used to run the model server.
tracing property writable #
Tracing configuration attached to this deployment.
training_dataset_version property #
training_dataset_version: int | None
Training dataset version whose statistics the deployment's transformations use.
transformer property writable #
Transformer configured in the predictor.
version property #
Number of the active configuration version of the deployment.
add_tag #
Attach a tag to a deployment.
A tag consists of a
| PARAMETER | DESCRIPTION |
|---|---|
name | Name of the tag to be added. TYPE: |
value | Value of the tag to be added. TYPE: |
| RAISES | DESCRIPTION |
|---|---|
hopsworks.client.exceptions.RestAPIError | in case the backend fails to add the tag. |
commit_feature_logs #
Run the commit job of the feature view this deployment logs through.
For a view on the "job" transport this commits every chunk that reached HopsFS, including what a stopped or killed replica left in the staging directory; for a "realtime" view it runs the materialization job.
| PARAMETER | DESCRIPTION |
|---|---|
wait | Whether to wait for the job to finish. TYPE: |
| RETURNS | DESCRIPTION |
|---|---|
list[Any] | The jobs that were started. |
| RAISES | DESCRIPTION |
|---|---|
hopsworks.client.exceptions.ModelServingException | If the deployment serves no feature view with logging enabled. |
create_feature_monitoring #
create_feature_monitoring(
name: str,
description: str | None = None,
start_date_time: int | str | None = None,
end_date_time: int | str | None = None,
cron_expression: str | None = "0 0 12 ? * * *",
) -> Any
Create a feature monitoring config on the logging feature group of a feature view deployment.
Model deployments use create_model_monitoring() instead. Finish the returned builder with a detection window, a reference window or value, a comparison, and save().
| PARAMETER | DESCRIPTION |
|---|---|
name | Name of the feature monitoring configuration. TYPE: |
description | Description of the feature monitoring configuration. TYPE: |
start_date_time | Start date and time from which to start computing statistics. |
end_date_time | End date and time at which to stop computing statistics. |
cron_expression | Cron expression scheduling the job (UTC, Quartz). TYPE: |
| RETURNS | DESCRIPTION |
|---|---|
Any | A |
| RAISES | DESCRIPTION |
|---|---|
hopsworks.client.exceptions.ModelServingException | If this is a model deployment, or the feature view has no logging enabled. |
create_model_monitoring #
create_model_monitoring(
name: str,
description: str | None = None,
start_date_time: int | str | None = None,
end_date_time: int | str | None = None,
cron_expression: str | None = "0 0 12 ? * * *",
) -> FeatureMonitoringConfig
Create a model monitoring config bound to this deployment's model.
Resolves the model's parent feature view from the registered provenance and delegates to feature_view.create_model_monitoring with this deployment's model_name and model_version already filled in. The resulting config targets the FV's logging feature group, filters by this model+version, and defaults the reference training dataset to the version that was used to train the model.
Experimental
Public API is subject to change, this feature is not suitable for production use-cases.
Example
my_deployment = ms.get_deployment(name="my_deployment")
my_deployment.create_model_monitoring(
name="psi_drift",
).with_detection_window(
time_offset="1d", window_length="1d",
).with_reference_training_dataset( # defaults to model's TD version
).compare_on_distribution(
feature_name="amount", metric="PSI", threshold=0.2,
).save()
| PARAMETER | DESCRIPTION |
|---|---|
name | Name of the feature monitoring configuration. TYPE: |
description | Description of the feature monitoring configuration. TYPE: |
start_date_time | Start date and time from which to start computing statistics. |
end_date_time | End date and time at which to stop computing statistics. |
cron_expression | Cron expression scheduling the FM job (UTC, Quartz). TYPE: |
| RAISES | DESCRIPTION |
|---|---|
hopsworks.client.exceptions.ModelServingException | If the deployment's model has no parent feature view recorded in its provenance. |
| RETURNS | DESCRIPTION |
|---|---|
FeatureMonitoringConfig | A |
FeatureMonitoringConfig |
|
FeatureMonitoringConfig |
|
delete #
delete(force: bool = False)
Delete the deployment.
| PARAMETER | DESCRIPTION |
|---|---|
force | Force the deletion of the deployment. If the deployment is running, it will be stopped and deleted automatically. TYPE: |
Warning
A call to this method does not ask for a second confirmation.
| RAISES | DESCRIPTION |
|---|---|
hopsworks.client.exceptions.RestAPIError | In case the backend encounters an issue. |
delete_tag #
delete_tag(name: str)
Delete a tag attached to a deployment.
| PARAMETER | DESCRIPTION |
|---|---|
name | Name of the tag to be removed. TYPE: |
| RAISES | DESCRIPTION |
|---|---|
hopsworks.client.exceptions.RestAPIError | in case the backend fails to delete the tag. |
download_artifact_files #
download_artifact_files(local_path: str | None = None)
Download the artifact files served by the deployment.
| PARAMETER | DESCRIPTION |
|---|---|
local_path | Path where to download the artifact files in the local filesystem. TYPE: |
| RAISES | DESCRIPTION |
|---|---|
hopsworks.client.exceptions.RestAPIError | In case the backend encounters an issue. |
download_logs #
Download the archived logs of this deployment from HopsFS.
Each running instance archives its own output to the project's Logs dataset without being asked: the container copies its log to Logs/Serving/<deployment_name>/ when it exits, is restarted, or is stopped, so a pod removed by scale-to-zero still leaves its output behind. Files are named <UTC yyyyMMdd-HHmmss>_<pod>_<component>.log, one per instance run.
Downloading the most recent archives
| PARAMETER | DESCRIPTION |
|---|---|
path | Local directory to download the archives into; the current working directory is used when unset. TYPE: |
latest | Download only the most recent archives instead of all of them. TYPE: |
| RETURNS | DESCRIPTION |
|---|---|
list[str] | The local paths of the downloaded archive files. |
| RAISES | DESCRIPTION |
|---|---|
hopsworks.client.exceptions.ModelServingException | If |
hopsworks.client.exceptions.RestAPIError | In case the backend encounters an issue. |
get_endpoint_url #
get_endpoint_url() -> str | None
Get the base endpoint URL for this deployment.
Returns the base URL that can be used with external HTTP clients. This is the path-based routing base endpoint without any protocol-specific suffixes like :predict or /v1.
If Istio client is not available, returns None.
| RETURNS | DESCRIPTION |
|---|---|
str | None | Base endpoint URL, or |
Examples:
get_feature_view #
Retrieve the feature view this deployment serves, or None.
A feature view deployment names its view in the deployment env vars; a model deployment resolves it through the model's provenance.
| PARAMETER | DESCRIPTION |
|---|---|
init | Whether to initialise the view for serving with the deployment's training dataset version. TYPE: |
| RETURNS | DESCRIPTION |
|---|---|
Any | The feature view, or |
| RAISES | DESCRIPTION |
|---|---|
hopsworks.client.exceptions.FeatureStoreException | If no project connection is available to reach the feature store. |
get_inference_url #
get_inference_url() -> str | None
Get the KServe inference URL for standard model deployments.
Returns the full URL with :predict suffix for KServe inference protocol. This method only returns a URL for standard model deployments (non-vLLM, with a model attached).
If Istio client is not available, falls back to Hopsworks REST API path.
| RETURNS | DESCRIPTION |
|---|---|
str | None | Inference URL with |
Examples:
get_logs #
Prints the deployment logs of the predictor or transformer.
Only the live pods of a running deployment are read. Logs of a stopped deployment are whatever was saved to HopsFS beforehand, retrieved with meth:
download_logs or the "Log history" section in the UI.
.. note:: Legacy: this method prints to stdout and returns None. New code (and any agent / scripted use) should call meth:
read_logs for a string return value or meth:
tail_logs for incremental streaming.
| PARAMETER | DESCRIPTION |
|---|---|
component | Deployment component to get the logs from (e.g., predictor or transformer). TYPE: |
tail | Number of most recent lines to retrieve from the logs. TYPE: |
| RAISES | DESCRIPTION |
|---|---|
hopsworks.client.exceptions.RestAPIError | In case the backend encounters an issue. |
get_model #
Retrieve the metadata object for the model being used by this deployment, or None when it has no model.
get_monitoring_configs #
get_monitoring_configs() -> list[FeatureMonitoringConfig]
Get the feature monitoring configurations for the model deployed by this deployment.
For a model deployment these are the configs filtered by the model version; for a feature view deployment the configs on the feature view's logging feature group. Delegates to the underlying model's get_monitoring_configs method.
Example
| RETURNS | DESCRIPTION |
|---|---|
list[FeatureMonitoringConfig] | List of |
| RAISES | DESCRIPTION |
|---|---|
hopsworks.client.exceptions.RestAPIError | In case the backend encounters an issue. |
get_openai_url #
get_openai_url() -> str | None
Get the OpenAI-compatible API URL for vLLM deployments.
Returns the URL for OpenAI-compatible API endpoints (e.g., /v1/chat/completions). This method only returns a URL for LLM (vLLM) deployments.
| RETURNS | DESCRIPTION |
|---|---|
str | None | OpenAI-compatible URL (base URL + "/v1"), or |
Examples:
get_state #
get_state() -> PredictorState
Get the current state of the deployment.
| RETURNS | DESCRIPTION |
|---|---|
PredictorState | The state of the deployment. |
| RAISES | DESCRIPTION |
|---|---|
hopsworks.client.exceptions.RestAPIError | In case the backend encounters an issue. |
get_tag #
Get the value of a tag attached to a deployment.
| PARAMETER | DESCRIPTION |
|---|---|
name | Name of the tag to get. TYPE: |
| RETURNS | DESCRIPTION |
|---|---|
Any | None | tag value, or |
| RAISES | DESCRIPTION |
|---|---|
hopsworks.client.exceptions.RestAPIError | in case the backend fails to retrieve the tag. |
get_tags #
Retrieve all tags attached to a deployment.
| RETURNS | DESCRIPTION |
|---|---|
dict[str, Any] | Dictionary of tag name/values. |
| RAISES | DESCRIPTION |
|---|---|
hopsworks.client.exceptions.RestAPIError | in case the backend fails to retrieve the tags. |
get_versions #
get_versions() -> list[DeploymentVersion]
Get every configuration version this deployment has ever had, newest first.
| RETURNS | DESCRIPTION |
|---|---|
list[DeploymentVersion] | One |
| RAISES | DESCRIPTION |
|---|---|
hopsworks.client.exceptions.RestAPIError | In case the backend encounters an issue. |
hopsworks.client.exceptions.ModelServingException | If the deployment has not been saved yet or the backend does not support deployment versions. |
init_predict #
Prepare this deployment for inference, so the first predict() does not.
Downloads the schema the deployment's revision names and connects the transport its API protocol selects: the gRPC channel, or for REST an open connection to the model's endpoint, made with a metadata GET. Calling it is optional: predict() prepares the schema and the channel the same way on its first call, through the same code, and opens its connection with the request itself. Call it when the first request should not pay for discovery or the connection setup, for instance when a serving process starts before it takes traffic.
No prediction is sent. A prediction can log rows and have application side effects, so it is not used as a warm-up.
Concurrent callers prepare once and share the result. A failure leaves the deployment unprepared and is raised to the caller, so it can be tried again.
| RAISES | DESCRIPTION |
|---|---|
hopsworks.client.exceptions.RestAPIError | In case the backend encounters an issue. |
is_created #
is_created() -> bool
Check whether the deployment is created.
| RETURNS | DESCRIPTION |
|---|---|
bool | Whether the deployment is created or not. |
| RAISES | DESCRIPTION |
|---|---|
hopsworks.client.exceptions.RestAPIError | In case the backend encounters an issue. |
is_running #
Check whether the deployment is ready to handle inference requests.
| PARAMETER | DESCRIPTION |
|---|---|
or_idle | Whether the idle state is considered as running (default is True). TYPE: |
or_updating | Whether the updating state is considered as running (default is True). TYPE: |
| RETURNS | DESCRIPTION |
|---|---|
bool | Whether the deployment is ready or not. |
| RAISES | DESCRIPTION |
|---|---|
hopsworks.client.exceptions.RestAPIError | In case the backend encounters an issue. |
is_stopped #
Check whether the deployment is stopped.
| PARAMETER | DESCRIPTION |
|---|---|
or_created | Whether the creating and created state is considered as stopped (default is True). TYPE: |
| RETURNS | DESCRIPTION |
|---|---|
bool | Whether the deployment is stopped or not. |
| RAISES | DESCRIPTION |
|---|---|
hopsworks.client.exceptions.RestAPIError | In case the backend encounters an issue. |
predict #
predict(
data: dict | InferInput = None,
inputs: list | dict = None,
validate: bool = True,
) -> dict
Send inference requests to the deployment.
One of data or inputs parameters must be set. Setting both raises ModelServingException. When the deployment has a schema, the rows are encoded and validated against it before the request is sent, and a gRPC deployment sends them as one v2 tensor per field; the pod validates again regardless.
| PARAMETER | DESCRIPTION |
|---|---|
data | Payload dictionary for the inference request including the model input(s). TYPE: |
inputs | Model inputs used in the inference requests. |
validate | Whether to validate the rows against TYPE: |
| RETURNS | DESCRIPTION |
|---|---|
dict | Inference response. |
| RAISES | DESCRIPTION |
|---|---|
hopsworks.client.exceptions.RestAPIError | In case the backend encounters an issue. |
Examples:
# login into Hopsworks using hopsworks.login()
# get Hopsworks Model Serving handle
ms = project.get_model_serving()
# retrieve deployment by name
my_deployment = ms.get_deployment("my_deployment")
# (optional) retrieve model input example
my_model = project.get_model_registry() .get_model(my_deployment.model_name, my_deployment.model_version)
# make predictions using model inputs (single or batch)
predictions = my_deployment.predict(inputs=my_model.input_example)
# or using more sophisticated inference request payloads
data = { "instances": [ my_model.input_example ], "key2": "value2" }
predictions = my_deployment.predict(data)
read_logs #
read_logs(
component: str = "predictor",
tail: int = 100,
source: str = "kubernetes",
since: str | None = None,
until: str | None = None,
pod: str | None = None,
) -> str
Return deployment logs as a single plain-text string.
Programmatic counterpart to meth:
get_logs. Suitable for agents and scripts: never prints, never short-circuits on deployment state. The default source="kubernetes" reads the live pods only; logs of a stopped deployment are whatever was saved beforehand and are archived to HopsFS and retrieved with meth:
download_logs or the "Log history" section in the UI.
| PARAMETER | DESCRIPTION |
|---|---|
component |
TYPE: |
tail | Most-recent lines to retrieve. Capped server-side. TYPE: |
source |
TYPE: |
since | ISO-8601 lower bound on log timestamp. Ignored on the Kubernetes path. TYPE: |
until | ISO-8601 upper bound on log timestamp. Ignored on the Kubernetes path. TYPE: |
pod | Restrict to one instance / container name. TYPE: |
| RETURNS | DESCRIPTION |
|---|---|
str | The joined logs as plain text. Empty string when there are no |
str | matching lines; |
str | multiple instances are present. |
reinfer_schema #
reinfer_schema() -> DeploymentSchema
Re-infer the deployment schema from the current feature view and mark it pending.
Use after enabling logging or changing the view; save() then publishes the new schema as a new revision. Passed features are kept.
Example
| RETURNS | DESCRIPTION |
|---|---|
DeploymentSchema | The re-inferred schema, also set on the deployment. |
| RAISES | DESCRIPTION |
|---|---|
hopsworks.client.exceptions.ModelServingException | If the deployment is not served by the default predictor and has no schema to refine. |
restart #
Restart the deployment so it picks up the latest code and environment state.
If the deployment is already stopped, it is started in place.
| PARAMETER | DESCRIPTION |
|---|---|
await_stopped | Awaiting time (seconds) for the deployment to stop. TYPE: |
await_running | Awaiting time (seconds) for the deployment to start again. TYPE: |
| RAISES | DESCRIPTION |
|---|---|
hopsworks.client.exceptions.RestAPIError | In case the backend encounters an issue. |
rollback #
rollback(
version: int | DeploymentVersion,
await_update: int | None = 600,
) -> None
Make an earlier version of this deployment the active one again.
Nothing is copied: the version's files are still where they were, and it keeps its number. This object is updated to the reactivated configuration. Rolling back to the version that is already active does nothing.
Running instances are restarted
A running deployment is rolled to the version, so its pods restart and requests fail over during the rollout. The API protocol, request batching, inference logging, scheduling configuration and Knative mode are kept as they are, because they are not part of a version. A version that was edited in place after it was created comes back as edited, not as it was first saved.
Rolling back while a new version is updating replaces that rollout with one for the target version; the instances that were serving keep serving until it is ready. An in-place edit of the active version cannot be undone this way: stop the deployment or wait for it to fail, then save the previous configuration.
| PARAMETER | DESCRIPTION |
|---|---|
version | The version to activate, as its number or as a TYPE: |
await_update | If the deployment is running, awaiting time (seconds) for the running instances to be updated. If the running instances are not updated within this timespan, the call raises, while the update continues in the background. A value below 5 is rounded up to one 5-second poll. TYPE: |
| RAISES | DESCRIPTION |
|---|---|
hopsworks.client.exceptions.RestAPIError | In case the backend encounters an issue, including a version the deployment does not have. |
hopsworks.client.exceptions.ModelServingException | If the deployment is starting or stopping, has not been saved yet, or the backend does not support deployment versions, or if the running instances are not updated within |
save #
Persist this deployment including the predictor and metadata to Model Serving.
On an existing deployment the active version is edited in place by default and keeps its number. With new_version the configuration is stored as a new version, numbered one above the highest the deployment ever had, and made active. A save that changes nothing does nothing in either mode: no version is created and the running instances are not restarted. Two cases always count as a change: a script read from outside the version (a HopsFS mount path or git) and a changed environment image. Use Deployment.restart to refresh the running instances on purpose.
Older backends
A backend that predates deployment versions has no in-place edit. There a changed script is stored as a new version automatically, and new_version raises.
What a version holds
The predictor and transformer scripts, config file, resources, scaling, environment variables, environments, tracing, feature logging configuration, git source, vLLM settings and the model artifact belong to a version. The API protocol, request batching, inference logging, scheduling configuration and Knative mode do not. They are edited in place whichever way you save, and a rollback does not restore them.
| PARAMETER | DESCRIPTION |
|---|---|
await_update | If the deployment is running, awaiting time (seconds) for the running instances to be updated. If the running instances are not updated within this timespan, the call raises, while the update continues in the background. A value below 5 is rounded up to one 5-second poll. TYPE: |
new_version | Store this configuration as a new version of the deployment instead of editing the active one. TYPE: |
| RAISES | DESCRIPTION |
|---|---|
hopsworks.client.exceptions.RestAPIError | In case the backend encounters an issue, including a conflict when the deployment was rolled back or given a new version by someone else since it was read. |
hopsworks.client.exceptions.ModelServingException | If |
hopsworks.client.exceptions.ModelServingException | If the running instances are not updated within |
start #
start(await_running: int | None = 600)
Start the deployment.
| PARAMETER | DESCRIPTION |
|---|---|
await_running | Awaiting time (seconds) for the deployment to start. If the deployment has not started within this timespan, the call raises, while it deploys in the background. A value below 5 is rounded up to one 5-second poll. TYPE: |
| RAISES | DESCRIPTION |
|---|---|
hopsworks.client.exceptions.RestAPIError | In case the backend encounters an issue. |
hopsworks.client.exceptions.ModelServingException | If the deployment fails to start or has not started within |
stop #
stop(await_stopped: int | None = 600)
Stop the deployment.
| PARAMETER | DESCRIPTION |
|---|---|
await_stopped | Awaiting time (seconds) for the deployment to stop. If the deployment has not stopped within this timespan, the call raises, while it stops in the background. A value below 5 is rounded up to one 5-second poll. TYPE: |
| RAISES | DESCRIPTION |
|---|---|
hopsworks.client.exceptions.RestAPIError | In case the backend encounters an issue. |
hopsworks.client.exceptions.ModelServingException | If the deployment has not stopped within |
tail_logs #
tail_logs(
component: str = "predictor",
interval: float = 2.0,
source: str = "kubernetes",
since: str | None = "now",
timeout: float | None = None,
stop_on_status: str | None = None,
pod: str | None = None,
) -> Iterator[str]
Yield only newly observed log chunks as plain text.
Client-side polling, not server-streaming: each tick calls meth:
read_logs with a moving cursor and yields the portion not already seen. The Kubernetes path dedups per pod by overlapping the previous and current tail windows; the OpenSearch path (old backends) uses the timestamp + doc_id pair. Only the live pods are followed; the archived logs of a stopped deployment are retrieved with meth:
download_logs.
Following a deployment's live logs
| PARAMETER | DESCRIPTION |
|---|---|
component |
TYPE: |
interval | Seconds between polls. TYPE: |
source |
TYPE: |
since |
TYPE: |
timeout | Stop after this many seconds. TYPE: |
stop_on_status | Stop when TYPE: |
pod | Follow one specific instance by name. The backend reads the first eight replicas of a component per request, so a deployment scaled beyond that needs the later instances tailed one by one. TYPE: |
| YIELDS | DESCRIPTION |
|---|---|
str | Plain-text log chunks containing only newly observed content. |