# Hopsworks Documentation > Official documentation for Hopsworks and its Feature Store - an open source data-intensive AI platform used for the development and operation of machine learning models at scale. ================================================================================ # Home Source: https://docs.hopsworks.ai/latest/

Hopsworks Documentation

# Build, deploy & maintain AI systems

Features, training data, models and inference on one governed platform.

## Your first feature vector, in three steps Install the client, connect to a project with an [API key](user_guides/projects/api_key/create_api_key.md), write a feature group and read a feature vector back.
=== "Python" ```bash uv venv && source .venv/bin/activate uv pip install "hopsworks[python]" hops setup # opens a browser, picks a project, caches the key python # opens the interpreter, the Python lines below go there or in a notebook ``` ```python import hopsworks project = hopsworks.login() # uses the key hops setup cached, else prompts fs = project.get_feature_store() ``` === "CLI" ```bash uv venv && source .venv/bin/activate uv pip install "hopsworks[python]" hops setup # opens a browser, picks a project, caches the key hops fg list ```
=== "Python" ```python import pandas as pd df = pd.DataFrame( { "cc_num": [4467360740682089], "amount": [12.5], "event_time": pd.to_datetime(["2026-01-01T00:00:00Z"]), } ) fg = fs.get_or_create_feature_group( name="transactions", version=1, primary_key=["cc_num"], event_time="event_time", online_enabled=True, ) fg.insert(df) ``` === "CLI" ```bash hops fg create transactions --version 1 --primary-key cc_num \ --features "cc_num:bigint,amount:double" --online echo '[{"cc_num": 4467360740682089, "amount": 12.5}]' | hops fg insert transactions --version 1 ```
=== "Python" ```python fv = fs.get_or_create_feature_view( name="transactions_view", version=1, query=fg.select_all(), ) fv.get_feature_vector(entry={"cc_num": 4467360740682089}) ``` === "CLI" ```bash hops fv create transactions_view --version 1 --feature-group transactions hops fv get transactions_view --version 1 --entry "cc_num=4467360740682089" ```
Next: [create a feature group](user_guides/fs/feature_group/create.md), [create a feature view](user_guides/fs/feature_view/overview.md), [retrieve feature vectors](user_guides/fs/feature_view/feature-vectors.md), or browse the Python API. ## Where Hopsworks runs
- :material-cloud-outline: **Use the managed SaaS** --- Sign in to the Hopsworks serverless app and create a project. Nothing to install, free tier available. [Open run.hopsworks.ai ↗](https://run.hopsworks.ai) - :material-server: **Deploy on your cloud or on-prem** --- Managed Kubernetes on AWS, Azure or GCP, or an air-gapped data centre. Talk to us to size and install it. [Contact Hopsworks ↗](https://www.hopsworks.ai/contact) · [Deployment options](setup_installation/index.md)
## One architecture, three pipelines Independent [feature, training and inference pipelines](concepts/fti.md), connected by a shared feature store and model registry. --8<-- "index/one-architecture-three-pipelines.html" ## Find your path
:material-code-tags:{ .hops-role-ico } Developer { .hops-role-cap } - [Client installation](user_guides/client_installation/index.md) - [Python, SageMaker, Kubeflow](user_guides/integrations/python.md) - [Create a feature group](user_guides/fs/feature_group/create.md) - Python API reference
:material-chart-scatter-plot:{ .hops-role-ico } Data scientist { .hops-role-cap } - [Feature views](concepts/fs/feature_view/fv_overview.md) - [Training data](user_guides/fs/feature_view/training-data.md) - [Model serving](user_guides/mlops/serving/index.md)
:material-server:{ .hops-role-ico } Platform engineer { .hops-role-cap } - [Jobs](user_guides/projects/jobs/python_job.md) - [Airflow](user_guides/projects/airflow/airflow.md) - [Python environments](user_guides/projects/python/python_env_overview.md) - [Kubernetes scheduling](user_guides/projects/scheduling/kube_scheduler.md)
:material-shield-lock-outline:{ .hops-role-ico } Security engineer { .hops-role-cap } - [Configure authentication](setup_installation/admin/auth.md) - [API keys](user_guides/projects/api_key/create_api_key.md) - [IAM role chaining](setup_installation/admin/roleChaining.md) - [Audit logs](setup_installation/admin/audit/audit-logs.md)
:material-cog-outline:{ .hops-role-ico } Administrator { .hops-role-cap } - [Administration overview](setup_installation/admin/index.md) - [User management](setup_installation/admin/user.md) - [Alerts](setup_installation/admin/alert.md) - [HA and DR](setup_installation/admin/ha-dr/intro.md)
:material-compass-outline:{ .hops-role-ico } Evaluator { .hops-role-cap } - [What Hopsworks is](concepts/hopsworks.md) - [Feature store architecture](concepts/fs/index.md) - [Analytical and operational ML](concepts/mlops/prediction_services.md) - [Deployment options](setup_installation/index.md)
## By task | Task | Start here | | --- | --- | | :material-rocket-launch-outline: Deploy | [AWS](setup_installation/aws/getting_started.md), [Azure](setup_installation/azure/getting_started.md), [GCP](setup_installation/gcp/getting_started.md), [on-prem](setup_installation/on_prem/contact_hopsworks.md), [Helm values][helm-chart-values-reference] | | :material-monitor-dashboard: Operate | [Administration](setup_installation/admin/index.md), [monitoring](setup_installation/admin/monitoring/grafana.md), [alerts](setup_installation/admin/alert.md), [HA and DR](setup_installation/admin/ha-dr/intro.md), [service operations](setup_installation/admin/operationLogs.md) | | :material-wrench-outline: Troubleshoot | [Model serving](user_guides/mlops/serving/troubleshooting.md), [Python deployments](user_guides/projects/python-deployment/troubleshooting.md), [online ingestion](user_guides/fs/feature_group/online_ingestion_observability.md), [Jupyter session capacity](user_guides/projects/jupyter/session_capacity_warnings.md) | | :material-arrow-up-circle-outline: Upgrade | [3.x to 4.0 migration](user_guides/migration/40_migration.md), [Airflow 3 upgrade](user_guides/projects/airflow/airflow3_upgrade.md), [Airflow 3 operator notes](setup_installation/admin/airflow3.md) |
:material-api:{ .hops-colophon-ico } APIs { .hops-colophon-cap } - Python API - Java API - Machine-readable: llms.txt, llms-full.txt, or `/index.md`
:material-tune:{ .hops-colophon-ico } Configure and query { .hops-colophon-cap } - [Helm chart values][helm-chart-values-reference] - [Cluster configuration](setup_installation/admin/variables.md) - [Query engine (Trino)](user_guides/projects/trino/query_engine.md) - [Vector similarity search](user_guides/fs/vector_similarity_search.md)
:material-forum-outline:{ .hops-colophon-ico } Community and source { .hops-colophon-cap } - [Public Slack ↗](https://join.slack.com/t/public-hopsworks/shared_invite/zt-24fc3hhyq-VBEiN8UZlKsDrrLvtU4NaA) - [hopsworks-api on GitHub ↗](https://github.com/logicalclocks/hopsworks-api) - [Apache License 2.0 ↗](https://www.apache.org/licenses/LICENSE-2.0.html)
================================================================================ # Concepts Source: https://docs.hopsworks.ai/latest/concepts/ # Concepts This section explains what Hopsworks is and why it is built the way it is. It is reference and explanation, not step-by-step instructions. For the how-to, see the [Guides](../user_guides/index.md). ## Start here Read the [FTI Pipeline Architecture](fti.md) first. It is the one idea the rest of this section builds on: every AI system decomposes into feature, training, and inference pipelines, connected through a feature store and a model registry. Once you have that model, the other pages are the parts of it. ## Reading path - [Hopsworks Platform](hopsworks.md): the components of the platform and how they fit together. - [FTI Pipeline Architecture](fti.md): the architecture all AI systems share, and the four classes of AI system. - **Feature Store**: how feature pipelines write feature data ([feature groups](fs/feature_group/fg_overview.md)) and how training and inference pipelines read it ([feature views](fs/feature_view/fv_overview.md)). - **Projects**: the multi-tenant unit that owns your data and ML assets, with governance, sharing, and lineage. - **MLOps**: training, the model registry, serving, and monitoring, the inference side of an AI system. - **Development**: building and running pipelines inside and outside Hopsworks. ## How the section is organised The Feature Store pages follow the write path then the read path: you write features to feature groups, and you read them through feature views. The MLOps pages follow a model from training through registration, serving, and monitoring. Projects and Development cut across both. ================================================================================ # Hopsworks Platform Source: https://docs.hopsworks.ai/latest/concepts/hopsworks/ # The Hopsworks Platform Hopsworks is a **modular** MLOps platform with: - a feature store (available as standalone) - model registry and model serving based on KServe - vector index based on OpenSearch - a data science and data engineering platform MLOps is a set of best practices for the automated testing, versioning, and monitoring of the ML pipelines and ML assets that power AI systems. Hopsworks is modular, so you can adopt the feature store on its own or use the full platform across the MLOps lifecycle. --8<-- "concepts/hopsworks/the-hopsworks-platform.html" ## Standalone Feature Store Hopsworks was the first open-source and first enterprise feature store for ML. You can use Hopsworks as a standalone feature store with the Hopsworks API. ## Model Management Hopsworks includes support for model management, with model deployments using [the KServe framework](https://github.com/kserve/kserve) and a model registry designed for KServe. Hopsworks logs all inference requests to Kafka to enable easy monitoring of deployed models, and provides model metrics with grafana/prometheus. ## Vector Index A feature group with an embedding column can have a vector index, based on [OpenSearch kNN](https://opensearch.org/docs/latest/search-plugins/knn/index/) (on the [FAISS](https://ai.facebook.com/tools/faiss/) engine, which has replaced the deprecated [nmslib](https://github.com/nmslib/nmslib) engine as the default). The vector index includes out-of-the-box support for authentication, access control, filtering, backup-and-restore, and horizontal scalability. The Feature Store and its vector index are often used together to build scalable recommender systems, such as ranking-and-retrieval for real-time recommendations. ## Governance Hopsworks provides a data-mesh architecture for managing ML assets and teams, with multi-tenant projects. Not unlike a GitHub repository, a project is a sandbox containing team members, data, and ML assets. In Hopsworks, all ML assets (features, models, training data) are versioned, taggable, lineage-tracked, and support free-text search. Data can be also be securely shared between projects. ## Data Science Platform You can develop feature engineering, model training and inference pipelines in Hopsworks. There is support for version control (GitHub, GitLab, BitBucket), Jupyter notebooks, a shared distributed file system, many bundled modular project Python environments for managing Python dependencies without needing to write Dockerfiles, jobs (Python, Spark), and workflow orchestration with Airflow. ================================================================================ # FTI Pipeline Architecture Source: https://docs.hopsworks.ai/latest/concepts/fti/ # FTI Pipeline Architecture Hopsworks is built around a single architecture for AI systems: the decomposition of any AI system into **feature**, **training**, and **inference** (FTI) pipelines. This page defines that architecture. Every other concept in this section is a part of it, so read this first. ## The three pipelines An AI system decomposes naturally into three machine learning pipelines, each with clear inputs and outputs, each developed, tested, and operated independently. - A **feature pipeline** takes data as input and produces reusable feature data as output. - A **training pipeline** takes feature data as input, trains a model, and outputs the trained model. - An **inference pipeline** takes feature data and a model as input and outputs predictions and prediction logs. The three pipelines are independent programs. They are composed into a working system through a shared data layer: a [feature store](fs/index.md) and a [model registry](mlops/registry.md). --8<-- "concepts/fti/the-three-pipelines.html" Feature pipelines ingest both backfill and production data and compute feature data that is stored as tabular data in the feature store. Feature pipelines can be batch programs or stream processing programs. Training pipelines read training data from the feature store and store the models they produce in the model registry. Inference pipelines output predictions using a model, either downloaded from the model registry or served behind an API, together with new feature data that is precomputed in the feature store or computed from data available at prediction request time. ## Why this architecture The five common AI system architectures (batch, stateless real-time, stateful real-time, RAG, and agentic) are very different from one another. Moving from one to another, or transferring what you learned building one, is hard. The FTI decomposition gives you one architecture for all of them. Modularity is the reason. Splitting an AI system into independent, small, testable modules lets teams build higher-quality systems faster. It also splits the work cleanly: feature engineering can involve data engineers, model training is the realm of data scientists, and inference can involve operations. ## The shared data layer The feature store holds three stores of feature data, each serving a different pipeline. - A row-oriented online store for low latency access from online inference pipelines and agents. - A columnar offline store for training models and batch inference. - A [vector index](mlops/opensearch.md) over embeddings for inference pipelines and agents. The model registry holds the trained models and their assets, versioned, for inference pipelines to load. ## What an AI system is An AI system is a set of independent feature pipelines, training pipelines, and inference pipelines connected through a feature store and a model registry. An AI system is defined by how it computes its predictions, not by the type of application that consumes them. On that basis, AI systems built with a feature store fall into four classes. - **Real-time (interactive)** systems make predictions in response to user requests. They read precomputed features from the feature store and can also compute features on demand from the request parameters. - **Agentic workflows** achieve goals with some autonomy using LLMs and tools, drawing context from a vector index, the online and offline stores, and external APIs. - **Batch** systems run inference on a schedule and write predictions to a downstream store, called an inference store, for an application to consume later. - **Stream processing** systems use an embedded model to make predictions on streaming data without user input, often machine to machine. The inference pipeline is what determines the class. When you know how a system computes its predictions, you know which of these you are building. ## Where to go next - [Feature Store Architecture](fs/index.md) for the shared data layer in detail. - [Feature Groups](fs/feature_group/fg_overview.md) for how feature pipelines write feature data. - [Feature Views](fs/feature_view/fv_overview.md) for how training and inference pipelines read it. - [AI Systems](mlops/prediction_services.md) for the inference side and how each class is served. ================================================================================ # AI Systems Source: https://docs.hopsworks.ai/latest/concepts/mlops/prediction_services/ # AI Systems An AI system is a set of independent feature pipelines, training pipelines, and inference pipelines that are connected via a feature store and model registry. Each pipeline is a separate program with its own inputs and outputs, and the shared data layer is what lets them be developed, run, and scaled independently. An AI system is defined by how it computes its predictions, not by the type of application that consumes them. The inference pipeline determines the class of AI system you are building. There are four classes: - **Real-time (interactive)**: a client sends a prediction request and an online inference pipeline computes and returns a prediction with low latency. - **Batch**: an inference pipeline runs on a schedule, scores a set of entities, and writes the predictions to an inference store. - **Stream processing**: an inference pipeline computes predictions continuously over an event stream. - **Agentic workflows**: an LLM-driven control flow decides which steps to run, retrieving the context and features it needs from the feature store. See [Agents and LLM Systems](agents.md). Whatever the class, an AI system is composed of the same parts: - one or more feature pipeline(s) that keep the feature store up to date, - a training pipeline that produces a model in the model registry, - an inference pipeline that reads features and computes predictions, - a sink for the predictions, either an inference store or a user interface. The two figures below illustrate the two most common classes, batch and real-time. ## Batch AI systems In the figure below, feature pipelines update the feature store with new feature data on a schedule (e.g., hourly, daily). A batch inference pipeline also runs on a schedule, reads batch scoring data from the feature store, computes predictions with an embedded model, and writes those predictions to an inference store. The inference store is any data store that holds the predictions from batch inference pipelines. From there, the predictions are consumed by (predictive, prescriptive) analytical reports and/or to AI-enable operational services. --8<-- "concepts/mlops/prediction_services/batch-ai-systems.html" ## Real-time AI systems In the figure below, feature pipelines update the feature store with new feature data on a schedule (e.g., streaming, hourly, daily), and the operational service sends prediction requests to a model deployed on KServe via its secured Istio endpoint. A deployed model on KServe handles the prediction request by first retrieving pre-computed features from the feature store for the given request, and then building a feature vector that is scored by the model. The prediction result is returned to the client (the operational service). KServe logs both the feature values and the prediction results back to Hopsworks for further analysis and to help create new training data. --8<-- "concepts/mlops/prediction_services/real-time-ai-systems.html" ## MLOps Flywheel Once you have built your batch or real-time AI system, the MLOps flywheel is the path to building a self-managing system that automatically collects and processes feature logs, prediction logs, and outcomes to help create new training data for models. This enables a ML flywheel where new training data and insights are generated from your AI system, by feeding logs back into the feature store. More training data enables the training of better models, and with better models, you should hopefully improve your operational/batch services, so that you attract more clients, who in turn produce more data for training models. And, thus, the ML flywheel is bootstrapped and leads to a virtuous cycle of more data leading to better models and more models leading to more users, who produce more data, and so on. === "Offline path: the training loop" --8<-- "concepts/mlops/prediction_services/mlops-flywheel-offline.html" === "Online path: the serving loop" --8<-- "concepts/mlops/prediction_services/mlops-flywheel-online.html" ================================================================================ # Agents and LLM Systems Source: https://docs.hopsworks.ai/latest/concepts/mlops/agents/ # Agents and LLM Systems Agentic workflows are one of the four classes of AI system, alongside real-time, batch, and stream processing. Context engineering for an agent follows many of the same principles as feature engineering for a classical ML model, and the feature store is where the context comes from. ## The feature store as a retrieval source An LLM system retrieves the context it needs at inference time, and the feature store is a natural source for that context. Precomputed features are retrieved by entity ID, and embeddings are retrieved from a vector index by similarity search. The key requirement is that the entity IDs are provided in the user query, as part of the deployment API, so the system knows whose features to retrieve. This is retrieval-augmented generation (RAG) with a feature store: structured features by key, unstructured context by similarity. --8<-- "concepts/mlops/agents/the-feature-store-as-a-retrieval-source.html" ## Workflow or agent An LLM workflow has a control flow the developer designs: the steps and their order are fixed, and the LLM fills in each step. An agent decides its own control flow: the LLM chooses which steps to run and in what order, calling tools as it goes. A workflow is more predictable, an agent is more flexible, and most production systems start as workflows. ## MCP and A2A Two protocols connect the moving parts. MCP (Model Context Protocol) is how an agent calls its tools, the intra-agent interface to data sources and functions, including a feature store. A2A (Agent-to-Agent) is how agents talk to each other, the inter-agent interface. See the [agent guides](../../user_guides/agents/index.md) for how to build and deploy agents and agent tasks on Hopsworks. ================================================================================ # Projects and Governance Source: https://docs.hopsworks.ai/latest/concepts/projects/governance/ # Projects and Governance Hopsworks provides project-level multi-tenancy, a data mesh enabling technology. Think of it as a GitHub repository for your teams and ML assets. More specifically, a project is a sandbox for team members, ML assets (features, training data, models, vector index, model deployments), and optionally feature pipelines and training pipelines. The ML assets can only be accessed by project members, and there is role-based access control (RBAC) for project members within a project. --8<-- "concepts/projects/governance/projects-and-governance.html" ## Dev/Staging/Prod for Data Projects enable you to define development, staging, and even production projects on the same cluster. Often, companies deploy production projects on dedicated clusters, but development projects and staging projects on a shared cluster. This way, projects can be easily used to implement CI/CD workflows. ## Data Mesh of Feature Stores Projects enable you to move beyond the traditional dev/staging/prod ownership model for data. Different teams or lines of business can have their own private feature stores, you can mix them with a group-wide feature store, and feature stores can be securely shared between teams/organizations. Effectively, you can have decentralized ownership of feature stores, with domain-specific projects, and each project managing its own feature pipelines. Hopsworks provides data/feature sharing support between these self-service projects. ## Audit Logs with REST API Hopsworks stores audit logs for all calls on its REST API in its file system, HopsFS. The audit log can be used to analyze the historical usage of services by users. ================================================================================ # Data Storage and Sharing Source: https://docs.hopsworks.ai/latest/concepts/projects/storage/ # Data Storage and Sharing Every project in Hopsworks has its own private assets: - a Feature Store (including both Online and Offline Stores) - a Filesystem subtree (all directory and files under /Projects//) - a Model Registry - Model Deployments - Kafka topics - OpenSearch indexes (including kNN indexes, the vector index) - a Hive Database Access control to these assets is controlled using project membership ACLs (access-control lists). Users in a project who have a *Data Owner* role have read/write access to these assets. Users in a project who have a *Data Scientist* role have mostly read-only access to these assets, with the exception of the ability to write to well-known directories (Resources, Jupyter, Logs). However, it is often desirable to share assets between projects, with read-only, read/write privileges, and to restrict the privileges to specific role (e.g., Data Owners) in the target project. In Hopsworks, you can explicitly share assets between projects without copying the assets. Sharing is managed by ACLs in Hopsworks, see example below: --8<-- "concepts/projects/storage/data-storage-and-sharing.html" ================================================================================ # Architecture Source: https://docs.hopsworks.ai/latest/concepts/fs/ # Feature Store Architecture ## What is Hopsworks Feature Store? Hopsworks and its Feature Store are an open source data-intensive AI platform used for the development and operation of machine learning models at scale. The Hopsworks Feature Store provides the Hopsworks API to enable clients to write features to feature groups in the feature store, and to read features from feature views - either through a low latency Online API to retrieve pre-computed features for operational models or through a high throughput, latency insensitive Offline API, used to create training data and to retrieve batch data for scoring. --8<-- "concepts/fs/index/what-is-hopsworks-feature-store.html" ## Hopsworks API The Hopsworks API is how you, as a developer, will use the feature store. The Hopsworks API helps simplify some of the problems that feature stores address including: - consistent features for training and serving - centralized, secure access to features - point-in-time JOINs of features to create training data with no data leakage - easier connection and backfilling of features from external data sources - use of external tables as features - transparent computation of statistics and usage data for features. ## Write to feature groups, read from feature views You write to feature groups with a feature pipeline program. The program can be written in Python, Spark, or SQL. You read from views on top of the feature groups, called feature views. That is, a feature view does not store feature data, but is a logical grouping of features. Typically, you define a feature view because you want to train/deploy a model with exactly those features in the feature view. Feature views enable the reuse of feature data from different feature groups across different models. ================================================================================ # Features and Feature Groups Source: https://docs.hopsworks.ai/latest/concepts/fs/feature_group/fg_overview/ # Features and Feature Groups As a programmer, you can consider a feature, in machine learning, to be a variable associated with some entity that contains a value that is useful for helping train a model to solve a prediction problem. That is, the feature is just a variable with predictive power for a machine learning problem, or task. A feature group is a table of features. Each feature group has a primary key, and optionally an event_time column (indicating when the features in that row were observed), a partition key, and foreign keys that point to the primary keys of other feature groups. These are index columns, not features: they identify and join rows, and they are excluded when you select the features for a model. A feature group stores untransformed feature data, so the same feature can be reused across models that each transform it differently. ??? note "Partitioning" The partition key determines how the feature group rows are laid out on disk, so that queries using the partition key read only the data they need. For example, if the partition key is the day and you have hundreds of days of data, a query for a given day or a range of days reads only those days from disk. --8<-- "concepts/fs/feature_group/fg_overview/features-and-feature-groups.html" ## Online and offline Storage Feature groups can be stored in a low-latency "online" database and/or in low cost, high throughput "offline" storage, typically a data lake or data warehouse. A feature group with an embedding column can also have a vector index, for similarity search from inference pipelines and agents. --8<-- "concepts/fs/feature_group/fg_overview/online-and-offline-storage.html" ### Online Storage By default, the online store keeps only the latest values of features for a feature group. It serves those precomputed features to models at runtime, and is backed by [RonDB](https://www.rondb.com), a low latency, high throughput, high availability data store. By including an event_time column and a time-to-live (TTL), the online store can instead keep many rows per entity, which is what shift-right on-demand aggregations need. ### Offline Storage The offline store stores the historical values of features for a feature group so that it may store much more data than the online store. Offline feature groups are used, typically, to create training data for models, but also to retrieve data for batch scoring of models. In most cases, offline data is stored in Hopsworks, but through the implementation of data sources, it can reside in an external file system. The externally stored data can be managed by Hopsworks by defining ordinary feature groups or it can be used for reading only by defining [External Feature Group](external_fg.md). ================================================================================ # Write APIs Source: https://docs.hopsworks.ai/latest/concepts/fs/feature_group/write_apis/ # Write APIs You write to feature groups, and read from feature views. There are 3 APIs for writing to feature groups, as shown in the table below: | | Stream API | Batch API | Connector API | | --- | --- | --- | --- | | Python | X | - | - | | Spark | X | X | - | | External Table | - | - | X | ## Stream API The Stream API is the only API for Python clients, and is the preferred API for Spark, as it ensures consistent features between offline and online feature stores. The Stream API first writes data to be ingested to a Kafka topic, and then Hopsworks ensures that the data is synchronized to the Online and Offline Feature Groups through the OnlineFS service and Hudi DeltaStreamer jobs, respectively. The Kafka transport delivers at-least-once, and Hopsworks upgrades this to exactly-once through idempotent writes to the online feature group (only the latest values of features are stored there, and duplicates in Kafka only cause idempotent updates) and duplicate removal by Apache Hudi for the offline feature group. --8<-- "concepts/fs/feature_group/write_apis/stream-api.html" ## Batch API For very large updates to feature groups, such as when you are backfilling large amounts of data to an offline feature group, it is often preferential to write directly to the Hudi tables in Hopsworks, instead of via Kafka - thus reducing write amplification. Spark clients can write directly to Hudi tables on Hopsworks with Hopsworks libraries and certificates using a HDFS API. This requires network connectivity between the Spark clients and the datanodes in Hopsworks. --8<-- "concepts/fs/feature_group/write_apis/batch-api.html" ## Connector API Hopsworks supports external tables as feature groups. You can mount a table from an external database as an offline feature group using the Connector API: you create an external table using the connector, without ingesting the data into Hopsworks. This enables you to use features from your external data source (Snowflake, Redshift, Delta Lake, etc) as you would any feature in an offline feature group in Hopsworks. You can, for example, join features from different feature groups (external or not) together to create feature views and training data for models. See [External Feature Groups](external_fg.md) for the full list of supported data sources. ================================================================================ # External Feature Groups Source: https://docs.hopsworks.ai/latest/concepts/fs/feature_group/external_fg/ # External Feature Groups External feature groups are offline feature groups where their data is stored in an external table. An external table requires a data source, defined with the [Connector API](write_apis.md#connector-api) (or more typically in the user interface), to enable Hopsworks to retrieve data from the external table. An external feature group doesn't allow for offline data ingestion or modification; instead, it includes a user-defined SQL string for retrieving data. You can also perform SQL operations, including projections, aggregations, and so on. The SQL query is executed on-demand when Hopsworks retrieves data from the external Feature Group, for example, when creating training data using features in the external table. In the image below, we can see that Hopsworks currently supports a large number of data sources, including any JDBC-enabled source, Snowflake, Data Lake, Redshift, BigQuery, Databricks Unity Catalog (Delta tables on Databricks on AWS only), S3, ADLS, GCS, SQL, and Kafka. --8<-- "concepts/fs/feature_group/external_fg/external-feature-groups.html" ================================================================================ # Spine Group Source: https://docs.hopsworks.ai/latest/concepts/fs/feature_group/spine_group/ # Spine Group The default way to bring labels or prediction events is a label feature group: a regular feature group that holds the labels among its features, updated by a feature pipeline at a specific cadence. Sometimes it is more convenient to provide the training events or entities in a Dataframe instead, when reading feature data from the feature store through a feature view. We call such a Dataframe a Spine as it is the structure around which the training data or batch data is built. In order to retrieve the correct feature values for the entities in the Dataframe, using a point-in-time correct join, some additional metadata apart from the Dataframe schema is necessary. Namely, the information about which columns define the **primary key**, and which column indicates the **event time** at which the label was valid. The spine Dataframe together with this additional metadata is what we call a **Spine Group**. For example, in the following spine, we want to retrieve the features for the three locations, no later than the event time of each of the rainfall measurements, which is our prediction target: | location_id | event_time | rainfall (label) | | ----------- | ---------------- | -----------------| | 1 | 2022-06-01 13:11 | 44 | | 2 | 2022-06-01 09:14 | 5 | | 3 | 2022-06-01 06:36 | 2 | A Spine Group does not materialize any data to the feature store itself, and always needs to be provided when retrieving features from the [offline API](../feature_view/offline_api.md). You can think of it as a place holder or a temporary feature group, to be replaced by a Dataframe in point-in-time joins. When using the [online API](../feature_view/online_api.md), it is not necessary to provide the spine, since the online feature store contains only the latest feature values, and therefore no point in time join is required, the label is not required, as the inference pipeline is going to compute the prediction and the primary key values are specified when calling the online API. Prefer a label feature group where you can. A Spine Group adds complexity and pushes work onto the clients, which must supply the entities and their event times on every call, and it can only be the root or label feature group of a feature view. It is most appropriate for batch inference, where the set of entities to score is known only at request time. ================================================================================ # Feature Pipelines Source: https://docs.hopsworks.ai/latest/concepts/fs/feature_group/feature_pipelines/ # Feature Pipelines A feature pipeline is a program that orchestrates the execution of a dataflow graph of data validation, aggregation, dimensionality reduction, transformation, and other feature engineering steps on input data to create and/or update feature data. With Hopsworks, you can write feature pipelines in different languages as shown in the figure below. A feature pipeline can run on a schedule over a batch of data, or continuously over an event stream; see [Streaming Feature Pipelines](streaming_feature_pipelines.md). --8<-- "concepts/fs/feature_group/feature_pipelines/feature-pipelines.html" ## Data Sources Your feature pipeline needs to connect to some (external) data source to read the data to be processed. Python and Spark have connectors to a huge number of different data sources, while SQL feature pipelines are often restricted to a single data source (for example, your connector to SnowFlake only runs SQL on SnowFlake). SparkSQL, in contrast, can be used over tables that originate in different data sources. ## Data Validation In order to be able to train and serve models that you can rely on, you need clean, high quality features. Data validation operations include removing bad data, removing or imputing missing values, and identifying problems such as feature drift. Hopsworks supports Great Expectations to specify data validation rules that are executed in the client before features are written to the Feature Store. The validation results are collected and shown in Hopsworks. Data validation in ML is a shift-left property: data is validated before it is written to a feature group, since one bad data point could later fail a training or inference run. The default ingestion policy is STRICT, so a feature pipeline fails on a validation error rather than writing bad data. ## Aggregations Aggregations are used to summarize large datasets into more concise, signal-rich features. Popular aggregations include count(), sum(), mean(), median(), stddev(), min(), and max(). These aggregations produce a single number (a numerical feature) that captures information about a potentially large dataset. Both numerical and categorical features are often transformed before being used to train or serve models. ## Dimensionality Reduction If input data is impractically large or if it has a significant amount of redundancy, it can often be transformed into a reduced set of features with dimensionality reduction (often called feature extraction). Popular dimensionality algorithms include embedding algorithms, PCA, and TSNE. ## Transformations Transformations are covered in more detail in [training/inference pipelines](../feature_view/training_inference_pipelines.md), as transformations typically happen after the feature store. If you store transformed features in feature groups, the feature data is no longer useful for EDA (as it near to impossible for Data Scientists to understand the transformed values). It also makes it impossible for inference pipelines to log untransformed feature values and predictions for an operational model. There is one use case for storing transformed features in feature groups - when you need to have ultra low latency when reading precomputed features (and online transformations when reading features add too much latency for your use case). The figure below shows to include transformations in your feature pipelines. --8<-- "concepts/fs/feature_group/feature_pipelines/transformations.html" ## Feature Engineering in Python Python is the most widely used framework for feature engineering due to its extensive library support for aggregations (Pandas/Polars), data validation (Great Expectations), and dimensionality reduction (embeddings, PCA), and transformations (in Scikit-Learn, TensorFlow, PyTorch). Python also supports open-source feature engineering frameworks used for automated feature engineering, such as [featuretools](https://www.featuretools.com/) that supports relational and temporal sources. ## Feature Engineering in Spark/PySpark Spark is popular as a feature engineering framework as it can scale to process larger volumes of data than Python, and provides native support for aggregations, and it supports many of the same data validation (Great Expectations), and dimensionality reduction algorithms (embeddings, PCA) as Python. Spark also has native support for transformations, which are useful for analytical models (batch scoring), but less useful for operational models, where online transformations are required, and Spark environments are less common. Online model serving environments typically only support online transformations in Python. ## Feature Engineering in SQL SQL has grown in popularity for performing heavy lifting in feature pipelines - computing aggregates on data - when the input data already resides in a data warehouse. Data warehouses also support data validation, for example, through Great Expectations in DBT. However, SQL is not mature as a platform for transformations and dimensionality reductions, where UDFs are applied row-wise. You can do aggregation in SQL for data in your data warehouse or database. ## Feature Engineering in Beam Beam feature engineering pipelines are supported in Java/Scala only. ================================================================================ # Streaming Feature Pipelines Source: https://docs.hopsworks.ai/latest/concepts/fs/feature_group/streaming_feature_pipelines/ # Streaming Feature Pipelines A streaming feature pipeline processes an unbounded stream of events and keeps features fresh in near real-time, instead of running on a schedule over a batch of data. The same pipeline must also be able to run over historical data, to backfill a feature group when it is first created or after a schema change. This backfill-and-incremental duality is a defining property of a feature pipeline, not an afterthought. ## Feature freshness Feature freshness is the total time from when an event is first read by a feature pipeline to when the resulting feature is available to an inference pipeline. For interactive, real-time systems it is often the freshness of a feature, not the latency of the model, that decides whether a prediction is useful. --8<-- "concepts/fs/feature_group/streaming_feature_pipelines/freshness.html" ## Windows Streaming aggregations are computed over windows of the event stream: - **Tumbling** windows are fixed-size and non-overlapping, so each event falls in exactly one window. - **Hopping** windows are fixed-size but overlap, advancing by a hop smaller than the window. - **Rolling** (sliding) windows are recomputed continuously as events arrive. A watermark tells the pipeline how long to wait for late-arriving events before it closes a window and emits the aggregate. ## Streaming-native or hybrid A streaming-native pipeline computes all features directly on the stream, a Kappa-style architecture. A hybrid streaming-batch pipeline splits the work: a streaming job keeps the freshest features up to date while a batch job computes the heavier, less time-sensitive aggregations, a Lambda-style architecture. Prefer streaming-native where you can, since a single code path is simpler to keep consistent than two. A streaming feature pipeline can run in four operational modes: real-time processing of live events, stream replay, backfilling from historical data, and stream reprocessing after a logic change. ================================================================================ # Data Transformations Source: https://docs.hopsworks.ai/latest/concepts/mlops/data_transformations/ # Data Transformations [Data transformations](https://www.hopsworks.ai/dictionary/data-transformation) are integral to all AI applications. Data transformations produce new features that can enhance the performance of an AI application. However, [not all transformations in an AI application are equivalent](https://www.hopsworks.ai/post/a-taxonomy-for-data-transformations-in-ai-systems). Transformations like binning and aggregations typically create reusable features, while transformations like one-hot encoding, scaling and normalization often produce model-specific features. Additionally, in real-time AI systems, some features can only be computed during inference when the request is received, as they need request-time parameters to be computed. --8<-- "concepts/mlops/data_transformations/data-transformations.html" This classification of features can be used to create a taxonomy for data transformation that would apply to any scalable and modular AI system that aims to reuse features. The taxonomy helps identify which classes of data transformation can cause [online-offline](https://www.hopsworks.ai/dictionary/online-offline-feature-skew) skews in AI systems, allowing for their prevention. Hopsworks provides support for a feature view abstraction as well as model-dependent transformations and on-demand transformations to prevent online-offline skew. ## Data Transformation Taxonomy for AI Systems Transformation functions in an AI system can be classified into three types based on the nature of the input features they generate: [model-independent](https://www.hopsworks.ai/dictionary/model-independent-transformations), [model-dependent](https://www.hopsworks.ai/dictionary/model-dependent-transformations), and [on-demand](https://www.hopsworks.ai/dictionary/on-demand-transformation) transformations. --8<-- "concepts/mlops/data_transformations/transformation-taxonomy-2.html" **Model-independent transformations** create reusable features that can be utilized across one or more machine-learning models. These transformations include techniques such as grouped aggregations (e.g., minimum, maximum, or average of a variable), windowed aggregations (e.g., the number of clicks per day), and binning to generate categorical variables. Since the data produced by model-independent transformations are reusable, these features can be stored in a feature store. **Model-dependent transformations** generate features specific to one model. These include transformations that are unique to a particular model or are parameterized by the training dataset, making them model-specific. For instance, text tokenization is a transformation required by all large language models (LLMs) but each LLM has their own (unique) tokenizer. Other transformations, such as encoding categorical variables in a numerical representation or scaling/normalizing/standardizing numerical variables to enhance the performance of gradient-based models, are parameterized by the training dataset. Consequently, the features produced are applicable only to the model trained using that specific training dataset. Since these features are not reusable, there is no need to store them in a feature store. Also, storing encoded features in a feature store leads to write amplification, as every time feature values are written to a feature group, all existing rows in the feature group have to be re-encoded (and creation of a training dataset using a subset or rows in the feature group becomes impossible as they cannot be re-encoded). **On-demand transformations** are exclusive to [real-time AI systems](https://www.hopsworks.ai/dictionary/real-time-machine-learning), where predictions must be generated in real time based on incoming prediction requests. On-demand transformations compute on-demand features, which usually require at least one input parameter that is only available in a prediction request for their computation. These transformations can also combine request-time parameters with precomputed features from feature stores. Some examples include generating *zip_codes* from latitude and longitude received in the prediction request or calculating the *time_since_last_transaction* from a transaction request. The on-demand features produced can also be computed and [backfilled](https://www.hopsworks.ai/dictionary/backfill-features) into a feature store when the necessary historical data required for their computation becomes available. Backfilling on-demand features into the feature store eliminates the need to recompute them when creating training data. On-demand transformations are typically also model-independent transformations (model-dependent transformations can be applied after the on-demand transformation). Each of these transformations is employed within specific areas in a modular AI system and can be illustrated using the figure below. --8<-- "concepts/mlops/data_transformations/transformation-taxonomy.html" Model-independent transformations are utilized exclusively in areas where new and historical data arrives, typically within feature pipelines. Model-dependent transformations are necessary during the creation of training data, in training programs and must also be consistently applied in inference programs prior to making predictions. On-demand transformations are primarily employed in online inference programs, though they can also be integrated into feature engineering programs to backfill data into the feature store. The presence of model-dependent and on-demand transformations across different modules in a modular AI system introduces the potential for online-offline skew. Hopsworks provides support for model-dependent transformations and on-demand transformations to easily create modular skew-free AI pipelines. ## Hopsworks and the Data Transformation Taxonomy --8<-- "concepts/mlops/data_transformations/hopsworks-transformation-taxonomy-2.html" --8<-- "concepts/mlops/data_transformations/hopsworks-feature-store-storage.html" In Hopsworks, an AI system is typically decomposed into different [AI pipelines](https://www.hopsworks.ai/dictionary/ai-pipelines) and usually falls into either a [feature pipeline](https://www.hopsworks.ai/dictionary/feature-pipeline), a [training pipeline](https://www.hopsworks.ai/dictionary/training-pipeline), or an [inference pipeline](https://www.hopsworks.ai/dictionary/inference-pipeline). Hopsworks stores reusable feature data, created by model-independent transformations within the feature pipeline, into [feature groups](../fs/feature_group/fg_overview.md) (tables containing feature data in both offline and online stores). Model-independent transformations in Hopsworks can be performed using a wide range of commonly used data engineering tools and the generated features can be seamlessly inserted into feature groups. The figure below illustrates the different software tools supported by Hopsworks for creating reusable features through model-independent transformations. --8<-- "concepts/mlops/data_transformations/hopsworks-transformation-taxonomy.html" Additionally, Hopsworks provides a simple Python API to [create custom transformation functions](../../user_guides/fs/transformation_functions.md) as either Python or Pandas User-Defined Functions (UDFs). Pandas UDFs enable the vectorized execution of transformation functions, offering significantly higher throughput compared to Python UDFs for large volumes of data. They can also be scaled out across workers in a Spark program, allowing for scalability from gigabytes (GBs) to terabytes (TBs) or more. However, Python UDFs can be much faster for small volumes of data, such as in the case of online inference. Transformation functions defined in Hopsworks can then be attached to feature groups to [create on-demand transformation](../../user_guides/fs/feature_group/on_demand_transformations.md). On-demand transformations in feature groups are executed automatically whenever data is inserted into them to compute and backfill the on-demand features into the feature group. Backfilling on-demand features removes the need to recompute them while creating training and batch data. Hopsworks also provides a powerful abstraction known as [feature views](../fs/feature_view/fv_overview.md), which enables feature reuse and prevents skew between training and inference pipelines. A feature view is a meta-data-only selection of features, created from potentially different feature groups. It includes the input and output schema required for a model. This means that a feature view describes not only the input features but also the output targets, along with any helper columns necessary for training or inference of the model. This allows feature views to create consistent snapshots of data for both training and inference of a model. Additionally feature views, also compute and save statistics for the training datasets they create. Hopsworks supports attaching transformations functions to feature views to [create model-dependent transformations](../../user_guides/fs/feature_view/model-dependent-transformations.md) that have no online-offline skew. These transformations get access to the same training dataset statistics during both training and inference ensuring their consistency. Additionally, feature views through lineage get access to the on-demand transformation used to create on-demand features if any are selected during the creation of the feature view. The registration locus is the cleanest way to remember where each transformation lives: on-demand transformations are registered on feature groups, model-dependent transformations on feature views. A Hopsworks transformation function is also mixed-mode: the same decorated Python function runs as a Pandas UDF offline, to create training data, and as a Python UDF online, to build a single feature vector, so one definition serves both pipelines with no skew. This allows for the computation of on-demand features in real-time during online-inference. ================================================================================ # Overview Source: https://docs.hopsworks.ai/latest/concepts/fs/feature_view/fv_overview/ # Feature Views A feature view is a logical view over (or interface to) a set of features that may come from different feature groups. You create a feature view by selecting features, starting from a root feature group and following foreign keys to join in features from other feature groups. When the feature view has a label for supervised learning, the root feature group is the label feature group, the one feature group that holds the labels. Features are reachable by graph traversal: any feature group joined to the root can, in turn, have foreign keys to further feature groups whose features you can also select. A feature view does not have a primary key of its own; it has serving keys, the foreign keys of its label feature group, which you provide to retrieve feature vectors. In the illustration below, we can see that features are joined together from the two feature groups: seller_delivery_time_monthly and the seller_reviews_quarterly. You can also see that features in the feature view inherit not only the feature type from their feature groups, but also whether they are the primary key and/or the event_time. The image also includes transformation functions that are applied to individual features. Transformation functions are a part of the feature types included in the feature view. That is, a feature in a feature view is not only defined by its data type (int, string, etc) or its feature type (categorical, numerical, embedding), but also by its transformation. --8<-- "concepts/fs/feature_view/fv_overview/feature-views-2.html" Feature views can also include: - the label for the supervised ML problem - transformation functions that should be applied to specified features consistently between training and serving - the ability to create training data - the ability to retrieve a feature vector with the most recent feature values In the flow chart below, we can see the decisions that can be taken when creating (1) a feature view, and (2) creating training data with the feature view. --8<-- "concepts/fs/feature_view/fv_overview/feature-views.html" We can see here how the feature view is a representation for a model in the feature store - the same feature view is used to retrieve feature vectors for operational model that was created with training data from this feature view. As such, you can see that the most common use case for creating a feature view is to define the features that will be used in a model. In this way, feature views enable features from different feature groups to be reused across different models, and if features are stored untransformed in feature groups, they become even more reusable, as different feature views can apply different transformations to the same feature. ================================================================================ # Offline API Source: https://docs.hopsworks.ai/latest/concepts/fs/feature_view/offline_api/ # Offline API The feature view provides an *Offline API* for - creating training data - creating batch (scoring) data ## Training Data Training data is created using a feature view. You can create training data as either: - in-memory Pandas/Polars DataFrames, useful when you have a small amount of training data; - materialized training data in files, in a file format of your choice (such as `.tfrecord`, `.csv`, or `.parquet`). You can apply filters when creating training data from a feature view: - `start_time` and `end_time`, for example, to create the train-set from an earlier time range, and the test-set from a later (unseen) time range; - feature value features, for example, only train a model on customers from a particular country. Note that filters are not applied when retrieving feature vectors using feature views, as we only look up features for a specific entity, like a customer. In this case, the application should know that predictions for this customer should be made on the model trained on customers in USA, for example. Materialized training data is immutable: once created, a training dataset version is not appended to or modified. To retrain on new data, create a new training dataset version. If the new data needs to be computed continuously, for example a daily batch for a time-series model, do that computation once in a [derived feature group][assign-parents-to-a-feature-group] that is kept up to date as new data arrives, and create a new training dataset version from it whenever you need updated data. ### Point-in-time Correct Training Data When you create training data from features in different feature groups, it is possible that the feature groups are updated at different cadences. For example, maybe one feature group is updated hourly, while another feature group is updated daily. It is very complex to write code that joins features together from such feature groups and ensures there is no data leakage in the resultant training data. Hopsworks hides this complexity by performing the point-in-time JOIN transparently, similar to the illustration below: --8<-- "concepts/fs/feature_view/offline_api/point-in-time-correct-training-data-2.html" Hopsworks uses the event_time columns on both feature groups to determine the most recent (but not newer) feature values that are joined together with the feature values from the feature group containing the label. That is, the features in the feature group containing the label are the observation times for the features in the resulting training data, and we want feature values from the other feature groups that have the most recent timestamps, but not newer than the timestamp in the label-containing feature group. --8<-- "concepts/fs/feature_view/offline_api/point-in-time-correct-training-data.html" #### Spine Groups The left side of the point-in-time join is typically the set of training entities/primary key values for which the relevant features need to be retrieved. This left side of the join can also be replaced by a [spine group](../feature_group/spine_group.md). When using feature groups also so save labels/prediction targets, it can happen that you end up with the same entity multiple times in the training dataset depending on the cadence at which the label group was updated and the length of the event time interval that is being used to generate the training dataset. This can lead to bias in the training dataset and should be avoided. To avoid this kind of situation, users can either narrow down the event time interval during training dataset creation or use a spine in order to precisely define the entities to be included in the training dataset. This is just one example where spines are helpful. ### Splitting Training Data You can create random train/validation/test splits of your training data using the Hopsworks API. You can also time-based splits with the Hopsworks API. ### Evaluation Sets Test data can also be split into evaluation sets to help evaluate a model for potential bias. First, you have to identify the classes of samples that could be at risk of bias, and generate *evaluation sets* from your unseen test set - one evaluation set for each group of samples at risk of bias. For example, if you have a feature group of users, where one of the features is gender, and you want to evaluate the risk of bias due to gender, you can use filters to generate 3 evaluation sets from your test set - one for male, female, and non-binary. Then you score your model against all 3 evaluation sets to ensure that the prediction performance is comparable and non-biased across all 3 gender. ## Batch (Scoring) Data Batch data for scoring models is created using a feature view. Similar to training data, you can create batch data as either: - in-memory Pandas/Polars DataFrames, useful when you have a small amount of data to score; - materialized data in files, in a file format of your choice (such as `.tfrecord`, `.csv`, or `.parquet`) Batch data requires specification of a `start_time` for the start of the batch scoring data. You can also specify the `end_time` (default is the current date). --8<-- "concepts/fs/feature_view/offline_api/batch-scoring-data.html" ### Spine Dataframes Similar to training dataset generation, it might be helpful to specify a spine when retrieving features for batch inference. The only difference in this case is that the spine dataframe doesn't need to contain the label, as this will be the output of the inference pipeline. A typical use case is the handling of opt-ins, where certain customers have to be excluded from an inference pipeline due to a missing marketing opt-in. ================================================================================ # Online API Source: https://docs.hopsworks.ai/latest/concepts/fs/feature_view/online_api/ # Online API The Feature View provides an Online API to return an individual feature vector, or a batch of feature vectors, containing the latest feature values. To retrieve a feature vector, a client provides the feature view's serving keys. The serving keys are the foreign keys of the feature view's label feature group; a feature view does not have a primary key of its own. For example, if a feature view is built from the `customer_profile` and `customer_purchases` feature groups joined on `customer_id`, then `customer_id` is the serving key you provide to retrieve a feature vector. ## Feature Vectors A feature vector is a row of features (without the primary key(s) and event timestamp): --8<-- "concepts/fs/feature_view/online_api/feature-vectors.html" It may be the case that for any given feature vector, not all features will come pre-engineered from the feature store. Some features will be provided by the client (or at least the raw data to compute the feature will come from the client). We call these 'passed' features and, similar to precomputed features from the feature store, they can also be transformed by the Hopsworks client in the method: ```python feature_view.get_feature_vector(entry, passed_features={"pressure": 1013}) ``` When you call `get_feature_vector`, Hopsworks builds the vector in a fixed order: 1. retrieve the precomputed features from the online store using the serving keys, 2. merge in any passed features, 3. compute on-demand transformations (ODTs), 4. compute model-dependent transformations (MDTs), 5. drop the index and helper columns, 6. return the feature vector. This ordering is the composition constraint: on-demand transformations run before model-dependent ones, which are always last, just before the model is called. ================================================================================ # Consistent Transformations Source: https://docs.hopsworks.ai/latest/concepts/fs/feature_view/training_inference_pipelines/ # Consistent Transformations A *training pipeline* is a program that orchestrates the training of a machine learning model. For supervised machine learning, a training pipeline requires both features and labels, and these can typically be retrieved from the feature store as either in-memory Pandas/Polars DataFrames or read as training data files, created from the feature store. An *inference pipeline* is a program that takes user input, optionally enriches it with features from the feature store, and builds a feature vector (or batch of feature vectors) with with it uses a model to make a prediction. ## Transformations Feature transformations are mathematical operations that change feature values with the goal of improving model convergence or performance properties. Transformation functions take as input a single value (or small number of values), they often require state (such as the mean value of a feature to normalize the input), and they output a single value or a list of values. ## Offline-Online Feature Skew Offline-online feature skew is a difference between the transformation code that runs in an offline (training) pipeline and the transformation code that runs in the corresponding inference pipeline. It is a code difference, not a data difference, so it cannot be detected by comparing distributions; the only way to avoid it is to run the same code in both pipelines. In the image below, you can see that transformations happen after the Feature Store, but that the implementation of the transformation functions need to be consistent between the training and inference pipelines. --8<-- "concepts/fs/feature_view/training_inference_pipelines/offline-online-feature-skew.html" There are 3 main approaches to prevent offline-online feature skew that we support in Hopsworks. These are (1) perform transformations in models, (2) perform transformations in pipelines (sklearn, TF, PyTorch) and use the model registry to save the transformation pipeline so that the same transformation is used in your inference pipeline, and (3) use Hopsworks transformations, defined as UDFs in Python. ### Transformations as Pre-Processing Layers in Models Transformation functions can be implemented as preprocessing steps within a model. For example, you can write a transformation function as a pre-processing layer in Keras/TensorFlow. When you save the model, the preprocessing steps will also be saved as part of the model. Any state required to compute the transformation, such as the arithmetic mean of a numerical feature in the train set, is also stored with the function, enabling consistent transformations during inference. When data preprocessing is part of the model, users can just send the untransformed feature values to the model and the model itself will apply any transformation functions as preprocessing layers (such as encoding categorical variables or normalizing numerical variables). ### Transformation Pipelines in Scikit-Learn/TensorFlow/PyTorch You have to save your transformation pipeline (serialize the object or the parameters) and make sure you apply exactly the same transformations in your inference pipeline. This means you should version the transformations. In Hopsworks, you can store the transformations with your versioned models in the Model Registry, helping you to ensure the same transformation pipeline is applied to both training/serving for the same model version. ### Transformations as Python UDFs in Hopsworks Hopsworks feature store also supports consistent transformation functions by enabling a Python UDF, that implements a transformation, to be attached a to feature in a feature view. When training data is created with a feature view or when a feature vector is retrieved from a feature view, Hopsworks ensures that any transformation functions defined over any features will be applied before returning feature values. You can use built-in transformation objects in Hopsworks or write your own custom transformation functions as Python UDFs. The benefit of this approach is that transformations are applied consistently when creating training data and when retrieving feature data from the online feature store. Transformations no longer need to be included in either your training pipeline or inference pipeline, as they are applied transparently when creating training data and retrieving feature vectors. Hopsworks uses Spark to create training data as files, and any transformation functions for features are executed as Python UDFs in Spark - enabling transformation functions to be applied on large volumes of data and removing potentially CPU-intensive transformations from training pipelines. ================================================================================ # On-Demand Features Source: https://docs.hopsworks.ai/latest/concepts/fs/feature_group/on_demand_feature/ # On-demand features Features are defined as on-demand when their value cannot be pre-computed beforehand, rather they need to be computed in real-time during inference. This is achieved by implementing the on-demand features as a Python function in a Python module. Also ensure that the same version of the Python module is installed in both the feature and inference pipelines. The figure below shows an example from a housing price model: a zip code (or post code) computed on demand from longitude and latitude parameters. In your online application, longitude and latitude are provided as parameters to the application, and the same python function used to calculate the zip code in the feature pipeline is used to compute the zip code in the Online Inference pipeline. --8<-- "concepts/fs/feature_group/on_demand_feature/on-demand-features.html" ## Shift left or shift right Deciding to compute a feature on-demand is a shift-right decision, and it is one of the biggest feature-engineering choices you make. Shift left means precomputing a feature in a feature pipeline and storing it in the feature store for retrieval. Shift right means computing it at request time, in an on-demand or model-dependent transformation. Shift right when the feature depends on request-time input, such as the zip code computed from the longitude and latitude in a request, or when a precomputed value would be too stale to be useful. Shift left when the feature can be precomputed, to keep inference latency low and avoid repeating the computation on every request. The trade-off is latency and operational overhead against freshness. ================================================================================ # Data Validation, Statistics, Alerts Source: https://docs.hopsworks.ai/latest/concepts/fs/feature_group/fg_statistics/ # Data Validation, Statistics, and Alerts Hopsworks supports monitoring, validation, and alerting for features: - transparently compute statistics over features on writing to a feature group; - validation of data written to feature groups using Great Expectations - alerting users when there was a problem writing or update features. ## Statistics When you create a Feature Group in Hopsworks, you can configure it to compute statistics over the features inserted into the Feature Group by setting the `statistics_config` dict parameter, see [Feature Group Statistics](../../../user_guides/fs/feature_group/statistics.md) for details. Every time you write to the Feature Group, new statistics will be computed over all of the data in the Feature Group. ## Data Validation You can define expectation suites in Great Expectations and associate them with feature groups. When you write to a feature group, the expectations are executed, then you can define a policy on the feature group for what to do if any expectation fails. --8<-- "concepts/fs/feature_group/fg_statistics/data-validation.html" ## Alerting Hopsworks also supports alerts, that can be triggered when there are problems in your feature pipelines, for example, when a write fails due to an error or a failed expectation. You can send alerts to different alerting endpoints, such as email or Slack, that can be configured in the Hopsworks UI. For example, you can send a slack message if features being written to a feature group are missing some input data. ================================================================================ # Feature Monitoring Source: https://docs.hopsworks.ai/latest/concepts/fs/feature_group/feature_monitoring/ # Feature Monitoring Feature monitoring complements data validation by letting you monitor feature data after it has been ingested into the feature store. It computes statistics over a detection window of data, compares them against a reference window, and raises alerts when the comparison crosses a threshold. The comparison can be a single scalar metric (e.g., the mean) or the whole feature distribution using a distance metric such as PSI or KL divergence. You can monitor at two levels, and each level detects a different kind of change. ## Monitoring a feature group Monitoring a feature group watches the raw data as it is ingested, independently of any model. The reference window is usually an earlier window of the same feature group, so what you detect is data ingestion drift: a new batch that no longer looks like the data already in the feature group. After creating a feature group, you can schedule statistics over one or more features, computed on the whole feature data or on a subset defined by a detection window. You then enable a comparison against a reference window and define the criteria: which statistic to compare and the threshold that flags an anomaly. ## Monitoring a feature view Monitoring a feature view watches what a specific model actually sees, because a feature view backs the features served to a model. Here the reference window is typically the model's training dataset, so what you detect is feature drift: the served features drifting away from the distribution the model was trained on. The mechanism is the same scheduled statistics and distribution comparison as for a feature group, computed using the feature view query; only the reference changes. Comparing a model's logged inference data against its training dataset, and deciding when to retrain, is model monitoring; see [Model Monitoring](../../mlops/model_monitoring.md). ## Statistics on training data A feature view holds no statistics of its own, since it is only an interface over features and their transformations. Statistics are computed over a training dataset instead. Those training-dataset statistics are the reference that feature-view monitoring compares against, and some online transformations need them too (normalizing a numerical feature requires the training-set mean). !!! info "Feature Monitoring Guide" More information can be found in the [Feature monitoring guide](../../../user_guides/fs/feature_monitoring/index.md). ================================================================================ # Versioning Source: https://docs.hopsworks.ai/latest/concepts/fs/feature_group/versioning/ # Versioning Hopsworks versions the ML assets that make up an AI system, so that a model in production is reproducible and clients are protected from breaking changes. Feature groups, feature views, training data, and models are versioned; deployments are the one asset that is not. ## Feature group schema versioning The schema of feature groups is versioned. If you make a breaking change to the schema of a feature group, you need to increment the version of the feature group, and then backfill the new feature group. A breaking schema change is when you: - drop a column from the schema - add a new feature without any default value for the new feature - change how a feature is computed, such that, for training models, the data for the old feature is not compatible with the data for the new feature. For example, if you have an embedding as a feature and change the algorithm to compute that embedding, you probably should not mix feature values computed with the old embedding model with feature values computed with the new embedding model. --8<-- "concepts/fs/feature_group/versioning/feature-group-schema-versioning.html" ## Feature group data versioning Data versioning of a feature group tracks updates to the feature group, so that you can recover the state of the feature group at a given point-in-time in the past. --8<-- "concepts/fs/feature_group/versioning/feature-group-data-versioning.html" There are two points in time you can travel back to, and they answer different questions. As-of ingestion time reads the data as it had been written by a given moment, which gives reproducible training data. As-of event time reads the data as it was true in the world at a given moment, which gives point-in-time correct training data with no future leakage. ## Feature view and training data versioning Feature views are interfaces, and if there is a change in the interface (the types of the features, the transformations applied to the features), then you need to change the version, to prevent breaking existing clients. Training datasets are associated with a specific feature view version, and each training dataset also has its own version number. For example, online transformation functions often need training data statistics (e.g., normalizing a numerical feature requires you to divide the feature value by the mean value for that feature in the training dataset). As many training datasets can be created from a feature view, when you initialize the feature view you need to tell it which version of the training data to use: `feature_view.init(1)` means use version 1 of the training data for this feature view. --8<-- "concepts/fs/feature_group/versioning/feature-view-training-data-versioning.html" ## Models and deployments A model has its own version in the model registry. A deployment, however, is not versioned: it is the one mutable asset. A new deployment gets a new name, upgrades and rollbacks are done with blue/green deployments, and clients depend on the [deployment API](../../mlops/serving.md), not on a deployment version number. A model deployment is also tightly coupled to the versioned feature views that supply its pre-computed features, so versioning the model alone is not enough. ================================================================================ # Tags, Search, Lineage Source: https://docs.hopsworks.ai/latest/concepts/projects/search/ # Tags, Search, and Lineage ## Search { #search-concept } Hopsworks supports free-text search to discover machine-learning assets: - features - feature groups - feature views - training data - jobs, including apps - models - deployments, including agents You can use the search bar at the top of your project to free-text search for the names or descriptions of any ML asset. You can also search using keywords or tags that are attached to an ML asset. Tags are indexed for every asset in the list above, so a governance question can be asked once across the whole set rather than per asset type. Keywords apply to feature groups, feature views and training data only, so a keyword filter never matches a job, a model or a deployment. Apps and agents are reported as their own classes but are not stored as their own kind of asset. An app is a job whose type is PythonApp, and an agent is a deployment serving no registered model. Each is a narrowing of the class it belongs to, which is why they need no separate index and appear the moment the distinguishing property does. You can search for assets within a specific project or across all projects in a Hopsworks deployment, including those you are not a member of. This allows for easier discoverability and reusability of assets within an organization. To avoid users gaining unauthorized access to data, if a search result is in a project you are **not** a member of, the information displayed is limited to: names, descriptions, tags, asset creator and create date. If the search result is within a project you are a member of, you are also able to inspect recent activities on the asset as well as statistics. Searching across projects you are not a member of is on by default. An administrator running a multi-tenant deployment can turn it off, which restricts every search to the projects the caller can already access. See [search index administration][search-index-administration] for the setting. For how to use search, including filtering by a specific tag key and value, see the [search guide][search-guide]. ## Tags A keyword is a single user-defined word attached to an ML asset. Keywords can be used to help it make it easier to find ML assets or understand the context in which they should be used, for example, *PII* could be used to indicate that the ML asset is based on personally identifiable information. However, it may be preferable to have a stronger governance framework for ML assets than keywords alone. For this, you can define a *schematized tag*, defining a list of key/value tags along with a type for a value. In the figure below, you can see an example of a schematized tag with two key/value pairs: *pii* of type boolean (indicating if this feature group contains PII data), and *owner* of type string (indicating who the owner of the data in this feature group is). Keywords are not part of a schematized tag and are not shown in this panel. They are attached in the feature group header instead, for example a *eu_region* keyword indicating the data has its origins in the EU. Schematized tags can also enforce policy, not just aid discovery. You could, for example, require that a model cannot be created in the production model registry unless its EU AI Act tag is filled in correctly. ## Lineage Hopsworks tracks the lineage (or provenance) of ML assets automatically for you. The lineage chain runs end to end: data source, feature group, feature view, training data, model, deployment. This is what makes governance answerable: from a biased model you can trace back to the exact feature groups and data sources that fed it. You can see what features are used in which feature view or training dataset, and what training dataset was used to train a given model. For assets that are managed outside of Hopsworks, there is support for the explicit definition of lineage dependencies. --8<-- "concepts/projects/search/provenance-lineage.html" The lineage of an asset is shown in the Hopsworks UI, here the feature groups behind a feature view and the training data and models derived from it. ================================================================================ # CI/CD Source: https://docs.hopsworks.ai/latest/concepts/projects/cicd/ # CI/CD Support You can setup traditional development, staging, and production environment in Hopsworks using Projects. A project enables you provide access control for the different environments - just like a GitHub repository, owners of projects can add and remove members of projects and assign different roles to project members - the "data owner" role can write to feature store, while a "data scientist" can only read from the feature store and create training data. ## Dev, Staging, Prod You can create dev, staging, and prod projects - either on the same cluster, but mostly commonly, with production on its own cluster: --8<-- "concepts/projects/cicd/dev-staging-prod.html" ## Versioning Automated promotion across dev, staging, and prod relies on every ML asset being versioned. Hopsworks versions feature groups, feature views, training data, and models, while deployments stay mutable behind the deployment API. See [Versioning](../fs/feature_group/versioning.md) for what is versioned and how. ## Pytest for feature logic and feature pipeline tests Pytest and Great Expectations can be used for testing feature pipelines. Pytest is used to test feature logic and for end-to-end feature pipeline tests, while Great Expectations is used for data validation tests. Here, we can see how a feature pipeline test uses sample data to compute features and validate they have been written successfully, first to a development feature store, and then they can be pushed to a staging feature store, before finally being promoted to production. --8<-- "concepts/projects/cicd/pytest-feature-logic.html" ================================================================================ # Model Training Source: https://docs.hopsworks.ai/latest/concepts/mlops/training/ # Model Training A training pipeline is a program that orchestrates the training of a machine learning model, reading features and labels from the feature store as training data. Hopsworks supports running model training pipelines on any Python environment, whether on an external Python client or on a Hopsworks cluster. The outputs of a training pipeline are typically experiment results, including logs, and possibly a trained model. You can plugin your own experimentation tracking platform or model registry, or you can use Hopsworks. A training pipeline typically runs five steps: select a feature view and a training dataset version, train the model, evaluate it, validate it, and register it in the model registry if it passes. ## Evaluation and validation Model evaluation and model validation are not the same thing. Evaluation measures the model's performance on a held-out test set, using metrics such as accuracy or AUC. Validation is a pass/fail gate: the model is run against evaluation data, including bias slices of the holdout built with feature-view filters and training helper columns (a column such as gender used to slice results but dropped before training), and only a model that passes is registered. The output of validation is a model validation scorecard, and it is what decides whether the model reaches the registry. ## Training Pipelines on Hopsworks If you train models with Hopsworks, you can setup CI/CD pipelines as shown below, where the experiments are tracked by Hopsworks, and any model created is published to a model registry. Each project has its own private model registry, so when you are working in a development project, you typically publish models to your project's private development registry, and if all model validation tests pass, and the model performance is good enough, the same training pipeline can be submitted via a CI/CD pipeline (e.g., GitHub push request) to a staging project, and the same procedure can be repeated to push the training pipeline to a production project. --8<-- "concepts/mlops/training/training-pipelines-on-hopsworks.html" Hopsworks [Model Registry](registry.md) and [Model Serving](serving.md) capabilities can then be used to build a batch or online prediction service using the model. ================================================================================ # Model Registry Source: https://docs.hopsworks.ai/latest/concepts/mlops/registry/ # Model Registry Hopsworks Model Registry is designed with specific support for KServe and MLOps, through versioning. It enables developers to publish, test, monitor, govern and share models for collaboration with other teams. The model registry is where developers publish their models during the experimentation phase. The model registry can also be used to share models with the team and stakeholders. Like other project-based multi-tenant services in Hopsworks, a model registry is private to a project. That means you can easily add a development, staging, and production model registry to a cluster, and implement CI/CD processes for transitioning a model from development to staging to production. The model registry for KServe's capability are shown in the diagram below: --8<-- "concepts/mlops/registry/model-registry.html" The model registry centralizes model management, enabling models to be securely accessed and governed. Models are more than just the model itself - the registry also stores sample data for testing, configuration information, provenance information, environment variables, links to the code used to generate the model, the model version, and tags/descriptions). When you save a model, you can also save model metrics with the model, enabling users to understand, for example, performance of the model on test (or unseen) data. ## Model Package A ML model consists of a number of different components in a model package: - Model Input/Output Schema - Model artifacts - Model version information - Model format (based on the ML framework used to train the model - e.g., .pkl or .tb files) You can also optionally include in your packaged model: - Sample data (used to test the model in KServe) - The source notebook/program/experiment used to create the model ================================================================================ # Model Serving Source: https://docs.hopsworks.ai/latest/concepts/mlops/serving/ # Model Serving In Hopsworks, you can easily deploy models from the model registry using [KServe](https://kserve.github.io/website/latest/), the standard open-source framework for model serving on Kubernetes. You rarely deploy just a model. What you deploy is an online inference pipeline, of which the model is one part, alongside feature retrieval, transformations, and logging. You can deploy models programmatically using [`Model.deploy`][hsml.model.Model.deploy] or via the UI. A KServe model deployment can include the following components: **`Predictor (KServe component)`** : A predictor runs a model server (Python, TensorFlow Serving, or vLLM) that loads a trained model, handles inference requests and returns predictions. **`Transformer (KServe component)`** : A ^^pre-processing^^ and ^^post-processing^^ component that can transform model inputs before predictions are made, and predictions before these are delivered back to the client. Not available for vLLM deployments. **`Inference Logger`** : Hopsworks logs inputs and outputs of transformers and predictors to a ^^Kafka topic^^ that is part of the same project as the model. This is for storing inference requests and responses for later consumption and analysis, and is separate from the feature logging that powers [Model Monitoring](model_monitoring.md). Not available for vLLM deployments. **`Inference Batcher`** : Inference requests can be batched to improve throughput (at the cost of slightly higher latency). **`Istio Model Endpoint`** : You can publish a model over REST(HTTP) or gRPC using a Hopsworks API key, accessible via **path-based routing** through Istio. API keys have scopes to ensure the principle of least privilege access control to resources managed by Hopsworks. For more details on path-based routing of requests through Istio, see [REST API Guide](../../user_guides/mlops/serving/rest-api.md). !!! warning "Host-based routing" The Istio Model Endpoint supports host-based routing for inference requests; however, this approach is considered legacy. Path-based routing is recommended for new deployments. Models deployed on KServe in Hopsworks can be easily integrated with the Hopsworks Feature Store using either a Transformer or Predictor Python script, that builds the predictor's input feature vector using the application input and pre-computed features from the Feature Store. --8<-- "concepts/mlops/serving/model-serving.html" ## Deployment API The deployment API is the interface to the online inference pipeline that clients send prediction requests to. It is the deployment API, not the model signature, that clients should version against. The model signature (the input and output schema of the model) changes whenever you retrain with a different set of features, so coupling clients to it turns every model update into a breaking change. The deployment API is a stable contract that can stay the same across model versions. A client request to the deployment API carries two kinds of parameter: - **serving keys**: the entity IDs used to retrieve pre-computed features from the feature store. - **request parameters**: values known only at request time, sent in the request and used to build the feature vector or to compute on-demand features. Because clients depend on it, a deployment API should carry an SLO, typically a p99 latency target for online predictions. ## Testing model deployments Two release-safety mechanisms are often confused, because they test different things. A blue/green test tests the correctness and performance of the model deployment directly, running the new deployment alongside the old one so clients can be switched over with no risk. An A/B test does not test the deployment; it tests the model's effect on the application, measured against an application KPI, to decide whether the new model actually makes the product better. !!! info "Model Serving Guide" More information can be found in the [Model Serving guide](../../user_guides/mlops/serving/index.md). !!! tip "Python deployments" For deploying Python scripts without a model artifact, see the [Python Deployments](../../user_guides/projects/python-deployment/python-deployment.md) page. ================================================================================ # Model Monitoring Source: https://docs.hopsworks.ai/latest/concepts/mlops/model_monitoring/ # Model Monitoring Model monitoring lets you track how a deployed model behaves in production by comparing the data it serves against the data it was trained on. When a model runs in production, the statistical properties of its inputs and predictions can drift away from those of the training data. This degrades model quality silently, without any error being raised. Model monitoring detects this drift early so you can decide whether to retrain the model. ## How it works Model monitoring builds on two existing Hopsworks capabilities: - **Feature logging**: a model deployment logs the features it serves and its predictions to the feature view's logging feature group through the Feature View logging APIs. See the [Feature Logging guide](../../user_guides/fs/feature_view/feature_logging.md). - **Feature monitoring**: Hopsworks computes statistics over windows of feature data and compares them against a reference, optionally raising alerts on significant drift. See the [Feature Monitoring concept](../fs/feature_group/feature_monitoring.md). ??? note "Log untransformed and transformed features" Log both the untransformed and the transformed feature values. Untransformed features drive feature monitoring and debugging, since drift is easiest to read on the raw values. Transformed features drive model monitoring and SHAP explainability, since those are the values the model actually sees. !!! info "Feature logging vs. the inference logger" Hopsworks provides two separate inference logging mechanisms. The [inference logger](../../user_guides/mlops/serving/inference-logger.md) stores the model inputs and predictions from inference requests and responses into Kafka, for later consumption and analysis. [Feature logging](../../user_guides/fs/feature_view/feature_logging.md) supports more fine-grained logging of inference logs and features, enabling feature monitoring and model monitoring. Model monitoring relies on feature logging, not on the inference logger. A model monitoring configuration is a feature monitoring configuration over the logging feature group, filtered to a single model and version. The detection window covers the recently served inference data, and the reference defaults to the training dataset version that was used to train that model. By comparing the two, on a scalar metric or on the whole feature distribution, Hopsworks detects feature drift over time. This comparison detects drift, not skew: offline-online feature skew is a difference in the transformation code between the offline and inference pipelines, so it is invisible to a distribution comparison and is prevented, not monitored. Feature drift is one kind of drift among several. Concept drift, in particular, is not detected by comparing distributions: you detect it by comparing the actual outcomes against the model's past predictions, once those outcomes are known. --8<-- "concepts/mlops/model_monitoring/how-it-works.html" ## Where to configure it Because monitoring is anchored on the feature view that backs the model, you can configure model monitoring from whichever entity is most convenient: - a **model deployment**, when operating a model in production. - a **model** in the model registry. - a **feature view**, when working directly with the feature data. All three resolve to the same underlying configuration. !!! info "Model Monitoring Guide" More information can be found in the [Model Monitoring guide](../../user_guides/mlops/model_monitoring/index.md). ================================================================================ # Vector Index Source: https://docs.hopsworks.ai/latest/concepts/mlops/opensearch/ # Vector Index A vector index stores embeddings so you can retrieve the items most similar to a query vector, the retrieval half of a recommender or a RAG system. In Hopsworks, a vector index is a property of an online-enabled feature group: a feature group with an embedding column can be indexed for similarity search, alongside its online and offline stores. The vector index is backed by OpenSearch, included as a multi-tenant service in projects. OpenSearch provides the index through its k-NN plugin, which supports several engines for embedding indexes. Hopsworks creates its indexes on the FAISS engine, which is the default from Hopsworks 5.1. Earlier releases used the nmslib engine, which OpenSearch has deprecated and which does not accept the filter that Hopsworks 5.1 and later send inside the nearest-neighbor query, so the upgrade to Hopsworks 5.2 recreates those indexes on FAISS. The [OpenSearch upgrade guide][upgrading-opensearch] describes what that upgrade involves. Through Hopsworks, OpenSearch also provides enterprise capabilities, including authentication and access control to indexes (an index can be private to a Hopsworks project), filtering, scalability, high availability, and disaster recovery support. To learn how OpenSearch powers vector similarity search in Hopsworks, you can see [this guide](../../user_guides/fs/vector_similarity_search.md). --8<-- "concepts/mlops/opensearch/vector-index.html" ================================================================================ # Development Inside Hopsworks Source: https://docs.hopsworks.ai/latest/concepts/dev/inside/ # Development Inside Hopsworks Hopsworks provides a complete self-service development environment for feature engineering and model training. You can develop programs as Jupyter notebooks or jobs, customize the bundled FTI (feature, training and inference pipeline) python environments, you can manage your source code with Git, and you can orchestrate jobs with Airflow. A browser terminal runs inside the project with the Hopsworks CLI and coding agents preinstalled, and the Wizard uses it to build a system end to end from a few choices. --8<-- "concepts/dev/inside/development-inside-hopsworks.html" ## Jupyter Notebooks Hopsworks provides a Jupyter notebook development environment for programs written in Python, Spark, and SparkSQL. You can also develop in your IDE (PyCharm, IntelliJ, etc), test locally, and then run your programs as Jobs in Hopsworks. Jupyter notebooks can also be run as Jobs. ## Source Code Control Hopsworks provides source code control support using Git (GitHub, GitLab or BitBucket). You can securely check out code into your project and commit and push updates to your code to your source code repository. ## FTI Pipeline Environments Hopsworks postulates that building ML systems following the FTI pipeline architecture is best practice. This architecture consists of three independently developed and operated ML pipelines: - Feature pipeline: takes as input raw data that it transforms into features (and labels) - Training pipeline: takes as input features (and labels) and outputs a trained model - Inference pipeline: takes new feature data and a trained model and makes predictions In order to facilitate the development of these pipelines Hopsworks bundles several python environments containing necessary dependencies. Each of these environments may then also be customized further by cloning it and installing additional dependencies from PyPi, Wheel files, GitHub repos or a custom Dockerfile. Internal compute such as Jobs and Jupyter is run in one of these environments and changes are applied transparently when you install new libraries using our APIs. That is, there is no need to write a Dockerfile, users install libraries directly in one or more of the environments. You can setup custom development and production environments by creating separate projects or creating multiple clones of an environment within the same project. ## Jobs In Hopsworks, a Job is a schedulable program that is allocated compute and memory resources. You can run a Job in Hopsworks: - From the UI - Programmatically with the Hopsworks SDK (Python, Java) or REST API - From Airflow programs (either inside our outside Hopsworks) - From your IDE using a plugin ([PyCharm/IntelliJ plugin](https://plugins.jetbrains.com/plugin/15537-hopsworks)) ## Orchestration Airflow comes out-of-the box with Hopsworks, but you can also use an external Airflow cluster (with the Hopsworks Job operator) if you have one. Airflow can be used to schedule the execution of Jobs, individually or as part of Airflow DAGs. ## Terminal { #inside-terminal } Every project has a browser terminal: a shell running in a pod under your project user, with your HopsFS home mounted, and `hops`, `git`, Claude Code and Codex preinstalled and already connected to the project. It is the fastest way to work with a project from inside Hopsworks, and the seat the Wizard drives. See the [Terminal guide][terminal] and the [Hopsworks CLI guide][hopsworks-cli]. ## Wizard { #inside-wizard } The Wizard asks what you want to build, where the data comes from and what to predict, then writes a kickoff prompt and hands it to Claude in the terminal, which builds the feature pipeline, the model and the dashboard with `hops`. See the [Wizard guide][wizard]. ================================================================================ # Development Outside Hopsworks Source: https://docs.hopsworks.ai/latest/concepts/dev/outside/ # Development Outside Hopsworks You can write programs that use Hopsworks in any [Python, Spark, or PySpark environment](../../user_guides/integrations/index.md). Hopsworks also supports running SQL queries to compute features in external data warehouses. The Feature Store can also be queried with SQL. There is REST API for Hopsworks that can be used with a valid API key, generated in Hopsworks. However, it is often easier to develop your programs against the Hopsworks SDK, available in Python and Java/Scala, which covers the feature store, the model registry and model serving. The same library ships the `hops` command line, so a shell, a CI pipeline or a coding agent on your machine can read and write the project with the same API key; see the [Hopsworks CLI guide][hopsworks-cli]. --8<-- "concepts/dev/outside/development-outside-hopsworks.html" ================================================================================ # BI Tools Source: https://docs.hopsworks.ai/latest/concepts/mlops/bi_tools/ # BI Tools Feature groups have well-defined schemas and live in two stores, so any BI tool that speaks SQL can analyze features and build reports on them. - The offline store is queried through the [Query Engine](../../user_guides/projects/trino/query_engine.md), Trino, with one catalog per table format (`delta`, `iceberg`, `hudi`). Any tool with a Trino connector (JDBC or ODBC) can read it. - The online store, RonDB, is queried over the MySQL protocol, so any tool with a MySQL connector can read the latest feature values. Hopsworks bundles [Apache Superset](https://superset.apache.org/) as a project service, already connected to the project's Trino catalogs. Dashboards live inside the project and follow its access control.
Superset dashboards listed inside a Hopsworks project
Superset dashboards inside a project, with their public and shared status.
SQL Lab in Superset runs directly against the feature store: pick the project's Trino connection, then the catalog matching the feature group's table format.
Superset SQL Lab with the project's Trino connection and the delta, hudi and iceberg catalogs
SQL Lab on the project's Trino connection, one catalog per table format.
See the [Superset guide](../../user_guides/projects/superset/superset.md) for building dashboards on feature data, and the [Superset setup](../../setup_installation/admin/superset.md) page for enabling it on a cluster. ================================================================================ # How-To Guides Source: https://docs.hopsworks.ai/latest/user_guides/ # How-To Guides Task-focused guides for the Hopsworks UI and APIs, organised by the part of the platform you are working with. For what things are and why, see the [Concepts](../concepts/index.md).
- :material-console:{ .lg .middle } **Start here** --- Install the client and authenticate once. Every guide in this section runs from the same session. ```bash uv venv && source .venv/bin/activate uv pip install "hopsworks[python]" hops setup ``` [Client installation](client_installation/index.md) · [Create a project](projects/project/create_project.md) · [Create a feature group](fs/feature_group/create.md)
:material-database:{ .hops-role-ico } Feature Store { .hops-role-cap } - [Feature groups](fs/feature_group/index.md) Write features from a DataFrame, validate them, keep statistics. - [Feature views](fs/feature_view/index.md) Read training data, batch data and online feature vectors. - [Data sources](fs/data_source/index.md) Connect warehouses, object stores and databases as inputs. - [Feature monitoring](fs/feature_monitoring/index.md) Watch statistics over time and compare them to a reference. - [Transformations and integrations](fs/transformation_functions.md) Model-independent transformations, compute engines, external clients.
:material-rocket-launch-outline:{ .hops-role-ico } MLOps { .hops-role-cap } - [Model registry](mlops/registry/index.md) Register models with metrics, schema and evaluation artifacts. - [Model serving](mlops/serving/index.md) Deploy a model with a predictor, transformer, logging and autoscaling. - [Model monitoring](mlops/model_monitoring/index.md) Compare inference data against training data on a schedule. - [Agents](agents/index.md) Run agent tasks as jobs or serve interactive agents.
:material-folder-outline:{ .hops-role-ico } Projects and compute { .hops-role-cap } - [Projects](projects/index.md) Sign in, create a project, manage members, secrets, keys and alerts. - [Compute](compute/index.md) Jupyter, the terminal, jobs, Airflow and Python environments. - [Analytics](analytics/index.md) SQL over the offline store with Trino, dashboards in Superset.
:material-wrench-outline:{ .hops-role-ico } Platform { .hops-role-cap } - [Clients](client_installation/index.md) Python and Java libraries for your own environment, and the `hops` CLI. - [Setup and administration](../setup_installation/index.md) Install on a cloud or on-prem, manage users and operations. - [Migration 3.x to 4.0](migration/40_migration.md) What changed and how to move.
================================================================================ # Client Installation Guide Source: https://docs.hopsworks.ai/latest/user_guides/client_installation/ # Client Installation Guide ## Hopsworks Python library The Hopsworks Python client library is required to connect to Hopsworks from your local machine or any other Python environment such as Google Colab or AWS Sagemaker. Execute the following command to install the Hopsworks client library in your Python environment: !!! note "Virtual environment" It is recommended to use a virtual python environment instead of the system environment used by your operating system, in order to avoid any side effects regarding interfering dependencies. !!! attention "Windows/Conda Installation" On Windows systems you might need to install twofish manually before installing hopsworks, if you don't have the Microsoft Visual C++ Build Tools installed. In that case, it is recommended to use a conda environment and run the following commands: ```bash conda install twofish pip install hopsworks[python] ``` === "uv" ```bash uv venv && source .venv/bin/activate uv pip install "hopsworks[python]" ``` === "pip" ```bash python3 -m venv .venv && source .venv/bin/activate pip install "hopsworks[python]" ``` Supported versions of Python: 3.10, 3.11, 3.12, 3.13, 3.14 ([PyPI ↗](https://pypi.org/project/hopsworks/)) ### Profiles The Hopsworks library has several profiles that bring additional dependencies and enable additional functionalities: | Profile Name | Description | | --- | --- | | No Profile | This is the base installation. Supports interacting with the feature store metadata, model registry and deployments. It also supports reading and writing from the feature store from PySpark environments. | | `python` | This profile enables reading and writing from/to the feature store from a Python environment | | `great-expectations` | Installs [Great Expectations](https://greatexpectations.io/) and enables data validation on feature pipelines. Supports 0.18.12 and 1.17.1; 1.17.1 is recommended | | `polars` | This profile installs the [Polars](https://pola.rs/) library and enables reading and writing Polars DataFrames | You can install all the above profiles with the following command: ```bash uv pip install "hopsworks[python,great-expectations,polars]" ``` ## Skills and instructions for coding agents The Hopsworks Python library ships a set of skills for coding agents: Claude Code, Codex, GitHub Copilot and OpenCode. Inside a Hopsworks terminal they are available to every agent automatically. On your own machine, two commands make them available in the repository you are working in. ### Authenticate and write the agent instructions ```bash uv pip install "hopsworks[python]" cd hops setup --host https:// ``` Without `--host`, `hops setup` asks for the host and proposes `https://eu-west.cloud.hopsworks.ai`, the Hopsworks serverless endpoint; press Enter to accept it or type the address of your cluster. `hops setup` opens a browser page where you choose a project, creates an API key for it, and stores the key in `~/.hops.toml`. It then writes the following files into the current directory: | Path | Purpose | | --- | --- | | `AGENTS.md` | Instructions for the agent: the project you are connected to, where the `hopsworks` library is installed on this machine, and how to use the `hops` CLI and the skills. | | `.claude/skills/hops/SKILL.md` | A reference for the `hops` CLI. | | `.claude/commands/hops.md` | The `/hops` slash command for Claude Code. | | `.claude/agents/hops-fti.md` | A Claude Code sub-agent that reviews a project against the feature, training and inference pipeline pattern. | | `.claude/settings.local.json` | Allows `Bash(hops *)`, so Claude Code can run the CLI without asking before each command. | `AGENTS.md` is read by Claude Code, Codex, GitHub Copilot and OpenCode. The files under `.claude/` are read by Claude Code only. Running `hops setup` again in a directory that already has these files updates the files you have not edited and leaves the ones you have edited unchanged. Pass `--no-scaffold` to authenticate without writing any files. ### Add the Hopsworks skills ```bash hops skills install ``` `hops skills install` copies the Hopsworks skills into `.claude/skills/`, one directory per skill, which is where Claude Code discovers them. For another agent, pass `--agent`, which can be repeated: ```bash hops skills install --agent codex hops skills install --agent copilot hops skills install --agent opencode ``` The skills are written to `.codex/skills/`, `.agents/skills/` and `.opencode/skills/` respectively, and for OpenCode the path is also registered in `opencode.json`. An agent loads only the name and description of each skill when it starts and reads a skill in full when a task calls for it, so adding all of them costs a few kilobytes of context rather than the size of the skills themselves. Running `hops skills install` again after upgrading the `hopsworks` library updates the skills you have not edited, keeps the skills you have edited, and removes skills that the new version no longer ships. Pass `--force` to overwrite edited skills as well. To read the skills without adding them to a repository: ```bash hops skills list hops skills show hops-fg ``` ## Hopsworks Java Library If you want to interact with the Hopsworks Feature Store from environments such as Spark or Beam, you can use the Hopsworks Feature Store (Hopsworks) Java library. !!! note "Feature Store Only" The Java library only allows interaction with the Feature Store component of the Hopsworks platform. Additionally each environment might restrict the supported API operation. You can see which API operation is supported by which environment [here](../fs/compute_engines.md) The Hopsworks library is available on the Hopsworks' Maven repository. If you are using Maven as build tool, you can add the following in your `pom.xml` file: ```xml Hops Hops Repository https://archiva.hops.works/repository/Hops/ true true ``` The library has different builds targeting different environments: ### Hopsworks Java The `artifactId` for the Hopsworks Java build is `hsfs`, if you are using Maven as build tool, you can add the following dependency: ```xml com.logicalclocks hsfs ${hsfs.version} ``` ### Spark The `artifactId` for the Spark build is `hsfs-spark-spark{spark.version}`, if you are using Maven as build tool, you can add the following dependency: ```xml com.logicalclocks hsfs-spark-spark3.1 ${hsfs.version} ``` Hopsworks provides builds for Spark 3.1, 3.3 and 3.5. The builds are also provided as JAR files which can be downloaded from [Hopsworks repository](https://repo.hops.works/master/hsfs) ### Beam The `artifactId` for the Beam build is `hsfs-beam`, if you are using Maven as build tool, you can add the following dependency: ```xml com.logicalclocks hsfs-beam ${hsfs.version} ``` ## Next Steps If you are using a local python environment and want to connect to Hopsworks, you can follow the [Python Guide](../integrations/python.md#generate-an-api-key) section to create an API Key and to get started. If you use a coding agent, see [Skills and instructions for coding agents][skills-and-instructions-for-coding-agents] to give it the Hopsworks skills. ## Other environments The Hopsworks Feature Store client libraries can also be installed in external environments, such as Databricks, AWS Sagemaker, or Azure Machine Learning. For more information, see [Client Integrations](../integrations/index.md). ================================================================================ # Hopsworks CLI Source: https://docs.hopsworks.ai/latest/user_guides/client_installation/cli/ # Hopsworks CLI `hops` is the Hopsworks command line. It ships with the Hopsworks Python library, so anywhere the library is installed the command is available. The same commands run from your laptop, from a CI pipeline, from a coding agent, or inside the [project terminal][terminal], where it is already connected. --8<-- "user_guides/client_installation/cli/one-cli-two-seats.html" ## Install and connect Install the library with the `python` profile and the CLI comes with it. The profile brings the Arrow and Kafka dependencies the data commands (`fg preview`, `fv get`, `sql`) read and write through. === "uv" ```bash uv venv && source .venv/bin/activate uv pip install "hopsworks[python]" ``` === "pip" ```bash python3 -m venv .venv && source .venv/bin/activate pip install "hopsworks[python]" ``` On your own machine, `hops setup` opens a browser, lets you pick a project, creates an API key and caches it in `~/.hops.toml`: ```bash hops setup --host https://my.hopsworks.ai ``` Non-interactive environments such as CI use an existing API key instead: ```bash hops login --host https://my.hopsworks.ai --api-key "$HOPSWORKS_API_KEY" --project fraud_detection ``` Inside the [project terminal][terminal] no login is needed, `hops` is pointed at the project you opened it from. ## Explore and use the feature store Every asset type has a subcommand: `project`, `fg`, `fv`, `td`, `model`, `deployment`, `job`, `datasource`, `sql`. `hops --help` lists the verbs. ```bash hops project use fraud_detection hops fg list hops fg info transactions hops fg preview transactions --n 5 hops fv create transactions_fraud --feature-group transactions hops fv get transactions_fraud --entry "cc_num=4532015112830366" hops sql "select count(*) from transactions_1" ``` Add `--json` to any command for machine-readable output: ```bash hops fg info transactions --json ``` ```json { "id": 1080, "name": "transactions", "version": 1, "type": "cached", "online_enabled": false, "primary_key": ["tid"], "event_time": "datetime", "features": [ {"name": "tid", "type": "bigint", "primary": true} ] } ``` ## For coding agents The CLI is the simplest way to give an agent access to Hopsworks: allow it to run `hops` and it can read and write the project. `hops init` scaffolds the Hopsworks skill, slash command and sub-agent for Claude Code into a repository and allows `Bash(hops *)` there: ```bash hops init --dir . ``` `hops skills list` shows the Hopsworks skills the agent can load, feature groups, feature views, training, online inference, monitoring and more. The [project terminal][terminal] has Claude Code and Codex preinstalled with `hops` already connected, and the [Wizard][wizard] uses exactly this path to build a system end to end. ================================================================================ # Feature Store Guides Source: https://docs.hopsworks.ai/latest/user_guides/fs/ # Feature Store Guides Feature pipelines write to feature groups, training and inference pipelines read through feature views. The guides below follow that order: connect a source, write, validate, read, transform.
- :material-database-plus-outline:{ .lg .middle } **Start here** --- Create a feature group from a DataFrame and insert it. Everything else in this section builds on a feature group that exists. ```python fg = fs.get_or_create_feature_group( name="transactions", version=1, primary_key=["tid"], event_time="datetime", online_enabled=True, ) fg.insert(df) ``` [Create a feature group](feature_group/create.md) · [Create a feature view](feature_view/overview.md) · [Training data](feature_view/training-data.md)
:material-database-import-outline:{ .hops-role-ico } Write { .hops-role-cap } - [Data sources](data_source/index.md) Register warehouses, object stores and databases to read from and write to. - [Feature groups](feature_group/index.md) Create, insert, evolve the schema, set time to live, deprecate. - [External and spine groups](feature_group/create_external.md) Point at data that stays where it is, or supply keys and labels without storing them. - [Ingest with dltHub](feature_group/ingest_with_dlthub.md) Load from hundreds of sources through dlt pipelines.
:material-check-decagram-outline:{ .hops-role-ico } Trust { .hops-role-cap } - [Statistics](feature_group/statistics.md) Descriptive statistics on every insert, configurable per group. - [Data validation](feature_group/data_validation.md) Great Expectations suites run on insert, with a policy on failure. - [Feature monitoring](feature_monitoring/index.md) Scheduled statistics and drift detection against a reference window. - [Notifications and observability](feature_group/notification.md) Change notifications and online ingestion status.
:material-database-export-outline:{ .hops-role-ico } Read { .hops-role-cap } - [Feature views](feature_view/index.md) Select features across groups and read them the same way for training and inference. - [Training data](feature_view/training-data.md) Materialise splits as files or read them straight into memory. - [Batch and online reads](feature_view/batch-data.md) Batch inference data by time range, single vectors from the online store. - [Feature server](feature_view/feature-server.md) Serve feature vectors over REST without the Python client.
:material-function-variant:{ .hops-role-ico } Transform and run { .hops-role-cap } - [Transformation functions](transformation_functions.md) Model-independent functions applied on write, model-dependent on read. - [Compute engines](compute_engines.md) Which operations run on Python, Spark or Flink. - [Client integrations](../integrations/index.md) Databricks, SageMaker, EMR, Azure ML, Flink, Beam and more. - [Vector similarity search](vector_similarity_search.md) Embeddings in a feature group, nearest-neighbour queries.
================================================================================ # Data Source Guides Source: https://docs.hopsworks.ai/latest/user_guides/fs/data_source/ # Data Source Guides You can define data sources in Hopsworks for batch and streaming data sources. Data Sources securely store the authentication information about how to connect to an external data store. They can be used from programs within Hopsworks or externally. !!!warning In the previous versions of Hopsworks, this used to be called a storage connector. There are four main use cases for Data Sources: - Simply use it to read data from the storage into a dataframe. - [External (on-demand) Feature Groups](../../../concepts/fs/feature_group/external_fg.md) can be defined with data sources. This way, Hopsworks stores only the metadata about the features, but does not keep a copy of the data itself. This is also called the Connector API. - Write [training data](../../../concepts/fs/feature_view/offline_api.md) to an external storage system to make it accessible by third parties. - Managed [feature group](../../../user_guides/fs/feature_group/create.md) that stores offline data in an external storage system. Currently [S3](../data_source/creation/s3.md), [GCS](../data_source/creation/gcs.md) and [AWS Glue](../data_source/creation/glue.md) connectors are supported. Data Sources provide two main mechanisms for authentication: using credentials or an authentication role (IAM Role on AWS or Managed Identity on Azure). Hopsworks supports both a single IAM role (AWS) or Managed Identity (Azure) for the whole Hopsworks cluster or multiple IAM roles (AWS) or Managed Identities (Azure) that can only be assumed by users with a specific role in a specific project. By default, each project is created with three default Data Sources: A JDBC connector to the online feature store, a HopsFS connector to the Training Datasets directory of the project and a JDBC connector to the offline feature store.
![Image title](../../../assets/images/guides/fs/data_source/data_source_overview.png)
The Data Source View in the User Interface
## Cloud Agnostic Cloud agnostic storage systems:
- :simple-snowflake:{ .lg .middle style="color:#29B5E8" } **Snowflake** --- Query Snowflake databases and tables using SQL. [:octicons-arrow-right-24: Configure](creation/snowflake.md) - :simple-apachekafka:{ .lg .middle } **Kafka** --- Read from a Kafka cluster into a Spark Structured Streaming Dataframe. [:octicons-arrow-right-24: Configure](creation/kafka.md) - :simple-sap:{ .lg .middle style="color:#0FAAFF" } **SAP HANA** --- Query SAP HANA tenant databases using SQL. [:octicons-arrow-right-24: Configure][data-source-sap-hana] - :material-database:{ .lg .middle style="color:var(--hops-accent-text)" } **JDBC** --- Connect to any JDBC compatible database and query it using SQL. [:octicons-arrow-right-24: Configure](creation/jdbc.md) - :material-api:{ .lg .middle style="color:var(--hops-accent-text)" } **REST API** --- Connect to external HTTP APIs with configurable headers and authentication. [:octicons-arrow-right-24: Configure](creation/rest_api.md) - :material-chart-box-outline:{ .lg .middle style="color:var(--hops-accent-text)" } **CRM, Sales & Analytics** --- Connect to supported CRM, sales, and analytics platforms. [:octicons-arrow-right-24: Configure](creation/crm_sales_analytics.md) - :material-folder-network-outline:{ .lg .middle style="color:var(--hops-accent-text)" } **HopsFS** --- Connect and read from directories of Hopsworks' internal file system. [:octicons-arrow-right-24: Configure](creation/hopsfs.md)
## AWS For AWS the following storage systems are supported:
- :fontawesome-brands-aws:{ .lg .middle style="color:#FF9900" } **S3** --- Read file-based storage in S3 such as parquet or CSV. [:octicons-arrow-right-24: Configure](creation/s3.md) - :fontawesome-brands-aws:{ .lg .middle style="color:#FF9900" } **AWS Glue** --- Integrate with the Glue Data Catalog over S3, for Iceberg, Delta, Hudi and plain files. [:octicons-arrow-right-24: Configure](creation/glue.md) - :fontawesome-brands-aws:{ .lg .middle style="color:#FF9900" } **Redshift** --- Query Redshift databases and tables using SQL. [:octicons-arrow-right-24: Configure](creation/redshift.md) - :fontawesome-brands-aws:{ .lg .middle style="color:#FF9900" } **RDS (SQL)** --- Query the Amazon Relational Database Service using SQL. [:octicons-arrow-right-24: Configure](creation/sql.md)
## Azure For Azure the following storage systems are supported:
- :material-microsoft-azure:{ .lg .middle style="color:#0078D4" } **ADLS** --- Read file-based storage in ADLS such as parquet or CSV. [:octicons-arrow-right-24: Configure](creation/adls.md)
## GCP For GCP the following storage systems are supported:
- :simple-googlebigquery:{ .lg .middle style="color:#4285F4" } **BigQuery** --- Query BigQuery databases and tables using SQL. [:octicons-arrow-right-24: Configure](creation/bigquery.md) - :simple-googlecloudstorage:{ .lg .middle style="color:#4285F4" } **GCS** --- Read file-based storage in Google Cloud Storage such as parquet or CSV. [:octicons-arrow-right-24: Configure](creation/gcs.md)
## Databricks (AWS only) For Databricks **on AWS** the following storage systems are supported:
- :simple-databricks:{ .lg .middle style="color:#FF3621" } **Unity Catalog** --- Browse catalogs, schemas, and Delta tables, and mount them as external feature groups. [:octicons-arrow-right-24: Configure](creation/unity_catalog.md)
Databricks on Azure and Databricks on GCP are not supported yet. See the [Unity Catalog guide](creation/unity_catalog.md) for the specific reasons and the status of follow-up work. ## Next Steps Move on to the [Configuration and Creation Guides](creation/jdbc.md) to learn how to set up a data source. ================================================================================ # JDBC Source: https://docs.hopsworks.ai/latest/user_guides/fs/data_source/creation/jdbc/ # How-To set up a JDBC Data Source ## Introduction JDBC is an API provided by many database systems. Using JDBC connections one can query and update data in a database, usually oriented towards relational databases. Examples of databases you can connect to using JDBC are MySQL, Postgres, Oracle, DB2, MongoDB or Microsoft SQLServer. In this guide, you will configure a Data Source in Hopsworks to save all the authentication information needed in order to set up a JDBC connection to your database of choice. When you're finished, you'll be able to query the database using Spark through Hopsworks APIs. !!! note Currently, it is only possible to create data sources in the Hopsworks UI. You cannot create a data source programmatically. ## Prerequisites Before you begin this guide you'll need to retrieve the following information from your JDBC compatible database: - **JDBC Connection URL:** Consult the documentation of your target database to determine the correct JDBC URL and parameters. As an example, for MySQL the URL could be: ```plaintext jdbc:mysql://10.0.2.15:3306/[databaseName]?useSSL=false&allowPublicKeyRetrieval=true ``` - **Username and Password:** Typically, you will need to add username and password in your JDBC URL or as key/value parameters. So make sure you have retrieved a username and password with the suitable permissions for the database and table you want to query. ## Creation in the UI ### Step 1: Set up new Data Source Head to the `Data Sources` view on Hopsworks (1) and click `New data source` (2). The `Add a data source` catalog lists the available sources, grouped under `Object storage`, `Data warehouse`, `Database`, `Streaming` and `API & SaaS`. Pick the `JDBC` card to open the creation form.
![Data Source Creation](../../../../assets/images/guides/fs/data_source/data_source_overview.png)
The Data Source View in the User Interface
### Step 2: Enter JDBC Settings Enter the details for your JDBC enabled database.
![JDBC Connector Creation](../../../../assets/images/guides/fs/data_source/jdbc_creation.png)
JDBC Connector Creation Form
1. The form opens with `Source` set to `JDBC`. Click `Change source` to pick a different one. 2. Enter the JDBC connection url. This can for example also contain the username and password. 3. Add additional key/value arguments to be passed to the connection, such as username or password. These might differ by database. !!! note Driver class name is a mandatory argument even if using the default MySQL driver. Add it by specifying a property with the name `driver` and class name as value. The driver class name will differ based on the database. For MySQL databases, the class name is `com.mysql.cj.jdbc.Driver`, as shown in the example image. 4. Click on "Save Credentials". !!! note To be able to use the connector, you need to upload the driver JAR file to the [Jupyter configuration](../../../projects/jupyter/spark_notebook.md) or [Job configuration](../../../projects/jobs/pyspark_job.md) in `Additional Jars`. For MySQL connections the default JDBC driver is already included in Hopsworks so this step can be skipped. ## Next Steps Move on to the [usage guide for data sources](../usage.md) to see how you can use your newly created JDBC connector. ================================================================================ # Snowflake Source: https://docs.hopsworks.ai/latest/user_guides/fs/data_source/creation/snowflake/ # How-To set up a Snowflake Data Source ## Introduction Snowflake provides a cloud-based data storage and analytics service, used as a data warehouse in many enterprises. Data warehouses are often the source of raw data for feature engineering pipelines and Snowflake supports scalable feature computation with SQL. However, Snowflake is not viable as an online feature store that serves features to models in production, with its columnar database layout its latency is too high compared to OLTP databases or key-value stores. In this guide, you will configure a Data Source in Hopsworks to save all the authentication information needed in order to set up a connection to your Snowflake database. When you're finished, you'll be able to query the database using Spark through Hopsworks APIs. !!! note Currently, it is only possible to create data sources in the Hopsworks UI. You cannot create a data source programmatically. ## Prerequisites Before you begin this guide you'll need to retrieve the following information from your Snowflake account and database, the following options are **mandatory**: - **Snowflake Connection URL:** Consult the documentation of your target snowflake account to determine the correct connection URL. This is usually some form of your [Snowflake account identifier](https://docs.snowflake.com/en/user-guide/admin-account-identifier.html). For example: ```plaintext .snowflakecomputing.com ``` OR: ```plaintext https://-.snowflakecomputing.com ``` The account and organization details can be viewed in the Snowsight UI under **Admin > Account** or by querying it in SQL, as explained in [Snowflake documentation](https://docs.snowflake.com/en/user-guide/organizations-gs.html#viewing-the-name-of-your-organization-and-its-accounts). Below is an example of how to view the account and organization to get the account identifier from the Snowsight UI.
![Viewing Snowflake account identifier](../../../../assets/images/guides/fs/data_source/snowflake_account_url.png)
Viewing Snowflake account identifier
!!! note "Authentication methods" The Snowflake data source supports username and password, token-based and key-pair based authentication options. General information on snowflake key pair authentication and setup is at [Snowflake key-pair authentication](https://docs.snowflake.com/en/user-guide/key-pair-auth). - **Username and Password:** Login name for the Snowflake user and password. This is often also referred to as `sfUser` and `sfPassword`. - **Warehouse:** The warehouse to use for the session after connecting - **Database:** The database to use for the session after connecting. - **Schema:** The schema to use for the session after connecting. These are a few additional **optional** arguments: - **Role:** The role field can be used to specify which [Snowflake security role](https://docs.snowflake.com/en/user-guide/security-access-control-overview.html#system-defined-roles) to assume for the session after the connection is established. - **Application:** The application field can also be specified to have better observability in Snowflake with regards to which application is running which query. The application field can be a simple string like “Hopsworks” or, for instance, the project name, to track usage and queries from each Hopsworks project. ## Creation in the UI ### Step 1: Set up new Data Source Head to the `Data Sources` view on Hopsworks (1) and click `New data source` (2). The `Add a data source` catalog lists the available sources, grouped under `Object storage`, `Data warehouse`, `Database`, `Streaming` and `API & SaaS`. Pick the `Snowflake` card to open the creation form.
![Data Source Creation](../../../../assets/images/guides/fs/data_source/data_source_overview.png)
The Data Source View in the User Interface
### Step 2: Enter Snowflake Settings Enter the details for your Snowflake connector. Start by giving it a **name** and an optional **description**. 01. The form opens with `Source` set to `Snowflake`. Click `Change source` to pick a different one. 02. Specify the hostname for your account in the following format `.snowflakecomputing.com` or `https://-.snowflakecomputing.com`. 03. Login name for the Snowflake user. 04. **Authentication** Choose between user account Password, Token or Private Key options. In case of private key, upload your snowflake user Private Key file and set Passphrase if applicable. 05. The warehouse to connect to. 06. The database to use for the connection. 07. Add any additional optional arguments. For example, you can specify `Schema`, `Table`, `Role`, and `Application`. 08. Optional additional key/value arguments. 09. Click on "Save Credentials".
![Snowflake Connector Creation](../../../../assets/images/guides/fs/data_source/snowflake_creation.png)
Snowflake Connector Creation Form
## Next Steps Move on to the [usage guide for data sources](../usage.md) to see how you can use your newly created Snowflake connector. ================================================================================ # Kafka Source: https://docs.hopsworks.ai/latest/user_guides/fs/data_source/creation/kafka/ # How-To set up a Kafka Data Source ## Introduction Apache Kafka is a distributed event store and stream-processing platform. It's a very popular framework for handling realtime data streams and is often used as a message broker for events coming from production systems until they are being processed and either loaded into a data warehouse or aggregated into features for Machine Learning. In this guide, you will configure a Data Source in Hopsworks to save all the authentication information needed in order to set up a connection to your Kafka cluster. When you're finished, you'll be able to read from Kafka topics in your cluster using Spark through Hopsworks APIs. !!! note Currently, it is only possible to create data sources in the Hopsworks UI. You cannot create a data source programmatically. ## Prerequisites Before you begin this guide you'll need to retrieve the following information from Kafka cluster, the following options are **mandatory**: - **Kafka Bootstrap servers:** It is the url of one of the Kafka brokers which you give to fetch the initial metadata about your Kafka cluster. The metadata consists of the topics, their partitions, the leader brokers for those partitions etc. Depending upon this metadata your producer or consumer produces or consumes the data. - **Security Protocol:** The security protocol you want to use to authenticate with your Kafka cluster. Make sure the chosen protocol is supported by your cluster. For an overview of the available protocols, please see the [Confluent Kafka Documentation](https://docs.confluent.io/platform/current/kafka/overview-authentication-methods.html). - **Certificates:** Depending on the chosen security protocol, you might need TrustStore and KeyStore files along with the corresponding key password. Contact your Kafka administrator, if you don't know how to retrieve these. If you want to setup a data source to Hopsworks' internal Kafka cluster, you can download the needed certificates from the integration tab in your project settings. ## Creation in the UI ### Step 1: Set up new Data Source Head to the `Data Sources` view on Hopsworks (1) and click `New data source` (2). The `Add a data source` catalog lists the available sources, grouped under `Object storage`, `Data warehouse`, `Database`, `Streaming` and `API & SaaS`. Pick the `Kafka` card to open the creation form.
![Data Source Creation](../../../../assets/images/guides/fs/data_source/data_source_overview.png)
The Data Source View in the User Interface
### Step 2: Enter Kafka Settings Enter the details for your Kafka connector. Start by giving it a **name** and an optional **description**. 01. The form opens with `Source` set to `Kafka`. Click `Change source` to pick a different one. 02. Add all the bootstrap server addresses and ports that you want the consumers/producers to connect to. The client will make use of all servers irrespective of which servers are specified here for bootstrapping. This list only impacts the initial hosts used to discover the full set of servers. 03. Choose the Security protocol. !!! example "TSL/SSL" By default, Apache Kafka communicates in `PLAINTEXT`, which means that all data is sent in the clear. To encrypt communication, you should configure all the Confluent Platform components in your deployment to use TLS/SSL encryption. TLS uses private-key/certificate pairs, which are used during the TLS handshake process. Each broker needs its own private-key/certificate pair, and the client uses the certificate to authenticate the broker. Each logical client needs a private-key/certificate pair if client authentication is enabled, and the broker uses the certificate to authenticate the client. These are provided in the form of *TrustStore* and *KeyStore* `JKS` files together with a key password. For more information, refer to the official [Apacha Kafka Guide for TSL/SSL authentication](https://docs.confluent.io/platform/current/kafka/authentication_ssl.html). !!! example "SASL SSL or SASL plaintext" Apache Kafka brokers support client authentication using SASL. SASL authentication can be enabled concurrently with TLS/SSL encryption (TLS/SSL client authentication will be disabled). This authentication method often requires extra arguments depending on your setup. Make use of the optional additional key/value arguments (5) to provide these. SASL authentication can be enabled concurrently with TLS/SSL encryption (TLS/SSL client authentication will be disabled). For more information, please refer to the official [Apache Kafka Guide for SASL authentication](https://docs.confluent.io/platform/current/kafka/authentication_sasl/index.html). 04. The endpoint identification algorithm used by clients to validate server host name. The default value is `https`. Clients including client connections created by the broker for inter-broker communication verify that the broker host name matches the host name in the broker’s certificate. 05. Optional additional key/value arguments. 06. Click on "Save Credentials".
![Kafka Connector Creation](../../../../assets/images/guides/fs/data_source/kafka_creation.png)
Kafka Connector Creation Form
## Next Steps Move on to the [usage guide for Data Sources](../usage.md) to see how you can use your newly created Kafka connector. ================================================================================ # HopsFS Source: https://docs.hopsworks.ai/latest/user_guides/fs/data_source/creation/hopsfs/ # How-To set up a HopsFS Data Source ## Introduction HopsFS is a HDFS-compatible filesystem on AWS/Azure/on-premises for data analytics. HopsFS stores its data on object storage in the cloud (S3 in AWs and Blob storage on Azure) and on commodity servers on-premises, ensuring low-cost storage, high availability, and disaster recovery. In Hopsworks, you can access HopsFS natively in programs (Spark, TensorFlow, etc) without the need to define a Data Source. By default, every Project has a Data Source for Training Datasets. When you create training datasets from features in the Feature Store the HopsFS connector is the default Data Source. However, if you want to output data to a different dataset, you can define a new Data Source for that dataset. In this guide, you will configure a HopsFS Data Source in Hopsworks which points at a different directory on the file system than the Training Datasets directory. When you're finished, you'll be able to write training data to different locations in your cluster through Hopsworks APIs. !!! note Currently, it is only possible to create data sources in the Hopsworks UI. You cannot create a data source programmatically. ## Prerequisites Before you begin this guide you'll need to identify a **directory on the filesystem** of Hopsworks, to which you want to point the Data Source that you are going to create. ## Creation in the UI ### Step 1: Set up new Data Source Head to the `Data Sources` view on Hopsworks (1) and click `New data source` (2). The `Add a data source` catalog opens below. Pick the `HopsFS` card to open the creation form.
![Data Source Creation](../../../../assets/images/guides/fs/data_source/data_source_overview.png)
The Data Source View in the User Interface
### Step 2: Enter HopsFS Settings Enter the details for your HopsFS connector. Start by giving it a **name** and an optional **description**. 1. The form opens with `Source` set to `HopsFS`. Click `Change source` to pick a different one. 2. Select the top-level dataset to point the connector to. 3. Click on "Save Credentials".
![HopsFS Connector Creation](../../../../assets/images/guides/fs/data_source/hopsfs_creation.png)
HopsFS Connector Creation Form
## Next Steps Move on to the [usage guide for data sources](../usage.md) to see how you can use your newly created HopsFS connector. ================================================================================ # S3 Source: https://docs.hopsworks.ai/latest/user_guides/fs/data_source/creation/s3/ # How-To set up a S3 Data Source { #data-source-s3 } ## Introduction Amazon S3 or Amazon Simple Storage Service is a service offered by AWS that provides object storage. That means you can store arbitrary objects associated with a key. These kinds of storage systems are often used as Data Lakes with large volumes of unstructured data or file based storage. Popular file formats are `CSV` or `PARQUET`. There are so called Data Lake House technologies such as Delta Lake or Apache Hudi, building an additional layer on top of object based storage with files, to provide database semantics like ACID transactions among others. This has the advantage that cheap storage can be turned into a cloud native data warehouse. These kinds of storages are often the source for raw data from which features can be engineered. In this guide, you will configure a Data Source in Hopsworks to save all the authentication information needed in order to set up a connection to your AWS S3 bucket. When you're finished, you'll be able to read files using Spark through Hopsworks APIs. You can also use the connector to write out training data from the Feature Store, in order to make it accessible by third parties. !!! note Currently, it is only possible to create data sources in the Hopsworks UI. You cannot create a data source programmatically. ## Prerequisites Before you begin this guide you'll need to retrieve the following information from your AWS S3 account and bucket: - **Bucket:** You will need a S3 bucket that you have access to. The bucket is identified by its name. - **Path (Optional):** If needed, a path can be defined to ensure that all operations are restricted to a specific location within the bucket. - **Region (Optional):** You will need an S3 region to have complete control over data when managing the feature group that relies on this data source. The region is identified by its code. - **Authentication Method:** You can authenticate using Access Key/Secret, or use IAM roles. If you want to use an IAM role it either needs to be attached to the entire Hopsworks cluster or Hopsworks needs to be able to assume the role. See [IAM role documentation](../../../../setup_installation/admin/roleChaining.md) for more information. - **Server Side Encryption details:** If your bucket has server side encryption (SSE) enabled, make sure you know which algorithm it is using (AES256 or SSE-KMS). If you are using SSE-KMS, you need the resource ARN of the managed key. ## Creation in the UI ### Step 1: Set up new Data Source Head to the `Data Sources` view on Hopsworks (1) and click `New data source` (2). The `Add a data source` catalog lists the available sources, grouped under `Object storage`, `Data warehouse`, `Database`, `Streaming` and `API & SaaS`. Pick the `AWS S3` card to open the creation form.
![Data Source Creation](../../../../assets/images/guides/fs/data_source/data_source_overview.png)
The Data Source View in the User Interface
### Step 2: Enter Bucket Information Enter the details for your S3 connector. The `Source` line at the top of the form shows `AWS S3`, and `Change source` takes you back to the catalog. Start by giving it a **name** and an optional **description**. And set the name of the S3 Bucket you want to point the connector to. Optionally, specify the region if you wish to have a Hopsworks-managed feature group stored using this connector.
![S3 Connector Creation](../../../../assets/images/guides/fs/data_source/s3_creation.png)
S3 Connector Creation Form
### Step 3: Configure Authentication #### Instance Role Choose instance role if you have an EC2 instance profile attached to your Hopsworks cluster nodes with a role which grants you access to the specified bucket. #### Temporary Credentials Choose temporary credentials if you are using [AWS Role chaining](../../../../setup_installation/admin/roleChaining.md) to control the access permission on a project and user role base. Once you have selected *Temporary Credentials* select the role that give access to the specified bucket. For this role to appear in the list it needs to have been configured by an administrator, see the [AWS Role chaining documentation](../../../../setup_installation/admin/roleChaining.md) for more details. !!! warning "Session Duration" By default, the session duration that the role will be assumed for is 1 hour or 3600 seconds. This means if you want to use the data source for example to write [training data to S3](../usage.md#writing-training-data), the training dataset creation cannot take longer than one hour. Your administrator can change the default session duration for AWS data sources, by first [increasing the max session duration of the IAM Role](https://docs.aws.amazon.com/IAM/latest/UserGuide/id_roles_use.html#id_roles_use_view-role-max-session) that you are assuming. And then changing the `fs_data_source_session_duration` [configuration variable](../../../../setup_installation/admin/variables.md) to the appropriate value in seconds. #### Access Key/Secret The most simple authentication method are Access Key/Secret, choose this option to get started quickly, if you are able to retrieve the keys using the IAM user administration. ### Step 4: Configure Server Side Encryption Additionally, you can specify if your Bucket has SSE enabled. #### AES256 For AES256, there is nothing to do but enabling the encryption by toggling the `AES256` option. This is using S3-Managed Keys, also called [SSE-S3](https://docs.aws.amazon.com/AmazonS3/latest/userguide/serv-side-encryption.html). #### SSE-KMS With this option the [encryption key is managed by AWS KMS](https://docs.aws.amazon.com/AmazonS3/latest/userguide/serv-side-encryption.html), with some additional benefits and charges for using this service. The difference is that you need to provide the resource ARN of the key. If you have SSE-KMS enabled for your bucket, you can find the key ARN in the "Properties" section of the bucket details on AWS. ### Step 5: Add Spark Options (optional) Here you can specify any additional spark options that you wish to add to the spark context at runtime. Multiple options can be added as key - value pairs. To connect to a S3 compatible storage other than AWS S3, you can add the option with key as `fs.s3a.endpoint` and the endpoint you want to use as value. The data source will then be able to read from your specified S3 compatible storage. You can also add options to configure the S3A client. For example, to disable SSL certificate verification, you can add the option with key as `fs.s3a.connection.ssl.enabled` and value as `false`. You can also configure other options such as `fs.s3a.path.style.access` if you use s3 compliant storage which does not support virtual hosting. !!! warning "Spark Configuration" When using the data source within a Spark application, the credentials are set at application level. This allows users to access multiple buckets with the same data source within the same application (assuming the credentials allow it). You can disable this behaviour by setting the option `fs.s3a.global-conf` to `False`. If the `global-conf` option is disabled, the credentials are set on a per-bucket basis and users will be able to use the credentials to access data only from the bucket specified in the data source configuration. ### Step 6: Save changes Click on "Save Credentials". ## Next Steps Move on to the [usage guide for data sources](../usage.md) to see how you can use your newly created S3 connector. ================================================================================ # AWS Glue Source: https://docs.hopsworks.ai/latest/user_guides/fs/data_source/creation/glue/ # How-To set up an AWS Glue Data Source { #data-source-glue } ## Introduction The Glue Data Source integrates with the AWS Glue Data Catalog. It points at a Glue database backed by Amazon S3, where the data always lives. For this reason the Glue Data Source provides the same S3 credentials (`access_key`, `secret_key`, `session_token`, `region`) as the [S3 Data Source](s3.md). This works for any data format: Apache Iceberg, Delta Lake and Apache Hudi, as well as plain file formats such as `csv` and `parquet`. How the Glue Data Catalog itself is used depends on the format: - Iceberg: the catalog owns the table's current-metadata pointer, so reads and writes are mediated by the catalog (the table is addressed by `.`). - Delta and Hudi: the on-path transaction log or timeline stays authoritative; the catalog is a discoverability mirror that is registered on create and synced on write so external engines (Athena, EMR, ...) can find the table by name. - Plain file formats (`csv`, `parquet`, ...): the Data Source is used only for S3 access; nothing is registered in the catalog. In this guide, you will configure a Data Source in Hopsworks to save all the authentication information needed in order to set up a connection to your AWS Glue database. When you're finished, you'll be able to read tables using Spark through Hopsworks APIs, and to create managed feature groups whose offline data is stored in the Glue-registered location on S3. !!! note Currently, it is only possible to create data sources in the Hopsworks UI. You cannot create a data source programmatically. ## Prerequisites Before you begin this guide you'll need to retrieve the following information from your AWS Glue and S3 setup: - **Database:** You will need the name of the Glue database that contains, or will contain, your tables. - **Region:** You will need the AWS region in which the Glue Data Catalog and the backing S3 bucket reside. The region is identified by its code. - **Authentication Method:** You can authenticate using Access Key/Secret, or use IAM roles. If you want to use an IAM role it either needs to be attached to the entire Hopsworks cluster or Hopsworks needs to be able to assume the role. See [IAM role documentation](../../../../setup_installation/admin/roleChaining.md) for more information. The credentials must grant access both to the Glue Data Catalog and to the backing S3 bucket. ## Creation in the UI ### Step 1: Set up new Data Source Head to the `Data Sources` view on Hopsworks (1) and click `New data source` (2). The `Add a data source` catalog opens below. Pick the `AWS Glue` card to open the creation form. The card is only offered where the connector is supported, so it is absent or disabled on non-cloud clusters.
![Data Source Creation](../../../../assets/images/guides/fs/data_source/data_source_overview.png)
The Data Source View in the User Interface
### Step 2: Enter Glue Settings Enter the details for your Glue connector. 01. The form opens with `Source` set to `AWS Glue`. Click `Change source` to pick a different one. 02. Give the data source a **name** and an optional **description**. 03. Set the name of the Glue **database** you want to point the connector to. 04. Optionally set the **Catalog ID**, the AWS account ID that owns the Glue Data Catalog. Leave it empty to use the catalog of the account the credentials belong to. 05. Optionally set the AWS **region** of the Glue Data Catalog and its backing S3 bucket. 06. Choose the **Authentication method**, see the options below. 07. Optionally add **Spark options** as key-value pairs to pass to the Spark context at runtime. 08. Click on "Save Credentials".
![Glue Connector Creation](../../../../assets/images/guides/fs/data_source/glue_creation.png)
Glue Connector Creation Form
The credentials must grant access both to the Glue Data Catalog and to the backing S3 bucket. The available authentication methods are the same as for the [S3 Data Source](s3.md): #### Instance Role Choose instance role if you have an EC2 instance profile attached to your Hopsworks cluster nodes with a role which grants access to the Glue Data Catalog and the backing S3 bucket. #### Temporary Credentials Choose temporary credentials if you are using [AWS Role chaining](../../../../setup_installation/admin/roleChaining.md) to control the access permission on a project and user role base. Once you have selected *Temporary Credentials* select the role that gives access to the Glue Data Catalog and the backing S3 bucket. For this role to appear in the list it needs to have been configured by an administrator, see the [AWS Role chaining documentation](../../../../setup_installation/admin/roleChaining.md) for more details. !!! warning "Session Duration" By default, the session duration that the role will be assumed for is 1 hour or 3600 seconds. This means if you want to use the data source for example to write [training data to S3](../usage.md#writing-training-data), the training dataset creation cannot take longer than one hour. Your administrator can change the default session duration for AWS data sources, by first [increasing the max session duration of the IAM Role](https://docs.aws.amazon.com/IAM/latest/UserGuide/id_roles_use.html#id_roles_use_view-role-max-session) that you are assuming. And then changing the `fs_data_source_session_duration` [configuration variable](../../../../setup_installation/admin/variables.md) to the appropriate value in seconds. #### Access Key/Secret The most simple authentication method are Access Key/Secret, choose this option to get started quickly, if you are able to retrieve the keys using the IAM user administration. ## Feature group path When creating a feature group from this Data Source and the Glue database has a location, the feature group path is generated automatically by appending the new table to that database location, so no path needs to be set. The database location is the **Location** set on the Glue database in the AWS console.
![Glue Database Location](../../../../assets/images/guides/fs/data_source/glue_database_location.png)
The Location of a Glue database in the AWS console
Otherwise, the path must be set explicitly on the data source, for example: === "PySpark" ```python ds = fs.get_data_source("glue") ds.path = "s3://mybucket/iceberg-warehouse/myglue.db/fg_1/" ``` An explicitly set path always takes precedence over the generated one. ## Direct Spark or PyIceberg access For direct Spark or PyIceberg access outside the feature group APIs, the Data Source supplies the matching catalog properties. See [`GlueConnector.catalog_options`][hsfs.storage_connector.GlueConnector.catalog_options] (Spark) and [`GlueConnector.pyiceberg_catalog_options`][hsfs.storage_connector.GlueConnector.pyiceberg_catalog_options] (PyIceberg). ## Next Steps Move on to the [usage guide for data sources](../usage.md) to see how you can use your newly created Glue connector. ================================================================================ # Redshift Source: https://docs.hopsworks.ai/latest/user_guides/fs/data_source/creation/redshift/ # How-To set up a Redshift Data Source ## Introduction Amazon Redshift is a popular managed data warehouse on AWS, used as a data warehouse in many enterprises. Data warehouses are often the source of raw data for feature engineering pipelines and Redshift supports scalable feature computation with SQL. However, Redshift is not viable as an online feature store that serves features to models in production, with its columnar database layout its latency is too high compared to OLTP databases or key-value stores. In this guide, you will configure a Data Source in Hopsworks to save all the authentication information needed in order to set up a connection to your AWS Redshift cluster. When you're finished, you'll be able to query the database using Spark through Hopsworks APIs. !!! note Currently, it is only possible to create data sources in the Hopsworks UI. You cannot create a data source programmatically. ## Prerequisites Before you begin this guide you'll need to retrieve the following information from your AWS account and Redshift database, the following options are **mandatory**: - **Cluster identifier:** The name of the cluster. - **Database endpoint:** The endpoint for the database. Should be in the format of `[UUID].eu-west-1.redshift.amazonaws.com`. - **Database name:** The name of the database to query. - **Database port:** The port of the cluster. Defaults to 5349. - **Authentication method:** There are three options available for authenticating with the Redshift cluster. The first option is to configure a username and a password. The second option is to configure an IAM role. With IAM roles, Jobs or notebooks launched on Hopsworks do not need to explicitly authenticate with Redshift, as the Hopsworks library will transparently use the IAM role to acquire a temporary credential to authenticate the specified user. Read more about IAM roles in our [AWS credentials pass-through guide](../../../../setup_installation/admin/roleChaining.md). Lastly, option `Instance Role` will use the default ARN Role configured for the cluster instance. ## Creation in the UI ### Step 1: Set up new Data Source Head to the `Data Sources` view on Hopsworks (1) and click `New data source` (2). The `Add a data source` catalog lists the available sources, grouped under `Object storage`, `Data warehouse`, `Database`, `Streaming` and `API & SaaS`. Pick the `Redshift` card to open the creation form.
![Data Source Creation](../../../../assets/images/guides/fs/data_source/data_source_overview.png)
The Data Source View in the User Interface
### Step 2: Enter The Connector Information Enter the details for your Redshift connector. Start by giving it a **name** and an optional **description**. 01. The form opens with `Source` set to `Redshift`. Click `Change source` to pick a different one. 02. The name of the cluster. 03. The database endpoint. Should be in the format `[UUID].eu-west-1.redshift.amazonaws.com`. For example, if the endpoint info displayed in Redshift is `cluster-id.uuid.eu-north-1.redshift.amazonaws.com:5439/dev` the value to enter here is just `uuid.eu-north-1.redshift.amazonaws.com` 04. The database name. 05. The database port. 06. The database username, here you have the possibility to let Hopsworks auto-create the username for you. 07. Database Driver (optional): You can use the default JDBC Redshift Driver `com.amazon.redshift.jdbc42.Driver` included in Hopsworks or set a different driver (More on this later). 08. Optionally provide the database group and table for the connector. A database group is the group created for the user if applicable. More information, at [redshift documentation](https://docs.aws.amazon.com/redshift/latest/dg/r_Groups.html) 09. Set the appropriate authentication method. 10. Click on "Save Credentials".
![Redshift Connector Creation](../../../../assets/images/guides/fs/data_source/redshift_creation.png)
Redshift Connector Creation Form
!!! warning "Session Duration" By default, the session duration that the role will be assumed for is 1 hour or 3600 seconds. This means if you want to use the data source for example to [read or create an external Feature Group from Redshift](../usage.md#creating-an-external-feature-group), the operation cannot take longer than one hour. Your administrator can change the default session duration for AWS data sources, by first [increasing the max session duration of the IAM Role](https://docs.aws.amazon.com/IAM/latest/UserGuide/id_roles_use.html#id_roles_use_view-role-max-session) that you are assuming. And then changing the `fs_data_source_session_duration` [configuration property](../../../../setup_installation/admin/variables.md) to the appropriate value in seconds. ### Step 3: Upload the Redshift database driver (optional) The `redshift-jdbc42` JDBC driver is included by default in the Hopsworks distribution. If you wish to use a different driver, you need to upload it on Hopsworks and add it as a dependency of Jobs and Jupyter Notebooks that need it. First, you need to [download the library](https://docs.aws.amazon.com/redshift/latest/mgmt/jdbc20-download-driver.html). Select the driver version without the AWS SDK. #### Add the driver to Jupyter Notebooks and Spark jobs You can now add the driver file to the default job and Jupyter configuration. This way, all jobs and Jupyter instances in the project will have the driver attached so that Spark can access it. 1. Go into the Project's settings. 2. Select "Compute configuration". 3. Select "Spark". 4. Under "Additional Jars" choose "Upload new file" to upload the driver jar file.
![Redshift Driver Job and Jupyter Configuration](../../../../assets/images/guides/fs/data_source/jupyter_config.png)
Attaching the Redshift Driver to all Jobs and Jupyter Instances of the Project
Alternatively, you can choose the "From Project" option. You will first have to upload the jar file to the Project using the File Browser. After you have uploaded the jar file, you can select it using the "From Project" option. To upload the jar file to the Project through the File Browser, see the example below: 1. Open File Browser 2. Navigate to "Resources" directory 3. Upload the jar file
![Redshift Driver Upload](../../../../assets/images/guides/fs/data_source/driver_upload.png)
Redshift Driver Upload in the File Browser
!!! tip If you face network connectivity issues to your Redshift cluster, a common cause could be the cluster database port not being accessible from outside the Redshift cluster VPC network. A quick and dirty way to enable connectivity is to [Enable Publicly Accessible](https://aws.amazon.com/premiumsupport/knowledge-center/redshift-cluster-private-public/). However, in a production setting, you should use [VPC peering](https://docs.aws.amazon.com/vpc/latest/peering/what-is-vpc-peering.html) or some equivalent mechanism for connecting the clusters. ## Next Steps Move on to the [usage guide for data sources](../usage.md) to see how you can use your newly created Redshift connector. ================================================================================ # ADLS Source: https://docs.hopsworks.ai/latest/user_guides/fs/data_source/creation/adls/ # How-To set up a ADLS Data Source ## Introduction Azure Data Lake Storage (ADLS) Gen2 is a HDFS-compatible filesystem on Azure for data analytics. The ADLS Gen2 filesystem stores its data in Azure Blob storage, ensuring low-cost storage, high availability, and disaster recovery. In Hopsworks, you can access ADLS Gen2 by defining a Data Source and creating and granting permissions to a service principal. In this guide, you will configure a Data Source in Hopsworks to save all the authentication information needed in order to set up a connection to your Azure ADLS filesystem. When you're finished, you'll be able to read files using Spark through Hopsworks APIs. You can also use the connector to write out training data from the Feature Store, in order to make it accessible by third parties. !!! note Currently, it is only possible to create data sources in the Hopsworks UI. You cannot create a data source programmatically. ## Prerequisites Before you begin this guide you'll need to retrieve the following information from your Azure ADLS account: - **Data Lake Storage Gen2 Account:** Create an [Azure Data Lake Storage Gen2 account](https://docs.microsoft.com/azure/storage/data-lake-storage/quickstart-create-account) and [initialize a filesystem, enabling the hierarchical namespace](https://docs.microsoft.com/azure/storage/data-lake-storage/namespace). Note that your storage account must belong to an Azure resource group. - **Azure AD application and service principal:** [Create an Azure AD application and service principal](https://docs.microsoft.com/en-us/azure/active-directory/develop/howto-create-service-principal-portal) that can access your ADLS storage account and its resource group. - **Service Principal Registration:** Register the service principal, granting it a role assignment such as Storage Blob Data Contributor, on the Azure Data Lake Storage Gen2 account. !!! info When you specify the 'container name' in the ADLS data source, you need to have previously created that container - the Hopsworks Feature Store will not create that storage container for you. ## Creation in the UI ### Step 1: Set up new Data Source Head to the `Data Sources` view on Hopsworks (1) and click `New data source` (2). The `Add a data source` catalog opens below. Pick the `Azure Data Lake` card to open the creation form. The card is only offered where the connector is supported, so it is absent or disabled on non-cloud clusters.
![Data Source Creation](../../../../assets/images/guides/fs/data_source/data_source_overview.png)
The Data Source View in the User Interface
### Step 2: Enter ADLS Information Enter the details for your ADLS connector. Start by giving it a **name** and an optional **description**.
![ADLS Connector Creation](../../../../assets/images/guides/fs/data_source/adls_creation.png)
ADLS Connector Creation Form
1. The form opens with `Source` set to `Azure Data Lake`. Click `Change source` to pick a different one. 2. Set directory ID. 3. Enter the Application ID. 4. Paste the Service Credentials. 5. Specify account name. 6. Provide the container name. 7. Click on "Save Credentials". ### Step 3: Azure Create an ADLS Resource When programmatically signing in, you need to pass the tenant ID with your authentication request and the application ID. You also need a certificate or an authentication key (described in the following section). To get those values, use the following steps: 1. Select Azure Active Directory. 2. From App registrations in Azure AD, select your application. 3. Copy the Directory (tenant) ID and store it in your application code.
![ADLS select tenant-id](../../../../assets/images/guides/fs/data_source/adls-copy-tenant-id.png)
You need to copy the Directory (tenant) id and paste it to the Hopsworks ADLS Data Source "Directory id" text field.
4. Copy the Application ID and store it in your application code.
![ADLS select app-id](../../../../assets/images/guides/fs/data_source/adls-copy-app-id.png)
>You need to copy the Application id and paste it to the Hopsworks ADLS Data Source "Application id" text field.
5. Create an Application Secret and copy it into the Service Credential field.
![ADLS enter application secret](../../../../assets/images/guides/fs/data_source/adls-copy-secret.png)
You need to copy the Application Secret and paste it to the Hopsworks ADLS Data Source "Service Credential" text field.
#### Common Problems If you get a permission denied error when writing or reading to/from a ADLS container, it is often because the storage principal (app) does not have the correct permissions. Have you added the "Storage Blob Data Owner" or "Storage Blob Data Contributor" role to the resource group for your storage account (or the subscription for your storage group, if you apply roles at the subscription level)? Go to your resource group, then in "Access Control (IAM)", click the "Add" button to add a "role assignment". If you get an error "StatusCode=404 StatusDescription=The specified filesystem does not exist.", then maybe you have not created the storage account or the storage container. #### References - [How to create a service principal on Azure](https://docs.microsoft.com/en-us/azure/active-directory/develop/howto-create-service-principal-portal) ## Next Steps Move on to the [usage guide for data sources](../usage.md) to see how you can use your newly created ADLS connector. ================================================================================ # BigQuery Source: https://docs.hopsworks.ai/latest/user_guides/fs/data_source/creation/bigquery/ # How-To set up a BigQuery Data Source ## Introduction A BigQuery data source provides integration to Google Cloud BigQuery. BigQuery is Google Cloud's managed data warehouse supporting that lets you run analytics and execute SQL queries over large scale data. Such data warehouses are often the source of raw data for feature engineering pipelines. In this guide, you will configure a Data Source in Hopsworks to connect to your BigQuery project by saving the necessary information. When you're finished, you'll be able to execute queries and read results of BigQuery using Spark through Hopsworks APIs. The data source uses the Google `spark-bigquery-connector` behind the scenes. To read more about the spark connector, like the spark options or usage, check [Apache Spark SQL connector for Google BigQuery.](https://github.com/GoogleCloudDataproc/spark-bigquery-connector#usage 'github.com/GoogleCloudDataproc/spark-bigquery-connector') !!! note Currently, it is only possible to create data sources in the Hopsworks UI. You cannot create a data source programmatically. ## Prerequisites Before you begin this guide you'll need to retrieve the following information about your GCP account: - **BigQuery Project:** You need a BigQuery project, dataset and table created and have read access to it. Or, if you wish to query a public dataset you need its corresponding details. - **Authentication Method:** Authentication to GCP account is handled by uploading the `JSON keyfile for service account` to the Hopsworks Project. You will need to create this JSON keyfile from GCP. For more information on service accounts and creating keyfile in GCP, read [Google Cloud documentation.](https://cloud.google.com/docs/authentication/production#create_service_account 'creating service account keyfile') !!! note To read data, the BigQuery service account user needs permission to `create read session` which is available in **BigQuery Admin role**. ## Creation in the UI ### Step 1: Set up new Data Source Head to the `Data Sources` view on Hopsworks (1) and click `New data source` (2). The `Add a data source` catalog lists the available sources, grouped under `Object storage`, `Data warehouse`, `Database`, `Streaming` and `API & SaaS`. Pick the `Google BigQuery` card to open the creation form.
![Data Source Creation](../../../../assets/images/guides/fs/data_source/data_source_overview.png)
The Data Source View in the User Interface
### Step 2: Enter source details Enter the details for your BigQuery storage. Start by giving it a unique **name** and an optional **description**.
![BigQuery Creation](../../../../assets/images/guides/fs/data_source/bigquery_creation.png)
BigQuery Creation Form
1. The form opens with `Source` set to `Google BigQuery`. Click `Change source` to pick a different one. 2. Next, set the name of the parent BigQuery project. This is used for billing by GCP. 3. Authentication: Here you should upload your `JSON keyfile for service account` used for authentication. You can choose to either upload from your local using `Upload new file` or choose an existing file within project using `From Project`. 4. Read Options: In the UI set the below fields, 1. *BigQuery Project*: The BigQuery project to read 2. *BigQuery Dataset*: The dataset of the table (Optional) 3. *BigQuery Table*: The table to read (Optional) !!! note *Materialization Dataset*: Temporary dataset used by BigQuery for writing. It must be set to a dataset where the GCP user has table creation permission. The queried table must be in the same location as the `materializationDataset` (e.g 'EU' or 'US'). Also, if a table in the `SQL statement` is from project other than the `parentProject` then use the fully qualified table name i.e. `[project].[dataset].[table]`. For details, read the Google documentation on [usage of query for BigQuery Spark connector](https://github.com/GoogleCloudDataproc/spark-bigquery-connector#reading-data-from-a-bigquery-query). 5. Spark Options: Optionally, you can set additional spark options using the `Key - Value` pairs. 6. Click on "Save Credentials". ## Next Steps Move on to the [usage guide for data sources](../usage.md) to see how you can use your newly created BigQuery connector. ================================================================================ # GCS Source: https://docs.hopsworks.ai/latest/user_guides/fs/data_source/creation/gcs/ # How-To set up a GCS Data Source { #data-source-gcs } ## Introduction This particular type of Data Source provides integration to Google Cloud Storage (GCS). GCS is an object storage service offered by Google Cloud. An object could be simply any piece of immutable data consisting of a file of any format, for example a `CSV` or `PARQUET`. These objects are stored in containers called as `buckets`. These types of storages are often the source for raw data from which features can be engineered. In this guide, you will configure a Data Source in Hopsworks to connect to your GCS bucket by saving the necessary information. When you're finished, you'll be able to read files from the GCS bucket using Spark through Hopsworks APIs. The Data Source uses the Google `gcs-connector-hadoop` behind the scenes. For more information, check out [Google Cloud Data Source for Spark and Hadoop](https://github.com/GoogleCloudDataproc/hadoop-connectors/tree/master/gcs#google-cloud-storage-connector-for-spark-and-hadoop 'google-cloud-storage-connector-for-spark-and-hadoop'). !!! note Currently, it is only possible to create data sources in the Hopsworks UI. You cannot create a data source programmatically. ## Prerequisites Before you begin this guide you'll need to retrieve the following information about your GCP account and bucket: - **Bucket:** You need a GCS bucket created and have read access to it. The bucket is identified by its name. - **Authentication Method:** Authentication to GCP account is handled by uploading the `JSON keyfile for service account` to the Hopsworks Project. You will need to create this JSON keyfile from GCP. For more information on service accounts and creating keyfile in GCP, read [Google Cloud documentation.](https://cloud.google.com/docs/authentication/production#create_service_account 'creating service account keyfile') - **Server-side Encryption** GCS encrypts the data on server side by default. The connector additionally supports the optional encryption method `Customer Supplied Encryption Key` by GCP. You can choose the encryption option `AES-256` and provide AES-256 key and hash, encoded in standard Base64. The encryption details are stored as [Secrets](../../../projects/secrets/create_secret.md) in the Hopsworks for keeping it secure. Read more about encryption on [Google Documentation.](https://cloud.google.com/storage/docs/encryption/customer-supplied-keys) ## Creation in the UI ### Step 1: Set up new Data Source Head to the `Data Sources` view on Hopsworks (1) and click `New data source` (2). The `Add a data source` catalog lists the available sources, grouped under `Object storage`, `Data warehouse`, `Database`, `Streaming` and `API & SaaS`. Pick the `Google Cloud Storage` card to open the creation form.
![Data Source Creation](../../../../assets/images/guides/fs/data_source/data_source_overview.png)
The Data Source View in the User Interface
### Step 2: Enter connector details Enter the details for your GCS connector. Start by giving it a unique **name** and an optional **description**.
![GCS Connector Creation](../../../../assets/images/guides/fs/data_source/gcs_creation.png)
GCS Connector Creation Form
1. The form opens with `Source` set to `Google Cloud Storage`. Click `Change source` to pick a different one. 2. Next, set the name of the GCS Bucket you wish to connect with. 3. Authentication: Here you should upload your `JSON keyfile for service account` used for authentication. You can choose to either upload from your local using `Upload new file` or choose an existing file within project using `From Project`. 4. GCS Server Side Encryption: You can leave this to `Default Encryption` if you do not wish to provide explicit encrypting keys. Otherwise, optionally you can set the encryption setting for `AES-256` and provide the encryption key and hash when selected. 5. Click on `Save Credentials`. ## Next Steps Move on to the [usage guide for data sources](../usage.md) to see how you can use your newly created GCS connector. ================================================================================ # SQL Source: https://docs.hopsworks.ai/latest/user_guides/fs/data_source/creation/sql/ # How-To set up an SQL Data Source ## Introduction The SQL Data Source connects Hopsworks to a Relational Database Service. Supported database types are **MySQL**, **PostgreSQL**, and **Oracle**. Using this connector, you can query and update data in your relational database from Hopsworks. In this guide, you will configure a Data Source in Hopsworks to securely store the authentication information needed to set up a connection to your database instance. When you're finished, you'll be able to query your SQL database using Hopsworks APIs. !!! note Currently, it is only possible to create data sources in the Hopsworks UI. You cannot create a data source programmatically. ## Prerequisites Before you begin, ensure you have the following information from your database instance: - **Host:** The endpoint for your database instance. Example from AWS: 1. Go to the AWS Console → `Aurora and RDS` 2. Click on your DB instance. 3. Under `Connectivity & security`, you'll find the endpoint, e.g.: `mydb.abcdefg1234.us-west-2.rds.amazonaws.com` - **Database:** The name of the database to connect to. For Oracle, this is the **service name** (e.g. `ORCL` or a TNS alias). - **Port:** The port to connect to (e.g. `3306` for MySQL, `5432` for PostgreSQL, `1521` for Oracle). - **Username and Password:** A username and password with the necessary permissions to access the required tables. ### Optional: Oracle Wallet for mTLS Authentication If your Oracle database requires mutual TLS (mTLS) authentication, which is common with Oracle Autonomous Database and Oracle Cloud, you will also need: - **Wallet file:** A `.zip` file containing the wallet credentials (e.g. `cwallet.sso`, `tnsnames.ora`, `sqlnet.ora`). - **Wallet password:** The password for the wallet, if using a PKCS12 wallet (`ewallet.p12`). Auto-login wallets (`cwallet.sso`) do not require a password. !!! tip You can download the wallet zip from the Oracle Cloud Console under your Autonomous Database's **DB Connection** page. Upload the zip file to your Hopsworks project (e.g. to `Resources/`) before creating the data source. !!! warning "Leave the host empty when using a wallet" A host and a wallet are alternatives, not a pair. The wallet's `tnsnames.ora` supplies the host and port, and the database field is the alias to look up there. Supplying a host as well makes the driver connect directly, past the wallet, which a wallet-protected database refuses with a connection error that names neither the cause nor the fix. Hopsworks therefore refuses to save a data source with both a host and a wallet. ## Creation in the UI ### Step 1: Set up a new Data Source Head to the `Data Sources` view on Hopsworks (1) and click `New data source` (2). The `Add a data source` catalog lists the available sources, grouped under `Object storage`, `Data warehouse`, `Database`, `Streaming` and `API & SaaS`. Pick the `SQL` card to open the creation form.
![Data Source Creation](../../../../assets/images/guides/fs/data_source/data_source_overview.png)
The Data Source View in the User Interface
### Step 2: Enter SQL Settings Enter the details for your database. Start by giving the connector a **name** and an optional **description**. 1. The form opens with `Source` set to `SQL`. Click `Change source` to pick a different one. 2. Select the database type (MySQL, PostgreSQL, or Oracle). 3. Enter the host endpoint. Leave it empty when using an Oracle wallet: the wallet supplies the connection details, and the database field names the TNS alias to use. 4. Enter the database name (service name for Oracle). 5. Specify the port. 6. Provide the username and password. 7. For Oracle with mTLS, upload the wallet zip file and provide the wallet password (if required). 8. Click on "Save Credentials".
![SQL Connector Creation](../../../../assets/images/guides/fs/data_source/sql_creation.png)
SQL Connector Creation Form
## Oracle-Specific Notes The generic read, external feature group, and training data workflows are covered in the [usage guide for data sources][data-source-usage]. The following notes apply only to Oracle. ### JDBC driver on the Spark classpath The Oracle JDBC driver JAR (e.g. `ojdbc11.jar`) must be available on the Spark classpath. Upload it via the [Jupyter configuration][how-to-run-a-pyspark-notebook] or [Job configuration][how-to-run-a-pyspark-job] in `Additional Jars`. The MySQL and PostgreSQL drivers are included in Hopsworks by default. ### Spark JDBC limitations !!! warning "Oracle Spark JDBC limitations" - **Single-partition reads only.** All data is fetched through a single JDBC connection from the Spark driver. Spark's parallel JDBC read (via `numPartitions` / `partitionColumn`) is not supported. For very large tables, filter with a `WHERE` clause in your query. - **Wallet available on the driver only.** When using wallet-based authentication, the wallet zip is downloaded from HopsFS and extracted on the Spark driver node. This is sufficient because reads are single-partition (driver-only). - **Timestamp precision.** Spark JDBC supports timestamp precision up to seconds only. Sub-second precision from Oracle `TIMESTAMP` columns may be truncated. ### Python engine The Python engine reads Oracle via the Hopsworks Arrow Flight service, which handles the database connection server-side. No JDBC driver or wallet files are needed on the client, and the Spark JDBC limitations above do not apply. ## Next Steps Move on to the [usage guide for data sources][data-source-usage] to see how you can use your newly created SQL connector. You can also make the database queryable from the query engine by adding a [Trino catalog][trino-catalogs] derived from this data source. ================================================================================ # CRM, Sales & Analytics Source: https://docs.hopsworks.ai/latest/user_guides/fs/data_source/creation/crm_sales_analytics/ # How-To set up a CRM, Sales & Analytics Data Source ## Introduction The `CRM, Sales & Analytics` data source lets you connect Hopsworks to supported business applications and marketing platforms. The following sources are available: - Facebook Ads - Freshdesk - Google Ads - Google Analytics - HubSpot - Pipedrive - Salesforce - Shopify In this guide, you will configure a Data Source in Hopsworks by saving the credentials required by the selected source. !!! note Currently, it is only possible to create data sources in the Hopsworks UI. You cannot create a data source programmatically. ## Prerequisites Before you begin, make sure you have: - A unique name for the data source in Hopsworks. - Read credentials for the external system you want to connect. - Any source-specific identifiers required by that system, such as account, customer, property, or domain identifiers. - For Google Ads and Google Analytics, a service account JSON keyfile that can be uploaded to the Hopsworks project. ## Creation in the UI ### Step 1: Set up new Data Source Head to the `Data Sources` view on Hopsworks (1) and click `New data source` (2). The `Add a data source` catalog lists the available sources, grouped under `Object storage`, `Data warehouse`, `Database`, `Streaming` and `API & SaaS`. Pick the `CRM, Sales & Analytics` card to open the creation form.
![Data Source Creation](../../../../assets/images/guides/fs/data_source/data_source_overview.png)
The Data Source View in the User Interface
### Step 2: Name the data source and select the platform The form opens with `Source` set to `CRM, Sales & Analytics`, and `Change source` takes you back to the catalog. Enter a unique **Name**, an optional **Description**, and pick the platform you want to configure in the **Source** radio group of the connection section.
![CRM, Sales & Analytics - Facebook Ads](../../../../assets/images/guides/fs/data_source/crm_sales_analytics_facebook_ads.png)
CRM, Sales & Analytics data source selection
### Step 3: Enter source-specific credentials The required fields depend on the selected source. #### Facebook Ads Required fields: - **Access Token** - **Account Id**
![Facebook Ads Data Source](../../../../assets/images/guides/fs/data_source/crm_sales_analytics_facebook_ads.png)
Facebook Ads data source form
#### Freshdesk Required fields: - **API Key** - **Domain**
![Freshdesk Data Source](../../../../assets/images/guides/fs/data_source/crm_sales_analytics_freshdesk.png)
Freshdesk data source form
#### Google Ads Required fields: - **Authentication JSON Keyfile** - **Developer Token** - **Customer Id** - **Impersonated Email** The JSON keyfile can be selected either from an existing project file or uploaded as a new file.
![Google Ads Data Source](../../../../assets/images/guides/fs/data_source/crm_sales_analytics_google_ads.png)
Google Ads data source form
#### Google Analytics Required fields: - **Authentication JSON Keyfile** - **Property Id** The JSON keyfile can be selected either from an existing project file or uploaded as a new file.
![Google Analytics Data Source](../../../../assets/images/guides/fs/data_source/crm_sales_analytics_google_analytics.png)
Google Analytics data source form
#### HubSpot Required fields: - **API Key**
![HubSpot Data Source](../../../../assets/images/guides/fs/data_source/crm_sales_analytics_hubspot.png)
HubSpot data source form
#### Pipedrive Required fields: - **API Key**
![Pipedrive Data Source](../../../../assets/images/guides/fs/data_source/crm_sales_analytics_pipedrive.png)
Pipedrive data source form
#### Salesforce Required fields: - **Security Token** - **Username** - **Password**
![Salesforce Data Source](../../../../assets/images/guides/fs/data_source/crm_sales_analytics_salesforce.png)
Salesforce data source form
#### Shopify Required fields: - **Shop URL** - **Private App Password**
![Shopify Data Source](../../../../assets/images/guides/fs/data_source/crm_sales_analytics_shopify.png)
Shopify data source form
### Step 4: Save the credentials After entering the required fields for the selected source: 1. Click **Save Credentials**. 2. Click **Next: Select resource** to continue configuring the data source for downstream use. ## Next Steps Move on to the [usage guide for data sources](../usage.md) to see how you can use your newly created data source. ================================================================================ # REST API Source: https://docs.hopsworks.ai/latest/user_guides/fs/data_source/creation/rest_api/ # How-To set up a REST API Data Source ## Introduction The `REST API` data source lets you connect Hopsworks to external HTTP APIs. You can use it to store the base connection details, optional headers, and the authentication method required by the target API. In this guide, you will configure a REST API Data Source in the Hopsworks UI. !!! note Currently, it is only possible to create data sources in the Hopsworks UI. You cannot create a data source programmatically. ## Prerequisites Before you begin, make sure you have: - A unique name for the data source in Hopsworks. - The **Base URL** of the target API. - Any headers you want to send with requests. - The authentication details required by the target API. ## Creation in the UI ### Step 1: Set up new Data Source Head to the `Data Sources` view on Hopsworks (1) and click `New data source` (2). The `Add a data source` catalog lists the available sources, grouped under `Object storage`, `Data warehouse`, `Database`, `Streaming` and `API & SaaS`. Pick the `REST API` card to open the creation form.
![Data Source Creation](../../../../assets/images/guides/fs/data_source/data_source_overview.png)
The Data Source View in the User Interface
### Step 2: Enter REST API settings The form opens with `Source` set to `REST API`, and `Change source` takes you back to the catalog. Provide the common connection settings shown in the form: 1. **Name:** A unique name for the data source. 2. **Description:** Optional description. 3. **Base URL:** The base endpoint for the external API. 4. **Headers:** Optional header key-value pairs. Use the `+` button to add headers. 5. **Authentication:** Select the authentication mode required by the API. The following authentication modes are available in the UI: - `NONE` - `BEARER_TOKEN` - `API_KEY` - `HTTP_BASIC` - `OAUTH2_CLIENT`
![REST API Data Source](../../../../assets/images/guides/fs/data_source/rest_api_creation.png)
REST API data source form
!!! note The screenshot shows the form with `NONE` selected. When you choose another authentication mode, the form will prompt for the additional credentials required by that method. ### Step 3: Save the credentials After entering the connection details: 1. Click **Save Credentials**. 2. Click **Next: Select resource** to continue configuring the data source for downstream use. ## Next Steps Move on to the [usage guide for data sources](../usage.md) to see how you can use your newly created REST API data source. ================================================================================ # Unity Catalog Source: https://docs.hopsworks.ai/latest/user_guides/fs/data_source/creation/unity_catalog/ # How-To set up a Unity Catalog Data Source ## Introduction A Unity Catalog data source provides integration with [Databricks Unity Catalog](https://docs.databricks.com/aws/en/data-governance/unity-catalog/). Unity Catalog is Databricks' unified governance layer for data and AI assets, organised as a catalog → schema → table hierarchy. In this guide, you will configure a Data Source in Hopsworks that points at a Databricks workspace. Once configured, you can browse catalogs, schemas, and tables, and mount Delta tables as external Feature Groups whose data is read through the Arrow Flight query service. !!! warning "Databricks on AWS only" Unity Catalog is currently only supported on **Databricks on AWS**. Databricks on Azure and Databricks on GCP are not supported in this release. Their Unity Catalog temporary-table-credentials responses use cloud-specific credential shapes (Azure SAS tokens, GCP service-account tokens) that the Hopsworks Arrow Flight read path does not yet handle. If you point this connector at a non-AWS Databricks workspace the browse flow may succeed but every preview and feature-group read will fail. !!! note Only Delta-formatted Unity Catalog tables are supported in this release. Managed non-Delta tables, Iceberg tables, views, streaming tables, and materialised views are filtered out when browsing. !!! note Currently, it is only possible to create data sources in the Hopsworks UI. You cannot create a data source programmatically. !!! warning Direct Spark reads from Unity Catalog are not supported in this release. Reads flow through the Arrow Flight query service, which resolves each Delta table via the Unity Catalog REST API and reads it using the `deltalake` Python package (delta-rs) against the S3 location that Databricks returns. ## Prerequisites Before you begin you need all of the following. The first three are on the Databricks side and are the most common source of 400 / 403 errors from the read path. ### Databricks side - **External Data Access enabled on the metastore.** In the Databricks account console, go to Catalog → the metastore backing your workspace → Details, and turn on "External data access". Without this toggle, every call to `/api/2.1/unity-catalog/temporary-table-credentials` returns `403 Forbidden` for any principal. This is an account-admin setting, workspace admin alone cannot flip it. - **`EXTERNAL USE SCHEMA` grant on the schemas you want to read.** In Databricks SQL: ```sql GRANT EXTERNAL USE SCHEMA ON SCHEMA . TO ``; ``` where `` is the user (or service principal) that owns the PAT you are about to paste into Hopsworks. Without this grant the temporary-table-credentials endpoint returns `400 Bad Request`. - **`USE CATALOG`, `USE SCHEMA`, and `SELECT` grants** on the specific catalog / schema / tables you want to mount. - **Delta format.** Unity Catalog tables backed by Iceberg, non-Delta file formats, views, or streaming / materialised views cannot be read through this connector in v1. ### Hopsworks side - **Databricks workspace URL**, for example `https://.cloud.databricks.com`. - **Personal access token** for the principal to which the grants above were issued. - **A catalog name** containing the Delta tables you want to mount. It is optional but recommended to set this as the default catalog on the connector so the UI opens the browse view directly. - Optional: an **AWS region** (for example `us-west-2`). If you omit it, the backend guesses the region by parsing the STS session-token returned with the table credentials. For FIPS regions or for any workspace where the guess has been wrong once, set the region explicitly on the connector. The personal access token is stored encrypted in the Hopsworks `secrets` table. It is never written to the connector table in plaintext. PAT rotation is manual in this release: when the token expires, edit the connector and paste in a fresh one. Unity Catalog PATs are typically short-lived (hours to a few days), so expect to do this periodically. ## Feature flag Unity Catalog connectors are gated by the `enable_unity_catalog_storage_connectors` Hopsworks variable. An administrator must set it to `true` in the admin variables UI before the connector type appears in the create form. ## Creation in the UI ### Step 1: Set up new Data Source Head to the `Data Sources` view in Hopsworks (1) and click `New data source` (2). The `Add a data source` catalog opens below. Pick the `Databricks Unity Catalog` card to open the creation form.
![Data Source Creation](../../../../assets/images/guides/fs/data_source/data_source_overview.png)
The Data Source view in the user interface
### Step 2: Enter source details Enter the details for your Unity Catalog workspace. Start by giving it a unique **name** and an optional **description**. 1. The form opens with `Source` set to `Databricks Unity Catalog`. Click `Change source` to pick a different one. 2. **Databricks Workspace URL**: the full `https://` URL of your workspace. 3. **Access Token**: a Databricks personal access token; the field is masked and stored encrypted. 4. **Default Catalog**: optional; the Unity Catalog catalog to pre-select when browsing. 5. **AWS Region**: optional; set explicitly (for example `us-west-2`) when the backend's region guess from the STS session-token is wrong or your workspace is in a FIPS region. Leave empty to use the guess. 6. **Arguments**: optional key/value pairs passed through to the query service. 7. Click "Save Credentials". On save, Hopsworks calls the Unity Catalog `/catalogs` endpoint using the provided token; an HTTP 2xx response is required for the connector to be accepted. ### Step 3: Browse and mount a table After saving, open the connector and click **Configure**. The "Catalog" dropdown lists catalogs visible to your token; pick one and the schema/table browser lists Delta tables grouped by schema. Select a table to create an external Feature Group pointing at it. Reads will flow through the Arrow Flight query service. ## Next Steps Move on to the [usage guide for data sources](../usage.md) to see how you can use your newly created Unity Catalog connector. ================================================================================ # SAP HANA Source: https://docs.hopsworks.ai/latest/user_guides/fs/data_source/creation/sap_hana/ # How-To set up an SAP HANA Data Source { #data-source-sap-hana } ## Introduction SAP HANA is an in-memory relational database used by many enterprises as the system of record for ERP, CRM, and analytics workloads. An SAP HANA Data Source in Hopsworks stores the connection details required to read tables and views from a HANA tenant database. Once configured, you can use the same data source as the basis for an external (on-demand) Feature Group, or as the source for a dltHub-driven ingestion job that materialises HANA data into a managed Feature Group. In this guide, you will configure a Data Source in Hopsworks that holds the authentication information needed to connect to your SAP HANA database. !!! note Currently, it is only possible to create data sources in the Hopsworks UI. You cannot create a data source programmatically. ## Prerequisites Before you begin this guide you'll need to retrieve the following information from your SAP HANA tenant. The following options are **mandatory**: - **Host**: The hostname of the SAP HANA endpoint, for example `hxehost.example.com` for an on-premise instance or the endpoint shown in SAP BTP for SAP HANA Cloud. - **Port**: The SQL port of the tenant database. The default is `39015`, the SQL port for the first tenant database on a default multi-tenant or HANA Express (HXE) install (instance number 90). For a non-tenant single-host install (instance 00) use `30015`. SAP HANA Cloud typically uses `443`. Consult your DBA if you are unsure. - **User**: The HANA database user that the connector authenticates as. - **Password**: The password for that user. These are a few additional **optional** arguments: - **Database**: The tenant database name. Use this when your SAP HANA system hosts more than one tenant database and you need to target a specific one. - **Schema**: The default schema applied to unqualified queries on the connection. If you leave this empty, queries must fully qualify table names with the schema prefix. - **Table**: The default table the connector points at when no SQL query is provided. - **Application**: A short identifier surfaced in HANA's session tracing (`APPLICATION` session variable). This makes it easier to attribute load to Hopsworks in HANA monitoring tools. - **Additional arguments**: Free-form key/value options forwarded to the underlying SAP HANA Python driver (`hdbcli`) and the Spark JDBC reader. !!! info "Drivers" Hopsworks ships the SAP HANA drivers needed to read from HANA out of the box. The Hopsworks Spark image bundles the SAP `ngdbc` JDBC driver for Spark JDBC reads, and the dlt ingestion image and Arrow Flight server bundle SAP's `hdbcli` Python DBAPI driver. You do not need to install or upload the drivers yourself. ## Creation in the UI ### Step 1: Set up new Data Source Head to the `Data Sources` view on Hopsworks and click `New data source`. The `Add a data source` catalog opens below. Pick the `SAP HANA` card to open the creation form.
![Data Source Creation](../../../../assets/images/guides/fs/data_source/data_source_overview.png)
The Data Source View in the User Interface
### Step 2: Enter SAP HANA Settings Enter the details for your SAP HANA connector. Start by giving it a **name** and an optional **description**. 01. Select "SAP HANA" as storage. 02. Specify the **Host** of your SAP HANA endpoint. 03. Specify the **Port** the tenant SQL service listens on (default `39015`). 04. Provide the **User** name of the HANA database user. 05. Provide the **Password** for that user. 06. Optionally fill in **Database**, **Schema**, **Table**, and **Application**. 07. Optionally add additional key/value arguments. These are forwarded both to the Python driver used by the on-demand read path and to the Spark JDBC reader used by notebook jobs. 08. Click on "Save Credentials". ## Use it as an ingestion source Once the SAP HANA data source exists, you can also use it with the dltHub-based ingestion workflow described in [Ingest Data with dltHub][ingest-data-with-dlthub]. SAP HANA is treated as a SQL-like source, so the ingestion job supports both full and incremental loading. ## Type mapping Hopsworks reads each source column's HANA type from the cursor description and maps it to a Hopsworks offline feature type. The mapping preserves precision and scale where possible, so a source `DECIMAL(12, 2)` becomes a Hopsworks `decimal(12,2)` feature rather than collapsing to `bigint`. | SAP HANA type | Hopsworks offline feature type | | --- | --- | | `TINYINT` | `tinyint` | | `SMALLINT` | `smallint` | | `INTEGER` | `int` | | `BIGINT` | `bigint` | | `DECIMAL(p, s)` | `decimal(p,s)` | | `REAL` | `float` | | `DOUBLE` | `double` | | `BOOLEAN` | `boolean` | | `DATE` | `date` | | `TIME` | `timestamp` | | `TIMESTAMP` / `SECONDDATE` / `LONGDATE` | `timestamp` | | `CHAR` / `VARCHAR` / `NCHAR` / `NVARCHAR` / `TEXT` / `CLOB` / `NCLOB` / `ALPHANUM` | `string` | | `BINARY` / `VARBINARY` / `BLOB` | `binary` | ## Known limitations ### Avoid the `SYSTEM` schema for source tables Place tables you intend to ingest or expose as feature groups in a regular user schema (for example a project-specific `MYAPP` or `HOPSDEMO`). Tables created under the system-owned `SYSTEM` schema do not reflect cleanly through the SQLAlchemy HANA dialect that powers DLT ingestion. A typical setup is: ```sql CREATE SCHEMA HOPSDEMO; RENAME TABLE SYSTEM.MY_TABLE TO HOPSDEMO.MY_TABLE; ``` Then set **Schema** in the data source to `HOPSDEMO` (or pick it from the schema browser) and use that as the basis for any external feature group or DLT ingestion job. ### Online ingestion requires non-null primary keys When you create a managed Feature Group fed from SAP HANA via DLT and enable online serving, online ingestion validates that every row has a non-null value in the Feature Group's primary-key column. If the source rows can carry `NULL` in that column, either filter them out at source, pick a different primary key on the Feature Group, or disable online serving for the Feature Group. ### Authentication The SAP HANA data source currently supports username and password authentication. Certificate-based and JWT authentication are tracked as follow-up work. ## Next Steps Move on to the [usage guide for data sources][data-source-usage] to see how you can use your newly created SAP HANA connector. ================================================================================ # Usage Source: https://docs.hopsworks.ai/latest/user_guides/fs/data_source/usage/ # Data Source Usage Here, we look at how to use a Data Source after it has been created. Data Sources provide an important first step for integrating with external data. The 4 fundamental functionalities where data sources are used are: 1. Reading data into Spark Dataframes 2. Creating external feature groups 3. Writing training data 4. Creating managed feature groups We will walk through each functionality in the sections below. ## Retrieving a Data Source We retrieve a data source simply by its unique name. === "PySpark" ```python import hopsworks # Connect to the Hopsworks feature store project = hopsworks.login() feature_store = project.get_feature_store() # Retrieve data source ds = feature_store.get_data_source("data_source_name") ``` === "Scala" ```scala import com.logicalclocks.hsfs._ val connection = HopsworksConnection.builder().build(); val featureStore = connection.getFeatureStore(); // get directly via connector sub-type class, e.g., for GCS type val connector = featureStore.getGcsConnector("data_source_name") ``` ## Reading a Spark Dataframe from a Data Source One of the most common usages of a Data Source is to read data directly into a Spark Dataframe. It's achieved via the `read` API of the connector object, which hides all the complexity of authentication and integration with a data storage source. The `read` API primarily has two parameters for specifying the data source, `path` and `query`, depending on the data source type. The exact behaviour could change depending on the fdata source type, but broadly they could be classified as below ### Data lake/object based connectors For data sources based on object/file storage such as AWS S3, ADLS, GCS, we set the full object path in the `path` argument and users should pass a Spark data format (parquet, csv, orc, hudi, delta) to the `data_format` argument. === "PySpark" ```python # read data into dataframe using path df = connector.read( data_format="data_format", path="fileScheme://bucket/path/" ) ``` === "Scala" ```scala // read data into dataframe using path val df = connector.read("", "data_format", new HashMap(), "fileScheme://bucket/path/") ``` #### Prepare Spark API Additionally, for reading file based data sources, another way to read the data is using the `prepare_spark` method. This method can be used if you are reading the data directly through Spark. Firstly, it handles the setup of all Spark configurations or properties necessary for a particular type of connector and prepares the absolute path to read from, along with bucket name and the appropriate file scheme of the data source. A Spark session can handle only one configuration setup at a time, so Hopsworks cannot set the Spark configurations when retrieving the connector since it would lead to only always initialising the last connector being retrieved. Instead, user can do this setup explicitly with the `prepare_spark` method and therefore potentially use multiple connectors in one Spark session. `prepare_spark` handles only one bucket associated with that particular connector, however, it is possible to set up multiple connectors with different types as long as their Spark properties do not interfere with each other. So, for example a S3 connector and a Snowflake connector can be used in the same session, without calling `prepare_spark` multiple times, as the properties don’t interfere with each other. If the data source is used in another API call, `prepare_spark` gets implicitly invoked, for example, when a user materialises a training dataset using a data source or uses the data source to set up an External Feature Group. So users do not need to call `prepare_spark` every time they do an operation with a connector, it is only necessary when reading directly using Spark. Using `prepare_spark` is also not necessary when using the `read` API. For example, to read directly from a S3 connector, we use the `prepare_spark` as follows: === "PySpark" ```python connector.prepare_spark() spark.read.format("json").load("s3a://[bucket]/path") # or spark.read.format("json").load(connector.prepare_spark("s3a://[bucket]/path")) ``` ### Data warehouse/SQL based connectors For data sources accessed via SQL such as data warehouses and JDBC compliant databases, e.g., Redshift, Snowflake, BigQuery, JDBC, users pass the SQL query to read the data to the `query` argument. In most cases, this will be some form of a `SELECT` query. Depending on the connector type, users can also just set the table path and read the whole table without explicitly passing any SQL query to the `query` argument. This is mostly relevant for Google BigQuery. === "PySpark" ```python # read results from a SQL df = connector.read(query="SELECT * FROM TABLE") # or directly read a table if set on connector df = connector.read() ``` === "Scala" ```scala // read results from a SQL val df = connector.read("SELECT * FROM TABLE", "" , new HashMap(),"") ``` ### Streaming based connector For reading data streams, the Kafka Data Source supports reading a Kafka topic into Spark Structured Streaming Dataframes instead of a static Dataframe as in other connector types. === "PySpark" ```python df = connector.read_stream(topic="kafka_topic_name") ``` ## Creating an External Feature Group Another important aspect of a data source is its ability to facilitate creation of external feature groups with the [Connector API](../../../concepts/fs/feature_group/external_fg.md). [External feature groups](../feature_group/create_external.md) are basically offline feature groups and essentially stored as tables on external data sources. The `Connector API` relies on data sources behind the scenes to integrate with external datasource. This enables seamless integration with any data source as long as there is a data source defined. To create an external feature group, we use the `create_external_feature_group` API, also known as `Connector API`, and simply pass the data source created before to the `data_source` argument. Depending on the external source, we should set either the `query` argument for data warehouse based sources, or the `path` and `data_format` arguments for data lake based sources, similar to reading into dataframes as explained in above section. Example for any data warehouse/SQL based external sources, we set the desired SQL to `query` argument, and set the `data_source` argument to the data source object of desired data source. === "PySpark" ```python ds.query = "SELECT * FROM TABLE" fg = feature_store.create_external_feature_group( name="sales", version=1, description="Physical shop sales features", data_source=ds, primary_key=["ss_store_sk"], event_time="sale_date", ) ``` `Connector API` (external feature groups) only stores the metadata about the features within Hopsworks, while the actual data is still stored externally. This enables users to create feature groups within Hopsworks without the hassle of data migration. For more information on `Connector API`, read detailed guide about [external feature groups](../feature_group/create_external.md). When creating an external feature group from the UI, the review step also offers an **Add Trino Catalog** checkbox, which makes the same data source queryable from the query engine; see [Trino catalogs][trino-catalogs]. ## Ingesting Data into a Managed Feature Group Data Sources can also be used to create a managed feature group and ingest data from the source into Hopsworks. In this workflow, Hopsworks creates a sink-enabled feature group together with an ingestion job that copies data from the source into the feature group. This is different from an external feature group: - An **external feature group** keeps the data in the external source and stores only metadata in Hopsworks. - A **managed feature group with ingestion enabled** copies the source data into Hopsworks and can keep it synchronized through recurring ingestion jobs. This workflow is especially useful when you want to: - Materialize source data inside Hopsworks. - Schedule recurring ingestions. - Use full-load or incremental ingestion strategies. - Build managed feature groups from SQL, CRM, or REST API sources. For the full workflow, including schema selection, ingestion job configuration, loading strategies, and REST pagination, see [Ingest Data with dltHub][ingest-data-with-dlthub]. ## Writing Training Data Data Sources are also used while writing training data to external sources. While calling the [Feature View](../../../concepts/fs/feature_view/fv_overview.md) API `create_training_data`, we can pass the `data_source` argument which is necessary to materialise the data to external sources, as shown below. === "PySpark" ```python # materialise a training dataset version, job = feature_view.create_training_data( description="describe training data", data_format="spark_data_format", # e.g., data_format = "parquet" or data_format = "csv" write_options={"wait_for_job": False}, data_source=ds, ) ``` For a detailed walkthrough on managing and utilizing training data, refer to the [training data guide](../feature_view/training-data.md). ## Next Steps We have gone through the basic use cases of a data source. For more details about the API functionality for any specific connector type, checkout the [API section][hsfs.storage_connector.StorageConnector]. ================================================================================ # Feature Group User Guides Source: https://docs.hopsworks.ai/latest/user_guides/fs/feature_group/ # Feature Group User Guides A feature group is a table of features with a primary key and, usually, an event time. These guides cover creating one, keeping its data correct, and managing it over time.
- :material-table-plus:{ .lg .middle } **Start here** --- Create a feature group and insert a DataFrame. The schema is inferred from the DataFrame on the first insert. ```python fg = fs.get_or_create_feature_group( name="transactions", version=1, primary_key=["tid"], event_time="datetime", ) fg.insert(df) ``` [Create a feature group](create.md) · [Data types and schema](data_types.md) · [Statistics](statistics.md)
:material-table-plus:{ .hops-role-ico } Create and write { .hops-role-cap } - [Create a feature group](create.md) Offline and online tables, primary keys, event time, partitioning. - [External feature groups](create_external.md) Read data that stays in a warehouse or object store. - [Spine groups](create_spine.md) Supply keys, event times and labels without storing features. - [Ingest with dltHub](ingest_with_dlthub.md) Load from external sources through dlt pipelines. - [Data types and schema](data_types.md) Type mapping, adding features, schema versions.
:material-check-decagram-outline:{ .hops-role-ico } Trust { .hops-role-cap } - [Statistics](statistics.md) What is computed on insert and how to configure it. - [Data validation](data_validation.md) Great Expectations on insert, then the advanced guide and best practices. - [Feature monitoring](feature_monitoring.md) Scheduled statistics and comparison to a reference window. - [Online ingestion observability](online_ingestion_observability.md) Track rows arriving in the online store.
:material-cog-outline:{ .hops-role-ico } Manage { .hops-role-cap } - [On-demand transformations](on_demand_transformations.md) Compute features at request time from request parameters. - [Notifications](notification.md) Emit change events to a Kafka topic. - [Time to live](ttl.md) Expire rows after a retention period. - [Deprecate](deprecation.md) Mark a group as retired without deleting it.
================================================================================ # Create Source: https://docs.hopsworks.ai/latest/user_guides/fs/feature_group/create/ # How to create a Feature Group { #create-feature-group } ## Introduction In this guide you will learn how to create and register a feature group with Hopsworks. Feature groups are created from code with the Hopsworks APIs. The UI does not offer a creation flow; created feature groups appear in the project's `Catalog` section, where you can browse, edit and share them. ## Prerequisites Before you begin this guide we suggest you read the [Feature Group](../../../concepts/fs/feature_group/fg_overview.md) concept page to understand what a feature group is and how it fits in the ML pipeline. ## Create using the Hopsworks APIs To create a feature group using the Hopsworks APIs, you need to provide a Pandas, Polars or Spark DataFrame. The DataFrame will contain all the features you want to register within the feature group, as well as the primary key, event time and partition key. ### Create a Feature Group The first step to create a feature group is to create the API metadata object representing a feature group. Using the Hopsworks API you can execute: #### Batch Write API === "PySpark" ```python fg = feature_store.create_feature_group( name="weather", version=1, description="Weather Features", online_enabled=True, primary_key=["location_id"], partition_key=["day"], event_time="event_time", time_travel_format="DELTA", ) ``` You can read the full [`FeatureStore.create_feature_group`][hsfs.feature_store.FeatureStore.create_feature_group] documentation to get more details. If you need to create a feature group with vector similarity search supported, refer to [the vector similarity guide](../vector_similarity_search.md#extending-feature-groups-with-similarity-search). `name` is the only mandatory parameter of the `create_feature_group` and represents the name of the feature group. In the example above we created the first version of a feature group named *weather*, we provide a description to make it searchable to the other project members, as well as making the feature group available online. Additionally we specify which columns of the DataFrame will be used as primary key, partition key and event time. Composite primary key and multi level partitioning is also supported. The version number is optional, if you don't specify the version number the APIs will create a new version by default with a version number equals to the highest existing version number plus one. The last parameter used in the examples above is `stream`. The `stream` parameter controls whether to enable the streaming write APIs to the online and offline feature store. When using `time_travel_format="HUDI"` in a Python environment this behavior is the default. ##### Primary key A primary key is required when using a table format with time travel support (Hudi, Delta, or Iceberg) to store offline feature data. When inserting data in a feature group on the offline feature store, the DataFrame you are writing is checked against the existing data in the feature group. If a row with the same primary key is found in the feature group, the row will be updated. If the primary key is not found, the row is appended to the feature group. When writing data on the online feature store, existing rows with the same primary key will be overwritten by new rows with the same primary key. ##### Event time The event time column represents the time at which the event was generated. For example, with transaction data, the event time is the time at which a given transaction happened. In the context of feature pipelines, the event time is often also the end timestamp of the interval of events included in the feature computation. For example, computing the feature "number of purchases by customer last week", the event time should be the last day of this "last week" window. The event time is added to the primary key when writing to the offline feature store. This will make sure that the offline feature store has the entire history of feature values over time. As an example, if a user has made multiple purchases on a website, each of the purchases for a given user (identified by a user_id) will be saved in the feature group, with each purchase having a different event time (the combination of user_id and event_time makes up the primary key for the offline feature store). The event time **is not** part of the primary key when writing to the online feature store. This will ensure that the online feature store has the most recent version of the feature vector for each primary key. !!!note "Event time data type restriction" The supported data types for the event time column are: `timestamp`, `date` and `bigint`. ##### Partition key It is best practice to add a partition key. When you specify a partition key, the data in the feature group will be stored under multiple directories based on the value of the partition column(s). All the rows with a given value as partition key will be stored in the same directory. Choosing the correct partition key has significant impact on the query performance as the execution engine (Spark) will be able to skip listing and reading files belonging to partitions which are not included in the query. As an example, if you have partitioned your feature group by day and you are creating a training dataset that includes only the last year of data, Spark will read only 365 partitions and not the entire history of data. On the other hand, if the partition key is too fine grained (e.g., timestamp at millisecond resolution) - a large number of small partitions will be generated. This will slow down query execution as Spark will need to list and read a large amount of small directories/files. If you do not provide a partition key, all the feature data will be stored as files in a single directory. The system has a limit of 10240 direct children (files or other subdirectories) per directory. This means that, as you add new data to a non-partitioned feature group, new files will be created and you might reach the limit. If you do reach the limit, your feature engineering pipeline will fail with the following error: ```sh MaxDirectoryItemsExceededException - The directory item limit is exceeded: limit=10240 items=10240 ``` By using partitioning the system will write the feature data in different subdirectories, thus allowing you to write 10240 files per partition. `partition_key` is plain identity partitioning on existing columns. For partition transforms such as `day(ts)` or `bucket(16, customer_id)` on Iceberg and Hudi, liquid clustering on Delta (`clustered_by`), the Hudi bucket index (`bucket_index`), and z-ordering (`zorder_by`), see the [partitioning and clustering guide][partitioning-feature-group]. ##### Table format When you create a feature group, you can specify the table format you want to use to store the data in your feature group by setting the `time_travel_format` parameter. The currently supported values are `"HUDI"`, `"DELTA"`, `"ICEBERG"`, and `"NONE"` (which stores as Parquet without time travel support). The parameter defaults to `"DELTA"`. The feature group overview in the UI shows a **Table DDL** card with the generated Spark SQL `CREATE TABLE` statement for the offline table (including the table format and any partition columns), and, for online-enabled feature groups, the `CREATE TABLE` statement for the online (RonDB) table. ##### Data Source During the creation of a feature group, it is possible to define the `data_source` parameter, this allows for management of offline data in the desired table format outside the Hopsworks cluster. Currently, [S3][data-source-s3] and [GCS][data-source-gcs] connectors with `"DELTA"` or `"ICEBERG"` `time_travel_format` are supported. ##### Online Table Configuration When defining online-enabled feature groups it is also possible to configure the online table. You can specify [table options](https://docs.rondb.com/table_options/#table-options) by providing comments. Additionally, it is also possible to define whether online data is stored in memory or on disk using [table space](https://docs.rondb.com/disk_columns/#disk-columns). The code example shows the creation of an online-enabled feature group that stores online data on disk using `ts_1` table space and sets several table properties in the comment section. ```python fg = fs.create_feature_group( name="air_quality", description="Air Quality characteristics of each day", version=1, primary_key=["city", "date"], online_enabled=True, online_config={ "table_space": "ts_1", "online_comments": [ "NDB_TABLE=READ_BACKUP=1", "NDB_TABLE=PARTITION_BALANCE=FOR_RP_BY_LDM_X_2", ], }, ) ``` !!! note Table Space The table space needs to be provisioned at system level before it can be used. You can do so by adding the following parameters to the values.yaml file used for your deployment with the Helm Charts: ```yaml rondb: resources: requests: storage: diskColumnGiB: 2 ``` #### Streaming Write API As explained above, the stream parameter controls whether to enable the streaming write APIs to the online and offline feature store. For Python environments, only the stream API is supported (stream=True). === "Python" ```python fg = feature_store.create_feature_group( name="weather", version=1, description="Weather Features", online_enabled=True, primary_key=["location_id"], partition_key=["day"], event_time="event_time", time_travel_format="HUDI", ) ``` === "PySpark" ```python fg = feature_store.create_feature_group( name="weather", version=1, description="Weather Features", online_enabled=True, primary_key=["location_id"], partition_key=["day"], event_time="event_time", time_travel_format="HUDI", stream=True, ) ``` When using the streaming API, the data will be written directly to the online storage (if `online_enabled=True`). However, you can control when the sync to the offline storage is going to happen. You can do it synchronously after every call to `fg.insert()`, which is the default. Often, you defer writes to a later point in order to batch together multiple writes to the offline storage (useful to reduce the overhead of many small writes): ```python # run multiple inserts without starting the offline materialization job job, _ = fg.insert(df1, write_options={"start_offline_materialization": False}) job, _ = fg.insert(df2, write_options={"start_offline_materialization": False}) job, _ = fg.insert(df3, write_options={"start_offline_materialization": False}) # start the materialization job for all three inserts # note the job object is always the same, you don't need to call it three times job.run() ``` It is also possible to define the topics used for data ingestion, this can be done by setting the `topic_name` parameter with your preferred value. By default, feature groups in Hopsworks will share a project-wide topic. The topic can also be changed after the feature group has been created, see the [ingestion topic][feature-group-ingestion-topic] guide. #### Best Practices for Writing When designing a feature group, it is worth taking a look at how this feature group will be queried in the future, in order to optimize it for those query patterns. At the same time, Spark and Hudi tend to overpartition writes, creating too many small parquet files, which is inefficient and slows down writes. But they also slow down queries, because file listings take more time and reading many small files is slower than fewer larger files. The best practices described in this section hold both for the Streaming API and the Batch API. Four main considerations influence the write and the query performance: 1. Partitioning on a feature group level 2. Parquet file size within a feature group partition 3. Backfilling of feature group partitions 4. The choice of topic for data ingestion ##### Partitioning on a feature group level **Partitioning on the feature group level** allows Hopsworks and the table format (Hudi, Delta, or Iceberg) to push down filters to the filesystem when reading from feature groups. In practice that means fewer directories need to be listed and fewer files need to be read, speeding up queries. For example, most commonly, filtering is done on the event time column of a feature group when generating training data or batches of data: ```python query = fg.select_all() # create a simple feature view fv = fs.create_feature_view(name="transactions_view", query=query) # set up dates start_time = "2022-01-01" end_time = "2022-06-30" # create a training dataset version, job = fv.create_training_data( start_time=start_time, end_time=end_time, description="Description of a dataset", ) ``` Assuming the feature group was partitioned by a daily event time column, for example, the features are updated with a daily batch job, the feature store will only have to list and read the files in the directories of those six months that are being queried. !!! danger "Too granular event time columns" An event time column which is too granular, such as a timestamp, shouldn't be used as partition key. For example, a streaming pipeline generating features where the event time includes seconds, and therefore almost all event timestamps are unique can lead to many partition directories and small files, each of which contains only a few number of rows, which are inefficient to query even with pushed down filters. A good practice are partition keys with at most daily granularity, if they are based on time. Additionally, one can look at the size of a partition directory, which should be in the 100s of MB. Additionally, if you are commonly training models for different categories of your data, you can add another level of partitioning for this. That is, if the query contains an additional filter: ```python query = fg.select_all().filter(fg.country_code == "US") ``` The feature group can be created with the following partition key in order to push down filters also for the `country_code` category: ```python fg = feature_store.create_feature_group(... partition_key=['day', 'country_code'], event_time='day', ) ``` ##### Parquet file size within a feature group partition Once you have decided on the feature group level partitioning and you start inserting data to the feature group, there are multiple ways in order to influence how the table format (Hudi, Delta, or Iceberg) will **split the data between parquet files within the feature group partitions**. The two things that influence the number of parquet files per partition are 1. The number of feature group partitions written in a single insert 2. The shuffle parallelism used by the table format For example, the inserted dataframe (unique combination of partition key values) will be parallelized according to the following Hudi settings: !!! example "Default Hudi partitioning" ```python write_options = { "hoodie.bulkinsert.shuffle.parallelism": 5, "hoodie.insert.shuffle.parallelism": 5, "hoodie.upsert.shuffle.parallelism": 5, } ``` That means, using Spark, Hudi shuffles the data into five in-memory partitions, which each fill map to a task and finally a parquet file (see figure below). If the inserted Dataframe contains only a single feature group partition, this feature group partition will be written with five parquet files. If the inserted Dataframe contains multiple feature group partitions, the parquet files will be split among those partition, potentially more parquet files will be added. --8<-- "user_guides/fs/feature_group/create/partition-files.html" !!! tip "Setting shuffle parallelism" In practice that means the shuffle parallelism should be set equal to the number of feature group partitions in the inserted dataframe. This will create one parquet file per feature group partition, which in many cases is optimal. Theoretically, this rule holds up to a partition size of 2GB, which is the limit of Spark. However, one should bump this up accordingly already for smaller inputs. We recommend having shuffle parallelism `hoodie.[insert|upsert|bulkinsert].shuffle.parallelism` such that it's at least input_data_size/500MB. You can change the write options on every insert, depending also on the size of the data you are writing: ```python write_options = { "hoodie.bulkinsert.shuffle.parallelism": 5, "hoodie.insert.shuffle.parallelism": 5, "hoodie.upsert.shuffle.parallelism": 5, } fg.insert(df, write_options=write_options) ``` ##### Backfilling of feature group partitions Hudi scales well with the number of partitions to write, when performing backfilling of old feature partitions, meaning moving backwards in time with the event-time, it makes sense to **batch those feature group partitions** together into a single `fg.insert()` call. As shown in the figure above, the number of utilised executors you choose for the insert depends highly on the number of partitions and shuffle parallelism you are writing. So by writing multiple feature group partitions in a single insert, you can scale up your Spark application and fully utilise the workers. In that case you can increase the Hudi shuffle parallelism accordingly. !!! danger "Concurrent feature group inserts" Hopsworks 3.1 and earlier, currently does not support concurrent inserts to feature groups. This means that if your feature pipeline writes to one feature group partition at a time, you cannot run it multiple times in parallel for backfilling. The recommended approach is to unionise the dataframes and insert them with a single `fg.insert()` instead. For clients that write with the Stream API, it is enough to defer starting the backfill job until after multiple inserts, [as described above](#streaming-write-api). ##### The choice of topic for data ingestion When creating a feature group that uses streaming write APIs for data ingestion it is possible to define the Kafka topics that should be utilized. The default approach of using a project-wide topic functions great for use cases involving little to no overlap when producing data. However, concurrently inserting into multiple feature groups could cause read amplification for the offline materialization job (e.g., Hudi Delta Streamer). The job of a feature group consumes every record the shared topic received since its last run, and only then discards the records whose `featureGroupId` header belongs to another feature group. One large or frequently written feature group therefore slows down the materialization job of every other feature group on its topic, in proportion to how much it writes. Therefore, it is advised to utilize separate topics when ingestions overlap or there is a large frequently running insertion into a specific feature group. If you only notice the read amplification once the feature group is in use, the [ingestion topic][feature-group-ingestion-topic] guide explains how to move it to its own topic. ### Register the metadata and save the feature data The snippet above only created the metadata object on the Python interpreter running the code. To register the feature group metadata and to save the feature data with Hopsworks, you should invoke the `insert` method: ```python fg.insert(df) ``` The save method takes in input a Pandas, Polars or Spark DataFrame. Hopsworks will use the DataFrame columns and types to determine the name and types of features, primary key, partition key and event time. The DataFrame *must* contain the columns specified as primary keys, partition key and event time in the `create_feature_group` call. If a feature group is online enabled, the `insert` method will store the feature data to both the online and offline storage. !!! api "API reference" - [`FeatureStore.create_feature_group`][hsfs.feature_store.FeatureStore.create_feature_group] - [`FeatureStore.get_or_create_feature_group`][hsfs.feature_store.FeatureStore.get_or_create_feature_group] - [`FeatureGroup`][hsfs.feature_group.FeatureGroup] - [`insert`][hsfs.feature_group.FeatureGroup.insert] - [`read`][hsfs.feature_group.FeatureGroup.read] - [`select_all`][hsfs.feature_group.FeatureGroupBase.select_all] - [`filter`][hsfs.feature_group.FeatureGroupBase.filter] Browse the full Python API :material-arrow-right: ## Find your feature group in the UI Feature groups created through the APIs appear in the `Catalog` section of the project sidebar. From there you can inspect features and statistics, edit metadata, and manage sharing and tags.
The Catalog listing the project's feature groups with their format, online and shared badges
The Catalog lists every feature group with its table format, online status and version.
================================================================================ # Partitioning and Clustering Source: https://docs.hopsworks.ai/latest/user_guides/fs/feature_group/partitioning/ # How to partition and cluster a Feature Group { #partitioning-feature-group } ## Introduction In this guide you will learn how to lay out the offline data of a feature group. Each layout mechanism is its own creation-time parameter, and each parameter maps to exactly one native mechanism of the table format: | Parameter | What it configures | Formats | | --- | --- | --- | | `partitioned_by` | A native partition specification from transform expressions, such as `["day(ts)", "bucket(16, customer_id)"]`. | `ICEBERG` (full transform set), `HUDI` (identity and time grains). Rejected on `DELTA`, which has no partition transforms. | | `clustered_by` | Delta liquid clustering columns, such as `["ts", "customer_id"]`. | `DELTA` only. | | `bucket_index` | The Hudi bucket index, as `{"field": , "num_buckets": N}`. | `HUDI` only. | | `zorder_by` | Columns to z-order the data files by. | `ICEBERG` (applied by [`FeatureGroup.optimize`][hsfs.feature_group.FeatureGroup.optimize]), `HUDI` (inline clustering). Rejected on `DELTA`, where `clustered_by` covers the use case. | | `sort_order` | A persistent write sort order, such as `["merchant_id asc", "amount desc nulls last"]`. | `ICEBERG` only, and requires a Spark environment at creation. | | `partition_key` | Plain identity partitioning on existing columns. | All formats; mutually exclusive with `partitioned_by` and with `clustered_by`. | ## Prerequisites Before you begin this guide we suggest you read the [Feature Group](../../../concepts/fs/feature_group/fg_overview.md) concept page and the [create feature group][create-feature-group] guide, which covers `partition_key`, `event_time`, and `time_travel_format`. ## Partition transforms Each element of `partitioned_by` is one transform expression: | Expression | Behaviour | Typical use | | --- | --- | --- | | `col` or `identity(col)` | The original column value. | Low-cardinality columns. | | `bucket(N, col)` | Murmur3 hash of the column modulo N. | High-cardinality ids. | | `truncate(W, col)` | The value truncated to width W (numeric ranges, string prefixes). | Range-like grouping. | | `year(col)` | Year of a timestamp or date column. | Very large historical tables. | | `month(col)` | Month. | Monthly queries. | | `day(col)` | Calendar day. | Event and transaction tables. | | `hour(col)` | Hour of a timestamp column. | High-volume streams. | | `week(col)` | ISO week. | HUDI only: Iceberg has no week transform. | | `void(col)` | Always null. | ICEBERG only: partition spec evolution placeholder. | Parsing is whitespace tolerant and the expressions are stored in a canonical lowercase form without spaces, so `bucket(16, customer_id)` reads back as `bucket(16,customer_id)`. A column that happens to be named after a grain (`year`, `month`, `week`, `day`, `hour`) must use the explicit form `identity(year)`, because the bare name is reserved for the legacy grain migration error; the canonical form keeps the explicit `identity()` for these columns. Transforms combine freely, for example one temporal transform plus one bucket, with the format-specific restrictions listed below. On Iceberg an expression can name its partition field with an alias, for example `"bucket(16, customer_id) as shard"`; Iceberg metadata tables and partition spec evolution refer to fields by name. Without an alias the field gets Iceberg's generated default name (`customer_id_bucket`, `ts_day`, and so on), and all field names in one spec must be unique, so two bucket transforms on the same column need an alias on one of them. Aliases are rejected on `HUDI`, whose partition paths are always named after the source column or grain. Not every transform is available on every format, and `DELTA` rejects `partitioned_by` entirely: | Transform | ICEBERG | HUDI | | --- | --- | --- | | `identity` | yes | yes | | `bucket` | yes | no, use `bucket_index` | | `truncate` | yes | no | | `year`/`month`/`day`/`hour` | yes | yes, and the column must be the event time | | `week` | no | yes, and the column must be the event time | | `void` | yes | no | ## Iceberg: hidden partitioning On Iceberg the transform list compiles into the table's native partition spec. No derived columns are added to the feature group schema, and your DataFrame carries only the real columns. Because the partitioning is hidden, every engine reading the table prunes partitions from ordinary predicates on the source columns. ```python fg = feature_store.create_feature_group( name="orders", version=1, description="Order events, day-partitioned and bucketed by customer", primary_key=["order_id"], event_time="order_ts", time_travel_format="ICEBERG", partitioned_by=["day(order_ts)", "bucket(16, customer_id)"], ) fg.insert(orders_df) ``` A read that filters on the source columns plans only the matching day partition and, within it, one bucket of sixteen: ```python df = ( fg.select_all() .filter(fg.get_feature("order_ts") >= "2026-07-03") .filter(fg.get_feature("order_ts") < "2026-07-04") .filter(fg.get_feature("customer_id") == 1042) .read() ) ``` Iceberg allows at most one temporal transform per source column, because a finer grain already supports coarser pruning: use `day(ts)` alone rather than `year(ts)` plus `day(ts)`. ### Persistent sort order `sort_order` sets the Iceberg table's persistent write sort order, so every write organizes rows before laying out files. Each element is a column with an optional direction and null ordering: `"col"`, `"col desc"`, `"col asc nulls last"`. Defaults follow Iceberg: ascending, nulls first when ascending and nulls last when descending. ```python fg = feature_store.create_feature_group( name="payments", version=1, primary_key=["payment_id"], event_time="ts", time_travel_format="ICEBERG", partitioned_by=["day(ts)"], sort_order=["merchant_id asc", "amount desc nulls last"], ) ``` A sort order is applied through the Iceberg Java API, so creating a feature group with `sort_order` requires a Spark environment; the pure Python and catalog creation paths reject it rather than silently dropping it. `sort_order` is the right choice when reads consistently filter or join on the same columns; `zorder_by` covers multi-dimensional point lookups instead. The two are mutually exclusive as stored defaults, because they prescribe conflicting file layouts (writes sorted linearly while maintenance rewrites on the z-curve); a one-off z-order on a sorted table stays available through `optimize(strategy="zorder", columns=[...])`. ### Layout for point-in-time training data Training data from a feature view runs a point-in-time join: each feature group is joined to the label side on the primary key with an event-time inequality, and a rank window keeps the latest row per label. Inside that query shape, an event-time bound on the feature group is pushed into the Iceberg scan, so `day(event_ts)` partitioning prunes the history scan to the bounded window. Set the bound with a `lookback` on the feature view read (or an explicit `event_ts >=` query filter); without one, point-in-time correctness requires scanning all history, and no partitioning can prune it. The layout that serves this access pattern: ```python fg = feature_store.create_feature_group( ..., time_travel_format="ICEBERG", partitioned_by=["day(event_ts)", "bucket(16, customer_id)"], sort_order=["customer_id asc", "event_ts desc"], ) ``` - `day(event_ts)` prunes the scan to the lookback window. - `bucket(N, primary_key)` keeps each day's data grouped by key, bounding the rows any one join task reads. - `sort_order` (or `zorder_by` plus a scheduled `optimize()`) clusters rows by key inside each file, so file-level min/max statistics skip files for keys not present in the label set. - Prefer `insert` (append) over upserts for event history: the Iceberg upsert rewrites the table and discards the maintained clustering until the next `optimize()`. Size the partitioning to the data volume: each `(day, bucket)` combination becomes at least one file, so a small feature group with fine-grained partitioning produces many tiny files and the task overhead outweighs the pruning. As a rule of thumb, choose the day grain and bucket count so partitions land in the hundreds of megabytes; for small feature groups skip `bucket()` or use a coarser time grain. ## Delta: liquid clustering with clustered_by Delta has no partition transforms, so `partitioned_by` is rejected there; the layout mechanism is liquid clustering, configured with `clustered_by` as a plain column list. This matches Delta's own API, where `clusterBy(...)` and `partitionedBy(...)` are different layout mechanisms. Liquid clustering supports at most 4 columns, and `clustered_by` cannot be combined with `partition_key`, because Delta does not support clustering a hive-partitioned table. ```python fg = feature_store.create_feature_group( name="transactions", version=1, primary_key=["tx_id"], event_time="ts", time_travel_format="DELTA", clustered_by=["ts", "customer_id"], ) ``` Data skipping on the clustered columns replaces directory-style partitioning, so no derived columns are added and reads need no special predicates. Delta only clusters on columns that have data-skipping statistics, which it collects for the first 32 columns by default; when a clustering column sits past that position Hopsworks widens `delta.dataSkippingNumIndexedCols` to cover it, so any schema column can be clustered. !!!warning "Clustered Delta feature groups are writable by Spark only" Liquid clustering uses the Clustering and DomainMetadata Delta writer table features, which delta-rs (the pure Python write path) does not implement, so treating them as optional would corrupt the table contract. Creating a clustered feature group from a Python environment with `stream=False` fails, and delta-rs writes, deletes, and optimize raise on clustered feature groups. From Python, pass `stream=True` so writes go through the Spark materialization job, or write from a Spark job or notebook. ## Hudi: grain columns and the bucket index Hudi has no hidden partitioning, so temporal transforms materialize as integer partition columns named after the grain (`year`, `month`, ...), derived from the event time on every write. Your DataFrame must not contain these columns; the write path computes them. The grain columns appear in the feature group schema flagged as partition columns, and by default they are stored offline only (set `online_partition_columns=True` to include them online). Bucketing on Hudi is not a partition transform: it is the Hudi bucket index, which hashes a primary key field into a fixed number of buckets, configured with the `bucket_index` parameter: ```python fg = feature_store.create_feature_group( name="transactions", version=1, primary_key=["tx_id"], event_time="ts", time_travel_format="HUDI", partitioned_by=["year(ts)", "month(ts)"], bucket_index={"field": "tx_id", "num_buckets": 16}, ) ``` `bucket_index` injects the corresponding write options: ```text hoodie.index.type=BUCKET hoodie.bucket.index.num.buckets=N hoodie.bucket.index.hash.field=col ``` The optional `"engine"` key accepts only `"simple"` (the default): Hudi's consistent-hashing bucket engine requires a merge-on-read table with a clustering lifecycle, and Hopsworks Hudi feature groups are copy-on-write. You can equally set these options, including a partition-level bucket index with per-partition `hoodie.bucket.index.*` overrides, directly through `write_options` on insert. On Hudi the platform rewrites `event_time` range filters into grain-column predicates at query time, so reads that filter on the event time prune partitions without referencing the grain columns. ## Z-ordering with zorder_by `zorder_by` records up to 4 columns to z-order the data files by, so point lookups on those columns inside a partition skip most files. It is supported for `ICEBERG` and `HUDI`; on `DELTA` it is rejected because `clustered_by` covers the same use case. ```python fg = feature_store.create_feature_group( name="ad_clicks", version=1, description="Click stream, hour-partitioned, z-ordered by user and campaign", primary_key=["click_id"], event_time="click_ts", time_travel_format="ICEBERG", partitioned_by=["hour(click_ts)"], zorder_by=["user_id", "campaign_id"], ) fg.insert(clicks_df) fg.optimize() ``` On Iceberg, z-order is a rewrite strategy rather than a write-time property: writes stay cheap, and calling [`FeatureGroup.optimize`][hsfs.feature_group.FeatureGroup.optimize] rewrites the data files ordered on the z-curve of the `zorder_by` columns. Run it from a scheduled job after heavy ingestion. The initial z-order over a backfill needs `optimize(rewrite_all=True)` to rewrite every existing file; routine maintenance calls default to `rewrite_all=False` so they never rewrite the whole table by accident. On Hudi, `zorder_by` configures inline clustering, which applies the z-order layout as part of the write pipeline, so no explicit call is needed. ## Optimizing the layout [`FeatureGroup.optimize`][hsfs.feature_group.FeatureGroup.optimize] rewrites the offline data files to apply the feature group's layout and returns the format's rewrite metrics: | `time_travel_format` | What `optimize()` does | | --- | --- | | `ICEBERG` | An Iceberg `rewriteDataFiles` action; requires a Spark environment. `strategy` picks `"zorder"` (over `columns`, defaulting to `zorder_by`), `"sort"` (the persistent `sort_order`), or `"binpack"`; unset, it follows the stored layout in that order. `rewrite_all=True` rewrites every file regardless of the planner thresholds (defaults to False for every strategy, so a routine call is incremental; pass it for the initial full z-order), `target_file_size_mb` overrides the target file size, and `where` restricts the rewrite to the matching files through an Iceberg filter expression over the feature group's columns. | | `DELTA` | `OPTIMIZE`, which incrementally clusters a liquid-clustered table; `optimize(full=True)` runs `OPTIMIZE FULL` to recluster all existing data after the clustering columns changed (clustered tables only), and `where` restricts the rewrite with a predicate. `strategy="zorder"` with `columns` runs the legacy `OPTIMIZE ... ZORDER BY`, which Delta only supports on unclustered tables, because z-order and liquid clustering are incompatible. Clustered feature groups require Spark; from pure Python only unclustered compaction is available. | | `HUDI` | Rejected: layout maintenance runs through inline clustering on writes. | ### Catalog-backed Iceberg tables The Iceberg feature is fully supported on the default path-based (`HadoopTables`) layout. Some operations are not yet available when the table is backed by an external catalog: | Operation | Path-based | Glue Data Catalog | User-provided catalog (`iceberg.catalog`) | | --- | --- | --- | --- | | Create with `partitioned_by` | yes | yes | yes | | Create with `sort_order` | yes | no (rejected at creation) | no | | `optimize()` | yes | no (run the catalog's `rewrite_data_files` procedure) | no | | `update_partition_spec()` | yes | no (evolve through the catalog) | no | | Introspection (`get_partition_spec`, `describe_layout`, ...) | yes | yes | no (inspect through the catalog) | Where an operation is unavailable the call raises with a pointer to the catalog-side equivalent rather than silently doing nothing. ## Evolving the layout Layout is not fixed at creation; each format's native evolution is exposed, and all three methods require a Spark environment (from pure Python they raise, run them from a Spark job or notebook): - [`FeatureGroup.update_partition_spec`][hsfs.feature_group.FeatureGroup.update_partition_spec] evolves an Iceberg partition spec. Evolution is metadata-only: existing data keeps its old layout and new writes use the evolved spec, so no data is rewritten. ```python fg.update_partition_spec(add=["hour(ts)"], remove=["day(ts)"]) ``` - [`FeatureGroup.update_clustering`][hsfs.feature_group.FeatureGroup.update_clustering] changes the Delta clustering columns. The change affects new writes; run `optimize(full=True)` to recluster existing data. ```python fg.update_clustering(["ts", "customer_id"]) fg.optimize(full=True) ``` - [`FeatureGroup.disable_clustering`][hsfs.feature_group.FeatureGroup.disable_clustering] turns Delta clustering off (`CLUSTER BY NONE`); existing data keeps its layout. Hudi partitions are physical directories and cannot evolve; `update_partition_spec` is rejected there. Partition spec evolution requires a feature group created with `partitioned_by`; a feature group using `partition_key` is rejected, because its identity partitions are recorded on the features themselves and cannot be restated as an evolvable spec. After each evolution the committed spec is read back from the table and persisted as the feature group's stored metadata (`partitioned_by`, `clustered_by`), so read-back always reflects the current layout, including changes made by external engines. If the table change succeeds but the metadata update fails, the error says exactly that and names the applied layout; calling `update_partition_spec()` with no arguments re-syncs the stored metadata from the table without changing it. ## Inspecting the layout The stored metadata describes what was requested; the table itself is the source of truth for what is physically there. Five methods read the actual table state, so drift (for example, an external engine evolving the spec) is visible: - [`FeatureGroup.get_partition_spec`][hsfs.feature_group.FeatureGroup.get_partition_spec] returns the current Iceberg partition spec fields (`ICEBERG` only). - [`FeatureGroup.get_partition_specs`][hsfs.feature_group.FeatureGroup.get_partition_specs] returns the Iceberg spec history, oldest first (`ICEBERG` only). - [`FeatureGroup.get_sort_order`][hsfs.feature_group.FeatureGroup.get_sort_order] returns the persistent sort order (`ICEBERG` only). - [`FeatureGroup.get_clustering_columns`][hsfs.feature_group.FeatureGroup.get_clustering_columns] returns the actual Delta clustering columns (`DELTA` only, Spark required). - [`FeatureGroup.describe_layout`][hsfs.feature_group.FeatureGroup.describe_layout] returns the stored metadata next to the actual table state for any format. The format-specific getters raise on the wrong format rather than returning `None`, so an unconfigured layout is never confused with an unsupported operation. For Iceberg feature groups written through a user-provided catalog (the `iceberg.catalog` write option), the current-metadata pointer lives in that catalog, so inspect the layout through the catalog instead; Glue-backed feature groups are inspected through the Glue Data Catalog automatically. ## Migrating from the grain list form Earlier Hopsworks versions accepted `partitioned_by` as a list of bare grain names, for example `["year", "month"]`. That form is no longer valid, because a bare element now means an identity transform on a column of that name. Creation fails with an error pointing at the replacement: write the grain as a transform on your event time column, for example `["year(event_ts)", "month(event_ts)"]`. Feature groups created with the old form keep working for reads; their stored layout is unchanged. Writes and deletes to such feature groups fail with a migration-required error, because the new write path no longer derives the grain values and continuing would silently change the physical layout; recreate the feature group with the transform grammar to write again. ## Restrictions - `partitioned_by` requires `time_travel_format` `ICEBERG` or `HUDI`; `clustered_by` requires `DELTA`; `bucket_index` requires `HUDI`; `zorder_by` requires `ICEBERG` or `HUDI`; `sort_order` requires `ICEBERG` and a Spark environment at creation. - `partitioned_by` and `partition_key` cannot be combined, and neither can `clustered_by` and `partition_key`, nor `sort_order` and `zorder_by`. - Temporal transforms require a `date` or `timestamp` source column, and `hour` requires a `timestamp`, because a date has no sub-day resolution. - `bucket` and `truncate` follow the Iceberg source-type restrictions: `float`, `double`, and `boolean` sources are rejected, and `truncate` additionally rejects `date` and `timestamp`. - Partition-field aliases are Iceberg-only, and all field names in one spec (explicit aliases and generated defaults) must be unique. - The `bucket_index` field must be part of the primary key, because the Hudi bucket index hashes the record key; its `engine` accepts only `"simple"`. - Clustered Delta feature groups are writable by Spark only; from Python use `stream=True`. - Stream feature groups support `partitioned_by` and `clustered_by` on `ICEBERG` and `DELTA`, but `partitioned_by` is rejected on `HUDI`. - Online-enabled feature groups support `partitioned_by` on `ICEBERG`, but not on `HUDI`, because the Hudi grain columns are not part of the online schema. ================================================================================ # Delta Maintenance Source: https://docs.hopsworks.ai/latest/user_guides/fs/feature_group/delta_maintenance/ # How to maintain a Delta Feature Group { #delta-maintenance-feature-group } ## Introduction A Delta table that is written to repeatedly accumulates two things: data files and log entries. Every commit writes at least one new data file, and every reader opens all of them. Every commit also appends to the `_delta_log`, and a reader replays that log from the last checkpoint. Neither is reclaimed on its own, and on a table written from Python neither is bounded on its own either: Spark writes a checkpoint every `delta.checkpointInterval` commits, delta-rs writes none. Four methods on a feature group bound them. They apply only to feature groups with `time_travel_format="DELTA"` and return `None` for any other format. | Method | What it does | | --- | --- | | `delta_optimize` | Rewrites many small files into fewer large ones. Also available as `delta_compact`. | | `delta_checkpoint` | Writes a checkpoint, so readers stop replaying the log from commit zero. | | `delta_cleanup_metadata` | Expires the log entries a checkpoint already covers. | | [`delta_vacuum`][hsfs.feature_group.FeatureGroup.delta_vacuum] | Deletes the data files no retained version references. | Each dispatches on the engine, so the same call works from a Python client with delta-rs and from a PySpark job with Delta Spark. The first three are rendered as plain code rather than API links until the client release that ships them, because the docs build resolves cross-references against the released client. ## Prerequisites Before you begin this guide we suggest you read the [Feature Group](../../../concepts/fs/feature_group/fg_overview.md) concept page and the [create feature group][create-feature-group] guide. ## The maintenance sequence Run them in this order. ```python fg = fs.get_feature_group("transactions", version=1) fg.delta_optimize(max_concurrent_tasks=1) fg.delta_checkpoint() fg.delta_cleanup_metadata() fg.delta_vacuum(retention_hours=168) ``` The order is what makes each step safe. Compaction replaces many small files with few large ones and leaves the old ones on disk, still referenced by older versions. The checkpoint goes next, so the smaller file list is recorded before anything is deleted. Only then the two deletions: the log entries the checkpoint now covers, and the data files the compaction orphaned. ## Choosing a retention `delta_vacuum` deletes files that versions inside the retention window no longer reference. A query that is already running holds no lock on those files, so the retention has to stay comfortably longer than the longest query that runs against the group. It is also the time travel window: a version whose files have been vacuumed cannot be read, which is why a compaction has to be followed by a checkpoint. The effect of a short retention is not that a vacuum deletes more, but that it deletes sooner. A run reclaims what earlier runs orphaned rather than its own rewrite, whose files are seconds old. !!! warning "Delta's own floor" Delta refuses a retention under seven days unless its retention check is disabled. Hopsworks disables that check for you so a shorter retention takes effect, which means the value you pass is the value that applies. Pick it against your own readers rather than relying on the engine to refuse a bad one. ## Compacting only what changed On a table partitioned by a date column, `after_ingest_date` bounds the rewrite to partitions at or after that date. ```python fg.delta_optimize(after_ingest_date="2026-09-10") ``` Use it for anything that runs on a schedule. Only files written since the last compaction need rewriting, and on a date-partitioned table they are all at or after that date, so bounding the rewrite this way keeps its cost flat. Without it every run rewrites the whole table, including everything earlier runs already compacted, and the cost grows with the table forever. Leave a day of slack for rows that arrived late. Only a partition column can select files without reading them, so this is refused on a group that is not partitioned by a date. Compact the whole table by leaving `after_ingest_date` unset. ## When to run them For an append-heavy table, compact when the active file count crosses a threshold and otherwise once a day. Around 100 files is the low hundreds of megabytes at typical commit sizes, near the engine's own target file size. Read the last compaction time from the table's own history rather than keeping state, so the schedule survives restarts and multiple writers. These can run from a [Hopsworks job](../../projects/jobs/pyspark_job.md) on a schedule. Compaction is the only one of the four that a deployment reading the same table notices: measured beside live traffic it roughly doubled p99 for the few seconds it ran, while the median moved by a tenth of a millisecond. The other three sat where the deployment sat with nothing running. ================================================================================ # Create External Source: https://docs.hopsworks.ai/latest/user_guides/fs/feature_group/create_external/ # How to create an External Feature Group { #create-external-feature-group } ## Introduction In this guide you will learn how to create and register an external feature group with Hopsworks. This guide covers creating an external feature group using the Hopsworks APIs as well as the user interface. ## Prerequisites Before you begin this guide we suggest you read the [External Feature Group](../../../concepts/fs/feature_group/external_fg.md) concept page to understand what a feature group is and how it fits in the ML pipeline. ## Create using the Hopsworks APIs ### Retrieve the Data Source To create an external feature group using the Hopsworks APIs you need to provide an existing [data source](../data_source/index.md). === "Python" ```python ds = feature_store.get_data_source("data_source_name") ``` ### Create an External Feature Group The first step is to instantiate the metadata through the `create_external_feature_group` method. Once you have defined the metadata, you can [persist the metadata and create the feature group](#register-the-metadata) in Hopsworks by calling `fg.save()`. #### SQL based external feature group === "Python" ```python query = """ SELECT TO_NUMERIC(ss_store_sk) AS ss_store_sk , AVG(ss_net_profit) AS avg_ss_net_profit , SUM(ss_net_profit) AS total_ss_net_profit , AVG(ss_list_price) AS avg_ss_list_price , AVG(ss_coupon_amt) AS avg_ss_coupon_amt , sale_date , ss_store_sk FROM STORE_SALES GROUP BY ss_store_sk, sales_date """ fg = feature_store.create_external_feature_group( name="sales", version=1, description="Physical shop sales features", query=query, data_source=ds, primary_key=["ss_store_sk"], event_time="sale_date", ) fg.save() ``` #### Data Lake based external feature group === "Python" ```python fg = feature_store.create_external_feature_group( name="sales", version=1, description="Physical shop sales features", data_format="parquet", data_source=ds, primary_key=["ss_store_sk"], event_time="sale_date", ) fg.save() ``` You can read the full [`FeatureStore.create_external_feature_group`][hsfs.feature_store.FeatureStore.create_external_feature_group] documentation for more details. `name` is a mandatory parameter of the `create_external_feature_group` and represents the name of the feature group. The version number is optional, if you don't specify the version number the APIs will create a new version by default with a version number equals to the highest existing version number plus one. If the data source is defined for a data warehouse (e.g., JDBC, Snowflake, Redshift) you need to provide a SQL statement that will be executed to compute the features. If the data source is defined for a data lake, the location of the data as well as the format need to be provided. Additionally we specify which columns of the DataFrame will be used as primary key, and event time. Composite primary keys are also supported. ### Register the metadata In the snippet above it's important that the created metadata object gets registered in Hopsworks. To do so, you should invoke the `save` method: === "Python" ```python fg.save() ``` ### Enable online storage You can enable online storage for external feature groups, however, the sync from the external storage to Hopsworks online storage is not automatic and needs to be setup manually. For an external feature group to be available online, during the creation of the feature group, the `online_enabled` option needs to be set to `True`. === "Python" ```python external_fg = fs.create_external_feature_group( name="sales", version=1, description="Physical shop sales features", query=query, data_source=ds, primary_key=["ss_store_sk"], event_time="sale_date", online_enabled=True, ) external_fg.save() # read from external storage and filter data to sync to online df = external_fg.read().filter(external_fg.customer_status == "active") # insert to online storage external_fg.insert(df) ``` The `insert()` method takes a DataFrame as parameter and writes it _only_ to the online feature store. Users can select which subset of the feature group data they want to make available on the online feature store by using the [query APIs][hsfs.constructor.query.Query]. ### Limitations Hopsworks Feature Store does not support time-travel queries on external feature groups. Additionally, support for `.read()` and `.show()` methods when using by the Python engine is limited to external feature groups defined on BigQuery and Snowflake and only through the ArrowFlight Server with DuckDB, which Hopsworks enables by default. Nevertheless, external feature groups defined top of any data source can be used to create a training dataset from a Python environment invoking one of the following methods: [`FeatureView.create_training_data`][hsfs.feature_view.FeatureView.create_training_data], [`FeatureView.create_train_test_split`][hsfs.feature_view.FeatureView.create_train_test_split] or [`FeatureView.create_train_validation_test_split`][hsfs.feature_view.FeatureView.create_train_validation_test_split]. !!! api "API reference" - [`FeatureStore.get_data_source`][hsfs.feature_store.FeatureStore.get_data_source] - [`FeatureStore.create_external_feature_group`][hsfs.feature_store.FeatureStore.create_external_feature_group] - [`ExternalFeatureGroup`][hsfs.feature_group.ExternalFeatureGroup] - [`save`][hsfs.feature_group.ExternalFeatureGroup.save] - [`insert`][hsfs.feature_group.ExternalFeatureGroup.insert] - [`read`][hsfs.feature_group.ExternalFeatureGroup.read] Browse the full Python API :material-arrow-right: ## Create using the UI You can also create a new feature group through the UI. For this, navigate to the `Data Sources` section and make sure you have a data source for the desired platform, or create a [new](../data_source/index.md) one. Table browsing is available for database and warehouse sources such as Snowflake, BigQuery, Redshift and SQL databases; the built-in HopsFS and JDBC sources of a project do not offer it.

Data Sources list

Open the data source with the pencil at the end of its row and click `Next: Select Tables` at the bottom of the form.

Edit data source form with the Next: Select Tables button

In the UI you can either select one or more tables or define a custom SQL query. ### Option A: Select tables The database navigation structure depends on your specific data source. You'll navigate through the appropriate hierarchy for your platform, such as Database → Schema → Table for Snowflake, or Project → Dataset → Table for BigQuery. Select one or more tables. For each selected table, you must designate one or more columns as primary keys before proceeding. You can also optionally select a single column as the event time for the row (supported types are timestamp, date and bigint), and edit names and data types of the individual columns you want to include. `Preview Metadata` and `Preview Data` show the source schema and a sample of rows before you commit to anything.

Select a table in the data source and configure its columns

### Option B: Define a SQL query Instead of selecting a table, you can write a custom SQL query to define the feature group. This is useful when you need to join multiple tables or apply transformations at read time. Click `Fetch Schema` to resolve the columns of the query, then, as with the table option, designate one or more columns as primary keys, optionally pick an event time column and give the feature group a name.

Define a SQL query in the data source and configure its columns

Complete the creation by clicking `Next: Review Configuration` at the bottom of the page. As the last step, you will be able to rename the feature groups and confirm their creation.

Confirm the creation of a new feature group

================================================================================ # Ingest Data with dltHub Source: https://docs.hopsworks.ai/latest/user_guides/fs/feature_group/ingest_with_dlthub/ # How to ingest data into a Feature Group with dltHub { #ingest-data-with-dlthub } ## Introduction Hopsworks can copy data from an existing data source into a new managed feature group using dltHub. This workflow creates: - A new feature group in Hopsworks. - An ingestion job that copies data from the selected source into that feature group. This is different from creating an external feature group. An external feature group keeps the data in the source system, while the dltHub ingestion flow copies the data into Hopsworks. !!! note You can configure this workflow both in the Hopsworks UI and with the Hopsworks Python APIs. ## When to use this workflow Use `Ingest Data to New Feature Group` when you want to: - Copy data from source into Hopsworks. - Schedule recurring ingestion jobs. - Use incremental loading for supported source types. ## Supported source types This ingestion flow supports multiple data sources: - SQL-like sources can either create an external feature group or ingest data into a new feature group. - The SQL family currently includes Snowflake, BigQuery, Redshift, generic JDBC (MySQL, PostgreSQL, Oracle), and SAP HANA. - CRM and REST API sources use the ingestion path only. - Incremental loading is available for SQL and REST API sources. - CRM sources currently use full-load ingestion. ## Step 1: Open the Data Source and start Feature Group creation Navigate to the data source you want to use and start the feature-group creation flow from the UI. For SQL-based sources, open the data source, click `Next: Select Tables`, select a database and a table, then choose `Ingest Data to New Feature Group`. Once the ingest option is selected, the column table gains a `Partition key` column and an `Add a feature` button for features that do not exist in the source; the transformation script computes their values.
![dltHub SQL Feature Group Selection](../../../assets/images/guides/fs/feature_group/dlthub_select_sql_table.png)
Select a source table, set the keys and choose Ingest Data to New Feature Group
For CRM sources, choose the source resource, click `Fetch Schema` and then configure the feature schema for the new feature group the same way. For REST API sources, first configure the endpoint before fetching the schema. ### REST endpoint pagination REST sources require endpoint configuration up front so Hopsworks can fetch the schema correctly. In this step, define: - **Resource**: Any unique identifier for the endpoint. - **Relative URL**: The endpoint path relative to the configured REST data source base URL. - **Request Parameters**: Optional query parameters sent with the request. - **Pagination Configuration**: The pagination mode and its parameters, if the API returns paged results. The REST pagination form supports these modes: - `NONE` - `HEADER_CURSOR` - `HEADER_LINK` - `JSON_CURSOR` - `JSON_LINK` - `OFFSET` - `PAGE_NUMBER` - `SINGLE_PAGE` For example, `PAGE_NUMBER` pagination exposes: - **Page Parameter Name**: Name of the request parameter that contains the page number. - **Base Page**: Starting page number used by the API, for example `0` or `1`. - **Total Pages Path**: Response path containing the total number of pages.
![dltHub REST Page Number Pagination](../../../assets/images/guides/fs/feature_group/dlthub_rest_page_number_pagination.png)
REST API pagination configuration using PAGE_NUMBER
Other pagination modes expose their own source-specific fields in the form: - `OFFSET`: offset parameter name, limit parameter name, limit value, total-items path, and has-more path. - `JSON_CURSOR`: cursor parameter name and cursor path. - `HEADER_CURSOR`: cursor header key and cursor path. - `HEADER_LINK`: next-link header key. - `JSON_LINK`: next URL path. For more details on how these pagination strategies work in dltHub, see the [dltHub REST API pagination documentation](https://dlthub.com/docs/dlt-ecosystem/verified-sources/rest_api/basic#pagination). ## Step 2: Configure the feature group schema After fetching metadata from the source, Hopsworks shows the feature selection table. At this stage you can: - Set the **Feature Group Name**. - Include or exclude columns. - Edit feature names and data types. - Mark one or more features as **Primary key**. - Optionally select a **Partition key**. - Optionally select an **Event time** column. - Preview metadata and preview data before continuing. When you are ready, click `Next: Configure Ingestion Job`. !!! note `Create External Feature Group` is not supported for CRM and REST connectors. !!! note For CRM and REST sources, schema fetching reads only a small sample of records from the source. Hopsworks uses this sample to infer the feature-group schema before you create the ingestion job. ## Step 3: Configure the dltHub ingestion job The next page configures the ingestion job that will populate the feature group.
![dltHub SQL Job Configuration](../../../assets/images/guides/fs/feature_group/dlthub_configure_job_sql.png)
Configure the dltHub ingestion job: job settings, transformation, resources and loading strategy
### Common job settings The following fields are available in the job configuration: - **Job Name**: Name of the ingestion job created in Hopsworks. - **Source Read Parallelism**: Number of parallel readers that pull data from the source database or API. Increase it to speed up ingestion if the source can handle the extra load. - **Data processing parallelism**: Number of parallel processes that prepare and transform data before loading it into the feature group. Increase it if processing is slow and CPU is available. - **Destination Write Batch Size**: Number of records written to the feature group in each batch during ingestion. - **Max Write Batch Size (MB)**: Maximum file size, in megabytes, when writing data to the feature group. - **Write Mode**: Controls whether incoming data is appended as-is or merged with existing rows using the primary key. - **Environment**: Python environment the job runs in, `dlthub-ingestion-pipeline` by default. - **Start the job after creation**: Starts the ingestion job immediately after the resources are created. - **Data Transformation**: Optional Python script, picked from the project or uploaded, that transforms rows before they are written and computes any extra features added to the schema. - **Memory (in MB)** and **CPU Cores**: Resources allocated to the ingestion job; `Estimate resources` proposes values from the source size. - **Schedule**: Optional recurring schedule for future ingestion runs. - **Alerts**: Optional alerting configuration for the ingestion job. ### SQL-only settings For SQL sources, the job configuration also includes: - **Source Read Batch Size**: Number of records fetched per read from the SQL source. - **Source Table Partitions**: Number of partitions used when reading from SQL sources. For very large tables, increase this value to split the read into smaller chunks that fit the allocated memory. These options control how data is read from the source table during ingestion. ### Write modes Two write modes are available: - **APPEND**: Appends new data without merging with existing rows. This greatly speeds up writes and uses less memory, but can result in duplicate rows. If you are ingesting a large amount of data, this is the recommended mode and duplicates can be handled later in a separate pipeline step. - **MERGE**: Merges incoming data with existing rows using the feature-group primary key. This avoids duplicate rows, but slows down ingestion and requires more memory, especially for large ingestions. Use it when ingesting smaller amounts of data. ## Step 4: Choose a loading strategy The `Loading Strategy` section controls whether the pipeline reads the entire source or only new data. The following strategies are available in the UI: - `FULL_LOAD` - `INCREMENTAL_ID` - `INCREMENTAL_TIMESTAMP` - `INCREMENTAL_DATE` --8<-- "user_guides/fs/feature_group/ingest_with_dlthub/loading-strategies.html" ### Full load `FULL_LOAD` is available for all sources in this workflow. With a full load, the ingestion job reads the complete dataset from the source and writes it again to the destination feature group. In practice, this means the target feature group is refreshed from scratch for the same feature-group name and version. Any data already stored in that feature group version is removed and replaced by the newly ingested data from the source. Use `FULL_LOAD` when you want the feature group to be a complete copy of the source at the time of ingestion, rather than an incremental continuation of previous runs. This is useful when: - The source does not provide a reliable incremental cursor. - You want to rebuild the feature group from a clean state. - The source data can change retroactively and you want to re-sync the full table or endpoint. Because a full load rewrites the destination dataset, it is typically more expensive than incremental ingestion for large sources. For recurring pipelines, prefer an incremental strategy when the source supports it and when you only need newly added or updated records. For SQL sources, you can also optionally define: - **Source Cursor Field**: A field used to efficiently synchronize only new or changed data from the source into the destination feature group. - **Initial Value**: Starting value for the selected source cursor field. This can be used to split or optimize the load when the source table has a monotonic column, even though the ingestion mode remains a full refresh of the feature group. ### Incremental loading Incremental loading is available for SQL and REST API sources. With incremental loading, the ingestion job does not re-copy the full source on every run. Instead, it keeps track of a cursor value and only fetches records that are newer than, or come after, the last processed value. This makes incremental loading the preferred option for recurring ingestion jobs when the source exposes a stable field that can be used to identify new or updated data. Typical cursor fields are: - Increasing numeric identifiers. - Update timestamps. - Event dates. Compared to `FULL_LOAD`, incremental loading typically: - Reduces the amount of data read from the source. - Shortens ingestion time. - Lowers resource usage. - Avoids rebuilding the destination feature group from scratch on every run. To work reliably, the selected cursor field should be monotonic or consistently ordered for the records you want to ingest. If the source does not provide such a field, `FULL_LOAD` is usually the safer option. The common incremental field is: - **Source Cursor Field**: A field used to efficiently synchronize only new or changed data from the source into the destination feature group. Depending on the strategy, you must also define: - **INCREMENTAL_ID**: **Initial Value**, the numeric starting value for incremental reads. - **INCREMENTAL_TIMESTAMP**: **Initial Value**, the starting Unix timestamp for incremental reads. - **INCREMENTAL_DATE**: **Initial Date**, the starting date and time for incremental reads. The initial value defines where the first run starts. After that, subsequent runs continue from the last successfully processed cursor value. For REST API sources, incremental loading also requires: - **REST Filter Param**: The actual API parameter used to request only new data since the last run, for example `start_date`, `updated_at`, or `since`. Choose the incremental strategy that matches the source cursor type: - `INCREMENTAL_ID` for sources with increasing numeric identifiers. - `INCREMENTAL_TIMESTAMP` for sources that expose Unix timestamps. - `INCREMENTAL_DATE` for sources that filter by date or datetime values.
![dltHub incremental loading](../../../assets/images/guides/fs/feature_group/dlthub_configure_job_incremental.png)
Incremental loading by id, with tid as the source cursor field
## Step 5: Review and create After configuring the ingestion job, click `Next: Review Configuration`. The review dialog shows: - The source schema, table, connector, or resource. - The final feature group name. - Whether sink ingestion is enabled. - The ingestion job name. - The number of selected features. You can still edit the feature-group name and ingestion-job name in this step before creating the resources.
![dltHub Review Configuration](../../../assets/images/guides/fs/feature_group/dlthub_review_modal.png)
Review the feature group and ingestion job before creation
Click `Create` to create the feature group and the dltHub ingestion job. ## Result After creation: - The feature group is registered in Hopsworks. - The ingestion job is available under project jobs. - If `Start the job after creation` is enabled, the initial ingestion starts immediately. - If a schedule is configured, future synchronizations will run automatically. ## Next Steps - Use the [Feature Group creation guide][create-feature-group] to understand managed feature groups in more detail. - Use the [External Feature Group guide][create-external-feature-group] if you want to query the source in place without copying data into Hopsworks. - Use the [Online Ingestion Observability guide][online-ingestion-observability] to monitor ingestion behavior for online-enabled feature groups. ## API support You can also configure data source ingestion programmatically with the Hopsworks Python APIs. This is done by creating a sink-enabled feature group and passing a sink job configuration, including loading strategy and, for REST sources, endpoint and pagination settings. ### Example: create a sink-enabled feature group ```python from hopsworks_common.core import sink_job_configuration fs = project.get_feature_store() data_source = fs.get_data_source("my_sql_source").get_tables()[0] data = data_source.get_data(use_cached=False) sink_job_conf = sink_job_configuration.SinkJobConfiguration( name="sql_to_fg_ingestion", write_mode=sink_job_configuration.WriteMode.APPEND, ) fg = fs.get_or_create_feature_group( name="transactions_fg", version=1, description="Managed feature group populated from a data source.", primary_key=[data.features[0]["name"]], features=data.features, data_source=data_source, time_travel_format="DELTA", sink_enabled=True, sink_job_conf=sink_job_conf, ) fg.save() # Run the ingestion job fg.sink_job.run(await_termination=True) ``` ### Example: REST ingestion with incremental loading ```python from hopsworks_common.core import rest_endpoint, sink_job_configuration from hsfs.core import data_source as ds fs = project.get_feature_store() parent_data_source = fs.get_data_source("my_rest_source") endpoint_config = rest_endpoint.RestEndpointConfig( relative_url="/transactions", query_params={"page_size": 100}, pagination_config=rest_endpoint.PageNumberPaginationConfig( base_page=1, page_param="page", total_path="total", stop_after_empty_page=True, ), ) rest_data_source = ds.DataSource( table="transactions_rest", rest_endpoint=endpoint_config, storage_connector=parent_data_source.storage_connector, ) rest_data = rest_data_source.get_data(use_cached=False) loading_config = sink_job_configuration.LoadingConfig( loading_strategy=sink_job_configuration.LoadingStrategy.INCREMENTAL_DATE, source_cursor_field="timestamp", initial_value="2024-01-01T00:00:00Z", rest_filter_param="start_time", ) sink_job_conf = sink_job_configuration.SinkJobConfiguration( name="rest_to_fg_ingestion", loading_config=loading_config, ) fg = fs.get_or_create_feature_group( name="transactions_rest_fg", version=1, description="Managed feature group populated from a REST source.", primary_key=["id"], features=rest_data.features, data_source=rest_data_source, time_travel_format="DELTA", sink_enabled=True, sink_job_conf=sink_job_conf, ) fg.save() ``` ================================================================================ # Create Spine Source: https://docs.hopsworks.ai/latest/user_guides/fs/feature_group/create_spine/ # How to create Spine Group ## Introduction In this guide you will learn how to create and register a Spine Group with Hopsworks. ## Prerequisites Before you begin this guide we suggest you read the [Spine Group](../../../concepts/fs/feature_group/spine_group.md) concept page to understand what a Spine Group is and how it fits in the ML pipeline. ## Create using the Hopsworks APIs ### Create a Spine Group Instead of using a feature group to save the label, you can also use a spine to use a Dataframe containing the labels on the fly. A spine is essentially a metadata object similar to a Feature Group, which tells the feature store the relevant event time column and primary key columns to perform point-in-time correct joins. Additionally, apart from primary key and event time information, a Spark dataframe is required in order to infer the schema of the group from. === "Python" ```python trans_spine = fs.get_or_create_spine_group( name="spine_transactions", version=1, description="Transaction data", primary_key=["cc_num"], event_time="datetime", dataframe=trans_df, ) ``` Once created, note that you can inspect the dataframe in the Spine Group: === "Python" ```python trans_spine.dataframe.show() ``` And you can always also replace the dataframe contained within the Spine Group. You just need to make sure it has the same schema. === "Python" ```python trans_spine.dataframe = new_df ``` ### Limitations !!! warning "Python support" Currently the Hopsworks library does not support usage of Spine Groups for training data creation or batch data retrieval in the Python engine. However, it is supported to create Spine Groups from the Python engine. !!! api "API reference" - [`FeatureStore.get_or_create_spine_group`][hsfs.feature_store.FeatureStore.get_or_create_spine_group] - [`SpineGroup`][hsfs.feature_group.SpineGroup] Browse the full Python API :material-arrow-right: ================================================================================ # Deprecate Source: https://docs.hopsworks.ai/latest/user_guides/fs/feature_group/deprecation/ # How to deprecate a Feature Group ## Introduction To discourage the usage of specific feature groups it is possible to deprecate them. When a feature group is deprecated, user will be warned when they try to use it or use a feature view that depends on it. In this guide you will learn how to deprecate a feature group within Hopsworks, showing examples in Hopsworks APIs as well as the user interface. ## Prerequisites Before you begin this guide it is expected that there is an existing feature group in your project. You can familiarize yourself with [the creation of a feature group](./create.md) in the user guide. ## Deprecate using the Hopsworks APIs ### Retrieve the feature group To deprecate a feature group using the Hopsworks APIs you need to provide a [Feature Group](../../../concepts/fs/feature_group/fg_overview.md). === "Python" ```python fg = fs.get_feature_group( name="feature_group_name", version=feature_group_version ) ``` ### Deprecate Feature Group Feature group deprecation occurs by calling the `update_deprecated` method on the feature group. === "Python" ```python fg.update_deprecated() ``` Users can also un-deprecate the feature group if need be, by setting the `deprecate` parameter to False. === "Python" ```python fg.update_deprecated(deprecate=False) ``` ## Deprecate using the UI You can deprecate/de-deprecate feature groups through the UI. For this, navigate to the `Feature Groups` section and select a feature group.

List of Feature Groups

Subsequently, make sure that the necessary feature group version is picked.

Feature Group version selection

Finally, click on the button with three vertical dots in the right corner and select `Deprecate`.

Deprecate Feature Group

The Feature group can be de-deprecated by selecting the `Undeprecate` option on a deprecated feature group. ================================================================================ # Data Types and Schema management Source: https://docs.hopsworks.ai/latest/user_guides/fs/feature_group/data_types/ # How to manage schema and feature data types ## Introduction In this guide, you will learn how to manage the feature group schema and control the data type of the features in a feature group. ## Prerequisites Before you begin this guide we suggest you read the [Feature Group](../../../concepts/fs/feature_group/fg_overview.md) concept page to understand what a feature group is and how it fits in the ML pipeline. We also suggest you familiarize yourself with the APIs to [create a feature group](./create.md). ## Feature group schema When a feature is stored in both the online and offline feature stores, it will be stored in a data type native to each store. - **[Offline data type](#offline-data-types)**: The data type of the feature when stored on the offline feature store. The offline feature store is based on Apache Hudi and Hive Metastore, as such, [Hive Data Types](https://cwiki.apache.org/confluence/display/Hive/LanguageManual+Types) can be leveraged. - **[Online data type](#online-data-types)**: The data type of the feature when stored on the online feature store. The online storage is based on RonDB and hence, [MySQL Data Types](https://dev.mysql.com/doc/refman/8.0/en/data-types.html) can be leveraged. The offline data type is always required, even if the feature group is stored only online. On the other hand, if the feature group is not *online_enabled*, its features will not have an online data type. The offline and online types for each feature are automatically inferred from the Spark or Pandas types of the input DataFrame as outlined in the following two sections. The default mapping, however, can be overwritten by using an [explicit schema definition](#explicit-schema-definition). ### Offline data types When registering a [Spark](https://spark.apache.org/docs/latest/sql-ref-datatypes.html) DataFrame in a PySpark environment (S), or a [Pandas](https://pandas.pydata.org/) DataFrame, or a [Polars](https://pola.rs/) DataFrame in a Python-only environment (P) the following default mapping to offline feature types applies: | Spark Type (S) | Pandas Type (P) |Polars Type (P) | Offline Feature Type | Remarks | |----------------|------------------------------------|-----------------------------------|-------------------------------|----------------------------------------------------------------| | BooleanType | bool, object(bool) |Boolean | BOOLEAN | | | ByteType | int8, Int8 |Int8 | TINYINT or INT | INT when time_travel_type="HUDI" | | ShortType | uint8, int16, Int16 |UInt8, Int16 | SMALLINT or INT | INT when time_travel_type="HUDI" | | IntegerType | uint16, int32, Int32 |UInt16, Int32 | INT | | | LongType | int, uint32, int64, Int64 |UInt32, Int64 | BIGINT | | | FloatType | float, float16, float32 |Float32 | FLOAT | | | DoubleType | float64 |Float64 | DOUBLE | | | DecimalType | decimal.decimal |Decimal | DECIMAL(PREC, SCALE) | Not supported in PO env. when time_travel_type="HUDI" | | TimestampType | datetime64[ns], datetime64[ns, tz] |Datetime | TIMESTAMP | s. [Timestamps and Timezones](#timestamps-and-timezones) | | DateType | object (datetime.date) |Date | DATE | | | StringType | object (str), object(np.unicode) |String, Utf8 | STRING | | | ArrayType | object (list), object (np.ndarray) |List | ARRAY<TYPE> | | | StructType | object (dict) |Struct | STRUCT<NAME: TYPE, ...> | | | BinaryType | object (binary) |Binary | BINARY | | | MapType | - |- | MAP<String,TYPE> | Only when time_travel_type!="HUDI"; Only string keys permitted | When registering a Pandas DataFrame in a PySpark environment (S) the Pandas DataFrame is first converted to a Spark DataFrame, using Spark's [default conversion](https://spark.apache.org/docs/3.1.1/api/python/reference/api/pyspark.sql.SparkSession.createDataFrame.html). It results in a less fine-grained mapping between Python and Spark types: | Pandas Type (S) | Spark Type | Remarks | |-------------------------------------------------------|---------------|----------------------------------------------------------| | bool | BooleanType | | | int8, uint8, int16, uint16, int32, int, uint32, int64 | LongType | | | float, float16, float32, float64 | DoubleType | | | object (decimal.decimal) | DecimalType | | | datetime64[ns], datetime64[ns, tz] | TimestampType | s. [Timestamps and Timezones](#timestamps-and-timezones) | | object (datetime.date) | DateType | | | object (str), object(np.unicode) | StringType | | | object (list), object (np.ndarray) | - | Not supported | | object (dict) | StructType | | | object (binary) | BinaryType | | ### Online data types The online data type is determined based on the offline type according to the following mapping, regardless of which environment the data originated from. Only a subset of the data types can be used as primary key, as indicated in the table as well: | Offline Feature Type | Online Feature Type | Primary Key | Remarks | |-------------------------------|----------------------|-------------|----------------------------------------------------------| | BOOLEAN | TINYINT | x | | | TINYINT | TINYINT | x | | | SMALLINT | SMALLINT | x | | | INT | INT | x | Also supports: TINYINT, SMALLINT | | BIGINT | BIGINT | x | | | FLOAT | FLOAT | | | | DOUBLE | DOUBLE | | | | DECIMAL(PREC, SCALE) | DECIMAL(PREC, SCALE) | | e.g. DECIMAL(38, 18) | | TIMESTAMP | TIMESTAMP | | s. [Timestamps and Timezones](#timestamps-and-timezones) | | DATE | DATE | x | | | STRING | VARCHAR(100) | x | Also supports: TEXT | | ARRAY<TYPE> | VARBINARY(100) | x | Also supports: BLOB | | STRUCT<NAME: TYPE, ...> | VARBINARY(100) | x | Also supports: BLOB | | BINARY | VARBINARY(100) | x | Also supports: BLOB | | MAP<String,TYPE> | VARBINARY(100) | x | Also supports: BLOB | More on how Hopsworks handles [string types](#string-online-data-types), [complex data types](#complex-online-data-types) and the online restrictions for [primary keys](#online-restrictions-for-primary-key-data-types) and [row size](#online-restrictions-for-row-size) in the following sections. #### String online data types String types are stored as *VARCHAR(100)* by default. This type is fixed-size, meaning it can only hold as many characters as specified in the argument (e.g., VARCHAR(100) can hold up to 100 unicode characters). The size should thus be within the maximum string length of the input data. Furthermore, the VARCHAR size has to be in line with the [online restrictions for row size](#online-restrictions-for-row-size). If the string size exceeds 100 characters, a larger type (e.g., VARCHAR(500)) can be specified via an [explicit schema definition](#explicit-schema-definition). If the string size is unknown or if it exceeds the maximum row size, then the [TEXT type](https://docs.rondb.com/blobs/) can be used instead. String data that exceeds the specified VARCHAR size will lead to an error when data gets written to the online feature store. When in doubt, use the TEXT type instead, but note that it comes with a potential performance overhead. #### Complex online data types Hopsworks allows users to store complex types (e.g. *ARRAY*) in the online feature store. Hopsworks serializes the complex features transparently and stores them as VARBINARY in the online feature store. The serialization happens when calling the [`FeatureGroup.save`][hsfs.feature_group.FeatureGroup.save], [`FeatureGroup.insert`][hsfs.feature_group.FeatureGroup.insert] or [`FeatureGroup.insert_stream`][hsfs.feature_group.FeatureGroup.insert_stream] methods. The deserialization will be executed when calling the [`TrainingDataset.get_serving_vector`][hsfs.training_dataset.TrainingDataset.get_serving_vector] method to retrieve data from the online feature store. If users query directly the online feature store, for instance using the `fs.sql("SELECT ...", online=True)` statement, it will return a binary blob. On the feature store UI, the online feature type for complex features will be reported as *VARBINARY*. If the binary size exceeds 100 bytes, a larger type (e.g., VARBINARY(500)) can be specified via an [explicit schema definition](#explicit-schema-definition). If the binary size is unknown of if it exceeds the maximum row size, then the [BLOB type](https://docs.rondb.com/blobs/) can be used instead. Binary data that exceeds the specified VARBINARY size will lead to an error when data gets written to the online feature store. When in doubt, use the BLOB type instead, but note that it comes with a potential performance overhead. #### Online restrictions for primary key data types When a feature is being used as a primary key, certain types are not allowed. Examples of such types are *FLOAT*, *DOUBLE*, *TEXT* and *BLOB*. Additionally, the size of the sum of the primary key online data types storage requirements **should not exceed 4KB**. #### Online restrictions for row size The online feature store supports **up to 500 columns** and all column types combined **should not exceed 30000 Bytes**. The byte size of each column is determined by its data type and calculated as follows: | Online Data Type | Byte Size | |---------------------------------|--------------| | TINYINT | 1 | | SMALLINT | 2 | | INT | 4 | | BIGINT | 8 | | FLOAT | 4 | | DOUBLE | 8 | | DECIMAL(PREC, SCALE) | 16 | | TIMESTAMP | 8 | | DATE | 8 | | VARCHAR(LENGTH) | LENGTH * 4 | | VARCHAR(LENGTH) charset latin1; | LENGTH * 1 | | TEXT | 256 | | VARBINARY(LENGTH) | LENGTH | | BLOB | 256 | | other | 8 | !!! note "VARCHAR / VARBINARY overhead" For VARCHAR and VARBINARY data types, an additional 1 byte is required if the size is less than 256 bytes. If the size is 256 bytes or greater, 2 additional bytes are required. Memory allocation is performed in groups of 4 bytes. For example, a VARBINARY(100) requires 104 bytes of memory: - 100 bytes for the data itself - 1 byte of overhead - Total = 101 bytes Since memory is allocated in 4-byte groups, storing 101 bytes requires 26 groups (26 × 4 = 104 bytes) of allocated memory. #### Pre-insert schema validation for online feature groups For online enabled feature groups, the dataframe to be ingested needs to adhere to the online schema definitions. The input dataframe is validated for schema checks accordingly. The validation is enabled by default and can be disabled by setting below key word argument when calling `insert()` === "Python" ```python feature_group.insert( df, validation_options={"online_schema_validation": False} ) ``` The most important validation checks or error messages are mentioned below along with possible corrective actions. 1. Primary key contains null values - **Rule** Primary key column should not contain any null values. - **Example correction** Drop the rows containing null primary keys. Alternatively, find the null values and assign them an unique value as per preferred strategy for data imputation. ```python # Drop rows: assuming 'id' is the primary key column df = df.dropna(subset=["id"]) # For composite keys df = df.dropna(subset=["id1", "id2"]) # Data imputation: replace null values with incrementing last integer id # existing max id max_id = df["id"].max() # counter to generate new id next_id = max_id + 1 # for each null id, assign the next id incrementally for idx in df[df["id"].isna()].index: df.loc[idx, "id"] = next_id next_id += 1 ``` 2. Primary key column missing - **Rule** The dataframe to be inserted must contain all the columns defined as primary key(s) in the feature group. - **Example correction** Add all the primary key columns in the dataframe. ```python # incrementing primary key upto the length of dataframe df["id"] = range(1, len(df) + 1) ``` 3. String length exceeded - **Rule** The character length of a string should be within the maximum length capacity in the online schema type of a feature. If the feature group is not created and explicit feature schema was not provided, the limit will be auto-increased to the maximum length found in a string column in the dataframe. - **Example correction** - Trim the string values to fit within maximum limit set during feature group creation. ```python max_length = 100 df["text_column"] = df["text_column"].str.slice(0, max_length) ``` - Another option is to simply [create new version of the feature group][hsfs.feature_store.FeatureStore.get_or_create_feature_group] and insert the dataframe. !!! note The total row size limit should be less than 30kb as per [row size restrictions](#online-restrictions-for-row-size). In such cases it is possible to define the feature as **TEXT** or **BLOB**. Below is an example of explicitly defining the string column as TEXT as online type. ```python import pandas as pd # example dummy dataframe with the string column df = pd.DataFrame(columns=["id", "string_col"]) from hsfs.feature import Feature features = [ Feature(name="id", type="bigint", online_type="bigint"), Feature(name="string_col", type="string", online_type="text"), ] fg = fs.get_or_create_feature_group( name="fg_manual_text_schema", version=1, features=features, online_enabled=True, primary_key=["id"], ) fg.insert(df) ``` ### Timestamps and Timezones All timestamp features are stored in Hopsworks in UTC time. Also, all timestamp-based functions (such as [point-in-time joins](../../../concepts/fs/feature_view/offline_api.md#point-in-time-correct-training-data)) use UTC time. This ensures consistency of timestamp features across different client timezones and simplifies working with timestamp-based functions in general. When ingesting timestamp features, the [`FeatureGroup.insert`][hsfs.feature_group.FeatureGroup.insert] will automatically handle the conversion to UTC, if necessary. The following table summarizes how different timestamp types are handled: | Data Frame (Data Type) | Environment | Handling | | --- | --- | --- | | Pandas DataFrame (datetime64[ns]) | Python-only and PySpark | interpreted as UTC, independent of the client's timezone | | Pandas DataFrame (datetime64[ns, tz]) | Python-only and PySpark | timezone-sensitive conversion from 'tz' to UTC | | Spark (TimestampType) | PySpark and Spark | interpreted as UTC, independent of the client's timezone | Timestamp features retrieved from the Feature Store, e.g., using the [Feature Store Read API][hsfs.feature_group.FeatureGroup.read], use a timezone-unaware format: | Data Frame (Data Type) | Environment | Timezone | |---------------------------------------|-------------------------|------------------------| | Pandas DataFrame (datetime64[ns]) | Python-only | timezone-unaware (UTC) | | Spark (TimestampType) | PySpark and Spark | timezone-unaware (UTC) | Note that our PySpark/Spark client automatically sets the Spark SQL session's timezone to UTC. This ensures that Spark SQL will correctly interpret all timestamps as UTC. The setting will only apply to the client's session, and you don't have to worry about setting/unsetting the configuration yourself. ## Explicit schema definition When creating a feature group it is possible for the user to control both the offline and online data type of each column. If users explicitly define the schema for the feature group, Hopsworks is going to use that schema to create the feature group, without performing any type mapping. You can explicitly define the feature group schema as follows: === "Python" ```python from hsfs.feature import Feature features = [ Feature(name="id", type="int", online_type="int"), Feature(name="name", type="string", online_type="varchar(20)"), ] fg = fs.create_feature_group( name="fg_manual_schema", features=features, online_enabled=True ) fg.save(features) ``` ## Append features to existing feature groups Hopsworks supports appending additional features to an existing feature group. Adding additional features to an existing feature group is not considered a breaking change. === "Python" ```python from hsfs.feature import Feature features = [ Feature(name="id", type="int", online_type="int"), Feature(name="name", type="string", online_type="varchar(20)"), ] fg = fs.get_feature_group(name="example", version=1) fg.append_features(features) ``` When adding additional features to a feature group, you can provide a default values for existing entries in the feature group. You can also backfill the new features for existing entries by running an `insert()` operation and update all existing combinations of *primary key* - *event time*. ### Appending features to an external feature group For an [external feature group](create_external.md), appending a feature only updates the Hopsworks-side metadata; it does not add a column to the external table itself. Reading online (`read(online=True)`) is unaffected, since Hopsworks owns the online table schema and can add the column there directly. Reading offline (`read()`) queries the external source directly, so it fails with a column-not-found error until the external table itself gains a matching column. Until then, exclude the appended feature from the offline read with `select_except`: === "Python" ```python fg = fs.get_feature_group(name="example", version=1) # "name" was appended but is not yet a column in the external table df = fg.select_except(["name"]).read() ``` Update the external table's schema, or keep excluding the appended feature, then read offline again once the column is present. ================================================================================ # Statistics Source: https://docs.hopsworks.ai/latest/user_guides/fs/feature_group/statistics/ # How to compute statistics on feature data ## Introduction In this guide you will learn how to configure, compute and visualize statistics for the features registered with Hopsworks. Hopsworks groups statistics in four categories: - **Descriptive**: These are the basic statistics Hopsworks computes. They include an _approximate_ count of the distinctive values and the completeness (i.e., the percentage of non null values). For numerical features Hopsworks also computes the minimum, maximum, mean, standard deviation and the sum of each feature. Enabled by default. - **Histograms**: Hopsworks computes the distribution of the values of a feature. Exact histograms are computed as long as the number of distinct values is less than 20. If a feature has a numerical data type (e.g., integer, float, double, ...) and has more than 20 unique values, then the values are bucketed in 20 buckets and the histogram represents the distribution of values in those buckets. By default histograms are disabled. - **Correlation**: If enabled, Hopsworks computes the Pearson correlation between features of numerical data type within a feature group. By default correlation is disabled. - **Exact Statistics**: Exact statistics are an enhancement of the descriptive statistics that provide an exact count of distinctive values, entropy, uniqueness and distinctiveness of the value of a feature. These statistics are more expensive to compute as they take into consideration all the values and they don't use approximations. By default they are disabled. When statistics are enabled, they are computed every time new data is written into the _offline_ storage of a feature group. Statistics are then displayed on the Hopsworks UI and users can track how data has changed over time. ## Prerequisites Before you begin this guide we suggest you read the [Feature Group](../../../concepts/fs/feature_group/fg_overview.md) concept page to understand what a feature group is and how it fits in the ML pipeline. We also suggest you familiarize with the APIs to [create a feature group](./create.md). ## Enable statistics when creating a feature group As mentioned above, by default only descriptive statistics are enabled when creating a feature group. To enable histograms, correlations or exact statistics the `statistics_config` configuration parameter can be provided in the create statement. The `statistics_config` parameter takes a dictionary with the keys: `enabled`, `correlations`, `histograms` and `exact_uniqueness` and, as values, a boolean to describe whether or not to compute the specific class of statistics. Additionally it is possible to restrict the statistics computation to only a subset of columns. This is configurable by adding a `columns` key to the `statistics_config` parameter. The key should contain the list of columns for which to compute statistics. By default the value is empty list `[]` and the statistics are computed for all columns in the feature group. === "Python" ```python fg = feature_store.create_feature_group( name="weather", version=1, description="Weather Features", online_enabled=True, primary_key=["location_id"], partition_key=["day"], event_time="event_time", statistics_config={ "enabled": True, "histograms": True, "correlations": True, "exact_uniqueness": False, "columns": [], }, ) ``` ## Enable statistics after creating a feature group You can change the statistics configuration after a feature group was created, to add or remove a class of statistics or to change the set of features for which to compute them. In the UI, open the feature group in the `Catalog` and click the edit icon; the statistics configuration sits at the top of the edit page.
Edit Feature Group page with the statistics configuration checkboxes
Statistics configuration on the Edit Feature Group page.
=== "Python" ```python fg.statistics_config = { "enabled": True, "histograms": False, "correlations": False, "exact_uniqueness": False, "columns": ["location_id", "min_temp", "max_temp"], } fg.update_statistics_config() ``` ## Explicitly compute statistics As mentioned above, the statistics are computed every time new data is written into the _offline_ storage of a feature group. By invoking the `compute_statistics` method, users can trigger explicitly the statistics computation for the data available in a feature group. This is useful when a feature group is receiving frequent updates. Users can schedule periodic statistics computation that take into consideration several data commits. By default, the `compute_statistics` method computes statistics on the most recent version of the data available in a feature group. Users can provide a specific time using the `wallclock_time` parameter, to compute the statistics for a previous version of the data. === "Python" ```python fg.compute_statistics(wallclock_time="20220611 20:00") ``` ### External feature groups External feature groups own the same built-in `ingestion_stats` configuration as cached and stream feature groups, but no data is ingested into Hopsworks for them, so it never runs on its own. Calling `compute_statistics` on the external feature group, or clicking "Compute statistics" in the UI, runs it: the statistics job reads the external source and profiles it like an internal feature group. Saving an external feature group with statistics enabled runs it once as well. External feature groups have no commit history, so their statistics carry the computation time only and `compute_statistics` takes no time argument. === "Python" ```python external_fg.compute_statistics() ``` ## Inspect statistics Open the feature group in the `Catalog` and select `Feature Statistics` in its sidebar. The page lists every feature with its count, completeness, min, max, mean and standard deviation, plus a histogram per numerical feature, for the latest commit or any earlier one you pick. The same numbers are available from the API with `fg.get_statistics()`.
Feature Statistics page listing every feature with its descriptive statistics
Feature Statistics for a feature group, one card per feature.
================================================================================ # Getting started Source: https://docs.hopsworks.ai/latest/user_guides/fs/feature_group/data_validation/ # Data Validation --8<-- "user_guides/fs/feature_group/data_validation/validation-on-insert.html" ## Introduction Clean, high quality feature data is of paramount importance to being able to train and serve high quality models. Hopsworks offers integration with [Great Expectations](https://greatexpectations.io/) to enable a smooth data validation workflow. This guide is designed to help you integrate a data validation step when inserting new DataFrames into a Feature Group. Note that validation is performed inline as part of your feature pipeline (on the client machine) - it is not executed by Hopsworks after writing features. ## UI ### Create a Feature Group (Pre-requisite) In the UI, you must create a Feature Group first before attaching an Expectation Suite. You can find out more information about [creating a Feature Group](create.md). You can attach at most one expectation suite to a Feature Group. Data validation is an optional step and is not required to write to a Feature Group. ### Step 1: Find and Edit Feature Group Click on the Feature Group section in the navigation menu. Find your Feature Group in the list and click on its name to access the Feature Group page. Select `edit` in the top right corner or scroll to the Expectations section and click on `Edit Expectation Suite`. ### Step 2: Edit General Expectation Suite Settings Scroll to the Expectation Suite section. Click add Expectation Suite and edit its metadata: - Choose a name for your expectation suite. - Checkbox enabled. This controls whether the Expectation Suite will be used to validate a Dataframe automatically upon insertion into a Feature Group. Note that validation is executed by the client. Disabling validation allows you to skip the validation step without deleting the Expectation Suite. - 'ALWAYS' vs. 'STRICT' mode. This option controls what happens after validation. Hopsworks defaults to 'ALWAYS', where data is written to the Feature Group regardless of the validation result. This means that even if expectations are failing or throw an exception, Hopsworks will attempt to insert the data into the Feature Group. In 'STRICT' mode, Hopsworks will only write data to the Feature Group if each individual expectation has been successful. ### Step 3: Add new expectations By clicking on `Add expectation` one can choose an expectation type from a searchable dropdown menu. Currently, only the built-in expectations from the Great Expectations framework are supported. For user-defined expectations, please use the Rest API or python client. All default kwargs associated to the selected expectation type are populated as a json below the dropdown menu. Edit the arguments in the json to configure the Expectation. In particular, arguments such as `column`, `columnA`, `columnB`, `column_set` and `column_list` require valid feature name(s). Click the tick button to save the expectation configuration and append it to the Expectation Suite locally. !!! info Click the `Save feature group` button to persist your changes! You can use the button `Clear Expectation Suite` to clean up before saving changes if you changed your mind. If the Expectation Suite is already registered, it will instead show a button to delete the Expectation Suite.
Expectation Suite section of the Edit Feature Group page with three expectations and the STRICT policy selected
The Expectation Suite editor: name, enabled flag, ingestion policy, and one row per expectation.
### Step 4: Save new data to a Feature Group Use the python client to write a DataFrame to the Feature Group. Note that if an expectation suite is enabled for a Feature Group, calling the `insert` method will run validation and default to uploading the corresponding validation report to Hopsworks. The report is uploaded even if validation fails and 'STRICT' mode is selected. ### Step 5: Check Validation Results Summary Hopsworks shows a visual summary of validation reports. To check it out, go to your Feature Group overview and scroll to the expectation section. Click on the `Validation Results` tab and check that all went according to plan. Each row corresponds to an expectation in the suite. Features can have several corresponding expectations and the same type of expectation can be applied to different features. You can navigate to older reports using the dropdown menu. Should you need more than the information displayed in the UI for e.g., debugging, the full report can be downloaded by clicking on the corresponding button. ### Step 6: Check Validation History The `Validation Reports` tab in the Expectations section displays a brief history of recent validations. Each row corresponds to a validation report, with some summary information about the success of the validation step. You can download the full report by clicking the download icon button that appears at the end of the row.
Expectations section of a feature group showing the validation reports table
The Expectations section on the feature group page, with the validation reports history.
## Code Hopsworks python client interfaces with the Great Expectations library to enable you to add data validation to your feature engineering pipeline. In this section, we show you how in a single line you enable automatic validation on each insertion of new data into your Feature Group. Whether you have an existing Feature Group you want to add validation to or Follow the guide or get your hands dirty by running our [tutorial data validation notebook](https://colab.research.google.com/github/logicalclocks/hopsworks-tutorials/blob/master/integrations/great_expectations/fraud_batch_data_validation.ipynb) in google colab. First checkout the pre-requisite and Hopsworks setup to follow the guide below. Create a project, install the hopsworks client and connect via the generated API key. You are ready to load your data in a DataFrame. The second step is a short introduction to the relevant Great Expectations API to build data validation suited to your data. Third and final step shows how to attach your Expectation Suite to the Feature Group to benefit from automatic validation on insertion capabilities. ### Step 1: Pre-requisite In order to define and validate an expectation when writing to a Feature Group, you will need: - A Hopsworks project. If you don't have a project yet you can go to [run.hopsworks.ai](https://run.hopsworks.ai), signup with your email and create your first project. - An API key, you can get one by going to "Account Settings" on [run.hopsworks.ai](https://run.hopsworks.ai). - The [Hopsworks Python library](https://pypi.org/project/hopsworks) installed in your client. See the [installation guide](../../client_installation/index.md). #### Connect your notebook to Hopsworks Connect the client running your notebooks to Hopsworks. ```python import hopsworks project = hopsworks.login() fs = project.get_feature_store() ``` You will be prompt to paste your API key to connect the notebook to your project. The `fs` Feature Store entity is now ready to be used to insert or read data from Hopsworks. #### Import your data Load your data in a DataFrame using the usual pandas API. ```python import pandas as pd df = pd.read_csv( "https://repo.hops.works/master/hopsworks-tutorials/data/card_fraud_data/transactions.csv", parse_dates=["datetime"], ) df.head(3) ``` ### Step 2: Great Expectation Introduction To validate the data, we will use the [Great Expectations](https://greatexpectations.io/) library. Below is a short introduction on how to build an Expectation Suite to validate your data. Everything is done using the Great Expectations API so you can re-use any prior knowledge you may have of the library. The Hopsworks `great-expectations` extra supports Great Expectations 0.18.12 and 1.17.1. We recommend 1.17.1, and examples in this guide target that version. #### Create an Expectation Suite Create (or import an existing) expectation suite using the Great Expectations library. This suite will hold all the validation tests we want to perform on our data before inserting them into Hopsworks. ```python import great_expectations as gx expectation_suite = gx.ExpectationSuite(name="validate_on_insert_suite") ``` #### Add Expectations in the Source Code Add some expectations to your suite. Each expectation configuration corresponds to a validation test to be run against your data. ```python from great_expectations.expectations.expectation_configuration import ( ExpectationConfiguration, ) expectation_suite.add_expectation_configuration( ExpectationConfiguration( type="expect_column_min_to_be_between", kwargs={"column": "foo_id", "min_value": 0, "max_value": 1}, ) ) expectation_suite.add_expectation_configuration( ExpectationConfiguration( type="expect_column_value_lengths_to_be_between", kwargs={"column": "bar_name", "min_value": 3, "max_value": 10}, ) ) ``` !!! info "Migrating from Great Expectations 0.18.x" The constructor argument was renamed from `expectation_suite_name=` to `name=` in 1.0. `ExpectationConfiguration` now takes `type=` instead of `expectation_type=` and was moved out of `great_expectations.core` to `great_expectations.expectations.expectation_configuration`. The Hopsworks SDK normalizes both shapes on the wire, so suites stored under either version remain readable. #### Build a Suite with Typed Expectation Classes Great Expectations 1.x also exposes a typed class for each expectation, which gives you IDE autocomplete on the kwargs. You can mix typed instances and `ExpectationConfiguration` instances in the same suite. ```python import great_expectations.expectations as gxe typed_suite = gx.ExpectationSuite( name="validate_on_insert_suite", expectations=[ gxe.ExpectColumnMinToBeBetween( column="foo_id", min_value=0, max_value=1 ), gxe.ExpectColumnValueLengthsToBeBetween( column="bar_name", min_value=3, max_value=10 ), ], ) ``` Once you have built an Expectation Suite you are satisfied with, it is time to create your first validation-enabled Feature Group. ### Step 3: Attach an Expectation Suite to your Feature Group to enable Automatic Validation on Insertion Writing data in Hopsworks is done using Feature Groups. Once a Feature Group is registered in the Feature Store, you can use it to insert your pandas DataFrames. For more information see [create Feature Group](create.md). To benefit from automatic validation on insertion, attach your newly created Expectation Suite when creating the Feature Group: ```python fg = fs.create_feature_group( "fg_with_data_validation", version=1, description="Validated data", primary_key=["foo_id"], online_enabled=False, expectation_suite=expectation_suite, ) ``` or, if the Feature Group already exist, you can simply run: ```python fg.save_expectation_suite(expectation_suite) ``` That is all there is to it. Hopsworks will now automatically use your suite to validate the DataFrames you want to write to the Feature Group. Try it out! ```python job, validation_report = fg.insert(df.head(5)) ``` As you can see, Hopsworks runs the validation in the client before attempting to insert the data. By default, Hopsworks will try to insert the data even if validation fails to prevent data loss. However it can be configured for production setup to be more restrictive, checkout the [data validation advanced guide](data_validation_advanced.md). !!!info Note that once the Expectation Suite is attached to the Feature Group, any subsequent attempt to insert to this Feature Group will apply the Data Validation step even from a different client or in a scheduled job. ### Step 4: Data Quality Monitoring Upon running validation, Great Expectations generates a report to help you assess the quality of your data. Nothing to do here, Hopsworks client automatically uploads the validation report to the backend when ingesting new data. It enables you to monitor the quality of the inserted data in the Feature Group over time. You can checkout a summary of the reports in the UI on your Feature Group page. As you can see, your Feature Group conveniently gather all in one place: your data, the Expectation Suite and the reports generated each time you inserted data! Hopsworks client API allows you to retrieve validation reports for further analysis. ```python # load multiple reports validation_reports = fg.get_all_validation_reports() # convenience method for rapid development ge_latest_report = fg.get_latest_validation_report() ``` Similarly you can retrieve the historic of validation results for a particular expectation, e.g to plot a time-series of a given expectation observed value over time. ```python validation_history = fg.get_validation_history(expectation_id=1) ``` You can find the expectation IDs in the UI or using `fg.get_expectation_suite()` and looking them up in the expectation's `meta` field under the `expectationId` key. !!! info If Validation Reports or Results are too long, they can be truncated to fit in the database. A full version of the reports can be downloaded from the UI. ## Conclusion The integration between Hopsworks and Great Expectations makes it simple to add a data validation step to your feature engineering pipeline. Build your Expectation Suite and attach it to your Feature Group with a single line of code. No need to add any code to your pipeline or job scripts, calling `fg.insert` will now automatically validate the data before inserting them in the Feature Group. The validation reports are stored along your data in Hopsworks allowing us to provide basic monitoring capabilities to quickly spot a data quality issue in the UI. ## Going Further If you wish to find out more about how to use the data validation API or best practices for development or production pipelines in Hopsworks, checkout the [advanced guide](data_validation_advanced.md) and [best practices guide](data_validation_best_practices.md). ================================================================================ # Advanced guide Source: https://docs.hopsworks.ai/latest/user_guides/fs/feature_group/data_validation_advanced/ # Advanced Data Validation Options and Best Practices The introduction to the data validation guide can be found in the [Data Validation Guide](data_validation.md). The notebook example to get started with Data Validation in Hopsworks can be found in the [Fraud Batch Data Validation Tutorial](https://colab.research.google.com/github/logicalclocks/hopsworks-tutorials/blob/master/integrations/great_expectations/fraud_batch_data_validation.ipynb). ## Data Validation Configuration Options in Hopsworks ### Validation Ingestion Policy Depending on your use case you can setup data validation as a monitoring or gatekeeping tool when trying to insert new data in your Feature Group. Switch behaviour by using the `validation_ingestion_policy` kwarg: - `"ALWAYS"` is the default option and will attempt to insert the data regardless of the validation result. Hassle free, it is ideal to monitor data ingestion in a development setup. - `"STRICT"` is the best option for production ready projects. This will prevent insertion of DataFrames which do not pass all data quality requirements. Ideal to avoid "garbage-in, garbage-out" scenarios, at the price of a potential loss of data. Check out the best practice section for more on that. #### Validation Ingestion Policy in UI Go to the Feature Group edit page, in the Expectation section you can choose between the options above. #### Validation Ingestion Policy in Python ```python fg.expectation_suite.validation_ingestion_policy = "ALWAYS" # "STRICT" ``` If your suite is registered with Hopsworks, it will persist the change to the server. ### Disable Data Validation Should you wish to do so, you can disable data validation on a punctual basis or until further notice. #### Disable Data Validation in UI You can do it in the UI in the Expectation section of the Feature Group edit page. Simply tick or untick the enabled checkbox. This will be used as the default option but can be overridden via the API. #### Disable Data Validation in Python To disable data validation until further notice in the API, you can update the `run_validation` field of the expectation suite. If your suite is registered with Hopsworks, this will persist the change to the server. ```python fg.expectation_suite.run_validation = False ``` If you wish to override the default behaviour of the suite when inserting data in the Feature Group, you can do so via the `validation_options` kwarg. The example below will enable validation for this insertion only. ```python fg.insert(df_to_validate, validation_options={"run_validation": True}) ``` We recommend to avoid using this option in scheduled job as it silently changes the expected behaviour that is displayed in the UI and prevents changes to the default behaviour to change the behaviour of the job. ### Edit Expectations The one constant in life is change. If you need to add, remove or edit an expectation you can do it both in the UI or via the python client. Note that changing the expectation type or its corresponding feature will throw an error in order to preserve a meaningful validation history. #### Edit Expectations in UI Go to the Feature Group edit page, in the expectation section. You can click on the expectation you want to edit and edit the json configuration. Check out Great Expectations documentation if you need more information on a particular expectation. #### Edit Expectations in Python There are several way to edit an Expectation in the python client. You can use Great Expectations API or directly go through Hopsworks. In the latter case, if you want to edit or remove an expectation, you will need the Hopsworks expectation ID. It can be found in the UI or in the meta field of an expectation. Note that you must have inserted data in the FG and attached the expectation suite to enable the Expectation API. Get an expectation with a given id: ```python my_expectation = fg.expectation_suite.get_expectation( expectation_id=my_expectation_id ) ``` Add a new expectation: ```python from great_expectations.expectations.expectation_configuration import ( ExpectationConfiguration, ) new_expectation = ExpectationConfiguration( type="expect_column_values_to_not_be_null", kwargs={"column": "foo_id", "mostly": 1}, ) fg.expectation_suite.add_expectation(new_expectation) ``` The single-expectation API on `fg.expectation_suite` accepts `ExpectationConfiguration` instances and plain dicts. On 0.18.x the same API still works with `ge.core.ExpectationConfiguration(expectation_type=...)`. Edit expectation kwargs of an existing expectation : ```python existing_expectation = fg.expectation_suite.get_expectation( expectation_id=existing_expectation_id ) existing_expectation.kwargs["mostly"] = 0.95 fg.expectation_suite.replace_expectation(existing_expectation) ``` Remove an expectation: ```python fg.expectation_suite.remove_expectation( expectation_id=id_of_expectation_to_delete ) ``` If you want to deal only with the Great Expectations API: ```python my_suite = fg.get_expectation_suite() my_suite.add_expectation_configuration(new_expectation) fg.save_expectation_suite(my_suite) ``` `add_expectation_configuration` is the right method to call when you only have the suite in hand (no `DataContext`). On 0.18.x the equivalent method was `my_suite.add_expectation(new_expectation)`; in 1.x `add_expectation` instead requires an active `DataContext` and is meant for context-managed suites. ### Save Validation Reports When running validation using Great Expectations, a validation report is generated containing all validation results for the different expectations. Each result provides information about whether the provided DataFrame conforms to the corresponding expectation. These reports can be stored in Hopsworks to save a validation history for the data written to a particular Feature Group. The boilerplate of uploading report on insertion is taken care of by hopsworks, however for custom pipelines we provide an alternative method in the python client. The UI does not currently support upload of a validation report. #### Save Validation Reports in Python ```python fg.save_validation_report(ge_report) ``` ### Monitor and Fetch Validation Reports A summary of uploaded reports will then be available via an API call or in the Hopsworks UI enabling easy monitoring. For in-depth analysis, it is possible to download the complete report from the UI. #### Monitor and Fetch Validation Reports in UI Open the Feature Group overview page and go to the Expectations section. One tab allows you to check the report history with general information, while the other tab allows you to explore a summary of the result for individual expectations. #### Monitor and Fetch Validation Reports in Python ```python # convenience method for rapid development ge_latest_report = fg.get_latest_validation_report() # fetching the latest summary prints a link to the UI # where you can download full report if summary is insufficient # or load multiple reports validation_history = fg.get_all_validation_reports() ``` ### Validate Your Data Manually While Hopsworks provides automatic validation on insertion logic, we recognise that some use cases may require a more fine-grained control over the validation process. Therefore, Feature Group objects offers a convenience wrapper around Great Expectations to manually trigger validation using the registered Expectation Suite. #### Validate Your Data Manually in UI You can validate data already ingested in the Feature Group by going to the Feature Group overview page. In the top right corner is a button to trigger a validation. The button will launch a job which will read the Feature Group data, run validation and persist the associated report. #### Validate Your Data Manually in Python ```python ge_report = fg.validate(df, ingestion_result="EXPERIMENT") # set the save_report parameter to False to skip uploading the report to Hopsworks # ge_report = fg.validate(df, save_report=False) ``` If you want to apply validation to the data already in the Feature Group you can call the `.validate` without providing data. It will read the data in the Feature Group. ```python report = fg.validate() ``` As validation objects returned by Hopsworks are native Great Expectations objects you can also run validation directly through the Great Expectations API. Great Expectations 1.x removed `ge.from_pandas` and replaced it with the `Context → DataSource → Asset → Batch` pattern: ```python import great_expectations as gx context = gx.get_context(mode="ephemeral") data_source = context.data_sources.add_pandas("hopsworks_pandas") asset = data_source.add_dataframe_asset("hopsworks_asset") batch_definition = asset.add_batch_definition_whole_dataframe("hopsworks_batch") batch = batch_definition.get_batch(batch_parameters={"dataframe": df}) ge_report = batch.validate(fg.get_expectation_suite()) ``` For most pipelines you should prefer `fg.validate(df)`: Hopsworks runs the same chain internally and uploads the report to the backend in a single call. Note that you should always use an expectation suite that has been saved to Hopsworks if you intend to upload the associated validation report. ================================================================================ # Best practices Source: https://docs.hopsworks.ai/latest/user_guides/fs/feature_group/data_validation_best_practices/ # Best practices Below is a set of recommendations and code snippets to help our users follow best practices when it comes to integrating a data validation step in your feature engineering pipelines. Rather than being prescriptive, we want to showcase how the API and configuration options can help adapt validation to your use-case. ## Development Data validation is generally considered to be a production-only feature and as such is often only setup once a project has reached the end of the development phase. At Hopsworks, we think there is a lot of value in setting up validation during early development. That's why we made it quick to get started and ensured that by default data validation is never an obstacle to inserting data. ### Validate Early As often with data validation, the best piece of advice is to set it up early in your development process. Use this phase to build a history you can then use when it becomes time to set quality requirements for a project in production. We made a code snippet to help you get started quickly: ```python import pandas as pd import great_expectations as gx from great_expectations.expectations.expectation_configuration import ( ExpectationConfiguration, ) # Load sample data. # Replace it with your own! my_data_df = pd.read_csv( "https://repo.hops.works/master/hopsworks-tutorials/data/card_fraud_data/credit_cards.csv" ) # Build a starter Expectation Suite that asserts every column exists and is # not null. This is a useful baseline; tighten it as you learn the data. # Build the list first and pass it to the constructor: GE 1.x's # add_expectation_configuration() deduplicates expect_column_to_exist entries, # but the constructor's expectations= argument preserves every entry. expectations = [] for column in my_data_df.columns: expectations.append( ExpectationConfiguration( type="expect_column_to_exist", kwargs={"column": column} ) ) expectations.append( ExpectationConfiguration( type="expect_column_values_to_not_be_null", kwargs={"column": column} ) ) expectation_suite = gx.ExpectationSuite( name="credit_cards_baseline", expectations=expectations ) # Create a Feature Group on Hopsworks with the suite attached. # Don't forget to change the primary key! my_validated_data_fg = fs.get_or_create_feature_group( name="my_validated_data_fg", version=1, description="My data", primary_key=["cc_num"], expectation_suite=expectation_suite, ) ``` Any data you insert in the Feature Group from now will be validated and a report will be uploaded to Hopsworks. ```python # Insert and validate your data insert_job, validation_report = my_validated_data_fg.insert(my_data_df) ``` Great Expectations 0.18.x shipped a `BasicSuiteBuilderProfiler` that auto-generated a starter suite from a sample DataFrame. That profiler was removed in 1.0 with no in-tree replacement, so the snippet above builds a minimal suite by hand. For real workloads, iterate on it: add `expect_column_(min/max/mean/stdev)_to_be_between`, `expect_column_values_to_be_unique`, and similar checks as you understand the distributions. Attaching the suite when creating the Feature Group ensures every piece of data finding its way into Hopsworks gets validated. Hopsworks defaults to its `"ALWAYS"` ingestion policy, meaning data is ingested whether validation succeeds or not. This way data validation is not a barrier, just a monitoring tool. ### Identify Unreliable Features Once you setup data validation, every insertion will upload a validation report to Hopsworks. Identifying Features which often have null values or wild statistical variations can help detecting unreliable Features that need refinements or should be avoided. Here are a few expectations you might find useful: - `expect_column_values_to_not_be_null` - `expect_column_(min/max/mean/stdev)_to_be_between` - `expect_column_values_to_be_unique` ### Get the stakeholders involved Hopsworks UI helps involve every project stakeholder by enabling both setting and monitoring of data quality requirements. No coding skills needed! You can monitor data quality requirements by checking out the validation reports and results on the Feature Group page. If you need to set or edit the existing requirements, you can go on the Feature Group edit page. The Expectation suite section allows you to edit individual expectations and set success parameters that match ever changing business requirements. ## Production Models in production require high-quality data to make accurate predictions for your customers. Hopsworks can use your Expectation Suite as a gatekeeper to make it simple to prevent low-quality data to make its way into production. Below are some simple tips and snippets to make the most of your data validation when your project is ready to enter its production phase. ### Be Strict in Production Whether you use an existing or create a new (recommended) Feature Group for production, we recommend you set the validation ingestion policy of your Expectation Suite to `"STRICT"`. ```python fg_prod.save_expectation_suite(my_suite, validation_ingestion_policy="STRICT") ``` In this setup, Hopsworks will abort inserting a DataFrame that does not successfully fulfill all expectations in the attached Expectation Suite. This ensures data quality standards are upheld for every insertion and provide downstream users with strong guarantees. ### Avoid Data Loss on materialization jobs Aborting insertions of DataFrames which do not satisfy the data quality standards can lead to data loss in your materialization job. To avoid such loss we recommend creating a duplicate Feature Group with the same Expectation Suite in `"ALWAYS"` mode which will hold the rejected data. ```python job, report = fg_prod.insert(df) if report["success"] is False: job, report = fg_rejected.insert(df) ``` ### Take Advantage of the Validation History You can easily retrieve the validation history of a specific expectation to export it to your favourite visualisation tool. You can filter on time and on whether insertion was successful or not. ```python validation_history = fg.get_validation_history( expectation_id=my_id, filter_by=["REJECTED", "UNKNOWN"], ge_type=False ) timeseries = pd.DataFrame( { "observed_value": [ res.result["observed_value"] for res in validation_history ], "validation_time": [res.validation_time for res in validation_history], } ) # export to your preferred Dashboard ``` ### Setup Alerts While checking your feature engineering pipeline executed properly in the morning can be good enough in the development phase, it won't make the cut for demanding production use-cases. In Hopsworks, you can setup alerts if ingestion fails or succeeds. First you will need to configure your preferred communication endpoint: slack, email or pagerduty. Check out [this page](../../../setup_installation/admin/alert.md) for more information on how to set it up. A typical use-case would be to add an alert on ingestion success to a Feature Group you created to hold data that failed validation. Here is a quick walkthrough: 1. Go the Feature Group page in the UI 2. Scroll down and click on the `Add an alert` button. 3. Choose the trigger, receiver and severity and click save. ## Conclusion Hopsworks extends Great Expectations by automatically running the validation, persisting the reports along your data and allowing you to monitor data quality in its UI. How you decide to make use of these tools depends on your application and requirements. Whether in development or in production, real-time or batch, we think there is configuration that will work for your team. Check out our [quick hands-on tutorial](https://colab.research.google.com/github/logicalclocks/hopsworks-tutorials/blob/master/integrations/great_expectations/fraud_batch_data_validation.ipynb) to start applying what you learned so far. ================================================================================ # Feature Monitoring Source: https://docs.hopsworks.ai/latest/user_guides/fs/feature_group/feature_monitoring/ # Feature Monitoring for Feature Groups Feature Monitoring complements the Hopsworks data validation capabilities for Feature Groups by allowing you to monitor your data once they have been ingested into the Feature Store. Hopsworks feature monitoring is centered around two functionalities: **scheduled statistics** and **statistics comparison**. Before continuing with this guide, see the [Feature monitoring guide](../feature_monitoring/index.md) to learn more about how feature monitoring works, and get familiar with the different use cases of feature monitoring for Feature Groups described in the **Use cases** sections of the [Scheduled statistics guide](../feature_monitoring/scheduled_statistics.md#use-cases) and [Statistics comparison guide](../feature_monitoring/statistics_comparison.md#use-cases). !!! warning "Limited UI support" Currently, feature monitoring can only be configured using the [Hopsworks Python library](https://pypi.org/project/hopsworks). However, you can enable/disable a feature monitoring configuration or trigger the statistics comparison manually from the UI. ## Code In this section, we show you how to setup feature monitoring in a Feature Group using the ==Hopsworks Python library==. Alternatively, you can get started quickly by running our [tutorial for feature monitoring](https://github.com/logicalclocks/hopsworks-tutorials/blob/master/api_examples/feature_monitoring.ipynb). First, checkout the pre-requisite and Hopsworks setup to follow the guide below. Create a project, install the [Hopsworks Python library](https://pypi.org/project/hopsworks) in your environment, connect via the generated API key. The second step is to start a new configuration for feature monitoring. After that, you can optionally define a detection window of data to compute statistics on, or use the default detection window (i.e., whole feature data). If you want to setup scheduled statistics alone, you can jump to the last step to save your configuration. Otherwise, the third and fourth steps are also optional and show you how to setup the comparison of statistics on a schedule by defining a reference window and specifying the statistics metric to monitor. ### Step 1: Pre-requisite In order to setup feature monitoring for a Feature Group, you will need: - A Hopsworks project. If you don't have a project yet you can go to [run.hopsworks.ai](https://run.hopsworks.ai), signup with your email and create your first project. - An API key, you can get one by going to "Account Settings" on [run.hopsworks.ai](https://run.hopsworks.ai). - The Hopsworks Python library installed in your client. See the [installation guide](../../client_installation/index.md). - A Feature Group #### Connect your notebook to Hopsworks Connect the client running your notebooks to Hopsworks. === "Python" ```python import hopsworks project = hopsworks.login() fs = project.get_feature_store() ``` See the API reference for [`hopsworks.login`][hopsworks.login] and [`Project.get_feature_store`][hopsworks_common.project.Project.get_feature_store]. You will be prompted to paste your API key to connect the notebook to your project. The `fs` Feature Store entity is now ready to be used to insert or read data from Hopsworks. #### Get or create a Feature Group Feature monitoring can be enabled on already created Feature Groups. We suggest you read the [Feature Group](../../../concepts/fs/feature_group/fg_overview.md) concept page to understand what a feature group is and how it fits in the ML pipeline. We also suggest you familiarize with the APIs to [create a feature group](./create.md). The following is a code example for getting or creating a Feature Group with name `trans_fg` for transaction data. === "Python" ```python # Retrieve an existing feature group trans_fg = fs.get_feature_group("trans_fg", version=1) # Or, create a new feature group with transactions trans_fg = fs.get_or_create_feature_group( name="trans_fg", version=1, description="Transaction data", primary_key=["cc_num"], event_time="datetime", ) trans_fg.insert(transactions_df) ``` See the API reference for [`FeatureStore.get_feature_group`][hsfs.feature_store.FeatureStore.get_feature_group] and [`FeatureStore.get_or_create_feature_group`][hsfs.feature_store.FeatureStore.get_or_create_feature_group]. ### Step 2: Initialize configuration #### Scheduled statistics You can setup statistics monitoring on a ==single feature or multiple features== of your Feature Group. === "Python" ```python # compute statistics for all the features fg_monitoring_config = trans_fg.create_scheduled_statistics( name="trans_fg_all_features_monitoring", description="Compute statistics on all data of all features of the Feature Group on a daily basis", ) # or for one or more specific features fg_monitoring_config = trans_fg.create_scheduled_statistics( name="trans_fg_amount_monitoring", description="Compute statistics on all data of selected features of the Feature Group on a daily basis", feature_names=["amount"], ) ``` See the API reference for [`FeatureGroup.create_scheduled_statistics`][hsfs.feature_group.FeatureGroup.create_scheduled_statistics]. #### Statistics comparison When enabling the comparison of statistics in a feature monitoring configuration, the feature to compare is selected later in the `compare_on` (or `compare_on_distribution`) method, not in `create_feature_monitoring`. You can create multiple feature monitoring configurations for the same Feature Group. === "Python" ```python fg_monitoring_config = trans_fg.create_feature_monitoring( name="trans_fg_amount_monitoring", description="Compute and compare descriptive statistics on the Feature Group on a daily basis", ) ``` See the API reference for [`FeatureGroup.create_feature_monitoring`][hsfs.feature_group.FeatureGroup.create_feature_monitoring]. #### Custom schedule By default, the computation of statistics is scheduled to run endlessly, every day at 12PM. You can modify the default schedule by adjusting the `cron_expression`, `start_date_time` and `end_date_time` parameters. To compute statistics on only a subset of the feature data, use the `row_percentage` parameter of `with_detection_window` (see Step 3). === "Python" ```python fg_monitoring_config = trans_fg.create_scheduled_statistics( name="trans_fg_all_features_monitoring", description="Compute statistics on all data of all features of the Feature Group on a weekly basis", cron_expression="0 0 12 ? * MON *", # weekly ) # or fg_monitoring_config = trans_fg.create_feature_monitoring( name="trans_fg_amount_monitoring", description="Compute and compare descriptive statistics on the Feature Group on a weekly basis", cron_expression="0 0 12 ? * MON *", # weekly ) ``` ### Step 3: (Optional) Define a detection window By default, the detection window is an _expanding window_ covering the whole Feature Group data. You can define a different detection window using the `window_length` and `time_offset` parameters provided in the `with_detection_window` method. Additionally, you can specify the percentage of feature data on which statistics will be computed using the `row_percentage` parameter. === "Python" ```python fm_monitoring_config.with_detection_window( window_length="1w", # data ingested during one week time_offset="1w", # starting from last week row_percentage=0.8, # use 80% of the data ) ``` See the API reference for [`FeatureMonitoringConfig.with_detection_window`][hsfs.core.feature_monitoring_config.FeatureMonitoringConfig.with_detection_window]. #### Time basis of the windows Rolling windows select rows by an event-time feature when the Feature Group declares one, and by commit time otherwise. The `event_time` parameter of `create_scheduled_statistics` and `create_feature_monitoring` overrides that default for the whole configuration, detection and reference windows alike. Pass a feature name to use another timestamp, date or epoch feature of the Feature Group, or `False` to select rows by commit time. === "Python" ```python # windows over the transaction time, the Feature Group event_time (default) fg_monitoring_config = trans_fg.create_feature_monitoring( name="trans_fg_amount_monitoring", ) # windows over another time feature of the Feature Group fg_monitoring_config = trans_fg.create_feature_monitoring( name="trans_fg_amount_monitoring_by_settlement", event_time="settlement_date", ) # windows over the time the rows were written (commit time) fg_monitoring_config = trans_fg.create_feature_monitoring( name="trans_fg_amount_monitoring_by_commit", event_time=False, ) ``` See [Time basis](../feature_monitoring/scheduled_statistics.md#time-basis) for how the two bases differ. ### Step 4: (Optional) Define a reference window When setting up feature monitoring for a Feature Group, you can compare the detection statistics against a reference window of feature data. A reference window is defined with the `with_reference_window` method. === "Python" ```python # compare statistics against a reference window fm_monitoring_config.with_reference_window( window_length="1w", # data ingested during one week time_offset="2w", # starting from two weeks ago row_percentage=0.8, # use 80% of the data ) ``` See the API reference for [`FeatureMonitoringConfig.with_reference_window`][hsfs.core.feature_monitoring_config.FeatureMonitoringConfig.with_reference_window]. !!! info "Comparing against a specific value" Instead of a reference window, you can compare the detection statistics against a fixed reference value (i.e., a window of size 1). In that case, skip this step and pass the `specific_value` parameter to `compare_on` in Step 5. ### Step 5.A: (Optional) Compare on a scalar metric In order to compare detection and reference statistics, you need to provide the criteria for such comparison. First, you select the feature and the metric to consider in the comparison using the `feature_name` and `metric` parameters. Then, you can define a relative or absolute threshold using the `threshold` and `relative` parameters. === "Python" ```python # compare against a reference window fm_monitoring_config.compare_on( feature_name="amount", # the feature to compare metric="mean", threshold=0.2, # a relative change over 20% is considered anomalous relative=True, # relative or absolute change strict=False, # strict or relaxed comparison ) # or compare against a specific value instead of a reference window fm_monitoring_config.compare_on( feature_name="amount", metric="mean", specific_value=100, threshold=0.2, relative=True, ) ``` See the API reference for [`FeatureMonitoringConfig.compare_on`][hsfs.core.feature_monitoring_config.FeatureMonitoringConfig.compare_on]. !!! info "Difference values and thresholds" For more information about the computation of difference values and the comparison against threshold bounds see the [Comparison criteria section](../feature_monitoring/statistics_comparison.md#comparison-criteria) in the Statistics comparison guide. ### Step 5.B: (Optional) Compare on the whole distribution Alternatively, instead of a single scalar metric, you can detect drift in the shape of a feature's distribution using `compare_on_distribution`. Select a distribution distance metric (e.g., `PSI`) and a threshold. A reference window (Step 4) is required for distribution comparison. === "Python" ```python fm_monitoring_config.compare_on_distribution( feature_name="amount", # the feature to compare metric="PSI", threshold=0.2, # a distance above 0.2 is considered a significant shift ) ``` See the API reference for [`FeatureMonitoringConfig.compare_on_distribution`][hsfs.core.feature_monitoring_config.FeatureMonitoringConfig.compare_on_distribution]. !!! tip "More distribution options" See the [Distribution comparison guide](../feature_monitoring/distribution_comparison.md) for the full list of metrics and binning strategies. ### Step 6: Save configuration Finally, you can save your feature monitoring configuration by calling the `save` method. Once the configuration is saved, the schedule for the statistics computation and comparison will be activated automatically. === "Python" ```python fm_monitoring_config.save() ``` See the API reference for [`FeatureMonitoringConfig.save`][hsfs.core.feature_monitoring_config.FeatureMonitoringConfig.save]. ### Retrieve configurations and history Once saved, you can retrieve your feature monitoring configurations and the results of past executions directly from the Feature Group. === "Python" ```python # fetch all configurations attached to the feature group configs = trans_fg.get_feature_monitoring_configs() # or a single configuration by name config = trans_fg.get_feature_monitoring_configs(name="trans_fg_amount_monitoring") # fetch the history of monitoring results (with computed statistics) history = trans_fg.get_feature_monitoring_history( config_name="trans_fg_amount_monitoring", with_statistics=True, ) ``` See the API reference for [`FeatureGroup.get_feature_monitoring_configs`][hsfs.feature_group.FeatureGroup.get_feature_monitoring_configs] and [`FeatureGroup.get_feature_monitoring_history`][hsfs.feature_group.FeatureGroup.get_feature_monitoring_history]. !!! info "Explore the API" The [`FeatureMonitoringConfig`][hsfs.core.feature_monitoring_config.FeatureMonitoringConfig] reference documents the full set of available methods, such as enabling or disabling a configuration, triggering it manually, or deleting it. ================================================================================ # Notification Source: https://docs.hopsworks.ai/latest/user_guides/fs/feature_group/notification/ # Change Data Capture for feature groups ## Introduction Changes to online-enabled feature groups can be captured by listening to events on specified topics. This optimizes the user experience by allowing users to proactively make predictions as soon as there is an update on the features. In this guide you will learn how to enable Change Data Capture (CDC) for online feature groups within Hopsworks, showing examples in Hopsworks APIs as well as the user interface. ## Prerequisites Before you begin this guide we suggest you read the [Feature Group](../../../concepts/fs/feature_group/fg_overview.md) concept page to understand what a feature group is and how it fits in the ML pipeline. Subsequently [create a Kafka topic](../../projects/kafka/create_topic.md), this topic will be used for storing Change Data Capture events. ## Using Hopsworks APIs ### Create a Feature Group with Change Data Capture using Python To enable Change Data Capture for an online-enabled feature group using the Hopsworks APIs you need to [create a feature group](./create.md) and set the `notification_topic_name` properties value to the previously created topic. === "Python" ```python fg = fs.create_feature_group( name="feature_group_name", version=feature_group_version, primary_key=feature_group_primary_keys, online_enabled=True, notification_topic_name="notification_topic_name", ) ``` ### Update Feature Group with Change Data Capture topic using Python The notification topic name can be changed after the creation of the feature group. By setting the `notification_topic_name` value to `None` or empty string notification will be disabled. With the default configuration, it can take up to 30 minutes for these changes to take place since the onlinefs service internally caches feature groups. === "Python" ```python fg.update_notification_topic_name( notification_topic_name="new_notification_topic_name" ) ``` ## Using UI ### Update Feature Group with Change Data Capture topic using UI The notification topic name can be changed after creation by editing the feature group. By setting the `CDC topic name` value to empty the notifications will be disabled. With the default configuration, it can take up to 30 minutes for these changes to take place since the onlinefs service internally caches feature groups.

Edit online enabled feature group

## Example of Change Data Capture event Once properly set up the online feature store service will produce events to the provided topic when data ingestion is completed for records. Here is an example output: ```jsonc { "projectName":"project_name", // name of the project the feature group belongs to "projectId":119, // id of the project the feature group belongs to "featureStoreId":67, // feature store where changes took place "featureGroupId":14, // id of the feature group "featureGroupName":"fg_name", // name of the feature group "featureGroupVersion":1, // version of the feature group "entry":{ // values of the affected feature group entry "id":"15", "text":"test" }, "featureViews":[ // list of feature views affected { "projectName":"project_name", // name of the project the feature view belongs to "id":9, // id of the feature view "name":"test", // name of the feature view "version":1, // version of the feature view "featurestoreId":67 // feature store where feature view resides } ] } ``` The list of `featureViews` in the event could be outdated for up to 10 minutes, due to internal logging in onlinefs service. ================================================================================ # Ingestion Topic Source: https://docs.hopsworks.ai/latest/user_guides/fs/feature_group/topic/ # How to configure the ingestion topic of a Feature Group { #feature-group-ingestion-topic } ## Introduction Feature groups written with the [streaming write API][streaming-write-api] do not write to the online and offline feature store directly. Every insert produces the rows to a Kafka topic, from which the OnlineFS service writes them to the online feature store and the offline materialization job writes them to the offline feature store. By default all feature groups in a project share a single topic. That works well until one feature group writes enough that the offline materialization jobs of the others spend their runs reading and discarding its records; [the choice of topic for data ingestion][the-choice-of-topic-for-data-ingestion] covers when a dedicated topic is worth it. In this guide you will learn how to give a feature group a topic of its own, and how to change that topic after the feature group has been created. ## Prerequisites Before you begin this guide we suggest you read the [create feature group][create-feature-group] guide, which covers the `topic_name` parameter. ## Which topic a feature group uses Hopsworks resolves the ingestion topic of a feature group in the following order: 1. The `topic_name` of the feature group, if one is set. 2. The topic of the project, if one is set. 3. The project default, which is `_onlinefs` for online-enabled feature groups and `` otherwise. Setting `topic_name` to an empty string clears the feature group override, so the feature group falls back to the project topic. !!! note "Topics of online-enabled feature groups must end in `_onlinefs`" The OnlineFS service subscribes to the topics matching the `.*_onlinefs` pattern, so a topic whose name does not match it is never consumed into the online feature store. Administrators can change the pattern with the `onlinefs/kafka_consumer/topic_pattern` configuration option, or replace it with an explicit topic list as described in the [external Kafka cluster][external-kafka-cluster] guide. ## Before you change the topic Changing the topic of a feature group that already holds data is not a migration, and there are two consequences to plan for. !!! warning "Pending data is not migrated automatically" Rows that were already inserted into the old topic but not yet consumed are never materialized to the new topic. Wait until all in-flight processing has completed before switching, for example by following the [online ingestion observability][online-ingestion-observability] of the feature group and letting the offline materialization job finish. !!! warning "Offline materialization restarts from the earliest offset" The offline materialization job stores the Kafka offsets it has consumed together with the name of the topic they belong to. When it detects that the topic has changed, those offsets are meaningless, so it starts from the earliest available offset of the new topic. This reprocesses everything the new topic still retains and can produce duplicates in the offline feature store. ## Using Hopsworks APIs ### Set the topic when creating the feature group Pass `topic_name` to `create_feature_group` to give the feature group its own topic from the start: === "Python" ```python fg = fs.create_feature_group( name="feature_group_name", version=1, primary_key=["id"], online_enabled=True, topic_name="feature_group_name_onlinefs", ) ``` ### Change the topic of an existing feature group Use [`FeatureGroup.update_topic_name`][hsfs.feature_group.FeatureGroup.update_topic_name] to point an existing feature group at a different topic: === "Python" ```python fg = fs.get_feature_group("feature_group_name", version=1) fg.update_topic_name(topic_name="feature_group_name_onlinefs") ``` The call emits the two warnings above as Python warnings before sending the request, and updates your local metadata object only once the backend has accepted the change. ## Using the UI ### Change the feature group topic Open the feature group, click `Edit`, and set the `Topic name` field. The field is shown for stream feature groups and for online-enabled feature groups, because those are the ones that ingest through Kafka. Saving a changed topic name asks you to confirm the two consequences described above before the update is sent. The topic a feature group currently uses is shown on its overview page. ### Change the project topic The project topic is the default for every feature group in the project that does not set its own. Navigate to `Project Settings` → `Kafka` and use `Edit project topic` in the `Project Topic` card. As with a feature group topic, you are asked to confirm before the change is applied. !!! note The `Project Topic` card is only shown when the cluster is configured to use an [external Kafka cluster][external-kafka-cluster], since that is the case in which the topic is not managed by Hopsworks. ## Topic creation When you set a topic that does not exist yet, Hopsworks creates it in the project with the cluster defaults for feature store topics. Topics count against the project's Kafka topic quota, which an administrator can raise as described in the [Kafka topics][kafka-topics] administration guide. When the cluster is configured to use an external Kafka cluster, Hopsworks does not provision topics. Create the topic in the external cluster first, otherwise ingestion fails as soon as the feature group starts producing to it. ## API Reference [`FeatureGroup`][hsfs.feature_group.FeatureGroup] ================================================================================ # On-Demand Transformations Source: https://docs.hopsworks.ai/latest/user_guides/fs/feature_group/on_demand_transformations/ # On-Demand Transformation Functions [On-demand transformations](https://www.hopsworks.ai/dictionary/on-demand-transformation) produce on-demand features, which usually require parameters accessible during inference for their calculation. Hopsworks facilitates the creation of on-demand transformations without introducing [online-offline skew](https://www.hopsworks.ai/dictionary/online-offline-feature-skew), ensuring consistency while allowing their dynamic computation during online inference. ## On Demand Transformation Function Creation An on-demand transformation function may be created by associating a [transformation function](../transformation_functions.md) with a feature group. Each on-demand transformation function can generate one or multiple on-demand features. If the on-demand transformation function returns a single feature, it is automatically assigned the same name as the transformation function. However, if it returns multiple features, they are by default named using the format `functionName_outputColumnNumber`. For instance, in the example below, the on-demand transformation function `transaction_age` produces an on-demand feature named `transaction_age` and the on-demand transformation function `stripped_strings` produces the on-demand features names `stripped_strings_0` and `stripped_strings_1`. Alternatively, the name of the resulting on-demand feature can be explicitly defined using the [`alias`](../transformation_functions.md#specifying-output-features-names-for-transformation-functions) function. !!! warning "On-demand transformation" All on-demand transformation functions attached to a feature group must have unique names and, in contrast to model-dependent transformations, they do not have access to training dataset statistics. Each on-demand transformation function can map specific features to its arguments by explicitly providing their names as arguments to the transformation function. If no feature names are provided, the transformation function will default to using features that match the name of the transformation function's argument. !!! example "Creating on-demand transformation functions." === "Python" ```python # Define transformation function @hopsworks.udf(return_type=int, drop=["current_date"]) def transaction_age(transaction_date, current_date): return (current_date - transaction_date).dt.days @hopsworks.udf(return_type=[str, str], drop=["current_date"]) def stripped_strings(country, city): return country.strip(), city.strip() # Attach transformation function to feature group to create on-demand transformation function. fg = feature_store.create_feature_group( name="fg_transactions", version=1, description="Transaction Features", online_enabled=True, primary_key=["id"], event_time="event_time", transformation_functions=[transaction_age, stripped_strings], ) ``` ### Specifying input features The features to be used by the on-demand transformation function can be specified by providing the feature names as input to the transformation functions. !!! example "Creating on-demand transformations by specifying features to be passed to transformation function." === "Python" ```python fg = feature_store.create_feature_group( name="fg_transactions", version=1, description="Transaction Features", online_enabled=True, primary_key=["id"], event_time="event_time", transformation_functions=[ age_transaction("transaction_time", "current_time") ], ) ``` ## Usage On-demand transformation functions attached to a feature group are automatically executed in the feature pipeline when you [insert data](./create.md#batch-write-api) into a feature group and [by the Python client while retrieving feature vectors](../feature_view/feature-vectors.md#retrieval) for online inference using feature views that contain on-demand features. The on-demand features computed by on-demand transformation functions are positioned after all other features in a feature group and are ordered alphabetically by their names. ### Inserting data All on-demand transformation functions attached to a feature group are executed whenever new data is inserted. This process computes on-demand features from historical data. The DataFrame used for insertion must include all features required for executing all on-demand transformation functions in the feature group. Inserting on-demand features as historical features saves time and computational resources by removing the need to compute all on-demand features while generating training or batch data. ### Accessing on-demand features in feature views A feature view can include on-demand features from feature groups by selecting them in the [query](../feature_view/query.md) used to create the feature view. These on-demand features are equivalent to regular features, and [model-dependent transformations](../feature_view/model-dependent-transformations.md) can be applied to them if required. !!! example "Creating feature view with on-demand features" === "Python" ```python # Selecting on-demand features in query query = fg.select( ["id", "feature1", "feature2", "on_demand_feature3", "on_demand_feature4"] ) # Creating a feature view using a query that contains on-demand transformations and model-dependent transformations feature_view = fs.create_feature_view( name="transactions_view", query=query, transformation_functions=[ min_max_scaler("feature1"), min_max_scaler("on_demand_feature3"), ], ) ``` ### Computing on-demand features On-demand features in the feature view are computed in real-time during online inference using the same on-demand transformation functions used to create them. Hopsworks, by default, automatically computes all on-demand features when retrieving feature view input features (feature vectors) with the functions `get_feature_vector` and `get_feature_vectors`. Additionally, on-demand features can be computed using the `compute_on_demand_features` function or by manually executing the same on-demand transformation function. The values for the input parameters required to compute on-demand features can be provided using the `request_parameters` argument. If values are not provided through the `request_parameters` argument, the transformation function will verify if the feature vector contains the necessary input parameters and will use those values instead. However, if the required input parameters are also not present in the feature vector, an error will be thrown. !!! note By default the functions `get_feature_vector` and `get_feature_vectors` will apply model-dependent transformation present in the feature view after computing on-demand features. #### Retrieving a feature vector The `get_feature_vector` function retrieves a single feature vector based on the feature view's serving key(s). The on-demand features in the feature vector can be computed using real-time data by passing a dictionary that associates the name of each input parameter needed for the on-demand transformation function with its respective new value to the `request_parameter` argument. !!! example "Computing on-demand features while retrieving a feature vector" === "Python" ```python feature_vector = feature_view.get_feature_vector( entry={"id": 1}, request_parameter={ "transaction_time": datetime(2022, 12, 28, 23, 55, 59), "current_time": datetime.now(), }, ) ``` #### Retrieving feature vectors The `get_feature_vectors` function retrieves multiple feature vectors using a list of feature view serving keys. The `request_parameter` in this case, can be a list of dictionaries that specifies the input parameters for the computation of on-demand features for each serving key or can be a dictionary if the on-demand transformations require the same parameters for all serving keys. !!! example "Computing on-demand features while retrieving a feature vectors" === "Python" ```python # Specify unique request parameters for each serving key. feature_vector = feature_view.get_feature_vectors( entry=[{"id": 1}, {"id": 2}], request_parameter=[ { "transaction_time": datetime(2022, 12, 28, 23, 55, 59), "current_time": datetime.now(), }, { "transaction_time": datetime(2022, 11, 20, 12, 50, 00), "current_time": datetime.now(), }, ], ) # Specify common request parameters for all serving key. feature_vector = feature_view.get_feature_vectors( entry=[{"id": 1}, {"id": 2}], request_parameter={ "transaction_time": datetime(2022, 12, 28, 23, 55, 59), "current_time": datetime.now(), }, ) ``` #### Retrieving feature vector without on-demand features The `get_feature_vector` and `get_feature_vectors` methods can return untransformed feature vectors without on-demand features by disabling model-dependent transformations and excluding on-demand features. To achieve this, set the parameters `transform` and `on_demand_features` to `False`. !!! example "Returning untransformed feature vectors" === "Python" ```python untransformed_feature_vector = feature_view.get_feature_vector( entry={"id": 1}, transform=False, on_demand_features=False ) untransformed_feature_vectors = feature_view.get_feature_vectors( entry=[{"id": 1}, {"id": 2}], transform=False, on_demand_features=False ) ``` #### Compute all on-demand features The `compute_on_demand_features` function computes all on-demand features attached to a feature view and adds them to the feature vectors provided as input to the function. This function does not apply model-dependent transformations to any of the features. The `transform` function can be used to apply model-dependent transformations to the returned values if required. The `request_parameter` in this case, can be a list of dictionaries that specifies the input parameters for the computation of on-demand features for each feature vector given as input to the function or can be a dictionary if the on-demand transformations require the same parameters for all input feature vectors. !!! example "Computing all on-demand features and manually applying model dependent transformations." === "Python" ```python # Specify request parameters for each serving key. untransformed_feature_vector = feature_view.get_feature_vector( entry={"id": 1}, transform=False, on_demand_features=False ) # re-compute and add on-demand features to the feature vector feature_vector_with_on_demand_features = fv.compute_on_demand_features( untransformed_feature_vector, request_parameter={ "transaction_time": datetime(2022, 12, 28, 23, 55, 59), "current_time": datetime.now(), }, ) # Applying model dependent transformations encoded_feature_vector = fv.transform(feature_vector_with_on_demand_features) # Specify request parameters for each serving key. untransformed_feature_vectors = feature_view.get_feature_vectors( entry=[{"id": 1}, {"id": 2}], transform=False, on_demand_features=False ) # re-compute and add on-demand features to the feature vectors - Specify unique request parameter for each feature vector feature_vectors_with_on_demand_features = fv.compute_on_demand_features( untransformed_feature_vectors, request_parameter=[ { "transaction_time": datetime(2022, 12, 28, 23, 55, 59), "current_time": datetime.now(), }, { "transaction_time": datetime(2022, 11, 20, 12, 50, 00), "current_time": datetime.now(), }, ], ) # re-compute and add on-demand feature to the feature vectors - Specify common request parameter for all feature vectors feature_vectors_with_on_demand_features = fv.compute_on_demand_features( untransformed_feature_vectors, request_parameter={ "transaction_time": datetime(2022, 12, 28, 23, 55, 59), "current_time": datetime.now(), }, ) # Applying model dependent transformations encoded_feature_vector = fv.transform(feature_vectors_with_on_demand_features) ``` #### Compute one on-demand feature On-demand transformation functions can also be accessed and executed as normal functions by using the dictionary `on_demand_transformations` that maps the on-demand features to their corresponding on-demand transformation function. !!! example "Executing each on-demand transformation function" === "Python" ```python # Specify request parameters for each serving key. feature_vector = feature_view.get_feature_vector( entry={"id": 1}, transform=False, on_demand_features=False, return_type="pandas", ) # Applying model dependent transformations feature_vector["on_demand_feature1"] = fv.on_demand_transformations[ "on_demand_feature1" ](feature_vector["transaction_time"], datetime.now()) ``` ## Chaining On-Demand Transformations On-demand transformations attached to the same feature group can be chained: one transformation's output column can serve as another transformation's input. The execution order is resolved automatically, and the resulting DAG is visible from the feature group overview page in the Hopsworks UI. !!! example "On-demand transformation that consumes an upstream output" === "Python" ```python from hopsworks import udf @udf(int, drop=["raw"]) def add_one(raw): return raw + 1 @udf(int, drop=["col"]) def double(col): return col * 2 fg = fs.create_feature_group( name="chained_odt_fg", version=1, primary_key=["id"], transformation_functions=[ add_one("raw").alias("raw_plus_one"), double("raw_plus_one").alias("raw_plus_one_doubled"), ], ) ``` Columns consumed only by the chain can be dropped, as the raw input `raw` and the intermediate `raw_plus_one` are in the example, leaving `raw_plus_one_doubled` as the only stored output. The full chain still executes during online serving, and dropped columns never become stored features. An on-demand transformation's output column becomes a regular feature in the feature group, which a downstream feature view can consume and pass into a model-dependent transformation. This is the implicit chaining path between on-demand and model-dependent transformations, with no additional setup on either side. ================================================================================ # Online Ingestion Observability Source: https://docs.hopsworks.ai/latest/user_guides/fs/feature_group/online_ingestion_observability/ # Online ingestion observability { #online-ingestion-observability } ## Introduction Knowing when ingested data becomes available for online serving, and understanding the cause of any ingestion failures, is crucial for users. To address this, the Hopsworks API provides observability features for online ingestion, allowing you to monitor ingestion status and troubleshoot issues. This guide explains how to use these observability features for online feature groups in Hopsworks, with examples using both the Hopsworks APIs and the user interface. ## Prerequisites Before you begin this guide we suggest you read the [Feature Group](../../../concepts/fs/feature_group/fg_overview.md) concept page to understand what a feature group is and how it fits in the ML pipeline. ## Using the Hopsworks API ### Create a Feature Group and Ingest Data First, create an online-enabled feature group and insert data into it: === "Python" ```python fg = fs.create_feature_group( name="feature_group_name", version=feature_group_version, primary_key=feature_group_primary_keys, online_enabled=True, ) fg.insert(fg_df) ``` ### Retrieve Online Ingestion Status After inserting data, you can monitor the ingestion progress: #### Get the latest ingestion instance === "Python" ```python oi = fg.get_latest_online_ingestion() ``` #### Get a specific ingestion by its ID === "Python" ```python oi = fg.get_online_ingestion(ingestion_id) ``` ### Use the Online Ingestion Object The online ingestion object provides methods to track and debug the ingestion process: #### Wait for completion Wait for the online ingestion to finish (equivalent to `fg.insert(fg_df, wait=True)`): === "Python" ```python oi.wait_for_completion() ``` #### Print mini-batch results Check the results of the ingestion. If the status is `UPSERTED` and the number of rows matches your data, the ingestion was successful: === "Python" ```python print([result.to_dict() for result in oi.results]) # Example output: [{'onlineIngestionId': 1, 'status': 'UPSERTED', 'rows': 10}] ``` #### Print ingestion service logs Retrieve logs from the online ingestion service to diagnose any issues: === "Python" ```python oi.print_logs(priority="error", size=5) ``` ## Using the UI ### Viewing Online Ingestion Status After inserting data into an online-enabled feature group, you can track the ingestion progress in the `Recent activities` section of the feature group in the Hopsworks UI.

See online ingestion status

================================================================================ # Time-To-Live (TTL) Source: https://docs.hopsworks.ai/latest/user_guides/fs/feature_group/ttl/ ## Feature Group TTL Usage Guide Time To Live (TTL) is a feature that automatically expires data in feature groups after a specified time period. This guide explains when and how to use TTL in your feature groups. ### Use Case: When to Use TTL TTL is particularly useful for feature groups that contain time-sensitive data that becomes stale or irrelevant after a certain period. Common use cases include: - **Regulatory compliance**: Data that must be automatically purged after a retention period for privacy or compliance reasons (e.g., GDPR, HIPAA) - **Cost optimization**: Reducing storage costs by automatically removing outdated data that is no longer needed for model inference - **Data freshness**: Ensuring that only recent, relevant data is available for online serving, preventing models from using stale features For example, if you're building a recommendation system, you might want user interaction features (like "items viewed in the last hour") to automatically expire after 1 hour, ensuring your model only uses current, relevant data. --- ## Getting Started ### Creating a Feature Group with TTL When creating a new feature group, you can enable TTL by specifying the `ttl` parameter. The TTL value determines how long data will remain in the feature group before being automatically expired. The TTL is calculated based on the `event_time` column. Data rows where `event_time` is older than the TTL period will be automatically removed. ```python from datetime import datetime, timezone import pandas as pd # Assume you already have a feature store handle # fs = ... now = datetime.now(timezone.utc) df = pd.DataFrame( { "id": [0, 1, 2], "timestamp": [now, now, now], "feature1": [10, 20, 30], "feature2": ["a", "b", "c"], } ) # Create a feature group with TTL enabled (60 seconds) fg = fs.create_feature_group( name="fg_ttl_example", version=1, primary_key=["id"], event_time="timestamp", online_enabled=True, ttl=60, # TTL in seconds - data will expire after 60 seconds ) fg.insert( df, write_options={ "start_offline_materialization": False, "wait_for_online_ingestion": True, }, ) # After 60 seconds, reading online will return empty data fg.read(online=True) # Returns empty DataFrame after TTL expires ``` For detailed API reference on all possible types of TTL values, see the [FeatureStore.create_feature_group API documentation][hsfs.feature_store.FeatureStore.create_feature_group]. --- ## Managing TTL on Existing Feature Groups ### Updating the TTL Value You can change the TTL value for an existing feature group at any time. This is useful when you need to adjust the retention period based on changing requirements. ```python # Get your existing feature group fg = fs.get_feature_group( name="fg_ttl_example", version=1, ) # Update TTL to a new value (120 seconds = 2 minutes) fg.enable_ttl(ttl=120) ``` After updating the TTL, the new retention period will apply to all future data insertions and will affect when existing data expires. --- ### Disabling and Re-enabling TTL You can temporarily disable TTL on a feature group if you need to retain data indefinitely, and then re-enable it later. #### Disabling TTL ```python # Disable TTL - data will no longer expire automatically fg.disable_ttl() ``` #### Re-enabling TTL When re-enabling TTL, you have two options: 1. **Re-enable with the previous TTL value**: If you don't specify a TTL value, the feature group will use the last TTL value that was set. ```python # Re-enable TTL using the previous TTL value fg.enable_ttl() ``` 2. **Re-enable with a new TTL value**: Specify a new TTL value when re-enabling. ```python # Re-enable TTL with a new value (90 seconds) fg.enable_ttl(ttl=90) ``` **Important**: If TTL was never set on the feature group before, you must provide a TTL value when enabling it. Otherwise, TTL cannot be enabled. --- ### Enabling TTL on an Existing Feature Group If you created a feature group without TTL initially, you can enable it later: ```python # Get an existing feature group that was created without TTL fg = fs.get_feature_group( name="fg_existing_no_ttl", version=1, ) # Enable TTL for the first time (60 seconds) fg.enable_ttl(ttl=60) ``` Once enabled, TTL will apply to all data in the feature group based on the `event_time` column. For detailed API reference on all possible types of TTL values and additional options, see the [FeatureGroup.enable_ttl API documentation][hsfs.feature_group.FeatureGroup.enable_ttl]. --- ## Monitoring TTL Purging Expired rows stop appearing in query results as soon as their TTL passes. Deleting them from storage happens separately, in the background. A purge worker inside each RonDB REST Server (RDRS) process walks every TTL-enabled online table one partition at a time, deleting a batch of expired rows on each pass. Hopsworks reports what that worker is doing in two places. ### On the Feature Group Page A **TTL purge** card appears on the feature group overview whenever the feature group is online enabled and has a TTL.

TTL purge card on the feature group overview

The summary row describes the table as a whole: | Field | Meaning | | --- | --- | | Online table | The online table backing this feature group, as `database.table` | | TTL | The retention period the purge worker read from the table's schema | | Rows purged | Rows deleted from this table, summed over the nodes that answered | | RDRS nodes | How many RonDB REST Server processes reported on this table | One row follows per RDRS node, because each node runs its own worker over its own partitions: | Field | Meaning | | --- | --- | | RDRS node | The node these counters came from | | Rows purged | Rows this node deleted from the table | | Partition | Where this node's cursor sits in the table's partition rotation | | Batch size | Rows attempted per partition visit, which the worker adapts on its own | | Last visited | When this node last visited the table | | Process started | When this node's RDRS process last started, so you can tell how much history its row count covers | A feature group created moments ago is not listed straight away. RDRS discovers TTL-enabled tables on a periodic schema scan, so for the first few seconds the card reports that no purge worker is tracking the feature group yet. It starts reporting counters on the next scan. ### Cluster-Wide Administrators can see every RDRS node's purge worker under **Settings → TTL Purge**.

Cluster-wide TTL purge page in Hopsworks settings

Each node reports a state: | State | Meaning | | --- | --- | | `running` | Actively purging | | `paused` | Healthy, but no TTL-enabled tables exist to work on | | `disabled` | Purging is switched off by configuration | | `outside window` | Outside the configured daily purge window | | `stopped` | Not started yet | | `error` | The worker hit an error, and the RDRS log has the detail | Only `error` indicates a fault. A cluster with no TTL-enabled feature groups sits in `paused`, which is the healthy idle state. Alongside the state, each node reports its counters (tables tracked, rows purged, rounds completed), the configuration it is running with (batch size range, sleep interval), and when its process last started. A restart count sits next to that timestamp, counting restarts of the container within its current pod; replacing the pod, as a redeploy does, starts a fresh count, so the start time is the figure to trust. Both views poll every ten seconds, show when the next refresh is due, and offer a **Refresh** button for an immediate read. ### Reading the Numbers A few properties of these counters are worth knowing before you draw conclusions from them. **The numbers are per RDRS node, and a node is not a datanode.** A node here is a RonDB REST Server process. Scaling RDRS changes how many rows the views list; adding datanodes does not, and shows up instead as a larger partition count. Each node keeps its counters in memory and starts again from zero when its process restarts, and nothing is persisted. Because different nodes purge different partitions, their per-table numbers legitimately differ. The cumulative row count is the only figure that is summed across nodes. **A counter is only as old as the process reporting it.** Nothing is persisted, so every figure on these views runs from the moment that node's RDRS process last started, which both views report as **Process started**. Read a low row count against that time rather than on its own: a worker that has been up for a minute and one that has been quietly idle for a week look identical without it. **A round that deleted nothing still counts as activity.** The last round timestamp advances on every pass, including passes that found nothing to delete. It tells you the worker is alive, not that rows were removed. **Rows that are already expired when you insert them never reach the online store.** Rows whose `event_time` is older than the TTL at insert time are filtered out before they are written, so they are never counted as purged. To watch the purge worker at work, insert rows that expire after they are written: ```python from datetime import datetime, timedelta, timezone import pandas as pd # Assume you already have a feature group with a TTL # fg = ... size = 100 now = datetime.now(timezone.utc) df = pd.DataFrame( { "id": range(size), # One second apart, so the rows come up for purging at a steady rate # instead of the whole batch expiring at once. "timestamp": pd.date_range(now, periods=size, freq=timedelta(seconds=1)), "feature1": range(size), } ) fg.insert(df) ``` **A batch size sitting at its configured maximum means the worker is behind.** The worker raises the batch size while there is a backlog and lowers it once it catches up, so a value pinned at the maximum shown on the cluster-wide page is the clearest sign that purging is not keeping up with expiry. ================================================================================ # Feature View User Guides Source: https://docs.hopsworks.ai/latest/user_guides/fs/feature_view/ # Feature View User Guides A feature view is a query over feature groups plus the metadata a model needs to read it consistently. These guides cover creating one, reading training and inference data, and keeping the two aligned.
- :material-eye-plus-outline:{ .lg .middle } **Start here** --- Select features from one or more feature groups and save the selection as a feature view. ```python query = trans_fg.select_all().join(profile_fg.select(["age"])) fv = fs.get_or_create_feature_view( name="transactions_fraud", version=1, query=query, labels=["fraud_label"], ) ``` [Create a feature view](overview.md) · [Training data](training-data.md) · [Feature vectors](feature-vectors.md)
:material-eye-plus-outline:{ .hops-role-ico } Create { .hops-role-cap } - [Create a feature view](overview.md) Select, join, filter and label, then save a version. - [Query](query.md) Joins, filters and point-in-time correctness. - [Helper columns](helper-columns.md) Columns for training or inference logic that are not model inputs. - [Spines](spine-query.md) Bring your own keys and labels at read time. - [Model-dependent transformations](model-dependent-transformations.md) Scaling and encoding fitted on training data, applied on read.
:material-database-export-outline:{ .hops-role-ico } Read { .hops-role-cap } - [Training data](training-data.md) Splits by ratio or time, materialised or in memory. - [Batch data](batch-data.md) Inference data for a time range, with transformations applied. - [Feature vectors](feature-vectors.md) Single or batched online lookups by serving key. - [Feature server](feature-server.md) Online lookups over REST, without the Python client.
:material-monitor-eye:{ .hops-role-ico } Observe { .hops-role-cap } - [Feature monitoring](feature_monitoring.md) Compare new data against a training dataset. - [Feature logging](feature_logging.md) Log the features a model actually saw at inference.
================================================================================ # Overview Source: https://docs.hopsworks.ai/latest/user_guides/fs/feature_view/overview/ # Feature View A feature view is a set of features that come from one or more feature groups. It is a logical view over the feature groups, as the feature data is only stored in feature groups. Feature views are used to read feature data for both training and serving (online and batch). You can create [training datasets](training-data.md), create [batch data](batch-data.md) and get [feature vectors](feature-vectors.md). If you want to understand more about the concept of feature view, you can refer to the [Feature View Overview](../../../concepts/fs/feature_view/fv_overview.md). ## Feature View Creation [Query](./query.md) and [transformation function](./model-dependent-transformations.md) are the building blocks of a feature view. You can define your set of features by building a `query`. You can also define which columns in your feature view are the `labels`, which is useful for supervised machine learning tasks. Furthermore, in python client, each feature can be attached to its own transformation function. This way, when a feature is read (for training or scoring), the transformation is executed on-demand - just before the feature data is returned. For example, when a client reads a numerical feature, the feature value could be normalized by a StandardScalar transformation function before it is returned to the client. === "Python" ```python # create a simple feature view feature_view = fs.create_feature_view(name="transactions_view", query=query) # create a feature view with transformation and label feature_view = fs.create_feature_view( name="transactions_view", query=query, labels=["fraud_label"], transformation_functions={ "amount": fs.get_transformation_function( name="standard_scaler", version=1 ) }, ) ``` === "Java" ```java // create a simple feature view FeatureView featureView = featureStore.createFeatureView() .name("transactions_view") .query(query) .build(); // create a feature view with label FeatureView featureView = featureStore.createFeatureView() .name("transactions_view") .query(query) .labels(Lists.newArrayList("fraud_label")) .build(); ``` You can refer to [query](./query.md) and [transformation function](./model-dependent-transformations.md) for creating `query` and `transformation_function`. To see a full example of how to create a feature view, you can read [this notebook](https://github.com/logicalclocks/hopsworks-tutorials/blob/master/batch-ai-systems/fraud_batch/2_fraud_batch_training_pipeline.ipynb). ## Retrieval Once you have created a feature view, you can retrieve it by its name and version. === "Python" ```python feature_view = fs.get_feature_view(name="transactions_view", version=1) ``` === "Java" ```java FeatureView featureView = featureStore.getFeatureView("transactions_view", 1) ``` ## Deletion If there are some feature view instances which you do not use anymore, you can delete a feature view. It is important to mention that all training datasets (include all materialised hopsfs training data) will be deleted along with the feature view. === "Python" ```python feature_view.delete() ``` === "Java" ```java featureView.delete() ``` ## Tags Feature views also support tags. You can attach, get, and remove tags. You can learn more in [Tags Guide](../tags/tags.md). === "Python" ```python # attach feature_view.add_tag(name="tag_schema", value={"key": "value"}) # get feature_view.get_tag(name="tag_schema") # remove feature_view.delete_tag(name="tag_schema") ``` === "Java" ```java // attach Map tag = Maps.newHashMap(); tag.put("key", "value"); featureView.addTag("tag_schema", tag) // get featureView.getTag("tag_schema") // remove featureView.deleteTag("tag_schema") ``` ## Next Once you have created a feature view, you can now [create training data](./training-data.md) ================================================================================ # Training data Source: https://docs.hopsworks.ai/latest/user_guides/fs/feature_view/training-data/ # Training data Training data can be created from the feature view and used by different ML libraries for training different models. You can read [training data concepts](../../../concepts/fs/feature_view/offline_api.md) for more details. To see a full example of how to create training data, you can read [this notebook](https://github.com/logicalclocks/hopsworks-tutorials/blob/master/batch-ai-systems/fraud_batch/2_fraud_batch_training_pipeline.ipynb). [](){ #arrowflight-server-with-duckdb } Python clients read and create in-memory training data through the ArrowFlight Server with DuckDB, which Hopsworks enables by default. For small and moderately sized datasets (what fits in a pandas DataFrame) it avoids the start-up cost of a Spark job; larger datasets can still be created with Spark by setting `read_options={"use_hive": True}`. ## Creation It can be created as in-memory DataFrames or materialised as `tfrecords`, `parquet`, `csv`, or `tsv` files to HopsFS or in all other locations, for example, S3, GCS. If you materialise a training dataset, a `PySparkJob` will be launched. By default, `create_training_data` waits for the job to finish. However, you can run the job asynchronously by passing `write_options={"wait_for_job": False}`. You can monitor the job status in the [jobs overview UI](../../projects/jobs/pyspark_job.md#step-1-jobs-overview). ```python # create a training dataset as dataframe feature_df, label_df = feature_view.training_data( description="transactions fraud batch training dataset", ) # materialise a training dataset version, job = feature_view.create_training_data( description="transactions fraud batch training dataset", data_format="csv", write_options={"wait_for_job": False}, ) # By default, it is materialised to HopsFS print(job.id) # get the job's id and view the job status in the UI ``` !!! note "Growing a training dataset over time" A materialized training dataset version cannot be appended to or modified in place. To retrain on new data, create a new training dataset version. If you need training data to keep growing, for example with a daily batch for a time-series model, do that computation once in a [derived feature group][assign-parents-to-a-feature-group] that is kept up to date as new data arrives, then create a new training dataset version from it whenever you need updated data. ### Extra filters {#training-data-extra-filters} Sometimes data scientists need to train different models using subsets of a dataset. For example, there can be different models for different countries, seasons, and different groups. One way is to create different feature views for training different models. Another way is to add extra filters on top of the feature view when creating training data. In the [transaction fraud example](https://github.com/logicalclocks/hopsworks-tutorials/blob/master/batch-ai-systems/fraud_batch/1_fraud_batch_feature_pipeline.ipynb), there are different transaction categories, for example: "Health/Beauty", "Restaurant/Cafeteria", "Holliday/Travel" etc. Examples below show how to create training data for different transaction categories. ```python # Create a training dataset for Health/Beauty df_health = feature_view.training_data( description="transactions fraud batch training dataset for Health/Beauty", extra_filter=trans_fg.category == "Health/Beauty", ) # Create a training dataset for Restaurant/Cafeteria and Holliday/Travel df_restaurant_travel = feature_view.training_data( description="transactions fraud batch training dataset for Restaurant/Cafeteria and Holliday/Travel", extra_filter=trans_fg.category == "Restaurant/Cafeteria" and trans_fg.category == "Holliday/Travel", ) ``` ### Lookback window for PIT joins {#training-data-lookback} When training data is materialised from a Feature View that joins multiple Feature Groups, the PIT join scans every historical partition of the root and every joined Feature Group. The `lookback` argument caps how far back the join is allowed to consider rows from the root and each joined Feature Group, so the engine can prune partitions before reading any files. Apply the same window uniformly with `FeatureGroupLookback`, or use `Lookback` for per-Feature-Group control; the argument shape mirrors the one accepted by `get_batch_data` (see [the batch-data lookback section][batch-data-lookback]). ```python import datetime from hsfs.constructor.lookback import FeatureGroupLookback version, job = feature_view.create_training_data( start_time=datetime.date(2026, 5, 10), end_time=datetime.date(2026, 5, 17), description="fraud batch training data, weekly partition pruning", lookback=FeatureGroupLookback( key="PARTITION_KEY", start=datetime.date(2026, 5, 10), end=datetime.date(2026, 5, 17), ), ) ``` Equivalent dict form (no `FeatureGroupLookback` import, but `datetime` is still required for the bound values): ```python import datetime version, job = feature_view.create_training_data( start_time=datetime.date(2026, 5, 10), end_time=datetime.date(2026, 5, 17), description="fraud batch training data, weekly partition pruning", lookback={ "key": "PARTITION_KEY", "start": datetime.date(2026, 5, 10), "end": datetime.date(2026, 5, 17), }, ) ``` For different lookbacks per joined Feature Group, pass a `Lookback`. See the [per-feature-group lookback section][batch-data-lookback] of the batch-data guide for the full shape. The resolved window is persisted with the training dataset, so re-reading the same training dataset version reconstructs the same per-join predicate. The same parameter is accepted by `create_train_test_split` and `create_train_validation_test_split`. ### Train/Validation/Test Splits In most cases, ML practitioners want to slice a dataset into multiple splits, most commonly train-test splits or train-validation-test splits, so that they can train and test their models. Feature view provides a sklearn-like API for this purpose, so it is very easy to create a training dataset with different splits. Create a training dataset (as in-memory DataFrames) or materialise a training dataset with train and test splits. ```python # create a training dataset X_train, X_test, y_train, y_test = feature_view.train_test_split(test_size=0.2) # materialise a training dataset version, job = feature_view.create_train_test_split( test_size=0.2, description="transactions fraud batch training dataset", data_format="csv", ) ``` Create a training dataset (as in-memory DataFrames) or materialise a training dataset with train, validation, and test splits. ```python # create a training dataset as DataFrame X_train, X_val, X_test, y_train, y_val, y_test = ( feature_view.train_validation_test_split( validation_size=0.3, test_size=0.2 ) ) # materialise a training dataset version, job = feature_view.create_train_validation_test_split( validation_size=0.3, test_size=0.2, description="transactions fraud batch training dataset", data_format="csv", ) ``` To create a particular in-memory training dataset with Spark instead of the ArrowFlight Server with DuckDB, set `read_options={"use_hive": True}`. ```python # create a training dataset as DataFrame with Hive X_train, X_test, y_train, y_test = feature_view.train_test_split( test_size=0.2, read_options={"use_hive": True} ) ``` ## Read Training Data Once you have created a training dataset, all its metadata are saved in Hopsworks. This enables you to reproduce exactly the same dataset at a later point in time. This holds for training data as both DataFrames or files. That is, you can delete the training data files (for example, to reduce storage costs), but still reproduce the training data files later on if you need to. ```python # get a training dataset feature_df, label_df = feature_view.get_training_data( training_dataset_version=1 ) # get a training dataset with train and test splits X_train, X_test, y_train, y_test = feature_view.get_train_test_split( training_dataset_version=1 ) # get a training dataset with train, validation and test splits X_train, X_val, X_test, y_train, y_val, y_test = ( feature_view.get_train_validation_test_split(training_dataset_version=1) ) ``` ## Passing Context Variables to Transformation Functions Once you have [defined a transformation function using a context variable](../transformation_functions.md#passing-context-variables-to-transformation-function), you can pass the required context variables using the `transformation_context` parameter when generating IN-MEMORY training data or materializing a training dataset. !!! note Passing context variables for materializing a training dataset is only supported in the PySpark Kernel. !!! example "Passing context variables while creating training data." === "Python" ```python # Passing context variable to IN-MEMORY Training Dataset. X_train, X_test, y_train, y_test = feature_view.get_train_test_split( training_dataset_version=1, primary_key=True, event_time=True, transformation_context={"context_parameter": 10}, ) # Passing context variable to Materialized Training Dataset. version, job = feature_view.get_train_test_split( training_dataset_version=1, primary_key=True, event_time=True, transformation_context={"context_parameter": 10}, ) ``` ## Read training data with primary key(s) and event time For certain use cases, e.g., time series models, the input data needs to be sorted according to the primary key(s) and event time combination. Primary key(s) and event time are not usually included in the feature view query as they are not features used for training. To retrieve the primary key(s) and/or event time when retrieving training data, you need to set the parameters `primary_key=True` and/or `event_time=True`. ```python # get a training dataset X_train, X_test, y_train, y_test = feature_view.get_train_test_split( training_dataset_version=1, primary_key=True, event_time=True, ) ``` !!! note All primary and event time columns of all the feature groups included in the feature view will be returned. If they have the same names across feature groups and the join prefix was not provided then reading operation will fail with ambiguous column exception. Make sure to define the join prefix if primary key and event time columns have the same names across feature groups. To use primary key(s) and event time column with materialized training datasets it needs to be created with `primary_key=True` and/or `with_event_time=True`. ## Deletion To clean up unused training data, you can delete all training data or for a particular version. Note that all metadata of training data and materialised files stored in HopsFS will be deleted and cannot be recreated anymore. ```python # delete a training data version feature_view.delete_training_dataset(training_dataset_version=1) # delete all training datasets feature_view.delete_all_training_datasets() ``` It is also possible to keep the metadata and delete only the materialised files. Then you can recreate the deleted files by just specifying a version, and you get back the exact same dataset again. This is useful when you are running out of storage. ```python # delete files of a training data version feature_view.purge_training_data(training_dataset_version=1) # delete files of all training datasets feature_view.purge_all_training_data() ``` To recreate a training dataset: ```python feature_view.recreate_training_dataset(training_dataset_version=1) ``` ## Tags Similar to feature view, You can attach, get, and remove tags. You can learn more in [Tags Guide](../tags/tags.md). ```python # attach feature_view.add_training_dataset_tag( training_dataset_version=1, name="tag_schema", value={"key": "value"} ) # get feature_view.get_training_dataset_tag( training_dataset_version=1, name="tag_schema" ) # remove feature_view.delete_training_dataset_tag( training_dataset_version=1, name="tag_schema" ) ``` ## Next Once you have created a training dataset and trained your model, you can deploy your model in a "batch" or "online" setting. Next, you can learn how to create [batch data](./batch-data.md) and get [feature vectors](./feature-vectors.md). ================================================================================ # Batch data Source: https://docs.hopsworks.ai/latest/user_guides/fs/feature_view/batch-data/ # Batch data (analytical ML systems) ## Creation It is very common that ML models are deployed in a "batch" setting where ML pipelines score incoming new data at a regular interval, for example, daily or weekly. Feature views support batch prediction by returning batch data as a DataFrame over a time range, by `start_time` and `end_time`. The resultant DataFrame (or batch-scoring DataFrame) can then be fed to models to make predictions. === "Python" ```python # get batch data df = feature_view.get_batch_data( start_time="20220620", end_time="20220627" ) # return a dataframe ``` === "Java" ```java Dataset ds = featureView.getBatchData("20220620", "20220627") ``` ## Retrieve batch data with primary keys and event time For certain use cases, e.g., time series models, the input data needs to be sorted according to the primary key(s) and event time combination. Or one might want to merge predictions back with the original input data for postmortem analysis. Primary key(s) and event time are not usually included in the feature view query as they are not features used for training. To retrieve the primary key(s) and/or event time when retrieving batch data for inference, you need to set the parameters `primary_key=True` and/or `event_time=True`. === "Python" ```python # get batch data df = feature_view.get_batch_data( start_time="20220620", end_time="20220627", primary_key=True, event_time=True, ) # return a dataframe with primary keys and event time ``` !!! note All primary and event time columns of all the feature groups included in the feature view will be returned. If they have the same names across feature groups and the join prefix was not provided then reading operation will fail with ambiguous column exception. Make sure to define the join prefix if primary key and event time columns have the same names across feature groups. Python clients read batch data through the ArrowFlight Server with DuckDB, which Hopsworks enables by default and which is much faster than Spark for small and moderately sized data. To read this particular batch data with Spark instead, set the read options to `{"use_hive": True}`. ```python # get batch data with Hive df = feature_view.get_batch_data( start_time="20220620", end_time="20220627", read_options={"use_hive": True} ) ``` ## Extra filters {#batch-data-extra-filters} `get_batch_data` accepts an `extra_filter` argument that lets you apply an arbitrary filter on top of the Feature View's own query filter and any training-dataset filter inherited from `init_batch_scoring`. Filters are combined with `AND`, pushed down to the storage layer, and apply equally to `get_batch_data`, `get_batch_query`, and `get_batch_query_string`. The simplest form uses a Feature Group handle to build the predicate: ```python df = feature_view.get_batch_data( start_time="20220620", end_time="20220627", extra_filter=(trans_fg.category == "Health/Beauty"), ) ``` Combine multiple predicates with `&` (AND) and `|` (OR): ```python df = feature_view.get_batch_data( extra_filter=(trans_fg.amount > 100) & (trans_fg.country.isin(["SE", "NO"])), ) ``` The Feature View's own query filter and any training-dataset filter from `init_batch_scoring` are AND-combined with `extra_filter` before the read. This is the same parameter that [training-data extra filters][training-data-extra-filters] exposes on training-data creation, so a filter expression works the same way in both APIs. ### Building filters from the Feature View alone When you do not have the Feature Group handle in scope (for example when reading a Feature View in an inference script), use `feature_view.get_feature(name)` to obtain a `Feature` directly from the Feature View's query. The returned `Feature` supports the same comparison operators (`==`, `!=`, `<`, `<=`, `>`, `>=`) and helper methods (`.like`, `.isin`, `.contains`), so it slots into `extra_filter` the same way: ```python df = feature_view.get_batch_data( extra_filter=(feature_view.get_feature("amount") > 100), ) ``` For Feature Views built from a join, `get_feature` accepts either the bare name or the prefixed name produced by the join. Bare names resolve against the left Feature Group when more than one side has the column; the prefixed form forces resolution against the joined Feature Group: ```python df = feature_view.get_batch_data( # `category` exists on both sides of the join, `sec_` selects the joined FG. extra_filter=(feature_view.get_feature("sec_category") == "A"), ) ``` If a bare name is ambiguous and no prefix is supplied, `get_feature` raises a `FeatureStoreException` listing the matching Feature Groups. ## Lookback window for PIT joins {#batch-data-lookback} Point-in-time (PIT) joins use the condition `feature_fg.event_time <= root_fg.event_time` to pick the latest matching record from each joined Feature Group. That predicate is a range comparison, not an equality, so partition pruning is defeated and every historical partition of every joined Feature Group is scanned on every read. As Feature Groups grow with daily ingestion, this scan grows unboundedly. The `lookback` argument lets you cap how far back the join is allowed to consider rows from each joined Feature Group. Hopsworks turns the window into a constant-bound predicate on the joined Feature Group so the ArrowFlight Server with DuckDB and Spark Catalyst pushdown can prune partitions before opening any files. ### Uniform lookback Apply the same window to every joined Feature Group with a `FeatureGroupLookback` instance from `hsfs.constructor.lookback`, or the equivalent dict. Both forms accept `date` and `datetime` values. ```python import datetime from hsfs.constructor.lookback import FeatureGroupLookback df = feature_view.get_batch_data( start_time=datetime.date(2026, 5, 10), end_time=datetime.date(2026, 5, 17), lookback=FeatureGroupLookback( key="PARTITION_KEY", start=datetime.date(2026, 5, 10), end=datetime.date(2026, 5, 17), ), ) ``` Equivalent dict form, no `FeatureGroupLookback` import required (you still need `datetime` for the bound values): ```python import datetime df = feature_view.get_batch_data( start_time=datetime.date(2026, 5, 10), end_time=datetime.date(2026, 5, 17), lookback={ "key": "PARTITION_KEY", "start": datetime.date(2026, 5, 10), "end": datetime.date(2026, 5, 17), }, ) ``` `key` selects which column the predicate is emitted against. `"PARTITION_KEY"` targets the Feature Group's partition column so the engine can prune partitions before reading files; the Feature Group must have a single DATE partition column. `"EVENT_TIME"` targets the Feature Group's `event_time` column and guarantees row-level correctness but offers only engine-dependent file pruning (Hudi, Delta, or Iceberg column-stats indexing). `start` is required and emits a `>=` predicate. `end` is optional and emits a `<=` predicate when present. When `end` is omitted, only the lower bound is emitted, making the short form below valid: the root Feature Group and every joined Feature Group get ` >= '2026-05-10'` (where `` is each Feature Group's own DATE partition column) and nothing else. ```python import datetime df = feature_view.get_batch_data( lookback={ "key": "PARTITION_KEY", "start": datetime.date(2026, 5, 10), }, ) ``` ### Per-feature-group lookback When different Feature Groups need different windows, use `Lookback` to bind a `FeatureGroupLookback` to specific joined Feature Groups. An optional `default` applies to every Feature Group not listed in `feature_group_lookbacks`. ```python import datetime from hsfs.constructor.lookback import FeatureGroupLookback, Lookback df = feature_view.get_batch_data( start_time=datetime.date(2026, 5, 11), end_time=datetime.date(2026, 5, 17), lookback=Lookback( default=FeatureGroupLookback( key="PARTITION_KEY", start=datetime.date(2026, 5, 5), end=datetime.date(2026, 5, 17), ), feature_group_lookbacks={ "transactions": FeatureGroupLookback( key="EVENT_TIME", start=datetime.datetime(2026, 5, 1, tzinfo=datetime.timezone.utc), ), }, ), ) ``` Skip the `default` to apply lookbacks only to the listed Feature Groups; unlisted Feature Groups receive no lookback for that call. ```python df = feature_view.get_batch_data( start_time=datetime.date(2026, 5, 11), end_time=datetime.date(2026, 5, 17), lookback=Lookback( feature_group_lookbacks={ "transactions": FeatureGroupLookback( key="PARTITION_KEY", start=datetime.date(2026, 5, 5) ), } ), ) ``` `feature_group_lookbacks` keys identify a Feature Group in one of two ways: by name (a bare string matches every version of the named Feature Group at any join site in the Feature View) or by passing the Feature Group instance itself (matches the exact `(name, version)` so a specific version can be targeted when multiple versions of the same Feature Group are joined). When both forms are supplied for the same name, the instance entry wins at its specific join site and the bare-string entry still applies elsewhere. Equivalent dict form: ```python import datetime df = feature_view.get_batch_data( start_time=datetime.date(2026, 5, 11), end_time=datetime.date(2026, 5, 17), lookback={ "default": { "key": "PARTITION_KEY", "start": datetime.date(2026, 5, 5), "end": datetime.date(2026, 5, 17), }, "feature_group_lookbacks": { "transactions": { "key": "EVENT_TIME", "start": datetime.datetime(2026, 5, 1, tzinfo=datetime.timezone.utc), }, }, }, ) ``` ### Combining `lookback` with other filters The `lookback` predicate combines with filters declared on the Query, but where the filter is attached changes whether the engine can prune partitions on the root Feature Group. Filters attached to a sub-query (`fg.select(...).filter(...)`) always prune on that Feature Group regardless of which Feature Group they reference. Filters attached to the outer query (`query.filter(...)` after the join, or `extra_filter` on `get_batch_data`) prune the root only when every referenced feature belongs to the root Feature Group. A mixed-Feature-Group outer filter still produces correct results, because the predicates apply at the outer level, but the root's partitions are no longer pruned at file-listing time. ```python # Root sub-query filter: lookback prunes both root and joined Feature Groups. query = root.select_all().filter(root.amount > 100).join(dim.select_all()) # Joined sub-query filter: lookback still prunes both sides. query = root.select_all().join(dim.select_all().filter(dim.category == "X")) # Outer filter referencing a joined Feature Group: root pruning is lost; # joined Feature Groups still prune via their own predicates. query = root.select_all().join(dim.select_all()).filter(dim.category == "X") ``` For best pruning, keep call-site filters at the sub-query level when their predicate references only one Feature Group. The same `lookback` argument is supported on `create_training_data` (see [the training-data section][training-data-lookback]). Both `extra_filter` and `lookback` can be combined. ## Creation with transformation If you have specified transformation functions when creating a feature view, you will get back transformed batch data as well. If your transformation functions require statistics of training dataset, you must also provide the training data version. `init_batch_scoring` will then fetch the statistics and initialize the functions with required statistics. Then you can follow the above examples and create the batch data. Please note that transformed batch data can only be returned in the python client but not in the java client. ```python feature_view.init_batch_scoring(training_dataset_version=1) ``` It is important to note that in addition to the filters defined in Feature View, [extra filters][training-data-extra-filters] will be applied if they are defined in the given training dataset version. ## Retrieving untransformed batch data By default, the `get_batch_data` function returns batch data with model-dependent transformations applied. However, you can retrieve untransformed batch data, while still including on-demand features, by setting the `transform` parameter to `False`. !!! example "Returning untransformed batch data" === "Python" ```python # Fetching untransformed batch data. untransformed_batch_data = feature_view.get_batch_data(transform=False) ``` ## Passing Context Variables to Transformation Functions After [defining a transformation function using a context variable](../transformation_functions.md#passing-context-variables-to-transformation-function), you can pass the necessary context variables through the `transformation_context` parameter when fetching batch data. !!! example "Passing context variables while fetching batch data." === "Python" ```python # Passing context variable to IN-MEMORY Training Dataset. batch_data = feature_view.get_batch_data( transformation_context={"context_parameter": 10} ) ``` ================================================================================ # Feature vectors Source: https://docs.hopsworks.ai/latest/user_guides/fs/feature_view/feature-vectors/ # Feature Vectors The Hopsworks Platform integrates real-time capabilities with its Online Store. Based on [RonDB](https://www.rondb.com/), your feature vectors are served at scale at in-memory latency (~1-10ms). Checkout [the benchmarks results](https://www.hopsworks.ai/post/feature-store-benchmark-comparison-hopsworks-and-feast#images-2) and [the benchmark code](https://github.com/featurestoreorg/featurestore-benchmarks). The same Feature View which was used to create training datasets can be used to retrieve feature vectors for real-time predictions. This allows you to serve the same features to your model in training and serving, ensuring consistency and reducing boilerplate. Whether you are either inside the Hopsworks platform, a model serving platform, or in an external environment, such as your application server. Below is a practical guide on how to use the Online Store Python and Java Client. The aim is to get you started quickly by providing code snippets which illustrate various use cases and functionalities of the clients. If you need to get more familiar with the concept of feature vectors, you can read this [short introduction](../../../concepts/fs/feature_view/online_api.md) first. ## Retrieval You can get back feature vectors from either python or java client by providing the primary key value(s) for the feature view. Note that filters defined in feature view and training data will not be applied when feature vectors are returned. If you need to retrieve a complete value of feature vectors without missing values, the required `entry` are [FeatureView.primary_keys][hsfs.feature_view.FeatureView.primary_keys]. Alternative, you can provide the primary key of the feature groups as the key of the entry. It is also possible to provide a subset of the entry, which will be discussed [below](#partial-feature-retrieval). === "Python" ```python # get a single vector feature_view.get_feature_vector(entry={"pk1": 1, "pk2": 2}) # get multiple vectors feature_view.get_feature_vectors( entry=[{"pk1": 1, "pk2": 2}, {"pk1": 3, "pk2": 4}, {"pk1": 5, "pk2": 6}] ) ``` === "Java" ```java // get a single vector Map entry1 = Maps.newHashMap(); entry1.put("pk1", 1); entry1.put("pk2", 2); featureView.getFeatureVector(entry1); // get multiple vectors Map entry2 = Maps.newHashMap(); entry2.put("pk1", 3); entry2.put("pk2", 4); featureView.getFeatureVectors(Lists.newArrayList(entry1, entry2)); ``` ### Required entry Starting from python client v3.4, you can specify different values for the primary key of the same name which exists in multiple feature groups but are not joint by the same name. The table below summarises the value of `primary_keys` in different settings. Considering that you are joining 2 feature groups, namely, `left_fg` and `right_fg`, the feature groups have different primary keys, and features (`feature_*`) in each setting. Also, the 2 feature groups are [joint][hsfs.constructor.query.Query.join] on different *join conditions* and *prefix* as `left_fg.join(right_fg, , prefix=)`. For java client, and python client before v3.4, the `primary_keys` are the set of primary key of all the feature groups in the query. Python client is backward compatible. It means that the `primary_keys` used before v3.4 can be applied to python client of later versions as well. The serving keys follow four rules, one per branch of the flow below: - A `left_fg` primary key is always a serving key, under its own name. - A `right_fg` primary key that the join matches to a `left_fg` primary key is covered by that key. - A `right_fg` primary key the join does not match becomes a serving key under its own name, if that name is still free. - If the name is already taken, the serving key is the join prefix plus the name, or `fgId___` plus the name when the join has no prefix. `` is `right_fg.id` and `` is the position of the feature group in the join, 1 for the first join. === "As a flow" --8<-- "user_guides/fs/feature_view/feature-vectors/serving-keys.html" === "As a table" `id = user_id` stands for `left_on=["id"], right_on=["user_id"]`, and `id = id` for `on=["id"]`. | `left_fg` keys | `right_fg` keys | join | prefix | serving keys | | --- | --- | --- | --- | --- | | id | id | `id = id` | | id | | id1 | id2 | `id1 = id2` | | id1 | | id1, id2 | id1 | `id1 = id1` | | id1, id2 | | id, user_id | id | `user_id = id` | | id, user_id | | id1 | id1, id2 | `id1 = id1` | | id1, id2 | | id | id, user_id | `id = user_id` | `right_` | id, `right_id` | | id | id, user_id | `id = user_id` | | id, `fgId___id` | | id | id | `id = feature_1` | `right_` | id, `right_id` | | id | id | `id = feature_1` | | id, `fgId___id` | | id | id | `feature_1 = id` | `right_` | id, `right_id` | | id | id | `feature_1 = id` | | id, `fgId___id` | | user, year | user, year | `user = user` | `right_` | user, year, `right_year` | | user, year | user, year | `user = user` | | user, year, `fgId___year` | For example, joining two feature groups that both have `id` as primary key on `left_on=["id"], right_on=["user_id"]` with `prefix="right_"` gives the serving keys `id` and `right_id`: ```python query = left_fg.select_all().join( right_fg.select_all(), left_on=["id"], right_on=["user_id"], prefix="right_" ) feature_view = fs.create_feature_view(name="fv", query=query) feature_view.get_feature_vector({"id": 42, "right_id": 7}) ``` ### Missing Primary Key Entries It can happen that some of the primary key entries are not available in some or all of the feature groups used by a feature view. Take the above example assuming the feature view consists of two joined feature groups, first one with primary key column `pk1`, the second feature group with primary key column `pk2`. === "Python" ```python # get a single vector feature_view.get_feature_vector(entry={"pk1": 1, "pk2": 2}) ``` === "Java" ```java // get a single vector Map entry1 = Maps.newHashMap(); entry1.put("pk1", 1); entry1.put("pk2", 2); featureView.getFeatureVector(entry1); ``` This call will raise an exception if `pk1 = 1` OR `pk2 = 2` can't be found but also if `pk1 = 1` AND `pk2 = 2` can't be found, meaning, it will not return a partial or empty feature vector. When retrieving a batch of vectors, the behaviour is slightly different. === "Python" ```python # get multiple vectors feature_view.get_feature_vectors( entry=[{"pk1": 1, "pk2": 2}, {"pk1": 3, "pk2": 4}, {"pk1": 5, "pk2": 6}] ) ``` === "Java" ```java // get multiple vectors Map entry2 = Maps.newHashMap(); entry2.put("pk1", 3); entry2.put("pk2", 4); Map entry3 = Maps.newHashMap(); entry3.put("pk1", 5); entry3.put("pk2", 6); featureView.getFeatureVectors(Lists.newArrayList(entry1, entry2, entry3)); ``` This call will raise an exception if for example for the third entry `pk1 = 5` OR `pk2 = 6` can't be found, however, it will simply not return a vector for this entry if `pk1 = 5` AND `pk2 = 6` can't be found. That means, `get_feature_vectors` will never return partial feature vector, but will omit empty feature vectors. If you are aware of missing features, you can use the [*passed features*](#passed-features) or [Partial feature retrieval](#partial-feature-retrieval) functionality, described down below. ### Partial feature retrieval If your model can handle missing value or if you want to impute the missing value, you can get back feature vectors with partial values using python client starting from version 3.4 (Note that this does not apply to java client.). In the example below, let's say you join 2 feature groups by `fg1.join(fg2, left_on=["pk1"], right_on=["pk2"])`, required keys of the `entry` are `pk1` and `pk2`. If `pk2` is not provided, this returns feature values from the first feature group and null values from the second feature group when using the option `allow_missing=True`, otherwise it raises exception. === "Python" ```python # get a single vector with feature_view.get_feature_vector(entry={"pk1": 1}, allow_missing=True) # get multiple vectors feature_view.get_feature_vectors( entry=[ {"pk1": 1}, {"pk1": 3}, ], allow_missing=True, ) ``` ### Retrieval with transformation If you have specified transformation functions when creating a feature view, you receive transformed feature vectors. If your transformation functions require statistics of training dataset, you must also provide the training data version. `init_serving` will then fetch the statistics and initialize the functions with the required statistics. Then you can follow the above examples and retrieve the feature vectors. Please note that transformed feature vectors can only be returned in the python client but not in the java client. === "Python" ```python feature_view.init_serving(training_dataset_version=1) ``` ## Passed features If some of the features values are only known at prediction time and cannot be computed and cached in the online feature store, you can provide those values as `passed_features` option. The `get_feature_vector` method is going to use the passed values to construct the final feature vector to submit to the model. You can use the `passed_features` parameter to overwrite individual features being retrieved from the online feature store. The feature view will apply the necessary transformations to the passed features as it does for the feature data retrieved from the online feature store. Please note that passed features is only available in the python client but not in the java client. === "Python" ```python # get a single vector feature_view.get_feature_vector( entry={"pk1": 1, "pk2": 2}, passed_features={"feature_a": "value_a"} ) # get multiple vectors feature_view.get_feature_vectors( entry=[{"pk1": 1, "pk2": 2}, {"pk1": 3, "pk2": 4}, {"pk1": 5, "pk2": 6}], passed_features=[ {"feature_a": "value_a1"}, {"feature_a": "value_a2"}, {"feature_a": "value_a3"}, ], ) ``` You can also use the parameter to provide values for all the features which are part of a specific feature group and used in the feature view. In this second case, you do not have to provide the primary key value for that feature group as no data needs to be retrieved from the online feature store. === "Python" ```python # get a single vector, replace values from an entire feature group # note how in this example you don't have to provide the value of # pk2, but you need to provide the features coming from that feature group # in this case feature_b and feature_c feature_view.get_feature_vector( entry={"pk1": 1}, passed_features={ "feature_a": "value_a", "feature_b": "value_b", "feature_c": "value_c", }, ) ``` ## Retrieving untransformed feature vectors By default, the `get_feature_vector` and `get_feature_vectors` functions return transformed feature vectors, which has model-dependent transformations applied and includes on-demand features. However, you can retrieve the untransformed feature vectors without applying model-dependent transformations while still including on-demand features by setting the `transform` parameter to False. !!! example "Returning untransformed feature vectors" === "Python" ```python # Fetching untransformed feature vector. untransformed_feature_vector = feature_view.get_feature_vector( entry={"id": 1}, transform=False ) # Fetching untransformed feature vectors. untransformed_feature_vectors = feature_view.get_feature_vectors( entry=[{"id": 1}, {"id": 2}], transform=False ) ``` ## Retrieving feature vector without on-demand features The `get_feature_vector` and `get_feature_vectors` methods can also return untransformed feature vectors without on-demand features by disabling model-dependent transformations and excluding on-demand features. To achieve this, set the parameters `transform` and `on_demand_features` to `False`. !!! example "Returning untransformed feature vectors" === "Python" ```python untransformed_feature_vector = feature_view.get_feature_vector( entry={"id": 1}, transform=False, on_demand_features=False ) untransformed_feature_vectors = feature_view.get_feature_vectors( entry=[{"id": 1}, {"id": 2}], transform=False, on_demand_features=False ) ``` ## Passing Context Variables to Transformation Functions After [defining a transformation function using a context variable](../transformation_functions.md#passing-context-variables-to-transformation-function), you can pass the required context variables using the `transformation_context` parameter when fetching the feature vectors. !!! example "Passing context variables while fetching batch data." === "Python" ```python # Passing context variable to IN-MEMORY Training Dataset. batch_data = feature_view.get_feature_vectors( entry=[{"pk1": 1}], transformation_context={"context_parameter": 10} ) ``` ## Retrieving feature vectors without blocking `get_feature_vector` and `get_feature_vectors` block the calling thread for the whole round trip to the online store. Inside a serving deployment, or anywhere else that runs an event loop, that stops every other request while the lookup is in flight. `get_feature_vector_async` and `get_feature_vectors_async` take the same arguments and return the same values, awaited instead. ```python vector = await my_feature_view.get_feature_vector_async(entry={"pk1": 1, "pk2": 2}) vectors = await my_feature_view.get_feature_vectors_async( entry=[{"pk1": 1, "pk2": 2}, {"pk1": 3, "pk2": 4}] ) ``` The statements are awaited on the caller's own event loop, against a connection pool belonging to that loop, so several lookups are in flight at once. On a measured deployment this raised throughput from 218 to 270 requests per second and cut p99 latency by 72 percent. The awaited path applies to the SQL client. A deployment reading through the REST client falls back to the blocking call, since there is nothing there to overlap. Each event loop gets its own connection pool, and that pool is released when its loop is collected. A process that creates a loop per lookup, for example by calling `asyncio.run` in a loop, therefore does not accumulate connections that way. The default predictor a deployment gets from `model.deploy()` or `feature_view.deploy()` already awaits its lookup. ## Choose the right Client The Online Store can be accessed via the **Python** or **Java** client allowing you to use your language of choice to connect to the Online Store. Additionally, the Python client provides two different implementations to fetch data: **SQL** or **REST**. The SQL client is the default implementation. It requires a direct SQL connection to your RonDB cluster and uses python asyncio to offer high performance even when your Feature View rows involve querying multiple different tables. The REST client is an alternative implementation connecting to [RonDB Feature Vector Server](./feature-server.md). Perfect if you want to avoid exposing ports of your database cluster directly to clients. This implementation is available as of Hopsworks 3.7. Initialise the client by calling the `init_serving` method on the Feature View object before starting to fetch feature vectors. This will initialise the chosen client, test the connection, and initialise the transformation functions registered with the Feature View. Note to use the REST client in the Hopsworks Cluster python environment you will need to provide an API key explicitly as JWT authentication is not yet supported. More configuration options can be found in the [API documentation][hsfs.feature_view.FeatureView.init_serving]. === "Python" ```python # initialize the SQL client to fetch feature vectors from the Online Store my_feature_view.init_serving() # or use the REST client my_feature_view.init_serving( init_rest_client=True, config_rest_client={ "api_key": "your_api_key", }, ) ``` Once the client is initialised, you can start fetching feature vector(s) via the Feature View methods: `get_feature_vector(s)`. You can initialise both clients for a given Feature View and switch between them by using the force flags in the get_feature_vector(s) methods. === "Python" ```python # initialize both clients and set the default to REST my_feature_view.init_serving( init_rest_client=True, init_sql_client=True, config_rest_client={ "api_key": "your_api_key", }, default_client="rest", ) # this will fetch a feature vector via REST try: my_feature_view.get_feature_vector( entry={"pk1": 1, "pk2": 2}, ) except TimeoutException: # if the REST client times out, the SQL client will be used my_feature_view.get_feature_vector( entry={"pk1": 1, "pk2": 2}, force_sql=True ) ``` ## Feature Server In addition to Python/Java clients, from Hopsworks 3.3, a new [feature server](./feature-server.md) implemented in Go is introduced. With this new API, single or batch feature vectors can be retrieved in any programming language. Note that you can connect to the Feature Vector Server via any REST client. However registered transformation function will not be applied to values in the JSON response and values stored in Feature Groups which contain embeddings will be missing. ================================================================================ # Feature server Source: https://docs.hopsworks.ai/latest/user_guides/fs/feature_view/feature-server/ # Feature Store REST API Server This API server allows users to retrieve single/batch feature vectors from a feature view. ## How to use From Hopsworks 3.3, you can connect to the Feature Vector Server via any REST client which supports POST requests. Set the `X-API-KEY` to your Hopsworks API Key and send the request with a JSON body, [single](#single-feature-vector-request) or [batch](#batch-feature-vectors-request). By default, the server listens on the `0.0.0.0:4406` and the api version is set to `0.1.0`. Please refer to `/srv/hops/mysql-cluster/rdrs_config.json` config file located on machines running the REST Server for additional configuration parameters. In Hopsworks 3.7, we introduced a python client for the Online Store REST API Server. The python client is available in the `hsfs` module and can be installed using `pip install hsfs`. This client can be used instead of the Online Store SQL client in the `FeatureView.get_feature_vector(s)` methods. Check the corresponding [documentation](./feature-vectors.md) for these methods. ## Single Feature Vector ### Single Feature Vector Request `POST /{api-version}/feature_store` #### Single Feature Vector Request Body ```json { "featureStoreName": "fsdb002", "featureViewName": "sample_2", "featureViewVersion": 1, "passedFeatures": {}, "entries": { "id1": 36 }, "metadataOptions": { "featureName": true, "featureType": true }, "options": { "validatePassedFeatures": true, "includeDetailedStatus": true } } ``` #### Single Feature Vector Request Parameters | **parameter** | **type** | **note** | | --- | --- | --- | | featureStoreName | string | | | featureViewName | string | | | featureViewVersion | number(int) | | | entries | objects | Map of serving key of feature view as key and value of serving key as value. Serving key are a set of the primary key of feature groups which are included in the feature view query. If feature groups are joint with prefix, the primary key needs to be attached with prefix. | | passedFeatures | objects | Optional. Map of feature name as key and feature value as value. This overwrites feature values in the response. | | metadataOptions | objects | Optional. Map of metadataoption as key and boolean as value. Default metadata option is false. Metadata is returned on request. Metadata options available: 1\. featureName 2\. featureType | | options | objects | Optional. Map of option as key and boolean as value. Default option is false. Options available: 1\. validatePassedFeatures 2\. includeDetailedStatus | ### Single Feature Vector Response ```json { "features": [ 36, "2022-01-24", "int24", "str14" ], "metadata": [ { "featureName": "id1", "featureType": "bigint" }, { "featureName": "ts", "featureType": "date" }, { "featureName": "data1", "featureType": "string" }, { "featureName": "data2", "featureType": "string" } ], "status": "COMPLETE", "detailedStatus": [ { "featureGroupId": 1, "httpStatus": 200, }, { "featureGroupId": 2, "httpStatus": 200, }, ] } ``` ### Single Feature Vector Errors | **Code** | **reason** | **response** | | -------- | ------------------------------------- | ------------------------------------ | | 200 | | | | 400 | Requested metadata does not exist | | | 400 | Error in pk or passed feature value | | | 401 | Access denied | Access unshared feature store failed | | 500 | Failed to read feature store metadata | | #### Response with PK/pass feature error ```json { "code": 12, "message": "Wrong primay-key column. Column: ts", "reason": "Incorrect primary key." } ``` #### Response with metadata error ```json { "code": 2, "message": "", "reason": "Feature store does not exist." } ``` #### PK value no match ```json { "features": [ 9876543, null, null, null ], "metadata": null, "status": "MISSING" } ``` #### Detailed Status If `includeDetailedStatus` option is set to true, detailed status is returned in the response. Detailed status is a list of feature group id and http status code, corresponding to each read operations perform internally by RonDB. Meaning is as follows: - `featureGroupId`: Id of the feature group, used to identify which table the operation correspond from. - `httpStatus`: Http status code of the operation. - 200 means success - 400 means bad request, likely pk name is wrong or pk is incomplete. In particular, if pk for this table/feature group is not provided in the request, this http status is returned. - 404 means no row corresponding to PK - 500 means internal error. Both `404` and `400` set the status to `MISSING` in the response. Examples below corresponds respectively to missing row and bad request. Missing Row: The PK name-value pair was correctly passed, but the corresponding row was not found in the feature group. ```json { "features": [ 36, "2022-01-24", null, null ], "status": "MISSING", "detailedStatus": [ { "featureGroupId": 1, "httpStatus": 200, }, { "featureGroupId": 2, "httpStatus": 404, }, ] } ``` Bad Request, e.g., when PK name-value pair for FG2 not provided or the corresponding column names was incorrect: ```json { "features": [ 36, "2022-01-24", null, null ], "status": "MISSING", "detailedStatus": [ { "featureGroupId": 1, "httpStatus": 200, }, { "featureGroupId": 2, "httpStatus": 400, }, ] } ``` ## Batch Feature Vectors ### Batch Feature Vectors Request `POST /{api-version}/batch_feature_store` #### Batch Feature Vectors Request Body ```json { "featureStoreName": "fsdb002", "featureViewName": "sample_2", "featureViewVersion": 1, "passedFeatures": [], "entries": [ { "id1": 16 }, { "id1": 36 }, { "id1": 71 }, { "id1": 48 }, { "id1": 29 } ], "requestId": null, "metadataOptions": { "featureName": true, "featureType": true }, "options": { "validatePassedFeatures": true, "includeDetailedStatus": true } } ``` #### Batch Feature Vectors Request Parameters | **parameter** | **type** | **note** | | --- | --- | --- | | featureStoreName | string | | | featureViewName | string | | | featureViewVersion | number(int) | | | entries | `array` | Each items is a map of serving key as key and value of serving key as value. Serving key of feature view. | | passedFeatures | `array` | Optional. Each items is a map of feature name as key and feature value as value. This overwrites feature values in the response. If provided, its size and order has to be equal to the size of entries. Item can be null. | | metadataOptions | objects | Optional. Map of metadataoption as key and boolean as value. Default metadata option is false. Metadata is returned on request. Metadata options available: 1\. featureName 2\. featureType | | options | objects | Optional. Map of option as key and boolean as value. Default option is false. Options available: 1\. validatePassedFeatures 2\. includeDetailedStatus | ### Batch Feature Vectors Response ```json { "features": [ [ 16, "2022-01-27", "int31", "str24" ], [ 36, "2022-01-24", "int24", "str14" ], [ 71, null, null, null ], [ 48, "2022-01-26", "int92", "str31" ], [ 29, "2022-01-03", "int53", "str91" ] ], "metadata": [ { "featureName": "id1", "featureType": "bigint" }, { "featureName": "ts", "featureType": "date" }, { "featureName": "data1", "featureType": "string" }, { "featureName": "data2", "featureType": "string" } ], "status": [ "COMPLETE", "COMPLETE", "MISSING", "COMPLETE", "COMPLETE" ], "detailedStatus": [ [{ "featureGroupId": 1, "httpStatus": 200, }], [{ "featureGroupId": 1, "httpStatus": 200, }], [{ "featureGroupId": 1, "httpStatus": 404, }], [{ "featureGroupId": 1, "httpStatus": 200, }], [{ "featureGroupId": 1, "httpStatus": 200, }] ] } ``` note: Order of the returned features are the same as the order of entries in the request. ### Batch Feature Vectors Errors | **Code** | **reason** | **response** | | -------- | ------------------------------------- | ------------------------------------ | | 200 | | | | 400 | Requested metadata does not exist | | | 404 | Missing row corresponding to pk value | | | 401 | Access denied | Access unshared feature store failed | | 500 | Failed to read feature store metadata | | #### Response with partial failure ```json { "features": [ [ 81, "id81", "2022-01-29 00:00:00", 6 ], null, [ 51, null, null, null, ] ], "metadata": null, "status": [ "COMPLETE", "ERROR", "MISSING" ], "detailedStatus": [ [{ "featureGroupId": 1, "httpStatus": 200, }], [{ "featureGroupId": 1, "httpStatus": 400, }], [{ "featureGroupId": 1, "httpStatus": 404, }] ] } ``` ## Access control to feature store Currently, the REST API server only supports Hopsworks API Keys for authentication and authorization. Add the API key to the HTTP requests using the `X-API-KEY` header. ================================================================================ # Query Source: https://docs.hopsworks.ai/latest/user_guides/fs/feature_view/query/ # Query vs DataFrame Hopsworks provides a DataFrame API to ingest data into the Hopsworks Feature Store. You can also retrieve feature data in a DataFrame, that can either be used directly to train models or [materialized to file(s)](./training-data.md) for later use to train models. The idea of the Feature Store is to have pre-computed features available for both training and serving models. The key functionality required to generate training datasets from reusable features are: feature selection, joins, filters, and point in time queries. The Query object enables you to select features from different feature groups to join together to be used in a feature view. The joining functionality is heavily inspired by the APIs used by Pandas to merge DataFrames. The APIs allow you to specify which features to select from which feature group, how to join them and which features to use in join conditions. === "Python" ```python fs = ... credit_card_transactions_fg = fs.get_feature_group(name="credit_card_transactions", version=1) account_details_fg = fs.get_feature_group(name="account_details", version=1) merchant_details_fg = fs.get_feature_group(name="merchant_details", version=1) # create a query selected_features = credit_card_transactions_fg.select_all() \ .join(account_details_fg.select_all(), on=["cc_num"]) \ .join(merchant_details_fg.select_all()) # save the query to feature view feature_view = fs.create_feature_view( version=1, name='credit_card_fraud', labels=["is_fraud"], query=selected_features ) # retrieve the query back from the feature view feature_view = fs.get_feature_view(“credit_card_fraud”, version=1) query = feature_view.query ``` === "Scala" ```scala val fs = ... val creditCardTransactionsFg = fs.getFeatureGroup("credit_card_transactions", 1) val accountDetailsFg = fs.getFeatureGroup(name="account_details", version=1) val merchantDetailsFg = fs.getFeatureGroup("merchant_details", 1) // create a query val selectedFeatures = (creditCardTransactionsFg.selectAll() .join(accountDetailsFg.selectAll(), on=Seq("cc_num")) .join(merchantDetailsFg.selectAll())) val featureView = featureStore.createFeatureView() .name("credit_card_fraud") .query(selectedFeatures) .build(); // retrieve the query back from the feature view val featureView = fs.getFeatureView(“credit_card_fraud”, 1) val query = featureView.getQuery() ``` If a data scientist wants to modify a new feature that is not available in the feature store, she can write code to compute the new feature (using existing features or external data) and ingest the new feature values into the feature store. If the new feature is based solely on existing feature values in the Feature Store, we call it a derived feature. The same Hopsworks APIs can be used to compute derived features as well as features using external data sources. ## The Query Abstraction Most operations performed on `FeatureGroup` metadata objects will return a `Query` with the applied operation. ### Examples Selecting features from a feature group is a lazy operation, returning a query with the selected features only: === "Python" ```python credit_card_transactions_fg = fs.get_feature_group("credit_card_transactions") # Returns Query selected_features = credit_card_transactions_fg.select( ["amount", "latitude", "longitude"] ) ``` === "Scala" ```scala val creditCardTransactionsFg = fs.getFeatureGroup("credit_card_transactions") # Returns Query val selectedFeatures = creditCardTransactionsFg.select(Seq("amount", "latitude", "longitude")) ``` #### Join Similarly, joins return query objects. The simplest join in one where we join all of the features together from two different feature groups without specifying a join key - `Hopsworks` will infer the join key as a common primary key between the two feature groups. By default, Hopsworks will use the maximal matching subset of the primary keys of the two feature groups as joining key(s), if not specified otherwise. === "Python" ```python # Returns Query selected_features = credit_card_transactions_fg.join(account_details_fg) ``` === "Scala" ```scala // Returns Query val selectedFeatures = creditCardTransactionsFg.join(accountDetailsFg) ``` More complex joins are possible by selecting subsets of features from the joined feature groups and by specifying a join key and type. Possible join types are "inner", "left" or "right". By default`join_type` is `"left". Furthermore, it is possible to specify different features for the join key of the left and right feature group. The join key lists should contain the names of the features to join on. === "Python" ```python selected_features = ( credit_card_transactions_fg.select_all() .join(account_details_fg.select_all(), on=["cc_num"]) .join( merchant_details_fg.select_all(), left_on=["merchant_id"], right_on=["id"], join_type="inner", ) ) ``` === "Scala" ```scala val selectedFeatures = (creditCardTransactionsFg.selectAll() .join(accountDetailsFg.selectAll(), Seq("cc_num")) .join(merchantDetailsFg.selectAll(), Seq("merchant_id"), Seq("id"), "inner")) ``` !!! warning If there is feature name clash in the query then prefixes will be automatically generated and applied. Generated prefix is feature group alias in the query (e.g., fg1, fg2). Prefix is applied to the right feature group of the query. ### Data modeling in Hopsworks Since v4.0 Hopsworks Feature selection API supports both Star and Snowflake Schema data models. #### Star schema data model When choosing Star Schema data model all tables are children of the parent (the left most) feature group, which has all foreign keys for its child feature groups. --8<-- "user_guides/fs/feature_view/query/star-schema.html" === "Python" ```python selected_features = credit_card_transactions.select_all() .join(aggregated_cc_transactions.select_all()) .join(account_details.select_all()) .join(merchant_details.select_all()) .join(cc_issuer_details.select_all()) ``` In online inference, when you want to retrieve features in your online model, you have to provide all foreign key values, known as the serving_keys, from the parent feature group to retrieve your precomputed feature values using the feature view. === "Python" ```python feature vector = feature_view.get_feature_vector({ ‘cc_num’: “1234 5555 3333 8888”, ‘issuer_id’: 20440455, ‘merchant_id’: 44208484, ‘account_id’: 84403331 }) ``` #### Snowflake schema Hopsworks also provides the possibility to define a feature view that consists of a nested tree of children (to up to a depth of 20) from the root (left most) feature group. This is called Snowflake Schema data model where you need to build nested tables (subtrees) using joins, and then join the subtrees to their parents iteratively until you reach the root node (the leftmost feature group in the feature selection): --8<-- "user_guides/fs/feature_view/query/snowflake-schema.html" === "Python" ```python nested_selection = aggregated_cc_transactions.select_all() .join(account_details.select_all()) .join(cc_issuer_details.select_all()) selected_features = credit_card_transactions.select_all() .join(nested_selection) .join(merchant_details.select_all()) ``` Now, you have the benefit that in online inference you only need to pass two serving key values (the foreign keys of the leftmost feature group) to retrieve the precomputed features: === "Python" ```python feature vector = feature_view.get_feature_vector({ ‘cc_num’: “1234 5555 3333 8888”, ‘merchant_id’: 44208484, }) ``` #### Filter In the same way as joins, applying filters to feature groups creates a query with the applied filter. Filters are constructed with Python Operators `==`, `>=`, `<=`, `!=`, `>`, `<` and additionally with the methods `isin` and `like`. Bitwise Operators `&` and `|` are used to construct conjunctions. For the Scala part of the API, equivalent methods are available in the `Feature` and `Filter` classes. === "Python" ```python filtered_credit_card_transactions = credit_card_transactions_fg.filter( credit_card_transactions_fg.category == "Grocery" ) ``` === "Scala" ```scala val filteredCreditCardTransactions = creditCardTransactionsFg.filter(creditCardTransactionsFg.getFeature("category").eq("Grocery")) ``` Filters are fully compatible with joins: === "Python" ```python selected_features = ( credit_card_transactions_fg.select_all() .join(account_details_fg.select_all(), on=["cc_num"]) .join( merchant_details_fg.select_all(), left_on=["merchant_id"], right_on=["id"], ) .filter( (credit_card_transactions_fg.category == "Grocery") | (credit_card_transactions_fg.category == "Restaurant/Cafeteria") ) ) ``` === "Scala" ```scala val selectedFeatures = (creditCardTransactionsFg.selectAll() .join(accountDetailsFg.selectAll(), Seq("cc_num")) .join(merchantDetailsFg.selectAll(), Seq("merchant_id"), Seq("id"), "left") .filter(creditCardTransactionsFg.getFeature("category").eq("Grocery").or(creditCardTransactionsFg.getFeature("category").eq("Restaurant/Cafeteria")))) ``` The filters can be applied at any point of the query: === "Python" ```python selected_features = ( credit_card_transactions_fg.select_all() .join( accountDetails_fg.select_all().filter( accountDetails_fg.avg_temp >= 22 ), on=["cc_num"], ) .join( merchant_details_fg.select_all(), left_on=["merchant_id"], right_on=["id"], ) .filter(credit_card_transactions_fg.category == "Grocery") ) ``` === "Scala" ```scala val selectedFeatures = (creditCardTransactionsFg.selectAll() .join(accountDetailsFg.selectAll().filter(accountDetailsFg.getFeature("avg_temp").ge(22)), Seq("cc_num")) .join(merchantDetailsFg.selectAll(), Seq("merchant_id"), Seq("id"), "left") .filter(creditCardTransactionsFg.getFeature("category").eq("Grocery"))) ``` #### Joins and/or Filters on feature view query The query retrieved from a feature view can be extended with new joins and/or new filters. However, this operation will not update the metadata and persist the updated query of the feature view itself. This query can then be used to create a new feature view. === "Python" ```python fs = ... merchant_details_fg = fs.get_feature_group(name="merchant_details", version=1) credit_card_transactions_fg = fs.get_feature_group(name="credit_card_transactions", version=1) feature_view = fs.get_feature_view(“credit_card_fraud”, version=1) feature_view.query \ .join(merchant_details_fg.select_all()) \ .filter(credit_card_transactions_fg.category == "Cash Withdrawal") ``` === "Scala" ```scala val fs = ... val merchantDetailsFg = fs.getFeatureGroup("merchant_details", 1) val creditCardTransactionsFg = fs.getFeatureGroup("credit_card_transactions", 1) val featureView = fs.getFeatureView(“credit_card_fraud”, 1) featureView.getQuery() .join(merchantDetailsFg.selectAll()) .filter(creditCardTransactionsFg.getFeature("category").eq("Cash Withdrawal")) ``` !!! warning Every join/filter operation applied to an existing feature view query instance will update its state and accumulate. To successfully apply new join/filter logic it is recommended to refresh the query instance by re-fetching the feature view: === "Python" ```python fs = ... merchant_details_fg = fs.get_feature_group(name="merchant_details", version=1) account_details_fg = fs.get_feature_group(name="account_details", version=1) credit_card_transactions_fg = fs.get_feature_group(name="credit_card_transactions", version=1) # fetch new feature view and its query instance feature_view = fs.get_feature_view(“credit_card_fraud”, version=1) # apply join/filter logic based on purchase type feature_view.query.join(merchant_details_fg.select_all()) \ .filter(credit_card_transactions_fg.category == "Cash Withdrawal") # to apply new logic independent of purchase type from above # re-fetch new feature view and its query instance feature_view = fs.get_feature_view(“credit_card_fraud”, version=1) # apply new join/filter logic based on account details feature_view.query.join(merchant_details_fg.select_all()) \ .filter(account_details_fg.gender == "F") ``` === "Scala" ```scala fs = ... merchantDetailsFg = fs.getFeatureGroup("merchant_details", 1) accountDetailsFg = fs.getFeatureGroup("account_details", 1) creditCardTransactionsFg = fs.getFeatureGroup("credit_card_transactions", 1) // fetch new feature view and its query instance val featureView = fs.getFeatureView(“credit_card_fraud”, version=1) // apply join/filter logic based on purchase type featureView.getQuery.join(merchantDetailsFg.selectAll()) .filter(creditCardTransactionsFg.getFeature("category").eq("Cash Withdrawal")) // to apply new logic independent of purchase type from above // re-fetch new feature view and its query instance val featureView = fs.getFeatureView(“credit_card_fraud”, 1) // apply new join/filter logic based on account details featureView.getQuery.join(merchantDetailsFg.selectAll()) .filter(accountDetailsFg.getFeature("gender").eq("F")) ``` ================================================================================ # Helper Columns Source: https://docs.hopsworks.ai/latest/user_guides/fs/feature_view/helper-columns/ # Helper columns Hopsworks Feature Store provides a functionality to define two types of helper columns `inference_helper_columns` and `training_helper_columns` for [feature views](./overview.md). !!! note Both inference and training helper column name(s) must be part of the `Query` object. If helper column name(s) belong to feature group that is part of a `Join` with `prefix` defined, then this prefix needs to prepended to the original column name when defining helper column list. ## Inference Helper columns `inference_helper_columns` are a list of feature names that are not used for training the model itself but are used for extra information during online or batch inference. For example, computing an [on-demand feature](../../../concepts/fs/feature_group/on_demand_feature.md) such as `days_valid` (days left that a credit card is valid at the time of the transaction) in a credit card fraud detection system. The feature `days_valid` will be computed using the credit card expiry date that needs to be fetched from the feature store and compared to the transaction date that the transaction is performed on (`days_valid` = `expiry_date` - `current_date`). In this use case `expiry_date` is an inference helper column. It is not used for training but is necessary for computing the [on-demand feature](../../../concepts/fs/feature_group/on_demand_feature.md)`days_valid` feature. !!! example "Define inference columns for feature views." === "Python" ```python # define query object query = label_fg.select("fraud_label").join( trans_fg.select(["amount", "days_valid", "expiry_date", "category"]) ) # define feature view with helper columns feature_view = fs.get_or_create_feature_view( name="fv_with_helper_col", version=1, query=query, labels=["fraud_label"], transformation_functions=transformation_functions, inference_helper_columns=["expiry_date"], ) ``` ### Inference Data Retrieval When retrieving data for model inference, helper columns will be omitted. However, they can be optionally fetched with inference or training data. #### Batch inference !!! example "Fetch inference helper column values and compute on-demand features during batch inference." === "Python" ```python # import feature functions from feature_functions import time_delta # Fetch feature view object feature_view = fs.get_feature_view( name="fv_with_helper_col", version=1, ) # Fetch feature data for batch inference with helper columns df = feature_view.get_batch_data( start_time=start_time, end_time=end_time, inference_helpers=True, event_time=True, ) # compute location delta df["days_valid"] = df.apply( lambda row: time_delta(row["expiry_date"], row["transaction_date"]), axis=1 ) # prepare datatame for prediction df = df[ [ f.name for f in feature_view.features if not ( f.label or f.inference_helper_column or f.training_helper_column ) ] ] ``` #### Online inference !!! example "Fetch inference helper column values and compute on-demand features during online inference." === "Python" ```python from feature_functions import time_delta # Fetch feature view object feature_view = fs.get_feature_view( name="fv_with_helper_col", version=1, ) # Fetch feature data for batch inference without helper columns df_without_inference_helpers = feature_view.get_batch_data() # Fetch feature data for batch inference with helper columns df_with_inference_helpers = feature_view.get_batch_data(inference_helpers=True) # here cc_num, longitude and latitude are provided as parameters to the application cc_num = ... transaction_date = ... # get previous transaction location of this credit card inference_helper = feature_view.get_inference_helper( {"cc_num": cc_num}, return_type="dict" ) # compute location delta days_valid = time_delta(transaction_date, inference_helper["expiry_date"]) # Now get assembled feature vector for prediction feature_vector = feature_view.get_feature_vector( {"cc_num": cc_num}, passed_features={"days_valid": days_valid}, ) ``` ## Training Helper columns `training_helper_columns` are a list of feature names that are not the part of the model schema itself but are used during training for the extra information. For example one might want to use feature like `category` of the purchased product to assign different weights. !!! example "Define training helper columns for feature views." === "Python" ```python # define query object query = label_fg.select("fraud_label").join( trans_fg.select(["amount", "days_valid", "expiry_date", "category"]) ) # define feature view with helper columns feature_view = fs.get_or_create_feature_view( name="fv_with_helper_col", version=1, query=query, labels=["fraud_label"], transformation_functions=transformation_functions, training_helper_columns=["category"], ) ``` ### Training Data Retrieval When retrieving training data helper columns will be omitted. However, they can be optionally fetched. !!! example "Fetch training data with or without inference helper column values." === "Python" ```python # import feature functions from feature_functions import location_delta, time_delta # Fetch feature view object feature_view = fs.get_feature_view( name="fv_with_helper_col", version=1, ) # Create and training data with training helper columns TEST_SIZE = 0.2 X_train, X_test, y_train, y_test = feature_view.train_test_split( description="transactions fraud training dataset", test_size=TEST_SIZE, training_helper_columns=True, ) # Get existing training data with training helper columns X_train, X_test, y_train, y_test = feature_view.get_train_test_split( training_dataset_version=1, training_helper_columns=True ) ``` !!! note To use helper columns with materialized training dataset it needs to be created with `training_helper_columns=True`. ================================================================================ # Model-Dependent Transformation Functions Source: https://docs.hopsworks.ai/latest/user_guides/fs/feature_view/model-dependent-transformations/ # Model Dependent Transformation Functions [Model-dependent transformations](https://www.hopsworks.ai/dictionary/model-dependent-transformations) transform feature data for a specific model. Feature encoding is one example of such a transformations. Feature encoding is parameterized by statistics from the training dataset, and, as such, many model-dependent transformations require the training dataset statistics as a parameter. Hopsworks enhances the robustness of AI pipelines by preventing [training-inference skew](https://www.hopsworks.ai/dictionary/training-inference-skew) by ensuring that the same model-dependent transformations and statistical parameters are used during both training dataset generation and online inference. Additionally, Hopsworks offers built-in model-dependent transformation functions, such as `min_max_scaler`, `standard_scaler`, `robust_scaler`, `label_encoder`, and `one_hot_encoder`, which can be easily imported and declaratively applied to features in a feature view. ## Model Dependent Transformation Function Creation Hopsworks allows you to create a model-dependent transformation function by attaching a [transformation function](../transformation_functions.md) to a feature view. The attached transformation function can be a simple function that takes one feature as input and outputs the transformed feature data. For example, in the case of min-max scaling a numerical feature, you will have a number as input parameter to the transformation function and a number as output. However, in the case of one-hot encoding a categorical variable, you will have a string as input and an array of 1s and 0s and output. You can also have transformation functions that take multiple features as input and produce one or more values as output. That is, transformation functions can be one-to-one, one-to-many, many-to-one, or many-to-many. Each model-dependent transformation function can map specific features to its arguments by explicitly providing their names as arguments to the transformation function. If no feature names are provided, the transformation function will default to using features from the feature view that match the name of the transformation function's argument. Hopsworks by default generates default names of transformed features output by a model-dependent transformation function. The generated names follows a naming convention structured as `functionName_features_outputColumnNumber` if the transformation function outputs multiple columns and `functionName_features` if the transformation function outputs one column. For instance, for the function named `add_one_multiple` that outputs multiple columns in the example given below, produces output columns that would be labeled as  `add_one_multiple_feature1_feature2_feature3_0`,  `add_one_multiple_feature1_feature2_feature3_1` and  `add_one_multiple_feature1_feature2_feature3_2`. The function named `add_two` that outputs a single column in the example given below, produces a single output column names as `add_two_feature`. Additionally, Hopsworks also allows users to specify custom names for transformed feature using the [`alias`](../transformation_functions.md#specifying-output-features-names-for-transformation-functions) function. !!! example "Creating model-dependent transformation functions" === "Python" ```python # Defining a many to many transformation function. @udf(return_type=[int, int, int], drop=["feature1", "feature3"]) def add_one_multiple(feature1, feature2, feature3): return pd.DataFrame( { "add_one_feature1": feature1 + 1, "add_one_feature2": feature2 + 1, "add_one_feature3": feature3 + 1, } ) # Defining a one to one transformation function. @udf(return_type=int) def add_two(feature): return feature + 2 # Creating model-dependent transformations by attaching transformation functions to feature views. feature_view = fs.create_feature_view( name="transactions_view", query=query, labels=["fraud_label"], transformation_functions=[add_two, add_one_multiple], ) ``` ### Specifying input features The features to be used by a model-dependent transformation function can be specified by providing the feature names (from the feature view / feature group) as input to the transformation functions. !!! example "Specifying input features to be passed to a model-dependent transformation function" === "Python" ```python feature_view = fs.create_feature_view( name="transactions_view", query=query, labels=["fraud_label"], transformation_functions=[ add_two("feature_1"), add_two("feature_2"), add_one_multiple("feature_5", "feature_6", "feature_7"), ], ) ``` ### Using built-in transformations Built-in transformation functions are attached in the same way. The only difference is that they can either be retrieved from the Hopsworks or imported from the `hopsworks` module. !!! example "Creating model-dependent transformation using built-in transformation functions retrieved from Hopsworks" === "Python" ```python min_max_scaler = fs.get_transformation_function(name="min_max_scaler") standard_scaler = fs.get_transformation_function(name="standard_scaler") robust_scaler = fs.get_transformation_function(name="robust_scaler") label_encoder = fs.get_transformation_function(name="label_encoder") feature_view = fs.create_feature_view( name="transactions_view", query=query, labels=["fraud_label"], transformation_functions=[ label_encoder("category"), robust_scaler("amount"), min_max_scaler("loc_delta"), standard_scaler("age_at_transaction"), ], ) ``` To attach built-in transformation functions from the `hopsworks` module they can be directly imported into the code from `hopsworks.builtin_transformations`. !!! example "Creating model-dependent transformation using built-in transformation functions imported from hopsworks" === "Python" ```python from hopsworks.hsfs.builtin_transformations import ( label_encoder, min_max_scaler, robust_scaler, standard_scaler, ) feature_view = fs.create_feature_view( name="transactions_view", query=query, labels=["fraud_label"], transformation_functions=[ label_encoder("category"), robust_scaler("amount"), min_max_scaler("loc_delta"), standard_scaler("age_at_transaction"), ], ) ``` ## Using Model Dependent Transformations Model-dependent transformations attached to a feature view are automatically applied when you [create training data](./training-data.md#creation), [read training data](./training-data.md#read-training-data), [read batch inference data](./batch-data.md#creation-with-transformation), or [get feature vectors](./feature-vectors.md#retrieval-with-transformation). The generated data includes untransformed features, on-demand features, if any, and the transformed features. The transformed features are organized by their output column names in alphabetical order and are positioned after the untransformed and on-demand features. Model-dependent transformation functions can also be manually applied to a feature vector using the `transform` function. !!! example "Manually applying model-dependent transformations during online inference" === "Python" ```python # Initialize the feature view with the correct training dataset version used for model-dependent transformations fv.init_serving(training_dataset_version) # Get untransformed feature Vector feature_vector = fv.get_feature_vector( entry={"index": 10}, transform=False, return_type="pandas" ) # Apply Model Dependent transformations encoded_feature_vector = fv.transform(feature_vector) ``` ### Retrieving untransformed feature vector and batch inference data The `get_feature_vector`, `get_feature_vectors`, and `get_batch_data` methods can return untransformed feature vectors and batch data without applying model-dependent transformations while still including on-demand features. To achieve this, set the `transform` parameter to False. !!! example "Returning untransformed feature vectors and batch data." === "Python" ```python # Fetching untransformed feature vector. untransformed_feature_vector = feature_view.get_feature_vector( entry={"id": 1}, transform=False ) # Fetching untransformed feature vectors. untransformed_feature_vectors = feature_view.get_feature_vectors( entry=[{"id": 1}, {"id": 2}], transform=False ) # Fetching untransformed batch data. untransformed_batch_data = feature_view.get_batch_data(transform=False) ``` ## Chaining Model-Dependent Transformations A model-dependent transformation (MDT) can consume another MDT's output as its input. The DAG is resolved automatically at execution time, so producers always run before consumers. !!! example "Chaining two increments and a sum" === "Python" ```python from hopsworks import udf @udf(int) def add_one(col): return col + 1 @udf(int) def add(a, b): return a + b fv = fs.create_feature_view( name="chained_mdt_fv", query=fg.select_all(), transformation_functions=[ add_one("data1").alias("data1_plus_one"), add_one("data2").alias("data2_plus_one"), add("data1_plus_one", "data2_plus_one").alias("sum_plus_two"), ], version=1, ) ``` ### Statistics over chained transformations Statistics-based transformations participate in chains like any other transformation. A transformation that requires statistics on another transformation's output, such as a min-max scaler applied to an imputed column, is fit on that intermediate output rather than on the raw feature. During training dataset creation the statistics are computed in dependency order on the train split, each transformation executes exactly once, and the fitted statistics are persisted so that online serving applies the same values. See [Transformation Functions Performance Tuning][transformation-functions-performance-tuning] for `n_processes` semantics on chained DAGs. ================================================================================ # Spines Source: https://docs.hopsworks.ai/latest/user_guides/fs/feature_view/spine-query/ # Using Spines In this section we will illustrate how to use a [Spine Group](../../../concepts/fs/feature_group/spine_group.md) instead of a regular Feature Group for performing point-in-time joins when reading batch data for inference or when creating training datasets. ## Prerequisites 1. Make sure you have read the [concept section about spines](../../../concepts/fs/feature_group/spine_group.md) in feature and inference pipelines. 2. Make sure you have gone through the [Spine Group creation guide](../feature_group/create_spine.md). 3. Make sure you understand the [concept of feature views](../../../concepts/fs/feature_view/fv_overview.md) and how to create them using the [query abstraction](../feature_view/query.md) ## Feature View with a Spine Group ### Step 1: Query Definition The first step before creating a Feature View, is to construct the query by selecting the label and features which are needed: ```python # Select features for training data. ds_query = trans_fg.select(["fraud_label"]).join( window_aggs_fg.select_except(["cc_num"]), on="cc_num" ) ds_query.show(5) ``` Similarly you can construct the query using a previously created spine equivalent. However, there are two thing to note: 1. **If you want to use the query for a feature view to be used for online serving, you can only select the "label" or target feature from the spine.** 2. **Spine groups can only be used on the left side of the join.** Think of the left side of the join as the base set of entities that should be included in you batch of data or training dataset, which we enrich with the relevant and point-in-time correct feature values. ```python trans_spine = fs.get_or_create_spine_group( name="spine_transactions", version=1, description="Transaction data", primary_key=["cc_num"], event_time="datetime", dataframe=trans_df, ) # Select features for training data. ds_query_spine = trans_spine.select(["fraud_label"]).join( window_aggs_fg.select_except(["cc_num"]), on="cc_num" ) ``` Calling the `show()` or `read()` method of this query object will use the spine dataframe included in the Spine Group object to perform the join. ```python ds_query_spine.show(10) ``` ### Step 2: Feature View Creation With the above defined query, we can continue to create the Feature View in the same way we would do it also without a spine: ```python feature_view_spine = fs.get_or_create_feature_view( name="transactions_view_spine", query=ds_query_spine, version=1, labels=["fraud_label"], ) ``` ### Step 3: Training Dataset Creation With the regular feature view, the labels are fetched from the feature store, but with the feature view created with a spine, you need to provide the dataframe. Here you have the chance to pass a different set of entities to generate the training dataset. ```python X_train, X_test, y_train, y_test = feature_view_spine.train_test_split( 0.2, spine=new_entities_df ) X_train.show() ``` ### Step 4: Retrieving New Batches Inference Data You can now use the offline and online API of the feature stores to read features for inference. Similarly to training dataset creation, every time you read up a new batch of data, you can pass a different spine dataframe. ```python feature_view_spine.get_batch_data(spine=scoring_spine_df).show() ``` ### Step 5: Online Feature Lookup For the online lookup, the label is not required, therefore it was important to only select label from the left spine group, so that we don't need to provide a spine for online serving: ```python # Note: no spine needs to be passed feature_view.get_feature_vector({"cc_num": 4473593503484549}) ``` ## Replacing a Regular Feature Group with a Spine at Serving Time In the case where you create a feature view with a regular feature group, but you would like to retrieve batch inference data using IDs (primary key values), you can use a spine to replace the left feature group. To do this, you can pass the Spine Group instead of a dataframe. ```python # Note: here feature_view was created with regular feature groups only # and trans_spine is of type SpineGroup instead of a dataframe feature_view.get_batch_data(spine=trans_spine).show() ``` ================================================================================ # Feature Monitoring Source: https://docs.hopsworks.ai/latest/user_guides/fs/feature_view/feature_monitoring/ # Feature Monitoring for Feature Views Feature Monitoring complements the Hopsworks data validation capabilities for Feature Group data by allowing you to monitor your data once they have been ingested into the Feature Store. Hopsworks feature monitoring is centered around two functionalities: **scheduled statistics** and **statistics comparison**. Before continuing with this guide, see the [Feature monitoring guide](../feature_monitoring/index.md) to learn more about how feature monitoring works, and get familiar with the different use cases of feature monitoring for Feature Views described in the **Use cases** sections of the [Scheduled statistics guide](../feature_monitoring/scheduled_statistics.md#use-cases) and [Statistics comparison guide](../feature_monitoring/statistics_comparison.md#use-cases). !!! warning "Limited UI support" Currently, feature monitoring can only be configured using the [Hopsworks Python library](https://pypi.org/project/hopsworks). However, you can enable/disable a feature monitoring configuration or trigger the statistics comparison manually from the UI. ## Code In this section, we show you how to set up feature monitoring on a Feature View using the ==Hopsworks Python library==. Alternatively, you can get started quickly by running our [tutorial for feature monitoring](https://github.com/logicalclocks/hopsworks-tutorials/blob/master/api_examples/feature_monitoring.ipynb). !!! info "Prerequisites" - A Hopsworks project. If you don't have one yet, go to [run.hopsworks.ai](https://run.hopsworks.ai), sign up with your email and create your first project. - An API key, which you can get from "Account Settings" on [run.hopsworks.ai](https://run.hopsworks.ai). - The [Hopsworks Python library](https://pypi.org/project/hopsworks) installed in your client. See the [installation guide](../../client_installation/index.md). - A Feature View and a Training Dataset. ### Step 1: Connect to Hopsworks Connect the client running your notebook to Hopsworks. You will be prompted to paste your API key to connect the notebook to your project. === "Python" ```python import hopsworks project = hopsworks.login() fs = project.get_feature_store() ``` See the API reference for [`hopsworks.login`][hopsworks.login] and [`Project.get_feature_store`][hopsworks_common.project.Project.get_feature_store]. ### Step 2: Get or create a Feature View Feature monitoring can be enabled on already created Feature Views. We suggest you read the [Feature View](../../../concepts/fs/feature_view/fv_overview.md) concept page and familiarize yourself with the APIs to [create a feature view](overview.md) using the [query abstraction](query.md). === "Python" ```python # Retrieve an existing feature view trans_fv = fs.get_feature_view("trans_fv", version=1) # Or, create a new feature view query = trans_fg.select(["fraud_label", "amount", "cc_num"]) trans_fv = fs.create_feature_view( name="trans_fv", version=1, query=query, labels=["fraud_label"], ) ``` See the API reference for [`FeatureStore.get_feature_view`][hsfs.feature_store.FeatureStore.get_feature_view] and [`FeatureStore.create_feature_view`][hsfs.feature_store.FeatureStore.create_feature_view]. ### Step 3: Get or create a Training Dataset A Training Dataset can be used later as a reference window to compare against (see Step 6). === "Python" ```python # Create a training dataset with train and test splits _, _ = trans_fv.create_train_validation_test_split( description="transactions fraud batch training dataset", data_format="csv", validation_size=0.2, test_size=0.1, ) ``` See the API reference for [`FeatureView.create_train_validation_test_split`][hsfs.feature_view.FeatureView.create_train_validation_test_split]. ### Step 4: Create a monitoring configuration Start a new configuration on the Feature View. Use `create_scheduled_statistics` to only compute statistics on a schedule, or `create_feature_monitoring` to also compare them against a reference. === "Scheduled statistics" ```python # compute statistics on one or more features on a schedule fm_monitoring_config = trans_fv.create_scheduled_statistics( name="trans_fv_all_features_monitoring", description="Compute statistics on the Feature View data on a daily basis", feature_names=["amount"], # omit to monitor all features ) ``` === "Statistics comparison" ```python # the feature to compare is selected later in # compare_on / compare_on_distribution (Step 7.A / 7.B), not here fm_monitoring_config = trans_fv.create_feature_monitoring( name="trans_fv_amount_monitoring", description="Compute and compare descriptive statistics on the Feature View data on a daily basis", ) ``` See the API reference for [`FeatureView.create_scheduled_statistics`][hsfs.feature_view.FeatureView.create_scheduled_statistics] and [`FeatureView.create_feature_monitoring`][hsfs.feature_view.FeatureView.create_feature_monitoring]. !!! info "Custom schedule" By default, the computation of statistics is scheduled to run endlessly, every day at 12PM. You can modify the default schedule by adjusting the `cron_expression`, `start_date_time` and `end_date_time` parameters (e.g., `cron_expression="0 0 12 ? * MON *"` for a weekly run). To compute statistics on only a subset of the feature data, use the `row_percentage` parameter of `with_detection_window` (see Step 5). ### Step 5: (Optional) Define a detection window By default, the detection window is an _expanding window_ covering the whole Feature Group data. You can define a different detection window using the `window_length` and `time_offset` parameters of the `with_detection_window` method. Additionally, you can specify the percentage of feature data on which statistics will be computed using the `row_percentage` parameter. === "Python" ```python fm_monitoring_config.with_detection_window( window_length="1w", # data ingested during one week time_offset="1w", # starting from last week row_percentage=0.8, # use 80% of the data ) ``` See the API reference for [`FeatureMonitoringConfig.with_detection_window`][hsfs.core.feature_monitoring_config.FeatureMonitoringConfig.with_detection_window]. #### Time basis of the windows Rolling windows select rows by the event-time feature of the Feature View's left Feature Group when it declares one, and by commit time otherwise. The `event_time` parameter of `create_scheduled_statistics` and `create_feature_monitoring` overrides that default for the whole configuration, detection and reference windows alike. Pass the name, prefix included, of a timestamp, date or epoch feature that the Feature View selects, or `False` to select rows by commit time. The left Feature Group's event-time feature is also accepted when the Feature View does not select it. With event time the joined Feature Groups contribute their current rows, whereas with commit time the same commit interval is applied to every Feature Group in the query. === "Python" ```python # windows over a time feature of a joined Feature Group fm_monitoring_config = trans_fv.create_feature_monitoring( name="trans_fv_amount_monitoring_by_event_time", event_time="datetime", ) # windows over the time the rows were written (commit time) fm_monitoring_config = trans_fv.create_feature_monitoring( name="trans_fv_amount_monitoring_by_commit", event_time=False, ) ``` See [Time basis](../feature_monitoring/scheduled_statistics.md#time-basis) for how the two bases differ. ### Step 6: (Optional) Define a reference window When setting up feature monitoring for a Feature View, the reference can be either a reference window of feature data or a training dataset. !!! tip "Basis for Model Monitoring" Using a training dataset as the reference is the basis for [Model Monitoring](../../mlops/model_monitoring/index.md), where a model's production inference data is compared against the distribution of its training dataset. === "Python" ```python # compare statistics against a reference window fm_monitoring_config.with_reference_window( window_length="1w", # data ingested during one week time_offset="2w", # starting from two weeks ago row_percentage=0.8, # use 80% of the data ) # or a training dataset fm_monitoring_config.with_reference_training_dataset( training_dataset_version=1, # use the training dataset used to train your production model ) ``` See the API reference for [`FeatureMonitoringConfig.with_reference_window`][hsfs.core.feature_monitoring_config.FeatureMonitoringConfig.with_reference_window] and [`FeatureMonitoringConfig.with_reference_training_dataset`][hsfs.core.feature_monitoring_config.FeatureMonitoringConfig.with_reference_training_dataset]. !!! info "Comparing against a specific value" Instead of a reference window or training dataset, you can compare the detection statistics against a fixed reference value (i.e., a window of size 1). In that case, skip this step and pass the `specific_value` parameter to `compare_on` in Step 7. ### Step 7.A: (Optional) Compare on a scalar metric In order to compare detection and reference statistics, you need to provide the criteria for such comparison. First, you select the feature and the metric to consider in the comparison using the `feature_name` and `metric` parameters. Then, you can define a relative or absolute threshold using the `threshold` and `relative` parameters. === "Python" ```python fm_monitoring_config.compare_on( feature_name="amount", # the feature to compare metric="mean", threshold=0.2, # a relative change over 20% is considered anomalous relative=True, # relative or absolute change strict=False, # strict or relaxed comparison ) ``` See the API reference for [`FeatureMonitoringConfig.compare_on`][hsfs.core.feature_monitoring_config.FeatureMonitoringConfig.compare_on]. !!! info "Difference values and thresholds" For more information about the computation of difference values and the comparison against threshold bounds see the [Comparison criteria section](../feature_monitoring/statistics_comparison.md#comparison-criteria) in the Statistics comparison guide. ### Step 7.B: (Optional) Compare on the whole distribution Alternatively, instead of a single scalar metric, you can detect drift in the shape of a feature's distribution using `compare_on_distribution`. Select a distribution distance metric (e.g., `PSI`) and a threshold. A reference window or training dataset (Step 6) is required for distribution comparison. === "Python" ```python fm_monitoring_config.compare_on_distribution( feature_name="amount", # the feature to compare metric="PSI", threshold=0.2, # a distance above 0.2 is considered a significant shift ) ``` See the API reference for [`FeatureMonitoringConfig.compare_on_distribution`][hsfs.core.feature_monitoring_config.FeatureMonitoringConfig.compare_on_distribution]. !!! tip "More distribution options" See the [Distribution comparison guide](../feature_monitoring/distribution_comparison.md) for the full list of metrics and binning strategies. ### Step 8: Save the configuration Finally, you can save your feature monitoring configuration by calling the `save` method. Once the configuration is saved, the schedule for the statistics computation and comparison will be activated automatically. === "Python" ```python fm_monitoring_config.save() ``` See the API reference for [`FeatureMonitoringConfig.save`][hsfs.core.feature_monitoring_config.FeatureMonitoringConfig.save]. ### Step 9: Retrieve configurations and history Once saved, you can retrieve your feature monitoring configurations and the results of past executions directly from the Feature View. === "Python" ```python # fetch all configurations attached to the feature view configs = trans_fv.get_feature_monitoring_configs() # or a single configuration by name config = trans_fv.get_feature_monitoring_configs(name="trans_fv_amount_monitoring") # fetch the history of monitoring results (with computed statistics) history = trans_fv.get_feature_monitoring_history( config_name="trans_fv_amount_monitoring", with_statistics=True, ) ``` See the API reference for [`FeatureView.get_feature_monitoring_configs`][hsfs.feature_view.FeatureView.get_feature_monitoring_configs] and [`FeatureView.get_feature_monitoring_history`][hsfs.feature_view.FeatureView.get_feature_monitoring_history]. !!! api "API reference" - [`FeatureView`][hsfs.feature_view.FeatureView] - [`create_feature_monitoring`][hsfs.feature_view.FeatureView.create_feature_monitoring] - [`create_scheduled_statistics`][hsfs.feature_view.FeatureView.create_scheduled_statistics] - [`get_feature_monitoring_configs`][hsfs.feature_view.FeatureView.get_feature_monitoring_configs] - [`get_feature_monitoring_history`][hsfs.feature_view.FeatureView.get_feature_monitoring_history] - [`FeatureMonitoringConfig`][hsfs.core.feature_monitoring_config.FeatureMonitoringConfig] - [`with_detection_window`][hsfs.core.feature_monitoring_config.FeatureMonitoringConfig.with_detection_window] - [`with_reference_window`][hsfs.core.feature_monitoring_config.FeatureMonitoringConfig.with_reference_window] - [`with_reference_training_dataset`][hsfs.core.feature_monitoring_config.FeatureMonitoringConfig.with_reference_training_dataset] Browse the full Python API :material-arrow-right: ## Monitor a model in production A Feature View can also monitor the inference data of a model served in production, comparing it against the training dataset the model was trained on. This is the feature-view entry point to [Model Monitoring](../../mlops/model_monitoring/index.md). It targets the feature view's logging feature group, so feature logging must be enabled with `feature_view.enable_logging()`, and filters the detection window by the given model name and version. The windows select inference rows by their `log_time`, the time the prediction was logged. The reference defaults to the training dataset version used to train the model. === "Python" ```python fm_monitoring_config = trans_fv.create_model_monitoring( name="trans_fv_model_monitoring", model_name="my_model", model_version=1, ).with_detection_window( time_offset="1d", window_length="1d", ).with_reference_training_dataset( # omitted -> defaults to the model's training dataset version ).compare_on_distribution( feature_name="amount", metric="PSI", threshold=0.2, ).save() ``` See the API reference for [`FeatureView.create_model_monitoring`][hsfs.feature_view.FeatureView.create_model_monitoring]. !!! info "Explore the API" The [`FeatureMonitoringConfig`][hsfs.core.feature_monitoring_config.FeatureMonitoringConfig] reference documents the full set of available methods, such as enabling or disabling a configuration, triggering it manually, or deleting it. ================================================================================ # Feature Logging Source: https://docs.hopsworks.ai/latest/user_guides/fs/feature_view/feature_logging/ # User Guide: Feature and Prediction Logging with a Feature View Log features and predictions with a feature view, then retrieve them for debugging and monitoring. ## Feature and Prediction Logging After you have trained a model, you can log the features it uses and the predictions with the feature view used to create the training data for the model. You can log transformed features, untransformed features, or both. ### Enabling Feature Logging To enable logging, set `logging_enabled=True` when creating the feature view. One logging feature group stores transformed features, untransformed features, predictions, and logging metadata together. Older feature views can retain separate transformed and untransformed logging groups. The logged features are written to the offline feature store by a materialization job that is created automatically and runs on a schedule. ```python feature_view = fs.create_feature_view("name", query, logging_enabled=True) ``` Alternatively, you can enable logging on an existing feature view by calling `feature_view.enable_logging()`. Also, calling `feature_view.log()` will implicitly enable logging if it has not already been enabled. ### Choosing the Transport A feature view logs through one of two transports, and the layout of its logging feature group follows from the choice. | Transport | Path of a logged row | Readable | | --- | --- | --- | | `realtime` (default) | The deployment posts Arrow batches to its inference logger, which produces them to Kafka; the online store receives them within seconds and the materialization job appends them to the offline store on its schedule | Online at once with `read_log(online=True)` for the group's time to live, offline after materialization | | `job` | The deployment appends Arrow batches to a file buffer on its pod, rotates the buffer on size or age and uploads it to HopsFS; a scheduled commit job appends the uploaded chunks to an offline-only logging group | Offline after the commit job has run | A new feature view names its transport when logging is enabled: ```python feature_view = fs.create_feature_view( "name", query, logging_enabled=True, logging_transport="job" ) feature_view.feature_logging.transport # "job" ``` A feature view that does not log yet names it when logging is enabled, and the transport is read back from the view: ```python feature_view.enable_logging(transport="realtime") feature_view.feature_logging.transport # "realtime" ``` The two cannot be combined on one feature view: enabling the other transport while the view logs is refused. To move a view from one transport to the other, drop its log and recreate the logging group for the new transport with `feature_view.delete_log(transport="job")`. Deployments take the transport from the view; a `DeploymentLoggingConfig` that names a different one is rejected. The `job` transport keeps no online copy, so `read_log(online=True)` is refused for such a view, and a deployment that stops uploads what its buffer holds and starts the commit job before the pod exits. Run `deployment.commit_feature_logs()` or `feature_view.materialize_log()` to commit the uploaded chunks on demand, for example after a replica was killed. ### Choosing the Materialization Interval { #choosing-the-materialization-interval } The materialization job runs every hour or once a day. The platform default applies unless you choose one, at creation or later. ```python feature_view = fs.create_feature_view( "name", query, logging_enabled=True, logging_materialization_interval="day" ) feature_view.enable_logging(materialization_interval="hour") feature_view.set_log_materialization_interval("day") ``` The interval only sets how often logs reach the offline store. Run `feature_view.materialize_log()` to write them on demand between scheduled runs. On the `job` transport the interval schedules the commit job instead. ### Logging Features and Predictions You can log features and predictions by calling `feature_view.log`. The logged features are written periodically to the offline store. If you need it to be available immediately, call `feature_view.materialize_log`. You can log either transformed or/and untransformed features. To get untransformed features, you can specify `transform=False` in `feature_view.get_batch_data` or `feature_view.get_feature_vector(s)`. Inference helper columns are returned along with the untransformed features. If you have On-Demand features as well, call `feature_view.compute_on_demand_features` to get the on demand features before calling `feature_view.log`.To get the transformed features, you can call `feature_view.transform` and pass the untransformed feature with the on-demand feature. Predictions can be optionally provided as one or more columns in the DataFrame containing the features or separately in the `predictions` argument. There must be the same number of prediction columns as there are labels in the feature view. It is required to provide predictions in the `predictions` argument if you provide the features as `list` instead of pandas `dataframe`. The training dataset version will also be logged if you have called either `feature_view.init_serving(...)` or `feature_view.init_batch_scoring(...)` or if the provided model has a training dataset version. The wallclock time of calling `feature_view.log` is automatically logged, enabling filtering by logging time when retrieving logs. #### Example 1: Log Features Only You have a DataFrame of features you want to log. ```python import pandas as pd features = pd.DataFrame( {"feature1": [1.1, 2.2, 3.3], "feature2": [4.4, 5.5, 6.6]} ) # Log features feature_view.log(features) ``` #### Example 2: Log Features, Predictions, and Model You can also log predictions, and optionally the training dataset and the model used for prediction. ```python predictions = pd.DataFrame({"prediction": [0, 1, 0]}) # Log features and predictions feature_view.log( features, predictions=predictions, training_dataset_version=1, model=Model(1, "model", version=1), ) ``` #### Example 3: Log Both Transformed and Untransformed Features ##### Batch Features ```python untransformed_df = fv.get_batch_data(transformed=False) # then apply the transformations after: transformed_df = fv.transform(untransformed_df) # Log untransformed features feature_view.log(untransformed_df) # Log transformed features feature_view.log(transformed_features=transformed_df) ``` ##### Real-time Features ```python untransformed_vector = fv.get_feature_vector({"id": 1}, transform=False) # then apply the transformations after: transformed_vector = fv.transform(untransformed_vector) # Log untransformed features feature_view.log(untransformed_vector) # Log transformed features feature_view.log(transformed_features=transformed_vector) ``` ## Retrieving the Log Timeline To audit and review the feature/prediction logs, you might want to retrieve the timeline of log entries. This helps understand when data was logged and monitor the logs. ### Retrieve Log Timeline A log timeline is the hudi commit timeline of the logging feature group. ```python # Retrieve the latest 10 log entries log_timeline = feature_view.get_log_timeline(limit=10) print(log_timeline) ``` ## Reading Log Entries You may need to read specific log entries for analysis, such as entries within a particular time range or for a specific model version and training dataset version. ### Read all Log Entries Read all log entries for comprehensive analysis. The output will return all values of the same primary keys instead of just the latest value. ```python # Read all log entries log_entries = feature_view.read_log() print(log_entries) ``` ### Read Log Entries within a Time Range Focus on logs within a specific time range. You can specify `start_time` and `end_time` for filtering, but the time columns will not be returned in the DataFrame. You can provide the `start/end_time` as `datetime`, `date`, `int`, or `str` type. Accepted date format are: `%Y-%m-%d`, `%Y-%m-%d %H`, `%Y-%m-%d %H:%M`, `%Y-%m-%d %H:%M:%S`, or `%Y-%m-%d %H:%M:%S.%f` ```python # Read log entries from January 2022 log_entries = feature_view.read_log( start_time="2022-01-01", end_time="2022-01-31" ) print(log_entries) ``` ### Read Log Entries by Training Dataset Version Analyze logs from a particular version of the training dataset. The training dataset version column will be returned in the DataFrame. ```python # Read log entries of training dataset version 1 log_entries = feature_view.read_log(training_dataset_version=1) print(log_entries) ``` ### Read Log Entries by Model in Hopsworks Analyze logs from a particular name and version of the HSML model. The HSML model column will be returned in the DataFrame. ```python # Read log entries of a specific HSML model log_entries = feature_view.read_log(model=Model(1, "model", version=1)) print(log_entries) ``` ### Read Log Entries using a Custom Filter Provide filters which work similarly to the filter method in the `Query` class. The filter should be part of the query in the feature view. ```python # Read log entries where feature1 is greater than 0 log_entries = feature_view.read_log(filter=fg.feature1 > 0) print(log_entries) ``` ## Pausing and Resuming Logging During maintenance or updates, you might need to pause logging to save computation resources. ### Pause Logging Pause the schedule of the materialization job for writing logs to the offline store. ```python # Pause logging feature_view.pause_logging() ``` ### Resume Logging Resume the schedule of the materialization job for writing logs to the offline store. ```python # Resume logging feature_view.resume_logging() ``` ## Materializing Logs Besides the scheduled materialization job, you can materialize logs to the offline store on demand. On the `realtime` transport this reads the rows from Kafka. On the `job` transport this runs the commit job over the chunks that deployments uploaded to HopsFS. This does not pause the scheduled job. Materialization writes all columns of the logging group. The `transformed` selector applies only to older feature views with separate logging groups. ### Materialize Logs Materialize logs and optionally wait for the process to complete. ```python # Materialize logs and wait for completion materialization_result = feature_view.materialize_log(wait=True) ``` ## Monitoring Feature Logging A deployment that logs through the `realtime` transport reports what its inference logger is doing to Prometheus, and the deployment page shows it. Open the deployment and look at the Feature logging card. It shows four panels: rows logged per second by outcome, the time from a post to Kafka's acknowledgement, rows in flight, and posts per second by type and outcome. The Full dashboard link opens the Feature Logging dashboard in Grafana, filtered to the same deployment, which adds in-flight bytes, rejected posts and totals over the selected range. Two of these answer most questions. A non-zero rate of dropped or failed rows means the deployment logs faster than the inference logger can produce, or Kafka is refusing writes; the deployment logs name the reason. Rejected posts mean the batches the predictor builds do not match the logging group's schema, which happens after the feature view changed without a redeploy. For a feature view on the `job` transport the card shows the same rows per second and buffered rows, the upload latency of a buffer segment to HopsFS, the bytes awaiting upload and the chunks uploaded per second; the predictor publishes these itself, and the Full dashboard adds commit job triggers and writer restarts. The card is not shown for a view whose logging still runs through the row path of earlier releases. Those logs are covered by the commit job's or the materialization job's own execution history instead. ## Deleting Logs When log data is no longer needed, you might want to delete it to free up space and maintain data hygiene. This operation deletes the feature groups and recreates new ones. Scheduled materialization job and log timeline are reset as well. Pass `transport="realtime"` or `transport="job"` to recreate the logging group for the other transport. ### Delete Logs Remove all log entries. The `transformed` selector applies only to older feature views with separate logging groups. ```python # Delete all log entries feature_view.delete_log() ``` Restart serving revisions after recreating a logging group so they load its new schema and destination. ================================================================================ # Deployment Source: https://docs.hopsworks.ai/latest/user_guides/fs/feature_view/deployment/ # How To Deploy A Feature View { #feature-view-deployment } ## Introduction In this guide, you will learn how to serve a feature view without a model. A feature view deployment answers a prediction-style request with the transformed feature vector a model would receive. It uses the same request contract, feature lookup, transformations, logging, and monitoring as a model deployment served by the default predictor. See the [Deployment Schema Guide][deployment-schema] for the request contract and the error codes, which are shared with model deployments. Use it to serve features to a model that runs outside Hopsworks, to test transformations online before a model exists, or to give a feature vector API to another team. !!! warning "Serving identity" The deployment looks up features as the project's serving identity, not as the caller. Anyone allowed to call the deployment can obtain the transformed features of any entity the feature view can serve. ## Code ### Step 1: Connect to Hopsworks === "Python" ```python import hopsworks project = hopsworks.login() fs = project.get_feature_store() ``` ### Step 2: Pin a training dataset Model-dependent transformations that need statistics, such as `min_max_scaler`, take them from a training dataset. The deployment uses the training dataset you last read or created in this session, or the one you pass to `deploy()`. === "Python" ```python feature_view = fs.get_feature_view("transactions", version=1) # reading or creating a training dataset records it as the one to serve with X_train, X_test, y_train, y_test = feature_view.train_test_split(test_size=0.2) ``` If the feature view has such a transformation and no training dataset was read or created, `deploy()` refuses and names the transformation, because its statistics cannot be computed. ### Step 3: Deploy the feature view === "Python" ```python deployment = feature_view.deploy( name="transactionsfv", passed_features=["amount"], # features the client sends with each request ) deployment.start(await_running=600) ``` The deployment name defaults to the feature view name and version without special characters. The client publishes the deployment schema before the deployment is created, so `deployment.schema` describes the request immediately: === "Python" ```python deployment.schema.describe() print(deployment.schema.names) # the order of positional rows ``` ### Step 4: Request feature vectors Each row carries the serving keys, the passed features, the request parameters of on-demand transformations, and any extra logging columns. The response carries one transformed vector per row and the column names. === "Python" ```python response = deployment.predict( inputs=[{"cc_num": 4473593503484549, "amount": 12.5}] ) print(response["columns"]) # ["amount_scaled", "age_days", ...] print(response["predictions"]) # [[0.31, -1.2, ...]] ``` Rows can also be arrays in `deployment.schema.names` order. When every stored feature of the view is passed, the schema has no serving keys and the deployment only computes the on-demand features and applies the model-dependent transformations; see [Deployments without lookups][deployment-schema-no-lookup]. Invalid rows are refused before any feature is read; see [Errors][deployment-schema-errors]. ### Step 5: Inspect the deployment === "Python" ```python print(deployment.has_feature_view) # True print(deployment.feature_view_name, deployment.feature_view_version) print(deployment.training_dataset_version) # the pinned version feature_view = deployment.get_feature_view() # the FeatureView object print(deployment.get_model()) # None ``` ## Feature logging When logging is enabled on the feature view, every request is logged with the untransformed and transformed features, the request id, the training dataset version, and the reserved deployment columns `deployment_name`, `deployment_version`, `deployment_schema_id`, and `request_row`, when the logging feature group declares them. The model columns of the log are null, because there is no model. See [Feature logging in the Deployment Schema Guide][deployment-schema-feature-logging] for how requests are logged and for the reserved columns. === "Python" ```python feature_view.enable_logging( extra_log_columns=[ {"name": "deployment_name", "type": "string"}, {"name": "deployment_version", "type": "int"}, {"name": "deployment_schema_id", "type": "string"}, {"name": "request_row", "type": "int"}, ] ) deployment = feature_view.deploy(name="transactionsfv", passed_features=["amount"]) ``` ## Feature monitoring A feature view deployment has no model, so `deployment.create_model_monitoring()` raises. Use `deployment.create_feature_monitoring()`, which attaches a feature monitoring configuration to the logging feature group of the view: === "Python" ```python config = ( deployment.create_feature_monitoring(name="amount_drift") .with_detection_window(time_offset="1d", window_length="1d") .with_reference_window(time_offset="8d", window_length="7d") .compare_on(metric="MEAN", threshold=10.0, feature_name="amount") .save() ) deployment.get_monitoring_configs() ``` A distribution comparison (`compare_on_distribution`) over rolling windows needs KLL statistics on the logging feature group, which it does not keep by default; enable them in the logging feature group's statistics configuration first. Two deployments of the same feature view version log to the same feature group; their rows are told apart by the reserved deployment columns, but a monitoring configuration sees both. Deploy a separate version of the feature view when the statistics of one deployment must not include another's traffic. ## Custom predictor script To post-process the vectors or to change how they are looked up, subclass the default predictor and pass the script to `deploy()`. The script must end with the hand-over to the serving wrapper: === "Python" ```python from hsml.default_predictor import DefaultPredict, run_kserve_wrapper class Predict(DefaultPredict): def model_predict(self, feature_vectors): # a feature view deployment has no model: return the vectors return feature_vectors.round(3) if __name__ == "__main__": run_kserve_wrapper() ``` Backends that support the `SERVING_SCRIPT_KIND=predictor` marker set by `deploy()` start the serving wrapper directly and never run the `__main__` block. Older backends start the script with `python`, and the block hands over to the wrapper. `deploy(script_file=...)` refuses a local script without it; a script already in HopsFS is not checked client-side. ## REST access The deployment answers on the KServe V1 route of the Istio ingress, `/v1/models/:predict`, and through the Hopsworks REST API at `/project//inference/serving/:predict`. See the [REST API Guide][hopsworks-model-serving-rest-api] for authentication and the base URL. `deployment.get_inference_url()` returns the Istio URL, or `None` when the Istio ingress is not configured for external access. Use the Hopsworks REST API path above when it does. ## CLI ```bash hops fv deploy transactions --passed-feature amount hops deployment schema transactionsfv --openapi ``` !!! api "API reference" - [`FeatureView.deploy`][hsfs.feature_view.FeatureView.deploy] - [`Deployment`][hsml.deployment.Deployment] - [`start`][hsml.deployment.Deployment.start] - [`predict`][hsml.deployment.Deployment.predict] - [`create_feature_monitoring`][hsml.deployment.Deployment.create_feature_monitoring] - [`schema`][hsml.deployment.Deployment.schema] - [`training_dataset_version`][hsml.deployment.Deployment.training_dataset_version] - [`DeploymentSchema`][hsml.deployment_schema.DeploymentSchema] - [`describe`][hsml.deployment_schema.DeploymentSchema.describe] Browse the full Python API :material-arrow-right: ================================================================================ # Vector Similarity Search Source: https://docs.hopsworks.ai/latest/user_guides/fs/vector_similarity_search/ ## Introduction Vector similarity search (also called similarity search) is a technique enabling the retrieval of similar items based on their vector embeddings or representations. Its applications range across various domains, from recommendation systems to image similarity and beyond. In Hopsworks, vector similarity search is enabled by extending an online feature group with approximate nearest neighbor search capabilities through a vector database, such as Opensearch. This guide provides a detailed walkthrough on how to leverage Hopsworks for vector similarity search. ## Extending Feature Groups with Similarity Search In Hopsworks, each vector embedding in a feature group is stored in an index within the backing vector database. By default, vector embeddings are stored in the default index for the project (created for every project in Hopsworks), but you have the option to create a new index for a feature group if needed. Creating a separate index per feature group is particularly useful for large volumes of data, ensuring that when a feature group is deleted, its associated index is also removed. For feature groups that use the default project index, the index will only be removed when the project is deleted - not when the feature group is deleted. The index will store all the vector embeddings defined in that feature group, if you have more than one vector embedding in the feature group. In the following example, we explicitly define an index for the feature group: ```aidl from hsfs import embedding # Specify optionally the index in the vector database emb = embedding.EmbeddingIndex(index_name="news_fg") ``` Then, add one or more embedding features to the index. Name and dimension of the embedding features are required for identifying which features should be indexed for k-nearest neighbor (KNN) search. In this example, we get the dimension of the embedding by taking the length of the value of the `embedding_heading` column in the first row of the dataframe `df`. Optionally, you can specify the similarity function among `l2_norm`, `cosine`, and `dot_product`. Refer to [`EmbeddingIndex.add_embedding`][hsfs.embedding.EmbeddingIndex.add_embedding] for the full list of arguments. ```aidl # Add embedding feature to the index emb.add_embedding("embedding_heading", len(df["embedding_heading"][0])) ``` Next, you create a feature group with the `embedding_index` and ingest data to the feature group. When the `embedding_index` is provided, the vector database is used as online feature store. That is, all the features in the feature group are stored **exclusively** in the vector database. The advantage of storing all features in the vector database is that it enables similarity search, and push-down filtering for all feature values. ```aidl # Create a feature group with the embedding index news_fg = fs.get_or_create_feature_group( name=f"news_fg", embedding_index=emb, # Provide the embedding index created primary_key=["news_id"], version=version, online_enabled=True ) # Write a DataFrame to the feature group, including the offline store and the ANN index (in the Vector Database) news_fg.insert(df) ``` ## Similarity Search for Feature Groups using Vector Embeddings You provide a vector embedding as a parameter to the search query using [`FeatureGroup.find_neighbors`][hsfs.feature_group.FeatureGroup.find_neighbors], and it returns the rows in the online feature group that have vector embedding values most similar to the provided vector embedding. It is also possible to filter rows by specifying a filter on any of the features in the feature group. The filter is pushed down to the vector database to improve query performance. In the first code snippet below, `find_neighbor`s returns 3 rows in `news_fg` that have the closest `news_description` values to the provided `news_description`. In the second code snippet below, we only return news articles with a `newstype` of `sports`. ```aidl # Search neighbor embedding with k=3 news_fg.find_neighbors(model.encode(news_description), k=3) # Filter and search news_fg.find_neighbors(model.encode(news_description), k=3, filter=news_fg.newstype == "sports") ``` To analyze feature values at specific points in time, you can utilize time travel functionality: ```aidl # Time travel and read from the offline feature store news_fg.as_of(time_in_past).read() ``` ## Querying Similar Embeddings with Additional features You can also use similarity search for vector embedding features in feature views. In the code snippet below, we create a feature view by selecting features from the earlier `news_fg` and a new feature group `view_fg`. If you include a feature group with vector embedding features in a feature view, **whether or not the vector embedding features are selected**, you can call `find_neighbors` on the feature view, and it will return rows containing all the feature values in the feature view. In the example below, a list of `heading` and `view_cnt` will be returned for the news articles which are closet to provided `news_description`. ```aidl view_fg = fs.get_or_create_feature_group( name="view_fg", primary_key=["news_id"], version=version, online_enabled=True ) fv = fs.get_or_create_feature_view( "news_view", version=version, query=news_fg.select(["heading"]).join(view_fg.select(["view_cnt"])) ) fv.find_neighbors(model.encode(news_description), k=5) ``` Note that you can use similarity search from the feature view **only if** the feature group which you are querying with `find_neighbors` has **all** the primary keys of the other feature groups. In the example above, you are querying against the feature group `news_fg` which has the vector embedding features, and it has the feature "news_id" which is the primary key of the feature group `view_fg`. But if `page_fg` is used as illustrated below, `find_neighbors` will fail to return any features because primary key `page_id` does not exist in `news_fg`. --8<-- "user_guides/fs/vector_similarity_search/find-neighbors.html" It is also possible to get back feature vector by providing the primary keys, but it is not recommended as explained in the next section. The client fetches feature vector from the vector store and the online store for `news_fg` and `view_fg` respectively. ```aidl fv.get_feature_vector({"news_id": 1}) ``` ## Performance considerations for Feature Groups with Embeddings ### Choose Features for Vector Store While it is possible to update feature value in vector store, updating feature value in online store is more efficient. If you have features which are frequently being updated and do not require for filtering, consider storing them separately in a different feature group. As shown in the previous example, `view_cnt` is updated frequently and stored separately. You can then get all the required features by using feature view. ### Choose the Appropriate Online Feature Stores There are 2 types of online feature stores in Hopsworks: online store (RonDB) and vector store (Opensearch). Online store is designed for retrieving feature vectors efficiently with low latency. Vector store is designed for finding similar embedding efficiently. If similarity search is not required, using online store is recommended for low latency retrieval of feature values including embedding. ### Use New Index per Feature Group Create a new index per feature group to optimize retrieval performance. ## Next steps Explore the [news search example](https://github.com/logicalclocks/hopsworks-tutorials/blob/master/api_examples/vector_similarity_search/1_feature_group_embeddings_api.ipynb), demonstrating how to use Hopsworks for implementing a news search application using natural language in the application. Additionally, you can see the application of querying similar embeddings with additional features in this [news rank example](https://github.com/logicalclocks/hopsworks-tutorials/blob/master/api_examples/vector_similarity_search/2_feature_view_embeddings_api.ipynb). ================================================================================ # Transformation Functions Source: https://docs.hopsworks.ai/latest/user_guides/fs/transformation_functions/ # Transformation Functions In AI systems, [transformation functions](https://www.hopsworks.ai/dictionary/transformation) transform data to create features, the inputs to machine learning models (in both training and inference). The [taxonomy of data transformations](../../concepts/mlops/data_transformations.md) introduces three types of data transformation prevalent in all AI systems. Hopsworks offers simple Python APIs to define custom transformation functions. These can be used along with [feature groups](./feature_group/index.md) and [feature views](./feature_view/overview.md) to create [on-demand transformations](./feature_group/on_demand_transformations.md) and [model-dependent transformations](./feature_view/model-dependent-transformations.md), producing modular AI pipelines that are skew-free. ## Custom Transformation Function Creation User-defined transformation functions can be created in Hopsworks using the [`@udf`][hsfs.hopsworks_udf.udf] decorator. These functions can be either implemented as pure Python UDFs or Pandas UDFs (User-Defined Functions). Hopsworks offers three execution modes to control the execution of transformation functions during training dataset creation, batch inference, and online inference. By default, Hopsworks executes transformation functions as Python UDFs for [feature vector retrieval](feature_view/feature-vectors.md) in online inference pipelines and as Pandas UDFs for both [batch data retrieval](feature_view/batch-data.md) in batch inference pipelines and [training dataset creation](feature_view/training-data.md) in training pipelines. Python UDFs are optimized for smaller data volumes, while Pandas UDFs provide better performance on larger datasets. This execution mode provides the optimal balance based on the data size across training dataset generations, batch inference, and online inference. Additionally, Hopsworks allows you to explicitly set the execution mode for a transformation function to `python` or `pandas`, forcing the transformation function to always run as either a Python or Pandas UDF as specified. A Pandas UDF in Hopsworks accepts one or more Pandas Series as input and can return either one or more Series or a Pandas DataFrame. When integrated with PySpark applications, Hopsworks automatically executes Pandas UDFs using PySpark’s [`pandas_udf`](https://spark.apache.org/docs/3.4.1/api/python/reference/pyspark.sql/api/pyspark.sql.functions.pandas_udf.html), enabling the transformation functions to efficiently scale for large datasets. !!! warning "Java/Scala support" Hopsworks supports transformations functions in Python (Pandas UDFs, Python UDFs). Transformations functions can also be executed in Python-based DataFrame frameworks (PySpark, Pandas). There is currently no support for transformation functions in SQL or Java-based feature pipelines. Transformation functions created in Hopsworks can be directly attached to feature views or feature groups or stored in the feature store for later retrieval. These functions can be part of a library [installed](../../user_guides/projects/python/python_install.md) in Hopsworks or be defined in a [Jupyter notebook](../../user_guides/projects/jupyter/python_notebook.md) running a Python kernel or added when starting a Jupyter notebook or [Hopsworks job](../../user_guides/projects/jobs/spark_job.md). !!! warning "PySpark Kernels" Definition transformation function within a Jupyter notebook is only supported in Python Kernel. In a PySpark Kernel transformation function have to defined as modules or added when starting a Jupyter notebook. The `@udf` decorator in Hopsworks creates a metadata class called [`HopsworksUdf`][hsfs.hopsworks_udf.HopsworksUdf]. This class manages the necessary operations to execute the transformation function. The decorator accepts three parameters: - **`return_type`** (required): Specifies the data type(s) of the features returned by the transformation function. It can be a single Python type if the function returns one transformed feature, or a list of Python types if it returns multiple transformed features. The supported Python types that be used with the `return_type` argument are provided in the table below: | Supported Python Types | | :--------------------: | | str | | int | | float | | bool | | datetime.datetime | | datetime.date | | datetime.time | - **`drop`** (optional): Identifies input arguments to exclude from the output after transformations are applied. By default, all inputs are retained in the output. Further details on this argument can be found [below](#dropping-input-features). - **`mode`** (optional): Determines the execution mode of the transformation function. The argument accepts three values: `default`, `python`, or `pandas`. By default, the `mode` is set to `default`. Further details on this argument can be found [below](#specifying-execution-modes). Hopsworks supports four types of transformation functions across all execution modes: 1. One-to-one: Transforms one feature into one transformed feature. 2. One-to-many: Transforms one feature into multiple transformed features. 3. Many-to-one: Transforms multiple features into one transformed feature. 4. Many-to-many: Transforms multiple features into multiple transformed features. ### One-to-one transformations To create a one-to-one transformation function, the Hopsworks `@udf` decorator must be provided with the `return_type` as a single Python type. The transformation function should take one argument as input and return a Pandas Series. !!! example "Creation of a one-to-one transformation function in Hopsworks." === "Python" ```python from hopsworks import udf @udf(return_type=int) def add_one(feature): return feature + 1 ``` ### Many-to-one transformations The creation of many-to-one transformation functions is similar to that of a one-to-one transformation function, the only difference being that the transformation function accepts multiple features as input. !!! example "Creation of a many-to-one transformation function in Hopsworks." === "Python" ```python from hopsworks import udf @udf(return_type=int) def add_features(feature1, feature2, feature3): return feature1 + feature2 + feature3 ``` ### One-to-many transformations To create a one-to-many transformation function, the Hopsworks `@udf` decorator must be provided with the `return_type` as a list of Python types, and the transformation function should take one argument as input and return multiple features as a Pandas DataFrame. The return types provided to the decorator must match the types of each column in the returned Pandas DataFrame. !!! example "Creation of a one-to-many transformation function in Hopsworks." === "Python" ```python from hopsworks import udf @udf(return_type=[int, int]) def add_one_and_two(feature1): return feature1 + 1, feature1 + 2 ``` ### Many-to-many transformations The creation of a many-to-many transformation function is similar to that of a one-to-many transformation function, the only difference being that the transformation function accepts multiple features as input. !!! example "Creation of a many-to-many transformation function in Hopsworks." === "Python" ```python from hopsworks import udf @udf(return_type=[int, int, int]) def add_one_multiple(feature1, feature2, feature3): return feature1 + 1, feature2 + 1, feature3 + 1 ``` ### Specifying execution modes The `mode` parameter of the `@udf` decorator can be used to specify the execution mode of the transformation function. It accepts three possible values `default`, `python` and `pandas`. Each mode is explained in more detail below: #### Default Mode This execution mode assumes that the transformation function can be executed as either a Pandas UDF or a Python UDF. It serves as the default mode used when the `mode` parameter is not specified. In this mode, the transformation function is executed as a Pandas UDF during training and in the batch inference pipeline, while it operates as a Python UDF during online inference. !!! example "Creating a many to many transformations function using the default execution mode" === "Python" ```python from hopsworks import udf # "default" mode is used if the parameter `mode` is not explicitly set. @udf(return_type=[int, int, int]) def add_one_multiple(feature1, feature2, feature3): return feature1 + 1, feature2 + 1, feature3 + 1 @udf(return_type=[int, int, int], mode="default") def add_two_multiple(feature1, feature2, feature3): return feature1 + 2, feature2 + 2, feature3 + 2 ``` #### Python Mode The transformation function can be configured to always execute as a Python UDF by setting the `mode` parameter of the `@udf` decorator to `python`. !!! example "Creating a many to many transformation function as a Python UDF" === "Python" ```python from hopsworks import udf @udf(return_type=[int, int, int], mode="python") def add_one_multiple(feature1, feature2, feature3): return feature1 + 1, feature2 + 1, feature3 + 1 ``` #### Pandas Mode The transformation function can be configured to always execute as a Pandas UDF by setting the `mode` parameter of the `@udf` decorator to `pandas`. !!! example "Creating a many to many transformations function as a Pandas UDF" === "Python" ```python import pandas as pd from hopsworks import udf # A Pandas UDF returning a Pandas DataFrame @udf(return_type=[int, int, int], mode="pandas") def add_one_multiple(feature1, feature2, feature3): return pd.DataFrame( { "add_one_feature1": feature1 + 1, "add_one_feature2": feature2 + 1, "add_one_feature3": feature3 + 1, } ) # A Pandas UDF returning multiple Pandas Series @udf(return_type=[int, int, int], mode="pandas") def add_two_multiple(feature1, feature2, feature3): return feature1 + 2, feature2 + 2, feature3 + 2 ``` ### Dropping input features The `drop` parameter of the `@udf` decorator is used to drop specific columns in the input DataFrame after transformation. If any argument of the transformation function is passed to the `drop` parameter, then the column mapped to the argument is dropped after the transformation functions are applied. In the example below, the columns mapped to the arguments `feature1` and `feature3` are dropped after the application of all transformation functions. !!! example "Specify arguments to drop after transformation" === "Python" ```python from hopsworks import udf @udf(return_type=[int, int, int], drop=["feature1", "feature3"]) def add_one_multiple(feature1, feature2, feature3): return feature1 + 1, feature2 + 1, feature3 + 1 ``` ### Specifying output features names for transformation functions The [`TransformationFunction.alias`][hsfs.transformation_function.TransformationFunction.alias] function of a transformation function allows the specification of names of transformed features generated by the transformation function. Each name must be uniques and should be at-most 63 characters long. If no name is provided via the `alias` function, Hopsworks generates default output feature names when [on-demand](./feature_group/on_demand_transformations.md) or [model-dependent](./feature_view/model-dependent-transformations.md) transformation functions are created. !!! example "Specifying output column names for transformation functions." === "Python" ```python from hopsworks import udf @udf(return_type=[int, int, int], drop=["feature1", "feature3"]) def add_one_multiple(feature1, feature2, feature3): return feature1 + 1, feature2 + 1, feature3 + 1 # Specifying output feature names of the transformation function. add_one_multiple.alias( "transformed_feature1", "transformed_feature2", "transformed_feature3" ) ``` ### Training dataset statistics A keyword argument `statistics` can be defined in the transformation function if it requires training dataset statistics for any of its arguments. The `statistics` argument must be assigned an instance of the class [`TransformationStatistics`][hsfs.transformation_statistics.TransformationStatistics] as the default value. The `TransformationStatistics` instance must be initialized using the names of the arguments requiring statistics. !!! warning "Transformation Statistics" The statistics provided to the transformation function is the statistics computed using [the train set](https://www.hopsworks.ai/dictionary/train-training-set). Training dataset statistics are not available for on-demand transformations. The `TransformationStatistics` instance contains separate objects with the same name as the arguments used to initialize it. These objects encapsulate statistics related to the argument as instances of the class [`FeatureTransformationStatistics`][hsfs.transformation_statistics.FeatureTransformationStatistics]. Upon instantiation, instances of `FeatureTransformationStatistics` contain `None` values and are updated with the required statistics after the creation of a training dataset. !!! example "Creation of a transformation function in Hopsworks that uses training dataset statistics" === "Python" ```python from hopsworks import udf from hopsworks.transformation_statistics import TransformationStatistics stats = TransformationStatistics("argument1", "argument2", "argument3") @udf(int) def add_features(argument1, argument2, argument3, statistics=stats): return ( argument1 + argument2 + argument3 + statistics.argument1.mean + statistics.argument2.mean + statistics.argument3.mean ) ``` ### Passing context variables to transformation function The `context` keyword argument can be defined in a transformation function to access shared context variables. These variables contain common data used across transformation functions. By including the context argument, you can pass the necessary data as a dictionary into the into the `context` argument of the transformation function during [training dataset creation](feature_view/training-data.md#passing-context-variables-to-transformation-functions) or [feature vector retrieval](feature_view/feature-vectors.md#passing-context-variables-to-transformation-functions) or [batch data retrieval](feature_view/batch-data.md#passing-context-variables-to-transformation-functions). !!! example "Creation of a transformation function in Hopsworks that accepts context variables" === "Python" ```python from hopsworks import udf @udf(int) def add_features(argument1, context): return argument1 + context["value_to_add"] ``` ## Saving to the Feature Store To save a transformation function to the feature store, use the function `create_transformation_function`. It creates a [`TransformationFunction`][hsfs.transformation_function.TransformationFunction] object which can then be saved by calling the save function. The save function will throw an error if another transformation function with the same name and version is already saved in the feature store. !!! example "Register transformation function `add_one` in the Hopsworks feature store" === "Python" ```python plus_one_meta = fs.create_transformation_function( transformation_function=add_one, version=1 ) plus_one_meta.save() ``` ## Retrieval from the Feature Store To retrieve all transformation functions from the feature store, use the function `get_transformation_functions`, which returns the list of `TransformationFunction` objects. A specific transformation function can be retrieved using its `name` and `version` with the function `get_transformation_function`. If only the `name` is provided, then the version will default to 1. !!! example "Retrieving transformation functions from the feature store" === "Python" ```python # get all transformation functions fs.get_transformation_functions() # get transformation function by name. This will default to version 1 plus_one_fn = fs.get_transformation_function(name="plus_one") # get transformation function by name and version. plus_one_fn = fs.get_transformation_function(name="plus_one", version=2) ``` ## Using transformation functions Transformation functions can be used by attaching it to a feature view to [create model-dependent transformations](./feature_view/model-dependent-transformations.md) or attached to feature groups to [create on-demand transformations](./feature_group/on_demand_transformations.md) ## Chained Transformation Functions Transformation functions can be chained: the output column of one transformation function can serve as the input to another. Hopsworks resolves the execution order automatically using a topological sort of the resulting DAG, so dependencies always run before their consumers. Chaining works for both on-demand transformations attached to a feature group and model-dependent transformations attached to a feature view. !!! example "Chained model-dependent transformations on a feature view" === "Python" ```python from hopsworks import udf @udf(int) def add_one(col): return col + 1 @udf(int) def add(a, b): return a + b fv = fs.create_feature_view( name="chained_mdts_fv", query=fg.select_all(), transformation_functions=[ add_one("data1").alias("data1_plus_one"), add_one("data2").alias("data2_plus_one"), add("data1_plus_one", "data2_plus_one").alias("sum_plus_two"), ], version=1, ) ``` The same DAG drives offline training data generation and online feature vector retrieval, so chains apply uniformly across both paths. Statistics-based transformations participate in chains too: a transformation that requires statistics on another transformation's output is fit on that intermediate output, as described in [model-dependent transformations][chaining-model-dependent-transformations]. Chaining also works across the two transformation types without additional setup: an on-demand transformation's output column becomes a feature in its feature group, which a feature view can consume and feed into a model-dependent transformation. A configuration with no valid execution order is rejected: a duplicate output column or a cycle between transformation functions raises an error naming the offending functions, which can be fixed by renaming outputs with `.alias()`. ### Visualizing the execution DAG The execution DAG is shown in the Hopsworks UI on the feature view and feature group overview pages under "Transformation execution DAG." The same graph can be rendered from the SDK with `visualize_transformations()`, available on both feature views and feature groups. It renders as a Mermaid flowchart in Jupyter and as text elsewhere. !!! example "Visualizing transformation DAGs" === "Python" ```python # Render both the model-dependent and on-demand DAGs. fv.visualize_transformations() # Render only the model-dependent DAG, top-to-bottom layout. fv.visualize_transformations(kind="model_dependent", orient="TB") # Render the on-demand DAG of a feature group. fg.visualize_transformations() ``` ### Transformation Functions Performance Tuning Transformation functions execute sequentially unless the `n_processes` argument requests worker processes. The argument is accepted by the feature view and feature group entry points that execute transformations, such as `get_feature_vector`, `get_feature_vectors`, `get_batch_data`, `training_data`, and `transform`. Parallelism is strictly opt-in because whether the worker-pool overhead pays off depends on the cost of your transformation functions. With more than one worker process, independent transformation functions in the DAG run concurrently, while a chained sequence always runs in dependency order. On the Spark engine `n_processes` is ignored because the whole DAG is pushed down to Spark, which distributes the work itself. For batch and offline calls such as `get_feature_vectors`, `get_batch_data`, and `training_data` with CPU-heavy functions benefit from `n_processes >= 2`, for vectorized Pandas UDFs on small inputs, sequential execution is at least as fast because the pool overhead dominates. For online serving, spawning the worker pool during the first request would add the pool startup cost to that request's latency. Passing `n_processes` to `init_serving` or `init_batch_scoring` pre-spawns the pool at initialization time and makes that value the default for subsequent retrieval calls; an explicit `n_processes` on an individual call still takes precedence. !!! example "Pre-spawning the worker pool for online serving" === "Python" ```python fv.init_serving(training_dataset_version=1, n_processes=2) # Served using the pool of two workers spawned at init time. vector = fv.get_feature_vector(entry={"id": 1}) ``` The worker pool start method defaults to `fork` on Linux and `spawn` on macOS and Windows. Set the `HOPSWORKS_TF_POOL_START_METHOD` environment variable to `fork`, `forkserver`, or `spawn` to override it. ================================================================================ # Compute Engines Source: https://docs.hopsworks.ai/latest/user_guides/fs/compute_engines/ ## Compute Engines In order to execute a feature pipeline to write to the Feature Store, as well as to retrieve data from the Feature Store, you need a compute engine. Hopsworks Feature Store APIs are built around dataframes, that means feature data is inserted into the Feature Store from a Dataframe and likewise when reading data from the Feature Store, it is returned as a Dataframe. As such, Hopsworks supports four computational engines: 1. [Apache Spark](https://spark.apache.org): Spark Dataframes and Spark Structured Streaming Dataframes are supported, both from Python environments (PySpark) and from Scala environments. 2. [Python](https://www.python.org/): For pure Python environments without dependencies on Spark, Hopsworks supports [Pandas Dataframes](https://pandas.pydata.org/) and [Polars Dataframes](https://pola.rs/). 3. [Apache Beam](https://beam.apache.org/) *experimental*: Beam Data Streams are currently supported as an experimental feature from Java/Scala environments. 4. [Java](https://www.java.com): For pure Java environments without dependencies on Spark, Hopsworks supports writing using List of POJO Objects. Hopsworks supports running [compute on the platform itself](../../concepts/dev/inside.md) in the form of [Jobs](../projects/jobs/pyspark_job.md) or in [Jupyter Notebooks](../projects/jupyter/python_notebook.md). Alternatively, you can also connect to Hopsworks using Python or Spark from [external environments](../../concepts/dev/outside.md), given that there is network connectivity. ## Functionality Support Hopsworks is aiming to provide functional parity between the computational engines, however, there are certain Hopsworks functionalities which are exclusive to the engines. | Functionality | Method | Spark | Python | Beam | Java | Comment | | --- | --- | --- | --- | --- | --- | --- | | Feature Group Creation from dataframes | [`FeatureStore.create_feature_group`][hsfs.feature_store.FeatureStore.create_feature_group] | :white_check_mark: | :white_check_mark: | - | - | Currently Beam/Java doesn't support registering feature group metadata. Thus it needs to be pre-registered before you can write real time features computed by Beam. | | Training Dataset Creation from dataframes | [`TrainingDataset.save`][hsfs.training_dataset.TrainingDataset.save] | :white_check_mark: | - | - | - | Functionality was deprecated in version 3.0 | | Data validation using Great Expectations for streaming dataframes | [`FeatureGroup.validate`][hsfs.feature_group.FeatureGroup.validate]
[`FeatureGroup.insert_stream`][hsfs.feature_group.FeatureGroup.insert_stream] | - | - | - | - | `insert_stream` does not perform any data validation even when a expectation suite is attached. | | Stream ingestion | [`FeatureGroup.insert_stream`][hsfs.feature_group.FeatureGroup.insert_stream] | :white_check_mark: | - | :white_check_mark: | :white_check_mark: | Python/Pandas/Polars has currently no notion of streaming. | | Reading from Streaming Storage Connectors | [`KafkaConnector.read_stream`][hsfs.storage_connector.KafkaConnector.read_stream] | :white_check_mark: | - | - | - | Python/Pandas/Polars has currently no notion of streaming. For Beam/Java only write operations are supported | | Reading training data from external storage other than S3 | [`FeatureView.get_training_data`][hsfs.feature_view.FeatureView.get_training_data] | :white_check_mark: | - | - | - | Reading training data that was written to external storage using a Storage Connector other than S3 can currently not be read using Hopsworks APIs, instead you will have to use the storage's native client. | | Reading External Feature Groups into Dataframe | [`ExternalFeatureGroup.read`][hsfs.feature_group.ExternalFeatureGroup.read] | :white_check_mark: | - | - | - | Reading an External Feature Group directly into a Pandas/Polars Dataframe is not supported, however, you can use the [Query API][hsfs.constructor.query.Query] to create Feature Views/Training Data containing External Feature Groups. | | Read Queries containing External Feature Groups into Dataframe | [`Query.read`][hsfs.constructor.query.Query.read] | :white_check_mark: | - | - | - | Reading a Query containing an External Feature Group directly into a Pandas/Polars Dataframe is not supported, however, you can use the Query to create Feature Views/Training Data and write the data to a Storage Connector, from where you can read up the data into a Pandas/Polars Dataframe. | ## Python ### Python Inside Hopsworks If you are using Spark or Python within Hopsworks, there is no further configuration required. Head over to the [Getting Started Guide](https://colab.research.google.com/github/logicalclocks/hopsworks-tutorials/blob/master/quickstart.ipynb){:target="_blank"}. ### Python Outside Hopsworks Connecting to the Feature Store from any Python environment, such as your local environment or Google Colab, requires setting up an API Key and installing the Hopsworks Python client library. The [Python integration guide](../integrations/python.md) explains step by step how to connect to the Feature Store from any Python environment. ## Spark ### Spark Inside Hopsworks If you are using Spark or Python within Hopsworks, there is no further configuration required. Head over to the [Getting Started Guide](https://colab.research.google.com/github/logicalclocks/hopsworks-tutorials/blob/master/quickstart.ipynb){:target="_blank"}. ### Spark Outside Hopsworks Connecting to the Feature Store from an external Spark cluster, such as Cloudera or Databricks, requires configuring it with the Hopsworks client jars, configuration and certificates. The [Spark integration guide](../integrations/spark.md) explains step by step how to connect to the Feature Store from an external Spark cluster. ## Beam ### Beam Inside Hopsworks Beam is only supported as an external client. ### Beam Outside Hopsworks Connecting to the Feature Store from Beam DataFlowRunner, requires configuring the Hopsworks certificates. The [Beam integration guide](../integrations/beam.md) explains step by step how to connect to the Feature Store from Beam Dataflow Runner. !!! warning Apache Beam integration with Hopsworks feature store was only tested using Dataflow Runner. For more details head over to the [Getting Started Guide](https://github.com/logicalclocks/hopsworks-tutorials/tree/master/integrations/java/beam). ## Java It is also possible to interact to Hopsworks feature store using pure Java environments without dependencies on Spark or Beam. For more details head over to the [Getting Started Guide](https://github.com/logicalclocks/hopsworks-tutorials/tree/master/java). ================================================================================ # Client Integrations Source: https://docs.hopsworks.ai/latest/user_guides/integrations/ # Client Integrations Hopsworks is an open platform, reachable from the tools you already use. Pick the client you connect from.
- **Python** --- Any Python environment, including SageMaker, Google Colab and Kubeflow. [Connect from Python](python.md) - **Java** --- Java and Scala clients. [Connect from Java](java.md) - **Databricks** --- Connect a Databricks workspace. [Connect from Databricks](databricks/networking.md) - **AWS EMR** --- Connect an EMR cluster. [Connect from AWS EMR](emr/emr_configuration.md) - **Azure HDInsight** --- Connect an HDInsight cluster. [Connect from Azure HDInsight](hdinsight.md) - **Azure Machine Learning** --- ML Studio designer and notebooks. [Connect from Azure Machine Learning](mlstudio_designer.md) - **Apache Spark** --- Connect an external Spark cluster. [Connect from Apache Spark](spark.md) - **Apache Beam** --- Feature pipelines on Beam. [Connect from Apache Beam](beam.md)
================================================================================ # Python / SageMaker / Kubeflow Source: https://docs.hopsworks.ai/latest/user_guides/integrations/python/ # Python Environments (Local, AWS SageMaker, Google Colab or Kubeflow) This guide explains step by step how to connect to Hopsworks from any Python environment such as your local environment, AWS SageMaker, Google Colab or Kubeflow. ## Install Python Library To be able to interact with Hopsworks from a Python environment you need to install the `Hopsworks` Python library. The library is available on [PyPi](https://pypi.org/project/hopsworks/) and is installed with the `python` profile: === "uv" ```bash uv pip install "hopsworks[python]~=[HOPSWORKS_VERSION]" ``` === "pip" ```bash pip install "hopsworks[python]~=[HOPSWORKS_VERSION]" ``` !!! attention "Python Profile" A bare `hopsworks` install does not bring the dependencies needed to use the library from a pure Python environment. Always install with the `python` profile, `hopsworks[python]`. !!! attention "Matching Hopsworks version" We recommend that the major and minor version of the Python library match the major and minor version of the Hopsworks deployment.

The library version needs to match the major version of Hopsworks
You find the Hopsworks version at the bottom of the help menu in the top navigation bar

## Generate an API key For instructions on how to generate an API key follow this [user guide](../projects/api_key/create_api_key.md). For the Python client to work correctly make sure you add the following scopes to your API key: 1. featurestore 2. project 3. job 4. kafka ## Connect to the Feature Store You are now ready to connect to Hopsworks from your Python environment: ```python import hopsworks project = hopsworks.login( host="my_instance", # DNS of your Hopsworks instance port=443, # Port to reach your Hopsworks instance, defaults to 443 project="my_project", # Name of your Hopsworks project api_key_value="apikey", # The API key to authenticate with Hopsworks engine="python", # Use the Python engine ) fs = project.get_feature_store() # Get the project's default feature store ``` !!! note "Engine" `Hopsworks` leverages several engines depending on whether you are running using Apache Spark or Pandas/Polars. The default behaviour of the library is to use the `spark` engine if you do not specify any `engine` option in the `login` method and if the `PySpark` library is available in the environment. Please refer to the [Spark integration guide](spark.md) to configure your PySpark cluster to interact with Hopsworks. ## Next Steps For more information on how to use the Hopsworks API check out the other guides or the [Login API][hopsworks.login]. ================================================================================ # Networking Source: https://docs.hopsworks.ai/latest/user_guides/integrations/emr/networking/ # Networking In order for Spark to communicate with the Hopsworks Feature Store from EMR, networking needs to be set up correctly. This includes deploying the Hopsworks Feature Store to either the same VPC or enable VPC peering between the VPC of the EMR cluster and the Hopsworks Feature Store. ## Step 1: Ensure network connectivity The DataFrame API needs to be able to connect directly to the IP on which the Feature Store is listening. This means that if you deploy the Feature Store on AWS you will either need to deploy the Feature Store in the same VPC as your EMR cluster or to set up [VPC Peering](https://docs.aws.amazon.com/vpc/latest/peering/create-vpc-peering-connection.html) between your EMR VPC and the Feature Store VPC. ### Option 1: Deploy the Feature Store in the EMR VPC When deploying the Hopsworks Feature Store, select the EMR *VPC* and *Availability Zone* as the VPC and Availability Zone of your Feature Store. Identify your EMR VPC in the Summary of your EMR cluster:

Identify the EMR VPC
Identify the EMR VPC

Identify the EMR VPC
Identify the EMR VPC

### Option 2: Set up VPC peering Follow the guide [VPC Peering](https://docs.aws.amazon.com/vpc/latest/peering/create-vpc-peering-connection.html) to set up VPC peering between the Feature Store and EMR. Get your Feature Store *VPC ID* and *CIDR* by searching for the Feature Store VPC in the AWS Management Console:

Identify the Feature Store VPC
Identify the Feature Store VPC

## Step 2: Configure the Security Group The Feature Store *Security Group* needs to be configured to allow traffic from your EMR clusters to be able to connect to the Feature Store. Open your feature store instance under EC2 in the AWS Management Console and ensure that ports *443*, *3306*, *9083*, *9085*, *8020* and *30010* (443,3306,8020,30010,9083,9085) are reachable from the EMR Security Group:

Hopsworks Feature Store Security Group
Hopsworks Feature Store Security Group

Connectivity from the EMR Security Group can be allowed by opening the Security Group, adding a port to the Inbound rules and setting the EMR master and core security group as source:

Hopsworks Feature Store Security Group details
Hopsworks Feature Store Security Group details

You can find your EMR security groups in the EMR cluster summary:

EMR Security Groups
EMR Security Groups

## Next Steps Continue with the [Configure EMR for the Hopsworks Feature Store](emr_configuration.md), in order to be able to use the Hopsworks Feature Store. ================================================================================ # Configure EMR for Hopsworks Source: https://docs.hopsworks.ai/latest/user_guides/integrations/emr/emr_configuration/ # Configure EMR for the Hopsworks Feature Store To enable EMR to access the Hopsworks Feature Store, you need to set up a Hopsworks API key, add a bootstrap action and configurations to your EMR cluster. !!! info Ensure [Networking](networking.md) is set up correctly before proceeding with this guide. ## Step 1: Set up a Hopsworks API key For instructions on how to generate an API key follow this [user guide](../../projects/api_key/create_api_key.md). For the EMR integration to work correctly make sure you add the following scopes to your API key: 1. featurestore 2. project 3. job 4. kafka ### Store the API key in the AWS Secrets Manager In the AWS management console ensure that your active region is the region you use for EMR. Go to the *AWS Secrets Manager* and select *Store new secret*. Select *Other type of secrets* and add *api-key* as the key and paste the API key created in the previous step as the value. Click next.

Store a Hopsworks API key in the Secrets Manager
Store a Hopsworks API key in the Secrets Manager

As a secret name, enter *hopsworks/featurestore*. Select next twice and finally store the secret. Then click on the secret in the secrets list and take note of the *Secret ARN*.

Name the secret
Name the secret

### Grant access to the secret to the EMR EC2 instance profile Identify your EMR EC2 instance profile in the EMR cluster summary:

Identify your EMR EC2 instance profile
Identify your EMR EC2 instance profile

In the AWS Management Console, go to *IAM*, select *Roles* and then the EC2 instance profile used by your EMR cluster. Select *Add inline policy*. Choose *Secrets Manager* as a service, expand the *Read* access level and check *GetSecretValue*. Expand Resources and select *Add ARN*. Paste the ARN of the secret created in the previous step. Click on *Review*, give the policy a name and click on *Create policy*.

Configure the access policy for the Secrets Manager
Configure the access policy for the Secrets Manager

## Step 2: Configure your EMR cluster ### Add the Hopsworks Feature Store configuration to your EMR cluster In order for EMR to be able to talk to the Feature Store, you need to update the Hadoop and Spark configurations. Copy the configuration below and replace ip-XXX-XX-XX-XXX.XX-XXXX-X.compute.internal with the private DNS name of your Hopsworks master node. ```json [ { "Classification": "hadoop-env", "Properties": { }, "Configurations": [ { "Classification": "export", "Properties": { "HADOOP_CLASSPATH": "$HADOOP_CLASSPATH:/usr/lib/hopsworks/client/*" }, "Configurations": [ ] } ] }, { "Classification": "spark-defaults", "Properties": { "spark.hadoop.hops.ipc.server.ssl.enabled": true, "spark.hadoop.fs.hopsfs.impl": "io.hops.hopsfs.client.HopsFileSystem", "spark.hadoop.client.rpc.ssl.enabled.protocol": "TLSv1.2", "spark.hadoop.hops.ssl.hostname.verifier": "ALLOW_ALL", "spark.hadoop.hops.rpc.socket.factory.class.default": "io.hops.hadoop.shaded.org.apache.hadoop.net.HopsSSLSocketFactory", "spark.hadoop.hops.ssl.keystores.passwd.name": "/usr/lib/hopsworks/material_passwd", "spark.hadoop.hops.ssl.keystore.name": "/usr/lib/hopsworks/keyStore.jks", "spark.hadoop.hops.ssl.trustore.name": "/usr/lib/hopsworks/trustStore.jks", "spark.serializer": "org.apache.spark.serializer.KryoSerializer", "spark.executor.extraClassPath": "/usr/lib/hopsworks/client/*", "spark.driver.extraClassPath": "/usr/lib/hopsworks/client/*", "spark.sql.hive.metastore.jars": "path", "spark.sql.hive.metastore.jars.path": "/usr/lib/hopsworks/apache-hive-bin/lib/*", "spark.hadoop.hive.metastore.uris": "thrift://ip-XXX-XX-XX-XXX.XX-XXXX-X.compute.internal:9083" } }, ] ``` When you create your EMR cluster, add the configuration: !!! note Don't forget to replace ip-XXX-XX-XX-XXX.XX-XXXX-X.compute.internal with the private DNS name of your Hopsworks master node.

Configure EMR to access the Feature Store
Configure EMR to access the Feature Store

### Add the Bootstrap Action to your EMR cluster EMR requires Hopsworks connectors to be able to communicate with the Hopsworks Feature Store. These connectors can be installed with the bootstrap action shown below. Copy the content into a file and name the file `hopsworks.sh`. Copy that file into any S3 bucket that is readable by your EMR clusters and take note of the S3 URI of that file e.g., `s3://my-emr-init/hopsworks.sh`. ```bash #!/bin/bash set -e if [ "$#" -ne 3 ]; then echo "Usage hopsworks.sh HOPSWORKS_API_KEY_SECRET, HOPSWORKS_HOST, PROJECT_NAME" exit 1 fi SECRET_NAME=$1 HOST=$2 PROJECT=$3 API_KEY=$(aws secretsmanager get-secret-value --secret-id $SECRET_NAME | jq -r .SecretString | jq -r '.["api-key"]') PROJECT_ID=$(curl -H "Authorization: ApiKey ${API_KEY}" https://$HOST/hopsworks-api/api/project/getProjectInfo/$PROJECT | jq -r .projectId) sudo yum -y install python3-devel.x86_64 || true sudo mkdir /usr/lib/hopsworks sudo chown hadoop:hadoop /usr/lib/hopsworks cd /usr/lib/hopsworks curl -o client.tar.gz -H "Authorization: ApiKey ${API_KEY}" https://$HOST/hopsworks-api/api/project/$PROJECT_ID/client tar -xvf client.tar.gz tar -xzf client/apache-hive-*-bin.tar.gz || true mv apache-hive-*-bin apache-hive-bin rm client.tar.gz rm client/apache-hive-*-bin.tar.gz curl -H "Authorization: ApiKey ${API_KEY}" https://$HOST/hopsworks-api/api/project/$PROJECT_ID/credentials | jq -r .kStore | base64 -d > keyStore.jks curl -H "Authorization: ApiKey ${API_KEY}" https://$HOST/hopsworks-api/api/project/$PROJECT_ID/credentials | jq -r .tStore | base64 -d > trustStore.jks echo -n $(curl -H "Authorization: ApiKey ${API_KEY}" https://$HOST/hopsworks-api/api/project/$PROJECT_ID/credentials | jq -r .password) > material_passwd chmod -R o-rwx /usr/lib/hopsworks sudo pip3 install --upgrade hopsworks~=X.X.0 ``` !!! attention "Matching Hopsworks version" We recommend that the major and minor version of the Python library match the major and minor version of the Hopsworks deployment.

The library version needs to match the major version of Hopsworks
You find the Hopsworks version at the bottom of the help menu in the top navigation bar

Add the bootstrap actions when configuring your EMR cluster. Provide 3 arguments to the bootstrap action: The name of the API key secret e.g., `hopsworks/featurestore`, the public DNS name of your Hopsworks cluster, such as `ad005770-33b5-11eb-b5a7-bfabd757769f.cloud.hopsworks.ai`, and the name of your Hopsworks project, e.g. `demo_fs_meb10179`.

Set the bootstrap action for EMR
Set the bootstrap action for EMR

Your EMR cluster will now be able to access your Hopsworks Feature Store. ## Next Steps Use the [Login API][hopsworks.login] to connect to the Hopsworks Feature Store. For more information about how to use the Feature Store, see the [Quickstart Guide](https://colab.research.google.com/github/logicalclocks/hopsworks-tutorials/blob/master/quickstart.ipynb){:target="_blank"}. ================================================================================ # Azure HDInsight Source: https://docs.hopsworks.ai/latest/user_guides/integrations/hdinsight/ # Configure HDInsight for the Hopsworks Feature Store To enable HDInsight to access the Hopsworks Feature Store, you need to set up a Hopsworks API key, add a script action and configurations to your HDInsight cluster. !!! info "Prerequisites" A HDInsight cluster with cluster type Spark is required to connect to the Feature Store. You can either use an existing cluster or create a new one. !!! info "Network Connectivity" To be able to connect to the Feature Store, please ensure that your HDInsight cluster and the Hopsworks Feature Store are either in the same [Virtual Network](https://docs.microsoft.com/en-us/azure/virtual-network/virtual-networks-overview) or [Virtual Network Peering](https://docs.microsoft.com/en-us/azure/virtual-network/virtual-network-manage-peering) is set up between the different networks. In addition, ensure that the Network Security Group of your Hopsworks instance is configured to allow incoming traffic from your HDInsight cluster on ports 443, 3306, 8020, 30010, 9083 and 9085 (443,3306,8020,30010,9083,9085). See [Network security groups](https://docs.microsoft.com/en-us/azure/virtual-network/network-security-groups-overview) for more information. ## Step 1: Set up a Hopsworks API key For instructions on how to generate an API key follow this [user guide](../projects/api_key/create_api_key.md). For the HDInsight integration to work correctly make sure you add the following scopes to your API key: 1. featurestore 2. project 3. job 4. kafka ## Step 2: Use a script action to install the Feature Store connector HDInsight requires Hopsworks connectors to be able to communicate with the Hopsworks Feature Store. These connectors can be installed with the script action shown below. Copy the content into a file, name the file `hopsworks.sh` and replace MY_INSTANCE, MY_PROJECT, MY_VERSION, MY_API_KEY and MY_CONDA_ENV with your values. Copy the `hopsworks.sh` file into any storage that is readable by your HDInsight clusters and take note of the URI of that file e.g., `https://account.blob.core.windows.net/scripts/hopsworks.sh`. The script action needs to be applied head and worker nodes and can be applied during cluster creation or to an existing cluster. Ensure to persist the script action so that it is run on newly created nodes. For more information about how to use script actions, see [Customize Azure HDInsight clusters by using script actions](https://docs.microsoft.com/en-us/azure/hdinsight/hdinsight-hadoop-customize-cluster-linux). !!! attention "Matching Hopsworks version" We recommend that the major and minor version of the Python library match the major and minor version of the Hopsworks deployment.

The library version needs to match the major version of Hopsworks
You find the Hopsworks version at the bottom of the help menu in the top navigation bar

Feature Store script action: ```bash set -e HOST="MY_INSTANCE.cloud.hopsworks.ai" # DNS of your Feature Store instance PROJECT="MY_PROJECT" # Port to reach your Hopsworks instance, defaults to 443 HOPSWORKS_VERSION="MY_VERSION" # The major version of Hopsworks library needs to match the major version of Hopsworks API_KEY="MY_API_KEY" # The API key to authenticate with Hopsworks CONDA_ENV="MY_CONDA_ENV" # py35 is the default for HDI 3.6 apt-get --assume-yes install python3-dev apt-get --assume-yes install jq /usr/bin/anaconda/envs/$CONDA_ENV/bin/pip install hopsworks==$HOPSWORKS_VERSION PROJECT_ID=$(curl -H "Authorization: ApiKey ${API_KEY}" https://$HOST/hopsworks-api/api/project/getProjectInfo/$PROJECT | jq -r .projectId) mkdir -p /usr/lib/hopsworks chown root:hadoop /usr/lib/hopsworks cd /usr/lib/hopsworks curl -o client.tar.gz -H "Authorization: ApiKey ${API_KEY}" https://$HOST/hopsworks-api/api/project/$PROJECT_ID/client tar -xvf client.tar.gz tar -xzf client/apache-hive-*-bin.tar.gz mv apache-hive-*-bin apache-hive-bin rm client.tar.gz rm client/apache-hive-*-bin.tar.gz curl -H "Authorization: ApiKey ${API_KEY}" https://$HOST/hopsworks-api/api/project/$PROJECT_ID/credentials | jq -r .kStore | base64 -d > keyStore.jks curl -H "Authorization: ApiKey ${API_KEY}" https://$HOST/hopsworks-api/api/project/$PROJECT_ID/credentials | jq -r .tStore | base64 -d > trustStore.jks echo -n $(curl -H "Authorization: ApiKey ${API_KEY}" https://$HOST/hopsworks-api/api/project/$PROJECT_ID/credentials | jq -r .password) > material_passwd chown -R root:hadoop /usr/lib/hopsworks ``` ## Step 3: Configure HDInsight for Feature Store access The Hadoop and Spark installations of the HDInsight cluster need to be configured in order to access the Feature Store. This can be achieved either by using a [bootstrap script](https://docs.microsoft.com/en-us/azure/hdinsight/hdinsight-hadoop-customize-cluster-bootstrap) when creating clusters or using [Ambari](https://docs.microsoft.com/en-us/azure/hdinsight/hdinsight-hadoop-manage-ambari) on existing clusters. Apply the following configurations to your HDInsight cluster. !!! attention "Using Hive and the Feature Store" HDInsight clusters cannot use their local Hive when being configured for the Feature Store as the Feature Store relies on custom Hive binaries and its own Metastore which will overwrite the local one. If you rely on Hive for feature engineering then it is advised to write your data to an external data storage such as ADLS from your main HDInsight cluster and in the Feature Store, create an [on-demand](../../concepts/fs/feature_group/on_demand_feature.md) Feature Group on the storage container in ADLS. Hadoop hadoop-env.sh: ```sh export HADOOP_CLASSPATH=$HADOOP_CLASSPATH:/usr/lib/hopsworks/client/* ``` Hadoop core-site.xml: ```ini hops.ipc.server.ssl.enabled=true fs.hopsfs.impl=io.hops.hopsfs.client.HopsFileSystem client.rpc.ssl.enabled.protocol=TLSv1.2 hops.ssl.keystore.name=/usr/lib/hopsworks/keyStore.jks hops.rpc.socket.factory.class.default=io.hops.hadoop.shaded.org.apache.hadoop.net.HopsSSLSocketFactory hops.ssl.keystores.passwd.name=/usr/lib/hopsworks/material_passwd hops.ssl.hostname.verifier=ALLOW_ALL hops.ssl.trustore.name=/usr/lib/hopsworks/trustStore.jks ``` Spark spark-defaults.conf: ```ini spark.executor.extraClassPath=/usr/lib/hopsworks/client/* spark.driver.extraClassPath=/usr/lib/hopsworks/client/* spark.sql.hive.metastore.jars=path spark.sql.hive.metastore.jars.path=/usr/lib/hopsworks/apache-hive-bin/lib/* ``` Spark hive-site.xml: ```ini hive.metastore.uris=thrift://MY_HOPSWORKS_INSTANCE_PRIVATE_IP:9083 ``` !!! info Replace MY_HOPSWORKS_INSTANCE_PRIVATE_IP with the private IP address of you Hopsworks Feature Store. ## Step 5: Connect to the Feature Store You are now ready to connect to the Hopsworks Feature Store, for instance using a Jupyter notebook in HDInsight with a PySpark3 kernel: ```python import hopsworks # Put the API key into Key Vault for any production setup: # See, https://azure.microsoft.com/en-us/services/key-vault/ secret_value = "MY_API_KEY" # Create a connection project = hopsworks.login( host="MY_INSTANCE.cloud.hopsworks.ai", # DNS of your Feature Store instance port=443, # Port to reach your Hopsworks instance, defaults to 443 project="MY_PROJECT", # Name of your Hopsworks project api_key_value=secret_value, # The API key to authenticate with Hopsworks hostname_verification=True, # Disable for self-signed certificates ) # Get the feature store handle for the project's feature store fs = project.get_feature_store() ``` ## Next Steps For more information on how to use the Hopsworks API check out the other guides or the [Login API][hopsworks.login]. ================================================================================ # Designer Source: https://docs.hopsworks.ai/latest/user_guides/integrations/mlstudio_designer/ # Azure Machine Learning Designer Integration Connecting to Hopsworks from the Azure Machine Learning Designer requires setting up a Hopsworks API key for the Designer and installing the **Hopsworks** Python library on the Designer. This guide explains step by step how to connect to the Feature Store from Azure Machine Learning Designer. !!! info "Network Connectivity" To be able to connect to the Feature Store, please ensure that the Network Security Group of your Hopsworks instance on Azure is configured to allow incoming traffic from your compute target on ports 443, 9083 and 9085 (443,9083,9085). See [Network security groups](https://docs.microsoft.com/en-us/azure/virtual-network/network-security-groups-overview) for more information. If your compute target is not in the same VNet as your Hopsworks instance and the Hopsworks instance is not accessible from the internet then you will need to configure [Virtual Network Peering](https://docs.microsoft.com/en-us/azure/virtual-network/virtual-network-manage-peering). ## Generate an API key For instructions on how to generate an API key follow this [user guide](../projects/api_key/create_api_key.md). For the Azure ML Designer integration to work correctly make sure you add the following scopes to your API key: 1. featurestore 2. project 3. job 4. kafka ## Connect to Hopsworks To connect to Hopsworks from the Azure Machine Learning Designer, create a new pipeline or open an existing one:

Add an Execute Python Script step
Add an Execute Python Script step

In the pipeline, add a new `Execute Python Script` step and replace the Python script from the next step:

Add the code to access the Hopsworks
Add the code to access the Hopsworks

!!! info "Updating the script" Replace MY_VERSION, MY_API_KEY, MY_INSTANCE, MY_PROJECT and MY_FEATURE_GROUP with the respective values. The major version set for MY_VERSION needs to match the major version of Hopsworks. Check [PyPI](https://pypi.org/project/hopsworks/#history) for available releases.

Hopsworks version needs to match the major version of Hopsworks
You find the Hopsworks version at the bottom of the help menu in the top navigation bar

```python import importlib.util import os package_name = "hopsworks" version = "MY_VERSION" spec = importlib.util.find_spec(package_name) if spec is None: import os os.system(f"pip install %s[python]==%s" % (package_name, version)) # Put the API key into Key Vault for any production setup: # See, https://docs.microsoft.com/en-us/azure/machine-learning/how-to-use-secrets-in-runs # from azureml.core import Experiment, Run # run = Run.get_context() # secret_value = run.get_secret(name="fs-api-key") secret_value = "MY_API_KEY" def azureml_main(dataframe1=None, dataframe2=None): import hopsworks project = hopsworks.login( host="MY_INSTANCE.cloud.hopsworks.ai", # DNS of your Hopsworks instance port=443, # Port to reach your Hopsworks instance, defaults to 443 project="MY_PROJECT", # Name of your Hopsworks project api_key_value=secret_value, # The API key to authenticate with Hopsworks hostname_verification=True, # Disable for self-signed certificates engine="python", # Choose python as engine ) fs = project.get_feature_store() # Get the project's default feature store return (fs.get_feature_group("MY_FEATURE_GROUP", version=1).read(),) ``` Select a compute target and save the step. The step is now ready to use:

Select a compute target
Select a compute target

As a next step, you have to connect the previously created `Execute Python Script` step with the next step in the pipeline. For instance, to export the features to a CSV file, create a `Export Data` step:

Add an Export Data step
Add an Export Data step

Configure the `Export Data` step to write to you data store of choice:

Configure the Export Data step
Configure the Export Data step

Connect the to steps by drawing a line between them:

Connect the steps
Connect the steps

Finally, submit the pipeline and wait for it to finish: !!! info "Performance on the first execution" The `Execute Python Script` step can be slow when being executed for the first time as the Hopsworks library needs to be installed on the compute target. Subsequent executions on the same compute target should use the already installed library.

Execute the pipeline
Execute the pipeline

## Next Steps For more information on how to use the Hopsworks API check out the other guides or the [Login API][hopsworks.login]. ================================================================================ # Notebooks Source: https://docs.hopsworks.ai/latest/user_guides/integrations/mlstudio_notebooks/ # Azure Machine Learning Notebooks Integration Connecting to the Hopsworks from Azure Machine Learning Notebooks requires setting up a Hopsworks API key for Azure Machine Learning Notebooks and installing the **Hopsworks** Python library on the notebook. This guide explains step by step how to connect to the Hopsworks from Azure Machine Learning Notebooks. !!! info "Network Connectivity" To be able to connect to the Feature Store, please ensure that the Network Security Group of your Hopsworks instance on Azure is configured to allow incoming traffic from your compute target on ports 443, 9083 and 9085 (443,9083,9085). See [Network security groups](https://docs.microsoft.com/en-us/azure/virtual-network/network-security-groups-overview) for more information. If your compute target is not in the same VNet as your Hopsworks instance and the Hopsworks instance is not accessible from the internet then you will need to configure [Virtual Network Peering](https://docs.microsoft.com/en-us/azure/virtual-network/virtual-network-manage-peering). ## Install Hopsworks Python Library To be able to interact with Hopsworks from a Python environment you need to install the `Hopsworks` Python library. The library is available on [PyPi](https://pypi.org/project/hopsworks/) and is installed with the `python` profile: === "uv" ```bash uv pip install "hopsworks[python]~=[HOPSWORKS_VERSION]" ``` === "pip" ```bash pip install "hopsworks[python]~=[HOPSWORKS_VERSION]" ``` !!! attention "Python Profile" A bare `hopsworks` install does not bring the dependencies needed to use the library from a local Python environment. Always install with the `python` profile, `hopsworks[python]`. !!! attention "Matching Hopsworks version" We recommend that the major and minor version of the Python library match the major and minor version of the Hopsworks deployment.

The library version needs to match the major version of Hopsworks
You find the Hopsworks version at the bottom of the help menu in the top navigation bar

## Generate an API key For instructions on how to generate an API key follow this [user guide](../projects/api_key/create_api_key.md). For the Azure ML Notebooks integration to work correctly make sure you add the following scopes to your API key: 1. featurestore 2. project 3. job 4. kafka ## Connect from an Azure Machine Learning Notebook To access Hopsworks from Azure Machine Learning, open a Python notebook and proceed with the following steps to install Hopsworks and connect to the Feature Store:

Connecting from an Azure Machine Learning Notebook
Connecting from an Azure Machine Learning Notebook

### Connect to Hopsworks You are now ready to connect to Hopsworks Feature Store from the notebook: ```python import hopsworks # Put the API key into Key Vault for any production setup: # See, https://docs.microsoft.com/en-us/azure/machine-learning/how-to-use-secrets-in-runs # from azureml.core import Experiment, Run # run = Run.get_context() # secret_value = run.get_secret(name="fs-api-key") secret_value = "MY_API_KEY" # Create a connection project = hopsworks.login( host="MY_INSTANCE.cloud.hopsworks.ai", # DNS of your Hopsworks instance port=443, # Port to reach your Hopsworks instance, defaults to 443 project="MY_PROJECT", # Name of your Hopsworks project api_key_value=secret_value, # The API key to authenticate with Hopsworks hostname_verification=True, # Disable for self-signed certificates engine="python", # Choose Python as engine ) # Get the feature store handle for the project's feature store fs = project.get_feature_store() ``` ## Next Steps For more information on how to use the Hopsworks API check out the other guides or the [Login API][hopsworks.login]. ================================================================================ # Apache Spark Source: https://docs.hopsworks.ai/latest/user_guides/integrations/spark/ # Spark Integration Connecting to the Feature Store from an external Spark cluster, such as Cloudera, requires configuring it with the Hopsworks client jars and configuration. This guide explains step by step how to connect to the Feature Store from an external Spark cluster. ## Download the Hopsworks Client Jars In the *Project Settings*, select the *integration* tab and scroll to the *Configure Spark Integration* section. Click on *Download client Jars*. This will start the download of the *client.tar.gz* archive. The archive contains two jar files for HopsFS, the Apache Hudi jar and the Java version of the Hopsworks library. You should upload these libraries to your Spark cluster and attach them as local resources to your Job. If you are using `spark-submit`, you should specify the `--jar` option. For more details see: [Spark Dependency Management](https://spark.apache.org/docs/latest/submitting-applications.html#advanced-dependency-management).

Spark integration tab
The Spark Integration gives access to Jars and configuration for an external Spark cluster

## Download the certificates Download the certificates from the same section as above. Hopsworks uses X.509 certificates for authentication and authorization. If you are interested in the Hopsworks security model, you can read more about it in this [blog post](https://www.logicalclocks.com/blog/how-we-secure-your-data-with-hopsworks). The certificates are composed of three different components: the `keyStore.jks` containing the private key and the certificate for your project user, the `trustStore.jks` containing the certificates for the Hopsworks certificates authority, and a password to unlock the private key in the `keyStore.jks`. The password is displayed in a pop-up when downloading the certificate and should be saved in a file named `material_passwd`. !!! warning When you copy-paste the password to the `material_passwd` file, pay attention to not introduce additional empty spaces or new lines. The three files (`keyStore.jks`, `trustStore.jks` and `material_passwd`) should be attached as resources to your Spark application as well. ## Configure your Spark cluster !!! warning "Spark version limitation" Currently Spark version 3.3.x is suggested to be able to use the full suite of Hopsworks Feature Store capabilities. Add the following configuration to the Spark application: ```plaintext spark.hadoop.fs.hopsfs.impl io.hops.hopsfs.client.HopsFileSystem spark.hadoop.hops.ipc.server.ssl.enabled true spark.hadoop.hops.ssl.hostname.verifier ALLOW_ALL spark.hadoop.hops.rpc.socket.factory.class.default io.hops.hadoop.shaded.org.apache.hadoop.net.HopsSSLSocketFactory spark.hadoop.client.rpc.ssl.enabled.protocol TLSv1.2 spark.hadoop.hops.ssl.keystores.passwd.name material_passwd spark.hadoop.hops.ssl.keystore.name keyStore.jks spark.hadoop.hops.ssl.trustore.name trustStore.jks spark.sql.hive.metastore.jars path spark.sql.hive.metastore.jars.path [Path to the Hopsworks Hive Jars] spark.hadoop.hive.metastore.uris thrift://[metastore_ip]:[metastore_port] ``` `spark.sql.hive.metastore.jars.path` should point to the path with the jars from the uncompressed Hive archive you can find in *clients.tar.gz*. ## PySpark To use PySpark, install the Hopsworks Python library which can be found on [PyPi](https://pypi.org/project/hsfs/). !!! attention "Matching Hopsworks version" The **major version of `Hopsworks`** needs to match the **major version of Hopsworks**. ## Generating an API Key For instructions on how to generate an API key follow this [user guide](../projects/api_key/create_api_key.md). For the Spark integration to work correctly make sure you add the following scopes to your API key: 1. featurestore 2. project 3. job 4. kafka ## Connecting to the Feature Store You are now ready to connect to the Hopsworks Feature Store from Spark: ```python import hopsworks project = hopsworks.login( host="my_instance", # DNS of your Feature Store instance port=443, # Port to reach your Hopsworks instance, defaults to 443 project="my_project", # Name of your Hopsworks Feature Store project api_key_value="api_key", # The API key to authenticate with the feature store hostname_verification=True, # Disable for self-signed certificates ) fs = project.get_feature_store() # Get the project's default feature store ``` !!! note "Engine" `Hopsworks` leverages several engines depending on whether you are running using Apache Spark or Pandas/Polars. The default behaviour of the library is to use the `spark` engine if you do not specify any `engine` option in the `login` method and if the `PySpark` library is available in the environment. ## Next Steps For more information about how to connect, see the [Login API][hopsworks.login]. Or continue with the Data Source guide to import your own data to the Feature Store. ================================================================================ # Apache Beam Source: https://docs.hopsworks.ai/latest/user_guides/integrations/beam/ # Apache Beam Dataflow Runner Connecting to the Feature Store from an Apache Beam Dataflow Runner, requires configuring the Hopsworks certificates. For this in your Beam Java application `pom.xml` file include following snippet: ```xml java.io.tmpdir **/*.jks ``` ## Generating an API Key For instructions on how to generate an API key follow this [user guide](../projects/api_key/create_api_key.md). For the Beam integration to work correctly make sure you add the following scopes to your API key: 1. featurestore 2. project 3. job 4. kafka ## Connecting to the Feature Store You are now ready to connect to the Hopsworks Feature Store from Beam: ```Java //Establish connection with Hopsworks. HopsworksConnection hopsworksConnection = HopsworksConnection.builder() .host("my_instance") // DNS of your Feature Store instance .port(443) // Port to reach your Hopsworks instance, defaults to 443 .project("my_project") // Name of your Hopsworks Feature Store project .apiKeyValue("api_key") // The API key to authenticate with the feature store .hostnameVerification(false) // Disable for self-signed certificates .build(); //get feature store handle FeatureStore fs = hopsworksConnection.getFeatureStore(); ``` ## Next Steps For more information and how to integrate Beam feature pipeline to the Hopsworks Feature store follow the [tutorial](https://github.com/logicalclocks/hopsworks-tutorials/tree/master/integrations/java/beam). ================================================================================ # Java Source: https://docs.hopsworks.ai/latest/user_guides/integrations/java/ # Java client This guide explains step by step how to connect to Hopsworks from a Java client. ## Generate an API key For instructions on how to generate an API key follow this [user guide](../projects/api_key/create_api_key.md). For the Java client to work correctly make sure you add the following scopes to your API key: 1. featurestore 2. project 3. job 4. kafka ## Connecting to the Feature Store You are now ready to connect to the Hopsworks Feature Store from a Java client: ```Java //Import necessary classes import com.logicalclocks.hsfs.FeatureStore; import com.logicalclocks.hsfs.FeatureView; import com.logicalclocks.hsfs.HopsworksConnection; //Establish connection with Hopsworks. HopsworksConnection hopsworksConnection = HopsworksConnection.builder() .host("my_instance") // DNS of your Feature Store instance .port(443) // Port to reach your Hopsworks instance, defaults to 443 .project("my_project") // Name of your Hopsworks Feature Store project .apiKeyValue("api_key") // The API key to authenticate with the feature store .hostnameVerification(false) // Disable for self-signed certificates .build(); //get feature store handle FeatureStore fs = hopsworksConnection.getFeatureStore(); //get feature view handle FeatureView fv = fs.getFeatureView(fvName, fvVersion); // get feature vector List singleVector = fv.getFeatureVector(new HashMap() {{ put("id", 100); }}); ``` ## Next Steps For more information how to interact from Java client with the Hopsworks Feature store follow this [tutorial](https://github.com/logicalclocks/hopsworks-tutorials/tree/java_engine/java). ================================================================================ # Create Schema Source: https://docs.hopsworks.ai/latest/user_guides/projects/kafka/create_schema/ # How To Create A Kafka Schema ## Introduction ## Code In this guide, you will learn how to create a Kafka Avro Schema in the Hopsworks Schema Registry. ### Step 1: Get the Kafka API ```python import hopsworks project = hopsworks.login() kafka_api = project.get_kafka_api() ``` ### Step 2: Define the schema Define the Avro Schema, see [types](https://avro.apache.org/docs/current/spec.html#schema_primitive) for the format of the schema. ```python schema = { "type": "record", "name": "tutorial", "fields": [ {"name": "id", "type": "int"}, {"name": "data", "type": "string"}, ], } ``` ### Step 3: Create the schema Create the schema in the Schema Registry. ```python SCHEMA_NAME = "schema_example" my_schema = kafka_api.create_schema(SCHEMA_NAME, schema) ``` !!! api "API reference" - [`KafkaApi`][hopsworks_common.core.kafka_api.KafkaApi] - [`create_schema`][hopsworks_common.core.kafka_api.KafkaApi.create_schema] - [`KafkaSchema`][hopsworks_common.kafka_schema.KafkaSchema] Browse the full Python API :material-arrow-right: ================================================================================ # Create Topic Source: https://docs.hopsworks.ai/latest/user_guides/projects/kafka/create_topic/ # How To Create A Kafka Topic ## Introduction A Topic is a queue to which records are stored and published. Producer applications write data to topics and consumer applications read from topics. ## Prerequisites This guide requires that you have 'Data owner' role and have previously created a [Kafka Schema](create_schema.md) to be used for the topic. ## Code In this guide, you will learn how to create a Kafka Topic. ### Step 1: Get the Kafka API ```python import hopsworks project = hopsworks.login() kafka_api = project.get_kafka_api() ``` ### Step 2: Define the schema ```python TOPIC_NAME = "topic_example" SCHEMA_NAME = "schema_example" my_topic = kafka_api.create_topic( TOPIC_NAME, SCHEMA_NAME, 1, replicas=1, partitions=1 ) ``` !!! api "API reference" - [`KafkaApi`][hopsworks_common.core.kafka_api.KafkaApi] - [`create_topic`][hopsworks_common.core.kafka_api.KafkaApi.create_topic] - [`KafkaTopic`][hopsworks_common.kafka_topic.KafkaTopic] Browse the full Python API :material-arrow-right: ================================================================================ # Produce messages Source: https://docs.hopsworks.ai/latest/user_guides/projects/kafka/produce_messages/ # How To Produce To A Topic ## Introduction A Producer is a process which produces messages to a Kafka topic. In Hopsworks, only users with the 'Data owner' role are capable of performing the 'Write' action on Kafka topics within the project that they are a member of. ## Prerequisites This guide requires that you have 'Data owner' role and have previously created a [Kafka Topic](create_topic.md). ## Code In this guide, you will learn how to produce messages to a kafka topic. ### Step 1: Get the Kafka API ```python import hopsworks project = hopsworks.login() kafka_api = project.get_kafka_api() ``` ### Step 2: Configure confluent-kafka client ```python producer_config = kafka_api.get_default_config() from confluent_kafka import Producer producer = Producer(producer_config) ``` ### Step 3: Produce messages to topic ```python import json import uuid # Send a few messages for i in range(0, 10): producer.produce( "my_topic", json.dumps({"id": i, "data": str(uuid.uuid1())}), "key" ) # Trigger the sending of all messages to the brokers, 10 sec timeout producer.flush(10) ``` !!! api "API reference" - [`KafkaApi`][hopsworks_common.core.kafka_api.KafkaApi] - [`get_default_config`][hopsworks_common.core.kafka_api.KafkaApi.get_default_config] - [`KafkaTopic`][hopsworks_common.kafka_topic.KafkaTopic] Browse the full Python API :material-arrow-right: ## Going Further Now you can create a [Consumer](consume_messages.md) to read the messages from the topic. ================================================================================ # Consume messages Source: https://docs.hopsworks.ai/latest/user_guides/projects/kafka/consume_messages/ # How To Consume Message From A Topic ## Introduction A Consumer is a process which reads messages from a Kafka topic. In Hopsworks, all user roles are capable of performing 'Read' and 'Describe' actions on Kafka topics within projects that they are a member of or are shared with them. ## Prerequisites This guide requires that you have previously [produced](produce_messages.md) messages to a kafka topic. ## Code In this guide, you will learn how to consume messages from a kafka topic. ### Step 1: Get the Kafka API ```python import hopsworks project = hopsworks.login() kafka_api = project.get_kafka_api() ``` ### Step 2: Configure confluent-kafka client ```python consumer_config = kafka_api.get_default_config() consumer_config["default.topic.config"] = {"auto.offset.reset": "earliest"} from confluent_kafka import Consumer consumer = Consumer(consumer_config) ``` ### Step 3: Consume messages from a topic ```python # Subscribe to topic consumer.subscribe(["my_topic"]) for i in range(0, 10): msg = consumer.poll(timeout=10.0) print(msg.value()) ``` !!! api "API reference" - [`KafkaApi`][hopsworks_common.core.kafka_api.KafkaApi] - [`get_default_config`][hopsworks_common.core.kafka_api.KafkaApi.get_default_config] - [`KafkaTopic`][hopsworks_common.kafka_topic.KafkaTopic] Browse the full Python API :material-arrow-right: ================================================================================ # Sharing Source: https://docs.hopsworks.ai/latest/user_guides/fs/sharing/sharing/ # Sharing ## Introduction Hopsworks allows artifacts (such as feature groups and feature views) to be shared between projects. There are two main use cases for sharing features: - **Cross-team collaboration:** When multiple teams work on the same Hopsworks deployment, each team typically has its own set of projects. If team A wants to leverage features built by team B, team B can share their feature groups with team A's project. - **Environment isolation:** By creating separate projects for different stages of the development lifecycle (development, testing, and production), you can ensure that changes in the development project don't impact production features. At the same time, you can share production features to use them when developing new models or additional features. ## Seeing what is shared Open the project that owns the feature store. In `Project Settings`, the `Feature store sharing` section lists everything the feature store shares, and a sentence at the top summarizes it. Only data owners can share, unshare and revoke, so only they get the `Share` button and the trash icons. Observers see the same lists read-only, and other members see a note instead, because only data owners and observers can read a project's sharing. The `Projects` tab lists every project the feature store is shared with. A project marked `Entire feature store` can read every feature group, including feature groups created after the share. Any other project lists the feature groups shared with it, with either `all features` or the features it can read.

The Projects tab of the Feature store sharing section, with one project sharing the entire feature store and one with two feature groups
Feature store sharing section in Project Settings

The `Restricted users` tab lists every member with the [Feature store restricted][feature-store-restricted] role who has been granted access, with the feature groups and features each of them can read.

The Restricted users tab of the Feature store sharing section, with one user granted two feature groups
Restricted users granted access to feature groups

Each feature group name links to that feature group's `Sharing` tab. ## Sharing the entire feature store You can share your project's entire feature store with another project, granting read-only access to all feature groups. ### Step 1: Open the share dialog In `Project Settings`, click `Share` in the `Feature store sharing` section. ### Step 2: Share the feature store In `With a project`, select the target project, keep `The entire feature store, including feature groups created later` selected, and click `Share`.

The share dialog with a target project selected and the entire feature store option checked
Share the entire feature store

!!! note "Read-only access" Shared feature stores are always read-only. Members of the target project cannot modify any data in the shared feature store. A project holds either the entire feature store or a set of individual feature groups, not both, because the entire feature store already includes every feature group. A project that already has feature groups shared with it is listed in the dialog but cannot be selected; unshare those feature groups first. After clicking `Share`, the project appears in the `Projects` tab marked `Entire feature store`. ### Using the API to share the feature store ```python import hopsworks project = hopsworks.login() fs = project.get_feature_store() # Share the whole feature store fs.share("target_project") # List projects it's shared with for share in fs.shared_with(): print(share["sharedWithProject"]["name"], share["sharedOn"]) # Revoke a share fs.unshare("target_project") ``` ## Sharing a feature group with selected features For more granular control, you can share individual feature groups and select which features to expose. This allows you to share specific data without granting access to your entire feature store. ### Step 1: Navigate to the feature group In the `Feature Groups` section, select the feature group you want to share and click the `Sharing` tab.

The Sharing tab of a feature group that is not shared yet, with the Share and Grant access buttons
Feature group sharing tab

You can also start from `Project Settings`: click `Share` in the `Feature store sharing` section, select `One feature group, entirely or some of its features`, and pick the feature group. ### Step 2: Share the feature group You can share with either a project or an individual user with the [Feature store restricted][feature-store-restricted] role. Under `Features to share`, all features are selected; clear the ones the target should not read. The primary key and the event time are always shared, because the target needs them to join the feature group and to make point-in-time correct reads. === "Share with a project" Click `Share`, select the target project in `With a project`, choose the features, and click `Share`.

The share dialog of a feature group with a target project selected and one feature cleared
Share a feature group with a project

The project appears under `Shared with projects` on the feature group's `Sharing` tab, and in the `Projects` tab of `Project Settings`.

The Shared with projects card listing one project that can read all features
Projects the feature group is shared with

=== "Share with a user" Click `Grant access`, select the user in `With a user`, choose the features, and click `Share`. The list offers the project members with the [Feature store restricted][feature-store-restricted] role.

The share dialog of a feature group with a restricted user selected and one feature cleared
Grant a restricted user access to a feature group

The user appears under `Users with restricted access` on the feature group's `Sharing` tab, and in the `Restricted users` tab of `Project Settings`.

The Users with restricted access card listing one user and the features granted to them
Users with restricted access to the feature group

### Using the API to share a feature group === "Share with a project" ```python fg = fs.get_feature_group("feature_group_name", version=1) # Share the whole feature group fg.share("target_project") # Or share only selected columns (primary key + event time are always included) fg.share("target_project", features=["amount", "country"]) # List projects it's shared with for share in fg.shared_with(): print(share["sharedWithProject"]["name"], share["sharedEntirely"]) # Revoke a share fg.unshare("target_project") ``` === "Share with a user" A [Feature store restricted][feature-store-restricted] member has no feature store access by default; access must be granted per feature group (or per feature), from within their own project. ```python fg = fs.get_feature_group("feature_group_name", version=1) # Grant access to the whole feature group fg.grant_restricted_access("restricted_user@example.com") # Or grant access to only selected columns fg.grant_restricted_access("restricted_user@example.com", features=["amount"]) # List who has been granted access for grant in fg.get_restricted_access(): print(grant["grantedToUser"], grant["grantedEntirely"]) # Revoke access fg.revoke_restricted_access("restricted_user@example.com") ``` ## Unsharing In the `Feature store sharing` section of `Project Settings`, click the trash icon next to what you want to stop sharing: - Next to a project marked `Entire feature store`, to unshare the entire feature store from that project. - Next to a feature group in the `Projects` tab, to unshare that feature group from that project. - Next to a feature group in the `Restricted users` tab, to revoke that user's access to it. Feature views and training data built on the unshared data stop working for the project or user that loses access, so the dialog asks you to type `confirm` first. The same trash icons are on a feature group's `Sharing` tab. From the API, use `fs.unshare`, `fg.unshare` and `fg.revoke_restricted_access` as shown above. ## Using shared features Once features have been shared with your project, you can access them through the UI or the API. ### Using the UI Navigate to the project that has access to shared features. In the `Feature Groups` section, use the dropdown in the upper right corner to select which feature store to view.

View shared feature groups
Selecting a shared feature store in the UI

### Using the API To access features from a shared feature store programmatically, retrieve the handle for the shared feature store using the Hopsworks API. #### Step 1: Get feature store handles Use the `get_feature_store()` method with the name of the shared feature store: ```python import hopsworks project = hopsworks.login() # Get your project's feature store project_feature_store = project.get_feature_store() # Get the shared feature store by name shared_feature_store = project.get_feature_store(name="name_of_shared_feature_store") ``` #### Step 2: Fetch feature groups ```python # Fetch a feature group from the shared feature store shared_fg = shared_feature_store.get_feature_group( name="shared_fg_name", version=1 ) # Fetch a feature group from your project's feature store fg = project_feature_store.get_or_create_feature_group( name="feature_group_name", version=1 ) ``` ================================================================================ # Tags Source: https://docs.hopsworks.ai/latest/user_guides/fs/tags/tags/ # Tags { #tags-guide } ## Introduction Hopsworks feature store enables users to attach tags to artifacts, such as feature groups, feature views, training datasets, jobs, apps, models or deployments. A tag is a `{key: value}` pair which provides additional information about the data managed by Hopsworks. Tags allow you to design custom metadata for your artifacts. For example, you could design a tag schema that encodes governance rules for your feature store, such as classifying data as personally identifiable, defining a data retention period for the data, and defining who signed off on the creation of some feature. ## Prerequisites Tags have a schema. Before you can attach a tag to an artifact and fill in the tag values, you first need to select an existing tag schema or create a new tag schema. Tag schemas can be defined by Hopsworks administrator in the `Cluster settings` section of the platform. Schemas are defined globally across all projects. When users attach tags to an artifact, the tag will be validated against a specific schema. This allows tags to be consistent no matter the project or the team generating them. !!! warning "Immutable" Tag schemas are immutable. Once defined, a tag schema cannot be edited nor deleted. ## Step 1: Define a tag schema Tag schemas can be defined using the UI wizard in the `Cluster settings` > `Tag schemas` section. Tag schemas have a name, the name is used to uniquely identify the schema. You can also provide an optional description. You can define a schema by using the UI tool or by providing the schema in JSON format. If you use the UI tool, you should provide the name of the property in the schema, the type of the property, whether or not the property is required and an optional description.

UI tag schema definition
UI tag schema definition

The UI tool allows you to define simple not-nested schemas. For more advanced use cases, more complex schemas (e.g., nested schemas) might be required to fully express the content of a given artifact. In such cases it is possible to provide the schema directly as JSON string. The JSON should follow the standard [https://json-schema.org](https://json-schema.org). An example of complex schema is the following: ```json { "type" : "object", "properties" : { "first_name" : { "type" : "string" }, "last_name" : { "type" : "string" }, "age" : { "type" : "integer" }, "hobbies" : { "type" : "array", "items" : { "type" : "string" } } }, "required" : ["first_name", "last_name", "age"], "additionalProperties": false } ``` Additionally it is also possible to define a single property as tag. You can achieve this by defining a JSON schema like the following: ```json { "type" : "string" } ``` Where the type is a valid primitive type: `string`, `boolean`, `integer`, `number`. ### Archiving a tag for analytics Most tags are only ever read as they are now: who owns this feature group, whether it holds PII. Some are interesting over time, and for those the current value is the least useful part. Marking a schema as archived says that attachments of this tag are worth keeping once they stop being current, so the tag's history can be analysed and not just its present state. Tick `Archive tag history` when defining the schema, or pass `archive=True` through the API: === "Python" ```python from hopsworks.core.tag_schemas_api import TagSchemasApi schema = { "type": "object", "properties": {"state": {"type": "string"}}, "required": ["state"], "additionalProperties": False, } TagSchemasApi().create("asset_lifecycle", schema, archive=True) ``` The flag is a property of the schema rather than of any one attachment, which is why it is set where the schema is defined and applies to every tag attached with it afterwards. It defaults to `False`, which discards an attachment once it stops being current. Registering a schema requires administrator privileges, as it does without the flag. #### What it is for Take an `asset_lifecycle` tag whose `state` is `dev`, `qa` or `prod`. Every artifact carries it, and the value moves forward as the artifact is promoted. Read as an ordinary tag it answers one question, which is where an artifact is now. The questions worth asking are about the pipeline rather than the artifact: how long does something sit in `qa` before it reaches `prod`, is that getting slower, whose artifacts stall. Those become answerable once the tag's history is kept as one record per value, each with the time that value became current. The analysis is then a group-by over the artifact: order its `dev`, `qa` and `prod` records by time, and the gaps between them are how long it spent in each stage. Aggregated across every artifact, that gives the promotion times for the deployment as a whole and how they are moving. Without archiving, promoting an artifact to `prod` discards the record that it was ever in `qa`, and the question stops being answerable at all. So the flag is worth setting on a tag whose values are states an artifact passes through, rather than facts about it. This history is not the same thing as the [attachment time][when-a-tag-was-attached] on the live tag. That timestamp deliberately stays at the first attachment when a value is corrected, so it records when an artifact was first classified and not when it entered its current state. The per-value history is what the archive is for. !!! note "Set it before you need it" Setting `archive` makes Hopsworks record every change to the tag's values, which is what makes the analysis above possible. See [Archive tag history][archive-tag-history] for what gets recorded, how to read it, and how to turn it on for a schema that already exists. Set it on the schemas whose history you expect to want. The flag cannot reconstruct changes that happened while it was off, because the live tag keeps only its current value. ## Step 2: Attach a tag to an artifact Once the tag schema has been created, you can attach a tag with that schema to a feature group, feature view, training dataset, job, app, model or deployment either using the APIs, or by using the UI. ### Using the API You can attach tags to feature groups and feature views by using the `add_tag()` method of the feature store APIs: === "Python" ```python # Retrieve the feature group fg = fs.get_feature_group("transactions_4h_aggs_fraud_batch_fg", version=1) # Define the tag tag = { "business_unit": "Fraud", "data_owner": "email@hopsworks.ai", "pii": True, } # Attach the tag fg.add_tag("data_privacy", tag) ``` You can see the list of tags attached to a given artifact by using the `get_tags()` method: === "Python" ```python # Retrieve the feature group fg = fs.get_feature_group("transactions_4h_aggs_fraud_batch_fg", version=1) # Retrieve the tags for this feature group fg.get_tags() ``` Finally you can remove a tag from a given artifact by calling the `delete_tag()` method: === "Python" ```python # Retrieve the feature group fg = fs.get_feature_group("transactions_4h_aggs_fraud_batch_fg", version=1) # Retrieve the tags for this feature group fg.delete_tag("data_privacy") ``` The same APIs work for feature views, training datasets, models and deployments alike. #### Jobs and apps Jobs and apps carry tags through the same three methods, reached from the job handle rather than the feature store: === "Python" ```python jobs_api = project.get_jobs_api() job = jobs_api.get_job("transactions_ingest") job.add_tag("data_privacy", {"business_unit": "Fraud", "pii": True}) job.get_tags() job.delete_tag("data_privacy") ``` An app is a job whose type is `PYTHON_APP`, so an app is tagged exactly the same way, through the handle its name resolves to. The CLI covers the same three operations: ```bash hops job tags transactions_ingest hops job add-tag transactions_ingest data_privacy --value '{"business_unit": "Fraud", "pii": true}' hops job remove-tag transactions_ingest data_privacy ``` `--value` takes JSON for a schema with properties, or a plain string for a single-property schema. Tags on a job are also editable in the UI, on the job's `Tags` section, and when creating or editing the job. For an app, the equivalent section is on the app overview page. Deleting a job deletes its tags with it. They are not restored by creating a new job under the same name, because the tags belong to the job that was deleted and not to its name. ### Using the UI You can attach tags to feature groups and feature views directly from the UI. You can navigate on the artifact page and click on the `Add tags` button. From there you can select the tag schema of the tag you want to attach and populate the values as shown in the gif below.

Attach tag to a feature group
Attach tag to a feature group

## When a tag was attached Hopsworks records the time each tag was attached and reports it alongside the value. `get_tags()` returns values only, so read the attachment time through the `_metadata` variants, which return `Tag` objects instead of bare values: === "Python" ```python fg = fs.get_feature_group("transactions_4h_aggs_fraud_batch_fg", version=1) tag = fg.get_tag_metadata("data_privacy") print(tag.value, tag.created_on) # every tag on the artifact, keyed by name for name, attached in fg.get_tags_metadata().items(): print(name, attached.created_on) ``` `created_on` is an aware UTC `datetime`, and the same methods exist on feature views, training datasets and jobs. The timestamp records when the tag was **attached**, not when its value last changed. Re-attaching a tag to change its value keeps the original attachment time, so the value can be corrected without losing the record of when the artifact was first classified. `created_on` is `None` when the attachment time is unknown rather than recent. That happens for tags attached before the cluster recorded attachment times, and for legacy per-file dataset tags, which are stored as HopsFS extended attributes and carry no timestamp. ## Archive tag history By default a tag records only its current value: reading it tells you where an artifact is now, not where it has been. Turning on **Archive tag history** for a schema makes Hopsworks additionally record every change to that tag's values, so you can ask how long an artifact spent in each state. The flag is set per schema, in `Cluster settings` > `Tag schemas`, when the schema is created. To turn it on or off for a schema that already exists, a cluster administrator calls `PUT /hopsworks-api/api/tags/{name}/archive?value=true`. It applies to every artifact the tag is attached to: feature groups, feature views, training datasets, jobs, models and deployments. Turning it off ends every interval the tag still has open, at the moment you turn it off, and keeps everything already recorded. The recorded rows stay because they are still true; the open intervals have to be ended because nothing would ever end them once recording stops, and the last value of every artifact would otherwise read as current forever. Turning it back on starts a fresh interval at that moment rather than pretending the gap was observed. Two things are worth knowing before you turn it on: - **History starts when you turn it on.** Changes made before that are not recoverable, because the live tag keeps only its current value. Attachments that already exist are backfilled with the state they are in, timed from when they were attached. That start is the attachment time and not the moment you turned archiving on, so the first interval of an attachment that already existed covers time that was never observed, and it credits the current value with all of it. A tag attached as `dev` in January, changed to `prod` in February with nothing recording, and archived in March reports `prod` as current since January. The figure is therefore an upper bound on how long that state has really held, never a lower one, and it raises the average time-in-state of any report that includes it. - **Turning it off stops recording but keeps what was recorded.** The rows already written are still true, and the tag is still attached, so nothing is deleted. History is recorded per key of the schema, not per tag. Changing one key of a multi-key tag records a change to that key alone and leaves the others untouched, so a correction to one field does not make every other field look like it changed at the same moment. ### Turning it on for a schema archived before the history existed Releases before 5.2 accepted `archive` when a schema was created, and stored it, but recorded nothing: the flag had no reader. A schema created with it on before upgrading therefore has the flag set and no history, and the upgrade does not start one. The baseline is written when the flag is set, and an upgrade sets nothing, so such a schema stays silent until each artifact's tag next changes, and the state it held before that change is gone. After upgrading, a cluster administrator turns it on once more for each schema that already had it, with the same call used to turn it on for any existing schema, `PUT /hopsworks-api/api/tags/{name}/archive?value=true`. `GET /hopsworks-api/api/tags` lists the schemas with their `archive` flag, which is how to find the ones to repeat it for. Repeating the call costs nothing on a schema that is already recording. The backfill works per key, not per attachment: it opens only the keys that do not already have an open interval, so a schema part-way through is completed rather than duplicated, and one that is fully recorded gets no new rows. A key whose last recorded event closed an interval is opened again at this point, since nothing is known about the stretch when recording was off. On a widely attached schema the call can be refused rather than run. It counts the events the backfill would write first, one per tag key of every attachment, and refuses above `tag_history_archive_max_events`, which defaults to 20000. The error names the count and the limit. Raising it is an administrator decision that belongs with NDB's `MaxNoOfConcurrentOperations`, because the backfill is one transaction and is bounded by both. ### Reading the history The history is stored in the `tag_history` table of the Hopsworks metadata database, one row per transition: the value became current, or it stopped being current. A value change writes both at the same instant, so one interval's end is the next one's start. It is read with SQL rather than through the tag APIs, which continue to return the current value. On a cluster with the `hopsworks_analytics` project enabled, its Superset connection can query the table directly, and `create_tag_history_dashboard.py` in the `okr-dashboards` repository builds a "Tag Lifecycle" dashboard over it: time spent in each state, whether that is increasing, what is in each state now, and what has been in one state longest. The table stores events rather than intervals. To get `added_on` and `removed_at`, take the next event's time for the same artifact, tag and key: ```sql SELECT e.artifact_type, e.artifact_id, e.tag_name, e.tag_key, e.tag_value, e.added_on, e.removed_at FROM ( SELECT artifact_type, artifact_id, tag_name, tag_key, tag_value, event_type, event_time AS added_on, LEAD(event_time) OVER ( PARTITION BY artifact_type, artifact_id, tag_name, tag_key ORDER BY event_time, id ) AS removed_at FROM hopsworks.tag_history ) e WHERE e.event_type = 'OPENED' ``` Two details in that query are easy to get wrong and produce numbers that look reasonable: - The `OPENED` filter has to be in the outer query. SQL applies `WHERE` before window functions, so filtering inside would hide every `CLOSED` row from `LEAD`, and anything that ended without a successor, a detached tag or a deleted artifact, would report as still current with its duration growing forever. - The tie-break has to be `id`, and not the event type. Both halves of a value change share one `event_time` by design, so the ordering needs one. `id` is the insertion order and the writer emits a change as `CLOSED` then `OPENED` inside one transaction, so `id` already puts the two halves in the order they happened. Forcing `CLOSED` first instead breaks an attach and a detach that share a millisecond: it orders that `CLOSED` ahead of the `OPENED` it followed, `LEAD` leaves the `OPENED` with no `removed_at`, and a tag that was removed reads as current from then on. One case is still open. RonDB allocates `id` per SQL node, so two events written a millisecond apart through different nodes can order arbitrarily with respect to each other. Ordering those exactly needs a per-key sequence rather than a tie-break. A `removed_at` of `NULL` means the artifact is still in that state. An `added_on` of `NULL` means the tag was attached before Hopsworks began recording attachment times, so the start is unknown; it is left empty rather than filled with a guess. The history outlives what it describes. Deleting the artifact, the tag schema or the project closes the open intervals and keeps the rows, so a report over a past quarter still returns what was true then. ## Step 3: Search Hopsworks indexes the tags attached to feature groups, feature views, training datasets, jobs, models and deployments. The tags are then searchable using the free text search box located at the top of the UI, and can be filtered on directly. See the [tag and keyword search guide][search-with-tags-and-keywords] for filtering by a specific tag key and value rather than by free text. Tags on artifacts of every indexed class are searchable, so a governance question such as "which artifacts are missing a data owner" can be answered across feature groups, jobs and deployments in one query.

Search for tags in the feature store
Search for tags in the feature store

================================================================================ # Mandatory Tags Source: https://docs.hopsworks.ai/latest/user_guides/fs/tags/mandatory_tags/ # Mandatory Tags ## Introduction Mandatory tags let a Hopsworks administrator require that specific tag schemas are populated on artifacts. They build on top of [tags](tags.md) and are used to enforce governance rules, such as requiring every model to declare a data owner. A mandatory tag is a tag schema that has been marked as required for one or more artifact types. The supported artifact types are feature groups, feature views, training datasets, models and deployments. ## Prerequisites A mandatory tag references an existing tag schema. Define the tag schema first, as described in the [Tags](tags.md) guide, before marking it mandatory. Only administrators can configure mandatory tags. Attaching the tag values afterwards is done by any project member with write access to the artifact. ## Configure mandatory tags Mandatory tags are configured in two scopes. - Cluster-wide mandatory tags apply to every project on the cluster. They are configured by an administrator in the `Cluster settings` > `Tag schemas` section. Configure cluster-wide mandatory tags under Cluster settings, Tag Schemas - Project-specific mandatory tags apply only to a single project. They are configured by an administrator in the `Cluster settings` > `Projects` section, by selecting the project. Configure project-specific mandatory tags under Cluster settings, Projects, by editing a project For each mandatory tag you select the artifact types it applies to. A tag schema can be mandatory for any combination of feature groups, feature views, training datasets, models and deployments. For example, a `data_owner` schema can be marked mandatory for models and deployments only, leaving feature groups, feature views and training datasets unaffected. ## Enforcement per artifact type All five artifact types, feature groups, feature views, training datasets, models and deployments, enforce mandatory tags the same way. The create request is validated against the configured mandatory tags. If any mandatory tag is missing from the tags provided at creation, the artifact is not created and the request is rejected with an HTTP 400 error that lists the missing tag names. The bulk tag endpoint applies the same validation. An artifact created before a tag was marked mandatory stays valid and is not deleted. The missing tag is surfaced on read through the `missing_mandatory_tags` property, described in [Missing mandatory tags on pre-existing artifacts](#missing-mandatory-tags-on-pre-existing-artifacts), and on the artifact page in the UI. ## Attach mandatory tags at creation Pass the mandatory tag values in the `tags` argument of the create call so the create request carries them and passes validation. The `tags` argument takes the same shape as feature group tags: a `{"name": ..., "value": ...}` dictionary, a list of such dictionaries, or `Tag` objects, where the value is a single primitive or a dictionary matching the tag schema. === "Feature Group (Python)" ```python fs = project.get_feature_store() # data_owner is mandatory for feature groups; pass it at creation fg = fs.create_feature_group( name="transactions", version=1, primary_key=["id"], tags=[{"name": "data_owner", "value": "email@hopsworks.ai"}], ) fg.insert(df) ``` === "Feature View (Python)" ```python fs = project.get_feature_store() fg = fs.get_feature_group("transactions", version=1) # data_owner is mandatory for feature views; pass it at creation fv = fs.create_feature_view( name="transactions_fv", version=1, query=fg.select_all(), tags=[{"name": "data_owner", "value": "email@hopsworks.ai"}], ) ``` === "Training Dataset (Python)" ```python fs = project.get_feature_store() fv = fs.get_feature_view("transactions_fv", version=1) # data_owner is mandatory for training datasets; pass it at creation td_version, td_job = fv.create_training_data( tags=[{"name": "data_owner", "value": "email@hopsworks.ai"}], ) ``` === "Model (Python)" ```python mr = project.get_model_registry() # data_owner is mandatory for models; pass it at creation model = mr.python.create_model( name="fraud_model", metrics={"accuracy": 0.94}, tags=[{"name": "data_owner", "value": "email@hopsworks.ai"}], ) model.save("/path/to/model_artifacts") # Deploy the model; pass the mandatory deployment tag on the deploy call deployment = model.deploy( name="fraudmodeldeployment", tags=[{"name": "data_owner", "value": "email@hopsworks.ai"}], ) ``` === "Deployment (Python)" ```python ms = project.get_model_serving() # Build a predictor for an already-saved model predictor = ms.create_predictor(model) # data_owner is mandatory for deployments; pass it at creation deployment = ms.create_deployment( predictor, name="fraudmodeldeployment", tags=[{"name": "data_owner", "value": "email@hopsworks.ai"}], ) deployment.save() ``` Omitting a mandatory tag from the `tags` argument rejects the create request with an HTTP 400 error that lists the missing tag names. ## Missing mandatory tags on pre-existing artifacts Marking a tag mandatory does not retroactively reject artifacts that already exist without it. An artifact created before the tag became mandatory stays valid, and Hopsworks surfaces the gap on read rather than deleting the artifact. Fetching such an artifact emits a Python `UserWarning` listing the missing tag names. The same list is available on the fetched object through the `missing_mandatory_tags` property and is shown on the artifact page in the UI. The property is populated when the object is fetched from the backend and reflects the state at fetch time. Adding a tag does not update the property on the object in hand, so fetch the artifact again to see the updated list. === "Model (Python)" ```python mr = project.get_model_registry() # If data_owner is mandatory but was never set on this model, # the fetch emits: UserWarning: Missing mandatory tags: ['data_owner'] model = mr.get_model("fraud_model", version=1) # Names of mandatory tags that are required for this model but not yet set missing = [tag["name"] for tag in model.missing_mandatory_tags] # Set the missing tag if "data_owner" in missing: model.add_tag("data_owner", "email@hopsworks.ai") # missing_mandatory_tags reflects the state at fetch time, # so fetch the model again to see the updated list (no warning now) model = mr.get_model("fraud_model", version=1) print(model.missing_mandatory_tags) # [] ``` === "Deployment (Python)" ```python ms = project.get_model_serving() # If data_owner is mandatory but was never set on this deployment, # the fetch emits: UserWarning: Missing mandatory tags: ['data_owner'] deployment = ms.get_deployment("fraudmodeldeployment") missing = [tag["name"] for tag in deployment.missing_mandatory_tags] # Set the missing tag if "data_owner" in missing: deployment.add_tag("data_owner", "email@hopsworks.ai") # Fetch the deployment again to see the updated list (no warning now) deployment = ms.get_deployment("fraudmodeldeployment") print(deployment.missing_mandatory_tags) # [] ``` After the tag is set, it no longer appears in `missing_mandatory_tags` on the next fetch, and the fetch no longer warns. ================================================================================ # Keywords Source: https://docs.hopsworks.ai/latest/user_guides/fs/tags/keywords/ # Keywords { #keywords-guide } ## Introduction A keyword is a single free-form word attached to a feature group, a feature view or a training dataset. Keywords need no schema and no administrator: any project member with write access can invent one and attach it. That is the difference from [tags][tags-guide], and it decides which to reach for. A tag is validated against a schema and is the right tool for governance, where the set of allowed keys and values has to be agreed in advance. A keyword is the right tool for discovery, where the point is to label something now and find it later. Keywords apply to feature groups, feature views and training datasets only. Jobs, apps, models and deployments take tags but not keywords, so a keyword filter never matches them. ## Read the keywords of an artifact === "Python" ```python fg = fs.get_feature_group("transactions_4h_aggs_fraud_batch_fg", version=1) fg.get_keywords() # ['fraud', 'aggregations', 'hourly'] ``` To see when each keyword was attached, use `get_keywords_metadata()`, which returns a dict of keyword to attachment time: === "Python" ```python for keyword, attached in fg.get_keywords_metadata().items(): print(keyword, attached) ``` The attachment time is an aware UTC `datetime`, or `None` when it is unknown. It is `None` for keywords attached before the cluster began recording attachment times, so `None` does not mean the keyword is new. ## Add, replace and delete keywords Three methods change the keyword set, and they differ in what they do to the keywords already there. `add_keywords()` adds to the set and leaves the rest alone. It accepts one keyword or a list: === "Python" ```python fg.add_keywords("fraud") fg.add_keywords(["aggregations", "hourly"]) ``` `set_keywords()` replaces the whole set. Anything not in the list you pass is removed, so use it when you intend the artifact to end up with exactly these keywords and nothing else: === "Python" ```python fg.set_keywords(["fraud", "hourly"]) ``` `delete_keyword()` removes a single keyword: === "Python" ```python fg.delete_keyword("hourly") ``` All three return the resulting keyword set, so a read-back is not needed to see the effect. The same methods exist on feature views. A training dataset's keywords are reached through its feature view, because a training dataset is identified by the feature view it was created from plus its own version: === "Python" ```python fv = fs.get_feature_view("fraud_detection", version=1) fv.add_training_dataset_keywords(1, "baseline") fv.get_training_dataset_keywords(1) fv.delete_training_dataset_keyword(1, "baseline") ``` ## The cluster vocabulary Keywords are free-form, which makes them prone to near-duplicates: `fraud`, `Fraud` and `fraud_detection` are three separate keywords that fragment the same idea. To let you reuse a word someone has already chosen, the feature store can list every keyword in use: === "Python" ```python fs.get_all_keywords() ``` The vocabulary is cluster-wide rather than project-scoped, so it shows words in use in projects you are not a member of. Only the words are returned, never which artifact or project they came from, so this discloses no artifact you could not otherwise see. ## Command line The CLI covers the same operations: ```bash # feature group keywords hops fg keywords transactions_fg --version 1 hops fg add-keyword transactions_fg fraud --version 1 hops fg remove-keyword transactions_fg fraud --version 1 # feature view keywords hops fv keywords fraud_detection --version 1 hops fv add-keyword fraud_detection baseline --version 1 hops fv remove-keyword fraud_detection baseline --version 1 # training dataset keywords, addressed by feature view name and td version hops td keywords fraud_detection 1 --fv-version 1 hops td add-keyword fraud_detection 1 baseline --fv-version 1 hops td remove-keyword fraud_detection 1 baseline --fv-version 1 ``` The listing commands show the attachment time next to each keyword. !!! warning "The `*-keyword` commands changed meaning" They used to operate on tags, which are name and value pairs. They now operate on keywords, which are plain labels, and tags moved to `hops fg tags`, `hops fg add-tag` and `hops fg remove-tag`. A script that passed a tag value to `add-keyword` needs to move to `add-tag`, because a keyword has no value to pass. The commands print this notice to stderr when they run. ## Searching by keyword Keywords are indexed, and can be filtered on directly rather than only matched as free text. See the [tag and keyword search guide][search-with-tags-and-keywords]. ================================================================================ # Provenance Source: https://docs.hopsworks.ai/latest/user_guides/fs/provenance/provenance/ # Provenance ## Introduction Hopsworks allows users to track provenance (lineage) between: - data sources - feature groups - feature views - training datasets - models In the provenance pages we will call a provenance artifact or shortly artifact, any of the five entities above. With the following provenance graph: ```plaintext data source -> feature group -> feature group -> feature view -> training dataset -> model ``` we will call the parent, the artifact to the left, and the child, the artifact to the right. So a feature view has a number of feature groups as parents and can have a number of training datasets as children. Tracking provenance allows users to determine where and if an artifact is being used. You can track, for example, if feature groups are being used to create additional (derived) feature groups or feature views, or if their data is eventually used to train models. You can interact with the provenance graph using the UI or the APIs. ## Step 1: Data Source lineage The relationship between data sources and feature groups is captured automatically when you create an external feature group. You can inspect the relationship between data sources and feature groups using the APIs. === "Python" ```python # Retrieve the data source ds = fs.get_data_source("snowflake_sc") ds.query = "SELECT * FROM USER_PROFILES" # Create the user profiles feature group user_profiles_fg = fs.create_external_feature_group( name="user_profiles", version=1, data_source=ds ) user_profiles_fg.save() ``` ### Step 1, Using Python Starting from a feature group metadata object, you can traverse upstream the provenance graph to retrieve the metadata objects of the data sources that are part of the feature group. To do so, you can use the [`FeatureGroup.get_data_source_provenance`][hsfs.feature_group.FeatureGroup.get_data_source_provenance] method. === "Python" ```python # Returns all data sources linked to the provided feature group lineage = user_profiles_fg.get_data_source_provenance() # List all accessible parent data sources lineage.accessible # List all deleted parent data sources lineage.deleted # List all the inaccessible parent data sources lineage.inaccessible ``` === "Python" ```python # Returns an accessible data source linked to the feature group (if it exists) user_profiles_fg.get_data_source() ``` To traverse the provenance graph in the opposite direction (i.e., from the data source to the feature group), you can use the [`StorageConnector.get_feature_groups_provenance`][hsfs.storage_connector.StorageConnector.get_feature_groups_provenance] method. When navigating the provenance graph downstream, the `deleted` feature groups are not tracked by provenance, as such, the `deleted` property will always return an empty list. === "Python" ```python # Returns all feature groups linked to the provided data source lineage = snowflake_sc.get_feature_groups_provenance() # List all accessible downstream feature groups lineage.accessible # List all the inaccessible downstream feature groups lineage.inaccessible ``` === "Python" ```python # Returns all accessible feature groups linked to the data source (if any exists) snowflake_sc.get_feature_groups() ``` ## Step 2: Feature group lineage ### Assign parents to a feature group When creating a feature group, it is possible to specify a list of feature groups used to create the derived features. For example, you could have an external feature group defined over a Snowflake or Redshift table, which you use to compute the features and save them in a feature group. You can mark the external feature group as parent of the feature group you are creating by using the `parents` parameter in the [`FeatureStore.get_or_create_feature_group`][hsfs.feature_store.FeatureStore.get_or_create_feature_group] or [`FeatureStore.create_feature_group`][hsfs.feature_store.FeatureStore.create_feature_group] methods: === "Python" ```python # Retrieve the feature group profiles_fg = fs.get_external_feature_group("user_profiles", version=1) # Do feature engineering age_df = transaction_df.merge(profiles_fg.read(), on="cc_num", how="left") transaction_df["age_at_transaction"] = ( age_df["datetime"] - age_df["birthdate"] ) / np.timedelta64(1, "Y") # Create the transaction feature group transaction_fg = fs.get_or_create_feature_group( name="transaction_fraud_batch", version=1, description="Transaction features", primary_key=["cc_num"], event_time="datetime", parents=[profiles_fg], ) transaction_fg.insert(transaction_df) ``` Another example use case for derived feature group is if you have a feature group containing features with daily resolution and you are using the content of that feature group to populate a second feature group with monthly resolution: === "Python" ```python # Retrieve the feature group daily_transaction_fg = fs.get_feature_group("daily_transaction", version=1) daily_transaction_df = daily_transaction_fg.read() # Do feature engineering cc_group = ( daily_transaction_df[["cc_num", "amount", "datetime"]] .groupby("cc_num") .rolling("1M", on="datetime") ) monthly_transaction_df = pd.DataFrame(cc_group.mean()) # Create the transaction feature group monthly_transaction_fg = fs.get_or_create_feature_group( name="monthly_transaction_fraud_batch", version=1, description="Transaction features - monthly aggregates", primary_key=["cc_num"], event_time="datetime", parents=[daily_transaction_fg], ) monthly_transaction_fg.insert(monthly_transaction_df) ``` ### List feature group parents You can query the provenance graph of a feature group using the UI and the APIs. From the APIs you can list the parent feature groups by calling the method [`FeatureGroup.get_parent_feature_groups`][hsfs.feature_group.FeatureGroup.get_parent_feature_groups] === "Python" ```python lineage = transaction_fg.get_parent_feature_groups() # List all accessible parent feature groups lineage.accessible # List all deleted parent feature groups lineage.deleted # List all the inaccessible parent feature groups lineage.inaccessible ``` A parent is marked as `deleted` (and added to the deleted list) if the parent feature group was deleted. `inaccessible` if you no longer have access to the parent feature group (e.g., the parent feature group belongs to a project you no longer have access to). To traverse the provenance graph in the opposite direction (i.e., from the parent feature group to the child), you can use the [`FeatureGroup.get_generated_feature_groups`][hsfs.feature_group.FeatureGroup.get_generated_feature_groups] method. When navigating the provenance graph downstream, the `deleted` feature groups are not tracked by provenance, as such, the `deleted` property will always return an empty list. === "Python" ```python lineage = transaction_fg.get_generated_feature_groups() # List all accessible child feature groups lineage.accessible # List all the inaccessible child feature groups lineage.inaccessible ``` You can also visualize the relationship between the parent and child feature groups in the UI. In each feature group overview page you can find a provenance section with the graph of parent data source/feature groups and child feature groups/feature views.

Derived feature group provenance graph
Provenance graph of derived feature groups

## Step 3: Feature view lineage The relationship between feature groups and feature views is captured automatically when you create a feature view. You can inspect the relationship between feature groups and feature views using the APIs or the UI. ### Step 3, Using Python Starting from a feature view metadata object, you can traverse upstream the provenance graph to retrieve the metadata objects of the feature groups that are part of the feature view. To do so, you can use the [`FeatureView.get_parent_feature_groups`][hsfs.feature_view.FeatureView.get_parent_feature_groups] method. === "Python" ```python lineage = fraud_fv.get_parent_feature_groups() # List all accessible parent feature groups lineage.accessible # List all deleted parent feature groups lineage.deleted # List all the inaccessible parent feature groups lineage.inaccessible ``` You can also traverse the provenance graph in the opposite direction. Starting from a feature group you can navigate downstream and list all the feature views the feature group is used in. As for the derived feature group example above, when navigating the provenance graph downstream `deleted` feature views are not tracked. As such, the `deleted` property will always be empty. === "Python" ```python lineage = transaction_fg.get_generated_feature_views() # List all accessible downstream feature views lineage.accessible # List all the inaccessible downstream feature views lineage.inaccessible ``` Users can call the [`FeatureView.get_models_provenance`][hsfs.feature_view.FeatureView.get_models_provenance] method which will return a [provenance Link object](#provenance-links). You can also retrieve directly the accessible models, without the need to extract them from the provenance links object: === "Python" ```python #List all accessible models models = fraud_fv.get_models() #List accessible models trained from a specific training dataset version models = fraud_fv.get_models(training_dataset_version: 1) ``` Also we added a utility method to retrieve from the user's accessible models, the last trained one. Last is determined based on timestamp when it was saved into the model registry. === "Python" ```python #Retrieve newest model from all user's accessible models based on this feature view model = fraud_fv.get_newest_model() #Retrieve newest model from all user's accessible models based on this training dataset version model = fraud_fv.get_newest_model(training_dataset_version: 1) ``` ### Step 3, Using UI In the feature view overview UI you can explore the provenance graph of the feature view:

Feature view provenance graph
Feature view provenance graph

## Provenance Links All the `_provenance` methods return a `Link` dictionary object that contains `accessible`, `inaccessible`, `deleted` lists. - `accessible` - contains any artifact from the result, that the user has access to. - `inaccessible` - contains any artifacts that might have been shared at some point in the past, but where this sharing was retracted. Since the relation between artifacts is still maintained in the provenance, the user will only have access to limited metadata and the artifacts will be included in this `inaccessible` list. - `deleted` - contains artifacts that are deleted with children still present in the system. There is minimum amount of metadata for the deleted allowing for some limited human readable identification. ================================================================================ # Feature Monitoring Source: https://docs.hopsworks.ai/latest/user_guides/fs/feature_monitoring/ # Feature Monitoring ## Introduction Feature Monitoring complements the Hopsworks data validation capabilities by allowing you to monitor your data once they have been ingested into the Feature Store. Hopsworks feature monitoring user interface is centered around three functionalities: - **Scheduled Statistics**: The user defines a _detection window_ over its data for which Hopsworks will compute the statistics on a regular basis. The results are stored in Hopsworks and enable the user to visualise the temporal evolution of statistical metrics on its data. This can be enabled for a whole Feature Group or Feature View, or for a particular Feature. For more details, see the [Scheduled statistics guide](scheduled_statistics.md). - **Statistics Comparison**: This variant allows the user to schedule the statistics computation on both a _detection_ and a _reference window_, and compare them on a selected feature using a single scalar metric (e.g., the mean). By providing information about how to compare those statistics, you can setup alerts to quickly detect critical change in the data. For more details, see the [Statistics comparison guide](statistics_comparison.md). - **Data Distribution Comparison**: Instead of a single scalar metric, this variant compares the whole distribution of a feature between the _detection_ and _reference windows_ using distance metrics such as PSI or KL divergence. This helps detect changes in the shape of the data that a single metric might miss. For more details, see the [Distribution comparison guide](distribution_comparison.md). ## Define windows over feature data Windows define the boundaries of the feature data on which Hopsworks operates. Both statistics and data distributions are computed over the feature data delimited by a window. There are different types of windows depending on how they evolve over time. A window can have either a _fixed_ length (e.g., static window) or _variable_ length (e.g., expanding window). Moreover, windows can stick to a _specific point in time_ (e.g., static window) or _move_ over time (e.g., sliding or rolling window). --8<-- "user_guides/fs/feature_monitoring/index/types-of-windows.html" !!! info "Specific values" A specific value can be seen as a window of length 1 where the start and end of the window have the same value. These types of windows apply to both _detection_ and _reference_ windows. Different types of windows allows for different use cases depending on whether you enable feature monitoring on your Feature Groups or Feature Views. See more details about _detection_ and _reference_ windows in the [Detection windows](./scheduled_statistics.md#detection-windows) and [Reference windows](./statistics_comparison.md#reference-windows) guides. ## Visualize metrics on a time series Hopsworks provides an interactive graph to make the exploration of statistics and metrics (e.g., distribution-based distances) more efficient and help you find unexpected trends or anomalous values faster. See the [Interactive graph guide](interactive_graph.md) for more information. ![Feature monitoring graph](../../../assets/images/guides/fs/feature_monitoring/fm-show-shifted-points.png) ## Alerting Moreover, feature monitoring integrates with the Hopsworks built-in system for [alerts](../../../setup_installation/admin/alert.md), enabling you to setup alerts that will notify you as soon as shift is detected in your feature values. You can setup alerts for feature monitoring at a Feature Group, Feature View, and project level. !!! tip "Select the correct trigger" When configuring alerts for feature monitoring, make sure you select the `data shift detected` or `data shift undetected` trigger. ![Feature monitoring alerts](../../../assets/images/guides/fs/feature_monitoring/fm-alerts.png) ================================================================================ # Scheduled Statistics Source: https://docs.hopsworks.ai/latest/user_guides/fs/feature_monitoring/scheduled_statistics/ Hopsworks scheduled statistics allows you to monitor your feature data once they have been ingested into the Feature Store. You can define a ==detection window== over your data for which Hopsworks will compute the statistics on a regular basis. Statistics can be computed on all or a subset of feature values, and on one or more features simultaneously. Hopsworks stores the computed statistics and enable you to visualise the temporal evolution of statistical metrics on your data. ![Detection statistics visualization](../../../assets/images/guides/fs/feature_monitoring/fm-multiple-metrics.png) !!! tip "Interactive graph" See the [Interactive graph guide](interactive_graph.md) to learn how to explore statistics more efficiently. ## Use cases Scheduled statistics monitoring is a powerful tool that allows you to monitor your data over time and detect anomalies in your feature data at a glance by visualizing the evolution of the statistics properties of your data in a time series. It can be enabled in both Feature Groups and Feature Views, but for different purposes. For **Feature Groups**, scheduled statistics enables you to analyze how your Feature Group data evolve over time, and leverage your intuition to identify trends or detect noisy values in the inserted feature data. See the [Feature Monitoring for Feature Groups](../feature_group/feature_monitoring.md) guide to configure it. For **Feature Views**, scheduled statistics enables you to analyze the statistical properties of potentially new training dataset versions without having to actually create new training datasets and, thus, helping you decide when your training data show sufficient significant changes to create a new version. See the [Feature Monitoring for Feature Views](../feature_view/feature_monitoring.md) guide to configure it. ## Detection windows Statistics are computed in a scheduled basis on a pre-defined detection window of feature data. Detection windows can be defined on the whole feature data or a subset of feature data depending on the `time_offset` and `window_length` parameters of the `with_detection_window` method. --8<-- "user_guides/fs/feature_monitoring/scheduled_statistics/detection-windows.html" In [a previous section](index.md#define-windows-over-feature-data) we described different types of windows available. Taking a Feature Group as an example, the figure above describes how these windows are applied to Feature Group data, resulting in three different applications: - A _expanding window_ covering the whole Feature Group data from its creation until the time when statistics are computing. It can be seen as an snapshot of the **latest version of your feature data**. - A _rolling window_ covering a variable subset of feature data (e.g., feature data written last week). It helps you analyze the properties of **newly inserted feature data**. ### Time basis A rolling window needs a **notion of time** to decide which rows fall inside it. Hopsworks supports two bases, chosen once per configuration and **shared by the detection and reference windows**: - _Event time_: rows are selected by the value of an event-time feature, so a window such as "last week" contains the rows whose event time falls in that week regardless of when they were written. This is the default for Feature Groups and Feature Views that declare an `event_time`. - _Commit time_: rows are selected by the time they were written to the Feature Group, using time travel. A window such as "last week" contains the rows committed during that week. This is the default when no event-time feature is declared, and it requires a time-travel enabled Feature Group. A rolling event-time window is anchored on the time the schedule fires. Each run selects the rows whose event time is inside the window at that moment and stores their statistics. A row that lands after the run covering its event time is not added to that run's statistics, and every later window starts after its event time, so no window counts it. This happens with backfills and with the materialization lag of streaming pipelines. For backfill-heavy pipelines, use an expanding window, which has no lower bound, or the commit-time basis, which selects rows by when they were written. !!! tip "Leave room for late rows" `time_offset` sets where the window starts, counted back from the run, and `window_length` sets how long it lasts, so a `time_offset` longer than the `window_length` ends the window before the run. For example, a daily schedule with `time_offset="25h"` and `window_length="24h"` that runs at 12:00 on Tuesday covers event times from 11:00 on Monday to 11:00 on Tuesday, and the next run covers 11:00 on Tuesday to 11:00 on Wednesday. Consecutive windows meet, and each row has one hour to land before the window that covers it runs. Size the gap to the longest delay with which rows land, such as the materialization interval of the pipeline. !!! note "Updated rows" An event-time window reads the current snapshot of the data and filters it on the event-time feature, so it sees only the latest version of each row. A commit-time window reads the commits in its range, so each version of an updated row is counted in the window of the commit that wrote it. See more details on how to define a detection window for your Feature Groups and Feature Views in the Feature Monitoring Guides for [Feature Groups](../feature_group/feature_monitoring.md) and [Feature Views](../feature_view/feature_monitoring.md). !!! info "Next steps" You can also define a reference window to be used as a baseline to compare against the detection window. See more details in the [Statistics comparison guide](statistics_comparison.md). ================================================================================ # Statistics Comparison Source: https://docs.hopsworks.ai/latest/user_guides/fs/feature_monitoring/statistics_comparison/ Hopsworks feature monitoring allows you to monitor your feature data once they have been ingested into the Feature Store. You can define ==detection and reference windows== over your data for which Hopsworks will compute the statistics on a regular basis, compare them, and optionally trigger alerts when significant differences are detected. Statistics can be computed on all or a subset of feature values, and on one or more features simultaneously. Also, you can specify the criteria under which statistics will be compared and set thresholds used to classify feature values as anomalous. Hopsworks stores both detection and reference statistics and enable you to visualise the temporal evolution of statistical metrics. ![Reference statistics visualization](../../../assets/images/guides/fs/feature_monitoring/fm-show-reference.png) !!! tip "Interactive graph" See the [Interactive graph guide](interactive_graph.md) to learn how to explore statistics and comparison results more efficiently. ## Use cases Feature monitoring is a powerful tool that allows you to monitor your data over time and quickly detect anomalies in your feature data by comparing statistics computed on different windows of your feature data, notifying you about anomalies, and/or visualizing the evolution of these statistics and comparison results in a time series. It can be enabled in both Feature Groups and Feature Views, but for different purposes. For **Feature Groups**, feature monitoring helps you rapidly identify unexpected trends or anomalous values in your Feature Group data, facilitating the debugging of possible root causes such as newly introduced changes in your feature pipelines. See the [Feature Monitoring for Feature Groups](../feature_group/feature_monitoring.md) guide to configure it. For **Feature Views**, feature monitoring helps you quickly detect when newly inserted Feature Group data differs statistically from your existing training datasets, and decide whether to retrain your ML models using a new training dataset version or analyze possible issues in your feature pipelines or inference pipelines. See the [Feature Monitoring for Feature Views](../feature_view/feature_monitoring.md) guide to configure it. ## Reference windows To compare statistics computed on a _detection window_ against a baseline, you need to define a _reference window_ of feature data. Reference windows can be defined in different ways depending on whether you are configuring feature monitoring on a Feature Group or Feature View. --8<-- "user_guides/fs/feature_monitoring/statistics_comparison/reference-windows.html" In [a previous section](index.md#define-windows-over-feature-data) we described different types of windows available. Taking a Feature View as an example, the figure above describes how these windows are applied to Feature Group data read by a Feature View query and Training data, resulting in the following applications: - A _expanding window_ covering the whole Feature Group data from its creation until the time when statistics are computing. It can be seen as an snapshot of the latest version of your feature data. This reference window is useful when you want to compare the statistics of **newly inserted feature data against all the Feature Group data**. - A _rolling window_ covering a variable subset of feature data (e.g., feature data written last week). It helps you compare the properties of **feature data inserted at different cadences** (e.g., feature data inserted last month and two months ago). - A _static window_ representing a snapshot of Feature Group data read using the Feature View query at a specific point in time (i.e., Training Dataset). It helps you compare **newly inserted feature data** into your Feature Groups **against a Training Dataset version**. - A _specific value_. It helps you target the analysis of feature data to a **specific feature and statistics metric**. Rolling and expanding reference windows use the same time basis as the detection window of the configuration, either the event-time feature or the commit time. See [Time basis](scheduled_statistics.md#time-basis) in the scheduled statistics guide. See more details on how to define a reference window for your Feature Groups and Training Datasets in the Feature Monitoring guides for [Feature Groups](../feature_group/feature_monitoring.md) and [Feature Views](../feature_view/feature_monitoring.md). ## Comparison criteria After defining the detection and reference windows, you can specify the criteria under which computed statistics will be compared. The criteria described below apply to the comparison of a single scalar metric using the `compare_on` method. !!! tip "Distribution comparison" Alternatively, you can compare the whole distribution of a feature between the detection and reference windows using metrics such as PSI or KL divergence. See the [Distribution comparison guide](distribution_comparison.md) for details. ??? no-icon "Statistics metric" Although all descriptive statistics are computed on the pre-defined windows of feature data, the comparison of statistics is performed only on one of the statistics metrics. In other words, **difference values are only computed for a single statistics metric**. You can select the targeted statistics metric using the `metric` parameter when calling the `compare_on` method. ??? no-icon "Threshold bounds" Threshold bounds are used to classify feature values as anomalous, by comparing them against the difference values computed on a specific statistics metric. You can defined a threshold value using the `threshold` parameter when calling the `compare_on` method. ??? no-icon "Relative or absolute" _Difference_ values represent the amount of change in the detection statistics with regards to the reference values. They can be computed in absolute or relative terms, as specified in the `relative` boolean parameter when calling the `compare_on` method. - **Absolute**: _$detection value - reference value$_ - **Relative**: _$(detection value - reference value) / reference value$_ ??? no-icon "Strict or relaxed" Threshold bounds set the limits under which the amount of change between detection and reference values is ==normal==. These bounds can be strict (`<` or `>`) or relaxed (`<=` and `=>`), as defined in the `strict` parameter when calling the `compare_on` method. Hopsworks stores the results of each statistics comparison and enables you to visualise them together with the detection and reference values in a time series graph. ![Threshold and shift visualization](../../../assets/images/guides/fs/feature_monitoring/fm-show-shifted-points.png) !!! info "Next steps" You can setup alerts that will notify you whenever anomalies are detected on your feature data. See more details in the [alerting section](index.md#alerting) of the feature monitoring guide. ================================================================================ # Distribution Comparison Source: https://docs.hopsworks.ai/latest/user_guides/fs/feature_monitoring/distribution_comparison/ # Distribution Comparison Distribution comparison lets you detect drift in the **shape of a feature's distribution** between a detection window and a reference window, rather than comparing a single scalar metric such as the mean. It is configured with the `compare_on_distribution` method, as an alternative to [`compare_on`](statistics_comparison.md#comparison-criteria). A single scalar metric can miss meaningful changes. For example, the mean of a feature can stay constant while its variance grows or while a unimodal distribution becomes bimodal. Distribution comparison captures these changes by computing a distance between the detection and reference distributions and comparing it against a threshold. !!! info "Reference window required" Distribution comparison always compares two windows, so a reference window is mandatory. Define it with `with_reference_window` or `with_reference_training_dataset` before calling `compare_on_distribution`. !!! tip "Reference training dataset" On a Feature View, you can use a training dataset as the reference distribution with `with_reference_training_dataset(training_dataset_version=...)`. This is the basis for [Model Monitoring](../../mlops/model_monitoring/index.md), where a model's production inference data is compared against the distribution of its training dataset. ## Use cases Distribution comparison is most valuable when changes in your data are not captured by a single scalar metric, but by a change in the overall shape of the feature distribution. It can be enabled on both Feature Groups and Feature Views, but for different purposes. For **Feature Groups**, distribution comparison helps you detect when newly ingested feature data drifts in shape from a baseline window of historical data, surfacing issues such as a new category becoming dominant or a numeric feature shifting from unimodal to multimodal, even when its mean stays stable. See the [Feature Monitoring for Feature Groups](../feature_group/feature_monitoring.md) guide to configure it. For **Feature Views**, distribution comparison helps you quantify how much the distribution of newly inserted Feature Group data has drifted from your training dataset, and decide whether to retrain your ML models on a new training dataset version before the drift degrades model performance. See the [Feature Monitoring for Feature Views](../feature_view/feature_monitoring.md) guide to configure it. ## Distance metrics You select the distance metric with the `metric` parameter. The following metrics are available: - **PSI** (Population Stability Index): the default metric, widely used to monitor drift in production. It is the only metric with a built-in default threshold of `0.2`. - **KL_DIVERGENCE**: Kullback–Leibler divergence; asymmetric, sensitive to regions where the reference has low probability. - **JS_DIVERGENCE**: Jensen–Shannon divergence; a symmetric, bounded smoothing of KL divergence. - **HELLINGER**: Hellinger distance; symmetric and bounded in `[0, 1]`. - **WASSERSTEIN**: Wasserstein (earth mover's) distance; numeric features only. - **KOLMOGOROV_SMIRNOV**: Kolmogorov–Smirnov statistic; numeric features only. For every metric other than PSI, you must provide a `threshold` explicitly. !!! warning "Numeric-only metrics" `WASSERSTEIN` and `KOLMOGOROV_SMIRNOV` require a numeric feature. Applying them to a categorical feature raises an error. ## Binning To compute a distance, Hopsworks first discretizes the feature values into a probability distribution over bins. You control the binning with the following parameters: - `binning_strategy`: how to build the bins. One of `EQUI_WIDTH` (equal-width bins), `EQUI_FREQUENCY` (equal-frequency / quantile bins), `CUSTOM_EDGES` (user-provided bin edges) or `CATEGORICAL` (one bin per category). Defaults to `EQUI_FREQUENCY` for numeric features and `CATEGORICAL` otherwise. - `bin_count`: the number of bins for numeric strategies. Defaults to `10`. - `custom_bin_edges`: the list of bin edges, required when `binning_strategy` is `CUSTOM_EDGES`. - `smoothing_epsilon`: a small additive constant applied to bins to avoid `log(0)` in log-based metrics such as PSI and KL divergence. Defaults to `1e-6`. !!! info "Next steps" Distribution comparison results integrate with the same [alerting](index.md#alerting) and [interactive graph](interactive_graph.md) tooling as scalar statistics comparison. ================================================================================ # Interactive Graph Source: https://docs.hopsworks.ai/latest/user_guides/fs/feature_monitoring/interactive_graph/ Hopsworks provides an *interactive graph* to help you explore the statistics computed on your feature data more efficiently and help you identify anomalies faster. The graph lives on the ^^Feature Monitoring^^ tab of a Feature Group or Feature View, one page per feature monitoring configuration. ### Select a feature monitoring configuration First, you need to select a feature monitoring configuration to visualize. The dropdown in the page header lists every configuration defined on the Feature Group or Feature View. ![Select feature monitoring config](../../../assets/images/guides/fs/feature_monitoring/fm-select-config.png) ### Select a statistics metric to visualize Below the graph, the feature table has one checkbox per feature and statistics metric. Ticking a checkbox plots that metric over time. ![Select statistics metric](../../../assets/images/guides/fs/feature_monitoring/fm-select-metric.png) ### Visualize multiple metrics simultaneously Several metrics can be visualized at the same time on the graph. Tick more than one checkbox in the feature table, and each metric gets its own colour and legend entry. ![Select multiple metrics](../../../assets/images/guides/fs/feature_monitoring/fm-multiple-metrics.png) ### Show reference statistics In feature monitoring configurations with reference windows, you can also visualize the reference values by enabling ^^Reference^^ under ^^Show^^ above the graph. Reference values are drawn as a dashed line: statistics computed over time, or a horizontal line for a specific value. !!! note The same statistics metric is visualized for both detection and reference values. ![Show reference values](../../../assets/images/guides/fs/feature_monitoring/fm-show-reference.png) !!! info More details about reference windows can be found in [Reference windows](statistics_comparison.md#reference-windows). ### Show threshold bounds In addition to reference windows, you can define thresholds to automate the identification of data points as anomalous values. A threshold can be absolute, or relative to the statistics values under comparison. You can visualize the threshold bounds as a band around the reference line by enabling ^^Threshold^^ under ^^View^^. ![Show threshold bounds](../../../assets/images/guides/fs/feature_monitoring/fm-show-threshold.png) !!! info More details about statistics comparison options can be found in [Comparison criteria](statistics_comparison.md#comparison-criteria). ### Highlight shifted data points If a reference window and threshold are provided, data points that fall out of the threshold bounds are considered anomalous values. You can highlight these data points by enabling ^^Shift detected^^ under ^^Show^^. The feature table below the graph flags the same features in its ^^Shift^^ column. ![Highlight shifted data points](../../../assets/images/guides/fs/feature_monitoring/fm-show-shifted-points.png) ### Visualize the computed differences between statistics Alternatively, you can change the time series to show the differences computed between detection and reference statistics rather than the statistics values themselves. You can achieve that by enabling ^^Difference^^ under ^^View^^. The threshold then shows as a horizontal line. ![Show difference between statistics](../../../assets/images/guides/fs/feature_monitoring/fm-show-diff.png) ### Configuration summary and controls The card at the top of a configuration page summarizes the detection and reference windows, the statistics comparison criteria and the job schedule. From there you can trigger the statistics comparison manually with ^^Run once^^, or pause the schedule of the feature monitoring job with ^^Disable^^. !!! note Triggering the statistics comparison manually does not affect the schedule of the feature monitoring. ![Feature monitoring configuration summary](../../../assets/images/guides/fs/feature_monitoring/fm-config-summary.png) ### List of configurations The ^^Feature Monitoring^^ tab itself lists all feature monitoring configurations defined for the Feature Group or Feature View, with their status and next scheduled check. ![List of feature monitoring configs](../../../assets/images/guides/fs/feature_monitoring/fm-list-configs.png) ================================================================================ # Projects Guides Source: https://docs.hopsworks.ai/latest/user_guides/projects/ # Projects Guides A project is the unit of ownership and access in Hopsworks: who is in it, what it can reach, and what it shares. These guides cover signing in, creating and running a project, and the settings that hang off it.
- :material-folder-plus-outline:{ .lg .middle } **Start here** --- Create a project, mint an API key, and connect from any Python environment. ```python import hopsworks project = hopsworks.login() fs = project.get_feature_store() ``` [Create a project](project/create_project.md) · [Create an API key](api_key/create_api_key.md) · [Wizard](wizard.md)
:material-account-key-outline:{ .hops-role-ico } Get in { .hops-role-cap } - [Sign in](auth/login.md) Register and log in, or use OAuth2, LDAP or Kerberos when your cluster is configured for it. - [API keys](api_key/create_api_key.md) Authenticate from outside the cluster: laptops, CI, agents, each key limited to its [scopes](api_key/api_key_scopes.md). - [Manage a project](project/create_project.md) Create projects and add members with roles. - [Wizard](wizard.md) Scaffold a whole system from a description, in a fresh project. - [Search](search.md) Find feature groups, feature views and models, and follow their lineage.
:material-cog-outline:{ .hops-role-ico } Configure { .hops-role-cap } - [Secrets and environment variables](secrets/create_secret.md) Store credentials once and read them from notebooks and jobs. - [Git providers](git/configure_git_provider.md) Connect GitHub, GitLab or Bitbucket, then clone and push from the project. - [Dataset sharing](datasets/sharing.md) Share datasets across projects. - [AWS IAM roles](iam_role/iam_role_chaining.md) Assume roles from the project to reach AWS resources. - [Alerts](alerts/index.md) Get notified on job failures, validation results and data shift.
================================================================================ # Registration Source: https://docs.hopsworks.ai/latest/user_guides/projects/auth/registration/ # Register A New Account On Hopsworks ## Introduction Hopsworks supports different methods of authentication. To use username and password as the method of authentication, you first need to register. ## Prerequisites Registration enabled Hopsworks cluster. The process for registering a new account is as follows ### Step 1: Register a new account Click on the _Register_ button on the login page and register your email address and details.
Register
Register new account
### Step 2: Validate your email address Validate your email address by clicking on the link in the validation email you received. After your account is created an administrator needs to validate your account before you can log in.
Register
Account created
## Two-factor authentication Two-factor authentication is enabled from your account settings after you log in, not during registration. See [how to enable a second factor][step-3-enablereset-two-factor-authentication] in your profile settings. ================================================================================ # Login Source: https://docs.hopsworks.ai/latest/user_guides/projects/auth/login/ # Log in To Hopsworks ## Introduction Hopsworks supports different methods of authentication. Here we will look at authentication using username and password. ## Prerequisites An account on a Hopsworks cluster. ### Step 1: Log in with email and password After your account is validated by an administrator you can use your email and password to login.
Login
Login with password
### Step 2: Two-factor authentication If two-factor authentication is enabled you will be presented with a two-factor authentication window after you enter your password. Use your authenticator app (example. [Google Authenticator](https://play.google.com/store/apps/details?id=com.google.android.apps.authenticator2&hl=en&gl=US)) on your phone to get a one-time password.
Two-factor
One time password
Upon successful login, you will arrive at the landing page:
landing page
Landing page
In the landing page, you will find two buttons. Use these buttons to either create a _demo project_ or [a new project](../../projects/project/create_project.md). ================================================================================ # Password Recovery Source: https://docs.hopsworks.ai/latest/user_guides/projects/auth/recovery/ # Password Recovery ## Introduction This topic describes how to recover a forgotten password. ## Prerequisites An account on a Hopsworks cluster. ### Step 1: Request password reset If you forget your password start by clicking on **Forgot password** on the login page. Enter your email and click on the **Send reset link** button.
Recover password
Password reset
### Step 2: Use the password reset link A password reset link will be sent to the email address you entered if the email is found in the system. Click on the reset link to set your new password. ================================================================================ # OAuth2 Authentication Source: https://docs.hopsworks.ai/latest/user_guides/projects/auth/oauth/ # Login Using A Third-party Identity Provider ## Introduction Hopsworks supports different methods of authentication. Here we will look at authentication using Third-party Identity Provider. ## Prerequisites A Hopsworks cluster with OAuth authentication. See [Configure OAuth2](../../../setup_installation/admin/oauth2/create-client.md) on how to configure OAuth on your cluster. ### Step 1: Log in with OAuth If OAuth is configured a **Login with** button will appear in the login page. Use this button to log in to Hopsworks using your OAuth credentials.
OAuth2 login
Login with OAuth2
### Step 2: Give consent When logging in with OAuth for the first time Hopsworks will retrieve and save consented claims (firstname, lastname and email), about the logged in end-user.
OAuth2 consent
Give consent
After clicking on **Register** you will be redirected to the landing page:
landing page
Landing page
In the landing page, you will find two buttons. Use these buttons to either create a _demo project_ or [a new project](../../projects/project/create_project.md). ================================================================================ # LDAP Authentication Source: https://docs.hopsworks.ai/latest/user_guides/projects/auth/ldap/ # Login using LDAP ## Introduction Hopsworks supports different methods of authentication. Here we will look at authentication using LDAP. ## Prerequisites A Hopsworks cluster with LDAP authentication. See [Configure LDAP](../../../setup_installation/admin/ldap/configure-ldap.md) on how to configure LDAP on your cluster. ### Step 1: Log in with LDAP If LDAP is configured you will see a _Log in using_ alternative on the login page. Choose LDAP and type in your _username_ and _password_ then click on **Login**. Note that you need to use your LDAP credentials.
Log in using LDAP
Log in using LDAP
### Step 2: Give consent When logging in with LDAP for the first time Hopsworks will retrieve and save consented claims (firstname, lastname and email), about the logged in end-user. If you have multiple email addresses registered in LDAP you can choose one to use with Hopsworks. If you do not want your information to be saved in Hopsworks you can click **Cancel**. This will redirect you back to the login page.
OAuth2 consent
Give consent
After clicking on **Register** you will be redirected to the landing page:
landing page
Landing page
In the landing page, you will find two buttons. Use these buttons to either create a _demo project_ or [a new project](../../projects/project/create_project.md). ================================================================================ # Kerberos Authentication Source: https://docs.hopsworks.ai/latest/user_guides/projects/auth/krb/ # Login using Kerberos ## Introduction Hopsworks supports different methods of authentication. Here we will look at authentication using Kerberos. ## Prerequisites A Hopsworks cluster with Kerberos authentication. See [Configure Kerberos](../../../setup_installation/admin/ldap/configure-krb.md) on how to configure Kerberos on your cluster. ### Step 1: Log in with Kerberos If Kerberos is configured you will see a _Log in using_ alternative on the login page. Choose Kerberos and click on **Go to Hopsworks** to login.
Log in using Kerberos
Log in using Kerberos
If password login is disabled you only see the _Log in using Kerberos/SSO_ alternative. Click on **Go to Hopsworks** to login.
Kerberos only
Kerberos only authentication
To be able to authenticate with Kerberos you need to configure your browser to use Kerberos. Note that without a properly configured browser, the Kerberos token is not sent to the server and so SSO will not work. If Kerberos is not configured properly you will see an error such as **User Principal Name not set** when trying to log in.
Browser not configured
Missing Kerberos ticket
### Step 2: Give consent When logging in with Kerberos for the first time Hopsworks will retrieve and save consented claims (firstname, lastname and email), about the logged in end-user. If you have multiple email addresses registered in Kerberos you can choose one to use with Hopsworks. If you do not want your information to be saved in Hopsworks you can click **Cancel**. This will redirect you back to the login page.
OAuth2 consent
Give consent
After clicking on **Register** you will be redirected to the landing page:
landing page
Landing page
In the landing page, you will find two buttons. Use these buttons to either create a _demo project_ or [a new project](../../projects/project/create_project.md). ================================================================================ # Update Profile Source: https://docs.hopsworks.ai/latest/user_guides/projects/auth/profile/ # Update Your Profile and Credentials ## Introduction A profile is required to access Hopsworks. A profile is created when a user registers and can be updated via Account settings. ## Prerequisites An account on a Hopsworks cluster. Updating profile and credentials is not supported if you are using Third-party Identity Providers like Kerberos, LDAP, or OAuth to authenticate to Hopsworks. ### Step 1: Go to your Account settings After you have logged in, in the upper right-hand corner of the screen, you will see your name. Click on your name, then click on the menu item **Account settings**. The account settings page will open with profile tab selected. In this tab you can change your first and last name. You cannot change your email address and will need to create a new account if you wish to change your email address.
User profile
Update profile
### Step 2: Update credential To update your credential go to the **Authentication** tab as shown in the image below.
Update credentials
Update credential
### Step 3: Enable/Reset Two-factor Authentication You can also change your two-factor setting in the **Authentication** tab. Two-factor authentication is only available if it is enabled from the cluster administration page.
Two-factor Authentication
Enable Two-factor Authentication
After enabling or resetting two-factor you will be presented with a QR Code. You will then need to scan the QR code to add it on your phone's authenticator application (example. [Google Authenticator](https://play.google.com/store/apps/details?id=com.google.android.apps.authenticator2&hl=en&gl=US)). If you miss this step, you will have to recover your smartphone credentials at a later stage.
Register Two-factor Authentication
Register Two-factor Authentication
Use the one time password generated by your authenticator app to confirm the registration. ================================================================================ # Create Project Source: https://docs.hopsworks.ai/latest/user_guides/projects/project/create_project/ # How To Create A Project ## Introduction In this guide, you will learn how to create a new project. !!! notice "Project name validation rules" A valid project name can only contain characters a-z, A-Z, 0-9 and special characters ‘_’ and ‘.’ but not ‘__’ (double underscore). There is also a number of [reserved project names](#reserved-project-names) that can not be used. ## Web UI ### Step 1: Create a project If you log in to the platform and do not have any projects, you are presented with the following view. Click `Create a project to get started` to continue.

API Keys
Landing page

### Step 2: Project creation form In the creation form in which you enter the project name, an optional description and set of members to invite to the project.

API Keys
Project creation form

### Step 3: Project creation Then wait for the project creation process to finish.

API Keys
List of created API Keys

### Step 4: Project overview Once the project is created the overview page for it will appear.

API Keys
List of created API Keys

## Code ### Step 1: Connect to Hopsworks ```python import hopsworks hopsworks.login() ``` ### Step 2: Create project ```python project = hopsworks.create_project("my_project") ``` !!! api "API reference" - [`hopsworks.create_project`][hopsworks.create_project] - [`Project`][hopsworks_common.project.Project] Browse the full Python API :material-arrow-right: ## Reserved project names ```bash PROJECTS, HOPS-SYSTEM, HOPSWORKS, INFORMATION_SCHEMA, AIRFLOW, GLASSFISH_TIMERS, GRAFANA, HOPS, METASTORE, MYSQL, NDBINFO, RONDB_REPLICATION, PERFORMANCE_SCHEMA, SQOOP, SYS, GLASSFISH_TIMERS, GRAFANA, HOPS, METASTORE, MYSQL, NDBINFO, PERFORMANCE_SCHEMA, SQOOP, SYS, BIGINT, BINARY, BOOLEAN, BOTH, BY, CASE, CAST, CHAR, COLUMN, CONF, CREATE, CROSS, CUBE, CURRENT, CURRENT_DATE, CURRENT_TIMESTAMP, CURSOR, DATABASE, DATE, DECIMAL, DELETE, DESCRIBE, DISTINCT, DOUBLE, DROP, ELSE, END, EXCHANGE, EXISTS, EXTENDED, EXTERNAL, FALSE, FETCH, FLOAT, FOLLOWING, FOR, FROM, FULL, FUNCTION, GRANT, GROUP, GROUPING, HAVING, IF, IMPORT, IN, INNER, INSERT, INT, INTERSECT, INTERVAL, INTO, IS, JOIN, LATERAL, LEFT, LESS, LIKE, LOCAL, MACRO, MAP, MORE, NONE, NOT, NULL, OF, ON, OR, ORDER, OUT, OUTER, OVER, PARTIALSCAN, PARTITION, PERCENT, PRECEDING, PRESERVE, PROCEDURE, RANGE, READS, REDUCE, REVOKE, RIGHT, ROLLUP, ROW, ROWS, SELECT, SET, SMALLINT, TABLE, TABLESAMPLE, THEN, TIMESTAMP, TO, TRANSFORM, TRIGGER, TRUE, TRUNCATE, UNBOUNDED, UNION, UNIQUEJOIN, UPDATE, USER, USING, UTC_TMESTAMP, VALUES, VARCHAR, WHEN, WHERE, WINDOW, WITH, COMMIT, ONLY, REGEXP, RLIKE, ROLLBACK, START, CACHE, CONSTRAINT, FOREIGN, PRIMARY, REFERENCES, DAYOFWEEK, EXTRACT, FLOOR, INTEGER, PRECISION, VIEWS, TIME, NUMERIC, SYNC, BASE, PYTHON37, FILEBEAT. And any word containing _FEATURESTORE. ``` ================================================================================ # Manage Members Source: https://docs.hopsworks.ai/latest/user_guides/projects/project/manage_members/ # How To Manage Members To A Project ## Introduction In this guide, you will learn how to add new members to your project and understand the different roles available within a project. ## Step 1: View the members list Navigate to the `Project settings` page and locate the `General` section, which displays the current members of the project.

List of project members
List of project members

## Step 2: Add a new member Click `Add members` to open a dialog where you can invite users. Select one or more users to invite.

Add new member dialog
Add new member dialog

Each member can be assigned one of three roles, depending on the level of access they need. ### Data owner Data owners hold the highest authority in the project, with full control over its contents. They can: - Share the project with other projects - Manage project settings and members - Create, read, update, and delete all feature store resources (feature groups, feature views, training datasets, etc.) !!! note "Project author" The project creator is a special type of Data owner. Only the creator can delete the project, and their role cannot be changed. ### Data scientist Data scientists are consumers and creators of feature views and training datasets. They can: - Create feature views and training datasets using existing feature groups - Manage the feature views and training datasets they have created - Read feature groups created by Data owners ### Feature store restricted Feature store restricted users function similarly to Data scientists but with tighter restrictions on what data they can access. They are designed for users who should only work with features that have been explicitly shared with them. Key differences from Data scientist: - **No cross-project feature store access:** A Feature store restricted user cannot use feature groups from a shared project. They can only work within the feature groups of their own project. - **Explicit sharing required:** A Feature store restricted user can only see and use feature groups that have been explicitly shared with them individually, not all feature groups in their project. - **Feature view access is gated by feature access:** A Feature store restricted user can only interact with a feature view if they have access to every feature group the feature view depends on. If even one underlying feature group has not been shared with them, they cannot use that feature view. They can: - Access and use feature groups that have been explicitly shared with them - Create feature views and training datasets using only the features they have been granted access to ## Step 3: Confirm member invitation The invited user will now appear in the members list and will have immediate access to the project based on their assigned role.

Member added to project
List of project members

## Step 4: Manage members To change a member's role or remove them from the project, click the `Manage members` button. From there, you can modify roles or delete members as needed.

Manage project members dialog
Managing members of project

### What happens to a removed member's files Each member has a private home directory in the project, `/Projects//Users/`, holding their notebooks, their SSH key and their agent configuration. When a member is removed, that directory and everything under it is transferred to another data owner. The files keep their contents and their paths; only the owner changes. The removed member loses access, as they do to the rest of the project. The directory keeps the name of the member who had it, since the paths do not change. The new owner finds it in the project's `Users` dataset under that name, next to their own home directory. Nobody else sees it: home directories stay private to whoever owns them. Adding that member back to the project gives them a new, empty home directory. The files they left keep the data owner who took them over, and move to `Users//former-members/` to free the path. A hand-over runs in the background and is retried until it completes. If it is lost, which deleting the removed member's account before it runs does, the platform's periodic permissions check finds the directory and hands it to the longest-serving data owner instead of the one the removal chose. The remove dialog asks which data owner takes them, and starts on the data owner who has been in the project the longest. Only data owners are offered: a data scientist cannot manage members, so files handed to one would be out of reach of the people who can. Service accounts are never chosen. The transfer runs in the background. A member with a large home directory takes a moment to hand over, because every file and directory under it changes owner one at a time, and the removal does not wait for that to finish. Two cases where nothing is transferred, and one where the removal is refused: | Case | Result | | --- | --- | | The removal asks for the home directory to be deleted | The directory is deleted, so there is nothing to transfer | | The member being removed has no home directory | Nothing to transfer | | Removing the member would leave the project with no data owner | The removal is refused. Give another member the data owner role first | ## Python SDK ```python import hopsworks project = hopsworks.login() # Add a member project.add_member("alice@example.com", "Data scientist") # List members for member in project.get_members(): print(member.email, member.role) # Change a member's role project.get_members_api().update_role("alice@example.com", "Observer") # Remove a member. Their files go to the longest-serving data owner project.remove_member("alice@example.com") # Name the data owner that takes over their files project.remove_member("alice@example.com", new_file_owner="carol@example.com") # Delete their files instead of handing them over project.remove_member("alice@example.com", delete_home_dir=True) ``` Roles are the same as in the UI: `Data owner`, `Data scientist`, `Observer`, and `Feature store restricted`. A data scientist removing a member can only remove themselves; the project owner's role cannot be changed or removed. ================================================================================ # Wizard Source: https://docs.hopsworks.ai/latest/user_guides/projects/wizard/ # Wizard The Wizard turns a goal into a running AI system. It asks what you want to build, where the data comes from and what to predict, then writes a kickoff prompt and hands it to Claude Code in the [project terminal][terminal], where the agent builds the feature pipeline, the model and the dashboard with the [Hopsworks CLI][hopsworks-cli]. --8<-- "user_guides/projects/wizard/wizard-flow.html" ## Prerequisites The Wizard drives an agent in the project terminal, so the terminal must be enabled on the cluster (the `enable_terminal` [configuration variable][cluster-configuration]). You need a project role of Data Owner or Data Scientist. ## Start the Wizard Click **Wizard** in the project header. The Wizard also opens by itself on an empty project. It runs as a floating dialog, so you can keep it open while you work in the terminal.
The Wizard dialog asking what you want to build today, with five build types
The Wizard opens on what you want to build; greyed entries are not available yet.
## What you can build The first step asks what you want to build today. **Time-series prediction dashboard.** Forecasting, anomaly detection and multi-horizon predictions. The Wizard asks where the data comes from: - public data, with no setup, for a first run, - features already in the project catalog, - a file you upload (CSV, Parquet or JSON), which the Wizard inspects, - an external data source (S3, BigQuery, Snowflake, Kafka and more), which you connect and preview. It then asks which features to use and what you are trying to predict, opens the terminal with Claude if it is not running yet, and inserts the kickoff prompt.
The Bring data step of the Wizard with four data source options
Bring data: public data, the catalog, an upload or an external source.
**Auto-research.** An agent loops on a training script and keeps the improvements. You pick the data (a small public dataset for a test run, or feature groups from the catalog), CPU or GPU, and how long the agent should run: quick (about ten experiments), an evening (about thirty) or overnight (about a hundred), at roughly five minutes per experiment. The Wizard then inserts the kickoff prompt. Unstructured data, feature engineering and model A/B testing are listed in the Wizard but not available yet. ## After the kickoff From the kickoff prompt on, the agent works in the terminal like any session: it uses `hops` to create feature groups, feature views and models in the project, and you can watch, steer or stop it. The assets it creates are ordinary project assets, visible in the catalog, the model registry and the jobs list. ================================================================================ # Search Source: https://docs.hopsworks.ai/latest/user_guides/projects/search/ # Search { #search-guide } ## Introduction Hopsworks indexes your artifacts so you can find them by name, description, [tag][tags-guide] or [keyword][keywords-guide]. This guide covers the search UI and the equivalent REST call. For what search is and how its scope relates to project membership, see the [search concept page][search-concept]. ## What you can search Search returns ten classes of artifact, each on its own tab: | Tab | What it contains | | --- | --- | | All | Every class below, in one view. The default. | | Feature Groups | Feature groups. | | Feature Views | Feature views. | | Training Datasets | Training datasets. | | Features | Individual features, matched by feature name. | | Jobs | Jobs, excluding those that are apps. | | Apps | Jobs of type PythonApp. | | Models | Models in a model registry. | | Deployments | Deployments that serve a registered model. | | Agents | Deployments that serve no registered model. | Apps and agents are not separate kinds of artifact, which is why they are not separate tabs in the sense the others are. An app is a job, and an agent is a deployment. They appear as their own tabs because the question "which agents are tagged for production" is worth asking on its own, and each tab excludes the other: a job that is an app is reported under Apps and not under Jobs. Each tab shows the number of matches next to its name, so you can see where the results are without visiting every tab. ## Free-text search The search box at the top of the UI matches names and descriptions, and the content of tags and keywords. Matches are highlighted in the results, including the tag key and value that matched, so it is clear why a result was returned. ## Search with tags and keywords Free text cannot express "the tag `data_privacy` has `pii` set to `true`", because it matches text anywhere in the document. For that, turn on `Search with Tags & Keywords` above the results. The panel opens beside the results, and has four rows: - `Selected:` shows every filter currently applied, each removable on its own, with a `Clear search` action for all of them. - `Tag:` is a three-column browser: pick a tag schema, then a key within it, then a value for that key. - `Keyword:` takes a keyword and adds it with `Add keyword filter`. - `Free-text search:` adds words to the free-text part of the query, with `Add` or by pressing Enter. Filters combine, so a tag filter and a keyword filter together return only artifacts matching both. ### Only what exists is offered The three tag columns each have a `filter` box, which matters because a cluster can hold more values than are worth scrolling. Entries that no artifact you can see actually uses are greyed out and cannot be selected, with the hint `No available matching assets in Hopsworks that use this tag/key/value`. This is drawn from the tags in use on indexed artifacts, not from the schema definitions. A schema permitting a value nobody has ever attached will show that value as unavailable, which is deliberate: selecting it could only ever return nothing. The offered vocabulary respects the scope you are searching, so it never reveals a tag value used only in a project you cannot see. ### Clearing a search `Clear search results` above the results removes every filter and the free-text term, and resets the per-tab counts. The same action is in the filter panel as `Clear search`, but the panel is closed by default, so the button above the results is the one to reach for after searching from the text box. ## Search from the API The REST endpoint is per project and takes the class as `docType`: ```bash curl -H "Authorization: ApiKey $API_KEY" \ "https://$HOPSWORKS_HOST/hopsworks-api/api/project/$PROJECT_ID/elastic/featurestore?searchTerm=fraud&docType=ALL" ``` `docType` accepts `ALL`, `FEATURE`, `FEATUREGROUP`, `FEATUREVIEW`, `TRAININGDATASET`, `JOB`, `APP`, `MODEL`, `DEPLOYMENT` and `AGENT`, and defaults to `ALL`. `from` and `size` page the results, with `size` capped at 10000. Tag and keyword filters are JSON arrays in the `tags` and `keywords` query parameters. A tag filter names the schema, and optionally a key and a value within it: ```json [{"name": "data_privacy", "key": "pii", "value": "true"}] ``` A request must carry at least one of `searchTerm`, `tags` or `keywords`. Without any of them it is rejected with a `422`, because there is no "match everything" search: the result would be every artifact on the cluster. The response carries one bucket per class, each with its own total, for example `featuregroups` with `featuregroupsTotal` and `apps` with `appsTotal`. ### API key scopes { #search-api-key-scopes } Search results are filtered to the scopes of the API key you use, so a key cannot discover a class it was not minted for. A `FEATURESTORE` key sees feature groups, feature views, training datasets and features, `JOB` sees jobs and apps, `MODELREGISTRY` sees models, and `SERVING` sees deployments and agents. Naming a `docType` the key does not carry the scope for is rejected, rather than returned empty, so a missing scope is distinguishable from a genuinely empty result. A `docType=ALL` request is instead narrowed to the classes the key does carry, which is what makes `ALL` usable from a single-scope key. A search made with a JWT, as the UI and the Python client do after logging in, is not scope-restricted; it is limited by project membership alone. ### The tag vocabulary in use The endpoint behind the greyed-out entries is available directly: ```bash curl -H "Authorization: ApiKey $API_KEY" \ "https://$HOPSWORKS_HOST/hopsworks-api/api/project/$PROJECT_ID/elastic/featurestore/tagfacets" ``` It returns the tags, keys and values attached to artifacts within your search scope. The answer is read from a bounded number of documents, so on a large cluster it can be incomplete. When it is, the response sets `partial` to `true`, which means the vocabulary shown is a subset and a value missing from it may still exist. ================================================================================ # Dataset Sharing Source: https://docs.hopsworks.ai/latest/user_guides/projects/datasets/sharing/ # Sharing A Dataset ## Introduction Besides [sharing feature groups and feature views][sharing], you can share any dataset in your project's file browser (`Resources`, `Models`, `Jupyter`, etc.) with another project. This grants the target project's members read (or write) access to that directory, without exposing the rest of your project. !!! warning "Requires the Data owner role" Only a [Data owner][data-owner] in the project the dataset lives in can share or unshare a dataset, because sharing exposes project data to members outside the project. !!! note "Feature store datasets are always read-only" Feature store datasets can only be shared as `READ_ONLY`. To grant richer access to feature store data, use [feature store / feature group sharing][sharing] instead. ## UI In the `Files` view, select the dataset (top-level folder) you want to share and choose `Share` from its context menu. Choose the target project and the permission to grant, then confirm. To revoke a share later, choose `Unshare` on the same dataset and select the project to remove. ## Python SDK ```python import hopsworks project = hopsworks.login() dataset_api = project.get_dataset_api() # Share a dataset with another project (read-only by default) dataset_api.share("Resources/my_dir", target_project="other_project") # Or grant write access dataset_api.share( "Resources/my_dir", target_project="other_project", permission="EDITABLE" ) # Revoke a share dataset_api.unshare("Resources/my_dir", target_project="other_project") ``` ================================================================================ # Configure Git Provider Source: https://docs.hopsworks.ai/latest/user_guides/projects/git/configure_git_provider/ # How To Configure a Git Provider ## Introduction When you perform Git operations on Hopsworks that need to interact with the remote repository, Hopsworks relies on the Git HTTPS protocol to perform those operations. Authentication with the remote repository happens through a token generated by the Git repository hosting service (GitHub, GitLab, BitBucket). !!! notice "Token permissions" The token permissions should grant access to public and private repositories including read and write access to repository contents and commit statuses. If you are using the new GitHub access tokens, make sure you choose the correct `Resource owner` when generating the token for the repositories you will want to clone. For the `Repository permissions` of the new GitHub fine-grained token, you should at least give read and write access to `Commit statuses` and `Contents`. ## UI Documentation on how to generate a token for the supported Git hosting services is available here: - [GitHub](https://docs.github.com/en/authentication/keeping-your-account-and-data-secure/creating-a-personal-access-token) - [GitLab](https://docs.gitlab.com/ee/user/profile/personal_access_tokens.html) - [BitBucket](https://confluence.atlassian.com/bitbucketserver/http-access-tokens-939515499.html) ### Step 1: Navigate to Git Providers You can access the `Git Providers` page of your Hopsworks cluster by clicking on your name, in the top right corner, and choosing `Account Settings` from the dropdown menu. The `Git providers` section displays which providers have been already configured and can be used to clone new repositories.

Git provider configuration list
Git provider configuration list

### Step 2: Configure a provider Click on `Edit Configuration` to change a provider username or token, or to configure a new provider. Tick the checkbox next to the provider you want to configure, then click `Add host`. Each host is a row of three fields: the host itself (for example `github.com`), the username, and the token to use for that host. A provider can carry several hosts, so add one row per host you need to authenticate against.

Git provider configuration
Git provider configuration

Click `Save Configuration` to save the configuration. ### Step 3: Provider is configured The configured provider should now be marked as configured.

Git provider configured
Git provider configured

## Code You can also configure a git provider using the hopsworks git API in python. ### Step 1: Get the git API ```python import hopsworks project = hopsworks.login() git_api = project.get_git_api() ``` ### Step 2: Configure git provider ```python PROVIDER = "GitHub" GITHUB_USER = "my_user" API_TOKEN = "my_token" git_api.set_provider(PROVIDER, GITHUB_USER, API_TOKEN) ``` !!! api "API reference" - [`GitApi`][hopsworks_common.core.git_api.GitApi] - [`set_provider`][hopsworks_common.core.git_api.GitApi.set_provider] - [`GitProvider`][hopsworks_common.git_provider.GitProvider] Browse the full Python API :material-arrow-right: ## Going Further You can now use the credentials to [clone a repository](clone_repo.md) from the configured provider. ================================================================================ # Clone Repository Source: https://docs.hopsworks.ai/latest/user_guides/projects/git/clone_repo/ # How To Clone a Git Repository ## Introduction Repositories are cloned and managed within the scope of a project. The content of the repository will reside on the Hopsworks File System. The content of the repository can be edited from Jupyter notebooks and can for example be used to configure Jobs. Repositories can be managed from the Git section in the project settings. The Git overview in the project settings provides a list of repositories currently cloned within the project, the location of their content as well which branch and commit their HEAD is currently at. ## Prerequisites - For cloning a private repository, you should configure a [Git Provider](configure_git_provider.md) with your git credentials. You can clone a GitHub and GitLab public repository without configuring the provider. However, for BitBucket you always need to configure the username and token to clone a repository. ## UI ### Step 1: Navigate to repositories In the left-hand sidebar found in your project click on `Project settings`, and then navigate to the `Git` section. This page lists all the cloned git repositories under `Repositories`, while operations performed on those repositories, e.g `push`/`pull`/`commit` are listed under `Git Executions`.

Repository overview
Git repository overview

### Step 2: Clone a repository To clone a new repository, click on the `Clone repository` button on the Git overview page.

Clone a repository
Git clone

You should first choose the git provider e.g., GitHub, GitLab or BitBucket. If you are cloning a private repository, remember to configure the host, username and token for the provider first in [Git Provider](configure_git_provider.md). The clone dialog also asks you to specify the URL of the repository to clone. The supported protocol is HTTPS. As an example, if the repository is hosted on GitHub, the URL should look like: `https://github.com/logicalclocks/hops-examples.git`. Then specify which branch you want to clone. By default the `main` branch will be used, however a different branch or commit can be specified by selecting `Clone from a specific branch`. You can select the folder, within your project, in which the repository should be cloned. By default, the repository is going to be cloned within the `Jupyter` dataset. However, by clicking on the location button, a different location can be selected. Finally, click on the `Clone repository` button to trigger the cloning of the repository. ### Step 3: Track progress of the clone The progress of the git clone can be tracked under `Git Executions`.

Clone a repository
Track progress of clone

### Step 4: Browse repository files In the `File browser` page you can now browse the files of the cloned repository. In the figure below, the repository is located in `Jupyter/hops-examples` directory.

Browse repository files
Browse repository files

## Code You can also clone a repository through the hopsworks git API in python. ### Step 1: Get the git API ```python import hopsworks project = hopsworks.login() git_api = project.get_git_api() ``` ### Step 2: Clone the repository ```python REPO_URL = ( "https://github.com/logicalclocks/hops-examples.git" # git repository ) HOPSWORKS_FOLDER = "Jupyter" # path in Hopsworks filesystem to clone to PROVIDER = "GitHub" BRANCH = "master" # optional branch to clone examples_repo = git_api.clone( REPO_URL, HOPSWORKS_FOLDER, PROVIDER, branch=BRANCH ) ``` !!! api "API reference" - [`Project.get_git_api`][hopsworks_common.project.Project.get_git_api] - [`GitApi`][hopsworks_common.core.git_api.GitApi] - [`clone`][hopsworks_common.core.git_api.GitApi.clone] - [`GitRepo`][hopsworks_common.git_repo.GitRepo] - [Git management notebook](https://github.com/logicalclocks/hops-examples/blob/master/notebooks/services/git.ipynb) Browse the full Python API :material-arrow-right: ## Errors and Troubleshooting ### Invalid credentials This might happen when the credentials entered for the provider are incorrect. Try the following: - Confirm that the settings for the provider ( in Account Settings > Git providers) are correct. Each host row must carry the host, the username and the token. - Confirm that you have selected the correct Git provider when cloning the repository. - Ensure your personal access token has the correct repository access rights. - Ensure your personal access token has not expired. ### Timeout errors Cloning a large repo or checking out a large branch may hit timeout errors. You can try again later if the system was under heavy load at the time. ### Symlink errors Git repositories with symlinks are not yet supported, therefore cloning repositories with symlinks will fail. You can create a separate branch to remove the symlinks, and clone from this branch. ### TLS certificate errors Cloning from a self-hosted GitLab, GitHub Enterprise or Bitbucket whose certificate is issued by a private certificate authority fails with `server certificate verification failed`. An administrator can configure the certificates to trust in [Cluster Configuration](../../../setup_installation/admin/variables.md): | Variable | Default | Effect | | --- | --- | --- | | `git_custom_ca_configmap` | `""` | Name of a Kubernetes ConfigMap holding PEM-encoded CA certificates to trust. It must exist in every project namespace. | | `git_custom_ca_configmap_key` | `ca-bundle.crt` | The key within that ConfigMap. A single key may hold several concatenated certificates. | | `git_disable_tls_verification` | `false` | Skips certificate verification for every Git remote. | The configured certificates are trusted in addition to the public certificate authorities, so public providers keep working. These settings apply to HTTPS remotes only and have no effect on SSH. !!! warning `git_disable_tls_verification` disables verification for all Git remotes, leaving those connections open to interception. Prefer `git_custom_ca_configmap`. ### Proxy errors If the cluster reaches your Git provider through an HTTP proxy, an administrator can configure one per provider in [Cluster Configuration](../../../setup_installation/admin/variables.md): | Variable | Default | Effect | | --- | --- | --- | | `git_github_http_proxy`, `git_github_https_proxy` | `""` | Proxy for Git traffic to GitHub. | | `git_gitlab_http_proxy`, `git_gitlab_https_proxy` | `""` | Proxy for Git traffic to GitLab. | | `git_bitbucket_http_proxy`, `git_bitbucket_https_proxy` | `""` | Proxy for Git traffic to Bitbucket. | An empty value means a direct connection. ## Going Further You can now start [Jupyter](../jupyter/python_notebook.md) from the cloned git repository path to work with the files. ================================================================================ # Repository Actions Source: https://docs.hopsworks.ai/latest/user_guides/projects/git/repository_actions/ # Repository actions ## Introduction This section explains the git operations or commands you can perform on hopsworks git repositories. These commands include commit, pull, push, create branches and many more. !!! notice "Repository permissions" Git repositories are private. Only the owner of the repository can perform git actions on the repository such as commit, push, pull e.t.c. ## UI The operations to perform on the cloned repository can be found in the dropdown as shown below.

Repository actions on a repository
Repository actions

Note that some repository actions will require the username and token to be configured first depending on the provider. For example to be able to perform a push action in any repository, you must configure the provider for the repository first. To be able to perform a pull action for the for a GitLab repository, you must configure the GitLab provider first. When the provider is not configured, the actions that need it are greyed out in the actions menu, `Commit` and `Push` among them. The Git page also shows an info banner stating that public repositories can be cloned without authentication and pointing you to configure a GitHub, GitLab or BitBucket provider. Configure the provider as described in [Git Provider](configure_git_provider.md) to enable those actions. ## Read only repositories In read only repositories, the following actions are disabled: commit, push and file checkout. The read only property can be enabled or disabled in the Cluster settings > Configuration, by updating the `enable_read_only_git_repositories` variable to true or false. Note that you need administrator privileges to update this property. ## Code You can also perform the repository actions using the hopsworks git API in python. ### Step 1: Get the git API ```python import hopsworks project = hopsworks.login() git_api = project.get_git_api() ``` ### Step 2: Get the git repository ```python git_repo = git_api.get_repo(REPOSITORY_NAME) ``` ### Step 3: Perform the git repository action e.g commit ```python git_repo.commit("Test commit") ``` !!! api "API reference" - [`GitApi`][hopsworks_common.core.git_api.GitApi] - [`get_repo`][hopsworks_common.core.git_api.GitApi.get_repo] - [`GitRepo`][hopsworks_common.git_repo.GitRepo] - [`commit`][hopsworks_common.git_repo.GitRepo.commit] Browse the full Python API :material-arrow-right: ================================================================================ # Secrets Source: https://docs.hopsworks.ai/latest/user_guides/projects/secrets/create_secret/ # How To Create A Secret ## Introduction A Secret is a key-value pair used to store encrypted information accessible only to the owner of the secret. Also if you wish to, you can share the same secret API key with all the members of a Project. ## UI ### Step 1: Navigate to Secrets In the `Account Settings` page you can find the `Secrets` section showing a list of all secrets.

API Keys
List of secrets

### Step 2: Create a Secret Click `New Secret` to bring up the dialog for secret creation. Enter a name for the secret to be used for lookup, then provide the secret value in one of two ways: - `Text`: paste or type the value directly. This is the default mode for short tokens, passwords, and API keys. - `File`: upload a file from your machine. The contents are base64-encoded in the browser and stored as the secret value. Useful for small key files such as SSH keys, service account JSON, or short PEM-encoded certificates. A secret value cannot exceed 9000 characters (about 6.6 KB of raw file content). If the secret should be private to this user, select `Private`. To share the secret with all members of a project, select `Project` and enter the project name.

Create Secret
Create new secret dialog

### Step 3: Secret created After saving, the new secret appears in the list with its name, visibility, and creation date. Use the `Read` action to reveal the stored value at any time.

Secret Created
Secret is now created

## Code ### Step 1: Get secrets API ```python import hopsworks hopsworks.login() secrets_api = hopsworks.get_secrets_api() ``` ### Step 2: Create secret Create a secret from a string value: ```python secret = secrets_api.create_secret("my_secret", "Fk3MoPlQXCQvPo") ``` Create a secret from the contents of a local file: ```python import base64 secrets_api.create_secret_from_file("my_ssh_key", "~/.ssh/id_ed25519") raw_bytes = base64.b64decode(secrets_api.get("my_ssh_key")) ``` Reads return the base64 string, so the caller is responsible for decoding it back to bytes. !!! api "API reference" - [`hopsworks.get_secrets_api`][hopsworks.get_secrets_api] - [`SecretsApi`][hopsworks_common.core.secret_api.SecretsApi] - [`create_secret`][hopsworks_common.core.secret_api.SecretsApi.create_secret] - [`create_secret_from_file`][hopsworks_common.core.secret_api.SecretsApi.create_secret_from_file] - [`get`][hopsworks_common.core.secret_api.SecretsApi.get] Browse the full Python API :material-arrow-right: ================================================================================ # Mountable Secrets Source: https://docs.hopsworks.ai/latest/user_guides/projects/mountable_secrets/mountable_secrets/ # Mountable Secrets Some connectors authenticate with a file rather than with a password. An Oracle Autonomous Database over `tcps` needs a wallet directory, Elasticsearch, MongoDB and Cassandra can need a keystore, and BigQuery and GCS need a key file. A catalog property cannot name a path on the query engine's machines, so a project needs a way to put its own files where the connector will look for them. A **mountable secret** is a named bundle of files that belongs to your project. You upload the files once, then refer to the bundle by name from a catalog property, and Hopsworks substitutes the real location when the catalog is written for Trino. The files are stored where project members cannot read or write them directly, and a catalog can only ever reach its own project's bundles. Only a project Data Owner can list, create or delete a project's mountable secrets. Through the API the same endpoints need an API key with the `MOUNTABLE_SECRET` scope. Private catalogs use mountable secrets that belong to your account instead, described in [Mountable secrets for private catalogs][mountable-secrets-for-private-catalogs]. ## Creating a bundle Open **Project Settings**, then **Mountable Secrets**, and click **New**. Give the bundle a name and add its files, either by selecting the files individually or by uploading a single zip archive, which is the form a downloaded wallet usually arrives in.
Mountable secrets
The project's bundles, what each one holds, which catalogs use it, and how much of the budget is used
New mountable secret
Creating a bundle from individual files or from a zip
The whole bundle is created in one step. There is no way to add a file to a bundle afterwards, or to replace one, which is what makes a bundle safe to reference: it is either complete or absent, and it cannot change under a catalog that is using it. A name starts with a letter or a digit and continues with letters, digits, dots, underscores or hyphens, up to 63 characters. Names that could not be written into a catalog property are rejected, so a leading dot, a space or `..` will not be accepted. A zip must hold its files flat. An archive whose entries sit inside a directory is rejected, because the bundle is the directory, and a wallet nested one level down would not be found by a driver pointed at it. Empty files are rejected too, as is the same filename twice in one request. These limits apply per project. All of them are cluster settings an administrator can raise. | Limit | Default | | --- | --- | | Bundles per project | 10 | | Files per bundle | 32 | | Bytes per file | 1 MiB | | Bytes per project | 16 MiB | | Bytes per upload request | 32 MiB | The listing shows what a bundle holds, with a SHA-256 for every file, when it was uploaded, and which catalogs use it. File contents are never returned: once uploaded, a file can be referenced and deleted, but not read back. The hash is there so you can tell which file is present without reading it.
Files in a mountable secret
Names, sizes and hashes are visible; contents are not
!!! warning "Treat a bundle as readable by the cluster, not by your project alone" Names, sizes, hashes and timestamps are visible to every Data Owner in the project. More importantly, where the cluster runs a Trino test coordinator, every mountable secret on it is readable from that coordinator, because the store is mounted whole and the test coordinator connection-tests catalogs before they go live. Prefer a credential scoped to the data the project needs over an administrative one. ## Referencing a bundle from a catalog Two forms are available, and which one a connector wants depends on whether it reads a directory or a single named file. ```text ${HOPSWORKS_MOUNT:my_bundle} # the bundle directory ${HOPSWORKS_MOUNT:my_bundle/keystore.jks} # one file inside it ``` Type `${` in the catalog properties editor to pick a bundle, or a file inside one, from a list.
Referencing a bundle from a catalog property
Typing ${HOPSWORKS_MOUNT: offers the bundle directory and the files in it
A reference stands on its own and cannot be extended with a path. Writing `${HOPSWORKS_MOUNT:my_bundle}/keystore.jks` is rejected, because the file form above already expresses it, and allowing a path after a reference would let a property address something outside the bundle. For the same reason a reference cannot contain `..`. References are checked when you create or edit the catalog, and again when you test the connection. A bundle or a file that does not exist is reported at that point rather than at the next restart. ## Changing or removing a bundle To change a wallet, delete the bundle and create it again under the same name. The listing names the catalogs that reference a bundle, so check there before removing one. A deletion takes effect immediately and is never refused for being in use, and it does not wait for a restart to bite. The query engine sees the store through a live mount, and a connector that reads its files when it opens a connection, Oracle among them, will fail on its next connection or query. Recreating the bundle under the same name with the same filenames restores it, and no catalog has to be edited, because a catalog refers to the bundle by name. Queries can fail in the gap between the two. Where an interruption is unacceptable, do not replace a bundle in place. Create the new one under a new name, edit the catalog to reference it, and delete the old bundle once the query engine has restarted and the catalog is working. ## Worked example: an Oracle Autonomous Database Download the wallet from the OCI console, upload the zip as a bundle called `oracle_wallet`, then create an `oracle` catalog whose connection URL points `TNS_ADMIN` at the bundle directory. ```properties connection-url=jdbc:oracle:thin:@dbname_low?TNS_ADMIN=${HOPSWORKS_MOUNT:oracle_wallet} connection-user=TRINO connection-password=${HOPSWORKS_SECRET:oracle_password} ``` Three things about this URL cause most of the failures. **The name before `?` is a TNS alias, not a service name.** It has to be one of the aliases in the wallet's own `tnsnames.ora`, such as `dbname_low` or `dbname_high`, and not the service name shown in the OCI console. A name that is not in the file produces `Could not find alias in tnsnames.ora`, which is a wallet-contents problem rather than a connectivity one. **The database's access control list has to admit the cluster.** `ORA-12506` has two causes that look identical from the outside: the connection came from an address the Autonomous Database does not accept, or the client never loaded the wallet at all. Add the outbound addresses of every query engine pod, coordinator and workers, since a query runs on the workers. The Catalogs tab reports those addresses when an administrator has enabled the check. **A downloaded wallet retries by default.** Each alias in `tnsnames.ora` carries `(retry_count=20)(retry_delay=3)` inside its connect descriptor, so a refused connection waits about a minute before any error appears and a rejected address looks like a hang. The driver accepts a connect descriptor in place of an alias, so to get the real error at once, paste the descriptor from the alias you were using into `connection-url` and set `retry_count=0` there. The wallet still authenticates, through `TNS_ADMIN`. ```properties connection-url=jdbc:oracle:thin:@(description=(retry_count=0)(address=(protocol=tcps)(port=1522)(host=))(connect_data=(service_name=))(security=(ssl_server_dn_match=yes)))?TNS_ADMIN=${HOPSWORKS_MOUNT:oracle_wallet} connection-user= connection-password=${HOPSWORKS_SECRET:oracle_password} ``` Take the host, port and `service_name` from the alias's entry in the wallet's `tnsnames.ora`, and drop `retry_delay`, which means nothing once `retry_count` is zero. A catalog can keep this form, and doing so records which consumer group it connects to instead of leaving it to an alias name. ## Mountable secrets for private catalogs A [private catalog][private-catalogs] follows its owner into every project they are a member of, so it cannot reference a project's bundles: it would carry them into the owner's other projects. It references bundles that belong to your account instead. Open **Account Settings**, then **Secrets**, and use the **Mountable secrets** section below your secrets. Creating, listing and deleting work as for a project's bundles, and the same limits apply, counted per account rather than per project.
Mountable secrets on the account Secrets page
Bundles that belong to your account, for use by your private catalogs from any project
A private catalog references your bundles with the same `${HOPSWORKS_MOUNT:}` forms, and a project catalog cannot reference them. Names, sizes, hashes and timestamps of your bundles are visible to you alone. The same caution applies as for a project's bundles: where the cluster runs a Trino test coordinator, every mountable secret on it is readable from that coordinator. Through the API, your account's bundles are at `/users/mountable-secrets`, with the same operations as a project's and an API key with the `MOUNTABLE_SECRET` scope. When your account is deleted, your bundles are deleted with it. ## When the feature is unavailable An administrator can turn the store off for a whole cluster. While it is off, the Mountable Secrets page reports that it is not available, and creating or editing a catalog that references a bundle is refused. A catalog that already went live keeps its stored definition, and its reference still resolves to a location, but nothing populates that location any more. For a connector that reads its files when a connection is opened, such as Oracle, the query engine starts normally and queries fail. ================================================================================ # Environment Variables Source: https://docs.hopsworks.ai/latest/user_guides/projects/env_vars/create/ # Account-level Environment Variables ## Introduction Account-level environment variables are user-scoped, encrypted `name=value` pairs that Hopsworks injects into every runtime you start in any project where you are an active member: - Jobs (Python, Spark, Ray) - Jupyter notebooks - Streamlit / Python apps - Model deployments (KServe, sklearn, TensorFlow, Python predictors) - Python / agent serving - Terminal pods Values are encrypted at rest with the same mechanism as project [Secrets][how-to-create-a-secret]. A user can store up to 64 account-level variables. Names must match `^[A-Za-z_][A-Za-z0-9_]*$` and must not collide with platform-reserved names such as `API_KEY`, `MATERIAL_DIRECTORY`, `PROJECT_ID`, or any name starting with `HOPS_`, `HOPSWORKS_`, or `HOPSFS_`. ## Precedence Environment variables resolve in this order, **highest first**: 1. **Per-execution** variables passed at run time (`Execution.run_env_vars`) where supported. 2. **Per-runtime** variables supplied at create or update time (`Job.envVars`, deployment `predictor_env_vars` / `transformer_env_vars`, app `env_vars`). 3. **Account-level** variables defined here. 4. Platform-injected image / runtime defaults (protected by the reserved-name list at validation time). A value defined for a specific job, deployment, or app **always overrides** the account-level value with the same name for that one runtime. To clear an account-level value for a single run, set it to an empty string at the runtime level. ## UI ### Step 1: Open Account settings Click your avatar in the top right of Hopsworks and choose **Account settings**. ### Step 2: Open the Environment variables tab In **Account settings**, click the **Environment variables** tab. You'll see a list of your existing variables, one per row. ### Step 3: Add a variable Type a `NAME` and a `value` in the trailing empty row, then click the green checkmark to save. The value is hidden by default if the name contains `key` or `token` (case-insensitive); for those rows an eye icon toggles visibility. All other names render as plain text. To **edit** a value, change it inline and click the checkmark again. To **remove** one, click the trash icon. ### Pre-fill in New Job / New Deployment / New App When you create a new Job, Deployment, or App, the **Environment variables** section in those dialogs is pre-filled with your account-level variables so you can see exactly what will be injected. Any value you change in those dialogs becomes a per-runtime override (precedence #2 above) for that one runtime. Deleting a row from the dialog sends an empty value as a per-runtime override, effectively clearing the account-level value for that runtime only. To remove a variable everywhere, delete it from **Account settings → Environment variables**. ## Python SDK ```python import hopsworks hopsworks.login() api = hopsworks.get_env_vars_api() # Add api.create_env_var("OPENAI_API_KEY", "sk-...") # Add or update (idempotent, safe in setup scripts) api.set_env_var("HF_TOKEN", "hf_...") # Read api.get("OPENAI_API_KEY") # -> "sk-..." or None api.get_env_var("OPENAI_API_KEY") # -> EnvVar object or None # List for v in api.get_env_vars(): print(v.name) # Update api.update_env_var("OPENAI_API_KEY", "sk-new") # Remove api.delete_env_var("OPENAI_API_KEY") # Remove all api.delete_all() ``` `get_env_var` and `get` return `None` for missing names instead of raising, so they're convenient for "set if missing" patterns. `delete_env_var` raises `RestAPIError` with `ENV_VAR_NOT_FOUND` if the name doesn't exist. ## Notes - Account-level variables are **per-user**. They are never shared with other users in your project. To share configuration with project members, use project-scoped [Secrets][how-to-create-a-secret] instead. - Removing yourself from a project does **not** remove these variables; they follow your account, not your project membership. - Values are encrypted at rest, mirroring Secrets. They are not exposed in audit logs or error messages. ================================================================================ # Create API Key Source: https://docs.hopsworks.ai/latest/user_guides/projects/api_key/create_api_key/ # How To Create An API Key ## Introduction An API key allows a user or a program to make API calls without having to authenticate with a username and password. To access an endpoint using an API key, a client should send the access token using the ApiKey authentication scheme. The API Key can now be used when connecting to your Hopsworks instance using the `hopsworks`, `hsfs` or `hsml` python library or set in the `ApiKey` header for the REST API. ```bash GET /resource HTTP/1.1 Host: server.hopsworks.ai Authorization: ApiKey ``` ## UI In this guide, you will learn how to create an API key. ### Step 1: Navigate to API Keys In the _Account Settings_ page you can find the _API_ section showing a list of all API keys. The table shows each key's name, scope, prefix, creation date, last modification date, and expiration date. Keys with no expiration show _Never_ in the expiration column.

API Keys
List of API Keys

### Step 2: Create an API Key Click `New API key`, enter a name, optionally set an expiration, select the required scopes, and click `Create API key`. Each scope unlocks a group of REST endpoints; see [API Key Scopes][api-key-scopes] for what every scope grants. **Expiration options:** | Option | Description | | -------- | ------------- | | No expiration | The key never expires (default). | | 7 / 30 / 60 / 90 days | The key expires the selected number of days from now. | | Custom | Pick a specific date using the date picker. | To change the expiration date of an existing key you must regenerate it. Expiration cannot be updated after creation. Copy the key value and save it in a secure location, such as a password manager. It will not be shown again.

Create API Key
Create new API Key

## Login with API Key using SDK In this guide you learned how to create an API Key. You can now use the API Key to [login][hopsworks.login] using the `hopsworks` python SDK. ================================================================================ # API Key Scopes Source: https://docs.hopsworks.ai/latest/user_guides/projects/api_key/api_key_scopes/ # API Key Scopes Every API key carries a set of scopes. A scope unlocks a group of REST endpoints; a request made with a key that lacks the scope an endpoint requires is rejected before it reaches the endpoint. Scopes are chosen when a key is created and can be changed later from the key's edit page, without regenerating the secret. See [How To Create An API Key][how-to-create-an-api-key] for the UI walkthrough. A scope never grants more than the account itself may do. Endpoints still check the caller's role in the project, so a Data Scientist's key with the `FEATURESTORE` scope cannot do what a Data Owner's key with the same scope can. ## Scope reference | Scope | Grants access to | | --- | --- | | `FEATURESTORE` | Feature stores and everything inside them: feature groups, feature views, training datasets, data sources and storage connectors, transformation functions, statistics, data validation, feature monitoring, tags, keywords, provenance and feature store search. Also Hopsworks actions and the tag schema catalogue; creating or deleting a tag schema additionally requires the `HOPS_ADMIN` role. | | `PROJECT` | Project management: list, create, update and delete projects; read project information and client credentials; manage members; project alerts, receivers, routes and silences; cloud role mappings; the operation log; tutorials and product news. | | `JOB` | Jobs and executions: create, update, schedule, start, stop and delete jobs; read execution logs; default job configurations; job alerts and tags; Python apps; expectation suites and validation reports. | | `DATASET_VIEW` | Read access to project datasets: list datasets, browse and download files, and use global, project and dataset search. | | `DATASET_CREATE` | Create datasets and directories, upload files, and copy, move, zip or unzip them. | | `DATASET_DELETE` | Delete datasets, directories and files. | | `MODELREGISTRY` | Model registries and models: register, update and delete models; model tags and provenance; Hugging Face imports; generated deployment configurations. | | `SERVING` | Model deployments: create, start, stop and delete deployments; read deployment logs; send inference requests; deployment tags; OpenTelemetry traces and metrics. | | `KAFKA` | The project's Kafka topics and schema registry: topics, subjects, schema versions and compatibility settings. Also accepted, as an alternative to `FEATURESTORE` or `PROJECT`, by the few read endpoints a Kafka or OnlineFS client needs, such as listing projects and feature stores. | | `PYTHON_LIBRARIES` | Python environments: list, create and delete environments; install and uninstall pip, conda and npm packages; search package indexes; environment build commands, history and conflicts. | | `GIT` | Git repositories in the project: clone, branches, commits, remotes, repository actions and their executions, and the account's Git provider credentials. | | `TRINO` | The Trino query engine: submit and cancel SQL statements, read query, worker and cluster status, and manage Trino catalogs. | | `SUPERSET` | Superset dashboards: log in to Superset, list dashboards, create permalinks, make a dashboard public or share it with another project, and delete dashboards. | | `TERMINAL` | The web terminal: start, extend, stop and inspect terminal sessions, and mint the proxy tokens used to attach to them. | | `MOUNTABLE_SECRET` | The project's mountable secrets: named bundles of credential files (Oracle wallets, JKS keystores, service account JSON) that a service mounts read-only. Create, list and delete bundles. Contents are never returned. Requires the Data Owner role. | | `USER` | The account itself: profile, secrets, account environment variables, AI provider settings, and API keys. A key with this scope can create, edit and delete API keys, including keys carrying any other scope the account is allowed to hold, so treat it as equivalent to all of them. | | `ADMIN` | Cluster administration: the admin API (configuration variables, backups, projects, users, Trino, TTL purge, coding agent configuration, cloud role mappings, search reindexing, the operation log), compute resources and the UI theme. Privileged. | | `ADMINISTER_USERS` | User administration in the admin API: list, accept, reject, block, modify and delete users, change roles, reset passwords and sync remote groups. Privileged. | | `ADMINISTER_USERS_REGISTER` | Only the user registration endpoint of the admin API. Privileged. | | `AUTH` | The JWT service: issue, renew and invalidate tokens and remove signing keys. Privileged, and also available to accounts in the `AGENT` group. | ## Privileged scopes `ADMIN`, `ADMINISTER_USERS`, `ADMINISTER_USERS_REGISTER` and `AUTH` are privileged. Only accounts with the `HOPS_ADMIN` role can create keys carrying them, because the endpoints they unlock act on the whole cluster rather than on a project the caller is a member of. ## Scopes an account can select The set of scopes offered when creating or editing a key depends on the account's role. | Account role | Selectable scopes | | --- | --- | | `HOPS_ADMIN` | All scopes. | | `HOPS_USER` | All unprivileged scopes. | | `AGENT` | All unprivileged scopes plus `AUTH`. | | `HOPS_SERVICE_USER` | All unprivileged scopes except `GIT`. | The API key form preselects `FEATURESTORE`, `PROJECT`, `JOB`, `DATASET_VIEW`, `DATASET_CREATE`, `DATASET_DELETE`, `KAFKA`, `SERVING`, `MODELREGISTRY`, `USER` and `PYTHON_LIBRARIES`. Deselect what the key's consumer does not need. ## Scopes of a key created by hops setup `hops setup` creates its key through the browser token flow rather than the API key form, so the scopes are not chosen interactively. The key carries every scope a `hops` subcommand needs: `FEATURESTORE`, `PROJECT`, `JOB`, `DATASET_VIEW`, `DATASET_CREATE`, `DATASET_DELETE`, `MODELREGISTRY`, `SERVING`, `USER`, `KAFKA`, `TERMINAL`, `PYTHON_LIBRARIES`, `GIT`, `TRINO` and `SUPERSET`. A key created by an older release lacks the last six; edit it in the UI to add them, or run `hops setup --force` to mint a new one. ## Scope errors A request made with a key that lacks the required scope fails with HTTP 403 and error code 320004. The message names the scope the endpoint accepts. ```json { "errorCode": 320004, "usrMsg": "No valid scope found for this invocation. Valid scope for this invocation is: [PYTHON_LIBRARIES]", "errorMsg": "No valid scope found for this invocation" } ``` Add the named scope to the key from the _API_ section of _Account Settings_, or create a new key that has it. ================================================================================ # AWS IAM Roles Source: https://docs.hopsworks.ai/latest/user_guides/projects/iam_role/iam_role_chaining/ # How To Use AWS IAM Roles on EC2 instances ## Introduction When deploying Hopsworks on EC2 instances you might need to assume different roles to access resources on AWS. These roles can be configured in AWS and mapped to a project in Hopsworks. ## Prerequisites Before you begin this guide you'll need the following: - A Hopsworks cluster running on EC2. - [Role chaining](https://docs.aws.amazon.com/IAM/latest/UserGuide/id_roles_terms-and-concepts.html#iam-term-role-chaining) setup in AWS. - Configure role mappings in Hopsworks. For a guide on how to configure this see [AWS IAM Role Chaining](../../../setup_installation/admin/roleChaining.md). ## UI In this guide, you will learn how to use a mapped IAM role in your project. ### Step 1: Navigate to your project's IAM Role Chaining tab In the _Project Settings_ page you can find the _IAM Role Chaining_ section showing a list of all IAM roles mapped to your project.
Role Chaining
Role Chaining
### Step 2: Use the IAM role You can now use the IAM roles listed in your project when creating a Data Source with [Temporary Credentials](../../fs/data_source/creation/s3.md#temporary-credentials). ================================================================================ # Alerts Source: https://docs.hopsworks.ai/latest/user_guides/projects/alerts/ # Alerts ## Introduction Hopsworks can notify you when something happens in your project, such as a job failing or a feature ingestion succeeding. Alerts are delivered through Prometheus' [Alert manager](https://prometheus.io/docs/alerting/latest/alertmanager/) to receivers that you define per project. Email, Slack and PagerDuty alerts require an administrator to first configure the corresponding channel for the cluster, so that Hopsworks knows how to reach the provider. Webhook receivers can be created directly in a project without any cluster-level channel configuration. See [Configure Alerts](../../../setup_installation/admin/alert.md) for the administrator setup. You manage a project's alerts under _Project Settings_ → _Alerts_. ## Alert receivers A receiver is a destination that an alert is sent to, for example a Slack channel or a webhook URL. The _Alert receivers_ table lists the receivers available to the project, both the receivers you created in this project and the global receivers shared across the cluster.
Alert receivers with load status
Alert receivers with their load status
To add a receiver click _Add receiver_, choose a channel, give it a name and fill in the channel details. ### Receiver load status When you create or edit a receiver, Hopsworks writes the change to the Alert manager configuration and then the Alert manager loads it asynchronously. The receiver is not usable until it has been loaded, so each receiver shows a _Status_ that tells you whether it is ready. | Status | Meaning | | --- | --- | | Loaded | The receiver is active in the Alert manager and can be used in an alert. | | Pending | The receiver has been saved but the Alert manager has not loaded it yet. It usually becomes _Loaded_ within a minute. | | Warning | The receiver has stayed unloaded past the timeout. The Alert manager most likely rejected the configuration, for example because it is malformed. | A newly created receiver appears immediately in the list as _Pending_ and switches to _Loaded_ once the Alert manager has picked it up.
A pending receiver
A newly created receiver shown as Pending until the Alert manager loads it
If a receiver stays in _Warning_, edit it to correct the configuration or delete it. The status refreshes on its own after a create or edit, and you can trigger a refresh manually with the _Refresh_ button. ## Creating alerts Alerts are triggered by events such as a job succeeding or failing, a feature ingestion completing, or feature monitoring detecting a shift. You create project-wide alerts under _Global alerts_, and you can also attach alerts to an individual job or feature group from its own page. For each alert you choose a trigger, a severity and the receiver that the notification is sent to. Only receivers that are _Loaded_ can be selected. A receiver that is not yet loaded is shown as disabled in the receiver dropdown, because routing an alert to a receiver the Alert manager has not loaded would silently drop the notification.
A pending receiver disabled in the receiver dropdown
A pending receiver cannot be selected until it is loaded
Each configured global alert shows a _Route_ status that mirrors the receiver status, so you can confirm that the route is loaded in the Alert manager. ## Triggered alerts (debugging) The _Triggered alerts (debugging)_ panel shows the Alert manager's live view of the alerts currently firing for the project. Use it to understand why a triggered alert may not have been delivered. The panel shows the routing timing in effect (`group_wait`, `group_interval`, `repeat_interval` and `group_by`) and a table of the alerts the Alert manager currently holds.
Triggered alerts debugging panel
Live view of alerts firing for the project
If an alert you expected is missing or was not delivered, the panel explains the common reasons: - Not listed below: the alert is not firing, or its labels do not match any route. Compare against the `group_by` shown above. - Suppressed: the alert is muted by a silence or inhibited by another alert. See the _Suppressed by_ column. - Held back by timing: the first alert in a new group waits `group_wait`, and later alerts added to an existing group are sent on the next `group_interval`. - Grouped or already firing: alerts that share the same `group_by` labels are bundled into a single notification, and an alert whose labels match one already firing is not re-sent until `repeat_interval`. Either way several firings can arrive as one message. - Wrong receiver: the alert resolved to `default-receiver` or to a receiver outside this project, so its route or labels are misconfigured. Global receivers are valid for every project. ================================================================================ # Compute Guides Source: https://docs.hopsworks.ai/latest/user_guides/compute/ # Compute Guides Compute is where code runs inside a project: notebooks, a terminal, jobs, schedules and served apps, all on the project's Python environments. The guides go from writing code interactively to running it on a schedule.
- :material-play-circle-outline:{ .lg .middle } **Start here** --- Upload a script and run it as a job from any Python session. Inside the project, Jupyter and the terminal are already logged in. ```python jobs_api = project.get_job_api() config = jobs_api.get_configuration("PYTHON") config["appPath"] = "Resources/script.py" job = jobs_api.create_job("py_job", config) execution = job.run(await_termination=True) ``` [Run a Python job](../projects/jobs/python_job.md) · [Jupyter](../projects/jupyter/python_notebook.md) · [Python environments](../projects/python/python_env_overview.md)
:material-code-tags:{ .hops-role-ico } Write code { .hops-role-cap } - [Jupyter](../projects/jupyter/python_notebook.md) Python, PySpark and Ray notebooks running against the project. - [Terminal](terminal.md) A shell in the project with the CLI, Git and coding agents preinstalled. - [Python environments](../projects/python/python_env_overview.md) Install libraries, clone and export environments, run custom build commands. - [Apps](../projects/apps/index.md) Serve Streamlit, Gradio or any web app from the project.
:material-calendar-clock-outline:{ .hops-role-ico } Run and schedule { .hops-role-cap } - [Jobs](../projects/jobs/python_job.md) Run Python, notebook, PySpark, Spark or Ray code as a job, on demand or on a schedule. - [Airflow](../projects/airflow/airflow.md) Orchestrate jobs as DAGs with the Hopsworks operators. - [Kubernetes scheduling](../projects/scheduling/kube_scheduler.md) Node selectors, tolerations and Kueue queues for where work lands. - [Python deployments](../projects/python-deployment/python-deployment.md) Expose a Python function behind a REST endpoint.
================================================================================ # Environments Overview Source: https://docs.hopsworks.ai/latest/user_guides/projects/python/python_env_overview/ # Python Environments ## Introduction Hopsworks postulates that building ML systems following the FTI pipeline architecture is best practice. This architecture consists of three independently developed and operated ML pipelines: - Feature Pipeline: takes as input raw data that it transforms into features (and labels) - Training Pipeline: takes as input features (and labels) and outputs a trained model - Inference Pipeline: takes new feature data and a trained model and makes predictions. In order to facilitate the development of these pipelines Hopsworks bundles several python environments containing necessary dependencies. Each environment can also be customized further by installing additional dependencies from PyPi, Wheel files, GitHub repos or applying custom Dockerfiles on top. ### Step 1: Go to environments page Under the `Project settings` section you can find the `Python environment` setting. ### Step 2: List available environments The page is titled `Prebuilt and Custom Container Images` and lists the environments in three columns. Environments listed under `FEATURE ENGINEERING` correspond to environments you would use in a feature pipeline, `MODEL TRAINING` maps to environments used in a training pipeline, and `INFERENCE / AGENTS / APPS` are what you would use in inference pipelines, agents and applications.

Bundled python environments
Bundled python environments

!!! note "Python version" The python version used in all the environments is 3.13. ### Feature engineering The `FEATURE ENGINEERING` environments can be used in [Jupyter notebooks](../jupyter/python_notebook.md), a [Python job](../jobs/python_job.md) or a [PySpark job](../jobs/pyspark_job.md). - `agent-job` an AI agent runtime bundling Claude Code and OpenAI Codex, meant to be cloned and extended with your own libraries - `dlthub-ingestion-pipeline` for ingesting data from data sources into feature groups - `python-feature-pipeline` for writing feature pipelines using Python, with Pandas and Polars - `spark-feature-pipeline` for writing feature pipelines using PySpark - `dbt-pipeline` extends `python-feature-pipeline` with dbt and the Trino and DuckDB adapters, for transformations written as dbt models ### Model training The `MODEL TRAINING` environments can be used in [Jupyter notebooks](../jupyter/python_notebook.md) or a [Python job](../jobs/python_job.md). - `tensorflow-training-pipeline` to train TensorFlow models - `torch-training-pipeline` to train and fine-tune PyTorch models and LLMs - `pandas-training-pipeline` to train XGBoost, Catboost and Sklearn models ### Inference, agents and apps The `INFERENCE / AGENTS / APPS` environments can be used in a deployment using a custom predictor script, and for agents and applications. - `tensorflow-inference-pipeline` to load and serve TensorFlow models - `torch-inference-pipeline` to load and serve PyTorch models - `pandas-inference-pipeline` to load and serve XGBoost, Catboost and Sklearn models - `python-agent-pipeline` to build Python agents, bundling FastAPI, LlamaIndex and OpenTelemetry - `python-app-pipeline` to build interactive applications with Streamlit - `vllm-inference-pipeline` to load and serve LLMs with vLLM inference engine - `minimal-inference-pipeline` to install your own custom framework, contains a minimal set of dependencies ## Next steps In this guide you learned how to find the bundled python environments and where they can be used. Now you can test out the environment in a [Jupyter notebook](../jupyter/python_notebook.md). ================================================================================ # Clone Environment Source: https://docs.hopsworks.ai/latest/user_guides/projects/python/python_env_clone/ # How To Clone Python Environment ## Introduction Cloning an environment in Hopsworks means creating a copy of one of the base environments. The base environments are immutable, meaning that it is required to clone an environment before you can make any change to it, such as installing your own libraries. This ensures that the project maintains a set of stable environments that are tested with the capabilities of the platform, meanwhile through cloning, allowing users to further customize an environment without affecting the base environments. In this guide, you will learn how to clone an environment. ## Step 1: Select an environment Under the `Project settings` section you can find the `Python environment` setting. First select an environment, for example the `python-feature-pipeline`.

Select a base environment

## Step 2: Clone environment The environment can now be cloned by clicking `Clone env` and entering a name and description. The interface will show `Syncing packages` while creating the environment.

Create environment
Clone a base environment

## Step 3: Environment is now ready

Environment is now cloned

!!! notice "What does the CUSTOM mean?" Notice that the cloned environment is tagged as `CUSTOM`, it means that it is a base environment which has been cloned. !!! notice "Base environment also marked" When you select a `CUSTOM` environment the base environment it was cloned from is also shown. ## Concerning upgrades !!! warning "Please note" The base environments are automatically upgraded when Hopsworks is upgraded and application code should keep functioning provided that no breaking changes were made in the upgraded version of the environment. A `CUSTOM` environment is not automatically upgraded and the users is recommended to reapply the modifications to a base environment if they encounter issues after an upgrade. ## Next steps In this guide you learned how to clone a new environment. The next step is to [install](python_install.md) a library in the environment. ================================================================================ # Install Library Source: https://docs.hopsworks.ai/latest/user_guides/projects/python/python_install/ # How To Install Python Libraries ## Introduction Hopsworks comes with several prepackaged Python environments that contain libraries for data engineering, machine learning, and more general data science use-cases. Hopsworks also offers the ability to install additional packages from various sources, such as using the pip package manager and public or private git repository. In this guide, you will learn how to install Python packages using these different options. - PyPi, using pip package manager - Packages contained in .whl format - A public or private git repository - A requirements.txt file to install multiple libraries at the same time using pip - npm packages, installed globally in the environment image and on `PATH` for jobs, Jupyter and the terminal !!! warning "Conda, .egg and environment.yml installs have been removed" Installing from a conda channel, from an `.egg` distribution, or by importing an `environment.yml` file is no longer supported and is refused with an error naming the sources you can use instead. The base environments are built on a virtualenv that does not carry conda, so there is nothing for such an install to run. Removing a package is not affected. An environment created before this change can still carry conda-sourced packages, and uninstalling those keeps working, so an upgraded environment does not strand packages you can no longer remove. !!! notice "Notice" If your libraries require installing some extra OS-Level packages, refer to the guide custom commands guide on how to install OS-Level packages. ## Prerequisites In order to install a custom dependency one of the base environments must first be cloned, follow [this guide](python_env_clone.md) for that. ### Step 1: Go to environments page Under the `Project settings` section select the `Python environment` setting. ### Step 2: Select a CUSTOM environment Select the environment that you have previously cloned and want to modify. ### Step 3: Installation options #### Name and version Enter the name and, optionally, the desired version to install.

Installing library by name and version
Installing library by name and version

#### Search Enter the search term and select a library and version to install.

Installing library using search
Installing library using search

#### Distribution (.whl) Install a python package by uploading the corresponding package file and selecting it in the file browser.

Installing library from file
Installing library from file

#### Git source The URL you should provide is the same as you would enter on the command line using `pip install git+{repo_url}`, where `repo_url` is the part that you enter in `Git URL`. For example to install matplotlib 3.7.2, the following are correct inputs: `matplotlib @ git+https://github.com/matplotlib/matplotlib@v3.7.2` `git+https://github.com/matplotlib/matplotlib@v3.7.2` In the case of a private git repository, also select whether it is a GitHub or GitLab repository and the preconfigured access token for the repository. !!! notice "Keep your secrets safe" If you are installing from a git repository which is not GitHub or GitLab simply supply the access token in the URL. Keep in mind that in this case the access token may be visible in logs for other users in the same project to see.

Installing library from git repo
Installing library from git repo

#### npm packages Environments also carry npm packages, installed globally in the image so they are on `PATH` for jobs, Jupyter and the terminal. This is how you get a CLI tool or a JavaScript dependency into an environment. Open the **Installed npm Libraries** tab beside the Python one and use **Install npm package**. Installed npm packages are listed there and can be uninstalled the same way, so an environment records what it carries rather than accumulating changes nobody can see. What the platform accepts: | | | | --- | --- | | Versions | An exact version such as `1.3.0`, or a dist-tag such as `latest`. Ranges like `^1.0.0` are refused, because they are not reproducible. | | Scoped names | Supported, for example `@tsconfig/node20`. | | Flags | A fixed set: `--ignore-scripts`, `--legacy-peer-deps`, `--no-audit`, `--no-fund`, `--no-optional`, `--strict-peer-deps`. Anything else is refused and the message lists what is allowed. | Installing by dist-tag records the version the tag resolved to once the build finishes, so `latest` becomes the concrete version in the listing rather than staying as `latest`. A package the base image already ships cannot be installed over, and cannot be uninstalled. Clone an environment and install your own version there instead. Packages come from the registry the cluster is configured to use. Ask your administrator if you need an internal registry; it is a cluster-wide setting rather than a per-environment one. !!! note "Same name in both ecosystems" A name can exist on both PyPI and npm. The two are tracked separately, so installing `requests` from npm does not touch the Python package of the same name, and each is listed under its own tab. ## Going Further Now you can use the library in a [Jupyter notebook](../jupyter/python_notebook.md) or a [Job](../jobs/python_job.md). ================================================================================ # Export Environment Source: https://docs.hopsworks.ai/latest/user_guides/projects/python/python_env_export/ # How To Export Python Environment ## Introduction Each of the python environments in a project can be exported to an `environment.yml` file. It can be useful to export it to keep a snapshot of all the installed libraries and their versions. In this guide, you will learn how to export a python environment. ## Step 1: Go to environment Under the `Project settings` section you can find the `Python environment` setting. ## Step 2: Select a CUSTOM environment Select the environment that you have previously cloned and want to export. Only a `CUSTOM` environment can be exported. ## Step 3: Click Export env Clicking `Export env` will download the `environment.yml` file in your browser. !!! notice "The export is a snapshot, not an import format" Exporting keeps working and describes whatever the environment currently carries. Importing an `environment.yml` to create or modify an environment is no longer supported, so use the exported file as a record of what was installed rather than as a way to rebuild the environment elsewhere. To reproduce an environment, clone it or install the same packages from PyPI.

Export environment
Export environment

================================================================================ # Custom Commands Source: https://docs.hopsworks.ai/latest/user_guides/projects/python/custom_commands/ # Adding extra configuration with generic bash commands ## Introduction Hopsworks comes with several prepackaged Python environments that contain libraries for data engineering, machine learning, and more general data science use-cases. Hopsworks also offers the ability to install additional packages from various sources, such as using the pip package manager and public or private git repository. Some Python libraries require the installation of some OS-Level libraries. In some cases, you may need to add more complex configuration to your environment. This demands writing your own commands and executing them on top of the existing environment. In this guide, you will learn how to run custom bash commands that can be used to add more complex configuration to your environment e.g., installing OS-Level packages or configuring an oracle database. ## Prerequisites In order to install a custom dependency one of the base environments must first be cloned, follow [this guide](python_env_clone.md) for that. ## Running bash commands In this section, we will see how you can run custom bash commands in Hopsworks to configure your Python environment. In Hopsworks, we maintain a docker image built on top of Ubuntu Linux distribution. You can run generic bash commands on top of the project environment from the UI or REST API. ### Setting up the bash script and artifacts from the UI To use the UI, navigate to the Python environment in the Project settings. In the Python environment page, navigate to custom commands. From the UI, you can write the bash commands in the textbox provided. These bash commands will be uploaded and executed when building your new environment. You can include build artifacts e.g., binaries that you would like to execute or include when building the environment. See Figure 1.

Writing custom commands and uploading build artifacts in the UI
Figure 1: You can write custom commands and upload build artifacts from the UI

## Code You can also run the custom commands using the REST API. From the REST API, you should provide the path, in HOPSFS, to the bash script and the artifacts(comma separated string of paths in HopsFs). The REST API endpoint for running custom commands is: `hopsworks-api/api/project//python/environments//commands/custom` and the body should look like this: ```json { "commandsFile": "", "artifacts": "" } ``` ## What to include in the bash script There are few important things to be aware of when writing the bash script: - The first line of your bash script should always be `#!/bin/bash` (known as shebang) so that the script can be interpreted and executed using the Bash shell. - You can use `apt`, `apt-get` and `deb` commands to install packages. You should always run these commands with `sudo`. In some cases, these commands will ask for user input, therefore you should provide the input of what the command expects, e.g., `sudo apt -y install`, otherwise the build will fail. We have already configured `apt-get` to be non-interactive - The build artifacts will be copied to `srv/hops/build`. You can use them in your script via this path. This path is also available via the environment variable `BUILD_DIR`. If you want to use many artifacts it is advisable to create a zip file and upload it to HopsFS in one of your project datasets. You can then include the zip file as one of the artifacts. - The Python environment is located in `/srv/hops/anaconda/envs/hopsworks_environment`. It is a virtualenv rather than a conda environment, and the path is unchanged so existing scripts keep working. You can install or uninstall packages in it using pip like: `/srv/hops/anaconda/envs/hopsworks_environment/bin/pip install spotify==0.10.2`. If the command requires some input, write the command together with the expected input otherwise the build will fail. ## Installing a compiler The PyTorch and Ray environments contain a C and C++ compiler, since they compile code at runtime for `torch.compile` and for CUDA extensions. The other environments do not, because `build-essential`, and every Ubuntu `-dev` package, depends on the Linux kernel headers, which account for most of the vulnerabilities reported against the images and are never patched within an Ubuntu release. In an environment without a compiler, a library that is published only as a source distribution fails to install: ```text × Failed to build `thriftpy2==0.5.2` ╰─▶ Call to `setuptools.build_meta:__legacy__.build_wheel` failed (exit status: 1) error: [Errno 2] No such file or directory: 'cc' ``` Install the compiler, build the library and remove the compiler again in the same script. The library keeps working, since what it needs at runtime are the shared libraries it links against, and the environment does not end up carrying the kernel headers. ```bash #!/bin/bash set -e sudo apt-get update sudo apt-get install -y build-essential uv pip install --no-cache --python /srv/hops/anaconda/envs/hopsworks_environment/bin/python thriftpy2==0.5.2 sudo apt-get purge -y --auto-remove build-essential sudo apt-get clean ``` Leave the purge out if you want to install more such libraries from the UI afterwards, or if the library loads a shared library owned by a `-dev` package you installed, since `--auto-remove` takes that with it. Do not purge in the PyTorch and Ray environments, where it would remove the compiler they came with. ## Making custom-command builds faster A custom-command build repeats all of its work every time, including compilation and downloads that have not changed. Two things can be declared in the environment variables you supply alongside the script. Both are read as build directives and never become environment variables in the image. !!! note Both need a cluster where an administrator has enabled the persistent BuildKit daemon. Without it there is nothing for a cache to survive in. See [Python Environment Build Performance](../../../setup_installation/admin/build_performance.md). ### Caching a toolchain `HOPSWORKS_BUILD_CACHE` names the toolchain caches your script should get, comma separated: ```text HOPSWORKS_BUILD_CACHE=ccache,maven ``` | Name | Cached | | --- | --- | | `uv`, `pip` | Python package downloads | | `ccache`, `sccache` | C and C++ compiler output | | `maven`, `gradle` | Java and Scala dependencies | | `cargo` | Rust crates and build output | | `npm` | Node packages | | `go` | Go modules | Each one mounts a directory that survives between builds and points the tool at it. Your script does not need to configure anything; it just needs to use the tool normally. This speeds up the work inside the step even though the step itself still re-runs. That is the point: a script that compiles a native extension pays the download and compile cost once rather than on every build. Every cache is scoped to your project, and only one build in your project uses a given cache at a time, so two builds cannot corrupt a shared repository. An unrecognised name fails the build and tells you which names are accepted, rather than being ignored. ### Reusing the whole layer By default a custom-command build never reuses its previous result, because a script can fetch anything and nothing declares what it fetched. An unchanged script that installs `curl | bash` from a URL, or `apt-get install` from a moving repository, does not produce the same thing a month later. If your script genuinely fetches nothing that can change, you can say so: ```text HOPSWORKS_BUILD_HERMETIC=true ``` An unchanged script then reuses its previous layer outright, which takes the step to near zero. Everything else about the step is already accounted for: the base image, your script, and your uploaded artifacts all change the result when they change. What you are asserting is the one thing the platform cannot check for you, which is that nothing your script reaches out to will change underneath it. !!! warning This only takes effect if your administrator has allowed such assertions on the cluster. If it has not been enabled, the setting is ignored and the layer is rebuilt as usual. If you are not certain, leave it out and use `HOPSWORKS_BUILD_CACHE` instead. That speeds up the work without assuming anything about the outside world. ### Referencing a secret A value in the environment variables file becomes an `ENV` instruction, which is image configuration: readable with `docker inspect` by anyone who can pull the image. To pass a credential to your script without it entering the image, reference one of your own Hopsworks secrets: ```text MY_TOKEN=secret:my_secret_name ``` The value is mounted only for the step that runs your script and is exported into its environment. It never becomes an `ENV`, and it is not recorded in the image history. ================================================================================ # Environment History Source: https://docs.hopsworks.ai/latest/user_guides/projects/python/environment_history/ # Environment History Hopsworks comes with several prepackaged Python environments that contain libraries for data engineering, machine learning, and more general data science use-cases. Hopsworks also offers the ability to install additional packages from various sources, such as using the pip package manager and public or private git repository. The Python virtual environment is shared by different members of the project. When a member of the project introduces a change to the environment i.e., installs/uninstalls a library, a new environment is created and it becomes a defacto environment for everyone in the project. It is therefore important to track how the environment has been changing over time i.e., what libraries were installed, uninstalled, upgraded, or downgraded when the environment was created and who introduced the changes. In this guide, you will learn how you can track the changes of your Python environment. ## Viewing python environment history in the UI The Python environment evolves over time as libraries are installed, uninstalled, upgraded, and downgraded. To assist in tracking the changes in the environment, you can see the environment history in the UI. You can view what changes were introduced at each point a new environment was created. Hopsworks keeps a YAML file for each version of the environment, as a record of what that version contained. Importing such a file back into an environment is no longer supported, so to return to an earlier version, clone an environment and install the same packages again. To see the differences between environments click on the button as shown in figure 1. You will then see the difference between the environment and the previous environment it was created from.

Python environment history
Figure 1: You can see the difference between the two environments by clicking on the button pointed.

If you had built the environment using custom commands you can go back to see what commands were run during the build as shown in figure 2.

Python environment history with custom commands
Figure 2: You can see custom commands that were used to create the environment by clicking on the button pointed.

================================================================================ # Run Python Notebook Source: https://docs.hopsworks.ai/latest/user_guides/projects/jupyter/python_notebook/ # How To Run A Python Notebook ## Introduction Jupyter is provided as a service in Hopsworks, providing the same user experience and features as if run on your laptop. - Supports JupyterLab and the classic Jupyter front-end - Configured with Python3, PySpark and Ray kernels ## Step 1: Jupyter dashboard The image below shows the Jupyter service page in Hopsworks and is accessed by clicking `Jupyter` in the sidebar.

Jupyter dashboard in Hopsworks
Jupyter dashboard in Hopsworks

From this page, you can configure various options and settings to start Jupyter with as described in the sections below. ## Step 2 (Optional): Configure resources Next step is to configure Jupyter, Click `edit configuration` to get to the configuration page and select `Python`. - `Container cores`: Number of cores to allocate for the Jupyter instance - `Container memory`: Number of MBs to allocate for the Jupyter instance !!! notice "Configured resource pool is shared by all running kernels. If a kernel crashes while executing a cell, try increasing the Container memory."

Resource configuration for the Python kernel
Resource configuration for the Python kernel

Click `Save` to save the new configuration. ## Step 3 (Optional): Configure environment, root folder and automatic shutdown Before starting the server there are three additional configurations that can be set next to the `Run Jupyter` button. The environment that Jupyter should run in needs to be configured. Select the environment that contains the necessary dependencies for your code.

Configure environment
Configure environment

The runtime of the Jupyter instance can be configured, this is useful to ensure that idle instances will not be hanging around and keep allocating resources. If a limited runtime is not desirable, this can be disabled by setting `no limit`.

Configure maximum runtime
Configure maximum runtime

The root path from which to start the Jupyter instance can be configured. By default it starts by setting the `/Jupyter` folder as the root.

Configure root folder
Configure root folder

## Step 4: (Kueue enabled) Select a Queue If the cluster is installed with Kueue enabled, you will need to select a queue in which the notebook should run. This can be done from `Advanced options`, in the `Scheduler` section of the full configuration page. ![Default queue for job](../../../assets/images/guides/project/scheduler/job_queue.png) ## Step 5: Start Jupyter Start the Jupyter instance by clicking the `Run Jupyter` button.

Starting the Jupyter server for a Python notebook
Starting the Jupyter server for a Python notebook

!!! info "Account-level environment variables" Variables defined under [Account settings → Environment variables][account-level-environment-variables] are injected into every Jupyter session you start. Read them in your notebook with `os.environ["MY_KEY"]`. ## Accessing project data !!! notice "Recommended approach if `/hopsfs` is mounted" If your Hopsworks installation is configured to mount the project datasets under `/hopsfs`, which it is in most cases, then please refer to this section. If the file system is not mounted, then project files can be localized using the [download api][hopsworks_common.core.dataset_api.DatasetApi.download] to localize files in the current working directory. ### Absolute paths The project datasets are mounted under `/hopsfs`, so you can access `data.csv` from the `Resources` dataset using `/hopsfs/Resources/data.csv` in your notebook. Shared datasets are accessible at `/hopsfs/shared-datasets//` if HopsFS is mounted. The shared datasets directory is also available through the `SHARED_DATASETS_DIR` environment variable. ### Relative paths The notebook's working directory is the folder it is located in. For example, if it is located in the `Resources` dataset, and you have a file named `data.csv` in that dataset, you simply access it using `data.csv`. Also, if you write a local file, for example `output.txt`, it will be saved in the `Resources` dataset. ## Going Further You can learn how to [install a library](../python/python_install.md) so that it can be used in a notebook. ================================================================================ # Run PySpark Notebook Source: https://docs.hopsworks.ai/latest/user_guides/projects/jupyter/spark_notebook/ # How To Run A PySpark Notebook ## Introduction Jupyter is provided as a service in Hopsworks, providing the same user experience and features as if run on your laptop. - Supports JupyterLab and the classic Jupyter front-end - Configured with Python3, PySpark and Ray kernels ## Step 1: Jupyter dashboard The image below shows the Jupyter service page in Hopsworks and is accessed by clicking `Jupyter` in the sidebar.

Jupyter dashboard in Hopsworks
Jupyter dashboard in Hopsworks

From this page, you can configure various options and settings to start Jupyter with as described in the sections below. ## Step 2: A Spark environment must be configured The PySpark kernel will only be available if Jupyter is configured to use the `spark-feature-pipeline` or an environment cloned from it. You can easily refer to the green ticks as to what kernels are available in which environment.

Select an environment with PySpark kernel enabled
Select an environment with PySpark kernel enabled

## Step 3 (Optional): Configure spark properties Next step is to configure the Ray properties to be used in Jupyter, Click `edit configuration` to get to the configuration page and select `Ray`. ### Resource and compute Resource allocation for the Spark driver and executors can be configured, also the number of executors and whether dynamic execution should be enabled. - `Driver memory`: Number of cores to allocate for the Spark driver - `Driver virtual cores`: Number of MBs to allocate for the Spark driver - `Executor memory`: Number of cores to allocate for each Spark executor - `Executor virtual cores`: Number of MBs to allocate for each Spark executor - `Dynamic/Static`: Run the Spark application in static or dynamic allocation mode (see [spark docs](https://spark.apache.org/docs/latest/configuration.html#dynamic-allocation) for details).

Resource configuration for the Spark kernels
Resource configuration for the Spark kernels

### Attach files or dependencies Additional files or dependencies required for the Spark job can be configured. - `Additional archives`: List of zip or .tgz files that will be locally accessible by the application - `Additional jars`: List of .jar files to add to the CLASSPATH of the application - `Additional python dependencies`: List of .py, .zip or .egg files that will be locally accessible by the application - `Additional files`: List of files that will be locally accessible by the application

File configuration for the Spark kernels
File configuration for the Spark kernels

Line-separates [properties](https://spark.apache.org/docs/3.1.1/configuration.html) to be set for the Spark application. For example, changing the configuration variables for the Kryo Serializer or setting environment variables for the driver, you can set the properties as shown below.

File configuration for the Spark kernels
Additional Spark configuration

Click `Save` to save the new configuration. ## Step 4 (Optional): Configure root folder and automatic shutdown Before starting the server there are two additional configurations that can be set next to the `Run Jupyter` button. The runtime of the Jupyter instance can be configured, this is useful to ensure that idle instances will not be hanging around and keep allocating resources. If a limited runtime is not desirable, this can be disabled by setting `no limit`.

Configure maximum runtime
Configure maximum runtime

The root path from which to start the Jupyter instance can be configured. By default it starts by setting the `/Jupyter` folder as the root.

Configure root folder
Configure root folder

## Step 5: (Kueue enabled) Select a Queue Currently we do not have Kueue support for Spark. You do not need to select a queue to run the notebook in. ## Step 5: Start Jupyter Start the Jupyter instance by clicking the `Run Jupyter` button.

Starting the Jupyter server for a Spark notebook
Starting the Jupyter server for a Spark notebook

## Step 6: Access Spark UI Navigate back to Hopsworks and a Spark session will have appeared, click on the `Spark UI` button to go to the Spark UI.

Access Spark UI and see application logs
Access Spark UI and see application logs

## Step 7: Saved application logs While the Spark session runs, the `Live Logs` button on its row streams the driver and executor output. When the session ends, each executor copies its logs to the `Logs/Jupyter//spark/` directory of the project, and the Jupyter page lists the application under `Spark and Ray applications` in the `Saved Logs` card. Click `View` to read the executor files, one per executor and stream, or delete the directory when it is no longer needed. The driver runs inside the Jupyter server, so its output is part of the server log, which is saved to `Logs/Jupyter/` when the server stops. ## Accessing project data ### Read directly from the filesystem (recommended) To read a dataset in your project using Spark, use the full filesystem path where the data is stored. For example, to read a CSV file named `data.csv` located in the `Resources` dataset of a project called `my_project`: ```python df = spark.read.csv( "/Projects/my_project/Resources/data.csv", header=True, inferSchema=True ) df.show() ``` ### Additional files Different files can be attached to the jupyter session and made available in the `/srv/hops/artifacts` folder when the PySpark kernel is started. This configuration is mainly useful when you need to add additional configuration such as jars that needs to be added to the CLASSPATH. When reading data in your Spark application, it is recommended to use the Spark read API as previously demonstrated, since this reads from the filesystem directly, whereas `Additional files` configuration options will download the files in its entirety and is not a scalable option. ## Going Further You can learn how to [install a library](../python/python_install.md) so that it can be used in a notebook. ================================================================================ # Run Ray Notebook Source: https://docs.hopsworks.ai/latest/user_guides/projects/jupyter/ray_notebook/ # How To Run A Ray Notebook ## Introduction Jupyter is provided as a service in Hopsworks, providing the same user experience and features as if run on your laptop. - Supports JupyterLab and the classic Jupyter front-end - Configured with Python3, PySpark and Ray kernels !!!warning "Enable Ray" Ray is gated behind the `ray_enabled` [configuration variable](../../../setup_installation/admin/variables.md), which an administrator has to turn on. Until it is enabled, the Ray kernel and the Ray configuration are not offered in Jupyter. Support for Ray also needs to be explicitly enabled by adding the following option in the `values.yaml` file for the deployment: ```yaml global: ray: enabled: true ``` ## Step 1: Jupyter dashboard The image below shows the Jupyter service page in Hopsworks and is accessed by clicking `Jupyter` in the sidebar.

Jupyter dashboard in Hopsworks
Jupyter dashboard in Hopsworks

From this page, you can configure various options and settings to start Jupyter with as described in the sections below. ## Step 2 (Optional): Configure Ray Next step is to configure the Ray cluster configuration that will be created when you start Ray session later in Jupyter. Click `edit configuration` to get to the configuration page and select `Ray`. ### Resource and compute Resource allocation for the Driver and Workers can be configured. !!! notice "Using the resources in the Ray script" The resource configurations describe the cluster that will be provisioned when launching the Ray job. User can still provide extra configurations in the job script using `ScalingConfig`, i.e. `ScalingConfig(num_workers=4, trainer_resources={"CPU": 1}, use_gpu=True)`. - `Driver memory`: Memory in MBs to allocate for Driver - `Driver virtual cores`: Number of cores to allocate for the Driver - `Driver GPUs`: Number of GPUs to allocate for the Driver - `Worker memory`: Memory in MBs to allocate for each worker - `Worker virtual cores`: Number of cores to allocate for each worker - `Worker GPUs`: Number of GPUs to allocate for each worker - `Min workers`: Minimum number of workers to start with - `Max workers`: Maximum number of workers to scale up to

Resource configuration for
the Ray kernels
Resource configuration for the Ray kernels

Runtime environment and Additional files required for the Ray job can also be provided. - `Runtime Environment (Optional)`: A [runtime environment](https://docs.ray.io/en/latest/ray-core/handling-dependencies.html#runtime-environments) describes the dependencies required for the Ray job including files, packages, environment variables, and more. This is useful when you need to install specific packages and set environment variables for this particular Ray job. It should be provided as a YAML file. You can select the file from the project or upload a new one. - `Additional files`: List of other files required for the Ray job. These files will be placed in `/srv/hops/ray/job`.

Runtime
environment and additional files
Runtime configuration and additional files for Ray jupyter session

Click `Save` to save the new configuration. ## Step 3 (Optional): Configure max runtime and root path Before starting the server there are two additional configurations that can be set next to the `Run Jupyter` button. The runtime of the Jupyter instance can be configured, this is useful to ensure that idle instances will not be hanging around and keep allocating resources. If a limited runtime is not desirable, this can be disabled by setting `no limit`.

Configure maximum runtime
Configure maximum runtime

The root path from which to start the Jupyter instance can be configured. By default it starts by setting the `/Jupyter` folder as the root.

Configure root folder
Configure root folder

## Step 4: Select the environment Hopsworks provides a variety of environments to run Jupyter notebooks. Select the environment you want to use by clicking on the dropdown menu. In order to be able to run a Ray notebook, you need to select the environment that has the Ray kernel installed. Environment with Ray kernel have a `Ray Enabled` label next to them. ## Step 5: (Kueue enabled) Select a Queue If the cluster is installed with Kueue enabled, you will need to select a queue in which the notebook should run. This can be done from `Advanced options`, in the `Scheduler` section of the full configuration page. ![Default queue for job](../../../assets/images/guides/project/scheduler/job_queue.png) ## Step 6: Start Jupyter Start the Jupyter instance by clicking the `Run Jupyter` button. ## Running Ray code in Jupyter Once the Jupyter instance is started, you can create a new notebook by clicking on the `New` button and selecting `Ray` kernel. You can now write and run Ray code in the notebook. When you first run a cell with Ray code, a Ray session will be started and you can monitor the resources used by the job in the Ray dashboard.

Ray Kernel
Ray Kernel

## Step 7: Access Ray Dashboard When you start a Ray session in Jupyter, a new application will appear in the Jupyter page. The notebook name from which the session was started is displayed. You can access the Ray UI by clicking on the `Ray Dashboard` and a new tab will be opened. The Ray dashboard is only available while the Ray kernel is running. You can kill the Ray session to free up resources by shutting down the kernel in Jupyter. In the Ray Dashboard, you can monitor the resources used by code you are running, the number of workers, logs, and the tasks that are running.

Access Ray Dashboard
Access Ray Dashboard for Jupyter Ray session

Ray Dashboard cluster view
The Ray Dashboard for the Jupyter Ray session, cluster view

## Step 8: Saved application logs While the Ray session runs, the `Live Logs` button on its row streams the output of the notebook kernel and of the Ray head and worker pods. When the session ends, the driver and worker process logs of the Ray cluster are saved as `stdout.log` and `stderr.log` in the `Logs/Jupyter//ray/` directory of the project, the same two files a Ray job leaves in `Logs/Ray`. The Jupyter page lists the application under `Spark and Ray applications` in the `Saved Logs` card. Click `View` to read the files or delete the directory when it is no longer needed. The code you run in the notebook executes in the Jupyter server, so its output is part of the server log, which is saved to `Logs/Jupyter/` when the server stops. ## Accessing project data If HopsFS is mounted in the Ray containers, project datasets are available under `/hopsfs`, so you can access `data.csv` from the `Resources` dataset using `/hopsfs/Resources/data.csv`. Shared datasets are accessible at `/hopsfs/shared-datasets//`. The shared datasets directory is also available through the `SHARED_DATASETS_DIR` environment variable. ================================================================================ # Remote Filesystem Driver Source: https://docs.hopsworks.ai/latest/user_guides/projects/jupyter/remote_filesystem_driver/ # Configuring remote filesystem driver ## Introduction We provide two ways to access and persist files in HopsFS from a jupyter notebook: - `hdfscontentsmanager`: With `hdfscontentsmanager` you interact with the project datasets using the dataset api. When you start a notebook using the `hdfscontentsmanager` you will only see the files in the configured root path. - `hopsfsmount`: With `hopsfsmount` all the project datasets are available in the jupyter notebook as a local filesystem. This means you can use native Python file I/O operations (copy, move, create, open, etc.) to interact with the project datasets. When you open the jupyter notebook you will see all the project datasets. ## Configuring the driver To configure the driver you need to have admin role and set the `jupyter_remote_fs_driver` to either `hdfscontentsmanager` or `hopsfsmount`. The default driver is `hdfscontentsmanager`. ================================================================================ # Session Capacity Warnings Source: https://docs.hopsworks.ai/latest/user_guides/projects/jupyter/session_capacity_warnings/ # Session capacity warnings Hopsworks limits how many interactive sessions (Jupyter notebooks, terminals, Streamlit apps) can run concurrently on each cluster. When the limit is being approached, you see one or two badges on the Jupyter, Terminal, and Apps pages. This page explains what the badges mean and what to do when they appear. ## Where the badges appear The badges sit next to the action button on three pages: - **Jupyter**: next to the `Jupyter server` card title and the `Run Jupyter` button. ![WebSocket warnings on the Jupyter page](../../../assets/images/guides/jupyter/websocket-warnings-jupyter.png) - **Terminal**: next to the `Start Terminal` button inside the terminal panel. ![WebSocket warnings on the Terminal panel](../../../assets/images/guides/jupyter/websocket-warnings-terminal.png) - **Apps**: next to the Apps card title and the per-row Start button. ![WebSocket warnings on the Apps page](../../../assets/images/guides/jupyter/websocket-warnings-apps.png) ## What each badge means There are two badges with two scopes. Both can appear in orange (warning) or red (critical), independently of each other. | Scope | Badge text | Color | Meaning | | --- | --- | --- | --- | | Instance | `This instance is close to its limit` | Orange | The Hopsworks instance you would be routed to has few free session slots left. New sessions can still start but may fail soon. | | Instance | `This instance is full` | Red | The Hopsworks instance you would be routed to has no free session slots. New sessions may still start if another instance has capacity; refreshing or signing back in may land you on that pod. See [What happens when a badge turns red][what-happens-when-a-badge-turns-red] for which buttons disable. | | Cluster | `Cluster is close to its limit` | Orange | All Hopsworks instances together have few free session slots. New sessions may succeed on a less-busy pod, but a refresh or retry might land on a full one. | | Cluster | `Cluster is full` | Red | All Hopsworks instances together have no free session slots. New sessions are rejected platform-wide. | The instance badge tracks the pod that your browser session is currently bound to. The cluster badge sums across every Hopsworks instance pod. On a single-instance deployment the two badges always agree, since the only instance is the cluster. ## What happens when a badge turns red - `Run Jupyter`, `Start Terminal`, and per-row `Start App` buttons are disabled while the **cluster** badge is red (no instance has capacity to serve a new session). While only the instance badge is red the buttons stay enabled because refreshing or signing back in may land you on a different instance pod that still has capacity. - An already-running Jupyter server keeps working. Opening a new notebook tab inside a running Jupyter server may still fail if the pod that hosts it is at its per-session cap: the new tab's kernel WebSocket upgrade is closed with a `1013 TRY_AGAIN_LATER` close rather than attaching. This is distinct from starting a new session (a Jupyter server, terminal, or Streamlit app), which is gated up front and rejected with a `WebSocket pool full` (HTTP 503) error when the pod has no free slot. - An already-running terminal session keeps working. Reconnecting after a network blip while the badge is red surfaces the error rather than silently retrying. ## What to do - **Orange** is informational. Sessions can still be started. Save your work and avoid opening more tabs than you need. - **Red on the instance badge, orange or no cluster badge**: a different pod has capacity. Refresh the page or sign out and back in to land on a different instance pod. - **Red on both badges**: the platform is at capacity. Close notebook tabs, terminals, and apps you are no longer using to free sessions, or contact your administrator about raising the platform's session cap. By default the proxy does not reap idle sessions, so a slot is only released when its connection actually closes. ## Why the limits exist Each WebSocket session (a Jupyter kernel connection, a terminal connection, or a Streamlit app connection) counts as one open connection against a per-pod cap inside the Hopsworks instance. The cap is sized so the platform can stay responsive when many users are active at once. Letting new sessions queue indefinitely would leave users staring at a loading spinner that never resolves; rejecting cleanly with the badge is the safer behavior. Administrators can change the pool sizing per cluster. See [WebSocket Proxy Pool][websocket-proxy-pool] in the admin guide for the Helm values and Grafana panels that drive the limits. ================================================================================ # Browser Terminal Source: https://docs.hopsworks.ai/latest/user_guides/compute/terminal/ # Terminal ## Introduction Hopsworks provides a browser terminal that runs inside your project. The terminal is a dedicated pod running under your project user, with your HopsFS home directory mounted at `/hopsfs/Users/`. Files you create there are stored in the project file system and survive the terminal session. ## Prerequisites The terminal is disabled by default. An administrator must set the `enable_terminal` [configuration variable][cluster-configuration] to `true`. The related `terminal_*` variables (image, session length, shared memory size) are listed in the [configuration reference][cluster-configuration-variables-reference]. ## Start a terminal Click on `>_ Terminal` in the top navigation bar of your project. A panel opens on the right side of the page. Select the environment for the session, `Python` or `Spark`, and set the CPU cores and memory for the terminal pod. Then click `Start Terminal`.
Terminal start form
Configure and start a terminal session
The session starts a pod in your project namespace, so it is subject to the same resource quotas and scheduling as jobs and notebooks. A session expires after a fixed lifetime, four hours by default (`terminal_session_hours`). ## Use the terminal Once connected, the panel shows a shell in your HopsFS home directory, with live CPU and memory usage of the terminal pod in the header.
Connected terminal session
A connected terminal session
The session comes with tooling preinstalled: - `hops` initializes the Hopsworks CLI, already pointed at your project. - `git` works against your configured [Git providers][how-to-configure-a-git-provider]; upload an SSH key to `~/.ssh` or add a GitHub access token in Account Settings. - `claude` starts a Claude Code session and `codex` starts an OpenAI Codex CLI session; their logins persist in HopsFS across sessions. - Right-click opens a menu to split the terminal horizontally or vertically (tmux). - Selecting text with the left mouse button copies it to your clipboard. ## Session controls The panel header offers fullscreen, reconnect and collapse controls. Closing the panel does not stop the session; it keeps running until it expires or you stop it. Reopening the panel reconnects to the running session. ================================================================================ # Teleport Sessions Source: https://docs.hopsworks.ai/latest/user_guides/projects/terminal/teleport/ # Teleport Claude Code Sessions Every Hopsworks project has a web terminal that runs Claude Code in a per-user pod. The `hops session` command moves a Claude Code session between your laptop and that pod, so you can start work locally and continue it in the cluster, or the other way round. A session is handed off, not copied. By default `hops session push` moves the one canonical copy of the session to the pod and leaves a marker (a "baton") recording that the pod now holds it. `hops session pull` brings that copy back. This keeps a single source of truth for the transcript, so the two sides never diverge silently. Use `--fork` when you deliberately want a second, independent copy. Transcripts are staged in your own private area of the project's HopsFS (`Users//teleport/`, readable only by you), not in a project-wide location. A session transcript can contain code and file contents, so it is never exposed to other project members. ## Prerequisites - The `hopsworks` Python package on your laptop. Install it with `pip install hopsworks`. - An API key with the `TERMINAL` scope for the target cluster. Run `hops setup --host https://` once to create one and store it locally. A key created through `hops setup` already carries the `TERMINAL` scope, so it can start a terminal. - A running Claude Code session on your laptop for `push`, or a staged session in the project for `pull`. ## Commands | Command | What it does | | --- | --- | | `hops session push` | Hand the current session up to a terminal pod and open it in the browser. | | `hops session pull` | Bring a session back down to your laptop. | | `hops session new` | Start a fresh session directly on a terminal pod. | | `hops session list` | Show your staged sessions and where each one currently lives. | | `hops session mirror` | Stream the live pod terminal on your laptop, read-only by default. | | `hops session reset` | Forget how git sync authenticates the pod and remove the staged SSH keys, so the next push asks again. | | `hops session extend` | Give the terminal session more time before the cluster stops it. | | `hops session stop` | Stop this project's terminal pod and every tab in it. | ### Push a session to the cluster Run this from the directory where your Claude Code session lives: ```bash hops session push ``` The command uploads the transcript, starts the project's terminal pod if it is not already running, and opens the terminal in your browser. Once the Terminal tab is open, the session lands on the pod as its own tab and resumes there. Switch between tabs by clicking them, or from the keyboard with Alt+PgUp and Alt+PgDn; Ctrl+PgUp and Ctrl+PgDn do the same wherever the browser does not keep them for its own tabs, such as an installed app window or `hops session mirror`. Until a tab is open the push stays staged, and the command prints the terminal URL and the manual landing steps. A terminal started by `hops session push` or `hops session new` lasts 12 hours, longer than the web terminal's default, since a session handed to the pod is often left to run. The command prints how long the terminal has left; `hops session extend` adds the cluster's default extension, and `hops session extend --hours N` adds N hours, within the per-request limit the cluster sets. Running `hops session push` again is safe: it re-stages the same session. It refuses only when the staged copy holds lines this machine does not have, until you `hops session pull` them back or pass `--force`. Useful options: - `--fork` copies the session instead of handing it off, so your local copy stays canonical. - `--model ` resumes the session on the pod with a specific model. - `--prompt ""` feeds the resumed session a first instruction. - `--no-open` prints the terminal URL instead of opening a browser. ### Pull a session back to your laptop ```bash hops session pull ``` Run from the same directory you pushed from to reclaim that directory's session. To reclaim a session staged from a different directory, pass its id: `hops session pull `. `pull` refuses to take a session that a live pod has landed, so you do not accidentally split it in two. Stop the pod first, or pass `--force` to take the baton anyway. A push the pod never landed, because no Terminal tab was opened, is reclaimed without either. Running `hops session pull` again is safe: it finds nothing new and leaves everything as it is. If both sides changed, `pull` stops and asks you to pick with `--ours` or `--theirs`; the copy you do not keep is parked to a sidecar file rather than lost. ### Start a fresh session on the cluster ```bash hops session new ``` This starts a new Claude Code session directly on a terminal pod, with no local transcript. `--model` and `--prompt` work the same as for `push`. ### See where your sessions live ```bash hops session list ``` Add `--all` to list every session you have staged across all directories. The store is per-user and private, so you only ever see your own sessions. ### Watch a running session from your laptop ```bash hops session mirror ``` This attaches to the live pod terminal over its WebSocket and streams it on your laptop. It is read-only by default; pass `--write` to type into it. Press `Ctrl-]` to detach. ### Stop the terminal pod ```bash hops session stop ``` This stops the project's terminal pod and every tab in it, from your laptop, without needing Kubernetes access. ## Optional: sync your Git checkout to the pod When you push a session, `hops session` can also reproduce your local Git checkout on the pod, so the landed session starts in the same repository, on the same branch, at the commit you pushed. This applies when you run `hops session push` from the root directory of the repository. From anywhere else the command says it is not a git repository, skips the sync and pushes the session without a checkout. The pod clones the repository into your private HopsFS home, `/hopsfs/Users//`, the first time, then fetches from the branch's remote and checks out the branch you had locally; the session resumes inside that checkout. This is opt-in. The first time it applies, the CLI asks whether to sync (always, this time, not now, or never) and remembers your answer. Pressing Enter picks always, so later pushes sync without asking. It then asks how the pod should authenticate to your Git host, and remembers that too: | Choice | What happens | | --- | --- | | An existing SSH private key | You confirm the key path (the key `ssh` would use is suggested). The key is uploaded once into your private `Users//.ssh/` area and the pod clones over SSH. | | A new SSH key created for Hopsworks | The CLI runs `ssh-keygen` to create a passphrase-free `ed25519` key at `~/.ssh/hopsworks_teleport_ed25519`, adds its public key to your GitHub account with `gh ssh-key add` when the GitHub CLI is logged in (otherwise it prints the public key for you to add), and stages it like an existing key. Offered on Linux, macOS and WSL; on Windows without `ssh-keygen`, point to an existing key instead. | | A personal access token | The token is registered with Hopsworks (see below) and the pod clones over HTTPS with it, so no key leaves your machine. This is the only option when your remote already uses HTTPS; an SSH remote is rewritten to its HTTPS form for the pod. | Commit and push your local work first, since the pod fetches from the remote, not from your laptop. The CLI offers to stage tracked files and commit and push if the tree is dirty. !!! note "Passphrase-protected keys" The pod runs no SSH agent, so a key that needs a passphrase is not supported. Use a passphrase-free key or a personal access token. ### Register a Git provider token Hopsworks keeps one personal access token per Git provider host in your account, used by jobs and the web terminal for HTTPS git operations. The teleport flow registers one for you when you choose the token option; you can also manage it directly: ```bash hops git provider list hops git provider set --provider github --username hops git provider delete --provider github ``` `set` prompts for the token without echoing it, defaults the host to the provider's public host (`github.com`, `gitlab.com`, `bitbucket.org`), and leaves an existing registration alone unless you pass `--force`. The same token can be registered in the UI under **Account settings > Git providers**. ### Start over with a different key or token ```bash hops session reset ``` `reset` forgets the remembered consent, authentication method and key path, and removes the SSH keys staged in your private `Users//.ssh/` area. The next `hops session push` asks the questions again and uploads the key you pick afresh, which is how you replace a key that was staged under the same name. Staged sessions are untouched, and so is a registered provider token; remove that with `hops git provider delete`. A terminal pod that is already running keeps the copy of the key it made when it first landed a session, so run `hops session stop` before the next push if you want it to pick up the new key. ## How teleport keeps a session private - Transcripts are staged in your own HopsFS home under `Users//teleport/`, created mode `0700`, so only you can read them. - Attaching to a terminal from the CLI uses a short-lived, one-time token that is bound to the specific terminal session it was minted for. It cannot be reused and cannot be presented against another session. - Staged transcripts are removed by a background cleaner after a retention period, so old sessions do not accumulate in your home. The default is 7 days; an administrator sets it with the `teleport_ttl_days` cluster variable. ================================================================================ # Run Python Job Source: https://docs.hopsworks.ai/latest/user_guides/projects/jobs/python_job/ # How To Run A Python Job ## Introduction This guide will describe how to configure a job to execute a python script inside the cluster. All members of a project in Hopsworks can launch the following types of applications through a project's Jobs service: - Python - Apache Spark - Ray Launching a job of any type is very similar process, what mostly differs between job types is the various configuration parameters each job type comes with. Hopsworks support scheduling jobs to run on a regular basis, e.g backfilling a Feature Group by running your feature engineering pipeline nightly. Scheduling can be done both through the UI and the python API, checkout [our Scheduling guide](schedule_job.md). ## UI ### Step 1: Jobs overview The image below shows the Jobs overview page in Hopsworks and is accessed by clicking `Jobs` in the sidebar.

Jobs overview
Jobs overview

### Step 2: Create new job dialog Click `New Job` and the following dialog will appear.

Create new job dialog
Create new job dialog

### Step 3: Set the job type The `Type` radio offers `PYTHON` and `SPARK`, and `PYTHON` is selected by default. Leave it on `PYTHON` to configure a Python job.

Select Python job type
Select Python job type

### Step 4: Set the script Next step is to select the python script to run. You can either select `From project`, if the file was previously uploaded to Hopsworks, or `Upload new file` which lets you select a file from your local filesystem as demonstrated below. By default, the job name is the same as the file name, but you can customize it as shown.

Configure program
Configure program

### Step 5 (optional): Set the Python script arguments In the job settings, you can specify arguments for your Python script. Remember to handle the arguments inside your Python script.

Configure Python script arguments
Configure Python script arguments

### Step 6 (optional): Additional configuration Click `Advanced options` in the dialog to open the full job configuration page. There you can also set the following configuration settings for a `PYTHON` job. - `Environment`: The python environment to use - `Container memory`: The amount of memory in MB to be allocated to the Python script - `Container cores`: The number of cores to be allocated for the Python script - `Additional files`: List of files that will be locally accessible in the working directory of the application. Only recommended to use if project datasets are not mounted under `/hopsfs`. You can always modify the arguments in the job settings. - `Environment variables`: Custom `KEY=VALUE` pairs that are set on the container for every execution of this job. Read them in your Python code with `os.environ["MY_KEY"]`. See [Environment variables](#environment-variables) below.

Additional configuration
Additional configuration

### Step 7: (Kueue enabled) Select a Queue If the cluster is installed with Kueue enabled, you will need to select a queue in which the job should run. This can be done from `Advanced options`, in the `Scheduler` section. ![Default queue for job](../../../assets/images/guides/project/scheduler/job_queue.png) ### Step 8: Execute the job Now click the `Run` button to start the execution of the job. Then open the `Executions` tab to see the list of all executions. Once the execution is finished, click on `Logs` to see the logs for the execution.

Start job execution
Start job execution

## Code ### Step 1: Upload the Python script This snippet assumes the python script is in the current working directory and named `script.py`. It will upload the python script to the `Resources` dataset in your project. ```python import hopsworks project = hopsworks.login() dataset_api = project.get_dataset_api() uploaded_file_path = dataset_api.upload("script.py", "Resources") ``` ### Step 2: Create Python job In this snippet we get the `JobsApi` object to get the default job configuration for a `PYTHON` job, set the python script and override the environment to run in, and finally create the `Job` object. ```python jobs_api = project.get_job_api() py_job_config = jobs_api.get_configuration("PYTHON") # Set the application file py_job_config["appPath"] = uploaded_file_path # Override the python job environment py_job_config["environmentName"] = "python-feature-pipeline" job = jobs_api.create_job("py_job", py_job_config) ``` ### Step 3: Execute the job In this snippet we execute the job synchronously, that is wait until it reaches a terminal state, and then download and print the logs. ```python # Run the job execution = job.run(await_termination=True) # Download logs out, err = execution.download_logs() f_out = open(out, "r") print(f_out.read()) f_err = open(err, "r") print(f_err.read()) ``` ## Configuration The following table describes the job configuration parameters for a PYTHON job. `conf = jobs_api.get_configuration("PYTHON")` | Field | Type | Description | Default | | --- | --- | --- | --- | | `conf['type']` | string | Type of the job configuration | `"pythonJobConfiguration"` | | `conf['appPath']` | string | Project relative path to script (e.g., `Resources/foo.py`) | `null` | | `conf['defaultArgs']` | string | Arguments to pass to the script. Will be overridden if arguments are passed explicitly via `Job.run(args="...")` | `null` | | `conf['environmentName']` | string | Name of the project Python environment to use | `"pandas-training-pipeline"` | | `conf['resourceConfig']['cores']` | float | Number of CPU cores to be allocated | `1.0` | | `conf['resourceConfig']['memory']` | int | Number of MBs to be allocated | `2048` | | `conf['resourceConfig']['gpus']` | int | Number of GPUs to be allocated | `0` | | `conf['logRedirection']` | boolean | Whether logs are redirected | `true` | | `conf['jobType']` | string | Type of job | `"PYTHON"` | | `conf['files']` | string | Comma-separated string of HDFS path(s) to files to be made available to the application. Example: `hdfs:///Project//Resources/file1.py,...` | `null` | ## Accessing project data !!! notice "Recommended approach if `/hopsfs` is mounted" If your Hopsworks installation is configured to mount the project datasets under `/hopsfs`, which it is in most cases, then please refer to this section instead of the `Additional files` property to reference file resources. ### Absolute paths The project datasets are mounted under `/hopsfs`, so you can access `data.csv` from the `Resources` dataset using `/hopsfs/Resources/data.csv` in your script. Shared datasets are accessible at `/hopsfs/shared-datasets//` if HopsFS is mounted. The shared datasets directory is also available through the `SHARED_DATASETS_DIR` environment variable. ### Relative paths The script's working directory is the folder it is located in. For example, if it is located in the `Resources` dataset, and you have a file named `data.csv` in that dataset, you simply access it using `data.csv`. Also, if you write a local file, for example `output.txt`, it will be saved in the `Resources` dataset. ## Environment variables User-defined environment variables can be attached to a Python job under the *Environment variables* panel of the advanced options page. Each entry is a `KEY=VALUE` pair that is set on the container for every execution of the job. !!! info "Account-level variables also apply" Variables defined under [Account settings → Environment variables][account-level-environment-variables] are also injected into this job. A value set on the job overrides the account-level value with the same name for this job only. ```python # inside your Python job script import os api_key = os.environ["MY_API_KEY"] region = os.environ.get("AWS_REGION", "us-east-1") print(f"Using {api_key[:4]}*** in {region}") ``` Scheduled and backfill runs also receive `HOPS_LOGICAL_DATE`, `HOPS_START_TIME` and `HOPS_END_TIME` describing the data interval they should process. See [Scheduling][logical-time-and-data-intervals] and [Batch feature pipelines][batch-feature-pipelines]. !!! warning "Reserved names" Names starting with `HOPS_` are reserved by the scheduler. Setting them in the *Environment variables* panel overrides the scheduler value for every execution. The UI shows a warning callout when you do this. !!! api "API reference" - [`Project.get_job_api`][hopsworks_common.project.Project.get_job_api] - [`JobsApi`][hopsworks.core.job_api.JobsApi] - [`get_configuration`][hopsworks.core.job_api.JobsApi.get_configuration] - [`create_job`][hopsworks.core.job_api.JobsApi.create_job] - [`Job`][hopsworks_common.job.Job] - [`run`][hopsworks_common.job.Job.run] - [`Execution`][hopsworks_common.execution.Execution] - [`download_logs`][hopsworks_common.execution.Execution.download_logs] - [`DatasetApi.upload`][hopsworks_common.core.dataset_api.DatasetApi.upload] Browse the full Python API :material-arrow-right: ================================================================================ # Run Jupyter Notebook Job Source: https://docs.hopsworks.ai/latest/user_guides/projects/jobs/notebook_job/ # How To Run A Jupyter Notebook Job ## Introduction This guide describes how to configure a job to execute a Jupyter Notebook (.ipynb) and visualize the evaluated notebook. All members of a project in Hopsworks can launch the following types of applications through a project's Jobs service: - Python - Apache Spark - Ray Launching a job of any type is very similar process, what mostly differs between job types is the various configuration parameters each job type comes with. Hopsworks support scheduling jobs to run on a regular basis, e.g backfilling a Feature Group by running your feature engineering pipeline nightly. Scheduling can be done both through the UI and the python API, checkout [our Scheduling guide](schedule_job.md). ## UI ### Step 1: Jobs overview The image below shows the Jobs overview page in Hopsworks and is accessed by clicking `Jobs` in the sidebar.

Jobs overview
Jobs overview

### Step 2: Create new job dialog Click `New Job` and the following dialog will appear.

Create new job dialog
Create new job dialog

### Step 3: Set the job type The `Type` radio offers `PYTHON` and `SPARK`, and `PYTHON` is selected by default. Leave it on `PYTHON` to configure a Jupyter Notebook job.

Select Python job type
Select Python job type

### Step 4: Set the notebook Next step is to select the Jupyter Notebook to run. You can either select `From project`, if the file was previously uploaded to Hopsworks, or `Upload new file` which lets you select a file from your local filesystem as demonstrated below. By default, the job name is the same as the file name, but you can customize it as shown.

Configure program
Configure program

Then click `Create job` to create the job. ### Step 5 (optional): Set the Jupyter Notebook arguments In the job settings, you can specify arguments for your notebook script. Arguments must be in the format of `-p arg1 value1 -p arg2 value2`. For each argument, you must first provide `-p`, followed by the parameter name (e.g. `arg1`), followed by its value (e.g. `value1`). The next step is to read the arguments in the notebook which is explained in this [guide](https://papermill.readthedocs.io/en/latest/usage-parameterize.html).

Configure notebook arguments
Configure notebook arguments

### Step 6 (optional): Additional configuration Click `Advanced options` in the dialog to open the full job configuration page. There you can also set the following configuration settings for a `PYTHON` job. - `Environment`: The python environment to use - `Container memory`: The amount of memory in MB to be allocated to the Jupyter Notebook script - `Container cores`: The number of cores to be allocated for the Jupyter Notebook script - `Additional files`: List of files that will be locally accessible in the working directory of the application. Only recommended to use if project datasets are not mounted under `/hopsfs`. You can always modify the arguments in the job settings.

Set the job type
Set the job type

### Step 7: (Kueue enabled) Select a Queue If the cluster is installed with Kueue enabled, you will need to select a queue in which the job should run. This can be done from `Advanced options`, in the `Scheduler` section. ![Default queue for job](../../../assets/images/guides/project/scheduler/job_queue.png) ### Step 8: Execute the job Now click the `Run` button to start the execution of the job. Then open the `Executions` tab to see the list of all executions.

Start job execution
Start job execution

### Step 9: Visualize output notebook Once the execution is finished, click `Logs` and then `notebook out` to see the logs for the execution.

Visualize output notebook
Visualize output notebook

You can directly edit and save the output notebook by clicking `Open Notebook`. ## Code ### Step 1: Upload the Jupyter Notebook script This snippet assumes the Jupyter Notebook script is in the current working directory and named `notebook.ipynb`. It will upload the Jupyter Notebook script to the `Resources` dataset in your project. ```python import hopsworks project = hopsworks.login() dataset_api = project.get_dataset_api() uploaded_file_path = dataset_api.upload("notebook.ipynb", "Resources") ``` ### Step 2: Create Jupyter Notebook job In this snippet we get the `JobsApi` object to get the default job configuration for a `PYTHON` job, set the jupyter notebook file and override the environment to run in, and finally create the `Job` object. ```python jobs_api = project.get_job_api() notebook_job_config = jobs_api.get_configuration("PYTHON") # Set the application file notebook_job_config["appPath"] = uploaded_file_path # Override the python job environment notebook_job_config["environmentName"] = "python-feature-pipeline" job = jobs_api.create_job("notebook_job", notebook_job_config) ``` ### Step 3: Execute the job In this code snippet, we execute the job with arguments and wait until it reaches a terminal state. ```python # Run the job execution = job.run(args="-p a 2 -p b 5", await_termination=True) ``` ## Configuration The following table describes the job configuration parameters for a PYTHON job. `conf = jobs_api.get_configuration("PYTHON")` | Field | Type | Description | Default | | --- | --- | --- | --- | | `conf['type']` | string | Type of the job configuration | `"pythonJobConfiguration"` | | `conf['appPath']` | string | Project relative path to notebook (e.g., `Resources/foo.ipynb`) | `null` | | `conf['defaultArgs']` | string | Arguments to pass to the notebook.
Will be overridden if arguments are passed explicitly via `Job.run(args="...")`.
Must conform to Papermill format `-p arg1 val1` | `null` | | `conf['environmentName']` | string | Name of the project Python environment to use | `"pandas-training-pipeline"` | | `conf['resourceConfig']['cores']` | float | Number of CPU cores to be allocated | `1.0` | | `conf['resourceConfig']['memory']` | int | Number of MBs to be allocated | `2048` | | `conf['resourceConfig']['gpus']` | int | Number of GPUs to be allocated | `0` | | `conf['logRedirection']` | boolean | Whether logs are redirected | `true` | | `conf['jobType']` | string | Type of job | `"PYTHON"` | | `conf['files']` | string | Comma-separated string of HDFS path(s) to files to be made available to the application. Example: `hdfs:///Project//Resources/file1.py,...` | `null` | ## Accessing project data !!! notice "Recommended approach if `/hopsfs` is mounted" If your Hopsworks installation is configured to mount the project datasets under `/hopsfs`, which it is in most cases, then please refer to this section instead of the `Additional files` property to reference file resources. ### Absolute paths The project datasets are mounted under `/hopsfs`, so you can access `data.csv` from the `Resources` dataset using `/hopsfs/Resources/data.csv` in your notebook. Shared datasets are accessible at `/hopsfs/shared-datasets//` if HopsFS is mounted. The shared datasets directory is also available through the `SHARED_DATASETS_DIR` environment variable. ### Relative paths The notebook's working directory is the folder it is located in. For example, if it is located in the `Resources` dataset, and you have a file named `data.csv` in that dataset, you simply access it using `data.csv`. Also, if you write a local file, for example `output.txt`, it will be saved in the `Resources` dataset. !!! api "API reference" - [`Project.get_job_api`][hopsworks_common.project.Project.get_job_api] - [`JobsApi`][hopsworks.core.job_api.JobsApi] - [`get_configuration`][hopsworks.core.job_api.JobsApi.get_configuration] - [`create_job`][hopsworks.core.job_api.JobsApi.create_job] - [`Job`][hopsworks_common.job.Job] - [`run`][hopsworks_common.job.Job.run] - [`DatasetApi.upload`][hopsworks_common.core.dataset_api.DatasetApi.upload] Browse the full Python API :material-arrow-right: ================================================================================ # Run PySpark Job Source: https://docs.hopsworks.ai/latest/user_guides/projects/jobs/pyspark_job/ # How To Run A PySpark Job ## Introduction This guide will describe how to configure a job to execute a pyspark script inside the cluster. All members of a project in Hopsworks can launch the following types of applications through a project's Jobs service: - Python - Apache Spark - Ray Launching a job of any type is very similar process, what mostly differs between job types is the various configuration parameters each job type comes with. Hopsworks clusters support scheduling to run jobs on a regular basis, e.g backfilling a Feature Group by running your feature engineering pipeline nightly. Scheduling can be done both through the UI and the python API, checkout [our Scheduling guide](schedule_job.md). PySpark program can either be a `.py` script or a `.ipynb` file, however be mindful of how to access/create the spark session based on the extension you provide. !!! notice "Instantiate the SparkSession" For a `.py` file, remember to instantiate the SparkSession i.e `spark=SparkSession.builder.getOrCreate()` For a `.ipynb` file, the `SparkSession` is already available as `spark` when the job is started. ## UI ### Step 1: Jobs overview The image below shows the Jobs overview page in Hopsworks and is accessed by clicking `Jobs` in the sidebar.

Jobs overview
Jobs overview

### Step 2: Create new job dialog Click `New Job` and the following dialog will appear.

Create new job dialog
Create new job dialog

### Step 3: Set the job type The `Type` radio offers `PYTHON` and `SPARK`, and `PYTHON` is selected by default. Select `SPARK` to configure a PySpark job. ### Step 4: Set the script Next step is to select the program to run. You can either select `From project`, if the file was previously uploaded to Hopsworks, or `Upload new file` which lets you select a file from your local filesystem as demonstrated below. By default, the job name is the same as the file name, but you can customize it as shown.

Configure program
Configure program

Then click `Create job` to create the job. ### Step 5 (optional): Set the PySpark script arguments In the job settings, you can specify arguments for your PySpark script. Remember to handle the arguments inside your PySpark script.

Configure PySpark script arguments
Configure PySpark script arguments

### Step 6 (optional): Advanced configuration Resource allocation for the Spark driver and executors can be configured, also the number of executors and whether dynamic execution should be enabled. - `Environment`: The python environment to use, must be based on `spark-feature-pipeline` - `Driver memory`: Number of MBs to allocate for the Spark driver - `Driver virtual cores`: Number of cores to allocate for the Spark driver - `Driver overhead factor`: Fraction of the driver memory that Spark adds to the driver pod as non-heap headroom. Prefilled with 0.40, which is also what Spark applies to a PySpark job. Clear the field to let Spark decide. - `Executor memory`: Number of MBs to allocate for each Spark executor - `Executor virtual cores`: Number of cores to allocate for each Spark executor - `Executor overhead factor`: Fraction of the executor memory that Spark adds to each executor pod as non-heap headroom. Prefilled with 0.40, which is also what Spark applies to a PySpark job. Raise it when executors are OOMKilled. - `Dynamic/Static`: Run the Spark application in static or dynamic allocation mode (see [spark docs](https://spark.apache.org/docs/latest/configuration.html#dynamic-allocation) for details).

Resource configuration for the PySpark job
Resource configuration for the PySpark job

Additional files or dependencies required for the Spark job can be configured. - `Additional archives`: List of archives to be extracted into the working directory of each executor. - `Additional jars`: List of jars to be placed in the working directory of each executor. - `Additional python dependencies`: List of python files and archives to be placed on each executor and added to PATH. - `Additional files`: List of files to be placed in the working directory of each executor.

File configuration for the PySpark job
File configuration for the PySpark job

Line-separates [properties](https://spark.apache.org/docs/3.1.1/configuration.html) to be set for the Spark application. For example, changing the configuration variables for the Kryo Serializer or setting environment variables for the driver, you can set the properties as shown below.

Additional Spark configuration
Additional Spark configuration

### Step 7: (Kueue enabled) Select a Queue Currently we do not have Kueue support for Spark. You do not need to select a queue to run the job in. ### Step 8: Execute the job Now click the `Run` button to start the execution of the job. Then open the `Executions` tab to see the list of all executions.

Start job execution
Start job execution

### Step 9: Application logs Each execution row of a Spark job carries `Spark UI`, `Metrics`, `Monitor`, `Logs` and `Live Logs`. While the execution is running, click `Spark UI` to open the Spark UI in a separate tab, or `Live Logs` to follow the logs as they are produced. Once the execution is finished, you can click on `Logs` to see the full logs for execution.

Access Spark logs
Access Spark logs

## Code ### Step 1: Upload the PySpark program This snippet assumes the program to run is in the current working directory and named `script.py`. It will upload the python script to the `Resources` dataset in your project. ```python import hopsworks project = hopsworks.login() dataset_api = project.get_dataset_api() uploaded_file_path = dataset_api.upload("script.py", "Resources") ``` ### Step 2: Create PySpark job In this snippet we get the `JobsApi` object to get the default job configuration for a `PYSPARK` job, set the pyspark script and override the environment to run in, and finally create the `Job` object. ```python jobs_api = project.get_job_api() spark_config = jobs_api.get_configuration("PYSPARK") # Set the application file spark_config["appPath"] = uploaded_file_path # Override the python job environment spark_config["environmentName"] = "spark-feature-pipeline" job = jobs_api.create_job("pyspark_job", spark_config) ``` ### Step 3: Execute the job In this snippet we execute the job synchronously, that is wait until it reaches a terminal state, and then download and print the logs. ```python execution = job.run(await_termination=True) out, err = execution.download_logs() f_out = open(out, "r") print(f_out.read()) f_err = open(err, "r") print(f_err.read()) ``` ## Configuration The following table describes the job configuration parameters for a PYSPARK job. `conf = jobs_api.get_configuration("PYSPARK")` | Field | Type | Description | Default | | --- | --- | --- | --- | | `conf['type']` | string | Type of the job configuration | `"sparkJobConfiguration"` | | `conf['appPath']` | string | Project path to spark program (e.g `Resources/foo.py`) | `null` | | `conf['defaultArgs']` | string | Arguments to pass to the program. Will be overridden if arguments are passed explicitly via `Job.run(args="...")` | `null` | | `conf['environmentName']` | string | Name of the project spark environment to use | `"spark-feature-pipeline"` | | `conf['spark.driver.cores']` | float | Number of CPU cores allocated for the driver | `1.0` | | `conf['spark.driver.memory']` | int | Memory allocated for the driver (in MB) | `2048` | | `conf['spark.executor.instances']` | int | Number of executor instances | `1` | | `conf['spark.executor.cores']` | float | Number of CPU cores per executor | `1.0` | | `conf['spark.executor.memory']` | int | Memory allocated per executor (in MB) | `4096` | | `conf['spark.driver.memoryOverheadFactor']` | float | Fraction of the driver memory added to the driver pod as non-heap overhead. `null` keeps Spark's default: 0.10 for Spark jobs, 0.40 for PySpark | `null` | | `conf['spark.executor.memoryOverheadFactor']` | float | Fraction of the executor memory added to each executor pod as non-heap overhead. `null` keeps Spark's default: 0.10 for Spark jobs, 0.40 for PySpark | `null` | | `conf['spark.dynamicAllocation.enabled']` | boolean | Enable dynamic allocation of executors | `true` | | `conf['spark.dynamicAllocation.minExecutors']` | int | Minimum number of executors with dynamic allocation | `1` | | `conf['spark.dynamicAllocation.maxExecutors']` | int | Maximum number of executors with dynamic allocation | `2` | | `conf['spark.dynamicAllocation.initialExecutors']` | int | Initial number of executors with dynamic allocation | `1` | | `conf['spark.blacklist.enabled']` | boolean | Whether executor/node blacklisting is enabled | `false` | | `conf['files']` | string | Comma-separated string of HDFS path(s) to files to be made available to the application. Example: `hdfs:///Project//Resources/file1.py,...` | `null` | | `conf['pyFiles']` | string | Comma-separated string of HDFS path(s) to python modules to be made available to the application. Example: `hdfs:///Project//Resources/file1.py,...` | `null` | | `conf['jars']` | string | Comma-separated string of HDFS path(s) to jars to be included in CLASSPATH. Example: `hdfs:///Project//Resources/app.jar,...` | `null` | | `conf['archives']` | string | Comma-separated string of HDFS path(s) to archives to be made available to the application. Example: `hdfs:///Project//Resources/archive.zip,...` | `null` | | `conf['properties']` | string | A new line separated (`\n`) list of properties to pass to the Spark application. The properties should be in the format `name=value` | `null` | ## Accessing project data ### Read directly from the filesystem (recommended) To read a dataset in your project using Spark, use the full filesystem path where the data is stored. For example, to read a CSV file named `data.csv` located in the `Resources` dataset of a project called `my_project`: ```python df = spark.read.csv( "/Projects/my_project/Resources/data.csv", header=True, inferSchema=True ) df.show() ``` ### Additional files Different file types can be attached to the spark job and made available in the `/srv/hops/artifacts` folder when the PySpark job is started. This configuration is mainly useful when you need to add additional setup, such as jars that needs to be added to the CLASSPATH. When reading data in your Spark job it is recommended to use the Spark read API as previously demonstrated, since this reads from the filesystem directly, whereas `Additional files` configuration options will download the files in its entirety and is not a scalable option. !!! api "API reference" - [`Project.get_job_api`][hopsworks_common.project.Project.get_job_api] - [`JobsApi`][hopsworks.core.job_api.JobsApi] - [`get_configuration`][hopsworks.core.job_api.JobsApi.get_configuration] - [`create_job`][hopsworks.core.job_api.JobsApi.create_job] - [`Job`][hopsworks_common.job.Job] - [`run`][hopsworks_common.job.Job.run] - [`Execution`][hopsworks_common.execution.Execution] - [`download_logs`][hopsworks_common.execution.Execution.download_logs] - [`DatasetApi.upload`][hopsworks_common.core.dataset_api.DatasetApi.upload] Browse the full Python API :material-arrow-right: ================================================================================ # Run Spark Job Source: https://docs.hopsworks.ai/latest/user_guides/projects/jobs/spark_job/ # How To Run A Spark Job ## Introduction This guide will describe how to configure a job to execute a spark program inside the cluster. All members of a project in Hopsworks can launch the following types of applications through a project's Jobs service: - Python - Apache Spark - Ray Launching a job of any type is very similar process, what mostly differs between job types is the various configuration parameters each job type comes with. Hopsworks support scheduling to run jobs on a regular basis, e.g backfilling a Feature Group by running your feature engineering pipeline nightly. Scheduling can be done both through the UI and the python API, checkout [our Scheduling guide](schedule_job.md). ## UI ### Step 1: Jobs overview The image below shows the Jobs overview page in Hopsworks and is accessed by clicking `Jobs` in the sidebar.

Jobs overview
Jobs overview

### Step 2: Create new job dialog Click `New Job` and the following dialog will appear.

Create new job dialog
Create new job dialog

### Step 3: Set the job type The `Type` radio offers `PYTHON` and `SPARK`, and `PYTHON` is selected by default. Select `SPARK` to configure a Spark job. ### Step 4: Set the jar Next step is to select the program to run. You can either select `From project`, if the file was previously uploaded to Hopsworks, or `Upload new file` which lets you select a file from your local filesystem as demonstrated below. After that set the name for the job. By default, the job name is the same as the file name, but you can customize it here.

Configure program
Configure program

### Step 5: Set the main class Next step is to set the main class for the application. Then specify [advanced configuration](#step-7-optional-advanced-configuration) or click `Create New Job` to create the job.

Set the main class
Set the main class

Then click `Create job` to create the job. ### Step 6 (optional): Set the Spark script arguments In the job settings, you can specify arguments for your Spark script. Remember to handle the arguments inside your Spark script.

Configure Spark script arguments
Configure Spark script arguments

### Step 7 (optional): Advanced configuration Resource allocation for the Spark driver and executors can be configured, also the number of executors and whether dynamic execution should be enabled. - `Environment`: The environment to use, must be based on `spark-feature-pipeline` - `Driver memory`: Number of MBs to allocate for the Spark driver - `Driver virtual cores`: Number of cores to allocate for the Spark driver - `Driver overhead factor`: Fraction of the driver memory that Spark adds to the driver pod as non-heap headroom. Prefilled with 0.40. Clear the field to let Spark decide, which for a jar job is 0.10. - `Executor memory`: Number of MBs to allocate for each Spark executor - `Executor virtual cores`: Number of cores to allocate for each Spark executor - `Executor overhead factor`: Fraction of the executor memory that Spark adds to each executor pod as non-heap headroom. Prefilled with 0.40. Clear the field to let Spark decide, which for a jar job is 0.10. Raise it when executors are OOMKilled. - `Dynamic/Static`: Run the Spark application in static or dynamic allocation mode (see [spark docs](https://spark.apache.org/docs/latest/configuration.html#dynamic-allocation) for details).

Resource configuration for the Spark job
Resource configuration for the Spark job

Additional files or dependencies required for the Spark job can be configured. - `Additional archives`: List of archives to be extracted into the working directory of each executor. - `Additional jars`: List of jars to be placed in the working directory of each executor. - `Additional python dependencies`: List of python files and archives to be placed on each executor and added to PATH. - `Additional files`: List of files to be placed in the working directory of each executor.

File configuration for the Spark job
File configuration for the Spark job

Line-separates [properties](https://spark.apache.org/docs/3.1.1/configuration.html) to be set for the Spark application. For example, changing the configuration variables for the Kryo Serializer or setting environment variables for the driver, you can set the properties as shown below.

File configuration for the Spark job
Additional Spark configuration

### Step 8: (Kueue enabled) Select a Queue Currently we do not have Kueue support for Spark. You do not need to select a queue to run the job in. ### Step 9: Execute the job Now click the `Run` button to start the execution of the job, and then click on `Executions` to see the list of all executions.

Start job execution
Start job execution

### Step 10: Application logs Each execution row of a Spark job carries `Spark UI`, `Metrics`, `Monitor`, `Logs` and `Live Logs`. While the execution is running, click `Spark UI` to open the Spark UI in a separate tab, or `Live Logs` to follow the logs as they are produced. Once the execution is finished, you can click on `Logs` to see the full logs for execution.

Access Spark logs
Access Spark logs

## Code ### Step 1: Upload the Spark jar This snippet assumes the Spark program is in the current working directory and named `sparkpi.jar`. It will upload the jar to the `Resources` dataset in your project. ```python import hopsworks project = hopsworks.login() dataset_api = project.get_dataset_api() uploaded_file_path = dataset_api.upload("sparkpi.jar", "Resources") ``` ### Step 2: Create Spark job In this snippet we get the `JobsApi` object to get the default job configuration for a `SPARK` job, set the python script to run and create the `Job` object. ```python jobs_api = project.get_job_api() spark_config = jobs_api.get_configuration("SPARK") spark_config["appPath"] = uploaded_file_path spark_config["mainClass"] = "org.apache.spark.examples.SparkPi" job = jobs_api.create_job("pyspark_job", spark_config) ``` ### Step 3: Execute the job In this snippet we execute the job synchronously, that is wait until it reaches a terminal state, and then download and print the logs. ```python execution = job.run(await_termination=True) out, err = execution.download_logs() f_out = open(out, "r") print(f_out.read()) f_err = open(err, "r") print(f_err.read()) ``` ## Configuration The following table describes the job configuration parameters for a SPARK job. `conf = jobs_api.get_configuration("SPARK")` | Field | Type | Description | Default | | --- | --- | --- | --- | | `conf['type']` | string | Type of the job configuration | `"sparkJobConfiguration"` | | `conf['appPath']` | string | Project path to spark program (e.g., `Resources/foo.jar`) | `null` | | `conf['mainClass']` | string | Name of the main class to run (e.g., `org.company.Main`) | `null` | | `conf['defaultArgs']` | string | Arguments to pass to the program. Will be overridden if arguments are passed explicitly via `Job.run(args="...")` | `null` | | `conf['environmentName']` | string | Name of the project spark environment to use | `"spark-feature-pipeline"` | | `conf['spark.driver.cores']` | float | Number of CPU cores allocated for the driver | `1.0` | | `conf['spark.driver.memory']` | int | Memory allocated for the driver (in MB) | `2048` | | `conf['spark.executor.instances']` | int | Number of executor instances | `1` | | `conf['spark.executor.cores']` | float | Number of CPU cores per executor | `1.0` | | `conf['spark.executor.memory']` | int | Memory allocated per executor (in MB) | `4096` | | `conf['spark.driver.memoryOverheadFactor']` | float | Fraction of the driver memory added to the driver pod as non-heap overhead. `null` keeps Spark's default: 0.10 for Spark jobs, 0.40 for PySpark | `null` | | `conf['spark.executor.memoryOverheadFactor']` | float | Fraction of the executor memory added to each executor pod as non-heap overhead. `null` keeps Spark's default: 0.10 for Spark jobs, 0.40 for PySpark | `null` | | `conf['spark.dynamicAllocation.enabled']` | boolean | Enable dynamic allocation of executors | `true` | | `conf['spark.dynamicAllocation.minExecutors']` | int | Minimum number of executors with dynamic allocation | `1` | | `conf['spark.dynamicAllocation.maxExecutors']` | int | Maximum number of executors with dynamic allocation | `2` | | `conf['spark.dynamicAllocation.initialExecutors']` | int | Initial number of executors with dynamic allocation | `1` | | `conf['spark.blacklist.enabled']` | boolean | Whether executor/node blacklisting is enabled | `false` | | `conf['files']` | string | Comma-separated string of HDFS path(s) to files to be made available to the application. Example: `hdfs:///Project//Resources/file1.py,...` | `null` | | `conf['pyFiles']` | string | Comma-separated string of HDFS path(s) to Python modules to be made available to the application. Example: `hdfs:///Project//Resources/file1.py,...` | `null` | | `conf['jars']` | string | Comma-separated string of HDFS path(s) to jars to be included in CLASSPATH. Example: `hdfs:///Project//Resources/app.jar,...` | `null` | | `conf['archives']` | string | Comma-separated string of HDFS path(s) to archives to be made available to the application. Example: `hdfs:///Project//Resources/archive.zip,...` | `null` | | `conf['properties']` | string | A new line separated (`\n`) list of properties to pass to the Spark application. The properties should be in the format `name=value` | `null` | ## Accessing project data ### Read directly from the filesystem (recommended) To read a dataset in your project using Spark, use the full filesystem path where the data is stored. For example, to read a CSV file named `data.csv` located in the `Resources` dataset of a project called `my_project`: ```java Dataset df = spark.read() .option("header", "true") // CSV has header .option("inferSchema", "true") // Infer data types .csv("/Projects/my_project/Resources/data.csv"); df.show(); ``` ### Additional files Different file types can be attached to the spark job and made available in the `/srv/hops/artifacts` folder when the Spark job is started. This configuration is mainly useful when you need to add additional configuration such as jars that needs to be added to the CLASSPATH. When reading data in your Spark job it is recommended to use the Spark read API as previously demonstrated, since this reads from the filesystem directly, whereas `Additional files` configuration options will download the files in its entirety and is not a scalable option. ## Environment variables User-defined environment variables can be attached to a Spark job under the *Environment variables* panel of the advanced configuration. Each entry is a `KEY=VALUE` pair that is set on the Spark driver container for every execution of the job. !!! info "Account-level variables also apply" Variables defined under [Account settings → Environment variables][account-level-environment-variables] are also injected into the Spark driver and executors. A value set on the job overrides the account-level value with the same name for this job only. ```python import os from pyspark.sql import SparkSession spark = SparkSession.builder.getOrCreate() bucket = os.environ["MY_S3_BUCKET"] df = spark.read.parquet(f"s3a://{bucket}/events/") df.printSchema() ``` Scheduled and backfill runs also receive `HOPS_LOGICAL_DATE`, `HOPS_START_TIME` and `HOPS_END_TIME` describing the data interval the run should process. See [Scheduling][logical-time-and-data-intervals] and [Batch feature pipelines][batch-feature-pipelines]. !!! warning "Reserved names" Names starting with `HOPS_` are reserved by the scheduler. Setting them in the *Environment variables* panel overrides the scheduler value for every execution. The UI shows a warning callout when you do this. !!! api "API reference" - [`Project.get_job_api`][hopsworks_common.project.Project.get_job_api] - [`JobsApi`][hopsworks.core.job_api.JobsApi] - [`get_configuration`][hopsworks.core.job_api.JobsApi.get_configuration] - [`create_job`][hopsworks.core.job_api.JobsApi.create_job] - [`Job`][hopsworks_common.job.Job] - [`run`][hopsworks_common.job.Job.run] - [`Execution`][hopsworks_common.execution.Execution] - [`download_logs`][hopsworks_common.execution.Execution.download_logs] - [`DatasetApi.upload`][hopsworks_common.core.dataset_api.DatasetApi.upload] Browse the full Python API :material-arrow-right: ================================================================================ # Run Ray Job Source: https://docs.hopsworks.ai/latest/user_guides/projects/jobs/ray_job/ # How To Run A Ray Job ## Introduction This guide will describe how to configure a job to execute a ray program inside the cluster. All members of a project in Hopsworks can launch the following types of applications through a project's Jobs service: - Python - Apache Spark - Ray Launching a job of any type is very similar process, what mostly differs between job types is the various configuration parameters each job type comes with. Hopsworks support scheduling to run jobs on a regular basis, e.g backfilling a Feature Group by running your feature engineering pipeline nightly. Scheduling can be done both through the UI and the python API, checkout [our Scheduling guide](schedule_job.md). !!!warning "Enable Ray" Ray is gated behind the `ray_enabled` [configuration variable](../../../setup_installation/admin/variables.md), which an administrator has to turn on. Until it is enabled, `RAY` is not offered as a job type. Support for Ray also needs to be explicitly enabled by adding the following option in the `values.yaml` file for the deployment: ```yaml global: ray: enabled: true ``` ## UI ### Step 1: Jobs overview The image below shows the Jobs overview page in Hopsworks and is accessed by clicking `Jobs` in the sidebar.

Jobs overview
Jobs overview

### Step 2: Create new job dialog Click `New Job` and the following dialog will appear.

Create new job dialog
Create new job dialog

### Step 3: Set the job type The `Type` radio offers `PYTHON` and `SPARK`, and `PYTHON` is selected by default. On clusters where Ray is enabled, `RAY` is offered as well. Select `RAY` to configure a Ray job. ### Step 4: Set the script Next step is to select the program to run. You can either select `From project`, if the file was previously uploaded to Hopsworks, or `Upload new file` which lets you select a file from your local filesystem as demonstrated below. After that set the name for the job. By default, the job name is the same as the file name, but you can customize it here.

Configure program
Configure program

### Step 5 (optional): Advanced configuration Resource allocation for the Driver and Workers can be configured. !!! notice "Using the resources in the Ray script" The resource configurations describe the cluster that will be provisioned when launching the Ray job. User can still provide extra configurations in the job script using `ScalingConfig`, i.e. `ScalingConfig(num_workers=4, trainer_resources={"CPU": 1}, use_gpu=True)`. - `Driver memory`: Memory in MBs to allocate for Driver - `Driver virtual cores`: Number of cores to allocate for the Driver - `Driver GPUs`: Number of GPUs to allocate for the Driver - `Worker memory`: Memory in MBs to allocate for each worker - `Worker virtual cores`: Number of cores to allocate for each worker - `Worker GPUs`: Number of GPUs to allocate for each worker - `Min workers`: Minimum number of workers to start with - `Max workers`: Maximum number of workers to scale up to

Resource configuration
for Ray Job
Resource configuration for the Ray Job

Runtime environment and Additional files required for the Ray job can also be provided. - `Runtime Environment (Optional)`: A [runtime environment](https://docs.ray.io/en/latest/ray-core/handling-dependencies.html#runtime-environments) describes the dependencies required for the Ray job including files, packages, environment variables, and more. This is useful when you need to install specific packages and set environment variables for this particular Ray job. It should be provided as a YAML file. You can select the file from the project or upload a new one. - `Additional files`: List of other files required for the Ray job. These files will be placed in `/srv/hops/ray/job`.

Runtime
environment and additional files
Runtime configuration and additional files for Ray job

### Step 6: (Kueue enabled) Select a Queue If the cluster is installed with Kueue enabled, you will need to select a queue in which the job should run. This can be done from `Advanced options`, in the `Scheduler` section of the full configuration page. ![Default queue for job](../../../assets/images/guides/project/scheduler/job_queue.png) ### Step 7: Execute the job Now click the `Run` button to start the execution of the job, and then click on `Executions` to see the list of all executions.

Start job execution
Start job execution

### Step 8: Ray Dashboard When the Ray job is running, you can access the Ray dashboard to monitor the job. The Ray dashboard is accessible from the `Executions` page. Please note that the Ray dashboard is only available when the job execution is running. In the Ray Dashboard, you can monitor the resources used by the job, the number of workers, logs, and the tasks that are running.

Access Ray Dashboard
Access Ray Dashboard

### Step 9: Application logs Once the execution is finished, you can click on `Logs` to see the full logs for execution.

Access Ray logs
Access Ray job execution logs

## Code ### Step 1: Upload the Ray script This snippet assumes the Ray program is in the current working directory and named `ray_job.py`. If the file is already in the project, you can skip this step. It will upload the jar to the `Resources` dataset in your project. ```python import hopsworks project = hopsworks.login() dataset_api = project.get_dataset_api() uploaded_file_path = dataset_api.upload("ray_job.py", "Resources") ``` ### Step 2: Create Ray job In this snippet we get the `JobsApi` object to get the default job configuration for a `RAY` job, set the python script to run and create the `Job` object. ```python jobs_api = project.get_job_api() ray_config = jobs_api.get_configuration("RAY") ray_config["appPath"] = uploaded_file_path ray_config["environmentName"] = "ray-training-pipeline" ray_config["driverCores"] = 2 ray_config["driverMemory"] = 4096 ray_config["workerCores"] = 2 ray_config["workerMemory"] = 4096 ray_config["minWorkers"] = 1 ray_config["maxWorkers"] = 4 job = jobs_api.create_job("ray_job", ray_config) ``` ### Step 3: Execute the job In this snippet we execute the job synchronously, that is wait until it reaches a terminal state, and then download and print the logs. ```python execution = job.run(await_termination=True) out, err = execution.download_logs() f_out = open(out, "r") print(f_out.read()) f_err = open(err, "r") print(f_err.read()) ``` ## Configuration The following table describes the job configuration parameters for a RAY job. `conf = jobs_api.get_configuration("RAY")` | Field | Type | Description | Default | | --- | --- | --- | --- | | `conf['type']` | string | Type of the job configuration | `"rayJobConfiguration"` | | `conf['appPath']` | string | Project relative path to script (e.g., `Resources/foo.py`) | `null` | | `conf['defaultArgs']` | string | Arguments to pass to the script. Will be overridden if arguments are passed explicitly via `Job.run(args="...")` | `null` | | `conf['environmentName']` | string | Name of the project Python environment to use | `"pandas-training-pipeline"` | | `conf['driverCores']` | float | Number of CPU cores to be allocated for the Ray head process | `1.0` | | `conf['driverMemory']` | int | Number of MBs to be allocated for the Ray head process | `4096` | | `conf['driverGpus']` | int | Number of GPUs to be allocated for the Ray head process | `0` | | `conf['workerCores']` | float | Number of CPU cores to be allocated for each Ray worker process | `1.0` | | `conf['workerMemory']` | int | Number of MBs to be allocated for each Ray worker process | `4096` | | `conf['workerGpus']` | int | Number of GPUs to be allocated for each Ray worker process | `0` | | `conf['workerMinInstances']` | int | Minimum number of Ray workers | `1` | | `conf['workerMaxInstances']` | int | Maximum number of Ray workers | `1` | | `conf['jobType']` | string | Type of job | `"RAY"` | | `conf['files']` | string | Comma-separated string of HDFS path(s) to files to be made available to the application. Example: `hdfs:///Project//Resources/file1.py,...` | `null` | ## Accessing project data If HopsFS is mounted, project datasets are available under `/hopsfs`, so you can access `data.csv` from the `Resources` dataset using `/hopsfs/Resources/data.csv` in your script. Shared datasets are accessible at `/hopsfs/shared-datasets//`. The shared datasets directory is also available through the `SHARED_DATASETS_DIR` environment variable. !!! api "API reference" - [`Project.get_job_api`][hopsworks_common.project.Project.get_job_api] - [`JobsApi`][hopsworks.core.job_api.JobsApi] - [`get_configuration`][hopsworks.core.job_api.JobsApi.get_configuration] - [`create_job`][hopsworks.core.job_api.JobsApi.create_job] - [`Job`][hopsworks_common.job.Job] - [`run`][hopsworks_common.job.Job.run] - [`Execution`][hopsworks_common.execution.Execution] - [`download_logs`][hopsworks_common.execution.Execution.download_logs] - [`DatasetApi.upload`][hopsworks_common.core.dataset_api.DatasetApi.upload] Browse the full Python API :material-arrow-right: ================================================================================ # Scheduling Source: https://docs.hopsworks.ai/latest/user_guides/projects/jobs/schedule_job/ # How To Schedule a Job ## Introduction Hopsworks clusters can run jobs on a schedule, allowing you to automate the execution. Whether you need to backfill your feature groups on a nightly basis or run a model training pipeline every week, the Hopsworks scheduler will help you automate these tasks. Each job can be configured to have a single schedule. For more advanced use cases, Hopsworks integrates with any DAG manager and directly with the open-source [Apache Airflow](https://airflow.apache.org/use-cases/); see our [Airflow Guide][orchestrate-jobs-using-apache-airflow]. Schedules can be defined using the drop-down menus in the UI or a Quartz [cron](https://en.wikipedia.org/wiki/Cron) expression. !!! note "Schedule frequency" The Hopsworks scheduler runs every minute. As such, the scheduling frequency should be of at least 1 minute. !!! note "Parallel executions" By default, at most one execution of a job runs at a time. If a fire time is reached while a previous execution is still running, the scheduler waits for the previous run to finish before starting the new one. You can raise this cap with `max_active_runs` (see [Scheduling fields](#scheduling-fields)). ## Logical time and data intervals When the scheduler fires a job, it attaches a **data window** to the execution, expressed through three environment variables: - `HOPS_START_TIME`: start of the data window. Two modes: - **Last execution time** *(default, `start_time_offset_seconds = null`)*: the previous cron fire. Adapts to the schedule's cadence, so the window is always "everything since the previous run". - **Before last execution (hh:mm)**: `previous_fire − (hh*3600 + mm*60)` seconds. Stored as a positive integer; the backend subtracts. - `HOPS_END_TIME`: end of the data window. **Always `HOPS_START_TIME + cron interval`** so consecutive scheduled runs tile the timeline with no gaps between them. `end_time_offset_seconds` is kept on the DTO for backward compatibility but is ignored. - `HOPS_LOGICAL_DATE`: the scheduler's stable identifier for this interval (Airflow-style *start of interval* = previous cron fire). Used for dedup and retries. With the defaults on an hourly schedule firing at 10:00, the window is `[09:00, 10:00)`, the last hour of data. The same defaults on a daily schedule firing at 00:00 give `[yesterday 00:00, today 00:00)`, the last day. Change the modes (see below) to shape a different window. --8<-- "user_guides/projects/jobs/schedule_job/data-windows.html" These three values are injected into the job container as **environment variables** on every scheduled execution. In your program, read them like any other env var: === "Python" ```python import os from datetime import datetime start = datetime.fromisoformat(os.environ["HOPS_START_TIME"]) end = datetime.fromisoformat(os.environ["HOPS_END_TIME"]) print(f"Processing rows in [{start}, {end})") ``` === "PySpark" ```python import os from pyspark.sql import SparkSession from pyspark.sql import functions as F spark = SparkSession.builder.getOrCreate() start = os.environ["HOPS_START_TIME"] end = os.environ["HOPS_END_TIME"] df = spark.read.table("events").where( (F.col("event_ts") >= F.lit(start)) & (F.col("event_ts") < F.lit(end)) ) ``` ### Why intervals and not just "now"? A scheduled run represents a *time slice of data to process*, not the wall-clock moment the container happened to start. Decoupling the two matters: - If the scheduler is down for six hours and an hourly job resumes, each catch-up run still sees the interval it was supposed to cover, not six copies of "now". - Backfills work naturally: replaying an interval produces the same output as the original run. - Downstream consumers can rely on the interval end being monotonic even when upstream is delayed. ### Legacy `-start_time` argument For backwards compatibility the scheduler also appends `-start_time ` to the job arguments (or `-p start_time ` for Papermill notebooks). New code should prefer the `HOPS_*` environment variables: they're present on every job type and don't require parsing CLI args. ## UI ### Scheduling a job You can define a schedule for a job during the creation of the job itself or after the job has been created, from the job overview UI.

Schedule a Job
Schedule a Job

The *add schedule* prompt requires you to select a frequency either through the drop-down menus or by using a cron expression. You can also provide a start time to specify when the schedule should take effect. The start time can be in the past. You can optionally provide an end time, in which case the scheduler stops firing after that point. In the job overview you can see the current scheduling configuration, whether it is enabled, and when the next execution is planned for. All times are in UTC.

Job scheduling overview
Job scheduling overview

### Scheduling fields The Schedule form exposes these fields for controlling the data window, concurrency and catch-up behaviour:

Schedule form with a start offset and catch-up enabled
Start offset and catch-up options in the Schedule form

| Field | Default | Description | | --- | --- | --- | | `max_active_runs` | `1` | Upper bound on concurrent executions for this job. | | `start_time_offset_seconds` | `null` *(last execution time)* | `null` → `HOPS_START_TIME = previous cron fire`. Positive integer → `HOPS_START_TIME = previous fire − seconds` (shifts the window earlier). Must be ≥ 0. | | `end_time_offset_seconds` | *legacy; ignored* | Kept on the DTO for backward compatibility. The backend always sets `HOPS_END_TIME = HOPS_START_TIME + cron interval`, so consecutive runs tile the timeline with no gaps. | | `catchup` | `false` | If `true`, on recovery after a scheduler outage the runs for every missed interval are created. If `false`, only the most recent missed interval is created. | | `skip_to_date` | *unset* | When `catchup=true`, missed intervals strictly before this date are skipped during reconciliation. | | `max_catchup_runs` | *unset* | When `catchup=true`, caps how many missed intervals are replayed, keeping the most recent. | #### Example: "process the previous 25 hours, not the previous hour" Defaults give you the natural cron interval (`[previous fire, current fire)`). To process a 25-hour trailing window every hour (e.g. to include a little overlap for late-arriving data), shift `start` 24 hours earlier. The end stays at the current fire because it's `start + cron interval = start + 1 h`: ```text start_time_offset_seconds = 25 * 3600 # HOPS_START_TIME = previous fire − 25 h # end_time_offset_seconds is ignored; HOPS_END_TIME = HOPS_START_TIME + 1 h ``` At 11:00 UTC on an hourly schedule the container sees `HOPS_START_TIME = 09:00 (previous day)` and `HOPS_END_TIME = 10:00 (previous day)`. #### Example: `catchup=false` after a 6-hour outage With an hourly schedule and `catchup=false`, if the scheduler is unreachable from 00:00 to 06:00, on recovery the scheduler creates one execution, the 06:00 interval, not six. This matches Airflow's `catchup_by_default=False` semantics. #### Example: bounded replay With `catchup=true`, `max_catchup_runs=24` and a 1-week outage on an hourly schedule, only the most recent 24 missed intervals are replayed; older ones are dropped. ### Disable / enable a schedule You can pause a schedule to prevent further executions. When re-enabled, the next fire uses the current time as the starting point (pause is not the same as an outage, paused time is not considered "missed"). Use `skip_to_date` (advanced) if you want a paused-and-re-enabled schedule with `catchup=true` to skip over the pause window. ### Delete a schedule You can remove the schedule for a job using the UI and by clicking on the trash icon on the schedule section of the job overview. If you re-schedule a job after having deleted the previous schedule, even with the same options, previous scheduled executions are not considered. ## Python API ```python from datetime import UTC, datetime import hopsworks project = hopsworks.login() job = project.get_job_api().get_job("my_feature_pipeline") # Defaults (None, None): HOPS_START_TIME = previous cron fire (last execution time), # HOPS_END_TIME = cron fire time. Adapts to any cron: hourly gives the last hour, # daily gives the last day, etc. job.schedule( cron_expression="0 0 * ? * * *", start_time=datetime(2026, 1, 1, tzinfo=UTC), ) # Shift the window 1 hour earlier so each hourly run sees the previous-to-previous # hour. HOPS_END_TIME is always HOPS_START_TIME + cron interval (1 h here), so the # window remains 1 hour wide, only the anchor moves. job.schedule( cron_expression="0 0 * ? * * *", start_time_offset_seconds=3600, # previous fire − 1 h catchup=True, max_active_runs=2, max_catchup_runs=24, ) # Inspect schedule = job.job_schedule print(schedule.cron_expression, schedule.catchup, schedule.max_active_runs) print(schedule.next_execution_date_time) # Remove job.unschedule() ``` See also [Batch feature pipelines][batch-feature-pipelines] for one-shot backfill runs and the `Job.run(start_time=..., end_time=...)` API. ## Differences from Apache Airflow Hopsworks follows the same [data-interval model as Airflow](https://airflow.apache.org/docs/apache-airflow/stable/authoring-and-scheduling/timetable.html): `logical_date` is the start of the data interval, `data_interval_end` is the cron fire time, runs are created at or after the interval end, and `catchup` controls whether missed intervals are replayed. However, a few behaviours differ. If you are coming from Airflow, these are the ones worth knowing. ### First-run interval starts at `start_date`, not the next cron boundary If a schedule has `start_date = 10:05` and an hourly cron firing on `:00`, Airflow aligns `start_date` forward to `11:00` and the first run has `data_interval = [11:00, 12:00)`. Hopsworks clamps the previous fire to `start_date`, so the first run fires at `11:00` with `data_interval = [10:05, 11:00)`, a short first interval starting at `start_date` itself. If that offset would break the first run's computation (for example, a job that expects full hourly windows), set `start_date` to a cron-aligned time. ### Manual runs don't automatically get a data interval Airflow always fills a data interval for manually triggered DagRuns via `infer_manual_data_interval`. Hopsworks sets `HOPS_LOGICAL_DATE` / `HOPS_START_TIME` / `HOPS_END_TIME` on a manual run **only if you explicitly pass them**, via the UI's *Run with time window* dialog, the Python `job.run(start_time=..., end_time=...)` kwargs, or the JSON execution endpoint. A plain "Run" with no time window leaves the env vars unset. This keeps the contract honest: a manual one-off either processes a specific window (opt in) or processes "whatever your code defines, independent of any window". It does mean job code shouldn't assume `HOPS_START_TIME` is always present, so check for the var and fall back as appropriate. ### Per-tick draining vs. one-per-loop Airflow's scheduler creates at most one DagRun per DAG per loop iteration, to keep many DAGs progressing in parallel. The Hopsworks scheduler will drain up to 20 past intervals for one schedule in a single timer tick (each tick is ~1 minute), then move on. In practice this means a schedule with a large backlog catches up in batches of 20 per minute rather than one per minute, while not blocking other schedules for more than one tick. You typically don't observe this unless `catchup=true` after a long outage: expect the backlog to drain at ~20 intervals/minute per schedule, not one. ### No DST fold-hour handling Airflow treats cron schedules as timezone-aware and has explicit logic for the DST "fold hour" when clocks go back. Hopsworks evaluates schedules in UTC. If you need non-UTC semantics, express your cron in UTC. There is no equivalent of Airflow's timezone-aware timetable for DST. For almost all users this is a non-issue. ### Reconciliation on schedule create / update, not just outage In Airflow, `catchup=true` kicks in only when the scheduler detects past intervals at runtime. Hopsworks also replays missed intervals at *creation* time: if you create a new schedule with `catchup=true` and a `start_date` in the past, past intervals run immediately rather than being implicitly skipped by the first-fire advancement. The same happens if you flip `catchup` from `false` to `true`, or re-enable a paused `catchup=true` schedule. This matches what most users expect ("turn on catchup ⇒ see the backfill"), but is a deliberate divergence. Airflow's equivalent behaviour is "the next DagRun after now is the earliest missed one, and subsequent loops fill in the rest over time". ### Leader election via Payara, not advisory locks Airflow coordinates multiple schedulers with PostgreSQL advisory locks (`pg_try_advisory_xact_lock`). Hopsworks uses Payara's cluster primary election: the scheduler timer only fires on the primary node. Reconciliation happens on the first primary-owned tick after startup. Correctness is backed at the application layer: `executeWithCron` seeds `startingActive = countActiveByJob(...)` once per tick and tracks an additive `firedThisTick` counter, which is monotonic and independent of RonDB COUNT read-after-write lag. This replaces an earlier `(job_id, logical_date)` unique index that was too strict: it blocked legitimate retry / manual-rerun / re-backfill flows that share a logical_date. ================================================================================ # Batch Feature Pipelines Source: https://docs.hopsworks.ai/latest/user_guides/projects/jobs/batch_feature_pipeline/ # Batch Feature Pipelines A **batch feature pipeline** is a job that processes a *time slice of data*, for example "the rows for the previous hour" or "every day's transactions", and writes the resulting features to a feature group. Hopsworks jobs carry this time slice as environment variables so your code does not need to guess the window from wall-clock time. Two operating modes cover the common cases: - **Incremental**: the pipeline runs on a recurring schedule. Each run sees a fresh `[HOPS_START_TIME, HOPS_END_TIME)` computed as offsets from the cron fire time. Use this for production pipelines. - **Backfill**: a one-shot run over an explicit absolute time window. Use this to populate history or re-process a specific interval after a bug fix. Both modes emit the same `HOPS_*` environment variables, so the same pipeline code handles both. --8<-- "user_guides/projects/jobs/batch_feature_pipeline/data-windows.html" ## Environment variables On every scheduled or backfill execution, Hopsworks injects: | Variable | Meaning | | --- | --- | | `HOPS_START_TIME` | `start_time_offset_seconds = null` *(default)* → previous cron fire (last execution time). Positive integer → `previous fire − seconds` (shifts the start earlier). Must be ≥ 0. | | `HOPS_END_TIME` | Always `HOPS_START_TIME + cron interval` so consecutive runs tile the timeline with no gaps. `end_time_offset_seconds` is kept on the DTO for backward compatibility but ignored. | | `HOPS_LOGICAL_DATE` | Stable identifier for this interval (Airflow-style start of interval = previous cron fire). Used for dedup and retries. | For a manual (non-scheduled) run, these variables are only set if you explicitly pass a time window via the UI or API (see [Backfill](#backfill-one-shot-absolute-window) below). !!! tip "User-defined env vars override the scheduler" If you set `HOPS_START_TIME` or `HOPS_END_TIME` in the job's envVars (or at launch time), your value wins over the scheduler-computed one for that execution. This is how backfill runs override the cron-derived interval. !!! info "Account-level variables also apply" Variables defined under [Account settings → Environment variables][account-level-environment-variables] are also injected into every execution. A value set on the job overrides the account-level value with the same name for this job only. ## Incremental (recurring) Create or edit a job and configure its schedule under **Advanced scheduling**. Typical settings for a batch feature pipeline: - `cron_expression`: how often to run (e.g. `0 0 * ? * * *` for hourly). - `start_time_offset_seconds`: default `null` gives the natural cron interval (`[previous fire, current fire)`). Set a positive integer to anchor the window earlier than the previous fire (e.g. `3600` makes each run see the window starting an hour before its predecessor); the window remains exactly one cron interval wide because `HOPS_END_TIME = HOPS_START_TIME + cron interval`. - `catchup`: on by default *off*. Enable it if missed runs during an outage should be replayed one-per-missed-interval. - `max_active_runs`: raise above 1 if runs can safely execute in parallel. See [How to schedule a job][scheduling-fields] for the full field reference. ### Reading the interval in your code ```python import os from datetime import datetime import hopsworks project = hopsworks.login() fs = project.get_feature_store() fg = fs.get_feature_group("page_views", version=1) source = fs.get_storage_connector("my_source") start = datetime.fromisoformat(os.environ["HOPS_START_TIME"]) end = datetime.fromisoformat(os.environ["HOPS_END_TIME"]) df = source.read( query=( "SELECT * FROM events " f"WHERE ts >= '{start.isoformat()}' AND ts < '{end.isoformat()}'" ), ) fg.insert(df) ``` This code works identically whether the job runs on its schedule or as a one-shot backfill. ## Backfill (one-shot, absolute window) Use backfill when you need to re-process historical data or seed a feature group before turning on the incremental schedule. No schedule change is required: you supply the window at launch time. ### From the UI On the job's **Run** dialog, tick **Run with time window (one-shot backfill)** and pick the `HOPS_START_TIME` / `HOPS_END_TIME` datetimes (UTC). The execution is submitted with those env vars set; any schedule-derived values are overridden for this run.

Run job dialog with a time window
The Run dialog with a one-shot backfill window

### From the Python SDK ```python from datetime import UTC, datetime import hopsworks project = hopsworks.login() job = project.get_job_api().get_job("my_feature_pipeline") # Backfill March 2026 job.run( start_time=datetime(2026, 3, 1, tzinfo=UTC), end_time=datetime(2026, 4, 1, tzinfo=UTC), ) ``` `Job.run()` accepts the following optional logical-time kwargs: | kwarg | Maps to | | --- | --- | | `start_time` | `HOPS_START_TIME` + `logical_date` (data interval start) | | `end_time` | `HOPS_END_TIME` + `data_interval_end` | | `logical_date` | Override `HOPS_LOGICAL_DATE` independently (rare). | | `env_vars` | Arbitrary per-run env vars; highest precedence. | ### Chained monthly backfill ```python from datetime import UTC, datetime, timedelta import hopsworks project = hopsworks.login() job = project.get_job_api().get_job("my_feature_pipeline") start = datetime(2024, 1, 1, tzinfo=UTC) end = datetime(2026, 1, 1, tzinfo=UTC) cursor = start while cursor < end: nxt = (cursor.replace(day=1) + timedelta(days=32)).replace(day=1) job.run(start_time=cursor, end_time=nxt, await_termination=True) cursor = nxt ``` ### Batched backfill at job creation When creating a new job in the UI, the **Backfill** card lets you split one window into **N equal sub-windows** and fire one execution per sub-window. Tick *Run job on creation* to have the sub-windows fired as soon as the job is saved:

Backfill card on the New Job form
The Backfill card on the New Job form, four batch jobs over one window

- **Number of Batch Jobs**: how many sub-windows. `[start, end)` is tiled with no gaps or overlaps; the last sub-window absorbs any integer-division remainder so the union is exactly the original window. `1` means one execution covering the whole window (the default). - **Max parallel executions**: must be `≥ Number of Batch Jobs` today. Runtime concurrency enforcement (pause the next batch until a running one completes) is on the roadmap; until then the backend rejects smaller values with a `400` rather than silently over-firing. Setting it equal to the batch count fires everything in parallel. Each sub-window run receives `HOPS_START_TIME` / `HOPS_END_TIME` for its slice, so the same pipeline code used by the incremental schedule works unchanged. ## Precedence summary Several sources can set the same env var. The rule is "most specific wins": 1. **Per-execution env vars** supplied to `job.run(env_vars=...)` or the UI time-window dialog. 2. **Job-config env vars** (the Environment variables panel on the Advanced job form). 3. **Scheduler-computed** `HOPS_*` from the data interval + offsets. 4. **Hopsworks defaults** (e.g. `HADOOP_HOME`). So setting `HOPS_END_TIME` in the Environment variables panel pins it for every execution of the job; passing `env_vars={"HOPS_END_TIME": "..."}` to a single `job.run()` pins it only for that run. ## See also - [Schedule a job][how-to-schedule-a-job]: full reference for cron + advanced scheduling fields. - [Python job][how-to-run-a-python-job] / [Spark job][how-to-run-a-spark-job]: adding generic env vars to a job configuration. ================================================================================ # Base Source: https://docs.hopsworks.ai/latest/user_guides/projects/scheduling/kube_scheduler/ # Scheduler ## Introduction Hopsworks allows users to configure some Kubernetes scheduler abstractions, such as [Affinity](https://kubernetes.io/docs/tasks/configure-pod-container/assign-pods-nodes-using-node-affinity/) and [Priority Classes](https://kubernetes.io/docs/concepts/scheduling-eviction/pod-priority-preemption/#priorityclass). Hopsworks also supports additional scheduling abstractions backed by Kueue. This includes [Queues](https://kueue.sigs.k8s.io/docs/concepts/cluster_queue/), [Cohorts](https://kueue.sigs.k8s.io/docs/concepts/cohort/) and [Topologies](https://kueue.sigs.k8s.io/docs/concepts/topology_aware_scheduling/). All these scheduling abstractions are supported in jobs, jupyter notebooks and model deployments. Kueue abstractions however, are currently not supported for Spark jobs. Hopsworks Admins can control which labels and priority classes can be used in the cluster (see [Admin configuration](#admin-configuration) section) and by which project (see [Project Configuration](#project-configuration) section) Within a project, data owners can set defaults for jobs and Jupyter notebooks running within that project (see: [Project defaults](#project-defaults) section). ### Node Labels, Node Affinity and Node Anti-Affinity Labels in Kubernetes are key-value pairs used to organize and select resources. Hopsworks relies on labels applied to nodes for pod-node affinity to determine where the pod can (or cannot) run. Some uses cases where labels and affinity can be used include: - Hardware constraints (GPU, SSD) - Environment separation (prod/dev) - Co-locating related pods - Spreading pods for high availability Hopsworks uses the node affinity `IN` operator for the Hopsworks Node Affinity and the `NOT IN` operator for the Hopsworks Node Anti Affinity. For more information on Kubernetes Affinity, you can check the Kubernetes [Affinity documentation](https://kubernetes.io/docs/tasks/configure-pod-container/assign-pods-nodes-using-node-affinity/) page. ### Priority Classes Priority classes in Kubernetes determine the scheduling and eviction priority of pods. Pods with higher priority: - Get scheduled first - Can preempt (evict) lower priority pods - Less likely to be evicted under resource pressure Common uses: - Protecting critical workloads - Ensuring core services stay running - Managing resource competition - Guaranteeing QoS for important applications For more information on Priority Classes, you can check the Kubernetes [Priority Classes documentation](https://kubernetes.io/docs/concepts/scheduling-eviction/pod-priority-preemption/#priorityclass) page. ## Kueue Hopsworks adds the integration with Kueue to offer more advanced scheduling abstractions such as queues, cohorts and topologies. For a more detailed view on how Hopsworks uses the Kueue abstractions you can check the [Kueue details](./kueue_details.md) section. ### Queues, Cohorts Jobs, notebooks and model deployments are submitted to these queues. Hopsworks administrator can define quotas on how many resources a queue can use. Queues can be grouped together in cohorts in order to add the ability to borrow resources from each other when the other queue does not use its resources. When creating a new job, the user can select a queue for the job in the `Kubernetes scheduling constraints` section of the job configuration page. ![Queue selection for a job](../../../assets/images/guides/project/scheduler/job_queue.png) ### Topologies The integration of Hopsworks with Kueue, also provides access to the topology abstraction. Topologies can be defined, so that the user can decide for the pods of jobs or model deployments to run somehow grouped together. The user could decide for example, that all pods of a job should run on the same host, because the pods need to transfer a lot of data between each other, and we want to avoid network traffic to lower the latency. The user can select the topology unit for jobs, notebooks and model deployments in the `Kubernetes scheduling constraints` section of the configuration page. ![Topology unit selection for a job](../../../assets/images/guides/project/scheduler/job_topology_unit.png) ## Admin configuration ### Affinity and priority classes Hopsworks admins can control the affinity labels and priority classes available on the Hopsworks cluster from the `Cluster Settings -> Scheduler` page: ![Cluster Configuration - Node Labels and Priority Classes](../../../assets/images/guides/project/scheduler/admin_cluster_scheduler.png) Hopsworks Cluster can run within a shared Kubernetes Cluster. The first configuration level is to limit the subset of labels and priority classes that can be used within the Hopsworks Cluster. This can be done from the `Available in Hopsworks` sub-section. !!! note "Permissions" In order to be able to list all the Kubernetes Node Labels, Hopsworks requires the following cluster role: ```yaml - apiGroups: [""] resources: ["nodes"] verbs: ["get", "list"] ``` In order to be able to list all the Kubernetes Cluster Priority Classes, Hopsworsk requires this cluster role: ```yaml - apiGroups: ["scheduling.k8s.io"] resources: ["priorityclasses"] verbs: ["get", "list"] ``` If the roles above are configured properly (default behaviour), admins can only select values from the drop down menu. If the roles are missing, admins would be required to enter them as free text and should be careful about typos. Any typos here will be propagated in the other configuration and use levels leading to errors or missbehaviour when running computation. ### Queues Every new project gets automatic access to the default Hopsworks queue. An administrator can define the default queue for projects user jobs and system jobs. ![Default queue for user and system jobs](../../../assets/images/guides/project/scheduler/default_queue.png) ## Project Configuration Hopsworks admins can configure the labels and priority classes that can be used by default within a project. This will be a subset of the ones configured for Hopsworks. In the figure above, in the sub-section `Available in Project` Hopsworks admins can configure the labels and priority classes available by default in any Hopsworks Project. Hopsworks admins can also override the default project configuration on a per-project basis. That is, Hopsworks admins can make certain labels and priority classes available only to certain projects. This can be achieved from the `Cluster Settings -> Project -> -> edit configuration` configuration page: ![Custom Project Configuration - Node Labels and Priority Classes](../../../assets/images/guides/project/scheduler/admin_project_scheduler.png) ## Project defaults Within a project, different jobs, Jupyter notebooks and model deployments can run with different labels and/or priority classes. `Data Owners` in a project can specify the default values from the project settings: The default Label will be used for the default Node Affinity for jobs, notebooks, and model deployments. ![ Project Default - Labels and Priority Classes](../../../assets/images/guides/project/scheduler/project_default.png) ## Configuration of Jobs, Notebooks, and Deployments In the advanced configuration sections for job, notebook, and model deployments, users can set affinity, anti affinity and priority class. The Affinity and Anti Affinity can be selected from the list of allowed labels. `Affinity` configures on which nodes this pod can run. If a node has any of the labels present in the Affinity option, the pod can be scheduler to run to run there. `Anti Affinity` configures on which nodes this pod will not run on. If a node has any of the labels present in the Anti Affinity option, the pod will not be scheduler to run there. `Priority Class` specifies with which priority a pod will run. ![ Job Configuration - Affinity and Priority Classes](../../../assets/images/guides/project/scheduler/job_configuration.png) ================================================================================ # Kueue Source: https://docs.hopsworks.ai/latest/user_guides/projects/scheduling/kueue_details/ # Kueue ## Introduction Hopsworks provides the integration with Kueue to provide the additional scheduling abstractions. Hopsworks currently acts only as a "reader" to the Kueue abstractions and currently does not manage the lifecycle of Kueue abstraction with the exception of the default localqueue for each namespace. All the other abstractions are expected to be managed by the administrators of Hopsworks, directly on the Kubernetes cluster. However Hopsworks and Kueue integration currently only supports frameworks python and ray for jobs, notebooks and model deployments. The same queues are also used for Hopsworks internal jobs (zipping, git operations, python library installation). Spark is currently not supported, and thus will not be managed by Kueue for scheduling, and instead it will bypass the queues setup (important to note when thinking about queue quotas) and instead are managed directly by the Kubernetes Scheduler. ### Resource flavors When trying to define queues in Kueue, the first abstraction that needs to be defined is a [Resource Flavor](https://kueue.sigs.k8s.io/docs/concepts/resource_flavor/). The resource flavor defines the resources that a queue will later manage. Hopsworks helm chart installs and uses a default ResourceFlavor ```yaml apiVersion: kueue.x-k8s.io/v1beta1 kind: ResourceFlavor metadata: name: default-flavor spec: nodeLabels: cloud.provider.com/region: europe topologyName: default ``` Node labels filter the available nodes to this resource flavor and is required for [topologies](#topologies) ### Cluster Queues [Cluster Queues](https://kueue.sigs.k8s.io/docs/concepts/cluster_queue/) are the actual queues for submitting jobs and model deployments to. The default hopsworks queue looks like: ```yaml apiVersion: kueue.x-k8s.io/v1beta1 kind: ClusterQueue metadata: name: other spec: cohort: cluster namespaceSelector: {} preemption: borrowWithinCohort: policy: Never reclaimWithinCohort: Never withinClusterQueue: Never queueingStrategy: BestEffortFIFO resourceGroups: - coveredResources: - cpu - memory - pods - nvidia.com/gpu flavors: - name: default-flavor resources: - name: cpu nominalQuota: "0" - name: memory nominalQuota: "0" - name: pods nominalQuota: "0" - name: nvidia.com/gpu nominalQuota: "0" ``` The [preemption](https://kueue.sigs.k8s.io/docs/concepts/cluster_queue/#preemption) and [nominal quotas](https://kueue.sigs.k8s.io/docs/concepts/cluster_queue/#flavors-and-resources) are set to the minimal as this queue is designed to have lowest priority in getting resources allocated. If a cluster is underutilized and there are resources available, it can still borrow up to the maximum resources present in the parent cohort, but by design this queue has no dedicated resources. The presumption is that other, more important queues, defined by the cluster administrator will have higher preference in getting resources. ### Local Queues [Local Queues](https://kueue.sigs.k8s.io/docs/concepts/local_queue/) are the mechanism to provide access to a queue (cluster queue) to a specific project in Hopsworks (Kubernetes namespace). Every new project gets automatic access to the default Hopsworks queue. An administrator can define the default queue for projects user jobs and system jobs. ![Default queue for user and system jobs](../../../assets/images/guides/project/scheduler/default_queue.png) ### Cohorts [Cohorts](https://kueue.sigs.k8s.io/docs/concepts/cohort/) are groupings of cluster queues that have some meaning together and can share resources. Hopsworks defines a default `cluster` cohort ```yaml apiVersion: kueue.x-k8s.io/v1alpha1 kind: Cohort metadata: name: cluster spec: resourceGroups: - coveredResources: - cpu - memory - pods - nvidia.com/gpu flavors: - name: default-flavor resources: - name: cpu nominalQuota: 100 - name: memory nominalQuota: 200Gi - name: pods nominalQuota: 100 - name: nvidia.com/gpu nominalQuota: 50 ``` Cohorts can contain other cohorts and thus you can create a hierarchy of cohorts. Cohorts can set [fair sharing weight](https://kueue.sigs.k8s.io/docs/concepts/admission_fair_sharing/) where using ```yaml fairSharing: weight ``` in the definition of a cohort, the user can control a priority towards borrowing resources from other cohorts. ### Topologies [Topologies](https://kueue.sigs.k8s.io/docs/concepts/topology_aware_scheduling/) defines a way of grouping together pods belonging to the same job/deployment so that they are colocated within the same topology unit. Hopsworks defines a default topology: ```yaml apiVersion: kueue.x-k8s.io/v1alpha1 kind: Topology metadata: name: default spec: levels: - nodeLabel: cloud.provider.com/region - nodeLabel: cloud.provider.com/zone - nodeLabel: kubernetes.io/hostname ``` The topology is defined in the Resource Flavor used by a Cluster Queue. When creating a new job, the user can select a topology unit for the job to run in and thus decide if all pods of a job should run on the same hostname, in the same zone or in the same region. The user can select the topology for jobs, notebooks and deployments from `Advanced options`, in the `Scheduler` section of the full configuration page. ![Default queue for user and system jobs](../../../assets/images/guides/project/scheduler/job_topology_unit.png) ================================================================================ # Overview Source: https://docs.hopsworks.ai/latest/user_guides/projects/airflow/airflow/ # Orchestrate Jobs using Apache Airflow ## Introduction Hopsworks jobs can be orchestrated using [Apache Airflow](https://airflow.apache.org/). You can define an Airflow DAG (Directed Acyclic Graph) containing the dependencies between Hopsworks jobs. You can then schedule the DAG to be executed at a specific schedule using a [cron](https://en.wikipedia.org/wiki/Cron) expression. Airflow DAGs are defined as Python files. Within the Python file, different operators can be used to trigger different actions. Hopsworks provides operators to execute jobs on Hopsworks and sensors to wait for a specific job to finish or for a HopsFS path to appear. Hopsworks ships **Airflow 3.0.6**. The DAG-authoring API differs significantly from Airflow 1.x; see [the Airflow 3 upgrade page](airflow3_upgrade.md) for a porting checklist. ### Use Apache Airflow in Hopsworks Hopsworks deployments include a deployment of Apache Airflow. You can access it from the Hopsworks UI by clicking on the _Airflow_ button on the left menu. Authorization is per Hopsworks project: admins on Hopsworks (`HOPS_ADMIN`) have access to all DAGs; regular users see only DAGs of the projects they are a member of (DAGs, runs, logs, triggers, edits: all surfaces). The Audit Log is row-filtered for non-admins to events for DAGs in their projects. See the [security model](security_model.md) for the full surface-by-surface contract. The Hopsworks UI's Airflow page shows each DAG's most recent runs as colored squares in a **Last runs** column (green = success, red = failed, blue = running, yellow = queued / scheduled, gray = other). Clicking anywhere on a DAG row opens the DAG in the Airflow UI. The pencil at the row's end opens the generated Python file in an in-app editor. The trash icon deletes the DAG: a click-confirm dialog appears, and on confirm the Python file is removed from the project's `Airflow/` HopsFS dataset, the per-DAG `hopsworks_api_key_` Variable is deleted from Airflow, the row in `dag_project_index` is dropped, and `airflow.api.common.delete_dag.delete_dag` is called so the `dag`, `dag_run`, `task_instance`, `xcom`, `log`, and related rows go with it. After delete the page reloads to reflect the new state. #### Hopsworks DAG Builder
Airflow DAG Builder
Airflow DAG Builder
You can create a new Airflow DAG to orchestrate jobs using the Hopsworks DAG builder tool. Click on _New Workflow_ to create a new Airflow DAG. You should provide a name for the DAG as well as a schedule interval. You can define the schedule using the dropdown menus or by providing a cron expression. The schedule `@continuous` is rejected by both the UI form and the backend. A continuous DAG re-runs as soon as the previous run finishes, so a DAG that errors at parse time (for example, missing the per-DAG API key Variable) loops at wall-clock speed and OOM-kills the shared scheduler pod, taking every other project's DAGs down with it. Use a cron expression for periodic runs, or `@once` for one-shot DAGs. You can add to the DAG Hopsworks operators and sensors: - **Operator**: The operator is used to trigger a job execution. When configuring the operator you select the job you want to execute and you can optionally provide execution arguments. You can decide whether or not the operator should wait for the execution to be completed. If you select the _wait_ option, the operator will block and Airflow will not execute any parallel task. If you select the _wait_ option the Airflow task fails if the job fails. If you want to execute tasks in parallel, you should not select the _wait_ option but instead use the sensor. When configuring the operator, you can can also provide which other Airflow tasks it depends on. If you add a dependency, the task will be executed only after the upstream tasks have been executed successfully. - **Sensor**: The sensor can be used to wait for executions to be completed. Similarly to the _wait_ option of the operator, the sensor blocks until the job execution is completed. The sensor can be used to launch several jobs in parallel and wait for their execution to be completed. Please note that the sensor is defined at the job level rather than the execution level. The sensor will wait for the most recent execution to be completed and it will fail the Airflow task if the execution was not successful. You can then create the DAG and Hopsworks will generate the Python file. #### Write your own DAG If you prefer to code the DAGs or you want to edit a DAG built with the builder tool, you can do so. The Airflow DAGs are stored in the _Airflow_ dataset which you can access using the file browser in the project settings. The Hopsworks operators and sensors are shipped in the `apache-airflow-providers-hopsworks` provider package that is preinstalled on Hopsworks-managed Airflow. Import them from the standard provider namespace: ```python from hopsworks.airflow.operators import HopsworksLaunchOperator # noqa: F401 from hopsworks.airflow.sensors import ( # noqa: F401 HopsworksHdfsSensor, HopsworksJobSuccessSensor, ) ``` Launch a Hopsworks job: ```python from hopsworks.airflow.operators import HopsworksLaunchOperator HopsworksLaunchOperator( task_id="profiles_fg_0", project_id=42, job_name="profiles_fg", args="", wait_for_completion=True, ) ``` Provide the Airflow task name (`task_id`) and the Hopsworks job information (`project_id`, `job_name`, `args`). Set `wait_for_completion=True` to block until the job execution finishes. Wait for a job's most recent execution to be successful: ```python from hopsworks.airflow.sensors import HopsworksJobSuccessSensor HopsworksJobSuccessSensor( task_id="wait_for_profiles_fg", project_id=42, job_name="profiles_fg", ) ``` Wait for a HopsFS path to exist (replaces the Airflow 1.x `HopsworksHdfsSensor` plugin): ```python from hopsworks.airflow.sensors import HopsworksHdfsSensor HopsworksHdfsSensor( task_id="wait_for_arrival", project_id=42, path="Resources/landing/2026-05-22/_SUCCESS", ) ``` `project_id` can be replaced with `project_name=` if you prefer name-based lookup. !!! note "Authorization is automatic" Airflow 3 on Hopsworks does not use Airflow's `access_control` parameter. DAG visibility is enforced from the dag_id-to-project mapping that Hopsworks writes when the DAG is composed (see the [security model](security_model.md)); editing the DAG file cannot change ownership. Hopsworks admins (`HOPS_ADMIN`) have full access to every DAG; project members access DAGs of their own projects. #### Manage Airflow DAGs using Git You can leverage the [Git integration](../git/clone_repo.md) to track your Airflow DAGs in a git repository. Airflow will only consider the DAG files which are stored in the _Airflow_ Dataset in Hopsworks. After cloning the git repository in Hopsworks, you can automate the process of copying the DAG file in the _Airflow_ Dataset using [`DatasetApi.copy`][hopsworks_common.core.dataset_api.DatasetApi.copy] of the Hopsworks API. ================================================================================ # Airflow 3 upgrade Source: https://docs.hopsworks.ai/latest/user_guides/projects/airflow/airflow3_upgrade/ # Airflow 3 in Hopsworks Hopsworks now ships Apache Airflow 3.0.6 as its workflow engine. Airflow 3 is a major release with breaking changes to the DAG authoring API; the old 1.10-era DAGs do not run on it. This page covers what changed, what you need to do to your DAGs, and what the new per-project security model guarantees. ## Per-project DAG isolation A non-admin Hopsworks user can only see, trigger, edit, pause, clear, or read logs of DAGs that belong to a Hopsworks project they are a member of. The boundary is enforced on every authenticated request to the Airflow API server, every navigation in the Airflow UI, and every CLI call. The **Audit Log** is visible to every authenticated user but its rows are filtered server-side: non-admin users see only events whose `dag_id` belongs to one of their projects. The Hopsworks platform admin (`HOPS_ADMIN`) sees the unfiltered log. What this **does not** isolate in this release: - **Execution-time data access.** DAG tasks run in one shared scheduler process (LocalExecutor). A task can in principle read any Airflow Variable, Connection, or XCom row, regardless of project. Treat Airflow Variables, Connections, and Pools as cluster-wide. - **DAG parsing.** DAGs from all projects are parsed in one shared process. Module-top-level code in a DAG file runs with the dag-processor's privileges; treat it as cluster-wide too. These are tracked for a future release that switches to KubernetesExecutor plus per-team dag-processors. Until then, do not put project-private secrets in Airflow Variables or Connections. The per-DAG Hopsworks API key written by Hopsworks (see [API key for operators](#api-key-for-operators-no-embed)) is the exception, written by the platform itself rather than by users. ## What changed in the DAG API You **must rewrite** your existing 1.10 DAGs for Airflow 3. No automated rewrite tool ships with this release. Concrete things to change: | Old (Airflow 1.10) | New (Airflow 3.0.6) | | --- | --- | | `schedule_interval='@daily'` | `schedule='@daily'` | | `provide_context=True` | implicit, remove the argument | | `execution_date` in a template | `logical_date` | | `from airflow.operators.python_operator import ...` | `from airflow.operators.python import ...` | | `@apply_defaults` on custom operators | removed; declare `__init__` params normally | | `SubDagOperator` | TaskGroups + Assets | | `from airflow.models import BaseOperator` | `from airflow.sdk.bases.operator import BaseOperator` | | Custom Hopsworks operators imported via plugins | Provider package `apache-airflow-providers-hopsworks` | | Default `catchup_by_default = True` | Default `catchup=False`; set explicitly | | `schedule_interval='@continuous'` | Rejected by Hopsworks; use cron or `@once` | The Hopsworks-provided operators are now exposed via a standard provider: ```python from hopsworks.airflow.operators import HopsworksLaunchOperator # noqa: F401 from hopsworks.airflow.sensors import ( # noqa: F401 HopsworksHdfsSensor, HopsworksJobSuccessSensor, ) ``` `HopsworksHdfsSensor` replaces the legacy `HopsworksHdfsSensor` plugin from the 1.x shim. It polls `/hopsworks-api/api/project//dataset/?action=stat` and accepts either `project_id` or `project_name`. ## API key for operators (no embed) Hopsworks operators and sensors authenticate via the `HopsworksHook`, which resolves a credential in this order: 1. **Task token exchange**: the scheduler signs a per-task RS256 token; the hook POSTs it to `/api/auth/airflow-task-exchange/exchange` on Hopsworks and gets a project-scoped JWT back. 2. **Per-DAG Airflow Variable**: Hopsworks writes a per-DAG API key into an Airflow Variable named `hopsworks_api_key_` (Fernet-encrypted at rest) on every DAG compose. The hook reads it at task runtime via `Variable.get(...)` and uses it as `Authorization: ApiKey `. 3. **Airflow Connection** `hopsworks_default`: `conn.password` is read as a literal API key. Useful for out-of-cluster operators. 4. **`HOPSWORKS_API_KEY` env var**: manual override for power users. The generated DAG file **never carries the API key**. The secret lives only in the Airflow Variables table (admin-only via `HopsworksAuthManager`), so the DAG `.py` is safe to inspect, version-control, or share. Re-generate the DAG from the Hopsworks UI to rotate the key. DAG files composed before this change still embed `os.environ.setdefault("HOPSWORKS_API_KEY", "")` near the top of the file. They continue to work because the env-var path is the fourth tier in the hook's fallback, but the secret is in the file. Regenerate the DAG from the Hopsworks UI to drop the embed and switch to the Variable-fetch path. ## DAG identity The Airflow `dag_id` for a Hopsworks-composed DAG is now: ```text p____ ``` For example, project `acme` (id `42`) with a DAG named `daily_ingest` becomes `p_acme_42__daily_ingest` in the Airflow UI. The Hopsworks UI hides the prefix when displaying DAG names. If you reference your DAGs by `dag_id` from external code (XCom pulls across DAGs, `TriggerDagRunOperator`, REST API integrations), update those references to the new format. ## REST authentication The Airflow API server is reached through the standard Hopsworks reverse proxy at `https:///hopsworks-api/airflow/`. Browsers carry an HttpOnly `_token` cookie set by the auth manager; external clients use bearer tokens. To obtain a bearer token from a Hopsworks JWT: ```bash curl -X POST "https:///hopsworks-api/airflow/auth/token" \ -H "Content-Type: application/json" \ -d '{"hopsworks_jwt": ""}' ``` Response: ```json {"access_token": "", "token_type": "Bearer", "expires_in": 3600} ``` Use the returned `access_token` in `Authorization: Bearer` on all `/api/v2/*` calls. ## Recent runs in the Hopsworks UI The Hopsworks Airflow page lists each DAG with a **Last runs** column that renders the most recent runs as colored squares (green = success, red = failed, blue = running, yellow = queued / scheduled, gray = other). Hover any square for the run id, state, and start time. The data is read from a project-scoped Hopsworks endpoint that proxies to the auth manager and walks the same `dag_run` table the Airflow UI does, so the two views stay consistent. Clicking anywhere on a DAG row opens the DAG in the Airflow UI (in a new tab). The pencil column at the row's end opens the generated Python file in a Hopsworks editor without leaving Hopsworks. ## Metadata DB on upgrade Upgrading from Airflow 1.10 to Airflow 3 drops and recreates the Airflow metadata DB. Historical DAG-run records, task logs in the DB, and any ad-hoc Variables / Connections you had configured are not preserved. Snapshot HopsFS `Projects/

/Airflow/` before upgrade so you can roll back DAG files; the DB itself is not recoverable through the chart. ## Task logs across pod restarts Airflow 3 with the LocalExecutor writes task logs to the scheduler pod's local filesystem and serves them on port 8793 to the api-server pod. The scheduler records the source endpoint (host + port) on the task instance row. Hopsworks configures the scheduler with `[core] hostname_callable = airflow.utils.net.get_host_ip_address` so that endpoint is the pod IP (routable across pods), not the pod's DNS hostname (not resolvable from sibling pods). Logs from runs that started before the scheduler pod was last restarted are unrecoverable: the pod's filesystem is ephemeral. Re-trigger the DAG to regenerate task logs if you need them. ================================================================================ # Security model Source: https://docs.hopsworks.ai/latest/user_guides/projects/airflow/security_model/ # Airflow Security Model Hopsworks deploys Apache Airflow 3 with a custom auth manager that enforces per-Hopsworks-project DAG visibility. This page is the authoritative reference for what the security boundary covers and what it does not. ## What is isolated Every authenticated request to the Airflow API server, every page in the Airflow UI, every CLI call: | Surface | Behavior for a non-admin user | | --- | --- | | `GET /api/v2/dags` | Returns only DAGs in the user's projects | | `GET /api/v2/dags/{dag_id}` | 404 for cross-project DAGs | | `POST /api/v2/dags/{dag_id}/dagRuns` (trigger) | 403 for cross-project DAGs | | `POST /api/v2/dags/{dag_id}/pause` | 403 for cross-project | | `GET /api/v2/dags/.../taskInstances/.../logs` | 403 for cross-project | | Airflow UI menu | Connections / Variables / Pools / Config / Plugins / Providers / Cluster Activity stripped (admin-only). Audit Log is visible to all users but rows are filtered server-side to `dag_id IN (user's project dag_ids)`. | | `GET /api/v2/dagSources/{dag_id}` | 403 for cross-project | | Direct HopsFS access to `Projects/

/Airflow/` | POSIX ACLs deny cross-project read | The dag-to-project map is written by the Hopsworks backend whenever a DAG is composed or deleted, and stored in a table inside Airflow's own metadata DB. Editing a DAG file directly (e.g. changing `tags=[...]`) cannot move the DAG to a different project's namespace. ### Active-project scoping Opening Airflow from a project's UI narrows visibility further to that project alone, even for users who are members of several. The Hopsworks proxy forwards the project context via a `?hopsworks_project=` query parameter that the auth manager turns into an `active_project_id` claim on the issued Airflow JWT. The DAG list, the per-DAG endpoints, and the Audit Log are all filtered against the active project for the lifetime of the session. Switching project in the Hopsworks UI re-mints the Airflow JWT with the new active project; the previous session's cookie remains valid for its TTL but is scoped to the previous project. A Hopsworks admin opening Airflow without a project context sees every DAG; opening from a project still scopes to that project. ## What is **not** isolated The shared `dag-processor` parses DAGs from all projects. The shared LocalExecutor runs tasks from all projects in one process tree. As a consequence: - Tasks can read any Variable, Connection, or Pool present in the metadata DB. - Tasks can read XCom rows belonging to other projects' tasks. - DAG-author code at module top level runs in the shared parser with metadata DB credentials. - A task can in principle read environment variables of co-running tasks on the same scheduler pod. A future release switches to KubernetesExecutor with per-team dag-processors to close these gaps. Until then: - **Do not store project-private secrets in Airflow Variables or Connections.** Use the Hopsworks-side secrets API for per-project credentials. Hopsworks operators obtain a short-lived project-scoped token at task runtime via the Execution-API token exchange, so the primary auth path does not depend on Airflow Variables. - **The per-DAG `hopsworks_api_key_` Variable is a fallback only, and is not per-DAG isolated at runtime.** Hopsworks writes it during compose so the task-token-exchange path has a working credential to fall back to if the task-instance JWT is unreachable for any reason. The Variables UI is admin-only via `HopsworksAuthManager`, but inside a running task `Variable.get("hopsworks_api_key_")` is not blocked. `dag_id` is non-secret and the hash is reproducible, so DAG code that calls `Variable.get` for another DAG's hashed name can read that key. Until per-task secret isolation lands with KubernetesExecutor, treat the fallback path as a shared credential surface within the cluster: don't run untrusted DAG code, and prefer the task-token-exchange path which is signed per-task and not stored in the metadata DB. - **Do not assume DAG code is sandboxed at parse time.** Module-top-level code in a customer DAG can execute network calls and reach anything the dag-processor's ServiceAccount can reach. ## Token + cookie behavior The auth manager sets a `_token` cookie on UI logins. The cookie is `HttpOnly`, `Secure` (under TLS), `SameSite=Lax`, and scoped to `/hopsworks-api/airflow`. The Hopsworks proxy passes it through unchanged; cookie path scoping is set by Airflow itself from the `[api] base_url` config. The same auth manager mints bearer tokens for external clients via `POST /hopsworks-api/airflow/auth/token` (the auth manager's `/auth` router is mounted under the Airflow `[api] base_url`, which the Hopsworks chart pins to `/hopsworks-api/airflow/`). The bearer token format and validation are identical to the cookie-borne token; only the carrier differs. The Airflow JWT carries the user's `project_ids`, `project_roles`, and `is_admin` flag at mint time. Membership is **not** refreshed per request, because Airflow 3.0.x does not expose a synchronous refresh hook on `BaseAuthManager`, so a user's authorization is stable for the cookie's TTL (1 hour by default). ## Project membership changes When a Hopsworks project's membership changes (add member, remove member, role change), the Hopsworks backend immediately invalidates the affected user's entry in the Airflow auth manager's cache via `/auth/internal/invalidate`. The next login by that user re-fetches their project list and reflects the new membership. The cache also has a 60-second safety-net TTL even without an explicit invalidation. ================================================================================ # Apps Source: https://docs.hopsworks.ai/latest/user_guides/projects/apps/ # Apps Apps are long-running applications that run as managed services in Hopsworks. Use them for Streamlit dashboards or custom web apps such as Flask, FastAPI, Gradio, or JavaScript apps like Express. Common uses include: - Interactive data apps that visualize model predictions or project data. - Predictive analytics dashboards. - Chatbot UIs for internal or external assistants. - GenAI front ends, such as a Gradio or Streamlit RAG assistant. - FastAPI services that expose inference or other application endpoints. Each app is backed by a Hopsworks job and a Kubernetes deployment, so it can be started, stopped, redeployed, and deleted like any other project service. ## Where to find Apps 1. Open your project in Hopsworks. 2. In the current sidebar, open **AI/ML** and click **Apps**. 3. Click **New App** to create one. The Apps page lists each app with its name, owner, state, UI link, uptime, and action buttons.

Apps list
The Apps page with one app serving

!!! note "Shared WebSocket capacity" Apps share the same per-pod WebSocket session pool as Jupyter and terminals. If you see capacity warnings, close unused sessions or see [Session Capacity Warnings](../jupyter/session_capacity_warnings.md). ## Creating an app The create dialog lets you choose the app type, source, runtime environment, app base path, readiness probe path, resources, monitoring, and per-app environment variables.

New App form
The New App form: type, source, app file and resources

| Setting | Typical value | | --- | --- | | App type | `STREAMLIT` or `CUSTOM` | | Proxy routing mode | `Root routing` for new apps, `Compatibility prefix` for legacy apps | | App base path | `/` for the app root, `/myapp` for a subpath | | Readiness probe path | Leave empty to use the platform default | | Source | Project file or Git repository | | Environment | `python-app-pipeline` | | Memory | `2048` MB | | CPU cores | `1.0` | | Custom app port | `8080` | ### App types - **Streamlit** apps are the default choice for dashboards and interactive ML UIs. - **Custom** apps are any web service that listens on `APP_PORT`. Common choices are Flask, FastAPI, Gradio, and JavaScript frameworks such as Express. Use `App base path` to choose where Hopsworks mounts the app. Set it to `/` for a root-based app or `/myapp` for a subpath. Legacy prefix routing is only for older apps that still depend on `APP_BASE_URL_PATH`. Use `Proxy routing mode` to switch between `Root routing` and `Compatibility prefix`. The compatibility mode is only for migrating older apps. ### App sources - **Project file** means a file already stored in HopsFS or the project file browser. - **Git repository** means the app source is cloned on every start. This is useful when you want a proper Git-backed CI/CD flow and when you want local file edits not to affect a running production app. The deployed app only sees the repository contents that are present when it starts, so changes in your working tree stay local until you commit, push, and redeploy. For Streamlit apps, a project file must be a `.py` file. For Git-backed Streamlit apps, you also need to provide the entrypoint script relative to the repository root. For custom apps, the entrypoint command is required and the app file is optional. #### Auto-redeploy on new commits Git-backed apps can roll themselves to the branch HEAD whenever a new commit is pushed. Enable **Auto-redeploy on new commits** in the app settings, or pass `git_auto_redeploy=True` to the SDK. Hopsworks polls the remote branch and, when it moves, rolls the app onto the new commit. The running app keeps serving until the new version is ready, so there is no gap in availability. While the roll is in progress the app shows **Redeploying** in the apps list. The setting only applies to Git-backed apps. An app without a Git source has nothing to poll, and Hopsworks rejects the flag in that case. If you do not set a branch, the clone follows the repository's default branch, and the app details page shows the branch it resolved to once the app has run. ### Routing and readiness The browser URL for every app is the public mount point under `/hopsworks-api/pythonapp///`. If you set `App base path` to `/myapp`, the full public URL becomes `/hopsworks-api/pythonapp///myapp/`. Hopsworks strips the public mount prefix before forwarding requests upstream, so app code can stay root-based. Hopsworks forwards `X-Forwarded-Prefix` for frameworks that need to generate absolute links. The proxy routing mode controls whether Hopsworks strips that prefix or preserves the legacy browser path. Open **App Settings** and change `Proxy routing mode` under **App routing and readiness** to switch an app. Readiness is separate from browser routing. Streamlit defaults to `/_stcore/health`. Custom apps default to `/`. You can override the readiness probe path in the app settings dialog or API when needed. ### Example structure Most apps follow the same pattern: 1. Put the app code in your project or Git repository. 2. Choose the `python-app-pipeline` environment or clone it if you need extra libraries. 3. Set the resources the pod should reserve. 4. Add monitoring routes if you want Envoy metrics for specific paths. 5. Start the app and wait for it to reach `Serving`. !!! tip "Default environment" The `python-app-pipeline` environment is the default runtime for Apps. If your app needs additional dependencies, clone that environment and install the extra packages there instead of modifying the base image. ## Writing app code ### Streamlit apps

A Streamlit app served by Hopsworks
A Streamlit dashboard reading a feature group, served through the Hopsworks proxy

Streamlit apps are launched with `streamlit run` behind the Hopsworks proxy. Hopsworks manages the mount prefix for you, so Streamlit apps can stay root-based. If the app is Git-backed, the entrypoint script is relative to the repository root. The default readiness probe for Streamlit is `/_stcore/health`. ### Custom apps Custom apps should bind to `0.0.0.0` and use the injected `APP_PORT`. Define routes at `/` and `/health`. Use legacy prefix-aware routes only while migrating an older app that still depends on `APP_BASE_URL_PATH`. ```python import os import uvicorn from fastapi import FastAPI app = FastAPI() @app.get("/health") def health(): return {"status": "ok"} @app.get("/") def home(): return {"status": "ready"} if __name__ == "__main__": uvicorn.run(app, host="0.0.0.0", port=int(os.environ["APP_PORT"])) ``` ### App base path `APP_BASE_URL_PATH` is deprecated and should not be used in new apps. For new apps, use `App base path` in the UI or API and let Hopsworks handle the browser mount prefix. Keep `Proxy routing mode` set to `Root routing` for new apps. Examples: - FastAPI: `@app.get("/health")` - Flask: `Blueprint(..., url_prefix="/")` - Gradio: `demo.launch(..., root_path=None)` - Express: `app.use("/", router)` If you are migrating an older app that still depends on `APP_BASE_URL_PATH`, keep the legacy pattern only until the app code can move to root-based routing. In that case, use `Compatibility prefix` in the app settings dialog during the migration period. See the matching examples in [appshopsworkstests](https://github.com/gibchikafa/appshopsworkstests). ## Runtime and environment variables Apps run inside a project environment and receive the same project context as other Hopsworks services. Per-app environment variables are applied every time the app starts. The platform injects app-specific variables such as: - `APP_BASE_URL_PATH` for legacy prefix-mode apps only - `APP_PORT` - `STREAMLIT_BASE_URL_PATH` - `STREAMLIT_PORT` - `APP_FILE` - `APP_PATH` - `APP_KIND` - `APP_ARGS` Git-backed apps also receive the Git URL, provider, branch, and Streamlit entrypoint script. `APP_BASE_URL_PATH` is only injected for legacy prefix-mode apps during migration. Some runtime names are reserved by the platform and cannot be overridden in the UI. That includes the app-path and routing variables above, plus other platform-managed names that start with `HOPS_`, `HOPSWORKS_`, `HOPSFS_`, or `AGENT_`. ## Database and feature store access Every app can reach the project's feature store data from inside the pod: - the project's **online feature store database** on RonDB (MySQL protocol), where the online feature group tables live and where the app can keep its own tables: sessions, settings, agent memory, job results. This is on by default and controlled by **Database access** in the create dialog, `db_access` in the SDK and `--no-db-access` in the CLI; - the project's **offline feature groups**. A Python app reads them with the Hopsworks Python SDK, as any other Hopsworks client does: [feature group](../../fs/feature_group/index.md) reads and [feature view](../../fs/feature_view/batch-data.md) batch data. When Trino is enabled on the cluster, the app also gets the [Trino query engine](../trino/query_engine.md) as an SQL path to the same tables, which is what an app in another language uses. This does not depend on the database access flag: an app created with `db_access=False` still gets the Trino variables. The database is created on demand the first time an app with database access starts, so it works in a project that never created an online feature group. The app finds everything in its environment; nothing has to be configured. | Variable | Value | | --- | --- | | `MYSQL_HOST`, `MYSQL_PORT` | the online feature store MySQL server | | `MYSQL_DB` | the project database, the project name in lowercase | | `MYSQL_USER` | the MySQL user of the person who **started** the app | | `MYSQL_PASSWORD_SECRET_NAME` | the Hopsworks secret holding that user's password | | `TRINO_HOST`, `TRINO_PORT` | the Trino coordinator (HTTPS), when Trino is enabled | | `TRINO_USER` | the Trino user of the person who started the app, `__` | | `TRINO_PASSWORD_SECRET_NAME` | the Hopsworks secret holding that user's Trino password | | `TRINO_SCHEMA` | the project's offline feature store schema, `_featurestore` | | `LIBHDFS_ROOT_CA_BUNDLE`, `NODE_EXTRA_CA_CERTS` | the cluster CA as a PEM file, so any HTTP client can verify the Hopsworks REST API and the Trino coordinator | Passwords are never placed in the environment. They are private secrets of the user who started the app, and the app reads them with its own credentials, through the Python SDK or the REST API. The privileges follow that user's project role: an app started by a Data Owner can create tables and write, an app started by a Data Scientist has read-only access. There is no `TRINO_CATALOG` because the catalog depends on each feature group's format: `delta` for Delta feature groups, `hudi` for Hudi ones. The variables are listed on the app details page under **Environment variables**, next to the per-app ones. ### Python Read offline feature groups with the Hopsworks Python SDK first. It is what the rest of the platform uses, it knows the feature group's format and location, and a feature view adds point-in-time joins and the model's transformations, so a dashboard and the model it shows stay consistent. ```python import hopsworks project = hopsworks.login() # in-cluster: no prompt fs = project.get_feature_store() # Offline feature group: a DataFrame, filtered and projected on the server side transactions = fs.get_feature_group("transactions", version=1) recent = ( transactions.select(["cc_num", "amount", "event_time"]) .filter(transactions.event_time >= "2025-01-01") .read() ) # Feature view: batch data with the point-in-time joins and transformations of the model fv = fs.get_feature_view("fraud_model", version=1) batch = fv.get_batch_data(start_time="2025-01-01", end_time="2025-02-01") ``` The online database is for the app's own tables and for primary-key lookups on the online feature group tables: ```python import os import pymysql password = project.get_secrets_api().get(os.environ["MYSQL_PASSWORD_SECRET_NAME"]) conn = pymysql.connect( host=os.environ["MYSQL_HOST"], port=int(os.environ["MYSQL_PORT"]), user=os.environ["MYSQL_USER"], password=password, database=os.environ["MYSQL_DB"], ) ``` Create the app's own tables with an explicit `ENGINE=NDBCLUSTER`, a primary key, and an `app_` prefix so they never collide with feature group tables (`_`). Write features through `feature_group.insert()`, not straight into the online tables; reading them with SQL is fine. Trino is an extra option for a Python app: ad-hoc SQL over the offline tables, a join with another Trino catalog, or a query the SDK does not express. The SDK wraps the connection: ```python trino = project.get_trino_api().connect( catalog="delta", schema=os.environ["TRINO_SCHEMA"] ) cursor = trino.cursor() cursor.execute( "SELECT * FROM transactions_1 WHERE event_time >= DATE '2025-01-01' LIMIT 100" ) rows = cursor.fetchall() ``` ### JavaScript A [custom app](#custom-apps) can run Node.js: the `python-app-pipeline` environment ships Node and the `@hopsworks/app` module, which turns the variables above into ready-to-use connections. Import it from any app without adding it to `package.json`. ```js import mysql from "mysql2/promise"; import { mysqlConfig, trinoClient, query, getSecret } from "@hopsworks/app"; // Online feature store / the app's own tables const pool = mysql.createPool({ ...(await mysqlConfig()), connectionLimit: 5 }); const [rows] = await pool.execute("SELECT * FROM transactions_1 WHERE cc_num = ?", [ccNum]); // Offline feature groups through Trino; the catalog is the feature group's format const trino = await trinoClient({ catalog: "delta" }); const recent = await query(trino, "SELECT cc_num, amount, event_time FROM transactions_1 WHERE event_time >= DATE '2025-01-01' LIMIT 100"); // Any other secret of the user who started the app const apiKey = await getSecret("openai_api_key"); ``` `mysqlConfig()` returns `{ host, port, user, password, database }` for `mysql2`, `mysql` or `knex`. `trinoClient()` returns a client authenticated as the starting user that speaks the [Trino REST protocol](https://trino.io/docs/current/develop/client-protocol.html); `query()` collects a result as an array of row objects and `streamQuery()` yields rows page by page for large results, cancelling the query if you stop early. Integers above 2^53 come back as `BigInt`, so identifiers are never rounded. Both resolve the password once per process from the Hopsworks secret, and TLS to the platform works out of the box through `NODE_EXTRA_CA_CERTS`. Outside a Hopsworks pod the functions throw an error naming the missing variable; guard local development on `inHopsworks()`. ```python node_app = apps.create_app( "node_api", app_kind="CUSTOM", git_url="https://github.com/my-org/node-api.git", git_provider="GitHub", entrypoint_command='bash -lc "npm ci --omit=dev && exec node server.js"', app_port=8080, ) ``` Trino is the right path for scans and aggregations over the offline tables; primary-key lookups belong on the online tables. Trino's HTTP protocol has no bound parameters, so never interpolate user input into SQL text. ## Managing an app The Apps list and the app details page expose the same lifecycle actions: - **Start** launches the current app configuration. - **Stop** stops the running execution. - **Restart / Redeploy** rolls the Kubernetes deployment and starts a fresh execution with the same configuration. - **Open App** opens the serving URL in a new tab. - **Logs** shows stdout and stderr for the latest execution. - **Kubernetes status** shows the deployment and pod health. - **Edit** opens the app settings dialog. - **Delete** removes the app entirely. Edit and delete are only available when the app is stopped. The app URL is only shown once the backend confirms that the app is actually serving. That means `RUNNING` and `Serving` are not the same thing: `Serving` is the state where the Hopsworks proxy can reach the app end to end. ### What the app details page shows

App details page
The app details page: lifecycle actions, source, URL and metrics

The details page includes: - App details, type, and description - App URL - App base path, proxy routing mode, and readiness probe path - Public access status for Streamlit apps - Git source details, if the app is Git-backed, including the branch, the deployed commit, and whether auto-redeploy is enabled - Monitoring configuration - Resource requests - Runtime environment - Environment variables: the per-app ones and the database and feature store access variables the platform injects - App metrics and Kubernetes health ## Public access for Streamlit apps !!! warning "Public Streamlit links are not read-only" Anyone with the link can use the app with the app's own credentials, secrets, and data access. Disabling public access revokes all live links immediately. Public access is only available when all of the following are true: - The app is a Streamlit app. - You are a Data Owner in the project. - The administrator has enabled the `streamlit_sharing` feature flag. - The app is currently serving. When public access is enabled, Hopsworks shows a share link in the UI. The link is built through the Hopsworks proxy, so users still access the app through the platform rather than directly. ## Monitoring and metrics Apps can publish Envoy-based request metrics. Monitoring is enabled by default, and route filters are optional. - Use exact or prefix route matches to narrow the traffic that gets counted. - Leave routes empty if you want the default behavior. - For Streamlit apps, the platform automatically ignores framework noise such as static assets, health checks, and websocket handshake traffic. - For custom apps, route filters are useful for paths such as `/api` or `/predict`. The app details page embeds metrics for request count, request rate, latency, CPU usage, and memory usage. ## Python SDK Use the Python SDK when you want to create or manage apps from code. ```python import hopsworks project = hopsworks.login() apps = project.get_app_api() app = apps.create_app( "customer_dashboard", app_path="Resources/app.py", db_access=True, # default: the online database, see above ) app.run() print(app.app_url) ``` A Git-backed app that redeploys itself on every push: ```python app = apps.create_app( "customer_dashboard", app_kind="STREAMLIT", git_url="https://github.com/my-org/my-app.git", git_provider="GitHub", git_branch="main", git_auto_redeploy=True, entrypoint_script="src/app.py", ) ``` Common SDK methods: - `project.get_app_api()` - `apps.create_app(...)` - `app.run()` - `app.redeploy()` - `app.stop()` - `app.delete()` - `app.make_public()` / `app.make_private()` - `app.app_url` ## CLI The `hops` CLI wraps the same app lifecycle API: ```bash hops app list hops app info hops app url hops app create --path /Projects//Resources/app.py --start hops app start hops app redeploy hops app stop hops app logs hops app delete --yes ``` Use `--git-url` and `--entrypoint-script` for Git-backed Streamlit apps. Use `--entrypoint-command` and `--app-port` for custom apps. Add `--git-auto-redeploy` to roll a Git-backed app onto every new commit. Pass `--no-db-access` for an app that must not get the online database variables; Trino access does not depend on it. ## See also - [Python Environments](../python/python_env_overview.md) - [Clone a Python Environment](../python/python_env_clone.md) - [Python Deployment](../python-deployment/python-deployment.md) - [Session Capacity Warnings](../jupyter/session_capacity_warnings.md) - [Superset](../superset/superset.md) - [Query Engine (Trino)](../trino/query_engine.md) - [Secrets](../secrets/create_secret.md) ================================================================================ # Deployment Creation Source: https://docs.hopsworks.ai/latest/user_guides/projects/python-deployment/python-deployment/ # Python Deployment { #python-deployment } ## Introduction Python deployments allow you to deploy a Python script as a service without requiring a model artifact in the Model Registry. This is useful for custom inference pipelines, feature view deployments, or any Python-based program that needs to be served behind an HTTP endpoint. !!! warning "Incoming requests are directed to port 8080" Python deployments run your script directly on port 8080. Therefore, make sure your implementation listens to 8080 port for handling incoming requests. !!! info "gRPC protocol not supported" !!! tip "Use your favourite HTTP server" There are no constraints on the framework or library used: you can use Flask, FastAPI, or any other HTTP server. In each Python deployment, you can configure the following: !!! info "" 1. [Python environments](#python-environments) 2. [Resources](#resources) 3. [Autoscaling](#autoscaling) 4. [Scheduling](#scheduling) Like model deployments, Python deployments keep a numbered history of their configuration, so a change to the script, environment, resources or scaling can be saved as a new version and rolled back, see [Versions](#versions). ## Web UI A Python deployment and an agent deployment are the same thing: a Python script served as a long-running service without a model artifact. Whether the script is a plain HTTP server or an autonomous agent, it is created and managed the same way. In the current UI both are listed under `Agent Deployments` and created from the same form. ### Step 1: Create a new deployment Navigate to `Agent Deployments` under the `Agents` section of the navigation menu on the left, then click on `New agent`.

Agent Deployments navigation tab
Agent Deployments navigation tab

### Step 2: Configure the deployment Choose a name for your deployment. Under `Source`, keep `Project file` to run a script already in the project, or select `Git repository` to clone the script on every start. Then provide the script under `Agent script file` by clicking on `From project`, `Upload new file` or `Create new file`. ### Step 3 (Optional): Change Python environment The script runs in one of the [Python Environments](../../projects/python/python_env_overview.md) available in your project. This environment must have all the necessary dependencies for your Python program. Select an environment from the `Python Environment` dropdown. Hopsworks provides built-in environments such as `python-agent-pipeline` and the `*-inference-pipeline` environments, each with a different set of libraries pre-installed. To create your own environment it is recommended to [clone](../../projects/python/python_env_clone.md) the `minimal-inference-pipeline` or `pandas-inference-pipeline` environment and install additional dependencies needed for your Python program.

Python environment in the deployment form
Select a Python environment for the program

### Step 4 (Optional): Advanced configuration Click on `advanced options` to configure the deployment further, including: !!! info "" 1. [Resources](#resources) 2. [Autoscaling](#autoscaling) 3. [Scheduling](#scheduling) Once you are done with the changes, click on `Create` at the bottom of the form to create the deployment. ## Code ### Step 1: Connect to Hopsworks === "Python" ```python import hopsworks project = hopsworks.login() # get Hopsworks Model Serving handle ms = project.get_model_serving() ``` ### Step 2: Implement a Python script === "Python" ```python import uvicorn from fastapi import FastAPI app = FastAPI() @app.get("/ping") async def ping(): return {"status": "ready"} @app.post("/echo") async def echo(data: dict): return data if __name__ == "__main__": uvicorn.run(app, host="0.0.0.0", port=8080) ``` !!! info "Jupyter magic" In a jupyter notebook, you can add `%%writefile python_server.py` at the top of the cell to save it as a local file. ### Step 3: Upload the script to your project === "Python" ```python import os dataset_api = project.get_dataset_api() uploaded_file_path = dataset_api.upload("python_server.py", "Resources", overwrite=True) script_path = os.path.join("/Projects", project.name, uploaded_file_path) ``` ### Step 4: Create a deployment === "Python" ```python py_server = ms.create_endpoint( name="pyserver", script_file=script_path ) py_deployment = py_server.deploy() ``` ### Step 5: Send requests === "Python" ```python import requests url = py_deployment.get_endpoint_url() response = requests.post(f"{url}/echo", json={"key": "value"}) print(response.json()) ``` ## Environment variables A number of different environment variables is available in the Python deployment to ease its implementation. !!! tip "Available environment variables" === "Deployment" These variables are available in all deployments. | Name | Description | | --------------------- | -------------------------------- | | `DEPLOYMENT_NAME` | Name of the current deployment | | `DEPLOYMENT_VERSION` | Version of the deployment | | `ARTIFACT_FILES_PATH` | Local path to the artifact files | === "Python deployment" These variables are specific to Python deployments. | Name | Description | | ------------------ | -------------------------------------------------- | | `SCRIPT_PATH` | Full path to the Python script | | `SCRIPT_NAME` | Prefixed filename of the Python script | | `CONFIG_FILE_PATH` | Local path to the configuration file (if provided) | === "Others" These variables are available in all deployments. | Name | Description | | ------------------------ | -------------------------------------------------- | | `REST_ENDPOINT` | Hopsworks REST API endpoint | | `HOPSWORKS_PROJECT_ID` | ID of the project | | `HOPSWORKS_PROJECT_NAME` | Name of the project | | `HOPSWORKS_PUBLIC_HOST` | Hopsworks public hostname | | `API_KEY` | API key for authenticating with Hopsworks services | | `PROJECT_ID` | Project ID (for Feature Store access) | | `PROJECT_NAME` | Project name (for Feature Store access) | | `SECRETS_DIR` | Path to secrets directory (`/keys`) | | `MATERIAL_DIRECTORY` | Path to TLS certificates (`/certs`) | | `REQUESTS_VERIFY` | SSL verification setting | ## Python environments Python deployments run in one of the `*-inference-pipeline` Python environments available in your project. Hopsworks provides built-in environments like `minimal-inference-pipeline`, `pandas-inference-pipeline` or `torch-inference-pipeline` with different sets of libraries pre-installed. By default, the `pandas-inference-pipeline` environment is used. To create your own environment, it is recommended to [clone](../../projects/python/python_env_clone.md) the `minimal-inference-pipeline` or `pandas-inference-pipeline` environment and install additional dependencies needed for your Python program. To learn more about Python environments, see [Python Environments](../../projects/python/python_env_overview.md). ## Resources Configure CPU, memory, and GPU allocation for your Python deployment. Each deployment component has separate request and limit values. For full details on resource configuration, see the [Resources Guide](../../mlops/serving/resources.md). ## Autoscaling Deployments use **Knative Pod Autoscaler (KPA)** to automatically scale the number of replicas based on traffic. You can configure the minimum and maximum number of instances as well as the scale metric (requests per second or concurrency). For full details on autoscaling parameters, see the [Autoscaling Guide](../../mlops/serving/autoscaling.md). ## Scheduling !!! info "Kueue is required" This feature requires Kueue to be enabled in your cluster. If Kueue is not available, queue and topology options will not be accessible. If the cluster has Kueue enabled, you can select a queue for your deployment from the advanced configuration. Queues control resource allocation and scheduling priority across the cluster. For full details on scheduling configuration, see the [Scheduling Guide](../../mlops/serving/scheduling.md). ## Versions A Python deployment keeps its configuration in numbered versions, the same way a model deployment does. In the edit form, `Save` edits the active version in place and `Save as new version` stores the changes as a new version and activates it. The `Versions` card on the overview page lists the versions, shows the configuration of each one, and rolls back to an earlier one. !!! info "Not covered by versions" Scheduling and Knative mode are not part of a version, so a rollback keeps their current values. A script read from a HopsFS path or a git repository is stored in the version as that path or repository, so a rollback does not bring back earlier code, and the deployment restarts on every save to pick up the current one. For full details, see the [Deployment Versions Guide][deployment-versions]. ================================================================================ # REST API Source: https://docs.hopsworks.ai/latest/user_guides/projects/python-deployment/rest-api/ # Python Deployment REST API ## Introduction Python deployments are accessible via REST API through the [Istio](https://istio.io/) ingress gateway. This document explains how to send requests to a Python deployment. !!! tip "Tutorials" End-to-end examples are available in the [hopsworks-tutorials](https://github.com/logicalclocks/hopsworks-tutorials/tree/master) repository. ## Sending Requests through Istio Ingress The full URL path is constructed by combining a base path with a resource path available in the Python server. See [URL Paths](#url-paths) for the complete URL format and examples. ### Authentication All requests must include an API Key for authentication. You can create an API key by following this [guide](../../projects/api_key/create_api_key.md). Include the key in the `authorization` header: ```text authorization: ApiKey ``` ### Headers | Header | Description | Example Value | | --------------- | --------------------------- | ----------------------- | | `authorization` | API key for authentication. | `ApiKey ` | | `content-type` | Request payload type. | `application/json` | ## URL Paths Python deployments are accessible through the ==Istio ingress gateway== using **path-based** routing. The full URL is constructed by combining the base URL with the paths defined in your Python server. !!! example "" **`/`** Where `` depends entirely on the routes defined in your Python server implementation (e.g., `/echo`, `/predict`, `/health`). ### Base URL The base URL is composed of the **Istio ingress gateway IP**, the **project name**, and the **deployment name**. !!! example "" **`https:///v1//`** !!! warning "Host-based routing (legacy)" Prior to path-based routing, requests were routed using a `Host` header matching the deployment hostname, and **`https://`** as base url. ``` Host: .. ``` Each deployment gets its own Knative-generated hostname, and routing depends on the `Host` header matching Istio ingress gateway rules. Path-based routing (described above) is the preferred method for external access. !!! tip "Hopsworks Python API" The endpoint URL can be retrieved using the `Deployment` class. ```python # Returns: https:///v1// endpoint_url = deployment.get_endpoint_url() ``` ## Request Format The request format depends entirely on your Python server implementation. There are no framework or protocol constraints: your server defines the expected HTTP methods, paths, and payload format. !!! example "REST API example" === "Python" ```python import requests url = deployment.get_endpoint_url() response = requests.post(f"{url}/echo", json={"key": "value"}) print(response.json()) ``` === "Curl" ```bash curl -X POST "https:///v1/my_project/pyserver/echo" \ -H "authorization: ApiKey " \ -H "content-type: application/json" \ -d '{"key": "value"}' ``` ## CORS The Istio EnvoyFilter handles CORS preflight (`OPTIONS`) requests automatically. Allowed origins can be configured via `istio.envoyFilter.corsAllowedOrigins` in the Helm chart configuration. ## Response The response format depends on your Python server implementation. ================================================================================ # Troubleshooting Source: https://docs.hopsworks.ai/latest/user_guides/projects/python-deployment/troubleshooting/ # How To Troubleshoot A Python Deployment ## Introduction In this guide, you will learn how to troubleshoot a deployment that is having issues running. But before that, it is important to understand how [deployment states](../../mlops/serving/deployment-state.md) are defined and the possible transitions between conditions. Before a deployment starts, it goes through a `CREATING` phase where deployment artifacts are prepared. When a deployment is starting, it follows an ordered sequence of [states](../../mlops/serving/deployment-state.md#deployment-conditions) before becoming ready for handling requests. Similarly, it follows an ordered sequence of states when being stopped, although with fewer steps. !!! warning "`FAILED` is a terminal state" If a deployment reaches the `FAILED` state, it cannot recover on its own. You must stop and restart the deployment to attempt recovery. ## Web UI ### Step 1: Inspect deployment status If you have at least one deployment already created, navigate to the deployments page by clicking on the `Deployments` tab on the navigation menu on the left.

Deployments navigation tab
Deployments navigation tab

Once in the deployments page, find the deployment you want to inspect. Next to the actions buttons, you can find an indicator showing the current status of the deployment. For a more descriptive representation, this indicator changes its color based on the status. To inspect the condition of the deployment, click on the name of the deployment to open the deployment overview page. ### Step 2: Inspect condition At the top of page, you can find the same status indicator mentioned in the previous step. Below it, a one-line message is shown with a more detailed description of the deployment status. This message is built using the current status [condition](../../mlops/serving/deployment-state.md#deployment-conditions) of the deployment. Oftentimes, the status and the one-line description are enough to understand the current state of a deployment. For instance, when the cluster lacks enough allocatable resources to meet the deployment requirements, a meaningful error message will be shown with the root cause.

Deployment failed to schedule condition
Condition of a deployment that cannot be scheduled

However, when the deployment fails to start further details might be needed depending on the source of failure. For example, failures in the initialization or starting steps will show a less relevant message. In those cases, you can explore the deployments logs in search of the cause of the problem.

Deployment failed to start condition
Condition of a deployment that fails to start

### Step 3: Explore transient logs Each deployment is composed of several components depending on its configuration. Transient logs refer to component-specific logs that are read directly from the running component. Therefore, these logs can only be retrieved as long as the deployment components are running. !!! info "" Transient logs are informative and fast to retrieve, facilitating the troubleshooting of deployment components at a glance. Transient logs are convenient when access to the most recent logs of a deployment is needed. To follow them in the UI, click the `Logs` button at the top of the deployment overview page. The pane tails the selected component every two seconds and lets you search, copy and download what it has buffered. You can also read them with the Hopsworks Machine Learning Python library, as shown in [Step 4](#step-4-explore-transient-logs) of the code section. !!! info When a deployment is in idle state, there are no components running (i.e., scaled to zero) and, thus, no transient logs are available. Use historical logs to inspect an instance that is already gone. !!! note Standard output and standard error arrive as a single interleaved stream. Kubernetes merges them at the container runtime, so the two cannot be separated after the fact. ### Step 4: Explore historical logs Historical logs are archives that each instance writes to the project's `Logs` dataset from inside its own container. An instance archives its output when it exits, is restarted, or is stopped, which means an instance removed by scale-to-zero or replaced by a new deployment revision still leaves its logs behind. !!! info "" Historical logs are convenient when a deployment fails occasionally, or when the instance you need to inspect is no longer running. Archives are written to `Logs/Serving//` and named `__.log`, one file per instance run. Browse them under the `Logs` section of the deployment overview page, or in the `Logs` dataset, and open one to read it. Historical logs are only written for components that have disk logging enabled. See [configuring disk logging](#configuring-disk-logging) below, and note that only Python deployments support it: Python predictors have it on by default. !!! warning The number of archives kept per deployment is capped by the `log_history_limit` cluster variable, which defaults to 30. Once the cap is reached, the oldest archive is deleted each time a new one is written, so long-lived deployments do not fill the project with logs. To retrieve archives with the Python library, use `deployment.download_logs()`, shown in [Step 5](#step-5-download-historical-logs) below. ### Configuring disk logging Disk logging controls whether a component archives its output to the project's `Logs` dataset. It is a per-component checkbox under `Disk logging` in the advanced options of the deployment form. When it is on, each instance keeps its output on local disk while it runs and uploads it when it stops, to a separate file distinguished by pod name. This covers stops the platform initiates on its own, such as scale-to-zero and revision replacement, not only stops a user asks for. When it is off, nothing is written. Disk logging is only available for deployments whose serving container runs a Hopsworks inference pipeline image, because the upload runs the Hopsworks Python library from inside that container. That means Python deployments, agent deployments included, and their transformers; Python predictors have it on by default. TensorFlow Serving and vLLM do not support it, and neither does a KServe Python deployment with no predictor script, which runs the sklearnserver runtime image. The API rejects the setting for those rather than deploying something that cannot archive. !!! note There is no single-instance mode. All instances of a deployment share one pod template, so they either all archive or none do. !!! note Changing disk logging starts a new deployment revision, because it changes the pod template. ## Code ### Step 1: Connect to Hopsworks === "Python" ```python import hopsworks project = hopsworks.login() # get Hopsworks Model Serving handle ms = project.get_model_serving() ``` ### Step 2: Retrieve an existing deployment === "Python" ```python deployment = ms.get_deployment("mydeployment") ``` ### Step 3: Get current deployment state === "Python" ```python state = deployment.get_state() state.describe() ``` ### Step 4: Explore transient logs === "Python" ```python deployment.get_logs(component="predictor|transformer", tail=10) ``` To follow a running deployment instead of taking a single snapshot, use `tail_logs`. It returns a generator that yields new lines as they arrive, skipping what it has already yielded. === "Python" ```python for chunk in deployment.tail_logs(component="predictor"): print(chunk, end="") ``` ### Step 5: Download historical logs === "Python" ```python local_paths = deployment.download_logs(latest=True) for local_path in local_paths: with open(local_path) as archive: print(archive.read()) ``` Omit `latest` to download every archive the deployment has kept. !!! api "API reference" - [`ModelServing.get_deployment`][hsml.model_serving.ModelServing.get_deployment] - [`Deployment`][hsml.deployment.Deployment] - [`get_state`][hsml.deployment.Deployment.get_state] - [`get_logs`][hsml.deployment.Deployment.get_logs] - [`tail_logs`][hsml.deployment.Deployment.tail_logs] - [`download_logs`][hsml.deployment.Deployment.download_logs] - [`PredictorState`][hsml.predictor_state.PredictorState] - [`describe`][hsml.predictor_state.PredictorState.describe] Browse the full Python API :material-arrow-right: ================================================================================ # Analytics Guides Source: https://docs.hopsworks.ai/latest/user_guides/analytics/ # Analytics Guides Analytics is the SQL and dashboard layer over the feature store: Trino queries the offline data, Superset charts it. Start with a query, then turn it into a dashboard.
- :material-chart-box-outline:{ .lg .middle } **Start here** --- Open the Query Engine from the project sidebar and run SQL against a feature group. One catalog per table format, so pick `delta`, `iceberg` or `hudi` first. ```sql SELECT customer_id, avg(amount) AS avg_amount FROM delta.fraud_featurestore.transactions_1 GROUP BY customer_id ORDER BY avg_amount DESC LIMIT 10 ``` [Query engine](../projects/trino/query_engine.md) · [Superset](../projects/superset/superset.md)
:material-database-search-outline:{ .hops-role-ico } Query { .hops-role-cap } - [Query engine](../projects/trino/query_engine.md) Run interactive SQL over feature groups, with query history and cluster status. - [Trino catalogs](../projects/trino/catalogs.md) Expose a data source as a catalog so external tables join feature data. - [External BI tools](../../concepts/mlops/bi_tools.md) How a BI tool connects to the offline store through Trino.
:material-view-dashboard-outline:{ .hops-role-ico } Visualize { .hops-role-cap } - [Superset](../projects/superset/superset.md) SQL Lab, datasets, charts and dashboards inside the project. - [Share a dashboard](../projects/superset/superset.md#sharing-and-collaboration) Give project members access to a dashboard. - [Enable Superset](../../setup_installation/admin/superset.md) Administrator setup, users and roles.
================================================================================ # Query Engine Source: https://docs.hopsworks.ai/latest/user_guides/projects/trino/query_engine/ # Query Engine (Trino) The Query Engine in Hopsworks is powered by Trino, a distributed SQL query engine that allows you to run interactive analytics on your data. Use it to explore feature groups, run ad-hoc queries, and analyze data across your project. ## Accessing the Query Engine Open **Queries** under **Analytics** in your project's left sidebar. The page carries four tabs: **SQL Runner**, **Cluster Overview**, **Catalogs** and **Queries**.
Query Engine
The Query Engine, open on the SQL Runner tab
## SQL Runner The SQL runner is where you write and execute SQL queries against your data. **To run a query:** 1. Pick a **Catalog** and a **Schema** in the **Explorer**. The tables in that schema are then listed below, and you can write the query against bare table names instead of qualifying every one. 2. Write the query in the editor, or click a table to start from `SELECT * FROM
`. 3. Choose a row limit, which is appended to the query as a `LIMIT`. 4. Click **Run query**, or press `Ctrl/Cmd + Enter`. The results appear below the editor, with each column's type under its name. The header reports the state of the run, how long it took, how many rows it processed and how many it returned. The editor auto-completes catalogs, schemas, tables and columns with `Ctrl + Space`, and **Add query** opens a second tab so several queries can be kept side by side.
The SQL runner with a query over the tpch catalog and its results
SQL runner
The expand icon in the results header opens the results in a full-size dialog, with the same run status, so a wide table can be read without scrolling sideways.
The results of a query opened in the full-size results dialog
Results in full
### Statements The SQL runner runs one statement at a time, and a trailing semicolon is removed before the statement is sent. A statement that changes data or the catalog reports what it did instead of a table of results: `CREATE TABLE succeeded`, or for `INSERT`, `UPDATE`, `DELETE`, `MERGE` and `CREATE TABLE ... AS SELECT` the number of rows it wrote, such as `INSERT succeeded (2 rows)`.
A CREATE TABLE statement reported as succeeded
A statement that returns no rows
Every statement runs in a session of its own. Statements that only change the session (`USE`, `SET SESSION`, `RESET SESSION`, `SET ROLE`, `SET PATH`, `SET TIME ZONE`, `PREPARE`, `DEALLOCATE`, and the transaction statements) therefore succeed and change nothing for the next statement, and the SQL runner says so. Choose the catalog and schema in the **Explorer** instead of with `USE`.
A USE statement reported as having no effect
A session statement
### Records End a statement with `\G` instead of `;` to see each row as a record, one column and value per line, the way the Trino CLI prints it. `\G` is not SQL: like a trailing semicolon, it is removed before the statement is sent. Records suit wide rows and long values; the records view shows the first 200 rows, so end the statement with `;` to see more in the grid.
Query results shown as records after ending the statement with \G
Results as records
### Query plans `EXPLAIN` shows the plan Trino would run for a statement. A plain `EXPLAIN` shows the plan as text, the way Trino prints it. `EXPLAIN (FORMAT GRAPHVIZ)` opens on a **Graph** view that draws the operators and the data flowing between them, grouped by fragment, which can be zoomed and moved around; **Text** beside it shows the Graphviz source. `EXPLAIN (FORMAT JSON)` and `EXPLAIN (TYPE IO)` are shown as indented JSON.
The Graph view of an EXPLAIN (FORMAT GRAPHVIZ) plan in the full-size results dialog
A query plan as a graph
### SQL Statement Syntax Help The help icon in the editor opens the syntax of every SQL statement Trino supports, from `ALTER` to `VALUES`, including the clauses of `SELECT` such as time travel with `FOR TIMESTAMP | VERSION AS OF`, `TABLESAMPLE`, `UNNEST` and `JSON_TABLE`. It also lists the keyboard shortcuts and the statements that have no effect in the SQL runner.
The SQL syntax help open on the SELECT statement
SQL statement syntax
## Cluster Overview The cluster overview reports the query engine's version, environment and uptime, then a tile per metric, each with a sparkline of its recent history: - **Running queries**, **Queued queries** and **Blocked Queries** - **Active workers** and **Worker Parallelism** - **Runnable drivers**, **Input Rows/s** and **Input Bytes/s** - **Reserved Memory** Together they say whether the cluster is busy and whether a slow query is competing for capacity.
cluster overview
Query Engine cluster overview
## Managing Catalogs The Catalogs tab is where a project makes an external data source queryable from the SQL runner. Creating a catalog from a data source or by hand, referencing credentials, testing the connection, and when a change reaches the query engine are all covered in [Trino Catalogs][trino-catalogs]. ## Queries The Queries tab lists the queries the project has run, filterable by state and sortable, with the query text alongside each entry. A card carries its id and state with a progress bar, the user and source that submitted it, the resource group, its split counts, wall, total and CPU time, and its reserved, peak and cumulative memory. A query that is still running is listed the same way and updates in place. Click a query id to open its details.
queries
The project’s query history
## Query Details Clicking on a query opens the detailed view with comprehensive execution information. ### Overview The overview tab shows query metadata, execution timeline, and performance metrics including: - Query text - Execution time - Data processed - Rows returned - Resource consumption
Query details
Query details
### Live Plan The live plan draws the query's stages and the operators inside them, with each stage's state and its CPU time, memory, drivers and tasks, updating while the query runs. The graph is usually taller than the page, so it can be laid out **Vertical** or **Horizontal**, panned by dragging, and zoomed by scrolling.
Query details live plan
Query details: live plan
### Stage performance This view takes one stage at a time, chosen with the **Stage** selector, and draws its pipelines operator by operator. Each operator reports its throughput, output rows and bytes, driver count, and CPU, wall and blocked time, which is what locates a bottleneck inside a stage rather than merely between stages.
Query details stages
Query details: stage performance
### Splits Splits show how Trino parallelizes query execution. Each split represents a portion of data processed by a worker. View split-level metrics to understand query parallelism and data distribution.
Query details split
Query details: split
### References The references tab lists the tables the query read, each with the user it was authorized as and whether the query named it directly, and the routines it called.
Query details references
Query details: references
### JSON The JSON view provides the complete query execution plan and statistics in JSON format, useful for programmatic analysis or debugging.
Query details json
Query details: json
## Best Practices - **Limit result sets**: Use `LIMIT` clauses for exploratory queries to reduce resource usage - **Filter early**: Apply `WHERE` clauses to reduce data scanned - **Monitor query performance**: Check the Queries tab to identify slow or failed queries - **Use the live plan**: For complex queries, review the execution plan to optimize performance - **Check cluster status**: Ensure adequate resources are available before running large queries ================================================================================ # Trino Catalogs Source: https://docs.hopsworks.ai/latest/user_guides/projects/trino/catalogs/ # Trino Catalogs A Trino catalog makes an external data source queryable from the query engine. Each catalog names a Trino connector and the properties that connector needs to reach the source, such as a connection URL and credentials. Once a catalog is live, its databases and tables can be queried from the SQL runner alongside your feature groups. Navigate to **Query Engine** → **Catalogs** in your project to see the project's catalogs and your own private catalogs. The cluster's shared default catalogs are listed once you add **Default** to the **Type** filter. A catalog belongs either to the project or to you: - A **project catalog** is named `__` and is queryable inside the project. The project's Data Owners create, edit, delete and share it. - A **private catalog** is named `___` and follows you into every project you are a member of. Only you edit, delete or share it. See [Private catalogs][private-catalogs]. Only a Data Owner of the current project can create a catalog of either kind. A catalog is queryable outside its own project, or outside your projects for a private one, only once it is shared, as described in [Sharing Catalogs and Feature Groups][sharing-catalogs-and-feature-groups].
Catalogs list
The project's catalogs, a private catalog, and the cluster's shared default catalogs
A catalog change is recorded immediately, but it reaches the query engine only when the engine restarts, because Trino reads catalogs at startup. [When the catalog goes live][when-the-catalog-goes-live] covers when that happens and who can bring it forward. ## Creating a catalog from a data source The recommended way to create a catalog is from a data source, because the data source already holds everything the catalog needs: the host, the database, and the credential. Hopsworks derives the connector type and properties for you, so you do not need to know the property names of the underlying Trino connector. When you browse a data source's databases and tables to create an external feature group or ingest data, there are two ways in: - A standalone **Add Trino Catalog** button on the configuration screen, which creates the catalog on its own, without creating or ingesting any feature group. - An **Add Trino Catalog** checkbox on the review step, checked by default, which creates the catalog as the first step of setting up the feature groups. Both are disabled with the reason when the data source cannot be mapped to a Trino connector, and show an **Already added** state when a catalog for this source already exists. The catalog is created first, so you can review its properties before the feature groups and their ingestion are set up. The **Create Trino Catalog** dialog opens pre-filled with a suggested name, the connector type derived from the data source, and the connector properties derived from its settings. You can edit them and add any property the connector supports that the data source does not carry. ### Credentials are references, not copies Credential properties are pre-filled as [references][referencing-a-credential-instead-of-typing-it] rather than values. The credential is read from the data source server-side and stored as a secret owned by you, or as a bundle of files where the credential is a file, and the catalog keeps only the reference. An Oracle data source that authenticates with a wallet is delivered the second way: the bundle is built from the data source's own wallet, and the catalog's `TNS_ADMIN` property points at it. No credential is ever sent to the browser, and rotating one stays a single operation on the secret rather than an edit of every catalog that uses it. ## Creating a catalog by hand A source that has no Hopsworks data source is added by hand. Click **Create catalog**, give the catalog a short name, choose the connector, and write the properties one `key=value` per line. The `__` prefix is added for you.
Create catalog
Creating a catalog by hand, with the name prefix added automatically
The picker offers every connector installed in the query engine, so a connector it does not list cannot be used: `bigquery`, `cassandra`, `clickhouse`, `delta_lake`, `druid`, `duckdb`, `elasticsearch`, `exasol`, `faker`, `gsheets`, `hive`, `hudi`, `iceberg`, `ignite`, `kafka`, `lakehouse`, `loki`, `mariadb`, `mongodb`, `mysql`, `opensearch`, `oracle`, `pinot`, `postgresql`, `prometheus`, `redis`, `redshift`, `singlestore`, `snowflake`, `sqlserver`, `trino_thrift`. Connectors that expose no external data source are rejected, because a catalog on one would either read the query engine's own internals or hold nothing: `system`, `jmx`, `memory`, `blackhole`, `datasketches`, `ai`, and the `tpch` and `tpcds` sample generators. The last two are already available to every project as shared catalogs. Properties must address the data source over the network, for example `jdbc:`, `thrift:`, `https:` or `s3:`. A property that points at a file path on the query engine's own machines is rejected, because you cannot place files there and the only files such a path could reach belong to the cluster itself. Two limits apply, because every project's catalogs share a fixed amount of storage in the cluster. A project may create a set number of catalogs, ten by default, and a single definition may not exceed a set size, 16 KiB by default, measured after any references are resolved. Both are cluster settings an administrator can raise, and neither is close to what an ordinary catalog needs: a few hundred bytes is typical. Property values must also be latin1 text, which is what the definition is stored as, so a credential containing characters outside it has to come from a Hopsworks secret rather than being typed into a property. ### Referencing a credential instead of typing it Two reference forms keep a credential out of the stored catalog definition. Type `${` in the properties editor to pick from either. - `${HOPSWORKS_SECRET:}` for a value, such as a password. - `${HOPSWORKS_MOUNT:}` or `${HOPSWORKS_MOUNT:/}` for a file, such as an Oracle wallet or a Java keystore. See [Mountable Secrets][mountable-secrets]. A secret reference resolves against the secrets of the person who created the catalog, so you can only reference your own. Naming a colleague's secret does not work, even if you can both see the catalog. Two consequences follow. A referenced secret cannot be deleted while a catalog still uses it, and the catalog keeps working after you leave the project only if the secret still exists. If a catalog needs to outlive your account, have someone recreate it under theirs. !!! warning "Only secrets created from typed text can be referenced" A secret created by uploading a file holds the base64 encoding of that file's contents, not the contents themselves. A reference to such a secret puts that base64 text into the catalog, and the connector then fails, because it receives an encoded string where it expects a password, a key, or a JSON document. Nothing records how a secret was created, so Hopsworks cannot detect the case and decode it for you, and the resulting error comes from the connector rather than from Hopsworks. Create the secret by typing or pasting the value as text when you intend to reference it from a catalog. For a credential that is naturally a file, use a mountable secret instead. ## Private catalogs Choose **Private** as the owner when creating a catalog to make it yours rather than the project's. The name prefix becomes `___`, for example `_meb10000__sales`, and the catalog is listed in every project you are a member of.
Creating a private catalog
A private catalog takes your username as its prefix and can reference only your own secrets
A private catalog differs from a project catalog in what it can reach and who controls it: - It can reference only your own Hopsworks secrets and your own mountable secrets, which you manage under **Account Settings** → **Secrets**. A project's mountable secrets are not available to it, because the catalog would carry them into every other project you are a member of. See [Mountable secrets for private catalogs][mountable-secrets-for-private-catalogs]. - It cannot be created from a data source, because a data source's credential belongs to the project. Enter the connection details yourself instead. - Only you can edit, delete or share it, from any of your projects and whatever your role there, so being made a Data Scientist in a project does not lock you out of your own catalogs. Other members of your projects cannot query it unless you share it with their project. - You write to it only from a project where you are a Data Owner, and read it from any other. In a project where you are a Data Scientist you can only read that project's data, and a private catalog writable from there would let you copy the data into a catalog you then read from your other projects. - The number of private catalogs you can own has the same limit as a project's catalogs, ten by default. When your account is deleted, your private catalogs are marked for removal and stop being queryable at once, and their shares are removed with them. They stay denied to everyone, including a later account with the same username, until they are removed like any deleted catalog, as described in [When the catalog goes live][when-the-catalog-goes-live]. ## Testing the connection **Create** stays disabled until **Test connection** succeeds, so a catalog that cannot reach its source is caught now rather than after a restart. The test creates a temporary catalog on the cluster's test coordinator, lists its schemas, and reports the result. A failed test shows the query engine's own error, which is what says how to fix the properties.
Test connection failed
A failed test reports the error the query engine saw, and Create stays disabled
Test connection succeeded
Create is enabled once the connection test passes
On a cluster where the administrator has turned the test coordinator off, the test reports that testing is unavailable and **Create** is enabled without it. ## When the catalog goes live A new catalog is saved with the status **Pending approval**. It does not appear as a target in the SQL runner until the query engine restarts and loads it. The **Catalog created** dialog says when that happens, and what it says depends on whether you can restart the query engine yourself. A project user is told the next scheduled restart, in their own timezone.
Catalog pending approval
A project user is told when the scheduled restart will make the catalog queryable
A cluster administrator is offered **Restart query engine now** instead, together with the number of queries currently running and queued. A restart cancels those queries everywhere on the cluster, so the dialog reports the activity to let the administrator judge whether restarting now is safe.
Catalog created, as a cluster administrator
A cluster administrator can apply the change immediately
Two cluster settings change this. Where the administrator has enabled an eager restart, the dialog also says the catalog may go live earlier than the scheduled time, because the query engine restarts as soon as it is idle while changes are pending. Where the administrator has instead required approval for all catalog changes, there is no schedule at all, and the dialog says to ask an administrator. Both are described in the [administrator guide][lifecycle-settings]. Deleting a catalog follows the same path in reverse. The catalog is marked **Pending removal** immediately, and stops being queryable at the next restart, because a running query engine keeps the catalogs it started with. A catalog you have deleted can therefore still answer queries for a while. Editing one works the same way: the stored definition changes at once, and the loaded catalog changes at the next restart. Secret-bearing property values come back masked, and a masked value left unchanged keeps the stored secret, so you do not retype credentials to change an unrelated property. The Catalogs tab shows where each catalog stands. | Status | Meaning | | --- | --- | | Approved | Loaded by the query engine and queryable. | | Pending approval | Saved, not yet written for the query engine. | | Pending restart | Written, waiting for the restart that loads it. | | Pending removal | Deleted, still loaded until the next restart. | | Failed | The query engine could not load it. See below. | ## When a catalog fails to load A catalog can be valid to save and still be rejected by the query engine, for example when a property name is not one the connector accepts. The query engine reads catalogs only at startup and refuses to start if it cannot load one, so such a catalog is removed from the engine automatically and marked **Failed**, which keeps the query engine available for everyone. The status carries the error the query engine reported, which says what to correct. Edit the catalog to fix the definition, and it returns to Pending approval and follows the normal flow again. Testing the connection before saving catches most of these earlier.
Failed catalog
A catalog the query engine could not load, with the reason it reported
## Who can query a catalog Inside the project that owns a project catalog, its Data Owners can read and write, and its Data Scientists can read. A private catalog can be read by you from any of your projects, and written only from a project where you are a Data Owner. Other projects can read a catalog only through a share, which grants read access to the whole catalog, one schema, one table, or some columns of a table, optionally with masked values. See [Sharing Catalogs and Feature Groups][sharing-catalogs-and-feature-groups]. The query engine reads the external system as the database user in the connection credentials, so no share can expose more than those credentials allow. Scoping that database user at the source remains the strongest limit on what a catalog can reach. A Data Owner of the project can also run a JDBC catalog's `system.query` table function, which passes a query to the source database as that database user. The query engine sends it as a subquery, so the database rejects a statement that changes data, but a database function that changes data as a side effect still runs. Give the database user only the privileges the catalog's Data Owners should have. The receiving project of a share cannot run the catalog's functions at all. Read access to an Iceberg or Delta Lake table also allows its table procedures, `ALTER TABLE ... EXECUTE` with `optimize`, `expire_snapshots`, `remove_orphan_files` or `rollback_to_snapshot`, because the query engine does not check them against the access rules. A Data Scientist of the project can therefore rewrite, expire or roll back the tables of the project's Iceberg and Delta Lake catalogs, and so can you on your private catalog from a project where you are a Data Scientist, although neither can write rows. Set the connector's `iceberg.security` or `delta.security` property to `read_only` on a catalog that must not be changed this way; the query engine then refuses table procedures to everyone and reading is unchanged. `CALL` procedures, such as `system.unregister_table`, are refused to everyone. ## Creating a catalog from the Python client The same operations are available from the Python client, so making a data source queryable can be scripted from a job or a notebook. ```python import hopsworks project = hopsworks.login() fs = project.get_feature_store() sc = fs.get_data_source("my_snowflake_db").storage_connector template = sc.get_trino_catalog_template() if template["supported"]: catalog = sc.create_trino_catalog() print(catalog["name"], catalog["status"]) else: print("cannot be mapped:", template["reason"]) ``` `get_trino_catalog_template()` returns the proposal without creating anything, so it can be reviewed or adjusted first. `create_trino_catalog()` accepts an optional `name` and a `properties` override, and tests the connection before creating where the cluster supports it, so an unreachable source raises before the catalog exists. Credentials are resolved server-side into references exactly as in the UI path. The full catalog lifecycle is available on the project: ```python import hopsworks project = hopsworks.login() catalogs = project.get_trino_catalog_api() for catalog in catalogs.get_catalogs(): print(catalog["name"], catalog["status"]) print(catalogs.get_capabilities()["nextScheduledRestart"]) ``` A cluster administrator can apply every pending catalog change immediately instead of waiting for the schedule: ```python import hopsworks project = hopsworks.login() catalogs = project.get_trino_catalog_api() result = catalogs.restart() print(result["restarted"], result["quarantined"]) ``` The restart interrupts queries running anywhere on the cluster, so prefer waiting for the scheduled restart unless the change is needed sooner. ================================================================================ # Sharing Source: https://docs.hopsworks.ai/latest/user_guides/projects/trino/sharing/ # Sharing Catalogs and Feature Groups The query engine enforces who can read what, so data one project owns is not visible to another project until it is shared. Two kinds of share reach the query engine: - A **catalog share** gives another project read access to one of your Trino catalogs: the whole catalog, or chosen schemas and tables, down to some columns of a table, optionally with masked values. - A **feature group share**, made from the feature store, also makes the shared feature group queryable through the query engine in the receiving project. A share always grants read access, and always to the receiving project's Data Owners and Data Scientists. It grants tables only: the receiving project cannot run the catalog's functions, such as `system.query` of a JDBC catalog, which would run any SQL on the source database as the catalog's database user. Changes take effect within seconds, without restarting the query engine. ## Sharing a catalog A project catalog is shared by a Data Owner of the project, and a private catalog by its owner, from any of their projects. A catalog can be shared once it is **Approved** and while it is not being deleted, because a catalog the query engine has not loaded has nothing to share yet. It cannot be shared with the project that owns it. While a catalog is shared, its connector cannot change, because a narrowed share denies the hidden columns of the connector the query engine runs, and the engine switches connector only at its next restart. Revoke its shares first; its other properties can be edited as usual. Click the share icon on the catalog's row in **Query Engine** → **Catalogs** to open its sharing page. The page lists every project the catalog is shared with, what it shares with each, and whether each share is live.
Sharing page of a catalog
A catalog shared whole with one project, and one table with two of its three columns, one of them masked, with another
Click **Share** and choose the project. A project holds one share of a catalog, which covers everything it receives from the catalog, so a project the catalog is shared with already is not offered: edit its share instead. Then check what to share in the tree of the catalog: - Check the catalog to share every schema and table in it, including ones created later. - Check a schema to share every table in it, including tables created in it later. - Expand a schema and check some of its tables to share only those tables. - Expand a table and uncheck some of its columns to share only the checked columns. A checked column can carry a mask. Unchecking something inside a checked schema or catalog keeps the rest of it: uncheck `sales.salaries` in a checked schema `sales`, and every other table of `sales` stays shared. A partly checked box shares only what is checked under it, so a table created later in a partly checked schema is not shared. The panel beside the tree shows what the receiving project will see. Click a table name to see its first rows as they will read them, with unchecked columns left out and masks applied. The sample is read as you, before anything is saved.
Sharing one table with some columns
Sharing one table, with one column left out and one column masked
The schemas, tables and columns offered are the ones you can see in the catalog yourself, read through the catalog's own connection. ### Sharing some columns of a table Expand a table in the tree to choose its columns. An unchecked column cannot be read by the receiving project, and a query that selects it, or selects `*`, is refused. The table's other ways of revealing a column are closed too: the connector's hidden columns, such as `$path` and `$partition`, or Elasticsearch's `_source`, which holds the whole document, and the table's metadata tables, such as `
$partitions`, are denied on a narrowed table. Tables of a Kafka, Redis, MongoDB, Cassandra or Thrift catalog cannot be narrowed to some columns, because those connectors can hide columns that are defined outside the query engine and cannot all be denied. Share such a table whole, or leave it out. The receiving project reads only the columns you checked. Hopsworks reads the table's columns each time it updates the rules, and at least every five minutes by default, and denies every column you did not check, so a column added to the table later, or renamed at the source, is not shared. Between the change at the source and the next update, the new column is readable. A renamed column loses its mask with its old name and is denied under the new one. The sharing page marks such a share with the number of new columns and names them; edit the share and check them to share them. While the query engine is restarting, a narrowed table the rules already held keeps what they denied before, as well as the columns recorded when the share was saved, until it is back. A narrowed table the rules did not hold yet, because its share is new or was wider, or because it was left out before, stays out until the query engine is back, and a narrowed table whose columns cannot be read is left out of the share until they can. Iceberg and Delta Lake tables can also be read as of an earlier version, with `FOR VERSION AS OF` or `FOR TIMESTAMP AS OF`, and such a read has the columns the table had then. Hopsworks denies those columns too: every column the table has had at a version that can still be read, so a column dropped or renamed at the source stays denied under its old name. This covers Iceberg and Delta Lake catalogs, and the Iceberg and Delta Lake tables of a Lakehouse catalog. Hive and Hudi tables cannot be read as of an earlier version. An Iceberg table whose metadata file is outside HopsFS, larger than 64 MiB, or not readable by the catalog's owner has its earlier columns read one snapshot at a time, 50 snapshots per update, and is left out of the share until all of them have been read; the Hopsworks log names the table.
Editing the columns of a share
Editing a share: a table with two columns shared, one of them masked, and one left out
### Masking a column A shared column can carry a mask, which replaces the value the receiving project reads. A mask is one SQL expression over the row, for example `'***'` or `regexp_replace(email, '.+@', '***@')`, and it must return the column's own type. A mask runs as the person querying, not as you, so it can use only what they can read: the checked columns of the same table. A mask that refers to an unchecked column is refused when you save the share. A mask that reads another table is accepted, because it is checked as you, but it fails at query time for anyone in the receiving project who cannot read that table. It cannot contain `;` or a comment, and it is limited to 2000 characters. The mask is checked against the table when the share is saved, so an expression the query engine cannot evaluate is reported then rather than when someone queries the table.
Querying a masked column
The receiving project reads the masked column as ***
You always read your own catalog unmasked, and a share never narrows your own access: in a project a private catalog is shared with, its owner keeps the access described in [Private catalogs][private-catalogs]. ### Editing a share Click the edit icon on a share to change what it covers: add or remove schemas and tables, narrow tables to some columns, or go from chosen schemas and tables to the whole catalog. Saving reads the narrowed tables' columns again, and the change is live within seconds. ### The status of a share A share is live once the query engine has loaded the rules that include it, which takes a few seconds. The sharing page shows where each share stands and refreshes itself while a change is being applied. | Status | Meaning | | --- | --- | | Applying | Saved, and being made live. | | Active | Live: the receiving project can read what it covers. | | Revoking | Being removed. The receiving project loses access when this finishes. | | Failed | Could not be made live. The status carries the reason; edit the share and save it to retry. | ### Objects that no longer exist A share names schemas and tables, and an object it names may be dropped at the source later. The share keeps it, and the sharing page lists it as **Not found**. If an object with the same name is created again, the share applies to it. To remove an object that is gone for good, click **Remove them from the share**, or edit the share and uncheck the object, which the tree shows as not found at the source. ### Revoking a share Click the delete icon on a share to revoke it. The receiving project loses access to everything the share covered as soon as the revoke is live, within seconds, and the share disappears from the page once it is. The catalog cannot be shared with the same project again until the revoke has finished. ### What a share does not restrict A share grants reading, but the query engine does not check table procedures against it. Anyone in the receiving project can run `ALTER TABLE ... EXECUTE` on a shared table of an Iceberg or Delta Lake catalog, for example `optimize`, `expire_snapshots` or `rollback_to_snapshot`, and change the table with the catalog's own credentials. `rollback_to_snapshot` returns the table to an earlier version and drops every write made since. A table with a masked column is the exception: the query engine refuses table procedures on it. Share such a catalog only with projects you trust with its tables, or set the connector's own `iceberg.security` or `delta.security` property to `read_only` on the catalog, which refuses table procedures to everyone, you included, and leaves reading unchanged. Deleting a catalog revokes all of its shares at once, and so does deleting the receiving project. ### Shares your project received The **Catalogs** tab lists, under **Shared with this project**, every catalog share your project received, who it comes from, and what it covers. It shows how many columns of a table are masked, but not the mask expressions, which can hold values such as a salt. A shared catalog is queried by its own name, like any other catalog, with the SQL runner or any Trino client.
Shares received by a project
A project catalog shared whole, and one table of another user's private catalog
## Sharing feature groups through the query engine A feature group shared with another project, as described in [Sharing a feature group with selected features][sharing-a-feature-group-with-selected-features], is also queryable by that project through the query engine. Where it is read from depends on whether the whole feature group was shared or a subset of its features. | Shared | Catalog | Readable | | --- | --- | --- | | The whole feature group, or the whole feature store | The catalog named after the feature group's format: `delta`, `hudi`, `iceberg` or `hive` | Every feature | | A subset of the features | The shared catalog of its format: `delta_shared`, `hudi_shared` or `iceberg_shared` | The shared features only | The schema is the owning project's feature store, `_featurestore`, and the table is `_`. The share dialog says which catalog the other project will read from, or that the share is not queryable through the query engine.
Sharing a subset of a feature group
A subset of the features is read through delta_shared
### A subset of the features A subset share is read through the shared catalog of the feature group's format. The primary key and the event time are always shared, and the receiving project reads the other shared features alongside them. ```sql SELECT datetime, cc_num, category, amount FROM delta_shared.fraud_featurestore.transactions_1; ```
Querying a subset share
Only the shared features are listed, and the preview selects them by name
Everything else in the table is denied: the unshared features, the connector's hidden columns such as `$path`, and the table's metadata tables such as `transactions_1$history`. A query that selects an unshared feature, or selects `*`, is refused.
Querying an unshared feature
Selecting a feature that was not shared is refused
The shared catalogs are visible only to projects that received a subset share, and a project sees only the feature groups shared with it there. A subset share is not readable through the catalog named after the format, and a whole share is not readable through the shared catalog. Adding features to a shared feature group does not widen the share: a new feature stays unreadable by the receiving project until you share it. The access rules are updated before the request that adds the features returns. A column that reaches the table some other way, such as a Delta write that merges a new column into the schema, is denied too, once Hopsworks next reads the table's columns: at least every five minutes by default, and on every share change. If the table's columns cannot be read, the share is left out of the rules until they can. The columns an Iceberg or Delta Lake feature group had at an earlier version are denied as well, as described in [Sharing some columns of a table][sharing-some-columns-of-a-table]. Unsharing the feature group, or deleting it, removes access within seconds. ### Feature groups that are not queryable Two kinds of feature group cannot be read through the query engine by the receiving project: - A feature group on an external data source, whether shared whole or in part, because its data does not live in the feature store's tables. - A subset of a feature group stored as a plain Hive table, without Delta, Hudi or Iceberg, because there is no shared catalog for that format. Such a feature group shared whole is read through the `hive` catalog. The share dialog says so for these feature groups, and the feature group remains shared and readable through the feature store APIs. ================================================================================ # Superset Source: https://docs.hopsworks.ai/latest/user_guides/projects/superset/superset/ # Data exploration and data visualization (Superset) Apache Superset is a modern data exploration and visualization platform integrated with Hopsworks. It provides an intuitive interface for creating interactive dashboards, running SQL queries, and building advanced visualizations on top of your feature data. ## Prerequisites To use Superset, you need: - Membership in a Hopsworks project with Superset enabled - Appropriate project role permissions (Data Owner or Data Scientist) ## Accessing Superset Access Superset from your Hopsworks project: 1. Navigate to your project in Hopsworks 2. In the left sidebar, locate the **Analytics** section 3. Click on **Superset**
Superset in project menu
Accessing Superset from a Hopsworks project
This opens the Superset dashboards page, which lists all dashboards available in your project.
Superset dashboards page
Superset dashboards page, with per-dashboard public/shared status and actions
To open a specific dashboard, click its name in the list. The dashboard opens in a new browser tab. To open the full Superset application, click **Open Superset**. Superset opens in a new browser tab. ## Superset Interface The Superset home page provides access to all major features:
Superset home page
Superset home page
### Main Navigation - **Home**: Landing page with recent activity and quick access - **Dashboards**: View and create interactive dashboards - **Charts**: Individual visualization components - **Datasets**: Configure and manage data sources from Hopsworks - **SQL**: Access SQL Lab for ad-hoc queries ### Recent Activity The home page displays your recent activity organized by: - **Viewed**: Recently accessed dashboards and charts - **Edited**: Items you have recently modified - **Created**: Your recently created content This provides quick access to your most frequently used visualizations. ## SQL Lab SQL Lab is an interactive SQL query interface for exploring your feature data. ### Opening SQL Lab 1. Click **SQL Lab** in the top navigation 2. Select your database from the **Database** dropdown 3. For the Trino connection, select the catalog matching your table format (`hive`, `delta`, `iceberg`, or `hudi`) from the **Catalog** dropdown 4. Select your project's Feature Store schema from the **Schema** dropdown
SQL Lab interface
SQL Lab for running queries
For the Trino connection, SQL Lab shows a **Catalog** dropdown between **Database** and **Schema**. The dropdown appears because the connection has *Allow changing catalogs* enabled, and it lists every Trino catalog you can read (`hive`, `delta`, `iceberg`, `hudi`, your project's catalogs, and catalogs shared with your project). Selecting the catalog that matches your table format lets you query the table without prefixing the catalog in the SQL.
Trino catalog selector in SQL Lab
Selecting a Trino catalog from the Catalog dropdown in SQL Lab
### Running Queries To execute a query: 1. Type your SQL query in the editor 2. Click **Run** or press `Ctrl+Enter` (Windows/Linux) or `Cmd+Enter` (Mac) 3. View results in the **Results** tab below the editor 4. Check the **Query History** tab to see previous queries ### Best Practices for SQL Lab - **Limit result sets**: Use `LIMIT` clauses to avoid retrieving excessive data - **Test queries incrementally**: Start with small queries and build up complexity - **Save important queries**: Use the save feature for queries you'll reuse - **Use proper formatting**: Format SQL for readability and maintainability ## Creating Datasets Datasets define the data sources for your charts and dashboards. They typically map to Feature Groups in your Hopsworks project. ### Adding a Dataset 1. Navigate to **Datasets** 2. Click **+ Dataset** in the top right 3. Select your database (e.g., `Trino_____superset`) 4. For the Trino connection, select the catalog matching your table format (`hive`, `delta`, `iceberg`, or `hudi`) from the **Catalog** dropdown 5. Select your schema (your project's Feature Store schema) 6. Select the table (your Feature Group or Feature View) 7. Click **Create Dataset and Create Chart** !!! note "Using Different Table Formats" The Trino connection has *Allow changing catalogs* enabled, so you switch table formats by picking the catalog from the **Catalog** dropdown rather than hardcoding it in the table name. Each table format has its own catalog: `delta`, `iceberg`, `hudi`, and `hive`. Select the catalog that matches your tables, then choose the schema and table as usual. If you prefer to define the dataset from SQL instead, click **Create dataset from SQL query** and prefix the table with the matching catalog: - For Delta format: `SELECT * FROM delta._featurestore.` - For Iceberg format: `SELECT * FROM iceberg._featurestore.` - For Hudi format: `SELECT * FROM hudi._featurestore.` If you're unsure which format your tables use, contact your Hopsworks administrator. ## Creating Charts Charts are individual visualization components that can be added to dashboards. ### Creating a New Chart 1. Navigate to **Charts** 2. Click **+ Chart** 3. Select a dataset 4. Choose a visualization type (e.g., Bar Chart, Line Chart, Pie Chart, Table) 5. Click **Create new chart** ### Configuring Charts The chart editor provides: - **Data tab**: Configure metrics, dimensions, and filters - **Customize tab**: Adjust colors, labels, and styling - **Preview**: Real-time preview of your chart - **Save**: Save the chart for use in dashboards ## Creating Dashboards Dashboards combine multiple charts into interactive visualization reports. ### Creating a New Dashboard 1. Navigate to **Dashboards** 2. Click **+ Dashboard** 3. Enter a dashboard name 4. Click **Save** ### Adding Charts to Dashboards 1. Open your dashboard in edit mode 2. Click **Edit dashboard** 3. Drag and drop charts from the chart list onto the dashboard 4. Resize and arrange charts as needed 5. Click **Save** to publish the dashboard ## Sharing and Collaboration ### Sharing Dashboards By default, a dashboard is only accessible to its owner. The project's **Superset** page in Hopsworks lists every dashboard with two visibility columns and per-row actions for changing who can see it: - **public** column: `not public`, `project public` (visible to every member of the current project), or `public` (visible to unauthenticated viewers). - **shared** column: `Not shared`, or `Shared` (hover to see which other projects the dashboard is shared with). The per-row action icons, from left to right, are: share with other projects, share with all project members, make public, copy permalink, and delete. These sharing and visibility actions, along with delete, are available only to the dashboard's owner. Other users can still open a dashboard and copy its permalink. All sharing actions require the dashboard to be **published** in Superset first. A draft is invisible to the roles being granted access, so the sharing icons stay disabled until you publish it. You can also share with specific users or roles from Superset's native dashboard **Access** dialog, as described next. #### Adding Users to Share a Dashboard To share a dashboard with specific users or roles: 1. Navigate to **Dashboards** and locate your dashboard 2. Click the **edit icon** menu next to the dashboard name 3. In the **Access** section, you can grant access in two ways: - **Owners**: Add individual users to the owners list to grant them full access - **Roles**: Add a role that users are members of to grant access to all users with that role Users added as owners or through roles can view and interact with the dashboard. #### Sharing with All Project Members To make a dashboard visible to every member of the current project, click the **share with all project members** icon on the dashboard's row. Hopsworks grants the project's role access to the dashboard, and the **public** column changes to `project public`. Click the icon again to remove access; the column returns to `not public`. Project members can then open the dashboard from their own **Superset** page or through a [permalink][sharing-dashboard-permalinks]. #### Sharing with Other Projects You can grant members of other projects access to a dashboard without making it public: 1. Click the **share with other projects** icon on the dashboard's row 2. In the **Share with other projects** dialog, find the target project by name 3. Tick a project to share the dashboard with all of its members; untick it to stop sharing The **shared** column shows `Shared` once at least one other project has access. Hover over it to see the list of projects. You do not need to be a member of a project to share a dashboard with it. Sharing attaches that project's role to the dashboard, so its members can open it from their own **Superset** page. #### Making Dashboards Public To make a dashboard accessible to unauthenticated viewers (anyone with the link, no Superset login), use the **make public** icon on the project's Superset dashboards page in Hopsworks: 1. Open the **Superset** page in your project 2. Locate the dashboard in the list 3. Click the **make public** icon for that dashboard The dashboard is marked **public** and becomes viewable without a Superset login. Click the icon again to remove public access. The icon is enabled only when Superset's built-in **Public** role exists. If the role is missing, the icon is disabled and its tooltip reads `The Superset Public role does not exist`, and you should contact your Hopsworks administrator. You can also add or remove the Public role from Superset's native dashboard **Access** dialog.
Adding public role to dashboard
Adding the Public role to make a dashboard accessible to unauthenticated users
!!! warning "Security Considerations" Only make a dashboard public when it contains non-sensitive data. Ensure the dashboard does not expose confidential or project-specific information. ##### Anonymous access requirements Marking a dashboard public requires only the Public role to exist. For unauthenticated viewers to actually open public dashboards and permalinks, the Superset deployment must satisfy two further conditions. The dashboards page shows a warning that lists any condition that is not met: - The built-in **Public** role exists. Without it, dashboards cannot be made public. - Unauthenticated access to the dashboard list is enabled. This requires `AUTH_ROLE_PUBLIC = "Public"` in `superset_config.py`. - The Public role has the required permissions. Running `superset init` syncs the permissions, and the Public role must also have dataset access. If the warning appears, contact your Hopsworks administrator to complete the configuration. See the Apache Superset [public role documentation](https://superset.apache.org/admin-docs/security/#public) for background. ### Sharing Dashboard Permalinks A permalink is a direct link to a dashboard that preserves the current view state, including applied filters and parameter values. From the project's Superset dashboards page in Hopsworks, click the copy-permalink action on a dashboard to copy its permalink to your clipboard. You can then share the copied URL with other users. To create a permalink from within Superset instead: 1. Open the dashboard you want to share 2. Apply any filters or configure the view as desired 3. Click the **Share** button (share icon) in the top right of the dashboard 4. Select **Copy permalink to clipboard** 5. Share the copied URL with other users The permalink captures: - Current dashboard state and selected filters - Active parameter values - Visible chart configurations - Dashboard layout and zoom level !!! note "Permission Requirements" **For project members:** Users must have access to the dashboard (either as owners or through roles) to view it via permalink. **For unauthenticated users:** To share permalinks with anonymous or unauthenticated users, the dashboard must have the Public role assigned (see [Making Dashboards Public][making-dashboards-public] above). If the Public role option is not available, contact your Hopsworks administrator to enable it. ### Exporting Dashboards You can export dashboards for reporting or archival purposes: 1. Open the dashboard 2. Click the **...** menu in the top right 3. Select **Download** to export as PDF or image 4. Choose your preferred format and resolution 5. Save the exported file Exported dashboards capture the current state of all charts and visualizations, making them suitable for reports, presentations, or documentation. ## Working with Feature Store Data ### Querying Feature Groups Feature Groups are available as tables in SQL Lab: ```sql SELECT feature1, feature2, COUNT(*) as count FROM delta._featurestore.my_feature_group GROUP BY feature1, feature2 ORDER BY count DESC LIMIT 100 ``` ### Joining Multiple Feature Groups Combine data from multiple Feature Groups: ```sql SELECT fg1.user_id, fg1.user_features, fg2.transaction_features FROM delta._featurestore.user_features fg1 JOIN delta._featurestore.transaction_features fg2 ON fg1.user_id = fg2.user_id WHERE fg1.created_date >= CURRENT_DATE - INTERVAL '30' DAY LIMIT 100 ``` ## Performance Tips ### Query Optimization - **Use selective filters**: Filter data as early as possible in queries - **Limit result sets**: Always use `LIMIT` to prevent excessive data transfers - **Use partition pruning**: Filter on partition columns when possible to reduce the amount of data scanned - **Avoid `SELECT *`**: Select only the columns you need - **Use statistics and compaction where applicable**: Keep table metadata and file layouts optimized for faster queries ### Chart Performance - **Reduce data points**: Use aggregation to reduce the number of data points - **Enable caching**: Cache chart results for frequently accessed dashboards - **Set appropriate refresh intervals**: Don't refresh more frequently than necessary - **Use time-based filters**: Limit data to relevant time periods ## Troubleshooting ### Common Issues **Cannot see my Feature Groups:** - Verify you are connected to the correct database and schema - Ensure you have appropriate project permissions - Refresh the schema list if Feature Groups were recently created **Queries timing out:** - Reduce the amount of data being queried using filters - Add `LIMIT` clauses to restrict result set size - Break complex queries into smaller steps - Contact your administrator about timeout settings **Charts not loading:** - Check that the underlying dataset is still available - Verify you have permissions to access the data source - Refresh the page or clear browser cache - Check query filters for errors **Dashboard filters not working:** - Ensure filters are properly connected to charts - Verify filter types match the chart data types - Check for filter conflicts or invalid filter combinations ### Trino table type error ```text trino error: TrinoExternalError(type=EXTERNAL, name=UNSUPPORTED_TABLE_TYPE, message="Cannot query Delta Lake table '_featurestore.'", query_id=20260414_155131_00049_qgyn2) This may be triggered by: Issue 1002 - The database returned an unexpected error. ``` - This error occurs when the selected catalog does not match your table's format (Delta, Iceberg, or Hudi) - Solution: Select the catalog matching your table format (`hive`, `delta`, `iceberg`, or `hudi`) from the **Catalog** dropdown in SQL Lab or when adding a dataset. The Trino connection has *Allow changing catalogs* enabled, so the dropdown lets you switch catalogs without editing the query - Alternatively, prefix the table with the matching catalog in your query (e.g., `delta._featurestore.`) - See [Using Different Table Formats][adding-a-dataset] for detailed instructions on creating datasets with the correct catalog - If you're unsure which catalog to use, contact your Hopsworks administrator ## Best Practices ### Organizing Content - **Use descriptive names**: Name dashboards and charts clearly - **Tag and categorize**: Use tags to organize related content - **Create folders**: Group related dashboards together - **Document dashboards**: Add descriptions explaining the purpose and key metrics ### Performance - **Keep queries simple**: Complex queries should be tested in SQL Lab first - **Use appropriate chart types**: Choose visualizations that match your data - **Limit dashboard size**: Too many charts can slow down loading - **Schedule refreshes**: Use automated refresh during off-peak hours ### Data Quality - **Validate queries**: Test queries thoroughly before creating charts - **Handle null values**: Account for missing data in visualizations - **Use appropriate aggregations**: Choose meaningful aggregation functions - **Document data sources**: Note which Feature Groups are used in each dashboard ### Security - **Respect data privacy**: Only share dashboards with authorized project members - **Avoid sensitive data in names**: Don't include sensitive information in dashboard titles - **Regular reviews**: Periodically review and clean up unused dashboards - **Follow project policies**: Adhere to your organization's data governance policies ================================================================================ # Model Registry & Serving Guides Source: https://docs.hopsworks.ai/latest/user_guides/mlops/ # Model Registry & Serving Guides A model goes from training into the registry, out through a deployment, and stays under monitoring. The guides below follow that path.
- :material-package-variant-closed:{ .lg .middle } **Start here** --- Register a trained model with its metrics, then deploy it in one call. ```python mr = project.get_model_registry() model = mr.python.create_model( name="fraud_detector", metrics={{"f1": 0.92}}, ) model.save("model_dir") deployment = model.deploy() ``` [Model registry](registry/index.md) · [Deployment creation](serving/deployment.md) · [Model monitoring](model_monitoring/index.md)
:material-package-variant:{ .hops-role-ico } Register { .hops-role-cap } - [Model registry](registry/index.md) Save a model with metrics, a schema and an input example, one version per save. - [Frameworks](registry/frameworks/python.md) TensorFlow, PyTorch, scikit-learn, LLM and plain Python models. - [Import from Hugging Face](registry/import_huggingface.md) Bring a Hub model into the registry without training it here. - [Evaluation images](registry/model_evaluation_images.md) Attach plots and confusion matrices to a model version.
:material-cloud-upload-outline:{ .hops-role-ico } Serve { .hops-role-cap } - [Deployment creation](serving/deployment.md) Deploy a registered model and check its state. - [Predictor and transformer](serving/predictor.md) Custom inference code and pre/post-processing on KServe. - [Logging and batching](serving/inference-logger.md) Log requests to a feature group, batch them for throughput. - [Resources and autoscaling](serving/resources.md) CPU, memory, GPU and replica bounds, with scheduling constraints. - [Reach the endpoint](serving/rest-api.md) API protocol, REST access from outside, troubleshooting.
:material-monitor-eye:{ .hops-role-ico } Observe { .hops-role-cap } - [Model monitoring](model_monitoring/index.md) Compare logged inference data against the training dataset on a schedule. - [Provenance](provenance/provenance.md) Trace a model back to its training data and features. - [Vector database](../fs/vector_similarity_search.md) Similarity search over embeddings stored in the feature store. - [Agents](../agents/index.md) Agent tasks as jobs and served interactive agents.
================================================================================ # Model Registry Guides Source: https://docs.hopsworks.ai/latest/user_guides/mlops/registry/ # Model Registry Guides **Hopsworks Model Registry** is a centralized repository, within an organization, to manage machine learning models. A model is the product of training a machine learning algorithm with training data. This section provides guides for creating models and publish them to the Model Registry to make them available for download for batch predictions, or deployed to serve realtime applications. ## Exporting a model Follow these framework-specific guides to export a Model to the Model Registry.
- :simple-tensorflow:{ .lg .middle style="color:#FF6F00" } **TensorFlow** --- Export a TensorFlow or Keras model. [:octicons-arrow-right-24: Export guide](frameworks/tf.md) - :simple-pytorch:{ .lg .middle style="color:#EE4C2C" } **Torch** --- Export a PyTorch model. [:octicons-arrow-right-24: Export guide](frameworks/tch.md) - :simple-scikitlearn:{ .lg .middle style="color:#F7931E" } **Scikit-learn** --- Export a scikit-learn model. [:octicons-arrow-right-24: Export guide](frameworks/skl.md) - :material-robot-outline:{ .lg .middle style="color:var(--hops-accent-text)" } **LLM** --- Export a large language model. [:octicons-arrow-right-24: Export guide](frameworks/llm.md) - :simple-python:{ .lg .middle style="color:#3776AB" } **Other Python frameworks** --- Export any other Python model, such as XGBoost or LightGBM. [:octicons-arrow-right-24: Export guide](frameworks/python.md)
## Importing a model from HuggingFace You can also import a model directly from the [HuggingFace Hub](https://huggingface.co) through the Hopsworks UI. The download runs server-side and the model is registered automatically. See [Import from HuggingFace][how-to-import-a-model-from-huggingface]. ## Model Schema A [Model schema](model_schema.md) describes the input and outputs for a model. It provides a functional description of the model which makes it simpler to get started working with it. For example if the model inputs a tensor, the model schema can define the shape and data type of the tensor. ## Input Example An [Input example](input_example.md) provides an instance of a valid model input. Input examples are stored with the model as separate artifacts. ================================================================================ # TensorFlow Source: https://docs.hopsworks.ai/latest/user_guides/mlops/registry/frameworks/tf/ # How To Export a TensorFlow Model ## Introduction In this guide you will learn how to export a TensorFlow model and register it in the Model Registry. !!! notice "Save in SavedModel format" Make sure the model is saved in the [SavedModel](https://github.com/tensorflow/tensorflow/blob/master/tensorflow/python/saved_model/README.md) format to be able to deploy it on TensorFlow Serving. ## Code ### Step 1: Connect to Hopsworks === "Python" ```python import hopsworks project = hopsworks.login() # get Hopsworks Model Registry handle mr = project.get_model_registry() ``` ### Step 2: Train Define your TensorFlow model and run the training loop. === "Python" ```python # Define a model model = tf.keras.Sequential() # Add layers model.add(..) # Compile the model. model.compile(..) # Train the model model.fit(..) ``` ### Step 3: Export to local path Export the TensorFlow model to a directory on the local filesystem. === "Python" ```python model_dir = "./model" tf.saved_model.save(model, model_dir) ``` ### Step 4: Register model in registry Use the `ModelRegistry.tensorflow.create_model(..)` function to register a model as a TensorFlow model. Define a name, and attach optional metrics for your model, then invoke the `save()` function with the parameter being the path to the local directory where the model was exported to. === "Python" ```python # Model evaluation metrics metrics = {"accuracy": 0.92} tf_model = mr.tensorflow.create_model("tf_model", metrics=metrics) tf_model.save(model_dir) ``` ## Going Further You can attach an [Input Example](../input_example.md) and a [Model Schema](../model_schema.md) to your model to document the shape and type of the data the model was trained on. ================================================================================ # Torch Source: https://docs.hopsworks.ai/latest/user_guides/mlops/registry/frameworks/tch/ # How To Export a Torch Model ## Introduction In this guide you will learn how to export a Torch model and register it in the Model Registry. ## Code ### Step 1: Connect to Hopsworks === "Python" ```python import hopsworks project = hopsworks.login() # get Hopsworks Model Registry handle mr = project.get_model_registry() ``` ### Step 2: Train Define your Torch model and run the training loop. === "Python" ```python # Define the model architecture class Net(nn.Module): def __init__(self): super().__init__() self.conv1 = nn.Conv2d(3, 6, 5) ... def forward(self, x): x = self.pool(F.relu(self.conv1(x))) ... return x # Instantiate the model net = Net() # Run the training loop for epoch in range(n): ... ``` ### Step 3: Export to local path Export the Torch model to a directory on the local filesystem. === "Python" ```python model_dir = "./model" torch.save(net.state_dict(), model_dir) ``` ### Step 4: Register model in registry Use the `ModelRegistry.torch.create_model(..)` function to register a model as a Torch model. Define a name, and attach optional metrics for your model, then invoke the `save()` function with the parameter being the path to the local directory where the model was exported to. === "Python" ```python # Model evaluation metrics metrics = {"accuracy": 0.92} tch_model = mr.torch.create_model("tch_model", metrics=metrics) tch_model.save(model_dir) ``` ## Going Further You can attach an [Input Example](../input_example.md) and a [Model Schema](../model_schema.md) to your model to document the shape and type of the data the model was trained on. ================================================================================ # Scikit-learn Source: https://docs.hopsworks.ai/latest/user_guides/mlops/registry/frameworks/skl/ # How To Export a Scikit-learn Model ## Introduction In this guide you will learn how to export a Scikit-learn model and register it in the Model Registry. ## Code ### Step 1: Connect to Hopsworks === "Python" ```python import hopsworks project = hopsworks.login() # get Hopsworks Model Registry handle mr = project.get_model_registry() ``` ### Step 2: Train Define your Scikit-learn model and run the training loop. === "Python" ```python # Define a model iris_knn = KNeighborsClassifier(..) iris_knn.fit(..) ``` ### Step 3: Export to local path Export the Scikit-learn model to a directory on the local filesystem. === "Python" ```python model_file = "skl_knn.pkl" joblib.dump(iris_knn, model_file) ``` ### Step 4: Register model in registry Use the `ModelRegistry.sklearn.create_model(..)` function to register a model as a Scikit-learn model. Define a name, and attach optional metrics for your model, then invoke the `save()` function with the parameter being the path to the local directory where the model was exported to. === "Python" ```python # Model evaluation metrics metrics = {"accuracy": 0.92} skl_model = mr.sklearn.create_model("skl_model", metrics=metrics) skl_model.save(model_file) ``` ## Going Further You can attach an [Input Example](../input_example.md) and a [Model Schema](../model_schema.md) to your model to document the shape and type of the data the model was trained on. ================================================================================ # LLM Source: https://docs.hopsworks.ai/latest/user_guides/mlops/registry/frameworks/llm/ # How To Export a Large Language Model (LLM) ## Introduction In this guide you will learn how to export a [Large Language Model (LLM)](https://www.hopsworks.ai/dictionary/llms-large-language-models) and register it in the Model Registry. ## Code ### Step 1: Connect to Hopsworks === "Python" ```python import hopsworks project = hopsworks.login() # get Hopsworks Model Registry handle mr = project.get_model_registry() ``` ### Step 2: Download the LLM Download your base or fine-tuned LLM. LLMs can typically be downloaded using the official frameworks provided by their creators (e.g., HuggingFace, Ollama, ...) === "Python" ```python # Download LLM (e.g., using huggingface to download Llama-3.1-8B base model) from huggingface_hub import snapshot_download model_dir = snapshot_download("meta-llama/Llama-3.1-8B", ignore_patterns="original/*") ``` ### Step 3: (Optional) Fine-tune LLM If necessary, fine-tune your LLM with an [instruction set](https://www.hopsworks.ai/dictionary/instruction-datasets-for-fine-tuning-llms). A LLM can be fine-tuned fully or using [Parameter Efficient Fine Tuning (PEFT)](https://www.hopsworks.ai/dictionary/parameter-efficient-fine-tuning-of-llms) methods such as LoRA or QLoRA. === "Python" ```python # Fine-tune LLM using PEFT (LoRA, QLoRA) or other methods model_dir = ... ``` ### Step 4: Register model in registry Use the `ModelRegistry.llm.create_model(..)` function to register a model as LLM. Define a name, and attach optional metrics for your model, then invoke the `save()` function with the parameter being the path to the local directory where the model was exported to. === "Python" ```python # Model evaluation metrics metrics = {"f1-score": 0.8, "perplexity": 31.62, "bleu-score": 0.73} llm_model = mr.llm.create_model("llm_model", metrics=metrics) llm_model.save(model_dir) ``` ## Going Further You can attach an [Input Example](../input_example.md) and a [Model Schema](../model_schema.md) to your model to document the shape and type of the data the model was trained on. ================================================================================ # Python Source: https://docs.hopsworks.ai/latest/user_guides/mlops/registry/frameworks/python/ # How To Export a Python Model ## Introduction In this guide you will learn how to export a generic Python model and register it in the Model Registry. ## Code ### Step 1: Connect to Hopsworks === "Python" ```python import hopsworks project = hopsworks.login() # get Hopsworks Model Registry handle mr = project.get_model_registry() ``` ### Step 2: Train Define your XGBoost model and run the training loop. === "Python" ```python # Define a model model = XGBClassifier() # Train model model.fit(X_train, y_train) ``` ### Step 3: Export to local path Export the XGBoost model to a directory on the local filesystem. === "Python" ```python model_file = "model.json" model.save_model(model_file) ``` ### Step 4: Register model in registry Use the `ModelRegistry.python.create_model(..)` function to register a model as a Python model. Define a name, and attach optional metrics for your model, then invoke the `save()` function with the parameter being the path to the local directory where the model was exported to. === "Python" ```python # Model evaluation metrics metrics = {"accuracy": 0.92} py_model = mr.python.create_model("py_model", metrics=metrics) py_model.save(model_dir) ``` ## Going Further You can attach an [Input Example](../input_example.md) and a [Model Schema](../model_schema.md) to your model to document the shape and type of the data the model was trained on. ================================================================================ # Import from HuggingFace Source: https://docs.hopsworks.ai/latest/user_guides/mlops/registry/import_huggingface/ # How To Import a Model from HuggingFace ## Introduction In this guide you will learn how to import a model from the [HuggingFace Hub](https://huggingface.co) directly into the Hopsworks Model Registry. The download runs server-side, the files stream straight into HopsFS under the project's `Models` dataset, and the model version is registered with a framework that Hopsworks auto-detects from the repository's metadata and file list. !!! info "Availability" This feature is part of the Data Science profile and is only available when the `DATA_SCIENCE_PROFILE` feature flag is enabled on the cluster. You also need the `Data Scientist` or `Data Owner` project role. ## UI flow ### Step 1: Open the Model Registry From the project sidebar, open **Data Science → Model Registry**. The list page shows all models registered in this project. ### Step 2: Start a HuggingFace import Click the **Import from HuggingFace** button in the toolbar at the top of the model list. A modal opens asking for the model identifier. ### Step 3: Enter the model ID (and token, if needed) - **Model ID or URL** (required). Accepts either a plain `owner/repo` slug (e.g. `Qwen/Qwen2.5-0.5B`) or the full HuggingFace URL (e.g. `https://huggingface.co/Qwen/Qwen2.5-0.5B`). - **Access token** (optional). Only needed for gated or private repositories. See HuggingFace's [access token docs](https://huggingface.co/settings/tokens) for how to create one. !!! tip "Gated models" If the model requires an access token and you don't supply one, the import fails fast and the modal prompts you to paste a token and retry, so no download time is wasted.

HuggingFace import modal with the model ID field
Enter the HuggingFace model ID

Click **Next** to inspect the repo on HuggingFace. ### Step 4: Choose which weight formats to import Many HuggingFace repos ship the same weights in several interchangeable formats (Safetensors, PyTorch, TensorFlow, Flax, GGUF, ONNX, OpenVINO, Core ML, TensorFlow Lite). Downloading all of them wastes HopsFS storage and bandwidth, and typically you only need one. The modal shows a checkbox for each weight format detected in the repo, with the following default precedence pre-selected: 1. Safetensors 2. PyTorch (`pytorch_model*.bin`) 3. TensorFlow (`tf_model.h5`, `saved_model.pb`) 4. Flax (`flax_model*.msgpack`) 5. Otherwise the first format found (GGUF, ONNX, OpenVINO, Core ML, TensorFlow Lite) You can tick additional formats if you need more than one. Config, tokenizer, README and other small auxiliary files are always imported regardless of your selection.

Weight format selection with a checkbox per format
Pick the weight formats to import; the total download size updates as you tick

#### Quantization variants Some repos, most commonly GGUF builds from `unsloth`, ship the same weights in several quantizations. The variants can be packaged either as files (`Qwen-Q4_K_M.gguf`, `Kimi.IQ4_XS.gguf`) or as subdirectories named after the quant (`UD-Q2_K_XL/`, `UD-Q4_K_XL/`). When variants are detected the modal adds a second checkbox group ("Quantization variant"). Tick the variants you want: files belonging to unticked variants are skipped. Files that don't carry a variant tag (e.g., the root README or a base GGUF in the repo root) are not affected by the variant filter. The recogniser knows the common llama.cpp / unsloth tags: `BF16`, `F16`, `F32`, `FP16`, `FP32`, `Q2_K*`, `Q3_K*`, `Q4_0`, `Q4_K*`, `Q5_K*`, `Q6_K*`, `Q8_0`, `Q8_K*`, `IQ1_S`, `IQ1_M`, `IQ2_*`, `IQ3_*`, `IQ4_*`, `IQ5_K*`, `IQ6_K`, the `Q4_0_*_*` packed variants, and the `UD-Q*` unsloth dynamic series. Click **Import** to start the download. ### Step 5: Monitor progress The modal switches to a progress view polling the backend every few seconds. You'll see: - a progress bar with the overall percentage, - the current file being downloaded, - the file counter (`{completed} / {total}`). You can close the modal and the download continues in the background; re-opening the modal or navigating back to the Model Registry will not interrupt it. The job state is held for one hour after it finishes so you can still view the final status. ### Step 6: Cancel (optional) Click **Cancel** while the download is in progress. You'll be asked whether to **Delete partially downloaded files**: - Leave the box **unchecked** to keep the files that have already been written to HopsFS (useful if you want to inspect what was downloaded so far). - Tick the box to have the server remove the partial `Models/{name}/{version}/` directory. If you close the modal (X or click outside) while the download is in progress, Hopsworks prompts you to either **Continue in background**, which closes the modal and keeps the download running server-side, or **Cancel download**. ### Step 7: Success When all files have been downloaded, the model version is automatically registered in the Model Registry with an auto-detected framework. The modal shows a success screen and the new version appears in the Model Registry list.

Import success screen
The model has been registered and appears in the Model Registry

## Framework auto-detection Hopsworks picks the framework in this order: 1. **HuggingFace `pipeline_tag`** from the repo metadata. Generative tags such as `text-generation`, `text2text-generation`, `conversational`, `image-text-to-text`, `visual-question-answering`, and `document-question-answering` map to `LLM`. 2. **HuggingFace `tags`** array. If it contains `llm` or `large-language-model`, the framework is set to `LLM`. 3. **README / model card**. Phrases such as *"large language model"*, standalone *"LLM"*, or a front-matter `pipeline_tag: text-generation` also map to `LLM`. 4. **File extensions** in the downloaded repo: - `saved_model.pb` → `TENSORFLOW` - `.pkl` or `.joblib` → `SKLEARN` - `pytorch_model.bin`, `.safetensors`, `.pt`, `.pth` → `TORCH` 5. **Fallback** → `PYTHON`. If the detected framework isn't right, you can change it later from the model's detail page. See [Editing the framework][editing-the-framework] below. !!! info "vLLM config for LLMs" When the framework is detected as `LLM`, Hopsworks writes a default `vllmconfig.yaml` alongside the model files (`dtype: "half"`, `gpu_memory_utilization: 0.96`). `max_model_len` is intentionally left out so vLLM uses the context window declared in the model's own `config.json`. Edit this file if you need a smaller context to fit your GPU. ## Editing the framework The framework is shown as a dropdown on the **Summary** panel of the model version page. Select a different value (`TENSORFLOW`, `TORCH`, `SKLEARN`, `LLM`, `PYTHON`, or *No framework*) and the change is persisted immediately. The dropdown is read-only for users with the `Observer` role. The framework is more than a label: the **Deploy this version** button on the same page uses it to pre-fill the deployment form: | Framework | Model server | Config file | Environment match | | ------------ | ------------------- | ---------------- | ----------------- | | `LLM` | vLLM | `vllmconfig.yaml`| `vllm` | | `TENSORFLOW` | TensorFlow Serving | `predictor.py` | `tensorflow` | | `SKLEARN` | Python | `predictor.py` | `sklearn` | | Other | Python | `predictor.py` | `python` | If the corresponding config file (`vllmconfig.yaml` or `predictor.py`) already exists under the model's `Files/` directory, it is auto-selected in the deployment form. ## Managing vLLM configs For models with the `LLM` framework, the model version page shows a **VLLM Configs** button that opens a dialog listing every `*-vllmconfig.yaml` file under the model's `Files/` directory, one per GPU type. A file named plain `vllmconfig.yaml` (no prefix) is labelled *Default*; others are keyed by the GPU they target, for example `NVIDIA-RTX-6000-vllmconfig.yaml`. Each entry can be inspected in-place and switched to edit mode. Saving an edit replaces the file in HopsFS and invalidates the cached content. ### Generate a vLLM config with Platform Intelligence If the Platform Intelligence feature is enabled on the cluster, the model version page also shows a **Generate VLLM Config** button. Pick a GPU type from the dropdown (populated from the cluster's available GPUs) and Platform Intelligence returns a minimal YAML (`dtype`, `max_model_len`, `gpu_memory_utilization`) tuned for that GPU. The result is saved as `-vllmconfig.yaml` next to the model files, and any legacy variants of the same name (with spaces or underscores) are cleaned up. The generator chat session is deleted as soon as the YAML has been uploaded, even if the generation fails. ## Model naming The imported model is registered under the **repo** portion of the HuggingFace ID, with any non-alphanumeric characters replaced by `_`. For example: | HuggingFace ID | Hopsworks model name | | --------------------------- | -------------------- | | `Qwen/Qwen2.5-0.5B` | `Qwen2_5_0_5B` | | `meta-llama/Llama-3.2-1B` | `Llama_3_2_1B` | | `prajjwal1/bert-tiny` | `bert_tiny` | If the model name already exists in the project, the import creates the next integer version; otherwise it creates version `1`. ## Errors If the import fails, the modal shows one of these reasons: | Error | Meaning | | ---------------------------- | ----------------------------------------------------------------------------------------------- | | `auth_required` | A token was supplied but rejected by HuggingFace (invalid, expired, or no grant for this repo). | | `not_found_or_auth_required` | No token supplied and HuggingFace returned 401 (ID is wrong, or the repo is gated/private). | | `model_not_found` | HuggingFace returned 404: the repo does not exist. | | `no_disk_space` | The project ran out of HopsFS storage quota while downloading. | | `download_failed: ` | A specific file failed to download (e.g. transient network issue). | | `invalid_filename: ` | The repo contained a filename with disallowed characters (e.g. `..`). | | `registration_failed: ...` | The files downloaded but the final registration step failed. | For all terminal failures Hopsworks removes the partial model directory from HopsFS on a best-effort basis. ## Going Further - Attach an [Input Example][how-to-attach-an-input-example] and [Model Schema][how-to-attach-a-model-schema] to your imported model. - Serve the imported model with [Model Serving][model-serving-guide]: LLMs auto-pick up the generated `vllmconfig.yaml`. ================================================================================ # Model Schema Source: https://docs.hopsworks.ai/latest/user_guides/mlops/registry/model_schema/ # How To Attach A Model Schema ## Introduction !!! warning "Deprecated" `ModelSchema` is deprecated and will be removed in a future release. A model registered with `create_model(feature_view=...)` gets its input and output schema from the feature view's training dataset, and a deployment describes its requests with the [deployment schema][deployment-schema]. The default predictor still reads a legacy model schema to select the model's input columns; a `DefaultPredict` subclass that overrides `model_predict` replaces that. For a model without a feature view, name its input columns with `passed_features=` on `deploy()`. In this guide you will learn how to attach a model schema to your model. A model schema, describes the type and shape of inputs and outputs (predictions) for your model. Attaching a model schema to your model will give other users a better understanding of what data it expects. !!! info "Model schema and deployment schema" A model registered with `feature_view=` gets its model schema inferred from the feature view's training dataset schema when it is saved. The default predictor checks at pod start that every model input column is served by the feature view, and the deployment schema describes what clients send. See the [Deployment Schema Guide][deployment-schema]. ## Code ### Step 1: Connect to Hopsworks === "Python" ```python import hopsworks project = hopsworks.login() # get Hopsworks Model Registry handle mr = project.get_model_registry() ``` ### Step 2: Create ModelSchema Create a ModelSchema for your inputs and outputs by passing in an example that your model is trained on and a valid prediction. Currently, we support `pandas.DataFrame, pandas.Series, numpy.ndarray, list`. === "Python" ```python # Import a Schema and ModelSchema definition from hsml.model_schema import ModelSchema from hsml.schema import Schema # Model inputs for MNIST dataset inputs = [ { "type": "uint8", "shape": [28, 28, 1], "description": "grayscale representation of 28x28 MNIST images", } ] # Build the input schema input_schema = Schema(inputs) # Model outputs outputs = [{"type": "float32", "shape": [10]}] # Build the output schema output_schema = Schema(outputs) # Create ModelSchema object model_schema = ModelSchema(input_schema=input_schema, output_schema=output_schema) ``` ### Step 3: Set model_schema parameter Set the `model_schema` parameter in the `create_model` function and call `save()` to attaching it to the model and register it in the registry. === "Python" ```python model = mr.tensorflow.create_model(name="mnist", model_schema=model_schema) model.save("./model") ``` ================================================================================ # Input Example Source: https://docs.hopsworks.ai/latest/user_guides/mlops/registry/input_example/ # How To Attach An Input Example ## Introduction In this guide you will learn how to attach an input example to a model. An input example is simply an instance of a valid model input. Attaching an input example to your model will give other users a better understanding of what data it expects. ## Code ### Step 1: Connect to Hopsworks === "Python" ```python import hopsworks project = hopsworks.login() # get Hopsworks Model Registry handle mr = project.get_model_registry() ``` ### Step 2: Generate an input example Generate an input example which corresponds to a valid input to your model. Currently we support `pandas.DataFrame, pandas.Series, numpy.ndarray, list` to be passed as input example. === "Python" ```python import numpy as np input_example = np.random.randint(0, high=256, size=784, dtype=np.uint8) ``` ### Step 3: Set input_example parameter Set the `input_example` parameter in the `create_model` function and call `save()` to attaching it to the model and register it in the registry. === "Python" ```python model = mr.tensorflow.create_model(name="mnist", input_example=input_example) model.save("./model") ``` ================================================================================ # Model Evaluation Images Source: https://docs.hopsworks.ai/latest/user_guides/mlops/registry/model_evaluation_images/ # How To Save Model Evaluation Images ## Introduction In this guide, you will learn how to attach ==model evaluation images== to a model. Model evaluation images are images that visually describe model performance metrics. For example, **confusion matrices**, **ROC curves**, **model bias tests**, and **training loss curves** are examples of common model evaluation images. By attaching model evaluation images to your versioned model, other users can better understand the model performance and evaluation metrics. ## Code ### Step 1: Connect to Hopsworks === "Python" ```python import hopsworks project = hopsworks.login() # get Hopsworks Model Registry handle mr = project.get_model_registry() ``` ### Step 2: Generate model evaluation images Generate an image that visualizes model performance and evaluation metrics === "Python" ```python import seaborn from sklearn.metrics import confusion_matrix # Predict the training data using the trained model y_pred_train = model.predict(X_train) # Predict the test data using the trained model y_pred_test = model.predict(X_test) # Calculate and print the confusion matrix for the test predictions results = confusion_matrix(y_test, y_pred_test) # Create a DataFrame for the confusion matrix results df_confusion_matrix = pd.DataFrame( results, ["True Normal", "True Fraud"], ["Pred Normal", "Pred Fraud"], ) # Create a heatmap using seaborn with annotations heatmap = seaborn.heatmap(df_confusion_matrix, annot=True) # Get the figure and display it fig = heatmap.get_figure() fig.show() ``` ### Step 3: Save the figure to a file inside the model directory Save the figure to a file with a common filename extension (for example, .png or .jpeg), and place it in a directory called `images` - a subdirectory of the model directory that is registered to Hopsworks. === "Python" ```python # Specify the directory name for saving the model and related artifacts model_dir = "./model" # Create a subdirectory of model_dir called 'images' for saving the model evaluation images model_images_dir = model_dir + "/images" if not os.path.exists(model_images_dir): os.mkdir(model_images_dir) # Save the figure to an image file in the images directory fig.savefig(model_images_dir + "/confusion_matrix.png") # Register the model py_model = mr.python.create_model(name="py_model") py_model.save("./model") ``` ================================================================================ # Model Serving Guide Source: https://docs.hopsworks.ai/latest/user_guides/mlops/serving/ # Model Serving Guide ## Deployment Assuming you have already created a model in the [Model Registry](../registry/index.md), a deployment can now be created to prepare a model artifact for this model and make it accessible for running predictions behind a REST or gRPC endpoint. Refer to the [Deployment Creation Guide](deployment.md) for step-by-step instructions on creating a deployment for your model. For details on monitoring the status and lifecycle of an existing deployment, see the [Deployment State Guide](deployment-state.md). !!! tip "Python deployments" If you want to deploy a Python script without a model artifact, see the [Python Deployments](../../projects/python-deployment/python-deployment.md) page. ### Versioning A deployment keeps a numbered history of its configuration. Save changes in place or as a new version, inspect earlier versions and roll back to one, for model deployments and Python deployments alike, see the [Deployment Versions Guide][deployment-versions]. ### Deployment Schema and default predictor Describe the prediction request a deployment accepts, validate requests against it, and serve a model without writing a predictor script, see the [Deployment Schema Guide][deployment-schema]. A feature view can be deployed on its own with the same contract, see the [Feature View Deployment Guide][feature-view-deployment]. ### Predictor (KServe component) Predictors are responsible for running a model server that loads a trained model, handles inference requests and returns predictions, see the [Predictor Guide](predictor.md). ### Transformer (KServe component) Transformers are used to apply transformations on the model inputs before sending them to the predictor for making predictions using the model, see the [Transformer Guide](transformer.md). ### Inference Batcher Configure the predictor to batch inference requests, see the [Inference Batcher Guide](inference-batcher.md). ### Inference Logger Configure the predictor to log inference requests and predictions, see the [Inference Logger Guide](inference-logger.md). ### Resources Configure the resources to be allocated for predictor and transformer in a model deployment, see the [Resources Guide](resources.md). ### Autoscaling Configure autoscaling for your model deployment, including scale metrics and scaling parameters, see the [Autoscaling Guide](autoscaling.md). How a deployment scales depends on whether it runs in KServe Knative or Standard mode, and scale-to-zero is available in Knative mode only, see [Deployment mode](autoscaling.md#deployment-mode). ### Scheduling !!! info "Kueue is required" This feature requires Kueue to be enabled in your cluster. If Kueue is not available, queue and topology options will not be accessible. Configure scheduling for your model deployment using Kueue queues, see the [Scheduling Guide](scheduling.md). ### API Protocol Choose between REST and gRPC API protocols for your model deployment, see the [API Protocol Guide](api-protocol.md). ### REST API Send inference requests to deployed models using REST API, see the [REST API Guide](rest-api.md). ### Troubleshooting Inspect the model server logs to troubleshoot your model deployments, see the [Troubleshooting Guide](troubleshooting.md). ### External access Grant users authenticated by an external Identity Provider access to model deployments, see the [External Access Guide](external-access.md). ================================================================================ # Deployment Creation Source: https://docs.hopsworks.ai/latest/user_guides/mlops/serving/deployment/ # How To Create A Model Deployment ## Introduction In this guide, you will learn how to create a new deployment for a trained model. !!! info This guide covers model deployments, which require a model saved in the Model Registry. To learn how to create a model in the Model Registry, see [Model Registry Guide](../registry/index.md#exporting-a-model). For Python deployments (running a Python script without a model artifact), see [Python Deployments](../../projects/python-deployment/python-deployment.md). Model deployments are used to unify the different components involved in making one or more trained models online and accessible to compute predictions on demand. For each model deployment, there are four concepts to understand: !!! info "" 1. [Model files](#model-files) 2. [Artifact files](#artifact-files) 3. [Predictor](#predictor) 4. [Transformer](#transformer) ## Web UI ### Step 1: Create a deployment If you have at least one model already trained and saved in the Model Registry, navigate to the model deployments page by clicking on `Model Deployments` in the navigation menu on the left.

Deployments navigation tab
Deployments navigation tab

Once in the model deployments page, click on `New model deployment` at the top of the page to open the deployment creation form. ### Step 2: Basic deployment configuration A simplified creation form will appear including the most common deployment fields from all available configurations. We provide default values for the rest of the fields, adjusted to the type of deployment you want to create. In the simplified form, choose the model server that will be used to serve your model.

Select the model server
Select the model server

Then, select the model you want to deploy from the list of available models under `pick a model`.

Select the model
Select the model

After selecting the model, select a model version and give your model deployment a name. !!! info "Deployment name validation rules" A valid deployment name can only contain characters a-z, A-Z and 0-9. !!! info "Predictor script for Python models" For Python models, you must select a custom [predictor script](#predictor) that loads and runs the trained model by clicking on `From project`, `Upload new file` or `Create new file`, to choose an existing script in the project file system, upload a new script, or write one in place, respectively. !!! info "Server configuration file for vLLM" For vLLM deployments, a server configuration file is required. See the [Predictor Guide](predictor.md#server-configuration-file) for more details. Lastly, click on `Create` to create the deployment for your model. ### Step 3 (Optional): Advanced configuration Optionally, you can access and adjust other parameters of the deployment configuration by clicking on `advanced options`.

Advance options
Advanced options. Go to advanced deployment creation form

You will be redirected to a full-page deployment creation form, where you can review all default configuration values and customize them to fit your requirements. In addition to the basic settings, this form allows you to further configure the [Predictor](#predictor) and [Transformer](#transformer) KServe components of your model deployment. Once you are done with the changes, click on `Create new model deployment` at the bottom of the page to create the deployment for your model. ### Step 4: Deployment creation Wait for the deployment creation process to finish.

Creating new deployment
Deployment creation in progress

### Step 5: Deployment overview Once the deployment is created, you will be redirected to the list of all your existing deployments in the project. You can use the filters on the top of the page to easily locate your new deployment.

List of deployments
List of deployments

After that, click on the new deployment to access the overview page.

Deployment overview
Deployment overview

## Code ### Step 1: Connect to Hopsworks === "Python" ```python import hopsworks project = hopsworks.login() # get Hopsworks Model Registry handle mr = project.get_model_registry() ``` ### Step 2: Retrieve your trained model Retrieve the trained model you want to deploy using the Model Registry handle. === "Python" ```python my_model = mr.get_model("my_model", version=1) ``` ### Step 3: Deploy your trained model Create a deployment for your model by calling `.deploy()` on the model metadata object. This will create a deployment for your model with default values. === "Python" ```python my_deployment = my_model.deploy() # optionally, start your model deployment my_deployment.start() ``` !!! info "Predictor script and server configuration file" You can provide a predictor script and a server configuration file directly in the `.deploy()` method using the `script_file` and `config_file` parameters. See the [Predictor Guide](predictor.md) for more details. ### Step 3b: Deploy a model with its feature view A Python model registered with `feature_view=` deploys without a predictor script. The default predictor looks up and transforms the features by serving key, runs the model, logs the request when the feature view has logging enabled, and validates every request against the deployment schema, which the client infers from the feature view. Clients send only the serving keys and the features named in `passed_features`. === "Python" ```python fs = project.get_feature_store() feature_view = fs.get_feature_view("transactions", version=1) # register the trained model with the feature view it was trained on fraud_model = mr.python.create_model(name="fraud", feature_view=feature_view) fraud_model.save("model_dir") # one .pkl or .joblib file inside fraud_deployment = fraud_model.deploy( name="fraud", passed_features=["amount"], # sent by the client, not read from the online store ) fraud_deployment.start(await_running=600) fraud_deployment.schema.describe() # the request contract fraud_deployment.predict(inputs=[{"cc_num": 4473593503484549, "amount": 12.5}]) ``` See the [Deployment Schema Guide][deployment-schema] for the request contract, the error codes, feature logging, and custom predictor scripts that subclass the default predictor. ### Step 3c: Deploy a feature view without a model A feature view deploys on its own and returns the transformed feature vectors, for callers that run the model elsewhere. The deployment pins the training dataset whose statistics the transformations use. === "Python" ```python feature_view = fs.get_feature_view("transactions", version=1) X_train, X_test, y_train, y_test = feature_view.train_test_split(test_size=0.2) fv_deployment = feature_view.deploy( name="transactionsfv", passed_features=["amount"], ) fv_deployment.start(await_running=600) response = fv_deployment.predict(inputs=[{"cc_num": 4473593503484549, "amount": 12.5}]) print(response["columns"], response["predictions"]) ``` See the [Feature View Deployment Guide][feature-view-deployment]. !!! api "API reference" - [`ModelRegistry.get_model`][hsml.model_registry.ModelRegistry.get_model] - [`Model.deploy`][hsml.model.Model.deploy] - [`Deployment`][hsml.deployment.Deployment] - [`start`][hsml.deployment.Deployment.start] - [`predict`][hsml.deployment.Deployment.predict] Browse the full Python API :material-arrow-right: ## Model Files Model files are the files exported when a specific version of a model is saved to the model registry (see [Model Registry](../registry/index.md)). These files are ==unique for each model version, but shared across model deployments== created for the same version of the model. Inside a model deployment, the local path to the model files is stored in the `MODEL_FILES_PATH` environment variable (see [environment variables](../serving/predictor.md#environment-variables)). Moreover, you can explore the model files under the `/Models///Files` directory using the File Browser. !!! warning All files under `/Models` and `/Deployments` are managed by Hopsworks. Manual changes to these files cannot be reverted and can have an impact on existing model deployments. ## Artifact Files Artifact files are essential for the proper initialization and operation of a model deployment. The most critical artifact files are the **predictor** and **transformer scripts**. The predictor script loads the trained model and handles prediction requests, while the transformer script applies any necessary input transformations before inference. Predictor and transformer scripts run on separate components and, therefore, scale independently of each other. !!! tip Whenever you provide a predictor script, you can include the transformations of model inputs in the same script as far as they don't need to be scaled independently from the model inference process. Additionally, artifact files can also contain a **server configuration file** that helps detach configuration used within the model deployment from the model server or the implementation of the predictor and transformer scripts. Inside a model deployment, the local path to the configuration file is stored in the `CONFIG_FILE_PATH` environment variable (see [environment variables](../serving/predictor.md#environment-variables)). Each deployment keeps its configuration, including its artifact files, in numbered ==deployment versions==. A save edits the active version in place unless you save it as a new version, and earlier versions can be reactivated with a rollback, see the [Deployment Versions Guide][deployment-versions]. Inside a model deployment, the local path to the artifact files is stored in the `ARTIFACT_FILES_PATH` environment variable (see [environment variables](../serving/predictor.md#environment-variables)). Deployments with a deployment schema also keep the schema documents under `/Deployments//resources/schema/`, one set of files per schema content id. These files are never modified, so a deployment revision always finds the schema it was created with; see [Revisions in the Deployment Schema Guide][deployment-schema-revisions]. !!! warning All files under `/Models` and `/Deployments` are managed by Hopsworks. Manual changes to these files cannot be reverted and can have an impact on existing model deployments. ## Predictor Predictors are responsible for running the model server that loads the trained model, handles inference requests and returns prediction results. To learn more about predictors and how to configure them, including [environment variables](predictor.md#environment-variables), [resources](predictor.md#resources), and [autoscaling](predictor.md#autoscaling), see the [Predictor (KServe) Guide](predictor.md). ## Transformer Transformers are used to apply transformations on the model inputs before sending them to the predictor for making predictions using the model. To learn more about transformers and how to configure them, including [environment variables](transformer.md#environment-variables), [resources](transformer.md#resources), and [autoscaling](transformer.md#autoscaling), see the [Transformer (KServe) Guide](transformer.md). !!! warning Transformers are not available for vLLM deployments. ================================================================================ # Deployment State Source: https://docs.hopsworks.ai/latest/user_guides/mlops/serving/deployment-state/ # How To Inspect A Deployment State ## Introduction In this guide, you will learn how to inspect the state of a deployment. A state can be seen as a snapshot of the current inner workings of a deployment. The following is the state transition diagram for deployments. --8<-- "user_guides/mlops/serving/deployment-state/status-transitions.html" States are composed of a [status](#deployment-status) and a [condition](#deployment-conditions). While a status represents a high-level view of the state, conditions contain more detailed information closely related to infrastructure terms. ## Web UI ### Step 1: Inspect deployment status If you have at least one deployment already created, navigate to the deployments page by clicking on the `Deployments` tab on the navigation menu on the left.

Deployments navigation tab
Deployments navigation tab

Once in the deployments page, find the deployment you want to inspect. Next to the actions buttons, you can find an indicator showing the current status of the deployment. This indicator changes its color based on the status. To inspect the condition of the deployment, click on the name of the deployment to open the deployment overview page. ### Step 2: Inspect condition Once in the deployment overview page, you can find the aforementioned status indicator at the top of page. Below it, a one-line message is shown with a more detailed description of the deployment status. This message is built using the current [condition](#deployment-conditions) of the deployment.

Deployment status condition
Deployments status condition

### Step 3: Check nº of running instances Additionally, you can find the nº of instances currently running by scrolling down to the `Resources per Instance` section, or by picking it in the left navigation under the deployment.

Resource allocation for a deployment
Resource allocation for a deployment

!!! info "Scale-to-zero capabilities" If scale-to-zero capabilities are enabled, you can see how the nº of instances of a running deployment goes to zero and the status changes to `idle`. To enable scale-to-zero in a deployment, see [Resources Guide](resources.md). Scale-to-zero, and therefore the `idle` status, apply to deployments in ==Knative== mode only, see [Deployment mode](autoscaling.md#deployment-mode). ## Code ### Step 1: Connect to Hopsworks === "Python" ```python import hopsworks project = hopsworks.login() # get Hopsworks Model Serving handle ms = project.get_model_serving() ``` ### Step 2: Retrieve an existing deployment === "Python" ```python deployment = ms.get_deployment("mydeployment") ``` ### Step 3: Inspect deployment state === "Python" ```python state = deployment.get_state() state.describe() ``` ### Step 4: Check nº of running instances === "Python" ```python # nº of predictor instances deployment.resources.describe() # nº of transformer instances deployment.transformer.resources.describe() ``` !!! api "API reference" - [`ModelServing.get_deployment`][hsml.model_serving.ModelServing.get_deployment] - [`Deployment`][hsml.deployment.Deployment] - [`get_state`][hsml.deployment.Deployment.get_state] - [`PredictorState`][hsml.predictor_state.PredictorState] - [`describe`][hsml.predictor_state.PredictorState.describe] Browse the full Python API :material-arrow-right: ## Deployment status The status of a deployment is a high-level description of its current state. ??? info "Deployment statuses" | Status | Description | | -------- | -------------------------------------------------------------------------------------------------------------------------------------------- | | CREATING | Deployment artifacts are being prepared | | CREATED | Deployment has never been started | | STARTING | Deployment is starting | | RUNNING | Deployment is ready and running. Predictions are served without additional latencies. | | IDLE | Deployment is ready but scaled to zero or has no active replicas. Higher latencies (cold-start) are expected on the first inference request. Knative mode only. | | FAILED | Terminal state. The deployment has encountered an unrecoverable error. More details can be found in the status condition. | | UPDATING | Deployment is applying updates to the running instances | | STOPPING | Deployment is stopping | | STOPPED | Deployment has been stopped | ## How States Are Determined Deployment state is determined from multiple sources: the database state (whether the deployment has been deployed and its revision), KServe InferenceService conditions, pod presence (available replicas for predictor and transformer), and the artifact filesystem (whether the deployment artifact files are ready). A revision ID and deployment version are used to distinguish between STARTING (first generation) and UPDATING (subsequent changes to a running deployment). ## Deployment conditions A condition contains more specific information about the status of the deployment. They are mainly useful to track the progress of starting or stopping deployments. Status conditions contain three pieces of information: type, status and reason. While the type describes the purpose of the condition, the status represents its progress. Additionally, a reason field is provided with a more descriptive message of the status. ??? info "Deployment conditions" | Type | Status | Description | | ----------- | --------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------- | | STOPPED | `Unknown` | Deployment is stopping. | | | `True` | Deployment is stopped. Therefore, no instances are running and no resources are allocated. | | SCHEDULED | `Unknown` | Deployment is being scheduled | | | `False` | Deployment failed to be scheduled. This is commonly due to insufficient resources to satisfy the deployment requirements | | | `True` | Deployment has been scheduled successfully. At this point, resources have been already allocated for the deployment. | | INITIALIZED | `Unknown` | Deployment is initializing. This step involves initialization tasks such as pulling docker images or mounting data volumes | | | `False` | Deployment failed to initialized | | | `True` | Deployment has been initialized successfully. At this point, the docker images have been pulled and data volumes mounted | | STARTED | `Unknown` | Deployment is starting. In this step, the model server is started and predictor / transformer scripts (if provided) are executed | | | `False` | Deployment failed to start. This can be due to errors in the predictor / transformer script, missing dependencies or model server incompatibilities. | | | `True` | Deployment has been started successfully. At this point, the model server has been started and the predictor / transformer scripts (if provided) executed. | | READY | `Unknown` | Connectivity is being set up. | | | `False` | Connectivity failed to be set up, mainly due to networking issues. | | | `True` | Connectivity has been set up and the deployment is ready | Condition transitions while a deployment starts: --8<-- "user_guides/mlops/serving/deployment-state/conditions-starting.html" Condition transitions while a deployment stops: --8<-- "user_guides/mlops/serving/deployment-state/conditions-stopping.html" ================================================================================ # Deployment Versions Source: https://docs.hopsworks.ai/latest/user_guides/mlops/serving/deployment-versions/ # How To Manage Deployment Versions { #deployment-versions } ## Introduction In this guide, you will learn how a deployment keeps a numbered history of its configuration, and how to save, inspect and roll back versions. Every deployment has one or more versions. A new deployment starts at version 1, and exactly ==one version is active at a time==, the one the deployment runs. This applies to model deployments and to [Python deployments][python-deployment], which run a script without a model. When you save changes to a deployment, you choose how they are stored: - **Save** edits the active version in place, and it keeps its number. - **Save as new version** stores the configuration as a new version, numbered one above the highest the deployment ever had, and makes it active. The previous versions are kept with their files, so a **rollback** reactivates an earlier version without copying anything. Running instances restart only when the active configuration changes, so a save without changes neither restarts the deployment nor creates a version. Two cases count as a change even when the form is unchanged: a script read from outside the version, from a HopsFS path or a git repository, and a Python environment whose libraries changed since the deployment started. !!! tip "Restarting a deployment" To restart a deployment without saving, you can stop and start it. Saving in place keeps the version number, so the `DEPLOYMENT_VERSION` environment variable and the `deployment_version` column of logged features stay the same across the change. Use **Save as new version** when you need to tell the predictions of the old and new configuration apart. ## What a version holds A version holds the configuration of the predictor and the transformer: | Configuration | Details | | ------------- | ------- | | **Model** | The model and model version, for model deployments. | | **Artifact files** | The predictor and transformer scripts, and the server configuration file, as described in [Artifact files of a version](#artifact-files-of-a-version). | | **vLLM** | The vLLM settings for LLM deployments. | | **Python environments** | The environments the predictor and the transformer run in. | | **Resources and scaling** | Resources, scaling configuration and environment variables. | | **Source and observability** | The git source, tracing and feature logging configuration. | The **API protocol**, **request batching**, **inference logging**, **scheduling configuration** and **Knative mode** belong to the deployment rather than to a version. They are edited in place whichever way you save, and a rollback does not restore them, because they apply to the endpoint and the infrastructure it runs on rather than to what it serves. !!! info "Versions edited in place" A version that was edited in place with `Save` after it was created comes back as edited, when you roll back to it. Use `Save as new version` when you want to be able to return to the current configuration. ## Web UI ### Step 1: Save a new version Open the deployment and go to its edit page. After making your changes, open the menu next to the `Save` button at the bottom of the form and click on `Save as new version`. Clicking on `Save` instead edits the active version in place.

Save as new version option in the Save menu
Save as new version option in the Save menu

### Step 2: Inspect the versions The `Versions` card on the deployment overview page lists every time a version became the active one, with when, by whom and why: creation, new version or rollback. The version the deployment runs is marked as `active`. Click on `Details` next to a version to see its configuration, including its model, scripts, Python environments, scaling and resources. The scripts can be downloaded from there, and the model and the Python environments link to their pages.

Details button in the Versions card
Details button in the Versions card

### Step 3: Roll back Click on `Roll back` next to an earlier version and confirm. A running deployment restarts with the configuration of that version. Rolling back requires the Data owner role in the project, and is disabled while the deployment is starting or stopping. It stays available while a new version is updating, so a version that fails to come up can be rolled back right away. The rollback then starts a new rollout with the configuration of the earlier version, and the instances that were serving keep serving until it is ready. An edit saved in place cannot be undone with a rollback, because it changed the active version itself: stop the deployment or wait for it to fail, then save the previous configuration again.

Roll back button in the Versions card
Roll back button in the Versions card

## Code ### Step 1: Connect to Hopsworks === "Python" ```python import hopsworks project = hopsworks.login() # get Hopsworks Model Serving handle ms = project.get_model_serving() my_deployment = ms.get_deployment("mydeployment") ``` ### Step 2: Save a new version Change the configuration and call `.save()` with `new_version=True`. Without it, the active version is edited in place. === "Python" ```python my_deployment.predictor.scaling_configuration.min_instances = 2 my_deployment.predictor.scaling_configuration.max_instances = 2 my_deployment.save(new_version=True) print(my_deployment.version) # number of the active version ``` ### Step 3: List the versions === "Python" ```python for version in my_deployment.get_versions(): # newest first print(version.version, version.active, version.created, version.created_by) ``` ### Step 4: Roll back === "Python" ```python my_deployment.rollback(1) ``` !!! api "API reference" - [`ModelServing.get_deployment`][hsml.model_serving.ModelServing.get_deployment] - [`Deployment`][hsml.deployment.Deployment] - [`save`][hsml.deployment.Deployment.save] - [`get_versions`][hsml.deployment.Deployment.get_versions] - [`rollback`][hsml.deployment.Deployment.rollback] - [`DeploymentVersion`][hsml.deployment_version.DeploymentVersion] Browse the full Python API :material-arrow-right: ## Artifact files of a version The artifact files of each version, such as the predictor and transformer scripts and the server configuration file, are stored under `/Deployments///`. A script read from a HopsFS path or a git repository is stored in the version as that path or repository, not copied, so a rollback restores the configuration but not the code. Inside a deployment, the active version number is available in the `DEPLOYMENT_VERSION` environment variable, and the local path to its artifact files in `ARTIFACT_FILES_PATH`. !!! warning All files under `/Models` and `/Deployments` are managed by Hopsworks. Manual changes to these files cannot be reverted and can have an impact on existing model deployments. ================================================================================ # Deployment Schema Source: https://docs.hopsworks.ai/latest/user_guides/mlops/serving/deployment-schema/ # How To Use A Deployment Schema { #deployment-schema } ## Introduction In this guide, you will learn how a deployment describes the prediction requests it accepts, how clients read that contract, and how to serve a model without writing a predictor script. A deployment schema lists the fields a client sends with each request, with their types, nullability, and order, plus the shape of the response. It is inferred from the feature view the model was registered with: the serving keys, the features you pass with the request, the request parameters of on-demand transformations, and the extra columns of feature logging. It is published as a JSON Schema and an OpenAPI document, so clients in any language can validate requests before sending them. Every REST V1 request is validated in the pod before any predictor code runs, and rejected with a structured error when it does not match. The wrapper that enforces it covers the KServe REST V1 protocol only. A gRPC deployment served by the default predictor is still validated, by the predictor itself on the rows it decodes from the request tensors. A gRPC deployment running your own script is not validated on either side, and the pod logs a warning at startup. The **default predictor** is the library class that serves such a deployment: it looks up and transforms the features by serving key, runs the model, and logs the request when the feature view has logging enabled. A [feature view can be deployed on its own][feature-view-deployment] with the same class and the same contract, returning the transformed feature vector instead of a prediction. ## Code ### Step 1: Connect to Hopsworks === "Python" ```python import hopsworks project = hopsworks.login() fs = project.get_feature_store() mr = project.get_model_registry() ``` ### Step 2: Register the model with its feature view Register the model with `feature_view=` so the deployment knows where its features come from. The training dataset version is taken from the training dataset you last read or created with that feature view in this session, and the model schema is inferred from the feature view's training dataset schema. === "Python" ```python feature_view = fs.get_feature_view("transactions", version=1) X_train, X_test, y_train, y_test = feature_view.train_test_split(test_size=0.2) # ... train and pickle the model into model_dir ... model = mr.python.create_model( name="fraud", feature_view=feature_view, ) model.save("model_dir") ``` The default predictor loads a single `.pkl`, `.pickle`, or `.joblib` file from the model directory. ### Step 3: Deploy without a predictor script === "Python" ```python deployment = model.deploy( name="fraud", passed_features=["amount"], # features the client sends with each request ) deployment.start(await_running=600) ``` The default predictor is used when the model is a Python model registered with a feature view, no `script_file` or transformer is given, and the deployment uses KServe. Pass `default_predictor=True` to force it, for instance for a scikit-learn model, or `default_predictor=False` to keep the plain model server. Such a deployment serves either [API protocol][api-protocol-guide]. It defaults to REST, which is what the `curl` example and the OpenAPI document below use. Pass `api_protocol="GRPC"` to serve gRPC instead, which costs less per request under concurrency: the library owns both ends of the encoding, so the rows travel as one KServe v2 tensor per schema field and `deployment.predict()` returns the same dictionary it returns over REST. A deployment serves one protocol, not both, so a gRPC deployment answers no HTTP and neither `curl` nor the OpenAPI document below reaches it. At pod start the predictor checks that every input column of the model schema is served by the feature view, with a compatible type. A mismatch fails the deployment with the offending columns in `deployment.get_logs()`, instead of serving wrong predictions. If the feature view has a transformation that needs training dataset statistics, such as `min_max_scaler`, and the model has no training dataset version, `deploy()` refuses and names the transformation. ### Step 4: Read the contract === "Python" ```python schema = deployment.schema schema.describe() # one row per field: group, name, type, nullable print(schema.names) # the order of positional rows print(schema.unresolved) # fields whose type is not known, such as request parameters json_schema = schema.to_json_schema() # {"request": ..., "response": ...} openapi = schema.to_openapi("fraud", url=deployment.get_inference_url()) ``` Fields belong to one of four groups: | group | source | required | | --- | --- | --- | | serving keys | the feature view's serving keys, present only when a feature is looked up | yes, non-null | | passed features | `passed_features=` | yes, nullable when the feature is | | request parameters | arguments of on-demand transformations that are not features | yes | | extra logging features | extra columns of the feature view's logging, minus the reserved ones | no | Request parameter types come from the annotations of the transformation function's arguments: `def amount_ratio(amount: float, budget: float)` records `budget` as `double`. An unannotated argument is reported as unresolved, and any value is accepted for it unless you refine the schema (Step 6). ### Step 5: Send requests Rows are objects keyed by field name, or arrays in `schema.names` order. A request holds one to `schema.max_batch_rows` rows (default 512). The limit is part of the schema, so the published JSON Schema, the client, the transformer, and the predictor all apply the same one. Set `SERVING_MAX_BATCH_ROWS` in `env_vars=` to change it; the change publishes a new schema id. A request carries either `instances` or `inputs`, never both. === "Python" ```python deployment.predict(inputs=[{"cc_num": 4473593503484549, "amount": 12.5}]) deployment.predict(inputs=[[4473593503484549, 12.5]]) ``` The client validates the rows against the schema before sending and raises `ModelServingException` with every problem found; pass `validate=False` to skip that and let the pod answer. A batch is all objects or all arrays, in the published JSON Schema as in the pod. === "curl" ```bash # INFERENCE_URL is deployment.get_inference_url(), also shown on the deployment page curl -X POST "$INFERENCE_URL" \ -H "Authorization: ApiKey $API_KEY" -H "Content-Type: application/json" \ -d '{"instances": [{"cc_num": 4473593503484549, "amount": 12.5}]}' ``` ### Step 6: Refine or replace the schema Pass `schema=` to `deploy()` to refine the inferred schema, for instance to give a request parameter a type. A refinement keeps the inferred fields; adding or removing one is refused. === "Python" ```python from hsml.deployment_schema import DeploymentSchema deployment = model.deploy(name="fraud", passed_features=["amount"]) inferred = deployment.schema refined = DeploymentSchema( serving_keys=inferred.serving_keys, passed_features=inferred.passed_features, request_parameters=[{"name": "rate", "type": "double", "nullable": False}], extra_logging_features=inferred.extra_logging_features, feature_view=inferred.feature_view, training_dataset_version=inferred.training_dataset_version, output=inferred.output, ) deployment.schema = refined deployment.save() ``` A custom predictor script deployed with `schema=` (or `passed_features=`) gets the same validation in the pod, before its `predict()` is called, for REST V1 requests. A custom script asked to serve gRPC is checked on neither side, and has to read v2 tensors itself. ### Step 7: Republish after changing the feature view The served contract does not change when you enable logging, add logging columns, or change the feature view. Re-infer and save to publish the new contract as a new revision: === "Python" ```python deployment.reinfer_schema() deployment.save() ``` ## Revisions { #deployment-schema-revisions } Every schema is content-addressed: `deployment.schema_id` is a hash of its content, and equal schemas have equal ids. The client writes the schema and its JSON Schema and OpenAPI renderings to `/Deployments//resources/schema/.*` before the deployment is created or updated, and records the id in the environment variable `SERVING_SCHEMA_ID` of the predictor and of the transformer, when there is one. `SERVING_SCHEMA_ENFORCER` on the same components records which of the two validates requests for that revision. The files are never modified or removed while the deployment exists. A deployment revision therefore always enforces the exact schema it was created with, and answers discovery consistently with what it enforces. The pod also takes its model, feature view, and training dataset version from its own revision, and refuses to start when the schema was published for another training dataset version than the one it would serve. Updating the schema rolls the instances: old pods keep the old contract until they are replaced. Rolling back is `deployment.schema = previous_schema; deployment.save()`, which points the revision at a file that is still there. ## Discovery for non-Python clients The Hopsworks REST API serves the three documents to any client with an API key that has the `SERVING` scope, so prediction access implies discovery access: ```bash SERVING_ID=$(curl -s -H "Authorization: ApiKey $API_KEY" \ "https://$HOST/hopsworks-api/api/project/$PROJECT_ID/serving?name=fraud" | jq .id) curl -s -H "Authorization: ApiKey $API_KEY" \ "https://$HOST/hopsworks-api/api/project/$PROJECT_ID/serving/$SERVING_ID/schema?format=openapi" ``` `format` is `schema` (default), `jsonschema`, or `openapi`. `schemaId=` returns the documents of an earlier revision. A deployment without a schema, or an unknown id, answers `404` with error code `240037`. ## Type encoding The JSON Schema fragment, the accepted JSON values, and the encoding the Python client applies follow one table. | feature type | JSON Schema | accepted JSON | Python client sends | | --- | --- | --- | --- | | `tinyint`, `smallint`, `int` | `{"type": "integer"}` | integer | `int` | | `bigint` | integer or decimal string | integer, or a decimal string for values beyond 2^53 | `int` | | `float`, `double` | `{"type": "number"}` | finite number | `int`, `float` | | `decimal(p,s)` | number or string | number, or decimal string | `Decimal` as string | | `string`, `varchar(n)`, `char(n)` | `{"type": "string"}` | string | `str` | | `boolean` | `{"type": "boolean"}` | boolean | `bool` | | `timestamp` | RFC 3339 string or integer | RFC 3339 string, or epoch milliseconds | `datetime` as RFC 3339 UTC | | `date` | date string or integer | `YYYY-MM-DD`, or days since epoch | `date` as `YYYY-MM-DD` | | `binary` | base64 string | base64 string | `bytes` as base64 | | `array` | array of `T` | array | list | | `struct<...>` | object with exactly those fields | object | dict | | `map` | object with values of `V` | object | dict | | unresolved | `{}` | anything | unchanged | ## Errors { #deployment-schema-errors } Errors raised by the default predictor and by the schema enforcement carry a structured `detail`: ```json {"detail": { "code": "SCHEMA_VALIDATION", "message": "Prediction request does not match the deployment schema of 'fraud'.", "schema_id": "3f9a1c2b7d4e6f80", "errors": [ {"row": 0, "field": "amount", "reason": "missing"}, {"row": 2, "field": "cc_num", "reason": "must not be null"} ]}} ``` | status | code | when | | --- | --- | --- | | 400 | `SCHEMA_VALIDATION` | the request does not match the schema; `errors` names every row and field | | 400 | `FEATURE_LOOKUP_FAILED` | the feature store rejected the lookup for a reason other than a missing entity | | 404 | `ENTITY_NOT_FOUND` | at least one row's serving keys match no entity; the batch is rejected and `errors` names the rows | | 413 | `BATCH_TOO_LARGE` | more than `SERVING_MAX_BATCH_ROWS` rows | | 422 | `TRANSFORMATION_FAILED` | a transformation raised; `field` is the transformation name | | 500 | `MODEL_FAILED` | the model raised | | 500 | `CONTRACT_VIOLATION` | the pod produced a result that does not match the contract: a different number of vectors or predictions than rows, or other feature vector columns than published | | 503 | `FEATURE_STORE_UNAVAILABLE` | the online store or the feature store API could not be reached | A request is all or nothing: either every row gets a prediction, in request order, or the whole request fails and no row is logged. Error responses never include feature values or exception text: a failure names the exception type only, and `detail.request_id` carries the correlation id (the `x-request-id` header, or one generated for the request) under which the pod log holds the full error. The Hopsworks REST inference proxy does not forward `x-request-id`; send it through the Istio URL when the id must be yours. Responses with a 5xx status may be retried; a retry may read newer features and always produces another log row, so reuse the `x-request-id` header to tie the rows together. ## Feature logging and monitoring { #deployment-schema-feature-logging } When the feature view has logging enabled, the default predictor logs every request with the untransformed and transformed features, the predictions, the request id, the training dataset version, and the model name and version, so `deployment.create_model_monitoring()` works with no extra code. Declare the reserved extra logging columns on the feature view and the predictor fills them, which tells deployments and revisions apart in the log: | column | type | value | | --- | --- | --- | | `deployment_name` | `string` | the deployment name | | `deployment_version` | `int` | the deployment version | | `deployment_schema_id` | `string` | the schema id of the revision that served the request | | `request_row` | `int` | the row's index in its request | Any other extra logging column becomes a request field that clients may send. ### Configuring feature logging per deployment { #deployment-schema-feature-logging-config } The predictor coalesces log rows into Arrow batches and hands them to the transport the feature view logs through: on `realtime` it posts them to the deployment's inference logger, which produces them to Kafka, and they reach the online store within seconds and the offline store on the materialization schedule; on `job` it appends them to a file buffer on the pod, which is uploaded to HopsFS and committed to the offline store by the view's commit job. Both sides take their limits from platform variables that an administrator sets, and a deployment can override any of them with a `DeploymentLoggingConfig` (`hsml.deployment_logging_config`) passed to `deploy()`, `create_predictor()` or `feature_view.deploy()`, or set on the deployment before it starts. ```python from hsml.deployment_logging_config import DeploymentLoggingConfig deployment = model.deploy( feature_logging=DeploymentLoggingConfig( batch_bytes=256 * 1024, # post once a quarter megabyte is waiting batch_seconds=2, # or after two seconds under load ) ) deployment.start() ``` A dict with the same field names is accepted wherever the object is. Fields left unset keep the platform default. The values are read when the pods start: edit `deployment.feature_logging`, call `deployment.save()`, and `deployment.restart()` a running deployment to apply them. ```python deployment.feature_logging.batch_seconds = 1 deployment.save() deployment.restart() ``` | Field | Controls | Platform default | Variable | | --- | --- | --- | --- | | `transport` | The feature view's transport, `realtime` or `job`; a deployment cannot choose the other one, the field only documents or checks it | the view's | `serving_feature_logging_transport` (the default for new views) | | `batch_bytes` | Coalesced bytes that force a post while the predictor has a backlog; an idle predictor posts at once | 1 MiB | `serving_feature_logger_batch_bytes` | | `batch_seconds` | Longest time the predictor holds a partial batch before posting it | 5 | `serving_feature_logger_batch_seconds` | | `batch_rows` | Most rows one post carries; the platform default is the inference logger's own limit, so a lower value only makes posts smaller | 512 | `serving_feature_logger_max_event_rows` | | `queue_size` | Rows the predictor keeps queued for logging, including rows in flight, beyond which rows are dropped and counted; the queue holds the requests' rows as received, so wide rows hold more memory per row | 1000 | `serving_feature_logger_queue_size` | | `max_event_bytes` | Largest single post; a group of requests larger than this goes out as several posts | 8 MiB | `serving_feature_logger_max_event_bytes` | | `sidecar_cpu` | CPU request of the sidecar container, in cores | from the chart | inference logger values | | `sidecar_memory_mb` | Memory request of the sidecar container, in MiB | from the chart | inference logger values | The `job` transport adds `flush_bytes` (1 MiB) and `flush_interval_seconds` (300), the size and age at which the pod's buffer segment is closed and uploaded, `max_buffer_bytes` (64 MiB), beyond which new rows are dropped while uploads fail, and `shutdown_seconds`, the budget a stopping pod has to upload its buffer and start the commit job; setting them on a `realtime` deployment is rejected. Their variables are `serving_feature_logger_flush_bytes`, `serving_feature_logger_flush_interval_seconds`, `serving_feature_logger_max_buffer_bytes` and `serving_feature_logger_shutdown_seconds`. A stopping pod is given a termination grace period of 30 seconds for the Knative drain plus `shutdown_seconds` plus 5, so a stop takes about that long to complete. A pod that has uploaded 32 MiB since the last run asks the commit job to run ahead of its schedule, at most once every five minutes. The object rejects non-positive values and inconsistent pairs: `batch_rows` cannot exceed `queue_size`, and `batch_bytes` cannot exceed `max_event_bytes`. A post is closed as soon as the next request would take it past `batch_rows` or `max_event_bytes`, so a backlog is posted in receiver-sized pieces. Under load a lower `batch_bytes` or `batch_seconds` shortens the time a row waits in the predictor at the cost of more posts; when the deployment is idle every row is posted as soon as it is built, whatever the values. Stopping a deployment posts whatever the predictor still holds before the pod exits; on the `job` transport it uploads the buffer and starts the commit job, and `deployment.commit_feature_logs()` runs that job on demand. Logging is asynchronous and a logging failure does not fail prediction. How often the rows reach the offline store is a property of the feature view, not the deployment; see [Choosing the Materialization Interval][choosing-the-materialization-interval]. A deployment that logs features serves on one worker process by default, whatever its CPU limit. The rows are safe with several: each worker buffers separately and the buffer directory lock arbitrates who adopts a dead worker's segments. The metrics are not: they are this process's counters, read when Prometheus scrapes, so with several workers behind one port a scrape reaches one of them and the Feature logging card reports a fraction of the rows. Set `KSERVE_WORKERS` in `env_vars=` to serve on more than one anyway, and read the card as a sample rather than a total. ## Deployments without lookups { #deployment-schema-no-lookup } Nothing is looked up in the online store when every stored feature of the view arrives with the request, or when the model has no feature view. In both cases the schema has no serving keys. ### Every feature passed Pass every non-label stored feature of the view in `passed_features=`; on-demand features are computed from the request parameters, so they are never looked up. The view still computes the on-demand features and applies the model-dependent transformations with the pinned training dataset's statistics, and feature logging and monitoring work as for any other deployment. === "Python" ```python stored = [ f.name for f in feature_view.features if not f.label and f.on_demand_transformation_function is None ] deployment = model.deploy(name="fraud_passed", passed_features=stored) print(deployment.schema.serving_keys) # [] ``` The same applies to `feature_view.deploy(passed_features=stored)`, which then returns the transformed vectors of the passed features. ### A model without a feature view A Python model registered without a feature view deploys with the default predictor when you pass `default_predictor=True` and name its input columns with `passed_features=`, in the order the model expects. Those columns are the whole request: the schema has no serving keys and no request parameters, nothing is looked up or transformed, and there is no feature logging or monitoring because there is no feature view. The types are unresolved, so any JSON value is accepted for them; refine the schema with `schema=` to pin them down. === "Python" ```python model = mr.python.create_model(name="fraud_plain") model.save("model_dir") deployment = model.deploy( default_predictor=True, passed_features=["amount", "age_days"], # the model's input columns, in its order ) deployment.schema.describe() # passed features only, types unresolved deployment.predict(inputs=[{"amount": 12.5, "age_days": 41}]) ``` ## Access control Prediction through the Hopsworks REST API and through the Istio ingress requires an API key with the `SERVING` scope and the Data Owner or Data Scientist role in the project. The pod looks up features as the project's serving identity, not as the caller. Anyone allowed to call `:predict` can therefore obtain the transformed features of any entity the feature view can serve, and a feature view deployment returns those features directly. Log rows contain feature values and are governed by the logging feature group's permissions. ## Environment variables | variable | set by | meaning | | --- | --- | --- | | `SERVING_SCHEMA_ID` | the client | the schema the revision serves | | `SERVING_FEATURE_VIEW_NAME`, `SERVING_FEATURE_VIEW_VERSION` | the client, feature view deployments | the feature view served | | `SERVING_TRAINING_DATASET_VERSION` | the client, feature view deployments | the pinned training dataset | | `SERVING_SCHEMA_ENFORCER` | the client | `predictor` or `transformer`: the component of the revision that validates requests | | `SERVING_MAX_BATCH_ROWS` | you, through `env_vars=` | rows accepted per request, default 512; recorded in the schema at publication | | `SERVING_PREDICTOR_ASYNC_LOOKUP` | you, through `env_vars=` | `false` returns the default predictor to the blocking online store lookup | | `KSERVE_WORKERS` | you, through `env_vars=` | uvicorn worker processes; the default is one per whole core, capped at 4, and one on a deployment that logs features | | `HOPSWORKS_FEATURE_LOGGING_TRANSPORT` | the backend | the transport the view's logging group uses, set only on a deployment that logs | The `HOPSWORKS_*` names are reserved and refused in `env_vars=`, as are `SERVING_SCHEMA_ID`, `SERVING_FEATURE_VIEW_NAME`, `SERVING_FEATURE_VIEW_VERSION`, `SERVING_TRAINING_DATASET_VERSION` and `SERVING_SCHEMA_ENFORCER`. The logging limits are set through `DeploymentLoggingConfig` rather than through `env_vars=`: `FEATURE_LOGGER_QUEUE_SIZE`, `FEATURE_LOGGER_BATCH_BYTES` and `FEATURE_LOGGER_BATCH_SECONDS` are reserved, so a value set there is refused. !!! api "API reference" - [`Model.deploy`][hsml.model.Model.deploy] - [`FeatureView.deploy`][hsfs.feature_view.FeatureView.deploy] - [`Deployment`][hsml.deployment.Deployment] - [`predict`][hsml.deployment.Deployment.predict] - [`reinfer_schema`][hsml.deployment.Deployment.reinfer_schema] - [`get_logs`][hsml.deployment.Deployment.get_logs] - [`schema`][hsml.deployment.Deployment.schema] - [`DeploymentSchema`][hsml.deployment_schema.DeploymentSchema] - [`describe`][hsml.deployment_schema.DeploymentSchema.describe] - [`to_openapi`][hsml.deployment_schema.DeploymentSchema.to_openapi] - [`to_json_schema`][hsml.deployment_schema.DeploymentSchema.to_json_schema] - [`DefaultPredict`][hsml.default_predictor.DefaultPredict] Browse the full Python API :material-arrow-right: ================================================================================ # Predictor (KServe) Source: https://docs.hopsworks.ai/latest/user_guides/mlops/serving/predictor/ # How To Configure A Predictor ## Introduction In this guide, you will learn how to configure a predictor for a trained model. !!! warning This guide assumes that a model has already been trained and saved into the Model Registry. To learn how to create a model in the Model Registry, see [Model Registry Guide](../registry/frameworks/tf.md) Predictors are the main component of deployments. They are responsible for running a model server that loads a trained model, handles inference requests and returns predictions. They can be configured to use different model servers, different resources or scale differently. In each predictor, you can decide the following configuration: !!! info "" 1. [Model server](#model-server) 2. [Predictor script](#predictor-script) 3. [Server configuration file](#server-configuration-file) 4. [vLLM variant and image version](#vllm-variant-and-image-version) 5. [Python environments](#python-environments) 6. [Transformer script](#transformer-script) 7. [Inference Logger](#inference-logger) 8. [Inference Batcher](#inference-batcher) 9. [Resources](#resources) 10. [Autoscaling](#autoscaling) 11. [Scheduling](#scheduling) 12. [API protocol](#api-protocol) ## Web UI ### Step 1: Create new deployment If you have at least one model already trained and saved in the Model Registry, navigate to the model deployments page by clicking on `Model Deployments` in the navigation menu on the left.

Deployments navigation tab
Deployments navigation tab

Once in the model deployments page, click on `New model deployment` at the top of the page to open the deployment creation form. ### Step 2: Choose a model server A simplified creation form will appear, including the most common deployment fields from all available configurations. The first step is to choose a ==model server== for your model deployment. The model server will filter the models shown below according to the framework that the model was registered with in the model registry. For example if you registered the model as a TensorFlow model using `ModelRegistry.tensorflow.create_model(...)` you select `Tensorflow Serving` in the dropdown.

Select the model server
Select the model server

All models compatible with the selected model server will be listed in the model dropdown.

Select the model
Select the model

After selecting a model from the dropdown, you can optionally choose a predictor script, modify the predictor environment, add a configuration file, or adjust other advanced settings as described in the optional steps below. Otherwise, click on `Create` to create the deployment for your model with default values. ### Step 3 (Optional): Select a predictor script For python models, to select a [predictor script](#predictor-script) click on `From project` and navigate through the file system to find it, click on `Upload new file` to upload a predictor script now, or click on `Create new file` to write one in place.

Predictor script in the simplified deployment form
Select a predictor script in the simplified deployment form

### Step 4 (Optional): Change predictor environment If you are using a predictor script, it is required to select an inference environment for the predictor. This environment needs to have all the necessary dependencies installed to run your predictor script. Hopsworks provide a collection of built-in environments like `minimal-inference-pipeline`, `pandas-inference-pipeline` or `torch-inference-pipeline` with different sets of libraries pre-installed. By default, the `pandas-inference-pipeline` Python environment is used. To create your own it is recommended to [clone](../../projects/python/python_env_clone.md) the `pandas-inference-pipeline` and install additional dependencies for your use-case.

Predictor script in the simplified deployment form
Select an environment for the predictor script

### Step 5 (Optional): Select a configuration file You can select a configuration file to be added to the [artifact files](deployment.md#artifact-files). In Python model deployments, this configuration file will be available inside the model deployment at the local path stored in the `CONFIG_FILE_PATH` environment variable. In vLLM deployments, this configuration file will be directly passed to the vLLM server. You can find all configuration parameters supported by the vLLM server in the [vLLM documentation](https://docs.vllm.ai/en/v0.28.0/cli/serve.html). !!! info Configuration files are required for vLLM deployments as they are used to define the configuration for the vLLM server.

Server configuration file in the simplified deployment form
Select a configuration file in the simplified deployment form

### Step 6 (Optional): Other advanced options To access the advanced deployment configuration, click on `advanced options`.

Advance options
Advanced options. Go to advanced deployment creation form

Here, you can further change the default values of the predictor: !!! info "Predictor configuration" 1. [vLLM variant and image version](#vllm-variant-and-image-version) 2. [Transformer](#transformer-script) 3. [Inference logger](#inference-logger) 4. [Inference batcher](#inference-batcher) 5. [Resources](#resources) 6. [Autoscaling](#autoscaling) 7. [Scheduling](#scheduling) 8. [API protocol](#api-protocol) Once you are done with the changes, click on `Create new model deployment` at the bottom of the page to create the deployment for your model. ## Code ### Step 1: Connect to Hopsworks === "Python" ```python import hopsworks project = hopsworks.login() # get Hopsworks Model Registry handle mr = project.get_model_registry() # get Hopsworks Model Serving handle ms = project.get_model_serving() ``` ### Step 2: Choose the predictor A Python model registered with `feature_view=` needs no predictor script when it takes a feature vector of the feature view and is stored as a single pickle or joblib file. The library's default predictor validates the request against the deployment schema, looks up and transforms the features, runs the model, and logs the request when the feature view has logging enabled. === "Python" ```python my_model = mr.get_model("my_model", version=1) my_deployment = my_model.deploy(passed_features=["amount"]) my_deployment.schema.describe() ``` To customise it, subclass it in your own script and deploy with `default_predictor=True`, so the schema is still inferred: === "Python" ```python from hsml.default_predictor import DefaultPredict class Predict(DefaultPredict): def load_model(self, model_files_path): # anything the default loader does not handle ... def model_predict(self, feature_vectors): return self.model.predict_proba(feature_vectors[self.model_input_columns]) ``` The serving wrapper imports a model deployment's script itself, so the script needs no `__main__` block. Only a [feature view deployment][feature-view-deployment] script, which may be started as a plain script, hands over to the wrapper. See the [Deployment Schema Guide][deployment-schema] for the request contract, the error codes, and the feature logging guarantees. !!! note "The default predictor's `predict` is a coroutine" `DefaultPredict.predict` is `async def`, and it awaits the online store lookup rather than blocking on it. The lookup is a round trip, so on a blocking predictor it held the server's event loop and stopped every other request in the deployment for its duration: measured, that capped a deployment at 218 requests per second where the same deployment with nothing to look up reached 310. Awaiting it raised throughput by 23.6 percent and cut p99 latency by 72 percent. A subclass overriding `model_predict` or `load_model` is unaffected, since neither is a coroutine. A subclass overriding `predict` itself must declare it `async def` and `await super().predict(...)`. To drive the predictor from a script or a notebook, where there is no event loop to hold up, call `predict_blocking` instead. It serves the same request through the same body and returns the same value; called from inside a running loop it refuses rather than deadlocks. ```python predictor = Predict() result = predictor.predict_blocking([{"cc_num": 1234}]) ``` Set `SERVING_PREDICTOR_ASYNC_LOOKUP=false` on the deployment to go back to the blocking lookup. That is worth doing only when the deployment reads the online store through the REST client, where there is nothing to overlap. To serve the model with your own code instead, implement a predictor script (Steps 2.1 and 2.2). ### Step 2.1 (Optional): Implement a predictor script For Python model deployments that the default predictor does not cover, implement a predictor script that loads and serves your model. A script deployed with `schema=` or `passed_features=` still gets every request validated against the deployment schema by the serving wrapper before `predict()` is called. === "Predictor" ``` python class Predictor: def __init__(self): """Initialization code goes here""" # Optional __init__ params: project, deployment, model, async_logger # Model files can be found at os.environ["MODEL_FILES_PATH"] # self.model = ... # load your model def predict(self, inputs): """Serve predictions using the trained model""" # Use the model to make predictions # return self.model.predict(inputs) ``` === "Predictor with Feature Logging" ``` python class Predictor: def __init__(self, async_logger, model, project): """Initializes the serving state, reads a trained model""" # Get feature view attached to model ## self.model = model ## self.feature_view = model.get_feature_view() # Initialize feature view with async feature logger ## self.feature_view.init_feature_logger(feature_logger=async_logger) def predict(self, inputs): """Serves a prediction request usign a trained model""" # Extract serving keys and request parameters from inputs ## serving_keys = ... ## request_parameters = ... # Fetch feature vector with logging metadata ## vector = self.feature_view.get_feature_vector(serving_keys, ## request_parameters=request_parameters, ## logging_data=True) # Make predictions ## predictions = model.predict(vector) # Log Predictions ## self.feature_view.log(vector, ## predictions=predictions, ## model = self.model) # Predictions ## return predictions ``` === "Async Predictor" ``` python class Predictor: def __init__(self): """Initialization code goes here""" # Optional __init__ params: project, deployment, model, async_logger # Model files can be found at os.environ["MODEL_FILES_PATH"] # self.model = ... # load your model async def predict(self, inputs): """Asynchronously serve predictions using the trained model""" # Perform async operations that required # result = await some_async_preprocessing(inputs) # Use the model to make predictions # return self.model.predict(result) ``` !!! tip "Optional `__init__` parameters" The `__init__` method supports optional parameters that are automatically injected at runtime: | Parameter | Class | Description | | -------------- | -------------------- | ------------------------------------------------------ | | `project` | `Project` | Hopsworks project handle | | `deployment` | `Deployment` | Current model deployment handle | | `model` | `Model` | Model handle | | `async_logger` | `AsyncFeatureLogger` | Async feature logger for logging features to Hopsworks | You can add any combination of these parameters to your `__init__` method: ```python class Predictor: def __init__(self, project, model): # Access the project and model directly self.project = project self.model = model ``` !!! tip "Feature logging" The `async_logger` parameter enables asynchronous logging of features and predictions from your predictor script via the Feature View API. This is useful for debugging, monitoring, and auditing the data your models use in production. Logged features are periodically materialized to the offline feature store, and can be retrieved, filtered, and managed through the feature view. See the [Feature and Prediction Logging](../../fs/feature_view/feature_logging.md) guide for details on enabling logging, retrieving logs, and managing the log lifecycle. !!! info "Jupyter magic" In a jupyter notebook, you can add `%%writefile my_predictor.py` at the top of the cell to save it as a local file. ### Step 2.2 (Optional): Upload the script to your project !!! info "You can also use the UI to upload your predictor script. See [above](#step-3-optional-select-a-predictor-script)" === "Python" ```python uploaded_file_path = dataset_api.upload( "my_predictor.py", "Resources", overwrite=True ) predictor_script_path = os.path.join( "/Projects", project.name, uploaded_file_path ) ``` ### Step 3: Pass predictor configuration to model deployment You can customize the default predictor settings when creating a model deployment. === "Python" ```python my_model = mr.get_model("my_model", version=1) my_deployment = my_model.deploy( # predictor configuration model_server="PYTHON", script_file=predictor_script_path, ) ``` !!! api "API reference" - [`Predictor`][hsml.predictor.Predictor] - [`Model.deploy`][hsml.model.Model.deploy] - [`Model.get_feature_view`][hsml.model.Model.get_feature_view] - [`FeatureView`][hsfs.feature_view.FeatureView] - [`init_feature_logger`][hsfs.feature_view.FeatureView.init_feature_logger] - [`get_feature_vector`][hsfs.feature_view.FeatureView.get_feature_vector] - [`log`][hsfs.feature_view.FeatureView.log] Browse the full Python API :material-arrow-right: ## Model Server Hopsworks Model Serving supports deploying models with a Python model server for python-based models (scikit-learn, XGBoost , pytorch...), TensorFlow Serving for TensorFlow / Keras models and vLLM for Large Language Models (LLMs). !!! info "Supported model servers" | Model Server | Backend | ML Models and Frameworks | | -------------------- | --------------------------------------------- | ------------------------------------------------------------------------------------------------ | | Python | Any `*-inference-pipeline` Python environment | Python-based (scikit-learn, XGBoost , pytorch...) | | KServe sklearnserver | Sklearn built-in KServe runtime | Scikit-learn, XGBoost | | TensorFlow Serving | TensorFlow Serving runtime | Keras, TensorFlow | | vLLM | vLLM openai-compatible server | vLLM-supported models (see [list](https://docs.vllm.ai/en/v0.28.0/models/supported_models.html)) | !!! note "vLLM variants" The vLLM model server is available in two variants, standard **vLLM** and **vLLM-Omni**, and the image version can be pinned per deployment. See [vLLM variant and image version](#vllm-variant-and-image-version). Each model server has specific requirements and supports different types of model artifacts, file formats, and configuration options. When deploying a model, ensure that your model files and configuration align with the expectations of the selected server. !!! info "Model artifact requirements" | Model server | Model files | | -------------------- | --------------------------------------------------------- | | Python | Any model file format | | KServe sklearnserver | Files with extensions `.joblib`, `.pkl`, `.pickle` | | TensorFlow Serving | Model artifact needs `variables/` and `.pb` file | | vLLM | Model files supported by vLLM engine (e.g., .safetensors) | All deployments use [KServe](https://kserve.github.io/website/latest/) as the serving platform, providing autoscaling, fine-grained resource allocation, inference logging, inference batching, and transformers. KServe runs a deployment in either Knative or Standard mode, see [Deployment mode](autoscaling.md#deployment-mode). ## Predictor script For **Python model deployments** ==only==, you can provide a custom Python script, called a predictor script, to load your model and serve predictions. This script is included in the [artifact files](../serving/deployment.md#artifact-files) of the deployment. The script must follow a specific template, as shown in [Step 2](#step-21-optional-implement-a-predictor-script). ## Server configuration file For **Python model deployments**, you can provide a server configuration file to separate deployment-specific settings from the logic in your predictor or transformer scripts. This approach allows you to update configuration parameters without modifying the code. Within the deployment, the configuration file is accessible at the path specified by the `CONFIG_FILE_PATH` environment variable (see [environment variables](#environment-variables)). For **vLLM deployments**, the server configuration file is ==required== and is used to configure the vLLM server. For example, you can use this configuration file to specify the chat template or LoRA modules to be loaded by the vLLM server. See all available parameters in the [official documentation](https://docs.vllm.ai/en/v0.28.0/serving/openai_compatible_server.html). !!! warning "Configuration file format" The configuration file can be of any format, except in **vLLM deployments** for which a YAML file (`.yml`/`.yaml`) is ==required==. When a predictor script is provided, any format is allowed as users can load it as necessary. ## vLLM variant and image version For **vLLM deployments**, you can choose which **variant** of vLLM to run and, optionally, which **image version** to use. Both settings are stored on the deployment and can be edited later. They are first-class fields on the predictor configuration. ### Variant | Variant | Runtime | When to use | | ---------------- | --------------------------------------------------------------------------------------- | ---------------------------------------------------------------------- | | `VLLM` (default) | Standard `vllm-openai` server ([docker hub](https://hub.docker.com/r/vllm/vllm-openai)) | OpenAI-compatible inference with the upstream vLLM engine. | | `VLLM_OMNI` | Standard `vllm-omni` server ([docker hub](https://hub.docker.com/r/vllm/vllm-omni)) | OpenAI-compatible multi-modal inference with the upstream vLLM engine. | Both the `VLLM` and `VLLM_OMNI` variants run the official upstream vLLM images published on [Docker Hub](https://hub.docker.com): [`vllm/vllm-openai`](https://hub.docker.com/r/vllm/vllm-openai) and [`vllm/vllm-omni`](https://hub.docker.com/r/vllm/vllm-omni) respectively. !!! tip "Python SDK" From the Python SDK, the variant is selected with the `vllm_variant` keyword argument. See [`Model.deploy()`][hsml.model.Model.deploy] in the API reference. ### Image version The image version is the runtime image tag (for example `v0.28.0`) used for the vLLM container. If not set, Hopsworks picks the **highest** image version advertised for the chosen variant at creation time. Set it explicitly to pin a deployment to a specific image, which is useful when you need a stable, reproducible runtime. The list of available versions per variant is advertised by the cluster administrator through the `kube_serving_vllm_versions` and `kube_serving_vllm_omni_versions` Hopsworks variables. Hopsworks ships exactly one default image version for each of the **vLLM** and **vLLM-Omni** variants. Any additional versions exposed in the **Version** dropdown are managed by the cluster administrator. If a version you need is not listed, contact your administrator. !!! warning "Cluster registry access" For any additional image version advertised in the **Version** dropdown to actually start, the official `docker.io` registry and the `vllm/vllm-openai` / `vllm/vllm-omni` repositories must be pullable from inside the cluster. If the cluster has no egress to Docker Hub, ask your administrator to either grant access or mirror the upstream tags into the cluster's internal registry under the same repository path, otherwise the deployment will fail to pull the image. !!! note "Upstream tags" Hopsworks does not repackage these images: the tag you select in the **Version** dropdown is the upstream ==Docker Hub tag==. !!! tip "Python SDK" From the Python SDK, the image version is selected with the `vllm_image_tag` keyword argument. See [`Model.deploy()`][hsml.model.Model.deploy] in the API reference. ## Environment variables Two kinds of environment variables are available in the predictor: user-defined variables that you set per deployment or at the account level, and built-in variables that Hopsworks sets automatically at runtime. ### User-defined environment variables In the advanced deployment form, the *Predictor environment variables* section is pre-filled with your account-level [Environment variables](../../projects/env_vars/create.md). Values you change in the form become per-deployment overrides for that deployment only. Account-level values for names you don't include in the form continue to apply at runtime. !!! info "Account-level variables also apply" Variables defined under [Account settings → Environment variables](../../projects/env_vars/create.md) are injected into every model deployment you start. A value set on the deployment overrides the account-level value with the same name for that deployment only. ### Built-in environment variables A number of different environment variables are available in the predictor to ease its implementation. !!! warning Built-in environment variables cannot be overwritten as they are meant for internal use. !!! tip "Available environment variables" === "Deployment" These variables are available in all deployments. | Name | Description | | --------------------- | -------------------------------- | | `DEPLOYMENT_NAME` | Name of the current deployment | | `DEPLOYMENT_VERSION` | Version of the deployment | | `ARTIFACT_FILES_PATH` | Local path to the artifact files | === "Predictor" These variables are set for predictor components. | Name | Description | | ------------------ | -------------------------------------------------- | | `SCRIPT_PATH` | Full path to the predictor script | | `SCRIPT_NAME` | Prefixed filename of the predictor script | | `CONFIG_FILE_PATH` | Local path to the configuration file (if provided) | | `IS_PREDICTOR` | Set to `true` for predictor components | === "Model" | Name | Description | | ------------------ | ---------------------------------------------------------------- | | `MODEL_FILES_PATH` | Local path to the model files (`/var/lib/hopsworks/model_files`) | | `MODEL_NAME` | Name of the model being served by the current deployment | | `MODEL_VERSION` | Version of the model being served by the current deployment | === "Others" These variables are available in all deployments. | Name | Description | | ------------------------ | -------------------------------------------------- | | `REST_ENDPOINT` | Hopsworks REST API endpoint | | `HOPSWORKS_PROJECT_ID` | ID of the project | | `HOPSWORKS_PROJECT_NAME` | Name of the project | | `HOPSWORKS_PUBLIC_HOST` | Hopsworks public hostname | | `API_KEY` | API key for authenticating with Hopsworks services | | `PROJECT_ID` | Project ID (for Feature Store access) | | `PROJECT_NAME` | Project name (for Feature Store access) | | `SECRETS_DIR` | Path to secrets directory (`/keys`) | | `MATERIAL_DIRECTORY` | Path to TLS certificates (`/certs`) | | `REQUESTS_VERIFY` | SSL verification setting | ## Python environments Based on the model server used in the model deployment, you can select the Python environment where the predictor and transformer scripts will run. The predictor and the transformer each run their own environment, selected independently. To create a new Python environment see [Python Environments](../../projects/python/python_env_overview.md). !!! info "Supported Python environments" | Model server | Predictor | Transformer | | -------------------- | -------------------------------- | -------------------------------- | | Python | any `*-inference-pipeline` image | any `*-inference-pipeline` image | | KServe sklearnserver | `sklearnserver` | any `*-inference-pipeline` image | | TensorFlow Serving | `tensorflow/serving` | any `*-inference-pipeline` image | | vLLM | `vllm-openai` | Not supported | A transformer that does not name an environment runs the predictor's. When the predictor runs a fixed runtime image, such as TensorFlow Serving or vLLM, it has no environment of its own and the transformer runs the one the deployment names. === "Python" ```python ms = project.get_model_serving() transformer = ms.create_transformer( script_file="my_transformer.py", environment="minimal-inference-pipeline", ) deployment = my_model.deploy( name="mydeployment", script_file="my_predictor.py", environment="pandas-inference-pipeline", transformer=transformer, ) ``` !!! note Deployments created before Hopsworks 5.2 run both components on the environment they were configured with. Change either one to move that component on its own. ## Transformer script Transformer scripts are Python scripts used to apply transformations on the model inputs before sending them to the predictor for making predictions using the model. To learn more about transformers, see the [Transformer (KServe) Guide](transformer.md). !!! note Transformer scripts are ==not== supported in **vLLM deployments** ## Inference logger Inference loggers are deployment components that log inference requests into a Kafka topic for later analysis. To learn about the different logging modes, see the [Inference Logger Guide](inference-logger.md) ## Inference batcher Inference batcher are deployment component that apply batching to the incoming inference requests for a better throughput-latency trade-off. To learn about the different configuration available for the inference batcher, see the [Inference Batcher Guide](inference-batcher.md). ## Resources Resources include the number of replicas for the deployment as well as the resources (i.e., memory, CPU, GPU) to be allocated per replica. To learn about the different combinations available, see the [Resources Guide](resources.md). ## Autoscaling Deployments can automatically scale the number of replicas, and how they scale depends on the [deployment mode](autoscaling.md#deployment-mode). To learn about the deployment modes and the different autoscaling parameters, see the [Autoscaling Guide](autoscaling.md). ## Scheduling !!! info "Kueue is required" This feature requires Kueue to be enabled in your cluster. If Kueue is not available, queue and topology options will not be accessible. If the cluster has Kueue enabled, you can select a queue for your deployment from the advanced configuration. Queues control resource allocation and scheduling priority across the cluster. For full details on scheduling configuration, see the [Scheduling Guide](scheduling.md). ## API protocol Depending on the model server, Hopsworks supports both REST and gRPC as the API protocols to send inference requests to model deployments. In general, you use gRPC when you need lower latency inference requests. To learn more about the REST and gRPC API protocols for model deployments, see the [API Protocol Guide](api-protocol.md). ================================================================================ # Transformer (KServe) Source: https://docs.hopsworks.ai/latest/user_guides/mlops/serving/transformer/ # How To Configure A Transformer ## Introduction In this guide, you will learn how to configure a transformer script in a model deployment. Transformer scripts are used to apply transformations on the model inputs before sending them to the predictor for making predictions using the model. They are user-provided Python scripts (`.py` or `.ipynb`) implementing the [Transformer class](#step-2-implement-transformer-script). !!! info "Transformer scripts are not supported in vLLM deployments." !!! tip "Independent scaling" The transformer has independent resources and autoscaling configuration from the predictor. This allows you to scale the pre/post-processing separately from the model inference. A transformer has the following configurable components: !!! info "" 1. [Transformer script](#transformer-script) 2. [Resources](#resources) 3. [Autoscaling](#autoscaling) 4. [Python environments](#python-environments) 5. [Environment variables](#environment-variables) See examples of transformer scripts in the serving [example notebooks](https://github.com/logicalclocks/hops-examples/blob/master/notebooks/ml/serving). ## Web UI ### Step 1: Create new deployment If you have at least one model already trained and saved in the Model Registry, navigate to the deployments page by clicking on the `Deployments` tab on the navigation menu on the left.

Deployments navigation tab
Deployments navigation tab

Once in the deployments page, you can create a new deployment by either clicking on `New deployment` (if there are no existing deployments) or on `Create new deployment` it the top-right corner. Both options will open the deployment creation form. ### Step 2: Go to advanced options A simplified creation form will appear including the most common deployment fields from all available configurations. Transformers are part of the advanced options of a deployment. To navigate to the advanced creation form, click on `Advanced options`.

Advance options
Advanced options. Go to advanced deployment creation form

### Step 3: Select a transformer script If the transformer script is already located in Hopsworks, click on `From project` and navigate through the file system to find your script. Otherwise, you can click on `Upload new file` to upload the transformer script now.

Transformer script in advanced deployment form
Choose a transformer script in the advanced deployment form

After selecting the transformer script, you can optionally configure resources and autoscaling for your transformer (see [Step 4](#step-4-optional-other-advanced-options)). Otherwise, click on `Create new deployment` to create the deployment for your model. ### Step 4 (Optional): Other advanced options In this page, you can also configure the [Python environment](#python-environments) the transformer runs in, the [resources](resources.md) to be allocated for the transformer, as well as the [autoscaling](autoscaling.md) parameters to control how the transformer scales based on traffic. The transformer has its own environment field, separate from the predictor's.

Resource allocation for the transformer
Resource allocation for the transformer

Once you are done with the changes, click on `Create new deployment` at the bottom of the page to create the deployment for your model. ## Code ### Step 1: Connect to Hopsworks === "Python" ```python import hopsworks project = hopsworks.login() # get Dataset API instance dataset_api = project.get_dataset_api() # get Hopsworks Model Registry handle mr = project.get_model_registry() ``` ### Step 2: Implement transformer script === "Transformer" ```python class Transformer: def __init__(self): """Initialization code goes here""" # Optional __init__ params: project, deployment, model, async_logger pass def preprocess(self, inputs): """Transform the requests inputs here. The object returned by this method will be used as model input to make predictions.""" return inputs def postprocess(self, outputs): """Transform the predictions computed by the model before returning a response""" return outputs ``` !!! tip "Optional `__init__` parameters" The `__init__` method supports optional parameters that are automatically injected at runtime: | Parameter | Class | Description | | -------------- | -------------------- | ------------------------------------------------------ | | `project` | `Project` | Hopsworks project handle | | `deployment` | `Deployment` | Current model deployment handle | | `model` | `Model` | Model handle | | `async_logger` | `AsyncFeatureLogger` | Async feature logger for logging features to Hopsworks | You can add any combination of these parameters to your `__init__` method: ```python class Transformer: def __init__(self, project, model): # Access the project and model directly self.project = project self.model_metadata = model ``` !!! info "Jupyter magic" In a jupyter notebook, you can add `%%writefile my_transformer.py` at the top of the cell to save it as a local file. ### Step 3: Upload the script to your project !!! info "You can also use the UI to upload your transformer script. See [above](#step-3-select-a-transformer-script)" === "Python" ```python uploaded_file_path = dataset_api.upload( "my_transformer.py", "Resources", overwrite=True ) transformer_script_path = os.path.join( "/Projects", project.name, uploaded_file_path ) ``` ### Step 4: Define a transformer === "Python" ```python my_transformer = ms.create_transformer(script_file=uploaded_file_path) # or from hsml.transformer import Transformer my_transformer = Transformer(script_file) ``` ### Step 5: Create a deployment with the transformer Use the `transformer` parameter to set the transformer configuration when creating the model deployment. === "Python" ```python my_model = mr.get_model("my_model", version=1) my_deployment = my_model.deploy( transformer=my_transformer ) ``` !!! api "API reference" - [`Transformer`][hsml.transformer.Transformer] - [`ModelServing.create_transformer`][hsml.model_serving.ModelServing.create_transformer] - [`Model.deploy`][hsml.model.Model.deploy] - [`DatasetApi.upload`][hopsworks_common.core.dataset_api.DatasetApi.upload] Browse the full Python API :material-arrow-right: ## Transformer script A transformer script is a custom Python script to apply pre/post-processing on the model inputs and outputs. This script is included in the [artifact files](../serving/deployment.md#artifact-files) of the deployment. The script must implement the `Transformer` class, as shown in [Step 2](#step-2-implement-transformer-script). !!! info "Transformer scripts are not supported in vLLM deployments." ## Resources Resources include the number of replicas for the deployment as well as the resources (i.e., memory, CPU, GPU) to be allocated per replica. To learn about the different combinations available, see the [Resources Guide](resources.md). ## Autoscaling The transformer has independent autoscaling from the predictor. How it scales depends on the [deployment mode](autoscaling.md#deployment-mode) of the deployment: on request traffic including scale-to-zero in Knative mode, and on CPU or memory utilization with at least one instance running in Standard mode. To learn about the different autoscaling parameters, see the [Autoscaling Guide](autoscaling.md). ## Environment variables Two kinds of environment variables are available in the transformer: user-defined variables that you set per deployment or at the account level, and built-in variables that Hopsworks sets automatically at runtime. ### User-defined environment variables In the advanced deployment form, the *Transformer environment variables* section lets you define environment variables specific to the transformer component. Values you set in the form become per-deployment overrides for that deployment only. Account-level [Environment variables](../../projects/env_vars/create.md) for names you don't include in the form continue to apply at runtime. !!! info "Account-level variables also apply" Variables defined under [Account settings → Environment variables](../../projects/env_vars/create.md) are injected into every model deployment you start. A value set on the deployment overrides the account-level value with the same name for that deployment only. ### Built-in environment variables A number of different environment variables are available in the transformer to ease its implementation. !!! warning Built-in environment variables cannot be overwritten as they are meant for internal use. !!! tip "Available environment variables" === "Deployment" These variables are available in all deployments. | Name | Description | | --------------------- | -------------------------------- | | `DEPLOYMENT_NAME` | Name of the current deployment | | `DEPLOYMENT_VERSION` | Version of the deployment | | `ARTIFACT_FILES_PATH` | Local path to the artifact files | === "Transformer" These variables are set for transformer components. | Name | Description | | ------------------ | -------------------------------------------------- | | `SCRIPT_PATH` | Full path to the transformer script | | `SCRIPT_NAME` | Prefixed filename of the transformer script | | `CONFIG_FILE_PATH` | Local path to the configuration file (if provided) | | `IS_TRANSFORMER` | Set to `true` for transformer components | === "Model" | Name | Description | | --------------- | ----------------------------------------------------------- | | `MODEL_NAME` | Name of the model being served by the current deployment | | `MODEL_VERSION` | Version of the model being served by the current deployment | === "Others" These variables are available in all deployments. | Name | Description | | ------------------------ | -------------------------------------------------- | | `REST_ENDPOINT` | Hopsworks REST API endpoint | | `HOPSWORKS_PROJECT_ID` | ID of the project | | `HOPSWORKS_PROJECT_NAME` | Name of the project | | `HOPSWORKS_PUBLIC_HOST` | Hopsworks public hostname | | `API_KEY` | API key for authenticating with Hopsworks services | | `PROJECT_ID` | Project ID (for Feature Store access) | | `PROJECT_NAME` | Project name (for Feature Store access) | | `SECRETS_DIR` | Path to secrets directory (`/keys`) | | `MATERIAL_DIRECTORY` | Path to TLS certificates (`/certs`) | | `REQUESTS_VERIFY` | SSL verification setting | ## Python environments Transformer scripts always run on `*-inference-pipeline` Python environments. The transformer's environment is selected independently of the predictor's, so the two components can run different sets of dependencies. A transformer that does not name one runs the predictor's environment. When the predictor runs a fixed runtime image, such as TensorFlow Serving or vLLM, it has no environment of its own and the transformer runs the one the deployment names. To create a new Python environment see [Python Environments](../../projects/python/python_env_overview.md). === "Python" ```python ms = project.get_model_serving() transformer = ms.create_transformer( script_file="my_transformer.py", environment="minimal-inference-pipeline", ) ``` !!! info "Supported Python environments" | Model server | Predictor | Transformer | | -------------------- | -------------------------------- | -------------------------------- | | Python | any `*-inference-pipeline` image | any `*-inference-pipeline` image | | KServe sklearnserver | `sklearnserver` | any `*-inference-pipeline` image | | TensorFlow Serving | `tensorflow/serving` | any `*-inference-pipeline` image | | vLLM | `vllm-openai` | Not supported | ================================================================================ # Inference Logger Source: https://docs.hopsworks.ai/latest/user_guides/mlops/serving/inference-logger/ # How To Configure Inference Logging ## Introduction Once a model is deployed and starts making predictions as inference requests arrive, logging model inputs and predictions becomes essential to monitor the health of the model and take action if the model's performance degrades over time. Hopsworks supports logging both inference requests and predictions as events to a Kafka topic for analysis. !!! warning "Inference logging is not supported for vLLM deployments." !!! info "Inference logger vs. feature logging" The inference logger described here stores model inputs and predictions from inference requests and responses into a Kafka topic, for later consumption and analysis. It is separate from [feature logging](../../fs/feature_view/feature_logging.md), which supports more fine-grained logging of inference logs and features and powers [feature monitoring](../../fs/feature_monitoring/index.md) and [model monitoring](../model_monitoring/index.md). !!! info "Logging modes" Three logging modes are available: | Mode | Logger Mode | Description | | ------------ | ----------- | --------------------------- | | ALL | `all` | Log both inputs and outputs | | PREDICTIONS | `response` | Log model outputs only | | MODEL_INPUTS | `request` | Log model inputs only | !!! note "Kafka topic requirements" The Kafka topic must use the `inferenceschema` subject. Schema v4+ is required for KServe topics. ## Web UI ### Step 1: Create new deployment If you have at least one model already trained and saved in the Model Registry, navigate to the deployments page by clicking on the `Deployments` tab on the navigation menu on the left.

Deployments navigation tab
Deployments navigation tab

Once in the deployments page, you can create a new deployment by either clicking on `New deployment` (if there are no existing deployments) or on `Create new deployment` it the top-right corner. Both options will open the deployment creation form. ### Step 2: Go to advanced options A simplified creation form will appear including the most common deployment fields from all available configurations. Inference logging is part of the advanced options of a deployment. To navigate to the advanced creation form, click on `Advanced options`.

Advance options
Advanced options. Go to advanced deployment creation form

### Step 3: Configure inference logging To enable inference logging, choose `CREATE` as Kafka topic name to create a new topic, or select an existing topic. If you prefer, you can disable inference logging by selecting `NONE`. If you decide to create a new topic, select the number of partitions and number of replicas for your topic, or use the default values.

Inference logger in advanced deployment form
Inference logging configuration with a new kafka topic

If the deployment is created with KServe enabled, you can specify which inference logs you want to send to the Kafka topic (i.e., `MODEL_INPUTS`, `PREDICTIONS` or both) Once you are done with the changes, click on `Create new deployment` at the bottom of the page to create the deployment for your model. ## Code ### Step 1: Connect to Hopsworks === "Python" ```python import hopsworks project = hopsworks.login() # get Hopsworks Model Registry handle mr = project.get_model_registry() # get Hopsworks Model Serving handle ms = project.get_model_serving() ``` ### Step 2: Define an inference logger === "Python" ```python from hsml.inference_logger import InferenceLogger from hsml.kafka_topic import KafkaTopic new_topic = KafkaTopic( name="CREATE", # optional num_partitions=1, num_replicas=1, ) my_logger = InferenceLogger(kafka_topic=new_topic, mode="ALL") ``` !!! notice "Use dict for simpler code" Similarly, you can create the same logger with: ```python my_logger = InferenceLogger(kafka_topic={"name": "CREATE"}, mode="ALL") ``` ### Step 3: Create a deployment with the inference logger === "Python" ```python my_model = mr.get_model("my_model", version=1) my_model.deploy(inference_logger=my_logger) ``` !!! api "API reference" - [`InferenceLogger`][hsml.inference_logger.InferenceLogger] - [`Model.deploy`][hsml.model.Model.deploy] Browse the full Python API :material-arrow-right: ## Topic schema Model inputs and predictions are logged in separate events, sharing the same `requestId` field. !!! example "Kafka topic schema" ``` json { "fields": [ { "name": "servingId", "type": "int" }, { "name": "modelName", "type": "string" }, { "name": "modelVersion", "type": "int" }, { "name": "requestTimestamp", "type": "long" }, { "name": "responseHttpCode", "type": "int" }, { "name": "inferenceId", "type": "string" }, { "name": "messageType", "type": "string" }, { "name": "payload", "type": "string" } ], "name": "inferencelog", "type": "record" } ``` ================================================================================ # Inference Batcher Source: https://docs.hopsworks.ai/latest/user_guides/mlops/serving/inference-batcher/ # How To Configure Inference Batcher ## Introduction Inference batching can be enabled to increase inference request throughput at the cost of higher latencies. The configuration of the inference batcher depends on the model server used in the deployment. !!! warning "Inference batching is not supported for vLLM deployments." ## Web UI ### Step 1: Create new deployment If you have at least one model already trained and saved in the Model Registry, navigate to the deployments page by clicking on the `Deployments` tab on the navigation menu on the left.

Deployments navigation tab
Deployments navigation tab

Once in the deployments page, you can create a new deployment by either clicking on `New deployment` (if there are no existing deployments) or on `Create new deployment` it the top-right corner. Both options will open the deployment creation form. ### Step 2: Go to advanced options A simplified creation form will appear including the most common deployment fields from all available configurations. Inference batching is part of the advanced options of a deployment. To navigate to the advanced creation form, click on `Advanced options`.

Advance options
Advanced options. Go to advanced deployment creation form

### Step 3: Configure inference batching To enable inference batching, click on the `Request batching` checkbox.

Inference batcher in advanced deployment form
Inference batching configuration (default values)

If your deployment uses KServe, you can optionally set three additional parameters for the inference batcher: maximum batch size, maximum latency (ms) and timeout (s). !!! note "Timeout parameter" The `timeout` parameter sets the request timeout in seconds for the inference batcher. If a batch is not filled within this time, the available requests are sent as a partial batch. Once you are done with the changes, click on `Create new deployment` at the bottom of the page to create the deployment for your model. ## Code ### Step 1: Connect to Hopsworks === "Python" ```python import hopsworks project = hopsworks.login() # get Hopsworks Model Registry handle mr = project.get_model_registry() # get Hopsworks Model Serving handle ms = project.get_model_serving() ``` ### Step 2: Define an inference logger === "Python" ```python from hsml.inference_batcher import InferenceBatcher my_batcher = InferenceBatcher( enabled=True, # optional max_batch_size=32, max_latency=5000, # milliseconds timeout=5, # seconds ) ``` ### Step 3: Create a deployment with the inference batcher === "Python" ```python my_model = mr.get_model("my_model", version=1) my_model.deploy(inference_batcher=my_batcher) ``` !!! api "API reference" - [`InferenceBatcher`][hsml.inference_batcher.InferenceBatcher] - [`Model.deploy`][hsml.model.Model.deploy] Browse the full Python API :material-arrow-right: ================================================================================ # Resources Source: https://docs.hopsworks.ai/latest/user_guides/mlops/serving/resources/ # How To Allocate Resources To A Model Deployment ## Introduction Resource allocation can be configured ==per component== (predictor and transformer) in a deployment, allowing you to specify how many CPUs, GPUs, and memory are allocated. For each component, you can set minimum (requests) and maximum (limits) resources, as well as the number of instances. ??? info "Resource defaults" | Field | Default Request | Default Limit | Validation | | ------------------ | --------------- | -------------- | -------------------------------------------------- | | CPU (cores) | 0.2 | -1 (unlimited) | Request cannot exceed limit (unless -1, unlimited) | | Memory (MB) | 32 | -1 (unlimited) | Request cannot exceed limit (unless -1, unlimited) | | GPUs | 0 | 0 | Request must equal limit | | Shared Memory (MB) | 128 | n/a | n/a | !!! tip "Automatic downscale of inactive instances" Setting the number of instances to **0** for a component (predictor or transformer) enables **scale-to-zero**. This means that all instances of the component will automatically scale down to zero after a default period of inactivity of 30 seconds. Scale-to-zero is available in ==Knative== mode only. A deployment in ==Standard== mode always keeps at least one instance of each component running, see [Deployment mode](autoscaling.md#deployment-mode). ## Web UI ### Step 1: Create new deployment If you have at least one model already trained and saved in the Model Registry, navigate to the deployments page by clicking on the `Deployments` tab on the navigation menu on the left.

Deployments navigation tab
Deployments navigation tab

Once in the deployments page, you can create a new deployment by either clicking on `New deployment` (if there are no existing deployments) or on `Create new deployment` it the top-right corner. Both options will open the deployment creation form. ### Step 2: Go to advanced options A simplified creation form will appear including the most common deployment fields from all available configurations. Resource allocation is part of the advanced options of a deployment. To navigate to the advanced creation form, click on `Advanced options`.

Advance options
Advanced options. Go to advanced deployment creation form

### Step 3: Configure resources In the `Resource allocation` section of the form, you can optionally set the resources to be allocated to the predictor and/or the transformer (if available). Moreover, you can choose the minimum number of replicas for each of these components. !!! note "Scale-to-zero capabilities" Set the number of instances to **0** to enable scale-to-zero on the component. This requires the deployment to run in ==Knative== mode, see [Deployment mode](autoscaling.md#deployment-mode).

Resource allocation for the predictor and transformer components
Resource allocation for the predictor and transformer

Once you are done with the changes, click on `Create new deployment` at the bottom of the page to create the deployment for your model. ## Code ### Step 1: Connect to Hopsworks === "Python" ```python import hopsworks project = hopsworks.login() # get Hopsworks Model Registry handle mr = project.get_model_registry() # get Hopsworks Model Serving handle ms = project.get_model_serving() ``` ### Step 2: Define the predictor resource configuration === "Python" ```python from hsml.resources import PredictorResources, Resources minimum_res = Resources(cores=1, memory=128, gpus=1) maximum_res = Resources(cores=2, memory=256, gpus=1) predictor_res = PredictorResources( num_instances=1, requests=minimum_res, limits=maximum_res ) ``` ### Step 3 (Optional): Define the transformer resource configuration === "Python" ```python from hsml.resources import TransformerResources minimum_res = Resources(cores=1, memory=128, gpus=1) maximum_res = Resources(cores=2, memory=256, gpus=1) transformer_res = TransformerResources( num_instances=2, requests=minimum_res, limits=maximum_res ) ``` ### Step 4: Create a deployment with the resource configuration === "Python" ```python my_model = mr.get_model("my_model", version=1) my_predictor = ms.create_predictor( my_model, resources=predictor_res, # transformer=Transformer(script_file, # resources=transformer_res) ) my_predictor.deploy() # or my_deployment = ms.create_deployment(my_predictor) my_deployment.save() ``` !!! api "API reference" - [`Resources`][hsml.resources.Resources] - [`ModelServing.create_predictor`][hsml.model_serving.ModelServing.create_predictor] - [`ModelServing.create_deployment`][hsml.model_serving.ModelServing.create_deployment] - [`Predictor.deploy`][hsml.predictor.Predictor.deploy] Browse the full Python API :material-arrow-right: ## Autoscaling Deployments can be configured to automatically scale the number of replicas, on traffic in Knative mode and on CPU or memory utilization in Standard mode. To learn about the different autoscaling parameters, see the [Autoscaling Guide](autoscaling.md). ================================================================================ # Autoscaling Source: https://docs.hopsworks.ai/latest/user_guides/mlops/serving/autoscaling/ # How To Configure Scaling For A Deployment ## Introduction This guide explains how to set up **autoscaling** for model deployments using either the [web UI](#web-ui) or the [Python API](#code). Autoscaling enables the deployment to use resources more efficiently, by growing and shrinking the allocated resources according to its actual, real-time usage. How a deployment scales depends on the mode it runs in. See [Deployment mode](#deployment-mode), [Scale metrics](#scale-metrics) and [Scaling parameters](#scaling-parameters) for details on the available scaling options in each mode. ## Web UI ### Step 1: Create new deployment If you have at least one model already trained and saved in the Model Registry, navigate to the deployments page by clicking on the `Deployments` tab on the navigation menu on the left.

Deployments navigation tab
Deployments navigation tab

Once in the deployments page, you can create a new deployment by either clicking on `New deployment` (if there are no existing deployments) or on `Create new deployment` it the top-right corner. Both options will open the deployment creation form. ### Step 2: Go to advanced options A simplified creation form will appear including the most common deployment fields from all available configurations. Autoscaling is part of the advanced options of a deployment. To navigate to the advanced creation form, click on `Advanced options`.

Advance options
Advanced options. Go to advanced deployment creation form

### Step 3: Configure autoscaling In the `Autoscaling` section of the advanced form, you can configure the scaling parameters for the predictor and/or the transformer (if available). The `Knative` checkbox at the top of the section selects the [deployment mode](#deployment-mode) and decides which of the fields below are shown. With it checked, you can set the scale metric, target value, minimum and maximum instances, as well as the panic and stable window parameters. With it cleared, the deployment runs in Standard mode and you can set the minimum and maximum instances plus a CPU or memory scale metric and its target. The checkbox is disabled while the deployment is running, because the mode of a running deployment cannot be changed.

Autoscaling configuration for the predictor and transformer components in Knative mode
Autoscaling configuration for the predictor and transformer in Knative mode

Autoscaling configuration for the predictor and transformer components in Standard mode
Autoscaling configuration for the predictor and transformer in Standard mode

Once you are done with the changes, click on `Create new deployment` at the bottom of the page to create the deployment for your model. ## Code ### Step 1: Connect to Hopsworks === "Python" ```python import hopsworks project = hopsworks.login() # get Hopsworks Model Registry handle mr = project.get_model_registry() # get Hopsworks Model Serving handle ms = project.get_model_serving() ``` ### Step 2: Define the predictor scaling configuration You can use the [`PredictorScalingConfig`][hsml.scaling_config.PredictorScalingConfig] class to configure the scaling options according to your preferences. Default values for scaling metrics and parameters are listed in the [Scale metrics](#scale-metrics) and [Scaling parameters](#scaling-parameters) sections above. Make sure the metric and parameters you set are valid for the [deployment mode](#deployment-mode) you deploy in. === "Knative mode" ```python from hsml.scaling_config import PredictorScalingConfig predictor_scaling = PredictorScalingConfig( min_instances=1, max_instances=5, scale_metric="RPS", target=100 ) ``` === "Standard mode" ```python from hsml.scaling_config import PredictorScalingConfig predictor_scaling = PredictorScalingConfig( min_instances=1, max_instances=5, scale_metric="CPU", target=80 ) ``` ### Step 3 (Optional): Define the transformer scaling configuration If a transformer script is also provided, you can use the [`TransformerScalingConfig`][hsml.scaling_config.TransformerScalingConfig] class to configure the scaling options according to your preferences. Default values for scaling metrics and parameters are listed in the [Scale metrics](#scale-metrics) and [Scaling parameters](#scaling-parameters) sections above. === "Knative mode" ```python from hsml.scaling_config import TransformerScalingConfig transformer_scaling = TransformerScalingConfig( min_instances=1, max_instances=3, scale_metric="CONCURRENCY", target=50 ) ``` === "Standard mode" ```python from hsml.scaling_config import TransformerScalingConfig transformer_scaling = TransformerScalingConfig( min_instances=1, max_instances=3, scale_metric="CPU", target=80 ) ``` ### Step 4: Create a deployment with the scaling configuration === "Knative mode" ```python my_model = mr.get_model("my_model", version=1) # optional my_transformer = ms.create_transformer( script_file="Resources/my_transformer.py", scaling_configuration=transformer_scaling, ) my_deployment = my_model.deploy( scaling_configuration=predictor_scaling, # optional: transformer=my_transformer, ) ``` === "Standard mode" ```python my_model = mr.get_model("my_model", version=1) # optional my_transformer = ms.create_transformer( script_file="Resources/my_transformer.py", scaling_configuration=transformer_scaling, ) my_deployment = my_model.deploy( scaling_configuration=predictor_scaling, knative_mode=False, # optional: transformer=my_transformer, ) ``` !!! note "Match the scaling configuration to the mode" The `knative_mode` argument selects the [deployment mode](#deployment-mode) of the deployment. Leaving it unset deploys in Knative mode for every model server except vLLM, so a Standard scaling configuration needs `knative_mode=False` passed explicitly. A scaling configuration that uses fields the chosen mode does not support is rejected, so set both together. !!! api "API reference" - [`PredictorScalingConfig`][hsml.scaling_config.PredictorScalingConfig] - [`TransformerScalingConfig`][hsml.scaling_config.TransformerScalingConfig] - [`ModelServing.create_transformer`][hsml.model_serving.ModelServing.create_transformer] - [`Model.deploy`][hsml.model.Model.deploy] Browse the full Python API :material-arrow-right: ## Deployment mode Every deployment runs on KServe in one of two modes. In ==Knative== mode, the deployment is backed by a Knative service that scales on request traffic and can scale to zero. In ==Standard== mode, the deployment is backed by a plain Kubernetes Deployment that always keeps at least one instance running and scales on CPU or memory utilization. Inference requests reach both modes through the same Istio endpoint, so the mode does not change how you send predictions to a deployment. !!! info "Knative and Standard mode" | Mode | Backed by | Scale-to-zero | Autoscaler | | -------- | --------------------- | ------------- | -------------------------------------------------------------------------------------------- | | Knative | Knative service | Yes | [Knative Pod Autoscaler (KPA)](https://knative.dev/docs/serving/autoscaling/) | | Standard | Kubernetes Deployment | No | [Kubernetes HPA](https://kubernetes.io/docs/tasks/run-application/horizontal-pod-autoscale/) | LLM deployments default to Standard mode, and deployments on every other model server default to Knative mode. LLMs load large model files and start slowly, so scaling them to zero costs a long cold start on the next request, and their throughput is bounded by GPU memory rather than by the number of concurrent requests. The two modes are backed by different Kubernetes resources, so the mode of a deployment cannot be changed while it is running. Stop the deployment, change the mode, and start it again. Switching the mode clears the scaling settings that do not carry over and re-derives the defaults of the new mode. !!! tip "Python SDK" From the Python SDK, the mode is selected with the `knative_mode` keyword argument: `True` for Knative mode, `False` for Standard mode. Leaving it unset uses the default for the model server when creating a deployment, and keeps the stored mode when updating one. See [`Model.deploy()`][hsml.model.Model.deploy] in the API reference. ## Scale metrics The metric a deployment can scale on depends on its [deployment mode](#deployment-mode), because the two modes are driven by different autoscalers. Setting a metric that the mode does not support is rejected. | Scale Metric | Mode | Default Target | Description | | ------------ | -------- | -------------- | ------------------------------- | | RPS | Knative | 200 | Requests per second per replica | | CONCURRENCY | Knative | 100 | Concurrent requests per replica | | CPU | Standard | 80 | CPU utilization percentage | | MEMORY | Standard | 80 | Memory utilization percentage | Knative deployments default to `CONCURRENCY`. See [Knative autoscaling metrics](https://knative.dev/docs/serving/autoscaling/autoscaling-metrics/) for more details on the Knative metrics. Standard deployments default to `CPU`, and scale between the minimum and maximum instances. Setting the minimum and maximum instances to the same value runs a fixed number of replicas instead, with no autoscaler and no scale metric. LLM deployments default to a fixed replica count, because every extra replica needs a GPU and GPU-bound pods scale poorly on CPU or memory utilization. ## Scaling parameters The following parameters can be used to fine-tune the autoscaling behavior. See [scale bounds](https://knative.dev/docs/serving/autoscaling/scale-bounds/), [autoscaling concepts](https://knative.dev/docs/serving/autoscaling/autoscaling-concepts/) and [scale-to-zero](https://knative.dev/docs/serving/autoscaling/scale-to-zero/) in the Knative documentation for more details. | Parameter | Mode | Default | Range | Description | | ----------------------------- | ------- | ------- | ------ | ------------------------------------------------------------- | | `minInstances` | Both | n/a | ≥ 0 | Minimum replicas (0 enables scale-to-zero, Knative mode only) | | `maxInstances` | Both | n/a | ≥ 1 | Maximum replicas (cannot be less than min) | | `panicWindowPercentage` | Knative | 10.0 | 1–100 | Panic window as percentage of stable window | | `stableWindowSeconds` | Knative | 60 | 6–3600 | Stable window duration in seconds | | `panicThresholdPercentage` | Knative | 200.0 | > 0 | Traffic threshold to trigger panic mode | | `scaleToZeroRetentionSeconds` | Knative | 0 | ≥ 0 | Time to retain pods before scaling to zero | The panic, stable window and scale-to-zero retention parameters are Knative Pod Autoscaler settings with no equivalent in a Kubernetes HPA, so a Standard deployment rejects them. A Standard deployment also requires at least one instance. !!! note "Cluster-level constraints" ==Administrators== can set cluster-wide limits on the maximum and minimum number of instances. When the minimum is set to 0, scale-to-zero is enforced for all Knative deployments. ================================================================================ # Scheduling Source: https://docs.hopsworks.ai/latest/user_guides/mlops/serving/scheduling/ # How To Configure Scheduling For A Model Deployment ## Introduction Scheduling configuration determines how and where your model deployment pods are placed in the Kubernetes cluster. Hopsworks supports Kubernetes scheduler abstractions such as [node affinity](https://kubernetes.io/docs/concepts/scheduling-eviction/assign-pod-node/#node-affinity), [anti-affinity](https://kubernetes.io/docs/concepts/scheduling-eviction/assign-pod-node/#inter-pod-affinity-and-anti-affinity), and [priority classes](https://kubernetes.io/docs/concepts/scheduling-eviction/pod-priority-preemption/), as well as advanced scheduling with [Kueue queues](https://kueue.sigs.k8s.io/docs/concepts/local_queue/) and [topologies](https://kueue.sigs.k8s.io/docs/concepts/topology_aware_scheduling/). !!! tip "Scheduling available for all workloads" In addition to model deployments, all scheduling options are also available for jobs, Jupyter notebooks, and Python deployments. ## Web UI ### Step 1: Create new deployment If you have at least one model already trained and saved in the Model Registry, navigate to the deployments page by clicking on the `Deployments` tab on the navigation menu on the left.

Deployments navigation tab
Deployments navigation tab

Once in the deployments page, you can create a new deployment by either clicking on `New deployment` (if there are no existing deployments) or on `Create new deployment` it the top-right corner. Both options will open the deployment creation form. ### Step 2: Go to advanced options A simplified creation form will appear including the most common deployment fields from all available configurations. Scheduling is part of the advanced options of a deployment. To navigate to the advanced creation form, click on `Advanced options`.

Advance options
Advanced options. Go to advanced deployment creation form

### Step 3: Configure scheduling In the advanced creation form, go to the **Scheduler** section to set up scheduling options for your deployment. Here, you can specify [affinity, anti-affinity, and priority classes](#affinity-anti-affinity-and-priority-classes) to control how your deployment pods are scheduled within the cluster.

Affinity and Priority Classes
Configure affinity and priority classes for the model deployment

If Kueue is ==enabled==, you can also select a [queue and topology](#queues-and-topologies) for your deployment.

Select a queue for the deployment
Select a queue for the model deployment

Select a topology unit for the deployment
Select a topology unit for the model deployment

Once you are done with the changes, click on `Create new deployment` at the bottom of the page to create the deployment for your model. ## Affinity, Anti-Affinity, and Priority Classes You can configure [node affinity](https://kubernetes.io/docs/concepts/scheduling-eviction/assign-pod-node/#node-affinity), [anti-affinity](https://kubernetes.io/docs/concepts/scheduling-eviction/assign-pod-node/#inter-pod-affinity-and-anti-affinity), and [priority classes](https://kubernetes.io/docs/concepts/scheduling-eviction/pod-priority-preemption/) to control pod placement and scheduling priority for your deployment. - **Affinity**: Constrains which nodes the deployment pods can run on based on node labels (e.g., GPU nodes, specific zones). - **Anti-Affinity**: Prevents pods from running on nodes with specific labels. - **Priority Class**: Determines the scheduling and eviction priority of pods. Higher priority pods are scheduled first and can preempt lower priority pods. ## Queues and Topologies !!! warning "Kueue is required" This feature requires Kueue to be enabled in your cluster. If Kueue is not available, queue and topology options will not be accessible. If the cluster has Kueue enabled, you can select a queue for your deployment. [Queues](https://kueue.sigs.k8s.io/docs/concepts/local_queue/) control resource allocation and scheduling priority across the cluster. Administrators define quotas on how many resources a queue can use, and queues can be grouped in cohorts to borrow resources from each other. You can also select a [topology](https://kueue.sigs.k8s.io/docs/concepts/topology_aware_scheduling/) unit to control how deployment pods are co-located. For example, you can require all pods to run on the same host to minimize network latency. ## Learn more For detailed documentation on scheduling abstractions and cluster-level configuration, see the following guides: - [Scheduler](../../projects/scheduling/kube_scheduler.md): Affinity, anti-affinity, priority classes, and project-level defaults - [Kueue Details](../../projects/scheduling/kueue_details.md): Queues, cohorts, topologies, and resource flavors ================================================================================ # API Protocol Source: https://docs.hopsworks.ai/latest/user_guides/mlops/serving/api-protocol/ # How to Select the API protocol for a Deployment { #api-protocol-guide } ## Introduction Hopsworks supports both REST and gRPC as API protocols for sending inference requests to model deployments. While REST API protocol is supported in all types of model deployments, gRPC is currently supported for **Python model deployments** only. The protocol is chosen per deployment with `api_protocol`, in the creation form or in the Python API, and defaults to REST. REST is what `curl`, the published OpenAPI document and any client that is not the Python library use. gRPC costs less per request under concurrency. On a four-client benchmark it served 20 to 30 percent more requests per second and cut p99 latency by around 3 ms, and the gain grows with the batch size. It is worth choosing when the Python library is the only client. A deployment served by the [default predictor][deployment-schema] supports both protocols, because the library encodes the request and decodes the response at both ends. On gRPC the rows travel as one KServe v2 tensor per schema field, and `deployment.predict()` returns the same dictionary it returns over REST. A deployment that runs your own predictor script has to stay on REST unless the script is written for gRPC. Under gRPC the model server hands `predict()` KServe v2 tensors rather than rows, which a script written for REST cannot read. ## Web UI ### Step 1: Create a new deployment If you have at least one model already trained and saved in the Model Registry, navigate to the deployments page by clicking on the `Deployments` tab on the navigation menu on the left.

Deployments navigation tab
Deployments navigation tab

Once in the deployments page, you can create a new deployment by either clicking on `New deployment` (if there are no existing deployments) or on `Create new deployment` it the top-right corner. Both options will open the deployment creation form. ### Step 2: Go to advanced options A simplified creation form will appear including the most common deployment fields from all available configurations. Resource allocation is part of the advanced options of a deployment. To navigate to the advanced creation form, click on `Advanced options`.

Advance options
Advanced options. Go to advanced deployment creation form

### Step 3: Select the API protocol You can select the API protocol to be enabled in your model deployment in the advanced deployment form.

Select gRPC API protocol
Select gRPC API protocol

!!! info "Only one API protocol can be enabled in a model deployment (they cannot support both gRPC and REST)" Currently, KServe model deployments are limited to one API protocol at a time. Therefore, only one of REST or gRPC API protocols can be enabled at the same time on the same model deployment. You cannot change the API protocol of existing deployments. A gRPC deployment answers no HTTP requests, so `curl` cannot test it and the deployment page shows no curl example and no OpenAPI reference for it. Once you are done with the changes, click on `Create new deployment` at the bottom of the page to create the deployment for your model. ## Code ### Step 1: Connect to Hopsworks === "Python" ```python import hopsworks project = hopsworks.login() # get Hopsworks Model Registry handle mr = project.get_model_registry() # get Hopsworks Model Serving handle ms = project.get_model_serving() ``` ### Step 2: Create a deployment with a specific API protocol === "Python" ```python my_model = mr.get_model("my_model", version=1) my_predictor = ms.create_predictor( my_model, api_protocol="GRPC", # defaults to "REST" ) my_predictor.deploy() # or my_deployment = ms.create_deployment(my_predictor) my_deployment.save() ``` !!! api "API reference" - [`ModelServing.create_predictor`][hsml.model_serving.ModelServing.create_predictor] - [`ModelServing.create_deployment`][hsml.model_serving.ModelServing.create_deployment] - [`Deployment`][hsml.deployment.Deployment] - [`api_protocol`][hsml.deployment.Deployment.api_protocol] Browse the full Python API :material-arrow-right: ================================================================================ # REST API Source: https://docs.hopsworks.ai/latest/user_guides/mlops/serving/rest-api/ # Hopsworks Model Serving REST API ## Introduction Hopsworks provides ==model serving capabilities== by leveraging [KServe](https://kserve.github.io/website/) as the model serving platform and [Istio](https://istio.io/) as the ingress gateway to the model deployments. This document explains how to interact with a model deployment via REST API. !!! tip "Tutorials" End-to-end examples are available in the [hopsworks-tutorials](https://github.com/logicalclocks/hopsworks-tutorials/tree/master) repository. ## Sending Inference Requests through Istio Ingress The full inference URL is constructed by combining a base path with a model server-specific suffix. See [URL Paths](#url-paths) for the complete URL format and examples. ### Authentication All requests must include an API Key for authentication. You can create an API key by following this [guide](../../projects/api_key/create_api_key.md). Include the key in the `authorization` header: ```text authorization: ApiKey ``` ### Headers | Header | Description | Example Value | | --------------- | ----------------------------------- | ----------------------- | | `authorization` | API key for authentication. | `ApiKey ` | | `content-type` | Request payload type (always JSON). | `application/json` | ## URL Paths Deployed models are accessible through the ==Istio ingress gateway== using **path-based** routing. The full URL is constructed by combining the base path with a model server-specific suffix. This URL is also provided on the model deployment page in the Hopsworks UI. !!! example "" **`/`** Where `server-specific_suffix` depends on the model server type (see [ML Inference Paths](#ml-inference) or [OpenAI-compatible Paths](#openai-compatible)). ### Base URL The base URL is composed of the **Istio ingress gateway IP**, the **project name**, and the **deployment name**. !!! example "" **`https:///v1//`** !!! warning "Host-based routing (legacy)" Prior to path-based routing, requests were routed using a `Host` header matching the model deployment hostname, and **`https://`** as base url. ``` Host: .. ``` Each model deployment gets its own Knative-generated hostname, and routing depends on the `Host` header matching Istio ingress gateway rules. Path-based routing (described above) is the preferred method for external access. ### ML Inference For model deployments using Python, KServe sklearnserver, or TensorFlow Serving, the URL follows the KServe V1 inference protocol. !!! info "Supported verbs and path format" | Model Server | Supported Verbs | Path Format | | -------------------- | -------------------------------- | ------------------------------------ | | Python | `predict` | `/v1/models/:` | | KServe sklearnserver | `predict` | `/v1/models/:` | | TensorFlow Serving | `predict`, `classify`, `regress` | `/v1/models/:` | !!! tip "Hopsworks Python API" ML inference urls can be retrieved using the `Deployment` class. ```python # Returns: https:///v1///v1/models/:predict inference_url = deployment.get_inference_url() ``` ### Deployment schema discovery Deployments that carry a deployment schema describe their request and response contract as JSON Schema and OpenAPI, and validate every request against it before any predictor code runs. The documents are served by the Hopsworks REST API (not the Istio ingress), with an API key that has the `SERVING` scope: !!! example "" **`GET https:///hopsworks-api/api/project//serving//schema?format=openapi`** `format` is `schema` (default), `jsonschema`, or `openapi`; `schemaId=` returns the contract of an earlier revision. The serving id is the `id` field of `GET .../project//serving?name=`. A deployment without a schema, or an unknown id, answers `404` with error code `240037`. Feature view deployments answer on the same `/v1/models/:predict` route as Python model deployments. See the [Deployment Schema Guide][deployment-schema]. ### OpenAI-compatible ==vLLM deployments== provide an OpenAI API-compatible endpoint at `/v1/`, allowing you to send any standard OpenAI API request to the vLLM server. !!! example "e.g., Chat Completions endpoint" **`/v1/chat/completions`** Refer to the official [vLLM OpenAI-compatible server documentation](https://docs.vllm.ai/en/v0.28.0/serving/openai_compatible_server.html) for details about the available APIs. !!! tip "Hopsworks Python API" OpenAI-compatible urls can be retrieved using the `Deployment` class. ```python # Returns: https:///v1///v1 # Append /chat/completions or /completions for specific endpoints openai_url = deployment.get_openai_url() ``` ## Request Format The request format depends on the model server being used. For predictive inference (TensorFlow, sklearn, or Python model server), the request must be sent as a JSON object containing an `inputs` or `instances` field. See [more information on the request format](https://kserve.github.io/website/docs/concepts/architecture/data-plane/v1-protocol#request-format). !!! example "REST API example for Predictive Inference (Tensorflow or SkLearn or Python Serving)" === "Python" ```python import requests data = {"inputs": [[4641025220953719, 4920355418495856]]} headers = {"authorization": "ApiKey ", "content-type": "application/json"} response = requests.post( "https:///v1/my_project/fraud/v1/models/fraud:predict", headers=headers, json=data, ) print(response.json()) ``` === "Curl" ```bash curl -X POST "https:///v1/my_project/fraud/v1/models/fraud:predict" \ -H "authorization: ApiKey " \ -H "content-type: application/json" \ -d '{ "inputs": [ [4641025220953719, 4920355418495856] ] }' ``` For generative inference (vLLM), the request follows the [OpenAI specification](https://docs.vllm.ai/en/v0.28.0/serving/openai_compatible_server.html) supported by the vLLM OpenAI-compatible server. !!! example "vLLM chat completions" === "Python" ```python import requests data = { "model": "my-llm", "messages": [{"role": "user", "content": "Hello, how are you?"}], } headers = {"authorization": "ApiKey ", "content-type": "application/json"} response = requests.post( "https:///v1/my_project/my-llm/v1/chat/completions", headers=headers, json=data, ) print(response.json()) ``` === "Curl" ```bash curl -X POST "https:///v1/my_project/my-llm/v1/chat/completions" \ -H "authorization: ApiKey " \ -H "content-type: application/json" \ -d '{ "model": "my-llm", "messages": [ {"role": "user", "content": "Hello, how are you?"} ] }' ``` ## CORS The Istio EnvoyFilter handles CORS preflight (`OPTIONS`) requests automatically. Allowed origins can be configured via `istio.envoyFilter.corsAllowedOrigins` in the Helm chart configuration. ## Response The model returns predictions in a JSON object. The response depends on the model server implementation. ================================================================================ # Troubleshooting Source: https://docs.hopsworks.ai/latest/user_guides/mlops/serving/troubleshooting/ # How To Troubleshoot A Model Deployment ## Introduction In this guide, you will learn how to troubleshoot a deployment that is having issues to serve a trained model. But before that, it is important to understand how [deployment states](deployment-state.md) are defined and the possible transitions between conditions. Before a deployment starts, it goes through a `CREATING` phase where deployment artifacts are prepared. When a deployment is starting, it follows an ordered sequence of [states](deployment-state.md#deployment-conditions) before becoming ready for serving predictions. Similarly, it follows an ordered sequence of states when being stopped, although with fewer steps. !!! warning "`FAILED` is a terminal state" If a deployment reaches the `FAILED` state, it cannot recover on its own. You must stop and restart the deployment to attempt recovery. ## Web UI ### Step 1: Inspect deployment status If you have at least one deployment already created, navigate to the deployments page by clicking on the `Deployments` tab on the navigation menu on the left.

Deployments navigation tab
Deployments navigation tab

Once in the deployments page, find the deployment you want to inspect. Next to the actions buttons, you can find an indicator showing the current status of the deployment. For a more descriptive representation, this indicator changes its color based on the status. To inspect the condition of the deployment, click on the name of the deployment to open the deployment overview page. ### Step 2: Inspect condition At the top of page, you can find the same status indicator mentioned in the previous step. Below it, a one-line message is shown with a more detailed description of the deployment status. This message is built using the current status [condition](deployment-state.md#deployment-conditions) of the deployment. Oftentimes, the status and the one-line description are enough to understand the current state of a deployment. For instance, when the cluster lacks enough allocatable resources to meet the deployment requirements, a meaningful error message will be shown with the root cause.

Deployment failed to schedule condition
Condition of a deployment that cannot be scheduled

However, when the deployment fails to start further details might be needed depending on the source of failure. For example, failures in the initialization or starting steps will show a less relevant message. In those cases, you can explore the deployments logs in search of the cause of the problem.

Deployment failed to start condition
Condition of a deployment that fails to start

### Step 3: Explore transient logs Each deployment is composed of several components depending on its configuration and the model being served. Transient logs refer to component-specific logs that are read directly from the running component. Therefore, these logs can only be retrieved as long as the deployment components are running. !!! info "" Transient logs are informative and fast to retrieve, facilitating the troubleshooting of deployment components at a glance. Transient logs are convenient when access to the most recent logs of a deployment is needed. To follow them in the UI, click the `Logs` button at the top of the deployment overview page. The pane tails the selected component every two seconds and lets you search, copy and download what it has buffered. You can also read them with the Hopsworks Machine Learning Python library, as shown in [Step 4](#step-4-explore-transient-logs) of the code section. !!! info When a deployment is in idle state, there are no components running (i.e., scaled to zero) and, thus, no transient logs are available. Use historical logs to inspect an instance that is already gone. !!! note Standard output and standard error arrive as a single interleaved stream. Kubernetes merges them at the container runtime, so the two cannot be separated after the fact. ### Step 4: Explore historical logs Historical logs are archives that each instance writes to the project's `Logs` dataset from inside its own container. An instance archives its output when it exits, is restarted, or is stopped, which means an instance removed by scale-to-zero or replaced by a new deployment revision still leaves its logs behind. !!! info "" Historical logs are convenient when a deployment fails occasionally, or when the instance you need to inspect is no longer running. Archives are written to `Logs/Serving//` and named `__.log`, one file per instance run. Browse them under the `Logs` section of the deployment overview page, or in the `Logs` dataset, and open one to read it. Historical logs are only written for components that have disk logging enabled. See [configuring disk logging](#configuring-disk-logging) below, and note that only Python deployments support it: Python predictors have it on by default. Every instance of the component writes its own archive, distinguished by pod name. !!! warning The number of archives kept per deployment is capped by the `log_history_limit` cluster variable, which defaults to 30. Once the cap is reached, the oldest archive is deleted each time a new one is written, so long-lived deployments do not fill the project with logs. To retrieve archives with the Python library, use `deployment.download_logs()`, shown in [Step 5](#step-5-download-historical-logs) below. ### Configuring disk logging Disk logging controls whether a component archives its output to the project's `Logs` dataset. It is a per-component checkbox under `Disk logging` in the advanced options of the deployment form. When it is on, each instance keeps its output on local disk while it runs and uploads it when it stops, to a separate file distinguished by pod name. This covers stops the platform initiates on its own, such as scale-to-zero and revision replacement, not only stops a user asks for. When it is off, nothing is written. Disk logging is only available for deployments whose serving container runs a Hopsworks inference pipeline image, because the upload runs the Hopsworks Python library from inside that container. That means Python deployments, agent deployments included, and their transformers; Python predictors have it on by default. TensorFlow Serving and vLLM do not support it, and neither does a KServe Python deployment with no predictor script, which runs the sklearnserver runtime image. The API rejects the setting for those rather than deploying something that cannot archive. !!! note There is no single-instance mode. All instances of a deployment share one pod template, so they either all archive or none do. !!! note Changing disk logging starts a new deployment revision, because it changes the pod template. ## Code ### Step 1: Connect to Hopsworks === "Python" ```python import hopsworks project = hopsworks.login() # get Hopsworks Model Serving handle ms = project.get_model_serving() ``` ### Step 2: Retrieve an existing deployment === "Python" ```python deployment = ms.get_deployment("mydeployment") ``` ### Step 3: Get current deployment's predictor state === "Python" ```python state = deployment.get_state() state.describe() ``` ### Step 4: Explore transient logs === "Python" ```python deployment.get_logs(component="predictor|transformer", tail=10) ``` To follow a running deployment instead of taking a single snapshot, use `tail_logs`. It returns a generator that yields new lines as they arrive, skipping what it has already yielded. === "Python" ```python for chunk in deployment.tail_logs(component="predictor"): print(chunk, end="") ``` ### Step 5: Download historical logs === "Python" ```python local_paths = deployment.download_logs(latest=True) for local_path in local_paths: with open(local_path) as archive: print(archive.read()) ``` Omit `latest` to download every archive the deployment has kept. !!! api "API reference" - [`ModelServing.get_deployment`][hsml.model_serving.ModelServing.get_deployment] - [`Deployment`][hsml.deployment.Deployment] - [`get_state`][hsml.deployment.Deployment.get_state] - [`get_logs`][hsml.deployment.Deployment.get_logs] - [`tail_logs`][hsml.deployment.Deployment.tail_logs] - [`download_logs`][hsml.deployment.Deployment.download_logs] - [`PredictorState`][hsml.predictor_state.PredictorState] - [`describe`][hsml.predictor_state.PredictorState.describe] Browse the full Python API :material-arrow-right: ================================================================================ # External Access Source: https://docs.hopsworks.ai/latest/user_guides/mlops/serving/external-access/ # How To Configure External Access To A Model Deployment ## Introduction Hopsworks supports **role-based access control (RBAC)** for project members within a project, where a project ML assets can only be accessed by Hopsworks users that are members of that project (See [governance](../../../concepts/projects/governance.md)). However, there are cases where you might want to grant ==external users== with access to specific model deployments without them having to register into Hopsworks or to join the project which will give them access to all project ML assets. For these cases, Hopsworks supports fine-grained access control to model deployments based on ==user groups== managed by an external Identity Provider. !!! info "Authentication methods" Hopsworks can be configured to use different types of authentication methods including OAuth2, LDAP and Kerberos. See the [Authentication Methods Guide](../../../setup_installation/admin/auth.md) for more information. ## Web UI (for Hopsworks users) ### Step 1: Navigate to a model deployment If you have at least one model deployment already created, navigate to the model deployments page by clicking on the `Deployments` tab on the navigation menu on the left.

Deployments navigation tab
Deployments navigation tab

Once in the model deployments page, find the model deployment you want to configure external access and click on the name of the deployment to open the model deployment overview page.

Deployment overview
Deployment overview

### Step 2: Go to External Access You can find the external access configuration by clicking on `External access` on the navigation menu on the left or scrolling down to the external access section.

Deployment external access
External access configuration

### Step 3: Add or remove user groups In this section, you can add and remove user groups by clicking on `edit external user groups` and typing the group name in the **text-free** input field or **selecting** one of the existing ones in the dropdown list. After that, click on the `save` button to persist the changes. !!! Warn "Case sensitivity" Inference requests are authorized using a ==case-sensitive exact match== between the group names of the user making the request and the group names granted access to the model deployment. Therefore, a user assigned to the group `lab1` won't have access to a model deployment accessible by group `LAB1`.

Deployment external access
External access configuration

## Web UI (for external users) ### Step 1: Login with the external identity provider Navigate to Hopsworks, and click on the `Login with` button to sign in using the configured external identity provider (e.g., Keycloak in this example).

Login external identity provider
Login with External Identity Provider

### Step 2: Explore the model deployments you are granted access to Once you sign in to Hopsworks, you can see the list of model deployments you are granted access to based on your assigned groups.

Deployments list
Deployments with external access

### Step 2: Inspect your current groups You can find the current groups you are assigned to at the top of the page.

External user groups
External user groups

### Step 3: Get an API key Inference requests to model deployments are authenticated and authorized based on your external user and user groups. You can create API keys to authenticate your inference requests by clicking on the `Create API Key` button. !!! info "Authorization header" API keys are set in the `authorization` header following the format `ApiKey `

Get API key
Get API key

### Step 4: Send inference requests The URI path for sending inference requests depends on the type of model deployment. For example, LLM deployments typically use `/chat/completions`, while traditional model deployments use `/predict`. You can find the exact URI path for each deployment on its model deployment card. For detailed instructions on constructing requests and handling authentication, refer to the [REST API Guide](rest-api.md). !!! tip "Code snippets" For clients sending inference requests using libraries similar to curl or OpenAI API-compatible libraries (e.g., LangChain), you can find code snippet examples by clicking on the `Curl >_` and `LangChain >_` buttons.

Deployment endpoint
Deployment endpoint

## Refreshing External User Groups Every time an external user signs in to Hopsworks using a pre-configured [authentication method](../../../setup_installation/admin/auth.md), Hopsworks fetches the external user groups and updates the internal state accordingly. Given that groups can be added/removed from users at any time by the Identity Provider, Hopsworks needs to periodically fetch the external user groups to keep the state updated. Therefore, external users that want to access model deployments are **required to login periodically** to ensure they are still part of the allowed groups. The timespan between logins is controlled by the configuration parameter `requireExternalUserLoginAfterHours` available during the Hopsworks installation and upgrade. The `requireExternalUserLoginAfterHours` configuration parameter controls the ==number of hours== after which external users are required to sign in to Hopsworks to refresh their external user groups. !!! info "Configuring `requireExternalUserLoginAfterHours`" Allowed values are -1, 0 and greater than 0, where -1 disables the periodic login requirement and 0 disables external access completely for every model deployment. ================================================================================ # Model Monitoring Source: https://docs.hopsworks.ai/latest/user_guides/mlops/model_monitoring/ # Model Monitoring Model monitoring lets you detect drift between the data a model was trained on and the data it serves in production. It is built on top of [feature monitoring](../../fs/feature_monitoring/index.md): a model monitoring configuration computes statistics over the model's logged inference data (the _detection window_) and compares them against the training dataset the model was trained on (the _reference window_). You can configure model monitoring from a [model deployment](model_monitoring.md), a [model](model_monitoring.md), or a [feature view](model_monitoring.md). All three resolve to the same underlying configuration on the feature view's logging feature group, filtered by the model name and version. ## Prerequisites To enable model monitoring, you need: - A **feature view** with **feature logging enabled**. Logging is what captures the inference data that monitoring analyzes. If logging is not enabled, enable it with `feature_view.enable_logging()`. - A **registered model** trained from that feature view, with a **recorded training dataset version**. The training dataset version is recorded automatically when the model is created from a feature view, and is used as the default reference distribution. - A **model deployment** that logs its inference data to the feature view via the feature view `log` methods (typically from its predictor script). Feature logging is what produces the data that monitoring analyzes, so the deployment must write its inference data to the feature view's logging feature group. This is currently only supported for model deployments with a predictor script. See the [Feature Logging guide](../../fs/feature_view/feature_logging.md). ## How it relates to feature monitoring Model monitoring is feature monitoring applied to a model's inference logs: - The **detection window** covers the data recently served by the model, read from the logging feature group and filtered by model name and version. - The **reference window** defaults to the training dataset version used to train the model, so you compare production data against training data out of the box. - The **comparison criteria** uses the same [scalar metric](../../fs/feature_monitoring/statistics_comparison.md) and [data distribution](../../fs/feature_monitoring/distribution_comparison.md) criteria as feature monitoring. !!! info "Next steps" See the [Model Monitoring Creation guide](model_monitoring.md) for code examples from a model deployment, a model, and a feature view. ================================================================================ # Model Monitoring Creation Source: https://docs.hopsworks.ai/latest/user_guides/mlops/model_monitoring/model_monitoring/ # Model Monitoring Creation This guide shows how to configure monitoring for a model in production using the ==Hopsworks Python library==. Make sure you have read the [Model Monitoring overview](index.md) first. !!! info "Prerequisites" Make sure you meet the [prerequisites](index.md#prerequisites): a feature view with feature logging enabled, a registered model with a recorded training dataset version, and a model deployment with a predictor script that logs its inference data. !!! warning "Limited UI support" Like feature monitoring, model monitoring can currently only be configured using the [Hopsworks Python library](https://pypi.org/project/hopsworks). ## Code In this section, we show you how to set up model monitoring from a model or a model deployment using the ==Hopsworks Python library==. ### Step 1: Connect to Hopsworks Connect the client running your notebook to Hopsworks and get the Model Registry and Model Serving handles. === "Python" ```python import hopsworks project = hopsworks.login() mr = project.get_model_registry() ms = project.get_model_serving() ``` See the API reference for [`hopsworks.login`][hopsworks.login], [`Project.get_model_registry`][hopsworks_common.project.Project.get_model_registry] and [`Project.get_model_serving`][hopsworks_common.project.Project.get_model_serving]. ### Step 2: Get a model or deployment Retrieve the model or the deployment whose inference data you want to monitor. === "From a deployment" ```python my_deployment = ms.get_deployment("my_deployment") ``` === "From a model" ```python my_model = mr.get_model("my_model", version=1) ``` See the API reference for [`ModelServing.get_deployment`][hsml.model_serving.ModelServing.get_deployment] and [`ModelRegistry.get_model`][hsml.model_registry.ModelRegistry.get_model]. ### Step 3: Create a model monitoring configuration Start a new configuration from the deployment or the model. Hopsworks resolves the model's parent feature view from its provenance and fills in the model name and version for you. === "From a deployment" ```python fm_monitoring_config = my_deployment.create_model_monitoring( name="model_psi_monitoring", ) ``` === "From a model" ```python fm_monitoring_config = my_model.create_model_monitoring( name="model_psi_monitoring", ) ``` See the API reference for [`Deployment.create_model_monitoring`][hsml.deployment.Deployment.create_model_monitoring] and [`Model.create_model_monitoring`][hsml.model.Model.create_model_monitoring]. !!! tip "Configuring from a feature view" You can also configure model monitoring directly from the feature view backing the model, using `feature_view.create_model_monitoring`. See the [Feature Monitoring guide for Feature Views](../../fs/feature_view/feature_monitoring.md#monitor-a-model-in-production). !!! info "Custom schedule" By default, monitoring runs every day at 12PM. You can modify the schedule by adjusting the `cron_expression`, `start_date_time` and `end_date_time` parameters of `create_model_monitoring` (UTC, Quartz specification). !!! warning "Sub-hourly schedules" The inference-log feature group materializes its offline data at most once per hour. A cron expression that fires more than once per hour produces redundant or incomplete detection windows and triggers a warning, so use a schedule that fires at most once per hour. ### Step 4: Define a detection window The detection window covers the inference data recently served by the model. Define it using the `window_length` and `time_offset` parameters of the `with_detection_window` method. === "Python" ```python fm_monitoring_config.with_detection_window( time_offset="1d", # data served by this model in the last day window_length="1d", ) ``` See the API reference for [`FeatureMonitoringConfig.with_detection_window`][hsfs.core.feature_monitoring_config.FeatureMonitoringConfig.with_detection_window]. ### Step 5: Define the reference training dataset By default, the reference is the training dataset version used to train the model, recorded at model registration time. You can call `with_reference_training_dataset()` without arguments to make this explicit. === "Python" ```python fm_monitoring_config.with_reference_training_dataset( # omitted -> defaults to the model's training dataset version ) ``` See the API reference for [`FeatureMonitoringConfig.with_reference_training_dataset`][hsfs.core.feature_monitoring_config.FeatureMonitoringConfig.with_reference_training_dataset]. !!! info "Training dataset version validation" If you pass a specific version, it must match the model's recorded training dataset version, otherwise the call raises an exception. This guards against accidentally comparing production data against a training dataset the model was never trained on. ### Step 6.A: Compare on a scalar metric Select the feature and the metric to compare, and define a relative or absolute threshold. === "Python" ```python fm_monitoring_config.compare_on( feature_name="amount", # the feature to compare metric="mean", threshold=0.2, # a relative change over 20% is considered anomalous relative=True, # relative or absolute change strict=False, # strict or relaxed comparison ) ``` See the API reference for [`FeatureMonitoringConfig.compare_on`][hsfs.core.feature_monitoring_config.FeatureMonitoringConfig.compare_on]. ### Step 6.B: Compare on the whole distribution Alternatively, instead of a single scalar metric, you can detect drift in the shape of the feature's distribution using `compare_on_distribution`. Select a distribution distance metric (e.g., `PSI`) and a threshold. === "Python" ```python fm_monitoring_config.compare_on_distribution( feature_name="amount", # the feature to compare metric="PSI", threshold=0.2, # a distance above 0.2 is considered a significant shift ) ``` See the API reference for [`FeatureMonitoringConfig.compare_on_distribution`][hsfs.core.feature_monitoring_config.FeatureMonitoringConfig.compare_on_distribution]. !!! tip "More distribution options" See the [Distribution comparison guide](../../fs/feature_monitoring/distribution_comparison.md) for the full list of metrics and binning strategies. ### Step 7: Save the configuration Finally, save the configuration by calling the `save` method. Once saved, the schedule for the statistics computation and comparison is activated automatically. === "Python" ```python fm_monitoring_config.save() ``` See the API reference for [`FeatureMonitoringConfig.save`][hsfs.core.feature_monitoring_config.FeatureMonitoringConfig.save]. ### Step 8: Retrieve configurations You can list the monitoring configurations attached to a model or deployment. === "From a deployment" ```python fm_configs = my_deployment.get_monitoring_configs() ``` === "From a model" ```python fm_configs = my_model.get_monitoring_configs() ``` See the API reference for [`Deployment.get_monitoring_configs`][hsml.deployment.Deployment.get_monitoring_configs] and [`Model.get_monitoring_configs`][hsml.model.Model.get_monitoring_configs]. !!! info "Next steps" Model monitoring results integrate with the same [alerting](../../fs/feature_monitoring/index.md#alerting) and [interactive graph](../../fs/feature_monitoring/interactive_graph.md) tooling as feature monitoring. See the [`FeatureMonitoringConfig`][hsfs.core.feature_monitoring_config.FeatureMonitoringConfig] reference to learn how to disable, manually trigger, or delete a configuration. ================================================================================ # Vector Similarity Search Source: https://docs.hopsworks.ai/latest/user_guides/fs/vector_similarity_search/ ## Introduction Vector similarity search (also called similarity search) is a technique enabling the retrieval of similar items based on their vector embeddings or representations. Its applications range across various domains, from recommendation systems to image similarity and beyond. In Hopsworks, vector similarity search is enabled by extending an online feature group with approximate nearest neighbor search capabilities through a vector database, such as Opensearch. This guide provides a detailed walkthrough on how to leverage Hopsworks for vector similarity search. ## Extending Feature Groups with Similarity Search In Hopsworks, each vector embedding in a feature group is stored in an index within the backing vector database. By default, vector embeddings are stored in the default index for the project (created for every project in Hopsworks), but you have the option to create a new index for a feature group if needed. Creating a separate index per feature group is particularly useful for large volumes of data, ensuring that when a feature group is deleted, its associated index is also removed. For feature groups that use the default project index, the index will only be removed when the project is deleted - not when the feature group is deleted. The index will store all the vector embeddings defined in that feature group, if you have more than one vector embedding in the feature group. In the following example, we explicitly define an index for the feature group: ```aidl from hsfs import embedding # Specify optionally the index in the vector database emb = embedding.EmbeddingIndex(index_name="news_fg") ``` Then, add one or more embedding features to the index. Name and dimension of the embedding features are required for identifying which features should be indexed for k-nearest neighbor (KNN) search. In this example, we get the dimension of the embedding by taking the length of the value of the `embedding_heading` column in the first row of the dataframe `df`. Optionally, you can specify the similarity function among `l2_norm`, `cosine`, and `dot_product`. Refer to [`EmbeddingIndex.add_embedding`][hsfs.embedding.EmbeddingIndex.add_embedding] for the full list of arguments. ```aidl # Add embedding feature to the index emb.add_embedding("embedding_heading", len(df["embedding_heading"][0])) ``` Next, you create a feature group with the `embedding_index` and ingest data to the feature group. When the `embedding_index` is provided, the vector database is used as online feature store. That is, all the features in the feature group are stored **exclusively** in the vector database. The advantage of storing all features in the vector database is that it enables similarity search, and push-down filtering for all feature values. ```aidl # Create a feature group with the embedding index news_fg = fs.get_or_create_feature_group( name=f"news_fg", embedding_index=emb, # Provide the embedding index created primary_key=["news_id"], version=version, online_enabled=True ) # Write a DataFrame to the feature group, including the offline store and the ANN index (in the Vector Database) news_fg.insert(df) ``` ## Similarity Search for Feature Groups using Vector Embeddings You provide a vector embedding as a parameter to the search query using [`FeatureGroup.find_neighbors`][hsfs.feature_group.FeatureGroup.find_neighbors], and it returns the rows in the online feature group that have vector embedding values most similar to the provided vector embedding. It is also possible to filter rows by specifying a filter on any of the features in the feature group. The filter is pushed down to the vector database to improve query performance. In the first code snippet below, `find_neighbor`s returns 3 rows in `news_fg` that have the closest `news_description` values to the provided `news_description`. In the second code snippet below, we only return news articles with a `newstype` of `sports`. ```aidl # Search neighbor embedding with k=3 news_fg.find_neighbors(model.encode(news_description), k=3) # Filter and search news_fg.find_neighbors(model.encode(news_description), k=3, filter=news_fg.newstype == "sports") ``` To analyze feature values at specific points in time, you can utilize time travel functionality: ```aidl # Time travel and read from the offline feature store news_fg.as_of(time_in_past).read() ``` ## Querying Similar Embeddings with Additional features You can also use similarity search for vector embedding features in feature views. In the code snippet below, we create a feature view by selecting features from the earlier `news_fg` and a new feature group `view_fg`. If you include a feature group with vector embedding features in a feature view, **whether or not the vector embedding features are selected**, you can call `find_neighbors` on the feature view, and it will return rows containing all the feature values in the feature view. In the example below, a list of `heading` and `view_cnt` will be returned for the news articles which are closet to provided `news_description`. ```aidl view_fg = fs.get_or_create_feature_group( name="view_fg", primary_key=["news_id"], version=version, online_enabled=True ) fv = fs.get_or_create_feature_view( "news_view", version=version, query=news_fg.select(["heading"]).join(view_fg.select(["view_cnt"])) ) fv.find_neighbors(model.encode(news_description), k=5) ``` Note that you can use similarity search from the feature view **only if** the feature group which you are querying with `find_neighbors` has **all** the primary keys of the other feature groups. In the example above, you are querying against the feature group `news_fg` which has the vector embedding features, and it has the feature "news_id" which is the primary key of the feature group `view_fg`. But if `page_fg` is used as illustrated below, `find_neighbors` will fail to return any features because primary key `page_id` does not exist in `news_fg`. --8<-- "user_guides/fs/vector_similarity_search/find-neighbors.html" It is also possible to get back feature vector by providing the primary keys, but it is not recommended as explained in the next section. The client fetches feature vector from the vector store and the online store for `news_fg` and `view_fg` respectively. ```aidl fv.get_feature_vector({"news_id": 1}) ``` ## Performance considerations for Feature Groups with Embeddings ### Choose Features for Vector Store While it is possible to update feature value in vector store, updating feature value in online store is more efficient. If you have features which are frequently being updated and do not require for filtering, consider storing them separately in a different feature group. As shown in the previous example, `view_cnt` is updated frequently and stored separately. You can then get all the required features by using feature view. ### Choose the Appropriate Online Feature Stores There are 2 types of online feature stores in Hopsworks: online store (RonDB) and vector store (Opensearch). Online store is designed for retrieving feature vectors efficiently with low latency. Vector store is designed for finding similar embedding efficiently. If similarity search is not required, using online store is recommended for low latency retrieval of feature values including embedding. ### Use New Index per Feature Group Create a new index per feature group to optimize retrieval performance. ## Next steps Explore the [news search example](https://github.com/logicalclocks/hopsworks-tutorials/blob/master/api_examples/vector_similarity_search/1_feature_group_embeddings_api.ipynb), demonstrating how to use Hopsworks for implementing a news search application using natural language in the application. Additionally, you can see the application of querying similar embeddings with additional features in this [news rank example](https://github.com/logicalclocks/hopsworks-tutorials/blob/master/api_examples/vector_similarity_search/2_feature_view_embeddings_api.ipynb). ================================================================================ # Connect Source: https://docs.hopsworks.ai/latest/user_guides/projects/opensearch/connect/ # How To Connect To OpenSearch ## Introduction Text here !!! notice "Limited to internal Jobs and Notebooks" Currently it's only possible to configure the opensearch-py client in a job or jupyter notebook running inside the Hopsworks cluster. ## Code In this guide, you will learn how to connect to the OpenSearch cluster using an [opensearch-py](https://opensearch.org/docs/1.3/clients/python/) client. ### Step 1: Get the OpenSearch API ```python import hopsworks project = hopsworks.login() opensearch_api = project.get_opensearch_api() ``` ### Step 2: Configure the opensearch-py client ```python from opensearchpy import OpenSearch client = OpenSearch(**opensearch_api.get_default_py_config()) ``` !!! api "API reference" - [`Project.get_opensearch_api`][hopsworks_common.project.Project.get_opensearch_api] - [`OpenSearchApi`][hopsworks_common.core.opensearch_api.OpenSearchApi] - [`get_default_py_config`][hopsworks_common.core.opensearch_api.OpenSearchApi.get_default_py_config] Browse the full Python API :material-arrow-right: ## Going Further You can now use the client to interact directly with the OpenSearch cluster, such as [vector database](../../../concepts/mlops/opensearch.md). ================================================================================ # KNN Source: https://docs.hopsworks.ai/latest/user_guides/projects/opensearch/knn/ # How To Use OpenSearch k-NN plugin ## Introduction The k-NN plugin enables users to search for the k-nearest neighbors to a query point across an index of vectors. To determine the neighbors, you can specify the space (the distance function) you want to use to measure the distance between points. Use cases include recommendations (for example, an “other songs you might like” feature in a music application), image recognition, and fraud detection. !!! notice "Limited to internal Jobs and Notebooks" Currently it's only possible to configure the opensearch-py client in a job or jupyter notebook running inside the Hopsworks cluster. ## Code In this guide, you will learn how to create a simple recommendation application, using the `k-NN plugin` in OpenSearch. ### Step 1: Get the OpenSearch API === "Python" ```python import hopsworks project = hopsworks.login() opensearch_api = project.get_opensearch_api() ``` ### Step 2: Configure the opensearch-py client === "Python" ```python from opensearchpy import OpenSearch client = OpenSearch(**opensearch_api.get_default_py_config()) ``` ### Step 3: Create an index Create an index to use by calling `opensearch_api.get_project_index(..)`. === "Python" ```python knn_index_name = opensearch_api.get_project_index("demo_knn_index") index_body = { "settings": { "knn": True, "knn.algo_param.ef_search": 100, }, "mappings": { "properties": {"my_vector1": {"type": "knn_vector", "dimension": 2}} }, } response = client.indices.create(knn_index_name, body=index_body) print(response) ``` ### Step 4: Bulk ingestion of vectors Ingest 10 vectors in a bulk fashion to the index. These vectors represent the list of vectors to calculate the similarity for. === "Python" ```python import random from opensearchpy.helpers import bulk actions = [ { "_index": knn_index_name, "_id": count, "_source": { "my_vector1": [random.uniform(0, 10), random.uniform(0, 10)], }, } for count in range(0, 10) ] bulk( client, actions, ) ``` ### Step 5: Score vector similarity Score the vector `[2.5, 3]` and find the 3 most similar vectors. === "Python" ```python # Define the search request query = { "size": 3, "query": {"knn": {"my_vector1": {"vector": [2.5, 3], "k": 3}}}, } # Perform the similarity search response = client.search(body=query, index=knn_index_name) # Pretty print response import pprint pp = pprint.PrettyPrinter() pp.pprint(response) ``` `Output` from the above script shows the score for each of the three most similar vectors that have been indexed. `[4.798869166444522, 4.069064892468535]` is the most similar vector to `[2.5, 3]` with a score of `0.1346312`. === "Bash" ```bash 2022-05-30 09:55:50,529 INFO: POST https://10.0.2.15:9200/my_project_demo_knn_index/_search [status:200 request:0.017s] {'_shards': {'failed': 0, 'skipped': 0, 'successful': 1, 'total': 1}, 'hits': {'hits': [{'_id': '9', '_index': 'my_project_demo_knn_index', '_score': 0.1346312, '_source': {'my_vector1': [4.798869166444522, 4.069064892468535]}, '_type': '_doc'}, {'_id': '0', '_index': 'my_project_demo_knn_index', '_score': 0.040784083, '_source': {'my_vector1': [6.267438489652193, 6.0538134453735175]}, '_type': '_doc'}, {'_id': '7', '_index': 'my_project_demo_knn_index', '_score': 0.03222388, '_source': {'my_vector1': [7.973873201006634, 2.7361877621502115]}, '_type': '_doc'}], 'max_score': 0.1346312, 'total': {'relation': 'eq', 'value': 3}}, 'timed_out': False, 'took': 9} ``` !!! api "API reference" - [`Project.get_opensearch_api`][hopsworks_common.project.Project.get_opensearch_api] - [`OpenSearchApi`][hopsworks_common.core.opensearch_api.OpenSearchApi] - [`get_default_py_config`][hopsworks_common.core.opensearch_api.OpenSearchApi.get_default_py_config] - [`get_project_index`][hopsworks_common.core.opensearch_api.OpenSearchApi.get_project_index] - [k-NN plugin](https://opensearch.org/docs/1.3/search-plugins/knn/knn-index/) Browse the full Python API :material-arrow-right: ================================================================================ # Provenance Source: https://docs.hopsworks.ai/latest/user_guides/mlops/provenance/provenance/ # Provenance ## Introduction Hopsworks allows users to track provenance (lineage) between: - data sources - feature groups - feature views - training datasets - models In the provenance pages we will call a provenance artifact or shortly artifact, any of the five entities above. With the following provenance graph: ```plaintext data source -> feature group -> feature group -> feature view -> training dataset -> model ``` we will call the parent, the artifact to the left, and the child, the artifact to the right. So a feature view has a number of feature groups as parents and can have a number of training datasets as children. Tracking provenance allows users to determine where and if an artifact is being used. You can track, for example, if feature groups are being used to create additional (derived) feature groups or feature views, or if their data is eventually used to train models. You can interact with the provenance graph using the UI or the APIs. ## Model provenance The relationship between feature views and models is captured in the [model][hsml.model.Model] constructor. If you do not provide at least the feature view object to the constructor, the provenance will not capture this relation and you will not be able to navigate from model to the feature view it used or from the feature view to this model. You can provide the feature view object and have the training dataset version be inferred. === "Python" ```python # this fv object will be provided to the model constructor fv = hsfs.get_feature_view(...) # when calling training data related methods on the feature view, the training dataset version is cached in the feature view and is implicitly provided to the model constructor X_train, X_test, y_train, y_test = feature_view.train_test_split(...) # provide the feature_view object in the model constructor hsml.model_registry.ModelRegistry.python.create_model( ... feature_view = fv ...) ``` You can of course explicitly provide the training dataset version. === "Python" ```python # this object will be provided to the model constructor fv = hsfs.get_feature_view(...) # this training dataset version will be provided to the model constructor X_train, X_test, y_train, y_test = feature_view.get_train_test_split(training_dataset_version=1) # provide the feature_view object in the model constructor hsml.model_registry.ModelRegistry.python.create_model( ... feature_view = fv, training_dataset_version = 1, ...) ``` Once the relation is stored in the provenance graph, you can navigate the graph from model to feature view or training dataset and the other way around. Users can call the [`Model.get_feature_view_provenance`][hsml.model.Model.get_feature_view_provenance] method or the [`Model.get_training_dataset_provenance`][hsml.model.Model.get_training_dataset_provenance] method which will each return a [provenance Link object](#provenance-links). You can also retrieve directly the parent feature view object, without the need to extract them from the provenance links object, using the [`Model.get_feature_view`][hsml.model.Model.get_feature_view] method. === "Python" ```python feature_view = model.get_feature_view() ``` This utility method also has the options to initialize the required components for batch or online retrieval of feature vectors. === "Python" ```python model.get_feature_view(init: bool = True, online: Optional[bool]: None) ``` By default, the base init for feature vector retrieval is enabled. In case you have a workflow that requires more particular options, you can disable this base init by setting the `init` to `false`. The method detects if it is running within a deployment and will initialize the feature vector retrieval for the serving. If the `online` argument is provided and `true` it will initialize for online feature vector retrieval. If the `online` argument is provided and `false` it will initialize the feature vector retrieval for batch scoring. ### Using the UI In the model overview UI you can explore the provenance graph of the model:

Model provenance graph
Provenance graph of derived feature groups

## Provenance Links All the `_provenance` methods return a `Link` dictionary object that contains `accessible`, `inaccessible`, `deleted` lists. - `accessible` - contains any artifact from the result, that the user has access to. - `inaccessible` - contains any artifacts that might have been shared at some point in the past, but where this sharing was retracted. Since the relation between artifacts is still maintained in the provenance, the user will only have access to limited metadata and the artifacts will be included in this `inaccessible` list. - `deleted` - contains artifacts that are deleted with children still present in the system. There is minimum amount of metadata for the deleted allowing for some limited human readable identification. ================================================================================ # Agents Source: https://docs.hopsworks.ai/latest/user_guides/agents/ # Agents This section serves to provide guides and examples for the common usage of agent tasks and agent deployments through the Hopsworks UI and APIs. - [Agent Tasks](tasks/index.md): Fire-and-forget agent executions that run as first-class Hopsworks jobs. - [Agent Deployments](deployments/index.md): Served interactive agents and LLM workflows. ================================================================================ # Agent Tasks Source: https://docs.hopsworks.ai/latest/user_guides/agents/tasks/ # How To Run An Agent Task ## Introduction Agent Tasks are **fire-and-forget AI agents** that run as first-class Hopsworks jobs. You write a prompt, select a provider (Claude Code or OpenAI Codex), grant the agent a set of permissions, point it at the project resources it needs, and launch it. The agent runs to completion inside a Kubernetes pod, writes its results to HopsFS, and the pod terminates. No interactive UI, no WebSocket - the same lifecycle as a Python or Spark job. Common use cases: - **Scheduled data-quality audits** - point the agent at a Feature Group and ask it to look for null spikes, schema drift, or freshness issues each night - **Model evaluation** - give it a Model + Deployment ref and have it grade recent predictions against a reference dataset - **Operational runbooks** - turn the wiki page you'd hand a new on-call into a prompt the agent runs whenever a `LONG_RUNNING` alert fires on a pipeline - **Ad-hoc exploration** - fire one off from the CLI with `hops job start` whenever you'd otherwise open Jupyter for a one-shot question Because Agent Tasks use the existing Jobs subsystem, you get the same scheduling, alerts, log capture, and execution history you do for any other Hopsworks job. ## Prerequisites 1. **Feature flag** - agent tasks are gated cluster-wide. An admin must set: ```sql UPDATE variables SET value='true' WHERE id='agent_jobs_enabled'; ``` 2. **Authentication credentials** for whichever provider you choose. The agent can pick these up from three sources (see [Authentication](#authentication) below) - the simplest is to log in to Claude or Codex once from an interactive Hopsworks Terminal session; the credentials persist into HopsFS and every subsequent Agent Task in any of your projects reuses them. 3. **The `agent-job` Python environment** (or a clone of it) - this is the runtime base image. It's installed by default; clone it in Project Settings -> Python Environments if you need to layer additional libraries on top. ## UI ### Step 1: Open Jobs and create a new Agent Task Click **Jobs** in the project sidebar, then **New Job**, and pick `AGENT` as the job type. ### Step 2: Write the prompt The prompt is the task description - what you want the agent to do. Be explicit about the inputs (which Feature Group, which Model) and the expected output (a report file at `${AGENT_OUTPUT_PATH}/result.md`, a Slack message, etc). ```text Inspect feature group `transactions` v3 for null spikes in the `amount` and `merchant_id` columns over the last 7 days. Report findings as Markdown at ${AGENT_OUTPUT_PATH}/result.md. If you find a column with more than 1% nulls, post a Slack alert. ``` ### Step 3: Pick the provider | Provider | When to use | | --- | --- | | **Claude** (default) | Most tasks; supports hooks, max-turns and max-budget controls. | | **Codex** | When you want OpenAI's coding agent. Supports free-form CLI flag overrides via `cliArgs`. | Codex-only fields (`cliArgs`) appear when you select Codex; Claude-only fields (`maxTurns`, `maxBudgetUsd`, `hooks`) are hidden in that case. ### Step 4: Configure permissions Permissions are an explicit allowlist of tools the agent may invoke - they map directly to the underlying CLI's tool patterns. Pick a preset for most jobs: | Preset | Effect | | --- | --- | | `READ_ONLY` (default) | Inspect feature groups, feature views, models - no writes. | | `OPERATOR` | Run any `hops` CLI command + `python *`, plus write under `/hopsfs/*`. | | `FULL` | A curated allowlist of common shell commands (`ls`, `cat`, `grep`, `find`, `curl`, `wget`, `echo`, `mkdir`, `cp`, `mv`), plus `Read(*)` and `Write(/hopsfs/*)`. **Not** unrestricted shell access - pick **Custom** if you need that. | | `Custom` | A list of patterns you provide explicitly (e.g. `Bash(hops fg info *)`, `Read(*)`, `Write(/hopsfs/Resources/*)`). | The agent runs with project-scoped JWT auth, so any data access still goes through normal Hopsworks ACLs - permissions just bound what the *agent* is allowed to attempt. ### Step 5: Attach resource references (optional) Refs inject Hopsworks resource metadata into the agent pod as JSON files under `/context/`. The agent can `cat` these to know which resources to operate on without you having to spell every detail out in the prompt: | Ref type | Required version? | Context file | | --- | --- | --- | | `feature_group` | yes | `/context/fg__v.json` | | `feature_view` | yes | `/context/fv__v.json` | | `model` | yes | `/context/model__v.json` | | `deployment` | no | `/context/deployment_.json` | | `job` | no | `/context/job_.json` | Refs are resolved server-side at launch time by the platform - you don't need to write any glue code. ### Step 6: Environment variables (optional) Use this card to supply credentials and other env vars the agent needs at runtime. Vars set here are merged on top of: 1. Platform-injected vars (project ID, JWT path, HopsFS user home, ...) - these are reserved and rejected if you try to set them 2. AI-provider secrets (from your AI provider secrets table, if any) 3. Your **account-level env vars** (visible across all your runtimes - manage them under Account -> Environment variables) So a per-job `ANTHROPIC_API_KEY` overrides the account-level value, which in turn overrides anything in the secrets table. Reserved-name prefixes that you may **not** set: `HOPS_`, `HOPSWORKS_`, `HOPSFS_`, `AGENT_`. The platform injects several `AGENT_*` variables that you can *read* from your prompt - most usefully `AGENT_OUTPUT_PATH` (where your final artifact must be written). They are reserved for the platform, so you can't override them, but referencing them inside the prompt (e.g. `${AGENT_OUTPUT_PATH}/result.md`) is the intended pattern. The platform also injects `SHARED_DATASETS_DIR` pointing at the project's shared-datasets folder under `/hopsfs/`, so prompts can refer to it as `${SHARED_DATASETS_DIR}/...` without hard-coding the path. ### Step 7: Resources, environment, alerts The bottom of the form is shared with other job types: - **Environment** - defaults to `agent-job`. Pick a clone if you've added Python libraries on top. - **Resource configuration** - CPU cores, memory, GPU count. Default is 0.5 cores / 1024 MB / 0 GPU; bump this for heavier tasks. - **Schedule** - same as other jobs. Cron expressions or interval runs. - **Alerts** - FINISHED / FAILED / KILLED / LONG_RUNNING, routed to a Slack / webhook / email receiver. Setting `passToAgent: true` on a Slack alert injects `SLACK_WEBHOOK_URL` and `SLACK_CHANNEL` into the agent pod's env so the agent can post during execution via the bundled `slack-post` helper. ### Step 8: Run it Click **Save**, then **Run**. The execution lands in the same Executions table as Python and Spark jobs. Click an execution to see its state, logs, and output. ## Authentication The agent pod resolves provider credentials in this priority order: 1. **Reused terminal-login credentials.** If you've opened a Hopsworks Terminal in any of your projects and run `claude login` (or `codex login`), the credentials persist in HopsFS at `/hopsfs/Users//.claude/.credentials.json` and `/hopsfs/Users//.codex/auth.json`. The agent pod symlinks `~/.claude` and `~/.codex` to those paths automatically - no API key needed. 2. **`ANTHROPIC_API_KEY` / `OPENAI_API_KEY` from env.** If terminal credentials are missing or invalid, the agent falls back to these env vars. Set them under Account -> Environment variables (applies to all your jobs) or in this job's Environment variables section (applies just to this job). 3. **AI provider secrets table.** If you've configured AI providers under Project Settings, those keys are also injected. Per-job env vars override them. Most users only need to do step 1 once. ## CLI You can also create and run agent tasks via the `hops` CLI: ```bash # Create from a JSON config cat > my-agent-task.json <// ├── result.md # primary, human-readable result - rendered by the UI ├── metadata.json # { "exit_code": 0, "completed_at": "" } └── (any other files the agent chose to write) ``` The agent's stdout and stderr are captured by the standard execution-log path, visible via the Logs button on the execution row. ## Limitations - **Non-interactive only.** The agent cannot ask follow-up questions - the prompt is the whole input. If you need interactive iteration, use a Terminal session instead. - **`maxTurns` is a hard safety net** (claude only) - set it high enough that reasonable runs don't trip it. Default 50. - **The image must be `agent-job` or a clone** - other Python environments don't ship the Claude/Codex CLIs and will be rejected at job-create time. - **Reserved env-var prefixes** - see [Step 6](#step-6-environment-variables-optional); `HOPS_`, `HOPSWORKS_`, `HOPSFS_`, and `AGENT_` are owned by the platform. ================================================================================ # Agent Deployments Source: https://docs.hopsworks.ai/latest/user_guides/agents/deployments/ # How To Run An Agent Deployment ## Introduction Agent Deployments are **server-only KServe deployments with no model attached**. You provide an entrypoint Python script that starts a REST server. Hopsworks exposes it behind an endpoint. Use Agent Deployments for interactive agents and LLM workflows. The service needs to stay up and answer requests. If you need a fire-and-forget background run, use [Agent Tasks](../tasks/index.md) instead. Common use cases: - **Interactive assistants** - a chat or tool-using agent that stays online - **LLM workflows** - a deterministic sequence of retrieval, reasoning, and generation steps that you want to expose as a service - **RAG-backed services** - an agent that reads Feature Store context and uses it to answer questions ## Where to find it in the UI In the project sidebar, go to **Agents** and then **Agent Deployments**. The list page shows the agent deployments in the project, along with the same type of detail page you use for model deployments. ## Create and manage an agent deployment The most common way to create one is with the SDK or the CLI: ```bash hops agent list hops agent create my_agent.py --name my_agent --requirements requirements.txt --environment my_agent hops agent start my_agent hops agent query my_agent --data '{"prompt": "hello"}' hops agent logs my_agent hops agent info my_agent hops agent stop my_agent hops agent delete my_agent --yes ``` Use `hops agent list` first to confirm auth and serving are reachable. ```python import hopsworks project = hopsworks.login() ms = project.get_model_serving() deployment = ms.deploy_agent( entry="my_agent.py", # .py file or a dir with pyproject.toml name="my_agent", requirements="requirements.txt", environment="my_agent", upload_dir="Resources/agents", # default ) deployment.start(await_running=600) print(deployment.predict(inputs={"prompt": "hello"})) # After editing the code: re-create, then deployment.restart() ``` After creation, the deployment appears in the Agent Deployments list where you can inspect its status, logs, endpoints, and configuration. ### Deploy from a Git repository An agent can be served from a Git repository instead of a project file. The repository is cloned every time the deployment starts, so a restart picks up whatever the branch points at. Supported providers are GitHub, GitLab, and BitBucket. Configure the provider credentials once under project settings, see [Configure a Git Provider](../../projects/git/configure_git_provider.md). ```python deployment = ms.deploy_agent( entry="src/agent.py", # path inside the repository name="my_agent", git_url="https://github.com/my-org/my-agent.git", git_provider="GitHub", git_branch="main", environment="my_agent", ) ``` `entry` is interpreted relative to the repository root rather than as a HopsFS path. If you leave `git_branch` unset, the clone follows the repository's default branch. ### Auto-redeploy on new commits A Git-backed agent can roll itself onto the branch HEAD whenever a new commit is pushed: ```python deployment = ms.deploy_agent( entry="src/agent.py", name="my_agent", git_url="https://github.com/my-org/my-agent.git", git_provider="GitHub", git_branch="main", git_auto_redeploy=True, environment="my_agent", ) ``` Hopsworks polls the remote branch and rolls the deployment when it moves. The running version keeps serving requests until the new one is ready. The flag only applies to Git-backed agents, and Hopsworks rejects it for an agent deployed from a project file. A stopped deployment is not rolled; it clones the branch HEAD on its next start. The deployment's **Artifact files** card shows the repository, the branch, the commit it is running, and whether auto-redeploy is enabled. The entrypoint links to the file in the repository at that commit. ## Small example The file below shows a simple agent program that uses LlamaIndex, FastAPI, and OpenTelemetry. Set `ANTHROPIC_API_KEY` in the deployment environment. Hopsworks injects the `OTEL_EXPORTER_OTLP_*` environment variables for the deployment, so the OpenTelemetry exporter can stay configuration-free. ```python import asyncio import os import uvicorn from fastapi import FastAPI from llama_index.core.agent.workflow import ReActAgent from llama_index.core.tools import FunctionTool from llama_index.llms.anthropic import Anthropic from openinference.instrumentation.llama_index import LlamaIndexInstrumentor from opentelemetry.exporter.otlp.proto.http.trace_exporter import OTLPSpanExporter from opentelemetry.instrumentation.fastapi import FastAPIInstrumentor from opentelemetry.sdk import trace as trace_sdk from opentelemetry.sdk.trace.export import SimpleSpanProcessor def add(a: float, b: float) -> float: """Adds two numbers.""" return a + b def subtract(a: float, b: float) -> float: """Subtracts two numbers.""" return a - b def multiply(a: float, b: float) -> float: """Multiplies two numbers.""" return a * b def divide(a: float, b: float) -> float: """Divides two numbers.""" if b == 0: raise ValueError("Cannot divide by zero.") return a / b def build_tracer_provider(): endpoint = os.environ.get( "OTEL_EXPORTER_OTLP_TRACES_ENDPOINT", "http://localhost:4318/v1/traces", ) tracer_provider = trace_sdk.TracerProvider() tracer_provider.add_span_processor( SimpleSpanProcessor(OTLPSpanExporter(endpoint=endpoint)) ) return tracer_provider class AgentPredictor: def __init__(self): self.tracer_provider = build_tracer_provider() LlamaIndexInstrumentor().instrument(tracer_provider=self.tracer_provider) llm = Anthropic( model="claude-haiku-4-5-20251001", max_tokens=1024, temperature=0.0, ) tools = [ FunctionTool.from_defaults(add), FunctionTool.from_defaults(subtract), FunctionTool.from_defaults(multiply), FunctionTool.from_defaults(divide), ] self.agent = ReActAgent( tools=tools, llm=llm, ) async def _predict_async(self, inputs): prompt = inputs.get("prompt", "") result = await self.agent.run(prompt) return {"answer": str(result)} def predict(self, inputs): return asyncio.run(self._predict_async(inputs)) predictor = AgentPredictor() agent_app = FastAPI() FastAPIInstrumentor.instrument_app( agent_app, tracer_provider=predictor.tracer_provider, ) @agent_app.post("/query") def query(payload: dict): return predictor.predict(payload) if __name__ == "__main__": uvicorn.run(agent_app, host="0.0.0.0", port=8080) ``` ## Tracing Agent Deployments can be configured with OpenTelemetry tracing. When tracing is enabled, Hopsworks automatically provisions four online, Delta-backed feature groups in the project's Feature Store: - `otel_spans` - root spans and trace summary fields - `otel_span_attributes` - span attributes as key-value pairs - `otel_events` - span events - `otel_event_attributes` - event attributes as key-value pairs The Traces UI reads from these feature groups, and the first traced deployment in a project creates them automatically if they do not already exist. After that, choose one of these storage modes: - `online` - the default; writes traces to only online - `offline` - writes traces to only offline. You will no be able to see the traces summaries in the UI, but you can use the hopsworks-api to read the offline feature groups and reconstruct the traces from there. - `both` - export traces to both online and offline feature groups. This is the recommended option for production deployments, as it allows you to see the traces in the UI and also have them stored cost-effectively for long-term retention. ## Next steps - Scheduled, non-interactive coding agent: [Agent Tasks](../tasks/index.md) - Model-backed online predictor: use the Model Deployments guides under MLOps - Agent-serving dependencies: see the environment guides for cloning Python environments and installing requirements ================================================================================ # Migration 3.X to 4.0 Source: https://docs.hopsworks.ai/latest/user_guides/migration/40_migration/ # 4.0 Migration Guide ## Breaking Changes With the release of Hopsworks 4.0, a number of necessary breaking changes have been put in place to improve the overall experience of using the Hopsworks platform. These breaking changes can be categorized in the following areas: - Python API - Multi-Environment Docker Images - On-Demand Transformation Functions ### Python API A number of significant changes have been made in the Python API Hopsworks 4.0. Previously, in Hopsworks 3.X, there were 3 python libraries used (“hopsworks”, “hsfs” & “hsml”) to develop feature, training & inference pipelines, with the 4.0 release there is now one single “hopsworks” python library that should be used. For backwards compatibility, it is still possible to import both the “hsfs” & “hsml” packages directly, but the proper way to import them is to use “hopsworks.hsfs” & “hopsworks.hsml”. The direct imports will be deprecated later. Another significant change in the Hopsworks Python API is the use of optional extras to allow a developer to easily import exactly what is needed as part of their work. The main ones are great-expectations and polars. It is arguable whether this is a breaking change but it is important to note depending on how a particular pipeline has been written which may encounter a problem when executing using Hopsworks 4.0. Finally, there are a number of relatively small breaking changes and deprecated methods to improve the developer experience, these include: - connection.init() is now considered deprecated - When loading arrow_flight_client, an OptionalDependencyNotFoundError can be now thrown providing more detailed information on the error than the previous ModuleNotFoundError in 3.X. - DatasetApi's zip and unzip will now return False when a timeout is exceeded instead of previously throwing an Exception ### Multi-Environment Docker Images As part of the Hopsworks 4.0 release, an engineering team using Hopsworks can now customize the docker images that they use for their feature, training and inference pipelines. By adding this flexibility, a set of breaking changes are necessary. Instead of having one common docker image for fti pipelines, with the release of 4.0 a number of specific docker images are provided to allow an engineering team using Hopsworks to install exactly what they need to get their feature, training and inference pipelines up and running. This breaking change will require existing customers running Hopsworks 3.X to test their existing pipelines using Hopsworks 4.0 before upgrading their production environments. ### On-Demand Transformation Functions A number of changes have been made to transformation functions in the last releases of Hopsworks. With 4.0, On-Demand Transformation Functions are now better supported which has resulted in some breaking changes. The following is how transformation functions were used in previous versions of Hopsworks and the how transformation functions are used in the 4.0 release. === "Pre-4.0" ```python ################################################# # Creating transformation function Hopsworks 3.8# ################################################# # Define custom transformation function def add_one(feature): return feature + 1 # Create transformation function add_one = fs.create_transformation_function( add_one, output_type=int, version=1, ) # Save transformation function add_one.save() # Retrieve transformation function scaler = fs.get_transformation_function( name="add_one", version=1, ) # Create feature view feature_view = fs.get_or_create_feature_view( name="serving_fv", version=1, query=selected_features, # Apply your custom transformation functions to the feature `feature_1` transformation_functions={ "feature_1": add_one, }, labels=["target"], ) ``` === "4.0" ```python ################################################# # Creating transformation function Hopsworks 4.0# ################################################# # Define custom transformation function @hopsworks.udf(int) def add_one(feature): return feature + 1 # Create feature view feature_view = fs.get_or_create_feature_view( name="serving_fv", version=1, query=selected_features, # Apply the custom transformation functions defined to the feature `feature_1` transformation_functions=[ add_one("feature_1"), ], labels=["target"], ) ``` Note that the number of lines of code required has been significantly reduced using the “@hopsworks.udf” python decorator. ================================================================================ # Setup and Administration Source: https://docs.hopsworks.ai/latest/setup_installation/ # Setup and Administration Hopsworks runs on Kubernetes, on a cloud provider or on your own hardware. This section covers installing the platform and administering it afterwards. For the client libraries, see the [Client Installation](../user_guides/client_installation/index.md) guide.
- :material-kubernetes:{ .lg .middle } **Start here** --- Pick your environment and follow its getting started guide. Each one takes you from an empty account to a running cluster with a first project. [AWS](aws/getting_started.md) · [Azure](azure/getting_started.md) · [GCP](gcp/getting_started.md) · [On-prem](on_prem/contact_hopsworks.md)
:material-cloud-outline:{ .hops-role-ico } Install { .hops-role-cap } - [AWS, Azure, GCP](aws/getting_started.md) Managed Kubernetes on each cloud, with the storage and network the cluster needs. - [On-prem](on_prem/contact_hopsworks.md) Your own Kubernetes, optionally with an external Kafka cluster. - [Cluster configuration](admin/variables.md) Configuration variables, with the full reference and build performance notes.
:material-account-group-outline:{ .hops-role-ico } Users and access { .hops-role-cap } - [Users and projects](admin/user.md) Approve users, assign roles, manage projects and quotas. - [Authentication](admin/auth.md) OAuth2 identity providers, LDAP and Kerberos, with project mapping. - [IAM role chaining](admin/roleChaining.md) Let projects assume AWS roles.
:material-shield-check-outline:{ .hops-role-ico } Operate { .hops-role-cap } - [Monitoring](admin/monitoring/grafana.md) Service dashboards, exported metrics and service logs. - [High availability and disaster recovery](admin/ha-dr/intro.md) Replicated services and backup, restore procedures. - [Audit and operations](admin/audit/audit-logs.md) Access audit logs, service operations, search index. - [Alerts, Trino, Superset](admin/alert.md) Cluster-wide alert receivers and the analytics services.
================================================================================ # AWS - Getting Started Source: https://docs.hopsworks.ai/latest/setup_installation/aws/getting_started/ # AWS - Getting started with EKS Kubernetes and Helm are used to install and run Hopsworks and the Feature Store in the cloud. They both integrate seamlessly with third-party platforms such as Databricks, SageMaker and Kubeflow. This guide shows how to set up the Hopsworks platform on Amazon EKS in your organization's AWS account. ## Prerequisites To follow the instructions on this page you will need the following: - Kubernetes version: Hopsworks supports EKS clusters running Kubernetes >= 1.27.0 and < 1.35.0 (the chart's supported range, enforced by its `kubeVersion` constraint). Pick a version in that range that is also under [AWS standard support](https://docs.aws.amazon.com/eks/latest/userguide/kubernetes-versions.html), for example 1.33. - [aws-cli](https://aws.amazon.com/cli/) to provision the AWS resources. - [eksctl](https://eksctl.io/) to create and manage the EKS cluster. - [kubectl](https://kubernetes.io/docs/tasks/tools/) to interact with the cluster. - [helm](https://helm.sh/) to deploy Hopsworks. ### Permissions - The deployment requires cluster admin access to create ClusterRoles, ServiceAccounts, and ClusterRoleBindings. - A namespace is required to deploy the Hopsworks stack. If you don't have permissions to create a namespace, ask your EKS administrator to provision one. ## Step 1: AWS EKS setup The following steps describe how to deploy an EKS cluster and the related resources so that it is compatible with Hopsworks. ### Step 1.1: Create an S3 bucket Create a bucket to store project data. ```bash aws s3 mb s3://BUCKET_NAME --region REGION --profile PROFILE ``` Enable versioning on the bucket. The RonDB and OpenSearch backups, and HopsFS, write to this bucket and rely on object versioning, so backups do not work correctly without it: ```bash aws s3api put-bucket-versioning --bucket BUCKET_NAME --versioning-configuration Status=Enabled --profile PROFILE ``` ### Step 1.2: Create an ECR repository Hopsworks allows users to customize the images used by Python jobs, Jupyter notebooks and (Py)Spark applications running in their projects. These images are stored in ECR, so Hopsworks needs access to an ECR repository to push the project images. Create the repository to host the project images: ```bash aws --profile PROFILE ecr create-repository --repository-name NAMESPACE/hopsworks-base --region REGION ``` ### Step 1.3: Create IAM policies Create a policy that grants access to the S3 bucket and the ECR repository. Save the following document as `policy.json`, replacing BUCKET_NAME, REGION, and ECR_AWS_ACCOUNT_ID with your values: ```json { "Version": "2012-10-17", "Statement": [ { "Sid": "hopsworksaiInstanceProfile", "Effect": "Allow", "Action": [ "s3:PutObject", "s3:ListBucket", "s3:GetObject", "s3:DeleteObject", "s3:AbortMultipartUpload", "s3:ListBucketMultipartUploads", "s3:PutLifecycleConfiguration", "s3:GetLifecycleConfiguration", "s3:PutBucketVersioning", "s3:GetBucketVersioning", "s3:ListBucketVersions", "s3:DeleteObjectVersion" ], "Resource": [ "arn:aws:s3:::BUCKET_NAME/*", "arn:aws:s3:::BUCKET_NAME" ] }, { "Sid": "AllowPushandPullImagesToUserRepo", "Effect": "Allow", "Action": [ "ecr:GetDownloadUrlForLayer", "ecr:BatchGetImage", "ecr:CompleteLayerUpload", "ecr:UploadLayerPart", "ecr:InitiateLayerUpload", "ecr:BatchCheckLayerAvailability", "ecr:PutImage", "ecr:ListImages", "ecr:BatchDeleteImage", "ecr:GetLifecyclePolicy", "ecr:PutLifecyclePolicy", "ecr:TagResource" ], "Resource": [ "arn:aws:ecr:REGION:ECR_AWS_ACCOUNT_ID:repository/*/hopsworks-base" ] } ] } ``` Create the policy: ```bash aws --profile PROFILE iam create-policy --policy-name POLICY_NAME --policy-document file://policy.json ``` This single policy grants both S3 (bucket) and ECR (repository) access. Because the Helm values do not store any S3 access keys, Hopsworks and HopsFS reach the bucket through the AWS default credential provider chain, which on EKS resolves to the worker node instance IAM role. Attaching this policy to the node group (Step 1.4) is therefore how Hopsworks is granted access to S3 and ECR: no access key or secret is stored in Kubernetes. ### Step 1.4: Create the EKS cluster using eksctl When creating the cluster with eksctl, the following are required in the cluster configuration file (`eksctl.yaml`): - `amiFamily` should be either `AmazonLinux2023` or `Ubuntu2404`. Amazon Linux 2 reaches end of support on June 30, 2026, so do not use it. - The instance type should be Intel or AMD based (that is, not ARM/Graviton). - The following managed policies are required (see [attaching policies by ARN](https://eksctl.io/usage/iam-policies/#attaching-policies-by-arn)): ```bash - arn:aws:iam::aws:policy/AmazonEKSWorkerNodePolicy - arn:aws:iam::aws:policy/AmazonEKS_CNI_Policy - arn:aws:iam::aws:policy/AmazonEC2ContainerRegistryReadOnly ``` To let Hopsworks provision the necessary load balancers through the [AWS Load Balancer Controller](https://kubernetes-sigs.github.io/aws-load-balancer-controller/latest/), grant the controller permissions by enabling the add-on policy on the node group: ```bash withAddonPolicies: awsLoadBalancerController: true ``` Update CLUSTER_NAME, REGION, ECR_AWS_ACCOUNT_ID, and the policy name (POLICY_NAME) created above, then create the configuration file: ```yaml apiVersion: eksctl.io/v1alpha5 kind: ClusterConfig metadata: name: CLUSTER_NAME region: REGION version: "1.33" iam: withOIDC: true managedNodeGroups: - name: ng-1 amiFamily: AmazonLinux2023 instanceType: m6i.2xlarge minSize: 1 maxSize: 5 desiredCapacity: 5 volumeSize: 100 # ssh: # optional: enable SSH access to the nodes # allow: true # publicKeyPath: ~/.ssh/id_ed25519.pub iam: attachPolicyARNs: - arn:aws:iam::aws:policy/AmazonEKSWorkerNodePolicy - arn:aws:iam::aws:policy/AmazonEKS_CNI_Policy - arn:aws:iam::aws:policy/AmazonEC2ContainerRegistryReadOnly - arn:aws:iam::ECR_AWS_ACCOUNT_ID:policy/POLICY_NAME withAddonPolicies: awsLoadBalancerController: true addons: - name: aws-ebs-csi-driver wellKnownPolicies: # add IAM role and service account ebsCSIController: true ``` Create the EKS cluster: ```bash eksctl create cluster -f eksctl.yaml --profile PROFILE ``` Once the cluster is created, verify that you can access it with the kubectl CLI tool: ```bash kubectl get nodes ``` You should see the list of nodes provisioned for the cluster. ### Step 1.5: Install the AWS Load Balancer Controller For Hopsworks to provision the network and application load balancers it needs, install the AWS Load Balancer Controller (see the [AWS documentation](https://docs.aws.amazon.com/eks/latest/userguide/lbc-helm.html)): ```bash helm repo add eks https://aws.github.io/eks-charts helm repo update eks helm install aws-load-balancer-controller eks/aws-load-balancer-controller -n kube-system --set clusterName=CLUSTER_NAME ``` The controller uses the IAM permissions attached to the node group through the `awsLoadBalancerController` add-on policy enabled above. Alternatively, you can run it with a dedicated service account using [IAM Roles for Service Accounts (IRSA)](https://docs.aws.amazon.com/eks/latest/userguide/lbc-helm.html). If the controller cannot auto-detect the VPC, pass `--set region=REGION` and `--set vpcId=VPC_ID`. ### Step 1.6: Create a GP3 storage class EKS clusters typically provide `gp2` as the default storage class. `gp3` is more cost effective and is the storage class referenced by the Helm values below (`storageClassName: ebs-gp3`), so create it: ```bash kubectl apply -f - <= 1.27.0. - An Azure resource group in which the Hopsworks cluster will be deployed. - The [azure CLI](https://learn.microsoft.com/en-us/cli/azure/install-azure-cli) installed and [logged in](https://docs.microsoft.com/en-us/cli/azure/authenticate-azure-cli). - kubectl (to manage the AKS cluster) - helm (to deploy Hopsworks) ### Permissions The deployment requires cluster admin access to create ClusterRoles, ServiceAccounts, and ClusterRoleBindings in AKS. A namespace is also required for deploying the Hopsworks stack. If you don’t have permissions to create a namespace, ask your AKS administrator to provision one for you. To run all the commands on this page the user needs to have at least the following permissions on the Azure resource group: You will also need to have a role such as *Application Administrator* on the Azure Active Directory to be able to create the hopsworks.ai service principal. ## Step 1: Azure Kubernetes Service (AKS) Setup ### Step 1.1: Create an Azure Blob Storage Account Create a storage account to host project data. Ensure that the storage account is in the same region as the AKS cluster for performance and cost reasons: ```bash az storage account create --name $STORAGE_ACCOUNT_NAME --resource-group $RESOURCE_GROUP --location $REGION ``` Also, create the corresponding container: ```bash az storage container create --account-name $STORAGE_ACCOUNT_NAME --name $CONTAINER_NAME ``` ### Step 1.2: Create an Azure Container Registry (ACR) Create an ACR to store the images used by Hopsworks: ```bash az acr create --resource-group $RESOURCE_GROUP --name $CONTAINER_REGISTRY_NAME --sku Basic --location $REGION export ACR_ID=`az acr show --name $CONTAINER_REGISTRY_NAME --resource-group $RESOURCE_GROUP --query "id" --output tsv` ``` ### Step 1.3: Create a User-Assigned Managed Identity Create a user-assigned managed identity to grant AKS access to the storage account and container registry: ```bash az identity create --name $UA_IDENTITY_NAME --resource-group $RESOURCE_GROUP export UA_IDENTITY_PRINCIPAL_ID=`az identity show --name $UA_IDENTITY_NAME --resource-group $RESOURCE_GROUP --query principalId --output tsv` export UA_IDENTITY_CLIENT_ID=`az identity show --name $UA_IDENTITY_NAME --resource-group $RESOURCE_GROUP --query clientId --output tsv` export UA_IDENTITY_RESOURCE_ID=`az identity show --name $UA_IDENTITY_NAME --resource-group $RESOURCE_GROUP --query id --output tsv` ``` ### Step 1.4: Grant permissions to the User-Assigned Managed Identity Create a custom role definition with the minimum permissions needed to read and write to the storage account: ```bash export STORAGE_ID=`az storage account show --name $STORAGE_ACCOUNT_NAME --resource-group $RESOURCE_GROUP --query "id" --output tsv` az role definition create --role-definition '{ "Name": "hopsfs-storage-permissions", "IsCustom": true, "Description": "Allow HopsFS to access the storage container", "Actions": [ "Microsoft.Storage/storageAccounts/blobServices/containers/write", "Microsoft.Storage/storageAccounts/blobServices/containers/read", "Microsoft.Storage/storageAccounts/blobServices/write", "Microsoft.Storage/storageAccounts/blobServices/read", "Microsoft.Storage/storageAccounts/listKeys/action" ], "NotActions": [], "DataActions": [ "Microsoft.Storage/storageAccounts/blobServices/containers/blobs/delete", "Microsoft.Storage/storageAccounts/blobServices/containers/blobs/read", "Microsoft.Storage/storageAccounts/blobServices/containers/blobs/move/action", "Microsoft.Storage/storageAccounts/blobServices/containers/blobs/write" ], "AssignableScopes": [ "'$STORAGE_ID'" ] }' az role assignment create --role hopsfs-storage-permissions --assignee-object-id $UA_IDENTITY_PRINCIPAL_ID --assignee-principal-type ServicePrincipal --scope $STORAGE_ID ``` ### Step 1.5: Create Service Principal for Hopsworks services Create a service principal to grant Hopsworks applications with access to the container registry. For example, Hopsworks uses this service principal to push new Python environments created via the Hopsworks UI. ```bash export SP_PASSWORD=`az ad sp create-for-rbac --name $SP_NAME --scopes $ACR_ID --role AcrPush --years 1 --query "password" --output tsv` export SP_USER_NAME=`az ad sp list --display-name $SP_NAME --query "[].appId" --output tsv` export SP_RESOURCE_ID=`az ad sp list --display-name $SP_NAME --query "[].id" --output tsv` az role assignment create --role AcrDelete --assignee-object-id $SP_RESOURCE_ID --assignee-principal-type ServicePrincipal --scope $ACR_ID ``` ### Step 1.6: Create an AKS Kubernetes Cluster Provision an AKS cluster with a number of nodes: ```bash az aks create --resource-group $RESOURCE_GROUP --name $KUBERNETES_CLUSTER_NAME --network-plugin azure \ --enable-cluster-autoscaler --min-count 1 --max-count 4 --node-count 3 --node-vm-size Standard_D8_v4 \ --attach-acr $CONTAINER_REGISTRY_NAME \ --assign-identity $UA_IDENTITY_RESOURCE_ID --assign-kubelet-identity $UA_IDENTITY_RESOURCE_ID \ --enable-managed-identity --generate-ssh-keys ``` ## Step 2: Configure kubectl ```bash az aks get-credentials --resource-group $RESOURCE_GROUP --name $KUBERNETES_CLUSTER_NAME --file ~/my-aks-kubeconfig.yaml export KUBECONFIG=~/my-aks-kubeconfig.yaml kubectl config current-context ``` ## Step 3: Create Secret for the Service Principal ### Step 3.1: Create Hopsworks namespace ```bash kubectl create namespace hopsworks ``` ### Step 3.2: Create secret ```bash kubectl create secret docker-registry azregcred \ --namespace hopsworks \ --docker-server=$CONTAINER_REGISTRY_NAME.azurecr.io \ --docker-username=$SP_USER_NAME \ --docker-password=$SP_PASSWORD ``` ## Step 4: Setup Hopsworks for Deployment ### Step 4.1: Add the Hopsworks Helm repository To obtain access to the Hopsworks helm chart repository, please [obtain](https://www.hopsworks.ai/try) an evaluation/startup licence. Once you have the helm chart repository URL, replace the environment variable $HOPSWORKS_REPO in the following command with this URL. ```bash helm repo add hopsworks $HOPSWORKS_REPO helm repo update hopsworks ``` ### Step 4.2: Create helm values file Below is a simplifield values.azure.yaml file to get started which can be updated for improved performance and further customisation. See the [Helm chart values reference][helm-chart-values-reference] for the full list of configurable values. ```yaml global: _hopsworks: storageClassName: null cloudProvider: "AZURE" managedDockerRegistery: enabled: true domain: "CONTAINER_REGISTRY_NAME.azurecr.io" namespace: "hopsworks" credHelper: enabled: false secretName: "" minio: enabled: false hopsworks: variables: docker_operations_managed_docker_secrets: &azregcred "azregcred" docker_operations_image_pull_secrets: *azregcred dockerRegistry: preset: usePullPush: false secrets: - *azregcred hopsfs: objectStorage: enabled: true provider: "AZURE" azure: storage: account: "STORAGE_ACCOUNT_NAME" container: "STORAGE_ACCOUNT_CONTAINER_NAME" identityClientId: "UA_IDENTITY_CLIENT_ID" ``` ## Step 5: Deploy Hopsworks Deploy Hopsworks in the created namespace. ```bash helm install hopsworks hopsworks/hopsworks --namespace hopsworks --values values.azure.yaml --timeout=600s ``` Check that Hopsworks is installing on your provisioned AKS cluster. ```bash kubectl get pods --namespace=hopsworks kubectl get svc -n hopsworks -o wide ``` Upon completion (circa 20 minutes), setup a load balancer to access Hopsworks: ```bash kubectl expose deployment hopsworks --type=LoadBalancer --name=hopsworks-service --namespace ``` ## Step 6: Next steps Check out our other guides for how to get started with Hopsworks and the Feature Store: - Get started with the [Hopsworks Feature Store](https://colab.research.google.com/github/logicalclocks/hopsworks-tutorials/blob/master/quickstart.ipynb){:target="_blank"} - Follow one of our [guides](../../user_guides/index.md) ================================================================================ # GCP - Getting Started Source: https://docs.hopsworks.ai/latest/setup_installation/gcp/getting_started/ # GCP - Getting started with GKE Kubernetes and Helm are used to install & run Hopsworks and the Feature Store in the cloud. They both integrate seamlessly with third-party platforms such as Databricks, SageMaker and KubeFlow. This guide shows how to set up the Hopsworks platform in your organization's Google Cloud Platform's (GCP) account. ## Prerequisites To follow the instruction on this page you will need the following: - Kubernetes Version: Hopsworks can be deployed on GKE clusters running Kubernetes >= 1.27.0. - [gcloud CLI](https://cloud.google.com/sdk/gcloud) to provision the GCP resources - [gke-gcloud-auth-plugin](https://cloud.google.com/blog/products/containers-kubernetes/kubectl-auth-changes-in-gke) to manage authentication with the GKE cluster - [helm](https://helm.sh/) to deploy Hopsworks ### Permissions - The deployment requires cluster admin access to create ClusterRoles, ServiceAccounts, and ClusterRoleBindings. - A namespace is required to deploy the Hopsworks stack. If you don’t have permissions to create a namespace, ask your GKE administrator to provision one. ## Step 1: GCP GKE Setup ### Step 1.1: Create a Google Cloud Storage (GCS) bucket Create a bucket to store project data. Ensure the bucket is in the same region as your GKE cluster for performance and cost optimization. ```bash gsutil mb -l $region gs://$bucket_name ``` ### Step 1.2: Create Service Account Create a file named `hopsworksai_role.yaml` with the following content: ```bash title: Hopsworks AI Instances description: Role that allows Hopsworks AI Instances to access resources stage: GA includedPermissions: - storage.buckets.get - storage.buckets.update - storage.multipartUploads.abort - storage.multipartUploads.create - storage.multipartUploads.list - storage.multipartUploads.listParts - storage.objects.create - storage.objects.delete - storage.objects.get - storage.objects.list - storage.objects.update - artifactregistry.repositories.create - artifactregistry.repositories.get - artifactregistry.repositories.uploadArtifacts - artifactregistry.repositories.downloadArtifacts - artifactregistry.tags.list - artifactregistry.tags.delete ``` Execute the following gcloud command to create a custom role from the file. Replace $PROJECT_ID with your GCP project id: ```bash gcloud iam roles create hopsworksai_instances \ --project=$PROJECT_ID \ --file=hopsworksai_role.yaml ``` Execute the following gcloud command to create a service account for Hopsworks AI instances. Replace $PROJECT_ID with your GCP project id: ```bash gcloud iam service-accounts create hopsworksai_instances \ --project=$PROJECT_ID \ --description="Service account for Hopsworks AI instances" \ --display-name="Hopsworks AI instances" ``` Execute the following gcloud command to bind the custom role to the service account. Replace all occurrences $PROJECT_ID with your GCP project id: ```bash gcloud projects add-iam-policy-binding $PROJECT_ID \ --member="serviceAccount:hopsworks-ai-instances@$PROJECT_ID.iam.gserviceaccount.com" \ --role="projects/$PROJECT_ID/roles/hopsworksai_instances" ``` ### Step 1.3: Create a GKE Cluster ```bash gcloud container clusters create \ --zone \ --machine-type n2-standard-8 \ --num-nodes 1 \ --enable-ip-alias \ --service-account my-service-account@my-project.iam.gserviceaccount.com ``` Once the creation process is completed, you should be able to access the cluster using the kubectl CLI tool: ```bash kubectl get nodes ``` ### Step 1.4: Create GCR repository Hopsworks allows users to customize images for Python jobs, Jupyter Notebooks, and (Py)Spark applications. These images should be stored in Google Container Registry (GCR). The GKE cluster needs access to a GCR repository to push project images. Enable Artifact Registry and create a GCR repository to store images: ```bash gcloud artifacts repositories create \ --repository-format=docker \ --location= ``` ## Step 3: Setup Hopsworks for Deployment ### Step 3.1: Add the Hopsworks Helm repository To obtain access to the Hopsworks helm chart repository, please [obtain](https://www.hopsworks.ai/try) an evaluation/startup licence. Once you have the helm chart repository URL, replace the environment variable $HOPSWORKS_REPO in the following command with this URL. ```bash helm repo add hopsworks $HOPSWORKS_REPO helm repo update hopsworks ``` ### Step 3.2: Create Hopsworks namespace ```bash kubectl create namespace hopsworks ``` ### Step 3.3: Create helm values file Below is a simplifield values.gcp.yaml file to get started which can be updated for improved performance and further customisation. See the [Helm chart values reference][helm-chart-values-reference] for the full list of configurable values. ```bash global: _hopsworks: storageClassName: null cloudProvider: "GCP" managedDockerRegistery: enabled: true domain: "europe-north1-docker.pkg.dev" namespace: "PROJECT_ID/hopsworks" credHelper: enabled: true secretName: &gcpregcred "gcpregcred" managedObjectStorage: enabled: true s3: bucket: name: &bucket "hopsworks" region: ®ion "europe-north1" endpoint: &gcpendpoint "https://storage.cloud.google.com" secret: name: &gcpcredentials "gcp-credentials" access_key_id: &gcpaccesskey "access-key-id" secret_key_id: &gcpsecretkey "secret-access-key" minio: enabled: false ``` ## Step 4: Deploy Hopsworks Deploy Hopsworks in the created namespace. ```bash helm install hopsworks hopsworks/hopsworks \ --namespace hopsworks \ --values values.gcp.yaml \ --timeout=600s ``` Check that Hopsworks is installing on your provisioned AKS cluster. ```bash kubectl get pods --namespace=hopsworks kubectl get svc -n hopsworks -o wide ``` Upon completion (circa 20 minutes), setup a load balancer to access Hopsworks: ```bash kubectl expose deployment hopsworks --type=LoadBalancer --name=hopsworks-service --namespace ``` ## Step 5: Next steps Check out our other guides for how to get started with Hopsworks and the Feature Store: - Get started with the [Hopsworks Feature Store](https://colab.research.google.com/github/logicalclocks/hopsworks-tutorials/blob/master/quickstart.ipynb){:target="_blank"} - Follow one of our [guides](../../user_guides/index.md) ================================================================================ # Background Source: https://docs.hopsworks.ai/latest/setup_installation/on_prem/contact_hopsworks/ # Hopsworks On-Premise Installation It is possible to use Hopsworks on-premises, which means that companies can run their machine learning workloads on their own hardware and infrastructure, rather than relying on a cloud provider. This can provide greater flexibility, control, and cost savings, as well as enabling companies to meet specific compliance and security requirements. Working on-premises with Hopsworks typically involves collaboration with the Hopsworks engineering teams, as each infrastructure is unique and requires a tailored approach to deployment and configuration. The process begins with an assessment of the company's existing infrastructure and requirements, including network topology, security policies, and hardware specifications. For further details about on-premise installations; [contact us](https://www.hopsworks.ai/contact). ================================================================================ # External Kafka cluster Source: https://docs.hopsworks.ai/latest/setup_installation/on_prem/external_kafka_cluster/ # External Kafka cluster Hopsworks uses [Apache Kafka](https://kafka.apache.org/) to ingest data to the feature store. Streaming applications and external clients send data to the Kafka cluster for ingestion to the online and offline feature store. By default, Hopsworks comes with an embedded Kafka cluster managed by Hopsworks itself, however, users can configure Hopsworks to leverage an existing external cluster. This guide will cover how to configure an Hopsworks cluster to leverage an external Kafka cluster. !!! warning "Broker version requirement" Hopsworks 5.2 connects to Kafka with kafka-clients 4.3.1, which only speaks to brokers running Kafka 2.1 or newer ([KIP-896](https://cwiki.apache.org/confluence/display/KAFKA/KIP-896%3A+Remove+old+client+protocol+API+versions+in+Kafka+4.0)). An external cluster on an older broker version stops working after the upgrade to 5.2, so upgrade the brokers first. See the [Kafka 4 and KRaft upgrade notes][kafka4-client-compatibility]. ## Configure the external Kafka cluster integration To enable the integration with an external Kafka cluster, you should set the `enable_bring_your_own_kafka` [configuration option](../admin/variables.md) to `true`. This can also be achieved in the cluster definition by setting the following attribute: ```yaml hopsworks: enable_bring_your_own_kafka: "true" ``` ### Online Feature Store service configuration In addition to the configuration changes above, you should also configure the Online Feature Store service (OnlineFS in short) to connect to the external Kafka cluster. This can be achieved by provisioning the necessary credentials for OnlineFS to subscribe and consume messages from Kafka topics used by the Hopsworks feature store. OnlineFs can be configured to use these credentials by adding the following configurations to the cluster definition used to deploy Hopsworks: ```yaml onlinefs: config_dir: "/home/ubuntu/cluster-definitions/byok" kafka_consumers: topic_list: "comma separated list of kafka topics to subscribe to" ``` In particular, the `onlinefs/config_dir` should contain the credentials necessary for the Kafka consumers to authenticate. Additionally the directory should contain a file name `onlinefs-kafka.properties` with the Kafka consumer configuration. The following is an example of the `onlinefs-kafka.properties` file: ```properties bootstrap.servers=cluster_identifier.us-east-2.aws.confluent.cloud:9092 security.protocol=SASL_SSL sasl.jaas.config=org.apache.kafka.common.security.plain.PlainLoginModule required username="username" password="password"; sasl.mechanism=PLAIN ``` !!! note "Hopsworks will not provision topics" Please note that when using an external Kafka cluster, Hopsworks will not provision the topics for the different projects. Users are responsible for provisioning the necessary topics and configure the projects accordingly (see next section). Users should also specify the list of topics OnlineFS should subscribe to by providing the `onlinefs/kafka_consumers/topic_list` option in the cluster definition. ### Project configuration #### Topic configuration As mentioned above, when configuring Hopsworks to use an external Kafka cluster, Hopsworks will not provision the topics for the different projects. Instead, when creating a project, users will be asked to provide the topic name to use for the feature store operations. The topic can be changed later from `Project Settings` → `Kafka`, and individual feature groups can override it, as described in the [ingestion topic][feature-group-ingestion-topic] guide.

Example project creation when using an external Kafka cluster
Example project creation when using an external Kafka cluster

#### Data Source configuration Users should create a [Kafka Data Source](../../user_guides/fs/data_source/creation/kafka.md) named `kafka_connector` which is going to be used by the feature store clients to configure the necessary Kafka producers to send data. The configuration is done for each project to ensure its members have the necessary authentication/authorization. If the data source is not found in the project, default values referring to Hopsworks managed Kafka will be used. ================================================================================ # Cluster Administration Source: https://docs.hopsworks.ai/latest/setup_installation/admin/ # Cluster Administration Hopsworks has a cluster management page that allows you, the administrator, to perform management actions, monitor and control Hopsworks. To access the cluster management page you should log in into Hopsworks using your administrator account. In the top right corner, click on your name in the top right corner of the navigation bar and choose Cluster Settings from the dropdown menu. ================================================================================ # Cluster Configuration Source: https://docs.hopsworks.ai/latest/setup_installation/admin/variables/ # Cluster Configuration ## Introduction Whether you run Hopsworks on-premise, or on the cloud using kubernetes, it is possible to change a variety of configurations on the cluster, changing its default behaviour. This section is not going into detail for every setting, since every Hopsworks cluster comes with a robust default setup. However, this guide is to explain where to find the configurations and if necessary, how to change them. !!! note In most cases you will be only be prompted to change these configurations by a Hopsworks Solutions Engineer or similar. ## Prerequisites An administrator account on a Hopsworks cluster. ### Step 1: The configuration page You can find the configuration page by navigating in the UI: 1. Click on your user name in the top right corner, then select *Cluster Settings*. 2. In the left sidebar of the cluster settings, under *Infrastructure*, select *Configuration*
Configuration Settings
Configuration settings
### Step 2: Editing existing configurations To edit an existing configuration, simply find the property using the search field, then click the *edit* button to change the value of the setting or its visibility. Once you have made the change, don't forget to click *save* to persist the changes. #### Visibility The visibility setting indicates whether a setting can be read only by **Hops Admins** or also by simple **Hops Users**, that is everyone. Additionally, you can also allow to read the setting even when **not authenticated**. If the setting contains a password or sensitive information, you can also hide the value so it's not shown in the UI. ### Step 3: Adding a new configuration In rare cases it might be necessary to add additional configurations. To do so, click on *New Variable*, where you can then configure the new setting with a key, value and visibility. Once you have set the desired properties, you can persist them by clicking *Create Configuration*
Adding a new configuration property
Adding a new configuration property
================================================================================ # Cluster Configuration Variables Reference Source: https://docs.hopsworks.ai/latest/setup_installation/admin/configuration_reference/ # Cluster Configuration Variables Reference This page lists every cluster configuration key declared in the Hopsworks server source code, with its type and default value. Descriptions are not included here. The source declarations carry no description field, so adding one would mean inventing text that was never reviewed against the actual behaviour of the setting. Adding descriptions is a planned enhancement, pending a source-side change that adds a description argument to the settings declarations. Some defaults below are computed Java expressions rather than literal values, for example a reference to a constant defined in another class. Those rows show the raw expression and are marked as computed rather than showing a guessed value. A small number of keys are declared independently in more than one source file with their own default value. Those are shown as separate rows and marked as duplicates rather than merged into one, since the two declarations can drift apart. See [Cluster Configuration][cluster-configuration] for how to view and change these variables from the Hopsworks UI. | Key | Type | Default | Source module | | --- | --- | --- | --- | | `action_attempt_limit` | Integer | `3` | Settings.java | | `admin_email` | String | `admin@hopsworks.ai` | Settings.java | | `agent_base_image_name` | String | `agent-job` | Settings.java | | `agent_default_max_turns` | Integer | `50` | Settings.java | | `agent_default_model` | String | `claude-sonnet-4-5` | Settings.java | | `agent_deployment_otel_cpu` | Double | `0.5` | Settings.java | | `agent_deployment_otel_enabled` | Boolean | `true` | Settings.java | | `agent_deployment_otel_image` | String | `docker.hops.works/hops-otel:5.0.0-SNAPSHOT` | Settings.java | | `agent_deployment_otel_input_token_price_per_million` | Double | `0.0` | Settings.java | | `agent_deployment_otel_memory_mb` | Integer | `1024` | Settings.java | | `agent_deployment_otel_output_token_price_per_million` | Double | `0.0` | Settings.java | | `agent_deployment_otel_ttl_seconds` | Integer | `86400` | Settings.java | | `agent_image` | String | `agent-job` | Settings.java | | `agent_jobs_enabled` | Boolean | `true` | Settings.java | | `airflow_enabled` | Boolean | `true` | Settings.java | | `airflow_user` | String | `airflow` | Settings.java | | `alert_manager_config_map` | String | `hopsworks-release-alertmanager` | KubeSettings.java | | `alert_receiver_load_timeout` | Integer | `300` | Settings.java | | `anaconda_enabled` | Boolean | `true` | Settings.java | | `app_kill_grace_period_seconds` | Long | `2` | Settings.java | | `application_certificate_validity_period` | String | `3d` | CAConf.java | | `arrow_flight_read_timeout_as_ms` | Long | `180000` | Settings.java | | `arrow_libhdfs_dir` | String | `/usr/local/bin/libhdfs-golang` | Settings.java | | `async_services_timer_batch_size` | Integer | `1000` | Settings.java | | `async_services_timer_delete_history_after_days` | Long | `7` | Settings.java | | `async_services_timer_enabled` | Boolean | `true` | Settings.java | | `async_services_timer_interval_ms` | Long | `15000` | Settings.java | | `audit_log_count` | Integer | `10` | VariablesHelper.java | | `audit_log_date_format` | String | `yyyy-MM-dd HH:mm:ss` | VariablesHelper.java | | `audit_log_file_format` | String | `server_audit_log%g.log` | VariablesHelper.java | | `audit_log_file_path` | String | `/audit-logs` | VariablesHelper.java | | `audit_log_file_type` | String | `SimpleFormatter.class.getName()` *(computed expression, not a literal)* | VariablesHelper.java | | `audit_log_size_limit` | Integer | `256000000` | VariablesHelper.java | | `base_image_name` | String | `hopsworks-base` | Settings.java | | `base_image_version` | String | `4.3.0` | Settings.java | | `cert_mater_delay` | String | `1m` | Settings.java | | `certs_dir` | String | `/srv/hops/certs-dir` *(duplicate key, also declared in Settings.java; values match)* | CAConf.java | | `certs_dir` | String | `/srv/hops/certs-dir` *(duplicate key, also declared in CAConf.java; values match)* | Settings.java | | `client_path` | String | `/srv/hops/client.tar.gz` | Settings.java | | `cloud` | String | *(empty string)* | Settings.java | | `cloud_events_endpoint` | String | *(empty string)* *(duplicate key, also declared in Settings.java; values match)* | CAConf.java | | `cloud_events_endpoint` | String | *(empty string)* *(duplicate key, also declared in CAConf.java; values match)* | Settings.java | | `cloud_events_endpoint_api_key` | String | *(empty string)* | Settings.java | | `command_agent_home_batch` | Integer | `20` | Settings.java | | `command_agent_home_claim_lease_as_ms` | Long | `600000` | Settings.java | | `command_agent_home_migration_period_as_ms` | Long | `3600000` | Settings.java | | `command_agent_home_process_timer_period_as_ms` | Long | `5000` | Settings.java | | `command_agent_home_retry_backoff_base_as_ms` | Long | `10000` | Settings.java | | `command_agent_home_retry_backoff_max_as_ms` | Long | `600000` | Settings.java | | `command_search_fs_history_clean_period_as_ms` | Long | `60000` | Settings.java | | `command_search_fs_history_enable` | Boolean | `false` | Settings.java | | `command_search_fs_history_window_as_s` | Long | `3600` | Settings.java | | `command_search_fs_process_timer_period_as_ms` | Long | `1000` | Settings.java | | `command_search_fs_retry_backoff_base_as_ms` | Long | `5000` | Settings.java | | `command_search_fs_retry_backoff_max_as_ms` | Long | `300000` | Settings.java | | `conda_default_repo` | String | `defaults` | Settings.java | | `conda_env_name` | String | `hopsworks_environment` | Settings.java | | `created_by_label_name` | String | `hw-created-by` | KubeSettings.java | | `created_by_label_value` | String | `hopsworks` | KubeSettings.java | | `cross_project_global_search_enabled` | Boolean | `true` | Settings.java | | `csi_driver_enabled` | Boolean | `false` | Settings.java | | `csi_sidecar_image` | String | `docker.hops.works/hopsworks/hopsfs-csi:0.1.0-SNAPSHOT` | Settings.java | | `databricks_account_host_allowlist` | String | `accounts.cloud.databricks.com,accounts.azuredatabricks.net,accounts.gcp.databricks.com` | Settings.java | | `databricks_oauth_allow_private_ranges` | Boolean | `false` | Settings.java | | `default_feature_store_project_id` | Integer | `-1` | Settings.java | | `default_featurestore_project_name` | String | `hopsworks_default` | Settings.java | | `default_jupyter_environment` | String | `pandas-training-pipeline` | Settings.java | | `default_python_job_environment` | String | `pandas-training-pipeline` | Settings.java | | `disable_password_login` | Boolean | `false` | Settings.java | | `disable_registration` | Boolean | `false` | Settings.java | | `disable_registration_ui` | Boolean | `false` | Settings.java | | `dlt_schema_fetch_job_deadline_seconds` | Long | `1800` | Settings.java | | `dlthub-warehouse` | String | `dlthub-warehouse` | Settings.java | | `dlthub_image_name` | String | `docker.hops.works/hopsworks/dlt` | Settings.java | | `docker_base_image_dbt` | String | `dbt-pipeline` | Settings.java | | `docker_base_image_dlthub` | String | `dlthub-ingestion-pipeline` | Settings.java | | `docker_base_image_minimal_inference` | String | `minimal-inference-pipeline` | Settings.java | | `docker_base_image_pandas_inference` | String | `pandas-inference-pipeline` | Settings.java | | `docker_base_image_pandas_training` | String | `pandas-training-pipeline` | Settings.java | | `docker_base_image_python` | String | `python-feature-pipeline` | Settings.java | | `docker_base_image_python_app` | String | `python-app-pipeline` | Settings.java | | `docker_base_image_python_version` | String | `3.13` | Settings.java | | `docker_base_image_ray_tensorflow_training` | String | `ray-tensorflow-training-pipeline` | Settings.java | | `docker_base_image_ray_torch_training` | String | `ray-torch-training-pipeline` | Settings.java | | `docker_base_image_ray_training` | String | `ray-training-pipeline` | Settings.java | | `docker_base_image_spark` | String | `spark-feature-pipeline` | Settings.java | | `docker_base_image_tensorflow_inference` | String | `tensorflow-inference-pipeline` | Settings.java | | `docker_base_image_tensorflow_training` | String | `tensorflow-training-pipeline` | Settings.java | | `docker_base_image_torch_inference` | String | `torch-inference-pipeline` | Settings.java | | `docker_base_image_torch_training` | String | `torch-training-pipeline` | Settings.java | | `docker_base_image_vllm_inference` | String | `vllm-inference-pipeline` | Settings.java | | `docker_cgroup_cpu_period` | String | `100000` | Settings.java | | `docker_cgroup_enabled` | Boolean | `false` | Settings.java | | `docker_cgroup_parent` | String | `docker.slice` | Settings.java | | `docker_job_mount_allowed` | Boolean | `true` | Settings.java | | `docker_job_mounts_list` | String | *(empty string)* | Settings.java | | `docker_job_uid_strict` | Boolean | `true` | Settings.java | | `docker_mounts` | String | `/srv/hops/hadoop/etc/hadoop,/srv/hops/spark,/srv/hops/flink` | Settings.java | | `docker_namespace` | String | *(empty string)* | Settings.java | | `docker_operations_allow_hermetic_custom_commands` | Boolean | `false` | Settings.java | | `docker_operations_backoff_limit` | Integer | `0` | Settings.java | | `docker_operations_build_metadata` | Boolean | `true` | Settings.java | | `docker_operations_build_start_grace_minutes` | Integer | `2` | Settings.java | | `docker_operations_buildkit_addr` | String | *(empty string)* | Settings.java | | `docker_operations_buildkit_backoff_limit` | Integer | `0` | Settings.java | | `docker_operations_buildkit_cache_scope` | String | `shared` | Settings.java | | `docker_operations_buildkit_extra_args` | String | *(empty string)* | Settings.java | | `docker_operations_buildkit_image_root` | String | `docker.hops.works/hopsworks/moby/buildkit:v0.31.2` | Settings.java | | `docker_operations_buildkit_limit_cpu` | Integer | `1` | Settings.java | | `docker_operations_buildkit_limit_memory` | String | `2Gi` | Settings.java | | `docker_operations_buildkit_priority_class` | String | `rondb-high-priority` | Settings.java | | `docker_operations_buildkit_replicas` | Integer | `1` | Settings.java | | `docker_operations_buildkit_request_cpu` | Integer | `1` | Settings.java | | `docker_operations_buildkit_request_memory` | String | `2Gi` | Settings.java | | `docker_operations_buildkit_storage` | String | `2Gi` | Settings.java | | `docker_operations_buildkit_tls_locality` | String | `buildkitd` | Settings.java | | `docker_operations_buildkit_tls_secret` | String | *(empty string)* | Settings.java | | `docker_operations_cert_name` | String | `kagent_certificate_bundle.pem` | Settings.java | | `docker_operations_context_orphan_minutes` | Integer | `120` | Settings.java | | `docker_operations_context_prefix` | String | *(empty string)* | Settings.java | | `docker_operations_context_url_ttl_minutes` | Integer | `10` | Settings.java | | `docker_operations_crane_extra_args` | String | *(empty string)* | Settings.java | | `docker_operations_crane_image` | String | `docker.hops.works/hopsworks/hwutils:0.3` | Settings.java | | `docker_operations_delete_jobs_add_description_if_fails` | Boolean | `false` | Settings.java | | `docker_operations_delete_jobs_on_completion` | Boolean | `false` | Settings.java | | `docker_operations_docker_context_builder` | String | `Auto` | Settings.java | | `docker_operations_docker_context_builder_s3_bucket` | String | `hopsworks` | Settings.java | | `docker_operations_docker_context_builder_s3_endpoint` | String | `http://minio.hopsworks.svc.cluster.local:9000` | Settings.java | | `docker_operations_docker_context_builder_s3_region` | String | `eu-west-1` | Settings.java | | `docker_operations_hopsworks_ca_secret_name` | String | `docker-registry-crypto-material` | Settings.java | | `docker_operations_image_builder_image` | String | `docker.hops.works/hopsworks/image-builder:0.2` | Settings.java | | `docker_operations_image_pull_secrets` | String | *(empty string)* | Settings.java | | `docker_operations_lock_dependencies` | Boolean | `false` | Settings.java | | `docker_operations_managed_docker_secrets` | String | *(empty string)* | Settings.java | | `docker_operations_multi_region_copy` | Boolean | `false` | Settings.java | | `docker_operations_oci_worker_snapshotter` | String | `auto` | Settings.java | | `docker_operations_push_insecure` | Boolean | `false` | Settings.java | | `docker_operations_registry_container` | String | `docker` | Settings.java | | `docker_operations_registry_http` | Boolean | `false` | Settings.java | | `docker_operations_registry_pod` | String | `docker-registry-0` | Settings.java | | `docker_operations_sidecar_image` | String | `docker.hops.works/hopsworks/hwutils:0.3` | Settings.java | | `docker_operations_suspend_jobs` | Boolean | `false` | Settings.java | | `docker_operations_timeout_check_minutes` | Integer | `5` | Settings.java | | `docker_operations_timeout_delete_minutes` | Integer | `10` | Settings.java | | `docker_operations_timeout_export_minutes` | Integer | `5` | Settings.java | | `docker_operations_timeout_listing_minutes` | Integer | `10` | Settings.java | | `docker_operations_timeout_minutes_buildkit` | Integer | `30` | Settings.java | | `docker_operations_timeout_tag_minutes` | Integer | `5` | Settings.java | | `download_allowed` | Boolean | `true` | Settings.java | | `elastic_admin_password` | String | *(empty string)* | Settings.java | | `elastic_admin_user` | String | *(empty string)* | Settings.java | | `elastic_https_enabled` | Boolean | *(empty string)* | Settings.java | | `elastic_jwt_enabled` | Boolean | *(empty string)* | Settings.java | | `elastic_jwt_exp_ms` | Long | *(empty string)* | Settings.java | | `elastic_jwt_url_parameter` | String | *(empty string)* | Settings.java | | `elastic_logs_index_expiration` | Long | `604800000` | Settings.java | | `elastic_max_page_size` | Integer | `10000` | Settings.java | | `elastic_opendistro_security_enabled` | Boolean | *(empty string)* | Settings.java | | `elastic_scroll_page_size` | Integer | `1000` | Settings.java | | `elastic_version` | String | *(empty string)* | Settings.java | | `enable_adls_storage_connectors` | Boolean | `false` | Settings.java | | `enable_bigquery_storage_connectors` | Boolean | `false` | Settings.java | | `enable_bring_your_own_kafka` | Boolean | `false` | Settings.java | | `enable_crm_storage_connectors` | Boolean | `true` | Settings.java | | `enable_custom_branding` | Boolean | `false` | Settings.java | | `enable_data_science_profile` | Boolean | `false` | Settings.java | | `enable_feature_monitoring` | Boolean | `false` | Settings.java | | `enable_gcs_storage_connectors` | Boolean | `false` | Settings.java | | `enable_glue_storage_connectors` | Boolean | `true` | Settings.java | | `enable_google_sheets_storage_connectors` | Boolean | `true` | Settings.java | | `enable_hopsfsmount_page_cache_in_jobs` | Boolean | `true` | Settings.java | | `enable_hopsfsmount_page_cache_in_jupyter` | Boolean | `false` | Settings.java | | `enable_kafka_storage_connectors` | Boolean | `true` | Settings.java | | `enable_mongodb_storage_connectors` | Boolean | `true` | Settings.java | | `enable_opensearch_storage_connectors` | Boolean | `false` | Settings.java | | `enable_project_observer_role` | Boolean | `false` | Settings.java | | `enable_read_only_git_repositories` | Boolean | `false` | Settings.java | | `enable_redshift_storage_connectors` | Boolean | `true` | Settings.java | | `enable_rest_storage_connectors` | Boolean | `true` | Settings.java | | `enable_sap_hana_storage_connectors` | Boolean | `true` | Settings.java | | `enable_snowflake_storage_connectors` | Boolean | `true` | Settings.java | | `enable_sql_storage_connectors` | Boolean | `true` | Settings.java | | `enable_terminal` | Boolean | `false` | Settings.java | | `enable_unity_catalog_storage_connectors` | Boolean | `true` | Settings.java | | `enable_user_search` | Boolean | `true` | Settings.java | | `executions_cleaner_batch_size` | Integer | `1000` | Settings.java | | `executions_cleaner_interval_ms` | Integer | `600000` | Settings.java | | `executions_per_job_limit` | Integer | `10000` | Settings.java | | `executions_ttl_days` | Integer | `90` | Settings.java | | `feature_monitoring_max_num_features` | Integer | `15` | Settings.java | | `featurestore_asof_spine_max_bytes` | Long | `1073741824` | Settings.java | | `featurestore_asof_spine_max_columns` | Integer | `256` | Settings.java | | `featurestore_asof_spine_max_file_age_ms` | Long | `86400000` | Settings.java | | `featurestore_asof_spine_max_rows` | Long | `1000000` | Settings.java | | `featurestore_db_admin_pass` | String | *(empty string)* | Settings.java | | `featurestore_db_admin_user` | String | *(empty string)* | Settings.java | | `featurestore_default_quota` | Long | `String.valueOf(HdfsConstants.QUOTA_DONT_SET)` *(computed expression, not a literal)* | Settings.java | | `featurestore_default_storage_format` | String | `ORC` | Settings.java | | `featurestore_jdbc_url` | String | `jdbc:mysql://onlinefs.mysql.service.consul:3306/` | Settings.java | | `featurestore_metrics_enabled` | Boolean | `true` | Settings.java | | `featurestore_metrics_max_concurrent_event_processors` | Integer | `5` | Settings.java | | `featurestore_metrics_online_ingestion_enabled` | Boolean | `false` | Settings.java | | `featurestore_online_enabled` | Boolean | `false` | Settings.java | | `featurestore_online_tablespace` | String | *(empty string)* | Settings.java | | `featurestore_schema_migration_timer_enabled` | Boolean | `true` | Settings.java | | `fg_preview_limit` | Integer | `100` | Settings.java | | `file_preview_image_size` | Integer | `10000000` | Settings.java | | `file_preview_txt_size` | Integer | `100` | Settings.java | | `first_time_login` | String | `0` | Settings.java | | `flink_dir` | String | `/srv/hops/flink` | Settings.java | | `flink_version` | String | *(empty string)* | Settings.java | | `fs_java_job_util` | String | `hdfs:///user/spark/hsfs-utils-2.1.0-SNAPSHOT.jar` | Settings.java | | `fs_py_job_util` | String | `hdfs:///user/spark/hsfs_util-2.1.0-SNAPSHOT.py` | Settings.java | | `fs_storage_connector_session_duration` | Integer | `3600` | Settings.java | | `git_bitbucket_http_proxy` | String | *(empty string)* | Settings.java | | `git_bitbucket_https_proxy` | String | *(empty string)* | Settings.java | | `git_command_timeout_minutes` | Integer | `60` | Settings.java | | `git_custom_ca_configmap` | String | *(empty string)* | Settings.java | | `git_custom_ca_configmap_key` | String | `ca-bundle.crt` | Settings.java | | `git_disable_tls_verification` | Boolean | `false` | Settings.java | | `git_github_http_proxy` | String | *(empty string)* | Settings.java | | `git_github_https_proxy` | String | *(empty string)* | Settings.java | | `git_gitlab_http_proxy` | String | *(empty string)* | Settings.java | | `git_gitlab_https_proxy` | String | *(empty string)* | Settings.java | | `git_image` | String | `docker.hops.works/hopsworks/git:0.7.0` | Settings.java | | `git_sync_poll_interval_seconds` | Integer | `60` | Settings.java | | `git_watcher_cpu` | Double | `0.2` | Settings.java | | `git_watcher_memory_mb` | Integer | `128` | Settings.java | | `grafana_version` | String | *(empty string)* | Settings.java | | `hadoop_configmap_name` | String | `hopsfs-config` | Settings.java | | `hadoop_dir` | String | `/srv/hops/hadoop` | Settings.java | | `hadoop_version` | String | `2.8.2` | Settings.java | | `hdfs_default_quota` | Long | `Long.toString(HdfsConstants.QUOTA_DONT_SET)` *(computed expression, not a literal)* | Settings.java | | `hdfs_file_op_job_driver_mem` | Integer | `2048` | Settings.java | | `hdfs_file_op_job_util` | String | `hdfs:///user/spark/hdfs_file_operations-0.2.0.py` | Settings.java | | `hdfs_log_storage_policy` | String | `DistributedFileSystemOps.StoragePolicy.DEFAULT.toString()` *(computed expression, not a literal)* | Settings.java | | `hdfs_user` | String | `hdfs` | Settings.java | | `hdfscontentsmanager_base_hopsfs_client` | String | `libhdfs-go` | Settings.java | | `hive2_version` | String | *(empty string)* | Settings.java | | `hive_conf_path` | String | `/srv/hops/apache-hive/conf/hive-site.xml` | Settings.java | | `hive_superuser` | String | `hive` | Settings.java | | `hive_warehouse` | String | `/apps/hive/warehouse` | Settings.java | | `hops_helm_install_namespace` | String | `hopsworks` *(duplicate key, also declared in Settings.java; values match)* | KubeSettings.java | | `hops_helm_install_namespace` | String | `hopsworks` *(duplicate key, also declared in KubeSettings.java; values match)* | Settings.java | | `hops_rpc_tls` | Boolean | `false` | Settings.java | | `hopsfs_mount_mount_path` | String | `/hopsfs` | Settings.java | | `hopsfsmount_log_level` | String | `warn` | Settings.java | | `hopsfsmount_nn_connections` | Integer | `4` | Settings.java | | `hopsfsmount_virtual_directories` | String | *(empty string)* | Settings.java | | `hopsworks_analytics` | Boolean | `false` | Settings.java | | `hopsworks_analytics_coding_agent` | String | `claude` | Settings.java | | `hopsworks_analytics_project_id` | String | *(empty string)* | Settings.java | | `hopsworks_analytics_ro_pass` | String | `${env:HOPSWORKS_ANALYTICS_RO_PASS}` | Settings.java | | `hopsworks_analytics_ro_user` | String | `hopsworks_ro` | Settings.java | | `hopsworks_analytics_setup_repo` | String | `https://github.com/logicalclocks/okr-dashboards` | Settings.java | | `hopsworks_dir` | String | `/srv/hops/domains` *(duplicate key, also declared in Settings.java; values diverge)* | CAConf.java | | `hopsworks_dir` | String | `/srv/hops/domains/domain1` *(duplicate key, also declared in CAConf.java; values diverge)* | Settings.java | | `hopsworks_engine` | String | `python` | Settings.java | | `hopsworks_enterprise` | Boolean | `false` | Settings.java | | `hopsworks_master_password` | String | `adminpw` *(duplicate key, also declared in Settings.java; values match)* | CAConf.java | | `hopsworks_master_password` | String | `adminpw` *(duplicate key, also declared in CAConf.java; values match)* | Settings.java | | `hopsworks_public_host` | String | *(empty string)* | Settings.java | | `hopsworks_public_proxy_url` | String | *(empty string)* | Settings.java | | `hopsworks_rest_log_level` | String | `PROD` *(duplicate key, also declared in Settings.java; values diverge)* | CAConf.java | | `hopsworks_rest_log_level` | String | `String.valueOf(RESTLogLevel.PROD.name())` *(computed expression, not a literal)* *(duplicate key, also declared in CAConf.java; values diverge)* | Settings.java | | `hopsworks_secret` | String | `hopsworks-secrets` | KubeSettings.java | | `hopsworks_user` | String | `glassfish` | Settings.java | | `hopsworks_version` | String | *(empty string)* | Settings.java | | `hw_group_mapping_sync_enabled` | Boolean | `false` | Settings.java | | `ingestion_job_cores` | Double | `1.0` | Settings.java | | `ingestion_job_gpus` | Integer | `0` | Settings.java | | `ingestion_job_memory` | Integer | `2048` | Settings.java | | `int_service_api_key` | String | *(empty string)* | Settings.java | | `job_name_validation_regex` | String | `^[a-zA-Z0-9_\\-]+$` | Settings.java | | `jupyter_allow_no_limit_shutdown` | Boolean | `true` | Settings.java | | `jupyter_dir` | String | `/srv/hops/jupyter` | Settings.java | | `jupyter_group` | String | `jupyter` | Settings.java | | `jupyter_host` | String | `localhost` | Settings.java | | `jupyter_hour_shutdown_options` | String | `8,24,72` | Settings.java | | `jupyter_origin_scheme` | String | `https` | Settings.java | | `jupyter_remote_fs_driver` | String | `hdfscontentsmanager` | Settings.java | | `jupyter_shell_command` | String | `/bin/bash` | Settings.java | | `jupyter_shutdown_timer_interval` | String | `30m` | Settings.java | | `jupyter_spark_notebook_server_memory_floor_mb` | Integer | `512` | Settings.java | | `jupyter_ws_ping_interval` | String | `10000` | Settings.java | | `jwt_exp_leeway_sec` | Long | `900` | Settings.java | | `jwt_issuer` | String | `hopsworks@logicalclocks.com` *(duplicate key, also declared in Settings.java; values match)* | CAConf.java | | `jwt_issuer` | String | `hopsworks@logicalclocks.com` *(duplicate key, also declared in CAConf.java; values match)* | Settings.java | | `jwt_lifetime_ms` | Long | `1800000` | Settings.java | | `jwt_signature_algorithm` | String | `HS512` | Settings.java | | `jwt_signing_key_name` | String | `apiKey` | Settings.java | | `kafka_enabled` | Boolean | `true` | Settings.java | | `kafka_max_num_topics` | Integer | `10` | Settings.java | | `kafka_num_partitions` | Integer | `2` | Settings.java | | `kafka_num_replicas` | Integer | `1` | Settings.java | | `kafka_version` | String | *(empty string)* | Settings.java | | `keepalive_timeout` | Integer | `30` | Settings.java | | `kerberos_auth` | Boolean | `false` | Settings.java | | `kibana_https_enabled` | Boolean | `false` | Settings.java | | `kibana_multi_tenancy_enabled` | Boolean | `false` | Settings.java | | `kibana_service_log_viewer` | String | *(empty string)* | Settings.java | | `kibana_version` | String | *(empty string)* | Settings.java | | `kube_api_max_attempts` | Integer | `12` | Settings.java | | `kube_ca_certfile` | String | `/srv/hops/certs-dir/certs/ca.cert.pem` | Settings.java | | `kube_ca_password` | String | `adminpw` | CAConf.java | | `kube_client_certfile` | String | `/srv/hops/certs-dir/kube/hopsworks/hopsworks.cert.pem` | Settings.java | | `kube_client_keyfile` | String | `/srv/hops/certs-dir/kube/hopsworks/hopsworks.key.pem` | Settings.java | | `kube_client_keypass` | String | `adminpw` | Settings.java | | `kube_cluster_domain` | String | `.svc.cluster.local` | KubeSettings.java | | `kube_hopsworks_default_service_account` | String | `hopsworks-default` | Settings.java | | `kube_hopsworks_user` | String | `hopsworks` | Settings.java | | `kube_image_builder_service_account` | String | `image-builder` | Settings.java | | `kube_img_pull_policy` | String | `Always` | Settings.java | | `kube_keystore_key` | String | `adminpw` | Settings.java | | `kube_keystore_path` | String | `/srv/hops/certs-dir/kube/hopsworks/hopsworks__kstore.jks` | Settings.java | | `kube_knative_lb_domain` | String | *(empty string)* | Settings.java | | `kube_kserve_installed` | Boolean | `false` | Settings.java | | `kube_kserve_tensorflow_version` | String | *(empty string)* | Settings.java | | `kube_master_url` | String | `https://192.168.68.102:6443` | Settings.java | | `kube_remove_job_when_completed` | Boolean | `true` | Settings.java | | `kube_scheduling_hopsfsmount_cpu_limits` | Double | `-1.0` | Settings.java | | `kube_scheduling_hopsfsmount_cpu_requests` | Double | `1.0` | Settings.java | | `kube_scheduling_hopsfsmount_memory_limits_mb` | Double | `2024.0` | Settings.java | | `kube_scheduling_jobinit_cpu_limits` | Double | `-1.0` | Settings.java | | `kube_scheduling_jobinit_cpu_requests` | Double | `0.5` | Settings.java | | `kube_scheduling_jobinit_memory_limits_mb` | Double | `512.0` | Settings.java | | `kube_scheduling_jobinit_memory_requests_mb` | Double | `256.0` | Settings.java | | `kube_serving_apikey_reaper_grace_minutes` | Integer | `10` | Settings.java | | `kube_serving_apikey_reaper_interval_minutes` | Integer | `15` | Settings.java | | `kube_serving_max_num_instances` | Integer | `-1` | Settings.java | | `kube_serving_min_num_instances` | Integer | `-1` | Settings.java | | `kube_serving_vllm_omni_versions` | String | *(empty string)* | Settings.java | | `kube_serving_vllm_versions` | String | *(empty string)* | Settings.java | | `kube_skip_namespace_creation` | Boolean | `false` | Settings.java | | `kube_truststore_key` | String | `adminpw` | Settings.java | | `kube_truststore_path` | String | `/srv/hops/certs-dir/kube/hopsworks/hopsworks__tstore.jks` | Settings.java | | `kube_type` | String | `local` *(duplicate key, also declared in Settings.java; values match)* | CAConf.java | | `kube_type` | String | `local` *(duplicate key, also declared in CAConf.java; values match)* | Settings.java | | `kube_user_workload_tolerations` | String | *(empty string)* | Settings.java | | `kubernetes_installed` | String | `false` *(duplicate key, also declared in Settings.java; values match)* | CAConf.java | | `kubernetes_installed` | Boolean | `false` *(duplicate key, also declared in CAConf.java; values match)* | Settings.java | | `kueue_installed` | Boolean | `false` | Settings.java | | `kueue_project_default_cluster_queue` | String | *(empty string)* | Settings.java | | `kueue_project_default_local_queue` | String | *(empty string)* | Settings.java | | `kueue_system_jobs_cluster_queue` | String | *(empty string)* | Settings.java | | `kueue_system_jobs_local_queue` | String | *(empty string)* | Settings.java | | `ldap_account_status` | Integer | `1` | Settings.java | | `ldap_attr_binary` | String | `java.naming.ldap.attributes.binary` | Settings.java | | `ldap_auth` | Boolean | `false` | Settings.java | | `ldap_dyn_group_target` | String | `memberOf` | Settings.java | | `ldap_group_dn` | String | *(empty string)* | Settings.java | | `ldap_group_mapping` | String | *(empty string)* | Settings.java | | `ldap_group_mapping_enabled` | Boolean | `true` | Settings.java | | `ldap_group_mapping_sync_enabled` | Boolean | `false` | Settings.java | | `ldap_group_mapping_sync_interval` | String | `0` | Settings.java | | `ldap_group_mapping_sync_limit` | Integer | `0` | Settings.java | | `ldap_group_members_filter` | String | `(&(objectCategory=user)(memberOf=%d))` | Settings.java | | `ldap_group_search_filter` | String | `member=%d` | Settings.java | | `ldap_group_target` | String | `cn` | Settings.java | | `ldap_groups_search_filter` | String | `(&(objectCategory=group)(cn=%c))` | Settings.java | | `ldap_groups_target` | String | `distinguishedName` | Settings.java | | `ldap_krb_user_search_filter` | String | `krbPrincipalName=%s` | Settings.java | | `ldap_user_dn` | String | *(empty string)* | Settings.java | | `ldap_user_email` | String | `mail` | Settings.java | | `ldap_user_givenName` | String | `givenName` | Settings.java | | `ldap_user_id` | String | `uid` | Settings.java | | `ldap_user_search_filter` | String | `uid=%s` | Settings.java | | `ldap_user_surname` | String | `sn` | Settings.java | | `library_install_timeout_minutes` | Integer | `60` | Settings.java | | `lifecycle_webhook_cluster_id` | String | *(empty string)* | Settings.java | | `lifecycle_webhook_secret` | String | *(empty string)* | Settings.java | | `lifecycle_webhook_url` | String | *(empty string)* | Settings.java | | `localhost` | Boolean | `false` | Settings.java | | `log_history_limit` | Integer | `30` | Settings.java | | `login_page_overwrite` | String | *(empty string)* | Settings.java | | `logstash_version` | String | *(empty string)* | Settings.java | | `managed_cloud_provider_name` | String | `hopsworks.ai` | Settings.java | | `managed_cloud_redirect_uri` | String | *(empty string)* | Settings.java | | `managed_docker_registry` | Boolean | `false` | Settings.java | | `management_mode` | String | `STANDALONE` | Settings.java | | `master_encryption_password_value` | String | `encryption_master_password` | KubeSettings.java | | `max_allowed_long_running_http_requests` | Integer | `50` | Settings.java | | `max_concurrent_base_sync_ops` | Integer | `6` | Settings.java | | `max_env_var_name_length` | Integer | `255` | Settings.java | | `max_env_var_value_length` | Integer | `8192` | Settings.java | | `max_env_vars_per_user` | Integer | `64` | Settings.java | | `max_env_yml_byte_size` | Integer | `20000` | Settings.java | | `max_mountable_secret_upload_bytes` | Long | `33554432` | Settings.java | | `max_num_proj_per_user` | Integer | `5` | Settings.java | | `max_ongoing_opensearch_doc_write` | Integer | `100` | Settings.java | | `max_project_cloned_environments` | Integer | `100` | Settings.java | | `max_status_poll_retry` | Integer | `5` | Settings.java | | `max_upload_request_bytes` | Long | `67108864` | Settings.java | | `mount_hopsfs_in_python_deployments` | Boolean | `true` | Settings.java | | `mount_hopsfs_in_python_job` | Boolean | `true` | Settings.java | | `mount_hopsfs_in_ray_job_container` | Boolean | `true` | Settings.java | | `mountable_secret_max_file_bytes` | Long | `1048576` | Settings.java | | `mountable_secret_max_files` | Integer | `32` | Settings.java | | `mountable_secret_max_per_project` | Integer | `10` | Settings.java | | `mountable_secret_max_project_bytes` | Long | `16777216` | Settings.java | | `mountable_secrets_enabled` | Boolean | `false` | Settings.java | | `mountable_secrets_path` | String | `/apps/mountable-secrets` | Settings.java | | `ndb_version` | String | *(empty string)* | Settings.java | | `news_webflow_api_key` | String | *(empty string)* | Settings.java | | `news_webflow_api_url` | String | *(empty string)* | Settings.java | | `notebook_converter_job_timeout_sec` | Long | `300` | Settings.java | | `npm_registry_url` | String | *(empty string)* | Settings.java | | `oauth_account_status` | Integer | `1` | Settings.java | | `oauth_enabled` | Boolean | `false` | Settings.java | | `oauth_group_mapping` | String | *(empty string)* | Settings.java | | `oauth_group_mapping_enabled` | Boolean | `true` | Settings.java | | `oauth_group_mapping_sync_enabled` | Boolean | `false` | Settings.java | | `oauth_logout_redirect_uri` | String | `hopsworks/` | Settings.java | | `oauth_redirect_uri` | String | `hopsworks/callback` | Settings.java | | `ongoing_backup` | Boolean | `false` | Settings.java | | `online_ingestion_max_per_featuregroup` | Integer | `1000` | Settings.java | | `onlinefs_service_thread_number` | Integer | `10` | Settings.java | | `opensearch_default_embedding_index` | String | *(empty string)* | Settings.java | | `opensearch_index_mapping_limit` | Integer | `1000` | Settings.java | | `opensearch_num_default_embedding_index` | Integer | `1` | Settings.java | | `openshift` | Boolean | `false` | Settings.java | | `payara_admin_password_value` | String | `admin_password` | KubeSettings.java | | `payment_type` | String | `NOLIMIT` | Settings.java | | `pki_ca_configuration` | String | *(empty string)* | CAConf.java | | `platform_intelligence_llm_api_key` | String | *(empty string)* | Settings.java | | `platform_intelligence_llm_base_url` | String | *(empty string)* | Settings.java | | `platform_intelligence_llm_model` | String | `gpt-5.4-mini` | Settings.java | | `preinstalled_npm_lib_names` | String | `corepack, npm` | Settings.java | | `preinstalled_python_lib_names` | String | `pydoop, pyspark, jupyterlab, hdfscontents, pyjks, hops-apache-beam, pyopenssl` | Settings.java | | `project_namespace_labels` | String | *(empty string)* | Settings.java | | `project_namespace_network_policy_allowed_namespaces` | String | *(empty string)* | Settings.java | | `project_namespace_network_policy_enabled` | Boolean | `true` | Settings.java | | `project_namespace_network_policy_reconcile_interval` | String | `1m` | Settings.java | | `provenance_graph_max_size` | Integer | `50` | Settings.java | | `public_projects` | String | *(empty string)* | Settings.java | | `pushgateway_cleaner_batch_size` | Integer | `100` | Settings.java | | `pushgateway_group_ttl_minutes` | Integer | `15` | Settings.java | | `pushgateway_monitor_interval_ms` | Integer | `300000` | Settings.java | | `pypi_indexer_timer_enabled` | Boolean | `true` | Settings.java | | `pypi_indexer_timer_interval` | String | `1d` | Settings.java | | `pypi_rest_endpoint` | String | `https://pypi.org/pypi/{package}/json` | Settings.java | | `pypi_simple_endpoint` | String | `https://pypi.org/simple/` | Settings.java | | `python_app_envoy_cpu` | Double | `0.1` | Settings.java | | `python_app_envoy_image` | String | `envoyproxy/envoy:v1.38.0` | Settings.java | | `python_app_envoy_memory_mb` | Integer | `128` | Settings.java | | `python_job_cores` | Double | `1.0` | Settings.java | | `python_job_gpus` | Integer | `0` | Settings.java | | `python_job_kube_waiting_timeout_ms` | Long | `300000` | Settings.java | | `python_job_memory` | Integer | `2048` | Settings.java | | `python_library_updates_monitor_interval` | String | `1d` | Settings.java | | `python_pod_kill_grace_period_seconds` | Long | `600` | Settings.java | | `pythonapp_cores` | Double | `1.0` | Settings.java | | `pythonapp_gpus` | Integer | `0` | Settings.java | | `pythonapp_memory` | Integer | `2048` | Settings.java | | `quotas_featuregroups_online_disabled` | Long | `-1` | Settings.java | | `quotas_featuregroups_online_enabled` | Long | `-1` | Settings.java | | `quotas_max_parallel_executions` | Long | `-1` | Settings.java | | `quotas_max_queued_executions_per_user_per_job` | Long | `10` | Settings.java | | `quotas_model_deployments_running` | Long | `-1` | Settings.java | | `quotas_model_deployments_total` | Long | `-1` | Settings.java | | `quotas_training_datasets` | Long | `-1` | Settings.java | | `ray_certs_dir` | String | `PlatformConstants.RAY_CERTS_DIR` *(computed expression, not a literal)* | Settings.java | | `ray_cluster_max_worker_replicas` | Long | `20` | Settings.java | | `ray_cluster_shutdown_after_completion` | Boolean | `true` | Settings.java | | `ray_cluster_start_wait_time_seconds` | Integer | `120` | Settings.java | | `ray_cluster_termination_grace_period_seconds` | Integer | `10` | Settings.java | | `ray_enabled` | Boolean | `true` | Settings.java | | `ray_job_active_deadline_seconds` | Integer | `120` | Settings.java | | `ray_job_driver_cores` | Double | `1.0` | Settings.java | | `ray_job_driver_gpus` | Integer | `0` | Settings.java | | `ray_job_driver_memory` | Integer | `4096` | Settings.java | | `ray_job_pod_kill_grace_period_seconds` | Integer | `300` | Settings.java | | `ray_job_worker_cores` | Double | `1.0` | Settings.java | | `ray_job_worker_gpus` | Integer | `0` | Settings.java | | `ray_job_worker_memory` | Integer | `4096` | Settings.java | | `ray_materialization_dir` | String | `/srv/hops/ray/job` | Settings.java | | `ray_version` | String | `2.9.0` | Settings.java | | `ray_warehouse_dir` | String | `ray-warehouse` | Settings.java | | `reject_remote_user_no_group` | Boolean | `false` | Settings.java | | `remote_auth_need_consent` | Boolean | `true` | Settings.java | | `remote_shuffle_services_storage_type` | String | `MEMORY_LOCALFILE` | Settings.java | | `replicated_kubernetes_ops_retention` | String | `P15D` | Settings.java | | `requests_verify` | Boolean | `false` | Settings.java | | `reserved_project_names` | String | `hopsworks,information_schema,airflow,glassfish_timers,grafana,hops,metastore,mysql,ndbinfo,performance_schema,sqoop,sys,base,python37,python38,python39,python310,filebeat,git,onlinefs,sklearnserver,rondb_replication,default,kube-system,kube-public,kube-node-lease,kube_system,kube_public,kube_node_lease,trino` | Settings.java | | `resource_dirs` | String | `".sparkStaging;spark-warehouse;.flinkStaging;.flinkCheckpoints;apps;jobs;" + DLT_WAREHOUSE_DIR.getDefaultValue() + ";" + RAY_WAREHOUSE_DIR.getDefaultValue()` *(computed expression, not a literal)* | Settings.java | | `rondb_mgmt_connection_timeout` | Integer | `5000` | Settings.java | | `rondb_mgmt_max_response_lines` | Integer | `10000` | Settings.java | | `rondb_mgmt_read_timeout` | Integer | `10000` | Settings.java | | `rondb_quotas` | String | *(empty string)* | Settings.java | | `rondb_usage_cache_ttl_seconds` | Integer | `60` | Settings.java | | `rondb_usage_query_timeout_seconds` | Integer | `10` | Settings.java | | `saas_entry_point_url` | String | *(empty string)* | Settings.java | | `service_discovery_domain` | String | `consul` *(duplicate key, also declared in Settings.java; values match)* | CAConf.java | | `service_discovery_domain` | String | `consul` *(duplicate key, also declared in CAConf.java; values match)* | Settings.java | | `service_jwt_exp_leeway_sec` | Integer | `43200` | Settings.java | | `service_jwt_lifetime_ms` | Long | `86400000` | Settings.java | | `service_key_rotation_enabled` | String | `false` | CAConf.java | | `service_key_rotation_interval` | String | `3d` | CAConf.java | | `serving_allow_stop_after_seconds` | Integer | `0` | Settings.java | | `serving_connection_pool_size` | Integer | `40` | Settings.java | | `serving_feature_log_materialization_cron` | String | `0 0 0 * * ? *` | Settings.java | | `serving_feature_log_materialization_row_limit` | Integer | `50000000` | Settings.java | | `serving_feature_log_online_ttl_hours` | Integer | `30` | Settings.java | | `serving_feature_logger_batch_bytes` | Integer | `1048576` | Settings.java | | `serving_feature_logger_batch_seconds` | Integer | `5` | Settings.java | | `serving_feature_logger_client_pool_size` | String | `3` | Settings.java | | `serving_feature_logger_client_req_timeout_seconds` | Integer | `3` | Settings.java | | `serving_feature_logger_flush_bytes` | Integer | `1048576` | Settings.java | | `serving_feature_logger_flush_interval_seconds` | Integer | `300` | Settings.java | | `serving_feature_logger_max_buffer_bytes` | Integer | `67108864` | Settings.java | | `serving_feature_logger_max_event_bytes` | Integer | `8388608` | Settings.java | | `serving_feature_logger_max_event_rows` | Integer | `512` | Settings.java | | `serving_feature_logger_queue_size` | Integer | `1000` | Settings.java | | `serving_feature_logger_shutdown_seconds` | Integer | `20` | Settings.java | | `serving_feature_logging_transport` | String | `realtime` | Settings.java | | `serving_max_route_connections` | Integer | `10` | Settings.java | | `serving_monitor_int` | String | `30s` | Settings.java | | `serving_redeploy_not_found_after_seconds` | Integer | `120` | Settings.java | | `serving_state_manager_batch_size` | Integer | `1000` | Settings.java | | `serving_state_manager_enabled` | Boolean | `true` | Settings.java | | `serving_state_manager_interval_ms` | Integer | `600000` | Settings.java | | `spark_configmap_name` | String | `spark` | Settings.java | | `spark_dir` | String | `/srv/hops/spark` | Settings.java | | `spark_executor_min_memory` | Integer | `1024` | Settings.java | | `spark_history_server_enabled` | Boolean | `true` | Settings.java | | `spark_job_driver_cores` | Double | `1.0` | Settings.java | | `spark_job_driver_memory` | Integer | `2048` | Settings.java | | `spark_job_executor_cores` | Double | `1.0` | Settings.java | | `spark_job_executor_memory` | Integer | `4096` | Settings.java | | `spark_kubernetes_materialisation_dir` | String | `/srv/hops/artifacts` | Settings.java | | `spark_launcher_sa_annotations` | String | *(empty string)* | Settings.java | | `spark_pod_kill_grace_period_seconds` | Integer | `1200` | Settings.java | | `spark_remove_job_when_completed` | Boolean | `true` | Settings.java | | `spark_resource_manager` | String | `kubernetes` | Settings.java | | `spark_ui_logs_offset` | Integer | `512000` | Settings.java | | `spark_user` | String | `spark` | Settings.java | | `spark_version` | String | *(empty string)* | Settings.java | | `sql_max_select_in` | Integer | `100` | Settings.java | | `staging_dir` | String | `/srv/hops/domains/domain1/staging` | Settings.java | | `statistics_cleaner_batch_size` | Integer | `1000` | Settings.java | | `statistics_cleaner_interval_ms` | Integer | `900000` | Settings.java | | `streamlit_sharing` | Boolean | `false` | Settings.java | | `sudoers_dir` | String | `/srv/hops/sbin` *(duplicate key, also declared in Settings.java; values match)* | CAConf.java | | `sudoers_dir` | String | `/srv/hops/sbin` *(duplicate key, also declared in CAConf.java; values match)* | Settings.java | | `superset_admin_roles` | String | `Admin` | Settings.java | | `superset_ca_cert_path` | String | `/srv/hops/super_crypto/superset/hops_ca_bundle.pem` | Settings.java | | `superset_enabled` | Boolean | `false` | Settings.java | | `superset_proxy_connect_timeout_ms` | Integer | `10000` | Settings.java | | `superset_proxy_connection_request_timeout_ms` | Integer | `10000` | Settings.java | | `superset_proxy_max_connections` | Integer | `50` | Settings.java | | `superset_proxy_read_timeout_ms` | Integer | `180000` | Settings.java | | `superset_secret` | String | `superset-admin-credentials` | KubeSettings.java | | `superset_user_roles` | String | `Gamma,sql_lab` | Settings.java | | `tag_history_archive_max_events` | Integer | `20000` | Settings.java | | `tag_history_cleaner_batch_size` | Integer | `1000` | Settings.java | | `tag_history_cleaner_interval_ms` | Integer | `86400000` | Settings.java | | `tag_history_retention_days` | Integer | `0` | Settings.java | | `teleport_cleaner_interval_ms` | Long | `86400000` | Settings.java | | `teleport_ttl_days` | Integer | `7` | Settings.java | | `tensorflow_version` | String | *(empty string)* | Settings.java | | `terminal_gpu_image` | String | `terminal-gpu` | Settings.java | | `terminal_image` | String | `terminal-server` | Settings.java | | `terminal_oom_guard_enabled` | Boolean | `true` | Settings.java | | `terminal_proxy_pod_app_labels` | String | `hopsworks-instance,hopsworks-admin` | Settings.java | | `terminal_proxy_token_ttl_ms` | Long | `60000` | Settings.java | | `terminal_session_hours` | Integer | `4` | Settings.java | | `terminal_session_max_hours` | Integer | `24` | Settings.java | | `terminal_shm_size` | String | `1Gi` | Settings.java | | `terminal_spark_image` | String | `terminal-spark` | Settings.java | | `testconnector_image` | String | `docker.hops.works/hopsworks/testconnector:0.2` | Settings.java | | `testconnector_launcher` | String | `testconnector-launch.sh` | Settings.java | | `trino_catalog_approval_required` | Boolean | `false` | Settings.java | | `trino_catalog_max_bytes` | Integer | `16384` | Settings.java | | `trino_catalog_max_per_project` | Integer | `10` | Settings.java | | `trino_catalog_reconcile_enabled` | Boolean | `false` | Settings.java | | `trino_catalogs_configmap` | String | `hopsworks-trino-catalogs` | KubeSettings.java | | `trino_connectors` | String | `ai,bigquery,blackhole,cassandra,clickhouse,datasketches,delta_lake,druid,duckdb,elasticsearch,exasol,faker,gsheets,hive,hudi,iceberg,ignite,jmx,kafka,lakehouse,loki,mariadb,memory,mongodb,mysql,opensearch,oracle,pinot,postgresql,prometheus,redis,redshift,singlestore,snowflake,sqlserver,tpcds,tpch,trino_thrift` | Settings.java | | `trino_coordinator_deployment` | String | `hopsworks-trino-coordinator` | KubeSettings.java | | `trino_credentials_secret` | String | `trino-admin-credentials` | KubeSettings.java | | `trino_default_catalog` | String | `hive` | Settings.java | | `trino_eager_restart` | Boolean | `false` | Settings.java | | `trino_eager_restart_poll_minutes` | Integer | `10` | Settings.java | | `trino_egress_probe_container` | String | `egress-probe` | KubeSettings.java | | `trino_enabled` | Boolean | `false` | Settings.java | | `trino_events_cleaner_batch_size` | Integer | `1000` | Settings.java | | `trino_events_delete_after_days` | Integer | `61` | Settings.java | | `trino_group_secret` | String | `trino-groups-file` | KubeSettings.java | | `trino_max_catalogs` | Integer | `250` | Settings.java | | `trino_mountable_secrets_root` | String | `/opt/hopsworks/mounts` | KubeSettings.java | | `trino_password_secret` | String | `trino-password-file` | KubeSettings.java | | `trino_reconcile_enabled` | Boolean | `true` | Settings.java | | `trino_reconcile_interval_ms` | Long | `300000` | Settings.java | | `trino_scheduled_restart_enabled` | Boolean | `true` | Settings.java | | `trino_scheduled_restart_idle_retry_minutes` | Integer | `5` | Settings.java | | `trino_scheduled_restart_idle_wait_minutes` | Integer | `60` | Settings.java | | `trino_scheduled_restart_interval_hours` | Integer | `24` | Settings.java | | `trino_scheduled_restart_time` | String | `02:00` | Settings.java | | `trino_test_coordinator_deployment` | String | `hopsworks-trino-test-coordinator` | KubeSettings.java | | `trino_test_coordinator_enabled` | Boolean | `false` | Settings.java | | `trino_user_catalogs_max_shards` | String | `2` | KubeSettings.java | | `trino_user_catalogs_secret_prefix` | String | `hopsworks-trino-catalogs-user-` | KubeSettings.java | | `trino_worker_deployment` | String | `hopsworks-trino-worker` | KubeSettings.java | | `twofactor-excluded-groups` | String | `AGENT;CLUSTER_AGENT` | Settings.java | | `twofactor_auth` | String | `false` | Settings.java | | `unity_catalog_oauth_m2m_enabled` | Boolean | `true` | Settings.java | | `upload_chunk_size` | Integer | `10485760` | Settings.java | | `upload_policy` | String | `enabled` | Settings.java | | `validate_remote_user_email_verified` | Boolean | `false` | Settings.java | | `velero_backup_main_schedule_name` | String | *(empty string)* | Settings.java | | `velero_backup_storage_location_name` | String | *(empty string)* | Settings.java | | `velero_backup_users_schedule_name` | String | *(empty string)* | Settings.java | | `velero_namespace` | String | *(empty string)* | Settings.java | | `whitelist_users` | String | `agent@hops.io` | Settings.java | | `yarn_app_uid` | Long | `1235` | Settings.java | | `yarn_default_quota` | Integer | `60000` | Settings.java | | `zookeeper_version` | String | *(empty string)* | Settings.java | ================================================================================ # Python Build Performance Source: https://docs.hopsworks.ai/latest/setup_installation/admin/build_performance/ # Python Environment Build Performance ## Introduction Every change to a project's Python environment produces a new container image. By default each build starts a private BuildKit daemon inside its own Kubernetes Job, on an `emptyDir`. That daemon re-pulls and re-unpacks the base image every time, and nothing it downloads survives the build. This guide covers the settings that change that: a long-lived BuildKit daemon, the package cache shared between builds, the toolchain caches a custom-command build can request, and layer reuse. All of these are off or conservative by default. Turn them on deliberately. ## Prerequisites An administrator account on a Hopsworks cluster, and Helm access to the deployment for the daemon itself. ## Persistent BuildKit daemon The daemon is deployed by the Helm chart and is disabled by default: ```yaml hopsworks: buildkitd: enabled: true ``` Enabling it deploys a StatefulSet with its own state volume. Once it is running, the chart also points the backend at it: `docker_operations_buildkit_addr` and the client TLS settings are filled in from `buildkitd.name`, `buildkitd.port`, `buildkitd.replicas` and `buildkitd.tls`. You do not set those variables yourself unless you are pointing builds at a daemon the chart does not manage. The daemon runs as root in a privileged container with `hostPID`. It is a shared service holding every project's build state and cache, so give it a dedicated node and taint that node so nothing else schedules there. A rootless mode exists (`buildkitd.rootless.enabled`) for installs that cannot grant a privileged container. It costs the process sandbox for `RUN` steps, so it is for a single trust zone only: do not use it where concurrent projects are mutually untrusted. It also needs a kernel that supports its snapshotter unprivileged. Rootless overlayfs needs **5.11 or later**; on older kernels such as RHEL 8's 4.18, set the snapshotter to `native` and expect slower builds and a larger state volume: ```yaml hopsworks: variables: docker_operations_oci_worker_snapshotter: "native" ``` fuse-overlayfs is the faster fallback on those kernels, and the daemon has to actually hold `/dev/fuse` to use it. Making the device node visible is not enough: access is governed by the container's device cgroup, so an unprivileged container with a `hostPath` still gets `Operation not permitted`. Kubernetes has no portable equivalent of `docker --device`, so name the mechanism your runtime provides: ```yaml hopsworks: buildkitd: rootless: enabled: true # crio emits the io.kubernetes.cri-o.Devices annotation CRI-O honours # devicePlugin requests devicePluginResource, for a device plugin or CDI provider # none no device; the snapshotter must then be native deviceInjection: "crio" variables: docker_operations_oci_worker_snapshotter: "fuse-overlayfs" ``` CRI-O ignores the device annotation unless the node allows it. Every node that can run the daemon needs a drop-in such as `/etc/crio/crio.conf.d/99-buildkitd-fuse.conf`, followed by `systemctl restart crio`: ```toml [crio.runtime] allowed_devices = ["/dev/fuse"] [crio.runtime.runtimes.runc] allowed_annotations = ["io.kubernetes.cri-o.Devices"] ``` To allow the annotation only for this daemon rather than for every pod on the node, use a CRI-O workload with an `activation_annotation` and set the matching annotation through `buildkitd.podAnnotations`; `values.rhel8.yaml` in the chart shows the pairing. Name the snapshotter explicitly rather than leaving it on `auto`. If FUSE turns out to be unavailable the daemon then fails where you can see it, instead of quietly selecting `native` and making every build slower and larger with no signal. Asking for `fuse-overlayfs` with `deviceInjection: none` is refused when the chart renders, as is `devicePlugin` with an empty `devicePluginResource`. A preflight init container (`buildkitd.rootless.preflight`, on by default) checks the node before the daemon starts: unprivileged user namespaces always, and `fuse` in `/proc/filesystems` when the configuration uses FUSE. The check that matters most, actually opening `/dev/fuse`, runs in the daemon container itself just before the daemon starts, because a device plugin assigns the device to the one container that requested it and an init container would never receive it. A daemon that cannot open the device exits immediately with the reason. Before trusting it, confirm on the node that `buildctl debug workers -v` reports `fuse-overlayfs` rather than `native`, that a second identical build reuses layers from the first, and that the cache survives a restart of the StatefulSet. Under enforcing SELinux the device may additionally need a matching SCC or targeted policy; do not disable SELinux labelling wholesale to get past it. On RHEL 8 the rootful daemon above remains the recommended configuration. ### Sizing the state volume ```yaml hopsworks: buildkitd: storage: 100Gi gc: totalKeepBytes: "80GB" cacheMountKeepBytes: "30GB" cacheMountKeepDuration: "168h" ``` `totalKeepBytes` must stay well above the total unpacked size of every base image in use. Below that, each build evicts the base image the next one needs and the daemon is slower than no daemon at all. Package caches get their own budget so that downloads cannot evict base image snapshots. !!! warning `gc.keySyntax` selects which GC key names to emit. BuildKit replaced `keepBytes` with `maxUsedSpace` in later releases, and an unknown key makes the daemon refuse to start. Set it to match the BuildKit version you deploy. ### Client certificates ```yaml hopsworks: buildkitd: tls: enabled: true ``` On by default, and only worth turning off on a single-trust-zone installation. This is not defence in depth. `buildctl` over TCP is unauthenticated, and BuildKit gives `RUN` steps host networking, so a custom-command script reaches the daemon on both the service address and `127.0.0.1` whatever pod network rules are in place. Client certificates are what stops it. The client key is mounted in the build Job's pod rather than in the `RUN` step's filesystem, so a build step cannot present it. Note that this authenticates but does not authorize: any certificate the Hopsworks CA issued is accepted. Separating tenants at the daemon level still requires a daemon per trust zone. ### Replicas ```yaml hopsworks: buildkitd: replicas: 3 ``` Each replica gets its own state volume, and a project is pinned to one of them by id so it keeps hitting the daemon that already unpacked its base image. This adds throughput; it is not what separates tenants, and a project is not failed over to another replica. ## Package cache ```text docker_operations_buildkit_cache_scope ``` | Value | Behaviour | | --- | --- | | `shared` | Default. Builds resolving through the same package-index configuration share one cache. | | `project` | Each project gets its own cache. No sharing at all. | | `off` | No cache mount. | A cache only survives between builds if the daemon does, so this pairs with `buildkitd.enabled`. `shared` is scoped by index configuration rather than by a single cluster-wide identifier. Two projects resolving through the same configuration share, which is the point; a project configured against a different private index gets a different cache, so it cannot receive an artifact fetched under someone else's credentials. Where the index configuration is cluster-wide, which is the normal setup, every project shares. Choose `project` if you want no sharing between projects under any circumstances. ## Layer reuse Two separate things decide whether a build step's *result* can be reused, as opposed to its downloads. ### Dependency locking ```text docker_operations_lock_dependencies ``` Off by default. When on, a `pip`, wheel or requirements install first resolves the complete dependency set with a hash per artifact, then installs only from that lock. This makes the resolved set reconstructible and fails the build if an index serves different bytes for a version it already served. A pinned version alone is not enough for a step's result to be reusable: `pandas==1.0.0` says nothing about the versions its transitive and build dependencies resolve to. Only a lock pins that, which is why locking is what allows the install layer to be reused. !!! warning Locking requires `uv` in the base image. A build on an image without it fails with a message naming the base image, rather than silently installing without a lock. Git installs are never locked: a git reference has no artifact hash to generate. ### Custom commands A custom-command layer is never reused by default, because the script can fetch anything and nothing declares what. See [Custom Commands](../../user_guides/projects/python/custom_commands.md) for the two directives a build can use, and: ```text docker_operations_allow_hermetic_custom_commands ``` Off by default. This decides whether a build's own assertion that its script fetches nothing mutable is allowed to control cache reuse. Only the script's author knows whether it is true; only you decide whether that claim counts. ## Cache keys and credentials BuildKit deliberately leaves the contents of build secrets out of a step's cache key. That keeps credentials out of the image, but on its own it means a rotated credential still matches the layer built with the old one, and two projects whose builds differ only by their index credentials produce the same key. Hopsworks adds an opaque tag to every reusable step covering the project it belongs to and the credential material it mounts, so a layer is never reused across projects and rotating a credential invalidates reuse. The tag is derived under the installation's own key and is not reversible into the credentials. If that key cannot be read, steps that consume credentials are marked uncacheable rather than reused. Builds get slower; they do not cross a trust boundary. ## Registry trust When the daemon pulls from a registry served with a private CA, it needs that CA. A per-build daemon inherits it from the build Job's pod; the shared daemon is a separate pod and does not. The chart configures trust for the cluster's own registry automatically. For any additional registry served over plain HTTP or with a certificate the daemon cannot verify: ```yaml hopsworks: buildkitd: insecureRegistries: - my-registry.example.com:5000 ``` ## Supply chain ### What a build fetches The build path adds no new outbound network dependency. `uv` is copied into the base image at base-image build time from a digest-pinned source, so no build downloads a toolchain. Python packages come from the indexes configured in `dockerImage.pypi.global_parameters` and nothing else. The only image the cluster needs beyond the ones it already pulled is the BuildKit daemon. That matters for air-gapped installations: mirror the BuildKit image and point the chart at the mirror, and the rest of the build path needs no egress beyond your own package index. ```yaml hopsworks: dockerRegistry: buildkit: image: my-mirror.example.com/moby/buildkit tag: v0.31.2 # keep in step with buildkitd.gc.keySyntax, below ``` ### BuildKit version The chart ships **v0.31.2**. Do not pin below it. Earlier releases carry published advisories, two of which matter more once the daemon is shared rather than started per build: - A path traversal in the Git URL subdirectory component. Reachable from an ordinary user action, because a library can be installed from a user-supplied Git URL. - A state-directory escape via a custom frontend. A shared daemon holds state for every project, so the blast radius is no longer one build. Both are fixed in v0.28.1; v0.31.2 also covers a Seccomp/AppArmor bypass, an unbounded-parsing denial of service, and a command injection through Git bundle checkout. A per-build daemon is affected by the same advisories, so this is not a reason to leave `buildkitd.enabled` off. Upgrading the image is the fix in both modes. If you repin to a different version, set `buildkitd.gc.keySyntax` to match it: ```yaml hopsworks: buildkitd: gc: keySyntax: maxUsedSpace # keepBytes for BuildKit up to ~v0.16 ``` Getting this wrong is silent rather than loud. A modern daemon still accepts `keepBytes`, but treats it as `reservedSpace`: the same number stops meaning "never exceed" and starts meaning "always keep", so the cache budget becomes a floor and the state volume fills. ### Attestations and signing Builds do not produce SBOM or provenance attestations, and images are not signed. BuildKit can generate both (`attest:sbom`, `attest:provenance`), but the SBOM scanner is itself an image that would have to be mirrored, so neither is enabled by default. If you enforce image policy, do it on the registry side against the pushed image rather than in the build: environment images are built continuously and per project, and each is recorded in the environment history with its output digest, which is the identifier to sign or admit against. ## Measuring The backend logs where the time in a build actually goes: ```text Build docker-build-xxxxxx: context packaged in N ms, uploaded in N ms Built in N ms (1 region job(s)) Registered : metadata read N ms, database registration N ms (metadata read from the image) ``` A slow environment build is either compilation, image transfer, or orchestration, and those have different fixes. Read the split rather than the total before changing any of the settings above. ================================================================================ # User Management Source: https://docs.hopsworks.ai/latest/setup_installation/admin/user/ # User Management ## Introduction Whether you run Hopsworks on-premise, or on the cloud using kubernetes, you have a Hopsworks cluster which contains all users and projects. ## Prerequisites Administrator account on a Hopsworks cluster. ### Step 1: Go to user management All the users of your Hopsworks instance have access to your cluster with different access rights. You can find them by clicking on your name in the top right corner of the navigation bar, choosing _Cluster Settings_ from the dropdown menu, then choosing _Users_ under _Security & Access_ in the left sidebar (You need to have _Admin_ role to get access to the _Cluster Settings_ page).
active users
Active Users
### Step 2: Manage user roles Roles let you manage the access rights of a user to the cluster. - User: users with this role are only allowed to use the cluster by creating a limited number of projects. - Admin: users with this role are allowed to manage the cluster. This includes accepting new users to the cluster or blocking them, managing user quota, [configure alerts](./alert.md) and setting up [authentication methods](./auth.md). You can change the role of a user by clicking on the _select dropdown_ that shows the current role of the user. ### Step 3: Validating and blocking users By default, a user who register on Hopsworks using their own credentials are not granted access to the cluster. First, a user with an admin role needs to validate their account. Users waiting for validation are listed at the top of the _Users_ page, above the search box, as shown in the image below.
request
Pending user requests
Each pending row carries an _Activate_ and a _Block_ button. A marker next to the email address shows that the address has not been validated yet. Similarly, if a user is no longer allowed access to the cluster you can block them. To keep consistency with the history of your datasets, a user can not be deleted but only blocked. If necessary a user can be deleted manually in the cluster using the command line. You can block a user by clicking on the block icon on the right side of the user in the list.
blocked users
Blocked Users
Blocked users will appear on the lower section of the page. Click on _Show blocked users_ to show all the blocked users in your cluster. If a user is blocked by mistake you can reactivate it by clicking on the check mark icon that corresponds to that user in the blocked users list. If there are too many users in your cluster, use the search box (available for blocked users too) to filter users by name or email. It is also possible to filter activated users by role. For example to see all administrators in you cluster click on the _select dropdown_ to the right of the search box and choose _Admin_. ### Step 4: Create a new users If you want to allow users to login without registering you can pre-create them by clicking on _New user_.
New user
Create new user
After setting the user's name and email chose the type of user you want to create (Hopsworks, Kerberos, LDAP or OAuth2). To create a Kerberos or LDAP user you need to get the users **UUID** from the Kerberos or LDAP server. _Max number of projects_ sets how many projects that user is allowed to create. Hopsworks user can also be assigned a _Cluster role_. Kerberos, LDAP and OAuth2 users on the other hand can only be assigned a role through group mapping. A temporary password will be generated and displayed when you click on _Create new user_. Copy the password and pass it securely to the user.
create user
Copy temporary password
### Step 5: Reset user password In the case where a user loses her/his password and can not recover it with the [password recovery](../../user_guides/projects/auth/recovery.md), an administrator can reset it for them. On the bottom of the _Users_ page, under _Advanced features_, click on _Reset a user password_. A popup window with a dropdown for searching users by name or email will open. Find the user and click on _Reset new password_.
reset password
Reset user password
A temporary password will be displayed. Copy the password and pass it to the user securely.
temp password
Copy temporary password
A user with a temporary password will see a warning message when going to _Account settings_ **Authentication** tab.
change password
Change password
!!! Note A temporary password should be changed as soon as possible. ## Python SDK !!! warning "Admin-only capability" The calling account must hold the `HOPS_ADMIN` platform role, since these calls manage users across the entire cluster. Admin accounts can also manage platform users programmatically: ```python import hopsworks hopsworks.login() users_api = hopsworks.get_users_api() # Register a new user (a temporary password is generated if not provided) new_user = users_api.register_user( email="alice@example.com", first_name="Alice", last_name="Smith", role="HOPS_USER", ) if new_user.password: print("temporary password:", new_user.password) # List / get users for user in users_api.get_users(): print(user.email, user.roles) user = users_api.get_user(new_user.id) # Activate / reject a registration request users_api.activate_user(new_user.id) users_api.reject_user(new_user.id) # Change platform role or project quota users_api.set_role(new_user.id, "HOPS_ADMIN") users_api.update_user(new_user.id, max_num_projects=10) # Delete a user (fails if they still own any projects) users_api.delete_user(new_user.id) ``` ================================================================================ # Project Management Source: https://docs.hopsworks.ai/latest/setup_installation/admin/project/ # Manage Projects Hopsworks provides an administrator with a view of the projects in a Hopsworks cluster. A Hopsworks administrator is not automatically a member of all the projects in a cluster. However, they can see which projects exist, who is the project owner, and they can limit the storage quota and compute quota for each project. ## Prerequisites You need to be an administrator on a Hopsworks cluster. ## Changing project quotas You can find the Project management page by clicking on your name, in the top right corner of the navigation bar, choosing _Cluster Settings_ from the dropdown menu, then choosing _Projects_ under _Projects & Data_ in the left sidebar.
Project page
Project page
This page will list all the projects in a cluster, their name, owner and when its quota was last updated. By clicking on the _edit configuration_ link of a project you will be able to edit the quotas of that project.
Project quotas
Project quotas
### Storage Storage quota represents the amount of data a project can store. The storage quota is broken down in two different areas: - **Feature Store**: This represents the storage quota for files and directories stored in the `_featurestore.db` dataset in the project. This dataset contains all the feature group offline data for the project. - **Project**: This represents the storage quota for all the data stored on any other dataset. Each storage quota is divided into space quota, i.e., how much space the files can consume, and namespace quota, i.e., how many files and directories there can be. If Hopsworks is deployed on-premise using hard drives to store the data, i.e., Hopsworks is not configured to store its data in a S3-compliant storage system, the data is replicated across multiple nodes (by default 3) and the space quota takes the replication factor into consideration. As an example, a 100MB file stored with a replication factor of 3, will consume 300MB of space quota. By default, all storage quotas are disabled and not enforced. Administrators can change this default by changing the following configuration in the [Configuration](../admin/variables.md) UI and/or the cluster definition: ```yaml hopsworks: featurestore_default_quota: [default quota in bytes, -1 to disable] hdfs_default_quota: [default quota in bytes, -1 to disable] ``` The values specified will be set during project creation and administrators will be able to customize each project using this UI. ### Compute Compute quotas represents the amount of compute a project can use to run Spark applications as well as Tez queries. Quota is expressed as number of seconds a container of size 1 CPU and 1GB of RAM can run for. If the Hopsworks cluster is connected to a Kubernetes cluster, Python jobs, Jupyter notebooks and KServe models are not subject to the compute quota. Currently, Hopsworks does not support defining quotas for compute scheduled on the connected Kubernetes cluster. By default, the compute quota is disabled. Administrators can change this default by changing the following configuration in the [Configuration](../admin/variables.md) UI and/or the cluster definition: ```yaml hopsworks: yarn_default_payment_type: [NOLIMIT to disable the quota, PREPAID to enable it] yarn_default_quota: [default quota in seconds] ``` The values specified will be set during project creation and administrators will be able to customize each project using this UI. ### Kafka Topics Kafka is used within Hopsworks to enable users to write data to the feature store in Real-Time and from a variety of different frameworks. If a user creates a feature group with the stream APIs enabled, then a Kafka topic will be created for that feature group. By default, a project can have up to 100 Kafka topics. Administrators can increase the number of Kafka topics a project is allowed to create by increasing the quota in the project admin UI. ## Force deleting a project Administrators have the option to force delete a project. This is useful if the project was not created or deleted properly, e.g., because of an error. ## Controlling who can create projects Every user on Hopsworks can create projects. By default, each user can create up to 10 projects. For production environments, the number of projects should be limited and controlled for resource allocation purposes as well as closer control over the data. Administrators can control how many projects a user can provision by setting the following configuration in the [Configuration](../admin/variables.md) UI and/or cluster definition: ```yaml hopsworks: max_num_proj_per_user: [Maximum number of projects each user can create] ``` This value will be set when the user is provisioned. Administrators can grant additional projects to a specific user through the [User Administration](../admin/user.md) UI. ================================================================================ # Configure Alerts Source: https://docs.hopsworks.ai/latest/setup_installation/admin/alert/ # Configure Alerts ## Introduction Alerts are sent from Hopsworks using Prometheus' [Alert manager](https://prometheus.io/docs/alerting/latest/alertmanager/). In order to send alerts we first need to configure the _Alert manager_. ## Prerequisites Administrator account on a Hopsworks cluster. ### Step 1: Go to alerts configuration To configure the _Alert manager_ click on your name in the top right corner of the navigation bar and choose Cluster Settings from the dropdown menu, then choose **Alerts** under _Observability_ in the left sidebar (`/settings/alerts`). On the **Alert global channels** page you can configure the alert manager to send alerts via email, slack, pagerduty or webhook.
Configure alerts
Configure alerts
### Step 2: Configure Email Alerts To send alerts via email you need to configure an SMTP server. Click on the _Configure_ button on the right side of the **email** row and fill out the form that pops up.
Configure Email Alerts
Configure Email Alerts
- _Default from_: the address used as sender in the alert email. - _SMTP smarthost_: the Simple Mail Transfer Protocol (SMTP) host through which emails are sent. - _Default hostname (optional)_: hostname to identify to the SMTP server. - _Authentication method_: how to authenticate to the SMTP server. CRAM-MD5, LOGIN or PLAIN. Optionally cluster wide Email alert receivers can be added in _Default receiver emails_. These receivers will be available to all users when they create event triggered [alerts](../../user_guides/fs/feature_group/data_validation_best_practices.md#setup-alerts). ### Step 3: Configure Slack Alerts Alerts can also be sent via Slack messages. To be able to send Slack messages you first need to configure a Slack webhook. Click on the _Configure_ button on the right side of the **slack** row and past in your [Slack webhook](https://api.slack.com/messaging/webhooks) in _Webhook_.
Configure slack Alerts
Configure slack Alerts
Optionally cluster wide Slack alert receivers can be added in _Slack channel/user_. These receivers will be available to all users when they create event triggered [alerts](../../user_guides/fs/feature_group/data_validation_best_practices.md#setup-alerts). ### Step 4: Configure Pagerduty Alerts Pagerduty is another way you can send alerts from Hopsworks. Click on the _Configure_ button on the right side of the **pager duty** row and fill out the form that pops up.
Configure Pagerduty Alerts
Configure Pagerduty Alerts
Fill in Pagerduty URL: the URL to send API requests to. Optionally cluster wide Pagerduty alert receivers can be added in _Service key/Routing key_. By first choosing the PagerDuty integration type: - _global event routing (routing_key)_: when using PagerDuty integration type `Events API v2`. - _service (service_key)_: when using PagerDuty integration type `Prometheus`. Then adding the Service key/Routing key of the receiver(s). PagerDuty provides [documentation](https://www.pagerduty.com/docs/guides/prometheus-integration-guide/) on how to integrate with Prometheus' Alert manager. ### Step 5: Configure Webhook Alerts You can also use webhooks to send alerts. A Webhook Alert is sent as an HTTP POST command with a JSON-encoded parameter payload. Click on the _Configure_ button on the right side of the **webhook** row and fill out the form that pops up.
Configure Webhook Alerts
Configure Webhook Alerts
Fill in the unique URL of your Webhook: the endpoint to send HTTP POST requests to. A global receiver is created when a webhook is configured and can be used by any project in the cluster. ### Step 6: Advanced configuration If you are familiar with Prometheus' [Alert manager](https://prometheus.io/docs/alerting/latest/alertmanager/) you can also configure alerts by editing the _yaml/json_ file directly. Click on _Advanced configuration_ at the top right of the Alerts page, then on the _Edit_ button. The advanced page shows the configuration currently loaded on the alert manager. After editing the configuration it takes some time to propagate changes to the alertmanager. To get back to the channel list, choose **Alerts** again in the left sidebar. The _Reload_ button at the bottom of the page can be used to validate the changes made to the configuration. It will try to load the new configuration to the alertmanager and show any errors that might prevent the configuration from being loaded.
Advanced configuration
Advanced configuration
A status indicator next to the _Advanced configuration_ heading shows whether the configuration Hopsworks has saved matches what the alert manager currently has loaded. - _Synced_: the loaded configuration matches the saved configuration. - _Pending reload_: a change has been saved but the alert manager has not loaded it yet. This is expected right after an edit because changes take some time to propagate. - _Warning_: the change has stayed unloaded past the timeout, which usually means the alert manager rejected the configuration. The reason reported by the alert manager is shown so you can find and fix the offending section.
Advanced configuration load status
Configuration load status next to the Advanced configuration heading
!!!warning If you make any changes to the configuration ensure that the changes are valid by reloading the configuration until the changes are loaded and visible in the advanced page. _Example:_ Adding the yaml snippet shown below in the global section of the alert manager configuration will have the same effect as creating the SMTP configuration as shown in [section 1](#step-2-configure-email-alerts) above. ```yaml global: smtp_smarthost: smtp.gmail.com:587 smtp_from: hopsworks@gmail.com smtp_auth_username: hopsworks@gmail.com smtp_auth_password: XXXXXXXXX smtp_auth_identity: hopsworks@gmail.com ... ``` To manage a project's receivers and create alerts on Jobs and Feature group validations see [Alerts](../../user_guides/projects/alerts/index.md). The yaml syntax in the UI is slightly different in that it does not allow double quotes (it will ignore the values but give no error). Below is an example configuration, that can be used in the UI, with both email and slack receivers configured for system alerts. ```yaml global: smtp_smarthost: smtp.gmail.com:587 smtp_from: hopsworks@gmail.com smtp_auth_username: hopsworks@gmail.com smtp_auth_password: XXXXXXXXX smtp_auth_identity: hopsworks@gmail.com resolveTimeout: 5m templates: - /srv/hops/alertmanager/alertmanager-0.17.0.linux-amd64/template/*.tmpl route: receiver: default routes: - receiver: email continue: true match: type: system-alert - receiver: slack continue: true match: type: system-alert groupBy: - alertname groupWait: 10s groupInterval: 10s receivers: - name: default - name: email emailConfigs: - to: someone@logicalclocks.com from: hopsworks@logicalclocks.com smarthost: mail.hello.com text: >- summary: {{ .CommonAnnotations.summary }} description: {{ .CommonAnnotations.description }} - name: slack slackConfigs: - apiUrl: >- https://hooks.slack.com/services/T00000000/B00000000/XXXXXXXXXXXXXXXXXXXXXXXX channel: '#general' text: >- summary: {{ .Annotations.summary }} description: {{ .Annotations.description }} ``` ================================================================================ # IAM Role Chaining Source: https://docs.hopsworks.ai/latest/setup_installation/admin/roleChaining/ # AWS IAM Role Chaining ## Introduction When running Hopsworks in Amazon EKS you have several options to give the Hopsworks user access to AWS resources. The simplest is to assign [Amazon EKS node IAM role](https://docs.aws.amazon.com/eks/latest/userguide/create-node-role.html) access to the resources. But, this will make these resources accessible by all users. To manage access to resources on a project base you need to use [Role chaining](https://docs.aws.amazon.com/IAM/latest/UserGuide/id_roles_terms-and-concepts.html#iam-term-role-chaining). In this document we will see how to configure AWS and Hopsworks to use Role chaining in your Hopsworks projects. ## Prerequisites Before you begin this guide you'll need the following: - A Hopsworks cluster running on EKS. - Enabled IAM [OpenID Connect (OIDC) provider](https://docs.aws.amazon.com/eks/latest/userguide/enable-iam-roles-for-service-accounts.html) for your cluster. - Administrator account on the Hopsworks cluster. ### Step 1: Create an IAM role and associate it with a Kubernetes service account To use role chaining the hopsworks instance pods need to be able to impersonate the roles you want to be linked to your project. For this you need to create an IAM role and associate it with your Kubernetes service accounts with assume role permissions and attach it to your hopsworks instance pods. For more details on how to create an IAM roles for Kubernetes service accounts see the [aws documentation](https://docs.aws.amazon.com/eks/latest/userguide/associate-service-account-role.html). !!!note To ensure that users can't use the service account role and impersonate the roles by their own means, you need to ensure that the service account is only attached to the hopsworks instance pods. ```sh account_id=$(aws sts get-caller-identity --query "Account" --output text) oidc_provider=$(aws eks describe-cluster --name my-cluster --region $AWS_REGION --query "cluster.identity.oidc.issuer" --output text | sed -e "s/^https:\/\///") ``` ```sh export namespace=hopsworks export service_account=my-service-account ``` ```json { "Version": "2012-10-17", "Statement": [ { "Effect": "Allow", "Principal": { "Federated": "arn:aws:iam::$account_id:oidc-provider/$oidc_provider" }, "Action": "sts:AssumeRoleWithWebIdentity", "Condition": { "StringEquals": { "$oidc_provider:aud": "sts.amazonaws.com", "$oidc_provider:sub": "system:serviceaccount:$namespace:$service_account" } } } ] } ```
Example trust policy for a service account.
```json { "Version": "2012-10-17", "Statement": [ { "Sid": "AssumeDataRoles", "Effect": "Allow", "Action": "sts:AssumeRole", "Resource": [ "arn:aws:iam::123456789011:role/my-role", "arn:aws:iam::xxxxxxxxxxxx:role/s3-role", "arn:aws:iam::xxxxxxxxxxxx:role/dev-s3-role", "arn:aws:iam::xxxxxxxxxxxx:role/redshift" ] } ] } ```
Example policy for assuming four roles.
The IAM role will need to add a trust policy to allow the service account to assume the role, and permissions to assume the different roles that will be used to access resources. To associate the IAM role with your Kubernetes service account you will need to annotate your service account with the Amazon Resource Name (ARN) of the IAM role that you want the service account to assume. ```sh kubectl annotate serviceaccount -n $namespace $service_account eks.amazonaws.com/role-arn=arn:aws:iam::$account_id:role/my-role ``` ### Step 2: Create the resource roles For the service account role to be able to impersonate the roles you also need to configure the roles themselves to allow it. This is done by adding the service account role to the role's [Trust relationships](https://docs.aws.amazon.com/directoryservice/latest/admin-guide/edit_trust.html). ```json { "Version": "2012-10-17", "Statement": [ { "Effect": "Allow", "Principal": { "AWS": "arn:aws:iam::xxxxxxxxxxxx:role/service-account-role" }, "Action": "sts:AssumeRole" } ] } ```
Example resource roles.
### Step 3: Create mappings Now that the service account IAM role can assume the roles we need to configure Hopsworks to delegate access to the roles on a project base. In Hopsworks, click on your name in the top right corner of the navigation bar and choose _Cluster Settings_ from the dropdown menu. In Cluster Settings, open _IAM Role Chaining_ under _Security & Access_ to configure the mappings between projects and IAM roles.
Role Chaining
Role Chaining
Add mappings by clicking on _New role chaining_. Enter the project name. Select the type of user that can assume the role. Enter the role ARN. And click on _Create new role chaining_
Create Role Chaining
Create Role Chaining
Project member can now create connectors using _temporary credentials_ to assume the role you configured. More details about using temporary credentials can be found in the [Temporary Credentials section](../../user_guides/fs/data_source/creation/s3.md#temporary-credentials) of the S3 datasource creation guide. Project member can see the list of role they can assume by going the _Project Settings_ -> [Assuming IAM Roles](../../user_guides/projects/iam_role/iam_role_chaining.md) page. ================================================================================ # Configure Project Mapping Source: https://docs.hopsworks.ai/latest/setup_installation/admin/configure-project-mapping/ # Configure group to project mapping ## Introduction A group-to-project mapping lets you automatically add all members of a Hopsworks group to a project, eliminating the need to add each user individually. To create a mapping, you simply select a Hopsworks group, choose the project it should be linked to, and assign the role that its members will have within that project. Once a mapping is created, project membership is controlled through Hopsworks group membership. Any updates made to the Hopsworks group, such as adding or removing users, will automatically be reflected in the project membership. For example, if a user is removed from the Hopsworks group, they will also be removed from the corresponding project. ## Prerequisites 1. Hopsworks group mapping sync enabled. This can be done by setting the variable ```hw_group_mapping_sync_enabled=true```. See [Cluster Configuration](./variables.md) on how to change variable values in Hopsworks.
Enable Hopsworks mapping
Enable Hopsworks mapping
If you can not find the variable ```hw_group_mapping_sync_enabled``` create it by clicking on **New variable**.
Create Hopsworks mapping enabled variable
Create Hopsworks group mapping enabled variable
### Step 1: Create a mapping To create a mapping go to **Cluster Settings** by clicking on your name in the top right corner of the navigation bar and choosing *Cluster Settings* from the dropdown menu. Choose *Project Mapping* under *Security & Access* in the left sidebar, then create a new mapping by clicking on *Create new mapping*.
Project mapping tab
Project mapping
This will take you to the create mapping page shown below
Create mapping
Create mapping
Here you can enter your Hopsworks group and map it to a project from the *Project* drop down list. You can also choose the *Project role* users will be assigned when they are added to the project. Finally, click on *Create mapping* and go back to mappings. You should see the newly created mapping(s) as shown below.
Project mappings
Project mappings
### Step 2: Edit a mapping From the list of mappings click on the edit button (:material-pencil:). This will open a popup that will allow you to change the *group*, *project name*, and *project role* of a mapping.
Edit mapping
Edit mapping
!!!Warning Updating a mapping's *group* or *project name* will remove all members of the previous group from the project. ### Step 3: Delete a mapping To delete a mapping click on the delete button. !!!Warning Deleting a mapping will remove all members of that group from the project. ================================================================================ # Airflow 3 operator notes Source: https://docs.hopsworks.ai/latest/setup_installation/admin/airflow3/ # Operator Notes: Airflow 3 on Hopsworks Administrative reference for cluster operators upgrading or installing Hopsworks with Airflow 3. ## What the chart deploys The Airflow subchart now creates four Kubernetes objects in addition to what the v1 chart deployed: 1. `dag-processor` Deployment: runs `airflow dag-processor`, parses DAGs listed in the manifest. Carries only validator keys (no private keys). 2. `keys-bootstrap` pre-install Job: generates two RSA 4096 keypairs (api-server + scheduler) and writes them to the `hopsworks-airflow-keys` Secret. Idempotent; re-runs are no-ops. 3. `db-reset` pre-install Job: drops and recreates the Airflow metadata database before migration. Install-only: gated by `hopsworkslib.isInstall`, so it never re-fires on `helm upgrade` or ArgoCD resync. Set `global._hopsworks.mode=install` on the first install only. 4. `hopsworks-airflow-keys` Secret: four PEM files (two private, two public). Pods project only the keys they need. The existing `webserver` (now Airflow `api-server`) and `scheduler` Deployments keep their resource names; only the container command and environment changed. ## Resource matrix | Component | Image | Command | Private keys mounted | | --- | --- | --- | --- | | `airflow-webserver` | `apache/airflow:3.0.6-python3.12` + Hopsworks layers | `airflow api-server --proxy-headers` | api-server-private | | `airflow-scheduler` | same | `airflow scheduler` | scheduler-private | | `airflow-dag-processor` | same | `airflow dag-processor` | none (validator-only) | All three pods render `/opt/airflow/airflow.cfg` from `airflow.cfg.template` at container start (the runtime MySQL password is substituted into the template before the main process execs). Liveness probes use a TCP check on the scheduler's serve_logs port (8793) and a `/proc/1/comm` read for the dag-processor; `pgrep` is unreliable under kernel hardening such as `kernel.yama.ptrace_scope >= 1` or `hidepid`. ## Key rotation Re-run the `keys-bootstrap` Job with the `--rotate` flag (or delete the `hopsworks-airflow-keys` Secret and run a `helm upgrade`), then rolling-restart in this order: ```text api-server → scheduler → dag-processor ``` Outstanding user cookies and outstanding Execution-API task tokens are invalidated; the proxy re-mints on 401. Plan for ~30s of UI unavailability during the api-server restart. ## Metadata DB on upgrade The v3 chart drops and recreates the Airflow metadata database **on install only**, gated by `hopsworkslib.isInstall` (set `global._hopsworks.mode=install` on the first install). There is no in-place 1.x → 3.x schema migration path; DAG-run history, ad-hoc Variables, ad-hoc Connections, and audit records from a 1.x deployment are not preserved by the cutover. Schema changes after install go through the existing `migrate` job's Alembic migrations. Customer DAG files in HopsFS at `Projects/

/Airflow/` are untouched. Snapshot HopsFS before the upgrade if you need a rollback path that preserves DAG sources. ## Reverse proxy contract `AirflowProxyServlet` in hopsworks-ee validates the Hopsworks JWT and forwards `Set-Cookie` from Airflow unchanged. The proxy does not rewrite cookie `Path=`; Airflow sets the cookie path from `[api] base_url` automatically. Membership is **not** refreshed on every forwarded request. The Airflow JWT carries `project_ids` / `project_roles` / `is_admin` at mint time and is stable for the cookie's TTL (1 hour by default). Real-time membership changes are propagated by the Hopsworks backend pushing to `POST /auth/internal/invalidate`, which evicts the affected user's cached entry so the next login re-fetches the membership. A 60-second safety-net TTL on the cache catches drift even without an explicit invalidation. See [Airflow Security Model](../../user_guides/projects/airflow/security_model.md#token--cookie-behavior) for the full description. ## DAG reconciler `AirflowDagReconciler` is a Hopsworks-side singleton EJB that runs every 60 s (with a 30 s initial delay) on the Hopsworks admin pod. It walks `Projects/

/Airflow/*.py` for every project, derives the canonical `dag_id` (`p____`), and reconciles `dag_project_index` against the on-disk truth: - A `.py` present on HopsFS without a matching index row triggers a row insert. This is the path that picks up files uploaded via the Hopsworks File Browser or copied via `DatasetApi.copy`, without needing a backend restart. - An index row whose `.py` is gone triggers a row delete plus the same `airflow.api.common.delete_dag.delete_dag` cleanup the explicit delete button uses. The reconciler runs under `@TransactionAttribute(NOT_SUPPORTED)` because the Airflow auth-manager HTTP calls it makes do not participate in the EJB global transaction. Its logs appear in the admin pod under the logger `io.hops.hopsworks.common.airflow.AirflowDagReconciler`. ## Orphan cleanup CronJob The chart deploys an `airflow-orphan-cleanup` CronJob (gated by `airflow.enabled`, no separate enable flag) that runs the SQL in `cleanup_orphans.sql` against the Airflow metadata DB. It deletes orphan rows in `dag_run`, `task_instance`, `task_instance_history`, `xcom`, `log`, `dag_warning`, `asset_dag_run_queue` (`target_dag_id`), `task_outlet_asset_reference` (`dag_id`), and `deadline` (both `dag_id` and `dagrun_id`) that point at a `dag_id` no longer in the `dag` table. This is the cleanup path Airflow itself does not run automatically when a DAG is hard-deleted out of band. Only the most recent successful run of the CronJob is retained in-namespace; older Pods are reaped by the CronJob's history limit. ## OpenShift compatibility The airflow image is built to OpenShift's arbitrary-UID + GID-0 contract. `/etc/airflow` and `launcher.sh` are group-owned by root (`chown :0`) with the user permission bits mirrored onto the group (`chmod g=u`), so OpenShift's per-namespace UID (which always has GID 0) can read, write, and execute everything the `airflow` UID can on vanilla Kubernetes. No `runAsUser` override is needed when deploying on OpenShift; the chart's pod-spec works unchanged. ## Metrics The legacy `airflow-exporter` 1.3.0 does not support Airflow 3. Metrics are now emitted via Airflow-native StatsD into a sidecar `statsd_exporter` Pod, scraped by Prometheus on its `/metrics` endpoint. The legacy `AllowMetricsSecurityManager` is gone. ## See also - User-facing release notes: `user_guides/projects/airflow/airflow3_upgrade.md` - Security model: `user_guides/projects/airflow/security_model.md` ================================================================================ # Kafka 4 and KRaft upgrade notes Source: https://docs.hopsworks.ai/latest/setup_installation/admin/kafka4/ # Operator Notes: Kafka 4 and KRaft in Hopsworks 5.2 Administrative reference for cluster operators upgrading a self-managed Hopsworks cluster to 5.2. Hopsworks 5.2 moves the bundled Kafka from Strimzi 0.45 with Kafka 3.9 to Strimzi 1.2 with Kafka 4.3.1. Kafka 4 has no ZooKeeper, so an existing cluster is migrated to KRaft as part of the upgrade. The chart performs that migration itself; this page covers what to prepare and what to expect. The chart's own README has the mechanics, the failure modes and the verification steps; read it with `helm show readme hopsworks/hopsworks --version <5.2 version>`. A fresh 5.2 install needs none of this: it starts on KRaft. ## What the upgrade does Two pre-upgrade hooks run before any new component is applied: 1. `kafka-kraft-migration` carries the cluster from ZooKeeper to KRaft under the operator that is still running. It adds the KRaft node pools, annotates the cluster for migration, waits for the metadata to move, and finalises. 2. `strimzi-crd-upgrade` converts the Strimzi custom resources from the `v1beta2` API to `v1` and brings the CRDs up to the new operator's version. The release then replaces the operator with Strimzi 1.2 and rolls the brokers to Kafka 4.3.1. ## What to expect - **The upgrade blocks while the migration runs.** The hook waits three times, each bounded by `kafka.migrationJob.timeoutSeconds` (default 1800), so the worst case is 3 x that value. Helm's `--timeout` applies to each hook on its own: give it at least that budget plus room for the rest of the release. On timeout the hook fails rather than hangs, and re-running the upgrade resumes from wherever the operator got to. - **Broker pods are replaced about seven times**: during the migration, when the finalised cluster drops its ZooKeeper-era settings, and when the brokers move to Kafka 4.3.1. A single-broker cluster is unavailable for the duration of each replacement; clusters with a replication factor above 1 roll without interruption. - **ZooKeeper and the new controllers run side by side** until the migration is finalised. The controller pool inherits its sizing from the ZooKeeper settings, so the namespace needs headroom for both during the upgrade. - **Finalising KRaft is irreversible.** Do not `helm rollback` after the migration has started, and do not run this upgrade with `--atomic`, which performs that rollback for you the moment a hook times out: the earlier revision puts the ZooKeeper-shaped Kafka resource back against a cluster that may already have finalised KRaft. Re-run the upgrade instead. - **One `helm upgrade` migrates the cluster.** The Strimzi CRDs finish moving on the next upgrade, whatever its reason; until then the 1.2.0 operator runs on the 0.51.0 schemas, which nothing the chart sets depends on. - **ArgoCD needs one manual step.** ArgoCD renders the chart from its cached list of cluster APIs before the hooks run, so the first sync converts the CRDs and then stops on purpose with a message in the `strimzi-crd-upgrade` Job. Then invalidate the cluster cache, hard-refresh the Application, and sync again; the second sync completes the upgrade. In the UI: Settings, Clusters, the cluster, Invalidate Cache; then Refresh (Hard) and Sync on the Application. From a terminal: ```sh curl -k -X POST "$ARGOCD_URL/api/v1/clusters/https%3A%2F%2Fkubernetes.default.svc/invalidate-cache?id.type=server" \ -H "Authorization: Bearer $ARGOCD_TOKEN" -H "Content-Type: application/json" -d '{}' argocd app get hopsworks --hard-refresh argocd app sync hopsworks ``` The Application also needs an `ignoreDifferences` entry for the Kafka CRD, on a fresh install as much as on an upgrade. Upstream's 1.2.0 Kafka CRD declares an empty `properties: {}` map under `status.clusterSecurity`, and the apiserver drops an empty map when it stores a CRD, so the live object never matches the render: without the entry the Application never reaches Synced and selfHeal re-applies that one CRD every five minutes. ```yaml spec: ignoreDifferences: - group: apiextensions.k8s.io kind: CustomResourceDefinition name: kafkas.kafka.strimzi.io jqPathExpressions: - .spec.versions[].schema.openAPIV3Schema.properties.status.properties.clusterSecurity.properties ``` ## Before you start - **Airgapped clusters** mirror three images with the rest: `strimzi/operator`, `strimzi/kafka` and `strimzi/crds`. The chart's `vendor_images.sh` includes them; nothing is fetched from the internet during the upgrade. - **A satellite cluster that shares the central operator** is upgraded before central. The CRD conversion re-stores every Kafka resource on the Kubernetes cluster, and a satellite still on ZooKeeper would lose its ZooKeeper configuration in the process. The conversion refuses while any other Kafka is not on KRaft, so the wrong order fails safely, but it blocks central's upgrade until the satellite is done. - **Clusters first installed on Hopsworks 4.3 or 4.5** still run Strimzi 0.39 CRDs, whose status schema cannot carry the migration state the hook reads, so the hook refuses rather than guess. Apply the 0.45.2 CRD bundle first: ```sh kubectl apply --server-side -f \ https://github.com/strimzi/strimzi-kafka-operator/releases/download/0.45.2/strimzi-crds-0.45.2.yaml ``` Wait for the operator to reconcile, then run the upgrade. Do not apply the 5.2 chart's own CRDs for that: the apiserver rejects them while `v1beta2` is still a stored version. An airgapped cluster has to mirror that bundle too; the `strimzi/crds` image carries only the 0.51.0 and 1.2.0 bundles. - **Record the topics**, so you can confirm nothing was lost: ```sh kubectl exec -n -kafka-0 -c kafka -- bash -c ' topics=$(ls /var/lib/kafka/data-0/kafka-log0 | grep -E -- "-[0-9]+$" | grep -v "^__cluster_metadata" | sed -E "s/-[0-9]+$//" | sort -u) printf "%s\n" "$topics" printf "%s\n" "$topics" | wc -l' ``` ## Verify ```sh kubectl get strimzipodset -zookeeper -n --ignore-not-found kubectl get pods -n -l strimzi.io/name=-zookeeper kubectl get pods -n -l strimzi.io/controller-role=true ``` Nothing for the first two, at least one controller pod for the third, and the same topic list as before. If the upgrade fails part-way, re-run it. The hook's log names what it was waiting for and the last state it saw, a timeout corrupts nothing because the operator migrates on its own schedule, and the re-run resumes from the live state. The one state the hook refuses outright is a cluster whose CRDs are too old to report `.status.kafkaMetadataState`, which is the 4.3 and 4.5 case above. ## Client compatibility {#kafka4-client-compatibility} Kafka 4 dropped support for old protocol versions ([KIP-896](https://cwiki.apache.org/confluence/display/KAFKA/KIP-896%3A+Remove+old+client+protocol+API+versions+in+Kafka+4.0)). The 4.3.1 brokers accept Java clients from 2.1 and librdkafka-based clients (Python, Go, .NET) from 1.8.2. The clients shipped with Hopsworks 5.2 are within those ranges; check any producer or consumer you build against the cluster yourself. The same cut applies the other way round for a [bring-your-own Kafka cluster][external-kafka-cluster]: Hopsworks 5.2 connects with kafka-clients 4.3.1, so the external brokers must run Kafka 2.1 or newer. ================================================================================ # Search Index Source: https://docs.hopsworks.ai/latest/setup_installation/admin/search_index/ # Search Index Administration { #search-index-administration } ## Introduction Hopsworks keeps a search index of the artifacts in a deployment: feature groups, feature views, training datasets, jobs, models and deployments. The index is updated from the database, not written to directly. Every change to a searchable property queues a command in the same transaction that made the change, and a background executor applies the queued commands to OpenSearch. That design is what makes the index eventually consistent rather than immediately correct, and it is the thing to understand before diagnosing a stale search result. A tag that was attached seconds ago and does not appear in search yet is normal. A tag that has not appeared after minutes means a command is stuck, which this page covers. Only administrators can reach any of this. ## Commands needing attention Go to `Cluster Settings` > `Service Operations` and open the `OpenSearch Index Commands` tab. The page opens on the `Service Operations` tab, which carries the platform's operation logs; the index commands share the page because both answer whether the platform is digesting what it was asked to do. You do not need to check the tab on the off-chance. While any command needs attention, the `Service Operations` entry in the settings menu shows a pulsing orange dot and the `OpenSearch Index Commands` tab is outlined in orange. Both appear and clear on their own as the queue fails and drains, so no dot means there is nothing to look at. The tab lists every command that has failed at least once, oldest first. For each it shows the document, the artifact type, the operation, the number of attempts made, when the next attempt is due, the project, and the last error. When the list is empty, every queued update has been applied. The reason a stuck command matters more than one failure suggests is ordering. Commands for a single document are applied in order, so while one keeps failing, every later update to that same document waits behind it, and that artifact's search result stays as it was. Other documents are unaffected. Retries are automatic and unbounded, with the delay growing between attempts. Unbounded is deliberate: abandoning a command would not lose only its own update, it would strand every later update to that document. So a command in this list is usually a transient failure that will clear itself, and the list is worth acting on when an entry stops making progress. The page makes that call for you at 20 attempts. A command that has failed 20 times or more is marked `stuck`, with a red border and a callout explaining what to do. The threshold is where the retry delay has been at its cap for a while, so crossing it means roughly an hour of the same failure: whatever a retry could outwait has had its chance. Read the stuck command's error first, because a failure whose cause has since been fixed clears itself on the next retry and needs nothing from you. Otherwise cancel it, as described below. ### Cancelling a command Cancel a command only to let the ones behind it through. The update the cancelled command carried is lost. The document keeps whatever the index already held, so cancelling an `UPDATE_TAGS` leaves the old tags visible in search until something changes that artifact again and queues a fresh command. That is why cancelling is an explicit operator action and never something the executor does on its own. A command can only be cancelled while it is `FAILED`, which is the state between attempts where nothing owns the row. A command that has just been picked up for another attempt cannot be cancelled, and the request is refused with that reason rather than interrupting the attempt. Wait for the attempt to finish and retry. Each cancellation is logged with the command id, the document, the attempt count, the requesting administrator and the last error, because once the row is gone the log is the only record it existed. ## Rebuilding the index A reindex rebuilds the search index from the database. It is the recovery path after the index has been lost or has diverged, not routine maintenance. ```bash # queue a reindex, returns the run id curl -X POST -H "Authorization: Bearer $JWT" \ "https://$HOPSWORKS_HOST/hopsworks-api/api/admin/search/featurestore/reindex" # follow that run curl -H "Authorization: Bearer $JWT" \ "https://$HOPSWORKS_HOST/hopsworks-api/api/admin/search/featurestore/reindex/$RUN_ID" ``` The `POST` returns `200` with the run id rather than an empty `204`. The run id is what makes progress observable: without it a caller can only watch the global command queue, which mixes the reindex with whatever ordinary use is queueing alongside it. Reindexing queues a command per document, so on a large deployment it takes a while and shows up as a long queue. That is expected, and ordinary updates continue to be applied while it drains. ## Configuration Set these in `Cluster Settings` > `Configuration`. | Variable | Default | Effect | | --- | --- | --- | | `cross_project_global_search_enabled` | `true` | Whether a search may return artifacts from projects the caller is not a member of. Set `false` on a multi-tenant deployment to restrict every search to the projects the caller can already access. | | `command_search_fs_retry_backoff_base_as_ms` | `5000` | The delay before the first retry of a failed command. The delay doubles per attempt from here. | | `command_search_fs_retry_backoff_max_as_ms` | `300000` | The ceiling on that growing delay, so a long-failing command keeps retrying at a fixed interval rather than drifting to never. | Lowering the backoff makes a transient failure clear sooner at the cost of more load against OpenSearch while it is failing. Raising it does the opposite. !!! warning "Turning off cross-project search does not rewrite history" `cross_project_global_search_enabled` filters searches as they are made. It does not remove anything from the index, so switching it off restricts what users can find from that point on and is not a way to redact something already indexed. ================================================================================ # Upgrading OpenSearch Source: https://docs.hopsworks.ai/latest/setup_installation/admin/upgrade_opensearch/ # Upgrading OpenSearch { #upgrading-opensearch } ## Introduction Each Hopsworks release ships a fixed OpenSearch version, and upgrading Hopsworks upgrades it. | Hopsworks | OpenSearch | | --------- | ---------- | | 4.8.x, 5.0.x | 1.3.x | | 5.1.x | 2.19.x | | 5.2.x | 3.8.0 | This page covers what changes for you on the 5.1 to 5.2 upgrade, because that is the one that can interrupt service, and the steps to run it. The commands assume Hopsworks is installed in the `hopsworks` namespace. ## Upgrade one minor release at a time Upgrade 5.0.x to 5.1.x, let it come up, then upgrade to 5.2.x. Going from 5.0.x straight to 5.2.x is not supported. OpenSearch 3.x refuses to start if the cluster still holds an index that was created before OpenSearch 2.0. OpenSearch 2.19 keeps serving indices created by 1.x without rewriting them, so a cluster that has passed through 5.1.x can still hold such indices. When a 3.x node meets one it stops during startup, while its container keeps reporting `Running`, so the first pod of the rolling upgrade never becomes ready and nothing in `kubectl get pods` says why. ## The index pass on 5.1.x to 5.2.x A `pre-upgrade` hook runs before the OpenSearch StatefulSet is touched, so the old cluster is still intact if it fails. It checks every index against the OpenSearch 2.0 floor. A cluster that has only ever held indices created by 2.x has nothing to fix, and the hook exits immediately. On a cluster installed before 5.1 the hook repairs what it finds: - It deletes only the indices the platform recreates or expires by itself: logs, audit logs, `pypi_libraries_*`, and plugin indices that the plugin repopulates. - It reindexes everything else with the name preserved, including `featurestore`, `projects`, the OpenSearch Dashboards saved objects, the index state management policies and the embedding indices behind similarity search. Nothing else is deleted, because anything the hook does not recognise might hold data nothing can rebuild. ### Plan for an outage While the hook copies indices, every client except the hook is refused, reads included. Hopsworks search, OpenSearch Dashboards, logstash and OnlineFS all receive HTTP 403 until the hook exits. The hook does this so that no write can land in an index while it is being copied and then be overwritten by the copy coming back. Rows written to feature groups with embeddings during the window are marked failed by OnlineFS rather than retried, because OnlineFS treats a rejected write as a permanent failure. Re-ingest those feature groups after the upgrade. Log lines shipped by logstash during the window may also be missing. The hook restores access when it exits, whether it succeeds or fails. If it is killed outright, for example by a node loss, clients stay refused until you run the upgrade again, because the next run picks up where it stopped and then lets clients back in. ### Set the timeout The hook is bounded by `olk.opensearch.indexUpgrade.activeDeadlineSeconds`, which defaults to 3600 seconds. Pass `helm upgrade --timeout` at least that long, or Helm gives up on the hook before the hook finishes. A reindex copies each index twice to keep its name, so the time grows with the size of your old indices, `featurestore` and the embedding indices in particular. As a guide, a copy ran at roughly 15 MB per second on a 16-CPU node, and slower disks or fewer CPUs take longer. Before it starts, the hook prints the size of every index it will copy, an estimate, and the configured deadline, and warns when the estimate exceeds the deadline. To read that report without changing anything, set `olk.opensearch.indexUpgrade.enabled` to `false`. The hook then lists the indices and fails, and you can raise the deadline and the timeout before running the real upgrade. Free disk space matters as well, because a copy needs room for a second complete copy of the largest index. The hook checks this on the smallest data node and stops if there is not enough. ## Embedding indices move to faiss OpenSearch indexes embeddings with an engine, and Hopsworks 5.1 and later create embedding indices on faiss. Indices created by Hopsworks 5.0 and earlier use nmslib, which OpenSearch has deprecated and which does not accept the filter that the Hopsworks 5.1 and later clients send with every similarity search. The index pass therefore recreates each embedding index on faiss while it copies it, so similarity search with filters works once the upgrade to 5.2 is done. Until then, on 5.1, filtered search fails on feature groups whose embedding index was created on 5.0, as described in the upgrade steps below. Approximate nearest neighbor results can shift slightly, because the graph is rebuilt on a different engine. Indices that you created yourself, outside the Hopsworks embedding naming, keep their engine. ## Satellite clusters A satellite cluster runs its own OpenSearch and gets its own index pass when it upgrades to 5.2. The outage is confined to that satellite's OpenSearch, and `activeDeadlineSeconds` is set per release, so a satellite holding large embedding indices sets its own value and its own `--timeout`. Upgrade every satellite to 5.1.x before the central cluster moves to 5.2.x. A satellite still on OpenSearch 1.3 would be two major versions behind a 5.2.x central cluster. A satellite on 5.2.x while the central cluster is still on 5.1.x should work, but it is untested and not the recommended order. ## Upgrade steps 1. Take a snapshot and check that it completed, because the index pass rewrites indices in place. 2. Upgrade to 5.1.x if you are not already on it, and confirm that OpenSearch 2.19 is live and green: ```bash kubectl exec -n hopsworks opensearch-0 -c opensearch -- bash -c \ 'curl -sk -u "admin:$ADMIN_PASSWORD" https://localhost:9200 | grep number' kubectl exec -n hopsworks opensearch-0 -c opensearch -- bash -c \ 'curl -sk -u "admin:$ADMIN_PASSWORD" https://localhost:9200/_cluster/health?pretty | grep status' ``` The indices are still the ones 1.x created at this point. That is expected, since 2.19 does not rewrite them. Two known issues apply to this upgrade: - On the first OpenSearch 2.19 start of each cluster, central or satellite, vector writes are rejected for about two minutes and the vectors that had not been flushed are lost. The chart fixes this on the 5.1 line, and a values file that copied the old `knn.circuit_breaker` settings (`triggered: true` or `percent: 0.75`) keeps the problem. - On 5.1, similarity search with a filter fails on feature groups whose embedding index was created on 5.0, because the 5.1 client sends the filter inside the nearest-neighbor query and the nmslib engine rejects it. A client fix and the pass in step 3 both resolve it. 3. Upgrade to 5.2.x with a `--timeout` at least as long as `olk.opensearch.indexUpgrade.activeDeadlineSeconds`, and watch the hook: ```bash kubectl logs -n hopsworks -l component=opensearch-index-upgrade -f ``` It ends with `Every index is now on OpenSearch 2.0 or later. Safe to upgrade.` 4. Verify the upgrade: ```bash kubectl get pods -n hopsworks -l app=opensearch kubectl logs -n hopsworks opensearch-0 -c opensearch | grep -c StartupException kubectl exec -n hopsworks opensearch-0 -c opensearch -- bash -c \ 'curl -sk -u "admin:$ADMIN_PASSWORD" https://localhost:9200 | grep -E "number|minimum_index"' ``` The `StartupException` count must be 0. Expect version `3.8.0`, `minimum_index_compatibility_version` `2.0.0` and green cluster health. Then run a search in the Hopsworks UI and a similarity search on a feature group with an embedding. 5. Re-ingest the feature groups that OnlineFS marked failed during the outage. ## If the hook fails The upgrade stops before the OpenSearch StatefulSet is touched, so the cluster is still on 5.1.x, and the hook lets clients back in on its way out. The hook log names the index and the reason. - **Not enough disk:** free space on the smallest data node and retry. The hook also stops when it cannot measure free disk or an index size, rather than skipping the check. - **A closed index:** open it, or reindex it by hand, and retry. The hook refuses to copy a closed index because it can neither size nor read it, and nothing in Hopsworks leaves an index closed, so this usually means a snapshot restore that died. Leaving it closed does not help, since a 3.x node still refuses to start. - **Timed out:** raise `activeDeadlineSeconds` and `--timeout` together and retry. The retry resumes any index that was in the middle of a copy. - **No answer from the cluster:** the hook prints the address and the last HTTP code. `000` is a DNS or connection failure, or a rejected admin certificate. `401` or `403` means the certificate's distinguished name is not in `plugins.security.authcz.admin_dn`. To reindex by hand instead, set `olk.opensearch.indexUpgrade.enabled` to `false`. The hook then reports the offending indices and fails, and for each one you run: ```text POST /_reindex {"source":{"index":""},"dest":{"index":"-tmp"}} DELETE / POST /_reindex {"source":{"index":"-tmp"},"dest":{"index":""}} DELETE /-tmp ``` Three things have to be right when you do it by hand: - Create the destination index first from the source's settings and mappings, because `_reindex` builds its destination from index templates and silently drops `index.knn` and `knn_vector` mappings. - For an embedding index, set `method.engine` to `faiss` in the destination mapping, as the hook does, for `l2`, `cosinesimil` and `innerproduct` fields. A destination that keeps nmslib is created on 2.x, so 3.8 opens it, but filtered similarity search on it still fails, and a 3.8 node on a CPU without AVX-512 crashes when it loads it. - Set `index.blocks.write` to `true` on the source before you copy it, so writes during the copy are refused instead of missed by it. `_reindex` reads its source through a scroll, which the block does not cover, so the copy still runs, whereas a block on the destination refuses it. Keep writers stopped from the `DELETE` until the recreated index is filled, because a write to the missing name creates it from the index templates. - Use the admin certificate for `.opendistro_security`, `.opendistro-ism-config` and `.plugins-ml-config`, because these protected system indices answer 403 to the `admin` user's password. ================================================================================ # Services Dashboards Source: https://docs.hopsworks.ai/latest/setup_installation/admin/monitoring/grafana/ # Services Dashboards ## Introduction The Hopsworks platform is composed of different services. Hopsworks uses Prometheus to collect health and performance metrics from the different services and Grafana to display them to the Hopsworks administrators. In this guide you will learn how to access the Grafana dashboards to monitor the health of the cluster or to troubleshoot performance issues. ## Prerequisites To access the services dashboards in Grafana, you need to have an administrator account on the Hopsworks cluster. ## Step 1: Access Grafana You can access the admin page of your Hopsworks cluster by clicking on your name, in the top right corner, and choosing _Cluster Settings_ from the dropdown menu. You can then choose _Monitoring_ under _Observability_ in the left sidebar. The _Monitoring_ page gives you access to several of the observability tools that are already deployed to help you manage the health of the cluster.

monitoring page
Monitoring page
Click on the _Grafana_ link to open the Grafana web application and navigate through the dashboards. ## Step 2: Navigate through the dashboards In the Grafana web application, you can click on the _Home_ button on the top left corner and navigate through the available dashboards. Dashboards are organized into three folders: - **Hops**: This folder contains all the dashboards of the Hopsworks services (e.g., the web application, the file system, resource manager) as well as the dashboards of the hosts (e.g., EC2 instances, virtual machines, servers) on which the cluster is deployed. - **RonDB**: This folder contains all the dashboard related to the database. The _Database_ dashboard contains a general overview of the RonDB cluster, while the remaining dashboards focus on specific items (e.g., thread activity, memory management, etc). - **Kubernetes**: If you have integrated Hopsworks with a Kubernetes cluster, this folder contains the dashboards to monitor the health of the Kubernetes cluster.
Grafana view
Grafana view
The default dashboards are read only and cannot be edited. Additional dashboards can be created by logging in to Grafana. You can log in into Grafana using the username and password specified in the cluster definition. !!! warning By default Hopsworks keeps metrics information only for the past 15 days. This means that, by default, you will not be able to access health and performance metrics which are older than 15 days. ## Going Further You can read [Grafana Documentation](https://grafana.com/docs/) to learn how to use it advancedly. ================================================================================ # Export metrics Source: https://docs.hopsworks.ai/latest/setup_installation/admin/monitoring/export-metrics/ # Exporting Hopsworks metrics ## Introduction Hopsworks services produce metrics which are centrally gathered by [Prometheus](https://prometheus.io/) and visualized in [Grafana](./grafana.md). Although the system is self-contained, it is possible for another *federated* Prometheus instance to scrape these metrics or directly push them to another system. This is useful if you have a centralized monitoring system with already configured alerts. ## Prerequisites In order to configure Prometheus to export metrics you need to have the right to change the remote Prometheus configuration. ## Exporting metrics Prometheus can be configured to export metrics to another Prometheus instance (cross-service federation) or to a custom service which knows how to handle them. ### Prometheus federation Prometheus servers can be federated to scale better or to just clone all metrics (cross-service federation). In the guide below we assume **Prometheus A** is the service running in Hopsworks and **Prometheus B** is the server you want to clone metrics to. #### Step 1 **Prometheus B** needs to be able to connect to TCP port `9090` of **Prometheus A** to scrape metrics. If you have any firewall (or Security Group) in place, allow ingress for that port. #### Step 2 The next step is to expose **Prometheus A** running inside Hopsworks Kubernetes cluster. If **Prometheus B** has direct access to **Prometheus A** then you can skip this step. We will create a Kubernetes *Service* of type *LoadBalancer* to expose port `9090` !!!Warning If you need to apply custom **annotations**, then modify the Manifest below. The example below assumes Hopsworks is **installed** at Namespace *hopsworks* ```bash kubectl apply -f - < monitoring page
Monitoring page
Click on the _Service Logs_ link to open the OpenSearch Dashboards web application and navigate through the logs. ## Step 2: Search the logs In the OpenSearch dashboard web application you will see by default all the logs generated by all monitored services in the last 15 minutes. You can filter the logs of a specific service by searching for the term `service:[service name]`. As shown in the picture below, you can search for the _namenode_ logs by querying `service:namenode`. Currently only the logs of the following services are collected and indexed: Hopsworks web application (called `domain1` in the log entries), namenodes, resource managers, datanodes, nodemanagers, Kafka brokers, Hive services and RonDB. These are the core component of the platform, additional logs will be added in the future.
OpenSearch Dashboards with services logs
OpenSearch Dashboards displaying the logs
!!! warning By default, logs are rotated automatically after 7 days. This means that by default, you will not be able to access logs through OpenSearch Dashboards which are older than 7 days. Depending on the service and on the Hopsworks configuration, you can still access the logs by SSH directly into the machines of the cluster. ## Going Further See [Service Log Labels](services-logs-labels.md) for the Kubernetes labels attached to each log document, and how to append to or replace that set. You can read [OpenSearch Dashboards Documentation](https://opensearch.org/docs/latest/dashboards/) to learn how to use them advancedly. ================================================================================ # Service Log Labels Source: https://docs.hopsworks.ai/latest/setup_installation/admin/monitoring/services-logs-labels/ # Service Log Labels ## Introduction Filebeat attaches Kubernetes metadata to every service log document before Logstash forwards it to OpenSearch. Only an explicit set of pod, namespace and node labels is kept. The set is bounded because every distinct label key becomes a field in the shared `.services-*` index mapping. OpenSearch rejects documents once an index exceeds `index.mapping.total_fields.limit`, which defaults to 1000 fields. A cluster that attaches every pod and node label, including the label sets that cloud providers and node feature discovery add, reaches that limit and then stops indexing service logs. In this guide you will learn which labels are kept by default, and how to append to or replace that set. ## Default pod labels | Label | Purpose | | --- | --- | | `name` | Excludes Filebeat's own logs from collection. | | `app` | Service identification, read by most pipelines. | | `app.kubernetes.io/name` | Service identification for charts using the recommended Kubernetes labels. | | `service` | Service name. | | `component` | Component within a service. | | `rondbService` | RonDB process type. | | `user` | Owner of a job, notebook or serving instance. | | `job-type` | Job type. | | `job-id` | Job identifier. | | `job-name` | Job name. | | `execution` | Execution identifier of a job run. | | `jupyter` | Marks a Jupyter pod. | | `jupyter-id` | Jupyter instance identifier. | | `jupyter-settings-id` | Jupyter settings identifier. | | `kernel-id` | Jupyter kernel identifier. | | `spark-role` | Driver or executor. | | `spark-app-selector` | Spark application identifier. | | `sparkoperator.k8s.io/launched-by-spark-operator` | Marks pods created by the Spark operator. | | `serving.hops.works/id` | Deployment identifier. | | `serving.hops.works/name` | Deployment name. | | `serving.hops.works/tool` | Serving tool. | | `serving.hops.works/model-name` | Model name. | | `serving.hops.works/model-version` | Model version. | | `serving.hops.works/model-server` | Model server. | | `serving.hops.works/project-id` | Project that owns the deployment. | ## Default namespace labels | Label | Purpose | | --- | --- | | `hopsworks.ai/project` | Marks a project namespace. | | `hopsworks.ai/onlinefs-cluster` | Marks an online feature store namespace. | Filebeat collects logs from the release namespace, from namespaces carrying either of these two labels, and from any namespace listed in `olk.filebeat.extraNamespaces`. ## Default node labels | Label | Purpose | | --- | --- | | `kubernetes.io/hostname` | Node name, read by the services, Spark, Python and serving pipelines. | ## How labels appear in a log document ```json { "kubernetes": { "labels": { "app": "namenode", "app_kubernetes_io/name": "hopsfs" }, "namespace_labels": { "hopsworks_ai/project": "demo" }, "node": { "labels": { "kubernetes_io/hostname": "worker-1" } } } } ``` Dots in a label key are replaced by underscores. In OpenSearch Dashboards, search for `kubernetes.labels.app_kubernetes_io/name`, not `kubernetes.labels.app.kubernetes.io/name`. ## Labels Hopsworks' log filtering reads Four defaults are unioned into the effective set whatever the lists below say, because losing one breaks log collection rather than degrading a field: | Label | Dimension | What breaks without it | | --- | --- | --- | | `name` | pod | Filebeat collects its own logs. It logs every OpenSearch rejection with the document embedded, so this feeds back on itself. | | `hopsworks.ai/project` | namespace | Project namespaces are no longer collected. | | `hopsworks.ai/onlinefs-cluster` | namespace | Online feature store namespaces are no longer collected. | | `kubernetes.io/hostname` | node | The services, Spark, Python and serving pipelines lose the node field. | `name` and `hopsworks.ai/onlinefs-cluster` are read by the chart's own Filebeat configuration rather than by a pipeline. `hopsworks.ai/project` is read by both: the Filebeat namespace gate and the `discriminator.conf` pipeline, which derives the project field from it. They cost eight fields of the 1000-field budget, not four: a dynamically mapped string label is indexed as `text` plus a `.keyword` sub-field, so every key counts twice. The same doubling applies to any label you append, which halves the effective budget. Every default is in this category: the audit of the shipped set found no key that is collected without something reading it. They live in the `logFiltering*` values, so the set is visible and auditable: ```yaml olk: filebeat: kubernetesMetadata: logFilteringPodLabels: [ ... ] logFilteringNamespaceLabels: - "hopsworks.ai/project" - "hopsworks.ai/onlinefs-cluster" logFilteringNodeLabels: - "kubernetes.io/hostname" ``` The effective set for a dimension is `logFiltering*` plus `extra*`, de-duplicated. Removing a key from `logFiltering*` is possible, since Helm cannot make a value read-only, and breaks whatever reads it: the four above stop log collection, and the rest stop a Logstash routing or enrichment branch from matching. A location that genuinely needs different metadata should set `addKubernetesMetadata: false` and supply its own processors instead. A log location that genuinely needs different metadata should instead set `addKubernetesMetadata: false` and supply its own processors, as described below. ## Metadata collection per log location Collection is switched on per entry of `olk.filebeat.logs_locations`: ```yaml olk: filebeat: logs_locations: - name: containerd path: /var/log/containers mountPaths: - /var/log/containers - /var/log/pods logtype: log glob: "/*.log" addKubernetesMetadata: true processors: [] ``` Helm replaces lists instead of merging them, so overriding `logs_locations` replaces the chart's entry in full, `addKubernetesMetadata` included. !!! warning An override that omits `addKubernetesMetadata` and does not supply its own `add_kubernetes_metadata` processor collects no Kubernetes metadata at all. Log lines still reach OpenSearch, but every pipeline branch that routes on `kubernetes.*` stops matching, and nothing reports an error. Set `addKubernetesMetadata: false` only for a location that supplies its own `add_kubernetes_metadata` in `processors`, which keeps two metadata processors from running on the same location. ## Append a label Use the `extra` lists to keep the defaults and add to them: ```yaml olk: filebeat: kubernetesMetadata: extraPodLabels: - "my.corp/team" extraNamespaceLabels: - "my.corp/cost-center" extraNodeLabels: - "topology.kubernetes.io/zone" ``` The extra lists are appended to the defaults and de-duplicated, so repeating a default is harmless. Label keys named by `olk.logstash.extendServicesPipeline` are added automatically and do not need an entry here. An appended label is stored on the log document and is searchable and aggregatable in OpenSearch Dashboards. Routing and field extraction are done by the Logstash pipelines, which read a fixed set of keys, so an appended label does not change how a log line is parsed. ## Replace the filtering set Override `logFiltering*` to replace the set that Hopsworks' filtering reads. Only do this if you know what stops working: ```yaml olk: filebeat: kubernetesMetadata: logFilteringPodLabels: - "app" - "my.corp/team" ``` The `logFiltering*` lists hold the keys Hopsworks' own log filtering reads, so removing one silently breaks whatever reads it. ## Collect annotations Annotations are not collected by default. Add the keys you need: ```yaml olk: filebeat: kubernetesMetadata: podAnnotations: - "my.corp/owner" namespaceAnnotations: [] nodeAnnotations: [] ``` !!! warning Annotations count towards the same 1000-field limit as labels. List individual keys rather than collecting all annotations. ## Deployment and cron job names Two toggles control whether `kubernetes.deployment.name` and `kubernetes.cronjob.name` are attached: ```yaml olk: filebeat: kubernetesMetadata: deployment: true cronjob: true ``` Set a toggle to `true` to add the field, `false` to drop it, or leave it unset to keep Filebeat's own default. !!! note Set the toggle explicitly if you depend on the field. Elastic documents the default inconsistently: the `add_kubernetes_metadata` processor reference says the name is added unless disabled, while the autodiscover provider reference says it is not added unless enabled. To see what your cluster does, look for `kubernetes.deployment.name` on a service log document in Dashboards. ## Going Further See [Services Logs](services-logs.md) for accessing the collected logs in OpenSearch Dashboards. ================================================================================ # WebSocket Proxy Pool Source: https://docs.hopsworks.ai/latest/setup_installation/admin/monitoring/websocket-pool/ # WebSocket Proxy Pool ## Introduction Jupyter, terminals, and Streamlit apps are reached through Hopsworks by way of a WebSocket-aware proxy on the `hopsworks-instance` pods. Each browser WebSocket that flows through that proxy is bridged to its upstream (the kernel, shell, or app pod) for the full lifetime of the connection, and every open bridge counts as one session against a per-pod cap. When a pod reaches that cap, the next upgrade is rejected and the user-facing component (for example, JupyterLab) cannot reach its kernel. This page describes how to monitor the proxy through Grafana, how to read the metrics that the platform exposes, and how to tune the cap from the Helm chart when the cluster has more concurrent notebook users than the defaults can serve. The same counts drive the user-visible capacity badges described in [Session Capacity Warnings][session-capacity-warnings]; admin tuning here shifts the threshold those badges report against. ## Prerequisites To access the Grafana dashboards described here, follow [Services Dashboards](./grafana.md) and confirm that you can open the `Hopsworks` dashboard under the `Hops` folder. To change the session cap or the proxy buffers, you need to be able to edit the Helm values file used to install the chart and run `helm upgrade`. ## How the cap relates to concurrent users Each open browser WebSocket the proxy is bridging counts as one in-flight session, and all three WebSocket-backed components (Jupyter kernels, terminals, and Streamlit apps) draw from the same per-pod budget. The pod saturates at `maxSessionsPerApp` concurrent connections, which defaults to `500`. Unlike the previous proxy, there is no two-threads-per-connection accounting: the cap is a direct count of open sessions, not a derived thread number. There is no queue. Once a pod is at the cap, the next upgrade is closed immediately with a WebSocket `1013 TRY_AGAIN_LATER` close whose reason string is `WebSocket session cap reached (maxSessionsPerApp)`, rather than blocking the user on a connection that will never attach. Rejecting cleanly protects the pod from connection-driven memory growth. ## Grafana panels The WebSocket panels live at the bottom of the `Hopsworks` dashboard. All read directly from metrics emitted by Payara on the `hopsworks-instance` pods. - **WebSocket - rejection rate** plots `sum(rate(ws_connection_rejections_total[1m]))`. Any non-zero value means at least one user was turned away from a Jupyter, terminal, or Streamlit session in the last minute. This is the alertable signal for "the cap is too low." This counter covers cap-reached rejections only; a connection that is refused because its upstream pod is unreachable is counted separately under `ws_upstream_connect_failures_total` (see the metrics table), so this signal stays a clean measure of saturation. - **WebSocket - duration (sliding window percentiles)** plots p50, p95, and p99 of the duration of WebSocket connections after they close. Values come from an exponentially-decaying sample reservoir with a roughly five-minute half-life, so closed long-lived sessions decay out of the percentiles quickly rather than dominating them forever. - **WebSocket - sessions** plots the current open-connection count per `hopsworks-instance` pod as solid lines, plus a single dashed red `Pool max (per instance)` line that shows the per-pod saturation point. A pod whose solid line meets the dashed line is rejecting new sessions. - **WebSocket - pool CPU** plots `ws_pool_cpu_cores` per pod, in CPU cores. A value of `1.0` means the proxy's forwarding threads are consuming one full CPU core on that pod. Compare against the pod's CPU request or limit to judge whether to scale the pod's CPU allocation up. - **WebSocket - pool allocation rate (GC pressure)** plots `ws_pool_alloc_bytes_per_second` per pod. This is heap allocation rate, not residency. A high rate means the forwarding threads are producing garbage, which translates to GC pressure on the JVM; it does not mean the proxy is holding that much memory. For "is the JVM heap getting full?" use the `Memory` panel above; for "is the pod approaching its k8s memory limit?" use the Kubernetes / Pods dashboard. The per-pod panels carry the pod short identifier in the legend (the `-` tail; the `hopsworks-instance-` prefix is stripped because every series on this dashboard comes from an instance pod by construction). ## Prometheus metric reference The metrics are emitted on the standard Payara MicroProfile Metrics endpoint (`/metrics`) under the `application` scope. They are scraped by the in-cluster Prometheus server alongside the other Hopsworks service metrics, so federation and `remote_write` (see [Exporting Hopsworks metrics](./export-metrics.md)) cover them automatically. The metric names are unchanged from the previous proxy implementation, so existing dashboards and alerts keep working; only the data source moved. | Metric | Type | Meaning | | --- | --- | --- | | `ws_connection_in_flight_count` | gauge | Inbound WebSocket sessions currently open on this pod, one per browser leg the proxy is bridging. Tracked directly from the bridge's session open/close, not derived from a thread count. | | `ws_connection_in_flight_max_age_seconds` | gauge | Age, in seconds, of the oldest currently open connection. | | `ws_pool_max_connections` | gauge | Configured saturation point for one pod: the `maxSessionsPerApp` cap above which new upgrades are rejected with `1013 TRY_AGAIN_LATER`. | | `ws_connection_rejections_total` | counter | Cumulative count of connections rejected because the cap was reached. | | `ws_upstream_connect_failures_total` | counter | Cumulative count of sessions closed because the proxy could not reach the upstream (the kernel, shell, or app pod is not accepting connections). Distinct from `ws_connection_rejections_total` even though both close the browser with `1013 TRY_AGAIN_LATER`: this one means "the upstream is not up", not "the pod is saturated". | | `ws_connection_duration_seconds` | summary | Count, sum, max, mean, and quantile values for the duration of closed connections, in seconds. | | `ws_pool_cpu_cores` | gauge | CPU cores currently in use by the proxy's shared client transport threads. Computed from per-thread `ThreadMXBean.getThreadCpuTime` deltas divided by wall-clock dt; `1.0` means one full CPU core. `0` if the JVM does not support per-thread CPU accounting or on the first scrape after startup. | | `ws_pool_alloc_bytes_per_second` | gauge | Heap allocation rate, in bytes per second, summed over the proxy's client transport threads. Computed from per-thread `getThreadAllocatedBytes` deltas divided by wall dt. This is GC pressure produced by the forwarding work, not the memory it currently holds; per-thread heap residency is not measurable through standard `ThreadMXBean` APIs. `0` on JVMs without `com.sun.management.ThreadMXBean`. | All metrics carry the standard Kubernetes pod labels added by Prometheus during scraping (`pod_name`, `namespace`, `node`). Dashboard queries that need to disambiguate per-pod series rewrite `pod_name` to a `pod_short` label that drops the `hopsworks-instance-` prefix. Cluster-wide aggregates (`sum()`, `max()`) collapse pod labels and present one line per metric. ## Tuning the proxy The proxy is configured at install time through `values.yaml`: ```yaml hopsworks: payara: websocketProxy: maxSessionsPerApp: 500 # concurrent inbound WS sessions per pod incomingBufferBytes: 33554432 # max single received frame, in bytes (32 MiB) grizzlyWorkerPoolMaxSize: 200 # cap on the shared client transport workers sessionIdleTimeoutMs: 0 # 0 = no idle reaper (sessions are long-lived) heartbeatIntervalMs: 20000 # keepalive ping toward the browser; 0 disables ``` Raise `maxSessionsPerApp` to serve more concurrent notebook, terminal, and Streamlit users per pod. The cap is a direct connection count, so the value is the number of concurrent connections you want each pod to serve. There is no longer a factor-of-two thread conversion. The dominant resource cost of raising the cap is the memory held by the in-flight connections and the CPU spent forwarding their traffic, not thread overhead. `grizzlyWorkerPoolMaxSize` bounds the shared client transport worker pool that carries the upstream-to-browser leg; leave it at the default unless the CPU panel shows the transport starved. `incomingBufferBytes` is the ceiling a single received WebSocket frame (for example, a large Jupyter cell output) must fit under; it is grown on demand rather than pre-allocated, and should stay at or above Jupyter's iopub rate-limit budget. After changing the cap, watch four panels on the `Hopsworks` dashboard: - **Memory**: confirm the `Used heap` line stays comfortably below the dashed `Max heap (-Xmx)` line after each GC. - **WebSocket - pool CPU**: confirm the proxy's CPU draw stays below the pod's CPU request or limit. - **WebSocket - pool allocation rate (GC pressure)**: check whether the higher session count is producing enough allocation to drive GC time up. - **GC**: confirm GC time stays within budget after the allocation rate moves. For pod-vs-Kubernetes-limit context (working set vs request and limit) see the `Kubernetes / Pods` dashboard. ## Idle connections By default `sessionIdleTimeoutMs` is `0`, which disables the proxy's idle reaper: proxied sessions are legitimately long-lived (an open notebook can sit idle for hours), so the proxy does not close them on inactivity. Set a positive value (in milliseconds) only if you want the proxy itself to drop idle connections; JupyterLab and the terminal client both reconnect transparently after such a close, and the kernel, shell, or Streamlit process running in its own pod is unaffected by the disconnect. ## Keepalive heartbeat The connection between the browser and the proxy runs through the ingress (nginx), whose default read timeout closes a WebSocket that carries no traffic for 60 seconds. An idle terminal or notebook would otherwise be dropped and reconnected about once a minute, because nothing keeps that hop active on its own: browsers do not send WebSocket pings, and the proxy answers the upstream's pings locally rather than relaying them to the browser. To prevent this, the proxy enables the Tyrus per-session heartbeat on the browser-facing connection: it sends an unsolicited pong every `heartbeatIntervalMs` (default `20000`, i.e. 20 seconds), which keeps the hop active so the ingress does not reap an idle session. Keep the interval comfortably below the ingress read timeout; set `heartbeatIntervalMs` to `0` to disable the heartbeat if your ingress does not impose an idle timeout. ================================================================================ # Configure Authentication Source: https://docs.hopsworks.ai/latest/setup_installation/admin/auth/ # Authentication Methods ## Introduction Hopsworks can be configured to use different types of authentication methods. In this guide we will look at the different authentication methods available in Hopsworks. ## Prerequisites Administrator account on a Hopsworks cluster. ### Step 1: Go to Authentication methods page To configure Authentication methods click on your name in the top right corner of the navigation bar and choose **Cluster Settings** from the dropdown menu, then choose **Authentication** under _Security & Access_ in the left sidebar. ### Step 2: Configure Authentication methods On the **Authentication configuration** page you can configure how users authenticate. Each control is applied as soon as you change it, there is no save step. 1. **Hopsworks accounts**: when checked, users can register and log in with credentials managed by Hopsworks itself. 2. **TOTP Two-factor Authentication**: can be _disabled_, _optional_ or _mandatory_. If set to mandatory all users are required to set up two-factor authentication when registering. !!! note If two-factor is set to _mandatory_ on a cluster with preexisting users all users will need to go through lost device recovery step to enable two-factor. So consider setting it to _optional_ first and allow users to enable it before setting it to mandatory. 3. **OAuth2**: if your organization already have an identity management system compatible with [OpenID Connect (OIDC)](https://openid.net/connect/) you can configure Hopsworks to use your identity provider by ticking the **OAuth** checkbox. After enabling OAuth you can register your identity provider by clicking on **Add Identity Provider** button. See [Create client](./oauth2/create-client.md) for details. 4. **LDAP/Kerberos**: if your organization is using LDAP or Kerberos to manage users and services you can configure Hopsworks to use it as the user management system. You can enable LDAP/Kerberos by ticking the **LDAP/Kerberos** checkbox and choosing LDAP or Kerberos. For more information on how to configure LDAP and Kerberos see [Configure LDAP](./ldap/configure-ldap.md) and [Configure Kerberos](./ldap/configure-krb.md).
Authentication config
Setup Authentication Methods
In the figure above we see a cluster with Hopsworks accounts enabled, Two-factor authentication disabled, and both OAuth and LDAP/Kerberos unchecked. ================================================================================ # Register an Identity Provider Source: https://docs.hopsworks.ai/latest/setup_installation/admin/oauth2/create-client/ # Register Identity Provider in Hopsworks ## Introduction Before registering your identity provider in Hopsworks you need to create a client application in your identity provider and acquire a _client id_ and a _client secret_. An example on how to create a client using [Okta](https://www.okta.com/) and [Azure Active Directory](https://portal.azure.com/#blade/Microsoft_AAD_IAM/ActiveDirectoryMenuBlade/Overview) identity providers can be found in the following guides: [Create Okta Client](./create-okta-client.md) and [Create Azure Client](./create-azure-client.md). ## Prerequisites Acquired a _client id_ and a _client secret_ from your identity provider. ### Step 1: Register a client After acquiring the _client id_ and _client secret_ create the client in Hopsworks by [enabling OAuth2](../auth.md) and clicking on _add another identity provider_ in the [Authentication configuration page](../auth.md). Then set base uri of your identity provider in _Connection URL_ give a name to your identity provider (the name will be used in the login page as an alternative login method) and set the _client id_ and _client secret_ in their respective fields, as shown in the figure below.
Application overview
Application overview
- _Connection URL_: (provider Uri) is the base uri of the identity provider's API (URI should contain scheme http:// or https://). Additional configuration can be set here: - _Verify email_: if checked only users with verified email address (in the identity provider) can log in to Hopsworks. - _Code challenge_: if your identity provider requires code challenge for authorization request check the _code challenge_ check box. This will allow you to choose code challenge method that can be either _plain_ or _S256_. - _Logo URL_: optionally a logo URL to an image can be added. The logo will be shown on the login page with the name as shown in the figure below. - Claim names for given name, family name, email and group can also be set here. If left empty the default openid claim names will be used. ### Step 2: Add Group mappings Optionally you can add a group mapping from your identity provider to Hopsworks groups, by clicking on your name in the top right corner of the navigation bar and choosing _Cluster Settings_ from the dropdown menu. On the _Configuration_ page, found under _Infrastructure_ in the left sidebar, search for _oauth\_group\_mapping_ and click on the edit button.
Set variables
Set Configuration variables
!!! Note Setting ```oauth_group_mapping``` to ```ANY_GROUP->HOPS_USER``` will assign the role *user* to any user from any group in your identity provider when they log into Hopsworks with OAuth for the first time. You can replace *ANY_GROUP* with the group of your choice in the identity provider. You can replace *HOPS_USER* by *HOPS_ADMIN* if you want the users of that group to be admins in Hopsworks. You can do several mappings by separating them with a semicolon. Group mapping can be disabled by setting ```oauth_group_mapping_enabled=false``` in the [Configuration](../variables.md) UI. When group mapping is disabled an administrator needs to activate each user from the [User Management](../user.md) page. If group mapping is disabled then ```oauth_account_status``` in the [Configuration](../variables.md) UI should be set to 1 (Verified). Users will now see a new button on the login page. The button has the name you set above for _Name_ and will redirect to your identity provider.
OAuth2 login
Login with OAuth2
!!! note When creating a client make sure you can access the provider metadata by making a GET request on the well known endpoint of the provider. The well-known URL, will typically be the _Connection URL_ plus `.well-known/openid-configuration`. For the above client it would be `https://dev-86723251.okta.com/.well-known/openid-configuration`. ================================================================================ # Create Okta Client Source: https://docs.hopsworks.ai/latest/setup_installation/admin/oauth2/create-okta-client/ # Create An Application in Okta ## Introduction This example uses an Okta development account to create an application that will represent a Hopsworks client in the identity provider. ## Prerequisites Okta development account. To create a developer account go to [Okta developer](https://developer.okta.com/signup/). ### Step 1: Register Hopsworks as an application in your identity provider After creating a developer account register a client by going to _Applications_ and click on **Create App Integration**.
Okta Applications
Okta Applications
This will open a popup as shown in the figure below. Select **OIDC** as _Sign-in-method_ and **Web Application** as _Application type_ and click next.
Create New Application
Create new Application
Give your application a name and select **Client credential** as _Grant Type_. Then add a _Sign-in redirect URI_ that is your Hopsworks cluster domain name (including the port number if needed) with path _/callback_, and a _Sign-out redirect URI_ that is Hopsworks cluster domain name (including the port number if needed) with no path.
New Application
New Application
If you want to limit who can access your Hopsworks cluster select _Limit access to selected groups_ and select group(s) you want to give access to. Here we will allow everyone in the organization to access the cluster.
Group assignment
Group assignment
## Group mapping You can also create mappings from groups in Okta to groups in Hopsworks. To achieve this you need to configure Okta to send _Groups_ with user information. To do this go to _Applications_ and select your application name. In the _Sign On_ tab click edit _OpenID Connect ID Token_ and select **Filter** for _Groups claim type_, then for _Groups claim filter_ add **groups** as the claim name, select **Match Regex** from the dropdown and .* (dot star) as Regex to match all groups. See [Group mapping](./create-client.md#step-2-add-group-mappings) on how to do the mapping in Hopsworks.
Group claim
Group claim
### Step 2: Get the necessary fields for client registration After the application is created go back to _Applications_ and click on the application you just created. Use the _Okta domain_ (_Connection URL_), _client id_ and _client secret_ generated for your app in the [Identity Provider registration](./create-client.md) in Hopsworks.
Application overview
Application overview
!!! note When copying the domain in the figure above make sure to add the url scheme (http:// or https://) when using it in the _Connection URL_ in the [Identity Provider registration form](./create-client.md). ================================================================================ # Create Azure Client Source: https://docs.hopsworks.ai/latest/setup_installation/admin/oauth2/create-azure-client/ # Create An Application in Azure Active Directory ## Introduction This example uses Azure Active Directory as the identity provider, but the same can be done with any identity provider supporting OAuth2 OpenID Connect protocol. ## Prerequisites Azure account. ### Step 1: Register Hopsworks as an application in your identity provider To use OAuth2 in Hopsworks you first need to create and configure an OAuth client in your identity provider. We will take the example of Azure AD for the remaining of this documentation, but equivalent steps can be taken on other identity providers. Navigate to the [Microsoft Azure Portal](https://portal.azure.com) and authenticate. Navigate to [Azure Active Directory](https://portal.azure.com/#blade/Microsoft_AAD_IAM/ActiveDirectoryMenuBlade/Overview). Click on [App Registrations](https://portal.azure.com/#blade/Microsoft_AAD_IAM/ActiveDirectoryMenuBlade/RegisteredApps). Click on *New Registration*.

Create application
Create application

Enter a name for the client such as *hopsworks_oauth_client*. Verify the Supported account type is set to *Accounts in this organizational directory only*. Click Register.

Name application
Name application

### Step 2: Get the necessary fields for client registration In the Overview section, copy the *Application (client) ID field*. We will use it in [Identity Provider registration](./create-client.md) under the name *Client id*.

Copy client ID
Copy client ID

Click on *Endpoints* and copy the *OpenId Connect metadata document* endpoint excluding the *.well-known/openid-configuration* part. We will use it in [Identity Provider registration](./create-client.md) under the name *Connection URL*.

Endpoint
Endpoint

!!! note If you have multiple tenants in your Azure Active Directory, the `OpenID Connect metadata document` endpoint might use `organizations` instead of a specific tenant ID. In such cases, replace `organizations` with your actual tenant ID to target a specific directory. example: ``` https://login.microsoftonline.com/organizations/oauth2/v2.0 --> https://login.microsoftonline.com//oauth2/v2.0 ``` Click on *Certificates & secrets*, then Click on *New client secret*.

New client secret
New client secret

Add a *description* of the secret. Select an expiration period. Click *Add*.

Client secret creation
Client secret creation

Copy the secret. This will be used in [Identity Provider registration](./create-client.md) under the name *Client Secret*.

Client secret creation
Client secret creation

Click on *Authentication*. Then click on *Add a platform*.

Add a platform
Add a platform

In *Configure platforms* click on *Web*.

Configure platform: Web
Configure platform: Web

Enter the *Redirect URI* and click on *Configure*. The redirect URI is *HOPSWORKS-URI/callback* with *HOPSWORKS-URI* the URI of your Hopsworks cluster.

Configure platform: Redirect
Configure platform: Redirect

================================================================================ # Configure Project Mapping Source: https://docs.hopsworks.ai/latest/setup_installation/admin/oauth2/configure-project-mapping/ # Configure OAuth2 group to project mapping ## Introduction A group-to-project mapping lets you automatically add all members of an OAuth2 group to a project, eliminating the need to add each user individually. To create a mapping, you simply select an OAuth2 group, choose the project it should be linked to, and assign the role that its members will have within that project. Once a mapping is created, project membership is controlled through OAuth2 group membership. Any updates made to the OAuth2 group, such as adding or removing users, will automatically be reflected in Hopsworks. For example, if a user is removed from the OAuth2 group, they will also be removed from the corresponding project. ## Prerequisites 1. A server configured with OAuth2. See [Register Identity Provider in Hopsworks](./create-client.md) for instructions on how to do this. 2. OAuth2 group mapping sync enabled. This can be done by setting the variable ```oauth_group_mapping_sync_enabled=true```. See [Cluster Configuration](../variables.md) on how to change variable values in Hopsworks.
Enable OAuth2 mapping
Enable OAuth2 mapping
If you can not find the variable ```oauth_group_mapping_sync_enabled``` create it by clicking on **New variable**.
Create OAuth2 mapping enabled variable
Create OAuth2 mapping enabled variable
### Step 1: Create a mapping To create a mapping go to **Cluster Settings** by clicking on your name in the top right corner of the navigation bar and choosing *Cluster Settings* from the dropdown menu. Choose *Project Mapping* under *Security & Access* in the left sidebar, then create a new mapping by clicking on *Create new mapping*.
Project mapping tab
Project mapping
This will take you to the create mapping page shown below
Create mapping
Create mapping
Here you can enter your OAuth2 group and map it to a project from the *Project* drop down list. You can also choose the *Project role* users will be assigned when they are added to the project. Finally, click on *Create mapping* and go back to mappings. You should see the newly created mapping(s) as shown below.
Project mappings
Project mappings
!!!Note Make sure the group names from your OAuth2 provider match the one you entered above. If your identity provider uses a claim name other than ```groups``` or ```roles``` to represent group information, be sure to specify that claim name in the **Group Claim** field when setting up your identity provider. ### Step 2: Edit a mapping From the list of mappings click on the edit button (:material-pencil:). This will open a popup that will allow you to change the *remote group*, *project name*, and *project role* of a mapping.
Edit mapping
Edit mapping
!!!Warning Updating a mapping's *remote group* or *project name* will remove all members of the previous group from the project. ### Step 3: Delete a mapping To delete a mapping click on the delete button. !!!Warning Deleting a mapping will remove all members of that group from the project. ================================================================================ # Configure LDAP Source: https://docs.hopsworks.ai/latest/setup_installation/admin/ldap/configure-ldap/ # Configure LDAP/Kerberos ## Introduction LDAP (Lightweight Directory Access Protocol) is a software protocol for enabling anyone in a network to gain access to resources such as files and devices. This tutorial shows an administrator how to configure LDAP authentication. LDAP need some server configuration before you can enable it from the UI. ## Prerequisites A server configured with LDAP. See [Server Configuration for LDAP](./configure-server.md#step-1-server-configuration-for-ldap) for instruction on how to do this. ### Step 1: Enable LDAP After configuring the server you can configure Authentication methods by clicking on your name in the top right corner of the navigation bar and choosing *Cluster Settings* from the dropdown menu. On the *Authentication* page, found under *Security & Access* in the left sidebar of **Cluster Settings**, you can enable LDAP by clicking on the LDAP checkbox. If LDAP/Kerberos checkbox is not checked make sure that you configured your application server and enable it by clicking on the checkbox.
Authentication config
Setup Authentication Methods
### Step 2: Edit configuration Finally, click on edit configuration and fill in the attributes.
LDAP config
Configure LDAP
- Account status: the status a user will be assigned when logging in for the first time. If a user is assigned a status different from *Activated* an admin needs to manually activate each user from the [User management](../user.md). - Group mapping: allows you to specify a mapping between LDAP groups and Hopsworks groups. The mapping is a semicolon separated string in the form ```Directory Administrators->HOPS_ADMIN;IT People-> HOPS_USER```. Default is empty. If no mapping is specified, users need to be assigned a role by an admin before they can log in. - User id: the id field in LDAP with a string placeholder. Default ```uid=%s```. - User given name: the given name field in LDAP. Default ```givenName```. - User surname: the surname field in LDAP. Default ```sn```. - User email: the email field in LDAP. Default ```mail```. - User search filter: the search filter for user. Default ```uid=%s```. - Group search filter: the search filter for groups. Default ```member=%d```. - Group target: the target to search for groups in the LDAP directory tree. Default ```cn```. - Dynamic group target: the target to search for dynamic groups in the LDAP directory tree. Default ```memberOf```. - User dn: specify the distinguished name (DN) of the container or base point where the users are stored. Default is empty. - Group dn: specify the DN of the container or base point where the groups are stored. Default is empty. All defaults are taken from [OpenLDAP](https://www.openldap.org/). The login page will now have the choice to use LDAP for authentication.
Log in using LDAP
Log in using LDAP
!!! note Group mapping can be disabled by setting ```ldap_group_mapping_enabled=false``` in the [Configuration](../variables.md) UI. When group mapping is disabled an administrator needs to activate each user from the [User Management](../user.md) page. If group mapping is disabled then Account status in LDAP configuration above should be set to ```Verified```. ================================================================================ # Configure Kerberos Source: https://docs.hopsworks.ai/latest/setup_installation/admin/ldap/configure-krb/ # Configure Kerberos ## Introduction Kerberos is a network authentication protocol that allow nodes to communicating over a non-secure network to prove their identity to one another in a secure manner. This tutorial shows an administrator how to configure Kerberos authentication. Kerberos need some server configuration before you can enable it from the UI. ## Prerequisites A server configured with Kerberos. See [Server Configuration for Kerberos](./configure-server.md#step-2-server-configuration-for-kerberos) for instruction on how to do this. ### Step 1: Enable Kerberos After configuring the server you can configure Authentication methods by clicking on your name in the top right corner of the navigation bar and choosing *Cluster Settings* from the dropdown menu. On the *Authentication* page, found under *Security & Access* in the left sidebar of **Cluster Settings**, you can enable Kerberos by clicking on the Kerberos checkbox. If LDAP/Kerberos checkbox is not checked, make sure that you configured your application server and enable it by clicking on the checkbox.
Authentication config
Setup Authentication Methods
### Step 2: Edit configuration Finally, click on edit configuration and fill in the attributes.
Kerberos config
Configure Kerberos
- Account status: the status a user will be assigned when logging in for the first time. If a user is assigned a status different from *Activated* an admin needs to manually activate each user from the [User management](../user.md). - Group mapping: allows you to specify a mapping between LDAP groups and Hopsworks groups. The mapping is a semicolon separated string in the form ```Directory Administrators->HOPS_ADMIN;IT People-> HOPS_USER```. Default is empty. If no mapping is specified, users need to be assigned a role by an admin before they can log in. - User id: the id field in LDAP with a string placeholder. Default ```uid=%s```. - User given name: the given name field in LDAP. Default ```givenName```. - User surname: the surname field in LDAP. Default ```sn```. - User email: the email field in LDAP. Default ```mail```. - User search filter: the search filter for user. Default ```uid=%s```. - Principal search filter: the search filter for principal name. Default ```krbPrincipalName=%s```. - Group search filter: the search filter for groups. Default ```member=%d```. - Group target: the target to search for groups in the LDAP directory tree. Default ```cn```. - Dynamic group target: the target to search for dynamic groups in the LDAP directory tree. Default ```memberOf```. - User dn: specify the distinguished name (DN) of the container or base point where the users are stored. Default is empty. - Group dn: specify the DN of the container or base point where the groups are stored. Default is empty. All defaults are taken from [OpenLDAP](https://www.openldap.org/). !!! note Group mapping can be disabled by setting ```ldap_group_mapping_enabled=false``` in the [Configuration](../variables.md) UI. When group mapping is disabled an administrator needs to activate each user from the [User Management](../user.md) page. If group mapping is disabled then Account status in the Kerberos configuration above should be set to ```Verified```. The login page will now have the choice to use Kerberos for authentication.
Log in using Kerberos
Log in using Kerberos
!!!Note Make sure that you have Kerberos properly configured on your computer and you are logged in. Kerberos support must also be configured on the browser to use Kerberos for authentication. ================================================================================ # Configure server for LDAP and Kerberos Source: https://docs.hopsworks.ai/latest/setup_installation/admin/ldap/configure-server/ # Configure Server for LDAP and Kerberos ## Introduction LDAP and Kerberos integration need some configuration in the helm charts for your cluster definition used to deploy your Hopsworks cluster. This tutorial shows an administrator how to configure the application server for LDAP and Kerberos integration. ## Prerequisites An accessible LDAP domain. A Kerberos Key Distribution Center (KDC) running on the same domain as Hopsworks (Only for Kerberos). ### Step 1: Server Configuration for LDAP The LDAP attributes below are used to configure JNDI external resource in Payara. The JNDI resource will communicate with your LDAP server to perform the authentication. ```yaml ldap: enabled: true jndilookupname: "dc=hopsworks,dc=ai" provider_url: "ldap://193.10.66.104:1389" attr_binary_val: "entryUUID" security_auth: "none" security_principal: "" security_credentials: "" referral: "ignore" additional_props: "" ``` - jndilookupname: should contain the LDAP domain. - attr_binary_val: is the binary unique identifier that will be used in subsequent logins to identify the user. - security_auth: how to authenticate to the LDAP server. - security_principal: contains the username of the user that will be used to query LDAP. - security_credentials: contains the password of the user that will be used to query LDAP. - referral: whether to follow or ignore an alternate location in which an LDAP Request may be processed. An already deployed instance can be configured to connect to LDAP. Go to the payara admin UI and create a new JNDI external resource. The name of the resource should be __ldap/LdapResource__.
LDAP Resource
LDAP Resource
This can also be achieved by running the below asadmin command. ```bash asadmin create-jndi-resource \ --restype javax.naming.ldap.LdapContext \ --factoryclass com.sun.jndi.ldap.LdapCtxFactory \ --jndilookupname dc\=hopsworks\,dc\=ai \ --property java.naming.provider.url=ldap\\://193\.10\.66\.104\\:1389:\ hopsworks.ldap.basedn=dc\\\=hopsworks\,dc\\\=ai:\ java.naming.ldap.attributes.binary=entryUUID:\ java.naming.security.authentication=simple:\ java.naming.security.principal=:\ java.naming.security.credentials=:\ java.naming.referral=ignore \ ldap/LdapResource ``` ### Step 2: Server Configuration for Kerberos The Kerberos attributes are used to configure [SPNEGO](https://spnego.sourceforge.net/). SPNEGO is used to establish a secure context between the requester and the application server when using Kerberos authentication. ```yaml kerberos: enabled: true krb_conf_path: "/etc/krb5.conf" krb_server_key_tab_path: "/etc/security/keytabs/service.keytab" krb_server_key_tab_name: "service.keytab" spnego_server_conf: '\nuseKeyTab=true\nprincipal=\"HTTP/server.hopsworks.ai@HOPSWORKS.AI\"\nstoreKey=true\nisInitiator=false' ldap: jndilookupname: "dc=hopsworks,dc=ai" provider_url: "ldap://193.10.66.104:1389" attr_binary_val: "objectGUID" security_auth: "none" security_principal: "" security_credentials: "" referral: "ignore" additional_props: "" ``` Both Kerberos and LDAP attributes need to be specified to configure Kerberos. The LDAP attributes are explained above. - krb_conf_path: contains the path to the krb5.conf used by SPNEGO to get information about the default domain and the location of the Kerberos KDC. The file is copied by the recipe in to /srv/hops/domains/domain1/config. - krb_server_key_tab_path: contains the path to the Kerberos service keytab. The keytab is copied by the recipe in to /srv/hops/domains/domain/config with the name set in the __krb_server_key_tab_name__ attribute. - spnego_server_conf: contains the configuration that will be appended to Payara's (application serve used to host hopsworks) login.conf. In particular, it should contain useKeyTab=true, and the principal name to be used in the authentication phase. Initiator should be set to false. ================================================================================ # Configure Project Mapping Source: https://docs.hopsworks.ai/latest/setup_installation/admin/ldap/configure-project-mapping/ # Configure LDAP/Kerberos group to project mapping ## Introduction A group-to-project mapping lets you automatically add all members of an LDAP group to a project, eliminating the need to add each user individually. To create a mapping, you simply select the LDAP group, choose the project it should be linked to, and assign the role that its members will have within that project. Once a mapping is created, project membership is controlled through LDAP group membership. Any updates made to the LDAP group, such as adding or removing users, will automatically be reflected in Hopsworks. For example, if a user is removed from the LDAP group, they will also be removed from the corresponding project. ## Prerequisites 1. A server configured with LDAP or Kerberos. See [Server Configuration for Kerberos](./configure-server.md#step-2-server-configuration-for-kerberos) and [Server Configuration for LDAP](./configure-server.md#step-1-server-configuration-for-ldap) for instructions on how to do this. 2. LDAP group mapping sync enabled. This can be done by setting the variable ```ldap_group_mapping_sync_enabled=true```. See [Cluster Configuration](../variables.md) on how to change variable values in Hopsworks.
Enable ldap mapping
Enable ldap mapping
If you can not find the variable ```ldap_group_mapping_sync_enabled``` create it by clicking on **New variable**.
Create ldap mapping enabled variable
Create ldap mapping enabled variable
### Step 1: Create a mapping To create a mapping go to **Cluster Settings** by clicking on your name in the top right corner of the navigation bar and choosing *Cluster Settings* from the dropdown menu. Choose *Project Mapping* under *Security & Access* in the left sidebar, then create a new mapping by clicking on *Create new mapping*.
Project mapping tab
Project mapping
This will take you to the create mapping page shown below
Create mapping
Create mapping
Here you can choose from your LDAP groups and map them to a project from the *Project* drop down list. You can also choose the *Project role* users will be assigned when they are added to the project. Finally, click on *Create mapping* and go back to mappings. You should see the newly created mapping(s) as shown below.
Project mappings
Project mappings
!!!Note If there are no groups in the *Remote group* drop down list check if **ldap_groups_search_filter** is correct by using the value in ```ldapsearch``` replacing ```%c``` with ```*```, as shown in the example below. ```bash ldapsearch -LLL -H ldap:/// -b '' -D '' -w '(&(objectClass=groupOfNames)(cn=*))' ``` This should return all the groups in your LDAP. See [Cluster Configuration](../variables.md) on how to find and update the value of this variable. ### Step 2: Edit a mapping From the list of mappings click on the edit button (:material-pencil:). This will open a popup that will allow you to change the *remote group*, *project name*, and *project role* of a mapping.
Edit mapping
Edit mapping
!!! Warning Updating a mapping's *remote group* or *project name* will remove all members of the previous group from the project. ### Step 3: Delete a mapping To delete a mapping click on the delete button. !!! Warning Deleting a mapping will remove all members of that group from the project. ### Step 4: Configure sync interval After configuring all the group mappings users will be added to or removed from the projects in the mapping when they login to Hopsworks. It is also possible to synchronize mappings without requiring users to log out. This can be done by setting ```ldap_group_mapping_sync_interval``` to an interval greater or equal to 2 minutes. If ```ldap_group_mapping_sync_interval``` is set group mapping sync will run periodically based on the interval and add or remove users from projects. ================================================================================ # Overview Source: https://docs.hopsworks.ai/latest/setup_installation/admin/ha-dr/intro/ # Hopsworks High Availability and Disaster Recovery Documentation The Hopsworks Feature Store is the underlying component powering enterprise ML pipelines as well as serving feature data to model making user facing predictions. Sometimes the Hopsworks cluster can experience hardware failures or power loss, to help you plan for these occasions and avoid Hopsworks Feature Store downtime, we put together this guide. This guide is divided into three sections: - **High availability**: deployment patterns and best practices to make sure individual component failures do not impact the availability of the Hopsworks cluster. - **Backup**: configuration policies and best practices to make sure you have have fresh copy of the data and metadata in case of necessity - **Restore**: procedures and best practices to restore a previous backup if needed. ================================================================================ # High Availability Source: https://docs.hopsworks.ai/latest/setup_installation/admin/ha-dr/ha/ # High Availability At a high level a Hopsworks cluster can be divided into 4 groups of nodes. Each node group should be deployed according to the requirements (e.g., 3/5/7 nodes for the head node group) to guarantee the availability of the components. - **Head nodes**: The head node is responsible for running all the metadata, public API, and user interface services that are required for Hopsworks to provide its functionality. They need to be deployed in an odd number (1, 3, 5) as the head nodes run services like Zookeeper and OpenSearch which enforce consistency through quorum based protocols. The head nodes are also responsible for managing the services running on the remaining group of nodes. - **Worker nodes**: The worker node is responsible for executing the feature engineering pipeline code as well as storing the data for the offline feature store (HopsFS). In an on-prem deployment, the data is stored and replicated on the workers’ local hard drives. By default the data is replicated across 3 workers. In a cloud deployment, HopsFS’ data is persisted in a cloud object store (Amazon S3, Azure Blob Storage, Google Cloud Blob Storage) and the HopsFS datanodes are responsible for persisting, retrieving and caching of blocks from the object store. - **RonDB Data nodes**: These nodes are responsible for storing the services’ metadata (Hopsworks, HopsFS, Hive Metastore, Airflow) as well as the data for the online feature store. For high availability, at least two data nodes should be deployed and RonDB is typically configured with a replication factor of 2, as it uses synchronous replication with 2-phase commit, not a quorum-based replication protocol. More advanced deployment patterns and best practices are covered in the [RonDB documentation](https://docs.rondb.com). - **Query brokers**: The query brokers are the entry point for querying the online feature store. They handle authentication, authorization and execution of the requests for online feature data being submitted from the feature store APIs. At least two query brokers should be deployed to achieve high availability. Query brokers are stateless. Additional query brokers should be deployed to handle additional load and clients. Example deployment:
Example HA deployment
Example High Available deployment
For higher availability, a Hopsworks cluster should be deployed across multiple availability zones, however, a single cluster cannot be deployed across multiple regions. Multiple region deployments are out of the scope of this guide. A different service placement is also possible, e.g., separating RonDB data nodes between metadata and online feature store or adding more replicas of a metadata service without necessarily adding a whole new head node, however, this is outside the scope of this guide. ================================================================================ # Disaster Recovery Source: https://docs.hopsworks.ai/latest/setup_installation/admin/ha-dr/dr/ # Disaster Recovery ## Backup The state of a Hopsworks cluster is split between data and metadata and distributed across multiple services. This section explains how to take consistent backups for the offline and online feature stores as well as cluster metadata. In Hopsworks, a consistent backup should back up the following services: - **RonDB**: cluster metadata and the online feature store data. - **HopsFS**: offline feature store data plus checkpoints and logs for feature engineering applications. - **Opensearch**: search metadata, logs, dashboards, and user embeddings. - **Superset**: dashboards, charts, saved queries, database connections, and users and roles, stored in the Superset metadata database, which is a separate MySQL from RonDB. - **Kubernetes objects**: cluster credentials, backup metadata, serving metadata, Trino authentication (password and group files plus the admin and monitoring credentials), and project namespaces with service accounts, roles, secrets, and configmaps. - **Python environments**: custom project environments are stored in your configured container registry. Back up the registry separately. If a project and its environment are deleted, you must recreate the environment after restore. Besides the above services, Hopsworks uses also Apache Kafka which carries in-flight data heading to the online feature store. In the event of a total cluster loss, running jobs with in-flight data must be replayed. ### Prerequisites When enabling backup in Hopsworks, cron jobs are configured for RonDB and Opensearch. For HopsFS, backups rely on versioning in the object store. For Kubernetes objects, Hopsworks uses Velero to snapshot the required resources. Before enabling backups: - Enable versioning on the S3-compatible bucket used for HopsFS. - Install and configure Velero with the AWS plugin (S3). #### Install Velero Velero provides backup and restore for Kubernetes resources. Install it with either the Velero CLI or Helm (Velero docs: [Velero basic install guide](https://velero.io/docs/v1.17/basic-install/)). - Using the Velero CLI, set up the CRDs and deployment: ```bash velero install \ --image velero/velero:v1.17.1 \ --plugins velero/velero-plugin-for-aws:v1.13.0 \ --no-default-backup-location \ --no-secret \ --use-volume-snapshots=false \ --wait ``` - Using the Velero Helm chart: ```bash helm repo add vmware-tanzu https://vmware-tanzu.github.io/helm-charts helm repo update helm install velero vmware-tanzu/velero \ --namespace velero \ --version 11.2.0 \ --create-namespace \ --set "initContainers[0].name=velero-plugin-for-aws" \ --set "initContainers[0].image=velero/velero-plugin-for-aws:v1.13.0" \ --set "initContainers[0].volumeMounts[0].mountPath=/target" \ --set "initContainers[0].volumeMounts[0].name=plugins" \ --set-json configuration.backupStorageLocation='[]' \ --set "credentials.useSecret=false" \ --set "snapshotsEnabled=false" \ --wait ``` ### Configuring Backup !!! Note Backup is only supported for clusters that use S3-compatible object storage. You can enable backups during installation or a later upgrade. Set the schedule with a cron expression in the values file: ```yaml global: _hopsworks: backups: enabled: true schedule: "@weekly" ``` After configuring backups, go to the cluster settings and choose Backup under Infrastructure in the left sidebar. You should see `enabled` at the top level and for all services if everything is configured correctly.
Backup overview page
Backup overview page
If any service is misconfigured, the backup status shows as `partial`. In the example below, Velero is disabled because it was not configured correctly. Fix partial backups before relying on them for recovery.
Backup overview page (partial setup)
Backup overview page (partial setup)
#### Cleanup Use the backup time-to-live (`ttl`) flag to automatically prune backups older than the configured duration. ```yaml global: _hopsworks: backups: enabled: true schedule: "@weekly" ttl: 60d ``` For S3 object storage, you can also configure a bucket lifecycle policy to expire old object versions. Example for AWS S3: ```json { "Rules": [ { "ID": "HopsFSBlocksRetentionPolicy", "Status": "Enabled", "Filter": {}, "Expiration": { "ExpiredObjectDeleteMarker": true }, "NoncurrentVersionExpiration": { "NoncurrentDays": 60 } } ] } ``` ### Superset { #superset-backup } Superset stores its state in its own MySQL database, which is separate from RonDB and is therefore not part of the RonDB backup. When backups are enabled, a `create-superset-backup` cron job takes a logical dump of the Superset database and uploads it to the same object storage as the other backups, under the `superset_backup//` prefix. The dump covers the whole `superset` schema, so it includes the database connections (connectors) along with dashboards, charts, saved queries, users and roles. That includes the connections Hopsworks creates itself, such as the per-project Trino connections, because they are rows in the same database. Connector credentials are stored encrypted with the Superset secret key, so they are only usable after a restore if that key is restored with the database, which is why the restore verifies it (see [Superset restore][superset-restore]). Each backup writes two objects: `superset.sql.gz` (the gzipped dump) and `manifest.json` (the dump checksum, the Superset image and schema version, and a fingerprint of the Superset secret key). The backup is also indexed in the `superset-backups-metadata` ConfigMap, which the Velero backup captures so the index is restored with the cluster. Superset's Kubernetes Secrets are captured by the Velero backup through the `backup.hops.works/include` label: - `superset-secret-key`: the Superset secret key. It must be restored together with the database, because Superset uses it to encrypt the database-connection passwords stored in the metadata database, so restoring the database with a different secret key leaves those connections undecryptable. - `superset-mysql-users-secrets`: the Superset MySQL credentials. - `superset-admin-credentials`: the Superset admin account. If you provide these Secrets yourself by setting `superset.auth.createSecrets: false`, you must add the label `backup.hops.works/include: "true"` to each of them, because the chart only labels the Secrets it creates and Velero selects Secrets by that label. To list the Superset backups that were taken, read the metadata ConfigMap: ```bash kubectl get configmap superset-backups-metadata -n hopsworks -o json \ | jq -r '.data | to_entries[] | select(.value | fromjson | .state == "SUCCESS") | .key' \ | sort -r ``` !!! note Backups taken before Superset backup was enabled do not contain Superset. Restoring from such a backup recovers the rest of the cluster but not Superset dashboards, charts, or users. ### Trino authentication Trino authenticates users against a password file and authorizes them against a group file. Both files, together with the Trino admin and monitoring credentials, live in Kubernetes Secrets that are labelled so the `k8s-backups-main` Velero schedule captures them. They are therefore included in the Kubernetes-objects backup and no extra configuration is required. After a restore, Hopsworks rebuilds the Trino password and group files from the restored RonDB state. The rebuild runs on the primary backend instance shortly after startup and then periodically, so an existing project user can authenticate to Trino and holds exactly the access their restored project role grants. This means a restore recovers Trino authentication automatically, without an operator step. !!! note This holds for backups taken after this feature was deployed. A backup taken before it does not contain the Trino Secrets, because they were not yet labelled for the `k8s-backups-main` schedule, and an existing cluster first has such a backup only after its next backup cycle. Restoring from an older backup still recovers all RonDB-derived authentication, because the reconciliation rebuilds the project users, their role groups, and the shared-dataset groups from the restored RonDB state. It does not recover the admin and monitoring credential entries: those are re-rendered from the chart values on a fresh install, so any out-of-band rotation of them is lost and Prometheus may need its Trino monitoring credentials re-aligned. The Trino internal shared secret is deliberately not backed up. It has no coupling to user state and is regenerated on a fresh install, and restoring an old value onto a running cluster would break Trino internal communication until every Trino pod restarts. If you set `createSharedSecret: false` and manage this secret yourself, restore it from your own source and then restart all Trino pods together so the coordinator and workers share the same value. !!! warning The backup contains the Trino admin credentials, which are only base64-encoded at the Kubernetes layer. Restrict access to the backup object store and enable encryption at rest so the backup does not expose usable credentials. !!! note On ArgoCD-managed clusters with automated self-healing, add `ignoreDifferences` on the `data` field of the four Trino auth Secrets (`trino-password-file`, `trino-groups-file`, `trino-admin-credentials`, `trino-monitoring-credentials`) and set `RespectIgnoreDifferences=true` in the Application sync options. Hopsworks writes project users and groups into these Secrets at runtime, so without this ArgoCD reapplies the chart defaults on every sync and overwrites the restored authentication data. ## Restore Hopsworks supports two restore modes: - **New cluster restore**: Install a fresh cluster and restore data from a backup during installation. - **In-place restore**: Restore data onto an existing running cluster via `helm upgrade`. !!! Note Use the exact Hopsworks version that was used to create the backup. ### New Cluster Restore The new cluster restore process has two phases: - Restore Kubernetes objects required for the cluster restore. - Install the cluster with Helm using the correct backup IDs. #### Restore Kubernetes objects Restore the Kubernetes objects that were backed up using Velero. - Ensure that Velero is installed and configured with the AWS plugin as described in the [prerequisites](#prerequisites). - Set up a [Velero backup storage location](https://velero.io/docs/v1.17/api-types/backupstoragelocation/) to point to the S3 bucket. - If you are using AWS S3 and access is controlled by an IAM role: ```bash kubectl apply -f - < hopsworks-bsl-credentials [default] aws_access_key_id=YOUR_ACCESS_KEY aws_secret_access_key=YOUR_SECRET_KEY EOF kubectl create secret generic -n velero hopsworks-bsl-credentials --from-file=cloud=hopsworks-bsl-credentials kubectl apply -f - </dev/null)" = "Available" ]; do echo "Still waiting..."; sleep 5; done echo "=== Waiting for Velero to sync the backups from hopsworks-bsl ===" until [ "$(kubectl get backups -n velero -ojson | jq -r '[.items[] | select(.spec.storageLocation == "hopsworks-bsl")] | length' 2>/dev/null)" != "0" ]; do echo "Still waiting..."; sleep 5; done # Restores the latest - if specific backup is needed then backupName instead echo "=== Creating Velero Restore object for k8s-backups-main ===" kubectl apply -f - </dev/null)" = "Completed" ]; do echo "Still waiting..."; sleep 5; done # Restores the latest - if specific backup is needed then backupName instead echo "=== Creating Velero Restore object for k8s-backups-users-resources ===" kubectl apply -f - </dev/null)" = "Completed" ]; do echo "Still waiting..."; sleep 5; done ``` After the restore completes, verify the restored resources in Kubernetes. RonDB and Opensearch store their backup metadata in the `rondb-backups-metadata` and `opensearch-backups-metadata` configmaps. Use the commands below to list successful backup IDs (newest first) that can be referenced during cluster installation. ```bash kubectl get configmap rondb-backups-metadata -n hopsworks -o json \ | jq -r '.data | to_entries[] | select(.value | fromjson | .state == "SUCCESS") | .key' \ | sort -nr kubectl get configmap opensearch-backups-metadata -n hopsworks -o json \ | jq -r '.data | to_entries[] | select(.value | fromjson | .state == "SUCCESS") | .key' \ | sort -nr ``` #### Restore on Cluster installation To restore a cluster during installation, configure the backup ID in the values YAML file: ```yaml global: _hopsworks: backups: enabled: true schedule: "@weekly" restoreFromBackup: backupId: "254811200" ``` ##### Customizations !!! Warning Even if you override the backup IDs for RonDB and Opensearch, you must still set `.global._hopsworks.restoreFromBackup.backupId` to ensure HopsFS is restored. To restore a different backup ID for RonDB: ```yaml global: _hopsworks: backups: enabled: true schedule: "@weekly" restoreFromBackup: backupId: "254811200" rondb: rondb: restoreFromBackup: backupId: "254811140" ``` To restore a different backup for Opensearch: ```yaml global: _hopsworks: backups: enabled: true schedule: "@weekly" restoreFromBackup: backupId: "254811200" olk: opensearch: restore: repositories: default: snapshots: default: snapshot_name: "254811140" ``` You can also customize the Opensearch restore process to skip specific indices: ```yaml global: _hopsworks: backups: enabled: true schedule: "@weekly" restoreFromBackup: backupId: "254811200" olk: opensearch: restore: repositories: default: snapshots: default: snapshot_name: "254811140" payload: indices: "-myindex" ``` ### In-Place Restore !!! Note In-place restore is available from Hopsworks version 4.8.0. In-place restore allows you to restore data onto an existing running cluster using `helm upgrade`. Unlike a new cluster restore, this does not require provisioning a fresh cluster: the existing stateful services are shut down, wiped if necessary, and restored from backup. !!! Warning In-place restore **replaces all existing data** in the cluster with the backup data. Any data written after the backup was taken will be lost. !!! Info After a fresh install from backup (new cluster restore), in-place restores can only be performed using backups taken **after** that fresh install, because the cluster certificates are regenerated during installation. To restore to a backup that was taken **before** the fresh install, you must perform another new cluster restore from that backup instead of an in-place restore. #### In-place restore prerequisites - A running Hopsworks cluster deployed via Helm. - A previously created backup with a known backup ID. - Object storage configured and accessible with the backup data. - Velero installed and configured as described in the [prerequisites](#prerequisites). #### Identify the backup ID Get the backup ID from the **Cluster Settings > Backup** page or by using the following commands. ```bash # RonDB backup IDs (newest first) kubectl get configmap rondb-backups-metadata -n hopsworks -o json \ | jq -r '.data | to_entries[] | select(.value | fromjson | .state == "SUCCESS") | .key' \ | sort -nr # Opensearch backup IDs (newest first) kubectl get configmap opensearch-backups-metadata -n hopsworks -o json \ | jq -r '.data | to_entries[] | select(.value | fromjson | .state == "SUCCESS") | .key' \ | sort -nr # Velero backup IDs for the main schedule (newest first) kubectl get backups -n velero -o json \ | jq -r '[.items[] | select(.spec.storageLocation == "hopsworks-bsl" and .metadata.labels["velero.io/schedule-name"] == "k8s-backups-main" and .status.phase == "Completed")] | sort_by(.status.completionTimestamp) | reverse[] | .metadata.name' # Velero backup IDs for the users schedule (newest first) kubectl get backups -n velero -o json \ | jq -r '[.items[] | select(.spec.storageLocation == "hopsworks-bsl" and .metadata.labels["velero.io/schedule-name"] == "k8s-backups-users-resources" and .status.phase == "Completed")] | sort_by(.status.completionTimestamp) | reverse[] | .metadata.name' ``` #### Run the in-place restore Configure the restore in the values file and run `helm upgrade`: ```yaml global: _hopsworks: backups: enabled: true schedule: "@weekly" restoreFromBackup: backupId: "254811200" inPlace: true forceDataClear: true # Optional: specify Velero backup IDs. If not set, the latest completed backup is used. hopsworks: velero: restore: mainScheduleBackupId: "k8s-backups-main-20260213T153627Z" usersScheduleBackupId: "k8s-backups-users-resources-20260213T153627Z" ``` Then run: ```bash helm upgrade hopsworks hopsworks/hopsworks --version \ --namespace hopsworks \ -f values.yaml \ --timeout 1200s ``` You can also pass the restore flags directly on the command line: ```bash helm upgrade hopsworks hopsworks/hopsworks --version \ --namespace hopsworks \ --set-string global._hopsworks.restoreFromBackup.backupId="254811200" \ --set global._hopsworks.restoreFromBackup.inPlace=true \ --set global._hopsworks.restoreFromBackup.forceDataClear=true \ --set-string hopsworks.velero.restore.mainScheduleBackupId="k8s-backups-main-20260213T153627Z" \ --set-string hopsworks.velero.restore.usersScheduleBackupId="k8s-backups-users-resources-20260213T153627Z" \ --timeout 1200s ``` The required flags are: | Parameter | Description | | --------- | ----------- | | `global._hopsworks.restoreFromBackup.backupId` | The backup ID to restore from. | | `global._hopsworks.restoreFromBackup.inPlace` | Must be `true` to enable in-place restore mode. | | `global._hopsworks.restoreFromBackup.forceDataClear` | Must be `true` to confirm that existing data will be replaced. This is a safety mechanism to prevent accidental data loss. | The following flags are optional. If not set, the latest available Velero backup will be used: | Parameter | Description | | --------- | ----------- | | `hopsworks.velero.restore.mainScheduleBackupId` | The Velero backup ID for the main schedule (`k8s-backups-main`). | | `hopsworks.velero.restore.usersScheduleBackupId` | The Velero backup ID for the users schedule (`k8s-backups-users-resources`). | !!! Important After a successful restore, remove the `restoreFromBackup` blocks from your values file and run `helm upgrade` to apply the change. If left in place, these blocks can cause subsequent upgrades to fail or behave unexpectedly. #### Re-running an in-place restore In-place restore creates marker resources to prevent accidental re-runs. If you need to run the restore again with the same backup ID, delete the marker resources first: ```bash # Delete the HopsFS restore job kubectl delete job hopsfs-inplace-restore- -n hopsworks --ignore-not-found=true # Delete the RonDB restore jobs kubectl delete job restore-native-backup- -n hopsworks --ignore-not-found=true kubectl delete job setup-mysqld-dont-remove- -n hopsworks --ignore-not-found=true # Delete the Opensearch restore job kubectl delete job opensearch-restore-default-default- -n hopsworks --ignore-not-found=true # Delete the velero restore objects, use the exact backup name or schedule name kubectl delete restore.velero.io k8s-backups-main -n velero --ignore-not-found=true kubectl delete restore.velero.io k8s-backups-users-resources -n velero --ignore-not-found=true ``` #### In-place restore customizations The same customization options for [RonDB and Opensearch](#customizations) backup IDs apply to in-place restore. You can override individual service backup IDs while keeping the global backup ID for HopsFS. ### Superset restore Superset is restored by reloading its database from a backup (see [Superset backup][superset-backup]). Nothing may write the database while it is reloaded, so the restore is a two-step operation: hold Superset at zero replicas and reload, then start it again. Find the Superset backup id to restore: ```bash kubectl get configmap superset-backups-metadata -n hopsworks -o json \ | jq -r '.data | to_entries[] | select(.value | fromjson | .state == "SUCCESS") | .key' \ | sort -r ``` Set the restore trigger, add the chart's `values.superset-restore.yaml`, which pins every Superset workload to zero replicas, and run `helm upgrade`: ```yaml global: _hopsworks: restoreFromBackup: superset: enabled: true backupId: "20260722215632-2116913658" # Recorded in the audit trail. Defaults to helm/@rev. initiatedBy: "ops:HWORKS-2973 alice" ``` ```bash helm upgrade hopsworks hopsworks/hopsworks --version \ --namespace hopsworks \ -f values.yaml \ -f values.superset-restore.yaml \ --timeout 1200s ``` The chart refuses to render the restore without the overlay: zero replicas has to be the desired state for the whole window, which is what keeps the barrier in place under both Helm and ArgoCD. It also refuses when Superset itself is not installed, rather than reporting success for a recovery that would restore nothing. The Superset init Job is skipped while the flag is set, so no schema work runs during the reload. The `superset-restore-` Job waits for the Superset pods to terminate, verifies the dump against its manifest (backup id, object path, size and sha256), waits until the live `superset-secret-key` matches the fingerprint recorded at backup time, and reloads the database. The fingerprint check is what makes the stored database connections usable afterwards: Superset encrypts their passwords with the secret key, so a database restored under a different key would leave every connection present but undecryptable. On a fresh cluster the key arrives with the platform Velero restore, and the Job waits for it (up to 30 minutes) rather than failing. Follow the Job and confirm it completed: ```bash kubectl logs -f job/superset-restore- -n hopsworks kubectl get job superset-restore- -n hopsworks ``` Then remove the overlay and clear the flag to start Superset again: ```yaml global: _hopsworks: restoreFromBackup: superset: enabled: false ``` ```bash helm upgrade hopsworks hopsworks/hopsworks --version \ --namespace hopsworks \ -f values.yaml \ --timeout 1200s ``` This second step is required: it is what returns the Superset workloads to their normal replica counts. The init Job runs on this upgrade and migrates the restored schema forward with `superset db upgrade`, as it does after an image upgrade, so a backup taken by an older Superset version is brought up to the running image. A backup taken by a newer Superset than the running image cannot be migrated down; the init Job fails with an Alembic error and the upgrade reports it. #### Audit trail and the backup fence Every restore appends to the `superset-restore-audit` ConfigMap: one entry per step (start, pods stopped, manifest verified, database reloaded, or the abort reason), each recording the time, the backup id and the initiator. The ConfigMap is retained across upgrades and carries the `backup.hops.works/include` label, so Velero captures it and the record survives the cluster it describes. The newest 50 entries are kept. The recorded initiator is deployment metadata, not an authenticated identity; correlate it with the Kubernetes audit log entry for the Helm or ArgoCD change to identify the actor. ```bash kubectl get configmap superset-restore-audit -n hopsworks -o json \ | jq -r '.data | to_entries | map(select(.key|startswith("h-"))) | sort_by(.key) | .[].value' ``` The same ConfigMap carries the fence: while a restore is running, or after one has failed, its `inProgress` key holds the backup id and scheduled Superset backups skip their run, so a half-reloaded database is never captured as a backup. A completed restore clears it. After a failed restore, verify or recover the database, then clear the fence: ```bash kubectl patch configmap superset-restore-audit -n hopsworks --type json \ -p '[{"op":"remove","path":"/data/inProgress"}]' ``` To run the same restore again, delete its Job first; the chart refuses to render a restore whose Job already exists: ```bash kubectl delete job superset-restore- -n hopsworks ``` A retry, whether by the Job's own back-off or by hand, repeats every step. That is safe because the reload drops and recreates the schema, and the fence stays set from the first attempt until a reload succeeds. #### Fresh-cluster restore On a brand-new cluster the sequence is the same two steps, with `helm install` in place of the first `helm upgrade`: 1. Restore the platform Velero backup so the Superset Secrets (including `superset-secret-key`) exist on the new cluster. 2. Install (or sync) the chart with the Superset restore flag set and `values.superset-restore.yaml` added. The Job reloads the database into the freshly created MySQL while Superset is held at zero replicas. 3. Wait for the `superset-restore-` Job to complete. 4. Remove the overlay, set `enabled: false` and upgrade, so Superset starts against the restored database and the init Job migrates the schema. If step 1 is skipped, the Job waits for the secret key and then fails with the mismatch recorded in the audit trail; nothing has been written to the database at that point. #### Limitations - The Superset restore is a logical reload of the metadata database, not a point-in-time snapshot coordinated with the RonDB or HopsFS backups. The Superset backup and the platform backup are taken independently, so a restore recovers each service to its own most recent backup, not to a single consistent instant across services. - The Superset MySQL credentials (`superset-mysql-users-secrets`) and the admin account (`superset-admin-credentials`) are fixed at install and are captured and restored as-is from the Velero backup. Rotating them is not supported for the lifetime of any cluster you intend to restore in place: an in-place restore rolls the Secrets back to the backed-up values, which would then be out of sync with the live MySQL grants written after a rotation. If a rotation is unavoidable, treat it as a re-baseline: rotate, then take a fresh backup, and discard backups taken before the rotation. - Enabling the Superset restore has no effect unless Superset itself is enabled (`global._hopsworks.superset.enabled=true`), and the restore reloads the bundled MySQL, so `superset.mysql.enabled` must be true (the chart rejects the restore at render time otherwise). - The restore does not flush Superset's Redis cache. Cached chart data and query results expire on their configured timeout, so entries cached before the restore may be served until then. - Connectors are restored as rows with their credentials, and the restore verifies that the secret key matches the backup so those credentials remain decryptable. What it cannot guarantee is that a credential is still the right one, for connectors whose password is owned by another service. The per-project Trino connections are the case that matters: Hopsworks stores each user's Trino password in its own secret store in RonDB, rebuilds the Trino password file from RonDB, and copies that same password into the Superset connection. So the Superset side holds a copy, and RonDB is the source of truth. If Superset and RonDB are restored to the same point, the copy matches and the connection works. If they are restored to different points, and that user's secret was recreated in between (which happens when a user is removed from a project and added again), the restored Superset connection carries a password Trino no longer accepts. The certificate material Superset uses to reach Trino is a CA bundle for verifying Trino's server certificate, not a credential, so its reissue by the certs-operator is expected and harmless. - A stale connector password does not repair itself, because Hopsworks only creates a connection when one is absent and skips when it already exists. After a fresh-cluster restore, confirm Trino access by opening a Trino-backed chart or running a query through a per-project Trino connection. If a user's Trino connection fails to authenticate, delete that connection in Superset and let Hopsworks recreate it from the current RonDB secret. A restored connection that decrypts is not proof that it still authenticates: the decryption only shows the secret key came back, while the password inside is a copy of a secret owned by RonDB. Run a query to confirm it. #### ArgoCD The Superset Secrets (`superset-secret-key`, `superset-mysql-users-secrets`, `superset-admin-credentials`) are generated once and preserved across upgrades using a `lookup` that returns nothing during an offline `helm template`. Under ArgoCD, which renders with `helm template`, a sync can regenerate these Secrets and overwrite the values a Velero restore put back, which would leave the restored database's encrypted connections undecryptable. Add an `ignoreDifferences` entry so ArgoCD ignores their data, and the `RespectIgnoreDifferences=true` sync option so it does not re-apply the rendered values during sync. By default `ignoreDifferences` only affects the diff ArgoCD shows; without `RespectIgnoreDifferences=true` a sync still applies the freshly-rendered Secret values and overwrites the restored ones: ```yaml spec: syncPolicy: syncOptions: - RespectIgnoreDifferences=true ignoreDifferences: - group: "" kind: Secret name: superset-secret-key jsonPointers: ["/data"] - group: "" kind: Secret name: superset-mysql-users-secrets jsonPointers: ["/data"] - group: "" kind: Secret name: superset-admin-credentials jsonPointers: ["/data"] ``` The zero-replica barrier is declarative: with `values.superset-restore.yaml` among the Application's value files, `helm template` renders every Superset workload at zero replicas, so that is the desired state and self-heal maintains it rather than fighting it. No auto-sync pause is needed during the reload. The two steps map directly to two syncs: add the overlay and the flag and sync (the Job reloads the database while Superset is held at zero), then remove both and sync (the workloads return to their normal replica counts and the init Job migrates the schema). Do not remove the overlay until the restore Job has completed. ================================================================================ # Access Audit Logs Source: https://docs.hopsworks.ai/latest/setup_installation/admin/audit/audit-logs/ # Access Audit Logs ## Introduction Hopsworks collects audit logs on all URL requests to the application server. These logs are saved in Payara log directory under ```/audit``` by default. ## Prerequisites In order to access the audit logs you need the following: - Administrator account on the Hopsworks cluster. - SSH access to the Hopsworks cluster with a user in the ```glassfish``` group. ## Step 1: Configure Audit logs Audit logs can be configured from the _Cluster Settings_ Configuration page. You can access it by clicking on your name, in the top right corner, choosing _Cluster Settings_ from the dropdown menu, then choosing _Configuration_ under _Infrastructure_ in the left sidebar.
Audit log configuration
Audit log configuration
Type _audit_ in the search box to see the configuration variables associated with audit logs. To edit a configuration variable, you can click on the edit button (:material-pencil:), insert the new value and save changes clicking on the check mark (:material-check:). !!! info "Audit logs configuration variables" | Name | Description | | --------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | audit_log_count | the number of files to keep when rotating logs (java.util.logging.FileHandler count) | | audit_log_file_format | log file name pattern. (java.util.logging.FileHandler.pattern) | | audit_log_file_path | the directory the audit log files are written to. | | audit_log_file_type | the output format of the log file. Can be one of java.util.logging.SimpleFormatter (default), io.hops.hopsworks.audit.helper.JSONLogFormatter, or io.hops.hopsworks.audit.helper.HtmlLogFormatter. | | audit_log_size_limit | the maximum number of bytes to write to any one file. (java.util.logging.FileHandler.limit) | !!! warning Hopsworks application needs to be reloaded for any changes to be applied. For doing that, go to the Payara admin panel (```https://:4848```), click on _Applications_ on the side menu and reload the _hopsworks-ear_ application. ## Step 2: Access the Logs To access the audit logs, SSH into the **instance pod** of your Hopsworks cluster and navigate to the path ```/opt/payara/appserver/glassfish/nodes///logs/audit```. Audit logs follow the format set in the _audit\_log\_file\_type_ configuration variable. !!! note "Example of audit logs using JSONLogFormatter" ```json {"className":"io.hops.hopsworks.api.user.AuthService","methodName":"login","parameters":"[admin@hopsworks.ai, org.apache.catalina.connector.ResponseFacade@2de6dd0b, org.apache.catalina.connector.RequestFacade@7a82f674]","outcome":"200","caller":{"username":null,"email":"admin@hopsworks.ai","userId":null},"clientIp":"10.0.2.2","userAgent":"Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/107.0.0.0 Safari/537.36","pathInfo":"/auth/login","dateTime":"2022-11-09 12:00:08"} ``` Regardless the format, each line in the audit logs can contain the following variables: !!! info "Audit log variables" | Name | Description | | ---------- | --------------------------------------------------------------------------- | | className | the class called by the request | | methodName | the method called by the request | | parameters | parameters sent from the client | | outcome | response code sent from the server | | caller | the logged in user that made the request. Can be username, email, or userId | | clientIp | the IP address of the client | | userAgent | the browser used by the client | | pathInfo | the URL path called by the client | | dateTime | time of the request | ## Going Further You can [export audit logs](../audit/export-audit-logs.md) to use them outside Hopsworks. ================================================================================ # Export Audit Logs Source: https://docs.hopsworks.ai/latest/setup_installation/admin/audit/export-audit-logs/ # Export Audit Logs ## Introduction Audit logs can be exported to your storage of preference. In case audit logs have not been configured yet in your Hopsworks cluster, please see [Access Audit Logs](../audit/audit-logs.md). !!! note As an example, in this guide we will show how to export audit logs to BigQuery using the ```bq``` command-line tool. ## Prerequisites In order to export audit logs you need SSH access to the Hopsworks cluster. ## Step 1: Create a BigQuery Table Create a dataset and a table in [BigQuery](https://cloud.google.com/bigquery/docs/datasets#console). The table schema is shown below. ```plaintext fullname mode type description pathInfo NULLABLE STRING methodName NULLABLE STRING caller NULLABLE RECORD dateTime NULLABLE TIMESTAMP bq-datetime userAgent NULLABLE STRING clientIp NULLABLE STRING outcome NULLABLE STRING parameters NULLABLE STRING className NULLABLE STRING caller.userId NULLABLE STRING caller.email NULLABLE STRING caller.username NULLABLE STRING ``` ## Step 2: Export Audit Logs to the BigQuery Table Audit logs can be exported in different formats. For instance, to export audit logs in JSON format set ```audit_log_file_type=io.hops.hopsworks.audit.helper.JSONLogFormatter```. !!! info For more information on how to configure the audit log file type see the ```audit_log_file_type``` configuration variable in [Audit logs](../audit/audit-logs.md#step-1-configure-audit-logs). To export the audit logs to the BigQuery table created in the previous step, run the following command. ```sh bq load --project_id \ --source_format=NEWLINE_DELIMITED_JSON \ \ /srv/hops/domains/domain1/logs/audit/server_audit_log0.log ``` !!! tip This command can be configured to run periodically on a given schedule by setting up a cronjob. ================================================================================ # Service Operations Source: https://docs.hopsworks.ai/latest/setup_installation/admin/operationLogs/ # Service Operations Service Operations provides a centralized view of asynchronous operations executed across Hopsworks services. Operations are handled by a timer-based system that manages execution, retries, and tracking across multiple service handlers. ## Overview Asynchronous operations in Hopsworks are processed by background timers rather than executing immediately. Each operation can be handled by multiple services, with each service registering its own handler. The Service Operations interface allows administrators to: - Monitor ongoing and completed operations - View operation status per service - Track retry attempts and failures - Reset failed operations - View operation history
Service Operations
Service Operations
## Operation Status Each operation displays status information for every service that has registered a handler. The status table shows: - **Operation ID**: Unique identifier for the operation - **Service**: The service handling this operation - **Status**: Current state (pending, running, completed, failed) - **Retry count**: Number of retry attempts made - **Last execution**: Timestamp of the most recent execution attempt Click on an operation to view detailed status messages and execution logs.
Service Operations status message
Service Operations status message
## Retry Mechanism When an operation fails for a service, the system automatically retries using exponential backoff: - **Exponential backoff**: Retry intervals increase exponentially (e.g., 1 min, 2 min, 4 min, 8 min, etc.) - **Daily limit**: Operations are retried up to a maximum number of times per day - **Successful operations**: Automatically removed from the active list and moved to history This approach prevents overwhelming services with continuous retries while allowing temporary failures to resolve automatically. ## Resetting Backoff For operations with more than 10 retry attempts, administrators can manually reset the backoff: 1. Locate the failed operation in the Service Operations list 2. Click on the operation to view details 3. If the retry count exceeds 10, a "Reset Backoff" button will be available 4. Click to reset the backoff to its initial value This allows the operation to retry sooner after you've resolved the underlying issue. ## Operation History Completed and failed operations are retained in the system for a configurable number of days. After the retention period expires, operations are automatically purged from history. This keeps the operation log manageable while preserving recent historical data for troubleshooting and auditing. ## Configuration Service Operations behavior can be customized through cluster configuration variables. To modify these settings, navigate to **Cluster Settings** → **Configuration** and search for the variable name. **Available Variables:** - **async_services_timer_enabled**: Enable or disable the async services timer (default: `true`) - **async_services_timer_interval_ms**: Timer execution interval in milliseconds (default: `15000` = 15 seconds) - **async_services_timer_delete_history_after_days**: Days to retain operation history (default: `7`) - **async_services_timer_batch_size**: Maximum operations processed per timer execution (default: `1000`) Adjust these values to balance system responsiveness, resource usage, and historical data retention. ## Best Practices - **Monitor regularly**: Check Service Operations to identify recurring failures - **Investigate failures**: Click on failed operations to review error messages and identify root causes - **Reset when appropriate**: Use backoff reset after fixing issues to speed up recovery - **Configure retention**: Set operation history retention based on compliance and troubleshooting needs - **Track patterns**: Look for patterns in failures that may indicate systemic issues ================================================================================ # Query Engine (Trino) Source: https://docs.hopsworks.ai/latest/setup_installation/admin/trino/ # Query Engine (Trino) As a Hopsworks administrator, you can monitor and manage the Trino cluster used for query execution across all projects. The admin interface provides cluster-wide visibility into resources, performance, and worker health. ## Cluster Overview The cluster overview provides a comprehensive view of your Trino deployment, including: - **Cluster status**: Overall health and availability - **Active queries**: Total number of running queries across all projects - **Worker nodes**: Number of active and total workers - **Resource utilization**: Cluster-wide CPU and memory usage - **Query throughput**: Average query execution times and data processed Use this dashboard to monitor overall cluster health and identify capacity issues.
cluster overview
Trino cluster overview
## Query History The query history shows all queries executed across the Trino cluster, regardless of project. This centralized view helps administrators: - **Monitor usage patterns**: Identify peak usage times and resource-intensive queries - **Troubleshoot issues**: Investigate failed or slow queries - **Audit activity**: Track query execution by project and user - **Optimize performance**: Identify queries that may need optimization Each query entry displays: - Query ID and text - Project and user who executed it - Status (running, completed, failed) - Execution time and resources consumed - Timestamp Click any query to view detailed execution information.
query history
Trino query history
## Managing Workers The workers view displays all Trino nodes in the cluster. For each node, you can see: - **Node IP**: IP address of the worker node - **Node version**: Trino version running on the node - **Coordinator or worker**: Role of the node (coordinator or worker) - **State**: Current state of the node (active, idle, or offline) This view helps you monitor the cluster topology and identify any nodes that may be offline or experiencing issues.
workers
Trino workers
### Worker Status Details Click on a worker to view detailed status information: - **Resource metrics**: Detailed CPU, memory, and network usage over time - **Task breakdown**: Types and number of tasks being executed - **Error logs**: Any errors or warnings from the worker - **Configuration**: Worker settings and assigned resources - **Performance history**: Historical performance trends Use this detailed view to diagnose worker-specific issues and optimize resource allocation.
worker status
Trino worker status
## Managing Catalogs Catalogs created by project Data Owners are saved to the database but are not loaded by the running cluster until they are applied. Two things apply them: the scheduled restart, which needs no administrator, and the Catalogs tab under Cluster Settings, Query Engine, where an administrator can apply them immediately or gate them behind approval. The tab lists every catalog waiting to be applied, along with its status and the operation to apply (create, update, or remove). Nothing notifies you when a Data Owner creates a catalog, and nothing notifies them when it is applied. With the schedule on, a pending catalog goes live at the next restart that finds it, so the tab is where you go to apply one sooner or to reject it; with approval required, nothing goes live until you act, so check the tab periodically or agree a cadence with your projects.
Pending catalogs
The lifecycle settings, the catalogs waiting to be applied, and the one action that applies them
### Applying pending requests Every pending request is selected by default. Clicking **Restart Trino** applies the selected ones in a single action: their definitions are written into the backend-owned Kubernetes Secrets that the cluster mounts at `/etc/trino/catalog`, and the query engine is restarted afterwards to load them, behind a dialog that confirms what is about to be applied. Both halves are needed and in that order, which is why they are one button: a restart on its own would load nothing, because a catalog is only a database row until its definition is written out. You can also **Delete** an individual pending request, which rejects that change without applying it. The restart interrupts queries running anywhere on the cluster, so check the reported activity before confirming. #### Where catalog credentials are stored A connector's credentials end up in the places below. Anyone who can read those places can read the credentials, so plan access to them accordingly. - A `${HOPSWORKS_SECRET:}` reference is stored verbatim in the `trino_catalog` database row and is resolved to its value only at approval time. The database row never holds the value. - A literal value typed straight into the properties editor is stored as-is in the `trino_catalog` database row, in cleartext, and is captured by database backups. Use a secret reference for any credential you do not want in the database. - Either way, the written file holds the resolved plaintext, because Trino reads the credential from the catalog file itself. That file lives in a Kubernetes Secret rather than a ConfigMap, so it is covered by the RBAC that applies to Secrets in the Hopsworks namespace and by etcd encryption-at-rest on clusters that enable it. ### Lifecycle settings The same tab carries the **Catalog lifecycle** card, where the whole schedule is configured and saved as one group: - **Scheduled restart every N hours or days.** The cadence is anchored at the configured time of day, so it keeps its phase across redeploys, and the next restart the schedule resolves to is shown next to the input. - **Require approval for all catalog changes.** Turning this on cancels the scheduled restart entirely, because approval means nothing goes live unattended. Pending requests then wait in the table until an administrator applies them with Restart Trino, or rejects them with Delete. - **Eager restart.** The query engine is checked every few minutes, and pending changes are applied ahead of the schedule the moment no query is running, queued, or blocked, so the restart lands in a moment with nothing to cancel. Users are told their catalog may go live earlier than the scheduled time. - **Maximum catalogs**, across every project, at most 250. Each catalog is a file the query engine loads at startup, so the deployment is sized for a bounded number; the setting may lower the bound but never raise it past the ceiling. Saving needs no redeploy: every Hopsworks instance derives its schedule from these settings and picks a change up within a minute, whichever instance served the save. ### A single project's allowance The cluster-wide maximum is a ceiling on the deployment; how many catalogs any one project may create is set per project. Open **Cluster Settings** → **Projects**, click **Edit configuration** on the project's row, and scroll to **Query engine**, **Trino catalogs**, just after the Kafka topic quota.
A project's Trino catalog allowance
The project's catalog count, its own limit, and the unlimited checkbox
A new project starts on the cluster default (`trino_catalog_max_per_project`, 10), so raising one project here raises that project only. Checking **unlimited** removes the project's own bound, leaving only the cluster-wide ceiling; a limit of 0 blocks new catalogs in the project. Both bounds apply to a create: the project must be under its own allowance, and the cluster must be under the ceiling. ### The wait for a quiet moment A due scheduled restart does not fire into a busy cluster immediately. It waits for the cluster to have no query running, queued, or blocked, re-checking every few minutes for up to an hour, and then restarts anyway: the wait buys a quiet moment when one exists, and the bounded give-up keeps a permanently busy cluster from deferring catalog changes forever. An activity count the query engine cannot report counts as busy rather than idle, so a failed reading never costs someone their query. The bounded wait is what makes that safe: a coordinator that is genuinely down never reports itself idle, and the restart that recovers it still happens when the window expires. ### Restarting Trino reads catalogs only at startup, so a catalog change takes effect on the next restart, whether the schedule performs it or an administrator does. Clicking "Restart Trino" applies the selected pending requests and rolls out the coordinator and workers. The confirmation dialog reports how many queries are currently running or queued, so you can choose a low-traffic window, and asks you to type `confirm` before it will proceed. The restart cancels those queries for **every project on the cluster**, not only the project whose catalog is being applied, and in-flight results are lost. Trino keeps recent query detail in the coordinator's memory, so after a restart the live query views show only what the new coordinator has seen; older queries remain in the query history, which is stored separately.
Restart confirmation
The confirmation names what the restart applies and what it interrupts
A restart is refused while another one is already running, so concurrent actions by different administrators cannot collide or trigger redundant restarts. If nothing is waiting to load or unload, the restart is skipped and reported as such rather than interrupting queries for no reason. ### Recovering a catalog Trino cannot load Trino reads its catalogs at startup and refuses to start if it cannot load one of them. A user catalog with an invalid definition therefore stops the whole query engine, coordinator and workers alike, and the pods stay in `CrashLoopBackOff`. Kubernetes keeps the previous pods serving while the new ones fail, so queries may keep working for a while and the rollout never completes. Click "Restart Trino" to recover. The button stays available when nothing is waiting to be applied, because this situation leaves no pending catalog to load: the restart itself is the repair.
Recover restart
A query engine that will not start is reported on the tab, and the restart action recovers it
Hopsworks reads the coordinator log, identifies the catalog Trino rejected, removes it from the mount, marks it **Failed**, and restarts so the cluster comes back without it. The result names the catalogs that were removed: > Removed 1 catalog Trino could not load. `project1__orders_pg`. Trino is restarting without them; the owners must fix the definitions. Only user-created catalogs are removed this way. A default catalog that fails to load is left in place, because that is a cluster configuration problem rather than something an administrator should resolve by deleting data. A removed catalog keeps its row, so its owner can see what happened on the project's Catalogs page along with the error Trino reported. Editing the definition returns it to Pending approval and it re-enters the normal flow. Failed catalogs are not listed under pending, because they no longer block anything and no administrator action can fix them. If Hopsworks cannot attribute the failure to a user catalog, it reports the connection error instead of removing anything, and the coordinator log is the place to look. ### Recovering catalog files lost from the mount The catalog definitions are stored in the Hopsworks database, and the Kubernetes Secrets mounted at `/etc/trino/catalog` are derived from it. A restore that brings back the database alone, a GitOps sync that prunes resources it does not manage, or a Secret deleted by hand therefore leaves catalogs that exist in Hopsworks with no file for Trino to read. `POST /hopsworks-api/api/admin/trino/catalogs/reconcile` repairs it. It writes the missing files back from the database, removes files that no catalog belongs to, which is also how a credential stops being mounted once its catalog is gone from the database, and reports what it changed. It is always available, since an administrator calling it has already established that the repair is needed. It also compares file contents, so it corrects a file whose name is right but whose content no longer matches the database. Set **trino_catalog_reconcile_enabled** to `true` to have the same repair run on a schedule instead, shortly after startup and on the reconcile interval thereafter. It is off by default because losing a shard Secret takes one of the events above rather than anything routine, so the repair belongs on a cluster that needs it rather than on every cluster. Each scheduled pass compares which catalog files the Secrets hold against which ones the database expects, and repairs only when they disagree. Comparing names rather than contents is what keeps a pass cheap enough for an interval, since rebuilding a file means decrypting every secret it references. The consequence is that the scheduled pass does not notice a file whose name is right and content is wrong; use the endpoint for that. Two things neither form does. Neither restarts Trino, so a restored catalog is in the mount but not loaded until the next restart, like any other catalog change. Neither touches a catalog that is pending approval, because that catalog's stored definition is the change an administrator has not approved yet, and applying it here would bypass that decision. Those catalogs are reported as still needing approval. A catalog whose `${HOPSWORKS_SECRET:}` reference no longer resolves cannot be rebuilt, since the file Trino reads has to hold the resolved value. The repair reports it, leaves any file it already has in place, because that copy resolved when it was approved and still works, and carries on with every other catalog. Its owner has to repoint the reference at an existing secret. ## Access control and sharing The query engine decides who can read what with Trino's file-based access control, from a rules file published into the Trino files store as `access-control/rules.json`. Hopsworks owns that file and rebuilds it whenever a share changes, and on a schedule every five minutes by default. The file is composed from two parts: - The base policy, from the Helm value `trino.accessControl.rules`, which the chart renders into the ConfigMap `hopsworks-trino-access-control-base`. It grants each project its own catalogs and feature store, and each user their private catalogs, written only from projects where the user is a Data Owner. Administrators see every catalog but read only `system`, `tpch` and `tpcds`, because the query engine's administrator is also the identity Hopsworks itself uses, and a view recorded as owned by it would otherwise read any project's data. The administrators' SQL console therefore cannot read a project's tables; query them as a member of the project. - One set of rules per share, for [catalog shares and feature group shares][sharing-catalogs-and-feature-groups]. A share names the receiving project's existing `__data_owner` and `__data_scientist` groups, so sharing never changes the group file. Change the base policy through the Helm value and an upgrade. An edit to the published `rules.json` is overwritten by the next rebuild, within minutes. Every rebuilt file is validated before it is published, and a file that fails validation is not published: the shares that caused it are marked **Failed** with the reason, and the file in place stays as it was. After publishing, Hopsworks checks that the query engine still answers once it has re-read the file, and restores the last file that worked if it does not, because Trino refuses every query while its rules file is unreadable. ### The shared feature store catalogs The chart ships two kinds of catalog over the feature store: - `delta`, `hudi`, `iceberg` and `hive` impersonate the querying user, so HopsFS permissions apply on top of the access-control rules. They serve a project's own feature groups, and feature groups or feature stores shared whole, which HopsFS grants the receiving project. - `delta_shared`, `hudi_shared` and `iceberg_shared` do not impersonate. They read HopsFS as the `trino` user, which is a HopsFS superuser, because a feature group shared with a subset of its features grants the receiving project no HopsFS access. For the second kind the access-control rules are the only gate. The base policy grants nobody access to them, not even administrators, and Hopsworks adds a rule per subset share that allows the receiving project the shared features of that one table and denies the rest. They are read-only at the connector as well, so no rule can let a query write through them. Do not add rules for these catalogs to the base policy: any rule that reaches one of them reads every project's feature store. Administrators are denied them because Trino runs a view as the user recorded as its owner, so a view recorded as owned by an administrator would reach them too. ### Reading the rules file The rules the query engine enforces can be read under **Cluster Settings** → **Query Engine** → **Files**, as `access-control/rules.json`. Beside it, `access-control/rules.json.last-good` is the last file the query engine loaded. They differ from a publish until Hopsworks confirms the query engine loaded the new file, a few seconds later. If they stay different, the new file is not confirmed yet, for example because the query engine was unreachable, and the next reconcile checks again. That check only happens while **trino_reconcile_enabled** is on. A file the query engine refused does not stay: the last good file goes back in its place, and the shares the refused file added are marked **Failed**. The groups the rules name are in `auth/group.db`. #### Who the rules match A query runs as a principal named `__`, for example `seeda__seed1000` for user `seed1000` in project `seeda`. Its groups are the member's role in that project, `__data_owner` or `__data_scientist`, and `__shared_featurestore` for each project `` whose feature store is shared with that project. Group `admin` has one member, the query engine administrator, which is the identity Hopsworks itself uses. Project names and usernames cannot contain `__`, so a pattern such as `.*__(.*)` splits a principal unambiguously, and `$1` in a later field stands for what the pattern captured. Each section of the file (`catalogs`, `schemas`, `tables`, `functions`, `queries`) is checked on its own. In a section, the first rule whose user, group and object all match decides, and a request no rule matches is denied. The order of the rules is therefore the policy: a broader rule placed first would answer before a narrower one. #### The order of the rules Every section keeps the same order, and the rules Hopsworks adds for shares go in one place in it: 1. The administrator rules. 2. A deny for each private catalog whose owner's account was deleted, until the catalog is removed. It comes before the private-owner rules because a later account with the same username would match them. 3. The private-owner rules. They come before the shares so that sharing a private catalog with a project the owner belongs to never narrows the owner's own access. 4. The share rules. 5. The rest of the base policy: every project's own catalogs and feature store. #### The base policy The `catalogs` section of the base policy, in order: | Rule | Effect | | --- | --- | | `group: admin`, `allow: none` on `iceberg_shared`, `delta_shared` and `hudi_shared` | The administrator never sees the shared feature store catalogs. | | `group: admin`, `catalog: .*`, `allow: read-only` | The administrator sees every other catalog, without writing to any. | | `user: .*__(.*)`, `group: .*__data_owner`, `catalog: _$1__.*`, `allow: all` | The owner of a private catalog reads and writes it from a project where they are a Data Owner. | | `user: .*__(.*)`, `catalog: _$1__.*`, `allow: read-only` | The owner reads it from any other project. | | `catalog: tpch` and `tpcds`, `allow: read-only` | Everyone reads the sample catalogs. | | `catalog: iceberg`, `delta`, `hive`, `hudi`, `allow: all` | Everyone reaches the feature store catalogs; the table rules decide what they read. | | `group: (.*)__data_owner`, `catalog: $1__.*`, `allow: all` | A project's Data Owners read and write its catalogs. | | `group: (.*)__data_scientist`, `catalog: $1__.*`, `allow: read-only` | Its Data Scientists read them. | | `catalog: system`, `allow: read-only` | Everyone reads the `system` catalog. | The `tables` section follows the same pattern: the administrator reads only `system`, `tpch` and `tpcds`, the private-owner rules mirror the catalog ones, and each project reaches the schema `_featurestore` in the feature store catalogs, all of it for its Data Owners, reading for its Data Scientists and for projects its feature store is shared with. The `schemas` section gives schema ownership, which is what creating and dropping schemas needs, to Data Owners only. The `functions` section lets everyone run builtin functions, and the Data Owners of a project run the `system` functions of their project's catalogs, such as `system.query` on a JDBC catalog. #### The rules a share adds Hopsworks writes the names in a share rule as literals between `\Q` and `\E`, so a name containing regular expression syntax matches only itself. A share names the receiving project's two role groups in one pattern, `\Q\E__data_(?:owner|scientist)`, so searching the file for `\Qseedc\E__data_` finds every rule a share to `seedc` added. A share of catalog `seeda__postgresql` with `seedc`, covering table `public.customers` with column `created` unchecked and column `name` masked, adds these rules: ```json {"catalogs": [ {"group": "\\Qseedc\\E__data_(?:owner|scientist)", "catalog": "\\Qseeda__postgresql\\E", "allow": "read-only"} ], "tables": [ {"group": "\\Qseedc\\E__data_(?:owner|scientist)", "catalog": "\\Qseeda__postgresql\\E", "schema": "\\Qpublic\\E", "table": "\\Qcustomers\\E", "privileges": ["SELECT"], "columns": [{"name": "created", "allow": false}, {"name": "$path", "allow": false}, {"name": "name", "mask": "'***'"}]}, {"group": "\\Qseedc\\E__data_(?:owner|scientist)", "catalog": "\\Qseeda__postgresql\\E", "schema": "\\Qpublic\\E", "table": "\\Qcustomers\\E\\$.*", "privileges": []} ], "functions": [ {"group": "\\Qseedc\\E__data_(?:owner|scientist)", "catalog": "\\Qseeda__postgresql\\E", "privileges": []} ]} ``` The list of denied columns in the rule is shortened here. - The catalog rule makes the catalog visible to the receiving project, read-only. - The table rule grants `SELECT` on the table and lists the columns it denies: the unchecked ones, the connector's hidden columns, such as `$path`, and, for an Iceberg or Delta Lake table, every column the table had at a version that can still be read. A table shared whole has a table rule without `columns`, a schema shared whole has `table: .*`, and a catalog shared whole has `schema: .*` too. - The rule after it, with no privileges, denies the table's metadata tables, such as `customers$partitions`, which Trino checks by their own name. - The function rule denies the receiving project the catalog's functions, which the base policy would otherwise give it. A share of a private catalog also adds, before that deny, a rule letting the owner keep running the catalog's `system` functions. A feature group shared whole adds one table rule on `hive|iceberg|delta|hudi` for its table in the owner's feature store. A feature group shared with a subset of its features adds a catalog rule on the shared catalog of its format, such as `delta_shared`, and a table rule there that denies every unshared feature and hidden column, followed by the metadata table deny. #### Debugging a share - The share is **Active** but a query is refused: find the share's rules by the receiving project's group, then look for a rule above them that matches the same principal and object first. - The share's rules are not in the file: the share is still **Applying**, or it is **Failed** and its status says why. - A column that should be hidden is readable: it is missing from the rule's `columns`. Each publish denies every column the table has at that moment except the shared ones, so a column added at the source is readable only until the next publish, at most one reconcile interval. - A narrowed table is refused although its share is **Active**: the last publish could not read the table's columns, or for an Iceberg or Delta Lake table the columns of its earlier versions, so it left the table out of the rules rather than grant it with columns it could not deny; the Hopsworks log names the table. An Iceberg table whose metadata file is outside HopsFS, over 64 MiB or unreadable by that project user is also left out until its earlier columns are read, 50 snapshots per publish. - `rules.json` and `rules.json.last-good` differ for minutes: the newest file is not confirmed, so check that the query engine is reachable; the reconcile verifies it again and restores the last good file if the query engine refuses it. ## Credential files a project supplies A connector that authenticates with a file, such as an Oracle wallet or a Java keystore, cannot be served by a catalog property alone. Projects supply those files as [mountable secrets][mountable-secrets], and this section covers what that adds to a cluster. A bundle is a directory of files in HopsFS under `mountable_secrets_path`, which defaults to `/apps/mountable-secrets`. It is keyed by **project id** rather than by project name, so a deleted project and a later project of the same name can never share a directory. The `charts/hopsfs` preset Job creates the root as `payara:hdfs` with mode `0750`. The backend checks that owner, group and mode against the filesystem on a project's first bundle, and again before it deletes a project's tree, and refuses if any of the three differs, so a root created by hand with the wrong mode is rejected rather than quietly widening access. Project members never reach those files directly. The path is outside any project's dataset, and a catalog can only ever name a bundle in its own project. A reference resolves to a path built from the project id, and a property that tries to extend a reference with a path, or to walk out of it with `..`, is refused when the catalog is created and again when it is resolved. ### How the files reach the query engine Each Trino pod, the coordinator, every worker and the test coordinator, runs a sidecar container named `mountable-secrets` from the `hopsfs-mount` image. It mounts the whole store read-only at `trino_mountable_secrets_root`, which defaults to `/opt/hopsworks/mounts`, with `ro`, `nosuid` and `nodev`. Entries appear as uid 0 with modes that let any user read them, which is what lets the unprivileged Trino process open a wallet. Two consequences of that sidecar are worth knowing before an upgrade. It is **privileged**, because FUSE requires it. On a cluster running the Kyverno restricted policies the chart ships a `PolicyException` for these pods, gated on Kyverno being enabled. The same privileged FUSE sidecar already runs on the three Airflow deployments in the release namespace, so this is not a new class of workload for the cluster. The Trino pods have their **own ServiceAccounts**, `hopsworks-trino` and `hopsworks-trino-test`, rather than the namespace default. An SCC or a cloud identity can therefore be granted to Trino narrowly. An upgrade from a release before this feature moves those pods off the `default` ServiceAccount, so any binding that named `default` to reach Trino has to be repointed. !!! warning "OpenShift is not supported" The sidecar has to run privileged and as root, so the default restricted SCC rejects it. `values.openshift.yaml` therefore turns the store off, and a project on such a cluster cannot supply credential files. Note that Trino was already rejected by the restricted SCC before this feature, because the subchart pins `runAsUser: 1000` regardless of `securityContextEnabled`, so the sidecar adds a second reason rather than a new break. ### Turning the store off Set `global._hopsworks.trino.mountableSecrets.enabled` to `false`, which seeds the `mountable_secrets_enabled` variable and stops the store being offered. Turning it off is not a single value. The sidecar entries live in untemplated subchart values, so the `initContainers` lists have to be restated without them, which is what `values.openshift.yaml` does and is the worked example to copy. The chart fails the render when the flag and the mount disagree, so a half-done change stops the upgrade instead of producing pods that mount nothing. An **already approved catalog keeps working only as far as its definition**. Its reference still resolves to a path, but nothing populates that path any more. For a connector that opens its files when a connection is made, such as Oracle, the coordinator starts cleanly and queries fail. Writing a catalog out does not consult the flag, by design, so switching the store off does not quarantine catalogs that already use it. ### Backup Bundles are HopsFS files. They are covered by the HopsFS backup, and **not** by the Kubernetes object backup that captures the catalog Secrets and the database. A restore that brings back the database and the Secrets without the HopsFS path leaves catalogs that reference bundles which no longer exist, and those catalogs fail to authenticate at the next restart. Recreating the bundle under the same name with the same filenames repairs it without editing any catalog. ### Diagnosing a bundle There is no admin API for the store, so the checks are on the cluster. ```bash # What the query engine can actually see for project kubectl exec -n hopsworks -c -- ls -l /opt/hopsworks/mounts// # The mount itself, including its options kubectl exec -n hopsworks -c -- grep /opt/hopsworks/mounts /proc/mounts # The source side kubectl exec -n hopsworks -- /srv/hops/hadoop/bin/hdfs dfs -ls /apps/mountable-secrets/ ``` Check the mount on a **worker** and not only on the coordinator, since a query reads the source from the workers. A missing mount is otherwise invisible: the sidecar mounts into its own filesystem, both containers report ready, and only a `${HOPSWORKS_MOUNT:...}` reference resolving to an empty directory gives it away. The outbound addresses a data source must admit are reported in the project's Catalogs tab only when `global._hopsworks.trino.mountableSecrets.egressProbe.echoUrl` is set. It is empty by default, because the probe otherwise calls a third-party service from every Trino pod on every start, and with it unset the UI reports that the addresses could not be determined. Without it: ```bash kubectl exec -n hopsworks -c -- curl -s https://ifconfig.me ``` ## Configuration Trino behavior can be customized through cluster configuration variables. To modify these settings, navigate to **Cluster Settings** → **Configuration** and search for the variable name. **Available Variables:** - **trino_enabled**: Enable or disable Trino cluster-wide (default: `false`) - **trino_default_catalog**: Default catalog of the Superset database connections created for new project members (default: `delta`). Connections created before a change keep the catalog they were created with. - **trino_test_coordinator_enabled**: Enable the optional test coordinator that backs the "Test connection" action for user-created catalogs (default: `true`) - **trino_reconcile_enabled**: Rebuild the login and group files from the database and republish the access-control rules on the reconcile interval (default: `true`). Do not disable it. It is what confirms a published rules file once the query engine was unreachable when Hopsworks first checked: without it, shares stay **Applying** or **Revoking** and the new file never becomes the last good one, until another share change publishes again. It is also what brings a changed base policy from a chart upgrade to the query engine, and what replaces a rules, login or group file that was edited, corrupted or deleted. - **trino_reconcile_interval_ms**: How often the reconcile runs, in milliseconds (default: `300000`) - **trino_catalog_reconcile_enabled**: Rebuild the user-catalog Secrets from the database on a schedule, for a cluster that has lost them (default: `false`, see [Recovering catalog files lost from the mount][recovering-catalog-files-lost-from-the-mount]) - **trino_catalog_max_per_project**: Catalogs a *newly created* project may create (default: `10`). It seeds each project's own allowance, which is then edited per project under Cluster Settings, Projects; changing it does not move the allowance of a project that already exists. - **trino_catalog_max_bytes**: Largest a single catalog definition may be once its secret references are resolved, in bytes (default: `16384`) - **trino_max_catalogs**: Catalogs the whole cluster may have, across every project (default: `250`, which is also the ceiling). Each catalog is a file the query engine loads at startup, so the setting may lower the bound but never raise it. - **trino_scheduled_restart_enabled**: Apply pending catalog changes with a scheduled restart (default: `true`). Safe to leave on, because the restart is skipped entirely when no catalog change is pending. - **trino_scheduled_restart_interval_hours**: How often the scheduled restart fires (default: `24`). Edited from the Catalog lifecycle card as "every N hours/days". - **trino_scheduled_restart_time**: Anchor time of day for the cadence, `HH:mm` in the server's timezone (default: `02:00`). Off-peak by default because the restart cancels every running query. - **trino_scheduled_restart_idle_wait_minutes**: How long a due scheduled restart waits for the cluster to go quiet before restarting anyway (default: `60`). Bounded, because a permanently busy cluster must not defer catalog changes forever. - **trino_scheduled_restart_idle_retry_minutes**: How long to wait between those quiet-moment re-checks (default: `5`). - **trino_eager_restart**: Restart ahead of the schedule the moment the query engine is idle while changes are pending (default: `false`). - **trino_eager_restart_poll_minutes**: How often the eager restart looks for that idle moment (default: `10`). - **trino_catalog_approval_required**: Require an administrator to apply every catalog change (default: `false`). Turning it on cancels the scheduled restart timers entirely, because approval means nothing goes live unattended. These settings control the availability and default behavior of the Trino query engine across your Hopsworks cluster. ### Mountable secret settings These are not all editable the same way, so they are listed apart from the variables above. Three are seeded by the chart and belong to Helm, not to the variables table. | Setting | Helm value | Seeded default | | --- | --- | --- | | `mountable_secrets_enabled` | `global._hopsworks.trino.mountableSecrets.enabled` | `true` | | `mountable_secrets_path` | `global._hopsworks.trino.mountableSecrets.storeRoot` | `/apps/mountable-secrets` | | `trino_mountable_secrets_root` | `global._hopsworks.trino.mountableSecrets.mountPath` | `/opt/hopsworks/mounts` | Change these through your Helm values and an upgrade. Editing the row instead moves only one end of the arrangement: the store root also presets the HopsFS directory and is passed to the mount sidecar as its source, and the mount root is what the Trino containers actually mount, so a row edited on its own points the backend at a path nothing is mounted from. The chart keeps the two ends together, and refuses to render when the flag and the mount disagree. Note also that the code's own fallback for the flag is `false`, which is what a cluster whose chart predates the row gets; the chart seeds `true`. The five per-project limits have no seeded row at all. The code's defaults apply until an administrator creates one, so searching for them in Cluster Settings finds nothing on a fresh cluster, which is expected rather than a fault. | Setting | Default | What it caps | | --- | --- | --- | | `mountable_secret_max_per_project` | `10` | bundles one project may hold | | `mountable_secret_max_files` | `32` | files in one bundle | | `mountable_secret_max_file_bytes` | `1048576` | largest single file, in bytes | | `mountable_secret_max_project_bytes` | `16777216` | a project's total across all its bundles, in bytes | | `max_mountable_secret_upload_bytes` | `33554432` | largest upload request, refused before the body is read | Turning the store off is described in [Turning the store off][turning-the-store-off]. #### How sharing scales Every change to a share, and every reconcile, publishes the whole rules file again. A publish reads the current columns of each table a share narrows to some of its columns, one statement per table, so its duration grows with the number of distinct narrowed tables across all shares. A narrowed Iceberg or Delta Lake table also has the columns of its earlier versions read: - An Iceberg table on HopsFS: one more statement, and a read of its current metadata file, which lists every schema the table has had. Hopsworks reads the file as the project user the table's columns are read as, only up to 64 MiB, and uses it only when it names the snapshot Trino reports for it. - Any other Iceberg table: two more statements, and one per snapshot not read before: one per schema the table has had, and every snapshot older than its metadata log, which keeps the last 100 entries by default. At most 50 snapshots are read per publish; a table with more is left out until later publishes have read them all. A table with more than 10,000 such snapshots cannot be listed and is left out; expiring old snapshots, or keeping the table on HopsFS, avoids it. - A Delta Lake table: one statement that reads the commits since the last publish, or every commit still in the table's log the first time, and on that first read one more for the oldest of them. The statement reads back from the newest commit to the first missing one, so versions before a gap in the log are not read; only log files removed by hand leave such a gap. Each Hopsworks instance keeps what it has read in memory, so after a restart its first publish reads each table's history in full once. It reads it again at least once a day: a Delta Lake table in full on that publish, an Iceberg table's snapshots 50 per publish while the earlier reads still count, so the table is never left out for it. A publish forgets the tables no share narrows any more. Measured on a development cluster, as extra time per narrowed table on top of reading its current columns: | Table | 1 commit | 100 commits | 300 commits | | --- | --- | --- | --- | | Delta Lake feature group, first read | 0.1 s | 1.9 s | 6.6 s | | Delta Lake feature group, later publishes | not measured | not measured | none measurable | | Iceberg table on HopsFS (metadata file size) | 0.1 s (4 KB) | 0.1 s (203 KB) | 0.3 s (556 KB) | | Iceberg table read through its snapshots, snapshots to read | 1 | 3 | 201 | A later publish of the 300-commit table, reading from the 290th or 299th commit, took as long as `SELECT 1`. Reading the current columns of the Delta Lake table also grew, from 0.5 s at 100 commits to 2.7 s at 300. Saving, editing or revoking a share returns once the share is recorded; the share shows **Applying** or **Revoking** until the publish has run and the query engine has loaded the file, about 15 seconds after the publish ends. Changes made while a publish runs are applied together by the next one. Measured on a development cluster with a PostgreSQL source, which has no earlier versions to read: four shares, each of a project catalog or a private catalog with one of two projects, each narrowed to N tables with two of their six columns shared. | Tables per share | Narrowed tables in the rules | Publish duration | `rules.json` size | Table rules | Median query time | | --- | --- | --- | --- | --- | --- | | 0 (one schema shared whole) | 0 | 5.6 s | 25 KB | 35 | 0.9 s | | 25 | 100 | 10.9 s | 204 KB | 231 | 0.9 s | | 100 | 400 | 21 s | 744 KB | 831 | 0.9 s | | 300 | 1,200 | not measured | 2.19 MB | 2,431 | 0.9 s | The query time is through the Hopsworks API, for a receiving project, the catalog's owner and `SELECT 1` alike, and did not change with the size of the file. The publish duration at 300 tables per share was not measured on its own. With the four shares saved one after another, a save took 5 to 7 seconds at 100 tables per share and 16 to 17 seconds at 300, mostly checking the tables and columns the share names, and all four shares were **Active** 41 and 119 seconds after the last save. A publish costs about 75 ms per distinct narrowed table on top of a fixed 5 seconds. The file grows by about 1.8 KB per narrowed table, mostly the hidden columns each narrowed rule denies. With the file above about a megabyte, a query once failed with `Invalid JSON file '/opt/hopsworks/trino/access-control/rules.json'` caused by `java.io.IOException: Input/output error`. The query engine reads the file through the HopsFS mount, and the read failed while a new file was replacing it; the file itself was complete. The query succeeds when run again. Hopsworks does not take such a read failure for a broken file, so it neither restores the last good file nor fails the shares being applied. ### Test coordinator resource cost `trino_test_coordinator_enabled` is on by default, and enabling it runs **an additional single-node Trino coordinator pod** for the lifetime of the cluster. It exists only to connection-test user catalogs before they are approved, so on a small or cost-sensitive cluster it is reasonable to turn it off. When it is off, "Test connection" reports that testing is unavailable and every other part of the catalog workflow is unaffected. ### Supported connectors A project can create a catalog on any connector installed in the Trino image. Connectors that expose no external data source are rejected: `system` and `jmx` (which would expose the query engine's own internals, including other projects' query text), `memory` and `blackhole` (which hold no data), `datasketches` and `ai` (function plugins), and `tpch` and `tpcds`, which already ship as shared read-only catalogs. The installed set is the cluster variable `trino_connectors`, whose default matches the Trino image the chart pins. The backend refuses a catalog on anything outside it, so a connector name that is not installed is rejected when the catalog is created rather than stopping the coordinator at the next restart. The connector picker in the project UI is served from the same list, so it offers exactly what the backend accepts. Change `trino_connectors` only when running an image with a different plugin set. Removing a connector from the list does not affect catalogs already created on it. ### Catalog storage capacity User-created catalogs are stored across a fixed number of Kubernetes Secrets, set by the Helm value `global._hopsworks.trino.userCatalogShards` (default: `2`). Each Secret holds up to roughly 800 KiB of catalog definitions, so the default gives about 1.6 MiB in total, which is a large number of catalogs. When they are full, an approval fails with an error naming the limit. Raise the value in your Helm values to add capacity. The chart mounts one source per shard and refuses to render if the two disagree, so a mismatch fails the upgrade rather than silently dropping catalogs. Two per-catalog limits keep one project from consuming that shared budget. `trino_catalog_max_per_project` caps how many catalogs a project may create, and `trino_catalog_max_bytes` caps how large a single definition may be. The size is measured after `${HOPSWORKS_SECRET:}` references are resolved, because the resolved form is what occupies a Secret: a stored definition is bounded by its database column, but a reference costs a couple of dozen characters and expands to a secret of up to about 10 KiB, and the same secret may be referenced repeatedly, so a row that fits its column can resolve to megabytes. The check therefore runs both when a catalog is created, so its owner hears about it, and again at approval, because a secret can be rotated to a larger value in between. Both defaults are generous against real catalogs, which are a few hundred bytes; the largest legitimate ones inline a service account JSON or a certificate pair and stay a few KiB. Raise them for a project with an unusual number of external sources, and remember that the product of the two bounds a single project's share of the shard budget. ## Best Practices for Trino Management - **Monitor regularly**: Check cluster overview daily to spot trends and issues early - **Review slow queries**: Investigate queries with long execution times in the query history - **Balance workload**: Ensure workers are evenly distributed and not overloaded - **Scale appropriately**: Add workers during peak usage periods if resources are constrained - **Track growth**: Monitor query volume trends to plan for future capacity needs ================================================================================ # Superset Source: https://docs.hopsworks.ai/latest/setup_installation/admin/superset/ # Superset As a Hopsworks administrator, you can manage the Apache Superset integration, which provides data visualization and business intelligence capabilities. Superset enables users across all projects to create interactive dashboards, explore data through SQL queries, and build advanced visualizations on top of feature data stored in Hopsworks. ## Accessing Superset Admin The Superset admin interface is accessible from the Hopsworks cluster settings. In **Cluster Settings**, choose **Superset** under _External_ in the left sidebar to access administrative controls for the Superset deployment.
Superset admin tab
Superset administration in cluster settings
From this interface, you can manage users, roles, and configure Superset settings that apply across all Hopsworks projects. ## Superset Overview Once you access the Superset interface, you will see the main dashboard that provides access to: - **Dashboards**: Interactive visualizations created by users across projects - **Charts**: Individual visualization components - **Datasets**: Configured data sources from Hopsworks Feature Store - **SQL Lab**: Interactive SQL query interface for data exploration
Superset home page
Superset home page
## Managing Users The user management interface allows you to view and manage all Superset users across the cluster. Navigate to **Settings** → **Users** in the Superset interface (or by navigating to `https:///hopsworks-api/superset/users/`) to access user administration. ### User Overview The users list displays: - **First name** and **Last name**: User identity information - **Username**: Unique identifier for each user - **Email**: Auto-generated email - **Is active?**: Whether the user account is enabled - **Roles**: Assigned roles determining permissions - **Groups**: Group memberships for organizational access control - **Actions**: Edit or delete user
Superset users list
Superset user management
### User Integration with Hopsworks Superset users are automatically synchronized with Hopsworks project members. When users are added to a project, their accounts are created automatically with appropriate permissions. For each Hopsworks user in a project, a corresponding Superset user is created with the naming pattern `___superset`. This per-user-per-project mapping ensures proper data isolation and access control between projects. ## Managing Roles Roles define the permissions and capabilities available to Superset users. Navigate to **Settings** → **Roles** (or by navigating to `https:///hopsworks-api/superset/roles/`) to manage role configurations. ### Default Roles Superset in Hopsworks includes several pre-configured roles: - **Admin**: Full administrative access to all Superset features - **Alpha**: Can create and edit all content types - **Gamma**: Read-only access to dashboards and charts - **sql_lab**: Access to SQL Lab for query execution - **Public**: Minimal read-only access for public dashboards - **Examples**: Role for example dashboards and datasets (added only if `superset.loadExamples` is set to `true`) - **Dataset**: Role for dataset management
Superset roles list
Superset role management
### Project- and User-Specific Roles Each Hopsworks project user has a dedicated Superset role with the naming pattern `HW_role____superset` that controls access to that project's data sources. These roles are automatically created and managed by Hopsworks to maintain proper access control and data isolation. ## Database Connections Superset uses database connections to access data from the Hopsworks Feature Store. Database connections are created automatically based on project configuration and Feature Store setup. Navigate to **Settings** → **Database connections** in the Superset interface to view and manage database connections.
Superset database connections
Database connections in Superset
### Connection Types Hopsworks projects can have up to two database connections, created based on Feature Store configuration: #### Online Feature Store Connection (MySQL) - **Naming pattern**: `___superset` - **Backend**: MySQL - **Purpose**: Access to the Online Feature Store for exploration - **Use cases**: Real-time dashboards, monitoring online features - **Created when**: An Online Feature Store is created in the project #### Offline Feature Store Connection (Trino) - **Naming pattern**: `Trino_____superset` - **Backend**: Trino - **Purpose**: Access to the Offline Feature Store for analytical queries - **Use cases**: Historical analysis, training data exploration, complex aggregations - **Created when**: Trino is enabled for the project - **Multi-catalog**: Enabled. The connection is created with Superset's **Allow changing catalogs** option turned on, so users can pick any Trino catalog (`hive`, `delta`, `iceberg`, `hudi`) from the **Catalog** selector in SQL Lab and when adding a dataset, rather than being limited to the one set in [`trino_default_catalog`][trino_default_catalog]. ### Connection Properties Database connections support various capabilities indicated in the connections list: - **AQE (Asynchronous Query Execution)**: Whether the connection supports async queries (false by default) - **DML (Data Manipulation Language)**: Whether INSERT/UPDATE/DELETE operations are allowed (true by default) - **File upload**: Whether data can be uploaded directly through this connection (false by default) - **Expose in SQL Lab**: Whether the connection is available in SQL Lab interface (true by default) ### Managing Connections Database connections are automatically created and configured by Hopsworks. Manual connection management is typically not required, but administrators can: - View connection details and properties - Monitor connection usage and performance - Test connection availability - Modify advanced connection settings if needed ## Configuration Superset behavior can be customized through Hopsworks cluster configuration variables. In **Cluster Settings**, choose **Configuration** under _Infrastructure_ in the left sidebar and search for `superset` to view all Superset-related variables.
Superset configuration variables
Superset configuration in Hopsworks cluster settings
### Available Variables #### superset_admin_roles - **Description**: Comma-separated list of Superset roles Hopsworks Admins should be assigned. - **Default**: `Admin` - **Format**: Comma-separated list of role names - Superset roles listed here are assigned to Hopsworks Admins. #### superset_enabled - **Description**: Enable or disable Superset cluster-wide - **Default**: `false` - **Values**: `true` or `false` - When set to `false`, Superset becomes unavailable for all projects across the cluster. #### superset_proxy_connect_timeout_ms - **Description**: Time the Hopsworks proxy waits to establish a connection to Superset, in milliseconds. - **Default**: `10000` - **Format**: Integer #### superset_proxy_connection_request_timeout_ms - **Description**: Time the Hopsworks proxy waits for a connection from its pool, in milliseconds. - **Default**: `10000` - **Format**: Integer #### superset_proxy_max_connections - **Description**: Maximum number of connections the Hopsworks proxy keeps open to Superset. - **Default**: `50` - **Format**: Integer #### superset_proxy_read_timeout_ms - **Description**: Time the Hopsworks proxy waits for a response from Superset, in milliseconds. - **Default**: `180000` - **Format**: Integer - Raise this if long running SQL Lab queries are cut off by the proxy. #### superset_user_roles - **Description**: Default roles automatically assigned to new Superset users - **Default**: `Gamma,sql_lab,Dataset` - **Format**: Comma-separated list of role names - These roles determine the initial permissions granted to users when they first access Superset from a project. #### trino_default_catalog - **Description**: Default catalog to use for the Offline Feature Store Connection - **Default**: `delta` - **Values**: `hive`, `delta`, `iceberg`, and `hudi`. The value is applied when a member's connection is created, so connections created before a change keep the catalog they were created with. Trino connections are created with multi-catalog enabled (Superset's **Allow changing catalogs** option), so users are not limited to this default. They select the catalog matching their table format from the **Catalog** dropdown in SQL Lab or when adding a dataset. Querying a table whose format does not match the selected catalog raises a Trino `UNSUPPORTED_TABLE_TYPE` error; see the [Trino table type error][trino-table-type-error] troubleshooting in the Superset user guide. ### Applying Configuration Changes After modifying any configuration variable: 1. Click **Save** next to the variable to save the change 2. Verify the change has taken effect by checking Superset behavior 3. Monitor Superset logs for any configuration-related errors ### Configuration Best Practices - **Limited admin access**: Only grant `superset_admin_roles` to trusted administrators - **Role defaults**: Set `superset_user_roles` to provide appropriate baseline permissions - **Testing changes**: Test configuration changes in a development environment first - **Documentation**: Document any custom configuration changes and update your Helm values accordingly. ================================================================================ # Helm chart values reference Source: https://docs.hopsworks.ai/latest/setup_installation/common/helm_chart_values/ # Helm chart values reference This section lists every value you can configure when deploying Hopsworks with the Hopsworks Helm chart, with one page per top-level key. It is generated from the `README.md` files of the chart and its subcharts, and for RonDB from the `values.schema.json` of the RonDB chart that Hopsworks pins. On a released version of the docs it matches the Hopsworks Helm chart for that release; on the development docs it reflects the latest chart published to the development channel. You set these values in the `values..yaml` file that you pass to `helm install`. For a guided, end-to-end setup, follow one of the cloud installation guides, such as the [AWS getting started guide][aws-getting-started-with-eks]. Only a small subset of these values is needed for a typical install: "Common values" below lists the ones the cloud installation guides set, and the pages are the exhaustive reference. "All values" lists the pages. Each page states when its subchart is deployed: when the condition names several values, Helm uses the first one that is set. "Upstream charts" are the charts a subchart installs from other Helm repositories, and each link opens a chart's documentation for the version Hopsworks pins. A page lists only the values Hopsworks sets for its upstream charts, except the RonDB page, which lists all of the RonDB chart's values. A default is the value Hopsworks deploys: the root chart's settings for a subchart are applied over the subchart's own defaults. Every value has its own link (the `#` next to its key), and each section has its defaults as a values file under "Defaults as YAML". _Generated from the Hopsworks Helm chart `5.2.0-alpha-1791549041` (Hopsworks `5.2.0`)._ ## Common values { #helm-values-common } | Value | What it sets | | --- | --- | | [`global._hopsworks.cloudProvider`](global.md#helm.global._hopsworks.cloudProvider) | The cloud the cluster runs on; HopsFS, Consul and Hopsworks configure themselves from it. | | [`global._hopsworks.storageClassName`](global.md#helm.global._hopsworks.storageClassName) | The storage class of every persistent volume. | | [`global._hopsworks.imageRegistry`](global.md#helm.global._hopsworks.imageRegistry) | The registry the Hopsworks images are pulled from. | | [`global._hopsworks.imagePullSecrets`](global.md#helm.global._hopsworks.imagePullSecrets) | The pull secrets for that registry. | | [`global._hopsworks.managedDockerRegistery`](global.md#helm.global._hopsworks.managedDockerRegistery) | Use the cloud provider's container registry for the images Hopsworks builds for users. | | [`global._hopsworks.managedObjectStorage`](global.md#helm.global._hopsworks.managedObjectStorage) | Use a cloud bucket for HopsFS data and for the RonDB and OpenSearch backups. | | [`global._hopsworks.minio.enabled`](global.md#helm.global._hopsworks.minio.enabled) | Deploy MinIO in the cluster; turn it off when a cloud bucket is used. | | [`global._hopsworks.externalLoadBalancers.enabled`](global.md#helm.global._hopsworks.externalLoadBalancers.enabled) | Expose services through LoadBalancer Services. | | [`hopsworks.ingress.host`](hopsworks.md#helm.hopsworks.ingress.host) | The host name of the Hopsworks UI and API. | | [`hopsworks.ingress.ingressClassName`](hopsworks.md#helm.hopsworks.ingress.ingressClassName) | The ingress controller that serves it. | | [`hopsworks.velero.backup.enabled`](hopsworks.md#helm.hopsworks.velero.backup.enabled) | Back up Kubernetes resources with Velero, which needs its own install. | | [`rondb.rondb.clusterSize`](rondb.md#helm.rondb.rondb.clusterSize) | The size of the RonDB cluster: data replicas, node groups, MySQL and REST API servers. | ## All values { #helm-values-pages } | Values | Upstream charts | Keys | | --- | --- | --- | | [`global`][helm-values-global] | | 145 | | [`airflow`][helm-values-airflow] | | 151 | | [`arrowflight`][helm-values-arrowflight] | | 46 | | [`certs-operator`][helm-values-certs-operator] | | 29 | | [`consul`][helm-values-consul] | [`consul` 1.8.16](https://artifacthub.io/packages/helm/hashicorp/consul/1.8.16) | 24 | | [`docker-registry`][helm-values-docker-registry] | | 90 | | [`grafana`][helm-values-grafana] | [`grafana` 7.0.17](https://artifacthub.io/packages/helm/grafana/grafana/7.0.17) | 7 | | [`hive`][helm-values-hive] | | 78 | | [`hopsfs`][helm-values-hopsfs] | | 221 | | [`hopsfs-csi`][helm-values-hopsfs-csi] | | 29 | | [`hopsworks`][helm-values-hopsworks] | | 934 | | [`hw-kueue`][helm-values-hw-kueue] | [`kueue` 0.12.2](https://github.com/kubernetes-sigs/kueue/blob/v0.12.2/charts/kueue/README.md) | 94 | | [`hw-kyverno`][helm-values-hw-kyverno] | | 58 | | [`judge`][helm-values-judge] | | 23 | | [`kafka`][helm-values-kafka] | [`strimzi-kafka-operator` 1.2.0](https://artifacthub.io/packages/helm/strimzi/strimzi-kafka-operator/1.2.0) | 102 | | [`kserve`][helm-values-kserve] | | 266 | | [`minio`][helm-values-minio] | | 42 | | [`olk`][helm-values-olk] | [`prometheus-elasticsearch-exporter` 5.8.0](https://artifacthub.io/packages/helm/prometheus-community/prometheus-elasticsearch-exporter/5.8.0) | 164 | | [`onlinefs`][helm-values-onlinefs] | | 93 | | [`prometheus`][helm-values-prometheus] | [`prometheus` 25.20.2](https://artifacthub.io/packages/helm/prometheus-community/prometheus/25.20.2), [`prometheus-adapter` 4.11.0](https://artifacthub.io/packages/helm/prometheus-community/prometheus-adapter/4.11.0) | 9 | | [`ray`][helm-values-ray] | [`kuberay-operator` 1.4.0](https://artifacthub.io/packages/helm/kuberay-operator/kuberay-operator/1.4.0) | 2 | | [`rondb`][helm-values-rondb] | [`rondb` 26.2.21](https://github.com/logicalclocks/rondb-helm/blob/v26.2.21/values.schema.json) | 464 | | [`spark`][helm-values-spark] | [`spark-operator` 2.5.1](https://github.com/kubeflow/spark-operator/blob/v2.5.1/charts/spark-operator-chart/README.md) | 160 | | [`superset`][helm-values-superset] | [`mysql` 12.3.5](https://artifacthub.io/packages/helm/bitnami/mysql/12.3.5), [`superset` 0.15.0](https://artifacthub.io/packages/helm/superset/superset/0.15.0) | 25 | | [`trino`][helm-values-trino] | [`trino` 1.41.0](https://artifacthub.io/packages/helm/trino/trino/1.41.0) | 38 | | [`vpa`][helm-values-vpa] | | 18 | ## Other values { #helm-values-other } ??? example "Defaults as YAML" ```yaml hopsworkslib: {} ```
`hopsworkslib` # { #helm.hopsworkslib } : Type `object`, default `{}`. override hopsworkslib values
================================================================================ # global Source: https://docs.hopsworks.ai/latest/setup_installation/common/helm_chart_values/global/ # Global values { #helm-values-global } Values under `global` are shared by every subchart: image registry and pull secrets, storage class, scheduling, cloud provider and which optional services are enabled. _Generated from the Hopsworks Helm chart `5.2.0-alpha-1791549041` (Hopsworks `5.2.0`)._ ## General { #helm-values-global-general } ??? example "Defaults as YAML" ```yaml global: _hopsworks: airflow: enabled: true keysSecretName: hopsworks-airflow-keys airflowApiKeySecretName: airflow-api-key autoscalers: {} buildkitd: enabled: false centralNamespace: hopsworks cloudProvider: '' consulDomainName: consul csi: driverName: '' enabled: true image: pullPolicy: IfNotPresent repository: hopsworks/hopsfs-csi tag: 0.1.0-SNAPSHOT executor_uid: 1235 full_platform: true grafana: extraDashboardProviders: [] imagePullPolicy: IfNotPresent imagePullSecrets: [] imageRegistry: docker.hops.works initContainerResources: runtime: limits: cpu: 1 memory: 1Gi requests: cpu: 250m memory: 512Mi tool: limits: cpu: 500m memory: 512Mi requests: cpu: 100m memory: 128Mi waiter: limits: cpu: 500m memory: 256Mi requests: cpu: 50m memory: 64Mi jobs: ttlSecondsAfterFinished: 86400 kafka: enabled: true kueue: enabled: false managedObjectStorage: enabled: false s3: null mode: auto mysql: hopsworksUser: hopsworksroot usersSecretname: mysql-users-secrets networkPolicy: rondbAccessLabels: access: mgmd-and-ndbmtd nodeSelector: {} onlinefs: email: onlinefs@hopsworks.ai password: onlinefspw opensearch: enabled: true openshift: enabled: false ray: enabled: false security: tls: enabled: true securityContextEnabled: true serviceAccount: annotations: {} create: true name: hopsworks-service-account serviceAccountAnnotations: {} skipDatabaseMigration: false spark: history: enabled: true storageClassName: null superset: enabled: true mysql: enabled: true redis: enabled: true tolerations: [] toolbox: image: hwutils tag: 1.10-SNAPSHOT topologySpreadConstraint: maxSkew: 1 nodeAffinityPolicy: Honor nodeTaintsPolicy: Honor topologyKey: topology.kubernetes.io/zone whenUnsatisfiable: ScheduleAnyway vpaEnabled: false wipeDataOnUninstall: true _kserve: servingruntime: vllmomni: tag: v0.28.0 vllmopenai: tag: v0.28.0 imageDigests: {} unmanagedLoadBalancers: {} ```
`global._hopsworks.airflow.enabled` # { #helm.global._hopsworks.airflow.enabled } : Type `bool`, default `true`. Enable or disable the installation of the airflow sub chart `global._hopsworks.airflow.keysSecretName` # { #helm.global._hopsworks.airflow.keysSecretName } : Type `string`, default `"hopsworks-airflow-keys"`. Name of the Secret holding the shared-bearer secret that hopsworks-instance uses to call the Airflow `/auth/internal/*` routes. Must match airflow.airflowApi.keysSecretName so the same Secret is mounted on both sides. The default `airflow-webserver-airflow-crypto-material` is the cert-only secret and does NOT contain `internal-shared-secret`, so the hopsworks-instance pod fails to mount on a fresh v3 install. `global._hopsworks.airflowApiKeySecretName` # { #helm.global._hopsworks.airflowApiKeySecretName } : Type `string`, default `"airflow-api-key"`. `global._hopsworks.autoscalers` # { #helm.global._hopsworks.autoscalers } : Type `object`, default `{}`. Map of autoscaler name to VPA configuration. Each key becomes a VerticalPodAutoscaler resource named -vpa. `global._hopsworks.buildkitd.enabled` # { #helm.global._hopsworks.buildkitd.enabled } : Type `bool`, default `false`. `global._hopsworks.centralNamespace` # { #helm.global._hopsworks.centralNamespace } : Type `string`, default `"hopsworks"`. Namespace to fetch OLK signing key from. Set to null to generate a new key. `global._hopsworks.cloudProvider` # { #helm.global._hopsworks.cloudProvider } : Type `string`, default `""`. cloud provider (AWS, AZURE, GCP, OVH). It is used by HopsFS, Consul, and Hopsworks. Hopsfs uses it to configure the object storage parameters. Consul uses it to configure coredns accordingly. Hopsworks uses it to determine how to store the users docker images within the same hopsworks-base repo as tags or in different repo per project, If cloud provider is set all the users docker images are stored as tags - it is an AWS limitation where only tags within the repo can share layers. `global._hopsworks.consulDomainName` # { #helm.global._hopsworks.consulDomainName } : Type `string`, default `"consul"`. The domain name for consul where it will answer DNS queries, e.g. `service-name.service.consul`. If changed, make sure to update consul.consul.global.domain to the same value. `global._hopsworks.csi` # { #helm.global._hopsworks.csi } : Type `object`. HopsFS CSI integration: the hopsfs-csi node plugin FUSE-mounts an inline ephemeral volume on every HopsFS-mounting pod (Hopsworks-generated workloads, Airflow, Trino) and hands the descriptor to an unprivileged `hopsfs-fuse` native sidecar in the pod, so no workload container is privileged. `enabled` installs the hopsfs-csi DaemonSet and CSIDriver and seeds `csi_driver_enabled` to the backend; requires Kubernetes >= 1.29 (native sidecars), which the render enforces with a message naming this value. Set it to false to keep the legacy privileged in-container HopsFS mount everywhere. Known limitation: a crashed `hopsfs-fuse` sidecar leaves that pod's mount dead (I/O fails with ENOTCONN) until the pod is recreated; there is no self-heal. `image` is the ONE place the hopsfs-csi image is named: the node plugin, the Airflow and Trino sidecars and the backend `csi_sidecar_image` variable are all built from `imageRegistry` + this block, because the fusermount3 proxy in the sidecar and the fd server in the plugin are two halves of one protocol and must never drift on a node. Pin a release tag here before GA. ??? note "Default" ```yaml driverName: '' enabled: true image: pullPolicy: IfNotPresent repository: hopsworks/hopsfs-csi tag: 0.1.0-SNAPSHOT ``` `global._hopsworks.csi.driverName` # { #helm.global._hopsworks.csi.driverName } : Type `string`, default `""`. The CSI driver name of this release. Empty (the default) derives `.hopsfs.csi.logicalclocks.com` from the release namespace, one name per release, which is what lets two Hopsworks installs share a cluster: kubelet routes CSI calls by driver name alone, so each release gets its own node plugin, socket, registration and CSIDriver object, also on nodes both installs use, and uninstalling one never touches the other's mounts (HWORKS-3297). Kubernetes caps a CSI driver name at 63 characters, so a namespace longer than 34 characters must set a shorter name here. Any name set here must be unique in the cluster: two releases on one name share one socket directory and one kubelet registration on every node they share, so neither is isolated from the other, which is the fault the per-release name removes; that is not a supported setup. Whatever the name, it is computed once (`hopsworkslib.csiDriverName`) for the plugin, the Airflow and Trino volumes and the backend `csi_driver_name` variable, so nothing can drift. `global._hopsworks.executor_uid` # { #helm.global._hopsworks.executor_uid } : Type `int`, default `1235`. User ID for the user running Hopsworks and Airflow containers `global._hopsworks.full_platform` # { #helm.global._hopsworks.full_platform } : Type `bool`, default `true`. Flag to indicate if the full platform is installed or just the online feature store infrastructure `global._hopsworks.grafana.extraDashboardProviders` # { #helm.global._hopsworks.grafana.extraDashboardProviders } : Type `list`, default `[]`. Extra Grafana dashboard providers, appended to the provider file the grafana chart renders. The dashboards themselves ship in the Grafana image, so this is the way to provision one without rebuilding the image: mount it (for example through `grafana.grafana.dashboardsConfigMaps`, which lands at /var/lib/grafana/dashboards/) and point a provider at the mount path. Entries are Grafana provider objects, so each needs at least `name`, `folder`, `type: file` and `options.path`; provider names must be unique across all providers. Paths under /usr/share/grafana/dashboards are the image's own and are verified by the verify-dashboards init container, so point elsewhere. Setting this changes the provider ConfigMap's content hash, which rolls the Grafana pod. `global._hopsworks.imagePullPolicy` # { #helm.global._hopsworks.imagePullPolicy } : Type `string`, default `"IfNotPresent"`. `global._hopsworks.imagePullSecrets` # { #helm.global._hopsworks.imagePullSecrets } : Type `list`, default `[]`. image pull secrets to be used by all the subcharts. Notice that for subcharts the have external dependices, you need to update the image pull secrets accordingly in those subcharts `global._hopsworks.imageRegistry` # { #helm.global._hopsworks.imageRegistry } : Type `string`, default `"docker.hops.works"`. `global._hopsworks.initContainerResources` # { #helm.global._hopsworks.initContainerResources } : Type `object`. Resource requests and limits for the chart's init containers, in three tiers by what the container does: `waiter` for shell wait loops, `tool` for file copies and small binaries, `runtime` for anything starting a JVM or a Python interpreter or pulling images. Every init container declares both requests and limits. A namespace `LimitRange` that supplies a `default` limit injects it into any container declaring none, and a pod's effective request is `max(sum of app containers, highest single init container)` where the init term is a floor for the pod's whole lifetime. An init container without resources can therefore pin a node's allocatable to the `LimitRange` default even after it has exited. Retune these if the target namespace has a `LimitRange` whose `min`/`max` would reject the defaults, since an out-of-range explicit value is rejected rather than defaulted. ??? note "Default" ```yaml runtime: limits: cpu: 1 memory: 1Gi requests: cpu: 250m memory: 512Mi tool: limits: cpu: 500m memory: 512Mi requests: cpu: 100m memory: 128Mi waiter: limits: cpu: 500m memory: 256Mi requests: cpu: 50m memory: 64Mi ``` `global._hopsworks.jobs` # { #helm.global._hopsworks.jobs } : Type `object`, default `{"ttlSecondsAfterFinished":86400}`. Global configuration for Kubernetes Jobs created by the chart `global._hopsworks.jobs.ttlSecondsAfterFinished` # { #helm.global._hopsworks.jobs.ttlSecondsAfterFinished } : Type `int`, default `86400`. Time in seconds after a finished Job is eligible for automatic cleanup. Applies to all Jobs unless overridden by a subchart-specific ttlSecondsAfterFinished value. `global._hopsworks.kafka.enabled` # { #helm.global._hopsworks.kafka.enabled } : Type `bool`, default `true`. Enable or disable the installation of the kafka sub chart. `global._hopsworks.kueue.enabled` # { #helm.global._hopsworks.kueue.enabled } : Type `bool`, default `false`. `global._hopsworks.managedObjectStorage` # { #helm.global._hopsworks.managedObjectStorage } : Type `object`, default `{"enabled":false,"s3":null}`. Configuration for managed object storage to be used by HopsFS, Opensearch, RonDB, and Hopsworks. Opensearch and RonDB uses this configuration to setup a remote sink for their backup, if a different remote sink is configured on the subchart then it will take precedence. Hopsworks uses the bucket configuration as context cache when building users' docker images, only S3 is supported at the moment. `global._hopsworks.managedObjectStorage.s3` # { #helm.global._hopsworks.managedObjectStorage.s3 } : Type `string`, default `nil`. S3 configuration `global._hopsworks.mode` # { #helm.global._hopsworks.mode } : Type `string`, default `"auto"`. Helm installation model. "auto" lets the chart decide based on the Release.IsInstall value. "install" forces installation mode, while "upgrade" forces upgrade mode. `global._hopsworks.mysql.hopsworksUser` # { #helm.global._hopsworks.mysql.hopsworksUser } : Type `string`, default `"hopsworksroot"`. `global._hopsworks.mysql.usersSecretname` # { #helm.global._hopsworks.mysql.usersSecretname } : Type `string`, default `"mysql-users-secrets"`. `global._hopsworks.networkPolicy.rondbAccessLabels.access` # { #helm.global._hopsworks.networkPolicy.rondbAccessLabels.access } : Type `string`, default `"mgmd-and-ndbmtd"`. `global._hopsworks.nodeSelector` # { #helm.global._hopsworks.nodeSelector } : Type `object`, default `{}`. Specifies the global nodeSelector settings applied across all subcharts unless explicitly overridden within a specific subchart. This ensures Kubernetes schedules Pods only onto nodes that match all the specified labels. Notice that some subcharts do not use this global variable, and you must manually override those by defining them in the values.yaml file, using anchors if necessary. `global._hopsworks.onlinefs.email` # { #helm.global._hopsworks.onlinefs.email } : Type `string`, default `"onlinefs@hopsworks.ai"`. `global._hopsworks.onlinefs.password` # { #helm.global._hopsworks.onlinefs.password } : Type `string`, default `"onlinefspw"`. `global._hopsworks.opensearch` # { #helm.global._hopsworks.opensearch } : Type `object`, default `{"enabled":true}`. Enable or disable the opensearch `global._hopsworks.opensearch.enabled` # { #helm.global._hopsworks.opensearch.enabled } : Type `bool`, default `true`. Enable or disable the opensearch `global._hopsworks.openshift.enabled` # { #helm.global._hopsworks.openshift.enabled } : Type `bool`, default `false`. Enable when installing on Openshift platform `global._hopsworks.ray.enabled` # { #helm.global._hopsworks.ray.enabled } : Type `bool`, default `false`. `global._hopsworks.security.tls.enabled` # { #helm.global._hopsworks.security.tls.enabled } : Type `bool`, default `true`. `global._hopsworks.securityContextEnabled` # { #helm.global._hopsworks.securityContextEnabled } : Type `bool`, default `true`. Flag to disable templating SecurityContext for Openshift `global._hopsworks.serviceAccount.annotations` # { #helm.global._hopsworks.serviceAccount.annotations } : Type `object`, default `{}`. custom annotations for the Hopsworks service account `global._hopsworks.serviceAccount.create` # { #helm.global._hopsworks.serviceAccount.create } : Type `bool`, default `true`. `global._hopsworks.serviceAccount.name` # { #helm.global._hopsworks.serviceAccount.name } : Type `string`, default `"hopsworks-service-account"`. `global._hopsworks.serviceAccountAnnotations` # { #helm.global._hopsworks.serviceAccountAnnotations } : Type `object`, default `{}`. Use it to annotate the serviceAccounts we create for Hopsworks `global._hopsworks.skipDatabaseMigration` # { #helm.global._hopsworks.skipDatabaseMigration } : Type `bool`, default `false`. Special flag which MUST be used only for 3.x -> 4.0 migrations (HWORKS-1600) `global._hopsworks.spark.history.enabled` # { #helm.global._hopsworks.spark.history.enabled } : Type `bool`, default `true`. Enable or disable installing of the spark history server `global._hopsworks.storageClassName` # { #helm.global._hopsworks.storageClassName } : Type `string`, default `nil`. global storage class name `global._hopsworks.superset.enabled` # { #helm.global._hopsworks.superset.enabled } : Type `bool`, default `true`. Enable or disable the installation of the superset sub chart. `global._hopsworks.superset.mysql.enabled` # { #helm.global._hopsworks.superset.mysql.enabled } : Type `bool`, default `true`. Enable or disable MySQL for Superset. Must match superset.mysql.enabled for consistent behavior across charts. `global._hopsworks.superset.redis.enabled` # { #helm.global._hopsworks.superset.redis.enabled } : Type `bool`, default `true`. Enable or disable Redis for Superset. Must match superset.superset.redis.enabled for consistent behavior across charts. `global._hopsworks.tolerations` # { #helm.global._hopsworks.tolerations } : Type `list`, default `[]`. Specifies the global tolerations settings applied to all subcharts unless explicitly overridden in a specific subchart. These tolerations allow Kubernetes to schedule Pods on nodes with matching taints, ensuring proper placement based on cluster policies. Notice that some subcharts do not use this global variable, and you must manually override those by defining them in the values.yaml file, using anchors if necessary. `global._hopsworks.toolbox.image` # { #helm.global._hopsworks.toolbox.image } : Type `string`, default `"hwutils"`. `global._hopsworks.toolbox.tag` # { #helm.global._hopsworks.toolbox.tag } : Type `string`, default `"1.10-SNAPSHOT"`. `global._hopsworks.topologySpreadConstraint` # { #helm.global._hopsworks.topologySpreadConstraint } : Type `object`. default topology spread constraint. If not defined the global topology spread constraint would be used ??? note "Default" ```yaml maxSkew: 1 nodeAffinityPolicy: Honor nodeTaintsPolicy: Honor topologyKey: topology.kubernetes.io/zone whenUnsatisfiable: ScheduleAnyway ``` `global._hopsworks.vpaEnabled` # { #helm.global._hopsworks.vpaEnabled } : Type `bool`, default `false`. `global._hopsworks.wipeDataOnUninstall` # { #helm.global._hopsworks.wipeDataOnUninstall } : Type `bool`, default `true`. on uninstall, delete data PVCs left behind by StatefulSets unless the PVC carries the label hopsworks.ai/keep=true. Consumed by subchart post-delete cleanup hooks via the hopsworkslib.wipeDataOnUninstall helper. `global._kserve` # { #helm.global._kserve } : Type `object`. Global KServe values shared between the kserve and hopsworks subcharts ??? note "Default" ```yaml servingruntime: vllmomni: tag: v0.28.0 vllmopenai: tag: v0.28.0 ``` `global._kserve.servingruntime` # { #helm.global._kserve.servingruntime } : Type `object`, default `{"vllmomni":{"tag":"v0.28.0"},"vllmopenai":{"tag":"v0.28.0"}}`. Mirrors the kserve subchart's `kserve.servingruntime` layout. Tags here drive both the kserve ClusterServingRuntime images and the kube_serving_vllm*_versions hopsworks variable seeds. `global._kserve.servingruntime.vllmomni` # { #helm.global._kserve.servingruntime.vllmomni } : Type `object`, default `{"tag":"v0.28.0"}`. vLLM-Omni runtime image tag. Drives both the kserve ClusterServingRuntime image and the kube_serving_vllm_omni_versions hopsworks variable seed. `global._kserve.servingruntime.vllmopenai` # { #helm.global._kserve.servingruntime.vllmopenai } : Type `object`, default `{"tag":"v0.28.0"}`. vLLM-OpenAI runtime image tag. Drives both the kserve ClusterServingRuntime image and the kube_serving_vllm_versions hopsworks variable seed. `global.imageDigests` # { #helm.global.imageDigests } : Type `object`, default `{}`. map image name to sha digest to be used instead of tags for reproducible deployment `global.unmanagedLoadBalancers` # { #helm.global.unmanagedLoadBalancers } : Type `object`, default `{}`. Load balancer configuration when using unmanaged LB, in AWS is the TargetGroup ARNs for each service
## backups { #helm-values-global-backups } ??? example "Defaults as YAML" ```yaml global: _hopsworks: backups: enabled: true metadataStore: configMap: opensearch: opensearch-backups-metadata ronDB: rondb-backups-metadata superset: superset-backups-metadata schedule: '@weekly' ttl: null ```
`global._hopsworks.backups` # { #helm.global._hopsworks.backups } : Type `object`. enable global backups configuration ??? note "Default" ```yaml enabled: true metadataStore: configMap: opensearch: opensearch-backups-metadata ronDB: rondb-backups-metadata superset: superset-backups-metadata schedule: '@weekly' ttl: null ``` `global._hopsworks.backups.enabled` # { #helm.global._hopsworks.backups.enabled } : Type `bool`, default `true`. enable global backup `global._hopsworks.backups.metadataStore` # { #helm.global._hopsworks.backups.metadataStore } : Type `object`. backups metadata store ??? note "Default" ```yaml configMap: opensearch: opensearch-backups-metadata ronDB: rondb-backups-metadata superset: superset-backups-metadata ``` `global._hopsworks.backups.metadataStore.configMap.opensearch` # { #helm.global._hopsworks.backups.metadataStore.configMap.opensearch } : Type `string`, default `"opensearch-backups-metadata"`. name of the configmap to store metadata information about opensearch backups `global._hopsworks.backups.metadataStore.configMap.ronDB` # { #helm.global._hopsworks.backups.metadataStore.configMap.ronDB } : Type `string`, default `"rondb-backups-metadata"`. name of the configmap to store metadata information about rondb backups `global._hopsworks.backups.metadataStore.configMap.superset` # { #helm.global._hopsworks.backups.metadataStore.configMap.superset } : Type `string`, default `"superset-backups-metadata"`. name of the configmap to store metadata information about superset backups `global._hopsworks.backups.schedule` # { #helm.global._hopsworks.backups.schedule } : Type `string`, default `"@weekly"`. cron schedule `global._hopsworks.backups.ttl` # { #helm.global._hopsworks.backups.ttl } : Type `string`, default `nil`. time to live to control when to clean up backups. It is a number followed by either d (days) or h (hours) suffix.
## externalLoadBalancers { #helm-values-global-externalloadbalancers } ??? example "Defaults as YAML" ```yaml global: _hopsworks: externalLoadBalancers: annotations: {} class: null enabled: true managed: true ```
`global._hopsworks.externalLoadBalancers` # { #helm.global._hopsworks.externalLoadBalancers } : Type `object`, default `{"annotations":{},"class":null,"enabled":true,"managed":true}`. Global load balancer configuration for external access to Hopsworks services: ArrowFlight, Kafka, and MySQL We fallback to this loadBalancerClass if the local loadBalancerClass is not defined `global._hopsworks.externalLoadBalancers.annotations` # { #helm.global._hopsworks.externalLoadBalancers.annotations } : Type `object`, default `{}`. Generic annotations attached to LoadBalancer objects `global._hopsworks.externalLoadBalancers.class` # { #helm.global._hopsworks.externalLoadBalancers.class } : Type `string`, default `nil`. Name of the LoadBalancer class `global._hopsworks.externalLoadBalancers.enabled` # { #helm.global._hopsworks.externalLoadBalancers.enabled } : Type `bool`, default `true`. Enable LoadBalancer Services `global._hopsworks.externalLoadBalancers.managed` # { #helm.global._hopsworks.externalLoadBalancers.managed } : Type `bool`, default `true`. Cloud provider provisions Load Balancers
## externalServices { #helm-values-global-externalservices } ??? example "Defaults as YAML" ```yaml global: _hopsworks: externalServices: hopsfs: external: false namenodeAddresses: [] opensearch: addresses: [] external: false prometheus: addresses: [] external: false rondb: external: false mgmdHostname: '' ```
`global._hopsworks.externalServices` # { #helm.global._hopsworks.externalServices } : Type `object`. Configuration for when services are external to this Kubernetes installation ??? note "Default" ```yaml hopsfs: external: false namenodeAddresses: [] opensearch: addresses: [] external: false prometheus: addresses: [] external: false rondb: external: false mgmdHostname: '' ``` `global._hopsworks.externalServices.hopsfs` # { #helm.global._hopsworks.externalServices.hopsfs } : Type `object`, default `{"external":false,"namenodeAddresses":[]}`. HopsFS configuration when it is installed externally `global._hopsworks.externalServices.hopsfs.external` # { #helm.global._hopsworks.externalServices.hopsfs.external } : Type `bool`, default `false`. Flag to indicate HopsFS is installed externally `global._hopsworks.externalServices.hopsfs.namenodeAddresses` # { #helm.global._hopsworks.externalServices.hopsfs.namenodeAddresses } : Type `list`, default `[]`. IP addresses where HopsFS Namenodes are installed `global._hopsworks.externalServices.opensearch` # { #helm.global._hopsworks.externalServices.opensearch } : Type `object`, default `{"addresses":[],"external":false}`. Opensearch configuration when it is installed externally `global._hopsworks.externalServices.opensearch.addresses` # { #helm.global._hopsworks.externalServices.opensearch.addresses } : Type `list`, default `[]`. IP addresses where OpenSearch is installed `global._hopsworks.externalServices.opensearch.external` # { #helm.global._hopsworks.externalServices.opensearch.external } : Type `bool`, default `false`. Flag to indicate Opensearch is installed externally `global._hopsworks.externalServices.prometheus` # { #helm.global._hopsworks.externalServices.prometheus } : Type `object`, default `{"addresses":[],"external":false}`. prometheus configuration when it is installed externally `global._hopsworks.externalServices.prometheus.addresses` # { #helm.global._hopsworks.externalServices.prometheus.addresses } : Type `list`, default `[]`. IP addresses where prometheus is installed `global._hopsworks.externalServices.prometheus.external` # { #helm.global._hopsworks.externalServices.prometheus.external } : Type `bool`, default `false`. Flag to indicate prometheus is installed externally `global._hopsworks.externalServices.rondb` # { #helm.global._hopsworks.externalServices.rondb } : Type `object`, default `{"external":false,"mgmdHostname":""}`. RonDB configuration when it is installed externally `global._hopsworks.externalServices.rondb.external` # { #helm.global._hopsworks.externalServices.rondb.external } : Type `bool`, default `false`. Flag to indicate RonDB is installed externally `global._hopsworks.externalServices.rondb.mgmdHostname` # { #helm.global._hopsworks.externalServices.rondb.mgmdHostname } : Type `string`, default `""`. Hostname of the machine where RonDB management service is running
## kyverno { #helm-values-global-kyverno } ??? example "Defaults as YAML" ```yaml global: _hopsworks: kyverno: enabled: false policies: addCertificatesVolume: enabled: false initContainers: annotation: key: kyverno-inject-certs-init value: enabled enabled: false mountPath: /etc/ssl/certs preconditions: annotation: key: kyverno-inject-certs value: enabled enabled: true ```
`global._hopsworks.kyverno.enabled` # { #helm.global._hopsworks.kyverno.enabled } : Type `bool`, default `false`. Enable or disable kyverno policies installation `global._hopsworks.kyverno.policies.addCertificatesVolume` # { #helm.global._hopsworks.kyverno.policies.addCertificatesVolume } : Type `object`. Configuration to add custom certificates to pods as a mounted volume ??? note "Default" ```yaml enabled: false initContainers: annotation: key: kyverno-inject-certs-init value: enabled enabled: false mountPath: /etc/ssl/certs preconditions: annotation: key: kyverno-inject-certs value: enabled enabled: true ``` `global._hopsworks.kyverno.policies.addCertificatesVolume.enabled` # { #helm.global._hopsworks.kyverno.policies.addCertificatesVolume.enabled } : Type `bool`, default `false`. Enable add certificates volume `global._hopsworks.kyverno.policies.addCertificatesVolume.initContainers` # { #helm.global._hopsworks.kyverno.policies.addCertificatesVolume.initContainers } : Type `object`. Opt-in for injecting the certificates volume into init containers. The main addCertificatesVolume policy only mutates spec.containers; init containers are intentionally left untouched because they are often injected by third-party operators (Istio, KServe, sidecar injectors) that ship minimal images where overwriting /etc/ssl/certs would break them. When enabled, a second mutate rule is rendered that iterates spec.initContainers and is gated by an OR of the dedicated init-container annotation below and any extraAnnotations / labels configured under the hw-kyverno subchart at policies.addCertificatesVolume.initContainers. The annotation is distinct from preconditions.annotation so authors can opt main containers and init containers in independently. ??? note "Default" ```yaml annotation: key: kyverno-inject-certs-init value: enabled enabled: false ``` `global._hopsworks.kyverno.policies.addCertificatesVolume.initContainers.annotation` # { #helm.global._hopsworks.kyverno.policies.addCertificatesVolume.initContainers.annotation } : Type `object`, default `{"key":"kyverno-inject-certs-init","value":"enabled"}`. The key and value of the annotation used to opt a pod's init containers into certificate volume injection. Distinct from preconditions.annotation so authors can opt main containers and init containers in independently. A pod opts in by matching any one of: this annotation, an entry in the hw-kyverno subchart's policies.addCertificatesVolume.initContainers.extraAnnotations, or an entry in policies.addCertificatesVolume.initContainers.labels. `global._hopsworks.kyverno.policies.addCertificatesVolume.initContainers.enabled` # { #helm.global._hopsworks.kyverno.policies.addCertificatesVolume.initContainers.enabled } : Type `bool`, default `false`. Render the init-container mutate rule. When false (default), the policy never touches init containers regardless of any annotation, label, or extra annotation set on a pod. `global._hopsworks.kyverno.policies.addCertificatesVolume.mountPath` # { #helm.global._hopsworks.kyverno.policies.addCertificatesVolume.mountPath } : Type `string`, default `"/etc/ssl/certs"`. Path to mount the certificates volume `global._hopsworks.kyverno.policies.addCertificatesVolume.preconditions.annotation` # { #helm.global._hopsworks.kyverno.policies.addCertificatesVolume.preconditions.annotation } : Type `object`, default `{"key":"kyverno-inject-certs","value":"enabled"}`. The key and value of the annotation used to inject kyverno certificates. If a pod has this annotation, it will be mutated. There are other options under hw-kyverno subchart to use labels and extra annotations as needed. `global._hopsworks.kyverno.policies.addCertificatesVolume.preconditions.enabled` # { #helm.global._hopsworks.kyverno.policies.addCertificatesVolume.preconditions.enabled } : Type `bool`, default `true`. Enable adding preconditions to the police. If disabled, the policy will apply to all the pods in the installation namespace and the hopsworks projects' namespaces created when crating a project.
## managedDockerRegistery { #helm-values-global-manageddockerregistery } ??? example "Defaults as YAML" ```yaml global: _hopsworks: managedDockerRegistery: credHelper: enabled: false secretName: '' domain: '' enabled: false namespace: '' port: null ```
`global._hopsworks.managedDockerRegistery` # { #helm.global._hopsworks.managedDockerRegistery } : Type `object`. configure managed docker registry for Hopsworks to store the user's docker image ??? note "Default" ```yaml credHelper: enabled: false secretName: '' domain: '' enabled: false namespace: '' port: null ``` `global._hopsworks.managedDockerRegistery.credHelper` # { #helm.global._hopsworks.managedDockerRegistery.credHelper } : Type `object`, default `{"enabled":false,"secretName":""}`. credentials helper to use for the managed docker registry. We only support cred helpers for [AWS](https://github.com/awslabs/amazon-ecr-credential-helper) and [GCP](https://github.com/GoogleCloudPlatform/docker-credential-gcr) `global._hopsworks.managedDockerRegistery.credHelper.secretName` # { #helm.global._hopsworks.managedDockerRegistery.credHelper.secretName } : Type `string`, default `""`. the name of the secret to be created with the credentials helper configuration `global._hopsworks.managedDockerRegistery.domain` # { #helm.global._hopsworks.managedDockerRegistery.domain } : Type `string`, default `""`. the managed docker registry domain name `global._hopsworks.managedDockerRegistery.namespace` # { #helm.global._hopsworks.managedDockerRegistery.namespace } : Type `string`, default `""`. the namespace to be used `global._hopsworks.managedDockerRegistery.port` # { #helm.global._hopsworks.managedDockerRegistery.port } : Type `string`, default `nil`. port number for the managed docker registry
## minio { #helm-values-global-minio } ??? example "Defaults as YAML" ```yaml global: _hopsworks: minio: enabled: true hopsfs: bucket: hopsfs enabled: true password: minioadmin region: eu-west-1 user: minioadmin ```
`global._hopsworks.minio.enabled` # { #helm.global._hopsworks.minio.enabled } : Type `bool`, default `true`. `global._hopsworks.minio.hopsfs.bucket` # { #helm.global._hopsworks.minio.hopsfs.bucket } : Type `string`, default `"hopsfs"`. `global._hopsworks.minio.hopsfs.enabled` # { #helm.global._hopsworks.minio.hopsfs.enabled } : Type `bool`, default `true`. `global._hopsworks.minio.password` # { #helm.global._hopsworks.minio.password } : Type `string`, default `"minioadmin"`. `global._hopsworks.minio.region` # { #helm.global._hopsworks.minio.region } : Type `string`, default `"eu-west-1"`. `global._hopsworks.minio.user` # { #helm.global._hopsworks.minio.user } : Type `string`, default `"minioadmin"`.
## restoreFromBackup { #helm-values-global-restorefrombackup } ??? example "Defaults as YAML" ```yaml global: _hopsworks: restoreFromBackup: backupId: null forceDataClear: false inPlace: false superset: activeDeadlineSeconds: 14400 backupId: null enabled: false initiatedBy: null ```
`global._hopsworks.restoreFromBackup` # { #helm.global._hopsworks.restoreFromBackup } : Type `object`. restore cluster from a backup id ??? note "Default" ```yaml backupId: null forceDataClear: false inPlace: false superset: activeDeadlineSeconds: 14400 backupId: null enabled: false initiatedBy: null ``` `global._hopsworks.restoreFromBackup.backupId` # { #helm.global._hopsworks.restoreFromBackup.backupId } : Type `string`, default `nil`. the backup id to restore `global._hopsworks.restoreFromBackup.forceDataClear` # { #helm.global._hopsworks.restoreFromBackup.forceDataClear } : Type `bool`, default `false`. flag to indicate if the data should be forcibly cleared before restore `global._hopsworks.restoreFromBackup.inPlace` # { #helm.global._hopsworks.restoreFromBackup.inPlace } : Type `bool`, default `false`. flag to indicate if the restore should be done in-place or to a new cluster `global._hopsworks.restoreFromBackup.superset` # { #helm.global._hopsworks.restoreFromBackup.superset } : Type `object`. Superset database restore (HWORKS-2973). Set enabled and backupId, add values.superset-restore.yaml to hold Superset at zero replicas, and upgrade; when the superset-restore- Job has completed, remove the file, set enabled=false and upgrade again. ??? note "Default" ```yaml activeDeadlineSeconds: 14400 backupId: null enabled: false initiatedBy: null ``` `global._hopsworks.restoreFromBackup.superset.activeDeadlineSeconds` # { #helm.global._hopsworks.restoreFromBackup.superset.activeDeadlineSeconds } : Type `int`, default `14400`. activeDeadlineSeconds for the restore Job, bounding the wait for the Superset pods and the secret key, the download and the reload `global._hopsworks.restoreFromBackup.superset.backupId` # { #helm.global._hopsworks.restoreFromBackup.superset.backupId } : Type `string`, default `nil`. the Superset backup id to restore (a key in the superset-backups-metadata ConfigMap). Required when enabled `global._hopsworks.restoreFromBackup.superset.enabled` # { #helm.global._hopsworks.restoreFromBackup.superset.enabled } : Type `bool`, default `false`. reload the Superset database from a backup `global._hopsworks.restoreFromBackup.superset.initiatedBy` # { #helm.global._hopsworks.restoreFromBackup.superset.initiatedBy } : Type `string`, default `nil`. who initiated this restore, recorded in the superset-restore-audit ConfigMap. Defaults to helm/@rev. Deployment metadata, not an authenticated identity: the Kubernetes audit log has the actor
## trino { #helm-values-global-trino } ??? example "Defaults as YAML" ```yaml global: _hopsworks: trino: auth: refreshPeriod: 5s egressProbe: echoUrl: '' enabled: true files: csi: defaultPermissions: false fdSocketVolume: trino-files-fuse-fd sidecarGid: 1000 sidecarUid: 1000 mountPath: /opt/hopsworks/trino mountWaitSeconds: 300 storeRoot: /apps/trino image: tag: 483-v1 mountRetryTimeLimit: 2m mountableSecrets: csi: defaultPermissions: false sidecarGid: 1000 sidecarUid: 1000 enabled: true image: repository: hopsworks/hopsfs-mount tag: 3.4.3.3-EE-RC1-1 mechanism: csi mountPath: /opt/hopsworks/mounts storeRoot: /apps/mountable-secrets testCoordinator: enabled: true userCatalogShards: 2 ```
`global._hopsworks.trino.auth` # { #helm.global._hopsworks.trino.auth } : Type `object`, default `{"refreshPeriod":"5s"}`. How Trino authenticates, now that the password and group files live in HopsFS. `global._hopsworks.trino.auth.refreshPeriod` # { #helm.global._hopsworks.trino.auth.refreshPeriod } : Type `string`, default `"5s"`. How often the coordinator re-reads password.db and group.db: the delay before a new project member can log in, and before a removed one is refused. `global._hopsworks.trino.egressProbe` # { #helm.global._hopsworks.trino.egressProbe } : Type `object`, default `{"echoUrl":""}`. Egress-address probe on the Trino pods. An init container prints the address the pod reaches the internet from, and the backend reads it back off the pod log so the catalog dialog can name the addresses to add to an external database's access control list. Nothing is stored: the log lives as long as the pod, and a terminated pod stops being listed. `global._hopsworks.trino.egressProbe.echoUrl` # { #helm.global._hopsworks.trino.egressProbe.echoUrl } : Type `string`, default `""`. URL of a service that echoes the caller's public IP address in its response body. **Empty by default, so the probe is opt-in**: it is the only outbound call this chart makes on its own, and a query engine reaching a third party on every pod start is not a default an on-premise cluster can be given without asking. While empty the init container prints `disabled` and exits without calling anything, the backend reports no addresses, and the catalog dialog tells the user to ask their administrator instead of naming them. Set it to turn the feature on: `https://ifconfig.me` and `https://api.ipify.org` both answer in the required shape, and an internal equivalent is preferable where one exists. The probe never fails a pod and never delays startup by more than its 5 second timeout. `global._hopsworks.trino.enabled` # { #helm.global._hopsworks.trino.enabled } : Type `bool`, default `true`. Enable or disable the installation of the trino sub chart. `global._hopsworks.trino.files` # { #helm.global._hopsworks.trino.files } : Type `object`. Trino's own platform files: password file, group file, access-control rules and user catalogs, as HopsFS files. No enable flag: Trino authenticates nobody without them. Delivered through the hopsfs-csi driver, so Trino requires global._hopsworks.csi.enabled. ??? note "Default" ```yaml csi: defaultPermissions: false fdSocketVolume: trino-files-fuse-fd sidecarGid: 1000 sidecarUid: 1000 mountPath: /opt/hopsworks/trino mountWaitSeconds: 300 storeRoot: /apps/trino ``` `global._hopsworks.trino.files.csi` # { #helm.global._hopsworks.trino.files.csi } : Type `object`. Settings for the hopsfs-csi transport, separate from mountableSecrets.csi because the two mounts need separate fd-handoff sockets. ??? note "Default" ```yaml defaultPermissions: false fdSocketVolume: trino-files-fuse-fd sidecarGid: 1000 sidecarUid: 1000 ``` `global._hopsworks.trino.files.csi.defaultPermissions` # { #helm.global._hopsworks.trino.files.csi.defaultPermissions } : Type `bool`, default `false`. FUSE default_permissions. Off: the tree presents as root:root 0750 in the pod and Trino runs as 1000, so a kernel check makes every read EPERM. HopsFS still authorizes server-side as `trino`, and the mount is read-only. `global._hopsworks.trino.files.csi.fdSocketVolume` # { #helm.global._hopsworks.trino.files.csi.fdSocketVolume } : Type `string`, default `"trino-files-fuse-fd"`. Name of the pod emptyDir carrying this mount's fd-handoff socket. Must differ from mountableSecrets' `hopsfs-fuse-fd`, or every Trino pod sits in ContainerCreating. Checked by charts/trino. `global._hopsworks.trino.files.csi.sidecarGid` # { #helm.global._hopsworks.trino.files.csi.sidecarGid } : Type `int`, default `1000`. gid counterpart of sidecarUid. `global._hopsworks.trino.files.csi.sidecarUid` # { #helm.global._hopsworks.trino.files.csi.sidecarUid } : Type `int`, default `1000`. uid the node plugin chowns the fd-handoff socket to. Must equal the sidecar's literal runAsUser in charts/trino/values.yaml; charts/trino checks it. `global._hopsworks.trino.files.mountPath` # { #helm.global._hopsworks.trino.files.mountPath } : Type `string`, default `"/opt/hopsworks/trino"`. Where the tree is mounted in the Trino pods, and the base of the paths in the password, group and access-control properties. `global._hopsworks.trino.files.mountWaitSeconds` # { #helm.global._hopsworks.trino.files.mountWaitSeconds } : Type `int`, default `300`. Seconds wait-trino-files waits for the mount to be served and seeded before failing the init container (kubelet then retries). `global._hopsworks.trino.files.storeRoot` # { #helm.global._hopsworks.trino.files.storeRoot } : Type `string`, default `"/apps/trino"`. Where the files live in HopsFS. Preset by charts/hopsfs, passed to the sidecar as `--srcDir` and seeded to the backend as `trino_files_path`. Moving it does not move existing files. `global._hopsworks.trino.image` # { #helm.global._hopsworks.trino.image } : Type `object`, default `{"tag":"483-v1"}`. The Trino image this release deploys, restated here for the backend. `global._hopsworks.trino.image.tag` # { #helm.global._hopsworks.trino.image.tag } : Type `string`, default `"483-v1"`. The Trino image tag, seeded to the backend as `trino_image_tag` so it can tell when its connector-property table is stale. Duplicates trino.image.tag in charts/trino/values.yaml, which cannot be templated; charts/trino fails the render when they disagree. Bump both. `global._hopsworks.trino.mountRetryTimeLimit` # { #helm.global._hopsworks.trino.mountRetryTimeLimit } : Type `string`, default `"2m"`. `-retryTimeLimit` for both HopsFS sidecars in the Trino pods: how long one filesystem operation blocks while the client retries an unreachable namenode. Effectively a startup setting, since Trino reads an I/O error on its config files as "does not exist" and exits. 2m rides out a namenode container restart (measured 2m14s) and caps the dead time after a namenode pod replacement, which hopsfs-mount does not follow. `global._hopsworks.trino.mountableSecrets` # { #helm.global._hopsworks.trino.mountableSecrets } : Type `object`. Per-project credential files (Oracle wallets, keystores) delivered to the Trino pods, so a connector property can name a real directory. Named after the capability rather than the transport. They arrive through the hopsfs-csi driver, like the platform files. ??? note "Default" ```yaml csi: defaultPermissions: false sidecarGid: 1000 sidecarUid: 1000 enabled: true image: repository: hopsworks/hopsfs-mount tag: 3.4.3.3-EE-RC1-1 mechanism: csi mountPath: /opt/hopsworks/mounts storeRoot: /apps/mountable-secrets ``` `global._hopsworks.trino.mountableSecrets.csi` # { #helm.global._hopsworks.trino.mountableSecrets.csi } : Type `object`, default `{"defaultPermissions":false,"sidecarGid":1000,"sidecarUid":1000}`. Settings for the hopsfs-csi transport: the identity the node plugin hands the mount to. `global._hopsworks.trino.mountableSecrets.csi.defaultPermissions` # { #helm.global._hopsworks.trino.mountableSecrets.csi.defaultPermissions } : Type `bool`, default `false`. Whether the kernel checks the file modes hopsfs-mount reports (FUSE default_permissions). Off for this mount: the readers are the Trino containers, whose uid is not a user in the sidecar image, and the tree's owners are backend service users that never resolve there either, so a kernel-side check can only refuse. The mount is read-only at the kernel and bounded to storeRoot, and HopsFS still enforces its own permissions server-side as the authenticated `trino` user, so nothing is lost by turning it off. `global._hopsworks.trino.mountableSecrets.csi.sidecarGid` # { #helm.global._hopsworks.trino.mountableSecrets.csi.sidecarGid } : Type `int`, default `1000`. gid the node plugin records on the FUSE mount and chowns the socket to; the sidecar's runAsGroup, with the same constraints as sidecarUid. `global._hopsworks.trino.mountableSecrets.csi.sidecarUid` # { #helm.global._hopsworks.trino.mountableSecrets.csi.sidecarUid } : Type `int`, default `1000`. uid the node plugin records on the FUSE mount and chowns the fd-handoff socket to. The sidecar in charts/trino/values.yaml MUST run as exactly this uid: runAsUser is an integer the upstream chart does not template, so it is a literal there, restated on OpenShift with an id from the namespace range (values.openshift.yaml), and charts/trino fails the render when the two differ. 1000 is the Trino user, so the pod has one identity. `global._hopsworks.trino.mountableSecrets.enabled` # { #helm.global._hopsworks.trino.mountableSecrets.enabled } : Type `bool`, default `true`. Whether the Trino pods are given the backend-owned /apps/mountable-secrets tree at all. This is trino's own declaration of intent. When false the backend variable mountable_secrets_enabled is unset, the feature is reported unavailable, and the sidecar entries must be removed from charts/trino/values.yaml and the `mountable-secrets` volume restated as an emptyDir (charts/trino fails the render otherwise: neither can be omitted by a conditional, since values.yaml is not templated by Helm, and a CSI volume left behind with no sidecar to serve it is a mount whose every access blocks forever). `global._hopsworks.trino.mountableSecrets.image.repository` # { #helm.global._hopsworks.trino.mountableSecrets.image.repository } : Type `string`, default `"hopsworks/hopsfs-mount"`. Repository of the small bash image the Trino init containers wait-trino-files and assemble-catalogs run in. The dedicated hopsfs-mount image built in docker-images (90.8 MB); the mount sidecars themselves run global._hopsworks.csi.image. `global._hopsworks.trino.mountableSecrets.image.tag` # { #helm.global._hopsworks.trino.mountableSecrets.image.tag } : Type `string`, default `"3.4.3.3-EE-RC1-1"`. Tag for the sidecar image, `-`. The first half is the hops-fuse-mount artifact version, not the platform version: the image carries the HopsFS FUSE client and nothing else, so it turns over with HopsFS. Keep that half in step with charts/hopsfs image.tag, since the FUSE client should match the HopsFS line it talks to. The second half moves when the image is rebuilt without the artifact changing, a base bump or a security rebuild, so such a rebuild cannot silently replace the bytes behind a tag already deployed. Both halves are pinned here on purpose. `global._hopsworks.trino.mountableSecrets.mountPath` # { #helm.global._hopsworks.trino.mountableSecrets.mountPath } : Type `string`, default `"/opt/hopsworks/mounts"`. Where the tree is mounted inside the Trino pods. Seeded to the backend as the `trino_mountable_secrets_root` variable and used for the sidecar's mount point, its preStop unmount and the Trino containers' mount, so one value drives both sides. They agreed only by both defaulting to the same literal before, which meant an operator moving the mount left the backend resolving `${HOPSWORKS_MOUNT:...}` under a path nothing was mounted at. `global._hopsworks.trino.mountableSecrets.storeRoot` # { #helm.global._hopsworks.trino.mountableSecrets.storeRoot } : Type `string`, default `"/apps/mountable-secrets"`. Where the store lives in HopsFS, the source side of the mount. One value, three consumers: charts/hopsfs presets the directory, the mount sidecar passes it as `--srcDir`, and it is seeded to the backend as the `mountable_secrets_path` variable so the backend writes bundles where the mount reads them. It was a literal in all three places before, agreeing only by coincidence, which is the trap `mountPath` had on the container side. Not a knob to reach for: moving it does not move the bundles already written under the old path, and the tree is backend-owned (payara:hdfs, 0750) rather than operator-managed. `global._hopsworks.trino.testCoordinator` # { #helm.global._hopsworks.trino.testCoordinator } : Type `object`, default `{"enabled":true}`. Optional dedicated Trino test coordinator used to connection-test user catalogs before they are synced to the production coordinator. `global._hopsworks.trino.testCoordinator.enabled` # { #helm.global._hopsworks.trino.testCoordinator.enabled } : Type `bool`, default `true`. Deploy an optional dedicated Trino coordinator (catalog.management=dynamic, writable catalog dir) used to connection-test user catalogs before syncing them to the production coordinator. When true, the backend variable trino_test_coordinator_enabled is set so the "test connection" feature becomes available. Enabled by default; set to false to skip the extra coordinator (the "test connection" action is then reported unavailable). `global._hopsworks.trino.mountableSecrets.mechanism` Deprecated # { #helm.global._hopsworks.trino.mountableSecrets.mechanism } : Type `string`, default `"csi"`. DEPRECATED and read by nothing. Selected between the hopsfs-csi transport and the privileged `hopsfsMount` sidecar; the sidecar is gone and Trino requires global._hopsworks.csi.enabled. Kept only so an override carried from 5.1 still validates. `global._hopsworks.trino.userCatalogShards` Deprecated # { #helm.global._hopsworks.trino.userCatalogShards } : Type `int`, default `2`. DEPRECATED and read by nothing: user catalogs are HopsFS files now, not sharded Secrets. Kept so an existing override still validates.
================================================================================ # airflow Source: https://docs.hopsworks.ai/latest/setup_installation/common/helm_chart_values/airflow/ # Airflow values { #helm-values-airflow } Values under `airflow` configure Apache Airflow, which schedules and orchestrates Hopsworks jobs. _Generated from the Hopsworks Helm chart `5.2.0-alpha-1791549041` (Hopsworks `5.2.0`)._ Deployed according to the first of these values that is set: [`global._hopsworks.airflow.enabled`](global.md#helm.global._hopsworks.airflow.enabled), [`global._hopsworks.full_platform`](global.md#helm.global._hopsworks.full_platform). ## General { #helm-values-airflow-general } ??? example "Defaults as YAML" ```yaml airflow: appName: airflow bidirectional_mount: mounted cleanupOnUninstall: enabled: true ttlSecondsAfterFinished: null csi: {} debug: false dependencies: hopsworks: consulServiceName: glassfish consulServiceTag: hopsworks port: 8182 mysql: consulServiceName: mysql ddlConsulServiceName: mysqlddl port: 3306 namenode: consulServiceName: namenode port: 8020 hopsfs_bin_url: null hopsworks: caBundlePath: /etc/airflow/ca/hopsworks-ca.crt enableApiKeyFallback: true hwJwtCacheTtlSeconds: 60 internalClientCn: hopsworks-ee.hopsworks.svc manifestPath: /shared-volume/hopsfs/.airflow/manifest.json membershipCacheTtlSeconds: 60 url: '' hopsworkslib: {} image: pullPolicy: IfNotPresent registry: docker.hops.works migrationBackOffLimit: 10 migrationJobTtlSecondsAfterFinished: null orphanCleanup: schedule: 42 2 * * * reset_db_if_error: false serviceAccount: annotations: {} serviceAccountName: airflow tls: true ```
`airflow` # { #helm.airflow } : Type `object`, default `{"csi":{}}`. override airflow values `airflow.appName` # { #helm.airflow.appName } : Type `string`, default `"airflow"`. app name label `airflow.bidirectional_mount` # { #helm.airflow.bidirectional_mount } : Type `string`, default `"mounted"`. the name of bidirectional mount `airflow.cleanupOnUninstall` # { #helm.airflow.cleanupOnUninstall } : Type `object`, default `{"enabled":true,"ttlSecondsAfterFinished":null}`. post-delete cleanup of Airflow leftovers. The keys Secret (airflowApi.keysSecretName) is created by the keys-bootstrap pre-install hook via kubectl, so Helm/ArgoCD never track it; this deletes it by name on uninstall. `airflow.cleanupOnUninstall.enabled` # { #helm.airflow.cleanupOnUninstall.enabled } : Type `bool`, default `true`. enable the post-delete Airflow keys-Secret cleanup hook `airflow.cleanupOnUninstall.ttlSecondsAfterFinished` # { #helm.airflow.cleanupOnUninstall.ttlSecondsAfterFinished } : Type `string`, default `nil`. ttlSecondsAfterFinished for the cleanup Job; null falls through to the global default `airflow.debug` # { #helm.airflow.debug } : Type `bool`, default `false`. enable or disable debug logging for hopsfs `airflow.dependencies.hopsworks` # { #helm.airflow.dependencies.hopsworks } : Type `object`, default `{"consulServiceName":"glassfish","consulServiceTag":"hopsworks","port":8182}`. hopsworks consul service `airflow.dependencies.mysql` # { #helm.airflow.dependencies.mysql } : Type `object`, default `{"consulServiceName":"mysql","ddlConsulServiceName":"mysqlddl","port":3306}`. mysql consul service `airflow.dependencies.namenode` # { #helm.airflow.dependencies.namenode } : Type `object`, default `{"consulServiceName":"namenode","port":8020}`. namenode consul service. This uses a headless ClusterIP underneath with `publishNotReadyAddresses: true` `airflow.hopsfs_bin_url` # { #helm.airflow.hopsfs_bin_url } : Type `string`, default `nil`. url to download a patched hopsfs mount for testing and development `airflow.hopsworks` # { #helm.airflow.hopsworks } : Type `object`. Hopsworks-specific Airflow 3 configuration. ??? note "Default" ```yaml caBundlePath: /etc/airflow/ca/hopsworks-ca.crt enableApiKeyFallback: true hwJwtCacheTtlSeconds: 60 internalClientCn: hopsworks-ee.hopsworks.svc manifestPath: /shared-volume/hopsfs/.airflow/manifest.json membershipCacheTtlSeconds: 60 url: '' ``` `airflow.hopsworkslib` # { #helm.airflow.hopsworkslib } : Type `object`, default `{}`. override hopsworkslib values `airflow.image` # { #helm.airflow.image } : Type `object`, default `{"pullPolicy":"IfNotPresent","registry":"docker.hops.works"}`. image configuration. The image tag is the .Chart.AppVersion `airflow.migrationBackOffLimit` # { #helm.airflow.migrationBackOffLimit } : Type `int`, default `10`. backoffLimit for airflow migration job `airflow.migrationJobTtlSecondsAfterFinished` # { #helm.airflow.migrationJobTtlSecondsAfterFinished } : Type `string`, default `nil`. TTL in seconds for the migrate-airflow Job. Overrides global default. `airflow.orphanCleanup` # { #helm.airflow.orphanCleanup } : Type `object`, default `{"schedule":"42 2 * * *"}`. periodic cleanup of orphaned Airflow rows. The metadata tables live in RonDB without enforced foreign keys (NDB cannot carry FKs on the blob/text tables), so a CronJob restores the dropped ON DELETE CASCADE semantics by deleting children whose parent `airflow db clean` removed. Gated by the airflow subchart's own enable condition in the umbrella `Chart.yaml` (`global._hopsworks.airflow.enabled,global._hopsworks.full_platform`): if airflow is disabled, this CronJob is not rendered. `airflow.reset_db_if_error` # { #helm.airflow.reset_db_if_error } : Type `bool`, default `false`. if true, will reset the database and create a new one in case of error. In Airflow 3 the db-reset job runs unconditionally before migrate; this key is retained for backwards compatibility but the v3 chart ignores it. `airflow.serviceAccount.annotations` # { #helm.airflow.serviceAccount.annotations } : Type `object`, default `{}`. service account annotations `airflow.serviceAccountName` # { #helm.airflow.serviceAccountName } : Type `string`, default `"airflow"`. service account name `airflow.tls` # { #helm.airflow.tls } : Type `bool`, default `true`. enable or disable TLS
## airflowApi { #helm-values-airflow-airflowapi } ??? example "Defaults as YAML" ```yaml airflow: airflowApi: basePath: /hopsworks-api/airflow bundleRoot: /opt/airflow/hopsworks-bundle/dags corsAllowedOrigins: '' jwtAlgorithm: RS256 jwtAudience: hopsworks-airflow jwtExpirationSeconds: 3600 jwtPrivateKeyPath: /etc/airflow/keys/api-server-private.pem jwtPublicKeyPath: /etc/airflow/keys/api-server-public.pem keysSecretName: hopsworks-airflow-keys schedulerPrivateKeyPath: /etc/airflow/keys/scheduler-private.pem schedulerPublicKeyPath: /etc/airflow/keys/scheduler-public.pem ```
`airflow.airflowApi.basePath` # { #helm.airflow.airflowApi.basePath } : Type `string`, default `"/hopsworks-api/airflow"`. `airflow.airflowApi.bundleRoot` # { #helm.airflow.airflowApi.bundleRoot } : Type `string`, default `"/opt/airflow/hopsworks-bundle/dags"`. `airflow.airflowApi.corsAllowedOrigins` # { #helm.airflow.airflowApi.corsAllowedOrigins } : Type `string`, default `""`. `airflow.airflowApi.jwtAlgorithm` # { #helm.airflow.airflowApi.jwtAlgorithm } : Type `string`, default `"RS256"`. `airflow.airflowApi.jwtAudience` # { #helm.airflow.airflowApi.jwtAudience } : Type `string`, default `"hopsworks-airflow"`. `airflow.airflowApi.jwtExpirationSeconds` # { #helm.airflow.airflowApi.jwtExpirationSeconds } : Type `int`, default `3600`. `airflow.airflowApi.jwtPrivateKeyPath` # { #helm.airflow.airflowApi.jwtPrivateKeyPath } : Type `string`, default `"/etc/airflow/keys/api-server-private.pem"`. `airflow.airflowApi.jwtPublicKeyPath` # { #helm.airflow.airflowApi.jwtPublicKeyPath } : Type `string`, default `"/etc/airflow/keys/api-server-public.pem"`. `airflow.airflowApi.keysSecretName` # { #helm.airflow.airflowApi.keysSecretName } : Type `string`, default `"hopsworks-airflow-keys"`. `airflow.airflowApi.schedulerPrivateKeyPath` # { #helm.airflow.airflowApi.schedulerPrivateKeyPath } : Type `string`, default `"/etc/airflow/keys/scheduler-private.pem"`. `airflow.airflowApi.schedulerPublicKeyPath` # { #helm.airflow.airflowApi.schedulerPublicKeyPath } : Type `string`, default `"/etc/airflow/keys/scheduler-public.pem"`.
## common { #helm-values-airflow-common } ??? example "Defaults as YAML" ```yaml airflow: common: celery: broker_url: rdis://:6379/0 celery_app_name: airflow.executors.celery_executor default_queue: default flower_port: 5555 flower_url_prefix: http://localhost/hopsworks-api/flower worker_concurrency: 8 worker_log_server_port: 8793 core_config: airflow_home: /airflow base_log_folder: /airflow/logs dag_concurrency: 16 dagbag_import_timeout: 60 dags_are_paused_at_creation: true dags_folder: /airflow/dags default_timezone: utc donot_pickle: false executor: LocalExecutor fernet_key: G3jB5--jCQpRYp7hwUtpfQ_S8zLRbRMwX8tr3dehnNU= hostname_callable: airflow.utils.net.get_host_ip_address load_examples: false logging_config_class: log_config.LOGGING_CONFIG max_active_runs_per_dag: 16 non_pooled_task_slot_count: 128 parallelism: 32 plugins_folder: /airflow/plugins sql_alchemy_max_overflow: 30 sql_alchemy_pool_pre_ping: true sql_alchemy_pool_recycle: 3600 sql_alchemy_pool_size: 10 group: airflow namenode_cluster_ip_name: namenode-cluster-ip smtp: smtp_host: localhost smtp_mail_from: admin@kth.se smtp_password: admin smtp_port: 25 smtp_ssl: false smtp_starttls: true smtp_user: admin@kth.se user: airflow ```
`airflow.common.celery.broker_url` # { #helm.airflow.common.celery.broker_url } : Type `string`, default `"rdis://:6379/0"`. `airflow.common.celery.celery_app_name` # { #helm.airflow.common.celery.celery_app_name } : Type `string`, default `"airflow.executors.celery_executor"`. `airflow.common.celery.default_queue` # { #helm.airflow.common.celery.default_queue } : Type `string`, default `"default"`. `airflow.common.celery.flower_port` # { #helm.airflow.common.celery.flower_port } : Type `int`, default `5555`. `airflow.common.celery.flower_url_prefix` # { #helm.airflow.common.celery.flower_url_prefix } : Type `string`, default `"http://localhost/hopsworks-api/flower"`. `airflow.common.celery.worker_concurrency` # { #helm.airflow.common.celery.worker_concurrency } : Type `int`, default `8`. `airflow.common.celery.worker_log_server_port` # { #helm.airflow.common.celery.worker_log_server_port } : Type `int`, default `8793`. `airflow.common.core_config.airflow_home` # { #helm.airflow.common.core_config.airflow_home } : Type `string`, default `"/airflow"`. `airflow.common.core_config.base_log_folder` # { #helm.airflow.common.core_config.base_log_folder } : Type `string`, default `"/airflow/logs"`. `airflow.common.core_config.dag_concurrency` # { #helm.airflow.common.core_config.dag_concurrency } : Type `int`, default `16`. `airflow.common.core_config.dagbag_import_timeout` # { #helm.airflow.common.core_config.dagbag_import_timeout } : Type `int`, default `60`. `airflow.common.core_config.dags_are_paused_at_creation` # { #helm.airflow.common.core_config.dags_are_paused_at_creation } : Type `bool`, default `true`. `airflow.common.core_config.dags_folder` # { #helm.airflow.common.core_config.dags_folder } : Type `string`, default `"/airflow/dags"`. `airflow.common.core_config.default_timezone` # { #helm.airflow.common.core_config.default_timezone } : Type `string`, default `"utc"`. `airflow.common.core_config.donot_pickle` # { #helm.airflow.common.core_config.donot_pickle } : Type `bool`, default `false`. `airflow.common.core_config.executor` # { #helm.airflow.common.core_config.executor } : Type `string`, default `"LocalExecutor"`. `airflow.common.core_config.fernet_key` # { #helm.airflow.common.core_config.fernet_key } : Type `string`, default `"G3jB5--jCQpRYp7hwUtpfQ_S8zLRbRMwX8tr3dehnNU="`. `airflow.common.core_config.hostname_callable` # { #helm.airflow.common.core_config.hostname_callable } : Type `string`, default `"airflow.utils.net.get_host_ip_address"`. `airflow.common.core_config.load_examples` # { #helm.airflow.common.core_config.load_examples } : Type `bool`, default `false`. `airflow.common.core_config.logging_config_class` # { #helm.airflow.common.core_config.logging_config_class } : Type `string`, default `"log_config.LOGGING_CONFIG"`. `airflow.common.core_config.max_active_runs_per_dag` # { #helm.airflow.common.core_config.max_active_runs_per_dag } : Type `int`, default `16`. `airflow.common.core_config.non_pooled_task_slot_count` # { #helm.airflow.common.core_config.non_pooled_task_slot_count } : Type `int`, default `128`. `airflow.common.core_config.parallelism` # { #helm.airflow.common.core_config.parallelism } : Type `int`, default `32`. `airflow.common.core_config.plugins_folder` # { #helm.airflow.common.core_config.plugins_folder } : Type `string`, default `"/airflow/plugins"`. `airflow.common.core_config.sql_alchemy_max_overflow` # { #helm.airflow.common.core_config.sql_alchemy_max_overflow } : Type `int`, default `30`. We set it high, not unlimited to account for memory leaks `airflow.common.core_config.sql_alchemy_pool_pre_ping` # { #helm.airflow.common.core_config.sql_alchemy_pool_pre_ping } : Type `bool`, default `true`. `airflow.common.core_config.sql_alchemy_pool_recycle` # { #helm.airflow.common.core_config.sql_alchemy_pool_recycle } : Type `int`, default `3600`. `airflow.common.core_config.sql_alchemy_pool_size` # { #helm.airflow.common.core_config.sql_alchemy_pool_size } : Type `int`, default `10`. `airflow.common.group` # { #helm.airflow.common.group } : Type `string`, default `"airflow"`. `airflow.common.namenode_cluster_ip_name` # { #helm.airflow.common.namenode_cluster_ip_name } : Type `string`, default `"namenode-cluster-ip"`. `airflow.common.smtp.smtp_host` # { #helm.airflow.common.smtp.smtp_host } : Type `string`, default `"localhost"`. `airflow.common.smtp.smtp_mail_from` # { #helm.airflow.common.smtp.smtp_mail_from } : Type `string`, default `"admin@kth.se"`. `airflow.common.smtp.smtp_password` # { #helm.airflow.common.smtp.smtp_password } : Type `string`, default `"admin"`. `airflow.common.smtp.smtp_port` # { #helm.airflow.common.smtp.smtp_port } : Type `int`, default `25`. `airflow.common.smtp.smtp_ssl` # { #helm.airflow.common.smtp.smtp_ssl } : Type `bool`, default `false`. `airflow.common.smtp.smtp_starttls` # { #helm.airflow.common.smtp.smtp_starttls } : Type `bool`, default `true`. `airflow.common.smtp.smtp_user` # { #helm.airflow.common.smtp.smtp_user } : Type `string`, default `"admin@kth.se"`. `airflow.common.user` # { #helm.airflow.common.user } : Type `string`, default `"airflow"`.
## csi { #helm-values-airflow-csi } ??? example "Defaults as YAML" ```yaml airflow: csi: clientCertificateSecretKey: '' clientKeySecretKey: '' defaultPermissions: true fallbackGroup: airflow fallbackUser: airflow imageTag: '' rootCABundleSecretKey: '' secretName: '' sidecarGid: 1508 sidecarResources: limits: cpu: '1' memory: 1024Mi requests: cpu: 200m memory: 256Mi sidecarUid: 1512 ```
`airflow.csi` # { #helm.airflow.csi } : Type `object`. CSI settings for the unprivileged HopsFS FUSE sidecar. When `tls=true`, the default secret selectors target the Airflow CSI-specific HopsworksCert secret. ??? note "Default" ```yaml clientCertificateSecretKey: '' clientKeySecretKey: '' defaultPermissions: true fallbackGroup: airflow fallbackUser: airflow imageTag: '' rootCABundleSecretKey: '' secretName: '' sidecarGid: 1508 sidecarResources: limits: cpu: '1' memory: 1024Mi requests: cpu: 200m memory: 256Mi sidecarUid: 1512 ``` `airflow.csi.clientCertificateSecretKey` # { #helm.airflow.csi.clientCertificateSecretKey } : Type `string`, default `""`. override the client certificate bundle file name inside the mounted TLS secret `airflow.csi.clientKeySecretKey` # { #helm.airflow.csi.clientKeySecretKey } : Type `string`, default `""`. override the client key file name inside the mounted TLS secret `airflow.csi.defaultPermissions` # { #helm.airflow.csi.defaultPermissions } : Type `bool`, default `true`. enable the kernel-side default_permissions check on the fuse mount; set false on OpenShift, where workload uids are arbitrary and can never match the owners hopsfs-mount reports (HDFS permissions are still enforced server-side) `airflow.csi.imageTag` # { #helm.airflow.csi.imageTag } : Type `string`, default `""`. override the tag of the hopsfs-csi image run as the unprivileged FUSE sidecar. Empty (the default) takes global._hopsworks.csi.image.tag, the one place the hopsfs-csi image is named, so the sidecar and the node plugin it fetches the mount from cannot drift; set only to test a sidecar build against a deployed plugin `airflow.csi.rootCABundleSecretKey` # { #helm.airflow.csi.rootCABundleSecretKey } : Type `string`, default `""`. override the root CA bundle file name inside the mounted TLS secret `airflow.csi.secretName` # { #helm.airflow.csi.secretName } : Type `string`, default `""`. override the Secret holding the PEM TLS material mounted into the FUSE sidecar (defaults to the csi HopsworksCert secret) `airflow.csi.sidecarGid` # { #helm.airflow.csi.sidecarGid } : Type `int`, default `1508`. gid recorded on the fuse mount; matches the airflow group precreated in the hopsfs-csi image `airflow.csi.sidecarResources` # { #helm.airflow.csi.sidecarResources } : Type `object`. resources for the unprivileged FUSE sidecar. All four slots are set on purpose: a namespace LimitRange fills any missing one with its own default, which for a limit-less request is rejected at admission when the default falls below the request (see HWORKS-3086). ??? note "Default" ```yaml limits: cpu: '1' memory: 1024Mi requests: cpu: 200m memory: 256Mi ``` `airflow.csi.sidecarUid` # { #helm.airflow.csi.sidecarUid } : Type `int`, default `1512`. uid recorded on the fuse mount and used to run the sidecar; matches the airflow user precreated in the hopsfs-csi image
## dagProcessor { #helm-values-airflow-dagprocessor } ??? example "Defaults as YAML" ```yaml airflow: dagProcessor: deployment: replicas: 1 name: airflow-dag-processor nodeSelector: {} probe: failureThreshold: 20 initialDelaySeconds: 30 periodSeconds: 10 timeoutSeconds: 10 refreshIntervalSeconds: 5 resources: limits: memory: 2000Mi requests: cpu: 200m memory: 500Mi securityContext: runAsGroup: 1508 runAsNonRoot: true runAsUser: 1512 tolerations: [] ```
`airflow.dagProcessor.deployment.replicas` # { #helm.airflow.dagProcessor.deployment.replicas } : Type `int`, default `1`. `airflow.dagProcessor.name` # { #helm.airflow.dagProcessor.name } : Type `string`, default `"airflow-dag-processor"`. `airflow.dagProcessor.nodeSelector` # { #helm.airflow.dagProcessor.nodeSelector } : Type `object`, default `{}`. `airflow.dagProcessor.probe.failureThreshold` # { #helm.airflow.dagProcessor.probe.failureThreshold } : Type `int`, default `20`. `airflow.dagProcessor.probe.initialDelaySeconds` # { #helm.airflow.dagProcessor.probe.initialDelaySeconds } : Type `int`, default `30`. `airflow.dagProcessor.probe.periodSeconds` # { #helm.airflow.dagProcessor.probe.periodSeconds } : Type `int`, default `10`. `airflow.dagProcessor.probe.timeoutSeconds` # { #helm.airflow.dagProcessor.probe.timeoutSeconds } : Type `int`, default `10`. `airflow.dagProcessor.refreshIntervalSeconds` # { #helm.airflow.dagProcessor.refreshIntervalSeconds } : Type `int`, default `5`. `airflow.dagProcessor.resources.limits.memory` # { #helm.airflow.dagProcessor.resources.limits.memory } : Type `string`, default `"2000Mi"`. `airflow.dagProcessor.resources.requests.cpu` # { #helm.airflow.dagProcessor.resources.requests.cpu } : Type `string`, default `"200m"`. `airflow.dagProcessor.resources.requests.memory` # { #helm.airflow.dagProcessor.resources.requests.memory } : Type `string`, default `"500Mi"`. `airflow.dagProcessor.securityContext.runAsGroup` # { #helm.airflow.dagProcessor.securityContext.runAsGroup } : Type `int`, default `1508`. `airflow.dagProcessor.securityContext.runAsNonRoot` # { #helm.airflow.dagProcessor.securityContext.runAsNonRoot } : Type `bool`, default `true`. `airflow.dagProcessor.securityContext.runAsUser` # { #helm.airflow.dagProcessor.securityContext.runAsUser } : Type `int`, default `1512`. `airflow.dagProcessor.tolerations` # { #helm.airflow.dagProcessor.tolerations } : Type `list`, default `[]`.
## scheduler { #helm-values-airflow-scheduler } ??? example "Defaults as YAML" ```yaml airflow: scheduler: config: dag_dir_list_interval: 40 job_heartbeat_sec: 5 max_threads: 2 min_file_process_interval: 10 print_stats_interval: 600 scheduler_zombie_task_threshold: 300 deployment: replicas: 1 is_tls: false name: airflow-scheduler nodeSelector: {} podDisruptionBudget: enabled: true minAvailable: 1 probe: failureThreshold: 20 initialDelaySeconds: 30 periodSeconds: 10 timeoutSeconds: 10 probeCheckWebserver: true probeCommand: set -e && pgrep -f "airflow scheduler" resources: limits: memory: 2000Mi requests: cpu: '1' memory: 1000Mi securityContext: runAsGroup: 1508 runAsNonRoot: true runAsUser: 1512 service: annotations: consul.hashicorp.com/service-name: airflow consul.hashicorp.com/service-tags: scheduler tolerations: [] topologySpreadConstraint: {} ```
`airflow.scheduler.config.dag_dir_list_interval` # { #helm.airflow.scheduler.config.dag_dir_list_interval } : Type `int`, default `40`. `airflow.scheduler.config.job_heartbeat_sec` # { #helm.airflow.scheduler.config.job_heartbeat_sec } : Type `int`, default `5`. `airflow.scheduler.config.max_threads` # { #helm.airflow.scheduler.config.max_threads } : Type `int`, default `2`. `airflow.scheduler.config.min_file_process_interval` # { #helm.airflow.scheduler.config.min_file_process_interval } : Type `int`, default `10`. `airflow.scheduler.config.print_stats_interval` # { #helm.airflow.scheduler.config.print_stats_interval } : Type `int`, default `600`. `airflow.scheduler.config.scheduler_zombie_task_threshold` # { #helm.airflow.scheduler.config.scheduler_zombie_task_threshold } : Type `int`, default `300`. `airflow.scheduler.deployment.replicas` # { #helm.airflow.scheduler.deployment.replicas } : Type `int`, default `1`. `airflow.scheduler.is_tls` # { #helm.airflow.scheduler.is_tls } : Type `bool`, default `false`. `airflow.scheduler.name` # { #helm.airflow.scheduler.name } : Type `string`, default `"airflow-scheduler"`. `airflow.scheduler.nodeSelector` # { #helm.airflow.scheduler.nodeSelector } : Type `object`, default `{}`. node selector configuration `airflow.scheduler.podDisruptionBudget.enabled` # { #helm.airflow.scheduler.podDisruptionBudget.enabled } : Type `bool`, default `true`. `airflow.scheduler.podDisruptionBudget.minAvailable` # { #helm.airflow.scheduler.podDisruptionBudget.minAvailable } : Type `int`, default `1`. `airflow.scheduler.probe.failureThreshold` # { #helm.airflow.scheduler.probe.failureThreshold } : Type `int`, default `20`. `airflow.scheduler.probe.initialDelaySeconds` # { #helm.airflow.scheduler.probe.initialDelaySeconds } : Type `int`, default `30`. `airflow.scheduler.probe.periodSeconds` # { #helm.airflow.scheduler.probe.periodSeconds } : Type `int`, default `10`. `airflow.scheduler.probe.timeoutSeconds` # { #helm.airflow.scheduler.probe.timeoutSeconds } : Type `int`, default `10`. `airflow.scheduler.probeCheckWebserver` # { #helm.airflow.scheduler.probeCheckWebserver } : Type `bool`, default `true`. `airflow.scheduler.probeCommand` # { #helm.airflow.scheduler.probeCommand } : Type `string`, default `"set -e && pgrep -f \"airflow scheduler\""`. The probe here will be concatenated with sleeping for the heartbeat time and then trying to reach the webserver. If the webserver is not reachable the scheduler should be restarted The scheduler goes to a non consistent state, then the health check of the webserver returns scheduler non-healthy and it fails `airflow.scheduler.resources.limits` # { #helm.airflow.scheduler.resources.limits } : Type `object`, default `{"memory":"2000Mi"}`. resources limits configuration `airflow.scheduler.resources.requests` # { #helm.airflow.scheduler.resources.requests } : Type `object`, default `{"cpu":"1","memory":"1000Mi"}`. resources requests configuration `airflow.scheduler.securityContext.runAsGroup` # { #helm.airflow.scheduler.securityContext.runAsGroup } : Type `int`, default `1508`. `airflow.scheduler.securityContext.runAsNonRoot` # { #helm.airflow.scheduler.securityContext.runAsNonRoot } : Type `bool`, default `true`. `airflow.scheduler.securityContext.runAsUser` # { #helm.airflow.scheduler.securityContext.runAsUser } : Type `int`, default `1512`. `airflow.scheduler.service.annotations."consul.hashicorp.com/service-name"` # { #helm.airflow.scheduler.service.annotations.consul.hashicorp.com-service-name } : Type `string`, default `"airflow"`. `airflow.scheduler.service.annotations."consul.hashicorp.com/service-tags"` # { #helm.airflow.scheduler.service.annotations.consul.hashicorp.com-service-tags } : Type `string`, default `"scheduler"`. `airflow.scheduler.tolerations` # { #helm.airflow.scheduler.tolerations } : Type `list`, default `[]`. `airflow.scheduler.topologySpreadConstraint` # { #helm.airflow.scheduler.topologySpreadConstraint } : Type `object`, default `{}`. The default topology spread constraint. If not defined the global topology spread constraint would be used instead.
## webserver { #helm-values-airflow-webserver } ??? example "Defaults as YAML" ```yaml airflow: webserver: config: authenticate: true expose_config: true rbac: true secret_key: temporary_key web_server_host: 0.0.0.0 web_server_port: 12358 web_server_worker_timeout: 120 worker_class: sync workers: 2 configHelm: base_path: /hopsworks-api/airflow deployment: replicas: 1 forwardedAllowIps: 10.0.0.0/8,172.16.0.0/12,192.168.0.0/16 heartbeat: 30 is_tls: false name: airflow-webserver nodeSelector: {} podDisruptionBudget: enabled: true minAvailable: 1 probe: failureThreshold: 10 initialDelaySeconds: 60 periodSeconds: 5 timeoutSeconds: 10 resources: limits: memory: 2000Mi requests: cpu: 100m memory: 1000Mi securityContext: runAsGroup: 1508 runAsUser: 1512 service: annotations: consul.hashicorp.com/service-name: airflow consul.hashicorp.com/service-port: server consul.hashicorp.com/service-tags: ui tolerations: [] topologySpreadConstraint: {} ```
`airflow.webserver.config.authenticate` # { #helm.airflow.webserver.config.authenticate } : Type `bool`, default `true`. `airflow.webserver.config.expose_config` # { #helm.airflow.webserver.config.expose_config } : Type `bool`, default `true`. `airflow.webserver.config.rbac` # { #helm.airflow.webserver.config.rbac } : Type `bool`, default `true`. `airflow.webserver.config.secret_key` # { #helm.airflow.webserver.config.secret_key } : Type `string`, default `"temporary_key"`. `airflow.webserver.config.web_server_host` # { #helm.airflow.webserver.config.web_server_host } : Type `string`, default `"0.0.0.0"`. `airflow.webserver.config.web_server_port` # { #helm.airflow.webserver.config.web_server_port } : Type `int`, default `12358`. `airflow.webserver.config.web_server_worker_timeout` # { #helm.airflow.webserver.config.web_server_worker_timeout } : Type `int`, default `120`. `airflow.webserver.config.worker_class` # { #helm.airflow.webserver.config.worker_class } : Type `string`, default `"sync"`. `airflow.webserver.config.workers` # { #helm.airflow.webserver.config.workers } : Type `int`, default `2`. `airflow.webserver.configHelm.base_path` # { #helm.airflow.webserver.configHelm.base_path } : Type `string`, default `"/hopsworks-api/airflow"`. `airflow.webserver.deployment.replicas` # { #helm.airflow.webserver.deployment.replicas } : Type `int`, default `1`. `airflow.webserver.forwardedAllowIps` # { #helm.airflow.webserver.forwardedAllowIps } : Type `string`, default `"10.0.0.0/8,172.16.0.0/12,192.168.0.0/16"`. `airflow.webserver.heartbeat` # { #helm.airflow.webserver.heartbeat } : Type `int`, default `30`. `airflow.webserver.is_tls` # { #helm.airflow.webserver.is_tls } : Type `bool`, default `false`. `airflow.webserver.name` # { #helm.airflow.webserver.name } : Type `string`, default `"airflow-webserver"`. `airflow.webserver.nodeSelector` # { #helm.airflow.webserver.nodeSelector } : Type `object`, default `{}`. node selector configuration `airflow.webserver.podDisruptionBudget.enabled` # { #helm.airflow.webserver.podDisruptionBudget.enabled } : Type `bool`, default `true`. `airflow.webserver.podDisruptionBudget.minAvailable` # { #helm.airflow.webserver.podDisruptionBudget.minAvailable } : Type `int`, default `1`. `airflow.webserver.probe.failureThreshold` # { #helm.airflow.webserver.probe.failureThreshold } : Type `int`, default `10`. `airflow.webserver.probe.initialDelaySeconds` # { #helm.airflow.webserver.probe.initialDelaySeconds } : Type `int`, default `60`. `airflow.webserver.probe.periodSeconds` # { #helm.airflow.webserver.probe.periodSeconds } : Type `int`, default `5`. `airflow.webserver.probe.timeoutSeconds` # { #helm.airflow.webserver.probe.timeoutSeconds } : Type `int`, default `10`. `airflow.webserver.resources.limits` # { #helm.airflow.webserver.resources.limits } : Type `object`, default `{"memory":"2000Mi"}`. resources limits configuration `airflow.webserver.resources.requests` # { #helm.airflow.webserver.resources.requests } : Type `object`, default `{"cpu":"100m","memory":"1000Mi"}`. resources requests configuration `airflow.webserver.securityContext.runAsGroup` # { #helm.airflow.webserver.securityContext.runAsGroup } : Type `int`, default `1508`. `airflow.webserver.securityContext.runAsUser` # { #helm.airflow.webserver.securityContext.runAsUser } : Type `int`, default `1512`. `airflow.webserver.service.annotations` # { #helm.airflow.webserver.service.annotations } : Type `object`. annotations on the airflow-webserver Service; the map is open, so an operator can add their own ??? note "Default" ```yaml consul.hashicorp.com/service-name: airflow consul.hashicorp.com/service-port: server consul.hashicorp.com/service-tags: ui ``` `airflow.webserver.tolerations` # { #helm.airflow.webserver.tolerations } : Type `list`, default `[]`. `airflow.webserver.topologySpreadConstraint` # { #helm.airflow.webserver.topologySpreadConstraint } : Type `object`, default `{}`. The default topology spread constraint. If not defined the global topology spread constraint would be used instead.
================================================================================ # arrowflight Source: https://docs.hopsworks.ai/latest/setup_installation/common/helm_chart_values/arrowflight/ # Arrow Flight values { #helm-values-arrowflight } Values under `arrowflight` configure the Arrow Flight server, which serves fast reads of feature groups and training datasets. _Generated from the Hopsworks Helm chart `5.2.0-alpha-1791549041` (Hopsworks `5.2.0`)._ Deployed when [`global._hopsworks.full_platform`](global.md#helm.global._hopsworks.full_platform) is `true`. ## General { #helm-values-arrowflight-general } ??? example "Defaults as YAML" ```yaml arrowflight: appName: arrowflight common: monitoring_port: 12810 port: 5005 configmap: name: arrowflight-configmap hopsworkslib: {} hpa: enabled: true flyingduck_queue_time_avg: 27 flyingduck_request_queue_gauge: 2 maxReplicas: 3 image: pullPolicy: IfNotPresent registry: docker.hops.works nodeSelector: {} podDisruptionBudget: enabled: true minAvailable: 1 spillVolume: size: 20Gi storageClassName: null tolerations: [] topologySpreadConstraint: {} ```
`arrowflight` # { #helm.arrowflight } : Type `object`, default `{}`. override arrow flight values `arrowflight.appName` # { #helm.arrowflight.appName } : Type `string`, default `"arrowflight"`. `arrowflight.common.monitoring_port` # { #helm.arrowflight.common.monitoring_port } : Type `int`, default `12810`. `arrowflight.common.port` # { #helm.arrowflight.common.port } : Type `int`, default `5005`. `arrowflight.configmap.name` # { #helm.arrowflight.configmap.name } : Type `string`, default `"arrowflight-configmap"`. `arrowflight.hopsworkslib` # { #helm.arrowflight.hopsworkslib } : Type `object`, default `{}`. override hopsworkslib values `arrowflight.hpa.enabled` # { #helm.arrowflight.hpa.enabled } : Type `bool`, default `true`. Autoscale the Deployment with an HPA (suppressed when a VPA targets it). While active, deployment.replicas is not rendered and the HPA owns spec.replicas. `arrowflight.hpa.flyingduck_queue_time_avg` # { #helm.arrowflight.hpa.flyingduck_queue_time_avg } : Type `int`, default `27`. `arrowflight.hpa.flyingduck_request_queue_gauge` # { #helm.arrowflight.hpa.flyingduck_request_queue_gauge } : Type `int`, default `2`. `arrowflight.hpa.maxReplicas` # { #helm.arrowflight.hpa.maxReplicas } : Type `int`, default `3`. `arrowflight.image` # { #helm.arrowflight.image } : Type `object`, default `{"pullPolicy":"IfNotPresent","registry":"docker.hops.works"}`. image configuration. The image tag is the .Chart.AppVersion `arrowflight.nodeSelector` # { #helm.arrowflight.nodeSelector } : Type `object`, default `{}`. node selector configuration `arrowflight.podDisruptionBudget.enabled` # { #helm.arrowflight.podDisruptionBudget.enabled } : Type `bool`, default `true`. `arrowflight.podDisruptionBudget.minAvailable` # { #helm.arrowflight.podDisruptionBudget.minAvailable } : Type `int`, default `1`. `arrowflight.spillVolume` # { #helm.arrowflight.spillVolume } : Type `object`, default `{"size":"20Gi","storageClassName":null}`. storage configuration for the ephemeral volume used by DuckDB for spilling on disk during query execution for spilling on disk during query execution for spilling on disk during query execution `arrowflight.spillVolume.size` # { #helm.arrowflight.spillVolume.size } : Type `string`, default `"20Gi"`. size of the ephemeral volume `arrowflight.spillVolume.storageClassName` # { #helm.arrowflight.spillVolume.storageClassName } : Type `string`, default `nil`. storage class name. If null, the default storage class will be used `arrowflight.tolerations` # { #helm.arrowflight.tolerations } : Type `list`, default `[]`. `arrowflight.topologySpreadConstraint` # { #helm.arrowflight.topologySpreadConstraint } : Type `object`, default `{}`. The default topology spread constraint. If not defined the global topology spread constraint would be used instead.
## deployment { #helm-values-arrowflight-deployment } ??? example "Defaults as YAML" ```yaml arrowflight: deployment: annotations: prometheus.io/path: /metrics prometheus.io/port: '12810' prometheus.io/scheme: http prometheus.io/scrape: 'true' name: arrowflight-deployment replicas: 1 resources: limits: memory: 8192Mi requests: cpu: '2' memory: 6553Mi security: runAsGroup: 1520 runAsUser: 1525 server: hopsfs_query_mode: view memory_limit: '6' queue_timeout: '600' ```
`arrowflight.deployment.annotations."prometheus.io/path"` # { #helm.arrowflight.deployment.annotations.prometheus.io-path } : Type `string`, default `"/metrics"`. `arrowflight.deployment.annotations."prometheus.io/port"` # { #helm.arrowflight.deployment.annotations.prometheus.io-port } : Type `string`, default `"12810"`. `arrowflight.deployment.annotations."prometheus.io/scheme"` # { #helm.arrowflight.deployment.annotations.prometheus.io-scheme } : Type `string`, default `"http"`. `arrowflight.deployment.annotations."prometheus.io/scrape"` # { #helm.arrowflight.deployment.annotations.prometheus.io-scrape } : Type `string`, default `"true"`. `arrowflight.deployment.name` # { #helm.arrowflight.deployment.name } : Type `string`, default `"arrowflight-deployment"`. `arrowflight.deployment.replicas` # { #helm.arrowflight.deployment.replicas } : Type `int`, default `1`. Number of replicas. Not rendered while the HPA is active (hpa.enabled and no VPA on this Deployment): the HPA then owns spec.replicas and this value is its minReplicas. Turning the HPA on for a running release drops the Deployment to 1 once, until the HPA scales it back up. `arrowflight.deployment.resources.limits` # { #helm.arrowflight.deployment.resources.limits } : Type `object`, default `{"memory":"8192Mi"}`. resources limits configuration `arrowflight.deployment.resources.requests` # { #helm.arrowflight.deployment.resources.requests } : Type `object`, default `{"cpu":"2","memory":"6553Mi"}`. resources requests configuration `arrowflight.deployment.security.runAsGroup` # { #helm.arrowflight.deployment.security.runAsGroup } : Type `int`, default `1520`. `arrowflight.deployment.security.runAsUser` # { #helm.arrowflight.deployment.security.runAsUser } : Type `int`, default `1525`. `arrowflight.deployment.server.hopsfs_query_mode` # { #helm.arrowflight.deployment.server.hopsfs_query_mode } : Type `string`, default `"view"`. `arrowflight.deployment.server.memory_limit` # { #helm.arrowflight.deployment.server.memory_limit } : Type `string`, default `"6"`. `arrowflight.deployment.server.queue_timeout` # { #helm.arrowflight.deployment.server.queue_timeout } : Type `string`, default `"600"`.
## externalLoadBalancer { #helm-values-arrowflight-externalloadbalancer } ??? example "Defaults as YAML" ```yaml arrowflight: externalLoadBalancer: annotations: {} class: null enabled: null managed: null nodePort: null nodeSelector: {} ```
`arrowflight.externalLoadBalancer.annotations` # { #helm.arrowflight.externalLoadBalancer.annotations } : Type `object`, default `{}`. annotations for load balancer `arrowflight.externalLoadBalancer.class` # { #helm.arrowflight.externalLoadBalancer.class } : Type `string`, default `nil`. load balancer class name `arrowflight.externalLoadBalancer.enabled` # { #helm.arrowflight.externalLoadBalancer.enabled } : Type `string`, default `nil`. Enable External Load Balancers for Arrowflight server. If not set the .global._hopsworks.externalLoadBalancers.enabled will be used instead `arrowflight.externalLoadBalancer.managed` # { #helm.arrowflight.externalLoadBalancer.managed } : Type `string`, default `nil`. Cloud provider provisions Load Balancers. If not set the .global._hopsworks.externalLoadBalancers.managed will be used instead `arrowflight.externalLoadBalancer.nodePort` # { #helm.arrowflight.externalLoadBalancer.nodePort } : Type `string`, default `nil`. Explicit nodePort for the external service when the load balancer is unmanaged (managed: false), so a load balancer outside Kubernetes can target a fixed port. Null lets Kubernetes allocate one from the cluster's node-port range; a set value must lie in that range (30000-32767 by default), which the API server enforces at install. `arrowflight.externalLoadBalancer.nodeSelector` # { #helm.arrowflight.externalLoadBalancer.nodeSelector } : Type `object`, default `{}`. selector for nodes the load balancer can use to route traffic
## service { #helm-values-arrowflight-service } ??? example "Defaults as YAML" ```yaml arrowflight: service: annotations: consul.hashicorp.com/service-name: flyingduck consul.hashicorp.com/service-port: server consul.hashicorp.com/service-tags: server prometheus.io/path: /metrics prometheus.io/port: '12810' prometheus.io/scheme: http prometheus.io/scrape: 'true' name: arrowflight-server ```
`arrowflight.service.annotations."consul.hashicorp.com/service-name"` # { #helm.arrowflight.service.annotations.consul.hashicorp.com-service-name } : Type `string`, default `"flyingduck"`. `arrowflight.service.annotations."consul.hashicorp.com/service-port"` # { #helm.arrowflight.service.annotations.consul.hashicorp.com-service-port } : Type `string`, default `"server"`. `arrowflight.service.annotations."consul.hashicorp.com/service-tags"` # { #helm.arrowflight.service.annotations.consul.hashicorp.com-service-tags } : Type `string`, default `"server"`. `arrowflight.service.annotations."prometheus.io/path"` # { #helm.arrowflight.service.annotations.prometheus.io-path } : Type `string`, default `"/metrics"`. `arrowflight.service.annotations."prometheus.io/port"` # { #helm.arrowflight.service.annotations.prometheus.io-port } : Type `string`, default `"12810"`. `arrowflight.service.annotations."prometheus.io/scheme"` # { #helm.arrowflight.service.annotations.prometheus.io-scheme } : Type `string`, default `"http"`. `arrowflight.service.annotations."prometheus.io/scrape"` # { #helm.arrowflight.service.annotations.prometheus.io-scrape } : Type `string`, default `"true"`. `arrowflight.service.name` # { #helm.arrowflight.service.name } : Type `string`, default `"arrowflight-server"`.
================================================================================ # certs-operator Source: https://docs.hopsworks.ai/latest/setup_installation/common/helm_chart_values/certs-operator/ # Certs operator values { #helm-values-certs-operator } Values under `certs-operator` configure the operator that issues the TLS certificates of the Hopsworks services and removes them on uninstall. _Generated from the Hopsworks Helm chart `5.2.0-alpha-1791549041` (Hopsworks `5.2.0`)._ Deployed when [`global._hopsworks.full_platform`](global.md#helm.global._hopsworks.full_platform) is `true`. ## General { #helm-values-certs-operator-general } ??? example "Defaults as YAML" ```yaml certs-operator: fullnameOverride: null gracefulCertTeardown: enabled: true timeoutSeconds: 120 ttlSecondsAfterFinished: null hopsworkslib: {} nameOverride: null nodeSelector: {} serviceAccount: create: false name: '' tolerations: [] topologySpreadConstraint: {} ```
`certs-operator` # { #helm.certs-operator } : Type `object`, default `{}`. override certs-operator values `certs-operator.fullnameOverride` # { #helm.certs-operator.fullnameOverride } : Type `string`, default `nil`. override app fully qualified name `certs-operator.gracefulCertTeardown` # { #helm.certs-operator.gracefulCertTeardown } : Type `object`, default `{"enabled":true,"timeoutSeconds":120,"ttlSecondsAfterFinished":null}`. ordered, narrow-RBAC teardown of HopsworksCerts on uninstall: a pre-delete hook deletes the certs while the certs-operator is still alive (so it releases its own finalizers), and a post-delete hook force-clears any finalizers left behind on runtimes that drop pre-delete hooks (ArgoCD < 3.3). Independent of the `wipe` job. `certs-operator.gracefulCertTeardown.enabled` # { #helm.certs-operator.gracefulCertTeardown.enabled } : Type `bool`, default `true`. enable the ordered HopsworksCert teardown hooks on uninstall `certs-operator.gracefulCertTeardown.timeoutSeconds` # { #helm.certs-operator.gracefulCertTeardown.timeoutSeconds } : Type `int`, default `120`. seconds the pre-delete drain waits for the operator to finalize the certs before deferring to the post-delete backstop `certs-operator.gracefulCertTeardown.ttlSecondsAfterFinished` # { #helm.certs-operator.gracefulCertTeardown.ttlSecondsAfterFinished } : Type `string`, default `nil`. ttlSecondsAfterFinished for the teardown Jobs; null falls through to the global default `certs-operator.hopsworkslib` # { #helm.certs-operator.hopsworkslib } : Type `object`, default `{}`. override hopsworkslib values `certs-operator.nameOverride` # { #helm.certs-operator.nameOverride } : Type `string`, default `nil`. override app chart name `certs-operator.nodeSelector` # { #helm.certs-operator.nodeSelector } : Type `object`, default `{}`. node selector configuration `certs-operator.serviceAccount.create` # { #helm.certs-operator.serviceAccount.create } : Type `bool`, default `false`. `certs-operator.serviceAccount.name` # { #helm.certs-operator.serviceAccount.name } : Type `string`, default `""`. `certs-operator.tolerations` # { #helm.certs-operator.tolerations } : Type `list`, default `[]`. `certs-operator.topologySpreadConstraint` # { #helm.certs-operator.topologySpreadConstraint } : Type `object`, default `{}`. The default topology spread constraint. If not defined the global topology spread constraint would be used instead.
## controller { #helm-values-certs-operator-controller } ??? example "Defaults as YAML" ```yaml certs-operator: controller: manager: resources: limits: cpu: 500m memory: 128Mi requests: cpu: 10m memory: 64Mi watchNamespaces: null serviceAccount: annotations: {} ```
`certs-operator.controller.manager.resources.limits.cpu` # { #helm.certs-operator.controller.manager.resources.limits.cpu } : Type `string`, default `"500m"`. `certs-operator.controller.manager.resources.limits.memory` # { #helm.certs-operator.controller.manager.resources.limits.memory } : Type `string`, default `"128Mi"`. `certs-operator.controller.manager.resources.requests.cpu` # { #helm.certs-operator.controller.manager.resources.requests.cpu } : Type `string`, default `"10m"`. `certs-operator.controller.manager.resources.requests.memory` # { #helm.certs-operator.controller.manager.resources.requests.memory } : Type `string`, default `"64Mi"`. `certs-operator.controller.manager.watchNamespaces` # { #helm.certs-operator.controller.manager.watchNamespaces } : Type `string`, default `nil`. Comma separated list of Namespaces to restrict certs-operator to watch for. If not set it will watch all Namespaces. `certs-operator.controller.serviceAccount.annotations` # { #helm.certs-operator.controller.serviceAccount.annotations } : Type `object`, default `{}`. service account annotations
## dependencies { #helm-values-certs-operator-dependencies } ??? example "Defaults as YAML" ```yaml certs-operator: dependencies: ca: apiKeySecretKey: key apiKeySecretName: hopsworks-api-key-auth authMethod: api_key consulServiceName: glassfish consulServiceTag: ca httpScheme: https httpTimeout: 60s password: adminpw port: 8182 user: agent@hops.io ```
`certs-operator.dependencies.ca.apiKeySecretKey` # { #helm.certs-operator.dependencies.ca.apiKeySecretKey } : Type `string`, default `"key"`. `certs-operator.dependencies.ca.apiKeySecretName` # { #helm.certs-operator.dependencies.ca.apiKeySecretName } : Type `string`, default `"hopsworks-api-key-auth"`. `certs-operator.dependencies.ca.authMethod` # { #helm.certs-operator.dependencies.ca.authMethod } : Type `string`, default `"api_key"`. `certs-operator.dependencies.ca.consulServiceName` # { #helm.certs-operator.dependencies.ca.consulServiceName } : Type `string`, default `"glassfish"`. `certs-operator.dependencies.ca.consulServiceTag` # { #helm.certs-operator.dependencies.ca.consulServiceTag } : Type `string`, default `"ca"`. `certs-operator.dependencies.ca.httpScheme` # { #helm.certs-operator.dependencies.ca.httpScheme } : Type `string`, default `"https"`. `certs-operator.dependencies.ca.httpTimeout` # { #helm.certs-operator.dependencies.ca.httpTimeout } : Type `string`, default `"60s"`. `certs-operator.dependencies.ca.password` # { #helm.certs-operator.dependencies.ca.password } : Type `string`, default `"adminpw"`. `certs-operator.dependencies.ca.port` # { #helm.certs-operator.dependencies.ca.port } : Type `int`, default `8182`. `certs-operator.dependencies.ca.user` # { #helm.certs-operator.dependencies.ca.user } : Type `string`, default `"agent@hops.io"`.
================================================================================ # consul Source: https://docs.hopsworks.ai/latest/setup_installation/common/helm_chart_values/consul/ # Consul values { #helm-values-consul } Values under `consul` configure Consul, which provides service discovery and DNS between the Hopsworks services. _Generated from the Hopsworks Helm chart `5.2.0-alpha-1791549041` (Hopsworks `5.2.0`)._ Deployed when [`global._hopsworks.full_platform`](global.md#helm.global._hopsworks.full_platform) is `true`. !!! info "Upstream charts" - Values under `consul.consul` go to [`consul` 1.8.16](https://artifacthub.io/packages/helm/hashicorp/consul/1.8.16) from `https://helm.releases.hashicorp.com`. Only the values Hopsworks sets under `consul.consul` are listed on this page. Any other value of the chart can be set there too; the link opens its documentation for the version Hopsworks pins. ## General { #helm-values-consul-general } ??? example "Defaults as YAML" ```yaml consul: autoConfigureCoreDNS: true autoConfigureCoreDNSPort: 53 autoConfigureJobResources: limits: cpu: 300m autoconfig: serviceAccount: annotations: {} ttlSecondsAfterFinished: null cleanupOnUninstall: enabled: true ttlSecondsAfterFinished: null consul: client: dnsPolicy: ClusterFirstWithHostNet hostNetwork: true connectInject: enabled: false dns: enabled: true global: acls: manageSystemACLs: true nodeSelector: null tolerations: '' datacenter: dc1 domain: consul enabled: true gossipEncryption: autoGenerate: true image: docker.hops.works/hashicorp/consul:1.16.4-alpine-h1 imageK8S: docker.hops.works/hashicorp/consul-k8s-control-plane:1.8.16 logLevel: info metrics: disableAgentHostName: true enableAgentMetrics: true enableGatewayMetrics: false enabled: true name: null tls: enabled: true metrics: enabled: true rbac: annotations: {} create: true extraRoleRules: [] name: consul-metrics-role useExistingRole: false serviceAccount: annotations: {} create: true name: consul-metrics-default server: connect: false enabled: true logLevel: info nodeSelector: null replicas: 3 resources: limits: cpu: 100m memory: 350Mi requests: cpu: 100m memory: 200Mi storageClass: null tolerations: '' topologySpreadConstraints: | - maxSkew: 1 topologyKey: topology.kubernetes.io/zone whenUnsatisfiable: ScheduleAnyway labelSelector: matchLabels: app: consul component: server syncCatalog: default: true enabled: true k8sAllowNamespaces: - '*' nodeSelector: null resources: limits: cpu: 50m memory: 100Mi requests: cpu: 50m memory: 100Mi toConsul: true toK8S: false tolerations: '' ui: enabled: true corednsConfigMapName: coredns corednsDeploymentName: coredns fullnameOverride: null hopsworkslib: {} nameOverride: null networkPolicy: enabled: false nodeSelector: {} tolerations: [] ```
`consul` # { #helm.consul } : Type `object`, default `{"consul":{"server":{"storageClass":null}}}`. override consul values `consul.autoConfigureCoreDNS` # { #helm.consul.autoConfigureCoreDNS } : Type `bool`, default `true`. `consul.autoConfigureCoreDNSPort` # { #helm.consul.autoConfigureCoreDNSPort } : Type `int`, default `53`. `consul.autoConfigureJobResources` # { #helm.consul.autoConfigureJobResources } : Type `object`, default `{"limits":{"cpu":"300m"}}`. resources configuration `consul.autoConfigureJobResources.limits` # { #helm.consul.autoConfigureJobResources.limits } : Type `object`, default `{"cpu":"300m"}`. resource limits configuration `consul.autoconfig.serviceAccount.annotations` # { #helm.consul.autoconfig.serviceAccount.annotations } : Type `object`, default `{}`. service account annotations `consul.autoconfig.ttlSecondsAfterFinished` # { #helm.consul.autoconfig.ttlSecondsAfterFinished } : Type `string`, default `nil`. TTL in seconds for the autoconfig coredns Job. Overrides global default. `consul.cleanupOnUninstall` # { #helm.consul.cleanupOnUninstall } : Type `object`, default `{"enabled":true,"ttlSecondsAfterFinished":null}`. post-delete cleanup of Consul runtime leftovers (consul-k8s ACL/gossip secrets, the metrics-acl-token secret, and the upstream install-hook RBAC) that Helm/ArgoCD never tracked and so never prune. Also deletes Consul server PVCs when global._hopsworks.wipeDataOnUninstall is enabled (except hopsworks.ai/keep=true). `consul.cleanupOnUninstall.enabled` # { #helm.consul.cleanupOnUninstall.enabled } : Type `bool`, default `true`. enable the post-delete Consul cleanup hook `consul.cleanupOnUninstall.ttlSecondsAfterFinished` # { #helm.consul.cleanupOnUninstall.ttlSecondsAfterFinished } : Type `string`, default `nil`. ttlSecondsAfterFinished for the cleanup Job; null falls through to the global default `consul.consul` # { #helm.consul.consul } : Type `object`, passed to the [`consul` 1.8.16](https://artifacthub.io/packages/helm/hashicorp/consul/1.8.16) chart, whose other values are documented there. override consul values ??? note "Default" ```yaml client: dnsPolicy: ClusterFirstWithHostNet hostNetwork: true connectInject: enabled: false dns: enabled: true global: acls: manageSystemACLs: true nodeSelector: null tolerations: '' datacenter: dc1 domain: consul enabled: true gossipEncryption: autoGenerate: true image: docker.hops.works/hashicorp/consul:1.16.4-alpine-h1 imageK8S: docker.hops.works/hashicorp/consul-k8s-control-plane:1.8.16 logLevel: info metrics: disableAgentHostName: true enableAgentMetrics: true enableGatewayMetrics: false enabled: true name: null tls: enabled: true metrics: enabled: true rbac: annotations: {} create: true extraRoleRules: [] name: consul-metrics-role useExistingRole: false serviceAccount: annotations: {} create: true name: consul-metrics-default server: connect: false enabled: true logLevel: info nodeSelector: null replicas: 3 resources: limits: cpu: 100m memory: 350Mi requests: cpu: 100m memory: 200Mi storageClass: null tolerations: '' topologySpreadConstraints: | - maxSkew: 1 topologyKey: topology.kubernetes.io/zone whenUnsatisfiable: ScheduleAnyway labelSelector: matchLabels: app: consul component: server syncCatalog: default: true enabled: true k8sAllowNamespaces: - '*' nodeSelector: null resources: limits: cpu: 50m memory: 100Mi requests: cpu: 50m memory: 100Mi toConsul: true toK8S: false tolerations: '' ui: enabled: true ``` `consul.corednsConfigMapName` # { #helm.consul.corednsConfigMapName } : Type `string`, default `"coredns"`. the name of the coredns configmap. This name is only used when global._hopsworks.cloudProvider does not equal to AZURE, OVH, or GCP. For AZURE and OVH, coredns-custom is being used by default and for GCP, kube-dns is being used by default. `consul.corednsDeploymentName` # { #helm.consul.corednsDeploymentName } : Type `string`, default `"coredns"`. the name of the coredns deployment. This name is only used when global._hopsworks.cloudProvider does not equal to AZURE, OVH, or GCP. `consul.fullnameOverride` # { #helm.consul.fullnameOverride } : Type `string`, default `nil`. override app fully qualified name `consul.hopsworkslib` # { #helm.consul.hopsworkslib } : Type `object`, default `{}`. override hopsworkslib values `consul.nameOverride` # { #helm.consul.nameOverride } : Type `string`, default `nil`. override app chart name `consul.networkPolicy.enabled` # { #helm.consul.networkPolicy.enabled } : Type `bool`, default `false`. `consul.nodeSelector` # { #helm.consul.nodeSelector } : Type `object`, default `{}`. node selector configuration used for the autoconfig job and register managed docker registry jobs `consul.tolerations` # { #helm.consul.tolerations } : Type `list`, default `[]`.
## manualServiceRegistration { #helm-values-consul-manualserviceregistration } ??? example "Defaults as YAML" ```yaml consul: manualServiceRegistration: resources: limits: cpu: 100m memory: 100Mi requests: cpu: 50m memory: 30Mi ttlSecondsAfterFinished: null ```
`consul.manualServiceRegistration.resources.limits.cpu` # { #helm.consul.manualServiceRegistration.resources.limits.cpu } : Type `string`, default `"100m"`. `consul.manualServiceRegistration.resources.limits.memory` # { #helm.consul.manualServiceRegistration.resources.limits.memory } : Type `string`, default `"100Mi"`. `consul.manualServiceRegistration.resources.requests.cpu` # { #helm.consul.manualServiceRegistration.resources.requests.cpu } : Type `string`, default `"50m"`. `consul.manualServiceRegistration.resources.requests.memory` # { #helm.consul.manualServiceRegistration.resources.requests.memory } : Type `string`, default `"30Mi"`. `consul.manualServiceRegistration.ttlSecondsAfterFinished` # { #helm.consul.manualServiceRegistration.ttlSecondsAfterFinished } : Type `string`, default `nil`. TTL in seconds for manual service registration Jobs. Overrides global default.
================================================================================ # docker-registry Source: https://docs.hopsworks.ai/latest/setup_installation/common/helm_chart_values/docker-registry/ # Docker registry values { #helm-values-docker-registry } Values under `docker-registry` configure the in-cluster Docker registry that stores the images Hopsworks builds, such as project Python environments. _Generated from the Hopsworks Helm chart `5.2.0-alpha-1791549041` (Hopsworks `5.2.0`)._ Deployed when [`global._hopsworks.full_platform`](global.md#helm.global._hopsworks.full_platform) is `true`. ## General { #helm-values-docker-registry-general } ??? example "Defaults as YAML" ```yaml docker-registry: affinity: {} appName: docker-registry debug: false dependencies: objectStorage: consulServiceName: minio port: 9000 enabled: true hopsworkslib: {} image: name: registry pullPolicy: IfNotPresent registry: docker.hops.works tag: 3.1.1 nodeSelector: {} podDisruptionBudget: enabled: true minAvailable: 1 replicas: 1 resources: limits: cpu: '2' memory: 10G requests: cpu: '2' memory: 4G security: secret: docker tls: enabled: true service: headless: name: registry-headless monitoringPort: 5001 nodePort: annotations: consul.hashicorp.com/service-name: registry port: 30443 storage: null storageClassName: null tolerations: [] ```
`docker-registry` # { #helm.docker-registry } : Type `object`, default `{"storageClassName":null}`. override docker-registry values `docker-registry.affinity` # { #helm.docker-registry.affinity } : Type `object`, default `{}`. affinity configuration `docker-registry.appName` # { #helm.docker-registry.appName } : Type `string`, default `"docker-registry"`. `docker-registry.debug` # { #helm.docker-registry.debug } : Type `bool`, default `false`. `docker-registry.dependencies.objectStorage.consulServiceName` # { #helm.docker-registry.dependencies.objectStorage.consulServiceName } : Type `string`, default `"minio"`. `docker-registry.dependencies.objectStorage.port` # { #helm.docker-registry.dependencies.objectStorage.port } : Type `int`, default `9000`. `docker-registry.enabled` # { #helm.docker-registry.enabled } : Type `bool`, default `true`. `docker-registry.hopsworkslib` # { #helm.docker-registry.hopsworkslib } : Type `object`, default `{}`. override hopsworkslib values `docker-registry.image.name` # { #helm.docker-registry.image.name } : Type `string`, default `"registry"`. `docker-registry.image.pullPolicy` # { #helm.docker-registry.image.pullPolicy } : Type `string`, default `"IfNotPresent"`. `docker-registry.image.registry` # { #helm.docker-registry.image.registry } : Type `string`, default `"docker.hops.works"`. `docker-registry.image.tag` # { #helm.docker-registry.image.tag } : Type `string`, default `"3.1.1"`. Distribution v3. The chart configures the registry only through `REGISTRY_*` environment variables and mounts no config file, so v3's new default config path does not affect it. `docker-registry.nodeSelector` # { #helm.docker-registry.nodeSelector } : Type `object`, default `{}`. node selector configuration `docker-registry.podDisruptionBudget.enabled` # { #helm.docker-registry.podDisruptionBudget.enabled } : Type `bool`, default `true`. `docker-registry.podDisruptionBudget.minAvailable` # { #helm.docker-registry.podDisruptionBudget.minAvailable } : Type `int`, default `1`. `docker-registry.replicas` # { #helm.docker-registry.replicas } : Type `int`, default `1`. Number of replicas. Not rendered while the HPA is active (hpa.enabled and no VPA on the registry): the HPA then owns spec.replicas, starting from hpa.minReplicas. Turning the HPA on for a running release drops the StatefulSet to 1 once, until the HPA scales it back up. `docker-registry.resources.limits.cpu` # { #helm.docker-registry.resources.limits.cpu } : Type `string`, default `"2"`. `docker-registry.resources.limits.memory` # { #helm.docker-registry.resources.limits.memory } : Type `string`, default `"10G"`. `docker-registry.resources.requests.cpu` # { #helm.docker-registry.resources.requests.cpu } : Type `string`, default `"2"`. `docker-registry.resources.requests.memory` # { #helm.docker-registry.resources.requests.memory } : Type `string`, default `"4G"`. `docker-registry.security.secret` # { #helm.docker-registry.security.secret } : Type `string`, default `"docker"`. `docker-registry.security.tls.enabled` # { #helm.docker-registry.security.tls.enabled } : Type `bool`, default `true`. `docker-registry.service.headless.name` # { #helm.docker-registry.service.headless.name } : Type `string`, default `"registry-headless"`. `docker-registry.service.monitoringPort` # { #helm.docker-registry.service.monitoringPort } : Type `int`, default `5001`. Port of the registry's debug server. Serves `/metrics`, and also `/debug/pprof`, which is why its Service is ClusterIP. `docker-registry.service.nodePort.annotations."consul.hashicorp.com/service-name"` # { #helm.docker-registry.service.nodePort.annotations.consul.hashicorp.com-service-name } : Type `string`, default `"registry"`. `docker-registry.service.port` # { #helm.docker-registry.service.port } : Type `int`, default `30443`. `docker-registry.storage` # { #helm.docker-registry.storage } : Type `string`, default `nil`. storage class name. Set if the REGISTRY_STORAGE is not S3 `docker-registry.storageClassName` # { #helm.docker-registry.storageClassName } : Type `string`, default `nil`. storage class name `docker-registry.tolerations` # { #helm.docker-registry.tolerations } : Type `list`, default `[]`.
## config { #helm-values-docker-registry-config } ??? example "Defaults as YAML" ```yaml docker-registry: config: - name: REGISTRY_STORAGE_DELETE_ENABLED value: 'true' - name: REGISTRY_STORAGE_CACHE_BLOBDESCRIPTOR value: inmemory - name: REGISTRY_STORAGE_CACHE_BLOBDESCRIPTORSIZE value: '10' - name: REGISTRY_STORAGE value: s3 - name: REGISTRY_STORAGE_S3_ACCESSKEY value: minioadmin - name: REGISTRY_STORAGE_S3_SECRETKEY value: minioadmin - name: REGISTRY_STORAGE_S3_REGION value: eu-west-1 - name: REGISTRY_STORAGE_S3_BUCKET value: public - name: REGISTRY_STORAGE_S3_ROOTDIRECTORY value: / - name: REGISTRY_STORAGE_S3_SECURE value: 'false' - name: REGISTRY_STORAGE_S3_V4AUTH value: 'true' - name: REGISTRY_STORAGE_S3_FORCEPATHSTYLE value: 'true' - name: REGISTRY_STORAGE_S3_LOGLEVEL value: debug - name: REGISTRY_STORAGE_S3_CHUNKSIZE value: '209715200' - name: REGISTRY_HTTP_DRAINTIMEOUT value: 10m - name: REGISTRY_VALIDATION_DISABLED value: 'true' - name: REGISTRY_STORAGE_REDIRECT_DISABLE value: 'true' - name: REGISTRY_LOG_LEVEL value: info - name: REGISTRY_HTTP_HEADERS value: '{X-Content-Type-Options: [nosniff]}' ```
`docker-registry.config[0].name` # { #helm.docker-registry.config.0.name } : Type `string`, default `"REGISTRY_STORAGE_DELETE_ENABLED"`. `docker-registry.config[0].value` # { #helm.docker-registry.config.0.value } : Type `string`, default `"true"`. `docker-registry.config[10].name` # { #helm.docker-registry.config.10.name } : Type `string`, default `"REGISTRY_STORAGE_S3_V4AUTH"`. `docker-registry.config[10].value` # { #helm.docker-registry.config.10.value } : Type `string`, default `"true"`. `docker-registry.config[11].name` # { #helm.docker-registry.config.11.name } : Type `string`, default `"REGISTRY_STORAGE_S3_FORCEPATHSTYLE"`. `docker-registry.config[11].value` # { #helm.docker-registry.config.11.value } : Type `string`, default `"true"`. `docker-registry.config[12].name` # { #helm.docker-registry.config.12.name } : Type `string`, default `"REGISTRY_STORAGE_S3_LOGLEVEL"`. `docker-registry.config[12].value` # { #helm.docker-registry.config.12.value } : Type `string`, default `"debug"`. `docker-registry.config[13].name` # { #helm.docker-registry.config.13.name } : Type `string`, default `"REGISTRY_STORAGE_S3_CHUNKSIZE"`. `docker-registry.config[13].value` # { #helm.docker-registry.config.13.value } : Type `string`, default `"209715200"`. `docker-registry.config[14].name` # { #helm.docker-registry.config.14.name } : Type `string`, default `"REGISTRY_HTTP_DRAINTIMEOUT"`. `docker-registry.config[14].value` # { #helm.docker-registry.config.14.value } : Type `string`, default `"10m"`. `docker-registry.config[15].name` # { #helm.docker-registry.config.15.name } : Type `string`, default `"REGISTRY_VALIDATION_DISABLED"`. `docker-registry.config[15].value` # { #helm.docker-registry.config.15.value } : Type `string`, default `"true"`. `docker-registry.config[16].name` # { #helm.docker-registry.config.16.name } : Type `string`, default `"REGISTRY_STORAGE_REDIRECT_DISABLE"`. `docker-registry.config[16].value` # { #helm.docker-registry.config.16.value } : Type `string`, default `"true"`. `docker-registry.config[17].name` # { #helm.docker-registry.config.17.name } : Type `string`, default `"REGISTRY_LOG_LEVEL"`. `docker-registry.config[17].value` # { #helm.docker-registry.config.17.value } : Type `string`, default `"info"`. `docker-registry.config[18].name` # { #helm.docker-registry.config.18.name } : Type `string`, default `"REGISTRY_HTTP_HEADERS"`. `docker-registry.config[18].value` # { #helm.docker-registry.config.18.value } : Type `string`, default `"{X-Content-Type-Options: [nosniff]}"`. `docker-registry.config[1].name` # { #helm.docker-registry.config.1.name } : Type `string`, default `"REGISTRY_STORAGE_CACHE_BLOBDESCRIPTOR"`. `docker-registry.config[1].value` # { #helm.docker-registry.config.1.value } : Type `string`, default `"inmemory"`. `docker-registry.config[2].name` # { #helm.docker-registry.config.2.name } : Type `string`, default `"REGISTRY_STORAGE_CACHE_BLOBDESCRIPTORSIZE"`. `docker-registry.config[2].value` # { #helm.docker-registry.config.2.value } : Type `string`, default `"10"`. `docker-registry.config[3].name` # { #helm.docker-registry.config.3.name } : Type `string`, default `"REGISTRY_STORAGE"`. `docker-registry.config[3].value` # { #helm.docker-registry.config.3.value } : Type `string`, default `"s3"`. `docker-registry.config[4].name` # { #helm.docker-registry.config.4.name } : Type `string`, default `"REGISTRY_STORAGE_S3_ACCESSKEY"`. `docker-registry.config[4].value` # { #helm.docker-registry.config.4.value } : Type `string`, default `"minioadmin"`. `docker-registry.config[5].name` # { #helm.docker-registry.config.5.name } : Type `string`, default `"REGISTRY_STORAGE_S3_SECRETKEY"`. `docker-registry.config[5].value` # { #helm.docker-registry.config.5.value } : Type `string`, default `"minioadmin"`. `docker-registry.config[6].name` # { #helm.docker-registry.config.6.name } : Type `string`, default `"REGISTRY_STORAGE_S3_REGION"`. `docker-registry.config[6].value` # { #helm.docker-registry.config.6.value } : Type `string`, default `"eu-west-1"`. `docker-registry.config[7].name` # { #helm.docker-registry.config.7.name } : Type `string`, default `"REGISTRY_STORAGE_S3_BUCKET"`. `docker-registry.config[7].value` # { #helm.docker-registry.config.7.value } : Type `string`, default `"public"`. `docker-registry.config[8].name` # { #helm.docker-registry.config.8.name } : Type `string`, default `"REGISTRY_STORAGE_S3_ROOTDIRECTORY"`. `docker-registry.config[8].value` # { #helm.docker-registry.config.8.value } : Type `string`, default `"/"`. `docker-registry.config[9].name` # { #helm.docker-registry.config.9.name } : Type `string`, default `"REGISTRY_STORAGE_S3_SECURE"`. `docker-registry.config[9].value` # { #helm.docker-registry.config.9.value } : Type `string`, default `"false"`.
## hpa { #helm-values-docker-registry-hpa } ??? example "Defaults as YAML" ```yaml docker-registry: hpa: enabled: true maxReplicas: 4 minReplicas: 1 targetCPUUtilizationPercentage: 60 targetMemoryUtilizationPercentage: 60 ```
`docker-registry.hpa.enabled` # { #helm.docker-registry.hpa.enabled } : Type `bool`, default `true`. Autoscale the registry with an HPA (suppressed when a VPA targets it). While active, replicas is not rendered and the HPA owns spec.replicas. `docker-registry.hpa.maxReplicas` # { #helm.docker-registry.hpa.maxReplicas } : Type `int`, default `4`. `docker-registry.hpa.minReplicas` # { #helm.docker-registry.hpa.minReplicas } : Type `int`, default `1`. `docker-registry.hpa.targetCPUUtilizationPercentage` # { #helm.docker-registry.hpa.targetCPUUtilizationPercentage } : Type `int`, default `60`. `docker-registry.hpa.targetMemoryUtilizationPercentage` # { #helm.docker-registry.hpa.targetMemoryUtilizationPercentage } : Type `int`, default `60`.
## tamperContainerEngine { #helm-values-docker-registry-tampercontainerengine } ??? example "Defaults as YAML" ```yaml docker-registry: tamperContainerEngine: containerdConfigFile: config.toml defaultServer: null enabled: true fallbackTo: https://docker.hops.works patchEngine: true patchEtcHosts: true patchOS: true registryProtocol: https rollbackChangesIfDeleted: true tolerations: - effect: NoSchedule operator: Exists - key: CriticalAddonsOnly operator: Exists - effect: NoExecute operator: Exists trustCA: USE_ROOT_HOPSWORKS_CA verifyTLS: true waitBeforeReset: 0 ```
`docker-registry.tamperContainerEngine.containerdConfigFile` # { #helm.docker-registry.tamperContainerEngine.containerdConfigFile } : Type `string`, default `"config.toml"`. `docker-registry.tamperContainerEngine.defaultServer` # { #helm.docker-registry.tamperContainerEngine.defaultServer } : Type `string`, default `nil`. default server to use in container engine. Change it to docker.hops.works if needed `docker-registry.tamperContainerEngine.enabled` # { #helm.docker-registry.tamperContainerEngine.enabled } : Type `bool`, default `true`. `docker-registry.tamperContainerEngine.fallbackTo` # { #helm.docker-registry.tamperContainerEngine.fallbackTo } : Type `string`, default `"https://docker.hops.works"`. `docker-registry.tamperContainerEngine.patchEngine` # { #helm.docker-registry.tamperContainerEngine.patchEngine } : Type `bool`, default `true`. `docker-registry.tamperContainerEngine.patchEtcHosts` # { #helm.docker-registry.tamperContainerEngine.patchEtcHosts } : Type `bool`, default `true`. `docker-registry.tamperContainerEngine.patchOS` # { #helm.docker-registry.tamperContainerEngine.patchOS } : Type `bool`, default `true`. `docker-registry.tamperContainerEngine.registryProtocol` # { #helm.docker-registry.tamperContainerEngine.registryProtocol } : Type `string`, default `"https"`. `docker-registry.tamperContainerEngine.rollbackChangesIfDeleted` # { #helm.docker-registry.tamperContainerEngine.rollbackChangesIfDeleted } : Type `bool`, default `true`. `docker-registry.tamperContainerEngine.tolerations[0].effect` # { #helm.docker-registry.tamperContainerEngine.tolerations.0.effect } : Type `string`, default `"NoSchedule"`. `docker-registry.tamperContainerEngine.tolerations[0].operator` # { #helm.docker-registry.tamperContainerEngine.tolerations.0.operator } : Type `string`, default `"Exists"`. `docker-registry.tamperContainerEngine.tolerations[1].key` # { #helm.docker-registry.tamperContainerEngine.tolerations.1.key } : Type `string`, default `"CriticalAddonsOnly"`. `docker-registry.tamperContainerEngine.tolerations[1].operator` # { #helm.docker-registry.tamperContainerEngine.tolerations.1.operator } : Type `string`, default `"Exists"`. `docker-registry.tamperContainerEngine.tolerations[2].effect` # { #helm.docker-registry.tamperContainerEngine.tolerations.2.effect } : Type `string`, default `"NoExecute"`. `docker-registry.tamperContainerEngine.tolerations[2].operator` # { #helm.docker-registry.tamperContainerEngine.tolerations.2.operator } : Type `string`, default `"Exists"`. `docker-registry.tamperContainerEngine.trustCA` # { #helm.docker-registry.tamperContainerEngine.trustCA } : Type `string`, default `"USE_ROOT_HOPSWORKS_CA"`. `docker-registry.tamperContainerEngine.verifyTLS` # { #helm.docker-registry.tamperContainerEngine.verifyTLS } : Type `bool`, default `true`. `docker-registry.tamperContainerEngine.waitBeforeReset` # { #helm.docker-registry.tamperContainerEngine.waitBeforeReset } : Type `int`, default `0`.
================================================================================ # grafana Source: https://docs.hopsworks.ai/latest/setup_installation/common/helm_chart_values/grafana/ # Grafana values { #helm-values-grafana } Values under `grafana` configure Grafana and the Hopsworks monitoring dashboards. _Generated from the Hopsworks Helm chart `5.2.0-alpha-1791549041` (Hopsworks `5.2.0`)._ Deployed when [`global._hopsworks.full_platform`](global.md#helm.global._hopsworks.full_platform) is `true`. !!! info "Upstream charts" - Values under `grafana.grafana` go to [`grafana` 7.0.17](https://artifacthub.io/packages/helm/grafana/grafana/7.0.17) from `https://grafana.github.io/helm-charts`. Only the values Hopsworks sets under `grafana.grafana` are listed on this page. Any other value of the chart can be set there too; the link opens its documentation for the version Hopsworks pins. ??? example "Defaults as YAML" ```yaml grafana: dependencies: prometheus: consulServiceName: prometheus consulServiceTag: prometheus port: 9090 grafana: dashboardProviders: dashboardproviders.yaml: apiVersion: 1 providers: - disableDeletion: true editable: false folder: Apps name: Apps options: path: /usr/share/grafana/dashboards/apps type: file updateIntervalSeconds: 10 - disableDeletion: true editable: false folder: Hops name: Hops options: path: /usr/share/grafana/dashboards/hops type: file updateIntervalSeconds: 10 - disableDeletion: false editable: false folder: RonDB name: RonDB options: path: /usr/share/grafana/dashboards/rondb type: file updateIntervalSeconds: 10 - disableDeletion: true editable: false folder: Overview name: Overview options: path: /usr/share/grafana/dashboards/overview type: file updateIntervalSeconds: 10 - disableDeletion: true editable: false folder: Kubernetes name: Kubernetes options: path: /usr/share/grafana/dashboards/kubernetes type: file updateIntervalSeconds: 10 - disableDeletion: true editable: false folder: ModelServing name: ModelServing options: path: /usr/share/grafana/dashboards/kserve type: file updateIntervalSeconds: 10 - disableDeletion: true editable: false folder: Ray name: Ray options: path: /usr/share/grafana/dashboards/ray type: file updateIntervalSeconds: 10 - disableDeletion: true editable: false folder: RSS name: RemoteShuffleService options: path: /usr/share/grafana/dashboards/rss type: file updateIntervalSeconds: 10 dashboardsConfigMaps: {} datasources: datasources.yaml: apiVersion: 1 datasources: - access: proxy editable: false isDefault: true name: Prometheus type: prometheus url: http://prometheus.prometheus.service.consul:9090 downloadDashboardsImage: pullPolicy: IfNotPresent registry: docker.hops.works repository: hopsworks/hwutils sha: '' tag: 1.10-SNAPSHOT extraConfigmapMounts: - configMap: '{{ include "hopsworks.grafana.conditionalProvidersName" . }}' mountPath: /etc/grafana/provisioning/dashboards/hopsworks-dashboardproviders.yaml name: hopsworks-dashboardproviders readOnly: true subPath: providers.yaml extraInitContainers: - command: - /bin/sh - -c - | set -eu checked=0 for f in /etc/grafana/provisioning/dashboards/*.yaml; do [ -f "$f" ] || continue for p in $(sed -n 's/^[[:space:]]*path:[[:space:]]*//p' "$f"); do case "$p" in /usr/share/grafana/dashboards/*) ;; *) echo "skip: $p is delivered by Helm, not by the image"; continue ;; esac checked=$((checked + 1)) n=$(find "$p" -name '*.json' 2>/dev/null | wc -l) if [ "$n" -eq 0 ]; then echo "FATAL: dashboard provider path $p (declared in $f) holds no dashboards." echo "The Hopsworks dashboards ship inside the Grafana image. This image does not" echo "carry them, so Grafana would start healthy with an empty dashboard list." echo "Use an image built with the dashboards, or override grafana.image.tag." exit 1 fi echo "ok: $p ($n dashboards)" done done if [ "$checked" -eq 0 ]; then echo "FATAL: no image dashboard provider paths found in /etc/grafana/provisioning/dashboards." echo "The chart always declares providers under /usr/share/grafana/dashboards, so either" echo "the provisioning files did not reach this container or every provider was replaced." exit 1 fi echo "verified $checked image dashboard provider paths" image: '{{ .Values.global.imageRegistry | default .Values.image.registry }}/{{ .Values.image.repository }}:{{ .Values.image.tag | default .Chart.AppVersion }}' imagePullPolicy: '{{ .Values.image.pullPolicy }}' name: verify-dashboards resources: limits: cpu: 100m memory: 64Mi requests: cpu: 10m memory: 32Mi securityContext: allowPrivilegeEscalation: false capabilities: drop: - ALL readOnlyRootFilesystem: true seccompProfile: type: RuntimeDefault volumeMounts: - mountPath: /etc/grafana/provisioning/dashboards/dashboardproviders.yaml name: config subPath: dashboardproviders.yaml - mountPath: /etc/grafana/provisioning/dashboards/hopsworks-dashboardproviders.yaml name: hopsworks-dashboardproviders subPath: providers.yaml global: imageRegistry: docker.hops.works grafana.ini: auth: disable_login_form: true disable_signout_menu: true auth.anonymous: enabled: false auth.basic: enabled: false auth.proxy: auto_sign_up: true enable_login_token: false enabled: true header_name: X-WEBAUTH-USER header_property: username headers: Name:X-WEBAUTH-NAME Role:X-WEBAUTH-ROLE Email:X-WEBAUTH-EMAIL headers_encoded: false sync_ttl: '60' whitelist: null rbac: enabled: true security: admin_password: adminpw allow_embedding: true strict_transport_security: false server: enforce_domain: false root_url: /hopsworks-api/grafana users: allow_org_create: false allow_sign_up: false auto_assign_org: true auto_assign_org_id: '1' auto_assign_org_role: Viewer default_theme: dark editors_can_admin: false home_page: /dashboards verify_email_enabled: false viewers_can_edit: false image: tag: 12.4.9-h10 nodeSelector: {} rbac: create: false resources: limits: cpu: 1 memory: 1000Mi requests: cpu: 200m memory: 200Mi service: annotations: consul.hashicorp.com/service-name: grafana consul.hashicorp.com/service-tags: grafana tolerations: [] topologySpreadConstraints: - labelSelector: matchLabels: app.kubernetes.io/instance: hopsworks app.kubernetes.io/name: grafana maxSkew: 1 topologyKey: topology.kubernetes.io/zone whenUnsatisfiable: ScheduleAnyway hopsworkslib: {} ```
`grafana` # { #helm.grafana } : Type `object`, default `{"grafana":{"global":{"imageRegistry":"docker.hops.works"}}}`. override grafana values `grafana.dependencies.prometheus.consulServiceName` # { #helm.grafana.dependencies.prometheus.consulServiceName } : Type `string`, default `"prometheus"`. `grafana.dependencies.prometheus.consulServiceTag` # { #helm.grafana.dependencies.prometheus.consulServiceTag } : Type `string`, default `"prometheus"`. `grafana.dependencies.prometheus.port` # { #helm.grafana.dependencies.prometheus.port } : Type `int`, default `9090`. `grafana.grafana` # { #helm.grafana.grafana } : Type `object`, passed to the [`grafana` 7.0.17](https://artifacthub.io/packages/helm/grafana/grafana/7.0.17) chart, whose other values are documented there. override grafana values ??? note "Default" ```yaml dashboardProviders: dashboardproviders.yaml: apiVersion: 1 providers: - disableDeletion: true editable: false folder: Apps name: Apps options: path: /usr/share/grafana/dashboards/apps type: file updateIntervalSeconds: 10 - disableDeletion: true editable: false folder: Hops name: Hops options: path: /usr/share/grafana/dashboards/hops type: file updateIntervalSeconds: 10 - disableDeletion: false editable: false folder: RonDB name: RonDB options: path: /usr/share/grafana/dashboards/rondb type: file updateIntervalSeconds: 10 - disableDeletion: true editable: false folder: Overview name: Overview options: path: /usr/share/grafana/dashboards/overview type: file updateIntervalSeconds: 10 - disableDeletion: true editable: false folder: Kubernetes name: Kubernetes options: path: /usr/share/grafana/dashboards/kubernetes type: file updateIntervalSeconds: 10 - disableDeletion: true editable: false folder: ModelServing name: ModelServing options: path: /usr/share/grafana/dashboards/kserve type: file updateIntervalSeconds: 10 - disableDeletion: true editable: false folder: Ray name: Ray options: path: /usr/share/grafana/dashboards/ray type: file updateIntervalSeconds: 10 - disableDeletion: true editable: false folder: RSS name: RemoteShuffleService options: path: /usr/share/grafana/dashboards/rss type: file updateIntervalSeconds: 10 dashboardsConfigMaps: {} datasources: datasources.yaml: apiVersion: 1 datasources: - access: proxy editable: false isDefault: true name: Prometheus type: prometheus url: http://prometheus.prometheus.service.consul:9090 downloadDashboardsImage: pullPolicy: IfNotPresent registry: docker.hops.works repository: hopsworks/hwutils sha: '' tag: 1.10-SNAPSHOT extraConfigmapMounts: - configMap: '{{ include "hopsworks.grafana.conditionalProvidersName" . }}' mountPath: /etc/grafana/provisioning/dashboards/hopsworks-dashboardproviders.yaml name: hopsworks-dashboardproviders readOnly: true subPath: providers.yaml extraInitContainers: - command: - /bin/sh - -c - | set -eu checked=0 for f in /etc/grafana/provisioning/dashboards/*.yaml; do [ -f "$f" ] || continue for p in $(sed -n 's/^[[:space:]]*path:[[:space:]]*//p' "$f"); do case "$p" in /usr/share/grafana/dashboards/*) ;; *) echo "skip: $p is delivered by Helm, not by the image"; continue ;; esac checked=$((checked + 1)) n=$(find "$p" -name '*.json' 2>/dev/null | wc -l) if [ "$n" -eq 0 ]; then echo "FATAL: dashboard provider path $p (declared in $f) holds no dashboards." echo "The Hopsworks dashboards ship inside the Grafana image. This image does not" echo "carry them, so Grafana would start healthy with an empty dashboard list." echo "Use an image built with the dashboards, or override grafana.image.tag." exit 1 fi echo "ok: $p ($n dashboards)" done done if [ "$checked" -eq 0 ]; then echo "FATAL: no image dashboard provider paths found in /etc/grafana/provisioning/dashboards." echo "The chart always declares providers under /usr/share/grafana/dashboards, so either" echo "the provisioning files did not reach this container or every provider was replaced." exit 1 fi echo "verified $checked image dashboard provider paths" image: '{{ .Values.global.imageRegistry | default .Values.image.registry }}/{{ .Values.image.repository }}:{{ .Values.image.tag | default .Chart.AppVersion }}' imagePullPolicy: '{{ .Values.image.pullPolicy }}' name: verify-dashboards resources: limits: cpu: 100m memory: 64Mi requests: cpu: 10m memory: 32Mi securityContext: allowPrivilegeEscalation: false capabilities: drop: - ALL readOnlyRootFilesystem: true seccompProfile: type: RuntimeDefault volumeMounts: - mountPath: /etc/grafana/provisioning/dashboards/dashboardproviders.yaml name: config subPath: dashboardproviders.yaml - mountPath: /etc/grafana/provisioning/dashboards/hopsworks-dashboardproviders.yaml name: hopsworks-dashboardproviders subPath: providers.yaml global: imageRegistry: docker.hops.works grafana.ini: auth: disable_login_form: true disable_signout_menu: true auth.anonymous: enabled: false auth.basic: enabled: false auth.proxy: auto_sign_up: true enable_login_token: false enabled: true header_name: X-WEBAUTH-USER header_property: username headers: Name:X-WEBAUTH-NAME Role:X-WEBAUTH-ROLE Email:X-WEBAUTH-EMAIL headers_encoded: false sync_ttl: '60' whitelist: null rbac: enabled: true security: admin_password: adminpw allow_embedding: true strict_transport_security: false server: enforce_domain: false root_url: /hopsworks-api/grafana users: allow_org_create: false allow_sign_up: false auto_assign_org: true auto_assign_org_id: '1' auto_assign_org_role: Viewer default_theme: dark editors_can_admin: false home_page: /dashboards verify_email_enabled: false viewers_can_edit: false image: tag: 12.4.9-h10 nodeSelector: {} rbac: create: false resources: limits: cpu: 1 memory: 1000Mi requests: cpu: 200m memory: 200Mi service: annotations: consul.hashicorp.com/service-name: grafana consul.hashicorp.com/service-tags: grafana tolerations: [] topologySpreadConstraints: - labelSelector: matchLabels: app.kubernetes.io/instance: hopsworks app.kubernetes.io/name: grafana maxSkew: 1 topologyKey: topology.kubernetes.io/zone whenUnsatisfiable: ScheduleAnyway ``` `grafana.hopsworkslib` # { #helm.grafana.hopsworkslib } : Type `object`, default `{}`. override hopsworkslib values `grafana.grafana.dashboardsConfigMaps` Deprecated # { #helm.grafana.grafana.dashboardsConfigMaps } : Type `object`, default `{}`. DEPRECATED and intentionally empty. The dashboards used to be inlined into nine ConfigMaps here, which put 1.9 MB of JSON into the rendered manifest and so into the 1 MiB Helm release Secret. They now ship in the Grafana image under /usr/share/grafana/dashboards and are provisioned by path. Kept as an empty map rather than removed so that an existing override does not become an unknown key.
================================================================================ # hive Source: https://docs.hopsworks.ai/latest/setup_installation/common/helm_chart_values/hive/ # Hive values { #helm-values-hive } Values under `hive` configure the Hive metastore, which holds the table metadata of the offline feature store. _Generated from the Hopsworks Helm chart `5.2.0-alpha-1791549041` (Hopsworks `5.2.0`)._ Deployed when [`global._hopsworks.full_platform`](global.md#helm.global._hopsworks.full_platform) is `true`. ## General { #helm-values-hive-general } ??? example "Defaults as YAML" ```yaml hive: debug: false hopsworkslib: {} name: hive ```
`hive` # { #helm.hive } : Type `object`, default `{}`. override hive values `hive.debug` # { #helm.hive.debug } : Type `bool`, default `false`. `hive.hopsworkslib` # { #helm.hive.hopsworkslib } : Type `object`, default `{}`. override hopsworkslib values `hive.name` # { #helm.hive.name } : Type `string`, default `"hive"`.
## externalLoadBalancer { #helm-values-hive-externalloadbalancer } ??? example "Defaults as YAML" ```yaml hive: externalLoadBalancer: annotations: {} class: null enabled: false managed: null nodePort: null nodeSelector: {} ```
`hive.externalLoadBalancer.annotations` # { #helm.hive.externalLoadBalancer.annotations } : Type `object`, default `{}`. annotations for load balancer `hive.externalLoadBalancer.class` # { #helm.hive.externalLoadBalancer.class } : Type `string`, default `nil`. load balancer class name `hive.externalLoadBalancer.enabled` # { #helm.hive.externalLoadBalancer.enabled } : Type `bool`, default `false`. Enable External Load Balancer for the Hive Metastore. If not set the .global._hopsworks.externalLoadBalancers.enabled will be used instead `hive.externalLoadBalancer.managed` # { #helm.hive.externalLoadBalancer.managed } : Type `string`, default `nil`. Cloud provider provisions Load Balancers. If not set the .global._hopsworks.externalLoadBalancers.managed will be used instead `hive.externalLoadBalancer.nodePort` # { #helm.hive.externalLoadBalancer.nodePort } : Type `string`, default `nil`. Explicit nodePort for the external service when the load balancer is unmanaged (managed: false), so a load balancer outside Kubernetes can target a fixed port. Null lets Kubernetes allocate one from the cluster's node-port range; a set value must lie in that range (30000-32767 by default), which the API server enforces at install. `hive.externalLoadBalancer.nodeSelector` # { #helm.hive.externalLoadBalancer.nodeSelector } : Type `object`, default `{}`. selector for nodes the load balancer can use to route traffic
## metastore { #helm-values-hive-metastore } ??? example "Defaults as YAML" ```yaml hive: metastore: appName: hivemetastore config: cm_rootdir: '' enforce_auth: true hopsfs_dir: /apps/hive/warehouse hudi_hadoop_version: 1.2.0 impersonation: hive,glassfish mapreduce_input_size: '134217728' repl_rootdir: '' scratch_dir: /tmp/hive configmap: name: hivemetastore-configmap create_secret: true dependencies: glassfish: consulServiceName: glassfish consulServiceTag: hopsworks port: 8182 mysql: consulServiceName: mysql port: 3306 namenode: consulServiceName: namenode consulServiceTag: rpc port: 8020 deployment: jvm_resources: xms: 2g xmx: 2g name: hivemetastore-deployment probes: liveness: initialDelaySeconds: 60 periodSeconds: 10 tcpSocket: port: 9083 timeoutSeconds: 10 readiness: initialDelaySeconds: 60 periodSeconds: 10 tcpSocket: port: 9083 timeoutSeconds: 60 startup: {} rbac: annotations: {} create: true extraRoleRules: [] name: hivemetastore-role useExistingRole: false replicas: 1 resources: limits: memory: 8192Mi requests: cpu: '2' memory: 6553Mi security: fsGroup: 1234 group: hive runAsGroup: 1234 runAsUser: 1516 user: hive serviceAccount: annotations: {} create: true name: hivemetastore-default hadoop_user: hive image: pullPolicy: IfNotPresent registry: docker.hops.works logLevel: INFO migration: resources: limits: cpu: 500m memory: 500Mi requests: cpu: 100m memory: 100Mi migrationBackOffLimit: 10 monitoring_port: 18002 nodeSelector: {} podDisruptionBudget: enabled: true minAvailable: 1 port: 9083 protocol: thrift service: annotations: consul.hashicorp.com/service-name: hive consul.hashicorp.com/service-tags: metastore,hiveserver2-tls,hiveserver2-plain name: metastore tolerations: [] topologySpreadConstraint: {} ```
`hive.metastore.appName` # { #helm.hive.metastore.appName } : Type `string`, default `"hivemetastore"`. `hive.metastore.config.cm_rootdir` # { #helm.hive.metastore.config.cm_rootdir } : Type `string`, default `""`. `hive.metastore.config.enforce_auth` # { #helm.hive.metastore.config.enforce_auth } : Type `bool`, default `true`. `hive.metastore.config.hopsfs_dir` # { #helm.hive.metastore.config.hopsfs_dir } : Type `string`, default `"/apps/hive/warehouse"`. `hive.metastore.config.hudi_hadoop_version` # { #helm.hive.metastore.config.hudi_hadoop_version } : Type `string`, default `"1.2.0"`. `hive.metastore.config.impersonation` # { #helm.hive.metastore.config.impersonation } : Type `string`, default `"hive,glassfish"`. `hive.metastore.config.mapreduce_input_size` # { #helm.hive.metastore.config.mapreduce_input_size } : Type `string`, default `"134217728"`. `hive.metastore.config.repl_rootdir` # { #helm.hive.metastore.config.repl_rootdir } : Type `string`, default `""`. `hive.metastore.config.scratch_dir` # { #helm.hive.metastore.config.scratch_dir } : Type `string`, default `"/tmp/hive"`. `hive.metastore.configmap.name` # { #helm.hive.metastore.configmap.name } : Type `string`, default `"hivemetastore-configmap"`. `hive.metastore.create_secret` # { #helm.hive.metastore.create_secret } : Type `bool`, default `true`. Create the mysql users secret for hive metastore. If false, you have to create the secret (hive-user-secrets) manually. `hive.metastore.dependencies.glassfish.consulServiceName` # { #helm.hive.metastore.dependencies.glassfish.consulServiceName } : Type `string`, default `"glassfish"`. `hive.metastore.dependencies.glassfish.consulServiceTag` # { #helm.hive.metastore.dependencies.glassfish.consulServiceTag } : Type `string`, default `"hopsworks"`. `hive.metastore.dependencies.glassfish.port` # { #helm.hive.metastore.dependencies.glassfish.port } : Type `int`, default `8182`. `hive.metastore.dependencies.mysql.consulServiceName` # { #helm.hive.metastore.dependencies.mysql.consulServiceName } : Type `string`, default `"mysql"`. `hive.metastore.dependencies.mysql.port` # { #helm.hive.metastore.dependencies.mysql.port } : Type `int`, default `3306`. `hive.metastore.dependencies.namenode.consulServiceName` # { #helm.hive.metastore.dependencies.namenode.consulServiceName } : Type `string`, default `"namenode"`. `hive.metastore.dependencies.namenode.consulServiceTag` # { #helm.hive.metastore.dependencies.namenode.consulServiceTag } : Type `string`, default `"rpc"`. `hive.metastore.dependencies.namenode.port` # { #helm.hive.metastore.dependencies.namenode.port } : Type `int`, default `8020`. `hive.metastore.deployment.jvm_resources` # { #helm.hive.metastore.deployment.jvm_resources } : Type `object`, default `{"xms":"2g","xmx":"2g"}`. Xms and Xmx parameters for the Hivemetastore JVM initialization. Notice that, if set, container resources will be ignored and automatically calculated based on the values provided by this property. `hive.metastore.deployment.jvm_resources.xms` # { #helm.hive.metastore.deployment.jvm_resources.xms } : Type `string`, default `"2g"`. Initial heap size for JVM (e.g., '0.5g', '1g', '1.5g') `hive.metastore.deployment.jvm_resources.xmx` # { #helm.hive.metastore.deployment.jvm_resources.xmx } : Type `string`, default `"2g"`. Maximum heap size for JVM (e.g., '0.5g', '1g', '1.5g') `hive.metastore.deployment.name` # { #helm.hive.metastore.deployment.name } : Type `string`, default `"hivemetastore-deployment"`. `hive.metastore.deployment.probes.liveness.initialDelaySeconds` # { #helm.hive.metastore.deployment.probes.liveness.initialDelaySeconds } : Type `int`, default `60`. `hive.metastore.deployment.probes.liveness.periodSeconds` # { #helm.hive.metastore.deployment.probes.liveness.periodSeconds } : Type `int`, default `10`. `hive.metastore.deployment.probes.liveness.tcpSocket.port` # { #helm.hive.metastore.deployment.probes.liveness.tcpSocket.port } : Type `int`, default `9083`. `hive.metastore.deployment.probes.liveness.timeoutSeconds` # { #helm.hive.metastore.deployment.probes.liveness.timeoutSeconds } : Type `int`, default `10`. `hive.metastore.deployment.probes.readiness.initialDelaySeconds` # { #helm.hive.metastore.deployment.probes.readiness.initialDelaySeconds } : Type `int`, default `60`. `hive.metastore.deployment.probes.readiness.periodSeconds` # { #helm.hive.metastore.deployment.probes.readiness.periodSeconds } : Type `int`, default `10`. `hive.metastore.deployment.probes.readiness.tcpSocket.port` # { #helm.hive.metastore.deployment.probes.readiness.tcpSocket.port } : Type `int`, default `9083`. `hive.metastore.deployment.probes.readiness.timeoutSeconds` # { #helm.hive.metastore.deployment.probes.readiness.timeoutSeconds } : Type `int`, default `60`. `hive.metastore.deployment.probes.startup` # { #helm.hive.metastore.deployment.probes.startup } : Type `object`, default `{}`. startup probes `hive.metastore.deployment.rbac.annotations` # { #helm.hive.metastore.deployment.rbac.annotations } : Type `object`, default `{}`. metrics rbac role annotations `hive.metastore.deployment.rbac.create` # { #helm.hive.metastore.deployment.rbac.create } : Type `bool`, default `true`. `hive.metastore.deployment.rbac.extraRoleRules` # { #helm.hive.metastore.deployment.rbac.extraRoleRules } : Type `list`, default `[]`. extra rules to attach to metrics acl role `hive.metastore.deployment.rbac.name` # { #helm.hive.metastore.deployment.rbac.name } : Type `string`, default `"hivemetastore-role"`. `hive.metastore.deployment.rbac.useExistingRole` # { #helm.hive.metastore.deployment.rbac.useExistingRole } : Type `bool`, default `false`. `hive.metastore.deployment.replicas` # { #helm.hive.metastore.deployment.replicas } : Type `int`, default `1`. `hive.metastore.deployment.resources.limits` # { #helm.hive.metastore.deployment.resources.limits } : Type `object`, default `{"memory":"8192Mi"}`. resources limits configuration `hive.metastore.deployment.resources.requests` # { #helm.hive.metastore.deployment.resources.requests } : Type `object`, default `{"cpu":"2","memory":"6553Mi"}`. resources requests configuration `hive.metastore.deployment.security.fsGroup` # { #helm.hive.metastore.deployment.security.fsGroup } : Type `int`, default `1234`. `hive.metastore.deployment.security.group` # { #helm.hive.metastore.deployment.security.group } : Type `string`, default `"hive"`. `hive.metastore.deployment.security.runAsGroup` # { #helm.hive.metastore.deployment.security.runAsGroup } : Type `int`, default `1234`. `hive.metastore.deployment.security.runAsUser` # { #helm.hive.metastore.deployment.security.runAsUser } : Type `int`, default `1516`. `hive.metastore.deployment.security.user` # { #helm.hive.metastore.deployment.security.user } : Type `string`, default `"hive"`. `hive.metastore.deployment.serviceAccount.annotations` # { #helm.hive.metastore.deployment.serviceAccount.annotations } : Type `object`, default `{}`. service account annotations `hive.metastore.deployment.serviceAccount.create` # { #helm.hive.metastore.deployment.serviceAccount.create } : Type `bool`, default `true`. `hive.metastore.deployment.serviceAccount.name` # { #helm.hive.metastore.deployment.serviceAccount.name } : Type `string`, default `"hivemetastore-default"`. `hive.metastore.hadoop_user` # { #helm.hive.metastore.hadoop_user } : Type `string`, default `"hive"`. `hive.metastore.image.pullPolicy` # { #helm.hive.metastore.image.pullPolicy } : Type `string`, default `"IfNotPresent"`. `hive.metastore.image.registry` # { #helm.hive.metastore.image.registry } : Type `string`, default `"docker.hops.works"`. `hive.metastore.logLevel` # { #helm.hive.metastore.logLevel } : Type `string`, default `"INFO"`. `hive.metastore.migration.resources.limits.cpu` # { #helm.hive.metastore.migration.resources.limits.cpu } : Type `string`, default `"500m"`. `hive.metastore.migration.resources.limits.memory` # { #helm.hive.metastore.migration.resources.limits.memory } : Type `string`, default `"500Mi"`. `hive.metastore.migration.resources.requests.cpu` # { #helm.hive.metastore.migration.resources.requests.cpu } : Type `string`, default `"100m"`. `hive.metastore.migration.resources.requests.memory` # { #helm.hive.metastore.migration.resources.requests.memory } : Type `string`, default `"100Mi"`. `hive.metastore.migrationBackOffLimit` # { #helm.hive.metastore.migrationBackOffLimit } : Type `int`, default `10`. backoffLimit for hive migration job `hive.metastore.monitoring_port` # { #helm.hive.metastore.monitoring_port } : Type `int`, default `18002`. `hive.metastore.nodeSelector` # { #helm.hive.metastore.nodeSelector } : Type `object`, default `{}`. node selector configuration `hive.metastore.podDisruptionBudget.enabled` # { #helm.hive.metastore.podDisruptionBudget.enabled } : Type `bool`, default `true`. `hive.metastore.podDisruptionBudget.minAvailable` # { #helm.hive.metastore.podDisruptionBudget.minAvailable } : Type `int`, default `1`. `hive.metastore.port` # { #helm.hive.metastore.port } : Type `int`, default `9083`. `hive.metastore.protocol` # { #helm.hive.metastore.protocol } : Type `string`, default `"thrift"`. `hive.metastore.service.annotations."consul.hashicorp.com/service-name"` # { #helm.hive.metastore.service.annotations.consul.hashicorp.com-service-name } : Type `string`, default `"hive"`. `hive.metastore.service.annotations."consul.hashicorp.com/service-tags"` # { #helm.hive.metastore.service.annotations.consul.hashicorp.com-service-tags } : Type `string`, default `"metastore,hiveserver2-tls,hiveserver2-plain"`. `hive.metastore.service.name` # { #helm.hive.metastore.service.name } : Type `string`, default `"metastore"`. `hive.metastore.tolerations` # { #helm.hive.metastore.tolerations } : Type `list`, default `[]`. `hive.metastore.topologySpreadConstraint` # { #helm.hive.metastore.topologySpreadConstraint } : Type `object`, default `{}`. The default topology spread constraint. If not defined the global topology spread constraint would be used instead.
================================================================================ # hopsfs Source: https://docs.hopsworks.ai/latest/setup_installation/common/helm_chart_values/hopsfs/ # HopsFS values { #helm-values-hopsfs } Values under `hopsfs` configure HopsFS, the distributed file system behind datasets and the offline feature store. _Generated from the Hopsworks Helm chart `5.2.0-alpha-1791549041` (Hopsworks `5.2.0`)._ Deployed when [`global._hopsworks.full_platform`](global.md#helm.global._hopsworks.full_platform) is `true`. ## General { #helm-values-hopsfs-general } ??? example "Defaults as YAML" ```yaml hopsfs: bypassBucketValidation: false cleanupOnUninstall: enabled: true ttlSecondsAfterFinished: null client: failureReplacementPolicy: NEVER locateFollowingBlockRetries: 10 dataDir: /srv/hops/hopsdata/hdfs datanode: count: 5 storage: size: 100Gi storageClassName: null hopsworkslib: {} image: name: hopsfs pullPolicy: IfNotPresent registry: '' tag: 3.4.3.3-EE-RC1 namenode: resources: limits: memory: 2048Mi requests: memory: 1024Mi nuke_db_in_retry: false objectStorage: enabled: true parallel_preset: 3 presetJob: clientRetryCount: 10 clientRetryIntervalSeconds: 10 runIndex: 0 ttlSecondsAfterFinished: null serviceAccount: annotations: {} setupJob: timeoutMinutes: 25 ```
`hopsfs` # { #helm.hopsfs } : Type `object`. override hopsfs values ??? note "Default" ```yaml datanode: count: 5 storage: size: 100Gi storageClassName: null namenode: resources: limits: memory: 2048Mi requests: memory: 1024Mi objectStorage: enabled: true ``` `hopsfs.bypassBucketValidation` # { #helm.hopsfs.bypassBucketValidation } : Type `bool`, default `false`. Do not run HopsFS bucket validation pre-install, pre-upgrade Job `hopsfs.cleanupOnUninstall` # { #helm.hopsfs.cleanupOnUninstall } : Type `object`, default `{"enabled":true,"ttlSecondsAfterFinished":null}`. post-delete cleanup of HopsFS leftovers. The datanode StatefulSet PVC survives uninstall; this deletes it by its labels (app=hopsfs, service=datanode), but only when global._hopsworks.wipeDataOnUninstall is enabled and never for PVCs labelled hopsworks.ai/keep=true. `hopsfs.cleanupOnUninstall.enabled` # { #helm.hopsfs.cleanupOnUninstall.enabled } : Type `bool`, default `true`. enable the post-delete HopsFS data-PVC cleanup hook (also requires global._hopsworks.wipeDataOnUninstall) `hopsfs.cleanupOnUninstall.ttlSecondsAfterFinished` # { #helm.hopsfs.cleanupOnUninstall.ttlSecondsAfterFinished } : Type `string`, default `nil`. ttlSecondsAfterFinished for the cleanup Job; null falls through to the global default `hopsfs.client.failureReplacementPolicy` # { #helm.hopsfs.client.failureReplacementPolicy } : Type `string`, default `"NEVER"`. `hopsfs.client.locateFollowingBlockRetries` # { #helm.hopsfs.client.locateFollowingBlockRetries } : Type `int`, default `10`. `hopsfs.dataDir` # { #helm.hopsfs.dataDir } : Type `string`, default `"/srv/hops/hopsdata/hdfs"`. `hopsfs.hopsworkslib` # { #helm.hopsfs.hopsworkslib } : Type `object`, default `{}`. override hopsworkslib values `hopsfs.image.name` # { #helm.hopsfs.image.name } : Type `string`, default `"hopsfs"`. `hopsfs.image.pullPolicy` # { #helm.hopsfs.image.pullPolicy } : Type `string`, default `"IfNotPresent"`. `hopsfs.image.registry` # { #helm.hopsfs.image.registry } : Type `string`, default `""`. Override the full registry+path prefix (must end with /). When set (non-empty), takes precedence over global._hopsworks.imageRegistry, bypassing the hardcoded /hopsworks/ segment. Use for custom HopsFS images at non-standard paths, e.g. "docker.hops.works/dev/salman/". Leave empty to use global._hopsworks.imageRegistry + "/hopsworks/". `hopsfs.image.tag` # { #helm.hopsfs.image.tag } : Type `string`, default `"3.4.3.3-EE-RC1"`. `hopsfs.nuke_db_in_retry` # { #helm.hopsfs.nuke_db_in_retry } : Type `bool`, default `false`. `hopsfs.parallel_preset` # { #helm.hopsfs.parallel_preset } : Type `int`, default `3`. `hopsfs.presetJob.clientRetryCount` # { #helm.hopsfs.presetJob.clientRetryCount } : Type `int`, default `10`. `hopsfs.presetJob.clientRetryIntervalSeconds` # { #helm.hopsfs.presetJob.clientRetryIntervalSeconds } : Type `int`, default `10`. `hopsfs.presetJob.runIndex` # { #helm.hopsfs.presetJob.runIndex } : Type `int`, default `0`. Index to make job names unique if necessary `hopsfs.presetJob.ttlSecondsAfterFinished` # { #helm.hopsfs.presetJob.ttlSecondsAfterFinished } : Type `string`, default `nil`. TTL in seconds for the namenode-preset-folders Job. Overrides global default. `hopsfs.serviceAccount.annotations` # { #helm.hopsfs.serviceAccount.annotations } : Type `object`, default `{}`. service account annotations `hopsfs.setupJob.timeoutMinutes` # { #helm.hopsfs.setupJob.timeoutMinutes } : Type `int`, default `25`.
## datanode { #helm-values-hopsfs-datanode } ??? example "Defaults as YAML" ```yaml hopsfs: datanode: count: 5 dataPort: 50010 gracefulDrain: enabled: true terminationGracePeriodSecs: null timeoutSecs: null uploadThroughputMiBps: 35 httpPort: 50075 httpsPort: 50475 ipcPort: 50020 jvmOpts: '' loadBalancer: annotations: {} enabled: null loadBalancerClass: null managed: null nodePort: null monitoringPort: 50076 nodeSelector: {} podDisruptionBudget: enabled: true minAvailable: 1 resources: limits: cpu: '4' memory: 2048Mi requests: cpu: '1' memory: 500Mi rpcHandlerCount: 20 storage: size: 100Gi storageClassName: null tls: extraDnsNames: [] extraIpAddresses: [] tolerations: [] topologySpreadConstraint: {} xmx: 1024 ```
`hopsfs.datanode.count` # { #helm.hopsfs.datanode.count } : Type `int`, default `5`. `hopsfs.datanode.dataPort` # { #helm.hopsfs.datanode.dataPort } : Type `int`, default `50010`. `hopsfs.datanode.gracefulDrain.enabled` # { #helm.hopsfs.datanode.gracefulDrain.enabled } : Type `bool`, default `true`. Enable the preStop drain hook when async cloud upload is on. Set false to opt out even when async is enabled; risks data loss under non-empty upload backlog. `hopsfs.datanode.gracefulDrain.terminationGracePeriodSecs` # { #helm.hopsfs.datanode.gracefulDrain.terminationGracePeriodSecs } : Type `string`, default `nil`. Explicit override for the pod terminationGracePeriodSeconds. When null, defaults to the calculated drain timeout. `hopsfs.datanode.gracefulDrain.timeoutSecs` # { #helm.hopsfs.datanode.gracefulDrain.timeoutSecs } : Type `string`, default `nil`. Explicit override for the dfsadmin -timeout passed to the preStop drain RPC. When null, calculated from queueCapacity, blockSizeMiB, and uploadThroughputMiBps. `hopsfs.datanode.gracefulDrain.uploadThroughputMiBps` # { #helm.hopsfs.datanode.gracefulDrain.uploadThroughputMiBps } : Type `int`, default `35`. Operator estimate of per-DN S3 upload throughput in MiB/s. Drives the calculated drain timeout when timeoutSecs is null. `hopsfs.datanode.httpPort` # { #helm.hopsfs.datanode.httpPort } : Type `int`, default `50075`. `hopsfs.datanode.httpsPort` # { #helm.hopsfs.datanode.httpsPort } : Type `int`, default `50475`. `hopsfs.datanode.ipcPort` # { #helm.hopsfs.datanode.ipcPort } : Type `int`, default `50020`. `hopsfs.datanode.jvmOpts` # { #helm.hopsfs.datanode.jvmOpts } : Type `string`, default `""`. `hopsfs.datanode.loadBalancer.annotations` # { #helm.hopsfs.datanode.loadBalancer.annotations } : Type `object`, default `{}`. annotations for load balancer `hopsfs.datanode.loadBalancer.enabled` # { #helm.hopsfs.datanode.loadBalancer.enabled } : Type `string`, default `nil`. Enable External Load Balancers for datanode. If not set the .global._hopsworks.externalLoadBalancers.enabled will be used instead `hopsfs.datanode.loadBalancer.loadBalancerClass` # { #helm.hopsfs.datanode.loadBalancer.loadBalancerClass } : Type `string`, default `nil`. load balancer class name `hopsfs.datanode.loadBalancer.managed` # { #helm.hopsfs.datanode.loadBalancer.managed } : Type `string`, default `nil`. Cloud provider provisions Load Balancers. If not set the .global._hopsworks.externalLoadBalancers.managed will be used instead `hopsfs.datanode.loadBalancer.nodePort` # { #helm.hopsfs.datanode.loadBalancer.nodePort } : Type `string`, default `nil`. Explicit nodePort for the external service when the load balancer is unmanaged (managed: false), so a load balancer outside Kubernetes can target a fixed port. Null lets Kubernetes allocate one from the cluster's node-port range; a set value must lie in that range (30000-32767 by default), which the API server enforces at install. `hopsfs.datanode.monitoringPort` # { #helm.hopsfs.datanode.monitoringPort } : Type `int`, default `50076`. `hopsfs.datanode.nodeSelector` # { #helm.hopsfs.datanode.nodeSelector } : Type `object`, default `{}`. node selector configuration `hopsfs.datanode.podDisruptionBudget.enabled` # { #helm.hopsfs.datanode.podDisruptionBudget.enabled } : Type `bool`, default `true`. Enable the DataNode PodDisruptionBudget. When enabled, the PDB hardcodes maxUnavailable=1 to prevent concurrent DN unavailability from exhausting HDFS client pipeline-recovery retries and failing in-flight writes. Not operator-tunable; this applies independently of async cloud upload. `hopsfs.datanode.resources.limits.cpu` # { #helm.hopsfs.datanode.resources.limits.cpu } : Type `string`, default `"4"`. `hopsfs.datanode.resources.limits.memory` # { #helm.hopsfs.datanode.resources.limits.memory } : Type `string`, default `"2048Mi"`. `hopsfs.datanode.resources.requests.cpu` # { #helm.hopsfs.datanode.resources.requests.cpu } : Type `string`, default `"1"`. `hopsfs.datanode.resources.requests.memory` # { #helm.hopsfs.datanode.resources.requests.memory } : Type `string`, default `"500Mi"`. `hopsfs.datanode.rpcHandlerCount` # { #helm.hopsfs.datanode.rpcHandlerCount } : Type `int`, default `20`. `hopsfs.datanode.storage.size` # { #helm.hopsfs.datanode.storage.size } : Type `string`, default `"100Gi"`. `hopsfs.datanode.storage.storageClassName` # { #helm.hopsfs.datanode.storage.storageClassName } : Type `string`, default `nil`. storage class name to request for volumes attached to the data nodes `hopsfs.datanode.tls.extraDnsNames` # { #helm.hopsfs.datanode.tls.extraDnsNames } : Type `list`, default `[]`. Additional DNS names to add as SAN to Datanode x.509 certificate `hopsfs.datanode.tls.extraIpAddresses` # { #helm.hopsfs.datanode.tls.extraIpAddresses } : Type `list`, default `[]`. Additional IP addresses to add as SAN to Datanode x.509 certificate `hopsfs.datanode.tolerations` # { #helm.hopsfs.datanode.tolerations } : Type `list`, default `[]`. `hopsfs.datanode.topologySpreadConstraint` # { #helm.hopsfs.datanode.topologySpreadConstraint } : Type `object`, default `{}`. The default topology spread constraint. If not defined the global topology spread constraint would be used instead. `hopsfs.datanode.xmx` # { #helm.hopsfs.datanode.xmx } : Type `int`, default `1024`. `hopsfs.datanode.podDisruptionBudget.minAvailable` Deprecated # { #helm.hopsfs.datanode.podDisruptionBudget.minAvailable } : Type `int`, default `1`. DEPRECATED. The DN PDB no longer uses minAvailable; the template hardcodes maxUnavailable=1. Retained only so existing values overrides do not fail schema validation on upgrade.
## dependencies { #helm-values-hopsfs-dependencies } ??? example "Defaults as YAML" ```yaml hopsfs: dependencies: glassfish: consulServiceName: glassfish consulServiceTag: ca port: 8182 mgmd: consulServiceName: mgmd port: 1186 mysql: consulServiceName: mysql port: 3306 objectStorage: consulServiceName: minio port: 9000 ```
`hopsfs.dependencies.glassfish.consulServiceName` # { #helm.hopsfs.dependencies.glassfish.consulServiceName } : Type `string`, default `"glassfish"`. `hopsfs.dependencies.glassfish.consulServiceTag` # { #helm.hopsfs.dependencies.glassfish.consulServiceTag } : Type `string`, default `"ca"`. `hopsfs.dependencies.glassfish.port` # { #helm.hopsfs.dependencies.glassfish.port } : Type `int`, default `8182`. `hopsfs.dependencies.mgmd.consulServiceName` # { #helm.hopsfs.dependencies.mgmd.consulServiceName } : Type `string`, default `"mgmd"`. `hopsfs.dependencies.mgmd.port` # { #helm.hopsfs.dependencies.mgmd.port } : Type `int`, default `1186`. `hopsfs.dependencies.mysql.consulServiceName` # { #helm.hopsfs.dependencies.mysql.consulServiceName } : Type `string`, default `"mysql"`. `hopsfs.dependencies.mysql.port` # { #helm.hopsfs.dependencies.mysql.port } : Type `int`, default `3306`. `hopsfs.dependencies.objectStorage.consulServiceName` # { #helm.hopsfs.dependencies.objectStorage.consulServiceName } : Type `string`, default `"minio"`. `hopsfs.dependencies.objectStorage.port` # { #helm.hopsfs.dependencies.objectStorage.port } : Type `int`, default `9000`.
## log4j { #helm-values-hopsfs-log4j } ??? example "Defaults as YAML" ```yaml hopsfs: log4j: console_logger_level: INFO enable_audit_log: true level: INFO rfa_logger_level: INFO rfa_max_backup_index: 10 rfa_max_file_size: 256MB root_level_logger: INFO ```
`hopsfs.log4j.console_logger_level` # { #helm.hopsfs.log4j.console_logger_level } : Type `string`, default `"INFO"`. Log level threshold for Console appender. Defaults to root_level_logger if not specified. `hopsfs.log4j.enable_audit_log` # { #helm.hopsfs.log4j.enable_audit_log } : Type `bool`, default `true`. `hopsfs.log4j.level` # { #helm.hopsfs.log4j.level } : Type `string`, default `"INFO"`. Backward compatibility field. Not used by log4j configuration. `hopsfs.log4j.rfa_logger_level` # { #helm.hopsfs.log4j.rfa_logger_level } : Type `string`, default `"INFO"`. Log level threshold for Rolling File Appender. Defaults to root_level_logger if not specified. `hopsfs.log4j.rfa_max_backup_index` # { #helm.hopsfs.log4j.rfa_max_backup_index } : Type `int`, default `10`. Maximum number of backup log files to keep `hopsfs.log4j.rfa_max_file_size` # { #helm.hopsfs.log4j.rfa_max_file_size } : Type `string`, default `"256MB"`. Maximum size of each log file before rotation (e.g., 256MB, 1GB) `hopsfs.log4j.root_level_logger` # { #helm.hopsfs.log4j.root_level_logger } : Type `string`, default `"INFO"`. Root logger level - acts as a global minimum level for all appenders
## namenode { #helm-values-hopsfs-namenode } ??? example "Defaults as YAML" ```yaml hopsfs: namenode: acls: enabled: true auditLog: enableListOperation: false enableStatOperation: false blockSizeMiB: 128 blockreportExecutorCount: 40 clusterIp: name: namenode-cluster-ip count: 1 customConfig: {} dirs: - group: hdfs mode: 1775 owner: payara path: /Projects - group: hadoop mode: 1775 owner: hdfs path: /user - group: hadoop mode: 1775 owner: hdfs path: /apps - group: hadoop mode: 1775 owner: hdfs path: /tmp - group: hadoop mode: 1777 owner: hive path: /tmp/hive httpPort: 50070 httpsPort: 50470 jvmOpts: '' loadBalancer: annotations: {} enabled: null loadBalancerClass: null managed: null nodePort: null maxBlocksPerFile: 10240 maxDirectMemorySize: 1024 maxDirectoryItems: 131072 monitoringPort: 50071 nodeSelector: {} numBlockReplicas: 1 podDisruptionBudget: enabled: true minAvailable: 1 quotaEnabled: true resources: limits: cpu: '4' memory: 2048Mi requests: cpu: '1' memory: 1024Mi rpcHandlerCount: 120 rpcPort: 8020 service: annotations: consul.hashicorp.com/service-name: namenode consul.hashicorp.com/service-port: rpc consul.hashicorp.com/service-tags: rpc,http prometheus.io/path: /metrics prometheus.io/port: 50071 prometheus.io/scheme: http prometheus.io/scrape: 'true' name: namenode setupBackOffLimit: 10 startupProbe: failureThreshold: 50 initialDelaySeconds: 10 periodSeconds: 15 timeoutSeconds: 10 stoCleanFaildOpsDelay: 600000 stoCleanSlowOpsDelay: 900000 subtreeExecutorCount: 40 tls: extraDnsNames: [] extraIpAddresses: [] tolerations: [] topologySpreadConstraint: {} txRetryCount: 10 users: - gid: 1508 group: airflow mode: 1750 name: airflow uid: 1512 - extra_folders: - /user/spark/applicationHistory - /user/spark/eventlog - /user/spark/share - /user/spark/spark-warehouse group: hadoop mode: 1777 name: spark uid: 1505 - extra_folders: - /apps/hive - /apps/hive/warehouse group: hadoop mode: 1755 name: hive - group: hdfs mode: 1750 name: payara - extra_folders: - /user/ray/applications group: hadoop mode: 1750 name: ray - group: hdfs mode: 1750 name: trino xattrs: enabled: true maxXAttrSize: 1039755 maxXAttrsPerInode: 32 xmx: 1024 ```
`hopsfs.namenode.acls.enabled` # { #helm.hopsfs.namenode.acls.enabled } : Type `bool`, default `true`. `hopsfs.namenode.auditLog.enableListOperation` # { #helm.hopsfs.namenode.auditLog.enableListOperation } : Type `bool`, default `false`. `hopsfs.namenode.auditLog.enableStatOperation` # { #helm.hopsfs.namenode.auditLog.enableStatOperation } : Type `bool`, default `false`. `hopsfs.namenode.blockSizeMiB` # { #helm.hopsfs.namenode.blockSizeMiB } : Type `int`, default `128`. HDFS block size in MiB. Templated into dfs.blocksize. Also drives the calculated DN graceful-drain timeout. `hopsfs.namenode.blockreportExecutorCount` # { #helm.hopsfs.namenode.blockreportExecutorCount } : Type `int`, default `40`. `hopsfs.namenode.clusterIp.name` # { #helm.hopsfs.namenode.clusterIp.name } : Type `string`, default `"namenode-cluster-ip"`. `hopsfs.namenode.count` # { #helm.hopsfs.namenode.count } : Type `int`, default `1`. `hopsfs.namenode.customConfig` # { #helm.hopsfs.namenode.customConfig } : Type `object`, default `{}`. Set/Override hdfs-site.xml properties. Each entry in this map will template a entry at the bottom of the hdfs-site.xml. The name of the property is the map entry key and the value of the property is the value of the map entry key. `hopsfs.namenode.dirs[0].group` # { #helm.hopsfs.namenode.dirs.0.group } : Type `string`, default `"hdfs"`. `hopsfs.namenode.dirs[0].mode` # { #helm.hopsfs.namenode.dirs.0.mode } : Type `int`, default `1775`. `hopsfs.namenode.dirs[0].owner` # { #helm.hopsfs.namenode.dirs.0.owner } : Type `string`, default `"payara"`. `hopsfs.namenode.dirs[0].path` # { #helm.hopsfs.namenode.dirs.0.path } : Type `string`, default `"/Projects"`. `hopsfs.namenode.dirs[1].group` # { #helm.hopsfs.namenode.dirs.1.group } : Type `string`, default `"hadoop"`. `hopsfs.namenode.dirs[1].mode` # { #helm.hopsfs.namenode.dirs.1.mode } : Type `int`, default `1775`. `hopsfs.namenode.dirs[1].owner` # { #helm.hopsfs.namenode.dirs.1.owner } : Type `string`, default `"hdfs"`. `hopsfs.namenode.dirs[1].path` # { #helm.hopsfs.namenode.dirs.1.path } : Type `string`, default `"/user"`. `hopsfs.namenode.dirs[2].group` # { #helm.hopsfs.namenode.dirs.2.group } : Type `string`, default `"hadoop"`. `hopsfs.namenode.dirs[2].mode` # { #helm.hopsfs.namenode.dirs.2.mode } : Type `int`, default `1775`. `hopsfs.namenode.dirs[2].owner` # { #helm.hopsfs.namenode.dirs.2.owner } : Type `string`, default `"hdfs"`. `hopsfs.namenode.dirs[2].path` # { #helm.hopsfs.namenode.dirs.2.path } : Type `string`, default `"/apps"`. `hopsfs.namenode.dirs[3].group` # { #helm.hopsfs.namenode.dirs.3.group } : Type `string`, default `"hadoop"`. `hopsfs.namenode.dirs[3].mode` # { #helm.hopsfs.namenode.dirs.3.mode } : Type `int`, default `1775`. `hopsfs.namenode.dirs[3].owner` # { #helm.hopsfs.namenode.dirs.3.owner } : Type `string`, default `"hdfs"`. `hopsfs.namenode.dirs[3].path` # { #helm.hopsfs.namenode.dirs.3.path } : Type `string`, default `"/tmp"`. `hopsfs.namenode.dirs[4].group` # { #helm.hopsfs.namenode.dirs.4.group } : Type `string`, default `"hadoop"`. `hopsfs.namenode.dirs[4].mode` # { #helm.hopsfs.namenode.dirs.4.mode } : Type `int`, default `1777`. `hopsfs.namenode.dirs[4].owner` # { #helm.hopsfs.namenode.dirs.4.owner } : Type `string`, default `"hive"`. `hopsfs.namenode.dirs[4].path` # { #helm.hopsfs.namenode.dirs.4.path } : Type `string`, default `"/tmp/hive"`. `hopsfs.namenode.httpPort` # { #helm.hopsfs.namenode.httpPort } : Type `int`, default `50070`. `hopsfs.namenode.httpsPort` # { #helm.hopsfs.namenode.httpsPort } : Type `int`, default `50470`. `hopsfs.namenode.jvmOpts` # { #helm.hopsfs.namenode.jvmOpts } : Type `string`, default `""`. `hopsfs.namenode.loadBalancer.annotations` # { #helm.hopsfs.namenode.loadBalancer.annotations } : Type `object`, default `{}`. annotations for load balancer `hopsfs.namenode.loadBalancer.enabled` # { #helm.hopsfs.namenode.loadBalancer.enabled } : Type `string`, default `nil`. Enable External Load Balancers for namenode. If not set the .global._hopsworks.externalLoadBalancers.enabled will be used instead `hopsfs.namenode.loadBalancer.loadBalancerClass` # { #helm.hopsfs.namenode.loadBalancer.loadBalancerClass } : Type `string`, default `nil`. load balancer class name `hopsfs.namenode.loadBalancer.managed` # { #helm.hopsfs.namenode.loadBalancer.managed } : Type `string`, default `nil`. Cloud provider provisions Load Balancers. If not set the .global._hopsworks.externalLoadBalancers.managed will be used instead `hopsfs.namenode.loadBalancer.nodePort` # { #helm.hopsfs.namenode.loadBalancer.nodePort } : Type `string`, default `nil`. Explicit nodePort for the external service when the load balancer is unmanaged (managed: false), so a load balancer outside Kubernetes can target a fixed port. Null lets Kubernetes allocate one from the cluster's node-port range; a set value must lie in that range (30000-32767 by default), which the API server enforces at install. `hopsfs.namenode.maxBlocksPerFile` # { #helm.hopsfs.namenode.maxBlocksPerFile } : Type `int`, default `10240`. Maximum number of blocks per file `hopsfs.namenode.maxDirectMemorySize` # { #helm.hopsfs.namenode.maxDirectMemorySize } : Type `int`, default `1024`. `hopsfs.namenode.maxDirectoryItems` # { #helm.hopsfs.namenode.maxDirectoryItems } : Type `int`, default `131072`. Maximum number of items (files and directories) allowed in a directory `hopsfs.namenode.monitoringPort` # { #helm.hopsfs.namenode.monitoringPort } : Type `int`, default `50071`. `hopsfs.namenode.nodeSelector` # { #helm.hopsfs.namenode.nodeSelector } : Type `object`, default `{}`. node selector configuration `hopsfs.namenode.numBlockReplicas` # { #helm.hopsfs.namenode.numBlockReplicas } : Type `int`, default `1`. `hopsfs.namenode.podDisruptionBudget.enabled` # { #helm.hopsfs.namenode.podDisruptionBudget.enabled } : Type `bool`, default `true`. `hopsfs.namenode.podDisruptionBudget.minAvailable` # { #helm.hopsfs.namenode.podDisruptionBudget.minAvailable } : Type `int`, default `1`. `hopsfs.namenode.quotaEnabled` # { #helm.hopsfs.namenode.quotaEnabled } : Type `bool`, default `true`. `hopsfs.namenode.resources.limits.cpu` # { #helm.hopsfs.namenode.resources.limits.cpu } : Type `string`, default `"4"`. `hopsfs.namenode.resources.limits.memory` # { #helm.hopsfs.namenode.resources.limits.memory } : Type `string`, default `"2048Mi"`. `hopsfs.namenode.resources.requests.cpu` # { #helm.hopsfs.namenode.resources.requests.cpu } : Type `string`, default `"1"`. `hopsfs.namenode.resources.requests.memory` # { #helm.hopsfs.namenode.resources.requests.memory } : Type `string`, default `"1024Mi"`. `hopsfs.namenode.rpcHandlerCount` # { #helm.hopsfs.namenode.rpcHandlerCount } : Type `int`, default `120`. `hopsfs.namenode.rpcPort` # { #helm.hopsfs.namenode.rpcPort } : Type `int`, default `8020`. `hopsfs.namenode.service.annotations."consul.hashicorp.com/service-name"` # { #helm.hopsfs.namenode.service.annotations.consul.hashicorp.com-service-name } : Type `string`, default `"namenode"`. `hopsfs.namenode.service.annotations."consul.hashicorp.com/service-port"` # { #helm.hopsfs.namenode.service.annotations.consul.hashicorp.com-service-port } : Type `string`, default `"rpc"`. `hopsfs.namenode.service.annotations."consul.hashicorp.com/service-tags"` # { #helm.hopsfs.namenode.service.annotations.consul.hashicorp.com-service-tags } : Type `string`, default `"rpc,http"`. `hopsfs.namenode.service.annotations."prometheus.io/path"` # { #helm.hopsfs.namenode.service.annotations.prometheus.io-path } : Type `string`, default `"/metrics"`. `hopsfs.namenode.service.annotations."prometheus.io/port"` # { #helm.hopsfs.namenode.service.annotations.prometheus.io-port } : Type `int`, default `50071`. `hopsfs.namenode.service.annotations."prometheus.io/scheme"` # { #helm.hopsfs.namenode.service.annotations.prometheus.io-scheme } : Type `string`, default `"http"`. `hopsfs.namenode.service.annotations."prometheus.io/scrape"` # { #helm.hopsfs.namenode.service.annotations.prometheus.io-scrape } : Type `string`, default `"true"`. `hopsfs.namenode.service.name` # { #helm.hopsfs.namenode.service.name } : Type `string`, default `"namenode"`. `hopsfs.namenode.setupBackOffLimit` # { #helm.hopsfs.namenode.setupBackOffLimit } : Type `int`, default `10`. Back off limit for database migration job. Default to be 10, i.e. 34 minutes, 30 seconds `hopsfs.namenode.startupProbe.failureThreshold` # { #helm.hopsfs.namenode.startupProbe.failureThreshold } : Type `int`, default `50`. `hopsfs.namenode.startupProbe.initialDelaySeconds` # { #helm.hopsfs.namenode.startupProbe.initialDelaySeconds } : Type `int`, default `10`. `hopsfs.namenode.startupProbe.periodSeconds` # { #helm.hopsfs.namenode.startupProbe.periodSeconds } : Type `int`, default `15`. `hopsfs.namenode.startupProbe.timeoutSeconds` # { #helm.hopsfs.namenode.startupProbe.timeoutSeconds } : Type `int`, default `10`. `hopsfs.namenode.stoCleanFaildOpsDelay` # { #helm.hopsfs.namenode.stoCleanFaildOpsDelay } : Type `int`, default `600000`. `hopsfs.namenode.stoCleanSlowOpsDelay` # { #helm.hopsfs.namenode.stoCleanSlowOpsDelay } : Type `int`, default `900000`. `hopsfs.namenode.subtreeExecutorCount` # { #helm.hopsfs.namenode.subtreeExecutorCount } : Type `int`, default `40`. `hopsfs.namenode.tls.extraDnsNames` # { #helm.hopsfs.namenode.tls.extraDnsNames } : Type `list`, default `[]`. Additional DNS names to add as SAN to Datanode x.509 certificate `hopsfs.namenode.tls.extraIpAddresses` # { #helm.hopsfs.namenode.tls.extraIpAddresses } : Type `list`, default `[]`. Additional IP addresses to add as SAN to Datanode x.509 certificate `hopsfs.namenode.tolerations` # { #helm.hopsfs.namenode.tolerations } : Type `list`, default `[]`. `hopsfs.namenode.topologySpreadConstraint` # { #helm.hopsfs.namenode.topologySpreadConstraint } : Type `object`, default `{}`. The default topology spread constraint. If not defined the global topology spread constraint would be used instead. `hopsfs.namenode.txRetryCount` # { #helm.hopsfs.namenode.txRetryCount } : Type `int`, default `10`. `hopsfs.namenode.users[0].gid` # { #helm.hopsfs.namenode.users.0.gid } : Type `int`, default `1508`. `hopsfs.namenode.users[0].group` # { #helm.hopsfs.namenode.users.0.group } : Type `string`, default `"airflow"`. `hopsfs.namenode.users[0].mode` # { #helm.hopsfs.namenode.users.0.mode } : Type `int`, default `1750`. `hopsfs.namenode.users[0].name` # { #helm.hopsfs.namenode.users.0.name } : Type `string`, default `"airflow"`. `hopsfs.namenode.users[0].uid` # { #helm.hopsfs.namenode.users.0.uid } : Type `int`, default `1512`. `hopsfs.namenode.users[1].extra_folders[0]` # { #helm.hopsfs.namenode.users.1.extra_folders.0 } : Type `string`, default `"/user/spark/applicationHistory"`. `hopsfs.namenode.users[1].extra_folders[1]` # { #helm.hopsfs.namenode.users.1.extra_folders.1 } : Type `string`, default `"/user/spark/eventlog"`. `hopsfs.namenode.users[1].extra_folders[2]` # { #helm.hopsfs.namenode.users.1.extra_folders.2 } : Type `string`, default `"/user/spark/share"`. `hopsfs.namenode.users[1].extra_folders[3]` # { #helm.hopsfs.namenode.users.1.extra_folders.3 } : Type `string`, default `"/user/spark/spark-warehouse"`. `hopsfs.namenode.users[1].group` # { #helm.hopsfs.namenode.users.1.group } : Type `string`, default `"hadoop"`. `hopsfs.namenode.users[1].mode` # { #helm.hopsfs.namenode.users.1.mode } : Type `int`, default `1777`. `hopsfs.namenode.users[1].name` # { #helm.hopsfs.namenode.users.1.name } : Type `string`, default `"spark"`. `hopsfs.namenode.users[1].uid` # { #helm.hopsfs.namenode.users.1.uid } : Type `int`, default `1505`. `hopsfs.namenode.users[2].extra_folders[0]` # { #helm.hopsfs.namenode.users.2.extra_folders.0 } : Type `string`, default `"/apps/hive"`. `hopsfs.namenode.users[2].extra_folders[1]` # { #helm.hopsfs.namenode.users.2.extra_folders.1 } : Type `string`, default `"/apps/hive/warehouse"`. `hopsfs.namenode.users[2].group` # { #helm.hopsfs.namenode.users.2.group } : Type `string`, default `"hadoop"`. `hopsfs.namenode.users[2].mode` # { #helm.hopsfs.namenode.users.2.mode } : Type `int`, default `1755`. `hopsfs.namenode.users[2].name` # { #helm.hopsfs.namenode.users.2.name } : Type `string`, default `"hive"`. `hopsfs.namenode.users[3].group` # { #helm.hopsfs.namenode.users.3.group } : Type `string`, default `"hdfs"`. `hopsfs.namenode.users[3].mode` # { #helm.hopsfs.namenode.users.3.mode } : Type `int`, default `1750`. `hopsfs.namenode.users[3].name` # { #helm.hopsfs.namenode.users.3.name } : Type `string`, default `"payara"`. `hopsfs.namenode.users[4].extra_folders[0]` # { #helm.hopsfs.namenode.users.4.extra_folders.0 } : Type `string`, default `"/user/ray/applications"`. `hopsfs.namenode.users[4].group` # { #helm.hopsfs.namenode.users.4.group } : Type `string`, default `"hadoop"`. `hopsfs.namenode.users[4].mode` # { #helm.hopsfs.namenode.users.4.mode } : Type `int`, default `1750`. `hopsfs.namenode.users[4].name` # { #helm.hopsfs.namenode.users.4.name } : Type `string`, default `"ray"`. `hopsfs.namenode.users[5].group` # { #helm.hopsfs.namenode.users.5.group } : Type `string`, default `"hdfs"`. `hopsfs.namenode.users[5].mode` # { #helm.hopsfs.namenode.users.5.mode } : Type `int`, default `1750`. `hopsfs.namenode.users[5].name` # { #helm.hopsfs.namenode.users.5.name } : Type `string`, default `"trino"`. `hopsfs.namenode.xattrs.enabled` # { #helm.hopsfs.namenode.xattrs.enabled } : Type `bool`, default `true`. `hopsfs.namenode.xattrs.maxXAttrSize` # { #helm.hopsfs.namenode.xattrs.maxXAttrSize } : Type `int`, default `1039755`. `hopsfs.namenode.xattrs.maxXAttrsPerInode` # { #helm.hopsfs.namenode.xattrs.maxXAttrsPerInode } : Type `int`, default `32`. `hopsfs.namenode.xmx` # { #helm.hopsfs.namenode.xmx } : Type `int`, default `1024`.
## objectStorage { #helm-values-hopsfs-objectstorage } ??? example "Defaults as YAML" ```yaml hopsfs: objectStorage: asyncUpload: enabled: false azure: storage: account: hopsfsdatastore container: hopsfs enableSoftDeletes: false identityClientId: 9cf39842-07c7-444a-8a44-63ee91411b70 softDeletesRetentionDays: 90 blockReportDelay: 21600000 datanode: asyncUpload: queueCapacity: 1000 retryCount: 3 retryIntervalMs: 30000 threads: 8 cache: bypass: false deleteActivationPercentage: 70 deleteWait: 1200000 maxUploadThreads: 20 enabled: true gcs: bucket: enableVersioning: true location: europe-north1 name: hopsfs markBlocksCorruptOrMissingAfter: 3600000 numCommittedAllowed: 1 provider: '' restoreFromBackup: null s3: bucket: aclBucketOwnerFullControl: false encryption: bucketKeyEnabled: true enabled: false mode: SSE-KMS userKeyARN: '' name: '' versioning: null bypassGovernanceRetention: false credentialsSecret: access_key_id: '' name: '' secret_key_id: '' disableCertChecking: false disableChecksumValidation: false endpoint: null region: '' signingRegion: null storeSmallFilesInDB: false ```
`hopsfs.objectStorage.asyncUpload.enabled` # { #helm.hopsfs.objectStorage.asyncUpload.enabled } : Type `bool`, default `false`. Enable async cloud upload. dfs.cloud.async.upload.enabled (shared NN + DN scope) `hopsfs.objectStorage.azure.storage.account` # { #helm.hopsfs.objectStorage.azure.storage.account } : Type `string`, default `"hopsfsdatastore"`. `hopsfs.objectStorage.azure.storage.container` # { #helm.hopsfs.objectStorage.azure.storage.container } : Type `string`, default `"hopsfs"`. `hopsfs.objectStorage.azure.storage.enableSoftDeletes` # { #helm.hopsfs.objectStorage.azure.storage.enableSoftDeletes } : Type `bool`, default `false`. `hopsfs.objectStorage.azure.storage.identityClientId` # { #helm.hopsfs.objectStorage.azure.storage.identityClientId } : Type `string`, default `"9cf39842-07c7-444a-8a44-63ee91411b70"`. `hopsfs.objectStorage.azure.storage.softDeletesRetentionDays` # { #helm.hopsfs.objectStorage.azure.storage.softDeletesRetentionDays } : Type `int`, default `90`. `hopsfs.objectStorage.blockReportDelay` # { #helm.hopsfs.objectStorage.blockReportDelay } : Type `int`, default `21600000`. dfs.cloud.block.report.delay `hopsfs.objectStorage.datanode.asyncUpload.queueCapacity` # { #helm.hopsfs.objectStorage.datanode.asyncUpload.queueCapacity } : Type `int`, default `1000`. Bounded async-upload queue size; overflow triggers sync fallback. dfs.cloud.dn.async.upload.queue.capacity `hopsfs.objectStorage.datanode.asyncUpload.retryCount` # { #helm.hopsfs.objectStorage.datanode.asyncUpload.retryCount } : Type `int`, default `3`. Retry attempts before re-enqueueing a failed upload at queue tail. dfs.cloud.dn.async.upload.retry.count `hopsfs.objectStorage.datanode.asyncUpload.retryIntervalMs` # { #helm.hopsfs.objectStorage.datanode.asyncUpload.retryIntervalMs } : Type `int`, default `30000`. Base backoff in ms between retries (exponential, capped at 5 min). dfs.cloud.dn.async.upload.retry.interval.ms `hopsfs.objectStorage.datanode.asyncUpload.threads` # { #helm.hopsfs.objectStorage.datanode.asyncUpload.threads } : Type `int`, default `8`. Async-upload worker pool size. dfs.cloud.dn.async.upload.threads `hopsfs.objectStorage.datanode.cache.bypass` # { #helm.hopsfs.objectStorage.datanode.cache.bypass } : Type `bool`, default `false`. `hopsfs.objectStorage.datanode.cache.deleteActivationPercentage` # { #helm.hopsfs.objectStorage.datanode.cache.deleteActivationPercentage } : Type `int`, default `70`. `hopsfs.objectStorage.datanode.cache.deleteWait` # { #helm.hopsfs.objectStorage.datanode.cache.deleteWait } : Type `int`, default `1200000`. Minimum block file age in milliseconds before a cached block is eligible for deletion. Prevents removing newly downloaded blocks that may still be served to remote clients. dfs.dn.cloud.cache.delete.wait `hopsfs.objectStorage.datanode.maxUploadThreads` # { #helm.hopsfs.objectStorage.datanode.maxUploadThreads } : Type `int`, default `20`. `hopsfs.objectStorage.enabled` # { #helm.hopsfs.objectStorage.enabled } : Type `bool`, default `true`. `hopsfs.objectStorage.gcs.bucket.enableVersioning` # { #helm.hopsfs.objectStorage.gcs.bucket.enableVersioning } : Type `bool`, default `true`. `hopsfs.objectStorage.gcs.bucket.location` # { #helm.hopsfs.objectStorage.gcs.bucket.location } : Type `string`, default `"europe-north1"`. `hopsfs.objectStorage.gcs.bucket.name` # { #helm.hopsfs.objectStorage.gcs.bucket.name } : Type `string`, default `"hopsfs"`. `hopsfs.objectStorage.markBlocksCorruptOrMissingAfter` # { #helm.hopsfs.objectStorage.markBlocksCorruptOrMissingAfter } : Type `int`, default `3600000`. dfs.cloud.mark.blocks.corrupt.or.missing.after `hopsfs.objectStorage.numCommittedAllowed` # { #helm.hopsfs.objectStorage.numCommittedAllowed } : Type `int`, default `1`. When using cloud storage, setting this to 1 allows NN to verify a block on cloud instead of waiting for DN signal. `hopsfs.objectStorage.provider` # { #helm.hopsfs.objectStorage.provider } : Type `string`, default `""`. object storage provicer, possible values are \["S3", "AZURE", "GCS"\] `hopsfs.objectStorage.restoreFromBackup` # { #helm.hopsfs.objectStorage.restoreFromBackup } : Type `string`, default `nil`. restore from backup flag which instructs HopsFS to roll backed deleted blocks `hopsfs.objectStorage.s3.bucket.aclBucketOwnerFullControl` # { #helm.hopsfs.objectStorage.s3.bucket.aclBucketOwnerFullControl } : Type `bool`, default `false`. `hopsfs.objectStorage.s3.bucket.encryption.bucketKeyEnabled` # { #helm.hopsfs.objectStorage.s3.bucket.encryption.bucketKeyEnabled } : Type `bool`, default `true`. `hopsfs.objectStorage.s3.bucket.encryption.enabled` # { #helm.hopsfs.objectStorage.s3.bucket.encryption.enabled } : Type `bool`, default `false`. `hopsfs.objectStorage.s3.bucket.encryption.mode` # { #helm.hopsfs.objectStorage.s3.bucket.encryption.mode } : Type `string`, default `"SSE-KMS"`. s3 encryption mode, possible values are \[SSE-S3, SSE-KMS\] `hopsfs.objectStorage.s3.bucket.encryption.userKeyARN` # { #helm.hopsfs.objectStorage.s3.bucket.encryption.userKeyARN } : Type `string`, default `""`. `hopsfs.objectStorage.s3.bucket.name` # { #helm.hopsfs.objectStorage.s3.bucket.name } : Type `string`, default `""`. `hopsfs.objectStorage.s3.bucket.versioning` # { #helm.hopsfs.objectStorage.s3.bucket.versioning } : Type `string`, default `nil`. versioning is enabled or disabled. If global backup is enabled then versioning will be set to be enabled by default unless overriden on the HopsFS subchart. `hopsfs.objectStorage.s3.bypassGovernanceRetention` # { #helm.hopsfs.objectStorage.s3.bypassGovernanceRetention } : Type `bool`, default `false`. `hopsfs.objectStorage.s3.credentialsSecret` # { #helm.hopsfs.objectStorage.s3.credentialsSecret } : Type `object`, default `{"access_key_id":"","name":"","secret_key_id":""}`. credentials secret configuration. If not defined the global._hopsworks.managedObjectStorage.s3.secret configuration will be used instead `hopsfs.objectStorage.s3.credentialsSecret.access_key_id` # { #helm.hopsfs.objectStorage.s3.credentialsSecret.access_key_id } : Type `string`, default `""`. Defaults to access-key-id `hopsfs.objectStorage.s3.credentialsSecret.name` # { #helm.hopsfs.objectStorage.s3.credentialsSecret.name } : Type `string`, default `""`. Default to aws-credentials `hopsfs.objectStorage.s3.credentialsSecret.secret_key_id` # { #helm.hopsfs.objectStorage.s3.credentialsSecret.secret_key_id } : Type `string`, default `""`. Defaults to secret-access-key `hopsfs.objectStorage.s3.disableCertChecking` # { #helm.hopsfs.objectStorage.s3.disableCertChecking } : Type `bool`, default `false`. disable TLS certificate checking for the S3 endpoint (useful with self-signed certs / MinIO) `hopsfs.objectStorage.s3.disableChecksumValidation` # { #helm.hopsfs.objectStorage.s3.disableChecksumValidation } : Type `bool`, default `false`. disable client-side checksum validation for S3 objects `hopsfs.objectStorage.s3.endpoint` # { #helm.hopsfs.objectStorage.s3.endpoint } : Type `string`, default `nil`. s3 provider endpoint `hopsfs.objectStorage.s3.region` # { #helm.hopsfs.objectStorage.s3.region } : Type `string`, default `""`. `hopsfs.objectStorage.s3.signingRegion` # { #helm.hopsfs.objectStorage.s3.signingRegion } : Type `string`, default `nil`. s3 provider signing region `hopsfs.objectStorage.storeSmallFilesInDB` # { #helm.hopsfs.objectStorage.storeSmallFilesInDB } : Type `bool`, default `false`.
## security { #helm-values-hopsfs-security } ??? example "Defaults as YAML" ```yaml hopsfs: security: fsGroup: 1234 runAsGroup: 1234 runAsUser: 1506 securityActions: maxConnectionsPerRoute: 30 tde: enabled: false kmsUri: '' tls: enabled: true ```
`hopsfs.security.fsGroup` # { #helm.hopsfs.security.fsGroup } : Type `int`, default `1234`. `hopsfs.security.runAsGroup` # { #helm.hopsfs.security.runAsGroup } : Type `int`, default `1234`. `hopsfs.security.runAsUser` # { #helm.hopsfs.security.runAsUser } : Type `int`, default `1506`. `hopsfs.security.securityActions.maxConnectionsPerRoute` # { #helm.hopsfs.security.securityActions.maxConnectionsPerRoute } : Type `int`, default `30`. Maximum number of concurrent HTTP connections to Hopsworks CA. Currently used only by WebHDFS `hopsfs.security.tde` # { #helm.hopsfs.security.tde } : Type `object`, default `{"enabled":false,"kmsUri":""}`. Configuration for Transparent Data Encryption `hopsfs.security.tde.enabled` # { #helm.hopsfs.security.tde.enabled } : Type `bool`, default `false`. Flag to enable encryption-at-rest `hopsfs.security.tde.kmsUri` # { #helm.hopsfs.security.tde.kmsUri } : Type `string`, default `""`. KMS URI when encryption-at-rest is enabled ie `hopsfs.security.tls.enabled` # { #helm.hopsfs.security.tls.enabled } : Type `bool`, default `true`.
================================================================================ # hopsfs-csi Source: https://docs.hopsworks.ai/latest/setup_installation/common/helm_chart_values/hopsfs-csi/ # HopsFS CSI driver values { #helm-values-hopsfs-csi } Values under `hopsfs-csi` configure the CSI driver that mounts HopsFS into pods. _Generated from the Hopsworks Helm chart `5.2.0-alpha-1791549041` (Hopsworks `5.2.0`)._ Deployed when `global._hopsworks.csi.enabled` is `true`. ## General { #helm-values-hopsfs-csi-general } ??? example "Defaults as YAML" ```yaml hopsfs-csi: cleanupOnUninstall: enabled: true ttlSecondsAfterFinished: null fullnameOverride: hopsfs-csi hopsworkslib: {} image: pullPolicy: IfNotPresent registry: docker.hops.works repository: hopsworks/hopsfs-csi tag: 0.1.0-SNAPSHOT nameOverride: hopsfs-csi nodeSelector: {} sidecars: livenessProbe: pullPolicy: IfNotPresent repository: registry.k8s.io/sig-storage/livenessprobe resources: limits: cpu: 100m memory: 128Mi requests: cpu: 10m memory: 64Mi tag: v2.16.0 nodeDriverRegistrar: pullPolicy: IfNotPresent repository: registry.k8s.io/sig-storage/csi-node-driver-registrar resources: limits: cpu: 400m memory: 512Mi requests: cpu: 40m memory: 128Mi tag: v2.14.0 registry: docker.hops.works tolerations: - operator: Exists ```
`hopsfs-csi` # { #helm.hopsfs-csi } : Type `object`, default `{}`. override hopsfs-csi values. Defaults live in charts/hopsfs-csi/values.yaml; the image comes from global._hopsworks.csi.image and global._hopsworks.imageRegistry, not from here. `hopsfs-csi.cleanupOnUninstall` # { #helm.hopsfs-csi.cleanupOnUninstall } : Type `object`, default `{"enabled":true,"ttlSecondsAfterFinished":null}`. post-delete cleanup of what `helm uninstall` does not own: the CSIDriver object (a pre-install hook, HWORKS-3285) and, on OpenShift, the node plugin's SCC. Best effort: a failed delete is logged, not a failed uninstall. The CSIDriver is cluster-scoped and shared by every release of this chart in the cluster, so it is deleted only when its Helm annotations name the uninstalled release and no other release's node plugin pods exist. The hw-kyverno PolicyException for the plugin, if enabled, is still left behind. A reinstall does not depend on this hook. `hopsfs-csi.cleanupOnUninstall.enabled` # { #helm.hopsfs-csi.cleanupOnUninstall.enabled } : Type `bool`, default `true`. enable the post-delete cleanup hook `hopsfs-csi.cleanupOnUninstall.ttlSecondsAfterFinished` # { #helm.hopsfs-csi.cleanupOnUninstall.ttlSecondsAfterFinished } : Type `int or null`, default `nil`. TTL of the finished hook Job; null falls through to the global default `hopsfs-csi.fullnameOverride` # { #helm.hopsfs-csi.fullnameOverride } : Type `string`, default `"hopsfs-csi"`. override the full chart name `hopsfs-csi.hopsworkslib` # { #helm.hopsfs-csi.hopsworkslib } : Type `object`, default `{}`. override hopsworkslib values `hopsfs-csi.image` # { #helm.hopsfs-csi.image } : Type `object`. The hopsfs-csi image, run by the node plugin. The same image is the workload FUSE sidecar (Hopsworks-generated pods, Airflow, Trino): the fusermount3 proxy in the sidecar and the fd server in the plugin are two halves of one protocol, so the two ends must never drift on a node. Under the umbrella chart every field here is overridden by the single `global._hopsworks.csi.image` block (and `global._hopsworks.imageRegistry` for the registry), which is also what the sidecars and the Hopsworks `csi_sidecar_image` variable are seeded from. The chart is only installed through the umbrella, so set the global rather than these. ??? note "Default" ```yaml pullPolicy: IfNotPresent registry: docker.hops.works repository: hopsworks/hopsfs-csi tag: 0.1.0-SNAPSHOT ``` `hopsfs-csi.nameOverride` # { #helm.hopsfs-csi.nameOverride } : Type `string`, default `"hopsfs-csi"`. override the chart name `hopsfs-csi.nodeSelector` # { #helm.hopsfs-csi.nodeSelector } : Type `object`, default `{}`. nodeSelector of the node plugin DaemonSet. Deliberately NOT inherited from global._hopsworks.nodeSelector: that selector pins the Hopsworks services to their nodes, while the plugin must run on every node a HopsFS-mounting pod can land on (job and GPU pools included), or those pods hang in ContainerCreating. Set it only to keep the plugin off nodes that will never run such a pod. `hopsfs-csi.sidecars` # { #helm.hopsfs-csi.sidecars } : Type `object`. The sig-storage helper containers next to the node plugin. ??? note "Default" ```yaml livenessProbe: pullPolicy: IfNotPresent repository: registry.k8s.io/sig-storage/livenessprobe resources: limits: cpu: 100m memory: 128Mi requests: cpu: 10m memory: 64Mi tag: v2.16.0 nodeDriverRegistrar: pullPolicy: IfNotPresent repository: registry.k8s.io/sig-storage/csi-node-driver-registrar resources: limits: cpu: 400m memory: 512Mi requests: cpu: 40m memory: 128Mi tag: v2.14.0 registry: docker.hops.works ``` `hopsfs-csi.sidecars.registry` # { #helm.hopsfs-csi.sidecars.registry } : Type `string`, default `"docker.hops.works"`. The mirror the sig-storage images are pulled from (they keep their upstream registry.k8s.io/... path under it). Overridden by global._hopsworks.imageRegistry under the umbrella chart, so an air-gapped install only sets the global. `hopsfs-csi.tolerations` # { #helm.hopsfs-csi.tolerations } : Type `list`, default `[{"operator":"Exists"}]`. tolerations of the node plugin DaemonSet. `operator: Exists` tolerates every taint for the reason above; narrow it only together with nodeSelector.
## node { #helm-values-hopsfs-csi-node } ??? example "Defaults as YAML" ```yaml hopsfs-csi: node: kubeletRootDir: /var/lib/kubelet priorityClassName: system-node-critical resources: limits: cpu: 500m memory: 256Mi requests: cpu: 50m memory: 64Mi serviceAccount: create: true name: hopsfs-csi-node-sa terminationGracePeriodSeconds: 300 ```
`hopsfs-csi.node.kubeletRootDir` # { #helm.hopsfs-csi.node.kubeletRootDir } : Type `string`, default `"/var/lib/kubelet"`. `hopsfs-csi.node.priorityClassName` # { #helm.hopsfs-csi.node.priorityClassName } : Type `string`, default `"system-node-critical"`. priorityClassName of the node plugin pods. system-node-critical so a full node never evicts the plugin (a node without it cannot start any CSI-mounted pod) and so it is scheduled before the workloads that depend on it. `hopsfs-csi.node.resources.limits.cpu` # { #helm.hopsfs-csi.node.resources.limits.cpu } : Type `string`, default `"500m"`. `hopsfs-csi.node.resources.limits.memory` # { #helm.hopsfs-csi.node.resources.limits.memory } : Type `string`, default `"256Mi"`. `hopsfs-csi.node.resources.requests.cpu` # { #helm.hopsfs-csi.node.resources.requests.cpu } : Type `string`, default `"50m"`. `hopsfs-csi.node.resources.requests.memory` # { #helm.hopsfs-csi.node.resources.requests.memory } : Type `string`, default `"64Mi"`. `hopsfs-csi.node.serviceAccount.create` # { #helm.hopsfs-csi.node.serviceAccount.create } : Type `bool`, default `true`. `hopsfs-csi.node.serviceAccount.name` # { #helm.hopsfs-csi.node.serviceAccount.name } : Type `string`, default `"hopsfs-csi-node-sa"`. `hopsfs-csi.node.terminationGracePeriodSeconds` # { #helm.hopsfs-csi.node.terminationGracePeriodSeconds } : Type `int`, default `300`. How long kubelet lets a stopping plugin pod run before killing it. When the plugin is removed together with the pods that mount through it (helm uninstall, namespace delete), it keeps serving kubelet's unmounts until none is left and exits then, so this is only ever used up when a consumer is itself stuck; a plugin that is merely being replaced (rollout, eviction) exits at once. Set it above the longest terminationGracePeriodSeconds of any HopsFS-mounting pod, or those pods hang in Terminating after the plugin is gone (HWORKS-3287).
## openshift { #helm-values-hopsfs-csi-openshift } ??? example "Defaults as YAML" ```yaml hopsfs-csi: openshift: csiEphemeralVolumeProfile: restricted existingSecurityContextConstraints: check: enabled: true ttlSecondsAfterFinished: null name: '' securityContextConstraints: false ```
`hopsfs-csi.openshift` # { #helm.hopsfs-csi.openshift } : Type `object`. OpenShift-only settings of the node plugin: the CSI Volume Admission profile label on the CSIDriver and whether to ship a SecurityContextConstraints for the plugin's service account. ??? note "Default" ```yaml csiEphemeralVolumeProfile: restricted existingSecurityContextConstraints: check: enabled: true ttlSecondsAfterFinished: null name: '' securityContextConstraints: false ``` `hopsfs-csi.openshift.csiEphemeralVolumeProfile` # { #helm.hopsfs-csi.openshift.csiEphemeralVolumeProfile } : Type `string`, default `"restricted"`. OpenShift CSI Volume Admission profile for inline ephemeral volumes, emitted as the security.openshift.io/csi-ephemeral-volume-profile label on the CSIDriver. Must be one of restricted, baseline or privileged and must not be empty on OpenShift: an absent label makes the CSIInlineVolumeSecurity admission plugin require enforce=privileged of every namespace that mounts the driver, which rejects every Airflow and Trino pod. `restricted` is correct for this driver (see templates/csidriver.yaml for why) and the label is inert off OpenShift. `hopsfs-csi.openshift.existingSecurityContextConstraints` # { #helm.hopsfs-csi.openshift.existingSecurityContextConstraints } : Type `object`, default `{"check":{"enabled":true,"ttlSecondsAfterFinished":null},"name":""}`. Let the cluster administrators own the node plugin's SCC instead of this chart. Off by default: the chart ships the SCC (securityContextConstraints above). Setting `name` to the SCC the administrators created turns the shipped one off, whatever securityContextConstraints and global._hopsworks.openshift.enabled say, keeps the post-delete cleanup off it, and adds a pre-install/pre-upgrade hook Job that fails the release early, with the reason in its log, when that SCC is missing or does not admit the plugin. Without the check a wrong SCC leaves the DaemonSet with no pods and every CSI-mounted workload in ContainerCreating, with no event naming the cause. The Job runs a PodSecurityPolicySubjectReview of the exact DaemonSet pod template for the plugin's service account (users, groups and RBAC `use` grants all count; the account need not exist yet), and a server-side dry-run create of a representative HopsFS-mounting workload pod in the release namespace: the unprivileged FUSE sidecar at global._hopsworks.executor_uid plus an inline csi volume of this driver, which also exercises the CSI Volume Admission plugin against the CSIDriver's profile label (HWORKS-3285) and restricted-v2's uid range. The reference manifest to hand to the administrators is the chart's own: `helm template . --set hopsfs-csi.openshift.securityContextConstraints=true --show-only charts/hopsfs-csi/templates/scc.yaml`; whatever they derive from it has to grant `system:serviceaccount::hopsfs-csi-node-sa`. `hopsfs-csi.openshift.existingSecurityContextConstraints.check` # { #helm.hopsfs-csi.openshift.existingSecurityContextConstraints.check } : Type `object`, default `{"enabled":true,"ttlSecondsAfterFinished":null}`. the pre-install/pre-upgrade admission check, run only when `name` is set `hopsfs-csi.openshift.existingSecurityContextConstraints.check.enabled` # { #helm.hopsfs-csi.openshift.existingSecurityContextConstraints.check.enabled } : Type `bool`, default `true`. run the check. Off only to install against an SCC the OpenShift review API cannot evaluate; the install then fails late, in the DaemonSet, if the SCC is wrong. `hopsfs-csi.openshift.existingSecurityContextConstraints.check.ttlSecondsAfterFinished` # { #helm.hopsfs-csi.openshift.existingSecurityContextConstraints.check.ttlSecondsAfterFinished } : Type `int or null`, default `nil`. TTL of the finished hook Job; null falls through to the global default `hopsfs-csi.openshift.existingSecurityContextConstraints.name` # { #helm.hopsfs-csi.openshift.existingSecurityContextConstraints.name } : Type `string`, default `""`. name of the administrator-created SecurityContextConstraints; empty, the default, means the chart ships its own `hopsfs-csi.openshift.securityContextConstraints` # { #helm.hopsfs-csi.openshift.securityContextConstraints } : Type `bool`, default `false`. Ship a SecurityContextConstraints granting the node plugin's service account what the DaemonSet needs on OpenShift (privileged, hostPath volumes, Bidirectional mount propagation, root); without it restricted-v2 rejects the plugin pod and every CSI-mounted workload sits in ContainerCreating. Rendered when this is true or global._hopsworks.openshift.enabled is true. Cluster-scoped and named after the release so two installs do not collide. Off, together with the global, when existingSecurityContextConstraints.name is set.
================================================================================ # hopsworks Source: https://docs.hopsworks.ai/latest/setup_installation/common/helm_chart_values/hopsworks/ # Hopsworks values { #helm-values-hopsworks } Values under `hopsworks` configure the Hopsworks backend: the Payara worker and admin deployments, ingress, the certificate authority, database migrations and backups. _Generated from the Hopsworks Helm chart `5.2.0-alpha-1791549041` (Hopsworks `5.2.0`)._ Deployed when [`global._hopsworks.full_platform`](global.md#helm.global._hopsworks.full_platform) is `true`. ## General { #helm-values-hopsworks-general } ??? example "Defaults as YAML" ```yaml hopsworks: agent_email: agent@hops.io agent_password: admin cleanupOnUninstall: enabled: true ttlSecondsAfterFinished: null create_secrets: true ddlConfigmapName: sql-ddl defaultServiceAccount: annotations: {} create: true dropDatabase: false flyway: replaceDDL: '' flywayconfigmapName: flyway-config fullnameOverride: null hopsfsMount: enabled: true mountPath: /hopsfs hopsworkslib: {} imageBuilder: image: image-builder tag: '0.2' imageBuilderServiceAccount: annotations: {} create: true name: image-builder instanceconfigmapName: instance-config jupyter: notebookConfig: allowOrigin: ${conf.allowOrigin} enableDownloads: true enableUploads: true lb: names: mysqld: mysqld-external rdrs: rdrs-external serviceAccount: annotations: {} nameOverride: null nodeSelector: {} objectStorageEnvInformation: null opensearchReindex: enabled: true payaraVersion: 6.2025.11-jdk21.0 payaraconfigmapName: post-boot-commands podAnnotations: {} podDisruptionBudget: hopsworksCA: enabled: true minAvailable: 1 worker: enabled: true minAvailable: 1 podLabels: {} replicaCount: worker: 2 resources: admin: auto_jvm: true jvm: memory: buffer: 2048 compressedClassSpaceSize: 512 heap: 4096 metaspace: 2048 nonMethodCodeHeapSize: 5 nonProfiledCodeHeapSize: 48 profiledCodeHeapSize: 48 worker: auto_jvm: true jvm: memory: buffer: 3072 compressedClassSpaceSize: 512 heap: 4096 metaspace: 2048 nonMethodCodeHeapSize: 5 nonProfiledCodeHeapSize: 48 profiledCodeHeapSize: 48 securityContext: {} setAdminHighPriority: true sparkConfigmapName: spark sqldmlconfigmapName: sql-dml sqlgrantsconfigmapName: sql-grants templatesconfigmapName: hopsworks-templates tolerations: [] topologySpreadConstraint: {} updateLoadBalancerDomains: resources: limits: cpu: 50m memory: 50M requests: cpu: 20m memory: 20M variables: kube_kserve_installed: true kube_serving_vllm_omni_versions: v0.28.0 kube_serving_vllm_versions: v0.28.0 volumeMounts: admin: [] migrate: [] worker: [] volumes: admin: [] migrate: [] worker: [] ```
`hopsworks` # { #helm.hopsworks } : Type `object`. override hopsworks values ??? note "Default" ```yaml lb: names: mysqld: mysqld-external rdrs: rdrs-external replicaCount: worker: 2 resources: admin: auto_jvm: true jvm: memory: buffer: 2048 compressedClassSpaceSize: 512 heap: 4096 metaspace: 2048 nonMethodCodeHeapSize: 5 nonProfiledCodeHeapSize: 48 profiledCodeHeapSize: 48 worker: auto_jvm: true jvm: memory: buffer: 3072 compressedClassSpaceSize: 512 heap: 4096 metaspace: 2048 nonMethodCodeHeapSize: 5 nonProfiledCodeHeapSize: 48 profiledCodeHeapSize: 48 variables: kube_kserve_installed: true kube_serving_vllm_omni_versions: v0.28.0 kube_serving_vllm_versions: v0.28.0 ``` `hopsworks.agent_email` # { #helm.hopsworks.agent_email } : Type `string`, default `"agent@hops.io"`. `hopsworks.agent_password` # { #helm.hopsworks.agent_password } : Type `string`, default `"admin"`. `hopsworks.cleanupOnUninstall` # { #helm.hopsworks.cleanupOnUninstall } : Type `object`, default `{"enabled":true,"ttlSecondsAfterFinished":null}`. post-delete cleanup of Hopsworks runtime leftovers (CA/setup-script Secrets & ConfigMaps annotated hopsworks.ai/project, plus the sql-dml/sql-grants/flyway-config hook ConfigMaps) that Helm/ArgoCD never tracked and so never prune. `hopsworks.cleanupOnUninstall.enabled` # { #helm.hopsworks.cleanupOnUninstall.enabled } : Type `bool`, default `true`. enable the post-delete Hopsworks cleanup hook `hopsworks.cleanupOnUninstall.ttlSecondsAfterFinished` # { #helm.hopsworks.cleanupOnUninstall.ttlSecondsAfterFinished } : Type `string`, default `nil`. ttlSecondsAfterFinished for the cleanup Job; null falls through to the global default `hopsworks.create_secrets` # { #helm.hopsworks.create_secrets } : Type `bool`, default `true`. If false, you have to create the secret for the hopsworks users (hopsworks-users-secrets) manually. If false and ldap/kerberos is enabled, the secret for ldap credentials must be created manually as well. `hopsworks.ddlConfigmapName` # { #helm.hopsworks.ddlConfigmapName } : Type `string`, default `"sql-ddl"`. `hopsworks.defaultServiceAccount.annotations` # { #helm.hopsworks.defaultServiceAccount.annotations } : Type `object`, default `{}`. annotations `hopsworks.defaultServiceAccount.create` # { #helm.hopsworks.defaultServiceAccount.create } : Type `bool`, default `true`. `hopsworks.dropDatabase` # { #helm.hopsworks.dropDatabase } : Type `bool`, default `false`. Should not be changed here. Leave this as a conscious choice for the user when installing. If set to true will drop databases when helm uninstall. `hopsworks.flyway.replaceDDL` # { #helm.hopsworks.flyway.replaceDDL } : Type `string`, default `""`. `hopsworks.flywayconfigmapName` # { #helm.hopsworks.flywayconfigmapName } : Type `string`, default `"flyway-config"`. `hopsworks.fullnameOverride` # { #helm.hopsworks.fullnameOverride } : Type `string`, default `nil`. override app fully qualified name `hopsworks.hopsfsMount.enabled` # { #helm.hopsworks.hopsfsMount.enabled } : Type `bool`, default `true`. `hopsworks.hopsfsMount.mountPath` # { #helm.hopsworks.hopsfsMount.mountPath } : Type `string`, default `"/hopsfs"`. `hopsworks.hopsworkslib` # { #helm.hopsworks.hopsworkslib } : Type `object`, default `{}`. override hopsworkslib values `hopsworks.imageBuilder.image` # { #helm.hopsworks.imageBuilder.image } : Type `string`, default `"image-builder"`. `hopsworks.imageBuilder.tag` # { #helm.hopsworks.imageBuilder.tag } : Type `string`, default `"0.2"`. 0.2 is a floor, not a preference. The OpenShift build path passes package-index credentials to buildah as -s id=NAME,src=PATH, and build-and-push.sh only learned that option in 0.2. Against 0.1 the script's getopts rejects it and every environment build on OpenShift fails, so this tag and the backend's docker_operations_image_builder_image move together. `hopsworks.imageBuilderServiceAccount` # { #helm.hopsworks.imageBuilderServiceAccount } : Type `object`, default `{"annotations":{},"create":true,"name":"image-builder"}`. Configuration for the Service Account running all user environment Image Building Jobs `hopsworks.instanceconfigmapName` # { #helm.hopsworks.instanceconfigmapName } : Type `string`, default `"instance-config"`. `hopsworks.jupyter.notebookConfig.allowOrigin` # { #helm.hopsworks.jupyter.notebookConfig.allowOrigin } : Type `string`, default `"${conf.allowOrigin}"`. `hopsworks.jupyter.notebookConfig.enableDownloads` # { #helm.hopsworks.jupyter.notebookConfig.enableDownloads } : Type `bool`, default `true`. Set to false to disable file downloads from the Jupyter UI. `hopsworks.jupyter.notebookConfig.enableUploads` # { #helm.hopsworks.jupyter.notebookConfig.enableUploads } : Type `bool`, default `true`. Set to false to disable file uploads from the Jupyter UI and reject base64 file uploads via the Contents API. Kernel/terminal file writes are unaffected. Requires JupyterLab >= 4.5. `hopsworks.lb.names.mysqld` # { #helm.hopsworks.lb.names.mysqld } : Type `string`, default `"mysqld-external"`. `hopsworks.lb.names.rdrs` # { #helm.hopsworks.lb.names.rdrs } : Type `string`, default `"rdrs-external"`. `hopsworks.lb.serviceAccount.annotations` # { #helm.hopsworks.lb.serviceAccount.annotations } : Type `object`, default `{}`. annotations `hopsworks.nameOverride` # { #helm.hopsworks.nameOverride } : Type `string`, default `nil`. override app chart name `hopsworks.nodeSelector` # { #helm.hopsworks.nodeSelector } : Type `object`, default `{}`. node selector configuration `hopsworks.objectStorageEnvInformation` # { #helm.hopsworks.objectStorageEnvInformation } : Type `string`, default `nil`. override object storage list of of environment variables `hopsworks.opensearchReindex.enabled` # { #helm.hopsworks.opensearchReindex.enabled } : Type `bool`, default `true`. Rebuild the featurestore search index after an upgrade, when what the index holds has changed since its last rebuild. The hook keys its request with the index generation, featurestore-index-5.1, which the chart bumps only with a change to what the backend indexes: a generation already rebuilt is answered with that run, so patch upgrades and ArgoCD syncs do not rebuild again, and a cluster that missed the rebuild gets it on its next upgrade. Runs requested by earlier charts carry no key, so the first upgrade with this chart rebuilds once. The rebuild runs in the backend after the upgrade and takes hours on a large cluster, with search incomplete until it finishes; its progress shows under Cluster Settings > Service Operations > OpenSearch Index Commands. A failed hook Job is removed after global._hopsworks.jobs.ttlSecondsAfterFinished in the default mode, or when the hook next renders; with global._hopsworks.mode set, as under ArgoCD, delete post-upgrade-opensearch-reindex-job by hand if you turn this off after a failed attempt. `hopsworks.payaraVersion` # { #helm.hopsworks.payaraVersion } : Type `string`, default `"6.2025.11-jdk21.0"`. `hopsworks.payaraconfigmapName` # { #helm.hopsworks.payaraconfigmapName } : Type `string`, default `"post-boot-commands"`. `hopsworks.podAnnotations` # { #helm.hopsworks.podAnnotations } : Type `object`, default `{}`. pod annotations `hopsworks.podDisruptionBudget.hopsworksCA.enabled` # { #helm.hopsworks.podDisruptionBudget.hopsworksCA.enabled } : Type `bool`, default `true`. `hopsworks.podDisruptionBudget.hopsworksCA.minAvailable` # { #helm.hopsworks.podDisruptionBudget.hopsworksCA.minAvailable } : Type `int`, default `1`. `hopsworks.podDisruptionBudget.worker.enabled` # { #helm.hopsworks.podDisruptionBudget.worker.enabled } : Type `bool`, default `true`. `hopsworks.podDisruptionBudget.worker.minAvailable` # { #helm.hopsworks.podDisruptionBudget.worker.minAvailable } : Type `int`, default `1`. `hopsworks.podLabels` # { #helm.hopsworks.podLabels } : Type `object`, default `{}`. pod labels `hopsworks.replicaCount.worker` # { #helm.hopsworks.replicaCount.worker } : Type `int`, default `2`. Number of Payara worker replicas. Not rendered while hpa.worker.enabled is true: the HPA then owns spec.replicas and this value is its minReplicas. Turning the HPA on for a running release drops the workers to 1 once, until the HPA scales them back up. `hopsworks.securityContext` # { #helm.hopsworks.securityContext } : Type `object`, default `{}`. custom security context for hopsworks admin and workers `hopsworks.setAdminHighPriority` # { #helm.hopsworks.setAdminHighPriority } : Type `bool`, default `true`. `hopsworks.sparkConfigmapName` # { #helm.hopsworks.sparkConfigmapName } : Type `string`, default `"spark"`. `hopsworks.sqldmlconfigmapName` # { #helm.hopsworks.sqldmlconfigmapName } : Type `string`, default `"sql-dml"`. `hopsworks.sqlgrantsconfigmapName` # { #helm.hopsworks.sqlgrantsconfigmapName } : Type `string`, default `"sql-grants"`. `hopsworks.templatesconfigmapName` # { #helm.hopsworks.templatesconfigmapName } : Type `string`, default `"hopsworks-templates"`. `hopsworks.tolerations` # { #helm.hopsworks.tolerations } : Type `list`, default `[]`. `hopsworks.topologySpreadConstraint` # { #helm.hopsworks.topologySpreadConstraint } : Type `object`, default `{}`. The default topology spread constraint. If not defined the global topology spread constraint would be used instead. `hopsworks.updateLoadBalancerDomains` # { #helm.hopsworks.updateLoadBalancerDomains } : Type `object`. update_load_balancer_domains Job configuration ??? note "Default" ```yaml resources: limits: cpu: 50m memory: 50M requests: cpu: 20m memory: 20M ``` `hopsworks.updateLoadBalancerDomains.resources` # { #helm.hopsworks.updateLoadBalancerDomains.resources } : Type `object`. resource limits configuration ??? note "Default" ```yaml limits: cpu: 50m memory: 50M requests: cpu: 20m memory: 20M ``` `hopsworks.volumeMounts.admin` # { #helm.hopsworks.volumeMounts.admin } : Type `list`, default `[]`. `hopsworks.volumeMounts.migrate` # { #helm.hopsworks.volumeMounts.migrate } : Type `list`, default `[]`. `hopsworks.volumeMounts.worker` # { #helm.hopsworks.volumeMounts.worker } : Type `list`, default `[]`. `hopsworks.volumes.admin` # { #helm.hopsworks.volumes.admin } : Type `list`, default `[]`. `hopsworks.volumes.migrate` # { #helm.hopsworks.volumes.migrate } : Type `list`, default `[]`. `hopsworks.volumes.worker` # { #helm.hopsworks.volumes.worker } : Type `list`, default `[]`.
## buildkitd { #helm-values-hopsworks-buildkitd } ??? example "Defaults as YAML" ```yaml hopsworks: buildkitd: gc: cacheMountKeepBytes: 20GiB cacheMountKeepDuration: 168h keySyntax: maxUsedSpace minFreeSpace: 10GiB totalKeepBytes: 60GiB image: '' insecureRegistries: [] maxParallelism: 4 name: buildkitd nodeSelector: {} podAnnotations: {} port: 1234 priorityClass: create: true value: 1000000 priorityClassName: '' registry: '' replicas: 1 resources: limits: cpu: '8' memory: 16Gi requests: cpu: '8' memory: 16Gi rootless: deviceInjection: none devicePluginResource: github.com/fuse enabled: false preflight: true tagSuffix: -rootless user: 1000 serviceAccountName: '' storage: 100Gi storageClassName: '' tag: '' tls: enabled: true locality: buildkitd tolerations: [] ```
`hopsworks.buildkitd.gc.cacheMountKeepBytes` # { #helm.hopsworks.buildkitd.gc.cacheMountKeepBytes } : Type `string`, default `"20GiB"`. Budget for package cache mounts, kept separate so they cannot evict base image snapshots. Scale by the number of projects that build custom environments, not by total projects. Every size here is binary and comparable to buildkitd.storage, since the two are weighed against each other. Spell it "GiB" and not the Kubernetes "Gi": BuildKit parses these with docker/go-units, which rejects a bare "Gi" and refuses to start the daemon. `hopsworks.buildkitd.gc.cacheMountKeepDuration` # { #helm.hopsworks.buildkitd.gc.cacheMountKeepDuration } : Type `string`, default `"168h"`. `hopsworks.buildkitd.gc.keySyntax` # { #helm.hopsworks.buildkitd.gc.keySyntax } : Type `string`, default `"maxUsedSpace"`. Which GC key names to emit: "keepBytes" for BuildKit up to ~v0.16, "maxUsedSpace" for later releases. Must match the deployed BuildKit, and getting it wrong is silent rather than loud: v0.31.2 still accepts keepBytes, but maps it to reservedSpace, which is a floor and not a ceiling (cmd/buildkitd/config: "Deprecated: use ReservedSpace instead"). Emitting keepBytes against a modern daemon therefore turns the budget below into an amount that is guaranteed to be kept rather than never exceeded, and the state volume fills. Explicit rather than inferred from the tag, because the tag can point at a mirror of any version. Constrained by the schema: a typo used to fall through to the maxUsedSpace branch silently, which is the failure this key exists to prevent. `hopsworks.buildkitd.gc.minFreeSpace` # { #helm.hopsworks.buildkitd.gc.minFreeSpace } : Type `string`, default `"10GiB"`. Free space the collector tries to leave on the volume, never going below what the policies above guarantee. The only budget here expressed against actual free space rather than against BuildKit's accounting of its own records, so it is the backstop for the daemon's other state on the volume, which no gcpolicy covers. Empty omits it. `hopsworks.buildkitd.gc.totalKeepBytes` # { #helm.hopsworks.buildkitd.gc.totalKeepBytes } : Type `string`, default `"60GiB"`. Total retained. Must stay well above the unpacked size of every base image served, and together with cacheMountKeepBytes must leave headroom under buildkitd.storage. GC is reactive, so a burst overshoots the threshold before collection catches up, and a full state volume gives ENOSPC on snapshot writes: on a ReadWriteOnce PVC recovery is volume expansion or deleting the PVC and cold-starting every base image snapshot and cache. `hopsworks.buildkitd.image` # { #helm.hopsworks.buildkitd.image } : Type `string`, default `""`. Daemon image name. Empty falls back to dockerRegistry.buildkit.image. `hopsworks.buildkitd.insecureRegistries` # { #helm.hopsworks.buildkitd.insecureRegistries } : Type `list`, default `[]`. Registries the daemon should treat as insecure, e.g. the in-cluster registry when it is served over plain HTTP. `hopsworks.buildkitd.maxParallelism` # { #helm.hopsworks.buildkitd.maxParallelism } : Type `int`, default `4`. Concurrent build steps the daemon will run. Bounds peak memory at roughly this many per-build peaks, so overload becomes slow instead of an OOMKill that fails every in-flight build. It does not bound the number of build Jobs: further builds still hold a client pod and a blocked backend thread, so this protects the daemon rather than the cluster. Size it against memory, not CPU. BuildKit's own documented example, and conservative against a heavy wheel or CUDA build peaking near 2Gi. `hopsworks.buildkitd.name` # { #helm.hopsworks.buildkitd.name } : Type `string`, default `"buildkitd"`. `hopsworks.buildkitd.nodeSelector` # { #helm.hopsworks.buildkitd.nodeSelector } : Type `object`, default `{}`. `hopsworks.buildkitd.podAnnotations` # { #helm.hopsworks.buildkitd.podAnnotations } : Type `object`, default `{}`. Extra annotations on the daemon pod. Exists mainly for CRI-O's workload activation route: when a node scopes allowed_annotations to a \[crio.runtime.workloads.*\] table rather than a runtime handler, the pod must carry that workload's activation_annotation for the device annotation to be honoured. See values.rhel8.yaml. `hopsworks.buildkitd.port` # { #helm.hopsworks.buildkitd.port } : Type `int`, default `1234`. `hopsworks.buildkitd.priorityClass.create` # { #helm.hopsworks.buildkitd.priorityClass.create } : Type `bool`, default `true`. Create a priority class for shared Hopsworks services that user workloads depend on. The daemon needs one: it is now the pod that does the building, while its own build Jobs run at docker_operations_buildkit_priority_class. Without it the daemon sits at priority 0 and a burst of build Jobs can preempt the very daemon they are about to talk to, which fails every in-flight build including the ones that displaced it. Turning this off is only safe if the build Jobs are left at priority 0 too. `hopsworks.buildkitd.priorityClass.value` # { #helm.hopsworks.buildkitd.priorityClass.value } : Type `int`, default `1000000`. Must not be below the build Jobs' priority (docker_operations_buildkit_priority_class, 0 by default): the scheduler only preempts a lower priority, so a daemon below its own clients can be evicted by a burst of build Jobs. `hopsworks.buildkitd.priorityClassName` # { #helm.hopsworks.buildkitd.priorityClassName } : Type `string`, default `""`. Priority class for the daemon. Empty uses the chart's own core-service class when priorityClass.create is on, which is the default; set a name to override it, or disable creation and leave this empty to run at priority 0 as before. `hopsworks.buildkitd.registry` # { #helm.hopsworks.buildkitd.registry } : Type `string`, default `""`. Registry for the daemon image. Empty uses the chart's usual registry. Set this when the daemon runs a different BuildKit version from the per-build image, which is mirrored separately. `hopsworks.buildkitd.replicas` # { #helm.hopsworks.buildkitd.replicas } : Type `int`, default `1`. Each replica gets its own ReadWriteOnce state volume. Projects are pinned to a replica by id, so a project keeps hitting the daemon that already unpacked its base image. The address and replica count are filled in from here. Pinning is what makes the cache useful and is also why this is not failover: a project is not retried against another replica, and a project that builds far more than the others stays on the one daemon. `hopsworks.buildkitd.resources` # { #helm.hopsworks.buildkitd.resources } : Type `object`, default `{"limits":{"cpu":"8","memory":"16Gi"},"requests":{"cpu":"8","memory":"16Gi"}}`. Requests equal limits, so the daemon is Guaranteed rather than Burstable. It is the shared build bottleneck for every project, and node-pressure eviction ranks by QoS class first and then by usage above request: Burstable with a 2 CPU request and an 8 CPU limit made it both first to evict and capped at 2 CPUs of sustained capacity on a contended node, which is the opposite of what the spread suggested. Raise both together on a dedicated builder node. `hopsworks.buildkitd.rootless.deviceInjection` # { #helm.hopsworks.buildkitd.rootless.deviceInjection } : Type `string`, default `"none"`. How the daemon is given /dev/fuse, which fuse-overlayfs needs. Kubernetes has no portable PodSpec equivalent of docker --device, so this is necessarily runtime-specific. none no device. The daemon must then use the native snapshotter, which copies whole layers instead of stacking them: correct, much slower, much larger. crio adds the io.kubernetes.cri-o.Devices annotation, which CRI-O honours. devicePlugin requests devicePluginResource below, for a device plugin or CDI that advertises /dev/fuse on upstream Kubernetes. A plain hostPath is deliberately not offered. It makes the device node visible and leaves every open returning EPERM, which looks like it works right up until a build runs. Measured on a containerd cluster: the node appears as crw-rw-rw- and the open fails. Asking for the fuse-overlayfs snapshotter with this set to none is refused at template time rather than deployed, since the daemon would either fail to start or quietly do something else. Nothing here is exercised on an enforcing-SELinux node yet; expect to need a matching SCC or policy there. `hopsworks.buildkitd.rootless.devicePluginResource` # { #helm.hopsworks.buildkitd.rootless.devicePluginResource } : Type `string`, default `"github.com/fuse"`. Resource requested when deviceInjection is devicePlugin. The name comes from whichever plugin or CDI provider is installed; there is no standard one. Must be a qualified extended-resource name (domain/name): an empty value would render a resources entry Kubernetes rejects, so the chart refuses it. `hopsworks.buildkitd.rootless.enabled` # { #helm.hopsworks.buildkitd.rootless.enabled } : Type `bool`, default `false`. Run the daemon as an unprivileged user instead of as root with privileged: true. Off by default. Switching an existing install is a one-way, disruptive change: the rootless daemon keeps its state in its own subdirectory of the volume, so it starts cold and unpacks every base image again, and the rootful tree stays on the volume taking space until the volume is replaced. Switching back has the mirror-image cost. Size buildkitd.storage for one tree, not two, and give the StatefulSet a fresh volume when the extra copy does not fit. What it costs, and this is more than a performance note. The daemon loses the process sandbox for RUN steps (--oci-worker-no-process-sandbox), which is the price of not needing privileged. A RUN step then shares the daemon's PID namespace at the same uid, so it can signal and potentially ptrace the daemon, and through /proc//root reach the daemon's mount namespace, where the mTLS private key is. One daemon also holds every project's build state and cache. So this mode is for a single trust zone only. Do not enable it where concurrent projects are mutually untrusted. Losing the daemon to a build step is an accepted denial of service; the credential reachability is what bounds where the mode may be used. It buys back: no host root, no privileged container, and no hostPID, which also means a daemon cannot outlive its container and strand the state lock. Requires a kernel that supports the configured snapshotter unprivileged. Rootless overlayfs needs 5.11 or later; below that, set variables.docker_operations_oci_worker_snapshotter to "native" and expect slower builds. fuse-overlayfs is the faster fallback on those kernels, and it needs the daemon to actually hold /dev/fuse. See deviceInjection below: a hostPath is not enough, because access is governed by the container's device cgroup rather than by the filesystem. `hopsworks.buildkitd.rootless.preflight` # { #helm.hopsworks.buildkitd.rootless.preflight } : Type `bool`, default `true`. Run a preflight check before the daemon starts, so a node that cannot support the requested configuration fails at once and says why, rather than degrading silently. Verifies node-level facts: unprivileged user namespaces, and kernel FUSE support when the configuration actually uses FUSE (fuse-overlayfs, or any device injection). The /dev/fuse open check is not here: it runs in the daemon container's own entrypoint, because a device plugin allocates per container and the device never reaches an init container. Only meaningful when rootless is on. `hopsworks.buildkitd.rootless.tagSuffix` # { #helm.hopsworks.buildkitd.rootless.tagSuffix } : Type `string`, default `"-rootless"`. Appended to the daemon image tag when rootless is on, since upstream ships the rootless daemon as a separate image variant. An explicit buildkitd.tag overrides this. `hopsworks.buildkitd.rootless.user` # { #helm.hopsworks.buildkitd.rootless.user } : Type `int`, default `1000`. uid and gid the rootless image runs as, and the pod's fsGroup, which is what makes the state volume and the certificate mount readable without an init container that chowns the whole volume. `hopsworks.buildkitd.serviceAccountName` # { #helm.hopsworks.buildkitd.serviceAccountName } : Type `string`, default `""`. Service account for the daemon pod. Empty runs it as the namespace default with no token mounted, which is what it needs to build. Set it only to give the daemon a cloud identity of its own, for instance an IRSA-annotated account so a registry cache export can reach S3. The daemon's identity is not the build pod's. Builds run as hopsworks-default, so an annotation applied through defaultServiceAccount.annotations reaches them and not this StatefulSet. That distinction only appears once the daemon moves out of the build pod, and it is the reason an S3 exporter that worked daemonless can stop working here. `hopsworks.buildkitd.storage` # { #helm.hopsworks.buildkitd.storage } : Type `string`, default `"100Gi"`. Must comfortably exceed the total unpacked size of the base images in use, plus the package cache. Below that, each build evicts the base image the next one needs and the cache is slower than none. `hopsworks.buildkitd.storageClassName` # { #helm.hopsworks.buildkitd.storageClassName } : Type `string`, default `""`. `hopsworks.buildkitd.tag` # { #helm.hopsworks.buildkitd.tag } : Type `string`, default `""`. Daemon image tag. Empty falls back to dockerRegistry.buildkit.tag. `hopsworks.buildkitd.tls.enabled` # { #helm.hopsworks.buildkitd.tls.enabled } : Type `bool`, default `true`. Require client certificates. On by default, and only worth turning off on a single-tenant install: buildctl over TCP is otherwise unauthenticated, and buildkitd gives RUN steps host networking, so a build step reaches the daemon on both the service address and 127.0.0.1 and the NetworkPolicy below does not contain it. Verified on a cluster: with mTLS the connection is closed before the API answers, and the client key lives in the build job pod rather than the RUN filesystem, so user code cannot present it. `hopsworks.buildkitd.tls.locality` # { #helm.hopsworks.buildkitd.tls.locality } : Type `string`, default `"buildkitd"`. `hopsworks.buildkitd.tolerations` # { #helm.hopsworks.buildkitd.tolerations } : Type `list`, default `[]`.
## create_certificate { #helm-values-hopsworks-create_certificate } ??? example "Defaults as YAML" ```yaml hopsworks: create_certificate: add: [] auto_generate: true cn: hopsworks.local.ai configmap: payara-cacerts enabled: false secretName: hopsworks-tls trustManager: aliasPrefix: trust-manager configMapName: '' enabled: false key: ca-bundle.crt ```
`hopsworks.create_certificate.add` # { #helm.hopsworks.create_certificate.add } : Type `list`, default `[]`. the following list is used to add additional certificates to trust. Notice this wont work in air gapped. An example - name: azure url: alias: digicertglobalrootg2 `hopsworks.create_certificate.auto_generate` # { #helm.hopsworks.create_certificate.auto_generate } : Type `bool`, default `true`. `hopsworks.create_certificate.cn` # { #helm.hopsworks.create_certificate.cn } : Type `string`, default `"hopsworks.local.ai"`. `hopsworks.create_certificate.configmap` # { #helm.hopsworks.create_certificate.configmap } : Type `string`, default `"payara-cacerts"`. `hopsworks.create_certificate.enabled` # { #helm.hopsworks.create_certificate.enabled } : Type `bool`, default `false`. `hopsworks.create_certificate.secretName` # { #helm.hopsworks.create_certificate.secretName } : Type `string`, default `"hopsworks-tls"`. `hopsworks.create_certificate.trustManager.aliasPrefix` # { #helm.hopsworks.create_certificate.trustManager.aliasPrefix } : Type `string`, default `"trust-manager"`. `hopsworks.create_certificate.trustManager.configMapName` # { #helm.hopsworks.create_certificate.trustManager.configMapName } : Type `string`, default `""`. `hopsworks.create_certificate.trustManager.enabled` # { #helm.hopsworks.create_certificate.trustManager.enabled } : Type `bool`, default `false`. `hopsworks.create_certificate.trustManager.key` # { #helm.hopsworks.create_certificate.trustManager.key } : Type `string`, default `"ca-bundle.crt"`.
## dependencies { #helm-values-hopsworks-dependencies } ??? example "Defaults as YAML" ```yaml hopsworks: dependencies: logstash: consulServiceName: logstash port: 5044 mysql: consulServiceName: mysql consulServiceTag: onlinefs port: 3306 objectStorage: consulServiceName: minio port: 9000 registry: consulServiceName: registry port: 30443 ```
`hopsworks.dependencies.logstash.consulServiceName` # { #helm.hopsworks.dependencies.logstash.consulServiceName } : Type `string`, default `"logstash"`. `hopsworks.dependencies.logstash.port` # { #helm.hopsworks.dependencies.logstash.port } : Type `int`, default `5044`. `hopsworks.dependencies.mysql.consulServiceName` # { #helm.hopsworks.dependencies.mysql.consulServiceName } : Type `string`, default `"mysql"`. `hopsworks.dependencies.mysql.consulServiceTag` # { #helm.hopsworks.dependencies.mysql.consulServiceTag } : Type `string`, default `"onlinefs"`. `hopsworks.dependencies.mysql.port` # { #helm.hopsworks.dependencies.mysql.port } : Type `int`, default `3306`. `hopsworks.dependencies.objectStorage.consulServiceName` # { #helm.hopsworks.dependencies.objectStorage.consulServiceName } : Type `string`, default `"minio"`. `hopsworks.dependencies.objectStorage.port` # { #helm.hopsworks.dependencies.objectStorage.port } : Type `int`, default `9000`. `hopsworks.dependencies.registry.consulServiceName` # { #helm.hopsworks.dependencies.registry.consulServiceName } : Type `string`, default `"registry"`. `hopsworks.dependencies.registry.port` # { #helm.hopsworks.dependencies.registry.port } : Type `int`, default `30443`.
## dockerImage { #helm-values-hopsworks-dockerimage } ??? example "Defaults as YAML" ```yaml hopsworks: dockerImage: apt: sources: [] conda: default_mirrors: [] envs_dirs: - /envs pkgs_dirs: - /pkgs proxy: enabled: false protocol: http url: http://proxy:3128 repo_data_ttl: 43200 ssl_verify: 'True' use_defaults: true configMap: name: docker-images-config packageAuth: caCertsConfigMap: '' secretName: '' pypi: global_parameters: null pythonDownloads: never ```
`hopsworks.dockerImage.apt.sources` # { #helm.hopsworks.dockerImage.apt.sources } : Type `list`, default `[]`. The sources will be injected in case the user wants to use a proxy repo. `hopsworks.dockerImage.conda.default_mirrors` # { #helm.hopsworks.dockerImage.conda.default_mirrors } : Type `list`, default `[]`. the following mirrors will be injected in case the user wants to use a custom registry `hopsworks.dockerImage.conda.envs_dirs[0]` # { #helm.hopsworks.dockerImage.conda.envs_dirs.0 } : Type `string`, default `"/envs"`. `hopsworks.dockerImage.conda.pkgs_dirs[0]` # { #helm.hopsworks.dockerImage.conda.pkgs_dirs.0 } : Type `string`, default `"/pkgs"`. `hopsworks.dockerImage.conda.proxy.enabled` # { #helm.hopsworks.dockerImage.conda.proxy.enabled } : Type `bool`, default `false`. `hopsworks.dockerImage.conda.proxy.protocol` # { #helm.hopsworks.dockerImage.conda.proxy.protocol } : Type `string`, default `"http"`. `hopsworks.dockerImage.conda.proxy.url` # { #helm.hopsworks.dockerImage.conda.proxy.url } : Type `string`, default `"http://proxy:3128"`. `hopsworks.dockerImage.conda.repo_data_ttl` # { #helm.hopsworks.dockerImage.conda.repo_data_ttl } : Type `int`, default `43200`. `hopsworks.dockerImage.conda.ssl_verify` # { #helm.hopsworks.dockerImage.conda.ssl_verify } : Type `string`, default `"True"`. `hopsworks.dockerImage.conda.use_defaults` # { #helm.hopsworks.dockerImage.conda.use_defaults } : Type `bool`, default `true`. `hopsworks.dockerImage.configMap.name` # { #helm.hopsworks.dockerImage.configMap.name } : Type `string`, default `"docker-images-config"`. `hopsworks.dockerImage.packageAuth` # { #helm.hopsworks.dockerImage.packageAuth } : Type `object`, default `{"caCertsConfigMap":"","secretName":""}`. Optional Secret with credentials for authenticated pip/apt mirrors. The Secret must be created out-of-band in the Hopsworks release namespace and may contain any of the following keys (both optional): netrc -- mounted at $HOME/package-auth/netrc, staged as .netrc into env-build contexts (used by pip/conda). apt-auth.conf -- mounted at $HOME/package-auth/apt-auth.conf, staged as /etc/apt/auth.conf.d/hopsworks.conf in env-build contexts. When the Secret is absent the env-build proceeds against anonymous mirrors. `hopsworks.dockerImage.packageAuth.caCertsConfigMap` # { #helm.hopsworks.dockerImage.packageAuth.caCertsConfigMap } : Type `string`, default `""`. Optional ConfigMap of trusted CA certificates to install into env-build Docker RUN steps. Each key in the ConfigMap is treated as a .crt / PEM file, bind-mounted into /usr/local/share/ca-certificates/ and installed via `update-ca-certificates` at the start of the RUN. Also exported as REQUESTS_CA_BUNDLE / SSL_CERT_FILE so pip/conda honour it. Leave empty to skip CA injection. `hopsworks.dockerImage.packageAuth.secretName` # { #helm.hopsworks.dockerImage.packageAuth.secretName } : Type `string`, default `""`. Name of an optional Secret in the Hopsworks release namespace containing credentials for authenticated pip/apt mirrors. Leave empty to skip credential injection. `hopsworks.dockerImage.pypi` # { #helm.hopsworks.dockerImage.pypi } : Type `object`, default `{"global_parameters":null,"pythonDownloads":"never"}`. pypi configurations global_parameters: trusted-host: pypi.org index-url: "" extra-index-url: "" proxy: "" Written verbatim to pip.conf, and translated into uv's own format in uv.toml because uv reads neither pip.conf nor PIP_INDEX_URL. These keys translate: trusted-host, proxy, index-url, extra-index-url, find-links, no-index, no-cache-dir, cache-dir, keyring-provider, index-strategy, require-hashes, pre. Anything else reaches pip only, and is listed in a comment at the end of uv.toml. timeout, retries, cert and client-cert have no uv equivalent; add trust roots through dockerImage.caCertsConfigMap instead of cert, which uv does honour. index-strategy is uv-only: pip searches every index and takes the highest version, while uv stops at the first index carrying the package. Set it to unsafe-best-match for pip's behaviour, at the cost of uv's dependency-confusion protection. `hopsworks.dockerImage.pypi.global_parameters` # { #helm.hopsworks.dockerImage.pypi.global_parameters } : Type `string`, default `nil`. pypi global parameters `hopsworks.dockerImage.pypi.pythonDownloads` # { #helm.hopsworks.dockerImage.pypi.pythonDownloads } : Type `string`, default `"never"`. uv's python-downloads. "never" is the default and the air-gap-safe value: an environment build installs against an interpreter the base image already has, so a request uv cannot satisfy locally means a broken base image, and the alternative to failing is uv silently fetching a standalone interpreter from GitHub. Relax this only on a cluster that is meant to reach the internet and wants uv to manage interpreters.
## dockerRegistry { #helm-values-hopsworks-dockerregistry } ??? example "Defaults as YAML" ```yaml hopsworks: dockerRegistry: buildkit: image: moby/buildkit tag: v0.32.2 preset: affinity: {} alternativeRegistry: null configMapName: docker-images-preset-config enabled: true env: [] extra_images: [] maxPushRetries: 5 nodeSelector: {} parallelPushes: 1 parallelism: 3 resources: limits: cpu: '1' memory: 3G requests: cpu: 100m memory: 500Mi restartPolicy: OnFailure retry: 50 runIndex: 0 secrets: [] serviceAccount: annotations: {} timeout: 3600 tolerations: [] ttlSecondsAfterFinished: null usePullPush: true security: password: null trust_registry: true user: null ```
`hopsworks.dockerRegistry.buildkit.image` # { #helm.hopsworks.dockerRegistry.buildkit.image } : Type `string`, default `"moby/buildkit"`. `hopsworks.dockerRegistry.buildkit.tag` # { #helm.hopsworks.dockerRegistry.buildkit.tag } : Type `string`, default `"v0.32.2"`. Do not go below v0.31.2. Earlier releases carry published advisories, two of which matter more with the persistent daemon: a Git URL subdirectory path traversal, reachable because a library can be installed from a user-supplied Git URL, and a state-directory escape via a custom frontend, whose blast radius is every project once one daemon holds the state for all of them. Both fixed in v0.28.1; v0.31.2 also covers a Seccomp/AppArmor bypass, an unbounded-parsing DoS and a command injection through Git bundle checkout. A per-build daemon is affected the same way, so this is not a reason to leave the persistent one off. Most of what a scanner reports on this image is the four `buildkit-cni-*` plugins, which no BuildKit release refreshes and which never run here: the image ships no CNI config, so buildkitd falls back to host networking. The rootless variant is a separate image, used only when buildkitd.rootless.enabled is on. Both move together. Moving this pin means revisiting buildkitd.gc.keySyntax, which is version dependent. `hopsworks.dockerRegistry.preset.affinity` # { #helm.hopsworks.dockerRegistry.preset.affinity } : Type `object`, default `{}`. affinity configuration `hopsworks.dockerRegistry.preset.alternativeRegistry` # { #helm.hopsworks.dockerRegistry.preset.alternativeRegistry } : Type `string`, default `nil`. Alternative registry URL for base images. It will override any other global registry URL configuration. `hopsworks.dockerRegistry.preset.configMapName` # { #helm.hopsworks.dockerRegistry.preset.configMapName } : Type `string`, default `"docker-images-preset-config"`. `hopsworks.dockerRegistry.preset.enabled` # { #helm.hopsworks.dockerRegistry.preset.enabled } : Type `bool`, default `true`. `hopsworks.dockerRegistry.preset.env` # { #helm.hopsworks.dockerRegistry.preset.env } : Type `list`, default `[]`. Additional env vars to be set in the preset docker images (e.g. Proxy configuration) `hopsworks.dockerRegistry.preset.extra_images` # { #helm.hopsworks.dockerRegistry.preset.extra_images } : Type `list`, default `[]`. `hopsworks.dockerRegistry.preset.maxPushRetries` # { #helm.hopsworks.dockerRegistry.preset.maxPushRetries } : Type `int`, default `5`. `hopsworks.dockerRegistry.preset.nodeSelector` # { #helm.hopsworks.dockerRegistry.preset.nodeSelector } : Type `object`, default `{}`. node selector configuration `hopsworks.dockerRegistry.preset.parallelPushes` # { #helm.hopsworks.dockerRegistry.preset.parallelPushes } : Type `int`, default `1`. `hopsworks.dockerRegistry.preset.parallelism` # { #helm.hopsworks.dockerRegistry.preset.parallelism } : Type `int`, default `3`. `hopsworks.dockerRegistry.preset.resources.limits.cpu` # { #helm.hopsworks.dockerRegistry.preset.resources.limits.cpu } : Type `string`, default `"1"`. `hopsworks.dockerRegistry.preset.resources.limits.memory` # { #helm.hopsworks.dockerRegistry.preset.resources.limits.memory } : Type `string`, default `"3G"`. `hopsworks.dockerRegistry.preset.resources.requests.cpu` # { #helm.hopsworks.dockerRegistry.preset.resources.requests.cpu } : Type `string`, default `"100m"`. `hopsworks.dockerRegistry.preset.resources.requests.memory` # { #helm.hopsworks.dockerRegistry.preset.resources.requests.memory } : Type `string`, default `"500Mi"`. `hopsworks.dockerRegistry.preset.restartPolicy` # { #helm.hopsworks.dockerRegistry.preset.restartPolicy } : Type `string`, default `"OnFailure"`. `hopsworks.dockerRegistry.preset.retry` # { #helm.hopsworks.dockerRegistry.preset.retry } : Type `int`, default `50`. `hopsworks.dockerRegistry.preset.runIndex` # { #helm.hopsworks.dockerRegistry.preset.runIndex } : Type `int`, default `0`. Index to make job names unique if necessary `hopsworks.dockerRegistry.preset.secrets` # { #helm.hopsworks.dockerRegistry.preset.secrets } : Type `list`, default `[]`. `hopsworks.dockerRegistry.preset.serviceAccount.annotations` # { #helm.hopsworks.dockerRegistry.preset.serviceAccount.annotations } : Type `object`, default `{}`. service account annotations `hopsworks.dockerRegistry.preset.timeout` # { #helm.hopsworks.dockerRegistry.preset.timeout } : Type `int`, default `3600`. `hopsworks.dockerRegistry.preset.tolerations` # { #helm.hopsworks.dockerRegistry.preset.tolerations } : Type `list`, default `[]`. `hopsworks.dockerRegistry.preset.ttlSecondsAfterFinished` # { #helm.hopsworks.dockerRegistry.preset.ttlSecondsAfterFinished } : Type `string`, default `nil`. TTL in seconds for the preset-images Job. Overrides global default. `hopsworks.dockerRegistry.preset.usePullPush` # { #helm.hopsworks.dockerRegistry.preset.usePullPush } : Type `bool`, default `true`. `hopsworks.dockerRegistry.security.password` # { #helm.hopsworks.dockerRegistry.security.password } : Type `string`, default `nil`. user password `hopsworks.dockerRegistry.security.trust_registry` # { #helm.hopsworks.dockerRegistry.security.trust_registry } : Type `bool`, default `true`. `hopsworks.dockerRegistry.security.user` # { #helm.hopsworks.dockerRegistry.security.user } : Type `string`, default `nil`. user name
## envs { #helm-values-hopsworks-envs } ??? example "Defaults as YAML" ```yaml hopsworks: envs: admin: {} postInstall: ASADMIN_DEPLOY_TIMEOUT_SEC: 900 ASADMIN_TIMEOUT_SEC: 120 CLEANUP_INTERVAL_SEC: 120 DEBUG: 'true' DEPLOY_RETRIES: 3 DEPLOY_RETRY_INTERVAL_SEC: 30 INITIAL_DELAY_SEC: 60 REGISTER_WAIT_SEC: 240 ```
`hopsworks.envs.admin` # { #helm.hopsworks.envs.admin } : Type `object`, default `{}`. admin environment variables `hopsworks.envs.postInstall.ASADMIN_DEPLOY_TIMEOUT_SEC` # { #helm.hopsworks.envs.postInstall.ASADMIN_DEPLOY_TIMEOUT_SEC } : Type `int`, default `900`. Timeout (seconds) for the asadmin deploy command. Raise it if a legitimate EAR deploy exceeds 15 minutes, but keep query + deploy timeouts well under the sidecar's 30-minute heartbeat window. `hopsworks.envs.postInstall.ASADMIN_TIMEOUT_SEC` # { #helm.hopsworks.envs.postInstall.ASADMIN_TIMEOUT_SEC } : Type `int`, default `120`. Timeout (seconds) for the admin sidecar's short asadmin calls (list-*, delete-instance), so one stuck call cannot hang its loop. The long-running deploy uses ASADMIN_DEPLOY_TIMEOUT_SEC instead. `hopsworks.envs.postInstall.CLEANUP_INTERVAL_SEC` # { #helm.hopsworks.envs.postInstall.CLEANUP_INTERVAL_SEC } : Type `int`, default `120`. `hopsworks.envs.postInstall.DEBUG` # { #helm.hopsworks.envs.postInstall.DEBUG } : Type `string`, default `"true"`. `hopsworks.envs.postInstall.DEPLOY_RETRY_INTERVAL_SEC` # { #helm.hopsworks.envs.postInstall.DEPLOY_RETRY_INTERVAL_SEC } : Type `int`, default `30`. Sleep (seconds) between deploy attempts in the admin sidecar. Pacing only; the sidecar keeps retrying across cycles. `hopsworks.envs.postInstall.INITIAL_DELAY_SEC` # { #helm.hopsworks.envs.postInstall.INITIAL_DELAY_SEC } : Type `int`, default `60`. `hopsworks.envs.postInstall.REGISTER_WAIT_SEC` # { #helm.hopsworks.envs.postInstall.REGISTER_WAIT_SEC } : Type `int`, default `240`. `hopsworks.envs.postInstall.DEPLOY_RETRIES` Deprecated # { #helm.hopsworks.envs.postInstall.DEPLOY_RETRIES } : Type `int`, default `3`. Deprecated. No longer consumed by the admin sidecar. Use DEPLOY_RETRY_INTERVAL_SEC to control the sleep between deploy attempts. Kept for backward compatibility; will be removed in a future major version.
## hopsworksAdmin { #helm-values-hopsworks-hopsworksadmin } ??? example "Defaults as YAML" ```yaml hopsworks: hopsworksAdmin: hopsworks_ear_download_url: null hopsworks_front_download_url: null hopsworks_realm_download_url: null mysql_connector_download_url: null ```
`hopsworks.hopsworksAdmin` # { #helm.hopsworks.hopsworksAdmin } : Type `object`. override hopsworks admin artifacts ??? note "Default" ```yaml hopsworks_ear_download_url: null hopsworks_front_download_url: null hopsworks_realm_download_url: null mysql_connector_download_url: null ``` `hopsworks.hopsworksAdmin.hopsworks_ear_download_url` # { #helm.hopsworks.hopsworksAdmin.hopsworks_ear_download_url } : Type `string`, default `nil`. use custom hopsworks ear `hopsworks.hopsworksAdmin.hopsworks_front_download_url` # { #helm.hopsworks.hopsworksAdmin.hopsworks_front_download_url } : Type `string`, default `nil`. use custom hopsworks front end `hopsworks.hopsworksAdmin.hopsworks_realm_download_url` # { #helm.hopsworks.hopsworksAdmin.hopsworks_realm_download_url } : Type `string`, default `nil`. use custom hopsworks realm jar file `hopsworks.hopsworksAdmin.mysql_connector_download_url` # { #helm.hopsworks.hopsworksAdmin.mysql_connector_download_url } : Type `string`, default `nil`. use custom mysql connector
## hopsworksCA { #helm-values-hopsworks-hopsworksca } ??? example "Defaults as YAML" ```yaml hopsworks: hopsworksCA: affinity: podAntiAffinity: requiredDuringSchedulingIgnoredDuringExecution: - labelSelector: matchExpressions: - key: app operator: In values: - hopsworks-ca topologyKey: kubernetes.io/hostname apiKey: secret_name: hopsworks-api-key-auth auto_jvm: true ca_count: 1 containerPort: 8182 download_url: null extraJavaToolOptions: '' image: pullPolicy: IfNotPresent tag: null internalCert: subject: locality: glassfishinternal organization: 0 jvm: garbageCollector: '' memory: buffer: 256 compressedClassSpaceSize: 256 heap: 1024 metaspace: 1024 nonMethodCodeHeapSize: 5 nonProfiledCodeHeapSize: 48 profiledCodeHeapSize: 48 livenessProbe: failureThreshold: 3 httpGet: path: /hopsworks-ca/v2/certificate/crl/intermediate port: 8182 scheme: HTTPS initialDelaySeconds: 600 periodSeconds: 20 timeoutSeconds: 60 name: hopsworks-ca readinessProbe: httpGet: path: /hopsworks-ca/v2/certificate/ready port: 8182 scheme: HTTPS initialDelaySeconds: 60 periodSeconds: 10 replaceEntryPoint: false resources: limits: cpu: 2000m memory: 2048Mi requests: cpu: 1000m memory: 2048Mi securityContext: {} service: annotations: consul.hashicorp.com/service-name: glassfish consul.hashicorp.com/service-tags: ca name: hopsworks-ca port: 8182 setupJob: backoffLimit: 10 shutdownWait: 30 startWaitTimeout: 300 ```
`hopsworks.hopsworksCA.affinity` # { #helm.hopsworks.hopsworksCA.affinity } : Type `object`. Ensure we send them to different machines ??? note "Default" ```yaml podAntiAffinity: requiredDuringSchedulingIgnoredDuringExecution: - labelSelector: matchExpressions: - key: app operator: In values: - hopsworks-ca topologyKey: kubernetes.io/hostname ``` `hopsworks.hopsworksCA.affinity.podAntiAffinity.requiredDuringSchedulingIgnoredDuringExecution` # { #helm.hopsworks.hopsworksCA.affinity.podAntiAffinity.requiredDuringSchedulingIgnoredDuringExecution } : Type `list`. podAntiAffinity.requiredDuringSchedulingIgnoredDuringExecution configuration ??? note "Default" ```yaml - labelSelector: matchExpressions: - key: app operator: In values: - hopsworks-ca topologyKey: kubernetes.io/hostname ``` `hopsworks.hopsworksCA.apiKey.secret_name` # { #helm.hopsworks.hopsworksCA.apiKey.secret_name } : Type `string`, default `"hopsworks-api-key-auth"`. `hopsworks.hopsworksCA.auto_jvm` # { #helm.hopsworks.hopsworksCA.auto_jvm } : Type `bool`, default `true`. `hopsworks.hopsworksCA.ca_count` # { #helm.hopsworks.hopsworksCA.ca_count } : Type `int`, default `1`. Number of hopsworks-ca replicas. Not rendered while hpa.ca.enabled is true; same rule as replicaCount.worker. `hopsworks.hopsworksCA.containerPort` # { #helm.hopsworks.hopsworksCA.containerPort } : Type `int`, default `8182`. `hopsworks.hopsworksCA.download_url` # { #helm.hopsworks.hopsworksCA.download_url } : Type `string`, default `nil`. use custom hopsworks ca war `hopsworks.hopsworksCA.extraJavaToolOptions` # { #helm.hopsworks.hopsworksCA.extraJavaToolOptions } : Type `string`, default `""`. `hopsworks.hopsworksCA.image.pullPolicy` # { #helm.hopsworks.hopsworksCA.image.pullPolicy } : Type `string`, default `"IfNotPresent"`. `hopsworks.hopsworksCA.image.tag` # { #helm.hopsworks.hopsworksCA.image.tag } : Type `string`, default `nil`. image tag. If not defined, the .Chart.AppVersion will be used `hopsworks.hopsworksCA.internalCert.subject.locality` # { #helm.hopsworks.hopsworksCA.internalCert.subject.locality } : Type `string`, default `"glassfishinternal"`. `hopsworks.hopsworksCA.internalCert.subject.organization` # { #helm.hopsworks.hopsworksCA.internalCert.subject.organization } : Type `int`, default `0`. `hopsworks.hopsworksCA.jvm.garbageCollector` # { #helm.hopsworks.hopsworksCA.jvm.garbageCollector } : Type `string`, default `""`. `hopsworks.hopsworksCA.jvm.memory.buffer` # { #helm.hopsworks.hopsworksCA.jvm.memory.buffer } : Type `int`, default `256`. `hopsworks.hopsworksCA.jvm.memory.compressedClassSpaceSize` # { #helm.hopsworks.hopsworksCA.jvm.memory.compressedClassSpaceSize } : Type `int`, default `256`. `hopsworks.hopsworksCA.jvm.memory.heap` # { #helm.hopsworks.hopsworksCA.jvm.memory.heap } : Type `int`, default `1024`. `hopsworks.hopsworksCA.jvm.memory.metaspace` # { #helm.hopsworks.hopsworksCA.jvm.memory.metaspace } : Type `int`, default `1024`. `hopsworks.hopsworksCA.jvm.memory.nonMethodCodeHeapSize` # { #helm.hopsworks.hopsworksCA.jvm.memory.nonMethodCodeHeapSize } : Type `int`, default `5`. `hopsworks.hopsworksCA.jvm.memory.nonProfiledCodeHeapSize` # { #helm.hopsworks.hopsworksCA.jvm.memory.nonProfiledCodeHeapSize } : Type `int`, default `48`. `hopsworks.hopsworksCA.jvm.memory.profiledCodeHeapSize` # { #helm.hopsworks.hopsworksCA.jvm.memory.profiledCodeHeapSize } : Type `int`, default `48`. `hopsworks.hopsworksCA.livenessProbe.failureThreshold` # { #helm.hopsworks.hopsworksCA.livenessProbe.failureThreshold } : Type `int`, default `3`. `hopsworks.hopsworksCA.livenessProbe.httpGet.path` # { #helm.hopsworks.hopsworksCA.livenessProbe.httpGet.path } : Type `string`, default `"/hopsworks-ca/v2/certificate/crl/intermediate"`. `hopsworks.hopsworksCA.livenessProbe.httpGet.port` # { #helm.hopsworks.hopsworksCA.livenessProbe.httpGet.port } : Type `int`, default `8182`. `hopsworks.hopsworksCA.livenessProbe.httpGet.scheme` # { #helm.hopsworks.hopsworksCA.livenessProbe.httpGet.scheme } : Type `string`, default `"HTTPS"`. `hopsworks.hopsworksCA.livenessProbe.initialDelaySeconds` # { #helm.hopsworks.hopsworksCA.livenessProbe.initialDelaySeconds } : Type `int`, default `600`. `hopsworks.hopsworksCA.livenessProbe.periodSeconds` # { #helm.hopsworks.hopsworksCA.livenessProbe.periodSeconds } : Type `int`, default `20`. `hopsworks.hopsworksCA.livenessProbe.timeoutSeconds` # { #helm.hopsworks.hopsworksCA.livenessProbe.timeoutSeconds } : Type `int`, default `60`. `hopsworks.hopsworksCA.name` # { #helm.hopsworks.hopsworksCA.name } : Type `string`, default `"hopsworks-ca"`. `hopsworks.hopsworksCA.readinessProbe.httpGet.path` # { #helm.hopsworks.hopsworksCA.readinessProbe.httpGet.path } : Type `string`, default `"/hopsworks-ca/v2/certificate/ready"`. `hopsworks.hopsworksCA.readinessProbe.httpGet.port` # { #helm.hopsworks.hopsworksCA.readinessProbe.httpGet.port } : Type `int`, default `8182`. `hopsworks.hopsworksCA.readinessProbe.httpGet.scheme` # { #helm.hopsworks.hopsworksCA.readinessProbe.httpGet.scheme } : Type `string`, default `"HTTPS"`. `hopsworks.hopsworksCA.readinessProbe.initialDelaySeconds` # { #helm.hopsworks.hopsworksCA.readinessProbe.initialDelaySeconds } : Type `int`, default `60`. `hopsworks.hopsworksCA.readinessProbe.periodSeconds` # { #helm.hopsworks.hopsworksCA.readinessProbe.periodSeconds } : Type `int`, default `10`. `hopsworks.hopsworksCA.replaceEntryPoint` # { #helm.hopsworks.hopsworksCA.replaceEntryPoint } : Type `bool`, default `false`. `hopsworks.hopsworksCA.resources.limits.cpu` # { #helm.hopsworks.hopsworksCA.resources.limits.cpu } : Type `string`, default `"2000m"`. `hopsworks.hopsworksCA.resources.limits.memory` # { #helm.hopsworks.hopsworksCA.resources.limits.memory } : Type `string`, default `"2048Mi"`. `hopsworks.hopsworksCA.resources.requests.cpu` # { #helm.hopsworks.hopsworksCA.resources.requests.cpu } : Type `string`, default `"1000m"`. `hopsworks.hopsworksCA.resources.requests.memory` # { #helm.hopsworks.hopsworksCA.resources.requests.memory } : Type `string`, default `"2048Mi"`. `hopsworks.hopsworksCA.securityContext` # { #helm.hopsworks.hopsworksCA.securityContext } : Type `object`, default `{}`. security context `hopsworks.hopsworksCA.service.annotations."consul.hashicorp.com/service-name"` # { #helm.hopsworks.hopsworksCA.service.annotations.consul.hashicorp.com-service-name } : Type `string`, default `"glassfish"`. `hopsworks.hopsworksCA.service.annotations."consul.hashicorp.com/service-tags"` # { #helm.hopsworks.hopsworksCA.service.annotations.consul.hashicorp.com-service-tags } : Type `string`, default `"ca"`. `hopsworks.hopsworksCA.service.name` # { #helm.hopsworks.hopsworksCA.service.name } : Type `string`, default `"hopsworks-ca"`. `hopsworks.hopsworksCA.service.port` # { #helm.hopsworks.hopsworksCA.service.port } : Type `int`, default `8182`. `hopsworks.hopsworksCA.setupJob.backoffLimit` # { #helm.hopsworks.hopsworksCA.setupJob.backoffLimit } : Type `int`, default `10`. `hopsworks.hopsworksCA.shutdownWait` # { #helm.hopsworks.hopsworksCA.shutdownWait } : Type `int`, default `30`. `hopsworks.hopsworksCA.startWaitTimeout` # { #helm.hopsworks.hopsworksCA.startWaitTimeout } : Type `int`, default `300`.
## hpa { #helm-values-hopsworks-hpa } ??? example "Defaults as YAML" ```yaml hopsworks: hpa: ca: enabled: false maxReplicas: 3 targetCPUUtilizationPercentage: 80 targetMemoryUtilizationPercentage: 95 worker: enabled: false maxReplicas: 3 targetCPUUtilizationPercentage: 80 targetMemoryUtilizationPercentage: 95 ```
`hopsworks.hpa.ca.enabled` # { #helm.hopsworks.hpa.ca.enabled } : Type `bool`, default `false`. Autoscale hopsworks-ca. While enabled, hopsworksCA.ca_count is not rendered and the HPA owns spec.replicas. `hopsworks.hpa.ca.maxReplicas` # { #helm.hopsworks.hpa.ca.maxReplicas } : Type `int`, default `3`. `hopsworks.hpa.ca.targetCPUUtilizationPercentage` # { #helm.hopsworks.hpa.ca.targetCPUUtilizationPercentage } : Type `int`, default `80`. `hopsworks.hpa.ca.targetMemoryUtilizationPercentage` # { #helm.hopsworks.hpa.ca.targetMemoryUtilizationPercentage } : Type `int`, default `95`. `hopsworks.hpa.worker.enabled` # { #helm.hopsworks.hpa.worker.enabled } : Type `bool`, default `false`. Autoscale the Payara workers. While enabled, replicaCount.worker is not rendered and the HPA owns spec.replicas. `hopsworks.hpa.worker.maxReplicas` # { #helm.hopsworks.hpa.worker.maxReplicas } : Type `int`, default `3`. `hopsworks.hpa.worker.targetCPUUtilizationPercentage` # { #helm.hopsworks.hpa.worker.targetCPUUtilizationPercentage } : Type `int`, default `80`. `hopsworks.hpa.worker.targetMemoryUtilizationPercentage` # { #helm.hopsworks.hpa.worker.targetMemoryUtilizationPercentage } : Type `int`, default `95`.
## image { #helm-values-hopsworks-image } ??? example "Defaults as YAML" ```yaml hopsworks: image: admin: imageName: hopsworks mysqlConnectorVersion: 8.0.21.1 mysqlStorageImageName: hopsworks-mysql-connector pullPolicy: IfNotPresent tag: null filebeat: imageName: filebeat tag: 8.19.21 migration: pullPolicy: IfNotPresent tag: null postInstall: pullPolicy: IfNotPresent tag: 6.2025.11-jdk21.0 registry: null worker: pullPolicy: IfNotPresent tag: 6.2025.11-jdk21.0 ```
`hopsworks.image.admin.imageName` # { #helm.hopsworks.image.admin.imageName } : Type `string`, default `"hopsworks"`. `hopsworks.image.admin.mysqlConnectorVersion` # { #helm.hopsworks.image.admin.mysqlConnectorVersion } : Type `string`, default `"8.0.21.1"`. `hopsworks.image.admin.mysqlStorageImageName` # { #helm.hopsworks.image.admin.mysqlStorageImageName } : Type `string`, default `"hopsworks-mysql-connector"`. `hopsworks.image.admin.pullPolicy` # { #helm.hopsworks.image.admin.pullPolicy } : Type `string`, default `"IfNotPresent"`. `hopsworks.image.admin.tag` # { #helm.hopsworks.image.admin.tag } : Type `string`, default `nil`. image tag. If not defined, the .Chart.AppVersion will be used `hopsworks.image.filebeat.imageName` # { #helm.hopsworks.image.filebeat.imageName } : Type `string`, default `"filebeat"`. `hopsworks.image.filebeat.tag` # { #helm.hopsworks.image.filebeat.tag } : Type `string`, default `"8.19.21"`. `hopsworks.image.migration.pullPolicy` # { #helm.hopsworks.image.migration.pullPolicy } : Type `string`, default `"IfNotPresent"`. `hopsworks.image.migration.tag` # { #helm.hopsworks.image.migration.tag } : Type `string`, default `nil`. image tag. If not defined, the .Chart.AppVersion will be used `hopsworks.image.postInstall.pullPolicy` # { #helm.hopsworks.image.postInstall.pullPolicy } : Type `string`, default `"IfNotPresent"`. pull policy for the admin sidecar's payara-node image. admindeployment.yaml already read this key before it was declared here, and its fallback cannot reach global._hopsworks.imagePullPolicy, so the sidecar was pinned to IfNotPresent with no way to override it. `hopsworks.image.postInstall.tag` # { #helm.hopsworks.image.postInstall.tag } : Type `string`, default `"6.2025.11-jdk21.0"`. `hopsworks.image.registry` # { #helm.hopsworks.image.registry } : Type `string`, default `nil`. image registry. If not defined, the global._hopsworks.imageRegistry will be used instead `hopsworks.image.worker.pullPolicy` # { #helm.hopsworks.image.worker.pullPolicy } : Type `string`, default `"IfNotPresent"`. `hopsworks.image.worker.tag` # { #helm.hopsworks.image.worker.tag } : Type `string`, default `"6.2025.11-jdk21.0"`.
## ingress { #helm-values-hopsworks-ingress } ??? example "Defaults as YAML" ```yaml hopsworks: ingress: annotations: nginx.ingress.kubernetes.io/affinity: cookie nginx.ingress.kubernetes.io/affinity-mode: persistent nginx.ingress.kubernetes.io/proxy-body-size: '0' nginx.ingress.kubernetes.io/proxy-redirect-from: 'http:' nginx.ingress.kubernetes.io/proxy-redirect-to: 'https:' nginx.ingress.kubernetes.io/session-cookie-expires: '5259600' nginx.ingress.kubernetes.io/session-cookie-max-age: '5259600' nginx.ingress.kubernetes.io/ssl-redirect: 'true' enabled: true extraLabels: {} extraPaths: [] host: hopsworks.ai.local hosts: - hopsworks.ai.local ingressClassName: nginx path: / pathType: Prefix secretName: hopsworks-ingress-crypto-material servicePort: 28080 sslPassthrough: false tls: - hosts: - hopsworks.ai.local secretName: hopsworks-ingress-crypto-material ```
`hopsworks.ingress.annotations` # { #helm.hopsworks.ingress.annotations } : Type `object`. ingress annotations ??? note "Default" ```yaml nginx.ingress.kubernetes.io/affinity: cookie nginx.ingress.kubernetes.io/affinity-mode: persistent nginx.ingress.kubernetes.io/proxy-body-size: '0' nginx.ingress.kubernetes.io/proxy-redirect-from: 'http:' nginx.ingress.kubernetes.io/proxy-redirect-to: 'https:' nginx.ingress.kubernetes.io/session-cookie-expires: '5259600' nginx.ingress.kubernetes.io/session-cookie-max-age: '5259600' nginx.ingress.kubernetes.io/ssl-redirect: 'true' ``` `hopsworks.ingress.enabled` # { #helm.hopsworks.ingress.enabled } : Type `bool`, default `true`. `hopsworks.ingress.extraLabels` # { #helm.hopsworks.ingress.extraLabels } : Type `object`, default `{}`. ingress extra labels `hopsworks.ingress.extraPaths` # { #helm.hopsworks.ingress.extraPaths } : Type `list`, default `[]`. `hopsworks.ingress.host` # { #helm.hopsworks.ingress.host } : Type `string`, default `"hopsworks.ai.local"`. `hopsworks.ingress.hosts` # { #helm.hopsworks.ingress.hosts } : Type `list`, default `["hopsworks.ai.local"]`. ingress hosts configuration `hopsworks.ingress.ingressClassName` # { #helm.hopsworks.ingress.ingressClassName } : Type `string`, default `"nginx"`. `hopsworks.ingress.path` # { #helm.hopsworks.ingress.path } : Type `string`, default `"/"`. `hopsworks.ingress.pathType` # { #helm.hopsworks.ingress.pathType } : Type `string`, default `"Prefix"`. `hopsworks.ingress.secretName` # { #helm.hopsworks.ingress.secretName } : Type `string`, default `"hopsworks-ingress-crypto-material"`. `hopsworks.ingress.servicePort` # { #helm.hopsworks.ingress.servicePort } : Type `int`, default `28080`. `hopsworks.ingress.sslPassthrough` # { #helm.hopsworks.ingress.sslPassthrough } : Type `bool`, default `false`. `hopsworks.ingress.tls` # { #helm.hopsworks.ingress.tls } : Type `list`. ingress tls configuration per host ??? note "Default" ```yaml - hosts: - hopsworks.ai.local secretName: hopsworks-ingress-crypto-material ```
## payara { #helm-values-hopsworks-payara } ??? example "Defaults as YAML" ```yaml hopsworks: payara: adminpwKeyName: admin_password adminuser: admin config: hopsworks-config debug: false deploymentGroup: hopsworks-dg disableMetricsServiceLogs: false disableXmlValidation: true encryptionpwKeyName: encryption_master_password http: keep_alive_timeout: 30 httpListener1Enabled: true httpthreadpool: idletimeout: 900 maxqueuesize: 4096 maxthreadpoolsize: 200 minthreadpoolsize: 5 kerberos: enabled: false keyTabName: service.keytab keyTabPath: /etc/security/keytabs keyTabPrincipal: HTTP/hopsworks.cluster.local@EXAMPLE.COM keyTabSecretName: keytab-secret krb5ConfigName: server-krb5-config ldap: enabled: false factory_class: com.sun.jndi.ldap.LdapCtxFactory jndilookupname: dc=example,dc=com property: additional_props: - key: hopsworks\.ldap\.basedn value: dc=example,dc=com attributes_binary: entryUUID provider_url: ldap://192.168.200.104:389 referral: ignore security: authentication: simple credentials: secret_key: credentials secret_name: ldap-credentials-secret principal: cn=admin,dc=example,dc=com res_type: javax.naming.ldap.LdapContext mail: email: smtp@gmail.com from: admin@hopsworks.ai password: secret_key: payara-mail-password smtp: smtp.gmail.com smtp_port: '587' smtp_ssl_port: '465' oauth: clients: [] enabled: false postbootCommands: /opt/payara/k8s/commands/post-boot-commands.asadmin prebootCommands: /opt/payara/k8s/commands/pre-boot-commands.asadmin uniformLogFormatter: false versionUpgrade: null websocketProxy: grizzlyWorkerPoolMaxSize: 200 heartbeatIntervalMs: 20000 incomingBufferBytes: 33554432 maxSessionsPerApp: 500 sessionIdleTimeoutMs: 0 ```
`hopsworks.payara.adminpwKeyName` # { #helm.hopsworks.payara.adminpwKeyName } : Type `string`, default `"admin_password"`. `hopsworks.payara.adminuser` # { #helm.hopsworks.payara.adminuser } : Type `string`, default `"admin"`. `hopsworks.payara.config` # { #helm.hopsworks.payara.config } : Type `string`, default `"hopsworks-config"`. `hopsworks.payara.debug` # { #helm.hopsworks.payara.debug } : Type `bool`, default `false`. `hopsworks.payara.deploymentGroup` # { #helm.hopsworks.payara.deploymentGroup } : Type `string`, default `"hopsworks-dg"`. `hopsworks.payara.disableMetricsServiceLogs` # { #helm.hopsworks.payara.disableMetricsServiceLogs } : Type `bool`, default `false`. `hopsworks.payara.disableXmlValidation` # { #helm.hopsworks.payara.disableXmlValidation } : Type `bool`, default `true`. `hopsworks.payara.encryptionpwKeyName` # { #helm.hopsworks.payara.encryptionpwKeyName } : Type `string`, default `"encryption_master_password"`. `hopsworks.payara.http.keep_alive_timeout` # { #helm.hopsworks.payara.http.keep_alive_timeout } : Type `int`, default `30`. `hopsworks.payara.httpListener1Enabled` # { #helm.hopsworks.payara.httpListener1Enabled } : Type `bool`, default `true`. `hopsworks.payara.httpthreadpool.idletimeout` # { #helm.hopsworks.payara.httpthreadpool.idletimeout } : Type `int`, default `900`. The maximum amount of time that a thread can remain idle in the pool. After this time expires, the thread is removed from the pool. `hopsworks.payara.httpthreadpool.maxqueuesize` # { #helm.hopsworks.payara.httpthreadpool.maxqueuesize } : Type `int`, default `4096`. The maximum number of threads in the queue. A value of -1 indicates that there is no limit to the queue size. `hopsworks.payara.httpthreadpool.maxthreadpoolsize` # { #helm.hopsworks.payara.httpthreadpool.maxthreadpoolsize } : Type `int`, default `200`. `hopsworks.payara.httpthreadpool.minthreadpoolsize` # { #helm.hopsworks.payara.httpthreadpool.minthreadpoolsize } : Type `int`, default `5`. `hopsworks.payara.kerberos.enabled` # { #helm.hopsworks.payara.kerberos.enabled } : Type `bool`, default `false`. `hopsworks.payara.kerberos.keyTabName` # { #helm.hopsworks.payara.kerberos.keyTabName } : Type `string`, default `"service.keytab"`. kerberos service principal keytab name. `hopsworks.payara.kerberos.keyTabPath` # { #helm.hopsworks.payara.kerberos.keyTabPath } : Type `string`, default `"/etc/security/keytabs"`. kerberos service principal keytab file path. `hopsworks.payara.kerberos.keyTabPrincipal` # { #helm.hopsworks.payara.kerberos.keyTabPrincipal } : Type `string`, default `"HTTP/hopsworks.cluster.local@EXAMPLE.COM"`. kerberos service principal name. This name is created by combining the string HTTP with the hostname 'HTTP@host_name'. The host name is the DNS name by which browsers contact the Web server. Use the fully qualified host name. `hopsworks.payara.kerberos.keyTabSecretName` # { #helm.hopsworks.payara.kerberos.keyTabSecretName } : Type `string`, default `"keytab-secret"`. keytab secret name. `hopsworks.payara.kerberos.krb5ConfigName` # { #helm.hopsworks.payara.kerberos.krb5ConfigName } : Type `string`, default `"server-krb5-config"`. krb5.conf config map name. `hopsworks.payara.ldap.enabled` # { #helm.hopsworks.payara.ldap.enabled } : Type `bool`, default `false`. `hopsworks.payara.ldap.factory_class` # { #helm.hopsworks.payara.ldap.factory_class } : Type `string`, default `"com.sun.jndi.ldap.LdapCtxFactory"`. Factory class for resource; implements javax.naming.spi.ObjectFactory. `hopsworks.payara.ldap.jndilookupname` # { #helm.hopsworks.payara.ldap.jndilookupname } : Type `string`, default `"dc=example,dc=com"`. Name used by the application to find the resource. `hopsworks.payara.ldap.property.additional_props[0].key` # { #helm.hopsworks.payara.ldap.property.additional_props.0.key } : Type `string`, default `"hopsworks\\.ldap\\.basedn"`. `hopsworks.payara.ldap.property.additional_props[0].value` # { #helm.hopsworks.payara.ldap.property.additional_props.0.value } : Type `string`, default `"dc=example,dc=com"`. `hopsworks.payara.ldap.property.attributes_binary` # { #helm.hopsworks.payara.ldap.property.attributes_binary } : Type `string`, default `"entryUUID"`. The binary unique identifier that will be used in subsequent logins to identify the user. `hopsworks.payara.ldap.property.provider_url` # { #helm.hopsworks.payara.ldap.property.provider_url } : Type `string`, default `"ldap://192.168.200.104:389"`. `hopsworks.payara.ldap.property.referral` # { #helm.hopsworks.payara.ldap.property.referral } : Type `string`, default `"ignore"`. Whether to follow or ignore an alternate location in which an LDAP request may be processed. `hopsworks.payara.ldap.property.security.authentication` # { #helm.hopsworks.payara.ldap.property.security.authentication } : Type `string`, default `"simple"`. `hopsworks.payara.ldap.property.security.credentials.secret_key` # { #helm.hopsworks.payara.ldap.property.security.credentials.secret_key } : Type `string`, default `"credentials"`. `hopsworks.payara.ldap.property.security.credentials.secret_name` # { #helm.hopsworks.payara.ldap.property.security.credentials.secret_name } : Type `string`, default `"ldap-credentials-secret"`. `hopsworks.payara.ldap.property.security.principal` # { #helm.hopsworks.payara.ldap.property.security.principal } : Type `string`, default `"cn=admin,dc=example,dc=com"`. `hopsworks.payara.ldap.res_type` # { #helm.hopsworks.payara.ldap.res_type } : Type `string`, default `"javax.naming.ldap.LdapContext"`. Resource Type. Enter a fully qualified type following the format xxx.xxx (for example, javax.jms.Topic) `hopsworks.payara.mail.email` # { #helm.hopsworks.payara.mail.email } : Type `string`, default `"smtp@gmail.com"`. `hopsworks.payara.mail.from` # { #helm.hopsworks.payara.mail.from } : Type `string`, default `"admin@hopsworks.ai"`. `hopsworks.payara.mail.password.secret_key` # { #helm.hopsworks.payara.mail.password.secret_key } : Type `string`, default `"payara-mail-password"`. `hopsworks.payara.mail.smtp` # { #helm.hopsworks.payara.mail.smtp } : Type `string`, default `"smtp.gmail.com"`. `hopsworks.payara.mail.smtp_port` # { #helm.hopsworks.payara.mail.smtp_port } : Type `string`, default `"587"`. `hopsworks.payara.mail.smtp_ssl_port` # { #helm.hopsworks.payara.mail.smtp_ssl_port } : Type `string`, default `"465"`. `hopsworks.payara.oauth.clients` # { #helm.hopsworks.payara.oauth.clients } : Type `list`, default `[]`. oauth clients `hopsworks.payara.oauth.enabled` # { #helm.hopsworks.payara.oauth.enabled } : Type `bool`, default `false`. `hopsworks.payara.postbootCommands` # { #helm.hopsworks.payara.postbootCommands } : Type `string`, default `"/opt/payara/k8s/commands/post-boot-commands.asadmin"`. `hopsworks.payara.prebootCommands` # { #helm.hopsworks.payara.prebootCommands } : Type `string`, default `"/opt/payara/k8s/commands/pre-boot-commands.asadmin"`. `hopsworks.payara.uniformLogFormatter` # { #helm.hopsworks.payara.uniformLogFormatter } : Type `bool`, default `false`. `hopsworks.payara.versionUpgrade` # { #helm.hopsworks.payara.versionUpgrade } : Type `string`, default `nil`. is Payara version upgrade. If null, auto detect based on current and target version `hopsworks.payara.websocketProxy` # { #helm.hopsworks.payara.websocketProxy } : Type `object`. Tyrus WebSocket-proxy tuning (jupyter / terminal / python-app). These five values are the single source of truth: post-boot-commands.txt passes each as a JVM system property (-D) which hopsworks-ee reads (WebSocketProxyConfig). ??? note "Default" ```yaml grizzlyWorkerPoolMaxSize: 200 heartbeatIntervalMs: 20000 incomingBufferBytes: 33554432 maxSessionsPerApp: 500 sessionIdleTimeoutMs: 0 ``` `hopsworks.payara.websocketProxy.grizzlyWorkerPoolMaxSize` # { #helm.hopsworks.payara.websocketProxy.grizzlyWorkerPoolMaxSize } : Type `int`, default `200`. Upper bound on the shared Tyrus/Grizzly client transport worker pool. Bounded to 1..Integer.MAX_VALUE (read via Integer.getInteger). `hopsworks.payara.websocketProxy.heartbeatIntervalMs` # { #helm.hopsworks.payara.websocketProxy.heartbeatIntervalMs } : Type `int`, default `20000`. Interval, in milliseconds, of the Tyrus per-session heartbeat on the inbound (browser) WebSocket leg. The browser hop traverses ingress-nginx, whose proxy_read_timeout (default 60s) reaps a WebSocket idle for that long; Tyrus emits an unsolicited pong every interval to keep it warm. Keep below the ingress timeout. 0 disables. Read via Long.getLong, so only a non-negative lower bound is enforced. `hopsworks.payara.websocketProxy.incomingBufferBytes` # { #helm.hopsworks.payara.websocketProxy.incomingBufferBytes } : Type `int`, default `33554432`. Max size, in bytes, of a single received WebSocket frame on the upstream leg — the ceiling a large Jupyter cell output must fit under. Grown on demand, not pre-allocated. Keep at or above jupyter's iopub rate-limit budget. Bounded to 1..Integer.MAX_VALUE (read via Integer.getInteger). `hopsworks.payara.websocketProxy.maxSessionsPerApp` # { #helm.hopsworks.payara.websocketProxy.maxSessionsPerApp } : Type `int`, default `500`. Max concurrent inbound WebSocket proxy sessions per pod (jupyter, terminal and python-app share this budget). The next upgrade past the cap is closed with 1013 TRY_AGAIN_LATER, protecting the pod from connection-driven OOM. Read via Integer.getInteger in hopsworks-ee, so values outside the int range would be silently ignored — bounded to 1..Integer.MAX_VALUE here. `hopsworks.payara.websocketProxy.sessionIdleTimeoutMs` # { #helm.hopsworks.payara.websocketProxy.sessionIdleTimeoutMs } : Type `int`, default `0`. Per-session idle timeout in milliseconds. 0 disables the idle reaper (proxy sessions are legitimately long-lived). Read via Long.getLong, so only a non-negative lower bound is enforced.
## probs { #helm-values-hopsworks-probs } ??? example "Defaults as YAML" ```yaml hopsworks: probs: admin: livenessProbe: exec: command: - /bin/sh - /opt/payara/k8s/ready.sh failureThreshold: 10 periodSeconds: 60 timeoutSeconds: 30 readinessProbe: exec: command: - /bin/sh - /opt/payara/k8s/ready.sh initialDelaySeconds: 90 periodSeconds: 10 timeoutSeconds: 30 startupProbe: exec: command: - /bin/sh - /opt/payara/k8s/ready.sh failureThreshold: 180 periodSeconds: 10 timeoutSeconds: 30 worker: livenessProbe: failureThreshold: 3 httpGet: path: /hopsworks-api/api/variables/versions port: 8182 scheme: HTTPS initialDelaySeconds: 150 periodSeconds: 20 timeoutSeconds: 20 readinessProbe: httpGet: path: /hopsworks-api/api/variables/versions port: 8182 scheme: HTTPS initialDelaySeconds: 150 periodSeconds: 10 startupProbe: failureThreshold: 10 httpGet: path: /health port: 8182 scheme: HTTPS initialDelaySeconds: 150 periodSeconds: 10 ```
`hopsworks.probs.admin.livenessProbe` # { #helm.hopsworks.probs.admin.livenessProbe } : Type `object`. liveness probe. Restarts the DAS if it stays unresponsive for ~10 minutes (10 x 60s); without it a hung DAS is never restarted. The long window is deliberate: it must not fire on slowness, and each run costs an asadmin JVM inside the DAS container. Set to null to disable. If you override with httpGet/tcpSocket, also set `exec: null` (Helm merges maps; a probe allows only one handler). ??? note "Default" ```yaml exec: command: - /bin/sh - /opt/payara/k8s/ready.sh failureThreshold: 10 periodSeconds: 60 timeoutSeconds: 30 ``` `hopsworks.probs.admin.livenessProbe.exec` # { #helm.hopsworks.probs.admin.livenessProbe.exec } : Type `object`, default `{"command":["/bin/sh","/opt/payara/k8s/ready.sh"]}`. exec probe handler. Set to null when overriding the probe with httpGet/tcpSocket (Helm merges maps; a probe allows only one handler). `hopsworks.probs.admin.readinessProbe` # { #helm.hopsworks.probs.admin.readinessProbe } : Type `object`. readiness probe. Controls the admin pod's Ready state and so the hopsworks-admin Service endpoint that workers use to reach the DAS. ready.sh only checks that the DAS admin interface responds. Short 10s period on purpose — this is the probe that must react fast; liveness samples slower. If you override with httpGet/tcpSocket, also set `exec: null`. ??? note "Default" ```yaml exec: command: - /bin/sh - /opt/payara/k8s/ready.sh initialDelaySeconds: 90 periodSeconds: 10 timeoutSeconds: 30 ``` `hopsworks.probs.admin.readinessProbe.exec` # { #helm.hopsworks.probs.admin.readinessProbe.exec } : Type `object`, default `{"command":["/bin/sh","/opt/payara/k8s/ready.sh"]}`. exec probe handler. Set to null when overriding the probe with httpGet/tcpSocket (Helm merges maps; a probe allows only one handler). `hopsworks.probs.admin.startupProbe` # { #helm.hopsworks.probs.admin.startupProbe } : Type `object`. startup probe. Gives the DAS up to 30 minutes (180 x 10s) to boot before liveness starts counting. Generous on purpose: container start includes artifact downloads, and a kill mid-download starts them over. Set to null to disable. If you override with httpGet/tcpSocket, also set `exec: null`. ??? note "Default" ```yaml exec: command: - /bin/sh - /opt/payara/k8s/ready.sh failureThreshold: 180 periodSeconds: 10 timeoutSeconds: 30 ``` `hopsworks.probs.admin.startupProbe.exec` # { #helm.hopsworks.probs.admin.startupProbe.exec } : Type `object`, default `{"command":["/bin/sh","/opt/payara/k8s/ready.sh"]}`. exec probe handler. Set to null when overriding the probe with httpGet/tcpSocket (Helm merges maps; a probe allows only one handler). `hopsworks.probs.worker.livenessProbe.failureThreshold` # { #helm.hopsworks.probs.worker.livenessProbe.failureThreshold } : Type `int`, default `3`. `hopsworks.probs.worker.livenessProbe.httpGet.path` # { #helm.hopsworks.probs.worker.livenessProbe.httpGet.path } : Type `string`, default `"/hopsworks-api/api/variables/versions"`. `hopsworks.probs.worker.livenessProbe.httpGet.port` # { #helm.hopsworks.probs.worker.livenessProbe.httpGet.port } : Type `int`, default `8182`. `hopsworks.probs.worker.livenessProbe.httpGet.scheme` # { #helm.hopsworks.probs.worker.livenessProbe.httpGet.scheme } : Type `string`, default `"HTTPS"`. `hopsworks.probs.worker.livenessProbe.initialDelaySeconds` # { #helm.hopsworks.probs.worker.livenessProbe.initialDelaySeconds } : Type `int`, default `150`. `hopsworks.probs.worker.livenessProbe.periodSeconds` # { #helm.hopsworks.probs.worker.livenessProbe.periodSeconds } : Type `int`, default `20`. `hopsworks.probs.worker.livenessProbe.timeoutSeconds` # { #helm.hopsworks.probs.worker.livenessProbe.timeoutSeconds } : Type `int`, default `20`. `hopsworks.probs.worker.readinessProbe.httpGet.path` # { #helm.hopsworks.probs.worker.readinessProbe.httpGet.path } : Type `string`, default `"/hopsworks-api/api/variables/versions"`. `hopsworks.probs.worker.readinessProbe.httpGet.port` # { #helm.hopsworks.probs.worker.readinessProbe.httpGet.port } : Type `int`, default `8182`. `hopsworks.probs.worker.readinessProbe.httpGet.scheme` # { #helm.hopsworks.probs.worker.readinessProbe.httpGet.scheme } : Type `string`, default `"HTTPS"`. `hopsworks.probs.worker.readinessProbe.initialDelaySeconds` # { #helm.hopsworks.probs.worker.readinessProbe.initialDelaySeconds } : Type `int`, default `150`. `hopsworks.probs.worker.readinessProbe.periodSeconds` # { #helm.hopsworks.probs.worker.readinessProbe.periodSeconds } : Type `int`, default `10`. `hopsworks.probs.worker.startupProbe.failureThreshold` # { #helm.hopsworks.probs.worker.startupProbe.failureThreshold } : Type `int`, default `10`. `hopsworks.probs.worker.startupProbe.httpGet.path` # { #helm.hopsworks.probs.worker.startupProbe.httpGet.path } : Type `string`, default `"/health"`. `hopsworks.probs.worker.startupProbe.httpGet.port` # { #helm.hopsworks.probs.worker.startupProbe.httpGet.port } : Type `int`, default `8182`. `hopsworks.probs.worker.startupProbe.httpGet.scheme` # { #helm.hopsworks.probs.worker.startupProbe.httpGet.scheme } : Type `string`, default `"HTTPS"`. `hopsworks.probs.worker.startupProbe.initialDelaySeconds` # { #helm.hopsworks.probs.worker.startupProbe.initialDelaySeconds } : Type `int`, default `150`. `hopsworks.probs.worker.startupProbe.periodSeconds` # { #helm.hopsworks.probs.worker.startupProbe.periodSeconds } : Type `int`, default `10`.
## rbac { #helm-values-hopsworks-rbac } ??? example "Defaults as YAML" ```yaml hopsworks: rbac: annotations: {} create: true extraRoleRules: - apiGroups: - discovery.k8s.io resources: - endpointslices verbs: - get - list - apiGroups: - sparkoperator.k8s.io resources: - sparkapplications verbs: - '*' - apiGroups: - ray.io resources: - rayjobs - rayclusters verbs: - '*' - apiGroups: - serving.knative.dev resources: - services verbs: - '*' - apiGroups: - serving.kserve.io resources: - inferenceservices - clusterservingruntimes verbs: - '*' - apiGroups: - scheduling.k8s.io resources: - priorityclasses verbs: - get - list - apiGroups: - metrics.k8s.io resources: - nodes - pods verbs: - get - list - apiGroups: - '' resources: - resourcequotas verbs: - get - list useExistingRole: false ```
`hopsworks.rbac.annotations` # { #helm.hopsworks.rbac.annotations } : Type `object`, default `{}`. annotations `hopsworks.rbac.create` # { #helm.hopsworks.rbac.create } : Type `bool`, default `true`. `hopsworks.rbac.extraRoleRules[0].apiGroups[0]` # { #helm.hopsworks.rbac.extraRoleRules.0.apiGroups.0 } : Type `string`, default `"discovery.k8s.io"`. `hopsworks.rbac.extraRoleRules[0].resources[0]` # { #helm.hopsworks.rbac.extraRoleRules.0.resources.0 } : Type `string`, default `"endpointslices"`. `hopsworks.rbac.extraRoleRules[0].verbs[0]` # { #helm.hopsworks.rbac.extraRoleRules.0.verbs.0 } : Type `string`, default `"get"`. `hopsworks.rbac.extraRoleRules[0].verbs[1]` # { #helm.hopsworks.rbac.extraRoleRules.0.verbs.1 } : Type `string`, default `"list"`. `hopsworks.rbac.extraRoleRules[1].apiGroups[0]` # { #helm.hopsworks.rbac.extraRoleRules.1.apiGroups.0 } : Type `string`, default `"sparkoperator.k8s.io"`. `hopsworks.rbac.extraRoleRules[1].resources[0]` # { #helm.hopsworks.rbac.extraRoleRules.1.resources.0 } : Type `string`, default `"sparkapplications"`. `hopsworks.rbac.extraRoleRules[1].verbs[0]` # { #helm.hopsworks.rbac.extraRoleRules.1.verbs.0 } : Type `string`, default `"*"`. `hopsworks.rbac.extraRoleRules[2].apiGroups[0]` # { #helm.hopsworks.rbac.extraRoleRules.2.apiGroups.0 } : Type `string`, default `"ray.io"`. `hopsworks.rbac.extraRoleRules[2].resources[0]` # { #helm.hopsworks.rbac.extraRoleRules.2.resources.0 } : Type `string`, default `"rayjobs"`. `hopsworks.rbac.extraRoleRules[2].resources[1]` # { #helm.hopsworks.rbac.extraRoleRules.2.resources.1 } : Type `string`, default `"rayclusters"`. `hopsworks.rbac.extraRoleRules[2].verbs[0]` # { #helm.hopsworks.rbac.extraRoleRules.2.verbs.0 } : Type `string`, default `"*"`. `hopsworks.rbac.extraRoleRules[3].apiGroups[0]` # { #helm.hopsworks.rbac.extraRoleRules.3.apiGroups.0 } : Type `string`, default `"serving.knative.dev"`. `hopsworks.rbac.extraRoleRules[3].resources[0]` # { #helm.hopsworks.rbac.extraRoleRules.3.resources.0 } : Type `string`, default `"services"`. `hopsworks.rbac.extraRoleRules[3].verbs[0]` # { #helm.hopsworks.rbac.extraRoleRules.3.verbs.0 } : Type `string`, default `"*"`. `hopsworks.rbac.extraRoleRules[4].apiGroups[0]` # { #helm.hopsworks.rbac.extraRoleRules.4.apiGroups.0 } : Type `string`, default `"serving.kserve.io"`. `hopsworks.rbac.extraRoleRules[4].resources[0]` # { #helm.hopsworks.rbac.extraRoleRules.4.resources.0 } : Type `string`, default `"inferenceservices"`. `hopsworks.rbac.extraRoleRules[4].resources[1]` # { #helm.hopsworks.rbac.extraRoleRules.4.resources.1 } : Type `string`, default `"clusterservingruntimes"`. `hopsworks.rbac.extraRoleRules[4].verbs[0]` # { #helm.hopsworks.rbac.extraRoleRules.4.verbs.0 } : Type `string`, default `"*"`. `hopsworks.rbac.extraRoleRules[5].apiGroups[0]` # { #helm.hopsworks.rbac.extraRoleRules.5.apiGroups.0 } : Type `string`, default `"scheduling.k8s.io"`. `hopsworks.rbac.extraRoleRules[5].resources[0]` # { #helm.hopsworks.rbac.extraRoleRules.5.resources.0 } : Type `string`, default `"priorityclasses"`. `hopsworks.rbac.extraRoleRules[5].verbs[0]` # { #helm.hopsworks.rbac.extraRoleRules.5.verbs.0 } : Type `string`, default `"get"`. `hopsworks.rbac.extraRoleRules[5].verbs[1]` # { #helm.hopsworks.rbac.extraRoleRules.5.verbs.1 } : Type `string`, default `"list"`. `hopsworks.rbac.extraRoleRules[6].apiGroups[0]` # { #helm.hopsworks.rbac.extraRoleRules.6.apiGroups.0 } : Type `string`, default `"metrics.k8s.io"`. `hopsworks.rbac.extraRoleRules[6].resources[0]` # { #helm.hopsworks.rbac.extraRoleRules.6.resources.0 } : Type `string`, default `"nodes"`. `hopsworks.rbac.extraRoleRules[6].resources[1]` # { #helm.hopsworks.rbac.extraRoleRules.6.resources.1 } : Type `string`, default `"pods"`. `hopsworks.rbac.extraRoleRules[6].verbs[0]` # { #helm.hopsworks.rbac.extraRoleRules.6.verbs.0 } : Type `string`, default `"get"`. `hopsworks.rbac.extraRoleRules[6].verbs[1]` # { #helm.hopsworks.rbac.extraRoleRules.6.verbs.1 } : Type `string`, default `"list"`. `hopsworks.rbac.extraRoleRules[7].apiGroups[0]` # { #helm.hopsworks.rbac.extraRoleRules.7.apiGroups.0 } : Type `string`, default `""`. `hopsworks.rbac.extraRoleRules[7].resources[0]` # { #helm.hopsworks.rbac.extraRoleRules.7.resources.0 } : Type `string`, default `"resourcequotas"`. `hopsworks.rbac.extraRoleRules[7].verbs[0]` # { #helm.hopsworks.rbac.extraRoleRules.7.verbs.0 } : Type `string`, default `"get"`. `hopsworks.rbac.extraRoleRules[7].verbs[1]` # { #helm.hopsworks.rbac.extraRoleRules.7.verbs.1 } : Type `string`, default `"list"`. `hopsworks.rbac.useExistingRole` # { #helm.hopsworks.rbac.useExistingRole } : Type `bool`, default `false`.
## resources { #helm-values-hopsworks-resources } ??? example "Defaults as YAML" ```yaml hopsworks: resources: admin: auto_jvm: true container: limits: cpu: 2000m memory: 2048Mi requests: cpu: 1000m memory: 2048Mi extraJavaToolOptions: '' jvm: garbageCollector: '' memory: buffer: 2048 compressedClassSpaceSize: 512 heap: 4096 metaspace: 2048 nonMethodCodeHeapSize: 5 nonProfiledCodeHeapSize: 48 profiledCodeHeapSize: 48 adminSidecar: requests: cpu: 200m memory: 256Mi worker: auto_jvm: true container: limits: cpu: 4000m memory: 8Gi requests: cpu: 2000m memory: 2048Mi extraJavaToolOptions: '' jvm: garbageCollector: '' memory: buffer: 3072 compressedClassSpaceSize: 512 heap: 4096 metaspace: 2048 nonMethodCodeHeapSize: 5 nonProfiledCodeHeapSize: 48 profiledCodeHeapSize: 48 ```
`hopsworks.resources.admin.auto_jvm` # { #helm.hopsworks.resources.admin.auto_jvm } : Type `bool`, default `true`. `hopsworks.resources.admin.container.limits.cpu` # { #helm.hopsworks.resources.admin.container.limits.cpu } : Type `string`, default `"2000m"`. `hopsworks.resources.admin.container.limits.memory` # { #helm.hopsworks.resources.admin.container.limits.memory } : Type `string`, default `"2048Mi"`. `hopsworks.resources.admin.container.requests.cpu` # { #helm.hopsworks.resources.admin.container.requests.cpu } : Type `string`, default `"1000m"`. `hopsworks.resources.admin.container.requests.memory` # { #helm.hopsworks.resources.admin.container.requests.memory } : Type `string`, default `"2048Mi"`. `hopsworks.resources.admin.extraJavaToolOptions` # { #helm.hopsworks.resources.admin.extraJavaToolOptions } : Type `string`, default `""`. `hopsworks.resources.admin.jvm.garbageCollector` # { #helm.hopsworks.resources.admin.jvm.garbageCollector } : Type `string`, default `""`. `hopsworks.resources.admin.jvm.memory.buffer` # { #helm.hopsworks.resources.admin.jvm.memory.buffer } : Type `int`, default `2048`. `hopsworks.resources.admin.jvm.memory.compressedClassSpaceSize` # { #helm.hopsworks.resources.admin.jvm.memory.compressedClassSpaceSize } : Type `int`, default `512`. `hopsworks.resources.admin.jvm.memory.heap` # { #helm.hopsworks.resources.admin.jvm.memory.heap } : Type `int`, default `4096`. `hopsworks.resources.admin.jvm.memory.metaspace` # { #helm.hopsworks.resources.admin.jvm.memory.metaspace } : Type `int`, default `2048`. `hopsworks.resources.admin.jvm.memory.nonMethodCodeHeapSize` # { #helm.hopsworks.resources.admin.jvm.memory.nonMethodCodeHeapSize } : Type `int`, default `5`. `hopsworks.resources.admin.jvm.memory.nonProfiledCodeHeapSize` # { #helm.hopsworks.resources.admin.jvm.memory.nonProfiledCodeHeapSize } : Type `int`, default `48`. `hopsworks.resources.admin.jvm.memory.profiledCodeHeapSize` # { #helm.hopsworks.resources.admin.jvm.memory.profiledCodeHeapSize } : Type `int`, default `48`. `hopsworks.resources.adminSidecar` # { #helm.hopsworks.resources.adminSidecar } : Type `object`, default `{"requests":{"cpu":"200m","memory":"256Mi"}}`. Admin side car resources `hopsworks.resources.worker.auto_jvm` # { #helm.hopsworks.resources.worker.auto_jvm } : Type `bool`, default `true`. `hopsworks.resources.worker.container.limits.cpu` # { #helm.hopsworks.resources.worker.container.limits.cpu } : Type `string`, default `"4000m"`. `hopsworks.resources.worker.container.limits.memory` # { #helm.hopsworks.resources.worker.container.limits.memory } : Type `string`, default `"8Gi"`. `hopsworks.resources.worker.container.requests.cpu` # { #helm.hopsworks.resources.worker.container.requests.cpu } : Type `string`, default `"2000m"`. `hopsworks.resources.worker.container.requests.memory` # { #helm.hopsworks.resources.worker.container.requests.memory } : Type `string`, default `"2048Mi"`. `hopsworks.resources.worker.extraJavaToolOptions` # { #helm.hopsworks.resources.worker.extraJavaToolOptions } : Type `string`, default `""`. `hopsworks.resources.worker.jvm.garbageCollector` # { #helm.hopsworks.resources.worker.jvm.garbageCollector } : Type `string`, default `""`. `hopsworks.resources.worker.jvm.memory.buffer` # { #helm.hopsworks.resources.worker.jvm.memory.buffer } : Type `int`, default `3072`. `hopsworks.resources.worker.jvm.memory.compressedClassSpaceSize` # { #helm.hopsworks.resources.worker.jvm.memory.compressedClassSpaceSize } : Type `int`, default `512`. `hopsworks.resources.worker.jvm.memory.heap` # { #helm.hopsworks.resources.worker.jvm.memory.heap } : Type `int`, default `4096`. `hopsworks.resources.worker.jvm.memory.metaspace` # { #helm.hopsworks.resources.worker.jvm.memory.metaspace } : Type `int`, default `2048`. `hopsworks.resources.worker.jvm.memory.nonMethodCodeHeapSize` # { #helm.hopsworks.resources.worker.jvm.memory.nonMethodCodeHeapSize } : Type `int`, default `5`. `hopsworks.resources.worker.jvm.memory.nonProfiledCodeHeapSize` # { #helm.hopsworks.resources.worker.jvm.memory.nonProfiledCodeHeapSize } : Type `int`, default `48`. `hopsworks.resources.worker.jvm.memory.profiledCodeHeapSize` # { #helm.hopsworks.resources.worker.jvm.memory.profiledCodeHeapSize } : Type `int`, default `48`.
## service { #helm-values-hopsworks-service } ??? example "Defaults as YAML" ```yaml hopsworks: service: admin: annotations: consul.hashicorp.com/service-name: glassfish consul.hashicorp.com/service-tags: admin name: admin port: 4848 type: ClusterIP worker: external: http: port: 28080 type: ClusterIP https: nodePort: null port: 28181 type: ClusterIP internal: annotations: consul.hashicorp.com/service-name: glassfish consul.hashicorp.com/service-tags: hopsworks prometheus.io/path: /metrics prometheus.io/port: 8182 prometheus.io/scheme: https prometheus.io/scrape: 'true' port: 8182 type: ClusterIP ```
`hopsworks.service.admin.annotations."consul.hashicorp.com/service-name"` # { #helm.hopsworks.service.admin.annotations.consul.hashicorp.com-service-name } : Type `string`, default `"glassfish"`. `hopsworks.service.admin.annotations."consul.hashicorp.com/service-tags"` # { #helm.hopsworks.service.admin.annotations.consul.hashicorp.com-service-tags } : Type `string`, default `"admin"`. `hopsworks.service.admin.name` # { #helm.hopsworks.service.admin.name } : Type `string`, default `"admin"`. `hopsworks.service.admin.port` # { #helm.hopsworks.service.admin.port } : Type `int`, default `4848`. `hopsworks.service.admin.type` # { #helm.hopsworks.service.admin.type } : Type `string`, default `"ClusterIP"`. `hopsworks.service.worker.external.http.port` # { #helm.hopsworks.service.worker.external.http.port } : Type `int`, default `28080`. `hopsworks.service.worker.external.http.type` # { #helm.hopsworks.service.worker.external.http.type } : Type `string`, default `"ClusterIP"`. `hopsworks.service.worker.external.https.nodePort` # { #helm.hopsworks.service.worker.external.https.nodePort } : Type `string`, default `nil`. Explicit nodePort for the https service when type is NodePort. Null lets Kubernetes allocate one from the cluster's node-port range; a set value must lie in that range (30000-32767 by default), which the API server enforces at install. `hopsworks.service.worker.external.https.port` # { #helm.hopsworks.service.worker.external.https.port } : Type `int`, default `28181`. `hopsworks.service.worker.external.https.type` # { #helm.hopsworks.service.worker.external.https.type } : Type `string`, default `"ClusterIP"`. `hopsworks.service.worker.internal.annotations."consul.hashicorp.com/service-name"` # { #helm.hopsworks.service.worker.internal.annotations.consul.hashicorp.com-service-name } : Type `string`, default `"glassfish"`. `hopsworks.service.worker.internal.annotations."consul.hashicorp.com/service-tags"` # { #helm.hopsworks.service.worker.internal.annotations.consul.hashicorp.com-service-tags } : Type `string`, default `"hopsworks"`. `hopsworks.service.worker.internal.annotations."prometheus.io/path"` # { #helm.hopsworks.service.worker.internal.annotations.prometheus.io-path } : Type `string`, default `"/metrics"`. `hopsworks.service.worker.internal.annotations."prometheus.io/port"` # { #helm.hopsworks.service.worker.internal.annotations.prometheus.io-port } : Type `int`, default `8182`. `hopsworks.service.worker.internal.annotations."prometheus.io/scheme"` # { #helm.hopsworks.service.worker.internal.annotations.prometheus.io-scheme } : Type `string`, default `"https"`. `hopsworks.service.worker.internal.annotations."prometheus.io/scrape"` # { #helm.hopsworks.service.worker.internal.annotations.prometheus.io-scrape } : Type `string`, default `"true"`. `hopsworks.service.worker.internal.port` # { #helm.hopsworks.service.worker.internal.port } : Type `int`, default `8182`. `hopsworks.service.worker.internal.type` # { #helm.hopsworks.service.worker.internal.type } : Type `string`, default `"ClusterIP"`.
## terminal { #helm-values-hopsworks-terminal } ??? example "Defaults as YAML" ```yaml hopsworks: terminal: enabled: false oomGuard: true proxyPodAppLabels: hopsworks-instance,hopsworks-admin proxyTokenTtlMs: '60000' teleportCleanerIntervalMs: '86400000' teleportTtlDays: 7 ```
`hopsworks.terminal.enabled` # { #helm.hopsworks.terminal.enabled } : Type `bool`, default `false`. `hopsworks.terminal.oomGuard` # { #helm.hopsworks.terminal.oomGuard } : Type `bool`, default `true`. In-image memory guard: kills the hungriest process in a terminal pod before the kernel group-OOM-kills the whole session (cgroup v2 only). False sets HOPS_OOM_GUARD_DISABLE=1 on new terminal pods. `hopsworks.terminal.proxyPodAppLabels` # { #helm.hopsworks.terminal.proxyPodAppLabels } : Type `string`, default `"hopsworks-instance,hopsworks-admin"`. Comma-separated `app` label values of the Payara pods allowed to reach a terminal pod's WebSocket port (the per-namespace terminal-isolation NetworkPolicy). Must match the chart's pod labels or every terminal is unreachable. `hopsworks.terminal.proxyTokenTtlMs` # { #helm.hopsworks.terminal.proxyTokenTtlMs } : Type `string`, default `"60000"`. Milliseconds a CLI terminal-attach proxy token stays valid before its single WebSocket handshake. `hopsworks.terminal.teleportCleanerIntervalMs` # { #helm.hopsworks.terminal.teleportCleanerIntervalMs } : Type `string`, default `"86400000"`. Milliseconds between teleport reaper sweeps (applied at Payara start). Quoted like the other *_ms settings: a bare integer this large renders as 8.64e+07 in the DML. `hopsworks.terminal.teleportTtlDays` # { #helm.hopsworks.terminal.teleportTtlDays } : Type `int`, default `7`. Days a staged hops-session teleport file (transcript, manifest, baton) survives in a user's HopsFS home before the reaper deletes it; 0 or less disables the reaper.
## variables { #helm-values-hopsworks-variables } ??? example "Defaults as YAML" ```yaml hopsworks: variables: admin_email: admin@hopsworks.ai admin_password: admin agent_deployment_otel_cpu: '0.5' agent_deployment_otel_enabled: 'true' agent_deployment_otel_image: otlp-sidecar agent_deployment_otel_memory_mb: '1024' agent_jobs_enabled: 'true' airflow_dir: /srv/hops/airflow airflow_enabled: true airflow_user: airflow airflow_user_email: airflow@hopsworks.ai alert_email_addrs: '' anaconda_dir: / anaconda_enabled: 'true' anaconda_env: '' anaconda_user: anaconda application_certificate_validity_period: 3650d apply_hopsfsmount_apparmor_profile_kube: 'false' async_services_timer_batch_size: '1000' async_services_timer_delete_history_after_days: '7' async_services_timer_enabled: 'true' async_services_timer_interval_ms: '15000' audit_log_count: '10' audit_log_file_format: server_audit_log%g.log audit_log_file_path: /audit-logs audit_log_file_type: io.hops.hopsworks.audit.helper.JSONLogFormatter audit_log_size_limit: '256000000' base_buildkit_image: docker.hops.works/hopsworks/moby/buildkit:v0.32.2-rootless base_image_name: hopsworks-base base_image_version: 5.2.0-SNAPSHOT cert_mater_delay: 3m certs_dir: /srv/hops/certs-dir check_nodemanagers_status: '' client_path: /srv/hops/clients-3.4.3 cloud: '' command_agent_home_batch: '20' command_agent_home_claim_lease_as_ms: '600000' command_agent_home_migration_period_as_ms: '3600000' command_agent_home_process_timer_period_as_ms: '5000' command_agent_home_retry_backoff_base_as_ms: '10000' command_agent_home_retry_backoff_max_as_ms: '600000' command_search_fs_history_clean_period_as_ms: '3600000' command_search_fs_history_enable: 'false' command_search_fs_history_window_as_s: '3600' command_search_fs_process_timer_period_as_ms: '1000' command_search_fs_reindex_queue_wait_as_ms: '1800000' command_search_fs_retry_per_clean_interval: '5' conda_default_repo: defaults default_jupyter_environment: pandas-training-pipeline default_python_job_environment: pandas-training-pipeline disable_password_login: 'false' disable_registration: 'false' dlt_schema_fetch_job_deadline_seconds: '1800' docker_base_image_python_version: '3.13' docker_cgroup_cpu_period: '100000' docker_cgroup_enabled: 'false' docker_cgroup_parent: docker.slice docker_job_mounts_allowed: 'false' docker_job_mounts_list: '' docker_job_uid_strict: 'true' docker_mounts: /srv/hops/hadoop/etc/hadoop,/srv/hops/spark,/srv/hops/flink,/srv/hops/apache-livy docker_operations_allow_hermetic_custom_commands: 'false' docker_operations_backoff_limit: '0' docker_operations_build_metadata: 'true' docker_operations_buildkit_addr: '' docker_operations_buildkit_backoff_limit: '0' docker_operations_buildkit_cache_scope: shared docker_operations_buildkit_extra_args: '' docker_operations_buildkit_limit_cpu: '2' docker_operations_buildkit_limit_memory: 4G docker_operations_buildkit_priority_class: '' docker_operations_buildkit_replicas: '1' docker_operations_buildkit_request_cpu: 200m docker_operations_buildkit_request_memory: 500Mi docker_operations_buildkit_storage: 70Gi docker_operations_buildkit_tls_locality: buildkitd docker_operations_buildkit_tls_secret: '' docker_operations_cert_name: kagent_certificate_bundle.pem docker_operations_context_orphan_minutes: '120' docker_operations_default_service_account: default docker_operations_delete_jobs_add_description_if_fails: false docker_operations_delete_jobs_on_completion: 'true' docker_operations_docker_context_builder: AUTO docker_operations_docker_context_builder_s3_bucket: ${env:S3_BUCKET} docker_operations_docker_context_builder_s3_endpoint: ${env:S3_ENDPOINT} docker_operations_docker_context_builder_s3_region: ${env:S3_REGION} docker_operations_hopsworks_ca_secret_name: docker-registry-crypto-material docker_operations_image_pull_secrets: '' docker_operations_lock_dependencies: 'false' docker_operations_managed_docker_secrets: '' docker_operations_multi_region_copy: 'false' docker_operations_oci_worker_snapshotter: auto docker_operations_push_insecure: 'false' docker_operations_registry_container: docker docker_operations_registry_http: 'false' docker_operations_registry_pod: docker-registry-0 docker_operations_suspend_jobs: 'false' docker_operations_timeout_check_minutes: '15' docker_operations_timeout_delete_minutes: '5' docker_operations_timeout_export_minutes: '15' docker_operations_timeout_listing_minutes: '15' docker_operations_timeout_minutes_buildkit: '120' docker_operations_timeout_tag_minutes: '5' download_allowed: 'true' elastic_dir: /srv/hops/elastic elastic_https_enabled: 'true' elastic_jwt_enabled: 'true' elastic_jwt_exp_ms: '1800000' elastic_jwt_url_parameter: jt elastic_logs_index_expiration: '604800000' elastic_opendistro_security_enabled: 'true' elastic_user: elastic elastic_version: 3.8.0 enable_adls_storage_connectors: 'false' enable_bigquery_storage_connectors: 'true' enable_bring_your_own_kafka: 'false' enable_feature_monitoring: 'true' enable_fix_receivers_timer: 'true' enable_gcs_storage_connectors: 'true' enable_jupyter_python_kernel_non_kubernetes: 'false' enable_kafka_storage_connectors: 'true' enable_metadata_designer: '' enable_opensearch_storage_connectors: 'true' enable_read_only_git_repositories: 'false' enable_redshift_storage_connectors: 'true' enable_snowflake_storage_connectors: 'true' enable_user_search: 'true' epipe_version: 0.20.0 executions_cleaner_batch_size: '50' executions_cleaner_interval_ms: '600000' executions_per_job_limit: '10000' feature_monitoring_max_num_features: '15' featurestore_asof_spine_max_bytes: '1073741824' featurestore_asof_spine_max_columns: '256' featurestore_asof_spine_max_file_age_ms: '86400000' featurestore_asof_spine_max_rows: '1000000' featurestore_db_admin_user: featurestore_admin_user featurestore_default_quota: -1L featurestore_default_storage_format: PARQUET featurestore_metrics_enabled: 'true' featurestore_metrics_online_ingestion_enabled: 'false' featurestore_online_enabled: 'true' featurestore_online_tablespace: '' file_preview_image_size: '10000000' file_preview_txt_size: '100' flink_dir: /srv/hops/flink flink_user: flink flink_version: 1.17.1.0 fs_job_activity_time: 5m fs_storage_connector_session_duration: '3600' git_bitbucket_http_proxy: '' git_bitbucket_https_proxy: '' git_command_timeout_minutes: '60' git_custom_ca_configmap: '' git_custom_ca_configmap_key: ca-bundle.crt git_disable_tls_verification: 'false' git_github_http_proxy: '' git_github_https_proxy: '' git_gitlab_http_proxy: '' git_gitlab_https_proxy: '' git_image_version: 1.6-SNAPSHOT grafana_version: 9.3.16 ha_enabled: 'true' hadoop_dir: /srv/hops/hadoop hadoop_version: 3.4.3.3-EE-RC1 hdfs_base_storage_policy: CLOUD hdfs_default_quota: -1L hdfs_log_storage_policy: CLOUD hdfs_user: hdfs hdfscontentsmanager_base_hopsfs_client: libhdfs-go hive2_version: 4.1.0.0-v1 hive_conf_path: /srv/hops/apache-hive/conf hive_superuser: hive hive_warehouse: /apps/hive/warehouse hiveserver_ext_hostname: '' hiveserver_ssl_hostname: '' hops_db: hops hops_rpc_tls: 'true' hopsexamples_version: '' hopsfsmount_apparmor_profile: '' hopsfsmount_log_level: warn hopsfsmount_nn_connections: '4' hopsworks_analytics: false hopsworks_analytics_coding_agent: claude hopsworks_analytics_ro_user: hopsworks_ro hopsworks_analytics_ro_user_adopt_existing: false hopsworks_analytics_setup_repo: https://github.com/logicalclocks/okr-dashboards hopsworks_db: hopsworks hopsworks_dir: /srv/hops/domains/domain1 hopsworks_enterprise: 'true' hopsworks_mysql_user: hopsworks hopsworks_public_proxy_url: '' hopsworks_rest_log_level: TEST hopsworks_user: payara hw_group_mapping_sync_enabled: 'false' ingestion_job_cores: '1.0' ingestion_job_gpus: '0' ingestion_job_memory: '2048' java_home: '' job_name_validation_regex: ^[a-zA-Z0-9_\-]+$ jupyter_allow_no_limit_shutdown: true jupyter_dir: /srv/hops/jupyter jupyter_group: hadoop jupyter_hour_shutdown_options: 8,24,72 jupyter_origin_scheme: https jupyter_shell_command: '["/bin/bash", "--login", "-c", "cd -L $JUPYTER_DATA_DIR || true && exec bash"]' jupyter_shutdown_timer_interval: 1m jupyter_spark_notebook_server_memory_floor_mb: '512' jupyter_ws_ping_interval: 10s jwt_exp_leeway_sec: '900' jwt_issuer: hopsworks@logicalclocks.com jwt_lifetime_ms: '86400000' jwt_signature_algorithm: HS512 jwt_signing_key_name: apiKey kafka_installed: true kafka_max_num_topics: '100' kafka_num_partitions: '1' kafka_num_replicas: '1' kafka_user: kafka kafka_version: 4.3.1 kibana_https_enabled: 'true' kibana_multi_tenancy_enabled: 'true' kibana_version: 3.8.0 kube_api_max_attempts: '20' kube_hopsworks_default_service_account: hopsworks-default kube_knative_domain_name: hopsworks.ai kube_knative_lb_domain: '' kube_kserve_installed: true kube_kserve_tensorflow_version: 2.20.0 kube_node_taints_monitor_interval: 10m kube_scheduling_hopsfsmount_cpu_limits: -1 kube_scheduling_hopsfsmount_cpu_requests: 1 kube_scheduling_hopsfsmount_memory_limits_mb: 1024 kube_scheduling_jobinit_cpu_limits: -1 kube_scheduling_jobinit_cpu_requests: 0.5 kube_scheduling_jobinit_memory_limits_mb: 512 kube_scheduling_jobinit_memory_requests_mb: 256 kube_serving_max_num_instances: '10' kube_serving_min_num_instances: '-1' kube_serving_vllm_omni_versions: v0.28.0 kube_serving_vllm_versions: v0.28.0 kube_skip_namespace_creation: false kube_tainted_nodes: '' kube_type: kube_cluster kube_user_workload_tolerations: '' kubernetes_installed: 'true' kueue_project_default_cluster_queue: other kueue_project_default_local_queue: other kueue_system_jobs_cluster_queue: '' kueue_system_jobs_local_queue: '' ldap_account_status: '2' ldap_attr_binary: java.naming.ldap.attributes.binary ldap_dyn_group_target: memberOf ldap_group_dn: '' ldap_group_mapping: ANY_GROUP->HOPS_USER ldap_group_mapping_sync_enabled: 'false' ldap_group_mapping_sync_interval: '0' ldap_group_search_filter: member=%d ldap_group_target: cn ldap_groups_search_filter: (&(objectCategory=group)(cn=%c)) ldap_krb_dyn_grp_search_filter: '' ldap_krb_search_filter: krbPrincipalName=%s ldap_user_dn: '' ldap_user_email: mail ldap_user_givenName: givenName ldap_user_id: uid ldap_user_search_filter: uid=%s ldap_user_surname: sn library_install_timeout_minutes: '60' lifecycle_webhook_cluster_id: '' lifecycle_webhook_secret: '' lifecycle_webhook_url: '' livy_startup_timeout: '240' livy_version: 0.8.4-incubating-SNAPSHOT-bin loadbalancer_external_domain_datanode: null loadbalancer_external_domain_feature_query: null loadbalancer_external_domain_mysqld: null loadbalancer_external_domain_namenode: null loadbalancer_external_domain_online_store_rest_server: null loadbalancer_external_domain_opensearch: null loadbalancer_external_domain_trino: null localhost: 'false' log_history_limit: '30' logstash_ip: '' logstash_port: '' logstash_port_beam_jobserver_local: '' logstash_port_serving: '' logstash_port_sklearn_serving: '' logstash_port_tf_serving: '' logstash_version: 7.16.3 managed_cloud_redirect_uri: '' managed_docker_registry: 'false' management_mode: '' max_allowed_long_running_http_requests: '50' max_concurrent_base_sync_ops: '5' max_env_yml_byte_size: '20000' max_num_proj_per_user: '10' max_status_poll_retry: '5' mount_hopsfs_in_python_job: true mount_hopsfs_ray_job_container: 'true' mr_user: mapred multiregion_watchdog_enabled: 'false' multiregion_watchdog_interval: 5s multiregion_watchdog_region: '' multiregion_watchdog_url: '' mysql_dir: /srv/hops/mysql ndb_dir: /srv/hop/mysql-cluster ndb_user: '' ndb_version: 21.04.15 ndbinfo_db: ndbinfo news_webflow_api_key: dcc84358bfd37ffc68dbf18c68f74f478ff160d2286094077a9415ee03fbc805 news_webflow_api_url: https://api.webflow.com/v2/collections/66bdd44475e24741477e1ae3/items notebook_converter_job_timeout_sec: '300' npm_registry_url: '' oauth_account_status: '1' oauth_group_mapping: '' oauth_group_mapping_enabled: 'false' oauth_group_mapping_sync_enabled: 'false' oauth_logout_redirect_uri: hopsworks/ oauth_redirect_uri: hopsworks/callback onlinefs_service_thread_number: '10' onlinefs_user_email: onlinefs@hopsworks.ai onlinefs_user_password: onlinefspw opensearch_default_embedding_index: '' opensearch_index_mapping_limit: '1000' opensearch_num_default_embedding_index: '1' payara_dir: /opt/payara/appserver/glassfish/domains/domain1 pki_ca_configuration: '{"rootCA":{},"intermediateCA":{},"kubernetesCA":{"subjectAlternativeName":{"dns":["hopsworks0.logicalclocks.com","hops-kubernetes","hops-kubernetes.default","hops-kubernetes.default.svc","hops-kubernetes.default.svc.cluster","hops-kubernetes.default.svc.cluster.local","*.hops-system.svc"],"ip":["10.244.0.1","192.168.30.101","127.0.0.1","10.96.0.10","10.96.0.1"]}}}' platform_intelligence_llm_api_key: '' platform_intelligence_llm_base_url: '' platform_intelligence_llm_model: '' preinstalled_python_lib_names: pydoop, pyspark, jupyterlab, sparkmagic, hdfscontents, pyjks, hops-apache-beam, pyopenssl project_namespace_labels: '' project_namespace_network_policy_allowed_namespaces: '' project_namespace_network_policy_enabled: true project_namespace_network_policy_reconcile_interval: 1m prometheus_port: '9089' provenance_archive_delay: '86400' provenance_archive_size: '10' provenance_cleaner_period: '3600' provenance_graph_max_size: '10000' provenance_type: FULL public_https_port: '' pushgateway_cleaner_batch_size: '100' pushgateway_group_ttl_minutes: '15' pushgateway_monitor_interval_ms: '300000' py4j_archive: '' pypi_indexer_timer_enabled: 'true' pypi_indexer_timer_interval: 1d pypi_rest_endpoint: https://pypi.org/pypi/{package}/json pypi_simple_endpoint: https://pypi.org/simple/ python_job_cores: '1.0' python_job_gpus: '0' python_job_kube_waiting_timeout_ms: '300000' python_job_memory: '2048' python_library_updates_monitor_interval: 1d python_pod_kill_grace_period_seconds: '60' pythonapp_cores: '1.0' pythonapp_gpus: '0' pythonapp_memory: '2048' quotas_featuregroups_online_disabled: '-1' quotas_featuregroups_online_enabled: '-1' quotas_max_parallel_executions: '-1' quotas_model_deployments_running: '-1' quotas_model_deployments_total: '-1' quotas_training_datasets: '-1' ray_cluster_max_worker_replicas: '20' ray_cluster_shutdown_after_completion: 'true' ray_cluster_start_wait_time_seconds: '360' ray_cluster_termination_grace_period_seconds: '10' ray_enabled: false ray_job_driver_cores: '1.0' ray_job_driver_gpus: '0' ray_job_driver_memory: '4096' ray_job_pod_kill_grace_period_seconds: '300' ray_job_worker_cores: '1.0' ray_job_worker_gpus: '0' ray_job_worker_memory: '4096' ray_materialization_dir: /srv/hops/ray/job ray_version: 2.58.0 recovery_path: '' reject_remote_user_no_group: 'false' remote_auth_need_consent: 'true' requests_verify: 'true' reserved_project_names: hopsworks,information_schema,airflow,glassfish_timers,grafana,hops,metastore,mysql,ndbinfo,performance_schema,sqoop,sys,base,python37,python38,python39,python310,filebeat,airflow,git,onlinefs,sklearnserver,rondb_replication,default,kube-system,kube-public,kube-node-lease,kube_system,kube_public,kube_node_lease rmyarn_user: rmyarn rondb_quotas: '' rondb_usage_cache_ttl_seconds: '60' rondb_usage_query_timeout_seconds: '10' saas_entry_point_url: '' scikit_learn_version: 1.3.2 service_jwt_exp_leeway_sec: '172800000' service_jwt_lifetime_ms: '604800000' service_key_rotation_enabled: 'false' service_key_rotation_interval: 2d serving_allow_stop_after_seconds: '30' serving_connection_pool_size: '40' serving_feature_log_materialization_cron: 0 0 0 * * ? * serving_feature_log_materialization_row_limit: '50000000' serving_feature_log_online_ttl_hours: '30' serving_feature_logger_batch_bytes: '1048576' serving_feature_logger_batch_seconds: '5' serving_feature_logger_client_pool_size: '3' serving_feature_logger_client_req_timeout_seconds: '3' serving_feature_logger_flush_bytes: '1048576' serving_feature_logger_flush_interval_seconds: '300' serving_feature_logger_max_buffer_bytes: '67108864' serving_feature_logger_max_event_bytes: '8388608' serving_feature_logger_max_event_rows: '512' serving_feature_logger_queue_size: '1000' serving_feature_logger_shutdown_seconds: '20' serving_feature_logging_transport: realtime serving_max_route_connections: '10' serving_redeploy_not_found_after_seconds: '120' serving_state_manager_batch_size: '25' serving_state_manager_enabled: 'true' serving_state_manager_interval_ms: '300000' spark_dir: /srv/hops/spark spark_executor_min_memory: '1024' spark_hops_utils_dir: /srv/hops/artifacts spark_job_driver_cores: '1.0' spark_job_driver_memory: '2048' spark_job_executor_cores: '1.0' spark_job_executor_memory: '4096' spark_launcher_sa_annotations: '' spark_pod_kill_grace_period_seconds: '1200' spark_remove_job_when_completed: 'true' spark_ui_logs_offset: '512000' spark_user: spark spark_version: 4.1.3.0 srvmanager_password: srvmanagerpwd staging_dir: /srv/hops/staging statistics_cleaner_batch_size: '1000' statistics_cleaner_interval_ms: '900000' streamlit_sharing: false sudoers_dir: /srv/hops/sbin superset_admin_roles: Admin superset_proxy_connect_timeout_ms: '10000' superset_proxy_connection_request_timeout_ms: '10000' superset_proxy_max_connections: '50' superset_proxy_read_timeout_ms: '180000' superset_user_roles: Gamma,sql_lab,Dataset support_email_addr: support@hopsworks.ai tag_history_archive_max_events: '20000' tag_history_cleaner_batch_size: '1000' tag_history_cleaner_interval_ms: '86400000' tag_history_retention_days: '0' tensorboard_max_last_accessed: '1140000' tensorboard_max_reload_threads: '1' tensorflow_version: 2.20.0 testconnector_image_version: '1.0' tf_spark_connector_version: '' trino_default_catalog: delta trino_events_cleaner_batch_size: '1000' trino_events_delete_after_days: '61' twofactor_auth: 'false' twofactor_excluded_groups: AGENT;CLUSTER_AGENT unix_usernames_conf: '{\"glassfish\":\"glassfish\",\"hdfs\":\"hdfs\",\"rmyarn\":\"rmyarn\",\"yarn\":\"yarn\",\"hive\":\"hive\",\"livy\":\"livy\",\"flink\":\"flink\",\"consul\":\"consul\",\"hopsmon\":\"hopsmon\",\"zookeeper\":\"zookeeper\",\"onlinefs\":\"onlinefs\",\"elastic\":\"elastic\",\"kagent\":\"kagent\",\"mysql\":\"mysql\",\"airflow\":\"airflow\"}' upload_chunk_size: '10485760' upload_policy: enabled user_cert_valid_days: '12' verification_path: hopsworks-api/api/auth/verify yarn_default_payment_type: NOLIMIT yarn_default_quota: '60000000' yarn_user: yarn zookeeper_version: 3.7.1 ```
`hopsworks.variables.admin_email` # { #helm.hopsworks.variables.admin_email } : Type `string`, default `"admin@hopsworks.ai"`. `hopsworks.variables.admin_password` # { #helm.hopsworks.variables.admin_password } : Type `string`, default `"admin"`. `hopsworks.variables.agent_deployment_otel_cpu` # { #helm.hopsworks.variables.agent_deployment_otel_cpu } : Type `string`, default `"0.5"`. `hopsworks.variables.agent_deployment_otel_enabled` # { #helm.hopsworks.variables.agent_deployment_otel_enabled } : Type `string`, default `"true"`. `hopsworks.variables.agent_deployment_otel_image` # { #helm.hopsworks.variables.agent_deployment_otel_image } : Type `string`, default `"otlp-sidecar"`. `hopsworks.variables.agent_deployment_otel_memory_mb` # { #helm.hopsworks.variables.agent_deployment_otel_memory_mb } : Type `string`, default `"1024"`. `hopsworks.variables.agent_jobs_enabled` # { #helm.hopsworks.variables.agent_jobs_enabled } : Type `string`, default `"true"`. `hopsworks.variables.airflow_dir` # { #helm.hopsworks.variables.airflow_dir } : Type `string`, default `"/srv/hops/airflow"`. `hopsworks.variables.airflow_enabled` # { #helm.hopsworks.variables.airflow_enabled } : Type `bool`, default `true`. `hopsworks.variables.airflow_user` # { #helm.hopsworks.variables.airflow_user } : Type `string`, default `"airflow"`. `hopsworks.variables.airflow_user_email` # { #helm.hopsworks.variables.airflow_user_email } : Type `string`, default `"airflow@hopsworks.ai"`. `hopsworks.variables.alert_email_addrs` # { #helm.hopsworks.variables.alert_email_addrs } : Type `string`, default `""`. `hopsworks.variables.anaconda_dir` # { #helm.hopsworks.variables.anaconda_dir } : Type `string`, default `"/"`. `hopsworks.variables.anaconda_enabled` # { #helm.hopsworks.variables.anaconda_enabled } : Type `string`, default `"true"`. `hopsworks.variables.anaconda_env` # { #helm.hopsworks.variables.anaconda_env } : Type `string`, default `""`. `hopsworks.variables.anaconda_user` # { #helm.hopsworks.variables.anaconda_user } : Type `string`, default `"anaconda"`. `hopsworks.variables.application_certificate_validity_period` # { #helm.hopsworks.variables.application_certificate_validity_period } : Type `string`, default `"3650d"`. `hopsworks.variables.apply_hopsfsmount_apparmor_profile_kube` # { #helm.hopsworks.variables.apply_hopsfsmount_apparmor_profile_kube } : Type `string`, default `"false"`. `hopsworks.variables.async_services_timer_batch_size` # { #helm.hopsworks.variables.async_services_timer_batch_size } : Type `string`, default `"1000"`. `hopsworks.variables.async_services_timer_delete_history_after_days` # { #helm.hopsworks.variables.async_services_timer_delete_history_after_days } : Type `string`, default `"7"`. `hopsworks.variables.async_services_timer_enabled` # { #helm.hopsworks.variables.async_services_timer_enabled } : Type `string`, default `"true"`. `hopsworks.variables.async_services_timer_interval_ms` # { #helm.hopsworks.variables.async_services_timer_interval_ms } : Type `string`, default `"15000"`. `hopsworks.variables.audit_log_count` # { #helm.hopsworks.variables.audit_log_count } : Type `string`, default `"10"`. `hopsworks.variables.audit_log_file_format` # { #helm.hopsworks.variables.audit_log_file_format } : Type `string`, default `"server_audit_log%g.log"`. `hopsworks.variables.audit_log_file_path` # { #helm.hopsworks.variables.audit_log_file_path } : Type `string`, default `"/audit-logs"`. `hopsworks.variables.audit_log_file_type` # { #helm.hopsworks.variables.audit_log_file_type } : Type `string`, default `"io.hops.hopsworks.audit.helper.JSONLogFormatter"`. `hopsworks.variables.audit_log_size_limit` # { #helm.hopsworks.variables.audit_log_size_limit } : Type `string`, default `"256000000"`. `hopsworks.variables.base_buildkit_image` # { #helm.hopsworks.variables.base_buildkit_image } : Type `string`, default `"docker.hops.works/hopsworks/moby/buildkit:v0.32.2-rootless"`. `hopsworks.variables.base_image_name` # { #helm.hopsworks.variables.base_image_name } : Type `string`, default `"hopsworks-base"`. `hopsworks.variables.base_image_version` # { #helm.hopsworks.variables.base_image_version } : Type `string`, default `"5.2.0-SNAPSHOT"`. `hopsworks.variables.cert_mater_delay` # { #helm.hopsworks.variables.cert_mater_delay } : Type `string`, default `"3m"`. `hopsworks.variables.certs_dir` # { #helm.hopsworks.variables.certs_dir } : Type `string`, default `"/srv/hops/certs-dir"`. `hopsworks.variables.check_nodemanagers_status` # { #helm.hopsworks.variables.check_nodemanagers_status } : Type `string`, default `""`. `hopsworks.variables.client_path` # { #helm.hopsworks.variables.client_path } : Type `string`, default `"/srv/hops/clients-3.4.3"`. `hopsworks.variables.cloud` # { #helm.hopsworks.variables.cloud } : Type `string`, default `""`. `hopsworks.variables.command_agent_home_batch` # { #helm.hopsworks.variables.command_agent_home_batch } : Type `string`, default `"20"`. `hopsworks.variables.command_agent_home_claim_lease_as_ms` # { #helm.hopsworks.variables.command_agent_home_claim_lease_as_ms } : Type `string`, default `"600000"`. `hopsworks.variables.command_agent_home_migration_period_as_ms` # { #helm.hopsworks.variables.command_agent_home_migration_period_as_ms } : Type `string`, default `"3600000"`. `hopsworks.variables.command_agent_home_process_timer_period_as_ms` # { #helm.hopsworks.variables.command_agent_home_process_timer_period_as_ms } : Type `string`, default `"5000"`. `hopsworks.variables.command_agent_home_retry_backoff_base_as_ms` # { #helm.hopsworks.variables.command_agent_home_retry_backoff_base_as_ms } : Type `string`, default `"10000"`. `hopsworks.variables.command_agent_home_retry_backoff_max_as_ms` # { #helm.hopsworks.variables.command_agent_home_retry_backoff_max_as_ms } : Type `string`, default `"600000"`. `hopsworks.variables.command_search_fs_history_clean_period_as_ms` # { #helm.hopsworks.variables.command_search_fs_history_clean_period_as_ms } : Type `string`, default `"3600000"`. `hopsworks.variables.command_search_fs_history_enable` # { #helm.hopsworks.variables.command_search_fs_history_enable } : Type `string`, default `"false"`. `hopsworks.variables.command_search_fs_history_window_as_s` # { #helm.hopsworks.variables.command_search_fs_history_window_as_s } : Type `string`, default `"3600"`. `hopsworks.variables.command_search_fs_process_timer_period_as_ms` # { #helm.hopsworks.variables.command_search_fs_process_timer_period_as_ms } : Type `string`, default `"1000"`. `hopsworks.variables.command_search_fs_reindex_queue_wait_as_ms` # { #helm.hopsworks.variables.command_search_fs_reindex_queue_wait_as_ms } : Type `string`, default `"1800000"`. How long a featurestore search reindex run waits for the search command queue to empty and for the featurestore index template to be installed before it is aborted, in milliseconds. The reindex empties the index first, so it only starts on an empty queue, and the template gives the new index its mappings. A run aborted this way is reported under Cluster Settings > Service Operations > OpenSearch Index Commands, where it can be requested again. `hopsworks.variables.command_search_fs_retry_per_clean_interval` # { #helm.hopsworks.variables.command_search_fs_retry_per_clean_interval } : Type `string`, default `"5"`. `hopsworks.variables.conda_default_repo` # { #helm.hopsworks.variables.conda_default_repo } : Type `string`, default `"defaults"`. `hopsworks.variables.default_jupyter_environment` # { #helm.hopsworks.variables.default_jupyter_environment } : Type `string`, default `"pandas-training-pipeline"`. `hopsworks.variables.default_python_job_environment` # { #helm.hopsworks.variables.default_python_job_environment } : Type `string`, default `"pandas-training-pipeline"`. `hopsworks.variables.disable_password_login` # { #helm.hopsworks.variables.disable_password_login } : Type `string`, default `"false"`. `hopsworks.variables.disable_registration` # { #helm.hopsworks.variables.disable_registration } : Type `string`, default `"false"`. `hopsworks.variables.dlt_schema_fetch_job_deadline_seconds` # { #helm.hopsworks.variables.dlt_schema_fetch_job_deadline_seconds } : Type `string`, default `"1800"`. activeDeadlineSeconds for dlthub schema-fetch Kubernetes Jobs. Kills schema-fetch pods that never get to run (unschedulable, volume mount failures), which would otherwise be reported as in-progress forever. `hopsworks.variables.docker_base_image_python_version` # { #helm.hopsworks.variables.docker_base_image_python_version } : Type `string`, default `"3.13"`. `hopsworks.variables.docker_cgroup_cpu_period` # { #helm.hopsworks.variables.docker_cgroup_cpu_period } : Type `string`, default `"100000"`. `hopsworks.variables.docker_cgroup_enabled` # { #helm.hopsworks.variables.docker_cgroup_enabled } : Type `string`, default `"false"`. `hopsworks.variables.docker_cgroup_parent` # { #helm.hopsworks.variables.docker_cgroup_parent } : Type `string`, default `"docker.slice"`. `hopsworks.variables.docker_job_mounts_allowed` # { #helm.hopsworks.variables.docker_job_mounts_allowed } : Type `string`, default `"false"`. `hopsworks.variables.docker_job_mounts_list` # { #helm.hopsworks.variables.docker_job_mounts_list } : Type `string`, default `""`. `hopsworks.variables.docker_job_uid_strict` # { #helm.hopsworks.variables.docker_job_uid_strict } : Type `string`, default `"true"`. `hopsworks.variables.docker_mounts` # { #helm.hopsworks.variables.docker_mounts } : Type `string`. ??? note "Default" ```yaml /srv/hops/hadoop/etc/hadoop,/srv/hops/spark,/srv/hops/flink,/srv/hops/apache-livy ``` `hopsworks.variables.docker_operations_allow_hermetic_custom_commands` # { #helm.hopsworks.variables.docker_operations_allow_hermetic_custom_commands } : Type `string`, default `"false"`. Whether a custom-commands build may declare itself hermetic, with HOPSWORKS_BUILD_HERMETIC=true in its environment file, and so keep its layer cache. Custom command layers are never reused otherwise, because the script can fetch anything and nothing declares what. Only the script's author knows whether that is true of their script, and only the operator decides whether that claim is allowed to control cache reuse. `hopsworks.variables.docker_operations_backoff_limit` # { #helm.hopsworks.variables.docker_operations_backoff_limit } : Type `string`, default `"0"`. `hopsworks.variables.docker_operations_build_metadata` # { #helm.hopsworks.variables.docker_operations_build_metadata } : Type `string`, default `"true"`. Capture the package list, environment export and pip check inside the image build so one post-build job reads them instead of three recomputing them. Images built before this existed fall back to the job-based path automatically. `hopsworks.variables.docker_operations_buildkit_addr` # { #helm.hopsworks.variables.docker_operations_buildkit_addr } : Type `string`, default `""`. Address of the persistent BuildKit daemon, e.g. tcp://buildkitd-0.buildkitd.hopsworks.svc.cluster.local:1234. Empty starts a private daemon inside each build job, which re-pulls and re-unpacks the base image every time. Leave empty and set global._hopsworks.buildkitd.enabled: the address of the daemon the chart deploys is filled in from buildkitd.name, buildkitd.port and the release namespace. Only set this to point builds at a daemon the chart does not manage. The example is a full pod DNS name rather than a short one because the client verifies the hostname it dials against the certificate: the chart's SANs cover the service and the per-pod names, not a bare "buildkitd", so a short-name address fails verification with mTLS on. `hopsworks.variables.docker_operations_buildkit_backoff_limit` # { #helm.hopsworks.variables.docker_operations_buildkit_backoff_limit } : Type `string`, default `"0"`. `hopsworks.variables.docker_operations_buildkit_cache_scope` # { #helm.hopsworks.variables.docker_operations_buildkit_cache_scope } : Type `string`, default `"shared"`. Scope of the package cache shared between builds: "off", "project" or "shared". Constrained by a pattern rather than an enum: helm-schema infers type: string from the non-empty default and then refuses enum and type together, while pattern coexists with it and rejects the same set of values. The backend still validates at runtime and falls back to "off", since an operator can set this in the variables table without going through the chart. "project" is safe for multi-tenant here because users cannot inject Dockerfile directives: the Dockerfile is generated by the backend, and custom commands supply a shell script that runs inside a RUN. "shared" gives every project one cache and is single-trust-zone only. This is also the off switch for the toolchain caches a custom-commands build can ask for with HOPSWORKS_BUILD_CACHE (uv, pip, ccache, sccache, maven, gradle, cargo, npm, go). Those are always scoped to the project whatever this is set to, since a script controls what goes into them. Anything other than "off" enables them. "shared" is the default because the package cache is scoped by index configuration, not by a single cluster-wide id: builds that resolve through the same configuration share, and a project using different index credentials gets a different cache. Every reusable build step also carries a per-project cache-key tag, so a layer is never reused across projects and rotating a credential invalidates reuse, which BuildKit does not do on its own because it leaves secret contents out of cache keys. `hopsworks.variables.docker_operations_buildkit_extra_args` # { #helm.hopsworks.variables.docker_operations_buildkit_extra_args } : Type `string`, default `""`. Extra arguments appended to every buildctl invocation. Empty by default. This previously shipped an S3 --export-cache/--import-cache pair. BuildKit resolves S3 credentials at the daemon rather than at the client, so whether it worked depended entirely on what identity the daemon had. On an EKS install with defaultServiceAccount annotations wired for IRSA (see values.aws.yaml) the daemonless build pod carried a web-identity token and the exporter worked. On a cluster with no AWS identity it could not: observed as "no EC2 IMDS role found ... context deadline exceeded" on every build, costing an IMDS timeout per build while caching nothing, with ignore-error hiding the failure rather than avoiding it. Defaulting it empty therefore removes a remote cache that some installs did have. That is deliberate, because it failed closed and expensively everywhere else, but it is a behaviour change on upgrade rather than the removal of something inert. Before re-enabling it, give the daemon real credentials and confirm the cache is actually being read. With the daemon enabled it is no longer the build pod, so the build pod's identity no longer applies to it: set buildkitd.serviceAccountName to an account carrying the cloud identity you want the exporter to use. A persistent daemon already avoids the base image re-pull this was reaching for, and does so without leaving the cluster. `hopsworks.variables.docker_operations_buildkit_limit_cpu` # { #helm.hopsworks.variables.docker_operations_buildkit_limit_cpu } : Type `string`, default `"2"`. `hopsworks.variables.docker_operations_buildkit_limit_memory` # { #helm.hopsworks.variables.docker_operations_buildkit_limit_memory } : Type `string`, default `"4G"`. `hopsworks.variables.docker_operations_buildkit_priority_class` # { #helm.hopsworks.variables.docker_operations_buildkit_priority_class } : Type `string`, default `""`. `hopsworks.variables.docker_operations_buildkit_replicas` # { #helm.hopsworks.variables.docker_operations_buildkit_replicas } : Type `string`, default `"1"`. Replica count behind a "%d" placeholder in docker_operations_buildkit_addr, e.g. tcp://buildkitd-%d.buildkitd.hopsworks.svc.cluster.local:1234 with replicas 3. Filled in from buildkitd.replicas unless set here. This spreads load and cache state across daemons. It is not high availability: a project is pinned to one replica by id and is not retried against another, so a project whose daemon is down waits for it to come back. It is not what separates tenants either; that is mTLS plus the per-project cache key. `hopsworks.variables.docker_operations_buildkit_request_cpu` # { #helm.hopsworks.variables.docker_operations_buildkit_request_cpu } : Type `string`, default `"200m"`. `hopsworks.variables.docker_operations_buildkit_request_memory` # { #helm.hopsworks.variables.docker_operations_buildkit_request_memory } : Type `string`, default `"500Mi"`. `hopsworks.variables.docker_operations_buildkit_storage` # { #helm.hopsworks.variables.docker_operations_buildkit_storage } : Type `string`, default `"70Gi"`. `hopsworks.variables.docker_operations_buildkit_tls_locality` # { #helm.hopsworks.variables.docker_operations_buildkit_tls_locality } : Type `string`, default `"buildkitd"`. Certificate locality, which is what names the key and certificate files inside that Secret. Must match buildkitd.tls.locality, and is filled in from it. `hopsworks.variables.docker_operations_buildkit_tls_secret` # { #helm.hopsworks.variables.docker_operations_buildkit_tls_secret } : Type `string`, default `""`. Secret holding the client certificate the build job presents to the persistent daemon. Empty means the client sends none, which only works against a daemon that requires no client certificate. Filled in from buildkitd.tls when that is enabled. `hopsworks.variables.docker_operations_cert_name` # { #helm.hopsworks.variables.docker_operations_cert_name } : Type `string`, default `"kagent_certificate_bundle.pem"`. `hopsworks.variables.docker_operations_context_orphan_minutes` # { #helm.hopsworks.variables.docker_operations_context_orphan_minutes } : Type `string`, default `"120"`. Age in minutes after which a build context in S3 whose build no longer exists is deleted. Covers builds that died with Payara or their node, which the per-build cleanup cannot. `hopsworks.variables.docker_operations_default_service_account` # { #helm.hopsworks.variables.docker_operations_default_service_account } : Type `string`, default `"default"`. `hopsworks.variables.docker_operations_delete_jobs_add_description_if_fails` # { #helm.hopsworks.variables.docker_operations_delete_jobs_add_description_if_fails } : Type `bool`, default `false`. `hopsworks.variables.docker_operations_delete_jobs_on_completion` # { #helm.hopsworks.variables.docker_operations_delete_jobs_on_completion } : Type `string`, default `"true"`. `hopsworks.variables.docker_operations_docker_context_builder` # { #helm.hopsworks.variables.docker_operations_docker_context_builder } : Type `string`, default `"AUTO"`. `hopsworks.variables.docker_operations_docker_context_builder_s3_bucket` # { #helm.hopsworks.variables.docker_operations_docker_context_builder_s3_bucket } : Type `string`, default `"${env:S3_BUCKET}"`. `hopsworks.variables.docker_operations_docker_context_builder_s3_endpoint` # { #helm.hopsworks.variables.docker_operations_docker_context_builder_s3_endpoint } : Type `string`, default `"${env:S3_ENDPOINT}"`. `hopsworks.variables.docker_operations_docker_context_builder_s3_region` # { #helm.hopsworks.variables.docker_operations_docker_context_builder_s3_region } : Type `string`, default `"${env:S3_REGION}"`. `hopsworks.variables.docker_operations_hopsworks_ca_secret_name` # { #helm.hopsworks.variables.docker_operations_hopsworks_ca_secret_name } : Type `string`, default `"docker-registry-crypto-material"`. `hopsworks.variables.docker_operations_image_pull_secrets` # { #helm.hopsworks.variables.docker_operations_image_pull_secrets } : Type `string`, default `""`. `hopsworks.variables.docker_operations_lock_dependencies` # { #helm.hopsworks.variables.docker_operations_lock_dependencies } : Type `string`, default `"false"`. Resolve a full dependency set with per-artifact hashes and install only from it. Makes the resolved set reconstructible and fails the build if an index serves different bytes for a version it already served. Needs uv in the base image; builds without it fall back. `hopsworks.variables.docker_operations_managed_docker_secrets` # { #helm.hopsworks.variables.docker_operations_managed_docker_secrets } : Type `string`, default `""`. `hopsworks.variables.docker_operations_multi_region_copy` # { #helm.hopsworks.variables.docker_operations_multi_region_copy } : Type `string`, default `"false"`. On a multi-region install, copy the built image to the secondary region instead of running the build again there. Building twice does the work twice and is not guaranteed to land the same image: a custom command or an unpinned package can resolve differently between the two runs, leaving the regions with different content under one tag. Off by default because the copy needs pull and push credentials for both registries in a single job. `hopsworks.variables.docker_operations_oci_worker_snapshotter` # { #helm.hopsworks.variables.docker_operations_oci_worker_snapshotter } : Type `string`, default `"auto"`. `hopsworks.variables.docker_operations_push_insecure` # { #helm.hopsworks.variables.docker_operations_push_insecure } : Type `string`, default `"false"`. `hopsworks.variables.docker_operations_registry_container` # { #helm.hopsworks.variables.docker_operations_registry_container } : Type `string`, default `"docker"`. `hopsworks.variables.docker_operations_registry_http` # { #helm.hopsworks.variables.docker_operations_registry_http } : Type `string`, default `"false"`. `hopsworks.variables.docker_operations_registry_pod` # { #helm.hopsworks.variables.docker_operations_registry_pod } : Type `string`, default `"docker-registry-0"`. `hopsworks.variables.docker_operations_suspend_jobs` # { #helm.hopsworks.variables.docker_operations_suspend_jobs } : Type `string`, default `"false"`. `hopsworks.variables.docker_operations_timeout_check_minutes` # { #helm.hopsworks.variables.docker_operations_timeout_check_minutes } : Type `string`, default `"15"`. `hopsworks.variables.docker_operations_timeout_delete_minutes` # { #helm.hopsworks.variables.docker_operations_timeout_delete_minutes } : Type `string`, default `"5"`. `hopsworks.variables.docker_operations_timeout_export_minutes` # { #helm.hopsworks.variables.docker_operations_timeout_export_minutes } : Type `string`, default `"15"`. `hopsworks.variables.docker_operations_timeout_listing_minutes` # { #helm.hopsworks.variables.docker_operations_timeout_listing_minutes } : Type `string`, default `"15"`. `hopsworks.variables.docker_operations_timeout_minutes_buildkit` # { #helm.hopsworks.variables.docker_operations_timeout_minutes_buildkit } : Type `string`, default `"120"`. `hopsworks.variables.docker_operations_timeout_tag_minutes` # { #helm.hopsworks.variables.docker_operations_timeout_tag_minutes } : Type `string`, default `"5"`. `hopsworks.variables.download_allowed` # { #helm.hopsworks.variables.download_allowed } : Type `string`, default `"true"`. `hopsworks.variables.elastic_dir` # { #helm.hopsworks.variables.elastic_dir } : Type `string`, default `"/srv/hops/elastic"`. `hopsworks.variables.elastic_https_enabled` # { #helm.hopsworks.variables.elastic_https_enabled } : Type `string`, default `"true"`. `hopsworks.variables.elastic_jwt_enabled` # { #helm.hopsworks.variables.elastic_jwt_enabled } : Type `string`, default `"true"`. `hopsworks.variables.elastic_jwt_exp_ms` # { #helm.hopsworks.variables.elastic_jwt_exp_ms } : Type `string`, default `"1800000"`. `hopsworks.variables.elastic_jwt_url_parameter` # { #helm.hopsworks.variables.elastic_jwt_url_parameter } : Type `string`, default `"jt"`. `hopsworks.variables.elastic_logs_index_expiration` # { #helm.hopsworks.variables.elastic_logs_index_expiration } : Type `string`, default `"604800000"`. `hopsworks.variables.elastic_opendistro_security_enabled` # { #helm.hopsworks.variables.elastic_opendistro_security_enabled } : Type `string`, default `"true"`. `hopsworks.variables.elastic_user` # { #helm.hopsworks.variables.elastic_user } : Type `string`, default `"elastic"`. `hopsworks.variables.elastic_version` # { #helm.hopsworks.variables.elastic_version } : Type `string`, default `"3.8.0"`. `hopsworks.variables.enable_adls_storage_connectors` # { #helm.hopsworks.variables.enable_adls_storage_connectors } : Type `string`, default `"false"`. `hopsworks.variables.enable_bigquery_storage_connectors` # { #helm.hopsworks.variables.enable_bigquery_storage_connectors } : Type `string`, default `"true"`. `hopsworks.variables.enable_bring_your_own_kafka` # { #helm.hopsworks.variables.enable_bring_your_own_kafka } : Type `string`, default `"false"`. `hopsworks.variables.enable_feature_monitoring` # { #helm.hopsworks.variables.enable_feature_monitoring } : Type `string`, default `"true"`. `hopsworks.variables.enable_fix_receivers_timer` # { #helm.hopsworks.variables.enable_fix_receivers_timer } : Type `string`, default `"true"`. `hopsworks.variables.enable_gcs_storage_connectors` # { #helm.hopsworks.variables.enable_gcs_storage_connectors } : Type `string`, default `"true"`. `hopsworks.variables.enable_jupyter_python_kernel_non_kubernetes` # { #helm.hopsworks.variables.enable_jupyter_python_kernel_non_kubernetes } : Type `string`, default `"false"`. `hopsworks.variables.enable_kafka_storage_connectors` # { #helm.hopsworks.variables.enable_kafka_storage_connectors } : Type `string`, default `"true"`. `hopsworks.variables.enable_metadata_designer` # { #helm.hopsworks.variables.enable_metadata_designer } : Type `string`, default `""`. `hopsworks.variables.enable_opensearch_storage_connectors` # { #helm.hopsworks.variables.enable_opensearch_storage_connectors } : Type `string`, default `"true"`. `hopsworks.variables.enable_read_only_git_repositories` # { #helm.hopsworks.variables.enable_read_only_git_repositories } : Type `string`, default `"false"`. `hopsworks.variables.enable_redshift_storage_connectors` # { #helm.hopsworks.variables.enable_redshift_storage_connectors } : Type `string`, default `"true"`. `hopsworks.variables.enable_snowflake_storage_connectors` # { #helm.hopsworks.variables.enable_snowflake_storage_connectors } : Type `string`, default `"true"`. `hopsworks.variables.enable_user_search` # { #helm.hopsworks.variables.enable_user_search } : Type `string`, default `"true"`. `hopsworks.variables.epipe_version` # { #helm.hopsworks.variables.epipe_version } : Type `string`, default `"0.20.0"`. `hopsworks.variables.executions_cleaner_batch_size` # { #helm.hopsworks.variables.executions_cleaner_batch_size } : Type `string`, default `"50"`. `hopsworks.variables.executions_cleaner_interval_ms` # { #helm.hopsworks.variables.executions_cleaner_interval_ms } : Type `string`, default `"600000"`. `hopsworks.variables.executions_per_job_limit` # { #helm.hopsworks.variables.executions_per_job_limit } : Type `string`, default `"10000"`. `hopsworks.variables.feature_monitoring_max_num_features` # { #helm.hopsworks.variables.feature_monitoring_max_num_features } : Type `string`, default `"15"`. `hopsworks.variables.featurestore_asof_spine_max_bytes` # { #helm.hopsworks.variables.featurestore_asof_spine_max_bytes } : Type `string`, default `"1073741824"`. `hopsworks.variables.featurestore_asof_spine_max_columns` # { #helm.hopsworks.variables.featurestore_asof_spine_max_columns } : Type `string`, default `"256"`. `hopsworks.variables.featurestore_asof_spine_max_file_age_ms` # { #helm.hopsworks.variables.featurestore_asof_spine_max_file_age_ms } : Type `string`, default `"86400000"`. `hopsworks.variables.featurestore_asof_spine_max_rows` # { #helm.hopsworks.variables.featurestore_asof_spine_max_rows } : Type `string`, default `"1000000"`. `hopsworks.variables.featurestore_db_admin_user` # { #helm.hopsworks.variables.featurestore_db_admin_user } : Type `string`, default `"featurestore_admin_user"`. `hopsworks.variables.featurestore_default_quota` # { #helm.hopsworks.variables.featurestore_default_quota } : Type `string`, default `"-1L"`. `hopsworks.variables.featurestore_default_storage_format` # { #helm.hopsworks.variables.featurestore_default_storage_format } : Type `string`, default `"PARQUET"`. `hopsworks.variables.featurestore_metrics_enabled` # { #helm.hopsworks.variables.featurestore_metrics_enabled } : Type `string`, default `"true"`. `hopsworks.variables.featurestore_metrics_online_ingestion_enabled` # { #helm.hopsworks.variables.featurestore_metrics_online_ingestion_enabled } : Type `string`, default `"false"`. `hopsworks.variables.featurestore_online_enabled` # { #helm.hopsworks.variables.featurestore_online_enabled } : Type `string`, default `"true"`. `hopsworks.variables.featurestore_online_tablespace` # { #helm.hopsworks.variables.featurestore_online_tablespace } : Type `string`, default `""`. `hopsworks.variables.file_preview_image_size` # { #helm.hopsworks.variables.file_preview_image_size } : Type `string`, default `"10000000"`. `hopsworks.variables.file_preview_txt_size` # { #helm.hopsworks.variables.file_preview_txt_size } : Type `string`, default `"100"`. `hopsworks.variables.flink_dir` # { #helm.hopsworks.variables.flink_dir } : Type `string`, default `"/srv/hops/flink"`. `hopsworks.variables.flink_user` # { #helm.hopsworks.variables.flink_user } : Type `string`, default `"flink"`. `hopsworks.variables.flink_version` # { #helm.hopsworks.variables.flink_version } : Type `string`, default `"1.17.1.0"`. `hopsworks.variables.fs_job_activity_time` # { #helm.hopsworks.variables.fs_job_activity_time } : Type `string`, default `"5m"`. `hopsworks.variables.fs_storage_connector_session_duration` # { #helm.hopsworks.variables.fs_storage_connector_session_duration } : Type `string`, default `"3600"`. `hopsworks.variables.git_bitbucket_http_proxy` # { #helm.hopsworks.variables.git_bitbucket_http_proxy } : Type `string`, default `""`. Proxy for git traffic to BitBucket over HTTP. Empty means a direct connection. `hopsworks.variables.git_bitbucket_https_proxy` # { #helm.hopsworks.variables.git_bitbucket_https_proxy } : Type `string`, default `""`. Proxy for git traffic to BitBucket over HTTPS. Empty means a direct connection. `hopsworks.variables.git_command_timeout_minutes` # { #helm.hopsworks.variables.git_command_timeout_minutes } : Type `string`, default `"60"`. `hopsworks.variables.git_custom_ca_configmap` # { #helm.hopsworks.variables.git_custom_ca_configmap } : Type `string`, default `""`. Name of a ConfigMap holding CA certificates that git operations should trust in addition to the public anchors the git image already ships, so repositories on a self-hosted GitLab / GitHub Enterprise / BitBucket behind a private CA can be cloned and pushed to over HTTPS. The ConfigMap must exist in **every project namespace**, since git runs there, and is expected to be distributed by the platform rather than by this chart. Leave empty to trust only the public anchors. Applies to HTTPS remotes only; SSH host-key verification is unaffected. This value is seeded on every install and upgrade, so edits made in the admin UI do not survive a `helm upgrade` -- configure it here. `hopsworks.variables.git_custom_ca_configmap_key` # { #helm.hopsworks.variables.git_custom_ca_configmap_key } : Type `string`, default `"ca-bundle.crt"`. Key within `git_custom_ca_configmap` holding the certificates. One key, which may hold several concatenated PEM certificates; an admin with several providers puts all their CAs in it. Ignored when `git_custom_ca_configmap` is empty, and must not be blank when it is set. `hopsworks.variables.git_disable_tls_verification` # { #helm.hopsworks.variables.git_disable_tls_verification } : Type `string`, default `"false"`. Disable TLS certificate verification for **all** git remotes, public ones included. Connections can then be intercepted without detection, so prefer naming the CA to trust in `git_custom_ca_configmap`; this exists for deployments where that is not workable. Applies cluster-wide, not per provider. Seeded on every install and upgrade, like the other git variables. `hopsworks.variables.git_github_http_proxy` # { #helm.hopsworks.variables.git_github_http_proxy } : Type `string`, default `""`. Proxy for git traffic to GitHub over HTTP, e.g. `http://proxy.corp:3128`. Empty means a direct connection. Configured per provider because a deployment may reach each of them by a different route. Traffic between the git container and Hopsworks itself never goes through these proxies. Seeded on every install and upgrade, like the other git variables. `hopsworks.variables.git_github_https_proxy` # { #helm.hopsworks.variables.git_github_https_proxy } : Type `string`, default `""`. Proxy for git traffic to GitHub over HTTPS. This is the one that matters in practice, since Hopsworks drives git over HTTPS remotes. `hopsworks.variables.git_gitlab_http_proxy` # { #helm.hopsworks.variables.git_gitlab_http_proxy } : Type `string`, default `""`. Proxy for git traffic to GitLab over HTTP. Empty means a direct connection. `hopsworks.variables.git_gitlab_https_proxy` # { #helm.hopsworks.variables.git_gitlab_https_proxy } : Type `string`, default `""`. Proxy for git traffic to GitLab over HTTPS. Empty means a direct connection. `hopsworks.variables.git_image_version` # { #helm.hopsworks.variables.git_image_version } : Type `string`, default `"1.6-SNAPSHOT"`. `hopsworks.variables.grafana_version` # { #helm.hopsworks.variables.grafana_version } : Type `string`, default `"9.3.16"`. `hopsworks.variables.ha_enabled` # { #helm.hopsworks.variables.ha_enabled } : Type `string`, default `"true"`. `hopsworks.variables.hadoop_dir` # { #helm.hopsworks.variables.hadoop_dir } : Type `string`, default `"/srv/hops/hadoop"`. `hopsworks.variables.hadoop_version` # { #helm.hopsworks.variables.hadoop_version } : Type `string`, default `"3.4.3.3-EE-RC1"`. `hopsworks.variables.hdfs_base_storage_policy` # { #helm.hopsworks.variables.hdfs_base_storage_policy } : Type `string`, default `"CLOUD"`. `hopsworks.variables.hdfs_default_quota` # { #helm.hopsworks.variables.hdfs_default_quota } : Type `string`, default `"-1L"`. `hopsworks.variables.hdfs_log_storage_policy` # { #helm.hopsworks.variables.hdfs_log_storage_policy } : Type `string`, default `"CLOUD"`. `hopsworks.variables.hdfs_user` # { #helm.hopsworks.variables.hdfs_user } : Type `string`, default `"hdfs"`. `hopsworks.variables.hdfscontentsmanager_base_hopsfs_client` # { #helm.hopsworks.variables.hdfscontentsmanager_base_hopsfs_client } : Type `string`, default `"libhdfs-go"`. `hopsworks.variables.hive2_version` # { #helm.hopsworks.variables.hive2_version } : Type `string`, default `"4.1.0.0-v1"`. `hopsworks.variables.hive_conf_path` # { #helm.hopsworks.variables.hive_conf_path } : Type `string`, default `"/srv/hops/apache-hive/conf"`. `hopsworks.variables.hive_superuser` # { #helm.hopsworks.variables.hive_superuser } : Type `string`, default `"hive"`. `hopsworks.variables.hive_warehouse` # { #helm.hopsworks.variables.hive_warehouse } : Type `string`, default `"/apps/hive/warehouse"`. `hopsworks.variables.hiveserver_ext_hostname` # { #helm.hopsworks.variables.hiveserver_ext_hostname } : Type `string`, default `""`. `hopsworks.variables.hiveserver_ssl_hostname` # { #helm.hopsworks.variables.hiveserver_ssl_hostname } : Type `string`, default `""`. `hopsworks.variables.hops_db` # { #helm.hopsworks.variables.hops_db } : Type `string`, default `"hops"`. `hopsworks.variables.hops_rpc_tls` # { #helm.hopsworks.variables.hops_rpc_tls } : Type `string`, default `"true"`. `hopsworks.variables.hopsexamples_version` # { #helm.hopsworks.variables.hopsexamples_version } : Type `string`, default `""`. `hopsworks.variables.hopsfsmount_apparmor_profile` # { #helm.hopsworks.variables.hopsfsmount_apparmor_profile } : Type `string`, default `""`. `hopsworks.variables.hopsfsmount_log_level` # { #helm.hopsworks.variables.hopsfsmount_log_level } : Type `string`, default `"warn"`. `hopsworks.variables.hopsfsmount_nn_connections` # { #helm.hopsworks.variables.hopsfsmount_nn_connections } : Type `string`, default `"4"`. `hopsworks.variables.hopsworks_analytics` # { #helm.hopsworks.variables.hopsworks_analytics } : Type `bool`, default `false`. When true, the backend creates a reserved 'hopsworks_analytics' project whose members are the cluster admins (HOPS_ADMIN), and attaches a read-only SQL (MySQL) data source onto the local 'hopsworks' database (user hopsworks_analytics_ro_user) to the 'hopsworks_analytics' project itself. Clearing it withdraws the read-only database account, so the two keys must stay required together. Enabling this on a cluster with create_secrets=false additionally requires adding a key named after hopsworks_analytics_ro_user to the hand-created hopsworks-users-secrets, holding the read-only account's password. This value is the only supported switch. The read-only account, its password and the grants are provisioned by the chart, so flipping the 'hopsworks_analytics' row from the admin variables page creates the project and its members but no data source, and the next upgrade sets the row back to this value. What the read-only account can read: a curated SELECT allow-list of metadata tables (see roTables in dml/grants.sql.template), not the whole database and never feature data. Several of those tables carry user identities as email addresses: project.username, project_team.team_member, jobs.creator, executions.user and dataset_request.user_email. Anyone with the shared Superset connection, which is every active HOPS_ADMIN, can read them, and the dashboard setup can sample rows from mounted tables into the configured model provider. Treat the account as exposing cluster users' addresses to cluster admins and to that provider. `hopsworks.variables.hopsworks_analytics_coding_agent` # { #helm.hopsworks.variables.hopsworks_analytics_coding_agent } : Type `string`, default `"claude"`. Which coding agent the Setup Analytics wizard runs: "claude" (Claude Code), "codex" (OpenAI Codex CLI), "copilot" (GitHub Copilot CLI) or "opencode". The wizard uses this without asking, so setting it is how a cluster standardises on one agent; an advanced option in the wizard lets a user run a different one for their own session, which does not write back here. All four are installed in the terminal image under exactly these command names, so the value is the command. Constrained by a pattern rather than an enum for the reason given on docker_operations_package_cache_scope. The frontend falls back to "claude" when the row is missing or holds a value it does not recognise. `hopsworks.variables.hopsworks_analytics_ro_user` # { #helm.hopsworks.variables.hopsworks_analytics_ro_user } : Type `string`, default `"hopsworks_ro"`. `hopsworks.variables.hopsworks_analytics_ro_user_adopt_existing` # { #helm.hopsworks.variables.hopsworks_analytics_ro_user_adopt_existing } : Type `bool`, default `false`. Take over a database account that already exists under hopsworks_analytics_ro_user but was not created by this chart. Off by default, and the install fails with the account name rather than adopting it, because adoption rewrites the account's password, strips its privileges and replaces them with the analytics allow-list. Set it to true only when that account is yours to hand over. The chart records the account it provisioned in the hopsworks_analytics_ro_user_managed variable and withdraws only that one. `hopsworks.variables.hopsworks_analytics_setup_repo` # { #helm.hopsworks.variables.hopsworks_analytics_setup_repo } : Type `string`, default `"https://github.com/logicalclocks/okr-dashboards"`. Repository the Setup Analytics wizard clones when the terminal image does not already carry it, so a cluster that cannot reach github.com can point the flow at a reachable mirror instead of failing in the Terminal with a git error. The checkout directory is derived from the last path segment. Constrained to characters that cannot change the meaning of the shell line the URL is spliced into; the frontend falls back to the default when the row is missing or fails that check. `hopsworks.variables.hopsworks_db` # { #helm.hopsworks.variables.hopsworks_db } : Type `string`, default `"hopsworks"`. `hopsworks.variables.hopsworks_dir` # { #helm.hopsworks.variables.hopsworks_dir } : Type `string`, default `"/srv/hops/domains/domain1"`. `hopsworks.variables.hopsworks_enterprise` # { #helm.hopsworks.variables.hopsworks_enterprise } : Type `string`, default `"true"`. `hopsworks.variables.hopsworks_mysql_user` # { #helm.hopsworks.variables.hopsworks_mysql_user } : Type `string`, default `"hopsworks"`. `hopsworks.variables.hopsworks_public_proxy_url` # { #helm.hopsworks.variables.hopsworks_public_proxy_url } : Type `string`, default `""`. `hopsworks.variables.hopsworks_rest_log_level` # { #helm.hopsworks.variables.hopsworks_rest_log_level } : Type `string`, default `"TEST"`. `hopsworks.variables.hopsworks_user` # { #helm.hopsworks.variables.hopsworks_user } : Type `string`, default `"payara"`. `hopsworks.variables.hw_group_mapping_sync_enabled` # { #helm.hopsworks.variables.hw_group_mapping_sync_enabled } : Type `string`, default `"false"`. `hopsworks.variables.ingestion_job_cores` # { #helm.hopsworks.variables.ingestion_job_cores } : Type `string`, default `"1.0"`. `hopsworks.variables.ingestion_job_gpus` # { #helm.hopsworks.variables.ingestion_job_gpus } : Type `string`, default `"0"`. `hopsworks.variables.ingestion_job_memory` # { #helm.hopsworks.variables.ingestion_job_memory } : Type `string`, default `"2048"`. `hopsworks.variables.java_home` # { #helm.hopsworks.variables.java_home } : Type `string`, default `""`. `hopsworks.variables.job_name_validation_regex` # { #helm.hopsworks.variables.job_name_validation_regex } : Type `string`, default `"^[a-zA-Z0-9_\\-]+$"`. `hopsworks.variables.jupyter_allow_no_limit_shutdown` # { #helm.hopsworks.variables.jupyter_allow_no_limit_shutdown } : Type `bool`, default `true`. `hopsworks.variables.jupyter_dir` # { #helm.hopsworks.variables.jupyter_dir } : Type `string`, default `"/srv/hops/jupyter"`. `hopsworks.variables.jupyter_group` # { #helm.hopsworks.variables.jupyter_group } : Type `string`, default `"hadoop"`. `hopsworks.variables.jupyter_hour_shutdown_options` # { #helm.hopsworks.variables.jupyter_hour_shutdown_options } : Type `string`, default `"8,24,72"`. `hopsworks.variables.jupyter_origin_scheme` # { #helm.hopsworks.variables.jupyter_origin_scheme } : Type `string`, default `"https"`. `hopsworks.variables.jupyter_shell_command` # { #helm.hopsworks.variables.jupyter_shell_command } : Type `string`. ??? note "Default" ```yaml '["/bin/bash", "--login", "-c", "cd -L $JUPYTER_DATA_DIR || true && exec bash"]' ``` `hopsworks.variables.jupyter_shutdown_timer_interval` # { #helm.hopsworks.variables.jupyter_shutdown_timer_interval } : Type `string`, default `"1m"`. `hopsworks.variables.jupyter_spark_notebook_server_memory_floor_mb` # { #helm.hopsworks.variables.jupyter_spark_notebook_server_memory_floor_mb } : Type `string`, default `"512"`. `hopsworks.variables.jupyter_ws_ping_interval` # { #helm.hopsworks.variables.jupyter_ws_ping_interval } : Type `string`, default `"10s"`. `hopsworks.variables.jwt_exp_leeway_sec` # { #helm.hopsworks.variables.jwt_exp_leeway_sec } : Type `string`, default `"900"`. `hopsworks.variables.jwt_issuer` # { #helm.hopsworks.variables.jwt_issuer } : Type `string`, default `"hopsworks@logicalclocks.com"`. `hopsworks.variables.jwt_lifetime_ms` # { #helm.hopsworks.variables.jwt_lifetime_ms } : Type `string`, default `"86400000"`. `hopsworks.variables.jwt_signature_algorithm` # { #helm.hopsworks.variables.jwt_signature_algorithm } : Type `string`, default `"HS512"`. `hopsworks.variables.jwt_signing_key_name` # { #helm.hopsworks.variables.jwt_signing_key_name } : Type `string`, default `"apiKey"`. `hopsworks.variables.kafka_installed` # { #helm.hopsworks.variables.kafka_installed } : Type `bool`, default `true`. `hopsworks.variables.kafka_max_num_topics` # { #helm.hopsworks.variables.kafka_max_num_topics } : Type `string`, default `"100"`. `hopsworks.variables.kafka_num_partitions` # { #helm.hopsworks.variables.kafka_num_partitions } : Type `string`, default `"1"`. `hopsworks.variables.kafka_num_replicas` # { #helm.hopsworks.variables.kafka_num_replicas } : Type `string`, default `"1"`. `hopsworks.variables.kafka_user` # { #helm.hopsworks.variables.kafka_user } : Type `string`, default `"kafka"`. `hopsworks.variables.kafka_version` # { #helm.hopsworks.variables.kafka_version } : Type `string`, default `"4.3.1"`. `hopsworks.variables.kibana_https_enabled` # { #helm.hopsworks.variables.kibana_https_enabled } : Type `string`, default `"true"`. `hopsworks.variables.kibana_multi_tenancy_enabled` # { #helm.hopsworks.variables.kibana_multi_tenancy_enabled } : Type `string`, default `"true"`. `hopsworks.variables.kibana_version` # { #helm.hopsworks.variables.kibana_version } : Type `string`, default `"3.8.0"`. `hopsworks.variables.kube_api_max_attempts` # { #helm.hopsworks.variables.kube_api_max_attempts } : Type `string`, default `"20"`. `hopsworks.variables.kube_hopsworks_default_service_account` # { #helm.hopsworks.variables.kube_hopsworks_default_service_account } : Type `string`, default `"hopsworks-default"`. `hopsworks.variables.kube_knative_domain_name` # { #helm.hopsworks.variables.kube_knative_domain_name } : Type `string`, default `"hopsworks.ai"`. `hopsworks.variables.kube_knative_lb_domain` # { #helm.hopsworks.variables.kube_knative_lb_domain } : Type `string`, default `""`. `hopsworks.variables.kube_kserve_installed` # { #helm.hopsworks.variables.kube_kserve_installed } : Type `bool`, default `true`. `hopsworks.variables.kube_kserve_tensorflow_version` # { #helm.hopsworks.variables.kube_kserve_tensorflow_version } : Type `string`, default `"2.20.0"`. `hopsworks.variables.kube_node_taints_monitor_interval` # { #helm.hopsworks.variables.kube_node_taints_monitor_interval } : Type `string`, default `"10m"`. `hopsworks.variables.kube_scheduling_hopsfsmount_cpu_limits` # { #helm.hopsworks.variables.kube_scheduling_hopsfsmount_cpu_limits } : Type `int`, default `-1`. `hopsworks.variables.kube_scheduling_hopsfsmount_cpu_requests` # { #helm.hopsworks.variables.kube_scheduling_hopsfsmount_cpu_requests } : Type `int`, default `1`. `hopsworks.variables.kube_scheduling_hopsfsmount_memory_limits_mb` # { #helm.hopsworks.variables.kube_scheduling_hopsfsmount_memory_limits_mb } : Type `int`, default `1024`. `hopsworks.variables.kube_scheduling_jobinit_cpu_limits` # { #helm.hopsworks.variables.kube_scheduling_jobinit_cpu_limits } : Type `int`, default `-1`. `hopsworks.variables.kube_scheduling_jobinit_cpu_requests` # { #helm.hopsworks.variables.kube_scheduling_jobinit_cpu_requests } : Type `float`, default `0.5`. `hopsworks.variables.kube_scheduling_jobinit_memory_limits_mb` # { #helm.hopsworks.variables.kube_scheduling_jobinit_memory_limits_mb } : Type `int`, default `512`. `hopsworks.variables.kube_scheduling_jobinit_memory_requests_mb` # { #helm.hopsworks.variables.kube_scheduling_jobinit_memory_requests_mb } : Type `int`, default `256`. `hopsworks.variables.kube_serving_max_num_instances` # { #helm.hopsworks.variables.kube_serving_max_num_instances } : Type `string`, default `"10"`. `hopsworks.variables.kube_serving_min_num_instances` # { #helm.hopsworks.variables.kube_serving_min_num_instances } : Type `string`, default `"-1"`. `hopsworks.variables.kube_serving_vllm_omni_versions` # { #helm.hopsworks.variables.kube_serving_vllm_omni_versions } : Type `string`, default `"v0.28.0"`. `hopsworks.variables.kube_serving_vllm_versions` # { #helm.hopsworks.variables.kube_serving_vllm_versions } : Type `string`, default `"v0.28.0"`. `hopsworks.variables.kube_skip_namespace_creation` # { #helm.hopsworks.variables.kube_skip_namespace_creation } : Type `bool`, default `false`. `hopsworks.variables.kube_tainted_nodes` # { #helm.hopsworks.variables.kube_tainted_nodes } : Type `string`, default `""`. `hopsworks.variables.kube_type` # { #helm.hopsworks.variables.kube_type } : Type `string`, default `"kube_cluster"`. `hopsworks.variables.kube_user_workload_tolerations` # { #helm.hopsworks.variables.kube_user_workload_tolerations } : Type `string`, default `""`. Kubernetes tolerations applied to all user-triggered compute pods. Comma-separated entries of the form key\[=value\]\[:effect\] matching kubectl taint grammar. Effect is optional (empty matches any effect). Example: 'nvidia.com/gpu=true:NoSchedule,dedicated=tenant-a:NoSchedule' `hopsworks.variables.kubernetes_installed` # { #helm.hopsworks.variables.kubernetes_installed } : Type `string`, default `"true"`. `hopsworks.variables.kueue_project_default_cluster_queue` # { #helm.hopsworks.variables.kueue_project_default_cluster_queue } : Type `string`, default `"other"`. `hopsworks.variables.kueue_project_default_local_queue` # { #helm.hopsworks.variables.kueue_project_default_local_queue } : Type `string`, default `"other"`. `hopsworks.variables.kueue_system_jobs_cluster_queue` # { #helm.hopsworks.variables.kueue_system_jobs_cluster_queue } : Type `string`, default `""`. `hopsworks.variables.kueue_system_jobs_local_queue` # { #helm.hopsworks.variables.kueue_system_jobs_local_queue } : Type `string`, default `""`. `hopsworks.variables.ldap_account_status` # { #helm.hopsworks.variables.ldap_account_status } : Type `string`, default `"2"`. `hopsworks.variables.ldap_attr_binary` # { #helm.hopsworks.variables.ldap_attr_binary } : Type `string`, default `"java.naming.ldap.attributes.binary"`. `hopsworks.variables.ldap_dyn_group_target` # { #helm.hopsworks.variables.ldap_dyn_group_target } : Type `string`, default `"memberOf"`. `hopsworks.variables.ldap_group_dn` # { #helm.hopsworks.variables.ldap_group_dn } : Type `string`, default `""`. `hopsworks.variables.ldap_group_mapping` # { #helm.hopsworks.variables.ldap_group_mapping } : Type `string`, default `"ANY_GROUP->HOPS_USER"`. `hopsworks.variables.ldap_group_mapping_sync_enabled` # { #helm.hopsworks.variables.ldap_group_mapping_sync_enabled } : Type `string`, default `"false"`. `hopsworks.variables.ldap_group_mapping_sync_interval` # { #helm.hopsworks.variables.ldap_group_mapping_sync_interval } : Type `string`, default `"0"`. `hopsworks.variables.ldap_group_search_filter` # { #helm.hopsworks.variables.ldap_group_search_filter } : Type `string`, default `"member=%d"`. `hopsworks.variables.ldap_group_target` # { #helm.hopsworks.variables.ldap_group_target } : Type `string`, default `"cn"`. `hopsworks.variables.ldap_groups_search_filter` # { #helm.hopsworks.variables.ldap_groups_search_filter } : Type `string`, default `"(&(objectCategory=group)(cn=%c))"`. `hopsworks.variables.ldap_krb_dyn_grp_search_filter` # { #helm.hopsworks.variables.ldap_krb_dyn_grp_search_filter } : Type `string`, default `""`. `hopsworks.variables.ldap_krb_search_filter` # { #helm.hopsworks.variables.ldap_krb_search_filter } : Type `string`, default `"krbPrincipalName=%s"`. `hopsworks.variables.ldap_user_dn` # { #helm.hopsworks.variables.ldap_user_dn } : Type `string`, default `""`. `hopsworks.variables.ldap_user_email` # { #helm.hopsworks.variables.ldap_user_email } : Type `string`, default `"mail"`. `hopsworks.variables.ldap_user_givenName` # { #helm.hopsworks.variables.ldap_user_givenName } : Type `string`, default `"givenName"`. `hopsworks.variables.ldap_user_id` # { #helm.hopsworks.variables.ldap_user_id } : Type `string`, default `"uid"`. `hopsworks.variables.ldap_user_search_filter` # { #helm.hopsworks.variables.ldap_user_search_filter } : Type `string`, default `"uid=%s"`. `hopsworks.variables.ldap_user_surname` # { #helm.hopsworks.variables.ldap_user_surname } : Type `string`, default `"sn"`. `hopsworks.variables.library_install_timeout_minutes` # { #helm.hopsworks.variables.library_install_timeout_minutes } : Type `string`, default `"60"`. `hopsworks.variables.lifecycle_webhook_cluster_id` # { #helm.hopsworks.variables.lifecycle_webhook_cluster_id } : Type `string`, default `""`. clusterId added to the lifecycle webhook envelope. Omitted when unset. `hopsworks.variables.lifecycle_webhook_secret` # { #helm.hopsworks.variables.lifecycle_webhook_secret } : Type `string`, default `""`. HMAC-SHA256 signing key for the lifecycle webhook (X-Hopsworks-Signature). Empty sends unsigned. Seeded hidden. `hopsworks.variables.lifecycle_webhook_url` # { #helm.hopsworks.variables.lifecycle_webhook_url } : Type `string`, default `""`. Receiver URL for user/project/membership lifecycle events (HTTP POST, at-least-once). Empty disables it. Usable on standalone clusters, not SAAS-only. `hopsworks.variables.livy_startup_timeout` # { #helm.hopsworks.variables.livy_startup_timeout } : Type `string`, default `"240"`. `hopsworks.variables.livy_version` # { #helm.hopsworks.variables.livy_version } : Type `string`, default `"0.8.4-incubating-SNAPSHOT-bin"`. `hopsworks.variables.loadbalancer_external_domain_datanode` # { #helm.hopsworks.variables.loadbalancer_external_domain_datanode } : Type `string`, default `nil`. The domain name of the external load balancer for the HopsFS datanodes. If the load balancer is pre-provisioned then set the domain here, otherwise the hopsworks-update-lb-domains job will discover the domain name and set automatically `hopsworks.variables.loadbalancer_external_domain_feature_query` # { #helm.hopsworks.variables.loadbalancer_external_domain_feature_query } : Type `string`, default `nil`. The domain name of the external load balancer for arrowflight. If the load balancer is pre-provisioned then set the domain here, otherwise the hopsworks-update-lb-domains job will discover the domain name and set automatically `hopsworks.variables.loadbalancer_external_domain_mysqld` # { #helm.hopsworks.variables.loadbalancer_external_domain_mysqld } : Type `string`, default `nil`. The domain name of the external load balancer for mysqld. If the load balancer is pre-provisioned then set the domain here, otherwise the hopsworks-update-lb-domains job will discover the domain name and set automatically `hopsworks.variables.loadbalancer_external_domain_namenode` # { #helm.hopsworks.variables.loadbalancer_external_domain_namenode } : Type `string`, default `nil`. The domain name of the external load balancer for the HopsFS namenode. If the load balancer is pre-provisioned then set the domain here, otherwise the hopsworks-update-lb-domains job will discover the domain name and set automatically `hopsworks.variables.loadbalancer_external_domain_online_store_rest_server` # { #helm.hopsworks.variables.loadbalancer_external_domain_online_store_rest_server } : Type `string`, default `nil`. The domain name of the external load balancer for RDRS. If the load balancer is pre-provisioned then set the domain here, otherwise the hopsworks-update-lb-domains job will discover the domain name and set automatically `hopsworks.variables.loadbalancer_external_domain_opensearch` # { #helm.hopsworks.variables.loadbalancer_external_domain_opensearch } : Type `string`, default `nil`. The domain name of the external load balancer for opensearch. If the load balancer is pre-provisioned then set the domain here, otherwise the hopsworks-update-lb-domains job will discover the domain name and set automatically `hopsworks.variables.loadbalancer_external_domain_trino` # { #helm.hopsworks.variables.loadbalancer_external_domain_trino } : Type `string`, default `nil`. The domain name of the external load balancer for Trino. If the load balancer is pre-provisioned then set the domain here, otherwise the hopsworks-update-lb-domains job will discover the domain name and set automatically `hopsworks.variables.localhost` # { #helm.hopsworks.variables.localhost } : Type `string`, default `"false"`. `hopsworks.variables.log_history_limit` # { #helm.hopsworks.variables.log_history_limit } : Type `string`, default `"30"`. Cap on entries kept in the Log History UI (serving/agent deployments, Jupyter). Archives beyond this count are deleted after each archive, oldest first. A value <= 0 disables rotation. `hopsworks.variables.logstash_ip` # { #helm.hopsworks.variables.logstash_ip } : Type `string`, default `""`. `hopsworks.variables.logstash_port` # { #helm.hopsworks.variables.logstash_port } : Type `string`, default `""`. `hopsworks.variables.logstash_port_beam_jobserver_local` # { #helm.hopsworks.variables.logstash_port_beam_jobserver_local } : Type `string`, default `""`. `hopsworks.variables.logstash_port_serving` # { #helm.hopsworks.variables.logstash_port_serving } : Type `string`, default `""`. `hopsworks.variables.logstash_port_sklearn_serving` # { #helm.hopsworks.variables.logstash_port_sklearn_serving } : Type `string`, default `""`. `hopsworks.variables.logstash_port_tf_serving` # { #helm.hopsworks.variables.logstash_port_tf_serving } : Type `string`, default `""`. `hopsworks.variables.logstash_version` # { #helm.hopsworks.variables.logstash_version } : Type `string`, default `"7.16.3"`. `hopsworks.variables.managed_cloud_redirect_uri` # { #helm.hopsworks.variables.managed_cloud_redirect_uri } : Type `string`, default `""`. `hopsworks.variables.managed_docker_registry` # { #helm.hopsworks.variables.managed_docker_registry } : Type `string`, default `"false"`. `hopsworks.variables.management_mode` # { #helm.hopsworks.variables.management_mode } : Type `string`, default `""`. SAAS bridge: 'STANDALONE' (default, stock Hopsworks) or 'SAAS_MANAGED' (auth and project quota delegated to hopsworks-saas). Unset falls back to STANDALONE. `hopsworks.variables.max_allowed_long_running_http_requests` # { #helm.hopsworks.variables.max_allowed_long_running_http_requests } : Type `string`, default `"50"`. `hopsworks.variables.max_concurrent_base_sync_ops` # { #helm.hopsworks.variables.max_concurrent_base_sync_ops } : Type `string`, default `"5"`. `hopsworks.variables.max_env_yml_byte_size` # { #helm.hopsworks.variables.max_env_yml_byte_size } : Type `string`, default `"20000"`. `hopsworks.variables.max_num_proj_per_user` # { #helm.hopsworks.variables.max_num_proj_per_user } : Type `string`, default `"10"`. `hopsworks.variables.max_status_poll_retry` # { #helm.hopsworks.variables.max_status_poll_retry } : Type `string`, default `"5"`. `hopsworks.variables.mount_hopsfs_ray_job_container` # { #helm.hopsworks.variables.mount_hopsfs_ray_job_container } : Type `string`, default `"true"`. `hopsworks.variables.mr_user` # { #helm.hopsworks.variables.mr_user } : Type `string`, default `"mapred"`. `hopsworks.variables.multiregion_watchdog_enabled` # { #helm.hopsworks.variables.multiregion_watchdog_enabled } : Type `string`, default `"false"`. `hopsworks.variables.multiregion_watchdog_interval` # { #helm.hopsworks.variables.multiregion_watchdog_interval } : Type `string`, default `"5s"`. `hopsworks.variables.multiregion_watchdog_region` # { #helm.hopsworks.variables.multiregion_watchdog_region } : Type `string`, default `""`. `hopsworks.variables.multiregion_watchdog_url` # { #helm.hopsworks.variables.multiregion_watchdog_url } : Type `string`, default `""`. `hopsworks.variables.mysql_dir` # { #helm.hopsworks.variables.mysql_dir } : Type `string`, default `"/srv/hops/mysql"`. `hopsworks.variables.ndb_dir` # { #helm.hopsworks.variables.ndb_dir } : Type `string`, default `"/srv/hop/mysql-cluster"`. `hopsworks.variables.ndb_user` # { #helm.hopsworks.variables.ndb_user } : Type `string`, default `""`. `hopsworks.variables.ndb_version` # { #helm.hopsworks.variables.ndb_version } : Type `string`, default `"21.04.15"`. `hopsworks.variables.ndbinfo_db` # { #helm.hopsworks.variables.ndbinfo_db } : Type `string`, default `"ndbinfo"`. `hopsworks.variables.news_webflow_api_key` # { #helm.hopsworks.variables.news_webflow_api_key } : Type `string`, default `"dcc84358bfd37ffc68dbf18c68f74f478ff160d2286094077a9415ee03fbc805"`. `hopsworks.variables.news_webflow_api_url` # { #helm.hopsworks.variables.news_webflow_api_url } : Type `string`, default `"https://api.webflow.com/v2/collections/66bdd44475e24741477e1ae3/items"`. `hopsworks.variables.notebook_converter_job_timeout_sec` # { #helm.hopsworks.variables.notebook_converter_job_timeout_sec } : Type `string`, default `"300"`. `hopsworks.variables.npm_registry_url` # { #helm.hopsworks.variables.npm_registry_url } : Type `string`, default `""`. Registry npm package installs resolve from, applied as `npm config set registry ` in the environment build so the built image also resolves from it at runtime. Empty leaves npm on its compiled-in default (registry.npmjs.org). `hopsworks.variables.oauth_account_status` # { #helm.hopsworks.variables.oauth_account_status } : Type `string`, default `"1"`. `hopsworks.variables.oauth_group_mapping` # { #helm.hopsworks.variables.oauth_group_mapping } : Type `string`, default `""`. `hopsworks.variables.oauth_group_mapping_enabled` # { #helm.hopsworks.variables.oauth_group_mapping_enabled } : Type `string`, default `"false"`. `hopsworks.variables.oauth_group_mapping_sync_enabled` # { #helm.hopsworks.variables.oauth_group_mapping_sync_enabled } : Type `string`, default `"false"`. `hopsworks.variables.oauth_logout_redirect_uri` # { #helm.hopsworks.variables.oauth_logout_redirect_uri } : Type `string`, default `"hopsworks/"`. `hopsworks.variables.oauth_redirect_uri` # { #helm.hopsworks.variables.oauth_redirect_uri } : Type `string`, default `"hopsworks/callback"`. `hopsworks.variables.onlinefs_service_thread_number` # { #helm.hopsworks.variables.onlinefs_service_thread_number } : Type `string`, default `"10"`. `hopsworks.variables.onlinefs_user_email` # { #helm.hopsworks.variables.onlinefs_user_email } : Type `string`, default `"onlinefs@hopsworks.ai"`. `hopsworks.variables.onlinefs_user_password` # { #helm.hopsworks.variables.onlinefs_user_password } : Type `string`, default `"onlinefspw"`. `hopsworks.variables.opensearch_default_embedding_index` # { #helm.hopsworks.variables.opensearch_default_embedding_index } : Type `string`, default `""`. `hopsworks.variables.opensearch_index_mapping_limit` # { #helm.hopsworks.variables.opensearch_index_mapping_limit } : Type `string`, default `"1000"`. `hopsworks.variables.opensearch_num_default_embedding_index` # { #helm.hopsworks.variables.opensearch_num_default_embedding_index } : Type `string`, default `"1"`. `hopsworks.variables.payara_dir` # { #helm.hopsworks.variables.payara_dir } : Type `string`, default `"/opt/payara/appserver/glassfish/domains/domain1"`. `hopsworks.variables.pki_ca_configuration` # { #helm.hopsworks.variables.pki_ca_configuration } : Type `string`. ??? note "Default" ```yaml '{"rootCA":{},"intermediateCA":{},"kubernetesCA":{"subjectAlternativeName":{"dns":["hopsworks0.logicalclocks.com","hops-kubernetes","hops-kubernetes.default","hops-kubernetes.default.svc","hops-kubernetes.default.svc.cluster","hops-kubernetes.default.svc.cluster.local","*.hops-system.svc"],"ip":["10.244.0.1","192.168.30.101","127.0.0.1","10.96.0.10","10.96.0.1"]}}}' ``` `hopsworks.variables.platform_intelligence_llm_api_key` # { #helm.hopsworks.variables.platform_intelligence_llm_api_key } : Type `string`, default `""`. `hopsworks.variables.platform_intelligence_llm_base_url` # { #helm.hopsworks.variables.platform_intelligence_llm_base_url } : Type `string`, default `""`. `hopsworks.variables.platform_intelligence_llm_model` # { #helm.hopsworks.variables.platform_intelligence_llm_model } : Type `string`, default `""`. `hopsworks.variables.preinstalled_python_lib_names` # { #helm.hopsworks.variables.preinstalled_python_lib_names } : Type `string`. ??? note "Default" ```yaml pydoop, pyspark, jupyterlab, sparkmagic, hdfscontents, pyjks, hops-apache-beam, pyopenssl ``` `hopsworks.variables.project_namespace_labels` # { #helm.hopsworks.variables.project_namespace_labels } : Type `string`, default `""`. `hopsworks.variables.project_namespace_network_policy_allowed_namespaces` # { #helm.hopsworks.variables.project_namespace_network_policy_allowed_namespaces } : Type `string`, default `""`. `hopsworks.variables.project_namespace_network_policy_enabled` # { #helm.hopsworks.variables.project_namespace_network_policy_enabled } : Type `bool`, default `true`. `hopsworks.variables.project_namespace_network_policy_reconcile_interval` # { #helm.hopsworks.variables.project_namespace_network_policy_reconcile_interval } : Type `string`, default `"1m"`. `hopsworks.variables.prometheus_port` # { #helm.hopsworks.variables.prometheus_port } : Type `string`, default `"9089"`. `hopsworks.variables.provenance_archive_delay` # { #helm.hopsworks.variables.provenance_archive_delay } : Type `string`, default `"86400"`. `hopsworks.variables.provenance_archive_size` # { #helm.hopsworks.variables.provenance_archive_size } : Type `string`, default `"10"`. `hopsworks.variables.provenance_cleaner_period` # { #helm.hopsworks.variables.provenance_cleaner_period } : Type `string`, default `"3600"`. `hopsworks.variables.provenance_graph_max_size` # { #helm.hopsworks.variables.provenance_graph_max_size } : Type `string`, default `"10000"`. `hopsworks.variables.provenance_type` # { #helm.hopsworks.variables.provenance_type } : Type `string`, default `"FULL"`. `hopsworks.variables.public_https_port` # { #helm.hopsworks.variables.public_https_port } : Type `string`, default `""`. `hopsworks.variables.pushgateway_cleaner_batch_size` # { #helm.hopsworks.variables.pushgateway_cleaner_batch_size } : Type `string`, default `"100"`. `hopsworks.variables.pushgateway_group_ttl_minutes` # { #helm.hopsworks.variables.pushgateway_group_ttl_minutes } : Type `string`, default `"15"`. `hopsworks.variables.pushgateway_monitor_interval_ms` # { #helm.hopsworks.variables.pushgateway_monitor_interval_ms } : Type `string`, default `"300000"`. `hopsworks.variables.py4j_archive` # { #helm.hopsworks.variables.py4j_archive } : Type `string`, default `""`. `hopsworks.variables.pypi_indexer_timer_enabled` # { #helm.hopsworks.variables.pypi_indexer_timer_enabled } : Type `string`, default `"true"`. `hopsworks.variables.pypi_indexer_timer_interval` # { #helm.hopsworks.variables.pypi_indexer_timer_interval } : Type `string`, default `"1d"`. `hopsworks.variables.pypi_rest_endpoint` # { #helm.hopsworks.variables.pypi_rest_endpoint } : Type `string`, default `"https://pypi.org/pypi/{package}/json"`. `hopsworks.variables.pypi_simple_endpoint` # { #helm.hopsworks.variables.pypi_simple_endpoint } : Type `string`, default `"https://pypi.org/simple/"`. `hopsworks.variables.python_job_cores` # { #helm.hopsworks.variables.python_job_cores } : Type `string`, default `"1.0"`. `hopsworks.variables.python_job_gpus` # { #helm.hopsworks.variables.python_job_gpus } : Type `string`, default `"0"`. `hopsworks.variables.python_job_kube_waiting_timeout_ms` # { #helm.hopsworks.variables.python_job_kube_waiting_timeout_ms } : Type `string`, default `"300000"`. `hopsworks.variables.python_job_memory` # { #helm.hopsworks.variables.python_job_memory } : Type `string`, default `"2048"`. `hopsworks.variables.python_library_updates_monitor_interval` # { #helm.hopsworks.variables.python_library_updates_monitor_interval } : Type `string`, default `"1d"`. `hopsworks.variables.python_pod_kill_grace_period_seconds` # { #helm.hopsworks.variables.python_pod_kill_grace_period_seconds } : Type `string`, default `"60"`. `hopsworks.variables.pythonapp_cores` # { #helm.hopsworks.variables.pythonapp_cores } : Type `string`, default `"1.0"`. `hopsworks.variables.pythonapp_gpus` # { #helm.hopsworks.variables.pythonapp_gpus } : Type `string`, default `"0"`. `hopsworks.variables.pythonapp_memory` # { #helm.hopsworks.variables.pythonapp_memory } : Type `string`, default `"2048"`. `hopsworks.variables.quotas_featuregroups_online_disabled` # { #helm.hopsworks.variables.quotas_featuregroups_online_disabled } : Type `string`, default `"-1"`. `hopsworks.variables.quotas_featuregroups_online_enabled` # { #helm.hopsworks.variables.quotas_featuregroups_online_enabled } : Type `string`, default `"-1"`. `hopsworks.variables.quotas_max_parallel_executions` # { #helm.hopsworks.variables.quotas_max_parallel_executions } : Type `string`, default `"-1"`. `hopsworks.variables.quotas_model_deployments_running` # { #helm.hopsworks.variables.quotas_model_deployments_running } : Type `string`, default `"-1"`. `hopsworks.variables.quotas_model_deployments_total` # { #helm.hopsworks.variables.quotas_model_deployments_total } : Type `string`, default `"-1"`. `hopsworks.variables.quotas_training_datasets` # { #helm.hopsworks.variables.quotas_training_datasets } : Type `string`, default `"-1"`. `hopsworks.variables.ray_cluster_max_worker_replicas` # { #helm.hopsworks.variables.ray_cluster_max_worker_replicas } : Type `string`, default `"20"`. `hopsworks.variables.ray_cluster_shutdown_after_completion` # { #helm.hopsworks.variables.ray_cluster_shutdown_after_completion } : Type `string`, default `"true"`. `hopsworks.variables.ray_cluster_start_wait_time_seconds` # { #helm.hopsworks.variables.ray_cluster_start_wait_time_seconds } : Type `string`, default `"360"`. `hopsworks.variables.ray_cluster_termination_grace_period_seconds` # { #helm.hopsworks.variables.ray_cluster_termination_grace_period_seconds } : Type `string`, default `"10"`. `hopsworks.variables.ray_enabled` # { #helm.hopsworks.variables.ray_enabled } : Type `bool`, default `false`. ray_enabled indicates if hopsworks configurations for ray should be applied `hopsworks.variables.ray_job_driver_cores` # { #helm.hopsworks.variables.ray_job_driver_cores } : Type `string`, default `"1.0"`. `hopsworks.variables.ray_job_driver_gpus` # { #helm.hopsworks.variables.ray_job_driver_gpus } : Type `string`, default `"0"`. `hopsworks.variables.ray_job_driver_memory` # { #helm.hopsworks.variables.ray_job_driver_memory } : Type `string`, default `"4096"`. `hopsworks.variables.ray_job_pod_kill_grace_period_seconds` # { #helm.hopsworks.variables.ray_job_pod_kill_grace_period_seconds } : Type `string`, default `"300"`. `hopsworks.variables.ray_job_worker_cores` # { #helm.hopsworks.variables.ray_job_worker_cores } : Type `string`, default `"1.0"`. `hopsworks.variables.ray_job_worker_gpus` # { #helm.hopsworks.variables.ray_job_worker_gpus } : Type `string`, default `"0"`. `hopsworks.variables.ray_job_worker_memory` # { #helm.hopsworks.variables.ray_job_worker_memory } : Type `string`, default `"4096"`. `hopsworks.variables.ray_materialization_dir` # { #helm.hopsworks.variables.ray_materialization_dir } : Type `string`, default `"/srv/hops/ray/job"`. `hopsworks.variables.ray_version` # { #helm.hopsworks.variables.ray_version } : Type `string`, default `"2.58.0"`. `hopsworks.variables.recovery_path` # { #helm.hopsworks.variables.recovery_path } : Type `string`, default `""`. `hopsworks.variables.reject_remote_user_no_group` # { #helm.hopsworks.variables.reject_remote_user_no_group } : Type `string`, default `"false"`. `hopsworks.variables.remote_auth_need_consent` # { #helm.hopsworks.variables.remote_auth_need_consent } : Type `string`, default `"true"`. `hopsworks.variables.requests_verify` # { #helm.hopsworks.variables.requests_verify } : Type `string`, default `"true"`. `hopsworks.variables.reserved_project_names` # { #helm.hopsworks.variables.reserved_project_names } : Type `string`. ??? note "Default" ```yaml hopsworks,information_schema,airflow,glassfish_timers,grafana,hops,metastore,mysql,ndbinfo,performance_schema,sqoop,sys,base,python37,python38,python39,python310,filebeat,airflow,git,onlinefs,sklearnserver,rondb_replication,default,kube-system,kube-public,kube-node-lease,kube_system,kube_public,kube_node_lease ``` `hopsworks.variables.rmyarn_user` # { #helm.hopsworks.variables.rmyarn_user } : Type `string`, default `"rmyarn"`. `hopsworks.variables.rondb_quotas` # { #helm.hopsworks.variables.rondb_quotas } : Type `string`, default `""`. RONDB quotas in csv format. e.g `rate-per-sec=1000,max-transaction-size=100`. Refer to for the full list of arguments and their definitions. `hopsworks.variables.rondb_usage_cache_ttl_seconds` # { #helm.hopsworks.variables.rondb_usage_cache_ttl_seconds } : Type `string`, default `"60"`. How often (seconds) the cached RonDB per-database memory usage snapshot is refreshed. The backing ndbinfo scan costs the same for one database as for all, so one snapshot serves every project. `hopsworks.variables.rondb_usage_query_timeout_seconds` # { #helm.hopsworks.variables.rondb_usage_query_timeout_seconds } : Type `string`, default `"10"`. Query timeout (seconds) for the ndbinfo memory usage scan. Scan duration grows with the cluster's table count; raise this on large clusters if usage stops showing in the project quotas UI. `hopsworks.variables.saas_entry_point_url` # { #helm.hopsworks.variables.saas_entry_point_url } : Type `string`, default `""`. SAAS bridge: auth entry-point URL the frontend redirects to when management_mode is SAAS_MANAGED. Ignored in STANDALONE. `hopsworks.variables.scikit_learn_version` # { #helm.hopsworks.variables.scikit_learn_version } : Type `string`, default `"1.3.2"`. `hopsworks.variables.service_jwt_exp_leeway_sec` # { #helm.hopsworks.variables.service_jwt_exp_leeway_sec } : Type `string`, default `"172800000"`. `hopsworks.variables.service_jwt_lifetime_ms` # { #helm.hopsworks.variables.service_jwt_lifetime_ms } : Type `string`, default `"604800000"`. `hopsworks.variables.service_key_rotation_enabled` # { #helm.hopsworks.variables.service_key_rotation_enabled } : Type `string`, default `"false"`. `hopsworks.variables.service_key_rotation_interval` # { #helm.hopsworks.variables.service_key_rotation_interval } : Type `string`, default `"2d"`. `hopsworks.variables.serving_allow_stop_after_seconds` # { #helm.hopsworks.variables.serving_allow_stop_after_seconds } : Type `string`, default `"30"`. `hopsworks.variables.serving_connection_pool_size` # { #helm.hopsworks.variables.serving_connection_pool_size } : Type `string`, default `"40"`. `hopsworks.variables.serving_feature_log_materialization_cron` # { #helm.hopsworks.variables.serving_feature_log_materialization_cron } : Type `string`, default `"0 0 0 * * ? *"`. `hopsworks.variables.serving_feature_log_materialization_row_limit` # { #helm.hopsworks.variables.serving_feature_log_materialization_row_limit } : Type `string`, default `"50000000"`. `hopsworks.variables.serving_feature_log_online_ttl_hours` # { #helm.hopsworks.variables.serving_feature_log_online_ttl_hours } : Type `string`, default `"30"`. `hopsworks.variables.serving_feature_logger_batch_bytes` # { #helm.hopsworks.variables.serving_feature_logger_batch_bytes } : Type `string`, default `"1048576"`. `hopsworks.variables.serving_feature_logger_batch_seconds` # { #helm.hopsworks.variables.serving_feature_logger_batch_seconds } : Type `string`, default `"5"`. `hopsworks.variables.serving_feature_logger_client_pool_size` # { #helm.hopsworks.variables.serving_feature_logger_client_pool_size } : Type `string`, default `"3"`. `hopsworks.variables.serving_feature_logger_client_req_timeout_seconds` # { #helm.hopsworks.variables.serving_feature_logger_client_req_timeout_seconds } : Type `string`, default `"3"`. `hopsworks.variables.serving_feature_logger_flush_bytes` # { #helm.hopsworks.variables.serving_feature_logger_flush_bytes } : Type `string`, default `"1048576"`. `hopsworks.variables.serving_feature_logger_flush_interval_seconds` # { #helm.hopsworks.variables.serving_feature_logger_flush_interval_seconds } : Type `string`, default `"300"`. `hopsworks.variables.serving_feature_logger_max_buffer_bytes` # { #helm.hopsworks.variables.serving_feature_logger_max_buffer_bytes } : Type `string`, default `"67108864"`. `hopsworks.variables.serving_feature_logger_max_event_bytes` # { #helm.hopsworks.variables.serving_feature_logger_max_event_bytes } : Type `string`, default `"8388608"`. `hopsworks.variables.serving_feature_logger_max_event_rows` # { #helm.hopsworks.variables.serving_feature_logger_max_event_rows } : Type `string`, default `"512"`. `hopsworks.variables.serving_feature_logger_queue_size` # { #helm.hopsworks.variables.serving_feature_logger_queue_size } : Type `string`, default `"1000"`. `hopsworks.variables.serving_feature_logger_shutdown_seconds` # { #helm.hopsworks.variables.serving_feature_logger_shutdown_seconds } : Type `string`, default `"20"`. `hopsworks.variables.serving_feature_logging_transport` # { #helm.hopsworks.variables.serving_feature_logging_transport } : Type `string`, default `"realtime"`. `hopsworks.variables.serving_max_route_connections` # { #helm.hopsworks.variables.serving_max_route_connections } : Type `string`, default `"10"`. `hopsworks.variables.serving_redeploy_not_found_after_seconds` # { #helm.hopsworks.variables.serving_redeploy_not_found_after_seconds } : Type `string`, default `"120"`. `hopsworks.variables.serving_state_manager_batch_size` # { #helm.hopsworks.variables.serving_state_manager_batch_size } : Type `string`, default `"25"`. `hopsworks.variables.serving_state_manager_enabled` # { #helm.hopsworks.variables.serving_state_manager_enabled } : Type `string`, default `"true"`. `hopsworks.variables.serving_state_manager_interval_ms` # { #helm.hopsworks.variables.serving_state_manager_interval_ms } : Type `string`, default `"300000"`. `hopsworks.variables.spark_dir` # { #helm.hopsworks.variables.spark_dir } : Type `string`, default `"/srv/hops/spark"`. `hopsworks.variables.spark_executor_min_memory` # { #helm.hopsworks.variables.spark_executor_min_memory } : Type `string`, default `"1024"`. `hopsworks.variables.spark_hops_utils_dir` # { #helm.hopsworks.variables.spark_hops_utils_dir } : Type `string`, default `"/srv/hops/artifacts"`. `hopsworks.variables.spark_job_driver_cores` # { #helm.hopsworks.variables.spark_job_driver_cores } : Type `string`, default `"1.0"`. `hopsworks.variables.spark_job_driver_memory` # { #helm.hopsworks.variables.spark_job_driver_memory } : Type `string`, default `"2048"`. `hopsworks.variables.spark_job_executor_cores` # { #helm.hopsworks.variables.spark_job_executor_cores } : Type `string`, default `"1.0"`. `hopsworks.variables.spark_job_executor_memory` # { #helm.hopsworks.variables.spark_job_executor_memory } : Type `string`, default `"4096"`. `hopsworks.variables.spark_launcher_sa_annotations` # { #helm.hopsworks.variables.spark_launcher_sa_annotations } : Type `string`, default `""`. `hopsworks.variables.spark_pod_kill_grace_period_seconds` # { #helm.hopsworks.variables.spark_pod_kill_grace_period_seconds } : Type `string`, default `"1200"`. `hopsworks.variables.spark_remove_job_when_completed` # { #helm.hopsworks.variables.spark_remove_job_when_completed } : Type `string`, default `"true"`. `hopsworks.variables.spark_ui_logs_offset` # { #helm.hopsworks.variables.spark_ui_logs_offset } : Type `string`, default `"512000"`. `hopsworks.variables.spark_user` # { #helm.hopsworks.variables.spark_user } : Type `string`, default `"spark"`. `hopsworks.variables.spark_version` # { #helm.hopsworks.variables.spark_version } : Type `string`, default `"4.1.3.0"`. `hopsworks.variables.srvmanager_password` # { #helm.hopsworks.variables.srvmanager_password } : Type `string`, default `"srvmanagerpwd"`. `hopsworks.variables.staging_dir` # { #helm.hopsworks.variables.staging_dir } : Type `string`, default `"/srv/hops/staging"`. `hopsworks.variables.statistics_cleaner_batch_size` # { #helm.hopsworks.variables.statistics_cleaner_batch_size } : Type `string`, default `"1000"`. `hopsworks.variables.statistics_cleaner_interval_ms` # { #helm.hopsworks.variables.statistics_cleaner_interval_ms } : Type `string`, default `"900000"`. `hopsworks.variables.streamlit_sharing` # { #helm.hopsworks.variables.streamlit_sharing } : Type `bool`, default `false`. `hopsworks.variables.sudoers_dir` # { #helm.hopsworks.variables.sudoers_dir } : Type `string`, default `"/srv/hops/sbin"`. `hopsworks.variables.superset_admin_roles` # { #helm.hopsworks.variables.superset_admin_roles } : Type `string`, default `"Admin"`. `hopsworks.variables.superset_proxy_connect_timeout_ms` # { #helm.hopsworks.variables.superset_proxy_connect_timeout_ms } : Type `string`, default `"10000"`. `hopsworks.variables.superset_proxy_connection_request_timeout_ms` # { #helm.hopsworks.variables.superset_proxy_connection_request_timeout_ms } : Type `string`, default `"10000"`. `hopsworks.variables.superset_proxy_max_connections` # { #helm.hopsworks.variables.superset_proxy_max_connections } : Type `string`, default `"50"`. `hopsworks.variables.superset_proxy_read_timeout_ms` # { #helm.hopsworks.variables.superset_proxy_read_timeout_ms } : Type `string`, default `"180000"`. `hopsworks.variables.superset_user_roles` # { #helm.hopsworks.variables.superset_user_roles } : Type `string`, default `"Gamma,sql_lab,Dataset"`. `hopsworks.variables.support_email_addr` # { #helm.hopsworks.variables.support_email_addr } : Type `string`, default `"support@hopsworks.ai"`. `hopsworks.variables.tag_history_archive_max_events` # { #helm.hopsworks.variables.tag_history_archive_max_events } : Type `string`, default `"20000"`. Largest number of tag history events an archive flip will write, one per tag key of every existing attachment. Turning archiving on or off writes them in a single transaction, which cannot be split without losing the baseline it exists to record, so the work is bounded by NDB's MaxNoOfConcurrentOperations and the request timeout rather than by paging. Counted in events because an attachment with several keys is several rows; a limit on attachments alone did not bound the work. Above this the call is refused with the count and this limit instead of rolling back with no readable cause. Raise it alongside MaxNoOfConcurrentOperations. `hopsworks.variables.tag_history_cleaner_batch_size` # { #helm.hopsworks.variables.tag_history_cleaner_batch_size } : Type `string`, default `"1000"`. Rows deleted per transaction by the tag history retention sweep. Bounded so a large sweep is a series of short transactions rather than one that runs into NDB's MaxNoOfConcurrentOperations. `hopsworks.variables.tag_history_cleaner_interval_ms` # { #helm.hopsworks.variables.tag_history_cleaner_interval_ms } : Type `string`, default `"86400000"`. How often the tag history retention sweep runs, in milliseconds. Values below 60000 are refused and fall back to 24h. `hopsworks.variables.tag_history_retention_days` # { #helm.hopsworks.variables.tag_history_retention_days } : Type `string`, default `"0"`. How many days of tag history to keep. tag_history is append-only and nothing else bounds it: clearing a schema's archive flag stops new rows but deletes none, so without retention the only way to reclaim the space is by hand. "0" keeps everything, which is the default because deleting analytics history as a side effect of an upgrade would surprise everyone reporting on it. Set a number of days to bound it. `hopsworks.variables.tensorboard_max_last_accessed` # { #helm.hopsworks.variables.tensorboard_max_last_accessed } : Type `string`, default `"1140000"`. `hopsworks.variables.tensorboard_max_reload_threads` # { #helm.hopsworks.variables.tensorboard_max_reload_threads } : Type `string`, default `"1"`. `hopsworks.variables.tensorflow_version` # { #helm.hopsworks.variables.tensorflow_version } : Type `string`, default `"2.20.0"`. `hopsworks.variables.testconnector_image_version` # { #helm.hopsworks.variables.testconnector_image_version } : Type `string`, default `"1.0"`. `hopsworks.variables.tf_spark_connector_version` # { #helm.hopsworks.variables.tf_spark_connector_version } : Type `string`, default `""`. `hopsworks.variables.trino_default_catalog` # { #helm.hopsworks.variables.trino_default_catalog } : Type `string`, default `"delta"`. `hopsworks.variables.trino_events_cleaner_batch_size` # { #helm.hopsworks.variables.trino_events_cleaner_batch_size } : Type `string`, default `"1000"`. `hopsworks.variables.trino_events_delete_after_days` # { #helm.hopsworks.variables.trino_events_delete_after_days } : Type `string`, default `"61"`. `hopsworks.variables.twofactor_auth` # { #helm.hopsworks.variables.twofactor_auth } : Type `string`, default `"false"`. `hopsworks.variables.twofactor_excluded_groups` # { #helm.hopsworks.variables.twofactor_excluded_groups } : Type `string`, default `"AGENT;CLUSTER_AGENT"`. `hopsworks.variables.unix_usernames_conf` # { #helm.hopsworks.variables.unix_usernames_conf } : Type `string`. ??? note "Default" ```yaml '{\"glassfish\":\"glassfish\",\"hdfs\":\"hdfs\",\"rmyarn\":\"rmyarn\",\"yarn\":\"yarn\",\"hive\":\"hive\",\"livy\":\"livy\",\"flink\":\"flink\",\"consul\":\"consul\",\"hopsmon\":\"hopsmon\",\"zookeeper\":\"zookeeper\",\"onlinefs\":\"onlinefs\",\"elastic\":\"elastic\",\"kagent\":\"kagent\",\"mysql\":\"mysql\",\"airflow\":\"airflow\"}' ``` `hopsworks.variables.upload_chunk_size` # { #helm.hopsworks.variables.upload_chunk_size } : Type `string`, default `"10485760"`. `hopsworks.variables.upload_policy` # { #helm.hopsworks.variables.upload_policy } : Type `string`, default `"enabled"`. Who may upload files into the cluster. `enabled` allows any user with write access to the destination dataset, `admins_only` restricts uploads to members of `HOPS_ADMIN`, `disabled` blocks everyone including administrators. Applies to the web UI and to clients such as the Python SDK. Unrecognised values fall back to `enabled`. Note that this governs uploading new files only: operations on files already in the cluster filesystem, such as installing a python library from an existing path, are unaffected. `hopsworks.variables.user_cert_valid_days` # { #helm.hopsworks.variables.user_cert_valid_days } : Type `string`, default `"12"`. `hopsworks.variables.verification_path` # { #helm.hopsworks.variables.verification_path } : Type `string`, default `"hopsworks-api/api/auth/verify"`. `hopsworks.variables.yarn_default_payment_type` # { #helm.hopsworks.variables.yarn_default_payment_type } : Type `string`, default `"NOLIMIT"`. `hopsworks.variables.yarn_default_quota` # { #helm.hopsworks.variables.yarn_default_quota } : Type `string`, default `"60000000"`. `hopsworks.variables.yarn_user` # { #helm.hopsworks.variables.yarn_user } : Type `string`, default `"yarn"`. `hopsworks.variables.zookeeper_version` # { #helm.hopsworks.variables.zookeeper_version } : Type `string`, default `"3.7.1"`. `hopsworks.variables.mount_hopsfs_in_python_job` Deprecated # { #helm.hopsworks.variables.mount_hopsfs_in_python_job } : Type `bool`, default `true`. Deprecated. Python Apps always mount HopsFS; this remains only for legacy Python jobs.
## velero { #helm-values-hopsworks-velero } ??? example "Defaults as YAML" ```yaml hopsworks: velero: backup: dynamicNamespaceFilter: enabled: true schedule: '@hourly' enabled: null excludedNamespaces: - kube-system - kube-public - kube-node-lease - argocd - default - ingress-nginx includedNamespaces: [] mainScheduleName: k8s-backups-main schedule: null storageLocation: create: true name: hopsworks-bsl s3ForcePathStyle: null storagePrefix: k8s_backup ttl: null usersScheduleName: k8s-backups-users-resources deploymentName: velero enforcePrerequisiteCheck: true namespace: velero restore: mainScheduleBackupId: null usersScheduleBackupId: null ttlSecondsAfterFinished: null ```
`hopsworks.velero.backup` # { #helm.hopsworks.velero.backup } : Type `object`. backup configuration ??? note "Default" ```yaml dynamicNamespaceFilter: enabled: true schedule: '@hourly' enabled: null excludedNamespaces: - kube-system - kube-public - kube-node-lease - argocd - default - ingress-nginx includedNamespaces: [] mainScheduleName: k8s-backups-main schedule: null storageLocation: create: true name: hopsworks-bsl s3ForcePathStyle: null storagePrefix: k8s_backup ttl: null usersScheduleName: k8s-backups-users-resources ``` `hopsworks.velero.backup.dynamicNamespaceFilter` # { #helm.hopsworks.velero.backup.dynamicNamespaceFilter } : Type `object`, default `{"enabled":true,"schedule":"@hourly"}`. Configure cron job to dynamically list the project namespaces and update the users backup schedule, and to discover satellite namespaces (labeled hopsworks.ai/onlinefs-cluster) and update the main backup schedule. If includedNamespaces is defined then this is disabled by default. `hopsworks.velero.backup.enabled` # { #helm.hopsworks.velero.backup.enabled } : Type `string`, default `nil`. Enable or disable velero for taking backups `hopsworks.velero.backup.excludedNamespaces` # { #helm.hopsworks.velero.backup.excludedNamespaces } : Type `list`. array of namespaces to exclude their resources in the backup. for ArgoCD, explicitly include all non-user namespaces since reconciliation makes dynamicNamespaceFilter ineffective. ??? note "Default" ```yaml - kube-system - kube-public - kube-node-lease - argocd - default - ingress-nginx ``` `hopsworks.velero.backup.includedNamespaces` # { #helm.hopsworks.velero.backup.includedNamespaces } : Type `list`, default `[]`. array of namespaces to include their resources in the backup `hopsworks.velero.backup.mainScheduleName` # { #helm.hopsworks.velero.backup.mainScheduleName } : Type `string`, default `"k8s-backups-main"`. The name of the main backup schedule that backs up the generated secrets from the cluster, the backups metadata configmaps, and serving configmaps and secrets. `hopsworks.velero.backup.schedule` # { #helm.hopsworks.velero.backup.schedule } : Type `string`, default `nil`. Backup schedule `hopsworks.velero.backup.storageLocation` # { #helm.hopsworks.velero.backup.storageLocation } : Type `object`. The storage backup location configuration ??? note "Default" ```yaml create: true name: hopsworks-bsl s3ForcePathStyle: null storagePrefix: k8s_backup ``` `hopsworks.velero.backup.storageLocation.create` # { #helm.hopsworks.velero.backup.storageLocation.create } : Type `bool`, default `true`. create a backup storage location. If disabled, we expect a backup storage location that exists with the name velero.backup.storageLocation.name. `hopsworks.velero.backup.storageLocation.name` # { #helm.hopsworks.velero.backup.storageLocation.name } : Type `string`, default `"hopsworks-bsl"`. storage backup location name. If create is enabled, we will create a new backup storage location otherwise we expect that the backup storage location to exist. `hopsworks.velero.backup.storageLocation.s3ForcePathStyle` # { #helm.hopsworks.velero.backup.storageLocation.s3ForcePathStyle } : Type `string`, default `nil`. Force S3 path-style addressing for the velero backup storage location. When unset (null), defaults to true for MinIO and unset for managed S3. Set explicitly to true for S3-compatible object storage that does not support virtual-hosted-style addressing (e.g. evroc, Ceph RadosGW). `hopsworks.velero.backup.storageLocation.storagePrefix` # { #helm.hopsworks.velero.backup.storageLocation.storagePrefix } : Type `string`, default `"k8s_backup"`. The storage prefix to store the backups under in the configured bucket `hopsworks.velero.backup.ttl` # { #helm.hopsworks.velero.backup.ttl } : Type `string`, default `nil`. The amount of time before backups created on this schedule are eligible for garbage collection. If not specified, a default value of 30 days will be used. The value must be a Go-style duration using hour/minute/second units (e.g. 24h, 168h, or 24h0m0s). Calendar units such as "d" or "w" are not supported. `hopsworks.velero.backup.usersScheduleName` # { #helm.hopsworks.velero.backup.usersScheduleName } : Type `string`, default `"k8s-backups-users-resources"`. The name of the users backup schedule that backs up the users' resources in their project namespaces. `hopsworks.velero.deploymentName` # { #helm.hopsworks.velero.deploymentName } : Type `string`, default `"velero"`. The name of the velero deployment `hopsworks.velero.enforcePrerequisiteCheck` # { #helm.hopsworks.velero.enforcePrerequisiteCheck } : Type `bool`, default `true`. When enabled, Helm will verify that the required Velero CRDs and Deployment already exist in the cluster using `lookup` and fail if they are missing. Disable this for offline rendering with `helm template`. `hopsworks.velero.namespace` # { #helm.hopsworks.velero.namespace } : Type `string`, default `"velero"`. The velero namespace. It should be on a different namespace other than the Hopsworks install namespace to avoid accidentaly deleting the Hopsworks namespace when uninstalling velero. `hopsworks.velero.restore.mainScheduleBackupId` # { #helm.hopsworks.velero.restore.mainScheduleBackupId } : Type `string`, default `nil`. The backup ID used for the restore operation of the main schedule. If unset, the latest backup from the main schedule will be used. This parameter is primarily used internally by the Helm chart during in-place restores, but also serves to record which backup ID was restored. `hopsworks.velero.restore.usersScheduleBackupId` # { #helm.hopsworks.velero.restore.usersScheduleBackupId } : Type `string`, default `nil`. The backup ID used for the restore operation of the users schedule. If unset, the latest backup from the users schedule will be used. This parameter is primarily used internally by the Helm chart during in-place restores, but also serves to record which backup ID was restored. `hopsworks.velero.ttlSecondsAfterFinished` # { #helm.hopsworks.velero.ttlSecondsAfterFinished } : Type `string`, default `nil`. TTL in seconds for the velero-create-cloud-credentials Job. Overrides global default.
## wipeWhenUninstall { #helm-values-hopsworks-wipewhenuninstall } ??? example "Defaults as YAML" ```yaml hopsworks: wipeWhenUninstall: - apiGroup: remoteshuffleservices.uniffle.apache.org resources: - remoteshuffleservices - apiGroup: sparkapplications.sparkoperator.k8s.io resources: - sparkapplications - apiGroup: hopsworkscerts.certs.hopsworks.ai resources: - hopsworkscerts - apiGroup: certs.hopsworks.ai resources: - hopsworkscerts - apiGroup: batch resources: - jobs - cronjobs - apiGroup: apps resources: - deployments - statefulsets - daemonsets - apiGroup: rbac.authorization.k8s.io resources: - roles - rolebindings - clusterroles - clusterrolebindings - apiGroup: networking.k8s.io resources: - networkpolicies - ingresses - apiGroup: storage.k8s.io resources: - storageclasses - apiGroup: scheduling.k8s.io resources: - priorityclasses - apiGroup: admissionregistration.k8s.io resources: - mutatingwebhookconfigurations - validatingwebhookconfigurations - apiGroup: '' resources: - secrets - serviceaccounts - configmaps - persistentvolumeclaims - persistentvolumes - pods - services - endpoints - namespaces ```
`hopsworks.wipeWhenUninstall[0].apiGroup` # { #helm.hopsworks.wipeWhenUninstall.0.apiGroup } : Type `string`, default `"remoteshuffleservices.uniffle.apache.org"`. `hopsworks.wipeWhenUninstall[0].resources[0]` # { #helm.hopsworks.wipeWhenUninstall.0.resources.0 } : Type `string`, default `"remoteshuffleservices"`. `hopsworks.wipeWhenUninstall[10].apiGroup` # { #helm.hopsworks.wipeWhenUninstall.10.apiGroup } : Type `string`, default `"admissionregistration.k8s.io"`. `hopsworks.wipeWhenUninstall[10].resources[0]` # { #helm.hopsworks.wipeWhenUninstall.10.resources.0 } : Type `string`, default `"mutatingwebhookconfigurations"`. `hopsworks.wipeWhenUninstall[10].resources[1]` # { #helm.hopsworks.wipeWhenUninstall.10.resources.1 } : Type `string`, default `"validatingwebhookconfigurations"`. `hopsworks.wipeWhenUninstall[11].apiGroup` # { #helm.hopsworks.wipeWhenUninstall.11.apiGroup } : Type `string`, default `""`. `hopsworks.wipeWhenUninstall[11].resources[0]` # { #helm.hopsworks.wipeWhenUninstall.11.resources.0 } : Type `string`, default `"secrets"`. `hopsworks.wipeWhenUninstall[11].resources[1]` # { #helm.hopsworks.wipeWhenUninstall.11.resources.1 } : Type `string`, default `"serviceaccounts"`. `hopsworks.wipeWhenUninstall[11].resources[2]` # { #helm.hopsworks.wipeWhenUninstall.11.resources.2 } : Type `string`, default `"configmaps"`. `hopsworks.wipeWhenUninstall[11].resources[3]` # { #helm.hopsworks.wipeWhenUninstall.11.resources.3 } : Type `string`, default `"persistentvolumeclaims"`. `hopsworks.wipeWhenUninstall[11].resources[4]` # { #helm.hopsworks.wipeWhenUninstall.11.resources.4 } : Type `string`, default `"persistentvolumes"`. `hopsworks.wipeWhenUninstall[11].resources[5]` # { #helm.hopsworks.wipeWhenUninstall.11.resources.5 } : Type `string`, default `"pods"`. `hopsworks.wipeWhenUninstall[11].resources[6]` # { #helm.hopsworks.wipeWhenUninstall.11.resources.6 } : Type `string`, default `"services"`. `hopsworks.wipeWhenUninstall[11].resources[7]` # { #helm.hopsworks.wipeWhenUninstall.11.resources.7 } : Type `string`, default `"endpoints"`. `hopsworks.wipeWhenUninstall[11].resources[8]` # { #helm.hopsworks.wipeWhenUninstall.11.resources.8 } : Type `string`, default `"namespaces"`. `hopsworks.wipeWhenUninstall[1].apiGroup` # { #helm.hopsworks.wipeWhenUninstall.1.apiGroup } : Type `string`, default `"sparkapplications.sparkoperator.k8s.io"`. `hopsworks.wipeWhenUninstall[1].resources[0]` # { #helm.hopsworks.wipeWhenUninstall.1.resources.0 } : Type `string`, default `"sparkapplications"`. `hopsworks.wipeWhenUninstall[2].apiGroup` # { #helm.hopsworks.wipeWhenUninstall.2.apiGroup } : Type `string`, default `"hopsworkscerts.certs.hopsworks.ai"`. `hopsworks.wipeWhenUninstall[2].resources[0]` # { #helm.hopsworks.wipeWhenUninstall.2.resources.0 } : Type `string`, default `"hopsworkscerts"`. `hopsworks.wipeWhenUninstall[3].apiGroup` # { #helm.hopsworks.wipeWhenUninstall.3.apiGroup } : Type `string`, default `"certs.hopsworks.ai"`. `hopsworks.wipeWhenUninstall[3].resources[0]` # { #helm.hopsworks.wipeWhenUninstall.3.resources.0 } : Type `string`, default `"hopsworkscerts"`. `hopsworks.wipeWhenUninstall[4].apiGroup` # { #helm.hopsworks.wipeWhenUninstall.4.apiGroup } : Type `string`, default `"batch"`. `hopsworks.wipeWhenUninstall[4].resources[0]` # { #helm.hopsworks.wipeWhenUninstall.4.resources.0 } : Type `string`, default `"jobs"`. `hopsworks.wipeWhenUninstall[4].resources[1]` # { #helm.hopsworks.wipeWhenUninstall.4.resources.1 } : Type `string`, default `"cronjobs"`. `hopsworks.wipeWhenUninstall[5].apiGroup` # { #helm.hopsworks.wipeWhenUninstall.5.apiGroup } : Type `string`, default `"apps"`. `hopsworks.wipeWhenUninstall[5].resources[0]` # { #helm.hopsworks.wipeWhenUninstall.5.resources.0 } : Type `string`, default `"deployments"`. `hopsworks.wipeWhenUninstall[5].resources[1]` # { #helm.hopsworks.wipeWhenUninstall.5.resources.1 } : Type `string`, default `"statefulsets"`. `hopsworks.wipeWhenUninstall[5].resources[2]` # { #helm.hopsworks.wipeWhenUninstall.5.resources.2 } : Type `string`, default `"daemonsets"`. `hopsworks.wipeWhenUninstall[6].apiGroup` # { #helm.hopsworks.wipeWhenUninstall.6.apiGroup } : Type `string`, default `"rbac.authorization.k8s.io"`. `hopsworks.wipeWhenUninstall[6].resources[0]` # { #helm.hopsworks.wipeWhenUninstall.6.resources.0 } : Type `string`, default `"roles"`. `hopsworks.wipeWhenUninstall[6].resources[1]` # { #helm.hopsworks.wipeWhenUninstall.6.resources.1 } : Type `string`, default `"rolebindings"`. `hopsworks.wipeWhenUninstall[6].resources[2]` # { #helm.hopsworks.wipeWhenUninstall.6.resources.2 } : Type `string`, default `"clusterroles"`. `hopsworks.wipeWhenUninstall[6].resources[3]` # { #helm.hopsworks.wipeWhenUninstall.6.resources.3 } : Type `string`, default `"clusterrolebindings"`. `hopsworks.wipeWhenUninstall[7].apiGroup` # { #helm.hopsworks.wipeWhenUninstall.7.apiGroup } : Type `string`, default `"networking.k8s.io"`. `hopsworks.wipeWhenUninstall[7].resources[0]` # { #helm.hopsworks.wipeWhenUninstall.7.resources.0 } : Type `string`, default `"networkpolicies"`. `hopsworks.wipeWhenUninstall[7].resources[1]` # { #helm.hopsworks.wipeWhenUninstall.7.resources.1 } : Type `string`, default `"ingresses"`. `hopsworks.wipeWhenUninstall[8].apiGroup` # { #helm.hopsworks.wipeWhenUninstall.8.apiGroup } : Type `string`, default `"storage.k8s.io"`. `hopsworks.wipeWhenUninstall[8].resources[0]` # { #helm.hopsworks.wipeWhenUninstall.8.resources.0 } : Type `string`, default `"storageclasses"`. `hopsworks.wipeWhenUninstall[9].apiGroup` # { #helm.hopsworks.wipeWhenUninstall.9.apiGroup } : Type `string`, default `"scheduling.k8s.io"`. `hopsworks.wipeWhenUninstall[9].resources[0]` # { #helm.hopsworks.wipeWhenUninstall.9.resources.0 } : Type `string`, default `"priorityclasses"`.
================================================================================ # hw-kueue Source: https://docs.hopsworks.ai/latest/setup_installation/common/helm_chart_values/hw-kueue/ # Kueue values { #helm-values-hw-kueue } Values under `hw-kueue` configure Kueue, the job queueing controller, and the queues Hopsworks schedules jobs with. _Generated from the Hopsworks Helm chart `5.2.0-alpha-1791549041` (Hopsworks `5.2.0`)._ Deployed according to the first of these values that is set: [`global._hopsworks.kueue.enabled`](global.md#helm.global._hopsworks.kueue.enabled), [`global._hopsworks.full_platform`](global.md#helm.global._hopsworks.full_platform). !!! info "Upstream charts" - Values under `hw-kueue.kueue` go to [`kueue` 0.12.2](https://github.com/kubernetes-sigs/kueue/blob/v0.12.2/charts/kueue/README.md) from `https://repo.hops.works/master/kueue/`. Only the values Hopsworks sets under `hw-kueue.kueue` are listed on this page. Any other value of the chart can be set there too; the link opens its documentation for the version Hopsworks pins. ## General { #helm-values-hw-kueue-general } ??? example "Defaults as YAML" ```yaml hw-kueue: fullnameOverride: kueue hopsworkslib: {} imagePullPolicy: Always kueue: controllerManager: featureGates: - enabled: false name: TopologyAwareScheduling imagePullSecrets: [] livenessProbe: failureThreshold: 3 initialDelaySeconds: 15 periodSeconds: 20 successThreshold: 1 timeoutSeconds: 1 manager: containerSecurityContext: allowPrivilegeEscalation: false capabilities: drop: - ALL image: pullPolicy: Always repository: docker.hops.works/registry.k8s.io/kueue/kueue podAnnotations: {} podSecurityContext: runAsNonRoot: true seccompProfile: type: RuntimeDefault resources: limits: cpu: '2' memory: 512Mi requests: cpu: 500m memory: 512Mi podDisruptionBudget: enabled: false minAvailable: 1 readinessProbe: failureThreshold: 3 initialDelaySeconds: 5 periodSeconds: 10 successThreshold: 1 timeoutSeconds: 1 replicas: 1 topologySpreadConstraints: [] enableCertManager: false enableKueueViz: false enablePrometheus: false enableVisibilityAPF: false enableVisibilityServerAuth: true fullnameOverride: kueue kubernetesClusterDomain: cluster.local metrics: prometheusNamespace: monitoring serviceMonitor: tlsConfig: insecureSkipVerify: true metricsService: annotations: {} ports: - name: metrics port: 8443 protocol: TCP targetPort: 8443 type: ClusterIP nameOverride: kueue webhookService: ipDualStack: enabled: false ipFamilies: - IPv6 - IPv4 ipFamilyPolicy: PreferDualStack ports: - port: 443 protocol: TCP targetPort: 9443 type: ClusterIP nameOverride: kueue resourceFlavours: - annotations: {} labels: {} name: default-flavor spec: nodeLabels: cloud.provider.com/region: europe nodeTaints: {} tolerations: {} topologyName: default topologies: - levels: - nodeLabel: cloud.provider.com/region - nodeLabel: cloud.provider.com/zone - nodeLabel: kubernetes.io/hostname name: default ```
`hw-kueue` # { #helm.hw-kueue } : Type `object`, default `{}`. override hw-kueue values `hw-kueue.fullnameOverride` # { #helm.hw-kueue.fullnameOverride } : Type `string`, default `"kueue"`. `hw-kueue.hopsworkslib` # { #helm.hw-kueue.hopsworkslib } : Type `object`, default `{}`. override hopsworkslib values `hw-kueue.imagePullPolicy` # { #helm.hw-kueue.imagePullPolicy } : Type `string`, default `"Always"`. `hw-kueue.kueue` # { #helm.hw-kueue.kueue } : Type `object`, passed to the [`kueue` 0.12.2](https://github.com/kubernetes-sigs/kueue/blob/v0.12.2/charts/kueue/README.md) chart, whose other values are documented there. override kueue values ??? note "Default" ```yaml controllerManager: featureGates: - enabled: false name: TopologyAwareScheduling imagePullSecrets: [] livenessProbe: failureThreshold: 3 initialDelaySeconds: 15 periodSeconds: 20 successThreshold: 1 timeoutSeconds: 1 manager: containerSecurityContext: allowPrivilegeEscalation: false capabilities: drop: - ALL image: pullPolicy: Always repository: docker.hops.works/registry.k8s.io/kueue/kueue podAnnotations: {} podSecurityContext: runAsNonRoot: true seccompProfile: type: RuntimeDefault resources: limits: cpu: '2' memory: 512Mi requests: cpu: 500m memory: 512Mi podDisruptionBudget: enabled: false minAvailable: 1 readinessProbe: failureThreshold: 3 initialDelaySeconds: 5 periodSeconds: 10 successThreshold: 1 timeoutSeconds: 1 replicas: 1 topologySpreadConstraints: [] enableCertManager: false enableKueueViz: false enablePrometheus: false enableVisibilityAPF: false enableVisibilityServerAuth: true fullnameOverride: kueue kubernetesClusterDomain: cluster.local metrics: prometheusNamespace: monitoring serviceMonitor: tlsConfig: insecureSkipVerify: true metricsService: annotations: {} ports: - name: metrics port: 8443 protocol: TCP targetPort: 8443 type: ClusterIP nameOverride: kueue webhookService: ipDualStack: enabled: false ipFamilies: - IPv6 - IPv4 ipFamilyPolicy: PreferDualStack ports: - port: 443 protocol: TCP targetPort: 9443 type: ClusterIP ``` `hw-kueue.nameOverride` # { #helm.hw-kueue.nameOverride } : Type `string`, default `"kueue"`. `hw-kueue.resourceFlavours` # { #helm.hw-kueue.resourceFlavours } : Type `list`. List of ResourceFlavors ??? note "Default" ```yaml - annotations: {} labels: {} name: default-flavor spec: nodeLabels: cloud.provider.com/region: europe nodeTaints: {} tolerations: {} topologyName: default ``` `hw-kueue.topologies[0].levels[0].nodeLabel` # { #helm.hw-kueue.topologies.0.levels.0.nodeLabel } : Type `string`, default `"cloud.provider.com/region"`. `hw-kueue.topologies[0].levels[1].nodeLabel` # { #helm.hw-kueue.topologies.0.levels.1.nodeLabel } : Type `string`, default `"cloud.provider.com/zone"`. `hw-kueue.topologies[0].levels[2].nodeLabel` # { #helm.hw-kueue.topologies.0.levels.2.nodeLabel } : Type `string`, default `"kubernetes.io/hostname"`. `hw-kueue.topologies[0].name` # { #helm.hw-kueue.topologies.0.name } : Type `string`, default `"default"`.
## clusterQueues { #helm-values-hw-kueue-clusterqueues } ??? example "Defaults as YAML" ```yaml hw-kueue: clusterQueues: - annotations: {} labels: {} name: other spec: cohort: cluster fairSharing: weight: '1' namespaceSelector: {} queueingStrategy: BestEffortFIFO resourceGroups: - coveredResources: - cpu - memory - pods - nvidia.com/gpu flavors: - name: default-flavor resources: - name: cpu nominalQuota: 0 - name: memory nominalQuota: 0 - name: pods nominalQuota: 0 - name: nvidia.com/gpu nominalQuota: 0 ```
`hw-kueue.clusterQueues[0].annotations` # { #helm.hw-kueue.clusterQueues.0.annotations } : Type `object`, default `{}`. `hw-kueue.clusterQueues[0].labels` # { #helm.hw-kueue.clusterQueues.0.labels } : Type `object`, default `{}`. `hw-kueue.clusterQueues[0].name` # { #helm.hw-kueue.clusterQueues.0.name } : Type `string`, default `"other"`. `hw-kueue.clusterQueues[0].spec.cohort` # { #helm.hw-kueue.clusterQueues.0.spec.cohort } : Type `string`, default `"cluster"`. `hw-kueue.clusterQueues[0].spec.fairSharing.weight` # { #helm.hw-kueue.clusterQueues.0.spec.fairSharing.weight } : Type `string`, default `"1"`. `hw-kueue.clusterQueues[0].spec.namespaceSelector` # { #helm.hw-kueue.clusterQueues.0.spec.namespaceSelector } : Type `object`, default `{}`. `hw-kueue.clusterQueues[0].spec.queueingStrategy` # { #helm.hw-kueue.clusterQueues.0.spec.queueingStrategy } : Type `string`, default `"BestEffortFIFO"`. `hw-kueue.clusterQueues[0].spec.resourceGroups[0].coveredResources[0]` # { #helm.hw-kueue.clusterQueues.0.spec.resourceGroups.0.coveredResources.0 } : Type `string`, default `"cpu"`. `hw-kueue.clusterQueues[0].spec.resourceGroups[0].coveredResources[1]` # { #helm.hw-kueue.clusterQueues.0.spec.resourceGroups.0.coveredResources.1 } : Type `string`, default `"memory"`. `hw-kueue.clusterQueues[0].spec.resourceGroups[0].coveredResources[2]` # { #helm.hw-kueue.clusterQueues.0.spec.resourceGroups.0.coveredResources.2 } : Type `string`, default `"pods"`. `hw-kueue.clusterQueues[0].spec.resourceGroups[0].coveredResources[3]` # { #helm.hw-kueue.clusterQueues.0.spec.resourceGroups.0.coveredResources.3 } : Type `string`, default `"nvidia.com/gpu"`. `hw-kueue.clusterQueues[0].spec.resourceGroups[0].flavors[0].name` # { #helm.hw-kueue.clusterQueues.0.spec.resourceGroups.0.flavors.0.name } : Type `string`, default `"default-flavor"`. `hw-kueue.clusterQueues[0].spec.resourceGroups[0].flavors[0].resources[0].name` # { #helm.hw-kueue.clusterQueues.0.spec.resourceGroups.0.flavors.0.resources.0.name } : Type `string`, default `"cpu"`. `hw-kueue.clusterQueues[0].spec.resourceGroups[0].flavors[0].resources[0].nominalQuota` # { #helm.hw-kueue.clusterQueues.0.spec.resourceGroups.0.flavors.0.resources.0.nominalQuota } : Type `int`, default `0`. `hw-kueue.clusterQueues[0].spec.resourceGroups[0].flavors[0].resources[1].name` # { #helm.hw-kueue.clusterQueues.0.spec.resourceGroups.0.flavors.0.resources.1.name } : Type `string`, default `"memory"`. `hw-kueue.clusterQueues[0].spec.resourceGroups[0].flavors[0].resources[1].nominalQuota` # { #helm.hw-kueue.clusterQueues.0.spec.resourceGroups.0.flavors.0.resources.1.nominalQuota } : Type `int`, default `0`. `hw-kueue.clusterQueues[0].spec.resourceGroups[0].flavors[0].resources[2].name` # { #helm.hw-kueue.clusterQueues.0.spec.resourceGroups.0.flavors.0.resources.2.name } : Type `string`, default `"pods"`. `hw-kueue.clusterQueues[0].spec.resourceGroups[0].flavors[0].resources[2].nominalQuota` # { #helm.hw-kueue.clusterQueues.0.spec.resourceGroups.0.flavors.0.resources.2.nominalQuota } : Type `int`, default `0`. `hw-kueue.clusterQueues[0].spec.resourceGroups[0].flavors[0].resources[3].name` # { #helm.hw-kueue.clusterQueues.0.spec.resourceGroups.0.flavors.0.resources.3.name } : Type `string`, default `"nvidia.com/gpu"`. `hw-kueue.clusterQueues[0].spec.resourceGroups[0].flavors[0].resources[3].nominalQuota` # { #helm.hw-kueue.clusterQueues.0.spec.resourceGroups.0.flavors.0.resources.3.nominalQuota } : Type `int`, default `0`.
## cohorts { #helm-values-hw-kueue-cohorts } ??? example "Defaults as YAML" ```yaml hw-kueue: cohorts: - annotations: {} labels: {} name: cluster spec: resourceGroups: - coveredResources: - cpu - memory - pods - nvidia.com/gpu flavors: - name: default-flavor resources: - name: cpu nominalQuota: 100 - name: memory nominalQuota: 200Gi - name: pods nominalQuota: 100 - name: nvidia.com/gpu nominalQuota: 50 ```
`hw-kueue.cohorts[0].annotations` # { #helm.hw-kueue.cohorts.0.annotations } : Type `object`, default `{}`. `hw-kueue.cohorts[0].labels` # { #helm.hw-kueue.cohorts.0.labels } : Type `object`, default `{}`. `hw-kueue.cohorts[0].name` # { #helm.hw-kueue.cohorts.0.name } : Type `string`, default `"cluster"`. `hw-kueue.cohorts[0].spec.resourceGroups[0].coveredResources[0]` # { #helm.hw-kueue.cohorts.0.spec.resourceGroups.0.coveredResources.0 } : Type `string`, default `"cpu"`. `hw-kueue.cohorts[0].spec.resourceGroups[0].coveredResources[1]` # { #helm.hw-kueue.cohorts.0.spec.resourceGroups.0.coveredResources.1 } : Type `string`, default `"memory"`. `hw-kueue.cohorts[0].spec.resourceGroups[0].coveredResources[2]` # { #helm.hw-kueue.cohorts.0.spec.resourceGroups.0.coveredResources.2 } : Type `string`, default `"pods"`. `hw-kueue.cohorts[0].spec.resourceGroups[0].coveredResources[3]` # { #helm.hw-kueue.cohorts.0.spec.resourceGroups.0.coveredResources.3 } : Type `string`, default `"nvidia.com/gpu"`. `hw-kueue.cohorts[0].spec.resourceGroups[0].flavors[0].name` # { #helm.hw-kueue.cohorts.0.spec.resourceGroups.0.flavors.0.name } : Type `string`, default `"default-flavor"`. `hw-kueue.cohorts[0].spec.resourceGroups[0].flavors[0].resources[0].name` # { #helm.hw-kueue.cohorts.0.spec.resourceGroups.0.flavors.0.resources.0.name } : Type `string`, default `"cpu"`. `hw-kueue.cohorts[0].spec.resourceGroups[0].flavors[0].resources[0].nominalQuota` # { #helm.hw-kueue.cohorts.0.spec.resourceGroups.0.flavors.0.resources.0.nominalQuota } : Type `int`, default `100`. `hw-kueue.cohorts[0].spec.resourceGroups[0].flavors[0].resources[1].name` # { #helm.hw-kueue.cohorts.0.spec.resourceGroups.0.flavors.0.resources.1.name } : Type `string`, default `"memory"`. `hw-kueue.cohorts[0].spec.resourceGroups[0].flavors[0].resources[1].nominalQuota` # { #helm.hw-kueue.cohorts.0.spec.resourceGroups.0.flavors.0.resources.1.nominalQuota } : Type `string`, default `"200Gi"`. `hw-kueue.cohorts[0].spec.resourceGroups[0].flavors[0].resources[2].name` # { #helm.hw-kueue.cohorts.0.spec.resourceGroups.0.flavors.0.resources.2.name } : Type `string`, default `"pods"`. `hw-kueue.cohorts[0].spec.resourceGroups[0].flavors[0].resources[2].nominalQuota` # { #helm.hw-kueue.cohorts.0.spec.resourceGroups.0.flavors.0.resources.2.nominalQuota } : Type `int`, default `100`. `hw-kueue.cohorts[0].spec.resourceGroups[0].flavors[0].resources[3].name` # { #helm.hw-kueue.cohorts.0.spec.resourceGroups.0.flavors.0.resources.3.name } : Type `string`, default `"nvidia.com/gpu"`. `hw-kueue.cohorts[0].spec.resourceGroups[0].flavors[0].resources[3].nominalQuota` # { #helm.hw-kueue.cohorts.0.spec.resourceGroups.0.flavors.0.resources.3.nominalQuota } : Type `int`, default `50`.
## installJob { #helm-values-hw-kueue-installjob } ??? example "Defaults as YAML" ```yaml hw-kueue: installJob: configMapMountPath: /mnt/kueue-resources configMapName: kueue-objects-install-job-resources name: kueue-obj resources: limits: cpu: 100m memory: 250Mi requests: cpu: 100m memory: 250Mi ttlSecondsAfterFinished: null ```
`hw-kueue.installJob.configMapMountPath` # { #helm.hw-kueue.installJob.configMapMountPath } : Type `string`, default `"/mnt/kueue-resources"`. `hw-kueue.installJob.configMapName` # { #helm.hw-kueue.installJob.configMapName } : Type `string`, default `"kueue-objects-install-job-resources"`. `hw-kueue.installJob.name` # { #helm.hw-kueue.installJob.name } : Type `string`, default `"kueue-obj"`. install job name `hw-kueue.installJob.resources.limits.cpu` # { #helm.hw-kueue.installJob.resources.limits.cpu } : Type `string`, default `"100m"`. `hw-kueue.installJob.resources.limits.memory` # { #helm.hw-kueue.installJob.resources.limits.memory } : Type `string`, default `"250Mi"`. `hw-kueue.installJob.resources.requests.cpu` # { #helm.hw-kueue.installJob.resources.requests.cpu } : Type `string`, default `"100m"`. `hw-kueue.installJob.resources.requests.memory` # { #helm.hw-kueue.installJob.resources.requests.memory } : Type `string`, default `"250Mi"`. `hw-kueue.installJob.ttlSecondsAfterFinished` # { #helm.hw-kueue.installJob.ttlSecondsAfterFinished } : Type `string`, default `nil`. TTL in seconds for the install-kueue-object Job. Overrides global default.
## managerConfig { #helm-values-hw-kueue-managerconfig } ??? example "Defaults as YAML" ```yaml hw-kueue: managerConfig: clientConnection: burst: 100 qps: 50 controller: groupKindConcurrency: ClusterQueue.kueue.x-k8s.io: 1 Job.batch: 5 LocalQueue.kueue.x-k8s.io: 1 Pod: 5 ResourceFlavor.kueue.x-k8s.io: 1 Workload.kueue.x-k8s.io: 5 excludedNamespaces: - kube-system - kueue-system - kyverno fairSharing: enable: true preemptionStrategies: - LessThanOrEqualToFinalShare - LessThanInitialShare health: healthProbeBindAddress: 8081 integrations: frameworks: - batch/job - kubeflow.org/mpijob - ray.io/rayjob - ray.io/raycluster - jobset.x-k8s.io/jobset - kubeflow.org/paddlejob - kubeflow.org/pytorchjob - kubeflow.org/tfjob - kubeflow.org/xgboostjob - workload.codeflare.dev/appwrapper - pod - deployment - statefulset leaderElection: leaderElect: true resourceName: c1f6bfd2.kueue.x-k8s.io manageJobsWithoutQueueName: false metrics: bindAddress: 8443 webhook: port: 9443 ```
`hw-kueue.managerConfig.clientConnection.burst` # { #helm.hw-kueue.managerConfig.clientConnection.burst } : Type `int`, default `100`. `hw-kueue.managerConfig.clientConnection.qps` # { #helm.hw-kueue.managerConfig.clientConnection.qps } : Type `int`, default `50`. `hw-kueue.managerConfig.controller.groupKindConcurrency."ClusterQueue.kueue.x-k8s.io"` # { #helm.hw-kueue.managerConfig.controller.groupKindConcurrency.ClusterQueue.kueue.x-k8s.io } : Type `int`, default `1`. `hw-kueue.managerConfig.controller.groupKindConcurrency."Job.batch"` # { #helm.hw-kueue.managerConfig.controller.groupKindConcurrency.Job.batch } : Type `int`, default `5`. `hw-kueue.managerConfig.controller.groupKindConcurrency."LocalQueue.kueue.x-k8s.io"` # { #helm.hw-kueue.managerConfig.controller.groupKindConcurrency.LocalQueue.kueue.x-k8s.io } : Type `int`, default `1`. `hw-kueue.managerConfig.controller.groupKindConcurrency."ResourceFlavor.kueue.x-k8s.io"` # { #helm.hw-kueue.managerConfig.controller.groupKindConcurrency.ResourceFlavor.kueue.x-k8s.io } : Type `int`, default `1`. `hw-kueue.managerConfig.controller.groupKindConcurrency."Workload.kueue.x-k8s.io"` # { #helm.hw-kueue.managerConfig.controller.groupKindConcurrency.Workload.kueue.x-k8s.io } : Type `int`, default `5`. `hw-kueue.managerConfig.controller.groupKindConcurrency.Pod` # { #helm.hw-kueue.managerConfig.controller.groupKindConcurrency.Pod } : Type `int`, default `5`. `hw-kueue.managerConfig.excludedNamespaces[0]` # { #helm.hw-kueue.managerConfig.excludedNamespaces.0 } : Type `string`, default `"kube-system"`. `hw-kueue.managerConfig.excludedNamespaces[1]` # { #helm.hw-kueue.managerConfig.excludedNamespaces.1 } : Type `string`, default `"kueue-system"`. `hw-kueue.managerConfig.excludedNamespaces[2]` # { #helm.hw-kueue.managerConfig.excludedNamespaces.2 } : Type `string`, default `"kyverno"`. `hw-kueue.managerConfig.fairSharing.enable` # { #helm.hw-kueue.managerConfig.fairSharing.enable } : Type `bool`, default `true`. `hw-kueue.managerConfig.fairSharing.preemptionStrategies[0]` # { #helm.hw-kueue.managerConfig.fairSharing.preemptionStrategies.0 } : Type `string`, default `"LessThanOrEqualToFinalShare"`. `hw-kueue.managerConfig.fairSharing.preemptionStrategies[1]` # { #helm.hw-kueue.managerConfig.fairSharing.preemptionStrategies.1 } : Type `string`, default `"LessThanInitialShare"`. `hw-kueue.managerConfig.health.healthProbeBindAddress` # { #helm.hw-kueue.managerConfig.health.healthProbeBindAddress } : Type `int`, default `8081`. `hw-kueue.managerConfig.integrations.frameworks[0]` # { #helm.hw-kueue.managerConfig.integrations.frameworks.0 } : Type `string`, default `"batch/job"`. `hw-kueue.managerConfig.integrations.frameworks[10]` # { #helm.hw-kueue.managerConfig.integrations.frameworks.10 } : Type `string`, default `"pod"`. `hw-kueue.managerConfig.integrations.frameworks[11]` # { #helm.hw-kueue.managerConfig.integrations.frameworks.11 } : Type `string`, default `"deployment"`. `hw-kueue.managerConfig.integrations.frameworks[12]` # { #helm.hw-kueue.managerConfig.integrations.frameworks.12 } : Type `string`, default `"statefulset"`. `hw-kueue.managerConfig.integrations.frameworks[1]` # { #helm.hw-kueue.managerConfig.integrations.frameworks.1 } : Type `string`, default `"kubeflow.org/mpijob"`. `hw-kueue.managerConfig.integrations.frameworks[2]` # { #helm.hw-kueue.managerConfig.integrations.frameworks.2 } : Type `string`, default `"ray.io/rayjob"`. `hw-kueue.managerConfig.integrations.frameworks[3]` # { #helm.hw-kueue.managerConfig.integrations.frameworks.3 } : Type `string`, default `"ray.io/raycluster"`. `hw-kueue.managerConfig.integrations.frameworks[4]` # { #helm.hw-kueue.managerConfig.integrations.frameworks.4 } : Type `string`, default `"jobset.x-k8s.io/jobset"`. `hw-kueue.managerConfig.integrations.frameworks[5]` # { #helm.hw-kueue.managerConfig.integrations.frameworks.5 } : Type `string`, default `"kubeflow.org/paddlejob"`. `hw-kueue.managerConfig.integrations.frameworks[6]` # { #helm.hw-kueue.managerConfig.integrations.frameworks.6 } : Type `string`, default `"kubeflow.org/pytorchjob"`. `hw-kueue.managerConfig.integrations.frameworks[7]` # { #helm.hw-kueue.managerConfig.integrations.frameworks.7 } : Type `string`, default `"kubeflow.org/tfjob"`. `hw-kueue.managerConfig.integrations.frameworks[8]` # { #helm.hw-kueue.managerConfig.integrations.frameworks.8 } : Type `string`, default `"kubeflow.org/xgboostjob"`. `hw-kueue.managerConfig.integrations.frameworks[9]` # { #helm.hw-kueue.managerConfig.integrations.frameworks.9 } : Type `string`, default `"workload.codeflare.dev/appwrapper"`. `hw-kueue.managerConfig.leaderElection.leaderElect` # { #helm.hw-kueue.managerConfig.leaderElection.leaderElect } : Type `bool`, default `true`. `hw-kueue.managerConfig.leaderElection.resourceName` # { #helm.hw-kueue.managerConfig.leaderElection.resourceName } : Type `string`, default `"c1f6bfd2.kueue.x-k8s.io"`. `hw-kueue.managerConfig.manageJobsWithoutQueueName` # { #helm.hw-kueue.managerConfig.manageJobsWithoutQueueName } : Type `bool`, default `false`. `hw-kueue.managerConfig.metrics.bindAddress` # { #helm.hw-kueue.managerConfig.metrics.bindAddress } : Type `int`, default `8443`. `hw-kueue.managerConfig.webhook.port` # { #helm.hw-kueue.managerConfig.webhook.port } : Type `int`, default `9443`.
## rbac { #helm-values-hw-kueue-rbac } ??? example "Defaults as YAML" ```yaml hw-kueue: rbac: clusterqueue: enabled: true cohort: enabled: true localqueue: enabled: true resourceFlavors: enabled: true topologies: enabled: true workload: enabled: true ```
`hw-kueue.rbac.clusterqueue.enabled` # { #helm.hw-kueue.rbac.clusterqueue.enabled } : Type `bool`, default `true`. `hw-kueue.rbac.cohort.enabled` # { #helm.hw-kueue.rbac.cohort.enabled } : Type `bool`, default `true`. `hw-kueue.rbac.localqueue.enabled` # { #helm.hw-kueue.rbac.localqueue.enabled } : Type `bool`, default `true`. `hw-kueue.rbac.resourceFlavors.enabled` # { #helm.hw-kueue.rbac.resourceFlavors.enabled } : Type `bool`, default `true`. `hw-kueue.rbac.topologies.enabled` # { #helm.hw-kueue.rbac.topologies.enabled } : Type `bool`, default `true`. `hw-kueue.rbac.workload.enabled` # { #helm.hw-kueue.rbac.workload.enabled } : Type `bool`, default `true`.
================================================================================ # hw-kyverno Source: https://docs.hopsworks.ai/latest/setup_installation/common/helm_chart_values/hw-kyverno/ # Kyverno values { #helm-values-hw-kyverno } Values under `hw-kyverno` configure the Kyverno cluster policies and the policy exceptions Hopsworks workloads need. _Generated from the Hopsworks Helm chart `5.2.0-alpha-1791549041` (Hopsworks `5.2.0`)._ Deployed according to the first of these values that is set: [`global._hopsworks.kyverno.enabled`](global.md#helm.global._hopsworks.kyverno.enabled), [`global._hopsworks.full_platform`](global.md#helm.global._hopsworks.full_platform). ## General { #helm-values-hw-kyverno-general } ??? example "Defaults as YAML" ```yaml hw-kyverno: hopsworkslib: {} ```
`hw-kyverno` # { #helm.hw-kyverno } : Type `object`, default `{}`. override hw-kyverno values `hw-kyverno.hopsworkslib` # { #helm.hw-kyverno.hopsworkslib } : Type `object`, default `{}`. override hopsworkslib values
## policies { #helm-values-hw-kyverno-policies } ??? example "Defaults as YAML" ```yaml hw-kyverno: policies: addCertificatesVolume: autogenControllers: DaemonSet,Deployment,Job,StatefulSet cacertsConfigMap: ca-pemstore enabled: false envs: [] extraAnnotations: [] hopsworksProjectLabelKey: hopsworks.ai/project initContainers: extraAnnotations: [] labels: [] labels: - key: job-type value: check-image-exist - key: job-type value: tag - key: job-type value: list-tags - key: job-type value: docker-build - key: job-type value: delete - key: job-type value: git-command mountPath: '' prohibitHostPath: enabled: false ```
`hw-kyverno.policies` # { #helm.hw-kyverno.policies } : Type `object`. Configuration for Hopsworks Kyverno policies and exceptions ??? note "Default" ```yaml addCertificatesVolume: autogenControllers: DaemonSet,Deployment,Job,StatefulSet cacertsConfigMap: ca-pemstore enabled: false envs: [] extraAnnotations: [] hopsworksProjectLabelKey: hopsworks.ai/project initContainers: extraAnnotations: [] labels: [] labels: - key: job-type value: check-image-exist - key: job-type value: tag - key: job-type value: list-tags - key: job-type value: docker-build - key: job-type value: delete - key: job-type value: git-command mountPath: '' prohibitHostPath: enabled: false ``` `hw-kyverno.policies.addCertificatesVolume` # { #helm.hw-kyverno.policies.addCertificatesVolume } : Type `object`. Configuration to add custom certificates to pods as a mounted volume ??? note "Default" ```yaml autogenControllers: DaemonSet,Deployment,Job,StatefulSet cacertsConfigMap: ca-pemstore enabled: false envs: [] extraAnnotations: [] hopsworksProjectLabelKey: hopsworks.ai/project initContainers: extraAnnotations: [] labels: [] labels: - key: job-type value: check-image-exist - key: job-type value: tag - key: job-type value: list-tags - key: job-type value: docker-build - key: job-type value: delete - key: job-type value: git-command mountPath: '' ``` `hw-kyverno.policies.addCertificatesVolume.autogenControllers` # { #helm.hw-kyverno.policies.addCertificatesVolume.autogenControllers } : Type `string`, default `"DaemonSet,Deployment,Job,StatefulSet"`. the list of controllers separated by comma to generate the policy for `hw-kyverno.policies.addCertificatesVolume.cacertsConfigMap` # { #helm.hw-kyverno.policies.addCertificatesVolume.cacertsConfigMap } : Type `string`, default `"ca-pemstore"`. The name of the configmap containing the certificates. The configmap should contains the ca-certificates.crt trusted by the user to replace the whole /etc/ssl/certs. It should also include the root certificate(s) if any defined for OAUTH or LDAP identity providers. Also, if using a hosted object storage with a custom CA, it should includes the java/cacerts to trust that object storage host and that require updating the hopsfs.namenode.jvmOpts = "-Djavax.net.ssl.trustStore=/etc/ssl/certs/my-cacerts -Djavax.net.ssl.trustStorePassword=changeit" and similarly for the hopsfs.datanode.jvmOpts. `hw-kyverno.policies.addCertificatesVolume.enabled` # { #helm.hw-kyverno.policies.addCertificatesVolume.enabled } : Type `bool`, default `false`. Enable add certificates volume `hw-kyverno.policies.addCertificatesVolume.envs` # { #helm.hw-kyverno.policies.addCertificatesVolume.envs } : Type `list`, default `[]`. The array of environment variables with their values that you would like to inject to the pods along side the certificates volume. `hw-kyverno.policies.addCertificatesVolume.extraAnnotations` # { #helm.hw-kyverno.policies.addCertificatesVolume.extraAnnotations } : Type `list`, default `[]`. The array of extra annotations to use when filtering which pods to inject the certificates volume into. `hw-kyverno.policies.addCertificatesVolume.hopsworksProjectLabelKey` # { #helm.hw-kyverno.policies.addCertificatesVolume.hopsworksProjectLabelKey } : Type `string`, default `"hopsworks.ai/project"`. the name of the label key to identify hopsworks user project namespaces `hw-kyverno.policies.addCertificatesVolume.initContainers` # { #helm.hw-kyverno.policies.addCertificatesVolume.initContainers } : Type `object`, default `{"extraAnnotations":[],"labels":[]}`. Opt-in for injecting the certificates volume into init containers. When enabled, a second mutate rule is rendered that targets spec.initContainers and is gated by an OR of the dedicated annotation (default kyverno-inject-certs-init=enabled), extraAnnotations, and labels configured below. Pods must opt in via at least one of these matchers. `hw-kyverno.policies.addCertificatesVolume.initContainers.extraAnnotations` # { #helm.hw-kyverno.policies.addCertificatesVolume.initContainers.extraAnnotations } : Type `list`, default `[]`. Init-specific extra annotations that opt a pod's init containers into certificate volume injection. ORed with the dedicated init-container annotation and labels below. Defaults to empty so init injection is opt-in by design. `hw-kyverno.policies.addCertificatesVolume.initContainers.labels` # { #helm.hw-kyverno.policies.addCertificatesVolume.initContainers.labels } : Type `list`, default `[]`. Init-specific labels that opt a pod's init containers into certificate volume injection. ORed with the dedicated init-container annotation and extraAnnotations. Defaults to empty so third-party operator-injected init containers are not silently mutated. `hw-kyverno.policies.addCertificatesVolume.labels` # { #helm.hw-kyverno.policies.addCertificatesVolume.labels } : Type `list`. The array of labels to use when filtering which pods to inject the certificates volume into. The default list includes job-type=check-image-exist, job-type=tag, job-type=list-tags, job-type=docker-build, job-type=delete for docker operations, that is needed if using a registry with custom CA. The default list also includes job-type=git-command for git operations, that is needed if using a git host with custom CA. ??? note "Default" ```yaml - key: job-type value: check-image-exist - key: job-type value: tag - key: job-type value: list-tags - key: job-type value: docker-build - key: job-type value: delete - key: job-type value: git-command ``` `hw-kyverno.policies.addCertificatesVolume.mountPath` # { #helm.hw-kyverno.policies.addCertificatesVolume.mountPath } : Type `string`, default `""`. Path to mount the certificates volume `hw-kyverno.policies.prohibitHostPath` # { #helm.hw-kyverno.policies.prohibitHostPath } : Type `object`, default `{"enabled":false}`. Configuration for disabling hostPath volumes in Pods `hw-kyverno.policies.prohibitHostPath.enabled` # { #helm.hw-kyverno.policies.prohibitHostPath.enabled } : Type `bool`, default `false`. Enable hostPath policy and relevant exceptions
## policyExceptions { #helm-values-hw-kyverno-policyexceptions } ??? example "Defaults as YAML" ```yaml hw-kyverno: policyExceptions: airflow: enabled: true buildkitd: appLabel: buildkitd enabled: true namePrefix: buildkitd rootless: false dockerRegistryConfigurer: enabled: true filebeat: enabled: true hopsfsCsi: enabled: true jobs: enabled: true jupyter: enabled: true knativeDryRun: enabled: true opensearch: enabled: true prometheusNodeExporter: enabled: true pythonDeployment: enabled: true pythonapp: enabled: true rondb: enabled: true spark: enabled: true rssAppName: rss-hops systemJobs: enabled: true terminal: enabled: true trino: enabled: true vllm: enabled: true ```
`hw-kyverno.policyExceptions` # { #helm.hw-kyverno.policyExceptions } : Type `object`. Configuration for PolicyExceptions to exempt specific services from Kyverno PSS Restricted policies ??? note "Default" ```yaml airflow: enabled: true buildkitd: appLabel: buildkitd enabled: true namePrefix: buildkitd rootless: false dockerRegistryConfigurer: enabled: true filebeat: enabled: true hopsfsCsi: enabled: true jobs: enabled: true jupyter: enabled: true knativeDryRun: enabled: true opensearch: enabled: true prometheusNodeExporter: enabled: true pythonDeployment: enabled: true pythonapp: enabled: true rondb: enabled: true spark: enabled: true rssAppName: rss-hops systemJobs: enabled: true terminal: enabled: true trino: enabled: true vllm: enabled: true ``` `hw-kyverno.policyExceptions.airflow` # { #helm.hw-kyverno.policyExceptions.airflow } : Type `object`, default `{"enabled":true}`. PolicyException for the Airflow Deployments. Under the legacy in-container mount (global._hopsworks.csi.enabled=false) Airflow has a mount-airflow-folders sidecar that requires privileged access for FUSE mounting of HopsFS. This exception is enabled by default and only renders when the airflow service and hw-kyverno are enabled AND the CSI integration is off: with hopsfs-csi the Airflow pods are restricted-compliant and need no exemption. `hw-kyverno.policyExceptions.airflow.enabled` # { #helm.hw-kyverno.policyExceptions.airflow.enabled } : Type `bool`, default `true`. Enable PolicyException for airflow-scheduler. Set to false to disable even when airflow service is enabled. `hw-kyverno.policyExceptions.buildkitd` # { #helm.hw-kyverno.policyExceptions.buildkitd } : Type `object`. PolicyException for the persistent BuildKit daemon (global._hopsworks.buildkitd.enabled). The system-jobs exception matches Jobs by job-type label and so never reaches this StatefulSet, which would then be rejected on a Kyverno cluster. ??? note "Default" ```yaml appLabel: buildkitd enabled: true namePrefix: buildkitd rootless: false ``` `hw-kyverno.policyExceptions.buildkitd.appLabel` # { #helm.hw-kyverno.policyExceptions.buildkitd.appLabel } : Type `string`, default `"buildkitd"`. Pod label the exception matches, together with the name prefix below. Tracks hopsworks.buildkitd.name, which is what the StatefulSet sets as its app label. `hw-kyverno.policyExceptions.buildkitd.enabled` # { #helm.hw-kyverno.policyExceptions.buildkitd.enabled } : Type `bool`, default `true`. Enable PolicyException for the persistent BuildKit daemon. Also gated on global._hopsworks.buildkitd.enabled, which is what turns the daemon itself on, so this defaults true without becoming a standing grant: the exception renders only where the daemon does. Turning it off on a Kyverno cluster that runs the daemon gets it rejected at admission, since the exception grants the rootful union (privileged, host namespaces, uid 0). `hw-kyverno.policyExceptions.buildkitd.namePrefix` # { #helm.hw-kyverno.policyExceptions.buildkitd.namePrefix } : Type `string`, default `"buildkitd"`. Name prefix the exception matches, so the grant is not reachable by anything that merely wears the app label. StatefulSet pods are -. `hw-kyverno.policyExceptions.buildkitd.rootless` # { #helm.hw-kyverno.policyExceptions.buildkitd.rootless } : Type `bool`, default `false`. Grant only the policies a rootless daemon needs, dropping the privileged-container and host-namespace exceptions. Defaults false, which grants the union, because a rootful daemon set to true is rejected at admission whereas a rootless daemon set to false merely carries two exceptions it does not use. Set true alongside hopsworks.buildkitd.rootless.enabled. `hw-kyverno.policyExceptions.dockerRegistryConfigurer` # { #helm.hw-kyverno.policyExceptions.dockerRegistryConfigurer } : Type `object`, default `{"enabled":true}`. PolicyException for docker-registry-configurer DaemonSet. This service requires privileged access, hostPID, hostNetwork, and hostPath to configure container runtimes on nodes. `hw-kyverno.policyExceptions.dockerRegistryConfigurer.enabled` # { #helm.hw-kyverno.policyExceptions.dockerRegistryConfigurer.enabled } : Type `bool`, default `true`. Enable PolicyException for docker-registry-configurer `hw-kyverno.policyExceptions.filebeat` # { #helm.hw-kyverno.policyExceptions.filebeat } : Type `object`, default `{"enabled":true}`. PolicyException for filebeat DaemonSet. Filebeat requires hostNetwork, hostPath volumes, and runs as root to collect logs from all nodes. `hw-kyverno.policyExceptions.filebeat.enabled` # { #helm.hw-kyverno.policyExceptions.filebeat.enabled } : Type `bool`, default `true`. Enable PolicyException for filebeat `hw-kyverno.policyExceptions.hopsfsCsi` # { #helm.hw-kyverno.policyExceptions.hopsfsCsi } : Type `object`, default `{"enabled":true}`. PolicyException for HopsFS CSI node DaemonSet. The csi-hopsfs-node-plugin container requires privileged access, root user, hostPath volumes and Bidirectional mount propagation to publish mounts via kubelet plugin directories. It adds no capability of its own; privileged already implies them. `hw-kyverno.policyExceptions.hopsfsCsi.enabled` # { #helm.hw-kyverno.policyExceptions.hopsfsCsi.enabled } : Type `bool`, default `true`. Enable PolicyException for HopsFS CSI node DaemonSet `hw-kyverno.policyExceptions.jobs` # { #helm.hw-kyverno.policyExceptions.jobs } : Type `object`, default `{"enabled":true}`. PolicyException for user jobs. These pods contain a hopsfsmount sidecar that requires privileged access for FUSE mounting of HopsFS. `hw-kyverno.policyExceptions.jobs.enabled` # { #helm.hw-kyverno.policyExceptions.jobs.enabled } : Type `bool`, default `true`. Enable PolicyException for user jobs `hw-kyverno.policyExceptions.jupyter` # { #helm.hw-kyverno.policyExceptions.jupyter } : Type `object`, default `{"enabled":true}`. PolicyException for Jupyter server deployments. These pods contain a hopsfsmount sidecar that requires privileged access for FUSE mounting of HopsFS. `hw-kyverno.policyExceptions.jupyter.enabled` # { #helm.hw-kyverno.policyExceptions.jupyter.enabled } : Type `bool`, default `true`. Enable PolicyException for Jupyter server deployments `hw-kyverno.policyExceptions.knativeDryRun` # { #helm.hw-kyverno.policyExceptions.knativeDryRun } : Type `object`, default `{"enabled":true}`. PolicyException for the throwaway Pod that Knative's webhook dry-run creates to validate a revision's pod spec. It carries no serving labels, so the label-scoped serving exceptions cannot match it, and a hopsfsmount FUSE sidecar makes it fail the restricted policies. `hw-kyverno.policyExceptions.knativeDryRun.enabled` # { #helm.hw-kyverno.policyExceptions.knativeDryRun.enabled } : Type `bool`, default `true`. Enable PolicyException for Knative pod-spec dry-run validation `hw-kyverno.policyExceptions.opensearch` # { #helm.hw-kyverno.policyExceptions.opensearch } : Type `object`, default `{"enabled":true}`. PolicyException for opensearch StatefulSet. OpenSearch requires a privileged init container to configure vm.max_map_count kernel parameter. Only needed when olk.opensearch.setVMMaxMapCount is true. `hw-kyverno.policyExceptions.opensearch.enabled` # { #helm.hw-kyverno.policyExceptions.opensearch.enabled } : Type `bool`, default `true`. Enable PolicyException for opensearch `hw-kyverno.policyExceptions.prometheusNodeExporter` # { #helm.hw-kyverno.policyExceptions.prometheusNodeExporter } : Type `object`, default `{"enabled":true}`. PolicyException for prometheus-node-exporter DaemonSet. Node exporter requires hostNetwork and hostPath to collect node-level metrics. `hw-kyverno.policyExceptions.prometheusNodeExporter.enabled` # { #helm.hw-kyverno.policyExceptions.prometheusNodeExporter.enabled } : Type `bool`, default `true`. Enable PolicyException for prometheus-node-exporter `hw-kyverno.policyExceptions.pythonDeployment` # { #helm.hw-kyverno.policyExceptions.pythonDeployment } : Type `object`, default `{"enabled":true}`. PolicyException for Python (model-server: python) KServe serving deployments. Agent deployments inject a root, privileged hopsfsmount FUSE sidecar (HWORKS-2871) that does not meet the restricted policy requirements. `hw-kyverno.policyExceptions.pythonDeployment.enabled` # { #helm.hw-kyverno.policyExceptions.pythonDeployment.enabled } : Type `bool`, default `true`. Enable PolicyException for Python model serving deployments `hw-kyverno.policyExceptions.pythonapp` # { #helm.hw-kyverno.policyExceptions.pythonapp } : Type `object`, default `{"enabled":true}`. PolicyException for Python app deployments (custom apps, Streamlit, Gradio). These pods contain a hopsfsmount sidecar that requires privileged access for FUSE mounting of HopsFS. `hw-kyverno.policyExceptions.pythonapp.enabled` # { #helm.hw-kyverno.policyExceptions.pythonapp.enabled } : Type `bool`, default `true`. Enable PolicyException for Python app deployments `hw-kyverno.policyExceptions.rondb` # { #helm.hw-kyverno.policyExceptions.rondb } : Type `object`, default `{"enabled":true}`. PolicyException for RonDB mysqlds StatefulSet. Mysqlds uses the SYS_NICE capability for process scheduling priority tuning. Only needed when rondb.rondb.meta.mysqld.addSysNiceCapability is true. `hw-kyverno.policyExceptions.rondb.enabled` # { #helm.hw-kyverno.policyExceptions.rondb.enabled } : Type `bool`, default `true`. Enable PolicyException for RonDB mysqlds `hw-kyverno.policyExceptions.spark` # { #helm.hw-kyverno.policyExceptions.spark } : Type `object`, default `{"enabled":true,"rssAppName":"rss-hops"}`. PolicyException for Spark-related resources including: (1) Spark driver and executor pods created by spark-operator - these pods have container-level security contexts but lack pod-level security context support in older spark-operator versions, (2) RSS (Remote Shuffle Service) coordinator and shuffle server Deployments/StatefulSets - these are dynamically created by the Uniffle controller with security contexts configured via CRD spec, and (3) the spark-operator Helm hook Job that applies CRDs on install/upgrade - the upstream chart hardcodes its securityContext without runAsNonRoot/seccompProfile and exposes no values to set them. `hw-kyverno.policyExceptions.spark.enabled` # { #helm.hw-kyverno.policyExceptions.spark.enabled } : Type `bool`, default `true`. Enable PolicyException for Spark-related resources (spark-operator driver/executor pods, RSS coordinator/shuffle server Deployments/StatefulSets, and the spark-operator CRD upgrade hook Job) `hw-kyverno.policyExceptions.spark.rssAppName` # { #helm.hw-kyverno.policyExceptions.spark.rssAppName } : Type `string`, default `"rss-hops"`. The app name of the RemoteShuffleService resource `hw-kyverno.policyExceptions.systemJobs` # { #helm.hw-kyverno.policyExceptions.systemJobs } : Type `object`, default `{"enabled":true}`. PolicyException for Hopsworks system jobs including docker operations (docker-build, check-image-exist, tag, delete, list-tags), image validation (check-image), and conda library operations (list-libraries, export-libraries, conda-search-libraries). These jobs require privileged access to run buildkit/podman for building container images. `hw-kyverno.policyExceptions.systemJobs.enabled` # { #helm.hw-kyverno.policyExceptions.systemJobs.enabled } : Type `bool`, default `true`. Enable PolicyException for Hopsworks system jobs `hw-kyverno.policyExceptions.terminal` # { #helm.hw-kyverno.policyExceptions.terminal } : Type `object`, default `{"enabled":true}`. PolicyException for terminal server deployments. These pods contain a hopsfsmount sidecar that requires privileged access for FUSE mounting of HopsFS. `hw-kyverno.policyExceptions.terminal.enabled` # { #helm.hw-kyverno.policyExceptions.terminal.enabled } : Type `bool`, default `true`. Enable PolicyException for terminal server deployments `hw-kyverno.policyExceptions.trino.enabled` # { #helm.hw-kyverno.policyExceptions.trino.enabled } : Type `bool`, default `true`. Enable PolicyException for trino `hw-kyverno.policyExceptions.vllm` # { #helm.hw-kyverno.policyExceptions.vllm } : Type `object`, default `{"enabled":true}`. PolicyException for vLLM model serving deployments created by KServe. These pods are deployed with default KServe/vLLM configurations that may not meet all restricted policy requirements. `hw-kyverno.policyExceptions.vllm.enabled` # { #helm.hw-kyverno.policyExceptions.vllm.enabled } : Type `bool`, default `true`. Enable PolicyException for vLLM model serving deployments `hw-kyverno.policyExceptions.trino` Deprecated # { #helm.hw-kyverno.policyExceptions.trino } : Type `object`, default `{"enabled":true}`. DEPRECATED and read by nothing. Excepted the Trino pods from the restricted policies while their mount sidecars were privileged root containers; they are unprivileged hopsfs-csi sidecars now and pass the policies as-is. Kept only so an override carried from 5.1 still validates.
================================================================================ # judge Source: https://docs.hopsworks.ai/latest/setup_installation/common/helm_chart_values/judge/ # Judge values { #helm-values-judge } Values under `judge` configure Judge, the active cluster arbitrator service. _Generated from the Hopsworks Helm chart `5.2.0-alpha-1791549041` (Hopsworks `5.2.0`)._ Deployed when [`global._hopsworks.full_platform`](global.md#helm.global._hopsworks.full_platform) is `true`. ??? example "Defaults as YAML" ```yaml judge: enabled: false hopsworkslib: {} image: name: /hopsworks/nginx pullPolicy: IfNotPresent registry: test-registry:6000 tag: stable-bookworm ingress: annotations: {} className: nginx host: judge.hopsworks.ai nodeSelector: {} port: 8080 regions: active: eu replica: us replicaCount: 1 resources: limits: cpu: 10m memory: 70M requests: cpu: 5m memory: 50M serviceAccount: name: hopsworks-judge serviceName: hopsworks-judge tolerations: [] topologySpreadConstraint: {} ```
`judge` # { #helm.judge } : Type `object`, default `{}`. override judge values `judge.enabled` # { #helm.judge.enabled } : Type `bool`, default `false`. `judge.hopsworkslib` # { #helm.judge.hopsworkslib } : Type `object`, default `{}`. override hopsworkslib values `judge.image.name` # { #helm.judge.image.name } : Type `string`, default `"/hopsworks/nginx"`. `judge.image.pullPolicy` # { #helm.judge.image.pullPolicy } : Type `string`, default `"IfNotPresent"`. `judge.image.registry` # { #helm.judge.image.registry } : Type `string`, default `"test-registry:6000"`. `judge.image.tag` # { #helm.judge.image.tag } : Type `string`, default `"stable-bookworm"`. `judge.ingress.annotations` # { #helm.judge.ingress.annotations } : Type `object`, default `{}`. ingress annotations `judge.ingress.className` # { #helm.judge.ingress.className } : Type `string`, default `"nginx"`. Name of the class implementing the Ingress controller `judge.ingress.host` # { #helm.judge.ingress.host } : Type `string`, default `"judge.hopsworks.ai"`. Rule for host based routing `judge.nodeSelector` # { #helm.judge.nodeSelector } : Type `object`, default `{}`. This ensures that Kubernetes schedules pods only onto nodes that match all the specified labels. `judge.port` # { #helm.judge.port } : Type `int`, default `8080`. Port Judge will be listening on internally `judge.regions.active` # { #helm.judge.regions.active } : Type `string`, default `"eu"`. Active region identifier `judge.regions.replica` # { #helm.judge.regions.replica } : Type `string`, default `"us"`. Replica region identifier `judge.replicaCount` # { #helm.judge.replicaCount } : Type `int`, default `1`. `judge.resources.limits.cpu` # { #helm.judge.resources.limits.cpu } : Type `string`, default `"10m"`. `judge.resources.limits.memory` # { #helm.judge.resources.limits.memory } : Type `string`, default `"70M"`. `judge.resources.requests.cpu` # { #helm.judge.resources.requests.cpu } : Type `string`, default `"5m"`. `judge.resources.requests.memory` # { #helm.judge.resources.requests.memory } : Type `string`, default `"50M"`. `judge.serviceAccount.name` # { #helm.judge.serviceAccount.name } : Type `string`, default `"hopsworks-judge"`. `judge.serviceName` # { #helm.judge.serviceName } : Type `string`, default `"hopsworks-judge"`. `judge.tolerations` # { #helm.judge.tolerations } : Type `list`, default `[]`. These tolerations allow Kubernetes to schedule pods on nodes with matching taints, ensuring proper placement based on cluster policies. `judge.topologySpreadConstraint` # { #helm.judge.topologySpreadConstraint } : Type `object`, default `{}`. The default topology spread constraint. If not defined the global topology spread constraint would be used instead.
================================================================================ # kafka Source: https://docs.hopsworks.ai/latest/setup_installation/common/helm_chart_values/kafka/ # Kafka values { #helm-values-kafka } Values under `kafka` configure Kafka, run by the Strimzi operator, which carries feature data on its way to the online feature store. _Generated from the Hopsworks Helm chart `5.2.0-alpha-1791549041` (Hopsworks `5.2.0`)._ Deployed when [`global._hopsworks.kafka.enabled`](global.md#helm.global._hopsworks.kafka.enabled) is `true`. !!! info "Upstream charts" - Values under `kafka.strimzi-kafka-operator` go to [`strimzi-kafka-operator` 1.2.0](https://artifacthub.io/packages/helm/strimzi/strimzi-kafka-operator/1.2.0) from `https://strimzi.io/charts/`. Only the values Hopsworks sets under `kafka.strimzi-kafka-operator` are listed on this page. Any other value of the chart can be set there too; the link opens its documentation for the version Hopsworks pins. ## General { #helm-values-kafka-general } ??? example "Defaults as YAML" ```yaml kafka: appName: kafka cleanupOnUninstall: enabled: true ttlSecondsAfterFinished: null cluster: kafka: storageClassName: null zookeeper: storageClassName: null configmap: name: kafka-configmap hopsworkslib: {} metrics: configMapKey: kafka-metrics-config enabled: true type: jmxPrometheusExporter strimzi-kafka-operator: createGlobalResources: true defaultImageRegistry: docker.hops.works defaultImageRepository: strimzi defaultImageTag: 1.2.0 enabled: true extraEnvs: - name: HOPSWORKS_FIRST_STRIMZICA_CA_CERT valueFrom: secretKeyRef: key: ca.crt name: kafka-cluster-cluster-ca-cert - name: HOPSWORKS_FIRST_STRIMZICA_CLIENT_CERT valueFrom: secretKeyRef: key: ca.crt name: kafka-cluster-clients-ca-cert - name: HOPSWORKS_FIRST_STRIMZICA_CLIENT_CA valueFrom: secretKeyRef: key: ca.key name: kafka-cluster-clients-ca - name: HOPSWORKS_FIRST_STRIMZICA_CA valueFrom: secretKeyRef: key: ca.key name: kafka-cluster-cluster-ca nodeSelector: {} podSecurityContext: fsGroup: 1000 runAsGroup: 1000 runAsNonRoot: true runAsUser: 1000 seccompProfile: type: RuntimeDefault securityContext: allowPrivilegeEscalation: false capabilities: drop: - ALL runAsGroup: 1000 runAsNonRoot: true runAsUser: 1000 seccompProfile: type: RuntimeDefault tolerations: [] watchAnyNamespace: false topics: [] users: [] ```
`kafka` # { #helm.kafka } : Type `object`. override kafka values ??? note "Default" ```yaml cluster: kafka: storageClassName: null zookeeper: storageClassName: null strimzi-kafka-operator: defaultImageRegistry: docker.hops.works ``` `kafka.appName` # { #helm.kafka.appName } : Type `string`, default `"kafka"`. `kafka.cleanupOnUninstall` # { #helm.kafka.cleanupOnUninstall } : Type `object`, default `{"enabled":true,"ttlSecondsAfterFinished":null}`. post-delete cleanup of Kafka leftovers. Strimzi's Kafka/Zookeeper StatefulSet PVCs survive uninstall; this deletes them by the strimzi.io/cluster label, but only when global._hopsworks.wipeDataOnUninstall is enabled and never for PVCs labelled hopsworks.ai/keep=true. `kafka.cleanupOnUninstall.enabled` # { #helm.kafka.cleanupOnUninstall.enabled } : Type `bool`, default `true`. enable the post-delete Kafka data-PVC cleanup hook (also requires global._hopsworks.wipeDataOnUninstall) `kafka.cleanupOnUninstall.ttlSecondsAfterFinished` # { #helm.kafka.cleanupOnUninstall.ttlSecondsAfterFinished } : Type `string`, default `nil`. ttlSecondsAfterFinished for the cleanup Job; null falls through to the global default `kafka.configmap.name` # { #helm.kafka.configmap.name } : Type `string`, default `"kafka-configmap"`. `kafka.hopsworkslib` # { #helm.kafka.hopsworkslib } : Type `object`, default `{}`. override hopsworkslib values `kafka.metrics.configMapKey` # { #helm.kafka.metrics.configMapKey } : Type `string`, default `"kafka-metrics-config"`. `kafka.metrics.enabled` # { #helm.kafka.metrics.enabled } : Type `bool`, default `true`. `kafka.metrics.type` # { #helm.kafka.metrics.type } : Type `string`, default `"jmxPrometheusExporter"`. `kafka.strimzi-kafka-operator` # { #helm.kafka.strimzi-kafka-operator } : Type `object`, passed to the [`strimzi-kafka-operator` 1.2.0](https://artifacthub.io/packages/helm/strimzi/strimzi-kafka-operator/1.2.0) chart, whose other values are documented there. override values for strimzi operator ??? note "Default" ```yaml createGlobalResources: true defaultImageRegistry: docker.hops.works defaultImageRepository: strimzi defaultImageTag: 1.2.0 enabled: true extraEnvs: - name: HOPSWORKS_FIRST_STRIMZICA_CA_CERT valueFrom: secretKeyRef: key: ca.crt name: kafka-cluster-cluster-ca-cert - name: HOPSWORKS_FIRST_STRIMZICA_CLIENT_CERT valueFrom: secretKeyRef: key: ca.crt name: kafka-cluster-clients-ca-cert - name: HOPSWORKS_FIRST_STRIMZICA_CLIENT_CA valueFrom: secretKeyRef: key: ca.key name: kafka-cluster-clients-ca - name: HOPSWORKS_FIRST_STRIMZICA_CA valueFrom: secretKeyRef: key: ca.key name: kafka-cluster-cluster-ca nodeSelector: {} podSecurityContext: fsGroup: 1000 runAsGroup: 1000 runAsNonRoot: true runAsUser: 1000 seccompProfile: type: RuntimeDefault securityContext: allowPrivilegeEscalation: false capabilities: drop: - ALL runAsGroup: 1000 runAsNonRoot: true runAsUser: 1000 seccompProfile: type: RuntimeDefault tolerations: [] watchAnyNamespace: false ``` `kafka.topics` # { #helm.kafka.topics } : Type `list`, default `[]`. A list of topics to provision in the Kafka cluster. `kafka.users` # { #helm.kafka.users } : Type `list`, default `[]`. A list of user to provision in the Kafka cluster.
## cluster { #helm-values-kafka-cluster } ??? example "Defaults as YAML" ```yaml kafka: cluster: clientsCa: validityDays: 365 clusterCa: validityDays: 365 entityOperator: {} kafka: authorizer: authorizerClass: io.hops.kafka.HopsAclAuthorizer principalBuilderClass: io.hops.kafka.HopsPrincipalBuilder superUsers: [] type: custom config: database: name: hopsworks log: retention: bytes: -1 checkIntervalMs: 300000 hours: 168 minInsyncReplicas: 1 replicationFactor: 1 dependencies: glassfish: consulServiceName: glassfish consulServiceTag: hopsworks mysql: consulServiceName: mysql port: 3306 onlinefs: consulServiceName: onlinefs externalLoadBalancer: annotations: {} bootstrapNodePort: null class: null dns: '' enabled: null finalizers: [] managed: null nodeSelector: {} startingAdvertisedPort: 9093 startingNodePort: null image: name: kafka tag: 1.2.0-kafka-4.3.1-h1 jvmOptions: {} nodeSelector: {} podAntiAffinity: required: false podSecurityContext: fsGroup: 1001 runAsGroup: 1001 runAsNonRoot: true runAsUser: 1001 seccompProfile: type: RuntimeDefault quotas: defaults: consumerByteRate: 9223372036854775807 enabled: false producerByteRate: 1048576 requestPercentage: 100 kafka: consumerByteRate: null controllerMutationRate: null producerByteRate: null requestPercentage: null minAvailableBytesPerVolume: 0 minAvailableRatioPerVolume: 0.01 strimzi: consumerByteRate: null excludedPrincipals: [] producerByteRate: null type: strimzi window: num: 11 sizeSeconds: 1 replicas: 1 resources: {} securityContext: allowPrivilegeEscalation: false capabilities: drop: - ALL runAsGroup: 1001 runAsNonRoot: true runAsUser: 1001 seccompProfile: type: RuntimeDefault services: brokers: annotations: consul.hashicorp.com/service-name: kafka consul.hashicorp.com/service-port: 9092 consul.hashicorp.com/service-tags: broker pods: annotations: prometheus.io/path: /metrics prometheus.io/port: 9404 prometheus.io/scheme: http prometheus.io/scrape: 'true' storageClassName: null storageSize: 30Gi tolerations: [] topologySpreadConstraint: {} version: 4.3.1 name: kafka-cluster tlscerts: annotations: certs.hopsworks.ai/owned-by: deployment-strimzi-cluster-operator zookeeper: nodeSelector: {} podAntiAffinity: required: false podSecurityContext: fsGroup: 1001 runAsGroup: 1001 runAsNonRoot: true runAsUser: 1001 seccompProfile: type: RuntimeDefault replicas: 1 resources: limits: cpu: '2' memory: 3Gi requests: cpu: 200m memory: 1Gi securityContext: allowPrivilegeEscalation: false capabilities: drop: - ALL runAsGroup: 1001 runAsNonRoot: true runAsUser: 1001 seccompProfile: type: RuntimeDefault services: client: annotations: consul.hashicorp.com/service-name: zookeeper consul.hashicorp.com/service-tags: client storageClassName: null storageSize: 5Gi tolerations: [] topologySpreadConstraint: {} ```
`kafka.cluster.clientsCa.validityDays` # { #helm.kafka.cluster.clientsCa.validityDays } : Type `int`, default `365`. Validity in days of the KafkaUser certificates Strimzi issues under the clients CA. Applies at the next issuance; already-issued certificates keep their dates. `kafka.cluster.clusterCa.validityDays` # { #helm.kafka.cluster.clusterCa.validityDays } : Type `int`, default `365`. Validity in days of the certificates Strimzi issues under the cluster CA (brokers, ZooKeeper, entity operator). Applies at the next issuance; already-issued certificates keep their dates. `kafka.cluster.entityOperator` # { #helm.kafka.cluster.entityOperator } : Type `object`, default `{}`. entity operator configuration for the cluster `kafka.cluster.kafka.authorizer.authorizerClass` # { #helm.kafka.cluster.kafka.authorizer.authorizerClass } : Type `string`, default `"io.hops.kafka.HopsAclAuthorizer"`. `kafka.cluster.kafka.authorizer.principalBuilderClass` # { #helm.kafka.cluster.kafka.authorizer.principalBuilderClass } : Type `string`, default `"io.hops.kafka.HopsPrincipalBuilder"`. `kafka.cluster.kafka.authorizer.superUsers` # { #helm.kafka.cluster.kafka.authorizer.superUsers } : Type `list`, default `[]`. `kafka.cluster.kafka.authorizer.type` # { #helm.kafka.cluster.kafka.authorizer.type } : Type `string`, default `"custom"`. `kafka.cluster.kafka.config.database.name` # { #helm.kafka.cluster.kafka.config.database.name } : Type `string`, default `"hopsworks"`. `kafka.cluster.kafka.config.log.retention.bytes` # { #helm.kafka.cluster.kafka.config.log.retention.bytes } : Type `int`, default `-1`. The maximum size of the log before deleting it. `kafka.cluster.kafka.config.log.retention.checkIntervalMs` # { #helm.kafka.cluster.kafka.config.log.retention.checkIntervalMs } : Type `int`, default `300000`. The interval at which log segments are checked for deletion. `kafka.cluster.kafka.config.log.retention.hours` # { #helm.kafka.cluster.kafka.config.log.retention.hours } : Type `int`, default `168`. The number of hours to keep a log segment before deleting it. `kafka.cluster.kafka.config.minInsyncReplicas` # { #helm.kafka.cluster.kafka.config.minInsyncReplicas } : Type `int`, default `1`. min.insync.replicas and transaction.state.log.min.isr: the replicas that must acknowledge a write from a producer using acks=all, the Kafka client default, before it is accepted. At most replicationFactor, and no more than the fewest replicas any existing topic has: a topic with fewer rejects acks=all writes. `kafka.cluster.kafka.config.replicationFactor` # { #helm.kafka.cluster.kafka.config.replicationFactor } : Type `int`, default `1`. Replication factor of the topics Kafka creates itself: __consumer_offsets, __transaction_state and any topic created without one. At most replicas. It applies to topics created after it is set; existing topics keep theirs. Topics Hopsworks creates take theirs from hopsworks.variables.kafka_num_replicas. `kafka.cluster.kafka.dependencies.glassfish.consulServiceName` # { #helm.kafka.cluster.kafka.dependencies.glassfish.consulServiceName } : Type `string`, default `"glassfish"`. `kafka.cluster.kafka.dependencies.glassfish.consulServiceTag` # { #helm.kafka.cluster.kafka.dependencies.glassfish.consulServiceTag } : Type `string`, default `"hopsworks"`. `kafka.cluster.kafka.dependencies.mysql.consulServiceName` # { #helm.kafka.cluster.kafka.dependencies.mysql.consulServiceName } : Type `string`, default `"mysql"`. `kafka.cluster.kafka.dependencies.mysql.port` # { #helm.kafka.cluster.kafka.dependencies.mysql.port } : Type `int`, default `3306`. `kafka.cluster.kafka.dependencies.onlinefs.consulServiceName` # { #helm.kafka.cluster.kafka.dependencies.onlinefs.consulServiceName } : Type `string`, default `"onlinefs"`. `kafka.cluster.kafka.externalLoadBalancer.annotations` # { #helm.kafka.cluster.kafka.externalLoadBalancer.annotations } : Type `object`, default `{}`. load balancer annotations `kafka.cluster.kafka.externalLoadBalancer.bootstrapNodePort` # { #helm.kafka.cluster.kafka.externalLoadBalancer.bootstrapNodePort } : Type `string`, default `nil`. Explicit nodePort for the bootstrap service when the load balancer is unmanaged (managed: false). Null lets Kubernetes allocate one from the cluster's node-port range; a set value must lie in that range (30000-32767 by default), which the API server enforces at install. `kafka.cluster.kafka.externalLoadBalancer.class` # { #helm.kafka.cluster.kafka.externalLoadBalancer.class } : Type `string`, default `nil`. load balancer class name `kafka.cluster.kafka.externalLoadBalancer.dns` # { #helm.kafka.cluster.kafka.externalLoadBalancer.dns } : Type `string`, default `""`. In case of unmanaged Load Balancers this is the advertised host for the brokers `kafka.cluster.kafka.externalLoadBalancer.enabled` # { #helm.kafka.cluster.kafka.externalLoadBalancer.enabled } : Type `string`, default `nil`. Enable External Load Balancers for Kafka cluster. If not set the .global._hopsworks.externalLoadBalancers.enabled will be used instead `kafka.cluster.kafka.externalLoadBalancer.finalizers` # { #helm.kafka.cluster.kafka.externalLoadBalancer.finalizers } : Type `list`, default `[]`. A list of finalizers which will be configured for the services created for the external listener. `kafka.cluster.kafka.externalLoadBalancer.managed` # { #helm.kafka.cluster.kafka.externalLoadBalancer.managed } : Type `string`, default `nil`. Cloud provider provisions Load Balancers. If not set the .global._hopsworks.externalLoadBalancers.managed will be used instead `kafka.cluster.kafka.externalLoadBalancer.nodeSelector` # { #helm.kafka.cluster.kafka.externalLoadBalancer.nodeSelector } : Type `object`, default `{}`. selector for nodes the load balancer can use to route traffic `kafka.cluster.kafka.externalLoadBalancer.startingAdvertisedPort` # { #helm.kafka.cluster.kafka.externalLoadBalancer.startingAdvertisedPort } : Type `int`, default `9093`. In case of unmanaged Load Balancers this is the starting advertised port for the brokers `kafka.cluster.kafka.externalLoadBalancer.startingNodePort` # { #helm.kafka.cluster.kafka.externalLoadBalancer.startingNodePort } : Type `string`, default `nil`. Explicit nodePort for broker 0 when the load balancer is unmanaged (managed: false); broker N gets startingNodePort + N, as with startingAdvertisedPort. Null lets Kubernetes allocate them from the cluster's node-port range; a set value must lie in that range (30000-32767 by default), which the API server enforces at install. `kafka.cluster.kafka.image` # { #helm.kafka.cluster.kafka.image } : Type `object`, default `{"name":"kafka","tag":"1.2.0-kafka-4.3.1-h1"}`. a tag docker.hops.works does not carry, because only the `-h` builds are published there. So clearing it also means pointing `strimzi-kafka-operator.defaultImageRegistry` at `quay.io` (or mirroring the plain upstream tags), and the same applies to the operator's Kafka Exporter and Cruise Control defaults if those are ever turned on. `kafka.cluster.kafka.jvmOptions` # { #helm.kafka.cluster.kafka.jvmOptions } : Type `object`, default `{}`. jvm options `kafka.cluster.kafka.nodeSelector` # { #helm.kafka.cluster.kafka.nodeSelector } : Type `object`, default `{}`. node selector configuration `kafka.cluster.kafka.podAntiAffinity.required` # { #helm.kafka.cluster.kafka.podAntiAffinity.required } : Type `bool`, default `false`. `kafka.cluster.kafka.podSecurityContext` # { #helm.kafka.cluster.kafka.podSecurityContext } : Type `object`. Pod-level security context for Kafka pods ??? note "Default" ```yaml fsGroup: 1001 runAsGroup: 1001 runAsNonRoot: true runAsUser: 1001 seccompProfile: type: RuntimeDefault ``` `kafka.cluster.kafka.quotas.kafka.consumerByteRate` # { #helm.kafka.cluster.kafka.quotas.kafka.consumerByteRate } : Type `string`, default `nil`. Built-in plugin only. Default bytes/sec each client may fetch from each broker. Per broker. Unset = none. `kafka.cluster.kafka.quotas.kafka.controllerMutationRate` # { #helm.kafka.cluster.kafka.quotas.kafka.controllerMutationRate } : Type `string`, default `nil`. Built-in plugin only. Default ceiling on partition create/delete mutations per second, per broker. Unset = none. `kafka.cluster.kafka.quotas.kafka.producerByteRate` # { #helm.kafka.cluster.kafka.quotas.kafka.producerByteRate } : Type `string`, default `nil`. Built-in plugin only. Default bytes/sec each client may publish to each broker. Per broker. Unset = none. `kafka.cluster.kafka.quotas.kafka.requestPercentage` # { #helm.kafka.cluster.kafka.quotas.kafka.requestPercentage } : Type `string`, default `nil`. Built-in plugin only. Default ceiling on each client's CPU use, as a percentage of a broker's network and I/O threads. Unset = none. `kafka.cluster.kafka.quotas.minAvailableBytesPerVolume` # { #helm.kafka.cluster.kafka.quotas.minAvailableBytesPerVolume } : Type `int`, default 0. Strimzi plugin only. Halt producers when available bytes per broker volume falls below this value; 0 = disabled `kafka.cluster.kafka.quotas.minAvailableRatioPerVolume` # { #helm.kafka.cluster.kafka.quotas.minAvailableRatioPerVolume } : Type `float`, default `0.01`. Strimzi plugin only. Halt producers when available ratio per broker volume falls below this value (0.0-1.0) `kafka.cluster.kafka.quotas.strimzi.consumerByteRate` # { #helm.kafka.cluster.kafka.quotas.strimzi.consumerByteRate } : Type `string`, default `nil`. Strimzi plugin only. Per-broker consumer byte-rate shared between all non-excluded clients on that broker. Unset = unlimited. `kafka.cluster.kafka.quotas.strimzi.excludedPrincipals` # { #helm.kafka.cluster.kafka.quotas.strimzi.excludedPrincipals } : Type `list`, default `[]`. Strimzi plugin only. Principals exempt from the rates above, each prefixed `User:`. Strimzi prepends its own broker and cruise-control principals. A superUser is exempt from ACLs but NOT from quotas, so without an entry here the platform's own clients are throttled by a rate meant for tenants. VERIFY ON THE CLUSTER; do not assume an entry took effect. Strimzi joins this list with ';' and the plugin splits on ';' with no escaping, so a principal whose own name contains ';' cannot be expressed. `authorizer.principalBuilderClass` decides the form: the default Kafka builder yields the certificate DN, which is expressible; this chart's default io.hops.kafka.HopsPrincipalBuilder yields the CN with every differing subject-alternative -name appended and ';'-joined, which is not. The chart fails the render on an entry that lacks the `User:` prefix or contains ';' rather than rendering an exemption that cannot match. Measured on 5.0: onlinefs stayed capped at the configured rate with its entry present. Confirm with kafka-consumer-perf-test.sh. `kafka.cluster.kafka.quotas.strimzi.producerByteRate` # { #helm.kafka.cluster.kafka.quotas.strimzi.producerByteRate } : Type `string`, default `nil`. Strimzi plugin only. Per-broker producer byte-rate shared between all non-excluded clients on that broker. Unset = unlimited. `kafka.cluster.kafka.quotas.type` # { #helm.kafka.cluster.kafka.quotas.type } : Type `string`, default `"strimzi"`. Quota plugin: "strimzi" or "kafka". "strimzi" installs Strimzi's StaticQuotaCallback - a storage guard plus per-broker rates shared between clients, with an exclusion list. "kafka" leaves Kafka's built-in plugin in place - per-user, per-broker limits resolved from config entities, and the only mode in which a KafkaUser's spec.quotas is enforced. Enabling one disables the other. `kafka.cluster.kafka.quotas.window.num` # { #helm.kafka.cluster.kafka.quotas.window.num } : Type `int`, default `11`. Applies under both plugins. Number of samples retained for the quota rate calculation. `kafka.cluster.kafka.quotas.window.sizeSeconds` # { #helm.kafka.cluster.kafka.quotas.window.sizeSeconds } : Type `int`, default `1`. Applies under both plugins. Duration of each quota sample window in seconds. `kafka.cluster.kafka.replicas` # { #helm.kafka.cluster.kafka.replicas } : Type `int`, default `1`. `kafka.cluster.kafka.resources` # { #helm.kafka.cluster.kafka.resources } : Type `object`, default `{}`. Resources of the broker pods. Applied through the broker KafkaNodePool; the v1 Kafka API has no spec.kafka.resources. `kafka.cluster.kafka.securityContext` # { #helm.kafka.cluster.kafka.securityContext } : Type `object`. Container-level security context for Kafka containers ??? note "Default" ```yaml allowPrivilegeEscalation: false capabilities: drop: - ALL runAsGroup: 1001 runAsNonRoot: true runAsUser: 1001 seccompProfile: type: RuntimeDefault ``` `kafka.cluster.kafka.services.brokers.annotations."consul.hashicorp.com/service-name"` # { #helm.kafka.cluster.kafka.services.brokers.annotations.consul.hashicorp.com-service-name } : Type `string`, default `"kafka"`. `kafka.cluster.kafka.services.brokers.annotations."consul.hashicorp.com/service-port"` # { #helm.kafka.cluster.kafka.services.brokers.annotations.consul.hashicorp.com-service-port } : Type `int`, default `9092`. `kafka.cluster.kafka.services.brokers.annotations."consul.hashicorp.com/service-tags"` # { #helm.kafka.cluster.kafka.services.brokers.annotations.consul.hashicorp.com-service-tags } : Type `string`, default `"broker"`. `kafka.cluster.kafka.services.pods.annotations."prometheus.io/path"` # { #helm.kafka.cluster.kafka.services.pods.annotations.prometheus.io-path } : Type `string`, default `"/metrics"`. `kafka.cluster.kafka.services.pods.annotations."prometheus.io/port"` # { #helm.kafka.cluster.kafka.services.pods.annotations.prometheus.io-port } : Type `int`, default `9404`. `kafka.cluster.kafka.services.pods.annotations."prometheus.io/scheme"` # { #helm.kafka.cluster.kafka.services.pods.annotations.prometheus.io-scheme } : Type `string`, default `"http"`. `kafka.cluster.kafka.services.pods.annotations."prometheus.io/scrape"` # { #helm.kafka.cluster.kafka.services.pods.annotations.prometheus.io-scrape } : Type `string`, default `"true"`. `kafka.cluster.kafka.storageClassName` # { #helm.kafka.cluster.kafka.storageClassName } : Type `string`, default `nil`. storage class name `kafka.cluster.kafka.storageSize` # { #helm.kafka.cluster.kafka.storageSize } : Type `string`, default `"30Gi"`. `kafka.cluster.kafka.tolerations` # { #helm.kafka.cluster.kafka.tolerations } : Type `list`, default `[]`. `kafka.cluster.kafka.topologySpreadConstraint` # { #helm.kafka.cluster.kafka.topologySpreadConstraint } : Type `object`, default `{}`. The default topology spread constraint. If not defined the global topology spread constraint would be used instead. `kafka.cluster.kafka.version` # { #helm.kafka.cluster.kafka.version } : Type `string`, default `"4.3.1"`. `kafka.cluster.name` # { #helm.kafka.cluster.name } : Type `string`, default `"kafka-cluster"`. `kafka.cluster.tlscerts.annotations."certs.hopsworks.ai/owned-by"` # { #helm.kafka.cluster.tlscerts.annotations.certs.hopsworks.ai-owned-by } : Type `string`, default `"deployment-strimzi-cluster-operator"`. `kafka.cluster.zookeeper` # { #helm.kafka.cluster.zookeeper } : Type `object`. Sizing and scheduling for the KRaft **controller pool**, despite the name. No ZooKeeper ensemble is deployed any more - Strimzi 1.x has no `.spec.zookeeper` - but controllers replace that ensemble one-for-one, so `kafka.controllerConfig` keeps reading these values: a 3-node quorum stays a 3-node quorum with the same storage, resources, scheduling and security context, and no values change is needed. The key keeps its name because renaming it would break every existing values file. ??? note "Default" ```yaml nodeSelector: {} podAntiAffinity: required: false podSecurityContext: fsGroup: 1001 runAsGroup: 1001 runAsNonRoot: true runAsUser: 1001 seccompProfile: type: RuntimeDefault replicas: 1 resources: limits: cpu: '2' memory: 3Gi requests: cpu: 200m memory: 1Gi securityContext: allowPrivilegeEscalation: false capabilities: drop: - ALL runAsGroup: 1001 runAsNonRoot: true runAsUser: 1001 seccompProfile: type: RuntimeDefault services: client: annotations: consul.hashicorp.com/service-name: zookeeper consul.hashicorp.com/service-tags: client storageClassName: null storageSize: 5Gi tolerations: [] topologySpreadConstraint: {} ``` `kafka.cluster.zookeeper.nodeSelector` # { #helm.kafka.cluster.zookeeper.nodeSelector } : Type `object`, default `{}`. node selector configuration `kafka.cluster.zookeeper.podSecurityContext` # { #helm.kafka.cluster.zookeeper.podSecurityContext } : Type `object`. Pod-level security context for Zookeeper pods ??? note "Default" ```yaml fsGroup: 1001 runAsGroup: 1001 runAsNonRoot: true runAsUser: 1001 seccompProfile: type: RuntimeDefault ``` `kafka.cluster.zookeeper.resources` # { #helm.kafka.cluster.zookeeper.resources } : Type `object`, default `{"limits":{"cpu":"2","memory":"3Gi"},"requests":{"cpu":"200m","memory":"1Gi"}}`. resources configuration `kafka.cluster.zookeeper.securityContext` # { #helm.kafka.cluster.zookeeper.securityContext } : Type `object`. Container-level security context for Zookeeper containers ??? note "Default" ```yaml allowPrivilegeEscalation: false capabilities: drop: - ALL runAsGroup: 1001 runAsNonRoot: true runAsUser: 1001 seccompProfile: type: RuntimeDefault ``` `kafka.cluster.zookeeper.storageClassName` # { #helm.kafka.cluster.zookeeper.storageClassName } : Type `string`, default `nil`. storage class name `kafka.cluster.zookeeper.topologySpreadConstraint` # { #helm.kafka.cluster.zookeeper.topologySpreadConstraint } : Type `object`, default `{}`. The default topology spread constraint. If not defined the global topology spread constraint would be used instead. `kafka.cluster.kafka.quotas.defaults.consumerByteRate` Deprecated # { #helm.kafka.cluster.kafka.quotas.defaults.consumerByteRate } : Type `int`, default `9223372036854775807`. Deprecated and no longer honoured. `kafka.cluster.kafka.quotas.defaults.enabled` Deprecated # { #helm.kafka.cluster.kafka.quotas.defaults.enabled } : Type `bool`, default `false`. Deprecated and no longer honoured. Use `quotas.kafka.*` with `type: kafka`, or `quotas.strimzi.*` with `type: strimzi`. Fails the render when combined with `type: kafka`; ignored under `type: strimzi`. `kafka.cluster.kafka.quotas.defaults.producerByteRate` Deprecated # { #helm.kafka.cluster.kafka.quotas.defaults.producerByteRate } : Type `int`, default `1048576`. Deprecated and no longer honoured. `kafka.cluster.kafka.quotas.defaults.requestPercentage` Deprecated # { #helm.kafka.cluster.kafka.quotas.defaults.requestPercentage } : Type `int`, default `100`. Deprecated and no longer honoured.
## crdUpgradeJob { #helm-values-kafka-crdupgradejob } ??? example "Defaults as YAML" ```yaml kafka: crdUpgradeJob: enabled: true image: name: crds tag: 1.2.0-h1 name: strimzi-crd-upgrade nodeSelector: {} resources: limits: cpu: 500m memory: 512Mi requests: cpu: 100m memory: 128Mi serviceAccount: annotations: {} tolerations: [] ttlSecondsAfterFinished: null ```
`kafka.crdUpgradeJob.enabled` # { #helm.kafka.crdUpgradeJob.enabled } : Type `bool`, default `true`. Run the Job that keeps the Strimzi CRDs in step with the operator this chart deploys, and preflights the KRaft migration. It runs pre-install and pre-upgrade: Helm applies the operator subchart's `crds/` on install only and skips any CRD that already exists, so every other path goes through here. A cluster whose objects are still stored as `v1beta2` is converted first (through the 0.51.0 bundle in `image`), which leaves both API versions served until the next upgrade renders `v1` and finishes the move to the 1.2.0 bundle. A failed *refresh* is not fatal by itself - it repairs rather than gates, so a failed apply warns and lets the upgrade continue - but the Job then refuses the upgrade when the live CRDs are too old for the Kafka resource this chart renders. The preflight checks are fatal: they only fire on states that would corrupt or roll back a cluster. `kafka.crdUpgradeJob.image` # { #helm.kafka.crdUpgradeJob.image } : Type `object`, default `{"name":"crds","tag":"1.2.0-h1"}`. Image carrying the CRD bundles the Job applies: `strimzi-crds-.yaml` for the Strimzi version this chart depends on, and the 0.51.0 bundle a cluster still storing `v1beta2` is converted through (Kubernetes refuses to drop a version still listed in a CRD's `status.storedVersions`, and 0.51.0 is the last Strimzi release to serve both). Built by `strimzi-crds` in the docker-images repo and pulled from the operator's `defaultImageRegistry`/`defaultImageRepository` like the broker image, so an airgapped install mirrors it with the rest (`vendor_images.sh` includes it). The tag's Strimzi part must match the dependency version: the Job refuses the upgrade when the bundle it needs is not in the image. `kafka.crdUpgradeJob.name` # { #helm.kafka.crdUpgradeJob.name } : Type `string`, default `"strimzi-crd-upgrade"`. `kafka.crdUpgradeJob.nodeSelector` # { #helm.kafka.crdUpgradeJob.nodeSelector } : Type `object`, default `{}`. node selector configuration `kafka.crdUpgradeJob.resources` # { #helm.kafka.crdUpgradeJob.resources } : Type `object`. resources configuration. Both halves are set deliberately: a namespace LimitRange defaults per resource, so a container that declares only requests still has limits injected - at whatever the operator's default is, which is high enough on some clusters to exhaust a ResourceQuota or stop the pod scheduling outright. The Job runs kubectl over about 1 MB of CRD YAML, so it is sized like the `tool` tier of hopsworkslib.initContainerResources. ??? note "Default" ```yaml limits: cpu: 500m memory: 512Mi requests: cpu: 100m memory: 128Mi ``` `kafka.crdUpgradeJob.serviceAccount.annotations` # { #helm.kafka.crdUpgradeJob.serviceAccount.annotations } : Type `object`, default `{}`. service account annotations `kafka.crdUpgradeJob.tolerations` # { #helm.kafka.crdUpgradeJob.tolerations } : Type `list`, default `[]`. `kafka.crdUpgradeJob.ttlSecondsAfterFinished` # { #helm.kafka.crdUpgradeJob.ttlSecondsAfterFinished } : Type `string`, default `nil`. ttlSecondsAfterFinished for the Job; null falls through to the global default
## migrationJob { #helm-values-kafka-migrationjob } ??? example "Defaults as YAML" ```yaml kafka: migrationJob: enabled: true name: kafka-kraft-migration nodeSelector: {} resources: limits: cpu: 500m memory: 512Mi requests: cpu: 100m memory: 128Mi serviceAccount: annotations: {} timeoutSeconds: 1800 tolerations: [] ttlSecondsAfterFinished: null ```
`kafka.migrationJob.enabled` # { #helm.kafka.migrationJob.enabled } : Type `bool`, default `true`. Run the pre-upgrade Job that carries a ZooKeeper-based cluster to KRaft before the rest of the upgrade is applied. It exits immediately on a cluster that is already on KRaft, so leaving it on costs one short Job per upgrade. Turning it off on a cluster that still needs migrating does not make the upgrade work - it makes it fail later, in the preflight, because the Strimzi this chart ships cannot serve a ZooKeeper cluster at all. `kafka.migrationJob.name` # { #helm.kafka.migrationJob.name } : Type `string`, default `"kafka-kraft-migration"`. `kafka.migrationJob.nodeSelector` # { #helm.kafka.migrationJob.nodeSelector } : Type `object`, default `{}`. node selector configuration `kafka.migrationJob.resources` # { #helm.kafka.migrationJob.resources } : Type `object`. resources configuration. Requests and limits both set, for the reason given on `crdUpgradeJob.resources`. This one only polls kubectl in a sleep loop. ??? note "Default" ```yaml limits: cpu: 500m memory: 512Mi requests: cpu: 100m memory: 128Mi ``` `kafka.migrationJob.serviceAccount.annotations` # { #helm.kafka.migrationJob.serviceAccount.annotations } : Type `object`, default `{}`. service account annotations `kafka.migrationJob.timeoutSeconds` # { #helm.kafka.migrationJob.timeoutSeconds } : Type `int`, default `1800`. How long to wait, in seconds, for Strimzi to reach KRaftPostMigration and then KRaft. Each wait gets this budget separately. On timeout the Job fails and so does the upgrade, but the cluster is not stranded: the operator keeps migrating regardless of this Job, so the next upgrade picks up wherever it got to. `kafka.migrationJob.tolerations` # { #helm.kafka.migrationJob.tolerations } : Type `list`, default `[]`. `kafka.migrationJob.ttlSecondsAfterFinished` # { #helm.kafka.migrationJob.ttlSecondsAfterFinished } : Type `string`, default `nil`. ttlSecondsAfterFinished for the Job; null falls through to the global default
================================================================================ # kserve Source: https://docs.hopsworks.ai/latest/setup_installation/common/helm_chart_values/kserve/ # KServe values { #helm-values-kserve } Values under `kserve` configure KServe and Knative Serving, which run model deployments. _Generated from the Hopsworks Helm chart `5.2.0-alpha-1791549041` (Hopsworks `5.2.0`)._ Deployed according to the first of these values that is set: [`hopsworks.variables.kube_kserve_installed`](hopsworks.md#helm.hopsworks.variables.kube_kserve_installed), [`global._hopsworks.full_platform`](global.md#helm.global._hopsworks.full_platform). ## General { #helm-values-kserve-general } ??? example "Defaults as YAML" ```yaml kserve: activator: minReplicas: 1 podDisruptionBudget: enabled: true minAvailable: 1 serviceAccount: annotations: {} cainjector: serviceAccount: annotations: {} cleanupOnUninstall: enabled: true ttlSecondsAfterFinished: null hopsworkslib: {} istioCRD: serviceAccount: annotations: {} kserve: servingruntime: vllmomni: imageRegistry: '' tag: v0.28.0 vllmopenai: imageRegistry: '' tag: v0.28.0 kserveUtilsImage: name: hopsworks/kserve-utils tag: 0.1.11 manager: serviceAccount: annotations: {} servingTerminationGracePeriodSeconds: 120 storageInitializer: image: hopsworks/storage-initializer tag: 5.2.0-SNAPSHOT webhook: minReplicas: 1 podDisruptionBudget: enabled: true minAvailable: 1 serviceAccount: annotations: {} ```
`kserve` # { #helm.kserve } : Type `object`. override kserve values ??? note "Default" ```yaml kserve: servingruntime: vllmomni: imageRegistry: '' tag: v0.28.0 vllmopenai: imageRegistry: '' tag: v0.28.0 ``` `kserve.activator.minReplicas` # { #helm.kserve.activator.minReplicas } : Type `int`, default `1`. `kserve.activator.podDisruptionBudget.enabled` # { #helm.kserve.activator.podDisruptionBudget.enabled } : Type `bool`, default `true`. `kserve.activator.podDisruptionBudget.minAvailable` # { #helm.kserve.activator.podDisruptionBudget.minAvailable } : Type `int`, default `1`. `kserve.activator.serviceAccount.annotations` # { #helm.kserve.activator.serviceAccount.annotations } : Type `object`, default `{}`. service account annotations `kserve.cainjector.serviceAccount.annotations` # { #helm.kserve.cainjector.serviceAccount.annotations } : Type `object`, default `{}`. service account annotations `kserve.cleanupOnUninstall` # { #helm.kserve.cleanupOnUninstall } : Type `object`, default `{"enabled":true,"ttlSecondsAfterFinished":null}`. post-delete cleanup of Istio/KServe leftovers (cert-manager-issued webhook cert Secrets, plus the istiod CA Secret and CA/leader-election ConfigMaps) that are created at runtime and so are never tracked or pruned by Helm/ArgoCD. `kserve.cleanupOnUninstall.enabled` # { #helm.kserve.cleanupOnUninstall.enabled } : Type `bool`, default `true`. enable the post-delete Istio/KServe cleanup hook `kserve.cleanupOnUninstall.ttlSecondsAfterFinished` # { #helm.kserve.cleanupOnUninstall.ttlSecondsAfterFinished } : Type `string`, default `nil`. ttlSecondsAfterFinished for the cleanup Job; null falls through to the global default `kserve.hopsworkslib` # { #helm.kserve.hopsworkslib } : Type `object`, default `{}`. override hopsworkslib values `kserve.istioCRD.serviceAccount.annotations` # { #helm.kserve.istioCRD.serviceAccount.annotations } : Type `object`, default `{}`. service account annotations `kserve.kserveUtilsImage.name` # { #helm.kserve.kserveUtilsImage.name } : Type `string`, default `"hopsworks/kserve-utils"`. `kserve.kserveUtilsImage.tag` # { #helm.kserve.kserveUtilsImage.tag } : Type `string`, default `"0.1.11"`. `kserve.manager.serviceAccount.annotations` # { #helm.kserve.manager.serviceAccount.annotations } : Type `object`, default `{}`. service account annotations `kserve.servingTerminationGracePeriodSeconds` # { #helm.kserve.servingTerminationGracePeriodSeconds } : Type `int`, default `120`. Grace period in seconds the model-serving-webhook stamps onto every serving pod, replacing the value Knative derives from the revision timeout. Must cover the in-pod log archive a disk-logging component uploads to the project's Logs dataset as it terminates (roughly fifteen seconds including SDK import and login). A ceiling, not a delay: a pod with nothing to archive exits as soon as its containers do. `kserve.storageInitializer.image` # { #helm.kserve.storageInitializer.image } : Type `string`, default `"hopsworks/storage-initializer"`. `kserve.storageInitializer.tag` # { #helm.kserve.storageInitializer.tag } : Type `string`, default `"5.2.0-SNAPSHOT"`. `kserve.webhook.minReplicas` # { #helm.kserve.webhook.minReplicas } : Type `int`, default `1`. `kserve.webhook.podDisruptionBudget.enabled` # { #helm.kserve.webhook.podDisruptionBudget.enabled } : Type `bool`, default `true`. `kserve.webhook.podDisruptionBudget.minAvailable` # { #helm.kserve.webhook.podDisruptionBudget.minAvailable } : Type `int`, default `1`. `kserve.webhook.serviceAccount.annotations` # { #helm.kserve.webhook.serviceAccount.annotations } : Type `object`, default `{}`. service account annotations
## cert_manager { #helm-values-kserve-cert_manager } ??? example "Defaults as YAML" ```yaml kserve: cert_manager: enabled: true jobResources: cainjector: limits: cpu: 200m controller: limits: cpu: 500m memory: 2Gi webhook: limits: cpu: 200m nodeSelector: {} tolerations: [] topologySpreadConstraint: {} version: v1.21.1 ```
`kserve.cert_manager.enabled` # { #helm.kserve.cert_manager.enabled } : Type `bool`, default `true`. `kserve.cert_manager.jobResources.cainjector` # { #helm.kserve.cert_manager.jobResources.cainjector } : Type `object`, default `{"limits":{"cpu":"200m"}}`. cainjector resources configuration `kserve.cert_manager.jobResources.cainjector.limits` # { #helm.kserve.cert_manager.jobResources.cainjector.limits } : Type `object`, default `{"cpu":"200m"}`. cainjector resources limits configuration `kserve.cert_manager.jobResources.controller` # { #helm.kserve.cert_manager.jobResources.controller } : Type `object`, default `{"limits":{"cpu":"500m","memory":"2Gi"}}`. controller resources configuration `kserve.cert_manager.jobResources.webhook` # { #helm.kserve.cert_manager.jobResources.webhook } : Type `object`, default `{"limits":{"cpu":"200m"}}`. webhook resources configuration `kserve.cert_manager.jobResources.webhook.limits` # { #helm.kserve.cert_manager.jobResources.webhook.limits } : Type `object`, default `{"cpu":"200m"}`. webhook resources limits configuration `kserve.cert_manager.nodeSelector` # { #helm.kserve.cert_manager.nodeSelector } : Type `object`, default `{}`. node selector configuration `kserve.cert_manager.tolerations` # { #helm.kserve.cert_manager.tolerations } : Type `list`, default `[]`. `kserve.cert_manager.topologySpreadConstraint` # { #helm.kserve.cert_manager.topologySpreadConstraint } : Type `object`, default `{}`. The default topology spread constraint. If not defined the global topology spread constraint would be used instead. `kserve.cert_manager.version` # { #helm.kserve.cert_manager.version } : Type `string`, default `"v1.21.1"`. cert-manager version. v1.21 supports Kubernetes 1.33-1.36; kserve-deps.env pins v1.17 (KServe CI), which is EOL and tops out at Kubernetes 1.33.
## crdUpgradeJob { #helm-values-kserve-crdupgradejob } ??? example "Defaults as YAML" ```yaml kserve: crdUpgradeJob: name: kserve-crd-upgrade nodeSelector: {} resources: requests: cpu: 100m memory: 128Mi serviceAccount: annotations: {} tolerations: [] ```
`kserve.crdUpgradeJob.name` # { #helm.kserve.crdUpgradeJob.name } : Type `string`, default `"kserve-crd-upgrade"`. `kserve.crdUpgradeJob.nodeSelector` # { #helm.kserve.crdUpgradeJob.nodeSelector } : Type `object`, default `{}`. node selector configuration `kserve.crdUpgradeJob.resources` # { #helm.kserve.crdUpgradeJob.resources } : Type `object`, default `{"requests":{"cpu":"100m","memory":"128Mi"}}`. resources configuration `kserve.crdUpgradeJob.serviceAccount.annotations` # { #helm.kserve.crdUpgradeJob.serviceAccount.annotations } : Type `object`, default `{}`. service account annotations `kserve.crdUpgradeJob.tolerations` # { #helm.kserve.crdUpgradeJob.tolerations } : Type `list`, default `[]`.
## dependencies { #helm-values-kserve-dependencies } ??? example "Defaults as YAML" ```yaml kserve: dependencies: hopsworks: consulServiceName: glassfish consulServiceTag: hopsworks port: 8182 registry: consulServiceName: registry port: 30443 registryProtocol: https ```
`kserve.dependencies.hopsworks.consulServiceName` # { #helm.kserve.dependencies.hopsworks.consulServiceName } : Type `string`, default `"glassfish"`. `kserve.dependencies.hopsworks.consulServiceTag` # { #helm.kserve.dependencies.hopsworks.consulServiceTag } : Type `string`, default `"hopsworks"`. `kserve.dependencies.hopsworks.port` # { #helm.kserve.dependencies.hopsworks.port } : Type `int`, default `8182`. `kserve.dependencies.registry.consulServiceName` # { #helm.kserve.dependencies.registry.consulServiceName } : Type `string`, default `"registry"`. `kserve.dependencies.registry.port` # { #helm.kserve.dependencies.registry.port } : Type `int`, default `30443`. `kserve.dependencies.registry.registryProtocol` # { #helm.kserve.dependencies.registry.registryProtocol } : Type `string`, default `"https"`.
## inferenceLogger { #helm-values-kserve-inferencelogger } ??? example "Defaults as YAML" ```yaml kserve: inferenceLogger: capabilities: {} digest: '' image: hopsworks/inference-logger limits: assemblyWorkers: 2 maxBatchRows: 512 maxBufferBytes: 134217728 maxDecodedBytes: 33554432 maxEventBytes: 8388608 maxInflightEvents: 16 maxQueuedRows: 1000 producerWorkers: 2 shutdownSeconds: 15 metrics: additionalLabels: {} enabled: false interval: 30s port: 9098 resources: limits: cpu: '1' memory: 1Gi requests: cpu: '0.1' memory: 128Mi tag: 5.2.0-SNAPSHOT ```
`kserve.inferenceLogger.capabilities` # { #helm.kserve.inferenceLogger.capabilities } : Type `object`, default `{}`. Map of verified full image references with digests to protocol tokens. Empty keeps legacy transport. `kserve.inferenceLogger.digest` # { #helm.kserve.inferenceLogger.digest } : Type `string`, default `""`. Immutable sha256 digest. When set, takes precedence over tag. `kserve.inferenceLogger.image` # { #helm.kserve.inferenceLogger.image } : Type `string`, default `"hopsworks/inference-logger"`. `kserve.inferenceLogger.limits.assemblyWorkers` # { #helm.kserve.inferenceLogger.limits.assemblyWorkers } : Type `int`, default `2`. `kserve.inferenceLogger.limits.maxBatchRows` # { #helm.kserve.inferenceLogger.limits.maxBatchRows } : Type `int`, default `512`. `kserve.inferenceLogger.limits.maxBufferBytes` # { #helm.kserve.inferenceLogger.limits.maxBufferBytes } : Type `int`, default `134217728`. `kserve.inferenceLogger.limits.maxDecodedBytes` # { #helm.kserve.inferenceLogger.limits.maxDecodedBytes } : Type `int`, default `33554432`. `kserve.inferenceLogger.limits.maxEventBytes` # { #helm.kserve.inferenceLogger.limits.maxEventBytes } : Type `int`, default `8388608`. `kserve.inferenceLogger.limits.maxInflightEvents` # { #helm.kserve.inferenceLogger.limits.maxInflightEvents } : Type `int`, default `16`. `kserve.inferenceLogger.limits.maxQueuedRows` # { #helm.kserve.inferenceLogger.limits.maxQueuedRows } : Type `int`, default `1000`. `kserve.inferenceLogger.limits.producerWorkers` # { #helm.kserve.inferenceLogger.limits.producerWorkers } : Type `int`, default `2`. `kserve.inferenceLogger.limits.shutdownSeconds` # { #helm.kserve.inferenceLogger.limits.shutdownSeconds } : Type `int`, default `15`. `kserve.inferenceLogger.metrics.additionalLabels` # { #helm.kserve.inferenceLogger.metrics.additionalLabels } : Type `object`, default `{}`. `kserve.inferenceLogger.metrics.enabled` # { #helm.kserve.inferenceLogger.metrics.enabled } : Type `bool`, default `false`. Create a PodMonitor when its CRD is available. KServe's existing scrape configuration is preserved. `kserve.inferenceLogger.metrics.interval` # { #helm.kserve.inferenceLogger.metrics.interval } : Type `string`, default `"30s"`. `kserve.inferenceLogger.metrics.port` # { #helm.kserve.inferenceLogger.metrics.port } : Type `int`, default `9098`. `kserve.inferenceLogger.resources.limits.cpu` # { #helm.kserve.inferenceLogger.resources.limits.cpu } : Type `string`, default `"1"`. `kserve.inferenceLogger.resources.limits.memory` # { #helm.kserve.inferenceLogger.resources.limits.memory } : Type `string`, default `"1Gi"`. `kserve.inferenceLogger.resources.requests.cpu` # { #helm.kserve.inferenceLogger.resources.requests.cpu } : Type `string`, default `"0.1"`. `kserve.inferenceLogger.resources.requests.memory` # { #helm.kserve.inferenceLogger.resources.requests.memory } : Type `string`, default `"128Mi"`. `kserve.inferenceLogger.tag` # { #helm.kserve.inferenceLogger.tag } : Type `string`, default `"5.2.0-SNAPSHOT"`.
## istio { #helm-values-kserve-istio } ??? example "Defaults as YAML" ```yaml kserve: istio: crdClassInstallServiceAccount: apiGroups: - networking.istio.io - security.istio.io - rbac.istio.io - authentication.istio.io name: istio-crd-class-install-sa resources: - '*' verbs: - get - list - watch - create - update - patch - delete envoyFilter: corsAllowedOrigins: - https://hopsworks.ai.local jobName: envoyfilterjob resources: limits: cpu: 200m gateways: clusterLocal: podDisruptionBudget: enabled: true minAvailable: 1 replicaCount: 1 http10: false ingress: http10: false http2Port: 32080 httpsPort: 32443 name: istio-ingressgateway podDisruptionBudget: enabled: true minAvailable: 1 replicaCount: 1 statusPort: 32021 jobName: gateway-operator-job operatorConfigFileName: istiooperator.yaml operatorConfigMountPath: /tmp ingressClass: enabled: true name: istio istioctl: currentVersion: 1.29.7 installJobName: istioctl-install-job jobRetries: 4 name: istioctl readinessTimeout: 15m0s resources: requests: cpu: 200m memory: 256Mi serviceAccount: annotations: {} targetVersion: 1.29.7 uninstallJobName: istioctl-uninstall-job nodeSelector: {} pilot: podDisruptionBudget: enabled: true minAvailable: 1 replicaCount: 1 tolerations: [] ```
`kserve.istio.crdClassInstallServiceAccount.apiGroups[0]` # { #helm.kserve.istio.crdClassInstallServiceAccount.apiGroups.0 } : Type `string`, default `"networking.istio.io"`. `kserve.istio.crdClassInstallServiceAccount.apiGroups[1]` # { #helm.kserve.istio.crdClassInstallServiceAccount.apiGroups.1 } : Type `string`, default `"security.istio.io"`. `kserve.istio.crdClassInstallServiceAccount.apiGroups[2]` # { #helm.kserve.istio.crdClassInstallServiceAccount.apiGroups.2 } : Type `string`, default `"rbac.istio.io"`. `kserve.istio.crdClassInstallServiceAccount.apiGroups[3]` # { #helm.kserve.istio.crdClassInstallServiceAccount.apiGroups.3 } : Type `string`, default `"authentication.istio.io"`. `kserve.istio.crdClassInstallServiceAccount.name` # { #helm.kserve.istio.crdClassInstallServiceAccount.name } : Type `string`, default `"istio-crd-class-install-sa"`. `kserve.istio.crdClassInstallServiceAccount.resources[0]` # { #helm.kserve.istio.crdClassInstallServiceAccount.resources.0 } : Type `string`, default `"*"`. `kserve.istio.crdClassInstallServiceAccount.verbs[0]` # { #helm.kserve.istio.crdClassInstallServiceAccount.verbs.0 } : Type `string`, default `"get"`. `kserve.istio.crdClassInstallServiceAccount.verbs[1]` # { #helm.kserve.istio.crdClassInstallServiceAccount.verbs.1 } : Type `string`, default `"list"`. `kserve.istio.crdClassInstallServiceAccount.verbs[2]` # { #helm.kserve.istio.crdClassInstallServiceAccount.verbs.2 } : Type `string`, default `"watch"`. `kserve.istio.crdClassInstallServiceAccount.verbs[3]` # { #helm.kserve.istio.crdClassInstallServiceAccount.verbs.3 } : Type `string`, default `"create"`. `kserve.istio.crdClassInstallServiceAccount.verbs[4]` # { #helm.kserve.istio.crdClassInstallServiceAccount.verbs.4 } : Type `string`, default `"update"`. `kserve.istio.crdClassInstallServiceAccount.verbs[5]` # { #helm.kserve.istio.crdClassInstallServiceAccount.verbs.5 } : Type `string`, default `"patch"`. `kserve.istio.crdClassInstallServiceAccount.verbs[6]` # { #helm.kserve.istio.crdClassInstallServiceAccount.verbs.6 } : Type `string`, default `"delete"`. `kserve.istio.envoyFilter.corsAllowedOrigins` # { #helm.kserve.istio.envoyFilter.corsAllowedOrigins } : Type `list`, default `["https://hopsworks.ai.local"]`. List of allowed origins for CORS requests. Defaults to the hopsworks ingress host. Update this when changing hopsworks.ingress.host. An empty list disables adding CORS headers (no CORS configuration will be applied). CORS responses for the listed origins are configured to allow credentials by default. Use \["*"\] to allow all origins (not recommended with credentials). `kserve.istio.envoyFilter.jobName` # { #helm.kserve.istio.envoyFilter.jobName } : Type `string`, default `"envoyfilterjob"`. `kserve.istio.envoyFilter.resources` # { #helm.kserve.istio.envoyFilter.resources } : Type `object`, default `{"limits":{"cpu":"200m"}}`. resources configuration `kserve.istio.envoyFilter.resources.limits` # { #helm.kserve.istio.envoyFilter.resources.limits } : Type `object`, default `{"cpu":"200m"}`. resources limits configuration `kserve.istio.gateways.clusterLocal.podDisruptionBudget.enabled` # { #helm.kserve.istio.gateways.clusterLocal.podDisruptionBudget.enabled } : Type `bool`, default `true`. `kserve.istio.gateways.clusterLocal.podDisruptionBudget.minAvailable` # { #helm.kserve.istio.gateways.clusterLocal.podDisruptionBudget.minAvailable } : Type `int`, default `1`. `kserve.istio.gateways.clusterLocal.replicaCount` # { #helm.kserve.istio.gateways.clusterLocal.replicaCount } : Type `int`, default `1`. `kserve.istio.gateways.ingress.http10` # { #helm.kserve.istio.gateways.ingress.http10 } : Type `bool`, default `false`. Allow HTTP/1.0 requests through the ingress gateway `kserve.istio.gateways.ingress.http2Port` # { #helm.kserve.istio.gateways.ingress.http2Port } : Type `int`, default `32080`. `kserve.istio.gateways.ingress.httpsPort` # { #helm.kserve.istio.gateways.ingress.httpsPort } : Type `int`, default `32443`. `kserve.istio.gateways.ingress.name` # { #helm.kserve.istio.gateways.ingress.name } : Type `string`, default `"istio-ingressgateway"`. `kserve.istio.gateways.ingress.podDisruptionBudget.enabled` # { #helm.kserve.istio.gateways.ingress.podDisruptionBudget.enabled } : Type `bool`, default `true`. `kserve.istio.gateways.ingress.podDisruptionBudget.minAvailable` # { #helm.kserve.istio.gateways.ingress.podDisruptionBudget.minAvailable } : Type `int`, default `1`. `kserve.istio.gateways.ingress.replicaCount` # { #helm.kserve.istio.gateways.ingress.replicaCount } : Type `int`, default `1`. `kserve.istio.gateways.ingress.statusPort` # { #helm.kserve.istio.gateways.ingress.statusPort } : Type `int`, default `32021`. `kserve.istio.gateways.jobName` # { #helm.kserve.istio.gateways.jobName } : Type `string`, default `"gateway-operator-job"`. `kserve.istio.gateways.operatorConfigFileName` # { #helm.kserve.istio.gateways.operatorConfigFileName } : Type `string`, default `"istiooperator.yaml"`. `kserve.istio.gateways.operatorConfigMountPath` # { #helm.kserve.istio.gateways.operatorConfigMountPath } : Type `string`, default `"/tmp"`. `kserve.istio.ingressClass.enabled` # { #helm.kserve.istio.ingressClass.enabled } : Type `bool`, default `true`. Whether to create a cluster-scoped `IngressClass` named after `name` below, for KServe Standard-mode InferenceServices (`serving.kserve.io/deploymentMode: Standard`). KServe's controller writes this class name into the `spec.ingressClassName` of the `networking.k8s.io/v1` Ingress objects it creates for those InferenceServices. `kserve.istio.ingressClass.name` # { #helm.kserve.istio.ingressClass.name } : Type `string`, default `"istio"`. Name of the IngressClass. Must match `kserve.controller.gateway.ingressGateway.className`, which is what KServe reads from the `inferenceservice-config` ConfigMap to populate `spec.ingressClassName`. `kserve.istio.istioctl.currentVersion` # { #helm.kserve.istio.istioctl.currentVersion } : Type `string`, default `"1.29.7"`. `kserve.istio.istioctl.installJobName` # { #helm.kserve.istio.istioctl.installJobName } : Type `string`, default `"istioctl-install-job"`. `kserve.istio.istioctl.jobRetries` # { #helm.kserve.istio.istioctl.jobRetries } : Type `int`, default `4`. `kserve.istio.istioctl.name` # { #helm.kserve.istio.istioctl.name } : Type `string`, default `"istioctl"`. `kserve.istio.istioctl.readinessTimeout` # { #helm.kserve.istio.istioctl.readinessTimeout } : Type `string`, default `"15m0s"`. `kserve.istio.istioctl.resources` # { #helm.kserve.istio.istioctl.resources } : Type `object`, default `{"requests":{"cpu":"200m","memory":"256Mi"}}`. resources configuration `kserve.istio.istioctl.serviceAccount.annotations` # { #helm.kserve.istio.istioctl.serviceAccount.annotations } : Type `object`, default `{}`. service account annotations `kserve.istio.istioctl.targetVersion` # { #helm.kserve.istio.istioctl.targetVersion } : Type `string`, default `"1.29.7"`. `kserve.istio.istioctl.uninstallJobName` # { #helm.kserve.istio.istioctl.uninstallJobName } : Type `string`, default `"istioctl-uninstall-job"`. `kserve.istio.nodeSelector` # { #helm.kserve.istio.nodeSelector } : Type `object`, default `{}`. node selector configuration `kserve.istio.pilot.podDisruptionBudget.enabled` # { #helm.kserve.istio.pilot.podDisruptionBudget.enabled } : Type `bool`, default `true`. `kserve.istio.pilot.podDisruptionBudget.minAvailable` # { #helm.kserve.istio.pilot.podDisruptionBudget.minAvailable } : Type `int`, default `1`. `kserve.istio.pilot.replicaCount` # { #helm.kserve.istio.pilot.replicaCount } : Type `int`, default `1`. `kserve.istio.tolerations` # { #helm.kserve.istio.tolerations } : Type `list`, default `[]`. `kserve.istio.gateways.http10` Deprecated # { #helm.kserve.istio.gateways.http10 } : Type `bool`, default `false`. Deprecated, use gateways.ingress.http10 instead. Still honoured: HTTP/1.0 is enabled when either this or gateways.ingress.http10 is true.
## knative { #helm-values-kserve-knative } ??? example "Defaults as YAML" ```yaml kserve: knative: autoscaler: scaleToZeroGracePeriod: 30s scaleToZeroPodRetentionPeriod: 0s configFeatures: affinity: enabled emptyDir: enabled initContainers: enabled nodeSelectors: enabled priorityClassName: enabled securePodDefaults: disabled securityContext: enabled tolerations: enabled volumesCsi: enabled volumesMountPropagation: enabled deployment: progressDeadline: 1800s domainName: hopsworks.ai gateways: jobName: knative-gateways-job resources: limits: cpu: 400m http: enabled: true https: credentialsName: '' enabled: false imageRegistry: '' initialDelaySeconds: 180 netIstioController: resources: limits: cpu: 300m memory: 400Mi requests: cpu: 30m memory: 40Mi nodeSelector: {} peerAuthenticationsJob: name: knative-serving-peer-authentications-job resources: {} queueSidecarImage: name: kserve/qpext tag: v0.21.0 registryCertMount: certFileName: ca.crt external: false mountDir: /etc/registry_certs volumeName: registry-certs serviceAccount: annotations: {} serviceAnnotations: prometheus.io/path: /metrics prometheus.io/port: 9090 prometheus.io/scheme: http prometheus.io/scrape: 'true' tolerations: [] topologySpreadConstraint: {} validatingWebhookDeleteJobName: knative-serving-validating-webhook-delete-job validatingWebhookName: validation.webhook.serving.knative.dev version: v1.23.0 ```
`kserve.knative.autoscaler.scaleToZeroGracePeriod` # { #helm.kserve.knative.autoscaler.scaleToZeroGracePeriod } : Type `string`, default `"30s"`. `kserve.knative.autoscaler.scaleToZeroPodRetentionPeriod` # { #helm.kserve.knative.autoscaler.scaleToZeroPodRetentionPeriod } : Type `string`, default `"0s"`. `kserve.knative.configFeatures.affinity` # { #helm.kserve.knative.configFeatures.affinity } : Type `string`, default `"enabled"`. `kserve.knative.configFeatures.emptyDir` # { #helm.kserve.knative.configFeatures.emptyDir } : Type `string`, default `"enabled"`. `kserve.knative.configFeatures.initContainers` # { #helm.kserve.knative.configFeatures.initContainers } : Type `string`, default `"enabled"`. `kserve.knative.configFeatures.nodeSelectors` # { #helm.kserve.knative.configFeatures.nodeSelectors } : Type `string`, default `"enabled"`. `kserve.knative.configFeatures.priorityClassName` # { #helm.kserve.knative.configFeatures.priorityClassName } : Type `string`, default `"enabled"`. `kserve.knative.configFeatures.securePodDefaults` # { #helm.kserve.knative.configFeatures.securePodDefaults } : Type `string`, default `"disabled"`. Knative "secure-pod-defaults" feature flag. Keep "disabled" so the HopsFS FUSE sidecar (runs as root/privileged) keeps working once Knative flips its built-in default from "disabled" to "AllowRootBounded" (planned ~1.22). Possible values are "disabled", "AllowRootBounded" or "enabled". `kserve.knative.configFeatures.securityContext` # { #helm.kserve.knative.configFeatures.securityContext } : Type `string`, default `"enabled"`. `kserve.knative.configFeatures.tolerations` # { #helm.kserve.knative.configFeatures.tolerations } : Type `string`, default `"enabled"`. `kserve.knative.configFeatures.volumesCsi` # { #helm.kserve.knative.configFeatures.volumesCsi } : Type `string`, default `"enabled"`. Knative "kubernetes.podspec-volumes-csi" feature gate, "disabled" by default in Knative. With the HopsFS CSI driver on (csi_driver_enabled), the backend gives Knative-mode agent and Python deployments the same inline csi: HopsFS volume plus unprivileged hopsfs-fuse sidecar that Jupyter, jobs, terminals and RawDeployment serving use, instead of the legacy root and privileged hopsfsmount container (HWORKS-3312). Knative's Revision validation only lets a csi: volume source through when this gate is enabled; with it disabled those deployments are rejected at admission. Possible values are "disabled" or "enabled". `kserve.knative.configFeatures.volumesMountPropagation` # { #helm.kserve.knative.configFeatures.volumesMountPropagation } : Type `string`, default `"enabled"`. Knative "kubernetes.podspec-volumes-mount-propagation" feature gate. The gate does not exist in Knative 1.13 (there was nothing to set) and defaults to "disabled" in 1.21. No Hopsworks serving path depends on it: the backend omits mountPropagation from InferenceService containers and the model-serving-webhook injects HostToContainer onto the pod, after Knative admission. It is enabled purely as headroom, so that an InferenceService carrying a permitted mountPropagation is not rejected. It does not rescue a spec written by a pre-upgrade backend: those set Bidirectional on the HopsFS mount plus privileged on the sidecar, and Knative rejects both however this gate is set (it accepts only None and HostToContainer). Possible values are "disabled" or "enabled". `kserve.knative.deployment.progressDeadline` # { #helm.kserve.knative.deployment.progressDeadline } : Type `string`, default `"1800s"`. `kserve.knative.domainName` # { #helm.kserve.knative.domainName } : Type `string`, default `"hopsworks.ai"`. `kserve.knative.gateways.jobName` # { #helm.kserve.knative.gateways.jobName } : Type `string`, default `"knative-gateways-job"`. `kserve.knative.gateways.resources` # { #helm.kserve.knative.gateways.resources } : Type `object`, default `{"limits":{"cpu":"400m"}}`. resources configuration `kserve.knative.gateways.resources.limits` # { #helm.kserve.knative.gateways.resources.limits } : Type `object`, default `{"cpu":"400m"}`. resources limits configuration `kserve.knative.http.enabled` # { #helm.kserve.knative.http.enabled } : Type `bool`, default `true`. `kserve.knative.https.credentialsName` # { #helm.kserve.knative.https.credentialsName } : Type `string`, default `""`. `kserve.knative.https.enabled` # { #helm.kserve.knative.https.enabled } : Type `bool`, default `false`. `kserve.knative.imageRegistry` # { #helm.kserve.knative.imageRegistry } : Type `string`, default `""`. Registry prefix for the seven upstream Knative images (serving controller, activator, autoscaler, webhook, queue; net-istio controller, webhook), which the chart references as /gcr.io/knative-releases/knative.dev/:. Empty means global._hopsworks.imageRegistry. Set it to pull a staged Knative build from another path without moving every other image, the same way servingruntime.*.imageRegistry does. `kserve.knative.initialDelaySeconds` # { #helm.kserve.knative.initialDelaySeconds } : Type `int`, default `180`. `kserve.knative.netIstioController.resources.limits.cpu` # { #helm.kserve.knative.netIstioController.resources.limits.cpu } : Type `string`, default `"300m"`. `kserve.knative.netIstioController.resources.limits.memory` # { #helm.kserve.knative.netIstioController.resources.limits.memory } : Type `string`, default `"400Mi"`. `kserve.knative.netIstioController.resources.requests.cpu` # { #helm.kserve.knative.netIstioController.resources.requests.cpu } : Type `string`, default `"30m"`. `kserve.knative.netIstioController.resources.requests.memory` # { #helm.kserve.knative.netIstioController.resources.requests.memory } : Type `string`, default `"40Mi"`. `kserve.knative.nodeSelector` # { #helm.kserve.knative.nodeSelector } : Type `object`, default `{}`. node selector configuration `kserve.knative.peerAuthenticationsJob.name` # { #helm.kserve.knative.peerAuthenticationsJob.name } : Type `string`, default `"knative-serving-peer-authentications-job"`. `kserve.knative.peerAuthenticationsJob.resources` # { #helm.kserve.knative.peerAuthenticationsJob.resources } : Type `object`, default `{}`. resources configuration `kserve.knative.queueSidecarImage.name` # { #helm.kserve.knative.queueSidecarImage.name } : Type `string`, default `"kserve/qpext"`. `kserve.knative.queueSidecarImage.tag` # { #helm.kserve.knative.queueSidecarImage.tag } : Type `string`, default `"v0.21.0"`. `kserve.knative.registryCertMount.certFileName` # { #helm.kserve.knative.registryCertMount.certFileName } : Type `string`, default `"ca.crt"`. `kserve.knative.registryCertMount.external` # { #helm.kserve.knative.registryCertMount.external } : Type `bool`, default `false`. `kserve.knative.registryCertMount.mountDir` # { #helm.kserve.knative.registryCertMount.mountDir } : Type `string`, default `"/etc/registry_certs"`. `kserve.knative.registryCertMount.volumeName` # { #helm.kserve.knative.registryCertMount.volumeName } : Type `string`, default `"registry-certs"`. `kserve.knative.serviceAccount.annotations` # { #helm.kserve.knative.serviceAccount.annotations } : Type `object`, default `{}`. service account annotations `kserve.knative.serviceAnnotations."prometheus.io/path"` # { #helm.kserve.knative.serviceAnnotations.prometheus.io-path } : Type `string`, default `"/metrics"`. `kserve.knative.serviceAnnotations."prometheus.io/port"` # { #helm.kserve.knative.serviceAnnotations.prometheus.io-port } : Type `int`, default `9090`. `kserve.knative.serviceAnnotations."prometheus.io/scheme"` # { #helm.kserve.knative.serviceAnnotations.prometheus.io-scheme } : Type `string`, default `"http"`. `kserve.knative.serviceAnnotations."prometheus.io/scrape"` # { #helm.kserve.knative.serviceAnnotations.prometheus.io-scrape } : Type `string`, default `"true"`. `kserve.knative.tolerations` # { #helm.kserve.knative.tolerations } : Type `list`, default `[]`. `kserve.knative.topologySpreadConstraint` # { #helm.kserve.knative.topologySpreadConstraint } : Type `object`, default `{}`. The default topology spread constraint. If not defined the global topology spread constraint would be used instead. `kserve.knative.validatingWebhookDeleteJobName` # { #helm.kserve.knative.validatingWebhookDeleteJobName } : Type `string`, default `"knative-serving-validating-webhook-delete-job"`. `kserve.knative.validatingWebhookName` # { #helm.kserve.knative.validatingWebhookName } : Type `string`, default `"validation.webhook.serving.knative.dev"`. `kserve.knative.version` # { #helm.kserve.knative.version } : Type `string`, default `"v1.23.0"`.
## kserve { #helm-values-kserve-kserve } ??? example "Defaults as YAML" ```yaml kserve: kserve: agent: image: kserve/agent tag: v0.21.0 controller: affinity: {} annotations: {} containerSecurityContext: allowPrivilegeEscalation: false capabilities: drop: - ALL privileged: false readOnlyRootFilesystem: true runAsNonRoot: true seccompProfile: type: RuntimeDefault deploymentMode: Knative gateway: additionalIngressDomains: [] disableIngressCreation: false disableIstioVirtualHost: false domain: '' domainTemplate: '{{ .Name }}.{{ .Namespace }}.{{ .IngressDomain }}' ingressGateway: className: istio enableGatewayApi: false gateway: knative-ingress-gateway localGateway: gateway: knative-local-gateway gatewayService: knative-local-gateway knativeGatewayService: '' urlScheme: http image: kserve/kserve-controller knativeAddressableResolver: enabled: false labels: {} nodeSelector: {} podAnnotations: {} podDisruptionBudget: enabled: true minAvailable: 1 podLabels: {} rbacProxy: resources: limits: cpu: 100m memory: 300Mi requests: cpu: 100m memory: 300Mi securityContext: allowPrivilegeEscalation: false capabilities: drop: - ALL privileged: false readOnlyRootFilesystem: true runAsNonRoot: true seccompProfile: type: RuntimeDefault rbacProxyImage: quay.io/brancz/kube-rbac-proxy:v0.18.0 replicaCount: 1 resources: limits: cpu: 100m memory: 300Mi requests: cpu: 100m memory: 300Mi securityContext: runAsNonRoot: true seccompProfile: type: RuntimeDefault serviceAccount: annotations: {} tag: v0.21.0 tolerations: [] topologySpreadConstraint: {} topologySpreadConstraints: [] localmodel: controller: image: kserve/kserve-localmodel-controller tag: v0.21.0 enabled: false jobNamespace: kserve-localmodel-jobs securityContext: FSGroup: 1000 serviceAccount: annotations: {} metricsaggregator: enableMetricAggregation: 'true' enablePrometheusScraping: 'true' nodeSelector: {} router: image: kserve/router tag: v0.21.0 serviceAccount: annotations: {} servingruntime: modelNamePlaceholder: '{{.Name}}' sklearnserver: image: hopsworks/sklearnserver tag: 0.21.0 tensorflow: image: tensorflow/serving tag: 2.20.0 vllmomni: env: [] image: vllm/vllm-omni imageRegistry: '' tag: v0.28.0 vllmopenai: env: [] image: vllm/vllm-openai imageRegistry: '' tag: v0.28.0 storage: caBundleConfigMapName: '' caBundleVolumeMountPath: /etc/ssl/custom-certs cpuModelcar: 10m enableModelcar: false image: hopsworks/storage-initializer memoryModelcar: 15Mi resources: limits: cpu: '1' memory: 1Gi requests: cpu: 100m memory: 100Mi s3: CABundle: '' accessKeyIdName: AWS_ACCESS_KEY_ID endpoint: '' region: '' secretAccessKeyName: AWS_SECRET_ACCESS_KEY useAnonymousCredential: '' useHttps: '' useVirtualBucket: '' verifySSL: '' storageSecretNameAnnotation: serving.kserve.io/secretName storageSpecSecretName: storage-config tag: 5.2.0-SNAPSHOT uidModelcar: 1010 tolerations: [] topologySpreadConstraint: {} version: v0.21.0 ```
`kserve.kserve.agent.image` # { #helm.kserve.kserve.agent.image } : Type `string`, default `"kserve/agent"`. `kserve.kserve.agent.tag` # { #helm.kserve.kserve.agent.tag } : Type `string`, default `"v0.21.0"`. `kserve.kserve.controller.affinity` # { #helm.kserve.kserve.controller.affinity } : Type `object`, default `{}`. A Kubernetes Affinity, if required. For more information, see [Affinity v1 core](https://kubernetes.io/docs/reference/generated/kubernetes-api/v1.27/#affinity-v1-core). For example: affinity: nodeAffinity: requiredDuringSchedulingIgnoredDuringExecution: nodeSelectorTerms: - matchExpressions: - key: foo.bar.com/role operator: In values: - master `kserve.kserve.controller.annotations` # { #helm.kserve.kserve.controller.annotations } : Type `object`, default `{}`. Optional additional annotations to add to the controller deployment. `kserve.kserve.controller.containerSecurityContext` # { #helm.kserve.kserve.controller.containerSecurityContext } : Type `object`. Container Security Context to be set on the controller component container. For more information, see [Configure a Security Context for a Pod or Container](https://kubernetes.io/docs/tasks/configure-pod-container/security-context/). ??? note "Default" ```yaml allowPrivilegeEscalation: false capabilities: drop: - ALL privileged: false readOnlyRootFilesystem: true runAsNonRoot: true seccompProfile: type: RuntimeDefault ``` `kserve.kserve.controller.deploymentMode` # { #helm.kserve.kserve.controller.deploymentMode } : Type `string`, default `"Knative"`. KServe deployment mode: "Knative" (formerly "Serverless"), "Standard" (formerly "RawDeployment"). `kserve.kserve.controller.gateway.additionalIngressDomains` # { #helm.kserve.kserve.controller.gateway.additionalIngressDomains } : Type `list`, default `[]`. Optional additional domains for ingress routing. `kserve.kserve.controller.gateway.disableIngressCreation` # { #helm.kserve.kserve.controller.gateway.disableIngressCreation } : Type `bool`, default `false`. Whether to disable ingress creation for RawDeployment mode. `kserve.kserve.controller.gateway.disableIstioVirtualHost` # { #helm.kserve.kserve.controller.gateway.disableIstioVirtualHost } : Type `bool`, default `false`. DisableIstioVirtualHost controls whether to use istio as network layer for top level component routing or path based routing. This configuration is only applicable for Serverless mode, when disabled Istio is no longer required. `kserve.kserve.controller.gateway.domain` # { #helm.kserve.kserve.controller.gateway.domain } : Type `string`, default `""`. Ingress domain for RawDeployment mode, for Serverless it is configured in Knative. Empty falls back to `knative.domainName` so Standard-mode InferenceService hosts match Knative-mode hosts. If set explicitly it must stay aligned with `knative.domainName`: the gateway's authority rewrite and the model-serving authenticator only know that domain. `kserve.kserve.controller.gateway.domainTemplate` # { #helm.kserve.kserve.controller.gateway.domainTemplate } : Type `string`, default `"{{ .Name }}.{{ .Namespace }}.{{ .IngressDomain }}"`. Ingress domain template for RawDeployment mode, for Serverless mode it is configured in Knative. Dot-separated to match the `..` host format Knative-mode uses. `kserve.kserve.controller.gateway.ingressGateway.className` # { #helm.kserve.kserve.controller.gateway.ingressGateway.className } : Type `string`, default `"istio"`. `kserve.kserve.controller.gateway.ingressGateway.enableGatewayApi` # { #helm.kserve.kserve.controller.gateway.ingressGateway.enableGatewayApi } : Type `bool`, default `false`. Whether to use the Gateway API for ingress routing instead of Kubernetes Ingress. Kept false for the Knative + Istio network layer. `kserve.kserve.controller.gateway.ingressGateway.gateway` # { #helm.kserve.kserve.controller.gateway.ingressGateway.gateway } : Type `string`, default `"knative-ingress-gateway"`. `kserve.kserve.controller.gateway.localGateway.gateway` # { #helm.kserve.kserve.controller.gateway.localGateway.gateway } : Type `string`, default `"knative-local-gateway"`. localGateway specifies the gateway which handles the network traffic within the cluster. `kserve.kserve.controller.gateway.localGateway.gatewayService` # { #helm.kserve.kserve.controller.gateway.localGateway.gatewayService } : Type `string`, default `"knative-local-gateway"`. localGatewayService specifies the hostname of the local gateway service. `kserve.kserve.controller.gateway.localGateway.knativeGatewayService` # { #helm.kserve.kserve.controller.gateway.localGateway.knativeGatewayService } : Type `string`, default `""`. knativeLocalGatewayService specifies the hostname of the Knative's local gateway service. When unset, the value of "localGatewayService" will be used. When enabling strict mTLS in Istio, KServe local gateway should be created and pointed to the Knative local gateway. `kserve.kserve.controller.gateway.urlScheme` # { #helm.kserve.kserve.controller.gateway.urlScheme } : Type `string`, default `"http"`. HTTP endpoint url scheme. `kserve.kserve.controller.image` # { #helm.kserve.kserve.controller.image } : Type `string`, default `"kserve/kserve-controller"`. KServe controller container image name. Mirrored from upstream kserve/kserve-controller (the Hopsworks idempotency patch is upstream as of 0.19, so the controller is no longer built from the fork). `kserve.kserve.controller.knativeAddressableResolver` # { #helm.kserve.kserve.controller.knativeAddressableResolver } : Type `object`, default `{"enabled":false}`. Indicates whether to create an addressable resolver ClusterRole for Knative Eventing. This ClusterRole grants the necessary permissions for the Knative's DomainMapping reconciler to resolve InferenceService addressables. `kserve.kserve.controller.labels` # { #helm.kserve.kserve.controller.labels } : Type `object`, default `{}`. Optional additional labels to add to the controller deployment. `kserve.kserve.controller.nodeSelector` # { #helm.kserve.kserve.controller.nodeSelector } : Type `object`, default `{}`. The nodeSelector on Pods tells Kubernetes to schedule Pods on the nodes with matching labels. For more information, see [Assigning Pods to Nodes](https://kubernetes.io/docs/concepts/scheduling-eviction/assign-pod-node/). `kserve.kserve.controller.podAnnotations` # { #helm.kserve.kserve.controller.podAnnotations } : Type `object`, default `{}`. Optional additional labels to add to the controller Pods. `kserve.kserve.controller.podDisruptionBudget.enabled` # { #helm.kserve.kserve.controller.podDisruptionBudget.enabled } : Type `bool`, default `true`. `kserve.kserve.controller.podDisruptionBudget.minAvailable` # { #helm.kserve.kserve.controller.podDisruptionBudget.minAvailable } : Type `int`, default `1`. `kserve.kserve.controller.podLabels` # { #helm.kserve.kserve.controller.podLabels } : Type `object`, default `{}`. Optional additional labels to add to the controller Pods. `kserve.kserve.controller.rbacProxy.resources.limits.cpu` # { #helm.kserve.kserve.controller.rbacProxy.resources.limits.cpu } : Type `string`, default `"100m"`. `kserve.kserve.controller.rbacProxy.resources.limits.memory` # { #helm.kserve.kserve.controller.rbacProxy.resources.limits.memory } : Type `string`, default `"300Mi"`. `kserve.kserve.controller.rbacProxy.resources.requests.cpu` # { #helm.kserve.kserve.controller.rbacProxy.resources.requests.cpu } : Type `string`, default `"100m"`. `kserve.kserve.controller.rbacProxy.resources.requests.memory` # { #helm.kserve.kserve.controller.rbacProxy.resources.requests.memory } : Type `string`, default `"300Mi"`. `kserve.kserve.controller.rbacProxy.securityContext` # { #helm.kserve.kserve.controller.rbacProxy.securityContext } : Type `object`. security context configuration ??? note "Default" ```yaml allowPrivilegeEscalation: false capabilities: drop: - ALL privileged: false readOnlyRootFilesystem: true runAsNonRoot: true seccompProfile: type: RuntimeDefault ``` `kserve.kserve.controller.rbacProxyImage` # { #helm.kserve.kserve.controller.rbacProxyImage } : Type `string`, default `"quay.io/brancz/kube-rbac-proxy:v0.18.0"`. KServe controller manager rbac proxy container image `kserve.kserve.controller.replicaCount` # { #helm.kserve.kserve.controller.replicaCount } : Type `int`, default `1`. `kserve.kserve.controller.resources` # { #helm.kserve.kserve.controller.resources } : Type `object`. Resources to provide to the kserve controller pod. For example: requests: cpu: 10m memory: 32Mi For more information, see [Resource Management for Pods and Containers](https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/). ??? note "Default" ```yaml limits: cpu: 100m memory: 300Mi requests: cpu: 100m memory: 300Mi ``` `kserve.kserve.controller.securityContext` # { #helm.kserve.kserve.controller.securityContext } : Type `object`, default `{"runAsNonRoot":true,"seccompProfile":{"type":"RuntimeDefault"}}`. Pod Security Context. For more information, see [Configure a Security Context for a Pod or Container](https://kubernetes.io/docs/tasks/configure-pod-container/security-context/). `kserve.kserve.controller.serviceAccount.annotations` # { #helm.kserve.kserve.controller.serviceAccount.annotations } : Type `object`, default `{}`. service account annotations `kserve.kserve.controller.tag` # { #helm.kserve.kserve.controller.tag } : Type `string`, default `"v0.21.0"`. KServe controller container image tag. `kserve.kserve.controller.tolerations` # { #helm.kserve.kserve.controller.tolerations } : Type `list`, default `[]`. A list of Kubernetes Tolerations, if required. For more information, see [Toleration v1 core](https://kubernetes.io/docs/reference/generated/kubernetes-api/v1.27/#toleration-v1-core). For example: tolerations: - key: foo.bar.com/role operator: Equal value: master effect: NoSchedule `kserve.kserve.controller.topologySpreadConstraint` # { #helm.kserve.kserve.controller.topologySpreadConstraint } : Type `object`, default `{}`. The default topology spread constraint. If not defined the global topology spread constraint would be used instead only if topologySpreadConstraints list is empty. `kserve.kserve.controller.topologySpreadConstraints` # { #helm.kserve.kserve.controller.topologySpreadConstraints } : Type `list`, default `[]`. A list of Kubernetes TopologySpreadConstraints, if required. For more information, see \[Topology spread constraint v1 core\]( For example: topologySpreadConstraints: - maxSkew: 2 topologyKey: topology.kubernetes.io/zone whenUnsatisfiable: ScheduleAnyway labelSelector: matchLabels: app.kubernetes.io/instance: kserve-controller-manager app.kubernetes.io/component: controller `kserve.kserve.localmodel.controller.image` # { #helm.kserve.kserve.localmodel.controller.image } : Type `string`, default `"kserve/kserve-localmodel-controller"`. `kserve.kserve.localmodel.controller.tag` # { #helm.kserve.kserve.localmodel.controller.tag } : Type `string`, default `"v0.21.0"`. `kserve.kserve.localmodel.enabled` # { #helm.kserve.kserve.localmodel.enabled } : Type `bool`, default `false`. Feeds the `localModel` block of `inferenceservice-config` only and must stay false: the chart ships no localmodel controller (the dormant copy vendored with 0.19 was removed with 0.21; the feature arrives with the upstream kserve-localmodel charts). Rendering fails on `true` so an old values file cannot enable the feature with no controller behind it. `controller` and `serviceAccount` are kept for values compatibility and render nothing. `kserve.kserve.localmodel.jobNamespace` # { #helm.kserve.kserve.localmodel.jobNamespace } : Type `string`, default `"kserve-localmodel-jobs"`. `kserve.kserve.localmodel.securityContext` # { #helm.kserve.kserve.localmodel.securityContext } : Type `object`, default `{"FSGroup":1000}`. security context configuration `kserve.kserve.localmodel.serviceAccount.annotations` # { #helm.kserve.kserve.localmodel.serviceAccount.annotations } : Type `object`, default `{}`. service account annotations `kserve.kserve.metricsaggregator.enableMetricAggregation` # { #helm.kserve.kserve.metricsaggregator.enableMetricAggregation } : Type `string`, default `"true"`. configures metric aggregation annotation. This adds the annotation serving.kserve.io/enable-metric-aggregation to every service with the specified boolean value. If true enables metric aggregation in queue-proxy by setting env vars in the queue proxy container to configure scraping ports. `kserve.kserve.metricsaggregator.enablePrometheusScraping` # { #helm.kserve.kserve.metricsaggregator.enablePrometheusScraping } : Type `string`, default `"true"`. If true, prometheus annotations are added to the pod to scrape the metrics. If serving.kserve.io/enable-metric-aggregation is false, the prometheus port is set with the default prometheus scraping port 9090, otherwise the prometheus port annotation is set with the metric aggregation port. `kserve.kserve.nodeSelector` # { #helm.kserve.kserve.nodeSelector } : Type `object`, default `{}`. node selector configuration `kserve.kserve.router.image` # { #helm.kserve.kserve.router.image } : Type `string`, default `"kserve/router"`. `kserve.kserve.router.tag` # { #helm.kserve.kserve.router.tag } : Type `string`, default `"v0.21.0"`. `kserve.kserve.serviceAccount.annotations` # { #helm.kserve.kserve.serviceAccount.annotations } : Type `object`, default `{}`. service account annotations `kserve.kserve.servingruntime.modelNamePlaceholder` # { #helm.kserve.kserve.servingruntime.modelNamePlaceholder } : Type `string`, default `"{{.Name}}"`. `kserve.kserve.servingruntime.sklearnserver.image` # { #helm.kserve.kserve.servingruntime.sklearnserver.image } : Type `string`, default `"hopsworks/sklearnserver"`. `kserve.kserve.servingruntime.sklearnserver.tag` # { #helm.kserve.kserve.servingruntime.sklearnserver.tag } : Type `string`, default `"0.21.0"`. `kserve.kserve.servingruntime.tensorflow.image` # { #helm.kserve.kserve.servingruntime.tensorflow.image } : Type `string`, default `"tensorflow/serving"`. `kserve.kserve.servingruntime.tensorflow.tag` # { #helm.kserve.kserve.servingruntime.tensorflow.tag } : Type `string`, default `"2.20.0"`. `kserve.kserve.servingruntime.vllmomni` # { #helm.kserve.kserve.servingruntime.vllmomni } : Type `object`, default `{"env":[],"image":"vllm/vllm-omni","imageRegistry":"","tag":"v0.28.0"}`. backs the `vllm-omni` ClusterServingRuntime (vLLM-Omni) `kserve.kserve.servingruntime.vllmomni.env` # { #helm.kserve.kserve.servingruntime.vllmomni.env } : Type `list`, default `[]`. Additional environment variables for vllm-omni container `kserve.kserve.servingruntime.vllmomni.imageRegistry` # { #helm.kserve.kserve.servingruntime.vllmomni.imageRegistry } : Type `string`, default `""`. Registry override for the vllm-omni image. Empty falls back to `global._hopsworks.imageRegistry`. `kserve.kserve.servingruntime.vllmopenai` # { #helm.kserve.kserve.servingruntime.vllmopenai } : Type `object`, default `{"env":[],"image":"vllm/vllm-openai","imageRegistry":"","tag":"v0.28.0"}`. backs the `vllm-openai` ClusterServingRuntime (standard vLLM) `kserve.kserve.servingruntime.vllmopenai.env` # { #helm.kserve.kserve.servingruntime.vllmopenai.env } : Type `list`, default `[]`. Additional environment variables for vllm-openai container `kserve.kserve.servingruntime.vllmopenai.imageRegistry` # { #helm.kserve.kserve.servingruntime.vllmopenai.imageRegistry } : Type `string`, default `""`. Registry override for the vllm-openai image. Empty falls back to `global._hopsworks.imageRegistry`. `kserve.kserve.storage.caBundleConfigMapName` # { #helm.kserve.kserve.storage.caBundleConfigMapName } : Type `string`, default `""`. Mounted CA bundle config map name for storage initializer. `kserve.kserve.storage.caBundleVolumeMountPath` # { #helm.kserve.kserve.storage.caBundleVolumeMountPath } : Type `string`, default `"/etc/ssl/custom-certs"`. Mounted path for CA bundle config map. `kserve.kserve.storage.cpuModelcar` # { #helm.kserve.kserve.storage.cpuModelcar } : Type `string`, default `"10m"`. Model sidecar cpu requirement. `kserve.kserve.storage.enableModelcar` # { #helm.kserve.kserve.storage.enableModelcar } : Type `bool`, default `false`. Flag for enabling model sidecar feature. `kserve.kserve.storage.image` # { #helm.kserve.kserve.storage.image } : Type `string`, default `"hopsworks/storage-initializer"`. `kserve.kserve.storage.memoryModelcar` # { #helm.kserve.kserve.storage.memoryModelcar } : Type `string`, default `"15Mi"`. Model sidecar memory requirement. `kserve.kserve.storage.resources` # { #helm.kserve.kserve.storage.resources } : Type `object`. Requests and limits KServe gives the storage-initializer init container it injects for storageUri models (the `storageInitializer` block of `inferenceservice-config`). ??? note "Default" ```yaml limits: cpu: '1' memory: 1Gi requests: cpu: 100m memory: 100Mi ``` `kserve.kserve.storage.resources.limits` # { #helm.kserve.kserve.storage.resources.limits } : Type `object`, default `{"cpu":"1","memory":"1Gi"}`. cpu and memory are read by the templates; other resource names pass through to the ClusterStorageContainer. `kserve.kserve.storage.resources.requests` # { #helm.kserve.kserve.storage.resources.requests } : Type `object`, default `{"cpu":"100m","memory":"100Mi"}`. cpu and memory are read by the templates; other resource names pass through to the ClusterStorageContainer. `kserve.kserve.storage.s3` # { #helm.kserve.kserve.storage.s3 } : Type `object`. Configurations for S3 storage ??? note "Default" ```yaml CABundle: '' accessKeyIdName: AWS_ACCESS_KEY_ID endpoint: '' region: '' secretAccessKeyName: AWS_SECRET_ACCESS_KEY useAnonymousCredential: '' useHttps: '' useVirtualBucket: '' verifySSL: '' ``` `kserve.kserve.storage.s3.CABundle` # { #helm.kserve.kserve.storage.s3.CABundle } : Type `string`, default `""`. The path to the certificate bundle to use for HTTPS certificate validation. `kserve.kserve.storage.s3.accessKeyIdName` # { #helm.kserve.kserve.storage.s3.accessKeyIdName } : Type `string`, default `"AWS_ACCESS_KEY_ID"`. AWS S3 static access key id. `kserve.kserve.storage.s3.endpoint` # { #helm.kserve.kserve.storage.s3.endpoint } : Type `string`, default `""`. AWS S3 endpoint. `kserve.kserve.storage.s3.region` # { #helm.kserve.kserve.storage.s3.region } : Type `string`, default `""`. Default region name of AWS S3. `kserve.kserve.storage.s3.secretAccessKeyName` # { #helm.kserve.kserve.storage.s3.secretAccessKeyName } : Type `string`, default `"AWS_SECRET_ACCESS_KEY"`. AWS S3 static secret access key. `kserve.kserve.storage.s3.useAnonymousCredential` # { #helm.kserve.kserve.storage.s3.useAnonymousCredential } : Type `string`, default `""`. Whether to use anonymous credentials to download the model or not, default to false. `kserve.kserve.storage.s3.useHttps` # { #helm.kserve.kserve.storage.s3.useHttps } : Type `string`, default `""`. Whether to use secured https or http to download models, allowed values are 0 and 1 and default to 1. `kserve.kserve.storage.s3.useVirtualBucket` # { #helm.kserve.kserve.storage.s3.useVirtualBucket } : Type `string`, default `""`. Whether to use virtual bucket or not, default to false. `kserve.kserve.storage.s3.verifySSL` # { #helm.kserve.kserve.storage.s3.verifySSL } : Type `string`, default `""`. Whether to verify the tls/ssl certificate, default to true. `kserve.kserve.storage.storageSecretNameAnnotation` # { #helm.kserve.kserve.storage.storageSecretNameAnnotation } : Type `string`, default `"serving.kserve.io/secretName"`. Storage secret name reference for storage initializer. `kserve.kserve.storage.storageSpecSecretName` # { #helm.kserve.kserve.storage.storageSpecSecretName } : Type `string`, default `"storage-config"`. Storage spec secret name. `kserve.kserve.storage.tag` # { #helm.kserve.kserve.storage.tag } : Type `string`, default `"5.2.0-SNAPSHOT"`. Hopsworks-built storage initializer tag. Pinned to the Hopsworks version (matching storageInitializer.tag and what the model-serving-webhook publishes), not the KServe version anchor. `kserve.kserve.storage.uidModelcar` # { #helm.kserve.kserve.storage.uidModelcar } : Type `int`, default `1010`. UID under which the modelcar process and the main container run. Some clusters require root (0). `kserve.kserve.tolerations` # { #helm.kserve.kserve.tolerations } : Type `list`, default `[]`. `kserve.kserve.topologySpreadConstraint` # { #helm.kserve.kserve.topologySpreadConstraint } : Type `object`, default `{}`. The default topology spread constraint. If not defined the global topology spread constraint would be used instead only if topologySpreadConstraints list is empty. `kserve.kserve.version` # { #helm.kserve.kserve.version } : Type `string`, default `"v0.21.0"`.
## servingAuthenticator { #helm-values-kserve-servingauthenticator } ??? example "Defaults as YAML" ```yaml kserve: servingAuthenticator: image: hopsworks/model-serving-authenticator name: model-serving-authenticator podDisruptionBudget: enabled: true minAvailable: 1 replicaCount: 1 requireExternalUserLoginAfterHours: 720 serviceAccount: annotations: {} tag: 5.2.0-SNAPSHOT ```
`kserve.servingAuthenticator.image` # { #helm.kserve.servingAuthenticator.image } : Type `string`, default `"hopsworks/model-serving-authenticator"`. `kserve.servingAuthenticator.name` # { #helm.kserve.servingAuthenticator.name } : Type `string`, default `"model-serving-authenticator"`. `kserve.servingAuthenticator.podDisruptionBudget.enabled` # { #helm.kserve.servingAuthenticator.podDisruptionBudget.enabled } : Type `bool`, default `true`. `kserve.servingAuthenticator.podDisruptionBudget.minAvailable` # { #helm.kserve.servingAuthenticator.podDisruptionBudget.minAvailable } : Type `int`, default `1`. `kserve.servingAuthenticator.replicaCount` # { #helm.kserve.servingAuthenticator.replicaCount } : Type `int`, default `1`. `kserve.servingAuthenticator.requireExternalUserLoginAfterHours` # { #helm.kserve.servingAuthenticator.requireExternalUserLoginAfterHours } : Type `int`, default `720`. Number of hours after which external users are required to sign in to Hopsworks to refresh their external user groups. Allowed values are -1, 0 and greater than 0, where -1 skips the periodic sign-in requirement and 0 disables external access completely. `kserve.servingAuthenticator.serviceAccount.annotations` # { #helm.kserve.servingAuthenticator.serviceAccount.annotations } : Type `object`, default `{}`. service account annotations `kserve.servingAuthenticator.tag` # { #helm.kserve.servingAuthenticator.tag } : Type `string`, default `"5.2.0-SNAPSHOT"`.
## servingWebhook { #helm-values-kserve-servingwebhook } ??? example "Defaults as YAML" ```yaml kserve: servingWebhook: ephemeralVolume: accessModes: - ReadWriteOnce enabled: false storageClassName: '' storageSize: 60Gi image: hopsworks/model-serving-webhook name: model-serving-webhook podDisruptionBudget: enabled: true minAvailable: 1 replicaCount: 1 serviceAccount: annotations: {} tag: 5.2.0-SNAPSHOT ttlSecondsAfterFinished: null ```
`kserve.servingWebhook.ephemeralVolume` # { #helm.kserve.servingWebhook.ephemeralVolume } : Type `object`. Configuration for ephemeral volumes in LLM deployments ??? note "Default" ```yaml accessModes: - ReadWriteOnce enabled: false storageClassName: '' storageSize: 60Gi ``` `kserve.servingWebhook.ephemeralVolume.accessModes` # { #helm.kserve.servingWebhook.ephemeralVolume.accessModes } : Type `list`, default `["ReadWriteOnce"]`. Access modes for the ephemeral volume `kserve.servingWebhook.ephemeralVolume.enabled` # { #helm.kserve.servingWebhook.ephemeralVolume.enabled } : Type `bool`, default `false`. Enable ephemeral volume injection for LLM deployments `kserve.servingWebhook.ephemeralVolume.storageClassName` # { #helm.kserve.servingWebhook.ephemeralVolume.storageClassName } : Type `string`, default `""`. Storage class name for the ephemeral volume `kserve.servingWebhook.ephemeralVolume.storageSize` # { #helm.kserve.servingWebhook.ephemeralVolume.storageSize } : Type `string`, default `"60Gi"`. Storage size for the ephemeral volume `kserve.servingWebhook.image` # { #helm.kserve.servingWebhook.image } : Type `string`, default `"hopsworks/model-serving-webhook"`. `kserve.servingWebhook.name` # { #helm.kserve.servingWebhook.name } : Type `string`, default `"model-serving-webhook"`. `kserve.servingWebhook.podDisruptionBudget.enabled` # { #helm.kserve.servingWebhook.podDisruptionBudget.enabled } : Type `bool`, default `true`. `kserve.servingWebhook.podDisruptionBudget.minAvailable` # { #helm.kserve.servingWebhook.podDisruptionBudget.minAvailable } : Type `int`, default `1`. `kserve.servingWebhook.replicaCount` # { #helm.kserve.servingWebhook.replicaCount } : Type `int`, default `1`. `kserve.servingWebhook.serviceAccount.annotations` # { #helm.kserve.servingWebhook.serviceAccount.annotations } : Type `object`, default `{}`. service account annotations `kserve.servingWebhook.tag` # { #helm.kserve.servingWebhook.tag } : Type `string`, default `"5.2.0-SNAPSHOT"`. `kserve.servingWebhook.ttlSecondsAfterFinished` # { #helm.kserve.servingWebhook.ttlSecondsAfterFinished } : Type `string`, default `nil`. TTL in seconds for the serving-mwh Job. Overrides global default.
================================================================================ # minio Source: https://docs.hopsworks.ai/latest/setup_installation/common/helm_chart_values/minio/ # MinIO values { #helm-values-minio } Values under `minio` configure MinIO, an S3-compatible object store deployed inside the cluster. _Generated from the Hopsworks Helm chart `5.2.0-alpha-1791549041` (Hopsworks `5.2.0`)._ Deployed when [`global._hopsworks.minio.enabled`](global.md#helm.global._hopsworks.minio.enabled) is `true`. ## General { #helm-values-minio-general } ??? example "Defaults as YAML" ```yaml minio: add_node_port: false affinity: {} buckets: - name: hopsworks cleanupOnUninstall: enabled: true ttlSecondsAfterFinished: null deploymentName: minio enabled: true hopsworkslib: {} image: name: minio pullPolicy: IfNotPresent registry: docker.hops.works tag: RELEASE.2024-03-26T22-10-45Z-cpuv1 nodeSelector: {} podDisruptionBudget: enabled: true minAvailable: 1 protocol: http publicBuckets: - name: public - name: buildkit replicas: 2 resources: limits: cpu: '4' memory: 8Gi requests: cpu: '1' memory: 4Gi storage: 100Gi storageClassName: null tests: minioBackup: bucket: '' enabled: false subfolder: '' tolerations: [] ```
`minio` # { #helm.minio } : Type `object`. override minio values ??? note "Default" ```yaml replicas: 2 resources: limits: cpu: '4' memory: 8Gi storage: 100Gi ``` `minio.add_node_port` # { #helm.minio.add_node_port } : Type `bool`, default `false`. `minio.affinity` # { #helm.minio.affinity } : Type `object`, default `{}`. affinity configuration `minio.buckets[0].name` # { #helm.minio.buckets.0.name } : Type `string`, default `"hopsworks"`. `minio.cleanupOnUninstall` # { #helm.minio.cleanupOnUninstall } : Type `object`, default `{"enabled":true,"ttlSecondsAfterFinished":null}`. post-delete cleanup of MinIO leftovers. The MinIO StatefulSet PVC survives uninstall; this deletes it by label (app=minio), but only when global._hopsworks.wipeDataOnUninstall is enabled and never for PVCs labelled hopsworks.ai/keep=true. `minio.cleanupOnUninstall.enabled` # { #helm.minio.cleanupOnUninstall.enabled } : Type `bool`, default `true`. enable the post-delete MinIO data-PVC cleanup hook (also requires global._hopsworks.wipeDataOnUninstall) `minio.cleanupOnUninstall.ttlSecondsAfterFinished` # { #helm.minio.cleanupOnUninstall.ttlSecondsAfterFinished } : Type `string`, default `nil`. ttlSecondsAfterFinished for the cleanup Job; null falls through to the global default `minio.deploymentName` # { #helm.minio.deploymentName } : Type `string`, default `"minio"`. `minio.enabled` # { #helm.minio.enabled } : Type `bool`, default `true`. `minio.hopsworkslib` # { #helm.minio.hopsworkslib } : Type `object`, default `{}`. override hopsworkslib values `minio.image.name` # { #helm.minio.image.name } : Type `string`, default `"minio"`. `minio.image.pullPolicy` # { #helm.minio.image.pullPolicy } : Type `string`, default `"IfNotPresent"`. `minio.image.registry` # { #helm.minio.image.registry } : Type `string`, default `"docker.hops.works"`. `minio.image.tag` # { #helm.minio.image.tag } : Type `string`, default `"RELEASE.2024-03-26T22-10-45Z-cpuv1"`. `minio.nodeSelector` # { #helm.minio.nodeSelector } : Type `object`, default `{}`. node selector configuration `minio.podDisruptionBudget.enabled` # { #helm.minio.podDisruptionBudget.enabled } : Type `bool`, default `true`. `minio.podDisruptionBudget.minAvailable` # { #helm.minio.podDisruptionBudget.minAvailable } : Type `int`, default `1`. `minio.protocol` # { #helm.minio.protocol } : Type `string`, default `"http"`. `minio.publicBuckets[0].name` # { #helm.minio.publicBuckets.0.name } : Type `string`, default `"public"`. `minio.publicBuckets[1].name` # { #helm.minio.publicBuckets.1.name } : Type `string`, default `"buildkit"`. `minio.replicas` # { #helm.minio.replicas } : Type `int`, default `2`. `minio.resources.limits.cpu` # { #helm.minio.resources.limits.cpu } : Type `string`, default `"4"`. `minio.resources.limits.memory` # { #helm.minio.resources.limits.memory } : Type `string`, default `"8Gi"`. `minio.resources.requests.cpu` # { #helm.minio.resources.requests.cpu } : Type `string`, default `"1"`. `minio.resources.requests.memory` # { #helm.minio.resources.requests.memory } : Type `string`, default `"4Gi"`. `minio.storage` # { #helm.minio.storage } : Type `string`, default `"100Gi"`. `minio.storageClassName` # { #helm.minio.storageClassName } : Type `string`, default `nil`. storage class name `minio.tests.minioBackup.bucket` # { #helm.minio.tests.minioBackup.bucket } : Type `string`, default `""`. `minio.tests.minioBackup.enabled` # { #helm.minio.tests.minioBackup.enabled } : Type `bool`, default `false`. `minio.tests.minioBackup.subfolder` # { #helm.minio.tests.minioBackup.subfolder } : Type `string`, default `""`. `minio.tolerations` # { #helm.minio.tolerations } : Type `list`, default `[]`.
## deployment { #helm-values-minio-deployment } ??? example "Defaults as YAML" ```yaml minio: deployment: env: MINIO_REGION: eu-west-1 MINIO_ROOT_PASSWORD: minioadmin MINIO_ROOT_USER: minioadmin extra_envs: MINIO_PROMETHEUS_AUTH_TYPE: public ports: console: 9001 http: 9000 ```
`minio.deployment.env.MINIO_REGION` # { #helm.minio.deployment.env.MINIO_REGION } : Type `string`, default `"eu-west-1"`. `minio.deployment.env.MINIO_ROOT_PASSWORD` # { #helm.minio.deployment.env.MINIO_ROOT_PASSWORD } : Type `string`, default `"minioadmin"`. `minio.deployment.env.MINIO_ROOT_USER` # { #helm.minio.deployment.env.MINIO_ROOT_USER } : Type `string`, default `"minioadmin"`. `minio.deployment.extra_envs` # { #helm.minio.deployment.extra_envs } : Type `object`, default `{"MINIO_PROMETHEUS_AUTH_TYPE":"public"}`. extra envs `minio.deployment.ports.console` # { #helm.minio.deployment.ports.console } : Type `int`, default `9001`. `minio.deployment.ports.http` # { #helm.minio.deployment.ports.http } : Type `int`, default `9000`.
## service { #helm-values-minio-service } ??? example "Defaults as YAML" ```yaml minio: service: annotations: consul.hashicorp.com/service-name: minio consul.hashicorp.com/service-port: http name: minio ports: console: 9001 http: 9000 ```
`minio.service.annotations."consul.hashicorp.com/service-name"` # { #helm.minio.service.annotations.consul.hashicorp.com-service-name } : Type `string`, default `"minio"`. `minio.service.annotations."consul.hashicorp.com/service-port"` # { #helm.minio.service.annotations.consul.hashicorp.com-service-port } : Type `string`, default `"http"`. `minio.service.name` # { #helm.minio.service.name } : Type `string`, default `"minio"`. `minio.service.ports.console` # { #helm.minio.service.ports.console } : Type `int`, default `9001`. `minio.service.ports.http` # { #helm.minio.service.ports.http } : Type `int`, default `9000`.
================================================================================ # olk Source: https://docs.hopsworks.ai/latest/setup_installation/common/helm_chart_values/olk/ # OpenSearch values { #helm-values-olk } Values under `olk` configure OpenSearch, OpenSearch Dashboards, Logstash and Filebeat, which provide search, the vector index and service logs. _Generated from the Hopsworks Helm chart `5.2.0-alpha-1791549041` (Hopsworks `5.2.0`)._ Deployed when [`global._hopsworks.opensearch.enabled`](global.md#helm.global._hopsworks.opensearch.enabled) is `true`. !!! info "Upstream charts" - Values under `olk.prometheus-elasticsearch-exporter` go to [`prometheus-elasticsearch-exporter` 5.8.0](https://artifacthub.io/packages/helm/prometheus-community/prometheus-elasticsearch-exporter/5.8.0) from `https://prometheus-community.github.io/helm-charts`. Only the values Hopsworks sets under `olk.prometheus-elasticsearch-exporter` are listed on this page. Any other value of the chart can be set there too; the link opens its documentation for the version Hopsworks pins. ## General { #helm-values-olk-general } ??? example "Defaults as YAML" ```yaml olk: appName: elk cleanupOnUninstall: enabled: true ttlSecondsAfterFinished: null hopsworkslib: {} image: registry: docker.hops.works opensearch: storageClassName: null prometheus-elasticsearch-exporter: es: sslSkipVerify: true uri: https://elastic_exporter:elastic_exporterpw@{{ include "olk.exporter.opensearchHost" . }}:9200 image: registry: docker.hops.works repository: prometheus/elasticsearch-exporter tag: 1.11.0-alpine-h1.1 nodeSelector: {} resources: limits: cpu: 800m memory: 200Mi requests: cpu: 300m memory: 128Mi service: annotations: prometheus.io/path: /metrics prometheus.io/port: '9108' prometheus.io/scheme: http prometheus.io/scrape: 'true' httpPort: 9108 tolerations: [] ```
`olk` # { #helm.olk } : Type `object`, default `{"opensearch":{"storageClassName":null}}`. override olk values `olk.appName` # { #helm.olk.appName } : Type `string`, default `"elk"`. `olk.cleanupOnUninstall` # { #helm.olk.cleanupOnUninstall } : Type `object`, default `{"enabled":true,"ttlSecondsAfterFinished":null}`. post-delete cleanup of OLK leftovers. Always deletes the signingjwt Secret (created centrally, or copied in satellite; not Helm-tracked). Also deletes the OpenSearch StatefulSet PVC by label (app=opensearch) when global._hopsworks.wipeDataOnUninstall is enabled, except PVCs labelled hopsworks.ai/keep=true. `olk.cleanupOnUninstall.enabled` # { #helm.olk.cleanupOnUninstall.enabled } : Type `bool`, default `true`. enable the post-delete OLK cleanup hook (signingjwt Secret always; OpenSearch PVC additionally requires global._hopsworks.wipeDataOnUninstall) `olk.cleanupOnUninstall.ttlSecondsAfterFinished` # { #helm.olk.cleanupOnUninstall.ttlSecondsAfterFinished } : Type `string`, default `nil`. ttlSecondsAfterFinished for the cleanup Job; null falls through to the global default `olk.hopsworkslib` # { #helm.olk.hopsworkslib } : Type `object`, default `{}`. override hopsworkslib values `olk.image.registry` # { #helm.olk.image.registry } : Type `string`, default `"docker.hops.works"`. `olk.prometheus-elasticsearch-exporter` # { #helm.olk.prometheus-elasticsearch-exporter } : Type `object`, passed to the [`prometheus-elasticsearch-exporter` 5.8.0](https://artifacthub.io/packages/helm/prometheus-community/prometheus-elasticsearch-exporter/5.8.0) chart, whose other values are documented there. override prometheus elasticsearch exporter values ??? note "Default" ```yaml es: sslSkipVerify: true uri: https://elastic_exporter:elastic_exporterpw@{{ include "olk.exporter.opensearchHost" . }}:9200 image: registry: docker.hops.works repository: prometheus/elasticsearch-exporter tag: 1.11.0-alpine-h1.1 nodeSelector: {} resources: limits: cpu: 800m memory: 200Mi requests: cpu: 300m memory: 128Mi service: annotations: prometheus.io/path: /metrics prometheus.io/port: '9108' prometheus.io/scheme: http prometheus.io/scrape: 'true' httpPort: 9108 tolerations: [] ```
## dashboard { #helm-values-olk-dashboard } ??? example "Defaults as YAML" ```yaml olk: dashboard: config: basePath: /hopsworks-api/kibana index: .kibana name: dashboard-config deleteKibanaIndexIfPrevious: false export: false jwt: roles_key: roles subject_key: sub url_parameter: jt name: opensearch-dashboard nodeSelector: {} port: 5601 resources: limits: cpu: 200m memory: 1Gi requests: cpu: 50m memory: 512Mi security: cookie_ttl: 1800000 session_ttl: 3600000 security_context: fsGroup: 1000 runAsGroup: 1000 runAsNonRoot: true runAsUser: 1000 service: annotations: consul.hashicorp.com/service-name: kibana shard_timeout: 10000 startup_timeout: 5000 tolerations: [] topologySpreadConstraint: {} version: 3.8.0.0 ```
`olk.dashboard.config.basePath` # { #helm.olk.dashboard.config.basePath } : Type `string`, default `"/hopsworks-api/kibana"`. `olk.dashboard.config.index` # { #helm.olk.dashboard.config.index } : Type `string`, default `".kibana"`. `olk.dashboard.config.name` # { #helm.olk.dashboard.config.name } : Type `string`, default `"dashboard-config"`. `olk.dashboard.deleteKibanaIndexIfPrevious` # { #helm.olk.dashboard.deleteKibanaIndexIfPrevious } : Type `bool`, default `false`. `olk.dashboard.export` # { #helm.olk.dashboard.export } : Type `bool`, default `false`. `olk.dashboard.jwt.roles_key` # { #helm.olk.dashboard.jwt.roles_key } : Type `string`, default `"roles"`. `olk.dashboard.jwt.subject_key` # { #helm.olk.dashboard.jwt.subject_key } : Type `string`, default `"sub"`. `olk.dashboard.jwt.url_parameter` # { #helm.olk.dashboard.jwt.url_parameter } : Type `string`, default `"jt"`. `olk.dashboard.name` # { #helm.olk.dashboard.name } : Type `string`, default `"opensearch-dashboard"`. `olk.dashboard.nodeSelector` # { #helm.olk.dashboard.nodeSelector } : Type `object`, default `{}`. node selector configuration `olk.dashboard.port` # { #helm.olk.dashboard.port } : Type `int`, default `5601`. `olk.dashboard.resources.limits.cpu` # { #helm.olk.dashboard.resources.limits.cpu } : Type `string`, default `"200m"`. `olk.dashboard.resources.limits.memory` # { #helm.olk.dashboard.resources.limits.memory } : Type `string`, default `"1Gi"`. `olk.dashboard.resources.requests.cpu` # { #helm.olk.dashboard.resources.requests.cpu } : Type `string`, default `"50m"`. `olk.dashboard.resources.requests.memory` # { #helm.olk.dashboard.resources.requests.memory } : Type `string`, default `"512Mi"`. `olk.dashboard.security.cookie_ttl` # { #helm.olk.dashboard.security.cookie_ttl } : Type `int`, default `1800000`. `olk.dashboard.security.session_ttl` # { #helm.olk.dashboard.security.session_ttl } : Type `int`, default `3600000`. `olk.dashboard.security_context.fsGroup` # { #helm.olk.dashboard.security_context.fsGroup } : Type `int`, default `1000`. `olk.dashboard.security_context.runAsGroup` # { #helm.olk.dashboard.security_context.runAsGroup } : Type `int`, default `1000`. `olk.dashboard.security_context.runAsNonRoot` # { #helm.olk.dashboard.security_context.runAsNonRoot } : Type `bool`, default `true`. `olk.dashboard.security_context.runAsUser` # { #helm.olk.dashboard.security_context.runAsUser } : Type `int`, default `1000`. `olk.dashboard.service.annotations."consul.hashicorp.com/service-name"` # { #helm.olk.dashboard.service.annotations.consul.hashicorp.com-service-name } : Type `string`, default `"kibana"`. `olk.dashboard.shard_timeout` # { #helm.olk.dashboard.shard_timeout } : Type `int`, default `10000`. `olk.dashboard.startup_timeout` # { #helm.olk.dashboard.startup_timeout } : Type `int`, default `5000`. `olk.dashboard.tolerations` # { #helm.olk.dashboard.tolerations } : Type `list`, default `[]`. `olk.dashboard.topologySpreadConstraint` # { #helm.olk.dashboard.topologySpreadConstraint } : Type `object`, default `{}`. The default topology spread constraint. If not defined the global topology spread constraint would be used instead. `olk.dashboard.version` # { #helm.olk.dashboard.version } : Type `string`, default `"3.8.0.0"`.
## filebeat { #helm-values-olk-filebeat } ??? example "Defaults as YAML" ```yaml olk: filebeat: config: name: filebeat-config enabled: true extraNamespaces: [] fullnameOverride: null image: tag: 8.19.21 kubernetesMetadata: cronjob: null deployment: null extraNamespaceLabels: [] extraNodeLabels: [] extraPodLabels: [] logFilteringNamespaceLabels: - hopsworks.ai/project - hopsworks.ai/onlinefs-cluster logFilteringNodeLabels: - kubernetes.io/hostname logFilteringPodLabels: - name - app - app.kubernetes.io/name - service - rondbService - component - user - job-type - job-id - job-name - execution - jupyter - jupyter-id - jupyter-settings-id - kernel-id - spark-role - spark-app-selector - sparkoperator.k8s.io/launched-by-spark-operator - serving.hops.works/id - serving.hops.works/name - serving.hops.works/tool - serving.hops.works/model-name - serving.hops.works/model-version - serving.hops.works/model-server - serving.hops.works/project-id namespaceAnnotations: [] nodeAnnotations: [] podAnnotations: [] logs_locations: - addKubernetesMetadata: true glob: /*.log logtype: log mountPaths: - /var/log/containers - /var/log/pods name: containerd path: /var/log/containers processors: [] name: filebeat resources: limits: memory: 400Mi requests: cpu: 100m memory: 200Mi serviceAccount: annotations: {} terminationgraceperiod: 30 tolerations: - effect: NoSchedule operator: Exists ```
`olk.filebeat.config.name` # { #helm.olk.filebeat.config.name } : Type `string`, default `"filebeat-config"`. `olk.filebeat.enabled` # { #helm.olk.filebeat.enabled } : Type `bool`, default `true`. `olk.filebeat.extraNamespaces` # { #helm.olk.filebeat.extraNamespaces } : Type `list`, default `[]`. extra namespaces to process their container logs. By default the release name space and the Hopsworks project namespaces are whitelisted. `olk.filebeat.fullnameOverride` # { #helm.olk.filebeat.fullnameOverride } : Type `string`, default `nil`. full name override `olk.filebeat.image.tag` # { #helm.olk.filebeat.image.tag } : Type `string`, default `"8.19.21"`. `olk.filebeat.kubernetesMetadata.cronjob` # { #helm.olk.filebeat.kubernetesMetadata.cronjob } : Type `bool|null`, default `nil`. attach kubernetes.cronjob.name; unset leaves filebeat's default `olk.filebeat.kubernetesMetadata.deployment` # { #helm.olk.filebeat.kubernetesMetadata.deployment } : Type `bool|null`, default `nil`. attach kubernetes.deployment.name; unset leaves filebeat's default `olk.filebeat.kubernetesMetadata.extraNamespaceLabels` # { #helm.olk.filebeat.kubernetesMetadata.extraNamespaceLabels } : Type `list`, default `[]`. extra namespace labels to keep, appended to logFilteringNamespaceLabels `olk.filebeat.kubernetesMetadata.extraNodeLabels` # { #helm.olk.filebeat.kubernetesMetadata.extraNodeLabels } : Type `list`, default `[]`. extra node labels to keep, appended to logFilteringNodeLabels `olk.filebeat.kubernetesMetadata.extraPodLabels` # { #helm.olk.filebeat.kubernetesMetadata.extraPodLabels } : Type `list`, default `[]`. extra pod labels to keep, appended to logFilteringPodLabels `olk.filebeat.kubernetesMetadata.logFilteringNamespaceLabels` # { #helm.olk.filebeat.kubernetesMetadata.logFilteringNamespaceLabels } : Type `list`, default `["hopsworks.ai/project","hopsworks.ai/onlinefs-cluster"]`. namespace labels Hopsworks' log filtering reads `olk.filebeat.kubernetesMetadata.logFilteringNodeLabels` # { #helm.olk.filebeat.kubernetesMetadata.logFilteringNodeLabels } : Type `list`, default `["kubernetes.io/hostname"]`. node labels Hopsworks' log filtering reads `olk.filebeat.kubernetesMetadata.logFilteringPodLabels` # { #helm.olk.filebeat.kubernetesMetadata.logFilteringPodLabels } : Type `list`. pod labels Hopsworks' log filtering reads ??? note "Default" ```yaml - name - app - app.kubernetes.io/name - service - rondbService - component - user - job-type - job-id - job-name - execution - jupyter - jupyter-id - jupyter-settings-id - kernel-id - spark-role - spark-app-selector - sparkoperator.k8s.io/launched-by-spark-operator - serving.hops.works/id - serving.hops.works/name - serving.hops.works/tool - serving.hops.works/model-name - serving.hops.works/model-version - serving.hops.works/model-server - serving.hops.works/project-id ``` `olk.filebeat.kubernetesMetadata.namespaceAnnotations` # { #helm.olk.filebeat.kubernetesMetadata.namespaceAnnotations } : Type `list`, default `[]`. namespace annotations to keep. Empty collects none. `olk.filebeat.kubernetesMetadata.nodeAnnotations` # { #helm.olk.filebeat.kubernetesMetadata.nodeAnnotations } : Type `list`, default `[]`. node annotations to keep. Empty collects none. `olk.filebeat.kubernetesMetadata.podAnnotations` # { #helm.olk.filebeat.kubernetesMetadata.podAnnotations } : Type `list`, default `[]`. pod annotations to keep. Empty collects none, filebeat's default. `olk.filebeat.logs_locations[0].addKubernetesMetadata` # { #helm.olk.filebeat.logs_locations.0.addKubernetesMetadata } : Type `bool`, default `true`. attach Kubernetes pod/namespace/node metadata to logs from this location, using the label whitelists in filebeat.kubernetesMetadata `olk.filebeat.logs_locations[0].glob` # { #helm.olk.filebeat.logs_locations.0.glob } : Type `string`, default `"/*.log"`. `olk.filebeat.logs_locations[0].logtype` # { #helm.olk.filebeat.logs_locations.0.logtype } : Type `string`, default `"log"`. `olk.filebeat.logs_locations[0].mountPaths[0]` # { #helm.olk.filebeat.logs_locations.0.mountPaths.0 } : Type `string`, default `"/var/log/containers"`. `olk.filebeat.logs_locations[0].mountPaths[1]` # { #helm.olk.filebeat.logs_locations.0.mountPaths.1 } : Type `string`, default `"/var/log/pods"`. `olk.filebeat.logs_locations[0].name` # { #helm.olk.filebeat.logs_locations.0.name } : Type `string`, default `"containerd"`. `olk.filebeat.logs_locations[0].path` # { #helm.olk.filebeat.logs_locations.0.path } : Type `string`, default `"/var/log/containers"`. `olk.filebeat.logs_locations[0].processors` # { #helm.olk.filebeat.logs_locations.0.processors } : Type `list`, default `[]`. extra processors for this location, appended after the metadata processor. Each entry is itself a list. `olk.filebeat.name` # { #helm.olk.filebeat.name } : Type `string`, default `"filebeat"`. `olk.filebeat.resources.limits` # { #helm.olk.filebeat.resources.limits } : Type `object`, default `{"memory":"400Mi"}`. resources limits configuration `olk.filebeat.resources.requests.cpu` # { #helm.olk.filebeat.resources.requests.cpu } : Type `string`, default `"100m"`. `olk.filebeat.resources.requests.memory` # { #helm.olk.filebeat.resources.requests.memory } : Type `string`, default `"200Mi"`. `olk.filebeat.serviceAccount.annotations` # { #helm.olk.filebeat.serviceAccount.annotations } : Type `object`, default `{}`. service account annotations `olk.filebeat.terminationgraceperiod` # { #helm.olk.filebeat.terminationgraceperiod } : Type `int`, default `30`. `olk.filebeat.tolerations` # { #helm.olk.filebeat.tolerations } : Type `list`, default `[{"effect":"NoSchedule","operator":"Exists"}]`. tolerations for filebeat daemonset
## logstash { #helm-values-olk-logstash } ??? example "Defaults as YAML" ```yaml olk: logstash: additional_audit_outputs: '' batch_delay: 100 batch_size: 100 config: name: logstash-config debug: false extendServicesPipeline: [] extraPipelines: [] extraServices: [] extraVolumeMounts: [] extraVolumes: [] hpa: enabled: false maxReplicas: 3 targetCPUUtilizationPercentage: 80 targetMemoryUtilizationPercentage: 80 index_pattern: .services-%{+YYYY.MM.dd} name: logstash nodeSelector: {} plugins: [] podDisruptionBudget: enabled: true minAvailable: 1 port: 5044 replicas: 1 resources: requests: cpu: 200m memory: 2048Mi security_context: fsGroup: 1000 runAsGroup: 1000 runAsNonRoot: true runAsUser: 1000 service: annotations: consul.hashicorp.com/service-name: logstash consul.hashicorp.com/service-tags: jupyter,pythonjobs type: ClusterIP tolerations: [] topologySpreadConstraint: {} version: 7.17.29.6 workers: 4 ```
`olk.logstash.additional_audit_outputs` # { #helm.olk.logstash.additional_audit_outputs } : Type `string`, default `""`. `olk.logstash.batch_delay` # { #helm.olk.logstash.batch_delay } : Type `int`, default `100`. `olk.logstash.batch_size` # { #helm.olk.logstash.batch_size } : Type `int`, default `100`. `olk.logstash.config.name` # { #helm.olk.logstash.config.name } : Type `string`, default `"logstash-config"`. `olk.logstash.debug` # { #helm.olk.logstash.debug } : Type `bool`, default `false`. `olk.logstash.extendServicesPipeline` # { #helm.olk.logstash.extendServicesPipeline } : Type `list`, default `[]`. extend the services pipeline to collect logs for other services besides the default. You need to define the kuberentes label key and value to identifiy the service. The serviceLabelValue will be used as the service name, and you can add extra mutation logic to rename some of the fields of the service log. `olk.logstash.extraPipelines` # { #helm.olk.logstash.extraPipelines } : Type `list`, default `[]`. logstash extraPipelines `olk.logstash.extraServices` # { #helm.olk.logstash.extraServices } : Type `list`, default `[]`. logstash extraServices `olk.logstash.extraVolumeMounts` # { #helm.olk.logstash.extraVolumeMounts } : Type `list`, default `[]`. `olk.logstash.extraVolumes` # { #helm.olk.logstash.extraVolumes } : Type `list`, default `[]`. `olk.logstash.hpa.enabled` # { #helm.olk.logstash.hpa.enabled } : Type `bool`, default `false`. Autoscale Logstash. While enabled, logstash.replicas is not rendered and the HPA owns spec.replicas. `olk.logstash.hpa.maxReplicas` # { #helm.olk.logstash.hpa.maxReplicas } : Type `int`, default `3`. `olk.logstash.hpa.targetCPUUtilizationPercentage` # { #helm.olk.logstash.hpa.targetCPUUtilizationPercentage } : Type `int`, default `80`. `olk.logstash.hpa.targetMemoryUtilizationPercentage` # { #helm.olk.logstash.hpa.targetMemoryUtilizationPercentage } : Type `int`, default `80`. `olk.logstash.index_pattern` # { #helm.olk.logstash.index_pattern } : Type `string`, default `".services-%{+YYYY.MM.dd}"`. `olk.logstash.name` # { #helm.olk.logstash.name } : Type `string`, default `"logstash"`. `olk.logstash.nodeSelector` # { #helm.olk.logstash.nodeSelector } : Type `object`, default `{}`. node selector configuration `olk.logstash.plugins` # { #helm.olk.logstash.plugins } : Type `list`, default `[]`. `olk.logstash.podDisruptionBudget.enabled` # { #helm.olk.logstash.podDisruptionBudget.enabled } : Type `bool`, default `true`. `olk.logstash.podDisruptionBudget.minAvailable` # { #helm.olk.logstash.podDisruptionBudget.minAvailable } : Type `int`, default `1`. `olk.logstash.port` # { #helm.olk.logstash.port } : Type `int`, default `5044`. `olk.logstash.replicas` # { #helm.olk.logstash.replicas } : Type `int`, default `1`. Number of Logstash replicas. Not rendered while logstash.hpa.enabled is true: the HPA then owns spec.replicas and this value is its minReplicas. Turning the HPA on for a running release drops Logstash to 1 once, until the HPA scales it back up. `olk.logstash.resources` # { #helm.olk.logstash.resources } : Type `object`, default `{"requests":{"cpu":"200m","memory":"2048Mi"}}`. resources configuration `olk.logstash.security_context.fsGroup` # { #helm.olk.logstash.security_context.fsGroup } : Type `int`, default `1000`. `olk.logstash.security_context.runAsGroup` # { #helm.olk.logstash.security_context.runAsGroup } : Type `int`, default `1000`. `olk.logstash.security_context.runAsNonRoot` # { #helm.olk.logstash.security_context.runAsNonRoot } : Type `bool`, default `true`. `olk.logstash.security_context.runAsUser` # { #helm.olk.logstash.security_context.runAsUser } : Type `int`, default `1000`. `olk.logstash.service.annotations."consul.hashicorp.com/service-name"` # { #helm.olk.logstash.service.annotations.consul.hashicorp.com-service-name } : Type `string`, default `"logstash"`. `olk.logstash.service.annotations."consul.hashicorp.com/service-tags"` # { #helm.olk.logstash.service.annotations.consul.hashicorp.com-service-tags } : Type `string`, default `"jupyter,pythonjobs"`. `olk.logstash.service.type` # { #helm.olk.logstash.service.type } : Type `string`, default `"ClusterIP"`. `olk.logstash.tolerations` # { #helm.olk.logstash.tolerations } : Type `list`, default `[]`. `olk.logstash.topologySpreadConstraint` # { #helm.olk.logstash.topologySpreadConstraint } : Type `object`, default `{}`. The default topology spread constraint. If not defined the global topology spread constraint would be used instead. `olk.logstash.version` # { #helm.olk.logstash.version } : Type `string`, default `"7.17.29.6"`. `olk.logstash.workers` # { #helm.olk.logstash.workers } : Type `int`, default `4`.
## opensearch { #helm-values-olk-opensearch } ??? example "Defaults as YAML" ```yaml olk: opensearch: audit: enable_rest: true enable_transport: false enabled: false retentionDays: 7 backup: enabled: null repositories: {} ttl: null config: name: opensearch-config cors: allow_origin: '*' externalLoadBalancer: annotations: {} class: null enabled: null managed: null nodePort: null nodeSelector: {} indexUpgrade: activeDeadlineSeconds: 3600 enabled: true jvmOpts: '' knn: cache_expire: true circuit_breaker: percent: 75 triggered: false index_threads: 1 memory: circuit_breaker: limit: 50% mandatoryIndices: - projects - featurestore max_shards_per_node: 3000 memory: xms: 2g xmx: 2g name: opensearch nodeSelector: {} plugins: [] podDisruptionBudget: enabled: true minAvailable: 1 port: 9200 protocol: https replicas: 1 resources: limits: memory: 4000Mi requests: cpu: 100m memory: 2000Mi restore: enabled: null list_snapshots: true pause_backup_while_restoring: SKIP repositories: {} security: admin_locality: elkadmin create_secret: true default_tenant: Private extraDnsNames: [] extraIpAddresses: [] locality: elastic multitenancy_enabled: true service: annotations: consul.hashicorp.com/service-name: elastic consul.hashicorp.com/service-port: http consul.hashicorp.com/service-tags: rest headlessName: opensearch-headless serviceAccount: annotations: {} serviceAccountName: opensearch setVMMaxMapCount: true storageClassName: null storage_size: 20Gi tolerations: [] topologySpreadConstraint: {} transport_port: 9300 ttlSecondsAfterFinished: null version: 3.8.0.0 ```
`olk.opensearch.audit.enable_rest` # { #helm.olk.opensearch.audit.enable_rest } : Type `bool`, default `true`. Audit the REST layer (plugins.security.audit.config.enable_rest). Only applies when audit.enabled is true. `olk.opensearch.audit.enable_transport` # { #helm.olk.opensearch.audit.enable_transport } : Type `bool`, default `false`. Audit the transport layer (plugins.security.audit.config.enable_transport). Only applies when audit.enabled is true. `olk.opensearch.audit.enabled` # { #helm.olk.opensearch.audit.enabled } : Type `bool`, default `false`. Enable OpenSearch security audit logging to an in-cluster `security-auditlog-*` index. Disabled by default. When enabled, the chart also installs an ISM policy that deletes these daily indices after `audit.retentionDays`, so they don't grow unbounded. On single-node clusters the audit index stays yellow (its replica cannot be allocated; harmless); on multi-node it is green. `olk.opensearch.audit.retentionDays` # { #helm.olk.opensearch.audit.retentionDays } : Type `int`, default `7`. Days to retain `security-auditlog-*` indices before ISM deletes them. Only applies when audit.enabled is true, and bounds the otherwise-unbounded daily audit index growth. NOTE: the `security-auditlog-retention` ISM policy is created once and never updated, so changing this value on an existing cluster has no effect on its own; to apply a new value, delete that ISM policy so it is recreated with the new setting on the next sync. `olk.opensearch.backup.enabled` # { #helm.olk.opensearch.backup.enabled } : Type `string`, default `nil`. `olk.opensearch.backup.repositories` # { #helm.olk.opensearch.backup.repositories } : Type `object`, default `{}`. backup repository configuration `olk.opensearch.backup.ttl` # { #helm.olk.opensearch.backup.ttl } : Type `string`, default `nil`. time to live to control when to clean up backups. It is a number followed by either d (days) or h (hours) suffix. `olk.opensearch.config.name` # { #helm.olk.opensearch.config.name } : Type `string`, default `"opensearch-config"`. `olk.opensearch.cors.allow_origin` # { #helm.olk.opensearch.cors.allow_origin } : Type `string`, default `"*"`. `olk.opensearch.externalLoadBalancer.annotations` # { #helm.olk.opensearch.externalLoadBalancer.annotations } : Type `object`, default `{}`. load balancer annotations `olk.opensearch.externalLoadBalancer.class` # { #helm.olk.opensearch.externalLoadBalancer.class } : Type `string`, default `nil`. load balancer class name `olk.opensearch.externalLoadBalancer.enabled` # { #helm.olk.opensearch.externalLoadBalancer.enabled } : Type `string`, default `nil`. Enable External Load Balancers for Opensearch. If not set the .global._hopsworks.externalLoadBalancers.enabled will be used instead `olk.opensearch.externalLoadBalancer.managed` # { #helm.olk.opensearch.externalLoadBalancer.managed } : Type `string`, default `nil`. Cloud provider provisions Load Balancers. If not set the .global._hopsworks.externalLoadBalancers.managed will be used instead `olk.opensearch.externalLoadBalancer.nodePort` # { #helm.olk.opensearch.externalLoadBalancer.nodePort } : Type `string`, default `nil`. Explicit nodePort for the external service when the load balancer is unmanaged (managed: false), so a load balancer outside Kubernetes can target a fixed port. Null lets Kubernetes allocate one from the cluster's node-port range; a set value must lie in that range (30000-32767 by default), which the API server enforces at install. `olk.opensearch.externalLoadBalancer.nodeSelector` # { #helm.olk.opensearch.externalLoadBalancer.nodeSelector } : Type `object`, default `{}`. selector for nodes the load balancer can use to route traffic `olk.opensearch.indexUpgrade.activeDeadlineSeconds` # { #helm.olk.opensearch.indexUpgrade.activeDeadlineSeconds } : Type `int`, default `3600`. Deadline for the whole pre-upgrade index pass. A reindex copies each index twice to preserve its name, so raise this for clusters holding large pre-2.0 indices; the hook prints the sizes and an estimate before it starts. Bounds a stuck copy so it fails the hook instead of hanging the upgrade. The timeout given to helm upgrade has to be at least this long, or helm gives up on the hook first. The deadline is per release: a satellite has its own OpenSearch and its own pass at its own upgrade, so it sets its own value. Upgrade every satellite to 5.1 before central moves to 5.2, since a satellite still on 1.3 is two majors behind a 5.2 central. A satellite moving to 5.2 ahead of central is untested. `olk.opensearch.indexUpgrade.enabled` # { #helm.olk.opensearch.indexUpgrade.enabled } : Type `bool`, default `true`. Let the pre-upgrade hook fix a cluster that still holds indices created before OpenSearch 2.0, which a 3.x node refuses to boot with. It reindexes the indices that carry data, preserving their names, and deletes only the ones the platform recreates or expires by itself (logs, audit, `pypi_libraries_*`, query insights and security analytics plugin indices); anything it does not recognise is reindexed, never deleted, including `featurestore`, `projects`, ISM policies and `.kibana*`. Embedding indices (`__embedding*`) come back on the faiss engine rather than nmslib, which is what Hopsworks 5.1 creates and the only one of the two that accepts the filter the 5.1+ client sends; approximate neighbour results can shift slightly. Every other index keeps its engine. While it copies, every client but the admin certificate is refused, reads included, so Hopsworks search, OpenSearch Dashboards, logstash and OnlineFS get a 403 until the hook exits; a hook killed outright leaves them refused until the upgrade is rerun. Set to false to have the hook report the offending indices and fail instead, leaving the remediation to the operator. A cluster still running OpenSearch 1.x is refused rather than fixed, whatever this is set to: reindexing on a 1.x node recreates the index pre-2.0 again, so a 1.x cluster has to be upgraded to a 2.x Hopsworks release first. `olk.opensearch.jvmOpts` # { #helm.olk.opensearch.jvmOpts } : Type `string`, default `""`. JVM Options to provide to Opensearch `olk.opensearch.knn.cache_expire` # { #helm.olk.opensearch.knn.cache_expire } : Type `bool`, default `true`. `olk.opensearch.knn.circuit_breaker.percent` # { #helm.olk.opensearch.knn.circuit_breaker.percent } : Type `float`, default `75`. knn.circuit_breaker.unset.percentage: a tripped k-NN breaker clears once the graph cache falls below this percentage (0 to 100) of knn.memory.circuit_breaker.limit `olk.opensearch.knn.circuit_breaker.triggered` # { #helm.olk.opensearch.knn.circuit_breaker.triggered } : Type `bool`, default `false`. knn.circuit_breaker.triggered is a flag the k-NN plugin sets on a memory trip and clears itself. Keep it false: OpenSearch 2.6 and later apply it at node start, and true rejects every vector write, translog replay included, until the plugin clears it `olk.opensearch.knn.index_threads` # { #helm.olk.opensearch.knn.index_threads } : Type `int`, default `1`. `olk.opensearch.knn.memory.circuit_breaker.limit` # { #helm.olk.opensearch.knn.memory.circuit_breaker.limit } : Type `string`, default `"50%"`. `olk.opensearch.mandatoryIndices[0]` # { #helm.olk.opensearch.mandatoryIndices.0 } : Type `string`, default `"projects"`. `olk.opensearch.mandatoryIndices[1]` # { #helm.olk.opensearch.mandatoryIndices.1 } : Type `string`, default `"featurestore"`. `olk.opensearch.max_shards_per_node` # { #helm.olk.opensearch.max_shards_per_node } : Type `int`, default `3000`. `olk.opensearch.memory.xms` # { #helm.olk.opensearch.memory.xms } : Type `string`, default `"2g"`. `olk.opensearch.memory.xmx` # { #helm.olk.opensearch.memory.xmx } : Type `string`, default `"2g"`. `olk.opensearch.name` # { #helm.olk.opensearch.name } : Type `string`, default `"opensearch"`. `olk.opensearch.nodeSelector` # { #helm.olk.opensearch.nodeSelector } : Type `object`, default `{}`. node selector configuration `olk.opensearch.plugins` # { #helm.olk.opensearch.plugins } : Type `list`, default `[]`. `olk.opensearch.podDisruptionBudget.enabled` # { #helm.olk.opensearch.podDisruptionBudget.enabled } : Type `bool`, default `true`. `olk.opensearch.podDisruptionBudget.minAvailable` # { #helm.olk.opensearch.podDisruptionBudget.minAvailable } : Type `int`, default `1`. `olk.opensearch.port` # { #helm.olk.opensearch.port } : Type `int`, default `9200`. `olk.opensearch.protocol` # { #helm.olk.opensearch.protocol } : Type `string`, default `"https"`. `olk.opensearch.replicas` # { #helm.olk.opensearch.replicas } : Type `int`, default `1`. `olk.opensearch.resources.limits` # { #helm.olk.opensearch.resources.limits } : Type `object`, default `{"memory":"4000Mi"}`. resources limits configuration `olk.opensearch.resources.requests.cpu` # { #helm.olk.opensearch.resources.requests.cpu } : Type `string`, default `"100m"`. `olk.opensearch.resources.requests.memory` # { #helm.olk.opensearch.resources.requests.memory } : Type `string`, default `"2000Mi"`. `olk.opensearch.restore.enabled` # { #helm.olk.opensearch.restore.enabled } : Type `string`, default `nil`. If true an initial job will be created to restore the data from the backup. Backup jobs will bypass creating snapshots while this job is running. `olk.opensearch.restore.list_snapshots` # { #helm.olk.opensearch.restore.list_snapshots } : Type `bool`, default `true`. `olk.opensearch.restore.pause_backup_while_restoring` # { #helm.olk.opensearch.restore.pause_backup_while_restoring } : Type `string`, default `"SKIP"`. If set, the backup jobs will be paused while the restore job is running. Values are SKIP or PAUSE, in case none, the backup will wait for the restore to finish `olk.opensearch.restore.repositories` # { #helm.olk.opensearch.restore.repositories } : Type `object`, default `{}`. restore repositories configuration `olk.opensearch.security.admin_locality` # { #helm.olk.opensearch.security.admin_locality } : Type `string`, default `"elkadmin"`. `olk.opensearch.security.create_secret` # { #helm.olk.opensearch.security.create_secret } : Type `bool`, default `true`. Create the opensearch-users-secrets secret with the opensearch internal users passwords `olk.opensearch.security.default_tenant` # { #helm.olk.opensearch.security.default_tenant } : Type `string`, default `"Private"`. Default OpenSearch tenant for Dashboards users. A non-empty value suppresses the OpenSearch Dashboards 2.x "Select your tenant" popup (the plugin only prompts when the reported default_tenant is empty). Hopsworks links open Dashboards with the Private tenant, so "Private" matches that and keeps multitenancy/index-pattern isolation intact. Set to "" to restore the popup / leave the default tenant unset. `olk.opensearch.security.extraDnsNames` # { #helm.olk.opensearch.security.extraDnsNames } : Type `list`, default `[]`. Additional DNS names to add as SAN to Datanode x.509 certificate `olk.opensearch.security.extraIpAddresses` # { #helm.olk.opensearch.security.extraIpAddresses } : Type `list`, default `[]`. Additional IP addresses to add as SAN to Datanode x.509 certificate `olk.opensearch.security.locality` # { #helm.olk.opensearch.security.locality } : Type `string`, default `"elastic"`. `olk.opensearch.security.multitenancy_enabled` # { #helm.olk.opensearch.security.multitenancy_enabled } : Type `bool`, default `true`. `olk.opensearch.service.annotations."consul.hashicorp.com/service-name"` # { #helm.olk.opensearch.service.annotations.consul.hashicorp.com-service-name } : Type `string`, default `"elastic"`. `olk.opensearch.service.annotations."consul.hashicorp.com/service-port"` # { #helm.olk.opensearch.service.annotations.consul.hashicorp.com-service-port } : Type `string`, default `"http"`. `olk.opensearch.service.annotations."consul.hashicorp.com/service-tags"` # { #helm.olk.opensearch.service.annotations.consul.hashicorp.com-service-tags } : Type `string`, default `"rest"`. `olk.opensearch.service.headlessName` # { #helm.olk.opensearch.service.headlessName } : Type `string`, default `"opensearch-headless"`. `olk.opensearch.serviceAccount.annotations` # { #helm.olk.opensearch.serviceAccount.annotations } : Type `object`, default `{}`. service account annotations `olk.opensearch.serviceAccountName` # { #helm.olk.opensearch.serviceAccountName } : Type `string`, default `"opensearch"`. `olk.opensearch.setVMMaxMapCount` # { #helm.olk.opensearch.setVMMaxMapCount } : Type `bool`, default `true`. `olk.opensearch.storageClassName` # { #helm.olk.opensearch.storageClassName } : Type `string`, default `nil`. storage class name `olk.opensearch.storage_size` # { #helm.olk.opensearch.storage_size } : Type `string`, default `"20Gi"`. `olk.opensearch.tolerations` # { #helm.olk.opensearch.tolerations } : Type `list`, default `[]`. `olk.opensearch.topologySpreadConstraint` # { #helm.olk.opensearch.topologySpreadConstraint } : Type `object`, default `{}`. The default topology spread constraint. If not defined the global topology spread constraint would be used instead. `olk.opensearch.transport_port` # { #helm.olk.opensearch.transport_port } : Type `int`, default `9300`. `olk.opensearch.ttlSecondsAfterFinished` # { #helm.olk.opensearch.ttlSecondsAfterFinished } : Type `string`, default `nil`. TTL in seconds for the create-repos Job. Overrides global default. `olk.opensearch.version` # { #helm.olk.opensearch.version } : Type `string`, default `"3.8.0.0"`.
================================================================================ # onlinefs Source: https://docs.hopsworks.ai/latest/setup_installation/common/helm_chart_values/onlinefs/ # OnlineFS values { #helm-values-onlinefs } Values under `onlinefs` configure OnlineFS, which consumes feature rows from Kafka and writes them to RonDB. _Generated from the Hopsworks Helm chart `5.2.0-alpha-1791549041` (Hopsworks `5.2.0`)._ Always deployed. ## General { #helm-values-onlinefs-general } ??? example "Defaults as YAML" ```yaml onlinefs: apiKey: email: onlinefs@hopsworks.ai password: onlinefspw secret_name: onlinefs-api-key-auth appName: onlinefs configmap: name: onlinefs-configmap debug: false hopsworkslib: {} hpa: enabled: false maxReplicas: 5 targetCPUUtilizationPercentage: 80 targetMemoryUtilizationPercentage: 80 nodeSelector: {} podDisruptionBudget: enabled: true minAvailable: 1 serviceAccount: annotations: {} create: true name: onlinefs-default setupBackOffLimit: 10 tlscerts: commonName: '' locality: onlinefs tolerations: [] topologySpreadConstraint: {} ```
`onlinefs` # { #helm.onlinefs } : Type `object`, default `{"debug":false}`. override onlinefs values `onlinefs.apiKey.email` # { #helm.onlinefs.apiKey.email } : Type `string`, default `"onlinefs@hopsworks.ai"`. `onlinefs.apiKey.password` # { #helm.onlinefs.apiKey.password } : Type `string`, default `"onlinefspw"`. `onlinefs.apiKey.secret_name` # { #helm.onlinefs.apiKey.secret_name } : Type `string`, default `"onlinefs-api-key-auth"`. `onlinefs.appName` # { #helm.onlinefs.appName } : Type `string`, default `"onlinefs"`. `onlinefs.configmap.name` # { #helm.onlinefs.configmap.name } : Type `string`, default `"onlinefs-configmap"`. `onlinefs.debug` # { #helm.onlinefs.debug } : Type `bool`, default `false`. `onlinefs.hopsworkslib` # { #helm.onlinefs.hopsworkslib } : Type `object`, default `{}`. override hopsworkslib values `onlinefs.hpa.enabled` # { #helm.onlinefs.hpa.enabled } : Type `bool`, default `false`. Autoscale the Deployment with an HPA. While enabled, deployment.replicas is not rendered and the HPA owns spec.replicas. `onlinefs.hpa.maxReplicas` # { #helm.onlinefs.hpa.maxReplicas } : Type `int`, default `5`. `onlinefs.hpa.targetCPUUtilizationPercentage` # { #helm.onlinefs.hpa.targetCPUUtilizationPercentage } : Type `int`, default `80`. `onlinefs.hpa.targetMemoryUtilizationPercentage` # { #helm.onlinefs.hpa.targetMemoryUtilizationPercentage } : Type `int`, default `80`. `onlinefs.nodeSelector` # { #helm.onlinefs.nodeSelector } : Type `object`, default `{}`. node selector configuration `onlinefs.podDisruptionBudget.enabled` # { #helm.onlinefs.podDisruptionBudget.enabled } : Type `bool`, default `true`. `onlinefs.podDisruptionBudget.minAvailable` # { #helm.onlinefs.podDisruptionBudget.minAvailable } : Type `int`, default `1`. `onlinefs.serviceAccount.annotations` # { #helm.onlinefs.serviceAccount.annotations } : Type `object`, default `{}`. service account annotations `onlinefs.serviceAccount.create` # { #helm.onlinefs.serviceAccount.create } : Type `bool`, default `true`. `onlinefs.serviceAccount.name` # { #helm.onlinefs.serviceAccount.name } : Type `string`, default `"onlinefs-default"`. `onlinefs.setupBackOffLimit` # { #helm.onlinefs.setupBackOffLimit } : Type `int`, default `10`. `onlinefs.tlscerts.commonName` # { #helm.onlinefs.tlscerts.commonName } : Type `string`, default `""`. `onlinefs.tlscerts.locality` # { #helm.onlinefs.tlscerts.locality } : Type `string`, default `"onlinefs"`. `onlinefs.tolerations` # { #helm.onlinefs.tolerations } : Type `list`, default `[]`. `onlinefs.topologySpreadConstraint` # { #helm.onlinefs.topologySpreadConstraint } : Type `object`, default `{}`. The default topology spread constraint. If not defined the global topology spread constraint would be used instead.
## configmapOverrides { #helm-values-onlinefs-configmapoverrides } ??? example "Defaults as YAML" ```yaml onlinefs: configmapOverrides: kafka: {} kafkaVectorDb: {} log4j: {} onlinefssite: hopsworks: tokenLocation: /onlinefs/etc/token kafkaConsumer: pollTimeoutMs: 1000 topicList: '' topicPattern: .*_onlinefs rondb: batchSize: 300 maxCachedInstances: 1024 maxCachedSessions: 20 maxTransactions: 1024 poolSize: 1 reconnectTimeout: 5 useDynamicObjectCache: false useSessionCache: false service: featureGroupCacheExpire: 30 featureStoreCacheExpire: 30 featureViewCacheExpire: 10 getSessionRetrySleepMs: 100 maxBlacklistSize: 100 maxFeatureGroupCacheSize: 1000 maxFeatureStoreCacheSize: 1000 maxFeatureViewCacheSize: 1000 pauseRetrySleepMs: 5000 ronDbThreadNumber: 10 threadNumber: 10 vectorDbThreadNumber: 10 producer: {} ```
`onlinefs.configmapOverrides.kafka` # { #helm.onlinefs.configmapOverrides.kafka } : Type `object`, default `{}`. Extra properties to add or overwrite in onlinefs-kafka.properties `onlinefs.configmapOverrides.kafkaVectorDb` # { #helm.onlinefs.configmapOverrides.kafkaVectorDb } : Type `object`, default `{}`. Extra properties to add or overwrite in onlinefs-kafka-vector-db.properties `onlinefs.configmapOverrides.log4j` # { #helm.onlinefs.configmapOverrides.log4j } : Type `object`, default `{}`. Extra properties to add or overwrite in log4j.properties `onlinefs.configmapOverrides.onlinefssite.hopsworks.tokenLocation` # { #helm.onlinefs.configmapOverrides.onlinefssite.hopsworks.tokenLocation } : Type `string`, default `"/onlinefs/etc/token"`. `onlinefs.configmapOverrides.onlinefssite.kafkaConsumer.pollTimeoutMs` # { #helm.onlinefs.configmapOverrides.onlinefssite.kafkaConsumer.pollTimeoutMs } : Type `int`, default `1000`. `onlinefs.configmapOverrides.onlinefssite.kafkaConsumer.topicList` # { #helm.onlinefs.configmapOverrides.onlinefssite.kafkaConsumer.topicList } : Type `string`, default `""`. `onlinefs.configmapOverrides.onlinefssite.kafkaConsumer.topicPattern` # { #helm.onlinefs.configmapOverrides.onlinefssite.kafkaConsumer.topicPattern } : Type `string`, default `".*_onlinefs"`. `onlinefs.configmapOverrides.onlinefssite.rondb.batchSize` # { #helm.onlinefs.configmapOverrides.onlinefssite.rondb.batchSize } : Type `int`, default `300`. `onlinefs.configmapOverrides.onlinefssite.rondb.maxCachedInstances` # { #helm.onlinefs.configmapOverrides.onlinefssite.rondb.maxCachedInstances } : Type `int`, default `1024`. `onlinefs.configmapOverrides.onlinefssite.rondb.maxCachedSessions` # { #helm.onlinefs.configmapOverrides.onlinefssite.rondb.maxCachedSessions } : Type `int`, default `20`. `onlinefs.configmapOverrides.onlinefssite.rondb.maxTransactions` # { #helm.onlinefs.configmapOverrides.onlinefssite.rondb.maxTransactions } : Type `int`, default `1024`. `onlinefs.configmapOverrides.onlinefssite.rondb.poolSize` # { #helm.onlinefs.configmapOverrides.onlinefssite.rondb.poolSize } : Type `int`, default `1`. `onlinefs.configmapOverrides.onlinefssite.rondb.reconnectTimeout` # { #helm.onlinefs.configmapOverrides.onlinefssite.rondb.reconnectTimeout } : Type `int`, default `5`. `onlinefs.configmapOverrides.onlinefssite.rondb.useDynamicObjectCache` # { #helm.onlinefs.configmapOverrides.onlinefssite.rondb.useDynamicObjectCache } : Type `bool`, default `false`. `onlinefs.configmapOverrides.onlinefssite.rondb.useSessionCache` # { #helm.onlinefs.configmapOverrides.onlinefssite.rondb.useSessionCache } : Type `bool`, default `false`. `onlinefs.configmapOverrides.onlinefssite.service.featureGroupCacheExpire` # { #helm.onlinefs.configmapOverrides.onlinefssite.service.featureGroupCacheExpire } : Type `int`, default `30`. `onlinefs.configmapOverrides.onlinefssite.service.featureStoreCacheExpire` # { #helm.onlinefs.configmapOverrides.onlinefssite.service.featureStoreCacheExpire } : Type `int`, default `30`. `onlinefs.configmapOverrides.onlinefssite.service.featureViewCacheExpire` # { #helm.onlinefs.configmapOverrides.onlinefssite.service.featureViewCacheExpire } : Type `int`, default `10`. `onlinefs.configmapOverrides.onlinefssite.service.getSessionRetrySleepMs` # { #helm.onlinefs.configmapOverrides.onlinefssite.service.getSessionRetrySleepMs } : Type `int`, default `100`. `onlinefs.configmapOverrides.onlinefssite.service.maxBlacklistSize` # { #helm.onlinefs.configmapOverrides.onlinefssite.service.maxBlacklistSize } : Type `int`, default `100`. `onlinefs.configmapOverrides.onlinefssite.service.maxFeatureGroupCacheSize` # { #helm.onlinefs.configmapOverrides.onlinefssite.service.maxFeatureGroupCacheSize } : Type `int`, default `1000`. `onlinefs.configmapOverrides.onlinefssite.service.maxFeatureStoreCacheSize` # { #helm.onlinefs.configmapOverrides.onlinefssite.service.maxFeatureStoreCacheSize } : Type `int`, default `1000`. `onlinefs.configmapOverrides.onlinefssite.service.maxFeatureViewCacheSize` # { #helm.onlinefs.configmapOverrides.onlinefssite.service.maxFeatureViewCacheSize } : Type `int`, default `1000`. `onlinefs.configmapOverrides.onlinefssite.service.pauseRetrySleepMs` # { #helm.onlinefs.configmapOverrides.onlinefssite.service.pauseRetrySleepMs } : Type `int`, default `5000`. `onlinefs.configmapOverrides.onlinefssite.service.ronDbThreadNumber` # { #helm.onlinefs.configmapOverrides.onlinefssite.service.ronDbThreadNumber } : Type `int`, default `10`. `onlinefs.configmapOverrides.onlinefssite.service.threadNumber` # { #helm.onlinefs.configmapOverrides.onlinefssite.service.threadNumber } : Type `int`, default `10`. `onlinefs.configmapOverrides.onlinefssite.service.vectorDbThreadNumber` # { #helm.onlinefs.configmapOverrides.onlinefssite.service.vectorDbThreadNumber } : Type `int`, default `10`. `onlinefs.configmapOverrides.producer` # { #helm.onlinefs.configmapOverrides.producer } : Type `object`, default `{}`. Extra properties to add or overwrite in producer.properties
## dependencies { #helm-values-onlinefs-dependencies } ??? example "Defaults as YAML" ```yaml onlinefs: dependencies: glassfish: consulServiceName: glassfish consulServiceTag: hopsworks port: 8182 kafka: consulServiceName: kafka consulServiceTag: broker port: 9092 securityProtocol: SSL url: '' mgmd: consulServiceName: mgmd port: 1186 url: '' opensearch: consulServiceName: elastic consulServiceTag: rest port: 9200 ```
`onlinefs.dependencies.glassfish.consulServiceName` # { #helm.onlinefs.dependencies.glassfish.consulServiceName } : Type `string`, default `"glassfish"`. `onlinefs.dependencies.glassfish.consulServiceTag` # { #helm.onlinefs.dependencies.glassfish.consulServiceTag } : Type `string`, default `"hopsworks"`. `onlinefs.dependencies.glassfish.port` # { #helm.onlinefs.dependencies.glassfish.port } : Type `int`, default `8182`. `onlinefs.dependencies.kafka.consulServiceName` # { #helm.onlinefs.dependencies.kafka.consulServiceName } : Type `string`, default `"kafka"`. `onlinefs.dependencies.kafka.consulServiceTag` # { #helm.onlinefs.dependencies.kafka.consulServiceTag } : Type `string`, default `"broker"`. `onlinefs.dependencies.kafka.port` # { #helm.onlinefs.dependencies.kafka.port } : Type `int`, default `9092`. `onlinefs.dependencies.kafka.securityProtocol` # { #helm.onlinefs.dependencies.kafka.securityProtocol } : Type `string`, default `"SSL"`. Kafka security.protocol. PLAINTEXT or SSL `onlinefs.dependencies.kafka.url` # { #helm.onlinefs.dependencies.kafka.url } : Type `string`, default `""`. The Kafka Brokers URL (can be comma-separated urls with port or a single url), if not set the consul DNS name will be used `onlinefs.dependencies.mgmd.consulServiceName` # { #helm.onlinefs.dependencies.mgmd.consulServiceName } : Type `string`, default `"mgmd"`. `onlinefs.dependencies.mgmd.port` # { #helm.onlinefs.dependencies.mgmd.port } : Type `int`, default `1186`. `onlinefs.dependencies.mgmd.url` # { #helm.onlinefs.dependencies.mgmd.url } : Type `string`, default `""`. The mgmd URL, if not set the consul DNS name will be used `onlinefs.dependencies.opensearch.consulServiceName` # { #helm.onlinefs.dependencies.opensearch.consulServiceName } : Type `string`, default `"elastic"`. `onlinefs.dependencies.opensearch.consulServiceTag` # { #helm.onlinefs.dependencies.opensearch.consulServiceTag } : Type `string`, default `"rest"`. `onlinefs.dependencies.opensearch.port` # { #helm.onlinefs.dependencies.opensearch.port } : Type `int`, default `9200`.
## deployment { #helm-values-onlinefs-deployment } ??? example "Defaults as YAML" ```yaml onlinefs: deployment: env: JAVA_OPTS: -agentlib:jdwp=transport=dt_socket,server=y,suspend=n,address=*:9029 name: onlinefs-deployment ports: debug: 9029 monitor: 12800 replicas: 1 resources: limits: cpu: '4' memory: 4096Mi requests: cpu: '2' memory: 1024Mi strategy: {} vectordb: enabled: true ```
`onlinefs.deployment.env.JAVA_OPTS` # { #helm.onlinefs.deployment.env.JAVA_OPTS } : Type `string`, default `"-agentlib:jdwp=transport=dt_socket,server=y,suspend=n,address=*:9029"`. `onlinefs.deployment.name` # { #helm.onlinefs.deployment.name } : Type `string`, default `"onlinefs-deployment"`. `onlinefs.deployment.ports.debug` # { #helm.onlinefs.deployment.ports.debug } : Type `int`, default `9029`. `onlinefs.deployment.ports.monitor` # { #helm.onlinefs.deployment.ports.monitor } : Type `int`, default `12800`. `onlinefs.deployment.replicas` # { #helm.onlinefs.deployment.replicas } : Type `int`, default `1`. Number of replicas. Not rendered while hpa.enabled is true: the HPA then owns spec.replicas and this value is its minReplicas. Turning the HPA on for a running release drops the Deployment to 1 once, until the HPA scales it back up. `onlinefs.deployment.resources.limits.cpu` # { #helm.onlinefs.deployment.resources.limits.cpu } : Type `string`, default `"4"`. `onlinefs.deployment.resources.limits.memory` # { #helm.onlinefs.deployment.resources.limits.memory } : Type `string`, default `"4096Mi"`. `onlinefs.deployment.resources.requests.cpu` # { #helm.onlinefs.deployment.resources.requests.cpu } : Type `string`, default `"2"`. `onlinefs.deployment.resources.requests.memory` # { #helm.onlinefs.deployment.resources.requests.memory } : Type `string`, default `"1024Mi"`. `onlinefs.deployment.strategy` # { #helm.onlinefs.deployment.strategy } : Type `object`, default `{}`. spec.strategy of the Deployment, rendered as given when set. Left empty, two or more replicas roll out by terminating a pod before creating its replacement (maxSurge 0, maxUnavailable 1), so an upgrade completes on a node pool with no room for an extra pod, and a single replica keeps the Kubernetes default. Set it explicitly for a single-replica install whose node pool cannot fit a surge pod, or to keep surge-first behaviour on an install with spare capacity that would rather not drop a consumer during a rollout. `onlinefs.deployment.vectordb.enabled` # { #helm.onlinefs.deployment.vectordb.enabled } : Type `bool`, default `true`.
## rbac { #helm-values-onlinefs-rbac } ??? example "Defaults as YAML" ```yaml onlinefs: rbac: annotations: {} create: true extraRoleRules: [] name: onlinefs-role useExistingRole: false ```
`onlinefs.rbac.annotations` # { #helm.onlinefs.rbac.annotations } : Type `object`, default `{}`. rbac annotations `onlinefs.rbac.create` # { #helm.onlinefs.rbac.create } : Type `bool`, default `true`. `onlinefs.rbac.extraRoleRules` # { #helm.onlinefs.rbac.extraRoleRules } : Type `list`, default `[]`. `onlinefs.rbac.name` # { #helm.onlinefs.rbac.name } : Type `string`, default `"onlinefs-role"`. `onlinefs.rbac.useExistingRole` # { #helm.onlinefs.rbac.useExistingRole } : Type `bool`, default `false`.
## service { #helm-values-onlinefs-service } ??? example "Defaults as YAML" ```yaml onlinefs: service: debug: annotations: consul.hashicorp.com/service-name: onlinefs consul.hashicorp.com/service-tags: debug name: onlinefs-debug port: 9029 monitoring: annotations: consul.hashicorp.com/service-name: onlinefs consul.hashicorp.com/service-tags: monitor prometheus.io/path: /metrics prometheus.io/port: 12800 prometheus.io/scheme: http prometheus.io/scrape: 'true' name: onlinefs-monitor port: 12800 ```
`onlinefs.service.debug.annotations."consul.hashicorp.com/service-name"` # { #helm.onlinefs.service.debug.annotations.consul.hashicorp.com-service-name } : Type `string`, default `"onlinefs"`. `onlinefs.service.debug.annotations."consul.hashicorp.com/service-tags"` # { #helm.onlinefs.service.debug.annotations.consul.hashicorp.com-service-tags } : Type `string`, default `"debug"`. `onlinefs.service.debug.name` # { #helm.onlinefs.service.debug.name } : Type `string`, default `"onlinefs-debug"`. `onlinefs.service.debug.port` # { #helm.onlinefs.service.debug.port } : Type `int`, default `9029`. `onlinefs.service.monitoring.annotations."consul.hashicorp.com/service-name"` # { #helm.onlinefs.service.monitoring.annotations.consul.hashicorp.com-service-name } : Type `string`, default `"onlinefs"`. `onlinefs.service.monitoring.annotations."consul.hashicorp.com/service-tags"` # { #helm.onlinefs.service.monitoring.annotations.consul.hashicorp.com-service-tags } : Type `string`, default `"monitor"`. `onlinefs.service.monitoring.annotations."prometheus.io/path"` # { #helm.onlinefs.service.monitoring.annotations.prometheus.io-path } : Type `string`, default `"/metrics"`. `onlinefs.service.monitoring.annotations."prometheus.io/port"` # { #helm.onlinefs.service.monitoring.annotations.prometheus.io-port } : Type `int`, default `12800`. `onlinefs.service.monitoring.annotations."prometheus.io/scheme"` # { #helm.onlinefs.service.monitoring.annotations.prometheus.io-scheme } : Type `string`, default `"http"`. `onlinefs.service.monitoring.annotations."prometheus.io/scrape"` # { #helm.onlinefs.service.monitoring.annotations.prometheus.io-scrape } : Type `string`, default `"true"`. `onlinefs.service.monitoring.name` # { #helm.onlinefs.service.monitoring.name } : Type `string`, default `"onlinefs-monitor"`. `onlinefs.service.monitoring.port` # { #helm.onlinefs.service.monitoring.port } : Type `int`, default `12800`.
================================================================================ # prometheus Source: https://docs.hopsworks.ai/latest/setup_installation/common/helm_chart_values/prometheus/ # Prometheus values { #helm-values-prometheus } Values under `prometheus` configure Prometheus, which collects cluster and service metrics, and the Prometheus adapter, which exposes them to autoscalers. _Generated from the Hopsworks Helm chart `5.2.0-alpha-1791549041` (Hopsworks `5.2.0`)._ Deployed when [`global._hopsworks.full_platform`](global.md#helm.global._hopsworks.full_platform) is `true`. !!! info "Upstream charts" - Values under `prometheus.prometheus` go to [`prometheus` 25.20.2](https://artifacthub.io/packages/helm/prometheus-community/prometheus/25.20.2) from `https://prometheus-community.github.io/helm-charts`. - Values under `prometheus.prometheus-adapter` go to [`prometheus-adapter` 4.11.0](https://artifacthub.io/packages/helm/prometheus-community/prometheus-adapter/4.11.0) from `https://prometheus-community.github.io/helm-charts`. Only the values Hopsworks sets under `prometheus.prometheus` and `prometheus.prometheus-adapter` are listed on this page. Any other value of the charts can be set under the same keys; each link opens the chart's documentation for the version Hopsworks pins. ??? example "Defaults as YAML" ```yaml prometheus: adapter: enabled: true configAnnotations: {} externalLoadBalancer: enabled: false managed: null nodeSelector: {} hopsworkslib: {} prometheus: server: persistentVolume: enabled: true storageClass: null prometheus-adapter: dnsConfig: searches: - service.consul image: pullPolicy: IfNotPresent repository: docker.hops.works/prometheus/prometheus-adapter tag: 0.12.0-alpine-h1.1 nameOverride: hopsworks-prometheus-adapter nodeSelector: {} prometheus: path: '' port: 9090 url: http://prometheus.service.consul resources: limits: cpu: 100m memory: 70Mi requests: cpu: 50m memory: 40Mi rules: existing: custom-prometheus-metrics-adapter tolerations: [] topologySpreadConstraints: - labelSelector: matchLabels: app.kubernetes.io/instance: hopsworks app.kubernetes.io/name: hopsworks-prometheus-adapter maxSkew: 1 topologyKey: topology.kubernetes.io/zone whenUnsatisfiable: ScheduleAnyway ```
`prometheus` # { #helm.prometheus } : Type `object`. override prometheus values ??? note "Default" ```yaml prometheus: server: persistentVolume: enabled: true storageClass: null ``` `prometheus.adapter.enabled` # { #helm.prometheus.adapter.enabled } : Type `bool`, default `true`. `prometheus.configAnnotations` # { #helm.prometheus.configAnnotations } : Type `object`, default `{}`. alertmanager-tmpl config map annotations `prometheus.externalLoadBalancer.enabled` # { #helm.prometheus.externalLoadBalancer.enabled } : Type `bool`, default `false`. Enable External Load Balancers for Prometheus server. If not set the .global._hopsworks.externalLoadBalancers.enabled will be used instead. If enabled AND you want the LoadBalancer to be **managed** you shall set prometheus.server.service.type: LoadBalancer - NOTE: Consul domain name will be a CNAME to LoadBalancer domain name If enabled AND you want the LoadBalancer to be **unmanged** you shall set prometheus.server.service.type: NodePort - NOTE: This only works in AWS `prometheus.externalLoadBalancer.managed` # { #helm.prometheus.externalLoadBalancer.managed } : Type `string`, default `nil`. Cloud provider provisions Load Balancers. If not set the .global._hopsworks.externalLoadBalancers.managed will be used instead `prometheus.externalLoadBalancer.nodeSelector` # { #helm.prometheus.externalLoadBalancer.nodeSelector } : Type `object`, default `{}`. selector for nodes the load balancer can use to route traffic `prometheus.hopsworkslib` # { #helm.prometheus.hopsworkslib } : Type `object`, default `{}`. override hopsworkslib values `prometheus.prometheus` # { #helm.prometheus.prometheus } : Type `object`, default check \[values.yaml\](./values.yaml) for more information, passed to the [`prometheus` 25.20.2](https://artifacthub.io/packages/helm/prometheus-community/prometheus/25.20.2) chart, whose other values are documented there. override prometheus values `prometheus.prometheus-adapter` # { #helm.prometheus.prometheus-adapter } : Type `object`, passed to the [`prometheus-adapter` 4.11.0](https://artifacthub.io/packages/helm/prometheus-community/prometheus-adapter/4.11.0) chart, whose other values are documented there. override prometheus-adapter values ??? note "Default" ```yaml dnsConfig: searches: - service.consul image: pullPolicy: IfNotPresent repository: docker.hops.works/prometheus/prometheus-adapter tag: 0.12.0-alpine-h1.1 nameOverride: hopsworks-prometheus-adapter nodeSelector: {} prometheus: path: '' port: 9090 url: http://prometheus.service.consul resources: limits: cpu: 100m memory: 70Mi requests: cpu: 50m memory: 40Mi rules: existing: custom-prometheus-metrics-adapter tolerations: [] topologySpreadConstraints: - labelSelector: matchLabels: app.kubernetes.io/instance: hopsworks app.kubernetes.io/name: hopsworks-prometheus-adapter maxSkew: 1 topologyKey: topology.kubernetes.io/zone whenUnsatisfiable: ScheduleAnyway ```
================================================================================ # ray Source: https://docs.hopsworks.ai/latest/setup_installation/common/helm_chart_values/ray/ # Ray values { #helm-values-ray } Values under `ray` configure the KubeRay operator, which runs the Ray clusters behind Ray jobs and notebooks. _Generated from the Hopsworks Helm chart `5.2.0-alpha-1791549041` (Hopsworks `5.2.0`)._ Deployed according to the first of these values that is set: [`global._hopsworks.ray.enabled`](global.md#helm.global._hopsworks.ray.enabled), [`global._hopsworks.full_platform`](global.md#helm.global._hopsworks.full_platform). !!! info "Upstream charts" - Values under `ray.kuberay-operator` go to [`kuberay-operator` 1.4.0](https://artifacthub.io/packages/helm/kuberay-operator/kuberay-operator/1.4.0) from `https://ray-project.github.io/kuberay-helm/`. Only the values Hopsworks sets under `ray.kuberay-operator` are listed on this page. Any other value of the chart can be set there too; the link opens its documentation for the version Hopsworks pins. ??? example "Defaults as YAML" ```yaml ray: kuberay-operator: batchScheduler: enabled: false name: '' crNamespacedRbacEnable: true env: null featureGates: - enabled: false name: RayClusterStatusConditions fullnameOverride: kuberay-operator image: pullPolicy: IfNotPresent repository: docker.hops.works/hopsworks/kuberay-operator tag: 1.4.0-h1 leaderElectionEnabled: true livenessProbe: failureThreshold: 5 initialDelaySeconds: 10 periodSeconds: 5 logging: baseDir: '' fileEncoder: '' fileName: '' stdoutEncoder: '' nameOverride: kuberay-operator podSecurityContext: {} rbacEnable: true readinessProbe: failureThreshold: 5 initialDelaySeconds: 10 periodSeconds: 5 resources: limits: cpu: 100m memory: 512Mi securityContext: allowPrivilegeEscalation: false capabilities: drop: - ALL readOnlyRootFilesystem: true runAsNonRoot: true seccompProfile: type: RuntimeDefault service: port: 8080 type: ClusterIP serviceAccount: create: true name: kuberay-operator singleNamespaceInstall: false ```
`ray` # { #helm.ray } : Type `object`, default `{}`. override ray values `ray.kuberay-operator` # { #helm.ray.kuberay-operator } : Type `object`, passed to the [`kuberay-operator` 1.4.0](https://artifacthub.io/packages/helm/kuberay-operator/kuberay-operator/1.4.0) chart, whose other values are documented there. override kuberay-operator values ??? note "Default" ```yaml batchScheduler: enabled: false name: '' crNamespacedRbacEnable: true env: null featureGates: - enabled: false name: RayClusterStatusConditions fullnameOverride: kuberay-operator image: pullPolicy: IfNotPresent repository: docker.hops.works/hopsworks/kuberay-operator tag: 1.4.0-h1 leaderElectionEnabled: true livenessProbe: failureThreshold: 5 initialDelaySeconds: 10 periodSeconds: 5 logging: baseDir: '' fileEncoder: '' fileName: '' stdoutEncoder: '' nameOverride: kuberay-operator podSecurityContext: {} rbacEnable: true readinessProbe: failureThreshold: 5 initialDelaySeconds: 10 periodSeconds: 5 resources: limits: cpu: 100m memory: 512Mi securityContext: allowPrivilegeEscalation: false capabilities: drop: - ALL readOnlyRootFilesystem: true runAsNonRoot: true seccompProfile: type: RuntimeDefault service: port: 8080 type: ClusterIP serviceAccount: create: true name: kuberay-operator singleNamespaceInstall: false ```
================================================================================ # rondb Source: https://docs.hopsworks.ai/latest/setup_installation/common/helm_chart_values/rondb/ # RonDB values { #helm-values-rondb } Values under `rondb` configure RonDB, the online feature store database, installed from the RonDB Helm chart. _Generated from the Hopsworks Helm chart `5.2.0-alpha-1791549041` (Hopsworks `5.2.0`)._ Always deployed. !!! info "Upstream charts" - Values under `rondb.rondb` go to [`rondb` 26.2.21](https://github.com/logicalclocks/rondb-helm/blob/v26.2.21/values.schema.json) from `https://logicalclocks.github.io/rondb-helm/`, and all of them are listed under [`rondb` chart values](#helm-values-rondb-rondb). ??? example "Defaults as YAML" ```yaml rondb: rondb: backups: enabled: null pathPrefix: rondb_backup clusterSize: activeDataReplicas: 2 maxNumMySQLServers: 1 maxNumRdrs: 2 minNumMySQLServers: 1 minNumRdrs: 1 numNodeGroups: 1 enableSecurityContext: true images: mysqldExporter: registry: docker.hops.works repository: hopsworks rondb: registry: docker.hops.works repository: hopsworks toolbox: name: hwutils registry: docker.hops.works repository: hopsworks tag: 1.10-SNAPSHOT meta: ddlMySQLd: clusterIp: annotations: consul.hashicorp.com/service-name: mysqlddl consul.hashicorp.com/service-tags: onlinefs enabled: true statefulSet: endToEndTls: enabled: false filenames: ca: hops_root_ca.pem cert: mysqlddl_certificate_bundle.pem key: mysqlddl_priv.pem secretName: mysqlddl-crypto-material supplyOwnSecret: true mgmd: headlessClusterIp: annotations: consul.hashicorp.com/service-name: mgmd mysqld: clusterIp: annotations: consul.hashicorp.com/service-name: mysql consul.hashicorp.com/service-tags: onlinefs externalLoadBalancer: annotations: {} class: null enabled: true name: mysqld-external statefulSet: endToEndTls: enabled: false filenames: ca: hops_root_ca.pem cert: mysql_certificate_bundle.pem key: mysql_priv.pem secretName: mysqld-crypto-material supplyOwnSecret: true rdrs: clusterIp: annotations: consul.hashicorp.com/service-name: rdrs prometheus.io/path: /metrics prometheus.io/port: '4406' prometheus.io/scheme: https prometheus.io/scrape: 'true' externalLoadBalancer: annotations: {} class: null enabled: true name: rdrs-external headlessClusterIpName: rdrs-cluster-ip ingress: enabled: false statefulSet: endToEndTls: enabled: true filenames: ca: hops_root_ca.pem cert: mysql_certificate_bundle.pem key: mysql_priv.pem secretName: rdrs-crypto-material supplyOwnSecret: true mysql: clusterUser: bench credentialsSecretName: mysql-users-secrets exporter: enabled: true users: - host: '%' privileges: - database: '*' privileges: - ALL table: '*' withGrantOption: true username: hopsworksroot networkPolicy: mgmds: ingressSelectors: - podSelector: matchLabels: access: mgmd-and-ndbmtd ndbmtds: ingressSelectors: - podSelector: matchLabels: access: mgmd-and-ndbmtd resources: requests: storage: classes: binlogFiles: null default: null diskColumns: null restoreFromBackup: backupId: null pathPrefix: rondb_backup serviceAccountAnnotations: {} ```
`rondb` # { #helm.rondb } : Type `object`. override rondb values ??? note "Default" ```yaml rondb: clusterSize: activeDataReplicas: 2 maxNumMySQLServers: 1 maxNumRdrs: 2 minNumMySQLServers: 1 minNumRdrs: 1 numNodeGroups: 1 enableSecurityContext: true images: mysqldExporter: registry: docker.hops.works rondb: registry: docker.hops.works toolbox: name: hwutils registry: docker.hops.works tag: 1.10-SNAPSHOT meta: mysqld: externalLoadBalancer: annotations: {} class: null enabled: true name: mysqld-external rdrs: externalLoadBalancer: annotations: {} class: null enabled: true name: rdrs-external statefulSet: endToEndTls: enabled: true mysql: credentialsSecretName: mysql-users-secrets exporter: enabled: true users: - host: '%' privileges: - database: '*' privileges: - ALL table: '*' withGrantOption: true username: hopsworksroot networkPolicy: mgmds: ingressSelectors: - podSelector: matchLabels: access: mgmd-and-ndbmtd ndbmtds: ingressSelectors: - podSelector: matchLabels: access: mgmd-and-ndbmtd resources: requests: storage: classes: binlogFiles: null default: null diskColumns: null serviceAccountAnnotations: {} ``` `rondb.rondb` # { #helm.rondb.rondb } : Type `object`, passed to the [`rondb` 26.2.21](https://github.com/logicalclocks/rondb-helm/blob/v26.2.21/values.schema.json) chart, whose values are listed under [`rondb` chart values](#helm-values-rondb-rondb). override rondb values ??? note "Default" ```yaml meta: mgmd: headlessClusterIp: annotations: consul.hashicorp.com/service-name: mgmd mysqld: statefulSet: endToEndTls: enabled: false secretName: mysqld-crypto-material supplyOwnSecret: true filenames: ca: hops_root_ca.pem cert: mysql_certificate_bundle.pem key: mysql_priv.pem clusterIp: annotations: consul.hashicorp.com/service-name: mysql consul.hashicorp.com/service-tags: onlinefs externalLoadBalancer: name: mysqld-external enabled: true class: null annotations: {} ddlMySQLd: enabled: true statefulSet: endToEndTls: enabled: false secretName: mysqlddl-crypto-material supplyOwnSecret: true filenames: ca: hops_root_ca.pem cert: mysqlddl_certificate_bundle.pem key: mysqlddl_priv.pem clusterIp: annotations: consul.hashicorp.com/service-name: mysqlddl consul.hashicorp.com/service-tags: onlinefs rdrs: statefulSet: endToEndTls: enabled: true secretName: rdrs-crypto-material supplyOwnSecret: true filenames: ca: hops_root_ca.pem cert: mysql_certificate_bundle.pem key: mysql_priv.pem clusterIp: annotations: consul.hashicorp.com/service-name: rdrs prometheus.io/path: /metrics prometheus.io/port: '4406' prometheus.io/scheme: https prometheus.io/scrape: 'true' headlessClusterIpName: rdrs-cluster-ip ingress: enabled: false externalLoadBalancer: name: rdrs-external enabled: true class: null annotations: {} images: rondb: registry: docker.hops.works repository: hopsworks toolbox: registry: docker.hops.works repository: hopsworks name: hwutils tag: 1.10-SNAPSHOT mysqldExporter: registry: docker.hops.works repository: hopsworks mysql: clusterUser: bench exporter: enabled: true users: - username: hopsworksroot host: '%' privileges: - database: '*' table: '*' withGrantOption: true privileges: - ALL credentialsSecretName: mysql-users-secrets clusterSize: activeDataReplicas: 2 numNodeGroups: 1 minNumMySQLServers: 1 maxNumMySQLServers: 1 minNumRdrs: 1 maxNumRdrs: 2 restoreFromBackup: backupId: null pathPrefix: rondb_backup backups: enabled: null pathPrefix: rondb_backup enableSecurityContext: true serviceAccountAnnotations: {} resources: requests: storage: classes: default: null diskColumns: null binlogFiles: null networkPolicy: mgmds: ingressSelectors: - podSelector: matchLabels: access: mgmd-and-ndbmtd ndbmtds: ingressSelectors: - podSelector: matchLabels: access: mgmd-and-ndbmtd ```
## `rondb` chart values { #helm-values-rondb-rondb } These are the values of the [`rondb` 26.2.21](https://github.com/logicalclocks/rondb-helm/blob/v26.2.21/values.schema.json) chart, set under `rondb.rondb`. The defaults are what Hopsworks deploys: the chart's own, with the `rondb` and `rondb.rondb` overrides above applied. Where Hopsworks overrides a value, the entry also gives the chart's own default. ### General { #helm-values-rondb-rondb-general } ??? example "Defaults as YAML" ```yaml rondb: rondb: createPriorityClass: true enableSecurityContext: true forceNodeGroupChange: false imagePullPolicy: IfNotPresent imagePullSecrets: [] isMultiNodeCluster: true mgmdPreUpgradeRolloutTimeout: 5m mode: '' priorityClass: rondb-high-priority serviceAccountAnnotations: {} skipMgmdPreUpgradeRollout: false staticCpuManagerPolicy: false tls: caSecretName: null ```
`rondb.rondb.createPriorityClass` # { #helm.rondb.rondb.createPriorityClass } : Type `boolean`, default `true`. Flag to disable the creation of PriorityClass in case it was created beforehand `rondb.rondb.enableSecurityContext` # { #helm.rondb.rondb.enableSecurityContext } : Type `boolean`, default `true`. `rondb.rondb.forceNodeGroupChange` # { #helm.rondb.rondb.forceNodeGroupChange } : Type `boolean`, default `false`. Skip the pre-upgrade immutability gate on clusterSize.numNodeGroups. RonDB still cannot add or remove node groups online, so an upgrade that changes numNodeGroups under this flag will break the cluster. `rondb.rondb.imagePullPolicy` # { #helm.rondb.rondb.imagePullPolicy } : Type `string`, default `"IfNotPresent"`. The Kubernetes image pull policy `rondb.rondb.imagePullSecrets` # { #helm.rondb.rondb.imagePullSecrets } : Type `array`, default `[]`. `rondb.rondb.imagePullSecrets[].name` # { #helm.rondb.rondb.imagePullSecrets.name } : Type `string`. The name of the Secret `rondb.rondb.isMultiNodeCluster` # { #helm.rondb.rondb.isMultiNodeCluster } : Type `boolean`, default `true`. Whether the Kubernetes cluster has multiple nodes; this will affect affinities and add requirements to scheduling. `rondb.rondb.mgmdPreUpgradeRolloutTimeout` # { #helm.rondb.rondb.mgmdPreUpgradeRolloutTimeout } : Type `string`, default `"5m"`. Timeout for `kubectl rollout status` inside the MGMd pre-upgrade hook (per attempt). Default 5m. Note: Helm's default `helm upgrade --timeout` is also 5m and bounds the whole upgrade including this hook — operators running with that default may see Helm's outer timeout fire first with a generic 'timed out waiting for the condition' message instead of this hook's specific diagnostic. Set Helm's `--timeout` higher (e.g. 15m) on slow networks or large image pulls. `rondb.rondb.mode` # { #helm.rondb.rondb.mode } : Type `enum`, default `""`. Mode of operation for the Helmchart. 'install' will always treat the deployment as a new installation, 'upgrade' will always treat it as an upgrade. 'auto' will treat it as an install if no previous release is found, otherwise as an upgrade. This flag is unset by default. When specified, it takes precedence over the value defined in .Values.global._hopsworks.mode One of: `"install"`, `"upgrade"`, `"auto"`, `""`. `rondb.rondb.priorityClass` # { #helm.rondb.rondb.priorityClass } : Type `string`, default `"rondb-high-priority"`. `rondb.rondb.serviceAccountAnnotations` # { #helm.rondb.rondb.serviceAccountAnnotations } : Type `object`, default `{}`. `rondb.rondb.skipMgmdPreUpgradeRollout` # { #helm.rondb.rondb.skipMgmdPreUpgradeRollout } : Type `boolean`, default `false`. If true, skip the pre-upgrade hook that rolls MGMd to the target image and config before the rest of the chart upgrades. The hook exists to keep at least one ArbitrationRank=1 node reachable while the API tier (mysqld, rdrs) is also rolling, preventing data-node arbitration loss (NDB Error 2305) on chart-wide upgrades. Set this to true only as an escape hatch if the hook itself is misbehaving on your cluster. `rondb.rondb.staticCpuManagerPolicy` # { #helm.rondb.rondb.staticCpuManagerPolicy } : Type `boolean`, default `false`. Whether the Kubernetes cluster has been configured with a static CPU manager policy. This is an optimization for RonDB data nodes. These have a scheduler which executes jobs within hundreds of microseconds (very quick). If the data nodes are pinned to CPUs, they can run CPU spnning to avoid context switching inbetween jobs. `rondb.rondb.tls` # { #helm.rondb.rondb.tls } : Type `object`. General settings how to use encrypted TLS connections with RonDB APIs (MySQL & RDRS). `rondb.rondb.tls.caSecretName` # { #helm.rondb.rondb.tls.caSecretName } : Type `string|null`, default `null`. User-provided TLS Secret. This can be used by the cert-manager as a base to sign further certificates. If not supplied, cert-manager will use a self-signed CA certificate.
### backups { #helm-values-rondb-rondb-backups } ??? example "Defaults as YAML" ```yaml rondb: rondb: backups: enabled: null metadataConfigmapName: null objectStorageProvider: s3 pathPrefix: rondb_backup s3: bucketName: null endpoint: null keyCredentialsSecret: key: null name: null provider: null region: null secretCredentialsSecret: key: null name: null serverSideEncryption: null schedule: null ttl: null ```
`rondb.rondb.backups` # { #helm.rondb.rondb.backups } : Type `object`. Whether, how and how often to run regular backups on the cluster `rondb.rondb.backups.enabled` # { #helm.rondb.rondb.backups.enabled } : Type `boolean|null`, default `null`. `rondb.rondb.backups.metadataConfigmapName` # { #helm.rondb.rondb.backups.metadataConfigmapName } : Type `string|null`, default `null`. The name of the configmap to be used to store the backups metadata information. `rondb.rondb.backups.objectStorageProvider` # { #helm.rondb.rondb.backups.objectStorageProvider } : Type `enum`, default `"s3"`. One of: `"s3"`. `rondb.rondb.backups.pathPrefix` # { #helm.rondb.rondb.backups.pathPrefix } : Type `string`, default `"rondb_backup"`. Prefix of RonDB backup in the configured bucket `rondb.rondb.backups.s3` # { #helm.rondb.rondb.backups.s3 } : Type `object`. `rondb.rondb.backups.s3.bucketName` # { #helm.rondb.rondb.backups.s3.bucketName } : Type `string|null`, default `null`. `rondb.rondb.backups.s3.endpoint` # { #helm.rondb.rondb.backups.s3.endpoint } : Type `string|null`, default `null`. `rondb.rondb.backups.s3.keyCredentialsSecret` # { #helm.rondb.rondb.backups.s3.keyCredentialsSecret } : Type `object`. `rondb.rondb.backups.s3.keyCredentialsSecret.key` # { #helm.rondb.rondb.backups.s3.keyCredentialsSecret.key } : Type `string|null`, default `null`. Key in the Secret `rondb.rondb.backups.s3.keyCredentialsSecret.name` # { #helm.rondb.rondb.backups.s3.keyCredentialsSecret.name } : Type `string|null`, default `null`. Name of the Secret `rondb.rondb.backups.s3.provider` # { #helm.rondb.rondb.backups.s3.provider } : Type `string|null`, default `null`. `rondb.rondb.backups.s3.region` # { #helm.rondb.rondb.backups.s3.region } : Type `string|null`, default `null`. `rondb.rondb.backups.s3.secretCredentialsSecret` # { #helm.rondb.rondb.backups.s3.secretCredentialsSecret } : Type `object`. `rondb.rondb.backups.s3.secretCredentialsSecret.key` # { #helm.rondb.rondb.backups.s3.secretCredentialsSecret.key } : Type `string|null`, default `null`. Key in the Secret `rondb.rondb.backups.s3.secretCredentialsSecret.name` # { #helm.rondb.rondb.backups.s3.secretCredentialsSecret.name } : Type `string|null`, default `null`. Name of the Secret `rondb.rondb.backups.s3.serverSideEncryption` # { #helm.rondb.rondb.backups.s3.serverSideEncryption } : Type `enum`, default `null`. One of: `"aws:kms"`, `"aws:kms:dsse"`, `"AES256"`, `null`. `rondb.rondb.backups.schedule` # { #helm.rondb.rondb.backups.schedule } : Type `string|null`, default `null`. Cron schedule for backups `rondb.rondb.backups.ttl` # { #helm.rondb.rondb.backups.ttl } : Type `string|null`, default `null`, pattern `^\d+[dh]$`. time to live to control when to clean up backups. It is a number followed by either d (days) or h (hours) suffix.
### benchmarking { #helm-values-rondb-rondb-benchmarking } ??? example "Defaults as YAML" ```yaml rondb: rondb: benchmarking: dbt2: numWarehouses: 4 runMulti: |- # NUM_MYSQL_SERVERS NUM_WAREHOUSES NUM_TERMINALS 2 1 1 2 2 1 2 2 2 runSingle: |- # NUM_MYSQL_SERVERS NUM_WAREHOUSES NUM_TERMINALS 1 1 1 1 2 1 1 4 1 1 4 2 enabled: false sysbench: minimizeBandwidth: false rows: 100000 threadCountsToRun: 1;2;4;8;12;16;24;32;64 type: sysbench ycsb: schemata: CREATE TABLE IF NOT EXISTS ycsb.usertable (YCSB_KEY VARCHAR(255) PRIMARY KEY, FIELD0 varbinary(4096)); ```
`rondb.rondb.benchmarking` # { #helm.rondb.rondb.benchmarking } : Type `object`. Whether, which and how to run a benchmarking job on the cluster `rondb.rondb.benchmarking.dbt2` # { #helm.rondb.rondb.benchmarking.dbt2 } : Type `object`. Configuration of DBT2 benchmarking job `rondb.rondb.benchmarking.dbt2.numWarehouses` # { #helm.rondb.rondb.benchmarking.dbt2.numWarehouses } : Type `integer`, default `4`, minimum `1`. `rondb.rondb.benchmarking.dbt2.runMulti` # { #helm.rondb.rondb.benchmarking.dbt2.runMulti } : Type `string`, default `"# NUM_MYSQL_SERVERS NUM_WAREHOUSES NUM_TERMINALS\n2 1 1\n2 2 1\n2 2 2"`. Table with columns: NUM_MYSQL_SERVERS, NUM_WAREHOUSES, NUM_TERMINALS `rondb.rondb.benchmarking.dbt2.runSingle` # { #helm.rondb.rondb.benchmarking.dbt2.runSingle } : Type `string`. Table with columns: NUM_MYSQL_SERVERS, NUM_WAREHOUSES, NUM_TERMINALS ??? note "Default" ```yaml |- # NUM_MYSQL_SERVERS NUM_WAREHOUSES NUM_TERMINALS 1 1 1 1 2 1 1 4 1 1 4 2 ``` `rondb.rondb.benchmarking.enabled` # { #helm.rondb.rondb.benchmarking.enabled } : Type `boolean`, default `false`. Whether to run a benchmarking job on the cluster `rondb.rondb.benchmarking.sysbench` # { #helm.rondb.rondb.benchmarking.sysbench } : Type `object`. Configuration of Sysbench benchmarking job `rondb.rondb.benchmarking.sysbench.minimizeBandwidth` # { #helm.rondb.rondb.benchmarking.sysbench.minimizeBandwidth } : Type `boolean`, default `false`. Whether to use filters to minimize bandwidth usage. This can be useful in cloud environments where bandwidth is expensive. `rondb.rondb.benchmarking.sysbench.rows` # { #helm.rondb.rondb.benchmarking.sysbench.rows } : Type `integer`, default `100000`. `rondb.rondb.benchmarking.sysbench.threadCountsToRun` # { #helm.rondb.rondb.benchmarking.sysbench.threadCountsToRun } : Type `string`, default `"1;2;4;8;12;16;24;32;64"`. Semi-colon-separated list of thread counts to run `rondb.rondb.benchmarking.type` # { #helm.rondb.rondb.benchmarking.type } : Type `enum`, default `"sysbench"`. Which benchmarking job to run. 'Multi' refers to multiple MySQLd servers to run against One of: `"sysbench"`, `"dbt2_single"`, `"dbt2_multi"`, `"ycsb"`. `rondb.rondb.benchmarking.ycsb` # { #helm.rondb.rondb.benchmarking.ycsb } : Type `object`. Configuration of YCSB benchmarking job `rondb.rondb.benchmarking.ycsb.schemata` # { #helm.rondb.rondb.benchmarking.ycsb.schemata } : Type `string`. MySQL table schema to use for YCSB. NOTE: The `ycsb` database is pre-created. ??? note "Default" ```yaml CREATE TABLE IF NOT EXISTS ycsb.usertable (YCSB_KEY VARCHAR(255) PRIMARY KEY, FIELD0 varbinary(4096)); ```
### clusterSize { #helm-values-rondb-rondb-clustersize } ??? example "Defaults as YAML" ```yaml rondb: rondb: clusterSize: activeDataReplicas: 2 maxNumMySQLServers: 1 maxNumRdrs: 2 minNumMySQLServers: 1 minNumRdrs: 1 numNodeGroups: 1 ```
`rondb.rondb.clusterSize` # { #helm.rondb.rondb.clusterSize } : Type `object`. Horizontal cluster size `rondb.rondb.clusterSize.activeDataReplicas` # { #helm.rondb.rondb.clusterSize.activeDataReplicas } : Type `integer`, default `2`, minimum `1`, maximum `3`. How many replicas each node group has. When descreasing this value, try going down by 1 at a time. Too many nodes leaving at once can cause issues with leader elections. `rondb.rondb.clusterSize.maxNumMySQLServers` # { #helm.rondb.rondb.clusterSize.maxNumMySQLServers } : Type `integer`, default `1`, Hopsworks overrides the `rondb` chart default `5`, minimum `1`. The maximum amount of MySQL servers we will auto-scale to. It is represented as number of slots in RonDB's config.ini. Changing this will only take effect if the ConfigMap for config.ini is updated and the MGMd is restarted. `rondb.rondb.clusterSize.maxNumRdrs` # { #helm.rondb.rondb.clusterSize.maxNumRdrs } : Type `integer`, default `2`, minimum `0`. `rondb.rondb.clusterSize.minNumMySQLServers` # { #helm.rondb.rondb.clusterSize.minNumMySQLServers } : Type `integer`, default `1`, minimum `1`. A minimum amount of MySQL servers `rondb.rondb.clusterSize.minNumRdrs` # { #helm.rondb.rondb.clusterSize.minNumRdrs } : Type `integer`, default `1`, minimum `0`. `rondb.rondb.clusterSize.numNodeGroups` # { #helm.rondb.rondb.clusterSize.numNodeGroups } : Type `integer`, default `1`, minimum `1`. Number of RonDB node groups. Immutable after install — RonDB has no online add/remove. Changing it requires uninstall and reinstall with restoreFromBackup.backupId. Set `mode: upgrade` for ArgoCD or other template-only renderers so the immutability check runs as an in-cluster Job.
### globalReplication { #helm-values-rondb-rondb-globalreplication } ??? example "Defaults as YAML" ```yaml rondb: rondb: globalReplication: clusterNumber: 1 primary: binlogFilename: binlog enabled: false expireBinlogsDays: 1.5 ignoreDatabases: [] includeDatabases: [] logReplicaUpdates: false maxNumBinlogServers: 2 numBinlogServers: 2 secondary: enabled: false replicateFrom: binlogServerHosts: [] clusterNumber: 2 ignoreDatabases: [] ignoreTables: [] includeDatabases: [] includeTables: [] useTlsConnection: false ```
`rondb.rondb.globalReplication` # { #helm.rondb.rondb.globalReplication } : Type `object`. `rondb.rondb.globalReplication.clusterNumber` # { #helm.rondb.rondb.globalReplication.clusterNumber } : Type `number`, default `1`. Determines the offset for global server IDs `rondb.rondb.globalReplication.primary` # { #helm.rondb.rondb.globalReplication.primary } : Type `object`. `rondb.rondb.globalReplication.primary.binlogFilename` # { #helm.rondb.rondb.globalReplication.primary.binlogFilename } : Type `string`, default `"binlog"`. `rondb.rondb.globalReplication.primary.enabled` # { #helm.rondb.rondb.globalReplication.primary.enabled } : Type `boolean`, default `false`. Specifies if the primary replication is enabled. This will actually create the binary log servers. Can be activated after an initial start. `rondb.rondb.globalReplication.primary.expireBinlogsDays` # { #helm.rondb.rondb.globalReplication.primary.expireBinlogsDays } : Type `number`, default `1.5`. `rondb.rondb.globalReplication.primary.ignoreDatabases` # { #helm.rondb.rondb.globalReplication.primary.ignoreDatabases } : Type `array`, default `[]`. `rondb.rondb.globalReplication.primary.includeDatabases` # { #helm.rondb.rondb.globalReplication.primary.includeDatabases } : Type `array`, default `[]`. `rondb.rondb.globalReplication.primary.logReplicaUpdates` # { #helm.rondb.rondb.globalReplication.primary.logReplicaUpdates } : Type `boolean`, default `false`. `rondb.rondb.globalReplication.primary.maxNumBinlogServers` # { #helm.rondb.rondb.globalReplication.primary.maxNumBinlogServers } : Type `number`, default `2`. Don't change this value. Determines how many binlog servers we can scale out to. Used to determine the server IDs for global replication. Even if replication is disabled, potential binlog servers will be written into the config.ini. `rondb.rondb.globalReplication.primary.numBinlogServers` # { #helm.rondb.rondb.globalReplication.primary.numBinlogServers } : Type `number`, default `2`. The current number of binlog servers, given global Replication is enabled. `rondb.rondb.globalReplication.secondary` # { #helm.rondb.rondb.globalReplication.secondary } : Type `object`. `rondb.rondb.globalReplication.secondary.enabled` # { #helm.rondb.rondb.globalReplication.secondary.enabled } : Type `boolean`, default `false`. `rondb.rondb.globalReplication.secondary.replicateFrom` # { #helm.rondb.rondb.globalReplication.secondary.replicateFrom } : Type `object`. RonDB clusters to replicate from. RonDB supports merge-replicating from multiple clusters, but we only support replicating from one cluster `rondb.rondb.globalReplication.secondary.replicateFrom.binlogServerHosts` # { #helm.rondb.rondb.globalReplication.secondary.replicateFrom.binlogServerHosts } : Type `array`, default `[]`. Hostnames of the binlog server to replicate from. The binlog servers have headless ClusterIPs and one LoadBalancer *per* binlog server. If we replicate across different K8s clusters, we can reference the External IPs of the binlog servers' LoadBalancers here. `rondb.rondb.globalReplication.secondary.replicateFrom.clusterNumber` # { #helm.rondb.rondb.globalReplication.secondary.replicateFrom.clusterNumber } : Type `number`, default `2`. A helper to understand from which serverIds to replicate from `rondb.rondb.globalReplication.secondary.replicateFrom.ignoreDatabases` # { #helm.rondb.rondb.globalReplication.secondary.replicateFrom.ignoreDatabases } : Type `array`, default `[]`. `rondb.rondb.globalReplication.secondary.replicateFrom.ignoreTables` # { #helm.rondb.rondb.globalReplication.secondary.replicateFrom.ignoreTables } : Type `array`, default `[]`. `rondb.rondb.globalReplication.secondary.replicateFrom.includeDatabases` # { #helm.rondb.rondb.globalReplication.secondary.replicateFrom.includeDatabases } : Type `array`, default `[]`. `rondb.rondb.globalReplication.secondary.replicateFrom.includeTables` # { #helm.rondb.rondb.globalReplication.secondary.replicateFrom.includeTables } : Type `array`, default `[]`. `rondb.rondb.globalReplication.secondary.replicateFrom.useTlsConnection` # { #helm.rondb.rondb.globalReplication.secondary.replicateFrom.useTlsConnection } : Type `boolean`, default `false`. Enable this for an encrypted replication channel. Important if the binlog server requires TLS connections
### images { #helm-values-rondb-rondb-images } ??? example "Defaults as YAML" ```yaml rondb: rondb: images: dataValidation: name: python registry: docker.io repository: '' tag: 3.12-slim mysqldExporter: name: mysqld_exporter registry: docker.hops.works repository: hopsworks tag: 0.11.5 rondb: name: rondb registry: docker.hops.works repository: hopsworks tag: 26.02.11 toolbox: name: hwutils registry: docker.hops.works repository: hopsworks tag: 1.10-SNAPSHOT upgrade2410revokegrants: name: rondb registry: docker.io repository: hopsworks tag: 22.10.13-0.7 ```
`rondb.rondb.images` # { #helm.rondb.rondb.images } : Type `object`. Information for Docker images used in the cluster `rondb.rondb.images.dataValidation` # { #helm.rondb.rondb.images.dataValidation } : Type `object`. `rondb.rondb.images.dataValidation.name` # { #helm.rondb.rondb.images.dataValidation.name } : Type `string`, default `"python"`. `rondb.rondb.images.dataValidation.registry` # { #helm.rondb.rondb.images.dataValidation.registry } : Type `string`, default `"docker.io"`. `rondb.rondb.images.dataValidation.repository` # { #helm.rondb.rondb.images.dataValidation.repository } : Type `string`, default `""`. `rondb.rondb.images.dataValidation.tag` # { #helm.rondb.rondb.images.dataValidation.tag } : Type `string`, default `"3.12-slim"`. `rondb.rondb.images.mysqldExporter` # { #helm.rondb.rondb.images.mysqldExporter } : Type `object`. `rondb.rondb.images.mysqldExporter.name` # { #helm.rondb.rondb.images.mysqldExporter.name } : Type `string`, default `"mysqld_exporter"`. `rondb.rondb.images.mysqldExporter.registry` # { #helm.rondb.rondb.images.mysqldExporter.registry } : Type `string`, default `"docker.hops.works"`. `rondb.rondb.images.mysqldExporter.repository` # { #helm.rondb.rondb.images.mysqldExporter.repository } : Type `string`, default `"hopsworks"`. `rondb.rondb.images.mysqldExporter.tag` # { #helm.rondb.rondb.images.mysqldExporter.tag } : Type `string`, default `"0.11.5"`. `rondb.rondb.images.rondb` # { #helm.rondb.rondb.images.rondb } : Type `object`. `rondb.rondb.images.rondb.name` # { #helm.rondb.rondb.images.rondb.name } : Type `string`, default `"rondb"`. `rondb.rondb.images.rondb.registry` # { #helm.rondb.rondb.images.rondb.registry } : Type `string`, default `"docker.hops.works"`, Hopsworks overrides the `rondb` chart default `"docker.io"`. `rondb.rondb.images.rondb.repository` # { #helm.rondb.rondb.images.rondb.repository } : Type `string`, default `"hopsworks"`. `rondb.rondb.images.rondb.tag` # { #helm.rondb.rondb.images.rondb.tag } : Type `string`, default `"26.02.11"`. The version of RonDB to use; This should always be equivalent to .Chart.AppVersion `rondb.rondb.images.toolbox` # { #helm.rondb.rondb.images.toolbox } : Type `object`. `rondb.rondb.images.toolbox.name` # { #helm.rondb.rondb.images.toolbox.name } : Type `string`, default `"hwutils"`. `rondb.rondb.images.toolbox.registry` # { #helm.rondb.rondb.images.toolbox.registry } : Type `string`, default `"docker.hops.works"`, Hopsworks overrides the `rondb` chart default `"docker.io"`. `rondb.rondb.images.toolbox.repository` # { #helm.rondb.rondb.images.toolbox.repository } : Type `string`, default `"hopsworks"`. `rondb.rondb.images.toolbox.tag` # { #helm.rondb.rondb.images.toolbox.tag } : Type `string`, default `"1.10-SNAPSHOT"`, Hopsworks overrides the `rondb` chart default `"1.7"`. `rondb.rondb.images.upgrade2410revokegrants` # { #helm.rondb.rondb.images.upgrade2410revokegrants } : Type `object`. `rondb.rondb.images.upgrade2410revokegrants.name` # { #helm.rondb.rondb.images.upgrade2410revokegrants.name } : Type `string`, default `"rondb"`. `rondb.rondb.images.upgrade2410revokegrants.registry` # { #helm.rondb.rondb.images.upgrade2410revokegrants.registry } : Type `string`, default `"docker.io"`. `rondb.rondb.images.upgrade2410revokegrants.repository` # { #helm.rondb.rondb.images.upgrade2410revokegrants.repository } : Type `string`, default `"hopsworks"`. `rondb.rondb.images.upgrade2410revokegrants.tag` # { #helm.rondb.rondb.images.upgrade2410revokegrants.tag } : Type `string`, default `"22.10.13-0.7"`. The version of RonDB to use when revoking bogus user privilege to upgrade from 22.10 to 24.10
### meta { #helm-values-rondb-rondb-meta } ??? example "Defaults as YAML" ```yaml rondb: rondb: meta: binlogServers: externalLoadBalancers: annotations: {} class: null enabled: true namePrefix: binlog-server port: 3306 headlessClusterIp: name: headless-binlog-servers port: 3306 statefulSet: endToEndTls: enabled: false filenames: ca: null cert: tls.crt key: tls.key secretName: binlog-end-to-end-tls supplyOwnSecret: false name: mysqld-binlog-servers ddlMySQLd: addSysNiceCapability: true clusterIp: annotations: consul.hashicorp.com/service-name: mysqlddl consul.hashicorp.com/service-tags: onlinefs name: ddl-mysqld port: 3306 enabled: true headlessClusterIp: name: headless-ddl-mysqld port: 3306 statefulSet: endToEndTls: enabled: false filenames: ca: hops_root_ca.pem cert: mysqlddl_certificate_bundle.pem key: mysqlddl_priv.pem secretName: mysqlddl-crypto-material supplyOwnSecret: true name: ddl-mysqld mgmd: headlessClusterIp: annotations: consul.hashicorp.com/service-name: mgmd name: headless-mgmds port: 1186 statefulSetName: mgmds mysqld: addSysNiceCapability: true clusterIp: annotations: consul.hashicorp.com/service-name: mysql consul.hashicorp.com/service-tags: onlinefs name: mysqld port: 3306 exporter: metricsPort: 9104 externalLoadBalancer: annotations: {} class: null enabled: true managed: true name: mysqld-external nodePort: null nodeSelector: {} port: 3306 headlessClusterIp: name: headless-mysqlds port: 3306 statefulSet: endToEndTls: enabled: false filenames: ca: hops_root_ca.pem cert: mysql_certificate_bundle.pem key: mysql_priv.pem secretName: mysqld-crypto-material supplyOwnSecret: true name: mysqlds ndbmtd: statefulSet: podAnnotations: {} rdrs: clusterIp: annotations: consul.hashicorp.com/service-name: rdrs prometheus.io/path: /metrics prometheus.io/port: '4406' prometheus.io/scheme: https prometheus.io/scrape: 'true' name: rdrs externalLoadBalancer: annotations: {} class: null enabled: true managed: true name: rdrs-external nodePort: null nodeSelector: {} headlessClusterIpName: rdrs-cluster-ip ingress: class: nginx dnsNames: [] enabled: false tls: enabled: true ipAddresses: [] useDefaultBackend: true statefulSet: endToEndTls: enabled: true filenames: ca: hops_root_ca.pem cert: mysql_certificate_bundle.pem key: mysql_priv.pem secretName: rdrs-crypto-material supplyOwnSecret: true name: rdrs replicaAppliers: headlessClusterIp: name: headless-replica-appliers port: 3306 statefulSet: endToEndTls: enabled: false filenames: ca: null cert: tls.crt key: tls.key secretName: replica-applier-end-to-end-tls supplyOwnSecret: false name: mysqld-replica-appliers ```
`rondb.rondb.meta` # { #helm.rondb.rondb.meta } : Type `object`. `rondb.rondb.meta.binlogServers` # { #helm.rondb.rondb.meta.binlogServers } : Type `object`. `rondb.rondb.meta.binlogServers.externalLoadBalancers` # { #helm.rondb.rondb.meta.binlogServers.externalLoadBalancers } : Type `object`. One LoadBalancer per binlog server; Creating an (alternative) Ingress per binlog server is difficult `rondb.rondb.meta.binlogServers.externalLoadBalancers.annotations` # { #helm.rondb.rondb.meta.binlogServers.externalLoadBalancers.annotations } : Type `object`, default `{}`. Cloud provider load balancer specific annotations. `rondb.rondb.meta.binlogServers.externalLoadBalancers.class` # { #helm.rondb.rondb.meta.binlogServers.externalLoadBalancers.class } : Type `string|null`, default `null`. `rondb.rondb.meta.binlogServers.externalLoadBalancers.enabled` # { #helm.rondb.rondb.meta.binlogServers.externalLoadBalancers.enabled } : Type `boolean`, default `true`. `rondb.rondb.meta.binlogServers.externalLoadBalancers.namePrefix` # { #helm.rondb.rondb.meta.binlogServers.externalLoadBalancers.namePrefix } : Type `string`, default `"binlog-server"`. `rondb.rondb.meta.binlogServers.externalLoadBalancers.port` # { #helm.rondb.rondb.meta.binlogServers.externalLoadBalancers.port } : Type `integer`, default `3306`. Port to expose the service on `rondb.rondb.meta.binlogServers.headlessClusterIp` # { #helm.rondb.rondb.meta.binlogServers.headlessClusterIp } : Type `object`. `rondb.rondb.meta.binlogServers.headlessClusterIp.name` # { #helm.rondb.rondb.meta.binlogServers.headlessClusterIp.name } : Type `string`, default `"headless-binlog-servers"`. `rondb.rondb.meta.binlogServers.headlessClusterIp.port` # { #helm.rondb.rondb.meta.binlogServers.headlessClusterIp.port } : Type `integer`, default `3306`. `rondb.rondb.meta.binlogServers.statefulSet` # { #helm.rondb.rondb.meta.binlogServers.statefulSet } : Type `object`. `rondb.rondb.meta.binlogServers.statefulSet.endToEndTls` # { #helm.rondb.rondb.meta.binlogServers.statefulSet.endToEndTls } : Type `object`. `rondb.rondb.meta.binlogServers.statefulSet.endToEndTls.enabled` # { #helm.rondb.rondb.meta.binlogServers.statefulSet.endToEndTls.enabled } : default `false`. Whether to use end-to-end encryption for MySQL binlog server Pods. This is recommended for high-security use cases. For MySQLds this is especially recommended since they use raw TCP connections and thereby by-pass TLS Ingress-rules. Otherwise, Ingress-TCP can also be configured directly via the Ingress controller (not in this Helmchart). `rondb.rondb.meta.binlogServers.statefulSet.endToEndTls.filenames` # { #helm.rondb.rondb.meta.binlogServers.statefulSet.endToEndTls.filenames } : Type `object`. `rondb.rondb.meta.binlogServers.statefulSet.endToEndTls.filenames.ca` # { #helm.rondb.rondb.meta.binlogServers.statefulSet.endToEndTls.filenames.ca } : Type `string|null`, default `null`. Name of the CA file in the Secret. ONLY overwrite this if you're not using standard TLS Secrets. `rondb.rondb.meta.binlogServers.statefulSet.endToEndTls.filenames.cert` # { #helm.rondb.rondb.meta.binlogServers.statefulSet.endToEndTls.filenames.cert } : Type `string`, default `"tls.crt"`. Name of the certificate file in the Secret. ONLY overwrite this if you're not using standard TLS Secrets. `rondb.rondb.meta.binlogServers.statefulSet.endToEndTls.filenames.key` # { #helm.rondb.rondb.meta.binlogServers.statefulSet.endToEndTls.filenames.key } : Type `string`, default `"tls.key"`. Name of the key file in the Secret. ONLY overwrite this if you're not using standard TLS Secrets. `rondb.rondb.meta.binlogServers.statefulSet.endToEndTls.secretName` # { #helm.rondb.rondb.meta.binlogServers.statefulSet.endToEndTls.secretName } : Type `string`, default `"binlog-end-to-end-tls"`. Name of the TLS Secret `rondb.rondb.meta.binlogServers.statefulSet.endToEndTls.supplyOwnSecret` # { #helm.rondb.rondb.meta.binlogServers.statefulSet.endToEndTls.supplyOwnSecret } : default `false`. Whether the Helmchart user will create a TLS Secret outside of this Helmchart. Otherwise we rely on cert-manager to create one. `rondb.rondb.meta.binlogServers.statefulSet.name` # { #helm.rondb.rondb.meta.binlogServers.statefulSet.name } : Type `string`, default `"mysqld-binlog-servers"`. `rondb.rondb.meta.ddlMySQLd` # { #helm.rondb.rondb.meta.ddlMySQLd } : Type `object`. `rondb.rondb.meta.ddlMySQLd.addSysNiceCapability` # { #helm.rondb.rondb.meta.ddlMySQLd.addSysNiceCapability } : Type `boolean`, default `true`. `rondb.rondb.meta.ddlMySQLd.clusterIp` # { #helm.rondb.rondb.meta.ddlMySQLd.clusterIp } : Type `object`. `rondb.rondb.meta.ddlMySQLd.clusterIp.annotations` # { #helm.rondb.rondb.meta.ddlMySQLd.clusterIp.annotations } : Type `object`, Hopsworks overrides the `rondb` chart default `{}`. ??? note "Default" ```yaml consul.hashicorp.com/service-name: mysqlddl consul.hashicorp.com/service-tags: onlinefs ``` `rondb.rondb.meta.ddlMySQLd.clusterIp.name` # { #helm.rondb.rondb.meta.ddlMySQLd.clusterIp.name } : Type `string`, default `"ddl-mysqld"`. `rondb.rondb.meta.ddlMySQLd.clusterIp.port` # { #helm.rondb.rondb.meta.ddlMySQLd.clusterIp.port } : Type `integer`, default `3306`. `rondb.rondb.meta.ddlMySQLd.enabled` # { #helm.rondb.rondb.meta.ddlMySQLd.enabled } : Type `boolean`, default `true`, Hopsworks overrides the `rondb` chart default `false`. Whether to deploy the DDL MySQLd StatefulSet and reserve its NDB node slot. `rondb.rondb.meta.ddlMySQLd.headlessClusterIp` # { #helm.rondb.rondb.meta.ddlMySQLd.headlessClusterIp } : Type `object`. `rondb.rondb.meta.ddlMySQLd.headlessClusterIp.name` # { #helm.rondb.rondb.meta.ddlMySQLd.headlessClusterIp.name } : Type `string`, default `"headless-ddl-mysqld"`. `rondb.rondb.meta.ddlMySQLd.headlessClusterIp.port` # { #helm.rondb.rondb.meta.ddlMySQLd.headlessClusterIp.port } : Type `integer`, default `3306`. `rondb.rondb.meta.ddlMySQLd.statefulSet` # { #helm.rondb.rondb.meta.ddlMySQLd.statefulSet } : Type `object`. `rondb.rondb.meta.ddlMySQLd.statefulSet.endToEndTls` # { #helm.rondb.rondb.meta.ddlMySQLd.statefulSet.endToEndTls } : Type `object`. `rondb.rondb.meta.ddlMySQLd.statefulSet.endToEndTls.enabled` # { #helm.rondb.rondb.meta.ddlMySQLd.statefulSet.endToEndTls.enabled } : Type `boolean`, default `false`. Whether to use end-to-end encryption for DDL MySQL server Pods. `rondb.rondb.meta.ddlMySQLd.statefulSet.endToEndTls.filenames` # { #helm.rondb.rondb.meta.ddlMySQLd.statefulSet.endToEndTls.filenames } : Type `object`. `rondb.rondb.meta.ddlMySQLd.statefulSet.endToEndTls.filenames.ca` # { #helm.rondb.rondb.meta.ddlMySQLd.statefulSet.endToEndTls.filenames.ca } : Type `string|null`, default `"hops_root_ca.pem"`, Hopsworks overrides the `rondb` chart default `null`. Name of the CA file in the Secret. ONLY overwrite this if you're not using standard TLS Secrets. `rondb.rondb.meta.ddlMySQLd.statefulSet.endToEndTls.filenames.cert` # { #helm.rondb.rondb.meta.ddlMySQLd.statefulSet.endToEndTls.filenames.cert } : Type `string`, default `"mysqlddl_certificate_bundle.pem"`, Hopsworks overrides the `rondb` chart default `"tls.crt"`. Name of the certificate file in the Secret. ONLY overwrite this if you're not using standard TLS Secrets. `rondb.rondb.meta.ddlMySQLd.statefulSet.endToEndTls.filenames.key` # { #helm.rondb.rondb.meta.ddlMySQLd.statefulSet.endToEndTls.filenames.key } : Type `string`, default `"mysqlddl_priv.pem"`, Hopsworks overrides the `rondb` chart default `"tls.key"`. Name of the key file in the Secret. ONLY overwrite this if you're not using standard TLS Secrets. `rondb.rondb.meta.ddlMySQLd.statefulSet.endToEndTls.secretName` # { #helm.rondb.rondb.meta.ddlMySQLd.statefulSet.endToEndTls.secretName } : Type `string`, default `"mysqlddl-crypto-material"`, Hopsworks overrides the `rondb` chart default `"ddl-mysqld-end-to-end-tls"`. Name of the TLS Secret `rondb.rondb.meta.ddlMySQLd.statefulSet.endToEndTls.supplyOwnSecret` # { #helm.rondb.rondb.meta.ddlMySQLd.statefulSet.endToEndTls.supplyOwnSecret } : Type `boolean`, default `true`, Hopsworks overrides the `rondb` chart default `false`. Whether the Helmchart user will create a TLS Secret outside of this Helmchart. Otherwise we rely on cert-manager to create one. `rondb.rondb.meta.ddlMySQLd.statefulSet.name` # { #helm.rondb.rondb.meta.ddlMySQLd.statefulSet.name } : Type `string`, default `"ddl-mysqld"`. `rondb.rondb.meta.mgmd` # { #helm.rondb.rondb.meta.mgmd } : Type `object`. `rondb.rondb.meta.mgmd.headlessClusterIp` # { #helm.rondb.rondb.meta.mgmd.headlessClusterIp } : Type `object`. `rondb.rondb.meta.mgmd.headlessClusterIp.annotations` # { #helm.rondb.rondb.meta.mgmd.headlessClusterIp.annotations } : Type `object`, default `{"consul.hashicorp.com/service-name":"mgmd"}`, Hopsworks overrides the `rondb` chart default `{}`. `rondb.rondb.meta.mgmd.headlessClusterIp.name` # { #helm.rondb.rondb.meta.mgmd.headlessClusterIp.name } : Type `string`, default `"headless-mgmds"`. `rondb.rondb.meta.mgmd.headlessClusterIp.port` # { #helm.rondb.rondb.meta.mgmd.headlessClusterIp.port } : Type `integer`, default `1186`. `rondb.rondb.meta.mgmd.statefulSetName` # { #helm.rondb.rondb.meta.mgmd.statefulSetName } : Type `string`, default `"mgmds"`. `rondb.rondb.meta.mysqld` # { #helm.rondb.rondb.meta.mysqld } : Type `object`. `rondb.rondb.meta.mysqld.addSysNiceCapability` # { #helm.rondb.rondb.meta.mysqld.addSysNiceCapability } : Type `boolean`, default `true`. `rondb.rondb.meta.mysqld.clusterIp` # { #helm.rondb.rondb.meta.mysqld.clusterIp } : Type `object`. `rondb.rondb.meta.mysqld.clusterIp.annotations` # { #helm.rondb.rondb.meta.mysqld.clusterIp.annotations } : Type `object`, Hopsworks overrides the `rondb` chart default `{}`. ??? note "Default" ```yaml consul.hashicorp.com/service-name: mysql consul.hashicorp.com/service-tags: onlinefs ``` `rondb.rondb.meta.mysqld.clusterIp.name` # { #helm.rondb.rondb.meta.mysqld.clusterIp.name } : Type `string`, default `"mysqld"`. `rondb.rondb.meta.mysqld.clusterIp.port` # { #helm.rondb.rondb.meta.mysqld.clusterIp.port } : Type `integer`, default `3306`. `rondb.rondb.meta.mysqld.exporter` # { #helm.rondb.rondb.meta.mysqld.exporter } : Type `object`. Configuration for mysqld exporter `rondb.rondb.meta.mysqld.exporter.metricsPort` # { #helm.rondb.rondb.meta.mysqld.exporter.metricsPort } : Type `integer`, default `9104`. `rondb.rondb.meta.mysqld.externalLoadBalancer` # { #helm.rondb.rondb.meta.mysqld.externalLoadBalancer } : Type `object`. Configuration for load balancer service to be used for external access `rondb.rondb.meta.mysqld.externalLoadBalancer.annotations` # { #helm.rondb.rondb.meta.mysqld.externalLoadBalancer.annotations } : Type `object`, default `{}`. Cloud provider load balancer specific annotations. `rondb.rondb.meta.mysqld.externalLoadBalancer.class` # { #helm.rondb.rondb.meta.mysqld.externalLoadBalancer.class } : Type `string|null`, default `null`. `rondb.rondb.meta.mysqld.externalLoadBalancer.enabled` # { #helm.rondb.rondb.meta.mysqld.externalLoadBalancer.enabled } : Type `boolean`, default `true`, Hopsworks overrides the `rondb` chart default `false`. `rondb.rondb.meta.mysqld.externalLoadBalancer.managed` # { #helm.rondb.rondb.meta.mysqld.externalLoadBalancer.managed } : Type `boolean`, default `true`. LoadBalancer is managed by provider `rondb.rondb.meta.mysqld.externalLoadBalancer.name` # { #helm.rondb.rondb.meta.mysqld.externalLoadBalancer.name } : Type `string`, default `"mysqld-external"`. `rondb.rondb.meta.mysqld.externalLoadBalancer.nodePort` # { #helm.rondb.rondb.meta.mysqld.externalLoadBalancer.nodePort } : Type `integer|null`, default `null`, minimum `1`, maximum `65535`. Explicit nodePort for the service when the load balancer is unmanaged (managed: false), so a load balancer outside Kubernetes can target a fixed port. Null lets Kubernetes allocate one from the cluster's node-port range. `rondb.rondb.meta.mysqld.externalLoadBalancer.nodeSelector` # { #helm.rondb.rondb.meta.mysqld.externalLoadBalancer.nodeSelector } : Type `object`, default `{}`. selector for nodes the load balancer can use to route traffic `rondb.rondb.meta.mysqld.externalLoadBalancer.port` # { #helm.rondb.rondb.meta.mysqld.externalLoadBalancer.port } : Type `integer`, default `3306`. Port to expose the service on `rondb.rondb.meta.mysqld.headlessClusterIp` # { #helm.rondb.rondb.meta.mysqld.headlessClusterIp } : Type `object`. `rondb.rondb.meta.mysqld.headlessClusterIp.name` # { #helm.rondb.rondb.meta.mysqld.headlessClusterIp.name } : Type `string`, default `"headless-mysqlds"`. `rondb.rondb.meta.mysqld.headlessClusterIp.port` # { #helm.rondb.rondb.meta.mysqld.headlessClusterIp.port } : Type `integer`, default `3306`. `rondb.rondb.meta.mysqld.statefulSet` # { #helm.rondb.rondb.meta.mysqld.statefulSet } : Type `object`. `rondb.rondb.meta.mysqld.statefulSet.endToEndTls` # { #helm.rondb.rondb.meta.mysqld.statefulSet.endToEndTls } : Type `object`. `rondb.rondb.meta.mysqld.statefulSet.endToEndTls.enabled` # { #helm.rondb.rondb.meta.mysqld.statefulSet.endToEndTls.enabled } : default `false`. Whether to use end-to-end encryption for MySQLd Pods. This is recommended for high-security use cases. For MySQLds this is especially recommended since they use raw TCP connections and thereby by-pass TLS Ingress-rules. Otherwise, Ingress-TCP can also be configured directly via the Ingress controller (not in this Helmchart). `rondb.rondb.meta.mysqld.statefulSet.endToEndTls.filenames` # { #helm.rondb.rondb.meta.mysqld.statefulSet.endToEndTls.filenames } : Type `object`. `rondb.rondb.meta.mysqld.statefulSet.endToEndTls.filenames.ca` # { #helm.rondb.rondb.meta.mysqld.statefulSet.endToEndTls.filenames.ca } : Type `string|null`, default `"hops_root_ca.pem"`, Hopsworks overrides the `rondb` chart default `null`. Name of the CA file in the Secret. ONLY overwrite this if you're not using standard TLS Secrets. `rondb.rondb.meta.mysqld.statefulSet.endToEndTls.filenames.cert` # { #helm.rondb.rondb.meta.mysqld.statefulSet.endToEndTls.filenames.cert } : Type `string`, default `"mysql_certificate_bundle.pem"`, Hopsworks overrides the `rondb` chart default `"tls.crt"`. Name of the certificate file in the Secret. ONLY overwrite this if you're not using standard TLS Secrets. `rondb.rondb.meta.mysqld.statefulSet.endToEndTls.filenames.key` # { #helm.rondb.rondb.meta.mysqld.statefulSet.endToEndTls.filenames.key } : Type `string`, default `"mysql_priv.pem"`, Hopsworks overrides the `rondb` chart default `"tls.key"`. Name of the key file in the Secret. ONLY overwrite this if you're not using standard TLS Secrets. `rondb.rondb.meta.mysqld.statefulSet.endToEndTls.secretName` # { #helm.rondb.rondb.meta.mysqld.statefulSet.endToEndTls.secretName } : Type `string`, default `"mysqld-crypto-material"`, Hopsworks overrides the `rondb` chart default `"mysqld-end-to-end-tls"`. Name of the TLS Secret `rondb.rondb.meta.mysqld.statefulSet.endToEndTls.supplyOwnSecret` # { #helm.rondb.rondb.meta.mysqld.statefulSet.endToEndTls.supplyOwnSecret } : default `true`, Hopsworks overrides the `rondb` chart default `false`. Whether the Helmchart user will create a TLS Secret outside of this Helmchart. Otherwise we rely on cert-manager to create one. `rondb.rondb.meta.mysqld.statefulSet.name` # { #helm.rondb.rondb.meta.mysqld.statefulSet.name } : Type `string`, default `"mysqlds"`. `rondb.rondb.meta.ndbmtd` # { #helm.rondb.rondb.meta.ndbmtd } : Type `object`. `rondb.rondb.meta.ndbmtd.statefulSet` # { #helm.rondb.rondb.meta.ndbmtd.statefulSet } : Type `object`. `rondb.rondb.meta.ndbmtd.statefulSet.podAnnotations` # { #helm.rondb.rondb.meta.ndbmtd.statefulSet.podAnnotations } : Type `object`, default `{}`. `rondb.rondb.meta.rdrs` # { #helm.rondb.rondb.meta.rdrs } : Type `object`. `rondb.rondb.meta.rdrs.clusterIp` # { #helm.rondb.rondb.meta.rdrs.clusterIp } : Type `object`. `rondb.rondb.meta.rdrs.clusterIp.annotations` # { #helm.rondb.rondb.meta.rdrs.clusterIp.annotations } : Type `object`. ??? note "Default" ```yaml consul.hashicorp.com/service-name: rdrs prometheus.io/path: /metrics prometheus.io/port: '4406' prometheus.io/scheme: https prometheus.io/scrape: 'true' ``` `rondb.rondb.meta.rdrs.clusterIp.name` # { #helm.rondb.rondb.meta.rdrs.clusterIp.name } : Type `string`, default `"rdrs"`. `rondb.rondb.meta.rdrs.externalLoadBalancer` # { #helm.rondb.rondb.meta.rdrs.externalLoadBalancer } : Type `object`. Configuration for load balancer service to be used for external access `rondb.rondb.meta.rdrs.externalLoadBalancer.annotations` # { #helm.rondb.rondb.meta.rdrs.externalLoadBalancer.annotations } : Type `object`, default `{}`. Cloud provider load balancer specific annotations. `rondb.rondb.meta.rdrs.externalLoadBalancer.class` # { #helm.rondb.rondb.meta.rdrs.externalLoadBalancer.class } : Type `string|null`, default `null`. `rondb.rondb.meta.rdrs.externalLoadBalancer.enabled` # { #helm.rondb.rondb.meta.rdrs.externalLoadBalancer.enabled } : Type `boolean`, default `true`, Hopsworks overrides the `rondb` chart default `false`. `rondb.rondb.meta.rdrs.externalLoadBalancer.managed` # { #helm.rondb.rondb.meta.rdrs.externalLoadBalancer.managed } : Type `boolean`, default `true`. LoadBalancer is managed by provider `rondb.rondb.meta.rdrs.externalLoadBalancer.name` # { #helm.rondb.rondb.meta.rdrs.externalLoadBalancer.name } : Type `string`, default `"rdrs-external"`. `rondb.rondb.meta.rdrs.externalLoadBalancer.nodePort` # { #helm.rondb.rondb.meta.rdrs.externalLoadBalancer.nodePort } : Type `integer|null`, default `null`, minimum `1`, maximum `65535`. Explicit nodePort for the service when the load balancer is unmanaged (managed: false), so a load balancer outside Kubernetes can target a fixed port. Null lets Kubernetes allocate one from the cluster's node-port range. `rondb.rondb.meta.rdrs.externalLoadBalancer.nodeSelector` # { #helm.rondb.rondb.meta.rdrs.externalLoadBalancer.nodeSelector } : Type `object`, default `{}`. selector for nodes the load balancer can use to route traffic `rondb.rondb.meta.rdrs.headlessClusterIpName` # { #helm.rondb.rondb.meta.rdrs.headlessClusterIpName } : Type `string`, default `"rdrs-cluster-ip"`. `rondb.rondb.meta.rdrs.ingress` # { #helm.rondb.rondb.meta.rdrs.ingress } : Type `object`. Configuration of Ingress for RDRS `rondb.rondb.meta.rdrs.ingress.class` # { #helm.rondb.rondb.meta.rdrs.ingress.class } : Type `string`, default `"nginx"`. `rondb.rondb.meta.rdrs.ingress.dnsNames` # { #helm.rondb.rondb.meta.rdrs.ingress.dnsNames } : Type `array`, default `[]`, example `"rondb.com"`. `rondb.rondb.meta.rdrs.ingress.enabled` # { #helm.rondb.rondb.meta.rdrs.ingress.enabled } : Type `boolean`, default `false`. `rondb.rondb.meta.rdrs.ingress.tls` # { #helm.rondb.rondb.meta.rdrs.ingress.tls } : Type `object`. `rondb.rondb.meta.rdrs.ingress.tls.enabled` # { #helm.rondb.rondb.meta.rdrs.ingress.tls.enabled } : Type `boolean`, default `true`. WARN: Nginx-ingress will always use encryption even if this is disabled. By enabling this, we simply have more control over the TLS Secret. The TLS Secrets are placed onto the Ingress controller instance. Ingress TLS is currently only supported with cert-manager (RonDB-standalone) `rondb.rondb.meta.rdrs.ingress.tls.ipAddresses` # { #helm.rondb.rondb.meta.rdrs.ingress.tls.ipAddresses } : Type `array`, default `[]`, example `"127.0.0.1"`. `rondb.rondb.meta.rdrs.ingress.useDefaultBackend` # { #helm.rondb.rondb.meta.rdrs.ingress.useDefaultBackend } : Type `boolean`, default `true`. Whether to use the RDRS as a default backend for the Ingress; this makes it reachable without specifying a subdomain; i.e. an IP can be used instead. `rondb.rondb.meta.rdrs.statefulSet` # { #helm.rondb.rondb.meta.rdrs.statefulSet } : Type `object`. `rondb.rondb.meta.rdrs.statefulSet.endToEndTls` # { #helm.rondb.rondb.meta.rdrs.statefulSet.endToEndTls } : Type `object`. `rondb.rondb.meta.rdrs.statefulSet.endToEndTls.enabled` # { #helm.rondb.rondb.meta.rdrs.statefulSet.endToEndTls.enabled } : default `true`, Hopsworks overrides the `rondb` chart default `false`. Whether to use end-to-end encryption for RDRS Pods. This is recommended for high-security use cases. `rondb.rondb.meta.rdrs.statefulSet.endToEndTls.filenames` # { #helm.rondb.rondb.meta.rdrs.statefulSet.endToEndTls.filenames } : Type `object`. `rondb.rondb.meta.rdrs.statefulSet.endToEndTls.filenames.ca` # { #helm.rondb.rondb.meta.rdrs.statefulSet.endToEndTls.filenames.ca } : Type `string|null`, default `"hops_root_ca.pem"`, Hopsworks overrides the `rondb` chart default `null`. Name of the CA file in the Secret. ONLY overwrite this if you're not using standard TLS Secrets. `rondb.rondb.meta.rdrs.statefulSet.endToEndTls.filenames.cert` # { #helm.rondb.rondb.meta.rdrs.statefulSet.endToEndTls.filenames.cert } : Type `string`, default `"mysql_certificate_bundle.pem"`, Hopsworks overrides the `rondb` chart default `"tls.crt"`. Name of the certificate file in the Secret. ONLY overwrite this if you're not using standard TLS Secrets. `rondb.rondb.meta.rdrs.statefulSet.endToEndTls.filenames.key` # { #helm.rondb.rondb.meta.rdrs.statefulSet.endToEndTls.filenames.key } : Type `string`, default `"mysql_priv.pem"`, Hopsworks overrides the `rondb` chart default `"tls.key"`. Name of the key file in the Secret. ONLY overwrite this if you're not using standard TLS Secrets. `rondb.rondb.meta.rdrs.statefulSet.endToEndTls.secretName` # { #helm.rondb.rondb.meta.rdrs.statefulSet.endToEndTls.secretName } : Type `string`, default `"rdrs-crypto-material"`, Hopsworks overrides the `rondb` chart default `"rdrs-end-to-end-tls"`. Name of the TLS Secret `rondb.rondb.meta.rdrs.statefulSet.endToEndTls.supplyOwnSecret` # { #helm.rondb.rondb.meta.rdrs.statefulSet.endToEndTls.supplyOwnSecret } : default `true`, Hopsworks overrides the `rondb` chart default `false`. Whether the Helmchart user will create a TLS Secret outside of this Helmchart. Otherwise we rely on cert-manager to create one. `rondb.rondb.meta.rdrs.statefulSet.name` # { #helm.rondb.rondb.meta.rdrs.statefulSet.name } : Type `string`, default `"rdrs"`. `rondb.rondb.meta.replicaAppliers` # { #helm.rondb.rondb.meta.replicaAppliers } : Type `object`. `rondb.rondb.meta.replicaAppliers.headlessClusterIp` # { #helm.rondb.rondb.meta.replicaAppliers.headlessClusterIp } : Type `object`. `rondb.rondb.meta.replicaAppliers.headlessClusterIp.name` # { #helm.rondb.rondb.meta.replicaAppliers.headlessClusterIp.name } : Type `string`, default `"headless-replica-appliers"`. `rondb.rondb.meta.replicaAppliers.headlessClusterIp.port` # { #helm.rondb.rondb.meta.replicaAppliers.headlessClusterIp.port } : Type `integer`, default `3306`. `rondb.rondb.meta.replicaAppliers.statefulSet` # { #helm.rondb.rondb.meta.replicaAppliers.statefulSet } : Type `object`. `rondb.rondb.meta.replicaAppliers.statefulSet.endToEndTls` # { #helm.rondb.rondb.meta.replicaAppliers.statefulSet.endToEndTls } : Type `object`. `rondb.rondb.meta.replicaAppliers.statefulSet.endToEndTls.enabled` # { #helm.rondb.rondb.meta.replicaAppliers.statefulSet.endToEndTls.enabled } : default `false`. Whether to use end-to-end encryption for MySQL replica applier Pods. This is recommended for high-security use cases. For MySQLds this is especially recommended since they use raw TCP connections and thereby by-pass TLS Ingress-rules. Otherwise, Ingress-TCP can also be configured directly via the Ingress controller (not in this Helmchart). `rondb.rondb.meta.replicaAppliers.statefulSet.endToEndTls.filenames` # { #helm.rondb.rondb.meta.replicaAppliers.statefulSet.endToEndTls.filenames } : Type `object`. `rondb.rondb.meta.replicaAppliers.statefulSet.endToEndTls.filenames.ca` # { #helm.rondb.rondb.meta.replicaAppliers.statefulSet.endToEndTls.filenames.ca } : Type `string|null`, default `null`. Name of the CA file in the Secret. ONLY overwrite this if you're not using standard TLS Secrets. `rondb.rondb.meta.replicaAppliers.statefulSet.endToEndTls.filenames.cert` # { #helm.rondb.rondb.meta.replicaAppliers.statefulSet.endToEndTls.filenames.cert } : Type `string`, default `"tls.crt"`. Name of the certificate file in the Secret. ONLY overwrite this if you're not using standard TLS Secrets. `rondb.rondb.meta.replicaAppliers.statefulSet.endToEndTls.filenames.key` # { #helm.rondb.rondb.meta.replicaAppliers.statefulSet.endToEndTls.filenames.key } : Type `string`, default `"tls.key"`. Name of the key file in the Secret. ONLY overwrite this if you're not using standard TLS Secrets. `rondb.rondb.meta.replicaAppliers.statefulSet.endToEndTls.secretName` # { #helm.rondb.rondb.meta.replicaAppliers.statefulSet.endToEndTls.secretName } : Type `string`, default `"replica-applier-end-to-end-tls"`. Name of the TLS Secret `rondb.rondb.meta.replicaAppliers.statefulSet.endToEndTls.supplyOwnSecret` # { #helm.rondb.rondb.meta.replicaAppliers.statefulSet.endToEndTls.supplyOwnSecret } : default `false`. Whether the Helmchart user will create a TLS Secret outside of this Helmchart. Otherwise we rely on cert-manager to create one. `rondb.rondb.meta.replicaAppliers.statefulSet.name` # { #helm.rondb.rondb.meta.replicaAppliers.statefulSet.name } : Type `string`, default `"mysqld-replica-appliers"`.
### mysql { #helm-values-rondb-rondb-mysql } ??? example "Defaults as YAML" ```yaml rondb: rondb: mysql: clusterUser: bench config: maxConnectErrors: '9223372036854775807' maxConnections: 512 maxPreparedStmtCount: 65530 credentialsSecretName: mysql-users-secrets exporter: enabled: true maxOpenConnections: 1 maxUserConnections: 3 username: exporter force2410UserGrantMigration: false sqlInitContent: {} supplyOwnSecret: false users: - host: '%' privileges: - database: '*' privileges: - ALL table: '*' withGrantOption: true username: hopsworksroot ```
`rondb.rondb.mysql` # { #helm.rondb.rondb.mysql } : Type `object`. How to initialize MySQL `rondb.rondb.mysql.clusterUser` # { #helm.rondb.rondb.mysql.clusterUser } : Type `string`, default `"bench"`, Hopsworks overrides the `rondb` chart default `"helm"`. The MySQL user for K8s probes, benchmarks and for the standard my.cnf file. `rondb.rondb.mysql.config` # { #helm.rondb.rondb.mysql.config } : Type `object`. MySQL configuration options `rondb.rondb.mysql.config.maxConnectErrors` # { #helm.rondb.rondb.mysql.config.maxConnectErrors } : Type `string`, default `"9223372036854775807"`. Maximum number of connection errors before the MySQL server blocks the host. `rondb.rondb.mysql.config.maxConnections` # { #helm.rondb.rondb.mysql.config.maxConnections } : Type `integer`, default `512`. Maximum number of connections to the MySQL server. `rondb.rondb.mysql.config.maxPreparedStmtCount` # { #helm.rondb.rondb.mysql.config.maxPreparedStmtCount } : Type `integer`, default `65530`. Maximum number of prepared statements. `rondb.rondb.mysql.credentialsSecretName` # { #helm.rondb.rondb.mysql.credentialsSecretName } : Type `string`, default `"mysql-users-secrets"`, Hopsworks overrides the `rondb` chart default `"mysql-passwords"`. Secret name for MySQL users' passwords `rondb.rondb.mysql.exporter` # { #helm.rondb.rondb.mysql.exporter } : Type `object`. MySQL exporter configuration `rondb.rondb.mysql.exporter.enabled` # { #helm.rondb.rondb.mysql.exporter.enabled } : Type `boolean`, default `true`, Hopsworks overrides the `rondb` chart default `false`. `rondb.rondb.mysql.exporter.maxOpenConnections` # { #helm.rondb.rondb.mysql.exporter.maxOpenConnections } : Type `integer`, default `1`, minimum `1`. Maximum number of open connections to the database per scrape. 1 (default) serializes all collectors onto a single connection; higher values let independent collectors run concurrently but may incur higher load on the database. Should not exceed maxUserConnections. `rondb.rondb.mysql.exporter.maxUserConnections` # { #helm.rondb.rondb.mysql.exporter.maxUserConnections } : Type `integer`, default `3`. `rondb.rondb.mysql.exporter.username` # { #helm.rondb.rondb.mysql.exporter.username } : Type `string`, default `"exporter"`. `rondb.rondb.mysql.force2410UserGrantMigration` # { #helm.rondb.rondb.mysql.force2410UserGrantMigration } : Type `boolean`, default `false`. Whether to force the execution of the 24.10 user privilege migration job that revokes the SET_USER_ID privilege from users that should not have it. This is useful when the automatic detection fails. `rondb.rondb.mysql.sqlInitContent` # { #helm.rondb.rondb.mysql.sqlInitContent } : Type `object`, default `{}`, example `{"createMyUser":"CREATE USER foo IF NOT EXISTS;"}`. SQL to run *only once* at cluster startup. Try to make these scripts idempotent, in case they are re-run by accident. Do so by e.g. using `IF NOT EXISTS` in the SQL commands. `rondb.rondb.mysql.supplyOwnSecret` # { #helm.rondb.rondb.mysql.supplyOwnSecret } : Type `boolean`, default `false`. If set to false, the Helmchart will auto-generate MySQL passwords. When running Global Replication as a secondary cluster, this should be set to true. Otherwise, the replication of ALTER root password will fail the cluster. `rondb.rondb.mysql.users` # { #helm.rondb.rondb.mysql.users } : Type `array`, Hopsworks overrides the `rondb` chart default `[]`. A list of MySQL users and their privileges. ??? note "Default" ```yaml - username: hopsworksroot host: '%' privileges: - database: '*' table: '*' withGrantOption: true privileges: - ALL ``` `rondb.rondb.mysql.users[].existingSecret` # { #helm.rondb.rondb.mysql.users.existingSecret } : Type `object`. Reference to a pre-created Kubernetes Secret holding this user's password. When set, the Helmchart does not generate or store a password for this user in mysql.credentialsSecretName; the referenced Secret must exist in the release namespace before install/upgrade. `rondb.rondb.mysql.users[].existingSecret.key` # { #helm.rondb.rondb.mysql.users.existingSecret.key } : Type `string`. Key within the Secret whose value is the user's password. `rondb.rondb.mysql.users[].existingSecret.name` # { #helm.rondb.rondb.mysql.users.existingSecret.name } : Type `string`. Name of the Kubernetes Secret containing the user's password. `rondb.rondb.mysql.users[].existingSecret.rotationId` # { #helm.rondb.rondb.mysql.users.existingSecret.rotationId } : Type `string`. Opaque rotation marker. Bump this (e.g. to a date or counter) whenever the Secret's password is rotated. It participates in the user-setup Job's name, forcing a re-run that converges the MySQL password to the Secret (ALTER USER). Required for rotation under Argo CD / template-only rendering, where the release revision is constant; plain `helm upgrade` re-runs the Job on every upgrade regardless. `rondb.rondb.mysql.users[].host` # { #helm.rondb.rondb.mysql.users.host } : Type `string`, example `"%"`. The host from which the MySQL user can connect. `rondb.rondb.mysql.users[].privileges` # { #helm.rondb.rondb.mysql.users.privileges } : Type `array`. Privileges assigned to the MySQL user. `rondb.rondb.mysql.users[].privileges[].database` # { #helm.rondb.rondb.mysql.users.privileges.database } : Type `string`, example `"*"`. The MySQL database to which the privileges apply. `rondb.rondb.mysql.users[].privileges[].privileges` # { #helm.rondb.rondb.mysql.users.privileges.privileges } : Type `array`. `rondb.rondb.mysql.users[].privileges[].table` # { #helm.rondb.rondb.mysql.users.privileges.table } : Type `string`, example `"*"`. The MySQL table to which the privileges apply. `rondb.rondb.mysql.users[].privileges[].withGrantOption` # { #helm.rondb.rondb.mysql.users.privileges.withGrantOption } : Type `boolean`, default `false`. Whether the user has the GRANT OPTION privilege. `rondb.rondb.mysql.users[].username` # { #helm.rondb.rondb.mysql.users.username } : Type `string`. The username of the MySQL database user.
### ndbmtdSequencedRollout { #helm-values-rondb-rondb-ndbmtdsequencedrollout } ??? example "Defaults as YAML" ```yaml rondb: rondb: ndbmtdSequencedRollout: enabled: false perGroupStallTimeoutMinutes: 0 reconcileIntervalMinutes: 3 suspendWhenIdle: true ```
`rondb.rondb.ndbmtdSequencedRollout` # { #helm.rondb.rondb.ndbmtdSequencedRollout } : Type `object`. Sequenced (one node group at a time) data node rollouts during upgrades. When enabled, the node-group StatefulSets render updateStrategy.rollingUpdate.partition equal to their replica count, so 'helm upgrade' lands new pod templates without restarting anything; a rollout CronJob then unfreezes one node group at a time, waiting for each to converge (all pods on the new revision and ready) before the next. This bounds upgrade exposure to a single node group and stops a bad rollout at the first group. Only takes effect with more than one node group (clusterSize.numNodeGroups > 1): with a single group there is nothing to sequence across and the chart behaves exactly as if disabled. Not applied during in-place restores or on externally managed clusters. Argo CD users with selfHeal enabled must add ignoreDifferences for .spec.updateStrategy.rollingUpdate.partition on node-group-* StatefulSets. `rondb.rondb.ndbmtdSequencedRollout.enabled` # { #helm.rondb.rondb.ndbmtdSequencedRollout.enabled } : Type `boolean`, default `false`. Enable the partition freeze and the rollout CronJob (they always render together). With this on, a successful 'helm upgrade' means the new specs are recorded and the rollout is in progress; completion is observable on the StatefulSets (updateRevision == currentRevision on every node group) rather than in the release status. Disabling the flag removes the partition field, so any still-pending update then rolls all node groups concurrently (the pre-feature behavior) - disable during a quiet period or after confirming no update is pending. `rondb.rondb.ndbmtdSequencedRollout.perGroupStallTimeoutMinutes` # { #helm.rondb.rondb.ndbmtdSequencedRollout.perGroupStallTimeoutMinutes } : Type `integer`, default `0`, minimum `0`. Minutes an unfrozen node group may stay unconverged before the rollout CronJob reports a stall (logged on every run) and pauses; no further groups are unfrozen while the stalled group keeps trying. Note the next 'helm upgrade' (e.g. a fixed image) re-freezes every group, the stalled one included (Helm restores the rendered partition over the live value); the CronJob then delivers the fix by unfreezing the broken group again on the next run, in preference to healthy ones, and the StatefulSet controller replaces its dead pod. This is an alerting threshold - nothing is killed when it fires. 0 derives it from activeDataReplicas x timeoutsMinutes.ndbmtdStartupProbe plus 30 minutes slack. `rondb.rondb.ndbmtdSequencedRollout.reconcileIntervalMinutes` # { #helm.rondb.rondb.ndbmtdSequencedRollout.reconcileIntervalMinutes } : Type `integer`, default `3`, minimum `1`, maximum `30`. How often the rollout CronJob runs, in minutes. Each run re-freezes converged node groups and unfreezes at most one pending group, so this adds at most one interval of latency per node-group boundary - noise against recovery times measured in hours. `rondb.rondb.ndbmtdSequencedRollout.suspendWhenIdle` # { #helm.rondb.rondb.ndbmtdSequencedRollout.suspendWhenIdle } : Type `boolean`, default `true`. When a run finds every node group converged and frozen, the rollout CronJob suspends itself so nothing at all runs while idle. Any 'helm upgrade' or 'helm rollback' turns it back on: the chart renders suspend: false and Helm resets live fields to their rendered values. Only takes effect when 'mode' is unset or 'auto' (plain Helm or Flux); when 'mode' is set - the Argo CD convention - the CronJob stays always-on, because Argo's sync would either keep re-enabling it (saving nothing) or, with ignoreDifferences on the CronJob's .spec.suspend, never re-enable it. WARNING: while suspended, a node-group spec change applied outside Helm (plain kubectl) stays frozen until the next helm operation or a manual wake: kubectl patch cronjob rondb-ndbmtd-sequenced-rollout --type merge -p '{"spec":{"suspend":false}}'. Pair with a revision-age alert (kube_statefulset_status_update_revision != kube_statefulset_status_current_revision with an age threshold) as the safety net. Set false to keep the CronJob always running.
### ndbmtdSettleWait { #helm-values-rondb-rondb-ndbmtdsettlewait } ??? example "Defaults as YAML" ```yaml rondb: rondb: ndbmtdSettleWait: fallbackSeconds: 15 maxWaitSeconds: 30 probeTimeoutSeconds: 5 quietSeconds: 8 ```
`rondb.rondb.ndbmtdSettleWait` # { #helm.rondb.rondb.ndbmtdSettleWait } : Type `object`. Pre-start settle wait for data nodes, protecting rolling restarts. During a roll, Kubernetes deletes a node group's second pod only after observing the first replacement Ready; that observation spread (measured 2.6-10.4s and set by the kubelet/API-server publication cycle, which the chart cannot tune) can overlap the window in which an earlier replacement has connected to the cluster but not yet reached the phase-110 restart barrier, killing it with error 2308 ('Another node failed during system restart'). Before starting the kernel on a non-initial start, the entrypoint polls the MGMd once a second and proceeds only once no data node has departed the cluster for quietSeconds (nodes reconnecting do not reset the timer: only a departing peer endangers a climbing node, and during a round every replacement's reconnect would otherwise extend the wait to its cap), so the node begins its climb only after the deletion wave has passed. The wait is bounded by maxWaitSeconds and never blocks a start: an unreachable or hanging MGMd degrades to a fixed fallbackSeconds sleep. Skipped entirely on initial starts. `rondb.rondb.ndbmtdSettleWait.fallbackSeconds` # { #helm.rondb.rondb.ndbmtdSettleWait.fallbackSeconds } : Type `integer`, default `15`, minimum `0`. Fixed sleep used instead of the adaptive wait when the MGMd never answers: 3 consecutive failed probes (unreachable, hanging, or empty output) with no successful probe in this wait. A transient MGMd outage after a successful probe does not trigger the fallback - the wait keeps retrying under maxWaitSeconds, with failed probes resetting the quiet timer. `rondb.rondb.ndbmtdSettleWait.maxWaitSeconds` # { #helm.rondb.rondb.ndbmtdSettleWait.maxWaitSeconds } : Type `integer`, default `30`, minimum `0`. Cap on the whole settle wait, including the final quietSeconds of quiet. 0 disables the settle wait entirely. `rondb.rondb.ndbmtdSettleWait.probeTimeoutSeconds` # { #helm.rondb.rondb.ndbmtdSettleWait.probeTimeoutSeconds } : Type `integer`, default `5`, minimum `1`. Timeout for each ndb_mgm membership probe. `rondb.rondb.ndbmtdSettleWait.quietSeconds` # { #helm.rondb.rondb.ndbmtdSettleWait.quietSeconds } : Type `integer`, default `8`, minimum `1`. Seconds without any data-node departure before the node starts (arrivals are ignored). Must exceed the largest gap between consecutive pod deletions within a rollout round, or the wait provides no protection. The default of 8 is calibrated to ONE measured cluster (max 7.4s gap over 66 rolls at 10 node groups on MicroK8s); the gap is set by the control plane's readiness-observation latency, not by RonDB, so it MUST be re-derived on a different control plane, larger cluster, or slower hardware - measure the deletion-timestamp gaps within one rollout round and set this above the maximum.
### networkPolicy { #helm-values-rondb-rondb-networkpolicy } ??? example "Defaults as YAML" ```yaml rondb: rondb: networkPolicy: mgmds: enabled: true ingressSelectors: - podSelector: matchLabels: access: mgmd-and-ndbmtd ndbmtds: enabled: true ingressSelectors: - podSelector: matchLabels: access: mgmd-and-ndbmtd ```
`rondb.rondb.networkPolicy` # { #helm.rondb.rondb.networkPolicy } : Type `object`. `rondb.rondb.networkPolicy.mgmds` # { #helm.rondb.rondb.networkPolicy.mgmds } : Type `object`. `rondb.rondb.networkPolicy.mgmds.enabled` # { #helm.rondb.rondb.networkPolicy.mgmds.enabled } : Type `boolean`, default `true`. Whether to limit ingress for MGMd pods `rondb.rondb.networkPolicy.mgmds.ingressSelectors` # { #helm.rondb.rondb.networkPolicy.mgmds.ingressSelectors } : Type `array`, default `[{"podSelector":{"matchLabels":{"access":"mgmd-and-ndbmtd"}}}]`, Hopsworks overrides the `rondb` chart default `[]`. `rondb.rondb.networkPolicy.mgmds.ingressSelectors[].namespaceSelector` # { #helm.rondb.rondb.networkPolicy.mgmds.ingressSelectors.namespaceSelector } : Type `object`. `rondb.rondb.networkPolicy.mgmds.ingressSelectors[].namespaceSelector.matchExpressions` # { #helm.rondb.rondb.networkPolicy.mgmds.ingressSelectors.namespaceSelector.matchExpressions } : Type `array`, default `[]`. `rondb.rondb.networkPolicy.mgmds.ingressSelectors[].namespaceSelector.matchExpressions[].key` # { #helm.rondb.rondb.networkPolicy.mgmds.ingressSelectors.namespaceSelector.matchExpressions.key } : Type `string`. `rondb.rondb.networkPolicy.mgmds.ingressSelectors[].namespaceSelector.matchExpressions[].operator` # { #helm.rondb.rondb.networkPolicy.mgmds.ingressSelectors.namespaceSelector.matchExpressions.operator } : Type `string`. `rondb.rondb.networkPolicy.mgmds.ingressSelectors[].namespaceSelector.matchExpressions[].values` # { #helm.rondb.rondb.networkPolicy.mgmds.ingressSelectors.namespaceSelector.matchExpressions.values } : Type `array`. `rondb.rondb.networkPolicy.mgmds.ingressSelectors[].namespaceSelector.matchLabels` # { #helm.rondb.rondb.networkPolicy.mgmds.ingressSelectors.namespaceSelector.matchLabels } : Type `object`, default `{}`, example `{"app":"rondb"}`. `rondb.rondb.networkPolicy.mgmds.ingressSelectors[].podSelector` # { #helm.rondb.rondb.networkPolicy.mgmds.ingressSelectors.podSelector } : Type `object`. `rondb.rondb.networkPolicy.mgmds.ingressSelectors[].podSelector.matchExpressions` # { #helm.rondb.rondb.networkPolicy.mgmds.ingressSelectors.podSelector.matchExpressions } : Type `array`, default `[]`. `rondb.rondb.networkPolicy.mgmds.ingressSelectors[].podSelector.matchExpressions[].key` # { #helm.rondb.rondb.networkPolicy.mgmds.ingressSelectors.podSelector.matchExpressions.key } : Type `string`. `rondb.rondb.networkPolicy.mgmds.ingressSelectors[].podSelector.matchExpressions[].operator` # { #helm.rondb.rondb.networkPolicy.mgmds.ingressSelectors.podSelector.matchExpressions.operator } : Type `string`. `rondb.rondb.networkPolicy.mgmds.ingressSelectors[].podSelector.matchExpressions[].values` # { #helm.rondb.rondb.networkPolicy.mgmds.ingressSelectors.podSelector.matchExpressions.values } : Type `array`. `rondb.rondb.networkPolicy.mgmds.ingressSelectors[].podSelector.matchLabels` # { #helm.rondb.rondb.networkPolicy.mgmds.ingressSelectors.podSelector.matchLabels } : Type `object`, default `{}`, example `{"app":"rondb"}`. `rondb.rondb.networkPolicy.ndbmtds` # { #helm.rondb.rondb.networkPolicy.ndbmtds } : Type `object`. `rondb.rondb.networkPolicy.ndbmtds.enabled` # { #helm.rondb.rondb.networkPolicy.ndbmtds.enabled } : Type `boolean`, default `true`. Whether to limit ingress for data node pods. If there is an empty API slot in the config.ini, any host can connect to them. `rondb.rondb.networkPolicy.ndbmtds.ingressSelectors` # { #helm.rondb.rondb.networkPolicy.ndbmtds.ingressSelectors } : Type `array`, default `[{"podSelector":{"matchLabels":{"access":"mgmd-and-ndbmtd"}}}]`, Hopsworks overrides the `rondb` chart default `[]`. `rondb.rondb.networkPolicy.ndbmtds.ingressSelectors[].namespaceSelector` # { #helm.rondb.rondb.networkPolicy.ndbmtds.ingressSelectors.namespaceSelector } : Type `object`. `rondb.rondb.networkPolicy.ndbmtds.ingressSelectors[].namespaceSelector.matchExpressions` # { #helm.rondb.rondb.networkPolicy.ndbmtds.ingressSelectors.namespaceSelector.matchExpressions } : Type `array`, default `[]`. `rondb.rondb.networkPolicy.ndbmtds.ingressSelectors[].namespaceSelector.matchExpressions[].key` # { #helm.rondb.rondb.networkPolicy.ndbmtds.ingressSelectors.namespaceSelector.matchExpressions.key } : Type `string`. `rondb.rondb.networkPolicy.ndbmtds.ingressSelectors[].namespaceSelector.matchExpressions[].operator` # { #helm.rondb.rondb.networkPolicy.ndbmtds.ingressSelectors.namespaceSelector.matchExpressions.operator } : Type `string`. `rondb.rondb.networkPolicy.ndbmtds.ingressSelectors[].namespaceSelector.matchExpressions[].values` # { #helm.rondb.rondb.networkPolicy.ndbmtds.ingressSelectors.namespaceSelector.matchExpressions.values } : Type `array`. `rondb.rondb.networkPolicy.ndbmtds.ingressSelectors[].namespaceSelector.matchLabels` # { #helm.rondb.rondb.networkPolicy.ndbmtds.ingressSelectors.namespaceSelector.matchLabels } : Type `object`, default `{}`, example `{"app":"rondb"}`. `rondb.rondb.networkPolicy.ndbmtds.ingressSelectors[].podSelector` # { #helm.rondb.rondb.networkPolicy.ndbmtds.ingressSelectors.podSelector } : Type `object`. `rondb.rondb.networkPolicy.ndbmtds.ingressSelectors[].podSelector.matchExpressions` # { #helm.rondb.rondb.networkPolicy.ndbmtds.ingressSelectors.podSelector.matchExpressions } : Type `array`, default `[]`. `rondb.rondb.networkPolicy.ndbmtds.ingressSelectors[].podSelector.matchExpressions[].key` # { #helm.rondb.rondb.networkPolicy.ndbmtds.ingressSelectors.podSelector.matchExpressions.key } : Type `string`. `rondb.rondb.networkPolicy.ndbmtds.ingressSelectors[].podSelector.matchExpressions[].operator` # { #helm.rondb.rondb.networkPolicy.ndbmtds.ingressSelectors.podSelector.matchExpressions.operator } : Type `string`. `rondb.rondb.networkPolicy.ndbmtds.ingressSelectors[].podSelector.matchExpressions[].values` # { #helm.rondb.rondb.networkPolicy.ndbmtds.ingressSelectors.podSelector.matchExpressions.values } : Type `array`. `rondb.rondb.networkPolicy.ndbmtds.ingressSelectors[].podSelector.matchLabels` # { #helm.rondb.rondb.networkPolicy.ndbmtds.ingressSelectors.podSelector.matchLabels } : Type `object`, default `{}`, example `{"app":"rondb"}`.
### nodeSelector { #helm-values-rondb-rondb-nodeselector } ??? example "Defaults as YAML" ```yaml rondb: rondb: nodeSelector: backup: {} mgmd: {} mysqld: {} ndbmtd: {} rdrs: {} ```
`rondb.rondb.nodeSelector` # { #helm.rondb.rondb.nodeSelector } : Type `object`. This ensures that Kubernetes schedules pods only onto nodes that match all the specified labels. `rondb.rondb.nodeSelector.backup` # { #helm.rondb.rondb.nodeSelector.backup } : Type `object`, default `{}`. `rondb.rondb.nodeSelector.mgmd` # { #helm.rondb.rondb.nodeSelector.mgmd } : Type `object`, default `{}`. `rondb.rondb.nodeSelector.mysqld` # { #helm.rondb.rondb.nodeSelector.mysqld } : Type `object`, default `{}`. `rondb.rondb.nodeSelector.ndbmtd` # { #helm.rondb.rondb.nodeSelector.ndbmtd } : Type `object`, default `{}`. `rondb.rondb.nodeSelector.rdrs` # { #helm.rondb.rondb.nodeSelector.rdrs } : Type `object`, default `{}`.
### podDisruptionBudget { #helm-values-rondb-rondb-poddisruptionbudget } ??? example "Defaults as YAML" ```yaml rondb: rondb: podDisruptionBudget: mysqld: enabled: true minAvailable: 1 ndbmtd: enabled: true minAvailable: 1 rdrs: enabled: true minAvailable: 1 ```
`rondb.rondb.podDisruptionBudget` # { #helm.rondb.rondb.podDisruptionBudget } : Type `object`. PodDisruptionBudget configuration per service. Ensures pod availability during voluntary disruptions like node drains and EKS node group upgrades. For single-replica deployments (effective replicas == 1), the PDB is not created even if minAvailable is 1, to avoid blocking node drains. For multi-replica deployments (effective replicas > 1), template rendering will fail if minAvailable is greater than or equal to the effective replica count, because such a PDB would prevent voluntary disruptions. The effective replica count used in these checks is the configured minimum replica count for each service (for example, clusterSize.minNumMySQLServers / clusterSize.minNumRdrs), not any dynamically autoscaled replica count at runtime. `rondb.rondb.podDisruptionBudget.mysqld` # { #helm.rondb.rondb.podDisruptionBudget.mysqld } : Type `object`. `rondb.rondb.podDisruptionBudget.mysqld.enabled` # { #helm.rondb.rondb.podDisruptionBudget.mysqld.enabled } : Type `boolean`, default `true`. `rondb.rondb.podDisruptionBudget.mysqld.minAvailable` # { #helm.rondb.rondb.podDisruptionBudget.mysqld.minAvailable } : Type `integer`, default `1`, minimum `1`. `rondb.rondb.podDisruptionBudget.ndbmtd` # { #helm.rondb.rondb.podDisruptionBudget.ndbmtd } : Type `object`. `rondb.rondb.podDisruptionBudget.ndbmtd.enabled` # { #helm.rondb.rondb.podDisruptionBudget.ndbmtd.enabled } : Type `boolean`, default `true`. `rondb.rondb.podDisruptionBudget.ndbmtd.minAvailable` # { #helm.rondb.rondb.podDisruptionBudget.ndbmtd.minAvailable } : Type `integer`, default `1`, minimum `1`. `rondb.rondb.podDisruptionBudget.rdrs` # { #helm.rondb.rondb.podDisruptionBudget.rdrs } : Type `object`. `rondb.rondb.podDisruptionBudget.rdrs.enabled` # { #helm.rondb.rondb.podDisruptionBudget.rdrs.enabled } : Type `boolean`, default `true`. `rondb.rondb.podDisruptionBudget.rdrs.minAvailable` # { #helm.rondb.rondb.podDisruptionBudget.rdrs.minAvailable } : Type `integer`, default `1`, minimum `1`.
### rdrs { #helm-values-rondb-rondb-rdrs } ??? example "Defaults as YAML" ```yaml rondb: rondb: rdrs: externalMetadataCluster: mgmds: [] slotsPerNode: 1 hpa: additionalMetrics: [] maxKeepaliveRequests: 0 probePort: enabled: true port: 4407 probes: liveness: failureThreshold: 12 initialDelaySeconds: 5 periodSeconds: 10 timeoutSeconds: 5 readiness: failureThreshold: 3 initialDelaySeconds: 5 periodSeconds: 5 timeoutSeconds: 3 startup: failureThreshold: 11 initialDelaySeconds: 5 periodSeconds: 5 timeoutSeconds: 2 security: apiKey: cacheRefreshIntervalMS: 180000 ttlPurge: activeWindow: null enable: null uploadPath: /tmp/rdrs-uploads ```
`rondb.rondb.rdrs` # { #helm.rondb.rondb.rdrs } : Type `object`. `rondb.rondb.rdrs.externalMetadataCluster` # { #helm.rondb.rondb.rdrs.externalMetadataCluster } : Type `object`. RDRSs will always be in the cluster of the data, not the metadata `rondb.rondb.rdrs.externalMetadataCluster.mgmds` # { #helm.rondb.rondb.rdrs.externalMetadataCluster.mgmds } : Type `array`, default `[]`. `rondb.rondb.rdrs.externalMetadataCluster.mgmds[].ip` # { #helm.rondb.rondb.rdrs.externalMetadataCluster.mgmds.ip } : Type `string`. `rondb.rondb.rdrs.externalMetadataCluster.mgmds[].port` # { #helm.rondb.rondb.rdrs.externalMetadataCluster.mgmds.port } : Type `integer`, default `1186`. `rondb.rondb.rdrs.externalMetadataCluster.slotsPerNode` # { #helm.rondb.rondb.rdrs.externalMetadataCluster.slotsPerNode } : Type `integer`, default `1`, minimum `1`, maximum `1`. `rondb.rondb.rdrs.hpa` # { #helm.rondb.rondb.rdrs.hpa } : Type `object`. Horizontal Pod Autoscaler for RDRS `rondb.rondb.rdrs.hpa.additionalMetrics` # { #helm.rondb.rondb.rdrs.hpa.additionalMetrics } : Type `array`, default `[]`. Additional metrics to use for the HPA. This is useful for custom metrics that are not supported by default. `rondb.rondb.rdrs.maxKeepaliveRequests` # { #helm.rondb.rondb.rdrs.maxKeepaliveRequests } : Type `integer`, default `0`, minimum `0`, maximum `4294967295`, example `1000`. Maximum number of requests served on one keep-alive connection to the RDRS main port; after it the connection is closed gracefully (Connection: close). 0 (the default) disables the limit. A Kubernetes Service balances per TCP connection, so long-lived connections keep the skew that builds up after a rolling restart; bounding their lifetime lets clients re-balance. Each reconnect costs a TCP+TLS handshake: use a high value (1000 recommended, not below 500) and only where post-restart skew is observed. Does not apply to the probe port. Requires an RDRS image with REST.MaxKeepaliveRequests support (releases 26.02.11 and newer on the 26.02 line); at 0 the key is not emitted, so older images keep working. `rondb.rondb.rdrs.probePort` # { #helm.rondb.rondb.rdrs.probePort } : Type `object`. Dedicated RDRS probe listener. It serves only the ping and health endpoints, from its own thread, so Kubernetes probes keep being answered while every worker thread is blocked on data-node operations (a stalled data node otherwise fails the liveness probe of all RDRS pods at once). The probe port performs NO authentication, regardless of the PingRequiresAuth/HealthRequiresAuth settings. It is not published as a port of any Service or Ingress, but like any pod port it is reachable via pod IPs and the headless Service's DNS records; where isolation is required, enforce it with a NetworkPolicy. Requires an RDRS image with REST.ProbePort support (releases 26.02.11 and newer on the 26.02 line): older images reject the unknown config keys at startup, so set enabled to false for pinned older images. Trade-off: the probe port answers as long as the process lives, so an RDRS whose worker threads are permanently wedged is not restarted by liveness. `rondb.rondb.rdrs.probePort.enabled` # { #helm.rondb.rondb.rdrs.probePort.enabled } : Type `boolean`, default `true`. Serve ping/health on the dedicated probe port and point the startup, liveness and readiness probes at it (ping answers 503 until the main port accepts connections, so startup semantics are unchanged). false emits no probe configuration keys and points all probes back at the main port: the escape hatch for pinned RDRS images that predate REST.ProbePort. `rondb.rondb.rdrs.probePort.port` # { #helm.rondb.rondb.rdrs.probePort.port } : Type `integer`, default `4407`, minimum `1`, maximum `65535`. TCP port of the dedicated probe listener. Must differ from the main REST port (4406). `rondb.rondb.rdrs.probes` # { #helm.rondb.rondb.rdrs.probes } : Type `object`. Timings of the RDRS container probes; the HTTP path, port and scheme are set by the chart. When installed through the Hopsworks chart, the path is rondb.rondb.rdrs.probes. A probe gives up after roughly (failureThreshold - 1) * periodSeconds + timeoutSeconds of consecutive failures. `rondb.rondb.rdrs.probes.liveness` # { #helm.rondb.rondb.rdrs.probes.liveness } : Type `object`. Checks /ping and restarts RDRS when it keeps failing. With rdrs.probePort.enabled (the default) /ping is answered from the dedicated probe thread, which keeps responding through data-node failures, so the defaults are ample. With probePort.enabled=false /ping shares the worker threads: while a failed data node has not yet been declared dead, requests touching it block and /ping cannot answer. Detection takes ~25 seconds per data node, so in that mode size (failureThreshold - 1) * periodSeconds + timeoutSeconds above ~25 seconds per data node that can go silent at once, plus ~40 seconds. The defaults give ~115 seconds; restarting the data nodes of 8 node groups in parallel needs failureThreshold 30 (~295 seconds). `rondb.rondb.rdrs.probes.liveness.failureThreshold` # { #helm.rondb.rondb.rdrs.probes.liveness.failureThreshold } : Type `integer`, default `12`, minimum `1`. `rondb.rondb.rdrs.probes.liveness.initialDelaySeconds` # { #helm.rondb.rondb.rdrs.probes.liveness.initialDelaySeconds } : Type `integer`, default `5`, minimum `0`. `rondb.rondb.rdrs.probes.liveness.periodSeconds` # { #helm.rondb.rondb.rdrs.probes.liveness.periodSeconds } : Type `integer`, default `10`, minimum `1`. `rondb.rondb.rdrs.probes.liveness.timeoutSeconds` # { #helm.rondb.rondb.rdrs.probes.liveness.timeoutSeconds } : Type `integer`, default `5`, minimum `1`. `rondb.rondb.rdrs.probes.readiness` # { #helm.rondb.rondb.rdrs.probes.readiness } : Type `object`. Checks /health. A pod that keeps failing is removed from the Service after ~13 to 18 seconds with the defaults; connections it already holds stay open. More than one failure is required so that a single slow check under load does not take the pod out of the Service. `rondb.rondb.rdrs.probes.readiness.failureThreshold` # { #helm.rondb.rondb.rdrs.probes.readiness.failureThreshold } : Type `integer`, default `3`, minimum `1`. `rondb.rondb.rdrs.probes.readiness.initialDelaySeconds` # { #helm.rondb.rondb.rdrs.probes.readiness.initialDelaySeconds } : Type `integer`, default `5`, minimum `0`. `rondb.rondb.rdrs.probes.readiness.periodSeconds` # { #helm.rondb.rondb.rdrs.probes.readiness.periodSeconds } : Type `integer`, default `5`, minimum `1`. `rondb.rondb.rdrs.probes.readiness.timeoutSeconds` # { #helm.rondb.rondb.rdrs.probes.readiness.timeoutSeconds } : Type `integer`, default `3`, minimum `1`. `rondb.rondb.rdrs.probes.startup` # { #helm.rondb.rondb.rdrs.probes.startup } : Type `object`. Checks /ping until RDRS first answers; liveness and readiness only start after it passes. RDRS opens its port once it has preloaded its caches from RonDB, about 25 seconds with a few hundred feature views. Raise failureThreshold if that preload takes longer than the ~57 seconds allowed. `rondb.rondb.rdrs.probes.startup.failureThreshold` # { #helm.rondb.rondb.rdrs.probes.startup.failureThreshold } : Type `integer`, default `11`, minimum `1`. `rondb.rondb.rdrs.probes.startup.initialDelaySeconds` # { #helm.rondb.rondb.rdrs.probes.startup.initialDelaySeconds } : Type `integer`, default `5`, minimum `0`. `rondb.rondb.rdrs.probes.startup.periodSeconds` # { #helm.rondb.rondb.rdrs.probes.startup.periodSeconds } : Type `integer`, default `5`, minimum `1`. `rondb.rondb.rdrs.probes.startup.timeoutSeconds` # { #helm.rondb.rondb.rdrs.probes.startup.timeoutSeconds } : Type `integer`, default `2`, minimum `1`. `rondb.rondb.rdrs.security` # { #helm.rondb.rondb.rdrs.security } : Type `object`. `rondb.rondb.rdrs.security.apiKey` # { #helm.rondb.rondb.rdrs.security.apiKey } : Type `object`. `rondb.rondb.rdrs.security.apiKey.cacheRefreshIntervalMS` # { #helm.rondb.rondb.rdrs.security.apiKey.cacheRefreshIntervalMS } : Type `integer`, default `180000`, minimum `1000`. How often the API key cache refreshes project associations from the database (in milliseconds). Lower values reduce staleness when project memberships change but increase database load. `rondb.rondb.rdrs.ttlPurge` # { #helm.rondb.rondb.rdrs.ttlPurge } : Type `object`. TTL purge worker settings of every RDRS pod, rendered as the TTLPurge section of rest_api.json. The section is rendered only when at least one field is set, because RDRS older than RonDB 26.02.9 rejects the TTLPurge key and does not start. Changes reach running pods only when they restart; a cluster-wide window in the mysql.ttl_purge_ctrl table takes precedence over activeWindow and needs no restart. When installed through the Hopsworks chart, the path is rondb.rondb.rdrs.ttlPurge. Unknown fields fail the render, so a misspelled field cannot silently leave purging running around the clock. `rondb.rondb.rdrs.ttlPurge.activeWindow` # { #helm.rondb.rondb.rdrs.ttlPurge.activeWindow } : Type `string|null`, default `null`, pattern `^$|^([01][0-9]|2[0-3]):[0-5][0-9]-([01][0-9]|2[0-3]):[0-5][0-9]$`. Daily UTC window during which the TTL purge worker deletes expired rows, formatted "HH:MM-HH:MM" (e.g. "03:00-05:00"); it wraps past midnight when start > end (e.g. "23:00-02:00"). Start and end must differ. When null or empty, purging runs around the clock. A valid window in mysql.ttl_purge_ctrl (ctrl_id 2/3) takes precedence. `rondb.rondb.rdrs.ttlPurge.enable` # { #helm.rondb.rondb.rdrs.ttlPurge.enable } : Type `boolean|null`, default `null`. Whether the RDRS pods run the TTL purge worker, which deletes expired rows of TTL tables. When null, RDRS decides (enabled). Can still be changed per pod at runtime through PUT /0.1.0/ttl-purge/config, until the pod restarts. `rondb.rondb.rdrs.uploadPath` # { #helm.rondb.rondb.rdrs.uploadPath } : Type `string`, default `"/tmp/rdrs-uploads"`. Writable directory where RDRS buffers HTTP request bodies larger than 64KiB. The container's working directory is not writable (RDRS runs as uid 1000): without this, every startup logs 256 'Permission denied' errors and oversized bodies are silently read as empty. Requires an RDRS image with REST.UploadPath support (releases 26.02.11 and newer on the 26.02 line); set to the empty string for pinned older images, which reject the unknown key.
### resources { #helm-values-rondb-rondb-resources } ??? example "Defaults as YAML" ```yaml rondb: rondb: resources: limits: cpus: benchs: 2 mgmds: 0.2 mysqldExporters: 0.2 mysqlds: 2 ndbmtds: 2 rdrs: 2 restore: 1 memory: benchsMiB: 500 mysqldExportersMiB: 100 mysqldMiB: 1400 ndbmtdsMiB: 5000 rdrsMiB: 500 requests: cpus: benchs: 1 mgmds: 0.2 mysqldExporters: 0.02 mysqlds: 1 rdrs: 1 memory: benchsMiB: 100 mysqldExportersMiB: 50 mysqldMiB: 650 rdrsMiB: 100 storage: binlogGiB: 4 classes: binlogFiles: null default: null diskColumns: null mgmd: null diskColumnGiB: 2 logGiB: 2 mgmdGiB: 1 ndbmtdGiB: 30 redoLogGiB: 4 relayLogGiB: 2 undoLogsGiB: 4 ```
`rondb.rondb.resources` # { #helm.rondb.rondb.resources } : Type `object`. Vertical cluster size `rondb.rondb.resources.limits` # { #helm.rondb.rondb.resources.limits } : Type `object`. Kubernetes resource limits `rondb.rondb.resources.limits.cpus` # { #helm.rondb.rondb.resources.limits.cpus } : Type `object`. CPU resources per RonDB service type `rondb.rondb.resources.limits.cpus.benchs` # { #helm.rondb.rondb.resources.limits.cpus.benchs } : Type `number`, default `2`. `rondb.rondb.resources.limits.cpus.mgmds` # { #helm.rondb.rondb.resources.limits.cpus.mgmds } : Type `number`, default `0.2`. `rondb.rondb.resources.limits.cpus.mysqldExporters` # { #helm.rondb.rondb.resources.limits.cpus.mysqldExporters } : Type `number`, default `0.2`. `rondb.rondb.resources.limits.cpus.mysqlds` # { #helm.rondb.rondb.resources.limits.cpus.mysqlds } : Type `number`, default `2`. `rondb.rondb.resources.limits.cpus.ndbmtds` # { #helm.rondb.rondb.resources.limits.cpus.ndbmtds } : Type `number`, default `2`. `rondb.rondb.resources.limits.cpus.rdrs` # { #helm.rondb.rondb.resources.limits.cpus.rdrs } : Type `number`, default `2`. `rondb.rondb.resources.limits.cpus.restore` # { #helm.rondb.rondb.resources.limits.cpus.restore } : Type `number`, default `1`. `rondb.rondb.resources.limits.memory` # { #helm.rondb.rondb.resources.limits.memory } : Type `object`. Memory resources per RonDB service type `rondb.rondb.resources.limits.memory.benchsMiB` # { #helm.rondb.rondb.resources.limits.memory.benchsMiB } : Type `integer`, default `500`. `rondb.rondb.resources.limits.memory.mysqldExportersMiB` # { #helm.rondb.rondb.resources.limits.memory.mysqldExportersMiB } : Type `integer`, default `100`. `rondb.rondb.resources.limits.memory.mysqldMiB` # { #helm.rondb.rondb.resources.limits.memory.mysqldMiB } : Type `integer`, default `1400`. This can usually be kept at the default independent of the load `rondb.rondb.resources.limits.memory.ndbmtdsMiB` # { #helm.rondb.rondb.resources.limits.memory.ndbmtdsMiB } : Type `integer`, default `5000`, minimum `2800`. It is recommended to keep this above 5GiB, otherwise some memory parts will be configured manually. `rondb.rondb.resources.limits.memory.rdrsMiB` # { #helm.rondb.rondb.resources.limits.memory.rdrsMiB } : Type `integer`, default `500`. `rondb.rondb.resources.requests` # { #helm.rondb.rondb.resources.requests } : Type `object`. Kubernetes resource requests; Note that data nodes will only apply limits, not requests `rondb.rondb.resources.requests.cpus` # { #helm.rondb.rondb.resources.requests.cpus } : Type `object`. CPU resources per RonDB service type `rondb.rondb.resources.requests.cpus.benchs` # { #helm.rondb.rondb.resources.requests.cpus.benchs } : Type `number`, default `1`. `rondb.rondb.resources.requests.cpus.mgmds` # { #helm.rondb.rondb.resources.requests.cpus.mgmds } : Type `number`, default `0.2`. `rondb.rondb.resources.requests.cpus.mysqldExporters` # { #helm.rondb.rondb.resources.requests.cpus.mysqldExporters } : Type `number`, default `0.02`. `rondb.rondb.resources.requests.cpus.mysqlds` # { #helm.rondb.rondb.resources.requests.cpus.mysqlds } : Type `number`, default `1`. `rondb.rondb.resources.requests.cpus.rdrs` # { #helm.rondb.rondb.resources.requests.cpus.rdrs } : Type `number`, default `1`. `rondb.rondb.resources.requests.memory` # { #helm.rondb.rondb.resources.requests.memory } : Type `object`. Memory resources per RonDB service type `rondb.rondb.resources.requests.memory.benchsMiB` # { #helm.rondb.rondb.resources.requests.memory.benchsMiB } : Type `integer`, default `100`. `rondb.rondb.resources.requests.memory.mysqldExportersMiB` # { #helm.rondb.rondb.resources.requests.memory.mysqldExportersMiB } : Type `integer`, default `50`. `rondb.rondb.resources.requests.memory.mysqldMiB` # { #helm.rondb.rondb.resources.requests.memory.mysqldMiB } : Type `integer`, default `650`. This can usually be kept at the default independent of the load `rondb.rondb.resources.requests.memory.rdrsMiB` # { #helm.rondb.rondb.resources.requests.memory.rdrsMiB } : Type `integer`, default `100`. `rondb.rondb.resources.requests.storage` # { #helm.rondb.rondb.resources.requests.storage } : Type `object`. Volume specifications `rondb.rondb.resources.requests.storage.binlogGiB` # { #helm.rondb.rondb.resources.requests.storage.binlogGiB } : Type `integer`, default `4`, minimum `1`. Keep in mind that this size needs to survive the retention period of the binlog files `rondb.rondb.resources.requests.storage.classes` # { #helm.rondb.rondb.resources.requests.storage.classes } : Type `object`. Storage classes `rondb.rondb.resources.requests.storage.classes.binlogFiles` # { #helm.rondb.rondb.resources.requests.storage.classes.binlogFiles } : Type `string|null`, default `null`. Storage class name for MySQLd binlog volumes in global replication `rondb.rondb.resources.requests.storage.classes.default` # { #helm.rondb.rondb.resources.requests.storage.classes.default } : Type `string|null`, default `null`. Default storage class name for all volumes `rondb.rondb.resources.requests.storage.classes.diskColumns` # { #helm.rondb.rondb.resources.requests.storage.classes.diskColumns } : Type `string|null`, default `null`. Storage class name for the data node disk columns volume `rondb.rondb.resources.requests.storage.classes.mgmd` # { #helm.rondb.rondb.resources.requests.storage.classes.mgmd } : Type `string|null`, default `null`. Storage class name for Management server `rondb.rondb.resources.requests.storage.diskColumnGiB` # { #helm.rondb.rondb.resources.requests.storage.diskColumnGiB } : Type `integer`, default `2`, minimum `1`. This depends on how much data the user is expecting to place on disk `rondb.rondb.resources.requests.storage.logGiB` # { #helm.rondb.rondb.resources.requests.storage.logGiB } : Type `integer`, default `2`. `rondb.rondb.resources.requests.storage.mgmdGiB` # { #helm.rondb.rondb.resources.requests.storage.mgmdGiB } : Type `integer`, default `1`, minimum `1`. The size of the statefulSet persistent volume for the RonDB management node. It will be used in new installations or in case of lookup function failing during upgrade in argo deployments. `rondb.rondb.resources.requests.storage.ndbmtdGiB` # { #helm.rondb.rondb.resources.requests.storage.ndbmtdGiB } : Type `integer`, default `30`, minimum `0`. The size of the statefulSet persistent volume for the RonDB data nodes. It will be used in new installations or in case of lookup function failing during upgrade in argo deployments. `rondb.rondb.resources.requests.storage.redoLogGiB` # { #helm.rondb.rondb.resources.requests.storage.redoLogGiB } : Type `integer`, default `4`, minimum `2`, maximum `64`. 64GiB is recommended for optimal performance `rondb.rondb.resources.requests.storage.relayLogGiB` # { #helm.rondb.rondb.resources.requests.storage.relayLogGiB } : Type `integer`, default `2`, minimum `1`. `rondb.rondb.resources.requests.storage.undoLogsGiB` # { #helm.rondb.rondb.resources.requests.storage.undoLogsGiB } : Type `integer`, default `4`, minimum `1`, maximum `128`. 64GiB is recommended for good performance
### restoreFromBackup { #helm-values-rondb-rondb-restorefrombackup } ??? example "Defaults as YAML" ```yaml rondb: rondb: restoreFromBackup: backupId: null excludeDatabases: [] excludeTables: [] forceDataClear: null inPlace: null objectStorageProvider: s3 pathPrefix: rondb_backup s3: bucketName: null endpoint: null keyCredentialsSecret: key: null name: null provider: null region: null secretCredentialsSecret: key: null name: null serverSideEncryption: null ```
`rondb.rondb.restoreFromBackup` # { #helm.rondb.rondb.restoreFromBackup } : Type `object`. Whether to restore a backup on the cluster `rondb.rondb.restoreFromBackup.backupId` # { #helm.rondb.rondb.restoreFromBackup.backupId } : Type `string|null`, default `null`. The native backup ID for the backup to restore `rondb.rondb.restoreFromBackup.excludeDatabases` # { #helm.rondb.rondb.restoreFromBackup.excludeDatabases } : Type `array`, default `[]`. Which databases to exclude from the native backup `rondb.rondb.restoreFromBackup.excludeTables` # { #helm.rondb.rondb.restoreFromBackup.excludeTables } : Type `array`, default `[]`. Which tables to exclude from the native backup. Use the format: database.table `rondb.rondb.restoreFromBackup.forceDataClear` # { #helm.rondb.rondb.restoreFromBackup.forceDataClear } : Type `boolean|null`, default `null`. Confirm that existing data will be destroyed during in-place restore. `rondb.rondb.restoreFromBackup.inPlace` # { #helm.rondb.rondb.restoreFromBackup.inPlace } : Type `boolean|null`, default `null`. Enable in-place restore on existing cluster. Requires forceDataClear=true. `rondb.rondb.restoreFromBackup.objectStorageProvider` # { #helm.rondb.rondb.restoreFromBackup.objectStorageProvider } : Type `enum`, default `"s3"`. One of: `"s3"`. `rondb.rondb.restoreFromBackup.pathPrefix` # { #helm.rondb.rondb.restoreFromBackup.pathPrefix } : Type `string`, default `"rondb_backup"`. Prefix of RonDB backup in the configured bucket `rondb.rondb.restoreFromBackup.s3` # { #helm.rondb.rondb.restoreFromBackup.s3 } : Type `object`. `rondb.rondb.restoreFromBackup.s3.bucketName` # { #helm.rondb.rondb.restoreFromBackup.s3.bucketName } : Type `string|null`, default `null`. `rondb.rondb.restoreFromBackup.s3.endpoint` # { #helm.rondb.rondb.restoreFromBackup.s3.endpoint } : Type `string|null`, default `null`. `rondb.rondb.restoreFromBackup.s3.keyCredentialsSecret` # { #helm.rondb.rondb.restoreFromBackup.s3.keyCredentialsSecret } : Type `object`. `rondb.rondb.restoreFromBackup.s3.keyCredentialsSecret.key` # { #helm.rondb.rondb.restoreFromBackup.s3.keyCredentialsSecret.key } : Type `string|null`, default `null`. Key in the Secret `rondb.rondb.restoreFromBackup.s3.keyCredentialsSecret.name` # { #helm.rondb.rondb.restoreFromBackup.s3.keyCredentialsSecret.name } : Type `string|null`, default `null`. Name of the Secret `rondb.rondb.restoreFromBackup.s3.provider` # { #helm.rondb.rondb.restoreFromBackup.s3.provider } : Type `string|null`, default `null`. `rondb.rondb.restoreFromBackup.s3.region` # { #helm.rondb.rondb.restoreFromBackup.s3.region } : Type `string|null`, default `null`. `rondb.rondb.restoreFromBackup.s3.secretCredentialsSecret` # { #helm.rondb.rondb.restoreFromBackup.s3.secretCredentialsSecret } : Type `object`. `rondb.rondb.restoreFromBackup.s3.secretCredentialsSecret.key` # { #helm.rondb.rondb.restoreFromBackup.s3.secretCredentialsSecret.key } : Type `string|null`, default `null`. Key in the Secret `rondb.rondb.restoreFromBackup.s3.secretCredentialsSecret.name` # { #helm.rondb.rondb.restoreFromBackup.s3.secretCredentialsSecret.name } : Type `string|null`, default `null`. Name of the Secret `rondb.rondb.restoreFromBackup.s3.serverSideEncryption` # { #helm.rondb.rondb.restoreFromBackup.s3.serverSideEncryption } : Type `enum`, default `null`. One of: `"aws:kms"`, `"aws:kms:dsse"`, `"AES256"`, `null`.
### rondbConfig { #helm-values-rondb-rondb-rondbconfig } ??? example "Defaults as YAML" ```yaml rondb: rondb: rondbConfig: ActivateRateLimits: 0 BackupLogBufferSize: 16M DataMemory: null DiskPageBufferMemory: null EmptyApiSlots: 8 FullRestartLogs: null HeartbeatIntervalDbApi: 5000 HeartbeatIntervalDbDb: 5000 InitialTablespaceSizeGiB: -1 LongMessageBuffer: null MaxDMLOperationsPerTransaction: 32768 MaxDiskWriteSpeed: null MaxNoOfAttributes: null MaxNoOfConcurrentOperations: 65536 MaxNoOfSchemaObjects: null MaxNoOfTables: null MaxNoOfTriggers: null MaxRRGroupSize: null MySQLdSlotsPerNode: 4 OsCpuOverhead: null OsStaticOverhead: null PartitionsPerNode: null RdrsMetadataSlotsPerNode: 1 RdrsSlotsPerNode: 1 RedoBuffer: null ReplicationMemory: null ReservedConcurrentOperations: null SchemaMemory: null SharedGlobalMemory: null TimeBetweenGlobalCheckpoints: null TotalMemoryConfig: null TransactionDeadlockDetectionTimeout: 1500 TransactionInactiveTimeout: 15000 TransactionMemory: null UseOnlyIPv4: null UseTcInRRGroup: null ```
`rondb.rondb.rondbConfig` # { #helm.rondb.rondb.rondbConfig } : Type `object`. Configurations for RonDB's config.ini. Memory configurations are in binary SI units (i.e. 1G = 1GiB = 1024MiB). `rondb.rondb.rondbConfig.ActivateRateLimits` # { #helm.rondb.rondb.rondbConfig.ActivateRateLimits } : Type `integer`, default `0`, minimum `0`, maximum `1`. Activate rate limits handling `rondb.rondb.rondbConfig.BackupLogBufferSize` # { #helm.rondb.rondb.rondbConfig.BackupLogBufferSize } : Type `string`, default `"16M"`. Buffer that records writes occurring during an online backup; the backup aborts if it fills. Default 16M; increase (e.g. 256M) for write-heavy clusters. See . `rondb.rondb.rondbConfig.DataMemory` # { #helm.rondb.rondb.rondbConfig.DataMemory } : Type `string|null`, default `null`, example `"4G"`. Memory available for storing in-memory database records on each data node, in bytes (with optional binary SI suffix, e.g. '4G'). When unset, RonDB's AutomaticMemoryConfig sizes this from the container's memory; only set explicitly to override automatic sizing. Only applied if explicitly set. `rondb.rondb.rondbConfig.DiskPageBufferMemory` # { #helm.rondb.rondb.rondbConfig.DiskPageBufferMemory } : Type `string|null`, default `null`. DiskPageBufferMemory are used by disk columns. This is the page cache that contains disk pages when they are in memory. By default it is not defined. `rondb.rondb.rondbConfig.EmptyApiSlots` # { #helm.rondb.rondb.rondbConfig.EmptyApiSlots } : Type `integer`, default `8`, minimum `1`. We need at least 1 for the MySQLd setup job. Otherwise this is for services that are not handled here, e.g. HopsFS `rondb.rondb.rondbConfig.FullRestartLogs` # { #helm.rondb.rondb.rondbConfig.FullRestartLogs } : Type `boolean|null`, default `null`. Enable full restart logs (RonDB-specific). RonDB's release-build default is false. `rondb.rondb.rondbConfig.HeartbeatIntervalDbApi` # { #helm.rondb.rondb.rondbConfig.HeartbeatIntervalDbApi } : Type `integer`, default `5000`. Each data node sends heartbeat signals to each MySQL server (SQL node) to ensure that it remains in contact. If a MySQL server fails to send a heartbeat in time it is declared “dead,” in which case all ongoing transactions are completed and all resources released. The SQL node cannot reconnect until all activities initiated by the previous MySQL instance have been completed. The three-heartbeat criteria for this determination are the same as described for HeartbeatIntervalDbDb. `rondb.rondb.rondbConfig.HeartbeatIntervalDbDb` # { #helm.rondb.rondb.rondbConfig.HeartbeatIntervalDbDb } : Type `integer`, default `5000`. One of the primary methods of discovering failed nodes is by the use of heartbeats. This parameter states how often heartbeat signals are sent and how often to expect to receive them. Heartbeats cannot be disabled. After missing four heartbeat intervals in a row, the node is declared dead. Thus, the maximum time for discovering a failure through the heartbeat mechanism is five times the heartbeat interval. `rondb.rondb.rondbConfig.InitialTablespaceSizeGiB` # { #helm.rondb.rondb.rondbConfig.InitialTablespaceSizeGiB } : Type `integer`, default `-1`. InitialTableSpace size in GiB. By default, it is set to -1, which enforces the initial tablespace to use the entire diskColumnGiB space. If set to 0, the same behaviour applies and the diskColumnGiB is used. `rondb.rondb.rondbConfig.LongMessageBuffer` # { #helm.rondb.rondb.rondbConfig.LongMessageBuffer } : Type `string|null`, default `null`, example `"64M"`. Internal buffer used for passing long messages within and between nodes, in bytes (with optional binary SI suffix). NDB's built-in default is 64M. Only applied if explicitly set. `rondb.rondb.rondbConfig.MaxDMLOperationsPerTransaction` # { #helm.rondb.rondb.rondbConfig.MaxDMLOperationsPerTransaction } : Type `integer`, default `32768`, minimum `32`, maximum `4294967039`. `rondb.rondb.rondbConfig.MaxDiskWriteSpeed` # { #helm.rondb.rondb.rondbConfig.MaxDiskWriteSpeed } : Type `string|null`, default `null`, example `"20M"`. Maximum disk write speed (bytes/sec, total across all ldm threads) used by local checkpoints during normal operation, i.e. when no node is restarting. The adaptive algorithm may write more slowly than this ceiling based on CPU usage and REDO log IO lag, but will not exceed it. Only applied if explicitly set; otherwise RonDB's built-in default is used. See . `rondb.rondb.rondbConfig.MaxNoOfAttributes` # { #helm.rondb.rondb.rondbConfig.MaxNoOfAttributes } : Type `integer|null`, default `null`. `rondb.rondb.rondbConfig.MaxNoOfConcurrentOperations` # { #helm.rondb.rondb.rondbConfig.MaxNoOfConcurrentOperations } : Type `integer`, default `65536`, minimum `32`, maximum `4294967039`. `rondb.rondb.rondbConfig.MaxNoOfSchemaObjects` # { #helm.rondb.rondb.rondbConfig.MaxNoOfSchemaObjects } : Type `integer|null`, default `null`, minimum `20320`, maximum `200000`. Maximum total number of schema objects (tables, ordered indexes, unique hash indexes, etc.) that a data node may allocate. Increase this when the application needs more than the default cap of 20,320 table objects. Only applied if explicitly set; otherwise RonDB's built-in default is used. See . `rondb.rondb.rondbConfig.MaxNoOfTables` # { #helm.rondb.rondb.rondbConfig.MaxNoOfTables } : Type `integer|null`, default `null`. `rondb.rondb.rondbConfig.MaxNoOfTriggers` # { #helm.rondb.rondb.rondbConfig.MaxNoOfTriggers } : Type `integer|null`, default `null`. `rondb.rondb.rondbConfig.MaxRRGroupSize` # { #helm.rondb.rondb.rondbConfig.MaxRRGroupSize } : Type `integer|null`, default `null`, minimum `8`, maximum `32`. Max size of a Round Robin (RR) group. RonDB default is 8; allowed range is 8 to 32. `rondb.rondb.rondbConfig.MySQLdSlotsPerNode` # { #helm.rondb.rondb.rondbConfig.MySQLdSlotsPerNode } : Type `integer`, default `4`, minimum `1`, maximum `4`. `rondb.rondb.rondbConfig.OsCpuOverhead` # { #helm.rondb.rondb.rondbConfig.OsCpuOverhead } : Type `string|null`, default `null`, example `"100M"`. Additional OS memory overhead for the RonDB datanode containers. This is multiplied by the number of CPUs. Also only set this if the container has less than 5GiB of memory. `rondb.rondb.rondbConfig.OsStaticOverhead` # { #helm.rondb.rondb.rondbConfig.OsStaticOverhead } : Type `string|null`, default `null`, example `"1400M"`. Memory overhead for the RonDB datanode containers. Only set this if the container has less than 5GiB of memory. `rondb.rondb.rondbConfig.PartitionsPerNode` # { #helm.rondb.rondb.rondbConfig.PartitionsPerNode } : Type `integer|null`, default `null`. Number of partitions per data node used when creating new tables. When null, RonDB chooses based on AutomaticThreadConfig / NumCPUs (typically the number of LDM threads). Only affects tables created after the change; existing tables are not repartitioned. `rondb.rondb.rondbConfig.RdrsMetadataSlotsPerNode` # { #helm.rondb.rondb.rondbConfig.RdrsMetadataSlotsPerNode } : Type `integer`, default `1`, minimum `1`, maximum `1`. We use additional cluster connections for metadata `rondb.rondb.rondbConfig.RdrsSlotsPerNode` # { #helm.rondb.rondb.rondbConfig.RdrsSlotsPerNode } : Type `integer`, default `1`, minimum `1`, maximum `1`. The number of cluster connections we support via the RDRS `rondb.rondb.rondbConfig.RedoBuffer` # { #helm.rondb.rondb.rondbConfig.RedoBuffer } : Type `string|null`, default `null`. `rondb.rondb.rondbConfig.ReplicationMemory` # { #helm.rondb.rondb.rondbConfig.ReplicationMemory } : Type `string|null`, default `null`. `rondb.rondb.rondbConfig.ReservedConcurrentOperations` # { #helm.rondb.rondb.rondbConfig.ReservedConcurrentOperations } : Type `integer|null`, default `null`. `rondb.rondb.rondbConfig.SchemaMemory` # { #helm.rondb.rondb.rondbConfig.SchemaMemory } : Type `string|null`, default `null`. `rondb.rondb.rondbConfig.SharedGlobalMemory` # { #helm.rondb.rondb.rondbConfig.SharedGlobalMemory } : Type `string|null`, default `null`. `rondb.rondb.rondbConfig.TimeBetweenGlobalCheckpoints` # { #helm.rondb.rondb.rondbConfig.TimeBetweenGlobalCheckpoints } : Type `integer|null`, default `null`, minimum `20`, maximum `32000`, example `2000`. Time in milliseconds between group commits of transactions to disk. RonDB's built-in default is 2000 ms. Only applied if explicitly set. `rondb.rondb.rondbConfig.TotalMemoryConfig` # { #helm.rondb.rondb.rondbConfig.TotalMemoryConfig } : Type `string|null`, default `null`. The total memory configured by RonDB datanode. By default it is not defined and memory is calculated automatically. `rondb.rondb.rondbConfig.TransactionDeadlockDetectionTimeout` # { #helm.rondb.rondb.rondbConfig.TransactionDeadlockDetectionTimeout } : Type `integer`, default `1500`. Time transaction can spend executing within data node `rondb.rondb.rondbConfig.TransactionInactiveTimeout` # { #helm.rondb.rondb.rondbConfig.TransactionInactiveTimeout } : Type `integer`, default `15000`. Milliseconds that application waits before executing another part of transaction `rondb.rondb.rondbConfig.TransactionMemory` # { #helm.rondb.rondb.rondbConfig.TransactionMemory } : Type `string|null`, default `null`. `rondb.rondb.rondbConfig.UseOnlyIPv4` # { #helm.rondb.rondb.rondbConfig.UseOnlyIPv4 } : Type `boolean|null`, default `null`. When true, restricts all RonDB cluster connections to IPv4 sockets only. Applied under both \[NDBD DEFAULT\] and \[MYSQLD DEFAULT\] so it covers data nodes, MySQLds (including binlog/replica appliers/DDL), RDRS, and other API clients. Introduced in RonDB 21.04.5 for Dolphin SuperSockets compatibility. `rondb.rondb.rondbConfig.UseTcInRRGroup` # { #helm.rondb.rondb.rondbConfig.UseTcInRRGroup } : Type `boolean|null`, default `null`. When true, each recv thread distributes connections only to TC threads within the same Round Robin (RR) group; when false, it distributes them across all TC threads. RonDB default is true.
### terminationGracePeriodSeconds { #helm-values-rondb-rondb-terminationgraceperiodseconds } ??? example "Defaults as YAML" ```yaml rondb: rondb: terminationGracePeriodSeconds: binlogServers: 30 mgmds: 30 mysqlds: 30 ndbmtds: 300 rdrs: 30 replicaAppliers: 30 ```
`rondb.rondb.terminationGracePeriodSeconds` # { #helm.rondb.rondb.terminationGracePeriodSeconds } : Type `object|integer`, minimum `10`. Pod terminationGracePeriodSeconds per RonDB service type. The chart's daemons stop on SIGTERM, so this is a ceiling, not a wait: a pod is removed as soon as its processes have exited. The legacy integer form (charts up to 26.2.19) still validates (minimum 10) and keeps its meaning: it overrides the data nodes only, every other component keeps 30. Helm drops null keys before schema validation: a null whole key falls back to the defaults, a null component key is rejected (all six keys are required; partial values files still work because Helm merges in the chart defaults). `rondb.rondb.terminationGracePeriodSeconds.binlogServers` # { #helm.rondb.rondb.terminationGracePeriodSeconds.binlogServers } : Type `integer`, default `30`, minimum `10`. Binlog server MySQLds; see mysqlds. `rondb.rondb.terminationGracePeriodSeconds.mgmds` # { #helm.rondb.rondb.terminationGracePeriodSeconds.mgmds } : Type `integer`, default `30`, minimum `10`. MGMd stops within seconds; this matches the Kubernetes default. `rondb.rondb.terminationGracePeriodSeconds.mysqlds` # { #helm.rondb.rondb.terminationGracePeriodSeconds.mysqlds } : Type `integer`, default `30`, minimum `10`. Used by MySQLds and DDL MySQLds; mysqld stops within seconds under no load, but allow time for open transactions to close. `rondb.rondb.terminationGracePeriodSeconds.ndbmtds` # { #helm.rondb.rondb.terminationGracePeriodSeconds.ndbmtds } : Type `integer`, default `300`, minimum `30`. Data nodes run a managed stop on SIGTERM: deactivate through the MGMd, then node shutdown with handover. Large nodes additionally need roughly 2 minutes per TB of data node memory for the kernel to tear the process down, so raise this for nodes above 1TB. Stops of several data nodes serialize at the MGMd. When deleting the entire cluster the MGMd may already be gone; data nodes then retry the deactivate for up to 60s before stopping directly. `rondb.rondb.terminationGracePeriodSeconds.rdrs` # { #helm.rondb.rondb.terminationGracePeriodSeconds.rdrs } : Type `integer`, default `30`, minimum `10`. RDRS stops within a few seconds; this matches the Kubernetes default. `rondb.rondb.terminationGracePeriodSeconds.replicaAppliers` # { #helm.rondb.rondb.terminationGracePeriodSeconds.replicaAppliers } : Type `integer`, default `30`, minimum `10`. Replica applier pods: the controller container stops its run_applier.sh worker on SIGTERM, the pod's mysqld receives SIGTERM directly.
### timeoutsMinutes { #helm-values-rondb-rondb-timeoutsminutes } ??? example "Defaults as YAML" ```yaml rondb: rondb: timeoutsMinutes: mysqldStartupProbe: 20 ndbmtdStartupProbe: 240 restoreNativeBackup: 120 singleSetupMySQLds: 5 ```
`rondb.rondb.timeoutsMinutes` # { #helm.rondb.rondb.timeoutsMinutes } : Type `object`. Using minutes for easier template addition `rondb.rondb.timeoutsMinutes.mysqldStartupProbe` # { #helm.rondb.rondb.timeoutsMinutes.mysqldStartupProbe } : Type `integer`, default `20`, minimum `1`. Maximum time a MySQLd pod (mysqld, ddl mysqld, binlog server, replica applier) is given to start up before Kubernetes restarts its container. `rondb.rondb.timeoutsMinutes.ndbmtdStartupProbe` # { #helm.rondb.rondb.timeoutsMinutes.ndbmtdStartupProbe } : Type `integer`, default `240`, minimum `1`. Maximum time a ndbmtd pod is given to start up before Kubernetes restarts its container. Node restarts with a lot of data to restore can take long. Must clear the worst-case data reload of the largest supported cluster - a node killed mid-recovery restarts from scratch and loops forever. `rondb.rondb.timeoutsMinutes.restoreNativeBackup` # { #helm.rondb.rondb.timeoutsMinutes.restoreNativeBackup } : Type `integer`, default `120`. This does not include *downloading* the native backups. IMPORTANT; Restoring native backups is NOT done in parallel yet; TODO: Make this dependent on amount of data `rondb.rondb.timeoutsMinutes.singleSetupMySQLds` # { #helm.rondb.rondb.timeoutsMinutes.singleSetupMySQLds } : Type `integer`, default `5`. This includes the time to start a single MySQLd pod, restoring MySQL metadata and running user-defined MySQL init scripts.
### tolerations { #helm-values-rondb-rondb-tolerations } ??? example "Defaults as YAML" ```yaml rondb: rondb: tolerations: backup: [] mgmd: [] mysqld: [] ndbmtd: [] rdrs: [] ```
`rondb.rondb.tolerations` # { #helm.rondb.rondb.tolerations } : Type `object`. These tolerations allow Kubernetes to schedule pods on nodes with matching taints, ensuring proper placement based on cluster policies. `rondb.rondb.tolerations.backup` # { #helm.rondb.rondb.tolerations.backup } : Type `array`, default `[]`. `rondb.rondb.tolerations.mgmd` # { #helm.rondb.rondb.tolerations.mgmd } : Type `array`, default `[]`. `rondb.rondb.tolerations.mysqld` # { #helm.rondb.rondb.tolerations.mysqld } : Type `array`, default `[]`. `rondb.rondb.tolerations.ndbmtd` # { #helm.rondb.rondb.tolerations.ndbmtd } : Type `array`, default `[]`. `rondb.rondb.tolerations.rdrs` # { #helm.rondb.rondb.tolerations.rdrs } : Type `array`, default `[]`.
================================================================================ # spark Source: https://docs.hopsworks.ai/latest/setup_installation/common/helm_chart_values/spark/ # Spark values { #helm-values-spark } Values under `spark` configure the Spark operator, the Spark history server and the remote shuffle service. _Generated from the Hopsworks Helm chart `5.2.0-alpha-1791549041` (Hopsworks `5.2.0`)._ Deployed when [`global._hopsworks.full_platform`](global.md#helm.global._hopsworks.full_platform) is `true`. !!! info "Upstream charts" - Values under `spark.spark-operator` go to [`spark-operator` 2.5.1](https://github.com/kubeflow/spark-operator/blob/v2.5.1/charts/spark-operator-chart/README.md) from `https://kubeflow.github.io/spark-operator/`. Only the values Hopsworks sets under `spark.spark-operator` are listed on this page. Any other value of the chart can be set there too; the link opens its documentation for the version Hopsworks pins. ## General { #helm-values-spark-general } ??? example "Defaults as YAML" ```yaml spark: cleanupOnUninstall: enabled: true ttlSecondsAfterFinished: null historyServer: certsDir: /srv/hops/super_crypto/spark cleaner: enabled: true interval: 1d maxAge: 7d deploymentName: spark-history-server-deployment hadoopHome: /srv/hops/hadoop image: pullPolicy: Always repository: sparkhistoryserver tag: 4.1.3.0 name: spark-history-server nodeSelector: {} probes: liveness: failureThreshold: 10 httpGet: path: / port: http scheme: HTTPS initialDelaySeconds: 8 periodSeconds: 5 timeoutSeconds: 10 readiness: failureThreshold: 10 httpGet: path: / port: http scheme: HTTPS initialDelaySeconds: 8 periodSeconds: 5 timeoutSeconds: 10 replicaCount: 1 resources: limits: cpu: 2000m memory: 3Gi requests: cpu: 500m memory: 1Gi service: annotations: consul.hashicorp.com/service-name: sparkhistoryserver externalPort: 80 internalPort: 18080 name: sparkhistoryserver type: ClusterIP sparkHome: /srv/hops/spark tolerations: [] topologySpreadConstraint: {} hopsworkslib: {} sparkJobDebugLevel: INFO sql: ansiEnabled: false ```
`spark` # { #helm.spark } : Type `object`, default `{}`. override spark values `spark.cleanupOnUninstall` # { #helm.spark.cleanupOnUninstall } : Type `object`, default `{"enabled":true,"ttlSecondsAfterFinished":null}`. post-delete cleanup of Spark/RSS leftovers: the spark-operator and rss-webhook TLS Secrets (generated at runtime by the operators, not Helm-tracked), and the rss shuffle-server data PVCs. PVCs are only deleted when global._hopsworks.wipeDataOnUninstall is enabled and never for PVCs labelled hopsworks.ai/keep=true. `spark.cleanupOnUninstall.enabled` # { #helm.spark.cleanupOnUninstall.enabled } : Type `bool`, default `true`. enable the post-delete Spark/RSS cleanup hook `spark.cleanupOnUninstall.ttlSecondsAfterFinished` # { #helm.spark.cleanupOnUninstall.ttlSecondsAfterFinished } : Type `string`, default `nil`. ttlSecondsAfterFinished for the cleanup Job; null falls through to the global default `spark.historyServer` # { #helm.spark.historyServer } : Type `object`. The configuration for the Spark history server ??? note "Default" ```yaml certsDir: /srv/hops/super_crypto/spark cleaner: enabled: true interval: 1d maxAge: 7d deploymentName: spark-history-server-deployment hadoopHome: /srv/hops/hadoop image: pullPolicy: Always repository: sparkhistoryserver tag: 4.1.3.0 name: spark-history-server nodeSelector: {} probes: liveness: failureThreshold: 10 httpGet: path: / port: http scheme: HTTPS initialDelaySeconds: 8 periodSeconds: 5 timeoutSeconds: 10 readiness: failureThreshold: 10 httpGet: path: / port: http scheme: HTTPS initialDelaySeconds: 8 periodSeconds: 5 timeoutSeconds: 10 replicaCount: 1 resources: limits: cpu: 2000m memory: 3Gi requests: cpu: 500m memory: 1Gi service: annotations: consul.hashicorp.com/service-name: sparkhistoryserver externalPort: 80 internalPort: 18080 name: sparkhistoryserver type: ClusterIP sparkHome: /srv/hops/spark tolerations: [] topologySpreadConstraint: {} ``` `spark.historyServer.nodeSelector` # { #helm.spark.historyServer.nodeSelector } : Type `object`, default `{}`. node selector configuration `spark.historyServer.topologySpreadConstraint` # { #helm.spark.historyServer.topologySpreadConstraint } : Type `object`, default `{}`. The default topology spread constraint. If not defined the global topology spread constraint would be used instead. `spark.hopsworkslib` # { #helm.spark.hopsworkslib } : Type `object`, default `{}`. override hopsworkslib values `spark.sparkJobDebugLevel` # { #helm.spark.sparkJobDebugLevel } : Type `string`, default `"INFO"`. `spark.sql` # { #helm.spark.sql } : Type `object`, default `{"ansiEnabled":false}`. Spark SQL defaults applied cluster-wide through `spark-defaults.conf`. `spark.sql.ansiEnabled` # { #helm.spark.sql.ansiEnabled } : Type `bool`, default `false`. Whether to enable ANSI SQL mode (`spark.sql.ansi.enabled`). Spark 4 changed the upstream default to `true`, which turns arithmetic overflow and invalid casts into runtime errors instead of returning `null`. Hopsworks pins it to `false` so SQL that ran on Spark 3.x keeps the same semantics after the upgrade. Set to `true` to opt in to standards-compliant behavior cluster-wide; individual jobs can still override `spark.sql.ansi.enabled` themselves.
## dependencies { #helm-values-spark-dependencies } ??? example "Defaults as YAML" ```yaml spark: dependencies: hive: consulServiceName: hive consulServiceTag: metastore port: 9083 namenode: consulServiceName: namenode consulServiceTag: rpc port: 8020 prometheusPushgateway: consulServiceName: prometheus consulServiceTag: pushgateway port: 9091 protocol: http ```
`spark.dependencies.hive.consulServiceName` # { #helm.spark.dependencies.hive.consulServiceName } : Type `string`, default `"hive"`. `spark.dependencies.hive.consulServiceTag` # { #helm.spark.dependencies.hive.consulServiceTag } : Type `string`, default `"metastore"`. `spark.dependencies.hive.port` # { #helm.spark.dependencies.hive.port } : Type `int`, default `9083`. `spark.dependencies.namenode.consulServiceName` # { #helm.spark.dependencies.namenode.consulServiceName } : Type `string`, default `"namenode"`. `spark.dependencies.namenode.consulServiceTag` # { #helm.spark.dependencies.namenode.consulServiceTag } : Type `string`, default `"rpc"`. `spark.dependencies.namenode.port` # { #helm.spark.dependencies.namenode.port } : Type `int`, default `8020`. `spark.dependencies.prometheusPushgateway.consulServiceName` # { #helm.spark.dependencies.prometheusPushgateway.consulServiceName } : Type `string`, default `"prometheus"`. `spark.dependencies.prometheusPushgateway.consulServiceTag` # { #helm.spark.dependencies.prometheusPushgateway.consulServiceTag } : Type `string`, default `"pushgateway"`. `spark.dependencies.prometheusPushgateway.port` # { #helm.spark.dependencies.prometheusPushgateway.port } : Type `int`, default `9091`. `spark.dependencies.prometheusPushgateway.protocol` # { #helm.spark.dependencies.prometheusPushgateway.protocol } : Type `string`, default `"http"`.
## rss { #helm-values-spark-rss } ??? example "Defaults as YAML" ```yaml spark: rss: appName: rss-hops configDir: /data/rssadmin/rss/conf configmap: name: rss-configuration controller: containerPort: 9876 image: rss-controller replicas: 1 resources: limits: cpu: '1' memory: 512Mi requests: cpu: 100m memory: 150Mi serviceAccount: annotations: {} coordinator: count: 2 dynamicClientConfigMapName: rss-dynamic-client-configuration dynamicClientConfigMountPath: /tmp httpPort: 19996 labels: role: rss-hops-coordinator replicas: 1 resources: limits: cpu: 500m memory: 2Gi requests: cpu: 500m rpcPort: 19997 xmxSizeMemoryExtraPercentage: 30 dashboard: enabled: false httpPort: 19997 name: uniffle-dashboard resources: limits: cpu: 500m memory: 2Gi requests: cpu: 250m xmxSizeMemoryExtraPercentage: 30 dynamicClient: readBufferSize: 14m storageType: MEMORY_LOCALFILE fullnameOverride: null image: image: rss initImage: hops-rss-init initImageVersion: 0.9.2 pullPolicy: Always namespaceSelector: '' nodeSelector: {} resources: jobs: limits: cpu: 200m rss: limits: cpu: 500m memory: 1Gi serviceAccount: annotations: {} shuffleServer: bufferCapacity: -1 bufferCapacityRatio: 0.6 diskCapacity: -1 diskCapacityRatio: 0.8 httpPort: 19998 nettyPort: 20000 readBufferCapacity: -1 readBufferCapacityRatio: 0.1 replicas: 3 resources: limits: cpu: 2000m memory: 3Gi requests: cpu: 200m rpcPort: 19999 upgradeStrategy: FullUpgrade storage: size: 10Gi storageClassName: null volumeNameTemplate: rss-storage tolerations: [] topologySpreadConstraint: {} ttlSecondsAfterFinished: null version: 0.11.1 webhook: app: rss-webhook resources: limits: cpu: '1' memory: 512Mi requests: cpu: 100m memory: 150Mi service: name: rss-webhook port: 443 targetPort: 9876 webhookName: rss-webhook ```
`spark.rss` # { #helm.spark.rss } : Type `object`. The configuration for the uniffle remote shuffle service ??? note "Default" ```yaml appName: rss-hops configDir: /data/rssadmin/rss/conf configmap: name: rss-configuration controller: containerPort: 9876 image: rss-controller replicas: 1 resources: limits: cpu: '1' memory: 512Mi requests: cpu: 100m memory: 150Mi serviceAccount: annotations: {} coordinator: count: 2 dynamicClientConfigMapName: rss-dynamic-client-configuration dynamicClientConfigMountPath: /tmp httpPort: 19996 labels: role: rss-hops-coordinator replicas: 1 resources: limits: cpu: 500m memory: 2Gi requests: cpu: 500m rpcPort: 19997 xmxSizeMemoryExtraPercentage: 30 dashboard: enabled: false httpPort: 19997 name: uniffle-dashboard resources: limits: cpu: 500m memory: 2Gi requests: cpu: 250m xmxSizeMemoryExtraPercentage: 30 dynamicClient: readBufferSize: 14m storageType: MEMORY_LOCALFILE fullnameOverride: null image: image: rss initImage: hops-rss-init initImageVersion: 0.9.2 pullPolicy: Always namespaceSelector: '' nodeSelector: {} resources: jobs: limits: cpu: 200m rss: limits: cpu: 500m memory: 1Gi serviceAccount: annotations: {} shuffleServer: bufferCapacity: -1 bufferCapacityRatio: 0.6 diskCapacity: -1 diskCapacityRatio: 0.8 httpPort: 19998 nettyPort: 20000 readBufferCapacity: -1 readBufferCapacityRatio: 0.1 replicas: 3 resources: limits: cpu: 2000m memory: 3Gi requests: cpu: 200m rpcPort: 19999 upgradeStrategy: FullUpgrade storage: size: 10Gi storageClassName: null volumeNameTemplate: rss-storage tolerations: [] topologySpreadConstraint: {} ttlSecondsAfterFinished: null version: 0.11.1 webhook: app: rss-webhook resources: limits: cpu: '1' memory: 512Mi requests: cpu: 100m memory: 150Mi service: name: rss-webhook port: 443 targetPort: 9876 webhookName: rss-webhook ``` `spark.rss.controller.serviceAccount.annotations` # { #helm.spark.rss.controller.serviceAccount.annotations } : Type `object`, default `{}`. service account annotations `spark.rss.coordinator.resources.requests` # { #helm.spark.rss.coordinator.resources.requests } : Type `object`, default `{"cpu":"500m"}`. memory is set to be equal to the limit `spark.rss.dashboard.resources.requests` # { #helm.spark.rss.dashboard.resources.requests } : Type `object`, default `{"cpu":"250m"}`. memory is set to be equal to the limit `spark.rss.fullnameOverride` # { #helm.spark.rss.fullnameOverride } : Type `string`, default `nil`. fullnameOverride `spark.rss.namespaceSelector` # { #helm.spark.rss.namespaceSelector } : Type `string`, default `""`. Label selector (key=value or key1=value1,key2=value2) for namespaces this instance should manage. Only events from matching namespaces will be processed by the controller and webhook. If empty, all namespaces are managed. `spark.rss.nodeSelector` # { #helm.spark.rss.nodeSelector } : Type `object`, default `{}`. node selector configuration `spark.rss.resources.jobs` # { #helm.spark.rss.resources.jobs } : Type `object`, default `{"limits":{"cpu":"200m"}}`. jobs resources `spark.rss.resources.rss` # { #helm.spark.rss.resources.rss } : Type `object`, default `{"limits":{"cpu":"500m","memory":"1Gi"}}`. rss resources `spark.rss.serviceAccount.annotations` # { #helm.spark.rss.serviceAccount.annotations } : Type `object`, default `{}`. service account annotations `spark.rss.shuffleServer.resources.requests` # { #helm.spark.rss.shuffleServer.resources.requests } : Type `object`, default `{"cpu":"200m"}`. memory is set to be equal to the limit `spark.rss.storage.storageClassName` # { #helm.spark.rss.storage.storageClassName } : Type `string`, default `nil`. storage class name to request for volumes attached to rss `spark.rss.topologySpreadConstraint` # { #helm.spark.rss.topologySpreadConstraint } : Type `object`, default `{}`. The default topology spread constraint. If not defined the global topology spread constraint would be used instead. `spark.rss.ttlSecondsAfterFinished` # { #helm.spark.rss.ttlSecondsAfterFinished } : Type `string`, default `nil`. TTL in seconds for the rss-config Job. Overrides global default. `spark.rss.webhookName` # { #helm.spark.rss.webhookName } : Type `string`, default `"rss-webhook"`. Name of the MutatingWebhookConfiguration and ValidatingWebhookConfiguration. Override when running multiple instances on the same cluster to avoid name collisions.
## spark-operator { #helm-values-spark-spark-operator } ??? example "Defaults as YAML" ```yaml spark: spark-operator: certManager: duration: '' enable: false issuerRef: {} renewBefore: '' commonLabels: {} controller: affinity: {} annotations: {} batchScheduler: default: '' enable: false kubeSchedulerNames: [] driverPodCreationGracePeriod: 10s env: [] envFrom: [] labels: {} leaderElection: enable: true logLevel: info maxTrackedExecutorPerApp: 1000 nodeSelector: {} podDisruptionBudget: enable: false minAvailable: 1 podSecurityContext: fsGroup: 185 pprof: enable: false port: 6060 portName: pprof priorityClassName: '' rbac: annotations: {} create: true replicas: 1 resources: requests: cpu: 300m memory: 512Mi securityContext: allowPrivilegeEscalation: false capabilities: drop: - ALL privileged: false readOnlyRootFilesystem: true runAsNonRoot: true seccompProfile: type: RuntimeDefault serviceAccount: annotations: {} automountServiceAccountToken: true create: true name: '' sidecars: [] tolerations: [] topologySpreadConstraints: [] uiIngress: annotations: {} enable: false ingressClassName: '' tls: [] urlFormat: '' uiService: enable: true volumeMounts: - mountPath: /tmp name: tmp readOnly: false volumes: - emptyDir: sizeLimit: 1Gi name: tmp workers: 10 workqueueRateLimiter: bucketQPS: 50 bucketSize: 500 maxDelay: duration: 6h enable: true fullnameOverride: '' hook: affinity: {} image: registry: docker.hops.works repository: hopsworks/spark-operator-crds tag: 2.5.1-h1-1.9 nodeSelector: {} tolerations: [] upgradeCrd: true image: pullPolicy: IfNotPresent pullSecrets: [] registry: docker.hops.works repository: hopsworks/spark-operator tag: 2.5.1-h1 nameOverride: '' podSecurityContext: fsGroup: 185 runAsGroup: 185 runAsNonRoot: true runAsUser: 185 seccompProfile: type: RuntimeDefault prometheus: metrics: enable: true endpoint: /metrics jobStartLatencyBuckets: 30,60,90,120,150,180,210,240,270,300 port: 8080 portName: metrics prefix: '' podMonitor: create: false jobLabel: spark-operator-podmonitor labels: {} podMetricsEndpoint: interval: 5s scheme: http securityContext: allowPrivilegeEscalation: false capabilities: drop: - ALL runAsGroup: 185 runAsNonRoot: true runAsUser: 185 seccompProfile: type: RuntimeDefault spark: jobNamespaces: [] rbac: annotations: {} create: true serviceAccount: annotations: {} automountServiceAccountToken: true create: true name: '' webhook: affinity: {} annotations: {} enable: true env: [] envFrom: [] failurePolicy: Fail labels: {} leaderElection: enable: true logLevel: info nodeSelector: {} podDisruptionBudget: enable: false minAvailable: 1 podSecurityContext: fsGroup: 185 port: 9443 portName: webhook priorityClassName: '' rbac: annotations: {} create: true replicas: 1 resourceQuotaEnforcement: enable: false resources: requests: cpu: 300m memory: 512Mi securityContext: allowPrivilegeEscalation: false capabilities: drop: - ALL privileged: false readOnlyRootFilesystem: true runAsNonRoot: true seccompProfile: type: RuntimeDefault serviceAccount: annotations: {} automountServiceAccountToken: true create: true name: '' sidecars: [] timeoutSeconds: 10 tolerations: [] topologySpreadConstraints: [] volumeMounts: - mountPath: /etc/k8s-webhook-server/serving-certs name: serving-certs readOnly: false subPath: serving-certs - mountPath: /tmp name: tmp volumes: - emptyDir: sizeLimit: 500Mi name: serving-certs - emptyDir: {} name: tmp ```
`spark.spark-operator` # { #helm.spark.spark-operator } : Type `object`, passed to the [`spark-operator` 2.5.1](https://github.com/kubeflow/spark-operator/blob/v2.5.1/charts/spark-operator-chart/README.md) chart, whose other values are documented there. override spark operator values ??? note "Default" ```yaml certManager: duration: '' enable: false issuerRef: {} renewBefore: '' commonLabels: {} controller: affinity: {} annotations: {} batchScheduler: default: '' enable: false kubeSchedulerNames: [] driverPodCreationGracePeriod: 10s env: [] envFrom: [] labels: {} leaderElection: enable: true logLevel: info maxTrackedExecutorPerApp: 1000 nodeSelector: {} podDisruptionBudget: enable: false minAvailable: 1 podSecurityContext: fsGroup: 185 pprof: enable: false port: 6060 portName: pprof priorityClassName: '' rbac: annotations: {} create: true replicas: 1 resources: requests: cpu: 300m memory: 512Mi securityContext: allowPrivilegeEscalation: false capabilities: drop: - ALL privileged: false readOnlyRootFilesystem: true runAsNonRoot: true seccompProfile: type: RuntimeDefault serviceAccount: annotations: {} automountServiceAccountToken: true create: true name: '' sidecars: [] tolerations: [] topologySpreadConstraints: [] uiIngress: annotations: {} enable: false ingressClassName: '' tls: [] urlFormat: '' uiService: enable: true volumeMounts: - mountPath: /tmp name: tmp readOnly: false volumes: - emptyDir: sizeLimit: 1Gi name: tmp workers: 10 workqueueRateLimiter: bucketQPS: 50 bucketSize: 500 maxDelay: duration: 6h enable: true fullnameOverride: '' hook: affinity: {} image: registry: docker.hops.works repository: hopsworks/spark-operator-crds tag: 2.5.1-h1-1.9 nodeSelector: {} tolerations: [] upgradeCrd: true image: pullPolicy: IfNotPresent pullSecrets: [] registry: docker.hops.works repository: hopsworks/spark-operator tag: 2.5.1-h1 nameOverride: '' podSecurityContext: fsGroup: 185 runAsGroup: 185 runAsNonRoot: true runAsUser: 185 seccompProfile: type: RuntimeDefault prometheus: metrics: enable: true endpoint: /metrics jobStartLatencyBuckets: 30,60,90,120,150,180,210,240,270,300 port: 8080 portName: metrics prefix: '' podMonitor: create: false jobLabel: spark-operator-podmonitor labels: {} podMetricsEndpoint: interval: 5s scheme: http securityContext: allowPrivilegeEscalation: false capabilities: drop: - ALL runAsGroup: 185 runAsNonRoot: true runAsUser: 185 seccompProfile: type: RuntimeDefault spark: jobNamespaces: [] rbac: annotations: {} create: true serviceAccount: annotations: {} automountServiceAccountToken: true create: true name: '' webhook: affinity: {} annotations: {} enable: true env: [] envFrom: [] failurePolicy: Fail labels: {} leaderElection: enable: true logLevel: info nodeSelector: {} podDisruptionBudget: enable: false minAvailable: 1 podSecurityContext: fsGroup: 185 port: 9443 portName: webhook priorityClassName: '' rbac: annotations: {} create: true replicas: 1 resourceQuotaEnforcement: enable: false resources: requests: cpu: 300m memory: 512Mi securityContext: allowPrivilegeEscalation: false capabilities: drop: - ALL privileged: false readOnlyRootFilesystem: true runAsNonRoot: true seccompProfile: type: RuntimeDefault serviceAccount: annotations: {} automountServiceAccountToken: true create: true name: '' sidecars: [] timeoutSeconds: 10 tolerations: [] topologySpreadConstraints: [] volumeMounts: - mountPath: /etc/k8s-webhook-server/serving-certs name: serving-certs readOnly: false subPath: serving-certs - mountPath: /tmp name: tmp volumes: - emptyDir: sizeLimit: 500Mi name: serving-certs - emptyDir: {} name: tmp ``` `spark.spark-operator.certManager.duration` # { #helm.spark.spark-operator.certManager.duration } : Type `string`, default `2160h` (90 days) will be used if not specified.. The duration of the certificate validity (e.g. `2160h`). See [cert-manager.io/v1.Certificate](https://cert-manager.io/docs/reference/api-docs/#cert-manager.io/v1.Certificate). `spark.spark-operator.certManager.enable` # { #helm.spark.spark-operator.certManager.enable } : Type `bool`, default `false`. Specifies whether to use [cert-manager](https://cert-manager.io) to generate certificate for webhook. `webhook.enable` must be set to `true` to enable cert-manager. `spark.spark-operator.certManager.issuerRef` # { #helm.spark.spark-operator.certManager.issuerRef } : Type `object`, default A self-signed issuer will be created and used if not specified.. The reference to the issuer. `spark.spark-operator.certManager.renewBefore` # { #helm.spark.spark-operator.certManager.renewBefore } : Type `string`, default 1/3 of issued certificate’s lifetime.. The duration before the certificate expiration to renew the certificate (e.g. `720h`). See [cert-manager.io/v1.Certificate](https://cert-manager.io/docs/reference/api-docs/#cert-manager.io/v1.Certificate). `spark.spark-operator.commonLabels` # { #helm.spark.spark-operator.commonLabels } : Type `object`, default `{}`. Common labels to add to the resources. `spark.spark-operator.controller.affinity` # { #helm.spark.spark-operator.controller.affinity } : Type `object`, default `{}`. Affinity for controller pods. `spark.spark-operator.controller.annotations` # { #helm.spark.spark-operator.controller.annotations } : Type `object`, default `{}`. Extra annotations for controller pods. `spark.spark-operator.controller.batchScheduler.default` # { #helm.spark.spark-operator.controller.batchScheduler.default } : Type `string`, default `""`. Default batch scheduler to be used if not specified by the user. If specified, this value must be either "volcano" or "yunikorn". Specifying any other value will cause the controller to error on startup. `spark.spark-operator.controller.batchScheduler.enable` # { #helm.spark.spark-operator.controller.batchScheduler.enable } : Type `bool`, default `false`. Specifies whether to enable batch scheduler for spark jobs scheduling. If enabled, users can specify batch scheduler name in spark application. `spark.spark-operator.controller.batchScheduler.kubeSchedulerNames` # { #helm.spark.spark-operator.controller.batchScheduler.kubeSchedulerNames } : Type `list`, default `[]`. Specifies a list of kube-scheduler names for scheduling Spark pods. `spark.spark-operator.controller.driverPodCreationGracePeriod` # { #helm.spark.spark-operator.controller.driverPodCreationGracePeriod } : Type `string`, default `"10s"`. Grace period after a successful spark-submit when driver pod not found errors will be retried. Useful if the driver pod can take some time to be created. `spark.spark-operator.controller.env` # { #helm.spark.spark-operator.controller.env } : Type `list`, default `[]`. Environment variables for controller containers. `spark.spark-operator.controller.envFrom` # { #helm.spark.spark-operator.controller.envFrom } : Type `list`, default `[]`. Environment variable sources for controller containers. `spark.spark-operator.controller.labels` # { #helm.spark.spark-operator.controller.labels } : Type `object`, default `{}`. Extra labels for controller pods. `spark.spark-operator.controller.leaderElection.enable` # { #helm.spark.spark-operator.controller.leaderElection.enable } : Type `bool`, default `true`. Specifies whether to enable leader election for controller. `spark.spark-operator.controller.logLevel` # { #helm.spark.spark-operator.controller.logLevel } : Type `string`, default `"info"`. Configure the verbosity of logging, can be one of `debug`, `info`, `error`. `spark.spark-operator.controller.maxTrackedExecutorPerApp` # { #helm.spark.spark-operator.controller.maxTrackedExecutorPerApp } : Type `int`, default `1000`. Specifies the maximum number of Executor pods that can be tracked by the controller per SparkApplication. `spark.spark-operator.controller.nodeSelector` # { #helm.spark.spark-operator.controller.nodeSelector } : Type `object`, default `{}`. Node selector for controller pods. `spark.spark-operator.controller.podDisruptionBudget.enable` # { #helm.spark.spark-operator.controller.podDisruptionBudget.enable } : Type `bool`, default `false`. Specifies whether to create pod disruption budget for controller. Ref: [Specifying a Disruption Budget for your Application](https://kubernetes.io/docs/tasks/run-application/configure-pdb/) `spark.spark-operator.controller.podDisruptionBudget.minAvailable` # { #helm.spark.spark-operator.controller.podDisruptionBudget.minAvailable } : Type `int`, default `1`. The number of pods that must be available. Require `controller.replicas` to be greater than 1 `spark.spark-operator.controller.podSecurityContext` # { #helm.spark.spark-operator.controller.podSecurityContext } : Type `object`, default `{"fsGroup":185}`. Security context for controller pods. `spark.spark-operator.controller.pprof.enable` # { #helm.spark.spark-operator.controller.pprof.enable } : Type `bool`, default `false`. Specifies whether to enable pprof. `spark.spark-operator.controller.pprof.port` # { #helm.spark.spark-operator.controller.pprof.port } : Type `int`, default `6060`. Specifies pprof port. `spark.spark-operator.controller.pprof.portName` # { #helm.spark.spark-operator.controller.pprof.portName } : Type `string`, default `"pprof"`. Specifies pprof service port name. `spark.spark-operator.controller.priorityClassName` # { #helm.spark.spark-operator.controller.priorityClassName } : Type `string`, default `""`. Priority class for controller pods. `spark.spark-operator.controller.rbac.annotations` # { #helm.spark.spark-operator.controller.rbac.annotations } : Type `object`, default `{}`. Extra annotations for the controller RBAC resources. `spark.spark-operator.controller.rbac.create` # { #helm.spark.spark-operator.controller.rbac.create } : Type `bool`, default `true`. Specifies whether to create RBAC resources for the controller. `spark.spark-operator.controller.replicas` # { #helm.spark.spark-operator.controller.replicas } : Type `int`, default `1`. Number of replicas of controller. `spark.spark-operator.controller.resources` # { #helm.spark.spark-operator.controller.resources } : Type `object`, default `{"requests":{"cpu":"300m","memory":"512Mi"}}`. Pod resource requests and limits for controller containers. Note, that each job submission will spawn a JVM within the controller pods using "/usr/local/openjdk-11/bin/java -Xmx128m". Kubernetes may kill these Java processes at will to enforce resource limits. When that happens, you will see the following error: 'failed to run spark-submit for SparkApplication \[...\]: signal: killed' - when this happens, you may want to increase memory limits. `spark.spark-operator.controller.securityContext` # { #helm.spark.spark-operator.controller.securityContext } : Type `object`. Security context for controller containers. ??? note "Default" ```yaml allowPrivilegeEscalation: false capabilities: drop: - ALL privileged: false readOnlyRootFilesystem: true runAsNonRoot: true seccompProfile: type: RuntimeDefault ``` `spark.spark-operator.controller.serviceAccount.annotations` # { #helm.spark.spark-operator.controller.serviceAccount.annotations } : Type `object`, default `{}`. Extra annotations for the controller service account. `spark.spark-operator.controller.serviceAccount.automountServiceAccountToken` # { #helm.spark.spark-operator.controller.serviceAccount.automountServiceAccountToken } : Type `bool`, default `true`. Auto-mount service account token to the controller pods. `spark.spark-operator.controller.serviceAccount.create` # { #helm.spark.spark-operator.controller.serviceAccount.create } : Type `bool`, default `true`. Specifies whether to create a service account for the controller. `spark.spark-operator.controller.serviceAccount.name` # { #helm.spark.spark-operator.controller.serviceAccount.name } : Type `string`, default `""`. Optional name for the controller service account. `spark.spark-operator.controller.sidecars` # { #helm.spark.spark-operator.controller.sidecars } : Type `list`, default `[]`. Sidecar containers for controller pods. `spark.spark-operator.controller.tolerations` # { #helm.spark.spark-operator.controller.tolerations } : Type `list`, default `[]`. List of node taints to tolerate for controller pods. `spark.spark-operator.controller.topologySpreadConstraints` # { #helm.spark.spark-operator.controller.topologySpreadConstraints } : Type `list`, default `[]`. Topology spread constraints rely on node labels to identify the topology domain(s) that each Node is in. Ref: [Pod Topology Spread Constraints](https://kubernetes.io/docs/concepts/workloads/pods/pod-topology-spread-constraints/). The labelSelector field in topology spread constraint will be set to the selector labels for controller pods if not specified. `spark.spark-operator.controller.uiIngress.annotations` # { #helm.spark.spark-operator.controller.uiIngress.annotations } : Type `object`, default `{}`. Optionally set default ingress annotations for the Spark UI's ingress. `ingressAnnotations` in the SparkApplication spec overrides this. `spark.spark-operator.controller.uiIngress.enable` # { #helm.spark.spark-operator.controller.uiIngress.enable } : Type `bool`, default `false`. Specifies whether to create ingress for Spark web UI. `controller.uiService.enable` must be `true` to enable ingress. `spark.spark-operator.controller.uiIngress.ingressClassName` # { #helm.spark.spark-operator.controller.uiIngress.ingressClassName } : Type `string`, default `""`. Optionally set the ingressClassName. `spark.spark-operator.controller.uiIngress.tls` # { #helm.spark.spark-operator.controller.uiIngress.tls } : Type `list`, default `[]`. Optionally set default TLS configuration for the Spark UI's ingress. `ingressTLS` in the SparkApplication spec overrides this. `spark.spark-operator.controller.uiIngress.urlFormat` # { #helm.spark.spark-operator.controller.uiIngress.urlFormat } : Type `string`, default `""`. Ingress URL format. Required if `controller.uiIngress.enable` is true. `spark.spark-operator.controller.uiService.enable` # { #helm.spark.spark-operator.controller.uiService.enable } : Type `bool`, default `true`. Specifies whether to create service for Spark web UI. `spark.spark-operator.controller.volumeMounts` # { #helm.spark.spark-operator.controller.volumeMounts } : Type `list`, default `[{"mountPath":"/tmp","name":"tmp","readOnly":false}]`. Volume mounts for controller containers. `spark.spark-operator.controller.volumes` # { #helm.spark.spark-operator.controller.volumes } : Type `list`, default `[{"emptyDir":{"sizeLimit":"1Gi"},"name":"tmp"}]`. Volumes for controller pods. `spark.spark-operator.controller.workers` # { #helm.spark.spark-operator.controller.workers } : Type `int`, default `10`. Reconcile concurrency, higher values might increase memory usage. `spark.spark-operator.controller.workqueueRateLimiter.bucketQPS` # { #helm.spark.spark-operator.controller.workqueueRateLimiter.bucketQPS } : Type `int`, default `50`. Specifies the average rate of items process by the workqueue rate limiter. `spark.spark-operator.controller.workqueueRateLimiter.bucketSize` # { #helm.spark.spark-operator.controller.workqueueRateLimiter.bucketSize } : Type `int`, default `500`. Specifies the maximum number of items that can be in the workqueue at any given time. `spark.spark-operator.controller.workqueueRateLimiter.maxDelay.duration` # { #helm.spark.spark-operator.controller.workqueueRateLimiter.maxDelay.duration } : Type `string`, default `"6h"`. Specifies the maximum delay duration for the workqueue rate limiter. `spark.spark-operator.controller.workqueueRateLimiter.maxDelay.enable` # { #helm.spark.spark-operator.controller.workqueueRateLimiter.maxDelay.enable } : Type `bool`, default `true`. Specifies whether to enable max delay for the workqueue rate limiter. This is useful to avoid losing events when the workqueue is full. `spark.spark-operator.fullnameOverride` # { #helm.spark.spark-operator.fullnameOverride } : Type `string`, default `""`. String to fully override release name. `spark.spark-operator.hook.affinity` # { #helm.spark.spark-operator.hook.affinity } : Type `object`, default `{}`. Affinity for the Helm hook Job. `spark.spark-operator.hook.image.registry` # { #helm.spark.spark-operator.hook.image.registry } : Type `string`, default `"docker.hops.works"`. Image registry. `spark.spark-operator.hook.image.repository` # { #helm.spark.spark-operator.hook.image.repository } : Type `string`, default `"hopsworks/spark-operator-crds"`. Image repository. `spark.spark-operator.hook.image.tag` # { #helm.spark.spark-operator.hook.image.tag } : Type `string`, default If not set, the chart appVersion will be used.. Image tag. `spark.spark-operator.hook.nodeSelector` # { #helm.spark.spark-operator.hook.nodeSelector } : Type `object`, default `{}`. Node selector for the Helm hook Job. `spark.spark-operator.hook.tolerations` # { #helm.spark.spark-operator.hook.tolerations } : Type `list`, default `[]`. List of node taints to tolerate for the Helm hook Job. `spark.spark-operator.hook.upgradeCrd` # { #helm.spark.spark-operator.hook.upgradeCrd } : Type `bool`, default `true`. Whether to create a Helm pre-install/pre-upgrade hook Job to update CRDs. `spark.spark-operator.image.pullPolicy` # { #helm.spark.spark-operator.image.pullPolicy } : Type `string`, default `"IfNotPresent"`. Image pull policy. `spark.spark-operator.image.pullSecrets` # { #helm.spark.spark-operator.image.pullSecrets } : Type `list`, default `[]`. Image pull secrets for private image registry. `spark.spark-operator.image.registry` # { #helm.spark.spark-operator.image.registry } : Type `string`, default `"docker.hops.works"`. Image registry. `spark.spark-operator.image.repository` # { #helm.spark.spark-operator.image.repository } : Type `string`, default `"hopsworks/spark-operator"`. Image repository. `spark.spark-operator.image.tag` # { #helm.spark.spark-operator.image.tag } : Type `string`, default If not set, the chart appVersion will be used.. Image tag. `spark.spark-operator.nameOverride` # { #helm.spark.spark-operator.nameOverride } : Type `string`, default `""`. String to partially override release name. `spark.spark-operator.podSecurityContext` # { #helm.spark.spark-operator.podSecurityContext } : Type `object`. Pod-level security context for spark-operator ??? note "Default" ```yaml fsGroup: 185 runAsGroup: 185 runAsNonRoot: true runAsUser: 185 seccompProfile: type: RuntimeDefault ``` `spark.spark-operator.prometheus.metrics.enable` # { #helm.spark.spark-operator.prometheus.metrics.enable } : Type `bool`, default `true`. Specifies whether to enable prometheus metrics scraping. `spark.spark-operator.prometheus.metrics.endpoint` # { #helm.spark.spark-operator.prometheus.metrics.endpoint } : Type `string`, default `"/metrics"`. Metrics serving endpoint. `spark.spark-operator.prometheus.metrics.jobStartLatencyBuckets` # { #helm.spark.spark-operator.prometheus.metrics.jobStartLatencyBuckets } : Type `string`, default `"30,60,90,120,150,180,210,240,270,300"`. Job Start Latency histogram buckets. Specified in seconds. `spark.spark-operator.prometheus.metrics.port` # { #helm.spark.spark-operator.prometheus.metrics.port } : Type `int`, default `8080`. Metrics port. `spark.spark-operator.prometheus.metrics.portName` # { #helm.spark.spark-operator.prometheus.metrics.portName } : Type `string`, default `"metrics"`. Metrics port name. `spark.spark-operator.prometheus.metrics.prefix` # { #helm.spark.spark-operator.prometheus.metrics.prefix } : Type `string`, default `""`. Metrics prefix, will be added to all exported metrics. `spark.spark-operator.prometheus.podMonitor.create` # { #helm.spark.spark-operator.prometheus.podMonitor.create } : Type `bool`, default `false`. Specifies whether to create pod monitor. Note that prometheus metrics should be enabled as well. `spark.spark-operator.prometheus.podMonitor.jobLabel` # { #helm.spark.spark-operator.prometheus.podMonitor.jobLabel } : Type `string`, default `"spark-operator-podmonitor"`. The label to use to retrieve the job name from `spark.spark-operator.prometheus.podMonitor.labels` # { #helm.spark.spark-operator.prometheus.podMonitor.labels } : Type `object`, default `{}`. Pod monitor labels `spark.spark-operator.prometheus.podMonitor.podMetricsEndpoint` # { #helm.spark.spark-operator.prometheus.podMonitor.podMetricsEndpoint } : Type `object`, default `{"interval":"5s","scheme":"http"}`. Prometheus metrics endpoint properties. `metrics.portName` will be used as a port `spark.spark-operator.securityContext` # { #helm.spark.spark-operator.securityContext } : Type `object`. Container-level security context for spark-operator ??? note "Default" ```yaml allowPrivilegeEscalation: false capabilities: drop: - ALL runAsGroup: 185 runAsNonRoot: true runAsUser: 185 seccompProfile: type: RuntimeDefault ``` `spark.spark-operator.spark.jobNamespaces` # { #helm.spark.spark-operator.spark.jobNamespaces } : Type `list`, default `[]`. List of namespaces where to run spark jobs. If empty string is included, all namespaces will be allowed. Make sure the namespaces have already existed. `spark.spark-operator.spark.rbac.annotations` # { #helm.spark.spark-operator.spark.rbac.annotations } : Type `object`, default `{}`. Optional annotations for the spark application RBAC resources. `spark.spark-operator.spark.rbac.create` # { #helm.spark.spark-operator.spark.rbac.create } : Type `bool`, default `true`. Specifies whether to create RBAC resources for spark applications. `spark.spark-operator.spark.serviceAccount.annotations` # { #helm.spark.spark-operator.spark.serviceAccount.annotations } : Type `object`, default `{}`. Optional annotations for the spark service account. `spark.spark-operator.spark.serviceAccount.automountServiceAccountToken` # { #helm.spark.spark-operator.spark.serviceAccount.automountServiceAccountToken } : Type `bool`, default `true`. Auto-mount service account token to the spark applications pods. `spark.spark-operator.spark.serviceAccount.create` # { #helm.spark.spark-operator.spark.serviceAccount.create } : Type `bool`, default `true`. Specifies whether to create a service account for spark applications. `spark.spark-operator.spark.serviceAccount.name` # { #helm.spark.spark-operator.spark.serviceAccount.name } : Type `string`, default `""`. Optional name for the spark service account. `spark.spark-operator.webhook.affinity` # { #helm.spark.spark-operator.webhook.affinity } : Type `object`, default `{}`. Affinity for webhook pods. `spark.spark-operator.webhook.annotations` # { #helm.spark.spark-operator.webhook.annotations } : Type `object`, default `{}`. Extra annotations for webhook pods. `spark.spark-operator.webhook.enable` # { #helm.spark.spark-operator.webhook.enable } : Type `bool`, default `true`. Specifies whether to enable webhook. `spark.spark-operator.webhook.env` # { #helm.spark.spark-operator.webhook.env } : Type `list`, default `[]`. Environment variables for webhook containers. `spark.spark-operator.webhook.envFrom` # { #helm.spark.spark-operator.webhook.envFrom } : Type `list`, default `[]`. Environment variable sources for webhook containers. `spark.spark-operator.webhook.failurePolicy` # { #helm.spark.spark-operator.webhook.failurePolicy } : Type `string`, default `"Fail"`. Specifies how unrecognized errors are handled. Available options are `Ignore` or `Fail`. `spark.spark-operator.webhook.labels` # { #helm.spark.spark-operator.webhook.labels } : Type `object`, default `{}`. Extra labels for webhook pods. `spark.spark-operator.webhook.leaderElection.enable` # { #helm.spark.spark-operator.webhook.leaderElection.enable } : Type `bool`, default `true`. Specifies whether to enable leader election for webhook. `spark.spark-operator.webhook.logLevel` # { #helm.spark.spark-operator.webhook.logLevel } : Type `string`, default `"info"`. Configure the verbosity of logging, can be one of `debug`, `info`, `error`. `spark.spark-operator.webhook.nodeSelector` # { #helm.spark.spark-operator.webhook.nodeSelector } : Type `object`, default `{}`. Node selector for webhook pods. `spark.spark-operator.webhook.podDisruptionBudget.enable` # { #helm.spark.spark-operator.webhook.podDisruptionBudget.enable } : Type `bool`, default `false`. Specifies whether to create pod disruption budget for webhook. Ref: [Specifying a Disruption Budget for your Application](https://kubernetes.io/docs/tasks/run-application/configure-pdb/) `spark.spark-operator.webhook.podDisruptionBudget.minAvailable` # { #helm.spark.spark-operator.webhook.podDisruptionBudget.minAvailable } : Type `int`, default `1`. The number of pods that must be available. Require `webhook.replicas` to be greater than 1 `spark.spark-operator.webhook.podSecurityContext` # { #helm.spark.spark-operator.webhook.podSecurityContext } : Type `object`, default `{"fsGroup":185}`. Security context for webhook pods. `spark.spark-operator.webhook.port` # { #helm.spark.spark-operator.webhook.port } : Type `int`, default `9443`. Specifies webhook port. `spark.spark-operator.webhook.portName` # { #helm.spark.spark-operator.webhook.portName } : Type `string`, default `"webhook"`. Specifies webhook service port name. `spark.spark-operator.webhook.priorityClassName` # { #helm.spark.spark-operator.webhook.priorityClassName } : Type `string`, default `""`. Priority class for webhook pods. `spark.spark-operator.webhook.rbac.annotations` # { #helm.spark.spark-operator.webhook.rbac.annotations } : Type `object`, default `{}`. Extra annotations for the webhook RBAC resources. `spark.spark-operator.webhook.rbac.create` # { #helm.spark.spark-operator.webhook.rbac.create } : Type `bool`, default `true`. Specifies whether to create RBAC resources for the webhook. `spark.spark-operator.webhook.replicas` # { #helm.spark.spark-operator.webhook.replicas } : Type `int`, default `1`. Number of replicas of webhook server. `spark.spark-operator.webhook.resourceQuotaEnforcement.enable` # { #helm.spark.spark-operator.webhook.resourceQuotaEnforcement.enable } : Type `bool`, default `false`. Specifies whether to enable the ResourceQuota enforcement for SparkApplication resources. `spark.spark-operator.webhook.resources` # { #helm.spark.spark-operator.webhook.resources } : Type `object`, default `{"requests":{"cpu":"300m","memory":"512Mi"}}`. Pod resource requests and limits for webhook pods. `spark.spark-operator.webhook.securityContext` # { #helm.spark.spark-operator.webhook.securityContext } : Type `object`. Security context for webhook containers. ??? note "Default" ```yaml allowPrivilegeEscalation: false capabilities: drop: - ALL privileged: false readOnlyRootFilesystem: true runAsNonRoot: true seccompProfile: type: RuntimeDefault ``` `spark.spark-operator.webhook.serviceAccount.annotations` # { #helm.spark.spark-operator.webhook.serviceAccount.annotations } : Type `object`, default `{}`. Extra annotations for the webhook service account. `spark.spark-operator.webhook.serviceAccount.automountServiceAccountToken` # { #helm.spark.spark-operator.webhook.serviceAccount.automountServiceAccountToken } : Type `bool`, default `true`. Auto-mount service account token to the webhook pods. `spark.spark-operator.webhook.serviceAccount.create` # { #helm.spark.spark-operator.webhook.serviceAccount.create } : Type `bool`, default `true`. Specifies whether to create a service account for the webhook. `spark.spark-operator.webhook.serviceAccount.name` # { #helm.spark.spark-operator.webhook.serviceAccount.name } : Type `string`, default `""`. Optional name for the webhook service account. `spark.spark-operator.webhook.sidecars` # { #helm.spark.spark-operator.webhook.sidecars } : Type `list`, default `[]`. Sidecar containers for webhook pods. `spark.spark-operator.webhook.timeoutSeconds` # { #helm.spark.spark-operator.webhook.timeoutSeconds } : Type `int`, default `10`. Specifies the timeout seconds of the webhook, the value must be between 1 and 30. `spark.spark-operator.webhook.tolerations` # { #helm.spark.spark-operator.webhook.tolerations } : Type `list`, default `[]`. List of node taints to tolerate for webhook pods. `spark.spark-operator.webhook.topologySpreadConstraints` # { #helm.spark.spark-operator.webhook.topologySpreadConstraints } : Type `list`, default `[]`. Topology spread constraints rely on node labels to identify the topology domain(s) that each Node is in. Ref: [Pod Topology Spread Constraints](https://kubernetes.io/docs/concepts/workloads/pods/pod-topology-spread-constraints/). The labelSelector field in topology spread constraint will be set to the selector labels for webhook pods if not specified. `spark.spark-operator.webhook.volumeMounts` # { #helm.spark.spark-operator.webhook.volumeMounts } : Type `list`. Volume mounts for webhook containers. ??? note "Default" ```yaml - mountPath: /etc/k8s-webhook-server/serving-certs name: serving-certs readOnly: false subPath: serving-certs - mountPath: /tmp name: tmp ``` `spark.spark-operator.webhook.volumes` # { #helm.spark.spark-operator.webhook.volumes } : Type `list`. Volumes for webhook pods. ??? note "Default" ```yaml - emptyDir: sizeLimit: 500Mi name: serving-certs - emptyDir: {} name: tmp ```
## sparkOperatorUpgradeJob { #helm-values-spark-sparkoperatorupgradejob } ??? example "Defaults as YAML" ```yaml spark: sparkOperatorUpgradeJob: crdImage: repository: spark-operator-crds tag: 2.2.1-1.4 imagePullPolicy: Always leaderElectionLockName: spark-operator-lock leaderElectionLockNamespace: '' name: spark-operator-upgrade-job nodeSelector: {} podMonitorName: spark-operator-podmonitor resources: limits: cpu: 200m memory: 200M requests: cpu: 100m memory: 100M serviceAccount: annotations: {} sparkJobNamespace: '' tolerations: [] ```
`spark.sparkOperatorUpgradeJob` # { #helm.spark.sparkOperatorUpgradeJob } : Type `object`. Job configuration to clean up old spark-operator resources before upgrading. ??? note "Default" ```yaml crdImage: repository: spark-operator-crds tag: 2.2.1-1.4 imagePullPolicy: Always leaderElectionLockName: spark-operator-lock leaderElectionLockNamespace: '' name: spark-operator-upgrade-job nodeSelector: {} podMonitorName: spark-operator-podmonitor resources: limits: cpu: 200m memory: 200M requests: cpu: 100m memory: 100M serviceAccount: annotations: {} sparkJobNamespace: '' tolerations: [] ``` `spark.sparkOperatorUpgradeJob.crdImage` # { #helm.spark.sparkOperatorUpgradeJob.crdImage } : Type `object`, default `{"repository":"spark-operator-crds","tag":"2.2.1-1.4"}`. Image for the pre-upgrade Job. Must ship the spark-operator CRDs at `/opt/spark-operator-crds/` and provide `bash` and `kubectl`. `spark.sparkOperatorUpgradeJob.leaderElectionLockName` # { #helm.spark.sparkOperatorUpgradeJob.leaderElectionLockName } : Type `string`, default `"spark-operator-lock"`. Leader election lease name used by the old chart (only if replicaCount > 1). `spark.sparkOperatorUpgradeJob.leaderElectionLockNamespace` # { #helm.spark.sparkOperatorUpgradeJob.leaderElectionLockNamespace } : Type `string`, default `""`. Optional leader election lease namespace (defaults to release namespace). `spark.sparkOperatorUpgradeJob.nodeSelector` # { #helm.spark.sparkOperatorUpgradeJob.nodeSelector } : Type `object`, default `{}`. node selector configuration `spark.sparkOperatorUpgradeJob.podMonitorName` # { #helm.spark.sparkOperatorUpgradeJob.podMonitorName } : Type `string`, default `"spark-operator-podmonitor"`. PodMonitor name used by the old chart (only if enabled). `spark.sparkOperatorUpgradeJob.serviceAccount.annotations` # { #helm.spark.sparkOperatorUpgradeJob.serviceAccount.annotations } : Type `object`, default `{}`. service account annotations `spark.sparkOperatorUpgradeJob.sparkJobNamespace` # { #helm.spark.sparkOperatorUpgradeJob.sparkJobNamespace } : Type `string`, default `""`. Namespace where spark app RBAC/SA were created by the old chart.
================================================================================ # superset Source: https://docs.hopsworks.ai/latest/setup_installation/common/helm_chart_values/superset/ # Superset values { #helm-values-superset } Values under `superset` configure Apache Superset, the BI dashboards over feature store data. _Generated from the Hopsworks Helm chart `5.2.0-alpha-1791549041` (Hopsworks `5.2.0`)._ Deployed according to the first of these values that is set: [`global._hopsworks.superset.enabled`](global.md#helm.global._hopsworks.superset.enabled), [`global._hopsworks.full_platform`](global.md#helm.global._hopsworks.full_platform). !!! info "Upstream charts" - Values under `superset.mysql` go to [`mysql` 12.3.5](https://artifacthub.io/packages/helm/bitnami/mysql/12.3.5) from `oci://registry-1.docker.io/bitnamicharts`. - Values under `superset.superset` go to [`superset` 0.15.0](https://artifacthub.io/packages/helm/superset/superset/0.15.0) from `https://apache.github.io/superset`. Only the values Hopsworks sets under `superset.mysql` and `superset.superset` are listed on this page. Any other value of the charts can be set under the same keys; each link opens the chart's documentation for the version Hopsworks pins. ## General { #helm-values-superset-general } ??? example "Defaults as YAML" ```yaml superset: cleanupOnUninstall: enabled: true ttlSecondsAfterFinished: null hopsworkslib: {} imageRegistry: docker.hops.works/superset mysql: architecture: standalone auth: createDatabase: true customPasswordFiles: {} database: superset existingSecret: superset-mysql-users-secrets password: temp-value replicationPassword: '' replicationUser: replicator rootPassword: temp-value usePasswordFiles: false username: superset enabled: true global: security: allowInsecureImages: true image: digest: '' pullPolicy: IfNotPresent registry: docker.hops.works/superset repository: mysql tag: 8.4.11-ubuntu24.04-h4 metrics: enabled: true image: digest: '' pullPolicy: IfNotPresent registry: docker.hops.works/superset repository: mysqld-exporter tag: 0.20.0-alpine-h1.1 resources: limits: cpu: 100m memory: 64Mi requests: cpu: 50m memory: 64Mi service: annotations: prometheus.io/port: '{{ .Values.metrics.service.port }}' prometheus.io/scrape: 'true' primary: persistentVolumeClaimRetentionPolicy: enabled: true whenDeleted: Delete whenScaled: Retain volumePermissions: enabled: false image: digest: '' pullPolicy: IfNotPresent registry: docker.hops.works/superset repository: os-shell tag: 12-alpine-h1.1 superset: _publicRoleDefaultName: Public allowAnonymousAccess: true configOverrides: feature_flags: | FEATURE_FLAGS = {"ALERT_REPORTS": True, "DASHBOARD_RBAC": True} flask_app_configuration: | from flask import session from flask import Flask from datetime import timedelta def make_session_permanent(): ''' Enable maxAge for the cookie 'session' ''' session.permanent = True # Set up max age of session to 24 hours PERMANENT_SESSION_LIFETIME = timedelta(hours=24) # (default: "31 days") SESSION_REFRESH_EACH_REQUEST = True # Default: True def FLASK_APP_MUTATOR(app: Flask) -> None: app.before_request_funcs.setdefault(None, []).append(make_session_permanent) mysql: | SQLALCHEMY_DATABASE_URI = f"mysql+mysqldb://{os.getenv('DB_USER')}:{os.getenv('DB_PASS')}@{os.getenv('DB_HOST')}:{os.getenv('DB_PORT')}/{os.getenv('DB_NAME')}" proxyConfig: | ENABLE_PROXY_FIX = True APP_ICON = "/hopsworks-api/superset/static/assets/images/superset-logo-horiz.png" LOGO_TARGET_PATH = "/hopsworks-api/superset/superset/welcome/" {{- if .Values.init.loadExamples }} PREVENT_UNSAFE_DB_CONNECTIONS = False {{- else }} PREVENT_UNSAFE_DB_CONNECTIONS = True {{- end }} public_role: | {{- if .Values.publicRoleLike }} PUBLIC_ROLE_LIKE = {{ .Values.publicRoleLike | quote }} {{- end }} {{- if .Values.allowAnonymousAccess }} AUTH_ROLE_PUBLIC = {{ .Values._publicRoleDefaultName | quote }} {{- end }} extraEnv: SUPERSET_APP_ROOT: /hopsworks-api/superset extraEnvRaw: - name: SUPERSET_USER valueFrom: secretKeyRef: key: username name: superset-admin-credentials - name: SUPERSET_PASS valueFrom: secretKeyRef: key: password name: superset-admin-credentials - name: DB_PASS valueFrom: secretKeyRef: key: mysql-password name: superset-mysql-users-secrets - name: SUPERSET_SECRET_KEY valueFrom: secretKeyRef: key: secret-key name: superset-secret-key extraRoles: - name: Dataset permissions: - - can_duplicate - Dataset - - can_write - Dataset - - can_get_or_create_dataset - Dataset - - can_warm_up_cache - Dataset extraVolumeMounts: - mountPath: /srv/hops/super_crypto/superset name: super-crypto-material readOnly: true extraVolumes: - name: super-crypto-material secret: optional: true secretName: hopsworks-superset-crypto-material fullnameOverride: hopsworks-superset image: pullPolicy: IfNotPresent repository: docker.hops.works/hopsworks/superset tag: 6.0.0p4.4 init: adminUser: email: admin@superset.com firstname: Superset lastname: Admin password: $SUPERSET_PASS username: $SUPERSET_USER command: - /bin/sh - -c - | {{- if ((((.Values.global | default dict)._hopsworks | default dict).restoreFromBackup | default dict).superset | default dict).enabled }} echo "Superset restore in progress (global._hopsworks.restoreFromBackup.superset.enabled): skipping init so no schema work runs while the database is reloaded. It runs on the upgrade that clears the flag."; exit 0 {{- end }} . {{ .Values.configMountPath }}/superset_bootstrap.sh; . {{ .Values.configMountPath }}/superset_init.sh; {{- if .Values.extraRoles }} {{- range $role := .Values.extraRoles }} python /scripts/create_role.py --role-name {{ $role.name }} --permissions '{{ $role.permissions | toJson | b64enc }}'; {{- end }} {{- end }} {{- if and .Values.publicRolePermissions (eq .Values.publicRoleLike .Values._publicRoleDefaultName) }} python /scripts/create_role.py --role-name {{ .Values._publicRoleDefaultName }} --permissions '{{ .Values.publicRolePermissions | toJson | b64enc }}'; {{- end }} containerSecurityContext: allowPrivilegeEscalation: false capabilities: drop: - ALL initContainers: - command: - /bin/sh - -c - dockerize -wait "tcp://$DB_HOST:$DB_PORT" -timeout 120s envFrom: - secretRef: name: '{{ tpl .Values.envFromSecret . }}' image: '{{ .Values.initImage.repository }}:{{ .Values.initImage.tag }}' imagePullPolicy: '{{ .Values.initImage.pullPolicy }}' name: wait-for-database resources: limits: cpu: 500m memory: 256Mi requests: cpu: 250m memory: 128Mi securityContext: allowPrivilegeEscalation: false capabilities: drop: - ALL loadExamples: false podSecurityContext: runAsNonRoot: true seccompProfile: type: RuntimeDefault initImage: pullPolicy: IfNotPresent repository: docker.hops.works/superset/dockerize tag: 0.14.0-alpine-h1.1 postgresql: enabled: false publicRoleLike: Public publicRolePermissions: - - can_read - Dashboard - - can_read - Chart - - can_dashboard - Superset - - can_slice - Superset - - can_explore_json - Superset - - can_dashboard_permalink - Superset - - can_read - DashboardPermalinkRestApi - - can_read - DashboardFilterStateRestApi - - can_write - DashboardFilterStateRestApi - - can_time_range - Api - - can_query_form_data - Api - - can_query - Api - - can_read - CssTemplate - - can_read - Theme - - can_read - EmbeddedDashboard - - can_read - CurrentUserRestApi - - can_get - Datasource - - can_external_metadata - Datasource - - can_read - Annotation - - can_read - AnnotationLayerRestApi - - can_read - ExplorePermalinkRestApi redis: enabled: true image: registry: docker.hops.works/superset repository: redis tag: 7.4.11-alpine-h1 master: configuration: |- maxmemory 256mb maxmemory-policy allkeys-lru containerSecurityContext: allowPrivilegeEscalation: false capabilities: drop: - ALL enabled: true readOnlyRootFilesystem: true runAsGroup: 1001 runAsNonRoot: true runAsUser: 1001 seLinuxOptions: {} seccompProfile: type: RuntimeDefault podSecurityContext: enabled: true fsGroup: 1001 fsGroupChangePolicy: Always runAsNonRoot: true runAsUser: 1001 seccompProfile: type: RuntimeDefault supplementalGroups: [] sysctls: [] resources: limits: cpu: 500m memory: 384Mi requests: cpu: 100m memory: 384Mi metrics: containerSecurityContext: allowPrivilegeEscalation: false capabilities: drop: - ALL enabled: true readOnlyRootFilesystem: true runAsGroup: 1001 runAsNonRoot: true runAsUser: 1001 seLinuxOptions: {} seccompProfile: type: RuntimeDefault enabled: true image: digest: '' pullPolicy: IfNotPresent registry: docker.hops.works/superset repository: redis-exporter tag: 1.90.0-alpine-h1.1 resources: limits: cpu: 100m memory: 64Mi requests: cpu: 50m memory: 64Mi service: annotations: prometheus.io/port: '{{ .Values.metrics.service.port }}' prometheus.io/scrape: 'true' runAsUser: 1000 secretEnv: create: false service: annotations: consul.hashicorp.com/service-name: superset consul.hashicorp.com/service-tags: app supersetNode: connections: db_host: '{{ .Release.Name }}-mysql' db_name: superset db_pass: superset db_port: '3306' db_type: mysql db_user: superset containerSecurityContext: allowPrivilegeEscalation: false capabilities: drop: - ALL initContainers: - command: - /bin/sh - -c - dockerize -wait "tcp://$DB_HOST:$DB_PORT" -timeout 120s envFrom: - secretRef: name: '{{ tpl .Values.envFromSecret . }}' image: '{{ .Values.initImage.repository }}:{{ .Values.initImage.tag }}' imagePullPolicy: '{{ .Values.initImage.pullPolicy }}' name: wait-for-db resources: limits: cpu: 500m memory: 256Mi requests: cpu: 250m memory: 128Mi securityContext: allowPrivilegeEscalation: false capabilities: drop: - ALL livenessProbe: failureThreshold: 3 httpGet: path: /hopsworks-api/superset/health port: http initialDelaySeconds: 15 periodSeconds: 15 successThreshold: 1 timeoutSeconds: 1 podSecurityContext: runAsNonRoot: true seccompProfile: type: RuntimeDefault readinessProbe: failureThreshold: 3 httpGet: path: /hopsworks-api/superset/health port: http initialDelaySeconds: 15 periodSeconds: 15 successThreshold: 1 timeoutSeconds: 1 replicas: enabled: true replicaCount: 1 startupProbe: failureThreshold: 60 httpGet: path: /hopsworks-api/superset/health port: http initialDelaySeconds: 15 periodSeconds: 5 successThreshold: 1 timeoutSeconds: 1 supersetWorker: command: - /bin/sh - -c - . {{ .Values.configMountPath }}/superset_bootstrap.sh; celery --app=superset.tasks.celery_app:app worker --pool=prefork -O fair -c 4 containerSecurityContext: allowPrivilegeEscalation: false capabilities: drop: - ALL initContainers: - command: - /bin/sh - -c - dockerize -wait "tcp://$DB_HOST:$DB_PORT" -wait "tcp://$REDIS_HOST:$REDIS_PORT" -timeout 120s envFrom: - secretRef: name: '{{ tpl .Values.envFromSecret . }}' image: '{{ .Values.initImage.repository }}:{{ .Values.initImage.tag }}' imagePullPolicy: '{{ .Values.initImage.pullPolicy }}' name: wait-for-db-redis resources: limits: cpu: 500m memory: 256Mi requests: cpu: 250m memory: 128Mi securityContext: allowPrivilegeEscalation: false capabilities: drop: - ALL podSecurityContext: runAsNonRoot: true seccompProfile: type: RuntimeDefault replicas: enabled: false resources: limits: cpu: 500m memory: 1024Mi requests: cpu: 500m memory: 1024Mi ```
`superset` # { #helm.superset } : Type `object`, default `{"mysql":{"enabled":true},"superset":{"redis":{"enabled":true}}}`. override superset values `superset.cleanupOnUninstall` # { #helm.superset.cleanupOnUninstall } : Type `object`, default `{"enabled":true,"ttlSecondsAfterFinished":null}`. post-delete cleanup of the Superset MySQL data PVC for existing clusters whose StatefulSet predates the persistentVolumeClaimRetentionPolicy fix. Gated by global._hopsworks.wipeDataOnUninstall and honors the hopsworks.ai/keep=true label, like the other data-PVC teardowns. New clusters clean up natively via the retention policy (which deletes unconditionally). `superset.cleanupOnUninstall.enabled` # { #helm.superset.cleanupOnUninstall.enabled } : Type `bool`, default `true`. enable the post-delete Superset MySQL PVC cleanup hook (also requires global._hopsworks.wipeDataOnUninstall) `superset.cleanupOnUninstall.ttlSecondsAfterFinished` # { #helm.superset.cleanupOnUninstall.ttlSecondsAfterFinished } : Type `string`, default `nil`. ttlSecondsAfterFinished for the cleanup Job; null falls through to the global default `superset.hopsworkslib` # { #helm.superset.hopsworkslib } : Type `object`, default `{}`. override hopsworkslib values `superset.imageRegistry` # { #helm.superset.imageRegistry } : Type `string`, default `"docker.hops.works/superset"`. `superset.mysql` # { #helm.superset.mysql } : Type `object`, passed to the [`mysql` 12.3.5](https://artifacthub.io/packages/helm/bitnami/mysql/12.3.5) chart, whose other values are documented there. override superset mysql values ??? note "Default" ```yaml architecture: standalone auth: createDatabase: true customPasswordFiles: {} database: superset existingSecret: superset-mysql-users-secrets password: temp-value replicationPassword: '' replicationUser: replicator rootPassword: temp-value usePasswordFiles: false username: superset enabled: true global: security: allowInsecureImages: true image: digest: '' pullPolicy: IfNotPresent registry: docker.hops.works/superset repository: mysql tag: 8.4.11-ubuntu24.04-h4 metrics: enabled: true image: digest: '' pullPolicy: IfNotPresent registry: docker.hops.works/superset repository: mysqld-exporter tag: 0.20.0-alpine-h1.1 resources: limits: cpu: 100m memory: 64Mi requests: cpu: 50m memory: 64Mi service: annotations: prometheus.io/port: '{{ .Values.metrics.service.port }}' prometheus.io/scrape: 'true' primary: persistentVolumeClaimRetentionPolicy: enabled: true whenDeleted: Delete whenScaled: Retain volumePermissions: enabled: false image: digest: '' pullPolicy: IfNotPresent registry: docker.hops.works/superset repository: os-shell tag: 12-alpine-h1.1 ``` `superset.superset` # { #helm.superset.superset } : Type `object`, passed to the [`superset` 0.15.0](https://artifacthub.io/packages/helm/superset/superset/0.15.0) chart, whose other values are documented there. override superset values ??? note "Default" ```yaml _publicRoleDefaultName: Public allowAnonymousAccess: true configOverrides: feature_flags: | FEATURE_FLAGS = {"ALERT_REPORTS": True, "DASHBOARD_RBAC": True} flask_app_configuration: | from flask import session from flask import Flask from datetime import timedelta def make_session_permanent(): ''' Enable maxAge for the cookie 'session' ''' session.permanent = True # Set up max age of session to 24 hours PERMANENT_SESSION_LIFETIME = timedelta(hours=24) # (default: "31 days") SESSION_REFRESH_EACH_REQUEST = True # Default: True def FLASK_APP_MUTATOR(app: Flask) -> None: app.before_request_funcs.setdefault(None, []).append(make_session_permanent) mysql: | SQLALCHEMY_DATABASE_URI = f"mysql+mysqldb://{os.getenv('DB_USER')}:{os.getenv('DB_PASS')}@{os.getenv('DB_HOST')}:{os.getenv('DB_PORT')}/{os.getenv('DB_NAME')}" proxyConfig: | ENABLE_PROXY_FIX = True APP_ICON = "/hopsworks-api/superset/static/assets/images/superset-logo-horiz.png" LOGO_TARGET_PATH = "/hopsworks-api/superset/superset/welcome/" {{- if .Values.init.loadExamples }} PREVENT_UNSAFE_DB_CONNECTIONS = False {{- else }} PREVENT_UNSAFE_DB_CONNECTIONS = True {{- end }} public_role: | {{- if .Values.publicRoleLike }} PUBLIC_ROLE_LIKE = {{ .Values.publicRoleLike | quote }} {{- end }} {{- if .Values.allowAnonymousAccess }} AUTH_ROLE_PUBLIC = {{ .Values._publicRoleDefaultName | quote }} {{- end }} extraEnv: SUPERSET_APP_ROOT: /hopsworks-api/superset extraEnvRaw: - name: SUPERSET_USER valueFrom: secretKeyRef: key: username name: superset-admin-credentials - name: SUPERSET_PASS valueFrom: secretKeyRef: key: password name: superset-admin-credentials - name: DB_PASS valueFrom: secretKeyRef: key: mysql-password name: superset-mysql-users-secrets - name: SUPERSET_SECRET_KEY valueFrom: secretKeyRef: key: secret-key name: superset-secret-key extraRoles: - name: Dataset permissions: - - can_duplicate - Dataset - - can_write - Dataset - - can_get_or_create_dataset - Dataset - - can_warm_up_cache - Dataset extraVolumeMounts: - mountPath: /srv/hops/super_crypto/superset name: super-crypto-material readOnly: true extraVolumes: - name: super-crypto-material secret: optional: true secretName: hopsworks-superset-crypto-material fullnameOverride: hopsworks-superset image: pullPolicy: IfNotPresent repository: docker.hops.works/hopsworks/superset tag: 6.0.0p4.4 init: adminUser: email: admin@superset.com firstname: Superset lastname: Admin password: $SUPERSET_PASS username: $SUPERSET_USER command: - /bin/sh - -c - | {{- if ((((.Values.global | default dict)._hopsworks | default dict).restoreFromBackup | default dict).superset | default dict).enabled }} echo "Superset restore in progress (global._hopsworks.restoreFromBackup.superset.enabled): skipping init so no schema work runs while the database is reloaded. It runs on the upgrade that clears the flag."; exit 0 {{- end }} . {{ .Values.configMountPath }}/superset_bootstrap.sh; . {{ .Values.configMountPath }}/superset_init.sh; {{- if .Values.extraRoles }} {{- range $role := .Values.extraRoles }} python /scripts/create_role.py --role-name {{ $role.name }} --permissions '{{ $role.permissions | toJson | b64enc }}'; {{- end }} {{- end }} {{- if and .Values.publicRolePermissions (eq .Values.publicRoleLike .Values._publicRoleDefaultName) }} python /scripts/create_role.py --role-name {{ .Values._publicRoleDefaultName }} --permissions '{{ .Values.publicRolePermissions | toJson | b64enc }}'; {{- end }} containerSecurityContext: allowPrivilegeEscalation: false capabilities: drop: - ALL initContainers: - command: - /bin/sh - -c - dockerize -wait "tcp://$DB_HOST:$DB_PORT" -timeout 120s envFrom: - secretRef: name: '{{ tpl .Values.envFromSecret . }}' image: '{{ .Values.initImage.repository }}:{{ .Values.initImage.tag }}' imagePullPolicy: '{{ .Values.initImage.pullPolicy }}' name: wait-for-database resources: limits: cpu: 500m memory: 256Mi requests: cpu: 250m memory: 128Mi securityContext: allowPrivilegeEscalation: false capabilities: drop: - ALL loadExamples: false podSecurityContext: runAsNonRoot: true seccompProfile: type: RuntimeDefault initImage: pullPolicy: IfNotPresent repository: docker.hops.works/superset/dockerize tag: 0.14.0-alpine-h1.1 postgresql: enabled: false publicRoleLike: Public publicRolePermissions: - - can_read - Dashboard - - can_read - Chart - - can_dashboard - Superset - - can_slice - Superset - - can_explore_json - Superset - - can_dashboard_permalink - Superset - - can_read - DashboardPermalinkRestApi - - can_read - DashboardFilterStateRestApi - - can_write - DashboardFilterStateRestApi - - can_time_range - Api - - can_query_form_data - Api - - can_query - Api - - can_read - CssTemplate - - can_read - Theme - - can_read - EmbeddedDashboard - - can_read - CurrentUserRestApi - - can_get - Datasource - - can_external_metadata - Datasource - - can_read - Annotation - - can_read - AnnotationLayerRestApi - - can_read - ExplorePermalinkRestApi redis: enabled: true image: registry: docker.hops.works/superset repository: redis tag: 7.4.11-alpine-h1 master: configuration: |- maxmemory 256mb maxmemory-policy allkeys-lru containerSecurityContext: allowPrivilegeEscalation: false capabilities: drop: - ALL enabled: true readOnlyRootFilesystem: true runAsGroup: 1001 runAsNonRoot: true runAsUser: 1001 seLinuxOptions: {} seccompProfile: type: RuntimeDefault podSecurityContext: enabled: true fsGroup: 1001 fsGroupChangePolicy: Always runAsNonRoot: true runAsUser: 1001 seccompProfile: type: RuntimeDefault supplementalGroups: [] sysctls: [] resources: limits: cpu: 500m memory: 384Mi requests: cpu: 100m memory: 384Mi metrics: containerSecurityContext: allowPrivilegeEscalation: false capabilities: drop: - ALL enabled: true readOnlyRootFilesystem: true runAsGroup: 1001 runAsNonRoot: true runAsUser: 1001 seLinuxOptions: {} seccompProfile: type: RuntimeDefault enabled: true image: digest: '' pullPolicy: IfNotPresent registry: docker.hops.works/superset repository: redis-exporter tag: 1.90.0-alpine-h1.1 resources: limits: cpu: 100m memory: 64Mi requests: cpu: 50m memory: 64Mi service: annotations: prometheus.io/port: '{{ .Values.metrics.service.port }}' prometheus.io/scrape: 'true' runAsUser: 1000 secretEnv: create: false service: annotations: consul.hashicorp.com/service-name: superset consul.hashicorp.com/service-tags: app supersetNode: connections: db_host: '{{ .Release.Name }}-mysql' db_name: superset db_pass: superset db_port: '3306' db_type: mysql db_user: superset containerSecurityContext: allowPrivilegeEscalation: false capabilities: drop: - ALL initContainers: - command: - /bin/sh - -c - dockerize -wait "tcp://$DB_HOST:$DB_PORT" -timeout 120s envFrom: - secretRef: name: '{{ tpl .Values.envFromSecret . }}' image: '{{ .Values.initImage.repository }}:{{ .Values.initImage.tag }}' imagePullPolicy: '{{ .Values.initImage.pullPolicy }}' name: wait-for-db resources: limits: cpu: 500m memory: 256Mi requests: cpu: 250m memory: 128Mi securityContext: allowPrivilegeEscalation: false capabilities: drop: - ALL livenessProbe: failureThreshold: 3 httpGet: path: /hopsworks-api/superset/health port: http initialDelaySeconds: 15 periodSeconds: 15 successThreshold: 1 timeoutSeconds: 1 podSecurityContext: runAsNonRoot: true seccompProfile: type: RuntimeDefault readinessProbe: failureThreshold: 3 httpGet: path: /hopsworks-api/superset/health port: http initialDelaySeconds: 15 periodSeconds: 15 successThreshold: 1 timeoutSeconds: 1 replicas: enabled: true replicaCount: 1 startupProbe: failureThreshold: 60 httpGet: path: /hopsworks-api/superset/health port: http initialDelaySeconds: 15 periodSeconds: 5 successThreshold: 1 timeoutSeconds: 1 supersetWorker: command: - /bin/sh - -c - . {{ .Values.configMountPath }}/superset_bootstrap.sh; celery --app=superset.tasks.celery_app:app worker --pool=prefork -O fair -c 4 containerSecurityContext: allowPrivilegeEscalation: false capabilities: drop: - ALL initContainers: - command: - /bin/sh - -c - dockerize -wait "tcp://$DB_HOST:$DB_PORT" -wait "tcp://$REDIS_HOST:$REDIS_PORT" -timeout 120s envFrom: - secretRef: name: '{{ tpl .Values.envFromSecret . }}' image: '{{ .Values.initImage.repository }}:{{ .Values.initImage.tag }}' imagePullPolicy: '{{ .Values.initImage.pullPolicy }}' name: wait-for-db-redis resources: limits: cpu: 500m memory: 256Mi requests: cpu: 250m memory: 128Mi securityContext: allowPrivilegeEscalation: false capabilities: drop: - ALL podSecurityContext: runAsNonRoot: true seccompProfile: type: RuntimeDefault replicas: enabled: false resources: limits: cpu: 500m memory: 1024Mi requests: cpu: 500m memory: 1024Mi ``` `superset.superset.fullnameOverride` # { #helm.superset.superset.fullnameOverride } : Type `string`, default `"hopsworks-superset"`. Provide a name to override the full names of resources `superset.superset.runAsUser` # { #helm.superset.superset.runAsUser } : Type `int`, default `1000`. User ID directive. This user must have enough permissions to run the bootstrap script Running containers as root is not recommended in production. Change this to another UID - e.g. 1000 to be more secure
## auth { #helm-values-superset-auth } ??? example "Defaults as YAML" ```yaml superset: auth: adminUserSecret: superset-admin-credentials adminUsername: adminuser createSecretEnvs: true createSecrets: true mysqlUserSecretName: superset-mysql-users-secrets secretKeySecretName: superset-secret-key ```
`superset.auth.adminUserSecret` # { #helm.superset.auth.adminUserSecret } : Type `string`, default `"superset-admin-credentials"`. `superset.auth.adminUsername` # { #helm.superset.auth.adminUsername } : Type `string`, default `"adminuser"`. `superset.auth.createSecretEnvs` # { #helm.superset.auth.createSecretEnvs } : Type `bool`, default `true`. `superset.auth.createSecrets` # { #helm.superset.auth.createSecrets } : Type `bool`, default `true`. `superset.auth.mysqlUserSecretName` # { #helm.superset.auth.mysqlUserSecretName } : Type `string`, default `"superset-mysql-users-secrets"`. `superset.auth.secretKeySecretName` # { #helm.superset.auth.secretKeySecretName } : Type `string`, default `"superset-secret-key"`.
## backups { #helm-values-superset-backups } ??? example "Defaults as YAML" ```yaml superset: backups: activeDeadlineSeconds: 3600 backoffLimit: 1 enabled: true pathPrefix: superset_backup schedule: null tmpSizeLimit: 2Gi ttl: null ttlSecondsAfterFinished: null ```
`superset.backups` # { #helm.superset.backups } : Type `object`. Superset MySQL backup (HWORKS-2973): a scheduled logical mysqldump of the `superset` schema to the platform object store (same bucket as the other backups), plus an authoritative manifest and a Velero-captured metadata index. Rendered only when global._hopsworks.backups.enabled is true and an object store is configured. ??? note "Default" ```yaml activeDeadlineSeconds: 3600 backoffLimit: 1 enabled: true pathPrefix: superset_backup schedule: null tmpSizeLimit: 2Gi ttl: null ttlSecondsAfterFinished: null ``` `superset.backups.activeDeadlineSeconds` # { #helm.superset.backups.activeDeadlineSeconds } : Type `int`, default `3600`. activeDeadlineSeconds for the backup Job; bounds a hung run so it cannot block later schedules under concurrencyPolicy Forbid, and bounds the dump and upload `superset.backups.backoffLimit` # { #helm.superset.backups.backoffLimit } : Type `int`, default `1`. backoffLimit for the backup Job `superset.backups.enabled` # { #helm.superset.backups.enabled } : Type `bool`, default `true`. enable the scheduled Superset database backup CronJob `superset.backups.pathPrefix` # { #helm.superset.backups.pathPrefix } : Type `string`, default `"superset_backup"`. object-path prefix within the backup bucket `superset.backups.schedule` # { #helm.superset.backups.schedule } : Type `string`, default `nil`. backup schedule (cron or @weekly). null inherits global._hopsworks.backups.schedule, else @weekly `superset.backups.tmpSizeLimit` # { #helm.superset.backups.tmpSizeLimit } : Type `string`, default `"2Gi"`. size limit for the temporary dump volume and the container ephemeral-storage request/limit `superset.backups.ttl` # { #helm.superset.backups.ttl } : Type `string`, default `nil`. retention (Go duration, e.g. 60d) for pruning whole backup sets older than the value. null inherits the global setting `superset.backups.ttlSecondsAfterFinished` # { #helm.superset.backups.ttlSecondsAfterFinished } : Type `string`, default `nil`. ttlSecondsAfterFinished for finished backup Job pods (removes the pod, and its node-local dump scratch, after completion). null falls through to the global default
================================================================================ # trino Source: https://docs.hopsworks.ai/latest/setup_installation/common/helm_chart_values/trino/ # Trino values { #helm-values-trino } Values under `trino` configure Trino, the SQL query engine, and its test coordinator for user catalogs. _Generated from the Hopsworks Helm chart `5.2.0-alpha-1791549041` (Hopsworks `5.2.0`)._ Deployed according to the first of these values that is set: [`global._hopsworks.trino.enabled`](global.md#helm.global._hopsworks.trino.enabled), [`global._hopsworks.full_platform`](global.md#helm.global._hopsworks.full_platform). !!! info "Upstream charts" - Values under `trino.trino` go to [`trino` 1.41.0](https://artifacthub.io/packages/helm/trino/trino/1.41.0) from `https://trinodb.github.io/charts`. - Values under `trino.trinotest` go to [`trino` 1.41.0](https://artifacthub.io/packages/helm/trino/trino/1.41.0) from `https://trinodb.github.io/charts`. Only the values Hopsworks sets under `trino.trino` and `trino.trinotest` are listed on this page. Any other value of the charts can be set under the same keys; each link opens the chart's documentation for the version Hopsworks pins. ## General { #helm-values-trino-general } ??? example "Defaults as YAML" ```yaml trino: catalogsConfigmapName: hopsworks-trino-catalogs hopsworkslib: {} trinotest: initContainers: coordinator: - null - null - name: mountable-secrets ```
`trino` # { #helm.trino } : Type `object`, default `{}`. override trino values `trino.catalogsConfigmapName` # { #helm.trino.catalogsConfigmapName } : Type `string`, default `"hopsworks-trino-catalogs"`. `trino.hopsworkslib` # { #helm.trino.hopsworkslib } : Type `object`, default `{}`. override hopsworkslib values `trino.trinotest` # { #helm.trino.trinotest } : Type `object`, default check \[values.yaml\](./values.yaml) for more information, passed to the [`trino` 1.41.0](https://artifacthub.io/packages/helm/trino/trino/1.41.0) chart, whose other values are documented there. override trinotest values. Rendered only when global._hopsworks.trino.testCoordinator.enabled is true (see Chart.yaml). An optional single-node coordinator running catalog.management=dynamic with a WRITABLE catalog dir, used by the backend to connection-test user catalogs (CREATE CATALOG / SHOW SCHEMAS / DROP CATALOG) before they are synced to the production coordinator. Mirrors the production coordinator's image, certs, TLS, and PASSWORD auth so the backend's admin credentials and discovery work identically. `trino.trinotest.initContainers.coordinator[2].name` # { #helm.trino.trinotest.initContainers.coordinator.2.name } : Type `string`, default `"mountable-secrets"`. readOnly is not optional. Without it the query engine gains a write channel into HopsFS as the trino service user; the volume's readOnly makes the kernel refuse writes too. --srcDir bounds what the mount exposes to the backend-owned tree; the mount authenticates as trino, which is in the hdfs superuser group, so srcDir is the only thing containing it. No preStop unmount: the node plugin owns the mount and unmounts it in NodeUnpublishVolume, and hopsfs-sidecar.sh exits on SIGTERM, so the container terminates cleanly on its own.
## auth { #helm-values-trino-auth } ??? example "Defaults as YAML" ```yaml trino: auth: adminUserPwd: '' adminUserPwdHash: null adminUsername: trino createSecrets: true createSharedSecret: true monitoringUser: prometheus monitoringUserPwd: '' monitoringUserPwdHash: null ```
`trino.auth.adminUserPwd` # { #helm.trino.auth.adminUserPwd } : Type `string`, default `""`. Plaintext password for `adminUsername`, read by the backend and bcrypted into password.db by the seeder Job. Empty generates one on first install and reuses it from the existing Secret; empty in a non-auto mode fails the render, since the existing Secret cannot be read back. `trino.auth.adminUsername` # { #helm.trino.auth.adminUsername } : Type `string`, default `"trino"`. Trino admin username. Should not be changed. Used in hadoop.proxyuser.trino `trino.auth.createSecrets` # { #helm.trino.auth.createSecrets } : Type `bool`, default `true`. If createSecrets is false, you must manually create the following Kubernetes Secrets: 1. trino-admin-credentials (keys `username`, `password`): the backend's principal. 2. trino-monitoring-credentials (same keys): the principal Prometheus scrapes with. The seeder Job bcrypts both into password.db in HopsFS. Label both `backup.hops.works/include: "true"` so Velero captures them. In any non-auto `global._hopsworks.mode` (ArgoCD) the existing Secret cannot be read back, so the render fails unless both passwords below are set or this is false. `trino.auth.createSharedSecret` # { #helm.trino.auth.createSharedSecret } : Type `bool`, default `true`. If `createSharedSecret` is set to `false`, you must generate an `internal-communication.shared-secret` value and store it in a Kubernetes Secret named `trino-internal-secret` under the key `shared-secret`. trino-internal-secret is intentionally NOT captured by the Velero backup (it has no coupling to user state and is regenerated on a fresh install). With createSharedSecret=false it is operator-managed, so disaster recovery must restore it from an independent source, followed by a coordinated restart of all Trino pods so coordinator and workers share the same value. `trino.auth.monitoringUser` # { #helm.trino.auth.monitoringUser } : Type `string`, default `"prometheus"`. Username used by Prometheus for scraping Trino metrics. If you override this value, you MUST also: 1. Update Trino access control to grant this user the required permissions (e.g. adjust `accessControl.rules.rules.json` accordingly). 2. Update the Prometheus scrape configuration so that the same username is used: set `.Values.prometheus.prometheus.serverFiles.prometheus.yml.scrape_configs[*].basic_auth.username` for the scrape job with `job_name: "trino"` to match this value. `trino.auth.monitoringUserPwd` # { #helm.trino.auth.monitoringUserPwd } : Type `string`, default `""`. Plaintext password for `monitoringUser`, read by Prometheus and bcrypted into password.db. Generated like `adminUserPwd`. `trino.auth.adminUserPwdHash` Deprecated # { #helm.trino.auth.adminUserPwdHash } : Type `string`, default `nil`. DEPRECATED, read by nothing: the seeder derives the hash from `adminUserPwd`. Kept so an existing override still validates. `trino.auth.monitoringUserPwdHash` Deprecated # { #helm.trino.auth.monitoringUserPwdHash } : Type `string`, default `nil`. DEPRECATED, read by nothing, like `adminUserPwdHash`.
## authSeeder { #helm-values-trino-authseeder } ??? example "Defaults as YAML" ```yaml trino: authSeeder: image: name: hopsfs registry: '' tag: 3.4.3.3-EE-RC1 ttlSecondsAfterFinished: 3600 waitSeconds: 600 ```
`trino.authSeeder` # { #helm.trino.authSeeder } : Type `object`. The Job that seeds password.db and group.db into HopsFS, and the pre-upgrade hook that migrates them from the legacy Secrets. Not optional: Trino cannot start without them. ??? note "Default" ```yaml image: name: hopsfs registry: '' tag: 3.4.3.3-EE-RC1 ttlSecondsAfterFinished: 3600 waitSeconds: 600 ``` `trino.authSeeder.image` # { #helm.trino.authSeeder.image } : Type `object`, default `{"name":"hopsfs","registry":"","tag":"3.4.3.3-EE-RC1"}`. The HopsFS client image. The seeder writes with `hdfs dfs`; the trino-files mount is read-only. `trino.authSeeder.image.name` # { #helm.trino.authSeeder.image.name } : Type `string`, default `"hopsfs"`. Image name. `trino.authSeeder.image.registry` # { #helm.trino.authSeeder.image.registry } : Type `string`, default `""`. Full registry+path prefix, ending with `/`. Empty uses `global._hopsworks.imageRegistry` plus `/hopsworks/`, matching charts/hopsfs. `trino.authSeeder.image.tag` # { #helm.trino.authSeeder.image.tag } : Type `string`, default `"3.4.3.3-EE-RC1"`. Image tag. Keep in step with `hopsfs.image.tag` in charts/hopsfs/values.yaml. `trino.authSeeder.ttlSecondsAfterFinished` # { #helm.trino.authSeeder.ttlSecondsAfterFinished } : Type `int`, default `3600`. Seconds to keep the finished Jobs; their logs record what was seeded. `trino.authSeeder.waitSeconds` # { #helm.trino.authSeeder.waitSeconds } : Type `int`, default `600`. Seconds the seeder waits for HopsFS before failing.
## dependencies { #helm-values-trino-dependencies } ??? example "Defaults as YAML" ```yaml trino: dependencies: hive: consulServiceName: hive consulServiceTag: metastore port: 9083 mysql: consulServiceName: mysql port: 3306 ```
`trino.dependencies.hive.consulServiceName` # { #helm.trino.dependencies.hive.consulServiceName } : Type `string`, default `"hive"`. `trino.dependencies.hive.consulServiceTag` # { #helm.trino.dependencies.hive.consulServiceTag } : Type `string`, default `"metastore"`. `trino.dependencies.hive.port` # { #helm.trino.dependencies.hive.port } : Type `int`, default `9083`. `trino.dependencies.mysql.consulServiceName` # { #helm.trino.dependencies.mysql.consulServiceName } : Type `string`, default `"mysql"`. `trino.dependencies.mysql.port` # { #helm.trino.dependencies.mysql.port } : Type `int`, default `3306`.
## externalLoadBalancer { #helm-values-trino-externalloadbalancer } ??? example "Defaults as YAML" ```yaml trino: externalLoadBalancer: annotations: {} class: null enabled: null managed: null nodePort: null nodeSelector: {} ```
`trino.externalLoadBalancer.annotations` # { #helm.trino.externalLoadBalancer.annotations } : Type `object`, default `{}`. annotations for load balancer `trino.externalLoadBalancer.class` # { #helm.trino.externalLoadBalancer.class } : Type `string`, default `nil`. load balancer class name `trino.externalLoadBalancer.enabled` # { #helm.trino.externalLoadBalancer.enabled } : Type `string`, default `nil`. Enable External Load Balancers for the Trino coordinator/service. If not set the .global._hopsworks.externalLoadBalancers.enabled will be used instead `trino.externalLoadBalancer.managed` # { #helm.trino.externalLoadBalancer.managed } : Type `string`, default `nil`. Cloud provider provisions Load Balancers. If not set the .global._hopsworks.externalLoadBalancers.managed will be used instead `trino.externalLoadBalancer.nodePort` # { #helm.trino.externalLoadBalancer.nodePort } : Type `string`, default `nil`. Explicit nodePort for the external service when the load balancer is unmanaged (managed: false), so a load balancer outside Kubernetes can target a fixed port. Null lets Kubernetes allocate one from the cluster's node-port range; a set value must lie in that range (30000-32767 by default), which the API server enforces at install. `trino.externalLoadBalancer.nodeSelector` # { #helm.trino.externalLoadBalancer.nodeSelector } : Type `object`, default `{}`. selector for nodes the load balancer can use to route traffic
## trino { #helm-values-trino-trino } ??? example "Defaults as YAML" ```yaml trino: trino: coordinator: additionalJVMConfig: - --add-opens=java.base/java.nio=ALL-UNNAMED image: registry: docker.hops.works repository: hopsworks/trino tag: 483-v1 initContainers: coordinator: - null - null - name: mountable-secrets worker: - null - name: mountable-secrets ```
`trino.trino` # { #helm.trino.trino } : Type `object`, default check \[values.yaml\](./values.yaml) for more information, passed to the [`trino` 1.41.0](https://artifacthub.io/packages/helm/trino/trino/1.41.0) chart, whose other values are documented there. override trino values `trino.trino.coordinator.additionalJVMConfig` # { #helm.trino.trino.coordinator.additionalJVMConfig } : Type `list`, default `["--add-opens=java.base/java.nio=ALL-UNNAMED"]`. add-opens for java.nio, and the failure is a CONFIGURATION error raised while the catalog is being loaded. On this coordinator that is fatal rather than local, because Trino runs catalog.management=static and exits when a catalog file fails to load -- so one project's Snowflake catalog would stop the whole cluster from starting. Set here rather than left to the operator because the connector is in TRINO_CONNECTORS, i.e. the backend offers it to every project. Other Arrow-based connectors need the same opens, so this is not Snowflake-specific. `trino.trino.image.registry` # { #helm.trino.trino.image.registry } : Type `string`, default `"docker.hops.works"`. Image registry, defaults to empty, which results in DockerHub usage `trino.trino.image.repository` # { #helm.trino.trino.image.repository } : Type `string`, default `"hopsworks/trino"`. Repository location of the Trino image, typically `organization/imagename` `trino.trino.image.tag` # { #helm.trino.trino.image.tag } : Type `string`, default `"483-v1"`. Image tag for the Trino image. This value is explicitly pinned here and overrides any defaulting to `appVersion` from Chart.yaml. `trino.trino.initContainers.coordinator[2].name` # { #helm.trino.trino.initContainers.coordinator.2.name } : Type `string`, default `"mountable-secrets"`. readOnly is not optional. Without it the query engine gains a write channel into HopsFS as the trino service user; the volume's readOnly makes the kernel refuse writes too. --srcDir bounds what the mount exposes to the backend-owned tree; the mount authenticates as trino, which is in the hdfs superuser group, so srcDir is the only thing containing it. No preStop unmount: the node plugin owns the mount and unmounts it in NodeUnpublishVolume, and hopsfs-sidecar.sh exits on SIGTERM, so the container terminates cleanly on its own. `trino.trino.initContainers.worker[1].name` # { #helm.trino.trino.initContainers.worker.1.name } : Type `string`, default `"mountable-secrets"`. readOnly is not optional. Without it the query engine gains a write channel into HopsFS as the trino service user; the volume's readOnly makes the kernel refuse writes too. --srcDir bounds what the mount exposes to the backend-owned tree; the mount authenticates as trino, which is in the hdfs superuser group, so srcDir is the only thing containing it. No preStop unmount: the node plugin owns the mount and unmounts it in NodeUnpublishVolume, and hopsfs-sidecar.sh exits on SIGTERM, so the container terminates cleanly on its own.
================================================================================ # vpa Source: https://docs.hopsworks.ai/latest/setup_installation/common/helm_chart_values/vpa/ # Vertical Pod Autoscaler values { #helm-values-vpa } Values under `vpa` configure the Vertical Pod Autoscaler: its admission controller, recommender and updater. _Generated from the Hopsworks Helm chart `5.2.0-alpha-1791549041` (Hopsworks `5.2.0`)._ Deployed according to the first of these values that is set: [`global._hopsworks.vpaEnabled`](global.md#helm.global._hopsworks.vpaEnabled), [`global._hopsworks.full_platform`](global.md#helm.global._hopsworks.full_platform). ??? example "Defaults as YAML" ```yaml vpa: allowOnlyOneReplica: true cleanupOnUninstall: enabled: true ttlSecondsAfterFinished: null enablePrometheusRecommender: true enablePrometheusScrapper: true enabled: true hopsworkslib: {} image: name: k8s-vpa tag: 1.7.1-2 imagesTag: 1.7.1 installController: true recommenderSettings: inRecommendationBoundsEvictionLifetimeThreshold: 12h podUpdateThreshold: '0.1' teardownWhenInstalling: true teardownWhenUninstall: true ```
`vpa` # { #helm.vpa } : Type `object`, default `{}`. override vpa values `vpa.allowOnlyOneReplica` # { #helm.vpa.allowOnlyOneReplica } : Type `bool`, default `true`. If true the vpa updater is configured to update resources even when only one replica is running `vpa.cleanupOnUninstall` # { #helm.vpa.cleanupOnUninstall } : Type `object`, default `{"enabled":true,"ttlSecondsAfterFinished":null}`. post-delete cleanup of the VPA ConfigMap. cm.yaml defines `vpa` as a Helm hook with no delete-policy, so Helm never removes it; this deletes it by name on uninstall. `vpa.cleanupOnUninstall.enabled` # { #helm.vpa.cleanupOnUninstall.enabled } : Type `bool`, default `true`. enable the post-delete VPA ConfigMap cleanup hook `vpa.cleanupOnUninstall.ttlSecondsAfterFinished` # { #helm.vpa.cleanupOnUninstall.ttlSecondsAfterFinished } : Type `string`, default `nil`. ttlSecondsAfterFinished for the cleanup Job; null falls through to the global default `vpa.enablePrometheusRecommender` # { #helm.vpa.enablePrometheusRecommender } : Type `bool`, default `true`. If true the recommender will be connected to prometheus `vpa.enablePrometheusScrapper` # { #helm.vpa.enablePrometheusScrapper } : Type `bool`, default `true`. If true the vpa metrics will be connected to prometheus `vpa.enabled` # { #helm.vpa.enabled } : Type `bool`, default `true`. `vpa.hopsworkslib` # { #helm.vpa.hopsworkslib } : Type `object`, default `{}`. override hopsworkslib values `vpa.image.name` # { #helm.vpa.image.name } : Type `string`, default `"k8s-vpa"`. `vpa.image.tag` # { #helm.vpa.image.tag } : Type `string`, default `"1.7.1-2"`. Installer toolbox version. Independent of imagesTag. `vpa.imagesTag` # { #helm.vpa.imagesTag } : Type `string`, default `"1.7.1"`. Upstream VPA release. Selects the admission-controller, recommender and updater images, and must match the VPA version the installer image was built from. Not tied to image.tag. `vpa.installController` # { #helm.vpa.installController } : Type `bool`, default `true`. If true, install the VPA controller (CRDs, admission controller, recommender, updater). Set to false if VPA is already installed in the cluster. `vpa.recommenderSettings` # { #helm.vpa.recommenderSettings } : Type `object`. Recommender settings ??? note "Default" ```yaml inRecommendationBoundsEvictionLifetimeThreshold: 12h podUpdateThreshold: '0.1' ``` `vpa.recommenderSettings.inRecommendationBoundsEvictionLifetimeThreshold` # { #helm.vpa.recommenderSettings.inRecommendationBoundsEvictionLifetimeThreshold } : Type `string`, default `"12h"`. The default value is 12h. This is the time needed to do a downscale `vpa.recommenderSettings.podUpdateThreshold` # { #helm.vpa.recommenderSettings.podUpdateThreshold } : Type `string`, default `"0.1"`. The default value is 0.1 Set to 0 to always update even if the recommendation is in bounds Notice it needs to be an string `vpa.teardownWhenInstalling` # { #helm.vpa.teardownWhenInstalling } : Type `bool`, default `true`. If true, the vpa will be deleted if already installed `vpa.teardownWhenUninstall` # { #helm.vpa.teardownWhenUninstall } : Type `bool`, default `true`. if true, the vpa will be deleted if already installed when doing helm uninstall
================================================================================ # hopsworks Source: https://docs.hopsworks.ai/latest/python-api/hopsworks/ ::: hopsworks options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hopsworks.login options: heading_level: 2 show_root_heading: true ::: hopsworks.create_project options: heading_level: 2 show_root_heading: true ::: hopsworks.disable_usage_logging options: heading_level: 2 show_root_heading: true ::: hopsworks.get_current_project options: heading_level: 2 show_root_heading: true ::: hopsworks.get_env_vars_api options: heading_level: 2 show_root_heading: true ::: hopsworks.get_sdk_info options: heading_level: 2 show_root_heading: true ::: hopsworks.get_secrets_api options: heading_level: 2 show_root_heading: true ::: hopsworks.get_users_api options: heading_level: 2 show_root_heading: true ::: hopsworks.logout options: heading_level: 2 show_root_heading: true ================================================================================ # alert Source: https://docs.hopsworks.ai/latest/python-api/hopsworks/alert/ ::: hopsworks.alert options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hopsworks_common.alert.Alert options: heading_level: 2 show_root_heading: true ::: hopsworks_common.alert.FeatureGroupAlert options: heading_level: 2 show_root_heading: true ::: hopsworks_common.alert.FeatureViewAlert options: heading_level: 2 show_root_heading: true ::: hopsworks_common.alert.JobAlert options: heading_level: 2 show_root_heading: true ::: hopsworks_common.alert.ProjectAlert options: heading_level: 2 show_root_heading: true ================================================================================ # alert​_receiver Source: https://docs.hopsworks.ai/latest/python-api/hopsworks/alert_receiver/ ::: hopsworks.alert_receiver options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hopsworks_common.alert_receiver.AlertReceiver options: heading_level: 2 show_root_heading: true ::: hopsworks_common.alert_receiver.EmailConfig options: heading_level: 2 show_root_heading: true ::: hopsworks_common.alert_receiver.PagerDutyConfig options: heading_level: 2 show_root_heading: true ::: hopsworks_common.alert_receiver.SlackConfig options: heading_level: 2 show_root_heading: true ::: hopsworks_common.alert_receiver.WebhookConfig options: heading_level: 2 show_root_heading: true ================================================================================ # app Source: https://docs.hopsworks.ai/latest/python-api/hopsworks/app/ ::: hopsworks.app options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hopsworks_common.app.App options: heading_level: 2 show_root_heading: true ================================================================================ # exceptions Source: https://docs.hopsworks.ai/latest/python-api/hopsworks/client/exceptions/ ::: hopsworks.client.exceptions options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hopsworks_common.client.exceptions.DataSourceException options: heading_level: 2 show_root_heading: true ::: hopsworks_common.client.exceptions.DataValidationException options: heading_level: 2 show_root_heading: true ::: hopsworks_common.client.exceptions.DatasetException options: heading_level: 2 show_root_heading: true ::: hopsworks_common.client.exceptions.EnvironmentException options: heading_level: 2 show_root_heading: true ::: hopsworks_common.client.exceptions.ExternalClientError options: heading_level: 2 show_root_heading: true ::: hopsworks_common.client.exceptions.FeatureStoreException options: heading_level: 2 show_root_heading: true ::: hopsworks_common.client.exceptions.GitException options: heading_level: 2 show_root_heading: true ::: hopsworks_common.client.exceptions.HopsworksSSLClientError options: heading_level: 2 show_root_heading: true ::: hopsworks_common.client.exceptions.HuggingFaceImportException options: heading_level: 2 show_root_heading: true ::: hopsworks_common.client.exceptions.JobException options: heading_level: 2 show_root_heading: true ::: hopsworks_common.client.exceptions.JobExecutionException options: heading_level: 2 show_root_heading: true ::: hopsworks_common.client.exceptions.KafkaException options: heading_level: 2 show_root_heading: true ::: hopsworks_common.client.exceptions.OpenSearchException options: heading_level: 2 show_root_heading: true ::: hopsworks_common.client.exceptions.PlatformIntelligenceException options: heading_level: 2 show_root_heading: true ::: hopsworks_common.client.exceptions.ProjectException options: heading_level: 2 show_root_heading: true ::: hopsworks_common.client.exceptions.RestAPIError options: heading_level: 2 show_root_heading: true ::: hopsworks_common.client.exceptions.TransformationFunctionException options: heading_level: 2 show_root_heading: true ::: hopsworks_common.client.exceptions.TrinoException options: heading_level: 2 show_root_heading: true ::: hopsworks_common.client.exceptions.UnknownSecretStorageError options: heading_level: 2 show_root_heading: true ::: hopsworks_common.client.exceptions.VectorDatabaseException options: heading_level: 2 show_root_heading: true ================================================================================ # hopsworks.core Source: https://docs.hopsworks.ai/latest/python-api/hopsworks/core/ ::: hopsworks.core options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hopsworks_common.core.execution_pod_log.ExecutionPodLog options: heading_level: 2 show_root_heading: true ::: hopsworks_common.core.sink_job_configuration.FeatureColumnMapping options: heading_level: 2 show_root_heading: true ::: hopsworks_common.core.sink_job_configuration.FullLoadConfig options: heading_level: 2 show_root_heading: true ::: hopsworks_common.core.sink_job_configuration.LoadingConfig options: heading_level: 2 show_root_heading: true ::: hopsworks_common.core.sink_job_configuration.SinkJobConfiguration options: heading_level: 2 show_root_heading: true ::: hopsworks_common.core.sink_job_configuration.TableIngestionTarget options: heading_level: 2 show_root_heading: true ::: hsfs.core.multi_table_ingestion.MultiTableIngestionJob options: heading_level: 2 show_root_heading: true ================================================================================ # alerts​_api Source: https://docs.hopsworks.ai/latest/python-api/hopsworks/core/alerts_api/ ::: hopsworks.core.alerts_api options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hopsworks_common.core.alerts_api.AlertsApi options: heading_level: 2 show_root_heading: true ================================================================================ # app​_api Source: https://docs.hopsworks.ai/latest/python-api/hopsworks/core/app_api/ ::: hopsworks.core.app_api options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hopsworks_common.core.app_api.AppApi options: heading_level: 2 show_root_heading: true ================================================================================ # dataset​_api Source: https://docs.hopsworks.ai/latest/python-api/hopsworks/core/dataset_api/ ::: hopsworks.core.dataset_api options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hopsworks_common.core.dataset_api.DatasetApi options: heading_level: 2 show_root_heading: true ================================================================================ # env​_var​_api Source: https://docs.hopsworks.ai/latest/python-api/hopsworks/core/env_var_api/ ::: hopsworks.core.env_var_api options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hopsworks_common.core.env_var_api.EnvVarsApi options: heading_level: 2 show_root_heading: true ================================================================================ # environment​_api Source: https://docs.hopsworks.ai/latest/python-api/hopsworks/core/environment_api/ ::: hopsworks.core.environment_api options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hopsworks_common.core.environment_api.EnvironmentApi options: heading_level: 2 show_root_heading: true ================================================================================ # git​_api Source: https://docs.hopsworks.ai/latest/python-api/hopsworks/core/git_api/ ::: hopsworks.core.git_api options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hopsworks_common.core.git_api.GitApi options: heading_level: 2 show_root_heading: true ================================================================================ # job​_api Source: https://docs.hopsworks.ai/latest/python-api/hopsworks/core/job_api/ ::: hopsworks.core.job_api options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hopsworks_common.core.job_api.JobApi options: heading_level: 2 show_root_heading: true ================================================================================ # job​_configuration Source: https://docs.hopsworks.ai/latest/python-api/hopsworks/core/job_configuration/ ::: hopsworks.core.job_configuration options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hopsworks_common.core.job_configuration.JobConfiguration options: heading_level: 2 show_root_heading: true ================================================================================ # kafka​_api Source: https://docs.hopsworks.ai/latest/python-api/hopsworks/core/kafka_api/ ::: hopsworks.core.kafka_api options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hopsworks_common.core.kafka_api.KafkaApi options: heading_level: 2 show_root_heading: true ================================================================================ # opensearch Source: https://docs.hopsworks.ai/latest/python-api/hopsworks/core/opensearch/ ::: hopsworks.core.opensearch options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hopsworks_common.core.opensearch.OpensearchRequestOption options: heading_level: 2 show_root_heading: true ================================================================================ # opensearch​_api Source: https://docs.hopsworks.ai/latest/python-api/hopsworks/core/opensearch_api/ ::: hopsworks.core.opensearch_api options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hopsworks_common.core.opensearch_api.OpenSearchApi options: heading_level: 2 show_root_heading: true ================================================================================ # project​_members​_api Source: https://docs.hopsworks.ai/latest/python-api/hopsworks/core/project_members_api/ ::: hopsworks.core.project_members_api options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hopsworks_common.core.project_members_api.ProjectMembersApi options: heading_level: 2 show_root_heading: true ================================================================================ # rest​_endpoint Source: https://docs.hopsworks.ai/latest/python-api/hopsworks/core/rest_endpoint/ ::: hopsworks.core.rest_endpoint options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hopsworks_common.core.rest_endpoint.AutoPaginationConfig options: heading_level: 2 show_root_heading: true ::: hopsworks_common.core.rest_endpoint.CursorPaginationConfig options: heading_level: 2 show_root_heading: true ::: hopsworks_common.core.rest_endpoint.HeaderCursorPaginationConfig options: heading_level: 2 show_root_heading: true ::: hopsworks_common.core.rest_endpoint.HeaderLinkPaginationConfig options: heading_level: 2 show_root_heading: true ::: hopsworks_common.core.rest_endpoint.JsonLinkPaginationConfig options: heading_level: 2 show_root_heading: true ::: hopsworks_common.core.rest_endpoint.OffsetPaginationConfig options: heading_level: 2 show_root_heading: true ::: hopsworks_common.core.rest_endpoint.PageNumberPaginationConfig options: heading_level: 2 show_root_heading: true ::: hopsworks_common.core.rest_endpoint.PaginationConfig options: heading_level: 2 show_root_heading: true ::: hopsworks_common.core.rest_endpoint.QueryParams options: heading_level: 2 show_root_heading: true ::: hopsworks_common.core.rest_endpoint.RestEndpointConfig options: heading_level: 2 show_root_heading: true ::: hopsworks_common.core.rest_endpoint.SinglePagePaginationConfig options: heading_level: 2 show_root_heading: true ================================================================================ # search​_api Source: https://docs.hopsworks.ai/latest/python-api/hopsworks/core/search_api/ ::: hopsworks.core.search_api options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hopsworks_common.core.search_api.SearchApi options: heading_level: 2 show_root_heading: true ::: hopsworks_common.core.search_api.TagSearchFilter options: heading_level: 2 show_root_heading: true ::: hopsworks_common.search_results.FeatureGroupSearchResult options: heading_level: 2 show_root_heading: true ::: hopsworks_common.search_results.FeatureSearchResult options: heading_level: 2 show_root_heading: true ::: hopsworks_common.search_results.FeatureViewSearchResult options: heading_level: 2 show_root_heading: true ::: hopsworks_common.search_results.FeaturestoreSearchResult options: heading_level: 2 show_root_heading: true ::: hopsworks_common.search_results.Highlights options: heading_level: 2 show_root_heading: true ::: hopsworks_common.search_results.Project options: heading_level: 2 show_root_heading: true ::: hopsworks_common.search_results.SearchResultItem options: heading_level: 2 show_root_heading: true ::: hopsworks_common.search_results.TrainingDatasetSearchResult options: heading_level: 2 show_root_heading: true ================================================================================ # secret​_api Source: https://docs.hopsworks.ai/latest/python-api/hopsworks/core/secret_api/ ::: hopsworks.core.secret_api options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hopsworks_common.core.secret_api.SecretsApi options: heading_level: 2 show_root_heading: true ================================================================================ # superset​_api Source: https://docs.hopsworks.ai/latest/python-api/hopsworks/core/superset_api/ ::: hopsworks.core.superset_api options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hopsworks_common.core.superset_api.SupersetApi options: heading_level: 2 show_root_heading: true ================================================================================ # tag​_schemas​_api Source: https://docs.hopsworks.ai/latest/python-api/hopsworks/core/tag_schemas_api/ ::: hopsworks.core.tag_schemas_api options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hopsworks_common.core.tag_schemas_api.TagSchemasApi options: heading_level: 2 show_root_heading: true ================================================================================ # trino​_api Source: https://docs.hopsworks.ai/latest/python-api/hopsworks/core/trino_api/ ::: hopsworks.core.trino_api options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hopsworks_common.core.trino_api.TrinoApi options: heading_level: 2 show_root_heading: true ================================================================================ # trino​_catalog​_api Source: https://docs.hopsworks.ai/latest/python-api/hopsworks/core/trino_catalog_api/ ::: hopsworks.core.trino_catalog_api options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hopsworks_common.core.trino_catalog_api.TrinoCatalogApi options: heading_level: 2 show_root_heading: true ================================================================================ # users​_api Source: https://docs.hopsworks.ai/latest/python-api/hopsworks/core/users_api/ ::: hopsworks.core.users_api options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hopsworks_common.core.users_api.UsersApi options: heading_level: 2 show_root_heading: true ================================================================================ # env​_var Source: https://docs.hopsworks.ai/latest/python-api/hopsworks/env_var/ ::: hopsworks.env_var options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hopsworks_common.env_var.EnvVar options: heading_level: 2 show_root_heading: true ================================================================================ # environment Source: https://docs.hopsworks.ai/latest/python-api/hopsworks/environment/ ::: hopsworks.environment options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hopsworks_common.environment.Environment options: heading_level: 2 show_root_heading: true ================================================================================ # execution Source: https://docs.hopsworks.ai/latest/python-api/hopsworks/execution/ ::: hopsworks.execution options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hopsworks_common.execution.Execution options: heading_level: 2 show_root_heading: true ================================================================================ # git​_commit Source: https://docs.hopsworks.ai/latest/python-api/hopsworks/git_commit/ ::: hopsworks.git_commit options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hopsworks_common.git_commit.GitCommit options: heading_level: 2 show_root_heading: true ================================================================================ # git​_file​_status Source: https://docs.hopsworks.ai/latest/python-api/hopsworks/git_file_status/ ::: hopsworks.git_file_status options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hopsworks_common.git_file_status.GitFileStatus options: heading_level: 2 show_root_heading: true ================================================================================ # git​_provider Source: https://docs.hopsworks.ai/latest/python-api/hopsworks/git_provider/ ::: hopsworks.git_provider options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hopsworks_common.git_provider.GitProvider options: heading_level: 2 show_root_heading: true ================================================================================ # git​_remote Source: https://docs.hopsworks.ai/latest/python-api/hopsworks/git_remote/ ::: hopsworks.git_remote options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hopsworks_common.git_remote.GitRemote options: heading_level: 2 show_root_heading: true ================================================================================ # git​_repo Source: https://docs.hopsworks.ai/latest/python-api/hopsworks/git_repo/ ::: hopsworks.git_repo options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hopsworks_common.git_repo.GitRepo options: heading_level: 2 show_root_heading: true ================================================================================ # job Source: https://docs.hopsworks.ai/latest/python-api/hopsworks/job/ ::: hopsworks.job options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hopsworks_common.job.Job options: heading_level: 2 show_root_heading: true ================================================================================ # job​_schedule Source: https://docs.hopsworks.ai/latest/python-api/hopsworks/job_schedule/ ::: hopsworks.job_schedule options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hopsworks_common.job_schedule.JobSchedule options: heading_level: 2 show_root_heading: true ================================================================================ # kafka​_schema Source: https://docs.hopsworks.ai/latest/python-api/hopsworks/kafka_schema/ ::: hopsworks.kafka_schema options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hopsworks_common.kafka_schema.KafkaSchema options: heading_level: 2 show_root_heading: true ================================================================================ # kafka​_topic Source: https://docs.hopsworks.ai/latest/python-api/hopsworks/kafka_topic/ ::: hopsworks.kafka_topic options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hopsworks_common.kafka_topic.KafkaTopic options: heading_level: 2 show_root_heading: true ================================================================================ # project Source: https://docs.hopsworks.ai/latest/python-api/hopsworks/project/ ::: hopsworks.project options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hopsworks_common.project.Project options: heading_level: 2 show_root_heading: true ================================================================================ # project​_member Source: https://docs.hopsworks.ai/latest/python-api/hopsworks/project_member/ ::: hopsworks.project_member options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hopsworks_common.project_member.ProjectMember options: heading_level: 2 show_root_heading: true ================================================================================ # secret Source: https://docs.hopsworks.ai/latest/python-api/hopsworks/secret/ ::: hopsworks.secret options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hopsworks_common.secret.Secret options: heading_level: 2 show_root_heading: true ================================================================================ # spark Source: https://docs.hopsworks.ai/latest/python-api/hopsworks/spark/ ::: hopsworks.spark options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hopsworks.spark.build_spark options: heading_level: 2 show_root_heading: true ================================================================================ # tag Source: https://docs.hopsworks.ai/latest/python-api/hopsworks/tag/ ::: hopsworks.tag options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hopsworks_common.tag.Tag options: heading_level: 2 show_root_heading: true ================================================================================ # triggered​_alert Source: https://docs.hopsworks.ai/latest/python-api/hopsworks/triggered_alert/ ::: hopsworks.triggered_alert options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hopsworks_common.triggered_alert.TriggeredAlert options: heading_level: 2 show_root_heading: true ================================================================================ # user Source: https://docs.hopsworks.ai/latest/python-api/hopsworks/user/ ::: hopsworks.user options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hopsworks_common.user.AdminUser options: heading_level: 2 show_root_heading: true ::: hopsworks_common.user.User options: heading_level: 2 show_root_heading: true ================================================================================ # hsfs Source: https://docs.hopsworks.ai/latest/python-api/hsfs/ ::: hsfs options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hsfs.disable_usage_logging options: heading_level: 2 show_root_heading: true ::: hsfs.get_sdk_info options: heading_level: 2 show_root_heading: true ================================================================================ # builtin​_transformations Source: https://docs.hopsworks.ai/latest/python-api/hsfs/builtin_transformations/ ::: hsfs.builtin_transformations options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hsfs.builtin_transformations.equal_frequency_binner options: heading_level: 2 show_root_heading: true ::: hsfs.builtin_transformations.equal_width_binner options: heading_level: 2 show_root_heading: true ::: hsfs.builtin_transformations.impute_category options: heading_level: 2 show_root_heading: true ::: hsfs.builtin_transformations.impute_constant options: heading_level: 2 show_root_heading: true ::: hsfs.builtin_transformations.impute_mean options: heading_level: 2 show_root_heading: true ::: hsfs.builtin_transformations.impute_median options: heading_level: 2 show_root_heading: true ::: hsfs.builtin_transformations.impute_mode options: heading_level: 2 show_root_heading: true ::: hsfs.builtin_transformations.label_encoder options: heading_level: 2 show_root_heading: true ::: hsfs.builtin_transformations.log_transform options: heading_level: 2 show_root_heading: true ::: hsfs.builtin_transformations.min_max_scaler options: heading_level: 2 show_root_heading: true ::: hsfs.builtin_transformations.one_hot_encoder options: heading_level: 2 show_root_heading: true ::: hsfs.builtin_transformations.quantile_binner options: heading_level: 2 show_root_heading: true ::: hsfs.builtin_transformations.quantile_transformer options: heading_level: 2 show_root_heading: true ::: hsfs.builtin_transformations.rank_normalizer options: heading_level: 2 show_root_heading: true ::: hsfs.builtin_transformations.robust_scaler options: heading_level: 2 show_root_heading: true ::: hsfs.builtin_transformations.standard_scaler options: heading_level: 2 show_root_heading: true ::: hsfs.builtin_transformations.top_k_categorical_binner options: heading_level: 2 show_root_heading: true ::: hsfs.builtin_transformations.winsorize options: heading_level: 2 show_root_heading: true ================================================================================ # filter Source: https://docs.hopsworks.ai/latest/python-api/hsfs/constructor/filter/ ::: hsfs.constructor.filter options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hsfs.constructor.filter.Filter options: heading_level: 2 show_root_heading: true ::: hsfs.constructor.filter.Logic options: heading_level: 2 show_root_heading: true ================================================================================ # join Source: https://docs.hopsworks.ai/latest/python-api/hsfs/constructor/join/ ::: hsfs.constructor.join options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hsfs.constructor.join.Join options: heading_level: 2 show_root_heading: true ================================================================================ # lookback Source: https://docs.hopsworks.ai/latest/python-api/hsfs/constructor/lookback/ ::: hsfs.constructor.lookback options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hsfs.constructor.lookback.FeatureGroupLookback options: heading_level: 2 show_root_heading: true ::: hsfs.constructor.lookback.Lookback options: heading_level: 2 show_root_heading: true ================================================================================ # prediction​_times Source: https://docs.hopsworks.ai/latest/python-api/hsfs/constructor/prediction_times/ ::: hsfs.constructor.prediction_times options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hsfs.constructor.prediction_times.PredictionTimes options: heading_level: 2 show_root_heading: true ================================================================================ # query Source: https://docs.hopsworks.ai/latest/python-api/hsfs/constructor/query/ ::: hsfs.constructor.query options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hsfs.constructor.query.Query options: heading_level: 2 show_root_heading: true ================================================================================ # chart Source: https://docs.hopsworks.ai/latest/python-api/hsfs/core/chart/ ::: hsfs.core.chart options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hsfs.core.chart.Chart options: heading_level: 2 show_root_heading: true ================================================================================ # dashboard Source: https://docs.hopsworks.ai/latest/python-api/hsfs/core/dashboard/ ::: hsfs.core.dashboard options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hsfs.core.dashboard.Dashboard options: heading_level: 2 show_root_heading: true ================================================================================ # data​_source Source: https://docs.hopsworks.ai/latest/python-api/hsfs/core/data_source/ ::: hsfs.core.data_source options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hsfs.core.data_source.DataSource options: heading_level: 2 show_root_heading: true ================================================================================ # data​_source​_data Source: https://docs.hopsworks.ai/latest/python-api/hsfs/core/data_source_data/ ::: hsfs.core.data_source_data options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hsfs.core.data_source_data.DataSourceData options: heading_level: 2 show_root_heading: true ================================================================================ # explicit​_provenance Source: https://docs.hopsworks.ai/latest/python-api/hsfs/core/explicit_provenance/ ::: hsfs.core.explicit_provenance options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hsfs.core.explicit_provenance.Artifact options: heading_level: 2 show_root_heading: true ::: hsfs.core.explicit_provenance.Links options: heading_level: 2 show_root_heading: true ================================================================================ # feature​_descriptive​_statistics Source: https://docs.hopsworks.ai/latest/python-api/hsfs/core/feature_descriptive_statistics/ ::: hsfs.core.feature_descriptive_statistics options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hsfs.core.feature_descriptive_statistics.FeatureDescriptiveStatistics options: heading_level: 2 show_root_heading: true ================================================================================ # feature​_logging Source: https://docs.hopsworks.ai/latest/python-api/hsfs/core/feature_logging/ ::: hsfs.core.feature_logging options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hsfs.core.feature_logging.FeatureLogging options: heading_level: 2 show_root_heading: true ================================================================================ # feature​_monitoring​_config Source: https://docs.hopsworks.ai/latest/python-api/hsfs/core/feature_monitoring_config/ ::: hsfs.core.feature_monitoring_config options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hsfs.core.feature_monitoring_config.FeatureMonitoringConfig options: heading_level: 2 show_root_heading: true ================================================================================ # feature​_monitoring​_result Source: https://docs.hopsworks.ai/latest/python-api/hsfs/core/feature_monitoring_result/ ::: hsfs.core.feature_monitoring_result options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hsfs.core.feature_monitoring_result.FeatureMonitoringResult options: heading_level: 2 show_root_heading: true ================================================================================ # feature​_statistics​_config Source: https://docs.hopsworks.ai/latest/python-api/hsfs/core/feature_statistics_config/ ::: hsfs.core.feature_statistics_config options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hsfs.core.feature_statistics_config.FeatureStatisticsConfig options: heading_level: 2 show_root_heading: true ================================================================================ # feature​_statistics​_result Source: https://docs.hopsworks.ai/latest/python-api/hsfs/core/feature_statistics_result/ ::: hsfs.core.feature_statistics_result options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hsfs.core.feature_statistics_result.FeatureStatisticsResult options: heading_level: 2 show_root_heading: true ================================================================================ # inferred​_metadata Source: https://docs.hopsworks.ai/latest/python-api/hsfs/core/inferred_metadata/ ::: hsfs.core.inferred_metadata options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hsfs.core.inferred_metadata.InferredFeature options: heading_level: 2 show_root_heading: true ::: hsfs.core.inferred_metadata.InferredMetadata options: heading_level: 2 show_root_heading: true ================================================================================ # monitoring​_window​_config Source: https://docs.hopsworks.ai/latest/python-api/hsfs/core/monitoring_window_config/ ::: hsfs.core.monitoring_window_config options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hsfs.core.monitoring_window_config.MonitoringWindowConfig options: heading_level: 2 show_root_heading: true ::: hsfs.core.monitoring_window_config.WindowConfigType options: heading_level: 2 show_root_heading: true ================================================================================ # online​_ingestion Source: https://docs.hopsworks.ai/latest/python-api/hsfs/core/online_ingestion/ ::: hsfs.core.online_ingestion options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hsfs.core.online_ingestion.OnlineIngestion options: heading_level: 2 show_root_heading: true ================================================================================ # online​_ingestion​_failure Source: https://docs.hopsworks.ai/latest/python-api/hsfs/core/online_ingestion_failure/ ::: hsfs.core.online_ingestion_failure options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hsfs.core.online_ingestion_failure.OnlineIngestionFailure options: heading_level: 2 show_root_heading: true ================================================================================ # online​_ingestion​_result Source: https://docs.hopsworks.ai/latest/python-api/hsfs/core/online_ingestion_result/ ::: hsfs.core.online_ingestion_result options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hsfs.core.online_ingestion_result.OnlineIngestionResult options: heading_level: 2 show_root_heading: true ================================================================================ # statistics​_comparison​_config Source: https://docs.hopsworks.ai/latest/python-api/hsfs/core/statistics_comparison_config/ ::: hsfs.core.statistics_comparison_config options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hsfs.core.statistics_comparison_config.StatisticsComparisonConfig options: heading_level: 2 show_root_heading: true ================================================================================ # statistics​_comparison​_result Source: https://docs.hopsworks.ai/latest/python-api/hsfs/core/statistics_comparison_result/ ::: hsfs.core.statistics_comparison_result options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hsfs.core.statistics_comparison_result.StatisticsComparisonResult options: heading_level: 2 show_root_heading: true ================================================================================ # embedding Source: https://docs.hopsworks.ai/latest/python-api/hsfs/embedding/ ::: hsfs.embedding options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hsfs.embedding.EmbeddingFeature options: heading_level: 2 show_root_heading: true ::: hsfs.embedding.EmbeddingIndex options: heading_level: 2 show_root_heading: true ::: hsfs.embedding.SimilarityFunctionType options: heading_level: 2 show_root_heading: true ================================================================================ # expectation​_suite Source: https://docs.hopsworks.ai/latest/python-api/hsfs/expectation_suite/ ::: hsfs.expectation_suite options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hsfs.expectation_suite.ExpectationSuite options: heading_level: 2 show_root_heading: true ================================================================================ # feature Source: https://docs.hopsworks.ai/latest/python-api/hsfs/feature/ ::: hsfs.feature options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hsfs.feature.Feature options: heading_level: 2 show_root_heading: true ================================================================================ # feature​_group Source: https://docs.hopsworks.ai/latest/python-api/hsfs/feature_group/ ::: hsfs.feature_group options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hsfs.feature_group.ExternalFeatureGroup options: heading_level: 2 show_root_heading: true ::: hsfs.feature_group.FeatureGroup options: heading_level: 2 show_root_heading: true ::: hsfs.feature_group.FeatureGroupBase options: heading_level: 2 show_root_heading: true ::: hsfs.feature_group.SpineGroup options: heading_level: 2 show_root_heading: true ================================================================================ # feature​_logger Source: https://docs.hopsworks.ai/latest/python-api/hsfs/feature_logger/ ::: hsfs.feature_logger options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hsfs.feature_logger.FeatureLogger options: heading_level: 2 show_root_heading: true ================================================================================ # feature​_logger​_async Source: https://docs.hopsworks.ai/latest/python-api/hsfs/feature_logger_async/ ::: hsfs.feature_logger_async options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hsfs.feature_logger_async.AsyncFeatureLogger options: heading_level: 2 show_root_heading: true ================================================================================ # feature​_store Source: https://docs.hopsworks.ai/latest/python-api/hsfs/feature_store/ ::: hsfs.feature_store options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hsfs.feature_store.FeatureStore options: heading_level: 2 show_root_heading: true ================================================================================ # feature​_view Source: https://docs.hopsworks.ai/latest/python-api/hsfs/feature_view/ ::: hsfs.feature_view options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hsfs.feature_view.FeatureView options: heading_level: 2 show_root_heading: true ================================================================================ # ge​_expectation Source: https://docs.hopsworks.ai/latest/python-api/hsfs/ge_expectation/ ::: hsfs.ge_expectation options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hsfs.ge_expectation.GeExpectation options: heading_level: 2 show_root_heading: true ================================================================================ # ge​_validation​_result Source: https://docs.hopsworks.ai/latest/python-api/hsfs/ge_validation_result/ ::: hsfs.ge_validation_result options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hsfs.ge_validation_result.ValidationResult options: heading_level: 2 show_root_heading: true ================================================================================ # hopsworks​_udf Source: https://docs.hopsworks.ai/latest/python-api/hsfs/hopsworks_udf/ ::: hsfs.hopsworks_udf options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hsfs.hopsworks_udf.HopsworksUdf options: heading_level: 2 show_root_heading: true ::: hsfs.hopsworks_udf.TransformationFeature options: heading_level: 2 show_root_heading: true ::: hsfs.hopsworks_udf.udf options: heading_level: 2 show_root_heading: true ================================================================================ # online​_config Source: https://docs.hopsworks.ai/latest/python-api/hsfs/online_config/ ::: hsfs.online_config options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hsfs.online_config.OnlineConfig options: heading_level: 2 show_root_heading: true ================================================================================ # serving​_key Source: https://docs.hopsworks.ai/latest/python-api/hsfs/serving_key/ ::: hsfs.serving_key options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hsfs.serving_key.ServingKey options: heading_level: 2 show_root_heading: true ================================================================================ # split​_statistics Source: https://docs.hopsworks.ai/latest/python-api/hsfs/split_statistics/ ::: hsfs.split_statistics options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hsfs.split_statistics.SplitStatistics options: heading_level: 2 show_root_heading: true ================================================================================ # statistics Source: https://docs.hopsworks.ai/latest/python-api/hsfs/statistics/ ::: hsfs.statistics options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hsfs.statistics.Statistics options: heading_level: 2 show_root_heading: true ================================================================================ # statistics​_config Source: https://docs.hopsworks.ai/latest/python-api/hsfs/statistics_config/ ::: hsfs.statistics_config options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hsfs.statistics_config.StatisticsConfig options: heading_level: 2 show_root_heading: true ================================================================================ # storage​_connector Source: https://docs.hopsworks.ai/latest/python-api/hsfs/storage_connector/ ::: hsfs.storage_connector options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hsfs.storage_connector.StorageConnector options: heading_level: 2 show_root_heading: true ::: hsfs.storage_connector.AdlsConnector options: heading_level: 2 show_root_heading: true ::: hsfs.storage_connector.BigQueryConnector options: heading_level: 2 show_root_heading: true ::: hsfs.storage_connector.GcsConnector options: heading_level: 2 show_root_heading: true ::: hsfs.storage_connector.GlueConnector options: heading_level: 2 show_root_heading: true ::: hsfs.storage_connector.GoogleSheetsConnector options: heading_level: 2 show_root_heading: true ::: hsfs.storage_connector.HopsFSConnector options: heading_level: 2 show_root_heading: true ::: hsfs.storage_connector.JdbcConnector options: heading_level: 2 show_root_heading: true ::: hsfs.storage_connector.KafkaConnector options: heading_level: 2 show_root_heading: true ::: hsfs.storage_connector.MongoDBConnector options: heading_level: 2 show_root_heading: true ::: hsfs.storage_connector.RedshiftConnector options: heading_level: 2 show_root_heading: true ::: hsfs.storage_connector.S3Connector options: heading_level: 2 show_root_heading: true ::: hsfs.storage_connector.SapHanaConnector options: heading_level: 2 show_root_heading: true ::: hsfs.storage_connector.SnowflakeConnector options: heading_level: 2 show_root_heading: true ::: hsfs.storage_connector.UnityCatalogConnector options: heading_level: 2 show_root_heading: true ::: hsfs.storage_connector.UnityCatalogSparkOptions options: heading_level: 2 show_root_heading: true ================================================================================ # training​_dataset Source: https://docs.hopsworks.ai/latest/python-api/hsfs/training_dataset/ ::: hsfs.training_dataset options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hsfs.training_dataset.TrainingDataset options: heading_level: 2 show_root_heading: true ================================================================================ # training​_dataset​_feature Source: https://docs.hopsworks.ai/latest/python-api/hsfs/training_dataset_feature/ ::: hsfs.training_dataset_feature options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hsfs.training_dataset_feature.TrainingDatasetFeature options: heading_level: 2 show_root_heading: true ================================================================================ # transformation​_function Source: https://docs.hopsworks.ai/latest/python-api/hsfs/transformation_function/ ::: hsfs.transformation_function options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hsfs.transformation_function.TransformationFunction options: heading_level: 2 show_root_heading: true ::: hsfs.transformation_function.TransformationType options: heading_level: 2 show_root_heading: true ================================================================================ # transformation​_statistics Source: https://docs.hopsworks.ai/latest/python-api/hsfs/transformation_statistics/ ::: hsfs.transformation_statistics options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hsfs.transformation_statistics.FeatureTransformationStatistics options: heading_level: 2 show_root_heading: true ::: hsfs.transformation_statistics.TransformationStatistics options: heading_level: 2 show_root_heading: true ================================================================================ # validation​_report Source: https://docs.hopsworks.ai/latest/python-api/hsfs/validation_report/ ::: hsfs.validation_report options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hsfs.validation_report.ValidationReport options: heading_level: 2 show_root_heading: true ================================================================================ # exceptions Source: https://docs.hopsworks.ai/latest/python-api/hsml/client/exceptions/ ::: hsml.client.exceptions options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hopsworks_common.client.exceptions.ModelRegistryException options: heading_level: 2 show_root_heading: true ::: hopsworks_common.client.exceptions.ModelServingException options: heading_level: 2 show_root_heading: true ================================================================================ # explicit​_provenance Source: https://docs.hopsworks.ai/latest/python-api/hsml/core/explicit_provenance/ ::: hsml.core.explicit_provenance options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hsml.core.explicit_provenance.Artifact options: heading_level: 2 show_root_heading: true ::: hsml.core.explicit_provenance.Links options: heading_level: 2 show_root_heading: true ================================================================================ # default​_predictor Source: https://docs.hopsworks.ai/latest/python-api/hsml/default_predictor/ ::: hsml.default_predictor options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hsml.default_predictor.DefaultPredict options: heading_level: 2 show_root_heading: true ::: hsml.default_predictor.PredictionError options: heading_level: 2 show_root_heading: true ::: hsml.default_predictor.run_kserve_wrapper options: heading_level: 2 show_root_heading: true ================================================================================ # deployable​_component Source: https://docs.hopsworks.ai/latest/python-api/hsml/deployable_component/ ::: hsml.deployable_component options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hsml.deployable_component.DeployableComponent options: heading_level: 2 show_root_heading: true ================================================================================ # deployment Source: https://docs.hopsworks.ai/latest/python-api/hsml/deployment/ ::: hsml.deployment options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hsml.deployment.Deployment options: heading_level: 2 show_root_heading: true ================================================================================ # deployment​_logging​_config Source: https://docs.hopsworks.ai/latest/python-api/hsml/deployment_logging_config/ ::: hsml.deployment_logging_config options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hsml.deployment_logging_config.DeploymentLoggingConfig options: heading_level: 2 show_root_heading: true ================================================================================ # deployment​_schema Source: https://docs.hopsworks.ai/latest/python-api/hsml/deployment_schema/ ::: hsml.deployment_schema options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hsml.deployment_schema.DeploymentSchema options: heading_level: 2 show_root_heading: true ::: hsml.deployment_schema.DeploymentSchemaError options: heading_level: 2 show_root_heading: true ::: hsml.deployment_schema.SchemaField options: heading_level: 2 show_root_heading: true ================================================================================ # deployment​_tracing​_config Source: https://docs.hopsworks.ai/latest/python-api/hsml/deployment_tracing_config/ ::: hsml.deployment_tracing_config options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hsml.deployment_tracing_config.DeploymentTracingConfig options: heading_level: 2 show_root_heading: true ================================================================================ # deployment​_version Source: https://docs.hopsworks.ai/latest/python-api/hsml/deployment_version/ ::: hsml.deployment_version options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hsml.deployment_version.DeploymentVersion options: heading_level: 2 show_root_heading: true ================================================================================ # inference​_batcher Source: https://docs.hopsworks.ai/latest/python-api/hsml/inference_batcher/ ::: hsml.inference_batcher options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hsml.inference_batcher.InferenceBatcher options: heading_level: 2 show_root_heading: true ================================================================================ # inference​_logger Source: https://docs.hopsworks.ai/latest/python-api/hsml/inference_logger/ ::: hsml.inference_logger options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hsml.inference_logger.InferenceLogger options: heading_level: 2 show_root_heading: true ================================================================================ # signature Source: https://docs.hopsworks.ai/latest/python-api/hsml/llm/signature/ ::: hsml.llm.signature options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hsml.llm.signature.create_model options: heading_level: 2 show_root_heading: true ================================================================================ # model Source: https://docs.hopsworks.ai/latest/python-api/hsml/model/ ::: hsml.model options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hsml.model.Model options: heading_level: 2 show_root_heading: true ================================================================================ # model​_registry Source: https://docs.hopsworks.ai/latest/python-api/hsml/model_registry/ ::: hsml.model_registry options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hsml.model_registry.ModelRegistry options: heading_level: 2 show_root_heading: true ================================================================================ # model​_schema Source: https://docs.hopsworks.ai/latest/python-api/hsml/model_schema/ ::: hsml.model_schema options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hsml.model_schema.ModelSchema options: heading_level: 2 show_root_heading: true ================================================================================ # model​_serving Source: https://docs.hopsworks.ai/latest/python-api/hsml/model_serving/ ::: hsml.model_serving options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hsml.model_serving.ModelServing options: heading_level: 2 show_root_heading: true ================================================================================ # predictor Source: https://docs.hopsworks.ai/latest/python-api/hsml/predictor/ ::: hsml.predictor options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hsml.predictor.Predictor options: heading_level: 2 show_root_heading: true ================================================================================ # predictor​_state Source: https://docs.hopsworks.ai/latest/python-api/hsml/predictor_state/ ::: hsml.predictor_state options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hsml.predictor_state.PredictorState options: heading_level: 2 show_root_heading: true ================================================================================ # predictor​_state​_condition Source: https://docs.hopsworks.ai/latest/python-api/hsml/predictor_state_condition/ ::: hsml.predictor_state_condition options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hsml.predictor_state_condition.PredictorStateCondition options: heading_level: 2 show_root_heading: true ================================================================================ # signature Source: https://docs.hopsworks.ai/latest/python-api/hsml/python/signature/ ::: hsml.python.signature options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hsml.python.signature.create_model options: heading_level: 2 show_root_heading: true ================================================================================ # resources Source: https://docs.hopsworks.ai/latest/python-api/hsml/resources/ ::: hsml.resources options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hsml.resources.Resources options: heading_level: 2 show_root_heading: true ::: hsml.resources.ComponentResources options: heading_level: 2 show_root_heading: true ::: hsml.resources.PredictorResources options: heading_level: 2 show_root_heading: true ::: hsml.resources.TransformerResources options: heading_level: 2 show_root_heading: true ================================================================================ # scaling​_config Source: https://docs.hopsworks.ai/latest/python-api/hsml/scaling_config/ ::: hsml.scaling_config options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hsml.scaling_config.ComponentScalingConfig options: heading_level: 2 show_root_heading: true ::: hsml.scaling_config.LogPersistence options: heading_level: 2 show_root_heading: true ::: hsml.scaling_config.PredictorScalingConfig options: heading_level: 2 show_root_heading: true ::: hsml.scaling_config.ScaleMetric options: heading_level: 2 show_root_heading: true ::: hsml.scaling_config.TransformerScalingConfig options: heading_level: 2 show_root_heading: true ================================================================================ # schema Source: https://docs.hopsworks.ai/latest/python-api/hsml/schema/ ::: hsml.schema options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hsml.schema.Schema options: heading_level: 2 show_root_heading: true ================================================================================ # signature Source: https://docs.hopsworks.ai/latest/python-api/hsml/sklearn/signature/ ::: hsml.sklearn.signature options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hsml.sklearn.signature.create_model options: heading_level: 2 show_root_heading: true ================================================================================ # signature Source: https://docs.hopsworks.ai/latest/python-api/hsml/tensorflow/signature/ ::: hsml.tensorflow.signature options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hsml.tensorflow.signature.create_model options: heading_level: 2 show_root_heading: true ================================================================================ # signature Source: https://docs.hopsworks.ai/latest/python-api/hsml/torch/signature/ ::: hsml.torch.signature options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hsml.torch.signature.create_model options: heading_level: 2 show_root_heading: true ================================================================================ # transformer Source: https://docs.hopsworks.ai/latest/python-api/hsml/transformer/ ::: hsml.transformer options: heading_level: 1 members: false show_root_full_path: true show_root_heading: true ::: hsml.transformer.Transformer options: heading_level: 2 show_root_heading: true ================================================================================ # REST API Status Codes Source: https://docs.hopsworks.ai/latest/reference/rest_error_codes/ # REST API Status Codes Hopsworks REST API responses carry a numeric status code alongside the HTTP status. The code is namespaced by resource category so that the same HTTP status (for example `400 BAD_REQUEST`) can be distinguished by the failure it actually represents. Most codes are errors, but a few categories also carry success codes for informational responses, so both the code and the HTTP status are shown for every entry. The numbering convention is a total of 6 digits: the first 2 digits indicate the category and the last 4 the code within that category. - Service error codes start with `10` - Dataset error codes start with `11` - Generic error codes start with `12` - Job error codes start with `13` - Request error codes start with `14` - Project error codes start with `15` - User and Security error codes start with `20` - Dela error codes start with `17` - Metadata error codes start with `18` - Kafka error codes start with `19` - CA error codes start with `22` - DelaCSR error codes start with `23` - Serving error codes start with `24` - Inference error codes start with `25` - Activities error codes start with `26` - Featurestore error codes start with `27` - Python error codes start with `28` The Schema Registry category is a documented exception to this convention: it mirrors the Confluent Schema Registry's own error codes, which are 5 digits starting with the HTTP status code (for example `50001` for `500 Internal Server Error`). It is listed last on this page for that reason. This page is generated from `RESTCodes.java` in the `hopsworks-ee` product source by `scripts/gen_error_codes.py`. Do not hand-edit the tables below; regenerate them instead. ## ServiceErrorCode | Code | Name | HTTP status | Message | | --- | --- | --- | --- | | 100001 | `JUPYTER_ADD_FAILURE` | 400 BAD_REQUEST | Failed to create Jupyter notebook dir. Jupyter will not work properly. Try recreating the following dir manually. | | 100002 | `OPENSEARCH_SERVER_NOT_AVAILABLE` | 400 BAD_REQUEST | The OpenSearch Server is either down or misconfigured. | | 100003 | `OPENSEARCH_SERVER_NOT_FOUND` | 503 SERVICE_UNAVAILABLE | Problem when reaching the OpenSearch server | | 100004 | `HIVE_ADD_FAILURE` | 400 BAD_REQUEST | Failed to create the Hive database | | 100005 | `LLAP_STATUS_INVALID` | 400 BAD_REQUEST | Unrecognized new LLAP status | | 100006 | `LLAP_CLUSTER_ALREADY_UP` | 400 BAD_REQUEST | LLAP cluster already up | | 100007 | `LLAP_CLUSTER_ALREADY_DOWN` | 400 BAD_REQUEST | LLAP cluster already down | | 100008 | `DATABASE_UNAVAILABLE` | 503 SERVICE_UNAVAILABLE | The database is temporarily unavailable. Please try again later | | 100010 | `ZOOKEEPER_SERVICE_UNAVAILABLE` | 503 SERVICE_UNAVAILABLE | ZooKeeper service unavailable | | 100011 | `ANACONDA_NODES_UNAVAILABLE` | 503 SERVICE_UNAVAILABLE | No conda machine is enabled. Contact the administrator. | | 100012 | `OPENSEARCH_INDEX_CREATION_ERROR` | 500 INTERNAL_SERVER_ERROR | Error while creating index in opensearch | | 100013 | `ANACONDA_LIST_LIB_FORMAT_ERROR` | 500 INTERNAL_SERVER_ERROR | Problem listing libraries. Did conda get upgraded and change its output format? | | 100014 | `ANACONDA_LIST_LIB_ERROR` | 500 INTERNAL_SERVER_ERROR | Problem listing libraries. Please contact the Administrator | | 100016 | `JUPYTER_HOME_ERROR` | 500 INTERNAL_SERVER_ERROR | Couldn't resolve JUPYTER_HOME using DB. | | 100017 | `JUPYTER_STOP_ERROR` | 500 INTERNAL_SERVER_ERROR | Couldn't stop Jupyter Notebook Server. | | 100018 | `INVALID_YML` | 400 BAD_REQUEST | Invalid .yml file | | 100019 | `INVALID_YML_SIZE` | 500 INTERNAL_SERVER_ERROR | .yml file too large. Please set a higher value for variable max_env_yml_byte_size | | 100020 | `ANACONDA_FROM_YML_ERROR` | 500 INTERNAL_SERVER_ERROR | Failed to create Anaconda environment from .yml file. | | 100021 | `PYTHON_INVALID_VERSION` | 400 BAD_REQUEST | Invalid version of python (valid: '3.7' | | 100022 | `ANACONDA_REPO_ERROR` | 500 INTERNAL_SERVER_ERROR | Problem adding the repo. | | 100023 | `ANACONDA_OP_IN_PROGRESS` | 412 PRECONDITION_FAILED | A conda environment operation is currently executing (create/remove/list). Wait for it to finish or clear it first. | | 100024 | `HOST_TYPE_NOT_FOUND` | 412 PRECONDITION_FAILED | No hosts with the desired capability. | | 100025 | `HOST_NOT_FOUND` | 404 NOT_FOUND | Host was not found. | | 100026 | `HOST_NOT_REGISTERED` | 404 NOT_FOUND | Host has not registered. | | 100027 | `ANACONDA_DEP_REMOVE_FORBIDDEN` | 400 BAD_REQUEST | Could not uninstall library, it is a mandatory dependency | | 100028 | `ANACONDA_DEP_INSTALL_FORBIDDEN` | 409 CONFLICT | Library is already installed | | 100029 | `ANACONDA_EXPORT_ERROR` | 500 INTERNAL_SERVER_ERROR | Failed to export Anaconda environment. | | 100030 | `ANACONDA_LIST_LIB_NOT_FOUND` | 204 NO_CONTENT | No results found | | 100031 | `OPENSEARCH_INDEX_NOT_FOUND` | 404 NOT_FOUND | Index was not found in OpenSearch | | 100032 | `OPENSEARCH_INDEX_TYPE_NOT_FOUND` | 404 NOT_FOUND | Index type was not found in OpenSearch | | 100033 | `JUPYTER_SERVERS_NOT_FOUND` | 404 NOT_FOUND | Could not find any Jupyter notebook servers for this project. | | 100034 | `JUPYTER_SERVERS_NOT_RUNNING` | 412 PRECONDITION_FAILED | Could not find any Jupyter notebook servers for this project. | | 100035 | `JUPYTER_START_ERROR` | 500 INTERNAL_SERVER_ERROR | Jupyter server could not start. | | 100036 | `JUPYTER_SAVE_SETTINGS_ERROR` | 500 INTERNAL_SERVER_ERROR | Could not save Jupyter Settings. | | 100037 | `IPYTHON_CONVERT_ERROR` | 500 INTERNAL_SERVER_ERROR | Problem converting ipython notebook to python program | | 100039 | `EMAIL_SENDING_FAILURE` | 500 INTERNAL_SERVER_ERROR | Could not send email | | 100040 | `HOST_EXISTS` | 409 CONFLICT | Host exists | | 100041 | `TENSORFLOW_VERSION_NOT_SUPPORTED` | 400 BAD_REQUEST | We currently do not support this version of TensorFlow. Update to a newer version or contact an admin | | 100042 | `SERVICE_GENERIC_ERROR` | 500 INTERNAL_SERVER_ERROR | Generic error while enabling the service | | 100043 | `JUPYTER_SERVER_ALREADY_RUNNING` | 400 BAD_REQUEST | Jupyter Notebook Server is already running | | 100044 | `ERROR_EXECUTING_REMOTE_COMMAND` | 500 INTERNAL_SERVER_ERROR | Error executing command over SSH | | 100045 | `OPERATION_NOT_SUPPORTED` | 400 BAD_REQUEST | Supplied operation is not supported | | 100046 | `GIT_COMMAND_FAILURE` | 400 BAD_REQUEST | Git command failed to execute | | 100047 | `JUPYTER_NOTEBOOK_VERSIONING_FAILED` | 500 INTERNAL_SERVER_ERROR | Failed to version notebook | | 100048 | `SERVICE_NOT_FOUND` | 404 NOT_FOUND | Service not found | | 100049 | `ACTION_FORBIDDEN` | 400 BAD_REQUEST | Action forbidden | | 100050 | `VARIABLE_NOT_FOUND` | 404 NOT_FOUND | Requested variable not found | | 100051 | `DOCKER_IMAGE_CREATION_ERROR` | 500 INTERNAL_SERVER_ERROR | Error while creating the docker image | | 100052 | `METASTORE_CONNECTION_ERROR` | 500 INTERNAL_SERVER_ERROR | Error opening connection with the Hive metastore | | 100053 | `SERVICE_DISCOVERY_ERROR` | 500 INTERNAL_SERVER_ERROR | Service not found | | 100054 | `WRONG_HDFS_USERNAME_PROVIDED_FOR_ATTACHING_JUPYTER_CONFIGURATION_TO_NOTEBOOK` | 400 BAD_REQUEST | Failed to attach jupyter configuration to notebook. Wrong hdfs username provided | | 100055 | `ATTACHING_JUPYTER_CONFIG_TO_NOTEBOOK_FAILED` | 500 INTERNAL_SERVER_ERROR | Failed to attach jupyter configuration to notebook | | 100056 | `RM_METRICS_ERROR` | 500 INTERNAL_SERVER_ERROR | Failed to fetch utilization metrics | | 100057 | `PROMETHEUS_QUERY_ERROR` | 500 INTERNAL_SERVER_ERROR | Failed to execute prometheus query | | 100058 | `GRAFANA_PROXY_ERROR` | 500 INTERNAL_SERVER_ERROR | Unauthorized access to dashboard | | 100059 | `DOCKER_ERROR` | 500 INTERNAL_SERVER_ERROR | Failed to run docker command | | 100060 | `LOCAL_FILESYSTEM_ERROR` | 500 INTERNAL_SERVER_ERROR | Failed to write to local filesystem | | 100061 | `INVALID_DOCKER_COMMAND_FILE` | 400 BAD_REQUEST | Invalid commands file provided | | 100062 | `INVALID_ARTIFACT_FOR_DOCKER_COMMANDS` | 400 BAD_REQUEST | Invalid artifact provided for docker commands | | 100063 | `ENVIRONMENT_YAML_READ_ERROR` | 500 INTERNAL_SERVER_ERROR | Failed to read yaml file | | 100064 | `ENVIRONMENT_BUILD_NOT_FOUND` | 404 NOT_FOUND | Build not found in environment history | | 100065 | `ENVIRONMENT_HISTORY_READ_ERROR` | 500 INTERNAL_SERVER_ERROR | Failed to read environment history record from database | | 100066 | `ENVIRONMENT_HISTORY_CUSTOM_COMMANDS_FILE_READ_ERROR` | 500 INTERNAL_SERVER_ERROR | Failed to read custom command file | | 100067 | `KUBE_CLIENT_ERROR` | 500 INTERNAL_SERVER_ERROR | Kubernetes client error | | 100068 | `WRONG_HDFS_USERNAME_PROVIDED_FOR_SAVING_SPARK_SESSION_FROM_NOTEBOOK` | 400 BAD_REQUEST | Failed to save spark session from jupyter server | | 100069 | `WRONG_HDFS_USERNAME_PROVIDED_FOR_JUPYTER_RAY_SESSION_OPERATION` | 400 BAD_REQUEST | Hdfs username provided for jupyter ray session operation does not exist | | 100070 | `KERNEL_NOT_FOUND` | 404 NOT_FOUND | Kernel id not found in the running jupyter notebook server | | 100071 | `RAY_SESSION_NOT_FOUND` | 404 NOT_FOUND | Ray session not found | | 100072 | `INVALID_CUSTOM_COMMAND_ENV_VARIABLES` | 400 BAD_REQUEST | Invalid custom command environment variables | | 100073 | `TERMINAL_ERROR` | 500 INTERNAL_SERVER_ERROR | Terminal error | | 100074 | `WEBSOCKET_POOL_FULL` | 503 SERVICE_UNAVAILABLE | The cluster has reached the limit for active sessions across Jupyter notebooks, terminals, and apps. Starting a new session will fail until existing sessions are closed. If this happens often, contact your administrator to raise the limit. | | 100075 | `NPM_REGISTRY_UNAVAILABLE` | 503 SERVICE_UNAVAILABLE | The npm registry could not be reached. | | 100076 | `NPM_PACKAGE_NOT_FOUND` | 404 NOT_FOUND | No such npm package. | | 100077 | `RDRS_UNAVAILABLE` | 503 SERVICE_UNAVAILABLE | The RonDB REST Server could not be reached. | | 100078 | `TTL_PURGE_TABLE_NOT_FOUND` | 404 NOT_FOUND | The TTL purge worker is not tracking this table. | | 100079 | `TERMINAL_INVALID_HOURS` | 400 BAD_REQUEST | The requested terminal session length is out of range. | ## DatasetErrorCode | Code | Name | HTTP status | Message | | --- | --- | --- | --- | | 110000 | `DATASET_OPERATION_FORBIDDEN` | 403 FORBIDDEN | Dataset/content operation forbidden | | 110001 | `DATASET_OPERATION_INVALID` | 400 BAD_REQUEST | Operation cannot be performed. | | 110002 | `DATASET_OPERATION_ERROR` | 500 INTERNAL_SERVER_ERROR | Dataset operation failed. | | 110003 | `DATASET_ALREADY_SHARED_WITH_PROJECT` | 400 BAD_REQUEST | Dataset already shared with project | | 110004 | `DATASET_NOT_SHARED_WITH_PROJECT` | 400 BAD_REQUEST | Dataset is not shared with project. | | 110005 | `DATASET_NAME_EMPTY` | 400 BAD_REQUEST | DataSet name cannot be empty. | | 110006 | `FILE_CORRUPTED_REMOVED_FROM_HDFS` | 400 BAD_REQUEST | Corrupted file removed from hdfs. | | 110007 | `INODE_DELETION_ERROR` | 500 INTERNAL_SERVER_ERROR | File/Dir could not be deleted. | | 110008 | `INODE_NOT_FOUND` | 404 NOT_FOUND | File not found. | | 110009 | `DATASET_REMOVED_FROM_HDFS` | 400 BAD_REQUEST | DataSet removed from hdfs. | | 110010 | `SHARED_DATASET_REMOVED` | 400 BAD_REQUEST | The shared dataset has been removed from this project. | | 110011 | `DATASET_NOT_FOUND` | 400 BAD_REQUEST | DataSet not found. | | 110012 | `DESTINATION_EXISTS` | 400 BAD_REQUEST | Destination already exists. | | 110013 | `DATASET_ALREADY_PUBLIC` | 409 CONFLICT | Dataset is already public. | | 110014 | `DATASET_ALREADY_IN_PROJECT` | 400 BAD_REQUEST | Dataset is already in project. | | 110015 | `DATASET_NOT_PUBLIC` | 400 BAD_REQUEST | DataSet is not public. | | 110016 | `DATASET_NOT_EDITABLE` | 400 BAD_REQUEST | DataSet is not editable. | | 110017 | `DATASET_PENDING` | 400 BAD_REQUEST | DataSet is not yet accessible. Accept the share request to access it. | | 110018 | `PATH_NOT_FOUND` | 400 BAD_REQUEST | Path not found | | 110019 | `PATH_NOT_DIRECTORY` | 400 BAD_REQUEST | Requested path is not a directory | | 110020 | `PATH_IS_DIRECTORY` | 400 BAD_REQUEST | Requested path is a directory | | 110021 | `DOWNLOAD_ERROR` | 400 BAD_REQUEST | Failed to download. | | 110022 | `DOWNLOAD_PERMISSION_ERROR` | 400 BAD_REQUEST | Your role does not allow to download this file | | 110023 | `DATASET_PERMISSION_ERROR` | 500 INTERNAL_SERVER_ERROR | Could not update dataset permissions | | 110024 | `COMPRESSION_ERROR` | 500 INTERNAL_SERVER_ERROR | Error while performing a (un)compress operation | | 110025 | `DATASET_OWNER_ERROR` | 400 BAD_REQUEST | You cannot perform this action on a dataset you are not the owner | | 110026 | `DATASET_PUBLIC_IMMUTABLE` | 400 BAD_REQUEST | Public datasets are immutable. | | 110028 | `DATASET_NAME_INVALID` | 400 BAD_REQUEST | Name of dir is invalid | | 110029 | `IMAGE_SIZE_INVALID` | 400 BAD_REQUEST | Image is too big to display please download it by double-clicking it instead | | 110030 | `FILE_PREVIEW_ERROR` | 400 BAD_REQUEST | README.md too large to be previewd | | 110031 | `DATASET_PARAMETERS_INVALID` | 400 BAD_REQUEST | Invalid parameters for requested dataset operation | | 110032 | `EMPTY_PATH` | 400 BAD_REQUEST | Empty path requested | | 110033 | `ONGOING_PERMISSION_OPERATION` | 409 CONFLICT | There is an ongoing permission operation | | 110035 | `UPLOAD_PATH_NOT_SPECIFIED` | 400 BAD_REQUEST | The path to upload the template was not specified | | 110036 | `README_NOT_ACCESSIBLE` | 401 UNAUTHORIZED | Readme not accessible. | | 110037 | `COMPRESSION_SIZE_ERROR` | 412 PRECONDITION_FAILED | Not enough free space on the local scratch directory to download and unzip this file. Talk to your admin to increase disk space at the path: hopsworks/staging_dir | | 110038 | `INVALID_PATH_FILE` | 400 BAD_REQUEST | The requested path does not resolve to a valid file | | 110039 | `INVALID_PATH_DIR` | 400 BAD_REQUEST | The requested path does not resolve to a valid directory | | 110040 | `UPLOAD_DIR_CREATE_ERROR` | 500 INTERNAL_SERVER_ERROR | Uploads directory could not be created in the file system | | 110041 | `UPLOAD_CONCURRENT_ERROR` | 412 PRECONDITION_FAILED | A file with the same name is being uploaded | | 110042 | `UPLOAD_RESUMABLEINFO_INVALID` | 400 BAD_REQUEST | ResumableInfo is invalid | | 110043 | `UPLOAD_ERROR` | 500 INTERNAL_SERVER_ERROR | Error occurred while uploading file | | 110044 | `DATASET_REQUEST_EXISTS` | 409 CONFLICT | Request for this dataset from this project already exists. | | 110045 | `COPY_FROM_PROJECT` | 403 FORBIDDEN | Cannot copy file/folder from another project | | 110046 | `COPY_TO_PUBLIC_DS` | 403 FORBIDDEN | Can not copy to a public dataset. | | 110047 | `DATASET_SUBDIR_ALREADY_EXISTS` | 400 BAD_REQUEST | A sub-directory with the same name already exists. | | 110048 | `DOWNLOAD_NOT_ALLOWED` | 403 FORBIDDEN | Downloading files is not allowed. Please contact the system administrator for further information. | | 110049 | `DATASET_REQUEST_ERROR` | 500 INTERNAL_SERVER_ERROR | Could not send dataset request | | 110050 | `DATASET_ACCESS_PERMISSION_DENIED` | 403 FORBIDDEN | Permission denied. | | 110051 | `PATH_ENCODING_NOT_SUPPORTED` | 400 BAD_REQUEST | Unsupported encoding. | | 110052 | `ATTACH_XATTR_ERROR` | 500 INTERNAL_SERVER_ERROR | Failed to attach Xattr. | | 110053 | `TARGET_PROJECT_NOT_FOUND` | 500 INTERNAL_SERVER_ERROR | Target project not found. | | 110054 | `DATASET_PERMISSION_IMMUTABLE` | 400 BAD_REQUEST | Internal datasets permission can not be changed. | | 110055 | `UPLOAD_DISK_SPACE_ERROR` | 500 INTERNAL_SERVER_ERROR | Upload failed: HopsFS storage is full. Please contact your administrator to free up disk space. | | 110056 | `UPLOAD_NOT_ALLOWED` | 403 FORBIDDEN | Uploading files is not allowed by the cluster upload policy. Please contact your administrator. | ## GenericErrorCode | Code | Name | HTTP status | Message | | --- | --- | --- | --- | | 120000 | `UNKNOWN_ERROR` | 500 INTERNAL_SERVER_ERROR | A generic error occurred. | | 120001 | `ILLEGAL_ARGUMENT` | 422 UNPROCESSABLE_ENTITY | An argument was not provided or it was malformed. | | 120002 | `ILLEGAL_STATE` | 400 BAD_REQUEST | A runtime error occurred. | | 120003 | `ROLLBACK` | 500 INTERNAL_SERVER_ERROR | The last transaction did not complete as expected | | 120004 | `WEBAPPLICATION` | dynamic (from wrapped exception) | Web application exception occurred | | 120005 | `PERSISTENCE_ERROR` | 500 INTERNAL_SERVER_ERROR | Persistence error occurred | | 120006 | `UNKNOWN_ACTION` | 400 BAD_REQUEST | This action can not be applied on this resource. | | 120007 | `INCOMPLETE_REQUEST` | 400 BAD_REQUEST | Some parameters were not provided or were not in the required format. | | 120008 | `SECURITY_EXCEPTION` | 500 INTERNAL_SERVER_ERROR | A Java security error occurred. | | 120009 | `ENDPOINT_ANNOTATION_MISSING` | 503 SERVICE_UNAVAILABLE | The requested endpoint did not have any project role annotation | | 120010 | `ENTERPRISE_FEATURE` | 400 BAD_REQUEST | This feature is only available in the enterprise edition | | 120011 | `NOT_AUTHORIZED_TO_ACCESS` | 400 BAD_REQUEST | Project not accessible to user | | 120012 | `FEATURE_FLAG_NOT_ENABLED` | 400 BAD_REQUEST | Platform feature not enabled | | 120013 | `REQUEST_TOO_LARGE` | 413 REQUEST_ENTITY_TOO_LARGE | The request body is larger than this endpoint accepts | | 120014 | `LENGTH_REQUIRED` | 411 LENGTH_REQUIRED | This endpoint requires a Content-Length | ## JobErrorCode | Code | Name | HTTP status | Message | | --- | --- | --- | --- | | 130000 | `JOB_START_FAILED` | 400 BAD_REQUEST | An error occurred while trying to start this job. Check the job logs for details | | 130001 | `JOB_STOP_FAILED` | 400 BAD_REQUEST | An error occurred while trying to stop this job. | | 130002 | `JOB_TYPE_UNSUPPORTED` | 400 BAD_REQUEST | Unsupported job type. | | 130003 | `JOB_ACTION_UNSUPPORTED` | 400 BAD_REQUEST | Unsupported action type. | | 130005 | `JOB_NAME_EMPTY` | 400 BAD_REQUEST | Job name is not set. | | 130006 | `JOB_NAME_INVALID` | 400 BAD_REQUEST | Job name is invalid. Invalid charater(s) in job name | | 130007 | `JOB_EXECUTION_NOT_FOUND` | 404 NOT_FOUND | Execution not found. | | 130008 | `JOB_EXECUTION_TRACKING_URL_NOT_FOUND` | 400 BAD_REQUEST | Tracking url not found. | | 130009 | `JOB_NOT_FOUND` | 404 NOT_FOUND | Job not found. | | 130010 | `JOB_EXECUTION_INVALID_STATE` | 400 BAD_REQUEST | Execution state is invalid. | | 130011 | `JOB_LOG` | 400 BAD_REQUEST | Job log error. | | 130012 | `JOB_DELETION_ERROR` | 400 BAD_REQUEST | Error while deleting job. | | 130013 | `JOB_CREATION_ERROR` | 400 BAD_REQUEST | Error while creating job. | | 130014 | `OPENSEARCH_INDEX_NOT_FOUND` | 400 BAD_REQUEST | OpenSearch indices do not exist | | 130015 | `OPENSEARCH_TYPE_NOT_FOUND` | 400 BAD_REQUEST | OpenSearch type does not exist | | 130016 | `TENSORBOARD_ERROR` | 204 NO_CONTENT | Error getting the TensorBoard(s) for this application | | 130017 | `APPLICATIONID_NOT_FOUND` | 400 BAD_REQUEST | Error while deleting job. | | 130018 | `JOB_ACCESS_ERROR` | 403 FORBIDDEN | Cannot access job | | 130019 | `LOG_AGGREGATION_NOT_ENABLED` | 503 SERVICE_UNAVAILABLE | YARN log aggregation is not enabled | | 130020 | `LOG_RETRIEVAL_ERROR` | 500 INTERNAL_SERVER_ERROR | Error while retrieving YARN logs | | 130021 | `JOB_SCHEDULE_UPDATE` | 500 INTERNAL_SERVER_ERROR | Could not update schedule. | | 130022 | `JAR_INSPECTION_ERROR` | 500 INTERNAL_SERVER_ERROR | Could not inspect jar file. | | 130023 | `PROXY_ERROR` | 500 INTERNAL_SERVER_ERROR | Could not get proxy user. | | 130024 | `JOB_CONFIGURATION_CONVERT_TO_JSON_ERROR` | 400 BAD_REQUEST | Could not convert JobConfiguration to json | | 130025 | `JOB_DELETION_FORBIDDEN` | 403 FORBIDDEN | Your role does not allow to delete this job. | | 130026 | `UNAUTHORIZED_EXECUTION_ACCESS` | 403 FORBIDDEN | This execution does not belong to a job of this project. | | 130027 | `APPID_NOT_FOUND` | 404 NOT_FOUND | AppId not found. | | 130028 | `JOB_PROGRAM_VERSIONING_FAILED` | 500 INTERNAL_SERVER_ERROR | Failed to version application program | | 130029 | `INSUFFICIENT_EXECUTOR_MEMORY` | 400 BAD_REQUEST | Insufficient executor memory provided. | | 130030 | `NODEMANAGERS_OFFLINE` | 503 SERVICE_UNAVAILABLE | Nodemanagers are offline | | 130031 | `DOCKER_MOUNT_NOT_ALLOWED` | 400 BAD_REQUEST | It is not allowed to mount volumes. | | 130032 | `DOCKER_MOUNT_DIR_NOT_ALLOWED` | 400 BAD_REQUEST | It is not allowed to mount this directory. | | 130033 | `DOCKER_UID_GID_STRICT` | 400 BAD_REQUEST | Docker jobs run in uid/gid strict mode. It it now allowed to set uid/gid. If you remove the uid/gid, the job will run with a default user. Please ask an administrator to update the setting if necessary. | | 130034 | `JOB_ALERT_NOT_FOUND` | 404 NOT_FOUND | Job alert not found | | 130035 | `JOB_ALERT_ILLEGAL_ARGUMENT` | 400 BAD_REQUEST | Job alert missing argument. | | 130036 | `JOB_ALERT_ALREADY_EXISTS` | 400 BAD_REQUEST | Job alert with the same status already exists. | | 130037 | `DOCKER_INVALID_JOB_PROPERTIES` | 400 BAD_REQUEST | Received invalid job property values | | 130038 | `FAILED_TO_CREATE_ROUTE` | 400 BAD_REQUEST | Failed to create route. | | 130039 | `FAILED_TO_DELETE_ROUTE` | 400 BAD_REQUEST | Failed to delete route. | | 130040 | `EXECUTIONS_LIMIT_REACHED` | 400 BAD_REQUEST | Job reached the maximum number of executions. | | 130041 | `JOB_ALREADY_EXISTS` | 400 BAD_REQUEST | Job with this name already exists. | | 130042 | `JOB_SCHEDULE_NOT_FOUND` | 404 NOT_FOUND | Cannot find the job schedule. | | 130043 | `UNMATCHED_JOB_NAME` | 400 BAD_REQUEST | Provided job names do not match. | | 130044 | `UNMATCHED_JOB_SCHEDULE_AND_JOB_NAME` | 400 BAD_REQUEST | Requested job schedule id does not match the job name. | | 130045 | `INVALID_RAY_JOB_ENVIRONMENT_YAML_FILE` | 400 BAD_REQUEST | Invalid Ray job environment yaml file. | | 130046 | `JOB_DEPENDENCY_INSPECTION_ERROR` | 500 INTERNAL_SERVER_ERROR | Could not inspect job dependencies. | | 130047 | `JOB_DEPENDENCY_NOT_FOUND` | 400 BAD_REQUEST | Job dependency not found. | | 130048 | `DLT_JOB_FEATURE_GROUP_NOT_FOUND` | 400 BAD_REQUEST | Feature group not found. | | 130049 | `DLT_JOB_FEATURE_GROUP_NO_DATASOURCE` | 400 BAD_REQUEST | The feature group does not have a datasource. | | 130050 | `DLT_JOB_UNSUPPORTED_CONNECTOR` | 400 BAD_REQUEST | The feature group connector does not support DLT sink functionality. | | 130051 | `DLT_JOB_CONFIGS_CREATION_ERROR` | 500 INTERNAL_SERVER_ERROR | Failed to create config for job. | | 130052 | `DLT_JOB_FEATURESTORE_CONNECTOR_NOT_FOUND` | 400 BAD_REQUEST | Featurestore connector not found. | | 130053 | `DLT_JOB_STATE_DIR_CREATION_FAILED` | 500 INTERNAL_SERVER_ERROR | Failed to create state directory for DLT job. | | 130054 | `DLT_JOB_FEATURE_GROUP_NO_STORAGE_CONNECTOR` | 400 BAD_REQUEST | The feature group does not have a storage connector. | | 130055 | `DLT_JOB_FEATURE_STORE_NOT_FOUND` | 400 BAD_REQUEST | Feature store not found. | | 130056 | `DLT_JOB_FEATURE_GROUP_NO_DATA_SOURCE` | 400 BAD_REQUEST | The feature group does not have a data source. | | 130057 | `DLT_JOB_FEATURE_GROUP_ID_MISMATCH` | 400 BAD_REQUEST | Ingestion feature group cannot be changed | | 130058 | `DLT_JOB_FEATURE_STORE_ID_MISMATCH` | 400 BAD_REQUEST | Ingestion feature group cannot be changed | | 130059 | `DLT_JOB_FEATURE_STORE_PROJECT_MISMATCH` | 400 BAD_REQUEST | Ingestion job can only be created in the same project as the feature store. | | 130060 | `DLT_JOB_ALREADY_RUNNING` | 400 BAD_REQUEST | A DLT job for this feature group is already running. | | 130061 | `DLT_JOB_CLEAR_FEATUREGROUP_FAILED` | 500 INTERNAL_SERVER_ERROR | Failed to clear feature group data before DLT ingestion. | | 130062 | `PYTHON_APP_ALREADY_RUNNING` | 400 BAD_REQUEST | A Python App is already running for this job. Stop it before starting a new one. | | 130063 | `RESERVED_ENV_VAR_NAME` | 400 BAD_REQUEST | One or more environment variable names are reserved by the Hopsworks platform. | | 130064 | `AGENT_INVALID_CONFIGURATION` | 400 BAD_REQUEST | Invalid agent job configuration. | | 130065 | `MAX_QUEUED_EXECUTIONS_REACHED` | 429 TOO_MANY_REQUESTS | Max queued executions reached for this job. Wait for some to start before submitting more. | | 130066 | `PYTHON_APP_READINESS_PROBE_INVALID` | 400 BAD_REQUEST | Readiness probe path must be a safe absolute path. | | 130067 | `PYTHON_APP_BASE_PATH_INVALID` | 400 BAD_REQUEST | App base path must be a safe absolute path. | | 130068 | `DLT_JOB_MIXED_CONNECTORS` | 400 BAD_REQUEST | All tables in an ingestion job must use the same data source connector. | | 130069 | `DLT_JOB_DUPLICATE_FEATURE_GROUP` | 400 BAD_REQUEST | The same feature group cannot appear in multiple targets of an ingestion job. | | 130070 | `DLT_JOB_INVALID_TABLE_PARALLELISM` | 400 BAD_REQUEST | tableParallelism must be between 1 and the number of ingestion targets. | | 130071 | `DLT_JOB_NO_TARGETS` | 400 BAD_REQUEST | An ingestion job must have at least one table target. | | 130072 | `DLT_JOB_NO_ENABLED_TARGETS` | 400 BAD_REQUEST | At least one table must be enabled for ingestion. | | 130073 | `DLT_JOB_TABLE_NOT_RUNNING` | 400 BAD_REQUEST | No running table with the given index for this execution. | | 130074 | `GIT_COMMIT_NOT_VALID` | 400 BAD_REQUEST | Git commit hash is not valid | | 130075 | `GIT_SYNC_NOT_ENABLED` | 400 BAD_REQUEST | Git auto-redeploy is not enabled for this app | | 130076 | `GIT_AUTO_REDEPLOY_NOT_SUPPORTED` | 400 BAD_REQUEST | Git auto-redeploy is only supported for apps deployed from a git repository | | 130077 | `JOB_DESCRIPTION_TOO_LONG` | 400 BAD_REQUEST | Job description is too long. | | 130078 | `INVALID_MEMORY_OVERHEAD_FACTOR` | 400 BAD_REQUEST | Memory overhead factor must be greater than zero. | ## RequestErrorCode | Code | Name | HTTP status | Message | | --- | --- | --- | --- | | 140000 | `MESSAGE_ACCESS_NOT_ALLOWED` | 403 FORBIDDEN | Message not allowed. | | 140001 | `EMAIL_EMPTY` | 400 BAD_REQUEST | Email cannot be empty. | | 140002 | `EMAIL_INVALID` | 400 BAD_REQUEST | Not a valid email address. | | 140003 | `DATASET_REQUEST_ERROR` | 400 BAD_REQUEST | Error while submitting dataset request | | 140004 | `REQUEST_UNKNOWN_ACTION` | 400 BAD_REQUEST | Unknown request action | | 140005 | `MESSAGE_NOT_FOUND` | 404 NOT_FOUND | Message was not found | ## ProjectErrorCode | Code | Name | HTTP status | Message | | --- | --- | --- | --- | | 150001 | `PROJECT_EXISTS` | 409 CONFLICT | Project with the same name already exists. | | 150002 | `NUM_PROJECTS_LIMIT_REACHED` | 400 BAD_REQUEST | You have reached the maximum number of projects you could create. Contact an administrator to increase your limit or delete some of the existing projects. | | 150003 | `INVALID_PROJECT_NAME` | 400 BAD_REQUEST | Invalid project name, valid characters: \[a-zA-Z0-9\]((?!__)\[_a-zA-Z0-9\]){0,62} | | 150004 | `PROJECT_NOT_FOUND` | 400 BAD_REQUEST | Project wasn't found. | | 150005 | `PROJECT_NOT_REMOVED` | 400 BAD_REQUEST | Project wasn't removed. | | 150007 | `PROJECT_FOLDER_NOT_CREATED` | 500 INTERNAL_SERVER_ERROR | Project folder could not be created in HDFS. | | 150008 | `STARTER_PROJECT_BAD_REQUEST` | 400 BAD_REQUEST | Type of starter project is not valid | | 150009 | `PROJECT_FOLDER_NOT_REMOVED` | 400 BAD_REQUEST | Project folder could not be removed from HDFS. | | 150010 | `PROJECT_REMOVAL_NOT_ALLOWED` | 403 FORBIDDEN | Project can only be deleted by its owner. | | 150011 | `PROJECT_MEMBER_NOT_REMOVED` | 500 INTERNAL_SERVER_ERROR | Failed to remove team member. | | 150012 | `MEMBER_REMOVAL_NOT_ALLOWED` | 403 FORBIDDEN | Your project role does not allow to remove other members from this project. | | 150013 | `PROJECT_OWNER_NOT_ALLOWED` | 403 FORBIDDEN | Removing the project owner is not allowed. | | 150014 | `PROJECT_OWNER_ROLE_NOT_ALLOWED` | 403 FORBIDDEN | Changing the role of the project owner is not allowed. | | 150015 | `FOLDER_INODE_NOT_CREATED` | 400 BAD_REQUEST | Folder Inode could not be created in DB. | | 150016 | `FOLDER_NAME_NOT_SET` | 400 BAD_REQUEST | Name cannot be empty. | | 150017 | `FOLDER_NAME_TOO_LONG` | 400 BAD_REQUEST | Name cannot be longer than 88 characters. | | 150018 | `FOLDER_NAME_CONTAIN_DISALLOWED_CHARS` | 400 BAD_REQUEST | Name cannot contain any of the characters | | 150019 | `FILE_NAME_EXIST` | 400 BAD_REQUEST | File with the same name already exists. | | 150020 | `FILE_NOT_FOUND` | 400 BAD_REQUEST | File not found. | | 150021 | `NO_MEMBER_TO_ADD` | 400 BAD_REQUEST | No member to add. | | 150022 | `NO_MEMBER_ADD` | 400 BAD_REQUEST | No member added. | | 150023 | `TEAM_MEMBER_NOT_FOUND` | 404 NOT_FOUND | The selected user is not a team member in this project. | | 150024 | `TEAM_MEMBER_ALREADY_EXISTS` | 400 BAD_REQUEST | The selected user is already a team member of this project. | | 150025 | `ROLE_NOT_SET` | 400 BAD_REQUEST | Role cannot be empty. | | 150026 | `PROJECT_NOT_SELECTED` | 400 BAD_REQUEST | No project selected | | 150027 | `QUOTA_NOT_FOUND` | 400 BAD_REQUEST | Quota information not found. | | 150028 | `QUOTA_ERROR` | 400 BAD_REQUEST | Quota create or update error. | | 150029 | `PROJECT_QUOTA_ERROR` | 412 PRECONDITION_FAILED | This project is out of credits. | | 150030 | `PROJECT_CREATED` | 201 CREATED | Project created successfully. | | 150031 | `PROJECT_DESCRIPTION_CHANGED` | 200 OK | Project description changed. | | 150033 | `PROJECT_SERVICE_ADDED` | 200 OK | Project service added | | 150034 | `PROJECT_SERVICE_ADD_FAILURE` | 500 INTERNAL_SERVER_ERROR | Failure adding service | | 150035 | `PROJECT_REMOVED` | 200 OK | The project and all related files were removed successfully. | | 150036 | `PROJECT_REMOVED_NOT_FOLDER` | 500 INTERNAL_SERVER_ERROR | The project was removed successfully. But its datasets have not been deleted. | | 150037 | `PROJECT_MEMBER_REMOVED` | 200 OK | Member removed successfully | | 150038 | `PROJECT_MEMBERS_ADDED` | 200 OK | Members added successfully | | 150039 | `PROJECT_MEMBER_ADDED` | 200 OK | One member added successfully | | 150040 | `MEMBER_ROLE_UPDATED` | 200 OK | Role updated successfully. | | 150041 | `MEMBER_REMOVED_FROM_TEAM` | 200 OK | Member removed from team. | | 150042 | `PROJECT_INODE_CREATION_ERROR` | 500 INTERNAL_SERVER_ERROR | Could not create dummy Inode | | 150043 | `PROJECT_FOLDER_EXISTS` | 409 CONFLICT | A folder with same name as the project already exists in the system. | | 150044 | `PROJECT_USER_EXISTS` | 409 CONFLICT | Filesystem user(s) already exists in the system. | | 150045 | `PROJECT_GROUP_EXISTS` | 409 CONFLICT | Filesystem group(s) already exists in the system. | | 150046 | `PROJECT_CERTIFICATES_EXISTS` | 409 CONFLICT | Certificates for this project already exist in the system. | | 150047 | `PROJECT_QUOTA_EXISTS` | 409 CONFLICT | Quotas corresponding to this project already exist in the system. | | 150048 | `PROJECT_LOGS_EXIST` | 409 CONFLICT | Logs corresponding to this project already exist in the system. | | 150049 | `PROJECT_VERIFICATIONS_FAILED` | 500 INTERNAL_SERVER_ERROR | Error occurred while running verifications | | 150050 | `PROJECT_SET_PERMISSIONS_ERROR` | 500 INTERNAL_SERVER_ERROR | Error occurred while setting permissions for project folders. | | 150051 | `PROJECT_HANDLER_PRECREATE_ERROR` | 500 INTERNAL_SERVER_ERROR | Error occurred during project precreate handler. | | 150052 | `PROJECT_HANDLER_POSTCREATE_ERROR` | 500 INTERNAL_SERVER_ERROR | Error occurred during project postcreate handler. | | 150053 | `PROJECT_HANDLER_PREDELETE_ERROR` | 500 INTERNAL_SERVER_ERROR | Error occurred during project predelete handler. | | 150054 | `PROJECT_HANDLER_POSTDELETE_ERROR` | 500 INTERNAL_SERVER_ERROR | Error occurred during project postdelete handler. | | 150055 | `PROJECT_TOUR_FILES_ERROR` | 500 INTERNAL_SERVER_ERROR | Error while adding tour files to project. | | 150056 | `PROJECT_KIBANA_CREATE_INDEX_ERROR` | 500 INTERNAL_SERVER_ERROR | Could not create kibana index-pattern for project | | 150057 | `PROJECT_KIBANA_CREATE_SEARCH_ERROR` | 500 INTERNAL_SERVER_ERROR | Could not create kibana search for project | | 150058 | `PROJECT_KIBANA_CREATE_DASHBOARD_ERROR` | 500 INTERNAL_SERVER_ERROR | Could not create kibana dashboard for project | | 150060 | `PROJECT_CONDA_LIBS_NOT_FOUND` | 404 NOT_FOUND | No preinstalled anaconda libs found. | | 150061 | `KILL_MEMBER_JOBS` | 500 INTERNAL_SERVER_ERROR | Could not kill user's yarn applications | | 150062 | `JUPYTER_SERVER_NOT_FOUND` | 404 NOT_FOUND | Could not find Jupyter entry for user in this project. | | 150063 | `PYTHON_LIB_ALREADY_INSTALLED` | 304 NOT_MODIFIED | This python library is already installed on this project | | 150064 | `PYTHON_LIB_NOT_INSTALLED` | 304 NOT_MODIFIED | This python library is not installed for this project. Cannot remove/upgrade op | | 150066 | `ANACONDA_NOT_ENABLED` | 412 PRECONDITION_FAILED | First enable Anaconda. Click on 'Python' -> Activate Anaconda | | 150067 | `TENSORBOARD_OPENSEARCH_INDEX_NOT_FOUND` | 404 NOT_FOUND | Could not find OpenSearch index for TensorBoard. | | 150068 | `PROJECT_ROLE_FORBIDDEN` | 403 FORBIDDEN | Your project role does not allow to perform this action. | | 150069 | `FOLDER_NAME_ENDS_WITH_DOT` | 400 BAD_REQUEST | Name cannot end in a period. | | 150070 | `FOLDER_NAME_EXISTS` | 400 BAD_REQUEST | A directory with the same name already exists. If you want to replace it delete it first then try recreating. | | 150071 | `PROJECT_SERVICE_NOT_FOUND` | 400 BAD_REQUEST | service was not found. | | 150072 | `QUOTA_REQUEST_NOT_COMPLETE` | 400 BAD_REQUEST | Please specify both namespace and space quota. | | 150073 | `RESERVED_PROJECT_NAME` | 400 BAD_REQUEST | Not allowed - reserved project name, pick another project name. | | 150074 | `PROJECT_ANACONDA_ENABLE_ERROR` | 500 INTERNAL_SERVER_ERROR | Failed to enable conda. | | 150075 | `PROJECT_NAME_TOO_LONG` | 400 BAD_REQUEST | Project name is too long - cannot be longer than 25 characters. | | 150076 | `PROJECT_DOCKER_VERSION_EXTRACT_ERROR` | 500 INTERNAL_SERVER_ERROR | Failed to extract the hopsworks version of the docker image for this project. | | 150077 | `PROJECT_DEFAULT_JOB_CONFIG_NOT_FOUND` | 404 NOT_FOUND | Default job config not found | | 150078 | `ALERT_NOT_FOUND` | 404 NOT_FOUND | Alert not found | | 150079 | `ALERT_ILLEGAL_ARGUMENT` | 400 BAD_REQUEST | Alert missing argument. | | 150080 | `ALERT_ALREADY_EXISTS` | 400 BAD_REQUEST | Alert with the same status already exists. | | 150081 | `FAILED_TO_ADD_MEMBER` | 400 BAD_REQUEST | Failed to add member. | | 150082 | `FAILED_TO_CREATE_ROUTE` | 400 BAD_REQUEST | Failed to create route. | | 150083 | `FAILED_TO_DELETE_ROUTE` | 400 BAD_REQUEST | Failed to delete route. | | 150084 | `PROJECT_TEAM_ROLE_HANDLER_ADD_MEMBER_ERROR` | 500 INTERNAL_SERVER_ERROR | Error occurred during project team role add handler. | | 150085 | `PROJECT_TEAM_ROLE_HANDLER_UPDATE_MEMBERS_ERROR` | 500 INTERNAL_SERVER_ERROR | Error occurred during project team role update handler. | | 150086 | `PROJECT_TEAM_ROLE_HANDLER_REMOVE_MEMBER_ERROR` | 500 INTERNAL_SERVER_ERROR | Error occurred during project team role remove handler. | | 150087 | `PROJECT_TEAM_ROLE_NOT_SUPPORTED` | 400 BAD_REQUEST | Role not supported. | | 150088 | `PROJECT_NAMESPACE_ERROR` | 500 INTERNAL_SERVER_ERROR | Error occurred when using kubernetes namespace in project | | 150090 | `PROJECT_SERVICE_NOT_ALLOWED` | 400 BAD_REQUEST | Project service not allowed. | | 150091 | `PROJECT_MAPPING_ILLEGAL_ARGUMENT` | 400 BAD_REQUEST | Illegal argument. | | 150092 | `PROJECT_MAPPING_NOT_ALLOWED` | 400 BAD_REQUEST | Operation not allowed. | | 150093 | `PROJECT_MAPPING_NOT_FOUND` | 400 BAD_REQUEST | Mapping not found. | | 150094 | `PROJECT_MAPPING_DUPLICATE_ENTRY` | 400 BAD_REQUEST | Duplicate entry. | | 150095 | `MEMBER_MANAGEMENT_NOT_ALLOWED` | 400 BAD_REQUEST | Member management not allowed. | | 150096 | `MCP_SERVER_NOT_FOUND` | 404 NOT_FOUND | MCP server not found. | | 150097 | `MCP_SERVER_VALIDATION` | 400 BAD_REQUEST | MCP server validation error. | | 150098 | `TOO_MANY_MEMBERS_AT_ONCE` | 400 BAD_REQUEST | Too many members in one request. | | 150099 | `LAST_DATA_OWNER_NOT_ALLOWED` | 403 FORBIDDEN | Removing the last data owner of the project is not allowed. | | 150100 | `FILE_OWNER_NOT_DATA_OWNER` | 400 BAD_REQUEST | The member chosen to take over the removed member's files must be a data owner in this project. | | 150101 | `SYSTEM_PROJECT_REMOVAL_NOT_ALLOWED` | 400 BAD_REQUEST | This project is managed by Hopsworks and cannot be deleted. | ## UserErrorCode | Code | Name | HTTP status | Message | | --- | --- | --- | --- | | 160000 | `NO_ROLE_FOUND` | 401 UNAUTHORIZED | No valid role found for this user | | 160001 | `USER_DOES_NOT_EXIST` | 400 BAD_REQUEST | User does not exist. | | 160002 | `USER_WAS_NOT_FOUND` | 404 NOT_FOUND | User not found | | 160003 | `USER_EXISTS` | 409 CONFLICT | There is an existing account associated with this email | | 160004 | `ACCOUNT_REQUEST` | 401 UNAUTHORIZED | Your account has not yet been approved. | | 160005 | `ACCOUNT_DEACTIVATED` | 401 UNAUTHORIZED | This account has been deactivated. | | 160006 | `ACCOUNT_VERIFICATION` | 400 BAD_REQUEST | You need to verify your account. | | 160007 | `ACCOUNT_BLOCKED` | 401 UNAUTHORIZED | Your account has been blocked. Contact the administrator. | | 160008 | `AUTHENTICATION_FAILURE` | 401 UNAUTHORIZED | Authentication failed, invalid credentials | | 160009 | `LOGOUT_FAILURE` | 400 BAD_REQUEST | Logout failed on backend. | | 160014 | `PASSWORD_EMPTY` | 400 BAD_REQUEST | Password cannot be empty. | | 160015 | `PASSWORD_TOO_SHORT` | 400 BAD_REQUEST | Password too short. | | 160016 | `PASSWORD_TOO_LONG` | 400 BAD_REQUEST | Password too long. | | 160017 | `PASSWORD_INCORRECT` | 400 BAD_REQUEST | Password incorrect | | 160018 | `PASSWORD_PATTERN_NOT_CORRECT` | 400 BAD_REQUEST | Password should include one uppercase letter, one special character and/or alphanumeric characters. | | 160019 | `INCORRECT_PASSWORD` | 401 UNAUTHORIZED | The password is incorrect. Please try again | | 160020 | `PASSWORD_MISS_MATCH` | 400 BAD_REQUEST | Passwords do not match - typo? | | 160021 | `TOS_NOT_AGREED` | 400 BAD_REQUEST | You must agree to our terms of use. | | 160022 | `CERT_DOWNLOAD_DENIED` | 400 BAD_REQUEST | Admin is not allowed to download certificates | | 160023 | `CREATED_ACCOUNT` | 400 BAD_REQUEST | You have successfully created an account but you might need to wait until your account has been approved before you can login. | | 160024 | `PASSWORD_RESET_SUCCESSFUL` | 400 BAD_REQUEST | Your password was successfully reset your new password have been sent to your email. | | 160025 | `PASSWORD_RESET_UNSUCCESSFUL` | 400 BAD_REQUEST | Your password could not be reset. Please try again later or contact support. | | 160026 | `PASSWORD_CHANGED` | 400 BAD_REQUEST | Your password was successfully changed. | | 160028 | `PROFILE_UPDATED` | 400 BAD_REQUEST | Your profile was updated successfully. | | 160029 | `SSH_KEY_REMOVED` | 400 BAD_REQUEST | Your ssh key was deleted successfully. | | 160030 | `NOTHING_TO_UPDATE` | 400 BAD_REQUEST | Nothing to update | | 160031 | `CREATE_USER_ERROR` | 500 INTERNAL_SERVER_ERROR | Error while creating user | | 160032 | `CERT_AUTHORIZATION_ERROR` | 401 UNAUTHORIZED | Certificate CN does not match the username provided. | | 160033 | `PROJECT_USER_CERT_NOT_FOUND` | 403 FORBIDDEN | Could not find exactly one certificate for user in project. | | 160034 | `ACCOUNT_INACTIVE` | 401 UNAUTHORIZED | This account has not been activated | | 160035 | `ACCOUNT_LOST_DEVICE` | 401 UNAUTHORIZED | This account has registered a lost device. | | 160036 | `ACCOUNT_NOT_APPROVED` | 401 UNAUTHORIZED | This account has not yet been approved | | 160037 | `INVALID_EMAIL` | 400 BAD_REQUEST | Invalid email format. | | 160038 | `INCORRECT_DEACTIVATION_LENGTH` | 400 BAD_REQUEST | The message should have a length between 5 and 500 characters | | 160039 | `TMP_CODE_INVALID` | 401 UNAUTHORIZED | The temporary code was wrong. | | 160040 | `INCORRECT_CREDENTIALS` | 400 BAD_REQUEST | Incorrect email or password. | | 160041 | `INCORRECT_VALIDATION_KEY` | 400 BAD_REQUEST | Incorrect validation key | | 160042 | `ACCOUNT_ALREADY_VERIFIED` | 409 CONFLICT | User is already verified | | 160043 | `TWO_FA_ENABLE_ERROR` | 500 INTERNAL_SERVER_ERROR | Cannot enable 2-factor authentication. | | 160044 | `ACCOUNT_REGISTRATION_ERROR` | 500 INTERNAL_SERVER_ERROR | Account registration error. | | 160045 | `TWO_FA_DISABLED` | 412 PRECONDITION_FAILED | 2-factor authentication is disabled. | | 160046 | `TRANSITION_STATUS_ERROR` | 400 BAD_REQUEST | The user can't transition from current status to requested status | | 160047 | `ACCESS_CONTROL` | 403 FORBIDDEN | Client not authorized for this invocation. | | 160048 | `SECRET_EMPTY` | 404 NOT_FOUND | Secret is empty | | 160049 | `SECRET_EXISTS` | 409 CONFLICT | Same Secret already exists | | 160050 | `SECRET_ENCRYPTION_ERROR` | 500 INTERNAL_SERVER_ERROR | Error encrypting/decrypting Secret | | 160051 | `ACCOUNT_NOT_ACTIVE` | 400 BAD_REQUEST | This account is not active | | 160052 | `ACCOUNT_ACTIVATION_FAILED` | 400 BAD_REQUEST | Account activation failed | | 160053 | `ROLE_NOT_FOUND` | 400 BAD_REQUEST | Role not found | | 160054 | `ACCOUNT_DELETION_ERROR` | 400 BAD_REQUEST | Failed to delete account. | | 160055 | `USER_NAME_NOT_SET` | 400 BAD_REQUEST | User name not set. | | 160056 | `SECRET_DELETION_FAILED` | 400 BAD_REQUEST | Failed to delete secret. | | 160057 | `USER_SEARCH_NOT_ALLOWED` | 400 BAD_REQUEST | Search not allowed. | | 160058 | `FAILED_TO_GENERATE_QR_CODE` | 417 EXPECTATION_FAILED | Failed to generate QR code. | | 160059 | `INVALID_OTP` | 400 BAD_REQUEST | Invalid OTP. | | 160060 | `USER_ACCOUNT_HANDLER_CREATE_ERROR` | 500 INTERNAL_SERVER_ERROR | Error occurred during user account create handler. | | 160061 | `USER_ACCOUNT_HANDLER_UPDATE_ERROR` | 500 INTERNAL_SERVER_ERROR | Error occurred during user account update handler. | | 160062 | `USER_ACCOUNT_HANDLER_REMOVE_ERROR` | 500 INTERNAL_SERVER_ERROR | Error occurred during user account remove handler. | | 160063 | `OPERATION_NOT_ALLOWED` | 400 BAD_REQUEST | Operation not allowed on user | | 160064 | `ACCOUNT_REJECTION_FAILED` | 400 BAD_REQUEST | Account rejection failed | | 160065 | `SECRET_CREATION_FAILED` | 500 INTERNAL_SERVER_ERROR | Secret creation failed | | 160066 | `ENV_VAR_INVALID_NAME` | 400 BAD_REQUEST | Environment variable name is invalid. | | 160067 | `ENV_VAR_RESERVED_NAME` | 400 BAD_REQUEST | Environment variable name is reserved. | | 160068 | `ENV_VAR_VALUE_TOO_LARGE` | 400 BAD_REQUEST | Environment variable value is too large. | | 160069 | `ENV_VAR_LIMIT_EXCEEDED` | 400 BAD_REQUEST | Environment variable limit exceeded. | | 160070 | `ENV_VAR_NOT_FOUND` | 404 NOT_FOUND | Environment variable was not found. | | 160071 | `ENV_VAR_ENCRYPTION_ERROR` | 500 INTERNAL_SERVER_ERROR | Error encrypting/decrypting environment variable. | | 160072 | `SECRET_VALUE_TOO_LARGE` | 400 BAD_REQUEST | Secret value is too large. | | 160073 | `ENV_VAR_INVALID_VALUE` | 400 BAD_REQUEST | Environment variable value is invalid. | | 160074 | `ACCOUNT_DELETION_PENDING_CLEANUP` | 409 CONFLICT | Account still has records awaiting background cleanup. Retry shortly. | ## MetadataErrorCode | Code | Name | HTTP status | Message | | --- | --- | --- | --- | | 180000 | `TEMPLATE_ALREADY_AVAILABLE` | 400 BAD_REQUEST | The template is already available | | 180001 | `TEMPLATE_INODEID_EMPTY` | 400 BAD_REQUEST | The template id is empty | | 180002 | `TEMPLATE_NOT_ATTACHED` | 400 BAD_REQUEST | The template could not be attached to a file | | 180003 | `DATASET_TEMPLATE_INFO_MISSING` | 400 BAD_REQUEST | Template info is missing. Please provide InodeDTO path and templateId. | | 180004 | `NO_METADATA_EXISTS` | 400 BAD_REQUEST | No metadata found | | 180005 | `METADATA_MAX_SIZE_EXCEEDED` | 400 BAD_REQUEST | Metadata is too large | | 180006 | `METADATA_MISSING_FIELD` | 400 BAD_REQUEST | Metadata missing attributed name. | | 180007 | `METADATA_ERROR` | 500 INTERNAL_SERVER_ERROR | Error while processing the extended metadata. | | 180008 | `METADATA_ILLEGAL_NAME` | 400 BAD_REQUEST | Metadata name is illegal. | ## KafkaErrorCode | Code | Name | HTTP status | Message | | --- | --- | --- | --- | | 190000 | `TOPIC_NOT_FOUND` | 404 NOT_FOUND | No topics found | | 190001 | `BROKER_METADATA_ERROR` | 500 INTERNAL_SERVER_ERROR | An error occurred while retrieving topic metadata from broker | | 190002 | `TOPIC_ALREADY_EXISTS` | 409 CONFLICT | Kafka topic already exists in database. Pick a different topic name | | 190003 | `TOPIC_ALREADY_EXISTS_IN_ZOOKEEPER` | 409 CONFLICT | Kafka topic already exists in ZooKeeper. Pick a different topic name | | 190004 | `TOPIC_LIMIT_REACHED` | 412 PRECONDITION_FAILED | Topic limit reached. Contact your administrator to increase the number of topics that can be created for this project. | | 190005 | `TOPIC_REPLICATION_ERROR` | 400 BAD_REQUEST | Maximum topic replication factor exceeded | | 190006 | `SCHEMA_NOT_FOUND` | 404 NOT_FOUND | Topic has no schema attached to it. | | 190007 | `KAFKA_GENERIC_ERROR` | 500 INTERNAL_SERVER_ERROR | An error occurred while retrieving information about Kafka | | 190008 | `DESTINATION_PROJECT_IS_TOPIC_OWNER` | 400 BAD_REQUEST | Destination projet is topic owner | | 190009 | `TOPIC_ALREADY_SHARED` | 400 BAD_REQUEST | Topic is already shared | | 190010 | `TOPIC_NOT_SHARED` | 404 NOT_FOUND | Topic is not shared with project | | 190011 | `ACL_ALREADY_EXISTS` | 409 CONFLICT | ACL already exists. | | 190012 | `ACL_NOT_FOUND` | 404 NOT_FOUND | ACL not found. | | 190013 | `ACL_NOT_FOR_TOPIC` | 400 BAD_REQUEST | ACL does not belong to the specified topic | | 190014 | `SCHEMA_IN_USE` | 412 PRECONDITION_FAILED | Schema is currently used by topics. topic | | 190015 | `BAD_NUM_PARTITION` | 400 BAD_REQUEST | Invalid number of partitions | | 190016 | `CREATE_SUBJECT_RESERVED_NAME` | 405 METHOD_NOT_ALLOWED | The provided subject name is reserved for system calls | | 190017 | `DELETE_RESERVED_SCHEMA` | 405 METHOD_NOT_ALLOWED | The schema is reserved and cannot be deleted | | 190018 | `SCHEMA_VERSION_NOT_FOUND` | 404 NOT_FOUND | Specified version of the schema not found | | 190019 | `PROJECT_IS_NOT_THE_OWNER_OF_THE_TOPIC` | 400 BAD_REQUEST | Specified project is not the owner of the topic | | 190020 | `ACL_FOR_ANY_USER` | 400 BAD_REQUEST | Cannot create an ACL for user with email '*' | | 190021 | `KAFKA_UNAVAILABLE` | 503 SERVICE_UNAVAILABLE | Kafka is temporarily unavailable. Please try again later | | 190022 | `TOPIC_DELETION_FAILED` | 500 INTERNAL_SERVER_ERROR | Could not delete Kafka topics. | | 190023 | `TOPIC_FETCH_FAILED` | 500 INTERNAL_SERVER_ERROR | Could not fetch topic details. | | 190024 | `TOPIC_CREATION_FAILED` | 500 INTERNAL_SERVER_ERROR | Could not create topic. | | 190025 | `BROKER_MISSING` | 404 NOT_FOUND | Could not find a broker endpoint. | ## SecurityErrorCode | Code | Name | HTTP status | Message | | --- | --- | --- | --- | | 200001 | `MASTER_ENCRYPTION_PASSWORD_CHANGE` | 400 BAD_REQUEST | Master password change procedure started. Check your inbox for final status | | 200002 | `HDFS_ACCESS_CONTROL` | 403 FORBIDDEN | Access error while trying to access hdfs resource | | 200003 | `EJB_ACCESS_LOCAL` | 401 UNAUTHORIZED | Unauthorized invocation | | 200004 | `CERT_CREATION_ERROR` | 500 INTERNAL_SERVER_ERROR | Error while generating certificates. | | 200005 | `CERT_CN_EXTRACT_ERROR` | 500 INTERNAL_SERVER_ERROR | Error while extracting CN from certificate. | | 200006 | `CERT_ERROR` | 401 UNAUTHORIZED | Certificate could not be validated. | | 200007 | `CERT_ACCESS_DENIED` | 403 FORBIDDEN | Certificate access denied. | | 200008 | `CSR_ERROR` | 500 INTERNAL_SERVER_ERROR | Error while signing CSR. | | 200009 | `CERT_APP_REVOKE_ERROR` | 500 INTERNAL_SERVER_ERROR | Error while revoking application certificate, check the logs | | 200010 | `CERT_MATERIALIZATION_ERROR` | 500 INTERNAL_SERVER_ERROR | CertificateMaterializer error, could not materialize certificates | | 200011 | `MASTER_ENCRYPTION_PASSWORD_ACCESS_ERROR` | 500 INTERNAL_SERVER_ERROR | Could not read master encryption password. | | 200012 | `NOT_RENEWABLE_TOKEN` | 400 BAD_REQUEST | Token can not be renewed. | | 200013 | `INVALIDATION_ERROR` | 417 EXPECTATION_FAILED | Error while invalidating token. | | 200014 | `REST_ACCESS_CONTROL` | 403 FORBIDDEN | Client not authorized for this invocation. | | 200015 | `DUPLICATE_KEY_ERROR` | 409 CONFLICT | A signing key with the same name already exists. | | 200016 | `CERTIFICATE_REVOKATION_ERROR` | 500 INTERNAL_SERVER_ERROR | Error revoking the certificate | | 200017 | `CERTIFICATE_NOT_FOUND` | 400 BAD_REQUEST | Could not find the certificate | | 200018 | `CERTIFICATE_REVOKATION_USER_ERR` | 400 BAD_REQUEST | Error revoking the certificate | | 200019 | `CERTIFICATE_SIGN_USER_ERR` | 400 BAD_REQUEST | Error signing the certificate | | 200020 | `MASTER_ENCRYPTION_PASSWORD_RESET_ERROR` | 500 INTERNAL_SERVER_ERROR | Error resetting master encryption password. | ## CAErrorCode | Code | Name | HTTP status | Message | | --- | --- | --- | --- | | 220000 | `BADSIGNREQUEST` | 400 BAD_REQUEST | No CSR provided or CSR is malformed | | 220001 | `BADREVOKATIONREQUEST` | 400 BAD_REQUEST | No certificate identifier provided | | 220002 | `CERTNOTFOUND` | 204 NO_CONTENT | Certificate not found | | 220003 | `CERTEXISTS` | 400 BAD_REQUEST | Certificate with the same identifier already exists | | 220004 | `BAD_SUBJECT_NAME` | 400 BAD_REQUEST | Invalid certificate subject name | | 220005 | `CERTIFICATE_DECODING_ERROR` | 500 INTERNAL_SERVER_ERROR | Could not decode certificate | | 220006 | `CERTIFICATE_REVOCATION_FAILURE` | 500 INTERNAL_SERVER_ERROR | Failed to revoke certificate | | 220007 | `CERTIFICATE_REVOCATION_LIST_READ` | 500 INTERNAL_SERVER_ERROR | Failed to read Certificate Revocation List | | 220008 | `CSR_GENERIC_ERROR` | 500 INTERNAL_SERVER_ERROR | Error handling certificate signing request | | 220009 | `CSR_SIGNING_ERROR` | 500 INTERNAL_SERVER_ERROR | Could not sign Certificate Signing Request | | 220010 | `CA_INITIALIZATION_ERROR` | 500 INTERNAL_SERVER_ERROR | Error while initializing Certificate Authorities | | 220011 | `PKI_GENERIC_ERROR` | 500 INTERNAL_SERVER_ERROR | Generic PKI error | ## DelaCSRErrorCode | Code | Name | HTTP status | Message | | --- | --- | --- | --- | | 230000 | `BADREQUEST` | 400 BAD_REQUEST | User or CS not set | | 230001 | `EMAIL` | 401 UNAUTHORIZED | CSR email not set or does not match user | | 230003 | `CN` | 400 BAD_REQUEST | CSR common name not set | | 230004 | `O` | 400 BAD_REQUEST | CSR organization name not set | | 230005 | `OU` | 400 BAD_REQUEST | CSR organization unit name not set | | 230006 | `NOTFOUND` | 400 BAD_REQUEST | No cluster registered with the given organization name and organizational unit | | 230007 | `SERIALNUMBER` | 400 BAD_REQUEST | Cluster has already a signed certificate | | 230008 | `CNNOTFOUND` | 400 BAD_REQUEST | No cluster registered with the CSR common name | | 230009 | `AGENTIDNOTFOUND` | 401 UNAUTHORIZED | No cluster registered for the user | ## ServingErrorCode | Code | Name | HTTP status | Message | | --- | --- | --- | --- | | 240000 | `INSTANCE_NOT_FOUND` | 404 NOT_FOUND | Serving instance not found | | 240001 | `DELETION_ERROR` | 500 INTERNAL_SERVER_ERROR | Serving instance could not be deleted | | 240002 | `UPDATE_ERROR` | 500 INTERNAL_SERVER_ERROR | Serving instance could not be updated | | 240003 | `LIFECYCLE_ERROR` | 400 BAD_REQUEST | Serving instance could not be started/stopped | | 240004 | `LIFECYCLE_ERROR_INT` | 500 INTERNAL_SERVER_ERROR | Serving instance could not be started/stopped | | 240005 | `STATUS_ERROR` | 500 INTERNAL_SERVER_ERROR | Error getting model server instance status | | 240006 | `MODEL_PATH_NOT_FOUND` | 400 BAD_REQUEST | Model path not found | | 240007 | `COMMAND_NOT_RECOGNIZED` | 400 BAD_REQUEST | Command not recognized | | 240008 | `COMMAND_NOT_PROVIDED` | 400 BAD_REQUEST | Command not provided | | 240009 | `SPEC_NOT_PROVIDED` | 400 BAD_REQUEST | TFServing spec not provided | | 240010 | `BAD_TOPIC` | 400 BAD_REQUEST | Topic provided cannot be used for Serving logging | | 240011 | `DUPLICATED_ENTRY` | 400 BAD_REQUEST | An entry with the same name already exists in this project | | 240012 | `PYTHON_ENVIRONMENT_NOT_ENABLED` | 400 BAD_REQUEST | Python environment has not been enabled in this project, which is required for serving SkLearn Models | | 240013 | `UPDATE_MODEL_SERVER_ERROR` | 400 BAD_REQUEST | The model server of a deployment cannot be updated. | | 240014 | `KUBERNETES_NOT_INSTALLED` | 400 BAD_REQUEST | Kubernetes is not installed | | 240015 | `KSERVE_NOT_ENABLED` | 400 BAD_REQUEST | KServe is not installed or disabled | | 240016 | `SCRIPT_NOT_FOUND` | 400 BAD_REQUEST | Script not found | | 240017 | `MODEL_FILES_STRUCTURE_NOT_VALID` | 400 BAD_REQUEST | Model path does not have a valid file structure | | 240018 | `MODEL_ARTIFACT_NOT_VALID` | 400 BAD_REQUEST | Model artifact not valid | | 240019 | `MODEL_ARTIFACT_OPERATION_ERROR` | 400 BAD_REQUEST | Model artifact cannot be created or changed | | 240020 | `PREDICTOR_NOT_SUPPORTED` | 400 BAD_REQUEST | Predictors not supported | | 240021 | `TRANSFORMER_NOT_SUPPORTED` | 400 BAD_REQUEST | Transformers not supported | | 240022 | `KAFKA_TOPIC_NOT_FOUND` | 400 BAD_REQUEST | Kafka topic not found | | 240023 | `KAFKA_TOPIC_NOT_VALID` | 400 BAD_REQUEST | Kafka topic not valid | | 240024 | `FINEGRAINED_INF_LOGGING_NOT_SUPPORTED` | 400 BAD_REQUEST | Fine-grained inference logging not supported | | 240025 | `REQUEST_BATCHING_NOT_SUPPORTED` | 400 BAD_REQUEST | Request batching not supported | | 240026 | `CREATE_ERROR` | 400 BAD_REQUEST | Serving instance could not be created | | 240027 | `SERVER_LOGS_NOT_AVAILABLE` | 404 NOT_FOUND | Server logs not available | | 240028 | `API_PROTOCOL_NOT_SUPPORTED` | 400 BAD_REQUEST | GRPC only supported in KServe deployments | | 240029 | `SCHEDULING_CONFIG_ERROR` | 400 BAD_REQUEST | Scheduling configuration error | | 240030 | `UNSUPPORTED_MODELLESS_SERVING_TYPE` | 400 BAD_REQUEST | Modelless serving type not supported | | 240031 | `RESERVED_ENV_VAR_NAME` | 400 BAD_REQUEST | One or more environment variable names are reserved by the Hopsworks platform. | | 240032 | `VLLM_VERSION_NOT_AVAILABLE` | 400 BAD_REQUEST | vLLM version not available | | 240033 | `GIT_SYNC_NOT_ENABLED` | 400 BAD_REQUEST | Git auto-redeploy is not enabled for this deployment | | 240034 | `GIT_COMMIT_NOT_VALID` | 400 BAD_REQUEST | Git commit hash is not valid | | 240035 | `GIT_AUTO_REDEPLOY_NOT_SUPPORTED` | 400 BAD_REQUEST | Git auto-redeploy is only supported for Git-backed agent deployments | | 240036 | `BLOCKED_ENV_VAR_NAME` | 400 BAD_REQUEST | One or more environment variable names are not allowed on model deployments. Remove them from the deployment and retry. | | 240037 | `SCHEMA_NOT_FOUND` | 404 NOT_FOUND | Deployment schema not found | | 240038 | `SCHEMA_READ_ERROR` | 500 INTERNAL_SERVER_ERROR | Deployment schema could not be read | | 240039 | `UPDATE_DEPLOYMENT_MODE_ERROR` | 400 BAD_REQUEST | The deployment mode (Knative or Standard) cannot be changed while the deployment is running. Stop the deployment first. | | 240040 | `LOG_PERSISTENCE_NOT_SUPPORTED` | 400 BAD_REQUEST | Disk logging is only supported for Python model deployments. | | 240041 | `INVALID_FEATURE_LOGGING_CONFIG` | 400 BAD_REQUEST | Invalid feature logging configuration | | 240052 | `DEPLOYMENT_VERSION_NOT_FOUND` | 404 NOT_FOUND | Deployment version not found | | 240053 | `DEPLOYMENT_VERSION_MISMATCH` | 409 CONFLICT | The deployment changed since it was read. Reload it and retry. | | 240054 | `DEPLOYMENT_VERSION_CONFLICT` | 409 CONFLICT | A deployment version with this number was created concurrently. Retry. | | 240055 | `DEPLOYMENT_FILE_NOT_FOUND` | 400 BAD_REQUEST | Deployment artifact file not found | ## InferenceErrorCode | Code | Name | HTTP status | Message | | --- | --- | --- | --- | | 250000 | `SERVING_NOT_FOUND` | 404 NOT_FOUND | Serving instance not found | | 250001 | `SERVING_NOT_RUNNING` | 400 BAD_REQUEST | Serving instance not running | | 250002 | `REQUEST_ERROR` | 500 INTERNAL_SERVER_ERROR | Error contacting the serving server | | 250003 | `EMPTY_RESPONSE` | 500 INTERNAL_SERVER_ERROR | Empty response from the serving server | | 250004 | `BAD_REQUEST` | 400 BAD_REQUEST | Request malformed | | 250005 | `MISSING_VERB` | 400 BAD_REQUEST | Verb is missing | | 250006 | `ERROR_READING_RESPONSE` | 500 INTERNAL_SERVER_ERROR | Error while reading the response | | 250007 | `SERVING_INSTANCE_INTERNAL` | 500 INTERNAL_SERVER_ERROR | Serving instance internal error | | 250008 | `SERVING_INSTANCE_BAD_REQUEST` | 400 BAD_REQUEST | Serving instance bad request error | | 250009 | `REQUEST_AUTH_TYPE_NOT_SUPPORTED` | 400 BAD_REQUEST | Authentication type not supported | | 250010 | `UNAUTHORIZED` | 401 UNAUTHORIZED | Unauthorized request | | 250011 | `FORBIDDEN` | 403 FORBIDDEN | Forbidden request | | 250012 | `ENDPOINT_NOT_FOUND` | 500 INTERNAL_SERVER_ERROR | Inference endpoint not found | ## ActivitiesErrorCode | Code | Name | HTTP status | Message | | --- | --- | --- | --- | | 260000 | `FORBIDDEN` | 403 FORBIDDEN | You are not allow to perform this action. | | 260001 | `ACTIVITY_NOT_FOUND` | 404 NOT_FOUND | Activity instance not found | | 260002 | `ACTIVITY_NOT_SUPPORTED` | 400 BAD_REQUEST | Activity type not supported | ## FeaturestoreErrorCode | Code | Name | HTTP status | Message | | --- | --- | --- | --- | | 270001 | `COULD_NOT_CREATE_FEATUREGROUP` | 500 INTERNAL_SERVER_ERROR | Could not create feature group and corresponding online/offline store. | | 270002 | `FEATURESTORE_ID_NOT_PROVIDED` | 400 BAD_REQUEST | Featurestore Id was not provided | | 270003 | `FEATUREGROUP_ID_NOT_PROVIDED` | 400 BAD_REQUEST | Featuregroup Id was not provided | | 270004 | `FEATUREGROUP_VERSION_NOT_PROVIDED` | 400 BAD_REQUEST | Featuregroup version was not provided | | 270005 | `COULD_NOT_DELETE_FEATUREGROUP` | 500 INTERNAL_SERVER_ERROR | Could not delete feature group and corresponding Hive table | | 270006 | `COULD_NOT_CREATE_FEATURESTORE` | 500 INTERNAL_SERVER_ERROR | Could not create feature store and corresponding Hive database | | 270007 | `COULD_NOT_PREVIEW_FEATUREGROUP` | 500 INTERNAL_SERVER_ERROR | Could not preview the contents of the feature group | | 270008 | `FEATURESTORE_NOT_FOUND` | 404 NOT_FOUND | Featurestore wasn't found. | | 270009 | `FEATUREGROUP_NOT_FOUND` | 404 NOT_FOUND | Featuregroup wasn't found. | | 270010 | `COULD_NOT_FETCH_FEATUREGROUP_SHOW_CREATE_SCHEMA` | 500 INTERNAL_SERVER_ERROR | The query SHOW CREATE SCHEMA for the featuregroup in Hive failed. | | 270011 | `FEATURE_STORE_NOT_SHARED` | 400 BAD_REQUEST | Trying to un-share a featurestore that is not shared | | 270012 | `TRAINING_DATASET_NOT_FOUND` | 404 NOT_FOUND | Training dataset wasn't found. | | 270013 | `TRAINING_DATASET_ID_NOT_PROVIDED` | 400 BAD_REQUEST | Training dataset Id was not provided | | 270014 | `COULD_NOT_DELETE_TRAINING_DATASET` | 500 INTERNAL_SERVER_ERROR | Could not delete training dataset | | 270015 | `CLONE_ID_NOT_PROVIDED` | 400 BAD_REQUEST | Clone Id not provided despite requesting to clone feature group version | | 270016 | `TRAINING_DATASET_ALREADY_EXISTS` | 400 BAD_REQUEST | The provided training dataset name already exists | | 270017 | `NO_PRIMARY_KEY_SPECIFIED` | 400 BAD_REQUEST | A feature group or training dataset must have a primary key specified | | 270018 | `CERTIFICATES_NOT_FOUND` | 500 INTERNAL_SERVER_ERROR | Could not find user certificates for authenticating with Hive Feature Store | | 270019 | `COULD_NOT_INITIATE_HIVE_CONNECTION` | 500 INTERNAL_SERVER_ERROR | Could not initiate connection to Hive Server | | 270020 | `HIVE_UPDATE_STATEMENT_ERROR` | 500 INTERNAL_SERVER_ERROR | Hive Update Statement failed | | 270021 | `HIVE_READ_QUERY_ERROR` | 500 INTERNAL_SERVER_ERROR | Hive Read Query failed | | 270023 | `FEATURESTORE_NAME_NOT_PROVIDED` | 400 BAD_REQUEST | Featurestore name was not provided | | 270024 | `FORBIDDEN_FEATURESTORE_OPERATION` | 403 FORBIDDEN | User is forbidden to enact these changes | | 270025 | `STORAGE_CONNECTOR_ID_NOT_PROVIDED` | 400 BAD_REQUEST | Storage backend id not provided | | 270026 | `CANNOT_FETCH_HIVE_SCHEMA_FOR_ON_DEMAND_FEATUREGROUPS` | 400 BAD_REQUEST | Fetching Hive Schema of On-demand feature groups is not supported | | 270027 | `ON_DEMAND_FEATUREGROUP_JDBC_CONNECTOR_NOT_FOUND` | 404 NOT_FOUND | The JDBC Connector for the on-demand feature group could not be found | | 270028 | `PREVIEW_NOT_SUPPORTED_FOR_ON_DEMAND_FEATUREGROUPS` | 400 BAD_REQUEST | Fetching Hive Schema of On-demand feature groups is not supported | | 270029 | `CLEAR_OPERATION_NOT_SUPPORTED_FOR_ON_DEMAND_FEATUREGROUPS` | 400 BAD_REQUEST | Clearing Feature Group contents is not supported for on-demand feature groups | | 270030 | `ILLEGAL_STORAGE_CONNECTOR_NAME` | 400 BAD_REQUEST | Illegal storage connector name | | 270031 | `ILLEGAL_STORAGE_CONNECTOR_DESCRIPTION` | 400 BAD_REQUEST | Illegal storage connector description | | 270032 | `ILLEGAL_JDBC_CONNECTION_STRING` | 400 BAD_REQUEST | Illegal JDBC Connection String | | 270033 | `ILLEGAL_JDBC_CONNECTION_ARGUMENTS` | 400 BAD_REQUEST | Illegal JDBC Connection Arguments | | 270034 | `ILLEGAL_S3_CONNECTOR_BUCKET` | 400 BAD_REQUEST | Illegal S3 connector bucket | | 270035 | `ILLEGAL_S3_CONNECTOR_ACCESS_KEY` | 400 BAD_REQUEST | Illegal S3 connector access key | | 270036 | `ILLEGAL_S3_CONNECTOR_SECRET_KEY` | 400 BAD_REQUEST | Illegal S3 connector secret key | | 270037 | `ILLEGAL_HOPSFS_CONNECTOR_DATASET` | 400 BAD_REQUEST | Illegal Hopsfs connector dataset | | 270040 | `ILLEGAL_FEATURE_NAME` | 400 BAD_REQUEST | Illegal feature name | | 270041 | `ILLEGAL_FEATURE_DESCRIPTION` | 400 BAD_REQUEST | Illegal feature description | | 270042 | `CONNECTOR_NOT_FOUND` | 404 NOT_FOUND | Connector not found | | 270043 | `CONNECTOR_ID_NOT_PROVIDED` | 400 BAD_REQUEST | Connector Id was not provided | | 270044 | `INVALID_SQL_QUERY` | 400 BAD_REQUEST | Invalid SQL query | | 270046 | `HOPSFS_CONNECTOR_NOT_FOUND` | 404 NOT_FOUND | HopsFs Connector not found | | 270047 | `STORAGE_CONNECTOR_TYPE_NOT_PROVIDED` | 400 BAD_REQUEST | Storage Connector Type was not provided | | 270048 | `COULD_NOT_CLEAR_FEATUREGROUP` | 500 INTERNAL_SERVER_ERROR | Could not clear contents of feature group | | 270049 | `ILLEGAL_FEATUREGROUP_TYPE` | 400 BAD_REQUEST | The provided feature group type was not recognized | | 270050 | `ILLEGAL_TRAINING_DATASET_TYPE` | 400 BAD_REQUEST | The provided training dataset type was not recognized | | 270051 | `CAN_ONLY_GET_INODE_FOR_HOPSFS_TRAINING_DATASETS` | 400 BAD_REQUEST | Getting the inode id of a non-hopsfs training dataset is not supported | | 270052 | `TRAINING_DATASET_VERSION_NOT_PROVIDED` | 400 BAD_REQUEST | Training Dataset version was not provided | | 270054 | `S3_CONNECTOR_ID_NOT_PROVIDED` | 400 BAD_REQUEST | S3 Connector Id was not provided | | 270055 | `HOPSFS_CONNECTOR_ID_NOT_PROVIDED` | 400 BAD_REQUEST | HopsFS Connector Id was not provided | | 270057 | `ILLEGAL_TRAINING_DATASET_DATA_FORMAT` | 400 BAD_REQUEST | Illegal training dataset data format | | 270058 | `ILLEGAL_TRAINING_DATASET_VERSION` | 400 BAD_REQUEST | Illegal training dataset version | | 270059 | `ILLEGAL_FEATUREGROUP_VERSION` | 400 BAD_REQUEST | Illegal feature group version | | 270060 | `ILLEGAL_STORAGE_CONNECTOR_TYPE` | 400 BAD_REQUEST | The provided storage connector type is not valid | | 270061 | `FEATURESTORE_INITIALIZATION_ERROR` | 500 INTERNAL_SERVER_ERROR | Featurestore Initialization Error | | 270062 | `FEATURESTORE_UTIL_ARGS_FAILURE` | 500 INTERNAL_SERVER_ERROR | Could not write featurestore util args to HDFS | | 270063 | `FEATURESTORE_ONLINE_SECRETS_ERROR` | 500 INTERNAL_SERVER_ERROR | Could not get JDBC connection for the online featurestore | | 270064 | `FEATURESTORE_ONLINE_NOT_ENABLED` | 400 BAD_REQUEST | Online featurestore not enabled | | 270065 | `SYNC_TABLE_NOT_FOUND` | 400 BAD_REQUEST | The Hive Table to Sync with the feature store was not found in the metastore | | 270066 | `COULD_NOT_INITIATE_MYSQL_CONNECTION_TO_ONLINE_FEATURESTORE` | 500 INTERNAL_SERVER_ERROR | Could not initiate connection to MySQL Server | | 270067 | `MYSQL_JDBC_UPDATE_STATEMENT_ERROR` | 500 INTERNAL_SERVER_ERROR | MySQL JDBC Update Statement failed | | 270068 | `MYSQL_JDBC_READ_QUERY_ERROR` | 500 INTERNAL_SERVER_ERROR | MySQL JDBC Read Query failed | | 270069 | `ONLINE_FEATURE_SERVING_NOT_SUPPORTED_FOR_ON_DEMAND_FEATUREGROUPS` | 400 BAD_REQUEST | Online Feature Serving is onlysupported for feature groups that are cached inside Hopsworks | | 270070 | `ERROR_CREATING_ONLINE_FEATURESTORE_DB` | 500 INTERNAL_SERVER_ERROR | An error occurred when trying to create the MySQL database for an online feature store | | 270071 | `ERROR_CREATING_ONLINE_FEATURESTORE_USER` | 500 INTERNAL_SERVER_ERROR | An error occurred when trying to create the MySQL database user for an online feature store | | 270072 | `ERROR_DELETING_ONLINE_FEATURESTORE_DB` | 500 INTERNAL_SERVER_ERROR | An error occurred when trying to delete the MySQL database for an online feature store | | 270073 | `ERROR_DELETING_ONLINE_FEATURESTORE_USER` | 500 INTERNAL_SERVER_ERROR | An error occurred when trying to delete the MySQL user for an online feature store | | 270074 | `ERROR_GRANTING_ONLINE_FEATURESTORE_USER_PRIVILEGES` | 500 INTERNAL_SERVER_ERROR | An error occurred when trying to grant/revoke privileges to a MySQL user for an online feature store | | 270075 | `ONLINE_FEATUREGROUP_CANNOT_BE_PARTITIONED` | 500 INTERNAL_SERVER_ERROR | An error occurred when trying to create the MySQL table for the online feature group. User-defined partitioning is not supported for MySQL tables | | 270076 | `COULD_NOT_CREATE_DATA_VALIDATION_RULES` | 500 INTERNAL_SERVER_ERROR | Failed to create data validation rules | | 270077 | `COULD_NOT_READ_DATA_VALIDATION_RESULT` | 500 INTERNAL_SERVER_ERROR | Failed to read data validation result | | 270078 | `IMPORT_JOB_ALREADY_RUNNING` | 400 BAD_REQUEST | A job to import this featuregroup is already running | | 270079 | `IMPORT_CONF_ERROR` | 500 INTERNAL_SERVER_ERROR | Error writing import job configuration | | 270080 | `TRAININGDATASETJOB_FAILURE` | 500 INTERNAL_SERVER_ERROR | Could not write featurestore cloud args to HDFS | | 270081 | `TRAININGDATASETJOB_DUPLICATE_FEATURE` | 400 BAD_REQUEST | Feature list contains duplicate | | 270082 | `FEATURE_DOES_NOT_EXIST` | 400 BAD_REQUEST | Feature does not exist | | 270083 | `TRAININGDATASETJOB_FEATUREGROUP_DUPLICATE` | 400 BAD_REQUEST | Multiple featuregroups contain feature | | 270084 | `TRAININGDATASETJOB_TRAININGDATASET_VERSION_EXISTS` | 400 BAD_REQUEST | Illegal training dataset name - version combination | | 270085 | `TRAININGDATASETJOB_CONF_ERROR` | 500 INTERNAL_SERVER_ERROR | Error writing training dataset job configuration to hdfs | | 270086 | `S3_KEYS_FORBIDDEN` | 400 BAD_REQUEST | IAM role is configured for this instance. AWS access/secret keys are not allowed | | 270087 | `MISSING_REDSHIFT_DRIVER` | 400 BAD_REQUEST | Could not find Redshift JDBC driver. Please upload it in Resources/RedshiftJDBC42-no-awssdk.jar | | 270088 | `TRAININGDATASETJOB_MISSPECIFICATION` | 400 BAD_REQUEST | Training dataset job is misspecified and cannot be created | | 270089 | `FEATUREGROUP_EXISTS` | 400 BAD_REQUEST | The feature group you are trying to create does already exist. | | 270090 | `XATTRS_OPERATIONS_ONLY_SUPPORTED_FOR_CACHED_FEATUREGROUPS` | 400 BAD_REQUEST | Attaching extended attributes is only supported for cached featuregroups. | | 270091 | `ILLEGAL_ENTITY_NAME` | 400 BAD_REQUEST | Illegal feature store entity name | | 270092 | `ILLEGAL_ENTITY_DESCRIPTION` | 400 BAD_REQUEST | Illegal featurestore entity description | | 270094 | `FEATUREGROUP_NAME_NOT_PROVIDED` | 400 BAD_REQUEST | Feature group name was not provided | | 270095 | `TRAINING_DATASET_NAME_NOT_PROVIDED` | 400 BAD_REQUEST | Training dataset name was not provided | | 270096 | `NO_PK_JOINING_KEYS` | 400 BAD_REQUEST | Could not find any matching feature to join | | 270097 | `LEFT_RIGHT_ON_DIFF_SIZES` | 400 BAD_REQUEST | LeftOn and RightOn have different sizes | | 270098 | `ILLEGAL_TRAINING_DATASET_SPLIT_NAME` | 400 BAD_REQUEST | Illegal training dataset split name | | 270099 | `ILLEGAL_TRAINING_DATASET_SPLIT_PERCENTAGE` | 400 BAD_REQUEST | Illegal training dataset split percentage | | 270100 | `TAG_NOT_ALLOWED` | 400 BAD_REQUEST | The provided tag is not allowed | | 270101 | `TAG_NOT_FOUND` | 404 NOT_FOUND | The provided tag is not attached | | 270102 | `FEATUREGROUP_NOT_ONLINE` | 400 BAD_REQUEST | The feature group is not available online | | 270103 | `FEATUREGROUP_ONDEMAND_NO_PARTS` | 400 BAD_REQUEST | Partitions not available for on demand feature group | | 270104 | `ILLEGAL_S3_CONNECTOR_SERVER_ENCRYPTION_ALGORITHM` | 400 BAD_REQUEST | Illegal server encryption algorithm provided | | 270105 | `ILLEGAL_S3_CONNECTOR_SERVER_ENCRYPTION_KEY` | 400 BAD_REQUEST | Illegal server encryption key provided | | 270106 | `TRAINING_DATASET_DUPLICATE_SPLIT_NAMES` | 400 BAD_REQUEST | Duplicate split names in training dataset provided. | | 270107 | `STATISTICS_READ_ERROR` | 500 INTERNAL_SERVER_ERROR | Error reading the statistics | | 270108 | `ILLEGAL_STATISTICS_CONFIG` | 400 BAD_REQUEST | Illegal statistics config | | 270109 | `ERROR_DELETING_STATISTICS` | 500 INTERNAL_SERVER_ERROR | Error deleting the statistics of a feature store entity | | 270110 | `ERROR_GETTING_S3_CONNECTOR_ACCESS_AND_SECRET_KEY_FROM_SECRET` | 500 INTERNAL_SERVER_ERROR | Could not get access and secret key from the user secret | | 270111 | `TRAINING_DATASET_NO_QUERY` | 400 BAD_REQUEST | The training dataset wasn't generated from a query | | 270112 | `TRAINING_DATASET_NO_SCHEMA` | 400 BAD_REQUEST | No query or feature schema provided | | 270113 | `QUERY_FAILED_FG_DELETED` | 400 BAD_REQUEST | Cannot generate query, some feature groups were deleted | | 270114 | `ILLEGAL_FEATUREGROUP_UPDATE` | 400 BAD_REQUEST | Illegal feature group update | | 270115 | `COULD_NOT_ALTER_FEAUTURE_GROUP_METADATA` | 500 INTERNAL_SERVER_ERROR | Failed to alter feature group meta data | | 270116 | `COULD_NOT_GET_FEATURE_GROUP_METADATA` | 500 INTERNAL_SERVER_ERROR | Failed to retrieve feature group meta data | | 270117 | `ERROR_CREATING_HIVE_METASTORE_CLIENT` | 500 INTERNAL_SERVER_ERROR | Failed to open Hive Metastore client | | 270118 | `NO_DATA_AVAILABLE_FEATUREGROUP_COMMITDATE` | 404 NOT_FOUND | No data is available for feature group with this commit date | | 270119 | `PROVIDED_DATE_FORMAT_NOT_SUPPORTED` | 400 BAD_REQUEST | Invalid date format | | 270120 | `ONLINE_FEATURESTORE_JDBC_CONNECTOR_NOT_FOUND` | 500 INTERNAL_SERVER_ERROR | Online featurestore JDBC connector not found | | 270121 | `PRIMARY_KEY_REQUIRED` | 400 BAD_REQUEST | Primary key is required when using Hudi time travel format | | 270122 | `DATABRICKS_INSTANCE_ALREADY_EXISTS` | 409 CONFLICT | Databricks Instance already registered | | 270123 | `DATABRICKS_INSTANCE_NOT_EXISTS` | 404 NOT_FOUND | Databricks Instance doesn't exists | | 270124 | `DATABRICKS_CANNOT_START_CLUSTER` | 500 INTERNAL_SERVER_ERROR | Could not start Databricks cluster | | 270125 | `DATABRICKS_ERROR` | 500 INTERNAL_SERVER_ERROR | Error communicating with Databricks | | 270126 | `STORAGE_CONNECTOR_GET_ERROR` | 500 INTERNAL_SERVER_ERROR | Error retrieving the storage connector | | 270127 | `ERROR_ONLINE_FEATURES` | 500 INTERNAL_SERVER_ERROR | Error retrieving online features | | 270128 | `ERROR_ONLINE_USERS` | 500 INTERNAL_SERVER_ERROR | Error getting database users | | 270129 | `ERROR_ONLINE_GENERIC` | 500 INTERNAL_SERVER_ERROR | Error communicating with the online feature store | | 270130 | `COULD_NOT_CREATE_ON_DEMAND_FEATUREGROUP` | 500 INTERNAL_SERVER_ERROR | Could not create on demand feature group | | 270131 | `COULD_NOT_DELETE_ON_DEMAND_FEATUREGROUP` | 500 INTERNAL_SERVER_ERROR | Could not delete on demand feature group | | 270132 | `ILLEGAL_FEATURE_GROUP_FEATURE_DEFAULT_VALUE` | 400 BAD_REQUEST | Illegal feature default value | | 270133 | `KEYWORD_ERROR` | 500 INTERNAL_SERVER_ERROR | Keyword error for feature group/training dataset | | 270134 | `KEYWORD_FORMAT_ERROR` | 400 BAD_REQUEST | Keyword format error | | 270135 | `REDSHIFT_CONNECTOR_NOT_FOUND` | 404 NOT_FOUND | Redshift Connector not found | | 270136 | `ILLEGAL_STORAGE_CONNECTOR_ARG` | 400 BAD_REQUEST | Illegal storage connector argument | | 270137 | `ERROR_SAVING_STATISTICS` | 400 BAD_REQUEST | Error saving statistics | | 270138 | `FILTER_CONSTRUCTION_ERROR` | 500 INTERNAL_SERVER_ERROR | Failed to construct filter condition | | 270139 | `ILLEGAL_FILTER_ARGUMENTS` | 400 BAD_REQUEST | Malformed filter conditions for Query | | 270140 | `ILLEGAL_ON_DEMAND_DATA_FORMAT` | 400 BAD_REQUEST | Illegal on-demand feature group data format | | 270141 | `ERROR_JOB_SETUP` | 500 INTERNAL_SERVER_ERROR | Error setting up feature store job | | 270142 | `LABEL_NOT_FOUND` | 404 NOT_FOUND | Could not find label in training dataset schema | | 270143 | `DATA_VALIDATION_RESULTS_NOT_FOUND` | 404 NOT_FOUND | Could not find feature group validation results. Make sure the results file was not manually removed from the dataset | | 270144 | `DATA_VALIDATION_NOT_FOUND` | 404 NOT_FOUND | Could not find feature group validation. | | 270145 | `FEATURE_STORE_EXPECTATION_NOT_FOUND` | 404 NOT_FOUND | Could not find feature store expectation. | | 270146 | `FEATURE_GROUP_EXPECTATION_NOT_FOUND` | 404 NOT_FOUND | Could not find feature group expectation. | | 270147 | `FEATURE_GROUP_EXPECTATION_FEATURE_NOT_FOUND` | 404 NOT_FOUND | Could not find expectation feature(s) in feature group expectation. | | 270148 | `FEATURE_STORE_RULE_NOT_FOUND` | 404 NOT_FOUND | Could not find feature store data validation rule. | | 270149 | `FEATURE_GROUP_CHECKS_FAILED` | 417 EXPECTATION_FAILED | Feature group validation checks did not pass, will not persist the data. | | 270150 | `RULE_NOT_FOUND` | 404 NOT_FOUND | Rule with provided name was not found. | | 270151 | `AVRO_PRIMITIVE_TYPE_NOT_SUPPORTED` | 400 BAD_REQUEST | Error converting Hive Type to Avro primitive type | | 270152 | `AVRO_MAP_STRING_KEY` | 400 BAD_REQUEST | Map types are only supported with STRING type keys | | 270153 | `AVRO_MALFORMED_SCHEMA` | 500 INTERNAL_SERVER_ERROR | Error converting Hive schema to Avro | | 270154 | `FEATURE_GROUP_EXPECTATION_FEATURE_TYPE_INVALID` | 400 BAD_REQUEST | Could not attach expectation because some feature types did not match rule types. | | 270155 | `ALERT_NOT_FOUND` | 404 NOT_FOUND | Alert not found | | 270156 | `ALERT_ILLEGAL_ARGUMENT` | 400 BAD_REQUEST | Alert missing argument. | | 270157 | `ALERT_ALREADY_EXISTS` | 400 BAD_REQUEST | Alert with the same status already exists. | | 270158 | `ERROR_DELETING_TRANSFORMERFUNCTION` | 500 INTERNAL_SERVER_ERROR | Error deleting the transformer function of a feature store entity | | 270159 | `TRANSFORMATION_FUNCTION_ALREADY_EXISTS` | 400 BAD_REQUEST | The provided transformation function name and version already exists | | 270160 | `TRANSFORMATION_FUNCTION_DOES_NOT_EXIST` | 400 BAD_REQUEST | Transformation function does not exist | | 270161 | `TRANSFORMATION_FUNCTION_READ_ERROR` | 500 INTERNAL_SERVER_ERROR | Error reading the transformation function | | 270162 | `TRANSFORMATION_FUNCTION_VERSION` | 400 BAD_REQUEST | Illegal transformation function version | | 270163 | `ILLEGAL_TRANSFORMATION_FUNCTION_OUTPUT_TYPE` | 400 BAD_REQUEST | Illegal transformation function output type | | 270164 | `FEATURE_WITH_TRANSFORMATION_NOT_FOUND` | 404 NOT_FOUND | Could not find feature in training dataset schema | | 270165 | `ILLEGAL_PREFIX_NAME` | 400 BAD_REQUEST | Illegal feature name | | 270169 | `FAILED_TO_CREATE_ROUTE` | 400 BAD_REQUEST | Failed to create route. | | 270170 | `FAILED_TO_DELETE_ROUTE` | 400 BAD_REQUEST | Failed to delete route. | | 270171 | `ILLEGAL_EVENT_TIME_FEATURE_TYPE` | 400 BAD_REQUEST | Illegal event time feature type | | 270172 | `EVENT_TIME_FEATURE_NOT_FOUND` | 400 BAD_REQUEST | Event time feature not found | | 270173 | `FEATURE_GROUP_MISSING_EVENT_TIME` | 400 BAD_REQUEST | Feature group is not event time enabled | | 270174 | `JOIN_OPERATOR_MISMATCH` | 400 BAD_REQUEST | Join features and operator list have different sizes | | 270175 | `VALIDATION_RULE_INCOMPLETE` | 400 BAD_REQUEST | Rule is missing a required field. | | 270176 | `COULD_NOT_CREATE_ONLINE_FEATUREGROUP` | 400 BAD_REQUEST | Could not create online feature group | | 270177 | `COULD_NOT_GET_QUERY_FILTER` | 500 INTERNAL_SERVER_ERROR | Error getting query filter | | 270178 | `ERROR_REGISTER_BUILTIN_TRANSFORMATION_FUNCTION` | 500 INTERNAL_SERVER_ERROR | This branch should not be reached. Please fix automatic registering of the built-in transformation functions upon project creation | | 270179 | `FEATURE_VIEW_ALREADY_EXISTS` | 400 BAD_REQUEST | The provided feature view name and version already exists | | 270180 | `FEATURE_VIEW_CREATION_ERROR` | 400 BAD_REQUEST | Cannot create feature view. | | 270181 | `FEATURE_VIEW_NOT_FOUND` | 404 NOT_FOUND | Feature view wasn't found. | | 270182 | `KAFKA_STORAGE_CONNECTOR_STORE_NOT_EXISTING` | 400 BAD_REQUEST | Provided certificate store location does not exist | | 270183 | `VALIDATION_NOT_SUPPORTED` | 400 BAD_REQUEST | Rule is not supported. | | 270184 | `STREAM_FEATURE_GROUP_ONLINE_DISABLE_ENABLE` | 400 BAD_REQUEST | Stream feature group cannot be online enabled if it was created as offline only. | | 270185 | `GCS_FIELD_MISSING` | 400 BAD_REQUEST | Field missing | | 270186 | `TRAINING_DATASET_COULD_NOT_BE_CREATED` | 500 INTERNAL_SERVER_ERROR | Could not create training dataset | | 270187 | `NESTED_JOIN_NOT_ALLOWED` | 400 BAD_REQUEST | Nested join is not supported. | | 270188 | `FEATURE_NOT_FOUND` | 404 NOT_FOUND | Could not find feature. | | 270189 | `EXPECTATION_TYPE_NOT_FOUND` | 404 NOT_FOUND | Expectation type not supported. | | 270190 | `EXPECTATION_NOT_FOUND` | 404 NOT_FOUND | Expectation not found. | | 270191 | `NO_EXPECTATION_SUITE_ATTACHED_TO_THIS_FEATUREGROUP` | 404 NOT_FOUND | No Expectation Suite attached to this feature group. Use fg.save_expectation_suite to attach a Great Expectations suite to your FeatureGroup. | | 270192 | `VALIDATION_REPORT_NOT_FOUND` | 404 NOT_FOUND | Validation report not found. | | 270193 | `FAILED_TO_PARSE_EXPECTATION_CONFIG_TO_JSON` | 400 BAD_REQUEST | Failed to parse expectation config field to json. Expectation config must be a valid json to fetch the expectationId from the meta field. | | 270194 | `KEY_NOT_FOUND_OR_INVALID_VALUE_TYPE_IN_JSON_OBJECT` | 400 BAD_REQUEST | Requested key has not been found in Json object or associated value is not of required type. | | 270195 | `FAILED_TO_PARSE_VALIDATION_RESULT_FOR_OBSERVED_VALUE` | 400 BAD_REQUEST | Failed to parse result json to get observed_value field | | 270196 | `FAILED_TO_PARSE_EXPECTATION_META_FIELD` | 400 BAD_REQUEST | Failed to parse expectation meta field. | | 270197 | `VALIDATION_REPORT_IS_NOT_VALID_JSON` | 400 BAD_REQUEST | Validation report is not a valid JSON. | | 270198 | `ERROR_SAVING_ON_DISK_VALIDATION_REPORT` | 500 INTERNAL_SERVER_ERROR | Error saving full json report to disk. | | 270199 | `ERROR_DELETING_ON_DISK_VALIDATION_REPORT` | 500 INTERNAL_SERVER_ERROR | Error deleting on-disk validation report. You can delete the report manually using the file browser in the project setting tab. Reports are stored by default in the DataValidation directory, under the corresponding feature group name and version subdirectories. | | 270200 | `INPUT_FIELD_EXCEEDS_MAX_ALLOWED_CHARACTER` | 400 BAD_REQUEST | Input field length exceeds max allowed characters. | | 270201 | `INPUT_FIELD_IS_NOT_VALID_JSON` | 400 BAD_REQUEST | Input field fail to be parsed to valid Json. | | 270202 | `INPUT_FIELD_IS_NOT_NULLABLE` | 400 BAD_REQUEST | Input field is not nullable. | | 270203 | `ERROR_INFERRING_INGESTION_RESULT` | 400 BAD_REQUEST | Could not infer ingestion result from validation ingestion policy and validation success. | | 270204 | `FAILED_TO_DELETE_TD_DATA` | 400 BAD_REQUEST | Failed to delete training dataset. | | 270205 | `ERROR_DELETING_FEATURE_VIEW` | 500 INTERNAL_SERVER_ERROR | Error deleting feature view. | | 270206 | `ILLEGAL_TRAINING_DATASET_TIME_SERIES_SPLIT` | 400 BAD_REQUEST | Illegal training dataset time series split. | | 270207 | `ILLEGAL_EXPECTATION_UPDATE` | 400 BAD_REQUEST | Illegal Expectation update. To preserve the validation history this update is not allowed. Create a new Expectation by removing expectationId from meta field instead. | | 270208 | `EXPECTATION_SUITE_ALREADY_EXISTS` | 409 CONFLICT | An expectation suite is already attached to this feature group. Either update the existing suite via the update endpoint or delete it first. | | 270209 | `FAILURE_HDFS_USER_OPERATION` | 500 INTERNAL_SERVER_ERROR | HDFS user operation failure | | 270210 | `FEATURE_NAME_NOT_FOUND` | 400 BAD_REQUEST | The Feature Name was not found in this version of the Feature Group. | | 270211 | `VALIDATION_RESULT_IS_NOT_VALID_JSON` | 400 BAD_REQUEST | The validation result is not a valid json. | | 270212 | `FEATURE_OFFLINE_TYPE_NOT_PROVIDED` | 400 BAD_REQUEST | Feature offline type cannot be null or empty. | | 270213 | `AMBIGUOUS_FEATURE_ERROR` | 400 BAD_REQUEST | Feature name is ambiguous. | | 270214 | `STORAGE_CONNECTOR_TYPE_NOT_ENABLED` | 400 BAD_REQUEST | Storage connector type not enabled | | 270215 | `COULD_NOT_SHARE_FEATURE_STORE` | 500 INTERNAL_SERVER_ERROR | Could not share feature store | | 270216 | `FILE_DELETION_ERROR` | 400 BAD_REQUEST | Failed to delete file | | 270217 | `FILE_READ_ERROR` | 400 BAD_REQUEST | Failed to read file | | 270218 | `DOCKER_FULLNAME_ERROR` | 400 BAD_REQUEST | Failed to retrieve full docker image name | | 270219 | `CONNECTION_CHECKER_LAUNCH_ERROR` | 400 BAD_REQUEST | Failed to launch process to start docker container for testing connection | | 270220 | `CONNECTION_CHECKER_ERROR` | 400 BAD_REQUEST | Failure in testing connection for storage connector | | 270221 | `ERROR_CREATING_ONLINE_FEATURESTORE_KAFKA_OFFSET_TABLE` | 500 INTERNAL_SERVER_ERROR | An error occurred when trying to create the kafka offset table for an online feature store | | 270222 | `ERROR_CONSTRUCTING_VALIDATION_REPORT_DIRECTORY_PATH` | 500 INTERNAL_SERVER_ERROR | An error occurred while constructing validation report directory path | | 270223 | `SPINE_GROUP_ON_RIGHT_SIDE_OF_JOIN_NOT_ALLOWED` | 400 BAD_REQUEST | Spine groups cannot be used on the right sideof a feature view join. | | 270224 | `FEATURE_GROUP_DUPLICATE_FEATURE` | 400 BAD_REQUEST | Feature list contains duplicate | | 270225 | `HELPER_COL_NOT_FOUND` | 404 NOT_FOUND | Could not find helper column in feature view schema | | 270226 | `OPENSEARCH_DEFAULT_EMBEDDING_INDEX_SUFFIX_NOT_DEFINED` | 500 INTERNAL_SERVER_ERROR | Opensearch default embedding index not defined | | 270227 | `FEATURE_GROUP_COMMIT_NOT_FOUND` | 400 BAD_REQUEST | Feature group commit not found | | 270228 | `STATISTICS_NOT_FOUND` | 404 NOT_FOUND | Statistics wasn't found. | | 270229 | `INVALID_STATISTICS_WINDOW_TIMES` | 400 BAD_REQUEST | Window times provided are invalid | | 270230 | `COULD_NOT_DELETE_VECTOR_DB_INDEX` | 500 INTERNAL_SERVER_ERROR | Could not delete index from vector db. | | 270231 | `COULD_NOT_INITIATE_ARROW_FLIGHT_CONNECTION` | 500 INTERNAL_SERVER_ERROR | Could not initiate connection to Arrow Flight server | | 270232 | `ARROW_FLIGHT_READ_QUERY_ERROR` | 400 BAD_REQUEST | Arrow Flight server Read Query failed | | 270233 | `FEATURE_MONITORING_ENTITY_NOT_FOUND` | 404 NOT_FOUND | Feature Monitoring entity not found. | | 270234 | `FEATURE_MONITORING_NOT_ENABLED` | 400 BAD_REQUEST | Feature monitoring is not enabled. | | 270235 | `FEATURE_NOT_FOUND_IN_VECTOR_DB` | 500 INTERNAL_SERVER_ERROR | Feature not found in vector db. | | 270236 | `COULD_NOT_PREVIEW_DATA_IN_VECTOR_DB` | 500 INTERNAL_SERVER_ERROR | Could not preview data in vector database. | | 270237 | `EMBEDDING_FEATURE_NOT_FOUND` | 400 BAD_REQUEST | Embedding feature cannot be found in feature group. | | 270238 | `COULD_NOT_GET_VECTOR_DB_INDEX` | 500 INTERNAL_SERVER_ERROR | Could not get index from vector db. | | 270239 | `EMBEDDING_INDEX_EXISTED` | 400 BAD_REQUEST | Embedding index already exists. | | 270240 | `INVALID_EMBEDDING_INDEX_NAME` | 400 BAD_REQUEST | Embedding index name is not valid. | | 270241 | `VECTOR_DATABASE_INDEX_MAPPING_LIMIT_EXCEEDED` | 400 BAD_REQUEST | Index mapping limit exceeded. | | 270242 | `VECTOR_DATABASE_DATA_TYPE_NOT_SUPPORTED` | 400 BAD_REQUEST | Provided data type is not supported by vector database. | | 270243 | `PREVIEW_NOT_SUPPORTED` | 400 BAD_REQUEST | Preview is not supported | | 270244 | `FOREIGN_KEY_NOT_PRIMARY_KEY` | 400 BAD_REQUEST | foreign key from the left feature group is not a primary key in the right feature group | | 270245 | `NESTED_JOINS_RECURSION_LIMIT_EXCEEDED` | 500 INTERNAL_SERVER_ERROR | Could not construct nested query, recursion limit exceeded | | 270246 | `ERROR_DELETING_TRANSFORMATION_FUNCTION_ATTACHED` | 500 INTERNAL_SERVER_ERROR | Cannot delete transformation function attached to a feature view | | 270247 | `HOSWORKS_ACTION_TASK_SERIALIZATION_ERROR` | 500 INTERNAL_SERVER_ERROR | Hopsworks action task serialization error. | | 270248 | `FEATURE_VIEW_LOGGING_DOES_NOT_EXIST` | 400 BAD_REQUEST | Feature view logging does not exist | | 270249 | `JOIN_ON_PARTIAL_PRIMARY_KEY` | 400 BAD_REQUEST | the join lacks a key which is part of the primary key of the feature group | | 270250 | `COULD_NOT_INITIATE_S3_CLIENT` | 500 INTERNAL_SERVER_ERROR | Could not initiate connection to s3 | | 270251 | `CONNECTOR_FIELD_MISSING` | 400 BAD_REQUEST | Field missing in storage connector | | 270252 | `COULD_NOT_INIT_VECTOR_DB` | 500 INTERNAL_SERVER_ERROR | Could not initiate vector database. | | 270253 | `TIME_TRAVEL_FORMAT_NOT_SUPPORTED` | 400 BAD_REQUEST | Not supported time travel format | | 270254 | `ERROR_CREATING_ONLINE_INGESTION_RESULT_TABLE` | 500 INTERNAL_SERVER_ERROR | An error occurred when trying to create the Online Ingestion Result table for an online feature store | | 270255 | `INVALID_OUTPUT_NAME_TRANSFORMATION_FUNCTION` | 400 BAD_REQUEST | Invalid output feature name specified for transformation function | | 270256 | `INVALID_ONLINE_DATA_TYPE` | 400 BAD_REQUEST | The provided online type is invalid | | 270257 | `DATABASE_NOT_SPECIFIED` | 400 BAD_REQUEST | The database was not specified | | 270258 | `DATABASE_CANNOT_BE_CHANGED` | 400 BAD_REQUEST | The database cannot be specified | | 270259 | `ARROW_FLIGHT_ERROR` | 500 INTERNAL_SERVER_ERROR | Arrow Flight error | | 270260 | `COULD_NOT_ENABLE_TTL` | 400 BAD_REQUEST | Could not enable ttl | | 270261 | `CHART_NOT_FOUND` | 404 NOT_FOUND | Chart wasn't found. | | 270262 | `DASHBOARD_NOT_FOUND` | 404 NOT_FOUND | Dashboard wasn't found. | | 270263 | `COULD_NOT_DELETE_FEATURE_GROUP` | 400 BAD_REQUEST | Could not delete feature group | | 270264 | `FEATURE_STORE_ALREADY_SHARED` | 400 BAD_REQUEST | Feature store already shared with project | | 270265 | `COULD_NOT_SHARE_FEATURE_GROUP` | 500 INTERNAL_SERVER_ERROR | Could not share feature group | | 270266 | `FEATURE_GROUP_NOT_SHARED` | 400 BAD_REQUEST | The feature group is not shared | | 270267 | `FEATURE_GROUP_ALREADY_SHARED` | 400 BAD_REQUEST | The feature group is already shared with the project | | 270268 | `FEATURE_NOT_SHARED` | 400 BAD_REQUEST | The feature is not shared | | 270269 | `ERROR_SIGNING_QUERY` | 500 INTERNAL_SERVER_ERROR | Could not sign the query | | 270270 | `RESTRICTED_ACCESS_ALREADY_GRANTED` | 400 BAD_REQUEST | Restricted access to this feature group is already granted to the user | | 270271 | `RESTRICTED_ACCESS_NOT_GRANTED` | 404 NOT_FOUND | Restricted access to this feature group is not granted to the user | | 270272 | `FEATUREGROUP_NO_ACCESSIBLE_FEATURES` | 400 BAD_REQUEST | No accessible features in this feature group for the current user | | 270273 | `FEATURE_GROUP_DUPLICATE_PATH` | 400 BAD_REQUEST | A feature group with the same path already exists | | 270274 | `FAILED_TO_DELETE_SINK_JOB_FOR_FEATURE_GROUP` | 500 INTERNAL_SERVER_ERROR | Failed to delete sink job for feature group | | 270275 | `COULD_NOT_DELETE_FEATURE_VIEW` | 409 CONFLICT | Could not delete feature view | | 270276 | `INVALID_ONLINE_CONFIG_PRIMARY_KEY_INDEX_TYPE` | 400 BAD_REQUEST | The provided online config primary key index type is invalid. | | 270277 | `LOOKBACK_WINDOW_MISSING_START` | 400 BAD_REQUEST | Lookback window requires `start` to be set. | | 270278 | `LOOKBACK_WINDOW_INVERTED_RANGE` | 400 BAD_REQUEST | Lookback window `start` must be strictly earlier than `end`. | | 270279 | `LOOKBACK_WINDOW_NO_EVENT_TIME` | 400 BAD_REQUEST | Lookback window `key=EVENT_TIME` requires the joined feature group to declare an event_time column. | | 270280 | `LOOKBACK_WINDOW_NO_PARTITION_KEY` | 400 BAD_REQUEST | Lookback window `key=PARTITION_KEY` requires the joined feature group to have a single DATE partition column. | | 270281 | `LOOKBACK_WINDOW_INVALID_KEY` | 400 BAD_REQUEST | Lookback window `key` must be `EVENT_TIME` or `PARTITION_KEY`. | | 270282 | `LOOKBACK_WINDOW_UNKNOWN_JOIN` | 400 BAD_REQUEST | Lookback override does not match any joined feature group. | | 270283 | `LOOKBACK_WINDOW_DUPLICATE_JOIN` | 400 BAD_REQUEST | Multiple lookback entries for the same joined feature group. | | 270284 | `DEFAULT_FEATURESTORE_NOT_CONFIGURED` | 404 NOT_FOUND | Default featurestore project is not configured. | | 270285 | `INVALID_ONLINE_CONFIG_SECONDARY_INDEX` | 400 BAD_REQUEST | The provided online config secondary index is invalid. | | 270286 | `TRANSFORMATION_FUNCTION_INPUT_TYPE_UNRESOLVABLE` | 400 BAD_REQUEST | Transformation function input feature type could not be resolved. | | 270287 | `FEATURE_MONITORING_INPUT_VALIDATION` | 400 BAD_REQUEST | Invalid feature monitoring input. | | 270288 | `PARTITIONED_BY_EMPTY` | 400 BAD_REQUEST | partitioned_by must be a non-empty list when set. | | 270289 | `PARTITIONED_BY_INVALID_GRAIN` | 400 BAD_REQUEST | partitioned_by contains an invalid or unsupported transform expression. | | 270290 | `PARTITIONED_BY_DUPLICATE` | 400 BAD_REQUEST | partitioned_by contains duplicate or redundant transforms. | | 270291 | `PARTITIONED_BY_CONFLICTS_WITH_PARTITION_KEY` | 400 BAD_REQUEST | partitioned_by cannot be set together with partition_key. | | 270292 | `PARTITIONED_BY_REQUIRES_EVENT_TIME` | 400 BAD_REQUEST | partitioned_by temporal transforms on HUDI must use the event_time column. | | 270293 | `PARTITIONED_BY_COLLIDES_WITH_EVENT_TIME` | 400 BAD_REQUEST | event_time column name collides with a partitioned_by grain. | | 270294 | `PARTITIONED_BY_COLLIDES_WITH_FEATURE` | 400 BAD_REQUEST | partitioned_by references a column that does not exist or collides with a feature name. | | 270295 | `PARTITIONED_BY_ONLINE_NOT_SUPPORTED` | 400 BAD_REQUEST | partitioned_by is not supported on online-enabled HUDI feature groups yet. | | 270296 | `PARTITIONED_BY_UNSUPPORTED_FORMAT` | 400 BAD_REQUEST | partitioned_by requires a time_travel_format that supports partition transforms (ICEBERG or HUDI); DELTA uses clustered_by. | | 270297 | `PARTITIONED_BY_HOUR_REQUIRES_TIMESTAMP` | 400 BAD_REQUEST | temporal partition transforms require a date or timestamp source column; 'hour' requires a timestamp (a date has no sub-day resolution). | | 270298 | `ONLINE_FEATUREGROUP_OFFLINE_ONLY_KEY_COLUMN` | 400 BAD_REQUEST | a primary key, event_time, or secondary-index column cannot be offline_only: the online table excludes offline_only columns, so it would reference a column that is absent from the online schema. | | 270300 | `ZORDER_BY_INVALID` | 400 BAD_REQUEST | zorder_by is empty, has duplicates, exceeds the column cap, or references a column that does not exist. | | 270301 | `ZORDER_BY_UNSUPPORTED_FORMAT` | 400 BAD_REQUEST | zorder_by requires time_travel_format ICEBERG or HUDI (DELTA uses clustered_by for liquid clustering). | | 270302 | `CLUSTERED_BY_INVALID` | 400 BAD_REQUEST | clustered_by is empty, has duplicates, exceeds the column cap, conflicts with partition_key, or references a column that does not exist. | | 270303 | `CLUSTERED_BY_UNSUPPORTED_FORMAT` | 400 BAD_REQUEST | clustered_by is Delta liquid clustering and requires time_travel_format DELTA. | | 270304 | `BUCKET_INDEX_INVALID` | 400 BAD_REQUEST | bucket_index requires a primary key field and a positive number of buckets. | | 270305 | `BUCKET_INDEX_UNSUPPORTED_FORMAT` | 400 BAD_REQUEST | bucket_index configures the Hudi bucket index and requires time_travel_format HUDI. | | 270306 | `SORT_ORDER_INVALID` | 400 BAD_REQUEST | sort_order is empty, malformed, repeats a column, or references a column that does not exist. | | 270307 | `SORT_ORDER_UNSUPPORTED_FORMAT` | 400 BAD_REQUEST | sort_order is a persistent Iceberg table sort order and requires time_travel_format ICEBERG. | | 270308 | `INVALID_STATISTICS` | 400 BAD_REQUEST | Statistics provided for registration are invalid. | | 270309 | `DATA_SOURCE_CONNECT_TIMEOUT` | 504 GATEWAY_TIMEOUT | Data source did not respond in time | | 270310 | `FEATURE_VIEW_LOGGING_INVALID_INTERVAL` | 400 BAD_REQUEST | Feature view logging materialization interval must be 'hour' or 'day' | | 270311 | `FEATURE_VIEW_LOGGING_INVALID_TRANSPORT` | 400 BAD_REQUEST | Feature view logging transport must be 'realtime' or 'job' | | 270312 | `FEATURE_VIEW_LOGGING_TRANSPORT_CONFLICT` | 400 BAD_REQUEST | Feature view already logs through the other transport; delete its log to switch | | 270313 | `FEATURE_VIEW_LOGGING_MATERIALIZATION_UNSUPPORTED` | 400 BAD_REQUEST | Feature view logging through this transport is not materialized by the platform | | 270322 | `ASOF_SPINE_EMPTY` | 400 BAD_REQUEST | The inference spine has no rows or no columns. | | 270323 | `ASOF_SPINE_NO_BINDABLE_COLUMN` | 400 BAD_REQUEST | No column of the inference spine matches a serving key, a root feature or the event time of the feature view. | | 270324 | `ASOF_SPINE_UNKNOWN_COLUMN` | 400 BAD_REQUEST | A column of the inference spine matches nothing in the feature view. | | 270325 | `ASOF_SPINE_TIME_SOURCE` | 400 BAD_REQUEST | The inference spine does not carry a usable prediction time. | | 270326 | `ASOF_SPINE_TOO_LARGE` | 400 BAD_REQUEST | The inference spine exceeds a configured limit. | | 270327 | `ASOF_SPINE_TYPE_MISMATCH` | 400 BAD_REQUEST | An inference spine column does not convert to the feature's type. | | 270328 | `ASOF_SPINE_NOT_SUPPORTED` | 400 BAD_REQUEST | The feature view cannot be read with an inference spine. | | 270329 | `ASOF_SPINE_BAD_NAME` | 400 BAD_REQUEST | The inference spine table or file name is not acceptable. | | 270330 | `ASOF_SPINE_FILE_UNREADABLE` | 400 BAD_REQUEST | The inference spine file does not exist or cannot be read. | | 270331 | `ASOF_SPINE_FILTER_ON_SPINE_COLUMN` | 400 BAD_REQUEST | A filter references a column the inference spine supplies. | | 270332 | `ASOF_SPINE_NO_FEATURE_VIEW` | 400 BAD_REQUEST | The inference spine does not name a feature view that can be read. | | 270333 | `ASOF_MAX_FEATURE_AGE_INVALID` | 400 BAD_REQUEST | The feature view's max feature age is not a positive duration. | ## AirflowErrorCode | Code | Name | HTTP status | Message | | --- | --- | --- | --- | | 290001 | `JWT_NOT_CREATED` | 500 INTERNAL_SERVER_ERROR | JWT for Airflow service could not be created | | 290002 | `JWT_NOT_STORED` | 500 INTERNAL_SERVER_ERROR | JWT for Airflow service could not be stored | | 290003 | `AIRFLOW_DIRS_NOT_CREATED` | 500 INTERNAL_SERVER_ERROR | Airflow internal directories could not be created | | 290004 | `DAG_NOT_TEMPLATED` | 500 INTERNAL_SERVER_ERROR | Could not template DAG file | | 290005 | `AIRFLOW_MANAGER_UNINITIALIZED` | 500 INTERNAL_SERVER_ERROR | AirflowManager is not initialized | | 290006 | `DAG_NAME_INVALID` | 400 BAD_REQUEST | DAG definition failed validation | | 290007 | `AIRFLOW_GENERIC_ERROR` | 500 INTERNAL_SERVER_ERROR | Airflow internal error | | 290008 | `AIRFLOW_DISABLED` | 400 BAD_REQUEST | Airflow is disabled | ## PythonErrorCode | Code | Name | HTTP status | Message | | --- | --- | --- | --- | | 300000 | `ANACONDA_ENVIRONMENT_NOT_FOUND` | 404 NOT_FOUND | Could not find the environment. | | 300001 | `ANACONDA_ENVIRONMENT_ALREADY_INITIALIZED` | 409 CONFLICT | Anaconda environment already created for this project. | | 300002 | `PYTHON_SEARCH_TYPE_NOT_SUPPORTED` | 400 BAD_REQUEST | The supplied search is not supported, only pip and conda are currently supported. | | 300003 | `PYTHON_LIBRARY_NOT_FOUND` | 404 NOT_FOUND | Library could not be found. | | 300004 | `YML_FILE_MISSING_PYTHON_VERSION` | 400 BAD_REQUEST | No python binary version was found in the environment yaml file. | | 300005 | `NOT_MATCHING_PYTHON_VERSIONS` | 400 BAD_REQUEST | The supplied yaml files have mismatching python versions. | | 300007 | `INSTALL_TYPE_NOT_SUPPORTED` | 400 BAD_REQUEST | The provided install type is not supported | | 300008 | `CONDA_COMMAND_NOT_FOUND` | 400 BAD_REQUEST | Command not found. | | 300009 | `MACHINE_TYPE_NOT_SPECIFIED` | 400 BAD_REQUEST | Machine type not specified. | | 300010 | `VERSION_NOT_SPECIFIED` | 400 BAD_REQUEST | Version not specified. | | 300011 | `ANACONDA_ENVIRONMENT_INITIALIZING` | 400 BAD_REQUEST | The project's Python environment is currently being initialized. Please try again later. | | 300012 | `ANACONDA_ENVIRONMENT_FILE_INVALID` | 400 BAD_REQUEST | Path is not a valid environment file, must be Anaconda .yml or requirements.txt | | 300013 | `ANACONDA_PIP_CHECK_FAILED` | 500 INTERNAL_SERVER_ERROR | pip check command failed | | 300014 | `ANACONDA_ENVIRONMENT_FAILED_INITIALIZATION` | 500 INTERNAL_SERVER_ERROR | The project's Python environment failed to initialize, please recreate the environment. | | 300015 | `ANACONDA_ENVIRONMENT_REMOVAL_FAILED` | 500 INTERNAL_SERVER_ERROR | Deletion of the project's Python environment encountered an issue | | 300016 | `CONDA_COMMAND_DELETE_ERROR` | 400 BAD_REQUEST | Failed to delete a command | | 300018 | `INVALID_ENVIRONMENT_NAME` | 400 BAD_REQUEST | The name is not correct and does not match a valid environment | | 300019 | `PREINSTALLED_ENVIRONMENT_CAN_NOT_BE_DELETED` | 403 FORBIDDEN | It is not possible to delete a preinstalled environment | | 300020 | `CAN_NOT_MODIFY_BASE_ENVIRONMENT` | 403 FORBIDDEN | It is not possible to modify the base environment, create your own environment instead | | 300021 | `INCORRECT_ENVIRONMENT` | 400 BAD_REQUEST | The configured environment is not compatible. | | 300022 | `ENVIRONMENT_IN_USE` | 400 BAD_REQUEST | This environment is currently in use. | | 300023 | `INVALID_ENVIRONMENT_NAME_INPUT` | 400 BAD_REQUEST | This environment is currently in use. | | 300024 | `ENVIRONMENT_NOT_FOUND` | 404 NOT_FOUND | The environment was not found. | | 300025 | `PROJECT_ENVIRONMENT_QUOTA_REACHED` | 403 FORBIDDEN | This project has reached the maximum number of environments that can be cloned | | 300026 | `INVALID_BUILD_CACHE` | 400 BAD_REQUEST | The custom commands declared a build cache that does not exist. | | 300027 | `NPM_INSTALL_INVALID` | 400 BAD_REQUEST | The npm install request is not valid. | | 300028 | `PYTHON_LIBRARY_AMBIGUOUS` | 400 BAD_REQUEST | The library name exists in more than one package ecosystem; pass packageSource to say which one is meant. | ## ResourceErrorCode | Code | Name | HTTP status | Message | | --- | --- | --- | --- | | 310000 | `INVALID_QUERY_PARAMETER` | 404 NOT_FOUND | Invalid query. | ## ApiKeyErrorCode | Code | Name | HTTP status | Message | | --- | --- | --- | --- | | 320001 | `KEY_NOT_CREATED` | 400 BAD_REQUEST | Api key could not be created | | 320002 | `KEY_NOT_FOUND` | 401 UNAUTHORIZED | Api key not found | | 320003 | `KEY_ROLE_CONTROL_EXCEPTION` | 403 FORBIDDEN | No valid role found for this invocation | | 320004 | `KEY_SCOPE_CONTROL_EXCEPTION` | 403 FORBIDDEN | No valid scope found for this invocation | | 320005 | `KEY_SCOPE_NOT_SPECIFIED` | 400 BAD_REQUEST | Api key scope can not be empty | | 320006 | `KEY_SCOPE_EMPTY` | 400 BAD_REQUEST | Api key scope can not be empty | | 320007 | `KEY_NAME_EXIST` | 400 BAD_REQUEST | Api key name already exists | | 320008 | `KEY_NAME_NOT_SPECIFIED` | 400 BAD_REQUEST | Api key name not specified | | 320009 | `KEY_NAME_NOT_VALID` | 400 BAD_REQUEST | Api key name not valid | | 320010 | `KEY_HANDLER_CREATE_ERROR` | 500 INTERNAL_SERVER_ERROR | Error occurred during apikey create handler. | | 320011 | `KEY_HANDLER_DELETE_ERROR` | 500 INTERNAL_SERVER_ERROR | Error occurred during apikey delete handler. | | 320012 | `KEY_INVALID` | 401 UNAUTHORIZED | Invalid or incorrect API key. | | 320013 | `KEY_NOT_FOUND_IN_DATABASE` | 401 UNAUTHORIZED | API key not found in the database | | 320014 | `KEY_EXPIRED` | 401 UNAUTHORIZED | Api key has expired | | 320015 | `KEY_EXPIRY_DATE_INVALID` | 400 BAD_REQUEST | Api key expiry date is invalid | ## OpenSearchErrorCode | Code | Name | HTTP status | Message | | --- | --- | --- | --- | | 330000 | `SIGNING_KEY_ERROR` | 500 INTERNAL_SERVER_ERROR | Couldn't get or create the elk signing key | | 330001 | `JWT_NOT_CREATED` | 500 INTERNAL_SERVER_ERROR | Jwt for elk couldn't be created | | 330002 | `KIBANA_REQ_ERROR` | 400 BAD_REQUEST | Error while executing Kibana request | | 330003 | `OPENSEARCH_CONNECTION_ERROR` | 503 SERVICE_UNAVAILABLE | Couldn't connect to OpenSearch | | 330004 | `OPENSEARCH_INTERNAL_REQ_ERROR` | 500 INTERNAL_SERVER_ERROR | Error while executing OpenSearch request | | 330005 | `OPENSEARCH_QUERY_ERROR` | 400 BAD_REQUEST | Error while executing a user query on OpenSearch | | 330006 | `INVALID_OPENSEARCH_ROLE` | 500 INTERNAL_SERVER_ERROR | Invalid OpenSearch security role | | 330007 | `INVALID_OPENSEARCH_ROLE_USER` | 401 UNAUTHORIZED | Invalid OpenSearch security role for a user | | 330008 | `OPENSEARCH_QUERY_NO_MAPPING` | 400 BAD_REQUEST | OpenSearch query uses a field that is not in the mapping of the index | | 330009 | `OPENSEARCH_INDEX_NOT_FOUND` | 404 NOT_FOUND | OpenSearch index not found | ## ProvenanceErrorCode | Code | Name | HTTP status | Message | | --- | --- | --- | --- | | 340001 | `MALFORMED_ENTRY` | 500 INTERNAL_SERVER_ERROR | Provenance entry is malformed | | 340002 | `BAD_REQUEST` | 500 INTERNAL_SERVER_ERROR | Provenance query request is malformed | | 340003 | `UNSUPPORTED` | 400 BAD_REQUEST | Provenance query is not supported | | 340004 | `INTERNAL_ERROR` | 500 INTERNAL_SERVER_ERROR | Provenance logical error | | 340005 | `ARCHIVAL_STORE` | 500 INTERNAL_SERVER_ERROR | Provenance archival store error | | 340006 | `FS_ERROR` | 500 INTERNAL_SERVER_ERROR | Provenance xattr - file system error | ## ModelRegistryErrorCode | Code | Name | HTTP status | Message | | --- | --- | --- | --- | | 360000 | `MODEL_NOT_FOUND` | 404 NOT_FOUND | No model found for provided name and version. | | 360001 | `KEY_NOT_STRING` | 404 NOT_FOUND | metrics key is not a string. | | 360002 | `METRIC_NOT_NUMBER` | 400 BAD_REQUEST | Could not cast provided metric to double. | | 360003 | `MODEL_LIST_FAILED` | 500 INTERNAL_SERVER_ERROR | Error occurred when fetching models. | | 360004 | `MODEL_MARSHALLING_FAILED` | 500 INTERNAL_SERVER_ERROR | Error occurred during marshalling/unmarshalling of model json. | | 360005 | `MODEL_REGISTRY_ID_NOT_PROVIDED` | 400 BAD_REQUEST | Model Registry Id was not provided. | | 360006 | `MODEL_REGISTRY_ID_NOT_FOUND` | 400 BAD_REQUEST | Model Registry Id was not found. | | 360007 | `MODEL_REGISTRY_ACCESS_DENIED` | 403 FORBIDDEN | Model Registry not accessible. | | 360008 | `MODEL_REGISTRY_MODELS_DATASET_NOT_FOUND` | 500 INTERNAL_SERVER_ERROR | Models dataset does not exist in project. | | 360009 | `MODEL_CANNOT_BE_DELETED` | 400 BAD_REQUEST | Could not delete the model | | 360010 | `HUGGINGFACE_AUTH_REQUIRED` | 401 UNAUTHORIZED | HuggingFace access token is required for this model. Please provide a valid token. | | 360011 | `HUGGINGFACE_MODEL_NOT_FOUND` | 404 NOT_FOUND | HuggingFace model not found. Please verify the model ID. | | 360012 | `HUGGINGFACE_DOWNLOAD_FAILED` | 500 INTERNAL_SERVER_ERROR | Failed to download model from HuggingFace. | | 360013 | `HUGGINGFACE_INVALID_REQUEST` | 400 BAD_REQUEST | Invalid HuggingFace import request. | ## SchematizedTagErrorCode | Code | Name | HTTP status | Message | | --- | --- | --- | --- | | 370000 | `TAG_SCHEMA_NOT_FOUND` | 404 NOT_FOUND | No schema found for provided name | | 370001 | `INVALID_TAG_SCHEMA` | 400 BAD_REQUEST | Invalid tag schema. | | 370002 | `TAG_NOT_FOUND` | 404 NOT_FOUND | No tag found for provided name. | | 370003 | `TAG_ALREADY_EXISTS` | 409 CONFLICT | Tag with the same name already exists. | | 370004 | `INVALID_TAG_NAME` | 400 BAD_REQUEST | Invalid tag name. | | 370005 | `INVALID_TAG_VALUE` | 400 BAD_REQUEST | Invalid tag value. | | 370006 | `TAG_NOT_ALLOWED` | 400 BAD_REQUEST | The provided tag is not allowed | | 370007 | `INTERNAL_PROCESSING_ERROR` | 500 INTERNAL_SERVER_ERROR | Internal error while processing tag | | 370008 | `INVALID_MANDATORY_TAG` | 400 BAD_REQUEST | Invalid mandatory tag | | 370009 | `TAG_MIGRATION_ONGOING` | 503 SERVICE_UNAVAILABLE | Tag migration in progress, wait for it to finish before issuing tag operations. | | 370010 | `TAG_ARCHIVE_TOO_MANY_ATTACHMENTS` | 400 BAD_REQUEST | The schema is attached to more artifacts than one archive transaction can cover. | ## CloudErrorCode | Code | Name | HTTP status | Message | | --- | --- | --- | --- | | 380000 | `CLOUD_FEATURE` | 405 METHOD_NOT_ALLOWED | This method is only available in cloud deployments. | | 380001 | `FAILED_TO_ASSUME_ROLE` | 400 BAD_REQUEST | Failed to assume role. | | 380002 | `ACCESS_CONTROL_EXCEPTION` | 403 FORBIDDEN | You are not allowed to assume this role. | | 380003 | `MAPPING_NOT_FOUND` | 400 BAD_REQUEST | Mapping not found. | | 380004 | `MAPPING_ALREADY_EXISTS` | 400 BAD_REQUEST | Mapping for the given project and role already exists. | | 380005 | `FAILED_TO_GET_CLUSTER_CRED` | 400 BAD_REQUEST | Failed to get cluster credential. | ## AlertErrorCode | Code | Name | HTTP status | Message | | --- | --- | --- | --- | | 390000 | `ALERT_CREATION_FAILED` | 400 BAD_REQUEST | Failed to create alert | | 390001 | `RECEIVER_EXIST` | 400 BAD_REQUEST | A receiver already exists. | | 390002 | `ROUTE_EXIST` | 400 BAD_REQUEST | A route already exists. | | 390003 | `RECEIVER_NOT_FOUND` | 400 BAD_REQUEST | Receiver not found. | | 390004 | `ROUTE_NOT_FOUND` | 400 BAD_REQUEST | Route not found. | | 390005 | `SILENCE_NOT_FOUND` | 400 BAD_REQUEST | Silence not found. | | 390006 | `ILLEGAL_ARGUMENT` | 400 BAD_REQUEST | Illegal argument. | | 390007 | `FAILED_TO_CONNECT` | 412 PRECONDITION_FAILED | Failed to connect to alert manager. | | 390008 | `FAILED_TO_UPDATE_AM_CONFIG` | 412 PRECONDITION_FAILED | Failed to update alert manager configuration. | | 390009 | `RESPONSE_ERROR` | 400 BAD_REQUEST | Alert manager response error. | | 390010 | `ACCESS_CONTROL_EXCEPTION` | 403 FORBIDDEN | You are not allowed to access this resource. | | 390011 | `FAILED_TO_READ_CONFIGURATION` | 412 PRECONDITION_FAILED | Failed to read alert manager configuration. | | 390012 | `FAILED_TO_CLEAN` | 412 PRECONDITION_FAILED | Failed to clean project from alert manager config. | | 390013 | `AM_CONFIG_NOT_UPDATED` | 409 CONFLICT | Alert manager config not updated. | ## RemoteAuthErrorCode | Code | Name | HTTP status | Message | | --- | --- | --- | --- | | 400000 | `NOT_FOUND` | 404 NOT_FOUND | Not found. | | 400001 | `DUPLICATE_ENTRY` | 400 BAD_REQUEST | Duplicate entry. | | 400002 | `ILLEGAL_ARGUMENT` | 400 BAD_REQUEST | Illegal argument. | | 400003 | `WRONG_CONFIG` | 412 PRECONDITION_FAILED | Wrong configuration. | | 400004 | `TOKEN_PARSE_EXCEPTION` | 417 EXPECTATION_FAILED | Token ParseException. | | 400005 | `NOT_ALLOWED` | 400 BAD_REQUEST | Operation not allowed. | ## CommandErrorCode | Code | Name | HTTP status | Message | | --- | --- | --- | --- | | 410000 | `INTERNAL_SERVER_ERROR` | 500 INTERNAL_SERVER_ERROR | Something went wrong executing command | | 410001 | `INVALID_SQL_QUERY` | 400 BAD_REQUEST | Invalid sql query for command | | 410002 | `ILLEGAL_ARGUMENT` | 400 BAD_REQUEST | Illegal argument in command | | 410003 | `FILESYSTEM_ACCESS_ERROR` | 500 INTERNAL_SERVER_ERROR | Error accessing files system for components | | 410004 | `FEATURESTORE_ACCESS_ERROR` | 500 INTERNAL_SERVER_ERROR | Error accessing the featurestore | | 410005 | `OPENSEARCH_ACCESS_ERROR` | 500 INTERNAL_SERVER_ERROR | Error acccessing OpenSearch | | 410006 | `SERIALIZATION_ERROR` | 500 INTERNAL_SERVER_ERROR | Error serializing content | | 410007 | `NOT_IMPLEMENTED` | 501 NOT_IMPLEMENTED | Internal error dealing with new artifact type | | 410008 | `DB_QUERY_ERROR` | 500 INTERNAL_SERVER_ERROR | DB error on query | | 410009 | `ARTIFACT_DELETED` | 500 INTERNAL_SERVER_ERROR | Artifact was deleted before command could be executed | ## GitOpErrorCode | Code | Name | HTTP status | Message | | --- | --- | --- | --- | | 500000 | `SIGNING_KEY_ERROR` | 500 INTERNAL_SERVER_ERROR | Couldn't get or create the GIT signing key. | | 500001 | `JWT_NOT_CREATED` | 500 INTERNAL_SERVER_ERROR | Jwt for GIT could not be created. | | 500003 | `INVALID_GIT_ROLE_USER` | 401 UNAUTHORIZED | Invalid git security role for a user. | | 500004 | `JWT_MATERIALIZATION_ERROR` | 500 INTERNAL_SERVER_ERROR | Could not materialize jwt. | | 500005 | `GIT_PATHS_CREATION_ERROR` | 500 INTERNAL_SERVER_ERROR | Could not create git paths. | | 500006 | `REPOSITORY_URL_NOT_PROVIDED` | 400 BAD_REQUEST | Repository url not provided. | | 500007 | `DIRECTORY_PATH_NOT_PROVIDED` | 400 BAD_REQUEST | Path to directory not provided. | | 500008 | `DIRECTORY_PATH_DOES_NOT_EXIST` | 400 BAD_REQUEST | The directory does not exist. | | 500009 | `PATH_IS_NOT_DIRECTORY` | 400 BAD_REQUEST | Path is not a directory. | | 500010 | `GIT_HOME_ERROR` | 500 INTERNAL_SERVER_ERROR | Could not resolve GIT_HOME using DB. | | 500012 | `GIT_CONTAINER_LAUNCH_ERROR` | 500 INTERNAL_SERVER_ERROR | Could not launch the git container. | | 500013 | `EXECUTION_OBJECT_NOT_FOUND` | 404 NOT_FOUND | Execution object with the id not found. | | 500014 | `INVALID_AUTHENTICATION_METHOD` | 400 BAD_REQUEST | Unknown authentication method. | | 500015 | `INVALID_GITHUB_USERNAME` | 400 BAD_REQUEST | Invalid git username. | | 500016 | `USER_DOES_NOT_HAVE_PERMISSIONS_TO_GIT_DIR` | 400 BAD_REQUEST | Git directory security error. | | 500017 | `COMMIT_MESSAGE_IS_EMPTY` | 400 BAD_REQUEST | Commit command message should not be empty. | | 500018 | `INVALID_BRANCH_NAME` | 400 BAD_REQUEST | Branch name should not be empty. | | 500019 | `REPOSITORY_NOT_FOUND` | 404 NOT_FOUND | Repository not found. | | 500020 | `GIT_PROVIDER_NOT_PROVIDED` | 400 BAD_REQUEST | Git provider not provided. | | 500021 | `INVALID_REPOSITORY_URL` | 400 BAD_REQUEST | Invalid repository url provided | | 500022 | `DIRECTORY_IS_ALREADY_GIT_REPO` | 400 BAD_REQUEST | Directory is already a git repository | | 500023 | `INVALID_BRANCH_ACTION` | 400 BAD_REQUEST | Invalid branch action provided. | | 500024 | `INVALID_REMOTES_ACTION` | 400 BAD_REQUEST | Invalid remotes action provided | | 500025 | `INVALID_REMOTE_NAME` | 400 BAD_REQUEST | Invalid remote name provided. Remote name should not be empty. | | 500026 | `INVALID_REMOTE_URL_PROVIDED` | 400 BAD_REQUEST | Invalid remote url provided. Remote url should not be empty. | | 500027 | `GIT_OPERATION_ERROR` | 500 INTERNAL_SERVER_ERROR | Git operation error. | | 500028 | `GIT_REPOSITORIES_NOT_FOUND` | 404 NOT_FOUND | No git repository found in project | | 500029 | `GIT_USERNAME_AND_PASSWORD_NOT_SET` | 400 BAD_REQUEST | Git username and password not set | | 500030 | `COMMIT_FILES_EMPTY` | 400 BAD_REQUEST | Files to add and commit is empty. | | 500031 | `INVALID_REPOSITORY_ACTION` | 400 BAD_REQUEST | Invalid repository action. | | 500032 | `REMOTE_NOT_FOUND` | 404 NOT_FOUND | Git remote not found. | | 500033 | `INVALID_BRANCH_AND_COMMIT_CHECKOUT_COMBINATION` | 400 BAD_REQUEST | Branch and Hash are mutually exclusive. | | 500034 | `INVALID_GIT_COMMAND_CONFIGURATION` | 500 INTERNAL_SERVER_ERROR | Invalid git command operation | | 500035 | `USER_IS_NOT_REPOSITORY_OWNER` | 403 FORBIDDEN | User not allowed to perform operation in repository | | 500036 | `READ_ONLY_REPOSITORY` | 400 BAD_REQUEST | Repository is read only | | 500037 | `ERROR_VALIDATING_REPOSITORY_PATH` | 500 INTERNAL_SERVER_ERROR | Error validating git repository path | | 500038 | `ERROR_CANCELLING_GIT_EXECUTION` | 400 BAD_REQUEST | Failed to cancel git execution | | 500039 | `INVALID_HOSTNAME` | 400 BAD_REQUEST | Hostname provided is not valid | | 500040 | `TOKEN_NOT_PROVIDED` | 400 BAD_REQUEST | Token not provided | | 500041 | `DUPLICATE_HOSTS` | 400 BAD_REQUEST | Duplicate hosts provided | ## KubeErrorCode | Code | Name | HTTP status | Message | | --- | --- | --- | --- | | 510000 | `INVALID_INPUT` | 400 BAD_REQUEST | The request contains parameters not valid | | 510001 | `INTERNAL_ERROR_MISSING` | 500 INTERNAL_SERVER_ERROR | Internal error - missing | | 510002 | `LOCAL_QUEUE_ALREADY_EXISTS` | 400 BAD_REQUEST | Local queue already exists | | 510003 | `LOCAL_QUEUE_CREATION_FAILED` | 500 INTERNAL_SERVER_ERROR | Local queue creation failed | | 510004 | `ERROR_FETCHING_QUEUE` | 500 INTERNAL_SERVER_ERROR | Error fetching configured queue | ## PlatformIntelligenceErrorCode | Code | Name | HTTP status | Message | | --- | --- | --- | --- | | 520012 | `LLM_NOT_CONFIGURED` | 400 BAD_REQUEST | LLM is not configured. | | 520013 | `METADATA_INFERENCE_FAILED` | 500 INTERNAL_SERVER_ERROR | Metadata inference failed. | | 520014 | `VLLM_CONFIG_GENERATION_FAILED` | 500 INTERNAL_SERVER_ERROR | vLLM config generation failed. | | 520015 | `DEPLOYMENT_GENERATION_FAILED` | 500 INTERNAL_SERVER_ERROR | Deployment generation failed. | ## FeatureStoreMetricsErrorCode | Code | Name | HTTP status | Message | | --- | --- | --- | --- | | 530000 | `METRIC_DOES_NOT_EXIST` | 404 NOT_FOUND | Metric does not exist. | | 530001 | `FEATURE_STORE_REQUIRED_FOR_METRIC` | 400 BAD_REQUEST | Feature store is required for metric. | | 530002 | `EVENT_NOT_FOUND` | 404 NOT_FOUND | Event not found. | ## TrinoErrorCode | Code | Name | HTTP status | Message | | --- | --- | --- | --- | | 540000 | `AUTHENTICATION_ERROR` | 417 EXPECTATION_FAILED | Trino authentication error | | 540001 | `CONNECTION_ERROR` | 503 SERVICE_UNAVAILABLE | Could not connect to Trino server | | 540002 | `QUERY_EXECUTION_ERROR` | 400 BAD_REQUEST | Error executing Trino query | | 540003 | `TRINO_NOT_ENABLED` | 400 BAD_REQUEST | Trino is disabled | | 540004 | `ILLEGAL_ARGUMENT` | 400 BAD_REQUEST | Illegal argument provided | | 540005 | `CATALOG_INVALID_NAME` | 400 BAD_REQUEST | Invalid Trino catalog name | | 540006 | `CATALOG_ALREADY_EXISTS` | 409 CONFLICT | A Trino catalog with this name already exists | | 540007 | `CATALOG_NOT_FOUND` | 404 NOT_FOUND | Trino catalog not found | | 540008 | `CATALOG_INVALID_CONFIG` | 400 BAD_REQUEST | Invalid Trino catalog configuration | | 540009 | `CATALOG_VALIDATION_FAILED` | 400 BAD_REQUEST | Trino catalog connectivity validation failed | | 540010 | `CATALOG_STORAGE_ERROR` | 500 INTERNAL_SERVER_ERROR | Failed to read or write the Trino catalog storage | | 540011 | `CATALOG_RESTART_ERROR` | 500 INTERNAL_SERVER_ERROR | Failed to restart Trino | | 540012 | `CATALOG_SECRET_ERROR` | 400 BAD_REQUEST | Could not resolve a Hopsworks secret referenced by the Trino catalog | | 540013 | `CATALOG_TEST_UNAVAILABLE` | 503 SERVICE_UNAVAILABLE | The Trino test coordinator is not deployed; connection testing is unavailable | | 540014 | `TRINO_RESTARTING` | 503 SERVICE_UNAVAILABLE | The Trino query engine is restarting; try again once it is back up | | 540015 | `CATALOG_OPERATION_IN_PROGRESS` | 503 SERVICE_UNAVAILABLE | Another Trino catalog operation is in progress; try again shortly | | 540016 | `CATALOG_LIMIT_EXCEEDED` | 400 BAD_REQUEST | The Trino catalog exceeds a configured limit | | 540017 | `CATALOG_MOUNT_UNAVAILABLE` | 503 SERVICE_UNAVAILABLE | The Trino pods do not mount the credential-file store | | 540018 | `CATALOG_PROJECT_NOT_READY` | 503 SERVICE_UNAVAILABLE | A previous project of this name is still being deleted; catalogs cannot be created until that finishes | | 540019 | `DATASOURCE_NOT_MAPPABLE` | 400 BAD_REQUEST | This data source cannot be mapped to a Trino catalog | ## AiProviderErrorCode | Code | Name | HTTP status | Message | | --- | --- | --- | --- | | 550000 | `NOT_FOUND` | 404 NOT_FOUND | AI provider not found | | 550001 | `VALIDATION` | 400 BAD_REQUEST | AI provider validation error | | 550002 | `SECRET_ERROR` | 500 INTERNAL_SERVER_ERROR | AI provider secret operation failed | ## SupersetErrorCode | Code | Name | HTTP status | Message | | --- | --- | --- | --- | | 560000 | `AUTHENTICATION_ERROR` | 417 EXPECTATION_FAILED | Superset authentication error | | 560001 | `CONNECTION_ERROR` | 503 SERVICE_UNAVAILABLE | Could not connect to Superset server | | 560002 | `API_REQUEST_ERROR` | 400 BAD_REQUEST | Error executing Superset API request | | 560003 | `SUPERSET_DISABLED` | 400 BAD_REQUEST | Superset is disabled | | 560004 | `FORBIDDEN` | 403 FORBIDDEN | Forbidden from accessing Superset resource | | 560005 | `INVALID_PARAMETER` | 400 BAD_REQUEST | Invalid parameter provided for Superset API request | ## ProxyErrorCode | Code | Name | HTTP status | Message | | --- | --- | --- | --- | | 570000 | `PROXY_UNAUTHORIZED` | 401 UNAUTHORIZED | Proxy authentication failed | | 570001 | `PROXY_FORBIDDEN` | 403 FORBIDDEN | Not authorized to access this proxied resource | | 570002 | `PROXY_UPSTREAM_NOT_FOUND` | 404 NOT_FOUND | Upstream proxied service not found | | 570003 | `PROXY_UPSTREAM_ERROR` | 502 BAD_GATEWAY | Error forwarding request to upstream proxied service | ## MountableSecretErrorCode | Code | Name | HTTP status | Message | | --- | --- | --- | --- | | 580000 | `MOUNTABLE_SECRETS_NOT_ENABLED` | 503 SERVICE_UNAVAILABLE | Mountable secrets are not available on this cluster | | 580001 | `MOUNTABLE_SECRET_NOT_FOUND` | 404 NOT_FOUND | Mountable secret not found | | 580002 | `MOUNTABLE_SECRET_FILE_NOT_FOUND` | 404 NOT_FOUND | File not found in the mountable secret | | 580003 | `MOUNTABLE_SECRET_ALREADY_EXISTS` | 409 CONFLICT | A mountable secret with this name already exists | | 580004 | `MOUNTABLE_SECRET_INVALID_NAME` | 400 BAD_REQUEST | Invalid mountable secret or file name | | 580005 | `MOUNTABLE_SECRET_LIMIT_EXCEEDED` | 400 BAD_REQUEST | The mountable secret exceeds a configured limit | | 580006 | `MOUNTABLE_SECRET_STORAGE_ERROR` | 500 INTERNAL_SERVER_ERROR | Failed to read or write the mountable secret store | | 580007 | `ILLEGAL_ARGUMENT` | 400 BAD_REQUEST | Illegal argument provided | | 580008 | `MOUNTABLE_SECRET_INVALID_ARCHIVE` | 400 BAD_REQUEST | The uploaded archive could not be expanded into a mountable secret | | 580009 | `MOUNTABLE_SECRETS_STORE_MISSING` | 503 SERVICE_UNAVAILABLE | The mountable secret store does not exist on this cluster | | 580010 | `MOUNTABLE_SECRET_PROJECT_DELETED` | 410 GONE | The project was deleted while the mountable secret was being uploaded | ## SchemaRegistryErrorCode This category does not follow the 6-digit convention described at the top of this page; it mirrors the Confluent Schema Registry's own error codes instead. | Code | Name | HTTP status | Message | | --- | --- | --- | --- | | 40401 | `SUBJECT_NOT_FOUND` | 404 NOT_FOUND | Subject not found | | 40402 | `VERSION_NOT_FOUND` | 404 NOT_FOUND | Version not found | | 40403 | `SCHEMA_NOT_FOUND` | 404 NOT_FOUND | Schema not found | | 40901 | `INCOMPATIBLE_AVRO_SCHEMA` | 409 CONFLICT | Incompatible Avro schema | | 42201 | `INVALID_AVRO_SCHEMA` | 422 UNPROCESSABLE_ENTITY | Invalid Avro schema | | 42202 | `INVALID_VERSION` | 422 UNPROCESSABLE_ENTITY | Invalid version | | 42203 | `INVALID_COMPATIBILITY` | 422 UNPROCESSABLE_ENTITY | Invalid compatibility level | | 50001 | `INTERNAL_SERVER_ERROR` | 500 INTERNAL_SERVER_ERROR | Error in the backend datastore | | 50002 | `OPERATION_TIMED_OUT` | 500 INTERNAL_SERVER_ERROR | Operation timed out | | 50003 | `ERROR_FORWARDING_REQUEST` | 500 INTERNAL_SERVER_ERROR | Error while forwarding the request to the primary | ================================================================================ # Docs for AI agents Source: https://docs.hopsworks.ai/latest/ai/ # Docs for AI agents The Hopsworks documentation is published in machine-readable form so agents and LLM tools can consume it directly. There are two ways to use it: a live MCP server, and a set of static text artifacts. ## MCP server `mcp.hopsworks.ai` is a read-only [Model Context Protocol](https://modelcontextprotocol.io) server over this documentation. It indexes the same Markdown that builds this site and exposes retrieval tools to any MCP-capable client. It is read-only: there is no write, create, or delete path, and it makes no outbound network calls. The endpoint is rate-limited per client IP. The tools it exposes: | Tool | Purpose | | ---- | ------- | | `search_docs(query, limit)` | Full-text (BM25) search across all pages; returns page ids, URLs, and snippets. | | `get_page(page_id)` | Full Markdown of a page by its canonical id. | | `list_sections(page_id)` | Heading structure and anchors of a page. | | `get_section(page_id, anchor)` | One section of a page by anchor. | | `list_pages(prefix)` | Browse the doc map, optionally scoped by a path prefix. | A `page_id` is the path of a page without the `.md` extension, for example `concepts/fs/feature_group/fg_overview`. ### Connect from Claude Code Add the server at user scope so it is available in every project: ```bash claude mcp add --transport http hopsworks-docs -s user https://mcp.hopsworks.ai/mcp ``` ### Connect from Claude Desktop Add the server to `claude_desktop_config.json`: ```json { "mcpServers": { "hopsworks-docs": { "type": "http", "url": "https://mcp.hopsworks.ai/mcp" } } } ``` ### Any other MCP client Point the client at the streamable-HTTP endpoint `https://mcp.hopsworks.ai/mcp`. Clients that only speak stdio can run the server locally instead: see the `mcp-server/` directory in the [documentation repository](https://github.com/logicalclocks/logicalclocks.github.io) for the self-host command. ## Text artifacts Every build also emits static files, so an agent can ingest the docs without an MCP client. - [`llms.txt`](https://docs.hopsworks.ai/llms.txt) is a curated index of the documentation that mirrors the site navigation, following the [llmstxt.org](https://llmstxt.org) convention. - [`llms-full.txt`](https://docs.hopsworks.ai/llms-full.txt) is the full-text Markdown corpus of every page, for bulk ingestion. - Every page has a raw Markdown sibling: append `.md` to any page URL, for example `https://docs.hopsworks.ai/concepts/fti_pipelines.md`. ## Copy for LLM Every page carries a **Copy for LLM** button at the top. It copies the page's raw Markdown to your clipboard, ready to paste into a chat or prompt.