# Hopsworks Documentation
> Official documentation for Hopsworks and its Feature Store - an open source data-intensive AI platform used for the development and operation of machine learning models at scale.
================================================================================
# Home
Source: https://docs.hopsworks.ai/latest/
Hopsworks Documentation
# Build, deploy & maintain AI systems
Features, training data, models and inference on one governed platform.
## Your first feature vector, in three steps
Install the client, connect to a project with an [API key](user_guides/projects/api_key/create_api_key.md), write a feature group and read a feature vector back.
=== "Python"
```bash
uv venv && source .venv/bin/activate
uv pip install "hopsworks[python]"
hops setup # opens a browser, picks a project, caches the key
python # opens the interpreter, the Python lines below go there or in a notebook
```
```python
import hopsworks
project = hopsworks.login() # uses the key hops setup cached, else prompts
fs = project.get_feature_store()
```
=== "CLI"
```bash
uv venv && source .venv/bin/activate
uv pip install "hopsworks[python]"
hops setup # opens a browser, picks a project, caches the key
hops fg list
```
Next: [create a feature group](user_guides/fs/feature_group/create.md),
[create a feature view](user_guides/fs/feature_view/overview.md),
[retrieve feature vectors](user_guides/fs/feature_view/feature-vectors.md),
or browse the Python API.
## Where Hopsworks runs
- :material-cloud-outline: **Use the managed SaaS**
---
Sign in to the Hopsworks serverless app and create a project.
Nothing to install, free tier available.
[Open run.hopsworks.ai ↗](https://run.hopsworks.ai)
- :material-server: **Deploy on your cloud or on-prem**
---
Managed Kubernetes on AWS, Azure or GCP, or an air-gapped data centre.
Talk to us to size and install it.
[Contact Hopsworks ↗](https://www.hopsworks.ai/contact) ·
[Deployment options](setup_installation/index.md)
## One architecture, three pipelines
Independent [feature, training and inference pipelines](concepts/fti.md), connected by a shared feature store and model registry.
--8<-- "index/one-architecture-three-pipelines.html"
## Find your path
:material-forum-outline:{ .hops-colophon-ico } Community and source
{ .hops-colophon-cap }
- [Public Slack ↗](https://join.slack.com/t/public-hopsworks/shared_invite/zt-24fc3hhyq-VBEiN8UZlKsDrrLvtU4NaA)
- [hopsworks-api on GitHub ↗](https://github.com/logicalclocks/hopsworks-api)
- [Apache License 2.0 ↗](https://www.apache.org/licenses/LICENSE-2.0.html)
================================================================================
# Concepts
Source: https://docs.hopsworks.ai/latest/concepts/
# Concepts
This section explains what Hopsworks is and why it is built the way it is.
It is reference and explanation, not step-by-step instructions.
For the how-to, see the [Guides](../user_guides/index.md).
## Start here
Read the [FTI Pipeline Architecture](fti.md) first.
It is the one idea the rest of this section builds on: every AI system decomposes into feature, training, and inference pipelines, connected through a feature store and a model registry.
Once you have that model, the other pages are the parts of it.
## Reading path
- [Hopsworks Platform](hopsworks.md): the components of the platform and how they fit together.
- [FTI Pipeline Architecture](fti.md): the architecture all AI systems share, and the four classes of AI system.
- **Feature Store**: how feature pipelines write feature data ([feature groups](fs/feature_group/fg_overview.md)) and how training and inference pipelines read it ([feature views](fs/feature_view/fv_overview.md)).
- **Projects**: the multi-tenant unit that owns your data and ML assets, with governance, sharing, and lineage.
- **MLOps**: training, the model registry, serving, and monitoring, the inference side of an AI system.
- **Development**: building and running pipelines inside and outside Hopsworks.
## How the section is organised
The Feature Store pages follow the write path then the read path: you write features to feature groups, and you read them through feature views.
The MLOps pages follow a model from training through registration, serving, and monitoring.
Projects and Development cut across both.
================================================================================
# Hopsworks Platform
Source: https://docs.hopsworks.ai/latest/concepts/hopsworks/
# The Hopsworks Platform
Hopsworks is a **modular** MLOps platform with:
- a feature store (available as standalone)
- model registry and model serving based on KServe
- vector index based on OpenSearch
- a data science and data engineering platform
MLOps is a set of best practices for the automated testing, versioning, and monitoring of the ML pipelines and ML assets that power AI systems.
Hopsworks is modular, so you can adopt the feature store on its own or use the full platform across the MLOps lifecycle.
--8<-- "concepts/hopsworks/the-hopsworks-platform.html"
## Standalone Feature Store
Hopsworks was the first open-source and first enterprise feature store for ML. You can use Hopsworks as a standalone feature store with the Hopsworks API.
## Model Management
Hopsworks includes support for model management, with model deployments using [the KServe framework](https://github.com/kserve/kserve) and a model registry designed for KServe.
Hopsworks logs all inference requests to Kafka to enable easy monitoring of deployed models, and provides model metrics with grafana/prometheus.
## Vector Index
A feature group with an embedding column can have a vector index, based on [OpenSearch kNN](https://opensearch.org/docs/latest/search-plugins/knn/index/) (on the [FAISS](https://ai.facebook.com/tools/faiss/) engine, which has replaced the deprecated [nmslib](https://github.com/nmslib/nmslib) engine as the default).
The vector index includes out-of-the-box support for authentication, access control, filtering, backup-and-restore, and horizontal scalability.
The Feature Store and its vector index are often used together to build scalable recommender systems, such as ranking-and-retrieval for real-time recommendations.
## Governance
Hopsworks provides a data-mesh architecture for managing ML assets and teams, with multi-tenant projects.
Not unlike a GitHub repository, a project is a sandbox containing team members, data, and ML assets.
In Hopsworks, all ML assets (features, models, training data) are versioned, taggable, lineage-tracked, and support free-text search.
Data can be also be securely shared between projects.
## Data Science Platform
You can develop feature engineering, model training and inference pipelines in Hopsworks.
There is support for version control (GitHub, GitLab, BitBucket), Jupyter notebooks, a shared distributed file system, many bundled modular project Python environments for managing Python dependencies without needing to write Dockerfiles, jobs (Python, Spark), and workflow orchestration with Airflow.
================================================================================
# FTI Pipeline Architecture
Source: https://docs.hopsworks.ai/latest/concepts/fti/
# FTI Pipeline Architecture
Hopsworks is built around a single architecture for AI systems: the decomposition of any AI system into **feature**, **training**, and **inference** (FTI) pipelines.
This page defines that architecture.
Every other concept in this section is a part of it, so read this first.
## The three pipelines
An AI system decomposes naturally into three machine learning pipelines, each with clear inputs and outputs, each developed, tested, and operated independently.
- A **feature pipeline** takes data as input and produces reusable feature data as output.
- A **training pipeline** takes feature data as input, trains a model, and outputs the trained model.
- An **inference pipeline** takes feature data and a model as input and outputs predictions and prediction logs.
The three pipelines are independent programs.
They are composed into a working system through a shared data layer: a [feature store](fs/index.md) and a [model registry](mlops/registry.md).
--8<-- "concepts/fti/the-three-pipelines.html"
Feature pipelines ingest both backfill and production data and compute feature data that is stored as tabular data in the feature store.
Feature pipelines can be batch programs or stream processing programs.
Training pipelines read training data from the feature store and store the models they produce in the model registry.
Inference pipelines output predictions using a model, either downloaded from the model registry or served behind an API, together with new feature data that is precomputed in the feature store or computed from data available at prediction request time.
## Why this architecture
The five common AI system architectures (batch, stateless real-time, stateful real-time, RAG, and agentic) are very different from one another.
Moving from one to another, or transferring what you learned building one, is hard.
The FTI decomposition gives you one architecture for all of them.
Modularity is the reason.
Splitting an AI system into independent, small, testable modules lets teams build higher-quality systems faster.
It also splits the work cleanly: feature engineering can involve data engineers, model training is the realm of data scientists, and inference can involve operations.
## The shared data layer
The feature store holds three stores of feature data, each serving a different pipeline.
- A row-oriented online store for low latency access from online inference pipelines and agents.
- A columnar offline store for training models and batch inference.
- A [vector index](mlops/opensearch.md) over embeddings for inference pipelines and agents.
The model registry holds the trained models and their assets, versioned, for inference pipelines to load.
## What an AI system is
An AI system is a set of independent feature pipelines, training pipelines, and inference pipelines connected through a feature store and a model registry.
An AI system is defined by how it computes its predictions, not by the type of application that consumes them.
On that basis, AI systems built with a feature store fall into four classes.
- **Real-time (interactive)** systems make predictions in response to user requests.
They read precomputed features from the feature store and can also compute features on demand from the request parameters.
- **Agentic workflows** achieve goals with some autonomy using LLMs and tools, drawing context from a vector index, the online and offline stores, and external APIs.
- **Batch** systems run inference on a schedule and write predictions to a downstream store, called an inference store, for an application to consume later.
- **Stream processing** systems use an embedded model to make predictions on streaming data without user input, often machine to machine.
The inference pipeline is what determines the class.
When you know how a system computes its predictions, you know which of these you are building.
## Where to go next
- [Feature Store Architecture](fs/index.md) for the shared data layer in detail.
- [Feature Groups](fs/feature_group/fg_overview.md) for how feature pipelines write feature data.
- [Feature Views](fs/feature_view/fv_overview.md) for how training and inference pipelines read it.
- [AI Systems](mlops/prediction_services.md) for the inference side and how each class is served.
================================================================================
# AI Systems
Source: https://docs.hopsworks.ai/latest/concepts/mlops/prediction_services/
# AI Systems
An AI system is a set of independent feature pipelines, training pipelines, and inference pipelines that are connected via a feature store and model registry.
Each pipeline is a separate program with its own inputs and outputs, and the shared data layer is what lets them be developed, run, and scaled independently.
An AI system is defined by how it computes its predictions, not by the type of application that consumes them.
The inference pipeline determines the class of AI system you are building.
There are four classes:
- **Real-time (interactive)**: a client sends a prediction request and an online inference pipeline computes and returns a prediction with low latency.
- **Batch**: an inference pipeline runs on a schedule, scores a set of entities, and writes the predictions to an inference store.
- **Stream processing**: an inference pipeline computes predictions continuously over an event stream.
- **Agentic workflows**: an LLM-driven control flow decides which steps to run, retrieving the context and features it needs from the feature store. See [Agents and LLM Systems](agents.md).
Whatever the class, an AI system is composed of the same parts:
- one or more feature pipeline(s) that keep the feature store up to date,
- a training pipeline that produces a model in the model registry,
- an inference pipeline that reads features and computes predictions,
- a sink for the predictions, either an inference store or a user interface.
The two figures below illustrate the two most common classes, batch and real-time.
## Batch AI systems
In the figure below, feature pipelines update the feature store with new feature data on a schedule (e.g., hourly, daily).
A batch inference pipeline also runs on a schedule, reads batch scoring data from the feature store, computes predictions with an embedded model, and writes those predictions to an inference store.
The inference store is any data store that holds the predictions from batch inference pipelines.
From there, the predictions are consumed by (predictive, prescriptive) analytical reports and/or to AI-enable operational services.
--8<-- "concepts/mlops/prediction_services/batch-ai-systems.html"
## Real-time AI systems
In the figure below, feature pipelines update the feature store with new feature data on a schedule (e.g., streaming, hourly, daily), and the operational service sends prediction requests to a model deployed on KServe via its secured Istio endpoint.
A deployed model on KServe handles the prediction request by first retrieving pre-computed features from the feature store for the given request, and then building a feature vector that is scored by the model.
The prediction result is returned to the client (the operational service).
KServe logs both the feature values and the prediction results back to Hopsworks for further analysis and to help create new training data.
--8<-- "concepts/mlops/prediction_services/real-time-ai-systems.html"
## MLOps Flywheel
Once you have built your batch or real-time AI system, the MLOps flywheel is the path to building a self-managing system that automatically collects and processes feature logs, prediction logs, and outcomes to help create new training data for models.
This enables a ML flywheel where new training data and insights are generated from your AI system, by feeding logs back into the feature store.
More training data enables the training of better models, and with better models, you should hopefully improve your operational/batch services, so that you attract more clients, who in turn produce more data for training models.
And, thus, the ML flywheel is bootstrapped and leads to a virtuous cycle of more data leading to better models and more models leading to more users, who produce more data, and so on.
=== "Offline path: the training loop"
--8<-- "concepts/mlops/prediction_services/mlops-flywheel-offline.html"
=== "Online path: the serving loop"
--8<-- "concepts/mlops/prediction_services/mlops-flywheel-online.html"
================================================================================
# Agents and LLM Systems
Source: https://docs.hopsworks.ai/latest/concepts/mlops/agents/
# Agents and LLM Systems
Agentic workflows are one of the four classes of AI system, alongside real-time, batch, and stream processing.
Context engineering for an agent follows many of the same principles as feature engineering for a classical ML model, and the feature store is where the context comes from.
## The feature store as a retrieval source
An LLM system retrieves the context it needs at inference time, and the feature store is a natural source for that context.
Precomputed features are retrieved by entity ID, and embeddings are retrieved from a vector index by similarity search.
The key requirement is that the entity IDs are provided in the user query, as part of the deployment API, so the system knows whose features to retrieve.
This is retrieval-augmented generation (RAG) with a feature store: structured features by key, unstructured context by similarity.
--8<-- "concepts/mlops/agents/the-feature-store-as-a-retrieval-source.html"
## Workflow or agent
An LLM workflow has a control flow the developer designs: the steps and their order are fixed, and the LLM fills in each step.
An agent decides its own control flow: the LLM chooses which steps to run and in what order, calling tools as it goes.
A workflow is more predictable, an agent is more flexible, and most production systems start as workflows.
## MCP and A2A
Two protocols connect the moving parts.
MCP (Model Context Protocol) is how an agent calls its tools, the intra-agent interface to data sources and functions, including a feature store.
A2A (Agent-to-Agent) is how agents talk to each other, the inter-agent interface.
See the [agent guides](../../user_guides/agents/index.md) for how to build and deploy agents and agent tasks on Hopsworks.
================================================================================
# Projects and Governance
Source: https://docs.hopsworks.ai/latest/concepts/projects/governance/
# Projects and Governance
Hopsworks provides project-level multi-tenancy, a data mesh enabling technology.
Think of it as a GitHub repository for your teams and ML assets.
More specifically, a project is a sandbox for team members, ML assets (features, training data, models, vector index, model deployments), and optionally feature pipelines and training pipelines.
The ML assets can only be accessed by project members, and there is role-based access control (RBAC) for project members within a project.
--8<-- "concepts/projects/governance/projects-and-governance.html"
## Dev/Staging/Prod for Data
Projects enable you to define development, staging, and even production projects on the same cluster.
Often, companies deploy production projects on dedicated clusters, but development projects and staging projects on a shared cluster.
This way, projects can be easily used to implement CI/CD workflows.
## Data Mesh of Feature Stores
Projects enable you to move beyond the traditional dev/staging/prod ownership model for data.
Different teams or lines of business can have their own private feature stores, you can mix them with a group-wide feature store, and feature stores can be securely shared between teams/organizations.
Effectively, you can have decentralized ownership of feature stores, with domain-specific projects, and each project managing its own feature pipelines.
Hopsworks provides data/feature sharing support between these self-service projects.
## Audit Logs with REST API
Hopsworks stores audit logs for all calls on its REST API in its file system, HopsFS.
The audit log can be used to analyze the historical usage of services by users.
================================================================================
# Data Storage and Sharing
Source: https://docs.hopsworks.ai/latest/concepts/projects/storage/
# Data Storage and Sharing
Every project in Hopsworks has its own private assets:
- a Feature Store (including both Online and Offline Stores)
- a Filesystem subtree (all directory and files under /Projects//)
- a Model Registry
- Model Deployments
- Kafka topics
- OpenSearch indexes (including kNN indexes, the vector index)
- a Hive Database
Access control to these assets is controlled using project membership ACLs (access-control lists).
Users in a project who have a *Data Owner* role have read/write access to these assets. Users in a project who have a *Data Scientist* role have mostly read-only access to these assets, with the exception of the ability to write to well-known directories (Resources, Jupyter, Logs).
However, it is often desirable to share assets between projects, with read-only, read/write privileges, and to restrict the privileges to specific role (e.g., Data Owners) in the target project.
In Hopsworks, you can explicitly share assets between projects without copying the assets.
Sharing is managed by ACLs in Hopsworks, see example below:
--8<-- "concepts/projects/storage/data-storage-and-sharing.html"
================================================================================
# Architecture
Source: https://docs.hopsworks.ai/latest/concepts/fs/
# Feature Store Architecture
## What is Hopsworks Feature Store?
Hopsworks and its Feature Store are an open source data-intensive AI platform used for the development and operation of machine learning models at scale.
The Hopsworks Feature Store provides the Hopsworks API to enable clients to write features to feature groups in the feature store, and to read features from feature views - either through a low latency Online API to retrieve pre-computed features for operational models or through a high throughput, latency insensitive Offline API, used to create training data and to retrieve batch data for scoring.
--8<-- "concepts/fs/index/what-is-hopsworks-feature-store.html"
## Hopsworks API
The Hopsworks API is how you, as a developer, will use the feature store.
The Hopsworks API helps simplify some of the problems that feature stores address including:
- consistent features for training and serving
- centralized, secure access to features
- point-in-time JOINs of features to create training data with no data leakage
- easier connection and backfilling of features from external data sources
- use of external tables as features
- transparent computation of statistics and usage data for features.
## Write to feature groups, read from feature views
You write to feature groups with a feature pipeline program.
The program can be written in Python, Spark, or SQL.
You read from views on top of the feature groups, called feature views.
That is, a feature view does not store feature data, but is a logical grouping of features.
Typically, you define a feature view because you want to train/deploy a model with exactly those features in the feature view.
Feature views enable the reuse of feature data from different feature groups across different models.
================================================================================
# Features and Feature Groups
Source: https://docs.hopsworks.ai/latest/concepts/fs/feature_group/fg_overview/
# Features and Feature Groups
As a programmer, you can consider a feature, in machine learning, to be a variable associated with some entity that contains a value that is useful for helping train a model to solve a prediction problem.
That is, the feature is just a variable with predictive power for a machine learning problem, or task.
A feature group is a table of features.
Each feature group has a primary key, and optionally an event_time column (indicating when the features in that row were observed), a partition key, and foreign keys that point to the primary keys of other feature groups.
These are index columns, not features: they identify and join rows, and they are excluded when you select the features for a model.
A feature group stores untransformed feature data, so the same feature can be reused across models that each transform it differently.
??? note "Partitioning"
The partition key determines how the feature group rows are laid out on disk, so that queries using the partition key read only the data they need.
For example, if the partition key is the day and you have hundreds of days of data, a query for a given day or a range of days reads only those days from disk.
--8<-- "concepts/fs/feature_group/fg_overview/features-and-feature-groups.html"
## Online and offline Storage
Feature groups can be stored in a low-latency "online" database and/or in low cost, high throughput "offline" storage, typically a data lake or data warehouse.
A feature group with an embedding column can also have a vector index, for similarity search from inference pipelines and agents.
--8<-- "concepts/fs/feature_group/fg_overview/online-and-offline-storage.html"
### Online Storage
By default, the online store keeps only the latest values of features for a feature group.
It serves those precomputed features to models at runtime, and is backed by [RonDB](https://www.rondb.com), a low latency, high throughput, high availability data store.
By including an event_time column and a time-to-live (TTL), the online store can instead keep many rows per entity, which is what shift-right on-demand aggregations need.
### Offline Storage
The offline store stores the historical values of features for a feature group so that it may store much more data than the online store.
Offline feature groups are used, typically, to create training data for models, but also to retrieve data for batch scoring of models.
In most cases, offline data is stored in Hopsworks, but through the implementation of data sources, it can reside in an external file system.
The externally stored data can be managed by Hopsworks by defining ordinary feature groups or it can be used for reading only by defining [External Feature Group](external_fg.md).
================================================================================
# Write APIs
Source: https://docs.hopsworks.ai/latest/concepts/fs/feature_group/write_apis/
# Write APIs
You write to feature groups, and read from feature views.
There are 3 APIs for writing to feature groups, as shown in the table below:
| | Stream API | Batch API | Connector API |
| --- | --- | --- | --- |
| Python | X | - | - |
| Spark | X | X | - |
| External Table | - | - | X |
## Stream API
The Stream API is the only API for Python clients, and is
the preferred API for Spark, as it ensures consistent features between offline and online feature stores.
The Stream API first writes data to be ingested to a Kafka topic, and then Hopsworks ensures that the data is synchronized to the Online and Offline Feature Groups through the OnlineFS service and Hudi DeltaStreamer jobs, respectively.
The Kafka transport delivers at-least-once, and Hopsworks upgrades this to exactly-once through idempotent writes to the online feature group (only the latest values of features are stored there, and duplicates in Kafka only cause idempotent updates) and duplicate removal by Apache Hudi for the offline feature group.
--8<-- "concepts/fs/feature_group/write_apis/stream-api.html"
## Batch API
For very large updates to feature groups, such as when you are backfilling large amounts of data to an offline feature group, it is often preferential to write directly to the Hudi tables in Hopsworks, instead of via Kafka - thus reducing write amplification.
Spark clients can write directly to Hudi tables on Hopsworks with Hopsworks libraries and certificates using a HDFS API.
This requires network connectivity between the Spark clients and the datanodes in Hopsworks.
--8<-- "concepts/fs/feature_group/write_apis/batch-api.html"
## Connector API
Hopsworks supports external tables as feature groups.
You can mount a table from an external database as an offline feature group using the Connector API: you create an external table using the connector, without ingesting the data into Hopsworks.
This enables you to use features from your external data source (Snowflake, Redshift, Delta Lake, etc) as you would any feature in an offline feature group in Hopsworks.
You can, for example, join features from different feature groups (external or not) together to create feature views and training data for models.
See [External Feature Groups](external_fg.md) for the full list of supported data sources.
================================================================================
# External Feature Groups
Source: https://docs.hopsworks.ai/latest/concepts/fs/feature_group/external_fg/
# External Feature Groups
External feature groups are offline feature groups where their data is stored in an external table.
An external table requires a data source, defined with the [Connector API](write_apis.md#connector-api) (or more typically in the user interface), to enable Hopsworks to retrieve data from the external table.
An external feature group doesn't allow for offline data ingestion or modification; instead, it includes a user-defined SQL string for retrieving data.
You can also perform SQL operations, including projections, aggregations, and so on.
The SQL query is executed on-demand when Hopsworks retrieves data from the external Feature Group, for example, when creating training data using features in the external table.
In the image below, we can see that Hopsworks currently supports a large number of data sources, including any JDBC-enabled source, Snowflake, Data Lake, Redshift, BigQuery, Databricks Unity Catalog (Delta tables on Databricks on AWS only), S3, ADLS, GCS, SQL, and Kafka.
--8<-- "concepts/fs/feature_group/external_fg/external-feature-groups.html"
================================================================================
# Spine Group
Source: https://docs.hopsworks.ai/latest/concepts/fs/feature_group/spine_group/
# Spine Group
The default way to bring labels or prediction events is a label feature group: a regular feature group that holds the labels among its features, updated by a feature pipeline at a specific cadence.
Sometimes it is more convenient to provide the training events or entities in a Dataframe instead, when reading
feature data from the feature store through a feature view.
We call such a Dataframe a Spine as it is the structure around which
the training data or batch data is built.
In order to retrieve the correct feature values for the entities in the Dataframe, using
a point-in-time correct join, some additional metadata apart from the Dataframe schema is necessary.
Namely, the information about which
columns define the **primary key**, and which column indicates the **event time** at which the label was valid.
The spine Dataframe together with this additional metadata is what we call a **Spine Group**.
For example, in the following spine, we want to retrieve the features for the three locations, no later than the event time of each of the rainfall
measurements, which is our prediction target:
| location_id | event_time | rainfall (label) |
| ----------- | ---------------- | -----------------|
| 1 | 2022-06-01 13:11 | 44 |
| 2 | 2022-06-01 09:14 | 5 |
| 3 | 2022-06-01 06:36 | 2 |
A Spine Group does not materialize any data to the feature store itself, and always needs to be provided when retrieving features from the [offline API](../feature_view/offline_api.md).
You can think of it as a place holder or a temporary feature group, to be replaced by a Dataframe in point-in-time joins.
When using the [online API](../feature_view/online_api.md), it is not necessary to provide the spine, since the online feature store contains only the latest feature values, and therefore
no point in time join is required, the label is not required, as the inference pipeline is going to compute the prediction
and the primary key values are specified when calling the online API.
Prefer a label feature group where you can.
A Spine Group adds complexity and pushes work onto the clients, which must supply the entities and their event times on every call, and it can only be the root or label feature group of a feature view.
It is most appropriate for batch inference, where the set of entities to score is known only at request time.
================================================================================
# Feature Pipelines
Source: https://docs.hopsworks.ai/latest/concepts/fs/feature_group/feature_pipelines/
# Feature Pipelines
A feature pipeline is a program that orchestrates the execution of a dataflow graph of data validation, aggregation, dimensionality reduction, transformation, and other feature engineering steps on input data to create and/or update feature data.
With Hopsworks, you can write feature pipelines in different languages as shown in the figure below.
A feature pipeline can run on a schedule over a batch of data, or continuously over an event stream; see [Streaming Feature Pipelines](streaming_feature_pipelines.md).
--8<-- "concepts/fs/feature_group/feature_pipelines/feature-pipelines.html"
## Data Sources
Your feature pipeline needs to connect to some (external) data source to read the data to be processed.
Python and Spark have connectors to a huge number of different data sources, while SQL feature pipelines are often restricted to a single data source (for example, your connector to SnowFlake only runs SQL on SnowFlake).
SparkSQL, in contrast, can be used over tables that originate in different data sources.
## Data Validation
In order to be able to train and serve models that you can rely on, you need clean, high quality features.
Data validation operations include removing bad data, removing or imputing missing values, and identifying problems such as feature drift.
Hopsworks supports Great Expectations to specify data validation rules that are executed in the client before features are written to the Feature Store.
The validation results are collected and shown in Hopsworks.
Data validation in ML is a shift-left property: data is validated before it is written to a feature group, since one bad data point could later fail a training or inference run.
The default ingestion policy is STRICT, so a feature pipeline fails on a validation error rather than writing bad data.
## Aggregations
Aggregations are used to summarize large datasets into more concise, signal-rich features.
Popular aggregations include count(), sum(), mean(), median(), stddev(), min(), and max().
These aggregations produce a single number (a numerical feature) that captures information about a potentially large dataset.
Both numerical and categorical features are often transformed before being used to train or serve models.
## Dimensionality Reduction
If input data is impractically large or if it has a significant amount of redundancy, it can often be transformed into a reduced set of features with dimensionality reduction (often called feature extraction).
Popular dimensionality algorithms include embedding algorithms, PCA, and TSNE.
## Transformations
Transformations are covered in more detail in [training/inference pipelines](../feature_view/training_inference_pipelines.md), as transformations typically happen after the feature store.
If you store transformed features in feature groups, the feature data is no longer useful for EDA (as it near to impossible for Data Scientists to understand the transformed values).
It also makes it impossible for inference pipelines to log untransformed feature values and predictions for an operational model.
There is one use case for storing transformed features in feature groups - when you need to have ultra low latency when reading precomputed features (and online transformations when reading features add too much latency for your use case).
The figure below shows to include transformations in your feature pipelines.
--8<-- "concepts/fs/feature_group/feature_pipelines/transformations.html"
## Feature Engineering in Python
Python is the most widely used framework for feature engineering due to its extensive library support for aggregations (Pandas/Polars), data validation (Great Expectations), and dimensionality reduction (embeddings, PCA), and transformations (in Scikit-Learn, TensorFlow, PyTorch).
Python also supports open-source feature engineering frameworks used for automated feature engineering, such as [featuretools](https://www.featuretools.com/) that supports relational and temporal sources.
## Feature Engineering in Spark/PySpark
Spark is popular as a feature engineering framework as it can scale to process larger volumes of data than Python, and provides native support for aggregations, and it supports many of the same data validation (Great Expectations), and dimensionality reduction algorithms (embeddings, PCA) as Python.
Spark also has native support for transformations, which are useful for analytical models (batch scoring), but less useful for operational models, where online transformations are required, and Spark environments are less common.
Online model serving environments typically only support online transformations in Python.
## Feature Engineering in SQL
SQL has grown in popularity for performing heavy lifting in feature pipelines - computing aggregates on data - when the input data already resides in a data warehouse.
Data warehouses also support data validation, for example, through Great Expectations in DBT.
However, SQL is not mature as a platform for transformations and dimensionality reductions, where UDFs are applied row-wise.
You can do aggregation in SQL for data in your data warehouse or database.
## Feature Engineering in Beam
Beam feature engineering pipelines are supported in Java/Scala only.
================================================================================
# Streaming Feature Pipelines
Source: https://docs.hopsworks.ai/latest/concepts/fs/feature_group/streaming_feature_pipelines/
# Streaming Feature Pipelines
A streaming feature pipeline processes an unbounded stream of events and keeps features fresh in near real-time, instead of running on a schedule over a batch of data.
The same pipeline must also be able to run over historical data, to backfill a feature group when it is first created or after a schema change.
This backfill-and-incremental duality is a defining property of a feature pipeline, not an afterthought.
## Feature freshness
Feature freshness is the total time from when an event is first read by a feature pipeline to when the resulting feature is available to an inference pipeline.
For interactive, real-time systems it is often the freshness of a feature, not the latency of the model, that decides whether a prediction is useful.
--8<-- "concepts/fs/feature_group/streaming_feature_pipelines/freshness.html"
## Windows
Streaming aggregations are computed over windows of the event stream:
- **Tumbling** windows are fixed-size and non-overlapping, so each event falls in exactly one window.
- **Hopping** windows are fixed-size but overlap, advancing by a hop smaller than the window.
- **Rolling** (sliding) windows are recomputed continuously as events arrive.
A watermark tells the pipeline how long to wait for late-arriving events before it closes a window and emits the aggregate.
## Streaming-native or hybrid
A streaming-native pipeline computes all features directly on the stream, a Kappa-style architecture.
A hybrid streaming-batch pipeline splits the work: a streaming job keeps the freshest features up to date while a batch job computes the heavier, less time-sensitive aggregations, a Lambda-style architecture.
Prefer streaming-native where you can, since a single code path is simpler to keep consistent than two.
A streaming feature pipeline can run in four operational modes: real-time processing of live events, stream replay, backfilling from historical data, and stream reprocessing after a logic change.
================================================================================
# Data Transformations
Source: https://docs.hopsworks.ai/latest/concepts/mlops/data_transformations/
# Data Transformations
[Data transformations](https://www.hopsworks.ai/dictionary/data-transformation) are integral to all AI applications.
Data transformations produce new features that can enhance the performance of an AI application.
However, [not all transformations in an AI application are equivalent](https://www.hopsworks.ai/post/a-taxonomy-for-data-transformations-in-ai-systems).
Transformations like binning and aggregations typically create reusable features, while transformations like one-hot encoding, scaling and normalization often produce model-specific features.
Additionally, in real-time AI systems, some features can only be computed during inference when the request is received, as they need request-time parameters to be computed.
--8<-- "concepts/mlops/data_transformations/data-transformations.html"
This classification of features can be used to create a taxonomy for data transformation that would apply to any scalable and modular AI system that aims to reuse features.
The taxonomy helps identify which classes of data transformation can cause [online-offline](https://www.hopsworks.ai/dictionary/online-offline-feature-skew) skews in AI systems, allowing for their prevention.
Hopsworks provides support for a feature view abstraction as well as model-dependent transformations and on-demand transformations to prevent online-offline skew.
## Data Transformation Taxonomy for AI Systems
Transformation functions in an AI system can be classified into three types based on the nature of the input features they generate: [model-independent](https://www.hopsworks.ai/dictionary/model-independent-transformations), [model-dependent](https://www.hopsworks.ai/dictionary/model-dependent-transformations), and [on-demand](https://www.hopsworks.ai/dictionary/on-demand-transformation) transformations.
--8<-- "concepts/mlops/data_transformations/transformation-taxonomy-2.html"
**Model-independent transformations** create reusable features that can be utilized across one or more machine-learning models.
These transformations include techniques such as grouped aggregations (e.g., minimum, maximum, or average of a variable), windowed aggregations (e.g., the number of clicks per day), and binning to generate categorical variables.
Since the data produced by model-independent transformations are reusable, these features can be stored in a feature store.
**Model-dependent transformations** generate features specific to one model.
These include transformations that are unique to a particular model or are parameterized by the training dataset, making them model-specific.
For instance, text tokenization is a transformation required by all large language models (LLMs) but each LLM has their own (unique) tokenizer.
Other transformations, such as encoding categorical variables in a numerical representation or scaling/normalizing/standardizing numerical variables to enhance the performance of gradient-based models, are parameterized by the training dataset.
Consequently, the features produced are applicable only to the model trained using that specific training dataset.
Since these features are not reusable, there is no need to store them in a feature store.
Also, storing encoded features in a feature store leads to write amplification, as every time feature values are written to a feature group, all existing rows in the feature group have to be re-encoded (and creation of a training dataset using a subset or rows in the feature group becomes impossible as they cannot be re-encoded).
**On-demand transformations** are exclusive to [real-time AI systems](https://www.hopsworks.ai/dictionary/real-time-machine-learning), where predictions must be generated in real time based on incoming prediction requests.
On-demand transformations compute on-demand features, which usually require at least one input parameter that is only available in a prediction request for their computation.
These transformations can also combine request-time parameters with precomputed features from feature stores.
Some examples include generating *zip_codes* from latitude and longitude received in the prediction request or calculating the *time_since_last_transaction* from a transaction request.
The on-demand features produced can also be computed and [backfilled](https://www.hopsworks.ai/dictionary/backfill-features) into a feature store when the necessary historical data required for their computation becomes available.
Backfilling on-demand features into the feature store eliminates the need to recompute them when creating training data.
On-demand transformations are typically also model-independent transformations (model-dependent transformations can be applied after the on-demand transformation).
Each of these transformations is employed within specific areas in a modular AI system and can be illustrated using the figure below.
--8<-- "concepts/mlops/data_transformations/transformation-taxonomy.html"
Model-independent transformations are utilized exclusively in areas where new and historical data arrives, typically within feature pipelines.
Model-dependent transformations are necessary during the creation of training data, in training programs and must also be consistently applied in inference programs prior to making predictions.
On-demand transformations are primarily employed in online inference programs, though they can also be integrated into feature engineering programs to backfill data into the feature store.
The presence of model-dependent and on-demand transformations across different modules in a modular AI system introduces the potential for online-offline skew.
Hopsworks provides support for model-dependent transformations and on-demand transformations to easily create modular skew-free AI pipelines.
## Hopsworks and the Data Transformation Taxonomy
--8<-- "concepts/mlops/data_transformations/hopsworks-transformation-taxonomy-2.html"
--8<-- "concepts/mlops/data_transformations/hopsworks-feature-store-storage.html"
In Hopsworks, an AI system is typically decomposed into different [AI pipelines](https://www.hopsworks.ai/dictionary/ai-pipelines) and usually falls into either a [feature pipeline](https://www.hopsworks.ai/dictionary/feature-pipeline), a [training pipeline](https://www.hopsworks.ai/dictionary/training-pipeline), or an [inference pipeline](https://www.hopsworks.ai/dictionary/inference-pipeline).
Hopsworks stores reusable feature data, created by model-independent transformations within the feature pipeline, into [feature groups](../fs/feature_group/fg_overview.md) (tables containing feature data in both offline and online stores).
Model-independent transformations in Hopsworks can be performed using a wide range of commonly used data engineering tools and the generated features can be seamlessly inserted into feature groups.
The figure below illustrates the different software tools supported by Hopsworks for creating reusable features through model-independent transformations.
--8<-- "concepts/mlops/data_transformations/hopsworks-transformation-taxonomy.html"
Additionally, Hopsworks provides a simple Python API to [create custom transformation functions](../../user_guides/fs/transformation_functions.md) as either Python or Pandas User-Defined Functions (UDFs).
Pandas UDFs enable the vectorized execution of transformation functions, offering significantly higher throughput compared to Python UDFs for large volumes of data.
They can also be scaled out across workers in a Spark program, allowing for scalability from gigabytes (GBs) to terabytes (TBs) or more.
However, Python UDFs can be much faster for small volumes of data, such as in the case of online inference.
Transformation functions defined in Hopsworks can then be attached to feature groups to [create on-demand transformation](../../user_guides/fs/feature_group/on_demand_transformations.md).
On-demand transformations in feature groups are executed automatically whenever data is inserted into them to compute and backfill the on-demand features into the feature group.
Backfilling on-demand features removes the need to recompute them while creating training and batch data.
Hopsworks also provides a powerful abstraction known as [feature views](../fs/feature_view/fv_overview.md), which enables feature reuse and prevents skew between training and inference pipelines.
A feature view is a meta-data-only selection of features, created from potentially different feature groups.
It includes the input and output schema required for a model.
This means that a feature view describes not only the input features but also the output targets, along with any helper columns necessary for training or inference of the model.
This allows feature views to create consistent snapshots of data for both training and inference of a model.
Additionally feature views, also compute and save statistics for the training datasets they create.
Hopsworks supports attaching transformations functions to feature views to [create model-dependent transformations](../../user_guides/fs/feature_view/model-dependent-transformations.md) that have no online-offline skew.
These transformations get access to the same training dataset statistics during both training and inference ensuring their consistency.
Additionally, feature views through lineage get access to the on-demand transformation used to create on-demand features if any are selected during the creation of the feature view.
The registration locus is the cleanest way to remember where each transformation lives: on-demand transformations are registered on feature groups, model-dependent transformations on feature views.
A Hopsworks transformation function is also mixed-mode: the same decorated Python function runs as a Pandas UDF offline, to create training data, and as a Python UDF online, to build a single feature vector, so one definition serves both pipelines with no skew.
This allows for the computation of on-demand features in real-time during online-inference.
================================================================================
# Overview
Source: https://docs.hopsworks.ai/latest/concepts/fs/feature_view/fv_overview/
# Feature Views
A feature view is a logical view over (or interface to) a set of features that may come from different feature groups.
You create a feature view by selecting features, starting from a root feature group and following foreign keys to join in features from other feature groups.
When the feature view has a label for supervised learning, the root feature group is the label feature group, the one feature group that holds the labels.
Features are reachable by graph traversal: any feature group joined to the root can, in turn, have foreign keys to further feature groups whose features you can also select.
A feature view does not have a primary key of its own; it has serving keys, the foreign keys of its label feature group, which you provide to retrieve feature vectors.
In the illustration below, we can see that features are joined together from the two feature groups: seller_delivery_time_monthly and the seller_reviews_quarterly.
You can also see that features in the feature view inherit not only the feature type from their feature groups, but also whether they are the primary key and/or the event_time.
The image also includes transformation functions that are applied to individual features.
Transformation functions are a part of the feature types included in the feature view.
That is, a feature in a feature view is not only defined by its data type (int, string, etc) or its feature type (categorical, numerical, embedding), but also by its transformation.
--8<-- "concepts/fs/feature_view/fv_overview/feature-views-2.html"
Feature views can also include:
- the label for the supervised ML problem
- transformation functions that should be applied to specified features consistently between training and serving
- the ability to create training data
- the ability to retrieve a feature vector with the most recent feature values
In the flow chart below, we can see the decisions that can be taken when creating (1) a feature view, and (2) creating training data with the feature view.
--8<-- "concepts/fs/feature_view/fv_overview/feature-views.html"
We can see here how the feature view is a representation for a model in the feature store - the same feature view is used to retrieve feature vectors for operational model that was created with training data from this feature view.
As such, you can see that the most common use case for creating a feature view is to define the features that will be used in a model.
In this way, feature views enable features from different feature groups to be reused across different models, and if features are stored untransformed in feature groups, they become even more reusable, as different feature views can apply different transformations to the same feature.
================================================================================
# Offline API
Source: https://docs.hopsworks.ai/latest/concepts/fs/feature_view/offline_api/
# Offline API
The feature view provides an *Offline API* for
- creating training data
- creating batch (scoring) data
## Training Data
Training data is created using a feature view.
You can create training data as either:
- in-memory Pandas/Polars DataFrames, useful when you have a small amount of training data;
- materialized training data in files, in a file format of your choice (such as `.tfrecord`, `.csv`, or `.parquet`).
You can apply filters when creating training data from a feature view:
- `start_time` and `end_time`, for example, to create the train-set from an earlier time range, and the test-set from a later (unseen) time range;
- feature value features, for example, only train a model on customers from a particular country.
Note that filters are not applied when retrieving feature vectors using feature views, as we only look up features for a specific entity, like a customer.
In this case, the application should know that predictions for this customer should be made on the model trained on customers in USA, for example.
Materialized training data is immutable: once created, a training dataset version is not appended to or modified.
To retrain on new data, create a new training dataset version.
If the new data needs to be computed continuously, for example a daily batch for a time-series model, do that computation once in a [derived feature group][assign-parents-to-a-feature-group] that is kept up to date as new data arrives, and create a new training dataset version from it whenever you need updated data.
### Point-in-time Correct Training Data
When you create training data from features in different feature groups, it is possible that the feature groups are updated at different cadences.
For example, maybe one feature group is updated hourly, while another feature group is updated daily.
It is very complex to write code that joins features together from such feature groups and ensures there is no data leakage in the resultant training data.
Hopsworks hides this complexity by performing the point-in-time JOIN transparently, similar to the illustration below:
--8<-- "concepts/fs/feature_view/offline_api/point-in-time-correct-training-data-2.html"
Hopsworks uses the event_time columns on both feature groups to determine the most recent (but not newer) feature values that are joined together with the feature values from the feature group containing the label.
That is, the features in the feature group containing the label are the observation times for the features in the resulting training data, and we want feature values from the other feature groups that have the most recent timestamps, but not newer than the timestamp in the label-containing feature group.
--8<-- "concepts/fs/feature_view/offline_api/point-in-time-correct-training-data.html"
#### Spine Groups
The left side of the point-in-time join is typically the set of training entities/primary key values for which the relevant features need to be retrieved.
This left side of the join can also be replaced by a [spine group](../feature_group/spine_group.md).
When using feature groups also so save labels/prediction targets, it can happen that you end up with the same entity multiple times in the training dataset depending on the cadence at which the label group was updated and the length of the event time interval
that is being used to generate the training dataset.
This can lead to bias in the training dataset and should be avoided.
To avoid this kind of situation, users can either narrow down the event time interval during training dataset creation or use a spine
in order to precisely define the entities to be included in the training dataset.
This is just one example where spines are helpful.
### Splitting Training Data
You can create random train/validation/test splits of your training data using the Hopsworks API.
You can also time-based splits with the Hopsworks API.
### Evaluation Sets
Test data can also be split into evaluation sets to help evaluate a model for potential bias.
First, you have to identify the classes of samples that could be at risk of bias, and generate *evaluation sets* from your unseen test set - one evaluation set for each group of samples at risk of bias.
For example, if you have a feature group of users, where one of the features is gender, and you want to evaluate the risk of bias due to gender, you can use filters to generate 3 evaluation sets from your test set - one for male, female, and non-binary.
Then you score your model against all 3 evaluation sets to ensure that the prediction performance is comparable and non-biased across all 3 gender.
## Batch (Scoring) Data
Batch data for scoring models is created using a feature view.
Similar to training data, you can create batch data as either:
- in-memory Pandas/Polars DataFrames, useful when you have a small amount of data to score;
- materialized data in files, in a file format of your choice (such as `.tfrecord`, `.csv`, or `.parquet`)
Batch data requires specification of a `start_time` for the start of the batch scoring data.
You can also specify the `end_time` (default is the current date).
--8<-- "concepts/fs/feature_view/offline_api/batch-scoring-data.html"
### Spine Dataframes
Similar to training dataset generation, it might be helpful to specify a spine when retrieving features for batch inference.
The only difference in this case is that the spine dataframe doesn't
need to contain the label, as this will be the output of the inference pipeline.
A typical use case is the handling of opt-ins, where certain customers have to be excluded from an inference pipeline due to a missing marketing opt-in.
================================================================================
# Online API
Source: https://docs.hopsworks.ai/latest/concepts/fs/feature_view/online_api/
# Online API
The Feature View provides an Online API to return an individual feature vector, or a batch of feature vectors, containing the latest feature values.
To retrieve a feature vector, a client provides the feature view's serving keys.
The serving keys are the foreign keys of the feature view's label feature group; a feature view does not have a primary key of its own.
For example, if a feature view is built from the `customer_profile` and `customer_purchases` feature groups joined on `customer_id`, then `customer_id` is the serving key you provide to retrieve a feature vector.
## Feature Vectors
A feature vector is a row of features (without the primary key(s) and event timestamp):
--8<-- "concepts/fs/feature_view/online_api/feature-vectors.html"
It may be the case that for any given feature vector, not all features will come pre-engineered from the feature store.
Some features will be provided by the client (or at least the raw data to compute the feature will come from the client).
We call these 'passed' features and, similar to precomputed features from the feature store, they can also be transformed by the Hopsworks client in the method:
```python
feature_view.get_feature_vector(entry, passed_features={"pressure": 1013})
```
When you call `get_feature_vector`, Hopsworks builds the vector in a fixed order:
1. retrieve the precomputed features from the online store using the serving keys,
2. merge in any passed features,
3. compute on-demand transformations (ODTs),
4. compute model-dependent transformations (MDTs),
5. drop the index and helper columns,
6. return the feature vector.
This ordering is the composition constraint: on-demand transformations run before model-dependent ones, which are always last, just before the model is called.
================================================================================
# Consistent Transformations
Source: https://docs.hopsworks.ai/latest/concepts/fs/feature_view/training_inference_pipelines/
# Consistent Transformations
A *training pipeline* is a program that orchestrates the training of a machine learning model.
For supervised machine learning, a training pipeline requires both features and labels, and these can typically be retrieved from the feature store as either in-memory Pandas/Polars DataFrames or read as training data files, created from the feature store.
An *inference pipeline* is a program that takes user input, optionally enriches it with features from the feature store, and builds a feature vector (or batch of feature vectors) with with it uses a model to make a prediction.
## Transformations
Feature transformations are mathematical operations that change feature values with the goal of improving model convergence or performance properties.
Transformation functions take as input a single value (or small number of values), they often require state (such as the mean value of a feature to normalize the input), and they output a single value or a list of values.
## Offline-Online Feature Skew
Offline-online feature skew is a difference between the transformation code that runs in an offline (training) pipeline and the transformation code that runs in the corresponding inference pipeline.
It is a code difference, not a data difference, so it cannot be detected by comparing distributions; the only way to avoid it is to run the same code in both pipelines.
In the image below, you can see that transformations happen after the Feature Store, but that the implementation of the transformation functions need to be consistent between the training and inference pipelines.
--8<-- "concepts/fs/feature_view/training_inference_pipelines/offline-online-feature-skew.html"
There are 3 main approaches to prevent offline-online feature skew that we support in Hopsworks.
These are (1) perform transformations in models, (2) perform transformations in pipelines (sklearn, TF, PyTorch) and use the model registry to save the transformation pipeline so that the same transformation is used in your inference pipeline, and (3) use Hopsworks transformations, defined as UDFs in Python.
### Transformations as Pre-Processing Layers in Models
Transformation functions can be implemented as preprocessing steps within a model.
For example, you can write a transformation function as a pre-processing layer in Keras/TensorFlow.
When you save the model, the preprocessing steps will also be saved as part of the model.
Any state required to compute the transformation, such as the arithmetic mean of a numerical feature in the train set, is also stored with the function, enabling consistent transformations during inference. When data preprocessing is part of the model, users can just send the untransformed feature values to the model and the model itself will apply any transformation functions as preprocessing layers (such as encoding categorical variables or normalizing numerical variables).
### Transformation Pipelines in Scikit-Learn/TensorFlow/PyTorch
You have to save your transformation pipeline (serialize the object or the parameters) and make sure you apply exactly the same transformations in your inference pipeline.
This means you should version the transformations.
In Hopsworks, you can store the transformations with your versioned models in the Model Registry, helping you to ensure the same transformation pipeline is applied to both training/serving for the same model version.
### Transformations as Python UDFs in Hopsworks
Hopsworks feature store also supports consistent transformation functions by enabling a Python UDF, that implements a transformation, to be attached a to feature in a feature view.
When training data is created with a feature view or when a feature vector is retrieved from a feature view, Hopsworks ensures that any transformation functions defined over any features will be applied before returning feature values.
You can use built-in transformation objects in Hopsworks or write your own custom transformation functions as Python UDFs.
The benefit of this approach is that transformations are applied consistently when creating training data and when retrieving feature data from the online feature store.
Transformations no longer need to be included in either your training pipeline or inference pipeline, as they are applied transparently when creating training data and retrieving feature vectors.
Hopsworks uses Spark to create training data as files, and any transformation functions for features are executed as Python UDFs in Spark - enabling transformation functions to be applied on large volumes of data and removing potentially CPU-intensive transformations from training pipelines.
================================================================================
# On-Demand Features
Source: https://docs.hopsworks.ai/latest/concepts/fs/feature_group/on_demand_feature/
# On-demand features
Features are defined as on-demand when their value cannot be pre-computed beforehand, rather they need to be computed in real-time during inference.
This is achieved by implementing the on-demand features as a Python function in a Python module.
Also ensure that the same version of the Python module is installed in both the feature and inference pipelines.
The figure below shows an example from a housing price model: a zip code (or post code) computed on demand from longitude and latitude parameters.
In your online application, longitude and latitude are provided as parameters to the application, and the same python function used to calculate the zip code in the feature pipeline is used to compute the zip code in the Online Inference pipeline.
--8<-- "concepts/fs/feature_group/on_demand_feature/on-demand-features.html"
## Shift left or shift right
Deciding to compute a feature on-demand is a shift-right decision, and it is one of the biggest feature-engineering choices you make.
Shift left means precomputing a feature in a feature pipeline and storing it in the feature store for retrieval.
Shift right means computing it at request time, in an on-demand or model-dependent transformation.
Shift right when the feature depends on request-time input, such as the zip code computed from the longitude and latitude in a request, or when a precomputed value would be too stale to be useful.
Shift left when the feature can be precomputed, to keep inference latency low and avoid repeating the computation on every request.
The trade-off is latency and operational overhead against freshness.
================================================================================
# Data Validation, Statistics, Alerts
Source: https://docs.hopsworks.ai/latest/concepts/fs/feature_group/fg_statistics/
# Data Validation, Statistics, and Alerts
Hopsworks supports monitoring, validation, and alerting for features:
- transparently compute statistics over features on writing to a feature group;
- validation of data written to feature groups using Great Expectations
- alerting users when there was a problem writing or update features.
## Statistics
When you create a Feature Group in Hopsworks, you can configure it to compute statistics over the features inserted into the Feature Group by setting the `statistics_config` dict parameter, see [Feature Group Statistics](../../../user_guides/fs/feature_group/statistics.md) for details.
Every time you write to the Feature Group, new statistics will be computed over all of the data in the Feature Group.
## Data Validation
You can define expectation suites in Great Expectations and associate them with feature groups.
When you write to a feature group, the expectations are executed, then you can define a policy on the feature group for what to do if any expectation fails.
--8<-- "concepts/fs/feature_group/fg_statistics/data-validation.html"
## Alerting
Hopsworks also supports alerts, that can be triggered when there are problems in your feature pipelines, for example, when a write fails due to an error or a failed expectation.
You can send alerts to different alerting endpoints, such as email or Slack, that can be configured in the Hopsworks UI.
For example, you can send a slack message if features being written to a feature group are missing some input data.
================================================================================
# Feature Monitoring
Source: https://docs.hopsworks.ai/latest/concepts/fs/feature_group/feature_monitoring/
# Feature Monitoring
Feature monitoring complements data validation by letting you monitor feature data after it has been ingested into the feature store.
It computes statistics over a detection window of data, compares them against a reference window, and raises alerts when the comparison crosses a threshold.
The comparison can be a single scalar metric (e.g., the mean) or the whole feature distribution using a distance metric such as PSI or KL divergence.
You can monitor at two levels, and each level detects a different kind of change.
## Monitoring a feature group
Monitoring a feature group watches the raw data as it is ingested, independently of any model.
The reference window is usually an earlier window of the same feature group, so what you detect is data ingestion drift: a new batch that no longer looks like the data already in the feature group.
After creating a feature group, you can schedule statistics over one or more features, computed on the whole feature data or on a subset defined by a detection window.
You then enable a comparison against a reference window and define the criteria: which statistic to compare and the threshold that flags an anomaly.
## Monitoring a feature view
Monitoring a feature view watches what a specific model actually sees, because a feature view backs the features served to a model.
Here the reference window is typically the model's training dataset, so what you detect is feature drift: the served features drifting away from the distribution the model was trained on.
The mechanism is the same scheduled statistics and distribution comparison as for a feature group, computed using the feature view query; only the reference changes.
Comparing a model's logged inference data against its training dataset, and deciding when to retrain, is model monitoring; see [Model Monitoring](../../mlops/model_monitoring.md).
## Statistics on training data
A feature view holds no statistics of its own, since it is only an interface over features and their transformations.
Statistics are computed over a training dataset instead.
Those training-dataset statistics are the reference that feature-view monitoring compares against, and some online transformations need them too (normalizing a numerical feature requires the training-set mean).
!!! info "Feature Monitoring Guide"
More information can be found in the [Feature monitoring guide](../../../user_guides/fs/feature_monitoring/index.md).
================================================================================
# Versioning
Source: https://docs.hopsworks.ai/latest/concepts/fs/feature_group/versioning/
# Versioning
Hopsworks versions the ML assets that make up an AI system, so that a model in production is reproducible and clients are protected from breaking changes.
Feature groups, feature views, training data, and models are versioned; deployments are the one asset that is not.
## Feature group schema versioning
The schema of feature groups is versioned.
If you make a breaking change to the schema of a feature group, you need to increment the version of the feature group, and then backfill the new feature group.
A breaking schema change is when you:
- drop a column from the schema
- add a new feature without any default value for the new feature
- change how a feature is computed, such that, for training models, the data for the old feature is not compatible with the data for the new feature.
For example, if you have an embedding as a feature and change the algorithm to compute that embedding, you probably should not mix feature values computed with the old embedding model with feature values computed with the new embedding model.
--8<-- "concepts/fs/feature_group/versioning/feature-group-schema-versioning.html"
## Feature group data versioning
Data versioning of a feature group tracks updates to the feature group, so that you can recover the state of the feature group at a given point-in-time in the past.
--8<-- "concepts/fs/feature_group/versioning/feature-group-data-versioning.html"
There are two points in time you can travel back to, and they answer different questions.
As-of ingestion time reads the data as it had been written by a given moment, which gives reproducible training data.
As-of event time reads the data as it was true in the world at a given moment, which gives point-in-time correct training data with no future leakage.
## Feature view and training data versioning
Feature views are interfaces, and if there is a change in the interface (the types of the features, the transformations applied to the features), then you need to change the version, to prevent breaking existing clients.
Training datasets are associated with a specific feature view version, and each training dataset also has its own version number.
For example, online transformation functions often need training data statistics (e.g., normalizing a numerical feature requires you to divide the feature value by the mean value for that feature in the training dataset).
As many training datasets can be created from a feature view, when you initialize the feature view you need to tell it which version of the training data to use: `feature_view.init(1)` means use version 1 of the training data for this feature view.
--8<-- "concepts/fs/feature_group/versioning/feature-view-training-data-versioning.html"
## Models and deployments
A model has its own version in the model registry.
A deployment, however, is not versioned: it is the one mutable asset.
A new deployment gets a new name, upgrades and rollbacks are done with blue/green deployments, and clients depend on the [deployment API](../../mlops/serving.md), not on a deployment version number.
A model deployment is also tightly coupled to the versioned feature views that supply its pre-computed features, so versioning the model alone is not enough.
================================================================================
# Tags, Search, Lineage
Source: https://docs.hopsworks.ai/latest/concepts/projects/search/
# Tags, Search, and Lineage
## Search { #search-concept }
Hopsworks supports free-text search to discover machine-learning assets:
- features
- feature groups
- feature views
- training data
- jobs, including apps
- models
- deployments, including agents
You can use the search bar at the top of your project to free-text search for the names or descriptions of any ML asset.
You can also search using keywords or tags that are attached to an ML asset.
Tags are indexed for every asset in the list above, so a governance question can be asked once across the whole set rather than per asset type.
Keywords apply to feature groups, feature views and training data only, so a keyword filter never matches a job, a model or a deployment.
Apps and agents are reported as their own classes but are not stored as their own kind of asset.
An app is a job whose type is PythonApp, and an agent is a deployment serving no registered model.
Each is a narrowing of the class it belongs to, which is why they need no separate index and appear the moment the distinguishing property does.
You can search for assets within a specific project or across all projects in a Hopsworks deployment, including those you are not a member of.
This allows for easier discoverability and reusability of assets within an organization.
To avoid users gaining unauthorized access to data, if a search result is in a project you are **not** a member of, the information displayed is limited to: names, descriptions, tags, asset creator and create date.
If the search result is within a project you are a member of, you are also able to inspect recent activities on the asset as well as statistics.
Searching across projects you are not a member of is on by default.
An administrator running a multi-tenant deployment can turn it off, which restricts every search to the projects the caller can already access.
See [search index administration][search-index-administration] for the setting.
For how to use search, including filtering by a specific tag key and value, see the [search guide][search-guide].
## Tags
A keyword is a single user-defined word attached to an ML asset.
Keywords can be used to help it make it easier to find ML assets or understand the context in which they should be used, for example, *PII* could be used to indicate that the ML asset is based on personally identifiable information.
However, it may be preferable to have a stronger governance framework for ML assets than keywords alone.
For this, you can define a *schematized tag*, defining a list of key/value tags along with a type for a value.
In the figure below, you can see an example of a schematized tag with two key/value pairs: *pii* of type boolean (indicating if this feature group contains PII data), and *owner* of type string (indicating who the owner of the data in this feature group is).
Keywords are not part of a schematized tag and are not shown in this panel.
They are attached in the feature group header instead, for example a *eu_region* keyword indicating the data has its origins in the EU.
Schematized tags can also enforce policy, not just aid discovery.
You could, for example, require that a model cannot be created in the production model registry unless its EU AI Act tag is filled in correctly.
## Lineage
Hopsworks tracks the lineage (or provenance) of ML assets automatically for you.
The lineage chain runs end to end: data source, feature group, feature view, training data, model, deployment.
This is what makes governance answerable: from a biased model you can trace back to the exact feature groups and data sources that fed it.
You can see what features are used in which feature view or training dataset, and what training dataset was used to train a given model.
For assets that are managed outside of Hopsworks, there is support for the explicit definition of lineage dependencies.
--8<-- "concepts/projects/search/provenance-lineage.html"
The lineage of an asset is shown in the Hopsworks UI, here the feature groups behind a feature view and the training data and models derived from it.
================================================================================
# CI/CD
Source: https://docs.hopsworks.ai/latest/concepts/projects/cicd/
# CI/CD Support
You can setup traditional development, staging, and production environment in Hopsworks using Projects.
A project enables you provide access control for the different environments - just like a GitHub repository, owners of projects can add and remove members of projects and assign different roles to project members - the "data owner" role can write to feature store, while a "data scientist" can only read from the feature store and create training data.
## Dev, Staging, Prod
You can create dev, staging, and prod projects - either on the same cluster, but mostly commonly, with production on its own cluster:
--8<-- "concepts/projects/cicd/dev-staging-prod.html"
## Versioning
Automated promotion across dev, staging, and prod relies on every ML asset being versioned.
Hopsworks versions feature groups, feature views, training data, and models, while deployments stay mutable behind the deployment API.
See [Versioning](../fs/feature_group/versioning.md) for what is versioned and how.
## Pytest for feature logic and feature pipeline tests
Pytest and Great Expectations can be used for testing feature pipelines.
Pytest is used to test feature logic and for end-to-end feature pipeline tests, while Great Expectations is used for data validation tests.
Here, we can see how a feature pipeline test uses sample data to compute features and validate they have been written successfully, first to a development feature store, and then they can be pushed to a staging feature store, before finally being promoted to production.
--8<-- "concepts/projects/cicd/pytest-feature-logic.html"
================================================================================
# Model Training
Source: https://docs.hopsworks.ai/latest/concepts/mlops/training/
# Model Training
A training pipeline is a program that orchestrates the training of a machine learning model, reading features and labels from the feature store as training data.
Hopsworks supports running model training pipelines on any Python environment, whether on an external Python client or on a Hopsworks cluster.
The outputs of a training pipeline are typically experiment results, including logs, and possibly a trained model.
You can plugin your own experimentation tracking platform or model registry, or you can use Hopsworks.
A training pipeline typically runs five steps: select a feature view and a training dataset version, train the model, evaluate it, validate it, and register it in the model registry if it passes.
## Evaluation and validation
Model evaluation and model validation are not the same thing.
Evaluation measures the model's performance on a held-out test set, using metrics such as accuracy or AUC.
Validation is a pass/fail gate: the model is run against evaluation data, including bias slices of the holdout built with feature-view filters and training helper columns (a column such as gender used to slice results but dropped before training), and only a model that passes is registered.
The output of validation is a model validation scorecard, and it is what decides whether the model reaches the registry.
## Training Pipelines on Hopsworks
If you train models with Hopsworks, you can setup CI/CD pipelines as shown below, where the experiments are tracked by Hopsworks, and any model created is published to a model registry.
Each project has its own private model registry, so when you are working in a development project, you typically publish models to your project's private development registry, and if all model validation tests pass, and the model performance is good enough, the same training pipeline can be submitted via a CI/CD pipeline (e.g., GitHub push request) to a staging project, and the same procedure can be repeated to push the training pipeline to a production project.
--8<-- "concepts/mlops/training/training-pipelines-on-hopsworks.html"
Hopsworks [Model Registry](registry.md) and [Model Serving](serving.md) capabilities can then be used to build a batch or online prediction service using the model.
================================================================================
# Model Registry
Source: https://docs.hopsworks.ai/latest/concepts/mlops/registry/
# Model Registry
Hopsworks Model Registry is designed with specific support for KServe and MLOps, through versioning.
It enables developers to publish, test, monitor, govern and share models for collaboration with other teams.
The model registry is where developers publish their models during the experimentation phase.
The model registry can also be used to share models with the team and stakeholders.
Like other project-based multi-tenant services in Hopsworks, a model registry is private to a project.
That means you can easily add a development, staging, and production model registry to a cluster, and implement CI/CD processes for transitioning a model from development to staging to production.
The model registry for KServe's capability are shown in the diagram below:
--8<-- "concepts/mlops/registry/model-registry.html"
The model registry centralizes model management, enabling models to be securely accessed and governed.
Models are more than just the model itself - the registry also stores sample data for testing, configuration information, provenance information, environment variables, links to the code used to generate the model, the model version, and tags/descriptions).
When you save a model, you can also save model metrics with the model, enabling users to understand, for example, performance of the model on test (or unseen) data.
## Model Package
A ML model consists of a number of different components in a model package:
- Model Input/Output Schema
- Model artifacts
- Model version information
- Model format (based on the ML framework used to train the model - e.g., .pkl or .tb files)
You can also optionally include in your packaged model:
- Sample data (used to test the model in KServe)
- The source notebook/program/experiment used to create the model
================================================================================
# Model Serving
Source: https://docs.hopsworks.ai/latest/concepts/mlops/serving/
# Model Serving
In Hopsworks, you can easily deploy models from the model registry using [KServe](https://kserve.github.io/website/latest/), the standard open-source framework for model serving on Kubernetes.
You rarely deploy just a model.
What you deploy is an online inference pipeline, of which the model is one part, alongside feature retrieval, transformations, and logging.
You can deploy models programmatically using [`Model.deploy`][hsml.model.Model.deploy] or via the UI.
A KServe model deployment can include the following components:
**`Predictor (KServe component)`**
: A predictor runs a model server (Python, TensorFlow Serving, or vLLM) that loads a trained model, handles inference requests and returns predictions.
**`Transformer (KServe component)`**
: A ^^pre-processing^^ and ^^post-processing^^ component that can transform model inputs before predictions are made, and predictions before these are delivered back to the client.
Not available for vLLM deployments.
**`Inference Logger`**
: Hopsworks logs inputs and outputs of transformers and predictors to a ^^Kafka topic^^ that is part of the same project as the model.
This is for storing inference requests and responses for later consumption and analysis, and is separate from the feature logging that powers [Model Monitoring](model_monitoring.md).
Not available for vLLM deployments.
**`Inference Batcher`**
: Inference requests can be batched to improve throughput (at the cost of slightly higher latency).
**`Istio Model Endpoint`**
: You can publish a model over REST(HTTP) or gRPC using a Hopsworks API key, accessible via **path-based routing** through Istio.
API keys have scopes to ensure the principle of least privilege access control to resources managed by Hopsworks.
For more details on path-based routing of requests through Istio, see [REST API Guide](../../user_guides/mlops/serving/rest-api.md).
!!! warning "Host-based routing"
The Istio Model Endpoint supports host-based routing for inference requests; however, this approach is considered legacy.
Path-based routing is recommended for new deployments.
Models deployed on KServe in Hopsworks can be easily integrated with the Hopsworks Feature Store using either a Transformer or Predictor Python script, that builds the predictor's input feature vector using the application input and pre-computed features from the Feature Store.
--8<-- "concepts/mlops/serving/model-serving.html"
## Deployment API
The deployment API is the interface to the online inference pipeline that clients send prediction requests to.
It is the deployment API, not the model signature, that clients should version against.
The model signature (the input and output schema of the model) changes whenever you retrain with a different set of features, so coupling clients to it turns every model update into a breaking change.
The deployment API is a stable contract that can stay the same across model versions.
A client request to the deployment API carries two kinds of parameter:
- **serving keys**: the entity IDs used to retrieve pre-computed features from the feature store.
- **request parameters**: values known only at request time, sent in the request and used to build the feature vector or to compute on-demand features.
Because clients depend on it, a deployment API should carry an SLO, typically a p99 latency target for online predictions.
## Testing model deployments
Two release-safety mechanisms are often confused, because they test different things.
A blue/green test tests the correctness and performance of the model deployment directly, running the new deployment alongside the old one so clients can be switched over with no risk.
An A/B test does not test the deployment; it tests the model's effect on the application, measured against an application KPI, to decide whether the new model actually makes the product better.
!!! info "Model Serving Guide"
More information can be found in the [Model Serving guide](../../user_guides/mlops/serving/index.md).
!!! tip "Python deployments"
For deploying Python scripts without a model artifact, see the [Python Deployments](../../user_guides/projects/python-deployment/python-deployment.md) page.
================================================================================
# Model Monitoring
Source: https://docs.hopsworks.ai/latest/concepts/mlops/model_monitoring/
# Model Monitoring
Model monitoring lets you track how a deployed model behaves in production by comparing the data it serves against the data it was trained on.
When a model runs in production, the statistical properties of its inputs and predictions can drift away from those of the training data.
This degrades model quality silently, without any error being raised.
Model monitoring detects this drift early so you can decide whether to retrain the model.
## How it works
Model monitoring builds on two existing Hopsworks capabilities:
- **Feature logging**: a model deployment logs the features it serves and its predictions to the feature view's logging feature group through the Feature View logging APIs.
See the [Feature Logging guide](../../user_guides/fs/feature_view/feature_logging.md).
- **Feature monitoring**: Hopsworks computes statistics over windows of feature data and compares them against a reference, optionally raising alerts on significant drift.
See the [Feature Monitoring concept](../fs/feature_group/feature_monitoring.md).
??? note "Log untransformed and transformed features"
Log both the untransformed and the transformed feature values.
Untransformed features drive feature monitoring and debugging, since drift is easiest to read on the raw values.
Transformed features drive model monitoring and SHAP explainability, since those are the values the model actually sees.
!!! info "Feature logging vs. the inference logger"
Hopsworks provides two separate inference logging mechanisms.
The [inference logger](../../user_guides/mlops/serving/inference-logger.md) stores the model inputs and predictions from inference requests and responses into Kafka, for later consumption and analysis.
[Feature logging](../../user_guides/fs/feature_view/feature_logging.md) supports more fine-grained logging of inference logs and features, enabling feature monitoring and model monitoring.
Model monitoring relies on feature logging, not on the inference logger.
A model monitoring configuration is a feature monitoring configuration over the logging feature group, filtered to a single model and version.
The detection window covers the recently served inference data, and the reference defaults to the training dataset version that was used to train that model.
By comparing the two, on a scalar metric or on the whole feature distribution, Hopsworks detects feature drift over time.
This comparison detects drift, not skew: offline-online feature skew is a difference in the transformation code between the offline and inference pipelines, so it is invisible to a distribution comparison and is prevented, not monitored.
Feature drift is one kind of drift among several.
Concept drift, in particular, is not detected by comparing distributions: you detect it by comparing the actual outcomes against the model's past predictions, once those outcomes are known.
--8<-- "concepts/mlops/model_monitoring/how-it-works.html"
## Where to configure it
Because monitoring is anchored on the feature view that backs the model, you can configure model monitoring from whichever entity is most convenient:
- a **model deployment**, when operating a model in production.
- a **model** in the model registry.
- a **feature view**, when working directly with the feature data.
All three resolve to the same underlying configuration.
!!! info "Model Monitoring Guide"
More information can be found in the [Model Monitoring guide](../../user_guides/mlops/model_monitoring/index.md).
================================================================================
# Vector Index
Source: https://docs.hopsworks.ai/latest/concepts/mlops/opensearch/
# Vector Index
A vector index stores embeddings so you can retrieve the items most similar to a query vector, the retrieval half of a recommender or a RAG system.
In Hopsworks, a vector index is a property of an online-enabled feature group: a feature group with an embedding column can be indexed for similarity search, alongside its online and offline stores.
The vector index is backed by OpenSearch, included as a multi-tenant service in projects.
OpenSearch provides the index through its k-NN plugin, which supports several engines for embedding indexes.
Hopsworks creates its indexes on the FAISS engine, which is the default from Hopsworks 5.1.
Earlier releases used the nmslib engine, which OpenSearch has deprecated and which does not accept the filter that Hopsworks 5.1 and later send inside the nearest-neighbor query, so the upgrade to Hopsworks 5.2 recreates those indexes on FAISS.
The [OpenSearch upgrade guide][upgrading-opensearch] describes what that upgrade involves.
Through Hopsworks, OpenSearch also provides enterprise capabilities, including authentication and access control to indexes (an index can be private to a Hopsworks project), filtering, scalability, high availability, and disaster recovery support.
To learn how OpenSearch powers vector similarity search in Hopsworks, you can see [this guide](../../user_guides/fs/vector_similarity_search.md).
--8<-- "concepts/mlops/opensearch/vector-index.html"
================================================================================
# Development Inside Hopsworks
Source: https://docs.hopsworks.ai/latest/concepts/dev/inside/
# Development Inside Hopsworks
Hopsworks provides a complete self-service development environment for feature engineering and model training.
You can develop programs as Jupyter notebooks or jobs, customize the bundled FTI (feature, training and inference pipeline) python environments, you can manage your source code with Git, and you can orchestrate jobs with Airflow.
A browser terminal runs inside the project with the Hopsworks CLI and coding agents preinstalled, and the Wizard uses it to build a system end to end from a few choices.
--8<-- "concepts/dev/inside/development-inside-hopsworks.html"
## Jupyter Notebooks
Hopsworks provides a Jupyter notebook development environment for programs written in Python, Spark, and SparkSQL.
You can also develop in your IDE (PyCharm, IntelliJ, etc), test locally, and then run your programs as Jobs in Hopsworks.
Jupyter notebooks can also be run as Jobs.
## Source Code Control
Hopsworks provides source code control support using Git (GitHub, GitLab or BitBucket).
You can securely check out code into your project and commit and push updates to your code to your source code repository.
## FTI Pipeline Environments
Hopsworks postulates that building ML systems following the FTI pipeline architecture is best practice.
This architecture consists of three independently developed and operated ML pipelines:
- Feature pipeline: takes as input raw data that it transforms into features (and labels)
- Training pipeline: takes as input features (and labels) and outputs a trained model
- Inference pipeline: takes new feature data and a trained model and makes predictions
In order to facilitate the development of these pipelines Hopsworks bundles several python environments containing necessary dependencies.
Each of these environments may then also be customized further by cloning it and installing additional dependencies from PyPi, Wheel files, GitHub repos or a custom Dockerfile.
Internal compute such as Jobs and Jupyter is run in one of these environments and changes are applied transparently when you install new libraries using our APIs.
That is, there is no need to write a Dockerfile, users install libraries directly in one or more of the environments.
You can setup custom development and production environments by creating separate projects or creating multiple clones of an environment within the same project.
## Jobs
In Hopsworks, a Job is a schedulable program that is allocated compute and memory resources.
You can run a Job in Hopsworks:
- From the UI
- Programmatically with the Hopsworks SDK (Python, Java) or REST API
- From Airflow programs (either inside our outside Hopsworks)
- From your IDE using a plugin ([PyCharm/IntelliJ plugin](https://plugins.jetbrains.com/plugin/15537-hopsworks))
## Orchestration
Airflow comes out-of-the box with Hopsworks, but you can also use an external Airflow cluster (with the Hopsworks Job operator) if you have one.
Airflow can be used to schedule the execution of Jobs, individually or as part of Airflow DAGs.
## Terminal { #inside-terminal }
Every project has a browser terminal: a shell running in a pod under your project user, with your HopsFS home mounted, and `hops`, `git`, Claude Code and Codex preinstalled and already connected to the project.
It is the fastest way to work with a project from inside Hopsworks, and the seat the Wizard drives.
See the [Terminal guide][terminal] and the [Hopsworks CLI guide][hopsworks-cli].
## Wizard { #inside-wizard }
The Wizard asks what you want to build, where the data comes from and what to predict, then writes a kickoff prompt and hands it to Claude in the terminal, which builds the feature pipeline, the model and the dashboard with `hops`.
See the [Wizard guide][wizard].
================================================================================
# Development Outside Hopsworks
Source: https://docs.hopsworks.ai/latest/concepts/dev/outside/
# Development Outside Hopsworks
You can write programs that use Hopsworks in any [Python, Spark, or PySpark environment](../../user_guides/integrations/index.md).
Hopsworks also supports running SQL queries to compute features in external data warehouses.
The Feature Store can also be queried with SQL.
There is REST API for Hopsworks that can be used with a valid API key, generated in Hopsworks.
However, it is often easier to develop your programs against the Hopsworks SDK, available in Python and Java/Scala, which covers the feature store, the model registry and model serving.
The same library ships the `hops` command line, so a shell, a CI pipeline or a coding agent on your machine can read and write the project with the same API key; see the [Hopsworks CLI guide][hopsworks-cli].
--8<-- "concepts/dev/outside/development-outside-hopsworks.html"
================================================================================
# BI Tools
Source: https://docs.hopsworks.ai/latest/concepts/mlops/bi_tools/
# BI Tools
Feature groups have well-defined schemas and live in two stores, so any BI tool that speaks SQL can analyze features and build reports on them.
- The offline store is queried through the [Query Engine](../../user_guides/projects/trino/query_engine.md), Trino, with one catalog per table format (`delta`, `iceberg`, `hudi`).
Any tool with a Trino connector (JDBC or ODBC) can read it.
- The online store, RonDB, is queried over the MySQL protocol, so any tool with a MySQL connector can read the latest feature values.
Hopsworks bundles [Apache Superset](https://superset.apache.org/) as a project service, already connected to the project's Trino catalogs.
Dashboards live inside the project and follow its access control.
Superset dashboards inside a project, with their public and shared status.
SQL Lab in Superset runs directly against the feature store: pick the project's Trino connection, then the catalog matching the feature group's table format.
SQL Lab on the project's Trino connection, one catalog per table format.
See the [Superset guide](../../user_guides/projects/superset/superset.md) for building dashboards on feature data, and the [Superset setup](../../setup_installation/admin/superset.md) page for enabling it on a cluster.
================================================================================
# How-To Guides
Source: https://docs.hopsworks.ai/latest/user_guides/
# How-To Guides
Task-focused guides for the Hopsworks UI and APIs, organised by the part of the platform you are working with.
For what things are and why, see the [Concepts](../concepts/index.md).
- :material-console:{ .lg .middle } **Start here**
---
Install the client and authenticate once.
Every guide in this section runs from the same session.
```bash
uv venv && source .venv/bin/activate
uv pip install "hopsworks[python]"
hops setup
```
[Client installation](client_installation/index.md) · [Create a project](projects/project/create_project.md) · [Create a feature group](fs/feature_group/create.md)
:material-database:{ .hops-role-ico } Feature Store
{ .hops-role-cap }
- [Feature groups](fs/feature_group/index.md)
Write features from a DataFrame, validate them, keep statistics.
- [Feature views](fs/feature_view/index.md)
Read training data, batch data and online feature vectors.
- [Data sources](fs/data_source/index.md)
Connect warehouses, object stores and databases as inputs.
- [Feature monitoring](fs/feature_monitoring/index.md)
Watch statistics over time and compare them to a reference.
- [Transformations and integrations](fs/transformation_functions.md)
Model-independent transformations, compute engines, external clients.
:material-rocket-launch-outline:{ .hops-role-ico } MLOps
{ .hops-role-cap }
- [Model registry](mlops/registry/index.md)
Register models with metrics, schema and evaluation artifacts.
- [Model serving](mlops/serving/index.md)
Deploy a model with a predictor, transformer, logging and autoscaling.
- [Model monitoring](mlops/model_monitoring/index.md)
Compare inference data against training data on a schedule.
- [Agents](agents/index.md)
Run agent tasks as jobs or serve interactive agents.
:material-folder-outline:{ .hops-role-ico } Projects and compute
{ .hops-role-cap }
- [Projects](projects/index.md)
Sign in, create a project, manage members, secrets, keys and alerts.
- [Compute](compute/index.md)
Jupyter, the terminal, jobs, Airflow and Python environments.
- [Analytics](analytics/index.md)
SQL over the offline store with Trino, dashboards in Superset.
:material-wrench-outline:{ .hops-role-ico } Platform
{ .hops-role-cap }
- [Clients](client_installation/index.md)
Python and Java libraries for your own environment, and the `hops` CLI.
- [Setup and administration](../setup_installation/index.md)
Install on a cloud or on-prem, manage users and operations.
- [Migration 3.x to 4.0](migration/40_migration.md)
What changed and how to move.
================================================================================
# Client Installation Guide
Source: https://docs.hopsworks.ai/latest/user_guides/client_installation/
# Client Installation Guide
## Hopsworks Python library
The Hopsworks Python client library is required to connect to Hopsworks from your local machine or any other Python environment such as Google Colab or AWS Sagemaker.
Execute the following command to install the Hopsworks client library in your Python environment:
!!! note "Virtual environment"
It is recommended to use a virtual python environment instead of the system environment used by your operating system, in order to avoid any side effects regarding interfering dependencies.
!!! attention "Windows/Conda Installation"
On Windows systems you might need to install twofish manually before installing hopsworks, if you don't have the Microsoft Visual C++ Build Tools installed.
In that case, it is recommended to use a conda environment and run the following commands:
```bash
conda install twofish
pip install hopsworks[python]
```
=== "uv"
```bash
uv venv && source .venv/bin/activate
uv pip install "hopsworks[python]"
```
=== "pip"
```bash
python3 -m venv .venv && source .venv/bin/activate
pip install "hopsworks[python]"
```
Supported versions of Python: 3.10, 3.11, 3.12, 3.13, 3.14 ([PyPI ↗](https://pypi.org/project/hopsworks/))
### Profiles
The Hopsworks library has several profiles that bring additional dependencies and enable additional functionalities:
| Profile Name | Description |
| --- | --- |
| No Profile | This is the base installation. Supports interacting with the feature store metadata, model registry and deployments. It also supports reading and writing from the feature store from PySpark environments. |
| `python` | This profile enables reading and writing from/to the feature store from a Python environment |
| `great-expectations` | Installs [Great Expectations](https://greatexpectations.io/) and enables data validation on feature pipelines. Supports 0.18.12 and 1.17.1; 1.17.1 is recommended |
| `polars` | This profile installs the [Polars](https://pola.rs/) library and enables reading and writing Polars DataFrames |
You can install all the above profiles with the following command:
```bash
uv pip install "hopsworks[python,great-expectations,polars]"
```
## Skills and instructions for coding agents
The Hopsworks Python library ships a set of skills for coding agents: Claude Code, Codex, GitHub Copilot and OpenCode.
Inside a Hopsworks terminal they are available to every agent automatically.
On your own machine, two commands make them available in the repository you are working in.
### Authenticate and write the agent instructions
```bash
uv pip install "hopsworks[python]"
cd
hops setup --host https://
```
Without `--host`, `hops setup` asks for the host and proposes `https://eu-west.cloud.hopsworks.ai`, the Hopsworks serverless endpoint; press Enter to accept it or type the address of your cluster.
`hops setup` opens a browser page where you choose a project, creates an API key for it, and stores the key in `~/.hops.toml`.
It then writes the following files into the current directory:
| Path | Purpose |
| --- | --- |
| `AGENTS.md` | Instructions for the agent: the project you are connected to, where the `hopsworks` library is installed on this machine, and how to use the `hops` CLI and the skills. |
| `.claude/skills/hops/SKILL.md` | A reference for the `hops` CLI. |
| `.claude/commands/hops.md` | The `/hops` slash command for Claude Code. |
| `.claude/agents/hops-fti.md` | A Claude Code sub-agent that reviews a project against the feature, training and inference pipeline pattern. |
| `.claude/settings.local.json` | Allows `Bash(hops *)`, so Claude Code can run the CLI without asking before each command. |
`AGENTS.md` is read by Claude Code, Codex, GitHub Copilot and OpenCode.
The files under `.claude/` are read by Claude Code only.
Running `hops setup` again in a directory that already has these files updates the files you have not edited and leaves the ones you have edited unchanged.
Pass `--no-scaffold` to authenticate without writing any files.
### Add the Hopsworks skills
```bash
hops skills install
```
`hops skills install` copies the Hopsworks skills into `.claude/skills/`, one directory per skill, which is where Claude Code discovers them.
For another agent, pass `--agent`, which can be repeated:
```bash
hops skills install --agent codex
hops skills install --agent copilot
hops skills install --agent opencode
```
The skills are written to `.codex/skills/`, `.agents/skills/` and `.opencode/skills/` respectively, and for OpenCode the path is also registered in `opencode.json`.
An agent loads only the name and description of each skill when it starts and reads a skill in full when a task calls for it, so adding all of them costs a few kilobytes of context rather than the size of the skills themselves.
Running `hops skills install` again after upgrading the `hopsworks` library updates the skills you have not edited, keeps the skills you have edited, and removes skills that the new version no longer ships.
Pass `--force` to overwrite edited skills as well.
To read the skills without adding them to a repository:
```bash
hops skills list
hops skills show hops-fg
```
## Hopsworks Java Library
If you want to interact with the Hopsworks Feature Store from environments such as Spark or Beam, you can use the Hopsworks Feature Store (Hopsworks) Java library.
!!! note "Feature Store Only"
The Java library only allows interaction with the Feature Store component of the Hopsworks platform.
Additionally each environment might restrict the supported API operation.
You can see which API operation is supported by which environment [here](../fs/compute_engines.md)
The Hopsworks library is available on the Hopsworks' Maven repository.
If you are using Maven as build tool, you can add the following in your `pom.xml` file:
```xml
HopsHops Repositoryhttps://archiva.hops.works/repository/Hops/truetrue
```
The library has different builds targeting different environments:
### Hopsworks Java
The `artifactId` for the Hopsworks Java build is `hsfs`, if you are using Maven as build tool, you can add the following dependency:
```xml
com.logicalclockshsfs${hsfs.version}
```
### Spark
The `artifactId` for the Spark build is `hsfs-spark-spark{spark.version}`, if you are using Maven as build tool, you can add the following dependency:
```xml
com.logicalclockshsfs-spark-spark3.1${hsfs.version}
```
Hopsworks provides builds for Spark 3.1, 3.3 and 3.5. The builds are also provided as JAR files which can be downloaded from [Hopsworks repository](https://repo.hops.works/master/hsfs)
### Beam
The `artifactId` for the Beam build is `hsfs-beam`, if you are using Maven as build tool, you can add the following dependency:
```xml
com.logicalclockshsfs-beam${hsfs.version}
```
## Next Steps
If you are using a local python environment and want to connect to Hopsworks, you can follow the [Python Guide](../integrations/python.md#generate-an-api-key) section to create an API Key and to get started.
If you use a coding agent, see [Skills and instructions for coding agents][skills-and-instructions-for-coding-agents] to give it the Hopsworks skills.
## Other environments
The Hopsworks Feature Store client libraries can also be installed in external environments, such as Databricks, AWS Sagemaker, or Azure Machine Learning.
For more information, see [Client Integrations](../integrations/index.md).
================================================================================
# Hopsworks CLI
Source: https://docs.hopsworks.ai/latest/user_guides/client_installation/cli/
# Hopsworks CLI
`hops` is the Hopsworks command line.
It ships with the Hopsworks Python library, so anywhere the library is installed the command is available.
The same commands run from your laptop, from a CI pipeline, from a coding agent, or inside the [project terminal][terminal], where it is already connected.
--8<-- "user_guides/client_installation/cli/one-cli-two-seats.html"
## Install and connect
Install the library with the `python` profile and the CLI comes with it.
The profile brings the Arrow and Kafka dependencies the data commands (`fg preview`, `fv get`, `sql`) read and write through.
=== "uv"
```bash
uv venv && source .venv/bin/activate
uv pip install "hopsworks[python]"
```
=== "pip"
```bash
python3 -m venv .venv && source .venv/bin/activate
pip install "hopsworks[python]"
```
On your own machine, `hops setup` opens a browser, lets you pick a project, creates an API key and caches it in `~/.hops.toml`:
```bash
hops setup --host https://my.hopsworks.ai
```
Non-interactive environments such as CI use an existing API key instead:
```bash
hops login --host https://my.hopsworks.ai --api-key "$HOPSWORKS_API_KEY" --project fraud_detection
```
Inside the [project terminal][terminal] no login is needed, `hops` is pointed at the project you opened it from.
## Explore and use the feature store
Every asset type has a subcommand: `project`, `fg`, `fv`, `td`, `model`, `deployment`, `job`, `datasource`, `sql`.
`hops --help` lists the verbs.
```bash
hops project use fraud_detection
hops fg list
hops fg info transactions
hops fg preview transactions --n 5
hops fv create transactions_fraud --feature-group transactions
hops fv get transactions_fraud --entry "cc_num=4532015112830366"
hops sql "select count(*) from transactions_1"
```
Add `--json` to any command for machine-readable output:
```bash
hops fg info transactions --json
```
```json
{
"id": 1080,
"name": "transactions",
"version": 1,
"type": "cached",
"online_enabled": false,
"primary_key": ["tid"],
"event_time": "datetime",
"features": [
{"name": "tid", "type": "bigint", "primary": true}
]
}
```
## For coding agents
The CLI is the simplest way to give an agent access to Hopsworks: allow it to run `hops` and it can read and write the project.
`hops init` scaffolds the Hopsworks skill, slash command and sub-agent for Claude Code into a repository and allows `Bash(hops *)` there:
```bash
hops init --dir .
```
`hops skills list` shows the Hopsworks skills the agent can load, feature groups, feature views, training, online inference, monitoring and more.
The [project terminal][terminal] has Claude Code and Codex preinstalled with `hops` already connected, and the [Wizard][wizard] uses exactly this path to build a system end to end.
================================================================================
# Feature Store Guides
Source: https://docs.hopsworks.ai/latest/user_guides/fs/
# Feature Store Guides
Feature pipelines write to feature groups, training and inference pipelines read through feature views.
The guides below follow that order: connect a source, write, validate, read, transform.
- :material-database-plus-outline:{ .lg .middle } **Start here**
---
Create a feature group from a DataFrame and insert it.
Everything else in this section builds on a feature group that exists.
```python
fg = fs.get_or_create_feature_group(
name="transactions",
version=1,
primary_key=["tid"],
event_time="datetime",
online_enabled=True,
)
fg.insert(df)
```
[Create a feature group](feature_group/create.md) · [Create a feature view](feature_view/overview.md) · [Training data](feature_view/training-data.md)
:material-database-import-outline:{ .hops-role-ico } Write
{ .hops-role-cap }
- [Data sources](data_source/index.md)
Register warehouses, object stores and databases to read from and write to.
- [Feature groups](feature_group/index.md)
Create, insert, evolve the schema, set time to live, deprecate.
- [External and spine groups](feature_group/create_external.md)
Point at data that stays where it is, or supply keys and labels without storing them.
- [Ingest with dltHub](feature_group/ingest_with_dlthub.md)
Load from hundreds of sources through dlt pipelines.
:material-check-decagram-outline:{ .hops-role-ico } Trust
{ .hops-role-cap }
- [Statistics](feature_group/statistics.md)
Descriptive statistics on every insert, configurable per group.
- [Data validation](feature_group/data_validation.md)
Great Expectations suites run on insert, with a policy on failure.
- [Feature monitoring](feature_monitoring/index.md)
Scheduled statistics and drift detection against a reference window.
- [Notifications and observability](feature_group/notification.md)
Change notifications and online ingestion status.
:material-database-export-outline:{ .hops-role-ico } Read
{ .hops-role-cap }
- [Feature views](feature_view/index.md)
Select features across groups and read them the same way for training and inference.
- [Training data](feature_view/training-data.md)
Materialise splits as files or read them straight into memory.
- [Batch and online reads](feature_view/batch-data.md)
Batch inference data by time range, single vectors from the online store.
- [Feature server](feature_view/feature-server.md)
Serve feature vectors over REST without the Python client.
:material-function-variant:{ .hops-role-ico } Transform and run
{ .hops-role-cap }
- [Transformation functions](transformation_functions.md)
Model-independent functions applied on write, model-dependent on read.
- [Compute engines](compute_engines.md)
Which operations run on Python, Spark or Flink.
- [Client integrations](../integrations/index.md)
Databricks, SageMaker, EMR, Azure ML, Flink, Beam and more.
- [Vector similarity search](vector_similarity_search.md)
Embeddings in a feature group, nearest-neighbour queries.
================================================================================
# Data Source Guides
Source: https://docs.hopsworks.ai/latest/user_guides/fs/data_source/
# Data Source Guides
You can define data sources in Hopsworks for batch and streaming data sources.
Data Sources securely store the authentication information about how to connect to an external data store.
They can be used from programs within Hopsworks or externally.
!!!warning
In the previous versions of Hopsworks, this used to be called a storage connector.
There are four main use cases for Data Sources:
- Simply use it to read data from the storage into a dataframe.
- [External (on-demand) Feature Groups](../../../concepts/fs/feature_group/external_fg.md) can be defined with data sources.
This way, Hopsworks stores only the metadata about the features, but does not keep a copy of the data itself.
This is also called the Connector API.
- Write [training data](../../../concepts/fs/feature_view/offline_api.md) to an external storage system to make it accessible by third parties.
- Managed [feature group](../../../user_guides/fs/feature_group/create.md) that stores offline data in an external storage system.
Currently [S3](../data_source/creation/s3.md), [GCS](../data_source/creation/gcs.md) and [AWS Glue](../data_source/creation/glue.md) connectors are supported.
Data Sources provide two main mechanisms for authentication: using credentials or an authentication role (IAM Role on AWS or Managed Identity on Azure).
Hopsworks supports both a single IAM role (AWS) or Managed Identity (Azure) for the whole Hopsworks cluster or multiple IAM roles (AWS) or Managed Identities (Azure) that can only be assumed by users with a specific role in a specific project.
By default, each project is created with three default Data Sources: A JDBC connector to the online feature store, a HopsFS connector to the Training Datasets directory of the project and a JDBC connector to the offline feature store.

The Data Source View in the User Interface
## Cloud Agnostic
Cloud agnostic storage systems:
- :simple-snowflake:{ .lg .middle style="color:#29B5E8" } **Snowflake**
---
Query Snowflake databases and tables using SQL.
[:octicons-arrow-right-24: Configure](creation/snowflake.md)
- :simple-apachekafka:{ .lg .middle } **Kafka**
---
Read from a Kafka cluster into a Spark Structured Streaming Dataframe.
[:octicons-arrow-right-24: Configure](creation/kafka.md)
- :simple-sap:{ .lg .middle style="color:#0FAAFF" } **SAP HANA**
---
Query SAP HANA tenant databases using SQL.
[:octicons-arrow-right-24: Configure][data-source-sap-hana]
- :material-database:{ .lg .middle style="color:var(--hops-accent-text)" } **JDBC**
---
Connect to any JDBC compatible database and query it using SQL.
[:octicons-arrow-right-24: Configure](creation/jdbc.md)
- :material-api:{ .lg .middle style="color:var(--hops-accent-text)" } **REST API**
---
Connect to external HTTP APIs with configurable headers and authentication.
[:octicons-arrow-right-24: Configure](creation/rest_api.md)
- :material-chart-box-outline:{ .lg .middle style="color:var(--hops-accent-text)" } **CRM, Sales & Analytics**
---
Connect to supported CRM, sales, and analytics platforms.
[:octicons-arrow-right-24: Configure](creation/crm_sales_analytics.md)
- :material-folder-network-outline:{ .lg .middle style="color:var(--hops-accent-text)" } **HopsFS**
---
Connect and read from directories of Hopsworks' internal file system.
[:octicons-arrow-right-24: Configure](creation/hopsfs.md)
## AWS
For AWS the following storage systems are supported:
- :fontawesome-brands-aws:{ .lg .middle style="color:#FF9900" } **S3**
---
Read file-based storage in S3 such as parquet or CSV.
[:octicons-arrow-right-24: Configure](creation/s3.md)
- :fontawesome-brands-aws:{ .lg .middle style="color:#FF9900" } **AWS Glue**
---
Integrate with the Glue Data Catalog over S3, for Iceberg, Delta, Hudi and plain files.
[:octicons-arrow-right-24: Configure](creation/glue.md)
- :fontawesome-brands-aws:{ .lg .middle style="color:#FF9900" } **Redshift**
---
Query Redshift databases and tables using SQL.
[:octicons-arrow-right-24: Configure](creation/redshift.md)
- :fontawesome-brands-aws:{ .lg .middle style="color:#FF9900" } **RDS (SQL)**
---
Query the Amazon Relational Database Service using SQL.
[:octicons-arrow-right-24: Configure](creation/sql.md)
## Azure
For Azure the following storage systems are supported:
- :material-microsoft-azure:{ .lg .middle style="color:#0078D4" } **ADLS**
---
Read file-based storage in ADLS such as parquet or CSV.
[:octicons-arrow-right-24: Configure](creation/adls.md)
## GCP
For GCP the following storage systems are supported:
- :simple-googlebigquery:{ .lg .middle style="color:#4285F4" } **BigQuery**
---
Query BigQuery databases and tables using SQL.
[:octicons-arrow-right-24: Configure](creation/bigquery.md)
- :simple-googlecloudstorage:{ .lg .middle style="color:#4285F4" } **GCS**
---
Read file-based storage in Google Cloud Storage such as parquet or CSV.
[:octicons-arrow-right-24: Configure](creation/gcs.md)
## Databricks (AWS only)
For Databricks **on AWS** the following storage systems are supported:
- :simple-databricks:{ .lg .middle style="color:#FF3621" } **Unity Catalog**
---
Browse catalogs, schemas, and Delta tables, and mount them as external feature groups.
[:octicons-arrow-right-24: Configure](creation/unity_catalog.md)
Databricks on Azure and Databricks on GCP are not supported yet. See the [Unity Catalog guide](creation/unity_catalog.md) for the specific reasons and the status of follow-up work.
## Next Steps
Move on to the [Configuration and Creation Guides](creation/jdbc.md) to learn how to set up a data source.
================================================================================
# JDBC
Source: https://docs.hopsworks.ai/latest/user_guides/fs/data_source/creation/jdbc/
# How-To set up a JDBC Data Source
## Introduction
JDBC is an API provided by many database systems.
Using JDBC connections one can query and update data in a database, usually oriented towards relational databases.
Examples of databases you can connect to using JDBC are MySQL, Postgres, Oracle, DB2, MongoDB or Microsoft SQLServer.
In this guide, you will configure a Data Source in Hopsworks to save all the authentication information needed in order to set up a JDBC connection to your database of choice.
When you're finished, you'll be able to query the database using Spark through Hopsworks APIs.
!!! note
Currently, it is only possible to create data sources in the Hopsworks UI.
You cannot create a data source programmatically.
## Prerequisites
Before you begin this guide you'll need to retrieve the following information from your JDBC compatible database:
- **JDBC Connection URL:** Consult the documentation of your target database to determine the correct JDBC URL and parameters.
As an example, for MySQL the URL could be:
```plaintext
jdbc:mysql://10.0.2.15:3306/[databaseName]?useSSL=false&allowPublicKeyRetrieval=true
```
- **Username and Password:** Typically, you will need to add username and password in your JDBC URL or as key/value parameters.
So make sure you have retrieved a username and password with the suitable permissions for the database and table you want to query.
## Creation in the UI
### Step 1: Set up new Data Source
Head to the `Data Sources` view on Hopsworks (1) and click `New data source` (2).
The `Add a data source` catalog lists the available sources, grouped under `Object storage`, `Data warehouse`, `Database`, `Streaming` and `API & SaaS`.
Pick the `JDBC` card to open the creation form.

The Data Source View in the User Interface
### Step 2: Enter JDBC Settings
Enter the details for your JDBC enabled database.

JDBC Connector Creation Form
1. The form opens with `Source` set to `JDBC`.
Click `Change source` to pick a different one.
2. Enter the JDBC connection url.
This can for example also contain the username and password.
3. Add additional key/value arguments to be passed to the connection, such as username or password.
These might differ by database.
!!! note
Driver class name is a mandatory argument even if using the default MySQL driver.
Add it by specifying a property with the name `driver` and class name as value.
The driver class name will differ based on the database.
For MySQL databases, the class name is `com.mysql.cj.jdbc.Driver`, as shown in the example image.
4. Click on "Save Credentials".
!!! note
To be able to use the connector, you need to upload the driver JAR file to the [Jupyter configuration](../../../projects/jupyter/spark_notebook.md) or [Job configuration](../../../projects/jobs/pyspark_job.md) in `Additional Jars`.
For MySQL connections the default JDBC driver is already included in Hopsworks so this step can be skipped.
## Next Steps
Move on to the [usage guide for data sources](../usage.md) to see how you can use your newly created JDBC connector.
================================================================================
# Snowflake
Source: https://docs.hopsworks.ai/latest/user_guides/fs/data_source/creation/snowflake/
# How-To set up a Snowflake Data Source
## Introduction
Snowflake provides a cloud-based data storage and analytics service, used as a data warehouse in many enterprises.
Data warehouses are often the source of raw data for feature engineering pipelines and Snowflake supports scalable feature computation with SQL.
However, Snowflake is not viable as an online feature store that serves features to models in production, with its columnar database layout its latency is too high compared to OLTP databases or key-value stores.
In this guide, you will configure a Data Source in Hopsworks to save all the authentication information needed in order to set up a connection to your Snowflake database.
When you're finished, you'll be able to query the database using Spark through Hopsworks APIs.
!!! note
Currently, it is only possible to create data sources in the Hopsworks UI.
You cannot create a data source programmatically.
## Prerequisites
Before you begin this guide you'll need to retrieve the following information from your Snowflake account and database, the following options are **mandatory**:
- **Snowflake Connection URL:** Consult the documentation of your target snowflake account to determine the correct connection URL.
This is usually some form of your [Snowflake account identifier](https://docs.snowflake.com/en/user-guide/admin-account-identifier.html).
For example:
```plaintext
.snowflakecomputing.com
```
OR:
```plaintext
https://-.snowflakecomputing.com
```
The account and organization details can be viewed in the Snowsight UI under **Admin > Account** or by querying it in
SQL, as explained in [Snowflake
documentation](https://docs.snowflake.com/en/user-guide/organizations-gs.html#viewing-the-name-of-your-organization-and-its-accounts).
Below is an example of how to view the account and organization to get the account identifier from the Snowsight UI.

Viewing Snowflake account identifier
!!! note "Authentication methods"
The Snowflake data source supports username and password, token-based and key-pair based authentication options. General information on snowflake key pair authentication and setup is at [Snowflake key-pair authentication](https://docs.snowflake.com/en/user-guide/key-pair-auth).
- **Username and Password:** Login name for the Snowflake user and password.
This is often also referred to as `sfUser` and `sfPassword`.
- **Warehouse:** The warehouse to use for the session after connecting
- **Database:** The database to use for the session after connecting.
- **Schema:** The schema to use for the session after connecting.
These are a few additional **optional** arguments:
- **Role:** The role field can be used to specify which [Snowflake security role](https://docs.snowflake.com/en/user-guide/security-access-control-overview.html#system-defined-roles) to assume for the session after the connection is established.
- **Application:** The application field can also be specified to have better observability in Snowflake with regards to which application is running which query.
The application field can be a simple string like “Hopsworks” or, for instance, the project name, to track usage and queries from each Hopsworks project.
## Creation in the UI
### Step 1: Set up new Data Source
Head to the `Data Sources` view on Hopsworks (1) and click `New data source` (2).
The `Add a data source` catalog lists the available sources, grouped under `Object storage`, `Data warehouse`, `Database`, `Streaming` and `API & SaaS`.
Pick the `Snowflake` card to open the creation form.

The Data Source View in the User Interface
### Step 2: Enter Snowflake Settings
Enter the details for your Snowflake connector.
Start by giving it a **name** and an optional **description**.
01. The form opens with `Source` set to `Snowflake`.
Click `Change source` to pick a different one.
02. Specify the hostname for your account in the following format `.snowflakecomputing.com` or `https://-.snowflakecomputing.com`.
03. Login name for the Snowflake user.
04. **Authentication** Choose between user account Password, Token or Private Key options.
In case of private key, upload your snowflake user Private Key file and set Passphrase if applicable.
05. The warehouse to connect to.
06. The database to use for the connection.
07. Add any additional optional arguments.
For example, you can specify `Schema`, `Table`, `Role`, and `Application`.
08. Optional additional key/value arguments.
09. Click on "Save Credentials".

Snowflake Connector Creation Form
## Next Steps
Move on to the [usage guide for data sources](../usage.md) to see how you can use your newly created Snowflake connector.
================================================================================
# Kafka
Source: https://docs.hopsworks.ai/latest/user_guides/fs/data_source/creation/kafka/
# How-To set up a Kafka Data Source
## Introduction
Apache Kafka is a distributed event store and stream-processing platform.
It's a very popular framework for handling realtime data streams and is often used as a message broker for events coming from production systems until they are being processed and either loaded into a data warehouse or aggregated into features for Machine Learning.
In this guide, you will configure a Data Source in Hopsworks to save all the authentication information needed in order to set up a connection to your Kafka cluster.
When you're finished, you'll be able to read from Kafka topics in your cluster using Spark through Hopsworks APIs.
!!! note
Currently, it is only possible to create data sources in the Hopsworks UI.
You cannot create a data source programmatically.
## Prerequisites
Before you begin this guide you'll need to retrieve the following information from Kafka cluster, the following options are **mandatory**:
- **Kafka Bootstrap servers:** It is the url of one of the Kafka brokers which you give to fetch the initial metadata about your Kafka cluster.
The metadata consists of the topics, their partitions, the leader brokers for those partitions etc.
Depending upon this metadata your producer or consumer produces or consumes the data.
- **Security Protocol:** The security protocol you want to use to authenticate with your Kafka cluster.
Make sure the chosen protocol is supported by your cluster.
For an overview of the available protocols, please see the [Confluent Kafka Documentation](https://docs.confluent.io/platform/current/kafka/overview-authentication-methods.html).
- **Certificates:** Depending on the chosen security protocol, you might need TrustStore and KeyStore files along with the corresponding key password.
Contact your Kafka administrator, if you don't know how to retrieve these.
If you want to setup a data source to Hopsworks' internal Kafka cluster, you can download the needed certificates from the integration tab in your project settings.
## Creation in the UI
### Step 1: Set up new Data Source
Head to the `Data Sources` view on Hopsworks (1) and click `New data source` (2).
The `Add a data source` catalog lists the available sources, grouped under `Object storage`, `Data warehouse`, `Database`, `Streaming` and `API & SaaS`.
Pick the `Kafka` card to open the creation form.

The Data Source View in the User Interface
### Step 2: Enter Kafka Settings
Enter the details for your Kafka connector.
Start by giving it a **name** and an optional **description**.
01. The form opens with `Source` set to `Kafka`.
Click `Change source` to pick a different one.
02. Add all the bootstrap server addresses and ports that you want the consumers/producers to connect to.
The client will make use of all servers irrespective of which servers are specified here for bootstrapping.
This list only impacts the initial hosts used to discover the full set of servers.
03. Choose the Security protocol.
!!! example "TSL/SSL"
By default, Apache Kafka communicates in `PLAINTEXT`, which means that all data is sent in the clear.
To encrypt communication, you should configure all the Confluent Platform components in your deployment to use TLS/SSL encryption.
TLS uses private-key/certificate pairs, which are used during the TLS handshake process.
Each broker needs its own private-key/certificate pair, and the client uses the certificate to authenticate the broker.
Each logical client needs a private-key/certificate pair if client authentication is enabled, and the broker uses the certificate to authenticate the client.
These are provided in the form of *TrustStore* and *KeyStore* `JKS` files together with a key password.
For more information, refer to the official [Apacha Kafka Guide for TSL/SSL authentication](https://docs.confluent.io/platform/current/kafka/authentication_ssl.html).
!!! example "SASL SSL or SASL plaintext"
Apache Kafka brokers support client authentication using SASL.
SASL authentication can be enabled concurrently with TLS/SSL encryption (TLS/SSL client authentication will be disabled).
This authentication method often requires extra arguments depending on your setup.
Make use of the optional additional key/value arguments (5) to provide these.
SASL authentication can be enabled concurrently with TLS/SSL encryption (TLS/SSL client authentication will be disabled).
For more information, please refer to the official [Apache Kafka Guide for SASL authentication](https://docs.confluent.io/platform/current/kafka/authentication_sasl/index.html).
04. The endpoint identification algorithm used by clients to validate server host name.
The default value is `https`.
Clients including client connections created by the broker for inter-broker communication verify that the broker host name matches the host name in the broker’s certificate.
05. Optional additional key/value arguments.
06. Click on "Save Credentials".

Kafka Connector Creation Form
## Next Steps
Move on to the [usage guide for Data Sources](../usage.md) to see how you can use your newly created Kafka connector.
================================================================================
# HopsFS
Source: https://docs.hopsworks.ai/latest/user_guides/fs/data_source/creation/hopsfs/
# How-To set up a HopsFS Data Source
## Introduction
HopsFS is a HDFS-compatible filesystem on AWS/Azure/on-premises for data analytics.
HopsFS stores its data on object storage in the cloud (S3 in AWs and Blob storage on Azure) and on commodity servers on-premises, ensuring low-cost storage, high availability, and disaster recovery.
In Hopsworks, you can access HopsFS natively in programs (Spark, TensorFlow, etc) without the need to define a Data Source.
By default, every Project has a Data Source for Training Datasets.
When you create training datasets from features in the Feature Store the HopsFS connector is the default Data Source.
However, if you want to output data to a different dataset, you can define a new Data Source for that dataset.
In this guide, you will configure a HopsFS Data Source in Hopsworks which points at a different directory on the file system than the Training Datasets directory.
When you're finished, you'll be able to write training data to different locations in your cluster through Hopsworks APIs.
!!! note
Currently, it is only possible to create data sources in the Hopsworks UI.
You cannot create a data source programmatically.
## Prerequisites
Before you begin this guide you'll need to identify a **directory on the filesystem** of Hopsworks, to which you want to point the Data Source that you are going to create.
## Creation in the UI
### Step 1: Set up new Data Source
Head to the `Data Sources` view on Hopsworks (1) and click `New data source` (2).
The `Add a data source` catalog opens below.
Pick the `HopsFS` card to open the creation form.

The Data Source View in the User Interface
### Step 2: Enter HopsFS Settings
Enter the details for your HopsFS connector.
Start by giving it a **name** and an optional **description**.
1. The form opens with `Source` set to `HopsFS`.
Click `Change source` to pick a different one.
2. Select the top-level dataset to point the connector to.
3. Click on "Save Credentials".

HopsFS Connector Creation Form
## Next Steps
Move on to the [usage guide for data sources](../usage.md) to see how you can use your newly created HopsFS connector.
================================================================================
# S3
Source: https://docs.hopsworks.ai/latest/user_guides/fs/data_source/creation/s3/
# How-To set up a S3 Data Source { #data-source-s3 }
## Introduction
Amazon S3 or Amazon Simple Storage Service is a service offered by AWS that provides object storage.
That means you can store arbitrary objects associated with a key.
These kinds of storage systems are often used as Data Lakes with large volumes of unstructured data or file based storage.
Popular file formats are `CSV` or `PARQUET`.
There are so called Data Lake House technologies such as Delta Lake or Apache Hudi, building an additional layer on top of object based storage with files, to provide database semantics like ACID transactions among others.
This has the advantage that cheap storage can be turned into a cloud native data warehouse.
These kinds of storages are often the source for raw data from which features can be engineered.
In this guide, you will configure a Data Source in Hopsworks to save all the authentication information needed in order to set up a connection to your AWS S3 bucket.
When you're finished, you'll be able to read files using Spark through Hopsworks APIs.
You can also use the connector to write out training data from the Feature Store, in order to make it accessible by third parties.
!!! note
Currently, it is only possible to create data sources in the Hopsworks UI.
You cannot create a data source programmatically.
## Prerequisites
Before you begin this guide you'll need to retrieve the following information from your AWS S3 account and bucket:
- **Bucket:** You will need a S3 bucket that you have access to.
The bucket is identified by its name.
- **Path (Optional):** If needed, a path can be defined to ensure that all operations are restricted to a specific location within the bucket.
- **Region (Optional):** You will need an S3 region to have complete control over data when managing the feature group that relies on this data source.
The region is identified by its code.
- **Authentication Method:** You can authenticate using Access Key/Secret, or use IAM roles.
If you want to use an IAM role it either needs to be attached to the entire Hopsworks cluster or Hopsworks needs to be able to assume the role.
See [IAM role documentation](../../../../setup_installation/admin/roleChaining.md) for more information.
- **Server Side Encryption details:** If your bucket has server side encryption (SSE) enabled, make sure you know which algorithm it is using (AES256 or SSE-KMS).
If you are using SSE-KMS, you need the resource ARN of the managed key.
## Creation in the UI
### Step 1: Set up new Data Source
Head to the `Data Sources` view on Hopsworks (1) and click `New data source` (2).
The `Add a data source` catalog lists the available sources, grouped under `Object storage`, `Data warehouse`, `Database`, `Streaming` and `API & SaaS`.
Pick the `AWS S3` card to open the creation form.

The Data Source View in the User Interface
### Step 2: Enter Bucket Information
Enter the details for your S3 connector.
The `Source` line at the top of the form shows `AWS S3`, and `Change source` takes you back to the catalog.
Start by giving it a **name** and an optional **description**.
And set the name of the S3 Bucket you want to point the connector to.
Optionally, specify the region if you wish to have a Hopsworks-managed feature group stored using this connector.

S3 Connector Creation Form
### Step 3: Configure Authentication
#### Instance Role
Choose instance role if you have an EC2 instance profile attached to your Hopsworks cluster nodes with a role which grants you access to the specified bucket.
#### Temporary Credentials
Choose temporary credentials if you are using [AWS Role chaining](../../../../setup_installation/admin/roleChaining.md) to control the access permission on a project and user role base.
Once you have selected *Temporary Credentials* select the role that give access to the specified bucket.
For this role to appear in the list it needs to have been configured by an administrator, see the [AWS Role chaining documentation](../../../../setup_installation/admin/roleChaining.md) for more details.
!!! warning "Session Duration"
By default, the session duration that the role will be assumed for is 1 hour or 3600 seconds.
This means if you want to use the data source for example to write [training data to S3](../usage.md#writing-training-data), the training dataset creation cannot take longer than one hour.
Your administrator can change the default session duration for AWS data sources, by first [increasing the max session duration of the IAM Role](https://docs.aws.amazon.com/IAM/latest/UserGuide/id_roles_use.html#id_roles_use_view-role-max-session) that you are assuming.
And then changing the `fs_data_source_session_duration` [configuration variable](../../../../setup_installation/admin/variables.md) to the appropriate value in seconds.
#### Access Key/Secret
The most simple authentication method are Access Key/Secret, choose this option to get started quickly, if you are able to retrieve the keys using the IAM user administration.
### Step 4: Configure Server Side Encryption
Additionally, you can specify if your Bucket has SSE enabled.
#### AES256
For AES256, there is nothing to do but enabling the encryption by toggling the `AES256` option.
This is using S3-Managed Keys, also called [SSE-S3](https://docs.aws.amazon.com/AmazonS3/latest/userguide/serv-side-encryption.html).
#### SSE-KMS
With this option the [encryption key is managed by AWS KMS](https://docs.aws.amazon.com/AmazonS3/latest/userguide/serv-side-encryption.html), with some additional benefits and charges for using this service.
The difference is that you need to provide the resource ARN of the key.
If you have SSE-KMS enabled for your bucket, you can find the key ARN in the "Properties" section of the bucket details on AWS.
### Step 5: Add Spark Options (optional)
Here you can specify any additional spark options that you wish to add to the spark context at runtime.
Multiple options can be added as key - value pairs.
To connect to a S3 compatible storage other than AWS S3, you can add the option with key as `fs.s3a.endpoint` and the endpoint you want to use as value.
The data source will then be able to read from your specified S3 compatible storage.
You can also add options to configure the S3A client. For example, to disable SSL certificate verification, you can add the option with key as `fs.s3a.connection.ssl.enabled` and value as `false`. You can also configure other options such as `fs.s3a.path.style.access` if you use s3 compliant storage which does not support virtual hosting.
!!! warning "Spark Configuration"
When using the data source within a Spark application, the credentials are set at application level.
This allows users to access multiple buckets with the same data source within the same application (assuming the credentials allow it).
You can disable this behaviour by setting the option `fs.s3a.global-conf` to `False`.
If the `global-conf` option is disabled, the credentials are set on a per-bucket basis and users will be able to use the credentials to access data only from the bucket specified in the data source configuration.
### Step 6: Save changes
Click on "Save Credentials".
## Next Steps
Move on to the [usage guide for data sources](../usage.md) to see how you can use your newly created S3 connector.
================================================================================
# AWS Glue
Source: https://docs.hopsworks.ai/latest/user_guides/fs/data_source/creation/glue/
# How-To set up an AWS Glue Data Source { #data-source-glue }
## Introduction
The Glue Data Source integrates with the AWS Glue Data Catalog.
It points at a Glue database backed by Amazon S3, where the data always lives.
For this reason the Glue Data Source provides the same S3 credentials (`access_key`, `secret_key`, `session_token`, `region`) as the [S3 Data Source](s3.md).
This works for any data format: Apache Iceberg, Delta Lake and Apache Hudi, as well as plain file formats such as `csv` and `parquet`.
How the Glue Data Catalog itself is used depends on the format:
- Iceberg: the catalog owns the table's current-metadata pointer, so reads and writes are mediated by the catalog (the table is addressed by `.
`).
- Delta and Hudi: the on-path transaction log or timeline stays authoritative; the catalog is a discoverability mirror that is registered on create and synced on write so external engines (Athena, EMR, ...) can find the table by name.
- Plain file formats (`csv`, `parquet`, ...): the Data Source is used only for S3 access; nothing is registered in the catalog.
In this guide, you will configure a Data Source in Hopsworks to save all the authentication information needed in order to set up a connection to your AWS Glue database.
When you're finished, you'll be able to read tables using Spark through Hopsworks APIs, and to create managed feature groups whose offline data is stored in the Glue-registered location on S3.
!!! note
Currently, it is only possible to create data sources in the Hopsworks UI.
You cannot create a data source programmatically.
## Prerequisites
Before you begin this guide you'll need to retrieve the following information from your AWS Glue and S3 setup:
- **Database:** You will need the name of the Glue database that contains, or will contain, your tables.
- **Region:** You will need the AWS region in which the Glue Data Catalog and the backing S3 bucket reside.
The region is identified by its code.
- **Authentication Method:** You can authenticate using Access Key/Secret, or use IAM roles.
If you want to use an IAM role it either needs to be attached to the entire Hopsworks cluster or Hopsworks needs to be able to assume the role.
See [IAM role documentation](../../../../setup_installation/admin/roleChaining.md) for more information.
The credentials must grant access both to the Glue Data Catalog and to the backing S3 bucket.
## Creation in the UI
### Step 1: Set up new Data Source
Head to the `Data Sources` view on Hopsworks (1) and click `New data source` (2).
The `Add a data source` catalog opens below.
Pick the `AWS Glue` card to open the creation form.
The card is only offered where the connector is supported, so it is absent or disabled on non-cloud clusters.

The Data Source View in the User Interface
### Step 2: Enter Glue Settings
Enter the details for your Glue connector.
01. The form opens with `Source` set to `AWS Glue`.
Click `Change source` to pick a different one.
02. Give the data source a **name** and an optional **description**.
03. Set the name of the Glue **database** you want to point the connector to.
04. Optionally set the **Catalog ID**, the AWS account ID that owns the Glue Data Catalog.
Leave it empty to use the catalog of the account the credentials belong to.
05. Optionally set the AWS **region** of the Glue Data Catalog and its backing S3 bucket.
06. Choose the **Authentication method**, see the options below.
07. Optionally add **Spark options** as key-value pairs to pass to the Spark context at runtime.
08. Click on "Save Credentials".

Glue Connector Creation Form
The credentials must grant access both to the Glue Data Catalog and to the backing S3 bucket.
The available authentication methods are the same as for the [S3 Data Source](s3.md):
#### Instance Role
Choose instance role if you have an EC2 instance profile attached to your Hopsworks cluster nodes with a role which grants access to the Glue Data Catalog and the backing S3 bucket.
#### Temporary Credentials
Choose temporary credentials if you are using [AWS Role chaining](../../../../setup_installation/admin/roleChaining.md) to control the access permission on a project and user role base.
Once you have selected *Temporary Credentials* select the role that gives access to the Glue Data Catalog and the backing S3 bucket.
For this role to appear in the list it needs to have been configured by an administrator, see the [AWS Role chaining documentation](../../../../setup_installation/admin/roleChaining.md) for more details.
!!! warning "Session Duration"
By default, the session duration that the role will be assumed for is 1 hour or 3600 seconds.
This means if you want to use the data source for example to write [training data to S3](../usage.md#writing-training-data), the training dataset creation cannot take longer than one hour.
Your administrator can change the default session duration for AWS data sources, by first [increasing the max session duration of the IAM Role](https://docs.aws.amazon.com/IAM/latest/UserGuide/id_roles_use.html#id_roles_use_view-role-max-session) that you are assuming.
And then changing the `fs_data_source_session_duration` [configuration variable](../../../../setup_installation/admin/variables.md) to the appropriate value in seconds.
#### Access Key/Secret
The most simple authentication method are Access Key/Secret, choose this option to get started quickly, if you are able to retrieve the keys using the IAM user administration.
## Feature group path
When creating a feature group from this Data Source and the Glue database has a location, the feature group path is generated automatically by appending the new table to that database location, so no path needs to be set.
The database location is the **Location** set on the Glue database in the AWS console.

The Location of a Glue database in the AWS console
Otherwise, the path must be set explicitly on the data source, for example:
=== "PySpark"
```python
ds = fs.get_data_source("glue")
ds.path = "s3://mybucket/iceberg-warehouse/myglue.db/fg_1/"
```
An explicitly set path always takes precedence over the generated one.
## Direct Spark or PyIceberg access
For direct Spark or PyIceberg access outside the feature group APIs, the Data Source supplies the matching catalog properties.
See [`GlueConnector.catalog_options`][hsfs.storage_connector.GlueConnector.catalog_options] (Spark) and [`GlueConnector.pyiceberg_catalog_options`][hsfs.storage_connector.GlueConnector.pyiceberg_catalog_options] (PyIceberg).
## Next Steps
Move on to the [usage guide for data sources](../usage.md) to see how you can use your newly created Glue connector.
================================================================================
# Redshift
Source: https://docs.hopsworks.ai/latest/user_guides/fs/data_source/creation/redshift/
# How-To set up a Redshift Data Source
## Introduction
Amazon Redshift is a popular managed data warehouse on AWS, used as a data warehouse in many enterprises.
Data warehouses are often the source of raw data for feature engineering pipelines and Redshift supports scalable feature computation with SQL.
However, Redshift is not viable as an online feature store that serves features to models in production, with its columnar database layout its latency is too high compared to OLTP databases or key-value stores.
In this guide, you will configure a Data Source in Hopsworks to save all the authentication information needed in order to set up a connection to your AWS Redshift cluster.
When you're finished, you'll be able to query the database using Spark through Hopsworks APIs.
!!! note
Currently, it is only possible to create data sources in the Hopsworks UI.
You cannot create a data source programmatically.
## Prerequisites
Before you begin this guide you'll need to retrieve the following information from your AWS account and Redshift database, the following options are **mandatory**:
- **Cluster identifier:** The name of the cluster.
- **Database endpoint:** The endpoint for the database.
Should be in the format of `[UUID].eu-west-1.redshift.amazonaws.com`.
- **Database name:** The name of the database to query.
- **Database port:** The port of the cluster.
Defaults to 5349.
- **Authentication method:** There are three options available for authenticating with the Redshift cluster.
The first option is to configure a username and a password.
The second option is to configure an IAM role.
With IAM roles, Jobs or notebooks launched on Hopsworks do not need to explicitly authenticate with Redshift, as the Hopsworks library will transparently use the IAM role to acquire a temporary credential to authenticate the specified user.
Read more about IAM roles in our [AWS credentials pass-through guide](../../../../setup_installation/admin/roleChaining.md).
Lastly, option `Instance Role` will use the default ARN Role configured for the cluster instance.
## Creation in the UI
### Step 1: Set up new Data Source
Head to the `Data Sources` view on Hopsworks (1) and click `New data source` (2).
The `Add a data source` catalog lists the available sources, grouped under `Object storage`, `Data warehouse`, `Database`, `Streaming` and `API & SaaS`.
Pick the `Redshift` card to open the creation form.

The Data Source View in the User Interface
### Step 2: Enter The Connector Information
Enter the details for your Redshift connector.
Start by giving it a **name** and an optional **description**.
01. The form opens with `Source` set to `Redshift`.
Click `Change source` to pick a different one.
02. The name of the cluster.
03. The database endpoint.
Should be in the format `[UUID].eu-west-1.redshift.amazonaws.com`.
For example, if the endpoint info displayed in Redshift is `cluster-id.uuid.eu-north-1.redshift.amazonaws.com:5439/dev` the value to enter here is just `uuid.eu-north-1.redshift.amazonaws.com`
04. The database name.
05. The database port.
06. The database username, here you have the possibility to let Hopsworks auto-create the username for you.
07. Database Driver (optional): You can use the default JDBC Redshift Driver `com.amazon.redshift.jdbc42.Driver` included in Hopsworks or set a different driver (More on this later).
08. Optionally provide the database group and table for the connector.
A database group is the group created for the user if applicable.
More information, at [redshift documentation](https://docs.aws.amazon.com/redshift/latest/dg/r_Groups.html)
09. Set the appropriate authentication method.
10. Click on "Save Credentials".

Redshift Connector Creation Form
!!! warning "Session Duration"
By default, the session duration that the role will be assumed for is 1 hour or 3600 seconds.
This means if you want to use the data source for example to [read or create an external Feature Group from Redshift](../usage.md#creating-an-external-feature-group), the operation cannot take longer than one hour.
Your administrator can change the default session duration for AWS data sources, by first [increasing the max session duration of the IAM Role](https://docs.aws.amazon.com/IAM/latest/UserGuide/id_roles_use.html#id_roles_use_view-role-max-session) that you are assuming.
And then changing the `fs_data_source_session_duration` [configuration property](../../../../setup_installation/admin/variables.md) to the appropriate value in seconds.
### Step 3: Upload the Redshift database driver (optional)
The `redshift-jdbc42` JDBC driver is included by default in the Hopsworks distribution.
If you wish to use a different driver, you need to upload it on Hopsworks and add it as a dependency of Jobs and Jupyter Notebooks that need it.
First, you need to [download the library](https://docs.aws.amazon.com/redshift/latest/mgmt/jdbc20-download-driver.html).
Select the driver version without the AWS SDK.
#### Add the driver to Jupyter Notebooks and Spark jobs
You can now add the driver file to the default job and Jupyter configuration.
This way, all jobs and Jupyter instances in the project will have the driver attached so that Spark can access it.
1. Go into the Project's settings.
2. Select "Compute configuration".
3. Select "Spark".
4. Under "Additional Jars" choose "Upload new file" to upload the driver jar file.

Attaching the Redshift Driver to all Jobs and Jupyter Instances of the Project
Alternatively, you can choose the "From Project" option.
You will first have to upload the jar file to the Project using the File Browser.
After you have uploaded the jar file, you can select it using the "From Project" option.
To upload the jar file to the Project through the File Browser, see the example below:
1. Open File Browser
2. Navigate to "Resources" directory
3. Upload the jar file

Redshift Driver Upload in the File Browser
!!! tip
If you face network connectivity issues to your Redshift cluster, a common cause could be the cluster database port not being accessible from outside the Redshift cluster VPC network.
A quick and dirty way to enable connectivity is to [Enable Publicly Accessible](https://aws.amazon.com/premiumsupport/knowledge-center/redshift-cluster-private-public/).
However, in a production setting, you should use [VPC peering](https://docs.aws.amazon.com/vpc/latest/peering/what-is-vpc-peering.html) or some equivalent mechanism for connecting the clusters.
## Next Steps
Move on to the [usage guide for data sources](../usage.md) to see how you can use your newly created Redshift connector.
================================================================================
# ADLS
Source: https://docs.hopsworks.ai/latest/user_guides/fs/data_source/creation/adls/
# How-To set up a ADLS Data Source
## Introduction
Azure Data Lake Storage (ADLS) Gen2 is a HDFS-compatible filesystem on Azure for data analytics.
The ADLS Gen2 filesystem stores its data in Azure Blob storage, ensuring low-cost storage, high availability, and disaster recovery.
In Hopsworks, you can access ADLS Gen2 by defining a Data Source and creating and granting permissions to a service principal.
In this guide, you will configure a Data Source in Hopsworks to save all the authentication information needed in order to set up a connection to your Azure ADLS filesystem.
When you're finished, you'll be able to read files using Spark through Hopsworks APIs.
You can also use the connector to write out training data from the Feature Store, in order to make it accessible by third parties.
!!! note
Currently, it is only possible to create data sources in the Hopsworks UI.
You cannot create a data source programmatically.
## Prerequisites
Before you begin this guide you'll need to retrieve the following information from your Azure ADLS account:
- **Data Lake Storage Gen2 Account:** Create an [Azure Data Lake Storage Gen2 account](https://docs.microsoft.com/azure/storage/data-lake-storage/quickstart-create-account) and [initialize a filesystem, enabling the hierarchical namespace](https://docs.microsoft.com/azure/storage/data-lake-storage/namespace).
Note that your storage account must belong to an Azure resource group.
- **Azure AD application and service principal:** [Create an Azure AD application and service principal](https://docs.microsoft.com/en-us/azure/active-directory/develop/howto-create-service-principal-portal) that can access your ADLS storage account and its resource group.
- **Service Principal Registration:** Register the service principal, granting it a role assignment such as Storage Blob Data Contributor, on the Azure Data Lake Storage Gen2 account.
!!! info
When you specify the 'container name' in the ADLS data source, you need to have previously created that container - the Hopsworks Feature Store will not create that storage container for you.
## Creation in the UI
### Step 1: Set up new Data Source
Head to the `Data Sources` view on Hopsworks (1) and click `New data source` (2).
The `Add a data source` catalog opens below.
Pick the `Azure Data Lake` card to open the creation form.
The card is only offered where the connector is supported, so it is absent or disabled on non-cloud clusters.

The Data Source View in the User Interface
### Step 2: Enter ADLS Information
Enter the details for your ADLS connector.
Start by giving it a **name** and an optional **description**.

ADLS Connector Creation Form
1. The form opens with `Source` set to `Azure Data Lake`.
Click `Change source` to pick a different one.
2. Set directory ID.
3. Enter the Application ID.
4. Paste the Service Credentials.
5. Specify account name.
6. Provide the container name.
7. Click on "Save Credentials".
### Step 3: Azure Create an ADLS Resource
When programmatically signing in, you need to pass the tenant ID with your authentication request and the application ID.
You also need a certificate or an authentication key (described in the following section).
To get those values, use the following steps:
1. Select Azure Active Directory.
2. From App registrations in Azure AD, select your application.
3. Copy the Directory (tenant) ID and store it in your application code.

You need to copy the Directory (tenant) id and paste it to the Hopsworks ADLS Data Source "Directory id" text field.
4. Copy the Application ID and store it in your application code.

>You need to copy the Application id and paste it to the Hopsworks ADLS Data Source "Application id" text field.
5. Create an Application Secret and copy it into the Service Credential field.

You need to copy the Application Secret and paste it to the Hopsworks ADLS Data Source "Service Credential" text field.
#### Common Problems
If you get a permission denied error when writing or reading to/from a ADLS container, it is often because the storage principal (app) does not have the correct permissions.
Have you added the "Storage Blob Data Owner" or "Storage Blob Data Contributor" role to the resource group for your storage account (or the subscription for your storage group, if you apply roles at the subscription level)?
Go to your resource group, then in "Access Control (IAM)", click the "Add" button to add a "role assignment".
If you get an error "StatusCode=404 StatusDescription=The specified filesystem does not exist.", then maybe you have not created the storage account or the storage container.
#### References
- [How to create a service principal on Azure](https://docs.microsoft.com/en-us/azure/active-directory/develop/howto-create-service-principal-portal)
## Next Steps
Move on to the [usage guide for data sources](../usage.md) to see how you can use your newly created ADLS connector.
================================================================================
# BigQuery
Source: https://docs.hopsworks.ai/latest/user_guides/fs/data_source/creation/bigquery/
# How-To set up a BigQuery Data Source
## Introduction
A BigQuery data source provides integration to Google Cloud BigQuery.
BigQuery is Google Cloud's managed data warehouse supporting that lets you run analytics and execute SQL queries over large scale data.
Such data warehouses are often the source of raw data for feature engineering pipelines.
In this guide, you will configure a Data Source in Hopsworks to connect to your BigQuery project by saving the necessary information.
When you're finished, you'll be able to execute queries and read results of BigQuery using Spark through Hopsworks APIs.
The data source uses the Google `spark-bigquery-connector` behind the scenes.
To read more about the spark connector, like the spark options or usage, check [Apache Spark SQL connector for Google BigQuery.](https://github.com/GoogleCloudDataproc/spark-bigquery-connector#usage 'github.com/GoogleCloudDataproc/spark-bigquery-connector')
!!! note
Currently, it is only possible to create data sources in the Hopsworks UI.
You cannot create a data source programmatically.
## Prerequisites
Before you begin this guide you'll need to retrieve the following information about your GCP account:
- **BigQuery Project:** You need a BigQuery project, dataset and table created and have read access to it.
Or, if you wish to query a public dataset you need its corresponding details.
- **Authentication Method:** Authentication to GCP account is handled by uploading the `JSON keyfile for service account` to the Hopsworks Project.
You will need to create this JSON keyfile from GCP.
For more information on service accounts and creating keyfile in GCP, read [Google Cloud documentation.](https://cloud.google.com/docs/authentication/production#create_service_account 'creating service account keyfile')
!!! note
To read data, the BigQuery service account user needs permission to `create read session` which is available in **BigQuery Admin role**.
## Creation in the UI
### Step 1: Set up new Data Source
Head to the `Data Sources` view on Hopsworks (1) and click `New data source` (2).
The `Add a data source` catalog lists the available sources, grouped under `Object storage`, `Data warehouse`, `Database`, `Streaming` and `API & SaaS`.
Pick the `Google BigQuery` card to open the creation form.

The Data Source View in the User Interface
### Step 2: Enter source details
Enter the details for your BigQuery storage.
Start by giving it a unique **name** and an optional **description**.

BigQuery Creation Form
1. The form opens with `Source` set to `Google BigQuery`.
Click `Change source` to pick a different one.
2. Next, set the name of the parent BigQuery project.
This is used for billing by GCP.
3. Authentication: Here you should upload your `JSON keyfile for service account` used for authentication.
You can choose to either upload from your local using `Upload new file` or choose an existing file within project using `From Project`.
4. Read Options:
In the UI set the below fields,
1. *BigQuery Project*: The BigQuery project to read
2. *BigQuery Dataset*: The dataset of the table (Optional)
3. *BigQuery Table*: The table to read (Optional)
!!! note
*Materialization Dataset*: Temporary dataset used by BigQuery for writing.
It must be set to a dataset where the GCP user has table creation permission.
The queried table must be in the same location as the `materializationDataset` (e.g 'EU' or 'US').
Also, if a table in the `SQL statement` is from project other than the `parentProject` then use the fully qualified table name i.e. `[project].[dataset].[table]`.
For details, read the Google documentation on [usage of query for BigQuery Spark connector](https://github.com/GoogleCloudDataproc/spark-bigquery-connector#reading-data-from-a-bigquery-query).
5. Spark Options: Optionally, you can set additional spark options using the `Key - Value` pairs.
6. Click on "Save Credentials".
## Next Steps
Move on to the [usage guide for data sources](../usage.md) to see how you can use your newly created BigQuery connector.
================================================================================
# GCS
Source: https://docs.hopsworks.ai/latest/user_guides/fs/data_source/creation/gcs/
# How-To set up a GCS Data Source { #data-source-gcs }
## Introduction
This particular type of Data Source provides integration to Google Cloud Storage (GCS).
GCS is an object storage service offered by Google Cloud.
An object could be simply any piece of immutable data consisting of a file of any format, for example a `CSV` or `PARQUET`.
These objects are stored in containers called as `buckets`.
These types of storages are often the source for raw data from which features can be engineered.
In this guide, you will configure a Data Source in Hopsworks to connect to your GCS bucket by saving the necessary information.
When you're finished, you'll be able to read files from the GCS bucket using Spark through Hopsworks APIs.
The Data Source uses the Google `gcs-connector-hadoop` behind the scenes.
For more information, check out [Google Cloud Data Source for Spark and Hadoop](https://github.com/GoogleCloudDataproc/hadoop-connectors/tree/master/gcs#google-cloud-storage-connector-for-spark-and-hadoop 'google-cloud-storage-connector-for-spark-and-hadoop').
!!! note
Currently, it is only possible to create data sources in the Hopsworks UI.
You cannot create a data source programmatically.
## Prerequisites
Before you begin this guide you'll need to retrieve the following information about your GCP account and bucket:
- **Bucket:** You need a GCS bucket created and have read access to it.
The bucket is identified by its name.
- **Authentication Method:** Authentication to GCP account is handled by uploading the `JSON keyfile for service account` to the Hopsworks Project.
You will need to create this JSON keyfile from GCP.
For more information on service accounts and creating keyfile in GCP, read [Google Cloud documentation.](https://cloud.google.com/docs/authentication/production#create_service_account 'creating service account keyfile')
- **Server-side Encryption** GCS encrypts the data on server side by default.
The connector additionally supports the optional encryption method `Customer Supplied Encryption Key` by GCP.
You can choose the encryption option `AES-256` and provide AES-256 key and hash, encoded in standard Base64.
The encryption details are stored as [Secrets](../../../projects/secrets/create_secret.md) in the Hopsworks for keeping it secure.
Read more about encryption on [Google Documentation.](https://cloud.google.com/storage/docs/encryption/customer-supplied-keys)
## Creation in the UI
### Step 1: Set up new Data Source
Head to the `Data Sources` view on Hopsworks (1) and click `New data source` (2).
The `Add a data source` catalog lists the available sources, grouped under `Object storage`, `Data warehouse`, `Database`, `Streaming` and `API & SaaS`.
Pick the `Google Cloud Storage` card to open the creation form.

The Data Source View in the User Interface
### Step 2: Enter connector details
Enter the details for your GCS connector.
Start by giving it a unique **name** and an optional **description**.

GCS Connector Creation Form
1. The form opens with `Source` set to `Google Cloud Storage`.
Click `Change source` to pick a different one.
2. Next, set the name of the GCS Bucket you wish to connect with.
3. Authentication: Here you should upload your `JSON keyfile for service account` used for authentication.
You can choose to either upload from your local using `Upload new file` or choose an existing file within project using `From Project`.
4. GCS Server Side Encryption: You can leave this to `Default Encryption` if you do not wish to provide explicit encrypting keys.
Otherwise, optionally you can set the encryption setting for `AES-256` and provide the encryption key and hash when selected.
5. Click on `Save Credentials`.
## Next Steps
Move on to the [usage guide for data sources](../usage.md) to see how you can use your newly created GCS
connector.
================================================================================
# SQL
Source: https://docs.hopsworks.ai/latest/user_guides/fs/data_source/creation/sql/
# How-To set up an SQL Data Source
## Introduction
The SQL Data Source connects Hopsworks to a Relational Database Service.
Supported database types are **MySQL**, **PostgreSQL**, and **Oracle**.
Using this connector, you can query and update data in your relational database from Hopsworks.
In this guide, you will configure a Data Source in Hopsworks to securely store the authentication information needed to set up a connection to your database instance.
When you're finished, you'll be able to query your SQL database using Hopsworks APIs.
!!! note
Currently, it is only possible to create data sources in the Hopsworks UI.
You cannot create a data source programmatically.
## Prerequisites
Before you begin, ensure you have the following information from your database instance:
- **Host:** The endpoint for your database instance.
Example from AWS:
1. Go to the AWS Console → `Aurora and RDS`
2. Click on your DB instance.
3. Under `Connectivity & security`, you'll find the endpoint, e.g.:
`mydb.abcdefg1234.us-west-2.rds.amazonaws.com`
- **Database:** The name of the database to connect to.
For Oracle, this is the **service name** (e.g. `ORCL` or a TNS alias).
- **Port:** The port to connect to (e.g. `3306` for MySQL, `5432` for PostgreSQL, `1521` for Oracle).
- **Username and Password:** A username and password with the necessary permissions to access the required tables.
### Optional: Oracle Wallet for mTLS Authentication
If your Oracle database requires mutual TLS (mTLS) authentication, which is common with Oracle Autonomous Database and Oracle Cloud, you will also need:
- **Wallet file:** A `.zip` file containing the wallet credentials (e.g. `cwallet.sso`, `tnsnames.ora`, `sqlnet.ora`).
- **Wallet password:** The password for the wallet, if using a PKCS12 wallet (`ewallet.p12`).
Auto-login wallets (`cwallet.sso`) do not require a password.
!!! tip
You can download the wallet zip from the Oracle Cloud Console under your Autonomous Database's **DB Connection** page.
Upload the zip file to your Hopsworks project (e.g. to `Resources/`) before creating the data source.
!!! warning "Leave the host empty when using a wallet"
A host and a wallet are alternatives, not a pair.
The wallet's `tnsnames.ora` supplies the host and port, and the database field is the alias to look up there.
Supplying a host as well makes the driver connect directly, past the wallet, which a wallet-protected database refuses with a connection error that names neither the cause nor the fix.
Hopsworks therefore refuses to save a data source with both a host and a wallet.
## Creation in the UI
### Step 1: Set up a new Data Source
Head to the `Data Sources` view on Hopsworks (1) and click `New data source` (2).
The `Add a data source` catalog lists the available sources, grouped under `Object storage`, `Data warehouse`, `Database`, `Streaming` and `API & SaaS`.
Pick the `SQL` card to open the creation form.

The Data Source View in the User Interface
### Step 2: Enter SQL Settings
Enter the details for your database.
Start by giving the connector a **name** and an optional **description**.
1. The form opens with `Source` set to `SQL`.
Click `Change source` to pick a different one.
2. Select the database type (MySQL, PostgreSQL, or Oracle).
3. Enter the host endpoint.
Leave it empty when using an Oracle wallet: the wallet supplies the connection details, and the database field names the TNS alias to use.
4. Enter the database name (service name for Oracle).
5. Specify the port.
6. Provide the username and password.
7. For Oracle with mTLS, upload the wallet zip file and provide the wallet password (if required).
8. Click on "Save Credentials".

SQL Connector Creation Form
## Oracle-Specific Notes
The generic read, external feature group, and training data workflows are covered in the [usage guide for data sources][data-source-usage].
The following notes apply only to Oracle.
### JDBC driver on the Spark classpath
The Oracle JDBC driver JAR (e.g. `ojdbc11.jar`) must be available on the Spark classpath.
Upload it via the [Jupyter configuration][how-to-run-a-pyspark-notebook] or [Job configuration][how-to-run-a-pyspark-job] in `Additional Jars`.
The MySQL and PostgreSQL drivers are included in Hopsworks by default.
### Spark JDBC limitations
!!! warning "Oracle Spark JDBC limitations"
- **Single-partition reads only.**
All data is fetched through a single JDBC connection from the Spark driver.
Spark's parallel JDBC read (via `numPartitions` / `partitionColumn`) is not supported.
For very large tables, filter with a `WHERE` clause in your query.
- **Wallet available on the driver only.**
When using wallet-based authentication, the wallet zip is downloaded from HopsFS and extracted on the Spark driver node.
This is sufficient because reads are single-partition (driver-only).
- **Timestamp precision.**
Spark JDBC supports timestamp precision up to seconds only.
Sub-second precision from Oracle `TIMESTAMP` columns may be truncated.
### Python engine
The Python engine reads Oracle via the Hopsworks Arrow Flight service, which handles the database connection server-side.
No JDBC driver or wallet files are needed on the client, and the Spark JDBC limitations above do not apply.
## Next Steps
Move on to the [usage guide for data sources][data-source-usage] to see how you can use your newly created SQL connector.
You can also make the database queryable from the query engine by adding a [Trino catalog][trino-catalogs] derived from this data source.
================================================================================
# CRM, Sales & Analytics
Source: https://docs.hopsworks.ai/latest/user_guides/fs/data_source/creation/crm_sales_analytics/
# How-To set up a CRM, Sales & Analytics Data Source
## Introduction
The `CRM, Sales & Analytics` data source lets you connect Hopsworks to supported business applications and marketing platforms.
The following sources are available:
- Facebook Ads
- Freshdesk
- Google Ads
- Google Analytics
- HubSpot
- Pipedrive
- Salesforce
- Shopify
In this guide, you will configure a Data Source in Hopsworks by saving the credentials required by the selected source.
!!! note
Currently, it is only possible to create data sources in the Hopsworks UI.
You cannot create a data source programmatically.
## Prerequisites
Before you begin, make sure you have:
- A unique name for the data source in Hopsworks.
- Read credentials for the external system you want to connect.
- Any source-specific identifiers required by that system, such as account, customer, property, or domain identifiers.
- For Google Ads and Google Analytics, a service account JSON keyfile that can be uploaded to the Hopsworks project.
## Creation in the UI
### Step 1: Set up new Data Source
Head to the `Data Sources` view on Hopsworks (1) and click `New data source` (2).
The `Add a data source` catalog lists the available sources, grouped under `Object storage`, `Data warehouse`, `Database`, `Streaming` and `API & SaaS`.
Pick the `CRM, Sales & Analytics` card to open the creation form.

The Data Source View in the User Interface
### Step 2: Name the data source and select the platform
The form opens with `Source` set to `CRM, Sales & Analytics`, and `Change source` takes you back to the catalog.
Enter a unique **Name**, an optional **Description**, and pick the platform you want to configure in the **Source** radio group of the connection section.

CRM, Sales & Analytics data source selection
### Step 3: Enter source-specific credentials
The required fields depend on the selected source.
#### Facebook Ads
Required fields:
- **Access Token**
- **Account Id**

Facebook Ads data source form
#### Freshdesk
Required fields:
- **API Key**
- **Domain**

Freshdesk data source form
#### Google Ads
Required fields:
- **Authentication JSON Keyfile**
- **Developer Token**
- **Customer Id**
- **Impersonated Email**
The JSON keyfile can be selected either from an existing project file or uploaded as a new file.

Google Ads data source form
#### Google Analytics
Required fields:
- **Authentication JSON Keyfile**
- **Property Id**
The JSON keyfile can be selected either from an existing project file or uploaded as a new file.

Google Analytics data source form
#### HubSpot
Required fields:
- **API Key**

HubSpot data source form
#### Pipedrive
Required fields:
- **API Key**

Pipedrive data source form
#### Salesforce
Required fields:
- **Security Token**
- **Username**
- **Password**

Salesforce data source form
#### Shopify
Required fields:
- **Shop URL**
- **Private App Password**

Shopify data source form
### Step 4: Save the credentials
After entering the required fields for the selected source:
1. Click **Save Credentials**.
2. Click **Next: Select resource** to continue configuring the data source for downstream use.
## Next Steps
Move on to the [usage guide for data sources](../usage.md) to see how you can use your newly created data source.
================================================================================
# REST API
Source: https://docs.hopsworks.ai/latest/user_guides/fs/data_source/creation/rest_api/
# How-To set up a REST API Data Source
## Introduction
The `REST API` data source lets you connect Hopsworks to external HTTP APIs.
You can use it to store the base connection details, optional headers, and the authentication method required by the target API.
In this guide, you will configure a REST API Data Source in the Hopsworks UI.
!!! note
Currently, it is only possible to create data sources in the Hopsworks UI.
You cannot create a data source programmatically.
## Prerequisites
Before you begin, make sure you have:
- A unique name for the data source in Hopsworks.
- The **Base URL** of the target API.
- Any headers you want to send with requests.
- The authentication details required by the target API.
## Creation in the UI
### Step 1: Set up new Data Source
Head to the `Data Sources` view on Hopsworks (1) and click `New data source` (2).
The `Add a data source` catalog lists the available sources, grouped under `Object storage`, `Data warehouse`, `Database`, `Streaming` and `API & SaaS`.
Pick the `REST API` card to open the creation form.

The Data Source View in the User Interface
### Step 2: Enter REST API settings
The form opens with `Source` set to `REST API`, and `Change source` takes you back to the catalog.
Provide the common connection settings shown in the form:
1. **Name:** A unique name for the data source.
2. **Description:** Optional description.
3. **Base URL:** The base endpoint for the external API.
4. **Headers:** Optional header key-value pairs. Use the `+` button to add headers.
5. **Authentication:** Select the authentication mode required by the API.
The following authentication modes are available in the UI:
- `NONE`
- `BEARER_TOKEN`
- `API_KEY`
- `HTTP_BASIC`
- `OAUTH2_CLIENT`

REST API data source form
!!! note
The screenshot shows the form with `NONE` selected.
When you choose another authentication mode, the form will prompt for the additional credentials required by that method.
### Step 3: Save the credentials
After entering the connection details:
1. Click **Save Credentials**.
2. Click **Next: Select resource** to continue configuring the data source for downstream use.
## Next Steps
Move on to the [usage guide for data sources](../usage.md) to see how you can use your newly created REST API data source.
================================================================================
# Unity Catalog
Source: https://docs.hopsworks.ai/latest/user_guides/fs/data_source/creation/unity_catalog/
# How-To set up a Unity Catalog Data Source
## Introduction
A Unity Catalog data source provides integration with [Databricks Unity Catalog](https://docs.databricks.com/aws/en/data-governance/unity-catalog/).
Unity Catalog is Databricks' unified governance layer for data and AI assets, organised as a catalog → schema → table hierarchy.
In this guide, you will configure a Data Source in Hopsworks that points at a Databricks workspace.
Once configured, you can browse catalogs, schemas, and tables, and mount Delta tables as external Feature Groups whose data is read through the Arrow Flight query service.
!!! warning "Databricks on AWS only"
Unity Catalog is currently only supported on **Databricks on AWS**.
Databricks on Azure and Databricks on GCP are not supported in this release.
Their Unity Catalog temporary-table-credentials responses use cloud-specific credential shapes (Azure SAS tokens, GCP service-account tokens) that the Hopsworks Arrow Flight read path does not yet handle.
If you point this connector at a non-AWS Databricks workspace the browse flow may succeed but every preview and feature-group read will fail.
!!! note
Only Delta-formatted Unity Catalog tables are supported in this release.
Managed non-Delta tables, Iceberg tables, views, streaming tables, and materialised views are filtered out when browsing.
!!! note
Currently, it is only possible to create data sources in the Hopsworks UI.
You cannot create a data source programmatically.
!!! warning
Direct Spark reads from Unity Catalog are not supported in this release.
Reads flow through the Arrow Flight query service, which resolves each Delta table via the Unity Catalog REST API and reads it using the `deltalake` Python package (delta-rs) against the S3 location that Databricks returns.
## Prerequisites
Before you begin you need all of the following.
The first three are on the Databricks side and are the most common source of 400 / 403 errors from the read path.
### Databricks side
- **External Data Access enabled on the metastore.**
In the Databricks account console, go to Catalog → the metastore backing your workspace → Details, and turn on "External data access".
Without this toggle, every call to `/api/2.1/unity-catalog/temporary-table-credentials` returns `403 Forbidden` for any principal.
This is an account-admin setting, workspace admin alone cannot flip it.
- **`EXTERNAL USE SCHEMA` grant on the schemas you want to read.** In Databricks SQL:
```sql
GRANT EXTERNAL USE SCHEMA ON SCHEMA . TO ``;
```
where `` is the user (or service principal) that owns the PAT you are about to paste into Hopsworks. Without this grant the temporary-table-credentials endpoint returns `400 Bad Request`.
- **`USE CATALOG`, `USE SCHEMA`, and `SELECT` grants** on the specific catalog / schema / tables you want to mount.
- **Delta format.** Unity Catalog tables backed by Iceberg, non-Delta file formats, views, or streaming / materialised views cannot be read through this connector in v1.
### Hopsworks side
- **Databricks workspace URL**, for example `https://.cloud.databricks.com`.
- **Personal access token** for the principal to which the grants above were issued.
- **A catalog name** containing the Delta tables you want to mount. It is optional but recommended to set this as the default catalog on the connector so the UI opens the browse view directly.
- Optional: an **AWS region** (for example `us-west-2`). If you omit it, the backend guesses the region by parsing the STS session-token returned with the table credentials. For FIPS regions or for any workspace where the guess has been wrong once, set the region explicitly on the connector.
The personal access token is stored encrypted in the Hopsworks `secrets` table.
It is never written to the connector table in plaintext.
PAT rotation is manual in this release: when the token expires, edit the connector and paste in a fresh one. Unity Catalog PATs are typically short-lived (hours to a few days), so expect to do this periodically.
## Feature flag
Unity Catalog connectors are gated by the `enable_unity_catalog_storage_connectors` Hopsworks variable.
An administrator must set it to `true` in the admin variables UI before the connector type appears in the create form.
## Creation in the UI
### Step 1: Set up new Data Source
Head to the `Data Sources` view in Hopsworks (1) and click `New data source` (2).
The `Add a data source` catalog opens below.
Pick the `Databricks Unity Catalog` card to open the creation form.

The Data Source view in the user interface
### Step 2: Enter source details
Enter the details for your Unity Catalog workspace.
Start by giving it a unique **name** and an optional **description**.
1. The form opens with `Source` set to `Databricks Unity Catalog`.
Click `Change source` to pick a different one.
2. **Databricks Workspace URL**: the full `https://` URL of your workspace.
3. **Access Token**: a Databricks personal access token; the field is masked and stored encrypted.
4. **Default Catalog**: optional; the Unity Catalog catalog to pre-select when browsing.
5. **AWS Region**: optional; set explicitly (for example `us-west-2`) when the backend's region guess from the STS session-token is wrong or your workspace is in a FIPS region. Leave empty to use the guess.
6. **Arguments**: optional key/value pairs passed through to the query service.
7. Click "Save Credentials".
On save, Hopsworks calls the Unity Catalog `/catalogs` endpoint using the provided token; an HTTP 2xx response is required for the connector to be accepted.
### Step 3: Browse and mount a table
After saving, open the connector and click **Configure**.
The "Catalog" dropdown lists catalogs visible to your token; pick one and the schema/table browser lists Delta tables grouped by schema.
Select a table to create an external Feature Group pointing at it.
Reads will flow through the Arrow Flight query service.
## Next Steps
Move on to the [usage guide for data sources](../usage.md) to see how you can use your newly created Unity Catalog connector.
================================================================================
# SAP HANA
Source: https://docs.hopsworks.ai/latest/user_guides/fs/data_source/creation/sap_hana/
# How-To set up an SAP HANA Data Source { #data-source-sap-hana }
## Introduction
SAP HANA is an in-memory relational database used by many enterprises as the system of record for ERP, CRM, and analytics workloads.
An SAP HANA Data Source in Hopsworks stores the connection details required to read tables and views from a HANA tenant database.
Once configured, you can use the same data source as the basis for an external (on-demand) Feature Group, or as the source for a dltHub-driven ingestion job that materialises HANA data into a managed Feature Group.
In this guide, you will configure a Data Source in Hopsworks that holds the authentication information needed to connect to your SAP HANA database.
!!! note
Currently, it is only possible to create data sources in the Hopsworks UI.
You cannot create a data source programmatically.
## Prerequisites
Before you begin this guide you'll need to retrieve the following information from your SAP HANA tenant.
The following options are **mandatory**:
- **Host**: The hostname of the SAP HANA endpoint, for example `hxehost.example.com` for an on-premise instance or the endpoint shown in SAP BTP for SAP HANA Cloud.
- **Port**: The SQL port of the tenant database.
The default is `39015`, the SQL port for the first tenant database on a default
multi-tenant or HANA Express (HXE) install (instance number 90).
For a non-tenant single-host install (instance 00) use `30015`.
SAP HANA Cloud typically uses `443`.
Consult your DBA if you are unsure.
- **User**: The HANA database user that the connector authenticates as.
- **Password**: The password for that user.
These are a few additional **optional** arguments:
- **Database**: The tenant database name.
Use this when your SAP HANA system hosts more than one tenant database and you need to target a specific one.
- **Schema**: The default schema applied to unqualified queries on the connection.
If you leave this empty, queries must fully qualify table names with the schema prefix.
- **Table**: The default table the connector points at when no SQL query is provided.
- **Application**: A short identifier surfaced in HANA's session tracing (`APPLICATION` session variable).
This makes it easier to attribute load to Hopsworks in HANA monitoring tools.
- **Additional arguments**: Free-form key/value options forwarded to the underlying SAP HANA Python driver (`hdbcli`) and the Spark JDBC reader.
!!! info "Drivers"
Hopsworks ships the SAP HANA drivers needed to read from HANA out of the box.
The Hopsworks Spark image bundles the SAP `ngdbc` JDBC driver for Spark JDBC reads, and the dlt ingestion image and Arrow Flight server bundle SAP's `hdbcli` Python DBAPI driver.
You do not need to install or upload the drivers yourself.
## Creation in the UI
### Step 1: Set up new Data Source
Head to the `Data Sources` view on Hopsworks and click `New data source`.
The `Add a data source` catalog opens below.
Pick the `SAP HANA` card to open the creation form.

The Data Source View in the User Interface
### Step 2: Enter SAP HANA Settings
Enter the details for your SAP HANA connector.
Start by giving it a **name** and an optional **description**.
01. Select "SAP HANA" as storage.
02. Specify the **Host** of your SAP HANA endpoint.
03. Specify the **Port** the tenant SQL service listens on (default `39015`).
04. Provide the **User** name of the HANA database user.
05. Provide the **Password** for that user.
06. Optionally fill in **Database**, **Schema**, **Table**, and **Application**.
07. Optionally add additional key/value arguments.
These are forwarded both to the Python driver used by the on-demand read path and to the Spark JDBC reader used by notebook jobs.
08. Click on "Save Credentials".
## Use it as an ingestion source
Once the SAP HANA data source exists, you can also use it with the dltHub-based ingestion workflow described in [Ingest Data with dltHub][ingest-data-with-dlthub].
SAP HANA is treated as a SQL-like source, so the ingestion job supports both full and incremental loading.
## Type mapping
Hopsworks reads each source column's HANA type from the cursor description and maps it to a Hopsworks offline feature type.
The mapping preserves precision and scale where possible, so a source `DECIMAL(12, 2)` becomes a Hopsworks `decimal(12,2)` feature rather than collapsing to `bigint`.
| SAP HANA type | Hopsworks offline feature type |
| --- | --- |
| `TINYINT` | `tinyint` |
| `SMALLINT` | `smallint` |
| `INTEGER` | `int` |
| `BIGINT` | `bigint` |
| `DECIMAL(p, s)` | `decimal(p,s)` |
| `REAL` | `float` |
| `DOUBLE` | `double` |
| `BOOLEAN` | `boolean` |
| `DATE` | `date` |
| `TIME` | `timestamp` |
| `TIMESTAMP` / `SECONDDATE` / `LONGDATE` | `timestamp` |
| `CHAR` / `VARCHAR` / `NCHAR` / `NVARCHAR` / `TEXT` / `CLOB` / `NCLOB` / `ALPHANUM` | `string` |
| `BINARY` / `VARBINARY` / `BLOB` | `binary` |
## Known limitations
### Avoid the `SYSTEM` schema for source tables
Place tables you intend to ingest or expose as feature groups in a regular user schema (for example a project-specific `MYAPP` or `HOPSDEMO`).
Tables created under the system-owned `SYSTEM` schema do not reflect cleanly through the SQLAlchemy HANA dialect that powers DLT ingestion.
A typical setup is:
```sql
CREATE SCHEMA HOPSDEMO;
RENAME TABLE SYSTEM.MY_TABLE TO HOPSDEMO.MY_TABLE;
```
Then set **Schema** in the data source to `HOPSDEMO` (or pick it from the schema browser) and use that as the basis for any external feature group or DLT ingestion job.
### Online ingestion requires non-null primary keys
When you create a managed Feature Group fed from SAP HANA via DLT and enable online serving, online ingestion validates that every row has a non-null value in the Feature Group's primary-key column.
If the source rows can carry `NULL` in that column, either filter them out at source, pick a different primary key on the Feature Group, or disable online serving for the Feature Group.
### Authentication
The SAP HANA data source currently supports username and password authentication.
Certificate-based and JWT authentication are tracked as follow-up work.
## Next Steps
Move on to the [usage guide for data sources][data-source-usage] to see how you can use your newly created SAP HANA connector.
================================================================================
# Usage
Source: https://docs.hopsworks.ai/latest/user_guides/fs/data_source/usage/
# Data Source Usage
Here, we look at how to use a Data Source after it has been created.
Data Sources provide an important first step for integrating with external data.
The 4 fundamental functionalities where data sources are used are:
1. Reading data into Spark Dataframes
2. Creating external feature groups
3. Writing training data
4. Creating managed feature groups
We will walk through each functionality in the sections below.
## Retrieving a Data Source
We retrieve a data source simply by its unique name.
=== "PySpark"
```python
import hopsworks
# Connect to the Hopsworks feature store
project = hopsworks.login()
feature_store = project.get_feature_store()
# Retrieve data source
ds = feature_store.get_data_source("data_source_name")
```
=== "Scala"
```scala
import com.logicalclocks.hsfs._
val connection = HopsworksConnection.builder().build();
val featureStore = connection.getFeatureStore();
// get directly via connector sub-type class, e.g., for GCS type
val connector = featureStore.getGcsConnector("data_source_name")
```
## Reading a Spark Dataframe from a Data Source
One of the most common usages of a Data Source is to read data directly into a Spark Dataframe.
It's achieved via the `read` API of the connector object, which hides all the complexity of authentication and integration with a data storage source.
The `read` API primarily has two parameters for specifying the data source, `path` and `query`, depending on the data source type.
The exact behaviour could change depending on the fdata source type, but broadly they could be classified as below
### Data lake/object based connectors
For data sources based on object/file storage such as AWS S3, ADLS, GCS, we set the full object path in the `path` argument and users should pass a Spark data format (parquet, csv, orc, hudi, delta) to the `data_format` argument.
=== "PySpark"
```python
# read data into dataframe using path
df = connector.read(
data_format="data_format", path="fileScheme://bucket/path/"
)
```
=== "Scala"
```scala
// read data into dataframe using path
val df = connector.read("", "data_format", new HashMap(), "fileScheme://bucket/path/")
```
#### Prepare Spark API
Additionally, for reading file based data sources, another way to read the data is using the `prepare_spark` method.
This method can be used if you are reading the data directly through Spark.
Firstly, it handles the setup of all Spark configurations or properties necessary for a particular type of connector and prepares the absolute path to read from, along with bucket name and the appropriate file scheme of the data source.
A Spark session can handle only one configuration setup at a time, so Hopsworks cannot set the Spark configurations when retrieving the connector since it would lead to only always initialising the last connector being retrieved.
Instead, user can do this setup explicitly with the `prepare_spark` method and therefore potentially use multiple connectors in one Spark session. `prepare_spark` handles only one bucket associated with that particular connector, however, it is possible to set up multiple connectors with different types as long as their Spark properties do not interfere with each other.
So, for example a S3 connector and a Snowflake connector can be used in the same session, without calling `prepare_spark` multiple times, as the properties don’t interfere with each other.
If the data source is used in another API call, `prepare_spark` gets implicitly invoked, for example,
when a user materialises a training dataset using a data source or uses the data source to set up an External Feature Group.
So users do not need to call `prepare_spark` every time they do an operation with a connector, it is only necessary when reading directly using Spark.
Using `prepare_spark` is also not necessary when using the `read` API.
For example, to read directly from a S3 connector, we use the `prepare_spark` as follows:
=== "PySpark"
```python
connector.prepare_spark()
spark.read.format("json").load("s3a://[bucket]/path")
# or
spark.read.format("json").load(connector.prepare_spark("s3a://[bucket]/path"))
```
### Data warehouse/SQL based connectors
For data sources accessed via SQL such as data warehouses and JDBC compliant databases, e.g., Redshift, Snowflake, BigQuery, JDBC, users pass the SQL query to read the data to the `query` argument.
In most cases, this will be some form of a `SELECT` query.
Depending on the connector type, users can also just set the table path and read the whole table without explicitly passing any SQL query to the `query` argument.
This is mostly relevant for Google BigQuery.
=== "PySpark"
```python
# read results from a SQL
df = connector.read(query="SELECT * FROM TABLE")
# or directly read a table if set on connector
df = connector.read()
```
=== "Scala"
```scala
// read results from a SQL
val df = connector.read("SELECT * FROM TABLE", "" , new HashMap(),"")
```
### Streaming based connector
For reading data streams, the Kafka Data Source supports reading a Kafka topic into Spark Structured Streaming Dataframes instead of a static Dataframe as in other connector types.
=== "PySpark"
```python
df = connector.read_stream(topic="kafka_topic_name")
```
## Creating an External Feature Group
Another important aspect of a data source is its ability to facilitate creation of external feature groups with the [Connector API](../../../concepts/fs/feature_group/external_fg.md). [External feature groups](../feature_group/create_external.md) are basically offline feature groups and essentially stored as tables on external data sources.
The `Connector API` relies on data sources behind the scenes to integrate with external datasource.
This enables seamless integration with any data source as long as there is a data source defined.
To create an external feature group, we use the `create_external_feature_group` API, also known as `Connector API`, and simply pass the data source created before to the `data_source` argument.
Depending on the external source, we should set either the `query` argument for data warehouse based sources, or the `path` and `data_format` arguments for data lake based sources, similar to reading into dataframes as explained in above section.
Example for any data warehouse/SQL based external sources, we set the desired SQL to `query` argument, and set the `data_source` argument to the data source object of desired data source.
=== "PySpark"
```python
ds.query = "SELECT * FROM TABLE"
fg = feature_store.create_external_feature_group(
name="sales",
version=1,
description="Physical shop sales features",
data_source=ds,
primary_key=["ss_store_sk"],
event_time="sale_date",
)
```
`Connector API` (external feature groups) only stores the metadata about the features within Hopsworks, while the actual data is still stored externally.
This enables users to create feature groups within Hopsworks without the hassle of data migration.
For more information on `Connector API`, read detailed guide about [external feature groups](../feature_group/create_external.md).
When creating an external feature group from the UI, the review step also offers an **Add Trino Catalog** checkbox, which makes the same data source queryable from the query engine; see [Trino catalogs][trino-catalogs].
## Ingesting Data into a Managed Feature Group
Data Sources can also be used to create a managed feature group and ingest data from the source into Hopsworks.
In this workflow, Hopsworks creates a sink-enabled feature group together with an ingestion job that copies data from the source into the feature group.
This is different from an external feature group:
- An **external feature group** keeps the data in the external source and stores only metadata in Hopsworks.
- A **managed feature group with ingestion enabled** copies the source data into Hopsworks and can keep it synchronized through recurring ingestion jobs.
This workflow is especially useful when you want to:
- Materialize source data inside Hopsworks.
- Schedule recurring ingestions.
- Use full-load or incremental ingestion strategies.
- Build managed feature groups from SQL, CRM, or REST API sources.
For the full workflow, including schema selection, ingestion job configuration, loading strategies, and REST pagination, see [Ingest Data with dltHub][ingest-data-with-dlthub].
## Writing Training Data
Data Sources are also used while writing training data to external sources.
While calling the [Feature View](../../../concepts/fs/feature_view/fv_overview.md) API `create_training_data`, we can pass the `data_source` argument which is necessary to materialise the data to external sources, as shown below.
=== "PySpark"
```python
# materialise a training dataset
version, job = feature_view.create_training_data(
description="describe training data",
data_format="spark_data_format", # e.g., data_format = "parquet" or data_format = "csv"
write_options={"wait_for_job": False},
data_source=ds,
)
```
For a detailed walkthrough on managing and utilizing training data, refer to the [training data guide](../feature_view/training-data.md).
## Next Steps
We have gone through the basic use cases of a data source.
For more details about the API functionality for any specific connector type, checkout the [API section][hsfs.storage_connector.StorageConnector].
================================================================================
# Feature Group User Guides
Source: https://docs.hopsworks.ai/latest/user_guides/fs/feature_group/
# Feature Group User Guides
A feature group is a table of features with a primary key and, usually, an event time.
These guides cover creating one, keeping its data correct, and managing it over time.
- :material-table-plus:{ .lg .middle } **Start here**
---
Create a feature group and insert a DataFrame.
The schema is inferred from the DataFrame on the first insert.
```python
fg = fs.get_or_create_feature_group(
name="transactions",
version=1,
primary_key=["tid"],
event_time="datetime",
)
fg.insert(df)
```
[Create a feature group](create.md) · [Data types and schema](data_types.md) · [Statistics](statistics.md)
:material-table-plus:{ .hops-role-ico } Create and write
{ .hops-role-cap }
- [Create a feature group](create.md)
Offline and online tables, primary keys, event time, partitioning.
- [External feature groups](create_external.md)
Read data that stays in a warehouse or object store.
- [Spine groups](create_spine.md)
Supply keys, event times and labels without storing features.
- [Ingest with dltHub](ingest_with_dlthub.md)
Load from external sources through dlt pipelines.
- [Data types and schema](data_types.md)
Type mapping, adding features, schema versions.
:material-check-decagram-outline:{ .hops-role-ico } Trust
{ .hops-role-cap }
- [Statistics](statistics.md)
What is computed on insert and how to configure it.
- [Data validation](data_validation.md)
Great Expectations on insert, then the advanced guide and best practices.
- [Feature monitoring](feature_monitoring.md)
Scheduled statistics and comparison to a reference window.
- [Online ingestion observability](online_ingestion_observability.md)
Track rows arriving in the online store.
:material-cog-outline:{ .hops-role-ico } Manage
{ .hops-role-cap }
- [On-demand transformations](on_demand_transformations.md)
Compute features at request time from request parameters.
- [Notifications](notification.md)
Emit change events to a Kafka topic.
- [Time to live](ttl.md)
Expire rows after a retention period.
- [Deprecate](deprecation.md)
Mark a group as retired without deleting it.
================================================================================
# Create
Source: https://docs.hopsworks.ai/latest/user_guides/fs/feature_group/create/
# How to create a Feature Group { #create-feature-group }
## Introduction
In this guide you will learn how to create and register a feature group with Hopsworks.
Feature groups are created from code with the Hopsworks APIs.
The UI does not offer a creation flow; created feature groups appear in the project's `Catalog` section, where you can browse, edit and share them.
## Prerequisites
Before you begin this guide we suggest you read the [Feature Group](../../../concepts/fs/feature_group/fg_overview.md) concept page to understand what a feature group is and how it fits in the ML pipeline.
## Create using the Hopsworks APIs
To create a feature group using the Hopsworks APIs, you need to provide a Pandas, Polars or Spark DataFrame.
The DataFrame will contain all the features you want to register within the feature group, as well as the primary key, event time and partition key.
### Create a Feature Group
The first step to create a feature group is to create the API metadata object representing a feature group.
Using the Hopsworks API you can execute:
#### Batch Write API
=== "PySpark"
```python
fg = feature_store.create_feature_group(
name="weather",
version=1,
description="Weather Features",
online_enabled=True,
primary_key=["location_id"],
partition_key=["day"],
event_time="event_time",
time_travel_format="DELTA",
)
```
You can read the full [`FeatureStore.create_feature_group`][hsfs.feature_store.FeatureStore.create_feature_group] documentation to get more details.
If you need to create a feature group with vector similarity search supported, refer to [the vector similarity guide](../vector_similarity_search.md#extending-feature-groups-with-similarity-search).
`name` is the only mandatory parameter of the `create_feature_group` and represents the name of the feature group.
In the example above we created the first version of a feature group named *weather*, we provide a description to make it searchable to the other project members, as well as making the feature group available online.
Additionally we specify which columns of the DataFrame will be used as primary key, partition key and event time.
Composite primary key and multi level partitioning is also supported.
The version number is optional, if you don't specify the version number the APIs will create a new version by default with a version number equals to the highest existing version number plus one.
The last parameter used in the examples above is `stream`.
The `stream` parameter controls whether to enable the streaming write APIs to the online and offline feature store.
When using `time_travel_format="HUDI"` in a Python environment this behavior is the default.
##### Primary key
A primary key is required when using a table format with time travel support (Hudi, Delta, or Iceberg) to store offline feature data.
When inserting data in a feature group on the offline feature store, the DataFrame you are writing is checked against the existing data in the feature group.
If a row with the same primary key is found in the feature group, the row will be updated.
If the primary key is not found, the row is appended to the feature group.
When writing data on the online feature store, existing rows with the same primary key will be overwritten by new rows with the same primary key.
##### Event time
The event time column represents the time at which the event was generated.
For example, with transaction data, the event time is the time at which a given transaction happened.
In the context of feature pipelines, the event time is often also the end timestamp of the interval of events included in the feature computation.
For example, computing the feature "number of purchases by customer last week", the event time should be the last day of this "last week" window.
The event time is added to the primary key when writing to the offline feature store.
This will make sure that the offline feature store has the entire history of feature values over time.
As an example, if a user has made multiple purchases on a website, each of the purchases for a given user (identified by a user_id) will be saved in the feature group, with each purchase having a different event time (the combination of user_id and event_time makes up the primary key for the offline feature store).
The event time **is not** part of the primary key when writing to the online feature store.
This will ensure that the online feature store has the most recent version of the feature vector for each primary key.
!!!note "Event time data type restriction"
The supported data types for the event time column are: `timestamp`, `date` and `bigint`.
##### Partition key
It is best practice to add a partition key.
When you specify a partition key, the data in the feature group will be stored under multiple directories based on the value of the partition column(s).
All the rows with a given value as partition key will be stored in the same directory.
Choosing the correct partition key has significant impact on the query performance as the execution engine (Spark) will be able to skip listing and reading files belonging to partitions which are not included in the query.
As an example, if you have partitioned your feature group by day and you are creating a training dataset that includes only the last year of data, Spark will read only 365 partitions and not the entire history of data.
On the other hand, if the partition key is too fine grained (e.g., timestamp at millisecond resolution) - a large number of small partitions will be generated.
This will slow down query execution as Spark will need to list and read a large amount of small directories/files.
If you do not provide a partition key, all the feature data will be stored as files in a single directory.
The system has a limit of 10240 direct children (files or other subdirectories) per directory.
This means that, as you add new data to a non-partitioned feature group, new files will be created and you might reach the limit.
If you do reach the limit, your feature engineering pipeline will fail with the following error:
```sh
MaxDirectoryItemsExceededException - The directory item limit is exceeded: limit=10240 items=10240
```
By using partitioning the system will write the feature data in different subdirectories, thus allowing you to write 10240 files per partition.
`partition_key` is plain identity partitioning on existing columns.
For partition transforms such as `day(ts)` or `bucket(16, customer_id)` on Iceberg and Hudi, liquid clustering on Delta (`clustered_by`), the Hudi bucket index (`bucket_index`), and z-ordering (`zorder_by`), see the [partitioning and clustering guide][partitioning-feature-group].
##### Table format
When you create a feature group, you can specify the table format you want to use to store the data in your feature group by setting the `time_travel_format` parameter.
The currently supported values are `"HUDI"`, `"DELTA"`, `"ICEBERG"`, and `"NONE"` (which stores as Parquet without time travel support).
The parameter defaults to `"DELTA"`.
The feature group overview in the UI shows a **Table DDL** card with the generated Spark SQL `CREATE TABLE` statement for the offline table (including the table format and any partition columns), and, for online-enabled feature groups, the `CREATE TABLE` statement for the online (RonDB) table.
##### Data Source
During the creation of a feature group, it is possible to define the `data_source` parameter, this allows for management of offline data in the desired table format outside the Hopsworks cluster.
Currently, [S3][data-source-s3] and [GCS][data-source-gcs] connectors with `"DELTA"` or `"ICEBERG"` `time_travel_format` are supported.
##### Online Table Configuration
When defining online-enabled feature groups it is also possible to configure the online table.
You can specify [table options](https://docs.rondb.com/table_options/#table-options) by providing comments.
Additionally, it is also possible to define whether online data is stored in memory or on disk using [table space](https://docs.rondb.com/disk_columns/#disk-columns).
The code example shows the creation of an online-enabled feature group that stores online data on disk using `ts_1` table space and sets several table properties in the comment section.
```python
fg = fs.create_feature_group(
name="air_quality",
description="Air Quality characteristics of each day",
version=1,
primary_key=["city", "date"],
online_enabled=True,
online_config={
"table_space": "ts_1",
"online_comments": [
"NDB_TABLE=READ_BACKUP=1",
"NDB_TABLE=PARTITION_BALANCE=FOR_RP_BY_LDM_X_2",
],
},
)
```
!!! note Table Space
The table space needs to be provisioned at system level before it can be used.
You can do so by adding the following parameters to the values.yaml file used for your deployment with the Helm Charts:
```yaml
rondb:
resources:
requests:
storage:
diskColumnGiB: 2
```
#### Streaming Write API
As explained above, the stream parameter controls whether to enable the streaming write APIs to the online and offline feature store.
For Python environments, only the stream API is supported (stream=True).
=== "Python"
```python
fg = feature_store.create_feature_group(
name="weather",
version=1,
description="Weather Features",
online_enabled=True,
primary_key=["location_id"],
partition_key=["day"],
event_time="event_time",
time_travel_format="HUDI",
)
```
=== "PySpark"
```python
fg = feature_store.create_feature_group(
name="weather",
version=1,
description="Weather Features",
online_enabled=True,
primary_key=["location_id"],
partition_key=["day"],
event_time="event_time",
time_travel_format="HUDI",
stream=True,
)
```
When using the streaming API, the data will be written directly to the online storage (if `online_enabled=True`).
However, you can control when the sync to
the offline storage is going to happen.
You can do it synchronously after every call to `fg.insert()`, which is the default.
Often, you defer writes to a later point in order to batch together multiple writes to the offline storage (useful to reduce the overhead of many small writes):
```python
# run multiple inserts without starting the offline materialization job
job, _ = fg.insert(df1, write_options={"start_offline_materialization": False})
job, _ = fg.insert(df2, write_options={"start_offline_materialization": False})
job, _ = fg.insert(df3, write_options={"start_offline_materialization": False})
# start the materialization job for all three inserts
# note the job object is always the same, you don't need to call it three times
job.run()
```
It is also possible to define the topics used for data ingestion, this can be done by setting the `topic_name` parameter with your preferred value.
By default, feature groups in Hopsworks will share a project-wide topic.
The topic can also be changed after the feature group has been created, see the [ingestion topic][feature-group-ingestion-topic] guide.
#### Best Practices for Writing
When designing a feature group, it is worth taking a look at how this feature group will be queried in the future, in order to optimize it for those query patterns.
At the same time, Spark and Hudi tend to overpartition writes, creating too many small parquet files, which is inefficient and slows down writes.
But they also slow down queries, because file listings take more time and reading many small files is slower than fewer larger files.
The best practices described in this section hold both for the Streaming API and the Batch API.
Four main considerations influence the write and the query performance:
1. Partitioning on a feature group level
2. Parquet file size within a feature group partition
3. Backfilling of feature group partitions
4. The choice of topic for data ingestion
##### Partitioning on a feature group level
**Partitioning on the feature group level** allows Hopsworks and the table format (Hudi, Delta, or Iceberg) to push down filters to the filesystem when reading from feature groups.
In practice that means fewer directories need to be listed and fewer files need to be read, speeding up queries.
For example, most commonly, filtering is done on the event time column of a feature group when generating training data or batches of data:
```python
query = fg.select_all()
# create a simple feature view
fv = fs.create_feature_view(name="transactions_view", query=query)
# set up dates
start_time = "2022-01-01"
end_time = "2022-06-30"
# create a training dataset
version, job = fv.create_training_data(
start_time=start_time,
end_time=end_time,
description="Description of a dataset",
)
```
Assuming the feature group was partitioned by a daily event time column, for example, the features are updated with a daily batch job, the feature store will only have to
list and read the files in the directories of those six months that are being queried.
!!! danger "Too granular event time columns"
An event time column which is too granular, such as a timestamp, shouldn't be used as partition key.
For example, a streaming pipeline generating features where the event time includes seconds, and therefore almost all
event timestamps are unique can lead to many partition directories and small files, each of which contains only a few number of rows,
which are inefficient to query even with pushed down filters.
A good practice are partition keys with at most daily granularity, if they are based on time.
Additionally, one can look at the size of a partition directory, which should be in the 100s of MB.
Additionally, if you are commonly training models for different categories of your data, you can add another level of partitioning for this.
That is, if the query contains
an additional filter:
```python
query = fg.select_all().filter(fg.country_code == "US")
```
The feature group can be created with the following partition key in order to push down filters also for the `country_code` category:
```python
fg = feature_store.create_feature_group(...
partition_key=['day', 'country_code'],
event_time='day',
)
```
##### Parquet file size within a feature group partition
Once you have decided on the feature group level partitioning and you start inserting data to the feature group, there are multiple ways in order to influence how the table format (Hudi, Delta, or Iceberg) will **split the data between parquet files within the feature group partitions**.
The two things that influence the number of parquet files per partition are
1. The number of feature group partitions written in a single insert
2. The shuffle parallelism used by the table format
For example, the inserted dataframe (unique combination of partition key values) will be parallelized according to the following Hudi settings:
!!! example "Default Hudi partitioning"
```python
write_options = {
"hoodie.bulkinsert.shuffle.parallelism": 5,
"hoodie.insert.shuffle.parallelism": 5,
"hoodie.upsert.shuffle.parallelism": 5,
}
```
That means, using Spark, Hudi shuffles the data into five in-memory partitions, which each fill map to a task and finally a parquet file (see figure below).
If the inserted Dataframe contains only a single feature group partition, this feature group partition will be written with five parquet files.
If the inserted Dataframe contains multiple feature group partitions, the parquet files will be split among those partition, potentially more parquet files will be added.
--8<-- "user_guides/fs/feature_group/create/partition-files.html"
!!! tip "Setting shuffle parallelism"
In practice that means the shuffle parallelism should be set equal to the number of feature group partitions in the inserted dataframe.
This will create one parquet file per feature group partition, which in many cases is optimal.
Theoretically, this rule holds up to a partition size of 2GB, which is the limit of Spark.
However, one should bump this up accordingly already for smaller inputs.
We recommend having shuffle parallelism `hoodie.[insert|upsert|bulkinsert].shuffle.parallelism` such that it's at least input_data_size/500MB.
You can change the write options on every insert, depending also on the size of the data you are writing:
```python
write_options = {
"hoodie.bulkinsert.shuffle.parallelism": 5,
"hoodie.insert.shuffle.parallelism": 5,
"hoodie.upsert.shuffle.parallelism": 5,
}
fg.insert(df, write_options=write_options)
```
##### Backfilling of feature group partitions
Hudi scales well with the number of partitions to write, when performing backfilling of old feature partitions, meaning moving backwards in time with the event-time, it makes sense to **batch those feature group partitions** together into a single `fg.insert()` call.
As shown in the figure above, the number of utilised executors you choose for the insert depends highly on the number of partitions and shuffle parallelism you are writing.
So by writing multiple feature group partitions in a single insert, you can scale up your Spark application and fully utilise the workers.
In that case you can increase the Hudi shuffle parallelism accordingly.
!!! danger "Concurrent feature group inserts"
Hopsworks 3.1 and earlier, currently does not support concurrent inserts to feature groups.
This means that if your feature pipeline writes to one feature group partition at a time,
you cannot run it multiple times in parallel for backfilling.
The recommended approach is to unionise the dataframes and insert them with a single `fg.insert()` instead.
For clients that write with the Stream API, it is enough to defer starting the backfill job until after multiple inserts,
[as described above](#streaming-write-api).
##### The choice of topic for data ingestion
When creating a feature group that uses streaming write APIs for data ingestion it is possible to define the Kafka topics that should be utilized.
The default approach of using a project-wide topic functions great for use cases involving little to no overlap when producing data.
However, concurrently inserting into multiple feature groups could cause read amplification for the offline materialization job (e.g., Hudi Delta Streamer).
The job of a feature group consumes every record the shared topic received since its last run, and only then discards the records whose `featureGroupId` header belongs to another feature group.
One large or frequently written feature group therefore slows down the materialization job of every other feature group on its topic, in proportion to how much it writes.
Therefore, it is advised to utilize separate topics when ingestions overlap or there is a large frequently running insertion into a specific feature group.
If you only notice the read amplification once the feature group is in use, the [ingestion topic][feature-group-ingestion-topic] guide explains how to move it to its own topic.
### Register the metadata and save the feature data
The snippet above only created the metadata object on the Python interpreter running the code.
To register the feature group metadata and to save the feature data with Hopsworks, you should invoke the `insert` method:
```python
fg.insert(df)
```
The save method takes in input a Pandas, Polars or Spark DataFrame.
Hopsworks will use the DataFrame columns and types to determine the name and types of features, primary key, partition key and event time.
The DataFrame *must* contain the columns specified as primary keys, partition key and event time in the `create_feature_group` call.
If a feature group is online enabled, the `insert` method will store the feature data to both the online and offline storage.
!!! api "API reference"
- [`FeatureStore.create_feature_group`][hsfs.feature_store.FeatureStore.create_feature_group]
- [`FeatureStore.get_or_create_feature_group`][hsfs.feature_store.FeatureStore.get_or_create_feature_group]
- [`FeatureGroup`][hsfs.feature_group.FeatureGroup]
- [`insert`][hsfs.feature_group.FeatureGroup.insert]
- [`read`][hsfs.feature_group.FeatureGroup.read]
- [`select_all`][hsfs.feature_group.FeatureGroupBase.select_all]
- [`filter`][hsfs.feature_group.FeatureGroupBase.filter]
Browse the full Python API :material-arrow-right:
## Find your feature group in the UI
Feature groups created through the APIs appear in the `Catalog` section of the project sidebar.
From there you can inspect features and statistics, edit metadata, and manage sharing and tags.
The Catalog lists every feature group with its table format, online status and version.
================================================================================
# Partitioning and Clustering
Source: https://docs.hopsworks.ai/latest/user_guides/fs/feature_group/partitioning/
# How to partition and cluster a Feature Group { #partitioning-feature-group }
## Introduction
In this guide you will learn how to lay out the offline data of a feature group.
Each layout mechanism is its own creation-time parameter, and each parameter maps to exactly one native mechanism of the table format:
| Parameter | What it configures | Formats |
| --- | --- | --- |
| `partitioned_by` | A native partition specification from transform expressions, such as `["day(ts)", "bucket(16, customer_id)"]`. | `ICEBERG` (full transform set), `HUDI` (identity and time grains). Rejected on `DELTA`, which has no partition transforms. |
| `clustered_by` | Delta liquid clustering columns, such as `["ts", "customer_id"]`. | `DELTA` only. |
| `bucket_index` | The Hudi bucket index, as `{"field": , "num_buckets": N}`. | `HUDI` only. |
| `zorder_by` | Columns to z-order the data files by. | `ICEBERG` (applied by [`FeatureGroup.optimize`][hsfs.feature_group.FeatureGroup.optimize]), `HUDI` (inline clustering). Rejected on `DELTA`, where `clustered_by` covers the use case. |
| `sort_order` | A persistent write sort order, such as `["merchant_id asc", "amount desc nulls last"]`. | `ICEBERG` only, and requires a Spark environment at creation. |
| `partition_key` | Plain identity partitioning on existing columns. | All formats; mutually exclusive with `partitioned_by` and with `clustered_by`. |
## Prerequisites
Before you begin this guide we suggest you read the [Feature Group](../../../concepts/fs/feature_group/fg_overview.md) concept page and the [create feature group][create-feature-group] guide, which covers `partition_key`, `event_time`, and `time_travel_format`.
## Partition transforms
Each element of `partitioned_by` is one transform expression:
| Expression | Behaviour | Typical use |
| --- | --- | --- |
| `col` or `identity(col)` | The original column value. | Low-cardinality columns. |
| `bucket(N, col)` | Murmur3 hash of the column modulo N. | High-cardinality ids. |
| `truncate(W, col)` | The value truncated to width W (numeric ranges, string prefixes). | Range-like grouping. |
| `year(col)` | Year of a timestamp or date column. | Very large historical tables. |
| `month(col)` | Month. | Monthly queries. |
| `day(col)` | Calendar day. | Event and transaction tables. |
| `hour(col)` | Hour of a timestamp column. | High-volume streams. |
| `week(col)` | ISO week. | HUDI only: Iceberg has no week transform. |
| `void(col)` | Always null. | ICEBERG only: partition spec evolution placeholder. |
Parsing is whitespace tolerant and the expressions are stored in a canonical lowercase form without spaces, so `bucket(16, customer_id)` reads back as `bucket(16,customer_id)`.
A column that happens to be named after a grain (`year`, `month`, `week`, `day`, `hour`) must use the explicit form `identity(year)`, because the bare name is reserved for the legacy grain migration error; the canonical form keeps the explicit `identity()` for these columns.
Transforms combine freely, for example one temporal transform plus one bucket, with the format-specific restrictions listed below.
On Iceberg an expression can name its partition field with an alias, for example `"bucket(16, customer_id) as shard"`; Iceberg metadata tables and partition spec evolution refer to fields by name.
Without an alias the field gets Iceberg's generated default name (`customer_id_bucket`, `ts_day`, and so on), and all field names in one spec must be unique, so two bucket transforms on the same column need an alias on one of them.
Aliases are rejected on `HUDI`, whose partition paths are always named after the source column or grain.
Not every transform is available on every format, and `DELTA` rejects `partitioned_by` entirely:
| Transform | ICEBERG | HUDI |
| --- | --- | --- |
| `identity` | yes | yes |
| `bucket` | yes | no, use `bucket_index` |
| `truncate` | yes | no |
| `year`/`month`/`day`/`hour` | yes | yes, and the column must be the event time |
| `week` | no | yes, and the column must be the event time |
| `void` | yes | no |
## Iceberg: hidden partitioning
On Iceberg the transform list compiles into the table's native partition spec.
No derived columns are added to the feature group schema, and your DataFrame carries only the real columns.
Because the partitioning is hidden, every engine reading the table prunes partitions from ordinary predicates on the source columns.
```python
fg = feature_store.create_feature_group(
name="orders",
version=1,
description="Order events, day-partitioned and bucketed by customer",
primary_key=["order_id"],
event_time="order_ts",
time_travel_format="ICEBERG",
partitioned_by=["day(order_ts)", "bucket(16, customer_id)"],
)
fg.insert(orders_df)
```
A read that filters on the source columns plans only the matching day partition and, within it, one bucket of sixteen:
```python
df = (
fg.select_all()
.filter(fg.get_feature("order_ts") >= "2026-07-03")
.filter(fg.get_feature("order_ts") < "2026-07-04")
.filter(fg.get_feature("customer_id") == 1042)
.read()
)
```
Iceberg allows at most one temporal transform per source column, because a finer grain already supports coarser pruning: use `day(ts)` alone rather than `year(ts)` plus `day(ts)`.
### Persistent sort order
`sort_order` sets the Iceberg table's persistent write sort order, so every write organizes rows before laying out files.
Each element is a column with an optional direction and null ordering: `"col"`, `"col desc"`, `"col asc nulls last"`.
Defaults follow Iceberg: ascending, nulls first when ascending and nulls last when descending.
```python
fg = feature_store.create_feature_group(
name="payments",
version=1,
primary_key=["payment_id"],
event_time="ts",
time_travel_format="ICEBERG",
partitioned_by=["day(ts)"],
sort_order=["merchant_id asc", "amount desc nulls last"],
)
```
A sort order is applied through the Iceberg Java API, so creating a feature group with `sort_order` requires a Spark environment; the pure Python and catalog creation paths reject it rather than silently dropping it.
`sort_order` is the right choice when reads consistently filter or join on the same columns; `zorder_by` covers multi-dimensional point lookups instead.
The two are mutually exclusive as stored defaults, because they prescribe conflicting file layouts (writes sorted linearly while maintenance rewrites on the z-curve); a one-off z-order on a sorted table stays available through `optimize(strategy="zorder", columns=[...])`.
### Layout for point-in-time training data
Training data from a feature view runs a point-in-time join: each feature group is joined to the label side on the primary key with an event-time inequality, and a rank window keeps the latest row per label.
Inside that query shape, an event-time bound on the feature group is pushed into the Iceberg scan, so `day(event_ts)` partitioning prunes the history scan to the bounded window.
Set the bound with a `lookback` on the feature view read (or an explicit `event_ts >=` query filter); without one, point-in-time correctness requires scanning all history, and no partitioning can prune it.
The layout that serves this access pattern:
```python
fg = feature_store.create_feature_group(
...,
time_travel_format="ICEBERG",
partitioned_by=["day(event_ts)", "bucket(16, customer_id)"],
sort_order=["customer_id asc", "event_ts desc"],
)
```
- `day(event_ts)` prunes the scan to the lookback window.
- `bucket(N, primary_key)` keeps each day's data grouped by key, bounding the rows any one join task reads.
- `sort_order` (or `zorder_by` plus a scheduled `optimize()`) clusters rows by key inside each file, so file-level min/max statistics skip files for keys not present in the label set.
- Prefer `insert` (append) over upserts for event history: the Iceberg upsert rewrites the table and discards the maintained clustering until the next `optimize()`.
Size the partitioning to the data volume: each `(day, bucket)` combination becomes at least one file, so a small feature group with fine-grained partitioning produces many tiny files and the task overhead outweighs the pruning.
As a rule of thumb, choose the day grain and bucket count so partitions land in the hundreds of megabytes; for small feature groups skip `bucket()` or use a coarser time grain.
## Delta: liquid clustering with clustered_by
Delta has no partition transforms, so `partitioned_by` is rejected there; the layout mechanism is liquid clustering, configured with `clustered_by` as a plain column list.
This matches Delta's own API, where `clusterBy(...)` and `partitionedBy(...)` are different layout mechanisms.
Liquid clustering supports at most 4 columns, and `clustered_by` cannot be combined with `partition_key`, because Delta does not support clustering a hive-partitioned table.
```python
fg = feature_store.create_feature_group(
name="transactions",
version=1,
primary_key=["tx_id"],
event_time="ts",
time_travel_format="DELTA",
clustered_by=["ts", "customer_id"],
)
```
Data skipping on the clustered columns replaces directory-style partitioning, so no derived columns are added and reads need no special predicates.
Delta only clusters on columns that have data-skipping statistics, which it collects for the first 32 columns by default; when a clustering column sits past that position Hopsworks widens `delta.dataSkippingNumIndexedCols` to cover it, so any schema column can be clustered.
!!!warning "Clustered Delta feature groups are writable by Spark only"
Liquid clustering uses the Clustering and DomainMetadata Delta writer table features, which delta-rs (the pure Python write path) does not implement, so treating them as optional would corrupt the table contract.
Creating a clustered feature group from a Python environment with `stream=False` fails, and delta-rs writes, deletes, and optimize raise on clustered feature groups.
From Python, pass `stream=True` so writes go through the Spark materialization job, or write from a Spark job or notebook.
## Hudi: grain columns and the bucket index
Hudi has no hidden partitioning, so temporal transforms materialize as integer partition columns named after the grain (`year`, `month`, ...), derived from the event time on every write.
Your DataFrame must not contain these columns; the write path computes them.
The grain columns appear in the feature group schema flagged as partition columns, and by default they are stored offline only (set `online_partition_columns=True` to include them online).
Bucketing on Hudi is not a partition transform: it is the Hudi bucket index, which hashes a primary key field into a fixed number of buckets, configured with the `bucket_index` parameter:
```python
fg = feature_store.create_feature_group(
name="transactions",
version=1,
primary_key=["tx_id"],
event_time="ts",
time_travel_format="HUDI",
partitioned_by=["year(ts)", "month(ts)"],
bucket_index={"field": "tx_id", "num_buckets": 16},
)
```
`bucket_index` injects the corresponding write options:
```text
hoodie.index.type=BUCKET
hoodie.bucket.index.num.buckets=N
hoodie.bucket.index.hash.field=col
```
The optional `"engine"` key accepts only `"simple"` (the default): Hudi's consistent-hashing bucket engine requires a merge-on-read table with a clustering lifecycle, and Hopsworks Hudi feature groups are copy-on-write.
You can equally set these options, including a partition-level bucket index with per-partition `hoodie.bucket.index.*` overrides, directly through `write_options` on insert.
On Hudi the platform rewrites `event_time` range filters into grain-column predicates at query time, so reads that filter on the event time prune partitions without referencing the grain columns.
## Z-ordering with zorder_by
`zorder_by` records up to 4 columns to z-order the data files by, so point lookups on those columns inside a partition skip most files.
It is supported for `ICEBERG` and `HUDI`; on `DELTA` it is rejected because `clustered_by` covers the same use case.
```python
fg = feature_store.create_feature_group(
name="ad_clicks",
version=1,
description="Click stream, hour-partitioned, z-ordered by user and campaign",
primary_key=["click_id"],
event_time="click_ts",
time_travel_format="ICEBERG",
partitioned_by=["hour(click_ts)"],
zorder_by=["user_id", "campaign_id"],
)
fg.insert(clicks_df)
fg.optimize()
```
On Iceberg, z-order is a rewrite strategy rather than a write-time property: writes stay cheap, and calling [`FeatureGroup.optimize`][hsfs.feature_group.FeatureGroup.optimize] rewrites the data files ordered on the z-curve of the `zorder_by` columns.
Run it from a scheduled job after heavy ingestion.
The initial z-order over a backfill needs `optimize(rewrite_all=True)` to rewrite every existing file; routine maintenance calls default to `rewrite_all=False` so they never rewrite the whole table by accident.
On Hudi, `zorder_by` configures inline clustering, which applies the z-order layout as part of the write pipeline, so no explicit call is needed.
## Optimizing the layout
[`FeatureGroup.optimize`][hsfs.feature_group.FeatureGroup.optimize] rewrites the offline data files to apply the feature group's layout and returns the format's rewrite metrics:
| `time_travel_format` | What `optimize()` does |
| --- | --- |
| `ICEBERG` | An Iceberg `rewriteDataFiles` action; requires a Spark environment. `strategy` picks `"zorder"` (over `columns`, defaulting to `zorder_by`), `"sort"` (the persistent `sort_order`), or `"binpack"`; unset, it follows the stored layout in that order. `rewrite_all=True` rewrites every file regardless of the planner thresholds (defaults to False for every strategy, so a routine call is incremental; pass it for the initial full z-order), `target_file_size_mb` overrides the target file size, and `where` restricts the rewrite to the matching files through an Iceberg filter expression over the feature group's columns. |
| `DELTA` | `OPTIMIZE`, which incrementally clusters a liquid-clustered table; `optimize(full=True)` runs `OPTIMIZE FULL` to recluster all existing data after the clustering columns changed (clustered tables only), and `where` restricts the rewrite with a predicate. `strategy="zorder"` with `columns` runs the legacy `OPTIMIZE ... ZORDER BY`, which Delta only supports on unclustered tables, because z-order and liquid clustering are incompatible. Clustered feature groups require Spark; from pure Python only unclustered compaction is available. |
| `HUDI` | Rejected: layout maintenance runs through inline clustering on writes. |
### Catalog-backed Iceberg tables
The Iceberg feature is fully supported on the default path-based (`HadoopTables`) layout. Some operations are not yet available when the table is backed by an external catalog:
| Operation | Path-based | Glue Data Catalog | User-provided catalog (`iceberg.catalog`) |
| --- | --- | --- | --- |
| Create with `partitioned_by` | yes | yes | yes |
| Create with `sort_order` | yes | no (rejected at creation) | no |
| `optimize()` | yes | no (run the catalog's `rewrite_data_files` procedure) | no |
| `update_partition_spec()` | yes | no (evolve through the catalog) | no |
| Introspection (`get_partition_spec`, `describe_layout`, ...) | yes | yes | no (inspect through the catalog) |
Where an operation is unavailable the call raises with a pointer to the catalog-side equivalent rather than silently doing nothing.
## Evolving the layout
Layout is not fixed at creation; each format's native evolution is exposed, and all three methods require a Spark environment (from pure Python they raise, run them from a Spark job or notebook):
- [`FeatureGroup.update_partition_spec`][hsfs.feature_group.FeatureGroup.update_partition_spec] evolves an Iceberg partition spec.
Evolution is metadata-only: existing data keeps its old layout and new writes use the evolved spec, so no data is rewritten.
```python
fg.update_partition_spec(add=["hour(ts)"], remove=["day(ts)"])
```
- [`FeatureGroup.update_clustering`][hsfs.feature_group.FeatureGroup.update_clustering] changes the Delta clustering columns.
The change affects new writes; run `optimize(full=True)` to recluster existing data.
```python
fg.update_clustering(["ts", "customer_id"])
fg.optimize(full=True)
```
- [`FeatureGroup.disable_clustering`][hsfs.feature_group.FeatureGroup.disable_clustering] turns Delta clustering off (`CLUSTER BY NONE`); existing data keeps its layout.
Hudi partitions are physical directories and cannot evolve; `update_partition_spec` is rejected there.
Partition spec evolution requires a feature group created with `partitioned_by`; a feature group using `partition_key` is rejected, because its identity partitions are recorded on the features themselves and cannot be restated as an evolvable spec.
After each evolution the committed spec is read back from the table and persisted as the feature group's stored metadata (`partitioned_by`, `clustered_by`), so read-back always reflects the current layout, including changes made by external engines.
If the table change succeeds but the metadata update fails, the error says exactly that and names the applied layout; calling `update_partition_spec()` with no arguments re-syncs the stored metadata from the table without changing it.
## Inspecting the layout
The stored metadata describes what was requested; the table itself is the source of truth for what is physically there.
Five methods read the actual table state, so drift (for example, an external engine evolving the spec) is visible:
- [`FeatureGroup.get_partition_spec`][hsfs.feature_group.FeatureGroup.get_partition_spec] returns the current Iceberg partition spec fields (`ICEBERG` only).
- [`FeatureGroup.get_partition_specs`][hsfs.feature_group.FeatureGroup.get_partition_specs] returns the Iceberg spec history, oldest first (`ICEBERG` only).
- [`FeatureGroup.get_sort_order`][hsfs.feature_group.FeatureGroup.get_sort_order] returns the persistent sort order (`ICEBERG` only).
- [`FeatureGroup.get_clustering_columns`][hsfs.feature_group.FeatureGroup.get_clustering_columns] returns the actual Delta clustering columns (`DELTA` only, Spark required).
- [`FeatureGroup.describe_layout`][hsfs.feature_group.FeatureGroup.describe_layout] returns the stored metadata next to the actual table state for any format.
The format-specific getters raise on the wrong format rather than returning `None`, so an unconfigured layout is never confused with an unsupported operation.
For Iceberg feature groups written through a user-provided catalog (the `iceberg.catalog` write option), the current-metadata pointer lives in that catalog, so inspect the layout through the catalog instead; Glue-backed feature groups are inspected through the Glue Data Catalog automatically.
## Migrating from the grain list form
Earlier Hopsworks versions accepted `partitioned_by` as a list of bare grain names, for example `["year", "month"]`.
That form is no longer valid, because a bare element now means an identity transform on a column of that name.
Creation fails with an error pointing at the replacement: write the grain as a transform on your event time column, for example `["year(event_ts)", "month(event_ts)"]`.
Feature groups created with the old form keep working for reads; their stored layout is unchanged.
Writes and deletes to such feature groups fail with a migration-required error, because the new write path no longer derives the grain values and continuing would silently change the physical layout; recreate the feature group with the transform grammar to write again.
## Restrictions
- `partitioned_by` requires `time_travel_format` `ICEBERG` or `HUDI`; `clustered_by` requires `DELTA`; `bucket_index` requires `HUDI`; `zorder_by` requires `ICEBERG` or `HUDI`; `sort_order` requires `ICEBERG` and a Spark environment at creation.
- `partitioned_by` and `partition_key` cannot be combined, and neither can `clustered_by` and `partition_key`, nor `sort_order` and `zorder_by`.
- Temporal transforms require a `date` or `timestamp` source column, and `hour` requires a `timestamp`, because a date has no sub-day resolution.
- `bucket` and `truncate` follow the Iceberg source-type restrictions: `float`, `double`, and `boolean` sources are rejected, and `truncate` additionally rejects `date` and `timestamp`.
- Partition-field aliases are Iceberg-only, and all field names in one spec (explicit aliases and generated defaults) must be unique.
- The `bucket_index` field must be part of the primary key, because the Hudi bucket index hashes the record key; its `engine` accepts only `"simple"`.
- Clustered Delta feature groups are writable by Spark only; from Python use `stream=True`.
- Stream feature groups support `partitioned_by` and `clustered_by` on `ICEBERG` and `DELTA`, but `partitioned_by` is rejected on `HUDI`.
- Online-enabled feature groups support `partitioned_by` on `ICEBERG`, but not on `HUDI`, because the Hudi grain columns are not part of the online schema.
================================================================================
# Delta Maintenance
Source: https://docs.hopsworks.ai/latest/user_guides/fs/feature_group/delta_maintenance/
# How to maintain a Delta Feature Group { #delta-maintenance-feature-group }
## Introduction
A Delta table that is written to repeatedly accumulates two things: data files and log entries.
Every commit writes at least one new data file, and every reader opens all of them.
Every commit also appends to the `_delta_log`, and a reader replays that log from the last checkpoint.
Neither is reclaimed on its own, and on a table written from Python neither is bounded on its own either: Spark writes a checkpoint every `delta.checkpointInterval` commits, delta-rs writes none.
Four methods on a feature group bound them.
They apply only to feature groups with `time_travel_format="DELTA"` and return `None` for any other format.
| Method | What it does |
| --- | --- |
| `delta_optimize` | Rewrites many small files into fewer large ones. Also available as `delta_compact`. |
| `delta_checkpoint` | Writes a checkpoint, so readers stop replaying the log from commit zero. |
| `delta_cleanup_metadata` | Expires the log entries a checkpoint already covers. |
| [`delta_vacuum`][hsfs.feature_group.FeatureGroup.delta_vacuum] | Deletes the data files no retained version references. |
Each dispatches on the engine, so the same call works from a Python client with
delta-rs and from a PySpark job with Delta Spark. The first three are rendered as
plain code rather than API links until the client release that ships them, because
the docs build resolves cross-references against the released client.
## Prerequisites
Before you begin this guide we suggest you read the [Feature Group](../../../concepts/fs/feature_group/fg_overview.md) concept page and the [create feature group][create-feature-group] guide.
## The maintenance sequence
Run them in this order.
```python
fg = fs.get_feature_group("transactions", version=1)
fg.delta_optimize(max_concurrent_tasks=1)
fg.delta_checkpoint()
fg.delta_cleanup_metadata()
fg.delta_vacuum(retention_hours=168)
```
The order is what makes each step safe.
Compaction replaces many small files with few large ones and leaves the old ones on disk, still referenced by older versions.
The checkpoint goes next, so the smaller file list is recorded before anything is deleted.
Only then the two deletions: the log entries the checkpoint now covers, and the data files the compaction orphaned.
## Choosing a retention
`delta_vacuum` deletes files that versions inside the retention window no longer reference.
A query that is already running holds no lock on those files, so the retention has to stay comfortably longer than the longest query that runs against the group.
It is also the time travel window: a version whose files have been vacuumed cannot be read, which is why a compaction has to be followed by a checkpoint.
The effect of a short retention is not that a vacuum deletes more, but that it deletes sooner.
A run reclaims what earlier runs orphaned rather than its own rewrite, whose files are seconds old.
!!! warning "Delta's own floor"
Delta refuses a retention under seven days unless its retention check is disabled.
Hopsworks disables that check for you so a shorter retention takes effect, which means the value you pass is the value that applies.
Pick it against your own readers rather than relying on the engine to refuse a bad one.
## Compacting only what changed
On a table partitioned by a date column, `after_ingest_date` bounds the rewrite to partitions at or after that date.
```python
fg.delta_optimize(after_ingest_date="2026-09-10")
```
Use it for anything that runs on a schedule.
Only files written since the last compaction need rewriting, and on a date-partitioned table they are all at or after that date, so bounding the rewrite this way keeps its cost flat.
Without it every run rewrites the whole table, including everything earlier runs already compacted, and the cost grows with the table forever.
Leave a day of slack for rows that arrived late.
Only a partition column can select files without reading them, so this is refused on a group that is not partitioned by a date.
Compact the whole table by leaving `after_ingest_date` unset.
## When to run them
For an append-heavy table, compact when the active file count crosses a threshold and otherwise once a day.
Around 100 files is the low hundreds of megabytes at typical commit sizes, near the engine's own target file size.
Read the last compaction time from the table's own history rather than keeping state, so the schedule survives restarts and multiple writers.
These can run from a [Hopsworks job](../../projects/jobs/pyspark_job.md) on a schedule.
Compaction is the only one of the four that a deployment reading the same table notices: measured beside live traffic it roughly doubled p99 for the few seconds it ran, while the median moved by a tenth of a millisecond.
The other three sat where the deployment sat with nothing running.
================================================================================
# Create External
Source: https://docs.hopsworks.ai/latest/user_guides/fs/feature_group/create_external/
# How to create an External Feature Group { #create-external-feature-group }
## Introduction
In this guide you will learn how to create and register an external feature group with Hopsworks.
This guide covers creating an external feature group using the Hopsworks APIs as well as the user interface.
## Prerequisites
Before you begin this guide we suggest you read the [External Feature Group](../../../concepts/fs/feature_group/external_fg.md) concept page to understand what a feature group is and how it fits in the ML pipeline.
## Create using the Hopsworks APIs
### Retrieve the Data Source
To create an external feature group using the Hopsworks APIs you need to provide an existing [data source](../data_source/index.md).
=== "Python"
```python
ds = feature_store.get_data_source("data_source_name")
```
### Create an External Feature Group
The first step is to instantiate the metadata through the `create_external_feature_group` method.
Once you have defined the metadata, you can
[persist the metadata and create the feature group](#register-the-metadata) in Hopsworks by calling `fg.save()`.
#### SQL based external feature group
=== "Python"
```python
query = """
SELECT TO_NUMERIC(ss_store_sk) AS ss_store_sk
, AVG(ss_net_profit) AS avg_ss_net_profit
, SUM(ss_net_profit) AS total_ss_net_profit
, AVG(ss_list_price) AS avg_ss_list_price
, AVG(ss_coupon_amt) AS avg_ss_coupon_amt
, sale_date
, ss_store_sk
FROM STORE_SALES
GROUP BY ss_store_sk, sales_date
"""
fg = feature_store.create_external_feature_group(
name="sales",
version=1,
description="Physical shop sales features",
query=query,
data_source=ds,
primary_key=["ss_store_sk"],
event_time="sale_date",
)
fg.save()
```
#### Data Lake based external feature group
=== "Python"
```python
fg = feature_store.create_external_feature_group(
name="sales",
version=1,
description="Physical shop sales features",
data_format="parquet",
data_source=ds,
primary_key=["ss_store_sk"],
event_time="sale_date",
)
fg.save()
```
You can read the full [`FeatureStore.create_external_feature_group`][hsfs.feature_store.FeatureStore.create_external_feature_group] documentation for more details.
`name` is a mandatory parameter of the `create_external_feature_group` and represents the name of the feature group.
The version number is optional, if you don't specify the version number the APIs will create a new version by default with a version number equals to the highest existing version number plus one.
If the data source is defined for a data warehouse (e.g., JDBC, Snowflake, Redshift) you need to provide a SQL statement that will be executed to compute the features.
If the data source is defined for a data lake, the location of the data as well as the format need to be provided.
Additionally we specify which columns of the DataFrame will be used as primary key, and event time.
Composite primary keys are also supported.
### Register the metadata
In the snippet above it's important that the created metadata object gets registered in Hopsworks.
To do so, you should invoke the `save` method:
=== "Python"
```python
fg.save()
```
### Enable online storage
You can enable online storage for external feature groups, however, the sync from the external storage to Hopsworks online storage is not automatic and needs to be setup manually.
For an external feature group to be available online, during the creation of the feature group, the `online_enabled` option needs to be set to `True`.
=== "Python"
```python
external_fg = fs.create_external_feature_group(
name="sales",
version=1,
description="Physical shop sales features",
query=query,
data_source=ds,
primary_key=["ss_store_sk"],
event_time="sale_date",
online_enabled=True,
)
external_fg.save()
# read from external storage and filter data to sync to online
df = external_fg.read().filter(external_fg.customer_status == "active")
# insert to online storage
external_fg.insert(df)
```
The `insert()` method takes a DataFrame as parameter and writes it _only_ to the online feature store.
Users can select which subset of the feature group data they want to make available on the online feature store by using the [query APIs][hsfs.constructor.query.Query].
### Limitations
Hopsworks Feature Store does not support time-travel queries on external feature groups.
Additionally, support for `.read()` and `.show()` methods when using by the Python engine is limited to external feature groups defined on BigQuery and Snowflake and only through the ArrowFlight Server with DuckDB, which Hopsworks enables by default.
Nevertheless, external feature groups defined top of any data source can be used to create a training dataset from a Python environment invoking one of the following methods: [`FeatureView.create_training_data`][hsfs.feature_view.FeatureView.create_training_data], [`FeatureView.create_train_test_split`][hsfs.feature_view.FeatureView.create_train_test_split] or [`FeatureView.create_train_validation_test_split`][hsfs.feature_view.FeatureView.create_train_validation_test_split].
!!! api "API reference"
- [`FeatureStore.get_data_source`][hsfs.feature_store.FeatureStore.get_data_source]
- [`FeatureStore.create_external_feature_group`][hsfs.feature_store.FeatureStore.create_external_feature_group]
- [`ExternalFeatureGroup`][hsfs.feature_group.ExternalFeatureGroup]
- [`save`][hsfs.feature_group.ExternalFeatureGroup.save]
- [`insert`][hsfs.feature_group.ExternalFeatureGroup.insert]
- [`read`][hsfs.feature_group.ExternalFeatureGroup.read]
Browse the full Python API :material-arrow-right:
## Create using the UI
You can also create a new feature group through the UI.
For this, navigate to the `Data Sources` section and make sure you have a data source for the desired platform, or create a [new](../data_source/index.md) one.
Table browsing is available for database and warehouse sources such as Snowflake, BigQuery, Redshift and SQL databases; the built-in HopsFS and JDBC sources of a project do not offer it.
Open the data source with the pencil at the end of its row and click `Next: Select Tables` at the bottom of the form.
In the UI you can either select one or more tables or define a custom SQL query.
### Option A: Select tables
The database navigation structure depends on your specific data source.
You'll navigate through the appropriate hierarchy for your platform, such as Database → Schema → Table for Snowflake, or Project → Dataset → Table for BigQuery.
Select one or more tables. For each selected table, you must designate one or more columns as primary keys before proceeding.
You can also optionally select a single column as the event time for the row (supported types are timestamp, date and bigint), and edit names and data types of the individual columns you want to include.
`Preview Metadata` and `Preview Data` show the source schema and a sample of rows before you commit to anything.
### Option B: Define a SQL query
Instead of selecting a table, you can write a custom SQL query to define the feature group.
This is useful when you need to join multiple tables or apply transformations at read time.
Click `Fetch Schema` to resolve the columns of the query, then, as with the table option, designate one or more columns as primary keys, optionally pick an event time column and give the feature group a name.
Complete the creation by clicking `Next: Review Configuration` at the bottom of the page.
As the last step, you will be able to rename the feature groups and confirm their creation.
================================================================================
# Ingest Data with dltHub
Source: https://docs.hopsworks.ai/latest/user_guides/fs/feature_group/ingest_with_dlthub/
# How to ingest data into a Feature Group with dltHub { #ingest-data-with-dlthub }
## Introduction
Hopsworks can copy data from an existing data source into a new managed feature group using dltHub.
This workflow creates:
- A new feature group in Hopsworks.
- An ingestion job that copies data from the selected source into that feature group.
This is different from creating an external feature group.
An external feature group keeps the data in the source system, while the dltHub ingestion flow copies the data into Hopsworks.
!!! note
You can configure this workflow both in the Hopsworks UI and with the Hopsworks Python APIs.
## When to use this workflow
Use `Ingest Data to New Feature Group` when you want to:
- Copy data from source into Hopsworks.
- Schedule recurring ingestion jobs.
- Use incremental loading for supported source types.
## Supported source types
This ingestion flow supports multiple data sources:
- SQL-like sources can either create an external feature group or ingest data into a new feature group.
- The SQL family currently includes Snowflake, BigQuery, Redshift, generic JDBC (MySQL, PostgreSQL, Oracle), and SAP HANA.
- CRM and REST API sources use the ingestion path only.
- Incremental loading is available for SQL and REST API sources.
- CRM sources currently use full-load ingestion.
## Step 1: Open the Data Source and start Feature Group creation
Navigate to the data source you want to use and start the feature-group creation flow from the UI.
For SQL-based sources, open the data source, click `Next: Select Tables`, select a database and a table, then choose `Ingest Data to New Feature Group`.
Once the ingest option is selected, the column table gains a `Partition key` column and an `Add a feature` button for features that do not exist in the source; the transformation script computes their values.

Select a source table, set the keys and choose Ingest Data to New Feature Group
For CRM sources, choose the source resource, click `Fetch Schema` and then configure the feature schema for the new feature group the same way.
For REST API sources, first configure the endpoint before fetching the schema.
### REST endpoint pagination
REST sources require endpoint configuration up front so Hopsworks can fetch the schema correctly.
In this step, define:
- **Resource**: Any unique identifier for the endpoint.
- **Relative URL**: The endpoint path relative to the configured REST data source base URL.
- **Request Parameters**: Optional query parameters sent with the request.
- **Pagination Configuration**: The pagination mode and its parameters, if the API returns paged results.
The REST pagination form supports these modes:
- `NONE`
- `HEADER_CURSOR`
- `HEADER_LINK`
- `JSON_CURSOR`
- `JSON_LINK`
- `OFFSET`
- `PAGE_NUMBER`
- `SINGLE_PAGE`
For example, `PAGE_NUMBER` pagination exposes:
- **Page Parameter Name**: Name of the request parameter that contains the page number.
- **Base Page**: Starting page number used by the API, for example `0` or `1`.
- **Total Pages Path**: Response path containing the total number of pages.

REST API pagination configuration using PAGE_NUMBER
Other pagination modes expose their own source-specific fields in the form:
- `OFFSET`: offset parameter name, limit parameter name, limit value, total-items path, and has-more path.
- `JSON_CURSOR`: cursor parameter name and cursor path.
- `HEADER_CURSOR`: cursor header key and cursor path.
- `HEADER_LINK`: next-link header key.
- `JSON_LINK`: next URL path.
For more details on how these pagination strategies work in dltHub, see the [dltHub REST API pagination documentation](https://dlthub.com/docs/dlt-ecosystem/verified-sources/rest_api/basic#pagination).
## Step 2: Configure the feature group schema
After fetching metadata from the source, Hopsworks shows the feature selection table.
At this stage you can:
- Set the **Feature Group Name**.
- Include or exclude columns.
- Edit feature names and data types.
- Mark one or more features as **Primary key**.
- Optionally select a **Partition key**.
- Optionally select an **Event time** column.
- Preview metadata and preview data before continuing.
When you are ready, click `Next: Configure Ingestion Job`.
!!! note
`Create External Feature Group` is not supported for CRM and REST connectors.
!!! note
For CRM and REST sources, schema fetching reads only a small sample of records from the source.
Hopsworks uses this sample to infer the feature-group schema before you create the ingestion job.
## Step 3: Configure the dltHub ingestion job
The next page configures the ingestion job that will populate the feature group.

Configure the dltHub ingestion job: job settings, transformation, resources and loading strategy
### Common job settings
The following fields are available in the job configuration:
- **Job Name**: Name of the ingestion job created in Hopsworks.
- **Source Read Parallelism**: Number of parallel readers that pull data from the source database or API. Increase it to speed up ingestion if the source can handle the extra load.
- **Data processing parallelism**: Number of parallel processes that prepare and transform data before loading it into the feature group. Increase it if processing is slow and CPU is available.
- **Destination Write Batch Size**: Number of records written to the feature group in each batch during ingestion.
- **Max Write Batch Size (MB)**: Maximum file size, in megabytes, when writing data to the feature group.
- **Write Mode**: Controls whether incoming data is appended as-is or merged with existing rows using the primary key.
- **Environment**: Python environment the job runs in, `dlthub-ingestion-pipeline` by default.
- **Start the job after creation**: Starts the ingestion job immediately after the resources are created.
- **Data Transformation**: Optional Python script, picked from the project or uploaded, that transforms rows before they are written and computes any extra features added to the schema.
- **Memory (in MB)** and **CPU Cores**: Resources allocated to the ingestion job; `Estimate resources` proposes values from the source size.
- **Schedule**: Optional recurring schedule for future ingestion runs.
- **Alerts**: Optional alerting configuration for the ingestion job.
### SQL-only settings
For SQL sources, the job configuration also includes:
- **Source Read Batch Size**: Number of records fetched per read from the SQL source.
- **Source Table Partitions**: Number of partitions used when reading from SQL sources. For very large tables, increase this value to split the read into smaller chunks that fit the allocated memory.
These options control how data is read from the source table during ingestion.
### Write modes
Two write modes are available:
- **APPEND**: Appends new data without merging with existing rows. This greatly speeds up writes and uses less memory, but can result in duplicate rows. If you are ingesting a large amount of data, this is the recommended mode and duplicates can be handled later in a separate pipeline step.
- **MERGE**: Merges incoming data with existing rows using the feature-group primary key. This avoids duplicate rows, but slows down ingestion and requires more memory, especially for large ingestions. Use it when ingesting smaller amounts of data.
## Step 4: Choose a loading strategy
The `Loading Strategy` section controls whether the pipeline reads the entire source or only new data.
The following strategies are available in the UI:
- `FULL_LOAD`
- `INCREMENTAL_ID`
- `INCREMENTAL_TIMESTAMP`
- `INCREMENTAL_DATE`
--8<-- "user_guides/fs/feature_group/ingest_with_dlthub/loading-strategies.html"
### Full load
`FULL_LOAD` is available for all sources in this workflow.
With a full load, the ingestion job reads the complete dataset from the source and writes it again to the destination feature group.
In practice, this means the target feature group is refreshed from scratch for the same feature-group name and version.
Any data already stored in that feature group version is removed and replaced by the newly ingested data from the source.
Use `FULL_LOAD` when you want the feature group to be a complete copy of the source at the time of ingestion, rather than an incremental continuation of previous runs.
This is useful when:
- The source does not provide a reliable incremental cursor.
- You want to rebuild the feature group from a clean state.
- The source data can change retroactively and you want to re-sync the full table or endpoint.
Because a full load rewrites the destination dataset, it is typically more expensive than incremental ingestion for large sources.
For recurring pipelines, prefer an incremental strategy when the source supports it and when you only need newly added or updated records.
For SQL sources, you can also optionally define:
- **Source Cursor Field**: A field used to efficiently synchronize only new or changed data from the source into the destination feature group.
- **Initial Value**: Starting value for the selected source cursor field.
This can be used to split or optimize the load when the source table has a monotonic column, even though the ingestion mode remains a full refresh of the feature group.
### Incremental loading
Incremental loading is available for SQL and REST API sources.
With incremental loading, the ingestion job does not re-copy the full source on every run.
Instead, it keeps track of a cursor value and only fetches records that are newer than, or come after, the last processed value.
This makes incremental loading the preferred option for recurring ingestion jobs when the source exposes a stable field that can be used to identify new or updated data.
Typical cursor fields are:
- Increasing numeric identifiers.
- Update timestamps.
- Event dates.
Compared to `FULL_LOAD`, incremental loading typically:
- Reduces the amount of data read from the source.
- Shortens ingestion time.
- Lowers resource usage.
- Avoids rebuilding the destination feature group from scratch on every run.
To work reliably, the selected cursor field should be monotonic or consistently ordered for the records you want to ingest.
If the source does not provide such a field, `FULL_LOAD` is usually the safer option.
The common incremental field is:
- **Source Cursor Field**: A field used to efficiently synchronize only new or changed data from the source into the destination feature group.
Depending on the strategy, you must also define:
- **INCREMENTAL_ID**: **Initial Value**, the numeric starting value for incremental reads.
- **INCREMENTAL_TIMESTAMP**: **Initial Value**, the starting Unix timestamp for incremental reads.
- **INCREMENTAL_DATE**: **Initial Date**, the starting date and time for incremental reads.
The initial value defines where the first run starts.
After that, subsequent runs continue from the last successfully processed cursor value.
For REST API sources, incremental loading also requires:
- **REST Filter Param**: The actual API parameter used to request only new data since the last run, for example `start_date`, `updated_at`, or `since`.
Choose the incremental strategy that matches the source cursor type:
- `INCREMENTAL_ID` for sources with increasing numeric identifiers.
- `INCREMENTAL_TIMESTAMP` for sources that expose Unix timestamps.
- `INCREMENTAL_DATE` for sources that filter by date or datetime values.

Incremental loading by id, with tid as the source cursor field
## Step 5: Review and create
After configuring the ingestion job, click `Next: Review Configuration`.
The review dialog shows:
- The source schema, table, connector, or resource.
- The final feature group name.
- Whether sink ingestion is enabled.
- The ingestion job name.
- The number of selected features.
You can still edit the feature-group name and ingestion-job name in this step before creating the resources.

Review the feature group and ingestion job before creation
Click `Create` to create the feature group and the dltHub ingestion job.
## Result
After creation:
- The feature group is registered in Hopsworks.
- The ingestion job is available under project jobs.
- If `Start the job after creation` is enabled, the initial ingestion starts immediately.
- If a schedule is configured, future synchronizations will run automatically.
## Next Steps
- Use the [Feature Group creation guide][create-feature-group] to understand managed feature groups in more detail.
- Use the [External Feature Group guide][create-external-feature-group] if you want to query the source in place without copying data into Hopsworks.
- Use the [Online Ingestion Observability guide][online-ingestion-observability] to monitor ingestion behavior for online-enabled feature groups.
## API support
You can also configure data source ingestion programmatically with the Hopsworks Python APIs.
This is done by creating a sink-enabled feature group and passing a sink job configuration, including loading strategy and, for REST sources, endpoint and pagination settings.
### Example: create a sink-enabled feature group
```python
from hopsworks_common.core import sink_job_configuration
fs = project.get_feature_store()
data_source = fs.get_data_source("my_sql_source").get_tables()[0]
data = data_source.get_data(use_cached=False)
sink_job_conf = sink_job_configuration.SinkJobConfiguration(
name="sql_to_fg_ingestion",
write_mode=sink_job_configuration.WriteMode.APPEND,
)
fg = fs.get_or_create_feature_group(
name="transactions_fg",
version=1,
description="Managed feature group populated from a data source.",
primary_key=[data.features[0]["name"]],
features=data.features,
data_source=data_source,
time_travel_format="DELTA",
sink_enabled=True,
sink_job_conf=sink_job_conf,
)
fg.save()
# Run the ingestion job
fg.sink_job.run(await_termination=True)
```
### Example: REST ingestion with incremental loading
```python
from hopsworks_common.core import rest_endpoint, sink_job_configuration
from hsfs.core import data_source as ds
fs = project.get_feature_store()
parent_data_source = fs.get_data_source("my_rest_source")
endpoint_config = rest_endpoint.RestEndpointConfig(
relative_url="/transactions",
query_params={"page_size": 100},
pagination_config=rest_endpoint.PageNumberPaginationConfig(
base_page=1,
page_param="page",
total_path="total",
stop_after_empty_page=True,
),
)
rest_data_source = ds.DataSource(
table="transactions_rest",
rest_endpoint=endpoint_config,
storage_connector=parent_data_source.storage_connector,
)
rest_data = rest_data_source.get_data(use_cached=False)
loading_config = sink_job_configuration.LoadingConfig(
loading_strategy=sink_job_configuration.LoadingStrategy.INCREMENTAL_DATE,
source_cursor_field="timestamp",
initial_value="2024-01-01T00:00:00Z",
rest_filter_param="start_time",
)
sink_job_conf = sink_job_configuration.SinkJobConfiguration(
name="rest_to_fg_ingestion",
loading_config=loading_config,
)
fg = fs.get_or_create_feature_group(
name="transactions_rest_fg",
version=1,
description="Managed feature group populated from a REST source.",
primary_key=["id"],
features=rest_data.features,
data_source=rest_data_source,
time_travel_format="DELTA",
sink_enabled=True,
sink_job_conf=sink_job_conf,
)
fg.save()
```
================================================================================
# Create Spine
Source: https://docs.hopsworks.ai/latest/user_guides/fs/feature_group/create_spine/
# How to create Spine Group
## Introduction
In this guide you will learn how to create and register a Spine Group with Hopsworks.
## Prerequisites
Before you begin this guide we suggest you read the [Spine Group](../../../concepts/fs/feature_group/spine_group.md) concept page to understand what a Spine Group is and how it fits in the ML pipeline.
## Create using the Hopsworks APIs
### Create a Spine Group
Instead of using a feature group to save the label, you can also use a spine to use a Dataframe containing the labels on the fly.
A spine is essentially a metadata object similar to a Feature Group, which tells the feature store the relevant event time column and primary key columns to perform point-in-time correct joins.
Additionally, apart from primary key and event time information, a Spark dataframe is required in order to infer the schema of the group from.
=== "Python"
```python
trans_spine = fs.get_or_create_spine_group(
name="spine_transactions",
version=1,
description="Transaction data",
primary_key=["cc_num"],
event_time="datetime",
dataframe=trans_df,
)
```
Once created, note that you can inspect the dataframe in the Spine Group:
=== "Python"
```python
trans_spine.dataframe.show()
```
And you can always also replace the dataframe contained within the Spine Group.
You just need to make sure it has the same schema.
=== "Python"
```python
trans_spine.dataframe = new_df
```
### Limitations
!!! warning "Python support"
Currently the Hopsworks library does not support usage of Spine Groups for training data creation or batch data retrieval in the Python engine.
However, it is supported to create Spine Groups from the Python engine.
!!! api "API reference"
- [`FeatureStore.get_or_create_spine_group`][hsfs.feature_store.FeatureStore.get_or_create_spine_group]
- [`SpineGroup`][hsfs.feature_group.SpineGroup]
Browse the full Python API :material-arrow-right:
================================================================================
# Deprecate
Source: https://docs.hopsworks.ai/latest/user_guides/fs/feature_group/deprecation/
# How to deprecate a Feature Group
## Introduction
To discourage the usage of specific feature groups it is possible to deprecate them.
When a feature group is deprecated, user will be warned when they try to use it or use a feature view that depends on it.
In this guide you will learn how to deprecate a feature group within Hopsworks, showing examples in Hopsworks APIs as well as the user interface.
## Prerequisites
Before you begin this guide it is expected that there is an existing feature group in your project.
You can familiarize yourself with [the creation of a feature group](./create.md) in the user guide.
## Deprecate using the Hopsworks APIs
### Retrieve the feature group
To deprecate a feature group using the Hopsworks APIs you need to provide a [Feature Group](../../../concepts/fs/feature_group/fg_overview.md).
=== "Python"
```python
fg = fs.get_feature_group(
name="feature_group_name", version=feature_group_version
)
```
### Deprecate Feature Group
Feature group deprecation occurs by calling the `update_deprecated` method on the feature group.
=== "Python"
```python
fg.update_deprecated()
```
Users can also un-deprecate the feature group if need be, by setting the `deprecate` parameter to False.
=== "Python"
```python
fg.update_deprecated(deprecate=False)
```
## Deprecate using the UI
You can deprecate/de-deprecate feature groups through the UI.
For this, navigate to the `Feature Groups` section and select a feature group.
Subsequently, make sure that the necessary feature group version is picked.
Finally, click on the button with three vertical dots in the right corner and select `Deprecate`.
The Feature group can be de-deprecated by selecting the `Undeprecate` option on a deprecated feature group.
================================================================================
# Data Types and Schema management
Source: https://docs.hopsworks.ai/latest/user_guides/fs/feature_group/data_types/
# How to manage schema and feature data types
## Introduction
In this guide, you will learn how to manage the feature group schema and control the data type of the features in a feature group.
## Prerequisites
Before you begin this guide we suggest you read the [Feature Group](../../../concepts/fs/feature_group/fg_overview.md) concept page to understand what a feature group is and how it fits in the ML pipeline.
We also suggest you familiarize yourself with the APIs to [create a feature group](./create.md).
## Feature group schema
When a feature is stored in both the online and offline feature stores, it will be stored in a data type native to each store.
- **[Offline data type](#offline-data-types)**: The data type of the feature when stored on the offline feature store.
The offline feature store is based on Apache Hudi and Hive Metastore, as such, [Hive Data Types](https://cwiki.apache.org/confluence/display/Hive/LanguageManual+Types) can be leveraged.
- **[Online data type](#online-data-types)**: The data type of the feature when stored on the online feature store.
The online storage is based on RonDB and hence, [MySQL Data Types](https://dev.mysql.com/doc/refman/8.0/en/data-types.html) can be leveraged.
The offline data type is always required, even if the feature group is stored only online.
On the other hand, if the feature group is not *online_enabled*, its features will not have an online data type.
The offline and online types for each feature are automatically inferred from the Spark or Pandas types of the input DataFrame as outlined in the following two sections.
The default mapping, however, can be overwritten by using an [explicit schema definition](#explicit-schema-definition).
### Offline data types
When registering a [Spark](https://spark.apache.org/docs/latest/sql-ref-datatypes.html) DataFrame in a PySpark environment (S),
or a [Pandas](https://pandas.pydata.org/) DataFrame, or a [Polars](https://pola.rs/) DataFrame in a Python-only environment (P) the following default mapping to offline feature types applies:
| Spark Type (S) | Pandas Type (P) |Polars Type (P) | Offline Feature Type | Remarks |
|----------------|------------------------------------|-----------------------------------|-------------------------------|----------------------------------------------------------------|
| BooleanType | bool, object(bool) |Boolean | BOOLEAN | |
| ByteType | int8, Int8 |Int8 | TINYINT or INT | INT when time_travel_type="HUDI" |
| ShortType | uint8, int16, Int16 |UInt8, Int16 | SMALLINT or INT | INT when time_travel_type="HUDI" |
| IntegerType | uint16, int32, Int32 |UInt16, Int32 | INT | |
| LongType | int, uint32, int64, Int64 |UInt32, Int64 | BIGINT | |
| FloatType | float, float16, float32 |Float32 | FLOAT | |
| DoubleType | float64 |Float64 | DOUBLE | |
| DecimalType | decimal.decimal |Decimal | DECIMAL(PREC, SCALE) | Not supported in PO env. when time_travel_type="HUDI" |
| TimestampType | datetime64[ns], datetime64[ns, tz] |Datetime | TIMESTAMP | s. [Timestamps and Timezones](#timestamps-and-timezones) |
| DateType | object (datetime.date) |Date | DATE | |
| StringType | object (str), object(np.unicode) |String, Utf8 | STRING | |
| ArrayType | object (list), object (np.ndarray) |List | ARRAY<TYPE> | |
| StructType | object (dict) |Struct | STRUCT<NAME: TYPE, ...> | |
| BinaryType | object (binary) |Binary | BINARY | |
| MapType | - |- | MAP<String,TYPE> | Only when time_travel_type!="HUDI"; Only string keys permitted |
When registering a Pandas DataFrame in a PySpark environment (S) the Pandas DataFrame is first converted to a Spark DataFrame, using Spark's [default conversion](https://spark.apache.org/docs/3.1.1/api/python/reference/api/pyspark.sql.SparkSession.createDataFrame.html).
It results in a less fine-grained mapping between Python and Spark types:
| Pandas Type (S) | Spark Type | Remarks |
|-------------------------------------------------------|---------------|----------------------------------------------------------|
| bool | BooleanType | |
| int8, uint8, int16, uint16, int32, int, uint32, int64 | LongType | |
| float, float16, float32, float64 | DoubleType | |
| object (decimal.decimal) | DecimalType | |
| datetime64[ns], datetime64[ns, tz] | TimestampType | s. [Timestamps and Timezones](#timestamps-and-timezones) |
| object (datetime.date) | DateType | |
| object (str), object(np.unicode) | StringType | |
| object (list), object (np.ndarray) | - | Not supported |
| object (dict) | StructType | |
| object (binary) | BinaryType | |
### Online data types
The online data type is determined based on the offline type according to the following mapping, regardless of which environment the data originated from.
Only a subset of the data types can be used as primary key, as indicated in the table as well:
| Offline Feature Type | Online Feature Type | Primary Key | Remarks |
|-------------------------------|----------------------|-------------|----------------------------------------------------------|
| BOOLEAN | TINYINT | x | |
| TINYINT | TINYINT | x | |
| SMALLINT | SMALLINT | x | |
| INT | INT | x | Also supports: TINYINT, SMALLINT |
| BIGINT | BIGINT | x | |
| FLOAT | FLOAT | | |
| DOUBLE | DOUBLE | | |
| DECIMAL(PREC, SCALE) | DECIMAL(PREC, SCALE) | | e.g. DECIMAL(38, 18) |
| TIMESTAMP | TIMESTAMP | | s. [Timestamps and Timezones](#timestamps-and-timezones) |
| DATE | DATE | x | |
| STRING | VARCHAR(100) | x | Also supports: TEXT |
| ARRAY<TYPE> | VARBINARY(100) | x | Also supports: BLOB |
| STRUCT<NAME: TYPE, ...> | VARBINARY(100) | x | Also supports: BLOB |
| BINARY | VARBINARY(100) | x | Also supports: BLOB |
| MAP<String,TYPE> | VARBINARY(100) | x | Also supports: BLOB |
More on how Hopsworks handles [string types](#string-online-data-types), [complex data types](#complex-online-data-types) and the online restrictions for [primary keys](#online-restrictions-for-primary-key-data-types) and [row size](#online-restrictions-for-row-size) in the following sections.
#### String online data types
String types are stored as *VARCHAR(100)* by default.
This type is fixed-size, meaning it can only hold as many characters as specified in the argument (e.g., VARCHAR(100) can hold up to 100 unicode characters).
The size should thus be within the maximum string length of the input data.
Furthermore, the VARCHAR size has to be in line with the [online restrictions for row size](#online-restrictions-for-row-size).
If the string size exceeds 100 characters, a larger type (e.g., VARCHAR(500)) can be specified via an [explicit schema definition](#explicit-schema-definition).
If the string size is unknown or if it exceeds the maximum row size, then the [TEXT type](https://docs.rondb.com/blobs/) can be used instead.
String data that exceeds the specified VARCHAR size will lead to an error when data gets written to the online feature store.
When in doubt, use the TEXT type instead, but note that it comes with a potential performance overhead.
#### Complex online data types
Hopsworks allows users to store complex types (e.g. *ARRAY*) in the online feature store.
Hopsworks serializes the complex features transparently and stores them as VARBINARY in the online feature store.
The serialization happens when calling the [`FeatureGroup.save`][hsfs.feature_group.FeatureGroup.save],
[`FeatureGroup.insert`][hsfs.feature_group.FeatureGroup.insert] or [`FeatureGroup.insert_stream`][hsfs.feature_group.FeatureGroup.insert_stream] methods.
The deserialization will be executed when calling the [`TrainingDataset.get_serving_vector`][hsfs.training_dataset.TrainingDataset.get_serving_vector] method to retrieve data from the online feature store.
If users query directly the online feature store, for instance using the `fs.sql("SELECT ...", online=True)` statement, it will return a binary blob.
On the feature store UI, the online feature type for complex features will be reported as *VARBINARY*.
If the binary size exceeds 100 bytes, a larger type (e.g., VARBINARY(500)) can be specified via an [explicit schema definition](#explicit-schema-definition).
If the binary size is unknown of if it exceeds the maximum row size, then the [BLOB type](https://docs.rondb.com/blobs/) can be used instead.
Binary data that exceeds the specified VARBINARY size will lead to an error when data gets written to the online feature store.
When in doubt, use the BLOB type instead, but note that it comes with a potential performance overhead.
#### Online restrictions for primary key data types
When a feature is being used as a primary key, certain types are not allowed.
Examples of such types are *FLOAT*, *DOUBLE*, *TEXT* and *BLOB*.
Additionally, the size of the sum of the primary key online data types storage requirements **should not exceed 4KB**.
#### Online restrictions for row size
The online feature store supports **up to 500 columns** and all column types combined **should not exceed 30000 Bytes**.
The byte size of each column is determined by its data type and calculated as follows:
| Online Data Type | Byte Size |
|---------------------------------|--------------|
| TINYINT | 1 |
| SMALLINT | 2 |
| INT | 4 |
| BIGINT | 8 |
| FLOAT | 4 |
| DOUBLE | 8 |
| DECIMAL(PREC, SCALE) | 16 |
| TIMESTAMP | 8 |
| DATE | 8 |
| VARCHAR(LENGTH) | LENGTH * 4 |
| VARCHAR(LENGTH) charset latin1; | LENGTH * 1 |
| TEXT | 256 |
| VARBINARY(LENGTH) | LENGTH |
| BLOB | 256 |
| other | 8 |
!!! note "VARCHAR / VARBINARY overhead"
For VARCHAR and VARBINARY data types, an additional 1 byte is required if the size is less than 256 bytes.
If the size is 256 bytes or greater, 2 additional bytes are required.
Memory allocation is performed in groups of 4 bytes.
For example, a VARBINARY(100) requires 104 bytes of memory:
- 100 bytes for the data itself
- 1 byte of overhead
- Total = 101 bytes
Since memory is allocated in 4-byte groups, storing 101 bytes requires 26 groups (26 × 4 = 104 bytes) of allocated memory.
#### Pre-insert schema validation for online feature groups
For online enabled feature groups, the dataframe to be ingested needs to adhere to the online schema definitions.
The input dataframe is validated for schema checks accordingly.
The validation is enabled by default and can be disabled by setting below key word argument when calling `insert()`
=== "Python"
```python
feature_group.insert(
df, validation_options={"online_schema_validation": False}
)
```
The most important validation checks or error messages are mentioned below along with possible corrective actions.
1. Primary key contains null values
- **Rule** Primary key column should not contain any null values.
- **Example correction** Drop the rows containing null primary keys.
Alternatively, find the null values and assign them an unique value as per preferred strategy for data imputation.
```python
# Drop rows: assuming 'id' is the primary key column
df = df.dropna(subset=["id"])
# For composite keys
df = df.dropna(subset=["id1", "id2"])
# Data imputation: replace null values with incrementing last integer id
# existing max id
max_id = df["id"].max()
# counter to generate new id
next_id = max_id + 1
# for each null id, assign the next id incrementally
for idx in df[df["id"].isna()].index:
df.loc[idx, "id"] = next_id
next_id += 1
```
2. Primary key column missing
- **Rule** The dataframe to be inserted must contain all the columns defined as primary key(s) in the feature group.
- **Example correction** Add all the primary key columns in the dataframe.
```python
# incrementing primary key upto the length of dataframe
df["id"] = range(1, len(df) + 1)
```
3. String length exceeded
- **Rule** The character length of a string should be within the maximum length capacity in the online schema type of a feature.
If the feature group is not created and explicit feature schema was not provided, the limit will be auto-increased to the maximum length found in a string column in the dataframe.
- **Example correction**
- Trim the string values to fit within maximum limit set during feature group creation.
```python
max_length = 100
df["text_column"] = df["text_column"].str.slice(0, max_length)
```
- Another option is to simply [create new version of the feature group][hsfs.feature_store.FeatureStore.get_or_create_feature_group] and insert the dataframe.
!!! note
The total row size limit should be less than 30kb as per [row size restrictions](#online-restrictions-for-row-size).
In such cases it is possible to define the feature as **TEXT** or **BLOB**.
Below is an example of explicitly defining the string column as TEXT as online type.
```python
import pandas as pd
# example dummy dataframe with the string column
df = pd.DataFrame(columns=["id", "string_col"])
from hsfs.feature import Feature
features = [
Feature(name="id", type="bigint", online_type="bigint"),
Feature(name="string_col", type="string", online_type="text"),
]
fg = fs.get_or_create_feature_group(
name="fg_manual_text_schema",
version=1,
features=features,
online_enabled=True,
primary_key=["id"],
)
fg.insert(df)
```
### Timestamps and Timezones
All timestamp features are stored in Hopsworks in UTC time.
Also, all timestamp-based functions (such as [point-in-time joins](../../../concepts/fs/feature_view/offline_api.md#point-in-time-correct-training-data)) use UTC time.
This ensures consistency of timestamp features across different client timezones and simplifies working with timestamp-based functions in general.
When ingesting timestamp features, the [`FeatureGroup.insert`][hsfs.feature_group.FeatureGroup.insert] will automatically handle the conversion to UTC, if necessary.
The following table summarizes how different timestamp types are handled:
| Data Frame (Data Type) | Environment | Handling |
| --- | --- | --- |
| Pandas DataFrame (datetime64[ns]) | Python-only and PySpark | interpreted as UTC, independent of the client's timezone |
| Pandas DataFrame (datetime64[ns, tz]) | Python-only and PySpark | timezone-sensitive conversion from 'tz' to UTC |
| Spark (TimestampType) | PySpark and Spark | interpreted as UTC, independent of the client's timezone |
Timestamp features retrieved from the Feature Store, e.g., using the [Feature Store Read API][hsfs.feature_group.FeatureGroup.read], use a timezone-unaware format:
| Data Frame (Data Type) | Environment | Timezone |
|---------------------------------------|-------------------------|------------------------|
| Pandas DataFrame (datetime64[ns]) | Python-only | timezone-unaware (UTC) |
| Spark (TimestampType) | PySpark and Spark | timezone-unaware (UTC) |
Note that our PySpark/Spark client automatically sets the Spark SQL session's timezone to UTC.
This ensures that Spark SQL will correctly interpret all timestamps as UTC.
The setting will only apply to the client's session, and you don't have to worry about setting/unsetting the configuration yourself.
## Explicit schema definition
When creating a feature group it is possible for the user to control both the offline and online data type of each column.
If users explicitly define the schema for the feature group, Hopsworks is going to use that schema to create the feature group, without performing any type mapping.
You can explicitly define the feature group schema as follows:
=== "Python"
```python
from hsfs.feature import Feature
features = [
Feature(name="id", type="int", online_type="int"),
Feature(name="name", type="string", online_type="varchar(20)"),
]
fg = fs.create_feature_group(
name="fg_manual_schema", features=features, online_enabled=True
)
fg.save(features)
```
## Append features to existing feature groups
Hopsworks supports appending additional features to an existing feature group.
Adding additional features to an existing feature group is not considered a breaking change.
=== "Python"
```python
from hsfs.feature import Feature
features = [
Feature(name="id", type="int", online_type="int"),
Feature(name="name", type="string", online_type="varchar(20)"),
]
fg = fs.get_feature_group(name="example", version=1)
fg.append_features(features)
```
When adding additional features to a feature group, you can provide a default values for existing entries in the feature group.
You can also backfill the new features for existing entries by running an `insert()` operation and update all existing combinations of *primary key* - *event time*.
### Appending features to an external feature group
For an [external feature group](create_external.md), appending a feature only updates the Hopsworks-side metadata; it does not add a column to the external table itself.
Reading online (`read(online=True)`) is unaffected, since Hopsworks owns the online table schema and can add the column there directly.
Reading offline (`read()`) queries the external source directly, so it fails with a column-not-found error until the external table itself gains a matching column.
Until then, exclude the appended feature from the offline read with `select_except`:
=== "Python"
```python
fg = fs.get_feature_group(name="example", version=1)
# "name" was appended but is not yet a column in the external table
df = fg.select_except(["name"]).read()
```
Update the external table's schema, or keep excluding the appended feature, then read offline again once the column is present.
================================================================================
# Statistics
Source: https://docs.hopsworks.ai/latest/user_guides/fs/feature_group/statistics/
# How to compute statistics on feature data
## Introduction
In this guide you will learn how to configure, compute and visualize statistics for the features registered with Hopsworks.
Hopsworks groups statistics in four categories:
- **Descriptive**: These are the basic statistics Hopsworks computes.
They include an _approximate_ count of the distinctive values and the completeness (i.e., the percentage of non null values).
For numerical features Hopsworks also computes the minimum, maximum, mean, standard deviation and the sum of each feature.
Enabled by default.
- **Histograms**: Hopsworks computes the distribution of the values of a feature.
Exact histograms are computed as long as the number of distinct values is less than 20. If a feature has a numerical data type (e.g., integer, float, double, ...) and has more than 20 unique values, then the values are bucketed in 20 buckets and the histogram represents the distribution of values in those buckets.
By default histograms are disabled.
- **Correlation**: If enabled, Hopsworks computes the Pearson correlation between features of numerical data type within a feature group.
By default correlation is disabled.
- **Exact Statistics**: Exact statistics are an enhancement of the descriptive statistics that provide an exact count of distinctive values, entropy, uniqueness and distinctiveness of the value of a feature.
These statistics are more expensive to compute as they take into consideration all the values and they don't use approximations.
By default they are disabled.
When statistics are enabled, they are computed every time new data is written into the _offline_ storage of a feature group.
Statistics are then displayed on the Hopsworks UI and users can track how data has changed over time.
## Prerequisites
Before you begin this guide we suggest you read the [Feature Group](../../../concepts/fs/feature_group/fg_overview.md) concept page to understand what a feature group is and how it fits in the ML pipeline.
We also suggest you familiarize with the APIs to [create a feature group](./create.md).
## Enable statistics when creating a feature group
As mentioned above, by default only descriptive statistics are enabled when creating a feature group.
To enable histograms, correlations or exact statistics the `statistics_config` configuration parameter can be provided in the create statement.
The `statistics_config` parameter takes a dictionary with the keys: `enabled`, `correlations`, `histograms` and `exact_uniqueness` and, as values, a boolean to describe whether or not to compute the specific class of statistics.
Additionally it is possible to restrict the statistics computation to only a subset of columns.
This is configurable by adding a `columns` key to the `statistics_config` parameter.
The key should contain the list of columns for which to compute statistics.
By default the value is empty list `[]` and the statistics are computed for all columns in the feature group.
=== "Python"
```python
fg = feature_store.create_feature_group(
name="weather",
version=1,
description="Weather Features",
online_enabled=True,
primary_key=["location_id"],
partition_key=["day"],
event_time="event_time",
statistics_config={
"enabled": True,
"histograms": True,
"correlations": True,
"exact_uniqueness": False,
"columns": [],
},
)
```
## Enable statistics after creating a feature group
You can change the statistics configuration after a feature group was created, to add or remove a class of statistics or to change the set of features for which to compute them.
In the UI, open the feature group in the `Catalog` and click the edit icon; the statistics configuration sits at the top of the edit page.
Statistics configuration on the Edit Feature Group page.
=== "Python"
```python
fg.statistics_config = {
"enabled": True,
"histograms": False,
"correlations": False,
"exact_uniqueness": False,
"columns": ["location_id", "min_temp", "max_temp"],
}
fg.update_statistics_config()
```
## Explicitly compute statistics
As mentioned above, the statistics are computed every time new data is written into the _offline_ storage of a feature group.
By invoking the `compute_statistics` method, users can trigger explicitly the statistics computation for the data available in a feature group.
This is useful when a feature group is receiving frequent updates.
Users can schedule periodic statistics computation that take into consideration several data commits.
By default, the `compute_statistics` method computes statistics on the most recent version of the data available in a feature group.
Users can provide a specific time using the `wallclock_time` parameter, to compute the statistics for a previous version of the data.
=== "Python"
```python
fg.compute_statistics(wallclock_time="20220611 20:00")
```
### External feature groups
External feature groups own the same built-in `ingestion_stats` configuration as cached and stream feature groups, but no data is ingested into Hopsworks for them, so it never runs on its own.
Calling `compute_statistics` on the external feature group, or clicking "Compute statistics" in the UI, runs it: the statistics job reads the external source and profiles it like an internal feature group.
Saving an external feature group with statistics enabled runs it once as well.
External feature groups have no commit history, so their statistics carry the computation time only and `compute_statistics` takes no time argument.
=== "Python"
```python
external_fg.compute_statistics()
```
## Inspect statistics
Open the feature group in the `Catalog` and select `Feature Statistics` in its sidebar.
The page lists every feature with its count, completeness, min, max, mean and standard deviation, plus a histogram per numerical feature, for the latest commit or any earlier one you pick.
The same numbers are available from the API with `fg.get_statistics()`.
Feature Statistics for a feature group, one card per feature.
================================================================================
# Getting started
Source: https://docs.hopsworks.ai/latest/user_guides/fs/feature_group/data_validation/
# Data Validation
--8<-- "user_guides/fs/feature_group/data_validation/validation-on-insert.html"
## Introduction
Clean, high quality feature data is of paramount importance to being able to train and serve high quality models.
Hopsworks offers integration with [Great Expectations](https://greatexpectations.io/) to enable a smooth data validation workflow.
This guide is designed to help you integrate a data validation step when inserting new DataFrames into a Feature Group.
Note that validation is performed inline as part of your feature pipeline (on the client machine) - it is not executed by Hopsworks after writing features.
## UI
### Create a Feature Group (Pre-requisite)
In the UI, you must create a Feature Group first before attaching an Expectation Suite.
You can find out more information about [creating a Feature Group](create.md).
You can attach at most one expectation suite to a Feature Group.
Data validation is an optional step and is not required to write to a Feature Group.
### Step 1: Find and Edit Feature Group
Click on the Feature Group section in the navigation menu.
Find your Feature Group in the list and click on its name to access the Feature Group page.
Select `edit` in the top right corner or scroll to the Expectations section and click on `Edit Expectation Suite`.
### Step 2: Edit General Expectation Suite Settings
Scroll to the Expectation Suite section.
Click add Expectation Suite and edit its metadata:
- Choose a name for your expectation suite.
- Checkbox enabled.
This controls whether the Expectation Suite will be used to validate a Dataframe automatically upon insertion into a Feature Group.
Note that validation is executed by the client.
Disabling validation allows you to skip the validation step without deleting the Expectation Suite.
- 'ALWAYS' vs. 'STRICT' mode.
This option controls what happens after validation.
Hopsworks defaults to 'ALWAYS', where data is written to the Feature Group regardless of the validation result.
This means that even if expectations are failing or throw an exception, Hopsworks will attempt to insert the data into the Feature Group.
In 'STRICT' mode, Hopsworks will only write data to the Feature Group if each individual expectation has been successful.
### Step 3: Add new expectations
By clicking on `Add expectation` one can choose an expectation type from a searchable dropdown menu.
Currently, only the built-in expectations from the Great Expectations framework are supported.
For user-defined expectations, please use the Rest API or python client.
All default kwargs associated to the selected expectation type are populated as a json below the dropdown menu.
Edit the arguments in the json to configure the Expectation.
In particular, arguments such as `column`, `columnA`, `columnB`, `column_set` and `column_list` require valid feature name(s).
Click the tick button to save the expectation configuration and append it to the Expectation Suite locally.
!!! info
Click the `Save feature group` button to persist your changes!
You can use the button `Clear Expectation Suite` to clean up before saving changes if you changed your mind.
If the Expectation Suite is already registered, it will instead show a button to delete the Expectation Suite.
The Expectation Suite editor: name, enabled flag, ingestion policy, and one row per expectation.
### Step 4: Save new data to a Feature Group
Use the python client to write a DataFrame to the Feature Group.
Note that if an expectation suite is enabled for a Feature Group, calling the `insert` method will run validation and default to uploading the corresponding validation report to Hopsworks.
The report is uploaded even if validation fails and 'STRICT' mode is selected.
### Step 5: Check Validation Results Summary
Hopsworks shows a visual summary of validation reports.
To check it out, go to your Feature Group overview and scroll to the expectation section.
Click on the `Validation Results` tab and check that all went according to plan.
Each row corresponds to an expectation in the suite.
Features can have several corresponding expectations and the same type of expectation can be applied to different features.
You can navigate to older reports using the dropdown menu.
Should you need more than the information displayed in the UI for e.g., debugging, the full report can be downloaded by clicking on the corresponding button.
### Step 6: Check Validation History
The `Validation Reports` tab in the Expectations section displays a brief history of recent validations.
Each row corresponds to a validation report, with some summary information about the success of the validation step.
You can download the full report by clicking the download icon button that appears at the end of the row.
The Expectations section on the feature group page, with the validation reports history.
## Code
Hopsworks python client interfaces with the Great Expectations library to enable you to add data validation to your feature engineering pipeline.
In this section, we show you how in a single line you enable automatic validation on each insertion of new data into your Feature Group.
Whether you have an existing Feature Group you want to add validation to or Follow the guide or get your hands dirty by running our [tutorial data validation notebook](https://colab.research.google.com/github/logicalclocks/hopsworks-tutorials/blob/master/integrations/great_expectations/fraud_batch_data_validation.ipynb) in google colab.
First checkout the pre-requisite and Hopsworks setup to follow the guide below.
Create a project, install the hopsworks client and connect via the generated API key.
You are ready to load your data in a DataFrame.
The second step is a short introduction to the relevant Great Expectations API to build data validation suited to your data.
Third and final step shows how to attach your Expectation Suite to the Feature Group to benefit from automatic validation on insertion capabilities.
### Step 1: Pre-requisite
In order to define and validate an expectation when writing to a Feature Group, you will need:
- A Hopsworks project.
If you don't have a project yet you can go to [run.hopsworks.ai](https://run.hopsworks.ai), signup with your email and create your first project.
- An API key, you can get one by going to "Account Settings" on [run.hopsworks.ai](https://run.hopsworks.ai).
- The [Hopsworks Python library](https://pypi.org/project/hopsworks) installed in your client.
See the [installation guide](../../client_installation/index.md).
#### Connect your notebook to Hopsworks
Connect the client running your notebooks to Hopsworks.
```python
import hopsworks
project = hopsworks.login()
fs = project.get_feature_store()
```
You will be prompt to paste your API key to connect the notebook to your project.
The `fs` Feature Store entity is now ready to be used to insert or read data from Hopsworks.
#### Import your data
Load your data in a DataFrame using the usual pandas API.
```python
import pandas as pd
df = pd.read_csv(
"https://repo.hops.works/master/hopsworks-tutorials/data/card_fraud_data/transactions.csv",
parse_dates=["datetime"],
)
df.head(3)
```
### Step 2: Great Expectation Introduction
To validate the data, we will use the [Great Expectations](https://greatexpectations.io/) library.
Below is a short introduction on how to build an Expectation Suite to validate your data.
Everything is done using the Great Expectations API so you can re-use any prior knowledge you may have of the library.
The Hopsworks `great-expectations` extra supports Great Expectations 0.18.12 and 1.17.1.
We recommend 1.17.1, and examples in this guide target that version.
#### Create an Expectation Suite
Create (or import an existing) expectation suite using the Great Expectations library.
This suite will hold all the validation tests we want to perform on our data before inserting them into Hopsworks.
```python
import great_expectations as gx
expectation_suite = gx.ExpectationSuite(name="validate_on_insert_suite")
```
#### Add Expectations in the Source Code
Add some expectations to your suite.
Each expectation configuration corresponds to a validation test to be run against your data.
```python
from great_expectations.expectations.expectation_configuration import (
ExpectationConfiguration,
)
expectation_suite.add_expectation_configuration(
ExpectationConfiguration(
type="expect_column_min_to_be_between",
kwargs={"column": "foo_id", "min_value": 0, "max_value": 1},
)
)
expectation_suite.add_expectation_configuration(
ExpectationConfiguration(
type="expect_column_value_lengths_to_be_between",
kwargs={"column": "bar_name", "min_value": 3, "max_value": 10},
)
)
```
!!! info "Migrating from Great Expectations 0.18.x"
The constructor argument was renamed from `expectation_suite_name=` to `name=` in 1.0.
`ExpectationConfiguration` now takes `type=` instead of `expectation_type=` and was moved out of `great_expectations.core` to `great_expectations.expectations.expectation_configuration`.
The Hopsworks SDK normalizes both shapes on the wire, so suites stored under either version remain readable.
#### Build a Suite with Typed Expectation Classes
Great Expectations 1.x also exposes a typed class for each expectation, which gives you IDE autocomplete on the kwargs.
You can mix typed instances and `ExpectationConfiguration` instances in the same suite.
```python
import great_expectations.expectations as gxe
typed_suite = gx.ExpectationSuite(
name="validate_on_insert_suite",
expectations=[
gxe.ExpectColumnMinToBeBetween(
column="foo_id", min_value=0, max_value=1
),
gxe.ExpectColumnValueLengthsToBeBetween(
column="bar_name", min_value=3, max_value=10
),
],
)
```
Once you have built an Expectation Suite you are satisfied with, it is time to create your first validation-enabled Feature Group.
### Step 3: Attach an Expectation Suite to your Feature Group to enable Automatic Validation on Insertion
Writing data in Hopsworks is done using Feature Groups.
Once a Feature Group is registered in the Feature Store, you can use it to insert your pandas DataFrames.
For more information see [create Feature Group](create.md).
To benefit from automatic validation on insertion, attach your newly created Expectation Suite when creating the Feature Group:
```python
fg = fs.create_feature_group(
"fg_with_data_validation",
version=1,
description="Validated data",
primary_key=["foo_id"],
online_enabled=False,
expectation_suite=expectation_suite,
)
```
or, if the Feature Group already exist, you can simply run:
```python
fg.save_expectation_suite(expectation_suite)
```
That is all there is to it.
Hopsworks will now automatically use your suite to validate the DataFrames you want to write to the Feature Group.
Try it out!
```python
job, validation_report = fg.insert(df.head(5))
```
As you can see, Hopsworks runs the validation in the client before attempting to insert the data.
By default, Hopsworks will try to insert the data even if validation fails to prevent data loss.
However it can be configured for production setup to be more restrictive, checkout the [data validation advanced guide](data_validation_advanced.md).
!!!info
Note that once the Expectation Suite is attached to the Feature Group, any subsequent attempt to insert to this Feature Group will apply the Data Validation step even from a different client or in a scheduled job.
### Step 4: Data Quality Monitoring
Upon running validation, Great Expectations generates a report to help you assess the quality of your data.
Nothing to do here, Hopsworks client automatically uploads the validation report to the backend when ingesting new data.
It enables you to monitor the quality of the inserted data in the Feature Group over time.
You can checkout a summary of the reports in the UI on your Feature Group page.
As you can see, your Feature Group conveniently gather all in one place: your data, the Expectation Suite and the reports generated each time you inserted data!
Hopsworks client API allows you to retrieve validation reports for further analysis.
```python
# load multiple reports
validation_reports = fg.get_all_validation_reports()
# convenience method for rapid development
ge_latest_report = fg.get_latest_validation_report()
```
Similarly you can retrieve the historic of validation results for a particular expectation, e.g to plot a time-series of a given expectation observed value over time.
```python
validation_history = fg.get_validation_history(expectation_id=1)
```
You can find the expectation IDs in the UI or using `fg.get_expectation_suite()` and looking them up in the expectation's `meta` field under the `expectationId` key.
!!! info
If Validation Reports or Results are too long, they can be truncated to fit in the database.
A full version of the reports can be downloaded from the UI.
## Conclusion
The integration between Hopsworks and Great Expectations makes it simple to add a data validation step to your feature engineering pipeline.
Build your Expectation Suite and attach it to your Feature Group with a single line of code.
No need to add any code to your pipeline or job scripts, calling `fg.insert` will now automatically validate the data before inserting them in the Feature Group.
The validation reports are stored along your data in Hopsworks allowing us to provide basic monitoring capabilities to quickly spot a data quality issue in the UI.
## Going Further
If you wish to find out more about how to use the data validation API or best practices for development or production pipelines in Hopsworks, checkout the [advanced guide](data_validation_advanced.md) and [best practices guide](data_validation_best_practices.md).
================================================================================
# Advanced guide
Source: https://docs.hopsworks.ai/latest/user_guides/fs/feature_group/data_validation_advanced/
# Advanced Data Validation Options and Best Practices
The introduction to the data validation guide can be found in the [Data Validation Guide](data_validation.md).
The notebook example to get started with Data Validation in Hopsworks can be found in the [Fraud Batch Data Validation Tutorial](https://colab.research.google.com/github/logicalclocks/hopsworks-tutorials/blob/master/integrations/great_expectations/fraud_batch_data_validation.ipynb).
## Data Validation Configuration Options in Hopsworks
### Validation Ingestion Policy
Depending on your use case you can setup data validation as a monitoring or gatekeeping tool when trying to insert new data in your Feature Group.
Switch behaviour by using the `validation_ingestion_policy` kwarg:
- `"ALWAYS"` is the default option and will attempt to insert the data regardless of the validation result.
Hassle free, it is ideal to monitor data ingestion in a development setup.
- `"STRICT"` is the best option for production ready projects.
This will prevent insertion of DataFrames which do not pass all data quality requirements.
Ideal to avoid "garbage-in, garbage-out" scenarios, at the price of a potential loss of data.
Check out the best practice section for more on that.
#### Validation Ingestion Policy in UI
Go to the Feature Group edit page, in the Expectation section you can choose between the options above.
#### Validation Ingestion Policy in Python
```python
fg.expectation_suite.validation_ingestion_policy = "ALWAYS" # "STRICT"
```
If your suite is registered with Hopsworks, it will persist the change to the server.
### Disable Data Validation
Should you wish to do so, you can disable data validation on a punctual basis or until further notice.
#### Disable Data Validation in UI
You can do it in the UI in the Expectation section of the Feature Group edit page.
Simply tick or untick the enabled checkbox.
This will be used as the default option but can be overridden via the API.
#### Disable Data Validation in Python
To disable data validation until further notice in the API, you can update the `run_validation` field of the expectation suite.
If your suite is registered with Hopsworks, this will persist the change to the server.
```python
fg.expectation_suite.run_validation = False
```
If you wish to override the default behaviour of the suite when inserting data in the Feature Group, you can do so via the `validation_options` kwarg.
The example below will enable validation for this insertion only.
```python
fg.insert(df_to_validate, validation_options={"run_validation": True})
```
We recommend to avoid using this option in scheduled job as it silently changes the expected behaviour that is displayed in the UI and prevents changes to the default behaviour to change the behaviour of the job.
### Edit Expectations
The one constant in life is change.
If you need to add, remove or edit an expectation you can do it both in the UI or via the python client.
Note that changing the expectation type or its corresponding feature will throw an error in order to preserve a meaningful validation history.
#### Edit Expectations in UI
Go to the Feature Group edit page, in the expectation section.
You can click on the expectation you want to edit and edit the json configuration.
Check out Great Expectations documentation if you need more information on a particular expectation.
#### Edit Expectations in Python
There are several way to edit an Expectation in the python client.
You can use Great Expectations API or directly go through Hopsworks.
In the latter case, if you want to edit or remove an expectation, you will need the Hopsworks expectation ID.
It can be found in the UI or in the meta field of an expectation.
Note that you must have inserted data in the FG and attached the expectation suite to enable the Expectation API.
Get an expectation with a given id:
```python
my_expectation = fg.expectation_suite.get_expectation(
expectation_id=my_expectation_id
)
```
Add a new expectation:
```python
from great_expectations.expectations.expectation_configuration import (
ExpectationConfiguration,
)
new_expectation = ExpectationConfiguration(
type="expect_column_values_to_not_be_null",
kwargs={"column": "foo_id", "mostly": 1},
)
fg.expectation_suite.add_expectation(new_expectation)
```
The single-expectation API on `fg.expectation_suite` accepts `ExpectationConfiguration` instances and plain dicts.
On 0.18.x the same API still works with `ge.core.ExpectationConfiguration(expectation_type=...)`.
Edit expectation kwargs of an existing expectation :
```python
existing_expectation = fg.expectation_suite.get_expectation(
expectation_id=existing_expectation_id
)
existing_expectation.kwargs["mostly"] = 0.95
fg.expectation_suite.replace_expectation(existing_expectation)
```
Remove an expectation:
```python
fg.expectation_suite.remove_expectation(
expectation_id=id_of_expectation_to_delete
)
```
If you want to deal only with the Great Expectations API:
```python
my_suite = fg.get_expectation_suite()
my_suite.add_expectation_configuration(new_expectation)
fg.save_expectation_suite(my_suite)
```
`add_expectation_configuration` is the right method to call when you only have the suite in hand (no `DataContext`).
On 0.18.x the equivalent method was `my_suite.add_expectation(new_expectation)`; in 1.x `add_expectation` instead requires an active `DataContext` and is meant for context-managed suites.
### Save Validation Reports
When running validation using Great Expectations, a validation report is generated containing all validation results for the different expectations.
Each result provides information about whether the provided DataFrame conforms to the corresponding expectation.
These reports can be stored in Hopsworks to save a validation history for the data written to a particular Feature Group.
The boilerplate of uploading report on insertion is taken care of by hopsworks, however for custom pipelines we provide an alternative method in the python client.
The UI does not currently support upload of a validation report.
#### Save Validation Reports in Python
```python
fg.save_validation_report(ge_report)
```
### Monitor and Fetch Validation Reports
A summary of uploaded reports will then be available via an API call or in the Hopsworks UI enabling easy monitoring.
For in-depth analysis, it is possible to download the complete report from the UI.
#### Monitor and Fetch Validation Reports in UI
Open the Feature Group overview page and go to the Expectations section.
One tab allows you to check the report history with general information, while the other tab allows you to explore a summary of the result for individual expectations.
#### Monitor and Fetch Validation Reports in Python
```python
# convenience method for rapid development
ge_latest_report = fg.get_latest_validation_report()
# fetching the latest summary prints a link to the UI
# where you can download full report if summary is insufficient
# or load multiple reports
validation_history = fg.get_all_validation_reports()
```
### Validate Your Data Manually
While Hopsworks provides automatic validation on insertion logic, we recognise that some use cases may require a more fine-grained control over the validation process.
Therefore, Feature Group objects offers a convenience wrapper around Great Expectations to manually trigger validation using the registered Expectation Suite.
#### Validate Your Data Manually in UI
You can validate data already ingested in the Feature Group by going to the Feature Group overview page.
In the top right corner is a button to trigger a validation.
The button will launch a job which will read the Feature Group data, run validation and persist the associated report.
#### Validate Your Data Manually in Python
```python
ge_report = fg.validate(df, ingestion_result="EXPERIMENT")
# set the save_report parameter to False to skip uploading the report to Hopsworks
# ge_report = fg.validate(df, save_report=False)
```
If you want to apply validation to the data already in the Feature Group you can call the `.validate` without providing data.
It will read the data in the Feature Group.
```python
report = fg.validate()
```
As validation objects returned by Hopsworks are native Great Expectations objects you can also run validation directly through the Great Expectations API.
Great Expectations 1.x removed `ge.from_pandas` and replaced it with the `Context → DataSource → Asset → Batch` pattern:
```python
import great_expectations as gx
context = gx.get_context(mode="ephemeral")
data_source = context.data_sources.add_pandas("hopsworks_pandas")
asset = data_source.add_dataframe_asset("hopsworks_asset")
batch_definition = asset.add_batch_definition_whole_dataframe("hopsworks_batch")
batch = batch_definition.get_batch(batch_parameters={"dataframe": df})
ge_report = batch.validate(fg.get_expectation_suite())
```
For most pipelines you should prefer `fg.validate(df)`: Hopsworks runs the same chain internally and uploads the report to the backend in a single call.
Note that you should always use an expectation suite that has been saved to Hopsworks if you intend to upload the associated validation report.
================================================================================
# Best practices
Source: https://docs.hopsworks.ai/latest/user_guides/fs/feature_group/data_validation_best_practices/
# Best practices
Below is a set of recommendations and code snippets to help our users follow best practices when it comes to integrating a data validation step in your feature engineering pipelines.
Rather than being prescriptive, we want to showcase how the API and configuration options can help adapt validation to your use-case.
## Development
Data validation is generally considered to be a production-only feature and as such is often only setup once a project has reached the end of the development phase.
At Hopsworks, we think there is a lot of value in setting up validation during early development.
That's why we made it quick to get started and ensured that by default data validation is never an obstacle to inserting data.
### Validate Early
As often with data validation, the best piece of advice is to set it up early in your development process.
Use this phase to build a history you can then use when it becomes time to set quality requirements for a project in production.
We made a code snippet to help you get started quickly:
```python
import pandas as pd
import great_expectations as gx
from great_expectations.expectations.expectation_configuration import (
ExpectationConfiguration,
)
# Load sample data.
# Replace it with your own!
my_data_df = pd.read_csv(
"https://repo.hops.works/master/hopsworks-tutorials/data/card_fraud_data/credit_cards.csv"
)
# Build a starter Expectation Suite that asserts every column exists and is
# not null. This is a useful baseline; tighten it as you learn the data.
# Build the list first and pass it to the constructor: GE 1.x's
# add_expectation_configuration() deduplicates expect_column_to_exist entries,
# but the constructor's expectations= argument preserves every entry.
expectations = []
for column in my_data_df.columns:
expectations.append(
ExpectationConfiguration(
type="expect_column_to_exist", kwargs={"column": column}
)
)
expectations.append(
ExpectationConfiguration(
type="expect_column_values_to_not_be_null", kwargs={"column": column}
)
)
expectation_suite = gx.ExpectationSuite(
name="credit_cards_baseline", expectations=expectations
)
# Create a Feature Group on Hopsworks with the suite attached.
# Don't forget to change the primary key!
my_validated_data_fg = fs.get_or_create_feature_group(
name="my_validated_data_fg",
version=1,
description="My data",
primary_key=["cc_num"],
expectation_suite=expectation_suite,
)
```
Any data you insert in the Feature Group from now will be validated and a report will be uploaded to Hopsworks.
```python
# Insert and validate your data
insert_job, validation_report = my_validated_data_fg.insert(my_data_df)
```
Great Expectations 0.18.x shipped a `BasicSuiteBuilderProfiler` that auto-generated a starter suite from a sample DataFrame.
That profiler was removed in 1.0 with no in-tree replacement, so the snippet above builds a minimal suite by hand.
For real workloads, iterate on it: add `expect_column_(min/max/mean/stdev)_to_be_between`, `expect_column_values_to_be_unique`, and similar checks as you understand the distributions.
Attaching the suite when creating the Feature Group ensures every piece of data finding its way into Hopsworks gets validated.
Hopsworks defaults to its `"ALWAYS"` ingestion policy, meaning data is ingested whether validation succeeds or not.
This way data validation is not a barrier, just a monitoring tool.
### Identify Unreliable Features
Once you setup data validation, every insertion will upload a validation report to Hopsworks.
Identifying Features which often have null values or wild statistical variations can help detecting unreliable Features that need refinements or should be avoided.
Here are a few expectations you might find useful:
- `expect_column_values_to_not_be_null`
- `expect_column_(min/max/mean/stdev)_to_be_between`
- `expect_column_values_to_be_unique`
### Get the stakeholders involved
Hopsworks UI helps involve every project stakeholder by enabling both setting and monitoring of data quality requirements.
No coding skills needed! You can monitor data quality requirements by checking out the validation reports and results on the Feature Group page.
If you need to set or edit the existing requirements, you can go on the Feature Group edit page.
The Expectation suite section allows you to edit individual expectations and set success parameters that match ever changing business requirements.
## Production
Models in production require high-quality data to make accurate predictions for your customers.
Hopsworks can use your Expectation Suite as a gatekeeper to make it simple to prevent low-quality data to make its way into production.
Below are some simple tips and snippets to make the most of your data validation when your project is ready to enter its production phase.
### Be Strict in Production
Whether you use an existing or create a new (recommended) Feature Group for production, we recommend you set the validation ingestion policy of your Expectation Suite to `"STRICT"`.
```python
fg_prod.save_expectation_suite(my_suite, validation_ingestion_policy="STRICT")
```
In this setup, Hopsworks will abort inserting a DataFrame that does not successfully fulfill all expectations in the attached Expectation Suite.
This ensures data quality standards are upheld for every insertion and provide downstream users with strong guarantees.
### Avoid Data Loss on materialization jobs
Aborting insertions of DataFrames which do not satisfy the data quality standards can lead to data loss in your materialization job.
To avoid such loss we recommend creating a duplicate Feature Group with the same Expectation Suite in `"ALWAYS"` mode which will hold the rejected data.
```python
job, report = fg_prod.insert(df)
if report["success"] is False:
job, report = fg_rejected.insert(df)
```
### Take Advantage of the Validation History
You can easily retrieve the validation history of a specific expectation to export it to your favourite visualisation tool.
You can filter on time and on whether insertion was successful or not.
```python
validation_history = fg.get_validation_history(
expectation_id=my_id, filter_by=["REJECTED", "UNKNOWN"], ge_type=False
)
timeseries = pd.DataFrame(
{
"observed_value": [
res.result["observed_value"] for res in validation_history
],
"validation_time": [res.validation_time for res in validation_history],
}
)
# export to your preferred Dashboard
```
### Setup Alerts
While checking your feature engineering pipeline executed properly in the morning can be good enough in the development phase, it won't make the cut for demanding production use-cases.
In Hopsworks, you can setup alerts if ingestion fails or succeeds.
First you will need to configure your preferred communication endpoint: slack, email or pagerduty.
Check out [this page](../../../setup_installation/admin/alert.md) for more information on how to set it up.
A typical use-case would be to add an alert on ingestion success to a Feature Group you created to hold data that failed validation.
Here is a quick walkthrough:
1. Go the Feature Group page in the UI
2. Scroll down and click on the `Add an alert` button.
3. Choose the trigger, receiver and severity and click save.
## Conclusion
Hopsworks extends Great Expectations by automatically running the validation, persisting the reports along your data and allowing you to monitor data quality in its UI.
How you decide to make use of these tools depends on your application and requirements.
Whether in development or in production, real-time or batch, we think there is configuration that will work for your team.
Check out our [quick hands-on tutorial](https://colab.research.google.com/github/logicalclocks/hopsworks-tutorials/blob/master/integrations/great_expectations/fraud_batch_data_validation.ipynb) to start applying what you learned so far.
================================================================================
# Feature Monitoring
Source: https://docs.hopsworks.ai/latest/user_guides/fs/feature_group/feature_monitoring/
# Feature Monitoring for Feature Groups
Feature Monitoring complements the Hopsworks data validation capabilities for Feature Groups by allowing you to monitor your data once they have been ingested into the Feature Store.
Hopsworks feature monitoring is centered around two functionalities: **scheduled statistics** and **statistics comparison**.
Before continuing with this guide, see the [Feature monitoring guide](../feature_monitoring/index.md) to learn more about how feature monitoring works, and get familiar with the different use cases of feature monitoring for Feature Groups described in the **Use cases** sections of the [Scheduled statistics guide](../feature_monitoring/scheduled_statistics.md#use-cases) and [Statistics comparison guide](../feature_monitoring/statistics_comparison.md#use-cases).
!!! warning "Limited UI support"
Currently, feature monitoring can only be configured using the [Hopsworks Python library](https://pypi.org/project/hopsworks).
However, you can enable/disable a feature monitoring configuration or trigger the statistics comparison manually from the UI.
## Code
In this section, we show you how to setup feature monitoring in a Feature Group using the ==Hopsworks Python library==.
Alternatively, you can get started quickly by running our [tutorial for feature monitoring](https://github.com/logicalclocks/hopsworks-tutorials/blob/master/api_examples/feature_monitoring.ipynb).
First, checkout the pre-requisite and Hopsworks setup to follow the guide below.
Create a project, install the [Hopsworks Python library](https://pypi.org/project/hopsworks) in your environment, connect via the generated API key.
The second step is to start a new configuration for feature monitoring.
After that, you can optionally define a detection window of data to compute statistics on, or use the default detection window (i.e., whole feature data).
If you want to setup scheduled statistics alone, you can jump to the last step to save your configuration.
Otherwise, the third and fourth steps are also optional and show you how to setup the comparison of statistics on a schedule by defining a reference window and specifying the statistics metric to monitor.
### Step 1: Pre-requisite
In order to setup feature monitoring for a Feature Group, you will need:
- A Hopsworks project.
If you don't have a project yet you can go to [run.hopsworks.ai](https://run.hopsworks.ai), signup with your email and create your first project.
- An API key, you can get one by going to "Account Settings" on [run.hopsworks.ai](https://run.hopsworks.ai).
- The Hopsworks Python library installed in your client.
See the [installation guide](../../client_installation/index.md).
- A Feature Group
#### Connect your notebook to Hopsworks
Connect the client running your notebooks to Hopsworks.
=== "Python"
```python
import hopsworks
project = hopsworks.login()
fs = project.get_feature_store()
```
See the API reference for [`hopsworks.login`][hopsworks.login] and [`Project.get_feature_store`][hopsworks_common.project.Project.get_feature_store].
You will be prompted to paste your API key to connect the notebook to your project.
The `fs` Feature Store entity is now ready to be used to insert or read data from Hopsworks.
#### Get or create a Feature Group
Feature monitoring can be enabled on already created Feature Groups.
We suggest you read the [Feature Group](../../../concepts/fs/feature_group/fg_overview.md) concept page to understand what a feature group is and how it fits in the ML pipeline.
We also suggest you familiarize with the APIs to [create a feature group](./create.md).
The following is a code example for getting or creating a Feature Group with name `trans_fg` for transaction data.
=== "Python"
```python
# Retrieve an existing feature group
trans_fg = fs.get_feature_group("trans_fg", version=1)
# Or, create a new feature group with transactions
trans_fg = fs.get_or_create_feature_group(
name="trans_fg",
version=1,
description="Transaction data",
primary_key=["cc_num"],
event_time="datetime",
)
trans_fg.insert(transactions_df)
```
See the API reference for [`FeatureStore.get_feature_group`][hsfs.feature_store.FeatureStore.get_feature_group] and [`FeatureStore.get_or_create_feature_group`][hsfs.feature_store.FeatureStore.get_or_create_feature_group].
### Step 2: Initialize configuration
#### Scheduled statistics
You can setup statistics monitoring on a ==single feature or multiple features== of your Feature Group.
=== "Python"
```python
# compute statistics for all the features
fg_monitoring_config = trans_fg.create_scheduled_statistics(
name="trans_fg_all_features_monitoring",
description="Compute statistics on all data of all features of the Feature Group on a daily basis",
)
# or for one or more specific features
fg_monitoring_config = trans_fg.create_scheduled_statistics(
name="trans_fg_amount_monitoring",
description="Compute statistics on all data of selected features of the Feature Group on a daily basis",
feature_names=["amount"],
)
```
See the API reference for [`FeatureGroup.create_scheduled_statistics`][hsfs.feature_group.FeatureGroup.create_scheduled_statistics].
#### Statistics comparison
When enabling the comparison of statistics in a feature monitoring configuration, the feature to compare is selected later in the `compare_on` (or `compare_on_distribution`) method, not in `create_feature_monitoring`.
You can create multiple feature monitoring configurations for the same Feature Group.
=== "Python"
```python
fg_monitoring_config = trans_fg.create_feature_monitoring(
name="trans_fg_amount_monitoring",
description="Compute and compare descriptive statistics on the Feature Group on a daily basis",
)
```
See the API reference for [`FeatureGroup.create_feature_monitoring`][hsfs.feature_group.FeatureGroup.create_feature_monitoring].
#### Custom schedule
By default, the computation of statistics is scheduled to run endlessly, every day at 12PM.
You can modify the default schedule by adjusting the `cron_expression`, `start_date_time` and `end_date_time` parameters.
To compute statistics on only a subset of the feature data, use the `row_percentage` parameter of `with_detection_window` (see Step 3).
=== "Python"
```python
fg_monitoring_config = trans_fg.create_scheduled_statistics(
name="trans_fg_all_features_monitoring",
description="Compute statistics on all data of all features of the Feature Group on a weekly basis",
cron_expression="0 0 12 ? * MON *", # weekly
)
# or
fg_monitoring_config = trans_fg.create_feature_monitoring(
name="trans_fg_amount_monitoring",
description="Compute and compare descriptive statistics on the Feature Group on a weekly basis",
cron_expression="0 0 12 ? * MON *", # weekly
)
```
### Step 3: (Optional) Define a detection window
By default, the detection window is an _expanding window_ covering the whole Feature Group data.
You can define a different detection window using the `window_length` and `time_offset` parameters provided in the `with_detection_window` method.
Additionally, you can specify the percentage of feature data on which statistics will be computed using the `row_percentage` parameter.
=== "Python"
```python
fm_monitoring_config.with_detection_window(
window_length="1w", # data ingested during one week
time_offset="1w", # starting from last week
row_percentage=0.8, # use 80% of the data
)
```
See the API reference for [`FeatureMonitoringConfig.with_detection_window`][hsfs.core.feature_monitoring_config.FeatureMonitoringConfig.with_detection_window].
#### Time basis of the windows
Rolling windows select rows by an event-time feature when the Feature Group declares one, and by commit time otherwise.
The `event_time` parameter of `create_scheduled_statistics` and `create_feature_monitoring` overrides that default for the whole configuration, detection and reference windows alike.
Pass a feature name to use another timestamp, date or epoch feature of the Feature Group, or `False` to select rows by commit time.
=== "Python"
```python
# windows over the transaction time, the Feature Group event_time (default)
fg_monitoring_config = trans_fg.create_feature_monitoring(
name="trans_fg_amount_monitoring",
)
# windows over another time feature of the Feature Group
fg_monitoring_config = trans_fg.create_feature_monitoring(
name="trans_fg_amount_monitoring_by_settlement",
event_time="settlement_date",
)
# windows over the time the rows were written (commit time)
fg_monitoring_config = trans_fg.create_feature_monitoring(
name="trans_fg_amount_monitoring_by_commit",
event_time=False,
)
```
See [Time basis](../feature_monitoring/scheduled_statistics.md#time-basis) for how the two bases differ.
### Step 4: (Optional) Define a reference window
When setting up feature monitoring for a Feature Group, you can compare the detection statistics against a reference window of feature data.
A reference window is defined with the `with_reference_window` method.
=== "Python"
```python
# compare statistics against a reference window
fm_monitoring_config.with_reference_window(
window_length="1w", # data ingested during one week
time_offset="2w", # starting from two weeks ago
row_percentage=0.8, # use 80% of the data
)
```
See the API reference for [`FeatureMonitoringConfig.with_reference_window`][hsfs.core.feature_monitoring_config.FeatureMonitoringConfig.with_reference_window].
!!! info "Comparing against a specific value"
Instead of a reference window, you can compare the detection statistics against a fixed reference value (i.e., a window of size 1).
In that case, skip this step and pass the `specific_value` parameter to `compare_on` in Step 5.
### Step 5.A: (Optional) Compare on a scalar metric
In order to compare detection and reference statistics, you need to provide the criteria for such comparison.
First, you select the feature and the metric to consider in the comparison using the `feature_name` and `metric` parameters.
Then, you can define a relative or absolute threshold using the `threshold` and `relative` parameters.
=== "Python"
```python
# compare against a reference window
fm_monitoring_config.compare_on(
feature_name="amount", # the feature to compare
metric="mean",
threshold=0.2, # a relative change over 20% is considered anomalous
relative=True, # relative or absolute change
strict=False, # strict or relaxed comparison
)
# or compare against a specific value instead of a reference window
fm_monitoring_config.compare_on(
feature_name="amount",
metric="mean",
specific_value=100,
threshold=0.2,
relative=True,
)
```
See the API reference for [`FeatureMonitoringConfig.compare_on`][hsfs.core.feature_monitoring_config.FeatureMonitoringConfig.compare_on].
!!! info "Difference values and thresholds"
For more information about the computation of difference values and the comparison against threshold bounds see the [Comparison criteria section](../feature_monitoring/statistics_comparison.md#comparison-criteria) in the Statistics comparison guide.
### Step 5.B: (Optional) Compare on the whole distribution
Alternatively, instead of a single scalar metric, you can detect drift in the shape of a feature's distribution using `compare_on_distribution`.
Select a distribution distance metric (e.g., `PSI`) and a threshold.
A reference window (Step 4) is required for distribution comparison.
=== "Python"
```python
fm_monitoring_config.compare_on_distribution(
feature_name="amount", # the feature to compare
metric="PSI",
threshold=0.2, # a distance above 0.2 is considered a significant shift
)
```
See the API reference for [`FeatureMonitoringConfig.compare_on_distribution`][hsfs.core.feature_monitoring_config.FeatureMonitoringConfig.compare_on_distribution].
!!! tip "More distribution options"
See the [Distribution comparison guide](../feature_monitoring/distribution_comparison.md) for the full list of metrics and binning strategies.
### Step 6: Save configuration
Finally, you can save your feature monitoring configuration by calling the `save` method.
Once the configuration is saved, the schedule for the statistics computation and comparison will be activated automatically.
=== "Python"
```python
fm_monitoring_config.save()
```
See the API reference for [`FeatureMonitoringConfig.save`][hsfs.core.feature_monitoring_config.FeatureMonitoringConfig.save].
### Retrieve configurations and history
Once saved, you can retrieve your feature monitoring configurations and the results of past executions directly from the Feature Group.
=== "Python"
```python
# fetch all configurations attached to the feature group
configs = trans_fg.get_feature_monitoring_configs()
# or a single configuration by name
config = trans_fg.get_feature_monitoring_configs(name="trans_fg_amount_monitoring")
# fetch the history of monitoring results (with computed statistics)
history = trans_fg.get_feature_monitoring_history(
config_name="trans_fg_amount_monitoring",
with_statistics=True,
)
```
See the API reference for [`FeatureGroup.get_feature_monitoring_configs`][hsfs.feature_group.FeatureGroup.get_feature_monitoring_configs] and [`FeatureGroup.get_feature_monitoring_history`][hsfs.feature_group.FeatureGroup.get_feature_monitoring_history].
!!! info "Explore the API"
The [`FeatureMonitoringConfig`][hsfs.core.feature_monitoring_config.FeatureMonitoringConfig] reference documents the full set of available methods, such as enabling or disabling a configuration, triggering it manually, or deleting it.
================================================================================
# Notification
Source: https://docs.hopsworks.ai/latest/user_guides/fs/feature_group/notification/
# Change Data Capture for feature groups
## Introduction
Changes to online-enabled feature groups can be captured by listening to events on specified topics.
This optimizes the user experience by allowing users to proactively make predictions as soon as there is an update on the features.
In this guide you will learn how to enable Change Data Capture (CDC) for online feature groups within Hopsworks, showing examples in Hopsworks APIs as well as the user interface.
## Prerequisites
Before you begin this guide we suggest you read the [Feature Group](../../../concepts/fs/feature_group/fg_overview.md) concept page to understand what a feature group is and how it fits in the ML pipeline.
Subsequently [create a Kafka topic](../../projects/kafka/create_topic.md), this topic will be used for storing Change Data Capture events.
## Using Hopsworks APIs
### Create a Feature Group with Change Data Capture using Python
To enable Change Data Capture for an online-enabled feature group using the Hopsworks APIs you need to [create a feature group](./create.md) and set the `notification_topic_name` properties value to the previously created topic.
=== "Python"
```python
fg = fs.create_feature_group(
name="feature_group_name",
version=feature_group_version,
primary_key=feature_group_primary_keys,
online_enabled=True,
notification_topic_name="notification_topic_name",
)
```
### Update Feature Group with Change Data Capture topic using Python
The notification topic name can be changed after the creation of the feature group.
By setting the `notification_topic_name` value to `None` or empty string notification will be disabled.
With the default configuration, it can take up to 30 minutes for these changes to take place since the onlinefs service internally caches feature groups.
=== "Python"
```python
fg.update_notification_topic_name(
notification_topic_name="new_notification_topic_name"
)
```
## Using UI
### Update Feature Group with Change Data Capture topic using UI
The notification topic name can be changed after creation by editing the feature group.
By setting the `CDC topic name` value to empty the notifications will be disabled.
With the default configuration, it can take up to 30 minutes for these changes to take place since the onlinefs service internally caches feature groups.
## Example of Change Data Capture event
Once properly set up the online feature store service will produce events to the provided topic when data ingestion is completed for records.
Here is an example output:
```jsonc
{
"projectName":"project_name", // name of the project the feature group belongs to
"projectId":119, // id of the project the feature group belongs to
"featureStoreId":67, // feature store where changes took place
"featureGroupId":14, // id of the feature group
"featureGroupName":"fg_name", // name of the feature group
"featureGroupVersion":1, // version of the feature group
"entry":{ // values of the affected feature group entry
"id":"15",
"text":"test"
},
"featureViews":[ // list of feature views affected
{
"projectName":"project_name", // name of the project the feature view belongs to
"id":9, // id of the feature view
"name":"test", // name of the feature view
"version":1, // version of the feature view
"featurestoreId":67 // feature store where feature view resides
}
]
}
```
The list of `featureViews` in the event could be outdated for up to 10 minutes, due to internal logging in onlinefs service.
================================================================================
# Ingestion Topic
Source: https://docs.hopsworks.ai/latest/user_guides/fs/feature_group/topic/
# How to configure the ingestion topic of a Feature Group { #feature-group-ingestion-topic }
## Introduction
Feature groups written with the [streaming write API][streaming-write-api] do not write to the online and offline feature store directly.
Every insert produces the rows to a Kafka topic, from which the OnlineFS service writes them to the online feature store and the offline materialization job writes them to the offline feature store.
By default all feature groups in a project share a single topic.
That works well until one feature group writes enough that the offline materialization jobs of the others spend their runs reading and discarding its records; [the choice of topic for data ingestion][the-choice-of-topic-for-data-ingestion] covers when a dedicated topic is worth it.
In this guide you will learn how to give a feature group a topic of its own, and how to change that topic after the feature group has been created.
## Prerequisites
Before you begin this guide we suggest you read the [create feature group][create-feature-group] guide, which covers the `topic_name` parameter.
## Which topic a feature group uses
Hopsworks resolves the ingestion topic of a feature group in the following order:
1. The `topic_name` of the feature group, if one is set.
2. The topic of the project, if one is set.
3. The project default, which is `_onlinefs` for online-enabled feature groups and `` otherwise.
Setting `topic_name` to an empty string clears the feature group override, so the feature group falls back to the project topic.
!!! note "Topics of online-enabled feature groups must end in `_onlinefs`"
The OnlineFS service subscribes to the topics matching the `.*_onlinefs` pattern, so a topic whose name does not match it is never consumed into the online feature store.
Administrators can change the pattern with the `onlinefs/kafka_consumer/topic_pattern` configuration option, or replace it with an explicit topic list as described in the [external Kafka cluster][external-kafka-cluster] guide.
## Before you change the topic
Changing the topic of a feature group that already holds data is not a migration, and there are two consequences to plan for.
!!! warning "Pending data is not migrated automatically"
Rows that were already inserted into the old topic but not yet consumed are never materialized to the new topic.
Wait until all in-flight processing has completed before switching, for example by following the [online ingestion observability][online-ingestion-observability] of the feature group and letting the offline materialization job finish.
!!! warning "Offline materialization restarts from the earliest offset"
The offline materialization job stores the Kafka offsets it has consumed together with the name of the topic they belong to.
When it detects that the topic has changed, those offsets are meaningless, so it starts from the earliest available offset of the new topic.
This reprocesses everything the new topic still retains and can produce duplicates in the offline feature store.
## Using Hopsworks APIs
### Set the topic when creating the feature group
Pass `topic_name` to `create_feature_group` to give the feature group its own topic from the start:
=== "Python"
```python
fg = fs.create_feature_group(
name="feature_group_name",
version=1,
primary_key=["id"],
online_enabled=True,
topic_name="feature_group_name_onlinefs",
)
```
### Change the topic of an existing feature group
Use [`FeatureGroup.update_topic_name`][hsfs.feature_group.FeatureGroup.update_topic_name] to point an existing feature group at a different topic:
=== "Python"
```python
fg = fs.get_feature_group("feature_group_name", version=1)
fg.update_topic_name(topic_name="feature_group_name_onlinefs")
```
The call emits the two warnings above as Python warnings before sending the request, and updates your local metadata object only once the backend has accepted the change.
## Using the UI
### Change the feature group topic
Open the feature group, click `Edit`, and set the `Topic name` field.
The field is shown for stream feature groups and for online-enabled feature groups, because those are the ones that ingest through Kafka.
Saving a changed topic name asks you to confirm the two consequences described above before the update is sent.
The topic a feature group currently uses is shown on its overview page.
### Change the project topic
The project topic is the default for every feature group in the project that does not set its own.
Navigate to `Project Settings` → `Kafka` and use `Edit project topic` in the `Project Topic` card.
As with a feature group topic, you are asked to confirm before the change is applied.
!!! note
The `Project Topic` card is only shown when the cluster is configured to use an [external Kafka cluster][external-kafka-cluster], since that is the case in which the topic is not managed by Hopsworks.
## Topic creation
When you set a topic that does not exist yet, Hopsworks creates it in the project with the cluster defaults for feature store topics.
Topics count against the project's Kafka topic quota, which an administrator can raise as described in the [Kafka topics][kafka-topics] administration guide.
When the cluster is configured to use an external Kafka cluster, Hopsworks does not provision topics.
Create the topic in the external cluster first, otherwise ingestion fails as soon as the feature group starts producing to it.
## API Reference
[`FeatureGroup`][hsfs.feature_group.FeatureGroup]
================================================================================
# On-Demand Transformations
Source: https://docs.hopsworks.ai/latest/user_guides/fs/feature_group/on_demand_transformations/
# On-Demand Transformation Functions
[On-demand transformations](https://www.hopsworks.ai/dictionary/on-demand-transformation) produce on-demand features, which usually require parameters accessible during inference for their calculation.
Hopsworks facilitates the creation of on-demand transformations without introducing [online-offline skew](https://www.hopsworks.ai/dictionary/online-offline-feature-skew), ensuring consistency while allowing their dynamic computation during online inference.
## On Demand Transformation Function Creation
An on-demand transformation function may be created by associating a [transformation function](../transformation_functions.md) with a feature group.
Each on-demand transformation function can generate one or multiple on-demand features.
If the on-demand transformation function returns a single feature, it is automatically assigned the same name as the transformation function.
However, if it returns multiple features, they are by default named using the format `functionName_outputColumnNumber`.
For instance, in the example below, the on-demand transformation function `transaction_age` produces an on-demand feature named `transaction_age` and the on-demand transformation function `stripped_strings` produces the on-demand features names `stripped_strings_0` and `stripped_strings_1`.
Alternatively, the name of the resulting on-demand feature can be explicitly defined using the [`alias`](../transformation_functions.md#specifying-output-features-names-for-transformation-functions) function.
!!! warning "On-demand transformation"
All on-demand transformation functions attached to a feature group must have unique names and, in contrast to model-dependent transformations, they do not have access to training dataset statistics.
Each on-demand transformation function can map specific features to its arguments by explicitly providing their names as arguments to the transformation function.
If no feature names are provided, the transformation function will default to using features that match the name of the transformation function's argument.
!!! example "Creating on-demand transformation functions."
=== "Python"
```python
# Define transformation function
@hopsworks.udf(return_type=int, drop=["current_date"])
def transaction_age(transaction_date, current_date):
return (current_date - transaction_date).dt.days
@hopsworks.udf(return_type=[str, str], drop=["current_date"])
def stripped_strings(country, city):
return country.strip(), city.strip()
# Attach transformation function to feature group to create on-demand transformation function.
fg = feature_store.create_feature_group(
name="fg_transactions",
version=1,
description="Transaction Features",
online_enabled=True,
primary_key=["id"],
event_time="event_time",
transformation_functions=[transaction_age, stripped_strings],
)
```
### Specifying input features
The features to be used by the on-demand transformation function can be specified by providing the feature names as input to the transformation functions.
!!! example "Creating on-demand transformations by specifying features to be passed to transformation function."
=== "Python"
```python
fg = feature_store.create_feature_group(
name="fg_transactions",
version=1,
description="Transaction Features",
online_enabled=True,
primary_key=["id"],
event_time="event_time",
transformation_functions=[
age_transaction("transaction_time", "current_time")
],
)
```
## Usage
On-demand transformation functions attached to a feature group are automatically executed in the feature pipeline when you [insert data](./create.md#batch-write-api) into a feature group and [by the Python client while retrieving feature vectors](../feature_view/feature-vectors.md#retrieval) for online inference using feature views that contain on-demand features.
The on-demand features computed by on-demand transformation functions are positioned after all other features in a feature group and are ordered alphabetically by their names.
### Inserting data
All on-demand transformation functions attached to a feature group are executed whenever new data is inserted.
This process computes on-demand features from historical data.
The DataFrame used for insertion must include all features required for executing all on-demand transformation functions in the feature group.
Inserting on-demand features as historical features saves time and computational resources by removing the need to compute all on-demand features while generating training or batch data.
### Accessing on-demand features in feature views
A feature view can include on-demand features from feature groups by selecting them in the [query](../feature_view/query.md) used to create the feature view.
These on-demand features are equivalent to regular features, and [model-dependent transformations](../feature_view/model-dependent-transformations.md) can be applied to them if required.
!!! example "Creating feature view with on-demand features"
=== "Python"
```python
# Selecting on-demand features in query
query = fg.select(
["id", "feature1", "feature2", "on_demand_feature3", "on_demand_feature4"]
)
# Creating a feature view using a query that contains on-demand transformations and model-dependent transformations
feature_view = fs.create_feature_view(
name="transactions_view",
query=query,
transformation_functions=[
min_max_scaler("feature1"),
min_max_scaler("on_demand_feature3"),
],
)
```
### Computing on-demand features
On-demand features in the feature view are computed in real-time during online inference using the same on-demand transformation functions used to create them.
Hopsworks, by default, automatically computes all on-demand features when retrieving feature view input features (feature vectors) with the functions `get_feature_vector` and `get_feature_vectors`.
Additionally, on-demand features can be computed using the `compute_on_demand_features` function or by manually executing the same on-demand transformation function.
The values for the input parameters required to compute on-demand features can be provided using the `request_parameters` argument.
If values are not provided through the `request_parameters` argument, the transformation function will verify if the feature vector contains the necessary input parameters and will use those values instead.
However, if the required input parameters are also not present in the feature vector, an error will be thrown.
!!! note
By default the functions `get_feature_vector` and `get_feature_vectors` will apply model-dependent transformation present in the feature view after computing on-demand features.
#### Retrieving a feature vector
The `get_feature_vector` function retrieves a single feature vector based on the feature view's serving key(s).
The on-demand features in the feature vector can be computed using real-time data by passing a dictionary that associates the name of each input parameter needed for the on-demand transformation function with its respective new value to the `request_parameter` argument.
!!! example "Computing on-demand features while retrieving a feature vector"
=== "Python"
```python
feature_vector = feature_view.get_feature_vector(
entry={"id": 1},
request_parameter={
"transaction_time": datetime(2022, 12, 28, 23, 55, 59),
"current_time": datetime.now(),
},
)
```
#### Retrieving feature vectors
The `get_feature_vectors` function retrieves multiple feature vectors using a list of feature view serving keys.
The `request_parameter` in this case, can be a list of dictionaries that specifies the input parameters for the computation of on-demand features for each serving key or can be a dictionary if the on-demand transformations require the same parameters for all serving keys.
!!! example "Computing on-demand features while retrieving a feature vectors"
=== "Python"
```python
# Specify unique request parameters for each serving key.
feature_vector = feature_view.get_feature_vectors(
entry=[{"id": 1}, {"id": 2}],
request_parameter=[
{
"transaction_time": datetime(2022, 12, 28, 23, 55, 59),
"current_time": datetime.now(),
},
{
"transaction_time": datetime(2022, 11, 20, 12, 50, 00),
"current_time": datetime.now(),
},
],
)
# Specify common request parameters for all serving key.
feature_vector = feature_view.get_feature_vectors(
entry=[{"id": 1}, {"id": 2}],
request_parameter={
"transaction_time": datetime(2022, 12, 28, 23, 55, 59),
"current_time": datetime.now(),
},
)
```
#### Retrieving feature vector without on-demand features
The `get_feature_vector` and `get_feature_vectors` methods can return untransformed feature vectors without on-demand features by disabling model-dependent transformations and excluding on-demand features.
To achieve this, set the parameters `transform` and `on_demand_features` to `False`.
!!! example "Returning untransformed feature vectors"
=== "Python"
```python
untransformed_feature_vector = feature_view.get_feature_vector(
entry={"id": 1}, transform=False, on_demand_features=False
)
untransformed_feature_vectors = feature_view.get_feature_vectors(
entry=[{"id": 1}, {"id": 2}], transform=False, on_demand_features=False
)
```
#### Compute all on-demand features
The `compute_on_demand_features` function computes all on-demand features attached to a feature view and adds them to the feature vectors provided as input to the function.
This function does not apply model-dependent transformations to any of the features.
The `transform` function can be used to apply model-dependent transformations to the returned values if required.
The `request_parameter` in this case, can be a list of dictionaries that specifies the input parameters for the computation of on-demand features for each feature vector given as input to the function or can be a dictionary if the on-demand transformations require the same parameters for all input feature vectors.
!!! example "Computing all on-demand features and manually applying model dependent transformations."
=== "Python"
```python
# Specify request parameters for each serving key.
untransformed_feature_vector = feature_view.get_feature_vector(
entry={"id": 1}, transform=False, on_demand_features=False
)
# re-compute and add on-demand features to the feature vector
feature_vector_with_on_demand_features = fv.compute_on_demand_features(
untransformed_feature_vector,
request_parameter={
"transaction_time": datetime(2022, 12, 28, 23, 55, 59),
"current_time": datetime.now(),
},
)
# Applying model dependent transformations
encoded_feature_vector = fv.transform(feature_vector_with_on_demand_features)
# Specify request parameters for each serving key.
untransformed_feature_vectors = feature_view.get_feature_vectors(
entry=[{"id": 1}, {"id": 2}], transform=False, on_demand_features=False
)
# re-compute and add on-demand features to the feature vectors - Specify unique request parameter for each feature vector
feature_vectors_with_on_demand_features = fv.compute_on_demand_features(
untransformed_feature_vectors,
request_parameter=[
{
"transaction_time": datetime(2022, 12, 28, 23, 55, 59),
"current_time": datetime.now(),
},
{
"transaction_time": datetime(2022, 11, 20, 12, 50, 00),
"current_time": datetime.now(),
},
],
)
# re-compute and add on-demand feature to the feature vectors - Specify common request parameter for all feature vectors
feature_vectors_with_on_demand_features = fv.compute_on_demand_features(
untransformed_feature_vectors,
request_parameter={
"transaction_time": datetime(2022, 12, 28, 23, 55, 59),
"current_time": datetime.now(),
},
)
# Applying model dependent transformations
encoded_feature_vector = fv.transform(feature_vectors_with_on_demand_features)
```
#### Compute one on-demand feature
On-demand transformation functions can also be accessed and executed as normal functions by using the dictionary `on_demand_transformations` that maps the on-demand features to their corresponding on-demand transformation function.
!!! example "Executing each on-demand transformation function"
=== "Python"
```python
# Specify request parameters for each serving key.
feature_vector = feature_view.get_feature_vector(
entry={"id": 1},
transform=False,
on_demand_features=False,
return_type="pandas",
)
# Applying model dependent transformations
feature_vector["on_demand_feature1"] = fv.on_demand_transformations[
"on_demand_feature1"
](feature_vector["transaction_time"], datetime.now())
```
## Chaining On-Demand Transformations
On-demand transformations attached to the same feature group can be chained: one transformation's output column can serve as another transformation's input.
The execution order is resolved automatically, and the resulting DAG is visible from the feature group overview page in the Hopsworks UI.
!!! example "On-demand transformation that consumes an upstream output"
=== "Python"
```python
from hopsworks import udf
@udf(int, drop=["raw"])
def add_one(raw):
return raw + 1
@udf(int, drop=["col"])
def double(col):
return col * 2
fg = fs.create_feature_group(
name="chained_odt_fg",
version=1,
primary_key=["id"],
transformation_functions=[
add_one("raw").alias("raw_plus_one"),
double("raw_plus_one").alias("raw_plus_one_doubled"),
],
)
```
Columns consumed only by the chain can be dropped, as the raw input `raw` and the intermediate `raw_plus_one` are in the example, leaving `raw_plus_one_doubled` as the only stored output.
The full chain still executes during online serving, and dropped columns never become stored features.
An on-demand transformation's output column becomes a regular feature in the feature group, which a downstream feature view can consume and pass into a model-dependent transformation.
This is the implicit chaining path between on-demand and model-dependent transformations, with no additional setup on either side.
================================================================================
# Online Ingestion Observability
Source: https://docs.hopsworks.ai/latest/user_guides/fs/feature_group/online_ingestion_observability/
# Online ingestion observability { #online-ingestion-observability }
## Introduction
Knowing when ingested data becomes available for online serving, and understanding the cause of any ingestion failures, is crucial for users.
To address this, the Hopsworks API provides observability features for online ingestion, allowing you to monitor ingestion status and troubleshoot issues.
This guide explains how to use these observability features for online feature groups in Hopsworks, with examples using both the Hopsworks APIs and the user interface.
## Prerequisites
Before you begin this guide we suggest you read the [Feature Group](../../../concepts/fs/feature_group/fg_overview.md) concept page to understand what a feature group is and how it fits in the ML pipeline.
## Using the Hopsworks API
### Create a Feature Group and Ingest Data
First, create an online-enabled feature group and insert data into it:
=== "Python"
```python
fg = fs.create_feature_group(
name="feature_group_name",
version=feature_group_version,
primary_key=feature_group_primary_keys,
online_enabled=True,
)
fg.insert(fg_df)
```
### Retrieve Online Ingestion Status
After inserting data, you can monitor the ingestion progress:
#### Get the latest ingestion instance
=== "Python"
```python
oi = fg.get_latest_online_ingestion()
```
#### Get a specific ingestion by its ID
=== "Python"
```python
oi = fg.get_online_ingestion(ingestion_id)
```
### Use the Online Ingestion Object
The online ingestion object provides methods to track and debug the ingestion process:
#### Wait for completion
Wait for the online ingestion to finish (equivalent to `fg.insert(fg_df, wait=True)`):
=== "Python"
```python
oi.wait_for_completion()
```
#### Print mini-batch results
Check the results of the ingestion.
If the status is `UPSERTED` and the number of rows matches your data, the ingestion was successful:
=== "Python"
```python
print([result.to_dict() for result in oi.results])
# Example output: [{'onlineIngestionId': 1, 'status': 'UPSERTED', 'rows': 10}]
```
#### Print ingestion service logs
Retrieve logs from the online ingestion service to diagnose any issues:
=== "Python"
```python
oi.print_logs(priority="error", size=5)
```
## Using the UI
### Viewing Online Ingestion Status
After inserting data into an online-enabled feature group, you can track the ingestion progress in the `Recent activities` section of the feature group in the Hopsworks UI.
================================================================================
# Time-To-Live (TTL)
Source: https://docs.hopsworks.ai/latest/user_guides/fs/feature_group/ttl/
## Feature Group TTL Usage Guide
Time To Live (TTL) is a feature that automatically expires data in feature groups after a specified time period.
This guide explains when and how to use TTL in your feature groups.
### Use Case: When to Use TTL
TTL is particularly useful for feature groups that contain time-sensitive data that becomes stale or irrelevant after a certain period.
Common use cases include:
- **Regulatory compliance**: Data that must be automatically purged after a retention period for privacy or compliance reasons (e.g., GDPR, HIPAA)
- **Cost optimization**: Reducing storage costs by automatically removing outdated data that is no longer needed for model inference
- **Data freshness**: Ensuring that only recent, relevant data is available for online serving, preventing models from using stale features
For example, if you're building a recommendation system, you might want user interaction features (like "items viewed in the last hour") to automatically expire after 1 hour, ensuring your model only uses current, relevant data.
---
## Getting Started
### Creating a Feature Group with TTL
When creating a new feature group, you can enable TTL by specifying the `ttl` parameter.
The TTL value determines how long data will remain in the feature group before being automatically expired.
The TTL is calculated based on the `event_time` column.
Data rows where `event_time` is older than the TTL period will be automatically removed.
```python
from datetime import datetime, timezone
import pandas as pd
# Assume you already have a feature store handle
# fs = ...
now = datetime.now(timezone.utc)
df = pd.DataFrame(
{
"id": [0, 1, 2],
"timestamp": [now, now, now],
"feature1": [10, 20, 30],
"feature2": ["a", "b", "c"],
}
)
# Create a feature group with TTL enabled (60 seconds)
fg = fs.create_feature_group(
name="fg_ttl_example",
version=1,
primary_key=["id"],
event_time="timestamp",
online_enabled=True,
ttl=60, # TTL in seconds - data will expire after 60 seconds
)
fg.insert(
df,
write_options={
"start_offline_materialization": False,
"wait_for_online_ingestion": True,
},
)
# After 60 seconds, reading online will return empty data
fg.read(online=True) # Returns empty DataFrame after TTL expires
```
For detailed API reference on all possible types of TTL values, see the [FeatureStore.create_feature_group API documentation][hsfs.feature_store.FeatureStore.create_feature_group].
---
## Managing TTL on Existing Feature Groups
### Updating the TTL Value
You can change the TTL value for an existing feature group at any time.
This is useful when you need to adjust the retention period based on changing requirements.
```python
# Get your existing feature group
fg = fs.get_feature_group(
name="fg_ttl_example",
version=1,
)
# Update TTL to a new value (120 seconds = 2 minutes)
fg.enable_ttl(ttl=120)
```
After updating the TTL, the new retention period will apply to all future data insertions and will affect when existing data expires.
---
### Disabling and Re-enabling TTL
You can temporarily disable TTL on a feature group if you need to retain data indefinitely, and then re-enable it later.
#### Disabling TTL
```python
# Disable TTL - data will no longer expire automatically
fg.disable_ttl()
```
#### Re-enabling TTL
When re-enabling TTL, you have two options:
1. **Re-enable with the previous TTL value**: If you don't specify a TTL value, the feature group will use the last TTL value that was set.
```python
# Re-enable TTL using the previous TTL value
fg.enable_ttl()
```
2. **Re-enable with a new TTL value**: Specify a new TTL value when re-enabling.
```python
# Re-enable TTL with a new value (90 seconds)
fg.enable_ttl(ttl=90)
```
**Important**: If TTL was never set on the feature group before, you must provide a TTL value when enabling it.
Otherwise, TTL cannot be enabled.
---
### Enabling TTL on an Existing Feature Group
If you created a feature group without TTL initially, you can enable it later:
```python
# Get an existing feature group that was created without TTL
fg = fs.get_feature_group(
name="fg_existing_no_ttl",
version=1,
)
# Enable TTL for the first time (60 seconds)
fg.enable_ttl(ttl=60)
```
Once enabled, TTL will apply to all data in the feature group based on the `event_time` column.
For detailed API reference on all possible types of TTL values and additional options, see the [FeatureGroup.enable_ttl API documentation][hsfs.feature_group.FeatureGroup.enable_ttl].
---
## Monitoring TTL Purging
Expired rows stop appearing in query results as soon as their TTL passes.
Deleting them from storage happens separately, in the background.
A purge worker inside each RonDB REST Server (RDRS) process walks every TTL-enabled online table one partition at a time, deleting a batch of expired rows on each pass.
Hopsworks reports what that worker is doing in two places.
### On the Feature Group Page
A **TTL purge** card appears on the feature group overview whenever the feature group is online enabled and has a TTL.
The summary row describes the table as a whole:
| Field | Meaning |
| --- | --- |
| Online table | The online table backing this feature group, as `database.table` |
| TTL | The retention period the purge worker read from the table's schema |
| Rows purged | Rows deleted from this table, summed over the nodes that answered |
| RDRS nodes | How many RonDB REST Server processes reported on this table |
One row follows per RDRS node, because each node runs its own worker over its own partitions:
| Field | Meaning |
| --- | --- |
| RDRS node | The node these counters came from |
| Rows purged | Rows this node deleted from the table |
| Partition | Where this node's cursor sits in the table's partition rotation |
| Batch size | Rows attempted per partition visit, which the worker adapts on its own |
| Last visited | When this node last visited the table |
| Process started | When this node's RDRS process last started, so you can tell how much history its row count covers |
A feature group created moments ago is not listed straight away.
RDRS discovers TTL-enabled tables on a periodic schema scan, so for the first few seconds the card reports that no purge worker is tracking the feature group yet.
It starts reporting counters on the next scan.
### Cluster-Wide
Administrators can see every RDRS node's purge worker under **Settings → TTL Purge**.
Each node reports a state:
| State | Meaning |
| --- | --- |
| `running` | Actively purging |
| `paused` | Healthy, but no TTL-enabled tables exist to work on |
| `disabled` | Purging is switched off by configuration |
| `outside window` | Outside the configured daily purge window |
| `stopped` | Not started yet |
| `error` | The worker hit an error, and the RDRS log has the detail |
Only `error` indicates a fault.
A cluster with no TTL-enabled feature groups sits in `paused`, which is the healthy idle state.
Alongside the state, each node reports its counters (tables tracked, rows purged, rounds completed), the configuration it is running with (batch size range, sleep interval), and when its process last started.
A restart count sits next to that timestamp, counting restarts of the container within its current pod; replacing the pod, as a redeploy does, starts a fresh count, so the start time is the figure to trust.
Both views poll every ten seconds, show when the next refresh is due, and offer a **Refresh** button for an immediate read.
### Reading the Numbers
A few properties of these counters are worth knowing before you draw conclusions from them.
**The numbers are per RDRS node, and a node is not a datanode.**
A node here is a RonDB REST Server process.
Scaling RDRS changes how many rows the views list; adding datanodes does not, and shows up instead as a larger partition count.
Each node keeps its counters in memory and starts again from zero when its process restarts, and nothing is persisted.
Because different nodes purge different partitions, their per-table numbers legitimately differ.
The cumulative row count is the only figure that is summed across nodes.
**A counter is only as old as the process reporting it.**
Nothing is persisted, so every figure on these views runs from the moment that node's RDRS process last started, which both views report as **Process started**.
Read a low row count against that time rather than on its own: a worker that has been up for a minute and one that has been quietly idle for a week look identical without it.
**A round that deleted nothing still counts as activity.**
The last round timestamp advances on every pass, including passes that found nothing to delete.
It tells you the worker is alive, not that rows were removed.
**Rows that are already expired when you insert them never reach the online store.**
Rows whose `event_time` is older than the TTL at insert time are filtered out before they are written, so they are never counted as purged.
To watch the purge worker at work, insert rows that expire after they are written:
```python
from datetime import datetime, timedelta, timezone
import pandas as pd
# Assume you already have a feature group with a TTL
# fg = ...
size = 100
now = datetime.now(timezone.utc)
df = pd.DataFrame(
{
"id": range(size),
# One second apart, so the rows come up for purging at a steady rate
# instead of the whole batch expiring at once.
"timestamp": pd.date_range(now, periods=size, freq=timedelta(seconds=1)),
"feature1": range(size),
}
)
fg.insert(df)
```
**A batch size sitting at its configured maximum means the worker is behind.**
The worker raises the batch size while there is a backlog and lowers it once it catches up, so a value pinned at the maximum shown on the cluster-wide page is the clearest sign that purging is not keeping up with expiry.
================================================================================
# Feature View User Guides
Source: https://docs.hopsworks.ai/latest/user_guides/fs/feature_view/
# Feature View User Guides
A feature view is a query over feature groups plus the metadata a model needs to read it consistently.
These guides cover creating one, reading training and inference data, and keeping the two aligned.
- :material-eye-plus-outline:{ .lg .middle } **Start here**
---
Select features from one or more feature groups and save the selection as a feature view.
```python
query = trans_fg.select_all().join(profile_fg.select(["age"]))
fv = fs.get_or_create_feature_view(
name="transactions_fraud",
version=1,
query=query,
labels=["fraud_label"],
)
```
[Create a feature view](overview.md) · [Training data](training-data.md) · [Feature vectors](feature-vectors.md)
:material-eye-plus-outline:{ .hops-role-ico } Create
{ .hops-role-cap }
- [Create a feature view](overview.md)
Select, join, filter and label, then save a version.
- [Query](query.md)
Joins, filters and point-in-time correctness.
- [Helper columns](helper-columns.md)
Columns for training or inference logic that are not model inputs.
- [Spines](spine-query.md)
Bring your own keys and labels at read time.
- [Model-dependent transformations](model-dependent-transformations.md)
Scaling and encoding fitted on training data, applied on read.
:material-database-export-outline:{ .hops-role-ico } Read
{ .hops-role-cap }
- [Training data](training-data.md)
Splits by ratio or time, materialised or in memory.
- [Batch data](batch-data.md)
Inference data for a time range, with transformations applied.
- [Feature vectors](feature-vectors.md)
Single or batched online lookups by serving key.
- [Feature server](feature-server.md)
Online lookups over REST, without the Python client.
:material-monitor-eye:{ .hops-role-ico } Observe
{ .hops-role-cap }
- [Feature monitoring](feature_monitoring.md)
Compare new data against a training dataset.
- [Feature logging](feature_logging.md)
Log the features a model actually saw at inference.
================================================================================
# Overview
Source: https://docs.hopsworks.ai/latest/user_guides/fs/feature_view/overview/
# Feature View
A feature view is a set of features that come from one or more feature groups.
It is a logical view over the feature groups, as the feature data is only stored in feature groups.
Feature views are used to read feature data for both training and serving (online and batch).
You can create [training datasets](training-data.md), create [batch data](batch-data.md) and get [feature vectors](feature-vectors.md).
If you want to understand more about the concept of feature view, you can refer to the [Feature View Overview](../../../concepts/fs/feature_view/fv_overview.md).
## Feature View Creation
[Query](./query.md) and [transformation function](./model-dependent-transformations.md) are the building blocks of a feature view.
You can define your set of features by building a `query`.
You can also define which columns in your feature view are the `labels`, which is useful for supervised machine learning tasks.
Furthermore, in python client, each feature can be attached to its own transformation function.
This way, when a feature is read (for training or scoring), the transformation is executed on-demand - just before the feature data is returned.
For example, when a client reads a numerical feature, the feature value could be normalized by a StandardScalar transformation function before it is returned to the client.
=== "Python"
```python
# create a simple feature view
feature_view = fs.create_feature_view(name="transactions_view", query=query)
# create a feature view with transformation and label
feature_view = fs.create_feature_view(
name="transactions_view",
query=query,
labels=["fraud_label"],
transformation_functions={
"amount": fs.get_transformation_function(
name="standard_scaler", version=1
)
},
)
```
=== "Java"
```java
// create a simple feature view
FeatureView featureView = featureStore.createFeatureView()
.name("transactions_view")
.query(query)
.build();
// create a feature view with label
FeatureView featureView = featureStore.createFeatureView()
.name("transactions_view")
.query(query)
.labels(Lists.newArrayList("fraud_label"))
.build();
```
You can refer to [query](./query.md) and [transformation function](./model-dependent-transformations.md) for creating `query` and `transformation_function`.
To see a full example of how to create a feature view, you can read [this notebook](https://github.com/logicalclocks/hopsworks-tutorials/blob/master/batch-ai-systems/fraud_batch/2_fraud_batch_training_pipeline.ipynb).
## Retrieval
Once you have created a feature view, you can retrieve it by its name and version.
=== "Python"
```python
feature_view = fs.get_feature_view(name="transactions_view", version=1)
```
=== "Java"
```java
FeatureView featureView = featureStore.getFeatureView("transactions_view", 1)
```
## Deletion
If there are some feature view instances which you do not use anymore, you can delete a feature view.
It is important to mention that all training datasets (include all materialised hopsfs training data) will be deleted along with the feature view.
=== "Python"
```python
feature_view.delete()
```
=== "Java"
```java
featureView.delete()
```
## Tags
Feature views also support tags.
You can attach, get, and remove tags.
You can learn more in [Tags Guide](../tags/tags.md).
=== "Python"
```python
# attach
feature_view.add_tag(name="tag_schema", value={"key": "value"})
# get
feature_view.get_tag(name="tag_schema")
# remove
feature_view.delete_tag(name="tag_schema")
```
=== "Java"
```java
// attach
Map tag = Maps.newHashMap();
tag.put("key", "value");
featureView.addTag("tag_schema", tag)
// get
featureView.getTag("tag_schema")
// remove
featureView.deleteTag("tag_schema")
```
## Next
Once you have created a feature view, you can now [create training data](./training-data.md)
================================================================================
# Training data
Source: https://docs.hopsworks.ai/latest/user_guides/fs/feature_view/training-data/
# Training data
Training data can be created from the feature view and used by different ML libraries for training different models.
You can read [training data concepts](../../../concepts/fs/feature_view/offline_api.md) for more details.
To see a full example of how to create training data, you can read [this notebook](https://github.com/logicalclocks/hopsworks-tutorials/blob/master/batch-ai-systems/fraud_batch/2_fraud_batch_training_pipeline.ipynb).
[](){ #arrowflight-server-with-duckdb }
Python clients read and create in-memory training data through the ArrowFlight Server with DuckDB, which Hopsworks enables by default.
For small and moderately sized datasets (what fits in a pandas DataFrame) it avoids the start-up cost of a Spark job; larger datasets can still be created with Spark by setting `read_options={"use_hive": True}`.
## Creation
It can be created as in-memory DataFrames or materialised as `tfrecords`, `parquet`, `csv`, or `tsv` files to HopsFS or in all other locations, for example, S3, GCS.
If you materialise a training dataset, a `PySparkJob` will be launched.
By default, `create_training_data` waits for the job to finish.
However, you can run the job asynchronously by passing `write_options={"wait_for_job": False}`.
You can monitor the job status in the [jobs overview UI](../../projects/jobs/pyspark_job.md#step-1-jobs-overview).
```python
# create a training dataset as dataframe
feature_df, label_df = feature_view.training_data(
description="transactions fraud batch training dataset",
)
# materialise a training dataset
version, job = feature_view.create_training_data(
description="transactions fraud batch training dataset",
data_format="csv",
write_options={"wait_for_job": False},
) # By default, it is materialised to HopsFS
print(job.id) # get the job's id and view the job status in the UI
```
!!! note "Growing a training dataset over time"
A materialized training dataset version cannot be appended to or modified in place.
To retrain on new data, create a new training dataset version.
If you need training data to keep growing, for example with a daily batch for a time-series model, do that computation once in a [derived feature group][assign-parents-to-a-feature-group] that is kept up to date as new data arrives, then create a new training dataset version from it whenever you need updated data.
### Extra filters {#training-data-extra-filters}
Sometimes data scientists need to train different models using subsets of a dataset.
For example, there can be different models for different countries, seasons, and different groups.
One way is to create different feature views for training different models.
Another way is to add extra filters on top of the feature view when creating training data.
In the [transaction fraud example](https://github.com/logicalclocks/hopsworks-tutorials/blob/master/batch-ai-systems/fraud_batch/1_fraud_batch_feature_pipeline.ipynb), there are different transaction categories, for example: "Health/Beauty", "Restaurant/Cafeteria", "Holliday/Travel" etc.
Examples below show how to create training data for different transaction categories.
```python
# Create a training dataset for Health/Beauty
df_health = feature_view.training_data(
description="transactions fraud batch training dataset for Health/Beauty",
extra_filter=trans_fg.category == "Health/Beauty",
)
# Create a training dataset for Restaurant/Cafeteria and Holliday/Travel
df_restaurant_travel = feature_view.training_data(
description="transactions fraud batch training dataset for Restaurant/Cafeteria and Holliday/Travel",
extra_filter=trans_fg.category == "Restaurant/Cafeteria"
and trans_fg.category == "Holliday/Travel",
)
```
### Lookback window for PIT joins {#training-data-lookback}
When training data is materialised from a Feature View that joins multiple Feature Groups, the PIT join scans every historical partition of the root and every joined Feature Group.
The `lookback` argument caps how far back the join is allowed to consider rows from the root and each joined Feature Group, so the engine can prune partitions before reading any files.
Apply the same window uniformly with `FeatureGroupLookback`, or use `Lookback` for per-Feature-Group control; the argument shape mirrors the one accepted by `get_batch_data` (see [the batch-data lookback section][batch-data-lookback]).
```python
import datetime
from hsfs.constructor.lookback import FeatureGroupLookback
version, job = feature_view.create_training_data(
start_time=datetime.date(2026, 5, 10),
end_time=datetime.date(2026, 5, 17),
description="fraud batch training data, weekly partition pruning",
lookback=FeatureGroupLookback(
key="PARTITION_KEY",
start=datetime.date(2026, 5, 10),
end=datetime.date(2026, 5, 17),
),
)
```
Equivalent dict form (no `FeatureGroupLookback` import, but `datetime` is still required for the bound values):
```python
import datetime
version, job = feature_view.create_training_data(
start_time=datetime.date(2026, 5, 10),
end_time=datetime.date(2026, 5, 17),
description="fraud batch training data, weekly partition pruning",
lookback={
"key": "PARTITION_KEY",
"start": datetime.date(2026, 5, 10),
"end": datetime.date(2026, 5, 17),
},
)
```
For different lookbacks per joined Feature Group, pass a `Lookback`. See the [per-feature-group lookback section][batch-data-lookback] of the batch-data guide for the full shape.
The resolved window is persisted with the training dataset, so re-reading the same training dataset version reconstructs the same per-join predicate.
The same parameter is accepted by `create_train_test_split` and `create_train_validation_test_split`.
### Train/Validation/Test Splits
In most cases, ML practitioners want to slice a dataset into multiple splits, most commonly train-test splits or train-validation-test splits, so that they can train and test their models.
Feature view provides a sklearn-like API for this purpose, so it is very easy to create a training dataset with different splits.
Create a training dataset (as in-memory DataFrames) or materialise a training dataset with train and test splits.
```python
# create a training dataset
X_train, X_test, y_train, y_test = feature_view.train_test_split(test_size=0.2)
# materialise a training dataset
version, job = feature_view.create_train_test_split(
test_size=0.2,
description="transactions fraud batch training dataset",
data_format="csv",
)
```
Create a training dataset (as in-memory DataFrames) or materialise a training dataset with train, validation, and test splits.
```python
# create a training dataset as DataFrame
X_train, X_val, X_test, y_train, y_val, y_test = (
feature_view.train_validation_test_split(
validation_size=0.3, test_size=0.2
)
)
# materialise a training dataset
version, job = feature_view.create_train_validation_test_split(
validation_size=0.3,
test_size=0.2,
description="transactions fraud batch training dataset",
data_format="csv",
)
```
To create a particular in-memory training dataset with Spark instead of the ArrowFlight Server with DuckDB, set `read_options={"use_hive": True}`.
```python
# create a training dataset as DataFrame with Hive
X_train, X_test, y_train, y_test = feature_view.train_test_split(
test_size=0.2, read_options={"use_hive": True}
)
```
## Read Training Data
Once you have created a training dataset, all its metadata are saved in Hopsworks.
This enables you to reproduce exactly the same dataset at a later point in time.
This holds for training data as both DataFrames or files.
That is, you can delete the training data files (for example, to reduce storage costs), but still reproduce the training data files later on if you need to.
```python
# get a training dataset
feature_df, label_df = feature_view.get_training_data(
training_dataset_version=1
)
# get a training dataset with train and test splits
X_train, X_test, y_train, y_test = feature_view.get_train_test_split(
training_dataset_version=1
)
# get a training dataset with train, validation and test splits
X_train, X_val, X_test, y_train, y_val, y_test = (
feature_view.get_train_validation_test_split(training_dataset_version=1)
)
```
## Passing Context Variables to Transformation Functions
Once you have [defined a transformation function using a context variable](../transformation_functions.md#passing-context-variables-to-transformation-function), you can pass the required context variables using the `transformation_context` parameter when generating IN-MEMORY training data or materializing a training dataset.
!!! note
Passing context variables for materializing a training dataset is only supported in the PySpark Kernel.
!!! example "Passing context variables while creating training data."
=== "Python"
```python
# Passing context variable to IN-MEMORY Training Dataset.
X_train, X_test, y_train, y_test = feature_view.get_train_test_split(
training_dataset_version=1,
primary_key=True,
event_time=True,
transformation_context={"context_parameter": 10},
)
# Passing context variable to Materialized Training Dataset.
version, job = feature_view.get_train_test_split(
training_dataset_version=1,
primary_key=True,
event_time=True,
transformation_context={"context_parameter": 10},
)
```
## Read training data with primary key(s) and event time
For certain use cases, e.g., time series models, the input data needs to be sorted according to the primary key(s) and event time combination.
Primary key(s) and event time are not usually included in the feature view query as they are not features used for training.
To retrieve the primary key(s) and/or event time when retrieving training data, you need to set the parameters `primary_key=True` and/or `event_time=True`.
```python
# get a training dataset
X_train, X_test, y_train, y_test = feature_view.get_train_test_split(
training_dataset_version=1,
primary_key=True,
event_time=True,
)
```
!!! note
All primary and event time columns of all the feature groups included in the feature view will be returned.
If they have the same names across feature groups and the join prefix was not provided then reading operation will fail with ambiguous column exception.
Make sure to define the join prefix if primary key and event time columns have the same names across feature groups.
To use primary key(s) and event time column with materialized training datasets it needs to be created with `primary_key=True` and/or `with_event_time=True`.
## Deletion
To clean up unused training data, you can delete all training data or for a particular version.
Note that all metadata of training data and materialised files stored in HopsFS will be deleted and cannot be recreated anymore.
```python
# delete a training data version
feature_view.delete_training_dataset(training_dataset_version=1)
# delete all training datasets
feature_view.delete_all_training_datasets()
```
It is also possible to keep the metadata and delete only the materialised files.
Then you can recreate the deleted files by just specifying a version, and you get back the exact same dataset again.
This is useful when you are running out of storage.
```python
# delete files of a training data version
feature_view.purge_training_data(training_dataset_version=1)
# delete files of all training datasets
feature_view.purge_all_training_data()
```
To recreate a training dataset:
```python
feature_view.recreate_training_dataset(training_dataset_version=1)
```
## Tags
Similar to feature view, You can attach, get, and remove tags.
You can learn more in [Tags Guide](../tags/tags.md).
```python
# attach
feature_view.add_training_dataset_tag(
training_dataset_version=1, name="tag_schema", value={"key": "value"}
)
# get
feature_view.get_training_dataset_tag(
training_dataset_version=1, name="tag_schema"
)
# remove
feature_view.delete_training_dataset_tag(
training_dataset_version=1, name="tag_schema"
)
```
## Next
Once you have created a training dataset and trained your model, you can deploy your model in a "batch" or "online" setting.
Next, you can learn how to create [batch data](./batch-data.md) and get [feature vectors](./feature-vectors.md).
================================================================================
# Batch data
Source: https://docs.hopsworks.ai/latest/user_guides/fs/feature_view/batch-data/
# Batch data (analytical ML systems)
## Creation
It is very common that ML models are deployed in a "batch" setting where ML pipelines score incoming new data at a regular interval, for example, daily or weekly.
Feature views support batch prediction by returning batch data as a DataFrame over a time range, by `start_time` and `end_time`.
The resultant DataFrame (or batch-scoring DataFrame) can then be fed to models to make predictions.
=== "Python"
```python
# get batch data
df = feature_view.get_batch_data(
start_time="20220620", end_time="20220627"
) # return a dataframe
```
=== "Java"
```java
Dataset ds = featureView.getBatchData("20220620", "20220627")
```
## Retrieve batch data with primary keys and event time
For certain use cases, e.g., time series models, the input data needs to be sorted according to the primary key(s) and event time combination.
Or one might want to merge predictions back with the original input data for postmortem analysis.
Primary key(s) and event time are not usually included in the feature view query as they are not features used for training.
To retrieve the primary key(s) and/or event time when retrieving batch data for inference, you need to set the parameters `primary_key=True` and/or `event_time=True`.
=== "Python"
```python
# get batch data
df = feature_view.get_batch_data(
start_time="20220620",
end_time="20220627",
primary_key=True,
event_time=True,
) # return a dataframe with primary keys and event time
```
!!! note
All primary and event time columns of all the feature groups included in the feature view will be returned.
If they have the same names across feature groups and the join prefix was not provided then reading operation will fail with ambiguous column exception.
Make sure to define the join prefix if primary key and event time columns have the same names across feature groups.
Python clients read batch data through the ArrowFlight Server with DuckDB, which Hopsworks enables by default and which is much faster than Spark for small and moderately sized data.
To read this particular batch data with Spark instead, set the read options to `{"use_hive": True}`.
```python
# get batch data with Hive
df = feature_view.get_batch_data(
start_time="20220620", end_time="20220627", read_options={"use_hive": True}
)
```
## Extra filters {#batch-data-extra-filters}
`get_batch_data` accepts an `extra_filter` argument that lets you apply an arbitrary filter on top of the Feature View's own query filter and any training-dataset filter inherited from `init_batch_scoring`.
Filters are combined with `AND`, pushed down to the storage layer, and apply equally to `get_batch_data`, `get_batch_query`, and `get_batch_query_string`.
The simplest form uses a Feature Group handle to build the predicate:
```python
df = feature_view.get_batch_data(
start_time="20220620",
end_time="20220627",
extra_filter=(trans_fg.category == "Health/Beauty"),
)
```
Combine multiple predicates with `&` (AND) and `|` (OR):
```python
df = feature_view.get_batch_data(
extra_filter=(trans_fg.amount > 100) & (trans_fg.country.isin(["SE", "NO"])),
)
```
The Feature View's own query filter and any training-dataset filter from `init_batch_scoring` are AND-combined with `extra_filter` before the read.
This is the same parameter that [training-data extra filters][training-data-extra-filters] exposes on training-data creation, so a filter expression works the same way in both APIs.
### Building filters from the Feature View alone
When you do not have the Feature Group handle in scope (for example when reading a Feature View in an inference script), use `feature_view.get_feature(name)` to obtain a `Feature` directly from the Feature View's query.
The returned `Feature` supports the same comparison operators (`==`, `!=`, `<`, `<=`, `>`, `>=`) and helper methods (`.like`, `.isin`, `.contains`), so it slots into `extra_filter` the same way:
```python
df = feature_view.get_batch_data(
extra_filter=(feature_view.get_feature("amount") > 100),
)
```
For Feature Views built from a join, `get_feature` accepts either the bare name or the prefixed name produced by the join.
Bare names resolve against the left Feature Group when more than one side has the column; the prefixed form forces resolution against the joined Feature Group:
```python
df = feature_view.get_batch_data(
# `category` exists on both sides of the join, `sec_` selects the joined FG.
extra_filter=(feature_view.get_feature("sec_category") == "A"),
)
```
If a bare name is ambiguous and no prefix is supplied, `get_feature` raises a `FeatureStoreException` listing the matching Feature Groups.
## Lookback window for PIT joins {#batch-data-lookback}
Point-in-time (PIT) joins use the condition `feature_fg.event_time <= root_fg.event_time` to pick the latest matching record from each joined Feature Group.
That predicate is a range comparison, not an equality, so partition pruning is defeated and every historical partition of every joined Feature Group is scanned on every read.
As Feature Groups grow with daily ingestion, this scan grows unboundedly.
The `lookback` argument lets you cap how far back the join is allowed to consider rows from each joined Feature Group.
Hopsworks turns the window into a constant-bound predicate on the joined Feature Group so the ArrowFlight Server with DuckDB and Spark Catalyst pushdown can prune partitions before opening any files.
### Uniform lookback
Apply the same window to every joined Feature Group with a `FeatureGroupLookback` instance from `hsfs.constructor.lookback`, or the equivalent dict.
Both forms accept `date` and `datetime` values.
```python
import datetime
from hsfs.constructor.lookback import FeatureGroupLookback
df = feature_view.get_batch_data(
start_time=datetime.date(2026, 5, 10),
end_time=datetime.date(2026, 5, 17),
lookback=FeatureGroupLookback(
key="PARTITION_KEY",
start=datetime.date(2026, 5, 10),
end=datetime.date(2026, 5, 17),
),
)
```
Equivalent dict form, no `FeatureGroupLookback` import required (you still need `datetime` for the bound values):
```python
import datetime
df = feature_view.get_batch_data(
start_time=datetime.date(2026, 5, 10),
end_time=datetime.date(2026, 5, 17),
lookback={
"key": "PARTITION_KEY",
"start": datetime.date(2026, 5, 10),
"end": datetime.date(2026, 5, 17),
},
)
```
`key` selects which column the predicate is emitted against.
`"PARTITION_KEY"` targets the Feature Group's partition column so the engine can prune partitions before reading files; the Feature Group must have a single DATE partition column.
`"EVENT_TIME"` targets the Feature Group's `event_time` column and guarantees row-level correctness but offers only engine-dependent file pruning (Hudi, Delta, or Iceberg column-stats indexing).
`start` is required and emits a `>=` predicate.
`end` is optional and emits a `<=` predicate when present.
When `end` is omitted, only the lower bound is emitted, making the short form below valid: the root Feature Group and every joined Feature Group get ` >= '2026-05-10'` (where `` is each Feature Group's own DATE partition column) and nothing else.
```python
import datetime
df = feature_view.get_batch_data(
lookback={
"key": "PARTITION_KEY",
"start": datetime.date(2026, 5, 10),
},
)
```
### Per-feature-group lookback
When different Feature Groups need different windows, use `Lookback` to bind a `FeatureGroupLookback` to specific joined Feature Groups.
An optional `default` applies to every Feature Group not listed in `feature_group_lookbacks`.
```python
import datetime
from hsfs.constructor.lookback import FeatureGroupLookback, Lookback
df = feature_view.get_batch_data(
start_time=datetime.date(2026, 5, 11),
end_time=datetime.date(2026, 5, 17),
lookback=Lookback(
default=FeatureGroupLookback(
key="PARTITION_KEY",
start=datetime.date(2026, 5, 5),
end=datetime.date(2026, 5, 17),
),
feature_group_lookbacks={
"transactions": FeatureGroupLookback(
key="EVENT_TIME",
start=datetime.datetime(2026, 5, 1, tzinfo=datetime.timezone.utc),
),
},
),
)
```
Skip the `default` to apply lookbacks only to the listed Feature Groups; unlisted Feature Groups receive no lookback for that call.
```python
df = feature_view.get_batch_data(
start_time=datetime.date(2026, 5, 11),
end_time=datetime.date(2026, 5, 17),
lookback=Lookback(
feature_group_lookbacks={
"transactions": FeatureGroupLookback(
key="PARTITION_KEY", start=datetime.date(2026, 5, 5)
),
}
),
)
```
`feature_group_lookbacks` keys identify a Feature Group in one of two ways: by name (a bare string matches every version of the named Feature Group at any join site in the Feature View) or by passing the Feature Group instance itself (matches the exact `(name, version)` so a specific version can be targeted when multiple versions of the same Feature Group are joined).
When both forms are supplied for the same name, the instance entry wins at its specific join site and the bare-string entry still applies elsewhere.
Equivalent dict form:
```python
import datetime
df = feature_view.get_batch_data(
start_time=datetime.date(2026, 5, 11),
end_time=datetime.date(2026, 5, 17),
lookback={
"default": {
"key": "PARTITION_KEY",
"start": datetime.date(2026, 5, 5),
"end": datetime.date(2026, 5, 17),
},
"feature_group_lookbacks": {
"transactions": {
"key": "EVENT_TIME",
"start": datetime.datetime(2026, 5, 1, tzinfo=datetime.timezone.utc),
},
},
},
)
```
### Combining `lookback` with other filters
The `lookback` predicate combines with filters declared on the Query, but where the filter is attached changes whether the engine can prune partitions on the root Feature Group.
Filters attached to a sub-query (`fg.select(...).filter(...)`) always prune on that Feature Group regardless of which Feature Group they reference.
Filters attached to the outer query (`query.filter(...)` after the join, or `extra_filter` on `get_batch_data`) prune the root only when every referenced feature belongs to the root Feature Group.
A mixed-Feature-Group outer filter still produces correct results, because the predicates apply at the outer level, but the root's partitions are no longer pruned at file-listing time.
```python
# Root sub-query filter: lookback prunes both root and joined Feature Groups.
query = root.select_all().filter(root.amount > 100).join(dim.select_all())
# Joined sub-query filter: lookback still prunes both sides.
query = root.select_all().join(dim.select_all().filter(dim.category == "X"))
# Outer filter referencing a joined Feature Group: root pruning is lost;
# joined Feature Groups still prune via their own predicates.
query = root.select_all().join(dim.select_all()).filter(dim.category == "X")
```
For best pruning, keep call-site filters at the sub-query level when their predicate references only one Feature Group.
The same `lookback` argument is supported on `create_training_data` (see [the training-data section][training-data-lookback]).
Both `extra_filter` and `lookback` can be combined.
## Creation with transformation
If you have specified transformation functions when creating a feature view, you will get back transformed batch data as well.
If your transformation functions require statistics of training dataset, you must also provide the training data version. `init_batch_scoring` will then fetch the statistics and initialize the functions with required statistics.
Then you can follow the above examples and create the batch data.
Please note that transformed batch data can only be returned in the python client but not in the java client.
```python
feature_view.init_batch_scoring(training_dataset_version=1)
```
It is important to note that in addition to the filters defined in Feature View, [extra filters][training-data-extra-filters] will be applied if they are defined in the given training dataset version.
## Retrieving untransformed batch data
By default, the `get_batch_data` function returns batch data with model-dependent transformations applied.
However, you can retrieve untransformed batch data, while still including on-demand features, by setting the `transform` parameter to `False`.
!!! example "Returning untransformed batch data"
=== "Python"
```python
# Fetching untransformed batch data.
untransformed_batch_data = feature_view.get_batch_data(transform=False)
```
## Passing Context Variables to Transformation Functions
After [defining a transformation function using a context variable](../transformation_functions.md#passing-context-variables-to-transformation-function), you can pass the necessary context variables through the `transformation_context` parameter when fetching batch data.
!!! example "Passing context variables while fetching batch data."
=== "Python"
```python
# Passing context variable to IN-MEMORY Training Dataset.
batch_data = feature_view.get_batch_data(
transformation_context={"context_parameter": 10}
)
```
================================================================================
# Feature vectors
Source: https://docs.hopsworks.ai/latest/user_guides/fs/feature_view/feature-vectors/
# Feature Vectors
The Hopsworks Platform integrates real-time capabilities with its Online Store.
Based on [RonDB](https://www.rondb.com/), your feature vectors are served at scale at in-memory latency (~1-10ms).
Checkout [the benchmarks results](https://www.hopsworks.ai/post/feature-store-benchmark-comparison-hopsworks-and-feast#images-2) and [the benchmark code](https://github.com/featurestoreorg/featurestore-benchmarks).
The same Feature View which was used to create training datasets can be used to retrieve feature vectors for real-time predictions.
This allows you to serve the same features to your model in training and serving, ensuring consistency and reducing boilerplate.
Whether you are either inside the Hopsworks platform, a model serving platform, or in an external environment, such as your application server.
Below is a practical guide on how to use the Online Store Python and Java Client.
The aim is to get you started quickly by providing code snippets which illustrate various use cases and functionalities of the clients.
If you need to get more familiar with the concept of feature vectors, you can read this [short introduction](../../../concepts/fs/feature_view/online_api.md) first.
## Retrieval
You can get back feature vectors from either python or java client by providing the primary key value(s) for the feature view.
Note that filters defined in feature view and training data will not be applied when feature vectors are returned.
If you need to retrieve a complete value of feature vectors without missing values, the required `entry` are [FeatureView.primary_keys][hsfs.feature_view.FeatureView.primary_keys].
Alternative, you can provide the primary key of the feature groups as the key of the entry.
It is also possible to provide a subset of the entry, which will be discussed [below](#partial-feature-retrieval).
=== "Python"
```python
# get a single vector
feature_view.get_feature_vector(entry={"pk1": 1, "pk2": 2})
# get multiple vectors
feature_view.get_feature_vectors(
entry=[{"pk1": 1, "pk2": 2}, {"pk1": 3, "pk2": 4}, {"pk1": 5, "pk2": 6}]
)
```
=== "Java"
```java
// get a single vector
Map entry1 = Maps.newHashMap();
entry1.put("pk1", 1);
entry1.put("pk2", 2);
featureView.getFeatureVector(entry1);
// get multiple vectors
Map entry2 = Maps.newHashMap();
entry2.put("pk1", 3);
entry2.put("pk2", 4);
featureView.getFeatureVectors(Lists.newArrayList(entry1, entry2));
```
### Required entry
Starting from python client v3.4, you can specify different values for the primary key of the same name which exists in multiple feature groups but are not joint by the same name.
The table below summarises the value of `primary_keys` in different settings.
Considering that you are joining 2 feature groups, namely, `left_fg` and `right_fg`, the feature groups have different primary keys, and features (`feature_*`) in each setting.
Also, the 2 feature groups are [joint][hsfs.constructor.query.Query.join] on different *join conditions* and *prefix* as `left_fg.join(right_fg, , prefix=)`.
For java client, and python client before v3.4, the `primary_keys` are the set of primary key of all the feature groups in the query.
Python client is backward compatible.
It means that the `primary_keys` used before v3.4 can be applied to python client of later versions as well.
The serving keys follow four rules, one per branch of the flow below:
- A `left_fg` primary key is always a serving key, under its own name.
- A `right_fg` primary key that the join matches to a `left_fg` primary key is covered by that key.
- A `right_fg` primary key the join does not match becomes a serving key under its own name, if that name is still free.
- If the name is already taken, the serving key is the join prefix plus the name, or `fgId___` plus the name when the join has no prefix.
`` is `right_fg.id` and `` is the position of the feature group in the join, 1 for the first join.
=== "As a flow"
--8<-- "user_guides/fs/feature_view/feature-vectors/serving-keys.html"
=== "As a table"
`id = user_id` stands for `left_on=["id"], right_on=["user_id"]`, and `id = id` for `on=["id"]`.
| `left_fg` keys | `right_fg` keys | join | prefix | serving keys |
| --- | --- | --- | --- | --- |
| id | id | `id = id` | | id |
| id1 | id2 | `id1 = id2` | | id1 |
| id1, id2 | id1 | `id1 = id1` | | id1, id2 |
| id, user_id | id | `user_id = id` | | id, user_id |
| id1 | id1, id2 | `id1 = id1` | | id1, id2 |
| id | id, user_id | `id = user_id` | `right_` | id, `right_id` |
| id | id, user_id | `id = user_id` | | id, `fgId___id` |
| id | id | `id = feature_1` | `right_` | id, `right_id` |
| id | id | `id = feature_1` | | id, `fgId___id` |
| id | id | `feature_1 = id` | `right_` | id, `right_id` |
| id | id | `feature_1 = id` | | id, `fgId___id` |
| user, year | user, year | `user = user` | `right_` | user, year, `right_year` |
| user, year | user, year | `user = user` | | user, year, `fgId___year` |
For example, joining two feature groups that both have `id` as primary key on `left_on=["id"], right_on=["user_id"]` with `prefix="right_"` gives the serving keys `id` and `right_id`:
```python
query = left_fg.select_all().join(
right_fg.select_all(), left_on=["id"], right_on=["user_id"], prefix="right_"
)
feature_view = fs.create_feature_view(name="fv", query=query)
feature_view.get_feature_vector({"id": 42, "right_id": 7})
```
### Missing Primary Key Entries
It can happen that some of the primary key entries are not available in some or all of the feature groups used by a feature view.
Take the above example assuming the feature view consists of two joined feature groups, first one with primary key column `pk1`, the second feature group with primary key column `pk2`.
=== "Python"
```python
# get a single vector
feature_view.get_feature_vector(entry={"pk1": 1, "pk2": 2})
```
=== "Java"
```java
// get a single vector
Map entry1 = Maps.newHashMap();
entry1.put("pk1", 1);
entry1.put("pk2", 2);
featureView.getFeatureVector(entry1);
```
This call will raise an exception if `pk1 = 1` OR `pk2 = 2` can't be found but also if `pk1 = 1` AND `pk2 = 2` can't be found, meaning, it will not return a partial or empty feature vector.
When retrieving a batch of vectors, the behaviour is slightly different.
=== "Python"
```python
# get multiple vectors
feature_view.get_feature_vectors(
entry=[{"pk1": 1, "pk2": 2}, {"pk1": 3, "pk2": 4}, {"pk1": 5, "pk2": 6}]
)
```
=== "Java"
```java
// get multiple vectors
Map entry2 = Maps.newHashMap();
entry2.put("pk1", 3);
entry2.put("pk2", 4);
Map entry3 = Maps.newHashMap();
entry3.put("pk1", 5);
entry3.put("pk2", 6);
featureView.getFeatureVectors(Lists.newArrayList(entry1, entry2, entry3));
```
This call will raise an exception if for example for the third entry `pk1 = 5` OR `pk2 = 6` can't be found, however, it will simply not return a vector for this entry if `pk1 = 5` AND `pk2 = 6`
can't be found.
That means, `get_feature_vectors` will never return partial feature vector, but will omit empty feature vectors.
If you are aware of missing features, you can use the [*passed features*](#passed-features) or [Partial feature retrieval](#partial-feature-retrieval) functionality, described down below.
### Partial feature retrieval
If your model can handle missing value or if you want to impute the missing value, you can get back feature vectors with partial values using python client starting from version 3.4 (Note that this does not apply to java client.).
In the example below, let's say you join 2 feature groups by `fg1.join(fg2, left_on=["pk1"], right_on=["pk2"])`, required keys of the `entry` are `pk1` and `pk2`.
If `pk2` is not provided, this returns feature values from the first feature group and null values from the second feature group when using the option `allow_missing=True`, otherwise it raises exception.
=== "Python"
```python
# get a single vector with
feature_view.get_feature_vector(entry={"pk1": 1}, allow_missing=True)
# get multiple vectors
feature_view.get_feature_vectors(
entry=[
{"pk1": 1},
{"pk1": 3},
],
allow_missing=True,
)
```
### Retrieval with transformation
If you have specified transformation functions when creating a feature view, you receive transformed feature vectors.
If your transformation functions require statistics of training dataset, you must also provide the training data version. `init_serving` will then fetch the statistics and initialize the functions with the required statistics.
Then you can follow the above examples and retrieve the feature vectors.
Please note that transformed feature vectors can only be returned in the python client but not in the java client.
=== "Python"
```python
feature_view.init_serving(training_dataset_version=1)
```
## Passed features
If some of the features values are only known at prediction time and cannot be computed and cached in the online feature store, you can provide those values as `passed_features` option.
The `get_feature_vector` method is going to use the passed values to construct the final feature vector to submit to the model.
You can use the `passed_features` parameter to overwrite individual features being retrieved from the online feature store.
The feature view will apply the necessary transformations to the passed features as it does for the feature data retrieved from the online feature store.
Please note that passed features is only available in the python client but not in the java client.
=== "Python"
```python
# get a single vector
feature_view.get_feature_vector(
entry={"pk1": 1, "pk2": 2}, passed_features={"feature_a": "value_a"}
)
# get multiple vectors
feature_view.get_feature_vectors(
entry=[{"pk1": 1, "pk2": 2}, {"pk1": 3, "pk2": 4}, {"pk1": 5, "pk2": 6}],
passed_features=[
{"feature_a": "value_a1"},
{"feature_a": "value_a2"},
{"feature_a": "value_a3"},
],
)
```
You can also use the parameter to provide values for all the features which are part of a specific feature group and used in the feature view.
In this second case, you do not have to provide the primary key value for that feature group as no data needs to be retrieved from the online feature store.
=== "Python"
```python
# get a single vector, replace values from an entire feature group
# note how in this example you don't have to provide the value of
# pk2, but you need to provide the features coming from that feature group
# in this case feature_b and feature_c
feature_view.get_feature_vector(
entry={"pk1": 1},
passed_features={
"feature_a": "value_a",
"feature_b": "value_b",
"feature_c": "value_c",
},
)
```
## Retrieving untransformed feature vectors
By default, the `get_feature_vector` and `get_feature_vectors` functions return transformed feature vectors, which has model-dependent transformations applied and includes on-demand features.
However, you can retrieve the untransformed feature vectors without applying model-dependent transformations while still including on-demand features by setting the `transform` parameter to False.
!!! example "Returning untransformed feature vectors"
=== "Python"
```python
# Fetching untransformed feature vector.
untransformed_feature_vector = feature_view.get_feature_vector(
entry={"id": 1}, transform=False
)
# Fetching untransformed feature vectors.
untransformed_feature_vectors = feature_view.get_feature_vectors(
entry=[{"id": 1}, {"id": 2}], transform=False
)
```
## Retrieving feature vector without on-demand features
The `get_feature_vector` and `get_feature_vectors` methods can also return untransformed feature vectors without on-demand features by disabling model-dependent transformations and excluding on-demand features.
To achieve this, set the parameters `transform` and `on_demand_features` to `False`.
!!! example "Returning untransformed feature vectors"
=== "Python"
```python
untransformed_feature_vector = feature_view.get_feature_vector(
entry={"id": 1}, transform=False, on_demand_features=False
)
untransformed_feature_vectors = feature_view.get_feature_vectors(
entry=[{"id": 1}, {"id": 2}], transform=False, on_demand_features=False
)
```
## Passing Context Variables to Transformation Functions
After [defining a transformation function using a context variable](../transformation_functions.md#passing-context-variables-to-transformation-function), you can pass the required context variables using the `transformation_context` parameter when fetching the feature vectors.
!!! example "Passing context variables while fetching batch data."
=== "Python"
```python
# Passing context variable to IN-MEMORY Training Dataset.
batch_data = feature_view.get_feature_vectors(
entry=[{"pk1": 1}], transformation_context={"context_parameter": 10}
)
```
## Retrieving feature vectors without blocking
`get_feature_vector` and `get_feature_vectors` block the calling thread for the whole round trip to the online store.
Inside a serving deployment, or anywhere else that runs an event loop, that stops every other request while the lookup is in flight.
`get_feature_vector_async` and `get_feature_vectors_async` take the same arguments and return the same values, awaited instead.
```python
vector = await my_feature_view.get_feature_vector_async(entry={"pk1": 1, "pk2": 2})
vectors = await my_feature_view.get_feature_vectors_async(
entry=[{"pk1": 1, "pk2": 2}, {"pk1": 3, "pk2": 4}]
)
```
The statements are awaited on the caller's own event loop, against a connection pool belonging to that loop, so several lookups are in flight at once.
On a measured deployment this raised throughput from 218 to 270 requests per second and cut p99 latency by 72 percent.
The awaited path applies to the SQL client.
A deployment reading through the REST client falls back to the blocking call, since there is nothing there to overlap.
Each event loop gets its own connection pool, and that pool is released when its loop is collected.
A process that creates a loop per lookup, for example by calling `asyncio.run` in a loop, therefore does not accumulate connections that way.
The default predictor a deployment gets from `model.deploy()` or `feature_view.deploy()` already awaits its lookup.
## Choose the right Client
The Online Store can be accessed via the **Python** or **Java** client allowing you to use your language of choice to connect to the Online Store.
Additionally, the Python client provides two different implementations to fetch data: **SQL** or **REST**.
The SQL client is the default implementation.
It requires a direct SQL connection to your RonDB cluster and uses python asyncio to offer high performance even when your Feature View rows involve querying multiple different tables.
The REST client is an alternative implementation connecting to [RonDB Feature Vector Server](./feature-server.md).
Perfect if you want to avoid exposing ports of your database cluster directly to clients.
This implementation is available as of Hopsworks 3.7.
Initialise the client by calling the `init_serving` method on the Feature View object before starting to fetch feature vectors.
This will initialise the chosen client, test the connection, and initialise the transformation functions registered with the Feature View.
Note to use the REST client in the Hopsworks Cluster python environment you will need to provide an API key explicitly as JWT authentication is not yet supported.
More configuration options can be found in the [API documentation][hsfs.feature_view.FeatureView.init_serving].
=== "Python"
```python
# initialize the SQL client to fetch feature vectors from the Online Store
my_feature_view.init_serving()
# or use the REST client
my_feature_view.init_serving(
init_rest_client=True,
config_rest_client={
"api_key": "your_api_key",
},
)
```
Once the client is initialised, you can start fetching feature vector(s) via the Feature View methods: `get_feature_vector(s)`.
You can initialise both clients for a given Feature View and switch between them by using the force flags in the get_feature_vector(s) methods.
=== "Python"
```python
# initialize both clients and set the default to REST
my_feature_view.init_serving(
init_rest_client=True,
init_sql_client=True,
config_rest_client={
"api_key": "your_api_key",
},
default_client="rest",
)
# this will fetch a feature vector via REST
try:
my_feature_view.get_feature_vector(
entry={"pk1": 1, "pk2": 2},
)
except TimeoutException:
# if the REST client times out, the SQL client will be used
my_feature_view.get_feature_vector(
entry={"pk1": 1, "pk2": 2}, force_sql=True
)
```
## Feature Server
In addition to Python/Java clients, from Hopsworks 3.3, a new [feature server](./feature-server.md) implemented in Go is introduced.
With this new API, single or batch feature vectors can be retrieved in any programming language.
Note that you can connect to the Feature Vector Server via any REST client.
However registered transformation function will not be applied to values in the JSON response and values stored in Feature Groups which contain embeddings will be missing.
================================================================================
# Feature server
Source: https://docs.hopsworks.ai/latest/user_guides/fs/feature_view/feature-server/
# Feature Store REST API Server
This API server allows users to retrieve single/batch feature vectors from a feature view.
## How to use
From Hopsworks 3.3, you can connect to the Feature Vector Server via any REST client which supports POST requests.
Set the `X-API-KEY` to your Hopsworks API Key and send the request with a JSON body, [single](#single-feature-vector-request) or [batch](#batch-feature-vectors-request).
By default, the server listens on the `0.0.0.0:4406` and the api version is set to `0.1.0`.
Please refer to `/srv/hops/mysql-cluster/rdrs_config.json` config file located on machines running the REST Server for additional configuration parameters.
In Hopsworks 3.7, we introduced a python client for the Online Store REST API Server.
The python client is available in the `hsfs` module and can be installed using `pip install hsfs`.
This client can be used instead of the Online Store SQL client in the `FeatureView.get_feature_vector(s)` methods.
Check the corresponding [documentation](./feature-vectors.md) for these methods.
## Single Feature Vector
### Single Feature Vector Request
`POST /{api-version}/feature_store`
#### Single Feature Vector Request Body
```json
{
"featureStoreName": "fsdb002",
"featureViewName": "sample_2",
"featureViewVersion": 1,
"passedFeatures": {},
"entries": {
"id1": 36
},
"metadataOptions": {
"featureName": true,
"featureType": true
},
"options": {
"validatePassedFeatures": true,
"includeDetailedStatus": true
}
}
```
#### Single Feature Vector Request Parameters
| **parameter** | **type** | **note** |
| --- | --- | --- |
| featureStoreName | string | |
| featureViewName | string | |
| featureViewVersion | number(int) | |
| entries | objects | Map of serving key of feature view as key and value of serving key as value. Serving key are a set of the primary key of feature groups which are included in the feature view query. If feature groups are joint with prefix, the primary key needs to be attached with prefix. |
| passedFeatures | objects | Optional. Map of feature name as key and feature value as value. This overwrites feature values in the response. |
| metadataOptions | objects | Optional. Map of metadataoption as key and boolean as value. Default metadata option is false. Metadata is returned on request. Metadata options available: 1\. featureName 2\. featureType |
| options | objects | Optional. Map of option as key and boolean as value. Default option is false. Options available: 1\. validatePassedFeatures 2\. includeDetailedStatus |
### Single Feature Vector Response
```json
{
"features": [
36,
"2022-01-24",
"int24",
"str14"
],
"metadata": [
{
"featureName": "id1",
"featureType": "bigint"
},
{
"featureName": "ts",
"featureType": "date"
},
{
"featureName": "data1",
"featureType": "string"
},
{
"featureName": "data2",
"featureType": "string"
}
],
"status": "COMPLETE",
"detailedStatus": [
{
"featureGroupId": 1,
"httpStatus": 200,
},
{
"featureGroupId": 2,
"httpStatus": 200,
},
]
}
```
### Single Feature Vector Errors
| **Code** | **reason** | **response** |
| -------- | ------------------------------------- | ------------------------------------ |
| 200 | | |
| 400 | Requested metadata does not exist | |
| 400 | Error in pk or passed feature value | |
| 401 | Access denied | Access unshared feature store failed |
| 500 | Failed to read feature store metadata | |
#### Response with PK/pass feature error
```json
{
"code": 12,
"message": "Wrong primay-key column. Column: ts",
"reason": "Incorrect primary key."
}
```
#### Response with metadata error
```json
{
"code": 2,
"message": "",
"reason": "Feature store does not exist."
}
```
#### PK value no match
```json
{
"features": [
9876543,
null,
null,
null
],
"metadata": null,
"status": "MISSING"
}
```
#### Detailed Status
If `includeDetailedStatus` option is set to true, detailed status is returned in the response.
Detailed status is a list of feature group id and http status code, corresponding to each read operations perform internally by RonDB.
Meaning is as follows:
- `featureGroupId`: Id of the feature group, used to identify which table the operation correspond from.
- `httpStatus`: Http status code of the operation.
- 200 means success
- 400 means bad request, likely pk name is wrong or pk is incomplete.
In particular, if pk for this table/feature group is not provided in the request, this http status is returned.
- 404 means no row corresponding to PK
- 500 means internal error.
Both `404` and `400` set the status to `MISSING` in the response.
Examples below corresponds respectively to missing row and bad request.
Missing Row: The PK name-value pair was correctly passed, but the corresponding row was not found in the feature group.
```json
{
"features": [
36,
"2022-01-24",
null,
null
],
"status": "MISSING",
"detailedStatus": [
{
"featureGroupId": 1,
"httpStatus": 200,
},
{
"featureGroupId": 2,
"httpStatus": 404,
},
]
}
```
Bad Request, e.g., when PK name-value pair for FG2 not provided or the corresponding column names was incorrect:
```json
{
"features": [
36,
"2022-01-24",
null,
null
],
"status": "MISSING",
"detailedStatus": [
{
"featureGroupId": 1,
"httpStatus": 200,
},
{
"featureGroupId": 2,
"httpStatus": 400,
},
]
}
```
## Batch Feature Vectors
### Batch Feature Vectors Request
`POST /{api-version}/batch_feature_store`
#### Batch Feature Vectors Request Body
```json
{
"featureStoreName": "fsdb002",
"featureViewName": "sample_2",
"featureViewVersion": 1,
"passedFeatures": [],
"entries": [
{
"id1": 16
},
{
"id1": 36
},
{
"id1": 71
},
{
"id1": 48
},
{
"id1": 29
}
],
"requestId": null,
"metadataOptions": {
"featureName": true,
"featureType": true
},
"options": {
"validatePassedFeatures": true,
"includeDetailedStatus": true
}
}
```
#### Batch Feature Vectors Request Parameters
| **parameter** | **type** | **note** |
| --- | --- | --- |
| featureStoreName | string | |
| featureViewName | string | |
| featureViewVersion | number(int) | |
| entries | `array` | Each items is a map of serving key as key and value of serving key as value. Serving key of feature view. |
| passedFeatures | `array` | Optional. Each items is a map of feature name as key and feature value as value. This overwrites feature values in the response. If provided, its size and order has to be equal to the size of entries. Item can be null. |
| metadataOptions | objects | Optional. Map of metadataoption as key and boolean as value. Default metadata option is false. Metadata is returned on request. Metadata options available: 1\. featureName 2\. featureType |
| options | objects | Optional. Map of option as key and boolean as value. Default option is false. Options available: 1\. validatePassedFeatures 2\. includeDetailedStatus |
### Batch Feature Vectors Response
```json
{
"features": [
[
16,
"2022-01-27",
"int31",
"str24"
],
[
36,
"2022-01-24",
"int24",
"str14"
],
[
71,
null,
null,
null
],
[
48,
"2022-01-26",
"int92",
"str31"
],
[
29,
"2022-01-03",
"int53",
"str91"
]
],
"metadata": [
{
"featureName": "id1",
"featureType": "bigint"
},
{
"featureName": "ts",
"featureType": "date"
},
{
"featureName": "data1",
"featureType": "string"
},
{
"featureName": "data2",
"featureType": "string"
}
],
"status": [
"COMPLETE",
"COMPLETE",
"MISSING",
"COMPLETE",
"COMPLETE"
],
"detailedStatus": [
[{
"featureGroupId": 1,
"httpStatus": 200,
}],
[{
"featureGroupId": 1,
"httpStatus": 200,
}],
[{
"featureGroupId": 1,
"httpStatus": 404,
}],
[{
"featureGroupId": 1,
"httpStatus": 200,
}],
[{
"featureGroupId": 1,
"httpStatus": 200,
}]
]
}
```
note: Order of the returned features are the same as the order of entries in the request.
### Batch Feature Vectors Errors
| **Code** | **reason** | **response** |
| -------- | ------------------------------------- | ------------------------------------ |
| 200 | | |
| 400 | Requested metadata does not exist | |
| 404 | Missing row corresponding to pk value | |
| 401 | Access denied | Access unshared feature store failed |
| 500 | Failed to read feature store metadata | |
#### Response with partial failure
```json
{
"features": [
[
81,
"id81",
"2022-01-29 00:00:00",
6
],
null,
[
51,
null,
null,
null,
]
],
"metadata": null,
"status": [
"COMPLETE",
"ERROR",
"MISSING"
],
"detailedStatus": [
[{
"featureGroupId": 1,
"httpStatus": 200,
}],
[{
"featureGroupId": 1,
"httpStatus": 400,
}],
[{
"featureGroupId": 1,
"httpStatus": 404,
}]
]
}
```
## Access control to feature store
Currently, the REST API server only supports Hopsworks API Keys for authentication and authorization.
Add the API key to the HTTP requests using the `X-API-KEY` header.
================================================================================
# Query
Source: https://docs.hopsworks.ai/latest/user_guides/fs/feature_view/query/
# Query vs DataFrame
Hopsworks provides a DataFrame API to ingest data into the Hopsworks Feature Store.
You can also retrieve feature data in a DataFrame, that can either be used directly to train models or [materialized to file(s)](./training-data.md) for later use to train models.
The idea of the Feature Store is to have pre-computed features available for both training and serving models.
The key functionality required to generate training datasets from reusable features are: feature selection, joins, filters, and point in time queries.
The Query object enables you to select features from different feature groups to join together to be used in a feature view.
The joining functionality is heavily inspired by the APIs used by Pandas to merge DataFrames.
The APIs allow you to specify which features to select from which feature group, how to join them and which features to use in join conditions.
=== "Python"
```python
fs = ...
credit_card_transactions_fg = fs.get_feature_group(name="credit_card_transactions", version=1)
account_details_fg = fs.get_feature_group(name="account_details", version=1)
merchant_details_fg = fs.get_feature_group(name="merchant_details", version=1)
# create a query
selected_features = credit_card_transactions_fg.select_all() \
.join(account_details_fg.select_all(), on=["cc_num"]) \
.join(merchant_details_fg.select_all())
# save the query to feature view
feature_view = fs.create_feature_view(
version=1,
name='credit_card_fraud',
labels=["is_fraud"],
query=selected_features
)
# retrieve the query back from the feature view
feature_view = fs.get_feature_view(“credit_card_fraud”, version=1)
query = feature_view.query
```
=== "Scala"
```scala
val fs = ...
val creditCardTransactionsFg = fs.getFeatureGroup("credit_card_transactions", 1)
val accountDetailsFg = fs.getFeatureGroup(name="account_details", version=1)
val merchantDetailsFg = fs.getFeatureGroup("merchant_details", 1)
// create a query
val selectedFeatures = (creditCardTransactionsFg.selectAll()
.join(accountDetailsFg.selectAll(), on=Seq("cc_num"))
.join(merchantDetailsFg.selectAll()))
val featureView = featureStore.createFeatureView()
.name("credit_card_fraud")
.query(selectedFeatures)
.build();
// retrieve the query back from the feature view
val featureView = fs.getFeatureView(“credit_card_fraud”, 1)
val query = featureView.getQuery()
```
If a data scientist wants to modify a new feature that is not available in the feature store, she can write code to compute the new feature (using existing features or external data) and ingest the new feature values into the feature store.
If the new feature is based solely on existing feature values in the Feature Store, we call it a derived feature.
The same Hopsworks APIs can be used to compute derived features as well as features using external data sources.
## The Query Abstraction
Most operations performed on `FeatureGroup` metadata objects will return a `Query` with the applied operation.
### Examples
Selecting features from a feature group is a lazy operation, returning a query with the selected features only:
=== "Python"
```python
credit_card_transactions_fg = fs.get_feature_group("credit_card_transactions")
# Returns Query
selected_features = credit_card_transactions_fg.select(
["amount", "latitude", "longitude"]
)
```
=== "Scala"
```scala
val creditCardTransactionsFg = fs.getFeatureGroup("credit_card_transactions")
# Returns Query
val selectedFeatures = creditCardTransactionsFg.select(Seq("amount", "latitude", "longitude"))
```
#### Join
Similarly, joins return query objects.
The simplest join in one where we join all of the features together from two different feature groups without specifying a join key - `Hopsworks` will infer the join key as a common primary key between the two feature groups.
By default, Hopsworks will use the maximal matching subset of the primary keys of the two feature groups as joining key(s), if not specified otherwise.
=== "Python"
```python
# Returns Query
selected_features = credit_card_transactions_fg.join(account_details_fg)
```
=== "Scala"
```scala
// Returns Query
val selectedFeatures = creditCardTransactionsFg.join(accountDetailsFg)
```
More complex joins are possible by selecting subsets of features from the joined feature groups and by specifying a join key and type.
Possible join types are "inner", "left" or "right".
By default`join_type` is `"left".
Furthermore, it is possible to specify different
features for the join key of the left and right feature group.
The join key lists should contain the names of the features to join on.
=== "Python"
```python
selected_features = (
credit_card_transactions_fg.select_all()
.join(account_details_fg.select_all(), on=["cc_num"])
.join(
merchant_details_fg.select_all(),
left_on=["merchant_id"],
right_on=["id"],
join_type="inner",
)
)
```
=== "Scala"
```scala
val selectedFeatures = (creditCardTransactionsFg.selectAll()
.join(accountDetailsFg.selectAll(), Seq("cc_num"))
.join(merchantDetailsFg.selectAll(), Seq("merchant_id"), Seq("id"), "inner"))
```
!!! warning
If there is feature name clash in the query then prefixes will be automatically generated and applied.
Generated prefix is feature group alias in the query (e.g., fg1, fg2).
Prefix is applied to the right feature group of the query.
### Data modeling in Hopsworks
Since v4.0 Hopsworks Feature selection API supports both Star and Snowflake Schema data models.
#### Star schema data model
When choosing Star Schema data model all tables are children of the parent (the left most) feature group, which has all
foreign keys for its child feature groups.
--8<-- "user_guides/fs/feature_view/query/star-schema.html"
=== "Python"
```python
selected_features = credit_card_transactions.select_all()
.join(aggregated_cc_transactions.select_all())
.join(account_details.select_all())
.join(merchant_details.select_all())
.join(cc_issuer_details.select_all())
```
In online inference, when you want to retrieve features in your online model, you have to provide all foreign key values,
known as the serving_keys, from the parent feature group to retrieve your precomputed feature values using the feature view.
=== "Python"
```python
feature vector = feature_view.get_feature_vector({
‘cc_num’: “1234 5555 3333 8888”,
‘issuer_id’: 20440455,
‘merchant_id’: 44208484,
‘account_id’: 84403331
})
```
#### Snowflake schema
Hopsworks also provides the possibility to define a feature view that consists of a nested tree of children (to up to a depth of 20) from the root (left most) feature group.
This is called Snowflake Schema data model where you need to build nested tables (subtrees) using joins, and then join the subtrees to their parents iteratively until you reach the root node (the leftmost feature group in the feature selection):
--8<-- "user_guides/fs/feature_view/query/snowflake-schema.html"
=== "Python"
```python
nested_selection = aggregated_cc_transactions.select_all()
.join(account_details.select_all())
.join(cc_issuer_details.select_all())
selected_features = credit_card_transactions.select_all()
.join(nested_selection)
.join(merchant_details.select_all())
```
Now, you have the benefit that in online inference you only need to pass two serving key values (the foreign keys of the leftmost feature group) to retrieve the precomputed features:
=== "Python"
```python
feature vector = feature_view.get_feature_vector({
‘cc_num’: “1234 5555 3333 8888”,
‘merchant_id’: 44208484,
})
```
#### Filter
In the same way as joins, applying filters to feature groups creates a query with the applied filter.
Filters are constructed with Python Operators `==`, `>=`, `<=`, `!=`, `>`, `<` and additionally with the methods `isin` and `like`.
Bitwise Operators `&` and `|` are used to construct conjunctions.
For the Scala part of the API, equivalent methods are available in the `Feature` and `Filter` classes.
=== "Python"
```python
filtered_credit_card_transactions = credit_card_transactions_fg.filter(
credit_card_transactions_fg.category == "Grocery"
)
```
=== "Scala"
```scala
val filteredCreditCardTransactions = creditCardTransactionsFg.filter(creditCardTransactionsFg.getFeature("category").eq("Grocery"))
```
Filters are fully compatible with joins:
=== "Python"
```python
selected_features = (
credit_card_transactions_fg.select_all()
.join(account_details_fg.select_all(), on=["cc_num"])
.join(
merchant_details_fg.select_all(),
left_on=["merchant_id"],
right_on=["id"],
)
.filter(
(credit_card_transactions_fg.category == "Grocery")
| (credit_card_transactions_fg.category == "Restaurant/Cafeteria")
)
)
```
=== "Scala"
```scala
val selectedFeatures = (creditCardTransactionsFg.selectAll()
.join(accountDetailsFg.selectAll(), Seq("cc_num"))
.join(merchantDetailsFg.selectAll(), Seq("merchant_id"), Seq("id"), "left")
.filter(creditCardTransactionsFg.getFeature("category").eq("Grocery").or(creditCardTransactionsFg.getFeature("category").eq("Restaurant/Cafeteria"))))
```
The filters can be applied at any point of the query:
=== "Python"
```python
selected_features = (
credit_card_transactions_fg.select_all()
.join(
accountDetails_fg.select_all().filter(
accountDetails_fg.avg_temp >= 22
),
on=["cc_num"],
)
.join(
merchant_details_fg.select_all(),
left_on=["merchant_id"],
right_on=["id"],
)
.filter(credit_card_transactions_fg.category == "Grocery")
)
```
=== "Scala"
```scala
val selectedFeatures = (creditCardTransactionsFg.selectAll()
.join(accountDetailsFg.selectAll().filter(accountDetailsFg.getFeature("avg_temp").ge(22)), Seq("cc_num"))
.join(merchantDetailsFg.selectAll(), Seq("merchant_id"), Seq("id"), "left")
.filter(creditCardTransactionsFg.getFeature("category").eq("Grocery")))
```
#### Joins and/or Filters on feature view query
The query retrieved from a feature view can be extended with new joins and/or new filters.
However, this operation will not update the metadata and persist the updated query of the feature view itself.
This query can then be used to create a new feature view.
=== "Python"
```python
fs = ...
merchant_details_fg = fs.get_feature_group(name="merchant_details", version=1)
credit_card_transactions_fg = fs.get_feature_group(name="credit_card_transactions", version=1)
feature_view = fs.get_feature_view(“credit_card_fraud”, version=1)
feature_view.query \
.join(merchant_details_fg.select_all()) \
.filter(credit_card_transactions_fg.category == "Cash Withdrawal")
```
=== "Scala"
```scala
val fs = ...
val merchantDetailsFg = fs.getFeatureGroup("merchant_details", 1)
val creditCardTransactionsFg = fs.getFeatureGroup("credit_card_transactions", 1)
val featureView = fs.getFeatureView(“credit_card_fraud”, 1)
featureView.getQuery()
.join(merchantDetailsFg.selectAll())
.filter(creditCardTransactionsFg.getFeature("category").eq("Cash Withdrawal"))
```
!!! warning
Every join/filter operation applied to an existing feature view query instance will update its state and accumulate.
To successfully apply new join/filter logic it is recommended to refresh the query instance by re-fetching the feature view:
=== "Python"
```python
fs = ...
merchant_details_fg = fs.get_feature_group(name="merchant_details", version=1)
account_details_fg = fs.get_feature_group(name="account_details", version=1)
credit_card_transactions_fg = fs.get_feature_group(name="credit_card_transactions", version=1)
# fetch new feature view and its query instance
feature_view = fs.get_feature_view(“credit_card_fraud”, version=1)
# apply join/filter logic based on purchase type
feature_view.query.join(merchant_details_fg.select_all()) \
.filter(credit_card_transactions_fg.category == "Cash Withdrawal")
# to apply new logic independent of purchase type from above
# re-fetch new feature view and its query instance
feature_view = fs.get_feature_view(“credit_card_fraud”, version=1)
# apply new join/filter logic based on account details
feature_view.query.join(merchant_details_fg.select_all()) \
.filter(account_details_fg.gender == "F")
```
=== "Scala"
```scala
fs = ...
merchantDetailsFg = fs.getFeatureGroup("merchant_details", 1)
accountDetailsFg = fs.getFeatureGroup("account_details", 1)
creditCardTransactionsFg = fs.getFeatureGroup("credit_card_transactions", 1)
// fetch new feature view and its query instance
val featureView = fs.getFeatureView(“credit_card_fraud”, version=1)
// apply join/filter logic based on purchase type
featureView.getQuery.join(merchantDetailsFg.selectAll())
.filter(creditCardTransactionsFg.getFeature("category").eq("Cash Withdrawal"))
// to apply new logic independent of purchase type from above
// re-fetch new feature view and its query instance
val featureView = fs.getFeatureView(“credit_card_fraud”, 1)
// apply new join/filter logic based on account details
featureView.getQuery.join(merchantDetailsFg.selectAll())
.filter(accountDetailsFg.getFeature("gender").eq("F"))
```
================================================================================
# Helper Columns
Source: https://docs.hopsworks.ai/latest/user_guides/fs/feature_view/helper-columns/
# Helper columns
Hopsworks Feature Store provides a functionality to define two types of helper columns `inference_helper_columns` and `training_helper_columns` for [feature views](./overview.md).
!!! note
Both inference and training helper column name(s) must be part of the `Query` object.
If helper column name(s) belong to feature group that is part of a `Join` with `prefix` defined, then this prefix needs to prepended
to the original column name when defining helper column list.
## Inference Helper columns
`inference_helper_columns` are a list of feature names that are not used for training the model itself but are used for extra information during online or batch inference.
For example, computing an [on-demand feature](../../../concepts/fs/feature_group/on_demand_feature.md) such as `days_valid` (days left that a credit card is valid at the time of the transaction)
in a credit card fraud detection system.
The feature `days_valid` will be computed using the credit card expiry date that needs to be fetched from the feature store and compared to the transaction
date that the transaction is performed on (`days_valid` = `expiry_date` - `current_date`).
In this use case `expiry_date` is an inference helper column.
It is not used for training but is necessary
for computing the [on-demand feature](../../../concepts/fs/feature_group/on_demand_feature.md)`days_valid` feature.
!!! example "Define inference columns for feature views."
=== "Python"
```python
# define query object
query = label_fg.select("fraud_label").join(
trans_fg.select(["amount", "days_valid", "expiry_date", "category"])
)
# define feature view with helper columns
feature_view = fs.get_or_create_feature_view(
name="fv_with_helper_col",
version=1,
query=query,
labels=["fraud_label"],
transformation_functions=transformation_functions,
inference_helper_columns=["expiry_date"],
)
```
### Inference Data Retrieval
When retrieving data for model inference, helper columns will be omitted.
However, they can be optionally fetched with inference or training data.
#### Batch inference
!!! example "Fetch inference helper column values and compute on-demand features during batch inference."
=== "Python"
```python
# import feature functions
from feature_functions import time_delta
# Fetch feature view object
feature_view = fs.get_feature_view(
name="fv_with_helper_col",
version=1,
)
# Fetch feature data for batch inference with helper columns
df = feature_view.get_batch_data(
start_time=start_time,
end_time=end_time,
inference_helpers=True,
event_time=True,
)
# compute location delta
df["days_valid"] = df.apply(
lambda row: time_delta(row["expiry_date"], row["transaction_date"]), axis=1
)
# prepare datatame for prediction
df = df[
[
f.name
for f in feature_view.features
if not (
f.label or f.inference_helper_column or f.training_helper_column
)
]
]
```
#### Online inference
!!! example "Fetch inference helper column values and compute on-demand features during online inference."
=== "Python"
```python
from feature_functions import time_delta
# Fetch feature view object
feature_view = fs.get_feature_view(
name="fv_with_helper_col",
version=1,
)
# Fetch feature data for batch inference without helper columns
df_without_inference_helpers = feature_view.get_batch_data()
# Fetch feature data for batch inference with helper columns
df_with_inference_helpers = feature_view.get_batch_data(inference_helpers=True)
# here cc_num, longitude and latitude are provided as parameters to the application
cc_num = ...
transaction_date = ...
# get previous transaction location of this credit card
inference_helper = feature_view.get_inference_helper(
{"cc_num": cc_num}, return_type="dict"
)
# compute location delta
days_valid = time_delta(transaction_date, inference_helper["expiry_date"])
# Now get assembled feature vector for prediction
feature_vector = feature_view.get_feature_vector(
{"cc_num": cc_num},
passed_features={"days_valid": days_valid},
)
```
## Training Helper columns
`training_helper_columns` are a list of feature names that are not the part of the model schema itself but are used during training for the extra information.
For example one might want to use feature like `category` of the purchased product to assign different weights.
!!! example "Define training helper columns for feature views."
=== "Python"
```python
# define query object
query = label_fg.select("fraud_label").join(
trans_fg.select(["amount", "days_valid", "expiry_date", "category"])
)
# define feature view with helper columns
feature_view = fs.get_or_create_feature_view(
name="fv_with_helper_col",
version=1,
query=query,
labels=["fraud_label"],
transformation_functions=transformation_functions,
training_helper_columns=["category"],
)
```
### Training Data Retrieval
When retrieving training data helper columns will be omitted.
However, they can be optionally fetched.
!!! example "Fetch training data with or without inference helper column values."
=== "Python"
```python
# import feature functions
from feature_functions import location_delta, time_delta
# Fetch feature view object
feature_view = fs.get_feature_view(
name="fv_with_helper_col",
version=1,
)
# Create and training data with training helper columns
TEST_SIZE = 0.2
X_train, X_test, y_train, y_test = feature_view.train_test_split(
description="transactions fraud training dataset",
test_size=TEST_SIZE,
training_helper_columns=True,
)
# Get existing training data with training helper columns
X_train, X_test, y_train, y_test = feature_view.get_train_test_split(
training_dataset_version=1, training_helper_columns=True
)
```
!!! note
To use helper columns with materialized training dataset it needs to be created with `training_helper_columns=True`.
================================================================================
# Model-Dependent Transformation Functions
Source: https://docs.hopsworks.ai/latest/user_guides/fs/feature_view/model-dependent-transformations/
# Model Dependent Transformation Functions
[Model-dependent transformations](https://www.hopsworks.ai/dictionary/model-dependent-transformations) transform feature data for a specific model.
Feature encoding is one example of such a transformations.
Feature encoding is parameterized by statistics from the training dataset, and, as such, many model-dependent transformations require the training dataset statistics as a parameter.
Hopsworks enhances the robustness of AI pipelines by preventing [training-inference skew](https://www.hopsworks.ai/dictionary/training-inference-skew) by ensuring that the same model-dependent transformations and statistical parameters are used during both training dataset generation and online inference.
Additionally, Hopsworks offers built-in model-dependent transformation functions, such as `min_max_scaler`, `standard_scaler`, `robust_scaler`, `label_encoder`, and `one_hot_encoder`, which can be easily imported and declaratively applied to features in a feature view.
## Model Dependent Transformation Function Creation
Hopsworks allows you to create a model-dependent transformation function by attaching a [transformation function](../transformation_functions.md) to a feature view.
The attached transformation function can be a simple function that takes one feature as input and outputs the transformed feature data.
For example, in the case of min-max scaling a numerical feature, you will have a number as input parameter to the transformation function and a number as output.
However, in the case of one-hot encoding a categorical variable, you will have a string as input and an array of 1s and 0s and output.
You can also have transformation functions that take multiple features as input and produce one or more values as output.
That is, transformation functions can be one-to-one, one-to-many, many-to-one, or many-to-many.
Each model-dependent transformation function can map specific features to its arguments by explicitly providing their names as arguments to the transformation function.
If no feature names are provided, the transformation function will default to using features from the feature view that match the name of the transformation function's argument.
Hopsworks by default generates default names of transformed features output by a model-dependent transformation function.
The generated names follows a naming convention structured as `functionName_features_outputColumnNumber` if the transformation function outputs multiple columns and `functionName_features` if the transformation function outputs one column.
For instance, for the function named `add_one_multiple` that outputs multiple columns in the example given below, produces output columns that would be labeled as `add_one_multiple_feature1_feature2_feature3_0`, `add_one_multiple_feature1_feature2_feature3_1` and `add_one_multiple_feature1_feature2_feature3_2`.
The function named `add_two` that outputs a single column in the example given below, produces a single output column names as `add_two_feature`.
Additionally, Hopsworks also allows users to specify custom names for transformed feature using the [`alias`](../transformation_functions.md#specifying-output-features-names-for-transformation-functions) function.
!!! example "Creating model-dependent transformation functions"
=== "Python"
```python
# Defining a many to many transformation function.
@udf(return_type=[int, int, int], drop=["feature1", "feature3"])
def add_one_multiple(feature1, feature2, feature3):
return pd.DataFrame(
{
"add_one_feature1": feature1 + 1,
"add_one_feature2": feature2 + 1,
"add_one_feature3": feature3 + 1,
}
)
# Defining a one to one transformation function.
@udf(return_type=int)
def add_two(feature):
return feature + 2
# Creating model-dependent transformations by attaching transformation functions to feature views.
feature_view = fs.create_feature_view(
name="transactions_view",
query=query,
labels=["fraud_label"],
transformation_functions=[add_two, add_one_multiple],
)
```
### Specifying input features
The features to be used by a model-dependent transformation function can be specified by providing the feature names (from the feature view / feature group) as input to the transformation functions.
!!! example "Specifying input features to be passed to a model-dependent transformation function"
=== "Python"
```python
feature_view = fs.create_feature_view(
name="transactions_view",
query=query,
labels=["fraud_label"],
transformation_functions=[
add_two("feature_1"),
add_two("feature_2"),
add_one_multiple("feature_5", "feature_6", "feature_7"),
],
)
```
### Using built-in transformations
Built-in transformation functions are attached in the same way.
The only difference is that they can either be retrieved from the Hopsworks or imported from the `hopsworks` module.
!!! example "Creating model-dependent transformation using built-in transformation functions retrieved from Hopsworks"
=== "Python"
```python
min_max_scaler = fs.get_transformation_function(name="min_max_scaler")
standard_scaler = fs.get_transformation_function(name="standard_scaler")
robust_scaler = fs.get_transformation_function(name="robust_scaler")
label_encoder = fs.get_transformation_function(name="label_encoder")
feature_view = fs.create_feature_view(
name="transactions_view",
query=query,
labels=["fraud_label"],
transformation_functions=[
label_encoder("category"),
robust_scaler("amount"),
min_max_scaler("loc_delta"),
standard_scaler("age_at_transaction"),
],
)
```
To attach built-in transformation functions from the `hopsworks` module they can be directly imported into the code from `hopsworks.builtin_transformations`.
!!! example "Creating model-dependent transformation using built-in transformation functions imported from hopsworks"
=== "Python"
```python
from hopsworks.hsfs.builtin_transformations import (
label_encoder,
min_max_scaler,
robust_scaler,
standard_scaler,
)
feature_view = fs.create_feature_view(
name="transactions_view",
query=query,
labels=["fraud_label"],
transformation_functions=[
label_encoder("category"),
robust_scaler("amount"),
min_max_scaler("loc_delta"),
standard_scaler("age_at_transaction"),
],
)
```
## Using Model Dependent Transformations
Model-dependent transformations attached to a feature view are automatically applied when you [create training data](./training-data.md#creation), [read training data](./training-data.md#read-training-data), [read batch inference data](./batch-data.md#creation-with-transformation), or [get feature vectors](./feature-vectors.md#retrieval-with-transformation).
The generated data includes untransformed features, on-demand features, if any, and the transformed features.
The transformed features are organized by their output column names in alphabetical order and are positioned after the untransformed and on-demand features.
Model-dependent transformation functions can also be manually applied to a feature vector using the `transform` function.
!!! example "Manually applying model-dependent transformations during online inference"
=== "Python"
```python
# Initialize the feature view with the correct training dataset version used for model-dependent transformations
fv.init_serving(training_dataset_version)
# Get untransformed feature Vector
feature_vector = fv.get_feature_vector(
entry={"index": 10}, transform=False, return_type="pandas"
)
# Apply Model Dependent transformations
encoded_feature_vector = fv.transform(feature_vector)
```
### Retrieving untransformed feature vector and batch inference data
The `get_feature_vector`, `get_feature_vectors`, and `get_batch_data` methods can return untransformed feature vectors and batch data without applying model-dependent transformations while still including on-demand features.
To achieve this, set the `transform` parameter to False.
!!! example "Returning untransformed feature vectors and batch data."
=== "Python"
```python
# Fetching untransformed feature vector.
untransformed_feature_vector = feature_view.get_feature_vector(
entry={"id": 1}, transform=False
)
# Fetching untransformed feature vectors.
untransformed_feature_vectors = feature_view.get_feature_vectors(
entry=[{"id": 1}, {"id": 2}], transform=False
)
# Fetching untransformed batch data.
untransformed_batch_data = feature_view.get_batch_data(transform=False)
```
## Chaining Model-Dependent Transformations
A model-dependent transformation (MDT) can consume another MDT's output as its input.
The DAG is resolved automatically at execution time, so producers always run before consumers.
!!! example "Chaining two increments and a sum"
=== "Python"
```python
from hopsworks import udf
@udf(int)
def add_one(col):
return col + 1
@udf(int)
def add(a, b):
return a + b
fv = fs.create_feature_view(
name="chained_mdt_fv",
query=fg.select_all(),
transformation_functions=[
add_one("data1").alias("data1_plus_one"),
add_one("data2").alias("data2_plus_one"),
add("data1_plus_one", "data2_plus_one").alias("sum_plus_two"),
],
version=1,
)
```
### Statistics over chained transformations
Statistics-based transformations participate in chains like any other transformation.
A transformation that requires statistics on another transformation's output, such as a min-max scaler applied to an imputed column, is fit on that intermediate output rather than on the raw feature.
During training dataset creation the statistics are computed in dependency order on the train split, each transformation executes exactly once, and the fitted statistics are persisted so that online serving applies the same values.
See [Transformation Functions Performance Tuning][transformation-functions-performance-tuning] for `n_processes` semantics on chained DAGs.
================================================================================
# Spines
Source: https://docs.hopsworks.ai/latest/user_guides/fs/feature_view/spine-query/
# Using Spines
In this section we will illustrate how to use a [Spine Group](../../../concepts/fs/feature_group/spine_group.md) instead of a regular Feature Group for performing
point-in-time joins when reading batch data for inference or when creating training datasets.
## Prerequisites
1. Make sure you have read the [concept section about spines](../../../concepts/fs/feature_group/spine_group.md) in feature and inference pipelines.
2. Make sure you have gone through the [Spine Group creation guide](../feature_group/create_spine.md).
3. Make sure you understand the [concept of feature views](../../../concepts/fs/feature_view/fv_overview.md) and how to create them using the [query abstraction](../feature_view/query.md)
## Feature View with a Spine Group
### Step 1: Query Definition
The first step before creating a Feature View, is to construct the query by selecting the label and features which are needed:
```python
# Select features for training data.
ds_query = trans_fg.select(["fraud_label"]).join(
window_aggs_fg.select_except(["cc_num"]), on="cc_num"
)
ds_query.show(5)
```
Similarly you can construct the query using a previously created spine equivalent.
However, there are two thing to note:
1. **If you want to use the query for a feature view to be used for online serving, you can only select the "label" or target feature from the spine.**
2. **Spine groups can only be used on the left side of the join.** Think of the left side of the join as the base set of entities that should be included in you batch of data or training dataset, which we enrich with the relevant and point-in-time correct feature values.
```python
trans_spine = fs.get_or_create_spine_group(
name="spine_transactions",
version=1,
description="Transaction data",
primary_key=["cc_num"],
event_time="datetime",
dataframe=trans_df,
)
# Select features for training data.
ds_query_spine = trans_spine.select(["fraud_label"]).join(
window_aggs_fg.select_except(["cc_num"]), on="cc_num"
)
```
Calling the `show()` or `read()` method of this query object will use the spine dataframe included in the Spine Group object to perform the join.
```python
ds_query_spine.show(10)
```
### Step 2: Feature View Creation
With the above defined query, we can continue to create the Feature View in the same way we would do it also without a spine:
```python
feature_view_spine = fs.get_or_create_feature_view(
name="transactions_view_spine",
query=ds_query_spine,
version=1,
labels=["fraud_label"],
)
```
### Step 3: Training Dataset Creation
With the regular feature view, the labels are fetched from the feature store, but with the feature view created with a spine, you need to provide the dataframe.
Here you have the chance to pass a different set of entities to generate the training dataset.
```python
X_train, X_test, y_train, y_test = feature_view_spine.train_test_split(
0.2, spine=new_entities_df
)
X_train.show()
```
### Step 4: Retrieving New Batches Inference Data
You can now use the offline and online API of the feature stores to read features for inference.
Similarly to training dataset creation, every time you read up a new batch of data, you can pass a different spine dataframe.
```python
feature_view_spine.get_batch_data(spine=scoring_spine_df).show()
```
### Step 5: Online Feature Lookup
For the online lookup, the label is not required, therefore it was important to only select label from the left spine group, so that we don't need to provide a spine for online serving:
```python
# Note: no spine needs to be passed
feature_view.get_feature_vector({"cc_num": 4473593503484549})
```
## Replacing a Regular Feature Group with a Spine at Serving Time
In the case where you create a feature view with a regular feature group, but you would like to retrieve batch inference data using IDs (primary key values), you can use a spine to replace the left feature group.
To do this, you can pass the Spine Group instead of a dataframe.
```python
# Note: here feature_view was created with regular feature groups only
# and trans_spine is of type SpineGroup instead of a dataframe
feature_view.get_batch_data(spine=trans_spine).show()
```
================================================================================
# Feature Monitoring
Source: https://docs.hopsworks.ai/latest/user_guides/fs/feature_view/feature_monitoring/
# Feature Monitoring for Feature Views
Feature Monitoring complements the Hopsworks data validation capabilities for Feature Group data by allowing you to monitor your data once they have been ingested into the Feature Store.
Hopsworks feature monitoring is centered around two functionalities: **scheduled statistics** and **statistics comparison**.
Before continuing with this guide, see the [Feature monitoring guide](../feature_monitoring/index.md) to learn more about how feature monitoring works, and get familiar with the different use cases of feature monitoring for Feature Views described in the **Use cases** sections of the [Scheduled statistics guide](../feature_monitoring/scheduled_statistics.md#use-cases) and [Statistics comparison guide](../feature_monitoring/statistics_comparison.md#use-cases).
!!! warning "Limited UI support"
Currently, feature monitoring can only be configured using the [Hopsworks Python library](https://pypi.org/project/hopsworks).
However, you can enable/disable a feature monitoring configuration or trigger the statistics comparison manually from the UI.
## Code
In this section, we show you how to set up feature monitoring on a Feature View using the ==Hopsworks Python library==.
Alternatively, you can get started quickly by running our [tutorial for feature monitoring](https://github.com/logicalclocks/hopsworks-tutorials/blob/master/api_examples/feature_monitoring.ipynb).
!!! info "Prerequisites"
- A Hopsworks project.
If you don't have one yet, go to [run.hopsworks.ai](https://run.hopsworks.ai), sign up with your email and create your first project.
- An API key, which you can get from "Account Settings" on [run.hopsworks.ai](https://run.hopsworks.ai).
- The [Hopsworks Python library](https://pypi.org/project/hopsworks) installed in your client.
See the [installation guide](../../client_installation/index.md).
- A Feature View and a Training Dataset.
### Step 1: Connect to Hopsworks
Connect the client running your notebook to Hopsworks.
You will be prompted to paste your API key to connect the notebook to your project.
=== "Python"
```python
import hopsworks
project = hopsworks.login()
fs = project.get_feature_store()
```
See the API reference for [`hopsworks.login`][hopsworks.login] and [`Project.get_feature_store`][hopsworks_common.project.Project.get_feature_store].
### Step 2: Get or create a Feature View
Feature monitoring can be enabled on already created Feature Views.
We suggest you read the [Feature View](../../../concepts/fs/feature_view/fv_overview.md) concept page and familiarize yourself with the APIs to [create a feature view](overview.md) using the [query abstraction](query.md).
=== "Python"
```python
# Retrieve an existing feature view
trans_fv = fs.get_feature_view("trans_fv", version=1)
# Or, create a new feature view
query = trans_fg.select(["fraud_label", "amount", "cc_num"])
trans_fv = fs.create_feature_view(
name="trans_fv",
version=1,
query=query,
labels=["fraud_label"],
)
```
See the API reference for [`FeatureStore.get_feature_view`][hsfs.feature_store.FeatureStore.get_feature_view] and [`FeatureStore.create_feature_view`][hsfs.feature_store.FeatureStore.create_feature_view].
### Step 3: Get or create a Training Dataset
A Training Dataset can be used later as a reference window to compare against (see Step 6).
=== "Python"
```python
# Create a training dataset with train and test splits
_, _ = trans_fv.create_train_validation_test_split(
description="transactions fraud batch training dataset",
data_format="csv",
validation_size=0.2,
test_size=0.1,
)
```
See the API reference for [`FeatureView.create_train_validation_test_split`][hsfs.feature_view.FeatureView.create_train_validation_test_split].
### Step 4: Create a monitoring configuration
Start a new configuration on the Feature View.
Use `create_scheduled_statistics` to only compute statistics on a schedule, or `create_feature_monitoring` to also compare them against a reference.
=== "Scheduled statistics"
```python
# compute statistics on one or more features on a schedule
fm_monitoring_config = trans_fv.create_scheduled_statistics(
name="trans_fv_all_features_monitoring",
description="Compute statistics on the Feature View data on a daily basis",
feature_names=["amount"], # omit to monitor all features
)
```
=== "Statistics comparison"
```python
# the feature to compare is selected later in
# compare_on / compare_on_distribution (Step 7.A / 7.B), not here
fm_monitoring_config = trans_fv.create_feature_monitoring(
name="trans_fv_amount_monitoring",
description="Compute and compare descriptive statistics on the Feature View data on a daily basis",
)
```
See the API reference for [`FeatureView.create_scheduled_statistics`][hsfs.feature_view.FeatureView.create_scheduled_statistics] and [`FeatureView.create_feature_monitoring`][hsfs.feature_view.FeatureView.create_feature_monitoring].
!!! info "Custom schedule"
By default, the computation of statistics is scheduled to run endlessly, every day at 12PM.
You can modify the default schedule by adjusting the `cron_expression`, `start_date_time` and `end_date_time` parameters (e.g., `cron_expression="0 0 12 ? * MON *"` for a weekly run).
To compute statistics on only a subset of the feature data, use the `row_percentage` parameter of `with_detection_window` (see Step 5).
### Step 5: (Optional) Define a detection window
By default, the detection window is an _expanding window_ covering the whole Feature Group data.
You can define a different detection window using the `window_length` and `time_offset` parameters of the `with_detection_window` method.
Additionally, you can specify the percentage of feature data on which statistics will be computed using the `row_percentage` parameter.
=== "Python"
```python
fm_monitoring_config.with_detection_window(
window_length="1w", # data ingested during one week
time_offset="1w", # starting from last week
row_percentage=0.8, # use 80% of the data
)
```
See the API reference for [`FeatureMonitoringConfig.with_detection_window`][hsfs.core.feature_monitoring_config.FeatureMonitoringConfig.with_detection_window].
#### Time basis of the windows
Rolling windows select rows by the event-time feature of the Feature View's left Feature Group when it declares one, and by commit time otherwise.
The `event_time` parameter of `create_scheduled_statistics` and `create_feature_monitoring` overrides that default for the whole configuration, detection and reference windows alike.
Pass the name, prefix included, of a timestamp, date or epoch feature that the Feature View selects, or `False` to select rows by commit time.
The left Feature Group's event-time feature is also accepted when the Feature View does not select it.
With event time the joined Feature Groups contribute their current rows, whereas with commit time the same commit interval is applied to every Feature Group in the query.
=== "Python"
```python
# windows over a time feature of a joined Feature Group
fm_monitoring_config = trans_fv.create_feature_monitoring(
name="trans_fv_amount_monitoring_by_event_time",
event_time="datetime",
)
# windows over the time the rows were written (commit time)
fm_monitoring_config = trans_fv.create_feature_monitoring(
name="trans_fv_amount_monitoring_by_commit",
event_time=False,
)
```
See [Time basis](../feature_monitoring/scheduled_statistics.md#time-basis) for how the two bases differ.
### Step 6: (Optional) Define a reference window
When setting up feature monitoring for a Feature View, the reference can be either a reference window of feature data or a training dataset.
!!! tip "Basis for Model Monitoring"
Using a training dataset as the reference is the basis for [Model Monitoring](../../mlops/model_monitoring/index.md), where a model's production inference data is compared against the distribution of its training dataset.
=== "Python"
```python
# compare statistics against a reference window
fm_monitoring_config.with_reference_window(
window_length="1w", # data ingested during one week
time_offset="2w", # starting from two weeks ago
row_percentage=0.8, # use 80% of the data
)
# or a training dataset
fm_monitoring_config.with_reference_training_dataset(
training_dataset_version=1, # use the training dataset used to train your production model
)
```
See the API reference for [`FeatureMonitoringConfig.with_reference_window`][hsfs.core.feature_monitoring_config.FeatureMonitoringConfig.with_reference_window] and [`FeatureMonitoringConfig.with_reference_training_dataset`][hsfs.core.feature_monitoring_config.FeatureMonitoringConfig.with_reference_training_dataset].
!!! info "Comparing against a specific value"
Instead of a reference window or training dataset, you can compare the detection statistics against a fixed reference value (i.e., a window of size 1).
In that case, skip this step and pass the `specific_value` parameter to `compare_on` in Step 7.
### Step 7.A: (Optional) Compare on a scalar metric
In order to compare detection and reference statistics, you need to provide the criteria for such comparison.
First, you select the feature and the metric to consider in the comparison using the `feature_name` and `metric` parameters.
Then, you can define a relative or absolute threshold using the `threshold` and `relative` parameters.
=== "Python"
```python
fm_monitoring_config.compare_on(
feature_name="amount", # the feature to compare
metric="mean",
threshold=0.2, # a relative change over 20% is considered anomalous
relative=True, # relative or absolute change
strict=False, # strict or relaxed comparison
)
```
See the API reference for [`FeatureMonitoringConfig.compare_on`][hsfs.core.feature_monitoring_config.FeatureMonitoringConfig.compare_on].
!!! info "Difference values and thresholds"
For more information about the computation of difference values and the comparison against threshold bounds see the [Comparison criteria section](../feature_monitoring/statistics_comparison.md#comparison-criteria) in the Statistics comparison guide.
### Step 7.B: (Optional) Compare on the whole distribution
Alternatively, instead of a single scalar metric, you can detect drift in the shape of a feature's distribution using `compare_on_distribution`.
Select a distribution distance metric (e.g., `PSI`) and a threshold.
A reference window or training dataset (Step 6) is required for distribution comparison.
=== "Python"
```python
fm_monitoring_config.compare_on_distribution(
feature_name="amount", # the feature to compare
metric="PSI",
threshold=0.2, # a distance above 0.2 is considered a significant shift
)
```
See the API reference for [`FeatureMonitoringConfig.compare_on_distribution`][hsfs.core.feature_monitoring_config.FeatureMonitoringConfig.compare_on_distribution].
!!! tip "More distribution options"
See the [Distribution comparison guide](../feature_monitoring/distribution_comparison.md) for the full list of metrics and binning strategies.
### Step 8: Save the configuration
Finally, you can save your feature monitoring configuration by calling the `save` method.
Once the configuration is saved, the schedule for the statistics computation and comparison will be activated automatically.
=== "Python"
```python
fm_monitoring_config.save()
```
See the API reference for [`FeatureMonitoringConfig.save`][hsfs.core.feature_monitoring_config.FeatureMonitoringConfig.save].
### Step 9: Retrieve configurations and history
Once saved, you can retrieve your feature monitoring configurations and the results of past executions directly from the Feature View.
=== "Python"
```python
# fetch all configurations attached to the feature view
configs = trans_fv.get_feature_monitoring_configs()
# or a single configuration by name
config = trans_fv.get_feature_monitoring_configs(name="trans_fv_amount_monitoring")
# fetch the history of monitoring results (with computed statistics)
history = trans_fv.get_feature_monitoring_history(
config_name="trans_fv_amount_monitoring",
with_statistics=True,
)
```
See the API reference for [`FeatureView.get_feature_monitoring_configs`][hsfs.feature_view.FeatureView.get_feature_monitoring_configs] and [`FeatureView.get_feature_monitoring_history`][hsfs.feature_view.FeatureView.get_feature_monitoring_history].
!!! api "API reference"
- [`FeatureView`][hsfs.feature_view.FeatureView]
- [`create_feature_monitoring`][hsfs.feature_view.FeatureView.create_feature_monitoring]
- [`create_scheduled_statistics`][hsfs.feature_view.FeatureView.create_scheduled_statistics]
- [`get_feature_monitoring_configs`][hsfs.feature_view.FeatureView.get_feature_monitoring_configs]
- [`get_feature_monitoring_history`][hsfs.feature_view.FeatureView.get_feature_monitoring_history]
- [`FeatureMonitoringConfig`][hsfs.core.feature_monitoring_config.FeatureMonitoringConfig]
- [`with_detection_window`][hsfs.core.feature_monitoring_config.FeatureMonitoringConfig.with_detection_window]
- [`with_reference_window`][hsfs.core.feature_monitoring_config.FeatureMonitoringConfig.with_reference_window]
- [`with_reference_training_dataset`][hsfs.core.feature_monitoring_config.FeatureMonitoringConfig.with_reference_training_dataset]
Browse the full Python API :material-arrow-right:
## Monitor a model in production
A Feature View can also monitor the inference data of a model served in production, comparing it against the training dataset the model was trained on.
This is the feature-view entry point to [Model Monitoring](../../mlops/model_monitoring/index.md).
It targets the feature view's logging feature group, so feature logging must be enabled with `feature_view.enable_logging()`, and filters the detection window by the given model name and version.
The windows select inference rows by their `log_time`, the time the prediction was logged.
The reference defaults to the training dataset version used to train the model.
=== "Python"
```python
fm_monitoring_config = trans_fv.create_model_monitoring(
name="trans_fv_model_monitoring",
model_name="my_model",
model_version=1,
).with_detection_window(
time_offset="1d",
window_length="1d",
).with_reference_training_dataset(
# omitted -> defaults to the model's training dataset version
).compare_on_distribution(
feature_name="amount",
metric="PSI",
threshold=0.2,
).save()
```
See the API reference for [`FeatureView.create_model_monitoring`][hsfs.feature_view.FeatureView.create_model_monitoring].
!!! info "Explore the API"
The [`FeatureMonitoringConfig`][hsfs.core.feature_monitoring_config.FeatureMonitoringConfig] reference documents the full set of available methods, such as enabling or disabling a configuration, triggering it manually, or deleting it.
================================================================================
# Feature Logging
Source: https://docs.hopsworks.ai/latest/user_guides/fs/feature_view/feature_logging/
# User Guide: Feature and Prediction Logging with a Feature View
Log features and predictions with a feature view, then retrieve them for debugging and monitoring.
## Feature and Prediction Logging
After you have trained a model, you can log the features it uses and the predictions with the feature view used to create the training data for the model.
You can log transformed features, untransformed features, or both.
### Enabling Feature Logging
To enable logging, set `logging_enabled=True` when creating the feature view.
One logging feature group stores transformed features, untransformed features, predictions, and logging metadata together.
Older feature views can retain separate transformed and untransformed logging groups.
The logged features are written to the offline feature store by a materialization job that is created automatically and runs on a schedule.
```python
feature_view = fs.create_feature_view("name", query, logging_enabled=True)
```
Alternatively, you can enable logging on an existing feature view by calling `feature_view.enable_logging()`.
Also, calling `feature_view.log()` will implicitly enable logging if it has not already been enabled.
### Choosing the Transport
A feature view logs through one of two transports, and the layout of its logging feature group follows from the choice.
| Transport | Path of a logged row | Readable |
| --- | --- | --- |
| `realtime` (default) | The deployment posts Arrow batches to its inference logger, which produces them to Kafka; the online store receives them within seconds and the materialization job appends them to the offline store on its schedule | Online at once with `read_log(online=True)` for the group's time to live, offline after materialization |
| `job` | The deployment appends Arrow batches to a file buffer on its pod, rotates the buffer on size or age and uploads it to HopsFS; a scheduled commit job appends the uploaded chunks to an offline-only logging group | Offline after the commit job has run |
A new feature view names its transport when logging is enabled:
```python
feature_view = fs.create_feature_view(
"name", query, logging_enabled=True, logging_transport="job"
)
feature_view.feature_logging.transport # "job"
```
A feature view that does not log yet names it when logging is enabled, and the transport is read back from the view:
```python
feature_view.enable_logging(transport="realtime")
feature_view.feature_logging.transport # "realtime"
```
The two cannot be combined on one feature view: enabling the other transport while the view logs is refused.
To move a view from one transport to the other, drop its log and recreate the logging group for the new transport with `feature_view.delete_log(transport="job")`.
Deployments take the transport from the view; a `DeploymentLoggingConfig` that names a different one is rejected.
The `job` transport keeps no online copy, so `read_log(online=True)` is refused for such a view, and a deployment that stops uploads what its buffer holds and starts the commit job before the pod exits.
Run `deployment.commit_feature_logs()` or `feature_view.materialize_log()` to commit the uploaded chunks on demand, for example after a replica was killed.
### Choosing the Materialization Interval { #choosing-the-materialization-interval }
The materialization job runs every hour or once a day.
The platform default applies unless you choose one, at creation or later.
```python
feature_view = fs.create_feature_view(
"name", query, logging_enabled=True, logging_materialization_interval="day"
)
feature_view.enable_logging(materialization_interval="hour")
feature_view.set_log_materialization_interval("day")
```
The interval only sets how often logs reach the offline store.
Run `feature_view.materialize_log()` to write them on demand between scheduled runs.
On the `job` transport the interval schedules the commit job instead.
### Logging Features and Predictions
You can log features and predictions by calling `feature_view.log`.
The logged features are written periodically to the offline store.
If you need it to be available immediately, call `feature_view.materialize_log`.
You can log either transformed or/and untransformed features.
To get untransformed features, you can specify `transform=False` in `feature_view.get_batch_data` or `feature_view.get_feature_vector(s)`.
Inference helper columns are returned along with the untransformed features.
If you have On-Demand features as well, call `feature_view.compute_on_demand_features` to get the on demand features before calling `feature_view.log`.To get the transformed features, you can call `feature_view.transform` and pass the untransformed feature with the on-demand feature.
Predictions can be optionally provided as one or more columns in the DataFrame containing the features or separately in the `predictions` argument.
There must be the same number of prediction columns as there are labels in the feature view.
It is required to provide predictions in the `predictions` argument if you provide the features as `list` instead of pandas `dataframe`.
The training dataset version will also be logged if you have called either `feature_view.init_serving(...)` or `feature_view.init_batch_scoring(...)` or if the provided model has a training dataset version.
The wallclock time of calling `feature_view.log` is automatically logged, enabling filtering by logging time when retrieving logs.
#### Example 1: Log Features Only
You have a DataFrame of features you want to log.
```python
import pandas as pd
features = pd.DataFrame(
{"feature1": [1.1, 2.2, 3.3], "feature2": [4.4, 5.5, 6.6]}
)
# Log features
feature_view.log(features)
```
#### Example 2: Log Features, Predictions, and Model
You can also log predictions, and optionally the training dataset and the model used for prediction.
```python
predictions = pd.DataFrame({"prediction": [0, 1, 0]})
# Log features and predictions
feature_view.log(
features,
predictions=predictions,
training_dataset_version=1,
model=Model(1, "model", version=1),
)
```
#### Example 3: Log Both Transformed and Untransformed Features
##### Batch Features
```python
untransformed_df = fv.get_batch_data(transformed=False)
# then apply the transformations after:
transformed_df = fv.transform(untransformed_df)
# Log untransformed features
feature_view.log(untransformed_df)
# Log transformed features
feature_view.log(transformed_features=transformed_df)
```
##### Real-time Features
```python
untransformed_vector = fv.get_feature_vector({"id": 1}, transform=False)
# then apply the transformations after:
transformed_vector = fv.transform(untransformed_vector)
# Log untransformed features
feature_view.log(untransformed_vector)
# Log transformed features
feature_view.log(transformed_features=transformed_vector)
```
## Retrieving the Log Timeline
To audit and review the feature/prediction logs, you might want to retrieve the timeline of log entries.
This helps understand when data was logged and monitor the logs.
### Retrieve Log Timeline
A log timeline is the hudi commit timeline of the logging feature group.
```python
# Retrieve the latest 10 log entries
log_timeline = feature_view.get_log_timeline(limit=10)
print(log_timeline)
```
## Reading Log Entries
You may need to read specific log entries for analysis, such as entries within a particular time range or for a specific model version and training dataset version.
### Read all Log Entries
Read all log entries for comprehensive analysis.
The output will return all values of the same primary keys instead of just the latest value.
```python
# Read all log entries
log_entries = feature_view.read_log()
print(log_entries)
```
### Read Log Entries within a Time Range
Focus on logs within a specific time range.
You can specify `start_time` and `end_time` for filtering, but the time columns will not be returned in the DataFrame.
You can provide the `start/end_time` as `datetime`, `date`, `int`, or `str` type.
Accepted date format are: `%Y-%m-%d`, `%Y-%m-%d %H`, `%Y-%m-%d %H:%M`, `%Y-%m-%d %H:%M:%S`, or `%Y-%m-%d %H:%M:%S.%f`
```python
# Read log entries from January 2022
log_entries = feature_view.read_log(
start_time="2022-01-01", end_time="2022-01-31"
)
print(log_entries)
```
### Read Log Entries by Training Dataset Version
Analyze logs from a particular version of the training dataset.
The training dataset version column will be returned in the DataFrame.
```python
# Read log entries of training dataset version 1
log_entries = feature_view.read_log(training_dataset_version=1)
print(log_entries)
```
### Read Log Entries by Model in Hopsworks
Analyze logs from a particular name and version of the HSML model.
The HSML model column will be returned in the DataFrame.
```python
# Read log entries of a specific HSML model
log_entries = feature_view.read_log(model=Model(1, "model", version=1))
print(log_entries)
```
### Read Log Entries using a Custom Filter
Provide filters which work similarly to the filter method in the `Query` class.
The filter should be part of the query in the feature view.
```python
# Read log entries where feature1 is greater than 0
log_entries = feature_view.read_log(filter=fg.feature1 > 0)
print(log_entries)
```
## Pausing and Resuming Logging
During maintenance or updates, you might need to pause logging to save computation resources.
### Pause Logging
Pause the schedule of the materialization job for writing logs to the offline store.
```python
# Pause logging
feature_view.pause_logging()
```
### Resume Logging
Resume the schedule of the materialization job for writing logs to the offline store.
```python
# Resume logging
feature_view.resume_logging()
```
## Materializing Logs
Besides the scheduled materialization job, you can materialize logs to the offline store on demand.
On the `realtime` transport this reads the rows from Kafka.
On the `job` transport this runs the commit job over the chunks that deployments uploaded to HopsFS.
This does not pause the scheduled job.
Materialization writes all columns of the logging group.
The `transformed` selector applies only to older feature views with separate logging groups.
### Materialize Logs
Materialize logs and optionally wait for the process to complete.
```python
# Materialize logs and wait for completion
materialization_result = feature_view.materialize_log(wait=True)
```
## Monitoring Feature Logging
A deployment that logs through the `realtime` transport reports what its inference logger is doing to Prometheus, and the deployment page shows it.
Open the deployment and look at the Feature logging card.
It shows four panels: rows logged per second by outcome, the time from a post to Kafka's acknowledgement, rows in flight, and posts per second by type and outcome.
The Full dashboard link opens the Feature Logging dashboard in Grafana, filtered to the same deployment, which adds in-flight bytes, rejected posts and totals over the selected range.
Two of these answer most questions.
A non-zero rate of dropped or failed rows means the deployment logs faster than the inference logger can produce, or Kafka is refusing writes; the deployment logs name the reason.
Rejected posts mean the batches the predictor builds do not match the logging group's schema, which happens after the feature view changed without a redeploy.
For a feature view on the `job` transport the card shows the same rows per second and buffered rows, the upload latency of a buffer segment to HopsFS, the bytes awaiting upload and the chunks uploaded per second; the predictor publishes these itself, and the Full dashboard adds commit job triggers and writer restarts.
The card is not shown for a view whose logging still runs through the row path of earlier releases.
Those logs are covered by the commit job's or the materialization job's own execution history instead.
## Deleting Logs
When log data is no longer needed, you might want to delete it to free up space and maintain data hygiene.
This operation deletes the feature groups and recreates new ones.
Scheduled materialization job and log timeline are reset as well.
Pass `transport="realtime"` or `transport="job"` to recreate the logging group for the other transport.
### Delete Logs
Remove all log entries.
The `transformed` selector applies only to older feature views with separate logging groups.
```python
# Delete all log entries
feature_view.delete_log()
```
Restart serving revisions after recreating a logging group so they load its new schema and destination.
================================================================================
# Deployment
Source: https://docs.hopsworks.ai/latest/user_guides/fs/feature_view/deployment/
# How To Deploy A Feature View { #feature-view-deployment }
## Introduction
In this guide, you will learn how to serve a feature view without a model.
A feature view deployment answers a prediction-style request with the transformed feature vector a model would receive.
It uses the same request contract, feature lookup, transformations, logging, and monitoring as a model deployment served by the default predictor.
See the [Deployment Schema Guide][deployment-schema] for the request contract and the error codes, which are shared with model deployments.
Use it to serve features to a model that runs outside Hopsworks, to test transformations online before a model exists, or to give a feature vector API to another team.
!!! warning "Serving identity"
The deployment looks up features as the project's serving identity, not as the caller.
Anyone allowed to call the deployment can obtain the transformed features of any entity the feature view can serve.
## Code
### Step 1: Connect to Hopsworks
=== "Python"
```python
import hopsworks
project = hopsworks.login()
fs = project.get_feature_store()
```
### Step 2: Pin a training dataset
Model-dependent transformations that need statistics, such as `min_max_scaler`, take them from a training dataset.
The deployment uses the training dataset you last read or created in this session, or the one you pass to `deploy()`.
=== "Python"
```python
feature_view = fs.get_feature_view("transactions", version=1)
# reading or creating a training dataset records it as the one to serve with
X_train, X_test, y_train, y_test = feature_view.train_test_split(test_size=0.2)
```
If the feature view has such a transformation and no training dataset was read or created, `deploy()` refuses and names the transformation, because its statistics cannot be computed.
### Step 3: Deploy the feature view
=== "Python"
```python
deployment = feature_view.deploy(
name="transactionsfv",
passed_features=["amount"], # features the client sends with each request
)
deployment.start(await_running=600)
```
The deployment name defaults to the feature view name and version without special characters.
The client publishes the deployment schema before the deployment is created, so `deployment.schema` describes the request immediately:
=== "Python"
```python
deployment.schema.describe()
print(deployment.schema.names) # the order of positional rows
```
### Step 4: Request feature vectors
Each row carries the serving keys, the passed features, the request parameters of on-demand transformations, and any extra logging columns.
The response carries one transformed vector per row and the column names.
=== "Python"
```python
response = deployment.predict(
inputs=[{"cc_num": 4473593503484549, "amount": 12.5}]
)
print(response["columns"]) # ["amount_scaled", "age_days", ...]
print(response["predictions"]) # [[0.31, -1.2, ...]]
```
Rows can also be arrays in `deployment.schema.names` order.
When every stored feature of the view is passed, the schema has no serving keys and the deployment only computes the on-demand features and applies the model-dependent transformations; see [Deployments without lookups][deployment-schema-no-lookup].
Invalid rows are refused before any feature is read; see [Errors][deployment-schema-errors].
### Step 5: Inspect the deployment
=== "Python"
```python
print(deployment.has_feature_view) # True
print(deployment.feature_view_name, deployment.feature_view_version)
print(deployment.training_dataset_version) # the pinned version
feature_view = deployment.get_feature_view() # the FeatureView object
print(deployment.get_model()) # None
```
## Feature logging
When logging is enabled on the feature view, every request is logged with the untransformed and transformed features, the request id, the training dataset version, and the reserved deployment columns `deployment_name`, `deployment_version`, `deployment_schema_id`, and `request_row`, when the logging feature group declares them.
The model columns of the log are null, because there is no model.
See [Feature logging in the Deployment Schema Guide][deployment-schema-feature-logging] for how requests are logged and for the reserved columns.
=== "Python"
```python
feature_view.enable_logging(
extra_log_columns=[
{"name": "deployment_name", "type": "string"},
{"name": "deployment_version", "type": "int"},
{"name": "deployment_schema_id", "type": "string"},
{"name": "request_row", "type": "int"},
]
)
deployment = feature_view.deploy(name="transactionsfv", passed_features=["amount"])
```
## Feature monitoring
A feature view deployment has no model, so `deployment.create_model_monitoring()` raises.
Use `deployment.create_feature_monitoring()`, which attaches a feature monitoring configuration to the logging feature group of the view:
=== "Python"
```python
config = (
deployment.create_feature_monitoring(name="amount_drift")
.with_detection_window(time_offset="1d", window_length="1d")
.with_reference_window(time_offset="8d", window_length="7d")
.compare_on(metric="MEAN", threshold=10.0, feature_name="amount")
.save()
)
deployment.get_monitoring_configs()
```
A distribution comparison (`compare_on_distribution`) over rolling windows needs KLL statistics on the logging feature group, which it does not keep by default; enable them in the logging feature group's statistics configuration first.
Two deployments of the same feature view version log to the same feature group; their rows are told apart by the reserved deployment columns, but a monitoring configuration sees both.
Deploy a separate version of the feature view when the statistics of one deployment must not include another's traffic.
## Custom predictor script
To post-process the vectors or to change how they are looked up, subclass the default predictor and pass the script to `deploy()`.
The script must end with the hand-over to the serving wrapper:
=== "Python"
```python
from hsml.default_predictor import DefaultPredict, run_kserve_wrapper
class Predict(DefaultPredict):
def model_predict(self, feature_vectors):
# a feature view deployment has no model: return the vectors
return feature_vectors.round(3)
if __name__ == "__main__":
run_kserve_wrapper()
```
Backends that support the `SERVING_SCRIPT_KIND=predictor` marker set by `deploy()` start the serving wrapper directly and never run the `__main__` block.
Older backends start the script with `python`, and the block hands over to the wrapper.
`deploy(script_file=...)` refuses a local script without it; a script already in HopsFS is not checked client-side.
## REST access
The deployment answers on the KServe V1 route of the Istio ingress, `/v1/models/:predict`, and through the Hopsworks REST API at `/project//inference/serving/:predict`.
See the [REST API Guide][hopsworks-model-serving-rest-api] for authentication and the base URL.
`deployment.get_inference_url()` returns the Istio URL, or `None` when the Istio ingress is not configured for external access.
Use the Hopsworks REST API path above when it does.
## CLI
```bash
hops fv deploy transactions --passed-feature amount
hops deployment schema transactionsfv --openapi
```
!!! api "API reference"
- [`FeatureView.deploy`][hsfs.feature_view.FeatureView.deploy]
- [`Deployment`][hsml.deployment.Deployment]
- [`start`][hsml.deployment.Deployment.start]
- [`predict`][hsml.deployment.Deployment.predict]
- [`create_feature_monitoring`][hsml.deployment.Deployment.create_feature_monitoring]
- [`schema`][hsml.deployment.Deployment.schema]
- [`training_dataset_version`][hsml.deployment.Deployment.training_dataset_version]
- [`DeploymentSchema`][hsml.deployment_schema.DeploymentSchema]
- [`describe`][hsml.deployment_schema.DeploymentSchema.describe]
Browse the full Python API :material-arrow-right:
================================================================================
# Vector Similarity Search
Source: https://docs.hopsworks.ai/latest/user_guides/fs/vector_similarity_search/
## Introduction
Vector similarity search (also called similarity search) is a technique enabling the retrieval of similar items based on their vector embeddings or representations.
Its applications range across various domains, from recommendation systems to image similarity and beyond.
In Hopsworks, vector similarity search is enabled by extending an online feature group with approximate nearest neighbor search capabilities through a vector database, such as Opensearch.
This guide provides a detailed walkthrough on how to leverage Hopsworks for vector similarity search.
## Extending Feature Groups with Similarity Search
In Hopsworks, each vector embedding in a feature group is stored in an index within the backing vector database.
By default, vector embeddings are stored in the default index for the project (created for every project in Hopsworks), but you have the option to create a new index for a feature group if needed.
Creating a separate index per feature group is particularly useful for large volumes of data, ensuring that when a feature group is deleted, its associated index is also removed.
For feature groups that use the default project index, the index will only be removed when the project is deleted - not when the feature group is deleted.
The index will store all the vector embeddings defined in that feature group, if you have more than one vector embedding in the feature group.
In the following example, we explicitly define an index for the feature group:
```aidl
from hsfs import embedding
# Specify optionally the index in the vector database
emb = embedding.EmbeddingIndex(index_name="news_fg")
```
Then, add one or more embedding features to the index.
Name and dimension of the embedding features are required for identifying which features should be indexed for k-nearest neighbor (KNN) search.
In this example, we get the dimension of the embedding by taking the length of the value of the `embedding_heading` column in the first row of the dataframe `df`.
Optionally, you can specify the similarity function among `l2_norm`, `cosine`, and `dot_product`.
Refer to [`EmbeddingIndex.add_embedding`][hsfs.embedding.EmbeddingIndex.add_embedding] for the full list of arguments.
```aidl
# Add embedding feature to the index
emb.add_embedding("embedding_heading", len(df["embedding_heading"][0]))
```
Next, you create a feature group with the `embedding_index` and ingest data to the feature group.
When the `embedding_index` is provided, the vector database is used as online feature store.
That is, all the features in the feature group are stored **exclusively** in the vector database.
The advantage of storing all features in the vector database is that it enables similarity search, and push-down filtering for all feature values.
```aidl
# Create a feature group with the embedding index
news_fg = fs.get_or_create_feature_group(
name=f"news_fg",
embedding_index=emb, # Provide the embedding index created
primary_key=["news_id"],
version=version,
online_enabled=True
)
# Write a DataFrame to the feature group, including the offline store and the ANN index (in the Vector Database)
news_fg.insert(df)
```
## Similarity Search for Feature Groups using Vector Embeddings
You provide a vector embedding as a parameter to the search query using [`FeatureGroup.find_neighbors`][hsfs.feature_group.FeatureGroup.find_neighbors], and it returns the rows in the online feature group that have vector embedding values most similar to the provided vector embedding.
It is also possible to filter rows by specifying a filter on any of the features in the feature group.
The filter is pushed down to the vector database to improve query performance.
In the first code snippet below, `find_neighbor`s returns 3 rows in `news_fg` that have the closest `news_description` values to the provided `news_description`.
In the second code snippet below, we only return news articles with a `newstype` of `sports`.
```aidl
# Search neighbor embedding with k=3
news_fg.find_neighbors(model.encode(news_description), k=3)
# Filter and search
news_fg.find_neighbors(model.encode(news_description), k=3, filter=news_fg.newstype == "sports")
```
To analyze feature values at specific points in time, you can utilize time travel functionality:
```aidl
# Time travel and read from the offline feature store
news_fg.as_of(time_in_past).read()
```
## Querying Similar Embeddings with Additional features
You can also use similarity search for vector embedding features in feature views.
In the code snippet below, we create a feature view by selecting features from the earlier `news_fg` and a new feature group `view_fg`.
If you include a feature group with vector embedding features in a feature view, **whether or not the vector embedding features are selected**, you can call `find_neighbors` on the feature view, and it will return rows containing all the feature values in the feature view.
In the example below, a list of `heading` and `view_cnt` will be returned for the news articles which are closet to provided `news_description`.
```aidl
view_fg = fs.get_or_create_feature_group(
name="view_fg",
primary_key=["news_id"],
version=version,
online_enabled=True
)
fv = fs.get_or_create_feature_view(
"news_view", version=version,
query=news_fg.select(["heading"]).join(view_fg.select(["view_cnt"]))
)
fv.find_neighbors(model.encode(news_description), k=5)
```
Note that you can use similarity search from the feature view **only if** the feature group which you are querying with `find_neighbors` has **all** the primary keys of the other feature groups.
In the example above, you are querying against the feature group `news_fg` which has the vector embedding features, and it has the feature "news_id" which is the primary key of the feature group `view_fg`.
But if `page_fg` is used as illustrated below, `find_neighbors` will fail to return any features because primary key `page_id` does not exist in `news_fg`.
--8<-- "user_guides/fs/vector_similarity_search/find-neighbors.html"
It is also possible to get back feature vector by providing the primary keys, but it is not recommended as explained in the next section.
The client fetches feature vector from the vector store and the online store for `news_fg` and `view_fg` respectively.
```aidl
fv.get_feature_vector({"news_id": 1})
```
## Performance considerations for Feature Groups with Embeddings
### Choose Features for Vector Store
While it is possible to update feature value in vector store, updating feature value in online store is more efficient.
If you have features which are frequently being updated and do not require for filtering, consider storing them separately in a different feature group.
As shown in the previous example, `view_cnt` is updated frequently and stored separately.
You can then get all the required features by using feature view.
### Choose the Appropriate Online Feature Stores
There are 2 types of online feature stores in Hopsworks: online store (RonDB) and vector store (Opensearch).
Online store is designed for retrieving feature vectors efficiently with low latency.
Vector store is designed for finding similar embedding efficiently.
If similarity search is not required, using online store is recommended for low latency retrieval of feature values including embedding.
### Use New Index per Feature Group
Create a new index per feature group to optimize retrieval performance.
## Next steps
Explore the [news search example](https://github.com/logicalclocks/hopsworks-tutorials/blob/master/api_examples/vector_similarity_search/1_feature_group_embeddings_api.ipynb), demonstrating how to use Hopsworks for implementing a news search application using natural language in the application.
Additionally, you can see the application of querying similar embeddings with additional features in this [news rank example](https://github.com/logicalclocks/hopsworks-tutorials/blob/master/api_examples/vector_similarity_search/2_feature_view_embeddings_api.ipynb).
================================================================================
# Transformation Functions
Source: https://docs.hopsworks.ai/latest/user_guides/fs/transformation_functions/
# Transformation Functions
In AI systems, [transformation functions](https://www.hopsworks.ai/dictionary/transformation) transform data to create features, the inputs to machine learning models (in both training and inference).
The [taxonomy of data transformations](../../concepts/mlops/data_transformations.md) introduces three types of data transformation prevalent in all AI systems.
Hopsworks offers simple Python APIs to define custom transformation functions.
These can be used along with [feature groups](./feature_group/index.md) and [feature views](./feature_view/overview.md) to create [on-demand transformations](./feature_group/on_demand_transformations.md) and [model-dependent transformations](./feature_view/model-dependent-transformations.md), producing modular AI pipelines that are skew-free.
## Custom Transformation Function Creation
User-defined transformation functions can be created in Hopsworks using the [`@udf`][hsfs.hopsworks_udf.udf] decorator.
These functions can be either implemented as pure Python UDFs or Pandas UDFs (User-Defined Functions).
Hopsworks offers three execution modes to control the execution of transformation functions during training dataset creation, batch inference, and online inference.
By default, Hopsworks executes transformation functions as Python UDFs for [feature vector retrieval](feature_view/feature-vectors.md) in online inference pipelines and as Pandas UDFs for both [batch data retrieval](feature_view/batch-data.md) in batch inference pipelines and [training dataset creation](feature_view/training-data.md) in training pipelines.
Python UDFs are optimized for smaller data volumes, while Pandas UDFs provide better performance on larger datasets.
This execution mode provides the optimal balance based on the data size across training dataset generations, batch inference, and online inference.
Additionally, Hopsworks allows you to explicitly set the execution mode for a transformation function to `python` or `pandas`, forcing the transformation function to always run as either a Python or Pandas UDF as specified.
A Pandas UDF in Hopsworks accepts one or more Pandas Series as input and can return either one or more Series or a Pandas DataFrame.
When integrated with PySpark applications, Hopsworks automatically executes Pandas UDFs using PySpark’s [`pandas_udf`](https://spark.apache.org/docs/3.4.1/api/python/reference/pyspark.sql/api/pyspark.sql.functions.pandas_udf.html), enabling the transformation functions to efficiently scale for large datasets.
!!! warning "Java/Scala support"
Hopsworks supports transformations functions in Python (Pandas UDFs, Python UDFs).
Transformations functions can also be executed in Python-based DataFrame frameworks (PySpark, Pandas).
There is currently no support for transformation functions in SQL or Java-based feature pipelines.
Transformation functions created in Hopsworks can be directly attached to feature views or feature groups or stored in the feature store for later retrieval.
These functions can be part of a library [installed](../../user_guides/projects/python/python_install.md) in Hopsworks or be defined in a [Jupyter notebook](../../user_guides/projects/jupyter/python_notebook.md) running a Python kernel or added when starting a Jupyter notebook or [Hopsworks job](../../user_guides/projects/jobs/spark_job.md).
!!! warning "PySpark Kernels"
Definition transformation function within a Jupyter notebook is only supported in Python Kernel.
In a PySpark Kernel transformation function have to defined as modules or added when starting a Jupyter notebook.
The `@udf` decorator in Hopsworks creates a metadata class called [`HopsworksUdf`][hsfs.hopsworks_udf.HopsworksUdf].
This class manages the necessary operations to execute the transformation function.
The decorator accepts three parameters:
- **`return_type`** (required): Specifies the data type(s) of the features returned by the transformation function.
It can be a single Python type if the function returns one transformed feature, or a list of Python types if it returns multiple transformed features.
The supported Python types that be used with the `return_type` argument are provided in the table below:
| Supported Python Types |
| :--------------------: |
| str |
| int |
| float |
| bool |
| datetime.datetime |
| datetime.date |
| datetime.time |
- **`drop`** (optional): Identifies input arguments to exclude from the output after transformations are applied.
By default, all inputs are retained in the output.
Further details on this argument can be found [below](#dropping-input-features).
- **`mode`** (optional): Determines the execution mode of the transformation function.
The argument accepts three values: `default`, `python`, or `pandas`.
By default, the `mode` is set to `default`. Further details on this argument can be found [below](#specifying-execution-modes).
Hopsworks supports four types of transformation functions across all execution modes:
1. One-to-one: Transforms one feature into one transformed feature.
2. One-to-many: Transforms one feature into multiple transformed features.
3. Many-to-one: Transforms multiple features into one transformed feature.
4. Many-to-many: Transforms multiple features into multiple transformed features.
### One-to-one transformations
To create a one-to-one transformation function, the Hopsworks `@udf` decorator must be provided with the `return_type` as a single Python type.
The transformation function should take one argument as input and return a Pandas Series.
!!! example "Creation of a one-to-one transformation function in Hopsworks."
=== "Python"
```python
from hopsworks import udf
@udf(return_type=int)
def add_one(feature):
return feature + 1
```
### Many-to-one transformations
The creation of many-to-one transformation functions is similar to that of a one-to-one transformation function, the only difference being that the transformation function accepts multiple features as input.
!!! example "Creation of a many-to-one transformation function in Hopsworks."
=== "Python"
```python
from hopsworks import udf
@udf(return_type=int)
def add_features(feature1, feature2, feature3):
return feature1 + feature2 + feature3
```
### One-to-many transformations
To create a one-to-many transformation function, the Hopsworks `@udf` decorator must be provided with the `return_type` as a list of Python types, and the transformation function should take one argument as input and return multiple features as a Pandas DataFrame.
The return types provided to the decorator must match the types of each column in the returned Pandas DataFrame.
!!! example "Creation of a one-to-many transformation function in Hopsworks."
=== "Python"
```python
from hopsworks import udf
@udf(return_type=[int, int])
def add_one_and_two(feature1):
return feature1 + 1, feature1 + 2
```
### Many-to-many transformations
The creation of a many-to-many transformation function is similar to that of a one-to-many transformation function, the only difference being that the transformation function accepts multiple features as input.
!!! example "Creation of a many-to-many transformation function in Hopsworks."
=== "Python"
```python
from hopsworks import udf
@udf(return_type=[int, int, int])
def add_one_multiple(feature1, feature2, feature3):
return feature1 + 1, feature2 + 1, feature3 + 1
```
### Specifying execution modes
The `mode` parameter of the `@udf` decorator can be used to specify the execution mode of the transformation function.
It accepts three possible values `default`, `python` and `pandas`. Each mode is explained in more detail below:
#### Default Mode
This execution mode assumes that the transformation function can be executed as either a Pandas UDF or a Python UDF.
It serves as the default mode used when the `mode` parameter is not specified.
In this mode, the transformation function is executed as a Pandas UDF during training and in the batch inference pipeline, while it operates as a Python UDF during online inference.
!!! example "Creating a many to many transformations function using the default execution mode"
=== "Python"
```python
from hopsworks import udf
# "default" mode is used if the parameter `mode` is not explicitly set.
@udf(return_type=[int, int, int])
def add_one_multiple(feature1, feature2, feature3):
return feature1 + 1, feature2 + 1, feature3 + 1
@udf(return_type=[int, int, int], mode="default")
def add_two_multiple(feature1, feature2, feature3):
return feature1 + 2, feature2 + 2, feature3 + 2
```
#### Python Mode
The transformation function can be configured to always execute as a Python UDF by setting the `mode` parameter of the `@udf` decorator to `python`.
!!! example "Creating a many to many transformation function as a Python UDF"
=== "Python"
```python
from hopsworks import udf
@udf(return_type=[int, int, int], mode="python")
def add_one_multiple(feature1, feature2, feature3):
return feature1 + 1, feature2 + 1, feature3 + 1
```
#### Pandas Mode
The transformation function can be configured to always execute as a Pandas UDF by setting the `mode` parameter of the `@udf` decorator to `pandas`.
!!! example "Creating a many to many transformations function as a Pandas UDF"
=== "Python"
```python
import pandas as pd
from hopsworks import udf
# A Pandas UDF returning a Pandas DataFrame
@udf(return_type=[int, int, int], mode="pandas")
def add_one_multiple(feature1, feature2, feature3):
return pd.DataFrame(
{
"add_one_feature1": feature1 + 1,
"add_one_feature2": feature2 + 1,
"add_one_feature3": feature3 + 1,
}
)
# A Pandas UDF returning multiple Pandas Series
@udf(return_type=[int, int, int], mode="pandas")
def add_two_multiple(feature1, feature2, feature3):
return feature1 + 2, feature2 + 2, feature3 + 2
```
### Dropping input features
The `drop` parameter of the `@udf` decorator is used to drop specific columns in the input DataFrame after transformation. If any argument of the transformation function is passed to the `drop` parameter, then the column mapped to the argument is dropped after the transformation functions are applied.
In the example below, the columns mapped to the arguments `feature1` and `feature3` are dropped after the application of all transformation functions.
!!! example "Specify arguments to drop after transformation"
=== "Python"
```python
from hopsworks import udf
@udf(return_type=[int, int, int], drop=["feature1", "feature3"])
def add_one_multiple(feature1, feature2, feature3):
return feature1 + 1, feature2 + 1, feature3 + 1
```
### Specifying output features names for transformation functions
The [`TransformationFunction.alias`][hsfs.transformation_function.TransformationFunction.alias] function of a transformation function allows the specification of names of transformed features generated by the transformation function.
Each name must be uniques and should be at-most 63 characters long.
If no name is provided via the `alias` function, Hopsworks generates default output feature names when [on-demand](./feature_group/on_demand_transformations.md) or [model-dependent](./feature_view/model-dependent-transformations.md) transformation functions are created.
!!! example "Specifying output column names for transformation functions."
=== "Python"
```python
from hopsworks import udf
@udf(return_type=[int, int, int], drop=["feature1", "feature3"])
def add_one_multiple(feature1, feature2, feature3):
return feature1 + 1, feature2 + 1, feature3 + 1
# Specifying output feature names of the transformation function.
add_one_multiple.alias(
"transformed_feature1", "transformed_feature2", "transformed_feature3"
)
```
### Training dataset statistics
A keyword argument `statistics` can be defined in the transformation function if it requires training dataset statistics for any of its arguments.
The `statistics` argument must be assigned an instance of the class [`TransformationStatistics`][hsfs.transformation_statistics.TransformationStatistics] as the default value.
The `TransformationStatistics` instance must be initialized using the names of the arguments requiring statistics.
!!! warning "Transformation Statistics"
The statistics provided to the transformation function is the statistics computed using [the train set](https://www.hopsworks.ai/dictionary/train-training-set).
Training dataset statistics are not available for on-demand transformations.
The `TransformationStatistics` instance contains separate objects with the same name as the arguments used to initialize it.
These objects encapsulate statistics related to the argument as instances of the class [`FeatureTransformationStatistics`][hsfs.transformation_statistics.FeatureTransformationStatistics].
Upon instantiation, instances of `FeatureTransformationStatistics` contain `None` values and are updated with the required statistics after the creation of a training dataset.
!!! example "Creation of a transformation function in Hopsworks that uses training dataset statistics"
=== "Python"
```python
from hopsworks import udf
from hopsworks.transformation_statistics import TransformationStatistics
stats = TransformationStatistics("argument1", "argument2", "argument3")
@udf(int)
def add_features(argument1, argument2, argument3, statistics=stats):
return (
argument1
+ argument2
+ argument3
+ statistics.argument1.mean
+ statistics.argument2.mean
+ statistics.argument3.mean
)
```
### Passing context variables to transformation function
The `context` keyword argument can be defined in a transformation function to access shared context variables.
These variables contain common data used across transformation functions.
By including the context argument, you can pass the necessary data as a dictionary into the into the `context` argument of the transformation function during [training dataset creation](feature_view/training-data.md#passing-context-variables-to-transformation-functions) or [feature vector retrieval](feature_view/feature-vectors.md#passing-context-variables-to-transformation-functions) or [batch data retrieval](feature_view/batch-data.md#passing-context-variables-to-transformation-functions).
!!! example "Creation of a transformation function in Hopsworks that accepts context variables"
=== "Python"
```python
from hopsworks import udf
@udf(int)
def add_features(argument1, context):
return argument1 + context["value_to_add"]
```
## Saving to the Feature Store
To save a transformation function to the feature store, use the function `create_transformation_function`. It creates a [`TransformationFunction`][hsfs.transformation_function.TransformationFunction] object which can then be saved by calling the save function.
The save function will throw an error if another transformation function with the same name and version is already saved in the feature store.
!!! example "Register transformation function `add_one` in the Hopsworks feature store"
=== "Python"
```python
plus_one_meta = fs.create_transformation_function(
transformation_function=add_one, version=1
)
plus_one_meta.save()
```
## Retrieval from the Feature Store
To retrieve all transformation functions from the feature store, use the function `get_transformation_functions`, which returns the list of `TransformationFunction` objects.
A specific transformation function can be retrieved using its `name` and `version` with the function `get_transformation_function`.
If only the `name` is provided, then the version will default to 1.
!!! example "Retrieving transformation functions from the feature store"
=== "Python"
```python
# get all transformation functions
fs.get_transformation_functions()
# get transformation function by name. This will default to version 1
plus_one_fn = fs.get_transformation_function(name="plus_one")
# get transformation function by name and version.
plus_one_fn = fs.get_transformation_function(name="plus_one", version=2)
```
## Using transformation functions
Transformation functions can be used by attaching it to a feature view to [create model-dependent transformations](./feature_view/model-dependent-transformations.md) or attached to feature groups to [create on-demand transformations](./feature_group/on_demand_transformations.md)
## Chained Transformation Functions
Transformation functions can be chained: the output column of one transformation function can serve as the input to another.
Hopsworks resolves the execution order automatically using a topological sort of the resulting DAG, so dependencies always run before their consumers.
Chaining works for both on-demand transformations attached to a feature group and model-dependent transformations attached to a feature view.
!!! example "Chained model-dependent transformations on a feature view"
=== "Python"
```python
from hopsworks import udf
@udf(int)
def add_one(col):
return col + 1
@udf(int)
def add(a, b):
return a + b
fv = fs.create_feature_view(
name="chained_mdts_fv",
query=fg.select_all(),
transformation_functions=[
add_one("data1").alias("data1_plus_one"),
add_one("data2").alias("data2_plus_one"),
add("data1_plus_one", "data2_plus_one").alias("sum_plus_two"),
],
version=1,
)
```
The same DAG drives offline training data generation and online feature vector retrieval, so chains apply uniformly across both paths.
Statistics-based transformations participate in chains too: a transformation that requires statistics on another transformation's output is fit on that intermediate output, as described in [model-dependent transformations][chaining-model-dependent-transformations].
Chaining also works across the two transformation types without additional setup: an on-demand transformation's output column becomes a feature in its feature group, which a feature view can consume and feed into a model-dependent transformation.
A configuration with no valid execution order is rejected: a duplicate output column or a cycle between transformation functions raises an error naming the offending functions, which can be fixed by renaming outputs with `.alias()`.
### Visualizing the execution DAG
The execution DAG is shown in the Hopsworks UI on the feature view and feature group overview pages under "Transformation execution DAG."
The same graph can be rendered from the SDK with `visualize_transformations()`, available on both feature views and feature groups.
It renders as a Mermaid flowchart in Jupyter and as text elsewhere.
!!! example "Visualizing transformation DAGs"
=== "Python"
```python
# Render both the model-dependent and on-demand DAGs.
fv.visualize_transformations()
# Render only the model-dependent DAG, top-to-bottom layout.
fv.visualize_transformations(kind="model_dependent", orient="TB")
# Render the on-demand DAG of a feature group.
fg.visualize_transformations()
```
### Transformation Functions Performance Tuning
Transformation functions execute sequentially unless the `n_processes` argument requests worker processes.
The argument is accepted by the feature view and feature group entry points that execute transformations, such as `get_feature_vector`, `get_feature_vectors`, `get_batch_data`, `training_data`, and `transform`.
Parallelism is strictly opt-in because whether the worker-pool overhead pays off depends on the cost of your transformation functions.
With more than one worker process, independent transformation functions in the DAG run concurrently, while a chained sequence always runs in dependency order. On the Spark engine `n_processes` is ignored because the whole DAG is pushed down to Spark, which distributes the work itself. For batch and offline calls such as `get_feature_vectors`, `get_batch_data`, and `training_data` with CPU-heavy functions benefit from `n_processes >= 2`, for vectorized Pandas UDFs on small inputs, sequential execution is at least as fast because the pool overhead dominates.
For online serving, spawning the worker pool during the first request would add the pool startup cost to that request's latency. Passing `n_processes` to `init_serving` or `init_batch_scoring` pre-spawns the pool at initialization time and makes that value the default for subsequent retrieval calls; an explicit `n_processes` on an individual call still takes precedence.
!!! example "Pre-spawning the worker pool for online serving"
=== "Python"
```python
fv.init_serving(training_dataset_version=1, n_processes=2)
# Served using the pool of two workers spawned at init time.
vector = fv.get_feature_vector(entry={"id": 1})
```
The worker pool start method defaults to `fork` on Linux and `spawn` on macOS and Windows. Set the `HOPSWORKS_TF_POOL_START_METHOD` environment variable to `fork`, `forkserver`, or `spawn` to override it.
================================================================================
# Compute Engines
Source: https://docs.hopsworks.ai/latest/user_guides/fs/compute_engines/
## Compute Engines
In order to execute a feature pipeline to write to the Feature Store, as well as to retrieve data from the Feature Store, you need a compute engine.
Hopsworks Feature Store APIs are built around dataframes, that means feature data is inserted into the Feature Store from a Dataframe and likewise when reading data from the Feature Store, it is returned
as a Dataframe.
As such, Hopsworks supports four computational engines:
1. [Apache Spark](https://spark.apache.org): Spark Dataframes and Spark Structured Streaming Dataframes are supported, both from Python environments (PySpark) and from Scala environments.
2. [Python](https://www.python.org/): For pure Python environments without dependencies on Spark, Hopsworks supports [Pandas Dataframes](https://pandas.pydata.org/) and [Polars Dataframes](https://pola.rs/).
3. [Apache Beam](https://beam.apache.org/) *experimental*: Beam Data Streams are currently supported as an experimental feature from Java/Scala environments.
4. [Java](https://www.java.com): For pure Java environments without dependencies on Spark, Hopsworks supports writing using List of POJO Objects.
Hopsworks supports running [compute on the platform itself](../../concepts/dev/inside.md) in the form of [Jobs](../projects/jobs/pyspark_job.md) or in [Jupyter Notebooks](../projects/jupyter/python_notebook.md).
Alternatively, you can also connect to Hopsworks using Python or Spark from [external environments](../../concepts/dev/outside.md), given that there is network connectivity.
## Functionality Support
Hopsworks is aiming to provide functional parity between the computational engines, however, there are certain Hopsworks functionalities which are exclusive to the engines.
| Functionality | Method | Spark | Python | Beam | Java | Comment |
| --- | --- | --- | --- | --- | --- | --- |
| Feature Group Creation from dataframes | [`FeatureStore.create_feature_group`][hsfs.feature_store.FeatureStore.create_feature_group] | :white_check_mark: | :white_check_mark: | - | - | Currently Beam/Java doesn't support registering feature group metadata. Thus it needs to be pre-registered before you can write real time features computed by Beam. |
| Training Dataset Creation from dataframes | [`TrainingDataset.save`][hsfs.training_dataset.TrainingDataset.save] | :white_check_mark: | - | - | - | Functionality was deprecated in version 3.0 |
| Data validation using Great Expectations for streaming dataframes | [`FeatureGroup.validate`][hsfs.feature_group.FeatureGroup.validate] [`FeatureGroup.insert_stream`][hsfs.feature_group.FeatureGroup.insert_stream] | - | - | - | - | `insert_stream` does not perform any data validation even when a expectation suite is attached. |
| Stream ingestion | [`FeatureGroup.insert_stream`][hsfs.feature_group.FeatureGroup.insert_stream] | :white_check_mark: | - | :white_check_mark: | :white_check_mark: | Python/Pandas/Polars has currently no notion of streaming. |
| Reading from Streaming Storage Connectors | [`KafkaConnector.read_stream`][hsfs.storage_connector.KafkaConnector.read_stream] | :white_check_mark: | - | - | - | Python/Pandas/Polars has currently no notion of streaming. For Beam/Java only write operations are supported |
| Reading training data from external storage other than S3 | [`FeatureView.get_training_data`][hsfs.feature_view.FeatureView.get_training_data] | :white_check_mark: | - | - | - | Reading training data that was written to external storage using a Storage Connector other than S3 can currently not be read using Hopsworks APIs, instead you will have to use the storage's native client. |
| Reading External Feature Groups into Dataframe | [`ExternalFeatureGroup.read`][hsfs.feature_group.ExternalFeatureGroup.read] | :white_check_mark: | - | - | - | Reading an External Feature Group directly into a Pandas/Polars Dataframe is not supported, however, you can use the [Query API][hsfs.constructor.query.Query] to create Feature Views/Training Data containing External Feature Groups. |
| Read Queries containing External Feature Groups into Dataframe | [`Query.read`][hsfs.constructor.query.Query.read] | :white_check_mark: | - | - | - | Reading a Query containing an External Feature Group directly into a Pandas/Polars Dataframe is not supported, however, you can use the Query to create Feature Views/Training Data and write the data to a Storage Connector, from where you can read up the data into a Pandas/Polars Dataframe. |
## Python
### Python Inside Hopsworks
If you are using Spark or Python within Hopsworks, there is no further configuration required.
Head over to the [Getting Started Guide](https://colab.research.google.com/github/logicalclocks/hopsworks-tutorials/blob/master/quickstart.ipynb){:target="_blank"}.
### Python Outside Hopsworks
Connecting to the Feature Store from any Python environment, such as your local environment or Google Colab, requires setting up an API Key and installing the Hopsworks Python client library.
The [Python integration guide](../integrations/python.md) explains step by step how to connect to the Feature Store from any Python environment.
## Spark
### Spark Inside Hopsworks
If you are using Spark or Python within Hopsworks, there is no further configuration required.
Head over to the [Getting Started Guide](https://colab.research.google.com/github/logicalclocks/hopsworks-tutorials/blob/master/quickstart.ipynb){:target="_blank"}.
### Spark Outside Hopsworks
Connecting to the Feature Store from an external Spark cluster, such as Cloudera or Databricks, requires configuring it with the Hopsworks client jars, configuration and certificates.
The [Spark integration guide](../integrations/spark.md) explains step by step how to connect to the Feature Store from an external Spark cluster.
## Beam
### Beam Inside Hopsworks
Beam is only supported as an external client.
### Beam Outside Hopsworks
Connecting to the Feature Store from Beam DataFlowRunner, requires configuring the Hopsworks certificates.
The [Beam integration guide](../integrations/beam.md) explains step by step how to connect to the Feature Store from Beam Dataflow Runner.
!!! warning
Apache Beam integration with Hopsworks feature store was only tested using Dataflow Runner.
For more details head over to the [Getting Started Guide](https://github.com/logicalclocks/hopsworks-tutorials/tree/master/integrations/java/beam).
## Java
It is also possible to interact to Hopsworks feature store using pure Java environments without dependencies on Spark or Beam.
For more details head over to the [Getting Started Guide](https://github.com/logicalclocks/hopsworks-tutorials/tree/master/java).
================================================================================
# Client Integrations
Source: https://docs.hopsworks.ai/latest/user_guides/integrations/
# Client Integrations
Hopsworks is an open platform, reachable from the tools you already use.
Pick the client you connect from.
- **Python**
---
Any Python environment, including SageMaker, Google Colab and Kubeflow.
[Connect from Python](python.md)
- **Java**
---
Java and Scala clients.
[Connect from Java](java.md)
- **Databricks**
---
Connect a Databricks workspace.
[Connect from Databricks](databricks/networking.md)
- **AWS EMR**
---
Connect an EMR cluster.
[Connect from AWS EMR](emr/emr_configuration.md)
- **Azure HDInsight**
---
Connect an HDInsight cluster.
[Connect from Azure HDInsight](hdinsight.md)
- **Azure Machine Learning**
---
ML Studio designer and notebooks.
[Connect from Azure Machine Learning](mlstudio_designer.md)
- **Apache Spark**
---
Connect an external Spark cluster.
[Connect from Apache Spark](spark.md)
- **Apache Beam**
---
Feature pipelines on Beam.
[Connect from Apache Beam](beam.md)
================================================================================
# Python / SageMaker / Kubeflow
Source: https://docs.hopsworks.ai/latest/user_guides/integrations/python/
# Python Environments (Local, AWS SageMaker, Google Colab or Kubeflow)
This guide explains step by step how to connect to Hopsworks from any Python environment such as your local environment, AWS SageMaker, Google Colab or Kubeflow.
## Install Python Library
To be able to interact with Hopsworks from a Python environment you need to install the `Hopsworks` Python library.
The library is available on [PyPi](https://pypi.org/project/hopsworks/) and is installed with the `python` profile:
=== "uv"
```bash
uv pip install "hopsworks[python]~=[HOPSWORKS_VERSION]"
```
=== "pip"
```bash
pip install "hopsworks[python]~=[HOPSWORKS_VERSION]"
```
!!! attention "Python Profile"
A bare `hopsworks` install does not bring the dependencies needed to use the library from a pure Python environment.
Always install with the `python` profile, `hopsworks[python]`.
!!! attention "Matching Hopsworks version"
We recommend that the major and minor version of the Python library match the major and minor version of the Hopsworks deployment.
You find the Hopsworks version at the bottom of the help menu in the top navigation bar
## Generate an API key
For instructions on how to generate an API key follow this [user guide](../projects/api_key/create_api_key.md).
For the Python client to work correctly make sure you add the following scopes to your API key:
1. featurestore
2. project
3. job
4. kafka
## Connect to the Feature Store
You are now ready to connect to Hopsworks from your Python environment:
```python
import hopsworks
project = hopsworks.login(
host="my_instance", # DNS of your Hopsworks instance
port=443, # Port to reach your Hopsworks instance, defaults to 443
project="my_project", # Name of your Hopsworks project
api_key_value="apikey", # The API key to authenticate with Hopsworks
engine="python", # Use the Python engine
)
fs = project.get_feature_store() # Get the project's default feature store
```
!!! note "Engine"
`Hopsworks` leverages several engines depending on whether you are running using Apache Spark or Pandas/Polars.
The default behaviour of the library is to use the `spark` engine if you do not specify any `engine` option in the `login` method and if the `PySpark` library is available in the environment.
Please refer to the [Spark integration guide](spark.md) to configure your PySpark cluster to interact with Hopsworks.
## Next Steps
For more information on how to use the Hopsworks API check out the other guides or the [Login API][hopsworks.login].
================================================================================
# Networking
Source: https://docs.hopsworks.ai/latest/user_guides/integrations/emr/networking/
# Networking
In order for Spark to communicate with the Hopsworks Feature Store from EMR, networking needs to be set up correctly.
This includes deploying the Hopsworks Feature Store to either the same VPC or enable VPC peering between the VPC of the EMR cluster and the Hopsworks Feature Store.
## Step 1: Ensure network connectivity
The DataFrame API needs to be able to connect directly to the IP on which the Feature Store is listening.
This means that if you deploy the Feature Store on AWS you will either need to deploy the Feature Store in the same VPC as your EMR
cluster or to set up [VPC Peering](https://docs.aws.amazon.com/vpc/latest/peering/create-vpc-peering-connection.html) between your EMR VPC and the Feature Store VPC.
### Option 1: Deploy the Feature Store in the EMR VPC
When deploying the Hopsworks Feature Store, select the EMR *VPC* and *Availability Zone* as the VPC and Availability Zone of your Feature Store.
Identify your EMR VPC in the Summary of your EMR cluster:
Identify the EMR VPC
Identify the EMR VPC
### Option 2: Set up VPC peering
Follow the guide [VPC Peering](https://docs.aws.amazon.com/vpc/latest/peering/create-vpc-peering-connection.html) to set up VPC peering between the Feature Store and EMR.
Get your Feature Store *VPC ID* and *CIDR* by searching for the Feature Store VPC in the AWS Management Console:
Identify the Feature Store VPC
## Step 2: Configure the Security Group
The Feature Store *Security Group* needs to be configured to allow traffic from your EMR clusters to be able to connect to the Feature Store.
Open your feature store instance under EC2 in the AWS Management Console and ensure that ports *443*, *3306*, *9083*, *9085*, *8020* and *30010* (443,3306,8020,30010,9083,9085) are reachable
from the EMR Security Group:
Hopsworks Feature Store Security Group
Connectivity from the EMR Security Group can be allowed by opening the Security Group, adding a port to the Inbound rules and setting the EMR master and core security group as source:
Hopsworks Feature Store Security Group details
You can find your EMR security groups in the EMR cluster summary:
EMR Security Groups
## Next Steps
Continue with the [Configure EMR for the Hopsworks Feature Store](emr_configuration.md), in order to be able to use the Hopsworks Feature Store.
================================================================================
# Configure EMR for Hopsworks
Source: https://docs.hopsworks.ai/latest/user_guides/integrations/emr/emr_configuration/
# Configure EMR for the Hopsworks Feature Store
To enable EMR to access the Hopsworks Feature Store, you need to set up a Hopsworks API key, add a bootstrap action and configurations to your EMR cluster.
!!! info
Ensure [Networking](networking.md) is set up correctly before proceeding with this guide.
## Step 1: Set up a Hopsworks API key
For instructions on how to generate an API key follow this [user guide](../../projects/api_key/create_api_key.md).
For the EMR integration to work correctly make sure you add the following scopes to your API key:
1. featurestore
2. project
3. job
4. kafka
### Store the API key in the AWS Secrets Manager
In the AWS management console ensure that your active region is the region you use for EMR.
Go to the *AWS Secrets Manager* and select *Store new secret*.
Select *Other type of secrets* and add *api-key*
as the key and paste the API key created in the previous step as the value.
Click next.
Store a Hopsworks API key in the Secrets Manager
As a secret name, enter *hopsworks/featurestore*.
Select next twice and finally store the secret.
Then click on the secret in the secrets list and take note of the *Secret ARN*.
Name the secret
### Grant access to the secret to the EMR EC2 instance profile
Identify your EMR EC2 instance profile in the EMR cluster summary:
Identify your EMR EC2 instance profile
In the AWS Management Console, go to *IAM*, select *Roles* and then the EC2 instance profile used by your EMR cluster.
Select *Add inline policy*.
Choose *Secrets Manager* as a service, expand the *Read* access level and check *GetSecretValue*.
Expand Resources and select *Add ARN*.
Paste the ARN of the secret created in the previous step.
Click on *Review*, give the policy a name and click on *Create policy*.
Configure the access policy for the Secrets Manager
## Step 2: Configure your EMR cluster
### Add the Hopsworks Feature Store configuration to your EMR cluster
In order for EMR to be able to talk to the Feature Store, you need to update the Hadoop and Spark configurations.
Copy the configuration below and replace ip-XXX-XX-XX-XXX.XX-XXXX-X.compute.internal with the private DNS name of your Hopsworks master node.
```json
[
{
"Classification": "hadoop-env",
"Properties": {
},
"Configurations": [
{
"Classification": "export",
"Properties": {
"HADOOP_CLASSPATH": "$HADOOP_CLASSPATH:/usr/lib/hopsworks/client/*"
},
"Configurations": [
]
}
]
},
{
"Classification": "spark-defaults",
"Properties": {
"spark.hadoop.hops.ipc.server.ssl.enabled": true,
"spark.hadoop.fs.hopsfs.impl": "io.hops.hopsfs.client.HopsFileSystem",
"spark.hadoop.client.rpc.ssl.enabled.protocol": "TLSv1.2",
"spark.hadoop.hops.ssl.hostname.verifier": "ALLOW_ALL",
"spark.hadoop.hops.rpc.socket.factory.class.default": "io.hops.hadoop.shaded.org.apache.hadoop.net.HopsSSLSocketFactory",
"spark.hadoop.hops.ssl.keystores.passwd.name": "/usr/lib/hopsworks/material_passwd",
"spark.hadoop.hops.ssl.keystore.name": "/usr/lib/hopsworks/keyStore.jks",
"spark.hadoop.hops.ssl.trustore.name": "/usr/lib/hopsworks/trustStore.jks",
"spark.serializer": "org.apache.spark.serializer.KryoSerializer",
"spark.executor.extraClassPath": "/usr/lib/hopsworks/client/*",
"spark.driver.extraClassPath": "/usr/lib/hopsworks/client/*",
"spark.sql.hive.metastore.jars": "path",
"spark.sql.hive.metastore.jars.path": "/usr/lib/hopsworks/apache-hive-bin/lib/*",
"spark.hadoop.hive.metastore.uris": "thrift://ip-XXX-XX-XX-XXX.XX-XXXX-X.compute.internal:9083"
}
},
]
```
When you create your EMR cluster, add the configuration:
!!! note
Don't forget to replace ip-XXX-XX-XX-XXX.XX-XXXX-X.compute.internal with the private DNS name of your Hopsworks master node.
Configure EMR to access the Feature Store
### Add the Bootstrap Action to your EMR cluster
EMR requires Hopsworks connectors to be able to communicate with the Hopsworks Feature Store.
These connectors can be installed with the
bootstrap action shown below.
Copy the content into a file and name the file `hopsworks.sh`.
Copy that file into any S3 bucket that
is readable by your EMR clusters and take note of the S3 URI of that file e.g., `s3://my-emr-init/hopsworks.sh`.
```bash
#!/bin/bash
set -e
if [ "$#" -ne 3 ]; then
echo "Usage hopsworks.sh HOPSWORKS_API_KEY_SECRET, HOPSWORKS_HOST, PROJECT_NAME"
exit 1
fi
SECRET_NAME=$1
HOST=$2
PROJECT=$3
API_KEY=$(aws secretsmanager get-secret-value --secret-id $SECRET_NAME | jq -r .SecretString | jq -r '.["api-key"]')
PROJECT_ID=$(curl -H "Authorization: ApiKey ${API_KEY}" https://$HOST/hopsworks-api/api/project/getProjectInfo/$PROJECT | jq -r .projectId)
sudo yum -y install python3-devel.x86_64 || true
sudo mkdir /usr/lib/hopsworks
sudo chown hadoop:hadoop /usr/lib/hopsworks
cd /usr/lib/hopsworks
curl -o client.tar.gz -H "Authorization: ApiKey ${API_KEY}" https://$HOST/hopsworks-api/api/project/$PROJECT_ID/client
tar -xvf client.tar.gz
tar -xzf client/apache-hive-*-bin.tar.gz || true
mv apache-hive-*-bin apache-hive-bin
rm client.tar.gz
rm client/apache-hive-*-bin.tar.gz
curl -H "Authorization: ApiKey ${API_KEY}" https://$HOST/hopsworks-api/api/project/$PROJECT_ID/credentials | jq -r .kStore | base64 -d > keyStore.jks
curl -H "Authorization: ApiKey ${API_KEY}" https://$HOST/hopsworks-api/api/project/$PROJECT_ID/credentials | jq -r .tStore | base64 -d > trustStore.jks
echo -n $(curl -H "Authorization: ApiKey ${API_KEY}" https://$HOST/hopsworks-api/api/project/$PROJECT_ID/credentials | jq -r .password) > material_passwd
chmod -R o-rwx /usr/lib/hopsworks
sudo pip3 install --upgrade hopsworks~=X.X.0
```
!!! attention "Matching Hopsworks version"
We recommend that the major and minor version of the Python library match the major and minor version of the Hopsworks deployment.
You find the Hopsworks version at the bottom of the help menu in the top navigation bar
Add the bootstrap actions when configuring your EMR cluster.
Provide 3 arguments to the bootstrap action: The name of the API key secret e.g., `hopsworks/featurestore`,
the public DNS name of your Hopsworks cluster, such as `ad005770-33b5-11eb-b5a7-bfabd757769f.cloud.hopsworks.ai`, and the name of your Hopsworks project, e.g. `demo_fs_meb10179`.
Set the bootstrap action for EMR
Your EMR cluster will now be able to access your Hopsworks Feature Store.
## Next Steps
Use the [Login API][hopsworks.login] to connect to the Hopsworks Feature Store.
For more information about how to use the Feature Store, see the [Quickstart Guide](https://colab.research.google.com/github/logicalclocks/hopsworks-tutorials/blob/master/quickstart.ipynb){:target="_blank"}.
================================================================================
# Azure HDInsight
Source: https://docs.hopsworks.ai/latest/user_guides/integrations/hdinsight/
# Configure HDInsight for the Hopsworks Feature Store
To enable HDInsight to access the Hopsworks Feature Store, you need to set up a Hopsworks API key, add a script action and configurations to your HDInsight cluster.
!!! info "Prerequisites"
A HDInsight cluster with cluster type Spark is required to connect to the Feature Store.
You can either use an existing cluster or create a new one.
!!! info "Network Connectivity"
To be able to connect to the Feature Store, please ensure that your HDInsight cluster and the Hopsworks Feature Store are either in the same [Virtual Network](https://docs.microsoft.com/en-us/azure/virtual-network/virtual-networks-overview) or [Virtual Network Peering](https://docs.microsoft.com/en-us/azure/virtual-network/virtual-network-manage-peering) is set up between the different networks.
In addition, ensure that the Network Security Group of your Hopsworks instance is configured to allow incoming traffic from your HDInsight cluster on ports 443, 3306, 8020, 30010, 9083 and 9085 (443,3306,8020,30010,9083,9085).
See [Network security groups](https://docs.microsoft.com/en-us/azure/virtual-network/network-security-groups-overview) for more information.
## Step 1: Set up a Hopsworks API key
For instructions on how to generate an API key follow this [user guide](../projects/api_key/create_api_key.md).
For the HDInsight integration to work correctly make sure you add the following scopes to your API key:
1. featurestore
2. project
3. job
4. kafka
## Step 2: Use a script action to install the Feature Store connector
HDInsight requires Hopsworks connectors to be able to communicate with the Hopsworks Feature Store.
These connectors can be installed with the script action shown below.
Copy the content into a file, name the file `hopsworks.sh` and replace MY_INSTANCE, MY_PROJECT, MY_VERSION, MY_API_KEY and MY_CONDA_ENV with your values.
Copy the `hopsworks.sh` file into any storage that is readable by your HDInsight clusters and take note of the URI of that file e.g., `https://account.blob.core.windows.net/scripts/hopsworks.sh`.
The script action needs to be applied head and worker nodes and can be applied during cluster creation or to an existing cluster.
Ensure to persist the script action so that it is run on newly created nodes.
For more information about how to use script actions, see [Customize Azure HDInsight clusters by using script actions](https://docs.microsoft.com/en-us/azure/hdinsight/hdinsight-hadoop-customize-cluster-linux).
!!! attention "Matching Hopsworks version"
We recommend that the major and minor version of the Python library match the major and minor version of the Hopsworks deployment.
You find the Hopsworks version at the bottom of the help menu in the top navigation bar
Feature Store script action:
```bash
set -e
HOST="MY_INSTANCE.cloud.hopsworks.ai" # DNS of your Feature Store instance
PROJECT="MY_PROJECT" # Port to reach your Hopsworks instance, defaults to 443
HOPSWORKS_VERSION="MY_VERSION" # The major version of Hopsworks library needs to match the major version of Hopsworks
API_KEY="MY_API_KEY" # The API key to authenticate with Hopsworks
CONDA_ENV="MY_CONDA_ENV" # py35 is the default for HDI 3.6
apt-get --assume-yes install python3-dev
apt-get --assume-yes install jq
/usr/bin/anaconda/envs/$CONDA_ENV/bin/pip install hopsworks==$HOPSWORKS_VERSION
PROJECT_ID=$(curl -H "Authorization: ApiKey ${API_KEY}" https://$HOST/hopsworks-api/api/project/getProjectInfo/$PROJECT | jq -r .projectId)
mkdir -p /usr/lib/hopsworks
chown root:hadoop /usr/lib/hopsworks
cd /usr/lib/hopsworks
curl -o client.tar.gz -H "Authorization: ApiKey ${API_KEY}" https://$HOST/hopsworks-api/api/project/$PROJECT_ID/client
tar -xvf client.tar.gz
tar -xzf client/apache-hive-*-bin.tar.gz
mv apache-hive-*-bin apache-hive-bin
rm client.tar.gz
rm client/apache-hive-*-bin.tar.gz
curl -H "Authorization: ApiKey ${API_KEY}" https://$HOST/hopsworks-api/api/project/$PROJECT_ID/credentials | jq -r .kStore | base64 -d > keyStore.jks
curl -H "Authorization: ApiKey ${API_KEY}" https://$HOST/hopsworks-api/api/project/$PROJECT_ID/credentials | jq -r .tStore | base64 -d > trustStore.jks
echo -n $(curl -H "Authorization: ApiKey ${API_KEY}" https://$HOST/hopsworks-api/api/project/$PROJECT_ID/credentials | jq -r .password) > material_passwd
chown -R root:hadoop /usr/lib/hopsworks
```
## Step 3: Configure HDInsight for Feature Store access
The Hadoop and Spark installations of the HDInsight cluster need to be configured in order to access the Feature Store.
This can be achieved either by using a [bootstrap script](https://docs.microsoft.com/en-us/azure/hdinsight/hdinsight-hadoop-customize-cluster-bootstrap) when creating clusters or using [Ambari](https://docs.microsoft.com/en-us/azure/hdinsight/hdinsight-hadoop-manage-ambari) on existing clusters.
Apply the following configurations to your HDInsight cluster.
!!! attention "Using Hive and the Feature Store"
HDInsight clusters cannot use their local Hive when being configured for the Feature Store as the Feature Store relies on custom Hive binaries and its own Metastore which will overwrite the local one.
If you rely on Hive for feature engineering then it is advised to write your data to an external data storage such as ADLS from your main HDInsight cluster and in the Feature Store, create an [on-demand](../../concepts/fs/feature_group/on_demand_feature.md) Feature Group on the storage container in ADLS.
Hadoop hadoop-env.sh:
```sh
export HADOOP_CLASSPATH=$HADOOP_CLASSPATH:/usr/lib/hopsworks/client/*
```
Hadoop core-site.xml:
```ini
hops.ipc.server.ssl.enabled=true
fs.hopsfs.impl=io.hops.hopsfs.client.HopsFileSystem
client.rpc.ssl.enabled.protocol=TLSv1.2
hops.ssl.keystore.name=/usr/lib/hopsworks/keyStore.jks
hops.rpc.socket.factory.class.default=io.hops.hadoop.shaded.org.apache.hadoop.net.HopsSSLSocketFactory
hops.ssl.keystores.passwd.name=/usr/lib/hopsworks/material_passwd
hops.ssl.hostname.verifier=ALLOW_ALL
hops.ssl.trustore.name=/usr/lib/hopsworks/trustStore.jks
```
Spark spark-defaults.conf:
```ini
spark.executor.extraClassPath=/usr/lib/hopsworks/client/*
spark.driver.extraClassPath=/usr/lib/hopsworks/client/*
spark.sql.hive.metastore.jars=path
spark.sql.hive.metastore.jars.path=/usr/lib/hopsworks/apache-hive-bin/lib/*
```
Spark hive-site.xml:
```ini
hive.metastore.uris=thrift://MY_HOPSWORKS_INSTANCE_PRIVATE_IP:9083
```
!!! info
Replace MY_HOPSWORKS_INSTANCE_PRIVATE_IP with the private IP address of you Hopsworks Feature Store.
## Step 5: Connect to the Feature Store
You are now ready to connect to the Hopsworks Feature Store, for instance using a Jupyter notebook in HDInsight with a PySpark3 kernel:
```python
import hopsworks
# Put the API key into Key Vault for any production setup:
# See, https://azure.microsoft.com/en-us/services/key-vault/
secret_value = "MY_API_KEY"
# Create a connection
project = hopsworks.login(
host="MY_INSTANCE.cloud.hopsworks.ai", # DNS of your Feature Store instance
port=443, # Port to reach your Hopsworks instance, defaults to 443
project="MY_PROJECT", # Name of your Hopsworks project
api_key_value=secret_value, # The API key to authenticate with Hopsworks
hostname_verification=True, # Disable for self-signed certificates
)
# Get the feature store handle for the project's feature store
fs = project.get_feature_store()
```
## Next Steps
For more information on how to use the Hopsworks API check out the other guides or the [Login API][hopsworks.login].
================================================================================
# Designer
Source: https://docs.hopsworks.ai/latest/user_guides/integrations/mlstudio_designer/
# Azure Machine Learning Designer Integration
Connecting to Hopsworks from the Azure Machine Learning Designer requires setting up a Hopsworks API key for the Designer and installing the **Hopsworks** Python library on the Designer.
This guide explains step by step how to connect to the Feature Store from Azure Machine Learning Designer.
!!! info "Network Connectivity"
To be able to connect to the Feature Store, please ensure that the Network Security Group of your Hopsworks instance on Azure is configured to allow incoming traffic from your compute target on ports 443, 9083 and 9085 (443,9083,9085).
See [Network security groups](https://docs.microsoft.com/en-us/azure/virtual-network/network-security-groups-overview) for more information.
If your compute target is not in the same VNet as your Hopsworks instance and the Hopsworks instance is not accessible from the internet then you will need to configure [Virtual Network Peering](https://docs.microsoft.com/en-us/azure/virtual-network/virtual-network-manage-peering).
## Generate an API key
For instructions on how to generate an API key follow this [user guide](../projects/api_key/create_api_key.md).
For the Azure ML Designer integration to work correctly make sure you add the following scopes to your API key:
1. featurestore
2. project
3. job
4. kafka
## Connect to Hopsworks
To connect to Hopsworks from the Azure Machine Learning Designer, create a new pipeline or open an existing one:
Add an Execute Python Script step
In the pipeline, add a new `Execute Python Script` step and replace the Python script from the next step:
Add the code to access the Hopsworks
!!! info "Updating the script"
Replace MY_VERSION, MY_API_KEY, MY_INSTANCE, MY_PROJECT and MY_FEATURE_GROUP with the respective values.
The major version set for MY_VERSION needs to match the major version of Hopsworks.
Check [PyPI](https://pypi.org/project/hopsworks/#history) for available releases.
You find the Hopsworks version at the bottom of the help menu in the top navigation bar
```python
import importlib.util
import os
package_name = "hopsworks"
version = "MY_VERSION"
spec = importlib.util.find_spec(package_name)
if spec is None:
import os
os.system(f"pip install %s[python]==%s" % (package_name, version))
# Put the API key into Key Vault for any production setup:
# See, https://docs.microsoft.com/en-us/azure/machine-learning/how-to-use-secrets-in-runs
# from azureml.core import Experiment, Run
# run = Run.get_context()
# secret_value = run.get_secret(name="fs-api-key")
secret_value = "MY_API_KEY"
def azureml_main(dataframe1=None, dataframe2=None):
import hopsworks
project = hopsworks.login(
host="MY_INSTANCE.cloud.hopsworks.ai", # DNS of your Hopsworks instance
port=443, # Port to reach your Hopsworks instance, defaults to 443
project="MY_PROJECT", # Name of your Hopsworks project
api_key_value=secret_value, # The API key to authenticate with Hopsworks
hostname_verification=True, # Disable for self-signed certificates
engine="python", # Choose python as engine
)
fs = project.get_feature_store() # Get the project's default feature store
return (fs.get_feature_group("MY_FEATURE_GROUP", version=1).read(),)
```
Select a compute target and save the step.
The step is now ready to use:
Select a compute target
As a next step, you have to connect the previously created `Execute Python Script` step with the next step in the pipeline.
For instance, to export the features to a CSV file, create a `Export Data` step:
Add an Export Data step
Configure the `Export Data` step to write to you data store of choice:
Configure the Export Data step
Connect the to steps by drawing a line between them:
Connect the steps
Finally, submit the pipeline and wait for it to finish:
!!! info "Performance on the first execution"
The `Execute Python Script` step can be slow when being executed for the first time as the Hopsworks library needs to be installed on the compute target.
Subsequent executions on the same compute target should use the already installed library.
Execute the pipeline
## Next Steps
For more information on how to use the Hopsworks API check out the other guides or the [Login API][hopsworks.login].
================================================================================
# Notebooks
Source: https://docs.hopsworks.ai/latest/user_guides/integrations/mlstudio_notebooks/
# Azure Machine Learning Notebooks Integration
Connecting to the Hopsworks from Azure Machine Learning Notebooks requires setting up a Hopsworks API key for Azure Machine Learning Notebooks and installing the **Hopsworks** Python library on the notebook.
This guide explains step by step how to connect to the Hopsworks from Azure Machine Learning Notebooks.
!!! info "Network Connectivity"
To be able to connect to the Feature Store, please ensure that the Network Security Group of your Hopsworks instance on Azure is configured to allow incoming traffic from your compute target on ports 443, 9083 and 9085 (443,9083,9085).
See [Network security groups](https://docs.microsoft.com/en-us/azure/virtual-network/network-security-groups-overview) for more information.
If your compute target is not in the same VNet as your Hopsworks instance and the Hopsworks instance is not accessible from the internet then you will need to configure [Virtual Network Peering](https://docs.microsoft.com/en-us/azure/virtual-network/virtual-network-manage-peering).
## Install Hopsworks Python Library
To be able to interact with Hopsworks from a Python environment you need to install the `Hopsworks` Python library.
The library is available on [PyPi](https://pypi.org/project/hopsworks/) and is installed with the `python` profile:
=== "uv"
```bash
uv pip install "hopsworks[python]~=[HOPSWORKS_VERSION]"
```
=== "pip"
```bash
pip install "hopsworks[python]~=[HOPSWORKS_VERSION]"
```
!!! attention "Python Profile"
A bare `hopsworks` install does not bring the dependencies needed to use the library from a local Python environment.
Always install with the `python` profile, `hopsworks[python]`.
!!! attention "Matching Hopsworks version"
We recommend that the major and minor version of the Python library match the major and minor version of the Hopsworks deployment.
You find the Hopsworks version at the bottom of the help menu in the top navigation bar
## Generate an API key
For instructions on how to generate an API key follow this [user guide](../projects/api_key/create_api_key.md).
For the Azure ML Notebooks integration to work correctly make sure you add the following scopes to your API key:
1. featurestore
2. project
3. job
4. kafka
## Connect from an Azure Machine Learning Notebook
To access Hopsworks from Azure Machine Learning, open a Python notebook and proceed with the following steps to install Hopsworks and connect to the Feature Store:
Connecting from an Azure Machine Learning Notebook
### Connect to Hopsworks
You are now ready to connect to Hopsworks Feature Store from the notebook:
```python
import hopsworks
# Put the API key into Key Vault for any production setup:
# See, https://docs.microsoft.com/en-us/azure/machine-learning/how-to-use-secrets-in-runs
# from azureml.core import Experiment, Run
# run = Run.get_context()
# secret_value = run.get_secret(name="fs-api-key")
secret_value = "MY_API_KEY"
# Create a connection
project = hopsworks.login(
host="MY_INSTANCE.cloud.hopsworks.ai", # DNS of your Hopsworks instance
port=443, # Port to reach your Hopsworks instance, defaults to 443
project="MY_PROJECT", # Name of your Hopsworks project
api_key_value=secret_value, # The API key to authenticate with Hopsworks
hostname_verification=True, # Disable for self-signed certificates
engine="python", # Choose Python as engine
)
# Get the feature store handle for the project's feature store
fs = project.get_feature_store()
```
## Next Steps
For more information on how to use the Hopsworks API check out the other guides or the [Login API][hopsworks.login].
================================================================================
# Apache Spark
Source: https://docs.hopsworks.ai/latest/user_guides/integrations/spark/
# Spark Integration
Connecting to the Feature Store from an external Spark cluster, such as Cloudera, requires configuring it with the Hopsworks client jars and configuration.
This guide explains step by step how to connect to the Feature Store from an external Spark cluster.
## Download the Hopsworks Client Jars
In the *Project Settings*, select the *integration* tab and scroll to the *Configure Spark Integration* section.
Click on *Download client Jars*.
This will start the download of the *client.tar.gz* archive.
The archive contains two jar files for HopsFS, the Apache Hudi jar and the Java version of the Hopsworks library.
You should upload these libraries to your Spark cluster and attach them as local resources to your Job.
If you are using `spark-submit`, you should specify the `--jar` option.
For more details see: [Spark Dependency Management](https://spark.apache.org/docs/latest/submitting-applications.html#advanced-dependency-management).
The Spark Integration gives access to Jars and configuration for an external Spark cluster
## Download the certificates
Download the certificates from the same section as above.
Hopsworks uses X.509 certificates for authentication and authorization.
If you are interested in the Hopsworks security model, you can read more about it in this [blog post](https://www.logicalclocks.com/blog/how-we-secure-your-data-with-hopsworks).
The certificates are composed of three different components: the `keyStore.jks` containing the private key and the certificate for your project user, the `trustStore.jks` containing the certificates for the Hopsworks certificates authority, and a password to unlock the private key in the `keyStore.jks`.
The password is displayed in a pop-up when downloading the certificate and should be saved in a file named `material_passwd`.
!!! warning
When you copy-paste the password to the `material_passwd` file, pay attention to not introduce additional empty spaces or new lines.
The three files (`keyStore.jks`, `trustStore.jks` and `material_passwd`) should be attached as resources to your Spark application as well.
## Configure your Spark cluster
!!! warning "Spark version limitation"
Currently Spark version 3.3.x is suggested to be able to use the full suite of Hopsworks Feature Store capabilities.
Add the following configuration to the Spark application:
```plaintext
spark.hadoop.fs.hopsfs.impl io.hops.hopsfs.client.HopsFileSystem
spark.hadoop.hops.ipc.server.ssl.enabled true
spark.hadoop.hops.ssl.hostname.verifier ALLOW_ALL
spark.hadoop.hops.rpc.socket.factory.class.default io.hops.hadoop.shaded.org.apache.hadoop.net.HopsSSLSocketFactory
spark.hadoop.client.rpc.ssl.enabled.protocol TLSv1.2
spark.hadoop.hops.ssl.keystores.passwd.name material_passwd
spark.hadoop.hops.ssl.keystore.name keyStore.jks
spark.hadoop.hops.ssl.trustore.name trustStore.jks
spark.sql.hive.metastore.jars path
spark.sql.hive.metastore.jars.path [Path to the Hopsworks Hive Jars]
spark.hadoop.hive.metastore.uris thrift://[metastore_ip]:[metastore_port]
```
`spark.sql.hive.metastore.jars.path` should point to the path with the jars from the uncompressed Hive archive you can find in *clients.tar.gz*.
## PySpark
To use PySpark, install the Hopsworks Python library which can be found on [PyPi](https://pypi.org/project/hsfs/).
!!! attention "Matching Hopsworks version"
The **major version of `Hopsworks`** needs to match the **major version of Hopsworks**.
## Generating an API Key
For instructions on how to generate an API key follow this [user guide](../projects/api_key/create_api_key.md).
For the Spark integration to work correctly make sure you add the following scopes to your API key:
1. featurestore
2. project
3. job
4. kafka
## Connecting to the Feature Store
You are now ready to connect to the Hopsworks Feature Store from Spark:
```python
import hopsworks
project = hopsworks.login(
host="my_instance", # DNS of your Feature Store instance
port=443, # Port to reach your Hopsworks instance, defaults to 443
project="my_project", # Name of your Hopsworks Feature Store project
api_key_value="api_key", # The API key to authenticate with the feature store
hostname_verification=True, # Disable for self-signed certificates
)
fs = project.get_feature_store() # Get the project's default feature store
```
!!! note "Engine"
`Hopsworks` leverages several engines depending on whether you are running using Apache Spark or Pandas/Polars.
The default behaviour of the library is to use the `spark` engine if you do not specify any `engine` option in the `login` method and if the `PySpark` library is available in the environment.
## Next Steps
For more information about how to connect, see the [Login API][hopsworks.login].
Or continue with the Data Source guide to import your own data to the Feature Store.
================================================================================
# Apache Beam
Source: https://docs.hopsworks.ai/latest/user_guides/integrations/beam/
# Apache Beam Dataflow Runner
Connecting to the Feature Store from an Apache Beam Dataflow Runner, requires configuring the Hopsworks certificates.
For this in your Beam Java application `pom.xml` file include following snippet:
```xml
java.io.tmpdir**/*.jks
```
## Generating an API Key
For instructions on how to generate an API key follow this [user guide](../projects/api_key/create_api_key.md).
For the Beam integration to work correctly make sure you add the following scopes to your API key:
1. featurestore
2. project
3. job
4. kafka
## Connecting to the Feature Store
You are now ready to connect to the Hopsworks Feature Store from Beam:
```Java
//Establish connection with Hopsworks.
HopsworksConnection hopsworksConnection = HopsworksConnection.builder()
.host("my_instance") // DNS of your Feature Store instance
.port(443) // Port to reach your Hopsworks instance, defaults to 443
.project("my_project") // Name of your Hopsworks Feature Store project
.apiKeyValue("api_key") // The API key to authenticate with the feature store
.hostnameVerification(false) // Disable for self-signed certificates
.build();
//get feature store handle
FeatureStore fs = hopsworksConnection.getFeatureStore();
```
## Next Steps
For more information and how to integrate Beam feature pipeline to the Hopsworks Feature store follow the [tutorial](https://github.com/logicalclocks/hopsworks-tutorials/tree/master/integrations/java/beam).
================================================================================
# Java
Source: https://docs.hopsworks.ai/latest/user_guides/integrations/java/
# Java client
This guide explains step by step how to connect to Hopsworks from a Java client.
## Generate an API key
For instructions on how to generate an API key follow this [user guide](../projects/api_key/create_api_key.md).
For the Java client to work correctly make sure you add the following scopes to your API key:
1. featurestore
2. project
3. job
4. kafka
## Connecting to the Feature Store
You are now ready to connect to the Hopsworks Feature Store from a Java client:
```Java
//Import necessary classes
import com.logicalclocks.hsfs.FeatureStore;
import com.logicalclocks.hsfs.FeatureView;
import com.logicalclocks.hsfs.HopsworksConnection;
//Establish connection with Hopsworks.
HopsworksConnection hopsworksConnection = HopsworksConnection.builder()
.host("my_instance") // DNS of your Feature Store instance
.port(443) // Port to reach your Hopsworks instance, defaults to 443
.project("my_project") // Name of your Hopsworks Feature Store project
.apiKeyValue("api_key") // The API key to authenticate with the feature store
.hostnameVerification(false) // Disable for self-signed certificates
.build();
//get feature store handle
FeatureStore fs = hopsworksConnection.getFeatureStore();
//get feature view handle
FeatureView fv = fs.getFeatureView(fvName, fvVersion);
// get feature vector
List