Skip to content

Features and Feature Groups#

As a programmer, you can consider a feature, in machine learning, to be a variable associated with some entity that contains a value that is useful for helping train a model to solve a prediction problem. That is, the feature is just a variable with predictive power for a machine learning problem, or task.

A feature group is a table of features. Each feature group has a primary key, and optionally an event_time column (indicating when the features in that row were observed), a partition key, and foreign keys that point to the primary keys of other feature groups. These are index columns, not features: they identify and join rows, and they are excluded when you select the features for a model. A feature group stores untransformed feature data, so the same feature can be reused across models that each transform it differently.

Partitioning

The partition key determines how the feature group rows are laid out on disk, so that queries using the partition key read only the data they need. For example, if the partition key is the day and you have hundreds of days of data, a query for a given day or a range of days reads only those days from disk.

columns primary key event time partition key feature feature location_id event_time day temperature rainfall 9844-3333 2022-06-01 13:11 2022-06-01 12.45 44 6783-9832 2022-06-01 09:14 2022-06-01 22.84 5 7538-1231 2022-06-01 06:34 2022-06-01 31.04 2 row

Online and offline Storage#

Feature groups can be stored in a low-latency "online" database and/or in low cost, high throughput "offline" storage, typically a data lake or data warehouse. A feature group with an embedding column can also have a vector index, for similarity search from inference pipelines and agents.

upsert append feature_group_v1 one schema Online store latest values only · RonDB 9844-3333 temp 12.45 · rain 44 6783-9832 temp 22.84 · rain 5 the upsert replaces the row for its key Offline store full history · Delta 9844-3333 · 2020-01-01 13.12 · 55 9844-3333 · 2021-01-01 14.45 · 34 9844-3333 · 2022-01-01 12.45 · 44 9844-3333 · 2023-01-01 13.02 · 41 the append preserves history for time travel write 13.02 write 13.02

Online Storage#

By default, the online store keeps only the latest values of features for a feature group. It serves those precomputed features to models at runtime, and is backed by RonDB, a low latency, high throughput, high availability data store. By including an event_time column and a time-to-live (TTL), the online store can instead keep many rows per entity, which is what shift-right on-demand aggregations need.

Offline Storage#

The offline store stores the historical values of features for a feature group so that it may store much more data than the online store. Offline feature groups are used, typically, to create training data for models, but also to retrieve data for batch scoring of models.

In most cases, offline data is stored in Hopsworks, but through the implementation of data sources, it can reside in an external file system. The externally stored data can be managed by Hopsworks by defining ordinary feature groups or it can be used for reading only by defining External Feature Group.