Feature Store Guides#
Feature pipelines write to feature groups, training and inference pipelines read through feature views. The guides below follow that order: connect a source, write, validate, read, transform.
-
Start here
Create a feature group from a DataFrame and insert it. Everything else in this section builds on a feature group that exists.
fg = fs.get_or_create_feature_group( name="transactions", version=1, primary_key=["tid"], event_time="datetime", online_enabled=True, ) fg.insert(df)Create a feature group · Create a feature view · Training data
Write
- Data sources Register warehouses, object stores and databases to read from and write to.
- Feature groups Create, insert, evolve the schema, set time to live, deprecate.
- External and spine groups Point at data that stays where it is, or supply keys and labels without storing them.
- Ingest with dltHub Load from hundreds of sources through dlt pipelines.
Trust
- Statistics Descriptive statistics on every insert, configurable per group.
- Data validation Great Expectations suites run on insert, with a policy on failure.
- Feature monitoring Scheduled statistics and drift detection against a reference window.
- Notifications and observability Change notifications and online ingestion status.
Read
- Feature views Select features across groups and read them the same way for training and inference.
- Training data Materialise splits as files or read them straight into memory.
- Batch and online reads Batch inference data by time range, single vectors from the online store.
- Feature server Serve feature vectors over REST without the Python client.
Transform and run
- Transformation functions Model-independent functions applied on write, model-dependent on read.
- Compute engines Which operations run on Python, Spark or Flink.
- Client integrations Databricks, SageMaker, EMR, Azure ML, Flink, Beam and more.
- Vector similarity search Embeddings in a feature group, nearest-neighbour queries.