Skip to content

hsfs.constructor.prediction_times #

PredictionTimes #

The timestamps a batch-inference read is anchored on.

Each timestamp becomes the prediction time of one row per entity: the feature values returned for it are the newest at or before it in every feature group of the view. The timestamps may be in the future, which is the point of the class.

Build one with cron, every, or of.

Example
entities = pd.DataFrame([{"country": "SE", "city": "Stockholm", "street": "Sveavagen"}])
schedule = PredictionTimes.every("daily", offset="08:00", count=7)
fv.get_batch_data(spine_df=schedule.cross(entities, event_time="date"))

timestamps property #

timestamps: list[datetime]

The resolved prediction times, timezone-aware in UTC, ascending and unique.

cron classmethod #

cron(
    expression: str,
    start: date | datetime | str | int | None = None,
    end: date | datetime | str | int | None = None,
    count: int | None = None,
    timezone: str = "UTC",
    max_horizon_days: int = DEFAULT_MAX_HORIZON_DAYS,
) -> PredictionTimes

Prediction times from a five-field cron expression.

The fields are minute hour day-of-month month day-of-week, in the Vixie cron dialect: *, an integer, a-b, a-b/n, */n, comma-separated lists, and the three-letter month and weekday names. Both 0 and 7 mean Sunday. When both day-of-month and day-of-week are restricted, a day matches if either matches. This is not the Quartz dialect used by Hopsworks job schedules, which has a seconds field and numbers Sunday as 1.

start is inclusive and defaults to now truncated to the minute; end is exclusive. Exactly one of end and count is required. A naive datetime and a date are read in timezone; an aware one is converted from its own. Expansion happens in timezone, so a daily schedule keeps its local time across a DST change.

Example
PredictionTimes.cron("0 8 * * MON-FRI", count=5)
PARAMETER DESCRIPTION
expression

The cron expression.

TYPE: str

start

First instant the schedule may fire at, inclusive.

TYPE: date | datetime | str | int | None DEFAULT: None

end

Instant the schedule stops before, exclusive.

TYPE: date | datetime | str | int | None DEFAULT: None

count

How many timestamps to produce.

TYPE: int | None DEFAULT: None

timezone

IANA time zone name the schedule is expressed in.

TYPE: str DEFAULT: 'UTC'

max_horizon_days

How far past start to search before giving up on count.

TYPE: int DEFAULT: DEFAULT_MAX_HORIZON_DAYS

RETURNS DESCRIPTION
PredictionTimes

The resolved prediction times.

RAISES DESCRIPTION
ValueError

If the expression is malformed, neither or both of end and count are given, or the schedule yields fewer than count timestamps within the horizon.

cross #

cross(spine_df: Any, event_time: str) -> Any

Cross a frame of entities with these times, one row per entity per time.

The ordering is the contract: entities in the order given, and within an entity ascending in time. A batch read returns its rows in that same order, so predictions zip back onto the frame positionally.

PARAMETER DESCRIPTION
spine_df

The entities, one row each.

TYPE: Any

event_time

Name of the feature view's event time column, which the result carries the times under.

TYPE: str

RETURNS DESCRIPTION
Any

A pandas DataFrame ready to pass as spine_df.

every classmethod #

every(
    interval: str,
    offset: str | None = None,
    start: date | datetime | str | int | None = None,
    end: date | datetime | str | int | None = None,
    count: int | None = None,
    timezone: str = "UTC",
    max_horizon_days: int = DEFAULT_MAX_HORIZON_DAYS,
) -> PredictionTimes

Prediction times from a named interval, which compiles to a cron expression.

interval is one of "hourly", "daily", "weekly" or "monthly". offset places the firing within the period and defaults to its start: a minute ("15") for hourly, "HH:MM" for daily, "DDD:HH:MM" with a three-letter weekday for weekly, and "D:HH:MM" with a day of the month for monthly. Every other parameter behaves as in cron.

Example
PredictionTimes.every("daily", offset="08:00", start=tomorrow, count=7)
PARAMETER DESCRIPTION
interval

The named interval.

TYPE: str

offset

Where in the period the schedule fires.

TYPE: str | None DEFAULT: None

start

First instant the schedule may fire at, inclusive.

TYPE: date | datetime | str | int | None DEFAULT: None

end

Instant the schedule stops before, exclusive.

TYPE: date | datetime | str | int | None DEFAULT: None

count

How many timestamps to produce.

TYPE: int | None DEFAULT: None

timezone

IANA time zone name the schedule is expressed in.

TYPE: str DEFAULT: 'UTC'

max_horizon_days

How far past start to search before giving up on count.

TYPE: int DEFAULT: DEFAULT_MAX_HORIZON_DAYS

RETURNS DESCRIPTION
PredictionTimes

The resolved prediction times.

RAISES DESCRIPTION
ValueError

If the interval or the offset is not recognized.

of classmethod #

of(
    timestamps: list[date | datetime | str | int],
    timezone: str = "UTC",
) -> PredictionTimes

Prediction times from an explicit list.

Accepts datetime, date, ISO-8601 strings and epoch milliseconds, the same unit an integer column of spine_df is read in. Naive values are read in timezone, at the precision given: an instant keeps its seconds. Use this for anything a schedule cannot express.

Example
PredictionTimes.of([datetime(2026, 9, 14, 8), datetime(2026, 9, 21, 8)])
PARAMETER DESCRIPTION
timestamps

The prediction times.

TYPE: list[date | datetime | str | int]

timezone

IANA time zone name naive values are read in.

TYPE: str DEFAULT: 'UTC'

RETURNS DESCRIPTION
PredictionTimes

The resolved prediction times, sorted ascending and deduplicated.

RAISES DESCRIPTION
ValueError

If the list is empty or contains a null.