Skip to main content
The Python SDK uses a familiar fit-and-predict interface. Fit a labeled table once, then predict new rows, get class probabilities, evaluate accuracy, or forecast a series.
This page covers the Python client. For the same calls over HTTP from another language, see the REST API.

Install

Install the wheel shown in the console under Playground → Code:
Hollerith is not published on PyPI. The package requires Python 3.11 or 3.12.

Configure

Set your API key and deployment URL:
Hollerith() reads both values automatically. You can also pass them directly:
Both values are required. Keep API keys out of source control. See Get your API key for issuing and rotating keys.

Choose a method

Fit a table

Pass a labeled frame with target=, or features and labels separately. Passing both, or neither, raises ValueError.
The task is inferred from the target: non-numeric and boolean targets are classification, and so is an integral target with 20 or fewer distinct values.
Check clf.task_ after fitting a numeric target with few distinct values. If it came back "classification" and you wanted regression, pass task="regression".
See Limits and quotas for table and class limits.

Predict rows

predict returns a Python list with one value per input row. A DataFrame must contain every column in feature_names_in_; extra columns are dropped and the rest are reordered. An array-like must already use the fitted column order.

Class probabilities

Classification only. Columns follow classes_.
predict_proba rewrites classes_ with the engine’s ordering so the array and the labels stay aligned. Read classes_ after the call, not before.

Prediction intervals

Regression only. With quantiles= you get a pandas.DataFrame instead of a list: a prediction column holding the mean, plus one column per level.
Quantile column labels are strings. frame[0.1] raises KeyError — index with frame["0.1"].

Evaluate a fit

evaluate takes no data. It re-sends the labeled training table and the service splits and scores it, k-fold at 10,000 rows or fewer and holdout above that. The result is also stored on evaluation_. fit(..., evaluate=True) does the same thing in one call. Either way it is a separately billed job. A client resumed with from_fitted_context has no training table in memory, so it cannot evaluate. See Evaluating accuracy for how to read the number.

Forecast a time series

Forecasting uses the history directly and does not require fit(). Pass exactly one of prediction_length or future. Every other column is treated as a covariate. You get back a DataFrame with a mean column plus one string-named column per quantile, indexed by timestamp — or by (item_id, timestamp) when you forecast many series at once. Output rows are capped at 200,000, counted as prediction_length × series count.

Reuse a fitted context

A fit produces a server-side context you can predict against later, from another process or another machine.
Contexts expire after 30 inactive days, and 90 days after creation at the latest. Successful predictions push the inactivity deadline out. Once a context expires, fit again — resuming an expired one raises NotFoundError. A resumed client carries the schema, not the training data. evaluate() raises RuntimeError, and classes_ only appears once predict_proba has run.

Run without blocking

Pass wait=False to get a handle instead of a result, and poll it when you are ready.
submit(data) also returns a prediction handle without polling. fit(wait=False) returns a FitHandle with id, context, refresh(), and wait().
fit(wait=False) does not bind the completed context to the original client. To reuse that context, resume it with from_fitted_context(). evaluate=True has no effect when wait=False; call evaluate() separately on the original client if needed.

Progress and cold starts

on_progress is called with the latest job view on submit and on every poll. It is accepted by predict, predict_proba, evaluate, forecast and PredictionHandle.wait(), but not by fit. on_warming receives a ServiceUnavailableError while the SDK waits for a cold worker. The callback takes the error as its only argument. If you know the shape of the table you are about to fit, hint for capacity ahead of time:
warm is a best-effort, non-blocking hint. It reserves nothing and returns the client.

Timeouts, polling and idempotency

Methods that wait for a job accept the same three controls.

Reading CSV files

The reader supports comma, tab, semicolon, and pipe delimiters. Structural problems raise ValidationError with code malformed_csv and identify the affected line or column without including dataset contents.
Every column comes back as str. No dtype is inferred, so a numeric target read this way looks like classification. Cast the columns you need, or pass task="regression" to fit.

Handling errors

Service errors inherit from HollerithError. Catch a specific subclass when the response needs different handling:
Every HollerithError includes problem, code, cause, fix, doc_url, retryable, and request_id. Branch on code, which is stable across changes to the explanatory text. Calling errors raise standard Python exceptions such as ValueError or TypeError. A polling deadline raises TimeoutError. Rate-limited requests are retried twice, honoring Retry-After. Cold workers are retried inside the poll loop. See Errors for the full catalog of codes and which ones are retryable.

Next