Skip to main content
Pass evaluate=True to fit() to evaluate the labeled training table. Hollerith stores the resulting score and validation method on clf.evaluation_.

Evaluation Result

evaluation_ records the metric, score and validation method used for the model call.
For classification, higher accuracy indicates better performance. For regression, RMSE is the metric used where lower indicates better performance. Classification fits return accuracy while regression fits return RMSE.

How validation works

Hollerith selects the validation method from the number of labeled rows passed to fit(). Tables with 10 000 rows or fewer use k-fold validation; larger tables use a holdout split. In k-fold validation, Hollerith divides the table into several parts, fits on all but one part, and scores the remaining part. It repeats this so each row is scored once, then combines the scores. folds records how many parts were used — usually 5, but it can be reduced toward 2 to stay within the evaluation cost limit. For larger tables, Hollerith uses a holdout split: it fits on 80% of the rows and scores the remaining 20%. In this case, folds is None. rows is the total number of labeled rows submitted for evaluation, not only the rows in the final scored fold or holdout set. For non-temporal data, shuffle the table before fitting and record the seed; holdout selection follows the uploaded row order.

Next steps

Usage and billing

Learn about Hollerith’s pricing structure

Python SDK

Complete API reference for the Hollerith SDK