Installing and configuring
| Symptom | Likely cause | Fix |
|---|---|---|
pip install hollerith fails | The SDK is not on PyPI. Each deployment serves its own wheel. | Install the wheel directly: pip install https://hollerith.monarcha.ai/sdk/hollerith-<version>-py3-none-any.whl. The exact filename is in the console Playground tab. |
Hollerith() raises a plain ValueError | No API origin is set, and there is no default. | Set base_url= or export HOLLERITH_BASE_URL="https://hollerith.monarcha.ai". Note: this is a bare ValueError, so except HollerithError will not catch it. |
invalid_api_key from a copied snippet | The docs and the Playground tab show the key as hk_live_… — a placeholder, not a real key. Pasted unchanged, it is sent as your key. | Replace the whole placeholder with your key from the API Keys tab. |
Requests are slow
| Symptom | Likely cause | Fix |
|---|---|---|
| The first request is slow | The single GPU worker is starting up, or has not yet claimed your job — an idle worker checks the queue every 60 seconds. | Wait — the SDK waits with you. Pass on_warming=lambda err: print("warming up") to see it happening. Blocking calls raise TimeoutError after timeout=900.0 seconds. |
predict takes as long as fit | No fitted context is bound, so predict re-sends the training table. Happens with server_context=False, fit(..., wait=False) (see wait=False skips steps), or a custom transport= without submit_fit. | Bind a context: handle.wait(), then clf = Hollerith.from_fitted_context(handle.id). |
predict is still slow with a context bound | The model reads the fitted table on every call, so prediction time scales with its size. | Fit on fewer rows when latency matters. See Introduction. |
wait=False skips steps
fit(..., wait=False) returns as soon as the job is submitted, and everything the SDK would
have done after the wait is skipped — with no warning.
| Symptom | Cause | Fix |
|---|---|---|
The next predict re-uploads the training table | The fitted context never binds. handle.wait() does not bind it either. | handle.wait(), then clf = Hollerith.from_fitted_context(handle.id). |
fit(evaluate=True) produced no evaluation | The evaluation job is silently dropped — never submitted, no error. | Fit with wait=True, then call clf.evaluate(). |
predict(quantiles=[...]) returned no DataFrame | The frame is only built after the wait. | Call .wait(), then build it from .predictions, .quantiles and .quantile_levels. |
Schema is unexpected
| Symptom | Likely cause | Fix |
|---|---|---|
| An integer column became a classification | An integer target with 20 or fewer distinct values is read as class labels — a 1–10 rating, a 0/1 flag or a small count, with no warning. | Pass task="regression". task= always wins over the inferred choice. |
predict says schema_mismatch | Your rows are missing columns fit saw. Extra columns are dropped silently; missing ones are the error. | The error’s cause field lists both column sets. An array is checked by column count only, in fit-time order. |
A forecast fails with schema_mismatch | The extra columns in your context rows must exactly match the declared covariateColumns. Usually from REST submits — the SDK derives the set from your context frame. | Declare every covariate you send. A future frame missing one fails earlier, as a client-side ValueError. |
The dataset is rejected
| Symptom | Likely cause | Fix |
|---|---|---|
dataset_too_large below 1 000 000 rows | Rows, columns and total cells each have a ceiling, and the table must clear all three. 100 000 rows × 2 000 columns breaks the 100 000 000-cell budget. | Stay under all three ceilings. See Limits. |
| The SDK rejects a dataset the deployment allows | The wheel ships its own copy of the limits. | Upgrade the wheel after your deployment’s limits are raised. |
payload_too_large arrives only after a long wait | The SDK encodes and compresses everything locally before the size check runs. | Nothing was uploaded. Reduce the payload and resubmit. |
Results look wrong
| Symptom | Likely cause | Fix |
|---|---|---|
result[0.1] raises a KeyError | Quantile column labels are strings, not floats. | Use result["0.1"]. predict(quantiles=[0.1, 0.5, 0.9]) returns columns ["prediction", "0.1", "0.5", "0.9"]. |
predict_proba changed the order of classes_ | It overwrites classes_ to match the engine’s ordering, so the two stay aligned. | Read classes_ after the call — column j matches clf.classes_[j]. |
| The holdout metric is much worse than the k-fold metric | The holdout is the last 20% of rows in file order — nothing shuffles. A file sorted by its target trains on some classes and scores on others. | Shuffle the file before evaluating. 10 000 rows or fewer use k-fold; larger tables use holdout. |
Fitted contexts
| Symptom | Likely cause | Fix |
|---|---|---|
fitted_context_expired (404) | A context expires after 30 days without a successful prediction, at most 90 days after creation. | Refit. fitted_context_expires_at_ holds the deadline. Identical active fits are reused within an organization — pass force_refit=True when you need a fresh one. |
evaluate() fails after from_fitted_context() | Evaluation needs the labeled training set, which a resumed client has never seen. It raises RuntimeError("call fit() before predict()"). | Call evaluate() from the client that ran fit. A resumed client also has no classes_ until predict_proba runs. |
Quota and retries
| Symptom | Likely cause | Fix |
|---|---|---|
quota_exceeded with quota left | Quota is reserved at submit and released at completion, so many jobs at once can claim the whole day’s allowance before any of them run. | Spread out your submits. Quota counts input rows only and resets at 00:00 UTC; a job stuck in the queue holds its reservation until then. |
| Retrying a 429 never succeeds | quota_exceeded is a 429 marked retryable: false. | Branch on retryable, not the status code. Only four codes are retryable, and at 429 only rate_limited is one. The SDK already retries it twice, waiting out the server’s Retry-After; a third rejection raises QuotaExceededError. |
Jobs that fail or vanish
| Symptom | Likely cause | Fix |
|---|---|---|
A job fails with internal_error and no detail | Two byte-identical payloads share one stored object, and the first job to finish deletes it out from under the second. | Give both calls the same idempotency_key, so the second replays the first job instead of racing it. |
A result is a 404 (training_ref_expired) an hour after it succeeded | Results are kept for 1 hour, then deleted. The job record stays. | Collect results within the hour, or re-run the job. |
worker_unavailable | No worker could take the job — it is starting, busy, or unreachable. A cold worker shows up as this too; nothing emits worker_warming_up. | Retry the call unchanged. Poll loops absorb it automatically — a fit wait only when you pass on_warming=. |
worker_oom | The dataset fit the published limits but not the GPU’s memory. | This is terminal and not retried. Reduce rows or columns and resubmit. |
Errors from the REST API
| Error | Cause | Fix |
|---|---|---|
GET /v1/predictions/{id}/result returns 404 (job_not_found) while the job runs | The result does not exist until the job succeeds. | Poll GET /v1/predictions/{id} for status, and read /result once it is succeeded. |
A bare {"error":"invalid_request"} with no code | The body matched neither of the two shapes POST /v1/predictions accepts, so the response cannot say which one failed. | Check trainRowCount and outputRowCount first — they are the fields that separate the two shapes. |
Next steps
Errors
Common errors and their causes
Limits
Learn about the constraints Hollerith is optimized for