> ## Documentation Index
> Fetch the complete documentation index at: https://docs.hollerith.monarcha.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Troubleshooting

Each entry is a symptom, the reason it happens, and what to do. If your error has a code,
[Errors](/reference/errors) is the reference; this page is the long tail around it.

## Installing and configuring

### Why does `pip install hollerith` not work?

The SDK is not on PyPI and is not planned to be. It is served as a wheel from the deployment
you use.

```bash theme={null}
pip install https://hollerith.monarcha.ai/sdk/hollerith-<version>-py3-none-any.whl
```

The exact filename is in the console Quickstart, and it is specific to that deployment.

### Why does `Hollerith()` raise a plain `ValueError`?

There is no default API origin. If neither `base_url=` nor `HOLLERITH_BASE_URL` is set, the
constructor raises a bare `ValueError` — not a `HollerithError`, so an `except HollerithError`
block will not catch it.

```bash theme={null}
export HOLLERITH_BASE_URL="https://hollerith.monarcha.ai"
```

### Why did I get `invalid_api_key` from a snippet I copied?

The console Quickstart shows the key as `hk_live_…` with a Unicode ellipsis, and the docs do
the same. Pasted unchanged, that string is sent as your key and fails
authentication. Replace the whole placeholder with the key from the API Keys tab.

## Speed

### Why is my first request slow?

One GPU worker serves the queue. If it is starting or rolling out, a poll returns a retryable
`503` and the SDK waits through it rather than raising. An idle worker also polls for queued
work on an interval — 60 seconds in the default configuration — so a job can sit queued that
long before it is claimed.

```python theme={null}
clf.predict(rows, on_warming=lambda err: print("warming up"))
```

Blocking calls give up at `timeout=900.0` seconds and raise `TimeoutError`.

### Why did `predict` take as long as `fit`?

Because no fitted context was bound, so `predict` re-sent the training table alongside the
rows to score. Three ways that happens:

* **`server_context=False`** on the constructor, which disables fitted contexts entirely.
* **`fit(..., wait=False)`**, which returns before the context binds. See the next entry.
* **A custom `transport=`** that does not implement `submit_fit`.

To recover a context from a `FitHandle`, build a client from its id.

```python theme={null}
handle = clf.fit(train, target="churned", wait=False)
handle.wait()
clf = Hollerith.from_fitted_context(handle.id)
```

Even with a context bound, prediction time still scales with the size of the fitted table.
The model reads it on every call. See [How Hollerith works](/concepts/how-hollerith-works).

### Why did `wait=False` drop what I asked for?

`wait=False` returns a handle the moment the job is submitted, and every step the SDK would
have run after the wait is skipped. Three things go with it, none of them warned about:

* **The fitted context is not bound.** `fitted_context_id_` stays `None`, so the next `predict`
  re-uploads the training table. `handle.wait()` does not bind it either.
* **`fit(evaluate=True)` is silently dropped.** The evaluation job is never submitted, and
  there is no error.
* **`predict(quantiles=[...])` returns no DataFrame.** Call `.wait()` and build the frame from
  `.predictions`, `.quantiles` and `.quantile_levels` yourself.

```python theme={null}
clf.fit(train, target="churned")   # wait=True
clf.evaluate()
```

## Task inference

### Why did my integer column become a classification?

An integral numeric target with 20 or fewer distinct values is read as class labels. A 1–10
rating, a 0/1 flag and a small count all become classification without a warning.

```python theme={null}
clf.fit(df, target="units_sold", task="regression")
```

`task=` always wins over the inferred choice.

### Why does `predict` say `schema_mismatch`?

The rows you passed do not carry every column `fit` saw. Extra columns are dropped silently;
missing ones are the error.

* **The `cause` field names both column sets** — the features `fit` saw, and the columns you sent.
* **An array is checked by column count only**, and assumes the order seen at fit time.

## Size and schema limits

### Why does `dataset_too_large` fire below 1,000,000 rows?

Rows, columns and their product each have a ceiling. A table has to be under all three, so
100,000 rows × 2,000 columns breaches the 100,000,000-cell budget even though neither the row
nor the column limit is reached. See [Limits and quotas](/reference/limits).

### Why does the SDK reject a dataset my deployment allows?

Client-side limits come from a snapshot bundled in the wheel, not from `GET /v1/model-limits`.
If your deployment's limits were raised, the wheel has to be upgraded before the SDK will
submit against them.

## Reading results

### Why does `result[0.1]` raise a `KeyError`?

Quantile column labels are strings, not floats. `predict(quantiles=[0.1, 0.5, 0.9])` returns
columns `["prediction", "0.1", "0.5", "0.9"]`.

```python theme={null}
result["0.1"]
```

### Why did `predict_proba` change the order of `classes_`?

It adopts the engine's class ordering and overwrites `classes_` so the two stay aligned.
Column `j` of the returned array matches `clf.classes_[j]` as read after the call, not before.

## Evaluation

### Why is my holdout metric so much worse than my k-fold metric?

Training takes the first 80% of rows in file order and the holdout is the last 20%. Nothing shuffles. A CSV sorted by
its target trains on some classes and scores on others, which makes the metric meaningless.

* **Method is chosen by row count** — 10,000 rows or fewer get k-fold, above that gets holdout.
* **Shuffle the file** before evaluating a large sorted table.

## Fitted contexts

### Why is my fitted context gone?

A context expires 7 days after the fit that created it. Using it does not extend that, and
after it lapses you get `fitted_context_expired` (404).

* **Every `fit` creates a new context.** Nothing is reused between fits.
* **`fitted_context_expires_at_`** holds the deadline for the one currently bound.

### Why can't I call `evaluate()` after `from_fitted_context()`?

Evaluation ships the labeled training set, and a client resumed from a context id has never
seen it. The call raises `RuntimeError("call fit() before predict()")`.

Such a client also has no `classes_` until `predict_proba` runs.

## Quota and retries

### Why did I get `quota_exceeded` when I have quota left?

Quota is reserved when a job is submitted and released when it completes, so a burst of
concurrent submits can exhaust the daily allowance before any of them run.

* **A job that never leaves the queue** holds its reservation for the rest of the UTC day.
* **The daily quota counts input rows only**, and resets at 00:00 UTC. Output rows bill but do
  not consume it.

### Why does retrying my 429 never succeed?

`quota_exceeded` is a 429 with `retryable: false`. Only four codes are retryable, and at 429
only `rate_limited` is one of them.

* **Branch on `retryable`**, not on the status code.
* **The SDK has no backoff-retry layer** of its own.

### Why was my upload rejected only after a long wait?

The SDK serializes and gzips the whole payload locally, then asks the control plane to
presign the upload. The size check happens at that point, so `payload_too_large` arrives
after the compression pass — no bytes are sent to storage.

## Jobs that fail or vanish

### Why did my job fail with no error detail?

Uploads are content-addressed, so two byte-identical payloads share one stored object. The
first job to finish schedules that object for deletion, and a second job still queued against
it fails.

The worker logs `staged object missing`, the job ends with no error envelope, and the SDK
raises `internal_error`. Submit both calls with the same `idempotency_key` so the second
replays the first job rather than queueing a second against the same object.

### Why is my result a 404 an hour after it succeeded?

Scored predictions and evaluation metrics are kept for 1 hour, then deleted. The job stays in
the ledger; only the result is gone, reported as `training_ref_expired` (404).

Collect results from long-lived handles inside the hour, or re-run the job.

### Why did I get `worker_unavailable`?

It is a retryable `503`, returned when no worker could accept the job. The SDK also raises it
itself when it cannot reach the API or object storage.

* **Inside a poll loop it is absorbed.** `predict`, `evaluate` and `forecast` keep waiting; a
  `fit` wait does so only when you pass `on_warming=`.
* **On the submitting call it reaches you.** Retry the call unchanged.

`worker_warming_up` is in the error contract but nothing emits it, so a cold worker arrives as
`worker_unavailable`.

### Why did my job fail with `worker_oom`?

The dataset fit the published limits but not the GPU memory available to it. This is a
terminal `500` and is not retried. Reduce rows or columns and resubmit.

### What should I quote when I ask for help?

The `requestId`. Every typed error envelope carries one, `req_` plus 16 hex characters, and
the API logs the same id server-side against the failing request.

```python theme={null}
try:
    clf.predict(rows)
except HollerithError as exc:
    print(exc.code, exc.request_id)
```

Two cases have none. A bare `invalid_json` or `invalid_request` response carries no id, and an
error the SDK raised before any request has `request_id` set to `None`.

## Working directly against the REST API

### Why does `GET /v1/predictions/{id}/result` return 404 while the job is running?

The result resource does not exist until the job succeeds, and a request for it before then
returns `job_not_found` — not a 409 or a 425. Poll `GET /v1/predictions/{id}` for status, and
read `/result` once the status is `succeeded`.

### Why did my request return `invalid_request` with no code?

A body that matches no accepted shape returns a bare `{"error":"invalid_request"}` — no
`code`, no `category`, no `requestId`. `POST /v1/predictions` accepts two shapes, and a typo
in the inline one falls through to the context-backed parser, so the response does not say
which shape failed.

Check `trainRowCount` and `outputRowCount` first; they are the fields that separate the two.

## Forecasting

### Why does my forecast fail with `schema_mismatch`?

The worker requires the set of non-`timestamp`/`target`/`item_id` keys in your context rows to
equal the declared `covariateColumns` exactly. The SDK derives that set from your context
frame, so this normally comes from submitting forecasts over REST.

A `future` frame missing a covariate is caught client-side as a `ValueError` before anything
is uploaded.

## Next

* [Errors](/reference/errors) — all 26 codes, their categories and which are retryable
* [Limits and quotas](/reference/limits) — every ceiling and the error you get at it
* [The fitted context](/concepts/fitted-context) — expiry, resuming, and skipping it by accident
* [Changelog](/help/changelog) — how versions work and what is guaranteed
