Limits & Timeouts (v2)

The numbers a client author needs to build against: request size, timeouts, pagination, batch and ingest bounds, and backpressure. These are enforced server-side โ€” design your chunking, retry, and pagination around them rather than discovering them at runtime. Values are the defaults for a standard Aito server; a managed tier may set stricter quotas (see Quotas).

When a limit produces an error, it uses the structured error envelope โ€” { "kind": "error", "data": { "code", "message" } } โ€” except the request timeout, which is enforced by the HTTP layer before the query handler runs.

Request size

  • Max request body: 10 MB (10,000,000 bytes) for every POST under /api/v1 and /api/v2, queries and data writes alike, except the v1 stream insert below. The body is read in full before the route runs; this includes /api/v2/data/{table}/batch, .../import and /api/v2/data/{table}/stream. A larger body is refused with 413 Payload Too Large (payload_too_large in the v2 error envelope), and the message names the limit.
  • POST /api/v1/data/{table}/stream has no total size limit. This v1 stream insert is shared by the v2 API on purpose: it is the route for loads bigger than one request. It reads newline-delimited JSON one object at a time and commits in chunks; each object can be up to 16 MB. It is also exempt from the request timeout, so a long upload is not cut off at 15 minutes. It streams whether the client sends a Content-Length or Transfer-Encoding: chunked.

Available since: v2.11.0. On v2.10.3 and earlier every route, including /api/v1/data/{table}/stream, is capped at 8 MiB (8,388,608 bytes), and an oversize body is answered with a 500 or a closed connection rather than a 413.

  • Response bodies are also capped at 10 MB. A query whose result would exceed that is truncated at the transport layer, so keep result sets bounded (see pagination) rather than fetching everything in one call.

To stay under the request limit: split large uploads into several requests (see Batch & ingest) or stream them to the v1 stream insert, and bound reads with limit.

Bulk loads without the request cap: use the file upload API (v1). It needs S3 storage configured on the deployment, and it does not accept tables with a primary key or identity column.

Timeouts

  • A single HTTP request has a 15-minute deadline (idle, request, and linger timeouts are all 15 minutes). A query still executing at 15 minutes has its connection closed by the server.
  • There is no separate per-query execution cap โ€” a query's server-side budget is the 15-minute request deadline together with the admission gate. Most queries return in milliseconds; the 15-minute ceiling exists for heavy analytics and ingest.

_evaluate and maxTime

_evaluate behaves differently depending on the engine of the from table:

  • On a legacy type: "table" (rep1) table, _evaluate honors maxTime (seconds): default 300 s, hard maximum 3600 s. On reaching maxTime it truncates gracefully โ€” returning the metrics computed so far plus a warnings note โ€” it does not error.
  • On a type: "collection" (rep2) table โ€” the v2 default โ€” _evaluate currently runs the full test set with no maxTime budget. It is bounded only by the 15-minute HTTP request timeout: if the test set is large enough to exceed 15 minutes on one request, the connection closes with no result.

Two ways to keep a rep2 evaluation inside the window:

  • Narrow testSource (a smaller held-out fold) so the run finishes well under 15 minutes.
  • Run it as a job (below) โ€” the way to evaluate a large test set without the 15-minute ceiling.

Long-running operations: jobs

Any v2 operation can be submitted as an asynchronous job instead of being held open on one request. This is the escape hatch for work that can exceed the 15-minute request timeout โ€” a large _evaluate, a bulk import/batch, an optimize.

POST   /api/v2/jobs/_evaluate          โ†’  201 { "id": "โ€ฆ", "path": "Evaluate", "startedAt": โ€ฆ }
GET    /api/v2/jobs                     โ†’  list running + recently-finished jobs
GET    /api/v2/jobs/{id}                โ†’  status; once done, adds "status": "ok" | "failed"
GET    /api/v2/jobs/{id}/result         โ†’  the result (byte-identical to the sync response);
                                           202 while still running
DELETE /api/v2/jobs/{id}                โ†’  cooperative cancel
  • The jobs/ prefix works in front of every v2 op โ€” query ops (jobs/_query, jobs/_predict, jobs/_evaluate, โ€ฆ) and write ops (jobs/data/{table}/batch, jobs/data/{table}/optimize, jobs/data/_delete, โ€ฆ).
  • A job's /result is byte-identical to what the synchronous endpoint would return, including the error envelope on failure โ€” a failed job reports "status": "failed" with the reason, so it never fails silently.
  • Poll GET /api/v2/jobs/{id} until a status field appears, then fetch /result. Jobs are retained for a while after finishing, then expire from the result cache (a later fetch is a 404).
  • Jobs are shared across API versions โ€” a job submitted at /api/v2/jobs is also visible and cancellable at /api/v1/jobs/{id}.
  • Cancellation is cooperative: _evaluate checks for cancellation per test row, so a DELETE stops it promptly; some other ops finish their current unit of work before stopping.

Result size & pagination

  • Default page size is 10 for every read (_search, _query, _predict, _recommend, _relate, $knn, $patterns). Set limit to change it.
  • There is no maximum-limit clamp โ€” a large limit is honored, bounded only by the 10 MB response cap. offset defaults to 0.
  • Paginate with offset + limit. The response carries offset, limit, and a real total.
  • total does not mean the same thing on every endpoint. On row endpoints it is the count of matching rows; on _predict/_recommend it is the number of distinct candidate values; on _match it is the matched-row count alongside a smaller hits array. Don't build a "showing 10 of N rows" pagination loop on total from a predict/recommend/match response.
  • SQL note: a SELECT with no LIMIT is unbounded (it fetches every matching row, subject to the 10 MB cap). Always add LIMIT in _sql.

Batch & ingest

_batch (an array of queries) is fail-fast. The sub-queries run in order within one transaction; if any one errors, the whole request returns that single error envelope โ€” there is no partial result array and no per-element error object. If query 3 of 10 fails, you get one error, not seven results and an error. Retry the whole batch once the cause is fixed.

Data writes are committed in chunks:

BoundValue
Chunk size (rows committed per flush)50,000 (raise with ?batchSize=, clamped to 250,000)
Flush window10 s
Max request body10 MB per request; no total limit (and no request timeout) on the v1 stream insert (see Request size)
Total rows in a tableunbounded across requests (subject to quotas)

Idempotency โ€” retrying a timed-out write. A write has no implicit deduplication: replaying an insert that timed out duplicates rows, unless the table declares a primaryKey and you pass ?on_conflict=update or ?on_conflict=ignore (the default is error). Define a key and an on_conflict policy before you build retry logic โ€” see Identity, keys & deduplication. Because a large ingest can approach the 15-minute request timeout, chunking it into smaller batches also makes each batch independently retryable.

Rate limiting & backpressure

  • There is no per-request rate limit by default. A token-bucket rate limiter exists but is opt-in per database/tenant (managed tiers may enable it, returning 429 with X-RateLimit-* headers).

  • Always-on backpressure is admission control. When the server's in-flight query weight exceeds its budget, or a request's projected wait exceeds ~20 s, it returns 429 with a Retry-After header. This is a back-off-and-retry signal, not a permanent failure. Heavier operations consume more of the budget:

    OperationAdmission weight
    _search, _predict, $similarity1
    _recommend, _match2
    _relate5
    _evaluate10

    So a burst of _evaluate/_relate calls hits the gate far sooner than the same number of _search calls.

  • On the public demo host, a shared key can be throttled โ€” treat a 429 as "retry with exponential backoff," and honor Retry-After.

Any client that issues concurrent requests should implement 429 + Retry-After handling. It is the one status you should expect under load even when every request is well-formed.

Field & schema limits

  • Links resolve multi-hop dotted paths on rep2 collections โ€” processor.department and processor.company.name both work, and where, select, and orderBy all traverse the full chain. (A separate, capped and off-by-default feature โ€” using deep linked attributes as prediction evidence โ€” is a different mechanism, not path resolution.)
  • Vector dimensions are fixed at schema-declaration time โ€” every vector written to a column must match the declared dimension, or the write is rejected (data.bad_request, 400).

Quotas

Row and disk quotas are off by default on a licensed self-hosted server (a table and the database as a whole are unbounded); the free Docker image is the exception, below. A managed tier may enforce them โ€” e.g. a free tier of 10,000 rows per table / 100 MB disk โ€” in which case a write past the quota fails loud with a 4xx rather than silently dropping rows. Check your plan's quotas before a bulk import.

The free Docker image is capped. The freely distributed Docker image (ghcr.io/aitohq/aito) enforces 10,000 rows per table and 50,000 rows in total. An insert past the cap returns HTTP 429 with the code row_limit_exceeded:

{
  "kind": "error",
  "data": {
    "code": "row_limit_exceeded",
    "message": "Free-tier cap reached. Get a Production License at https://aito.ai/docker",
    "limit": 10000,
    "table": "big",
    "hint": "Free-tier cap reached. Get a Production License at https://aito.ai/docker"
  }
}

A batch that would cross the cap is refused whole, and data.limit says which cap it hit (10000 per table, 50000 in total). Reads keep working at any size; the cap only stops new rows. Unlike the admission-control 429 above, this one is not transient, so a client should not retry it: branch on data.code. A production licence key, set as AITO_LICENSE_KEY on the container, lifts the cap.