Ablitron / HomePROJECT DOCUMENTATION

Built around
your intent.

Connect a text chat-completions client to Ablitron. Public access is in private preview; this documentation describes the implemented gateway.

Models

ablitron-core is the lower-cost access option. ablitron-pro provides access to a modified GLM-5.3 model. These are service aliases, not claims that Ablitron trained the underlying models. Small live tests verified tool-call responses and streaming on both routes; coding quality remains to be benchmarked.

1. Create an API key

When registration is enabled, open the console, register and verify your email. Create a key, copy it once and keep it in your server environment. Keys can be revoked from the same panel. Add prepaid or monthly credit when billing is enabled.

2. Send a request

Production base URL after deployment: https://ablitron.com/v1. For local development use http://localhost:4174/v1. Store your key in ABLITRON_API_KEY.

curl https://ablitron.com/v1/chat/completions \
  -H "Authorization: Bearer $ABLITRON_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"ablitron-core","messages":[{"role":"user","content":"Write a Python binary search."}],"max_tokens":2048}'

3. Stream or call tools

Add "stream": true to receive server-sent events with completion chunks, final usage and a [DONE] marker. Function definitions use the tools array; return tool results in messages with role tool and the matching tool_call_id. The client executes tools; Ablitron does not execute your code.

Input and output limits

The preview supports text messages only, at most 256 messages and 64 tools. The output limit is 1–16,384 tokens, default 2,048. Use either max_tokens or max_completion_tokens, not both. Input uses a conservative size limit of 65,536 estimated tokens including template overhead; this is not the full advertised model context window. Only one completion per request is supported.

Credit and billing

Requests reserve credit before inference, using conservative input size and the maximum requested output. The final charge uses reported token usage. Cached input is a subset of input, charged at its lower rate. Reasoning tokens are already included in output and are not charged twice. Monthly credit is spent first; refunds from a request keep the original credit expiry.

Errors and retries

400
Invalid or unsupported request. Check the message and output limits.
401
Missing, revoked or inactive key.
402
Insufficient available credit for the reservation. Add credit or lower the output limit.
429
Rate limit or four requests already in flight or pending review.
502
Model response or usage could not be confirmed. Check your console before retrying.
503
Inference or another required service is not enabled.

Save the X-Request-ID response header when reporting an issue. Inference is not automatically retried. If usage cannot be confirmed after a disconnect or server error, the reservation stays pending for operator reconciliation. A retry is a new request and may incur additional usage.

Data handled by the gateway

Requests are forwarded to an inference provider. The local usage ledger stores request IDs, timestamps, model aliases, token counts and charges; it does not store prompt or output text. Customer API keys are stored as hashes. Provider-side handling and applicable service terms must be reviewed before public launch; no end-to-end zero-retention claim is made.