Sable serves open and custom LLMs on a single OpenAI-compatible endpoint — sub-250ms first token, autoscaling to zero, and per-token billing. No cluster to babysit, no cold-start tax.
$ pip install sable · deploy in under 60 seconds · no card required
Serverless and warm. Swap models by changing a string — pay only for the tokens you generate. Prices below are per 1M output tokens.
Bring your own weights: upload a fine-tune or LoRA adapter and Sable serves it on the same API within 90 seconds. Benchmarks measured at 512-token prompts, us-east-1, batch 1.
A speculative-decoding runtime, continuous batching, and paged KV-cache keep GPUs saturated so your latency stays flat under load.
Drop-in OpenAI compatibility. Change the base URL and key — your existing client libraries, tools, and agents work unchanged.
Requests land on the nearest region with capacity. Speculative decoding drafts ahead while continuous batching packs the accelerator.
Paged KV-cache reuses shared prefixes across requests, so repeated system prompts and RAG contexts cost you almost nothing.
Serverless idles down between bursts. Need guaranteed headroom? Reserve dedicated H100s with per-second billing and no cold starts.
Start serverless and per-token. Graduate to reserved GPUs when your traffic is steady. No minimums, no egress fees within-region.
Per-token billing on shared capacity. Scales to zero automatically between requests.
Reserved accelerators, per-second billing, zero cold starts. A100 from $1.40/hr.
Private VPC deployments, dedicated regions, and a 99.99% contractual uptime SLA.
Six regions on NVMe-backed GPU fleets. Round-trip figures are intra-region p50, measured from a co-located client.
Zero data retention on inference by default — requests and completions are never logged or used for training. Encryption in transit and at rest, scoped API keys, and audited access on every path.
Create a key, paste it into any OpenAI client, and stream tokens in the next five minutes. $25 in credit, no card, no sales call.