AI-native cloud for builders and operators

One platform for
models and compute

Serverless model APIs, dedicated endpoints, GPU instances, and bare-metal clusters—delivered on a unified compute and inference stack.

500+Models available
80+Upstream providers
99.5%Target uptime
Model Market

Call every model through one API.
No inference infra to manage.

Access leading models via a unified, OpenAI-compatible interface. Billed by the token, production-ready.

01 · SERVERLESS API

One API for text and multimodal inference

Pay per token, not idle hours. The management plane orchestrates ServingApps while traffic goes straight to inference pods—fault domains stay isolated.

curl https://api.chariot.hk/userapi/v1/model/v1/chat/completions \
  -H "Authorization: Bearer $TOKEN" \
  -d '{"model":"anthropic/claude-opus-5","messages":[...]}'
02 · DEDICATED ENDPOINT

Private endpoints. Isolated compute. Steady latency.

Bind an Accelerator Pool and Serving Recipe for consistent throughput. Built for production ServingApps—no noisy neighbors.

BASE_URL · api.chariot.ai/ep-7f2a ● OPERATIONAL
Recipe · SGLang · Disaggregated serving H200 × 8
Compatibility · Pre-checks passed Ready
Compute Market

GPUs in seconds.
From instances to bare metal.

Built-in provisioning and scheduling across GPU servers, bare metal, serverless jobs, and multi-cloud price discovery.

1

GPU Instances

Fully managed GPU machines, ready in seconds. Docker / KVM with CUDA preinstalled—ideal for training and debugging. Prepaid or per-second billing.

2

Serverless GPU

Submit a job; we allocate accelerator pool capacity and scale to zero when done. No idle cost—built for batch inference and elastic serving.

3

Bare Metal

Physically isolated 100% compute with automated GPU drivers. IB / RoCEv2 networking for large-scale training and dedicated inference.

Why Chariot AI

Infrastructure built for AI workloads

Owned AIDC capacity plus a heterogeneous accelerator control plane—turning deployment from tribal knowledge into governed engineering.

Better price-performance

Self-owned liquid-cooled clusters and live multi-cloud pricing—target 50%+ higher useful output on the same hardware.

Production reliability

Management plane stays outside the GPU fault domain; inference traffic hits the cluster directly. Deep links into Prometheus / Grafana / Loki.

Full-stack platform

Model APIs, accelerator pools, Serving Recipes, and usage billing—connected in one operating surface.

Vendor-neutral abstraction

Accelerator inventory and compatibility decisions in one model—NVIDIA first, with room for AMD / Ascend / TPU.

Scale with demand

Grow from serverless APIs to dedicated endpoints and bare metal—with quotas, credits, and order workflows end to end.

Dedicated support

Built for operators and enterprises: modular governance across provisioning, scheduling, operations, and ops tooling.

Pricing

By token, by second, by instance—pay for what you use

Full-lifecycle token commerce with on-demand, reserved, or committed capacity—plus compute credits, vouchers, and allowances.

API

Model inference

Billed per million tokens. Example: Claude Opus 5 at $5.00 input / $25.00 output.

GPU

GPU instances

H200 from about $3.65–$4.23 / GPU / hr. Per-second billing, no minimum. Low-stock alerts and upgrade approvals when capacity tightens.

ENT

Enterprise

Private inference foundations, credit limits, invoicing, and contract billing. Talk to sales for a custom Accelerator Pool.

From your first API call to a dedicated cluster

Start with free developer credits, then scale to dedicated endpoints and bare metal. Docs, console, and Unified Compute Interface (UCI) are ready.

Start Building Read the docs
Contact

Get in touch

Docs and onboarding are available through our team. Reach out and we will help you get started.

jihao@chroit.hk