Inference proxy
The proxy is the inference entry point. It mirrors the upstream OpenAI/Anthropic API surface, with your token in the URL path.
Prefer to pay per request with no token at all? See Pay per request with x402.
https://api.usepod.ai/proxy/<token>/... # Anthropic-compatiblehttps://api.usepod.ai/proxy/<token>/v1/... # OpenAI-compatibleCommon endpoints behind the proxy:
| Endpoint | Surface |
|---|---|
/v1/chat/completions | OpenAI chat completions (streaming and non-streaming) |
/v1/messages | Anthropic messages |
/v1/models | Model listing |
/v1/responses (also /responses) | OpenAI Responses API, translated onto the same routing and billing as chat completions. Streaming and tool calls work; stateless — previous_response_id is rejected, so clients resend history — and hosted tools (web_search, file_search, tool_search) are ignored. Used by Codex CLI, which speaks only this surface. |
Request headers
Section titled “Request headers”| Header | Required | Meaning |
|---|---|---|
Content-Type: application/json | yes | Standard JSON body |
X-Pod-Max-Price-Input | no | Max price per million input tokens (USDC microunits) |
X-Pod-Max-Price-Output | no | Max price per million output tokens (USDC microunits) |
X-Pod-Routing-Mode | no | auto (default), marketplace-only, or centralized-only |
X-Pod-Providers | no | Comma-separated provider names to pin this request to |
The Authorization / api_key your SDK sends is ignored — auth is the token in
the path. See Spend controls for the price headers and
Routing & matching for the routing headers.
Response headers
Section titled “Response headers”| Header | Meaning |
|---|---|
X-Balance-Remaining | Remaining token balance after this request |
X-Pod-Route | Which source served the request: marketplace, key relay, or centralized |
X-Pod-Provider-Id | The specific provider that served it — a provider name for centralized routes, a provider UUID for marketplace and key-relay routes |
Errors
Section titled “Errors”- A token with no balance is rejected before any upstream call.
- An unknown token returns
401. - In marketplace-only mode, no provider at your price returns a dedicated no-provider result rather than billing at a higher price.
Streaming responses are relayed byte-for-byte from the selected source; token usage is extracted from the stream for settlement.