API & Docs
v1· Updated Jul 2026140+ models · 10 model families
WayJet is an OpenAI-compatible gateway. Point any OpenAI SDK at the base URL below, use a sk- API key, and call every provider through one unified API.
10 model families
OpenAI, Anthropic, Gemini, Groq, and more behind one API.
OpenAI-compatible
Point any OpenAI SDK at the base URL — no code changes.
Streaming
Token-by-token SSE on chat completions out of the box.
Keys & budgets
Scoped keys, per-period spend limits, and usage tracking.
Call GET /v1/models with your key to list the models your gateway serves.
Introduction#
Every request goes to a single base URL. Because the gateway speaks the OpenAI wire format, existing SDKs and tools work unchanged — you only swap the base URL and key.
https://api.wayjet.ai/v1Authentication#
Authenticate with a bearer token in the Authorization header. Create a sk--prefixed key on the API keys page — the secret is shown only once, so store it somewhere safe.
Authorization: Bearer sk-xxxxxxxxxxxxxxxxxxxxxxxx
Quickstart#
Send your first chat completion. Set WAYJET_API_KEY to your API key and run it as is, or swap gpt-5.1 for any model id from GET /v1/models, which lists the models your key can call.
curl https://api.wayjet.ai/v1/chat/completions \
-H "Authorization: Bearer $WAYJET_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-5.1",
"messages": [{"role": "user", "content": "Hello"}]
}'Chat completions#
POST /v1/chat/completions — the core endpoint. Requests and responses follow the OpenAI schema; the gateway adds provider and cache_status to the response.
| Parameter | Type | Description |
|---|---|---|
| model* | string | A model id from GET /v1/models (the gpt-5.1 in the examples) |
| messages* | array | Conversation messages (role + content) |
| stream | boolean | Stream tokens back as server-sent events |
| temperature | number | Sampling temperature, 0–2 (default 1) |
| max_tokens | integer | Maximum tokens to generate |
| top_p | number | Nucleus sampling probability mass |
| tools | array | Function/tool definitions the model may call |
| tool_choice | string | object | Force or constrain tool selection |
| response_format | object | e.g. { "type": "json_object" } for JSON mode |
| reasoning_effort | string | low · medium · high (reasoning models) |
| stop | string | array | Up to 4 stop sequences |
| seed | integer | Best-effort deterministic sampling |
| provider | string | Gateway-only — pin the request to one provider |
* required
Streaming#
Set stream: true to receive tokens as server-sent events. Each event is a data: line with a delta; the stream ends with data: [DONE].
curl https://api.wayjet.ai/v1/chat/completions \
-H "Authorization: Bearer $WAYJET_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-5.1",
"messages": [{"role": "user", "content": "Hello"}],
"stream": true
}'
# → server-sent events: lines of data: {...} terminated by data: [DONE]Prompt caching#
Several upstream providers cache the prefix of a prompt they have seen recently and charge less to reprocess it. On OpenAI, Gemini and DeepSeek this happens automatically: it needs no parameter from you, and neither you nor WayJet can turn it off. Anthropic caches only where a request asks it to.
On the platform key, WayJet only honours cache_control for accounts pinned to their own Anthropic workspace. Otherwise the marker is accepted and then removed before the request reaches Anthropic. The call still succeeds, but its usage carries no cache tokens and every input token is billed at the full input rate, which is also what Anthropic charges for it. With your own Anthropic key, the marker always goes through.
Every vendor draws the cache boundary at its own customer. On a WayJet platform key, that customer is WayJet — so on the three automatic providers, your requests share one cache namespace with other WayJet accounts for the reuse window below. The gateway responds with cache_status on chat completions, but that reports whether a prefix was reused, not whose it was.
| Provider | Caching | Cache boundary | Reuse window |
|---|---|---|---|
| OpenAI | Automatic | The organization — WayJet's, on the platform key | up to 30 min |
| Gemini | Automatic | The Google Cloud project — WayJet's, on the platform key | up to 24 h |
| DeepSeek | Automatic | The account — WayJet's, on the platform key | hours to days |
| Anthropic | On request (cache_control), with your own key or a pinned account | Your own account (BYOK) or your pinned workspace | 5 min |
| Claude on Bedrock / Vertex | On request, own credentials only | The AWS account / GCP project | 5 min |
What a shared cache does and does not mean
No prompt or completion content crosses a prompt cache. A cache hit is the vendor reprocessing a prefix the caller already sent, more cheaply — it does not hand anyone else's text to anyone. What a shared namespace does expose is timing: someone who already holds an exact prefix could, in principle, learn from a faster or cheaper response that it had been sent recently. They have to know the prefix first, which is what bounds this.
We publish this because you cannot observe it yourself. WayJet deliberately hides which upstream served a request, so nothing in the API or the dashboard would tell you where a vendor's cache boundary sits. If you resell inference to your own users, this is the answer to the same question one level down.
If you need a cache namespace of your own
Two routes, and they are the only honest ones — there is no per-tenant setting on the platform key that would do it:
- Bring your own key. Add your own provider credentials and your traffic runs under your account at the vendor, with their cache boundary drawn at you. This works on every provider and is the complete answer.
- Ask for an Anthropic workspace. For Anthropic specifically, we can pin your account to its own Anthropic workspace, which makes your prompt cache yours alone while you keep using the platform key. It is opt-in and enabled per account — there is no self-serve switch, and it does not apply to Claude reached through Bedrock or Vertex, which have no equivalent. It is also what turns Anthropic caching on for a platform-key account at all.
Embeddings#
POST /v1/embeddings — vectorize text for search and RAG. Accepts a string or an array of strings.
curl https://api.wayjet.ai/v1/embeddings \
-H "Authorization: Bearer $WAYJET_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "your-embedding-model",
"input": "The quick brown fox"
}'Audio (speech)#
POST /v1/audio/speech turns text into spoken audio (billed per input character), and POST /v1/audio/transcriptions turns an uploaded audio file into text (billed per audio-minute). Both are OpenAI-compatible.
Text-to-speech
curl https://api.wayjet.ai/v1/audio/speech \
-H "Authorization: Bearer $WAYJET_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "your-tts-model",
"input": "The quick brown fox jumped over the lazy dog.",
"voice": "alloy"
}' \
--output speech.mp3Speech-to-text
Upload the file as multipart/form-data. Set response_format to srt, vtt, or verbose_json for timestamped subtitles instead of plain text.
curl https://api.wayjet.ai/v1/audio/transcriptions \ -H "Authorization: Bearer $WAYJET_API_KEY" \ -F "model=your-stt-model" \ -F "file=@audio.mp3" \ -F "response_format=json" # json (default) · text · srt · vtt · verbose_json
Models#
GET /v1/models returns the catalog available to your key, each with pricing and capability metadata. Use an id from this list as the model in every example on this page.
curl https://api.wayjet.ai/v1/models \ -H "Authorization: Bearer $WAYJET_API_KEY"
CLI tools (Claude Code, Codex)#
The gateway speaks the Anthropic Messages API (POST /v1/messages) and the OpenAI Responses API (POST /v1/responses), so the popular coding CLIs point straight at WayJet and bill through your own key — no proxy, no per-vendor SDK. Point each tool at a model you can call.
What is a squad?
A squad is a model name you call: a named product on the gateway that answers under its own name. Use a squad name anywhere a model name goes. GET /v1/models lists the ones your key can call. The model names in the configs below are squads, so they work as pasted.
Claude Code
Add to ~/.claude/settings.json. The key rides ANTHROPIC_AUTH_TOKEN (or ANTHROPIC_API_KEY as the x-api-key header), and each tier maps to a model of your choice.
{
"hasCompletedOnboarding": true,
"env": {
"ANTHROPIC_BASE_URL": "https://api.wayjet.ai",
"ANTHROPIC_AUTH_TOKEN": "sk-your-key",
"ANTHROPIC_DEFAULT_OPUS_MODEL": "claude-opus-5",
"ANTHROPIC_DEFAULT_SONNET_MODEL": "claude-sonnet-5",
"ANTHROPIC_DEFAULT_HAIKU_MODEL": "claude-haiku-4-5-20251001"
}
}Codex
Codex uses the OpenAI Responses API by default — no wire_api = "chat" override needed.
model = "gpt-5.1" model_provider = "wayjet" [model_providers.wayjet] name = "WayJet" base_url = "https://api.wayjet.ai/v1" wire_api = "responses"
{
"auth_mode": "apikey",
"OPENAI_API_KEY": "sk-your-key"
}Claude Code and Codex send each request to the one model they name.
Errors#
Errors use standard HTTP status codes and an OpenAI-style envelope.
{
"error": {
"message": "Incorrect API key provided.",
"type": "authentication_error",
"code": null,
"param": null
}
}| Status | Type | Meaning |
|---|---|---|
| 200 | OK | The request succeeded |
| 400 | invalid_request_error | Malformed request or invalid parameters |
| 401 | authentication_error | Missing, invalid, or revoked API key |
| 403 | permission_error | The key is not allowed to perform this action |
| 404 | not_found_error | Unknown model or resource |
| 429 | rate_limit_error | Rate limit hit, or a budget was exceeded |
| 500 | api_error | An unexpected gateway error |
Endpoints#
The endpoints you can call with an API key. Keys, usage, and budgets are managed from the dashboard.
| Method | Endpoint | Description |
|---|---|---|
| POST | /v1/chat/completions | Chat completion (streaming + non-stream) |
| POST | /v1/messages | Anthropic Messages format (streaming + non-stream; x-api-key or Bearer) |
| POST | /v1/responses | OpenAI Responses format (streaming + non-stream). Stateless: previous_response_id is ignored, so send the full conversation each time |
| GET | /v1/generation | Cost and token usage for one request, by its X-Wayjet-Request-Id (?id=) |
| POST | /v1/embeddings | Create embeddings |
| POST | /v1/audio/speech | Text-to-speech (audio out) |
| POST | /v1/audio/transcriptions | Speech-to-text (transcription) |
| GET | /v1/models | List available models |
Changelog#
Notable changes to the public API. The base URL stays /v1 for backward compatibility.
- Jul 2026CLI tools: use Claude Code (/v1/messages) and Codex (/v1/responses) directly. Generation stats (/v1/generation).
- Jul 2026Audio: text-to-speech (/v1/audio/speech) and speech-to-text (/v1/audio/transcriptions, with srt/vtt/verbose_json formats).
- Jun 2026reasoning_effort added.
- May 2026Embeddings endpoint + JSON mode (response_format) support.
- Apr 2026Streaming (SSE) on chat completions; provider pinning.

