Back to WayJet

API & Docs

v1· Updated Jul 2026

140+ models · 10 model families

WayJet is an OpenAI-compatible gateway. Point any OpenAI SDK at the base URL below, use a sk- API key, and call every provider through one unified API.

10 model families

OpenAI, Anthropic, Gemini, Groq, and more behind one API.

OpenAI-compatible

Point any OpenAI SDK at the base URL — no code changes.

Streaming

Token-by-token SSE on chat completions out of the box.

Keys & budgets

Scoped keys, per-period spend limits, and usage tracking.

Call GET /v1/models with your key to list the models your gateway serves.

Introduction#

Every request goes to a single base URL. Because the gateway speaks the OpenAI wire format, existing SDKs and tools work unchanged — you only swap the base URL and key.

https://api.wayjet.ai/v1

Authentication#

Authenticate with a bearer token in the Authorization header. Create a sk--prefixed key on the API keys page — the secret is shown only once, so store it somewhere safe.

header
Authorization: Bearer sk-xxxxxxxxxxxxxxxxxxxxxxxx
Never expose a key in client-side code. Call the gateway from your server, or proxy it.
Create an API key

Quickstart#

Send your first chat completion. Set WAYJET_API_KEY to your API key and run it as is, or swap gpt-5.1 for any model id from GET /v1/models, which lists the models your key can call.

curl https://api.wayjet.ai/v1/chat/completions \
  -H "Authorization: Bearer $WAYJET_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-5.1",
    "messages": [{"role": "user", "content": "Hello"}]
  }'
Try it in the Playground

Chat completions#

POST /v1/chat/completions — the core endpoint. Requests and responses follow the OpenAI schema; the gateway adds provider and cache_status to the response.

ParameterTypeDescription
model*stringA model id from GET /v1/models (the gpt-5.1 in the examples)
messages*arrayConversation messages (role + content)
streambooleanStream tokens back as server-sent events
temperaturenumberSampling temperature, 0–2 (default 1)
max_tokensintegerMaximum tokens to generate
top_pnumberNucleus sampling probability mass
toolsarrayFunction/tool definitions the model may call
tool_choicestring | objectForce or constrain tool selection
response_formatobjecte.g. { "type": "json_object" } for JSON mode
reasoning_effortstringlow · medium · high (reasoning models)
stopstring | arrayUp to 4 stop sequences
seedintegerBest-effort deterministic sampling
providerstringGateway-only — pin the request to one provider

* required

Streaming#

Set stream: true to receive tokens as server-sent events. Each event is a data: line with a delta; the stream ends with data: [DONE].

curl https://api.wayjet.ai/v1/chat/completions \
  -H "Authorization: Bearer $WAYJET_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-5.1",
    "messages": [{"role": "user", "content": "Hello"}],
    "stream": true
  }'
# → server-sent events: lines of  data: {...}  terminated by  data: [DONE]

Prompt caching#

Several upstream providers cache the prefix of a prompt they have seen recently and charge less to reprocess it. On OpenAI, Gemini and DeepSeek this happens automatically: it needs no parameter from you, and neither you nor WayJet can turn it off. Anthropic caches only where a request asks it to.

On the platform key, WayJet only honours cache_control for accounts pinned to their own Anthropic workspace. Otherwise the marker is accepted and then removed before the request reaches Anthropic. The call still succeeds, but its usage carries no cache tokens and every input token is billed at the full input rate, which is also what Anthropic charges for it. With your own Anthropic key, the marker always goes through.

Every vendor draws the cache boundary at its own customer. On a WayJet platform key, that customer is WayJet — so on the three automatic providers, your requests share one cache namespace with other WayJet accounts for the reuse window below. The gateway responds with cache_status on chat completions, but that reports whether a prefix was reused, not whose it was.

ProviderCachingCache boundaryReuse window
OpenAIAutomaticThe organization — WayJet's, on the platform keyup to 30 min
GeminiAutomaticThe Google Cloud project — WayJet's, on the platform keyup to 24 h
DeepSeekAutomaticThe account — WayJet's, on the platform keyhours to days
AnthropicOn request (cache_control), with your own key or a pinned accountYour own account (BYOK) or your pinned workspace5 min
Claude on Bedrock / VertexOn request, own credentials onlyThe AWS account / GCP project5 min

What a shared cache does and does not mean

No prompt or completion content crosses a prompt cache. A cache hit is the vendor reprocessing a prefix the caller already sent, more cheaply — it does not hand anyone else's text to anyone. What a shared namespace does expose is timing: someone who already holds an exact prefix could, in principle, learn from a faster or cheaper response that it had been sent recently. They have to know the prefix first, which is what bounds this.

We publish this because you cannot observe it yourself. WayJet deliberately hides which upstream served a request, so nothing in the API or the dashboard would tell you where a vendor's cache boundary sits. If you resell inference to your own users, this is the answer to the same question one level down.

If you need a cache namespace of your own

Two routes, and they are the only honest ones — there is no per-tenant setting on the platform key that would do it:

  • Bring your own key. Add your own provider credentials and your traffic runs under your account at the vendor, with their cache boundary drawn at you. This works on every provider and is the complete answer.
  • Ask for an Anthropic workspace. For Anthropic specifically, we can pin your account to its own Anthropic workspace, which makes your prompt cache yours alone while you keep using the platform key. It is opt-in and enabled per account — there is no self-serve switch, and it does not apply to Claude reached through Bedrock or Vertex, which have no equivalent. It is also what turns Anthropic caching on for a platform-key account at all.

Embeddings#

POST /v1/embeddings — vectorize text for search and RAG. Accepts a string or an array of strings.

curl https://api.wayjet.ai/v1/embeddings \
  -H "Authorization: Bearer $WAYJET_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "your-embedding-model",
    "input": "The quick brown fox"
  }'

Audio (speech)#

POST /v1/audio/speech turns text into spoken audio (billed per input character), and POST /v1/audio/transcriptions turns an uploaded audio file into text (billed per audio-minute). Both are OpenAI-compatible.

Text-to-speech

curl https://api.wayjet.ai/v1/audio/speech \
  -H "Authorization: Bearer $WAYJET_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "your-tts-model",
    "input": "The quick brown fox jumped over the lazy dog.",
    "voice": "alloy"
  }' \
  --output speech.mp3

Speech-to-text

Upload the file as multipart/form-data. Set response_format to srt, vtt, or verbose_json for timestamped subtitles instead of plain text.

curl https://api.wayjet.ai/v1/audio/transcriptions \
  -H "Authorization: Bearer $WAYJET_API_KEY" \
  -F "model=your-stt-model" \
  -F "file=@audio.mp3" \
  -F "response_format=json"   # json (default) · text · srt · vtt · verbose_json

Models#

GET /v1/models returns the catalog available to your key, each with pricing and capability metadata. Use an id from this list as the model in every example on this page.

shell
curl https://api.wayjet.ai/v1/models \
  -H "Authorization: Bearer $WAYJET_API_KEY"
Browse all models

CLI tools (Claude Code, Codex)#

The gateway speaks the Anthropic Messages API (POST /v1/messages) and the OpenAI Responses API (POST /v1/responses), so the popular coding CLIs point straight at WayJet and bill through your own key — no proxy, no per-vendor SDK. Point each tool at a model you can call.

What is a squad?

A squad is a model name you call: a named product on the gateway that answers under its own name. Use a squad name anywhere a model name goes. GET /v1/models lists the ones your key can call. The model names in the configs below are squads, so they work as pasted.

Claude Code

Add to ~/.claude/settings.json. The key rides ANTHROPIC_AUTH_TOKEN (or ANTHROPIC_API_KEY as the x-api-key header), and each tier maps to a model of your choice.

~/.claude/settings.json
{
  "hasCompletedOnboarding": true,
  "env": {
    "ANTHROPIC_BASE_URL": "https://api.wayjet.ai",
    "ANTHROPIC_AUTH_TOKEN": "sk-your-key",
    "ANTHROPIC_DEFAULT_OPUS_MODEL": "claude-opus-5",
    "ANTHROPIC_DEFAULT_SONNET_MODEL": "claude-sonnet-5",
    "ANTHROPIC_DEFAULT_HAIKU_MODEL": "claude-haiku-4-5-20251001"
  }
}

Codex

Codex uses the OpenAI Responses API by default — no wire_api = "chat" override needed.

~/.codex/config.toml
model = "gpt-5.1"
model_provider = "wayjet"

[model_providers.wayjet]
name = "WayJet"
base_url = "https://api.wayjet.ai/v1"
wire_api = "responses"
~/.codex/auth.json
{
  "auth_mode": "apikey",
  "OPENAI_API_KEY": "sk-your-key"
}

Claude Code and Codex send each request to the one model they name.

Errors#

Errors use standard HTTP status codes and an OpenAI-style envelope.

json
{
  "error": {
    "message": "Incorrect API key provided.",
    "type": "authentication_error",
    "code": null,
    "param": null
  }
}
StatusTypeMeaning
200OKThe request succeeded
400invalid_request_errorMalformed request or invalid parameters
401authentication_errorMissing, invalid, or revoked API key
403permission_errorThe key is not allowed to perform this action
404not_found_errorUnknown model or resource
429rate_limit_errorRate limit hit, or a budget was exceeded
500api_errorAn unexpected gateway error
View your usage & rate limits

Endpoints#

The endpoints you can call with an API key. Keys, usage, and budgets are managed from the dashboard.

MethodEndpointDescription
POST/v1/chat/completionsChat completion (streaming + non-stream)
POST/v1/messagesAnthropic Messages format (streaming + non-stream; x-api-key or Bearer)
POST/v1/responsesOpenAI Responses format (streaming + non-stream). Stateless: previous_response_id is ignored, so send the full conversation each time
GET/v1/generationCost and token usage for one request, by its X-Wayjet-Request-Id (?id=)
POST/v1/embeddingsCreate embeddings
POST/v1/audio/speechText-to-speech (audio out)
POST/v1/audio/transcriptionsSpeech-to-text (transcription)
GET/v1/modelsList available models

Changelog#

Notable changes to the public API. The base URL stays /v1 for backward compatibility.

  • Jul 2026CLI tools: use Claude Code (/v1/messages) and Codex (/v1/responses) directly. Generation stats (/v1/generation).
  • Jul 2026Audio: text-to-speech (/v1/audio/speech) and speech-to-text (/v1/audio/transcriptions, with srt/vtt/verbose_json formats).
  • Jun 2026reasoning_effort added.
  • May 2026Embeddings endpoint + JSON mode (response_format) support.
  • Apr 2026Streaming (SSE) on chat completions; provider pinning.
Full changelog