AI proxy API
After this page you can call WAMP’s metered models from any OpenAI-compatible client, stream a response over standard server-sent events (SSE), continue a turn after a tool call, and handle authentication, rate limits, and usage.
The canonical base URL is https://api.vampikez.fun/v1. The public inference
surface is HTTP:
| Method | Endpoint | Contract |
|---|---|---|
GET |
/v1/models |
OpenAI model-list response |
POST |
/v1/responses |
OpenAI Responses API request and response |
POST |
/v1/chat/completions |
OpenAI Chat Completions request and response |
Use /v1/responses for new integrations. Use Chat Completions when a harness
already requires that contract. Both routes use the same authentication,
model catalog, metering, and error policy. Your client executes application
function tools.
Authenticate every request
Section titled “Authenticate every request”Send a bearer credential in the Authorization header. Credentials are
checked on every request, so revocation and capability changes take effect
without waiting for a connection to expire.
| Credential | Header | Scope |
|---|---|---|
| User session access token | Authorization: Bearer <jwt> |
A human user with AI permission |
| App installation access token | Authorization: Bearer <jwt> |
An installed app with wamp.ai.invoke in one organization |
| End-user token | Authorization: Bearer <jwt> |
One end user of an app, attributed to that user |
| Runtime credential | Authorization: Bearer <jwt> |
A short-lived credential issued for an agent runtime |
Keep App signing keys on your server. Browser and mobile clients should receive a short-lived end-user token from your backend. Your own backend documents App installation and end-user token minting.
An app installation must be authorized for the wamp.ai.invoke capability.
A validly signed token without that grant is rejected.
Discover available models
Section titled “Discover available models”The model catalog is deployment-specific. Call /v1/models after
authentication and cache the result. Do not hardcode provider names or assume
that every deployment has the same model aliases.
curl -sS https://api.vampikez.fun/v1/models \ -H "Authorization: Bearer $WAMP_TOKEN"The response follows the OpenAI list shape. Each model has id, object,
owned_by, and WAMP capability metadata such as context size, output limit,
tool support, input modalities, and pricing. Pass the returned id as
model in a request.
Make a Responses API request
Section titled “Make a Responses API request”A non-streaming Responses request uses the standard input and
instructions fields. The example uses a model id returned by the catalog.
curl -sS https://api.vampikez.fun/v1/responses \ -H "Authorization: Bearer $WAMP_TOKEN" \ -H 'Content-Type: application/json' \ -d '{ "model": "gpt-5.6-terra", "instructions": "Answer concisely.", "input": "What is the capital of Portugal?", "max_output_tokens": 256 }'The response is an OpenAI response object. Read the generated text from
output_text, or inspect output when the response contains function calls or
reasoning items.
{ "id": "resp_…", "object": "response", "status": "completed", "model": "gpt-5.6-terra", "output_text": "Lisbon.", "output": [ { "id": "msg_…", "type": "message", "status": "completed", "role": "assistant", "content": [{ "type": "output_text", "text": "Lisbon.", "annotations": [] }] } ], "usage": { "input_tokens": 18, "input_tokens_details": { "cached_tokens": 0 }, "output_tokens": 4, "output_tokens_details": { "reasoning_tokens": 0 }, "total_tokens": 22 }}The response status is completed, incomplete, or failed. A response that
reaches max_output_tokens has status incomplete and
incomplete_details.reason set to max_output_tokens.
Request fields
Section titled “Request fields”The proxy accepts the standard fields below. Provider capabilities still vary;
when a field has no valid meaning for the selected model, the request returns a
structured 400 rather than silently ignoring the field.
| Responses field | Purpose |
|---|---|
model |
Required catalog model id or alias |
input |
Required string or array of standard input items |
instructions |
Optional system/developer instruction |
stream |
Return SSE events instead of one JSON response |
tools, tool_choice, parallel_tool_calls |
Function and supported hosted-tool controls |
max_output_tokens |
Output limit |
temperature, top_p, service_tier |
Sampling and service controls where supported |
reasoning |
Whether the model accepts reasoning_effort |
text |
Responses text-format controls where supported |
store, include |
Responses persistence and encrypted reasoning controls where supported |
user, safety_identifier |
Opaque end-user attribution for app-level credentials |
prompt_cache_key |
Stable cache-affinity key |
metadata |
String metadata echoed on the response |
previous_response_id, conversation, and background: true are rejected.
Replay the relevant input and output items explicitly for another turn.
text, nonempty include, and store: true require a native Responses
provider; they are rejected on translated provider paths.
The proxy does not translate arbitrary provider-native request bodies. Use the OpenAI Responses or Chat Completions schema shown here.
Use Chat Completions when a harness requires it
Section titled “Use Chat Completions when a harness requires it”Chat Completions uses the standard messages array and returns the standard
chat.completion object.
const API = 'https://api.vampikez.fun/v1';
const response = await fetch(`${API}/chat/completions`, { method: 'POST', headers: { Authorization: `Bearer ${process.env.WAMP_TOKEN}`, 'Content-Type': 'application/json', }, body: JSON.stringify({ model: 'gpt-5.6-luna', messages: [ { role: 'system', content: 'Answer concisely.' }, { role: 'user', content: 'What is the capital of Portugal?' }, ], max_completion_tokens: 256, }),});
if (!response.ok) throw new Error(await response.text());const completion = await response.json();console.log(completion.choices[0].message.content);console.log(completion.usage);For Chat Completions, output text is in
choices[0].message.content. A tool turn is represented by
choices[0].message.tool_calls and finish_reason: "tool_calls".
Stream with standard SSE
Section titled “Stream with standard SSE”Set stream: true and keep the HTTP response open while consuming SSE. Each
SSE record has a data: line containing a JSON object. Do not parse proprietary
frames or invent a second message protocol.
Responses streaming uses named OpenAI events. A normal text stream includes
response.created, one or more response.output_text.delta events, and a
terminal response.completed event. A truncated stream ends with
response.incomplete; a provider failure ends with response.failed. The
terminal event contains the final response and usage. Responses streams do not
use a [DONE] sentinel.
const response = await fetch('https://api.vampikez.fun/v1/responses', { method: 'POST', headers: { Authorization: `Bearer ${process.env.WAMP_TOKEN}`, 'Content-Type': 'application/json', }, body: JSON.stringify({ model: 'gpt-5.6-terra', input: 'Summarize the release in one sentence.', stream: true, }),});
if (!response.ok || !response.body) throw new Error(await response.text());
const reader = response.body.pipeThrough(new TextDecoderStream()).getReader();let buffer = '';let completed = false;while (true) { const { value, done } = await reader.read(); if (done) break; buffer += value; const records = buffer.split('\n\n'); buffer = records.pop() ?? ''; for (const record of records) { const event = record.match(/^event: ([^\n]+)$/m)?.[1]; const data = record.match(/^data: (.+)$/m)?.[1]; if (!data) continue; const payload = JSON.parse(data); if (event === 'response.output_text.delta') process.stdout.write(payload.delta); if (event === 'response.completed') { completed = true; console.error('\nusage:', payload.response.usage); } if (event === 'response.incomplete' || event === 'response.failed') { throw new Error(JSON.stringify(payload.response)); } }}if (!completed) throw new Error('Stream ended without response.completed');Chat Completions streaming also uses SSE, but its records contain
chat.completion.chunk JSON objects without named event fields. Text arrives in
choices[0].delta.content, tool calls arrive in
choices[0].delta.tool_calls, and the final record has a finish reason. When
stream_options.include_usage is true, an additional chunk with an empty
choices array carries usage, followed by data: [DONE].
Continue after a tool call
Section titled “Continue after a tool call”WAMP does not execute your application tools. The model returns a standard function call; your code validates the arguments, executes the tool, and sends the call result in the next HTTP request.
With Responses, replay the original input, the returned output items, and a
function_call_output for each function call. This executable Node example
uses a local addition tool so it needs no external weather service:
const API = 'https://api.vampikez.fun/v1';const model = process.env.WAMP_MODEL;if (!process.env.WAMP_TOKEN || !model) throw new Error('Set WAMP_TOKEN and WAMP_MODEL');
async function respond(body) { const response = await fetch(`${API}/responses`, { method: 'POST', headers: { Authorization: `Bearer ${process.env.WAMP_TOKEN}`, 'Content-Type': 'application/json', }, body: JSON.stringify({ model, ...body }), }); if (!response.ok) throw new Error(await response.text()); const result = await response.json(); if (result.status !== 'completed') throw new Error(JSON.stringify(result)); return result;}
const tools = [{ type: 'function', name: 'add', description: 'Add two finite numbers.', parameters: { type: 'object', properties: { a: { type: 'number' }, b: { type: 'number' } }, required: ['a', 'b'], additionalProperties: false, },}];const input = [{ role: 'user', content: 'Use add to calculate 17 + 25.' }];const first = await respond({ input, tools, tool_choice: 'required' });const calls = first.output.filter((item) => item.type === 'function_call');if (calls.length === 0) throw new Error('Expected a function call');
const results = calls.map((call) => { if (call.name !== 'add') throw new Error('Unexpected tool'); const args = JSON.parse(call.arguments); if (!args || !Number.isFinite(args.a) || !Number.isFinite(args.b)) { throw new Error('add requires finite a and b'); } const sum = args.a + args.b; if (!Number.isFinite(sum)) throw new Error('Addition overflow'); return { type: 'function_call_output', call_id: call.call_id, output: JSON.stringify({ sum }), };});const second = await respond({ input: [...input, ...first.output, ...results], tools, tool_choice: 'none',});console.log(second.output_text);For Chat Completions, replay the assistant message with its tool_calls, then
append { role: "tool", tool_call_id, content } before making the next request.
Always validate and bound tool arguments before execution. Treat the model’s
arguments as untrusted input.
Errors and retries
Section titled “Errors and retries”Errors use the standard OpenAI envelope. param identifies the invalid field
when one exists; it is null for request-independent failures.
{ "error": { "message": "The model 'unknown-model' does not exist or is not available", "type": "invalid_request_error", "param": "model", "code": "model_not_found" }}| HTTP status | Code | Meaning | Retry |
|---|---|---|---|
400 |
invalid_request, unsupported_parameter, invalid_tool_arguments |
The request or its tool history is invalid | No; fix the request |
401 |
missing_api_key, invalid_api_key |
The bearer credential is absent, expired, revoked, or not accepted | Refresh or replace the credential |
404 |
model_not_found |
The model is not in this deployment’s catalog | Select an id from /v1/models |
429 |
too_many_concurrent |
Principal concurrency is exhausted | Wait; honor Retry-After |
503 |
service_unavailable |
The selected provider is not configured | Select an available model or wait for the deployment |
504 |
upstream_timeout |
The upstream sent no first data event within 45 seconds | Retry at most once if the operation is safe |
502 or 500 |
provider_error, provider_protocol_error, internal_error |
Upstream or server failure | Retry only when the operation is safe and the message indicates a transient failure |
Streaming errors follow the selected contract. Before the first SSE event they
are returned as an HTTP error. After streaming starts, Responses emits one
terminal response.failed event; Chat Completions emits an error object and
ends the stream. A Responses failure that is worth retrying carries the code
OpenAI uses for it: server_error when the upstream failed or its stream broke,
and rate_limit_exceeded when the upstream throttled the request. Other codes,
such as upstream_timeout, keep the meanings in the table above. Do not retry
automatically after a tool has produced an external side effect unless the
operation is idempotent.
Metering, attribution, and limits
Section titled “Metering, attribution, and limits”Admitted inference attempts write usage records, including provider failures.
For the default deepseek-v4-flash-0731 model, the gateway may try its hidden
OpenRouter route once when the primary fails before output starts. A failover
produces two Usage rows: a FAILED row for the primary attempt and a recorded
or failed row under the OpenRouter route id. Each row has its own latency; a
successful response is billed at the price of the route that served it. The
response model remains the requested id. A request-shaped upstream 400 or a
failure after output starts does not trigger failover.
After authentication, validation, model, capacity, and availability refusals
also appear as zero-token, unknown-cost FAILED rows. Their errorCode is
refused_<response code>; provider failures use upstream_429, upstream_5xx,
upstream_4xx, upstream_timeout, upstream_network, or stream_incomplete.
Client cancellation remains client_aborted and missing provider usage remains
usage_missing. Authentication failures write no row. Refusal diagnostics are
limited to 120 writes per minute per principal, so sustained rejected traffic
may omit excess rows. Each admitted authenticated refusal writes a structured
info log with request id, route, code, status, model and actor kind. Every gateway
refusal, including authentication failures and excess refusals, increments
ai_inference_refusals_total{code}; pre-authentication failures write no refusal
info log. A successful Responses response reports standard usage
fields:
input_tokensincludes fresh and cached input; cached input is also reported ininput_tokens_details.cached_tokens.output_tokensis the generated output; reasoning tokens are reported inoutput_tokens_details.reasoning_tokenswhen available.total_tokensis the sum of input and output tokens.
Chat Completions uses the corresponding prompt_tokens,
completion_tokens, and total_tokens fields. WAMP’s internal usage record
also retains cache creation/read counts, provider cost when available, model,
latency, principal, and end-user attribution. Provider cost is not added to the
OpenAI response object.
The deployment defaults are:
| Limit | Default behavior |
|---|---|
| Request body | 40 MiB maximum; an oversized body returns 413 |
| Concurrent requests | 64 per authenticated principal; excess returns 429 |
Operators can change the concurrency default. Provider-owned rate limits may
also apply; treat response status, error code, and Retry-After as the authority.
Use user or safety_identifier as an opaque attribution id when calling with
an app-level credential. For an end-user token, the signed subject is
authoritative: a request field that disagrees with it is rejected, so a client
cannot move usage from one end user to another.
Cancellation and timeouts
Section titled “Cancellation and timeouts”An HTTP request is canceled when the client aborts or disconnects. Use an
AbortController with fetch, and set a timeout appropriate to the model’s
reasoning behavior:
The gateway waits up to 45 seconds for the first upstream data event and 60
seconds between events. API Marketplace gpt-5.6-terra and gpt-5.6-sol get
180 seconds between events because their reasoning can arrive in one burst.
Empty data chunks count; keep-alive comments and Anthropic ping events do not.
An expiry before the response starts returns HTTP 504 upstream_timeout.
After streaming starts, Responses sends response.failed with that code and
Chat Completions sends an SSE error. WAMP’s engine retries this error at most
once, discarding any output the failed attempt had streamed; an external client
must choose its own retry policy.
const controller = new AbortController();const timeout = setTimeout(() => controller.abort(), 300_000);try { const response = await fetch('https://api.vampikez.fun/v1/responses', { method: 'POST', signal: controller.signal, headers: { Authorization: `Bearer ${process.env.WAMP_TOKEN}`, 'Content-Type': 'application/json', }, body: JSON.stringify({ model: 'gpt-5.6-terra', input: 'Reply once.' }), }); if (!response.ok) throw new Error(await response.text()); const result = await response.json(); // keep the timeout active through body consumption if (result.status !== 'completed') throw new Error(JSON.stringify(result)); console.log(result.output_text);} finally { clearTimeout(timeout);}A cancellation is not a successful response. Stop consuming the stream and do
not replay the request unless your application explicitly chose to retry it.
Use exponential backoff with jitter for retryable 429, 502, and 503
responses, bounded by your product’s deadline.