Skip to content

AI proxy API

After this page you can call WAMP’s metered models from any OpenAI-compatible client, stream a response over standard server-sent events (SSE), continue a turn after a tool call, and handle authentication, rate limits, and usage.

The canonical base URL is https://api.vampikez.fun/v1. The public inference surface is HTTP:

Method Endpoint Contract
GET /v1/models OpenAI model-list response
POST /v1/responses OpenAI Responses API request and response
POST /v1/chat/completions OpenAI Chat Completions request and response

Use /v1/responses for new integrations. Use Chat Completions when a harness already requires that contract. Both routes use the same authentication, model catalog, metering, and error policy. Your client executes application function tools.

Send a bearer credential in the Authorization header. Credentials are checked on every request, so revocation and capability changes take effect without waiting for a connection to expire.

Credential Header Scope
User session access token Authorization: Bearer <jwt> A human user with AI permission
App installation access token Authorization: Bearer <jwt> An installed app with wamp.ai.invoke in one organization
End-user token Authorization: Bearer <jwt> One end user of an app, attributed to that user
Runtime credential Authorization: Bearer <jwt> A short-lived credential issued for an agent runtime

Keep App signing keys on your server. Browser and mobile clients should receive a short-lived end-user token from your backend. Your own backend documents App installation and end-user token minting.

An app installation must be authorized for the wamp.ai.invoke capability. A validly signed token without that grant is rejected.

The model catalog is deployment-specific. Call /v1/models after authentication and cache the result. Do not hardcode provider names or assume that every deployment has the same model aliases.

Terminal window
curl -sS https://api.vampikez.fun/v1/models \
-H "Authorization: Bearer $WAMP_TOKEN"

The response follows the OpenAI list shape. Each model has id, object, owned_by, and WAMP capability metadata such as context size, output limit, tool support, input modalities, and pricing. Pass the returned id as model in a request.

A non-streaming Responses request uses the standard input and instructions fields. The example uses a model id returned by the catalog.

Terminal window
curl -sS https://api.vampikez.fun/v1/responses \
-H "Authorization: Bearer $WAMP_TOKEN" \
-H 'Content-Type: application/json' \
-d '{
"model": "gpt-5.6-terra",
"instructions": "Answer concisely.",
"input": "What is the capital of Portugal?",
"max_output_tokens": 256
}'

The response is an OpenAI response object. Read the generated text from output_text, or inspect output when the response contains function calls or reasoning items.

{
"id": "resp_…",
"object": "response",
"status": "completed",
"model": "gpt-5.6-terra",
"output_text": "Lisbon.",
"output": [
{
"id": "msg_…",
"type": "message",
"status": "completed",
"role": "assistant",
"content": [{ "type": "output_text", "text": "Lisbon.", "annotations": [] }]
}
],
"usage": {
"input_tokens": 18,
"input_tokens_details": { "cached_tokens": 0 },
"output_tokens": 4,
"output_tokens_details": { "reasoning_tokens": 0 },
"total_tokens": 22
}
}

The response status is completed, incomplete, or failed. A response that reaches max_output_tokens has status incomplete and incomplete_details.reason set to max_output_tokens.

The proxy accepts the standard fields below. Provider capabilities still vary; when a field has no valid meaning for the selected model, the request returns a structured 400 rather than silently ignoring the field.

Responses field Purpose
model Required catalog model id or alias
input Required string or array of standard input items
instructions Optional system/developer instruction
stream Return SSE events instead of one JSON response
tools, tool_choice, parallel_tool_calls Function and supported hosted-tool controls
max_output_tokens Output limit
temperature, top_p, service_tier Sampling and service controls where supported
reasoning Whether the model accepts reasoning_effort
text Responses text-format controls where supported
store, include Responses persistence and encrypted reasoning controls where supported
user, safety_identifier Opaque end-user attribution for app-level credentials
prompt_cache_key Stable cache-affinity key
metadata String metadata echoed on the response

previous_response_id, conversation, and background: true are rejected. Replay the relevant input and output items explicitly for another turn. text, nonempty include, and store: true require a native Responses provider; they are rejected on translated provider paths.

The proxy does not translate arbitrary provider-native request bodies. Use the OpenAI Responses or Chat Completions schema shown here.

Use Chat Completions when a harness requires it

Section titled “Use Chat Completions when a harness requires it”

Chat Completions uses the standard messages array and returns the standard chat.completion object.

const API = 'https://api.vampikez.fun/v1';
const response = await fetch(`${API}/chat/completions`, {
method: 'POST',
headers: {
Authorization: `Bearer ${process.env.WAMP_TOKEN}`,
'Content-Type': 'application/json',
},
body: JSON.stringify({
model: 'gpt-5.6-luna',
messages: [
{ role: 'system', content: 'Answer concisely.' },
{ role: 'user', content: 'What is the capital of Portugal?' },
],
max_completion_tokens: 256,
}),
});
if (!response.ok) throw new Error(await response.text());
const completion = await response.json();
console.log(completion.choices[0].message.content);
console.log(completion.usage);

For Chat Completions, output text is in choices[0].message.content. A tool turn is represented by choices[0].message.tool_calls and finish_reason: "tool_calls".

Set stream: true and keep the HTTP response open while consuming SSE. Each SSE record has a data: line containing a JSON object. Do not parse proprietary frames or invent a second message protocol.

Responses streaming uses named OpenAI events. A normal text stream includes response.created, one or more response.output_text.delta events, and a terminal response.completed event. A truncated stream ends with response.incomplete; a provider failure ends with response.failed. The terminal event contains the final response and usage. Responses streams do not use a [DONE] sentinel.

const response = await fetch('https://api.vampikez.fun/v1/responses', {
method: 'POST',
headers: {
Authorization: `Bearer ${process.env.WAMP_TOKEN}`,
'Content-Type': 'application/json',
},
body: JSON.stringify({
model: 'gpt-5.6-terra',
input: 'Summarize the release in one sentence.',
stream: true,
}),
});
if (!response.ok || !response.body) throw new Error(await response.text());
const reader = response.body.pipeThrough(new TextDecoderStream()).getReader();
let buffer = '';
let completed = false;
while (true) {
const { value, done } = await reader.read();
if (done) break;
buffer += value;
const records = buffer.split('\n\n');
buffer = records.pop() ?? '';
for (const record of records) {
const event = record.match(/^event: ([^\n]+)$/m)?.[1];
const data = record.match(/^data: (.+)$/m)?.[1];
if (!data) continue;
const payload = JSON.parse(data);
if (event === 'response.output_text.delta') process.stdout.write(payload.delta);
if (event === 'response.completed') {
completed = true;
console.error('\nusage:', payload.response.usage);
}
if (event === 'response.incomplete' || event === 'response.failed') {
throw new Error(JSON.stringify(payload.response));
}
}
}
if (!completed) throw new Error('Stream ended without response.completed');

Chat Completions streaming also uses SSE, but its records contain chat.completion.chunk JSON objects without named event fields. Text arrives in choices[0].delta.content, tool calls arrive in choices[0].delta.tool_calls, and the final record has a finish reason. When stream_options.include_usage is true, an additional chunk with an empty choices array carries usage, followed by data: [DONE].

WAMP does not execute your application tools. The model returns a standard function call; your code validates the arguments, executes the tool, and sends the call result in the next HTTP request.

With Responses, replay the original input, the returned output items, and a function_call_output for each function call. This executable Node example uses a local addition tool so it needs no external weather service:

const API = 'https://api.vampikez.fun/v1';
const model = process.env.WAMP_MODEL;
if (!process.env.WAMP_TOKEN || !model) throw new Error('Set WAMP_TOKEN and WAMP_MODEL');
async function respond(body) {
const response = await fetch(`${API}/responses`, {
method: 'POST',
headers: {
Authorization: `Bearer ${process.env.WAMP_TOKEN}`,
'Content-Type': 'application/json',
},
body: JSON.stringify({ model, ...body }),
});
if (!response.ok) throw new Error(await response.text());
const result = await response.json();
if (result.status !== 'completed') throw new Error(JSON.stringify(result));
return result;
}
const tools = [{
type: 'function',
name: 'add',
description: 'Add two finite numbers.',
parameters: {
type: 'object',
properties: { a: { type: 'number' }, b: { type: 'number' } },
required: ['a', 'b'],
additionalProperties: false,
},
}];
const input = [{ role: 'user', content: 'Use add to calculate 17 + 25.' }];
const first = await respond({ input, tools, tool_choice: 'required' });
const calls = first.output.filter((item) => item.type === 'function_call');
if (calls.length === 0) throw new Error('Expected a function call');
const results = calls.map((call) => {
if (call.name !== 'add') throw new Error('Unexpected tool');
const args = JSON.parse(call.arguments);
if (!args || !Number.isFinite(args.a) || !Number.isFinite(args.b)) {
throw new Error('add requires finite a and b');
}
const sum = args.a + args.b;
if (!Number.isFinite(sum)) throw new Error('Addition overflow');
return {
type: 'function_call_output',
call_id: call.call_id,
output: JSON.stringify({ sum }),
};
});
const second = await respond({
input: [...input, ...first.output, ...results],
tools,
tool_choice: 'none',
});
console.log(second.output_text);

For Chat Completions, replay the assistant message with its tool_calls, then append { role: "tool", tool_call_id, content } before making the next request. Always validate and bound tool arguments before execution. Treat the model’s arguments as untrusted input.

Errors use the standard OpenAI envelope. param identifies the invalid field when one exists; it is null for request-independent failures.

{
"error": {
"message": "The model 'unknown-model' does not exist or is not available",
"type": "invalid_request_error",
"param": "model",
"code": "model_not_found"
}
}
HTTP status Code Meaning Retry
400 invalid_request, unsupported_parameter, invalid_tool_arguments The request or its tool history is invalid No; fix the request
401 missing_api_key, invalid_api_key The bearer credential is absent, expired, revoked, or not accepted Refresh or replace the credential
404 model_not_found The model is not in this deployment’s catalog Select an id from /v1/models
429 too_many_concurrent Principal concurrency is exhausted Wait; honor Retry-After
503 service_unavailable The selected provider is not configured Select an available model or wait for the deployment
504 upstream_timeout The upstream sent no first data event within 45 seconds Retry at most once if the operation is safe
502 or 500 provider_error, provider_protocol_error, internal_error Upstream or server failure Retry only when the operation is safe and the message indicates a transient failure

Streaming errors follow the selected contract. Before the first SSE event they are returned as an HTTP error. After streaming starts, Responses emits one terminal response.failed event; Chat Completions emits an error object and ends the stream. A Responses failure that is worth retrying carries the code OpenAI uses for it: server_error when the upstream failed or its stream broke, and rate_limit_exceeded when the upstream throttled the request. Other codes, such as upstream_timeout, keep the meanings in the table above. Do not retry automatically after a tool has produced an external side effect unless the operation is idempotent.

Admitted inference attempts write usage records, including provider failures. For the default deepseek-v4-flash-0731 model, the gateway may try its hidden OpenRouter route once when the primary fails before output starts. A failover produces two Usage rows: a FAILED row for the primary attempt and a recorded or failed row under the OpenRouter route id. Each row has its own latency; a successful response is billed at the price of the route that served it. The response model remains the requested id. A request-shaped upstream 400 or a failure after output starts does not trigger failover.

After authentication, validation, model, capacity, and availability refusals also appear as zero-token, unknown-cost FAILED rows. Their errorCode is refused_<response code>; provider failures use upstream_429, upstream_5xx, upstream_4xx, upstream_timeout, upstream_network, or stream_incomplete. Client cancellation remains client_aborted and missing provider usage remains usage_missing. Authentication failures write no row. Refusal diagnostics are limited to 120 writes per minute per principal, so sustained rejected traffic may omit excess rows. Each admitted authenticated refusal writes a structured info log with request id, route, code, status, model and actor kind. Every gateway refusal, including authentication failures and excess refusals, increments ai_inference_refusals_total{code}; pre-authentication failures write no refusal info log. A successful Responses response reports standard usage fields:

  • input_tokens includes fresh and cached input; cached input is also reported in input_tokens_details.cached_tokens.
  • output_tokens is the generated output; reasoning tokens are reported in output_tokens_details.reasoning_tokens when available.
  • total_tokens is the sum of input and output tokens.

Chat Completions uses the corresponding prompt_tokens, completion_tokens, and total_tokens fields. WAMP’s internal usage record also retains cache creation/read counts, provider cost when available, model, latency, principal, and end-user attribution. Provider cost is not added to the OpenAI response object.

The deployment defaults are:

Limit Default behavior
Request body 40 MiB maximum; an oversized body returns 413
Concurrent requests 64 per authenticated principal; excess returns 429

Operators can change the concurrency default. Provider-owned rate limits may also apply; treat response status, error code, and Retry-After as the authority.

Use user or safety_identifier as an opaque attribution id when calling with an app-level credential. For an end-user token, the signed subject is authoritative: a request field that disagrees with it is rejected, so a client cannot move usage from one end user to another.

An HTTP request is canceled when the client aborts or disconnects. Use an AbortController with fetch, and set a timeout appropriate to the model’s reasoning behavior:

The gateway waits up to 45 seconds for the first upstream data event and 60 seconds between events. API Marketplace gpt-5.6-terra and gpt-5.6-sol get 180 seconds between events because their reasoning can arrive in one burst. Empty data chunks count; keep-alive comments and Anthropic ping events do not. An expiry before the response starts returns HTTP 504 upstream_timeout. After streaming starts, Responses sends response.failed with that code and Chat Completions sends an SSE error. WAMP’s engine retries this error at most once, discarding any output the failed attempt had streamed; an external client must choose its own retry policy.

const controller = new AbortController();
const timeout = setTimeout(() => controller.abort(), 300_000);
try {
const response = await fetch('https://api.vampikez.fun/v1/responses', {
method: 'POST',
signal: controller.signal,
headers: {
Authorization: `Bearer ${process.env.WAMP_TOKEN}`,
'Content-Type': 'application/json',
},
body: JSON.stringify({ model: 'gpt-5.6-terra', input: 'Reply once.' }),
});
if (!response.ok) throw new Error(await response.text());
const result = await response.json(); // keep the timeout active through body consumption
if (result.status !== 'completed') throw new Error(JSON.stringify(result));
console.log(result.output_text);
} finally {
clearTimeout(timeout);
}

A cancellation is not a successful response. Stop consuming the stream and do not replay the request unless your application explicitly chose to retry it. Use exponential backoff with jitter for retryable 429, 502, and 503 responses, bounded by your product’s deadline.