# AI proxy API Call WAMP's metered AI proxy through the OpenAI-compatible HTTP API with bearer authentication, standard responses, streaming, tools, errors, and usage. Source: https://docs.vampikez.fun/reference/ai-proxy/ After this page you can call WAMP's metered models from any OpenAI-compatible client, stream a response over standard server-sent events (SSE), continue a turn after a tool call, and handle authentication, rate limits, and usage. The canonical base URL is `https://api.vampikez.fun/v1`. The public inference surface is HTTP: | Method | Endpoint | Contract | |---|---|---| | `GET` | `/v1/models` | OpenAI model-list response | | `POST` | `/v1/responses` | OpenAI Responses API request and response | | `POST` | `/v1/chat/completions` | OpenAI Chat Completions request and response | Use `/v1/responses` for new integrations. Use Chat Completions when a harness already requires that contract. Both routes use the same authentication, model catalog, metering, and error policy. Your client executes application function tools. ## Authenticate every request Send a bearer credential in the `Authorization` header. Credentials are checked on every request, so revocation and capability changes take effect without waiting for a connection to expire. | Credential | Header | Scope | |---|---|---| | User session access token | `Authorization: Bearer ` | A human user with AI permission | | App installation access token | `Authorization: Bearer ` | An installed app with `wamp.ai.invoke` in one organization | | End-user token | `Authorization: Bearer ` | One end user of an app, attributed to that user | | Runtime credential | `Authorization: Bearer ` | A short-lived credential issued for an agent runtime | Keep App signing keys on your server. Browser and mobile clients should receive a short-lived end-user token from your backend. [Your own backend](/ship/own-backend/) documents App installation and end-user token minting. An app installation must be authorized for the `wamp.ai.invoke` capability. A validly signed token without that grant is rejected. ## Discover available models The model catalog is deployment-specific. Call `/v1/models` after authentication and cache the result. Do not hardcode provider names or assume that every deployment has the same model aliases. ```bash curl -sS https://api.vampikez.fun/v1/models \ -H "Authorization: Bearer $WAMP_TOKEN" ``` The response follows the OpenAI list shape. Each model has `id`, `object`, `owned_by`, and WAMP capability metadata such as context size, output limit, tool support, input modalities, and pricing. Pass the returned `id` as `model` in a request. ## Make a Responses API request A non-streaming Responses request uses the standard `input` and `instructions` fields. The example uses a model id returned by the catalog. ```bash curl -sS https://api.vampikez.fun/v1/responses \ -H "Authorization: Bearer $WAMP_TOKEN" \ -H 'Content-Type: application/json' \ -d '{ "model": "gpt-5.6-terra", "instructions": "Answer concisely.", "input": "What is the capital of Portugal?", "max_output_tokens": 256 }' ``` The response is an OpenAI `response` object. Read the generated text from `output_text`, or inspect `output` when the response contains function calls or reasoning items. ```json { "id": "resp_…", "object": "response", "status": "completed", "model": "gpt-5.6-terra", "output_text": "Lisbon.", "output": [ { "id": "msg_…", "type": "message", "status": "completed", "role": "assistant", "content": [{ "type": "output_text", "text": "Lisbon.", "annotations": [] }] } ], "usage": { "input_tokens": 18, "input_tokens_details": { "cached_tokens": 0 }, "output_tokens": 4, "output_tokens_details": { "reasoning_tokens": 0 }, "total_tokens": 22 } } ``` The response status is `completed`, `incomplete`, or `failed`. A response that reaches `max_output_tokens` has status `incomplete` and `incomplete_details.reason` set to `max_output_tokens`. ### Request fields The proxy accepts the standard fields below. Provider capabilities still vary; when a field has no valid meaning for the selected model, the request returns a structured `400` rather than silently ignoring the field. | Responses field | Purpose | |---|---| | `model` | Required catalog model id or alias | | `input` | Required string or array of standard input items | | `instructions` | Optional system/developer instruction | | `stream` | Return SSE events instead of one JSON response | | `tools`, `tool_choice`, `parallel_tool_calls` | Function and supported hosted-tool controls | | `max_output_tokens` | Output limit | | `temperature`, `top_p`, `service_tier` | Sampling and service controls where supported | | `reasoning` | Whether the model accepts `reasoning_effort` | | `text` | Responses text-format controls where supported | | `store`, `include` | Responses persistence and encrypted reasoning controls where supported | | `user`, `safety_identifier` | Opaque end-user attribution for app-level credentials | | `prompt_cache_key` | Stable cache-affinity key | | `metadata` | String metadata echoed on the response | `previous_response_id`, `conversation`, and `background: true` are rejected. Replay the relevant input and output items explicitly for another turn. `text`, nonempty `include`, and `store: true` require a native Responses provider; they are rejected on translated provider paths. The proxy does not translate arbitrary provider-native request bodies. Use the OpenAI Responses or Chat Completions schema shown here. ## Use Chat Completions when a harness requires it Chat Completions uses the standard `messages` array and returns the standard `chat.completion` object. ```js const API = 'https://api.vampikez.fun/v1'; const response = await fetch(`${API}/chat/completions`, { method: 'POST', headers: { Authorization: `Bearer ${process.env.WAMP_TOKEN}`, 'Content-Type': 'application/json', }, body: JSON.stringify({ model: 'gpt-5.6-luna', messages: [ { role: 'system', content: 'Answer concisely.' }, { role: 'user', content: 'What is the capital of Portugal?' }, ], max_completion_tokens: 256, }), }); if (!response.ok) throw new Error(await response.text()); const completion = await response.json(); console.log(completion.choices[0].message.content); console.log(completion.usage); ``` For Chat Completions, output text is in `choices[0].message.content`. A tool turn is represented by `choices[0].message.tool_calls` and `finish_reason: "tool_calls"`. ## Stream with standard SSE Set `stream: true` and keep the HTTP response open while consuming SSE. Each SSE record has a `data:` line containing a JSON object. Do not parse proprietary frames or invent a second message protocol. Responses streaming uses named OpenAI events. A normal text stream includes `response.created`, one or more `response.output_text.delta` events, and a terminal `response.completed` event. A truncated stream ends with `response.incomplete`; a provider failure ends with `response.failed`. The terminal event contains the final response and usage. Responses streams do not use a `[DONE]` sentinel. ```js const response = await fetch('https://api.vampikez.fun/v1/responses', { method: 'POST', headers: { Authorization: `Bearer ${process.env.WAMP_TOKEN}`, 'Content-Type': 'application/json', }, body: JSON.stringify({ model: 'gpt-5.6-terra', input: 'Summarize the release in one sentence.', stream: true, }), }); if (!response.ok || !response.body) throw new Error(await response.text()); const reader = response.body.pipeThrough(new TextDecoderStream()).getReader(); let buffer = ''; let completed = false; while (true) { const { value, done } = await reader.read(); if (done) break; buffer += value; const records = buffer.split('\n\n'); buffer = records.pop() ?? ''; for (const record of records) { const event = record.match(/^event: ([^\n]+)$/m)?.[1]; const data = record.match(/^data: (.+)$/m)?.[1]; if (!data) continue; const payload = JSON.parse(data); if (event === 'response.output_text.delta') process.stdout.write(payload.delta); if (event === 'response.completed') { completed = true; console.error('\nusage:', payload.response.usage); } if (event === 'response.incomplete' || event === 'response.failed') { throw new Error(JSON.stringify(payload.response)); } } } if (!completed) throw new Error('Stream ended without response.completed'); ``` Chat Completions streaming also uses SSE, but its records contain `chat.completion.chunk` JSON objects without named event fields. Text arrives in `choices[0].delta.content`, tool calls arrive in `choices[0].delta.tool_calls`, and the final record has a finish reason. When `stream_options.include_usage` is true, an additional chunk with an empty `choices` array carries usage, followed by `data: [DONE]`. ## Continue after a tool call WAMP does not execute your application tools. The model returns a standard function call; your code validates the arguments, executes the tool, and sends the call result in the next HTTP request. With Responses, replay the original input, the returned output items, and a `function_call_output` for each function call. This executable Node example uses a local addition tool so it needs no external weather service: ```js const API = 'https://api.vampikez.fun/v1'; const model = process.env.WAMP_MODEL; if (!process.env.WAMP_TOKEN || !model) throw new Error('Set WAMP_TOKEN and WAMP_MODEL'); async function respond(body) { const response = await fetch(`${API}/responses`, { method: 'POST', headers: { Authorization: `Bearer ${process.env.WAMP_TOKEN}`, 'Content-Type': 'application/json', }, body: JSON.stringify({ model, ...body }), }); if (!response.ok) throw new Error(await response.text()); const result = await response.json(); if (result.status !== 'completed') throw new Error(JSON.stringify(result)); return result; } const tools = [{ type: 'function', name: 'add', description: 'Add two finite numbers.', parameters: { type: 'object', properties: { a: { type: 'number' }, b: { type: 'number' } }, required: ['a', 'b'], additionalProperties: false, }, }]; const input = [{ role: 'user', content: 'Use add to calculate 17 + 25.' }]; const first = await respond({ input, tools, tool_choice: 'required' }); const calls = first.output.filter((item) => item.type === 'function_call'); if (calls.length === 0) throw new Error('Expected a function call'); const results = calls.map((call) => { if (call.name !== 'add') throw new Error('Unexpected tool'); const args = JSON.parse(call.arguments); if (!args || !Number.isFinite(args.a) || !Number.isFinite(args.b)) { throw new Error('add requires finite a and b'); } const sum = args.a + args.b; if (!Number.isFinite(sum)) throw new Error('Addition overflow'); return { type: 'function_call_output', call_id: call.call_id, output: JSON.stringify({ sum }), }; }); const second = await respond({ input: [...input, ...first.output, ...results], tools, tool_choice: 'none', }); console.log(second.output_text); ``` For Chat Completions, replay the assistant message with its `tool_calls`, then append `{ role: "tool", tool_call_id, content }` before making the next request. Always validate and bound tool arguments before execution. Treat the model's arguments as untrusted input. ## Errors and retries Errors use the standard OpenAI envelope. `param` identifies the invalid field when one exists; it is `null` for request-independent failures. ```json { "error": { "message": "The model 'unknown-model' does not exist or is not available", "type": "invalid_request_error", "param": "model", "code": "model_not_found" } } ``` | HTTP status | Code | Meaning | Retry | |---|---|---|---| | `400` | `invalid_request`, `unsupported_parameter`, `invalid_tool_arguments` | The request or its tool history is invalid | No; fix the request | | `401` | `missing_api_key`, `invalid_api_key` | The bearer credential is absent, expired, revoked, or not accepted | Refresh or replace the credential | | `404` | `model_not_found` | The model is not in this deployment's catalog | Select an id from `/v1/models` | | `429` | `too_many_concurrent` | Principal concurrency is exhausted | Wait; honor `Retry-After` | | `503` | `service_unavailable` | The selected provider is not configured | Select an available model or wait for the deployment | | `504` | `upstream_timeout` | The upstream sent no first data event within 45 seconds | Retry at most once if the operation is safe | | `502` or `500` | `provider_error`, `provider_protocol_error`, `internal_error` | Upstream or server failure | Retry only when the operation is safe and the message indicates a transient failure | Streaming errors follow the selected contract. Before the first SSE event they are returned as an HTTP error. After streaming starts, Responses emits one terminal `response.failed` event; Chat Completions emits an error object and ends the stream. A Responses failure that is worth retrying carries the code OpenAI uses for it: `server_error` when the upstream failed or its stream broke, and `rate_limit_exceeded` when the upstream throttled the request. Other codes, such as `upstream_timeout`, keep the meanings in the table above. Do not retry automatically after a tool has produced an external side effect unless the operation is idempotent. ## Metering, attribution, and limits Admitted inference attempts write usage records, including provider failures. For the default `deepseek-v4-flash-0731` model, the gateway may try its hidden OpenRouter route once when the primary fails before output starts. A failover produces two Usage rows: a `FAILED` row for the primary attempt and a recorded or failed row under the OpenRouter route id. Each row has its own latency; a successful response is billed at the price of the route that served it. The response `model` remains the requested id. A request-shaped upstream 400 or a failure after output starts does not trigger failover. After authentication, validation, model, capacity, and availability refusals also appear as zero-token, unknown-cost `FAILED` rows. Their `errorCode` is `refused_`; provider failures use `upstream_429`, `upstream_5xx`, `upstream_4xx`, `upstream_timeout`, `upstream_network`, or `stream_incomplete`. Client cancellation remains `client_aborted` and missing provider usage remains `usage_missing`. Authentication failures write no row. Refusal diagnostics are limited to 120 writes per minute per principal, so sustained rejected traffic may omit excess rows. Each admitted authenticated refusal writes a structured info log with request id, route, code, status, model and actor kind. Every gateway refusal, including authentication failures and excess refusals, increments `ai_inference_refusals_total{code}`; pre-authentication failures write no refusal info log. A successful Responses response reports standard `usage` fields: - `input_tokens` includes fresh and cached input; cached input is also reported in `input_tokens_details.cached_tokens`. - `output_tokens` is the generated output; reasoning tokens are reported in `output_tokens_details.reasoning_tokens` when available. - `total_tokens` is the sum of input and output tokens. Chat Completions uses the corresponding `prompt_tokens`, `completion_tokens`, and `total_tokens` fields. WAMP's internal usage record also retains cache creation/read counts, provider cost when available, model, latency, principal, and end-user attribution. Provider cost is not added to the OpenAI response object. The deployment defaults are: | Limit | Default behavior | |---|---| | Request body | 40 MiB maximum; an oversized body returns `413` | | Concurrent requests | 64 per authenticated principal; excess returns `429` | Operators can change the concurrency default. Provider-owned rate limits may also apply; treat response status, error code, and `Retry-After` as the authority. Use `user` or `safety_identifier` as an opaque attribution id when calling with an app-level credential. For an end-user token, the signed subject is authoritative: a request field that disagrees with it is rejected, so a client cannot move usage from one end user to another. ## Cancellation and timeouts An HTTP request is canceled when the client aborts or disconnects. Use an `AbortController` with `fetch`, and set a timeout appropriate to the model's reasoning behavior: The gateway waits up to 45 seconds for the first upstream data event and 60 seconds between events. API Marketplace `gpt-5.6-terra` and `gpt-5.6-sol` get 180 seconds between events because their reasoning can arrive in one burst. Empty data chunks count; keep-alive comments and Anthropic ping events do not. An expiry before the response starts returns HTTP 504 `upstream_timeout`. After streaming starts, Responses sends `response.failed` with that code and Chat Completions sends an SSE error. WAMP's engine retries this error at most once, discarding any output the failed attempt had streamed; an external client must choose its own retry policy. ```js const controller = new AbortController(); const timeout = setTimeout(() => controller.abort(), 300_000); try { const response = await fetch('https://api.vampikez.fun/v1/responses', { method: 'POST', signal: controller.signal, headers: { Authorization: `Bearer ${process.env.WAMP_TOKEN}`, 'Content-Type': 'application/json', }, body: JSON.stringify({ model: 'gpt-5.6-terra', input: 'Reply once.' }), }); if (!response.ok) throw new Error(await response.text()); const result = await response.json(); // keep the timeout active through body consumption if (result.status !== 'completed') throw new Error(JSON.stringify(result)); console.log(result.output_text); } finally { clearTimeout(timeout); } ``` A cancellation is not a successful response. Stop consuming the stream and do not replay the request unless your application explicitly chose to retry it. Use exponential backoff with jitter for retryable `429`, `502`, and `503` responses, bounded by your product's deadline.