Skip to Content
API AccessChat Completions

Chat completions

POST /v1/chat/completions Content-Type: application/json Authorization: Bearer sk-steinkauz-live-...

Request and response bodies follow the OpenAI Chat Completions API  for supported fields. For parameter reference, streaming behavior, and response structure, see the OpenAI documentation:

Steinkauz AI supports the common OpenAI fields, including sampling controls, token limits, stop sequences, penalties, function tools with strict, tool_choice, parallel_tool_calls, structured output, reasoning_effort, verbosity, user, store, and string metadata. Unsupported parameters return 400 invalid_request_error.

Non-streaming

curl -sS "$STEINKAUZ_BASE_URL/v1/chat/completions" \ -H "Content-Type: application/json" \ -H "Authorization: Bearer $STEINKAUZ_API_KEY" \ -d '{ "model": "vercel-ai-gateway/openai/gpt-5.5", "messages": [{ "role": "user", "content": "Hello 👋" }] }'

Streaming

Set "stream": true. The response is text/event-stream with chat.completion.chunk objects, ending with data: [DONE]. To receive token counts, also send "stream_options": { "include_usage": true }. Steinkauz then emits the finish chunk followed by one final chunk whose choices array is empty and whose usage object contains totals. Earlier chunks carry usage: null. See OpenAI streaming  for the event format.

curl -sS -N "$STEINKAUZ_BASE_URL/v1/chat/completions" \ -H "Content-Type: application/json" \ -H "Authorization: Bearer $STEINKAUZ_API_KEY" \ -d '{ "model": "vercel-ai-gateway/openai/gpt-5.5", "messages": [{ "role": "user", "content": "Hello 👋" }], "stream": true, "stream_options": { "include_usage": true } }'

The usage object includes prompt_tokens, completion_tokens, and total_tokens. When the provider reports them, it also includes prompt_tokens_details.cached_tokens and completion_tokens_details.reasoning_tokens. Non-streaming responses include the same usage shape directly on the completion object.

Tools

Function tools and tool_choice are supported where the underlying model and provider allow them. Behavior is consistent with the web chat.

OpenAI SDK (Python)

from openai import OpenAI client = OpenAI( api_key="sk-steinkauz-live-YOUR_SECRET", base_url="https://platform.steinkauz.ai/v1", ) completion = client.chat.completions.create( model="vercel-ai-gateway/openai/gpt-5.5", messages=[{"role": "user", "content": "Hello 👋"}], ) print(completion.choices[0].message.content)

OpenAI SDK (JavaScript)

import OpenAI from "openai"; const client = new OpenAI({ apiKey: process.env.STEINKAUZ_API_KEY, baseURL: "https://platform.steinkauz.ai/v1", }); const completion = await client.chat.completions.create({ model: "vercel-ai-gateway/openai/gpt-5.5", messages: [{ role: "user", content: "Hello 👋" }], }); console.log(completion.choices[0]?.message?.content);

Data sensitivity

Standard Completions uses the API key’s default data class. To raise sensitivity for one request, send a custom header (not a body field):

X-Steinkauz-Data-Sensitivity: D2_CONFIDENTIAL

The value cannot go below the key’s minimum floor. See API routing policy. Responses include X-Steinkauz-Data-Sensitivity and X-Steinkauz-Policy-Decision-Id.

Session and Turn ids

Completions may include optional session_id and turn_id (UUIDs) on the JSON body. Both are echoed on the non-stream completion object (next to id and usage) and on each stream chat.completion.chunk. Stock OpenAI SDKs ignore unknown fields; pass them with extra_body when you want to pin:

completion = client.chat.completions.create( model="vercel-ai-gateway/openai/gpt-5.5", messages=[{"role": "user", "content": "Hello 👋"}], extra_body={ "session_id": session_id, "turn_id": turn_id, }, ) session_id = completion.session_id turn_id = completion.turn_id

Omitting the ids still persists the transcript: ingest continues the Session whose full history is the longest prefix of messages[] among that API key’s 100 most recently updated API Sessions, or creates a new Session. A duplicate assistant row from a tool-loop persist race does not force a new Session. Tool-call assistants with content: "", content: null, or omitted content are treated as the same message, so the next POST continues the Session. Omitted ids only continue or create; they never fork. Rewind and divergence lineage (branched_from_session_id) require a named, owned session_id.

A Turn is one user-facing send. Tool-loop follow-ups that append only assistant/tool messages reuse turn_id; a new user message starts a new Turn. Each follow-up POST is one step; the step number comes from how many assistant messages follow the last user message in that POST, so two overlapping tool-loop requests cannot record the same step.

Usage tracking

Successful completions create activity records for your active organization, attributed to the API key used. Open Usage & Activity to view them alongside chat activity. You can filter by source (API), inspect model and provider, token counts, cost, duration, tool calls where applicable, and the API key label.

Usage is recorded for your active organization while you are billed by your providers directly. Optional BYOK Budgets apply when configured. A customer-owned Gateway connection still needs your Gateway API key. See Usage Budgets and Usage, costs, and activity details.

Last updated on