Chat completions
An OpenAI-compatible chat completions endpoint: point your existing client at AgentWorks.
POST /v1/chat/completions speaks the OpenAI Chat Completions format. If your code already uses an OpenAI SDK, change the base URL and the key and it works. Usage is paid from the workspace balance.
Scope: chat:write. Listing models needs chat:read.
Request
curl "$AGENTWORKS_API_URL/v1/chat/completions" \
-H "Authorization: Bearer $AGENTWORKS_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "MODEL_ID",
"messages": [
{"role": "system", "content": "You answer in one sentence."},
{"role": "user", "content": "What is a purchase order?"}
]
}'
| Field | |
|---|---|
model | Required. A model id your workspace may use; get the list from GET /v1/models. |
messages | Required. role is system, user or assistant; content is a non-empty string. |
stream | Optional. true streams the answer. |
temperature | Optional. |
Other OpenAI fields are accepted and ignored.
Response
{
"id": "chatcmpl-…",
"object": "chat.completion",
"created": 1730000000,
"model": "MODEL_ID",
"choices": [
{ "index": 0, "message": { "role": "assistant", "content": "…" }, "finish_reason": "stop" }
]
}
Streaming
With "stream": true the response is a text/event-stream of chat.completion.chunk objects, ending with data: [DONE] — the same wire format OpenAI clients expect.
With an OpenAI SDK
from openai import OpenAI
client = OpenAI(
base_url="https://api.agent-works.ai/v1",
api_key="aw_ak_…",
)
reply = client.chat.completions.create(
model="MODEL_ID",
messages=[{"role": "user", "content": "Hello"}],
)
print(reply.choices[0].message.content)
Good to know
- These calls are plain model calls. They do not appear in anyone's chat history and they do not use agents, actions or knowledge. To use those, see Ask an agent.
- Requests from a browser are refused. Call this endpoint from your server.
- Errors on
/v1use the OpenAI shape:{"error": {"type", "code", "message", "param"}}. - An empty balance answers
402before any model is called. - The limit is 300 requests per minute per workspace.