Chat completions

An OpenAI-compatible chat completions endpoint: point your existing client at AgentWorks.

POST /v1/chat/completions speaks the OpenAI Chat Completions format. If your code already uses an OpenAI SDK, change the base URL and the key and it works. Usage is paid from the workspace balance.

Scope: chat:write. Listing models needs chat:read.

Request

curl "$AGENTWORKS_API_URL/v1/chat/completions" \
  -H "Authorization: Bearer $AGENTWORKS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "MODEL_ID",
    "messages": [
      {"role": "system", "content": "You answer in one sentence."},
      {"role": "user", "content": "What is a purchase order?"}
    ]
  }'
Field
modelRequired. A model id your workspace may use; get the list from GET /v1/models.
messagesRequired. role is system, user or assistant; content is a non-empty string.
streamOptional. true streams the answer.
temperatureOptional.

Other OpenAI fields are accepted and ignored.

Response

{
  "id": "chatcmpl-…",
  "object": "chat.completion",
  "created": 1730000000,
  "model": "MODEL_ID",
  "choices": [
    { "index": 0, "message": { "role": "assistant", "content": "…" }, "finish_reason": "stop" }
  ]
}

Streaming

With "stream": true the response is a text/event-stream of chat.completion.chunk objects, ending with data: [DONE] — the same wire format OpenAI clients expect.

With an OpenAI SDK

from openai import OpenAI

client = OpenAI(
    base_url="https://api.agent-works.ai/v1",
    api_key="aw_ak_…",
)
reply = client.chat.completions.create(
    model="MODEL_ID",
    messages=[{"role": "user", "content": "Hello"}],
)
print(reply.choices[0].message.content)

Good to know

  • These calls are plain model calls. They do not appear in anyone's chat history and they do not use agents, actions or knowledge. To use those, see Ask an agent.
  • Requests from a browser are refused. Call this endpoint from your server.
  • Errors on /v1 use the OpenAI shape: {"error": {"type", "code", "message", "param"}}.
  • An empty balance answers 402 before any model is called.
  • The limit is 300 requests per minute per workspace.