API reference
The chat API is OpenAI-compatible. Point an existing client at https://api.nimychat.ir/v1, use your API key, and name the agent in the model field.
Authentication
Create a key under API keys. It is shown once and stored only as a hash - if you lose it, revoke it and make another.
Authorization: Bearer ask_...A key can be scoped to a single agent. A scoped key used against another agent is refused, even one you also own.
Quick start
import OpenAI from 'openai';
const client = new OpenAI({
apiKey: process.env.AGENT_API_KEY,
baseURL: 'https://api.nimychat.ir/v1',
});
const reply = await client.chat.completions.create({
model: 'your-agent-slug',
messages: [{ role: 'user', content: 'What are your opening hours?' }],
});
console.log(reply.choices[0].message.content);Two things that differ from OpenAI
model names your agent, not a model
The agent owns its model, system prompt and knowledge. Choosing the model per request would let a caller route your spend to something you never approved, so model takes the agent's slug and the agent's own configuration decides the rest. temperature, max_tokens and similar are accepted and ignored so existing clients do not break.
Only your last user message is sent
Earlier turns are the agent's own history, held server-side. Replaying a transcript would let a caller fabricate assistant turns the agent never said, which is a well-worn way to talk a model past its instructions - so the conversation belongs to the agent, not to the request body.
Streaming
Set stream: true for server-sent events. Chunks are chat.completion.chunk objects and the stream terminates with data: [DONE], exactly as an OpenAI client expects.
const stream = await client.chat.completions.create({
model: 'your-agent-slug',
stream: true,
messages: [{ role: 'user', content: 'Summarise your refund policy.' }],
});
for await (const chunk of stream) {
process.stdout.write(chunk.choices[0]?.delta?.content ?? '');
}Endpoints
| Method | Path | Purpose |
|---|---|---|
| GET | /v1/models | Your agents with the API channel enabled, listed as models |
| POST | /v1/chat/completions | Send a message. Streaming and non-streaming |
An agent only appears here once you switch on its API channel. A key scoped to one agent sees only that agent.
Errors
Errors use OpenAI's envelope, so a client that reads error.message keeps working.
{
"error": {
"message": "Incorrect API key provided.",
"type": "invalid_request_error",
"param": null,
"code": "invalid_api_key"
}
}| Status | Code | When |
|---|---|---|
| 401 | invalid_api_key | Missing, malformed, revoked, or the account is suspended |
| 403 | invalid_api_key | Key is scoped to a different agent |
| 404 | model_not_found | No such agent, or its API channel is off |
| 400 | invalid_request_error | No user message in the request |
| 429 | request_failed | The agent's daily spend cap is exhausted |
| 502 | upstream_error | The model failed part-way through a reply |
Billing
Calls are charged in credits against your balance, on the same ledger as every other channel. The usage object is zeroed deliberately: you are not billed per token, so token counts here would let a client compute a cost that has nothing to do with what you were charged.
Usage is the honest answer to what a call cost - per agent, per model, and with the provider cost beside the charge.
Rate limits
Every response reports its own limits in x-ratelimit-* headers, so read those rather than hard-coding a number - they are configurable and differ per route.
| Route | Burst | Sustained |
|---|---|---|
| /v1/chat/completions | 10 / 10s | 120 / min |
| /v1/models | 20 / s | 300 / min |
Your agent's daily spend cap and your credit balance apply underneath, and either can stop a reply before a rate limit does.