TomdacatAI
API Reference

TomdacatAI API docs

OpenAI-compatible chat & image API for Roblox games and beyond. Drop in your key and start generating — works with Claude, GPT, Gemini, and dozens more, all through one endpoint.

Get API Key
Powered by

Base URL

ai.tomdacat.com

Protocol

HTTPS only

Format

JSON

Auth

Bearer token

Step 1

Getting started

Three steps to your first AI response.

01

Get your API key

Log in with Discord and visit the Dashboard to copy your personal API key (starts with tdc_).

02

Send a request

POST to /api/chat with your key in the Authorization header and a messages array in the body.

03

Read the reply

The response is OpenAI-compatible — read choices[0].message.content and you're done.

Auth

Authentication

All requests require a Bearer token sent in the Authorization header.

Header
Authorization: Bearer tdc_YOUR_API_KEY

Keep your key secret

For Roblox games, store the key in a server-side Script only — never in a LocalScript or client code. Anyone who can read a LocalScript can steal your key.
Auth

Password-protected keys

API keys can optionally require a password. When enabled, every request must include a second header alongside Authorization.

Required headers for a password-protected key:

Headers
Authorization: Bearer tdc_YOUR_API_KEY
X-Api-Key-Password: your_password

Full curl example:

Shell
curl -X POST https://ai.tomdacat.com/api/chat \
  -H "Authorization: Bearer tdc_YOUR_API_KEY" \
  -H "X-Api-Key-Password: your_password" \
  -H "Content-Type: application/json" \
  -d '{"model":"gpt-5-mini","messages":[{"role":"user","content":"Hello"}]}'

Node.js / fetch:

JavaScript
const res = await fetch('https://ai.tomdacat.com/api/chat', {
  method: 'POST',
  headers: {
    'Authorization': 'Bearer tdc_YOUR_API_KEY',
    'X-Api-Key-Password': 'your_password',
    'Content-Type': 'application/json',
  },
  body: JSON.stringify({ model: 'gpt-5-mini', messages: [...] }),
});

Roblox Luau:

Luau
local HttpService = game:GetService("HttpService")

local API_KEY      = "tdc_YOUR_API_KEY"
local KEY_PASSWORD = "your_password"  -- store in ServerScriptService only

local result = HttpService:RequestAsync({
  Url    = "https://ai.tomdacat.com/api/chat",
  Method = "POST",
  Headers = {
    ["Authorization"]    = "Bearer " .. API_KEY,
    ["X-Api-Key-Password"] = KEY_PASSWORD,
    ["Content-Type"]     = "application/json",
  },
  Body = HttpService:JSONEncode({
    model    = "gpt-5-mini",
    messages = {{ role = "user", content = "Hello" }},
  }),
})

Error responses

KEY_PASSWORD_REQUIREDX-Api-Key-Password header missing

{ "error": "Password required for this API key", "code": "KEY_PASSWORD_REQUIRED" }

KEY_PASSWORD_INCORRECTHeader present but password is wrong

{ "error": "Invalid API key password", "code": "KEY_PASSWORD_INCORRECT" }

Password is one-way hashed

The password is stored using scrypt — it cannot be recovered from the dashboard. If you lose it, delete the key and create a new one. Use the “Generate” button in the dashboard to create a strong random password that is automatically copied to your clipboard.
Endpoint

Chat completions

OpenAI-compatible chat completions. Use the standard /v1/chat/completions path with any OpenAI SDK, or the /api/chat alias.

POSThttps://ai.tomdacat.com/v1/chat/completions

OpenAI-standard path. Set base_url="https://ai.tomdacat.com/v1" in any OpenAI SDK or tool. Supports streaming, tools, vision, and stream: true/false in one endpoint.

POSThttps://ai.tomdacat.com/api/chat

Tomdacat alias — non-streaming only. Kept for backward compatibility.

Request

Request body

JSON body with the following fields.

modelreq
string

Model ID to use. See the Models section for all available IDs (e.g. "gpt-5-mini").

messagesreq
array

Conversation history as an ordered array of message objects. Each object must have role ("system" | "user" | "assistant") and content (string).

messages[].rolereq
"system" | "user" | "assistant"

Who sent the message. Use "system" for the NPC personality prompt, "user" for player input, "assistant" for previous AI replies.

messages[].contentreq
string | array

Text string, or an array of content parts for vision. Each part is { type: "text", text: "..." } or { type: "image_url", image_url: { url: "..." } }.

temperature
number

Sampling temperature (0–2). Lower = more deterministic. Optional.

max_tokens
number

Maximum tokens to generate in the response. Optional.

reasoning_effort
"minimal" | "low" | "medium" | "high"

Controls how much the model thinks before answering. Use with reasoning-capable models (e.g. deepseek-v4-flash, minimax-m2-5). Optional.

tools
array

Array of function tool definitions for tool calling. Each entry: { type: "function", function: { name, description?, parameters? } }. See Tool calling section. Optional.

tool_choice
string | object

"auto" (default), "none", "required", or { type: "function", function: { name } } to force a specific tool. Optional.

Feature

Vision

Send images alongside text by using an array of content parts. Supports public image URLs and base64 data URIs.

Supported image formats

PNG, JPEG, GIF, WebP — via HTTPS URL or data:image/png;base64,... URI. Models with the vision capability badge support this.

Using an image URL

cURL
curl -X POST https://ai.tomdacat.com/api/chat \
  -H "Authorization: Bearer tdc_YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-5-mini",
    "messages": [
      {
        "role": "user",
        "content": [
          { "type": "text", "text": "What is in this image?" },
          { "type": "image_url", "image_url": { "url": "https://example.com/screenshot.png" } }
        ]
      }
    ]
  }'

Using base64 (Node.js)

JavaScript
// Convert an image file to base64 and send it
const fs = require('fs');

const imageBytes = fs.readFileSync('screenshot.png');
const base64 = imageBytes.toString('base64');
const dataUri = `data:image/png;base64,${base64}`;

const response = await fetch('https://ai.tomdacat.com/api/chat', {
  method: 'POST',
  headers: {
    'Authorization': 'Bearer tdc_YOUR_API_KEY',
    'Content-Type': 'application/json',
  },
  body: JSON.stringify({
    model: 'gpt-5-mini',
    messages: [
      {
        role: 'user',
        content: [
          { type: 'text', text: 'Describe what you see.' },
          { type: 'image_url', image_url: { url: dataUri } },
        ],
      },
    ],
  }),
});

const data = await response.json();
console.log(data.choices[0].message.content);

Roblox — sending an image URL

Luau
-- Send an image URL for vision analysis (server Script)
local HttpService = game:GetService("HttpService")
local API_KEY = "tdc_YOUR_API_KEY"
local API_URL = "https://ai.tomdacat.com/api/chat"

local function analyzeImage(imageUrl, question)
  local ok, result = pcall(HttpService.RequestAsync, HttpService, {
    Url    = API_URL,
    Method = "POST",
    Headers = {
      ["Authorization"] = "Bearer " .. API_KEY,
      ["Content-Type"]  = "application/json",
    },
    Body = HttpService:JSONEncode({
      model = "gpt-5-mini",
      messages = {
        {
          role = "user",
          content = {
            { type = "text",      text = question },
            { type = "image_url", image_url = { url = imageUrl } },
          },
        },
      },
    }),
  })

  if not ok or not result.Success then
    warn("[Vision] error:", ok and result.StatusCode or result)
    return nil
  end

  return HttpService:JSONDecode(result.Body).choices[1].message.content
end

-- Example: analyze a screenshot URL
-- local description = analyzeImage("https://example.com/map.png", "Describe this map.")
-- print(description)
Feature

Video input

Send a video alongside text using a video_url content part — same shape as image_url, one level up. Only models with the video capability accept this.

No models currently support video input

The video_url content part is supported by the API and works end-to-end, but no model on Tomdacat currently advertises the video capability — sending it to any model today returns a 400 error. This section documents the format for when a video-capable model is added.

Using a video URL

cURL
curl -X POST https://ai.tomdacat.com/api/chat \
  -H "Authorization: Bearer tdc_YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "your-video-capable-model",
    "messages": [
      {
        "role": "user",
        "content": [
          { "type": "text", "text": "What is happening in this clip?" },
          { "type": "video_url", "video_url": { "url": "https://example.com/clip.mp4" } }
        ]
      }
    ]
  }'

Using base64 (Node.js)

JavaScript
// Convert a video file to base64 and send it
const fs = require('fs');

const videoBytes = fs.readFileSync('clip.mp4');
const base64 = videoBytes.toString('base64');
const dataUri = `data:video/mp4;base64,${base64}`;

const response = await fetch('https://ai.tomdacat.com/api/chat', {
  method: 'POST',
  headers: {
    'Authorization': 'Bearer tdc_YOUR_API_KEY',
    'Content-Type': 'application/json',
  },
  body: JSON.stringify({
    model: 'your-video-capable-model',
    messages: [
      {
        role: 'user',
        content: [
          { type: 'text', text: 'Summarize what happens in this video.' },
          { type: 'video_url', video_url: { url: dataUri } },
        ],
      },
    ],
  }),
});

const data = await response.json();
console.log(data.choices[0].message.content);
Feature

Reasoning

Some models think step-by-step before answering. Whether that's on by default and how to control it varies by model family.

Reasoning models

Models with the reasoning capability can think before they answer, but the default and the field it shows up in both depend on the model family — see below. The reasoning_effort parameter ("minimal" | "low" | "medium" | "high") is the one that reliably moves the needle across model families.
cURL
curl -X POST https://ai.tomdacat.com/api/chat \
  -H "Authorization: Bearer tdc_YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "claude-opus-4-7",
    "reasoning_effort": "high",
    "messages": [
      { "role": "system", "content": "You are a game logic solver." },
      { "role": "user",   "content": "Design an optimal pathfinding algorithm for my Roblox RPG map with dynamic obstacles." }
    ]
  }'

Reading the model's thinking

DeepSeek and GLM models think by default and return it as a plain reasoning_content string on the message (and on each streamed delta) — readable chain of thought, separate from the final content. Claude models are the opposite: thinking is off by default, and raising reasoning_effort turns it on as a content_blocks array entry of { "type": "thinking" } — its presence tells you the model reasoned, but the thinking text itself comes back empty (redacted by the provider; only an opaque signature is included). Either way, we pass the response through exactly as the model provider sends it and bill reasoning as output tokens, the same as any other generated text.

Reducing or disabling reasoning

reasoning_effort: "minimal" is the one we've verified actually works: on deepseek-v4-flash-special and glm-5-2 it drops reasoning_content from the response entirely and cuts completion tokens sharply.

cURL
curl -X POST https://ai.tomdacat.com/v1/chat/completions \
  -H "Authorization: Bearer tdc_YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deepseek-v4-flash-special",
    "reasoning_effort": "minimal",
    "messages": [
      { "role": "user", "content": "What is 17 * 24?" }
    ]
  }'

DeepSeek's own API also documents a top-level thinking field for this — we forward it unchanged, exactly as sent:

Python
from openai import OpenAI

client = OpenAI(
    base_url="https://ai.tomdacat.com/v1",
    api_key="tdc_YOUR_API_KEY",
)

response = client.chat.completions.create(
    model="deepseek-v4-flash",
    messages=[{"role": "user", "content": "Hello!"}],
    extra_body={"thinking": {"type": "disabled"}},
)

thinking doesn't work on any DeepSeek route we offer

We tested thinking: { "type": "disabled" } against both deepseek-v4-flash and deepseek-v4-flash-special — it has no effect on either; reasoning_content keeps coming back unchanged. Use reasoning_effort instead. Separately, deepseek-v4-flash itself (unlike -special) doesn't respond to reasoning_effort either — as of this writing there's no known way to reduce its reasoning. This is a limitation on the model provider's side, not something configurable from here; we'll update this note if that changes.
Feature

Tool calling

Let the model call your own functions. The model decides when to call a tool, you run it and return the result, then the model produces a final reply.

How it works — 4 steps

  1. Send tools array describing your functions with the first message.
  2. If the model wants to call a function it returns finish_reason: "tool_calls" with a tool_calls array on the assistant message.
  3. Run the function yourself and add a role: "tool" message containing the result.
  4. Call the API again — the model reads the result and produces the final reply.

Supported models

claude-opus-5claude-opus-4-7claude-opus-4-6claude-sonnet-5claude-opus-5-5claude-sonnet-4-6claude-haiku-4-5claude-haiku-4-5-freegemini-3-1-pro-previewgemini-3-flashgemini-3-7-flashgemini-3-5-flashgemini-3-6-flashgemini-3-8-flashgemini-2-5-flash-litegemini-3-5-flash-litegemma-4-26bgemma-4-26b-rawgrok-4-2-reasoninggrok-4-2grok-4-3grok-4-5grok-4-6gpt-6-astragpt-5-5gpt-5-6-solgpt-5-6-terragpt-5-6-lunagpt-6-solgpt-6-lunagpt-5-chat-latestgpt-5-4gpt-5-4-minigpt-5-minigpt-5-mini-specialgpt-4o-minigpt-5-nanogpt-5-nano-specialgpt-5-4-nanogpt-5-4-nano-specialgpt-oss-20bdeepseek-v4-flashdeepseek-v4-flash-specialdeepseek-v4-prominimax-m3nemotron-lightning-3-5-30bqwen-3-8-maxkimi-k2-6kimi-k2-7-codekimi-k3glm-5-2glm-5-2-specialglm-5-3glm-5-3-flashnemotron-3-ultra-550bnemotron-3-super-120bnemotron-3-ultra-550b-nvidiadiffusiongemma-4-26bgpt-oss-120b-fastgpt-oss-20b-fastqwen-3-8-27bqwen-3-6-27bmimo-v2.5mimo-v2.5-proqwen-largemercury-2ling-3-0-flashceleris-1

Additional request fields

tools
array

Array of tool definitions. Each entry: { type: "function", function: { name, description?, parameters? } }. parameters follows JSON Schema.

tool_choice
"none" | "auto" | "required" | object

"auto" (default) lets the model decide. "required" forces a tool call. "none" disables tools. Pass { type: "function", function: { name } } to force a specific tool.

messages[].role = "tool"
"tool"

Role for tool result messages. Must include tool_call_id matching the id from the assistant's tool_calls entry, and content with the JSON-encoded result.

messages[].tool_call_id
string

The id from the assistant tool_calls entry this result belongs to.

Tool call response (step 2)

JSON
{
  "choices": [
    {
      "finish_reason": "tool_calls",
      "message": {
        "role": "assistant",
        "content": null,
        "tool_calls": [
          {
            "id": "call_abc123",
            "type": "function",
            "function": {
              "name": "get_player_stats",
              "arguments": "{\"username\":\"PlayerOne\"}"
            }
          }
        ]
      }
    }
  ]
}

Full roundtrip examples

cURL
# ── Step 1: send your tool definitions ───────────────────────────────────
curl -X POST https://ai.tomdacat.com/api/chat \
  -H "Authorization: Bearer tdc_YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-5-mini",
    "tools": [
      {
        "type": "function",
        "function": {
          "name": "get_player_stats",
          "description": "Return level, health and gold for a Roblox player.",
          "parameters": {
            "type": "object",
            "properties": {
              "username": { "type": "string", "description": "Roblox username" }
            },
            "required": ["username"]
          }
        }
      }
    ],
    "messages": [
      { "role": "user", "content": "What are the stats for PlayerOne?" }
    ]
  }'

# ── Step 2: model replies with a tool_call ────────────────────────────────
# choices[0].message.tool_calls[0]:
# { "id": "call_abc", "type": "function",
#   "function": { "name": "get_player_stats", "arguments": "{\"username\":\"PlayerOne\"}" } }

# ── Step 3: run your function locally and send the result ─────────────────
curl -X POST https://ai.tomdacat.com/api/chat \
  -H "Authorization: Bearer tdc_YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-5-mini",
    "tools": [ ... ],
    "messages": [
      { "role": "user", "content": "What are the stats for PlayerOne?" },
      { "role": "assistant", "content": null,
        "tool_calls": [{ "id": "call_abc", "type": "function",
          "function": { "name": "get_player_stats",
                        "arguments": "{\"username\":\"PlayerOne\"}" } }] },
      { "role": "tool", "tool_call_id": "call_abc",
        "content": "{\"level\":42,\"health\":100,\"gold\":3200}" }
    ]
  }'

# ── Step 4: model produces a final human-readable reply ───────────────────
JavaScript
// Full tool-use roundtrip in JavaScript
const BASE = 'https://ai.tomdacat.com/api/chat';
const KEY  = 'tdc_YOUR_API_KEY';
const MODEL = 'gpt-5-mini';

const tools = [
  {
    type: 'function',
    function: {
      name: 'get_player_stats',
      description: 'Return level, health and gold for a Roblox player.',
      parameters: {
        type: 'object',
        properties: {
          username: { type: 'string', description: 'Roblox username' },
        },
        required: ['username'],
      },
    },
  },
];

async function chat(messages) {
  const res = await fetch(BASE, {
    method: 'POST',
    headers: { Authorization: `Bearer ${KEY}`, 'Content-Type': 'application/json' },
    body: JSON.stringify({ model: MODEL, tools, messages }),
  });
  return res.json();
}

// ── Step 1: initial call ─────────────────────────────────────────────────
const messages = [{ role: 'user', content: 'What are the stats for PlayerOne?' }];
let data = await chat(messages);

// ── Step 2: handle tool_call ─────────────────────────────────────────────
while (data.choices[0].finish_reason === 'tool_calls') {
  const assistantMsg = data.choices[0].message;
  messages.push(assistantMsg);

  for (const tc of assistantMsg.tool_calls) {
    const args = JSON.parse(tc.function.arguments);
    // call your own function — e.g. look up the player in your database
    const result = await getPlayerStats(args.username);

    messages.push({
      role: 'tool',
      tool_call_id: tc.id,
      content: JSON.stringify(result),
    });
  }

  data = await chat(messages);
}

// ── Step 4: final reply ──────────────────────────────────────────────────
console.log(data.choices[0].message.content);
Python
import json, requests

BASE  = 'https://ai.tomdacat.com/api/chat'
KEY   = 'tdc_YOUR_API_KEY'
MODEL = 'gpt-5-mini'

tools = [
    {
        'type': 'function',
        'function': {
            'name': 'get_player_stats',
            'description': 'Return level, health and gold for a Roblox player.',
            'parameters': {
                'type': 'object',
                'properties': {
                    'username': {'type': 'string', 'description': 'Roblox username'},
                },
                'required': ['username'],
            },
        },
    }
]

def chat(messages):
    r = requests.post(BASE, headers={
        'Authorization': f'Bearer {KEY}',
        'Content-Type': 'application/json',
    }, json={'model': MODEL, 'tools': tools, 'messages': messages})
    return r.json()

# ── Step 1: initial call ──────────────────────────────────────────────────
messages = [{'role': 'user', 'content': 'What are the stats for PlayerOne?'}]
data = chat(messages)

# ── Step 2 & 3: handle tool calls ────────────────────────────────────────
while data['choices'][0]['finish_reason'] == 'tool_calls':
    assistant_msg = data['choices'][0]['message']
    messages.append(assistant_msg)

    for tc in assistant_msg['tool_calls']:
        args = json.loads(tc['function']['arguments'])
        # call your own function
        result = get_player_stats(args['username'])
        messages.append({
            'role': 'tool',
            'tool_call_id': tc['id'],
            'content': json.dumps(result),
        })

    data = chat(messages)

# ── Step 4: final reply ───────────────────────────────────────────────────
print(data['choices'][0]['message']['content'])
Luau (Roblox)
-- Tool calling roundtrip in Roblox Luau (server Script)
local HttpService = game:GetService("HttpService")
local API_KEY = "tdc_YOUR_API_KEY"
local API_URL = "https://ai.tomdacat.com/api/chat"

local tools = {
  {
    type = "function",
    ["function"] = {
      name = "get_player_stats",
      description = "Return level, health and gold for a Roblox player.",
      parameters = {
        type = "object",
        properties = {
          username = { type = "string", description = "Roblox username" },
        },
        required = { "username" },
      },
    },
  },
}

-- Your local function that returns real data
local function getPlayerStats(username)
  -- replace with your own DataStore / leaderstats lookup
  return { level = 42, health = 100, gold = 3200 }
end

local function chat(messages)
  local ok, result = pcall(HttpService.RequestAsync, HttpService, {
    Url    = API_URL,
    Method = "POST",
    Headers = {
      ["Authorization"] = "Bearer " .. API_KEY,
      ["Content-Type"]  = "application/json",
    },
    Body = HttpService:JSONEncode({ model = "gpt-5-mini", tools = tools, messages = messages }),
  })
  if not ok or not result.Success then
    warn("[TomAI] chat error:", result)
    return nil
  end
  return HttpService:JSONDecode(result.Body)
end

-- ── Step 1: initial call ──────────────────────────────────────────────────
local messages = {
  { role = "user", content = "What are the stats for PlayerOne?" },
}
local data = chat(messages)

-- ── Steps 2 & 3: handle tool calls ───────────────────────────────────────
while data and data.choices[1].finish_reason == "tool_calls" do
  local assistantMsg = data.choices[1].message
  table.insert(messages, assistantMsg)

  for _, tc in ipairs(assistantMsg.tool_calls) do
    local args = HttpService:JSONDecode(tc["function"].arguments)
    local result = getPlayerStats(args.username)

    table.insert(messages, {
      role        = "tool",
      tool_call_id = tc.id,
      content     = HttpService:JSONEncode(result),
    })
  end

  data = chat(messages)
end

-- ── Step 4: final reply ───────────────────────────────────────────────────
if data then
  print(data.choices[1].message.content)
end
Feature

Web search & fetch

Give the model live web access with two boolean flags — no tool-calling protocol to implement. We run the tool loop server-side and hand back a normal final answer.

Two flags, fully transparent

Set web_search: true and/or web_fetch: true on your request — either, both, or neither. When the model wants to search or read a page, we call it out to TinyFish, feed the result back in, and keep going — up to 5 rounds — until it has a real answer. You get back a normal choices[0].message.content, the same shape as any other reply. No tool schemas to write, no role: "tool" messages to send back — that's all handled for you.
web_search
boolean

Lets the model search the live web and get back ranked results (title, URL, snippet, date).

web_fetch
boolean

Lets the model fetch up to 10 URLs at once and read the extracted page text — a search result, or a URL you gave it directly.

Bringing your own tools too?

This only auto-resolves when your request has no tools of its own. If you pass custom tools alongside web_search/web_fetch, the model can still call them, but you get raw tool_calls back and handle everything yourself — same as any other tool. There's no way to silently resolve only some of a turn's tool calls while leaving yours unanswered, so we don't try.

Billing

Search and fetch themselves are free — nothing extra is charged for using them. A request that triggers a few rounds of tool use just bills normal per-token credits across however many turns it took, the same as if you'd run that multi-turn tool-calling loop yourself.

Example — Roblox / Luau

Luau (Roblox)
local HttpService = game:GetService("HttpService")

local response = HttpService:RequestAsync({
  Url    = "https://ai.tomdacat.com/api/chat",
  Method = "POST",
  Headers = {
    ["Authorization"] = "Bearer tdc_YOUR_API_KEY",
    ["Content-Type"]  = "application/json",
  },
  Body = HttpService:JSONEncode({
    model      = "gpt-5-mini",
    web_search = true,
    web_fetch  = true,
    messages   = {
      { role = "user", content = "What Roblox events are happening this week? Check a source and summarize." },
    },
  }),
})

local data = HttpService:JSONDecode(response.Body)
print(data.choices[1].message.content) -- final answer, web access already applied

Example — cURL

cURL
curl https://ai.tomdacat.com/v1/chat/completions \
  -H "Authorization: Bearer tdc_YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-5-mini",
    "web_search": true,
    "web_fetch": true,
    "messages": [
      { "role": "user", "content": "What Roblox events are happening this week? Check a source and summarize." }
    ]
  }'

Fetched page content reaches the model's context

Like any web-fetch-into-LLM tool, page text pulled by web_fetch lands directly in the conversation. Treat it the same as any other untrusted input — don't build flows where a fetched page's content can unilaterally trigger actions you wouldn't want a random website to be able to cause.
Examples

Chat code examples

The same request in four languages. Swap the model ID to use any available model.

cURL
curl -X POST https://ai.tomdacat.com/api/chat \
  -H "Authorization: Bearer tdc_YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-5-mini",
    "messages": [
      { "role": "system", "content": "You are a helpful NPC." },
      { "role": "user",   "content": "What quests do you have?" }
    ]
  }'
Roblox

Streaming in Roblox

Roblox added native HTTP streaming, but it currently only works inside Studio — here's what's real, and how to get a streaming-style chat UI in a published game today.

CreateWebStreamClient is Studio-only

In October 2025 Roblox shipped HttpService:CreateWebStreamClient(), which can open a real Server-Sent-Events connection and fire a MessageReceived event per chunk as a model's reply streams in. As of writing, this method only works inside Studio (plugins/testing) — "any CreateWebStreamClient() requests made in a live experience will be blocked," and Roblox hasn't announced a date for live-experience support. Don't ship code that depends on it; it will silently fail once your game is published.

Testing real streaming in Studio (plugin/local only)

Luau (Roblox)
local HttpService = game:GetService("HttpService")

local client = HttpService:CreateWebStreamClient(Enum.WebStreamClientType.SSE, {
  Url    = "https://ai.tomdacat.com/v1/chat/completions",
  Method = "POST",
  Headers = {
    ["Authorization"] = "Bearer tdc_YOUR_API_KEY",
    ["Content-Type"]  = "application/json",
  },
  Body = HttpService:JSONEncode({
    model    = "gpt-5-mini",
    messages = { { role = "user", content = "Tell me a short Roblox story." } },
    stream   = true,
  }),
})

client.MessageReceived:Connect(function(message)
  if message == "[DONE]" then return end
  local ok, chunk = pcall(HttpService.JSONDecode, HttpService, message)
  local delta = ok and chunk.choices and chunk.choices[1] and chunk.choices[1].delta and chunk.choices[1].delta.content
  if delta then
    print(delta) -- append to your UI as each piece arrives
  end
end)

client.Closed:Connect(function()
  print("Stream closed")
end)

In a published game: request normally, then "type out" the reply

Since live experiences can't consume a real stream yet, send a normal (non-streaming) request from a server Script and reveal the finished text character-by-character on the client — this is what most shipped Roblox AI chat NPCs do, and it reads as streaming to players.

Luau (Roblox)
local HttpService = game:GetService("HttpService")
local API_KEY = "tdc_YOUR_API_KEY" -- server-side Script only, never a LocalScript

local response = HttpService:RequestAsync({
  Url    = "https://ai.tomdacat.com/v1/chat/completions",
  Method = "POST",
  Headers = {
    ["Authorization"] = "Bearer " .. API_KEY,
    ["Content-Type"]  = "application/json",
  },
  Body = HttpService:JSONEncode({
    model    = "gpt-5-mini",
    messages = { { role = "user", content = "Tell me a short Roblox story." } },
    stream   = false, -- real streaming isn't available in live experiences yet
  }),
})

if response.Success then
  local data = HttpService:JSONDecode(response.Body)
  local text = data.choices[1].message.content

  -- Fire to the client and "type" it out — feels like streaming, works everywhere
  local revealEvent = game.ReplicatedStorage.RevealChatMessage -- a RemoteEvent
  revealEvent:FireClient(player, text)
end
Luau (Roblox)
local revealEvent = game.ReplicatedStorage.RevealChatMessage
local chatLabel = script.Parent.ChatLabel -- a TextLabel

revealEvent.OnClientEvent:Connect(function(text)
  chatLabel.Text = ""
  for i = 1, #text do
    chatLabel.Text = text:sub(1, i)
    task.wait(0.015) -- tune for typing speed
  end
end)
Response

Response object

Successful responses are 200 OK with an OpenAI-compatible chat completion object.

JSON
{
  "id": "chatcmpl-abc123",
  "object": "chat.completion",
  "model": "gpt-5-mini",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "I have a collection of swords, shields, and potions for sale!"
      },
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 22,
    "completion_tokens": 14,
    "total_tokens": 36
  }
}
choices[0].message.content
string

The AI-generated text reply.

choices[0].finish_reason
string

"stop" means a complete response. "length" means it was cut short.

usage.prompt_tokens
number

Tokens consumed by your input messages.

usage.completion_tokens
number

Tokens generated in the response.

usage.total_tokens
number

Sum of prompt and completion tokens. Used to calculate credit cost.

Feature

Prompt caching

Some models discount prompt tokens the provider already has cached from a recent, identical prefix. It's automatic — nothing to opt into on your end.

Automatic, no request changes needed

If a model supports it, repeating an identical prefix (a long system prompt, injected context, earlier turns of a conversation) gets billed at a cheaper cached rate for the portion the provider recognizes — you don't send anything special to enable it, and nothing about your request changes. Only models with a Cached in cr/M rate on the pricing page or models page support it — everything else bills at the normal input rate exactly as before.

Reading it back — OpenAI-shaped (/v1/chat/completions)

cached_tokens and cache_write_tokens are a subset of prompt_tokens, not additional tokens — they tell you how that same prompt was billed, not extra usage.

JSON
{
  "usage": {
    "prompt_tokens": 1562,
    "completion_tokens": 14,
    "total_tokens": 1576,
    "prompt_tokens_details": {
      "cached_tokens": 1552,    // subset of prompt_tokens billed at the model's cached rate
      "cache_write_tokens": 0   // subset billed at the cache-write rate (falls back to the
                                 // normal input rate if the model has no cache-write rate)
    }
  }
}

Reading it back — Anthropic-shaped (/v1/messages)

Same idea, using Anthropic's own field names — cache_read_input_tokens and cache_creation_input_tokens — both subsets of input_tokens.

JSON
{
  "usage": {
    "input_tokens": 1562,
    "output_tokens": 14,
    "cache_read_input_tokens": 1552,
    "cache_creation_input_tokens": 0
  }
}

Only /v1/chat/completions and /v1/messages

This applies to the credit-billed chat endpoints only — not Tomdacat Code, which is a flat-rate subscription and isn't billed per token to begin with.
Endpoint

POST /api/image

Generate images from a text prompt. Returns a permanent hosted image URL. Credits are charged per generation.

POSThttps://ai.tomdacat.com/api/image
Request

Image request body

JSON body sent to POST /api/image.

modelreq
string

Model ID, e.g. "dreamshaper-8-lcm" (default, cheapest), "flux-1-schnell", "gpt-image-1-mini", "mai-image-2-6", "gpt-image-1-5", "gpt-image-2". See the Models page for the full list.

promptreq
string

Text description of what to generate. Max 2000 characters.

width
number

Width in pixels (256–2048). Every current image model is fixed at 1024x1024, so this is accepted but has no effect.

height
number

Height in pixels (256–2048). Every current image model is fixed at 1024x1024, so this is accepted but has no effect.

seed
number

Integer seed. Currently not supported by any image model — accepted but ignored.

Examples

Image code examples

Generate images in four languages. Swap the model ID for different quality and speed.

cURL
curl -X POST https://ai.tomdacat.com/api/image \
  -H "Authorization: Bearer tdc_YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "dreamshaper-8-lcm",
    "prompt": "A neon cyberpunk cityscape at night, ultra detailed",
    "width": 1024,
    "height": 1024
  }'
JavaScript
const response = await fetch('https://ai.tomdacat.com/api/image', {
  method: 'POST',
  headers: {
    'Authorization': 'Bearer tdc_YOUR_API_KEY',
    'Content-Type': 'application/json',
  },
  body: JSON.stringify({
    model: 'mai-image-2-6',
    prompt: 'A dragon flying over a glowing Roblox city at sunset',
    width: 1024,
    height: 1024,
  }),
});

const data = await response.json();
console.log(data.url); // permanent hosted image URL
Python
import requests

response = requests.post(
    'https://ai.tomdacat.com/api/image',
    headers={
        'Authorization': 'Bearer tdc_YOUR_API_KEY',
        'Content-Type': 'application/json',
    },
    json={
        'model': 'gpt-image-1-mini',
        'prompt': 'A cute anime character in a magical forest',
        'width': 1024,
        'height': 1024,
    },
)

data = response.json()
print(data['url'])  # permanent hosted image URL — download separately if you need the bytes
Luau (Roblox)
local HttpService = game:GetService("HttpService")
local API_KEY = "tdc_YOUR_API_KEY"

local response = HttpService:RequestAsync({
  Url    = "https://ai.tomdacat.com/api/image",
  Method = "POST",
  Headers = {
    ["Authorization"] = "Bearer " .. API_KEY,
    ["Content-Type"]  = "application/json",
  },
  Body = HttpService:JSONEncode({
    model  = "dreamshaper-8-lcm",
    prompt = "A fantasy RPG sword with glowing runes, game asset style",
    width  = 1024,
    height = 1024,
  }),
})

if response.Success then
  local data = HttpService:JSONDecode(response.Body)
  print(data.url)        -- permanent hosted image URL
  print(data.tomdacat.creditsCharged)  -- 1
end
Response

Image response object

Successful responses return 200 OK with a permanent hosted URL to the generated image.

JSON
{
  "url": "https://media.pollinations.ai/HWSTclgEfU",
  "model": "dreamshaper-8-lcm",
  "prompt": "A neon cyberpunk cityscape at night",
  "width": 1024,
  "height": 1024,
  "seed": null,
  "tomdacat": {
    "creditsCharged": 1,
    "creditsRemaining": 242
  }
}
url
string

A permanent hosted URL to the generated image — not a temporary CDN link. Use directly as an <img> src or download it.

model
string

The model ID that was used.

prompt
string

The prompt that was sent.

width / height
number

Dimensions in pixels.

seed
number | null

Always null — no current image model supports a seed.

tomdacat.creditsCharged
number

Flat credit cost deducted for this generation.

tomdacat.creditsRemaining
number

Your remaining credit balance after this request.

Integrations

Coding Agents & AI Clients

TomdacatAI is OpenAI-compatible — connect any coding assistant that supports a custom base URL in seconds.

Two values, any tool

Base URL / API Base

https://ai.tomdacat.com/api

API Key

tdc_YOUR_API_KEY
🤖

Cline

VS Code extension — autonomous coding agent

Click the Cline icon in the sidebar → open Settings → choose OpenAI Compatible as the provider.

Setup
# 1. Install the Cline extension from the VS Code Marketplace
# 2. Click the Cline icon in the sidebar → open Settings (gear icon)
# 3. Select provider:  OpenAI Compatible
# 4. Enter your values and click Save

API Provider:  OpenAI Compatible
Base URL:      https://ai.tomdacat.com/api
API Key:       tdc_YOUR_API_KEY
Model ID:      gpt-5-mini          (or any Tomdacat model ID)
⌨️

Cursor

AI-first code editor

Go to Cursor Settings → Models, scroll to the OpenAI section (or Add Custom Model), and enter your base URL and key.

Setup
# Cursor Settings → Models → scroll to "OpenAI API Key" or "Add Custom Model"
# Enter the values below, then select the model in the editor:

Base URL:  https://ai.tomdacat.com/v1
API Key:   tdc_YOUR_API_KEY
Model:     gpt-5-mini

# Alt: set env vars before launching Cursor
OPENAI_BASE_URL=https://ai.tomdacat.com/v1
OPENAI_API_KEY=tdc_YOUR_API_KEY
🔵

Continue.dev

VS Code & JetBrains — chat + autocomplete

Add a model entry to ~/.continue/config.json or the equivalent YAML file.

JSON
// ~/.continue/config.json
{
  "models": [
    {
      "title": "TomdacatAI",
      "provider": "openai",
      "model": "gpt-5-mini",
      "apiBase": "https://ai.tomdacat.com/v1",
      "apiKey": "tdc_YOUR_API_KEY"
    }
  ],
  "tabAutocompleteModel": {
    "title": "TomdacatAI",
    "provider": "openai",
    "model": "gpt-5-mini",
    "apiBase": "https://ai.tomdacat.com/v1",
    "apiKey": "tdc_YOUR_API_KEY"
  }
}
YAML
# ~/.continue/config.yaml
models:
  - title: TomdacatAI
    provider: openai
    model: gpt-5-mini
    apiBase: https://ai.tomdacat.com/v1
    apiKey: tdc_YOUR_API_KEY

tabAutocompleteModel:
  title: TomdacatAI
  provider: openai
  model: gpt-5-mini
  apiBase: https://ai.tomdacat.com/v1
  apiKey: tdc_YOUR_API_KEY
🔧

Aider

Terminal pair programmer — AI edits your files

Pass flags on the CLI or put them in .aider.conf.yml in your project root.

Shell
# CLI flags
aider --openai-api-base https://ai.tomdacat.com/v1 \
      --openai-api-key  tdc_YOUR_API_KEY \
      --model           openai/gpt-5-mini
YAML
# .aider.conf.yml  (project root or ~/.aider.conf.yml)
openai-api-base: https://ai.tomdacat.com/v1
openai-api-key: tdc_YOUR_API_KEY
model: openai/gpt-5-mini
🟣

OpenCode

Terminal AI coding agent by SST

Add a custom provider to ~/.opencode.json or an opencode.json file in your project.

JSON
// ~/.opencode.json  (or opencode.json in your project root)
{
  "$schema": "https://opencode.ai/config.schema.json",
  "model": "tomdacat/gpt-5-mini",
  "provider": {
    "tomdacat": {
      "name": "TomdacatAI",
      "apiKey": "tdc_YOUR_API_KEY",
      "baseURL": "https://ai.tomdacat.com/v1"
    }
  }
}

Pick the right model

Use any Tomdacat model ID in the model field — see the Models section below for the full list. Free models like gpt-5-mini are great for everyday coding tasks. For complex refactors or architecture, try a reasoning model like claude-opus-4-7.
Anthropic-compatible

Claude Code & Anthropic SDK

TomdacatAI exposes a full Anthropic Messages API endpoint. Point any Anthropic SDK or Claude Code at it and every Tomdacat model works out of the box.

POSThttps://ai.tomdacat.com/v1/messages

Anthropic-standard path. Set base_url="https://ai.tomdacat.com" in any Anthropic SDK (the SDK appends /v1/messages automatically). Authentication accepts both x-api-key and Authorization: Bearer.

POSThttps://ai.tomdacat.com/api/anthropic/v1/messages

Legacy path — still works. Old base URL was https://ai.tomdacat.com/api/anthropic.

Setting up Claude Code

Set two environment variables — either in your shell, a .env file, or directly in ~/.claude/settings.json:

Shell
# Add these to your shell profile, .env, or Claude Code settings
ANTHROPIC_BASE_URL=https://ai.tomdacat.com
ANTHROPIC_API_KEY=tdc_YOUR_API_KEY
JSON
// ~/.claude/settings.json  (or project .claude/settings.json)
{
  "env": {
    "ANTHROPIC_BASE_URL": "https://ai.tomdacat.com",
    "ANTHROPIC_API_KEY": "tdc_YOUR_API_KEY"
  }
}
Shell
# Start Claude Code pointed at TomdacatAI
ANTHROPIC_BASE_URL=https://ai.tomdacat.com \
ANTHROPIC_API_KEY=tdc_YOUR_API_KEY \
claude --model gpt-5-mini
Pass any Tomdacat model ID via --model. Use gpt-5-mini for everyday tasks or a reasoning model like claude-opus-4-7 for complex refactors.

Python SDK

Python
import anthropic

client = anthropic.Anthropic(
    base_url="https://ai.tomdacat.com",
    api_key="tdc_YOUR_API_KEY",
)

# Use any Tomdacat model ID
message = client.messages.create(
    model="gpt-5-mini",
    max_tokens=1024,
    messages=[{"role": "user", "content": "What quests do you have?"}],
)
print(message.content[0].text)

# Streaming
with client.messages.stream(
    model="gpt-5-mini",
    max_tokens=1024,
    messages=[{"role": "user", "content": "Tell me a Roblox story."}],
) as stream:
    for text in stream.text_stream:
        print(text, end="", flush=True)

TypeScript / Node.js SDK

TypeScript
import Anthropic from '@anthropic-ai/sdk';

const client = new Anthropic({
  baseURL: 'https://ai.tomdacat.com',
  apiKey: 'tdc_YOUR_API_KEY',
});

// Use any Tomdacat model ID
const message = await client.messages.create({
  model: 'gpt-5-mini',
  max_tokens: 1024,
  messages: [{ role: 'user', content: 'What quests do you have?' }],
});
console.log(message.content[0].text);

// Streaming
const stream = client.messages.stream({
  model: 'gpt-5-mini',
  max_tokens: 1024,
  messages: [{ role: 'user', content: 'Tell me a Roblox story.' }],
});
for await (const chunk of stream) {
  if (chunk.type === 'content_block_delta' && chunk.delta.type === 'text_delta') {
    process.stdout.write(chunk.delta.text);
  }
}

Feature support

Non-streaming responses✅ Supported
Streaming (SSE)✅ Supported
System prompt✅ Supported
Vision (image blocks)✅ Supported (URL & base64)
Tool use / function calling✅ Supported
Multi-turn conversations✅ Supported
temperature, max_tokens✅ Supported
OAuth

OAuth Apps

Build apps that access TomdacatAI on behalf of users. Users authorize your app and receive an access token you can use to call the API with their balance.

How it works

  1. Register your app in the Developer Apps dashboard — you get a client_id and client_secret.
  2. Redirect users to /authorize with your client_id. They approve access and choose their limits.
  3. Exchange the code for an access_token at POST /api/oauth/token.
  4. Use the token as a Bearer token — calls bill the user's Tomdacat balance.

1Redirect to authorization

Send users to this URL. They will see a consent screen to choose scopes, model access, credit limit, and expiry.

text
https://ai.tomdacat.com/authorize?client_id=YOUR_CLIENT_ID&redirect_uri=https://yourapp.com/callback&scope=chat,image&state=RANDOM_STATE

scope — comma-separated: chat, image

state — random string for CSRF protection (recommended)

redirect_uri — must be registered in your app settings

2Exchange code for token

After the user approves, they are redirected to redirect_uri?code=AUTH_CODE&state=.... Exchange the code for a token:

bash
curl -X POST https://ai.tomdacat.com/api/oauth/token \
  -H "Content-Type: application/json" \
  -d '{
    "grant_type": "authorization_code",
    "client_id": "YOUR_CLIENT_ID",
    "client_secret": "YOUR_CLIENT_SECRET",
    "code": "AUTH_CODE",
    "redirect_uri": "https://yourapp.com/callback"
  }'

Response:

json
{
  "access_token": "tdcapp_...",
  "token_type": "Bearer"
}
Codes expire in 10 minutes and are single-use. Exchange immediately after receiving.

3Make API calls

Use the access_token exactly like a regular API key. All calls bill the authorizing user's Tomdacat balance.

bash
curl -X POST https://ai.tomdacat.com/api/chat/stream \
  -H "Authorization: Bearer tdcapp_..." \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gemini-2-5-flash-lite",
    "messages": [{ "role": "user", "content": "Hello!" }],
    "stream": true
  }'

Full example — Node.js

typescript
import express from 'express';
import crypto from 'crypto';

const CLIENT_ID = process.env.CLIENT_ID!;
const CLIENT_SECRET = process.env.CLIENT_SECRET!;
const REDIRECT_URI = 'http://localhost:3000/callback';

const app = express();
const states = new Map<string, boolean>();

// Step 1: redirect to Tomdacat authorization
app.get('/login', (_req, res) => {
  const state = crypto.randomBytes(16).toString('hex');
  states.set(state, true);
  const url = new URL('https://ai.tomdacat.com/authorize');
  url.searchParams.set('client_id', CLIENT_ID);
  url.searchParams.set('redirect_uri', REDIRECT_URI);
  url.searchParams.set('scope', 'chat,image');
  url.searchParams.set('state', state);
  res.redirect(url.toString());
});

// Step 2: handle callback + exchange code
app.get('/callback', async (req, res) => {
  const { code, state, error } = req.query as Record<string, string>;
  if (error || !states.has(state)) return res.send('Authorization failed.');
  states.delete(state);

  const resp = await fetch('https://ai.tomdacat.com/api/oauth/token', {
    method: 'POST',
    headers: { 'Content-Type': 'application/json' },
    body: JSON.stringify({
      grant_type: 'authorization_code',
      client_id: CLIENT_ID,
      client_secret: CLIENT_SECRET,
      code,
      redirect_uri: REDIRECT_URI,
    }),
  });
  const { access_token } = await resp.json();

  // Step 3: use the token
  const chat = await fetch('https://ai.tomdacat.com/api/chat/stream', {
    method: 'POST',
    headers: { Authorization: `Bearer ${access_token}`, 'Content-Type': 'application/json' },
    body: JSON.stringify({
      model: 'gemini-2-5-flash-lite',
      messages: [{ role: 'user', content: 'Hello from my app!' }],
      stream: false,
    }),
  });
  res.json(await chat.json());
});

app.listen(3000);

Token limits & scopes

When users authorize your app they can set restrictions that are enforced server-side:

  • Scopes — chat (chat API) and/or image (image generation). Calling outside the granted scope returns 403.
  • Model allowlist — user may restrict to specific models. Using a disallowed model returns 403.
  • Daily credit limit — once the daily cap is hit, requests return 402 until the next day.
  • Expiry — token automatically becomes invalid after the set date.
OAuth

Developer Earnings

Monetize your app by enabling a revenue share on every API call your users make through your access token.

How it works

  1. Enable earnings in your app settings in the Apps dashboard.
  2. Users accessing the API via your app's token are billed at 1.25× the standard credit cost.
  3. The extra 25% is automatically credited to your TomdacatAI account as lite credits after each interaction.
  4. No minimum thresholds — earnings are awarded per-request in real time.

Example

ItemAmount
Standard model cost1 credit
User is billed1.25 credits
You earn0.25 credits

Good to know

  • Earnings are lite credits added to your balance, not withdrawable cash.
  • The daily credit limit set by the user on authorization applies to the 1.25× billed amount.
  • You can toggle earnings on/off any time from the Apps dashboard; existing tokens pick up the change immediately.

Enabling via API

http
PUT /api/developer/apps/{app_id}
Authorization: Bearer <your-session-token>
Content-Type: application/json

{
  "developerEarningsEnabled": true
}
A separate product

Tomdacat Code API

A flat monthly Robux subscription for AI-assisted coding — its own base URL, its own API keys, and a rolling token allowance instead of per-token credit billing.

Two values, any tool

Base URL / API Base

https://ai.tomdacat.com/code/v1

API Key

tdcc_YOUR_API_KEY

Tomdacat Code is a completely separate product from the main credit-billed API: a flat monthly Robux plan that unlocks a fixed set of models plus a rolling 5-hour session and 7-day weekly token allowance. It doesn't touch your regular credit balance, and it uses its own tdcc_-prefixed API keys from the Code dashboard, not your main TomdacatAI key.

Plans are purchased and activated manually for now — join the Discord server to buy one. See ai.tomdacat.com/products/code for pricing.

Plans & models

PlanModelsWeekly tokens
Code LiteMimo v2.5, OpenAI GPT-5.6 Luna, Nemotron 3 Ultra 550B, deepseek-v4-flash-nvidia, Z.ai GLM 5.3 Flash~700K
Code StarterDeepSeek V4 Flash, Mimo v2.5, DeepSeek V4 Pro, Z.ai GLM 5.3 Flash, OpenAI GPT-5.6 Luna, Z.ai GLM 5.3, 6 Astra (direct)~1.5M
Code ProMimo v2.5, deepseek-v4-pro-2, kimi-k3-2, Claude Sonnet 5, Z.ai GLM 5.3, Z.ai GLM 5.3 Flash, 6 Astra (direct), 5.6 Sol (direct), Google Gemini 3.8 Flash~4M
Code StudioMimo v2.5, kimi-k3-2, Z.ai GLM 5.3, Z.ai GLM 5.3 Flash, deepseek-v4-pro-2, Claude Opus 5, Claude Sonnet 5, Nemotron 3 Ultra 550B, OpenAI GPT-5.6 Luna, OpenAI GPT-5.6 Sol, 6 Astra (direct), DeepSeek V4 Flash, 5.6 Sol (direct)~9M
Auth

Authentication

Same Bearer-token shape as the main API — a different key, from a different dashboard.

Header
Authorization: Bearer tdcc_YOUR_API_KEY

Not your main API key

Keys created on /dashboard/code start with tdcc_ and only work against /code/v1. Your main tdc_ key doesn't work here, and vice versa.
Endpoints

Endpoints

Five endpoints under /code/v1 — two ways to send a chat turn (OpenAI-shaped or Anthropic-shaped, so both OpenAI-style tools and Claude Code work), the models your plan unlocks, live usage, and a key-validity check. All auth the same tdcc_ key.

Chat completions — OpenAI-shaped

POSThttps://ai.tomdacat.com/code/v1/chat/completions
cURL
curl https://ai.tomdacat.com/code/v1/chat/completions \
  -H "Authorization: Bearer tdcc_YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deepseek-v4-flash",
    "messages": [
      { "role": "user", "content": "Write a Luau function that shuffles a table." }
    ]
  }'

Use this one for OpenAI-compatible tools (Cline, Cursor, Continue, Aider, OpenCode — see the setup guides below). The model field must be one your plan unlocks — using anything else returns a 403. Successful responses carry the same fields any OpenAI-compatible client already reads, plus a non-standard tomdacat_code object with your live session/weekly usage — safe to ignore if you don't need it.

JSON
{
  "id": "chatcmpl-...",
  "object": "chat.completion",
  "model": "deepseek-v4-flash",
  "choices": [
    { "index": 0, "message": { "role": "assistant", "content": "..." }, "finish_reason": "stop" }
  ],
  "usage": { "prompt_tokens": 42, "completion_tokens": 180, "total_tokens": 222 },

  // Extra field — not part of the OpenAI spec, safe to ignore if your
  // client only reads the standard fields above.
  "tomdacat_code": {
    "session": { "used": 4222, "limit": 300000, "resetsAt": "2026-07-27T14:00:00.000Z" },
    "weekly":  { "used": 88012, "limit": 4000000, "resetsAt": "2026-08-02T00:00:00.000Z" }
  }
}

Messages — Anthropic-shaped

POSThttps://ai.tomdacat.com/code/v1/messages
cURL
curl https://ai.tomdacat.com/code/v1/messages \
  -H "x-api-key: tdcc_YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deepseek-v4-flash",
    "max_tokens": 1024,
    "messages": [
      { "role": "user", "content": "Write a Luau function that shuffles a table." }
    ]
  }'

This is the one Claude Code needs — it only speaks Anthropic's wire format, not OpenAI's, so pointing it at /chat/completions above won't work. Same model restriction and plan limits as chat completions, translated to Anthropic's request/response shape (content blocks, stop_reason, tool_use blocks for tool calls). Accepts both x-api-key and Authorization: Bearer, same as the main product's Anthropic endpoint.

JSON
{
  "id": "msg_...",
  "type": "message",
  "role": "assistant",
  "model": "deepseek-v4-flash",
  "content": [
    { "type": "text", "text": "..." }
  ],
  "stop_reason": "end_turn",
  "usage": { "input_tokens": 24, "output_tokens": 180 },

  // Extra field — not part of the Anthropic spec, safe to ignore.
  "tomdacat_code": {
    "session": { "used": 4222, "limit": 300000, "resetsAt": "2026-07-27T14:00:00.000Z" },
    "weekly":  { "used": 88012, "limit": 4000000, "resetsAt": "2026-08-02T00:00:00.000Z" }
  }
}

Models

GEThttps://ai.tomdacat.com/code/v1/models
cURL
curl https://ai.tomdacat.com/code/v1/models \
  -H "Authorization: Bearer tdcc_YOUR_API_KEY"

Lists only the models your plan includes — useful for a coding agent to populate its model picker without hardcoding one. Each entry carries a non-standard tomdacat_code.premium flag marking whether it counts against the plan's separate flagship-model allowance, if it has one.

JSON
{
  "object": "list",
  "data": [
    {
      "id": "deepseek-v4-flash",
      "object": "model",
      "created": 1700000000,
      "owned_by": "tomdacat",
      "capabilities": ["tools", "vision"],

      // Extra field — flags whether this model also counts against the
      // plan's separate premium/flagship weekly allowance, if it has one.
      "tomdacat_code": { "premium": false }
    },
    {
      "id": "nemotron-3-ultra-550b",
      "object": "model",
      "created": 1700000000,
      "owned_by": "tomdacat",
      "capabilities": ["tools"],
      "tomdacat_code": { "premium": true }
    }
  ]
}

Usage

GEThttps://ai.tomdacat.com/code/v1/usage
cURL
curl https://ai.tomdacat.com/code/v1/usage \
  -H "Authorization: Bearer tdcc_YOUR_API_KEY"

Returns your current session (rolling 5h), weekly (rolling 7d), and — on plans with a flagship model — premium token allowance, each with used, limit, remaining, and resetsAt. Check this before a big request instead of finding out via a 429 mid-task.

JSON
{
  "plan": { "id": "code-pro", "name": "Code Pro" },
  "session": {
    "used": 42000,
    "limit": 300000,
    "remaining": 258000,
    "resetsAt": "2026-07-28T19:00:00.000Z"
  },
  "weekly": {
    "used": 812000,
    "limit": 4000000,
    "remaining": 3188000,
    "resetsAt": "2026-08-02T00:00:00.000Z"
  },

  // null on plans without a separate flagship-model allowance
  "premium": null
}

Verify

GEThttps://ai.tomdacat.com/code/v1/verify
cURL
curl https://ai.tomdacat.com/code/v1/verify \
  -H "Authorization: Bearer tdcc_YOUR_API_KEY"

Checks whether a key is valid without spending any tokens. Always answers with 200 — read the valid field instead of the status code, so an integration can branch on it directly. When valid, the response includes key metadata (name, prefix, timestamps) plus — if the account has an active plan — that plan's models, limits, and live usage in one call.

JSON
{
  "valid": true,
  "key": {
    "id": "cakey_9f2a1b7e3c5d",
    "name": "My laptop",
    "prefix": "tdcc_3f9a2c8b1d4e",
    "createdAt": "2026-05-01T10:00:00.000Z",
    "lastUsedAt": "2026-07-28T18:42:11.000Z"
  },
  "plan": {
    "id": "code-pro",
    "name": "Code Pro",
    "status": "active",
    "models": ["deepseek-v4-flash", "mimo-v2.5"],
    "premiumModels": [],
    "currentPeriodEnd": "2026-08-15T00:00:00.000Z"
  },
  "usage": {
    "session": { "used": 42000, "limit": 300000, "resetsAt": "..." },
    "weekly":  { "used": 812000, "limit": 4000000, "resetsAt": "..." },
    "premium": null
  }
}
JSON
{
  "valid": false,
  "error": {
    "message": "This key is invalid, revoked, or belongs to a suspended account.",
    "code": "invalid_api_key"
  }
}
Integrations

Coding agent setup

Claude Code uses the Messages endpoint above; every other tool here uses Chat completions — same OpenAI-compatible clients as the main API, pointed at the Code endpoint and key instead.

Two base URLs — don't mix them up

OpenAI-compatible tools (Cline, Cursor, Continue, Aider, OpenCode) want https://ai.tomdacat.com/code/v1 — the client appends /chat/completions itself. Claude Code wants https://ai.tomdacat.com/code — no /v1 — because the Anthropic SDK it's built on appends /v1/messages itself. Copy the value from the tool's own card below rather than reusing one you copied for a different tool.

Claude Code

Anthropic's own CLI coding agent

Set two environment variables — either in your shell, a .env file, or directly in ~/.claude/settings.json:

Shell
# Add these to your shell profile, .env, or Claude Code settings
ANTHROPIC_BASE_URL=https://ai.tomdacat.com/code
ANTHROPIC_AUTH_TOKEN=tdcc_YOUR_API_KEY
JSON
// ~/.claude/settings.json  (or project .claude/settings.json)
{
  "env": {
    "ANTHROPIC_BASE_URL": "https://ai.tomdacat.com/code",
    "ANTHROPIC_AUTH_TOKEN": "tdcc_YOUR_API_KEY"
  }
}
Shell
# Start Claude Code pointed at your Tomdacat Code plan
ANTHROPIC_BASE_URL=https://ai.tomdacat.com/code \
ANTHROPIC_AUTH_TOKEN=tdcc_YOUR_API_KEY \
claude --model deepseek-v4-flash
🤖

Cline

VS Code extension — autonomous coding agent

Setup
# 1. Install the Cline extension from the VS Code Marketplace
# 2. Click the Cline icon in the sidebar → open Settings (gear icon)
# 3. Select provider:  OpenAI Compatible
# 4. Enter your values and click Save

API Provider:  OpenAI Compatible
Base URL:      https://ai.tomdacat.com/code/v1
API Key:       tdcc_YOUR_API_KEY
Model ID:      deepseek-v4-flash   (or another model your plan unlocks)
⌨️

Cursor

AI-first code editor

Setup
# Cursor Settings → Models → scroll to "OpenAI API Key" or "Add Custom Model"
# Enter the values below, then select the model in the editor:

Base URL:  https://ai.tomdacat.com/code/v1
API Key:   tdcc_YOUR_API_KEY
Model:     deepseek-v4-flash

# Alt: set env vars before launching Cursor
OPENAI_BASE_URL=https://ai.tomdacat.com/code/v1
OPENAI_API_KEY=tdcc_YOUR_API_KEY
🔵

Continue.dev

VS Code & JetBrains — chat + autocomplete

JSON
// ~/.continue/config.json
{
  "models": [
    {
      "title": "Tomdacat Code",
      "provider": "openai",
      "model": "deepseek-v4-flash",
      "apiBase": "https://ai.tomdacat.com/code/v1",
      "apiKey": "tdcc_YOUR_API_KEY"
    }
  ]
}
🔧

Aider

Terminal pair programmer — AI edits your files

Shell
# CLI flags
aider --openai-api-base https://ai.tomdacat.com/code/v1 \
      --openai-api-key  tdcc_YOUR_API_KEY \
      --model           openai/deepseek-v4-flash
🟣

OpenCode

Terminal AI coding agent by SST

JSON
// ~/.opencode.json  (or opencode.json in your project root)
{
  "$schema": "https://opencode.ai/config.schema.json",
  "model": "tomdacat-code/deepseek-v4-flash",
  "provider": {
    "tomdacat-code": {
      "name": "Tomdacat Code",
      "apiKey": "tdcc_YOUR_API_KEY",
      "baseURL": "https://ai.tomdacat.com/code/v1"
    }
  }
}

Pick a model your plan includes

deepseek-v4-flash is in every plan, so it's the safest default above. Higher tiers add mimo-v2.5 and nemotron-3-ultra-550b — see the plan table above.
Errors

Limits & error codes

Hitting a token limit returns a normal OpenAI-shaped error, not a silent failure.

JSON
// Limit-reached error response
{
  "error": {
    "message": "Weekly token limit reached for this plan. Resets at 2026-08-02T00:00:00.000Z.",
    "type": "rate_limit_error",
    "code": "code_weekly_limit_reached"
  }
}
CodeStatusMeaning
401UnauthorizedMissing or invalid Authorization header, or the key isn’t a tdcc_ Tomdacat Code key.
402Payment RequiredNo active Tomdacat Code plan on this account.
403ForbiddenThe model isn’t included in your current plan, or you requested tool calling on a model that doesn’t support it.
404Not FoundUnknown model ID.
429Too Many RequestsSession, weekly, or flagship-model token limit reached. The response includes when it resets.
500Internal Server ErrorUpstream model provider error. Retry with exponential backoff.
Research preview · Playground only

Browser Use

Let any vision + tool-calling model drive a real Chromium browser from the Playground — navigate, click, type, read pages, and see screenshots as it happens.

Each session spins up a real remote Chromium sandbox that the model controls one action at a time. Click the globe icon in the Playground toolbar, confirm the research-preview notice, and start chatting — every navigate, click, or screenshot the model makes shows up inline as it happens.

This is not yet part of the public API — it only runs inside the Playground's own chat UI for now, not through /v1/chat/completions. See ai.tomdacat.com/products/browser-use for the full walkthrough and FAQ.

Requires vision + tool calling

The globe toggle only lights up once you've picked a model with both vision and tools capabilities — the model needs to see screenshots and call tools to drive the browser.
8 tools

What it can do

Granular actions rather than one mega-tool, so each step in the agent's log is easy to follow.

navigate

Loads any http(s) URL — file://, javascript:, and data: URLs are refused.

click

Clicks by CSS selector or plain visible text.

type

Fills a field (selector or label text), optionally pressing Enter to submit.

scroll

Scrolls the page up or down by a configurable pixel amount.

press_key

Sends a single keystroke — Enter, Escape, PageDown, and the like.

screenshot

Captures the page as JPEG — shown to you and fed back to the model.

read_text

Extracts visible page text (truncated to ~6,000 chars) — cheaper than a screenshot.

go_back

Navigates back one step in browser history.

Billing

Billed by the minute

0.005 credits per minute, ceiling-billed — enabling it charges a full minute immediately, then every minute after is metered as the sandbox keeps running.

Session lengthCost
1 min0.005 cr
5 min0.025 cr
15 min0.075 cr
30 min (session cap)0.15 cr

Billing draws from your normal credit balance — lite credits first, then paid — the same pool used by every other model call. Normal per-token model billing applies on top of the per-minute sandbox fee for whatever the model itself generates.

Ending the session (clicking the globe icon again) stops billing immediately. Running out of credits mid-session kills the sandbox automatically, and the model answers with whatever it had already found.

Known limits

Limits

This is an early research preview — a few hard limits keep it predictable while it's tested.

Research preview

Behavior, tools, and pricing may change — expect occasional bugs.

http(s) only

No local files, no script URLs — the browser can only reach real web pages.

30-minute session cap

Sessions don't renew — once 30 minutes pass, the sandbox self-destructs.

10 steps per message

Each reply gets up to 10 tool calls before the model has to answer with what it has.

Dashboard · Discord bots

Bot Hosting

Run a Discord bot powered by TomdacatAI without hosting anything yourself — from the Dashboard, provide a bot token and either configure a ready-made bot or bring your own code.

Add a bot from ai.tomdacat.com/dashboard/hosting with your Discord bot token — stored AES-256-GCM encrypted, decrypted only server-side when the bot is running. Two ways to define what it does:

Config

Pick a model, response mode, and system prompt — TomdacatAI runs a ready-made bot for you. Available on Basic and Pro.

Custom code

Upload your own Python file, or describe the bot to an AI assistant that writes it for you. Pro, chat bots only.

Pro · Chat bots

Custom code

Write your own disnake bot, or have TomdacatAI's AI assistant write one from a description — same review either way.

Your bot runs as its own isolated process with the token you provided. Conventions to follow (the generated code already does this): use disnake for the Discord connection, read config from the BOT_CONFIG environment variable, and call /api/chat with your own API key for AI responses — the same endpoint documented above.

To generate code with AI, describe the bot in the Code tab on your bot's settings page and pick a model — it defaults to mimo-v2.5, but any tool-calling model works. It drafts a full file, checks it against a static scan, revises if anything's flagged, then sends the result through the same safety review a manual upload gets.

Every version is reviewed

A bot can't start until its current code version has an approved review verdict. Uploading a new version, or asking the AI assistant to revise, always re-runs review.
Two stages

Safety review

A static scan and an AI reviewer run on every version before it's allowed to run — this is a filter against obviously unsafe code, not a security guarantee.

Static scan

Instant pattern checks — process spawning, eval/exec, network calls outside Discord and TomdacatAI, obfuscated payloads.

AI review

A fixed reviewer model — not the one you picked for the bot — reads the code and its static flags, then approves, rejects, or asks for a manual look.

Runtime limits

Approved code still runs under memory, file-descriptor, and CPU-time limits, with a watchdog that stops anything that breaches them.

Not a guarantee

Review and limits reduce risk — they don’t certify your code is correct or safe. You’re responsible for what it does.

Billing

A small process fee, usage billed like any other call

1 credit per day keeps the bot's process running. Every AI call it makes — chat or image — bills at the same per-token or per-image rate as any other request.

ChargeRate
Process fee (Basic)1 lite cr / day
Process fee (Pro)1 pro cr / day
Chat / image usagestandard per-token / per-image rate
Safety review (Custom code)standard per-token rate, one small call per version

The process fee is charged when the bot starts and renews daily while it keeps running. If your balance can't cover the next day, the bot stops automatically and its status shows the reason — restart once you've topped up.

Try it live

API Playground

Paste your API key, pick an endpoint, tweak the parameters, and send a real request — the exact raw response (including streamed chunks) shows up on the right. Your key is only ever sent from your browser to ai.tomdacat.com and is stored locally on this device.

Chat completions (OpenAI-compatible)

curl
curl -X POST "https://ai.tomdacat.com/v1/chat/completions" \
  -H "Content-Type: application/json" \
  -d '{
  "model": "claude-opus-5",
  "messages": [
    {
      "role": "system",
      "content": "You are a helpful assistant."
    },
    {
      "role": "user",
      "content": "Say hello in one sentence."
    }
  ],
  "stream": false,
  "temperature": 1,
  "max_tokens": 512
}'
Response

Send a request to see the exact raw response here.

Models

Available models

All models use the same endpoint. Prices are in credits per million tokens.

Claude Opus 5

Pro

Anthropic's most powerful Opus model. Best for complex reasoning, agentic tasks, and long-context work.

visionreasoningtools
claude-opus-5

500 in · 2500 out
credits / M tokens
50 cached in / M
625 cache write / M

Claude Opus 4.7

Pro

Anthropic's latest and most capable Opus model. Premium reasoning and agentic tasks.

visionreasoningtools
claude-opus-4-7

500 in · 2500 out
credits / M tokens
50 cached in / M
625 cache write / M

Claude Opus 4.6

Pro

Anthropic's most powerful model. Best for complex reasoning and agentic tasks.

visionreasoningtools
claude-opus-4-6

500 in · 2500 out
credits / M tokens
50 cached in / M
625 cache write / M

Claude Sonnet 5

Pro

Anthropic's latest Sonnet model. Balanced premium model for rich Roblox NPC and gameplay agents.

visiontools
claude-sonnet-5

300 in · 1500 out
credits / M tokens
20 cached in / M
250 cache write / M

Claude Opus 5.5

Pro

Anthropic Claude Opus 5.5 via Pollinations — flagship reasoning for demanding coding and long-context agentic work. Requires pro credits.

visiontools
claude-opus-5-5

400 in · 2000 out
credits / M tokens
20 cached in / M
500 cache write / M

Claude Sonnet 4.6

Pro

Balanced premium model for rich Roblox NPC and gameplay agents.

visiontools
claude-sonnet-4-6

300 in · 1500 out
credits / M tokens
30 cached in / M
375 cache write / M

Anthropic Claude Haiku 4.5

Pro

Quick premium responses with strong safety and instruction following.

visiontools
claude-haiku-4-5

100 in · 500 out
credits / M tokens
10 cached in / M
125 cache write / M

Claude Haiku 4.5 (Free)

Lite

Same Claude Haiku 4.5. Uses lite credits.

visiontools
claude-haiku-4-5-free

500 in · 500 out
credits / M tokens

Google Gemini 3.1 Pro Preview

Pro

Google's most advanced Gemini 3.1 Pro — frontier reasoning and multimodal capabilities. Requires pro credits.

visionreasoningtools
gemini-3-1-pro-preview

200 in · 1200 out
credits / M tokens
20 cached in / M
200 cache write / M

Google Gemini 3 Flash

Lite

Google Gemini 3 Flash — fast and capable multimodal model.

visionreasoningtools
gemini-3-flash

50 in · 300 out
credits / M tokens

Google Gemini 3.7 Flash

Lite

Google Gemini 3.7 Flash — enhanced multimodal model with improved reasoning.

visionreasoningtools
gemini-3-7-flash

75 in · 375 out
credits / M tokens

Google Gemini 3.5 Flash

Lite

Google Gemini 3.5 Flash — capable multimodal model with reasoning.

visionreasoningtools
gemini-3-5-flash

150 in · 900 out
credits / M tokens

Google Gemini 3.6 Flash

Lite

Google Gemini 3.6 Flash — capable multimodal model with reasoning.

visionreasoningtools
gemini-3-6-flash

150 in · 900 out
credits / M tokens

Google Gemini 3.8 Flash

Lite

Google Gemini 3.8 Flash — enhanced multimodal model with improved reasoning.

visionreasoningtools
gemini-3-8-flash

75 in · 375 out
credits / M tokens

Google Gemini 2.5 Flash Lite

Lite

Ultra-fast and cost-effective for high volume game traffic.

visiontools
gemini-2-5-flash-lite

10 in · 25 out
credits / M tokens

Gemini 3.1 Flash Lite

Lite

Google Gemini 3.1 Flash Lite — fast multimodal model.

textvision
gemini-3-1-flash-lite-preview

30 in · 160 out
credits / M tokens

Google Gemini 3.5 Flash Lite

Lite

Google Gemini 3.5 Flash Lite — fast multimodal model.

visiontools
gemini-3-5-flash-lite

30 in · 250 out
credits / M tokens

Google Gemma 4 31B

Lite

Google Gemma 4 31B instruction-tuned — powerful open model. Uses lite credits.

text
gemma-4-31b

30 in · 60 out
credits / M tokens

Google Gemma 4 26B

Lite

Google Gemma 4 26B MoE instruction-tuned — open model. Uses lite credits.

texttools
gemma-4-26b

20 in · 20 out
credits / M tokens

Google Gemma 4 26B (Raw)

Lite

Google Gemma 4 26B MoE instruction-tuned — raw, unmodified output including reasoning tokens. Uses lite credits.

textreasoningtools
gemma-4-26b-raw

20 in · 20 out
credits / M tokens

xAI Grok 4.2 Reasoning

Pro

xAI Grok 4.2 Reasoning — deep reasoning model with vision. Requires pro credits.

visionreasoningtools
grok-4-2-reasoning

200 in · 600 out
credits / M tokens
20 cached in / M

Grok 4.2

Lite

xAI Grok 4.2 (non-reasoning) with vision. Uses lite credits.

visiontools
grok-4-2

200 in · 600 out
credits / M tokens
20 cached in / M

xAI Grok 4.3

Lite

xAI Grok 4.3 with vision and reasoning. Uses lite credits.

visionreasoningtools
grok-4-3

125 in · 250 out
credits / M tokens
20 cached in / M

xAI Grok 4.5

Lite

xAI Grok 4.5 with reasoning and tool calling. Uses lite credits.

reasoningtools
grok-4-5

100 in · 300 out
credits / M tokens

xAI Grok 4.6

Lite

xAI Grok 4.6 with vision, reasoning and tool calling. Uses lite credits.

visionreasoningtools
grok-4-6

125 in · 250 out
credits / M tokens

OpenAI GPT-6 Astra

Lite

OpenAI GPT-6 Astra — frontier reasoning for complex agentic, coding, and multimodal work. Uses lite credits.

visionreasoningtools
gpt-6-astra

1000 in · 5000 out
credits / M tokens
100 cached in / M
1250 cache write / M

OpenAI GPT-5.5

Pro

OpenAI GPT-5.5 — frontier model with vision and reasoning. Requires pro credits.

visionreasoningtools
gpt-5-5

500 in · 3000 out
credits / M tokens
50 cached in / M

OpenAI GPT-5.6 Sol

Lite

OpenAI GPT-5.6 Sol — frontier model with vision and reasoning. Uses lite credits.

visionreasoningtools
gpt-5-6-sol

500 in · 3000 out
credits / M tokens
25 cached in / M
312.5 cache write / M

OpenAI GPT-5.6 Terra

Pro

OpenAI GPT-5.6 Terra — frontier model with vision and reasoning. Requires pro credits.

visionreasoningtools
gpt-5-6-terra

250 in · 1500 out
credits / M tokens
12.5 cached in / M
156.25 cache write / M

OpenAI GPT-5.6 Luna

Lite

OpenAI GPT-5.6 Luna — capable model with vision and reasoning. Uses lite credits.

visionreasoningtools
gpt-5-6-luna

100 in · 600 out
credits / M tokens
2 cached in / M
25 cache write / M

OpenAI GPT-6 Sol

Lite

OpenAI GPT-6 Sol — reasoning for complex coding and agentic workflows. Uses lite credits.

visionreasoningtools
gpt-6-sol

200 in · 1000 out
credits / M tokens
20 cached in / M
250 cache write / M

OpenAI GPT-6 Luna

Lite

OpenAI GPT-6 Luna — efficient reasoning for focused, high-volume tasks. Uses lite credits.

visionreasoningtools
gpt-6-luna

10 in · 50 out
credits / M tokens
1 cached in / M
12.5 cache write / M

GPT-5 Chat Latest

Pro

OpenAI GPT-5 Chat Latest — chat-tuned frontier model with vision. Requires pro credits.

visiontools
gpt-5-chat-latest

600 in · 600 out
credits / M tokens

OpenAI GPT-5.4

Pro

OpenAI GPT-5.4 — frontier model with vision and reasoning. Requires pro credits.

visionreasoningtools
gpt-5-4

250 in · 1500 out
credits / M tokens
25 cached in / M

OpenAI GPT-5.4 Mini

Lite

GPT-5.4 Mini — a capable mid-size model with vision support. Uses lite credits.

visiontools
gpt-5-4-mini

75 in · 450 out
credits / M tokens
5.625 cached in / M

OpenAI GPT-5 Mini

Lite

Strong default choice for smart assistants and game copilots.

visiontools
gpt-5-mini

15 in · 65 out
credits / M tokens
1.5 cached in / M

OpenAI GPT-5 Mini Special

Pro

Same GPT-5 Mini at a discounted special rate. Requires pro credits.

visiontools
gpt-5-mini-special

8 in · 33 out
credits / M tokens
1.5 cached in / M

GPT-4o Mini

Lite

OpenAI GPT-4o Mini — fast multimodal model. Uses lite credits.

visiontools
gpt-4o-mini

15 in · 60 out
credits / M tokens

OpenAI GPT-5 Nano

Lite

Fast lightweight model for responsive in-game conversations.

visiontools
gpt-5-nano

10 in · 50 out
credits / M tokens
0.375 cached in / M

OpenAI GPT-5 Nano Special

Pro

Same GPT-5 Nano at a discounted special rate. Requires pro credits.

visiontools
gpt-5-nano-special

3 in · 13 out
credits / M tokens
0.375 cached in / M

OpenAI GPT-5.4 Nano

Lite

GPT-5.4 Nano — fast and capable with vision support. Uses lite credits.

visiontools
gpt-5-4-nano

20 in · 125 out
credits / M tokens
1.5 cached in / M

OpenAI GPT-5.4 Nano Special

Pro

Same GPT-5.4 Nano at a discounted special rate. Requires pro credits.

visiontools
gpt-5-4-nano-special

10 in · 63 out
credits / M tokens
1.5 cached in / M

OpenAI GPT-OSS 120B

Lite

OpenAI's open-source 120B model for powerful NPC dialogue and game logic.

text
gpt-oss-120b

10 in · 15 out
credits / M tokens

OpenAI GPT-OSS 20B

Lite

OpenAI's open-source 20B model — lightweight NPC dialogue and game logic. Uses lite credits.

texttools
gpt-oss-20b

5 in · 5 out
credits / M tokens

DeepSeek V4 Flash

Lite

Agentic reasoning model for complex task orchestration.

reasoningtools
deepseek-v4-flash

30 in · 85 out
credits / M tokens

DeepSeek V4 Flash Special

Pro

Same DeepSeek V4 Flash at a discounted special rate. Requires pro credits.

reasoningtools
deepseek-v4-flash-special

15 in · 43 out
credits / M tokens
2.8 cached in / M

DeepSeek V4 Pro

Lite

DeepSeek V4 Pro. Uses lite credits.

reasoningtools
deepseek-v4-pro

174 in · 348 out
credits / M tokens

MiniMax M3

Lite

MiniMax M3 — vision-capable reasoning model. Lite credits.

visionreasoningtools
minimax-m3

15 in · 60 out
credits / M tokens
6 cached in / M

Nemotron Lightning 3.5 30B

Lite

Nemotron Lightning 3.5 30B — fast MoE model. Lite credits.

texttools
nemotron-lightning-3-5-30b

5 in · 30 out
credits / M tokens
1 cached in / M

Qwen 3.8 Max

Lite

Qwen 3.8 Max — large-scale flagship model with vision and reasoning. Lite credits.

visionreasoningtools
qwen-3-8-max

200 in · 600 out
credits / M tokens
25 cached in / M

Perplexity Sonar

Pro

Perplexity Sonar — live web search built in. Answers are grounded with up-to-date results from the internet. Requires pro credits.

web-search
perplexity-fast

25 in · 250 out
credits / M tokens
6 cached in / M
25 cache write / M

Moonshot Kimi K2.6

Lite

Long-context agentic model for deep gameplay workflows.

visionreasoningtools
kimi-k2-6

95 in · 400 out
credits / M tokens
16 cached in / M

Moonshot Kimi K2.7 Code

Pro

Kimi K2.7 Code — vision and reasoning model optimised for coding tasks. Requires pro credits.

visionreasoningtools
kimi-k2-7-code

100 in · 422 out
credits / M tokens
20 cached in / M

Moonshot Kimi K3

Lite

Moonshot Kimi K3.

visionreasoningtools
kimi-k3

300 in · 1500 out
credits / M tokens

Z.ai GLM 5.2

Pro

Z.ai GLM 5.2 — long-context reasoning model. Requires pro credits.

reasoningtools
glm-5-2

148 in · 464 out
credits / M tokens
27 cached in / M

Z.ai GLM 5.2 Special

Pro

Same GLM 5.2 at a discounted special rate. Requires pro credits.

reasoningtools
glm-5-2-special

120 in · 120 out
credits / M tokens

Z.ai GLM 5.3

Lite

Z.ai GLM 5.3 — better at complex coding and long-horizon tasks than GLM 5.2. Priced the same as GLM 5.2. Lite credits.

reasoningtools
glm-5-3

140 in · 440 out
credits / M tokens
14 cached in / M

Z.ai GLM 5.3 Flash

Lite

Z.ai GLM 5.3 Flash — faster, cheaper variant of GLM 5.3. Lite credits.

reasoningtools
glm-5-3-flash

15 in · 50 out
credits / M tokens
3 cached in / M

Muse Spark 1.3 Contributor

Lite

Muse Spark 1.3, contributor tier — vision and reasoning model. Note: this provider may use your prompts to help improve Meta's products. Uses lite credits.

visionreasoning
muse-spark-1-3-contributor

10 in · 10 out
credits / M tokens

Nemotron 3 Ultra 550B

Pro

NVIDIA's 550B parameter flagship model with tool calling. Requires pro credits.

texttools
nemotron-3-ultra-550b

60 in · 360 out
credits / M tokens

Nemotron 3 Super 120B

Lite

NVIDIA's 120B parameter model for high-quality game AI at low cost.

texttools
nemotron-3-super-120b

10 in · 20 out
credits / M tokens

Nemotron 3 Ultra 550B

Lite

NVIDIA's 550B parameter flagship model. Uses lite credits.

texttools
nemotron-3-ultra-550b-nvidia

20 in · 50 out
credits / M tokens

DiffusionGemma 4 26B

Lite

Google's diffusion-based Gemma 4 26B. Supports reasoning mode and tool calling. Uses lite credits.

textreasoningtools
diffusiongemma-4-26b

5 in · 5 out
credits / M tokens

Nemotron 3 nano omni reasoning

Lite

NVIDIA Nemotron 3 nano omni 30B reasoning — free-tier model. Note: this provider may log your prompts and responses to train the model.

textreasoning
nemotron-3-nano-omni-reasoning

5 in · 5 out
credits / M tokens

OpenAI GPT-OSS 120B Fast

Lite

OpenAI's open 120B model, run for speed — same weights as GPT-OSS 120B. Uses lite credits.

texttools
gpt-oss-120b-fast

15 in · 25 out
credits / M tokens

OpenAI GPT-OSS 20B Fast

Lite

OpenAI's open 20B model — the quickest route to a short answer. Uses lite credits.

texttools
gpt-oss-20b-fast

8 in · 8 out
credits / M tokens

Qwen 3.8 27B

Lite

Qwen 3.8 27B — vision, reasoning, and tool calling. Uses lite credits.

visionreasoningtools
qwen-3-8-27b

60 in · 250 out
credits / M tokens

Qwen 3.6 27B

Lite

Qwen 3.6 27B — vision, reasoning, and tool calling. Uses lite credits.

visionreasoningtools
qwen-3-6-27b

45 in · 190 out
credits / M tokens

Mimo v2.5

Lite

Mimo v2.5 — multimodal model. Uses lite credits.

visiontools
mimo-v2.5

85 in · 85 out
credits / M tokens

Mimo v2.5 Pro

Pro

Mimo v2.5 Pro — multimodal model. Requires pro credits.

visiontools
mimo-v2.5-pro

150 in · 150 out
credits / M tokens

Qwen3.7 Plus

Lite

Alibaba Qwen 3.7 Plus — vision-capable reasoning model with tool use and strong multilingual performance. Uses lite credits.

visionreasoningtools
qwen-large

40 in · 160 out
credits / M tokens
6.4 cached in / M
40 cache write / M

Ezra

Lite

Ezra — free AI model. Uses lite credits.

text
ezra

1 in · 1 out
credits / M tokens

Mercury 2

Lite

Mercury 2 by Inception Labs — diffusion-based language model with reasoning and tool support. Uses lite credits.

textreasoningtools
mercury-2

25 in · 75 out
credits / M tokens

Qwen3Guard 8B

Lite

Alibaba Qwen3Guard 8B safety classifier. Does not respond normally — returns a structured safety verdict: "Safety: Safe/Unsafe\nCategories: ..." instead of conversational text. Ideal for content moderation pipelines.

moderation
qwen3guard-8b

5 in · 5 out
credits / M tokens

Ling 3.0 Flash

Lite

Strong and cheap agentic model designed with token efficiency in mind. Great for coding and agentic work. Uses lite credits.

reasoningtools
ling-3-0-flash

10 in · 10 out
credits / M tokens

Celeris 1

Lite

Celeris 1 — a diffusion LLM tuned for very fast, short, structured responses like classification, extraction, and scoring. Uses lite credits.

texttools
celeris-1

20 in · 70 out
credits / M tokens

Image & video models

Flat rate per generation — no token counting.

DreamShaper 8 LCM

Lite

DreamShaper 8 LCM — near-instant images at rock-bottom cost, simpler detail than premium models.

text-to-image
dreamshaper-8-lcm1 credit / image

FLUX.1 Schnell

Lite

FLUX.1 Schnell — fast, high-quality images at a tiny cost.

text-to-image
flux-1-schnell2 credits / image

Z-Image Turbo

Lite

Z-Image Turbo — instant, budget-friendly images with crisp upscaled output.

text-to-image
z-image-turbo4 credits / image

FLUX.2 Klein 4B

Lite

FLUX.2 Klein 4B — fast generation and single-reference editing up to 2.4 megapixels.

text-to-image
flux-2-klein-4b5 credits / image

GPT Image 1 Mini

Lite

OpenAI GPT Image 1 Mini — affordable image creation and single-reference editing for everyday use.

text-to-image
gpt-image-1-mini7 credits / image

MAI Image 2.6 Flash

Lite

Microsoft MAI Image 2.6 Flash — photorealistic generation and single-reference editing with accurate text rendering.

text-to-image
mai-image-2-6-flash5 credits / image

MAI Image 2.5 Flash

Lite

Microsoft MAI Image 2.5 Flash — quick photorealistic generation and single-reference editing with accurate text rendering.

text-to-image
mai-image-2-5-flash5 credits / image

MAI Image 2.6

Lite

Microsoft MAI Image 2.6 — detailed photorealistic generation and single-reference editing with strong instruction following.

text-to-image
mai-image-2-610 credits / image

FLUX 1.1 Pro

Lite

FLUX 1.1 Pro — fast text-to-image generation with precise dimensions and reproducible seeds.

text-to-image
flux-1-1-pro13 credits / image

FLUX.1 Kontext Pro

Lite

FLUX.1 Kontext Pro — edits an existing image from plain instructions: swap, restyle, refine.

text-to-image
flux-1-kontext-pro13 credits / image

GPT Image 1.5

Lite

OpenAI GPT Image 1.5 — high-fidelity image generation and single-reference editing with fine detail.

text-to-image
gpt-image-1-514 credits / image

GPT Image 2

Lite

OpenAI GPT Image 2 — premium high-resolution images with excellent prompt following.

text-to-image
gpt-image-215 credits / image
Errors

Error codes

All errors return JSON with an error.message field describing what went wrong.

JSON
// Error response body
{
  "error": {
    "message": "Insufficient credits. Please top up your account.",
    "type": "insufficient_credits"
  }
}
CodeStatusMeaning
200OKRequest succeeded. Response body contains the completion.
400Bad RequestMissing or invalid request body fields (model, messages).
401UnauthorizedMissing or invalid Authorization header, or missing/incorrect X-Api-Key-Password for a password-protected key.
402Payment RequiredInsufficient credits. Top up via the dashboard.
404Not FoundUnknown model ID. Check /v1/models (auth required) or /api/models for valid IDs.
429Too Many RequestsRate limit exceeded. Back off and retry.
500Internal Server ErrorUpstream model provider error. Retry with exponential backoff.

Need help?

Join the Discord server and ask — or grab your API key from the dashboard and keep building.