SeeAI spendx, found

See AI cost by model, and stop runaway loops.

Cost by model, provider and conversation, with daily allowances per role and a budget that stops a looping agent.

How it works

Specs

AI spend, in detail

Delivery and data

Delivery
SaaS, from one login.
Isolation
Each customer runs in an isolated workspace with its own database.
Certifications
None held. Frameworks are mapped to and assessed against.

Frameworks

OWASP Top 10 for Agentic Applications
Assessed per agent against: Rate limits counted in each agent's assessment.
See the frameworks

Last reviewed 6 Oct 2026

Allowances at the gatewayIllustrative

Allowances at the gateway: Claims team to gateway; research-agent to gateway (same call, again); gateway to model A (allowed); gateway refused before model B (budget reached); gateway to cost record (tokens, cost).

In short

LLM cost management in ColossalX records what each AI request costs, by model, provider and conversation, from the traffic the gateway serves. Daily request and spend allowances apply per role, model and use, and a budget on turns and tool calls stops a looping agent early. A model with no known price is reported, never costed at zero.

An agent loops on the same call all weekend, and the invoice is the first warning.

ColossalX stops the agent at its loop budget, refuses calls past the allowance, and says why.

How it works

From a looping agent to a stopped one.

One research agent repeats the same tool call. Its loop budget stops it early, and the refusal says which limit it reached, before the bill grows.

Workflow · a looping agent, stopped earlyIllustrative

01 Allowance set

An access policy sets daily requests and spend for one role.

02 Agent loops

The agent repeats the same tool call, turn after turn, without progress.

03 Budget reached

Past its loop budget, the next call is refused with the reason.

04 On the record

The cost of the run and the refusal stay on the record.

What you see

Usage, hour by hour, from the gateway.

Cost is computed from the same requests the gateway serves: volume each hour, what was blocked, and each request's tokens, cost and model, so finance and security read the same figures.

  1. Cost per request
  2. Allowances by role
  3. A loop budget
  4. Cost by conversation
Read the detail, step by step4
  1. Cost per request. Each request records its tokens, its cost and the model that answered. Model prices come from provider discovery and a fallback catalogue. A model with no known price is reported as unpriced, never costed at zero.
  2. Allowances by role. Requests and spend per day, per role, model and use. A model access policy names a use, a role, a provider and model, allow or refuse, and limits: output tokens per request, requests per day and cost per day. The lowest priority number wins.
  3. A loop budget. An agent run has a budget of turns and tool calls. When an agent repeats itself past its budget, the run is stopped and the refusal names the limit. A guardrail profile can also set a daily spend ceiling and a per-answer token ceiling.
  4. Cost by conversation. The ColossalX Assistant shows usage and cost by model and conversation. Workforce chat in the ColossalX Assistant runs through the same gateway, with usage and cost by model and by conversation against a budget.
AI usage and spendIllustrative

An illustrative view of AI usage and spend: agents against their daily allowance, one agent repeating the same call until its allowance stops it and its owner is told, and the day's cost split by model.

How it connectsx, found

Where a cost figure goes next.

Spend is not a separate ledger. The same request records feed the controls and the people who answer for them.

  1. Allowances apply where requests run: one gateway across 31 provider families.

  2. Workforce chat usage and cost by model and conversation, against a budget.

  3. A runaway agent can be contained automatically on the limits you set.

  4. The fleet view shows tokens and cost beside violations and refusals.

Honest by design

What it does, and what it does not.

One cost recordIllustrative

One cost record: Model asked model C; Answered by model C; Tokens 1,240 in · 310 out; Price Unknown; Cost Not costed, reported. Unpriced, not zero.

What it does not do

x, not measured

Costs are as exact as model prices, which come from provider discovery and a fallback catalogue.

All 4 limits
  • Allowances are per day, counted from midnight UTC, not per hour.
  • The per-answer token ceiling caps the answer; it does not refuse the request.
  • Spend is counted for traffic through the ColossalX gateway and Assistant only.

How we know

  • A model with no known price is reported as unpriced, never costed at zero.
  • Daily allowances count the requests actually served, across everyone a policy covers.
  • A refused request says which allowance or budget it reached.
  • The allowances are read on each request, not reconciled at the end of the month.

Questions

Questions buyers ask

How do we track LLM costs by team and model?

Route AI traffic through the ColossalX gateway and each request records its tokens, its cost and the model that answered. Usage then rolls up by model, provider and conversation, and access policies tie it to roles and uses. The hourly series can be exported for finance, so both teams read the same figures.

How do daily request and spend allowances work?

A model access policy sets, for a use, a role and a model, how many requests a day and how much spend a day are allowed, plus a cap on output tokens per request. Allowances count requests actually served since midnight UTC. Once one is used up, further requests are refused with the reason.

How is a looping agent stopped before it runs up the bill?

Each agent run carries a budget of turns and tool calls. When an agent repeats itself past that budget, the next call is refused and the refusal names the limit, so the run stops early rather than at the end of the month. Automatic containment on your own limits is also available.

Can we set different allowances per role and per model?

Yes. Policies are ordered, and each one names a use, everyone or a named role, a provider and model, allow or refuse, and its own limits. The lowest priority number wins, so a premium model can carry a tight allowance for one role while a smaller model stays open for everyone.

Which model providers are covered?

ColossalX governs 31 provider families through one OpenAI-compatible gateway, including self-hosted runtimes such as Ollama, vLLM, LocalAI and LM Studio. Costs follow the prices discovered from each provider, with a fallback catalogue; a model with no known price is reported as unpriced.

Related

Next step

Know your x.

See what your own AI costs: by model, by role and by conversation, with the loops stopped early.

  1. 01Tell us what you run
  2. 02See the four verbs on it
  3. 03Decide where to start