← Back to Blog

July 25, 2026

LLM Cost Management for Agencies: Protect Your Margin with a Token Usage Dashboard

In the chatbot resale business, one factor directly drives your gross margin yet is routinely ignored: LLM (AI) usage cost.

You sell clients a fixed monthly fee — but AI cost varies with conversation volume. Agencies that can't see this structure end up scaling without knowing which clients are profitable and which are losses.

This guide covers how agencies track LLM cost per client and protect margin in practice.

Why LLM Cost Decides Your Margin

The P&L of chatbot resale is simple:

Gross margin = client price (fixed) − platform fee − LLM cost (variable)

Only the last item is variable. LLM cost accrues per "token" (the unit of text AI processes) and moves with:

  • Conversation volume — busier clients cost more
  • Document volume — more reference documents per answer means more processing
  • Model choice — flagship vs. lightweight model prices differ dramatically

So even if you assume "¥50,000/month price, ~¥10,000 average cost," it's entirely normal for one heavy-use client to quietly balloon to ¥30,000. If you can't see it, you can't act on it.

Step 1: Measure Usage Per Client, Per Model — Don't Estimate

Start with measurement, not guesswork.

OneBot's admin dashboard includes a token usage dashboard as standard, with cost visible by:

  • Project (= client) — who consumes how much AI cost
  • Model — which AI model costs what
  • Period — monthly trends to catch anomalies

Crucially, these figures come from the AI provider's official token-count API — measured, not estimated from character counts, so they don't drift from actual billing. When the AI makes multiple tool calls (e.g., via MCP integration), every token is counted.

Agency monthly routine: check per-client cost on the 1st → tabulate cost ratio against each client's price → flag clients above your threshold (say 40%) for action. Five minutes.

Step 2: Optimize with Model Mix

Your second lever is model selection. OneBot supports multiple AI models, assignable per use case:

Use caseModel tierRationale
Mostly routine FAQLightweight (e.g., Gemini 2.5 Flash-Lite)Low unit cost at high volume
Complex document reasoningStandard–flagship (Gemini 2.5 Flash, GPT-5)Answer quality first
Balance of bothMid-tier (e.g., GPT-5 mini)The realistic default for most SMB accounts

The point is not "one model for everyone" but matching the model to each client's inquiry profile. Switching an FAQ-heavy client to a lightweight model routinely cuts that client's LLM cost substantially.

Multi-model support is also a structural hedge: when a provider re-prices or a model's quality shifts, you can switch. Single-model platforms leave you absorbing every price increase.

Step 3: Quotas — Structurally Prevent Overuse

After measuring and optimizing, design the ceiling.

OneBot assigns each client (tenant) a message-volume plan, with three configurable overage behaviors:

Overage behaviorBest for
Block — pause at the capLow-price plans where cost must be fixed
Metered — bill overage per unitStandard plans monetizing extra usage
Notify-only — alert without stoppingKey accounts; observation periods

Threshold alerts email you automatically (EN/JA) — no more "costs doubled before anyone noticed."

This turns plan design into real product design: "Light = 1,000 messages/month, blocks on overage; Standard = 5,000, metered" — products with a knowable cost ceiling.

Worked Example: Margin Impact at 20 Clients

Say you run 20 clients averaging ¥50,000/month (¥1M revenue).

  • Unmanaged: cost ratios scatter from 10% to 60%; at 35% blended, margin is ¥650k — with loss-making clients invisible.
  • Managed: identify the 5 high-cost clients → switch 2 to lightweight models, re-plan 2, move 1 to metered overage. Blended cost ratio drops to 25% → margin ¥750k (+¥100k/month).

Illustrative numbers, universal structure: visibility → model mix → quota design turns margin from luck into engineering.

FAQ

Q1. What exactly is a token?

The smallest unit of text AI processes — roughly 1–2 Japanese characters (or ~4 English characters) per token. LLM fees are token count × unit price.

Q2. Are the usage numbers accurate?

Yes — OneBot uses the AI provider's official count API. Measured, not estimated, so they match billing.

Q3. Can each client run a different model?

Yes. Model and response parameters are configured per project, so you can match each client's inquiry profile.

Q4. Can we set our own overage pricing?

Yes — per plan, you choose the overage behavior (block/metered/notify) and unit price to fit your product design.

Q5. Can we use cost data in client reports?

Usage and conversation analytics are available per tenant and feed monthly reporting. See the monthly report automation guide for a deep dive.

Q6. What happens if a model gets re-priced?

Per-model usage visibility lets you size the impact immediately and switch affected clients' models to protect your cost structure.

Summary

  • LLM cost is the variable that moves resale margin — you can't manage what you can't measure
  • OneBot ships a token usage dashboard (per client, per model, measured) as standard
  • Model mix × three-mode quotas turn cost into a designed variable, not a surprise

See the token dashboard yourself in a trial environment — 14 days free, no credit card.

👉 Start your free trial