Skip to main content
Reporting and analytics

AI Usage Analytics

See what the AI costs and where the spend goes: calls, tokens, cache savings, estimated cost, a month-end projection, model fallbacks and the AI calls behind a single ticket.

Written By Chris Scaminaci

Last updated About 2 hours ago

Every ticket QuantumOps analyses spends AI tokens. AI Usage Analytics makes that spend visible: how many calls were made, how many tokens they used, what caching and batching saved, what it all cost, and where the money goes by model, function and ticket. Administrators use it to watch the bill and to find patterns that waste money.

Before you start:

  • You need the Administrator role. The AI Usage Analytics link under Analytics & Reporting in the sidebar is shown only to Administrators, and anyone else who types the address gets an access-denied page. See Access denied and sign-in errors.
  • The page covers AI calls that QuantumOps recorded for your organisation.

Choose the period

  1. Open AI Usage Analytics from Analytics & Reporting in the sidebar.
  2. Pick the dates in the Select Date Range field. The page opens on the last 30 days, ending today.
  3. Select Refresh to read the data again.

The subtitle names your organisation's time zone. The dates you pick are read in that zone. Three things use UTC instead: the daily buckets of the Daily Token Usage Trend chart, the daily costs behind the cost anomaly check, and the times in the ticket trace.

Read the headline figures

FigureWhat it tells you
API CallsHow many AI calls were made in the period.
Total TokensInput and output tokens together, with the split underneath.
Cache EfficiencyThe share of input served from the provider's cache, with cache reads and cache writes underneath. A higher share costs less.
Est. CostEstimated spend for the period. If some calls could not be priced, the tile says how many it left out. See How cost is priced.
Avg Tokens/CallInput plus output tokens, averaged per call.
Categorization Calls/TicketThe average number of categorization calls per ticket that had any. The line underneath gives the alarm threshold, more than 3 calls per ticket, or how many tickets went over it. The value shows a dash when there were none.
Cache SavingsWhat caching saved compared with paying the full input rate. A trailing est. means a model has no published cache price, so a standard ratio was used.
Batch SavingsThe estimated saving from calls sent through the provider's batch service. It says so when there were none.
Projected Month-EndThe estimated cost per day over the period you picked, multiplied by the days in the current month. A quiet week understates it, and a busy week overstates it.
Error RateThe share of calls whose response was not a success, with the median response time underneath. The tile turns red above 5%.

Alerts above the charts

These appear only when there is something to report.

Categorization re-fire alarm

The Categorization re-fire alarm names the tickets that needed more than 3 categorization calls in the period, as chips such as #12345 · 4 calls · 3.8K in. Up to 12 are shown, then +N more.

The alarm exists to catch a ticket whose categorization keeps firing, which adds cost. Trace the ticket to see each call. See Trace the AI calls for one ticket.

Monthly budget banner

When you set a Monthly AI Budget (USD), a banner compares the projected month-end cost with it. It reads "On track" with the share of the budget used, or says the projection exceeds the budget. With no budget set, there is no banner.

Set the budget in the AI Provider Configuration section of Tenant settings, which you open from the Settings button in the header (see Tenant settings). The steps are in Set a monthly AI budget. The budget is a reference figure. It does not limit or stop AI calls.

Cost anomaly

The Cost anomaly callout lists days on which one function cost much more than usual. Each chip shows the function, the day, the cost, how many standard deviations above normal it was, and the normal level in brackets. A day is flagged when all of these hold:

  • The function has at least 4 days of data in the period.
  • The day cost at least $0.01.
  • The cost is at least 2 standard deviations above the average of that function's other days in the period. If all the other days cost the same, any higher day is flagged and the chip says spike.

Because the comparison uses only the days you picked, a short period can hide an anomaly. Widen the period to check.

Break the spend down

Each chart answers one question. All of them use the period you picked.

ChartShows
Token Usage by ModelTokens per model, as a share of the total.
Daily Token Usage TrendInput and output tokens per day.
Usage by Operation TypeTokens per function, meaning the kind of work the call did, such as ticket analysis or a sanity check. Calls with no recorded function show as Unclassified.
Calls by ProviderCalls per AI provider.
Cost by Function (USD)Estimated cost per function. Select a bar to open the drill-down.
Cost by Model (USD)Estimated cost per model.
Top Calling Methods (by Token Volume)Which part of QuantumOps made the heaviest calls, shown by its short name. It is mostly useful when you work with support to trace heavy usage.

The cost charts leave out calls that could not be priced, so they can show less than the token charts suggest.

Review the function and model table

Cost by Function & Model lists one row per function and model. It is grouped by function, with a subtotal for each. The columns are Function, Model, Provider, Calls, Input Tokens, Output Tokens, Cache Reads, Cache Writes, Avg/Call, Est. Cost and Pricing Basis.

Pricing Basis says how the cost was found: Catalog rates, Unpriced (not in catalog), or Partial with the number of unpriced calls.

Select Export to download the table as an Excel file.

Drill into one function

  1. Select a bar in Cost by Function (USD). The Function Drilldown panel opens below the charts.
  2. Read the totals at the top: Calls, Cost (range, billed), Proj. month-end (billed), Avg in/call and Avg out/call.
  3. Compare the models the function ran on in Per-Model Breakdown, and read the Daily Cost Trend.
  4. To test a cheaper model, use What-If: reprice this function on another model. Pick a model from the list.
  5. Read the result: the estimated month-end cost on that model, with the difference from the current model. The panel says so when the model you picked has no price.
  6. Select Close to dismiss the panel.

The what-if prices both models the same way: with the function's actual average tokens per call, projected over a month of calls, at today's catalog rates. It does not compare against what you were billed, so picking the model already in use shows no difference. It estimates only the price. It cannot tell you whether the other model answers as well.

See model fallbacks and rerouted spend

Model Fallback / Reroute Spend shows calls that ran on a different model from the one configured, for example because a provider rate-limited or refused the call. The total is shown as rerouted spend. The table lists the Configured model, the Actual model, the Reason, the Calls, the Cost, and whether the cost was Priced (yes, or legacy).

When every call ran on its configured model, the section says so. For where models are chosen, see AI providers and models.

Trace the AI calls for one ticket

  1. Find the Per-Ticket AI Call Trace section at the bottom of the page.
  2. Type a HaloPSA ticket number in the Ticket id (e.g. 12345) field.
  3. Select Trace, or press Enter.
  4. Read the summary line (the ticket, the number of calls and their total cost) and the table: When (UTC), Model, Function, Method, In, Out, Cost, Latency and Status.

The trace lists every call recorded for that ticket, not only those in the period you picked. A warning mark beside a model means the call fell back to it from another model. If nothing was recorded, the page says "No AI calls recorded for ticket #12345 in this tenant."

How cost is priced

  • Each call is priced when it is made, using the rate in force at that moment. History does not change when a provider changes its prices.
  • Calls recorded before this pricing existed are priced at today's catalog rates. In the fallback table they show as legacy.
  • A call on a model that is missing from the price list is unpriced. It is left out of Est. Cost and counted in the tile's note. QuantumOps never guesses a price from the model's name or tier.
  • Cached tokens are priced in addition to the input tokens, not taken out of them.

These are estimates from QuantumOps's own records. Your provider's invoice is the authoritative figure.