Skip to main content
Q-Director and AI settings

AI providers and models

Switch on the AI providers QuantumOps uses, enter their keys, choose the model for each job, override a single operation, and set an optional monthly AI budget.

Written By Chris Scaminaci

Last updated About 2 hours ago

QuantumOps needs at least one AI provider to analyse tickets, answer chat questions and draft replies. This page covers the AI settings in Tenant settings: which providers are switched on and with which keys, which model does which job, how to override the model for a single operation, and the optional monthly AI budget. Only the Administrator role can open these settings.

Before you start:

  • You need the Administrator role.
  • Have an API key ready for each provider you want to use. You create these keys with the provider, not in QuantumOps.

Open the AI settings

  1. Select the gear icon in the page header (its tooltip is Settings). The page is titled Tenant Configuration.
  2. Scroll to AI Provider Configuration.

Everything on this page sits in that one section: the providers, AI Model Configuration, Monthly AI Budget (USD) and Per-Operation Model Overrides. The section has a single Save button that saves all of it. A Save on this page stores every section except Communication Policy, so a change you made in another section and have not checked, such as untested HaloPSA credentials, is saved with it. See Save your changes and Tenant settings for the rest of the page.

Switch on a provider

ProviderSwitchKey field
OpenAIEnable OpenAI IntegrationOpenAI API Key
Anthropic (Claude)Enable Anthropic (Claude) IntegrationAnthropic API Key
Together.aiEnable Together.ai IntegrationTogether.ai API Key
  1. Turn on the switch of the provider.
  2. Enter its key in the key field that appears. The field is masked.
  3. Select Save.

QuantumOps will not save while a switched-on provider has no key. It tells you which key is missing.

When you open the page, keys you saved earlier are already in their fields, masked. Keep these rules in mind:

  • If you turn OpenAI or Anthropic (Claude) off and select Save, QuantumOps removes the stored key for that provider. You enter the key again if you switch the provider back on.
  • A switched-on provider needs a key in its field every time you save. This is the same for OpenAI, Anthropic (Claude) and Together.ai.
  • Together.ai is the exception to the removal rule. Switching it off keeps the stored key, so a model slot or override that still names a Together.ai model keeps working (see Turn off Together.ai below). To remove the key on purpose, turn on Clear stored key on save (it is shown while a key is stored) and select Save. Switch Together.ai off first, because a switched-on provider without a key raises the AI configuration issue banner.

When Anthropic is switched on, a Use Voyage AI for Embeddings switch also appears under its key. It has no key field on this page. The embeddings behind semantic search are managed for you (see Vector Store Configuration in Tenant settings).

The AI configuration issue banner

When a provider is switched on but has no key, a red banner titled AI configuration issue appears at the top of every page, for everyone who is signed in. It reads, for example, "1 AI provider(s) enabled without API keys". Select Fix now to open Tenant settings, then enter the missing key or switch the provider off. The Dismiss button hides the banner for now, and it returns the next time the page loads while the key is still missing.

Turn off Together.ai

Switching Enable Together.ai Integration off does not stop calls to Together.ai by itself. QuantumOps picks the provider from the model chosen for each job, so any model slot or per-operation override that still points at a Together.ai model keeps calling Together.ai, and keeps being billed by it. When the integration is off but a key is still stored, an amber notice titled Together.ai is disabled but a key is still stored says whether any selection still points at a Together.ai model. To stop Together.ai completely:

  1. Change every model slot and every override that uses a Together.ai model to a model from another provider.
  2. Turn on Clear stored key on save.
  3. Select Save.

Set the Anthropic service tier and rate limits

Anthropic limits how much a key can use each minute, and the limit depends on your Anthropic service tier. Under the Anthropic key, find Rate Limiting Configuration and pick your tier in Service Tier. You can find your tier in the Anthropic Console, in its limits settings.

Service tierTokens per minuteRequests per minuteDescribed on screen as
Tier 14,0005Good for development
Tier 220,00025Moderate production use
Tier 340,00050Standard production
Tier 4200,0001,000High-volume workloads
Monthly Invoicing500,000 or more2,500 or moreEnterprise

Choosing a tier fills in the three number fields. You can change them afterwards.

FieldWhat it holds
Token Capacity/MinThe tokens per minute your Anthropic account can process (1,000 to 10,000,000).
Request Capacity/MinThe requests per minute your account allows (10 to 100,000).
Max Wait Time (minutes)How long a request may wait for capacity before it times out (1 to 30). Tiers 1 to 4 and Monthly Invoicing fill in 2, 3, 5, 8 and 10.
Enable Smart QueuingWhen a rate limit is reached, requests are queued and retried instead of failing. It is recommended for production.

Together.ai has its own Token Capacity/Min, Request Capacity/Min and Max Wait Time (minutes) fields under its key. It has no tier list.

Choose the model for each job

AI Model Configuration has one slot for each kind of job. Every slot except Sanity Check Provider opens the model picker described below. Sanity Check Provider is a plain list of the three providers.

SlotUsed for
Primary ModelThe default model for bulk ticket analysis, categorisation and sentiment. The model marked Recommended gives the best balance of quality and cost. If the slot is empty when you save, QuantumOps fills it with the Recommended model.
Secondary ModelTriage, incremental refresh and quick checks. A faster, cheaper model is a good fit. It is the same value as the AI Model field on the AI Processing tab (see Q-Director: AI Processing).
Triage/Dispatch ModelTriage recommendations and dispatch predictions. A fast model suits it. If you clear it, QuantumOps uses the Secondary Model.
Reasoning ModelComplex planning and deep analysis. A high-capability model suits it.
Large Context Model (1M)Work that needs more than 200K tokens of context, up to 1M. Only models with a native 1M context window are listed. You can clear it.
Sanity Check ProviderThe provider that cross-checks ticket analysis: OpenAI (GPT), Anthropic (Claude) or Together.ai. It starts as OpenAI. Changing it empties the sanity check model.
Sanity Check ModelThe model for those checks, chosen from the selected provider's models only. Leave it empty to use the provider's default, which the field describes as a fast model.

Use the model picker

Each slot that uses the picker shows a button with the model now chosen: its name, its provider, a short cost note, its context size and its cost tier, plus a star when it is recommended. Select the button, or Change, to open the picker. Its title names the slot, for example Select Primary Model.

  1. Search the list by typing part of a model's name, its model id or its provider.
  2. Narrow the list with the provider buttons at the top. All shows every provider.
  3. Scan the list. Recommended models are pinned at the top for this slot, then models are grouped by provider. Deprecated models are listed last in their group, and your saved choice carries a Current badge.
  4. Select a model to read its card. The card shows Context window, Max output, Knowledge cutoff, Pricing (per 1M tokens) for input and output, Capabilities, What to use it for, Strengths and Recommended for. A model the catalog marks as not active shows an Inactive badge.
  5. To give the same model to specific operations as well, expand the "Also assign to functions" section under the card and tick the operations. Each row shows the model the operation uses now, or Default.
  6. Confirm with Select. You can also press Enter, or double-click a model. Nothing is saved until you select Save in the section.

A few things to know:

  • Models from a provider that is not switched on are shown greyed out. You can still select one, but the card warns that you need to enable the provider before you rely on the model.
  • A Deprecated model shows the month it was deprecated and the reason. When a successor exists, the card offers a button to use it instead.
  • If a slot holds a model that is no longer in the list, the slot shows Unknown model and the card reads Unknown / legacy model. Pick a model from the list to replace it.
  • The Large Context Model (1M) slot lists only models with a native 1M window. A saved model that does not qualify shows a note on its card.
  • The Triage/Dispatch Model, Large Context Model (1M) and Sanity Check Model slots, and every override row, have a clear button (a cross) that empties the slot so the default applies.

Override the model for one operation

By default each operation uses the model of its group. The Per-Operation Model Overrides list shows the groups, with the model each one uses now:

  • Primary Model Operations use the Primary Model.
  • Secondary Model Operations use the Secondary Model.
OperationWhat it coversGroup
Ticket AnalysisFull ticket analysis and tonality assessmentPrimary
Chat & ConversationsInteractive chat, response synthesis and conversationPrimary
Research & EnrichmentDeep research and web search enrichmentPrimary
QBR ReportsQuarterly business review documents and insightsPrimary
Profile GenerationClient, agent and user profilesPrimary
Customer SuccessCSAT scoring, sentiment and customer success measuresPrimary
Support TrajectoryPredicting support paths and resolution patternsPrimary
Interactive GuideAI responses in the interactive guidePrimary
Agent RunnersAgent Runner goals run on their ownPrimary
Channel AssistantQubit answering in Slack and Microsoft TeamsPrimary
CategorizationTicket category classificationSecondary
Incremental RefreshLight re-analysis when new actions arriveSecondary
Stale Ticket DetectionFinding stale and idle ticketsSecondary
Utility TasksSmall jobs such as repairing JSON, reading times and pulling out keywordsSecondary
Dispatch Workspace SetupTurning a sentence on the dispatch board into workspace settingsSecondary

Triage and the sanity check are not in this list, because they have their own slots above.

Other pages speak of a "tenant default" model. For Qubit chat it is the Chat & Conversations row, for Agent Runners the Agent Runners row, and for the Slack and Teams assistant the Channel Assistant row. Each of these rows uses the Primary Model until you override it. A Slack or Teams channel can also have its own model; see Channel Assistant settings.

  1. On the operation's row, select Using Default. The row now shows the group's current model.
  2. Select that model to open the picker, choose another model, and confirm with Select.
  3. Select Save.

An override beats the slot for that one operation. To undo it, select the red cross on the row (its tooltip is Reset to default) and select Save.

Set a monthly AI budget

Monthly AI Budget (USD) is optional. Leave it empty and the field shows No budget configured. When you enter an amount, AI Usage Analytics compares the projected cost for the month with it and shows a budget banner: on track while the projected month-end cost is within the budget, and a warning when the projection goes over. The cost anomaly flags on that page do not depend on the budget.

The budget never stops AI calls. It only drives that banner. If an AI provider reports that your account has reached its own billing limit, QuantumOps pauses AI processing and e-mails the primary contact. See Pausing and resuming AI processing.