AI providers and models
Switch on the AI providers QuantumOps uses, enter their keys, choose the model for each job, override a single operation, and set an optional monthly AI budget.
Written By Chris Scaminaci
Last updated About 2 hours ago
QuantumOps needs at least one AI provider to analyse tickets, answer chat questions and draft replies. This page covers the AI settings in Tenant settings: which providers are switched on and with which keys, which model does which job, how to override the model for a single operation, and the optional monthly AI budget. Only the Administrator role can open these settings.
Before you start:
- You need the Administrator role.
- Have an API key ready for each provider you want to use. You create these keys with the provider, not in QuantumOps.
Open the AI settings
- Select the gear icon in the page header (its tooltip is Settings). The page is titled Tenant Configuration.
- Scroll to AI Provider Configuration.
Everything on this page sits in that one section: the providers, AI Model Configuration, Monthly AI Budget (USD) and Per-Operation Model Overrides. The section has a single Save button that saves all of it. A Save on this page stores every section except Communication Policy, so a change you made in another section and have not checked, such as untested HaloPSA credentials, is saved with it. See Save your changes and Tenant settings for the rest of the page.
Switch on a provider
- Turn on the switch of the provider.
- Enter its key in the key field that appears. The field is masked.
- Select Save.
QuantumOps will not save while a switched-on provider has no key. It tells you which key is missing.
When you open the page, keys you saved earlier are already in their fields, masked. Keep these rules in mind:
- If you turn OpenAI or Anthropic (Claude) off and select Save, QuantumOps removes the stored key for that provider. You enter the key again if you switch the provider back on.
- A switched-on provider needs a key in its field every time you save. This is the same for OpenAI, Anthropic (Claude) and Together.ai.
- Together.ai is the exception to the removal rule. Switching it off keeps the stored key, so a model slot or override that still names a Together.ai model keeps working (see Turn off Together.ai below). To remove the key on purpose, turn on Clear stored key on save (it is shown while a key is stored) and select Save. Switch Together.ai off first, because a switched-on provider without a key raises the AI configuration issue banner.
When Anthropic is switched on, a Use Voyage AI for Embeddings switch also appears under its key. It has no key field on this page. The embeddings behind semantic search are managed for you (see Vector Store Configuration in Tenant settings).
The AI configuration issue banner
When a provider is switched on but has no key, a red banner titled AI configuration issue appears at the top of every page, for everyone who is signed in. It reads, for example, "1 AI provider(s) enabled without API keys". Select Fix now to open Tenant settings, then enter the missing key or switch the provider off. The Dismiss button hides the banner for now, and it returns the next time the page loads while the key is still missing.
Turn off Together.ai
Switching Enable Together.ai Integration off does not stop calls to Together.ai by itself. QuantumOps picks the provider from the model chosen for each job, so any model slot or per-operation override that still points at a Together.ai model keeps calling Together.ai, and keeps being billed by it. When the integration is off but a key is still stored, an amber notice titled Together.ai is disabled but a key is still stored says whether any selection still points at a Together.ai model. To stop Together.ai completely:
- Change every model slot and every override that uses a Together.ai model to a model from another provider.
- Turn on Clear stored key on save.
- Select Save.
Set the Anthropic service tier and rate limits
Anthropic limits how much a key can use each minute, and the limit depends on your Anthropic service tier. Under the Anthropic key, find Rate Limiting Configuration and pick your tier in Service Tier. You can find your tier in the Anthropic Console, in its limits settings.
Choosing a tier fills in the three number fields. You can change them afterwards.
Together.ai has its own Token Capacity/Min, Request Capacity/Min and Max Wait Time (minutes) fields under its key. It has no tier list.
Choose the model for each job
AI Model Configuration has one slot for each kind of job. Every slot except Sanity Check Provider opens the model picker described below. Sanity Check Provider is a plain list of the three providers.
Use the model picker
Each slot that uses the picker shows a button with the model now chosen: its name, its provider, a short cost note, its context size and its cost tier, plus a star when it is recommended. Select the button, or Change, to open the picker. Its title names the slot, for example Select Primary Model.
- Search the list by typing part of a model's name, its model id or its provider.
- Narrow the list with the provider buttons at the top. All shows every provider.
- Scan the list. Recommended models are pinned at the top for this slot, then models are grouped by provider. Deprecated models are listed last in their group, and your saved choice carries a Current badge.
- Select a model to read its card. The card shows Context window, Max output, Knowledge cutoff, Pricing (per 1M tokens) for input and output, Capabilities, What to use it for, Strengths and Recommended for. A model the catalog marks as not active shows an Inactive badge.
- To give the same model to specific operations as well, expand the "Also assign to functions" section under the card and tick the operations. Each row shows the model the operation uses now, or Default.
- Confirm with Select. You can also press Enter, or double-click a model. Nothing is saved until you select Save in the section.
A few things to know:
- Models from a provider that is not switched on are shown greyed out. You can still select one, but the card warns that you need to enable the provider before you rely on the model.
- A Deprecated model shows the month it was deprecated and the reason. When a successor exists, the card offers a button to use it instead.
- If a slot holds a model that is no longer in the list, the slot shows Unknown model and the card reads Unknown / legacy model. Pick a model from the list to replace it.
- The Large Context Model (1M) slot lists only models with a native 1M window. A saved model that does not qualify shows a note on its card.
- The Triage/Dispatch Model, Large Context Model (1M) and Sanity Check Model slots, and every override row, have a clear button (a cross) that empties the slot so the default applies.
Override the model for one operation
By default each operation uses the model of its group. The Per-Operation Model Overrides list shows the groups, with the model each one uses now:
- Primary Model Operations use the Primary Model.
- Secondary Model Operations use the Secondary Model.
Triage and the sanity check are not in this list, because they have their own slots above.
Other pages speak of a "tenant default" model. For Qubit chat it is the Chat & Conversations row, for Agent Runners the Agent Runners row, and for the Slack and Teams assistant the Channel Assistant row. Each of these rows uses the Primary Model until you override it. A Slack or Teams channel can also have its own model; see Channel Assistant settings.
- On the operation's row, select Using Default. The row now shows the group's current model.
- Select that model to open the picker, choose another model, and confirm with Select.
- Select Save.
An override beats the slot for that one operation. To undo it, select the red cross on the row (its tooltip is Reset to default) and select Save.
Set a monthly AI budget
Monthly AI Budget (USD) is optional. Leave it empty and the field shows No budget configured. When you enter an amount, AI Usage Analytics compares the projected cost for the month with it and shows a budget banner: on track while the projected month-end cost is within the budget, and a warning when the projection goes over. The cost anomaly flags on that page do not depend on the budget.
The budget never stops AI calls. It only drives that banner. If an AI provider reports that your account has reached its own billing limit, QuantumOps pauses AI processing and e-mails the primary contact. See Pausing and resuming AI processing.
Related pages
- Tenant settings: the rest of the page, including billing details and saving.
- AI Usage Analytics: what the AI calls cost, and the budget banner.
- Q-Director: AI Processing: the analysis settings that use these models.
- Setup wizard: AI and Q-Director: the same choices during first setup.
- Pausing and resuming AI processing: what happens when a provider's limit is reached.
- Communication policy for AI replies: the rules every AI-drafted reply follows, on the same page.
Was this helpful?
Still need help? Ask the team