Models
Every turn, q15 decides which model to call. There is no model list in the configuration: each provider’s own roster is the source of truth, and the choice is made per request.
The roster
Section titled “The roster”The agent queries every configured provider for its model list at startup, enriches what it finds with capability metadata from models.dev, and refreshes it periodically. The result is the roster.
- A provider that is unreachable at startup contributes nothing, and the agent starts with the rest.
- An empty roster is a startup failure. If
q15-agentexits immediately, check the provider list and the matching secret file first.
See Agent config for the provider block that produces the roster.
Choosing a model
Section titled “Choosing a model”Selection is two stages: filter, then rank.
flowchart TD
Req[Turn arrives] --> Infer[Derive the request's requirements]
Infer --> Filter[Filter the roster by capability]
Filter --> Any{Any candidate left?}
Any -->|No| Fail[The turn fails]
Any -->|Yes| Rank[Rank deterministically]
Rank --> Call[Call the first model]
Call --> Ok{Call succeeded?}
Ok -->|Yes| Done[Answer]
Ok -->|No| Next{Another candidate?}
Next -->|Yes| Call
Next -->|No| Fail
Requirements. Text is a hard requirement. Media is not: an image or audio part is rendered down to a text hint for a model that cannot take it inline, rather than excluding that model from selection. Tool calling is required only when no candidate in the current set could fall back to a tool-free call — otherwise the agent may omit tools and call a model that does not support them.
Ranking. Candidates that survive the capability filter are ranked by cheapest known cost, then best known benchmark, then the order the provider returned them. The ranking is deterministic: the same roster and the same request produce the same first choice. When the top choice is an ambiguous tie, the agent logs the reason rather than guessing silently.
The current model is tried first on every path, then the rest of the ranked candidates. If a call fails, the next candidate is tried; a turn fails only when no candidate is left.
The current model is runtime state
Section titled “The current model is runtime state”The selection is stored on the agent, not in the config file. Two consequences:
- A pin that no longer exists in the roster does not stop the agent from starting — the seed is not validated at startup.
- Every change you make at runtime is validated against the live roster, so you cannot switch to a model that is not currently reachable.
The current pair persists across restarts.
You change it from chat:
| Tool | Effect |
|---|---|
list_providers |
Which providers are reachable, and their model counts. |
list_models |
The roster, with capabilities and context windows. |
switch_model |
The model that answers interactive turns. |
switch_cognition_model |
The model for one background cognition job. |
switch_cognition_model takes a job type — verification_review, working_memory.consolidate or
semantic_memory.extract. A cognition job uses its own override when one is set, and otherwise
inherits the interactive model. See Memory for what the jobs do.
Running a local model
Section titled “Running a local model”Local is the direction of the project, and one provider block is what stands between today and it:
providers: - name: local type: ollama base_url: http://ollama:11434The base_url is the part to get right. It has to resolve from inside the container, not from
your shell, so localhost is wrong unless Ollama runs in the same stack. A tested value for a
host-run Ollama is in development; the address that works is the one your container engine
publishes for the host.
Two things to know before you commit to it:
- The model has to fit the machine. The stack itself needs about 2 GB of RAM; a local model needs whatever that model needs, and running a large one on a small box is a slow experience rather than a broken one.
- Embeddings are separate. The search backend is a second provider block, and on the shipped
Compose stack it is the local
q15-teicontainer. See Memory.
Discovery
Section titled “Discovery”The catalog is built in memory at startup and refreshed periodically; it is not a file you edit and not a table you maintain. Provider types determine how a provider is reached and how it authenticates:
| Type | Authenticates with | Notes |
|---|---|---|
ollama |
nothing, or key_env for cloud |
base_url defaults to the local Ollama endpoint when omitted. |
openai-compatible |
key_env |
Any endpoint that speaks the OpenAI API, for example Moonshot. |
openai-codex |
auth.json produced by q15-auth |
The OpenAI OAuth flow. See Credentials. |