# q15, the whole manual
Generated from the documentation source. Every section is one page of https://q15.co.
Where a page names a command that does not exist yet, it says so.
---
# q15
_A self-hosted AI agent runtime for one person. Persistent memory, a durable workspace, real command execution, and two clients, in a browser and in Telegram._
## Install
```bash title="on your server"
curl -fsSL https://q15.co/install.sh | sh
q15 init
q15 up
q15 status
```
| Command | What it does |
| ------------ | ------------------------------------------------------------------------------ |
| the `curl` | One signed binary into `~/.local/bin`. No sudo, nothing installed system-wide. |
| `q15 init` | Names the agent, picks a model, sets the address. |
| `q15 up` | Renders the units, starts the stack, waits for health. |
| `q15 status` | What runs, and what the agent can reach. |
That is the whole install. `init` prints the address of the browser client: open it, create your
passkey, send a message.
There is no `q15` binary in a release. The tickets that produce it are filed and none has shipped,
so those four commands do not work today. Until they do, the published images are what runs, and
[the reference walks that path end to end](/reference/#run-it-today-end-to-end): six containers, a
Compose file and a handful of secret files.
q15 is an AI agent runtime you install on a server you control. One person runs one q15, keeps it
running, and talks to it from a browser or from Telegram. Everything it knows and everything it has
produced (memory, files, jobs, turn history) is stored as files in volumes you own, so restarts and
updates do not take its work away.
## What you get
A passkey-locked browser app with streaming chat, history and attachments, or a Telegram bot.
Both reach the same agent, the same transcript and the same memory.
Every command's outbound traffic goes through an egress proxy you configure, and the tokens for
what the agent reaches live there. The agent reads data; it never reads the key.
Identity, extracted knowledge, working state and the full turn history live under `/memory` as
files; `/workspace` is a project tree, not scratch space. Both carry across deploys.
[Scheduled jobs](/jobs/) run as their own isolated agent turns, with their own model pins and a
per-job tool allow-list, and report back to the chat that created them.
## What it looks like
One turn in the browser client: your question, what the agent did about it, and the answer. Tool
arguments render as fields instead of raw JSON, reasoning sits in a block you can collapse, and the
same conversation is what Telegram writes into.
## Who it is for
- You are comfortable with SSH, containers, DNS and TLS, and you are willing to operate a daemon.
- You want custody of your own memory, workspace, jobs and model choice, and to keep them after the
next update.
- You are one person. q15 has one owner and no multi-user model.
## Who it is not for
- Teams. There is no shared account, no roles and no per-user isolation.
- Anyone who wants an account on someone else's server rather than a stack on their own.
- Anyone who needs prompts kept away from model providers without running local models. If you use a
hosted provider, your prompts go to it.
- Anyone who wants install-and-forget. The installer is not built yet, and backups and reboots are
yours either way.
[Why q15](/why-q15/) states what the design does and does not claim, in nine lines.
## Where to start
| If you want to | Read |
| ---------------------------------- | ---------------------------------------------------------------------------------- |
| Run it on your own server | [Install q15](/install/) |
| Know what it does not claim first | [Why q15](/why-q15/) |
| Pick a path by what you want to do | [Start here by intent](/start-here/) |
| Reach it from your phone | [Chat](/chat/) then [Telegram](/telegram/) |
| Let it work without you | [Jobs](/jobs/) |
| Learn what it can actually do | [Tools](/tools/) and [Skills](/skills/) |
| Run a local model | [Models](/models/) |
| Read, search or move its memory | [Memory](/memory/) and [Workspace](/workspace/) |
| Understand the boundaries | [Architecture](/architecture/) and [Security](/security/) |
| Keep it running and keep the data | [Updating](/updating/), [Backups](/backups/), [Troubleshooting](/troubleshooting/) |
| Look up a key, a path or a command | [Reference](/reference/) |
## For coding agents
This site ships a plain-text index for agents: [`/llms.txt`](/llms.txt) lists every page with a one
line description, and [`/llms-full.txt`](/llms-full.txt) concatenates all of them. Every page is also
served as Markdown, by adding `.md` to its URL: [`/install.md`](/install.md), [`/reference.md`](/reference.md).
If you work with a coding agent of your own, it can read those instead of you pasting pages:
```
Read https://q15.co/llms.txt, then https://q15.co/llms-full.txt, and answer my
questions about installing and running q15 from them.
```
---
# Install q15
_Four commands on a server you control. The installer is in development; the published images are what runs today._
q15 installs with four commands.
```bash title="on your server"
curl -fsSL https://q15.co/install.sh | sh
q15 init
q15 up
q15 status
```
| Command | What it does |
| ------------ | ------------------------------------------------------------------------------ |
| the `curl` | One signed binary into `~/.local/bin`. No sudo, nothing installed system-wide. |
| `q15 init` | Names the agent, picks a model, sets the address. |
| `q15 up` | Renders the units, starts the stack, waits for health. |
| `q15 status` | What runs, and what the agent can reach. |
`init` prints the address of the browser client. Open it, create your passkey, send a message.
There is no `q15` binary in any release. The tickets that produce it are filed and none has
shipped, so those four commands do not work today. Until they do, [run the published
images](/reference/#run-it-today-end-to-end): that page has the Compose file, the volumes, the
secret files and the health checks, and it is how every deployment that exists was built.
## What `init` asks
- **The agent's name.** It names the units and the volumes, so renaming it later means a migration.
- **The address.** Your passkey is bound to its hostname, so decide it before you enrol. Changing it
later fails startup.
- **The model.** An Ollama API key, an Ollama server you run yourself, or any OpenAI-compatible
endpoint. This one is not final: the agent lists and switches models at runtime.
The name and the address are the two that do not change afterwards.
## Before you start
```
FOR THIS YOU NEED
· SSH and a shell
· A DNS record, only if you want a public hostname
AND THIS
· Linux with systemd and rootless podman 4.9 or newer
· 2 GB RAM for the stack, plus room for the models you pick
```
Tested minimum CPU, RAM and disk figures are **in development**. What is known: the images, the
embedding weights the stack downloads on first start (the project's own figure is about 1.2 GB for the
default model), and your growing `/workspace` and `/memory`.
**Synced passkeys are refused.** q15 needs a credential it can revoke for one device, so iCloud
Keychain and password-manager passkeys do not enrol. Use a security key with PIN verification, or a
device-bound platform authenticator. The reason is in [Chat](/chat/#enrolling-signing-in-and-revoking).
## What you end up with
Six containers: the four q15 images (`q15-web`, `q15-agent`, `q15-exec`, `q15-proxy`), a vector
database for search, and a model server. The four q15 images always carry the same release tag,
because the agent and the browser client share a bridge protocol.
Today's stack serves embeddings from a local TEI container and talks to a model provider you
configure. Shipping Ollama inside the stack, so that a local model is the default, is filed work.
The [reference](/reference/) describes the stack as it is.
Your data is files. `/workspace` is the project tree, `/memory` is the transcript and what the agent
has learned from it, and both survive updates and reboots. See [Backups](/backups/) for what to copy.
Only the browser client binds a port, and only on `127.0.0.1:8080`. Everything else is reachable only
inside the stack, and the credentials never reach the agent: every outbound request goes through
`q15-proxy`, which holds them.
A Telegram bot is optional and takes two values, a bot token and your numeric user ID. See
[Telegram](/telegram/).
## Serving it
The browser client binds loopback, so reach it over an SSH forward or put a TLS terminator in front.
The hostname you give `init` is the public one: your passkey is bound to it. A worked reverse-proxy
example is **in development**; the [reference](/reference/) lists what the stack expects.
## Keeping it running
```bash title="once the binary ships"
q15 status
q15 logs -f q15-agent
q15 doctor # what runs, which credentials resolve, what the agent can reach
q15 secret set # add or replace a credential
```
**Not built yet.** Those four are the goal, not the current state; today the same four things come from
the Compose commands in the [reference](/reference/#run-it-today-end-to-end).
The units are systemd user units, so the stack comes back after a reboot without you logging in. To
move to another release, install the newer `q15` and run `q15 up` again.
## Where to go next
- [Reference](/reference/) for the stack as it runs today: files, volumes, ports, tags, keys.
- [Updating](/updating/) for releases and what does not roll back.
- [Backups](/backups/) before you have anything to lose.
- [Troubleshooting](/troubleshooting/) when something is not healthy.
- [Security](/security/) before you give the agent a shell.
---
# Why q15
_What q15 claims, what each claim costs you, and what it does not claim at all._
q15 is for one person who wants an agent they keep. Custody is the whole argument: the memory, the
files, the jobs and the credentials stay where you can read them, and the machine keeps running when
a vendor changes its mind.
Everything below is a claim you can test against the code, followed by its limit. None of it is a
promise about how it will feel to use.
## Four properties
**1. Your memory is files you can open.**
The transcript is JSON under `/memory/history/`, the extracted knowledge is markdown under
`/memory/semantic/`, and the project tree is `/workspace`. Copy them, grep them, edit them, move them
to another machine.
_The limit:_ they are plaintext on your disk. Encrypt the disk, or encrypt the backups.
**2. The credentials the agent reaches with are not in the agent.**
The tokens for the hosts the agent calls through the shell are injected by `q15-proxy` at the network
layer, for the hosts your policy names. The agent's own environment does not carry them, so its prompt
and its transcript have nothing to leak for those hosts.
_The limit:_ two of them. The agent **does** hold the key for the model provider it calls, because it
is the thing that calls the provider, and a hosted provider sees the prompt either way. And it can
still **use** an injected credential by asking for a host the policy matches. Nothing to steal, plenty
to abuse, so keep the policy narrow. See
[Credentials](/credentials/).
**3. Chat bodies are sealed between the browser and the agent.**
Every frame that carries content is encrypted per connection, with keys the browser and the agent
derive and `q15-web` never sees. A carrier that copies your traffic gets ciphertext and frame sizes.
_The limit:_ this is not end-to-end encryption from the edge. Whatever serves the browser's
JavaScript, or replaces the key exchange, is outside the guarantee. A TLS terminator still sees who
connects, when, and how much.
**4. Policy is enforced in code, not in the prompt.**
Egress routing, file roots, the schedule tool's limits and the web tier's authorization are properties
of the deployment. A prompt that asks to skip them is asking a process that does not have the
option.
_The limit:_ the model is not the only untrusted input. Anything the agent reads is untrusted, and the
agent acts on it with the permissions the deployment gave it.
## What q15 does not claim
- **Not multi-user.** One owner, one agent, no roles, no isolation between people.
- **Not a sandbox against a hostile model.** Commands run in a container, so they are off your host,
but they can read every file under `/workspace`, `/memory` and `/skills`. That is your data.
- **Not private from a hosted provider.** Anything you send to a hosted model is in that provider's
hands.
- **Not private from Telegram**, if you enable that channel. Those messages travel through Telegram's
servers in the clear.
- **Not audited.** No third party has reviewed this code. It is one person's project, and the security
page describes what the design defends, not a certification.
- **No push notifications.** A closed browser tab hears nothing.
- **No recovery password.** Host access to the credential store is the only recovery path.
- **No management console.** The shipped web tier authorizes one scope: chat.
- **Not finished, and not installable by a stranger yet.** The installer is filed as work and not
shipped; the documentation says so wherever it matters.
## Who it is not for
Teams. Anyone who wants an account on someone else's server. Anyone who needs prompts kept away from
model providers without running local models. Anyone who wants install-and-forget.
If your case is one of those, this is the wrong tool, and it is better to find out here than on day
three.
## Where to go next
- [Security](/security/) for the threat model, and the class of attack the architecture is built
around.
- [Credentials](/credentials/) for where secrets actually live and how they are injected.
- [Architecture](/architecture/) for the four services and the storage contract.
- [Install q15](/install/) when you have decided.
---
# Start here by intent
_Pick the line closest to what you want and read three pages instead of nineteen, instead of starting at the top of a table of contents._
You do not have to read this site in order. Find the line closest to what you want, read the pages it
names, and come back when you want the next thing.
## I want to talk to my agent from my phone
1. [Install q15](/install/): the four commands the binary will bring, and the Compose path that
runs today.
2. [Chat](/chat/) for the browser app: passkey enrolment, sessions, attachments. On your phone it
installs to the home screen like any other web app, and it can receive a photo or a PDF straight
from Android's share sheet.
3. [Telegram](/telegram/) if you would rather answer there. It is a second client for the same
agent, not a second agent.
## I want it to work while I am not looking
1. [Jobs](/jobs/) for scheduled work: what a job is, what it costs, and how to see whether it ran.
2. [Tools](/tools/) for what it can actually do with the rest of its turn.
3. [Chat](/chat/) or [Telegram](/telegram/) for where a job's report arrives.
## I want to run a local model
1. [Models](/models/) for the provider block, the roster, and how the agent picks what answers.
2. [Reference](/reference/) for the exact configuration keys.
3. Give it room: the stack itself needs about 2 GB of RAM, and a local model needs whatever that model
needs on top.
The shipped Compose configuration points at a hosted Ollama endpoint. Running against your own
Ollama server is one provider block, and shipping Ollama inside the stack is filed work.
## I want my memory in files I own
1. [Memory](/memory/) for the layers, the transcript, and the background jobs that maintain them.
2. [Workspace](/workspace/) for the project tree, media store and what survives a rebuild.
3. [Backups](/backups/) for what to copy, in what order, and what not to restore.
## I want to give it a shell without giving it my credentials
1. [Security](/security/) for prompt injection, the class of attack the architecture is built around.
2. [Credentials](/credentials/) for the proxy, the injection rules, and the two credential stores.
3. [Architecture](/architecture/) for the four services and where the boundaries actually are.
4. [Why q15](/why-q15/) for what none of this protects.
## I want to know what it can do
1. [Tools](/tools/) for the surface: files, commands, search, media, subagents.
2. [Skills](/skills/) for teaching it procedures it can reuse.
3. [Jobs](/jobs/) for the work it does with nobody watching.
## I want to keep it running
- [Updating](/updating/) for the releases, the tag scheme and what does not roll back.
- [Backups](/backups/) before you need them.
- [Troubleshooting](/troubleshooting/) when something is not healthy.
- [Reference](/reference/) for every key, path and command in one place.
## Every page
**Start here** — [What is q15?](/) · [Install q15](/install/) · [Why q15](/why-q15/) · this page
**Using q15** — [Chat](/chat/) · [Telegram](/telegram/) · [Models](/models/) · [Memory](/memory/) ·
[Workspace](/workspace/) · [Tools](/tools/) · [Skills](/skills/) · [Jobs](/jobs/)
**How it works** — [Architecture](/architecture/) · [Security](/security/) ·
[Credentials](/credentials/)
**Operating it** — [Updating](/updating/) · [Backups](/backups/) ·
[Troubleshooting](/troubleshooting/) · [Reference](/reference/)
Nothing in the sidebar is hidden from this list, and nothing here is missing from the sidebar. If a
page you were sent to is not one of these, it has been folded into one that is.
---
# Chat
_Passkey sign-in, sealed chat, history and attachments, and the exact limits of what the browser client protects._
Chat is the browser client, served by the `q15-web` container. It authenticates exactly one owner with
WebAuthn, keeps the passkeys and sessions in its own store rather than in the agent's memory, and keeps
no transcript of its own: what you read in the browser is the agent's transcript, read back to you.
## What it serves
Chat is the app you look at: streaming replies that arrive as they are written, the tool activity
behind them, paged history, and attachments in both directions. The route list is in the
[reference](/reference/#routes).
## On your phone
The app is a web app, so install it from the browser's menu and it behaves like one you installed
yourself: full screen, no browser chrome, and a passkey tap instead of a login.
On Android it also registers as a **share target**. A scan from a document scanner, or a photo from
the gallery, can be shared straight into q15 and arrives in the composer's attachment tray, already
sealed and ready to send. Chrome is the browser that supports it today, and the app has to be
installed, not merely open in a tab.
## Enrolling, signing in, and revoking
Enrollment is a host operation, including for the first device. The commands use `podman compose`;
on Docker they are the same with `docker compose`.
```bash
podman compose exec q15-web /usr/local/bin/q15-web auth enroll 'Laptop'
podman compose exec q15-web /usr/local/bin/q15-web auth list
podman compose exec q15-web /usr/local/bin/q15-web auth revoke DEVICE_ID
```
The `enroll` command prints an enrollment blob and waits. You open the locked page in a browser, paste
the blob into its enrollment panel, create the device credential, and paste the one-line response
back into the waiting host command within five minutes. Only the host command submits the
registration, over a private Unix socket, so an origin that can run a browser ceremony still cannot
enroll itself. Enrollment can be run over SSH from the device you are enrolling.
Two constraints matter:
- Every credential requires user verification — a PIN or a biometric check.
- Synced and backup-eligible credentials are refused, because revoking a key that several devices share
cannot revoke one physical device. A device-bound platform authenticator or a security key with a PIN
is what works.
There is no recovery password. Host access is the recovery path: enroll a spare authenticator in
advance, and if one is lost, enroll a replacement and revoke the lost device.
## Sessions
A session cookie alone grants nothing. Each HTTP request and each socket upgrade must carry a fresh
proof signed by a session key that lives in the browser and is not exportable, so a copied cookie is
not enough to impersonate a device. Reloads reuse the session key without another passkey gesture.
Sessions expire server-side after **30 days**, with no sliding renewal; signing out deletes the session
server-side. Revoking a device deletes its credential and all of its sessions at once, and other
devices are unaffected. The hard limits are 32 devices, 128 sessions and 64 pending login challenges.
A login ceremony expires after two minutes and an enrollment ceremony after five.
The shipped tier authorizes the `chat` scope. The `console` scope is refused. A management console
would need its own fresh, time-boxed authorization before that scope could be added.
## Content sealing
Everything that carries content is sealed between your browser and the agent: chat bodies, reasoning,
tool arguments and results, live deltas and replayed history. A passive carrier gets ciphertext and
frame sizes, and `q15-web` itself holds no key and parses nothing. The algorithms, the key derivation
and the one cap that applies to attachments are in the [reference](/reference/#content-sealing).
## State and volumes
The web tier keeps exactly one private directory, `q15_web_state` mounted at `/var/lib/q15-web`: the
0600 credential store and a private admin socket. It holds no transcript. Keep it across updates and
rebuilds, because deleting it means enrolling every device again, and read the restore warning in
[Backups](/backups/) before reusing an old copy. The bridge socket it dials lives on a second volume,
mounted in the agent and the web tier only. Both are in the
[volume table](/reference/#volumes).
## What it does not protect
These are the shipped tier's own stated limits, and they are worth reading before you point a domain
at it:
- **A TLS terminator sees cookies and routing metadata.** Sealing hides content from a passive carrier.
It does not hide who is talking, or when, or how much.
- **It is not end-to-end encryption from the edge.** An origin or relay that serves modified
JavaScript, or that replaces the key exchange, is outside the guarantee. Authenticating the code your
browser runs would need a separate mechanism.
- **Host administrators and the web process itself are trusted.** Access to the admin socket or a
writable state directory can enroll an identity.
- **A compromised browser, operating system or authenticator is out of scope.**
- **It bounds work and state, not availability.** The authentication budget limits how much work a
client can cause; it is not a promise that the service stays up under a distributed attack.
See [Security](/security/) for how this fits into the rest of the architecture.
---
# Telegram
_A second client for the same agent — a bot token, your numeric user ID, and one shared conversation._
q15 can answer in Telegram. It is the same agent you reach in the browser: the same transcript, the
same memory, the same tools. Connecting it costs you two values and about ten minutes.
Be clear about the trade before you start: **Telegram sees your messages.** The browser channel seals
content between your browser and the agent; the Telegram channel does not, because the messages pass
through Telegram's servers by definition. Use it for the things you would type into any chat app.
Anything you want sealed stays in [the browser client](/chat/).
## What you need
- **A bot token** from Telegram's [@BotFather](https://t.me/BotFather). Send it `/newbot`, answer two
questions, and it prints a token.
- **Your numeric user ID.** Telegram identifies people by number, not by username. Any of the
well-known ID bots reports yours when you message it.
## Set it up
1. **Put the token in a file.**
The Compose stack mounts one file per secret and passes its path to the agent. Copy the tracked
template and fill it in:
```bash title="in your deployment directory"
cp secrets/q15_telegram_token.example secrets/q15_telegram_token
chmod 600 secrets/q15_telegram_token
```
It reaches the agent as `Q15_TELEGRAM_TOKEN_FILE`, pointing at
`/run/q15-secrets/q15_telegram_token` inside the container.
You should now have a second file beside the template. `ls -l secrets/` shows it mode 600, and
`secrets/q15_telegram_token.example` untouched.
2. **Put your user ID in a second file.**
```bash
cp secrets/q15_telegram_allowed_user_ids.example secrets/q15_telegram_allowed_user_ids
```
One numeric ID per line. This list is the only thing standing between your agent and anyone who
finds the bot, so it is not optional: q15 refuses to build a Telegram channel with an empty
allow-list rather than reading an empty list as "everyone".
You should now have both files in `secrets/`, each containing one line and nothing else. A file with
a trailing blank line is fine; a file with two IDs is not, unless both people should be able to talk
to your agent.
3. **Point the config at both.**
```yaml title="agent-config.yaml"
agent:
telegram:
token_env: Q15_TELEGRAM_TOKEN
allowed_user_ids_env: Q15_TELEGRAM_ALLOWED_USER_IDS
```
The stack sets the matching `*_FILE` variables, so the values themselves never appear in the YAML.
You should now see a third file under `secrets/`, and `grep -n TELEGRAM deploy/compose/docker-compose.image-first.yml`
shows the two `*_FILE` variables the agent will read.
4. **Restart the agent and send the bot a message.**
```bash
podman compose --env-file deploy/compose/release.env \
-f deploy/compose/docker-compose.image-first.yml up -d --force-recreate q15-agent
```
You should now see the agent's reply in the chat, and the same exchange in your browser client's
history a moment later.
If nothing arrives: the user ID is wrong, or the token file is still the empty template. Both are
logged on stdout by the agent, and the second is listed in
[Troubleshooting](/troubleshooting/).
Incoming messages are checked against the sender's user ID. If you add the bot to a group, a
message from an allowed person there is handled like any other and the answer lands in the group
where everyone can read it. Live reply drafts are private chats only.
## What it looks like in use
In a private chat, a long answer grows in place while it is being written. Telegram's native thinking
indicator shows while the model works, reasoning appears as a short excerpt when the model streams
it, and the finished answer replaces it. Updates reuse one message, are coalesced to at most one per
second, and refresh every 20 seconds during long waits. Telegram's Stop button cancels the run; the
unfinished draft is discarded.
How much of the machinery you see is your choice, per chat:
| Command | What the chat shows |
| -------------------- | ----------------------------------------------------------------------------------------------------------------------------- |
| `/progress quiet` | Thinking, partial replies and routine tool activity stay hidden. Only the long-wait notice appears. |
| `/progress progress` | The current action, with a short command, file or search preview (up to 320 characters, five lines). **This is the default.** |
| `/progress verbose` | The same action, with a longer preview (up to 640 characters, ten lines) and longer reasoning excerpts. |
Tool **output** is never dumped into a progress message; only what the agent asked for. Commands keep
their line breaks and shell syntax inside a code block.
## Attachments
Both directions work.
- **To the agent:** send a photo, a document, a voice note, a video, a sticker or a GIF. It is stored
in the agent's media store under a content hash and attached to your message, so the agent sees the
file, not a link.
- **From the agent:** when the agent attaches something, it arrives as the matching Telegram kind:
photo, audio, document, video, animation, sticker or video note.
A failed attachment does not lose your caption: the message still arrives, with a line saying what
could not be ingested.
## Limits
- **No sealing.** Everything on this channel passes through Telegram.
- **The token is a bearer credential.** Anyone holding it can act as your bot. It belongs in a secret
file, not in the YAML, and it belongs in your [backup](/backups/) only if that backup is encrypted.
- **Long polling, so no inbound port.** q15 asks Telegram for updates; it never accepts a connection.
There is nothing to expose, and it works behind NAT.
- **One bot per agent.** Two agents need two bots, and they keep two separate conversations.
- **The two clients share one transcript.** An exchange from Telegram shows up in the browser
history, because the agent keeps one conversation and both clients read it. What differs is
delivery: an answer written in Telegram is sent to Telegram, and is not repeated in the browser.
## Where to go next
- [Chat](/chat/) for the browser client that seals its payloads.
- [Jobs](/jobs/) for scheduled work; a job created here reports back here.
- [Credentials](/credentials/) for how the bot token is held.
---
# Models
_How q15 chooses which model answers a turn, how to change it at runtime, and how to point it at a model you run yourself._
Every turn, q15 decides which model to call. There is no model list in the configuration: each
provider's own roster is the source of truth, and the choice is made per request.
## The roster
The agent queries every configured provider for its model list at startup, enriches what it finds
with capability metadata from [models.dev](https://models.dev), and refreshes it periodically. The
result is the **roster**.
- A provider that is unreachable at startup contributes nothing, and the agent starts with the rest.
- **An empty roster is a startup failure.** If `q15-agent` exits immediately, check the provider list
and the matching secret file first.
See [Agent config](/reference/#configuration) for the provider block that produces the roster.
## Choosing a model
Selection is two stages: filter, then rank.
```mermaid
flowchart TD
Req[Turn arrives] --> Infer[Derive the request's requirements]
Infer --> Filter[Filter the roster by capability]
Filter --> Any{Any candidate left?}
Any -->|No| Fail[The turn fails]
Any -->|Yes| Rank[Rank deterministically]
Rank --> Call[Call the first model]
Call --> Ok{Call succeeded?}
Ok -->|Yes| Done[Answer]
Ok -->|No| Next{Another candidate?}
Next -->|Yes| Call
Next -->|No| Fail
```
**Requirements.** Text is a hard requirement. Media is not: an image or audio part is rendered down
to a text hint for a model that cannot take it inline, rather than excluding that model from
selection. Tool calling is required only when no candidate in the current set could fall back to a
tool-free call — otherwise the agent may omit tools and call a model that does not support them.
**Ranking.** Candidates that survive the capability filter are ranked by cheapest known cost, then
best known benchmark, then the order the provider returned them. The ranking is deterministic: the
same roster and the same request produce the same first choice. When the top choice is an ambiguous
tie, the agent logs the reason rather than guessing silently.
The current model is tried first on every path, then the rest of the ranked candidates. If a call
fails, the next candidate is tried; a turn fails only when no candidate is left.
## The current model is runtime state
The selection is stored on the agent, not in the config file. Two consequences:
- A pin that no longer exists in the roster does not stop the agent from starting — the seed is not
validated at startup.
- Every change you make at runtime **is** validated against the live roster, so you cannot switch to
a model that is not currently reachable.
The current pair persists across restarts.
You change it from chat:
| Tool | Effect |
| ------------------------ | ------------------------------------------------------ |
| `list_providers` | Which providers are reachable, and their model counts. |
| `list_models` | The roster, with capabilities and context windows. |
| `switch_model` | The model that answers interactive turns. |
| `switch_cognition_model` | The model for one background cognition job. |
`switch_cognition_model` takes a job type — `verification_review`, `working_memory.consolidate` or
`semantic_memory.extract`. A cognition job uses its own override when one is set, and otherwise
inherits the interactive model. See [Memory](/memory/) for what the jobs do.
## Running a local model
Local is the direction of the project, and one provider block is what stands between today and it:
```yaml
providers:
- name: local
type: ollama
base_url: http://ollama:11434
```
**The `base_url` is the part to get right.** It has to resolve from inside the container, not from
your shell, so `localhost` is wrong unless Ollama runs in the same stack. A tested value for a
host-run Ollama is **in development**; the address that works is the one your container engine
publishes for the host.
Two things to know before you commit to it:
- **The model has to fit the machine.** The stack itself needs about 2 GB of RAM; a local model needs
whatever that model needs, and running a large one on a small box is a slow experience rather than a
broken one.
- **Embeddings are separate.** The search backend is a second provider block, and on the shipped
Compose stack it is the local `q15-tei` container. See [Memory](/memory/#searching-your-own-files).
## Discovery
The catalog is built in memory at startup and refreshed periodically; it is not a file you edit and
not a table you maintain. Provider types determine how a provider is reached and how it
authenticates:
| Type | Authenticates with | Notes |
| ------------------- | ---------------------------------- | -------------------------------------------------------------- |
| `ollama` | nothing, or `key_env` for cloud | `base_url` defaults to the local Ollama endpoint when omitted. |
| `openai-compatible` | `key_env` | Any endpoint that speaks the OpenAI API, for example Moonshot. |
| `openai-codex` | `auth.json` produced by `q15-auth` | The OpenAI OAuth flow. See [Credentials](/credentials/). |
---
# Memory
_What q15 remembers, where each layer is stored, how it reaches the prompt, and how background jobs keep it current._
## The sessionless agent
q15 maintains one continuous conversation. There are no sessions to start, no context windows to manage, no "new chat" button. You send a message on Monday, another on Thursday, another next month — and the agent picks up where you left off because its memory persists across the gap.
What follows is what it remembers, where each layer lives, how it reaches the prompt, and how the
background jobs keep it current.
## The memory layers
| Layer | Path | In the prompt? | Kept by |
| ------------- | -------------------- | -------------------- | ------------------- |
| **Core** | `/memory/core/` | every turn | you |
| **Working** | `/memory/working/` | every turn | consolidation job |
| **Semantic** | `/memory/semantic/` | on demand | extraction job |
| **History** | `/memory/history/` | after the checkpoint | appended every turn |
| **Cognition** | `/memory/cognition/` | never | the job controller |
| **Notes** | `/memory/notes/` | on demand | the agent |
Each layer is a directory you can open. What is in them, and what maintains each one, is below.
### Core memory
Files under `/memory/core/` define who the agent is. These are always in the prompt:
- `AGENT.md` — role and behavioral protocol
- `USER.md` — user identity, preferences, communication norms
- `SOUL.md` — voice, teaching style, working principles
Core memory is operator-maintained. Cognition jobs never write to it. Your agent's identity is not something a background model call should rewrite.
### Working memory
`/memory/working/WORKING_MEMORY.md` holds bounded active state: current priorities, active tasks, open threads, recent progress, pending checks. It is the primary mechanism for sessionless continuity.
The consolidation job keeps it compact. Resolved tasks are removed. Abandoned threads are dropped. New constraints are captured. It is a scratch pad for the current moment, not a history log.
Only `WORKING_MEMORY.md` is auto-injected. Other files under `/memory/working/` are not prompt-visible.
### Semantic memory
Three canonical files hold durable extracted knowledge:
- `facts.md` — confirmed facts and grounded inferences
- `preferences.md` — user preferences and collaboration preferences
- `projects.md` — active projects and durable project knowledge
Semantic memory is not auto-injected. The agent fetches it on demand when durable context is relevant. The extraction job promotes only explicit, durable statements — not single-turn tasks or speculative guesses. If the verification review flags an entry as stale, the next extraction cycle removes or rewrites it.
### History
Completed turns are stored as JSON files under `/memory/history/turns/YYYY/MM/DD/`. The full transcript persists forever. Nothing is deleted.
Replay is checkpoint-aware. When consolidation succeeds, it advances a checkpoint. The next prompt replays only turns after that checkpoint. The semantic extraction job has its own independent checkpoint for processing a different (typically larger) transcript slice.
### Cognition
`/memory/cognition/` is system-owned state. The agent does not read or write here during interactive turns. It contains job trigger state, append-only run records with full provenance (model used, input/output sequence, success/failure), and the verification review artifact.
### Notes
`/memory/notes/` contains the auxiliary zettelkasten notebook: inbox, atomic notes, and structure maps. These are durable knowledge infrastructure that the agent can search and reference on demand.
## Searching your own files
When the embeddings tool is configured, the agent can index your own content and search it. Qdrant
stores the vectors. The **default backend is the local `q15-tei` container**, which the agent reaches
through the OpenAI-compatible embeddings provider and which needs no API key; a hosted backend is two
lines away.
1. **Declare sources.** A source names a collection, a source type and a path. The registry is a file:
`/workspace/.q15/embed/sources.json`.
2. **Sync.** `embed_sync` indexes each source, storing a dense vector per chunk plus a BM25 sparse
vector.
3. **Search.** `embed_search` runs hybrid search by default, combining semantic recall with lexical
matching.
| Collection | Purpose |
| -------------- | -------------------------------------------------------- |
| `library` | Books, articles and documents from your personal library |
| `zettelkasten` | Atomic knowledge notes |
| `semantic` | Extracted semantic memory content |
| `core` | Core identity and reference material |
The collection is a property of the source; the scanner behaviour comes from the source **type**,
which is `markdown_tree`, `markdown_file` or `chunked_markdown_tree` for pre-chunked content.
```yaml
agent:
tools:
embeddings:
qdrant_url_env: Q15_QDRANT_URL
provider: openai
base_url_env: Q15_EMBEDDINGS_BASE_URL
model: qwen3-embedding-0.6b
dimensions: 1024
batch_size: 128
```
**Changing the backend is a migration, not a config edit.** A collection's dimension is fixed when it
is created: change the model or the dimensions and the vector-version stamp changes, every stored
record is marked dirty, and the next sync recreates the collections. On a local backend that costs
time; on a hosted one it costs money. The runbook is in the repository's
[`deploy/compose/README.md`](https://github.com/q15co/q15/blob/main/deploy/compose/README.md).
## The prompt, assembled
Here is what goes into the model's context on every turn, ordered from most stable to least stable:
```mermaid
flowchart TD
subgraph prompt["System messages (cached prefix)"]
SP["System prompt code-owned execution policy"]
CM["Core memory AGENT.md · USER.md · SOUL.md"]
SC["Skill catalog available skills + descriptions"]
WM["Working memory WORKING_MEMORY.md"]
end
subgraph ctx["Conversation context"]
RT["Recent transcript turns after consolidation checkpoint (bounded replay window)"]
UM["Current user message + temporal metadata"]
end
SP --> CM --> SC --> WM --> RT --> UM
style SP fill:#89b4fa,stroke:#2563eb,color:#1e1e2e
style CM fill:#cba6f7,stroke:#8839ef,color:#1e1e2e
style SC fill:#fab387,stroke:#d97706,color:#1e1e2e
style WM fill:#a6e3a1,stroke:#16a34a,color:#1e1e2e
style RT fill:#313244,stroke:#89b4fa,color:#cdd6f4
style UM fill:#f38ba8,stroke:#dc2626,color:#1e1e2e
```
1. **System prompt** — Code-owned execution policy: autonomy rules, tool persistence, verification loop, output contract. Compiled into the binary. Changes only on upgrades. It also tells the agent where every memory layer lives and which paths are auto-injected versus tool-fetched.
2. **Core memory** — `AGENT.md`, `USER.md`, `SOUL.md`. Agent identity, user profile, voice and principles. Durable files that change rarely.
3. **Skill catalog** — Descriptions of installed skills. Changes when skills are added or removed.
4. **Working memory** — `WORKING_MEMORY.md`. Bounded active state maintained by the consolidation job.
5. **Recent transcript** — Turns after the consolidation checkpoint. The unconsolidated tail, bounded by the replay window.
6. **User message** — The current message with temporal metadata (local time, day of week, gap since last message).
The ordering lets providers cache the prefix. The system prompt and core memory rarely change, so the cached prefix can be reused even when working memory updates.
## How memory is maintained: cognition jobs
q15 does not compress your history. It keeps the transcript and maintains **structured memory artifacts** through background model calls called **cognition jobs**. These run on their own schedules and triggers and keep the memory layers current.
They are not the [scheduled jobs](/jobs/) you create in chat: those live in a different store, they do
not count against the job limit, and you cannot list or delete a cognition job. The one lever you have
is the model each of them runs on, which is what `switch_cognition_model` sets below.
The full transcript is always persisted as JSON files under `/memory/history/`. Nothing is deleted. But only the recent unconsolidated turns are replayed into the prompt. Everything before the consolidation checkpoint is represented by the working memory artifact, which was built from those turns by a background job.
```mermaid
flowchart TD
subgraph session["One continuous conversation (no sessions)"]
U1["User message Monday"] --> R1["Agent reply"]
R1 --> P1["Persist turn to /memory/history/"]
P1 --> D1["Dirty state +1 turn"]
D1 -->|"6+ dirty turns"| C1["Working memory consolidation job"]
C1 --> CP1["Advance consolidation checkpoint to turn N"]
CP1 --> WR["Working memory artifact updated with active state"]
WR -->|"Next reply"| N1["Replay only turns after checkpoint"]
U2["User message Thursday"] --> R2["Agent reply"]
R2 --> N1
N1 --> D2["Dirty state +1 turn"]
D2 -->|"12+ dirty turns"| C2["Semantic memory extraction job"]
C2 --> SM["facts.md, preferences.md, projects.md updated"]
SM -->|"Available via tools"| N2["Agent can fetch on demand"]
end
style C1 fill:#a6e3a1,stroke:#16a34a,color:#1e1e2e
style C2 fill:#cba6f7,stroke:#8839ef,color:#1e1e2e
style CP1 fill:#89b4fa,stroke:#2563eb,color:#1e1e2e
style WR fill:#313244,stroke:#89b4fa,color:#cdd6f4
style SM fill:#313244,stroke:#cba6f7,color:#cdd6f4
```
The key difference: **compaction replaces history with a summary. Cognition jobs extract structured state and advance a checkpoint.** The history stays on disk. The prompt carries the structured artifact plus only the recent tail that has not been consolidated yet.
### The three cognition jobs
Three background jobs maintain memory. Each runs on its own triggers and produces a specific artifact. Jobs can use a different model than the interactive loop — assign a stronger reasoning model to verification, a large-context model to extraction, and a lighter model to consolidation.
**Verification review** (`verification_review`) audits the other memory layers for accuracy. It reads core memory, working memory, and recent transcript turns, then produces a review artifact flagging stale, unsupported, or contradicted entries. It runs twice daily or after 8+ unconsolidated turns. Its output is consumed only by the other two jobs — never by the interactive agent directly.
**Working memory consolidation** (`working_memory.consolidate`) maintains `WORKING_MEMORY.md`, the bounded active-state artifact auto-injected into every prompt. It reads the current working memory, the verification review, and the last 16 transcript turns, then updates the file to reflect current priorities, active tasks, open threads, and recent progress. It runs daily or after 6+ unconsolidated turns. On success, it advances the consolidation checkpoint — the next reply replays only turns after that boundary.
**Semantic memory extraction** (`semantic_memory.extract`) maintains three durable knowledge files: `facts.md`, `preferences.md`, and `projects.md`. It reads all three files, the working memory, the verification review, and transcript turns since its last checkpoint, then promotes durable facts, preferences, and project knowledge. It runs daily or after 12+ unconsolidated turns. Semantic memory is not auto-injected — the agent fetches it on demand with `read_file`.
### How artifacts reach the agent
Cognition jobs produce artifacts that reach the interactive agent through four paths:
| Path | Artifacts | Mechanism |
| --------------------- | ---------------------------------------------- | --------------------------------------------------------------------------- |
| **Auto-injected** | Core memory, working memory | Injected into every prompt as system sections |
| **Tool-fetched** | Semantic memory, zettelkasten notes | Agent calls `read_file` on demand |
| **Checkpoint replay** | Recent transcript turns | Replayed verbatim for turns after the consolidation checkpoint |
| **Vector search** | All memory layers (when embeddings configured) | Agent calls `embed_search` for semantic retrieval across Qdrant collections |
The consolidation checkpoint marks the replay boundary. Everything before it is represented by the working memory artifact. Everything after it is replayed verbatim:
```mermaid
flowchart TD
subgraph disk["Full transcript on disk (never deleted)"]
Old["Turns 1 to 47 (before checkpoint)"]
New["Turns 48 to 50 (after checkpoint)"]
end
CP["Consolidation checkpoint at turn 47"]
CP -.->|"marks the boundary"| New
Old -.->|"represented by"| WM["WORKING_MEMORY.md"]
New -->|"replayed into"| Prompt["Goes into prompt"]
WM -->|"injected into"| Prompt
style CP fill:#89b4fa,stroke:#2563eb,color:#1e1e2e
style Old fill:#313244,stroke:#585b70,color:#cdd6f4
style New fill:#a6e3a1,stroke:#16a34a,color:#1e1e2e
style WM fill:#a6e3a1,stroke:#16a34a,color:#1e1e2e
style Prompt fill:#313244,stroke:#89b4fa,color:#cdd6f4
```
When embeddings are configured, the agent can also search memory semantically via Qdrant. The `semantic`, `core`, `library`, and `zettelkasten` collections are indexed with hybrid dense+sparse vectors. The agent calls `embed_search` with a natural-language query and receives ranked results with file paths and relevance scores. Embeddings are optional — when not configured, the agent relies on the other three paths. See [Searching your own files](#searching-your-own-files) above for the configuration of both.
### The correction cycle
The three jobs form a self-correcting loop. The verification job audits state; the consolidation and extraction jobs apply corrections. This cycle runs continuously in the background, keeping memory artifacts accurate without user intervention.
```mermaid
flowchart TD
VR["Verification review audits state"] -->|"flags stale unsupported contradicted entries"| WM["Working memory consolidation"]
VR -->|"flags stale unsupported contradicted entries"| SM["Semantic memory extraction"]
WM -->|"removes or rewrites flagged entries"| WMF["WORKING_MEMORY.md corrected"]
SM -->|"removes, rewrites, or downgrades"| SMF["facts.md preferences.md projects.md corrected"]
WM -->|"updates active state"| SM
WMF -->|"next turn"| Prompt["Prompt"]
SMF -.->|"on demand"| Prompt
style VR fill:#f9e2af,stroke:#df8e1d,color:#1e1e2e
style WM fill:#a6e3a1,stroke:#16a34a,color:#1e1e2e
style SM fill:#cba6f7,stroke:#8839ef,color:#1e1e2e
style WMF fill:#313244,stroke:#a6e3a1,color:#cdd6f4
style SMF fill:#313244,stroke:#cba6f7,color:#cdd6f4
style Prompt fill:#313244,stroke:#89b4fa,color:#cdd6f4
```
The ordering emerges from trigger thresholds and cron schedules, not from hardcoded sequencing:
1. **Verification review** runs first (8+ dirty turns, or twice daily). It produces a fresh review artifact.
2. **Working memory consolidation** runs next (6+ dirty turns, or daily). It reads the verification review and applies corrections.
3. **Semantic memory extraction** runs last (12+ dirty turns, or daily). It reads both the verification review and the freshly consolidated working memory, then applies corrections to the semantic files.
A burst of conversation can trigger all three jobs in sequence within minutes.
## Triggers, run records and models
Each job has two triggers: a **state** trigger (enough unconsolidated turns have accumulated) and a
**schedule** trigger (a fixed UTC time). The controller evaluates both after every turn and launches
a job that is due. A job that is already running is not started a second time.
| Job | Identifier | Schedule trigger | State trigger |
| ---------------------------- | ---------------------------- | ------------------- | ------------------------ |
| Verification review | `verification_review` | 03:00 and 15:00 UTC | 8+ unconsolidated turns |
| Working-memory consolidation | `working_memory.consolidate` | 04:00 UTC | 6+ unconsolidated turns |
| Semantic-memory extraction | `semantic_memory.extract` | 05:00 UTC | 12+ unconsolidated turns |
Every run is recorded under `/memory/cognition/runs/YYYY/MM/DD/` as a JSON file carrying the job
type, the trigger cause, the model and provider used, the start and end times, the outcome, and
whether the target files changed. The records are an audit trail: you can see when each job ran,
what triggered it, and what it did.
```
/memory/cognition/
├── state/ job state, checkpoints, and the latest verification review
└── runs/YYYY/MM/DD/ one JSON record per run
```
**Models.** A job runs on its own model when one is set, and otherwise inherits the interactive
model. Set a per-job model in chat with `switch_cognition_model`, naming the job type above.
Verification benefits from a careful model; consolidation is mechanical. See
[Model selection](/models/).
**Files.** Each job may write only the artifacts it owns: consolidation writes the working-memory
file, extraction writes the three semantic files, and verification writes its review. None of them
can write another job's artifact.
## Why not summarize it instead
The obvious alternative is to have a model compress older turns into a summary and keep only the
recent ones verbatim. q15 does not do that: a summary is lossy, unstructured and opaque, and the
transcript it replaces is worth keeping. The reasoning, and the comparison, are in
[Architecture](/architecture/#why-memory-is-not-a-summarized-transcript).
## What this means in practice
When you talk to q15, you are having one conversation with an agent that remembers. The working memory artifact carries your current thread forward. Semantic memory carries durable facts about you and your projects. The full transcript is on disk. Background jobs maintain all of this while you are not looking — not by compressing your words into a paragraph, but by extracting structured state that the next turn can use.
The verification job audits that state. The consolidation job keeps the active-state artifact current. The extraction job promotes durable knowledge. The correction cycle means memory does not just persist — it gets better over time, as stale entries are flagged and corrected without any user intervention.
The surface the agent acts with, and how embeddings fit into it.
The four services, and how memory fits into the turn flow.
The durable project tree that pairs with memory for long-running work.
---
# Workspace
_The durable project tree that survives restarts and upgrades._
`/workspace` is not ephemeral scratch space. It is the durable working state for a q15 stack.
## What lives here
- Your project files and working directory
- Embedding source registry at `/workspace/.q15/embed/sources.json`
- Embedding sync state at `/workspace/.q15/embed/state.jsonl`
- Any files the agent creates, edits, or downloads during its work
## Persistence model
| Deployment | Expectation |
| --------------------- | ----------------------------------- |
| **Compose** | Required named volume or bind mount |
| **Local development** | May start empty, populated later |
A newly created stack may attach an empty persistent volume. That empty initial state is valid. Operators may populate it later through normal agent work or manual setup.
The key expectation is that `/workspace` **remains attached to the same stack over time**. Restarts, redeployments, and upgrades should preserve the existing volume. Do not downgrade it to an ephemeral mount (`emptyDir`, anonymous volume, or temporary bind location).
## Files the agent sends and receives
Attachments and generated media live in the media store at `/media`, not in the project tree. When you
send a photo in the browser or in Telegram, the agent stores it under a content hash and attaches that
reference to your message; when the agent sends something back, it comes from the same store. The
project tree holds your work, and the media store holds what came through a chat window.
## What does not live here
- Agent identity and memory live under `/memory`
- Skill artifacts live under `/skills`
- Attachments and generated media live under `/media`
- Nix store state lives under `/nix`
- Proxy state lives under `/var/lib/q15/proxy`
These are separate persistent volumes, each with its own ownership and lifecycle. What each one costs
you is worth knowing before you decide what to copy:
- **`/nix`** is the executor's package store. It grows with every package the agent has fetched, and it
is the one volume you can throw away: the packages come back on demand, and the first command after
a restore is simply slower than the ones after it.
- **`/workspace` and `/memory`** are the two that cannot be recreated. They are what
[Backups](/backups/) exists for.
---
# Tools
_The agent's capabilities — file operations, shell execution, web search, embeddings, media, skills, and subagents._
The agent's tools fall into eight categories. Each is a capability the model can invoke during a
turn. Tools are wired by `q15-agent` and executed through the appropriate service boundary.
## Tool overview
| Tool | What it does | Execution boundary |
| -------------- | -------------------------------------------------------- | ---------------------------------------------------- |
| **Files** | Read, write, edit, and patch files | Agent (rooted to `/workspace`, `/memory`, `/skills`) |
| **Exec** | Run shell commands through Nix | `q15-exec` (gRPC) |
| **Web** | Search the web and fetch pages | Agent → `q15-exec` → `q15-proxy` |
| **Embeddings** | Sync and search vector collections | Agent (Qdrant + the configured backend) |
| **Media** | Load images for vision, attach images/audio for delivery | Agent |
| **Skills** | Validate skill directories | Agent |
| **Schedule** | Create and manage jobs that run without you | Agent |
| **Subagent** | Delegate work to isolated sub-agents | Agent (spawns independent model calls) |
## Files
The agent can read, write, edit, and patch UTF-8 text files. All file operations are rooted to `/workspace`, `/memory`, and `/skills`. The agent cannot access paths outside these roots.
- **`read_file`** — Read a file with optional line-based pagination (`offset_lines`, `limit_lines`).
- **`write_file`** — Create or fully replace a file.
- **`edit_file`** — Perform one exact text replacement in an existing file.
- **`apply_patch`** — Apply a multi-file patch using the Codex patch envelope.
## Exec
Shell commands run through `q15-exec` over gRPC. The exec service uses Nix to provide packages on demand — any package from nixpkgs is available without rebuilding the container image.
- Commands specify optional `packages` (nix installables) for dependencies.
- Sessions can be long-running with `wait_seconds` and `keep_stdin_open`.
- All egress from exec routes through `q15-proxy`.
See [Architecture](/architecture/#inside-the-executor) for the full execution model.
## Web
Two web tools give the agent internet access:
- **`web_search`** — Search the web using Brave Search. Returns titles, URLs, and snippets.
- **`web_fetch`** — Fetch a known URL and convert readable HTML to markdown.
Both route through `q15-proxy`. If the proxy has a credential rule for the target host, credentials are injected at the network layer — the agent never sees them.
Web search requires `BRAVE_API_KEY` in the agent config. Omit the `web_search` block to disable it.
## Embeddings
When configured, the agent can build and search Qdrant-backed embedding collections over local content. This powers semantic recall from the library, zettelkasten, and memory.
- **`embed_sources`** — Manage typed embedding sources (add, remove, enable, disable).
- **`embed_sync`** — Index source content into Qdrant.
- **`embed_search`** — Run hybrid (dense + sparse), dense-only, or sparse-only search.
- **`embed_status`** — Report source state and collection health.
- **`embed_job`** — Inspect or cancel a running sync job.
Requires `Q15_QDRANT_URL` and a backend. The shipped stack uses the local `q15-tei` container through
the OpenAI-compatible provider, which needs no key; a hosted backend such as Gemini needs its API
key. See [Embeddings and search](/memory/#searching-your-own-files) for the configuration of both.
## Media
Media tools handle images and audio for vision-capable models and user-facing delivery:
- **`load_image`** — Register an image for the model to inspect with vision on the next turn.
- **`attach_media`** — Deliver a file to the user: an image, audio, a document, a video, a sticker, an
animation, or a video note.
Media files must be under a shared runtime root (`/workspace`, `/memory`, `/skills` or `/media`).
The `/tmp` directory is not accessible for media delivery.
## Schedule
The schedule tool manages jobs that run without you:
- **`schedule_create`** — Add a one-shot or recurring job, on a UTC schedule, with its own turn
budget and tool allow-list.
- **`schedule_list`**, **`schedule_update`**, **`schedule_delete`** — Inspect and manage them.
Scheduled jobs are optional: omit the `schedule` block in the config to disable the tool. A job's
own tool list is decided when the job is created, not by the config. See [Jobs](/jobs/).
## Skills
The skills tool validates skill directories:
- **`validate_skill`** — Parse a skill's `SKILL.md`, check metadata, and report errors and warnings.
Skills themselves are not tools — they are workflow packages the agent reads and follows. The skills tool is the mechanism for checking that a skill is well-formed.
See [Skills](/skills/) for the full skill system.
## Subagent
The agent can delegate work to isolated sub-agents running on configured models:
- **`subagent`** — Start a delegated sub-agent with a specific model, task, tools allowlist, and optional skill injection.
- **`subagent_read`** — Read events from a running sub-agent session.
- **`subagent_write`** — Send a follow-up message to a running sub-agent.
- **`subagent_list`** — List sub-agent sessions.
- **`subagent_kill`** — Cancel a running sub-agent session.
Subagents run independently from the main conversation. They receive sanitized context — only what the parent agent explicitly provides. This keeps the main session context clean and allows using cheaper models for delegated work.
## Optional tools
Web search and embeddings are optional. Omit their config blocks to disable them:
```yaml
agent:
tools:
web_search: # omit to disable
brave_api_key_env: BRAVE_API_KEY
embeddings: # omit to disable
qdrant_url_env: Q15_QDRANT_URL
provider: openai
base_url_env: Q15_EMBEDDINGS_BASE_URL
model: qwen3-embedding-0.6b
dimensions: 1024
batch_size: 128
```
Files, exec, media, skills, schedule and subagent tools are wired by the agent itself; only web
search, embeddings and schedule are gated on config blocks.
---
# Skills
_Reusable workflow packages that extend what the agent can do — markdown-defined, version-controlled, and portable across stacks._
Skills are reusable workflow packages installed under `/skills`. They extend the agent's capabilities without changing its code. A skill is a directory containing instructions, scripts, and references that the agent reads and follows.
## What a skill is
A skill is a directory containing:
- A `SKILL.md` file with YAML frontmatter (name, description) and markdown instructions
- Supporting scripts, references, and assets
Skills are **not code that runs automatically**. They are instructions and context that the agent reads when the skill is relevant to the current task. The agent follows the procedures defined in the skill's markdown, running scripts and referencing documentation as directed.
## How skills work
When a skill is relevant, the agent reads the `SKILL.md` and follows its instructions. This might mean:
- Running a script from the skill directory
- Following a step-by-step procedure defined in markdown
- Referencing documentation bundled with the skill
- Using the skill's declared tools (if any)
Skills are discovered from the filesystem at runtime. The agent scans `/skills` for directories containing `SKILL.md` files and builds a catalog of available skills. There is no separate registry to maintain — the directory structure is the source of truth.
## Built-in vs. shared skills
Skills come from two roots:
| Root | Location | Persistence | Examples |
| ------------ | ------------------------------- | ------------------------ | -------------------------------------------------------- |
| **Built-in** | `/skills/@builtin/` (read-only) | Shipped with the runtime | `skill-creator`, `skill-discovery` |
| **Shared** | `/skills//` (read-write) | Persistent volume | `technical-writing`, `browser-use`, `youtube-transcript` |
Built-in skills are embedded in the agent binary and available on every deployment. They are read-only — the agent cannot modify them.
Shared skills are installed by the operator or by the agent itself during normal work. They persist across restarts and upgrades because `/skills` is a persistent volume.
## Skill catalog
The agent maintains an in-memory catalog of all available skills. Each entry includes:
- **Name** — the directory name (e.g., `technical-writing`)
- **Description** — from the SKILL.md frontmatter
- **Source** — `builtin` or `shared`
- **Path** — filesystem location
The catalog is rebuilt at startup and can be refreshed during a session. The agent uses the catalog to decide which skills are relevant to the current task.
## Skill injection
Skills can be injected into subagents. When you delegate work to a subagent, you can specify which skills it should have access to:
```
subagent(
model: "deepseek-v4-flash",
task: "Write a concept page...",
skills: ["technical-writing"],
tools: ["read_file", "write_file"]
)
```
The subagent receives the skill's `SKILL.md` in its system prompt and can use the skill's declared tools. This lets you compose capabilities: a subagent with the `technical-writing` skill writes documentation, while the parent agent orchestrates.
## Skill validation
The `validate_skill` tool checks that a skill directory is well-formed:
- `SKILL.md` exists and is valid markdown
- YAML frontmatter has required fields (`name`, `description`)
- Name matches the `^[a-z0-9]+(?:-[a-z0-9]+)*$` pattern
- Description is under 1024 characters
- Directory structure follows conventions
Use validation after creating or editing a skill to catch metadata or structure issues before the agent tries to use it.
## Installing skills
Skills can be installed through normal agent work — the agent fetches and places them under `/skills`. They can also be installed manually by the operator.
The directory structure under `/skills` is the source of truth:
```
/skills/
├── @builtin/
│ ├── skill-creator/
│ │ └── SKILL.md
│ └── skill-discovery/
│ └── SKILL.md
├── technical-writing/
│ ├── SKILL.md
│ └── references/
│ ├── style-guide.md
│ ├── documentation-patterns.md
│ ├── anti-patterns.md
│ └── diagram-guide.md
└── browser-use/
├── SKILL.md
└── scripts/
└── chrome-wrapper.sh
```
## Persistence
`/skills` is a persistent volume. Skills survive restarts and upgrades. In local development, skills may be reset, but persistence avoids repeated skill bootstrap in production deployments.
## Relationship to tools
Skills are not tools — they are workflow packages. The `validate_skill` tool is the mechanism for checking that a skill is well-formed, but the skill itself is content the agent reads and follows, not a function it calls.
See [Tools](/tools/) for the full list of agent capabilities.
---
# Jobs
_Scheduled work that runs without you, with its own turn budget, its own tools, and a report back to the chat that created it._
A job is a prompt you leave behind. You ask the agent for it once, in a chat, and from then on it runs
on its own: no conversation, no one watching, and, if the job was created with a notification
requested, a message when it has something to say.
Jobs are how q15 does the work you would otherwise have to remember to ask for. A morning summary of
a folder, a check on a page that changes, a weekly tidy-up of a workspace.
## What a job costs you
Read this before you create a dozen of them.
- **Model calls.** Every run is real turns against a real model, on your provider account. A job with
no tools is cheap; a job that reads twelve files and searches the web is not.
- **Five minutes, then it is killed.** A run is stopped after five minutes of wall time. That cap is
not configurable through the agent config today; the field request is
[filed](https://github.com/q15co/q15/issues/169).
- **A tool allow-list decided at creation.** A job can use only the tools named when it was created
or last updated. That is a security property, not a convenience: a job that should read a file has
no reason to hold your shell.
- **Ownership.** A job belongs to the chat that created it. Create it in Telegram and it reports in
Telegram; it is not visible from the browser client, and the browser client cannot run it.
## Create one
You do not write a file. You ask, in chat:
> Every weekday at 07:00 Berlin time, summarise anything new in `/workspace/inbox` and send me three
> lines. Do not reply if there is nothing new.
The agent calls `schedule_create`, which takes:
| Field | What it means |
| ----------------- | --------------------------------------------------------------------------------------------------------- |
| `name` | The label you will recognise in a list. |
| `prompt` | The task, written for an isolated agent that has no memory of this conversation. |
| `kind` | `oneshot` or `recurring`. |
| `run_at` | An RFC3339 **UTC** timestamp. Required for a one-shot. |
| `cron` | A strict five-field **UTC** cron expression. Required for a recurring job. |
| `notify` | Whether the run reports back. A silent job still writes its record. |
| `model` | Optional provider and model ref to pin the run. Omitted, it uses the model you are on when you create it. |
| `max_turns` | Optional turn cap for a run, bounded by the deployment. |
| `allowed_tools` | The exact tool names the job may use. An empty list means it needs none. |
| `context_profile` | `minimal` (default) or `agent`. See below. |
**Context profiles.** `minimal` is the default and the one you want most of the time: the job gets its
prompt and its tools, and nothing of your conversation. `agent` adds the stable agent context (core
memory and the skill catalog) but still never the interactive conversation. Neither profile changes
which tools the job may use; only `allowed_tools` does.
Local time is your problem, not the scheduler's: both `run_at` and `cron` are UTC. If you think in
Berlin time, subtract the offset, and remember it moves twice a year.
## Limits
| Limit | Default | Where it comes from |
| ----------------- | ------- | -------------------------------------------------------------------- |
| Jobs per agent | 64 | `agent.tools.schedule.max_jobs`, bounded at 1000 |
| Turns per run | 16 | `agent.tools.schedule.max_run_turns`, bounded at 128 |
| Wall time per run | 5 min | Fixed. Not in the YAML. |
| Concurrent runs | 1 | A job that is already running is skipped when its next slot arrives. |
Reaching `max_jobs` fails the create rather than silently dropping the oldest job.
## Watching them
From chat:
- `schedule_list` — the jobs this chat owns, with their next fire time and last status.
- `schedule_runs` — the archived history: exact counts by status, and the newest records, filtered by
job id, name, status or time window. Use it when you want to know whether the thing actually ran.
- `schedule_update` — change the prompt, the schedule, the tools or the model without changing the
job's identity or its history.
- `schedule_delete` — remove it. Its run records stay.
On disk, everything is a file you can read:
```
/var/lib/q15/agent/schedule/
├── jobs/.json one file per job: prompt, schedule, tools, last status
└── runs/YYYY/MM/DD/