Skip to content

Architecture

q15 is four services, one storage contract, and one boundary you can point at. Nothing about it is novel except where the lines are drawn.

sequenceDiagram
    participant You
    participant Web as q15-web
    participant Agent as q15-agent
    participant Exec as q15-exec
    participant Proxy as q15-proxy
    participant Model as Model provider

    You->>Web: chat message (sealed)
    Web->>Agent: bridge socket
    Agent->>Agent: assemble prompt, pick model
    Agent->>Model: prompt + tools
    Model-->>Agent: text or tool calls
    loop Tool execution
        Agent->>Exec: command (gRPC)
        Exec->>Proxy: outbound request
        Proxy->>Proxy: apply policy, inject credentials
        Proxy-->>Exec: response, credentials stripped
        Exec-->>Agent: command output
    end
    Agent->>Agent: persist the turn to /memory/history/
    Agent-->>Web: streamed reply
    Web-->>You: sealed frames

Telegram, if you enable it, enters at q15-agent directly and joins the same conversation and the same memory.

Service Owns Listens
q15-agent Prompt assembly, tool wiring, memory, scheduled jobs, Telegram I/O, file operations Telegram (long poll), bridge socket
q15-exec Command execution, session lifecycle, Nix package management :50051 (gRPC)
q15-proxy Credential injection, request mutation, egress policy, TLS interception :18080 (HTTP proxy), :50052
q15-web The browser client: passkey authentication, chat, history, attachments 127.0.0.1:8080 (published)
q15-qdrant Embedding collections (optional) :6334 inside the Compose network
q15-tei Local embedding backend (optional) :80 inside the Compose network

The agent reaches the executor at q15-exec:50051; the executor routes outbound traffic through q15-proxy:18080. The agent never talks to the proxy directly, because the executor is the only egress boundary. The browser client reaches the agent over a Unix socket, unix:///run/q15/bridge.sock, on a volume mounted only in those two services.

Most agent platforms run everything in one process: the model, the tools, the credentials and the network access share a trust domain. If the model is compromised through prompt injection, the attacker has everything.

q15 splits that into four services with hard boundaries between them:

  • Credentials live in the proxy, not the agent. The agent never reads an API token. When a command calls api.github.com, the proxy attaches the token at the network layer. A compromised prompt cannot exfiltrate a secret the agent never had.
  • Execution is isolated. Commands run through q15-exec, with file access rooted to /workspace, /memory and /skills.
  • Egress is unconditional. The agent cannot ask the executor to bypass the proxy; the routing is a property of the deployment, not of the prompt.
  • The browser client holds no agent credentials. It dials the bridge socket, keeps its own credential store, mounts no memory, and stores no transcript of its own.

The design target is that you never have to trust the model: the architecture constrains what a compromised model can actually do. What it does not protect is stated on Why q15.

q15-agent is the core of the stack. It owns prompt assembly, the tools, both clients, memory, and the file operations under a rooted path set.

Concern What it does
Prompt assembly Identity files from /memory/core/, working memory, the skill catalog, recent unconsolidated turns, and your message.
Tool wiring Files, command execution through q15-exec, web search and fetch, embeddings, media, skills, subagents, scheduling.
Clients A Telegram bot over long polling, and the bridge socket q15-web dials. Both join one conversation and one memory.
Memory Turns persisted to /memory/history/, plus the background jobs that maintain the working and semantic layers.
Rooted file access Reads and writes are confined to /workspace, /memory and /skills. Nothing above those roots is reachable.
Scheduled jobs Isolated runs with their own turn budget, tool allow-list and model pin.

Model choice is runtime state, not configuration. Each provider contributes a roster, which the agent discovers at startup and refreshes; each turn filters that roster by the capabilities it needs and ranks what is left. You change the model in chat. See Models.

Completed turns are JSON files under /memory/history/turns/YYYY/MM/DD/. A turn carries ordered messages[].parts[] entries of type text, reasoning, tool_call and tool_result, so reasoning and tool activity survive a restart and replay in the transcript. On startup the agent upgrades stored history to the current schema in place.

q15-exec owns command execution.

Commands run through nix shell, so any package from nixpkgs is available on demand without rebuilding an image:

Terminal window
nix shell nixpkgs#git --command bash -lc "git log --oneline -5"

/nix is persistent, so the first run of a package is slower than the ones after it. Sessions are held in memory inside the exec container, which is why one stack runs one executor: session routing stays local and storage ownership stays unambiguous.

The executor also deliberately passes --option build-users-group "" to nix shell, so builds run as the container user rather than dropping to a nixbld user that a container’s seccomp profile may block from executing anything. The security boundary for those builds is the container, not Nix user isolation inside it.

Every LLM has a fixed context window. When you send a message, the agent assembles a prompt: system instructions, identity, recent conversation history, and your new message. The model processes all of it at once.

If your conversation is 500 turns long, you cannot send all 500 turns every time. Something has to be dropped.

The standard approach is compaction. When the transcript exceeds a threshold, the agent asks the model to summarize the older portion into a compressed paragraph, then keeps only recent turns verbatim.

flowchart TD
    subgraph compaction["Compaction cycle"]
        T1["Turns 1–200<br/>(full transcript)"] --> S["Model summarizes<br/>turns 1–180"]
        S --> Summary["Compressed summary<br/>(~500 tokens of prose)"]
        Summary --> M1["Next prompt:<br/>summary + turns 181–200"]
        T2["Turns 201–400<br/>(grows again)"] --> S2["Summarize again:<br/>summary + turns 181–380"]
        S2 --> Summary2["New compressed summary"]
        Summary2 --> M2["Next prompt:<br/>new summary + turns 381–400"]
    end
    style S fill:#fab387,stroke:#d97706,color:#1e1e2e
    style S2 fill:#fab387,stroke:#d97706,color:#1e1e2e
    style Summary fill:#313244,stroke:#89b4fa,color:#cdd6f4
    style Summary2 fill:#313244,stroke:#89b4fa,color:#cdd6f4

Compaction has four problems:

  • Lossy and opaque. The summary is a paragraph of prose. If the model omitted a detail during summarization, it is gone from the prompt. The agent no longer knows it.
  • Compounding drift. Each compaction cycle summarizes the previous summary, not the original turns. After three cycles, the summary is a summary of a summary of a summary. Detail erodes. Errors compound.
  • Unstructured. The summary is free-form text with no schema, no separation between active state and durable facts, no way to distinguish “what we’re working on right now” from “the user’s name.”
  • Blocking. Compaction happens during the user’s turn. The model summarizes before it can answer, adding latency and risking mid-conversation failure.
Compaction Cognition jobs
What happens to old turns Summarized into prose, then dropped from prompt Persisted on disk; represented by structured artifacts
What the prompt carries Summary paragraph + recent turns Structured working memory + recent unconsolidated turns
Lossiness Lossy — detail erodes with each compaction cycle Non-lossy — full transcript on disk; extraction is additive
Structure Free-form prose, no schema Canonical headings, typed files, enforced section structure
Verifiability Cannot query what was lost Full run records with model, input/output seq, success/failure
Correction No mechanism — errors compound Verification job flags stale claims; next cycle corrects them
When it runs Blocking, during the user’s turn Background, between turns
Failure mode Conversation stalls if summarization fails Job fails, dirty state preserved, retries on next trigger
Model choice Same model as conversation Per-job model override or interactive model inheritance

Every stack owns persistent volumes, and they are the reason nothing is lost on an update:

Path Purpose
/workspace Durable project tree, working files, embedding sync state
/memory Identity, semantic knowledge, working state, turn history
/skills Installed skill artifacts
/media The media store, including chat attachments
/var/lib/q15/agent Scheduled-job definitions and run records
/var/lib/q15/proxy Proxy-owned durable state
/var/lib/q15-web The web tier’s credential store and admin socket
/nix Executor package store, so fetched packages stay warm

/workspace may start empty on a first deploy. That is a valid initial state, and it is expected to be long-lived from then on: nothing in the stack treats it as scratch space.

The agent’s memory and the web tier’s credential store are separate on purpose. The agent cannot read the credentials that prove you are you, and the web tier cannot read your memory. The volume names, mounts and contents are in the reference; what to copy and in which order is in Backups.

One stack is one q15-agent, one q15-exec, one q15-proxy and one q15-web, with a Qdrant container and an embedding backend alongside them. The Compose example in the repository is the supported deployment: published images, one pinned release tag, named volumes for every path above.

Two bootstraps happen on the host rather than in the containers, and they are easy to confuse:

  • q15-auth produces the OAuth credential q15 uses as a client of a model provider.
  • q15-web auth enroll registers you as the owner of the browser client, with a passkey.

Security

The threat model, and the class of attack the architecture is built around.

Credentials

Where secrets live, how they are injected, and the two credential stores that are not the same.

Reference

Every key, path, port and command in one place.