Skip to content

Memory

q15 maintains one continuous conversation. There are no sessions to start, no context windows to manage, no “new chat” button. You send a message on Monday, another on Thursday, another next month — and the agent picks up where you left off because its memory persists across the gap.

What follows is what it remembers, where each layer lives, how it reaches the prompt, and how the background jobs keep it current.

Layer Path In the prompt? Kept by
Core /memory/core/ every turn you
Working /memory/working/ every turn consolidation job
Semantic /memory/semantic/ on demand extraction job
History /memory/history/ after the checkpoint appended every turn
Cognition /memory/cognition/ never the job controller
Notes /memory/notes/ on demand the agent

Each layer is a directory you can open. What is in them, and what maintains each one, is below.

Files under /memory/core/ define who the agent is. These are always in the prompt:

  • AGENT.md — role and behavioral protocol
  • USER.md — user identity, preferences, communication norms
  • SOUL.md — voice, teaching style, working principles

Core memory is operator-maintained. Cognition jobs never write to it. Your agent’s identity is not something a background model call should rewrite.

/memory/working/WORKING_MEMORY.md holds bounded active state: current priorities, active tasks, open threads, recent progress, pending checks. It is the primary mechanism for sessionless continuity.

The consolidation job keeps it compact. Resolved tasks are removed. Abandoned threads are dropped. New constraints are captured. It is a scratch pad for the current moment, not a history log.

Only WORKING_MEMORY.md is auto-injected. Other files under /memory/working/ are not prompt-visible.

Three canonical files hold durable extracted knowledge:

  • facts.md — confirmed facts and grounded inferences
  • preferences.md — user preferences and collaboration preferences
  • projects.md — active projects and durable project knowledge

Semantic memory is not auto-injected. The agent fetches it on demand when durable context is relevant. The extraction job promotes only explicit, durable statements — not single-turn tasks or speculative guesses. If the verification review flags an entry as stale, the next extraction cycle removes or rewrites it.

Completed turns are stored as JSON files under /memory/history/turns/YYYY/MM/DD/. The full transcript persists forever. Nothing is deleted.

Replay is checkpoint-aware. When consolidation succeeds, it advances a checkpoint. The next prompt replays only turns after that checkpoint. The semantic extraction job has its own independent checkpoint for processing a different (typically larger) transcript slice.

/memory/cognition/ is system-owned state. The agent does not read or write here during interactive turns. It contains job trigger state, append-only run records with full provenance (model used, input/output sequence, success/failure), and the verification review artifact.

/memory/notes/ contains the auxiliary zettelkasten notebook: inbox, atomic notes, and structure maps. These are durable knowledge infrastructure that the agent can search and reference on demand.

When the embeddings tool is configured, the agent can index your own content and search it. Qdrant stores the vectors. The default backend is the local q15-tei container, which the agent reaches through the OpenAI-compatible embeddings provider and which needs no API key; a hosted backend is two lines away.

  1. Declare sources. A source names a collection, a source type and a path. The registry is a file: /workspace/.q15/embed/sources.json.
  2. Sync. embed_sync indexes each source, storing a dense vector per chunk plus a BM25 sparse vector.
  3. Search. embed_search runs hybrid search by default, combining semantic recall with lexical matching.
Collection Purpose
library Books, articles and documents from your personal library
zettelkasten Atomic knowledge notes
semantic Extracted semantic memory content
core Core identity and reference material

The collection is a property of the source; the scanner behaviour comes from the source type, which is markdown_tree, markdown_file or chunked_markdown_tree for pre-chunked content.

agent:
tools:
embeddings:
qdrant_url_env: Q15_QDRANT_URL
provider: openai
base_url_env: Q15_EMBEDDINGS_BASE_URL
model: qwen3-embedding-0.6b
dimensions: 1024
batch_size: 128

Changing the backend is a migration, not a config edit. A collection’s dimension is fixed when it is created: change the model or the dimensions and the vector-version stamp changes, every stored record is marked dirty, and the next sync recreates the collections. On a local backend that costs time; on a hosted one it costs money. The runbook is in the repository’s deploy/compose/README.md.

Here is what goes into the model’s context on every turn, ordered from most stable to least stable:

flowchart TD
    subgraph prompt["System messages (cached prefix)"]
        SP["System prompt<br/>code-owned execution policy"]
        CM["Core memory<br/>AGENT.md · USER.md · SOUL.md"]
        SC["Skill catalog<br/>available skills + descriptions"]
        WM["Working memory<br/>WORKING_MEMORY.md"]
    end
    subgraph ctx["Conversation context"]
        RT["Recent transcript<br/>turns after consolidation checkpoint<br/>(bounded replay window)"]
        UM["Current user message<br/>+ temporal metadata"]
    end
    SP --> CM --> SC --> WM --> RT --> UM
    style SP fill:#89b4fa,stroke:#2563eb,color:#1e1e2e
    style CM fill:#cba6f7,stroke:#8839ef,color:#1e1e2e
    style SC fill:#fab387,stroke:#d97706,color:#1e1e2e
    style WM fill:#a6e3a1,stroke:#16a34a,color:#1e1e2e
    style RT fill:#313244,stroke:#89b4fa,color:#cdd6f4
    style UM fill:#f38ba8,stroke:#dc2626,color:#1e1e2e
  1. System prompt — Code-owned execution policy: autonomy rules, tool persistence, verification loop, output contract. Compiled into the binary. Changes only on upgrades. It also tells the agent where every memory layer lives and which paths are auto-injected versus tool-fetched.
  2. Core memory — AGENT.md, USER.md, SOUL.md. Agent identity, user profile, voice and principles. Durable files that change rarely.
  3. Skill catalog — Descriptions of installed skills. Changes when skills are added or removed.
  4. Working memory — WORKING_MEMORY.md. Bounded active state maintained by the consolidation job.
  5. Recent transcript — Turns after the consolidation checkpoint. The unconsolidated tail, bounded by the replay window.
  6. User message — The current message with temporal metadata (local time, day of week, gap since last message).

The ordering lets providers cache the prefix. The system prompt and core memory rarely change, so the cached prefix can be reused even when working memory updates.

q15 does not compress your history. It keeps the transcript and maintains structured memory artifacts through background model calls called cognition jobs. These run on their own schedules and triggers and keep the memory layers current.

They are not the scheduled jobs you create in chat: those live in a different store, they do not count against the job limit, and you cannot list or delete a cognition job. The one lever you have is the model each of them runs on, which is what switch_cognition_model sets below.

The full transcript is always persisted as JSON files under /memory/history/. Nothing is deleted. But only the recent unconsolidated turns are replayed into the prompt. Everything before the consolidation checkpoint is represented by the working memory artifact, which was built from those turns by a background job.

flowchart TD
    subgraph session["One continuous conversation (no sessions)"]
        U1["User message<br/>Monday"] --> R1["Agent reply"]
        R1 --> P1["Persist turn to<br/>/memory/history/"]
        P1 --> D1["Dirty state<br/>+1 turn"]
        D1 -->|"6+ dirty turns"| C1["Working memory<br/>consolidation job"]
        C1 --> CP1["Advance consolidation<br/>checkpoint to turn N"]
        CP1 --> WR["Working memory artifact<br/>updated with active state"]
        WR -->|"Next reply"| N1["Replay only turns<br/>after checkpoint"]
        U2["User message<br/>Thursday"] --> R2["Agent reply"]
        R2 --> N1
        N1 --> D2["Dirty state<br/>+1 turn"]
        D2 -->|"12+ dirty turns"| C2["Semantic memory<br/>extraction job"]
        C2 --> SM["facts.md, preferences.md,<br/>projects.md updated"]
        SM -->|"Available via tools"| N2["Agent can fetch<br/>on demand"]
    end
    style C1 fill:#a6e3a1,stroke:#16a34a,color:#1e1e2e
    style C2 fill:#cba6f7,stroke:#8839ef,color:#1e1e2e
    style CP1 fill:#89b4fa,stroke:#2563eb,color:#1e1e2e
    style WR fill:#313244,stroke:#89b4fa,color:#cdd6f4
    style SM fill:#313244,stroke:#cba6f7,color:#cdd6f4

The key difference: compaction replaces history with a summary. Cognition jobs extract structured state and advance a checkpoint. The history stays on disk. The prompt carries the structured artifact plus only the recent tail that has not been consolidated yet.

Three background jobs maintain memory. Each runs on its own triggers and produces a specific artifact. Jobs can use a different model than the interactive loop — assign a stronger reasoning model to verification, a large-context model to extraction, and a lighter model to consolidation.

Verification review (verification_review) audits the other memory layers for accuracy. It reads core memory, working memory, and recent transcript turns, then produces a review artifact flagging stale, unsupported, or contradicted entries. It runs twice daily or after 8+ unconsolidated turns. Its output is consumed only by the other two jobs — never by the interactive agent directly.

Working memory consolidation (working_memory.consolidate) maintains WORKING_MEMORY.md, the bounded active-state artifact auto-injected into every prompt. It reads the current working memory, the verification review, and the last 16 transcript turns, then updates the file to reflect current priorities, active tasks, open threads, and recent progress. It runs daily or after 6+ unconsolidated turns. On success, it advances the consolidation checkpoint — the next reply replays only turns after that boundary.

Semantic memory extraction (semantic_memory.extract) maintains three durable knowledge files: facts.md, preferences.md, and projects.md. It reads all three files, the working memory, the verification review, and transcript turns since its last checkpoint, then promotes durable facts, preferences, and project knowledge. It runs daily or after 12+ unconsolidated turns. Semantic memory is not auto-injected — the agent fetches it on demand with read_file.

Cognition jobs produce artifacts that reach the interactive agent through four paths:

Path Artifacts Mechanism
Auto-injected Core memory, working memory Injected into every prompt as system sections
Tool-fetched Semantic memory, zettelkasten notes Agent calls read_file on demand
Checkpoint replay Recent transcript turns Replayed verbatim for turns after the consolidation checkpoint
Vector search All memory layers (when embeddings configured) Agent calls embed_search for semantic retrieval across Qdrant collections

The consolidation checkpoint marks the replay boundary. Everything before it is represented by the working memory artifact. Everything after it is replayed verbatim:

flowchart TD
    subgraph disk["Full transcript on disk (never deleted)"]
        Old["Turns 1 to 47<br/>(before checkpoint)"]
        New["Turns 48 to 50<br/>(after checkpoint)"]
    end
    CP["Consolidation checkpoint<br/>at turn 47"]
    CP -.->|"marks the boundary"| New
    Old -.->|"represented by"| WM["WORKING_MEMORY.md"]
    New -->|"replayed into"| Prompt["Goes into prompt"]
    WM -->|"injected into"| Prompt
    style CP fill:#89b4fa,stroke:#2563eb,color:#1e1e2e
    style Old fill:#313244,stroke:#585b70,color:#cdd6f4
    style New fill:#a6e3a1,stroke:#16a34a,color:#1e1e2e
    style WM fill:#a6e3a1,stroke:#16a34a,color:#1e1e2e
    style Prompt fill:#313244,stroke:#89b4fa,color:#cdd6f4

When embeddings are configured, the agent can also search memory semantically via Qdrant. The semantic, core, library, and zettelkasten collections are indexed with hybrid dense+sparse vectors. The agent calls embed_search with a natural-language query and receives ranked results with file paths and relevance scores. Embeddings are optional — when not configured, the agent relies on the other three paths. See Searching your own files above for the configuration of both.

The three jobs form a self-correcting loop. The verification job audits state; the consolidation and extraction jobs apply corrections. This cycle runs continuously in the background, keeping memory artifacts accurate without user intervention.

flowchart TD
    VR["Verification review<br/>audits state"] -->|"flags stale<br/>unsupported<br/>contradicted entries"| WM["Working memory<br/>consolidation"]
    VR -->|"flags stale<br/>unsupported<br/>contradicted entries"| SM["Semantic memory<br/>extraction"]
    WM -->|"removes or rewrites<br/>flagged entries"| WMF["WORKING_MEMORY.md<br/>corrected"]
    SM -->|"removes, rewrites,<br/>or downgrades"| SMF["facts.md<br/>preferences.md<br/>projects.md<br/>corrected"]
    WM -->|"updates active state"| SM
    WMF -->|"next turn"| Prompt["Prompt"]
    SMF -.->|"on demand"| Prompt
    style VR fill:#f9e2af,stroke:#df8e1d,color:#1e1e2e
    style WM fill:#a6e3a1,stroke:#16a34a,color:#1e1e2e
    style SM fill:#cba6f7,stroke:#8839ef,color:#1e1e2e
    style WMF fill:#313244,stroke:#a6e3a1,color:#cdd6f4
    style SMF fill:#313244,stroke:#cba6f7,color:#cdd6f4
    style Prompt fill:#313244,stroke:#89b4fa,color:#cdd6f4

The ordering emerges from trigger thresholds and cron schedules, not from hardcoded sequencing:

  1. Verification review runs first (8+ dirty turns, or twice daily). It produces a fresh review artifact.
  2. Working memory consolidation runs next (6+ dirty turns, or daily). It reads the verification review and applies corrections.
  3. Semantic memory extraction runs last (12+ dirty turns, or daily). It reads both the verification review and the freshly consolidated working memory, then applies corrections to the semantic files.

A burst of conversation can trigger all three jobs in sequence within minutes.

Each job has two triggers: a state trigger (enough unconsolidated turns have accumulated) and a schedule trigger (a fixed UTC time). The controller evaluates both after every turn and launches a job that is due. A job that is already running is not started a second time.

Job Identifier Schedule trigger State trigger
Verification review verification_review 03:00 and 15:00 UTC 8+ unconsolidated turns
Working-memory consolidation working_memory.consolidate 04:00 UTC 6+ unconsolidated turns
Semantic-memory extraction semantic_memory.extract 05:00 UTC 12+ unconsolidated turns

Every run is recorded under /memory/cognition/runs/YYYY/MM/DD/ as a JSON file carrying the job type, the trigger cause, the model and provider used, the start and end times, the outcome, and whether the target files changed. The records are an audit trail: you can see when each job ran, what triggered it, and what it did.

/memory/cognition/
├── state/ job state, checkpoints, and the latest verification review
└── runs/YYYY/MM/DD/ one JSON record per run

Models. A job runs on its own model when one is set, and otherwise inherits the interactive model. Set a per-job model in chat with switch_cognition_model, naming the job type above. Verification benefits from a careful model; consolidation is mechanical. See Model selection.

Files. Each job may write only the artifacts it owns: consolidation writes the working-memory file, extraction writes the three semantic files, and verification writes its review. None of them can write another job’s artifact.

The obvious alternative is to have a model compress older turns into a summary and keep only the recent ones verbatim. q15 does not do that: a summary is lossy, unstructured and opaque, and the transcript it replaces is worth keeping. The reasoning, and the comparison, are in Architecture.

When you talk to q15, you are having one conversation with an agent that remembers. The working memory artifact carries your current thread forward. Semantic memory carries durable facts about you and your projects. The full transcript is on disk. Background jobs maintain all of this while you are not looking — not by compressing your words into a paragraph, but by extracting structured state that the next turn can use.

The verification job audits that state. The consolidation job keeps the active-state artifact current. The extraction job promotes durable knowledge. The correction cycle means memory does not just persist — it gets better over time, as stale entries are flagged and corrected without any user intervention.

Tools

The surface the agent acts with, and how embeddings fit into it.

Architecture

The four services, and how memory fits into the turn flow.

Workspace

The durable project tree that pairs with memory for long-running work.