# WORK ORDER — Install GLAMMBRAIN, the naked memory stack (GLAMMBOX pack)

You are an AI agent with write access to one folder on the user's machine.
Execute this work order to stand up a real hybrid memory stack: a vector
store, a local embedder, a keyword index, and a cross-encoder reranker,
fused into one search command. An optional graph stays a separate relationship
tool. Ask before anything
that costs money (nothing here should). Everything in this pack is free and
open source.

## TRUTH CONTRACT — read this before you start
- **This ships EMPTY.** There is no pre-loaded data, no seeded brain, no
  "your files are already in here." Every fact in it comes from the feed
  script in Step 8, run by the user, on the user's own files.
- **Never say "one-click" or "no setup."** This is Docker, a Python
  environment, and several running services. It is real infrastructure
  work, done once, not a toggle.
- **Never say "never leaves your machine."** Everything here runs locally
  by default and nothing in this pack calls an external API — but that is
  a property of this configuration, not a law of the universe. If the user
  later points a model or a tool at a cloud service, that claim stops being
  true. State what is actually running, not a permanent guarantee.
- **The verified installer path is Linux with systemd or WSL2 + Ubuntu.**
  macOS is a guided beta, not a proven native installer: use
  `pack-glammbrain-macos.md` as the platform override and do not run Linux
  systemd steps unchanged. Windows native (no WSL2) is experimental.
- **A packaged tarball now exists; no public GitHub repo yet.** It is
  served from this same site at
  `/downloads/glammbrain-20260827-full-v4.tar.gz`
  (sha256 `c7fde89a22ba289891041c25718891921312f107d0b1a9e7184f6147299c2723`).
  Verify the hash before running anything extracted from it. Once
  downloaded and extracted, the tarball's own `brain-engine/AGENT-PACK.md`
  is the authoritative, up-to-date install recipe and supersedes the
  inline recipe in this file — hand the agent that file instead if it is
  available. This file remains the standalone fallback recipe.

## Goal
A running local stack — Qdrant (vectors), optionally Neo4j (a separate graph tool), Ollama
running embeddinggemma:300m (embedder), a BM25 keyword index, and a
reranker service on port 18793 — plus a feed script that loads one folder
of the user's files, and a search script that fuses meaning and exact-word results into ranked
results with a receipt (source file + snippet) behind every hit.

## Prerequisites
- Linux with systemd or WSL2 + Ubuntu (Windows). On macOS, stop and read
  `pack-glammbrain-macos.md` first; it replaces the Linux service and timer
  assumptions. Confirm with the user before proceeding on Windows without
  WSL2 — flag it as experimental and continue only if they accept that.
- Docker installed and running (`docker --version`). Free, from
  https://docs.docker.com/get-started/.
- Ollama installed and running (`ollama --version`). Free, from
  https://ollama.com/download.
- Python 3.10+ available (`python3 --version`).
- 16 GB RAM recommended, 8 GB is a tight floor. GPU optional — every model
  below has a CPU path, just slower.
- One folder of the user's own files (Markdown, text, transcripts, dated
  notes) to feed once the stack is up.

## Step 1 — confirm the source folder
Ask the user which folder is the source of truth. That folder is the truth
forever after — indexes get rebuilt from it; the folder is never edited to
match an index. Do not proceed until one real folder is named.

## Step 2 — docker-compose.yml (Qdrant + Neo4j)
Create `glammbrain/docker-compose.yml`:

```yaml
version: "3.9"
services:
  qdrant:
    image: qdrant/qdrant:latest
    container_name: glammbrain-qdrant
    ports:
      - "6333:6333"
      - "6334:6334"
    volumes:
      - ./data/qdrant:/qdrant/storage
    restart: unless-stopped

  neo4j:
    image: neo4j:5
    container_name: glammbrain-neo4j
    ports:
      - "7474:7474"
      - "7687:7687"
    environment:
      - NEO4J_AUTH=neo4j/CHANGE_THIS_PASSWORD
    volumes:
      - ./data/neo4j:/data
    restart: unless-stopped
```

Tell the user to replace `CHANGE_THIS_PASSWORD` with a real password of
their own before the first `docker compose up` — this env var sets the
Neo4j password directly instead of the interactive first-login flow.
Neo4j is optional: if the user only wants meaning + keyword search, skip
the `neo4j` service block entirely and move on.

Run it:
```
cd glammbrain
docker compose up -d
```
Confirm Qdrant answers: `curl http://localhost:6333/collections` should
return JSON, not a connection error. If Neo4j is included, confirm
http://localhost:7474 loads in a browser.

## Step 3 — pull the embedder
```
ollama pull embeddinggemma:300m
```
This is a small (300M parameter), free, local embedding model from Google,
served through Ollama. It produces 768-dimension vectors. Confirm it is
present: `ollama list` should show `embeddinggemma:300m`.

## Step 4 — Python environment
```
cd glammbrain
python3 -m venv venv
source venv/bin/activate        # Windows/WSL: same command inside WSL bash
pip install --upgrade pip
pip install qdrant-client neo4j rank-bm25 sentence-transformers fastapi uvicorn requests
```
`FlagEmbedding` is optional and only needed for the primary reranker path
in Step 7:
```
pip install FlagEmbedding
```
If that install fails or the machine is CPU-only with no GPU, skip it —
Step 7's fallback path covers that case without it.

## Step 5 — create the Qdrant collection (768-dim, cosine)
Save as `glammbrain/create_collection.py`:

```python
from qdrant_client import QdrantClient
from qdrant_client.models import Distance, VectorParams

client = QdrantClient(url="http://localhost:6333")
client.recreate_collection(
    collection_name="glammbrain",
    vectors_config=VectorParams(size=768, distance=Distance.COSINE),
)
print("collection 'glammbrain' ready: 768-dim, cosine")
```
Run: `python create_collection.py`

## Step 6 — chunk, embed, and feed the folder
Save as `glammbrain/feed.py`:

```python
import os, sys, glob, json, hashlib, requests
from qdrant_client import QdrantClient
from qdrant_client.models import PointStruct

SOURCE_DIR = sys.argv[1] if len(sys.argv) > 1 else "./source"
CHUNK_SIZE = 900       # characters, not tokens — simple and good enough to start
CHUNK_OVERLAP = 150

client = QdrantClient(url="http://localhost:6333")

def embed(text):
    r = requests.post(
        "http://localhost:11434/api/embeddings",
        json={"model": "embeddinggemma:300m", "prompt": text},
        timeout=60,
    )
    r.raise_for_status()
    return r.json()["embedding"]

def chunk(text):
    out, i = [], 0
    while i < len(text):
        out.append(text[i:i + CHUNK_SIZE])
        i += CHUNK_SIZE - CHUNK_OVERLAP
    return out

points, bm25_corpus = [], []
files = [f for f in glob.glob(f"{SOURCE_DIR}/**/*", recursive=True) if os.path.isfile(f)]
print(f"found {len(files)} files under {SOURCE_DIR}")

for path in files:
    if not path.lower().endswith((".md", ".txt")):
        continue
    with open(path, "r", errors="ignore") as fh:
        text = fh.read()
    for idx, piece in enumerate(chunk(text)):
        if not piece.strip():
            continue
        point_id = hashlib.sha256(f"{path}:{idx}".encode()).hexdigest()[:32]
        vector = embed(piece)
        points.append(PointStruct(
            id=point_id,
            vector=vector,
            payload={"path": path, "chunk": idx, "text": piece},
        ))
        bm25_corpus.append({"id": point_id, "path": path, "chunk": idx, "text": piece})

if points:
    client.upsert(collection_name="glammbrain", points=points)
print(f"upserted {len(points)} chunks into Qdrant")

with open("bm25_corpus.json", "w") as fh:
    json.dump(bm25_corpus, fh)
print("wrote bm25_corpus.json for the keyword index")
```
Run: `python feed.py /path/to/the/source/folder`

Nothing is indexed before this command runs. This is the step that turns
the empty stack into an actual memory.

## Step 7 — build the BM25 keyword index and the reranker service
Save as `glammbrain/bm25_index.py`:

```python
import json
from rank_bm25 import BM25Okapi

with open("bm25_corpus.json") as fh:
    corpus = json.load(fh)

tokenized = [c["text"].lower().split() for c in corpus]
bm25 = BM25Okapi(tokenized)

def search(query, top_k=10):
    scores = bm25.get_scores(query.lower().split())
    ranked = sorted(zip(scores, corpus), key=lambda x: x[0], reverse=True)[:top_k]
    return [{"score": float(s), **c} for s, c in ranked]
```

Save as `glammbrain/reranker_service.py` — a small FastAPI service on
port 18793. It tries the primary model first and falls back to the CPU
cross-encoder if that model or a GPU is unavailable:

```python
from fastapi import FastAPI
from pydantic import BaseModel

app = FastAPI()
mode = "primary"
reranker = None

try:
    from FlagEmbedding import FlagReranker
    reranker = FlagReranker("BAAI/bge-reranker-v2-m3", use_fp16=False)
except Exception:
    from sentence_transformers import CrossEncoder
    reranker = CrossEncoder("cross-encoder/ms-marco-MiniLM-L-6-v2")
    mode = "cpu-fallback"

class RerankRequest(BaseModel):
    query: str
    passages: list[str]

@app.get("/health")
def health():
    return {"status": "ok", "mode": mode}

@app.post("/rerank")
def rerank(req: RerankRequest):
    pairs = [[req.query, p] for p in req.passages]
    if mode == "primary":
        scores = reranker.compute_score(pairs)
    else:
        scores = reranker.predict(pairs).tolist()
    ranked = sorted(zip(scores, req.passages), key=lambda x: x[0], reverse=True)
    return {"mode": mode, "results": [{"score": float(s), "passage": p} for s, p in ranked]}
```
Run it: `uvicorn reranker_service:app --host 0.0.0.0 --port 18793`

Confirm: `curl http://localhost:18793/health` should return
`{"status":"ok","mode":"primary"}` (GPU/FlagEmbedding available) or
`{"status":"ok","mode":"cpu-fallback"}` (CPU-only machine — still correct,
just slower).

## Step 8 — one search command that fuses meaning + exact words, with source paths
Save as `glammbrain/search.py`:

```python
import sys, json, requests
from qdrant_client import QdrantClient
from bm25_index import search as bm25_search

QDRANT = QdrantClient(url="http://localhost:6333")

def embed(text):
    r = requests.post(
        "http://localhost:11434/api/embeddings",
        json={"model": "embeddinggemma:300m", "prompt": text},
        timeout=60,
    )
    r.raise_for_status()
    return r.json()["embedding"]

def vector_search(query, top_k=10):
    hits = QDRANT.search(collection_name="glammbrain", query_vector=embed(query), limit=top_k)
    return [{"score": h.score, "path": h.payload["path"], "text": h.payload["text"]} for h in hits]

def reciprocal_rank_fusion(*ranked_lists, k=60):
    scores = {}
    for ranked in ranked_lists:
        for rank, item in enumerate(ranked):
            key = (item["path"], item.get("chunk", item["text"][:40]))
            scores[key] = scores.get(key, {"item": item, "score": 0.0})
            scores[key]["score"] += 1.0 / (k + rank + 1)
    return sorted(scores.values(), key=lambda x: x["score"], reverse=True)

def search(query, top_k=8):
    vec_hits = vector_search(query, top_k=top_k)
    bm25_hits = bm25_search(query, top_k=top_k)
    fused = reciprocal_rank_fusion(vec_hits, bm25_hits)[:top_k]
    passages = [f["item"]["text"] for f in fused]

    try:
        r = requests.post(
            "http://localhost:18793/rerank",
            json={"query": query, "passages": passages},
            timeout=30,
        )
        reranked = r.json()["results"]
    except Exception:
        reranked = [{"score": f["score"], "passage": f["item"]["text"]} for f in fused]

    receipt = {"query": query, "top_hits": []}
    for r in reranked[:5]:
        source = next((f["item"] for f in fused if f["item"]["text"] == r["passage"]), None)
        receipt["top_hits"].append({
            "score": r["score"],
            "path": source["path"] if source else "unknown",
            "snippet": r["passage"][:200],
        })
    return receipt

if __name__ == "__main__":
    query = " ".join(sys.argv[1:]) or "test query"
    result = search(query)
    print(json.dumps(result, indent=2))
    with open("search_receipts.jsonl", "a") as fh:
        fh.write(json.dumps(result) + "\n")
```
Run: `python search.py your real question here`

This recipe appends a line to `search_receipts.jsonl` — query, top hits,
source paths and scores — so its own verification is repeatable. Other search
implementations may return source paths without keeping a permanent query log.

## Verification
1. Restart the containers (`docker compose restart`) and confirm the Qdrant
   collection still reports the same point count — the index must survive
   a restart because of the mounted volume.
2. Ask the search script one question the user already knows the answer to.
   The top hit must name the correct source file and a snippet that
   actually contains the answer. A fluent-sounding result with no matching
   source file is a fail — treat it as one.
3. If the index and a source file ever disagree, the file wins: fix or
   re-run the feed on that file, never hand-edit the index to match a
   stale answer.

## Honest costs
Every component here is free and open source: Docker, Qdrant, Neo4j,
Ollama, embeddinggemma:300m, rank-bm25, sentence-transformers, FastAPI.
The only costs are disk space, RAM, and the time to run these steps once.
This pack does not call any paid API and asks before anything that would.
The packaged tarball (see the TRUTH CONTRACT above) is MIT-licensed and
free to keep, fork, or hand to anyone; there is still no public GitHub
repo — the download and this pack remain the full recipe today.
