Files
NexusOS/CLAUDE.md
T
Athena Kaminsky 1449280fcd feat(cli): add interactive TUI chat
Add a Textual chat interface with threaded SSE streaming, slash commands, interrupt handling, and bare nexus dispatch. Package it behind the tui extra, document usage, and cover command routing, dependencies, and headless interaction with tests.
2026-08-26 08:17:31 -05:00

9.7 KiB

CLAUDE.md

This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.

What is NexusOS

NexusOS is a local AI assistant platform. It runs a Python/FastAPI backend (Synapse) that interfaces with a locally bundled Ollama instance, a dedicated memory microservice, and a React/Vite frontend. All AI inference runs through Ollama on localhost; no external AI provider is configured or called.

Running the Project

NexusOS runs single-process: the Synapse backend on port 8000 serves the built web UI (interface/web/dist) itself, so there is no separate Vite server at runtime. Ollama is not started on backend process startup itself, but ncp start/start -b bring it up right after via the /ollama/start endpoint; the sidebar Start/Stop AI button and ncp start --ai remain for toggling it independently once the backend is already up.

Windows (recommended):

powershell -ExecutionPolicy Bypass -File .\install-windows.ps1   # one-time native install
ncp web                                                          # backend :8000 + UI

ncp comes from management\ncp.cmd, which the installer puts on the machine PATH — open a NEW shell after installing. .\launch_nexus.ps1 is what the desktop shortcut runs and still works directly.

Linux install / update:

./install.sh              # wrapper over `python3 bin/sync.py restore`
./install.sh --check      # dry run
./install.sh --no-desktop # skip the XFCE wiring (test clones, non-XFCE boxes)

Linux full stack (dev):

./launch_nexus.sh

This activates the Promethean venv and starts the Synapse backend on port 8000 (which also serves the built UI). It additionally starts a Vite dev server for frontend hot-reload — a Linux-dev convenience, unlike the single-process Windows/production path where the backend serves dist/ alone. It does not start Ollama.

Individual services via CLI:

# From nexus-core/ with Promethean venv active:
source Promethean/bin/activate

# Backend (serves the built UI at :8000 too)
uvicorn synapse.main:sio_app --host 127.0.0.1 --port 8000 --reload

# Frontend DEV server (hot-reload) — only when editing the UI; production is the
# built dist/ served by the backend. Run `npm run build` to refresh dist/.
cd interface/web && npm run dev

Management CLI (nexus / ncp) — start/stop services with PID tracking, plus terminal access to the same features as the web UI (REST API on :8000):

./management/nexus-cli.sh start   # starts backend + frontend
./management/nexus-cli.sh stop
./management/nexus-cli.sh start --backend|-b / --frontend|-f / --memory|-m

# Interactive TUI (Hermes/OpenClaw-style; needs pip install 'nexusos-ai[tui]'):
nexus                             # bare command opens the Textual chat TUI
nexus tui                         # same, explicit

# Feature one-shots (dispatch to nexusos_cli/nexus_api.py — httpx):
nexus chat send "<message>"       # stream a reply (POST /chat/stream)
nexus memory list|add <text>|rm <id>
nexus playbook list|show <id>     # first playbook (*) is the active system prompt
nexus history [query]             # recent conversations
nexus monitor                     # ASCII status dashboard (no prompt)

The interactive TUI lives in nexusos_cli/tui_app.py (Textual, optional extra). One-shot subcommands and nexus monitor remain for scripts. The CLI package is nexusos_cli/ (what the wheel ships); management/ keeps desktop-only pieces — shell wrappers, Tk control panel, XFCE panel wiring. management/controlpanel.py (tkinter GUI, wired into the XFCE panel via bin/panel/nexus-popup.py) stays.

Checks (the release gate):

./bin/check.sh          # pytest + eslint + frontend tests + .ps1/.sh parse + wheel build

There is no hosted CI — the remote is self-hosted Gitea with no act_runner — so this script is the gate. Run it before tagging a release.

Frontend lint only:

cd interface/web && npm run lint

Frontend build:

cd interface/web && npm run build

Architecture

Python venv

All Python code runs inside Promethean/ (a local venv). Always activate it before running backend commands: source Promethean/bin/activate. Dependencies are layered: requirements-base.txt holds the GPU-agnostic core (nothing in it needs a GPU or imports torch), and a thin overlay per platform sets the right PyTorch package index — requirements-amd.txt (ROCm), requirements-nvidia.txt (CUDA, generated by bin/gen-nvidia-reqs.py), or requirements-windows.txt (CPU-only, standalone). bin/sync.py (requirements()) selects NVIDIA, AMD, or CPU/Windows requirements from the host and installs that alone by default — fast, no multi-GB downloads.

requirements-ml.txt is a separate, opt-in overlay for local ML inference (transformers/accelerate/bitsandbytes + torch/torchaudio/torchvision) — nothing in synapse/ imports any of it; Ollama does all inference over HTTP. Only pull it in for local model work outside Ollama: pip install -r requirements-amd.txt -r requirements-ml.txt (or -nvidia, or alone for CPU-only torch). Not installed by bin/sync.py/the installers.

Synapse Backend (synapse/)

FastAPI app at synapse/main.py. Key responsibilities:

  • /chat/stream — chat with Ollama; streaming uses SSE (the only chat endpoint — the non-stream /chat was removed). When a conversation goes idle the stream endpoint hands it to the in-process curator, which extracts persistent facts.
  • /playbooks — CRUD for playbooks stored as YAML files in data/playbooks/ via synapse/playbooks/store.py.
  • /memory — CRUD for persistent facts, backed by the same SQLite store the curator writes to.
  • /models — lists, pulls, and deletes Ollama models by proxying Ollama's HTTP API.
  • /settings and /ollama — persist runtime settings and control Ollama lifecycle.
  • /conversations — persists, retrieves, edits, deletes, and exports full chat history from SQLite.
  • /icons — lists local application icons and applies NexusOS branding.

System prompt assembly (in main.py chat_stream_endpoint): the final system prompt is built by layering the active playbook instructions → reference playbook context → persistent memory facts → relevant past conversation snippets retrieved by store.search_conversations.

Memory (synapse/memory/)

Runs in-process — there is no separate service and no second model. curator.py reads the messages a conversation has added since its watermark, extractor.py asks the chat model which permanent facts they contain, and store.py merges the results into memory.db. The backend schedules it when a conversation goes idle (_pending_extractions in main.py), so a half-said fact is never persisted mid-thought. The old :8001 FastAPI app held a second, smaller model that could not share the GPU with the chat model; the chat model is already resident, so the extra hop bought nothing.

Playbook System (synapse/playbooks/ + synapse/playbook_manager.py)

Playbooks are ordered records (title, goal, instructions, tags), each persisted as a {id}.yaml file in data/playbooks/ by PlaybookFileStore (the dir is PLAYBOOK_DIR in nexus_config.py). The first playbook by order is the active system prompt; all subsequent playbooks are injected as reference context. PlaybookManager is the thin class the backend uses to retrieve them and assemble the system prompt.

Ollama (ollama/bin/ollama)

A bundled Ollama binary lives at ollama/bin/ollama. OllamaManager in synapse/ollama_manager.py manages its lifecycle (start/stop/health-check) and selects the best available model. GPU detection uses Vulkan (vulkaninfo) to prefer discrete AMD/NVIDIA GPUs. The Ollama HTTP API is at http://127.0.0.1:11434 (overridable via OLLAMA_HOST env var).

Frontend (interface/web/)

React 19 + Vite. No routing library — App.jsx manages page state in a single currentPage useState. All API calls hit http://localhost:8000 (configured in src/config.js). Built to dist/ (gitignored) via npm run build and served by the backend at :8000 — the mount is in synapse/main.py (_DIST at /, guarded by is_dir()), so dist/ must be built for the UI to appear. Pages: Chatbot, Playbook editor, Conversation History, Models, Memory, Settings, Logs.

Persistent Storage

Most data lands in synapse/memory/memory.db (SQLite, WAL mode). Tables: memory facts, conversations, messages, app settings. synapse/memory/store.py (PersistentMemoryStore) owns the schema and all queries. Playbooks are the exception — they live as YAML files in data/playbooks/ (see Playbook System). nexus_config.py defines all paths; it also ensures all required directories exist on import.

Logs & Runtime State

  • runtime/backend.log, runtime/frontend.log, runtime/memory.log — service stdout
  • runtime/logs/ollama.log, runtime/logs/chat.log
  • runtime/pids/backend.pid, runtime/pids/frontend.pid — used by the management CLI

Key Config

Concern Location
Ollama host OLLAMA_HOST env var (default http://127.0.0.1:11434)
All filesystem paths synapse/nexus_config.py Settings class
Frontend API base URL interface/web/src/config.js
Default chat/memory models synapse/nexus_config.py DEFAULT_CHAT_MODEL / DEFAULT_MEMORY_MODEL
Python dependencies (base) requirements-base.txt
Python dependencies (AMD/ROCm) requirements-amd.txt
Python dependencies (NVIDIA/CUDA) requirements-nvidia.txt (generated by bin/gen-nvidia-reqs.py)
Python dependencies (Windows/CPU) requirements-windows.txt
Python dependencies (optional local ML/torch) requirements-ml.txt (opt-in, not installed by default)