Files
NexusOS/CLAUDE.md
T
Athena KaminskyandClaude Opus 5 affba1805c feat(chat): add run_snippet, an execution track beside the render track
render_preview validates markup and hands it to the browser, which renders it
in an opaque-origin sandboxed iframe. Nothing executes server-side. That model
fits HTML/SVG/JSX and cannot fit C, Rust or Erlang, which need a real
toolchain - so those get a second tool instead of a widened first one.

The split is the feature: the model picks a track by picking a tool, rather
than picking a `lang` value from an enum where half the entries run
server-side and half do not.

synapse/code_run.py compiles and runs one file in a throwaway directory and
returns a ```nexus-run fence carrying the source and its captured output
together, so a model cannot paste output without the code that produced it.
Backticks in the source are re-encoded as ` - still valid JSON, and it
cannot close the fence early.

It is not a sandbox, and the module docstring says so up front. What it gives
is containment by layers: consent (an action tool, gated by
action_tool_policy, per-call Approve/Deny on "ask"), static screening, a
scrubbed environment in a temp dir, wall-clock and POSIX rlimits, and a
network namespace on Linux where unprivileged userns are available. Screening
is a tripwire against a model reaching for `requests` out of habit, not a
boundary against an adversary; layers 1 and 3-5 are the load-bearing ones.

Backend RUN_LANGS and frontend run-langs.js are separate registries because
the two sides need different things - one executes, one labels - and neither
should depend on the other at runtime. tests/test_tools.py asserts the key
sets and the fence tag stay equal, so drift fails the gate instead of
rendering a run result under the wrong language.

tests/snippet_probes/ is a data catalog rather than inlined cases, so adding a
language is a data change and the meta-tests can assert every RUN_LANGS key
has both a smoke probe and a screening probe. Probes skip cleanly on hosts
without the toolchain.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-20 14:30:03 -05:00

154 lines
9.4 KiB
Markdown

# CLAUDE.md
This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.
## What is NexusOS
NexusOS is a local AI assistant platform. It runs a Python/FastAPI backend (Synapse) that interfaces with a locally bundled Ollama instance, a dedicated memory microservice, and a React/Vite frontend. All AI inference runs through Ollama on `localhost`; no external AI provider is configured or called.
## Running the Project
NexusOS runs **single-process**: the Synapse backend on port 8000 serves the
built web UI (`interface/web/dist`) itself, so there is no separate Vite server
at runtime. Ollama is started manually (sidebar **Start AI** / `nexus-cli.sh
start --ai`), not on backend startup.
**Windows (recommended):**
```powershell
powershell -ExecutionPolicy Bypass -File .\install-windows.ps1 # one-time native install
ncp web # memory :8001 + backend :8000 + UI
```
**Linux full stack (dev):**
```bash
./launch_nexus.sh
```
This activates the `Promethean` venv and starts the memory service on port 8001 and the Synapse backend on port 8000 (which also serves the built UI). It additionally starts a **Vite dev server** for frontend hot-reload — a Linux-dev convenience, unlike the single-process Windows/production path where the backend serves `dist/` alone. It does **not** start Ollama.
**Individual services via CLI:**
```bash
# From nexus-core/ with Promethean venv active:
source Promethean/bin/activate
# Backend (serves the built UI at :8000 too)
uvicorn synapse.main:sio_app --host 0.0.0.0 --port 8000 --reload
# Memory service
uvicorn synapse.memory.service:app --host 0.0.0.0 --port 8001 --reload
# Frontend DEV server (hot-reload) — only when editing the UI; production is the
# built dist/ served by the backend. Run `npm run build` to refresh dist/.
cd interface/web && npm run dev
```
**Management CLI** (`ncp`) — start/stop services with PID tracking, plus terminal
access to the same features as the web UI (all via the REST API on `:8000`):
```bash
./management/nexus-cli.sh start # starts backend + frontend
./management/nexus-cli.sh stop
./management/nexus-cli.sh start --backend|-b / --frontend|-f / --memory|-m
# Feature commands (dispatch to nexusos_cli/nexus_api.py — httpx, no TUI):
ncp chat "<message>" # stream a reply (POST /chat/stream)
ncp memory list|add <text>|rm <id>
ncp playbook list|show <id> # first playbook (*) is the active system prompt
ncp history [query] # recent conversations
```
The old curses TUIs (`nexus-chat.py`, `nexus-playbook.py`) were removed in favor of
these API-backed subcommands. The CLI covers chat, memory, playbooks, and history;
the web UI and control panel expose the remaining management features.
The CLI itself lives in `nexusos_cli/` (that is what the wheel ships and what
`nexus`/`ncp`/`nexusos` dispatch to); `management/` keeps the desktop-only
pieces — the shell wrappers, the Tk control panel, and the XFCE panel wiring.
`management/controlpanel.py` (tkinter GUI, wired into the XFCE panel via
`bin/panel/nexus-popup.py`) stays.
**Checks (the release gate):**
```bash
./bin/check.sh # pytest + eslint + frontend tests + .ps1/.sh parse + wheel build
```
There is no hosted CI — the remote is self-hosted Gitea with no act_runner — so
this script *is* the gate. Run it before tagging a release.
**Frontend lint only:**
```bash
cd interface/web && npm run lint
```
**Frontend build:**
```bash
cd interface/web && npm run build
```
## Architecture
### Python venv
All Python code runs inside `Promethean/` (a local venv). Always activate it before running backend commands: `source Promethean/bin/activate`. Dependencies are layered: `requirements-base.txt` holds the GPU-agnostic core, and a thin overlay pins the right PyTorch build for the target — `requirements-amd.txt` (ROCm), `requirements-nvidia.txt` (CUDA, generated by `bin/gen-nvidia-reqs.py`), or `requirements-windows.txt` (CPU-only). `bin/sync.py` (`requirements()`) selects NVIDIA, AMD, or CPU/Windows requirements from the host.
### Synapse Backend (`synapse/`)
FastAPI app at `synapse/main.py`. Key responsibilities:
- `/chat/stream` — chat with Ollama; streaming uses SSE (the only chat endpoint — the non-stream `/chat` was removed). After each exchange the stream endpoint calls the Memory Service to auto-extract persistent facts.
- `/playbooks` — CRUD for playbooks stored as YAML files in `data/playbooks/` via `synapse/playbooks/store.py`.
- `/memory` — CRUD for persistent facts (proxies the same SQLite store as the memory service).
- `/models` — lists, pulls, and deletes Ollama models by proxying Ollama's HTTP API.
- `/settings` and `/ollama` — persist runtime settings and control Ollama lifecycle.
- `/conversations` — persists, retrieves, edits, deletes, and exports full chat history from SQLite.
- `/icons` — lists local application icons and applies NexusOS branding.
**System prompt assembly** (in `main.py` `chat_stream_endpoint`): the final system prompt is built by layering the active playbook instructions → reference playbook context → persistent memory facts → relevant past conversation snippets retrieved by `store.search_conversations`.
### Memory Service (`synapse/memory/`)
A separate FastAPI app on port 8001. `service.py` exposes `/memories/extract` which calls `extractor.py` — an Ollama prompt that decides whether to persist a new fact from a conversation exchange. The main Synapse backend calls this asynchronously after each streaming response. Both services share the same SQLite database (`synapse/memory/memory.db`).
### Playbook System (`synapse/playbooks/` + `synapse/playbook_manager.py`)
Playbooks are ordered records (title, goal, instructions, tags), each persisted as a `{id}.yaml` file in `data/playbooks/` by `PlaybookFileStore` (the dir is `PLAYBOOK_DIR` in `nexus_config.py`). The **first** playbook by order is the active system prompt; all subsequent playbooks are injected as reference context. `PlaybookManager` is the thin class the backend uses to retrieve them and assemble the system prompt.
### Ollama (`ollama/bin/ollama`)
A bundled Ollama binary lives at `ollama/bin/ollama`. `OllamaManager` in `synapse/ollama_manager.py` manages its lifecycle (start/stop/health-check) and selects the best available model. GPU detection uses Vulkan (`vulkaninfo`) to prefer discrete AMD/NVIDIA GPUs. The Ollama HTTP API is at `http://127.0.0.1:11434` (overridable via `OLLAMA_HOST` env var).
### Frontend (`interface/web/`)
React 19 + Vite. No routing library — `App.jsx` manages page state in a single `currentPage` useState. All API calls hit `http://localhost:8000` (configured in `src/config.js`). Built to `dist/` (gitignored) via `npm run build` and served by the backend at `:8000` — the mount is in `synapse/main.py` (`_DIST` at `/`, guarded by `is_dir()`), so `dist/` must be built for the UI to appear. Pages: Chatbot, Playbook editor, Conversation History, Models, Memory, Settings, Logs.
### Code Tracks (`synapse/tools.py` + `synapse/code_run.py`)
Two separate tools, split by *where the code runs*:
- **`render_preview`** — validates markup and returns a fence the chat renders in
an opaque-origin `sandbox="allow-scripts"` iframe. Nothing executes
server-side. Languages: `PREVIEW_LANGS` in `synapse/tools.py`, mirrored by
`interface/web/src/preview/languages.js`.
- **`run_snippet`** — compiles and runs a single file on the host via
`synapse/code_run.py`, and returns a ```nexus-run fence carrying the source and
its captured output. Languages: `RUN_LANGS` in `synapse/code_run.py`, mirrored
by `interface/web/src/preview/run-langs.js`.
Each pair of registries is asserted equal by `tests/test_tools.py` — nothing
couples them at runtime, so drift fails the check gate instead of silently
degrading in the chat.
`run_snippet` is an **action tool**: `action_tool_policy` gates it (`off` by
default, `ask` = per-call Approve/Deny in chat). Read the `code_run.py` module
docstring before touching it — it runs code as the current user and is explicit
about which of its five layers are load-bearing and which are only a tripwire.
### Persistent Storage
Most data lands in `synapse/memory/memory.db` (SQLite, WAL mode). Tables: memory facts, conversations, messages, app settings. `synapse/memory/store.py` (`PersistentMemoryStore`) owns the schema and all queries. Playbooks are the exception — they live as YAML files in `data/playbooks/` (see Playbook System). `nexus_config.py` defines all paths; it also ensures all required directories exist on import.
### Logs & Runtime State
- `runtime/backend.log`, `runtime/frontend.log`, `runtime/memory.log` — service stdout
- `runtime/logs/ollama.log`, `runtime/logs/chat.log`
- `runtime/pids/backend.pid`, `runtime/pids/frontend.pid` — used by the management CLI
## Key Config
| Concern | Location |
|---|---|
| Ollama host | `OLLAMA_HOST` env var (default `http://127.0.0.1:11434`) |
| All filesystem paths | `synapse/nexus_config.py` `Settings` class |
| Frontend API base URL | `interface/web/src/config.js` |
| Default chat/memory models | `synapse/nexus_config.py` `DEFAULT_CHAT_MODEL` / `DEFAULT_MEMORY_MODEL` |
| Python dependencies (base) | `requirements-base.txt` |
| Python dependencies (AMD/ROCm) | `requirements-amd.txt` |
| Python dependencies (NVIDIA/CUDA) | `requirements-nvidia.txt` (generated by `bin/gen-nvidia-reqs.py`) |
| Python dependencies (Windows/CPU) | `requirements-windows.txt` |