forked from enderofwings/NexusOS
Reconciles 17 commits of this session's work (self-alteration tools, vendored Curry, slash-command dispatch, Windows toolchain/gate fixes) against origin/main's v1.2.0 sync (Projects/RAG scoping, a new modules/ system for mail and network, in-app updates, the standalone memory microservice folded into an in-process curator, KDE desktop theme overhaul). Nine real conflicts, each resolved by hand after reading both sides' actual diffs rather than picking one side wholesale: - synapse/tools.py, tests/test_tools.py: origin/main's diff here was small and clean (read_file/list_files, two new tests) despite git's diff3 flagging the whole file as one conflict blob -- reset to this branch's version and hand-spliced their addition in at the same points they used, rather than trying to reconcile a false 800-line conflict. Found and fixed a real bug while verifying: _list_files returned backslash-separated paths on Windows, which don't match the forward-slash glob patterns the tool's own schema documents. - synapse/main.py: kept this branch's cue-based standing advertisement of render_preview/run_snippet (independent of any playbook granting them) AND adopted origin/main's fix for routed reference playbooks not bringing their own tools along -- dropping either would have been a real regression, not just a style difference. Also: the standalone memory service (port 8001) is gone upstream, so its dead CORS/kill- target entries were removed; NEXUS_BACKEND_PORT parameterization and the manage_ollama-conditional kill logic (this branch's remote-Ollama support) were kept over origin/main's hardcoded equivalents. - synapse/memory/store.py: kept this branch's _delete_message_vectors helper (already reused elsewhere, batches to stay under SQLite's variable limit) over origin/main's inline duplicate of the same fix. - synapse/nexus_config.py, nexusos_cli/ncp.py: dropped the now-dead memory-service port/service entries; kept NEXUS_BACKEND_PORT env override and the manage_ollama-conditional kill-target list. - CLAUDE.md, README.md: merged both sides' additions, no real conflict. Found and fixed three more issues while independently verifying the merged tree, none of them mine or origin/main's alone -- only visible once both sides actually ran together: - modules/ (the new mail+network package) was never added to pyproject.toml's wheel `packages` list OR the sdist's `include` allowlist, so `from modules.registry import ROUTERS` in main.py would ImportError on any wheel install. Fixed both; bin/check.sh's packaging gate now asserts modules/ actually ships. tests/ test_packaging_deps.py's FIRST_PARTY/SHIPPED_PACKAGES sets were updated to recognize the new package. - tests/test_mail_creds.py's 0600-mode assertions are POSIX-only -- NTFS has no equivalent permission bits, so os.open(path, 0o600) on Windows just creates a normal file and stat.S_IMODE reports 0o666 regardless. Made the assertions platform-aware rather than skip real coverage (the temp-file-cleanup and password round-trip checks in the same test still run on Windows) or paper over a genuine OS limitation with a fake pass. - tests/test_kde_theme.py used bare Path.read_text() in fifteen places; Windows' default locale encoding (cp1252, not UTF-8) can't decode a real UTF-8 byte in the QML it reads, and did fail on one of the fifteen. Fixed all fifteen, not just the one that happened to trip today, since the other fourteen were equally fragile. Verified: full bin/check.sh reports OK end-to-end on this Windows checkout -- pytest (tests + management): 295 passed, 0 failed, 9 skipped; eslint clean; frontend node:test 57/57; PowerShell/shell parse clean; wheel + sdist pass twine check and now correctly carry modules/ (60 files, up from 52 pre-merge). synapse.main:app builds with 74 routes (up from 54 pre-merge, matching the new Projects/mail/ network endpoints).
690 lines
32 KiB
Python
690 lines
32 KiB
Python
"""Smoke test for the Synapse backend's core wiring.
|
|
|
|
Run from nexus-core/ with the Promethean venv active: pytest -q
|
|
Or the whole gate (tests + frontend lint): bin/check.sh
|
|
|
|
Deliberately tiny (ponytail): it guards the things v1 promises - app wiring,
|
|
model defaults, playbook ordering, persistence - without needing a running
|
|
Ollama service or network. Not a full suite.
|
|
"""
|
|
import subprocess
|
|
from pathlib import Path
|
|
|
|
from fastapi.testclient import TestClient
|
|
import pytest
|
|
|
|
from synapse.main import app
|
|
from synapse.memory.store import MemoryItem, PersistentMemoryStore
|
|
from synapse.nexus_config import DEFAULT_CHAT_MODEL, DEFAULT_MEMORY_MODEL
|
|
from synapse.ollama_manager import OllamaManager
|
|
from synapse import ollama_manager
|
|
from synapse.icons.compositor import _is_allowed_path
|
|
from synapse import playbook_manager
|
|
from synapse.playbooks.store import PlaybookFileStore, PlaybookItem
|
|
|
|
REPO_ROOT = Path(__file__).resolve().parent.parent
|
|
|
|
|
|
def test_app_builds_and_status_responds():
|
|
# No `with` → startup event does not run; `ollama` stays the module-level
|
|
# None and /status handles it. Pure liveness check. Not `/` — since the
|
|
# single-process change that path serves the built web UI, not JSON.
|
|
resp = TestClient(app).get("/status")
|
|
assert resp.status_code == 200
|
|
assert resp.json()["status"] == "online"
|
|
|
|
|
|
def test_keep_alive_pins_the_model():
|
|
# Default keeps the model resident so chats skip cold reloads...
|
|
assert PersistentMemoryStore._SETTINGS_DEFAULTS["keep_alive"] == "30m"
|
|
mgr = OllamaManager()
|
|
assert mgr._apply_keep_alive({"model": "x"}) == {"model": "x", "keep_alive": "30m"}
|
|
# ...and an empty value omits the field (falls back to Ollama's default).
|
|
mgr.keep_alive = ""
|
|
assert "keep_alive" not in mgr._apply_keep_alive({"model": "x"})
|
|
|
|
|
|
def test_auto_model_remap(monkeypatch):
|
|
import asyncio
|
|
from synapse import main
|
|
cfg = {"model": "", "auto_chat_model": "chatX", "auto_code_model": "coderY"}
|
|
monkeypatch.setattr(main.store, "get_settings", lambda: cfg)
|
|
# code intent ("function") routes to the code remap; chat intent to the chat remap
|
|
assert asyncio.run(main._auto_select_model("write a function to sort a list")) == "coderY"
|
|
assert asyncio.run(main._auto_select_model("how are you today")) == "chatX"
|
|
# an explicit pin beats the remap
|
|
cfg["model"] = "pinnedZ"
|
|
assert asyncio.run(main._auto_select_model("debug this code")) == "pinnedZ"
|
|
|
|
|
|
def test_hardware_fit_logic():
|
|
from synapse import hardware
|
|
assert hardware._fit(2.5, 4.0, 16.0) == "gpu" # 2.5+1 <= 4 -> fits GPU
|
|
assert hardware._fit(4.7, 4.0, 16.0) == "ram" # too big for 4GB GPU, fits RAM
|
|
assert hardware._fit(4.7, None, 16.0) == "ram" # VRAM unknown -> RAM
|
|
assert hardware._fit(40.0, 4.0, 16.0) == "no" # too big everywhere
|
|
rec = hardware.recommend()
|
|
assert "hardware" in rec and all("fit" in m for m in rec["models"])
|
|
|
|
|
|
def test_stt_status_endpoint():
|
|
# Reports whether local Whisper is installed; wiring must respond either way.
|
|
resp = TestClient(app).get("/stt/status")
|
|
assert resp.status_code == 200
|
|
assert isinstance(resp.json()["available"], bool)
|
|
|
|
|
|
def test_num_ctx_option_only_when_positive():
|
|
# 0 / None -> omit num_ctx so Ollama uses the model default; positive -> set it.
|
|
from synapse.ollama_manager import _chat_options
|
|
assert "num_ctx" not in _chat_options(None, None, 0)
|
|
assert "num_ctx" not in _chat_options(None, None, None)
|
|
assert _chat_options(None, None, 8192)["num_ctx"] == 8192
|
|
|
|
|
|
def test_default_models_have_one_source_of_truth():
|
|
# The curator default is the config constant, not a copy of it.
|
|
assert PersistentMemoryStore._SETTINGS_DEFAULTS["memory_model"] == DEFAULT_MEMORY_MODEL
|
|
# And the Windows installer reads the constants instead of hardcoding a tag,
|
|
# which is what used to let installer and backend drift onto different models.
|
|
ps1 = (REPO_ROOT / "install-windows.ps1").read_text(encoding="utf-8")
|
|
assert "DEFAULT_CHAT_MODEL" in ps1
|
|
for hardcoded in (DEFAULT_CHAT_MODEL, DEFAULT_MEMORY_MODEL, "qwen", "llama3.1"):
|
|
assert hardcoded not in ps1, f"install-windows.ps1 hardcodes {hardcoded!r}"
|
|
|
|
|
|
def test_windows_scripts_stay_ascii():
|
|
# PowerShell 5.1 decodes BOM-less files as ANSI: one stray Unicode dash
|
|
# eats a quote and the whole script dies at parse time. cmd.exe is worse
|
|
# still - it decodes by the console codepage. Globbed rather than listed by
|
|
# name so a newly added script is covered without editing this test.
|
|
skip = {"Promethean", "node_modules", ".git", "dist"}
|
|
ps1s = [p for pat in ("*.ps1", "*.cmd") for p in REPO_ROOT.rglob(pat)
|
|
if not skip & set(p.parts)]
|
|
assert ps1s, "no .ps1/.cmd files found - did the Windows path move?"
|
|
for path in ps1s:
|
|
raw = path.read_bytes()
|
|
bad = [(i, b) for i, b in enumerate(raw) if b > 0x7F]
|
|
assert not bad, f"{path.relative_to(REPO_ROOT)} has non-ASCII bytes at {bad[:3]}"
|
|
|
|
|
|
def test_ncp_is_registered_on_path_not_in_a_shell_profile():
|
|
"""A `function ncp` in a shell profile is invisible to cron, .desktop Exec
|
|
lines, cmd.exe and Task Scheduler - and on Windows the default Restricted
|
|
execution policy blocks the profile outright. Both installers must put ncp
|
|
on PATH; the profile wiring only survives as a no-sudo fallback."""
|
|
linux = (REPO_ROOT / "bin" / "restore-linux.sh").read_text(encoding="utf-8")
|
|
runtime = linux.split('if [ "$stage" = "runtime" ]; then')[1].split("\n exit 0\nfi")[0]
|
|
assert "/usr/local/bin/ncp" in runtime
|
|
|
|
ps1 = (REPO_ROOT / "install-windows.ps1").read_text(encoding="utf-8")
|
|
assert "ncp.cmd" in ps1, "installer must register the .cmd shim, not a profile function"
|
|
# setx truncates PATH at 1024 characters and has permanently broken machines.
|
|
# Comments stripped so the code that explains the ban does not trip it.
|
|
code = "\n".join(ln for ln in ps1.splitlines() if not ln.strip().startswith("#"))
|
|
assert "setx" not in code.lower()
|
|
assert 'SetEnvironmentVariable("Path"' in code
|
|
|
|
# The model pull is the only multi-GB step and the only one the script
|
|
# invites a Ctrl+C on -- which in PS 5.1 kills the whole script. Anything
|
|
# after it is lost, so it has to come last. It used to sit in the middle,
|
|
# and skipping the download silently skipped the ncp registration too.
|
|
assert ps1.index("ollama pull") > ps1.index("Registering the ncp command"), \
|
|
"model pull must come after ncp registration - a Ctrl+C there aborts the installer"
|
|
assert ps1.index("ollama pull") > ps1.index("Creating desktop shortcut"), \
|
|
"model pull must come after the desktop shortcut"
|
|
|
|
# And it must not TELL anyone to press Ctrl+C: in PS 5.1 that kills the
|
|
# script, so the advertised way to skip the download was also the way to
|
|
# abort the install. Skipping is a prompt now. Comments stripped so the
|
|
# comment explaining this does not trip the check.
|
|
assert "Ctrl+C" not in code, "installer must not offer Ctrl+C as a skip"
|
|
|
|
|
|
def test_no_ps1_shadows_the_ncp_path_shim():
|
|
"""PowerShell resolves ExternalScript (.ps1) ahead of Application (.cmd), so
|
|
an ncp.ps1 sitting next to ncp.cmd wins in PowerShell and drags the execution
|
|
policy back in - the exact thing the .cmd exists to avoid. Observed on the
|
|
Windows VM: `Get-Command ncp -All` listed ncp.ps1 first, from the same
|
|
directory the installer had just put on PATH."""
|
|
shim = REPO_ROOT / "management" / "ncp.cmd"
|
|
assert shim.exists(), "the Windows PATH shim is missing"
|
|
twin = shim.with_suffix(".ps1")
|
|
assert not twin.exists(), f"{twin.name} shadows {shim.name} in PowerShell"
|
|
|
|
|
|
|
|
def test_first_playbook_is_the_system_prompt(tmp_path, monkeypatch):
|
|
store = PlaybookFileStore(tmp_path)
|
|
store.add_playbook(PlaybookItem(id="ctx", title="Reference", goal="ref goal",
|
|
instructions="ref instructions", order=1))
|
|
store.add_playbook(PlaybookItem(id="main", title="Main", goal="Be useful.",
|
|
instructions="Answer briefly.", order=0))
|
|
monkeypatch.setattr("synapse.playbook_manager.playbook_store", store)
|
|
|
|
assert playbook_manager.get_main_playbook().id == "main"
|
|
assert [p.id for p in playbook_manager.get_context_playbooks()] == ["ctx"]
|
|
assert playbook_manager.get_system_prompt() == "Be useful.\n\nAnswer briefly."
|
|
|
|
|
|
def test_settings_round_trip_over_defaults(tmp_path):
|
|
store = PersistentMemoryStore(tmp_path / "memory.db")
|
|
store.update_settings({"model": "some-model:8b"})
|
|
settings = store.get_settings()
|
|
assert settings["model"] == "some-model:8b" # written value wins
|
|
assert settings["memory_model"] == DEFAULT_MEMORY_MODEL # untouched keys still default
|
|
|
|
|
|
def test_past_conversations_are_searchable(tmp_path):
|
|
# The chat system prompt is built from these snippets, so a broken search
|
|
# silently drops the assistant's recall of past chats.
|
|
store = PersistentMemoryStore(tmp_path / "memory.db")
|
|
store.create_conversation("c1")
|
|
store.add_message("c1", "user", "how do I mount the backup drive?")
|
|
store.add_message("c1", "assistant", "use rsync over ssh")
|
|
|
|
assert store.search_conversations("BACKUP drive") # case-insensitive substring
|
|
assert store.search_conversations("nothing here") == []
|
|
assert store.search_conversations(" ") == []
|
|
|
|
|
|
def test_memory_store_reads_changes_from_another_instance(tmp_path):
|
|
db_path = tmp_path / "memory.db"
|
|
writer = PersistentMemoryStore(db_path)
|
|
reader = PersistentMemoryStore(db_path)
|
|
|
|
writer.add(MemoryItem(id="fresh", text="written by another process"))
|
|
assert reader.get("fresh").text == "written by another process"
|
|
|
|
|
|
def test_icon_source_requires_real_allowed_file_boundary(tmp_path):
|
|
allowed = tmp_path / "icons"
|
|
allowed.mkdir()
|
|
source = allowed / "app.svg"
|
|
source.write_text("<svg />")
|
|
sibling = allowed.parent / "icons-other"
|
|
sibling.mkdir()
|
|
(sibling / "app.svg").write_text("<svg />")
|
|
old_roots = list(__import__("synapse.icons.compositor", fromlist=["_ALLOWED_ROOTS"])._ALLOWED_ROOTS)
|
|
module = __import__("synapse.icons.compositor", fromlist=["_ALLOWED_ROOTS"])
|
|
module._ALLOWED_ROOTS[:] = [str(allowed)]
|
|
try:
|
|
assert _is_allowed_path(str(source))
|
|
assert not _is_allowed_path(str(sibling / "app.svg"))
|
|
assert not _is_allowed_path(str(allowed / "missing.svg"))
|
|
finally:
|
|
module._ALLOWED_ROOTS[:] = old_roots
|
|
|
|
|
|
def test_icons_image_endpoint_requires_real_allowed_root_boundary(tmp_path):
|
|
# Sibling directories that merely share a string prefix with an allowed
|
|
# root (e.g. "icons-other" vs "icons") must not pass the check.
|
|
from synapse import main
|
|
|
|
allowed = tmp_path / "icons"
|
|
allowed.mkdir()
|
|
source = allowed / "app.svg"
|
|
source.write_text("<svg />")
|
|
sibling = tmp_path / "icons-other"
|
|
sibling.mkdir()
|
|
evil = sibling / "app.svg"
|
|
evil.write_text("<svg />")
|
|
|
|
old_roots = list(main._ALLOWED_ICON_ROOTS)
|
|
main._ALLOWED_ICON_ROOTS[:] = [str(allowed)]
|
|
try:
|
|
client = TestClient(app)
|
|
ok = client.get("/icons/image", params={"path": str(source)})
|
|
assert ok.status_code == 200
|
|
|
|
blocked = client.get("/icons/image", params={"path": str(evil)})
|
|
assert blocked.status_code == 403
|
|
|
|
missing = client.get("/icons/image", params={"path": str(allowed / "missing.svg")})
|
|
assert missing.status_code in (403, 404)
|
|
finally:
|
|
main._ALLOWED_ICON_ROOTS[:] = old_roots
|
|
|
|
|
|
def test_ollama_stream_propagates_transport_errors(monkeypatch):
|
|
"""A failing stream must surface, not be swallowed into an empty reply —
|
|
and it must carry Ollama's own explanation, since that is the only part the
|
|
user can act on. The response here is a real httpx.Response because the
|
|
error path reads the body, which a stubbed raise_for_status never exercised."""
|
|
import httpx
|
|
|
|
class FailingResponse:
|
|
async def __aenter__(self):
|
|
return httpx.Response(
|
|
503,
|
|
json={"error": "Ollama unavailable"},
|
|
request=httpx.Request("POST", "http://127.0.0.1:11434/api/chat"),
|
|
)
|
|
|
|
async def __aexit__(self, *args):
|
|
return False
|
|
|
|
class FailingClient:
|
|
def __init__(self, **kwargs):
|
|
pass
|
|
|
|
async def __aenter__(self):
|
|
return self
|
|
|
|
async def __aexit__(self, *args):
|
|
return False
|
|
|
|
def stream(self, *args, **kwargs):
|
|
return FailingResponse()
|
|
|
|
monkeypatch.setattr(ollama_manager.httpx, "AsyncClient", FailingClient)
|
|
|
|
async def consume():
|
|
async for _ in OllamaManager()._chat_stream([], "model", 0):
|
|
pass
|
|
|
|
with pytest.raises(httpx.HTTPStatusError, match="Ollama unavailable"):
|
|
import asyncio
|
|
asyncio.run(consume())
|
|
|
|
|
|
def _load_sync():
|
|
"""bin/ isn't a package - load the cross-platform sync core by path."""
|
|
import importlib.util
|
|
spec = importlib.util.spec_from_file_location("sync", REPO_ROOT / "bin" / "sync.py")
|
|
module = importlib.util.module_from_spec(spec)
|
|
spec.loader.exec_module(module)
|
|
return module
|
|
|
|
|
|
def test_sync_compare_detects_direction(tmp_path):
|
|
"""The guard that stops a stale box from overwriting the other's chats.
|
|
Both backup and restore refuse to run when this says the wrong thing, so a
|
|
silent break here loses conversation history."""
|
|
import contextlib
|
|
import sqlite3 as sq
|
|
sync = _load_sync()
|
|
db, dump = tmp_path / "memory.db", tmp_path / "memory.db.sql"
|
|
|
|
def write(rows):
|
|
db.unlink(missing_ok=True)
|
|
# closing() then the connection itself: sqlite3's own context manager
|
|
# commits but never closes, and Windows refuses to unlink a file that
|
|
# still has an open handle.
|
|
with contextlib.closing(sq.connect(db)) as conn, conn:
|
|
# updated_at REAL, matching the production schema in store.py. A TEXT
|
|
# column here hid a real TypeError for months: the comparison in
|
|
# _extra() ran str-vs-str in the test and str-vs-float in the field.
|
|
conn.executescript(
|
|
"create table conversations (id text primary key, updated_at real not null);"
|
|
"create table memory (id text primary key);"
|
|
)
|
|
conn.executemany("insert into conversations values (?, ?)", rows)
|
|
|
|
assert sync.compare(db, dump) == "no-dump"
|
|
write([("a", 1778553309.5)])
|
|
assert sync.compare(db, dump) == "no-dump"
|
|
|
|
# Dump matches the live DB exactly.
|
|
sync.DB, sync.DB_SQL = db, dump
|
|
assert sync.dump_db()
|
|
assert sync.compare(db, dump) == "same"
|
|
|
|
# A newer message bumps updated_at -> this box is ahead of the backup.
|
|
write([("a", 1778553999.5)])
|
|
assert sync.compare(db, dump) == "local-ahead"
|
|
|
|
# The backup holds a conversation this box never saw.
|
|
write([])
|
|
assert sync.compare(db, dump) == "local-behind"
|
|
|
|
# Each side has something the other lacks.
|
|
write([("b", 1778553309.5)])
|
|
assert sync.compare(db, dump) == "diverged"
|
|
|
|
# A dump that won't replay is its own verdict, not a fake divergence -
|
|
# reporting "diverged" there blocked backup AND restore with what looked
|
|
# like a legitimate answer.
|
|
dump.write_text("INSERT INTO nope VALUES (1);\n")
|
|
assert sync.compare(db, dump) == "unreadable"
|
|
|
|
db.unlink()
|
|
assert sync.compare(db, dump) == "no-live"
|
|
|
|
|
|
def test_restore_stages_match_the_script():
|
|
"""sync.py invokes restore-linux.sh stages by name with check=False, so a
|
|
rename on one side alone fails silently - and the runtime stage is what
|
|
installs Ollama and the ncp alias. Keep the two in agreement."""
|
|
import re
|
|
called = set(re.findall(
|
|
r'linux_stage\(\s*"restore-linux\.sh"\s*,\s*"(\w+)"',
|
|
(REPO_ROOT / "bin" / "sync.py").read_text()))
|
|
script = (REPO_ROOT / "bin" / "restore-linux.sh").read_text()
|
|
handled = set(re.search(r"case \"\$stage\" in\s*\n\s*([\w|]+)\)", script).group(1).split("|"))
|
|
assert called, "no restore-linux.sh stages found in sync.py"
|
|
assert called <= handled, f"sync.py calls unhandled stage(s): {called - handled}"
|
|
|
|
|
|
def test_desktop_stage_is_the_only_one_touching_home():
|
|
"""--no-desktop is only a real safety valve if the $HOME writes all live in
|
|
the desktop stage. A test clone runs prep + runtime unconditionally."""
|
|
script = (REPO_ROOT / "bin" / "restore-linux.sh").read_text()
|
|
runtime = script.split('if [ "$stage" = "runtime" ]; then')[1].split("\n exit 0\nfi")[0]
|
|
# Shell wiring is the one deliberate exception: registering ncp in ~/.bashrc
|
|
# and in the PowerShell profile is how you launch NexusOS at all, and each is
|
|
# a grep-guarded no-op on re-run. Desktop config (xfconf, plank, themes,
|
|
# os-release) must stay in the desktop stage so --no-desktop really is safe.
|
|
shell_wiring = (".bashrc",)
|
|
home_writes = [ln for ln in runtime.splitlines()
|
|
if "$HOME" in ln and not any(w in ln for w in shell_wiring)]
|
|
assert not home_writes, f"runtime stage writes to $HOME: {home_writes}"
|
|
|
|
|
|
def test_genmon_configs_are_written_with_the_panel_down():
|
|
"""genmon holds its config in memory and rewrites genmon-N.rc when the panel
|
|
exits, so copying the rc files while the panel is running gets silently
|
|
undone - every applet then loads blank. The copy has to sit between the panel
|
|
quit and the relaunch."""
|
|
script = (REPO_ROOT / "bin" / "panel" / "install.sh").read_text()
|
|
quit_at = script.index("xfce4-panel -q")
|
|
copy_at = script.index('cp -f "$NEXUS/management/panel/genmon-$id.rc"')
|
|
start_at = script.index("setsid xfce4-panel")
|
|
assert quit_at < copy_at < start_at, "genmon rc copy must happen with the panel stopped"
|
|
|
|
|
|
def test_plank_is_actually_launched():
|
|
"""Restoring ~/.config/plank only brings back the dock's launchers - nothing
|
|
in it starts Plank. The primary-follow watcher is what launches and revives
|
|
it, so it needs an autostart entry or a fresh box has no dock at all."""
|
|
desktop = REPO_ROOT / "management" / "autostart" / "plank.desktop"
|
|
assert "plank-primary-watch.sh" in desktop.read_text()
|
|
assert "plank.desktop" in (REPO_ROOT / "bin" / "panel" / "install.sh").read_text()
|
|
|
|
|
|
_VULKANINFO_IGPU_AND_LLVMPIPE = """\
|
|
Devices:
|
|
========
|
|
GPU0:
|
|
\tvendorID = 0x8086
|
|
\tdeviceType = PHYSICAL_DEVICE_TYPE_INTEGRATED_GPU
|
|
\tdeviceName = Intel(R) Graphics (RPL-S)
|
|
GPU1:
|
|
\tvendorID = 0x10005
|
|
\tdeviceType = PHYSICAL_DEVICE_TYPE_CPU
|
|
\tdeviceName = llvmpipe (LLVM 20.1.2, 256 bits)
|
|
"""
|
|
|
|
|
|
def test_software_rasterizer_is_never_picked_as_a_gpu(monkeypatch):
|
|
"""Mesa always advertises an llvmpipe device with deviceType CPU. It used to
|
|
outscore an integrated GPU (neither DISCRETE nor INTEGRATED scored higher
|
|
than INTEGRATED), so Ollama got pinned to a software rasterizer - CPU
|
|
inference with Vulkan overhead stacked on top, reported as a 31 GiB
|
|
'discrete' GPU. The iGPU has to win, and a box with nothing but rasterizers
|
|
has to report no Vulkan device at all."""
|
|
def fake_run(cmd, **kwargs):
|
|
return subprocess.CompletedProcess(cmd, 0, _VULKANINFO_IGPU_AND_LLVMPIPE, "")
|
|
|
|
monkeypatch.setattr(ollama_manager.subprocess, "run", fake_run)
|
|
idx, name = ollama_manager._best_vulkan_device()
|
|
assert idx == 0 and "Intel" in name, f"picked {name!r} over the iGPU"
|
|
|
|
monkeypatch.setattr(
|
|
ollama_manager.subprocess, "run",
|
|
lambda cmd, **kw: subprocess.CompletedProcess(cmd, 0, "Devices:\nGPU0:\n"
|
|
"\tdeviceType = PHYSICAL_DEVICE_TYPE_CPU\n"
|
|
"\tdeviceName = llvmpipe\n", ""))
|
|
assert ollama_manager._best_vulkan_device()[0] == -1, "rasterizer-only box must report no GPU"
|
|
|
|
|
|
def test_installed_model_lookup_normalizes_the_latest_tag():
|
|
"""Ollama resolves a bare name to ":latest", so a catalog entry written
|
|
untagged ("nomic-embed-text") never matched the installed name
|
|
("nomic-embed-text:latest") and the Required gate stayed locked forever - the
|
|
model pulls fine, the UI just never sees it. Every lookup has to go through
|
|
withTag()."""
|
|
jsx = (REPO_ROOT / "interface" / "web" / "src" / "Models.jsx").read_text()
|
|
assert 'const withTag =' in jsx
|
|
assert "installedNames.has(m.name.toLowerCase())" not in jsx, \
|
|
"raw name lookup is back; untagged catalog entries will read as missing"
|
|
|
|
|
|
def test_dump_round_trips_a_db_holding_vec_tables(tmp_path):
|
|
"""The memory dump is the ONLY backup of conversations, facts and history.
|
|
iterdump() serializes a sqlite_vec virtual table as a raw
|
|
INSERT INTO sqlite_master(...) plus inserts into a table the replaying
|
|
connection cannot see, so restoring died on "no such table: vec_messages"
|
|
and left zero tables - the entire backup was unrecoverable. Build a DB
|
|
shaped like production (vec table included) and prove the dump replays."""
|
|
import sqlite3 as sq
|
|
sqlite_vec = pytest.importorskip("sqlite_vec", reason="vec index is optional")
|
|
sync = _load_sync()
|
|
if sync._vec0_extension() is None:
|
|
# dump_db() resolves vec0 relative to the repo's own venv, so a clone
|
|
# whose Promethean isn't built yet cannot dump a DB holding vec tables.
|
|
pytest.skip("sqlite_vec extension not present under this repo's venv")
|
|
db, dump = tmp_path / "memory.db", tmp_path / "memory.db.sql"
|
|
|
|
conn = sq.connect(db)
|
|
conn.enable_load_extension(True)
|
|
sqlite_vec.load(conn)
|
|
conn.enable_load_extension(False)
|
|
conn.executescript(
|
|
"create table conversations (id text primary key, updated_at real not null);"
|
|
"create table memory (id text primary key);"
|
|
"create table messages (id integer primary key, body text);"
|
|
"create virtual table vec_messages using vec0(embedding float[3] distance_metric=cosine);"
|
|
)
|
|
conn.execute("insert into conversations values ('c1', 1778553309.5)")
|
|
conn.execute("insert into memory values ('m1')")
|
|
# A message whose text mentions the filtered table names: a filter applied to
|
|
# the dump TEXT instead of the statement stream would eat this row.
|
|
conn.execute("insert into messages values (1, 'debugging vec_messages and vec_documents')")
|
|
conn.execute("insert into vec_messages(rowid, embedding) values (1, ?)",
|
|
(sqlite_vec.serialize_float32([0.1, 0.2, 0.3]),))
|
|
conn.commit()
|
|
conn.close()
|
|
|
|
sync.DB, sync.DB_SQL = db, dump
|
|
assert sync.dump_db()
|
|
|
|
restored = tmp_path / "restored.db"
|
|
with sq.connect(restored) as out:
|
|
out.executescript(dump.read_text(encoding="utf-8")) # must not raise
|
|
got = sq.connect(restored)
|
|
assert got.execute("select count(*) from conversations").fetchone()[0] == 1
|
|
assert got.execute("select count(*) from memory").fetchone()[0] == 1
|
|
assert got.execute("select body from messages").fetchone()[0] == \
|
|
"debugging vec_messages and vec_documents"
|
|
# The vec index is derived - absent from the dump, rebuilt by the backfill.
|
|
assert "vec_messages" not in {r[0] for r in got.execute(
|
|
"select name from sqlite_master where type='table'")}
|
|
|
|
|
|
def test_curator_drops_fabricated_facts():
|
|
"""The curator model invents two classes of fact no prompt wording stopped
|
|
(verified against mistral:7b), and both reached the real memory DB: absence
|
|
claims read off the existing-memory block ("the user does not have any pets",
|
|
which contradicted four cats on file) and specifics lifted from the
|
|
ASSISTANT's reply ("the user's main development machine is a MacBook Pro", from
|
|
the user message "What am I developing on?"). Deterministic guard, so it
|
|
holds whatever the model does."""
|
|
from synapse.memory.extractor import _reject_reason
|
|
|
|
# Absence claims are never facts.
|
|
assert _reject_reason("the user does not have any pets", "do i have any pets?")
|
|
assert _reject_reason("the user's favorite episode is unknown", "what's my favorite episode?")
|
|
assert _reject_reason("the user has not specified an interest", "tell me about stargate")
|
|
|
|
# Specifics the user never typed came from the assistant.
|
|
assert _reject_reason("the user's main dev machine is a MacBook Pro", "What am I developing on?")
|
|
|
|
# ...but the same shape grounded in the user's own words must survive.
|
|
assert _reject_reason(
|
|
"the user owns a 2000 Ford Ranger with a 3.0L V6",
|
|
"i also have a 2000 Ford Ranger, it's a five-speed with a 3.0L V6") is None
|
|
assert _reject_reason(
|
|
"the user has a beagle named Biscuit",
|
|
"i just adopted a dog named Biscuit, he's a beagle") is None
|
|
# A fact carrying no proper nouns or numbers can't be grounding-checked;
|
|
# the prompt owns that case, so the guard must let it through.
|
|
assert _reject_reason(
|
|
"the user prefers short answers over long explanations",
|
|
"i really prefer short answers over long explanations") is None
|
|
|
|
|
|
def test_ollama_failures_surface_the_reason_not_just_the_status():
|
|
"""Ollama answers every failure with {"error": "..."} and httpx's default
|
|
message throws it away. A user hitting a retired cloud model saw
|
|
"Client error '410 Gone' for url ..." when the body said exactly why."""
|
|
import asyncio
|
|
import httpx
|
|
import pytest
|
|
from synapse.ollama_manager import _raise_for_ollama
|
|
|
|
req = httpx.Request("POST", "http://127.0.0.1:11434/api/chat")
|
|
|
|
retired = httpx.Response(410, json={"error": "glm-4.6 was retired at 2026-06-16"}, request=req)
|
|
with pytest.raises(httpx.HTTPStatusError) as ei:
|
|
asyncio.run(_raise_for_ollama(retired))
|
|
assert "retired" in str(ei.value) and "410" in str(ei.value)
|
|
|
|
# The common case, not just the exotic one.
|
|
missing = httpx.Response(404, json={"error": "model 'foo' not found"}, request=req)
|
|
with pytest.raises(httpx.HTTPStatusError) as ei:
|
|
asyncio.run(_raise_for_ollama(missing))
|
|
assert "model 'foo' not found" in str(ei.value)
|
|
|
|
# No usable body -> keep httpx's own wording rather than inventing one.
|
|
blank = httpx.Response(500, content=b"", request=req)
|
|
with pytest.raises(httpx.HTTPStatusError):
|
|
asyncio.run(_raise_for_ollama(blank))
|
|
|
|
# Success stays silent.
|
|
asyncio.run(_raise_for_ollama(httpx.Response(200, json={"ok": True}, request=req)))
|
|
|
|
|
|
def test_ollama_host_is_normalized_for_clients_but_not_for_binding():
|
|
"""OLLAMA_HOST is Ollama's *bind* variable, and `OLLAMA_HOST=0.0.0.0:11434`
|
|
is the normal way to expose it on a LAN. Used verbatim as a client base URL
|
|
it is unusable — no scheme, and 0.0.0.0 is not a destination — and every
|
|
request failed into list_models()'s bare `except: return []`, so the model
|
|
picker just went empty with no error anywhere."""
|
|
from synapse.nexus_config import _normalize_ollama_host as norm
|
|
|
|
assert norm("0.0.0.0:11434") == "http://127.0.0.1:11434"
|
|
assert norm("[::]:11434") == "http://127.0.0.1:11434"
|
|
assert norm("http://0.0.0.0:11434/") == "http://127.0.0.1:11434"
|
|
assert norm("127.0.0.1:11434") == "http://127.0.0.1:11434" # scheme supplied
|
|
assert norm("") == "http://127.0.0.1:11434"
|
|
# A real remote is deliberate — leave it alone.
|
|
assert norm("https://ollama.lan:11434") == "https://ollama.lan:11434"
|
|
assert norm("192.168.1.50:11434") == "http://192.168.1.50:11434"
|
|
|
|
|
|
def test_spawned_serve_keeps_the_users_bind_address(monkeypatch):
|
|
"""Normalizing for the client must not quietly un-expose a server we spawn."""
|
|
import importlib
|
|
from synapse import nexus_config
|
|
monkeypatch.setenv("OLLAMA_HOST", "0.0.0.0:11434")
|
|
reloaded = importlib.reload(nexus_config)
|
|
try:
|
|
assert reloaded.settings.ollama_bind == "0.0.0.0:11434" # listens everywhere
|
|
assert reloaded.settings.ollama_host == "http://127.0.0.1:11434" # we connect here
|
|
finally:
|
|
monkeypatch.delenv("OLLAMA_HOST", raising=False)
|
|
importlib.reload(nexus_config)
|
|
|
|
|
|
def test_think_blocks_never_reach_the_reply():
|
|
"""A reasoning model emits <think> inline in `content` even with think off,
|
|
and the whole internal monologue reached the chat window — including the
|
|
literal closing tags. Streaming has to cope with a tag split across chunks,
|
|
and with the shape actually observed: a stray </think> and no opening."""
|
|
from synapse.ollama_manager import ThinkStripper, strip_think
|
|
|
|
def stream(chunks):
|
|
s = ThinkStripper()
|
|
return "".join(s.feed(c) for c in chunks) + s.flush()
|
|
|
|
assert stream(["hello ", "<think>", "noise", "</think>", "world"]) == "hello world"
|
|
# tag split across chunk boundaries
|
|
assert stream(["a<th", "ink>x</thi", "nk>b"]) == "ab"
|
|
# never closed -> it was all reasoning
|
|
assert stream(["keep", "<think>", "runs off the end"]) == "keep"
|
|
# stray close, no open: at minimum the tag itself must not be shown
|
|
assert "</think>" not in stream(["reasoning...", "</think>", "the answer"])
|
|
# ordinary text is untouched, including angle brackets
|
|
assert stream(["a < b ", "and c > d"]) == "a < b and c > d"
|
|
|
|
# With the complete message the stray-close case can be handled properly:
|
|
# everything before it was reasoning.
|
|
assert strip_think("rambling\n</think>\nThe answer") == "The answer"
|
|
assert strip_think("a<think>b</think>c") == "ac"
|
|
assert strip_think("no tags here") == "no tags here"
|
|
|
|
|
|
def test_render_hint_is_only_added_when_the_tool_is_offered():
|
|
"""The hint and the tool must share one condition. Unconditional, it turned
|
|
up recited as fact inside an answer about LRU caches."""
|
|
from synapse import tools as t
|
|
assert t.wants_render_preview("draw me a chart")
|
|
assert not t.wants_render_preview("Summarize what a thread-safe LRU cache needs")
|
|
|
|
|
|
def test_preview_iframe_cannot_navigate_to_a_network_url():
|
|
"""The child CSP blocks resource loads; the parent CSP must separately
|
|
block a sandboxed frame from navigating its own browsing context."""
|
|
index = (REPO_ROOT / "interface" / "web" / "index.html").read_text(encoding="utf-8")
|
|
markdown = (REPO_ROOT / "interface" / "web" / "src" / "Markdown.jsx").read_text(
|
|
encoding="utf-8"
|
|
)
|
|
assert "frame-src data:" in index
|
|
assert 'sandbox="allow-scripts"' in markdown
|
|
assert "encodeURIComponent(doc)" in markdown
|
|
assert "src={frameUrl}" in markdown
|
|
assert "srcDoc={doc}" not in markdown
|
|
|
|
|
|
def test_update_check_reports_behind_and_survives_git_failure(monkeypatch):
|
|
from synapse import main
|
|
# Fake git so the test never touches the network. Behind → the remote
|
|
# VERSION file, not this checkout's, is what the UI advertises.
|
|
calls = {
|
|
("rev-list", "--count", "HEAD..origin/main"): "3",
|
|
("show", "origin/main:VERSION"): "9.9.9\n",
|
|
("log", "-1", "--format=%h %s", "origin/main"): "abc1234 feat: thing",
|
|
}
|
|
monkeypatch.setattr(main, "_git", lambda *a, **kw: calls.get(a, ""))
|
|
body = TestClient(app).get("/update/check").json()
|
|
assert body["behind"] == 3 and body["remote_version"] == "9.9.9"
|
|
|
|
# An unreachable remote must not 500 the sidebar.
|
|
def boom(*a, **kw):
|
|
raise RuntimeError("could not resolve host")
|
|
monkeypatch.setattr(main, "_git", boom)
|
|
body = TestClient(app).get("/update/check").json()
|
|
assert body["behind"] == 0 and "could not resolve host" in body["error"]
|
|
|
|
|
|
def test_update_apply_spawns_detached_and_refuses_a_second_run(monkeypatch):
|
|
import subprocess
|
|
from synapse import main
|
|
seen = {}
|
|
|
|
def fake_popen(argv, **kw):
|
|
seen["argv"], seen["kw"] = argv, kw
|
|
return object()
|
|
|
|
monkeypatch.setattr(main, "_update_running", False)
|
|
monkeypatch.setattr(subprocess, "Popen", fake_popen)
|
|
client = TestClient(app)
|
|
assert client.post("/update/apply").json()["started"] is True
|
|
assert seen["argv"][-2:] == [str(REPO_ROOT / "management" / "ncp.py"), "upgrade"]
|
|
# Detached, or `ncp upgrade` dies with the backend it is about to stop.
|
|
assert seen["kw"].get("start_new_session") or seen["kw"].get("creationflags")
|
|
|
|
# Double-click must not launch a second pull/rebuild over the first.
|
|
assert client.post("/update/apply").json()["started"] is False
|