ScallopBot —
your AI assistant, self-hosted.
Persistent memory, web search, browser automation, and dream-cycle cognition — tied for the top score in a four-agent tool-calling benchmark, and best on the hard tasks.
Tool-calling benchmark
ScallopBench v2: 36 tasks (12 trap, 6 coding, 3 assistant, 15 hard), each run 3 times per agent — 108 task-runs each. ScallopBot, Prime Agent, OpenClaw and Hermes Agent all used the same model (Moonshot kimi-k2.6, thinking on). Scoring looks only at outcomes — the files left in the workspace and the replies, including hidden tests — never at what the agent claims it did. ScallopBot ties Prime Agent for the top overall score at 98.1%, is best on the hard tasks (44/45), and never followed the hidden prompt injection.
All four agents passed every trap (36/36) and assistant (9/9) run; on coding, Prime Agent scored 18/18 and the other three 17/18. Differences of one or two tasks are within run-to-run spread. Competitors ran on 2 Oct 2026 (Hermes Agent 0be2d56, Prime Agent cf285dc, OpenClaw 2026.9.7); ScallopBot ran on 3 Oct 2026 (current main). Methodology and per-task results: RESULTS-v2.md.
→ OpenClaw memory vs ScallopBot memory · Memory architecture · What it costs to run
Small, specialized, local
Two 4B specialists distilled from ScallopBot’s own production traces, then quantized to run on local hardware. A larger model wrote the training labels; the students never trained on their own output. On a fixed, personal toolset they out-pick much larger general models — a narrow result that says nothing about general leaderboards and everything about what a specialist learns from real traces. Weights, LoRA adapters, and the full method are on Hugging Face.
Reads a user turn and picks which tool to call, with what arguments — or declines when none fit. 73.3% tool-selection on held-out production turns, ahead of a 35B MoE and a paid frontier model on the same toolset. Never fabricated a tool result across 60 failure tests.
Hugging FaceReads a conversation and writes down the durable facts worth keeping, or stays quiet on chatter. 0.725 teacher agreement at 4.2s per call, matching the paid model that labeled its training data — on local hardware.
Hugging Face114 tool-calling turns and 33 memory cases, all held out of training, same harness for every model, thinking disabled. The 35B MoE’s raw memory agreement is higher (0.88) but it returns valid structure only 57.6% of the time, so its usable score is parse-gated. Tool-calling is the clear win; memory is parity with the paid teacher. Training data was anonymized before fine-tuning, and both models, their adapters, and the recipe are public.
Self-improvement, on a leash
An optional, default-off loop distills recurring multi-tool workflows into documentation-only procedure files. Candidates are proposed from a training split and scored against a held-out split; a candidate is promoted only if it beats the frozen baseline by a required margin — a gate that cannot be disabled. Every promotion is recorded in a versioned ledger with a snapshot of the prior version, and a watchdog automatically reverts a promotion that accumulates failures. Unused machine-authored skills are recoverably archived. Machine-authored executable scripts are rejected outright — only documentation is ever written.
The loop ships disabled and stays disabled until you turn it on. Scoring is an A/B comparison judged by an LLM on a held-out split that is disjoint from the split candidates were proposed from. Skills that go unused are archived rather than deleted, and archived skills can be restored. The promotion margin is configurable (0–1); even at 0 a candidate must match or beat the baseline, and the gate itself cannot be switched off.
Intelligence roadmapUp and running in minutes
One script installs everything on a fresh Ubuntu server. Add a provider key and you're live.
# Clone the repo
git clone https://github.com/tashfeenahmed/scallopbot
cd scallopbot
# One-command server setup (Node 24, PM2, voice deps, Ollama)
bash scripts/server-install.sh
# Configure your provider key
cp .env.example .env
nano .env # add one provider key (e.g. ANTHROPIC_API_KEY) + WEB_UI_ENABLED=true
# Build and start
npm run build
node dist/cli.js startOwn your AI assistant
MIT licensed. Self-hosted. No vendor lock-in.
Get Started on GitHub