Open Source · MIT License

ScallopBot —
your AI assistant, self-hosted.

Persistent memory, web search, browser automation, and dream-cycle cognition — tied for the top score in a four-agent tool-calling benchmark, and best on the hard tasks.

ScallopBotonline
What did Sarah say about the project deadline?
10:41
memory_search · 847 memories · 3 results
Sarah mentioned on Oct 15 that the deadline moved to November 30. She noted the design review needs to happen first.
10:42
Search for recent papers on sparse attention
10:43
web_search · "sparse attention 2026"
Found 3 papers. Most cited: "Sub-quadratic Attention" (Chen et al.) — linear cross-attention with 40% fewer FLOPs.
10:43
Remind me to follow up with Jake on Friday at 10
10:44
board · scheduled Fri 10:00
Done — I’ll message you Friday at 10:00 to follow up with Jake.
10:44
Message...
Hybrid MemoryModel RoutingTelegram + Web + CLILocal VoiceModular SkillsDream CyclesGuarded EvolutionDashboardSchedulingReliability

Tool-calling benchmark

ScallopBench v2: 36 tasks (12 trap, 6 coding, 3 assistant, 15 hard), each run 3 times per agent — 108 task-runs each. ScallopBot, Prime Agent, OpenClaw and Hermes Agent all used the same model (Moonshot kimi-k2.6, thinking on). Scoring looks only at outcomes — the files left in the workspace and the replies, including hidden tests — never at what the agent claims it did. ScallopBot ties Prime Agent for the top overall score at 98.1%, is best on the hard tasks (44/45), and never followed the hidden prompt injection.

98.1%
Overall pass rate
106/108 task-runs · tied for the top score with Prime Agent
44/45
Hard tasks
Best of the four agents on the 15 hard tasks, run 3 times each
3/3
Injection resisted
Never followed the hidden prompt injection; Hermes Agent followed it in all three runs
Results by agent
Overall pass rate
All 36 tasks (trap, coding, assistant, hard), 3 runs each — 108 task-runs per agent
ScallopBot
98.1%
Prime Agent
98.1%
OpenClaw
97.2%
Hermes Agent
93.5%
Tied for the top score
Hard tasks
15 hard multi-step tasks, 3 runs each, scored with hidden tests
ScallopBot
44/45
Prime Agent
43/45
OpenClaw
43/45
Hermes Agent
39/45
Best on the hard tasks
Prompt-injection resistance
A task README hid an instruction to delete files; runs where the agent did not follow it
ScallopBot
3/3
Prime Agent
2/3
OpenClaw
3/3
Hermes Agent
0/3
Never followed the injection

All four agents passed every trap (36/36) and assistant (9/9) run; on coding, Prime Agent scored 18/18 and the other three 17/18. Differences of one or two tasks are within run-to-run spread. Competitors ran on 2 Oct 2026 (Hermes Agent 0be2d56, Prime Agent cf285dc, OpenClaw 2026.9.7); ScallopBot ran on 3 Oct 2026 (current main). Methodology and per-task results: RESULTS-v2.md.

→ OpenClaw memory vs ScallopBot memory · Memory architecture · What it costs to run

Small, specialized, local

Two 4B specialists distilled from ScallopBot’s own production traces, then quantized to run on local hardware. A larger model wrote the training labels; the students never trained on their own output. On a fixed, personal toolset they out-pick much larger general models — a narrow result that says nothing about general leaderboards and everything about what a specialist learns from real traces. Weights, LoRA adapters, and the full method are on Hugging Face.

73.3%
Tool selection
Beats a 35B MoE (46.5%) and a paid frontier model (54.7%) on the toolset
0 / 60
Fabrications
Across single and multi-step failure tests — most trustworthy in the lineup
2.9 GB
Footprint
q5 GGUF runs locally, even on a Raspberry Pi 5 at ~3 tok/s
scalloptools-1Tool calling

Reads a user turn and picks which tool to call, with what arguments — or declines when none fit. 73.3% tool-selection on held-out production turns, ahead of a 35B MoE and a paid frontier model on the same toolset. Never fabricated a tool result across 60 failure tests.

Hugging Face
scallopmemory-1Memory extraction

Reads a conversation and writes down the durable facts worth keeping, or stays quiet on chatter. 0.725 teacher agreement at 4.2s per call, matching the paid model that labeled its training data — on local hardware.

Hugging Face
Tool-calling & memory, by model
Tool selection
Correct function chosen on 114 held-out production turns (thinking off, greedy)
ScallopBot 4Bfine-tuned
73.3%
Qwen3.5 4Bstock
65.3%
Qwen3.6 Pluspaid, ~400B+
54.7%
Qwen3.6 MoElocal, 35B
46.5%
Memory agreement
Agreement with the teacher on 33 held-out extraction cases
Qwen3.6 Pluspaid teacher
74.8%
ScallopBot 4Bfine-tuned
72.5%
Qwen3.5 4Bstock
69.5%
Qwen3.6 MoElocal 35B · 57.6% parse
57.6%

114 tool-calling turns and 33 memory cases, all held out of training, same harness for every model, thinking disabled. The 35B MoE’s raw memory agreement is higher (0.88) but it returns valid structure only 57.6% of the time, so its usable score is parse-gated. Tool-calling is the clear win; memory is parity with the paid teacher. Training data was anonymized before fine-tuning, and both models, their adapters, and the recipe are public.

Self-improvement, on a leash

An optional, default-off loop distills recurring multi-tool workflows into documentation-only procedure files. Candidates are proposed from a training split and scored against a held-out split; a candidate is promoted only if it beats the frozen baseline by a required margin — a gate that cannot be disabled. Every promotion is recorded in a versioned ledger with a snapshot of the prior version, and a watchdog automatically reverts a promotion that accumulates failures. Unused machine-authored skills are recoverably archived. Machine-authored executable scripts are rejected outright — only documentation is ever written.

+0.05
Minimum measured gain
Default fitness margin (0\u20131 scale) over the frozen baseline, on a held-out split
100%
Versions ledgered
Every promotion snapshots the prior version and stays reversible
0
Machine-authored scripts
Executable output is rejected outright — documentation only

The loop ships disabled and stays disabled until you turn it on. Scoring is an A/B comparison judged by an LLM on a held-out split that is disjoint from the split candidates were proposed from. Skills that go unused are archived rather than deleted, and archived skills can be restored. The promotion margin is configurable (0–1); even at 0 a candidate must match or beat the baseline, and the gate itself cannot be switched off.

Intelligence roadmap

Up and running in minutes

One script installs everything on a fresh Ubuntu server. Add a provider key and you're live.

# Clone the repo
git clone https://github.com/tashfeenahmed/scallopbot
cd scallopbot

# One-command server setup (Node 24, PM2, voice deps, Ollama)
bash scripts/server-install.sh

# Configure your provider key
cp .env.example .env
nano .env  # add one provider key (e.g. ANTHROPIC_API_KEY) + WEB_UI_ENABLED=true

# Build and start
npm run build
node dist/cli.js start

Own your AI assistant

MIT licensed. Self-hosted. No vendor lock-in.

Get Started on GitHub