Same 30B coding model, same test — the retired array vs the new silicon.
Speed only earns a spot if tool-calls are 3/3 — dropped tool calls are what killed OpenClaw.
Twelve candidates researched → three worth downloading → one earned a slot.
Our biggest pool is 32GB. Weight size decides who even gets a tryout; the frontier darlings aren't runnable in this house.
Small-class paper cuts (fit, but no case): DeepSeek R1 distills & Llama 3.3 8B — qwen3.5:9b already beats their profile at 3/3 tools; Dolphin 3.0 (uncensored) — noted for the guardrails article, no fleet role. Frontier sizes are estimates.
The driveway watcher samples every 20 seconds — a model slower than the poll can never catch up.
| Model | Host | tok/s | Tools 3× | Strict JSON | Vision | Verdict |
|---|---|---|---|---|---|---|
| qwen3-coder:30b-peepers | Hommer | 90.5 | 3/3 | ✓ | — | kept — Peepers primary |
| gpt-oss:20b | Hommer | 122.3 | 2/3 | ✓ | — | removed — tool drops |
| qwen3.6:27b | Hommer | 6.2 | 3/3 | ✗ | — | removed — ROCm speed |
| qwen3-coder:30b | Pool | 58.3 | 3/3 | ✓ | — | kept — default + fallback |
| qwen3.6:27b | Pool | 23.6 | 3/3 | ✓ | 46s | added — deep/256K/vision |
| qwen3:14b | Pool | 43.1 | 3/3 | ✗ | — | kept — mid batch |
| gpt-oss:20b | Pool | 86.0 | 2/3 | ✗ | — | removed |
| qwen3.5:9b | 3060 | 59.8 | 3/3 | ✗ | — | kept — chat |
| gemma4:12b | 3060 | 36.5 | 3/3 | ✗ | 22.9s | removed — vision too slow |
| qwen2.5vl:7b | 3060 | 75.5 | — | — | 1.6s | kept — Farm Watch |
| nomic-embed-text | 3060 | instant | — | — | — | kept — reindex cost blocks swap |
Method: fleet fully quiesced (farm-watch paused, OpenDiz timer + chat bridge stopped, all models unloaded) per Jess. Harness: warm-speed on a 40-number listing · native tool-calls ×3 trials via /api/chat · strict-JSON extraction · vision on a live farm snapshot. Cost figures are estimates pending receipts.