Which agent / tool talks to which brain — solid comet = routed through the ollama-proxy · dotted = straight to the GPU. OpenDiz agents go through the proxy. Live routing as of 2026-07-22 — post-bench re-tier: ALL agent work = qwen3.6:35b @ the 5060 Ti pool; Hommer = fallback ONLY. Not shown: per-request fallback paths (alerts fire to Discord when used).
| Primary | qwen3-coder:30b-peepers — Hommer RX 7900 XTX (91 tok/s) |
| Fallbacks | qwen3-coder:30b — RainbowAI2 5060 Ti pool |
| Rule | Only agent allowed on Hommer; load → work → free VRAM |
| Runs | qwen3-coder:30b — 5060 Ti pool, pinned resident (re-tiered 07-21) |
| Chat | qwen3.5:9b — RTX 3060 :11435 (55 tok/s) |
| Fallback | qwen3-coder:30b — 5060 Ti pool (:11434) |
| Rule | Never Hommer — dedicated RainbowAI box only |
| VRAM | 24 GB · fastest single GPU (91 tok/s qwen3-coder) |
| Runs | Peepers primary + fallback, aivid WAN 2.2 video |
| Note | Jess's workstation — freed after each job, never pinned |
| Pool | 2× RTX 5060 Ti 16GB (32GB) · :11434 · ctx 65k · coder:30b resident @ 58 tok/s · qwen3.6:27b deep/256K |
| RTX 3060 | 12 GB · :11435 · qwen3.5:9b chat + qwen2.5vl vision + embeddings |
| Role | All non-Peepers inference + image gen |
| Routes | by model name → correct host/GPU |
| Injects | keep_alive per model; /no_think for extraction |
| Guards | tool_call arg normalize + upstream-error passthrough → clean fallback (added 2026-07-06) |
| Model | Fable 5 / Opus 4.8 — Anthropic cloud |
| Role | Designs, builds & fixes the whole fleet; not a local model |
| Orchestrates | N8N, the bots, and this map itself |