DizyDiz — LLM & Agent Map

Which agent / tool talks to which brain — solid comet = routed through the ollama-proxy · dotted = straight to the GPU. OpenDiz agents go through the proxy. Live routing as of 2026-07-22 — post-bench re-tier: ALL agent work = qwen3.6:35b @ the 5060 Ti pool; Hommer = fallback ONLY. Not shown: per-request fallback paths (alerts fire to Discord when used).

Mr. Peepers → Hommer Velma → 3060 Louis → 3060 Subagents → 5060 Ti pool LiteLLM/N8N → 5060 Ti pool Embeddings → 3060 aivid → Hommer SD images → 3060
CONSUMERS ROUTER MODELS · GPUs 🖥 HOMMER · RX 7900 XTX 24GB · :11434 (FALLBACK ONLY — display GPU; any use alerts Discord) 🌈 RAINBOWAI2 · 2× RTX 5060 Ti · 32GB POOL · :11434 🌈 RAINBOWAI · RTX 3060 12GB · :11435 🐹Mr. Peepers main agent · Discord #mrpeepers-bot → qwen3.6:35b @ RainbowAI pool 📝Velma content/research subagent → qwen3.5:9b @ RTX 3060 🖨Louis 3D-print/shop subagent → qwen3.5:9b @ RTX 3060 🧩Workflow subagents agents.defaults · tool tasks → qwen3-coder:30b @ 5060 Ti pool 🔀LiteLLM + N8N apps gateway .16:4000 · Mr Krabs, etc. → RainbowAI models 🧠Memory / embeddings memory-core (all agents) → nomic-embed-text @ 3060 (direct) 🎬aivid worker AI video · ComfyUI :8188 → WAN 2.2 14B @ Hommer (direct) 🎨SD image service sd-service :7860 → SDXL-Turbo / FLUX @ 3060 (direct) 🔁ollama-proxy OpenDiz VM (internal) • routes by model name → host • injects keep_alive (peepers 5m · qwen3.5 2h) • /no_think strip · tool_call normalize • error passthrough → fallback (fix 07-06) qwen3-coder:30b-peepers Peepers PRIMARY · 91 tok/s · agent runs FALLBACK ONLY: coder:30b-peepers · 9b · 2.5vl · nomic on-disk spares WAN 2.2 14B · ComfyUI aivid video render qwen3.6:35b — team work model (benched 07-22) WORK tier: Peepers+Velma+Louis runs · 109 tok/s · coder:30b installed for coding · 27b deprecated (misfits one card, 24 tok/s) qwen3.5:9b Chat lane (all agents) + judge + vision + embeds · ctx 32k qwen2.5vl:7b · nomic-embed Farm Watch vision · embeddings SDXL-Turbo / FLUX.1 sd-service (rebuilding)

🐹 Mr. Peepers MAIN

Primaryqwen3-coder:30b-peepers — Hommer RX 7900 XTX (91 tok/s)
Fallbacksqwen3-coder:30b — RainbowAI2 5060 Ti pool
RuleOnly agent allowed on Hommer; load → work → free VRAM

📝 Velma & 🖨 Louis RAINBOWAI

Runsqwen3-coder:30b — 5060 Ti pool, pinned resident (re-tiered 07-21)
Chatqwen3.5:9b — RTX 3060 :11435 (55 tok/s)
Fallbackqwen3-coder:30b — 5060 Ti pool (:11434)
RuleNever Hommer — dedicated RainbowAI box only

🖥 Hommer RX 7900 XTX

VRAM24 GB · fastest single GPU (91 tok/s qwen3-coder)
RunsPeepers primary + fallback, aivid WAN 2.2 video
NoteJess's workstation — freed after each job, never pinned

🌈 RainbowAI 3 GPUs

Pool2× RTX 5060 Ti 16GB (32GB) · :11434 · ctx 65k · coder:30b resident @ 58 tok/s · qwen3.6:27b deep/256K
RTX 306012 GB · :11435 · qwen3.5:9b chat + qwen2.5vl vision + embeddings
RoleAll non-Peepers inference + image gen

🔁 ollama-proxy .233

Routesby model name → correct host/GPU
Injectskeep_alive per model; /no_think for extraction
Guardstool_call arg normalize + upstream-error passthrough → clean fallback (added 2026-07-06)

☁️ Claude Code EXTERNAL

ModelFable 5 / Opus 4.8 — Anthropic cloud
RoleDesigns, builds & fixes the whole fleet; not a local model
OrchestratesN8N, the bots, and this map itself