← Back to DizyDiz

LLM Benchmark Archive

How our homelab's LLMs have evolved. We run the same 5-prompt benchmark suite every time a model changes — coding fixes, reasoning, tool-calling, persona adherence, summarization — on real hardware (Hommer's RX 7900 XTX and the RainbowAI GPU pool). This page charts what happened when the models were swapped, and where we landed after the July rebuild.
loading… · bench runs · swap decisions · bench dates
filters apply to every chart and the table below

Speed by run median tok/s across the 5-prompt suite

Each bar is one model on one host. Longer is faster. Colour and marker shape both encode the host era, so the grouping never depends on colour alone.

Internal IPs, credentials, and full thinking traces are stripped from sample outputs. Original NDJSON archive lives on Hommer for ops use.

Prompt-type profile tok/s per suite prompt

Where each run spends its speed. A model can be quick at short coding fixes and slow at long summaries — the suite average hides that; this doesn't.

Wall time vs speed every prompt, every run

One dot per prompt. Down-and-right is the good corner: fast tokens and a short wall clock. Dots up top took real seconds no matter the token rate.

Decisions

All runs

The table view — every value in the charts above is readable here, sortable by any column.

Date Model Host Role tok/s (med) Wall (s) Sample outputs