/evaluations
Evaluations
Tested local models and the accumulated evidence from WumboLabs hardware. Select a model to see its tested profiles, current findings, testing history, and canonical evidence.
Each WumboLabs-tested model has exactly one canonical Evaluation page answering one question: what have we learned about this model? Current findings come first; tested profiles, results, and the published testing history follow, with every value attributed to the profile and event that measured it.
Public scientific reports and profile metadata live together in WumboLabs/evaluations. Every model page links to its reports at exact commits. Profiles describe tested scientific surfaces, not separate repositories.
Current methodology and historical interpretation limits are explained in Methodology. Protocol revisions do not silently replace frozen model scores or verdicts.
Model evaluations
All evaluations are shown.
| Model | Producer | Artifact / quant | Latest evidence | Status | Evidence |
|---|---|---|---|---|---|
| Ling 3.0 Tiny | inclusionAI (Ant Group) | llama.cpp official Q8_0 (Reasoning On, deployment sampler, 32K) | 2026-09-25 | not ready | 4 events |
| NeoHorse-1-9B | TokenRhythm | llama.cpp official Q8_0 (Reasoning On = publisher default, deployment sampler, 32K) | 2026-09-25 | ready with guardrails | 3 events |
| Granite 4.2 8B | IBM (ibm-granite) | llama.cpp official Q4_K_M (reasoning-on, deployment sampler, 32K) | 2026-09-24 | not ready | 1 event |
| LFM2.5-8B-A1B | Liquid AI | llama.cpp upstream official Q6_K (reasoning-on, publisher sampler) | 2026-09-24 | not ready | 2 events |
| Spark-X2.5-4B | XHToken (SparkLLM Team, iFlytek) | llama.cpp upstream official Q8_0 (reasoning-on, vendor default) | 2026-09-20 | ready with guardrails | 1 event |
| Ternary Bonsai 2 27B | PrismML | llama.cpp PTQ1_0 (thinking-off) | 2026-09-20 | ready with guardrails | 2 events |
| Mellum2 12B-A2.5B | JetBrains | llama.cpp Q4_K_M (Instruct) | 2026-09-16 | limited role only | 5 events |
| Qwen3.6-35B-A3B | Alibaba (Qwen team) | llama.cpp UD-IQ2_M q8_0 KV (32K text-generation profile) | 2026-09-15 | limited role only | 3 events |
| Qwen3-14B | Alibaba (Qwen team) | llama.cpp Q4_K_M q8_0 KV (unsloth artifact, native-max 32K surface) | 2026-09-14 | ready with guardrails | 4 events |
| Gemma 4 12B IT | llama.cpp official QAT Q4_0 (google GGUF, UD-Q4_K_XL packaging) | 2026-09-12 | ready with guardrails | 6 events | |
| Qwen3.8-27B | Alibaba (Qwen team) | ExLlamaV3 H1 (SC_2.20bpw_H3_V3) | 2026-09-12 | completed deep evaluation | 4 events |
| Nemotron 3 Nano 4B | NVIDIA | llama.cpp Q4_K_M | 2026-09-11 | ready with guardrails | 2 events |
| Qwen3.5-4B | Alibaba (Qwen team) | llama.cpp BF16 GGUF | 2026-09-11 | evidence in progress | 1 event |
| Qwen3.5-9B | Alibaba (Qwen team) | llama.cpp Q8_0 GGUF | 2026-09-11 | evidence in progress | 1 event |
| Gemma 4 E4B | llama.cpp official QAT Q4_0 | 2026-09-10 | ready with guardrails / closed welp cross family control | 1 event | |
| MiniCPM5-2B | OpenBMB | contained vLLM BF16 | 2026-09-10 | ready with guardrails | 1 event |
| Ornith 1.5 9B | ornith-ai | llama.cpp Q4_K_M | 2026-09-02 | complete / final welp readiness not ready | 2 events |
| Qwen3.8-4B (Empero) | empero-ai (Qwen3.8 derivative) | llama.cpp Q4_K_M Distill (WumboServer RTX 2060S lane) | 2026-09-01 | evidence in progress | 2 events |
| Qwen3.8-9B (Empero) | empero-ai (Qwen3.8 derivative) | llama.cpp Distill Q4_K_M/Q5_K_M/Q6_K (WumboServer RTX 2060S lane) | 2026-09-01 | evidence in progress | 2 events |
| LFM2.5 1.2B | Liquid AI | llama.cpp QAD Q4_0 | 2026-08-26 | complete small model evaluation | 1 event |
| LFM2.5 2.6B | Liquid AI | llama.cpp QAD Q4_0 | 2026-08-26 | complete small model evaluation | 1 event |
| Apodex 1.1 mini | Apodex (Qwen3.5-35B-A3B MoE base) | llama.cpp IQ1_M (community conversion) | 2026-08-25 | bounded early stop / phase 3 do not advance | 1 event |
| Qwen3.8-2B (Empero) | empero-ai (Qwen3.8 derivative) | vLLM BF16 vendor-alignment battery | 2026-08-18 | evidence in progress | 1 event |
| Qwen2.5-3B Instruct | Alibaba (Qwen team) | vLLM BF16 (cross-runtime lane) | 2026-07-15 | evidence in progress | 2 events |
| Bonsai 27B (Q1_0) | Upstream identity unresolved (local artifact name) | llama.cpp Q1_0 (unscored smoke payload) | 2026-07-14 | evidence in progress | 1 event |
| Gemmable 4 12B MTP | Mia-AiLab (Gemma 4 12B MTP derivative) | llama.cpp Q4_K_M | 2026-07-04 | evidence in progress | 2 events |
| Grug 12B | kai-os | llama.cpp Q4_K_M | 2026-07-04 | evidence in progress | 2 events |
No evaluations match this search.