
Qwen3.8-27B vs Gemma 4 26B: RTX 4090 Benchmarks
Introduction I tested five current local vision-language models on one NVIDIA RTX 4090 across Ollama and direct llama.cpp, centered on Qwen3.8-27B and Gemma 4 26B A4B. The work covered installation, model downloads, coding-harness configuration, practical text and vision context ceilings, multi-token prediction (MTP), short and filled-context throughput, CPU offload, matched runtime performance, and 360 W versus 450 W GPU power limits. The two-model daily result is summarized next. The complete setup and all five-model evidence remain below so the conclusions are reproducible. ...
