Qwen3.8-27B Q6_K mit 156K Kontext auf einer RTX 5090: 140–190 tok/s via DFlash2
CompaniesHugging Face
Warum es zählt
Für lokale Agentic-Coding-Setups mit Tools und Vision auf Single-GPU: Die Kombination aus DFlash2-Spekulation, q8_0-KV-Cache und 16-GiB-RAM-Prompt-Cache ermöglicht lange Sessions ohne Re-Prefill. Konkrete llama.cpp-Konfiguration und GGUF-Links sind direkt nutzbar.
— Lumeric Redaktion
Decode-Throughput (tok/s, RTX 5090, greedy, thinking off) · Spitzenwert
165%
Qwen3.8-27B Q6_K + DFlash2 (Code)
Frag die KI zum Artikel
Folgefragen zu Headline, Quelle und Volltext — Antwort streamt in wenigen Sekunden.
Verwandte Beiträge
Qwen3.8-27B Q6_K mit 156K Kontext auf einer RTX 5090: 140–190 tok/s via DFlash2
CompaniesHugging Face
Warum es zählt
Für lokale Agentic-Coding-Setups mit Tools und Vision auf Single-GPU: Die Kombination aus DFlash2-Spekulation, q8_0-KV-Cache und 16-GiB-RAM-Prompt-Cache ermöglicht lange Sessions ohne Re-Prefill. Konkrete llama.cpp-Konfiguration und GGUF-Links sind direkt nutzbar.
— Lumeric Redaktion
Decode-Throughput (tok/s, RTX 5090, greedy, thinking off) · Spitzenwert
165%
Qwen3.8-27B Q6_K + DFlash2 (Code)
Frag die KI zum Artikel
Folgefragen zu Headline, Quelle und Volltext — Antwort streamt in wenigen Sekunden.