GLM-5.2 Long-Context-Bug in ik_llama.cpp: NaN-Crash ab 32k Token
ToolsLlama
Warum es zählt
GLM-5.2 ist für 1M-Kontext trainiert, bleibt aber in der Praxis auf kurze Kontexte beschränkt. Bestehende Workarounds (FA off, KV-Cache auf CPU, reduzierter Batch) helfen nicht — GPU-seitige f16-Overflows im DSA/Indexer-Pfad sind der Verdacht. CPU-only als einziger stabiler Pfad wäre ein erheblicher Performance-Verlust.
— Lumeric Redaktion
Frag die KI zum Artikel
Folgefragen zu Headline, Quelle und Volltext — Antwort streamt in wenigen Sekunden.
Verwandte Beiträge
- MEINUNGreddit.com2w
llama.cpp vs. vLLM für 200K+ Kontext: Praxistest mit Qwen3.8-Flash-Next
- BENCHMARKreddit.com3w
GLM-5.2 lokal: ubatch-Größe verdoppelt Prefill-Durchsatz bei langen Prompts
- BENCHMARKreddit.com1w
llama.cpp-Fork für AMD gfx906: +14% Prompt-Throughput, +9% Long-Context-Fill
- MEINUNGreddit.com2w
Ersteinstieg in lokale LLMs mit ik_llama auf 12 GB VRAM
GLM-5.2 Long-Context-Bug in ik_llama.cpp: NaN-Crash ab 32k Token
ToolsLlama
Warum es zählt
GLM-5.2 ist für 1M-Kontext trainiert, bleibt aber in der Praxis auf kurze Kontexte beschränkt. Bestehende Workarounds (FA off, KV-Cache auf CPU, reduzierter Batch) helfen nicht — GPU-seitige f16-Overflows im DSA/Indexer-Pfad sind der Verdacht. CPU-only als einziger stabiler Pfad wäre ein erheblicher Performance-Verlust.
— Lumeric Redaktion
Frag die KI zum Artikel
Folgefragen zu Headline, Quelle und Volltext — Antwort streamt in wenigen Sekunden.
Verwandte Beiträge
- MEINUNGreddit.com2w
llama.cpp vs. vLLM für 200K+ Kontext: Praxistest mit Qwen3.8-Flash-Next
- BENCHMARKreddit.com3w
GLM-5.2 lokal: ubatch-Größe verdoppelt Prefill-Durchsatz bei langen Prompts
- BENCHMARKreddit.com1w
llama.cpp-Fork für AMD gfx906: +14% Prompt-Throughput, +9% Long-Context-Fill
- MEINUNGreddit.com2w
Ersteinstieg in lokale LLMs mit ik_llama auf 12 GB VRAM