The slowdown is far more than 15% for token generation. Token generation is most...

DiabloD3 · 2025-09-13T06:27:20 1757744840

I suggest figuring out what your configuration problem is.

Which llama.cpp flags are you using, because I am absolutely not having the same bug you are.

EnPissant · 2025-09-13T06:44:24 1757745864

It's not a bug. It's the reality of token generation. It's bottlenecked by memory bandwidth.

Please publish your own benchmarks proving me wrong.

DiabloD3 · 2025-09-14T12:25:15 1757852715

I cannot reproduce your bug on AMD. I'm going to have to conclude this is a vendor issue.