Community performance summary

Meta Llama 3.1 8B

3 GPU configurations measured with this model. Compare hardware differences, quantizations and every recorded result.

6 signed runs · 0 verified · 3 testers
Median prompt510.02tok/s · 271.81 to 828.51
Median generation50.64tok/s · 17.97 to 77.02
Configurations1Q4_K_M
Last observedJul 20GCN / Vega + RDNA 2 · RDNA 2

What the results show

One clear view.
Every recorded run.

Measurements from comparable hardware and models are presented together. Medians show typical performance, while every signed run remains available below for inspection and sharing.

Recorded evidence

Inspect every run.

Open any result to review its exact hardware, model configuration, measurements and submission details.

1 to 8 of 6 configurations

Verified benchmark Community submission
Model / quantHardwareArchitecturePromptGenerationSource
Meta Llama 3.1 8BQ4_K_M · Dense · Signed ToshLLM app · amd-gpu · Metal · comparison
RDNA 2
828.51tok/s
77.02tok/s
Dolphins1972Jul 18, 2026
Meta Llama 3.1 8BQ4_K_M · Dense · Signed ToshLLM app · amd-gpu · Metal · comparison
RDNA 2
612.88tok/s
67tok/s
Flint IronstagJul 18, 2026
Meta Llama 3.1 8BQ4_K_M · Dense · Signed ToshLLM app · amd-gpu · Metal · comparison
RDNA 2
611.58tok/s
66.87tok/s
ToshLLM communityJul 18, 2026
Meta Llama 3.1 8BQ4_K_M · Dense · Signed ToshLLM app · amd-gpu · Metal · comparison
GCN / Vega + RDNA 2
277.72tok/s
34.4tok/s
ToshLLM communityJul 20, 2026
Meta Llama 3.1 8BQ4_K_M · Dense · Signed ToshLLM app · amd-gpu · Metal · comparison
GCN / Vega + RDNA 2
408.46tok/s
30.83tok/s
Flint IronstagJul 20, 2026
Meta Llama 3.1 8BQ4_K_M · Dense · Signed ToshLLM app · amd-gpu · Metal · comparison
GCN / Vega + RDNA 2
271.81tok/s
17.97tok/s
Flint IronstagJul 20, 2026