Community performance summary

LLaMA V2 7B

3 GPU configurations measured with this model. Compare hardware differences, quantizations and every recorded result.

3 signed runs · 1 verified · 3 testers
Median prompt655.45tok/s · 128.09 to 1508.5
Median generation72.73tok/s · 12.75 to 105.77
Configurations1Q4_0
Last observedAug 6Apple Silicon · RDNA · RDNA 2

What the results show

One clear view.
Every recorded run.

Measurements from comparable hardware and models are presented together. Medians show typical performance, while every signed run remains available below for inspection and sharing.

Hardware tested3 groups
AMD Radeon GFX10, 16 GB1 runAMD Radeon RX 6700 XT1 runApple M11 run
macOS

macOS 15.7.8 Sequoia · macOS 26.5.1 Tahoe · macOS 26.5.2 Tahoe

Recorded evidence

Inspect every run.

Open any result to review its exact hardware, model configuration, measurements and submission details.

1 to 8 of 3 configurations

Verified benchmark Community submission
Model / quantHardwareArchitecturePromptGenerationSource
LLaMA V2 7BQ4_0 · Dense · Signed ToshLLM app · amd-gpu · Metal · comparison
RDNA
1508.5tok/s
105.77tok/s
ToshLLM communityAug 6, 2026
LLaMA V2 7BQ4_0 · Dense · Signed ToshLLM app · amd-gpu · Metal · comparison
RDNA 2
655.45tok/s
72.73tok/s
ToshLLM LabJul 20, 2026
LLaMA V2 7BQ4_0 · Dense · Signed ToshLLM app · standard-auto · Metal · comparison
Apple M111.84 GB VRAM
Apple Silicon
128.09tok/s
12.75tok/s
bastenJul 18, 2026