Community performance summary

GLM 4.7 Flash 23B-A3B

2 GPU configurations measured with this model. Compare hardware differences, quantizations and every recorded result.

2 signed runs · 0 verified · 1 tester
Median prompt1133.79tok/s · 925.86 to 1341.71
Median generation63.75tok/s · 32.66 to 94.83
Configurations1Q4_K_M
Last observedAug 7RDNA 1 · Unknown

What the results show

One clear view.
Every recorded run.

Measurements from comparable hardware and models are presented together. Medians show typical performance, while every signed run remains available below for inspection and sharing.

Recorded evidence

Inspect every run.

Open any result to review its exact hardware, model configuration, measurements and submission details.

1 to 8 of 2 configurations

Verified benchmark Community submission
Model / quantHardwareArchitecturePromptGenerationSource
GLM 4.7 Flash 23B-A3BQ4_K_M · MoE · Signed ToshLLM app · ncmoe 1 · amd-gpu · Metal · comparison
Unknown
1341.71tok/s
94.83tok/s
ToshLLM communityAug 6, 2026
GLM 4.7 Flash 23B-A3BQ4_K_M · MoE · Signed ToshLLM app · amd-gpu · Metal · comparison
RDNA 1
925.86tok/s
32.66tok/s
ToshLLM communityAug 7, 2026