Prompt and generation answer different questions.
A multi-GPU split may improve prompt speed while leaving generation unchanged or slower. Always read the two values separately.
The server defines the shared run.
When you choose to share a benchmark, the app requests a short-lived challenge from ToshLLM. That response defines the runner, prompt-token count, generated-token count, repetition count and accepted consent version. The current comparable workload uses llama-bench; each repetition is launched independently so every row is preserved.
The published result stores every prompt and generation measurement, then presents the median. It also records the model artifact hashes, engine hash, hardware, operating system and effective runtime settings.
The server and benchmark share VRAM. Leaving both active can distort the result or make the workload fail.
Review first, sign second.
- Choose Share benchmark in the ToshLLM app.
- Read the consent summary and run the server-defined workload.
- Review the model, hardware, configuration, measurements and sanitized evidence.
- Sign and send the exact bytes shown in the review.
The private P-256 key stays in this Mac's Keychain. The public fingerprint groups submissions from the same installation and lets the API verify that the payload was not changed after review.
Compare like with like.
The public index groups runs by GPU and model so repeated submissions strengthen a comparison instead of creating a wall of nearly identical pages. Quantization, model family, backend, GPU count, context and runtime flags remain visible because they can materially change performance.
Open the community benchmark index to browse current results and the evidence behind each aggregate.