0810

ToshLLM field guide

Benchmark guide

Read prompt and generation measurements, compare configurations and share signed results from the app.

Two measurements

Prompt and generation answer different questions.

Prompt processingHow quickly the runtime reads the supplied context. Reported as tokens per second.
Token generationHow quickly the model produces new tokens after processing the prompt. Reported as tokens per second.

A multi-GPU split may improve prompt speed while leaving generation unchanged or slower. Always read the two values separately.

Comparable workload

The server defines the shared run.

When you choose to share a benchmark, the app requests a short-lived challenge from ToshLLM. That response defines the runner, prompt-token count, generated-token count, repetition count and accepted consent version. The current comparable workload uses llama-bench; each repetition is launched independently so every row is preserved.

The published result stores every prompt and generation measurement, then presents the median. It also records the model artifact hashes, engine hash, hardware, operating system and effective runtime settings.

Stop the running model server first.

The server and benchmark share VRAM. Leaving both active can distort the result or make the workload fail.

Signed sharing

Review first, sign second.

  1. Choose Share benchmark in the ToshLLM app.
  2. Read the consent summary and run the server-defined workload.
  3. Review the model, hardware, configuration, measurements and sanitized evidence.
  4. Sign and send the exact bytes shown in the review.

The private P-256 key stays in this Mac's Keychain. The public fingerprint groups submissions from the same installation and lets the API verify that the payload was not changed after review.

Reading the index

Compare like with like.

The public index groups runs by GPU and model so repeated submissions strengthen a comparison instead of creating a wall of nearly identical pages. Quantization, model family, backend, GPU count, context and runtime flags remain visible because they can materially change performance.

Open the community benchmark index to browse current results and the evidence behind each aggregate.

Trust and moderation

A signature proves origin, not performance.

The signature proves that the registered ToshLLM installation signed the exact payload accepted by the API. Automatic checks then compare the submission with compatible hardware, model and workload baselines.

Verified benchmarkA result reproduced or explicitly verified under ToshLLM's review process.
Community submissionA signed app result that passed validation but does not claim independent reproduction.
Automatic approvalThe measurements fit the expected range and required evidence is complete.
Manual reviewAn unusual value is held back so an administrator can inspect evidence before publication.
Normal variation

One run is an observation, not a law.

Thermal state, background GPU activity, memory pressure and first-load effects can move a result. The shared protocol keeps all independent repetitions and displays their median to reduce the influence of one unusually fast or slow measurement.

Use grouped pages to inspect the range across contributors. A narrow cluster is stronger evidence than a single leading result.

For teams and individuals

Make ToshLLM work for your team.

Need a tailored deployment, help choosing models, or guidance for a fleet of Intel Macs? Tell us what you are building and we will get back to you personally.

Found a reproducible bug? A public GitHub issue helps everyone follow the fix. Open an issue ↗
hello@toshllm.com

CONTACT / TOSHLLM

Your message goes directly to ToshLLM. Please do not include passwords, API keys, or private logs.