Start with the workload, then the size.
Model size alone does not describe how a model will feel. Architecture, quantization, context length, prompt size and the split between GPU and CPU all affect speed and memory use. ToshLLM recommendations combine the detected hardware with measured estimates, but they remain estimates rather than guarantees.
Trade memory for fidelity deliberately.
GGUF quantization reduces model size and memory pressure. Lower-bit variants usually fit more easily and load faster, while larger variants preserve more precision. Names such as Q4_K_M, Q6_K and Q8_0 identify different storage formats, not different parameter counts.
Compare results only when model, quantization and relevant runtime settings match. Two files with the same display name can still have different artifact hashes or internal metadata.
Prism ML's Ternary Bonsai 2 27B is supported in PQ2_0 and PTQ1_0, including its matching vision projector. These are distinct GGUF packings; compare the exact file and configuration when measuring them.
One GGUF may not be the complete capability.
The model catalog can associate the main GGUF with optional companion artifacts. A vision projector adds image input, while a compatible DFlash draft can accelerate speculative generation. Keep split GGUF parts and their companion files in the configured model directory.
mmproj lets a multimodal model read images. Disable the eye control to reclaim its VRAM for text-only work.Leave room for the runtime.
- GPU layers: the default value of 99 requests full GPU offload when the model permits it.
- CPU MoE experts: moving experts to the CPU can make a large MoE model fit, with a generation-speed tradeoff.
- VRAM reserve: ToshLLM reserves memory for macOS and runtime allocations instead of filling the GPU to the final megabyte.
- Context size: larger windows consume more KV-cache memory. Reduce context before reducing model quality when an oversized window is unnecessary.
- KV cache types: quantized cache formats reduce memory use, and the app applies the compatible attention route when needed.
Lower context, increase the VRAM reserve or move more MoE experts to CPU. This keeps the cause visible and makes the next benchmark meaningful.
Keep the fast path on the GPU.
The bundled engine includes AMD-specific Metal work that the app selects automatically where compatible.
Quantized KV formats require a compatible Flash Attention route. If a custom external engine lacks the ToshLLM AMD kernel, verify both speed and output before reusing the same cache settings.
Use every card with context.
ToshLLM can pin a model to one GPU or split it across selected AMD GPUs. Tensor splitting and direct peer transfers are available where the hardware supports them. Duo cards are still composed of multiple GPU devices, so the benchmark records the complete detected configuration.
More GPUs can improve prompt processing while cross-device synchronization can limit token generation. Measure both prompt and generation speed before deciding that a wider split is better.
Fit first. Then optimize speed.
- Load the model with the recommended profile and a practical context size.
- Confirm stable generation before changing GPU selection or expert placement.
- For MoE models, use the optimum finder to locate a safe CPU-expert value.
- Benchmark prompt and generation separately with the server stopped.
- Save the configuration as a profile only after it remains stable in real chat.
A benchmark can expose a fast configuration that is too close to the memory limit for long conversations. Leave headroom for KV cache growth, image projectors and other applications using the GPU.
Measure accepted work, not just the badge.
MTP and DFlash propose multiple tokens and keep only those confirmed by the main model, so accepted output does not change the answer the main model would have produced. Their benefit depends on acceptance rate, verification cost and memory pressure.
The live diagnostics show acceptance and the average accepted tokens per verification. A value below one can cost more than it saves. Compare real chat generation with speculation on and off instead of relying on llama-bench, which measures raw decode without the complete interactive path.
Save a known-good runtime.
A profile captures the selected model and runtime configuration so a stable setup can be restored after experimentation. Import and export in Settings transfers the app's resettable configuration as a versioned JSON file.
Imported engine settings take effect after the server restarts. API and MCP credentials remain protected separately and are not placed into the exported settings file.
Serve several downloaded models through one port.
Router mode publishes stable aliases through GET /v1/models and loads the model named in each request. The model limit controls how many remain resident, using least-recently-used unloading when another model is requested.
Start with a limit of one on a single GPU. Every routed model retains its own vision, MoE and speculative-decoding configuration, but switching still incurs model load time.
Thunderbolt changes the tradeoff.
External AMD GPUs have enough VRAM to make otherwise impossible models fit, but repeated transfers over Thunderbolt can reduce throughput. ToshLLM can force VRAM-resident private Metal buffers for an eGPU so weights do not stream from system memory for every operation.
Pin the intended eGPU explicitly when possible. If macOS chooses it automatically, use the matching external-GPU memory option and confirm the selected device in the server log.