1010

ToshLLM field guide

Troubleshooting

Fix first launch, CPU compatibility, memory pressure, model loading and local server connection problems.

First launch

macOS says the app cannot be opened.

The official ToshLLM release is signed and notarized. Confirm that you downloaded the correct DMG from the official GitHub releases page and moved the app to Applications. If macOS still shows a security warning, keep the exact message and report it rather than bypassing an unexpected warning.

CPU compatibility

The engine exits with an illegal instruction.

An illegal-instruction or SIGILL error usually means the processor lacks AVX2. Install the release whose filename includes noavx2. ToshLLM marks that variant internally so future update checks continue selecting the compatible package.

Memory pressure

The model fails while loading.

  1. Stop any server or benchmark already using VRAM.
  2. Reduce context size to shrink the KV cache.
  3. Increase the VRAM reserve so macOS and Metal allocations have headroom.
  4. For MoE models, move more experts to the CPU.
  5. Try a smaller quantization or model if the file still cannot fit.

Change one setting at a time and review Logs after every attempt. The app translates common engine failures into guidance, while the complete log preserves the underlying message.

Local API

A client cannot connect.

  • Confirm the server status is Running and verify its configured port. The default is 8080.
  • On the same Mac, use http://127.0.0.1:8080 rather than a public URL.
  • If API-key protection is enabled, include Authorization: Bearer ….
  • For another device, enable local-network discovery, allow the macOS network prompt and use the Mac's local address.
  • Restart the server after changing its listener or security settings.
Vision and attachments

The model ignores an image, audio file or video.

  • Images require the matching mmproj file and the vision eye control must be enabled.
  • Audio and video controls appear only when the selected model advertises those modalities.
  • Video input and MP4 export require ffmpeg and ffprobe.
  • Review the attachment-token warning. A file that fills the context can make the request fail before generation.
  • For scanned PDFs, wait until on-device OCR finishes before sending.
Agents and MCP

A tool never runs or repeats in a loop.

  • Enable Local agent tools and use a model with a compatible tool-calling template.
  • Inspect pending permission cards. A denied call is returned to the model as an error.
  • Test each MCP server from Chat Settings and confirm that tools are discovered.
  • Validate optional MCP headers as a JSON object and check the server URL and transport.
  • Reduce the maximum agent turns when a model keeps retrying an invalid call.

Start with a read-only tool. If that works but writes do not, the remaining issue is usually permission, runtime isolation or filesystem access rather than model connectivity.

Image and video studio

A creative run exhausts VRAM.

  1. Stop the chat server or another studio instance using the same GPU.
  2. Reduce image dimensions or video frame count.
  3. Choose a smaller model or enable CPU offload.
  4. On a multi-GPU Mac, move supporting components to a different card.
  5. Avoid running two heavy instances on one AMD GPU.

The image and video logs are separate from the language-model server log. Use the stage shown in the studio to identify whether loading, sampling or decoding failed.

Support

Bring evidence, not guesses.

When reporting a reproducible issue, include the ToshLLM version, Mac model, AMD GPU, macOS version, model filename, relevant settings and the smallest useful section of the log. Do not publish API keys, private paths or personal chat content.

The Logs view supports search, subsystem filters, automatic following, copying and a diagnostics export. Clear only resets the visible view; it does not delete the log file stored on disk.

Open a GitHub issue when the problem persists.

Fast diagnosis

Start with the subsystem that failed.

DownloadCheck available storage, the Hugging Face URL and whether every split GGUF part completed.
Model loadInspect VRAM, context size, GPU selection, quantization and the first engine error.
GenerationCheck the chat template, tool settings, context exhaustion and whether the server remains healthy.
IntegrationVerify the base URL, model ID, Bearer key and the endpoint expected by the client.

Use the Logs view for the full server lifecycle. Benchmark logs are stored separately so a failed or interrupted measurement can be diagnosed without mixing it into chat output.

For teams and individuals

Make ToshLLM work for your team.

Need a tailored deployment, help choosing models, or guidance for a fleet of Intel Macs? Tell us what you are building and we will get back to you personally.

Found a reproducible bug? A public GitHub issue helps everyone follow the fix. Open an issue ↗
hello@toshllm.com

CONTACT / TOSHLLM

Your message goes directly to ToshLLM. Please do not include passwords, API keys, or private logs.