0110

ToshLLM field guide

Getting started

Install ToshLLM, choose the correct build, pass the first-launch security check and start a local model.

Before you install

Built for a specific kind of Mac.

ToshLLM is a native macOS application for Intel Macs with AMD graphics. It requires macOS 14 or newer and at least 16 GB of memory. For larger mixture-of-experts models, 32 GB or more gives the model and macOS more room to work.

RequiredIntel Mac, AMD GPU with Metal support, macOS 14 or newer and 16 GB RAM.
Recommended32 GB RAM for larger models and enough free storage for GGUF model files.
Apple silicon is not the target.

ToshLLM exists to bring modern local inference to Intel Macs with AMD GPUs. The application and its bundled engines are built around that hardware.

Step 01

Choose the right download.

The standard release is the right choice for most supported Intel Macs. Older Intel processors without AVX2 require the separate no-AVX2 build. The updater keeps the two variants on separate channels so a no-AVX2 installation does not replace itself with an incompatible build.

  1. Download the appropriate DMG from the ToshLLM release page.
  2. Open the DMG and drag ToshLLM into Applications.
  3. Launch ToshLLM from Applications.

The inference engines are included. You do not need Homebrew, Python or a separate llama.cpp installation.

Step 02

Open the official release.

ToshLLM releases are signed and notarized by Apple. Download the DMG from the official GitHub releases page, move the app to Applications, then open it normally. macOS may ask you to confirm the first launch.

Check the source if you see an unexpected warning.

Make sure the DMG came from the official ToshLLM release. Do not bypass a security warning for a copy from an unknown source.

Step 03

Download and run a model.

Open Models to see recommendations for the detected GPU and available memory. Download a GGUF model, choose Use, then start the server from Home. ToshLLM stores models in ~/models by default; the folder can be changed in Settings.

Begin with a smaller model and its recommended quantization. Once it loads reliably, use the benchmark view and live memory readout to decide whether a larger model or context window fits your machine.

First successful run

Know when the setup is healthy.

  1. The selected model appears in the Home server card with its quantization.
  2. The GPU panel reports Metal memory use without exhausting the available reserve.
  3. The server reaches Running and the request count remains available.
  4. A short chat produces tokens and the toolbar reports prompt and generation activity.

After this first check, create a profile for settings you want to reuse. Keep a smaller known-good model installed so hardware and server changes can be tested independently of a demanding model.

Choose a workspace

Chat, create images or generate video.

The control above the main window switches between three independent local workflows. Chat uses the configured GGUF language model and server. Images and Video use their own catalogs and generation runtimes.

ChatConversations, projects, attachments, vision, local agent tools and MCP resources.
ImagesText to image, image to image, prompt queues and ESRGAN upscaling.
VideoText or image to local frame sequences with H.264 MP4 export.
ConfigurationModels, benchmarks, documentation, logs, chat controls and engine settings.

Open the chat guide or the image and video guide for the full workflow.

Updates and storage

Keep builds and models predictable.

ToshLLM can check GitHub releases at launch and periodically while it remains open. It does not install an update silently. The no-AVX2 build remains on its compatible update channel.

Changing the model directory does not move existing GGUF files. Move them manually if needed, then refresh the model catalog. Split GGUF models must keep all of their parts together in the same folder.

For teams and individuals

Make ToshLLM work for your team.

Need a tailored deployment, help choosing models, or guidance for a fleet of Intel Macs? Tell us what you are building and we will get back to you personally.

Found a reproducible bug? A public GitHub issue helps everyone follow the fix. Open an issue ↗
hello@toshllm.com

CONTACT / TOSHLLM

Your message goes directly to ToshLLM. Please do not include passwords, API keys, or private logs.