Built for a specific kind of Mac.
ToshLLM is a native macOS application for Intel Macs with AMD graphics. It requires macOS 14 or newer and at least 16 GB of memory. For larger mixture-of-experts models, 32 GB or more gives the model and macOS more room to work.
ToshLLM exists to bring modern local inference to Intel Macs with AMD GPUs. The application and its bundled engines are built around that hardware.
Choose the right download.
The standard release is the right choice for most supported Intel Macs. Older Intel processors without AVX2 require the separate no-AVX2 build. The updater keeps the two variants on separate channels so a no-AVX2 installation does not replace itself with an incompatible build.
- Download the appropriate DMG from the ToshLLM release page.
- Open the DMG and drag ToshLLM into Applications.
- Launch ToshLLM from Applications.
The inference engines are included. You do not need Homebrew, Python or a separate llama.cpp installation.
Open the official release.
ToshLLM releases are signed and notarized by Apple. Download the DMG from the official GitHub releases page, move the app to Applications, then open it normally. macOS may ask you to confirm the first launch.
Make sure the DMG came from the official ToshLLM release. Do not bypass a security warning for a copy from an unknown source.
Download and run a model.
Open Models to see recommendations for the detected GPU and available memory. Download a GGUF model, choose Use, then start the server from Home. ToshLLM stores models in ~/models by default; the folder can be changed in Settings.
Begin with a smaller model and its recommended quantization. Once it loads reliably, use the benchmark view and live memory readout to decide whether a larger model or context window fits your machine.
Know when the setup is healthy.
- The selected model appears in the Home server card with its quantization.
- The GPU panel reports Metal memory use without exhausting the available reserve.
- The server reaches Running and the request count remains available.
- A short chat produces tokens and the toolbar reports prompt and generation activity.
After this first check, create a profile for settings you want to reuse. Keep a smaller known-good model installed so hardware and server changes can be tested independently of a demanding model.
Chat, create images or generate video.
The control above the main window switches between three independent local workflows. Chat uses the configured GGUF language model and server. Images and Video use their own catalogs and generation runtimes.
Open the chat guide or the image and video guide for the full workflow.
Keep builds and models predictable.
ToshLLM can check GitHub releases at launch and periodically while it remains open. It does not install an update silently. The no-AVX2 build remains on its compatible update channel.
Changing the model directory does not move existing GGUF files. Move them manually if needed, then refresh the model catalog. Split GGUF models must keep all of their parts together in the same folder.