Use the same Mac for three local workflows.
The segmented control above the main window switches between Chat, Images and Video. Each studio has its own model catalog and runtime. Loading a language model is not required for image or video generation.
Let VRAM guide the first choice.
The catalog ranges from compact Stable Diffusion models to Z-Image, Flux and Qwen-Image variants. ToshLLM recommends a model from the selected GPU's usable memory and downloads every required component, including diffusion weights, VAE and text encoders.
Model licenses differ. The catalog labels large or restricted models, but you remain responsible for checking the upstream model license before commercial use or redistribution.
The 7B model can render legible text and edit from up to sixteen reference images. The Q3 variant starts at the 8 GB class; Q4 and Q6 target 12 GB cards, and Q8 targets 16 GB or more. Its Qwen3-VL encoder runs off the card. On a GPU driving the display, the app limits the long edge to 1920 pixels to keep the desktop responsive. This addition is in the current development build and may not be in the latest public download yet.
A model can contain several components that load at different stages. ToshLLM estimates the component that remains resident during sampling and accounts for transient encoders and VAE work separately.
Change one visual variable at a time.
- Aspect and size: available dimensions follow the model's latent grid, quality limits and detected VRAM.
- Steps: distilled and Turbo models expect few steps, while other models need a longer sampling run.
- CFG: use the model's recommended guidance before experimenting.
- Seed: keep a fixed value to compare prompt or setting changes; use
-1for a new result. - Image to image: a lower strength preserves more of the reference and a higher value reinvents it.
- Reference edits (development build): Qwen-Image 2.1 accepts up to sixteen reference images. Start with a small set so the requested edit is clear.
- Output: PNG preserves lossless output; JPG is smaller for photographic results.
A mismatch between the reference image and output aspect ratio can crop or distort the composition. Match both ratios before judging the model.
Separate work without competing for one card.
You can add image-generation instances with their own model, GPU and settings. The prompt queue assigns work only to idle instances and avoids starting two queued jobs on the same GPU.
An instance can also place the main diffusion model on one GPU and supporting components on another. CPU offload reduces VRAM pressure by moving components through system memory, with a speed cost.
The app warns about this configuration because simultaneous heavy Metal workloads can exhaust memory or destabilize the card.
Compare the source and result directly.
The Upscale view accepts batches and processes them one at a time. ToshLLM can download the curated ESRGAN model or use a compatible custom model. A 4x model produces native 4x output; the 2x option scales with the same model and resamples to half.
Results are stored beside the source by default and the comparison view keeps the original available for inspection. Upscaling adds pixels and detail, but it cannot reliably restore information absent from the source.
Budget memory across space and time.
The video catalog includes text-to-video and image-to-video models. Select a frame size, a valid frame count, steps, seed and GPU. The memory estimate accounts for resolution, number of frames and resident model size.
If a run approaches the memory limit, reduce frame count first, then resolution. The studio keeps generated frames as individual PNG files and previews them as an animation.
MP4 export and video attachments require ffmpeg and ffprobe. ToshLLM searches its bundled tools and common local installation paths.