0510

ToshLLM field guide

Local API

Use ToshLLM through OpenAI-compatible and Anthropic-compatible APIs from editors, agents, SDKs and other local applications.

Default connection

Two familiar APIs on one local server.

When the server is running, ToshLLM exposes OpenAI-compatible and Anthropic-compatible endpoints at http://127.0.0.1:8080 by default. The host remains limited to the Mac until local-network discovery is enabled.

OpenAI chatPOST /v1/chat/completions
OpenAI responsesPOST /v1/responses
Anthropic messagesPOST /v1/messages
Anthropic token countPOST /v1/messages/count_tokens
Available modelsGET /v1/models
EmbeddingsPOST /v1/embeddings when embeddings mode is enabled.
Web interfaceThe server root exposes the lightweight chat interface bundled with ToshLLM.
OpenAI Responses API

Use clients built for the newer response format.

The current ToshLLM engine accepts POST /v1/responses in addition to Chat Completions. This gives newer OpenAI-compatible clients a direct local route without translating them through a cloud service.

curl http://127.0.0.1:8080/v1/responses \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer YOUR_TOSHLLM_KEY" \
  -d '{
    "model": "MODEL_ID",
    "input": "Summarize the benefits of local inference.",
    "stream": true
  }'

Use the exact local model ID. Support covers the response surface implemented by the bundled engine, not cloud-only OpenAI services.

Anthropic Messages API

Send a native Anthropic-style message.

Use the server root as the Anthropic base URL. Do not append /v1 when configuring an Anthropic SDK because the client adds /v1/messages itself.

curl http://127.0.0.1:8080/v1/messages \
  -H "Content-Type: application/json" \
  -H "anthropic-version: 2023-06-01" \
  -H "x-api-key: YOUR_TOSHLLM_KEY" \
  -d '{
    "model": "MODEL_ID",
    "max_tokens": 256,
    "messages": [
      {"role": "user", "content": "Explain Metal in one paragraph."}
    ]
  }'

ToshLLM returns Anthropic message blocks, token usage and stop reasons. Add "stream": true to receive server-sent events. Tool calls are available when the selected model, chat template and Agent tools setting support them.

Anthropic SDK

Use the official Python client locally.

The official Anthropic SDK accepts a custom base_url. The model can be any local model ID returned by ToshLLM, not only a Claude model name.

from anthropic import Anthropic

client = Anthropic(
    base_url="http://127.0.0.1:8080",
    api_key="YOUR_TOSHLLM_KEY",
)

message = client.messages.create(
    model="MODEL_ID",
    max_tokens=256,
    messages=[
        {"role": "user", "content": "Write a concise project summary."}
    ],
)

print(message.content[0].text)

If API protection is disabled but the SDK requires a key, use toshllm-local. For a streamed response, use client.messages.stream(...).

Anthropic Python SDK reference ↗
Token counting

Measure a request before generation.

The Anthropic-compatible token-count endpoint accepts a Messages API request and returns its input-token count using the tokenizer for the selected local model.

curl http://127.0.0.1:8080/v1/messages/count_tokens \
  -H "Content-Type: application/json" \
  -H "anthropic-version: 2023-06-01" \
  -H "x-api-key: YOUR_TOSHLLM_KEY" \
  -d '{
    "model": "MODEL_ID",
    "messages": [{"role": "user", "content": "How many tokens are here?"}]
  }'

The result is local and model-specific. Do not compare it with a token count produced by a different cloud model or tokenizer.

Anthropic token-counting reference ↗
Compatibility boundary

Know exactly what is compatible.

MessagesText messages, system instructions and structured content through /v1/messages.
StreamingAnthropic-style server-sent events for text and tool-call output.
Tool useClient-defined tools when the local model and its chat template support reliable function calling.
Token countingLocal input-token measurement through /v1/messages/count_tokens.
Protocol compatibility is not cloud feature parity.

ToshLLM implements the Anthropic Messages surface used by local clients. Anthropic-hosted features such as cloud files, batch jobs, prompt caching and provider billing are not local ToshLLM services.

First request

Send a chat completion.

curl http://127.0.0.1:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "your-model",
    "messages": [
      {"role": "user", "content": "Explain Metal in one paragraph."}
    ]
  }'

Use the model identifier returned by GET /v1/models. In single-model mode, many compatible clients can use their default model field; router mode uses stable aliases derived from downloaded filenames.

Authentication

Protect clients beyond one process.

Enable Protect the API with a key in Settings to generate a key stored in macOS Keychain. ToshLLM then protects inference, embeddings, tools and other non-public operations. Health and model discovery remain readable so clients can detect the server and populate model selectors.

curl http://127.0.0.1:8080/v1/models \
  -H "Authorization: Bearer YOUR_TOSHLLM_KEY"

The in-app chat supplies the key automatically. Copy it only into clients you trust.

Discovery metadata remains public.

GET /health, GET /v1/health, GET /models and GET /v1/models do not require the API key. Do not expose the listener to an untrusted network if model names or server availability are sensitive.

Network and routing

Expose only what you intend to expose.

Enabling Discoverable on local network changes the listener from 127.0.0.1 to 0.0.0.0 and advertises a ToshLLM API service through Bonjour. Use it only on a trusted network and enable API-key protection first.

Router mode publishes multiple downloaded models behind the same endpoint and automatically loads them as requests arrive. The configured model limit controls how many remain loaded at once.

Open the integrations guide for Claude Code, Anthropic SDKs, VS Code, Zed, OpenCode, Aider and other clients.

Streaming and model discovery

Let the server describe itself.

Query GET /v1/models instead of guessing a filename. In router mode, send the returned alias in the model field so ToshLLM can load the correct preset. Compatible clients can request streamed chat completions using "stream": true.

curl http://127.0.0.1:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer YOUR_TOSHLLM_KEY" \
  -d '{
    "model": "MODEL_ID",
    "stream": true,
    "messages": [{"role": "user", "content": "Hello"}]
  }'
Runtime modes

Match the endpoint to the job.

Single modelBest when one model should remain loaded and every client uses the same runtime configuration.
RouterPublishes multiple downloaded models and loads the requested alias, subject to the configured model limit.
EmbeddingsEnables /v1/embeddings for a compatible embedding model and client.
Agent toolsEnables the tool runtime for models that reliably support function calling.
Endpoint reference

Use the smallest interface for the task.

MethodPathPurpose
GET/healthCheck whether the server is ready.
GET/v1/modelsDiscover model IDs and router aliases.
POST/v1/chat/completionsOpenAI-style chat with optional streaming and tools.
POST/v1/responsesOpenAI Responses-compatible generation.
POST/v1/completionsPlain text completion without a chat template.
POST/v1/messagesAnthropic Messages-compatible generation.
POST/v1/messages/count_tokensCount Anthropic message input tokens.
POST/v1/embeddingsCreate vectors when embeddings mode is active.
POST/v1/rerankScore and reorder documents with a compatible model.
POST/tokenizeConvert text to token IDs for the active model.
POST/detokenizeConvert token IDs back to text.
GET/propsInspect context size, chat template and model properties.
GET/metricsRead runtime metrics when metrics are enabled.

LoRA, slot persistence, router management and resumable-stream endpoints also exist for advanced runtime control. They are lower-level engine interfaces and can change more readily than the compatibility endpoints above.

For teams and individuals

Make ToshLLM work for your team.

Need a tailored deployment, help choosing models, or guidance for a fleet of Intel Macs? Tell us what you are building and we will get back to you personally.

Found a reproducible bug? A public GitHub issue helps everyone follow the fix. Open an issue ↗
hello@toshllm.com

CONTACT / TOSHLLM

Your message goes directly to ToshLLM. Please do not include passwords, API keys, or private logs.