Two familiar APIs on one local server.
When the server is running, ToshLLM exposes OpenAI-compatible and Anthropic-compatible endpoints at http://127.0.0.1:8080 by default. The host remains limited to the Mac until local-network discovery is enabled.
POST /v1/chat/completionsPOST /v1/responsesPOST /v1/messagesPOST /v1/messages/count_tokensGET /v1/modelsPOST /v1/embeddings when embeddings mode is enabled.Use clients built for the newer response format.
The current ToshLLM engine accepts POST /v1/responses in addition to Chat Completions. This gives newer OpenAI-compatible clients a direct local route without translating them through a cloud service.
curl http://127.0.0.1:8080/v1/responses \
-H "Content-Type: application/json" \
-H "Authorization: Bearer YOUR_TOSHLLM_KEY" \
-d '{
"model": "MODEL_ID",
"input": "Summarize the benefits of local inference.",
"stream": true
}'Use the exact local model ID. Support covers the response surface implemented by the bundled engine, not cloud-only OpenAI services.
Send a native Anthropic-style message.
Use the server root as the Anthropic base URL. Do not append /v1 when configuring an Anthropic SDK because the client adds /v1/messages itself.
curl http://127.0.0.1:8080/v1/messages \
-H "Content-Type: application/json" \
-H "anthropic-version: 2023-06-01" \
-H "x-api-key: YOUR_TOSHLLM_KEY" \
-d '{
"model": "MODEL_ID",
"max_tokens": 256,
"messages": [
{"role": "user", "content": "Explain Metal in one paragraph."}
]
}'ToshLLM returns Anthropic message blocks, token usage and stop reasons. Add "stream": true to receive server-sent events. Tool calls are available when the selected model, chat template and Agent tools setting support them.
Use the official Python client locally.
The official Anthropic SDK accepts a custom base_url. The model can be any local model ID returned by ToshLLM, not only a Claude model name.
from anthropic import Anthropic
client = Anthropic(
base_url="http://127.0.0.1:8080",
api_key="YOUR_TOSHLLM_KEY",
)
message = client.messages.create(
model="MODEL_ID",
max_tokens=256,
messages=[
{"role": "user", "content": "Write a concise project summary."}
],
)
print(message.content[0].text)If API protection is disabled but the SDK requires a key, use toshllm-local. For a streamed response, use client.messages.stream(...).
Measure a request before generation.
The Anthropic-compatible token-count endpoint accepts a Messages API request and returns its input-token count using the tokenizer for the selected local model.
curl http://127.0.0.1:8080/v1/messages/count_tokens \
-H "Content-Type: application/json" \
-H "anthropic-version: 2023-06-01" \
-H "x-api-key: YOUR_TOSHLLM_KEY" \
-d '{
"model": "MODEL_ID",
"messages": [{"role": "user", "content": "How many tokens are here?"}]
}'The result is local and model-specific. Do not compare it with a token count produced by a different cloud model or tokenizer.
Anthropic token-counting reference ↗Know exactly what is compatible.
/v1/messages./v1/messages/count_tokens.ToshLLM implements the Anthropic Messages surface used by local clients. Anthropic-hosted features such as cloud files, batch jobs, prompt caching and provider billing are not local ToshLLM services.
Send a chat completion.
curl http://127.0.0.1:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "your-model",
"messages": [
{"role": "user", "content": "Explain Metal in one paragraph."}
]
}'Use the model identifier returned by GET /v1/models. In single-model mode, many compatible clients can use their default model field; router mode uses stable aliases derived from downloaded filenames.
Protect clients beyond one process.
Enable Protect the API with a key in Settings to generate a key stored in macOS Keychain. ToshLLM then protects inference, embeddings, tools and other non-public operations. Health and model discovery remain readable so clients can detect the server and populate model selectors.
curl http://127.0.0.1:8080/v1/models \
-H "Authorization: Bearer YOUR_TOSHLLM_KEY"The in-app chat supplies the key automatically. Copy it only into clients you trust.
GET /health, GET /v1/health, GET /models and GET /v1/models do not require the API key. Do not expose the listener to an untrusted network if model names or server availability are sensitive.
Expose only what you intend to expose.
Enabling Discoverable on local network changes the listener from 127.0.0.1 to 0.0.0.0 and advertises a ToshLLM API service through Bonjour. Use it only on a trusted network and enable API-key protection first.
Router mode publishes multiple downloaded models behind the same endpoint and automatically loads them as requests arrive. The configured model limit controls how many remain loaded at once.
Open the integrations guide for Claude Code, Anthropic SDKs, VS Code, Zed, OpenCode, Aider and other clients.
Let the server describe itself.
Query GET /v1/models instead of guessing a filename. In router mode, send the returned alias in the model field so ToshLLM can load the correct preset. Compatible clients can request streamed chat completions using "stream": true.
curl http://127.0.0.1:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer YOUR_TOSHLLM_KEY" \
-d '{
"model": "MODEL_ID",
"stream": true,
"messages": [{"role": "user", "content": "Hello"}]
}'Match the endpoint to the job.
/v1/embeddings for a compatible embedding model and client.Use the smallest interface for the task.
GET/healthCheck whether the server is ready.GET/v1/modelsDiscover model IDs and router aliases.POST/v1/chat/completionsOpenAI-style chat with optional streaming and tools.POST/v1/responsesOpenAI Responses-compatible generation.POST/v1/completionsPlain text completion without a chat template.POST/v1/messagesAnthropic Messages-compatible generation.POST/v1/messages/count_tokensCount Anthropic message input tokens.POST/v1/embeddingsCreate vectors when embeddings mode is active.POST/v1/rerankScore and reorder documents with a compatible model.POST/tokenizeConvert text to token IDs for the active model.POST/detokenizeConvert token IDs back to text.GET/propsInspect context size, chat template and model properties.GET/metricsRead runtime metrics when metrics are enabled.LoRA, slot persistence, router management and resumable-stream endpoints also exist for advanced runtime control. They are lower-level engine interfaces and can change more readily than the compatibility endpoints above.