/connect to add a model server such as Ollama, LM Studio, vLLM, or llama.cpp. Start and configure the server before adding it to Embedder.
If Local models is unavailable, ask your Embedder administrator whether it can be enabled for your environment.
Add the endpoint
Run/connect, choose Local models, and enter:
- Name: a short label for the endpoint.
- Server: the model server you use.
- Base URL: its reachable HTTP or HTTPS address, with or without
/v1. - Model: the exact model identifier, or leave blank for discovery.
- Context window: the limit configured on that server.
- API key: the credential required by the server, if any.
/model.
For a server on another machine, use its reachable address. localhost refers to the machine running Embedder.
Match the context window
Use the context limit actually configured on the model server. Declaring a larger limit in Embedder does not increase the server’s capacity and can cause failures or loss of earlier context. For Ollama, setOLLAMA_CONTEXT_LENGTH in the environment that starts the server and restart it after changing that value.
Use advanced options
For options outside the form, edit~/.embedder/local-endpoints.json. For example:
enabled, discover, headers, timeoutMs, and dropParams for request fields the server does not support.
Use environment references such as ${LOCAL_LLM_KEY} for credentials instead of placing them in a shared configuration file.
Set a project default
Use the endpoint name and exact model ID:.embedder/models.json
Troubleshoot
- Connection test fails: confirm the server is running and the URL is reachable from the Embedder host.
- Model is missing: enter its exact ID if the server cannot list models.
- Earlier context is lost: match the declared window to the server’s configured limit.
- Request field is rejected: use the server’s supported configuration, or add that field to
dropParams.

