Skip to content

Two design improvements for local AI deployments — trusted HTTP endpoints and clearer Mode definitions #112

Description

@unfall103-debug

Environment

  • SLM version: 3.8.14
  • Deployment: LXC/Docker containers on private LAN
  • LLM: local OpenAI-compatible API (llama.cpp)
  • Embedding: local OpenAI-compatible API (Qwen3-Embedding)
  • Reranker: local rerank API (llama.cpp)
  • All services communicate through internal HTTP endpoints

Issue 1: Trusted private LAN HTTP endpoints rejected for reranker

Symptom

Running:

slm soft-prompts

returns:

Remote reranker not started — retrieval.cross_encoder_endpoint must use HTTPS for non-loopback hosts.
Plain HTTP is allowed only for localhost/loopback.
Reranking is DISABLED.

Manual verification

The reranker endpoint itself works correctly:

curl -X POST http://192.168.x.x:8041/v1/rerank \
  -H "Content-Type: application/json" \
  -d '{"query":"test","documents":["hello world"]}'

The endpoint returns valid reranking scores.

Root cause

The reranker service is not the problem. The issue is the security validation layer rejecting all non-loopback HTTP endpoints.

In local AI deployments, internal HTTP is commonly used between trusted services:

LXC containers

Docker containers

Homelab servers

llama.cpp

vLLM

OpenAI-compatible local inference servers


These endpoints are normally not exposed to the public internet.

Suggested improvement

Keep HTTPS as the default security policy, but provide an explicit opt-in mechanism for trusted local networks:

Examples:

allow_trusted_http_endpoints: true

CIDR whitelist for private networks

Per-service override:

embedding.allow_http

reranker.allow_http



This preserves security defaults while supporting advanced local deployments.


---

Issue 2: Mode definitions should describe user experience, not vendors

The current Mode design appears to mix user workflow and backend technology.

For example, Ollama is currently treated as the definition of local AI, but in advanced local AI deployments many users choose different backends:

llama.cpp

vLLM

OpenAI-compatible inference servers

custom model stacks


Reasons include:

multi-GPU support

deployment flexibility

hardware optimization

wider model selection


The following classification is based on the user deployment experience, not the technical implementation.


---

Proposed Mode redesign

A — Mathematical Mode

No LLM dependency

Pure mathematical / algorithmic memory processing

Maximum privacy

Lowest resource usage

No external model required


B — Default / Easy Local Mode

One-click local installation experience

SLM manages local dependencies automatically

Ollama + recommended models as the default implementation

Designed for users who want a simple local AI memory system


C — Custom Mode

User-controlled AI infrastructure

User selects Provider / Endpoint / Model

OpenAI-compatible APIs

llama.cpp

vLLM

Custom Ollama deployments

Local API servers

Cloud APIs


Designed for advanced users who already manage their own AI stack.


---

Why these changes matter

Both issues are related to the same long-term direction:

SLM should become a true "bring your own models" memory system.

Beginners should have a simple Ollama-based path.

Advanced users should be able to connect their own llama.cpp/vLLM infrastructure.

Non-English users should be able to choose specialized embedding and reranking models.

The system should avoid vendor assumptions while keeping safe defaults for normal users.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions