Engineering

Local Document Search with Open WebUI

Ask questions of your own documents with every step on your own hardware. Marigold serves the chat and embedding models; Open WebUI handles upload, chunking, storage and retrieval. This tutorial starts the simple-rag package, checks that the embedding settings line up, and shows where each piece of data lives.

This assumes the setup guide is done and the platform is running.

Start it

The simple-rag package has this wired up: a chat model and an embedding model in its models.yaml, webui in its compose_files, and Open WebUI’s RAG_EMBEDDING_MODEL pointed at the same embedding model through its [environment] table:

marigold package create marigold-examples/simple-rag -o /tmp
marigold package install /tmp/simple-rag-<version>.tar.gz
marigold cache populate simple-rag
marigold application start simple-rag

package create prints the archive path. cache populate downloads both models. application start adds Open WebUI to the platform, on the applications network as a client of the API, then runs the package’s application, which checks that both models are in the catalogue. marigold application status simple-rag shows exit code 0 when they are.

Open WebUI keeps its documents, vector store and chat history under the cache directory, in data/webui. It runs with the platform: marigold application stop simple-rag leaves it running, and marigold platform stop stops it.

Once it’s up, http://localhost:3000’s embedding settings (Admin Panel -> Settings -> Documents) should already show the right model – worth a glance, since if this is wrong, uploads still appear to succeed and chat still looks normal; the failure is silent rather than an error you’d notice.

The embedding setting

The package forwards one variable to Open WebUI:

[environment]
RAG_EMBEDDING_MODEL = "sentence-transformers/all-minilm-l6-v2"

Marigold’s CLI passes [environment] values through to the other services in the stack without interpreting them. To change the embedding model, change it here and in the package’s models.yaml together, then rebuild, install and populate the package.

Ask a question

Open http://localhost:3000, attach a document to a chat, and ask about something only that document contains. A correct answer that cites the file confirms the full path – embedding, storage, retrieval, generation – ran locally.

A general question is weak evidence: the model may answer from its training. The simple-rag example ships three invented documents and a question set that separates retrieval from recall, including one question whose correct answer is “I don’t know”.

Where each piece lives

Piece Served by Stored in
Chat model Marigold worker data/models
Embedding model Marigold worker data/models
Uploaded documents and chunks Open WebUI data/webui
Vector store Open WebUI’s local ChromaDB data/webui
Chat history Open WebUI data/webui

Marigold hosts no vector database. It supplies embedding vectors on request and generates the answer; everything else in the retrieval path belongs to Open WebUI.

Change the chat model

The chat model is an entry in the package’s models.yaml. Adding a new model covers the edit, rebuild and populate loop.

Stop

marigold application stop simple-rag    # the application container
marigold platform stop --applications   # everything, Open WebUI included

Try it yourself.

All example code is in the repo. Clone it and run this tutorial locally.