Comparison
An Ollama alternative for self-hosted inference.
Open-weight models on your own hardware, arranged as a platform.
Ollama runs open models on one machine, from a desktop app, a terminal or a coding agent. Marigold is an open source Ollama alternative arranged as a platform: a shared model cache, a queue, a worker, and applications that call one typed inference API, with text, image and audio models served side by side.
Both run open-weight models with no request leaving your hardware once the models are downloaded. The choice turns on what you serve, and how many applications call it.
Side by side
| Criterion | Ollama | Marigold |
|---|---|---|
| Model source | Ollama model library; imports GGUF and Safetensors | HuggingFace repositories declared in a package's models.yaml, loaded with Transformers; optional 4-bit loading |
| Model types | Chat, vision, embeddings, reasoning | Instruct, img2txt, text and image embedding, txt2img, tts, txt2audio, depth, img2mask, text, image and image-text evals |
| OpenAI-compatible API | /v1/chat/completions, /v1/completions, /v1/embeddings, /v1/models, /v1/responses; streaming and tools |
/v1/chat/completions, /v1/embeddings, /v1/models; no token-level streaming; one tool call per assistant turn |
| Other APIs | Native Ollama API; Anthropic-compatible clients | Typed submit-and-poll API per model type |
| Request handling | Request and streamed response | Jobs queued in Postgres, run by a worker; identical requests return the stored result |
| Multi-step work | In client code | Declarative workflows: parallel steps, branching on outputs, per-step audit trail |
| Hosted option | Cloud models, run without a local GPU | None; runs where you deploy it |
| Runs as | App or local server, port 11434 | Docker Compose: API on port 8000, CPU or NVIDIA GPU worker |
Ollama details from docs.ollama.com, checked 1 October 2026.
Measured on one consumer GPU
Three serving stacks, one 6GB NVIDIA card (about GBP 170), one model: Qwen2.5-1.5B-Instruct. 86 prompts at four concurrency levels, three repeats, 1,032 calls per stack.
| Stack | Tokens / sec | Time to first token | Mean duration |
|---|---|---|---|
| vLLM | 54.77 | 0.063s | 1.449s |
| Ollama | 39.03 | 1.528s | 2.196s |
| Marigold | 11.58 | not measured (no token streaming) | 8.015s |
Marigold was the slowest per request: it has no request batching. The full vLLM vs Ollama benchmark covers a 14.7x speedup from one attention configuration change, and a VRAM problem: every stack allocates GPU memory greedily, so whichever server starts second on a shared card fails to allocate.
Where Ollama is the right tool
One developer, one chat model
Ollama installs as an app and runs a model in one command. It returned responses faster than Marigold in our test.
GGUF and quantised weights
Ollama imports GGUF files prepared with llama.cpp tooling. Marigold loads HuggingFace repositories through Transformers.
Clients that need the full chat API
Coding agents and desktop clients that rely on token streaming, parallel tool calls, the Responses API or Anthropic-style clients work against Ollama today.
No local GPU
Ollama's cloud models run without one. Marigold has no hosted service; it runs on CPU or on your own GPU.
Where Marigold fits
More than chat
Speech, image generation, segmentation, depth and scoring models sit behind the same API as the chat model, loaded from one shared weight cache.
Pipelines across models
A workflow chains models -- extract, classify, summarise, convert to speech -- with parallel steps, branching on output values and a per-step audit trail.
Several applications on one host
Each package declares the models it needs. The platform downloads each model once; applications run in their own containers and reach models only through the API.
Air-gapped and regulated hosts
The worker makes no network calls and loads only from the cache. All state sits under one directory. What leaves the machine.
Trying it next to Ollama
pip install bayis-marigold
git clone https://github.com/bayinfosys/marigold-examples
marigold cache init
marigold platform start
marigold package create marigold-examples/chat -o /tmp
marigold package install /tmp/chat-<version>.tar.gz
marigold cache populate chat
marigold application start chat
An OpenAI-SDK client moves across by changing its base URL from
http://localhost:11434/v1 to
http://localhost:8000/v1. On a single GPU, stop Ollama
before starting the Marigold worker: both allocate VRAM greedily.
The local LLM server setup guide
covers each step.
Run it next to what you have.
Open source. Clone it, run it on your own hardware.