Comparison

An Ollama alternative for self-hosted inference.

Open-weight models on your own hardware, arranged as a platform.

Ollama runs open models on one machine, from a desktop app, a terminal or a coding agent. Marigold is an open source Ollama alternative arranged as a platform: a shared model cache, a queue, a worker, and applications that call one typed inference API, with text, image and audio models served side by side.

Both run open-weight models with no request leaving your hardware once the models are downloaded. The choice turns on what you serve, and how many applications call it.

Criterion Ollama Marigold
Model source Ollama model library; imports GGUF and Safetensors HuggingFace repositories declared in a package's models.yaml, loaded with Transformers; optional 4-bit loading
Model types Chat, vision, embeddings, reasoning Instruct, img2txt, text and image embedding, txt2img, tts, txt2audio, depth, img2mask, text, image and image-text evals
OpenAI-compatible API /v1/chat/completions, /v1/completions, /v1/embeddings, /v1/models, /v1/responses; streaming and tools /v1/chat/completions, /v1/embeddings, /v1/models; no token-level streaming; one tool call per assistant turn
Other APIs Native Ollama API; Anthropic-compatible clients Typed submit-and-poll API per model type
Request handling Request and streamed response Jobs queued in Postgres, run by a worker; identical requests return the stored result
Multi-step work In client code Declarative workflows: parallel steps, branching on outputs, per-step audit trail
Hosted option Cloud models, run without a local GPU None; runs where you deploy it
Runs as App or local server, port 11434 Docker Compose: API on port 8000, CPU or NVIDIA GPU worker

Ollama details from docs.ollama.com, checked 1 October 2026.

Three serving stacks, one 6GB NVIDIA card (about GBP 170), one model: Qwen2.5-1.5B-Instruct. 86 prompts at four concurrency levels, three repeats, 1,032 calls per stack.

Stack Tokens / sec Time to first token Mean duration
vLLM54.770.063s1.449s
Ollama39.031.528s2.196s
Marigold11.58not measured (no token streaming)8.015s

Marigold was the slowest per request: it has no request batching. The full vLLM vs Ollama benchmark covers a 14.7x speedup from one attention configuration change, and a VRAM problem: every stack allocates GPU memory greedily, so whichever server starts second on a shared card fails to allocate.

One developer, one chat model

Ollama installs as an app and runs a model in one command. It returned responses faster than Marigold in our test.

GGUF and quantised weights

Ollama imports GGUF files prepared with llama.cpp tooling. Marigold loads HuggingFace repositories through Transformers.

Clients that need the full chat API

Coding agents and desktop clients that rely on token streaming, parallel tool calls, the Responses API or Anthropic-style clients work against Ollama today.

No local GPU

Ollama's cloud models run without one. Marigold has no hosted service; it runs on CPU or on your own GPU.

More than chat

Speech, image generation, segmentation, depth and scoring models sit behind the same API as the chat model, loaded from one shared weight cache.

Pipelines across models

A workflow chains models -- extract, classify, summarise, convert to speech -- with parallel steps, branching on output values and a per-step audit trail.

Several applications on one host

Each package declares the models it needs. The platform downloads each model once; applications run in their own containers and reach models only through the API.

Air-gapped and regulated hosts

The worker makes no network calls and loads only from the cache. All state sits under one directory. What leaves the machine.

pip install bayis-marigold
git clone https://github.com/bayinfosys/marigold-examples
marigold cache init
marigold platform start
marigold package create marigold-examples/chat -o /tmp
marigold package install /tmp/chat-<version>.tar.gz
marigold cache populate chat
marigold application start chat

An OpenAI-SDK client moves across by changing its base URL from http://localhost:11434/v1 to http://localhost:8000/v1. On a single GPU, stop Ollama before starting the Marigold worker: both allocate VRAM greedily. The local LLM server setup guide covers each step.

Run it next to what you have.

Open source. Clone it, run it on your own hardware.