Open WebUI

Self-hosted AI chat interface for local models

AI & LLM BSD-3-Clause beginner ★ 153,674 stars

What is Open WebUI?

Open WebUI is an extensible, self-hosted web interface for running AI models locally. It works with Ollama and any OpenAI-compatible API, and supports RAG document chat, model management, image generation and multi-user access. Everything runs on your own hardware.

Best for

Anyone running Ollama or local models who wants a proper chat UI

Why choose Open WebUI

Open WebUI gives local language models the interface they deserve. It connects to Ollama and any OpenAI-compatible endpoint, so the same chat UI can talk to a model running on your own GPU or to a hosted API, and everything about the interaction stays on your machine. It supports document chat through retrieval, model management, image generation through connected backends, prompt presets and multi-user access with separate conversations. For anyone running local models, it converts a command-line experience into something a household or small team can actually use, with conversation history and the ability to share or export what matters. Because the models themselves run locally, nothing you type leaves your network, which is frequently the entire reason for doing this.

Replaces

  • ChatGPT
  • Claude
  • Google Gemini

Key features

  • Works with Ollama and OpenAI-compatible APIs
  • RAG over your own documents
  • Multi-user with role management
  • Model builder and fine-tuning support

What to watch out for

The quality of the experience is dominated by the model and the hardware, not by the interface. A small model on a CPU produces responses that are slow and frequently wrong, and no amount of good UI fixes that; expectations should be set by the hardware. The project moves quickly, so features and configuration options change between releases. Connected features — image generation, external APIs, web search — involve additional services and sometimes keys, and enabling them expands what the deployment touches. If you expose it beyond your network, its user management becomes a real authentication surface, and an open instance giving access to a model with system-level tool permissions is a genuine risk. Retrieval over documents works best when the documents are chunked and curated rather than dumped in wholesale.

How to deploy

  • Docker
  • Kubernetes
  • pip

Getting started

Get a model running and responding before installing an interface, so that any problem has an obvious place to look. Deploy it with persistent storage for its database of users and conversations, and include that in your backup routine. Set the first admin account immediately and disable open registration, because an instance reachable by others with registration open is a way to give strangers access to your hardware. Start with a single model and one conversation before enabling retrieval, image generation or web search, so each addition can be attributed when something misbehaves. Size your expectations against the hardware, and test whether the latency is acceptable to the people who will actually use it.

Typical setup

A model runs and responds before the interface is installed, so any problem has an obvious place to look. It is deployed with persistent storage for its database of users and conversations, and that storage is backed up. The first admin account is created immediately and open registration is disabled, because a reachable instance with registration enabled hands strangers access to the hardware. One model and one conversation are established before retrieval, image generation or web search is enabled, so each addition can be attributed when something misbehaves. Expectations are sized against the hardware, and latency is tested by the people who will actually use it.

Who should look elsewhere

Do not use it if you have no hardware capable of running a model at a usable speed, because the interface will not compensate and a hosted API is the honest choice for the same purpose. Avoid it if you expect frontier-model quality from local models, since open weights at runnable sizes are simply not equivalent. And if you are not prepared to keep its user management and exposure properly configured, running a chat interface that reaches into your machine is a risk worth weighing before you start.

Project health

  • GitHub stars: 153,674
  • Last code push: 2026-10-01
  • Open issues: 296
  • Status: actively developed

Figures pulled from the GitHub API and refreshed periodically.

Open WebUI as an alternative

More in AI & LLM