Langfuse
Observability and evaluation for LLM apps
What is Langfuse?
Langfuse is an open-source platform for tracing, evaluating and monitoring LLM applications. It captures traces of model calls, supports prompt management and datasets, and helps teams measure quality of AI features over time.
Best for
Developers shipping LLM features who need to debug and evaluate them
Why choose Langfuse
Langfuse answers the question every team building on top of language models eventually asks: what is actually happening in production? It records each request, what the model returned, how long it took, what it cost and how your prompts have changed over time. For anyone whose application's output is non-deterministic, that visibility is not a luxury — it is the only way to debug a bad answer or justify the monthly bill. Self-hosting keeps the prompts and the user data inside your own perimeter.
Replaces
- LangSmith
- Weights & Biases
- Helicone
Key features
- Full tracing of LLM calls
- Prompt management and versioning
- Evaluation datasets and scoring
- Cost and latency tracking
What to watch out for
Observability always has an adoption cost. You must instrument your code, and how much you get out depends on how much you put in: add tracing to a handful of calls and you will have a partial picture that can mislead. It is also another service to run, with its own database — ClickHouse and Postgres in the full setup — which is meaningful infrastructure. And costs scale with the volume you log, because every trace is stored data, so decide your retention policy before you turn on verbose logging.
How to deploy
- Docker Compose
- Kubernetes
Getting started
Deploy the full Compose stack on a machine with real memory, since the analytics database is not light. Instrument one endpoint first, send a few test traces and confirm they appear before rolling it out across the application. Set the retention window deliberately on day one; a tracing store that grows unbounded is a surprisingly common surprise. If prompts contain sensitive information, configure redaction before any real traffic reaches it, not after.
Typical setup
The full stack wants meaningful hardware — an analytical database, Postgres, Redis and the application itself — so it usually lands on a mid-sized VPS or a dedicated machine on an internal network. Teams instrument a development environment first, confirm traces look right, then roll the SDK out to production. The operational decisions that matter are retention and redaction, both of which are easiest to set before real traffic starts flowing.
Who should look elsewhere
Overkill if you make a handful of model calls a day and have no intention of debugging them systematically. It is also a mismatch for anyone without somewhere to run an analytical database, and for teams whose prompts contain data they cannot store at all — observability is only useful when you log what happened, and sometimes that is the problem.
Project health
- GitHub stars: 35,243
- Last code push: 2026-09-30
- Open issues: 984
- Status: actively developed
Figures pulled from the GitHub API and refreshed periodically.