SeaweedFS

Fast distributed object store and file system

Databases & Storage Apache-2.0 Advanced ★ 35,182 stars

What is SeaweedFS?

SeaweedFS is an object store, S3 gateway and distributed file system designed to handle huge numbers of small files efficiently, which is where naive filesystem storage performs worst.

Best for

Millions of small files without a filesystem melting

Why choose SeaweedFS

SeaweedFS is built around a specific observation: general-purpose filesystems and object stores handle enormous numbers of small files badly. It stores file metadata centrally while distributing file content across volume servers, which keeps lookups fast even when the count runs into the millions. It speaks the S3 API, can present itself as a distributed filesystem, and is designed to scale out by adding volume servers rather than by growing one machine. For photo libraries, document archives, media pipelines and any workload where the object count is the dominant problem, it addresses the bottleneck the mainstream options quietly ignore.

Replaces

  • AWS S3
  • MinIO
  • Ceph

Key features

  • S3 API alongside its own protocol
  • Efficient handling of very large file counts
  • Filer for POSIX-style access
  • Erasure coding and replication

What to watch out for

It is not a drop-in replacement for an all-purpose storage system, and the separation between master, volume and filer components is a distributed architecture you have to operate and understand — the simple single-node mode is not what you would run when the data matters. Documentation is uneven and the community is smaller than the alternatives, so diagnosing an unusual failure takes longer. Using it as a general network filesystem exposes you to POSIX semantics it does not fully emulate, and applications that expect real file locking or random writes may misbehave. Backup and recovery are your design decisions rather than a built-in feature.

How to deploy

  • Docker Compose for master, volume and filer
  • Choose replication level based on node count
  • Enable the S3 gateway for object clients

Getting started

Decide which interface you actually need first — S3, a mounted filesystem or the raw API — because that determines which components to run. Start on a single machine to learn the components, then plan the distributed layout with separate hosts for masters and volumes before storing anything you care about. Configure replication explicitly for each collection rather than accepting a default and assuming it is safe. Test that a volume server can be lost and the data still served. Point one real workload at it and watch the metadata store, since that is where growth concentrates. Document the topology, because recovery depends entirely on it.

Typical setup

The interface required — S3, a mounted filesystem, or the raw API — is chosen first, which determines which components run. A single machine is used to learn the components, then masters and volumes are separated across hosts before anything important is stored. Replication is configured explicitly per collection rather than left at a default. Losing a volume server is tested rather than feared. The metadata store is watched because growth concentrates there, and the topology is written down so a rebuild is possible.

Who should look elsewhere

Do not use it if your workload is a modest number of large files, where MinIO or a plain filesystem is simpler and better supported. Avoid it if you need POSIX semantics faithfully, since partial emulation causes problems that look like application bugs. And if nothing in your stack actually suffers from small-file counts, you would be adopting a distributed system to solve a problem you do not have.

Project health

  • GitHub stars: 35,182
  • Last code push: 2026-10-02
  • Open issues: 782
  • Status: actively developed

Figures pulled from the GitHub API and refreshed periodically.

More in Databases & Storage