Architecting Forgejo Storage: Managing the Database and Filesystem Split
Explore the architectural split between metadata databases and Git filesystems in Forgejo, including strategies to prevent orphaned repositories and manage inode exhaustion.
24 Jul 2025, 21:01 UTC

The Challenge of Decoupled Git Storage
Forgejo separates repository metadata from the actual Git object graph. While this allows the application to scale the web interface independently of the data, it introduces a synchronization risk: the database (PostgreSQL, MySQL, or SQLite) believes a repository exists, but the filesystem may be missing the directory, or vice versa. This "split-brain" state can lead to 404 errors for users or orphaned directories that consume disk space without being tracked.
The Minimal Design for Data Consistency
The most stable deployment uses a relational database for metadata and a high-performance local or network filesystem for the .git directories. The application interacts with the filesystem primarily through the Git binary, avoiding the loading of large object graphs into the application's memory heap.
To maintain consistency, Forgejo uses a repository abstraction. When a user creates a project, the system follows a strict sequence: it first allocates the directory on disk and then commits the metadata to the database. If the database commit fails, the system attempts to roll back the filesystem change.
Trust and Data Boundaries
Forgejo establishes a hard boundary between the web-facing API and the filesystem. No user-supplied input should ever directly map to a filesystem path. Access is mediated through two primary channels:
- HTTP/SSH Git Access: Authenticated via the database; the system then invokes the Git binary with a specific system user (typically
git) to interact with the filesystem. - Web Interface: The application reads metadata from the database and only queries the filesystem for specific data (like a file tree or a commit hash) using controlled API calls.
Operational Configuration and Verification
When deploying, ensure the git user has explicit ownership of the data root. A common failure point is running the Forgejo binary as root, which creates files that the Git SSH service cannot later modify.
Verification Check: To verify that your database and filesystem are synchronized, run a manual cross-reference check. On the server, identify a repository root path from the database and verify its existence on disk.
# 1. Connect to your database (example using psql)
psql -U forgejo -d forgejo
# 2. Query the root path of a specific repository
SELECT root FROM repository WHERE name = 'my-project';
# 3. Verify the path exists on the filesystem (run as system user)
ls -ld /var/lib/forgejo/git/repositories/user/my-project.git
Expected Result: The path returned by the SQL query must match the actual directory structure. If the database returns a path that does not exist, you have an orphaned record.
Failure Modes and Risk Mitigation
| Failure Mode | Cause | Impact | Mitigation |
|---|---|---|---|
| Orphaned Repositories | Interrupted migration or manual rm -rf |
DB entries point to non-existent disks | Use Forgejo's built-in cleanup tools; avoid manual disk edits |
| Inode Exhaustion | High volume of small Git objects | Disk reports space available, but cannot create files | Use filesystems like XFS or ZFS; monitor df -i |
| NFS Latency | Network filesystem bottlenecks | Slow git push and timeouts |
Use local SSDs or high-throughput NVMe-over-Fabric |
Conditions for Redesign
The current decoupled architecture is sufficient for most engineering teams. However, you should consider a different storage strategy (such as moving to a specialized distributed storage backend) if the following conditions are met:
- Extreme Horizontal Scaling: If the web tier grows beyond a few dozen nodes, the latency of a shared network filesystem (NFS) for Git objects becomes the primary bottleneck.
- Strict Atomic Requirements: If your workflow requires absolute atomicity between metadata updates and object writes across multiple geographic regions.
Rollback Procedure
If a filesystem migration leads to desynchronization, do not attempt to manually rename directories. Instead:
- Restore the database to the snapshot taken immediately before the migration.
- Restore the filesystem from the corresponding snapshot.
- Verify the
rootpaths in therepositorytable match the restored disk paths before restarting the service.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.