Architecting Weblate: Managing Asynchronous Translation and Data Consistency
Discover how Weblate separates UI and heavy file parsing with Celery, Redis, and PostgreSQL, and learn operational checks to keep the translation pipeline healthy.
20 Jan 2026, 15:41 UTC

The Problem: Heavy File Parsing in a Responsive UI
When a user pushes a repository containing thousands of translation files, processing those files inside a single web request leads to timeouts and a frozen interface. Weblate solves this by decoupling the UI from the heavy lifting.
Decoupled Architecture Overview
- Django Web Layer: Handles authentication, API requests, and the front‑end state. It never parses files.
- Redis Message Broker: Queues tasks generated by the web layer and hands them to workers.
- Celery Workers: Consume tasks, parse source files (PO, JSON, XLIFF), and write results to PostgreSQL.
Data Boundaries and State Management
The PostgreSQL database is the single source of truth for translation units. Weblate synchronizes file‑system state with database rows using a state‑machine approach:
- File components are split into units.
- Units are processed in batches to avoid race conditions.
- Row‑level locking in PostgreSQL ensures concurrent edits do not corrupt data.
Operational Checks for a Healthy Pipeline
Because processing is asynchronous, the health of the translation pipeline depends on Celery’s queue depth and worker heartbeats.
# Run on the server where Celery workers are hosted
celery -A weblate status
Interpretation:
- High
activetasks relative to worker count → scale workers or tune file‑parsing limits. - Workers reporting
stoppedoroffline→ investigate crashes or resource exhaustion.
Example Redis configuration for a production broker:
# settings.py
CELERY_BROKER_URL = 'redis://redis:6379/0'
CELERY_RESULT_BACKEND = 'redis://redis:6379/1'
CELERY_ACCEPT_CONTENT = ['json']
CELERY_TASK_SERIALIZER = 'json'
CELERY_RESULT_SERIALIZER = 'json'
Failure Modes and Mitigations
- Memory Exhaustion: Large binary or deeply nested JSON files can exceed worker memory. Mitigation: set
celeryd_concurrencyand use streaming parsers where possible. - Stale Cache States: A worker crash during a file write can leave Redis with an inconsistent status. Mitigation: make write operations idempotent and retry on failure.
- Lock Contention: High‑concurrency edits to the same component cause row‑level locks, slowing the pipeline. Mitigation: batch updates and use optimistic locking where feasible.
Verifying Data Integrity in Production
- Trigger a sync of a large component (e.g., > 500 translation units) via the UI.
- While the sync is in progress, restart the Celery worker service (e.g.,
systemctl restart celery). - After the sync completes, query the database for the affected component:
# psql -d weblate -c "SELECT count(*) FROM weblate_translationunit WHERE component_id = 42;" - Verify that the count matches the expected number of units and that no partial entries exist.
Conclusion
Weblate’s decoupled design—separating UI, task queue, workers, and database—enables responsive interfaces even when handling massive translation repositories. By monitoring Celery queue depth, ensuring idempotent writes, and tuning worker resources, administrators can keep the pipeline robust and avoid the common pitfalls of asynchronous processing.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.