Django bulk_create for high‑volume imports: benefits and caveats
Learn when Django's bulk_create cuts round‑trips for large imports, what it skips (signals, validation, auto_now), and how to use ignore_conflicts safely.
24 Jul 2025, 14:31 UTC

Problem: nightly import stalls on per‑row saves
A typical data pipeline that loops over thousands of rows and calls model.save() for each record generates one INSERT statement per row. The resulting round‑trip overhead, combined with Django’s model initialization and signal handling, can cause timeouts, excessive connection usage, and lock contention on the database. The useful takeaway is that Django’s bulk operations replace many small queries with far fewer batched statements, but they do so by intentionally skipping parts of the ORM lifecycle that you may rely on.
How bulk_create works (brief)
bulk_create builds one INSERT statement per batch (default size 100 for SQLite, 1000 for other backends) and sends it to the database. It does not invoke model.save(), so pre_save and post_save signals are not fired, auto_now and auto_now_add fields are not populated automatically, and model validation (full_clean, field validators) is bypassed. Primary keys are filled in for PostgreSQL and SQLite backends; MySQL may require the RETURNING clause (available via third‑party backends) to get IDs.
Worked example: idempotent import with ignore_conflicts
Suppose you have an Event model with a unique external_id and a created_at timestamp that must be set explicitly. You receive a CSV of events and want to insert new rows while silently ignoring duplicates.
# myapp/management/commands/import_events.py
from django.core.management.base import BaseCommand
from django.db import transaction
from myapp.models import Event
class Command(BaseCommand):
help = 'Import events idempotently using bulk_create'
def handle(self, *args, **options):
# Imagine `rows` is a list of dicts parsed from CSV
rows = [
{'external_id': 'evt-001', 'payload': 'A', 'ts': '2026-09-01T12:00:00Z'},
{'external_id': 'evt-002', 'payload': 'B', 'ts': '2026-09-01T12:05:00Z'},
# … many more …
]
objs = [Event(external_id=r['external_id'],
payload=r['payload'],
created_at=r['ts']) for r in rows]
# Wrap in a transaction for atomicity; requires DB write permission.
with transaction.atomic():
Event.objects.bulk_create(
objs,
batch_size=500,
ignore_conflicts=True, # works on PostgreSQL, SQLite, MySQL 8.0.19+
)
self.stdout.write(self.style.SUCCESS('Import finished'))
Notes for this pattern:
- Assign
created_atyourself;auto_now_addwill not be set bybulk_create. ignore_conflicts=Truecauses the database to skip rows that violate a unique constraint. You will not receive a list of which rows were skipped; you must query the table afterward if you need that information.- Batch size is a tuning knob. Larger batches reduce the number of round‑trips but increase memory usage and lock duration. You can verify the number of INSERT statements issued by examining
django.db.connection.queriesin DEBUG mode or using the Django Debug Toolbar. - Any audit trail or denormalized counter that lives in a
post_savesignal will not run. Move that logic into the import code or implement a database trigger if necessary.
Trade‑offs and limits
- Signals are bypassed. If your business logic, audit logging, or cache invalidation depends on
pre_save/post_save, it will not execute. - Timestamps and validation are skipped. You must set
auto_now/auto_now_addfields manually and validate data before building instances; invalid data can reach the database. - Relationships are not created. Foreign keys must point to existing rows, and many‑to‑many tables require a separate bulk operation after the parent rows are saved.
- Database limits apply. PostgreSQL caps the number of query parameters at 65535; MySQL is limited by
max_allowed_packet; SQLite defaults toSQLITE_MAX_VARIABLE_NUMBER=999. Very wide models may hit these limits even with modest batch sizes. - Partial failures inside
atomic()can leave the transaction unusable on some backends. Consider using savepoints or smaller batches to isolate faults.
Actionable closing checklist
- Identify any signal‑dependent logic before switching to bulk methods; refactor or replace with explicit calls.
- Explicitly set
auto_now,auto_now_add, and any default‑valued fields. - Validate input data in Python (e.g., using
full_cleanor a serializer) before constructing model instances. - Choose a batch size based on row width and your backend’s parameter limits, not just raw throughput; monitor memory and lock time.
- After the bulk insert, run a simple count query (
Event.objects.filter(...).count()) to confirm the expected number of rows were added. - Test the behavior on your target Django version and database backend; check release notes for changes to
ignore_conflictsor primary‑key population.
Use bulk operations when you control the data shape and can accept the omission of per‑instance ORM hooks. When you need signals, validation, or automatic timestamp handling, stay with model.save() and accept the higher per‑row cost.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.