Understanding Twitter's Snowflake IDs for Approximate Timestamp Extraction
Twitter generates unique 64-bit IDs locally using a timestamp, worker ID and sequence counter, enabling approximate time ordering without a central sequencer.
08 Oct 2025, 13:48 UTC

Generating a unique identifier for every tweet, like, and internal object at Twitter‑scale means you cannot ask a single database for the next number. A central sequencer becomes a latency bottleneck and a single point of failure for ingestion and fan‑out pipelines that must write from many machines and data centers. The practical decision is local generation with a time‑based, roughly ordered 64‑bit ID.
How the ID is built
Twitter’s Snowflake‑style ID packs three fields into a 64‑bit integer: a millisecond‑precision timestamp since a custom epoch, a worker identifier, and a per‑millisecond sequence counter. The timestamp occupies the most‑significant bits, guaranteeing that IDs increase with time for a given worker.
What it enables
Because each writer can mint IDs independently, there is no network round‑trip for every write. Clients can infer an approximate creation time simply by looking at the magnitude of the ID, which helps with client‑side sorting, storage partitioning and debugging without an extra lookup.
Operational considerations
Clock skew between workers can produce out‑of‑order IDs or, in the worst case, collisions if a worker’s clock moves backward. Production deployments rely on NTP synchronization and guards that pause generation on detected backward jumps.
The sequence counter limits how many IDs a single worker can emit in one millisecond. When the counter exhausts, the generator waits for the next tick, which can become a latency bottleneck under extreme bursts.
The worker‑ID space limits the number of independent generators that can be deployed; allocating new IDs requires central coordination when expanding regions or fleets.
Practical check: decoding a public tweet ID
To see the components, you can decode a publicly visible tweet ID with the known epoch and bit widths. The following Python snippet illustrates the operation; replace the constants with the values observed for your environment.
# Run locally, no special permissions required
# EPOCH_MS: milliseconds since the custom epoch (e.g., Twitter’s epoch 1288834974657)
# TIMESTAMP_SHIFT = WORKER_BITS + SEQUENCE_BITS
EPOCH_MS = 1288834974657
WORKER_BITS = 10
SEQUENCE_BITS = 12
def decode_snowflake(id_int):
timestamp = (id_int >> (WORKER_BITS + SEQUENCE_BITS)) + EPOCH_MS
worker_id = (id_int >> SEQUENCE_BITS) & ((1 << WORKER_BITS) - 1)
sequence = id_int & ((1 << SEQUENCE_BITS) - 1)
return timestamp, worker_id, sequence
# Example ID – replace with a real tweet ID you have collected
sample_id = 1234567890123456789
print(decode_snowflake(sample_id))
Expected checks: the returned timestamp should increase with the ID’s numeric value, worker IDs should stay constant for a given source, and sequence values should wrap within each millisecond. Risks: using an incorrect epoch or bit layout yields meaningless numbers, and decoding only reveals approximate time and worker, not the content of the tweet.
Actionable takeaway
If you need high‑throughput, roughly ordered identifiers without a central service, a time + worker + sequence scheme is a proven pattern. Plan for clock synchronization, allocate worker IDs centrally, monitor per‑millisecond sequence utilization, and accept that IDs leak time and origin. Document your epoch and bit layout, add alerts for clock drift and sequence stalls, and test generation behavior with simulated clock adjustments before rolling out to production.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.