Integrating Norg Transactional Messaging into a Microservice Architecture: An Architecture Note
Learn how to embed Norg’s exactly‑once transactional messaging into a microservice ecosystem, covering requirements, minimal design, trust boundaries, monitoring, and failure handling.
04 Jan 2026, 20:48 UTC

Problem Statement
Microservices that need to coordinate state changes across boundaries often rely on message queues. When those changes must be atomic—either all succeed or none do—exact‑once delivery and transactional semantics are required. Norg’s Enterprise edition offers such a feature via XA‑compliant transactions. This note explains how to embed that feature into a typical microservice stack, from prerequisites to operational health checks.
Requirements
- Java 11 or newer runtime for the client library.
- Access to a Norg Enterprise cluster that exposes XA support.
- Schema registry (e.g., Confluent Schema Registry) for validating message payloads.
- Mutual TLS (mTLS) certificates for all client connections.
- Monitoring tooling that can scrape Prometheus metrics exposed by Norg.
Transactional Guarantees Needed
We require:
- Exactly‑once delivery for every message sent within a transaction.
- Atomic publish/subscribe: all messages in the transaction commit together.
- Idempotent retry logic that does not produce duplicate side‑effects.
Minimal Design
The simplest, most resilient architecture uses a single Norg cluster and one topic per microservice. The cluster is configured for fault tolerance with a quorum of five nodes and replicated storage. Each service exposes a thin wrapper around the Norg client that participates in local XA transactions.
Cluster Layout
norg-cluster-1
norg-cluster-2
norg-cluster-3
norg-cluster-4
norg-cluster-5
All nodes run with the same configuration file (norg.properties), which includes:
# norg.properties
broker.id=1
listeners=PLAINTEXT://0.0.0.0:9092
security.inter.broker.protocol=PLAINTEXT
num.partitions=3
replication.factor=3
transaction.state.log.replication.factor=3
transaction.state.log.min.isr=2
Client Wrapper
Each service runs a lightweight TransactionalProducer that starts an XA transaction, sends messages to its dedicated topic, and commits or rolls back based on the business logic outcome.
public class TransactionalProducer {
private final KafkaProducer<String, byte[]> producer;
public TransactionalProducer(Properties props) {
this.producer = new KafkaProducer<>(props);
}
public void sendWithTransaction(String topic, String key, byte[] payload, Runnable businessLogic) {
producer.initTransactions();
producer.beginTransaction();
try {
producer.send(new ProducerRecord<>(topic, key, payload));
businessLogic.run();
producer.commitTransaction();
} catch (Exception e) {
producer.abortTransaction();
throw e;
}
}
}
Note that the producer must be configured with transactional.id and enable.idempotence=true to satisfy Norg’s XA contract.
Trust and Data Boundaries
Because Norg is a privileged component, the following boundaries must be respected:
- Authentication: Only clients presenting a signed client certificate that matches a whitelist in the broker’s
ssl.client.auth=requiredsetting may connect. - Encryption: All traffic uses TLS (port 9093) and data is stored encrypted at rest via disk encryption on each broker node.
- Schema Validation: Before sending, the producer serializes the payload using Avro and validates against the schema registry. The consumer also validates on receipt.
- Isolation: Each service writes to its own topic; cross‑service data is only visible via downstream consumers that explicitly subscribe to the topic.
Operational Checks
Monitoring is critical because a transaction failure can leave the system in an inconsistent state if not caught. The following metrics should be scraped and alerted on:
| Metric | Description | Alert Condition |
|---|---|---|
| norg_tx_commits_total | Total successful transaction commits. | None – used for trend analysis. |
| norg_tx_rollbacks_total | Total transaction rollbacks. | High rate (>5% of total tx) indicates instability. |
| norg_tx_commit_latency_ms | Latency of commit operation. | Average >200 ms triggers alert. |
| norg_partition_lag_ms | Consumer lag per partition. | Lag >1 min triggers alert. |
Health probes should also check the leader election status of each broker. A probe that cannot reach a majority of nodes should fail.
Sample Prometheus Alert
groups:
- name: norg-alerts
rules:
- alert: NorgHighRollbackRate
expr: rate(norg_tx_rollbacks_total[5m]) / rate(norg_tx_commits_total[5m]) > 0.05
for: 5m
labels:
severity: warning
annotations:
summary: "High rollback rate detected"
description: "More than 5% of transactions are rolling back in the last 5 minutes. Investigate potential network or application issues."
Failure Modes
- Broker Failure During Transaction: Norg rolls back the transaction. The client receives an
TransactionAbortedExceptionand must retry. Idempotence of the business logic prevents duplicate side‑effects. - Partition Rebalance: During a rebalance, temporary duplication may occur. Deduplication logic in the consumer (e.g., using a unique message ID in the key) ensures that duplicates are ignored.
- Network Partition: Intermittent connectivity can cause transaction timeouts. The client should implement exponential back‑off retries and log the failure for later analysis.
- Schema Mismatch: If the producer sends a payload that does not match the registered schema, the broker rejects the message. Validation should occur client‑side before sending.
Graceful Retry Strategy
public void safeSend(String topic, String key, byte[] payload) {
int attempts = 0;
while (attempts < 5) {
try {
sendWithTransaction(topic, key, payload, () -> {/* business logic */});
return;
} catch (TransactionAbortedException e) {
attempts++;
long backoff = (1L << attempts) * 100L; // 100 ms, 200 ms, 400 ms...
Thread.sleep(backoff);
}
}
throw new RuntimeException("Failed to commit transaction after retries");
}
Conditions That Prompt a Redesign
- Loss of XA Support: If Norg Enterprise drops XA support, the architecture must shift to a two‑phase commit model using a distributed transaction manager or switch to a system that natively supports atomic publish/subscribe (e.g., Pulsar with transaction support).
- Horizontal Scaling Beyond One Cluster: When the service load exceeds what a single Norg cluster can handle, a multi‑cluster federation or sharding strategy becomes necessary. This introduces cross‑cluster transaction coordination, which is outside the scope of the simple design.
- Regulatory or Compliance Constraints: If new regulations require stricter data residency or audit trails that Norg cannot provide, the messaging layer must be replaced or augmented with additional logging.
- Cost Constraints: Enterprise licensing costs may become prohibitive. In that case, a community edition or open‑source alternative lacking XA support would necessitate re‑architecting the transaction logic into the application layer.
Practical Verification Steps
- Deploy a three‑node Norg cluster with
transaction.state.log.replication.factor=3and verify thatkafka-topics.sh --describeshows thetransaction.state.logtopic with the correct replication. - Run the
TransactionalProducersample against a test topic and intentionally trigger a rollback (e.g., throw an exception after sending). Confirm thatnorg_tx_rollbacks_totalincrements and that no message appears in the consumer. - Simulate a broker failure by stopping one node during a transaction. Ensure that the client receives an abort exception and that the transaction is rolled back.
- Check the Prometheus dashboard for the metrics listed above and verify that the alerting rules fire under the simulated failure conditions.
Conclusion
By treating Norg as a privileged, transaction‑aware broker and wrapping its client calls within local XA transactions, a microservice can achieve exactly‑once semantics with minimal architectural overhead. The design relies on strong trust boundaries, rigorous monitoring, and clear failure handling. When the assumptions—such as XA support or cluster size—change, the architecture should be revisited to maintain consistency and performance.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.