ROS 2 DDS Middleware: Architecture Decisions for Reliable Robot Communication
ROS 2's DDS middleware replaces ROS 1's central master with decentralized discovery, but introduces QoS configuration complexity. This architecture note covers domain partitioning, per-topic QoS tuning, silent failure modes, and concrete verification commands for multi-robot deployments.
05 Jul 2026, 00:27 UTC

The Problem: Moving Beyond a Single Point of Failure
ROS 1's centralized master node created a single point of failure: if the master crashed, the entire robot system lost discovery and communication. ROS 2 replaces this with a decentralized Data Distribution Service (DDS) layer, but the shift introduces new configuration complexity. The practical takeaway: DDS gives you granular control over reliability and durability per topic, but mismatched Quality of Service (QoS) profiles between publisher and subscriber cause silent communication failures that are difficult to debug without the right tooling.
Requirements Driving the Design
Robot systems need:
- Decentralized discovery: No central broker to fail or become a bottleneck
- Per-topic transport tuning: Sensor streams (lidar, cameras) need different guarantees than command topics
- Network isolation: Multiple robots on the same LAN must not interfere
- Real-time constraints: Deterministic latency for control loops
DDS satisfies these by implementing the OMG DDS standard directly in the middleware layer, with ROS 2's rmw (ROS Middleware) abstraction allowing vendor swaps (Fast DDS, Cyclone DDS, Connext) without application changes.
Smallest Suitable Design: Publish-Subscribe with Domain Partitioning
The core pattern is straightforward: nodes publish to topics; subscribers receive matching topics. DDS handles serialization (CDR), transport (UDP/TCP/shared memory), and discovery. The minimal configuration requires only a Domain ID (integer 0–232) shared by all participating nodes. Nodes with different Domain IDs never discover each other, providing the first trust boundary.
Each topic independently declares a QoS profile combining:
- Reliability:
RELIABLE(ACK/NACK retransmission) vsBEST_EFFORT(fire-and-forget) - Durability:
TRANSIENT_LOCAL(late joiners get last value) vsVOLATILE(only live data) - History:
KEEP_LASTwith depth N vsKEEP_ALL - Deadline/Lifespan: Expected update frequency and data validity window
A publisher offering RELIABLE + TRANSIENT_LOCAL + KEEP_LAST(10) will not match a subscriber requesting BEST_EFFORT + VOLATILE. The middleware drops the connection silently—no error logs, no exceptions.
Trust and Data Boundaries
Three mechanisms enforce boundaries:
- Domain ID: Coarse isolation. Typical practice: Domain ID = robot ID for multi-robot fleets. Set via
ROS_DOMAIN_IDenvironment variable before launch. - Partition QoS: String-based logical partitions within a domain. A publisher with partition "sensors/lidar" and subscriber with "sensors/*" match; "actuators/*" does not.
- Security plugins: DDS Security spec (authentication, access control, encryption) available in vendor implementations but adds latency and configuration overhead. Most robotics deployments skip this and rely on physical network isolation (VLANs, dedicated Wi-Fi).
Data serialization uses CDR (Common Data Representation). ROS 2 defines standard message types (sensor_msgs/msg/LaserScan, geometry_msgs/msg/Twist) that map to IDL; custom types require .msg definitions and generate type support code at build time.
Operational Checks: Verifying Discovery and QoS Match
Run these on each machine participating in the ROS 2 system (requires sourced ROS 2 workspace, no elevated permissions):
1. Confirm Node Visibility Across Machines
ros2 node list
Expected: Nodes from all machines appear. If only local nodes show, check multicast routing (IGMP snooping on managed switches often blocks discovery traffic) and ROS_DOMAIN_ID consistency.
2. Inspect Active QoS Profiles
ros2 topic info --verbose /scan
Output shows publisher and subscriber QoS side by side. Look for mismatches in Reliability, Durability, History. Example mismatch:
Publisher QoS:
Reliability: RELIABLE
Durability: TRANSIENT_LOCAL
Subscriber QoS:
Reliability: BEST_EFFORT
Durability: VOLATILE
This pair will not communicate. Fix by aligning publisher to subscriber (for sensor streams) or subscriber to publisher (for commands).
3. Measure End-to-End Latency Under Load
ros2 topic hz /cmd_vel --window 1000
Run while robot executes typical motion. Compare against control loop period (e.g., 10 ms for 100 Hz). Spikes indicate queue buildup or DDS thread starvation.
Failure Modes and Mitigations
| Failure Mode | Symptom | Root Cause | Mitigation |
|---|---|---|---|
| Network partitioning | Nodes disappear from ros2 node list; topics stop updating | Multicast blocked, firewall, Wi-Fi roaming, Domain ID mismatch | Use FASTRTPS_DEFAULT_PROFILES_FILE to configure unicast discovery; set ROS_DISCOVERY_SERVER for centralized discovery relay |
| Silent QoS mismatch | Publisher runs, subscriber receives nothing, no errors | Incompatible Reliability/Durability/History | Enforce QoS policy in launch files; add CI check parsing ros2 topic info --verbose output |
| Resource exhaustion | OOM kill, latency spikes, dropped samples | Queue depth too high for high-frequency topics (e.g., 1 kHz IMU with KEEP_LAST(1000)) | Profile memory with ros2 topic echo --qos-profile sensor_data /imu; use BEST_EFFORT + KEEP_LAST(1) for high-rate sensors |
| Discovery storm | CPU spike at startup, delayed node readiness | Many nodes, large participant data, multicast flood | Limit participant lease duration; use static discovery XML for known topologies |
Concrete Configuration Example: Lidar vs Command Topics
A differential-drive robot with a 20 Hz lidar and 100 Hz control loop:
# Launch file snippet (YAML)
lidar_publisher:
ros__parameters:
qos_overrides:
/scan:
publisher:
reliability: best_effort
durability: volatile
history: keep_last
depth: 5
deadline:
sec: 0
nanosec: 50000000 # 50 ms
cmd_vel_publisher:
ros__parameters:
qos_overrides:
/cmd_vel:
publisher:
reliability: reliable
durability: transient_local
history: keep_last
depth: 10
lifespan:
sec: 1
nanosec: 0
Rationale: Lidar data is high-volume, periodic, and stale scans are useless—BEST_EFFORT avoids retransmission overhead. cmd_vel must arrive reliably; TRANSIENT_LOCAL ensures a late-joining controller gets the last command (safety stop).
Verify with:
ros2 topic info --verbose /scan
ros2 topic info --verbose /cmd_vel
Both should show matching publisher/subscriber QoS.
Conditions That Would Change the Design
- Hard real-time requirement: Switch to Cyclone DDS with
--realtimeflag or use ROS 2 Real-Time Executor; avoid Fast DDS's default thread pool. - Cross-subnet deployment: Default multicast discovery fails across routers. Configure
DiscoveryServer(Fast DDS) orDDS_Router(Cyclone) as relay; or use static peer lists in XML. - Security mandate: Enable DDS Security plugins (auth: PKI-DH, access: permissions XML, crypto: AES-GCM). Expect 15–30% latency increase; test on target hardware.
- Microcontroller targets: Use micro-ROS (DDS-XRCE) with agent on Linux host; the agent bridges microcontroller nodes to DDS domain. QoS options subset only.
Limitations and Verification Checklist
This analysis assumes ROS 2 Humble/Iron/Jazzy on Linux with Fast DDS or Cyclone DDS. Behavior varies by vendor:
- Fast DDS defaults to multicast discovery; Cyclone defaults to unicast with peer list.
- Shared memory transport (SHM) activates automatically on same host but requires
/dev/shmcapacity. - QoS compatibility rules follow DDS spec but vendor bugs exist (e.g., Fast DDS 2.10.x had a
TRANSIENT_LOCALdurability regression).
Before deploying to robot:
- Run
ros2 doctor --reporton each machine; resolve DDS middleware warnings. - Simulate 5% packet loss with
tc qdisc add dev eth0 root netem loss 5%(requires root) and verifyRELIABLEtopics recover,BEST_EFFORTtopics drop gracefully. - Stress test: publish 10 MB/s on 5 topics for 10 minutes; monitor memory growth with
pidstat -r -p $(pgrep -f 'my_node') 1.
If all checks pass, the DDS layer is configured for the intended operational envelope. Revisit when topology changes (new robot, new network segment, new real-time requirement).
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.