Designing a Minimal Oracle Data Guard Physical Standby for HA/DR
Learn the requirements, smallest viable architecture, trust boundaries, key operational checks, and failure modes for a basic Oracle Data Guard standby in maximum performance mode.
21 May 2026, 04:08 UTC

Problem and Takeaway
You need a reliable high‑availability and disaster‑recovery solution for an Oracle database without over‑engineering the setup. The smallest viable design is a primary database paired with a single physical standby operating in MAXPERFORMANCE (asynchronous) mode, with redo transport isolated from client traffic. This note outlines the requirements, that minimal architecture, trust/data boundaries, essential operational checks, typical failure modes, and the conditions that would force a redesign.
Requirements
- Primary database running a supported Oracle Enterprise Edition release (e.g., 19c or 21c).
- Standby host with identical Oracle binaries and patch level as the primary.
- Network connectivity that allows redo log transmission from primary to standby with sufficient bandwidth for the expected redo generation rate.
- Separate network paths (or VLANs) for redo transport and for client‑to‑database traffic to enforce trust boundaries.
- Sufficient storage on the standby to hold archived redo logs until they are applied.
- A backup strategy for the primary (RMAN) that can be used to reinstantiate the standby if needed.
Smallest Suitable Design
The design consists of three components:
- Primary database – generates redo and ships it to the standby.
- Physical standby database – receives redo via log transport services and applies it using the managed recovery process (MRP).
- Network isolation – a dedicated redo‑transport network (or VLAN) that carries only Oracle Net traffic for LOG_ARCHIVE_DEST_n; client connections use a separate network.
Configuration example (primary side):
# In the primary spfile or pfile
LOG_ARCHIVE_CONFIG = 'DG_CONFIG=(primary,standby)'
LOG_ARCHIVE_DEST_2 = 'SERVICE=standby LGWR ASYNC VALID_FOR=(ONLINE_LOGFILES,PRIMARY_ROLE) DB_UNIQUE_NAME=standby'
LOG_ARCHIVE_DEST_STATE_2 = ENABLE
FAL_SERVER = standby
FAL_CLIENT = primary
On the standby, the minimal parameters are:
LOG_ARCHIVE_CONFIG = 'DG_CONFIG=(primary,standby)'
LOG_ARCHIVE_DEST_1 = 'LOCATION=/u01/app/oracle/standby/archivelog VALID_FOR=(ALL_LOGFILES,ALL_ROLES) DB_UNIQUE_NAME=standby'
STANDBY_FILE_MANAGEMENT = AUTO
DB_UNIQUE_NAME = standby
The listener for redo transport should listen on a different IP address or port than the client listener, ensuring that a failure in the client network does not impede redo shipping.
Trust and Data Boundaries
Trust boundary: the redo‑transport network is considered trusted because it carries authenticated Oracle Net streams; only the primary and standby hosts are allowed to initiate connections. Data boundary: the standby does not accept DML directly; it only applies redo received from the primary. Any attempt to write to the standby while it is in managed recovery mode will be rejected, preserving the one‑way flow of data.
Operational Checks
Regular verification ensures the standby stays within the agreed RPO and RTO.
- Redo transport status – query V$ARCHIVE_DEST_STATUS on the primary:
SELECT DEST_ID, STATUS, LOG_SEQUENCE, LOG_ARRIVED, APPLIED_SEQUENCE
FROM V$ARCHIVE_DEST_STATUS
WHERE DEST_ID = 2;
Look for STATUS = VALID and LOG_ARRIVED increasing in step with the primary’s LOG_SEQUENCE.
- Apply lag – on the standby, check V$MANAGED_STANDBY:
SELECT PROCESS, STATUS, SEQUENCE#, DELAY_MINS
FROM V$MANAGED_STANDBY
WHERE PROCESS = 'MRP';
The DELAY_MINS column shows how many minutes the standby is behind the primary; keep it under your RPO threshold.
- Archived log availability – verify that archived logs are present on the standby:
SELECT SEQUENCE#, APPLIED
FROM V$ARCHIVED_LOG
WHERE DEST_ID = 1
ORDER BY SEQUENCE#;
All sequences shown as APPLIED='YES' indicate no gaps.
Perform a switchover test in a maintenance window:
ALTER DATABASE COMMIT TO SWITCHOVER TO PHYSICAL STANDBY;
After the command completes, verify the new primary is accepting connections and the former primary is now a standby with MRP running.
Failure Modes
- Network split‑brain – loss of redo‑transport connectivity causes the primary to stall if operating in MAXPROTECTION; in MAXPERFORMANCE the primary continues but redo accumulates, risking data loss if the outage exceeds the standby’s redo‑generation capacity.
- Log apply corruption** – damaged redo on the standby (e.g., due to storage I/O errors) can halt MRP, leading to increasing apply lag.
- Standby lag exceeding RPO** – prolonged apply lag means the standby cannot meet the agreed recovery point objective; a failover would lose transactions.
- Version/patch mismatch** – applying a patch to the primary but not the standby breaks log apply and forces a reinstantiation.
When the Design Must Change
- Zero data loss requirement** – regulatory or business needs that demand RPO = 0 force a switch to MAXPROTECTION (synchronous) mode, which may impact primary performance and requires a standby with guaranteed network availability.
- Geographic latency** – if the standby is located far enough that synchronous redo would violate RTO, you may need to introduce a far‑sync instance or accept asynchronous mode with a tighter monitoring regime.
- Licensing constraints** – Active Data Guard (real‑time query on the standby) requires an extra license; if unavailable, you cannot offload reporting queries to the standby and must size the primary accordingly.
- Multiple standbys for geographic distribution** – a single standby may not satisfy disaster‑recovery site separation; adding a second standby in a different region changes the network and configuration topology.
In each case, revisit the trust boundaries (e.g., add a far‑sync listener), adjust the redo‑transport mode, and update operational checks to reflect the new architecture.
Limitations and Practical Verification
The MAXPERFORMANCE mode does not guarantee that every committed transaction is shipped before a primary failure; therefore, verify that your RPO tolerates the possible loss of up to the last few seconds of redo. Monitor the V$ARCHIVE_DEST_STATUS LOG_ARRIVED vs. LOG_SEQUENCE delta; a consistently growing delta indicates a transport issue that must be addressed before relying on the standby for failover.
Finally, always keep a recent RMAN backup of the primary and a documented procedure to reinstantiate the standby from that backup if log apply becomes unrecoverable.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.