Shipping ONNX inference with OpenCV DNN without PyTorch or TensorFlow
OpenCV DNN can run ONNX models without PyTorch or TensorFlow at runtime. This post covers loading, preprocessing, inference and post-processing patterns for YOLO-style detectors, backend selection, NMS post-processing and thread safety trade-offs.
21 Sept 2026, 15:56 UTC

The problem is not model accuracy, it is shipping. You have a trained detector exported to ONNX and you need it running in a C++ or Python service that cannot pull in PyTorch or TensorFlow at runtime. OpenCV’s dnn::Net can load ONNX opset 9–17 and run inference on CPU, OpenCL, CUDA or Vulkan with a single library dependency. The useful takeaway: it works well for fixed-size models like YOLOv8 when you own preprocessing, backend selection and thread safety.
What OpenCV DNN actually gives you
cv::dnn::readNetFromONNX creates a dnn::Net from an ONNX file. The net is backend-agnostic. You choose the backend and target at runtime:
net.setPreferableBackend(cv2.dnn.DNN_BACKEND_OPENCV)
net.setPreferableTarget(cv2.dnn.DNN_TARGET_CPU)
On a build with WITH_CUDA you can select CUDA and optionally FP16. Backend availability is build-time, not install-time. Check it before assuming GPU:
python -c \"import cv2; print(cv2.dnn.getAvailableBackends())\"
Expected check: the list includes DNN_BACKEND_CUDA if OpenCV was built with CUDA and matching cuDNN. If not, inference silently falls back to CPU. That silent fallback is a common production surprise.
Preprocessing is manual but centralized. cv2.dnn.blobFromImage handles resize, scale factor, mean subtraction, channel swap BGR to RGB, and layout to NCHW. Output tensors arrive as cv::Mat with shapes like [1, C, H, W] or [1, N, 7] for YOLO-style heads.
Fixed pipeline for YOLO-style ONNX
YOLOv8/v9/v10 ONNX exports load directly. The pipeline is load, blob, forward, decode, NMS, scale back.
- Load net and confirm input name and size. For dynamic batch axes, set input explicitly after blob creation to avoid shape inference errors.
- Create a blob with the model’s expected size, scale and mean. Reuse the cv::Mat buffer across frames to avoid per-frame allocations in real-time loops.
- Run net.forward().
- Decode boxes, filter by confidence, run cv::dnn.NMSBoxes. NMSBoxes expects boxes in [x, y, w, h] and scores as float vector.
- Scale coordinates from input resolution back to original frame size.
Thread safety note: dnn::Net is not thread-safe for concurrent forward calls. Clone the net per thread or serialize access with a mutex. Python cv2.dnn shares the C++ backend but adds GIL overhead for high throughput.
Worked example: YOLOv8 ONNX with OpenCV Python
This is a reference pattern, not a tested run. Replace placeholders with your paths.
import cv2
import numpy as np
net = cv2.dnn.readNetFromONNX('model.onnx')
net.setPreferableBackend(cv2.dnn.DNN_BACKEND_OPENCV)
net.setPreferableTarget(cv2.dnn.DNN_TARGET_CPU)
img = cv2.imread('input.jpg')
h0, w0 = img.shape[:2]
blob = cv2.dnn.blobFromImage(img, scalefactor=1/255.0, size=(640,640), swapRB=True, crop=False)
net.setInput(blob)
out = net.forward()
# out shape depends on export; typical YOLOv8 head is [1, N, 84] with xywh + class scores
# Decode, confidence filter, build boxes in [x,y,w,h]
boxes = []
scores = []
# ... reshape and apply confidence threshold ...
idxs = cv2.dnn.NMSBoxes(boxes, scores, score_threshold=0.25, nms_threshold=0.45)
# Scale boxes back to w0,h0 and draw
Verification steps you can run locally: confirm ONNX loads, confirm backend, compare a single forward output against the original framework on the same input for parity, and profile forward time with time.perf_counter to check latency budget.
Trade-offs and limitations
OpenCV DNN lacks native support for dynamic shapes beyond batch. Models with variable H/W need fixed resize or rebuild.
ONNX opset >17 may not parse. Export with opset 13–16 for broad compatibility.
FP16 on CUDA requires GPU with tensor cores and OpenCV built with WITH_CUDA and matching cuDNN. Older GPUs fall back to FP32 silently.
Model optimization helps: run onnxsim to fuse Conv+BN+ReLU before import. This reduces node count and can improve latency.
Licensing is separate. OpenCV 4.x is Apache-2.0. The model license is not. YOLOv8 for example carries AGPL-3.0 in some releases. Verify the model card before commercial deployment.
Actionable closing checklist
- Check backend availability with getAvailableBackends before deploying.
- Pin input size and preprocessing parameters to match export.
- Clone nets per worker thread or guard forward with a mutex.
- Profile forward latency on target hardware and reuse blob buffers.
- Validate inference parity on a small golden set and document model license.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.