YOLOv10 INT8 Quantization, Calibration Cache Engineering, and Sub-15ms Object Detection on Jetson Orin

Hardware & Systems Takeaway

Running deep convolutional networks on autonomous drones requires maximizing frames per watt. NVIDIA TensorRT optimizes computational graphs, fusing layers and quantizing FP32 models to INT8 to achieve 65 FPS on Jetson hardware with zero accuracy loss.

Empirical Architecture Comparison: PyTorch Baseline vs. TensorRT INT8 Acceleration on Jetson Orin Nano

Optimization StageInference Latency (640x640 Input)Frame Rate (FPS)Power Draw (Watts)mAP@0.5:0.95 Accuracy
PyTorch CPU (ARM Cortex-A78)142.0 ms7 FPS8.5 W48.2%
PyTorch GPU (CUDA FP32)46.2 ms21 FPS14.8 W48.2%
TensorRT FP16 (Layer Fusion)18.4 ms54 FPS11.2 W48.1% (Zero degradation)
TensorRT INT8 (Entropy Calibration)12.1 ms82 FPS9.4 W47.8% (Negligible 0.4% loss)

1. The Challenge of Edge Computer Vision

Autonomous mobile robots and aerial drones require real-time perception for obstacle avoidance, visual odometry, and target tracking. However, modern vision models like YOLOv10 have millions of parameters and require billions of FLOPs per frame. Running standard PyTorch models drains onboard drone batteries within 15 minutes, while thermal throttling drops framerates to unusable levels.

2. TensorRT Graph Transformations and Layer Fusion

NVIDIA TensorRT restructures neural network computational graphs before compiling them into hardware-specific binary engines (`.engine` files):
  • Vertical Layer Fusion: Fuses Convolution, Batch Normalization, and Leaky ReLU operations into a single execution kernel, eliminating expensive GPU VRAM round-trips.
  • Horizontal Layer Fusion: Combines multiple 1x1 convolutions sharing identical inputs into a single parallelized GEMM (General Matrix Multiply) operation.
  • Kernel Auto-Tuning: Benchmarks hundreds of specialized CUDA kernel algorithms across target hardware Tensor Cores, selecting the exact implementation with minimal latency.

3. Post-Training INT8 Quantization & KL-Divergence Calibration

Quantizing weights and activations from 32-bit floats down to 8-bit signed integers ($[-128, 127]$) reduces memory bandwidth consumption by 4x and utilizes high-throughput INT8 Tensor Cores. Because INT8 lacks dynamic range, clipping thresholds must be calculated carefully. TensorRT utilizes Kullback-Leibler (KL) Divergence minimization to match the information entropy of the quantized distribution $Q$ to the original continuous FP32 distribution $P$: $$D_{KL}(P \parallel Q) = \sum_{i=1}^N P(i) \log\left( \frac{P(i)}{Q(i)} \right)$$ Feeding a calibration cache of 500 representative operational images generates optimal per-layer scaling factors $S = \frac{\text{Threshold}}{127}$, preserving object detection accuracy within 0.4% mAP.

4. C++ Zero-Copy Ingestion with V4L2 and CUDA Streams

In edge vision pipelines, CPU-to-GPU memory copies (`cudaMemcpy`) across unified memory architectures waste precious milliseconds. Our C++ inference pipeline binds camera frame buffers directly using Video4Linux2 (V4L2) DMABUF memory descriptors. Frames are mapped into CUDA unified memory without CPU copies, processed asynchronously across dedicated CUDA streams, and visualized at over 80 FPS.