A path toward 5G NR PHY acceleration: the LDPC offloading on BlueField-3 DPU
Carregando...
Data
Autores
Título da Revista
ISSN da Revista
Título de Volume
Editor
Universidade Federal de São Carlos
Resumo
The transition toward AI-native Radio Access Network (RAN) architectures further intensifies the demand for cloud infrastructure resources, as distributed intelligence, machine learning inference, and real-time data analytics become integral components of future RAN deployments. Containerized and virtualized network functions in Cloud RAN environments increasingly rely on hardware acceleration and workload offloading to meet the stringent latency, throughput and energy efficiency requirements of Fifth Generation (5G) Advanced and emerging Sixth Generation (6G) mobile networks. Within the 5G New Radio (NR) Physical Layer (PHY), Low-Density Parity-Check (LDPC) decoding represents a computationally intensive workload that depletes host CPU resources. While offloading to Data Processing Unit (DPU) offers general purpose core relief, it introduces systemic latency overheads. This research project evaluates an acceleration framework that offloads the 5G NR LDPC encoding and decoding pipeline from an OpenAirInterface (OAI) x86 host to an ARM-based NVIDIA BlueField-3 DPU utilizing the Arm RAN Acceleration Library (ArmRAL) over a Communication Channel. Hardware acceleration provides several benefits for 5G NR. Specialized accelerators such as Graphics Processor Unit (GPU), Field-Programmable Gate Array (FPGA), and DPU execute computationally intensive PHY functions in parallel, significantly increasing throughput while reducing processing latency to meet the stringent timing requirements of 5G. They also improve energy efficiency by delivering higher performance per watt than general-purpose CPUs. In addition, offloading functions such as LDPC decoding, Fast Fourier Transform (FFT), and beamforming relieves the host CPU, allowing it to focus on higher-layer protocol processing, scheduling, and network management. These advantages improve system scalability and make hardware acceleration a key enabler of cloud-native Open RAN deployments. In this work the ArmRAL LDPC kernels on the BlueField-3 ARM Cortex-A78AE cores are functionally correct, producing output bit-identical to the OAI software baseline as verified through Block Error Rate (BLER)–Signal-to-Noise Ratio (SNR) characterization under Additive White Gaussian Noise (AWGN) and Tapped Delay Line A (TDL-A30) fading channels, and execute competitively with the x86 software baseline. However, the host-visible per-CB Communication Channel (Comch) latency exceeds the slot budget by two orders of magnitude, producing a substantial Flow Completion Time (FCT) slowdown and energy efficiency degradation relative to the OAI software baseline across end-to-end experiments. Our results show that the bottleneck is the per-CB Comch framework overhead between the host and the DPU, rather than ARM compute capability or Peripheral Component Interconnect Express (PCIe) payload transfer, identifying host-visible Comch latency as the dominant architectural barrier to transparent lookaside PHY acceleration and quantifying its impact on FCT, throughput, and gross energy consumption relative to the OAI software baseline. Future work directions include persistent Comch contexts, Data Center Infrastructure-on-a-Chip Architecture (DOCA) Direct Memory Access (DMA)-based data paths, and slot-level OAI interfaces approaching an inline acceleration model.