Distributed AI-RAN Inference Placement for Joint Radio Optimization and Edge Workload Acceleration
DOI:
https://doi.org/10.64751/7ep35m50Abstract
The rapid evolution of Artificial Intelligence for Radio Access Networks (AI-RAN) has created new opportunities to optimize radio resource management while accelerating edge intelligence for next-generation wireless networks. However, centralized AI inference often introduces excessive latency, GPU resource imbalance, and increased network congestion, limiting the performance of real-time 5G and emerging 6G applications. This paper proposes a Distributed AI-RAN Inference Placement framework that jointly optimizes radio resource allocation and edge workload acceleration through intelligent inference placement across distributed GPU-enabled edge servers. The framework dynamically analyzes radio conditions, computational capacity, traffic demand, and latency requirements to determine the optimal execution location for AI models. By coordinating radio and compute resources simultaneously, the proposed approach reduces inference delay, improves GPU utilization, enhances throughput, lowers energy consumption, and maintains high Quality of Service (QoS). Experimental evaluation demonstrates significant improvements in scalability, resource efficiency, and real-time AIdriven radio optimization compared with conventional centralized AI-RAN deployment strategies.
Downloads
Published
Issue
Section
License

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.







