Sociedade Brasileira de Telecomunicações · desde 1983 secretaria@sbrt.org.br
← SBrT2020

Leveraging Reinforcement Learning for User Pairing in Full Duplex Networks

João Rafael Barbosa de Araujo, Francisco Rafael Marques Lima
Reinforcement LearningFull-DuplexMultiarmed bandit

Resumo

In this article we employ a reinforcement learning solution called Upper Confidence Bound (UCB) over the framework of Multi-Armed Bandit (MAB) to solve User Equipment (UE) pairing problem in Full Duplex (FD) network. In the context of the total data rate maximization problem, our proposed solution is capable of learning the best UE pair iteratively by exploring and exploiting the solution space. By the presented simulation results, we show that our proposed algorithm is more robust to the absence of knowledge about inter-UE Channel State Information (CSI). In the complete absence of CSI about inter-UE channel gains, our proposed solution overperforms the Maximum Rate (MR) solution by 26%.