Abstract
To keep up with the ever-increasing performance demand from wireless applications, wireless networks necessitate to operate following an efficient use of the available radio spectrum. Since the radio spectrum is a limited resource, the increasing demand for wireless services and applications to support a wide spectrum of users is a challenge in itself and therefore requires advanced techniques to manage and utilize radio resources efficiently. Dynamic spectrum access (DSA) and sharing play a crucial role in improving the utilization of the radio spectrum, as it departs from the traditional approach of static radio spectrum band allocation which usually leads to spectrum under-utilization. Cognitive radios (CRs) are considered the answer to effectively using the radio spectrum because of their ability to autonomously gain awareness of the wireless network environment and learn to adapt to changing conditions. The CRs have the ability to learn from its surroundings and adapt by making changes in its operating parameters dynamically to ensure reliable communication and efficient utilization and management of the radio spectrum. The cognitive functionality of the CRs can be achieved by leveraging machine learning (ML) techniques. Due to its model-free characteristic, reinforcement learning (RL) is a suitable ML candidate for resource allocation that naturally aligns with the CR paradigm. The work in this dissertation explores the application of different RL methodologies, namely Table-based Q-learning, Deep Q-learning (DQL), and Proximal Policy Optimization (PPO), to enable the cognition capacity of CRs in the challenging scenario of resource allocation in a distributed and uncoordinated heterogeneous underlay cognitive radio networks (CRNs). The application of RL in general wireless networking presents a key challenge. In a general scenario with distributed and uncoordinated transmissions (e.g., an ad-hoc network), the transmission from one CR affects the environment perceived by other CRs (appearing as interference) resulting in a multi-agent non-stationary environment. The multi-agent interaction with the non-stationary environment may not necessarily lead to the convergence of RL to the best policy even in the limit of arbitrarily long training time. In fact, it may be the case that it is not possible to define an optimal policy for all the multiple agents and, instead, it is necessary to consider convergence to some form of equilibrium state (e.g., Nash equilibrium). The work in this dissertation, considers this challenging case of a multi-agent non-stationary environment where nodes in a CR network share the spectrum with a primary network (PN) by operating in an uncoordinated distributed fashion following an underlay DSA scheme. The goal of this work is to develop a distributed and uncoordinated resource allocation RL mechanism for entangled CRs, that is, where a change of transmission parameters of one radio affects the operating environment and performance of the others. The proposed techniques does not rely on centralized training for convergence to optimal policy. The RL methodology is distributed, resulting in uncoordinated and dynamic interaction between the CRs. The first contribution in this dissertation introduces Uncoordinated and Distributed Multi-agent Deep Q-Learning (UDMA-DQL), a DQL technique that combines learning in exploration phases, and the use of a Best Reply Process with Inertia. Furthermore, in the study of UDMA-DQL herein, it is shown by considering aspects specific to deep reinforcement learning, that under an arbitrarily long time the UDMA-DQL technique converges with probability one to equilibrium policies in the non-stationary environment resulting from the distributed and uncoordinated operation of CRs. This analytical study is confirmed through simulation results showing that, in cases when an optimal policy can be identified, UDMA-DQL is able to find such policy in 99\% of cases for a sufficiently long learning time. Importantly, further simulations show that the presented UDMA-DQL approach achieves a much faster learning performance compared to an equivalent Table-based Q learning implementation. The second contribution in this dissertation introduces a PPO approach for UDMA resource allocation in the challenging case of non-stationary CR agents. The simulations demonstrate performance advantage of UDMA-PPO in the challenging case of a non-stationary environment because of multiple uncoordinated active CR agents that are learning through a shared wireless environment. The experimental results validate the stable and comparatively faster convergence to optimal solution of the UDMA-PPO with about 70\% and 25\% less training steps compared to equivalent Table-based and UDMA-DQL respectively. A third contribution in this dissertation is the incorporation of an attention mechanism into both the actor and critic networks to accelerate the learning performance of the PPO in the UDMA CRN setting. The results demonstrate that the attention mechanism accelerates learning resulting for the same number of exploration steps in PPO with attention achieving an optimal policy selection accuracy of 92\% compared to the 89\% accuracy without attention in a non-stationary environment, affirming that the PPO algorithm augmented with attention is capable of accelerated and accurate learning in a non-stationary environment.
Publication Date
6-2026
Document Type
Dissertation
Student Type
Graduate
Degree Name
Electrical and Computer Engineering (Ph.D)
College
Kate Gleason College of Engineering
Advisor
Andres Kwasinski
Advisor/Committee Member
Jamison Heard
Advisor/Committee Member
Panos Markopoulos
Recommended Citation
Tondwalkar, Ankita Vijay, "Reinforcement Learning-Enabled Resource Allocation for Distributed and Uncoordinated Cognitive Radio Networks" (2026). Thesis. Rochester Institute of Technology. Accessed from
https://repository.rit.edu/theses/12772
Campus
RIT – Main Campus
