| Dynamic spectrum allocation for interference avoidance is an important research direction in the field of electromagnetic countermeasures.In the current non-cooperative and strong electromagnetic interference environment,the existing wireless communication systems still widely used the fixed rules-based spectrum allocation schemes.Since these schemes lack flexibility and adaptability,they cannot meet the spectrum allocation requirements in the complex electromagnetic environments.Therefore,it is urgent to propose a spectrum dynamic allocation method that can adapt to various complex electromagnetic interference environments.This thesis focuses on the problem that traditional dynamic spectrum allocation methods are difficult to adapt to complex electromagnetic interference environments,combined with the reinforcement learning,the following research is conducted:1.A dynamic spectrum allocation method based on the basic reinforcement learning framework is proposed.The method applies the reinforcement learning to the field of interference-avoiding spectrum allocation,predicts the interference rules through reinforcement learning,and assigns the friendly spectrum in advance to avoid interference.Finally,compared with traditional statistically-based spectrum allocation methods,the reinforcement learning allocation method reduces the conflict rate by 57.57% under two kinds of single-frequency interference,two kinds of multi-frequency interference,and two kinds of sweep-frequency interference in simulations.2.In view of the problem of the high spectrum switching rate caused by directly applying reinforcement learning methods to the spectrum allocation field,which leads to frequency communication interruptions,this thesis proposes an interference-avoidance method based on cooperative reinforcement learning.The method defines two reinforcement learning processes: spectrum allocation process and collaborative process.The spectrum allocation process performs spectrum allocation to avoid interference,while the collaborative process modifies part of the reward signal of the spectrum allocation process by learning the statistical properties of the spectrum allocation process based on the trajectory of its statistical properties,so that the statistical properties of the spectrum allocation process can converge to low conflict rate and low spectrum switching rate.Compared with the reinforcement learning methods,the results suggest that the cooperative reinforcement learning methods can better balance the spectrum switching rate and the conflict rate.3.A multi-user spectrum allocation method based on distribution prediction reinforcement learning is proposed.In the current field of multi-user spectrum allocation,not only the avoidance of enemy interference but also the interference between friendly users need to be considered.However,commonly used multi-agent reinforcement learningbased multi-user spectrum allocation methods only focus on the avoidance of enemy interference and ignores the influence of the interference between users,leading to poor performance.To solve this problem,this thesis proposes a multi-user spectrum allocation method based on the distribution prediction reinforcement learning.The method predicts the spectrum distribution at the next moment to achieve multi-user spectrum allocation.Since different users cooperate to use frequencies through a spectrum distribution,it can effectively reduce interference between users and leads to better performance in the field of multi-user spectrum allocation than the multi-agent reinforcement learning.Finally,under simulated sweep-frequency interference,the method reduces the conflict rate by32.75% compared with the multi-agent reinforcement learning method based on independent Q-learning and by 12.62% compared with the multi-agent reinforcement learning method based on the value decomposition networks. |