Research on Intelligent UAV Path Navigation Technology Based on Radio Map Reconstruction

In the context of the rapid advancement of sixth-generation (6G) mobile communication systems, integrated terrestrial, aerial, and satellite networks have gradually become a key technology for achieving global seamless coverage and high-quality communication services. Within this architecture, unmanned aerial vehicles (UAVs) play an indispensable role as a bridge connecting the air, ground, and other domains. The study of communication technologies for these platforms is particularly important. UAVs offer unique advantages in providing flexible communication services, executing emergency rescue tasks, and supporting smart city construction, making them a vital driving force for the evolution of future communication networks.

In the integrated terrestrial-aerial-satellite framework, the development of UAV communication technologies can greatly expand the coverage of communication networks and enhance their flexibility and robustness. Moreover, the integration of communication and sensing, a significant direction in future communication systems, provides new insights for UAV path navigation. By enabling UAVs to sense the environment in real time and collect data, this integration supports adaptive adjustment and optimization of the network, further enhancing the intelligence and efficiency of communication systems.

Despite notable progress in UAV communication technology, challenges remain in path planning in complex environments and accurate prediction of signal propagation. Especially in dense urban environments, ensuring safe and efficient path navigation while maintaining stable communication is a pressing problem. To address these challenges, map-reconstruction-based UAV path navigation methods have emerged as a powerful solution. Such methods combine map reconstruction technology with deep reinforcement learning (DRL) algorithms to significantly improve the efficiency and stability of path navigation. By integrating image reconstruction techniques with path planning algorithms, it is possible to accurately predict UAV behavior in complex environments and optimize its path to maximize cellular connection efficiency.

In the context of 6G and integrated sensing and communication, combining radio map reconstruction with DRL algorithms for UAV path navigation and communication stability optimization is particularly important. By merging these two technologies, the navigation performance of UAVs in dynamically changing environments can be significantly improved while ensuring stable and efficient communication links. Compared with traditional path planning solutions, this approach adapts better to environmental changes, substantially reduces the risk of communication outages, and thereby increases the reliability and efficiency of UAV mission execution.

Deep learning and artificial intelligence (AI) have provided new perspectives and methods for research in UAV path navigation and communication. Deep learning, as an important branch of AI, automatically learns and extracts complex features from large amounts of data using deep neural networks. It has demonstrated extraordinary capabilities in image recognition, speech processing, natural language understanding, and many other fields. In particular, deep learning shows great potential for handling complex nonlinear problems, high-dimensional data analysis, and pattern recognition.

In the field of UAV path navigation and communication optimization, deep learning techniques can be used for high-precision radio map reconstruction, path planning in complex environments, and optimal allocation of communication signals. Through deep learning models, UAVs can more accurately understand and adapt to environmental changes, enabling dynamic planning and real-time decision-making, thereby improving the efficiency and safety of mission execution. In addition, the powerful computing and self-learning capabilities of AI can further enhance the autonomy and intelligence of the system, providing new directions and possibilities for the development of UAV technology.

Within this background, this thesis carries out research on intelligent UAV path navigation technology based on radio map reconstruction. The core of this research is to develop a novel radio-map-assisted path navigation method. The method first performs sparse sampling of the wireless environment using a UAV, then reconstructs a radio map based on the sampled data to estimate the outage probability at all positions in the target area. This approach greatly reduces the cost of map sampling. Through deep learning techniques, the UAV can optimize its path navigation strategy over multiple flights using DRL algorithms, significantly improving mission efficiency while reducing flight time and energy consumption. By combining radio map reconstruction with deep reinforcement learning, this thesis not only provides a new development path for UAV technology but also offers strong technical support for the design and implementation of integrated terrestrial-aerial-satellite communication systems.

Literature Review and Research Status

In this section, I first classify radio map reconstruction techniques and introduce several representative algorithms, summarizing the latest domestic and international research progress in this field. Then, I review related work on UAV path navigation to provide a solid foundation for further optimization of cellular-connected UAV online path navigation.

Radio Map Reconstruction

Radio map reconstruction is an important step in UAV path navigation. It aims to model and predict radio signals in the environment, enabling UAVs to better sense the surrounding environment and improving the accuracy and robustness of path planning. Existing radio map reconstruction methods can be divided into interpolation-based methods and statistical-model-based prediction methods. Interpolation-based methods collect actual signal data in the environment and use interpolation algorithms to estimate signal strength in unknown areas. Common interpolation algorithms include Kriging interpolation and inverse distance weighting interpolation. Statistical-model-based prediction methods, on the other hand, model the signal propagation characteristics in the environment and use statistical methods to predict signal propagation at different locations, thereby generating a map.

With the development of machine learning and deep learning technologies, more and more research has applied these methods to radio map reconstruction. For example, using deep learning models to model signal propagation in the environment can more accurately predict the distribution of signals in space, improving the accuracy and reliability of maps. In addition, big data and artificial intelligence-based radio map reconstruction methods have become a current research hotspot. By analyzing and mining large amounts of signal data, it is possible to better understand the propagation laws of signals in the environment, providing richer information for map reconstruction.

In wireless communication systems, understanding the distribution of wireless signals in space is critical for optimizing system performance. The complexity of wireless communication environments is one of the main challenges. Mobile terminals, building structures, human bodies, and other elements lead to non-uniform signal distribution in space. Precisely modeling and understanding wireless signal propagation characteristics has become a necessary prerequisite for communication system design. Traditional empirical models are insufficient when facing these challenges. Therefore, using deep learning methods to construct radio maps has become a cutting-edge research direction. By leveraging the powerful feature extraction and representation learning capabilities of neural networks, it is possible to capture the variation patterns of signals in space more accurately, describe the propagation characteristics of signals, and thereby provide finer spatial information for communication system optimization.

Deep learning-based radio maps not only have potential applications in indoor positioning and navigation, but also have broad application prospects in smart transportation, the Internet of Things (IoT), and industrial automation. By obtaining a more precise understanding of wireless signal distribution, it is possible to achieve positioning, tracking, and communication optimization for mobile terminals and devices, thereby promoting the future development of communication systems. In addition, deep learning methods can not only improve the modeling effect of static environments but also adapt to changes in dynamic environments. In wireless communication systems, especially for scenarios such as mobile users and UAVs, the adaptability and generalization capabilities of deep learning provide strong support for improving system performance.

In this research context, the problem of radio map reconstruction becomes particularly important, involving the application of image restoration techniques. However, in actual image restoration, learning-based methods and non-learning-based methods are two general categories commonly used. Learning-based methods are direct: they train deep convolutional networks using noisy images as input data and original images as output data. Non-learning-based methods, or manually crafted prior-based methods, forcibly add and inform the model of what types of images are natural and realistic. Mathematically expressing naturalness is very difficult. In the proposed untrained-network-prior-based radio map reconstruction, only a convolutional neural network is used to construct a new non-learning-based method that can bridge the gap between these two general image restoration approaches.

Specifically, deep convolutional networks have become a popular tool for image generation and restoration. By learning real image priors from a large number of example images, these networks achieve excellent performance. However, in contrast, the structure of the generator network is sufficient to capture a large amount of low-level image statistics before learning. This inherent prior property of the network provides a new solution for image reconstruction without requiring training sets or undamaged original images.

Although the field of radio map reconstruction has made certain progress, it still faces challenges. Among them, data sparsity is a prominent issue. Because signal data in the environment is affected by multiple factors, such as terrain and building occlusion, data collection is difficult, limiting the accuracy and coverage of map reconstruction. In addition, the dynamic nature of the environment also brings challenges to map reconstruction, requiring maps to be updated regularly to ensure they can accurately reflect the latest state of the environment.

In UAV path navigation, maps used by traditional navigation systems often cannot provide sufficient information when dealing with complex and changing environments. To solve this problem, introducing radio maps becomes an effective way to improve path navigation accuracy and can provide efficient and accurate navigation information for cellular-connected UAVs. Radio map reconstruction is a technology that uses wireless signal data to perform high-precision restoration and updates of geospatial information. The key lies in collecting and analyzing radio signals and then building an accurate map database, providing more precise position data and navigation assistance for UAV navigation, significantly improving navigation performance. This technology mainly covers sensor technology, signal processing and data fusion technology, as well as machine learning and artificial intelligence methods, together constituting the technical foundation of radio map reconstruction.

In terms of sensor technology, the integrated application of multiple sensors such as Global Positioning System (GPS), inertial navigation systems, and visual sensors provides UAVs with more comprehensive and multi-source map information. High-precision sensors can not only achieve accurate positioning but also capture diverse features of the surrounding environment, providing a basis for refined map reconstruction. In signal processing and data fusion, effective signal processing algorithms can reduce interference and improve signal quality. Data fusion integrates multi-source data, including sensor data, map databases, and communication signals, directly affecting the accuracy and robustness of the map. In addition, machine learning and artificial intelligence technologies play an increasingly important role in radio map reconstruction. Through deep learning and other technologies, systems can learn map features from large amounts of data, further improving map accuracy. Neural network models can adaptively adjust the way maps are constructed to adapt to navigation requirements in different environments.

Traditional image super-resolution and denoising techniques such as bicubic interpolation, nearest-neighbor interpolation, and non-local means have been widely used. However, these traditional methods have limitations in understanding image semantic information and detail restoration. With the development of deep learning, deep-learning-based image reconstruction and denoising methods have received attention due to their strong generalization and adaptability. These methods usually have stronger generalization ability and adaptability, but they also require larger computational resources and datasets. To address these challenges, image prior algorithms based on untrained network priors have emerged. This technique introduces deep learning ideas, using large-scale unlabeled image data for prior learning to improve the understanding and expression of image features.

In untrained-network-prior algorithms, image denoising extracts local texture and gradient distribution information, reducing noise without relying on pre-trained neural networks. Similarly, this technique has been applied to image super-resolution, generating high-resolution images from low-resolution images using the image’s own prior knowledge without requiring large pre-trained image pairs. Several studies have explored adaptive sparse domain selection and regularization for image deblurring and super-resolution, trainable nonlinear reaction diffusion models for image restoration, deep image prior-based design surface models for optical imaging, and unsupervised deep learning frameworks for dynamic positron emission tomography imaging. These works demonstrate the versatility and effectiveness of untrained-network-prior techniques.

The advantages of untrained-network-prior technology are significant. This algorithm is particularly suitable for specific fields or scenarios, effectively avoiding overfitting, and can function in environments with limited computational resources. It does not rely on large amounts of pre-trained data, showing strong generality. Its disadvantages are that its accuracy is usually lower than traditional deep learning methods, it may perform poorly in complex scenes, and it relies on manually extracted image priors, making it sensitive to prior selection. However, compared with traditional methods, untrained-network-prior algorithms can better handle image semantic information, improve the restoration of image details and textures, and make super-resolution results more natural and realistic.

UAV Path Navigation

UAV path navigation plays a vital role in UAV systems. It aims to plan and execute flight paths autonomously or semi-autonomously, improving autonomy, stability, and safety to meet various application requirements. In cellular-connected systems, UAVs can not only expand communication coverage but also improve task flexibility and efficiency. Therefore, UAVs have been widely used in cargo transportation, aerial video streaming, virtual reality, and augmented reality. By deeply integrating UAVs with ground base stations (GBSs), cellular-connected UAVs can efficiently perform intelligent network control and data processing. In addition, cellular-connected UAVs can achieve dense cellular communication coverage to meet communication network service requirements. First, for UAV-assisted cellular communication systems, UAVs can act as relays for communication connections. For example, when ground base stations fail, UAVs can be quickly deployed to provide emergency communication support for ground users. Second, for cellular-network-supported UAV systems, UAVs can complete flight tasks by maintaining communication with ground base stations. Considering the high mobility and high-speed flight of UAVs, as well as the large amount of data transmission between UAVs and ground users, establishing high-quality air-ground communication links is essential.

However, UAV flight time is often limited by battery constraints. To ensure stable and continuous communication connections between UAVs and ground base stations, and to ensure the reliable completion of UAV flight missions, it is necessary to study energy-efficient cellular-connected UAV systems. Moreover, traditional UAV path navigation methods usually rely on static maps and rule-based models, which are difficult to adapt to dynamically changing environments. Traditional communication technologies applied to UAV communication are also limited, affecting communication quality and real-time navigation. Existing UAV path navigation methods are mainly based on static maps or rules, and the application of cellular connections is mostly limited to data transmission. This leads to insufficient path navigation and communication management in complex dynamic environments.

Traditional navigation methods are usually based on GPS. However, in cities, forests, or other complex environments, GPS signals may be interfered with and cannot provide accurate position information. Therefore, researchers are actively exploring more advanced navigation systems, including but not limited to visual navigation, inertial navigation, and radio-map-assisted navigation. Visual navigation uses visual sensors, such as cameras, to analyze real-time images to perceive the surrounding environment, and uses computer vision techniques for landmark recognition and obstacle avoidance. Inertial navigation systems use accelerometers and gyroscopes to measure acceleration and angular velocity, providing relatively accurate navigation information when GPS signals are unavailable. Radio-map-assisted navigation uses signals transmitted by ground base stations or satellites, including global navigation satellite systems and other radio signal positioning technologies.

To achieve effective path navigation, researchers continue to develop various path navigation algorithms, taking into account mission requirements, environmental conditions, and UAV dynamics. These algorithms include search-based and optimization-based methods, such as A* algorithm, Dijkstra algorithm, genetic algorithm, and others. Meanwhile, continuous innovation in sensor technology is also a key factor for navigation. Researchers are committed to fusing high-precision sensors such as GPS, inertial measurement units (IMUs), and visual sensors to improve navigation robustness and accuracy.

In terms of communication technology and cellular network connectivity, the stability and reliability requirements of UAV communication systems are extremely high. With the development of 5G technology, higher communication bandwidth and lower latency are now available, providing better conditions for UAV communication. The combination of cellular connection technology allows UAVs to access real-time map data and weather information through mobile networks, enabling more intelligent path planning. Research on map data and radio maps is crucial for path navigation. High-quality map data is the foundation of path navigation. Therefore, researchers are committed to developing high-resolution, real-time updated map data to provide more accurate navigation support. Through radio maps, UAVs can perceive signal strength, interference, and other information in the environment, enabling them to better adapt to complex navigation environments.

To address the high complexity of UAV path navigation, many research methods based on deep reinforcement learning and radio maps have emerged. Some studies focus on improving the efficiency of UAV trajectory optimization, using deep learning schemes such as DRL-based algorithms to obtain optimal trajectories and throughput. Other studies consider multiple aspects, such as user coverage, energy saving, and geographic fairness, and use multi-agent DRL trajectory planning algorithms for optimization. Several papers have explored UAV navigation in various environments, such as dynamic unstructured environments, large-scale complex environments, and resource-limited environments. For instance, a hierarchical reinforcement learning framework combining high-level behavior selection with low-level obstacle avoidance and goal-driven control has been proposed, effectively eliminating the need for manually designed control rules and significantly enhancing autonomous navigation capabilities. Another study formulated navigation as a partially observable Markov decision process and developed an online DRL algorithm based on policy gradient theorem, realizing direct mapping from raw sensory data to navigation control signals. In resource-limited environments, a combination of proportional-integral-derivative control and proximal policy optimization demonstrated high performance in complex navigation tasks.

Radio maps play a crucial role in UAV path planning, especially in improving navigation accuracy and flight safety. First, with the help of precise radio map data, UAVs can gain an in-depth understanding of their flight environment, enabling higher-precision navigation and path planning. This is particularly important for improving flight safety and reliability. This capability is especially significant in challenging urban environments, helping UAVs avoid collisions and improve mission efficiency. Second, by integrating radio map information, UAVs can effectively identify and avoid obstacles, preventing potential collision risks. UAVs can accurately perceive the surrounding terrain, buildings, and obstacles, and proactively avoid possible dangers, ensuring flight safety. This capability is crucial for applications in scenarios such as urban logistics and emergency rescue, significantly improving the reliability of UAVs in complex environments.

At the technical implementation level, key technologies include map data integration, sensor fusion, obstacle avoidance algorithms, real-time updating and communication, and path planning algorithm optimization. These technologies jointly promote the accuracy and efficiency of UAV path navigation. Environmental perception and adaptability research focuses on the intelligent perception and adaptive adjustment of UAVs to dynamic environments. Through real-time map updating, sensor fusion, and adaptive algorithms, UAVs can flexibly adapt to different weather conditions, dynamic obstacles, and traffic flow, ensuring flight safety and efficiency. Specific application scenarios such as urban logistics, emergency rescue, and agricultural monitoring demonstrate the practical effects of these technologies. Through dynamic obstacle detection, UAVs in urban logistics can identify buildings, vehicles, and pedestrians in real time, improving efficiency and avoiding collisions. In emergency rescue, combined with real-time map updates and dynamic obstacle detection, UAVs adapt to changes in disaster areas, improving search and rescue efficiency. In agricultural monitoring, weather adaptability technologies ensure the accuracy of UAV missions under different weather conditions. These research and application examples highlight the indispensable importance of environmental perception and adaptability technologies for UAV path navigation with the assistance of radio maps.

Several recent studies have shown how radio map assistance can further enhance UAV communication performance, flight safety, and path navigation efficiency. Some works have explored constructing radio maps for air-to-ground communication channels, proposing a learning framework that decomposes the map into structural and non-structural components to predict path loss and improve the construction efficiency of virtual obstacle maps. Others have focused on fusing radio signal strength and environmental depth information to construct radio maps in flying UAV networks, predicting the received signal strength between UAV-mounted base stations and ground users. In the context of anti-jamming communications, signal power maps and signal-to-interference-plus-noise ratio (SINR) maps have been used to assist path planning algorithms, improving the quality of UAV communications. A simultaneous radio and physical mapping method has been proposed, utilizing UAV-collected radio and sensing data to build maps in parallel, effectively improving the accuracy and efficiency of map construction. Model-driven deep learning frameworks have also been introduced for UAV-assisted radio map learning and 3D environment reconstruction. Additionally, six-dimensional radio maps based on received signal strength measurements have been constructed to predict channel gains between arbitrary transmitter and receiver positions. Radio maps have been used to capture large-scale channel gain characteristics for UAV path planning, using shortest path problems and grid quantization methods to compress data and improve path planning efficiency. Under resource constraints, map-compression-based node scheduling trajectory optimization algorithms have been proposed to effectively solve trajectory optimization problems. Furthermore, radio map-assisted estimators have been used to derive optimal UAV trajectories while accelerating the learning process under task time constraints.

In summary, the combination of radio map reconstruction and deep reinforcement learning provides a powerful framework for intelligent UAV path navigation. This thesis builds on these foundations to propose a novel method that integrates untrained-network-prior radio map reconstruction with DRL-based path navigation.

Background Knowledge and Algorithm Overview

In this section, I introduce the relevant technical background knowledge for this thesis, including the principles of radio map reconstruction, data acquisition and processing methods, common modeling and prediction algorithms, and an overview of UAV path navigation technology, especially deep reinforcement learning approaches.

Fundamentals of Radio Map Reconstruction

Radio map reconstruction is a technology that combines data acquisition, signal feature extraction, and map construction. It creates an electromagnetic property map of a specific area by analyzing the wireless signal distribution in that area. This technology is crucial for UAV path navigation because it not only enhances the UAV’s ability to perceive the surrounding environment but also significantly improves the accuracy and robustness of path planning.

In the initial data acquisition stage, wireless sensor devices or sensors mounted on UAVs are typically used to collect wireless signal data in a geographical area. These data can include various types of signals such as Wi-Fi, cellular networks, GPS, and others. For example, some researchers have used GPS and Wi-Fi signal strength data collected by vehicles and pedestrians to obtain signal information at different locations. In the signal feature extraction stage, the collected raw signal data contains a large amount of information, but not all of it is useful for map reconstruction. Therefore, signal preprocessing and feature extraction become a crucial step. Gaussian process regression methods have been used for feature extraction and modeling of collected signal data to achieve accurate signal modeling in geographic areas.

Finally, in the map construction stage, interpolation algorithms or prediction models are used to estimate and reconstruct the signal distribution in the geographic area based on the extracted signal features and location information, thereby constructing a radio map. For instance, radial basis function interpolation methods have been applied to spatially interpolate signal strength data, yielding high-precision map reconstruction results. Radio maps provide UAVs with an efficient environmental awareness method, enabling them to make more accurate decisions during path planning and navigation tasks based on real-time signal environment information. This is of great significance not only for improving the autonomous flight capability of UAVs in complex environments but also for opening up new avenues for applications in urban planning, disaster response, and other fields.

Data Acquisition and Processing for Radio Map Reconstruction

In radio map reconstruction, data acquisition and processing are vital. Data acquisition methods include field measurement and database retrieval. Field measurement involves deploying sensor devices in the geographic area or using mobile devices for real-time collection. Database retrieval uses existing wireless signal databases for data acquisition. The data processing stage involves signal preprocessing, feature extraction, data cleaning, and calibration. Feature extraction and signal calibration are the key and difficult points. For example, deep learning methods have been used to extract features and model collected signal data, achieving efficient processing and reconstruction of signals in a geographic area. In addition, adaptive filtering and beamforming techniques have been used to preprocess and clean wireless signal data, improving data quality and reliability.

Radio-map-based UAV path navigation technology aims to improve the precision and robustness of UAV path navigation by sensing the radio signal map in the environment. The key technical implementation process includes signal acquisition and preprocessing, map construction and localization, path navigation and obstacle avoidance, real-time updating and adaptation, and communication and system integration.

First, a UAV equipped with a radio receiver performs signal acquisition to obtain radio signals in the environment. The collected signals are then preprocessed, such as filtering and denoising, to extract effective information. Second, a signal fingerprint map is built by mapping the signal strength and characteristics of different positions into a map database, enabling accurate localization of the UAV. Localization algorithms compare the real-time collected signal features with the information in the map database to determine the UAV’s position.

In the path navigation and obstacle avoidance stage, the signal fingerprint map is used for path planning, considering factors such as signal quality and obstruction. Obstacle avoidance strategies are integrated to dynamically adjust the path through the real-time sensed signal environment to avoid collisions with obstacles. Real-time updating and adaptation are important components, including continuously updating the signal fingerprint map during flight to adapt to environmental changes, and using adaptive algorithms to adjust positioning and path navigation strategies based on real-time signal information, improving navigation robustness.

Finally, an efficient UAV communication and system integration is essential. By integrating wireless communication modules, UAVs can transmit and receive signal information in real time. Additionally, integrating the signal navigation system with other UAV systems enables more comprehensive path navigation capabilities.

This radio-map-based UAV path navigation technology is suitable for urban or GPS-limited areas, has high adaptability and accuracy, and can effectively perform urban cruising and search-and-rescue missions. Future research directions include optimizing signal processing algorithms, improving system adaptability, and enhancing robustness to multi-path environments to further promote the development of this field.

Common Modeling and Prediction Algorithms in Radio Map Reconstruction

Common modeling and prediction algorithms in radio map reconstruction include Kriging interpolation, radial basis function interpolation, and Gaussian process regression. These algorithms use known signal strength data to infer signal strength at unknown locations and perform interpolation or prediction based on spatial relationships. In addition to Gaussian process regression, many other algorithms are widely applied in radio map reconstruction.

For traditional Kriging interpolation reconstruction algorithms, estimation errors are first given, spatial variable correlations are fully considered, and the clustering effects existing in the dataset are effectively compensated, resulting in high interpolation accuracy. As the core of geostatistics, the Kriging method is used to estimate the attribute values of unsampled positions. Its research object is the regionalized variable, and it is an optimal unbiased estimation method. The Kriging method mainly consists of two parts: the first part is quantifying the spatial correlation of known data, with the variogram as its tool. Generally, the closer the points, the more similar their attribute values; the farther the points, the more different their attribute values. However, for different research objects, the degree of variation with distance is not the same. The second part is determining the neighborhood range, searching for neighborhood points, establishing Kriging equations, solving the equations to obtain weight coefficients, and using the weighted summation to obtain the attribute value of the point to be estimated.

The first part is the most important part of the Kriging technique, covering three steps. First, the experimental variogram is calculated using the relative positions and attribute values of known points. Then, one or more theoretical variograms are selected from the existing theoretical variograms according to the shape of the experimental variogram scatter plot and the actual characteristics of the geological variables. Finally, the existing fitting algorithm and the selected theoretical variogram are used for fitting. The variogram ultimately reflects the spatial randomness and correlation of the studied variable. Therefore, the accuracy of each step affects the accuracy of the final interpolation result.

In the research area, the variable \(Z(x)\) takes values \(Z(x_i)\) at sampling points \(x_i (i=1,2,3,\ldots,n)\). The variable value at an unknown point \(Z(x)\) is obtained by the weighted sum of the above \(n\) known variable values, expressed as:

\[
Z(x) = \sum_{i=1}^{n} \lambda_i Z(x_i)
\]

where \(\lambda_i (i=1,2,3,\ldots,n)\) are the weight coefficients to be determined, representing the contribution of each spatial sample point observation value to the estimated value. The weight coefficients must satisfy two conditions: first, the estimate is unbiased, meaning the mathematical expectation of the deviation is zero; second, it is optimal, meaning the sum of squares of the differences between the estimated value and the actual value is minimized. That is, the Kriging system equations are expressed as:

\[
\begin{cases}
\sum_{j=1}^{n} \lambda_j \gamma(x_i, x_j) + \mu = \gamma(x_i, x) \quad i=1,2,\ldots,n \\
\sum_{i=1}^{n} \lambda_i = 1
\end{cases}
\]

where \(\lambda_i\) are the weight coefficients, \(\mu\) is the Lagrange constant, and \(x_i\) and \(x\) represent the positions of the \(i\)-th sample and the unknown point, respectively. They are used to calculate the semi-variogram function \(\gamma(h)\), which is expressed as:

\[
\gamma(h) = \frac{1}{2} E\left[ (Z(x) – Z(x+h))^2 \right]
\]

By solving the Kriging equations, the weight coefficients can be obtained, and the estimated value can be computed through the expression above. Compared with other interpolation methods, the Kriging technique has its own advantages. In the Kriging technique, the contribution of known points to the point to be interpolated is not only related to the distance between them but also to the spatial configuration of all known points. In addition, geostatisticians can incorporate their own understanding of the research variable during the data estimation process, thereby improving the model in certain steps.

With the development of technology, more advanced methods for improving the efficiency and accuracy of map reconstruction have emerged. For example, genetic algorithms and particle swarm optimization algorithms have been used to optimize model parameters, improving the reconstruction accuracy and generalization ability of models. Deep learning methods, such as convolutional neural networks and recurrent neural networks, have been used for modeling and prediction of signal data. These deep learning models can learn complex spatial features and achieve good results in map reconstruction. In addition, methods based on Bayesian inference and graph models, such as Bayesian networks and Markov random fields, have been applied to radio map reconstruction. These methods make full use of the spatial relationships and correlations between signals, thereby improving the accuracy and robustness of map reconstruction.

Radio Map Reconstruction Based on Untrained Network Priors

In this chapter, I focus on two innovative downsampling image map completion algorithms proposed in this thesis: the Radio Map Construction based on Deep Image Prior (RMC-DIP) and the Radio Map Construction based on Deep Denoising Regularization (RMC-DDR). The goal is to break the dependence of traditional map reconstruction techniques on large-scale labeled data and significantly enhance the efficiency and accuracy of map reconstruction. I first elaborate the basic principles and operational procedures of these techniques. Then, I discuss the implementation details and core technologies in depth. Finally, I analyze and evaluate the results through simulation experiments.

Map Reconstruction Model

In the radio map reconstruction process, the UAV first sparsely samples the actual environment and calculates the outage probability at the sampling points, then reconstructs and restores the radio map of the target area. Suppose the UAV randomly samples \(N\) data points in the target airspace, denoted as \(x_i (i=1,2,3,\ldots,N)\), where \(x_i\) represents the three-dimensional coordinates of the UAV. Then, according to the outage probability calculation formula, the outage probability at position \(x_i\), denoted as \(Z(x_i)\), is obtained. The set of outage probabilities obtained from \(N\) spatial positions in this area can be viewed as the pixels on the outage probability distribution map to be reconstructed. For simplicity, the degraded outage probability distribution map corresponding to the sparse sampling points is represented as \(y_0\). Since this chapter focuses on map reconstruction techniques and the next chapter introduces radio-map-assisted UAV path navigation, the specific outage probability calculation will be described in detail in the next chapter.

Based on the obtained outage probability set, I use the RMC-DIP algorithm and the RMC-DDR algorithm to reconstruct the radio map. RMC-DIP uses the random initialization of a neural network as an image prior. It utilizes the capacity of the network structure and the random initialization of parameters as prior knowledge to reconstruct a radio map highly consistent with the actual environment without supervision. RMC-DDR, on the other hand, combines the powerful modeling capability of RMC-DIP with the stability advantages of Regularization by Denoising (RED). Through the deep learning framework, RMC-DDR can not only adapt to the nonlinear propagation characteristics of wireless signals in complex environments but also effectively solve the data incompleteness problems introduced by sparse sampling and noise.

Problem Formulation

In the RMC-DIP algorithm, the radio map to be reconstructed is defined as an \(m \times n\) matrix \(\mathbf{R}\in\mathbb{R}^{m \times n}\), expressed as:

\[
\mathbf{R} = \begin{bmatrix}
R(x^{(0)},x^{(0)}) & R(x^{(0)},x^{(1)}) & \cdots & R(x^{(0)},x^{(n)}) \\
R(x^{(1)},x^{(0)}) & R(x^{(1)},x^{(1)}) & \cdots & R(x^{(1)},x^{(n)}) \\
\vdots & \vdots & \ddots & \vdots \\
R(x^{(m)},x^{(0)}) & R(x^{(m)},x^{(1)}) & \cdots & R(x^{(m)},x^{(n)})
\end{bmatrix}
\]

The sampling matrix obtained from the above radio map matrix can be expressed as:

\[
\mathbf{r} = \begin{bmatrix}
R(x^{(0)},x^{(0)}) & R(x^{(0)},x^{(1)}) & \cdots & R(x^{(0)},x^{(n)}) \\
R(x^{(1)},x^{(0)}) & R(x^{(1)},x^{(1)}) & \cdots & R(x^{(1)},x^{(n)}) \\
\vdots & \vdots & \ddots & \vdots \\
R(x^{(m)},x^{(0)}) & R(x^{(m)},x^{(1)}) & \cdots & R(x^{(m)},x^{(n)})
\end{bmatrix}
\]

The low-resolution input image is the sparse radio map sampled by the UAV. Define the sampling factor (the number of times the number of pixels in the length and width of the map to be reconstructed is reduced) as \(s\), then the sampled map is denoted as \(y_0 \in \mathbb{R}^{(m/s) \times (n/s)}\). Define the reconstruction factor (the ratio of the number of pixels after reconstruction to that before reconstruction) as \(u\), then the reconstructed map \(y\) is denoted as \(y \in \mathbb{R}^{(u m/s) \times (u n/s)}\). Therefore, the reconstruction task is to reconstruct the high-resolution image from the sampled radio map \(y_0\), and the reconstructed radio map \(y\) must be highly fitted to the actual environment.

RMC-DIP Algorithm

The degraded radio map obtained after sparse sampling is denoted as \(y_0\). Then, the UAV reconstructs the radio map based on \(y_0\). The radio map reconstruction is expressed as an energy minimization optimization problem, denoted as:

\[
y^* = \arg\min_{y} \Phi(y) + e(y; y_0)
\]

where \(e(\cdot)\) is the data term related to reconstruction. In the reconstruction task, the data term is:

\[
e(y; y_0) = \| d(y) – y_0 \|^2
\]

where \(d(\cdot): \mathbb{R}^{(u m/s) \times (u n/s)} \to \mathbb{R}^{(m/s) \times (n/s)}\) adjusts the image size to \((m/s) \times (n/s)\). Here, \(y\) represents the reconstructed radio map. The closer \(y\) is to \(y_0\), the smaller the value of \(e(y; y_0)\). The goal of map reconstruction is to find the optimal solution \(y^*\) of problem. In this thesis, the prior information hidden in the neural network is used to replace the regularization function, and the neural network mapping \(f_\theta(\cdot)\) is used to replace the map \(y\) to be reconstructed, that is:

\[
y^* = f_{\theta^*}(z)
\]

where

\[
\theta^* = \arg\min_{\theta} e(f_\theta(z); y_0)
\]

The optimal solution \(\theta^*\) can be obtained by stochastic gradient descent with random initialization of parameters, continuously iterating to find the optimal solution. Here, \(z\) is a fixed three-dimensional tensor containing 32 feature maps, and its spatial size is the same as \(y\). The input of the network is a randomly initialized \(z\). \(\theta\) is the network parameter, and the optimal value is obtained by training. After obtaining the optimal parameter \(\theta^*\), the input \(z\) gives the optimal \(y\), thereby obtaining the reconstructed radio map. The specific steps of the untrained-network-prior-based map reconstruction algorithm are shown in Algorithm 3.1.

Algorithm 3.1: RMC-DIP
1: Initialize the reconstruction matrix R, sampling factor s, reconstruction factor u, and deep convolutional network fθ with random parameters θ
2: Input the random encoding z of the degraded image y0 and the total number of iterations T
3: Feed the random code z into the deep convolutional neural network fθ(z)
4: Train parameters θ
5: Obtain the optimal solution for the problem min ||fθ(z) – y0||²
6: Complete the map reconstruction

The iterative process of the RMC-DIP algorithm is illustrated in the previous figure, which shows the partial iteration process of the radio map for the ground cellular network coverage in specific experiments. The figure depicts two cellular base stations, with blank areas representing coverage holes. The darker the color, the better the coverage. The input image is \(y_0\). As the network parameters \(\theta\) are continuously updated, the radio map mapped by the network becomes increasingly accurate.

RMC-DDR Algorithm

To further deepen the research on radio map reconstruction for UAV communications in complex environments, I further optimize the RMC-DIP algorithm proposed in the previous section and propose the RMC-DDR algorithm. RMC-DDR significantly improves the accuracy of radio map reconstruction from sparse sampling data by integrating deep denoising regularization technology with the Deep Image Prior (DIP) strategy.

The RMC-DDR algorithm builds on the deep image prior by introducing the powerful feature extraction capability of convolutional neural networks. It reconstructs radio maps through the image features learned by the network itself, without relying on external training datasets. This design takes into account the difficulty of obtaining large amounts of labeled training data in practical applications. At the same time, the algorithm effectively utilizes the local smoothness characteristics of images by combining denoising regularization strategies. With the help of a denoising function, it accurately captures this feature and guides the reconstruction process, thereby significantly reducing the impact of noise.

Considering the challenges in practical scenarios, including data sparsity and inevitable noise during measurement, the goal is to recover the accurate radio signal strength distribution \(y\) from damaged measurement data \(y_0\). Specifically, the given measurement data \(y_0\) can be modeled through the equation:

\[
y_0 = \mathbf{H} y + \nu
\]

where \(\mathbf{H}\) is the known linear degradation matrix that simulates the propagation and attenuation process of wireless signals in the environment, and \(\nu\) represents additive white Gaussian noise, reflecting random errors in the measurement process.

To recover \(y\) from these challenging data, we adopt the following optimization framework:

\[
\min_{y} \frac{1}{2} \| y_0 – \mathbf{H} y \|_2^2 + \lambda \Phi(y)
\]

where the first term is the data fidelity term that quantifies the match between the reconstructed signal and the observed data; \(\Phi(y)\) is the regularization term that stabilizes the solution of the inverse problem and suppresses the influence of noise by introducing prior knowledge of the signal \(y\) (such as smoothness, sparsity, or specific statistical properties); \(\lambda\) is a regularization parameter that balances the data fidelity term and the regularization term.

In the RMC-DIP algorithm, the regularization term \(\Phi(y)\) is removed, and the reconstruction problem is expressed as finding the minimum of \(\| y_0 – \mathbf{H} y \|_2^2\). It is assumed that \(y\) is the output of the network, expressed as:

\[
y = C_\theta(z)
\]

where \(z\) is a fixed random vector, and \(\theta\) denotes the network parameters to be learned. Therefore, the above expression can be transformed into:

\[
\min_{\theta} \| y_0 – \mathbf{H} C_\theta(z) \|_2^2
\]

However, for image reconstruction problems, effective regularization is very important. Since a small amount of noise at one point in the image may have a large impact on the reconstruction result, a regularization term \(\Phi(y)\) is added to the loss function to maintain image smoothness, representing the general prior on natural images. Therefore, based on RMC-DIP, we further investigate a method that uses denoising algorithms to construct regularization. Specifically, RED is combined with RMC-DIP. In RED, the regularization function is defined as:

\[
\Phi(y) = \frac{1}{2} y^T \left( y – f(y) \right)
\]

where \(f(\cdot)\) is the selected denoiser. In summary, the RED-DDR algorithm reformulates the problem as:

\[
\min_{\theta} \frac{1}{2} \| y_0 – \mathbf{H} C_\theta(z) \|_2^2 + \frac{\lambda}{2} \left[ C_\theta(z) \right]^T \left( C_\theta(z) – f\left( C_\theta(z) \right) \right)
\]

From the above content, it can be seen that if the problem is solved by backpropagation through \(C_\theta(z)\), it is necessary to differentiate the denoiser \(f(\cdot)\). For most denoisers, this is difficult to achieve. To solve this problem, the Alternating Direction Method of Multipliers (ADMM) is used, converting \(y = C_\theta(z)\) into a penalty term, using the Augmented Lagrangian (AL):

\[
\min_{\theta, y} \frac{1}{2} \| y_0 – \mathbf{H} y \|_2^2 + \frac{\lambda}{2} y^T \left( y – f(y) \right) + \frac{\mu}{2} \| y – C_\theta(z) \|_2^2 – v^T \left( y – C_\theta(z) \right)
\]

where \(v\) represents the Lagrange multiplier vector of the equality constraint set, and \(\mu\) is an optional parameter. By merging the last two terms, the scaled form of AL is obtained:

\[
\min_{\theta, y} \frac{1}{2} \| y_0 – \mathbf{H} y \|_2^2 + \frac{\lambda}{2} y^T \left( y – f(y) \right) + \frac{\mu}{2} \| y – C_\theta(z) – v \|_2^2
\]

The ADMM is equivalent to sequentially updating the three unknowns \(\theta\), \(y\), and \(v\) in this expression, finally obtaining the optimal solution. The iterative reconstruction process of the RMC-DDR algorithm is similar to the RMC-DIP process. Compared with RMC-DIP, RMC-DDR introduces residual calculation and parameter update mechanisms. These two steps are closely interwoven in each iteration to improve reconstruction accuracy and reduce noise impact.

In the residual calculation step, the algorithm evaluates the difference between the current network output and the target image, using this residual to guide the next parameter update. This not only accelerates the convergence process but also ensures that images can be effectively recovered even under sparse or noisy data conditions. In the parameter update stage, the RMC-DDR algorithm uses residual and regularized denoising information to optimize the network parameters, ensuring that each iteration moves in the direction of reducing the overall error. This strategy not only significantly improves the prediction accuracy of cellular network coverage but also enhances the accurate restoration of the outage probability distribution in the actual environment.

The specific steps of the RMC-DDR algorithm are shown in Algorithm 3.2.

Algorithm 3.2: RMC-DDR
1: Input the image to be reconstructed \(y_0\)
2: Initialize z
3: Alternately update \(C_\theta(z)\) to minimize the difference between the reconstructed image and the sampling data
4: Apply the denoising function \(f(\cdot)\) to update y, enhancing the smoothness of the reconstructed image
5: Update the Lagrange multiplier vector through the ADMM method to satisfy the constraint conditions

In the above Algorithm 3.2, the RMC-DDR algorithm is introduced. This method not only relies on the image features learned by the convolutional neural network itself for map reconstruction but also reduces the impact of noise through the denoising function, thus effectively recovering images under sparse or noisy data conditions. However, to handle data sparsity and inevitable noise in practical applications, an optimization framework that can precisely handle such constraint conditions is required. Algorithm 3.3 details the ADMM optimization process, describing how to sequentially update network parameters, image reconstruction variables, and Lagrange multipliers through ADMM to satisfy the equality constraints in the map reconstruction process.

Algorithm 3.3: ADMM Optimization Process
1: Input: damaged measurement data \(y_0\), linear degradation matrix \(\mathbf{H}\), additive white Gaussian noise level \(\nu\), regularization parameter \(\Phi\), ADMM parameter \(\mu\)
2: Initialization: set iteration count k=0, initialize network parameters \(\theta^0\), image reconstruction variable y, Lagrange multiplier \(v^0\)
3: Repeat the following steps until convergence:
a. Update network parameter \(\theta\): solve the optimization problem using current y and v to update \(\theta\)
b. Update image reconstruction variable y: based on the current \(\theta^{k+1}\) and \(v^k\), and denoising regularization, update y
c. Update Lagrange multiplier v: update v to further satisfy the constraint conditions
4: Output: optimized image reconstruction variable y and network parameter \(\theta\)

Simulation Design and Performance Analysis

To evaluate the performance of the proposed RMC-DIP and RMC-DDR algorithms in improving map reconstruction accuracy and efficiency, a series of simulation experiments were designed. A radio map of a 2 km × 2 km area containing high-rise buildings was used as the reconstruction input. It is assumed that multiple cellular base stations are distributed in the area. The pixel value in the radio map represents the probability of connection outage between the UAV and the ground cellular base station.

The experimental process follows two core steps: first, sparse sampling and outage probability calculation. In this stage, by simulating UAV sparse sampling in a specified environment, outage probability data at randomly sampled positions are collected, and an outage probability distribution map is constructed. Second, high-resolution reconstruction of the radio map is performed, applying the RMC-DIP and RMC-DDR algorithms respectively to achieve precise map reconstruction.

To ensure the reliability and fairness of the experimental results, the Kriging algorithm based on the spherical variogram model was introduced as a comparative algorithm to comprehensively evaluate the performance of the proposed algorithms. The evaluation metrics include Peak Signal-to-Noise Ratio (PSNR), Structural Similarity Index (SSIM), and Normalized Mean Squared Error (NMSE), used to quantify the similarity and accuracy between the reconstructed map and the actual environment map.

Figure 3.3 shows the UAV reconstructed maps under the RMC-DIP algorithm with reduction factors of 16, 8, 4, and 2 times respectively (denoted as DIP16, DIP8, DIP4, and DIP2), as well as the reconstruction map with a 4-fold reduction using the spherical variogram model Kriging algorithm (denoted as SPH4) and the actual environment map. The larger the undersampling factor, the lower the resolution of the reconstructed outage probability distribution map. As the undersampling factor increases, the resolution of the outage probability distribution map improves and can better match the outage probability in the actual environment. Comparing DIP4 and SPH4, it can be inferred that under the same undersampling factor, the quality of the outage probability distribution map reconstructed by the proposed RMC-DIP algorithm is significantly higher than that of the Kriging algorithm based on the spherical variogram model, and it can almost fit the actual outage probability distribution map.

Table 3.1 presents the performance comparison of Kriging, RMC-DIP, and RMC-DDR algorithms under different sampling rates, specifically 1%, 5%, 10%, and 25% sampling rates, with evaluation metrics including PSNR, SSIM, and NMSE.

Table 3.1: Performance comparison of Kriging, RMC-DIP, and RMC-DDR at different sampling rates
Sampling rate PSNR (dB) Kriging PSNR (dB) RMC-DIP PSNR (dB) RMC-DDR SSIM Kriging SSIM RMC-DIP SSIM RMC-DDR NMSE Kriging NMSE RMC-DIP NMSE RMC-DDR
1% 10.6870 18.1926 18.2038 0.0106 0.4479 0.4635 0.2540 0.1569 0.1558
5% 10.5420 20.0830 20.3540 0.0051 0.6395 0.6894 0.2627 0.1015 0.0954
10% 10.3642 21.6606 22.0747 0.0107 0.7719 0.7875 0.2736 0.0706 0.0641
25% 10.7023 23.8576 24.2434 0.0116 0.8625 0.8667 0.2378 0.0424 0.0389

From Table 3.1, it can be observed that at all sampling rates, the PSNR performance of RMC-DIP and RMC-DDR is significantly better than that of the Kriging algorithm. As the sampling rate increases, the PSNR values of these two algorithms increase significantly. RMC-DDR is always slightly better than RMC-DIP, especially at a 25% sampling rate, where RMC-DDR’s PSNR is 13.7114 dB higher than Kriging and 0.5777 dB higher than RMC-DIP. In terms of SSIM, RMC-DDR is slightly better than RMC-DIP at all sampling rates, and both are significantly better than Kriging. For example, at a 5% sampling rate, RMC-DDR’s SSIM is 0.6894, while Kriging is only 0.0051. For NMSE, all algorithms improve as the sampling rate increases, but RMC-DDR consistently shows the smallest error, indicating its superior performance in reconstruction accuracy. For instance, at a 10% sampling rate, RMC-DDR achieves an NMSE of 0.0641, while Kriging and RMC-DIP are 0.2736 and 0.0706, respectively.

Additionally, comparing the PSNR performance under different downsampling factors (SF10, 5, 3, and 2), the PSNR of the RMC-DDR algorithm stabilizes quickly after the initial 100 iterations, indicating rapid convergence characteristics. Moreover, compared with RMC-DIP, RMC-DDR consistently achieves higher final PSNR values under all downsampling conditions, demonstrating the clear advantage of RMC-DDR in image reconstruction quality. The SSIM comparison shows that RMC-DDR outperforms RMC-DIP under all downsampling factors, especially at lower factors (3 and 2), where RMC-DDR quickly converges to higher stable SSIM values. The NMSE comparison shows that after the initial rapid decline, RMC-DDR maintains lower NMSE values under all downsampling factors, showing a significant advantage in maintaining reconstruction accuracy.

The experimental results fully demonstrate the remarkable performance of the RMC-DDR algorithm in radio map reconstruction, especially in handling sparse data and noise. This validates the potential of the RMC-DDR algorithm in improving the stability and quality of UAV communications in complex environments, highlighting the application prospects of deep learning technology in the field of radio map reconstruction. With further development, RMC-DDR is expected to play an increasingly important role in optimizing UAV communication systems.

Deep Reinforcement Learning for Map-Assisted UAV Path Navigation

On the basis of the previous work, this chapter further considers UAV path navigation assisted by the reconstructed radio map. How to integrate the offline radio map with the online path navigation of UAVs and improve the path navigation performance is a problem worth studying. To address this need, an algorithm based on deep reinforcement learning is proposed to train the UAV path and learn the optimal flight path. The details are introduced in the following sections.

System Model

As shown in the system model diagram, the system includes a UAV and multiple ground cellular base stations. The UAV flies in the target airspace, and the base stations provide communication services for the UAV. Assume the UAV flight area is a cuboid, expressed as \([x_1, x_2] \times [y_1, y_2] \times [z_1, z_2]\), where 1 and 2 represent the lower and upper boundaries of the area, respectively. The UAV’s task is to fly from a random initial position to a fixed final position based on the radio map.

The position of the UAV at time \(t\) is expressed as \(\mathbf{l}(t)\), for \(0 \le t \le T\), with \(\mathbf{l}_I\) and \(\mathbf{l}_F\) representing the initial and final positions, respectively. Thus, we have \(\mathbf{l}(0) = \mathbf{l}_I\) and \(\mathbf{l}(T) = \mathbf{l}_F\). Assume there are \(C\) cellular base stations in the target area. Let \(h_c(t)\), \(1 \le c \le C\), denote the equivalent channel gain from base station \(c\) to the UAV at time \(t\). Therefore, the signal power received by the UAV from base station \(c\) at time \(t\) is expressed as:

\[
p_c(t) = P_c |h_c(t)|^2 = P_c \beta_c(\mathbf{l}(t)) G_c(\mathbf{l}(t)) |\tilde{h}_c(t)|^2, \quad c = 1, \ldots, C
\]

where \(P_c\) is the fixed transmit power of base station \(c\), \(\beta_c(\cdot)\) and \(G_c(\cdot)\) represent the large-scale channel gain and antenna gain of base station \(c\), respectively, and the random variable \(\tilde{h}_c(t)\) represents small-scale fading. Use \(c'(t) \in \{1, \ldots, C\}\) to denote the cellular base station connected to the UAV at time \(t\). When the received Signal-to-Interference Ratio (SIR) of the UAV is less than a threshold \(\gamma_{\mathrm{th}}\), i.e., \(\mathrm{SIR}(t) < \gamma_{\mathrm{th}}\), the connection between the UAV and the cellular network is judged to be in an outage state. The received SIR of the UAV at time \(t\) is expressed as:

\[
\mathrm{SIR}(t) = \frac{p_{c'(t)}(t)}{\sum_{c \neq c'(t)} p_c(t)}
\]

Due to the randomness of small-scale fading, at time \(t\), for any UAV position and the associated cell, the received signal-to-noise ratio is a random number. Therefore, the outage probability is a function of \(\mathbf{l}(t)\) and \(c'(t)\), expressed as:

\[
P_{\mathrm{out}}(\mathbf{l}(t), c'(t)) = \Pr\left[ \mathrm{SIR}(t) < \gamma_{\mathrm{th}} \right]
\]

According to the outage probability of the UAV, the outage time during task execution can be obtained as:

\[
T_{\mathrm{out}} = \int_0^T P_{\mathrm{out}}(\mathbf{l}(t), c'(t)) dt
\]

Let the total time cost of the UAV be the weighted sum of task completion time and outage time, expressed as:

\[
T_s = \alpha T + \beta T_{\mathrm{out}}
\]

where \(\alpha\) and \(\beta\) represent the weights for the total task completion time and the outage time during the total task completion time, respectively. Since the UAV is required to maintain good communication quality with the base station during flight, \(\beta\) is defined as a large constant to ensure a stable communication connection.

Problem Formulation

The energy consumption of UAV during mission execution usually includes flight propulsion energy and communication energy. Since the communication energy of UAV is much smaller than the propulsion energy, only the propulsion energy of the UAV is considered. The propulsion energy of a fixed-wing UAV can be expressed as:

\[
E_f \approx \int_0^T c_1 \left( 1 + \frac{\|\mathbf{a}(t)\|^2}{g^2} \right) \|\mathbf{v}(t)\|^2 + c_2 \|\mathbf{v}(t)\|^3 dt
\]

where \(c_1\) and \(c_2\) are fixed parameters related to air density, UAV weight, and wing area; \(\mathbf{v}(t)\) and \(\mathbf{a}(t)\) represent the velocity and acceleration of the UAV at time \(t\), respectively; \(g = 9.8 \, \mathrm{m/s^2}\) is the gravitational acceleration. Therefore, the flight energy consumption of the UAV depends on its velocity and acceleration. In this thesis, the UAV is assumed to fly at a constant speed with zero acceleration. Thus, the propulsion power is:

\[
P_u = c_1 \|\mathbf{v}(t)\| + c_2 \|\mathbf{v}(t)\|^3
\]

According to the above formula, the propulsion energy can be further expressed as:

\[
E_f = T_s \left( c_1 \|\mathbf{v}\| + c_2 \|\mathbf{v}\|^3 \right)
\]

To obtain the optimal flight path of the UAV, the flight energy consumption during mission execution is minimized under the constraint that the UAV-GBS connection quality is good. Since the main concern of this thesis is the energy minimization problem in UAV online trajectory planning, only the optimization problem of minimizing flight energy is solved. The optimization problem can be expressed as:

\[
\min_{\{\mathbf{l}(t), c'(t)\}} E_f
\]

\[
\text{s.t.} \quad \mathbf{l}(0) = \mathbf{l}_I, \quad \mathbf{l}(T) = \mathbf{l}_F
\]

\[
\mathbf{l}(t) \in \mathcal{L}, \quad \forall t \in [0, T]
\]

\[
c'(t) \in \{1, \ldots, C\}
\]

Map-Assisted UAV Path Navigation Algorithm Based on DRL

In this section, the D3QN algorithm is used to perform path navigation for UAVs with the assistance of the reconstructed map. The UAV learns from feedback by trying different actions, then reinforces actions until the best feedback is produced. The proposed map-reconstruction-based deep reinforcement learning path navigation method is illustrated in the flow chart.

In the scenario considered in this thesis, the UAV path navigation problem can be represented as a Markov Decision Process (MDP). Since MDP is defined in discrete time steps, the total flight time \(T\) is first discretized. Specifically, \(T\) is divided into \(M\) time steps, with interval \(\Delta t\), i.e., \(T = M \cdot \Delta t\). Here, \(\Delta t\) should be small enough so that the distance between the UAV and any base station in the considered area remains roughly unchanged during each movement, and the large-scale channel gain and antenna gain remain approximately constant. After discretization, the trajectory of the UAV is represented as \(\{\mathbf{l}(m)\}_{m=1}^M\).

Based on the above analysis, the total time cost of the UAV is converted to:

\[
T_s \approx \sum_{m=1}^M \alpha \Delta t + \beta P_{\mathrm{out}}(\mathbf{l}(m), c'(m)) \Delta t
\]

From the optimization problem, the flight energy of the UAV is related to the total flight time \(T\). Therefore, the optimization problem of minimizing the flight energy of the UAV can be transformed into minimizing the total task completion time under the constraint of satisfying the connection quality between the UAV and the cellular network. The UAV optimization problem can be transformed into:

\[
\max_{\{\mathbf{l}(m), c'(m)\}} \sum_{m=1}^M -\alpha – \beta P_{\mathrm{out}}(\mathbf{l}(m), c'(m))
\]

\[
\text{s.t.} \quad \mathbf{l}_{m+1} = \mathbf{l}_m + \Delta s \cdot \mathbf{v}_m
\]

\[
\mathbf{l}_0 = \mathbf{l}_I, \quad \mathbf{l}_M = \mathbf{l}_F
\]

\[
\mathbf{l}_m \in \mathcal{L}, \quad m = 1, 2, \ldots
\]

\[
c'(m) \in C
\]

Since \(\Delta t\) is a very small constant, it can be ignored. Here, \(\Delta s\) represents the moving distance of the UAV per time step, and \(\mathbf{v}_m\) represents the moving direction of the UAV per time step.

In the experiment, the goal is to maximize the feedback \(R\) in the MDP, which approximately equals the energy minimization problem under the connection quality constraint. An MDP is represented by a four-tuple variable: state \(S\), action \(A\), state transition probability \(P\), and feedback \(R\). The state space \(\mathcal{S} = \{\mathbf{l}: \mathbf{l} \in \mathcal{L}\}\) contains all possible positions of the UAV in the given flight area; the action space \(A\) contains the flight directions of the UAV; the state transition probability \(P\) is determined according to the current state and the subsequent flight direction; the feedback function \(R\) is defined as:

\[
R(\mathbf{l}) = -\mu – P_{\mathrm{out}}(\mathbf{l})
\]

where \(\mu\) is the penalty generated when the UAV stops, set to a large constant. The specific steps of the algorithm are shown in Algorithm 4.1.

Algorithm 4.1: D3QN-based UAV Online Path Navigation
1: Obtain the reconstructed radio map based on the untrained-network-prior map reconstruction algorithm
2: Initialize the current Q-network parameters \(\theta\). Initialize the target Q’ network parameters \(\theta’\), and assign Q network parameters to Q’ network. Initialize total iteration times epT, maximum steps per episode stepT, minimum interval to reach destination D, discount factor \(\gamma\), learning rate \(\varepsilon\), target Q network parameter update frequency f, random sample number m per episode
3: Initialize reward to reach destination R=0, out-of-bound penalty Pob, outage penalty weight \(\mu\), buffer D
4: Loop begins (iterating epT times)
5: Initialize the environment, obtain the current state S, done = False
6: According to the obtained state S, input the current Q network, calculate the Q value corresponding to each action, and use the greedy algorithm to select the corresponding action A in the current state S
7: Execute action A to obtain the new state S’. Obtain the outage probability Pout from the reconstructed radio map, reward R = -\(\mu\) – Pout. If the episode ends, done = True
8: If done == True, end the current loop
9: Store {S, S’, A, R, done} in D
10: Randomly sample m samples from D, {S_j, S’_j, A_j, R_j, done_j, j=1,2,…,m}, calculate the target y_j of the Q network
11: Use the mean squared loss function \(\frac{1}{m} \sum_{j=1}^{m} (y_j – Q(S_j, A_j; \theta))^2\) to calculate the loss, and back-propagate to update the parameter \(\theta\)
12: If epT % f == 0, update \(\theta’ = \theta\)
13: Set S = S’
14: Loop ends

Different from traditional methods, in Algorithm 4.1, the UAV does not need to interact directly with the environment. Instead, before the UAV performs the task, a radio map that is highly consistent with the actual environment is reconstructed. In reinforcement learning, the agent directly extracts data from the radio map to obtain the experienced outage probability, thereby obtaining the feedback value, and uses training data to adjust the UAV flight path.

Since the state space and action space in this problem are continuous, this thesis keeps the state space continuous while discretizing the action space \(A\) into four flight directions, i.e., \(A = \{v_1, v_2, v_3, v_4\}\). The discretization of the action space makes the state input of the action value function continuous and the action output discrete. The D3QN network architecture is adopted. In each step of each episode, the state of the UAV, i.e., its current position, is set as the input of the neural network, and the output is the flight direction of the UAV. Finally, based on the trained neural network, the UAV can select the best flight direction at any position based on the radio map, thereby completing the path navigation.

Simulation Design and Performance Analysis

In this section, simulation experiments are carried out to evaluate the proposed algorithm. First, the location and height distribution of buildings and the layout of base stations are generated using the statistical model proposed by the International Telecommunication Union (ITU). It is assumed that the proportion of land covered by buildings to the total land area is \(\alpha_{bd} = 0.3\); the average number of buildings per unit area is \(\beta_{bd} = 300\); the parameter of building height distribution is \(\gamma_{bd} = 50\) m, and the building height does not exceed 90 m. The connectivity weight with the base station is set to a large value to ensure a good communication connection with the ground.

Second, the path loss between the UAV and each base station is calculated. Considering the occlusion of buildings and the existence of Line-of-Sight (LoS) links, the signal transmission strength between the UAV and the base station at different positions is simulated. This step uses the 3GPP Urban Macro (UMa) model to generate the path loss from the base station to the UAV, combined with small-scale fading effects, to accurately simulate the channel environment.

Next, the received signal strength of the UAV from each base station is simulated. By judging whether a LoS link exists at a given UAV position, the connectivity between the UAV and the base station can be determined. Then, the Rayleigh fading model is used to simulate signal transmission in Non-Line-of-Sight (NLoS) situations, and the signal attenuation effect in LoS situations is considered. In the specific experiments, it is assumed that Rayleigh fading exists in the NLoS case and 15 dB Rayleigh fading exists in the LoS case.

The experiment considers a 2 km × 2 km area containing high-rise buildings. It is assumed that 2 GBSs are deployed in the area, with an antenna height of 25 m. In addition, the connectivity weight with the BS is set to a fixed value that is large enough to ensure a good communication connection with the ground. Finally, it is assumed that each dimension of the area has 201 data points, i.e., m = n = 201. Thus, the total number of data points is 201 × 201, covering different positions in each dimension of the area. These data points are used to evaluate and compare the performance of the proposed algorithm in different scenarios.

The experiments first compare the flight path of the UAV obtained by the D3QN algorithm under the reconstructed radio map and the direct flight path in the actual environment. The results show that the path trends in the two training cases are consistent. When selecting the path, the UAV tends to pass through the area with large communication coverage to reach the destination, and almost no steps are wasted. Therefore, the proposed map reconstruction method enables the UAV to learn a near-optimal path while maintaining a good connection with the cellular base station and reducing flight energy consumption.

The total outage time of each path for the UAV to reach the destination under different sampling factors and reconstruction methods is compared. The outage time is closely related to the reconstructed radio map. The outage time increases as the sampling factor increases. At the same sampling factor, the outage time of the UAV trajectory trained based on the RMC-DIP reconstructed map is shorter than that of the Kriging algorithm. Therefore, the radio map reconstructed by the proposed method can more accurately reflect the real wireless environment, allowing the UAV to maintain a better communication connection with the base station during flight.

The probability of successful arrival at the destination under different training episode numbers is shown. In the specific experiment, the success probability is calculated as the ratio of successful episodes to the total training episodes within the training set, and then accumulated. The results show that the highest success rate is achieved when the RMC-DIP algorithm has an undersampling factor of 2. When the undersampling rate is lower, the success rate decreases. The results also indicate that after about 3000 episodes of training, the UAV can fly to the destination under the constraints.

The moving average reward of the trajectory quality is also compared. When the undersampling factor is 2, the average reward of the proposed RMC-DIP algorithm is the closest to the actual trajectory, followed by a factor of 4, and the smallest is a factor of 16. When the undersampling factor is 4, the moving average reward of the spherical variogram model is only equivalent to that of the RMC-DIP algorithm with a factor of 8 to 16. This result clearly proves the effectiveness of map learning based on RMC-DIP, and the effect of trajectory learning is closely related to the quality of the reconstructed map.

Finally, the energy consumption of the UAV successfully reaching the destination during DRL training is examined. The comparison shows that the energy consumption trends under different sampling rates are the same. The maximum energy of the UAV under the cellular connection quality constraint is only about 2000 higher than that of the direct flight, and the average energy consumption is about 7000, which is a better trajectory flight state under good cellular connection conditions. Compared with direct flight, map-based UAV navigation under the constraint of good cellular network connection quality has stronger practical relevance.

Conclusion and Future Work

This thesis focuses on intelligent UAV path navigation technology based on radio map reconstruction, integrating radio map reconstruction with deep reinforcement learning algorithms. The main contributions and findings are summarized as follows:

First, I proposed two map reconstruction techniques based on untrained network priors. By combining sparse sampling and outage probability calculation, the UAV performs sparse sampling in the real environment, calculates the outage probability at sampling positions, and obtains an outage probability distribution map. Two downsampling image map completion algorithms were implemented: the RMC-DIP algorithm, which uses deep neural network structure and randomly initialized parameters as a strong prior for image content, enabling autonomous reconstruction without external labeled data; and the RMC-DDR algorithm, which incorporates deep denoising regularization to enhance image priors, improving robustness especially under sparse and noisy data conditions. Simulation results showed that the proposed algorithms achieve the highest PSNR and SSIM values and the smallest NMSE at the same sampling rate compared with traditional map reconstruction algorithms, allowing more accurate restoration of the outage probability distribution in actual environments while effectively avoiding dependence on large-scale labeled data.

Second, I implemented intelligent UAV path navigation through deep reinforcement learning. By designing appropriate reward and penalty mechanisms and using the Markov Decision Process for training, the D3QN algorithm was adopted to obtain the optimal path navigation scheme considering the outage probability. The optimization objective includes not only the total operation time of the UAV but also the impact of communication interruption time. Through weighted sum minimization, the UAV maintains stable communication connection with the ground cellular network while performing tasks, effectively reducing the probability of outage locations. Simulation results showed that the flight path of the UAV assisted by the radio map is almost consistent with the path directly flown in the actual environment, but the assistance of the radio map can effectively avoid uncertainties and risks in physical flight tests, reduce the need for trial and error in real environments, significantly improve work efficiency, and reduce task complexity.

Future research directions include:

1. Improving the map reconstruction algorithm by exploring more advanced machine learning and data processing methods, such as generative adversarial networks (GANs), to further enhance the accuracy and accuracy of radio maps; strengthening the handling of errors and uncertainties during reconstruction to improve the effect and stability.

2. Optimizing the path navigation algorithm, especially in dynamic environments and obstacle avoidance, to improve the efficiency and performance of UAV path navigation; considering more complex factors such as weather changes and wind speed to further enhance applicability and robustness.

3. Supplementing the experimental verification of the RMC-DDR algorithm. In the UAV path navigation simulation in Chapter 4, I mainly focused on the RMC-DIP map reconstruction technology. Since the current research stage aims to verify the basic feasibility of map reconstruction for UAV path navigation, RMC-DIP was chosen due to its relatively simple structure and mature implementation. In the future, I plan to conduct detailed verification of the RMC-DDR algorithm to evaluate its performance in radio-map-assisted UAV path navigation and explore its optimization potential and application scope.

In summary, the proposed map-reconstruction-based cellular-connected UAV online path navigation method has important theoretical and application value in the UAV field, providing new ideas and methods for the development and application of UAV technology. Future work will continue to deeply study and optimize the proposed methods to promote the advancement and application of UAV technology.

Scroll to Top