Visual Perception-Based Anti-UAV Technology: Development Dynamics and Trends

In the context of the continuous advancement of deep learning, visual perception-based object detection technology has achieved remarkable progress. Leveraging computer vision and image processing techniques to identify and track UAV targets, and subsequently predict their behavioral trends, can significantly enhance the accuracy and speed of UAV target detection. This holds substantial research significance and application value for ensuring national defense security. This article first elaborates on various categories of anti-UAV technologies domestically and internationally. Subsequently, it introduces related technologies and core algorithms from both traditional object detection methods and deep learning-based object detection methods. Finally, it analyzes future development trends and existing issues in visual perception-based anti-UAV technology, and prospects the future development of the anti-UAV field.

The current anti-UAV technology system can primarily be categorized into detection, tracking, early warning, damage, interference, and camouflage deception technologies. Detection and early warning technologies form the foundation and key for the latter three. As countries implement strict confidentiality measures on anti-UAV technologies, publicly accessible technical materials are limited. However, publicly available research materials on using UAVs as detection targets and employing visual perception techniques for detection and identification are relatively abundant. Visual perception technology based on deep learning refers to the technology that, targeting optical sensor information, draws on biological visual perception mechanisms, integrates computer vision processing methods and artificial intelligence algorithms, and uses computers or embedded devices to detect, understand, and predict graphical and video data. By analyzing UAV characteristics, shape, motion trajectories, and other information through object detection algorithms, it has become one of the primary methods for researching UAV targets in academia. Such methods enable real-time monitoring and effective identification of UAVs, providing target information for various anti-UAV systems and equipment.

Although deep learning-based visual perception technology holds great potential in the field of anti-UAV target detection, UAV technology itself continues to evolve, exhibiting trends such as miniaturization and concealment, thereby continuously raising requirements for visual perception-based anti-UAV technology. Deep learning object detection methods perform excellently in anti-UAV target detection tasks but also exhibit some notable drawbacks: they require large amounts of labeled data for training; they use end-to-end black-box models, making it difficult to interpret their decision-making processes; and they are prone to poor generalization performance.

1. Visual Perception-Based Anti-UAV Technology

1.1 Classification of Anti-UAV Technologies

When implementing anti-UAV operations, the primary task is to detect, track, and warn against UAVs. This process forms the basis of the entire anti-UAV system, utilizing monitoring equipment such as radar, optical sensors, and infrared sensors to ensure accurate awareness of UAV presence and activities. Once a UAV is detected, based on actual battlefield conditions, hard kill or soft kill measures can be chosen to counter the threat: implementing hard kill requires accurate target localization, conducting precise fire strikes by understanding the UAV’s position and trajectory to eliminate the threat; soft kill focuses on interference and paralysis, requiring knowledge of the UAV’s navigation and communication systems to selectively interfere with its control and communication signals, causing the UAV to lose navigation and control capabilities.

Furthermore, in anti-UAV operations, adopting proactive camouflage and protective methods is crucial. Currently, anti-UAV technologies developed by various countries include acoustic interference, signal interference, cyber attacks, laser interference, “anti-UAV” UAVs, and radio control seizure. Based on differences in interference techniques and suppression forms, these technologies can generally be classified into signal interference, damage interception, and network control categories. A comparison of operational effects of common anti-UAV means is shown in Table 1.

Anti-UAV Means Operational Effect Damage Effect Visibility Coverage Radius
Signal Interference Partial Disablement No Target Visibility Required ≤ 5 km
Damage Interception Target Destruction No Target Visibility Required Related to Weapon Range
Network Control Partial Disablement Network Connection Required ≤ 3 km

1.1.1 Signal Interference Anti-UAV Systems

Signal interference anti-UAV systems involve technologies such as electro-optical countermeasures, control information interference, and data link interference. These technologies can effectively interfere with UAV autopilot and control systems, communication systems, power systems, etc., reducing or disabling their primary operational functions. A typical application is emitting directional high-power interference radio frequency to cut off communication between the UAV and the remote controller, forcing the UAV to land or return autonomously. Another application direction is interfering with the UAV’s GPS signal receiver, causing it to rely solely on gyroscope-based inertial navigation systems, thereby losing precise navigation capability. Additionally, acoustic technology can interfere with UAV flight by causing gyroscope resonance and outputting erroneous information.

Electromagnetic interference anti-UAV systems employ electromagnetic pulses, high-power microwaves, etc., to burn out unprotected UAV electronic components or temporarily disable them, leading to UAV paralysis or crash. Although these interference technologies are relatively easy to implement and cost-effective, as UAV anti-interference capabilities improve, ordinary interference may fail to halt UAV actions, and the system cannot accurately predict the UAV’s next moves after interference. A device introduced by the United States, named “DroneDefender,” can quickly prevent UAVs from approaching. Users simply aim at the UAV and pull the trigger to shoot it down. However, this device is only effective against real-time remote-controlled UAVs or those relying on GPS navigation, with a strike range of about 400 meters. Russia’s “Banshee” UAV suppression system, introduced in 2023, features a lightweight design capable of interfering with UAVs and causing loss of control. Additionally, the system can deceive UAVs, guiding them to land at unknown locations, then reprogramming them to join friendly forces.

1.1.2 Damage Interception Anti-UAV Systems

Missiles, anti-aircraft guns, laser weapons, and other damage interception anti-UAV systems employ physical means to destroy UAVs, exhibiting good interception capabilities against high-speed, long-range flying UAVs. Missile interception systems use radar to detect and track UAV flight paths. Once a UAV is detected, the missile’s fire control system calculates parameters required for missile launch, including launch time, missile speed, and flight trajectory. After launch, the missile receives updated information from the fire control system via communication links during flight to ensure precise tracking of the target. When the missile approaches the UAV, it can employ different methods to destroy the target. The Micro Kinetic Kill Interceptor developed by Lockheed Martin in the United States is a pocket-sized interceptor that uses semi-active guidance, relying on radar to capture incoming targets, with the fire control system guiding the missile toward the target until the seeker detects the target’s reflected echo.

Moreover, anti-UAV laser weapons have become a representative of currently popular technologies. Laser weapon interception systems use laser beams to interfere with and degrade UAV performance. Laser weapons interfere with normal UAV operation or destroy key components by irradiating the UAV’s electro-optical system, navigation system, or transmission links. They exhibit exceptional precision and flexibility in countering UAV threats. Although laser weapons are limited by laser beam power and range, they achieve destruction or damage by releasing photons or particles moving at or near the speed of light toward the target, making it difficult for UAVs to evade attacks. Therefore, laser weapons, with their high precision, low cost, and rapid response, are ideal for countering small, low-altitude flying UAVs.

Boeing in the United States developed a laser weapon named “High Energy Laser Mobile Demonstrator (HELMD).” This weapon has a power of 10 kW and can shoot down slow, low-altitude flying UAVs within seconds of detection. HELMD successfully shot down over 150 simulated enemy targets in demonstration tests, including UAVs and rockets. Moreover, operating this laser weapon costs only the electricity for onboard equipment and required diesel fuel.

China’s “Silent Hunter” low-altitude laser air defense system has two modes: one is mobile, installed on wheeled vehicles for flexible deployment as needed; the other is fixed, deployed on building rooftops or other open areas. “Silent Hunter” has emission power levels of 5, 10, 20, and 30 kW, with effective interception distances of 200 to 4,000 meters. Theoretically, “Silent Hunter” can intercept all UAVs with wingspans not exceeding 2 meters and speeds not exceeding 60 m/s.

1.1.3 Network Control Anti-UAV Systems

Network control anti-UAV systems employ intrusion into UAV communication systems, ground control stations, or data links to implement interference, tampering, or gaining control, among other operations. The technical requirements for such systems are highly complex, as they need to block transmission of control signals to the UAV without damaging the UAV itself, and also disguise UAV control commands. Implementing network control is difficult for fully autonomous flying UAVs. Whether interference or direct destruction, both easily cause UAV crashes and bring additional impacts. To avoid such situations during UAV interception, there is a desire to intercept the transmission codes used by UAVs, thereby controlling the UAV or even guiding it to return. The United Kingdom developed a new system named “Anti-UAV Defense System,” which can interfere with signals sent by UAV operators by transmitting radio signals to the UAV’s directional antenna. Once the UAV receives signals from this system, it “freezes,” unable to determine direction, thus staying in the air.

In modern warfare environments, effective countermeasures against UAVs are crucial. Effective countermeasures require anti-UAV systems to possess autonomy, capable of quickly and accurately identifying UAV targets and taking proactive measures to address potential threats without human intervention. The prerequisite for achieving active defense is that anti-UAV systems and equipment can “see and distinguish” target UAVs. To achieve this, visual perception technology has become a core detection means for counter-UAV systems and equipment. Visual perception technology plays a key role in modern counter-UAV systems and equipment, with its high precision, real-time capability, non-invasiveness, and wide applicability making it an ideal choice for autonomous identification and response to UAV threats. Through images and videos provided by optical detection equipment, along with advanced object detection algorithms, visual perception technology is expected to provide more reliable and efficient UAV countermeasure solutions.

1.2 Traditional Anti-UAV Target Detection Methods

Traditional object detection methods mainly include steps such as data acquisition, data preprocessing, feature extraction, feature representation, object recognition, bounding box generation, and post-processing. The data acquisition stage requires obtaining images or videos containing UAVs, then performing preprocessing operations such as scaling, denoising, adjusting brightness and contrast on the images to ensure input data quality and consistency. The feature extraction stage extracts discriminative information from images through filters and image processing techniques, such as edges, textures, and colors, providing basis for subsequent object recognition. The object recognition stage uses these features to compare with predefined templates to determine if targets exist in the image and their positions. Once a target is recognized, the system generates a bounding box to mark the UAV’s position information and possible identification information. The framework of traditional object detection algorithms is shown in Figure 1.

1.2.1 SIFT Algorithm

Scale-Invariant Feature Transform (SIFT) is an image local feature extraction algorithm. It locates and determines principal directions by finding extreme points, i.e., feature points or keypoints, in different scale spaces, and constructs keypoint descriptors to extract features. The SIFT algorithm possesses scale invariance and rotation invariance, and is not interfered by factors such as illumination, affine transformation, and noise. It can find salient points in images that are less affected by illumination, affine transformation, and noise, such as corner points, edge points, dark area bright points, and bright area dark points. SIFT algorithm’s keypoint detection and description provide a reliable foundation for various visual perception tasks. Although the algorithm has scale, rotation, and illumination invariance, challenges remain in real-time performance and extracting feature points for smooth-edged targets.

The SIFT algorithm involves several steps, including scale-space extrema detection, keypoint localization, orientation assignment, and keypoint descriptor generation. The scale-space representation is constructed using Gaussian blur and difference-of-Gaussian (DoG) functions. The keypoint detection can be expressed as finding local extrema in the DoG function:

$$ D(x, y, \sigma) = (G(x, y, k\sigma) – G(x, y, \sigma)) * I(x, y) $$

where \( G(x, y, \sigma) \) is the Gaussian kernel, \( I(x, y) \) is the image, and \( k \) is a constant multiplicative factor. Keypoints are selected where \( D(x, y, \sigma) \) is a local extremum in both scale and space.

1.2.2 HOG Detection Algorithm

The Histogram of Oriented Gradients (HOG) detection algorithm’s main idea is to capture target texture and shape information by analyzing gradient directions in images, thereby achieving object detection. In the HOG detection algorithm, feature extraction, as a key step, includes operations such as gradient computation, cell division, block formation, gradient histogram normalization, and feature vector concatenation. Even without knowing corresponding gradient or edge positions, this method can well represent local object appearance and shape by dividing the image window into small spatial cells, with each cell accumulating a one-dimensional histogram of local gradient directions or edge directions over pixels.

The HOG feature calculation involves computing gradient magnitudes and orientations for each pixel:

$$ G_x(x, y) = I(x+1, y) – I(x-1, y) $$

$$ G_y(x, y) = I(x, y+1) – I(x, y-1) $$

$$ G(x, y) = \sqrt{G_x(x, y)^2 + G_y(x, y)^2} $$

$$ \theta(x, y) = \arctan\left(\frac{G_y(x, y)}{G_x(x, y)}\right) $$

These gradients are then accumulated into histograms within cells, and blocks are normalized to form the final feature vector.

1.2.3 DPM Algorithm

The Deformable Part Model (DPM) algorithm’s core idea is to represent target objects as models composed of multiple parts, and achieve object detection by learning relative positions and shapes between these parts. The DPM algorithm can effectively handle target diversity and variation, performing excellently in part-specific tasks. By decomposing targets into multi-part models and learning their position and shape relationships, it achieves adaptability to targets with different poses, scales, and shapes, possessing certain robustness. The DPM algorithm is computationally complex, especially in multi-model combination and sliding window search, and its application performance is inferior to deep learning methods for projects with high real-time requirements. Additionally, the DPM algorithm performs poorly when dealing with partially occluded or low-visibility targets; for target categories with significant natural variations, such as animals, large-scale training data and complex models are required to achieve satisfactory recognition results.

The DPM model can be formulated as a scoring function that combines root and part filters:

$$ score(p_0, \ldots, p_n) = \sum_{i=0}^n F_i \cdot \phi(H, p_i) – \sum_{i=1}^n d_i \cdot \psi(p_i – p_0) + b $$

where \( p_0 \) is the root location, \( p_i \) are part locations, \( F_i \) are filters, \( \phi(H, p_i) \) are feature vectors, \( d_i \) are deformation costs, \( \psi \) is a deformation feature, and \( b \) is a bias term.

1.2.4 Selective Search Algorithm

The selective search algorithm’s basic idea is to generate candidate object regions in images through multi-scale and multi-level region segmentation and merging, thereby reducing computational costs in subsequent object detection steps. The algorithm’s main steps include multi-scale segmentation, region merging, and candidate box generation. Selective search decomposes images into regions at different scales through multi-scale segmentation, merges similar regions into larger regions by measuring similarity between regions (such as color, texture, and edges), generates candidate regions that may contain objects, and finally converts these merged regions into candidate rectangular bounding boxes. The pixel count within rectangular bounding boxes is usually much smaller than the pixel count in the image, significantly reducing computational complexity in subsequent object detection.

Selective search can retain important target regions in images while reducing computational complexity, effectively reducing redundant candidate regions, and improving object detection efficiency and accuracy. Multi-scale and multi-level region segmentation and merging can effectively capture targets at different scales and possess robustness to complex scenes and occlusion situations. Additionally, selective search can be combined with deep learning methods to further enhance object detection accuracy. The algorithm’s drawback is that due to the limited number of generated candidate regions, small or sparse targets may be missed.

Although traditional object detection algorithms require less labeled data and computational resources, and can run on lower hardware requirements, they typically rely on manually designed feature extraction processes (such as HOG, SIFT, etc.) and use machine learning classifiers (such as support vector machines, decision trees, etc.) for object detection. These algorithms exhibit good performance in specific scenes and conditions, but their robustness in complex, diverse data and scenes is relatively limited. Deep learning object detection algorithms adopt a different approach, utilizing network structures such as Convolutional Neural Networks (CNN) to automatically learn feature representations from large amounts of labeled data, eliminating the tedious process of manual feature design.

A comparison of traditional object detection algorithms is summarized in Table 2.

Algorithm Key Features Advantages Disadvantages
SIFT Scale and rotation invariant, local feature extraction Robust to illumination and affine changes Computationally intensive, poor for real-time
HOG Gradient orientation histograms, texture and shape capture Effective for pedestrian detection, relatively fast Sensitive to occlusion, limited for small objects
DPM Deformable parts, multi-component models Handles deformation and occlusion well Complex training, slow inference
Selective Search Region proposal generation, hierarchical segmentation Reduces search space, good for objectness May miss small objects, not end-to-end

1.3 Deep Learning-Based Anti-UAV Target Detection Methods

Deep learning achieves data learning and pattern recognition by constructing and training neural networks. Compared to traditional object detection methods, deep learning possesses automatic feature learning capability, able to automatically learn feature representations from data, reducing dependence on manually designed features. Deep learning-based object detection algorithms are mainly divided into two categories. One is regression-based one-stage object detection algorithms, which do not require generating candidate boxes but directly treat target bounding box localization as a regression problem. This method can quickly detect target positions, improving detection speed. The other is candidate box-based two-stage object detection algorithms, which first generate a series of sample candidate boxes, then classify these candidate boxes through CNN. This method has higher detection accuracy and localization precision but relatively slower speed. The framework of deep learning object detection algorithms is shown in Figure 2.

1.3.1 One-Stage Detection Algorithms

In 2016, Redmon et al. proposed the YOLO (You Only Look Once) algorithm, which introduced the distinction between single-stage and two-stage for deep learning-based object detection algorithms, treating the entire detection process as an end-to-end network operation. Compared to traditional two-stage methods, the YOLO algorithm has significant advantages in speed, enabling real-time object detection, suitable for application scenarios with high real-time requirements. However, the YOLO algorithm also has some issues, including coarse target localization, low detection accuracy, and difficulty detecting small targets. YOLOv4 and YOLOv5 object detection algorithms, as improved versions of YOLO, exhibit higher detection accuracy and can achieve real-time or near-real-time object detection. In 2023, Ultralytics released YOLOv8, which adopts a lighter network structure and more efficient inference techniques while maintaining high accuracy.

The YOLO algorithm divides the input image into an \( S \times S \) grid. Each grid cell predicts \( B \) bounding boxes and confidence scores for those boxes, as well as \( C \) class probabilities. The confidence score reflects the probability that the box contains an object and the accuracy of the box. The output is a tensor of size \( S \times S \times (B \times 5 + C) \). The loss function combines localization, confidence, and classification errors:

$$ \lambda_{\text{coord}} \sum_{i=0}^{S^2} \sum_{j=0}^{B} \mathbb{1}_{ij}^{\text{obj}} \left[ (x_i – \hat{x}_i)^2 + (y_i – \hat{y}_i)^2 \right] + \lambda_{\text{coord}} \sum_{i=0}^{S^2} \sum_{j=0}^{B} \mathbb{1}_{ij}^{\text{obj}} \left[ (\sqrt{w_i} – \sqrt{\hat{w}_i})^2 + (\sqrt{h_i} – \sqrt{\hat{h}_i})^2 \right] $$

$$ + \sum_{i=0}^{S^2} \sum_{j=0}^{B} \mathbb{1}_{ij}^{\text{obj}} (C_i – \hat{C}_i)^2 + \lambda_{\text{noobj}} \sum_{i=0}^{S^2} \sum_{j=0}^{B} \mathbb{1}_{ij}^{\text{noobj}} (C_i – \hat{C}_i)^2 + \sum_{i=0}^{S^2} \mathbb{1}_{i}^{\text{obj}} \sum_{c \in \text{classes}} (p_i(c) – \hat{p}_i(c))^2 $$

The Single Shot MultiBox Detector (SSD) algorithm adopts a grid division approach similar to the YOLO algorithm, combining advantages of other algorithms, fully exploiting convolutional layer feature information, enabling it to ensure algorithm speed while meeting detection accuracy, to some extent overcoming YOLO’s difficulty in detecting small targets and inaccurate localization. SSD’s core idea is transforming object detection into regression and classification problems. The SSD algorithm can achieve real-time object detection, suitable for scenarios with high real-time requirements, but performs poorly in detecting small objects, prone to false detections, and requires extensive data augmentation and dataset filtering during training. The SSD algorithm has high hardware resource requirements, needing strong computing power and large storage space.

SSD uses a base network (e.g., VGG) and adds auxiliary convolutional layers to produce feature maps at multiple scales. Each feature map cell predicts a set of default boxes with different aspect ratios and scales. The predictions include offsets for localization and confidence scores for each class. The loss function is a weighted sum of localization loss (smooth L1) and confidence loss (softmax):

$$ L(x, c, l, g) = \frac{1}{N} (L_{\text{conf}}(x, c) + \alpha L_{\text{loc}}(x, l, g)) $$

where \( N \) is the number of matched default boxes, \( x \) is an indicator for matching, \( c \) is confidence, \( l \) is predicted box parameters, \( g \) is ground truth box parameters.

The Retina-NET algorithm utilizes a Feature Pyramid Network (FPN) structure, enabling object detection at different scales, allowing the model to have good perception ability for targets of different sizes. Retina-NET adds a subnetwork on each feature layer, which simultaneously performs object category classification and bounding box regression. Retina-NET introduces a new loss function—Focal Loss—to address class imbalance issues. Focal Loss adjusts class loss weights, overall equivalent to increasing the weight of inaccurately classified samples in the loss function, focusing attention on difficult-to-classify positive samples, thereby improving object detection performance. It should be noted that because small target features are difficult to extract and localize, Retina-NET may experience decreased detection performance when handling small targets.

Focal Loss is defined as:

$$ \text{FL}(p_t) = -\alpha_t (1 – p_t)^\gamma \log(p_t) $$

where \( p_t \) is the model’s estimated probability for the true class, \( \alpha_t \) is a balancing factor, and \( \gamma \) is a focusing parameter that reduces the loss for well-classified examples.

The CornerNet algorithm’s main idea is simplifying object detection to a pair of keypoints, i.e., the top-left and bottom-right corners of the target. The network introduces two branches, one responsible for predicting top-left corner heatmaps, the other for predicting bottom-right corner heatmaps. This design advantage greatly simplifies network output, no longer requiring complex anchor box design. Additionally, the network can predict embedding vectors for each detection point to corner points, used to match and group a pair of corner points belonging to the same target. CornerNet uses an hourglass network as the backbone network, where each module has its own corner pooling module, merging features in the hourglass network before predicting heatmaps, embeddings, and offsets. CornerNet has good robustness to target pose, occlusion, and scale changes, adapting to complex scenes, but due to lack of observation of global object information, it may detect many incorrect bounding boxes, especially when target intersection over union is small, this situation is more severe.

CornerNet predicts heatmaps for top-left and bottom-right corners, along with embeddings and offsets. The loss function includes terms for corner detection and grouping:

$$ L = L_{\text{det}} + \alpha L_{\text{pull}} + \beta L_{\text{push}} + \gamma L_{\text{off}} $$

where \( L_{\text{det}} \) is a focal loss for corner detection, \( L_{\text{pull}} \) and \( L_{\text{push}} \) are losses for grouping corners based on embeddings, and \( L_{\text{off}} \) is a smooth L1 loss for offset regression.

The DETR (Detection Transformers) algorithm is the first to apply Transformers to object detection. DETR’s core idea is transforming object detection tasks into set prediction problems, treating object detection as establishing a one-to-one matching relationship between input images and a set of predefined object sets. DETR’s decoder associates and matches elements in the object set through self-attention mechanisms, thereby predicting object detection results. DETR finds and matches targets through heuristic assignment rules similar to those in modern detectors, but unlike traditional methods, it does not require using anchor boxes, simplifying model design and training processes. However, the Transformer architecture is computationally complex, requiring considerable computational resources and time. Additionally, due to sensitivity to local details of small targets, DETR does not perform as well as some anchor-based methods in small target detection.

DETR uses a CNN backbone to extract features, then a Transformer encoder-decoder to produce a set of object predictions. The bipartite matching loss is used to assign predictions to ground truth:

$$ L_{\text{matching}} = \sum_{i=1}^N \left[ -\log \hat{p}_{\sigma(i)}(c_i) + \mathbb{1}_{\{c_i \neq \varnothing\}} L_{\text{box}}(b_i, \hat{b}_{\sigma(i)}) \right] $$

where \( \sigma \) is an optimal assignment computed with the Hungarian algorithm, \( \hat{p} \) are class probabilities, and \( L_{\text{box}} \) is a combination of L1 loss and generalized IoU loss for bounding boxes.

1.3.2 Two-Stage Detection Algorithms

The Region-based Convolutional Neural Network (R-CNN) object detection algorithm is a classic two-stage detection algorithm. R-CNN first extracts candidate boxes during detection, then uses CNN for feature extraction, and finally uses support vector machines to classify extracted features and perform box regression. Compared to traditional object detection algorithms, R-CNN introduces selective search to address the issue of excessive computation when generating candidate boxes with sliding windows. Additionally, R-CNN uses SVM classifiers to classify extracted features and employs regression algorithms to correct target bounding boxes, reducing error between regions of interest and actual target regions, improving detection accuracy. In complex backgrounds, R-CNN may struggle to generate effective candidate boxes, possibly leading to missed or false detections. R-CNN’s candidate regions undergo crop/warp for fixed size, cannot guarantee image non-deformation. Moreover, R-CNN has high time complexity in the feature extraction stage and may suffer from information loss.

Fast R-CNN is an improved version of the R-CNN series methods. Fast R-CNN improves object detection speed and accuracy by sharing convolutional feature extraction operations and reducing repetitive computations. Compared to older versions of R-CNN, Fast R-CNN has improvements in object detection speed.

Faster R-CNN is another improved version of the R-CNN series algorithms, achieving faster end-to-end object detection by introducing a Region Proposal Network (RPN). Faster R-CNN has higher accuracy compared to traditional object detection methods, and the network structure is more complex. The RPN structure is a subnetwork for generating candidate regions, its role is extracting candidate regions that may contain targets in images.

The R-CNN family involves region proposal, feature extraction, and classification. Faster R-CNN integrates RPN for proposal generation. The RPN uses anchors of various scales and aspect ratios at each sliding window location. The loss function for RPN includes classification (object vs. not object) and regression for anchor refinement:

$$ L(\{p_i\}, \{t_i\}) = \frac{1}{N_{\text{cls}}} \sum_i L_{\text{cls}}(p_i, p_i^*) + \lambda \frac{1}{N_{\text{reg}}} \sum_i p_i^* L_{\text{reg}}(t_i, t_i^*) $$

where \( p_i \) is predicted objectness score, \( p_i^* \) is ground truth label (1 for object, 0 for background), \( t_i \) is predicted bounding box parameters, \( t_i^* \) is ground truth parameters.

The team led by Kaiming He improved the R-CNN model and proposed the Spatial Pyramid Pooling Network (SPPNet) model, mainly addressing the issue of repeated feature extraction on images by R-CNN. The spatial pyramid pooling network connects a pyramid pooling module before fully connected layers to adapt to any size image input, solving the problem of information loss caused by normalization in R-CNN. Additionally, SPPNet connects a pyramid pooling layer after the last convolutional layer, allowing the network to input any image and generate fixed-size output. Compared to traditional object detection methods like R-CNN, SPPNet only requires one convolutional operation in the feature extraction stage, then extracts fixed-length feature vectors through the spatial pyramid pooling layer, reducing repetitive computations and lowering computational load. However, SPPNet’s spatial pyramid pooling layer divides feature maps into fixed-size sub-regions and generates fixed-length feature vectors, which may lead to poor detection performance for dense targets.

The Feature Pyramid Network (FPN) is a top-down feature fusion method, its principle is providing more comprehensive semantic and contextual information by constructing feature pyramids and fusing multi-scale feature information at different network levels. FPN differs from other algorithms in that it does not only use one feature prediction layer. Although some algorithms also use multi-scale feature fusion for object detection, they typically only utilize one scale of fused features, possibly introducing errors affecting detection accuracy. The FPN algorithm addresses this issue by allowing object prediction on multiple different scales of fused features to maximize detection accuracy. This means FPN can simultaneously utilize feature information at multiple scales, thereby improving object detection accuracy.

FPN constructs a feature pyramid with lateral connections from deeper to shallower layers. Each level of the pyramid can be used for object detection. The feature at level \( l \) is computed as:

$$ P_l = \text{Upsample}(P_{l+1}) + \text{Conv}(C_l) $$

where \( C_l \) is the feature map from the backbone at level \( l \), and Upsample is typically bilinear interpolation.

Deep learning object detection algorithms possess the ability to automatically learn and extract features from images. The uniqueness of these algorithms lies in their ability to extract multi-scale feature information from input images, thereby capturing target features at different scales and shapes, improving detection robustness, reducing computational complexity, significantly increasing detection speed, and achieving high-precision object detection and identification.

A summary of deep learning-based object detection algorithms is presented in Table 3.

Algorithm Category Key Features Advantages Disadvantages
YOLO series One-stage Unified detection, grid-based Very fast, real-time capable Lower accuracy for small objects
SSD One-stage Multi-scale feature maps, default boxes Good speed-accuracy trade-off Struggles with very small objects
Retina-NET One-stage FPN backbone, focal loss Handles class imbalance well Computationally heavy
CornerNet One-stage Corner keypoints, hourglass network Anchor-free, good for occlusion High false positives for low IoU
DETR One-stage Transformer based, set prediction End-to-end, no anchors needed Slow training, poor for small objects
R-CNN Two-stage Selective search, SVM classification High accuracy with good features Very slow, multi-stage pipeline
Fast R-CNN Two-stage ROI pooling, shared features Faster than R-CNN, single-stage training Depends on external proposals
Faster R-CNN Two-stage RPN for proposals, end-to-end High accuracy, integrated proposal Slower than one-stage methods
SPPNet Two-stage Spatial pyramid pooling, multi-scale Handles variable input sizes Complex pipeline, not end-to-end
FPN Two-stage/One-stage Feature pyramid, top-down pathway Excellent for multi-scale objects Increased memory and computation

2. Development Trends of Visual Perception-Based Anti-UAV Technology

The multi-category object recognition capability of deep learning enables anti-UAV systems to distinguish between different models and appearances of UAVs, which is crucial for dealing with diverse threats. Deep learning also enables anti-UAV systems to operate efficiently under real-time requirements; hardware acceleration and model optimization allow systems to quickly respond to UAV threats and take necessary countermeasures promptly. This not only helps achieve continuous tracking of targets, better understanding of UAV dynamic behavior, but also enhances system intelligence.

2.1 Analysis of Development Trends

Visual perception-based anti-UAV technology is a key area in the anti-UAV system, with many researchers contributing from different aspects, greatly promoting the development of anti-UAV technology. Although significant progress has been made, there are still some unresolved research issues in this field, and new challenges will continue to emerge with societal development. The following is an analysis of future development trends of this technology.

2.1.1 Multi-Modal Perception and Integration

In the future, anti-UAV technology will increasingly rely on multiple sensors, including vision, infrared, radar, sound, etc., to improve perception capability of UAVs. This multi-sensor application will help achieve all-weather, multi-angle perception, thereby assisting anti-UAV systems in more reliably detecting and identifying different types of UAVs, especially those with low-visibility stealth characteristics, and reducing false alarms.

Multi-modal data fusion technology will cover various data types such as images, videos, thermal infrared images, and sound signals. Systems can comprehensively understand target UAV characteristics and behavior by integrating these multi-modal data. Different sensors have complementarity under different conditions; for example, infrared sensors can detect UAV thermal radiation, radar can track UAV speed and position. Adaptive sensor selection will become a future trend: systems can automatically select the most suitable sensor combination based on current environmental conditions and UAV characteristics, thereby reducing system energy consumption and improving performance. With the development of deep learning and machine learning, data fusion algorithms will continue to evolve; these algorithms can integrate information from different sensors and data sources, providing anti-UAV systems with stronger analytical capabilities.

2.1.2 Real-Time Automated Decision-Making

Future anti-UAV systems will become more automated, achieving real-time automated decision-making, able to identify threats in real-time and take necessary countermeasures, such as interfering with communications, implementing interception, etc. This will reduce operator burden and improve system response speed. Specifically, systems will use sensors to obtain UAV visual information, and through steps such as data processing, target recognition, and classification, interpret and analyze this information in real-time. Based on preset strategies and rules, decision systems will autonomously make corresponding decisions and take actions. Achieving this real-time automated decision-making process relies on fast computing and real-time feedback. Systems will continuously monitor and evaluate decision effectiveness, and provide feedback and iteration based on real-time situations. Future technological development will further improve the accuracy, response speed, and intelligence level of this aspect, to achieve more efficient and reliable anti-UAV systems.

2.1.3 Anti-Interference and Counter-Countermeasure Technology

As anti-UAV technology develops, UAV manufacturers will continuously improve their products to enhance anti-interference capability and countermeasure technology levels. Therefore, anti-UAV systems need to possess anti-interference and counter-countermeasure technology, continuously upgrading to adapt to threats from new UAVs. Frequency scanning and adaptive signal processing, by scanning UAV communication frequencies and employing adaptive signal processing algorithms, improve system anti-interference capability; enhanced artificial intelligence and machine learning, by learning and identifying UAV behavior patterns, improve detection and response capability; cooperative anti-UAV systems through cooperation and coordination among multiple systems jointly address complex threats, improving system effectiveness and robustness. The comprehensive application of these technologies can enhance anti-UAV system countermeasure capability.

2.1.4 UAV Behavior Prediction

Anti-UAV systems need to be able to timely predict and classify UAV behavior to take corresponding measures in advance; behavior prediction and classification algorithms are key to achieving this goal. Systems collect UAV characteristic data for training and learning, based on which predictive models can be established. These characteristic data include UAV flight patterns, trajectories, communication frequencies, etc. By analyzing patterns and correlations in characteristic data, predictive models can classify UAV behavior, determining whether these behaviors are normal or potential threat behaviors. Thus, anti-UAV systems can detect threats beforehand and take appropriate countermeasures to prevent potential threats from occurring.

2.2 Existing Issues

2.2.1 Target Detection and Tracking in Complex Environments

When conducting UAV target detection and tracking in complex environments, faced complex environments include atmospheric disturbance, illumination condition changes, background interference, multi-target tracking, and target changes, etc. Complex environments contain numerous occlusions, such as buildings, trees, or other objects, which may cause UAV targets to be partially or completely occluded, making detection and tracking difficult. Moreover, complex environments are often dynamic, with frequent changes in objects and backgrounds, such as insects, birds, or other moving objects, requiring timely updates of target positions and attributes, which also poses challenges for target detection and tracking.

To address these difficulties, researchers have attempted some solutions. Using deep learning methods to learn target features and contextual information from large amounts of data can achieve more accurate detection and tracking in complex environments. Additionally, when targets are occluded or tracking is lost, target re-identification technology can model and match target appearance features to re-identify targets, maintaining continuous tracking of targets. Furthermore, through motion prediction and model updates, analyzing target motion patterns and behavior, timely updating models and tracking algorithms can better track targets and cope with dynamic environmental changes.

2.2.2 Detection and Tracking of “Low, Slow, and Small” UAVs

Detection and tracking of “low, slow, and small” UAVs face multiple difficulties. Since such UAVs are typically small in size, coupled with limited target pixel count and easy blending with background, long-distance visual detection is difficult. “Low, slow, and small” UAVs possess fast and agile motion capabilities, causing rapid position and shape changes in image sequences, prone to blurring and position instability, making tracking tasks more complex. Additionally, low-altitude high-speed flight environments may have low signal-to-noise ratio issues, sensor signals affected by noise and interference, reducing target visibility in images or sensor data, making target edges unclear. It should be noted that visual occlusion is also a challenge: small UAVs may be occluded by obstacles or objects in the background, reducing target visibility.

Therefore, research on prediction algorithms is needed to estimate UAV future positions and trajectories. Using high-resolution sensors can obtain more detailed information to enhance target visibility. Through sensor data fusion and motion model prediction, reducing interference from motion can mitigate the impact of complex environments on target detection and tracking. Comprehensively utilizing information from multiple perspectives or sensors can reduce the impact of visual occlusion, improving target detection and tracking performance.

2.2.3 Detection and Tracking of Stealth UAVs

Stealth UAVs possess low detectability; their shape, materials, or coatings may be similar to the surrounding environment, causing difficulty in distinguishing them from background in sensor data. Stealth UAVs typically use low-noise engines or motors to reduce sound and thermal signal generation. This makes traditional sound or thermal imaging sensors ineffective in detecting UAVs, increasing difficulty of detection and tracking.

Multi-spectral image sensors, infrared imaging sensors, radar, etc., can be utilized to improve detection capability for stealth UAVs. Multi-spectral image sensors can capture subtle differences between stealth UAVs and surrounding environment; infrared imaging sensors can rely on UAV thermal radiation for detection; radar can use echo signals to identify UAV presence. These advanced sensors can provide diverse data sources, increasing chances of detecting stealth UAVs.

3. Conclusion

Visual perception-based anti-UAV technology has always been a research hotspot in the fields of national security and public safety. This article summarizes the current research status of target detection in the anti-UAV technology field and prospects future development trends. Target detection technology plays a key role in anti-UAV technology. Currently, researchers domestically and internationally have achieved significant results, developing various target detection algorithms and applying them to actual anti-UAV systems. These algorithms possess characteristics such as strong adversariality, strong real-time capability, and high fragmentation, able to accurately detect and identify UAV targets in complex environments.

However, target detection technology still faces some challenges. First, UAV appearance and morphology are diverse, including various sizes, shapes, and different flight attitudes, requiring target detection algorithms to have good adaptability and generalization capability. Second, UAV flight speed is fast, requiring algorithms with strong real-time capability to achieve rapid target detection and tracking. Finally, UAVs often move in complex backgrounds, such as urban environments, forest areas, etc., requiring target detection algorithms to have high robustness, able to effectively suppress background interference and false detections.

With the continuous development of computer vision and deep learning, anti-UAV technology levels will further improve. More advanced and complex deep learning models can be adopted, such as high-level algorithms like target tracking and behavior analysis in object detection, to enhance perception and understanding capabilities of UAVs. Combining multi-modal perception, such as fusion of vision, sound, radar, and other sensors, can provide more comprehensive and accurate UAV information, enhancing anti-UAV system performance.

In summary, visual perception-based anti-UAV technology holds vast potential in areas such as multimodal fusion, automated decision-making, and predictive analysis of UAV behavior. Related technologies will provide robust support to the field of anti-UAV technology, enhancing perception and response capabilities to UAV threats. Continuous research and innovation in this domain are essential to address evolving UAV challenges and ensure effective countermeasures for security and defense applications.

Scroll to Top