The advent of intelligent manufacturing and advanced information technologies has ushered in a new era of automation and smart production for agriculture. Within this transformative landscape, the concept of the unmanned smart farm emerges as a critical frontier, presenting both significant opportunities and complex technical challenges. A key enabling technology for this vision is the coordinated operation of unmanned aerial vehicles (UAVs). Drone formation flight control technology provides indispensable technical support for various agricultural production tasks, including field inspection, herd management, and precision irrigation. The ability to deploy multiple UAVs in a synchronized manner amplifies operational efficiency, coverage, and data collection capabilities beyond what a single unit can achieve.

However, the development and testing of drone formation algorithms using physical hardware are often prohibitively expensive, time-consuming, and risky. Common simulation platforms may lack the necessary interfaces or flexibility for deep-level algorithm development and integration with standard robotic software frameworks. Therefore, there is a pressing need for a high-fidelity, versatile, and open-source simulation environment that can accurately replicate field conditions for agricultural UAVs and facilitate the optimization of coordinated flight control algorithms. Such a platform must offer unified interfaces, particularly with the Robot Operating System (ROS), which has become the de facto standard for robotics research and development.
This study focuses on the construction and evaluation of a simulation environment for multi-rotor drone formation cooperative flight, based on the integration of ROS, the PX4 autopilot, and the Gazebo simulator. We analyze the data interaction logic between these core components, build and optimize a multi-UAV simulation framework, and implement key functionalities including formation flight control, simultaneous localization and mapping (SLAM), and autonomous navigation. The primary objective is to demonstrate the feasibility and effectiveness of using simulated drone formation swarms to meet common operational demands in an unmanned smart farm context, thereby providing a robust technical foundation and valuable insights for further research and practical application.
1. System Architecture and Interaction Logic
The core of the simulation environment is the seamless interaction between three main software ecosystems: ROS (for high-level control and communication), PX4 (for low-level flight dynamics and control simulation), and Gazebo (for 3D world and physics simulation). Understanding their data flow is crucial for building a functional drone formation simulation.
1.1 Interaction Between ROS and the PX4 Autopilot
In a Software-In-The-Loop (SITL) setup, the PX4 autopilot software runs as a separate process, simulating the flight controller’s firmware. Communication between the ROS ecosystem and PX4 is bridged by the MAVROS package. MAVROS acts as a ROS node, translating ROS messages (such as pose, velocity, or actuator commands) into the MAVLink protocol understood by PX4, and vice-versa. This bidirectional link allows high-level control algorithms running in ROS to send setpoints (e.g., target position $$p_t = [x_t, y_t, z_t]^T$$) to PX4, which then calculates the required motor outputs using its internal control loops (e.g., a PID controller to minimize the error $$e = p_t – p_c$$, where $$p_c$$ is the current position). PX4 subsequently publishes the vehicle’s actual state (attitude, global/local position, battery status) back to ROS via MAVROS topics (e.g., `/mavros/state`, `/mavros/local_position/pose`), closing the perception-control loop for the simulation.
1.2 Interaction Between ROS and Gazebo
Gazebo is responsible for simulating the physical world, including the UAV models, sensors, and environmental dynamics. ROS interacts with Gazebo through its dedicated plugin interface and service calls. The UAV model, defined in URDF/Xacro format, is spawned into the Gazebo world by a ROS launch file. Gazebo’s physics engine calculates the forces and torques based on PX4’s motor commands and updates the model’s state. This updated state (pose, twist) is then published to specific ROS topics (e.g., `/gazebo/model_states`). Furthermore, sensor plugins (for IMU, camera, laser rangefinder) simulate data based on the model’s state and publish corresponding ROS sensor messages (e.g., `sensor_msgs/LaserScan`). These messages become the perceptual input for autonomy algorithms like SLAM and navigation within the ROS network. This interaction ensures that the drone formation perceives and reacts to a dynamically simulated environment.
The overall data interaction for a single UAV in the formation can be summarized by the following information flow cycle:
- Control Command Flow: ROS Node → MAVROS (ROS) → MAVLink → PX4 SITL → Actuator Commands → Gazebo Model.
- State Feedback Flow: Gazebo Model State & Sensor Data → Gazebo ROS Plugins → ROS Topics (Sensor Msgs, Model States) → ROS Perception/Navigation Nodes.
- Autopilot Feedback: PX4 SITL (Internal State) → MAVLink → MAVROS → ROS Topics (Vehicle State).
For a drone formation, this cycle is replicated for each vehicle, with the added complexity of inter-vehicle communication for coordination, typically managed through additional ROS topics or services in a distributed or centralized manner.
2. Construction of the Drone Formation Simulation Platform
2.1 Modeling the Individual UAV and the Formation
The physical properties of each UAV in the drone formation are defined using the Unified Robot Description Format (URDF) and its macro-enabled extension, Xacro. These files describe the kinematic tree of the vehicle.
Key components defined as “ elements include:
base_link: The root coordinate frame of the UAV.body: The central fuselage.arm_front_right,arm_front_left,arm_rear_right,arm_rear_left: The structural arms.motor_front_right, …,motor_rear_left: The motor models.prop_front_right, …,prop_rear_left: The spinning propellers.- Sensor links:
laser_link,imu_link,camera_link.
These links are connected by “ elements specifying their kinematic relationships (fixed, revolute). The inertial properties (mass, inertia tensor), visual geometry, and collision geometry are specified for each link. For a multi-rotor, the propeller joints are modeled as continuous revolute joints. Sensor plugins are declared to interface simulated sensors like the 2D LiDAR with Gazebo’s physics engine. The Xacro macro capability is then exploited to efficiently spawn multiple instances of this base model, creating the drone formation by parameterizing each drone’s initial spawn position $$S_i = [X_{0,i}, Y_{0,i}, Z_{0,i}]$$ within the Gazebo world, typically in a grid pattern.
| Component | Parameter | Typical Value / Description |
|---|---|---|
| Vehicle Dynamics | Mass ($$m$$) | 1.5 kg |
| Diagonal Inertia Matrix ($$I_{xx}, I_{yy}, I_{zz}$$) | (0.0081, 0.0081, 0.0142) kg·m² | |
| Motor Thrust Coefficient ($$k_F$$) | 8.54858e-6 N/(rad/s)² | |
| Control (PX4) | Position Control PID Gains ($$k_{P,pos}, k_{I,pos}, k_{D,pos}$$) | Tuned for stable flight in simulation |
| Velocity Control PID Gains | Tuned for stable flight in simulation | |
| Sensors (Simulated) | 2D LiDAR (Range, FOV, Samples) | 30 m, 270°, 1080 samples |
| IMU Noise (Gyro, Accel) | Gaussian noise model applied |
2.2 Multi-UAV Communication and Control Logic
Coordinating a drone formation requires a robust communication architecture. In our simulation, each UAV operates with its own instances of PX4 SITL and MAVROS nodes, creating independent communication channels. A centralized or decentralized control scheme can be implemented in ROS. A common approach uses a “leader-follower” strategy. In this model, a designated leader drone receives high-level navigation goals. Its resulting trajectory or state is then communicated to follower drones via ROS topics.
Each follower $$i$$ calculates its desired setpoint relative to the leader’s state. For a formation maintaining a fixed geometric shape, if the leader’s position at time $$t$$ is $$p_L(t)$$ and the desired relative offset for follower $$i$$ is $$d_i$$ (e.g., $$d_1 = [3, 0, 0]^T$$ meters for a drone to the right), then the follower’s target position is:
$$ p_{t,i}(t) = p_L(t) + R_L(t) \cdot d_i $$
where $$R_L(t)$$ is the rotation matrix representing the leader’s orientation, ensuring the formation rotates with the leader. These individual targets are sent to each UAV’s MAVROS node, which forwards them to the respective PX4 instance for low-level tracking. The global TF (transform) library in ROS is essential for maintaining the spatial relationships between all drones and the world frame.
2.3 Optimization of Communication Scripts
To manage the complexity of launching and communicating with a drone formation, shell and Python scripts are essential. The original scripts for managing multi-UAV communication were analyzed and optimized for better performance and maintainability.
| Script File | Initial Issue | Optimization Applied | Benefit |
|---|---|---|---|
multi_uav_communication.sh (Launch Script) |
Used multiple sequential `while` loops and hard-coded values for each UAV, leading to redundant code and poor scalability. | Introduced arrays to store vehicle types and IDs. Encapsulated launch logic into a function called within a single loop iterating over the array. | Improved code clarity, reduced redundancy, and enhanced scalability for larger formations. |
multi_uav_communication.py (Control Script) |
Used a complex message type (`PositionTarget`) with high parameter redundancy and lacked input validation. | Simplified internal data handling using `Vector3` messages. Added robust exception handling to filter invalid or malformed pose/velocity commands. | Increased code readability and robustness, preventing crashes from erroneous input data during formation maneuvers. |
3. Core Algorithms for Formation Flight and Navigation
3.1 Formation Flight Control
The fundamental test for the drone formation platform is basic coordinated movement. We implemented keyboard control to command the entire formation as a single entity. The control mapping translates key presses into velocity or position setpoints for the leader, which are then propagated to followers using the logic described in Section 2.2.
Key parameters defined for smooth control include:
- Maximum Linear Velocity ($$v_{max}$$): 1 m/s
- Maximum Angular Velocity ($$\omega_{max}$$): 0.3 rad/s
- Linear Velocity Increment ($$\Delta v$$): 0.01 m/s
- Angular Velocity Increment ($$\Delta \omega$$): 0.01 rad/s
The control law for the leader based on keyboard input can be expressed as a simple integration:
$$ p_L(t+\Delta t) = p_L(t) + v_c(t) \cdot \Delta t $$
where $$v_c(t)$$ is the commanded velocity vector derived from keyboard input.
A test was conducted with a 3×3 formation (9 drones). The formation was commanded to take off, hover, then move 1 meter sequentially along the positive X, Y, and Z axes. The position of each drone before and after the maneuver was recorded. The ability to maintain the relative grid formation during translation indicates successful low-level control and inter-drone communication. The Circular Error Probable (CEP) for a unit movement can be approximated by analyzing the final position errors. If the expected displacement is $$D$$ and the average positional error magnitude across the formation is $$\bar{\Delta E}$$, the unit flight precision $$P$$ is:
$$ P = \left(1 – \frac{\bar{\Delta E}}{D}\right) \times 100\% $$
Analysis of the test data yielded a unit flight precision of approximately 76% for the simulated drone formation under coordinated velocity commands.
3.2 2D Laser-Based SLAM for Map Construction
For autonomous operations in an unknown or partially known farm environment, the drone formation must be capable of building a map. We employed a scan-matching algorithm, specifically Hector SLAM, due to its computational efficiency and robustness in structured environments typical of agricultural settings (e.g., greenhouses, structured orchards).
Hector SLAM estimates the UAV’s pose in a 2D plane by aligning consecutive LiDAR scans with an evolving grid map. It optimizes for the pose $$\xi_t = (x, y, \psi)$$ (x, y, yaw) that maximizes the probability of the current scan $$z_t$$ given the map $$m_{t-1}$$ built from previous scans:
$$ \xi_t^* = \arg \max_{\xi_t} \, P(z_t | \xi_t, m_{t-1}) $$
This is achieved using a Gauss-Newton optimization on the occupancy grid. The map $$m$$ is an occupancy grid map where each cell holds a log-odds probability of being occupied. The update for a cell $$m_i$$ given a scan is:
$$ l(m_i | z_{1:t}, \xi_{1:t}) = l(m_i | z_{1:t-1}, \xi_{1:t-1}) + l(m_i | z_t, \xi_t) – l(m_0) $$
where $$l$$ denotes log-odds and $$l(m_0)$$ is the prior.
In simulation, a single drone from the formation, equipped with a simulated 2D LiDAR, was flown through a complex “figure-8” shaped corridor with obstacles, mimicking farm structures. The Hector SLAM algorithm, running as a ROS node, processed the laser scan data (`/scan` topic) and the drone’s odometry (from PX4, fused with IMU data), to produce a real-time 2D occupancy grid map published on the `/map` topic. The resulting map accurately reflected the corridors and obstacles, demonstrating the platform’s capability for environmental perception—a prerequisite for autonomous drone formation navigation in confined farm spaces.
| Algorithm | Sensor Required | Computational Load | Robustness to Motion | Suitability for Farm UAVs |
|---|---|---|---|---|
| Hector SLAM | 2D LiDAR | Low to Moderate | Moderate (Requires good odometry) | High for structured environments (e.g., greenhouses, rows). |
| Gmapping (RB PF) | 2D LiDAR + Odometry | High (Particle Filter) | High | Good, but heavier computation may limit swarm size. |
| Cartographer | 2D/3D LiDAR + Odometry | Moderate to High | Very High | Excellent for large-scale outdoor farms, but complex. |
| Visual SLAM (ORB-SLAM) | Monocular/Stereo Camera | Moderate | Moderate | Good for daylight operations, sensitive to lighting changes. |
3.3 Autonomous Navigation Using Global Path Planning
With a pre-built or dynamically built map, the next step for the drone formation is autonomous point-to-point navigation. This involves global path planning from a start pose to a goal pose within the map. The ROS navigation stack (`move_base`) was utilized for this purpose. It integrates a global planner, a local planner (for obstacle avoidance), and a recovery behavior system.
For global planning on the 2D occupancy grid, we compared two classic algorithms: Dijkstra’s and A*. Both find the shortest path but use different strategies. Dijkstra’s algorithm explores all nodes with the lowest cumulative cost $$g(n)$$ from the start. A* uses a heuristic $$h(n)$$ (often Euclidean distance to goal) to guide the search, evaluating nodes by $$f(n) = g(n) + h(n)$$. While A* is generally faster, our tests within the ROS `global_planner` module showed that the implemented A* sometimes produced paths with unnecessary sharp turns and small discontinuities compared to the smoother paths from Dijkstra’s algorithm. Therefore, Dijkstra was selected for its reliability in generating coherent paths for the UAV to follow. The cost function for traversing from node $$n_i$$ to $$n_j$$ is primarily based on the distance, but can be inflated near obstacles:
$$ \text{Cost}(n_i, n_j) = w_d \cdot \|n_j – n_i\| + w_o \cdot \text{Inflation}(n_j) $$
where $$w_d$$ and $$w_o$$ are weights for distance and obstacle inflation, respectively.
An autonomous navigation test was conducted. Ten consecutive goal points were set within the simulated “figure-8” environment using the `2D Nav Goal` tool in RViz. The drone successfully planned a path to each goal using the Dijkstra global planner and the built map, and the PX4 controller executed the path by following the sequence of intermediate poses. The actual reached coordinates were logged and compared against the target coordinates. The results show a predictable growth in cumulative positional error with increasing flight distance, which is characteristic of odometry drift in the absence of absolute global positioning (like GPS, which was disabled for this indoor-style test). A linear regression model of cumulative distance flown ($$D_{cum}$$) versus cumulative error ($$E_{cum}$$) yielded a strong correlation:
$$ E_{cum} = 0.0253 \cdot D_{cum} + 0.3512, \quad R^2 = 0.8309 $$
This relationship quantifies the navigation performance of a single agent within the drone formation framework under these specific simulation conditions.
| Goal # | Target Coordinates (x, y, z) m | Actual Coordinates (x, y, z) m | Error for This Goal (m) | Cumulative Flight Distance (m) | Cumulative Error (m) |
|---|---|---|---|---|---|
| 1 | (3.5, 4.0, 1.0) | (3.46, 4.32, 0.92) | 0.335 | 5.32 | 0.335 |
| 2 | (3.5, 8.0, 1.0) | (3.38, 7.26, 1.27) | 0.778 | 9.32 | 1.113 |
| 3 | (5.0, 9.0, 1.0) | (4.50, 9.33, 0.83) | 0.553 | 11.05 | 1.666 |
| … | … | … | … | … | … |
| 10 | (6.0, 15.0, 1.0) | (4.58, 14.01, 0.98) | 1.568 | 37.82 | 7.892 |
4. Discussion: Implications for Unmanned Smart Farms
The successful simulation of drone formation flight control, mapping, and navigation directly translates to potential applications in unmanned smart farms. The demonstrated capabilities form a foundational toolkit for more complex agricultural workflows.
Scalable Field Monitoring: A drone formation can disperse over a large field to simultaneously capture multispectral or thermal imagery from multiple points, drastically reducing the time for crop health assessment compared to a single UAV. The simulation platform allows for testing efficient area coverage patterns and communication strategies for data fusion before real-world deployment.
Precision Task Execution: Coordinated drone formations could perform synchronized precision spraying or granular application. One drone could identify a weed patch via real-time vision, and communicate its location to nearby sprayer drones in the formation. The simulation environment is ideal for developing and refining such cooperative control algorithms, including managing downwash effects and ensuring safe inter-drone spacing during close-proximity work.
Infrastructure Inspection: The SLAM and autonomous navigation capabilities enable a drone formation to inspect farm infrastructure like irrigation systems, fences, or solar panel arrays in GPS-denied or obstructed environments, such as under canopies or inside barns. The simulation allows risk-free testing of obstacle avoidance and formation re-shaping algorithms in such complex 3D environments.
Limitations and Future Work: The current study primarily validates individual capabilities (formation flying, SLAM, navigation) in sequence. The next critical step is integrated, cooperative autonomy for the entire drone formation. This includes:
- Collaborative SLAM (C-SLAM): Where multiple drones share map and pose information to build a consistent global map faster and more accurately.
- Distributed Task Allocation: Developing algorithms for the formation to dynamically assign sub-tasks (e.g., covering different field zones, responding to specific pest detections) based on individual capabilities and positions.
- Robust Cooperative Navigation: Implementing fault-tolerant navigation where the formation can reconfigure if one member fails or loses communication, ensuring mission continuity.
Testing these advanced behaviors first in the high-fidelity simulation platform described here significantly de-risks subsequent physical prototype development and field trials.
| Advantage | Description | Benefit for Farm Application Development |
|---|---|---|
| Open-Source & Low Cost | All core components (ROS, PX4, Gazebo) are free and open-source, eliminating software licensing costs. | Makes advanced R&D accessible to universities, research institutes, and small-to-medium agricultural tech enterprises. |
| High Fidelity & Extensibility | Gazebo provides realistic physics and sensor models. Models and worlds can be customized (e.g., adding wind, rain, specific crop models). | Allows testing under varied and controllable environmental conditions (windy days, low light) that are critical for robust agricultural UAV design. |
| Modularity & ROS Integration | The system is built on the modular ROS framework, allowing easy swapping of algorithms (e.g., different SLAM or planner modules). | Enables rapid prototyping and benchmarking of different autonomy stacks tailored for specific farm tasks (e.g., row-following vs. orchard navigation). |
| Hardware-in-the-Loop (HITL) Ready | The PX4 SITL can be replaced with a real PX4 flight controller connected to the simulation, bridging the gap to real hardware. | Facilitates testing of actual flight hardware and low-level controller tuning in a safe, simulated farm environment before field deployment. |
5. Conclusion
This study successfully established a comprehensive simulation environment for multi-rotor drone formation cooperative flight, leveraging the integration of ROS, PX4, and Gazebo. By elucidating the critical data interactions between these systems and optimizing the communication infrastructure for multi-UAV management, we constructed a robust platform capable of simulating key behaviors required for unmanned smart farm operations. The platform demonstrated effective drone formation flight control with a unit movement precision of approximately 76%, successful 2D environmental mapping using laser-based Hector SLAM, and reliable autonomous navigation via Dijkstra’s global path planning algorithm, with a predictable error growth characterized by $$E_{cum} = 0.0253D_{cum} + 0.3512$$.
The results affirm the technical feasibility of using coordinated drone formations for agricultural tasks such as synchronized field scouting and structured-environment inspection. The open-source, modular, and extensible nature of the simulation platform provides an invaluable sandbox for developing, testing, and de-risking advanced cooperative algorithms—such as collaborative SLAM, distributed task allocation, and fault-tolerant navigation—before committing to costly and risky physical field trials. This work not only demonstrates a practical pathway for drone formation application in agriculture but also provides a foundational reference and a versatile toolset for researchers and engineers aiming to push the boundaries of automation in the next generation of smart and sustainable farming systems.
