AGT-CV: Drexel's Air-Ground Perception Dataset Targets the Off-Road Blind Spot in Multi-Robot AI
Drexel's AGT-CV dataset pairs a Clearpath Husky UGV with an Autel EVO II UAV for real-world cross-view collaborative perception research.

Main Story
Collaborative perception between aerial and ground robots has long been constrained by a fundamental data problem: virtually no real-world datasets exist that capture overlapping, multi-modal observations from heterogeneous platforms navigating genuinely unstructured terrain. A new open dataset from Drexel University's robotics group directly addresses this gap.
The AGT-CV (Aerial-Ground Team Cross-View) dataset was collected using a Clearpath Husky UGV and an Autel EVO II UAV across diverse unstructured environments, including forest trails, rocky paths, muddy terrain, snow piles, and grass-covered fields.
The research team chose this pairing deliberately. Heterogeneous air-ground robot teams combine complementary sensing modalities, mobility characteristics, and spatial viewpoints that can significantly enhance perception in complex outdoor environments. Yet progress in multi-robot collaborative perception has been constrained by the lack of real-world datasets featuring overlapping multi-modal observations from platforms operating in unstructured terrain.
AGT-CV is designed to be the corrective. The dataset is collected across four unique environments, with over 13,000 synchronized frames spanning approximately 29 minutes of operation, and includes both SAM 3-based zero-shot segmentation and almost 8,000 manually labeled images.
A defining methodological choice was the collection timing. A unique aspect of the dataset is its early-spring collection period, during which sparse tree canopies allow the aerial robot to partially observe the ground robot and terrain through the trees, enabling occlusion-aware collaborative perception. This deliberate seasonal selection gives researchers a rare natural laboratory for studying partial occlusion — one of the hardest failure modes to replicate synthetically.
Unlike prior multi-robot datasets that primarily focus on SLAM or simulated cooperative driving, AGT-CV is specifically designed to support research on cross-view perception, air-ground viewpoint fusion, terrain-aware perception, and collaborative scene understanding in real off-road environments.
On the annotation side, the team deployed Meta's SAM 3, a zero-shot model released in November 2025 that detects, segments, and tracks objects in images and videos from concept prompts. SAM 3 introduces Promptable Concept Segmentation: given a short noun phrase, it returns unique masks and IDs for every matching instance at once — a capability where SAM 1 and 2 predicted only one object per prompt. Using SAM 3 for automated pre-annotation, alongside nearly 8,000 hand-labeled frames, gives AGT-CV a dual-track ground truth that future benchmarks can leverage at both ends of the labeling effort spectrum.
Technical Breakdown
Ground Platform — Clearpath Husky UGV The Clearpath Husky is a medium-sized unmanned ground vehicle designed for research and industrial applications in outdoor environments, with a steel chassis, all-wheel-drive electric drivetrain, and a payload capacity of up to 75 kg. It weighs 50 kg, reaches a maximum speed of 1.0 m/s, and offers 130 mm of ground clearance. In the AGT-CV configuration, the ground platform provides 3D LiDAR, stereo camera, IMU, and GPS data. The platform includes an onboard computer running Ubuntu and ROS, with pre-configured drivers for common sensors including Velodyne LiDAR and various camera systems.
Aerial Platform — Autel EVO II UAV The aerial platform contributes RGB imagery, thermal/infrared observations, and GPS from a complementary overhead viewpoint, enabling rich cross-modal and cross-view perception. The Autel EVO II Dual thermal variant combines an infrared imaging camera with an 8K visible-light camera. The infrared sensor features a vanadium oxide uncooled focal plane detector with 640×512 resolution, a 13 mm focal length, and 1–8× zoom. Maximum horizontal flight speed reaches 72 km/h and maximum flight time is 38 minutes.
Dataset Composition
- Environments: 4 distinct terrain types (forest trails, rocky paths, muddy terrain, snow, grass)
- Frames: >13,000 synchronized multi-modal frames
- Duration: ~29 minutes of operational recording
- Labels: ~8,000 manually annotated images + SAM 3 zero-shot segmentation masks
- Sensor modalities: 3D LiDAR, stereo camera, IMU, GPS (ground); RGB, thermal/IR, GPS (aerial)
Autonomy Level The dataset targets collaborative perception research rather than closed-loop autonomy. The platforms operate as a coordinated observational team, with cross-view and cross-modal data fusion being the primary research target.
Industry Impact
For Robotics Researchers and Dataset Curators AGT-CV fills a documented vacuum. Very few datasets are designed for heterogeneous multi-robot applications, and prior sensor modality setups were generally unsuitable for multi-robot collaborative perception — making AGT-CV among the first real-world heterogeneous air-ground collaborative perception datasets explicitly designed for multi-modal collaborative robot perception research. Teams working on field robotics, search-and-rescue, and autonomous inspection now have a benchmark that reflects the sensor diversity of operational deployments.
For UAV and UGV Manufacturers The pairing of a Clearpath Husky — the most widely deployed ROS-based outdoor UGV in research, used at hundreds of universities and research institutions worldwide — with the Autel EVO II's dual-sensor payload validates a practical, commercially available hardware stack. Manufacturers building multi-robot platforms for off-road applications can reference AGT-CV as a benchmark configuration.
For AI and Perception Engineers The dataset's dual annotation strategy — combining SAM 3 zero-shot masks with ~8,000 hand-labeled images — establishes a practical workflow for scaling perception ground truth at low cost. SAM 3 generally performs well in zero-shot settings but benefits from fine-tuning for niche domains — precisely the kind of domain-specific gap that AGT-CV's manual labels are positioned to close. Engineers developing terrain-aware segmentation and cross-view localization models gain a ready-made evaluation split that does not rely on simulation.
For Defence and Critical-Infrastructure Operators Off-road multi-robot perception is a core capability requirement for perimeter monitoring, disaster response, and infrastructure inspection in GPS-degraded, canopy-occluded environments. AGT-CV's explicit focus on occlusion-aware collaborative perception and early-spring canopy conditions makes it directly relevant to operational scenario planning in these sectors.
For Standards and Certification Bodies As regulators begin to consider certification frameworks for autonomous multi-robot systems operating beyond visual line of sight in complex terrain, curated real-world datasets like AGT-CV will be essential inputs to evidence-based performance standards. The dataset's synchronized, multi-modal structure and dual ground-truth methodology offer a reproducible evaluation baseline that standards bodies can reference.
