52.22%
Macro task success
Best reported result, achieved by PI0.5 after multi-task and task-specific fine-tuning.
Overview video coming soon
A concise tour of the tasks, aerial platforms, policy interfaces, and dynamics-critical evaluation.
AM-Bench is designed to systematically evaluate aerial manipulation across diverse dimensions: (Left) a range of platforms, including underactuated (UA-Quad, UA-Hexa) to fully/over actuated (FA-Hexa, Omni-Hexa) systems and different low-level controllers; (Center) a diverse suite of manipulation tasks with domain randomization in spatial poses and textures; (Right) various visuomotor policy architectures and realistic physical constraints, including aerodynamic disturbances (ground/wall effects) and actuator saturation. u, φ(·, ·), s, a, π(·), and o refer to control input, controller, state, high-level action, policy, and observation, respectively.
The current paper evaluates each high-level policy over 30 rollouts on each of 12 tasks using FA-Hexa and the EE-target IK–PID interface. These figures are a fixed paper snapshot, not a live leaderboard.
52.22%
Best reported result, achieved by PI0.5 after multi-task and task-specific fine-tuning.
67.31%
Best reported intermediate-progress result under the same evaluation setting.
+46.4 pp
PI0.5 macro success improvement from zero-shot evaluation to multi-task fine-tuning.
48%
Ground-effect modeling reduced simulation-to-real EE-height RMSE from 1.91 cm to 0.99 cm in the tested interval.
Twelve tasks span instantaneous interaction, object transport, articulated objects, and constrained contact.
Task 1: Press Button
Task 2: Peg in Hole
Task 3: Frame Assembly
Task 4: Cabinet Pick and Place
Task 5: Lemon Harvesting
Task 6: Rotate Valve
Task 7: Push Slider
Task 8: Pull Lever
Task 9: Open Door
Task 10: Wipe Window
Task 11: Toss Ball
Task 12: NDT
Configurable geometry, placement, textures, and object appearance expose policies to controlled visual and spatial variation.
Wall Texture
Wall Tilt
Aerodynamic and actuation models test policy behavior under wind, ground effect, near-wall interaction, and rotor saturation.
Wind Disturbance
Ground Effect Disturbance
Four multirotor embodiments cover underactuated, fully actuated, and overactuated aerial manipulation platforms under a shared benchmark interface.
UA-Quad
UA-Hexa
FA-Hexa
Omni-Hexa
The framework follows a modular pipeline composed of diverse task environments, robot models, disturbance, actuator saturation, low-level controller and high-level visuomotor policy, under a high-fidelity simulation environment.
Community benchmark
The public leaderboard will compare submissions under matched task, embodiment, controller, disturbance, rollout, and seed settings. The current paper results remain available as a fixed experimental snapshot.
Please use the provisional citation below. The archival paper link and publication record will replace it when they become available.
@article{wang2026ambench,
title = {{AM-Bench}: A Modular Simulation Suite and Benchmark for Aerial Manipulation Policy Learning},
author = {Wang, Yutong and Lee, Dongjae and Guo, Xiaofeng and Zhan, Yuanzhu and
Jiang, Yufei and Saravanan, Bavin and Cao, Muqing and Xie, Jia and
Mao, Chenyang and Scherer, Sebastian and Geng, Junyi and Shi, Guanya},
year = {2026},
note = {Preprint}
}