In order to use humanoid robots in industrial applications, they must meet criteria relevant to those applications. The Fraunhofer IPA has developed a benchmark for this purpose. This allows manufacturers and end users to have humanoid robots analyzed by an independent third party for energy efficiency, cleanroom compatibility, data security, and more. The Fraunhofer Institute for Manufacturing Engineering and Automation IPA has developed a comprehensive benchmark for the standardized analysis of humanoid robots.For the first time, this allows manufacturers and end users to have the actual capabilities, safety, and operational suitability of these robots objectively evaluated by an independent third party. The modular benchmark comprises six application-relevant criteria and is based on internationally recognized industry standards.
From Media Presence to Realistic AssessmentHumanoid robots are omnipresent in the media and fascinate with their human-like appearance. Yet there is a vast gap between spectacular presentations and actual capabilities.
“For end users and manufacturers alike, it is essential to look behind the facade sometimes constructed by marketing agencies,” explains Simon Schmidt, Head of the Automated Systems Division at Fraunhofer IPA. “The market is too volatile and opaque to allow for a well-founded assessment and reliable evaluation of humanoids for specific applications.”
What is the benchmark?The benchmark is a standardized service in which research teams at Fraunhofer IPA guide humanoid robots through various challenges and scientifically evaluate the results. The foundation for this was established thanks to funding from the Baden-Württemberg Ministry of Economic Affairs, Labor, and Tourism as part of the AI Advancement Center “Learning Systems and Cognitive Robotics.”
The benchmark’s modular structure enables manufacturers, end users, and software providers to specifically test the areas relevant to their applications. Where possible, the benchmarking is based on established industry standards that have been internationally recognized for decades—such as ISO 14644 for cleanroom compatibility or ISO 10218 and ISO TS 15066 for functional safety.
The benchmark is divided into six key areas:
- Technologies and Basic Capabilities: Examination of the built-in sensors, AI models, and gripper types, as well as tests of walking speed, gripping forces, and handleable loads. Objective measurements are recorded using a 3D tracking system and force sensors.
- Complex capabilities: Evaluation of practical, generic tasks such as stair climbing, obstacle navigation, movement and force precision, and reaction speed. The tests are deliberately designed to be challenging to ensure comparability with future model generations.
- Cleanroom compatibility: Evaluation of particle emission according to ISO 14644-14, outgassing behavior, and cleanability—critical for applications in the semiconductor, pharmaceutical, or food industries.
- Functional Safety: Central to human-robot collaboration. Tests include stability on various surfaces, force limitation during collisions, obstacle detection, and system behavior during failures. Collision tests are conducted using the same force sensors as for collaborative industrial robots.
- Cybersecurity (Security): Four modules assess vulnerability management, secure lifecycle, network security, and penetration resistance—a critical factor given increasing regulatory requirements.
- Energy efficiency: Measurement of battery life and power consumption in various scenarios (standing, walking, walking uphill, and walking while carrying a load). The results enable realistic operational planning and optimization of charging cycles.
Robots: Good Self-StabilizationUsing the Unitree G1 as an example, Fraunhofer IPA applied the benchmark comprehensively for the first time. The technical basis was a Unitree G1 EDU-4 delivered in May 2025, equipped with Dex3-1 3-finger hands and firmware version 1.04.
While the robot demonstrates good self-stabilization and could be suitable for ISO Class 5 cleanrooms, significant limitations also became apparent. In the event of collisions, forces exceeding 500 newtons can occur—far above the pain thresholds permitted by the standard.
In addition, the researchers identified a critical Bluetooth security vulnerability in the software version available at the time of testing, which allows attackers to take complete remote control. This vulnerability has since been fixed.
In terms of energy efficiency, maximum operating times were 2 hours and 49 minutes on a single battery charge while simply standing, and 1 hour and 49 minutes in a typical scenario involving both standing and walking.
The Relevance of the Benchmark for Businesses“Users can interpret the results directly and thus find the right humanoid for the right application,” emphasizes Werner Kraus, Head of Research at Fraunhofer IPA. The benchmark makes humanoids comparable not only with each other but also with proven automation components. This is particularly important because:
- demographic change is driving the adoption of automation in areas that were previously manual
- major investment decisions require well-founded, objective evaluation criteria
- safety standards for humanoids are not expected until 2028 (ISO 25785-1)
- regulatory requirements for cybersecurity are increasing
- sensitive production environments require reliable data to
prevent contamination
The benchmark provides transparency in an opaque market and enables companies to develop realistic expectations and minimize risks.
The Fraunhofer IPA plans to test additional humanoids and establish a comparative database. Manufacturers and users can now commission individual benchmark modules or comprehensive evaluations and benefit from the existing infrastructure and expertise.