AGIBOT's AI models: GO-1, GO-2, WITA-Omni and AGILE

Short answer
AGIBOT develops AI for three skills: movement, interaction and manipulation (grasping and using objects). There are dedicated models for this, such as GO-1 and GO-2 for tasks, WITA-Omni for seeing and talking, AGILE 2.0 for walking and GE-Act 2.0 as a world model. You notice them through software updates for your robot.
Overview
| Model | What for | What AGIBOT reports |
|---|---|---|
| GO-1 | general model for robot tasks | presented in 2025 as AGIBOT's first general model |
| GO-2 | planning and carrying out tasks | reasons in steps; scores 98.5% on the LIBERO benchmark |
| WITA-Omni | seeing, hearing and talking at the same time | 85.21% on the Daily-Omni benchmark; recognises who is speaking and turns towards the speaker |
| AGILE 2.0 | walking with vision in the control loop | X2 balances on a rolling ball and skips rope |
| GE-Act 2.0 | world model that predicts and acts | success rises with more training data, see world models |
| Genie Envisioner-Sim 2.0 | simulating and testing | number 1 on the WorldArena benchmark |
What does this mean for you?
- Better conversations: WITA-Omni makes interaction more natural, important for customer guidance.
- More stable walking: AGILE 2.0 helps in busy, unpredictable environments.
- New tasks faster: GO-2 and world models reduce training time, see robot foundation models.
- Benchmarks are not practice: always test on your own task.
Terms explained: VLA models, embodied AI and the glossary.
Source: GO-2, WITA-Omni, AGILE 2.0, GE-Act 2.0, Genie Envisioner-Sim 2.0. We summarise it in our own words.
Frequently asked questions
What is AGIBOT's GO-1?
AGIBOT's first general AI model for robot tasks, presented in 2025. It is used through Genie Studio.
What is WITA-Omni?
An AI model that processes images, sound and speech at the same time, so the robot responds more naturally in a conversation.
Will my robot get these models automatically?
New AI arrives through software updates. Which functions become available differs per model and version.