As embodied AI enters a period of rapid model advancement, reliable evaluation standards remain in short supply. What can these models actually do well, and where do they fall short?
Some results have been hidden because they may be inaccessible to you
Show inaccessible results