As embodied AI enters a period of rapid model advancement, reliable evaluation standards remain in short supply. What can these models actually do well, and where do they fall short?