Vision-language-action (VLA) models are changing how robots learn. Instead of relying on separately engineered systems for perception, task understanding, and control, VLAs combine these capabilities in a single learned policy that can interpret visual observations and natural-language instructions and translate them into actions.
A new survey paper by Chef Robotics Senior Staff AI Research Scientist Inkyu Sa examines the rapidly evolving field through one of robotics' hardest challenges: bimanual manipulation, or coordinating two arms to complete tasks such as folding fabric, assembling objects, and handling deformable materials.
Drawing on more than 200 sources, the paper compares 31 VLA methods across their architectures, training strategies, action representations, coordination approaches, and reported performance. It also examines real-world deployments in food assembly, manufacturing, logistics, homes, healthcare, laboratories, and agriculture.
Download the full survey paper
Download
Not all two-arm tasks are equally challenging
The paper's central argument is that the difficulty of a bimanual task depends less on the task category than on how closely the two arms must coordinate.
Some tasks can be divided into largely independent actions. Others require the arms to synchronize their timing, as in a handoff. The hardest tasks require both arms' movements and forces to remain continuously aligned (e.g., tensioning a piece of fabric without wrinkling or dropping it).
For these tightly coupled tasks, the survey finds that models generating both arms' actions jointly are better suited than approaches that produce commands sequentially or treat each arm independently. Flow-matching models currently offer the strongest reported balance between expressive action generation and the speed required for real-time control. However, the paper emphasizes that no standard bimanual benchmark exists, so results from different systems cannot yet be compared under controlled conditions.

Research systems have made rapid progress on tasks such as laundry folding, box assembly, household cleaning, and operation on previously unseen robot platforms. At the same time, the systems operating at the greatest commercial scale tend to perform narrower, more structured tasks. Production environments commonly require reliability above 99% over an entire shift, far beyond the success rates reported on many general-purpose real-robot evaluations.
This suggests that better models alone will not be enough. Hardware, sensing, data collection, inference speed, safety systems, evaluation methods, and the surrounding production workflow all shape whether a robot can operate reliably outside the lab.
Three priorities for the field
The survey identifies three areas that will be especially important for the next generation of bimanual systems:
- Better evaluation: The field needs a shared benchmark for two-arm manipulation, larger evaluation sets, and reporting standards that make results genuinely comparable.
- Dexterity and multimodal sensing: Vision-only systems with simple grippers cannot reliably measure the contact forces involved in many tightly coupled tasks. Tactile, force, and audio sensing will become increasingly important as robots take on more dexterous work.
- Production-grade safety and reliability: Commercial deployment requires runtime monitoring, constrained action generation, collision avoidance, and stronger evidence that systems can operate safely and consistently around people.
Bimanual manipulation has progressed remarkably quickly, moving from isolated demonstrations to autonomous operation in real homes in only a few years. The next challenge is closing the gap between robots that can perform impressive tasks and systems that can perform useful work safely, reliably, and at scale.
Download the full survey paper to explore the complete taxonomy, method comparisons, deployment analysis, and research agenda, and get in touch with our team.
Download the full survey paper
Download%20(1).jpg)


.jpg)