CoRE
Learning Collaboration-Role Experts for Decentralized Collaborative Manipulation with One Policy
1 School of Computer Science, The University of Sydney
2 Zhejiang University of Technology · * Corresponding author
Abstract
Collaborative manipulation requires robots to perform complementary actions as interactions unfold. We study single-policy decentralized collaboration: every robot runs the same policy from its visual observations and proprioception, without task prompts, identity labels, or inter-robot messages. The challenge is to learn complementary team behaviors within shared parameters and select appropriate actions from each robot's local observations.
We introduce CoRE, which learns Collaboration-Role Experts from pooled multi-task, multi-robot demonstrations. Fused appearance and geometry provide local interaction evidence. Query-conditioned cross-attention experts provide adaptable prediction paths, which a local router combines at each action-chunk position. During training, an action–expert alignment loss supervises expert selection using relative forced-route prediction errors against demonstrations under fixed inputs, without role labels.
Across simulation benchmarks, CoRE achieves the highest average performance among evaluated decentralized methods. Physical experiments demonstrate effective collaboration across diverse manipulation tasks and robustness to partner delays and slowdowns.
III / Method
Collaboration-role experts.
Local observations → shared experts → robot-specific actions.
IV / Experiments
Tasks and evaluation.
13 tasks · 3 benchmarks · 2–4 robots
The following videos are task demonstrations, not CoRE policy rollouts.
DuoBench
4 task demonstrations · 2 robotsBiCoord
4 task demonstrations · 2 robotsRQ1 / Single-policy collaboration
Does one shared policy collaborate?
Highest benchmark-average performance among the evaluated decentralized methods.
RoboFactory · SR
DuoBench · SSR
BiCoord · SR
100 fixed seeds per task. SR: full-task success; DuoBench SSR: stage progress.
Full task results Table I ↗
| Method | Carry Pot | Spring Door | Transfer Gate | Ball Maze | Balance Roller | Build Bridge | Extract Bottom Block to Top | Sweep Block | Lift Barrier | Camera Alignment | Three-Robot Stacking | Long Pipeline Delivery | Take Photo |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Independent Regression | 34.1% | 19.3% | 10.8% | 35.6% | 47/100 | 22/100 | 4/100 | 11/100 | 43/100 | 19/100 | 5/100 | 0/100 | 3/100 |
| CLS-DP | 63.3% | 42.3% | 5.0% | 52.0% | 22/100 | 34/100 | 0/100 | 13/100 | 97/100 | 96/100 | 53/100 | 68/100 | 55/100 |
| GauDP | 38.0% | 32.3% | 10.3% | 29.0% | 21/100 | 17/100 | 0/100 | 18/100 | 72/100 | 26/100 | 0/100 | 0/100 | 3/100 |
| CHORUS | 66.0% | 32.3% | 16.3% | 59.5% | 83/100 | 16/100 | 1/100 | 6/100 | 99/100 | 89/100 | 64/100 | 24/100 | 16/100 |
| CoRE | 75.3% | 37.7% | 13.3% | 64.0% | 93/100 | 67/100 | 70/100 | 72/100 | 99/100 | 100/100 | 99/100 | 94/100 | 29/100 |
| Method | Data | Policies | Avg. SR |
|---|---|---|---|
| RGB Regression | Task / robot | 16 | 14.0% |
| RGB Regression | Per task | 5 | 53.6% |
| RGB Regression | All tasks | 1 | 40.4% |
| CoRE (RGB) | All tasks | 1 | 72.4% |
| CoRE | All tasks | 1 | 84.2% |
Selected columns from Table II. Full results in the paper.
RQ2 / Component contributions
What makes the difference?
With spatial inputs and expert paths fixed, alignment raises LPD success from 15% to 94%.
| Variant | Lift | Camera | 3Stack | LPD | Photo | Avg. |
|---|---|---|---|---|---|---|
| RGB Regression | 87 | 98 | 3 | 0 | 14 | 40.4 |
| Stereo Regression | 92 | 100 | 42 | 20 | 29 | 56.6 |
| Stereo + MoE | 100 | 99 | 68 | 17 | 28 | 62.4 |
| Stereo + MoE + ARCA | 96 | 100 | 81 | 15 | 29 | 64.2 |
| CoRE (RGB) | 89 | 100 | 71 | 86 | 16 | 72.4 |
| CoRE | 99 | 100 | 99 | 94 | 29 | 84.2 |
| Experts | Top-K | Lift | Camera | 3Stack | LPD | Photo | Avg. |
|---|---|---|---|---|---|---|---|
| 2 | 2 | 96 | 100 | 99 | 65 | 41 | 80.2 |
| 4 | 1 | 87 | 96 | 82 | 33 | 15 | 62.6 |
| 4 (default) | 2 | 99 | 100 | 99 | 94 | 29 | 84.2 |
| 4 | 4 | 100 | 100 | 59 | 4 | 31 | 58.8 |
| 8 | 2 | 99 | 100 | 98 | 97 | 23 | 83.4 |
RQ3 / Shared expert use
How are experts used across robots?
Task demonstrations illustrate the interactions; expert trajectories are shown in Fig. 5.
RQ4 / Simulation robustness
When a partner’s timing changes.
Lift: random delays. ThreeStack: random slowdown to 0.25×.
Lift · Partner delay

100 trials per point · Bands: 95% Wilson intervals.
ThreeStack · Partner slowdown

100 trials per point · Bands: 95% Wilson intervals.
RQ5 / Physical collaboration
From simulation to real robots.
Five tasks · Two independent Piper arms · 30 trials per method and task
Box assembly

Pan transport

Bridge assembly

Block stacking

Block exchange

30 trials per method and task · Whiskers: 95% Wilson intervals.
Box assembly
Pan transport
Bridge assembly
Block stacking
Block exchange
Robustness on physical robots.
Box assembly · Partner pauses
30 trials per point · Bands: 95% Wilson intervals.
Pan transport · Partner slowdown
30 trials per point · Bands: 95% Wilson intervals.
Box · Partner pauses
Pan · Slower partner
Selected qualitative examples; failure method unspecified. Box recordings use different pause ranges. Quantitative protocols and 95% Wilson intervals are in Fig. 10.
Paper & resources
Cite CoRE BibTeX
@misc{zhou2026core,
title={CoRE: Learning Collaboration-Role Experts for
Decentralized Collaborative Manipulation with One Policy},
author={Zhou, Yanan and Qian, Zhaoyan and Li, Zihao and
Ba, Mingyuan and Qiu, Ranpeng and Zhi, Weiming},
year={2026},
note={Preprint}
}