CoRE

Learning Collaboration-Role Experts for Decentralized Collaborative Manipulation with One Policy

Yanan Zhou1, Zhaoyan Qian1, Zihao Li1, Mingyuan Ba1, Ranpeng Qiu2, Weiming Zhi1*

1 School of Computer Science, The University of Sydney
2 Zhejiang University of Technology · * Corresponding author

Fig. 1 · Multi-robot collaboration paradigmsPaper ↗

Abstract

Collaborative manipulation requires robots to perform complementary actions as interactions unfold. We study single-policy decentralized collaboration: every robot runs the same policy from its visual observations and proprioception, without task prompts, identity labels, or inter-robot messages. The challenge is to learn complementary team behaviors within shared parameters and select appropriate actions from each robot's local observations.

We introduce CoRE, which learns Collaboration-Role Experts from pooled multi-task, multi-robot demonstrations. Fused appearance and geometry provide local interaction evidence. Query-conditioned cross-attention experts provide adaptable prediction paths, which a local router combines at each action-chunk position. During training, an action–expert alignment loss supervises expert selection using relative forced-route prediction errors against demonstrations under fixed inputs, without role labels.

Across simulation benchmarks, CoRE achieves the highest average performance among evaluated decentralized methods. Physical experiments demonstrate effective collaboration across diverse manipulation tasks and robustness to partner delays and slowdowns.

III / Method

Collaboration-role experts.

Local observations → shared experts → robot-specific actions.

Fig. 2 · CoRE overviewPaper ↗

IV / Experiments

Tasks and evaluation.

13 tasks · 3 benchmarks · 2–4 robots

Fig. 3 · Simulation benchmarksPaper ↗

The following videos are task demonstrations, not CoRE policy rollouts.

RoboFactory

5 task demonstrations · 2–4 robots

DuoBench

4 task demonstrations · 2 robots

BiCoord

4 task demonstrations · 2 robots

RQ1 / Single-policy collaboration

Does one shared policy collaborate?

Highest benchmark-average performance among the evaluated decentralized methods.

RoboFactory · SR

Independent Regression14.0%
CLS-DP73.8%
GauDP20.2%
CHORUS58.4%
CoRE84.2%
0%100%

DuoBench · SSR

Independent Regression25.0%
CLS-DP40.7%
GauDP27.4%
CHORUS43.5%
CoRE47.6%
0%100%

BiCoord · SR

Independent Regression21.0%
CLS-DP17.3%
GauDP14.0%
CHORUS26.5%
CoRE75.5%
0%100%

100 fixed seeds per task. SR: full-task success; DuoBench SSR: stage progress.

Full task results Table I ↗
Table I · Decentralized methods; DuoBench SSR (%), BiCoord and RoboFactory SR (/100)
MethodCarry PotSpring DoorTransfer GateBall MazeBalance RollerBuild BridgeExtract Bottom Block to TopSweep BlockLift BarrierCamera AlignmentThree-Robot StackingLong Pipeline DeliveryTake Photo
Independent Regression34.1%19.3%10.8%35.6%47/10022/1004/10011/10043/10019/1005/1000/1003/100
CLS-DP63.3%42.3%5.0%52.0%22/10034/1000/10013/10097/10096/10053/10068/10055/100
GauDP38.0%32.3%10.3%29.0%21/10017/1000/10018/10072/10026/1000/1000/1003/100
CHORUS66.0%32.3%16.3%59.5%83/10016/1001/1006/10099/10089/10064/10024/10016/100
CoRE75.3%37.7%13.3%64.0%93/10067/10070/10072/10099/100100/10099/10094/10029/100
Fig. 4 · Sequential, coupled and asymmetric collaborationPaper ↗
Table II · Demonstration sharing on RoboFactory
MethodDataPoliciesAvg. SR
RGB RegressionTask / robot1614.0%
RGB RegressionPer task553.6%
RGB RegressionAll tasks140.4%
CoRE (RGB)All tasks172.4%
CoREAll tasks184.2%

Selected columns from Table II. Full results in the paper.

RQ2 / Component contributions

What makes the difference?

With spatial inputs and expert paths fixed, alignment raises LPD success from 15% to 94%.

Table III · Module ablations; SR (%)
VariantLiftCamera3StackLPDPhotoAvg.
RGB Regression8798301440.4
Stereo Regression9210042202956.6
Stereo + MoE1009968172862.4
Stereo + MoE + ARCA9610081152964.2
CoRE (RGB)8910071861672.4
CoRE9910099942984.2
Table IV · Expert count and routing sparsity; SR (%)
ExpertsTop-KLiftCamera3StackLPDPhotoAvg.
229610099654180.2
41879682331562.6
4 (default)29910099942984.2
441001005943158.8
829910098972383.4

RQ3 / Shared expert use

How are experts used across robots?

Fig. 5 · Expert mixtures during LPD and Take PhotoPaper ↗
Sequential reuse across robotsLPD task demonstration · 2×↗
Concurrent complementary actionsTake Photo task demonstration · 2×↗

Task demonstrations illustrate the interactions; expert trajectories are shown in Fig. 5.

RQ4 / Simulation robustness

When a partner’s timing changes.

Lift: random delays. ThreeStack: random slowdown to 0.25×.

Lift · Partner delay

Lift: matched-stage success and failure illustrations
Success (%)025507510099%96%89%61%49%05102030Perturbed time (%)

100 trials per point · Bands: 95% Wilson intervals.

ThreeStack · Partner slowdown

ThreeStack: matched-stage success and failure illustrations
Success (%)025507510099%97%82%69%44%010203050Perturbed time (%)

100 trials per point · Bands: 95% Wilson intervals.

Fig. 6 · Collaboration under partner perturbations

RQ5 / Physical collaboration

From simulation to real robots.

Five tasks · Two independent Piper arms · 30 trials per method and task

Box assembly

CLS-DP5/30
CHORUS16/30
CoRE28/30
0%100%
Box assembly: selected real-robot collaboration frame

Pan transport

CLS-DP19/30
CHORUS23/30
CoRE29/30
0%100%
Pan transport: selected real-robot collaboration frame

Bridge assembly

CLS-DP2/30
CHORUS7/30
CoRE25/30
0%100%
Bridge assembly: selected real-robot collaboration frame

Block stacking

CLS-DP9/30
CHORUS14/30
CoRE26/30
0%100%
Block stacking: selected real-robot collaboration frame

Block exchange

CLS-DP18/30
CHORUS24/30
CoRE29/30
0%100%
Block exchange: selected real-robot collaboration frame

30 trials per method and task · Whiskers: 95% Wilson intervals.

Fig. 7 · Real-world task performance

Robustness on physical robots.

Fig. 9 · Failure cases under partner perturbationsPaper ↗

Box assembly · Partner pauses

Success (%)025507510016/304/301/300/3028/3022/3015/3011/30Normal1.0 s1.5 s2.0 sPause duration

30 trials per point · Bands: 95% Wilson intervals.

Pan transport · Partner slowdown

Success (%)025507510023/3015/306/301/3029/3027/3024/3017/30Normal0.75×0.5×0.25×Right robot speed

30 trials per point · Bands: 95% Wilson intervals.

Fig. 10 · Success rates under partner pauses and slowdown

Box · Partner pauses

CoRE · Success

Box assemblyPerturbed execution · 4×↗

Failure example

Box assemblyPerturbed execution · 4×↗

Pan · Slower partner

CoRE · Success

Pan transportPerturbed execution · 4×↗

Failure example

Pan transportPerturbed execution · 4×↗

Selected qualitative examples; failure method unspecified. Box recordings use different pause ranges. Quantitative protocols and 95% Wilson intervals are in Fig. 10.

Paper & resources

Cite CoRE BibTeX
@misc{zhou2026core,
  title={CoRE: Learning Collaboration-Role Experts for
         Decentralized Collaborative Manipulation with One Policy},
  author={Zhou, Yanan and Qian, Zhaoyan and Li, Zihao and
          Ba, Mingyuan and Qiu, Ranpeng and Zhi, Weiming},
  year={2026},
  note={Preprint}
}

On this page

0%

Figure

Expanded research figure

CoRE / Research overview

English narration · 2:59 · Download video ↗