LOOK AHEAD. ACT TOGETHER.

MA-WAMMulti-Agent World-Action Model for Test-Time Planning

MA-WAM compares candidate joint futures before acting, using a routed world model to choose a coordinated action sequence at test time.

Guowei ZouHaitao WangGuoxin WangBeiwen ZhangZhiquan ChenGuojie WangHejun Wu

Sun Yat-sen University

Explore the methodPaper & code · coming soon
Frozen policies · Test-time coordination · Offline MARL

01 / Overview

A frozen policy proposes. A world model looks ahead.

A reactive policy executes one proposed joint action without comparing its consequences with other possibilities. MA-WAM lets a frozen multi-agent flow policy propose several joint-action sequences, then uses a Routed World Model (RWM) to predict their outcomes. The planner selects the sequence with the highest predicted team return, executes its first action, and replans from the next observation.

+22.0%Mean relative gainover direct execution across 30 settings
+25.6%Mean relative gainover uniform candidate selection
26 / 30Settings improvedcompared with the reactive policy
12.1 msAdded planning time2.5% of generation-and-scoring time on A100
A reactive policy commits to one proposal. MA-WAM evaluates multiple joint futures before selecting an action.
Comparing candidate futures before execution.A reactive policy commits to one proposal. MA-WAM evaluates multiple joint futures before selecting an action.

02 / Method

Plan over joint futures.

The world model conditions on the joint observation and action to represent dependencies among agents. Routing combines shared experts for dynamics and reward prediction. The flow policy and world model remain fixed during deployment.

The frozen policy generates candidates; RWM predicts their returns; the first action of the selected sequence is executed.
From joint proposals to coordinated execution.The frozen policy generates candidates; RWM predicts their returns; the first action of the selected sequence is executed.
01

Propose

Sample candidate joint-action sequences from the frozen flow policy.

02

Predict & rank

Roll out candidates with RWM and rank their predicted cumulative team returns.

03

Execute & replan

Execute the first joint action of the selected sequence and plan again at the next observation.

Inside MA-WAM · Task explorer

选择任务,看它如何先想再行动。

正在载入…

03 / Videos

Cooperative trajectories in motion.

Play the trajectory videos archived with this paper project.

Particle environments

MPE · Spread

Expert offline trajectory · selected rank 1

MPE · Tag

Expert offline trajectory · selected rank 1

MPE · World

Expert offline trajectory · selected rank 1

StarCraft multi-agent micromanagement

SMAC · 3m

Good dataset · episode 31669

SMAC · 2s3z

Good dataset · episode 13583

SMAC · 5m vs 6m

Good dataset · episode 16943

SMAC · 8m

Good dataset · episode 23654

These videos visualize recorded offline dataset trajectories. SMAC movement, allied actions and health follow the recordings, rendered with custom unit artwork; this is not StarCraft client footage or an evaluation of the proposed method. Enemy attack poses are inferred from allied health changes, and a terminal display frame may be appended for recorded wins.

04 / Results

Look-ahead planning improves frozen policies.

MA-WAM is evaluated on 30 task and dataset combinations across MAMuJoCo, SMAC, and MPE. It achieves higher mean return than direct execution in 26 settings. The reported gains compare planning with direct execution under the paper’s fixed-denoising protocol.

Bold and underlined entries retain the paper’s markings. Scroll wide tables horizontally.

Main results against published baselines

Main results across MPE, SMAC, and MA-MuJoCo. Entries are episode returns; MPE uses the normalized score scale, whereas SMAC and MA-MuJoCo use native returns. MA-WAM (ours) reports mean ± standard deviation across seeds, while published baselines retain their source-paper uncertainty conventions. Baseline abbreviations follow the manuscript.

MPE

TaskQualityBCMA-ICQMA-TD3+BCMA-CQLOMARMADiffMA-SfBCDOM2MA-WAM (ours)
SpreadExpert35.0±2.6104.0±3.4108.3±3.998.2±5.2114.9±2.695.0±5.387.5±7.388.7±6.3118.3±1.5
Md-Replay10.0±3.813.6±5.715.4±5.631.4±7.237.9±6.130.3±2.58.2±4.663.1±9.561.6±3.0
Medium31.6±4.829.3±5.539.4±3.634.1±7.247.9±18.964.9±7.751.6±14.278.6±8.186.7±2.8
Random-0.5±3.26.3±3.59.8±4.924.0±9.834.4±5.36.9±3.15.1±3.937.4±11.366.2±3.2
TagExpert40.0±9.6113.0±14.4115.2±12.8119.3±14.0123.9±10.5103.0±12.077.4±13.998.2±14.4132.9±5.5
Md-Replay0.9±1.434.5±27.828.7±20.941.7±15.347.1±15.353.9±11.412.7±7.368.2±16.777.4±5.0
Medium22.5±1.863.3±20.065.1±29.561.7±23.166.7±23.272.7±9.447.1±17.982.6±18.2116.7±4.6
Random1.2±0.52.2±1.36.0±2.111.1±2.811.1±2.84.6±2.611.6±5.129.6±8.150.1±2.9
WorldExpert33.0±9.9109.5±22.8110.3±21.3119.8±28.1110.4±25.7109.3±15.497.3±19.199.5±17.1148.3±7.1
Md-Replay2.3±1.512.0±9.117.4±8.119.3±18.342.9±19.519.8±6.29.1±5.965.9±10.660.4±3.5
Medium25.3±2.071.9±20.073.4±9.358.6±11.274.6±11.584.7±12.354.2±22.784.5±23.4138.0±6.8
Random-2.4±0.51.0±3.22.8±5.50.6±2.05.9±5.26.1±2.43.1±1.34.1±1.15.0±1.7

SMAC

TaskQualityBCMA-ICQMA-CQLMADTMADiffDoFFlow BCMAC-FlowVGM2PMA-WAM (ours)
3mGood16.0±1.018.8±0.619.0±0.319.6±0.719.3±0.519.8±0.220.0±0.019.8±0.219.5±0.720.0±0.3
Medium8.2±0.818.1±0.718.9±0.717.2±0.716.4±2.618.6±1.214.7±1.518.0±3.216.9±1.116.0±0.8
Poor4.4±0.114.4±1.25.8±0.48.9±0.310.3±6.110.9±1.14.5±0.110.6±2.214.9±1.512.8±0.9
2s3zGood18.2±0.419.6±0.319.1±0.819.4±0.115.9±1.218.5±0.819.5±0.119.5±0.519.9±0.120.0±0.1
Medium14.3±0.717.2±0.814.3±2.017.4±0.315.6±0.318.1±0.915.1±2.017.6±0.616.5±0.617.7±0.5
Poor6.7±0.312.1±0.410.1±0.79.9±0.28.5±1.310.0±1.16.9±0.88.5±0.67.9±0.710.0±0.1
5m_vs_6mGood16.6±0.616.3±0.913.8±3.118.0±1.016.5±2.817.7±1.114.7±2.118.6±3.517.6±1.317.6±0.6
Medium14.2±0.517.2±0.416.8±3.117.5±0.415.2±2.616.2±0.912.8±0.815.6±1.317.0±0.918.5±0.5
Poor7.5±0.29.4±0.410.4±1.08.9±0.38.9±1.310.8±0.37.7±0.89.8±2.110.7±1.111.1±0.5
8mGood16.7±0.419.6±0.313.1±6.119.2±0.118.9±1.119.6±0.319.5±0.219.7±0.319.7±0.420.0±0.2
Medium10.7±0.518.6±0.516.3±3.118.0±0.516.8±1.618.6±0.818.2±0.819.4±0.618.2±1.619.6±0.2
Poor5.3±0.110.8±0.84.6±2.45.1±0.19.8±0.912.0±1.24.9±0.111.5±0.84.9±0.17.1±0.1

MA-MuJoCo

TaskQualityBCMA-TD3+BCMA-CQLMADTMADiffMA-SfBCDOM2MA-WAM (ours)
2×AntGood2697±2672922±194464±4692940±563105±471764±4572187±1902642±92
Medium1145±126744±283799±1861210±891241±301038±2951432±3051330±63
Poor954±801256±122857±73902±241037±32883±373916±1821038±35
4×AntGood2802±1332628±971344±6313090±263087±321722±3921836±2423025±23
Medium1617±1531843±494929±3491697±431897±441529±3721692±1831933±70
Poor1033±1221075±96518±1121268±511332±45976±2411158±2251358±39
Means and standard deviations across evaluation seeds. Planning improves 26 settings and reduces return in 4.
Planning gains across 30 settings.Means and standard deviations across evaluation seeds. Planning improves 26 settings and reduces return in 4.

The manuscript has not yet been publicly released.

Planning versus direct execution

Per-setting planning gain with three denoising steps per candidate. Reported scores are means under the common evaluation protocol; MPE uses normalized score, while SMAC and MA-MuJoCo use raw return. Δ is Planning minus Reactive, computed before display rounding, and bold marks a positive Δ. Planning improves the mean on 26 of 30 settings.

TaskQualityReactivePlanningΔTaskQualityReactivePlanningΔ
MPE (normalized score)SMAC (episode return)
SpreadExpert111.6118.2+6.63mGood19.019.5+0.5
Md-Replay34.855.7+20.9Medium13.413.8+0.4
Medium54.084.7+30.7Poor10.011.0+0.9
Random37.464.4+27.02s3zGood19.919.5-0.4
TagExpert130.8132.0+1.1Medium17.317.2-0.1
Md-Replay69.773.2+3.5Poor9.39.7+0.4
Medium109.7110.6+0.85m_vs_6mGood17.818.4+0.5
Random24.549.2+24.7Medium15.916.1+0.3
WorldExpert143.7144.4+0.7Poor10.710.9+0.2
Md-Replay45.560.2+14.78mGood19.620.0+0.4
Medium136.5132.4-4.1Medium17.919.1+1.2
Random1.44.5+3.1Poor6.57.0+0.4
MA-MuJoCo, 2×Ant (per-agent return)MA-MuJoCo, 4×Ant
2×AntGood23912406+154×AntGood27342988+254
Medium10331256+223Medium17371635-102
Poor775928+153Poor10981305+206

Aggregate gains by benchmark

Benchmark-level seed-averaged gain summary with three denoising steps per candidate. Mean Δ is in each benchmark's native unit. Aggregate relative gain is computed after summing reactive and planning scores within each benchmark family. Mean and median setting-wise gains first compute a relative gain for each task--quality setting and then aggregate those percentages within the benchmark family.

BenchmarkSettingsPositive/Zero/ Negative ΔMean ΔAgg. rel. gainMean setting-wise gainMedian setting-wise gain
MPE1211/0/1+10.81+14.42%+46.27%+19.16%
SMAC1210/0/2+0.41+2.75%+3.29%+2.70%
MA-MuJoCo65/0/1+125.00+7.68%+10.70%+14.04%

Equal-budget ablation: RWM versus random selection

Full 30-setting equal-candidate selector control with three denoising steps per candidate. Each entry is a mean under the common evaluation protocol. Both arms use M=8 candidates; Random selects uniformly, whereas RWM ranks candidates by predicted return. MPE uses OMAR normalized score; SMAC and MA-MuJoCo use native returns. Within each block, Δ is RWM minus Random, computed before display rounding; bold marks a positive difference. RWM exceeds Random in mean on 23 settings.

TaskQualityRandom M=8RWM M=8ΔTaskQualityRandom M=8RWM M=8Δ
MPE (normalized score)SMAC (episode return)
SpreadExpert111.9118.2+6.33mGood19.519.5+0.0
Md-Replay34.255.7+21.4Medium14.913.8-1.1
Medium51.584.7+33.2Poor8.911.0+2.1
Random37.164.4+27.42s3zGood19.819.5-0.3
TagExpert133.7132.0-1.8Medium16.717.2+0.5
Md-Replay68.573.2+4.7Poor9.19.7+0.6
Medium109.4110.6+1.15m_vs_6mGood17.718.4+0.7
Random31.149.2+18.1Medium16.916.1-0.8
WorldExpert147.8144.4-3.5Poor10.310.9+0.6
Md-Replay44.060.2+16.28mGood19.920.0+0.1
Medium130.9132.4+1.5Medium17.619.1+1.5
Random1.04.5+3.5Poor6.47.0+0.6
MA-MuJoCo, 2Ant (per-agent return)MA-MuJoCo, 4Ant (per-agent return)
2AntGood27412406-3354AntGood29062988+82
Medium9811256+275Medium17121635-78
Poor820928+108Poor11901305+115

Same-state candidate-ranking reliability

Same-state counterfactual candidate-ranking diagnostic on MPE and MAMuJoCo Medium tasks using 64 anchor states per task. RWM is compared with the analytical chance references for Top-1, Top-2, and Spearman rank correlation, and with the 0.5 chance reference for PairAcc (pairwise ranking accuracy). NormReg (normalized selection regret) compares RWM selection with empirical uniform selection from the same candidate sets.

SettingTop-1↑Top-2↑Spearman↑PairAcc↑NormReg↓
RWMChanceRWMChanceRWMChanceRWMChanceRWMRand.
MPE
Spread-Medium0.4220.1250.7030.2500.5840.0000.7370.5000.2080.555
Tag-Medium0.2500.1250.4220.2500.2400.0000.5900.5000.5060.534
World-Medium0.2660.1250.4380.2500.2540.0000.6020.5000.4550.572
MAMuJoCo
2Ant-Medium0.3120.1250.5000.2500.2900.0000.6090.5000.2970.490
4Ant-Medium0.1250.1250.3280.2500.2200.0000.5830.5000.4030.452
Mean0.2750.1250.4780.2500.3180.0000.6240.5000.3740.521

Inference cost across candidate budgets

Candidate-count profiling on an RTX 3090 and an A100. Results are averaged equally over MPE Spread-Medium, MAMuJoCo 2Ant-Medium, and SMAC 3m-Medium. Candidates are generated sequentially and scored in one batched RWM pass. We use four parallel environments, K=3, and a requested rollout cap H=8; the effective action horizons are 8/8/3 in the listed task order. Each entry is the mean of three repeats, each with 10 warm-up and 30 timed steps per setting. Total is Gen. plus Score.

RTX 3090A100
MGen. (ms)Score (ms)Total (ms)ShareGen. (ms)Score (ms)Total (ms)Share
155.910.166.015.3%64.411.175.514.8%
2106.312.8119.110.7%120.512.5132.99.4%
4209.013.3222.36.0%234.012.2246.24.9%
8414.613.4428.03.1%461.612.1473.72.5%
16827.913.3841.21.6%931.012.6943.61.3%

Batched versus sequential generation

Batched versus sequential candidate-generation latency for six Medium settings at M=8. We use the protocol of the corresponding table in the paper: four parallel environments, K=3, a requested rollout cap of eight, and three repeats with 10 warm-up and 30 timed steps per setting. Batch folds the M candidates into one generation pass; ``×'' is the within-row generation speedup. Times are milliseconds. ``Qual.'' is the number of task--quality settings sharing that architecture. The effective horizon for SMAC 3m is three. The A100 measurements were taken on a shared node.

RTX 3090A100
SettingAQual.Seq (ms)Batch (ms)×Seq (ms)Batch (ms)×
MPE Spread34424.673.25.8456.588.25.2
MPE Tag34430.473.55.9454.887.75.2
MPE World34429.073.55.8480.895.25.1
MAMuJoCo 2Ant23408.470.45.8469.977.96.0
MAMuJoCo 4Ant43455.195.74.8470.0111.14.2
SMAC 3m33425.371.16.0453.185.85.3

SMAC success rates

SMAC win rate; higher is better. Reactive uses M=1, whereas planning uses M=8 under the horizon protocol in Appendix (paper). Both use the fixed three-step denoising budget of the corresponding table in the paper. Planning improves 8 of 12 settings and ties 2 at the win-rate floor. Δ is Planning minus Reactive, and bold marks an improvement.

TaskQualityReactivePlanningΔTaskQualityReactivePlanningΔ
3mGood0.920.96+0.045m_vs_6mGood0.760.82+0.06
Medium0.460.50+0.04Medium0.560.60+0.04
Poor0.220.30+0.08Poor0.100.14+0.04
2s3zGood0.980.92-0.068mGood0.941.00+0.06
Medium0.600.56-0.04Medium0.740.88+0.14
Poor0.000.000Poor0.000.000

05 / Analysis

Understand when planning helps.

Planning benefits depend on both the quality of the candidate set and the reliability of the world model’s ranking. The diagnostics relate planning gains to reactive-policy headroom and ranking reliability; they do not imply that planning improves every setting.

Settings are grouped by reactive-policy headroom and ranking reliability. Each cell gives its mean relative gain and setting count.
When does planning help?Settings are grouped by reactive-policy headroom and ranking reliability. Each cell gives its mean relative gain and setting count.

Further experiments

Key ablations and diagnostics.

Alternative candidate scorers
Alternative candidate scorers. RWM, monolithic world model, current-action Q-ranker, direct-return predictor and Independent-WM on ten settings. Error bars are standard deviations across evaluation seeds; Independent-WM is a single-run reference.
Ranking reliability across planning horizons
Ranking reliability across planning horizons. Spearman correlation, pairwise accuracy and normalized regret across benchmark settings at horizons 1, 4 and 8. The dashed pairwise-accuracy line marks chance.
Routing across all 30 settings
Routing across all 30 settings. Routing diagnostics compare context and agent-identity association, expert utilization and switching across the full benchmark.