MULTI-AGENT REINFORCEMENT LEARNING
Combat Learning Lab
A visible self-play environment for testing whether embodied agents develop robust navigation and combat behaviour without reward exploitation.
A visible self-play environment for testing whether embodied agents develop robust navigation and combat behaviour without reward exploitation.
The question
Multi agent self play is notorious for producing agents that are excellent at exploiting the reward function and mediocre at the behaviour you actually wanted. The question this lab is built to test: in a visible, inspectable combat environment, do embodied agents trained via self play converge on genuinely robust navigation and engagement behaviour (flanking, positioning, retreat), or do they find degenerate shortcuts that score well against the specific opponents they trained against and fall apart against anything else?