Finalized the MARL environment, set up the MAPPO training pipeline, and replanned the schedule after falling ~2 weeks behind.
Milestones
- Finalized
GGSwarmMarlEnv— multi-agent Crazyflie spawning, observation/action spaces, distance-based graph connectivity, and reward shaping - Set up MAPPO training pipeline in SKRL,
validated via
phase1_demo.py - Replanned Phase 2 — ~2 weeks behind the original Mar 3 target, but back on track to complete within days
Next Steps
- Complete Phase 2: Brain Development by training the GATv2 coordination policy
- Transition into Phase 3: Muscle Refinement — MINCO trajectory optimization and SwarmRaft decentralized consensus
Challenges
The sheer complexity of integrating Isaac Lab, PyTorch Geometric, SKRL, and MARL — all while learning them from scratch. Breaking through those integration hurdles has been tough but rewarding.