Herbák, Marcell and Kovásznai, Gergely and Adil, Ali Adil (2026) Optimizing Collaborative Learning: A Hard-Constraint Reinforcement Learning Approach to Fair Team Composition. In: Proceedings of the 13th International Conference on Applied Informatics. Líceum Kiadó, Eger, pp. 133-145. ISBN 9789634963271
|
Text
ICAI2026-pp133-145.pdf - Published Version Download (683kB) | Preview |
Abstract
Collaborative learning requires pedagogically balanced teams to function effectively. However, satisfying multiple strict constraints, such as specific gender ratios and minimized intra-group skill variance, transforms team composition into an NP-hard combinatorial optimization problem. Exact solving approaches, such as Optimization Modulo Theories (OMT), suffer from combinatorial explosion, taking hours to evaluate small cohorts (N = 40). We propose a Deep Reinforcement Learning (DRL) framework to resolve this bottleneck. Adapting the Long and Short-Term Constraints (LSTC) architecture, we enforce non-negotiable rules via dynamic action masking and cubic reward shaping. We utilize Deep Q-learning from Demonstrations (DQfD) via replay buffer pre-loading to initialize the policy with valid baselines. Empirical results show that our DRL agent achieves a significant computational speedup over the OMT-baseline while matching its engagement maximization. Furthermore, our approach eliminates the “failure clusters” prevalent in global optimization, improving worst-case team fairness. Robustness testing proves that the policy remains highly deterministic across randomized initializations.
| Item Type: | Book Section |
|---|---|
| Subjects: | Q Science / természettudomány > QA Mathematics / matematika > QA75 Electronic computers. Computer science / számítástechnika, számítógéptudomány |
| SWORD Depositor: | MTMT SWORD |
| Depositing User: | MTMT SWORD |
| Date Deposited: | 25 Sep 2026 12:29 |
| Last Modified: | 25 Sep 2026 12:29 |
| URI: | https://real.mtak.hu/id/eprint/247695 |
Actions (login required)
![]() |
View Item |




