Generative AI is rapidly entering introductory programming, yet evidence about how learners coordinate with AI, especially in dyads, remains limited, and open datasets that support reproducible, trace-based evaluation are scarce. I present ASTRA (Adaptive Socially-intelligent Team Reasoning Agents), a multi-agent tutoring prototype and benchmark framework for studying collaborative programming with socially differentiated agents. ASTRA supports three configurations: alone_tutor (one learner with a Tutor agent), pair_tutor (two learners with a Tutor agent), and pair_multiagent (two learners with Tutor and Facilitator agents, where the Facilitator prompts coordination and balanced participation). As access to research participants is not yet available, I release an open synthetic benchmark dataset that mirrors ASTRA’s logging schema and a prespecified between-subjects design (𝑁 = 540 participants; 360 sessions; 1440 task episodes) across a bank of 20 short Python programming tasks. The dataset includes turn-level dialogue traces and task-level artefacts designed to support log-operational research questions about interaction dynamics, participation balance and reciprocal engagement in dyads, and performance and verification behaviours. Descriptive summaries and illustrative models indicate that the benchmark yields measurable condition-differentiated patterns consistent with the simulation assumptions. I emphasise that these findings are simulated evidence intended for benchmarking, measurement feasibility, and reproducible pipeline development, not causal estimates of learning effects, while providing a transparent analysis blueprint for future ethics-approved validation studies.
Fulltext license: CC BY