Independent research / Sept 2026
Self-play vs. text pretraining
Independent follow-up to Cowsik et al.'s self-play pretraining research, adding text-trained baselines and new in-context learning evaluations.
Follow-up to Self-Play Pretraining with Zero Data by Cowsik et al.
timelineSep 2026 - Sep 2026
focusresearch, ml
stack
PythonPyTorchDCLMMatplotlib
problem
How does in-context learning from self-play differ from ordinary text pretraining, and which abilities transfer when text training starts from a self-play model?
approach
Reused the authors' released self-play checkpoints and architecture, trained DCLM text baselines at matched learner-token budgets, and compared procedural and word-level tasks across model sizes and training checkpoints.
implementation notes
- Reproduced the original raw-byte evaluation and added printable-text and word-level task suites.
- Tested self-play initialization for text training and released the experiment harness, checkpoints, results, and plots.
- The original self-play method, architecture, and self-play checkpoints are from Cowsik et al.; the text baselines, extended evaluations, and follow-up analysis are my contributions.
impact
- In these experiments, self-play performed better on procedural tasks, while text pretraining performed better on word classification and lookup.
- Models up to 24M parameters, with 1–2 text seeds per size. Learner tokens are matched; total compute is not.