Independent research / Sept 2026

Self-play vs. text pretraining

Independent follow-up to Cowsik et al.'s self-play pretraining research, adding text-trained baselines and new in-context learning evaluations.

Follow-up to Self-Play Pretraining with Zero Data by Cowsik et al.

timelineSep 2026 - Sep 2026
focusresearch, ml
stack
  • Python
  • PyTorch
  • DCLM
  • Matplotlib

problem

How does in-context learning from self-play differ from ordinary text pretraining, and which abilities transfer when text training starts from a self-play model?

approach

Reused the authors' released self-play checkpoints and architecture, trained DCLM text baselines at matched learner-token budgets, and compared procedural and word-level tasks across model sizes and training checkpoints.

implementation notes

  • Reproduced the original raw-byte evaluation and added printable-text and word-level task suites.
  • Tested self-play initialization for text training and released the experiment harness, checkpoints, results, and plots.
  • The original self-play method, architecture, and self-play checkpoints are from Cowsik et al.; the text baselines, extended evaluations, and follow-up analysis are my contributions.

impact

  • In these experiments, self-play performed better on procedural tasks, while text pretraining performed better on word classification and lookup.
  • Models up to 24M parameters, with 1–2 text seeds per size. Learner tokens are matched; total compute is not.

links