Scaling Laws for Mixture Pretraining Under Data Constraints

Apple Research Blog
Read full post
Researchers analyzed over 2,000 language model training runs to understand how mixing scarce target data with abundant generic data affects performance. They found that repeated use of limited target data can improve results up to 15-20 repetitions, depending on model size and compute budget. They propose a scaling law to optimize data mixtures for pretraining under data constraints.

More in LLM & Text Generation

DeepSeek V4.1 Flash now available on AI Gateway

Covered by 2 sources

Cohere Debuts Open-Weight 218B Mixture-of-Experts Machine Translation Model

Unite.AI