Mup
Machine Learning Model TrainingML Frameworks

Mup

maximal update parametrization (µP)

Open Source

About

In Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer, we show that optimal hyperparameters become stable across neural network sizes when we parametrize the model in maximal update parametrization (μP). This can be used to tune extremely large neural networks such as large pretrained transformers, as we have done in our work. More generally, μP reduces the fragility and uncertainty when transitioning from exploration to scaling up, which are not often talked about explicitly in the deep learning literature.

Open Source Health

Not enough history
Stars
1,755
Forks
105
License
MIT
Last commit
2 years ago
Jupyter Notebook

Related Categories