Loading the catalog…
Loading the catalog…
Summary The test loss of a language model follows a power law with respect to each of model size $N$, dataset size $D$, and training compute $C$, when the other two are not the bottleneck. In contrast, architecture shape such as depth, width, and the number of heads has almost no effect on loss as long as the non-embedding parameter count $N$ is the same. When $N$ and $D$ are scaled together, the degree of overfitting is determined by a single ratio, $N^{0.74}/D$. In other words, if the model is made 8 times larger, the data only needs to be increased by about 5 times. Under a fixed compute budget, the optimal strategy is to make the model much larger ($N \propto C^{0.73}$) and stop training well before convergence. 1. Introduction The performance of language modeling depends on the model architecture, the number of parameters, the amount of compute, the amount of available data, and so on. This paper targets Transformer language modeling and experimentally measures which of these factors are the main ones that actually determine performance. According to the results of several experiments conducted under varying conditions, performance scales as a power law with respect to training time, context length, dataset size, model size, and compute budget. Since the paper compares results across many conditions, it defines and uses several notations. $L$: Cross-entropy loss $N$: Number of model parameters $D$: Dataset size $C$: Total non-embedding training compute $C_{min}$: Estimate of minimum amount of non-embedding compute to reach a given loss 2. Background & Methods 2.1. Parameter and Compute Scaling of Transformers The hyperparameters of the Transformer architecture are set as follows. $n_{layer}$: number of layers $d_{model}$: dimension of the residual stream $d_{ff}$: intermediate dimension of the feed-forward layer $d_{attn}$: dimension of the attention output $n_{head}$: number of attention heads per layer 1 Model size $N$ A single layer consists of the following parameters. Q, K, V projections and output projection of attention: $4 \times d_{model} \times d_{attn}$ Two linear layers of the FFN: $2 \times d_{model} \times d_{ff}$ Therefore, the total number of parameters excluding embeddings and biases is $$ N \approx 2, d_{model}, n_{layer} \left(2 d_{attn} + d_{ff}\right) = 12, n_{layer}, d_{model}^2 \qquad (d_{attn} = d_{model} = d_{ff}/4) $$ 2 Compute The amount of computation (FLOPs) for the forward pass of a single token is $$ C_{forward} \approx 2N + 2, n_{layer}, n_{ctx}, d_{attn} $$ Here, the second term is a context-length-dependent term arising from the attention score computation, and when $d_{model} \gg n_{ctx}/12$ it becomes much smaller than $N$, so it can be ignored. In addition, since the compute of the backward pass is about twice that of the forward pass, the total training compute per token can be approximated as follows. $$ C \approx 6N $$ 3. Empirical Results and Basic Power Laws 3.1. Transformer Shape and Hyperparameter Independence To check the effect of the Transformer's shape on performance, the loss was measured while fixing the non-embedding parameter count $N$ and changing one of $n_{layer}$, $n_{head}$, and $d_{ff}$ at a time. The resulting loss is as follows. Even when the feed-forward ratio, aspect ratio ($d_{model}/n_{layer}$), and head dimension are varied over a wide range, the change in loss stays within a few percent, and even architectures whose aspect ratios differ by a factor of 40 achieve similar performance. In other words, once $N$ is fixed, the performance of the Transformer does not depend much on its shape. 3.2. Performance with Non-Embedding Parameter Count $N$ The value used to represent the size of a model is the number of parameters in the model. Here, parameters can be viewed from two main perspectives, depending on whether only the parameters trained for the actual objective function are counted, or whether embedding parameters are also included. Therefore, to express the relationship between parameter count and model performance more accurately, it must be decided which of the two values to use as the parameter count, and the following experiment was conducted as the verification process. Left: When embedding parameters are included in the count, the loss appears to depend not only on the number of parameters but also on $n_{layer}$. Right: When embedding parameters are excluded from the count, models with different depths converge onto a single line (power law). Only models with just one layer or with extreme depth/width ratios were exceptions. The reason it looks like the left plot is that the smaller the model, the larger the share of embedding in the total parameters. The parameter count looks large, but the part actually used for computation is small, so the loss comes out poor. For this reason, the paper uses the non-embedding parameter count as $N$ in all subsequent analyses. With this definition, the loss can be written as a function of $N$. $$ L(N) \approx \left(\frac{N_c}{N}\right)^{\alpha_N}, \qquad \alpha_N \approx 0.076,\quad N_c \approx 8.8 \times 10^{13} $$ 3.3. Performance with Dataset Size and Compute Dataset size $D$ To make the amount of data the bottleneck of the model, a large model ($n_{layer}=36$, $d_{model}=1280$) is trained on subsets of WebText2, and the lowest loss reachable with a given amount of data is measured. To compare the lowest losses, training is stopped once the test loss no longer decreases. The resulting relationship between loss and dataset size is as follows. $$ L(D) \approx \left(\frac{D_c}{D}\right)^{\alpha_D}, \qquad \alpha_D \approx 0.095,\quad D_c \approx 5.4 \times 10^{13} $$ Compute $C$ The training compute is $C = 6NBS$. With $C$ fixed, $N$ is varied to find the model that achieves the lowest loss within that compute. Connecting these optimal points gives the following relationship. $$ L(C) \approx \left(\frac{C_c}{C}\right)^{\alpha_C} $$ The values written as $X_c$ in each equation are obtained through fitting to express the relationship on a log-log graph. The relationships of loss with respect to parameters, dataset, and compute are summarized as follows. All three graphs are linear on log-log axes. The light blue curves in the left graph are the learning curves of models of different sizes, and their lower envelope (black line) forms the power law with respect to compute, $L = (C_{\min}/2.3 \cdot 10^8)^{-0.050}$. However, this power law holds only when the other two factors are not the bottleneck. 4. Charting the Infinite Data Limit and Overfitting Up to this point, only one variable was made the bottleneck at a time in order to check the independent effect of each variable. This section deals with how the loss behaves when the two values $N$ and $D$ are changed simultaneously, and in particular with the overfitting that occurs when data is insufficient. 4.1. Proposed $L(N, D)$ Equation $$ L(N, D) = \left[\left(\frac{N_c}{N}\right)^{\frac{\alpha_N}{\alpha_D}} + \frac{D_c}{D}\right]^{\alpha_D} $$ This loss as a function of $N$ and $D$ was chosen to satisfy the following principles. Rescaling : If the vocabulary size or tokenization changes, the overall loss is expected to change by a constant factor, so the equation should naturally allow for such rescaling. A limit imposed by the other variable even when one variable grows infinitely : If $D$ is fixed and $N \to \infty$, the loss should converge to $L(D)$, and if $N$ is fixed and $D \to \infty$, it should converge to $L(N)$. $L(N, D)$ is analytic at $D = \infty$ : It should be expandable as a series in integer powers of $1/D$. 4.2. Results Training was performed while varying $N$ and $D$, stopped once the test loss no longer decreased, and then the parameters of the above equation were fitted. Parameter $\alpha_N$ $\alpha_D$ $N_c$ $D_c$ Value 0.076 0.103 $6.4 \times 10^{13}$ $1.8 \times 10^{13}$ Overfitting How much is lost compared to the loss with infinite data, $L(N, \infty)$, is defined as $\delta L$. $$ \delta L \equiv \frac{L(N, D)}{L(N, \infty)} - 1 \approx \left(1 + \left(\frac{N}{N_c}\right)^{\frac{\alpha_N}{\alpha_D}} \frac{D_c}{D}\right)^{\alpha_D} - 1 $$ The larger $\delta L$ is, the more severe the overfitting. From the equation, $\delta L$ depends only on $N^{\alpha_N/\alpha_D}/D \approx N^{0.74}/D$ . Since the variation in loss due to the random seed is about 0.02, there is considered to be no overfitting when $\delta L < 0.02$. The condition that satisfies this is $$ D \gtrsim (5 \times 10^3), N^{0.74} $$ In other words, when increasing model size, overfitting can be avoided by increasing data only sub-linearly. In this way, not only can the loss equation be fitted with data $D$ and parameters $N$, but the relationship between $N$ and $D$ needed to prevent overfitting can also be identified. 5. Scaling Laws with Model Size and Training Time Now, the number of training steps $S$ is also taken into account, and the loss as a function of $N$ and training time is combined into a single equation. Before that, the critical batch size is introduced to correct for the step count and compute that vary with batch size. 5.1 Critical Batch Size $B_{crit}(L)$ Training time (number of steps) and compute vary with batch size $B$. $B \ll B_{crit}$: Increasing the batch reduces the number of steps almost proportionally. Compute efficiency is good. $B \gg B_{crit}$: Increasing the batch further barely reduces the number of steps. The number of steps is close to the minimum, but compute is wasted. The number of steps $S$ needed to reach a target loss and the amount of data processed $E = BS$ have the following relationship. Here, $E$ is the number of tokens actually processed, which differs from the dataset size $D$. $$ \left(\frac{S}{S_{\min}} - 1\right)\left(\frac{E}{E_{\min}} - 1\right) = 1 $$ $S_{\min}$: minimum number of steps needed to reach the target loss (when $B \to \infty$) $E_{\min}$: minimum amount of data needed to reach the target loss (wh
What RADAR observed and classified to build this opportunity. It is what the source published, not a verification that the offer is still active.
[Paper review] Scaling Laws for Neural Language Models. Summary The test loss of a language model follows a power law with respect to each of model size $N$, dataset size $D$, and training compute $C$, when the other two are not the bottleneck. In contrast, architecture shape such as depth, width, and the number of heads has almost no effect on loss as long as the non-embedding parameter count…
Open source