Outcome Prediction Quiz: Hyperparameter Scaling
Instructions
Predict the outcomes of the following experiments. From the assignment restates, for reference, the assignment setting the question starts from; nothing in it is new. This question changes lists every difference from that setup. Everything not listed there is unchanged.
Check one answer in every row. The correct option earns full points, and an adjacent interval earns half. Each question has three rationale fields: the assignment problem you cite, what that experiment showed (with numbers), and why it implies your answer. Detailed proper citation, with experimental details and sound rationale, can return 75% of missed points on a question.
Default setup
The standard setup of the Hyperparameter Scaling assignment: Llama-style architecture with 8 layers,
hidden size 512, 8 attention heads, QK-norm, unscaled RoPE, a 4k vocab tokenizer, and about 35M
parameters (d8); fresh data with the token budget stated in each question; AdamW with
$\beta_1, \beta_2 = (0.9, 0.95)$ and $\epsilon = 10^{-8}$, linear decay with 1% linear warmup,
peak learning rate 0.003, context length 1024, batch size 64, weight decay 0.1 (masked), and 1.0 grad
clipping; model and data seed 42. Lower validation loss is better.
Conventions
- Fitted optimum: the vertex of the quadratic in $\log(\mathrm{LR})$ through the best sampled LR and its two grid neighbours.
- Transfer penalty: the loss at a transferred LR minus the fitted minimum loss of the target curve.
Your result
0 / 140
This score covers answer choices only. Rationale credit requires human evaluation.