Outcome Prediction Quiz: Hyperparameter Scaling

140 points ยท 7 questions

Instructions

Predict the outcomes of the following experiments. From the assignment restates, for reference, the assignment setting the question starts from; nothing in it is new. This question changes lists every difference from that setup. Everything not listed there is unchanged.

Check one answer in every row. The correct option earns full points, and an adjacent interval earns half. Each question has three rationale fields: the assignment problem you cite, what that experiment showed (with numbers), and why it implies your answer. Detailed proper citation, with experimental details and sound rationale, can return 75% of missed points on a question.

Default setup

The standard setup of the Hyperparameter Scaling assignment: Llama-style architecture with 8 layers, hidden size 512, 8 attention heads, QK-norm, unscaled RoPE, a 4k vocab tokenizer, and about 35M parameters (d8); fresh data with the token budget stated in each question; AdamW with $\beta_1, \beta_2 = (0.9, 0.95)$ and $\epsilon = 10^{-8}$, linear decay with 1% linear warmup, peak learning rate 0.003, context length 1024, batch size 64, weight decay 0.1 (masked), and 1.0 grad clipping; model and data seed 42. Lower validation loss is better.

Conventions

Submitting reveals every answer and explanation. Your work is graded locally in this browser and is not sent to the course staff.