Outcome Prediction Quiz: Basics
Instructions
Predict the outcomes of the following experiments. You will have 40 minutes for
8 questions (5 minutes each). Each question changes the standard d8
setup from the assignment by applying its diff to the student repository
(deep-learning-alchemy/assignments, branch a1_release,
commit baafc3f). Briefly justify each prediction by citing your
assignment's experiments. You will get partial credit (up to 75% of the value of
the problem) if your rationale and citations are sound.
Grading and expectations
Outcome prediction is very difficult, even for experts. Do not panic if you do not have confidence in getting 90+%. Our expectation is that strong students will maybe be able to attain 80% and most scores will hover around 50%.
Due to the difficulty of this task, there are two sources of partial credit. For many problems where exact numerical prediction is challenging, you will get partial credit if you are only off by one cell. Each question's scoring line states whether this applies and how many points it is worth. The form automatically grades answer choices using these rules. Rationale credit is not auto-graded online.
Default setup
The standard setup of the Basics assignment: Llama-style architecture with 8 layers,
hidden size 512, 8 attention heads, a 4k vocab size tokenizer, and about 35M parameters
(d8); 614.4M training tokens of fresh data (no epoching); AdamW with
β1, β2 = (0.9, 0.95), linear decay learning rate schedule with
1% linear warmup, learning rate 0.003, context length 1024, batch size 64, weight decay
0.1 (masked), and 1.0 grad clipping.
Your result
0 / 95
This score covers answer choices only. Review your written reasoning against each explanation below; rationale credit requires human evaluation.