Intermediate
Open
Pro
A 'Reproducible' Training Run That Isn't Fully Reproducible
A researcher sets torch.manual_seed(42) and np.random.seed(42) at the top
of their training script, trains a model twice on the same machine with the
same data and code, and gets two different final loss values (differing in
the third decimal place). They ask you why "the seed didn't work."
- List the additional sources of randomness/nondeterminism their script is probably missing, beyond the two seed calls shown.
- Give the specific settings/flags needed to close the gap, and explain the cost of doing so.
- Suppose they fix all of that, and then a colleague reruns the exact same seeded, deterministic-flagged script on a different GPU model. Should they expect bit-for-bit identical output? Justify your answer.
Share this question