← Back to Lesson 44

Quiz 44: Precision and Parallel Training

Score: 0 / 4

. Why does casting a very small gradient value (e.g. around `1e-11`) from fp32 to fp16 cause training to fail, rather than just losing some precision?

Below fp16's representable floor, underflow means total information loss (exactly 0.0), not just reduced precision — and gradients in deep sigmoid networks routinely fall in that underflowing range.

. What are the two safeguards mixed-precision training uses to keep the speed of fp16 without hitting the underflow problem?

fp32 master weights avoid updates being rounded away, and loss scaling exploits gradients being linear in the loss to push otherwise-underflowing values back into fp16's representable range.

. In the data-parallelism demonstration, splitting a batch into 4 shards, computing gradients independently per shard, and averaging them afterward produces a result that:

This is the whole correctness argument for data parallelism: per-shard gradients averaged after the fact equal the gradient computed over the whole batch at once, up to ~1e-8 roundoff.

. What is the key difference between what limits data parallelism versus what limits (pipeline) model parallelism at scale?

Data parallelism's cost is the all-reduce communication step; model (pipeline) parallelism's cost is devices sitting idle unless multiple microbatches keep every stage busy.