Expand description
Gradient compression algorithms for distributed training.
Implements Top-K sparsification (Stich et al. 2018), Random-K sparsification, 1-bit gradient quantization (Seide et al. 2014), and PowerSGD-style low-rank approximation.
Structs§
- OneBit
Quantizer - 1-bit gradient quantization (Seide et al. 2014).
- Power
SgdConfig - Configuration for PowerSGD low-rank gradient compression.
- RandomK
Compressor - Random-K gradient sparsification.
- TopK
Compressor - Top-K gradient sparsification compressor.
- TopK
Config - Configuration for Top-K gradient sparsification (Stich et al. 2018).
Functions§
- low_
rank_ compress - Compress a gradient matrix
G(m × n) into P (m × r) and Q (n × r) such thatG ≈ P · Q^T. - low_
rank_ decompress - Decompress a low-rank gradient approximation back to a full matrix.