Mojo module
reducescatter
Multi-GPU reducescatter implementation for distributed tensor reduction across GPUs.
comptime valuesβ
elementwise_epilogue_typeβ
comptime elementwise_epilogue_type = def[dtype: DType, width: Int, *, alignment: Int, ?, .element_types.values: KGENParamList[CoordLike], .element_types`1: TypeList[values]](Coord[element_types], SIMD[dtype, width]) capturing -> None ``
Structsβ
- β
ReduceScatterConfig: Configuration for axis-aware reduce-scatter partitioning.
Functionsβ
- β
reducescatter: Per-device reducescatter operation with axis-aware scatter.
Was this page helpful?
Thank you! We'll create more content like this.
Thank you for helping us improve!