For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /max/get-started.md).
Mojo package
shmem
Implements a subset of OpenSHMEM for multi-node GPU communication.
This package abstracts over NVSHMEM (NVIDIA) and ROCSHMEM (AMD), exposing a
DeviceContext-compatible API backed by a symmetric heap that is accessible by
all GPUs in a job, both within a node and across nodes. It also includes
expert-parallelism (EP) communication kernels for MoE token dispatch and
combine operations.
Use SHMEMContext as a context manager to initialize the SHMEM runtime and
launch per-GPU threads, as shown below.
from std.testing import assert_equal
from shmem import shmem_my_pe, shmem_n_pes, shmem_p, SHMEMContext
def simple_shift_kernel(destination: UnsafePointer[Int32, _]):
var mype = shmem_my_pe()
var npes = shmem_n_pes()
var peer = (mype + 1) % npes
shmem_p(destination, mype, peer)
def main() raises:
with SHMEMContext() as ctx:
var destination = ctx.enqueue_create_buffer[DType.int32](1)
ctx.enqueue_function[simple_shift_kernel](
destination, grid_dim=1, block_dim=1
)
ctx.barrier_all()
var msg = Int32(0)
destination.enqueue_copy_to(UnsafePointer(to=msg))
ctx.synchronize()
print("PE:", ctx.my_pe(), "received message:", msg)
assert_equal(msg, (ctx.my_pe() + 1) % ctx.n_pes())Modulesβ
- β
ep: Helper functions for Expert Parallelism (EP) Communication Kernels. - β
ep_comm: Expert-parallelism (EP) communication kernels for MoE token dispatch and combine. - β
shmem_api: This is a work-in-progress implementation of the OpenSHMEM spec, in the future both ROCSHMEM and NVSHMEM will be supported. - β
shmem_buffer: Symmetric heap buffer type for OpenSHMEM-backed multi-GPU memory. - β
shmem_context: SHMEM runtime context and multi-GPU launch helper.
Was this page helpful?
Thank you! We'll create more content like this.
Thank you for helping us improve!