IMPORTANT: To view this page as Markdown, append `.md` to the URL (e.g. /get-started.md). For the complete documentation index, see llms.txt.
Skip to main content

Modular Documentation

The MAX framework accelerates AI inference and abstracts hardware complexity. Using our Docker container, you can deploy a GenAI model from Hugging Face with an OpenAI-compatible endpoint on a wide range of hardware.

And if you need to customize the model or tune a GPU kernel, MAX provides a depth of model extensibility and GPU programmability that you won't find anywhere else.

Cloud

Deploy MAX to your own hardware or use our fully managed cloud service.

Serving

High-performance, hardware-agnostic serving with OpenAI-compatible APIs.

Modeling

500+ models like DeepSeek, Gemma, and Kimi out of the box.

GPU programming

Extend or write custom GPU kernels that run on NVIDIA, AMD, and Apple GPUs.

Latest blog posts

Go to blog