Lab 12: Serve Models from a KitOps ModelKit on HAMi
Package a model as a KitOps ModelKit, pull it from Jozu Hub with an initContainer, and serve it locally with SGLang (and optionally vLLM) on HAMi GPU shares.
Package a model as a KitOps ModelKit, pull it from Jozu Hub with an initContainer, and serve it locally with SGLang (and optionally vLLM) on HAMi GPU shares.
Install HAMi on a GPU cluster and schedule vLLM inference services with GPU partitioning.