Lab 12: Serve Models from a KitOps ModelKit on HAMi
Package a model as a KitOps ModelKit, pull it from Jozu Hub with an initContainer, and serve it locally with SGLang (and optionally vLLM) on HAMi GPU shares.
Package a model as a KitOps ModelKit, pull it from Jozu Hub with an initContainer, and serve it locally with SGLang (and optionally vLLM) on HAMi GPU shares.
在已有 GPU 集群上安装 HAMi,并用 GPU 切分能力调度 vLLM 推理服务。