g

gpustack

Developed by gpustack

GPUStack is an open-source GPU cluster manager for AI model serving and GPU instance provisioning. It automatically configures and orchestrates high-performance inference engines like vLLM, SGLang, and TensorRT-LLM, allowing on-demand SSH-accessible GPU instances for development, fine-tuning, and interactive workloads. GPUStack supports multi-cluster GPU management across on-premises, Kubernetes, and cloud environments, offering enterprise-grade operations like failure recovery, load balancing, monitoring, and authentication. It delivers Model-as-a-Service via OpenAI-compatible APIs, optimizes inference performance, and supports diverse accelerators including NVIDIA, AMD, and Huawei Ascend, maximizing resource utilization.

  • Unified Multi-Cluster GPU Management & Orchestration
  • Pluggable High-Performance AI Inference Engine Support
  • Day 0 Model Deployment & Performance Optimization
  • On-Demand SSH-Accessible GPU Instance Provisioning
  • Enterprise-Grade Operations & OpenAI-Compatible API
webdesktop