
vllm-mlx
High-performance OpenAI and Anthropic compatible LLM inference server for Apple Silicon.
free(무료 요금제 있음)1.6k github stars
vllm-mlx 소개
High-performance OpenAI and Anthropic compatible LLM inference server for Apple Silicon. Native MLX, continuous batching, multimodal models, MCP tool calling, and Claude Code support.