
Run Model-as-a-Service on your own Kubernetes and GPUs. Deploy LLMs with vLLM in a few clicks and serve them behind one authenticated, OpenAI-compatible endpoint with per-model API keys, token quotas, and models that span several GPU nodes, powered by KServe.
One HTTPS endpoint for all models. Each model gets its own API key, checked at the edge - a key for one model cannot call another.
vLLM serves the standard /v1 API, so the OpenAI SDK, LangChain, and other OpenAI-compatible tools work without code changes.
Tensor parallelism across GPUs in a node and pipeline parallelism across nodes, so models bigger than one server still run as one endpoint.
Meter real prompt and completion tokens per API key, and enforce hourly budgets per model with a Redis or Valkey backend.
Platform teams curate the models and ship known-good presets; users deploy with one click or a short Instance manifest.
vLLM metrics through a managed PodMonitor, per-key token usage from the Gateway, and optional OpenTelemetry tracing.
Run on EKS, GKE, AKS, GPU clouds, or bare-metal clusters - your models, prompts, and data never leave your infrastructure.
Models on OpenEverest are served by KServe, the open-source, Kubernetes-native platform for generative and predictive AI inference.
vLLM is the high-throughput inference engine that runs the LLMs, with tensor and pipeline parallelism across GPUs and nodes.
Envoy AI Gateway routes requests by model name, checks API keys, and meters tokens at the edge.
base_url at the endpoint from the connection details and use the API key as the key. Anthropic-style clients can send the key in the x-api-key header.hf://), S3 (s3://), Google Cloud Storage (gs://), or a PersistentVolumeClaim (pvc://). The UI offers a curated catalog, and gated models work once you provide a Hugging Face token. Predictive models (scikit-learn, XGBoost, PyTorch, TensorFlow, ONNX, Triton) are supported too.401, and a key for a different model gets 403. Rotating a key takes seconds.Deploy OpenEverest on any Kubernetes cluster and serve your first model behind an authenticated, OpenAI-compatible API in minutes.
KServe, vLLM, and Envoy are trademarks of their respective owners. Model names are trademarks of their respective owners and are used for identification purposes only. OpenEverest is not affiliated with, endorsed by, or sponsored by these organizations.
