KServe logo

OpenEverest for KServe

Run Model-as-a-Service on your own Kubernetes and GPUs. Deploy LLMs with vLLM in a few clicks and serve them behind one authenticated, OpenAI-compatible endpoint with per-model API keys, token quotas, and models that span several GPU nodes, powered by KServe.

Key capabilities

Authenticated AI Gateway

One HTTPS endpoint for all models. Each model gets its own API key, checked at the edge - a key for one model cannot call another.

OpenAI-Compatible

vLLM serves the standard /v1 API, so the OpenAI SDK, LangChain, and other OpenAI-compatible tools work without code changes.

Multi-Node Models

Tensor parallelism across GPUs in a node and pipeline parallelism across nodes, so models bigger than one server still run as one endpoint.

Token Quotas

Meter real prompt and completion tokens per API key, and enforce hourly budgets per model with a Redis or Valkey backend.

Presets & Model Catalog

Platform teams curate the models and ship known-good presets; users deploy with one click or a short Instance manifest.

Observability

vLLM metrics through a managed PodMonitor, per-key token usage from the Gateway, and optional OpenTelemetry tracing.

Your GPUs, Anywhere

Run on EKS, GKE, AKS, GPU clouds, or bare-metal clusters - your models, prompts, and data never leave your infrastructure.

Powered by

kserve/kserve

Models on OpenEverest are served by KServe, the open-source, Kubernetes-native platform for generative and predictive AI inference.

vllm-project/vllm

vLLM is the high-throughput inference engine that runs the LLMs, with tensor and pipeline parallelism across GPUs and nodes.

envoyproxy/ai-gateway

Envoy AI Gateway routes requests by model name, checks API keys, and meters tokens at the edge.

Frequently asked questions

Is KServe free to run on OpenEverest?
Yes. OpenEverest, KServe, vLLM, and Envoy AI Gateway are all open-source with no licensing fees, and everything runs on your own Kubernetes cluster and GPUs.
What is Model-as-a-Service?
Model-as-a-Service (MaaS) means offering models to your teams the way a cloud AI API does - pick a model, get an endpoint and an API key - but running on your own infrastructure. OpenEverest does for models what it already does for databases.
Is the API compatible with OpenAI?
Yes. Models are served through vLLM’s OpenAI-compatible API. Point the OpenAI SDK’s base_url at the endpoint from the connection details and use the API key as the key. Anthropic-style clients can send the key in the x-api-key header.
Which models can I serve?
Any model vLLM supports, from Hugging Face (hf://), S3 (s3://), Google Cloud Storage (gs://), or a PersistentVolumeClaim (pvc://). The UI offers a curated catalog, and gated models work once you provide a Hugging Face token. Predictive models (scikit-learn, XGBoost, PyTorch, TensorFlow, ONNX, Triton) are supported too.
How are models secured?
The shared Gateway terminates HTTPS with a certificate from cert-manager. Each model gets its own generated API key: requests without a valid key get 401, and a key for a different model gets 403. Rotating a key takes seconds.
Can I serve a model that does not fit on one GPU or one node?
Yes. Use tensor parallelism to split a model across GPUs in one node, and set a worker count with pipeline parallelism to split it across nodes via LeaderWorkerSet. Clients still see one endpoint and one model name. See Serve a model that does not fit on one node.
Do I need GPUs?
GPUs are recommended for real workloads. For development and smoke tests, a CPU compute profile runs small models such as SmolLM2 without a GPU.

Ready to get started?

Deploy OpenEverest on any Kubernetes cluster and serve your first model behind an authenticated, OpenAI-compatible API in minutes.

KServe, vLLM, and Envoy are trademarks of their respective owners. Model names are trademarks of their respective owners and are used for identification purposes only. OpenEverest is not affiliated with, endorsed by, or sponsored by these organizations.