Skip to main content
VELORAlaunch home

AI Agents & Infrastructure · United States

LLMKube

Kubernetes operator for self-hosted LLM inference

Visit llmkube.com0 votes. Log in to vote for LLMKube

Community votes
0
DR by Ahrefs
20 no change since last check

Domain Rating by Ahrefs, checked 3 Oct 2026, changed 3 Oct 2026.

Match report

Deploy and manage production LLMs on your own hardware with Kubernetes-native orchestration, supporting multiple inference runtimes and GPU accelerators.

Post-match questions

What runtimes does LLMKube support?

LLMKube is runtime-agnostic and supports seven pluggable backend options: vLLM for throughput, llama.cpp for efficiency, TGI (HuggingFace Text Generation Inference) for flexibility, mlx-server and vllm-swift for Apple Silicon, Ollama backend for Apple Silicon, and support for custom containers. Users can choose the best inference engine for their specific workload without changing the LLMKube CRD.

Is LLMKube really free and open source?

Yes, LLMKube is completely free and open source under the Apache 2.0 license. The project is built in the open since 2025, created and maintained by Defilan Technologies LLC in Washington State. Enterprise features are on the roadmap but currently everything is free.

What GPU hardware does LLMKube support?

LLMKube supports a wide range of GPU hardware: NVIDIA GPUs with CUDA 13 including Blackwell, AMD GPUs via Vulkan and ROCm backends, Intel Arc and Data Center GPUs via oneAPI/SYCL, and Apple Silicon including M1, M2, M3, M4, and M5 with the Metal Agent. The operator automatically handles hardware-specific scheduling and optimization.

Can I deploy models in seconds?

Yes, LLMKube enables fast deployment through declarative YAML that feels native to Kubernetes developers. The operator handles the complexity below the API layer. Users can deploy an LLM in seconds using simple kubectl commands or the llmkube CLI tool with a single command like llmkube deploy llama-3.1-8b --gpu --runtime vllm.

What is Foreman and how does it work?

Foreman is a Kubernetes-native control plane that dispatches coder, verifier, and reviewer agents across a heterogeneous fleet of local models running on your own hardware. The agents fix issues, open pull requests, and gate their own work. The honest-verdict harness makes agents ground every claim so a GO means the change was verified. Agents and human contributors review and merge work together, with all code staying on your infrastructure.

Same division

Alternatives to LLMKube
PosLaunchDivisionForm (DR)Community votes
1 Number 1: AirEnterprise Readiness AI Agents & Infrastructure DR 62 no change since last check 0 votes. Log in to vote for Air
2 Number 2: BlandEnterprise Voice AI Platform for Phone Agents AI Agents & Infrastructure DR 72 no change since last check 0 votes. Log in to vote for Bland
3 Number 3: SynthflowAI Voice Agent Platform to Automate Your Phone Calls AI Agents & Infrastructure DR 73 no change since last check 0 votes. Log in to vote for Synthflow
4 Number 4: Retell AIAI Voice Agent Platform for Phone Call Automation AI Agents & Infrastructure DR 77 no change since last check 0 votes. Log in to vote for Retell AI

Key features

  1. Pluggable runtime backends for vLLM, TGI, llama.cpp, and custom containers
  2. HPA autoscaling based on real inference metrics
  3. GPU layer offloading with custom sharding splits
  4. Automatic model download and persistent caching
  5. Prometheus and Grafana dashboards for inference metrics
  6. CUDA 13 with NVIDIA Blackwell GPU support
  7. Multi-GPU tensor parallelism and layer sharding
  8. Foreman agentic coding with GitHub integration

More on this profile

LLMKube has a free listing. The owner can unlock the full profile for $5:

Link metrics

Majestic link data for llmkube.com, change since the previous check.

Top topic: Computers/Programming/Languages (17). Checked 3 Oct 2026, changed 3 Oct 2026. Link data by Majestic. Trust Flow, Citation Flow and Topical Trust Flow are trademarks of Majestic-12 Ltd.

Need help?