Skip to main content

Self-hosted AI / LLM ops

The runtime layer for self-hosted AI: model servers (Ollama for ease of use, vLLM for production throughput), gateways (LiteLLM as the unified API), chat UIs (Open WebUI, LibreChat, AnythingLLM), and agent / RAG platforms (Dify, Flowise).

Guides in this category cover Docker Compose deployment, GPU requirements (when needed), cost-vs-cloud-API breakeven, and integration patterns. For agent-style assistants that wrap these runtimes (OpenClaw, NemoClaw, Hermes), see Self-Hosted AI Assistants.

GPU workloads → Liquid Web GPU servers. CPU-only LLM gateways and chat UIs → Liquid Web Managed VPS. Pricing subject to change.

Articles in this section

Apify Affiliate Banner 728x90Apify Affiliate Banner 728x90Apify Affiliate Banner 300x50Apify Affiliate Banner 300x50