# getlilac.com > AI-optimized mirror of getlilac.com containing 22 pages totalling 9,385 words of clean markdown content, structured data, and semantic HTML. Original source: https://getlilac.com/. Last updated: 2026-06-13T23:53:09.886Z. Each page is available as HTML (with JSON-LD structured data) and Markdown (text-only, ideal for LLMs and RAG). ## Homepage - [Fast inference, no markup.](/content/site-root.html): Hosted inference for MiniMax M2.7, Kimi K2.6, GLM 5.1, Gemma 4 (31B) and more. Industry-leading speed, lower cost — because we route to idle enterprise GPUs that are already powered on. (254 words) ## Articles & Blog Posts - [Privacy Policy](/content/privacy/index.html): Read Lilac’s privacy policy for the website, inference API, and related services. (906 words) - [Lilac is now self-serve — plus GLM 5.1 and Gemma 4 are live](/content/blog/lilac-self-serve-glm-gemma/index.html): No more waitlist. Sign up, grab an API key, and start running inference. GLM 5.1 is live at $0.90/M input, and Gemma 4 is live at $0.11/M input. (533 words) - [The GPU Scarcity Paradox](/content/blog/gpu-scarcity-paradox/index.html): The GPU shortage isn't what you think. The industry doesn't have a supply problem — it has a utilization problem masquerading as one. (1,398 words) - [We're partnering with MiniMax to bring M2.7 to Lilac](/content/blog/minimax-m2-7-partnership/index.html): We are partnering with MiniMax to bring commercially licensed MiniMax M2.7 access to Lilac. (550 words) - [How to keep frontier open weights viable](/content/blog/commercial-licensing-open-weight-models/index.html): Why Lilac supports open-weight licensing, and why commercial rights can help more frontier models stay open. (552 words) - [Kimi K2.6 API](/content/kimi-k2-6-api/index.html): Kimi K2.6 API at $0.70/M input, $3.50/M output, and $0.20/M cached input. OpenAI-compatible, no contracts. (257 words) - [GPU Inference API Pricing Compared](/content/blog/gpu-inference-api-pricing/index.html): A direct comparison of GPU inference API pricing across major providers. How idle GPU economics enable Lilac to offer lower per-token rates. (685 words) - [GLM 5.1 Inference Benchmark](/content/blog/glm-5-1-benchmark/index.html): We benchmarked our GLM 5.1 endpoint against every GLM 5.1 provider listed on OpenRouter. Competitive throughput at the lowest per-token price in the comparison. (422 words) - [Cache read pricing is now live on Lilac](/content/blog/cache-read-pricing/index.html): Supported Lilac models now show lower cache read rates for repeated context, making long-context and agent workloads cheaper to run. (419 words) - [Introducing Lilac: Turn Idle GPU Capacity into Revenue](/content/blog/introducing-lilac/index.html): Most Kubernetes clusters run GPUs at 30-50% utilization. We built a single operator to change that. (382 words) - [Make money from idle GPUs.](/content/providers/index.html): Turn idle Kubernetes GPU capacity into revenue. Install one operator, keep workload priority, and earn from spare enterprise GPU hours — your jobs always come first. (468 words) - [We are partnering with **MiniMax** to bring MiniMax M2.7 to Lilac.](/content/openrouter-alternative/index.html): Skip the aggregator markup. Lilac offers direct OpenAI-compatible endpoints — same models, ~5% less per token. (265 words) - [Gemma 4 API](/content/gemma-4-api/index.html): Gemma 4 API at $0.11/M input, $0.35/M output (BF16). OpenAI-compatible, no contracts. (252 words) - [Kimi K2.6 is live on Lilac](/content/blog/kimi-k2-6-live/index.html): Kimi K2.6 is now available on Lilac with OpenAI-compatible chat completions, 262K context, cache-read pricing, and no commitments. (266 words) - [How Idle GPUs Make Cheap Inference Possible](/content/blog/idle-gpu-inference/index.html): Lilac serves Kimi K2.6 inference on idle enterprise GPUs with OpenAI-compatible, pay-per-token shared endpoints. (390 words) - [MiniMax M2.7 API.](/content/minimax-m2-7-api/index.html): MiniMax M2.7 API pricing on Lilac: $0.30/M input, $0.055/M cached input, and $1.20/M output for agentic coding and tool-use workloads. (263 words) - [GLM 5.1 API](/content/glm-5-1-api/index.html): GLM 5.1 API at $0.90/M input, $3.00/M output (FP8). OpenAI-compatible, no contracts, 0.58s TTFT on shared warm endpoints. (250 words) - [Blog](/content/blog/index.html): Lilac updates, partnership announcements, cache pricing, pricing snapshots, and technical deep dives on GPU inference and idle GPU monetization. (317 words) - [Cheap inference API.](/content/cheap-inference-api/index.html): Cheap inference API with OpenAI-compatible endpoints, visible token pricing, and no contracts. Multiple open-weight models available. (234 words) - [Serverless inference, no cold starts.](/content/serverless-inference-api/index.html): Serverless inference API with OpenAI-compatible endpoints, no cold starts, and no contracts. (238 words) - [Main Article Headline](/content/sitemap-xml.html) (84 words) ## Resources - [Full Page Index](/index.html): Browse all cached pages with rich metadata - [About This Cache](/content/about.html): Methodology, technical details, and usage guidelines - [XML Sitemap](/sitemap.xml): Machine-readable sitemap for crawler discovery - [Robots.txt](/robots.txt): Crawler directives