<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>Llm-Hosting on DevOpsTales</title>
    <link>https://devopstales.github.io/tags/llm-hosting/</link>
    <description>Recent content in Llm-Hosting on DevOpsTales</description>
    <generator>Hugo -- gohugo.io</generator>
    <language>en-US</language>
    <lastBuildDate>Mon, 14 Sep 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://devopstales.github.io/tags/llm-hosting/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>Ollama to vLLM: When Your Local Server Needs to Grow Up</title>
      <link>https://devopstales.github.io/ai/ollama-to-vllm-migration/</link>
      <pubDate>Mon, 14 Sep 2026 00:00:00 +0000</pubDate>
      
      <guid>https://devopstales.github.io/ai/ollama-to-vllm-migration/</guid>
      <description>Ollama is the easiest way to run a local model, until the experiment becomes a shared service. Requests queue, latency wobbles, prefixes recompute, and one GPU stops being enough. vLLM answers exactly those problems — but migration is a trade, not an upgrade: simplicity for scheduling, memory control, parallelism, and production operations.
</description>
      <enclosure url="https://devopstales.github.io/img/ollama-to-vllm-migration-icon.webp" length="3900" type="image/png" />
    </item>
    
  </channel>
</rss>
