{"id":2516,"date":"2026-08-31T20:12:45","date_gmt":"2026-08-31T20:12:45","guid":{"rendered":"https:\/\/convly.ai\/?p=2516"},"modified":"2026-08-31T20:12:45","modified_gmt":"2026-08-31T20:12:45","slug":"ollama-cloud-models","status":"publish","type":"post","link":"https:\/\/convly.ai\/de\/ollama-cloud-models\/","title":{"rendered":"Ollama-Cloud-Modelle: Ollama auf Cloud-Infrastruktur ausf\u00fchren"},"content":{"rendered":"<div class=\"convly-tldr\">\n<ul>\n<li>Ollama bietet keine cloudbasierten Modelle als API-Dienst an \u2013 es handelt sich um eine lokale Inferenz-Engine, die Sie auf Ihrer eigenen Hardware ausf\u00fchren.<\/li>\n<li>Sie k\u00f6nnen Ollama auf Cloud-VMs (AWS EC2, GCP Compute Engine, Azure) bereitstellen, um die Benutzerfreundlichkeit von Ollama mit der Skalierbarkeit der Cloud zu kombinieren.<\/li>\n<li>GPU-fokussierte Cloud-Anbieter wie Lambda Labs, Vast.ai und RunPod bieten kosteng\u00fcnstigere GPU-Instanzen als die gro\u00dfen Cloud-Anbieter f\u00fcr den Betrieb von Ollama.<\/li>\n<li>Die Kosten f\u00fcr Ollama in der Cloud liegen bei 0,50\u20133\u202f$\/Stunde f\u00fcr GPU-Instanzen, w\u00e4hrend kommerzielle APIs 0,50\u20135\u202f$ pro Million Tokens kosten \u2013 die Gewinnschwelle h\u00e4ngt vom Nutzungsvolumen ab.<\/li>\n<\/ul>\n<\/div>\n<p>Ollama stellt keine cloudbasierten Modelle als verwalteten API-Dienst zur Verf\u00fcgung. <a href=\"https:\/\/convly.ai\/de\/what-is-ollama-complete-guide-2026\/\">Ollama<\/a> ist eine lokale Inferenz-Engine, die Open-Source-Modelle auf Ihrer eigenen Hardware ausf\u00fchrt. Sie k\u00f6nnen Ollama jedoch auf Cloud-Infrastruktur \u2013 etwa AWS EC2, Google Cloud, Azure-VMs oder spezialisierten GPU-Anbietern \u2013 bereitstellen, um Rechenleistung aus der Cloud zu nutzen und gleichzeitig die einfache Benutzeroberfl\u00e4che von Ollama beizubehalten. Dieser Ansatz bietet Ihnen volle Kontrolle \u00fcber die Laufzeitumgebung und kann bei hohem Nutzungsvolumen kosteng\u00fcnstiger sein als kommerzielle APIs.<\/p>\n<div id=\"ez-toc-container\" class=\"ez-toc-v2_0_86 counter-flat ez-toc-counter ez-toc-container-direction\">\n<label for=\"ez-toc-cssicon-toggle-item-6a9604d095c8d\" class=\"ez-toc-cssicon-toggle-label\"><span class=\"\"><span class=\"eztoc-hide\" style=\"display:none;\">Umschalten<\/span><span class=\"ez-toc-icon-toggle-span\"><svg style=\"fill: #000000;color:#000000\" xmlns=\"http:\/\/www.w3.org\/2000\/svg\" class=\"list-377408\" width=\"20px\" height=\"20px\" viewbox=\"0 0 24 24\" fill=\"none\"><path d=\"M6 6H4v2h2V6zm14 0H8v2h12V6zM4 11h2v2H4v-2zm16 0H8v2h12v-2zM4 16h2v2H4v-2zm16 0H8v2h12v-2z\" fill=\"currentColor\"><\/path><\/svg><svg style=\"fill: #000000;color:#000000\" class=\"arrow-unsorted-368013\" xmlns=\"http:\/\/www.w3.org\/2000\/svg\" width=\"10px\" height=\"10px\" viewbox=\"0 0 24 24\" version=\"1.2\" baseprofile=\"tiny\"><path d=\"M18.2 9.3l-6.2-6.3-6.2 6.3c-.2.2-.3.4-.3.7s.1.5.3.7c.2.2.4.3.7.3h11c.3 0 .5-.1.7-.3.2-.2.3-.5.3-.7s-.1-.5-.3-.7zM5.8 14.7l6.2 6.3 6.2-6.3c.2-.2.3-.5.3-.7s-.1-.5-.3-.7c-.2-.2-.4-.3-.7-.3h-11c-.3 0-.5.1-.7.3-.2.2-.3.5-.3.7s.1.5.3.7z\"\/><\/svg><\/span><\/span><\/label><input type=\"checkbox\"  id=\"ez-toc-cssicon-toggle-item-6a9604d095c8d\"  aria-label=\"Umschalten\" \/><nav><ul class='ez-toc-list ez-toc-list-level-1' ><li class='ez-toc-page-1'><a class=\"ez-toc-link ez-toc-heading-1\" href=\"https:\/\/convly.ai\/de\/ollama-cloud-models\/#What_Ollama_Is_and_What_It_Isnt\" >Was Ollama ist (und was nicht)<\/a><\/li><li class='ez-toc-page-1'><a class=\"ez-toc-link ez-toc-heading-2\" href=\"https:\/\/convly.ai\/de\/ollama-cloud-models\/#Running_Ollama_on_Major_Cloud_Providers\" >Ollama auf gro\u00dfen Cloud-Anbietern betreiben<\/a><\/li><li class='ez-toc-page-1'><a class=\"ez-toc-link ez-toc-heading-3\" href=\"https:\/\/convly.ai\/de\/ollama-cloud-models\/#GPU-Focused_Cloud_Providers\" >GPU-fokussierte Cloud-Anbieter<\/a><\/li><li class='ez-toc-page-1'><a class=\"ez-toc-link ez-toc-heading-4\" href=\"https:\/\/convly.ai\/de\/ollama-cloud-models\/#Cost_Comparison_Cloud_Ollama_vs_API_Services\" >Kostenvergleich: Ollama in der Cloud vs. API-Dienste<\/a><\/li><li class='ez-toc-page-1'><a class=\"ez-toc-link ez-toc-heading-5\" href=\"https:\/\/convly.ai\/de\/ollama-cloud-models\/#Deployment_Patterns\" >Bereitstellungsmuster<\/a><\/li><li class='ez-toc-page-1'><a class=\"ez-toc-link ez-toc-heading-6\" href=\"https:\/\/convly.ai\/de\/ollama-cloud-models\/#Model_Selection_for_Cloud_Deployments\" >Modellauswahl f\u00fcr Cloud-Bereitstellungen<\/a><\/li><li class='ez-toc-page-1'><a class=\"ez-toc-link ez-toc-heading-7\" href=\"https:\/\/convly.ai\/de\/ollama-cloud-models\/#Performance_Optimization\" >Leistungsoptimierung<\/a><\/li><li class='ez-toc-page-1'><a class=\"ez-toc-link ez-toc-heading-8\" href=\"https:\/\/convly.ai\/de\/ollama-cloud-models\/#Frequently_Asked_Questions\" >H\u00e4ufig gestellte Fragen<\/a><\/li><\/ul><\/nav><\/div>\n<h2><span class=\"ez-toc-section\" id=\"What_Ollama_Is_and_What_It_Isnt\"><\/span>Was Ollama ist (und was nicht)<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>Ollama ist eine Open-Source-Inferenz-Engine, die gro\u00dfe Sprachmodelle lokal ausf\u00fchrt. Sie \u00fcbernimmt das Herunterladen von Modellen, deren Quantisierung sowie die Inferenz \u00fcber eine einfache CLI und eine REST-API. Gem\u00e4\u00df dem <a href=\"https:\/\/github.com\/ollama\/ollama\" rel=\"noopener\" target=\"_blank\">offiziellen Ollama-Repository<\/a>, it supports models from the Llama, Mistral, Gemma, Qwen, and DeepSeek families, among others.<\/p>\n<p>Ollama fungiert nicht als Cloud-Anbieter. Es hostet keine Modelle auf eigenen Servern und berechnet keine Preise pro Token. Wenn Sie <code>ollama run llama3.3<\/code>ausf\u00fchren, l\u00e4uft das Modell auf dem jeweiligen Ger\u00e4t, auf dem der Befehl ausgef\u00fchrt wird \u2013 Ihrem Laptop, einem Server im Rechenzentrum oder einer Cloud-VM, f\u00fcr die Sie separat bezahlen.<\/p>\n<p>The term &#8220;ollama cloud models&#8221; typically refers to one of three deployment patterns:<\/p>\n<ul>\n<li>Ausf\u00fchrung von Ollama auf einer GPU-beschleunigten Cloud-VM<\/li>\n<li>Bereitstellung von Ollama in einer Container-Orchestrierungsplattform (Kubernetes, ECS)<\/li>\n<li>Einsatz von Ollama auf einer dedizierten GPU-Cloud-Instanz f\u00fcr bedarfsgesteuerte Skalierung<\/li>\n<\/ul>\n<h2><span class=\"ez-toc-section\" id=\"Running_Ollama_on_Major_Cloud_Providers\"><\/span>Ollama auf gro\u00dfen Cloud-Anbietern betreiben<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<h3>AWS EC2-GPU-Instanzen<\/h3>\n<p>AWS bietet GPU-f\u00e4hige EC2-Instanzen in den Serien G, P und Inf. F\u00fcr Ollama-Workloads beginnt die g5.xlarge (1\u00d7 NVIDIA A10G, 24\u202fGB VRAM) bei etwa 1,01\u202f$\/Stunde On-Demand in us-east-1, w\u00e4hrend die g5.12xlarge (4\u00d7 A10G, 96\u202fGB VRAM) etwa 5,67\u202f$\/Stunde kostet.<\/p>\n<p>So stellen Sie Ollama auf EC2 bereit:<\/p>\n<pre><code># Ubuntu-22.04-Instanz vom Typ g5.xlarge mit Deep-Learning-AMI starten\n# Per SSH mit der Instanz verbinden\nsudo apt update\ncurl -fsSL https:\/\/ollama.com\/install.sh | sh\n\n# \u00dcberpr\u00fcfen, ob die GPU erkannt wird\nnvidia-smi\n\n# Ein Modell ausf\u00fchren\nollama run llama3.3:70b\n<\/code><\/pre>\n<p>A <a href=\"https:\/\/convly.ai\/de\/vram-requirements-every-major-llm-2026\/\">Llama 3.3 70B<\/a> model requires approximately 40 GB VRAM at 4-bit quantization, so you&#8217;d need a g5.12xlarge or larger. Llama 3.1 8B needs around 5 GB VRAM at 4-bit, fitting comfortably on a g5.xlarge.<\/p>\n<h3>Google Cloud Platform<\/h3>\n<p>GCP stellt GPU-Instanzen \u00fcber die Compute-Engine-Maschinenfamilien N1 und A2 bereit. Eine n1-standard-4 mit 1\u00d7 NVIDIA T4 (16\u202fGB VRAM) kostet etwa 0,62\u202f$\/Stunde, w\u00e4hrend eine a2-highgpu-1g mit 1\u00d7 A100 (40\u202fGB VRAM) in us-central1 etwa 3,67\u202f$\/Stunde kostet.<\/p>\n<pre><code># Instanz mit GPU erstellen\ngcloud compute instances create ollama-instance \n  --zone=us-central1-a \n  --machine-type=n1-standard-4 \n  --accelerator=type=nvidia-tesla-t4,count=1 \n  --image-family=ubuntu-2204-lts \n  --image-project=ubuntu-os-cloud \n  --maintenance-policy=TERMINATE\n\n# Per SSH verbinden und installieren\ngcloud compute ssh ollama-instance --zone=us-central1-a\ncurl -fsSL https:\/\/ollama.com\/install.sh | sh\nollama run mistral:7b\n<\/code><\/pre>\n<h3>Microsoft Azure<\/h3>\n<p>Azures NC-Serie-VMs bietet NVIDIA-GPUs f\u00fcr Inferenz-Workloads. Eine NC6s_v3 (1\u00d7 V100, 16\u202fGB VRAM) kostet etwa 3,06\u202f$\/Stunde, w\u00e4hrend die NC24ads_A100_v4 (1\u00d7 A100, 80\u202fGB VRAM) in der Region East US etwa 3,67\u202f$\/Stunde kostet.<\/p>\n<p>Installation follows the same pattern: provision a GPU-enabled Ubuntu VM, install NVIDIA drivers if not using a pre-configured image, and run the Ollama install script.<\/p>\n<h2><span class=\"ez-toc-section\" id=\"GPU-Focused_Cloud_Providers\"><\/span>GPU-fokussierte Cloud-Anbieter<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>Spezialisierte GPU-Cloud-Anbieter bieten f\u00fcr Ollama-Bereitstellungen h\u00e4ufig ein besseres Preis-Leistungs-Verh\u00e4ltnis als die gro\u00dfen Cloud-Anbieter:<\/p>\n<table>\n<thead>\n<tr>\n<th>Anbieter<\/th>\n<th>GPU<\/th>\n<th>VRAM<\/th>\n<th>Kosten\/Stunde<\/th>\n<th>Geeignet f\u00fcr<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Lambda Labs<\/td>\n<td>A100<\/td>\n<td>40 GB<\/td>\n<td>$1.10<\/td>\n<td>DeepSeek R1 Distill Llama 70B, Llama 3.3 70B<\/td>\n<\/tr>\n<tr>\n<td>Vast.ai<\/td>\n<td>RTX 4090<\/td>\n<td>24 GB<\/td>\n<td>$0.34-$0.54<\/td>\n<td>Mistral 7B, Llama 3.1 8B, Phi-4<\/td>\n<\/tr>\n<tr>\n<td>RunPod<\/td>\n<td>A40<\/td>\n<td>48 GB<\/td>\n<td>$0.79<\/td>\n<td>Mistral Large 3 (Single-Instanz), Llama 3.3 70B<\/td>\n<\/tr>\n<tr>\n<td>Paperspace<\/td>\n<td>A4000<\/td>\n<td>16 GB<\/td>\n<td>$0.76<\/td>\n<td>Kleinere Modelle bis ca. 14\u202fMilliarden Parameter<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>Lambda Labs bietet die einfachste Einrichtung \u2013 Instanzen werden mit vorinstallierten NVIDIA-Treibern und CUDA geliefert. Nach dem Start einer Instanz \u00fcber das Dashboard:<\/p>\n<pre><code>ssh ubuntu@\ncurl -fsSL https:\/\/ollama.com\/install.sh | sh\nollama run qwen3:32b\n<\/code><\/pre>\n<p>Vast.ai agiert als Marktplatz f\u00fcr ungenutzte GPU-Kapazit\u00e4t und bietet die niedrigsten Preise \u2013 allerdings mit wechselnder Verf\u00fcgbarkeit. Sie bieten Instanzen \u00fcber die Web-Oberfl\u00e4che entweder per Gebot oder als Miete an und installieren dann nach der SSH-Verbindung Ollama.<\/p>\n<h2><span class=\"ez-toc-section\" id=\"Cost_Comparison_Cloud_Ollama_vs_API_Services\"><\/span>Kostenvergleich: Ollama in der Cloud vs. API-Dienste<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>Ob eine cloudbasierte Ollama-Bereitstellung wirtschaftlich sinnvoll ist, h\u00e4ngt von Ihrem Nutzungsvolumen ab. Nutzen Sie den <a href=\"https:\/\/convly.ai\/de\/self-hosting-vs-api-calculator\/\">Selbsthosting vs. API-Rechner<\/a> um Ihren Break-Even-Punkt zu ermitteln.<\/p>\n<table>\n<thead>\n<tr>\n<th>Modell<\/th>\n<th>API-Kosten (pro 1 Mio. Tokens)<\/th>\n<th>Cloud-GPU<\/th>\n<th>Kosten pro Instanz pro Stunde<\/th>\n<th>Tokens pro Stunde bei der Break-Even-Menge<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Llama 3.3 70B<\/td>\n<td>0,10 USD Einzahlung \/ 0,32 USD Auszahlung<\/td>\n<td>Lambda A100 40 GB<\/td>\n<td>$1.10<\/td>\n<td>~3,4 Mio. Ausgabetokens<\/td>\n<\/tr>\n<tr>\n<td>Mistral Large 3<\/td>\n<td>2,00 USD Einzahlung \/ 6,00 USD Auszahlung<\/td>\n<td>RunPod 4\u00d7A40<\/td>\n<td>$3.16<\/td>\n<td>~530.000 Ausgabetokens<\/td>\n<\/tr>\n<tr>\n<td>Qwen3 32B<\/td>\n<td>0,08 USD Einzahlung \/ 0,28 USD Auszahlung<\/td>\n<td>Vast.ai RTX 4090<\/td>\n<td>$0.54<\/td>\n<td>~1,9 Mio. Ausgabetokens<\/td>\n<\/tr>\n<tr>\n<td>DeepSeek R1<\/td>\n<td>0,50 USD Einzahlung \/ 2,15 USD Auszahlung<\/td>\n<td>Lambda 4\u00d7A100<\/td>\n<td>$4.40<\/td>\n<td>~2 Mio. Ausgabetokens<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>If you&#8217;re processing more than the break-even token volume per hour, cloud Ollama becomes cheaper. For bursty workloads or low volumes, APIs like Claude Sonnet 5 ($2.00 in \/ $10.00 out per 1M tokens) or Gemini 3.6 Flash ($1.50 in \/ $7.50 out per 1M tokens) offer better economics with no idle time costs.<\/p>\n<h2><span class=\"ez-toc-section\" id=\"Deployment_Patterns\"><\/span>Bereitstellungsmuster<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<h3>On-Demand-Instanzen<\/h3>\n<p>Starten Sie eine GPU-Instanz nach Bedarf, f\u00fchren Sie Ihre Arbeitslast aus und beenden Sie sie anschlie\u00dfend. Dies eignet sich gut f\u00fcr Batch-Verarbeitung, Entwicklung oder gelegentliche Inferenz. Alle gro\u00dfen Cloud-Anbieter und GPU-Provider unterst\u00fctzen On-Demand-Preismodelle.<\/p>\n<h3>Spot-\/Pr\u00e4emptible-Instanzen<\/h3>\n<p>AWS Spot Instances, GCP Preemptible VMs und Azure Spot VMs bieten Rabatte von 60\u201390 %, k\u00f6nnen jedoch mit einer Vorwarnzeit von 30 Sekunden bis 2 Minuten beendet werden. Sie eignen sich f\u00fcr fehlertolerante Batch-Arbeitslasten, bei denen Fortschritte regelm\u00e4\u00dfig gesichert (\u201echeckpointed\u201c) werden k\u00f6nnen.<\/p>\n<p>Bei Vast.ai laufen unterbrechbare Instanzen zu Marktpreisen ohne Beendigungsgarantie \u2013 dies bietet die niedrigsten Preise, erfordert aber robuste Fehlerbehandlung.<\/p>\n<h3>Container-Orchestrierung<\/h3>\n<p>For production deployments, run Ollama in Kubernetes with GPU node pools. This enables auto-scaling, load balancing, and high availability. Use the official Ollama Docker image:<\/p>\n<pre><code>docker pull ollama\/ollama\ndocker run -d --gpus=all -v ollama:\/root\/.ollama -p 11434:11434 ollama\/ollama\n<\/code><\/pre>\n<p>Stellen Sie es dann in Ihrem Cluster bereit, wobei GPU-Ressourcenanforderungen in Ihrer Pod-Spezifikation definiert sind.<\/p>\n<h2><span class=\"ez-toc-section\" id=\"Model_Selection_for_Cloud_Deployments\"><\/span>Modellauswahl f\u00fcr Cloud-Bereitstellungen<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>W\u00e4hlen Sie Modelle entsprechend Ihrem GPU-VRAM-Budget aus. Nutzen Sie den <a href=\"https:\/\/convly.ai\/de\/llm-vram-calculator\/\">VRAM-Rechner<\/a> zur Absch\u00e4tzung des Ressourcenbedarfs:<\/p>\n<ul>\n<li><strong>16\u201324 GB VRAM<\/strong> (T4, RTX 4090, A10G): Mistral 7B (~4.5 GB at 4-bit), Llama 3.1 8B (~5 GB), Phi-4 (~9 GB), Qwen3 14B (~9 GB), Gemma 3 12B (~8 GB)<\/li>\n<li><strong>40\u201348 GB VRAM<\/strong> (A100 40 GB, A40): Llama 3.3 70B (~40 GB), DeepSeek R1 Distill Llama 70B (~40 GB), Mistral NeMo 12B mit Platz f\u00fcr gr\u00f6\u00dferen Kontext<\/li>\n<li><strong>80+ GB VRAM<\/strong> (A100 80 GB, H100): Llama 4 Scout (~65 GB), DeepSeek R1 (~400 GB \u2013 ben\u00f6tigt 4\u00d7A100 oder vergleichbar), Mistral Large 3 (~400 GB)<\/li>\n<\/ul>\n<p>Modelle mit einem VRAM-Bedarf von \u00fcber 80 GB erfordern Mehr-GPU-Konfigurationen oder Tensor-Parallelisierung \u2013 Funktionen, die Ollama ab Version 0.3 nicht nativ unterst\u00fctzt. F\u00fcr solche Modelle sollten Sie stattdessen Frameworks wie vLLM oder TGI oder kommerzielle APIs in Betracht ziehen.<\/p>\n<h2><span class=\"ez-toc-section\" id=\"Performance_Optimization\"><\/span>Leistungsoptimierung<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>Mehrere Einstellungen beeinflussen die Inferenzgeschwindigkeit von Ollama auf Cloud-GPUs:<\/p>\n<ul>\n<li><strong>Quantisierung<\/strong>Quantisierung: Ollama verwendet standardm\u00e4\u00dfig 4-Bit-Quantisierung (Q4_0). Verwenden Sie <code>ollama pull model:q8_0<\/code> f\u00fcr 8-Bit (bessere Qualit\u00e4t, doppelt so viel VRAM) oder <code>model:q2_K<\/code> f\u00fcr 2-Bit (schneller, geringere Qualit\u00e4t).<\/li>\n<li><strong>Kontextfenster<\/strong>Kontextl\u00e4nge: Festgelegt \u00fcber den <code>num_ctx<\/code> Parameter. Gr\u00f6\u00dfere Kontexte beanspruchen mehr VRAM \u2013 ein Kontext von 32 K Tokens ben\u00f6tigt deutlich mehr Speicher als ein 4-K-Context beim selben Modell.<\/li>\n<li><strong>Batch-Gr\u00f6\u00dfe<\/strong>: Erh\u00f6hen Sie den Wert von <code>num_batch<\/code> , um den Durchsatz bei latenzunempfindlichen Arbeitslasten zu steigern, bei denen Tokens pro Sekunde wichtiger ist als Reaktionszeit.<\/li>\n<li><strong>GPU-Schichten<\/strong>: Ollama l\u00e4dt standardm\u00e4\u00dfig alle Schichten auf die GPU. Bei Instanzen mit begrenztem VRAM wechselt es automatisch auf die CPU f\u00fcr Schichten, die nicht in den GPU-Speicher passen \u2013 was die Geschwindigkeit drastisch reduziert.<\/li>\n<\/ul>\n<h2><span class=\"ez-toc-section\" id=\"Frequently_Asked_Questions\"><\/span>H\u00e4ufig gestellte Fragen<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<h3>Bietet Ollama eigene Cloud-Hosting-Dienste an?<\/h3>\n<p>Nein. Ollama ist eine selbstgehostete Software, die Modelle auf Ihrer eigenen Infrastruktur ausf\u00fchrt. Es gibt keinen offiziellen Ollama-Cloud-Service oder eine verwaltete API. Wenn Personen von \u201eOllama-Cloud-Modellen\u201c sprechen, meinen sie damit, dass die Ollama-Software auf Cloud-Infrastruktur ausgef\u00fchrt wird, die sie selbst \u00fcber AWS, GCP, Azure oder GPU-Cloud-Anbieter bereitgestellt haben.<\/p>\n<h3>Kann ich das Ollama-API-Format mit Cloud-Modellen verwenden?<\/h3>\n<p>Ja. Sobald Ollama auf einer Cloud-VM l\u00e4uft, stellt es eine REST-API an Port 11434 bereit, die mit dem OpenAI-API-Format kompatibel ist. Sie k\u00f6nnen jeden OpenAI-kompatiblen Client auf die IP-Adresse und den Port Ihrer Cloud-Instanz ausrichten. Damit lassen sich Ollama-basierte Modelle in vielen Anwendungen als direkter Ersatz f\u00fcr kommerzielle APIs nutzen \u2013 allerdings tragen Sie selbst die Verantwortung f\u00fcr Verf\u00fcgbarkeit, Skalierung und Sicherheit.<\/p>\n<h3>Welcher Cloud-Anbieter ist am kosteng\u00fcnstigsten f\u00fcr den Betrieb von Ollama?<\/h3>\n<p>GPU-fokussierte Anbieter wie Lambda Labs (1,10\u202f$\/Stunde f\u00fcr A100 mit 40\u202fGB VRAM) und Vast.ai (0,34\u20130,54\u202f$\/Stunde f\u00fcr RTX 4090) liegen deutlich unter AWS, GCP und Azure bei vergleichbarer GPU-Speicherausstattung. Lambda bietet h\u00f6here Zuverl\u00e4ssigkeit und besseren Support; Vast.ai bietet die niedrigsten Preise, allerdings mit wechselnder Verf\u00fcgbarkeit. F\u00fcr Produktionsworkloads mit SLA-Anforderungen bieten die gro\u00dfen Cloud-Anbieter zwar bessere Garantien \u2013 allerdings zu h\u00f6heren Kosten.<\/p>\n<h3>Wie hoch sind die Kosten f\u00fcr den Betrieb von Llama 3.3 70B in der Cloud im Vergleich zu APIs?<\/h3>\n<p>Llama 3.3 70B ben\u00f6tigt bei 4-Bit-Quantisierung etwa 40 GB VRAM. Eine Lambda Labs A100-40-GB-Instanz kostet 1,10 USD pro Stunde. Wenn Sie pro Stunde 3,4 Millionen Ausgabetokens generieren (~945 Tokens\/Sekunde kontinuierlich), erreichen Sie die Break-Even-Menge gegen\u00fcber dem API-Preis von 0,32 USD pro Million Tokens. Darunter sind APIs g\u00fcnstiger; dar\u00fcber lohnt sich die Cloud-Instanz. Die meisten realen Arbeitslasten sind jedoch eher sprunghaft als kontinuierlich \u2013 daher bevorzugt die Preisgestaltung \u00fcber APIs die Mehrheit der Szenarien, es sei denn, Sie f\u00fchren kontinuierlich Batch-Verarbeitungsaufgaben aus.<\/p>\n<h3>Kann Ollama in der Cloud \u00fcber mehrere GPUs skaliert werden?<\/h3>\n<p>Ollama erkennt automatisch mehrere GPUs auf einer einzigen Instanz und nutzt sie auch \u2013 allerdings f\u00fchrt es jeweils ein eigenes Modell pro GPU aus, statt ein einzelnes Modell \u00fcber mehrere GPUs zu verteilen (Tensor-Parallelisierung). Das bedeutet: Eine 4\u00d7A100-Instanz kann vier separate Modellinstanzen ausf\u00fchren oder vier gleichzeitige Anfragen effizient verarbeiten, aber kein einzelnes 400-GB-Modell wie Mistral Large 3 laden, das die Kapazit\u00e4t einer einzelnen GPU \u00fcbersteigt. F\u00fcr Multi-GPU-Modellparallelisierung verwenden Sie stattdessen vLLM, TensorRT-LLM oder Text Generation Inference.<\/p>\n<h3>Ist der Betrieb von Ollama in der Cloud sicher?<\/h3>\n<p>Standardm\u00e4\u00dfig bindet Ollama an 0.0.0.0:11434 und macht seine API damit f\u00fcr beliebige Netzwerkclients zug\u00e4nglich. Auf Cloud-Instanzen mit \u00f6ffentlichen IPs bedeutet dies, dass Ihr Inferenz-Endpunkt ohne zus\u00e4tzliche Konfiguration von Firewallregeln, Sicherheitsgruppen oder VPN-Zugang \u00f6ffentlich erreichbar ist. Setzen Sie <code>OLLAMA_HOST=127.0.0.1:11434<\/code> , um den Zugriff auf localhost einzuschr\u00e4nken, und nutzen Sie dann SSH-Tunneling, einen Reverse-Proxy mit Authentifizierung oder ein VPN, um den Remote-Zugriff abzusichern. Cloud-Bereitstellungen sollten zudem Protokollierung eingehender Anfragen, Rate-Limiting und Eingabevalidierung implementieren, um Missbrauch zu verhindern.<\/p>","protected":false},"excerpt":{"rendered":"<p>Ollama doesn&#8217;t offer cloud-hosted models as an API service\u2014it&#8217;s a local inference engine you run on your own hardware. You [\u2026]<\/p>\n","protected":false},"author":1,"featured_media":2517,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"site-sidebar-layout":"default","site-content-layout":"","ast-site-content-layout":"default","site-content-style":"default","site-sidebar-style":"default","ast-global-header-display":"","ast-banner-title-visibility":"","ast-main-header-display":"","ast-hfb-above-header-display":"","ast-hfb-below-header-display":"","ast-hfb-mobile-header-display":"","site-post-title":"","ast-breadcrumbs-content":"","ast-featured-img":"","footer-sml-layout":"","ast-disable-related-posts":"","theme-transparent-header-meta":"","adv-header-id-meta":"","stick-header-meta":"","header-above-stick-meta":"","header-main-stick-meta":"","header-below-stick-meta":"","astra-migrate-meta-layouts":"default","ast-page-background-enabled":"default","ast-page-background-meta":{"desktop":{"background-color":"var(--ast-global-color-5)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"tablet":{"background-color":"","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"mobile":{"background-color":"","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""}},"ast-content-background-meta":{"desktop":{"background-color":"var(--ast-global-color-4)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"tablet":{"background-color":"var(--ast-global-color-4)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"mobile":{"background-color":"var(--ast-global-color-4)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""}},"footnotes":""},"categories":[7],"tags":[],"class_list":["post-2516","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-ai-news"],"_links":{"self":[{"href":"https:\/\/convly.ai\/de\/wp-json\/wp\/v2\/posts\/2516","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/convly.ai\/de\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/convly.ai\/de\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/convly.ai\/de\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/convly.ai\/de\/wp-json\/wp\/v2\/comments?post=2516"}],"version-history":[{"count":1,"href":"https:\/\/convly.ai\/de\/wp-json\/wp\/v2\/posts\/2516\/revisions"}],"predecessor-version":[{"id":2518,"href":"https:\/\/convly.ai\/de\/wp-json\/wp\/v2\/posts\/2516\/revisions\/2518"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/convly.ai\/de\/wp-json\/wp\/v2\/media\/2517"}],"wp:attachment":[{"href":"https:\/\/convly.ai\/de\/wp-json\/wp\/v2\/media?parent=2516"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/convly.ai\/de\/wp-json\/wp\/v2\/categories?post=2516"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/convly.ai\/de\/wp-json\/wp\/v2\/tags?post=2516"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}