Monday, 28 September 2026 | Updating Daily AI insight, written for builders

Ollama GPT OSS: Complete Guide to Running OpenAI’s Open Models

TL;DR

  • gpt-oss:20b runs in ~16GB memory (fits most gaming GPUs), gpt-oss:120b needs ~70GB (single 80GB GPU or split across consumer cards)
  • Install with ollama pull gpt-oss:20b or ollama pull gpt-oss:120b, then run with ollama run gpt-oss:20b
  • Both variants use MXFP4 quantisation at 4.25 bits per parameter and support 128K context windows
  • Released by OpenAI as open-weight models in partnership with Ollama, comparable to Llama 3.1 and Qwen 2.5 in quality

OpenAI released gpt-oss as open-weight models in August 2025, distributed exclusively through Ollama. The gpt-oss family consists of two variants: a 20-billion parameter model that runs comfortably on consumer hardware, and a 120-billion parameter model that delivers near-frontier performance on a single high-end GPU. Both use MXFP4 quantisation, packing model weights at 4.25 bits per parameter while maintaining quality close to full-precision inference.

What Are the GPT-OSS Models?

The gpt-oss models represent OpenAI’s first open-weight release, following industry pressure for transparency and reproducibility. Unlike GPT-4 or GPT-4o, these models ship with full weights, architecture details, and permissive licenses that allow commercial use without restriction.

Both models support a 128K token context window, handle multi-turn conversations, and perform instruction following comparable to other open models in their parameter class. The 20B variant competes directly with Llama 3.1 8B and Mistral 7B, while the 120B model sits between Llama 3.1 70B and full frontier systems like GPT-4.

OpenAI partnered with Ollama to handle distribution and inference optimization. This means you install and run gpt-oss the same way you would any other Ollama model, with no separate conversion or quantisation steps.

Hardware Requirements

The table below shows minimum and recommended hardware for each gpt-oss variant. Memory figures account for model weights plus a reasonable overhead for context and inference buffers.

Model Parameters Disk Size Minimum Memory Recommended Memory Example Hardware
gpt-oss:20b 20B 14GB 16GB 24GB RTX 4090, RTX 3090, A5000
gpt-oss:120b 120B 65GB 70GB 80GB A100 80GB, H100, 2× RTX 4090

The 20-billion parameter model fits comfortably on consumer GPUs released since 2020. If you have an RTX 3090 or 4090 with 24GB VRAM, you can run gpt-oss:20b with room for a 16K context window. With 16GB (RTX 4080, RTX 3080 Ti), you can still run the model but should limit context length to 8K tokens to avoid memory overflow.

The 120-billion parameter model requires datacenter hardware or a multi-GPU consumer setup. An A100 80GB handles it comfortably. Two RTX 4090s in a desktop can split the model across both cards, though inference speed drops slightly compared to single-GPU deployments. Use the VRAM calculator to estimate memory requirements for different context lengths.

CPU and System RAM

If you lack sufficient VRAM, Ollama will offload layers to system RAM and run them on the CPU. For gpt-oss:20b, you need at least 32GB system RAM if running entirely on CPU. For gpt-oss:120b, CPU inference is impractical—expect multiple seconds per token even on high-core-count systems. See the best GPUs for local LLMs guide if you’re building or upgrading hardware specifically for model inference.

Installation

First, install Ollama if you haven’t already. The process varies by platform:

macOS

Download the Ollama macOS app from ollama.com and drag it to /Applications. The installer sets up the ollama CLI automatically. Alternatively, install via Homebrew:

brew install ollama

Linux

Run the official install script, which places the binary at /usr/local/bin/ollama and sets up a systemd service:

curl -fsSL https://ollama.com/install.sh | sh

For manual installation or distro-specific instructions, see the Ollama installation guide.

Windows

Download the Windows installer from ollama.com and run it. The installer adds Ollama to your PATH and starts the background service automatically. After installation, open PowerShell or Command Prompt to verify:

ollama --version

Pulling and Running the Models

Once Ollama is installed, pull the model you want. For the 20-billion parameter variant:

ollama pull gpt-oss:20b

For the 120-billion parameter variant:

ollama pull gpt-oss:120b

Pulling downloads the model weights to ~/.ollama/models on macOS and Linux, or %USERPROFILE%.ollamamodels on Windows. The download is resumable, so you can interrupt and restart without losing progress.

After pulling, start an interactive session:

ollama run gpt-oss:20b

Or for the larger model:

ollama run gpt-oss:120b

This opens a chat interface. Type your prompt, press Enter, and the model streams its response. Type /bye to exit.

API Access

Ollama runs a local HTTP server at http://localhost:11434. You can send requests via curl, Python, or any HTTP client:

curl http://localhost:11434/api/generate -d '{
  "model": "gpt-oss:20b",
  "prompt": "Explain MXFP4 quantisation in one sentence."
}'

For structured applications, use the Ollama Python or JavaScript SDK. Both are available via standard package managers (pip install ollama, npm install ollama) and provide typed interfaces for generation, embedding, and model management.

Performance and Model Comparison

The 20-billion parameter model generates 15-30 tokens per second on an RTX 4090, depending on context length and prompt complexity. The 120-billion parameter model produces 8-15 tokens per second on an A100 80GB. Both figures assume default settings with no additional optimizations like speculative decoding.

Quality-wise, gpt-oss:20b performs comparably to Llama 3.1 8B and Mistral 7B on reasoning benchmarks like MMLU and GSM8K. It outperforms both on multi-turn instruction following, likely due to OpenAI’s reinforcement learning from human feedback (RLHF) tuning. The 120B variant approaches GPT-3.5 Turbo quality, making it the strongest open model under 200B parameters as of August 2026.

Compared to proprietary API-based models, running gpt-oss locally becomes cost-effective above approximately 2 million tokens per month for the 20B model, or 500K tokens per month for the 120B model, assuming you already own the hardware. Use the self-hosting vs API calculator to estimate your break-even point based on token volume and hardware costs.

MXFP4 Quantisation

Both gpt-oss models ship with MXFP4 quantisation baked in. MXFP4 is a microscaling floating-point format that represents weights at 4.25 bits per parameter on average, down from 16 bits in standard half-precision. Unlike older quantisation schemes like GPTQ or AWQ, MXFP4 preserves more dynamic range, reducing the accuracy loss typically associated with aggressive compression.

In practice, MXFP4-quantised models behave nearly identically to their full-precision counterparts on most tasks. The quantisation happens during training rather than post-hoc, which allows the model to adapt to the reduced precision. This approach reduces VRAM requirements by roughly 75% compared to FP16, making 120-billion parameter models viable on single consumer-grade GPUs.

Context Window and Memory Usage

Both gpt-oss variants support a 128K token context window, equivalent to roughly 100,000 words of English text. However, actually using the full context window requires significant additional memory beyond the base model weights.

For gpt-oss:20b, a 128K context adds approximately 8GB of memory overhead. With 24GB VRAM, you can comfortably use the full window. With 16GB, you should limit context to 32K-48K tokens to avoid running out of memory mid-generation.

For gpt-oss:120b, a 128K context adds roughly 20GB overhead. An 80GB GPU can handle this, but a 70GB deployment (two consumer GPUs) should cap context at 64K tokens. Check VRAM requirements by model for detailed memory scaling curves.

Frequently Asked Questions

Can I fine-tune gpt-oss models?

Yes. OpenAI released gpt-oss under a permissive license that allows fine-tuning for any purpose, including commercial applications. You can use standard fine-tuning frameworks like Axolotl or Hugging Face TRL. The model weights are compatible with the Llama architecture, so most Llama-tuned tooling works without modification.

How does gpt-oss:120b compare to Llama 3.1 70B?

gpt-oss:120b outperforms Llama 3.1 70B on instruction following and multi-turn conversations, particularly in tasks that require maintaining context across many exchanges. Llama 3.1 70B edges ahead on pure reasoning benchmarks like MATH and DROP. Both models are roughly equivalent on coding tasks. The 120B model requires more VRAM (70GB vs 40GB) but delivers a noticeable quality improvement in assistant-style interactions.

Can I run gpt-oss:120b on multiple consumer GPUs?

Yes. Ollama automatically splits the model across available GPUs if a single GPU lacks sufficient memory. Two RTX 4090s (48GB combined VRAM) can run gpt-oss:120b, though inference speed drops by 20-30% compared to a single 80GB GPU due to inter-GPU communication overhead. Three or more GPUs add diminishing returns—if you have the budget for that, a single A100 or H100 is more cost-effective.

Does gpt-oss work offline?

Yes, once you’ve pulled the model with ollama pull. Ollama stores the full model weights locally, and inference happens entirely on your machine. The only network access occurs during the initial download and when checking for model updates. You can disable automatic updates by setting OLLAMA_SKIP_UPGRADE=1 in your environment.

What license applies to outputs generated by gpt-oss?

OpenAI’s gpt-oss license does not claim ownership over generated outputs. You own the outputs you produce, and you may use them for any purpose, including commercial products and services. This differs from the terms of service for OpenAI’s API, which restrict certain use cases. Always review the full license text included with the model download to confirm the latest terms.

Can I use gpt-oss models commercially?

Yes. The license permits commercial use without restriction, including embedding the model in proprietary products, offering it as a service, or using it to generate content for sale. You are not required to open-source derivative works or disclose that you used gpt-oss in your product, though attribution is encouraged.

Written by Mustafa Ihsan

Mustafa Ihsan is the founder and editor of Convly.ai. He built and maintains the site's live AI models database, its price-performance index, and its free calculators for VRAM requirements, API costs and self-hosting economics. He writes about model pricing, benchmark results and the hardware needed to run AI models locally, and consistently prefers measured numbers to vendor claims.

Scroll to Top