← All templates

gpt-oss-20b (llama.cpp)

Published, not hardware-testedllamacppgguf-mxfp4LLM inference

OpenAI's open-weight 20B reasoning model on llama.cpp's OpenAI-compatible server.

Needs 15 GiB VRAM · CUDA on NVIDIA

No GPU big enough is free right now.

New machines appear as hosts connect them.

No speed benchmarks yet, so we match on price.

Template details
Rating
No ratings yet
Runs
0
Min VRAM
16 GB
Backend
CUDA · NVIDIA GPUs
Version
llamacpp-gpt-oss-20b-v1.0.0
Image digest
1f4b9cf58982…
Health check
SSH probe on the leased endpoint

OpenAI-compatible API on the leased HTTP endpoint (per-job API key from the rental's access panel). Weights (GGUF MXFP4, 12.1 GB) are warmed into the host cache before the first rental and mounted read-only; no download at start. Upstream image pinned by digest (ghcr.io/ggml-org/llama.cpp:server-cuda12-b11176@sha256:1f4b9cf58982dd4d7cc497aea31b1a456ca9a3a1f94f527d317d3fdee0d60ab6), verified via the registry API on 2026-09-26. Licence: MIT (https://github.com/ggml-org/llama.cpp/blob/master/LICENSE). Weights: Apache-2.0 (openai/gpt-oss-20b) (https://huggingface.co/openai/gpt-oss-20b). Needs an NVIDIA driver that supports CUDA 12.8 or newer. Status: not hardware-tested.

We match GPUs on their measured runtime and dedicated VRAM. A hardware report doesn't qualify another vendor's backend.

#llm #reasoning #openai-api #gguf

Reported on real GPUs

People report these themselves. We haven't checked them.

No reports yet.

Sign in to report how this template runs on your GPU.

Reviews

No reviews yet.

Sign in to leave a review.