Qwen2.5 Coder 7B Q4_K_M (llama.cpp)
Code model for editors and agents, 4-bit GGUF on llama.cpp.
Needs 7.5 GiB VRAM · CUDA on NVIDIA
No GPU big enough is free right now.
New machines appear as hosts connect them.
No speed benchmarks yet, so we match on price.
Template details
- Rating
- No ratings yet
- Runs
- 0
- Min VRAM
- 8.1 GB
- Backend
- CUDA · NVIDIA GPUs
- Version
- llamacpp-qwen2-5-coder-7b-v1.0.0
- Image digest
- 1f4b9cf58982…
- Health check
- SSH probe on the leased endpoint
OpenAI-compatible API on the leased HTTP endpoint (per-job API key from the rental's access panel). Weights (4.7 GB) are warmed into the host cache and mounted read-only. Upstream image pinned by digest (ghcr.io/ggml-org/llama.cpp:server-cuda12-b11176@sha256:1f4b9cf58982dd4d7cc497aea31b1a456ca9a3a1f94f527d317d3fdee0d60ab6), verified via the registry API on 2026-09-26. Licence: MIT (https://github.com/ggml-org/llama.cpp/blob/master/LICENSE). Weights: Apache-2.0 (Qwen/Qwen2.5-Coder-7B-Instruct-GGUF) (https://huggingface.co/Qwen/Qwen2.5-Coder-7B-Instruct-GGUF). Needs an NVIDIA driver that supports CUDA 12.8 or newer. Status: not hardware-tested.
We match GPUs on their measured runtime and dedicated VRAM. A hardware report doesn't qualify another vendor's backend.
#llm #code #openai-api #gguf
Reported on real GPUs
People report these themselves. We haven't checked them.
No reports yet.
Sign in to report how this template runs on your GPU.
Reviews
No reviews yet.
Sign in to leave a review.