vLLM workspace
SSH workspace with vLLM installed: serve any model you choose with `vllm serve`.
Needs 15 GiB VRAM · CUDA on NVIDIA
No GPU big enough is free right now.
New machines appear as hosts connect them.
No speed benchmarks yet, so we match on price.
Template details
- Rating
- No ratings yet
- Runs
- 0
- Min VRAM
- 16 GB
- Backend
- CUDA · NVIDIA GPUs
- Version
- vllm-workspace-v1.0.0
- Image digest
- d72482c035df…
- Health check
- SSH probe on the leased endpoint
SSH as `tenant` with your key; tools are on PATH. Models download into /work at your request (public egress is allowed; private networks are blocked). Reach a server on port 8000 with `ssh -L 8000:127.0.0.1:8000`. Thin AmpleRun wrapper (templates/images/workspace, variant ws-vllm) over vllm/vllm-openai:v0.30.0-cu129@sha256:a67f8f186d4567612ac37a55bd82295006002b18858eb12a8af2f05f86c2ae3f (digest verified 2026-09-26); NOT BUILT, so image_digest is a placeholder that can be neither published nor rented. Licence: Apache-2.0 (https://github.com/vllm-project/vllm/blob/main/LICENSE). Needs an NVIDIA driver that supports CUDA 12.9 or newer. Status: not hardware-tested.
We match GPUs on their measured runtime and dedicated VRAM. A hardware report doesn't qualify another vendor's backend.
#llm #vllm #workspace
Reported on real GPUs
People report these themselves. We haven't checked them.
No reports yet.
Sign in to report how this template runs on your GPU.
Reviews
No reviews yet.
Sign in to leave a review.