All articles

Run DeepSeek V4 Flash on your own Lium GPUs

Run DeepSeek-V4-Flash-0731 on a Lium GPU node with vLLM, then connect to your private, OpenAI-compatible endpoint from your laptop.
Lium team5 min read
deepseek
vllm
llm-inference
h100
DeepSeek V4 Flash doesn't have to live behind someone else's API. Rent a GPU node on Lium, serve the open weights yourself, and connect your app to a private endpoint. You choose the exact model release and serving settings without buying GPUs.
In this guide, we'll rent an 8x H100 node on Lium, serve the official DeepSeek-V4-Flash-0731 release with vLLM, and send a chat request from a laptop. You can follow the same steps and adapt them for your own workload.

1. Pick the GPUs

We used the official deepseek-ai/DeepSeek-V4-Flash-0731 release, not the earlier preview checkpoint. It's MIT-licensed, so you can run it, modify it, and use it commercially. The weights alone need about 167 GB of GPU memory.
4 H100s (320 GB) can hold the weights, but they leave little room for the KV cache, the memory that holds each conversation's context. On 8 H100s (640 GB), our server had room for about 469,000 tokens of cache, enough for many long requests at once. We only tested 8 GPUs.
Lium bills by the second. Check the current H100 price before you launch; this example uses a whole 8-GPU node.

2. Rent the node

You need a Lium account with enough balance for a few hours on the node, and Python 3 on your laptop.
Install the CLI and sign in once. lium init opens the browser and saves your API key and SSH key. lium ls shows free nodes, cheapest first; lium up starts the default PyTorch template. --no-ssh returns you to your laptop's prompt rather than opening a shell automatically. The --ttl flag removes the pod after three hours, so a forgotten test doesn't keep billing.
bash
curl -fsSL https://lium.io/install.sh | bash   # or: pip install lium.io
lium init
lium ls --gpu H100 --count 8
lium up --gpu H100 --count 8 --name dsv4 --ttl 3h --no-ssh
Do not add port 8000 to the pod's --internal-ports. That would publish the model server on the pod's public address. We'll reach it through SSH instead.

3. Serve the model with vLLM

We use vLLM 0.26.0, the version tested with the commands below. The vLLM recipe for this model reports the same 8x H100 setup failing to start on 0.27.0. Installing it with uv --torch-backend=auto selects a PyTorch build that matches the pod's CUDA driver.
The serve command uses a few DeepSeek-specific settings:
  • --tokenizer-mode deepseek_v4 is required for /v1/chat/completions; the model doesn't ship a chat template.
  • --reasoning-parser deepseek_v4 separates reasoning from the answer in the response.
  • --tensor-parallel-size 8 --enable-expert-parallel spreads the model across the 8 GPUs.
  • --kv-cache-dtype fp8 --block-size 256 sets the cache format and block size the model expects.
We limit context to 64K to leave memory headroom. The request below uses DeepSeek's recommended temperature 1.0 and top_p 1.0.
The server will listen on 127.0.0.1 inside the pod, and you'll reach it through SSH rather than publishing port 8000. We also generate an API key for client requests. Keep the server private even with the key: not every vLLM route requires it. Copy the key below for step 4.
On your laptop, open a shell on the pod:
bash
lium ssh dsv4
Then, in the pod:
bash
python3 -m venv ~/vllm-env && . ~/vllm-env/bin/activate
pip install uv
uv pip install vllm==0.26.0 --torch-backend=auto
export VLLM_API_KEY=$(openssl rand -hex 16)
echo "$VLLM_API_KEY"   # copy this key
VLLM_USE_BREAKABLE_CUDAGRAPH=0 vllm serve deepseek-ai/DeepSeek-V4-Flash-0731 \
  --trust-remote-code --tensor-parallel-size 8 --enable-expert-parallel \
  --kv-cache-dtype fp8 --block-size 256 \
  --tokenizer-mode deepseek_v4 --reasoning-parser deepseek_v4 \
  --max-model-len 65536 --gpu-memory-utilization 0.85 \
  --host 127.0.0.1 --port 8000 --api-key "$VLLM_API_KEY"
On a fresh pod, the first start took about 12 minutes in our run: around four to download the weights and eight more for kernel compilation and warm-up. The server is ready when the log prints Application startup complete.
VLLM_USE_BREAKABLE_CUDAGRAPH=0 disables an experimental CUDA-graph mode; it is part of the configuration we tested on H100.

4. Call the private endpoint

Leave vLLM running in the pod. On your laptop, open another terminal for an SSH tunnel. It sends your laptop's 127.0.0.1:8000 to the same address inside the pod without opening that port to the internet. The $(...) part prints the pod's SSH command from lium describe, such as ssh root@203.0.113.7 -p 40123, and the flags after it turn that into a tunnel. It works in bash and zsh. Leave it running.
bash
$(lium describe dsv4 --json | python3 -c 'import json,sys; print(json.load(sys.stdin)["access"]["ssh_cmd"])') \
  -N -L 8000:127.0.0.1:8000
Despite its name, lium port-forward isn't this kind of tunnel. It relays TCP to a port published on the pod's public address, so it can't reach the loopback-only server we started.
In one more laptop terminal, paste the key from step 3 at the prompt and make a request. read -rs keeps the key out of your shell history. You can also use an OpenAI-compatible client with base URL http://127.0.0.1:8000/v1 and the same API key.
bash
read -rs VLLM_API_KEY && export VLLM_API_KEY   # paste the key from step 3; it stays out of shell history
curl http://127.0.0.1:8000/v1/chat/completions \
  -H "Authorization: Bearer $VLLM_API_KEY" \
  -H 'Content-Type: application/json' \
  -d '{"model": "deepseek-ai/DeepSeek-V4-Flash-0731", "temperature": 1.0, "top_p": 1.0,
       "messages": [{"role": "user", "content": "Explain MoE models in two sentences."}]}'
The model's answer began:
Mixture-of-Experts (MoE) models divide a neural network into multiple specialized sub-networks (experts) and use a learned router to activate only a few of them per input token, drastically cutting compute cost while keeping a huge total parameter count.
The request runs without thinking mode. The vLLM recipe shows how to turn it on with reasoning_effort.
The curl example is just a starting point. Point any OpenAI-compatible client at http://127.0.0.1:8000/v1 with the same key, and try prompts from your own workload.

5. Shut it down when you're done

When you're finished, remove the pod; the three-hour TTL will do it automatically if you forget.
bash
lium rm dsv4
That's the full loop: rent the GPUs, run the model you chose, use it through a private endpoint, and release the node when you're finished.
Run it on Lium
See H100 pricing and availability before you start.
Browse GPUs