Fine-tune LLMs on Serverless GPUs | Beam
Fine-tune models on GPUs
Fine-tune LLMs and diffusion models on serverless H100s without touching infrastructure. Decorate your training function, run the file, and Beam handles the GPUs, storage, and scaling.
Fine-tuning without wrangling GPUs
Beam runs the compute, just bring your code. LoRA, QLoRA, and more with zero infrastructure to manage.
GPUs on-demand
Grab a B300 the moment you kick off a run and release it the second training ends.
Runs like local Python
Decorate your training function and call .remote(). Beam syncs your working directory and streams logs back to your shell.
Volumes for weights & datasets
Persistent storage volumes hold model weights, datasets, and checkpoints between runs. Storage is included, free of charge.
Secrets for gated models
Store your Hugging Face token once and pull gated models like Llama and Gemma directly inside the container.
No timeouts
Run tasks as long as you need to.
Train and deploy, all in one platform
Serve your fine-tuned weights on a managed inference endpoint.
Launch a training job
01
Upload weights and data to a volume
Create a volume and download your base model and dataset.
02
Run training with one decorator
Wrap your training loop in @function, declare the GPU inline, and run the file. Beam streams logs to your terminal.
03
Serve the result
Point an endpoint at the fine-tuned weights in the same volume and deploy it as an autoscaling API.
# 1. Create a volume and stage weights + data
$ beam volume create gemma-ft
$ python upload.py # snapshot_download() into the volume
# 2. Kick off training on an H100
$ python finetune.py
=> Running function: <finetune:fine_tune>
{'loss': 1.617, 'epoch': 0.01}
...
=> Function complete
# 3. Inspect the artifacts
$ beam ls gemma-ft/gemma-2b-finetuned
adapter_config.json
adapter_model.safetensors
Frequently asked questions
Which fine-tuning methods work?
Anything that runs in PyTorch: LoRA and QLoRA with PEFT, Unsloth, TRL, or your own training loop. Beam provides the managed GPU runtime — the framework is up to you.
Do I pay when I'm not training?
No. For serverless, billing is per-second and stops when the run ends. Storage for weights and checkpoints is included, free of charge.
Where do my checkpoints go?
To a persistent storage volume mounted into the container. Inspect files with beam ls, download them, or serve them directly from an inference endpoint.
Can I fine-tune gated models like Llama or Gemma?
Yes. Save your Hugging Face token with beam secret create HF_TOKEN, and the container can pull gated repos securely.
What GPUs can I train on?
T4s and A10Gs up through B200s and B300s. Pick the GPU in the decorator and change it any time with one line.