Getting Real Work Done on Colab's Free GPUs: Allocation, Verification, and the 12-Hour Clock
Colab's free GPU is one menu click away, but the real work is verifying what you got, checkpointing to Drive, and designing around the 12-hour session limit.
04 May 2026, 04:48 UTC

You opened a Colab notebook, ran your training cell, and watched an epoch crawl by on CPU. The fix is one menu click away — but the engineers who get burned on Colab aren't the ones who forgot to enable the GPU. They're the ones who assumed it stayed enabled, never checked what they actually got, and lost a trained model when the session died overnight.
The useful takeaway: treat Colab's GPU as a borrowed, time-boxed resource. Verify it before every long run, persist anything you care about to Drive, and design your training loop around the session limits rather than discovering them at hour eleven.
What you actually get (and how to check)
On the free tier, Colab typically assigns an NVIDIA Tesla T4 (16 GB VRAM) or an older K80 (12 GB), depending on availability. You don't choose — the backend assigns whatever is free. That matters because a script tuned to fill 16 GB will OOM on a K80.
Enable the GPU via Runtime → Change runtime type → Hardware accelerator → GPU. Then verify before training, because runtime type can silently reset when a session reconnects:
# Run in a notebook cell
!nvidia-smi
import torch
print(torch.cuda.is_available()) # expect True
print(torch.cuda.get_device_name(0)) # e.g. 'Tesla T4'nvidia-smi runs in the notebook's VM shell (no special permissions needed) and shows the GPU model, driver version, and current memory use. If torch.cuda.is_available() returns False after you selected GPU, the runtime didn't pick up the change — use Runtime → Restart runtime and re-check. TensorFlow users can run tf.config.list_physical_devices('GPU') for the equivalent check. Both frameworks pick up the GPU automatically; no CUDA installation is required in the standard Colab image.
Persist or lose it: mounting Drive
Everything under /content vanishes when the session ends. Mount Google Drive at the start of any notebook that trains for more than a few minutes:
from google.colab import drive
drive.mount('/content/drive')
CHECKPOINT_DIR = '/content/drive/MyDrive/my-project/checkpoints'The mount triggers an OAuth consent flow in the browser. Two practical notes: reading large datasets directly from Drive is slow (it's a network filesystem), so copy data to local disk first and only write checkpoints back. And checkpoint on a schedule — every N epochs — not just at the end.
A worked example: CIFAR-10 with a checkpointing loop
A small CNN on CIFAR-10 is a good sanity workload: on CPU an epoch can take many minutes, while on a T4 the same epoch typically finishes in seconds — the commonly cited figure is on the order of a 100x speedup for convolution-heavy training, though your exact ratio depends on the model and data pipeline. The point of the example isn't the benchmark; it's the structure:
import torch, time
device = 'cuda' if torch.cuda.is_available() else 'cpu'
model = build_cnn().to(device)
for epoch in range(EPOCHS):
t0 = time.time()
train_one_epoch(model, train_loader, device)
if epoch % 5 == 0:
torch.save(model.state_dict(),
f'{CHECKPOINT_DIR}/epoch_{epoch}.pt')
print(f'epoch {epoch}: {time.time()-t0:.1f}s')Run one epoch on CPU, note the time, then switch runtime type to GPU, restart, and re-run. That side-by-side comparison is the honest way to confirm acceleration on your own code rather than trusting a generic claim.
The trade-offs you design around
Three constraints shape real usage. First, free sessions cap out around 12 hours of continuous runtime — often less under heavy demand — and idle sessions disconnect sooner. Long training runs need resume-from-checkpoint logic, not just checkpointing. Second, you get one GPU per session; there is no multi-GPU training on the free tier. Third, VRAM is finite: if you hit CUDA out of memory, reduce batch size before reaching for gradient checkpointing or mixed precision, and watch actual usage with nvidia-smi in a separate cell.
Colab Pro and paid tiers soften these limits — longer sessions, priority access to newer GPUs — but the engineering habits are the same either way.
Before your next long run
Make this the first cell of every training notebook: check torch.cuda.is_available(), print the device name, mount Drive, and confirm your checkpoint directory exists. Thirty seconds of verification beats rediscovering at hour ten that you were training on CPU — or that your checkpoints went to a filesystem that no longer exists.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.