Optimizing Deep‑Learning Workflows in Google Colab: GPU Selection, Memory Management, and Persistent Storage
Learn how to pick the right GPU, avoid memory fragmentation, and keep your experiments reproducible in Google Colab. A step‑by‑step guide with code, trade‑offs, and practical checks.
02 Dec 2025, 00:31 UTC

Why GPU Choice and Persistence Matter in Colab
When you build a deep‑learning model in a notebook, the two biggest bottlenecks are compute speed and data loss. Google Colab gives you free access to GPUs, but the type of GPU you attach and how you store data can dramatically change training time and experiment reliability. This post walks through the practical steps to choose the right GPU, keep an eye on memory, and persist your checkpoints with Google Drive.
1. Picking the Right GPU Runtime
Colab offers three GPU families for free users: Tesla K80, T4, and P4. Pro users can also access T4, P100, and V100. The GPU you select determines:
- CUDA compute capability (affects library support)
- Memory size (K80: 12 GB, T4: 16 GB, P4: 8 GB, P100: 16 GB, V100: 16 GB)
- Clock speed and tensor core availability
To attach a GPU, open the Runtime ▸ Change runtime type menu, set Hardware accelerator to GPU, and click Save. Verify the device with:
!nvidia-smi
Check the output for the GPU model and memory. If your model requires more than 12 GB, avoid the K80 and use a T4 or higher.
2. Managing Memory to Avoid Fragmentation
Even if your peak usage stays below the RAM limit, memory fragmentation can trigger OutOfMemoryError in TensorFlow or PyTorch. This happens when large tensors are allocated and freed in a non‑linear fashion, leaving gaps that the runtime can’t fill.
Strategies to mitigate fragmentation:
- Clear cache after each epoch (PyTorch:
torch.cuda.empty_cache()). - Use
tf.keras.backend.clear_session()in TensorFlow after training a model. - Pre‑allocate tensors of fixed size if you know the shape ahead of time.
Monitor memory with:
!cat /proc/meminfo | grep MemTotal
!cat /proc/meminfo | grep MemFree
These commands show total and free RAM in the Colab instance. Keep MemFree above 1 GB to avoid sudden OOM errors.
3. Persisting Data with Google Drive
Colab’s local filesystem (/content) is volatile; it wipes clean when the session ends. Mounting Google Drive attaches a persistent volume at /content/drive. Steps:
from google.colab import drive drive.mount('/content/drive')- When prompted, paste the authorization code from the displayed link.
- Verify persistence: create a file, disconnect, reconnect, and read it back.
%%writefile /content/drive/MyDrive/test.txt Hello, Colab! # Later in a new session !cat /content/drive/MyDrive/test.txt
Place checkpoints, datasets, and logs inside the Drive folder. This ensures reproducibility: the same data and model weights are available whenever you resume the notebook.
4. A Concrete Example: Training ResNet on CIFAR‑10 with Checkpoints
Below is a minimal notebook snippet that demonstrates GPU selection, memory cleanup, and Drive persistence.
# 1. Attach GPU (run in Runtime menu first)
!nvidia-smi
# 2. Mount Drive
from google.colab import drive
drive.mount('/content/drive')
# 3. Load dataset (stored on Drive)
import tensorflow as tf
(x_train, y_train), (x_test, y_test) = tf.keras.datasets.cifar10.load_data()
# 4. Build model
model = tf.keras.applications.ResNet50(weights=None, input_shape=(32,32,3), classes=10)
model.compile(optimizer='adam', loss='sparse_categorical_crossentropy', metrics=['accuracy'])
# 5. Train with checkpoint callback
checkpoint_path = '/content/drive/MyDrive/resnet_cifar10.ckpt'
cp_callback = tf.keras.callbacks.ModelCheckpoint(filepath=checkpoint_path,
save_weights_only=True,
monitor='val_accuracy',
mode='max',
save_best_only=True)
model.fit(x_train, y_train, epochs=10, validation_data=(x_test, y_test), callbacks=[cp_callback])
# 6. Clear session after training
import tensorflow.keras.backend as K
K.clear_session()
After training, the checkpoint file lives in Drive and can be re‑loaded in any future session:
model.load_weights('/content/drive/MyDrive/resnet_cifar10.ckpt')
Trade‑Offs and Limitations
- Session Timeout: Free runtimes are capped at 12 hours and may disconnect if idle or if the system is overloaded. Pro users get 24 hours, but both can be pre‑empted.
- Pre‑emption Risk: When the host GPU is busy, Colab can kill your process mid‑run. Always checkpoint frequently.
- Drive I/O Speed: Drive reads/writes are slower than local SSD. For large datasets, consider downloading to
/tmpat the start of the session. - Memory Fragmentation: Even with 16 GB GPUs, complex models can hit fragmentation. Use the cache‑clearing techniques described above.
Actionable Checklist
- Open Runtime ▸ Change runtime type and select GPU. Verify with
!nvidia-smi. - Mount Drive once per session:
drive.mount('/content/drive'). - Monitor memory with
!cat /proc/meminfo | grep MemFreeevery few minutes. - Enable
ModelCheckpointortorch.saveafter each epoch. - After training, clear the session with
K.clear_session()(TensorFlow) ortorch.cuda.empty_cache()(PyTorch). - When a disconnection occurs, re‑mount Drive, reload the checkpoint, and resume training.
Following these steps keeps your Colab experiments fast, reproducible, and resilient against the platform’s inherent session limits.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.