Moving Terabytes Off Your File Server with Google Cloud Storage Transfer Service
gsutil rsync falls apart at multi-terabyte scale. Storage Transfer Service gives you incremental, scheduled, agent-based on-prem migrations to Cloud Storage — with honest trade-offs on bandwidth, metadata fidelity, and cost.
03 Aug 2026, 08:57 UTC

You have 40 TB of files on an on-premises NAS and a mandate to get them into Cloud Storage. The first instinct — a gsutil -m rsync from a jump box — works for a few hundred gigabytes, then falls apart: no resume across reboots, no scheduling, no audit trail, and one long-lived credential sitting on a machine you now have to babysit. The useful takeaway: Storage Transfer Service (STS) is Google's managed answer to exactly this problem, and for recurring or large migrations it removes most of the operational work you'd otherwise script yourself.
What STS actually does for on-premises moves
For on-premises sources, STS runs one or more transfer agents — Docker containers you run on machines near your data. The agents make outbound HTTPS connections to Google, pull the job definition, read your files, and stream them to the destination bucket. Because the connection is outbound-only, you don't need to punch inbound holes in your firewall or expose the file server.
The properties that matter for a real migration:
- Incremental by default. Re-running a job transfers only new or changed objects, so a multi-week initial sync followed by a final delta is a normal pattern, not a hack.
- Scheduling built in. Jobs can run once or recur (hourly, daily, weekly), which makes ongoing replication during a cutover period straightforward.
- Parallelism across agents. Running several agents on separate machines scales throughput horizontally; a single agent is limited by its host's CPU and NIC.
- Managed credentials. The destination writes happen under a Google-managed service account you grant
roles/storage.objectAdminon the target bucket — no service account key files on your NAS host.
A worked example: a scheduled sync job
Assume the Storage Transfer API is enabled and you've created a destination bucket. The flow, all runnable from Cloud Shell or any machine with the gcloud CLI authenticated:
1. Find the Google-managed service account and grant it write access to the destination bucket:
gcloud storage buckets add-iam-policy-binding gs://my-migration-bucket \
--member=serviceAccount:[contact removed] \
--role=roles/storage.objectAdminRetrieve the exact account email with gcloud transfer service-account get --project=YOUR_PROJECT. The numeric placeholder above is illustrative.
2. Create an agent pool and install agents. In the console (or via the REST API), create an agent pool, then run the agent container on a host that can read the source share, passing the pool name and credentials. Mount the source path read-only into the container — the agent only needs to read.
3. Create the job pointing at the agent pool's source path prefix and the destination bucket, with a daily schedule. In the console this is a guided form; the same job can be created via the API for repeatability.
4. Verify. Each run produces an operation record visible in the console and in Cloud Logging. Check three things after the first run: the job status is success (not "success with errors"), the copied-object count roughly matches a local count (find /mnt/share -type f | wc -l as a sanity check), and a spot-check of object checksums — STS records CRC32C/MD5 metadata you can compare against gsutil hash output on sample files.
Bandwidth, cost, and the trade-offs
The honest limitations:
- Your uplink is the bottleneck. 40 TB over a 1 Gbps connection takes roughly four days at full utilization — and you rarely get full utilization. Budget agent bandwidth (
--bandwidth-limiton the agent) so the migration doesn't starve production traffic, and consider whether a Cloud Interconnect or an offline transfer appliance makes more sense above ~100 TB. - Egress from your side isn't free. Your ISP or carrier costs apply, and while Google doesn't charge for ingress to Cloud Storage, STS itself has per-object operation costs for on-prem transfers. Estimate before starting, not after the bill arrives.
- File-system fidelity is limited. STS copies file content and basic metadata, but POSIX permissions, ACLs, symlinks, and hard links don't map cleanly onto object storage. If your workloads depend on them, plan a translation layer or reconsider whether object storage is the right target.
- Quotas exist. Very large jobs can hit project-level quotas on concurrent operations; check the Quotas page in the console and request increases ahead of the migration window rather than mid-run.
What to do this week
Run a pilot before committing: pick a 50 GB folder, set up one agent, run a one-time job, and verify object counts and checksums. Then re-run the same job to confirm the incremental pass copies nothing (or only files you changed). That two-hour exercise validates connectivity, IAM, and your verification method — and gives you a measured throughput number you can extrapolate into a realistic migration timeline. Only then schedule the full job, and keep the recurring sync running until the day you cut over.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.