Internet Access Whitelist Limits for Kaggle Notebooks
0 reputation · 14 Mar 2025, 03:38 UTC
0 reputation · 14 Mar 2025, 03:38 UTC
Users who rely on Kaggle Notebooks for reproducible research often need to call external APIs—for example, to fetch pretrained models, query public datasets, or interact with cloud services—while also wanting the session to survive the platform’s runtime limits. The notebook environment runs on Ubuntu 18.04 LTS and enforces a 9‑hour GPU or 12‑hour CPU maximum, after which the container is terminated and the /kaggle/working directory is cleared. However, outbound internet traffic is currently restricted to a whitelist of approved services, and the criteria for adding new endpoints remain under discussion, creating uncertainty about which external APIs can be relied upon without prior approval.
What factors does Kaggle evaluate when deciding to whitelist a new domain or API? Is there a public roadmap or timeline for expanding the whitelist to cover commonly used services? Can notebook authors request temporary access to specific external endpoints for a given session, and if so, what is the process?
29775 reputation · 14 Mar 2025, 11:01 UTC
Short answer: Kaggle approves whitelist additions based on security risk, abuse potential, terms-of-service compliance, and demonstrated community demand. There is no public roadmap or timeline for expansion, and there is no mechanism for temporary, per-session access — a domain must be permanently approved before any notebook can reach it.
Kaggle has not published a formal scoring rubric, but based on staff responses in Kaggle Meta discussions and support replies, the recurring factors are:
Treat this as the likely explanation rather than a confirmed checklist: the exact internal criteria are not publicly documented, so any claim of a definitive policy should be viewed with caution.
There is no published roadmap. Additions happen ad hoc after review, and community reports suggest requests can take anywhere from days to several weeks, with no guarantee of approval. Because of this, notebooks intended for reproducible research should be designed to work with currently approved services wherever possible, rather than assuming a pending request will succeed.
No such mechanism exists. The whitelist is global to the platform, not scoped to a user, notebook, or session. If your workflow depends on an unapproved endpoint, the practical workarounds are:
You can test from inside a notebook session:
curl -sS -m 10 -o /dev/null -w "%{http_code}\n" https://api.example.com/healthA timeout, Network is unreachable, or a connection reset generally indicates the domain is not whitelisted; an HTTP status code (even a 4xx) means network egress to that host works. Note that even whitelisted domains remain subject to rate limiting, and traffic patterns flagged as abusive can still be blocked.
Caveat: Kaggle's whitelist policy is not fully documented and may change; verify the current state with a live test and recent Kaggle Meta announcements before relying on any specific endpoint.
Use comments to ask for clarification. Post a solution as an answer.
No question comments on this page.