Breaking the Linear Bottleneck: Using GitLab DAGs to Speed Up Pipelines
Stop letting one slow build job block your entire test suite. Learn how to use GitLab's 'needs' keyword to implement Directed Acyclic Graphs (DAGs) for faster CI/CD execution.
30 May 2026, 20:11 UTC

The Stage Bottleneck Problem
In a standard GitLab CI/CD pipeline, jobs are organized into stages. By default, every job in the test stage must wait for every single job in the build stage to complete before any of them can start. This creates a "bottleneck" effect: if your build stage has ten fast jobs and one slow job that takes ten minutes, your entire testing suite is blocked for ten minutes, even if the specific tests only depend on the fast jobs.
The solution is to move from a linear pipeline to a Directed Acyclic Graph (DAG). By using the needs keyword, you can tell GitLab that a job should start as soon as its specific dependencies are finished, regardless of which stage those dependencies belong to.
How the 'needs' Keyword Works
The needs keyword allows you to define a job-level dependency. When you specify needs, you are explicitly telling the GitLab Runner: "Ignore the stage sequence; just wait for these specific jobs to succeed."
This decouples your pipeline. For example, if your linting job finishes in 30 seconds, a security-scan job in a later stage can start immediately, even while a heavy compile job in the first stage is still running.
Worked Example: Linear vs. DAG
Consider a scenario where you have a build process, a suite of tests, and a deployment to a staging environment. In a linear pipeline, the deployment must wait for every single test (unit, integration, and end-to-end) to pass.
In a DAG configuration, you can trigger the staging deployment as soon as the build and the critical integration tests are done, allowing end-to-end tests to continue running in the background.
# .gitlab-ci.yml
stages:
- build
- test
- deploy
build_app:
stage: build
script: echo "Building the application..."
unit_tests:
stage: test
script: echo "Running unit tests..."
integration_tests:
stage: test
script: echo "Running integration tests..."
e2e_tests:
stage: test
script: echo "Running slow end-to-end tests..."
deploy_staging:
stage: deploy
# This job starts as soon as build_app and integration_tests finish
# It does NOT wait for e2e_tests
needs: ["build_app", "integration_tests"]
script: echo "Deploying to staging..."
Execution Logic
- Linear Flow: build_app $\rightarrow$ (unit_tests + integration_tests + e2e_tests) $\rightarrow$ deploy_staging.
- DAG Flow: build_app $\rightarrow$ integration_tests $\rightarrow$ deploy_staging. (e2e_tests runs in parallel without blocking the deploy).
Implementation Risks and Constraints
While DAGs reduce total execution time, they introduce specific technical constraints that can lead to pipeline failures if ignored:
- Stage Order: A job cannot
needa job that is defined in a later stage. Dependencies must exist in the same stage or a previous one. - Circular Dependencies: If Job A needs Job B, and Job B needs Job A, GitLab will fail the pipeline validation immediately.
- Artifact Handling: By default, using
needsonly downloads artifacts from the jobs listed in theneedsarray. If your job requires artifacts from a job not listed inneeds, the pipeline will fail during the script execution phase.
Verifying the DAG Configuration
To verify that your DAG is functioning as intended, do not rely solely on the logs. Navigate to CI/CD > Pipelines in your GitLab project and click on the running pipeline. The Pipeline Graph UI will visually represent the dependencies. In a linear pipeline, you will see vertical columns of stages; in a DAG pipeline, you will see arrows skipping stages, indicating that jobs are triggering out of sequence.
Trade-offs in Pipeline Design
The primary trade-off is visibility vs. velocity. As you add more needs requirements, the pipeline graph becomes a "web" rather than a sequence. For small teams, this is a minor cost for faster feedback. For large organizations with hundreds of jobs, an overly complex DAG can make it difficult for new engineers to understand the actual deployment flow.
Start by identifying your slowest job in the early stages and determine which downstream jobs truly depend on it. Decouple the non-dependent jobs first to gain the most significant time savings without over-complicating the YAML configuration.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.