GitLab CI/CD From Zero to Pipeline Architect: Part 1
What is a Pipeline? At its heart, a CI/CD pipeline is just a series of automated steps that you want to run on your code. It’s like a digital assembly line.
Imagine a parcel delivery system:
When you buy something online, it doesn’t instantly appear at your door. There’s a process:
Pick & Pack: The warehouse picks your item and packs it in a box.
Ship: The box gets loaded onto a truck or plane and sent toward your city.
Deliver: A delivery van brings the package to your doorstep.
Each step depends on the previous one. You can’t deliver a package that hasn’t been picked or packed yet.
A GitLab pipeline works the same way. When you push code, it triggers a series of tasks:
Build: Compile your code into an executable file.
Test: Run your automated tests to make sure nothing is broken.
Deploy: Send your application to a server where users can access it.
This automation saves you time, reduces errors, and lets you focus on what you do best: writing code.
The Building Blocks: Stages and Jobs
A pipeline is made up of two main concepts: Stages and Jobs.
Stages: These are the high-level phases of your pipeline, like “Build,” “Test,” or “Deploy.” They execute one after another. The “Test” stage won’t start until the “Build” stage has successfully completed.
Jobs: These are the specific tasks within a stage. A single stage can have multiple jobs that run in parallel. For example, in your “Test” stage, you might have one job that runs unit tests and another that runs linting checks at the same time.
The Unsung Hero: The Runner
So, you’ve defined your stages and jobs in a configuration file. Who actually does the work? That’s where the Runner comes in.
A GitLab Runner is a separate machine (a server, a virtual machine, or a Docker container) that is registered with your GitLab instance. It picks up the jobs you’ve defined and executes them.
Think of it this way: your pipeline configuration is the delivery instructions, and the Runner is the delivery truck that actually moves the parcels.
A GitLab Runner is a separate machine (a server, a VM, or something running in Kubernetes) that:
- Takes a job from your pipeline.
- Checks out your code.
- Starts a container with the image you asked for.
- Runs the script commands.
- Sends the logs and results back to GitLab.
You can think of it like this:
- Your pipeline file is the list of tasks.
- A runner is the worker that actually does those tasks.
If no runner is available (or tagged correctly), your jobs will just sit in “pending” forever.
One pipeline, many runners
In real projects, you might have different types of work:
- Node.js frontend builds.
- Go or Java backend builds.
- Docker image building.
- Heavy integration tests.
- Lightweight linting jobs.
You don’t have to run everything on the same kind of runner.
There are two common ways to mix and match:
1. Different images per job (simple, most common)
Even on the same runner, each job can use a different Docker image:
stages:
- build
- testbuild_frontend:
stage: build
image: node:20-alpine
script:
- npm ci
- npm run buildbuild_backend:
stage: build
image: golang:1.22-alpine
script:
- go test ./...
- go build -o app ./cmd/apptest_frontend:
stage: test
image: node:20-alpine
script:
- npm testHere:
- The same runner can still pick up all jobs.
- The runner will start a different container image for each job.
This is usually enough to support multiple languages in the same pipeline.
2. Different runners by tags (advanced, more control)
Runners can be tagged, for example:
- docker – normal Docker runner.
- k8s – runner in Kubernetes with more memory.
- windows – runner on Windows.
- large – high-CPU/high-RAM runner for heavy jobs.
Then in your .gitlab-ci.yml, you match jobs to runners using tags:
default:
tags:
- docker # default runner poolstages:
- build
- test
- deploybuild_app:
stage: build
tags:
- docker # use the default Docker runners
script:
- ./scripts/build.shintegration_tests:
stage: test
tags:
- large # run on a more powerful runner
script:
- ./scripts/run-integration-tests.shdeploy_prod:
stage: deploy
tags:
- k8s # runner that has kubectl/helm access
script:
- ./scripts/deploy-prod.shNow:
- build_app goes to any runner tagged docker.
- integration_tests only runs on runners tagged large.
- deploy_prod only runs on runners tagged k8s.
This is how larger teams:
- keep light jobs on cheap runners,
- send heavy jobs to powerful machines,
- restrict deployment jobs to secure runners that can reach production.
Different executor types (just so you know)
Under the hood, a runner is configured with an executor. The big ones you’ll hear about are:
- shell — runs directly on the machine (not isolated, simplest).
- docker — runs each job inside a Docker container (most common).
- docker+machine — auto-spawns new VMs and runs Docker on them (auto-scaling).
- kubernetes — starts each job as a pod in a Kubernetes cluster.
As a pipeline author, you mostly just pick an image: and/or set tags:
lint:
stage: validate
image: python:3.12-alpine # Docker executor will use this image
script:
- pip install ruff
- ruff check .You don’t have to worry which VM or cluster is actually running it; that’s the runner’s job.
Should you run a separate runner for every step?
You don’t need a dedicated runner for each step. Most teams do this:
- a shared pool of general-purpose runners for most jobs,
- a few specialized runners for:
- heavy builds (large),
- GPU tests (gpu),
- deployments (deploy).
You control who runs what with tags, not by hardcoding machines.
A nice rule of thumb:
- Start with one shared runner pool.
- When a certain job type is slow or needs special access (like prod clusters), create a new runner type, tag it, and update only those jobs.
Your First .gitlab-ci.yml
Everything we’ve talked about is controlled by a single file in the root of your repository called .gitlab-ci.yml. This file tells GitLab what to do when you push code.
Let’s look at a very simple example:
# Define the stages in the order they should run
stages:
- build
- test# A job in the 'build' stage
compile_code:
stage: build
script:
- echo "Compiling the code..."
# In a real project, this would be a command like 'npm run build' or 'mvn package'# A job in the 'test' stage
run_tests:
stage: test
script:
- echo "Running tests..."
# In a real project, this would be a command like 'npm test' or 'pytest'When you commit this file, GitLab will create a pipeline with two stages: build and test. The compile_code job will run first, and if it succeeds, the run_tests job will run next.
Leveling Up: Artifacts and Variables
Our first pipeline is a good start, but it has a problem: jobs are isolated from each other. The run_tests job runs on a fresh, clean environment and doesn't know anything about what the compile_code job did.
If your build job creates a file (like a compiled binary or a packaged application), the test job won’t be able to see it. To solve this, we use Artifacts.
Artifacts: Passing the Baton
Artifacts are files or directories that are created by a job and are saved by GitLab. They can then be passed to subsequent stages in the pipeline.
This image illustrates how an artifact works:
- The Build stage creates an application binary.
- This binary is stored as an artifact in GitLab.
- The Deploy stage then downloads this artifact and uses it.
Variables: No More Hardcoding
Another best practice is to avoid hardcoding values in your scripts. What if you want to deploy to a different server depending on the branch? Or what if you need to use a secret API key?
This is where CI/CD variables come in. You can define them in your project settings or in the .gitlab-ci.yml file itself. GitLab also provides many predefined variables, like CI_COMMIT_SHA (the commit hash) or CI_COMMIT_REF_NAME (the branch name).
Your First Serious Pipeline
Let’s put it all together. We’ll build a pipeline with a clean stage layout that uses artifacts and variables. We’ll also introduce two special stages: .pre (runs before everything) and .post (runs after everything), which are great for setup and cleanup tasks.
# A clean, standard stage layout
stages:
- .pre
- build
- test
- deploy
- .post# A global variable
variables:
APP_VERSION: "1.0.0-${CI_COMMIT_SHORT_SHA}"# This job runs first
setup_environment:
stage: .pre
script:
- echo "Setting up the environment for version ${APP_VERSION}..."# The build job creates an artifact
build_app:
stage: build
script:
- echo "Building the application..."
- mkdir -p build/
- echo "This is the compiled app binary for version ${APP_VERSION}" > build/app-binary
artifacts:
paths:
- build/# The test job uses the artifact from the build job
test_app:
stage: test
script:
- echo "Testing the application..."
- if [ -f build/app-binary ]; then echo "Artifact found!"; else echo "Artifact not found!"; exit 1; fi
- cat build/app-binary# The deploy job also uses the artifact
deploy_to_staging:
stage: deploy
script:
- echo "Deploying version ${APP_VERSION} to staging server..."
- cat build/app-binary
# In a real scenario, you'd use scp, rsync, or a cloud CLI to deploy the file# This job runs last
cleanup_environment:
stage: .post
script:
- echo "Cleaning up environment..."In this pipeline:
We define a clear order of stages.
- We use a variable APP_VERSION that includes the commit's short hash, giving every build a unique version number.
- The build_app job creates a file build/app-binary and saves the entire build/ directory as an artifact.
- The test_app and deploy_to_staging jobs automatically download this artifact so they can use the binary.
- The .pre and .post jobs run at the absolute beginning and end, regardless of the order you define them in the stages list.
Conclusion
Congratulations! You’ve just gone from knowing that “GitLab has pipelines” to building a structured, multi-stage pipeline that uses variables and passes data between jobs. You now have the foundational knowledge to start automating your own projects.
Comments
Questions, corrections, war stories — all welcome. Sign in with GitHub to join the discussion.