Playbook · 5 minute read
How to Build a GitLab AI Agent
A GitLab AI agent reviews merge requests by reading the diff alongside pipeline results, diagnoses failed pipelines from job logs, triages and links issues, and drafts changes as merge requests under protected branch and approval rules that prevent it from merging its own work. Project or group access tokens scoped to the minimum role bound what it can do.
GitLab's advantage for an agent is that everything is in one place: the issue that asked for a change, the merge request that implements it, the pipeline that tested it, and the environment it deployed to. An agent can follow that thread end to end, which lets it review a change in context rather than as an isolated diff. This guide covers building one that helps developers without gaining the ability to merge its own work, drawing on FISTA Solutions' AI agents delivery in engineering platforms. It complements how to build a github ai agent and how to build a code review agent.
How should the agent authenticate?
With a project or group access token attached to a bot identity, at the lowest role that permits what it needs. Developer allows commenting, pushing to non-protected branches, and opening merge requests, which covers most agent functions. Maintainer and Owner allow changing protected branches and settings and should not be granted.
| Token type | Identity | Fits |
|---|---|---|
| Project access token, Developer | Bot per project | Single-project agents |
| Group access token, Developer | Bot across projects | Organisation-wide agents |
| Personal access token | An individual | Not for agents |
| Deploy token | Read repository and registry | Read-only analysis |
Webhooks on merge request, pipeline, and issue events trigger the agent, so it responds to activity rather than polling.
What makes merge request review useful?
Context. A review that reads only the diff comments on what a linter already checks. A review that reads the diff alongside the linked issue, the pipeline results, the project's conventions inferred from existing code, and the files the change touches can address the questions that matter: does this do what the issue asked, why did that test fail, does this fit how the codebase does things, and does it touch anything sensitive.
Review comments should be few and specific, attached to lines, with the reasoning stated. An agent that leaves twenty comments per merge request is ignored by the second week. See how to build a code review agent.
How does pipeline failure diagnosis work?
A failed pipeline produces a job log of thousands of lines, and the developer's first task is finding the actual error. The agent reads the failed job's log, identifies the error among the build noise, correlates it with the diff to say whether the change caused it, checks whether the same test has failed intermittently on other merge requests to flag flakiness, and comments on the merge request with the diagnosis and a suggested next step.
That is the fastest win for developer time in most GitLab projects, because pipeline failures happen constantly and reading logs is tedious. Distinguishing flaky tests from real failures, using pipeline history, is the part developers value most.
How is the agent prevented from merging?
Through GitLab's own controls, applied to the bot identity. Protected branch rules restrict who may push and merge to protected branches, and the bot should not be in the allowed set. Approval rules require a defined number of human approvers before merge, and the bot's approval, if it is permitted to approve at all, should not count toward the requirement.
The agent then pushes branches, opens merge requests, and comments; people review, approve, and merge. That preserves the normal discipline and is what makes the agent acceptable to security. See how to adopt ai coding agents safely.
What issue triage tasks fit?
Applying labels and milestones based on issue templates and content, using the taxonomy the team already maintains rather than inventing one. Detecting duplicates and linking related issues. Linking issues to the merge requests that address them where the developer forgot. Requesting missing information from reporters when a template field is empty. Summarising long issue discussions for someone arriving late.
Each is measurable against what the team did manually, and none changes an issue's substance.
What about drafting changes?
For well-scoped issues such as dependency updates, documentation fixes, and small changes with clear acceptance criteria, the agent can draft the change as a merge request with the issue linked, for a developer to review through the normal path. The pipeline runs, approvals apply, and a person merges.
Drafting should be restricted to issues where the acceptance criteria are explicit, because an agent implementing a vague issue produces a plausible wrong change that consumes review time. See spec-driven development with coding agents.
How is noise managed?
By writing once per event with substance. One review comment set per merge request revision, one diagnosis per pipeline failure, one triage comment per issue. Resolving its own outdated comments when a new revision addresses them. And staying silent when there is nothing useful to say, which is frequently the case and which an agent should recognise.
How is it evaluated?
Review comments by whether developers acted on them, sampled and judged by the team. Pipeline diagnoses by whether the identified error was the actual cause and whether flaky classifications were correct. Triage by label accuracy against the team's own labelling and duplicate detection precision. Drafted changes by review outcome and cycles to merge. And the developer-time measure: time from pipeline failure to fix, before and after.
What does the build sequence look like?
One week on bot identity, scoped tokens, webhooks, and confirming protected branch and approval rules exclude the bot. Two weeks on pipeline failure diagnosis, which is the quickest to prove. Two weeks on contextual merge request review with the team calibrating comment volume. One week on issue triage. Change drafting last, for explicitly scoped issues.
What goes wrong?
Personal tokens. Maintainer role granted for convenience. Bot included in protected branch allowlists. Approval rules the bot can satisfy. Diff-only review that duplicates the linter. Twenty comments per merge request. Drafting changes for vague issues. And diagnosis that quotes the log rather than identifying the error.
How FISTA Solutions helps
FISTA Solutions builds GitLab agents on scoped bot tokens, excluded from protected branch and approval rules so they cannot merge, with pipeline failure diagnosis that distinguishes flaky from real, contextual merge request review at a comment volume developers tolerate, and triage using the team's own taxonomy, through AI enablement, AI agents, and forward deployed engineers. The record behind the approach is 150+ projects for 50+ companies with 99.9% uptime.
To give developers an agent that reads the whole thread from issue to pipeline, message FISTA on WhatsApp, or read how to build a code review agent.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01How does an agent authenticate to GitLab?
Through project or group access tokens with a bot identity, granted the lowest role that permits the required actions, typically Developer for commenting and pushing branches and never Maintainer or Owner, and scoped to specific API capabilities. Personal access tokens tie the agent to an individual and should not be used.
02What makes merge request review useful?
Reading the diff together with the pipeline results, the linked issue, and the project's conventions, so the review addresses whether the change does what the issue asked, why a test failed, and whether it fits the codebase, rather than commenting on style that a linter already enforces.
03How does pipeline failure diagnosis work?
When a pipeline fails, the agent reads the failed job's log, identifies the actual error among the noise, correlates it with the diff, distinguishes a flaky test from a real failure, and comments on the merge request with the diagnosis and a suggested next step, saving the developer from reading thousands of log lines.
04How is the agent kept from merging its own changes?
Through protected branch rules that restrict who may merge and approval rules requiring human approvers, which apply to the agent's bot identity as to anyone. The agent pushes branches and opens merge requests; people approve and merge.
05Which issue triage tasks fit?
Applying labels and milestones from issue templates and content, detecting duplicates, linking related issues and merge requests, requesting missing information from reporters, and summarising long discussions, all using the label taxonomy the team already maintains.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. Weâll map the fastest credible path from intent to verified production.