How I Cut Manual Infrastructure Security Reviews by 80% with Checkov and GitLab CI

E

E.L.

Guest
Infrastructure as Code makes infrastructure changes reproducible, reviewable and auditable. It does not make them secure.

A perfectly valid Terraform configuration can still expose a database to the Internet, create an unencrypted disk, grant excessive IAM permissions or attach a public IP address to a workload that should never be publicly reachable. Terraform will apply all of it without complaining, because that is exactly what I told it to do. But that was not even the part that hurt most.

Our infrastructure was managed through Terraform and GitLab CI. Security engineers could review infrastructure changes, and for a while they did. Adding them as mandatory approvers on every Merge Request did not scale, and most of what reached them was not interesting: a tag, an instance size, a new output variable.

Changing a harmless tag should not require the same security review as opening a firewall rule to 0.0.0.0/0.

So I went looking for something in between:

automate the obvious security decisions and involve a security engineer only when a change actually crosses a security boundary.

I built that workflow around Checkov.

Why I Picked Checkov​


Checkov performs static analysis of Infrastructure as Code and supports Terraform among several other frameworks.

It ships with a large set of built-in policies, but the part that mattered to me was that I could write my own.

Custom policies can evaluate Terraform resource attributes and the relationships between resources, including logical combinations of conditions. Checkov supports both custom policy mechanisms and external check directories or repositories, so the checks could live in their own repository and be versioned like any other code.

That was the deciding factor. Built-in policies encode the industry’s idea of a mistake. Custom policies let me encode ours.

The Pipeline​


The basic flow was simple.

When a developer opened or updated a Merge Request, GitLab checked whether any Terraform files had changed.

Conceptually:

Code:
checkov:
  stage: security

  rules:
    - changes:
        - "**/*.tf"

  script:
    - checkov -d . --external-checks-dir ./checkov/custom-checks

The real pipeline had more configuration, custom checks and reporting around it, but the principle was the same:

do not run the Terraform security stage when Terraform has not changed. When .tf files changed, Checkov ran before the infrastructure could be merged.

That changes rule is doing far more work than it looks like it is.

If you enable a full scan across an existing Terraform codebase, day one gives you a wall of findings, almost all of them about infrastructure nobody is currently touching. That does not improve security. It blocks unrelated work, and it teaches people that the security stage is something you scroll past on the way to the merge button. Once a check loses credibility, you do not get it back cheaply.

So I never ran Checkov against the whole infrastructure at once. The scope was limited through the pipeline’s changes rule: the security stage ran only when Terraform files were actually modified, which in practice meant that only the infrastructure being changed right now was evaluated. I then expanded coverage gradually until the policies applied to all resources.

Existing infrastructure was never retroactively blocked. Legacy findings got addressed as those resources were touched for other reasons, instead of in one big remediation campaign that nobody has time for.

The workflow looked roughly like this:

HRS0SxJQeDVMsVv2lcvo3v6v0VV2-ce93bwb.jpeg



A Merge Request touches .tf files, the security stage runs Checkov, and every check that passes lets the Merge Request continue through the normal review flow. A check that fails adds one more requirement: approval from a security engineer before the change can be merged.

The distinction between PASS and FAIL mattered here, and not in the way people usually assume. A failed Checkov check did not automatically mean the change had to be rejected. Some infrastructure changes legitimately need configurations that would normally be considered risky. A public-facing service, for example, may need Internet exposure by design. In those cases I did not want the pipeline making the final call. I wanted it to ask for a second opinion.

So the logic became:

Code:
Checkov PASS
→ normal Merge Request flow

Checkov FAIL
→ security approval required

I want to be precise about how the approval part worked, because this is the piece people expect to be dynamic, and in our case it was not.

I did not generate approval rules from the result of the pipeline. We used GitLab Enterprise Merge Request approval rules, and those rules were static. Every repository had its own approval rule with an explicitly defined list of approvers, so the people who could sign off on an infrastructure change were known per repository rather than assigned globally or per Merge Request.

What the pipeline changed was not who could approve. It changed how often anyone had to.

That distinction is the whole trick. I automated the routine security checks without trying to automate the decisions that still needed human context.

Who Gets Involved, and When​


A failed check was first of all feedback for the author, not a ticket for the security team.

In most cases the author just fixed it: removed an unnecessary public IP, enabled encryption, narrowed an IAM policy. The next commit passed, and no security engineer ever opened the Merge Request. The same applied to accidental triggers. If a rule fired on something the author did not intend, they fixed it and the pending security item closed itself.

Security got involved only when the author explicitly decided not to change the configuration, because the flagged setting was intentional and required by the service design.

I also added a simple time-based fallback. If a finding stayed unresolved for several hours, a security engineer picked it up, with an expected response time of 24 hours. That gave engineers room to fix the easy cases themselves, while making sure deliberate exceptions still got reviewed instead of quietly aging out.

The key point is that escalation depended on intent, not only on the severity of the finding. A tool can tell you that a port is open. It cannot tell you whether it is supposed to be.

What I Checked​


I used both built-in Checkov policies and my own custom checks.

The exact rules depended on the resource type, but the common ones were:

  • public IP addresses where they were not expected;
  • unencrypted VM disks;
  • unrestricted network access;
  • overly permissive IAM permissions.

A Custom Check in Practice​


Code:
from typing import Any

from checkov.common.models.enums import CheckCategories, CheckResult
from checkov.terraform.checks.resource.base_resource_check import BaseResourceCheck


class EC2PublicIPCheck(BaseResourceCheck):
    def __init__(self) -> None:
        name = "EC2 instances must not receive a public IP"
        id = "CKV_CUSTOM_AWS_1"
        supported_resources = ("aws_instance",)
        categories = (CheckCategories.NETWORKING,)

        super().__init__(
            name=name,
            id=id,
            categories=categories,
            supported_resources=supported_resources,
        )

    def scan_resource_conf(
        self,
        conf: dict[str, list[Any]],
    ) -> CheckResult:

        public_ip = conf.get("associate_public_ip_address", [False])

        if public_ip == [True]:
            return CheckResult.FAILED

        return CheckResult.PASSED


check = EC2PublicIPCheck()

This one is deliberately simple. The check inspects aws_instance resources and fails when associate_public_ip_address is explicitly enabled.

Our production rules carried more context than that. A public IP is not automatically a vulnerability in every case, so the policy could take the environment, the resource type, the network or other attributes into account before deciding whether a change needed a second pair of eyes. I would still start simple. A check you can explain in one sentence is a check people will argue with productively.

What I Have Not Solved Yet​


Checkov supports inline suppressions (#checkov:skip=CKV_ID:reason). I have not built any control around them yet, which means a suppression can currently be added in the same Merge Request that introduces the finding. Treating a new suppression as its own review trigger is the obvious next step, and it is on my list.

Scoping the scan to changed files has a known blind spot too: a change to a variable or a module can affect resources declared somewhere else entirely, and those resources never get re-evaluated. I know about it. I have not fixed it.

Final Thoughts​


Checkov was useful to me because it reduced the amount of manual review. The standard cases got handled automatically, and the exceptions still went to the security team. It worked well for Terraform specifically because most of our security requirements were predictable enough to express as code.

Before this workflow, roughly 12–15 infrastructure Merge Requests per day went through manual security review. After it, 2–3 per day did, and in those cases the security engineer was not reviewing the change from scratch. The finding was already identified, the author had already stated that it was intentional, and the only thing left was a decision: approve the exception, or do not.

That is the part worth keeping. The automation never decided anything about risk. It removed everything that did not need a decision, so that the cases which did arrived with the context already attached.
 

Thread statistics

Created
E.L.,
Replies
0
Views
3
Back
Top