~/blog/gtm-as-code-tag-manager-in-pull-requests

gtm-as-code: Google Tag Manager Changes Belong in a Pull Request

A Terraform-style CLI for GTM and GA4: YAML config, a plan you review in the PR, apply from CI, and adoption of a live container without rewriting it by hand.

Krzysztof Słomka12 min read

Scattered switches and sliders drawn into one stream that passes through a lit, approved gate and leaves as an orderly stack of configuration pages

Open the version history of a Google Tag Manager container that has been in use for a couple of years and read it top to bottom. "Version 37", no notes. "Version 38", no notes. "fix", published on a Friday. Every one of those versions changed what a production website sends to an analytics property, and not one of them went through a review.

I don't think that is carelessness. It is what the tool invites. GTM is a web interface with a publish button, and that button never asks who looked at the change, which commit it belongs to, or what the container looked like before. The code that renders the page goes through a pull request, CI, and a preview deploy. The code that decides what the page reports about its visitors goes through a click.

I kept running into that gap on my own sites, so I built gtm-as-code, a CLI that treats a GTM container and a GA4 property the way Terraform treats cloud resources. This article covers what the tool manages, the validate, plan, apply, publish loop, how an existing container gets adopted without being rewritten by hand, the ownership rule that keeps it from touching anything it did not create, and the two bugs that only a real, live container could surface.

The container is production code with no review#

A GTM container holds three kinds of objects: tags (what gets sent), triggers (when), and variables (with what values). A GA4 property adds custom dimensions, metrics, key events, and a handful of settings. Change any one of them and the data changes. Sometimes nobody notices for months, until a conversion rate looks wrong in a report.

The GTM interface has workspaces and versions, which sounds like version control. It isn't. A version is a snapshot with an optional note, the diff view is per object, and nothing ties a version to the change in the application that needed it. When the frontend renames a data-analytics attribute and the tag still reads the old one, the two changes live in different systems with different histories.

A better GTM interface wouldn't fix that. Taking the interface out of the write path does.

What gtm-as-code is#

gtm-as-code is a Node.js CLI, gtm-code, published as @stackmade/gtm-as-code. You describe the container and the property in one YAML file, and the tool diffs that file against the live state through Google's Tag Manager and Analytics Admin APIs.

A trimmed version of the config that runs this site:

version: 1
project:
  name: slomka-pro
google:
  gtm:
    accountId: ${GTM_ACCOUNT_ID}
    containerId: ${GTM_CONTAINER_ID}
  ga4:
    propertyId: ${GA4_PROPERTY_ID}
gtm:
  variables:
    dlv_analytics_event:
      type: autoEventVariable
      varType: ATTRIBUTE
  triggers:
    click_tracked_cta:
      type: click
  tags:
    ga4_tracked_cta_click:
      type: ga4Event
      eventName: "{{DLV - analytics event}}"
      measurementId: G-Y48RF8DBKE
      trigger:
        - click_tracked_cta
      consent:
        status: notNeeded
ga4:
  keyEvents:
    cv_download: {}
    contact_email: {}

Two details matter here. The account, container and property IDs are ${...} references, resolved from the environment or a git-ignored analytics/.env.analytics, so the YAML never carries them. And every ga4Event and googleTag tag must declare consent explicitly. There is no default, because a default consent setting is exactly the kind of thing nobody reviews.

The loop: validate, plan, apply, publish#

I copied Terraform's workflow on purpose, so anyone who has run terraform plan can read the output without a manual.

reviewed and merged

bad release

analytics.yaml
in a pull request

gtm-code validate
offline, no credentials

gtm-code plan
read-only scopes

gtm-code apply
writes to the GTM workspace

gtm-code publish
new container version

gtm-code rollback

validate parses the file, checks the schema and dependency cycles, and needs no Google credentials, so it runs on every commit. plan authorizes with read-only scopes (tagmanager.readonly, analytics.readonly) and prints the diff:

Google Tag Manager

+ create variable   form
~ update trigger    generate_lead
    triggerCondition: {...} → {...}
- delete tag         GA4 - old_tag

Plan:
  1 to create
  1 to update
  1 to delete

The exit codes follow terraform plan -detailed-exitcode: 0 for no changes, 2 for pending changes, 1 for an error. With three codes, CI can tell "nothing to do" apart from "something broke".

apply runs the same diff and executes it. Deletes go first, tags before triggers before variables, then creates in dependency order, so a tag can reference a trigger created in the same run. It writes only to the GTM workspace. Nothing reaches visitors until publish creates a container version, and that version is named after the current commit's short SHA and subject line. The GTM version history stops saying "Version 38". It says which commit changed it.

rollback republishes the version that was live before the current one. It does not touch the workspace or the local state, so it costs almost nothing once there is history to roll back through.

Deletion needs its own permission#

I made the flags deliberately strict. A tool that can delete production tags from CI should be annoying about it.

--auto-approve skips the confirmation prompt. It does not allow deletes. A plan that removes anything also needs --allow-destroy, and a resource marked protected: true needs --allow-destroy-protected on top of that:

# GOOD: routine CI apply, fails if the plan deletes anything
gtm-code apply --auto-approve

# BAD: copied from a runbook once, now every accidental removal from YAML ships
gtm-code apply --auto-approve --allow-destroy

The flag your routine pipeline already passes can't quietly widen what it is allowed to do. With no terminal attached, apply does not hang waiting for an answer. It declines and names the flag it wanted.

apply also refuses to run if the workspace has changes it cannot safely merge, most likely because someone edited the same workspace in the GTM interface. It stops rather than overwrite them. You can still use the interface; it just stops being the silent source of truth.

Ownership: never touch what you did not create#

gtm-as-code only updates or deletes resources it knows it manages. Most of the other design choices follow from that.

For GTM objects, ownership is a marker written into each object's notes field. For GA4, which has no such field, it is recorded in .analytics/state.json, committed to the repository so a fresh CI checkout knows the same thing a laptop does. Anything without the marker is invisible to plan and apply. A tag a marketer added by hand last year is not a pending delete. It is simply not in scope.

The rule has a cost. Any container that already exists (and those are the ones worth managing) starts out entirely unmanaged. Rewriting it by hand in YAML would be slow, and I'd get small things wrong, so the tool reads it instead.

Adopting a container you already have#

pull reverse-generates YAML from the live container and property. It can pull everything, one resource (--resource tag:generate_lead_tag), or a GTM export file (--from-export), which needs no API permissions at all. Pulling is read-only. The YAML it writes is a proposal.

adopt is the one write in that flow. It stamps the ownership marker onto a resource that is already in the config, re-submitting the live GTM object with the marker in notes and no functional change, or recording a GA4 resource in the state file. It always asks Continue? [y/N], and there is no flag to skip that question. Adoption is a one-time decision per resource, and it should be made by a person.

I adopted this site's container that way on 2026-08-31: two variables, one trigger, two tags and a key event, resource by resource. The only real diff between the pulled YAML and live state was the explicit consent block the schema requires. After one apply, plan reported 0 to create, 0 to update, 0 to delete.

What I left out on purpose matters just as much. GA4's auto-created All Users and Purchasers audiences have empty filter clauses the schema does not accept, and the system "Default channel group" cannot be deleted through the API anyway. I didn't adopt any of them, so the tool leaves them alone.

What a real container found#

The unit tests said the diff engine was correct. The live container disagreed twice.

A diff that never went away#

After adoption, plan kept reporting ~ update on every variable and on the trigger, with no field-level diff printed underneath. Running apply changed nothing, and the next plan showed the same updates.

The cause was a mismatch between two representations of "absent". The config side, parsed with Zod, materialized an unset optional field like folder as a key with the value undefined. The live side never emitted the key at all. The equality check compared key counts first, so the two objects were different forever. The field-diff printer used JSON.stringify, which drops undefined, so it saw nothing to print. So one function reported a change and the other found nothing to show.

The fix in src/core/diff.ts drops undefined-valued keys before comparing. It shipped with a regression test as 0.8.1 the same day.

An API that says yes and does nothing#

GA4's enhanced measurement has flags for site search, video engagement and form interactions. The Admin API accepts a PATCH to them without an error. It does not apply the change. Declared in config, those three produce a permanent diff that no apply can close.

I can't fix that on my side. The config leaves those three unmanaged, with a comment explaining why, and manages the flags the API actually honours. A key event GA4 marks deletable: false gets similar treatment: removing it from config fails plan up front with a message naming it, instead of failing halfway through every apply.

Neither problem would have shown up against a mock, because a mock only knows what you told it the API does.

Authentication is the hard part#

The YAML is the easy part. Most of the setup time goes into Google auth.

gcloud auth application-default login looks like the obvious local setup, and it does not work: Google rejects the tagmanager and analytics scopes on gcloud's own OAuth client. What works is service-account impersonation:

gcloud auth application-default login \
  --impersonate-service-account=analytics-sync@<project>.iam.gserviceaccount.com

The service account needs Edit in GTM's user management and Editor on the GA4 property. Read is not enough, and GTM makes that confusing: its write endpoints return 404 Not found or permission denied instead of 403 when the caller can only read. The first time I hit it, I went looking for a wrong container ID. The ID was fine.

In CI there is no key file. The GitHub workflow authenticates through Workload Identity Federation, with the provider's attribute condition scoped to a single repository, and runs plan on every pull request that touches analytics/:

- uses: google-github-actions/auth@v3
  with:
    workload_identity_provider: ${{ secrets.WIF_PROVIDER }}
    service_account: ${{ secrets.WIF_SERVICE_ACCOUNT }}
- uses: StackMade/gtm-as-code-action@v0
  with:
    command: plan

The action wraps the CLI. With --format markdown, the plan comes out as a table that can go into a PR comment unchanged, so the reviewer sees what the container will do without leaving the pull request. That step is the reason I built the tool.

How it compares#

GTM interfaceContainer export in gitgtm-as-code
Source of truththe live containerthe live containerthe YAML file
Reviewable diff before publishper objectraw export JSONa plan, per resource
Tied to a commitnoonly by conventionversion named after the commit
Deletes need a separate opt-innonot applicableyes
Leaves unmanaged objects alonenot applicablenot applicableyes
GA4 dimensions and key eventsseparate interfacenosame file

The middle column is the obvious first attempt, so it's worth saying why it falls short. An exported container in git is a backup, not a workflow. Nobody reads a raw export diff line by line, and nothing applies it back.

What it does not do#

The scope is narrower than "all of GTM". I'd rather say so here than have someone find out halfway through a migration.

It does not create containers or properties. It manages resources inside ones that already exist. It does not support GTM custom templates or community gallery template tags, because their payloads carry fields I cannot verify against a sandbox container. It writes to one container and one property per run; several environments means several runs with --env. And it cannot tell whether your site actually fires the events it declares. The separate verify command queries the GA4 Data API for that, which is a different question from whether the configuration matches.

It also doesn't make GTM's interface go away, and people will keep opening it. What the tool does is make sure a change made there either shows up in the next plan or blocks apply instead of being overwritten.

When not to use this#

A container with one GA4 tag that nobody has touched in a year does not need a CLI, a service account, Workload Identity Federation and a GitHub workflow. The setup is mostly Google auth, and it pays back nothing if the container never changes.

It also fits poorly when the people changing tags are not the people who work in pull requests. If marketing owns the container and engineering never sees it, moving the write path into git moves it away from its owners. That is an organizational decision before it is a tooling one.

It earns its cost when the tracking plan changes together with the application, when a broken tag costs real money, or when staging and production need the same setup. At that point the container is code, and I want it reviewed like code.

StackMade/gtm-as-codeGoogle Tag Manager and GA4 configuration as code: declare resources in YAML, review the diff, apply Terraform-styleTypeScript★ 0⑂ 0view on GitHub →

$ related posts