Content engineering, and a system that writes from code
Content engineering is building the machine that makes the content, instead of making each piece by hand. I build one of those machines. It reads merged pull requests and writes posts and changelog entries from them, so this page has a real example in it and not just a diagram.
The short answer
What content engineering is
Content engineering is the practice of building systems that produce content, instead of producing each piece by hand. A content engineer decides where the raw material comes from, what shape it gets stored in, what writes the draft, who or what checks it, and where it goes. Then the system runs, and the person works on the system.
Ahrefs defines it almost exactly that way: the practice of building the systems that create content. AirOps puts the stress on teams, describing systems that help a team create, update, reuse and distribute content. Both companies sell AI content tools, so read the definitions with that in mind. They are still good definitions.
Same words, two jobs
The older meaning and the newer one
Search the term and you get two different jobs on the same page.
The older one is structured content. The Wikipedia entry describes organizing the shape and structure of content with content and metadata models, and Content Science Review frames it as designing, structuring, delivering and managing content. This is content models, reusable components, metadata, a headless CMS feeding a website and an app and a help center from one source. Documentation teams have done this for years.
The newer one is the marketing meaning. Here the system includes generation: a language model writes the draft from structured inputs, and the engineer owns the prompts, the inputs, the review step and the feedback loop. Jasper's piece on the content engineer is the clearest statement of that version. Ahrefs names both kinds, a structured content engineer and an AI pipeline engineer, and then says its article is about the second.
They are less different than they look. The newer meaning only works if you did the older one first. A model given messy, unstructured input writes messy, generic output, and the fix is almost always upstream of the prompt.
Not everyone is sold on the job, either. Ryan Law at Ahrefs wrote a post saying he wouldn't hire a content engineer, arguing that scaled AI content is losing its edge and that he would rather hire marketing experts who use AI than AI experts. Same company as the definition above, which I find kind of great. His point holds for systems that reshuffle what already exists. It holds less for systems built on a source only you have, which is where the worked example below lives.
Anatomy
The parts of a content system
Every content system I have looked at, including my own, has the same six parts. Most of the failures I have seen come from a missing one.
- A source. Something true that already exists: a product database, a pricing table, support tickets, a git history. The best sources are ones the business keeps accurate for its own reasons.
- A structure. The source turned into fields a machine can read. A record per product, per city, per release.
- A filter. The rule for what deserves a piece at all. This is the part people skip, and it is why so much generated content is noise.
- A writer. A template, a model, or both. The writer gets the structured record and the house rules for voice and format.
- A gate. A human check, an automated check, or a rule that says which pieces can skip the human. The gate is what lets you trust the volume.
- A measure. Some number that says whether the output is any good, read often enough that you change the system when it drops.
Distribution sits on top of all six. Where the piece goes, when, and to whom is a system decision too, and it is the one that GTM teams already automate well. The GTM workflows guide has five of those.
Worked example
A worked example: content from code
Here is the system I run, part by part. I built Merge & Tell because I had years of merged pull requests and a changelog I updated roughly never. I would ship something and then not tell anyone, which is a strange way to run a business.
The source is the merged pull request. A GitHub webhook fires when a PR merges. A PR is a better source than most marketing input, because it has a title, a description, a diff and a timestamp, and the engineering team keeps it accurate for reasons that have nothing to do with marketing.
The structure is a classification. Every PR gets a kind (feature, improvement, fix, perf, security, docs, refactor, chore, ci, test or revert), a visibility (user-facing, indirect or internal-only), and a magnitude (major, minor or trivial). Those three fields are what every later step reads.
The filter runs before anything is written. Bot PRs and anything labelled to skip never become a post. A diff under ten meaningful lines reads as trivial unless something says otherwise. Chores, refactors and CI changes stay internal. Most merges are not news, and the system agrees with that out loud.
The writer reads the diff, not just the title. One pass works out what changed, trimmed to the first 12,000 characters of diff, which matters if you merge a lockfile the size of a novel. A second pass writes per channel: a changelog entry, a Bluesky post, a LinkedIn post, each to its own network's length and rules.
The gate is a set of rules, and one hard line. Routing rules decide which networks a PR goes to, and a post can go out on its own when the rules say so. A post written in a person's voice never does. It waits for that person, always. I will let a brand account post on merge. I will not let a sentence appear under my own name that I have never read.
The measure is the kept rate. The share of drafts that got published, out of the ones somebody decided on. Undecided drafts stay out of the math, so a backlog does not flatter it or punish it.

The output lands in three places: posts on the connected networks, a public changelog page, and a changelog feed in Atom and JSON. The feed is the part a GTM engineer cares about, along with an API that answers "what did we ship this week that was worth talking about." That turns a content system into a signal source, which the product signals guide gets into.
Who does it
Content engineer, content marketer, GTM engineer
The titles overlap and nobody agrees on the borders yet. This is how the work tends to split, not a rule.
| Role | Owns | Typical output |
|---|---|---|
| Content marketer | The piece | An article, a launch post, a newsletter issue |
| Content engineer | The system that makes the pieces | A pipeline, a content model, prompts, a review step |
| GTM engineer | The system that finds and reaches buyers | Enrichment tables, signal routing, outbound sequences |
The content engineer and the GTM engineer use the same tools more often than you would guess: Clay, n8n, a language model, a lot of webhooks. The difference is the output. One ends in a published piece, the other ends in a conversation with a buyer. Teams with one person doing both are common, and in my experience that person ends up the most useful one in the building. If your readers are developers, the developer marketing guide covers which pieces are worth making at all.
The failure modes
Where content systems go wrong
No filter. The system writes about everything because it can. A feed that posts every dependency bump is worse than silence, and I say that as someone whose baseline was silence.
A source nobody keeps true. If the input is a spreadsheet a marketer updates when they remember, the output drifts the day they stop remembering. Pick a source that somebody else is forced to keep accurate.
Sameness. A model with no voice rules writes like every other model with no voice rules. Readers spot it in a sentence. Voice and banned phrases belong in the system, written down, not in somebody's head.
No measure. Without a number you find out the output got bad by noticing, slowly, months later. I would rather find out on a screen.
Honesty
What this page leaves out
I am a solo founder who built one content system for one kind of source. I have not run a documentation team's content model or migrated a company to a headless CMS, so the older meaning gets a paragraph here, not a guide. Merge & Tell also does not write articles like this one. It writes short posts and changelog entries from code, and I wrote this page by hand, which you can probably tell from the jokes.
Your git history is already a content source.
You were going to merge the PR anyway.