← Blog

AI coding agents

Code entropy: how CI checks keep AI from piling up legacy

AI didn't change the law that code complexity grows on its own — it just sped it up. Here are seven levels of CI checks that stop legacy from piling up, in the order I add them.

Entropy grows on its own. In physics this is the second law of thermodynamics. Software development looks similar: mess accumulates easily, and order takes effort.

It has always been this way. As a project grows, so does its complexity: legacy piles up, the project gets harder to maintain, every change gets heavier, and the risk of breaking something keeps rising. In 1974 Manny Lehman wrote this down as a law of software evolution: a system's complexity increases unless work is done to reduce it. AI didn't change that. It only made legacy pile up faster.

You want to keep code as simple as possible, and this is where CI checks help.

  1. 1Dead codefails the build
  2. 2Linters and formattersfails the build
  3. 3Static analysisfails the build
  4. 4Unit tests and coveragefails the build
  5. 5Complexity and duplicationfails the build
  6. 6Architecture and couplingfails the build
  7. 7Written rules (AGENTS.md)not enforced by CI
Cheapest checks at the bottom. Levels 1–6 fail the build every time; written rules on top guide the agent but can't stop it.

1. Dead code

The simplest step is to delete what is no longer used. This can be automated too: put a dead-code finder in CI (vulture for Python, knip for JS/TS), and unused code simply fails the build.

2. Linters and formatters

Next comes the linter: phpcs for PHP, ruff for Python, ESLint and Prettier for JS/TS, and so on. Every language has its own code style, and which one you pick doesn't matter. What matters is not spending time and money on reading code and on review comments like “change double quotes to single quotes”. Agree on the style once, and from then on CI checks it automatically.

3. Static analysis

The next level is static analyzers: PHPStan for PHP, mypy for Python, tsc for JS/TS. Strict static type checking catches simple mistakes: a string passed where a number was expected, or forgetting that a value can be empty. It won't check your logic, but it removes a whole class of errors. Empty values are the classic case: Tony Hoare, who invented the null reference, called it his “billion-dollar mistake”, and according to Harness's 2020 data, NullPointerException was still in the top 10 exceptions in 70% of the Java production environments they looked at.

4. Unit tests and 80% coverage

Why? Because bad code is very hard to unit-test. To make code coverable by unit tests, you have to split it into separate classes and functions, think about coupling and interfaces, decide what stays public and what stays private, use dependency injection, and avoid global state. Of course it's not a 100% guarantee, especially since AI easily inflates coverage with tests that check almost nothing. But one way or another, the threshold puts constraints on the design and raises the odds of good code.

5. Complexity and duplication

Human working memory holds only a few independent items at once: George Miller counted 7 ± 2 in 1956, and Nelson Cowan refined it to about four in 2001. Anything beyond that gets hard and goes to swap. I haven't seen research on this for AI, but if code is understandable to a human, the odds go up that AI won't hallucinate. And if code is so tangled that a human can't follow it, AI will most likely pile new layers of legacy on top.

So I set hard limits on duplication and complexity metrics. Limit exceeded → CI fails → you have to refactor and extract code into separate methods and classes. For Python, complexipy is one example.

If legacy has already piled up, you can start from where you are: set the limit at your most complex function, and things at least won't get worse. But that doesn't make the debt go away. You pay it down separately: set aside time for refactoring and lower the limit after each step to lock in the result. Checks don't clean things up on their own. They only stop the mess from growing.

6. Architecture and coupling

Hard boundaries reduce complexity too, one level up. It's like encapsulation in OOP: the system is built from large blocks whose internals you don't need to keep in your head; knowing their contract is enough. And a block's implementation can be replaced when needed, without rewriting the rest of the system. Instead of dozens of classes and connections, you reason about a few large blocks. Again you shrink the number of things you have to hold in your head at once, this time at every level of the architecture.

In practice this means Clean Architecture and Hexagonal Architecture: code is split into layers, and dependencies between them are strictly limited. The domain knows nothing about the database, HTTP, or external APIs, and outside code reaches it through predefined interfaces. This is checked automatically too (import-linter for Python, for example): a forbidden import appears, CI fails.

7. Written rules

The last level is written rules. These used to be guidelines and documentation for developers; now it's an AGENTS.md file and documentation for AI. Unlike the checks above, the outcome here isn't deterministic: AI can read a rule and still break it. So written rules are the last mile. They hold only what couldn't be covered by checks. All else equal, skip the prose and write a hard check in code.

AI keeps trying to switch off the check instead of fixing the code: raise the limit, add an exception, silence the rule with a comment. So my rules file says it on a separate line: weakening a check to make it pass is a workaround, not a fix.

The tools, by language

Here's what I use for each level; the specific tools matter less than having every level fail the build.
LevelPythonPHPJS/TS
Dead codevultureshipmonk/dead-code-detector (PHPStan)knip
Linters and formattersruffPHP_CodeSniffer (phpcs)ESLint + Prettier
Static analysismypy (strict)PHPStantsc (strict)
Unit tests and coveragepytest + coveragePHPUnitJest / Vitest
ComplexitycomplexipyPHPMD, cognitive-complexity (PHPStan)ESLint complexity
Duplicationpylint duplicate-codejscpdjscpd
Architecture and couplingimport-linterDeptracdependency-cruiser

Bottom line

Entropy will grow regardless; the only question is who holds it back. The more rules move from written guidelines into automated CI checks, the less it depends on a human paying attention in review. AI runs the checks itself, sees what failed, and fixes it: it can skip a line in AGENTS.md, but it can't skip a failed CI run. And the more of these checks you have, the more calmly you can hand code over to it.

If you're a founder and don't read code, this list still works as a set of questions for your team: which of these levels actually fail the build? Every “we catch that in review” answer is a rule that depends on someone's attention on a given day. Why agent-built code needs these guardrails in the first place is its own story — managing AI like a junior.

A template for Python

The approach doesn't depend on the language. For Python I've packaged all of it into a ready-made template; feel free to use it: python-guardrails-template.

Inside:

  • every check from this post behind a single make verify command in Docker, and the same command in GitHub Actions;
  • architecture layers with import checks;
  • a commit-time secret scanner;
  • an AGENTS.md where nearly every rule points to the check that enforces it;
  • a small example, so the checks pass right after the project is created.

A new project is created with one command:

copier copy gh:sg4tech/python-guardrails-template my-project

And when the template gets updated, copier update brings the new checks into existing projects and keeps your changes. No more setting everything up from scratch in every new project.

FAQ

Won't strict CI checks slow the team down?

Checks move the cost from review and production into a failing build, where it's cheapest. On a legacy codebase you start from the current level, so nothing blocks on day one.

Can I add these checks to an existing legacy project?

Yes: freeze the current level as the limit so nothing gets worse, then lower it step by step as you refactor.

Isn't code review enough to catch this?

Review depends on attention and doesn't scale with how fast agents write code. A check runs on every change and fails the same way every time.

Do I need all seven levels at once?

No. Add them in order: dead code, formatting and types first, because they're cheap. Complexity and architecture limits pay off once the codebase grows.