Won't strict CI checks slow the team down?
Checks move the cost from review and production into a failing build, where it's cheapest. On a legacy codebase you start from the current level, so nothing blocks on day one.
AI coding agents
AI didn't change the law that code complexity grows on its own — it just sped it up. Here are seven levels of CI checks that stop legacy from piling up, in the order I add them.
Entropy grows on its own. In physics this is the second law of thermodynamics. Software development looks similar: mess accumulates easily, and order takes effort.
It has always been this way. As a project grows, so does its complexity: legacy piles up, the project gets harder to maintain, every change gets heavier, and the risk of breaking something keeps rising. In 1974 Manny Lehman wrote this down as a law of software evolution: a system's complexity increases unless work is done to reduce it. AI didn't change that. It only made legacy pile up faster.
You want to keep code as simple as possible, and this is where CI checks help.
The simplest step is to delete what is no longer used. This can be automated too: put a dead-code finder in CI (vulture for Python, knip for JS/TS), and unused code simply fails the build.
Next comes the linter: phpcs for PHP, ruff for Python, ESLint and Prettier for JS/TS, and so on. Every language has its own code style, and which one you pick doesn't matter. What matters is not spending time and money on reading code and on review comments like “change double quotes to single quotes”. Agree on the style once, and from then on CI checks it automatically.
The next level is static analyzers: PHPStan for PHP, mypy for Python, tsc for JS/TS. Strict static type checking catches simple mistakes: a string passed where a number was expected, or forgetting that a value can be empty. It won't check your logic, but it removes a whole class of errors. Empty values are the classic case: Tony Hoare, who invented the null reference, called it his “billion-dollar mistake”, and according to Harness's 2020 data, NullPointerException was still in the top 10 exceptions in 70% of the Java production environments they looked at.
Why? Because bad code is very hard to unit-test. To make code coverable by unit tests, you have to split it into separate classes and functions, think about coupling and interfaces, decide what stays public and what stays private, use dependency injection, and avoid global state. Of course it's not a 100% guarantee, especially since AI easily inflates coverage with tests that check almost nothing. But one way or another, the threshold puts constraints on the design and raises the odds of good code.
Human working memory holds only a few independent items at once: George Miller counted 7 ± 2 in 1956, and Nelson Cowan refined it to about four in 2001. Anything beyond that gets hard and goes to swap. I haven't seen research on this for AI, but if code is understandable to a human, the odds go up that AI won't hallucinate. And if code is so tangled that a human can't follow it, AI will most likely pile new layers of legacy on top.
So I set hard limits on duplication and complexity metrics. Limit exceeded → CI fails → you have to refactor and extract code into separate methods and classes. For Python, complexipy is one example.
If legacy has already piled up, you can start from where you are: set the limit at your most complex function, and things at least won't get worse. But that doesn't make the debt go away. You pay it down separately: set aside time for refactoring and lower the limit after each step to lock in the result. Checks don't clean things up on their own. They only stop the mess from growing.
Hard boundaries reduce complexity too, one level up. It's like encapsulation in OOP: the system is built from large blocks whose internals you don't need to keep in your head; knowing their contract is enough. And a block's implementation can be replaced when needed, without rewriting the rest of the system. Instead of dozens of classes and connections, you reason about a few large blocks. Again you shrink the number of things you have to hold in your head at once, this time at every level of the architecture.
In practice this means Clean Architecture and Hexagonal Architecture: code is split into layers, and dependencies between them are strictly limited. The domain knows nothing about the database, HTTP, or external APIs, and outside code reaches it through predefined interfaces. This is checked automatically too (import-linter for Python, for example): a forbidden import appears, CI fails.
The last level is written rules. These used to be guidelines and documentation for developers; now it's an AGENTS.md file and documentation for AI. Unlike the checks above, the outcome here isn't deterministic: AI can read a rule and still break it. So written rules are the last mile. They hold only what couldn't be covered by checks. All else equal, skip the prose and write a hard check in code.
AI keeps trying to switch off the check instead of fixing the code: raise the limit, add an exception, silence the rule with a comment. So my rules file says it on a separate line: weakening a check to make it pass is a workaround, not a fix.
| Level | Python | PHP | JS/TS |
|---|---|---|---|
| Dead code | vulture | shipmonk/dead-code-detector (PHPStan) | knip |
| Linters and formatters | ruff | PHP_CodeSniffer (phpcs) | ESLint + Prettier |
| Static analysis | mypy (strict) | PHPStan | tsc (strict) |
| Unit tests and coverage | pytest + coverage | PHPUnit | Jest / Vitest |
| Complexity | complexipy | PHPMD, cognitive-complexity (PHPStan) | ESLint complexity |
| Duplication | pylint duplicate-code | jscpd | jscpd |
| Architecture and coupling | import-linter | Deptrac | dependency-cruiser |
Entropy will grow regardless; the only question is who holds it back. The more rules move from written guidelines into automated CI checks, the less it depends on a human paying attention in review. AI runs the checks itself, sees what failed, and fixes it: it can skip a line in AGENTS.md, but it can't skip a failed CI run. And the more of these checks you have, the more calmly you can hand code over to it.
If you're a founder and don't read code, this list still works as a set of questions for your team: which of these levels actually fail the build? Every “we catch that in review” answer is a rule that depends on someone's attention on a given day. Why agent-built code needs these guardrails in the first place is its own story — managing AI like a junior.
The approach doesn't depend on the language. For Python I've packaged all of it into a ready-made template; feel free to use it: python-guardrails-template.
Inside:
make verify command in Docker, and the same command in GitHub Actions;A new project is created with one command:
copier copy gh:sg4tech/python-guardrails-template my-projectAnd when the template gets updated, copier update brings the new checks into existing projects and keeps your changes. No more setting everything up from scratch in every new project.
Checks move the cost from review and production into a failing build, where it's cheapest. On a legacy codebase you start from the current level, so nothing blocks on day one.
Yes: freeze the current level as the limit so nothing gets worse, then lower it step by step as you refactor.
Review depends on attention and doesn't scale with how fast agents write code. A check runs on every change and fails the same way every time.
No. Add them in order: dead code, formatting and types first, because they're cheap. Complexity and architecture limits pay off once the codebase grows.