Fifty-odd automations and one registry to know they're alive
Every process that runs on its own has a name, a runner, a proof of life and an off switch. Why a dashboard without proof lies, and how we found out.
AI · immagine generataAn agency that builds its own tools sooner or later ends up with a lot of things running on their own. We have about fifty: scheduled jobs on the CRM and the websites, scripts attached to the mailbox, nightly database copies, native automations inside older tools, even scheduled tasks on an office computer. Each one is simple. The trouble is the whole: at some point nobody can say for certain which of them are running.
This piece is about how we answered that: one registry where every autonomous process has an entry, and one rule that keeps the page honest.
The registry, entry by entry
Every entry states the same things:
- a name a person can understand, not the technical path;
- where it runs (the CRM, a website, a script, a native automation, a computer), because that decides who can switch it off;
- its cadence, written the way a person would say it: "every 15 minutes", "nightly at 03:15";
- how soon the next sign of life is expected;
- what it does and who runs it, so that when it stops you know whom to ask;
- its risk: can it write to a client? does it only alert the team? does it only touch data?
- its proof of life, and above all what kind of proof it is.
That last point is the reason the registry exists.
Five kinds of proof, because a date isn't enough
At first we had a page that turned processes green or red based on the last recorded date. One showed as stopped for seven days and looked dead. It wasn't. That date wasn't the process's sign of life: it was a log, written only when there was work to do. The process ran every night and simply had nothing to write.
So we separated five kinds of proof:
| Kind | When it's written | If it's old, it means |
|---|---|---|
| heartbeat | on every run, even an empty one | it's broken |
| log | only when there was something to do | it's normal |
| configuration | only when someone changes it | the date tells you nothing |
| data | measured from what the process produced | the work done is the proof |
| none | leaves no trace | never green: "not measurable" |
Without this distinction a dashboard lies in the worst way. It raises red alarms on healthy systems until nobody looks at them any more. Then, when something really breaks, its red is one among many.
"No proof" isn't giving up
The last row of the table is the most useful. Some processes can't be observed today. Scheduled tasks on a computer don't run if the computer is off, and nobody finds out. Native automations inside an external tool expose no readable log. One copy runs on a service whose results we can't read.
Declaring these invisible, instead of leaving them out, has two effects. The page never shows them as green, so it promises nothing it can't know. And the list of invisible processes becomes the work queue for making them visible. A process that admits it's invisible is more useful than one that claims to be healthy without proof.
What the registry made us find
Taking a census of everything showed up things nobody was watching. We worked from the code in production and from real data, not from what was written elsewhere.
A check that had quietly stopped. A routine that re-read the automatic email classifications to correct them had worked through a few thousand messages and then stopped in mid-June. Nobody noticed, because no indicator said so. It ran as a manually launched session and left no trace.
A dormant but armed engine. An old script that classified emails with an AI model had been idle for months, but its key was still in place and the project still existed. Turning a trigger back on would have put it back in the chain, grabbing new emails before the right engine could. No error anywhere: just contacts that never get created. A registry that didn't list it was declaring a chain healthy with a competitor lying in wait.
Two emails that sent themselves. In July two messages went out to clients while the automation panel said nothing was sending. That incident gave us a master switch. Every automation that can write to a client now starts switched off.
One runner per automation
Moving automations from the old CRM to the new one, we adopted a strict rule: each automation has exactly one runner. The new version is born behind an off switch while its old twin keeps running. The handover is a double gesture: off over there, on over here, in the same minute. Until that happens, the registry shows the new version as "off", and that's the truth. Two runners switched on at once don't produce an error. They produce two different snapshots of the same day, or two reminders for the same thing.
Who watches the watcher
Every alarm in the CRM depends on a single process that refreshes the indicators. If it stops, every alarm stops with it, silently. So an external sentinel, on separate infrastructure, checks one thing every half hour: that this process has beaten recently. If it hasn't, the sentinel alerts on the same channel as the indicators. The sentinel leaves its own heartbeat, which the internal indicators watch. They keep an eye on each other.
The cron that ran but didn't exist
The latest mistake is only a few days old. In the new CRM every heartbeat is tied to the process that emits it: the database rejects a heartbeat from a process that isn't registered. We put a new scheduled job into production and forgot its registry row. The job ran normally, but its heartbeat was rejected, with a warning visible only in the logs. As far as the health indicator was concerned, the process didn't exist.
The constraint did its job: instead of accepting an orphan heartbeat, it refused it. But a refusal that ends up only in a log is still too quiet. We caught it because we were waiting for that row and it didn't arrive. Since then, the registry row is part of the same script that creates the job. After every deploy we check that the first heartbeat has actually arrived.
In short
You don't need an expensive monitoring product to know whether your automations are alive. You need:
- one list, written by looking at what actually runs;
- an honest distinction between kinds of proof;
- an off switch for every process;
- the habit of not trusting green when nobody can prove it.


