Hey, Luca here! Welcome to a new 🔒 weekly essay 🔒 from Refactoring.
Every week I write an article to 170K+ engineers about how to make good software together. To get access to all our articles, resources, and our private community, subscribe to the full version.
Lately, I have become more and more interested in how teams are using AI inside real production workflows.
Sure, your Claude skills are cool, and so is your AGENTS.md, but I want to see how this tech graduates into stuff we run on our servers, reliably, all the time, and creating value.
This is also not the first time I have written about this:
In May, I published How to Orchestrate AI Workflows, which is a primer for moving from AI loops to more structured orchestration.
Earlier this month I published Getting Started with AI Governance, to talk about security, control, and how “if you want the car to go fast, you need brakes!”
So today I want to go deeper into what it means to create good AI workflows at scale, based on conversations I had with the people who are doing this for real. I took a lot from my recent chat with Jan, from friends I met during recent travels, and especially from Ludovic and the guys at Kestra, that I am looping in again after the huge success of the first piece we wrote together!
Here is the agenda:
👁️ Observability — what to watch on agentic runs, plus a practical checklist.
🛡️ Boundaries — how keep sensitive data where it lives, and separating workers from orchestration.
🧑✈️ Human in the loop — when to put a gate, and how to resume without replaying an hour of work.
🔌 How to get started — prefer governed flows as tools for the agents you already trust.
📚 Resources — a few places to go deeper on the practices above.
Let’s dive in!
Disclaimer: I am thankful to Kestra for partnering on this and providing ideas and insights about the orchestration industry. I am a fan of what they build and you should check it out.
However, as always, I will only write my honest opinion on the practices and tools covered, Kestra included.
👁️ Observability
A surprisingly high number of the teams I spoke with built some AI observability tools in-house.
To me this feels like when, around 2022, I ran the first surveys about what tools teams used to measure engineering metrics, and about ~50% of respondents had patched something together themselves, e.g. pulling numbers from Github and putting them on Grafana.
Likewise, today many teams are logging basic info themselves. For every run, this is typically token spend and the final answer of the agent. This is good! You get your hands dirty, learn what you really need, and wrap your head around the idea of measuring this stuff.
However, just like with engineering metrics, these setups fail as soon as you need to drill down into specific steps of the agent runs. Which, as long as your runs are short, might not be a big deal, but as soon as you have a long (and expensive!) recurring job failing halfway through, you may want more control over what happened and why.
In my experience, good observability for such workflows means you can reliably get the following information, ideally from the same system:
Intermediate step outputs — not only the final blob. Each meaningful step should leave something inspectable: a fetched record, a classification, a draft that was about to be sent.
Tokens and cost per run — these are the basics everyone measures, and for good reason. This should be ideally broken down per step / model / tool.
Tool and MCP calls — which servers, which tools, with what inputs and outputs.
Who or what triggered the run — a human (manual trigger), a schedule, a webhook, or another agent session. Origin and session tags matter when you audit later.
State at failure — which step failed, what was already committed, what is safe to retry.
If these still feel a bit abstract, here are a few questions you can ask yourself (or your team!) about your setup:



