How a Bad git reset Built an Agent Coordination System
One fateful evening, when vibe coding a project by the name of Wopps, that would never see the light of day, one of my agents had the bright idea to git reset to solve a transient issue that it faced. This would have been a non-issue had I informed my agents to commit after finishing a task. Unfortunately, this was the git reset that made me add that instruction into all future AGENTS.md files, and so, in a blink of an eye, 4 days of hard handcrafted vibes disappeared into the digital oblivion of /dev/null.
This git reset, which set me back 4 days of work, gave me an opportunity; I knew what had to be built! And my agents were the bottleneck. Back then, and to this day, I’m one of these weirdos who like to watch their agents work. Partially because I have trust issues, and in large part because I use much cheaper and faster Chinese models, the fateful agent that caused this article to exist being an emotionally and neurologically unstable Mimo v2.5 Flash.
After about thirty minutes of depression and thinking about sleep, as this was at 4 a.m. and both my agents and I were near our context limit. I started telling my agents to rebuild what they had deleted, but this time I was able to graduate from managing 2 to 3 agents to 4 to 6 agents in parallel. This, like all changes in tech, resulted in a new issue: the agents started stumbling over each other.
You may ask, like most sane, rational people, “Why not just use worktrees?”
This is a very rational workflow that I don’t use much to this day, for 2 reasons. Firstly, I’m building in Elixir, which means my code hot reloads, yes, both front and backend.
I build and test constantly, and secondly, worktrees mean you end up with very large, messy merges that agents don’t always handle in the most graceful fashion. These constraints resulted in me starting to build a system for my agents to lock files to notify other agents where they are working. This worked surprisingly well. Whenever an agent would go to a file, they would lock it before writing. If another agent wanted to go there, their request to lock the file would be rejected, and they would know that there was another agent working there.
Here you can see my agent doing 3 steps:
- Create work: the agent becomes visible to the rest of the fleet.
- Retrieve skill: it gets procedural knowledge rather than reinventing the process.
- Retrieve memories: it gets project-specific historical context.
As my agents continued building, another realisation dawned on me. Using MCP tools for my agents made one thing clear: every request was an opportunity to coordinate with the agent.
This is where Steward changed from a traffic cone into a coordination system. The tool call wasn’t merely something Steward needed to answer with a yes or no. It was an opportunity to talk with the agent. Steward knew exactly what file the agent was about to edit and could give it warnings, instructions, and context from its forefathers.
Returning the {ok} from an application is a wasted opportunity. The agent is telling my app what it’s going to do. There are very few chances to interact with the agent’s context from outside when it’s working, and here is an opportunity to return context to the agent alongside the yes, being squandered with only a HTTP 200. Opportunities to manipulate context are rare, far between, and usually at the agent’s discretion. Here I could send high-quality, fresh context to my agents proactively when they went to a file, giving them relevant context right at the top of the context window. Unlike normal memory systems, Steward doesn’t have to wait for the agent to realise it needs context and ask for it; the agent’s actions themselves create the openings to push the right context.
This has resulted in a system with thousands of memories with agent-based memory quality control. There is an LLM that grades memories and filters out the noise, with up-to-date context, a skill and spec store, plus a system for wrapping an API surface in MCP tools so that the agents working on these applications can call the APIs of the applications that they are working on, get application logs, and many other features.
There are things that I use less nowadays, such as the API wrappers or getting the logs, since I’ve found other tools that solve these problems for me. But ive yet to find a better system to replace Steward with for memories, task tracking, and file locking. It’s a system that’s saved me billions of tokens wasted on my agents getting confused, wandering around and figuring out how the CI/CD works for the 100th time and allows me to manage a fleet of 12+ agents, with an agent satisfaction rating of 95.6%.
What in the hallucination is an agent satisfaction rating? Again, exploiting the MCP response alongside giving agents guidance as to what tools to use, Steward sends the agents surveys to fill out at the end of a task, like a good corporate cruise.
One of the realisations that dawned on me that changed how I built Steward is thinking of the agents I built Steward for as clients and sending them surveys to submit feedback, issues, and feature requests.
This, plus the Meta Harness approach, which is, in short, just gathering an ungodly amount of telemetry from the application, feeding it every hour into a poor, unsuspecting Chinese LLM to distill, then taking the distillations and giving them to an expensive American worker to act as if they discovered the insights. These insights are turned into Steward TO DOs that are then taken up by cheap models to implement, and the loop continues.
Steward, which originated from the neurotic misbehaviour of an agent that belongs in /dev/null, has become my most-used application and a primitive I build around whenever I need to give a team of agents shared context and tools to work together. The source is available on GitHub.
Enjoyed this? Subscribe on Substack for more essays on systems and operations.