Reinventing the wheel

MP 177: Building a custom harness in 2026?!

I've been using LLMs in some capacity since they came out a few years ago, but I haven't adopted agentic workflows at all. Using a simple chat interface has let me do everything I've needed to up until now, at a comfortable pace.

I have lots of things I want to build now though, so I want to use some automatic code-generation tools. But I don't trust AI companies as far as I can throw a datacenter, and I don't trust the tools they release either. So how do I start to use agentic workflows with a harness I fully trust?

(For the context of this post, a harness is a tool that manages your interactions with an LLM model, or set of models.)

Some quick AI thoughts

Whenever someone asks for my thoughts about AI, I start by clarifying that most of what we're all feeling and thinking is a reaction to AI as it's being rolled out in the current political climate. The tech itself is not the problem; the rash behavior of the companies pushing this tech into every corner of society as fast as they can, regardless of the consequences, is the problem. The roll out of LLM tech would look much different in a well-regulated society.

Like every other author and open-source maintainer, AI has been trained on my life's work. The fruit of that work is now being sold back to me, and to each of us in turn. That's an ugly hypocrisy; I can't walk into OpenAI's headquarters and profit directly off of their work, but they've managed to do that to the entire human population. It's a mess, and one I fully recognize while using this tech to some degree.

AI and boundaries

I became interested in building my own harness after thinking repeatedly about boundaries. AI companies are terrible at boundaries. They seem to want boundaries for you and me and everyone else, but none for themselves. When I think about using AI-related tools, I tend to think about what boundaries these tools should have, what boundaries they actually have, and how confident I am in how those boundaries are enforced.

Claude Code, and other current harnesses, are effective because they can do lots of things. They can generate code, read files, write files, edit files, remove files, run commands, and lots more. They can launch subagents to farm out tasks, allowing them to build large projects with an impressive degree of independence.

Most people I know who use Claude on a daily basis are doing so on someone else's system, and someone else's budget. I asked a friend what would happen if Claude overstepped a boundary and read or erased important parts of their system. Their response was roughly: "That's not really my problem. The company sets up the laptop, manages all the boundaries and access controls, and I just tell Claude what I want it to do." There was nothing flippant about their answer; they were just accurately describing their working environment.

Many of the people I talk to who use Claude at home are using it on dedicated systems much like that corporate laptop; if that dedicated system gets compromised, they just take it offline and rebuild it. I've talked to some people who run agents on their main system, who are vigilant about monitoring all the permission and privacy settings that tools like Claude offer.

I don't really want to manage all that. I don't want to buy a separate computer just for AI tools. I don't want to always work in a VM, or through a VPS. I just want to generate some code in a specific project directory.

A minimal custom harness?

As I've been looking more closely at harnesses, I realize my needs are much simpler than the audience most of these harnesses are trying to serve. I'm a solo developer. I do some open source work, so I collaborate at times, but I mostly work alone on my own projects. I have three core needs for a model harness: The model needs to be able to read certain files, edit those files, and write new files. There are more abilities that would be useful, but those three abilities will go a long way for me.

I can also specify some things I don't need or want. I don't want a tool that can run system commands. I don't want a tool making git commits or pushing deployments. I still want to take care of those tasks myself.

When people raise an eyebrow and ask why you'd write your own harness, I immediately think of all the people who've written a minimal text editor of their own over the past few decades. Most of those people didn't need to write their own editor. But I would guess that most of them enjoyed the process and learned a lot from it.

My recent journey

With all that in mind, I decided to spend a few days trying to write my own harness. I wanted to take it to the point of being able to read, edit, and write files within a project directory. I wanted to know exactly what its boundaries are, and how they're enforced. I wanted to treat it just like a web app, where no input is trusted until it's validated and sanitized.

I decided to use Textual to build the harness as a terminal app. It didn't take long to get something up and running. The first version that let me talk to an OpenAI model through their API was pretty satisfying to see:

TUI app with message "Hi. Are you there?" and response "Hi! I'm here. What can I help you with?"
My first conversation with an LLM through my custom harness.

I'm calling the harness Clarence, because it's a name I've always used in teaching examples, and it's a passing nod to Claude.

Getting a project listing

Once I could send a message and get a response, I wanted to focus on supporting project-based work. Rather than letting the model inspect the project by running system commands, I let it call a list_project_dir command, which indirectly runs a system command that I've defined. Here's that command:

$ tree -a --gitignore -I .git -I __pycache__

This gives the model a clear understanding of the project directory structure, and only shows it what I've committed in Git. That's a nice boundary; this is the same kind of information I've been comfortable pushing to a GitHub repository for years now. I'm not left wondering if the tool is going to try to read my entire filesystem, or anywhere outside the project directory. All it can see is the output of a command I've stored, that runs within the project directory. The project directory is visible in the header at all times, so it should be pretty straightforward to monitor that I'm working in the correct directory.

With this simple ability, the interaction is already much more interesting:

Same TUI app, with more boxes. Message asking about the project, and detailed response listing all that's inferred about the project from the directory listing.
Just by seeing the project's directory tree, the model is ready to do some work.

I asked the model what it can tell me about the project, and it infers a lot just from examining the project's directory tree.

Reading and editing files

I added two more abilities, so the model can read and edit files. To work with a specific file, it passes a relative path from the directory tree it examined. The validation work doesn't just reject any attempt to read or write a file outside the project directory. It also calls out that attempt, and asks the model why it's trying to access something external to the project. I'm really curious to see if any of those boundaries get crossed.

To edit a file, I'm not dealing with patches or diffs yet. Instead the model provides a file path it read from the directory listing, a target string, and a replacement string. The harness then handles updating the file. That's been working so far.

Using Clarence to build Clarence

I reached another satisfying milestone yesterday. A day into this project, I was able to use Clarence to add a new feature to Clarence: Instead of clicking submit after writing a message, I can press Ctrl-S to submit the message.

Using Clarence to build itself is a bit clunky at the moment, because I have to quit the TUI to see the effect of the changes. But this is exactly what I wanted: a tool that can generate code within a specific project, where I know exactly what boundaries have been established and how those boundaries are implemented.

What's next

I don't actually care to take this harness too far. I don't want to spend much time building a harness, I want to build other more meaningful projects. I just want a tool I can fully trust when working on those projects. Clarence seems to be heading in that direction. I'll clean it up steadily as I go, but it's never meant to be a drop-in replacement for Claude or any of the other major harnesses that are currently available. As I've been building this, I've wondered how many solo developers are out there working within a similar custom harness. I'm sure there's a nontrivial number of us.

I still want to own my code. I want to understand the overall structure of my projects, and I want to understand all the details of the critical parts of my projects. But I can also recognize sections of every project where I don't need to know every detail. I'm as excited to continue building things as I was when I was a kid and wrote my first functioning number-guessing game. The mechanics of building things has changed significantly, but the joy of building meaningful things hasn't changed at all.