All articles

I built forty AI agents and used almost none of them.

For two years I kept a personal AI system alive by hand: domain documents, a spreadsheet of prompts, forty named agents. It worked, and maintaining it cost more than it returned. What changed was not a smarter model. It was the agent getting permission to write.

In July 2025 I wrote a document describing my personal AI setup. It contained this sentence:

Less than 1% of AI users globally have a fully structured, domain-specific personal AI assistant system. This gives you an unfair productivity advantage.

I was the one being unfair to myself. The system was real: eight life domains, forty named agents, a map of how they connected. Trip Planner. Career Strategist. Running Coach. Meeting Notes Writer. A spreadsheet with sixty-five prompt rows and a column for every domain document each prompt depended on, so I could tell at a glance what to update when something about me changed.

I used it. That is the part worth being precise about, because the easy version of this story is that I built a toy and abandoned it, and that is not what happened. I pasted those prompts into ChatGPT and later into Gemini. I kept the list of agents open. It worked.

It also died, and it died of maintenance.

The bill nobody quotes you

Every one of those forty agents was a set of instructions pointing at documents about me: how I work, what I am aiming at, what I have already tried. The instructions were the cheap part. The documents were the expensive part, because they had to stay true.

So after a decision, I updated a document. After a trip, a doctor's appointment, a change of plan, a finished project, I opened the right file and typed the new state of things into it. Nothing about that is hard. It is just work, and it arrives every single day, and it competes with the work you actually wanted the system to help you with.

What keeping a personal AI system current actually costs

  • Every fact about you lives in two places: in your life, where it changed, and in a document, where it has not changed yet.
  • The gap between those two is invisible. Nothing breaks, nothing warns you. The answers just quietly get worse.
  • Closing the gap is manual, unrewarding and never finished, so it loses to whatever is urgent that day.

There is a version of my old knowledge base still sitting in that archive, written in six blocks with instructions to myself at the top of each one: paste this first, paste second, paste third. That was the interface. Before a serious conversation, I reassembled my own context by hand, in order, like loading a program from tape.

I want to be fair to that setup. It was not stupid, and for a while the trade was worth it. But the maintenance bill grew with the system, and the returns did not, and there came a point where doing the thing manually was simply cheaper than keeping the machine that did the thing.

That is the real reason people give up on this, and it is not the reason usually given. It is not that they are lazy or impatient. It is that the arithmetic stops working, and they are right to notice.

Two years spent circling one limitation

I did not fix this by finding a better model. I circled the same wall for two years.

First: ChatGPT with local files. It could read what I pointed it at and could not change any of it. Every result came back to me as text I had to file myself.

Then: Google Drive with Gemini, a whole tracked folder, better retrieval. Same wall. The model reads, the human writes.

Both eras produced the same artefact, a documented self that was accurate on the day I last had the patience to update it. My health, work, finance, travel, personal development and content domains all existed back then, with roughly the names they have now. They survived every tool migration. What never survived was their accuracy.

The wall was never intelligence. It was permission. The assistant could read my life and could not record it.

What changed

Now those same domains are plain Markdown files in a folder, and the agent I talk to can write to them.

That is the entire difference, and it sounds too small to matter until you notice what it removes. I no longer update documents. I have not typed a fact into a file in months. I say what happened, and the writing happens as a side effect of the conversation that was going to occur anyway.

The maintenance cost did not get reduced. It got moved onto the machine.

Everything else people call an AI operating system follows from that one permission. Files are not the point, and folders are definitely not the point. The point is that knowledge about you now accumulates instead of decaying, because capturing it stopped being a separate chore you have to remember to do.

What accumulation buys

Three ordinary things, and then one that is not ordinary at all.

A meeting ends. I open a window and talk for two minutes about how it went and what happens next. Underneath, the transcript is fetched and turned into notes, filed where that topic already lives, next to what was decided about it three weeks ago. Calendar entries appear. Tasks appear. What comes back is not a summary. It is a recommendation about what to do next, made by something that knows the history of that thread.

An inbox is a mess. I ask for it to be sorted, and thirty-odd filters get written. Doing that by hand is an afternoon.

My watch records a run. The data lands in the same system as everything else, so a training question is answered against my actual training, not against the general population.

Then there is the fourth one. I can ask a question that reaches across work, health, finances and travel at the same time, and get an answer shaped by all four.

That last one is not a faster version of anything. There is no prompt that produces it from a blank window, no matter how well written, because the raw material does not exist in the window. It exists only if something has been quietly accumulating it for a year. This is the payoff the first three are financing.

The day it stopped being a chat

For a long time the system was a thing I asked. Then I built a dashboard on top of the same files, and it became a thing I look at.

Nothing underneath changed. Same Markdown, same folders. But some things are understood by seeing rather than by reading, and the difference in practice was larger than I expected. Instead of interrogating the system to find out where things stood, the state of things was simply there, in a shape I had chosen, in colour, at a glance.

Now the day starts by looking at it. I read what is open, what has gone stale, what is waiting on me, and only then do I go and talk to the agent, already knowing what to ask. Fewer questions, better questions.

That is the maturity marker, and it is worth naming because it is not the one people expect. A personal AI system has arrived not when it answers well, but when you stop needing to ask.

What I had to unlearn

Two years of prompt engineering, mostly wasted. I spent a serious amount of time learning to open conversations with you are a senior software engineer, you are a strategic advisor, respond in the following format. I do not do any of that any more. I say what I want in ordinary language and it happens. That skill decayed almost entirely, and I would not spend the time again.

Manual note-taking went with it. The reflex to record something after it happened was a good habit for twenty years and is now just latency.

But the substantial thing to unlearn is not a technique. It is the mental model of a chat with an assistant. What is actually in front of you is closer to a small staff: several capable workers, working in parallel, on a body of knowledge you own, in files you can read yourself. That reframing is the whole skill, and it has nothing to do with writing code.

Which is why the people who do best with this are frequently not engineers. The barrier is not technical. It is being able to say clearly what you want, and business people tend to have had more practice at that than developers.

Where it still does not work

I would not trust an article like this one without this section.

The system is wrong regularly. Things get missed. Instructions get read too loosely, or a step gets done in a way I did not intend. This article went through several rounds of my editing before it said what I meant, and getting it right the first time is still an ambition rather than a description.

The barrier is low to start and real to get far. I said above that this is not technical, and I stand by it, but I have put in a lot of hours to reach the setup I have now, and pretending otherwise would be the kind of lie that makes people feel stupid when their first week is frustrating. It gets easier every few months. It is not yet easy.

And most of what I have built has been thrown away. A whole habit-tracking layer, unused and dead. Two knowledge bases archived two weeks after I created them. An entire bridge between tools, wired up carefully and switched off. That failure rate is not a flaw in the method, it is the method: when building something costs one sentence, the constraint stops being construction and becomes judgement. The scarce skill is no longer making things. It is noticing which of them actually gave you something, and deleting the rest without sentiment.

Why none of this is about Claude

The obvious objection: I have tied my working life to one vendor, and vendors get overtaken. Fair. Today it is Claude. In a year it may well be Codex, or whatever comes after it.

That objection would have been fatal to my 2025 setup. The prompts were shaped to one product, the agents lived inside a specific interface, and moving meant rebuilding, which is exactly what moving from ChatGPT to Gemini actually cost me.

It is not fatal now, because what I would carry over is not the tool. It is a folder. The knowledge is in Markdown, the working rules are in Markdown, the whole harness is text. Switching means pointing something new at the same directory. I do not have to explain who I am again.

So the useful question is not which model is best this quarter. It is where the last two years of your context ended up. If it is in a chat history, you are renting it, and most of it is already gone. If it is in files you own, the tool underneath is a detail, and you can afford to change your mind.

None of this required me to write code, and none of it required me to be clever. It required me to stop maintaining my own system by hand, and to let the thing that reads it also write it.