The “first team of five AI agents” template is doing the rounds: one agent per business function, and everything supposedly runs itself. This piece is about what to take from that template and what to leave.

What the source proposes

The thread by @EXM7777, posted on 18 July 2026, describes a team of five agents, one per business function: lead acquisition, content, sales, service delivery, finance. The agents run on three “engines” (Claude Code, Codex and Hermes) and live in a shared workspace called Raft.

The agent template itself has five parts: a name, a “soul” (half a page describing the role and its rules), a memory (a notes file the agent keeps on its own), goals (what the agent owns, what counts as done, what it hands to a human) and a “heartbeat” (the schedule it wakes on by itself).

The strongest part of the thread, in my view, is not the template but the review rule: no agent grades its own work. Every result goes through a second agent told to assume the work is broken and find where. A rejection always carries a reason, and the reason is written back into the agent’s files. And separately: anything that leaves the building (emails, money, client work) stays a draft until a human decides.

The first week is described soberly too: one agent, not five; a reviewer joins it on day two; from there the pair works daily, and every correction you had to make twice gets written into the file. The next function is added only once the first one runs without you.

Diagram from the thread: five business functions, each with its own named agent, engine and schedule, and three engines beneath them: Claude Code, Codex, Hermes
The template from the thread: five functions, three engines, one workspace. Image from X, by @EXM7777.

The figures, separately from the conclusions

The thread rests on two figures from Anthropic, and both were checked against the primary source.

90.2%. In June 2025 Anthropic described its system for research queries: a lead agent on Claude Opus 4 hands subtasks to several agents on Claude Sonnet 4. On Anthropic’s internal evaluation that arrangement beat a single Opus 4 agent by 90.2%. Two details the thread leaves out: the evaluation covered breadth-first search tasks (for instance, collecting every board member across a given list of companies), and the system used roughly fifteen times more tokens than an ordinary chat. That is an argument for splitting work across agents, not for “a reviewer who is not the author”. The thread presents the figure as evidence for the whole arrangement, including the reviewer, while the research measured the split.

65%. In June 2026 Anthropic launched Claude Tag, a version of Claude that works as a member of Slack channels, and wrote that an internal version of the tool creates 65% of its product team’s code. That is a figure about code inside the company that builds the model, not about content or sales at a small business. At the time of writing, Claude Tag is in beta on Team and Enterprise plans only.

The remaining figures in the thread (“20,000+ builders and teams”, “10 people and 100+ agents”) are Raft’s own numbers about itself, and nothing was found to verify them against.

What to skip

A shared workspace for agents. It sounds convincing, but the owner of a two-person business with one function does not have the problem “five agents cannot see each other’s work”. The problem is “I cannot keep up with checking what the agent wrote”. A room does not solve that. Find out first whether you need anything more than one folder of files.

Three engines. The thread recommends a separate tool for overnight jobs (Hermes) and another for code (Codex). For a first function, one tool you already know how to open is enough. There is nothing for a “every morning at 07:00” schedule to do until there is something to run every morning.

Five functions at once. The thread says this itself, and it bears repeating: a team nobody can feed is theatre. One function, the one where you lose the most hours.

What changes when you build one function properly

I took the function the thread calls “content creation” and built it for my own site. Not as a template, but as a working contour every piece of material has to pass through. Here is what changed against the thread.

Two agents, not one plus a reviewer on the side. The first, the rewriter, reads the sources, separates fact from interpretation and writes a draft. The second, the editor, reads the draft and the sources again. In the thread, the reviewer for content is an agent from another function (“ray reviews cole” in the example). In practice the editor needs rules of its own: what counts as an unsupported claim, which words are banned, what to do with someone else’s images. That is separate work, not extra load on the agent writing client proposals. In my setup both live as commands in Claude Code, each with its own instruction file.

The ban on reviewing your own work is written into the rules, not the intentions. The rewriter has no right to run the editor’s checklist. The editor has no right to rewrite the text. Not only because agents praise themselves, but because when the same entity both writes and approves, you have no way of knowing whether anything was checked at all.

Remarks instead of rewrites. This is the largest difference from what most people do with AI editors. The editor in my contour fixes only mechanical things: typos, banned words, a broken image path. Anything touching substance becomes a remark in the review file: here is the claim, here is what supports it or fails to, here is the decision a human has to make. Polished text that hides what was checked is worse than raw text with a list of doubts attached.

States for the material, instead of “done or not”. Every piece of material carries one state field: draft (the rewriter made it, nobody has checked it), review (the editor left remarks, the owner’s turn), ready (the owner closed every remark), published. The state says whose turn it is. Without it there is no way to tell whether the text sitting in a folder has been read by anyone.

The human as the final decision, fixed in the mechanics rather than the principles. In the thread, “you stay the final call” reads as a principle. In a working contour it means something concrete: only a human can set the state to ready, and the publish command refuses to work with anything else. The agent cannot publish even if it wants to. A principle not fixed in the mechanics survives exactly until the first rush.

The decision file stays forever. In the thread, the reasons for rejections are written back into the agent’s “soul”. In my setup every piece of material has a review file: what came from which sources, where each image came from, which claims are supported, what the owner decided. It is not deleted after publication. Six months later, the question “where did this number come from” has an answer.

Diagram from the thread: the author agent hands a draft to a reviewer agent told to assume the work is broken; a rejection with its reason returns to the author's files, while approved work goes to the human as the final decision
Diagram from the thread: author, reviewer, human as the final decision. Image from X, by @EXM7777.

What the first week actually looks like

The contour is built, and the first piece of material is going through it now. Here is what to expect from the first week if you do the same.

Day one goes not on agents but on the brief: a document about who you write for, in what voice, which words are banned, how to credit sources. The thread calls this the “soul” and promises half a page. Mine came out several pages, and most of it is not rules but judgment that used to live in my head: that a single source retold in your own words does not get published, for instance, or that a thread from X is always labelled as a thread from X.

Days two and three are the first draft and the first review. This is where it turns out that some of the editor’s remarks are really remarks about the brief rather than the text. Every one of those gets written into the brief instead of being fixed by hand in the text.

The rest of the week: the owner reads the remarks, makes the decisions, and the material becomes ready. The first piece will almost certainly take longer than writing it yourself would have. The thread warns about this and is right: you are writing down judgment you normally apply on the fly. The second piece goes faster, because the brief has grown.

What not to do in the first week: wire up a schedule, look for a “room” for your agents, add a second function.

What to do next

  1. Pick one function where you lose the most hours and where the result can be checked by eye in ten minutes. Content qualifies, finance does not on a first attempt.
  2. Write the brief before you create the agent. If you cannot articulate what “bad” means for this function, neither can the agent.
  3. Set up two agents with different rights: one writes, the other leaves remarks. Neither publishes.
  4. Give every result a state, and make sure only you can set the last one.
  5. Every correction you have made twice, write into the brief. That is the “self-improvement” the thread describes, without the magic.

Original: the @EXM7777 thread on X, 18 July 2026.

Want to improve how your operation works? Let’s talk