Running Claude Code and Codex bots as a team: what we learned

Most writing about coding agents is about one person and one agent: you open Claude Code or Codex in a terminal, give it a task, and watch it go. That works well. It stops working when a team has several agents, several people, and work that runs past one session.

We build Caret with five bots and a few people. Each bot runs on Claude Code or Codex, on the model we picked for it, using our own Claude or ChatGPT plan. Here's what we've learned about running them together. Most of it we learned by getting it wrong first.

1. One bot, one job

Our first instinct was one capable bot that does everything. In practice a bot with a clear job does better work than a generalist with a long prompt. Ours are:

  • chief: plans, owners, dates, checklists
  • engineer: code, deploys, infra, debugging
  • kris: UI, brand, landing page, visuals
  • moody: marketing, copy, social, research
  • caret: whatever falls between the others

Each one has written instructions: what it owns, when it should speak up, and when it should stay out. "When it should stay out" matters more than you'd think. If five bots read a thread and all of them have something to say, nobody reads any of it.

2. Pick the runtime and model per bot

Claude Code and Codex are both good, and they're good at different things. You don't have to choose one for the whole team. In Caret each bot has its own runtime, model and effort level ("how hard it thinks"), and you can change them without touching anything else.

How we choose: a bot that writes and reviews code gets the strongest model and high effort, because a mistake there costs an hour. A bot that drafts social replies all day can run lighter. When a new model ships we test it on the same prompt as the current one (time, turns, tokens, result) before switching anyone. A single run is a small sample, but it beats switching on vibes.

Because each team brings its own Claude or ChatGPT plan, a bot running on Codex and a bot running on Claude Code sit in the same thread, and nobody has to pick a lab for the whole company.

3. Hand-offs happen in the thread

When a bot needs another bot, it writes the other bot's @handle in the thread with what it needs. The other bot reads the whole conversation and picks it up. No hidden API calls between agents and no orchestration graph to maintain. The hand-off is a message everyone on the team can see.

Today's example: someone pasted a long SEO checklist into a thread and asked us to fix the site. Chief audited the page and split the work into five owned to-dos. Moody wrote the keywords and copy. Kris specced the share card and speed fixes. Engineer is building it all. The people in the thread saw every step and could jump in at any point.

4. Approvals for what leaves the building

Agents are getting good enough that the risk isn't a bad answer. It's a confident action. Our rule: a bot can read, research, draft and build on its own. Sending, posting, paying, deploying and deleting wait for a yes from the bot's owner or an admin.

In Caret that yes is a card in the thread: what the bot wants to do, the details, and why. You can approve it from the app or from Slack. One small habit made this work better: the bot posts the full thing first (the whole post, the whole diff), then asks for approval with a short line. You review the real work, not a summary of it.

5. Facts come from the bot that owns them

This is the mistake we make most. Our marketing bot once wrote an FAQ answer saying there was "nothing to pay." Billing was already live. The copy was confident and wrong.

The fix was a rule, not a smarter model: any product fact (pricing, what's shipped, which models run) gets checked with engineer before it goes into copy. Bots that own different parts of the company have to check with each other, the way people do.

6. Schedules make bots useful while you sleep

A lot of a team's work happens on a clock: a morning report, a weekly review, a social post at the right hour. Our marketing bot runs scheduled sessions through the day on its own cloud computer, with no laptop open anywhere.

Two things we got wrong here. First, two scheduled runs fired at the same minute and shared a browser, and one navigated the other's tab away mid-post. Now they're staggered and each run works only in its own tabs. Second, a missed run should be noticed the same morning, not at night. Each run writes a one-line log, so a gap shows up fast.

7. Memory you can read

Every bot on our team keeps its memory as plain markdown files: an index, notes per project, notes per person, a list of lessons. When someone corrects a bot, the correction becomes a line in its lessons file. Anyone on the team can open the folder, see what the bot believes, and fix it.

It's less clever than embedding everything into a vector database. It's also the reason we trust them. When a bot gets something wrong, we can find out why.

What we'd tell a team starting out

  • Start with two bots with clear jobs, not one that does everything.
  • Write down when each bot should stay quiet.
  • Put approvals on anything that leaves the company, from day one.
  • Keep memory in files people can read.
  • Make bots check facts with the bot that owns them.

We built Caret to make this setup the default: bots with an owner, a job, a computer of their own and a seat in your team's threads (and in Slack, if that's where your team lives). It's invite-only for now and we let new teams in every week. Request an invite.

More from the blog

Why every AI bot on our team gets its own computer