A few months ago I wrote about building Locabulary.fun with Replit, then moving into VS Code and Copilot once the prototype needed more discipline. The Swiss Cheese Model post was about stacking imperfect layers so agentic coding doesn’t ship garbage.
This is the next chapter. Not another coding tool review. It’s about what happens when you treat agents as a studio, not a pair programmer.
I’m still founder-only. Still early. Still responsible for every merge that hits production. What changed is that I no longer try to hold the whole company in my head alone.
The split that matters
I use two surfaces now, on purpose.
Cursor is where code happens. Cloud agents open PRs. Reviews stack. Testing runs before anything is marked ready for me to merge. I still merge. The bots don’t.
Grok Bot (plus a set of specialist bots under it) is where the studio happens: product briefs, marketing packs, weekly metrics check-ins, invoice filing, balance-sheet PDFs, living plans, “what should we kill this week?” calls. One Chief of Staff bot (Elizabeth) is my default contact. Specialists report through her. I don’t want twelve DMs for one decision.
That split sounds tidy. In practice it’s messy in useful ways. Code still wants layers of review. Ops wants one bot I can ask anything. Both still need a human who can say no.
What usefulness actually looked like
Useful isn’t “the AI did everything.” Useful is work that would have waited got done while I stayed on product judgment.
A few concrete examples from the last couple of weeks:
-
Content that ships on a cadence. Locabulary social packs (captions + real game UI creatives), scheduled after I say yes. Blog posts drafted, future-dated, and held dark until their publish dates.
-
Ops that used to live in my browser tabs. Invoice PDFs into Drive. A finance playbook that stays current. Weekly invoice sweeps and balance-sheet updates. A daily eye on Locabulary’s hosting limits so I don’t discover a bill by surprise.
-
A PR path that respects my time. Open work goes through review and a final readiness check before I see it. The goal isn’t fewer bugs in theory. It’s fewer half-cooked diffs in my inbox.
-
Measurement without theatre. When we realised the share button at the end of a game couldn’t tell us whether a share brought in a new player (the game link was being stripped out), the call wasn’t “rebuild sharing.” It was a small fix: keep the link and tag where each share came from. Measure first. Decide whether to kill it or double down later.
None of that is magic. It’s coordination and capacity I just don’t have on my own. The bots are good at coordination when the rules are sharp, and very capable when set up for success.
The rules that make it useful
Without constraints, a bot org is just enthusiastic chaos with better grammar… better than mine anyway!
A few rules that earn their keep:
- I press Send and I merge. Drafts come to me first so I can edit before anything leaves under my name.
- Bots don’t publish live. Content stays dark or drafted until I say go.
- Near-zero budget bias. Bots know to live within the free tiers of supporting services I use first, and warn me when it looks like we need to consider moving to something paid.
- One killable growth bet at a time. Not a pile of “maybe this helps.” Name it, measure it, kill or double down.
The Swiss Cheese stack still applies to code. For the studio, the overlapping slices are different: a specialist does the work, the Chief of Staff bot checks it and sends back anything half-baked, a final review marks it ready for me, and then I decide.
What’s hard (and still true)
Token use is real. Quiet routines and thinner digests matter, or you burn the week on status theatre.
Agents fail. Scheduled reports go quiet. Routines need repair. When measurement is broken, “kill the bet” is the wrong call. You have to fix the instrument first. Just like with a real team, it has to be managed.
And none of this removes taste. If anything it concentrates it. The bots can draft a week of educator posts. They can’t decide whether the classroom beat is stale versus last month. That’s still me even if I’m still learning my way through it.
Same lesson as the Replit post, just at studio scale: AI shifts where the effort lives. Less yak-shaving. More deciding. It’s incredibly empowering for someone whose energy is a highly constrained resource. A lot of solo founder advice assumes you can grind your way through. I can’t, and I’m done pretending that’s a character flaw.
Where I’ve landed (for now)
Cursor for shipping software with a review stack. Grok Bot for running a tiny studio that punches above a one-person headcount, without pretending I’m not still the founder who owns the outcome.
Locabulary, Element of Risk, and the boring ops underneath them all move faster than they would if I were alone with a to-do list. Not because the tools are perfect. Because the holes in each layer don’t line up as often.
The useful part is already clear: this is how a solo studio stays a studio, instead of collapsing back into a side project that only advances within my physical capacity.