01 — The fleet
Twenty-two agents, named and accounted for.
Not a free-roaming loop. Each one is a function a durable workflow invokes for a judgment-heavy sub-task: a system prompt, a curated tool subset, an output contract, a spend cap, and a defined escalation.
Domain16
Do the operating work.
sourcingBrief to ranked candidatesvettingFit score and brand-safety flagsoutreach_writerGrounded personalised outreachconversationClassify and answer repliesconversation_responderDraft the reply that goes outlogisticsCreate and track shipmentscontent_verifyFind the post, match the briefanalystCompile and narrate the reportresearchSite to sales analysisintakeOne line to a valid brieflead_outreach_writerOutbound to brandspayment_mandateAuthorise spend against a mandatecomplianceDisclosure and policy checkscreativeBrief-aligned creative directiona11yAccessible copy and structurecustomer_successActivation and retention
Meta3
Route, score, and tune the ones that do.
coordinatorRoute work across the fleetcriticScore before anything shipsoptimizerTune what the critic scored
Watchdog3
Watch drift, spend, and the trust boundary.
anomaly_watchFlag runs that driftcost_watchStop spend before the capsecurity_watchWatch the trust boundary
02 — The replacement ledger
Eight seats a person used to sit in.
The left column is quoted verbatim from our internal agent roster, written before this website existed. We did not rewrite them to read better. They read like someone's actual job because they were — ours.
- before · human
the human clicking Search + Step 2after · agentsourcingBrief in, ranked candidates out, de-duped against the blacklist and every prior campaign. - before · human
the human eyeballing Step 2's tableafter · agentvettingPulls each profile, computes engagement and average views, scores fit, raises brand-safety flags. - before · human
"AI 작성" button + the human reviewingafter · agentoutreach_writerWrites multiple angles, runs a four-judge tournament, checks spam score, returns one draft. - before · human
the human reading the Replies tabafter · agentconversationClassifies every reply, extracts the address or the rate, drafts the response or escalates. - before · human
the human in Step 4after · agentlogisticsCreates the shipment, watches tracking, handles the exceptions. - before · human
the human in Step 5after · agentcontent_verifyPolls for the post, matches it against the campaign brief, computes what it did. - before · human
the human in Step 6after · agentanalystCompiles the report and writes the narrative around it. - before · human
the 6-tab campaign-create formafter · agentintakeA short conversation that turns a one-line ask into a valid campaign brief.
…and 14 more agents doing work that had no human predecessor at all.
Quoted from our agent roster.
03 — How the work moves
Autonomy is a setting, not a personality.
Work moves through named stages, and every irreversible action sits behind a named gate. Gates ship on by default. An operator moves the whole workspace between three levels — and the level, not the mood of a model, decides what happens without a human.
Stages
- 01
overviewBrief accepted, plan generated - 02
sourcingFind and vet creators - 03
outreachWrite, send, handle replies - 04
shippingShip samples - 05
content_reviewVerify posts went live - 06
performanceCompile the report
Gates
approveShortlistafter sourcing + vettingapproveOutreachSendbefore sending each batchapproveReplyResponsebefore sending a drafted replyapproveShipmentrequiredbefore creating a shipmentapproveStageAdvancemoving the campaign to the next stageapproveContentrequiredfinal sign-off on a verified postapproveBudgetrequiredreleasing spend, executing a contract
Autonomy levels
copilotagents only propose; a human clicks every actioncheckpointeddefaultagents act, each stage has an approval gateautonomousagents act end to end; escalate on policy exceptions
04 — The operating layer
$2,400 and 45 hours became $7.40 and 2 hours.
The same piece of work, priced two ways. The left column is what an outside operator charges to run it, plus the hours it takes them. Ours is metered agent and infrastructure cost, plus the hours a person spends at the gates.
The management and operations layer only. The spend that passes through to third parties is identical either way and is excluded — including it would make the difference look larger than it is.
Outsourced to an operator
$2,400
operating cost
20% management fee on $12,000 of pass-through spend
45
human hours
Run by the fleet
$7.40
operating cost
metered agent and infrastructure cost
2
human hours
$12,000 — Pass-through spend — identical on both sides · excluded from this comparison
One engagement of the size a small team would outsource.
05 — How we staff
We don't hire people. We build agents.
That is a statement about where headcount goes, not about who is accountable. Two people run this company, and the boundary between what they decide and what the fleet executes is written down, versioned, and enforced in code.
What stays with a person
- Anything irreversible: releasing spend, signing a contract, final sign-off on published work.
- Setting the autonomy level, and moving it.
- Deciding what the company builds next.
- Everything on this page. Two names, at the bottom.
Gates a person can never delegate: approveShipmentapproveContentapproveBudget
What the fleet holds
- The operating roles a company this size would otherwise hire for.
- The work that runs on a schedule, and the work nobody wants to do twice.
- The checks on its own output — scoring, cost, drift, and the trust boundary.
06 — The company itself
The same rule applies inward.
Merging to main is the release — no human runs a deploy command. Deploys to the agent service land as a zero-traffic canary and promoting one stays a deliberate human decision: the same gate-shaped governance the products are built on.
Gates on every pull request5
- verify-build (lint, build, type-check)
- TypeScript unit suite
- offline Python suite + golden-eval holdout gate
- submission artifact suite
- checksum-pinned secret scan
Test suites that block the merge
- 3,000
- Python
- 596
- TypeScript
This website is in that loop too. The figures on this page are a JSON snapshot committed to a public repository and rendered on the server — nothing here is fetched after the page loads.
07 — What we operate
The fleet is not the product. It is how the products get run.
Each one makes its own case on its own site.
- Social Seedingsocialseed.ing
Creator campaigns, run end to end by the fleet. A brief goes in; sourcing, outreach, replies, shipping and verification come out.
- kbeauty.marketkbeauty.market
A buying service dressed as a shop: pick from a live catalogue, we buy at Seoul retail and ship one consolidated box. Checkout is not connected yet — browsing works, paying does not.
- Teslamteslam.io
Drive-to-earn for Tesla owners in Korea. The unit economics are public on the site; the settlement pipeline is a design, not a deployment. An independent project, not affiliated with Tesla, Inc.
08 — The lab
What else two people shipped.
The fleet frees the calendar, and the calendar fills with builds. Everything below is verifiable without taking our word for it — a live site, a public listing, a demo video, or a judged result.
Satellite
webThe live satellite catalogue on an interactive globe — close approaches, re-entries, space weather — briefed by agents in manual, assist and autopilot modes. EN·KO·JA.
Travel
webAn AI travel planner that plans routes and weather together, and refuses to invent what it cannot verify — no fabricated fares, no imaginary flight numbers. Booking hands off to the real engines.
Somm.dev
webSix AI sommeliers read a public repository in parallel — structure, quality, security, invention — and return one evaluation with its reasoning attached. Any public repo, no sign-in.
vibeDeploy
webOne click runs idea discovery, a live board ranks the candidates, and the chosen app is built and deployed. Took 1st place at DigitalOcean's Gradient AI Hackathon.
VibeCat
macOSA macOS desktop companion that watches the screen with you and suggests before you ask. Honorable Mention at Google Cloud's Gemini Live Agent Challenge.
Glasshat
webAn audit layer for AI judging: a six-perspective panel scores with evidence, then audits its own over-confidence and pulls the scores back — every step a visible trace.
Preview Forge
Claude CodeA Claude Code plugin that turns one line into twenty-six mockups, lets a person pick one, and drives 144 agents to a frozen full-stack app.
GitLab Atlas
GitLabOne command turns a GitLab repository into architecture documents that then defend themselves: every merge request is checked for drift, and every merge regenerates the baseline.
Memex
macOSLocal memory for coding agents — session history as navigable spatial memory, no LLM at runtime, everything on your machine. Its embedding fix was merged upstream into fastembed-rs.
He was Socrates
macOSA fullscreen Socratic bust that listens and only asks back, fully on device — zero network entitlements, first token in 192 ms median, the benchmark committed beside the code. Built for the Gemma 4 Good hackathon.
Find Your Bible Character
iOSTwenty-four questions read four spiritual tendencies back as one of twelve biblical characters, with the model running on the phone. It is on the App Store, and it deliberately gives no score and no ranking.
vibe-mod
RedditWrite a moderation rule in English; it compiles once into a deterministic rule with shadow mode and a 30-day undo. Live in the Reddit App Directory.
Also on the bench
- Fairthonthe six-hat evaluation system Glasshat grew out of, scoring pitch decks and repositories against one rubric
- VibeMeetinga macOS overlay that translates the other side of an English call and drafts three answers you can say yourself
- ClaudeSynckeeps AI coding environments in sync across Macs
- vibeVoicethe text-to-audio dashboard our demo narrations are made with
- openClawWorlda 2D world where people and agents share the map
- Memoed on your lifean evidence-first iPhone app that finds what changed in everyday memories
09 — The competition record
The whole record, results as they fell.
Public competitions are the cheapest neutral benchmark a two-person company can buy: outside judges, fixed deadlines, published winner lists. We enter with the fleet and publish the whole column — the ribbons and the losses alike.
- 12
- entries
- 1
- 1st place
- 1
- Honorable Mention
- 1
- Judging
- Gemini 3 HackathonGoogleSomm.devNo award
- PlayMCP “Player 10”Kakaokidsafe-mcpSelected
Program selection. No public roster page exists, so this row rests on our own records.
- Gradient AI HackathonDigitalOceanvibeDeploy1st place$8,000
- Gemini Live Agent ChallengeGoogle CloudVibeCatHonorable Mention$2,000
- GitLab AI HackathonGitLabGitLab AtlasNo award
- Built with Opus 4.7Cerebral Valley × AnthropicPreview ForgeNo award
- Mod Tools HackathonRedditvibe-modNo award
No award; the app itself passed review and is live in the Reddit App Directory.
- Vector Space DayQdrantMemexNo award
- Rapid Agent Hackathon · Arize trackGoogle CloudGlasshatNo award
- AI Agents Challenge · Track 3Google for StartupsSocialSeed.ingNo award
- Games with a HookRedditPipuzzleNo award
- Global AI Hackathon Series with Qwen CloudAlibabaCircle TakeJudging
The last row is still being judged — the result lands on Aug 21, 2026, and this table will carry it either way.
Judged once: CMUX × AIM Intelligence Hackathon, Seoul, April 2026 — our own record, not a published roster.
And two internal hackathons of our own, scored against the same kind of rubric.
10 — What we can hold for you
Three ways to put the fleet on your work.
Not a promise of outcomes — a choice of engagement shapes. Each one points at the part of this page that backs it.
Run
Creator campaigns, operated end to end: a brief goes in; sourcing, outreach, replies, shipping and verification come out. This is Social Seeding, already operating — section 07.
Build
An operating seat in your company, replaced the way ours were: a named roster, output contracts, spend caps, gates that default to on. Sections 01–06 are the spec.
Prove
Two weeks to a working system: fixed scope, a fixed fortnight, and at the end something you can operate. The cadence the competition record was built on — section 09.
11 — Where we operate
Operating work does not need a local office.
- SeoulKRHQ
- TokyoJP
- UlaanbaatarMN
- Mexico CityMX
- HanoiVN
- BangkokTH
- San FranciscoUS
- New YorkUS
- ParisFR
- MilanIT
12 — Intake
One line is enough.
The fleet has an agent named intake, and its whole job is turning one line into a valid brief. Send a line about your work — the thing that runs on a schedule, the thing nobody wants to do twice. Both founders read every line.
No form? The addresses in the footer work the same.
Grants & programs
Built with
- Google Cloud
- Google Gemini
- OpenAI
- Vercel
- Next.js
- MongoDB
- Go
- Tailwind CSS