AI delivery / Development harness
Building an AI development harness: From brief to verified product
How I built the context, reusable skills, scoped tasks and verification loop around Launcherry’s AI-delivered implementation.
Explore the Launcherry seriesThe work around the code determines what ships
A Launcherry founder brings in a product, receives analysis, reviews a campaign and chooses what to export. Research, UX, generation, approval and implementation all have to support that sequence. A convincing answer from a coding agent is useful only when it improves the product the founder can use.
I built an AI development harness around that responsibility. By harness, I mean the working system that gives agents current project context, reusable guidance, bounded tasks and a way to demonstrate completion. I own the strategy, workflow architecture and acceptance decisions; AI delivers the implementation under my direction.
The harness matters at the moment a change crosses a boundary. A copy-generation fix can affect validation, stored content and what the founder sees. A useful working system makes those relationships visible before I accept the change. It also preserves the evidence needed to revisit the decision later.
Give the agent the project it is actually changing
Launcherry’s project records give an agent the decisions, implementation context and verification checkpoints it needs before changing the code. That saves me from reopening settled decisions and helps the agent recognise which capabilities are still planned.
I give those records different jobs. Decisions explain intent; source code shows implementation; dated checks record what was exercised. Release records show what reached the public product. Mixing them up can produce a coherent change against the wrong version of the system.
For a new task, I want relevant context: the user outcome, affected part of the system, existing constraints and the evidence needed at the end. More documentation is useful only when it resolves a real uncertainty. The handoff should make the work easier to judge, without requiring the agent to reconstruct the entire project.
Make completion observable before implementation begins
A scoped task needs a result that someone can assess. In Launcherry, one concrete decision was that generated ad copy should remain complete within platform requirements. Cutting an overlong sentence to fit a field could satisfy a length check while damaging the meaning. I rejected that outcome.
The target was clear: preserve complete copy, keep platform limits and repair overlong output through generation and validation. I reviewed what a founder would receive. The agent’s explanation helped locate the change; the output showed whether it worked.
This is where acceptance criteria earn their place. They connect the business requirement to a check that can fail. They also constrain the solution: a fix that silently relaxes the field limit would miss the requirement, just as a fix that cuts away the sentence would. I describe this example further in evaluating AI outputs beyond passing validation.
Build, challenge, repair and verify
I organise AI-directed development as a feedback loop. A builder implements a bounded change. Critique challenges the result against the requirement. Tests and browser checks provide evidence about behaviour. A targeted repair then returns through the relevant checks before I decide whether to integrate it.
Each part answers a different question. Critique can expose a mistaken assumption or an omitted case. A regression test can catch a known failure returning. A browser check can show whether the intended action remains reachable in the interface. The product decision still needs an owner who can judge whether the combined result is useful.
The supporting Launcherry tooling includes a repeatable demo environment, regression coverage, browser checks and evaluation runners that exercise generation. I also built a bridge for iterating on real-model outputs. Its results inform output quality; production cost, latency and cache behaviour require separate checks against the production provider.
Turn findings into reusable project knowledge
The loop becomes more valuable when a finding changes the next task’s starting point. A missing platform requirement may belong in a skill. An unclear product boundary may need a recorded decision. An implementation defect may need a regression check. I choose the destination according to the failure, so the lesson can influence future work.
Launcherry’s reusable channel skills illustrate this. A review found gaps in platform guidance and drafts carrying internal goal labels into reader-facing copy. I directed a skill library with shared standards and channel-specific guidance, connected to generation and assessed separately. The reusable asset came from a diagnosed product problem.
Maintenance matters here. Guidance can drift away from the workflow that consumes it, and a test can survive after its premise changes. I need evidence about the current connection between instructions, implementation and output. Keeping that connection inspectable is part of the harness’s job.
Use the harness where uncertainty justifies it
This approach carries an overhead: tasks need framing, records need upkeep and evidence needs review. I scale that effort to the change. A small copy adjustment needs less machinery than a change to campaign approval, billing or generation. A useful harness helps me choose the smallest check that can establish the intended outcome.
At the 3 October 2026 local integration checkpoint, Launcherry’s core, web and API suites recorded 10,010 unit-test passes, with 71 API skips and one todo. I use that record alongside the specific behaviours checked and the release history. Test totals alone cannot tell me which features a founder can use today.
My practical recommendation is to start with one recurring change and trace it from brief to acceptance. Identify where context is lost, where judgment is required and which evidence would expose failure. Build the harness around those needs, then expand when the work warrants it. The Launcherry case study shows the product this delivery practice supports.