AI delivery / Output evaluation
Evaluating AI outputs: Beyond passing validation
Why I rejected truncated ad copy in Launcherry, and how I connect structural checks, real-model evaluation and product judgment.
Explore the Launcherry seriesThe copy fitted the field and failed the reader
In Launcherry, I rejected an approach that cut generated ad copy to fit platform limits. A field could meet its length requirement while carrying an unfinished thought. For the founder, the deliverable still needed repair. For the validation layer, the length problem could appear resolved.
That gap shaped how I think about evaluating AI outputs. Structural validity, useful communication and readiness for action need different evidence. I directed changes to generation and validation so overlong copy would be handled within the requirements, then checked the behaviour through regression coverage and real-model evaluation.
The case is specific, but the decision applies across AI products. A system can produce the required format without achieving the user’s purpose. My job is to define that purpose clearly enough that implementation and evaluation can expose the difference.
Define usefulness before choosing a score
For Launcherry’s campaign copy, usefulness includes a complete message, appropriate channel treatment and claims supported by the business context. Platform constraints remain part of that requirement. The founder should receive material they can meaningfully review, without first reconstructing a sentence or translating internal planning labels.
I begin with the failure and the intended result. In the truncation example, a useful criterion is whether the text remains complete within the allowed field. That is more actionable than an overall quality score: it identifies the problem, establishes the boundary and gives a reviewer something specific to examine.
Here is an illustrative example, not a recorded Launcherry output. Suppose an overlong draft says, “Plan your next campaign with recommendations tailored to your product.” Cutting it mid-sentence may satisfy the character counter. Writing a shorter, complete message requires a different transformation. The counter cannot decide whether the new message preserves the important meaning.
Use deterministic checks for requirements they can establish
Structural checks are valuable because some requirements can be assessed precisely. A response can be checked for required fields, valid values and platform limits. Regression coverage can exercise a known failure and show whether it returns after a change.
A length check answers “does it fit?” A schema check answers “can the system process it?” Persuasive writing and factual support need review of the content itself. I use each check for the question it can answer.
On 4 October 2026, local development records captured another source of truncated copy: a provider-facing length constraint could cut the text before local checks received it. I directed the removal of that upstream stop while retaining local output limits and stored-data constraints. Release availability for this change remains separate from the local verification.
That investigation matters because a downstream check can only examine what it receives. When a value has already been altered upstream, a valid-looking result may hide the original failure. I need to trace the output through the workflow before deciding which layer should change.
Exercise real generation alongside regression coverage
Deterministic tests give repeatable evidence about the code. Real-model evaluation adds evidence about how generation behaves with actual model outputs. Launcherry has evaluation runners that exercise the generation pipeline, and I built a bridge to support iteration on those outputs before a production-provider check.
That combination helps answer different questions. Did the repair logic behave as expected for a known case? Did the model produce appropriate material for the supplied business context? Did a prompt change improve the particular failure being investigated? None of those questions can be answered solely by the agent that implemented the change reporting success.
The bridge lets me assess generated output. Production cost, latency and caching are checked against the production provider. I retain the inputs, conditions and unresolved failures with each evaluation so the next decision starts from a usable record.
Judge the output as well as the instructions
The Launcherry skill library separates guidance used during generation from criteria used to assess the result. The authoring standard checks the skill files. Evaluation considers the outputs. I want each layer to be clear about what it is judging.
When a draft follows the instructions and still disappoints, I inspect the instructions, business context and model behaviour. Evaluation needs to reveal that failure, even when the draft faithfully echoes the guidance.
I also review the evaluation itself. A criterion that rewards the wrong signal can encourage a weak result. In the truncation case, satisfying a field limit was only one requirement. Preserving complete, useful copy had to remain visible in acceptance. Otherwise the evaluation would make the undesirable shortcut look successful.
Report what the checks establish
The 3 October 2026 local checkpoint records 10,010 passing unit tests across core, web and API, with 71 API skips and one todo. These are suite results at one date. I report them alongside the behaviours exercised; user-flow coverage and release readiness need their own evidence.
I treat generation evidence with the same discipline. A measured result should identify the sample and the behaviour assessed. A recorded failure deserves a concrete explanation and a targeted correction. Broad claims about “AI quality” hide too much of the work needed to judge them.
For an AI product team, I would start with an output that currently causes human rework. Define the unacceptable state, retain an example and create a check that could expose it again. Pair that with review of the actual deliverable. The development harness makes this loop repeatable; the Launcherry case shows why the loop belongs to product delivery.