Generative Engine Optimization (GEO) taught content authors how to structure their material so external AI engines retrieve, cite, and elevate it correctly. Schema markup. Citation engineering. Retrieval-aware structure. Evidence-anchored writing. Visibility metrics with drift monitoring. The discipline matured fast — 84% of marketers recognize the term, AI-referred sessions jumped 527% YoY, Gartner projects 25% organic-search decline by end of 2026.
Now turn the lens around. What if you took those exact practices and applied them to your own agent’s internal context infrastructure — the AGENTS.md files, the prompt structure, the retrieval mechanisms, the verification steps your own agent reads before it acts? Different consumer (your agent, not someone else’s), same structural problem: organize information so an AI retrieves and grounds correctly. That’s the framing this post walks through.
Where the field actually is
Harness engineering has the name now: OpenAI, Phil Schmid at Google DeepMind, Anthropic’s two reference papers on long-running agents and the Planner / Generator / Evaluator architecture, Martin Fowler, Atlan’s 2026 guide. The optimization layer has working frameworks — HARBOR formalizes constrained Bayesian optimization over harness configuration; Meta-Harness uses an outer-loop coding agent to read source, scores, and traces and write new harnesses. Context engineering crystallized as a discipline. The closest published neighbor to what we’re proposing is AEO (Agentic Engine Optimization, Addy Osmani / Google Cloud AI) — but AEO is your content optimized for external agents. The gap nobody has packaged is the explicit reverse-direction translation: GEO practices applied, deliberately and one-to-one, to your own agent’s internal harness.
The reverse-direction translation
Five mature GEO practices map cleanly onto five harness moves. The framing is the contribution; the practices on the right are real and consistent with current AGENTS.md and context-engineering literature. A practitioner who understands GEO can read this and orient on harness engineering in two minutes.

Outcome data, not benchmark data
Even with the right framework, you have to pick what signal to optimize against. Harbor anchors on benchmark pass rates from a fixed task suite. Meta-Harness reads per-task execution traces. Both implicitly assume the benchmark is a faithful proxy for what you care about. That assumption broke recently. OpenAI’s audit of SWE-bench Verified found that every frontier model could reproduce verbatim gold patches or problem-statement specifics for some tasks. The same models score ~46% on SWE-Bench Pro versus ~81% on Verified. Hamel Husain’s playbook — build internal evals from production traces — is now the only reliable pattern at the frontier.
The honest signal layer for harness optimization is what reaches the customer. Not commits authored, not PRs created, not synthetic task pass rates. What shipped, what required rework, where bottlenecks ate capacity, where the team’s interaction graph cleaved at the wrong points. Anchoring on outcome telemetry changes which agents get flagged, changes how recommendations are prioritized, and anchors a tighter L4 maturity definition: not automated drift detection on synthetic eval scores, but automated drift detection on customer-reaching delivery.

How we run it in practice
The methodology, reduced to seven operational moves:
- Pick the outcome telemetry layer. Leanmote, your DORA dashboard, your task-tracker analytics — the signals you want are rework rate, cycle-time anomalies, review concentration, collaboration breakdowns. Not commit count.
- Score the harness across five dimensions: Context, Versioning, Governance, Drift Detection, Distribution. Four rungs (L1 Artisanal → L4 Industrial). Take min(dim levels) as team score — a single weak link gates the setup.
- Score each agent on six operational axes: Context Retrieval, Instruction Engineering, Context Packing, Grounding/Verification, Output Feedback Loop, Collaboration. Three states: healthy / degraded / broken.
- Run a structured interview against each agent. Closed-stem questions for parsing, open-tail questions for unknown-unknowns. Three respondent modes: self-mode (agent answers), proxy-mode (peer agent reads its substrate), manager-mode (human owner answers).
- Triangulate channels. Outcome telemetry is strong-confidence. Interview is medium, upgraded on corroboration. Uploaded harness artifacts (AGENTS.md, prompt exports) are strong on Context and Versioning.
- Generate a punch list. Tag each finding Impact × Effort. Apply L+1 gating to dimension-level moves — no skipping maturity levels. Compounding pairs get higher weight.
- Split output by audience: manager-facing diagnostic with action plan, bot-facing self-heal procedure with backup + verification + reporting, full evidence trail for audit.
No optimizer. No ML infrastructure. A Jira project, the agent’s filesystem access, and a few days. The Markdown form of the methodology can be executed by any competent agent end-to-end.

A real diagnostic at Leanmote
Before the diagnostic was possible, the fleet had to be built. The Orchestrator–Developer–QA configuration emerged from a test-lab phase: a series of MVPs run through a paperclip tool — a minimal harness scaffold used to probe agent handoff behavior, delegation payload fidelity, and early context assumptions before committing to a production workload. That inception phase is where the first AGENTS.md conventions took shape, where inter-agent handoff slots were stress-tested, and where the earliest structural gaps first appeared. The diagnostic described here ran on the fleet after it had graduated from that early harness phase.
We ran this on our internal 3-agent dev fleet — Orchestrator, Developer, QA. The Orchestrator picks up Jira tickets, evaluates them, delegates implementation to the Developer, routes the Developer’s PRs through QA, and either ships to human review or re-delegates with feedback. A Leanmote anomaly had been firing for weeks: 17% of merged PRs were being iterated post-review. Nobody had a confident root-cause read. The hypothesis was “the Developer agent isn’t grounded enough; we need better prompts.”
Four interview rounds, four days, all async via Jira. ~70 trigger firings across the three agents. The 17% rate decomposed into five compounded structural failures — none of which prompt-tuning would have addressed:
- The Orchestrator’s delegation payload to Developer was referential (file paths, doc references) but lacked visual specs and negative examples — a citation engineering gap.
- The Developer’s AGENTS.md had no pre-handoff verification step. Straight from “implement” to “push + handoff” — an evidence-anchored content gap.
- The Developer literally couldn’t do visual verification — no headed browser, no MCP browser tools. Verification was structurally deferred to QA — a tooling gap that made the previous row unfixable from prompts alone.
- QA’s verification depth was single-layer (acceptance-criteria pass = approve). One canonical false-approve case: a ticket where QA verified tooltip popovers rendered with “Learn more” links but didn’t click through every link to verify destinations.
- No agent had cross-ticket memory. The same bug pattern caught twice never made it into an anti-patterns.md — a visibility-and-drift gap.

The 17% rate was the symptom of a verification pipeline missing layers 1–3, with a depth gap on layer 4 and a drift-monitoring absence on layer 5. Structural, not competence. The diagnostic also surfaced something the team hadn’t articulated: the fleet scored L1 maturity, gated by Governance (no review cadence on the Orchestrator’s PRs), Drift Detection (no aggregated quality dashboard), and Versioning (no .git anywhere in the agent directory). The agents had been deployed on a substrate nobody had thought of as needing software-engineering rigor — because it didn’t look like software, it looked like prompts. The output of the diagnostic is a prioritized punch list, ranked Impact × Effort, with each item rendered as a single decision-ready card.

Why this matters
For CTOs operating AI agents in production: the harness-optimization tooling becoming available is real and useful. But what signal you optimize against is upstream of which optimizer you use. If your eval suite is synthetic or inherited from a benchmark that has leaked into training data, you’re optimizing against a noisy proxy. Build internal evals from production traces. If you operate on top of an outcome-telemetry platform, anchor harness optimization on that signal layer. The translation table is the framework for thinking about what to change; outcome telemetry is the input layer that tells you which row matters most this quarter.
The methodology already exists, runs on Leanmote signal plus structured Jira interviews, no specialized tooling. We’re building the platform on top of it — a tool that ingests your team’s outcome telemetry and your agent’s harness files, runs the diagnostic without an analyst, generates the punch list, and works directly with your agent through the proxy-mode pattern to apply fixes that don’t need human authorization. If the iterated-PRs / inconsistent-output / hard-to-attribute-rework problem sounds familiar, we’re starting an early-access cohort — get in touch. The agent is fine. The harness is the work. The framing you bring to the harness — and the input signal you anchor it on — is the work upstream of the work.
References
- Anthropic — Building effective agents
- Anthropic — How we built our multi-agent research system
- OpenAI — Introducing SWE-bench Verified
- Phil Schmid — Notes on agent harnesses, Google DeepMind
- Martin Fowler — Articles on AI agents and engineering practice
- Atlan — 2026 guide to AI data and agent platforms
- Hamel Husain — Building production evals from real traces
- Addy Osmani — Agentic Engine Optimization (AEO)
- HARBOR — Constrained Bayesian optimization over harness configuration
- Meta-Harness — Outer-loop coding agent for harness synthesis
- SWE-Bench Pro — Augmented benchmark with reduced contamination
- Gartner — Forecast: organic-search decline driven by generative engines, 2026