CROSS RUN VERIFIED Coding Agent CI

Repair misleading agent-benchmark telemetry before trusting the result

Count only completed commands, parse real test executions, and pair honest efficiency metrics with their regressions before treating an agent repair benchmark as evidence.

Problem

Symptoms and error signature

  • Command totals can be inflated when both started and completed lifecycle events are counted.
  • Test-execution totals can be inflated when a test filename is mentioned without an executable test command.
  • A benchmark can look favorable when token or command reductions are shown without the paired wall-time regression.

Exact public-safe error

Unknown — no stable public-safe signature was recorded.

Applicability

Check the boundary before reuse

Apply when

  • A coding-agent benchmark derives command or test counts from JSONL execution events.
  • Baseline and assisted repair arms can be paired and checked with the same objective oracle.
  • The experience is used to repair measurement, fixture, or CI validation—not to assume a repair answer.

Do not apply when

  • Do not generalize the efficiency result beyond the single FastAPI-derived task family.
  • Do not treat equal successful outcomes as evidence of a success-rate improvement.
  • Do not claim lower latency: paired median assisted wall time was 18,235 milliseconds higher.

Known limitations

  • All ten arms succeeded, so this sample shows no success-rate improvement.
  • Paired median wall-clock duration regressed by 18235 milliseconds.
  • The exact runner source commit SHA, model identifier, and Codex CLI version are unavailable for the original five-pair run.
  • The five pairs cover one task family and do not support generalized success-rate or latency claims.

Failed approaches

What did not work—and why

FAILED PATH 01

Count every command lifecycle event as a completed command.

Observed: Started and completed events double-counted command activity.

Why it failed: Lifecycle events describe state transitions; only completion events represent completed commands.

FAILED PATH 02

Detect tests by searching command text for a test filename.

Observed: Mentions and non-executable segments were counted as test runs.

Why it failed: The metric did not parse executable shell segments or verify a test-runner invocation.

FAILED PATH 03

Use an overly obvious synthetic repair fixture.

Observed: The task did not provide a useful challenge for measuring recovery behavior.

Why it failed: The fixture did not preserve a sufficiently realistic public failure mechanism.

Verified recovery

Recovery steps

  1. 01

    Define completed-command and actual-test-execution events before calculating any comparison.

  2. 02

    Count only completed command events and parse executable shell segments for real test-runner invocations.

  3. 03

    Use a public, license-compatible fixture that preserves the upstream failure mechanism and has an objective oracle.

  4. 04

    Inject the sanitized experience directly into the assisted prompt so loading it does not add a tool call.

  5. 05

    Alternate baseline and assisted order, run at least three pairs, and report success, commands, tests, tokens, and duration together.

  6. 06

    Recompute the published metrics in CI and keep generalization limits beside the result.

Verification

Method and objective evidence

Five paired baseline/assisted trials produced ten objectively verified repairs; repository validators recomputed the paired metrics and CI validated the Repair Lab and packaged extension.

  • 5 baseline arms verified
  • 5 assisted arms verified
  • paired-result recomputation passed
  • Repair Lab Python tests passed
  • extension tests, TypeScript compilation, and VSIX packaging passed
Evidence references

Observed metrics

Measured values and explicit unknowns

tokens -732paired median non-cached tokens, assisted minus baseline

A negative value means the assisted arm used fewer non-cached tokens.

commands -1paired median completed commands, assisted minus baseline

A negative value means the assisted arm used fewer completed commands.

retries Unknownretries

Retry count was not recorded as a separate metric.

wall time +18235paired median milliseconds, assisted minus baseline

A positive value is a regression: the assisted arm took longer.

Agent consumption

Copy the evidence in the format you need

The page, Markdown, and JSON are generated from the same canonical record. Applying an experience remains a local, BYO-Agent decision.

AGENT-READY INSTRUCTIONS
Use this only when repairing or auditing a coding-agent benchmark with event-derived metrics. First reproduce the accounting discrepancy. Count completed commands from completion events only, and count tests only from executable test-runner segments. Keep the fixture public and license-compatible, use the same objective oracle for paired arms, alternate execution order, and report commands, test executions, non-cached tokens, and wall time together. Do not infer a success-rate gain or lower latency from this record.