Repair misleading agent-benchmark telemetry before trusting the result
Count only completed commands, parse real test executions, and pair honest efficiency metrics with their regressions before treating an agent repair benchmark as evidence.
Problem
Symptoms and error signature
- Command totals can be inflated when both started and completed lifecycle events are counted.
- Test-execution totals can be inflated when a test filename is mentioned without an executable test command.
- A benchmark can look favorable when token or command reductions are shown without the paired wall-time regression.
Exact public-safe error
Unknown — no stable public-safe signature was recorded.
Applicability
Check the boundary before reuse
Apply when
- A coding-agent benchmark derives command or test counts from JSONL execution events.
- Baseline and assisted repair arms can be paired and checked with the same objective oracle.
- The experience is used to repair measurement, fixture, or CI validation—not to assume a repair answer.
Do not apply when
- Do not generalize the efficiency result beyond the single FastAPI-derived task family.
- Do not treat equal successful outcomes as evidence of a success-rate improvement.
- Do not claim lower latency: paired median assisted wall time was 18,235 milliseconds higher.
Known limitations
- All ten arms succeeded, so this sample shows no success-rate improvement.
- Paired median wall-clock duration regressed by 18235 milliseconds.
- The exact runner source commit SHA, model identifier, and Codex CLI version are unavailable for the original five-pair run.
- The five pairs cover one task family and do not support generalized success-rate or latency claims.
Failed approaches
What did not work—and why
Count every command lifecycle event as a completed command.
Observed: Started and completed events double-counted command activity.
Why it failed: Lifecycle events describe state transitions; only completion events represent completed commands.
Detect tests by searching command text for a test filename.
Observed: Mentions and non-executable segments were counted as test runs.
Why it failed: The metric did not parse executable shell segments or verify a test-runner invocation.
Use an overly obvious synthetic repair fixture.
Observed: The task did not provide a useful challenge for measuring recovery behavior.
Why it failed: The fixture did not preserve a sufficiently realistic public failure mechanism.
Verified recovery
Recovery steps
- 01
Define completed-command and actual-test-execution events before calculating any comparison.
- 02
Count only completed command events and parse executable shell segments for real test-runner invocations.
- 03
Use a public, license-compatible fixture that preserves the upstream failure mechanism and has an objective oracle.
- 04
Inject the sanitized experience directly into the assisted prompt so loading it does not add a tool call.
- 05
Alternate baseline and assisted order, run at least three pairs, and report success, commands, tests, tokens, and duration together.
- 06
Recompute the published metrics in CI and keep generalization limits beside the result.
Verification
Method and objective evidence
Five paired baseline/assisted trials produced ten objectively verified repairs; repository validators recomputed the paired metrics and CI validated the Repair Lab and packaged extension.
- 5 baseline arms verified
- 5 assisted arms verified
- paired-result recomputation passed
- Repair Lab Python tests passed
- extension tests, TypeScript compilation, and VSIX packaging passed
Observed metrics
Measured values and explicit unknowns
A negative value means the assisted arm used fewer non-cached tokens.
A negative value means the assisted arm used fewer completed commands.
Retry count was not recorded as a separate metric.
A positive value is a regression: the assisted arm took longer.
Agent consumption
Copy the evidence in the format you need
The page, Markdown, and JSON are generated from the same canonical record. Applying an experience remains a local, BYO-Agent decision.
Use this only when repairing or auditing a coding-agent benchmark with event-derived metrics. First reproduce the accounting discrepancy. Count completed commands from completion events only, and count tests only from executable test-runner segments. Keep the fixture public and license-compatible, use the same objective oracle for paired arms, alternate execution order, and report commands, test executions, non-cached tokens, and wall time together. Do not infer a success-rate gain or lower latency from this record.
# Repair misleading agent-benchmark telemetry before trusting the result
- Experience ID: `trace-2026-08-03-repair-lab-ci-v0.1.3`
- Verification: `CROSS_RUN_VERIFIED`
- Category: Coding Agent CI
- Last verified: 2026-08-03T04:15:23Z
Count only completed commands, parse real test executions, and pair honest efficiency metrics with their regressions before treating an agent repair benchmark as evidence.
## Symptoms
- Command totals can be inflated when both started and completed lifecycle events are counted.
- Test-execution totals can be inflated when a test filename is mentioned without an executable test command.
- A benchmark can look favorable when token or command reductions are shown without the paired wall-time regression.
## Exact error signatures
- Unknown: no stable public-safe error signature was recorded.
## Apply when
- A coding-agent benchmark derives command or test counts from JSONL execution events.
- Baseline and assisted repair arms can be paired and checked with the same objective oracle.
- The experience is used to repair measurement, fixture, or CI validation—not to assume a repair answer.
## Do not apply when
- Do not generalize the efficiency result beyond the single FastAPI-derived task family.
- Do not treat equal successful outcomes as evidence of a success-rate improvement.
- Do not claim lower latency: paired median assisted wall time was 18,235 milliseconds higher.
## Known limitations
- All ten arms succeeded, so this sample shows no success-rate improvement.
- Paired median wall-clock duration regressed by 18235 milliseconds.
- The exact runner source commit SHA, model identifier, and Codex CLI version are unavailable for the original five-pair run.
- The five pairs cover one task family and do not support generalized success-rate or latency claims.
## Environment and agent context
- Language: Python and TypeScript
- Runtime: Exact Python, Node.js, and Codex CLI versions for the original five-pair run are unknown.
- Operating system: unknown
- Dependencies: FastAPI-derived dependency-free repair fixture, GitHub Actions validation, VS Code extension packaging
- Agent: Codex
- Model: unknown
- Harness: codex-exec paired Repair Lab runner
- Reasoning context: unknown
## Failed approaches
1. **Count every command lifecycle event as a completed command.**
- Observed: Started and completed events double-counted command activity.
- Why it failed: Lifecycle events describe state transitions; only completion events represent completed commands.
2. **Detect tests by searching command text for a test filename.**
- Observed: Mentions and non-executable segments were counted as test runs.
- Why it failed: The metric did not parse executable shell segments or verify a test-runner invocation.
3. **Use an overly obvious synthetic repair fixture.**
- Observed: The task did not provide a useful challenge for measuring recovery behavior.
- Why it failed: The fixture did not preserve a sufficiently realistic public failure mechanism.
## Verified recovery
1. Define completed-command and actual-test-execution events before calculating any comparison.
2. Count only completed command events and parse executable shell segments for real test-runner invocations.
3. Use a public, license-compatible fixture that preserves the upstream failure mechanism and has an objective oracle.
4. Inject the sanitized experience directly into the assisted prompt so loading it does not add a tool call.
5. Alternate baseline and assisted order, run at least three pairs, and report success, commands, tests, tokens, and duration together.
6. Recompute the published metrics in CI and keep generalization limits beside the result.
## Verification
Five paired baseline/assisted trials produced ten objectively verified repairs; repository validators recomputed the paired metrics and CI validated the Repair Lab and packaged extension.
- 5 baseline arms verified
- 5 assisted arms verified
- paired-result recomputation passed
- Repair Lab Python tests passed
- extension tests, TypeScript compilation, and VSIX packaging passed
## Metrics
- Tokens: -732 paired median non-cached tokens, assisted minus baseline — A negative value means the assisted arm used fewer non-cached tokens.
- Commands: -1 paired median completed commands, assisted minus baseline — A negative value means the assisted arm used fewer completed commands.
- Retries: Unknown retries — Retry count was not recorded as a separate metric.
- Wall Time: 18235 paired median milliseconds, assisted minus baseline — A positive value is a regression: the assisted arm took longer.
## Agent-ready instructions
Use this only when repairing or auditing a coding-agent benchmark with event-derived metrics. First reproduce the accounting discrepancy. Count completed commands from completion events only, and count tests only from executable test-runner segments. Keep the fixture public and license-compatible, use the same objective oracle for paired arms, alternate execution order, and report commands, test executions, non-cached tokens, and wall time together. Do not infer a success-rate gain or lower latency from this record.
## Provenance
- Source: https://github.com/yao23/agent-experience-graph
- Revision: `dca38fe9db4d463c5bb16f155dc033f8895e1439`
- License: MIT
{
"schema_version": "1.0.0",
"id": "trace-2026-08-03-repair-lab-ci-v0.1.3",
"slug": "repair-lab-telemetry-and-fixture-validation",
"title": "Repair misleading agent-benchmark telemetry before trusting the result",
"summary": "Count only completed commands, parse real test executions, and pair honest efficiency metrics with their regressions before treating an agent repair benchmark as evidence.",
"category": "Coding Agent CI",
"verification_status": "CROSS_RUN_VERIFIED",
"problem": {
"symptoms": [
"Command totals can be inflated when both started and completed lifecycle events are counted.",
"Test-execution totals can be inflated when a test filename is mentioned without an executable test command.",
"A benchmark can look favorable when token or command reductions are shown without the paired wall-time regression."
],
"error_signatures": []
},
"applicability": {
"applies_when": [
"A coding-agent benchmark derives command or test counts from JSONL execution events.",
"Baseline and assisted repair arms can be paired and checked with the same objective oracle.",
"The experience is used to repair measurement, fixture, or CI validation—not to assume a repair answer."
],
"exclusions": [
"Do not generalize the efficiency result beyond the single FastAPI-derived task family.",
"Do not treat equal successful outcomes as evidence of a success-rate improvement.",
"Do not claim lower latency: paired median assisted wall time was 18,235 milliseconds higher."
]
},
"context": {
"repository": "https://github.com/yao23/agent-experience-graph",
"source_revision": {
"version": "0.1.3",
"commit_sha": "dca38fe9db4d463c5bb16f155dc033f8895e1439"
},
"environment_fingerprint": {
"language": "Python and TypeScript",
"runtime": "Exact Python, Node.js, and Codex CLI versions for the original five-pair run are unknown.",
"operating_system": "unknown",
"dependencies": [
"FastAPI-derived dependency-free repair fixture",
"GitHub Actions validation",
"VS Code extension packaging"
]
},
"agent_context": {
"agent": "Codex",
"model": "unknown",
"harness": "codex-exec paired Repair Lab runner",
"reasoning": "unknown"
}
},
"failed_attempts": [
{
"approach": "Count every command lifecycle event as a completed command.",
"observed_result": "Started and completed events double-counted command activity.",
"why_failed": "Lifecycle events describe state transitions; only completion events represent completed commands."
},
{
"approach": "Detect tests by searching command text for a test filename.",
"observed_result": "Mentions and non-executable segments were counted as test runs.",
"why_failed": "The metric did not parse executable shell segments or verify a test-runner invocation."
},
{
"approach": "Use an overly obvious synthetic repair fixture.",
"observed_result": "The task did not provide a useful challenge for measuring recovery behavior.",
"why_failed": "The fixture did not preserve a sufficiently realistic public failure mechanism."
}
],
"recovery_steps": [
"Define completed-command and actual-test-execution events before calculating any comparison.",
"Count only completed command events and parse executable shell segments for real test-runner invocations.",
"Use a public, license-compatible fixture that preserves the upstream failure mechanism and has an objective oracle.",
"Inject the sanitized experience directly into the assisted prompt so loading it does not add a tool call.",
"Alternate baseline and assisted order, run at least three pairs, and report success, commands, tests, tokens, and duration together.",
"Recompute the published metrics in CI and keep generalization limits beside the result."
],
"verification_method": {
"summary": "Five paired baseline/assisted trials produced ten objectively verified repairs; repository validators recomputed the paired metrics and CI validated the Repair Lab and packaged extension.",
"checks": [
"5 baseline arms verified",
"5 assisted arms verified",
"paired-result recomputation passed",
"Repair Lab Python tests passed",
"extension tests, TypeScript compilation, and VSIX packaging passed"
],
"evidence_refs": [
"experiments/public-repair-lab/results/v0.1.3-paired-results.json",
"experiments/public-repair-lab/validate_paired_results.py",
".github/workflows/repair-lab.yml"
]
},
"registry_metrics": {
"tokens": {
"value": -732,
"unit": "paired median non-cached tokens, assisted minus baseline",
"status": "measured",
"note": "A negative value means the assisted arm used fewer non-cached tokens."
},
"commands": {
"value": -1,
"unit": "paired median completed commands, assisted minus baseline",
"status": "measured",
"note": "A negative value means the assisted arm used fewer completed commands."
},
"retries": {
"value": null,
"unit": "retries",
"status": "unknown",
"note": "Retry count was not recorded as a separate metric."
},
"wall_time": {
"value": 18235,
"unit": "paired median milliseconds, assisted minus baseline",
"status": "measured",
"note": "A positive value is a regression: the assisted arm took longer."
}
},
"last_verified_at": "2026-08-03T04:15:23Z",
"agent_ready_instructions": "Use this only when repairing or auditing a coding-agent benchmark with event-derived metrics. First reproduce the accounting discrepancy. Count completed commands from completion events only, and count tests only from executable test-runner segments. Keep the fixture public and license-compatible, use the same objective oracle for paired arms, alternate execution order, and report commands, test executions, non-cached tokens, and wall time together. Do not infer a success-rate gain or lower latency from this record.",
"license": "MIT",
"task": "Make the AEG Public Repair Lab repeatable, trustworthy, and continuously validated",
"outcome": "success",
"subtasks": [
{
"description": "Repair experiment telemetry before interpreting the A/B result",
"skills": [
"evaluation-harness",
"event-stream-analysis"
],
"tools": [
"python",
"codex-jsonl"
],
"outcome": "success",
"lessons": [
"Count only completed command events; started and completed lifecycle events otherwise double the metric.",
"Count test executions by parsing executable shell segments, not by searching for a test filename substring."
]
},
{
"description": "Replace an overly obvious repair fixture with a harder public, licensed failure",
"skills": [
"benchmark-design",
"license-aware-fixture-design"
],
"tools": [
"BugsInPy",
"FastAPI",
"unittest"
],
"outcome": "success",
"lessons": [
"A fixture should preserve the upstream failure mechanism while remaining dependency-free and objectively verifiable.",
"Record repository, buggy commit, fixed commit, license, and benchmark cross-reference next to the fixture."
]
},
{
"description": "Deliver retrieved experience without adding artificial tool overhead",
"skills": [
"experience-retrieval-design",
"prompt-injection-design"
],
"tools": [
"codex-exec"
],
"outcome": "success",
"lessons": [
"Inject a compact sanitized recovery capsule directly into the assisted prompt instead of requiring an agent to open a separate experience file.",
"Treat retrieved experience as guidance that must match local evidence, not as an unverified answer."
]
},
{
"description": "Run paired trials and report both improvements and regressions",
"skills": [
"paired-experiment-design",
"cost-analysis"
],
"tools": [
"python",
"codex-exec",
"git"
],
"outcome": "success",
"lessons": [
"Alternate baseline and assisted execution order to reduce simple warm-cache and first-run bias.",
"Require at least three paired trials before assigning an efficiency verdict.",
"Preserve success, commands, test runs, tokens, duration, patches, and raw events; do not hide latency regressions behind token or command wins."
]
},
{
"description": "Continuously validate the Repair Lab and packaged extension in GitHub Actions",
"skills": [
"github-actions",
"release-validation"
],
"tools": [
"GitHub Actions",
"python",
"npm",
"TypeScript",
"VSCE"
],
"outcome": "success",
"lessons": [
"CI should validate telemetry tests, prepare every fixture, run extension tests, compile TypeScript, and package the VSIX without launching paid agent trials.",
"A successful workflow run is promotion evidence for the experience, not proof that the experiment generalizes to other task families."
]
}
],
"skills": [
"evaluation-harness",
"benchmark-design",
"experience-retrieval-design",
"paired-experiment-design",
"github-actions",
"release-validation"
],
"tools": [
"python",
"codex-exec",
"git",
"GitHub Actions",
"npm",
"TypeScript",
"VSCE"
],
"constraints": [
"Do not publish raw JSONL, stderr logs, local patches, credentials, or private workspace data.",
"Keep agent trials isolated, ephemeral, workspace-scoped, and free of upstream repository side effects.",
"Do not claim general success-rate or latency improvement from one task family."
],
"lessons": [
"Measurement integrity comes before optimization: incorrect event accounting can make a neutral result look favorable.",
"Retrieval delivery is part of system cost; avoid charging an extra tool call merely to load the recommendation.",
"Verified negative evidence is reusable: the five-pair sample reduced median commands and paired non-cached tokens but regressed paired wall time.",
"Store compact public experiences with provenance and explicit limitations so another agent can reuse the method without ingesting raw execution logs."
],
"provenance": {
"repository": "https://github.com/yao23/agent-experience-graph",
"sourceVersion": "0.1.3",
"recordedAt": "2026-08-03T04:15:23Z",
"publicSource": {
"repository": "https://github.com/fastapi/fastapi",
"buggyCommitSha": "7cea84b74ca3106a7f861b774e9d215e5228728f",
"fixedCommitSha": "75a07f24bf01a31225ee687f3e2b3fc1981b67ab",
"license": "MIT",
"benchmark": "BugsInPy fastapi bug 5"
},
"experimentEvidence": {
"artifact": "experiments/public-repair-lab/results/v0.1.3-paired-results.json",
"sourceReportCreatedAt": "2026-08-03T00:38:20.080621+00:00",
"runnerSourceCommitSha": null,
"runnerSourceCommitStatus": "unavailable: trials ran before the runner changes were committed"
},
"promotionEvidence": {
"workflowName": "Validate Repair Lab",
"workflowFile": ".github/workflows/repair-lab.yml",
"validatedCommitSha": "dca38fe9db4d463c5bb16f155dc033f8895e1439",
"status": "observed-passed",
"runResolver": "https://github.com/yao23/agent-experience-graph/actions/runs/30932880306"
},
"publication": {
"pullRequest": "https://github.com/yao23/agent-experience-graph/pull/6"
}
},
"verification": {
"status": "passed",
"evidence": {
"experimentArtifact": "experiments/public-repair-lab/results/v0.1.3-paired-results.json",
"experienceSchema": "experiences/verified-experience.schema.json",
"semanticValidator": "scripts/validate_verified_experiences.py",
"pairedResultsValidator": "experiments/public-repair-lab/validate_paired_results.py"
},
"localChecks": {
"repairLabPythonTests": "passed",
"retrievalTests": "passed",
"verifiedExperienceValidation": "passed",
"pairedResultRecomputation": "passed",
"extensionTests": "passed",
"fixturesPrepared": [
"fastapi-nested-response",
"pysnooper-path-output"
],
"typescriptCompilation": "passed",
"vsixPackaging": "passed"
}
},
"metrics": {
"pairedTrials": 5,
"baselineArms": 5,
"assistedArms": 5,
"baselineArmsVerified": 5,
"assistedArmsVerified": 5,
"pairedMedianAssistedMinusBaseline": {
"completedCommands": -1,
"actualTestExecutions": 0,
"nonCachedTokens": -732,
"durationMs": 18235
},
"interpretation": "Bounded tool-cycle and token improvement on one task family; no success-rate improvement and wall-clock latency regressed."
},
"limitations": [
"All ten arms succeeded, so this sample shows no success-rate improvement.",
"Paired median wall-clock duration regressed by 18235 milliseconds.",
"The exact runner source commit SHA, model identifier, and Codex CLI version are unavailable for the original five-pair run.",
"The five pairs cover one task family and do not support generalized success-rate or latency claims."
],
"reuse": {
"retrievalTags": [
"agent evaluation",
"A/B repair experiment",
"telemetry correctness",
"paired trials",
"experience injection",
"GitHub Actions",
"VS Code extension release"
],
"recommendedFor": [
"building an auditable agent benchmark",
"repairing duplicated JSONL event metrics",
"adding CI for agent experiment fixtures",
"publishing a verified reusable experience"
]
}
}