Agent Experience Intelligence

Turn real agent runs into verified, reusable experience.

AEG captures what agents tried, what failed, what recovered, and what worked—then turns those runs into evidence-backed guidance for future tasks.

The problem

Agents keep paying the same learning cost.

Execution experience disappears across sessions: useful failures, recovery steps, cost, and validation are rarely captured in a form the next agent can trust.

01 / REDISCOVERY

Start from zero

Similar tasks are decomposed, explored, and debugged again—without access to prior execution lessons.

02 / LOST SIGNAL

Keep the answer, lose the path

Attempts, dead ends, recovery steps, latency, and token cost vanish even when the final patch survives.

03 / WEAK TRUST

Guidance lacks evidence

A plausible tip is not a verified experience. Agents need provenance, constraints, outcomes, and reproducibility.

The operating loop

Make execution evidence useful at the next decision.

AEG focuses first on coding agents. It is an evidence layer—not a generic agent marketplace, an orchestrator, or a replacement for observability platforms.

01 / CAPTURE

Record the run

Intent, context, steps, skills, artifacts, failures, recovery, outcome, and cost.

02 / NORMALIZE

Shape the signal

Sanitize and structure the reusable lesson without raw private work.

03 / VERIFY

Test the outcome

Attach objective checks, limitations, and reproducible provenance.

04 / RECOMMEND

Explain the match

Retrieve relevant lessons and show why they apply—or abstain.

05 / REUSE

Guide the next run

Offer a guarded capsule, then capture the new outcome and feedback.

Evidence, not theater

What exists. What is verified. What is still a hypothesis.

AEG’s public artifacts are useful precisely because their boundaries are visible.

Shipped infrastructure

Working developer surface

  • Local-first VS Code extension v0.1.5
  • Open trace schema and validator
  • Explainable recommender and demo
  • Two records in the verified library
Bounded verification

Auditable, narrow results

  • Five paired trials on one FastAPI repair family
  • All ten arms passed with the same fix
  • Median assisted runs used one fewer completed command.
  • Median assisted wall time regressed by 18,235 ms
Hypothesis under test

Does retrieval improve the next run?

Test whether relevant prior experience reduces retries, token usage, and execution time without lowering task success.

5 pairssingle repair family
10 / 10arms passed verification
1 fewermedian completed command
+18.2smedian wall-time regression

Source: public repair lab results. This does not establish success-rate improvement, generalized transfer, or product-market fit.

What AEG preserves

A compact receipt for agent work.

The reusable unit contains enough evidence to judge a recommendation without retaining raw conversations or proprietary code.

↳

Execution

Intent, context, decomposition, attempted skills and tools.

!

Recovery

Failures, rejected paths, guardrails, and the steps that recovered.

✓

Verification

Outcome, objective oracle, provenance, reproducibility, and limitations.

∑

Cost

Attempts, commands, tests, latency, and tokens when reliably observable.

Next validation milestone

Run a controlled retrieval test on fresh coding tasks.