Agent Experience Intelligence · 5-minute brief

Verified experience for the next agent run.

AEG turns real execution into reusable guidance—and tests whether that guidance actually improves the next task.

01 / Repeated-learning problem

Every new run starts too close to zero.

Teams keep code, traces, and documentation. They rarely keep a trustworthy account of which recovery worked, under what conditions, and whether reusing it helped.

OBSERVABILITY

Shows what happened

Useful traces do not decide what should guide the next run.

DOCUMENTATION

Says what should happen

General guidance rarely carries run-level outcome evidence.

AEG

Tests what transfers

Unlike generic memory, AEG keeps provenance and a verified outcome attached, then tests whether the experience should guide the next run.

02 / Operating loop

Turn execution into evidence, then test the reuse.

The cross-platform layer sits beside agents, models, tools, and observability. It captures a run, sanitizes the reusable signal, verifies the outcome, retrieves relevant experience or abstains, and measures what happens next.

CAPTURE

Run receipt

Attempts, failures, recovery, and outcome.

STRUCTURE

Safe record

Portable fields and sanitized content.

VERIFY

Evidence

Oracle, provenance, applicability, limits.

REUSE

Guidance

Explain relevance—or correctly abstain.

MEASURE

Comparison

Success, commands, tests, tokens, latency.

03 / Unit of learning

The experience record preserves the path, not just the answer.

A verified record is situated evidence. It is useful only when its provenance, applicability, and limitations remain attached.

01Attempt
02Failure
03Recovery
04Verified outcome
05Cost & latency
06Provenance
07Applicability

04 / Initial wedge

Start where transfer can be tested cleanly.

The first tasks are reproducible failures with an objective success check, comparable baseline and assisted runs, and low privacy or IP risk.

NOW

Engineers using coding agents

Real repositories, bounded failures, objective verification.

NEXT

AI dev-tool and OSS maintainers

Repeatable task families where recovery knowledge gets lost.

LATER

Enterprise AI-platform teams

Governed reuse across internal agent workflows—if transfer holds.

Long-term thesis

From Superintelligence to Distributed Intelligence

The future may not be one agent knowing everything. One possible topology is billions of human-agent systems learning locally and sharing selectively—when evidence, permission, and context support the transfer.

MODELS

General capability

What can an intelligent system broadly do?

CONTEXT / RAG

Current task knowledge

What does this task need to know now?

AEG

Verified situated experience

What worked before under comparable conditions?

  • attempts, failures, and recovery
  • verified outcome and cost/latency
  • provenance and applicability

05 / Shipped and measured

Capture, validation, and retrieval are shipped. Benefit remains bounded.

Repository evidence keeps shipped facts, positive results, negative results, and unanswered questions separate.

5 pairsone repair family
10 / 10arms passed
1 fewermedian completed command
+18.2 smedian assisted latency

Median assisted runs used one fewer completed command. This bounded five-pair result did not improve success, and latency regressed.

Claim Status Repository evidence Boundary
AEG can capture, validate, and retrieve experience. Shipped VS Code v0.1.5, schemas, validator, recommender Developer preview; local-first
Verified records can preserve auditable outcomes. Verified Two records in experiences/registry.json Outcome evidence ≠ causal AEG benefit
A bounded repair run used fewer median commands. Bounded Five pairs; all ten arms passed objective verification No success gain; median latency regressed 18,235 ms
Related transfer pairs produced positive retrieval evidence. No One neutral pair; one preregistered result not positive Negative and neutral evidence stays visible
The autonomous loop follows bounded state and approval rules. Validated Hash-chained transitions and approval-gate stop Workflow validation only; not retrieval benefit

06 / Evidence boundary

What the evidence does not say.

  • No demonstrated success-rate improvement
  • No reliable cross-project learning claim
  • No proven reduction in retries, tokens, cost, or latency
  • No customer adoption or product-market-fit evidence
  • No global reputation, popularity, or safety scoring
  • Not a generic marketplace or agent orchestrator
  • No LangSmith replacement claim

07 / Current market status

Stage A outreach is authorized. The public budget remains at zero.

Under the merged Stage A approval record, individually reviewed outreach is authorized for up to three voluntary seed participants. The public recruitment budget currently records 0 invitations and 0 enrolled participants. Task execution and AEG-assisted testing remain unauthorized.

08 / Next falsifiable milestone

Earn one controlled transfer result in a real workflow.

Recruit a small seed cohort, accept only reproducible low-risk tasks, then compare baseline and AEG-assisted runs with retrieval as the intentional difference.

Decision rule

Relevant prior experience must improve at least one preregistered efficiency measure—retries, commands, tests, tokens, or time—without lowering objective task success. Correct abstention, neutral results, and regressions remain evidence.

Design-partner experiment

Bring one reproducible task.

Bring a reproducible task that is public and license-compatible, explicitly authorized, or synthetic, with an objective success check. Any retained evidence will be sanitized. AEG will run a bounded baseline-versus-experience test and return an auditable result.

AEG contributes experiment design and implementation, the comparison, a sanitized record, and a clear result with limits.

09 / Technical architecture

A cross-platform execution-evidence layer.

AEG complements agent runtimes and observability. It connects execution inputs to governed experience records, retrieval with explanations or abstention, and measured outcomes.

Coding agents & IDE workflows
Tasks, traces & test artifacts
Developer feedback
AEGverified experience layer
Evidence schema & provenance
Retrieval, explanation & abstention
Outcome comparison & recapture