Strategy learning per pass
Every strategy is scored per objective: observed change minus the change its doctrine expected. The expectation itself learns from completed uses -- "learned" is what it actually does, "authored" is only where it started. A residual is the gap between the two.
| Strategy | Situations | Objective | Uses | Successes | Authored | Learned | Mean residual | Solo residual | Recent residuals |
|---|
Strategy interactions per pass
Pairs of strategies whose five-pass windows overlapped: what the pair achieved together against the sum of both strategies' own expectations.
| Strategy | With | Objective | Co-uses | Mean observed | Mean expected | Mean residual |
|---|
Plain-action baselines per pass
What each ordinary action, with no strategy chosen, has achieved in each active situation's objectives -- the yardstick a strategy's learned expectation is really being compared against.
| Situation | Action | Objective | Uses | Mean observed |
|---|
Stall recognition per pass
A stall is a gap since the last objective improvement longer than 90% of the gaps play has recorded. The threshold is learned, not set.
Advice standing per session
Every monster tactic and doctrine line Jev is shown carries its record inline, as "[shown N times: X died, Y survived]" -- plain counts, never a weight or ranking. Loaded once at session start; only changes once a run ends. Text is recovered from the current knowledge base and matched by the same hash the payload itself uses, so a record whose advice has since been edited or removed shows its counts with no text rather than a guess.
| Advice | Source | Shown | Outcomes |
|---|
Learned lessons per session
A death pattern (same monster or hazard, at least twice) becomes a sentence Jev is actually shown. Below it, the raw deaths it was built from.
Recent mistakes (source material for the lessons above)
Bandit posteriors per session
Which value of each tunable threshold (e.g. how low HP has to fall before a flee rule's advice fires) is sampled at session start, weighted by which values have produced better-scoring runs so far.
Promoted interpretations per pass
Pilot (2026-09-23): one hand-written derivation rule, proving the full lifecycle -- derive, corroborate, promote, contest -- end to end. Corroboration is a for/against count pair, never a single confidence number. Once promoted and matching the active situation, the claim reaches Jev's actual payload, not just this dashboard.
| Claim | Status | Situations | For | Against | Promoted |
|---|
Whether it's working
These read accumulated data and report on it. None of them change what Jev is offered or told -- they answer "is this getting better," not "what did the agent just learn."
Coverage trend
Of the item/monster/hazard kinds actually encountered in a run, what fraction match an entry in the hand-authored knowledge base. Not a measure of how much of the game is known -- the denominator is only what showed up that run, so a quiet run and a catalog gap can both read 100%.
Reward-shaping weights
| Dimension | Default | Current | Note |
|---|
Learning measurement
Three separate claims -- a growing store is not the same as better play. Read together, not as one score.
Accumulation
Reachability
Contributions by goal
Outcome improvement
Knowledge ledger static
Not learned. This is the hand-authored reference catalog itself and how much of it has been transcribed and verified -- a documentation audit, not a report on play.
Knowledge discovered
First-time-seen items, monsters, and hazards -- a log, not itself read back into any decision.
Run history
| # | Started | Outcome | Depth | Turns | Gold | Amulet | Escaped | Kills | Score | Cost |
|---|
Past training generations
Built, not reaching anything
Mechanisms whose write path, read path, or both are disabled or fail-closed by design -- with the real reason, not silence. Counted live below rather than asserted in prose, so an entry here corrects itself the day it's re-wired, the way the last two did.