Learning Detail

← Back to game
generation --

Strategy learning per pass

Every strategy is scored per objective: observed change minus the change its doctrine expected. The expectation itself learns from completed uses -- "learned" is what it actually does, "authored" is only where it started. A residual is the gap between the two.

Strategy Situations Objective Uses Successes Authored Learned Mean residual Solo residual Recent residuals

Strategy interactions per pass

Pairs of strategies whose five-pass windows overlapped: what the pair achieved together against the sum of both strategies' own expectations.

Strategy With Objective Co-uses Mean observed Mean expected Mean residual

Plain-action baselines per pass

What each ordinary action, with no strategy chosen, has achieved in each active situation's objectives -- the yardstick a strategy's learned expectation is really being compared against.

Situation Action Objective Uses Mean observed

Stall recognition per pass

A stall is a gap since the last objective improvement longer than 90% of the gaps play has recorded. The threshold is learned, not set.

Advice standing per session

Every monster tactic and doctrine line Jev is shown carries its record inline, as "[shown N times: X died, Y survived]" -- plain counts, never a weight or ranking. Loaded once at session start; only changes once a run ends. Text is recovered from the current knowledge base and matched by the same hash the payload itself uses, so a record whose advice has since been edited or removed shows its counts with no text rather than a guess.

Advice Source Shown Outcomes

Learned lessons per session

A death pattern (same monster or hazard, at least twice) becomes a sentence Jev is actually shown. Below it, the raw deaths it was built from.

    Recent mistakes (source material for the lessons above)

      Bandit posteriors per session

      Which value of each tunable threshold (e.g. how low HP has to fall before a flee rule's advice fires) is sampled at session start, weighted by which values have produced better-scoring runs so far.

      Promoted interpretations per pass

      Pilot (2026-09-23): one hand-written derivation rule, proving the full lifecycle -- derive, corroborate, promote, contest -- end to end. Corroboration is a for/against count pair, never a single confidence number. Once promoted and matching the active situation, the claim reaches Jev's actual payload, not just this dashboard.

      Claim Status Situations For Against Promoted

      Whether it's working

      These read accumulated data and report on it. None of them change what Jev is offered or told -- they answer "is this getting better," not "what did the agent just learn."

      Coverage trend

      Of the item/monster/hazard kinds actually encountered in a run, what fraction match an entry in the hand-authored knowledge base. Not a measure of how much of the game is known -- the denominator is only what showed up that run, so a quiet run and a catalog gap can both read 100%.

      Reward-shaping weights

      Dimension Default Current Note

      Learning measurement

      Three separate claims -- a growing store is not the same as better play. Read together, not as one score.

      Accumulation

      Reachability

      Contributions by goal

      Outcome improvement

      Knowledge ledger static

      Not learned. This is the hand-authored reference catalog itself and how much of it has been transcribed and verified -- a documentation audit, not a report on play.

        Knowledge discovered

        First-time-seen items, monsters, and hazards -- a log, not itself read back into any decision.

          Run history

          # Started Outcome Depth Turns Gold Amulet Escaped Kills Score Cost

          Past training generations

            Built, not reaching anything

            Mechanisms whose write path, read path, or both are disabled or fail-closed by design -- with the real reason, not silence. Counted live below rather than asserted in prose, so an entry here corrects itself the day it's re-wired, the way the last two did.