Release history

Changelog

A release-by-release record of changes to training behavior, orchestration, telemetry, and the dashboard.

125 releases documented

Release 16.4.0

16.4.0

Latest release No runs loaded

My Pokémon run history

  • Add a My Pokémon section to run details showing teammates throughout an attempt, including fainted Pokémon later deposited in the PC.
  • Preserve individual Pokémon history across party switches and evolution, with party/PC status and the recorded location of each Pokémon's faint.
  • Use available starter, party, and catch records for older attempts, clearly identifying missing status and faint locations.

Release 16.3.5

16.3.5

No runs loaded

Pre-Parcel route protection

  • Route every optional wild battle before Parcel delivery through the verified flee controller, with its last-resort attack fallback disabled so encounters cannot become an eight-win grinding loop.
  • Guide the policy down a collision-replayed Route 1 Parcel-return corridor and through the Forest south gate, while synchronizing stale semantic pointers only after the emulator proves each exact route transition.
  • Reject a completed batch before baseline collection when Parcel acquisitions fail to convert into deliveries, including a balanced per-starter check, so a locally regressed optimizer checkpoint cannot become the next resume point.

Release 16.3.4

16.3.4

No runs loaded

Affordability-aware Poké Ball reserves

  • Raise the opening reserve target from five to ten Poké Balls and, on later visits to an audited Mart, scale it to 15, 20, and 25 after badge milestones.
  • Restock only to the quantity current cash can actually cover, and stop a clerk loop after repeated no-progress transactions.
  • Enable the same verified purchase flow in the historically observed Pewter Mart, allowing another restock before Route 3 and Mt. Moon encounters.

Release 16.3.3

16.3.3

No runs loaded

Restock-safe Mart continuity

  • Restore unrestricted Viridian Mart entry so a depleted party can legally restock Poké Balls after its Route 1 encounter.
  • Convert only clerk interactions at or above the configured Poké Ball target into policy-owned idle actions, discouraging the repeated shopping loop without blocking legitimate purchases.

Release 16.3.2

16.3.2

No runs loaded

Post-encounter Route 2 continuity

  • Align stale post-Parcel waypoint guidance to the Viridian return phase only after the emulator proves the Route 1 encounter was spent and the exact northbound Route 1 transition was crossed.
  • Block pre-Brock Mart re-entry after that return, preventing the policy from replaying its learned shopping loop instead of taking Viridian's north exit.

Release 16.3.1

16.3.1

No runs loaded

Pre-Forest level-cap guard

  • Escape optional wild battles once a party member reaches one level below the active badge cap, while always preserving a legal policy-requested catch.
  • Stop wild victories from renewing the continuous-run progress lease and clamp a level-cap failure's complete episode return to a negative value.
  • Close only Viridian's southbound Route 1 transition after the catch loop has returned to town, leaving the required initial trip and later badge-gated travel available.

Release 16.3.0

16.3.0

No runs loaded

Balanced 12-worker full-run collection

  • Add a versioned ordinary-lane worker-pool contract with twelve training environments allocated 4/4/4 across Bulbasaur, Charmander, and Squirtle.
  • Preserve the 512-step recurrent horizon and 768-sample PPO minibatch while increasing each rollout from 4,608 to 6,144 transitions.
  • Keep in-flight nine-worker batches reconcilable, require new batches to bind the 12-worker config and contract digests, and provide an isolated one-rollout checkpoint-resume canary before production activation.

Release 16.2.0

16.2.0

No runs loaded

Parcel-delivery route continuity

  • Treat verified removal of Oak's Parcel in his Lab as authoritative route evidence, monotonically synchronizing stale geometric guidance to the post-delivery phase without issuing inputs or granting skipped rewards.
  • Require each starter cohort to demonstrate at least three verified deliveries and 75% delivery-to-post-delivery alignment before curriculum promotion, isolated to fresh 16.2 evidence.

Release 16.1.0

16.1.0

No runs loaded

Battle-menu synchronization and checkpoint safety

  • Require matching battle RAM and framebuffer evidence before issuing a wild- battle command, and restore a neutral input edge while following cursor transitions so a stale pre-battle cursor cannot misroute RUN.
  • Retain invalid flee-synchronization outcomes in the incremental health cursor, fail the active batch's gate immediately when one occurs, and cancel any semantic checkpoint promotion when that gate rolls back.
  • Bind health reconciliation to the completed batch's proven release instead of the scheduler's newly installed release, preserving correct validation across an in-flight minor-version upgrade.

Operational ownership and diagnostics

  • Add deterministic replay and trace support for bounded failure capsules; the captured flee-synchronization capsule now completes all 94 requested steps without a synchronization abort.
  • Project full-run status from the ordinary cycle owner, report the historical V20 guardian as a non-owner, and correct the local status-job manifest to use that operational source of truth.

Bounded training observability

  • Add a compact incremental episode index with general waypoint conversion, per-starter story funnels, first-hit timing, 64/128/512-episode statistical previews, and ranked failure fingerprints so routine diagnosis no longer rescans multi-gigabyte episode history.
  • Add bounded failure flight recorders paired with rolling diagnostic-only PyBoy checkpoints, sparse full policy action distributions, and an isolated capsule replay helper.
  • Attribute PPO rewards and computed advantages by story phase, starter, and policy/controller/rejection ownership without changing the rollout buffer or optimizer.
  • Report Cloudflare spool pressure, queue age, uploader freshness, and stable upload-failure categories; divert sampled traces into a bounded, restart-recoverable local emergency ring when the spool is saturated instead of silently dropping them.

Release 16.0.0

16.0.0

No runs loaded

Opening rival training continuity

  • Treat exactly one fully proven Oak's Lab opening-rival whiteout as a training-only, nonterminal event: omit its faint and blackout penalties, keep the starter out of the permanent-death ledger, advance the ordinary defeat dialogue and heal/warp, then resume the same episode.
  • Require the exact map, waypoint, first-battle state, single level-5 base starter, matching level-5 rival starter, zero badges and wins, and pre-Parcel story state; every near miss and every later death retains normal Nuzlocke penalties and termination.
  • Disable the exception in held-out evaluation and record dedicated ignored- loss telemetry so evaluation outcomes and training continuation remain independently auditable.

Release 15.1.9

15.1.9

No runs loaded

Viridian northbound and flee synchronization

  • Align the post-capture Viridian north-exit waypoint with the collision-safe x=18 approach and accept both legitimate north warp tiles, preventing a Route 2 transition from being entered without advancing the ordered route.
  • Normalize the wild-battle command cursor from exact live RAM after every bounded movement, recovering a dropped directional input while submitting RUN only after its exact cursor position is proven.

Release 15.1.8

15.1.8

No runs loaded

Forest and wild-battle controller recovery

  • Route Viridian Forest's southwest pocket through a collision-replayed south-gate suffix instead of snapping toward a geometrically close waypoint across an impassable wall, while retaining narrow branch selection around wall-separated starts.
  • Re-confirm a dropped wild-battle RUN command once with the robust button hold only when live menu RAM proves the cursor is still exactly on RUN, then retain the existing bounded fail-closed synchronization abort if the retry is not accepted.

Release 15.1.7

15.1.7

No runs loaded

Viridian Route 22 detour guard

  • Mask the four westbound Viridian-to-Route-22 warp tiles until eight badges, matching the existing Route 22 frontier phase while leaving the northbound Route 2 objective and late-game return fully policy-owned.
  • Attribute blocked westbound requests to the policy, apply the existing no-op/repeat shaping, and retain row-specific gate counts in episode telemetry so boundary stalls remain measurable.

Release 15.1.6

15.1.6

No runs loaded

Stall recovery and tactical intent

  • Record whether each unsafe opening-rival controller attack aligns with the policy's sampled attack head while retaining legal low-HP finisher behavior.
  • Use four Growls for Bulbasaur against the opening Charmander after the exact 999,972-step continuation's canonical-bedroom panel won 30/36 versus 28/36 for two Growls, improving Bulbasaur from 8/12 to 10/12 with no other-starter, item-legality, or move-execution regression.
  • Keep mandatory post-faint party replacement active when proactive low-HP switching is disabled, preventing recovery from looping indefinitely on a forced-switch menu instead of selecting a living teammate.
  • Report explicit reasons for masked Potion, trainer-catch, and trainer-item confirmation requests, and distinguish policy-aligned unsafe rival attacks from unrequested forced attacks without changing the ten-action checkpoint.

Route 2 continuation visibility

  • Treat the full-run ordered Viridian-to-Route-2 transition waypoint as exact Route 2 progress in web run analysis even when the conservative rewardable map frontier remains at Viridian Mart, and report how many post-delivery stalls reached Route 2 before ending.
  • Extend the ordinary v3 lane's 24,000-step rolling lease floor from Parcel delivery through the first Route 2 anchor while preserving the 16,000-step global zero-badge timeout and the isolated v4/v5/v6 continuation contracts.

Release 15.1.5

15.1.5

No runs loaded

Monotonic parcel-learning continuation

  • Authorize the exact canonical bedroom state for internal promotion evidence, so the per-starter Parcel conversion gate counts real clean-start episodes.
  • Continue completed batches from their final sealed periodic checkpoint and recover that checkpoint if a mutable `latest.zip` alias is replaced before the scheduler evaluates it, preventing accidental PPO-lineage rollback.

Release 15.1.4

15.1.4

No runs loaded

Verified Oak Parcel handoff

  • Let the existing policy-selected `USE_POTION` head trigger a bounded Oak interaction only from the exact delivery waypoint, prove facing through the occupied Oak tile, and require the real Parcel removal plus cleared dialogue before reporting success.
  • Retain interaction attempts, orientation proof, A presses, Parcel removal, dialogue cleanup, and failure reasons in internal episode telemetry without adding new fields or stages to the public dashboard.
  • Withhold promotion until each starter converts at least 50% of eligible current-release Parcel acquisitions into deliveries, while preserving the existing 24,000-step post-Parcel progress lease.

Release 15.1.3

15.1.3

No runs loaded

Bounded v6 Viridian handoff

  • Extend the local 24,000-step progress lease to waypoint 54, covering the measured Viridian north-exit bottleneck as well as Route 2 and the immediate Forest handoff, without widening the global zero-badge lease.
  • Queue one exact-parent, 500,000-step v6 tranche behind the active v3 batch; the scheduler seals v3 before launch and grants no automatic extension or promotion authority.
  • Freeze the current 512-episode v3 health baseline and stop v6 behind starter-balanced Forest-entry, Parcel-delivery, rival-safety, legal-item, and move-execution evidence gates.
  • Bind the handoff to the fully resolved inherited configuration, package source, scheduler source, and exact parent so drift fails closed before launch or terminal evaluation.
  • Add a recoverable handoff-retirement path: preserve the stopped v6 checkpoint, telemetry, and log with hashes, clear only the active handoff slot, and resume ordinary v3 learning from its verified continuation.

Release 15.1.2

15.1.2

No runs loaded

Recurrent v6 Route 2-to-Forest handoff

  • Keep the global zero-badge meaningful-progress lease at 16,000 steps while applying a 24,000-step floor only to the Route 2 approach, Forest South Gate, and immediate Forest-entry waypoints.
  • Synchronize stale geometric waypoints after exact Viridian-to-Route 2 and Route 2-to-Forest transitions only when Parcel delivery, Ball-era encounter activation, the spent Route 1 encounter, and the configured doorway all agree; skipped waypoint rewards remain uncredited.
  • Report Forest entries per million completed episode transitions overall and by starter in ordinary-cycle health summaries.
  • Declare the isolated, untrained `nuzlocke-full-run-recurrent-v6` continuation lane with the same exact hashed v3 parent checkpoint inherited through v5; this release records no v6 checkpoint, panel pass, promotion, or live launch.

Release 15.1.1

15.1.1

No runs loaded

Recurrent v5 active-HP reconciliation

  • Reconcile a transient zero active-battle HP byte against the indexed party slot for battle eligibility, tactical waits, safety, finishing attacks, and damage advice while preserving both raw values in telemetry.
  • Require a zero party slot, permanent-death record, or visible replacement menu before treating the active member as fainted; forbidden battle items and the recurrent observation contract remain unchanged.
  • Declare the isolated, untrained `nuzlocke-full-run-recurrent-v5` continuation lane with the same exact hashed v3 parent checkpoint inherited through v4; this release records no v5 checkpoint, panel pass, promotion, or live launch.

Release 15.1.0

15.1.0

No runs loaded

Recurrent v4 handoff

  • Add an isolated `nuzlocke-full-run-recurrent-v4` continuation lane rooted at the immutable v3 499,986-step checkpoint, with an exact path, size, and SHA-256 resume binding and unchanged recurrent observation/action contract.
  • Bound the first training tranche to 500,000 transitions with a `5e-6` learning rate, `0.001` entropy coefficient, one PPO epoch, preserved optimizer state, and the shared feature extractor frozen.
  • Promote only canonical bedroom starts, retain forbidden battle-item masking, and use the validated four-Growl Bulbasaur opening-rival setup.
  • Correct critical-reply attack safety and retain low-HP recovery-classified controller turns in terminal battle traces.
  • Pass the 36-attempt paired rival gate at 30 wins versus 27 for v3, paired Route 1 starter smoke at 3/3 in both lanes, and canonical-bedroom smoke at 3/3 with no blackouts, rejected battle actions, or move mismatches.

Release 15.0.12

15.0.12

No runs loaded

Story-first Parcel delivery

  • Remove pre-Parcel wild-battle damage, victory, and level-up rewards, add a per-step wild-battle cost, and reserve progress-lease renewal for navigation, story events, and required trainer battles.
  • End an unproductive pre-Parcel detour on its eighth optional wild win and clamp the failed episode's complete return to `-500` so earlier navigation rewards cannot make the grinding trajectory profitable.
  • Retain bounded post-delivery wild-battle shaping while applying one shared reward cap across damage and victory signals.
  • Fail full-run health validation on any level-cap episode or excessive pre-Parcel wild wins, and support audited quarantine plus immutable-root restart of unhealthy optimizer continuations.

Release 15.0.11

15.0.11

No runs loaded

Opening recovery and verified healing

  • Route an injured opening-rival survivor from Oak's Lab to Red's house and restore the living party through Mom's real dialogue interaction before the first Route 1 grass exposure.
  • Enable the audited Pallet–Route 1–Viridian recovery corridor while retaining standard Pokémon Center healing at Viridian; both facilities verify the resulting party health and refuse to revive permanent Nuzlocke deaths.
  • Add saved-state emulator replays for both the Pallet Home and Viridian Pokémon Center paths, including collision-safe Pallet doorway navigation.

Release 15.0.10

15.0.10

No runs loaded

Production-grounded Pallet rival rollback

  • Restore the two-Growl Bulbasaur plan after completed production episodes measured 6/9 rival wins on 15.0.8 versus 5/15 on the four-Growl 15.0.9 release; the larger setup was a live regression despite passing the original frozen post-starter panel.
  • Add a bedroom-start rival panel so validation now includes starter creation, variable starter stats, scripted cohort selection, and the full pre-battle RNG path instead of relying only on one frozen post-starter state.
  • In the new panel, two Growls passed at 30/36 wins versus 29/36 for four, including 11/12 versus 10/12 Bulbasaur wins with unchanged 10/12 Charmander and 9/12 Squirtle results, 289/289 matched controller moves, and zero rejected battle actions.
  • Reject an attempted damage/HP-adaptive setup rule after it regressed to 28/36 wins versus the four-Growl arm's 29/36 in the same bedroom panel.

Release 15.0.9

15.0.9

No runs loaded

Bulbasaur rival variance reduction

  • Extend the frozen Bulbasaur setup sweep through six Growls. Four Growls won 8/12 fixed-seed battles versus 7/12 for the 15.0.8 two-Growl plan; six tied at 8/12 but required extra setup turns, while five regressed to 5/12.
  • Promote the four-Growl schedule for Bulbasaur against the opening Charmander while retaining live cursor verification, zero battle items, and the critical-aware forced attack path.

Release 15.0.8

15.0.8

No runs loaded

Pallet rival strategy and move-menu correctness

  • Model Generation I critical-hit probability and damage separately from the normal damage path, including doubled critical level and unmodified combat stats, while retaining the unchanged recurrent observation contract.
  • Navigate the live one-based move-menu cursor instead of trying to normalize a wrapping menu with repeated Up inputs or verifying against the previous confirmed move.
  • Keep legal damaging-move authority through unsafe rival states and use a starter-scoped two-Growl Bulbasaur setup against Charmander; battle items remain forbidden.
  • Add a frozen Bulbasaur setup sweep: two Growls won 7/12 battles versus 4/12 with one and 6/12 with three, with 429/429 matched move executions.
  • Expand the frozen release panel to 12 seeds per starter. The 15.0.8 candidate passed at 30/36 wins versus 26/36 for the advisor baseline, including 7/12 Bulbasaur, 11/12 Charmander, and 12/12 Squirtle wins, with 329/329 matched controller moves and zero rejected battle actions.

Release 15.0.7

15.0.7

No runs loaded

Opening-rival tactical safety and battle observability

  • Add a read-only Generation I battle calculator and compact live advisor telemetry without changing the recurrent policy observation contract.
  • Mask forbidden trainer-battle item and catch inputs at execution source, and verify controller move submissions against Red's dedicated move-list cursor.
  • Add an opening-Lab-rival safety controller that uses only legal battle-menu inputs and records intended-versus-executed move evidence for every turn.
  • Add a frozen, starter-balanced, inference-only rival comparison panel; the 15.0.7 candidate passed at 6/9 wins versus 5/9 for the advisor-only baseline, with 59 verified controller turns and zero execution mismatches or rejected battle actions.
  • Preserve inherited evaluation configs as both raw and resolved frozen artifacts so replay and Cloudflare metadata uploads remain reproducible.
  • Rotate oversized legacy step logs atomically when telemetry columns change, preserving history without copying a 24 GB CSV during trainer startup.

Release 15.0.6

15.0.6

No runs loaded

V20 execution-lock schema admission

  • Admit V19 and V20 to the existing inode-bound truthful execution-lock schema used by prior frozen-root recovery generations.
  • Keep the exact V17 lease preimage, runtime-config canonicalization, and all one-shot development-only limits unchanged.

Release 15.0.5

15.0.5

No runs loaded

V19 execution-lock binding correction

  • Bind the replacement launch to the exact unheld V17 execution-lease bytes left by the interrupted V17 execution, instead of the older V13 lease.
  • Preserve the V18 runtime-config canonicalization fix and all existing one-shot development-only limits.

Release 15.0.4

15.0.4

No runs loaded

V18 runtime-attestation correction

  • Supersede the interrupted, unbound V17 execution after its guardian compared the source seed and YAML key types against the deterministic JSON runtime.
  • Canonicalize the expected runtime config and derive the execution-rotated seed before binding, while retaining the immutable V13 root and scope.

Release 15.0.3

15.0.3

No runs loaded

V17 scheduler-binding correction

  • Supersede the unbound V16 activation after its stale scheduler-source digest blocked binding; no trainer launch or optimizer transition was consumed.
  • Bind the exact V17 cron manifest and reconciler digests before activation, then retain the same frozen checkpoint, optimizer state, and training scope.

Release 15.0.2

15.0.2

No runs loaded

Isolated V16 recovery launch

  • Replace the consumed zero-transition V15 launch with one bounded V16 launch from the same immutable V13 terminal checkpoint and unchanged optimizer and training settings.
  • Verify V16 authorization and root records without importing unpinned sibling modules inside the trainer's isolated Python runtime.

Release 15.0.1

15.0.1

No runs loaded

V15 recovery after the zero-transition V14 launch failure

  • Admit the completed V13 terminal checkpoint through a distinct generic, create-only V15 frozen-root attestation after the V14 launch failed before execution binding or optimizer progress.
  • Add a fresh scheduler-owned V15 one-shot recovery generation with 737,280 optimizer transitions, nine 3/3/3 starter environments, preserved optimizer state, and no evaluation, qualification, retention, promotion, retry, fallback, or extension authority.
  • Fail closed on the consumed V14 attempt and preserve every V14 lifecycle record as immutable history.

Release 15.0.0

15.0.0

No runs loaded

Verified V14 training generation

  • Preserve the exact completed V13 terminal checkpoint at lineage 144,762,678 as an immutable, content-addressed initialization root for one new bounded development-only generation. V13 remains consumed history and grants no retry, evaluation, panel, qualification, retention, or promotion authority.
  • Advance the sole scheduled guardian to recovery generation V14 with one optimizer-preserving 737,280-transition execution, nine training environments split 3/3/3 across starters, zero inline evaluation, and the unchanged first-controllable-bedroom-to-Hall-of-Fame episode contract.
  • Require create-only release-transition, authority, scheduler-binding, launch, process-activation, and execution-binding evidence. Reuse the exact existing guardian and status jobs, and fail closed on source, scheduler, process, checkpoint, budget, or identity drift without retry or fallback.

Release 14.0.26

14.0.26

No runs loaded

Process-visible launch binding

  • Preserve the consumed 14.0.25 v2 recovery claim as immutable, poisoned historical evidence. Its trainer was spawned but never activated or bound: the macOS process-visible interpreter replaced the claimed virtualenv `argv[0]`, so the fail-closed guardian handed the process off before any execution manifest, PPO update, or training transition was admitted.
  • Treat cloud execution identity `35637737-a383-4cea-ac4d-1e88221b0592`, written during that refused startup, as an orphan telemetry pointer, never as a candidate, execution binding, checkpoint, or evidence of training. V2 has zero admitted executions and zero transitions; its claim cannot be retried.
  • Authorize a distinct v3 development-only generation from the same exact sealed 2,999,997-step root. Bind the logical launcher and an independently measured process-visible interpreter identity before launch, then reuse that exact binding for activation, active-runtime validation, and emergency handoff. Launch the trainer in isolated interpreter mode with exact module paths. In a two-phase gate, the trainer first validates the claim and exact source/config/authority/guardian/root handshake; the guardian then binds the exact Screen → login → bash → Python ancestry plus the `tee` sibling and atomically publishes activation, which the trainer must acknowledge before it may write cloud identity, create an execution manifest, or initialize environments.
  • Scope the handshake as fail-closed operational safety against accidental launch, configuration, and process drift in the managed lane. It is not authentication or containment against a hostile process running as the same operating-system user.
  • Reserve `1005200001`–`1005200100` for the separate v3 development panel. The v2 `1005100xxx` panel remains unused historical scope, and `100300xxx` remains untouched for a separately authorized future qualification.

Release 14.0.25

14.0.25

No runs loaded

Semantic-retention recovery

  • Preserve the terminal 14.0.24 semantic-quality hold and its rejected branch as immutable evidence. Start a new development-only recovery identity from the sealed 2,999,997-step checkpoint at lineage 142,015,509; the rejected branch tip and its 4M/5M descendants are never eligible resume fallbacks.
  • Preserve the selected checkpoint's optimizer while reducing value-loss pressure and freezing its shared feature extractor for one bounded 1,004,544-transition tranche. Qualification, retention certification, and promotion authority remain false.
  • Consume the recovery authority before process creation, then bind the exact process start, command, execution manifest, health identity, and sealed terminal checkpoint pair. Interrupted, drifting, unbound, overshooting, or repeated execution fails closed and cannot be retried.
  • Distinguish a verified terminal semantic hold from a stale live trainer, retain last-runtime provenance, and independently verify the existing public Cloudflare run-status readback without making telemetry availability a launch authority.

Release 14.0.24

14.0.24

No runs loaded

Post-confirmation training recovery

  • Preserve the poisoned v1 and v2 bounded-recovery journals, all 216 v2 inference rows, and their sealed candidate as immutable historical evidence. The failed confirmation grants no qualification, retention certification, or promotion authority and cannot be retried or reused.
  • Authorize a new ordinary development generation from the exact sealed lineage-139,015,512 candidate solely as an optimizer-preserving development initialization. The guardian remains the only trainer owner, and only future sealed descendants of this new recovery identity may be resumed.
  • Require the north-Pallet waypoint to observe Oak's actual event flag before advancing, and guide premature Lab entry back to Pallet only during that pre-starter phase.
  • Keep the existing nine-worker 3/3/3 starter cohort and append-only `runs/train` telemetry path so the singleton metadata uploader continues publishing the new execution and episodes to the consolidated Cloudflare D1 database.

Release 14.0.23

14.0.23

No runs loaded

Bounded recovery

  • Preserve the complete V10 qualification as immutable failed evidence. Every starter supplied 12/12 scripted-cohort authority proofs and battle exposures; Bulbasaur recorded three wins and nine blackouts, while Charmander and Squirtle each recorded nine wins and three blackouts. Bulbasaur missed the unchanged four-win/eight-blackout gates, so no V10 snapshot is valid.
  • Replace outcome-driven ancestor selection with one predeclared bounded development execution initialized from the failed V10 checkpoint at lineage 138,010,968. Run exactly 218 nine-worker PPO rollouts (1,004,544 optimizer transitions) to lineage 139,015,512 and permit exactly one candidate.
  • Fail closed on interruption, source drift, resume-after-start, a second execution, intermediate or replacement checkpoint selection, extension, or ancestor fallback. Preserve the exact optimizer state and conservative PPO settings while granting no ordinary training authority.
  • Reserve fresh confirmation seeds 100100001–100100036 for 36 stochastic attempts per starter. Require 36/36 authority proofs and battle exposures, 36/36 binary-resolved outcomes, at least 14 wins, and no more than 22 blackouts for every starter under the predeclared intersection-union gate. Failure is terminal; a pass still requires a deliberate new recovery-contract binding.

Operations and reporting

  • Teach the guardian-owned trainer path to seal one immutable pre-drain candidate at the exact optimizer boundary and poison any interrupted or drifting execution instead of resuming it.
  • Report bounded-development startup, active runtime binding, evaluation-lease waits, confirmation handoff, and terminal holds explicitly. Keep V10 visible as historical evidence without misclassifying its expected failure as a new live alert under the deliberate bounded contract.
  • Move production telemetry writes to a fresh D1 generation while preserving the capacity-bound predecessor database read-only, and keep aggregate spatial history exact across the cutover.

Web dashboard

  • Move the release changelog from the crowded run-atlas overview to a dedicated, prerendered page backed by a compact release-notes payload.
  • Expand run analysis into a zoomed-out, horizontally scrollable journey Sankey with ordered checkpoints, including Route 1 before Viridian City and an optional Route 22 branch.
  • Make each Sankey ribbon selectable and list its exact matching episodes below the chart, with direct access to the existing full run-detail view.
  • Limit visible-tab telemetry refreshes to one request cycle every five minutes while preserving an immediate initial load and a due refresh on return.
  • Replace full-table D1 episode cursors and version counts with indexed cursors and materialized totals, and publish atlas updates from an ingestion cursor.
  • Archive terminal episode telemetry after 30 days to checksum-verified, immutable R2 batches before bounded nightly D1 deletion.
  • Preserve the durable static/R2 best-run winner when a hot telemetry shard has no stronger semantic candidate, so a shard cutover cannot regress the public best-run view.

Release 14.0.22

14.0.22

No runs loaded

New recovery generation

  • Preserve the complete 72/72-row v9 qualification as immutable failed evidence. Every starter supplied 12/12 exact scripted-cohort proofs and battle exposures; Bulbasaur recorded three wins and nine blackouts, Charmander seven wins and five blackouts, and Squirtle nine wins and three blackouts. Bulbasaur alone missed the unchanged four-win/eight-blackout gates, so no v9 snapshot is valid.
  • Reject the failed 3,999,996-step generation and every pre-existing descendant, then select the next eligible earlier root under the exact public-telemetry ranking fixed before v8: the sealed 14.0.16 checkpoint at execution step 2,999,997, lineage step 138,010,968, SHA-256 `b6b22afd185139e3882f7054c5db3de6dc668424f8f4862dbe74f76b42791c9f`.
  • Bind recovery `full-run-v3-semantic-stability-v9-2999997` to that root with contract SHA-256 `996d42817e2ea108da9082b075a81c2136ba715965ed3181b74c65558a0caff3`, preserving its structurally healthy Adam state and the conservative PPO stability settings.

One-shot v10 qualification

  • Reserve fresh seeds 32101–32112 for a new inference-only v10 qualification, preserving the 12/12 starter-proof and exposure requirements, four-win and eight-blackout gates, and nonblocking deterministic diagnostics.
  • Preserve v9 as a strict, snapshot-free, poison-on-interruption historical failure. V10 cannot import or resume any v7–v9 evidence, and only its passing sealed result may authorize its immutable non-training snapshot.

Release 14.0.21

14.0.21

No runs loaded

New recovery generation

  • Preserve the complete v8 qualification as immutable failed evidence: Bulbasaur supplied all 12 exact starter proofs and battle exposures but recorded only three wins and nine blackouts, missing the unchanged four-win and eight-blackout gates. No v8 snapshot is valid.
  • Reject the failed 4,999,995-step generation and every pre-existing descendant, then select the next eligible earlier root under the exact public-telemetry ranking fixed before v8: the sealed 14.0.16 checkpoint at execution step 3,999,996, lineage step 139,010,967, SHA-256 `4d4972f995b89c08b3b2071d5d94c0ea774fc23ada7f2a0845306dfa042a1771`.
  • Bind recovery `full-run-v3-semantic-stability-v8-3999996` to that root while preserving its complete finite Adam state and the existing conservative PPO stability settings.

One-shot v9 qualification

  • Reserve fresh seeds 31101–31112 for a new inference-only v9 qualification, preserving the 12/12 starter-proof and exposure requirements, four-win and eight-blackout gates, and nonblocking deterministic diagnostics.
  • Keep v8's strict poison-on-interruption and exact source/runtime binding; extend the narrow historical-failure verifier so sealed failed evidence remains authentic after a deliberate later recovery source change.

Release 14.0.20

14.0.20

No runs loaded

New recovery generation

  • Preserve the completed v7 qualification as immutable failed evidence: the Squirtle cohort supplied only 10/12 exact scripted-cohort starter proofs and 10/12 battle exposures, so the otherwise satisfied win/blackout gates cannot authorize a snapshot or launch from the 11,999,988-step root.
  • Define recovery generation `full-run-v3-semantic-stability-v7-4999995` from the earlier sealed 14.0.16 periodic checkpoint at execution step 4,999,995, lineage step 140,010,966, and SHA-256 `10131272b8e3090bd856e56e8003d2e349751c0c34933820f65fc950693a8252`.
  • Select that root before v8 evaluation from healthy sealed periodic ancestors earlier than the failed v7 root, requiring Parcel delivery for every starter in the following production million and ranking exact starter-proof plus battle-exposure coverage before proof, opening-win, Parcel, aggregate coverage, and earlier-checkpoint tie breakers.
  • Record the selected root's public following-million evidence: exact proof and battle exposure were 14/14, 8/8, and 8/8; opening wins were 6/14, 5/8, and 6/8; and Parcel deliveries were 3/14, 4/8, and 3/8 for Bulbasaur, Charmander, and Squirtle. These observations select the root but do not qualify it.

One-shot v8 qualification

  • Reserve the fresh inference-only v8 panel at seeds 30101–30112 with the same predeclared four-win/eight-blackout gates, complete authority proof and battle-exposure prerequisites, and nonblocking deterministic diagnostics.
  • Keep the guardian stopped until the exact v8 matrix seals successfully and its immutable snapshot exists; v7 evidence is never retried, completed rows cannot be resumed into v8, and no failed qualification can create a snapshot.

Release 14.0.19

14.0.19

No runs loaded

Recovery qualification

  • Define recovery generation `full-run-v3-semantic-stability-v6-11999988` from the sealed 14.0.16 checkpoint at execution step 11,999,988, lineage step 147,010,959, and SHA-256 `d4a85c108f352bf901b7b5b29cda3623ea144f2b3abd829c4cfeb5eea582f646`.
  • Select that single root before evaluation by maximizing the minimum per-starter opening trainer-win fraction in the following production million among pre-e9 periodic roots with nonzero Parcel delivery for every starter; preserve its complete, finite Adam optimizer state.
  • Add the immutable inference-only v7 qualification with twelve previously unused stochastic certifying seeds, deterministic nonblocking diagnostics, and proportionally scaled four-win/eight-blackout per-starter gates. The historical 271xx and 281xx panels cannot be reused to choose a root, threshold, or retry.

Evaluation and guardian clarity

  • Keep PPO at nine balanced training workers and zero inline evaluation workers; expose frozen-checkpoint qualification as a separate external lane in the dashboard.
  • Preserve cryptographically verified failed qualification reasons as a fresh, exit-zero guardian hold even after source-only drift, while keeping missing, malformed, or tampered evidence as nonzero admission errors.
  • Forbid snapshot creation from a failed qualification and continue to block checkpoint selection and launch until the active root has both passing qualification evidence and its immutable snapshot.

Release 14.0.18

14.0.18

No runs loaded

Deliberate semantic recovery

  • Define reviewed recovery candidate `full-run-v3-semantic-stability-v5-12999987` from the sealed 14.0.16 root at execution step 12,999,987, lineage step 148,010,958, and SHA-256 `e9a01b6ba6400dca3552feb29d1ba00d8a17b130d53b46c83f4c9bd6dbdd183d`.
  • Review the earlier root using its immediate production-million proxy: 13/41 eligible episodes delivered the Parcel (Bulbasaur 4/15, Charmander 5/11, Squirtle 4/15), compared with 3/39 after the later 15,999,984-step root.
  • Preserve failed intermediate recovery `full-run-v3-semantic-stability-v4-15999984` and its v5 qualification as historical evidence: the root at execution step 15,999,984 and lineage step 151,010,955, with SHA-256 `acc890970cb7e16057bee25da6847c790c2952d09c73d5f53ead2f1a54ec402f` produced zero Bulbasaur wins and three blackouts in three attempts, failed qualification, and therefore never received a snapshot.
  • Reject that failed root and every 14.0.17 descendant, including the 999,999- and 1,999,998-step periodic checkpoints and the 2,744,757-step terminal checkpoint, instead of selecting a newer checkpoint from a regressed branch.
  • Clear the latched 14.0.17 semantic-quality hold only through this changed recovery contract; source-only releases and recovered rolling metrics remain unable to resume a rejected generation.

Optimizer stability

  • Preserve the candidate root's structurally healthy Adam optimizer state, class, and keyword arguments; verified descendants preserve their archived optimizer state on subsequent resumes.
  • Reduce PPO updates to a `1e-5` learning rate, one epoch, `target_kl: 0.0025`, `max_grad_norm: 0.3`, and `ent_coef: 0.005` while retaining the recurrent policy, observation, and action contracts.

Retention authority

  • Complete the v6 six-seed rows with stochastic wins of 1/6, 5/6, and 4/6 and blackouts of 5/6, 1/6, and 2/6 for Bulbasaur, Charmander, and Squirtle. Bulbasaur fails both the minimum-two-win and maximum-four-blackout gates, so no baseline snapshot is authorized.
  • Preserve the first completed v6 row set separately after final sealing rejected source drift caused by verifier code changing during evaluation; rerun the canonical frozen source to seal the same failed evidence.
  • Keep v4 and the failed, snapshot-free v5 qualification immutable as historical evidence with no authority to unlock the reviewed candidate.
  • Keep the current v5/e9 source contract review-only: v6 cannot authorize a launch, no snapshot may be created, and training remains held.

Release 14.0.17

14.0.17

No runs loaded

Semantic recovery generation

  • Start recovery generation `full-run-v3-semantic-stability-v3-15999984` from the audited sealed 14.0.16 checkpoint at execution step 15,999,984 and lineage step 151,010,955 while preserving its optimizer state and PPO stability settings.
  • Exclude the regressed 18,713,088-step predecessor terminal tip from resume selection; only the exact new root and its cryptographically verified new- generation descendants are eligible.
  • Reserve `runs/full_run_battle_retention/v4/latest` for fresh 14.0.17 retention evidence instead of inheriting the predecessor release's result.
  • Fail closed before installed-release checks, checkpoint selection, or launch until the public v4 verifier accepts the exact root's capability qualification and immutable non-training baseline snapshot.

Training control

  • Make the file-invoked full-run guardian import its shared semantic reporter from the repository root, matching the production cron entrypoint.
  • Treat unavailable semantic-quality evidence as an explicit non-launching blocked state instead of reporting the trainer healthy, with subprocess and active-runtime regression coverage for the production import context.
  • Keep a statistically separated fixed-window Parcel-delivery collapse in the regressed state when one late delivery enters the rolling window, so the guardian cannot miss the checkpoint-safe latch between polling ticks.
  • Scope a semantic-quality hold to its recovery contract across source-only releases, preventing a version bump from making the same regressed branch tip selectable; only a deliberate new recovery generation clears the hold.
  • Preserve a committed semantic-quality handoff across release, config, and source refreshes so the exact trainer is never signalled twice or relabelled as an ordinary restart.
  • Keep a latched semantic regression visible while battle-retention owns the execution lease, and report that evaluation returns to hold without launch.
  • Move the status-facing 14.0.16 terminal-retention retry to a fresh immutable output after the first complete matrix correctly failed final sealing when its audited worktree identity changed during evaluation.
  • Treat `pokemon_fainted` as the required companion marker of a bounded whiteout in battle-retention evidence, while continuing to reject genuine non-blackout safety failures.

Release 14.0.16

14.0.16

No runs loaded

Route 1 catch eligibility

  • Keep pre-Ball wild encounters out of the one-encounter-per-area ledger, then latch ordinary Nuzlocke spending when the first legal catch becomes possible.
  • Preserve a newly registered encounter's transient eligibility on its first policy frame so an immediate legal `THROW_POKEBALL` is not remapped to flee.
  • Revisit the audited Route 1 grass corridor immediately after the Mart's five-Ball gate, while leaving throw, battle, or flee entirely policy-owned.
  • Attest the catch unlock, post-Ball detour, and policy authority in guardian and status contracts, and report policy-selected catch macros as policy-owned.

Release 14.0.15

14.0.15

No runs loaded

Latched semantic-stability recovery

  • Start a new recovery generation from the sealed 14.0.13 periodic checkpoint at step 7,999,992 instead of treating the later 14.0.14 descendant as independent clean-start retention evidence.
  • Preserve the checkpoint optimizer state while pinning resumed PPO updates to a `2e-5` learning rate, three epochs, `target_kl: 0.005`, and a `0.3` maximum gradient norm.
  • Add a target-independent, one-shot 1,000 reward for the last policy-owned Oak escort transition into the controller handoff anchor, and give waypoint 14 a local 24,000-step progress lease for the opening rival battle.
  • Make semantic regression or plateau request one checkpoint-safe guardian handoff and latch the same release/recovery generation in a stopped hold even if the rolling metric subsequently appears healthy.
  • Keep controller-created starter party, level, and catch accounting reward-neutral, including a delayed Pokédex-owned bit that previously could race the rebase and masquerade as the policy's first capture.

Release 14.0.14

14.0.14

No runs loaded

Full-run semantic recovery

  • Treat scripted starter initialization as reward-free controller bookkeeping, rebasing party, Pokédex, and level trackers without catch or lease credit.
  • Remove the controller-owned starter-match reward while retaining exact target verification, mismatch penalties, and invalid terminal behavior.
  • Apply an audited PPO target-KL early-stop guard to fresh and resumed recurrent policies to contain destructive optimizer updates.
  • Resume from the last sealed periodic checkpoint before the observed policy collapse, then admit only cryptographically verified descendants of that recovery root instead of drifting back to the rejected lineage tip.

Release 14.0.13

14.0.13

No runs loaded

Battle-retention evaluation integrity

  • Reapply and attest each fixed evaluation seed after recurrent-checkpoint loading so candidate and baseline resets cannot silently inherit different saved training seeds.
  • Certify the policy's stochastic training-time behavior under a versioned retention contract while retaining deterministic inference as an explicit, non-certifying diagnostic.
  • Hold one typed execution lease for the complete evaluation and seal cycle, with atomic guardian arbitration that prevents an evaluator and trainer from acquiring emulators concurrently.
  • Keep verified retention outcomes as direct release-quality blockers without relabeling a fresh, exactly bound trainer as operationally unhealthy, and print the per-starter failure causes needed for remediation.

Release 14.0.12

14.0.12

No runs loaded

Audited Parcel-return guidance

  • Replace broad off-phase first-visit rewards with bounded projected-exit shaping: only a new best distance or a one-shot audited transition pays, and lease renewal remains interval-gated and episode-persistent against farming.
  • Add the trace-verified Blue's House exit plus southbound Route 1 and Oak's Lab approach anchors while preserving uninterrupted bedroom-start policy runs.
  • Report a config-derived Parcel-return funnel, per-starter acquisition and delivery conversion, and fixed non-overlapping 256-episode Wilson windows.
  • Preserve complete pre-selection retention failures as failed matrix cells, and report zero-decision failed cohorts without misclassifying their sealed evidence as corrupt.
  • Keep battle penalties unchanged so this navigation release remains separable from the pending battle-retention certification.

Release 14.0.11

14.0.11

No runs loaded

Parcel-return learning recovery

  • Restore positive, first-visit navigation feedback on audited off-phase detours after broad suppression caused Parcel-return performance to collapse.
  • Keep the audited exit observations and Route 22 eight-badge frontier gate, and pin the corrected reward contract in guardian and status validation.
  • Add coupled regression coverage so production detour guidance cannot again coexist with a fully suppressed navigation signal.
  • Retain completed, legal, exact-cohort semantic bests as passive content-addressed checkpoints across release handoffs, without changing the guardian's maximum-lineage resume selection.

Release 14.0.10

14.0.10

No runs loaded

Verified starter-phase alignment

  • Align stale opening waypoint guidance to Oak's Lab exit only after the bounded cohort controller proves a coherent, exact assigned-starter match, without awarding skipped waypoint rewards or executing additional inputs.
  • Pin the reward-free synchronization marker in guardian and status contract validation so target allocation alone can never masquerade as observed selection evidence.
  • Retain same-release Parcel acquisition evidence beyond the rolling semantic window so an aged loss of acquisition performance remains a regression.

Release 14.0.9

14.0.9

No runs loaded

Full-run semantic quality and status accuracy

  • Give only the audited post-Parcel Oak's Lab return waypoints a longer progress lease, and suppress positive exploration credit on maps outside the pending ordered phase without hiding their detour telemetry.
  • Keep same-release Parcel deliveries in the semantic history so a later loss of delivery performance is reported as a regression instead of a missing milestone.
  • Report active vector episodes independently of clean promotion authority, separate semantic-quality blockers from operational liveness, and describe absent battle-retention evidence during an exact active run as pending the guardian-sealed terminal candidate.
  • Publish installed release and source identities on healthy guardian ticks so release status no longer falls back to `unreported`.
  • Bind battle-retention evaluation exclusion to the live training execution lease, fail closed when that lease disagrees with its pending manifest, and ignore abandoned unlocked manifests during terminal-candidate preparation.

Release 14.0.8

14.0.8

No runs loaded

Parcel-delivery navigation

  • Guide the policy through Oak's Lab along the collision-checked return aisle, then admit every historically demonstrated tile adjacent to Oak while still requiring the real Parcel-delivery transition for semantic completion.
  • Charge the checkpoint-compatible out-of-Mart shopping action as a policy-owned no-op so repeated idle selections receive the configured action penalties without changing the recurrent policy's ten-action interface.

Release 14.0.7

14.0.7

No runs loaded

Parcel-return phase alignment

  • Synchronize a real, policy-owned Parcel acquisition to the Mart-exit return phase when geometric waypoint credit is stale, without executing controller inputs or awarding skipped-route rewards.
  • Record the monotonic semantic alignment for audit and isolate its evidence from the plateaued 14.0.6 cohort.

Release 14.0.6

14.0.6

No runs loaded

Parcel-return learning

  • Aim parcel-delivery guidance at Oak's audited interaction tile with exact distance shaping, while leaving the final facing and interaction inputs under uninterrupted policy control.
  • Isolate post-fix semantic evidence from the plateaued 14.0.5 cohort so a future delivery must be demonstrated by the corrected release itself.

Release 14.0.5

14.0.5

No runs loaded

Full-run training quality

  • Gate map-frontier reward by the active story phase and require eight badges before Route 22 can earn frontier credit, while retaining detours as telemetry.
  • Rank general checkpoints by ordered semantic waypoint progress before raw map priority so an accessible late-game side map cannot displace story progress.
  • Add audited off-phase exits for optional Viridian detours and region-aware Route 2 guidance back through Viridian Forest without rewinding the objective.
  • Report balanced current-release semantic progress windows and raise an explicit training-quality alert for regressions or sustained plateaus.

Release 14.0.4

14.0.4

No runs loaded

Opening-battle learning

  • Ignore transient, uninitialized opponent HP while establishing battle-damage baselines so real HP loss supplies the intended dense learning signal.
  • Stop penalizing directional battle-menu navigation, explicitly penalize rejected battle-item actions, strengthen bounded damage shaping, and reduce excess policy entropy while preserving policy-owned combat.
  • Remove the repeatable battle-entry bonus so controller-fled wild encounters cannot out-reward meaningful damage, trainer wins, or route progress.
  • Pin the corrected battle-learning settings in the full-run guardian and continue the recurrent v3 lineage from its gracefully sealed checkpoint.

Certification integrity and observability

  • Require starter-target and selection-authority proof from every fixed-seed candidate and retained-baseline battle attempt, including losing cells.
  • Bind the retention evaluator's exact checkpoint, scenario, seed, inference, and job matrix back to the canonical bedroom-rival contract, and reject non-blackout safety failures from either side of the paired cohort.
  • Preserve fresh/resume launch provenance from the bound execution manifest in healthy guardian status, rejecting missing or contradictory resume evidence.
  • Fail closed in the independent status report when the configured or guardian-reported entropy, battle-reward, launch, or config identity drifts from the bound 14.0.4 runtime.
  • Rebuild the run-level furthest-map frontier from completed episode telemetry on resume so a graceful handoff cannot make reported progress move backward.
  • Restore the canonical total-party-level frontier across execution handoffs instead of reading the similarly named legacy health field.
  • Bound interrupted subprocess-vector teardown after a partially consumed rollout step so an atomically sealed guardian handoff cannot hang forever.
  • Let the guardian safely retry a terminal-teardown signal that failed before delivery, while rechecking process start time to prevent PID-reuse races.

Release 14.0.3

14.0.3

No runs loaded

Starter controller hardening

  • Finalize the transient starter award with cancel-safe inputs so timing differences cannot enter the nickname keyboard or strand an uninitialized party slot.
  • Guard both sides of the starter table before policy input and keep controller aborts out of cohort mismatch binding, reward, and error counters.
  • Require the durable followed-Oak story flags before the controller can own the Lab interaction, and guide fresh-bedroom workers through the north Pallet trigger before any Lab waypoint. A pre-Oak visit to the identical table tile now remains ordinary policy-owned exploration.

Runtime and evidence integrity

  • Bind guardian adoption, handoff, health, manifests, and execution leases to the exact Python source-tree digest, with a fail-closed legacy migration and highest-lineage same-run checkpoint resume selection.
  • Recompute battle-retention certification from the fully sealed plan, checkpoint inputs, fixed matrix results, and diagnostic evidence; outer summary edits or missing/tampered frozen inputs can no longer certify a release.
  • Bind status health to the guardian's exact PID, run, execution, config, and source identity before reporting operational or release readiness.

Release 14.0.2

14.0.2

No runs loaded

Guaranteed starter assignment

  • Give Oak's Lab starter selection to a bounded, fail-closed cohort controller that uses ordinary game inputs, targets only the worker's assigned ball, and verifies the initialized party slot before returning control to the policy.
  • Record the pickup honestly as `scripted_cohort`, suppress sampled-action reward leakage during the macro, and terminate safely on any unexpected Lab state instead of permitting a wrong-ball fallback.
  • Protect the controller coordinates, timing bounds, authority, terminal contract, status metadata, and all three ROM pickup paths with regression tests.

Release 14.0.1

14.0.1

No runs loaded

Full-run starter cohorts

  • Keep each worker's Lab waypoint aligned to its assigned starter until the policy completes that pickup, instead of advancing target-agnostic guidance toward the exit while the party is still empty.
  • Defer starter-cohort scoring and Nuzlocke faint detection until the newly written party slot has coherent level and HP data, preventing transient starter-pickup frames from becoming false deaths or invalid attempts.

Release operations

  • Derive the full-run dashboard's no-heartbeat release fallback from the package version instead of pinning a release number in the component.
  • Gate every guardian launch on the installed distribution matching the source release, and bind graceful version handoffs to the adopted trainer's PID and execution identity so stale health cannot stop a newer process.
  • Report battle outcomes and rejected battle actions by starter, distinguish operational liveness from release readiness, and fail release evidence closed when retention artifacts are missing, stale, mismatched, or tampered.
  • Make the v14 full-run guardian and status reporter the documented operating path, and mark the former supervisor, specialist, canary, and hierarchical procedures as archived references that must not be revived.
  • Replace the README's stale scaffold goals and legacy June progress snapshot with the current v14 contract, required assets, safe verification commands, and an explicitly archived analyzer marker block.

Battle retention evaluation

  • Add a provenance-verified, inference-only full-run battle-retention lane for fixed all-starter bedroom-to-rival cohorts, including immutable baseline snapshots, paired win/blackout/action-rejection regression checks, a 40,000 step opening horizon, and a status-facing fail-closed summary with no training or promotion authority.
  • Require a same-run, earlier-lineage baseline before certifying retention; bind candidates to the sealed configured v3 latest model/config/algorithm/ observation identity, keep unpaired checks capability-only, and count battle wins relative to reset-time cumulative counters.
  • Preserve a digest-verified, content-addressed `best_battle` checkpoint lane independently for each starter, ranked by trainer wins before aggregate combat progress and published atomically without replacing navigation-owned `best_reward` selection.

Release 14.0.0

14.0.0

No runs loaded

Continuous full-run lineage

  • Add `nuzlocke-full-run-recurrent-v2`, a fresh recurrent lineage whose every episode loads only the audited first-controllable bedroom state and then runs continuously to Hall of Fame, blackout, a meaningful-progress stall, or an immediate Nuzlocke-invalid terminal action. Disable promoted-state resets, progressive replay, emulator-state capture, and inline deterministic eval.
  • Preserve infrequent policy-weight checkpoints for crash recovery without resetting episodes, diversify canonical resets with seeded pre-control idle frames, and persist policy-owned starter selection plus RNG attestation in exact episode telemetry.
  • Add the incompatible `nuzlocke-full-run-recurrent-v3` lineage with an explicit nine-worker target pool: three workers each for Bulbasaur, Charmander, and Squirtle. Condition the shared policy on a target-starter one-hot, require policy-owned selection to match that target, and use a 768-sample PPO minibatch over each 4,608-step rollout.
  • Cut unattended ownership over to the dedicated nine-worker, zero-evaluator guardian and retire the specialist curriculum, clean-start canary, hierarchical readiness poller, specialist reviewer, and stuck checker.

Sealed specialist predecessor

  • Preserve the exhausted Viridian-to-Pewter continuation and its final checkpoints as historical evidence only. It no longer owns training, evaluation, promotion, or reset-state authority after the full-run cutover.

Release 13.0.0

13.0.0

No runs loaded

Automatic cohort qualification

  • Reserve immutable challenger, baseline, scenario, and state identities for every bounded Stage-10 or full-game milestone cohort, then run the paired semantic evaluator automatically after the trainers become idle.
  • Fail closed on missing, partial, stale, or tampered evaluation evidence and issue at most one exact continuation token only after the challenger is safety-clean and non-regressing. Evaluation evidence never promotes a model directly.

Policy and specialist lineages

  • Make opening rehearsal genuinely clean-heavy with an adaptive 80%, 75%, or 70% clean-start reset share while retaining every required ladder rung in a spaced deterministic per-worker cycle and failing closed if a rung is absent.
  • Add matched fresh 73-feature feed-forward and recurrent PPO lineages with unit-normalized game state, plus recurrent hidden-state and episode-boundary handling across training, evaluation, canary, and hierarchical inference. Embed their ordered feature, normalization, lineage, algorithm, policy, and network contract inside each checkpoint and reject incompatible resumes.
  • Split specialist contracts into shared recurrent navigation and starter-scoped battle policies, require digest-bound trained promotion evidence for publication, and preflight the complete hierarchy before any emulator step. The monolithic end-to-end PPO remains benchmark-only.
  • Cut unattended training ownership over to a fail-closed specialist guardian: pin the reviewed shared-navigation inputs, resume only sealed exact-contract checkpoints toward one cumulative six-million-step target, bound unchanged launch retries, and keep the legacy monolithic supervisor disabled.

Shared-navigation canary cutover

  • Cut the scheduled singleton canary over from held starter-specific policies to the newest immutable, provenance-bound periodic checkpoint from the developmental shared recurrent Viridian-navigation lineage. Rotate all three clean starter states through the same upstream-trained navigation policy so newer canary snapshots reflect experience pooled during shared training; canary attempts themselves remain inference-only and never update one another or the trainer.
  • Preserve the source policy's Viridian-Mart objective, six-waypoint target set, 6,000-step horizon, and normalized 73-feature recurrent contract through the complete trained phase. Once the Mart milestone is reached, dynamically activate an evaluation-only route that acquires and delivers Oak's Parcel, exits the Lab, and continues under a 16,000-step meaningful-progress lease with a 500,000-step hard cap. Verified Center healing can renew the lease; merely occupying the Center cannot. Give the continuation a distinct series identity so historical Mart-terminal attempts cannot inflate its status or dashboard counts. Keep the entire canary allocation to one emulator outside the eight-worker plus one-periodic-evaluator training pool.
  • Label shared checkpoints as developmental and unpromoted, load no battle specialist, and grant no training, promotion, publication, hierarchical readiness, fallback, or production authority. Retain the prior starter-specific Boulder-Badge canary as a separately invoked manual legacy control whose results cannot be combined with the scheduled parcel-continuation series.

Bounded learning semantics

  • Preserve Stage-10 ordered partial-progress returns instead of clawing every ordinary failed rehearsal back to the same total; retain a small fixed failure cost and the existing hard Nuzlocke safety penalties while reducing the nearly uniform policy's entropy pressure.
  • Settle verified Route-3 handoffs when the map byte changes even if no blank transition frame is observed, preventing stale Center-exit coordinates from handing the remaining Pewter corridor back to PPO.
  • Let Route-3 readiness recovery recognize exact-full-health as the sole remaining objective condition and permit a second verified, action-bounded post-heal return, while retaining the one-return default and a hard cap of two.
  • Reserve the Route-3 canonical worker for promotion-authority states and use appended Stage-10 candidates only as non-authority learning-lane fallbacks until stable progressive replay states exist.

Reproducible training and evaluation telemetry

  • Add immutable per-execution manifests with source, dependency, input, resume-lineage, failure, and checkpoint identities, plus append-only PPO update history.
  • Add exact per-episode behavior, resource, reward-component, milestone, safety, recovery, and terminal-outcome telemetry; keep sampled step traces out of automatic action-rate tuning and stop cumulative species inflation.
  • Add fixed, immutable candidate-versus-baseline semantic evaluation cohorts, retention matrices, and opt-in hash-bound curriculum promotion gates.

Dashboard clarity

  • Explain the clean-start canary's evaluator and checkpoint versions beside their reported values, distinguishing the code used for the test from the frozen model weights and why their release numbers can differ.

Release 12.1.64

12.1.64

No runs loaded

Causal Route 3 readiness learning

  • Move the deterministic post-Brock Route 3 handoff into bounded reset preprocessing, preserve its semantic proof without PPO reward or clock credit, and permit only one verified post-Center rearm per episode.
  • Shape only stable, delta-based party readiness; renew bounded episodes only on durable party progress or confirmed wins, and prevent controller travel from manufacturing frontier, waypoint, or objective reward.
  • Assign the three learner workers distinct canonical, uniform-replay, and priority-replay roles; diversify recorded cohort seeds and tune resumed PPO exploration and critic weights for the readiness objective.
  • Record policy/controller/rejected action ownership and require sufficient policy-owned data and authoritative episodes before promotion or automatic milestone renewal.
  • Retain compatible 12.1.63 Route 3 battle states across the source-only release while filtering replay to healthy Route 3 battle boundaries.
  • Load the bounded Route 3 learners on CPU because the current Torch runtime aborts while restoring the Stage-9 archive directly onto MPS.

Release 12.1.63

12.1.63

No runs loaded

Elite Four handoff integrity

  • Add an explicit HM04 milestone after Surf: acquire Strength, teach move 70 to a living party member, and revalidate that living user before Victory Road can hand off to Indigo Plateau.
  • Require every newly written starter-track promotion, team signature, and safety total to come from the exact current release while preserving the separate historical-revalidation path.
  • Derive minimum-party-level progress only from a complete, positive live-party vector; incomplete vectors and unproven aggregate fields can no longer renew a cohort or authorize semantic progress.

Release 12.1.62

12.1.62

No runs loaded

Route 3 readiness and promotion integrity

  • Keep healthy post-Brock starts on a verified Gym/Center/Pewter handoff to Route 3, and rotate a safe under-level teammate into the lead before weak wild encounters can starve it of participation experience.
  • Route top-region Viridian Forest recovery north toward Pewter while retaining the southbound Viridian branch, covering the current Charmander poison and low-HP failure frontier without weakening recovery priority.
  • Include minimum party level in the clean durable milestone frontier and bind newly created promotion candidates to the exact current release, while preserving explicitly separate historical revalidation.
  • Correct the Silph Co frontier and specialist routing to start at canonical Silph Co 1F (map 181), not Saffron Mart (map 180).
  • Require Cut and Surf to be taught to a living party member before their HM milestones promote, retaining the party-wide proof in candidate sidecars.
  • Add code-owned recovery corridors from east Route 4 through Cerulean Center, and movement-only retreat routes for Nugget Bridge, Route 25, Bill's cottage, and Cerulean Gym.

Release 12.1.61

12.1.61

No runs loaded

Current-contract promotion authority

  • Require starter-tracked curricula to earn their promotion successes under the current model contract, preventing historical rows from an obsolete objective definition from ending a rematerialized cohort on its first vector step.

Release 12.1.60

12.1.60

No runs loaded

Elite Four milestone throughput

  • Publish the CLI-resolved worker count into every environment so a 3-worker cohort shards weighted static/replay cycles as three workers instead of the stale eight-worker YAML default.
  • Instantiate SB3's reward-only periodic evaluator only when enabled and turn it off for milestone stages, restoring the literal nine-learner plus one-canary emulator budget.
  • Make Route 3 an at-least-two-member handoff where every current party member is level 16 and fully healthy, with durable training replay; normalize Route 3/Mt. Moon battles through the bag-safe strongest-damaging-move controller while filtering stale off-route replay states.

Release 12.1.59

12.1.59

No runs loaded

Milestone replay integrity

  • Balance Mt. Moon's authoritative Route 3 predecessor and progressive replay states evenly, preserving clean static-start promotion authority while giving learned deep starts equal rehearsal weight.
  • Credit ordered warp waypoints only from an exact, internally produced macro proof that matches the waypoint index, source tile, direction, and sampled destination; policy-supplied or mismatched evidence remains ineligible.
  • Keep replay-started episode provenance permanently promotion-ineligible, even if a later reset path looks canonical, and require clean semantic authority when recording a promotion candidate.

Release 12.1.58

12.1.58

No runs loaded

Mt. Moon forward progress

  • Replace Route 3's unreachable RAM-injected `(66, 0)` waypoint with the real collision-walked `(59, 0)` north connection and extend ordered guidance through Route 4, Mt. Moon 1F/B1F/B2F, the mandatory trainer and fossil gates, and the genuine east exit.
  • Replay only clean, current-release, stage-local waypoint/battle/item/heal captures while retaining the authoritative Route-3 start for promotion, and checkpoint twice within each bounded milestone cohort so useful learning survives exits.
  • Permit one level-cap-safe damaging finisher against a nearly defeated ordinary trainer when no healthy preventive switch exists; Gym, wild, unsafe, and repeated attempts remain fail-closed.
  • Attribute milestone progress to the clean static-start semantic frontier when progressive replay is enabled, keeping rehearsal states out of review gates.

Release 12.1.57

12.1.57

No runs loaded

Milestone frontier integrity

  • Rank Route 4's reused west-side map before the Mt. Moon Center and cave floors, while retaining the exact east-exit coordinate as the completion proof. Entering the west cave approach can no longer pre-credit the final map frontier or its full bounded-cohort bonus.
  • Suspend ordered waypoint shaping only when deterministic recovery actually owns the action. A route-missing or out-of-range policy fallback can now approach the verified corridor and earn genuine semantic progress.
  • Include Pewter and its Pokemon Center in the Mt. Moon stage's required recovery-map preflight, so the complete Route 3 heal leg remains code-owned and fails closed if either verified map route regresses.

Supervisor fail-closed scheduling

  • Transfer the shared nine-environment pool only after every parsed process in the previous lane checkpoints and exits, derive availability and inventory from one process snapshot, and reserve legacy launches before spawning.
  • Validate either surviving supervisor ledger before scheduling, keep pending Stage-10 approval tokens bound to their starter, and give Stage-10 startups the same bounded fresh-health lease as milestone cohorts.
  • Prune review authority after a valid promotion, report checkpointed/held lanes from actual active work, and tie fresh `running` health to its parsed live stage process before attributing the runtime release.

Release 12.1.56

12.1.56

No runs loaded

Bounded milestone review

  • Reserve Route 3's west exit for verified recovery while the party is injured, preventing a healthy Mt. Moon cohort from immediately backtracking into Pewter from its Stage-9 handoff boundary.
  • Describe full-budget milestone review blocks as completed bounded cohorts in authoritative status instead of labeling them interrupted, keep explicitly authorized milestone trials out of the blocked set, and retain the Stage-10 semantic threshold while the milestone lane owns the supervisor status.
  • Replace a consumed Stage-10 or milestone approval with a fresh, one-use review block after the bounded worker completes, interrupts, disappears with running health, or exits without attributable current-release health.
  • Require the exact release/starter/block token for every Stage-10 or milestone trial, reject persistent bare starter IDs, and atomically reserve the token, launch attempt, and milestone startup lease before spawning a trainer.
  • Fail closed on malformed supervisor ledgers, reserve approvals for already active bounded cohorts before any early return, and reconcile a consumed Stage-10 result before returning the shared pool to milestone scheduling.
  • Preserve an exact Stage-10 approval across the complete milestone runner/worker checkpoint handoff, and select the newest matching-generation primary/last-known-good ledger in both the supervisor and status reporter.
  • Report exact current-release review-block tokens for both lanes without treating token availability as human authorization or a monitor threshold as every starter's hold trigger.

Release 12.1.55

12.1.55

No runs loaded

Recovery-safe waypoint progress

  • Suspend ordered waypoint rewards, index advancement, and progress-lease renewal on every recovery-navigation action, so a westbound Route 3 heal cannot impersonate eastbound milestone progress.
  • Preserve the active semantic waypoint through the complete recovery route and resume ordinary advancement only after control returns to the policy.

Release 12.1.54

12.1.54

No runs loaded

Route 3 progress guidance

  • Close the bounded Charmander cohort's Route 3 recovery gap at `(19, 4)` and replay that exact current-release battle-end state through a clean, battle-free Pewter Center heal.
  • Add ordered eastbound Route 3 waypoints plus post-heal Center and Pewter exit guidance while preserving the active semantic waypoint across recovery.
  • Distinguish the genuine east Route 4 cave exit from Route 4's west entrance: Mt. Moon completion now requires B2F entry and the post-cave `(24, 6)` tile.

Release 12.1.53

12.1.53

No runs loaded

Route 3 recovery closure

  • Extend the code-owned Route 3 recovery profile from the west transition to the north, middle, and south interior corridors observed in the bounded 12.1.52 milestone cohort, then retreat through Pewter to its registered Pokemon Center without opening field-item menus. Keep trainer-flag variants with no battle-free westbound path outside the authorized trial scope.
  • Cover every 12.1.52 `route_out_of_range` terminal coordinate with a verified branch and replay every Charmander Route 3 battle-end state, plus representative live states from the other map components, through a battle-free Center heal before authorizing its bounded milestone trial.

Release 12.1.52

12.1.52

No runs loaded

Milestone cohort safety

  • Read the Gen-I bag from its authoritative item count and `$FF` terminator, failing closed on malformed slots instead of retaining sentinel or stale storage as phantom inventory and item-transition evidence.
  • Block ordinary `A` confirmation on the trainer-battle ITEM command whenever battle items are disabled, preserving the existing ten-action checkpoint interface while preventing raw menu navigation from bypassing Nuzlocke item rules. Keep legal wild-catch macros and the RUN column unchanged.
  • Preserve navigation-first recovery when a weakened party is outside a verified corridor: return control to the policy under the bounded route-gap guard instead of immediately opening the less reliable field-healing menu.
  • Report battle-item and invalid-catch episode endings in milestone failure frontiers so semantic review cannot mistake those violations for an empty failure surface.
  • Preserve a fresh one-use milestone approval through prelaunch processing; the terminal health that created its review block now becomes the next cohort baseline instead of replacing the approved block before launch.

Release 12.1.51

12.1.51

No runs loaded

Verified Mt. Moon recovery

  • Close the Mt. Moon recovery preflight with ROM-replayed routes for all three cave floors and Route 4's west approach to the registered Center. Select disconnected B1F/B2F components independently and require Route 4 map 15 as part of the milestone's complete recovery chain.
  • Re-select a persisted recovery branch after a same-map warp coordinate settles, while preserving ordinary one-tile movement and two-tile ledges. Exhaustive component checks and genuine low-HP and poisoned-menu PyBoy replays now cover the cave-to-Center retreat without trainer engagement.

Canary runtime and telemetry

  • Resolve clean-start checkpoints from the held clean-opening lane instead of the milestone supervisor's top-level stage, so starters 1, 4, and 7 evaluate their Stage-10 checkpoints while the downstream lane is independently held.
  • Publish live canary status and run heartbeats with exact completed, error, interrupted, and incompatible counts, checkpoint provenance, coordinates, and distinct attempt outcomes. Keep the reviewed 16,000-step progress lease unchanged while labeling lease and step-limit endings as bounded recycling.
  • Atomically claim cloud outbox envelopes, preserve producer replacements, suppress superseded mutable run state, and identify uploader release/PID so the guardian can perform a verified two-tick release handoff.

Web canary status

  • Expose the latest canary heartbeat separately from training metadata and add a dashboard card for evaluator release, checkpoint model, freshness, current attempt, proof counts, and bounded-recycle outcomes.
  • Recover legacy endpoint coordinates from nested final locations, retain completed canary rows that are not spatially mappable, and keep the public endpoint-kind contract limited to objective, violation, or other.

Orchestration safety

  • Preserve milestone semantic-review blocks across release and configuration changes. A held starter now requires a one-use, release-scoped `POKEMON_MILESTONE_TRIAL_STARTERS` approval, and every approved bounded cohort returns to review instead of authorizing another launch.

Release 12.1.50

12.1.50

No runs loaded

Guarded Route 3 correction

  • Record the live 12.1.49 Route 3 failure mode: low-HP post-Brock starts could enter the field-Potion macro before their verified Pewter recovery route, and bounded menu cleanup could abort the attempt with movement still unsafe.
  • Make navigation-first recovery a code-owned profile invariant, so the real Bulbasaur, Charmander, and Squirtle Stage-9 starts follow their verified Gym and Pewter Center corridors before field healing is considered.
  • Define the corrected-release resume contract: the first 12.1.50 Route 3 cohort selects the validated immediate Stage-9 predecessor instead of any interrupted 12.1.49 same-stage checkpoint. Tracked milestone checkpoints are eligible only when health proves both their release and model provenance match the current source release; unknown or prior-release checkpoints fall back to the exact validated predecessor without deleting their artifacts.

Release 12.1.49

12.1.49

No runs loaded

Guarded post-Brock transfer

  • Pin every downstream starter track to the audited Stage-9 PPO interface: twelve ordered memory features over four frames, the same ten ordered actions and transition semantics, and a rollout-compatible batch size of 512. Reject an incompatible resume before creating any PyBoy worker and isolate ROM/RAM worker namespaces by starter and stage.
  • Require one candidate-backed, starter-bound model/state/promotion triple from the exact preceding milestone. Revalidate party identities and HP, encounter/catch/death ledgers, ruleset, release, transition evidence, and promotion authority before skip, routing, specialist publication, or the next stage; final consolidation also requires the complete prior chain.
  • Extend in-episode map, item, badge, Gym identity/roster, and level-cap proof through every configured milestone instead of accepting pre-satisfied endpoint states.

Recovery and supervisor fail-closed guards

  • Add PyBoy-verified recovery from all three real post-Brock Gym states and from Route 3's west entrance back through Pewter to the Center. Materialized stages now declare required recovery maps; Mt. Moon, Bill, and Misty remain explicitly blocked until each missing route is emulator-verified.
  • Account for every live trainer globally, checkpoint duplicate runners or workers and mixed/oversubscribed lanes, stop consolidation when its clean gate closes, and persist startup failures or ten-minute no-health timeouts so a release cannot crash-loop.
  • Scope Stage-10 approval to the exact durable semantic block, require a current-release fully finished bounded cohort before milestone renewal, and keep prior-release health from becoming a new baseline.

Release 12.1.48

12.1.48

No runs loaded

Post-Brock milestone lane

  • Decouple audited post-Brock milestone training from the strict clean-opening rehearsal gate, preserving Stage 10 as independent release and final consolidation evidence while starter-specific tracks advance in stage-major lockstep toward Misty and later milestones.
  • Seed the new versioned full-game tracks only from validated paired Stage-9 artifacts, reserve promotion authority for the immediate predecessor state, and retain older states as training-only rehearsal.
  • Bound each milestone launch to one rollout-aligned cohort below 500K steps, share the fixed nine-environment pool without oversubscription, and hold a non-advancing durable frontier for semantic review.

Transition and safety evidence

  • Record episode-start badges, map and item transitions, badge deltas, and recognized Gym Leader entry evidence in promotion sidecars and episode telemetry so pre-satisfied replay states cannot manufacture progress.
  • Require Misty promotion to prove a legal one-to-two badge transition after a recognized level-21-cap battle, and carry the verified Pewter recovery route into the first post-Brock stage.
  • Report clean-opening and downstream milestone lanes independently, including their active, held, and next-action state.

Release 12.1.47

12.1.47

No runs loaded

Stage-10 reporting clarity

  • Report completed `progress_stalled` attempts as bounded progress-lease recycling under monitoring, not as trainer-stuck alerts unless the supervisor has emitted an explicit current semantic-review block.
  • Show active clean semantic progress separately as provisional liveness that cannot satisfy the durable opening gate until a valid clean episode ends.
  • Distinguish prior Stage-9 Brock proof from current Stage-10 canary successes, legal opening rehearsals, and distinct-team requirements in authoritative and scheduled status language.

Stage-10 semantic phase control

  • Defer forced wild-battle readiness training until the audited first-capture handoff, so Route 1 encounters before waypoint 28 are fled instead of consuming the bounded opening cohort on premature starter grinding.
  • Preserve legal catch precedence at the handoff and fail closed when the live semantic waypoint cannot be interpreted.

Bounded flee recovery

  • Abort an attempt safely after three consecutive flee-menu synchronization timeouts in one wild battle, rather than allowing a bad command or cursor state to consume the remainder of its progress lease.
  • Report exhausted flee synchronization as its own stable environmental ending and cover command-menu, cursor-normalization, and outcome-timeout paths.

Release 12.1.46

12.1.46

No runs loaded

Bounded cohort integrity

  • Stop treating a live PID with terminal or mixed-release health as a healthy member of the starter pool; give a newly launched worker one observation to publish running health, then drain a persistently finalized worker so its allocation can be relaunched on the current release, while status reports the bounded recovery or release handoff as an expected stop.
  • Preserve each live Stage-10 worker's environment width when another starter exits, leaving freed slots idle instead of restarting survivors with another full budget; hold an interrupted current-release cohort for semantic review rather than automatically granting it another 497,664 steps.
  • Fail closed on current-release running health with no live worker and on an already oversubscribed Stage-10 pool, checkpointing the latter instead of accepting more than the fixed nine-environment budget.
  • Fit Stage 10's per-episode hard limit inside each worker's share of the fixed 497,664-step cohort, guaranteeing at least one completed episode per worker even when one starter receives all nine environments.
  • Keep the semantic plateau baseline strictly durable. A valid clean episode already active at the deadline receives one fair completion opportunity, while invalidation, terminal health, or a post-deadline retry cannot renew or repeatedly defer the review boundary.

Release 12.1.45

12.1.45

No runs loaded

Clean-frontier liveness

  • Publish a provisional semantic frontier only from active, rule-valid, promotion-authority clean episodes, allowing the aggregate Stage-10 plateau monitor to observe real progress before a vectorized episode finishes.
  • Keep provisional progress out of promotion and opening-gate evidence while using it for supervisor and hourly-monitor liveness, so replay workers still cannot advance or reset the clean semantic contract.
  • Persist clean map ID and progression priority for completed and in-flight clean episodes, and include map priority in plateau comparisons and status.

Release 12.1.44

12.1.44

No runs loaded

Final semantic safeguards

  • Treat a field heal as successful only after both the Potion effect and menu closure are verified; abort to a safe checkpoint if bounded cleanup leaves a menu open, before recovery navigation can emit directional input.
  • Distinguish Route 2's disconnected southern and northern regions even though they share map ID 13, routing a regressed northern-phase attempt back through Viridian Forest without changing its semantic waypoint index.
  • Match the clean-start canary's policy-visible episode-progress normalization to Stage 10's 30,000-step horizon, avoiding a canary-only observation shift.
  • Honor an explicitly supplied callback release throughout clean-frontier restoration and health, summary, and run-stat persistence.
  • Label a stopped trainer as an intentional semantic hold or release handoff only when the supervisor record is fresh, current-release, and explicit.

Release 12.1.43

12.1.43

No runs loaded

Stage-10 semantic recovery

  • Reset the meaningful-progress distance baseline after each ordered waypoint, and use monotonic four-tile approach credits so a completed landmark cannot suppress legitimate progress toward the next one.
  • Project waypoint features onto the latest phase-consistent transition when a policy regresses to another map, with explicit exits for Route 22 and optional Viridian interiors instead of comparing unrelated map coordinates.
  • Bound failed field-Potion attempts, fall back immediately to verified Center navigation, and apply a retry cooldown so an unavailable Start menu cannot monopolize the remainder of an episode.
  • Route low-HP Viridian Mart replay starts through the stable south exit before continuing along the existing Viridian-to-Center recovery corridor.
  • Start the repaired release from a fresh 497,664-step budget: exactly 108 complete nine-environment PPO rollouts, below the reviewed 500K ceiling.

Evaluation and supervision integrity

  • Persist a per-release semantic frontier assembled only from completed, rule-valid promotion-authority episodes, and use it for Stage-10 supervision, status, and stuckness checks so replay-ladder maxima cannot reset the gate.
  • Match the clean-start canary to Stage 10's 12K floor, 16K progress lease, four-tile distance interval, and story renewals while retaining its 60K observation scale and 500K absolute safety cap.
  • Exclude exhausted health from extension counts unless it belongs to the active release, preventing old terminal cohorts from consuming a new release's bounded retry allowance.
  • Report a trainer intentionally stopped for semantic review or release handoff as held instead of raising a false dead-screen runtime alarm.

Release 12.1.41

12.1.41

No runs loaded

Stage-10 semantic frontier

  • Hold the Parcel, delivery, and Poké Ball waypoints until their underlying story transitions actually complete, preserving route guidance through the Mart and Oak's Lab interactions.
  • Renew bounded episodes on story progress and allow a 16K-step progress gap, covering the measured clean Pallet-to-Lab frontier while keeping the current-release Charmander trial below its aggregate 500K ceiling.
  • Rank promotion-authority clean checkpoints by ordered waypoint progress before incidental map discovery, so detours cannot displace a stronger full-opening route checkpoint.
  • Record Parcel delivery once in the canary milestone buffer instead of repeating the durable flag on every subsequent step.
  • Distinguish a selected active bounded trial from the starters still held for semantic review in the authoritative next-action report.

Release 12.1.40

12.1.40

No runs loaded

Stage-10 semantic recovery

  • Restore the audited Mart, Pallet, Route 2, and Viridian Forest navigation frontiers to the full-opening route, and add the missing reach-Mart replay rung so the shared Viridian failure is practiced directly.
  • Preserve the first worker-sharded replay slots by leaving the initial reset to the vector environment, interleave clean starts throughout the weighted ladder, and recycle non-progressing openings within the bounded cohort.
  • Separate Brock-engagement shaping from the badge-confirmed terminal frontier, and classify battle wins, losses, and escapes from the game's battle-result byte instead of an opponent-HP proxy when available.

Evaluation integrity

  • Save the best promotion-authority clean-start checkpoint and evaluate that exact immutable checkpoint cohort, preventing stale or changing checkpoint failures from stopping a new model.
  • Scope clean-attempt telemetry to the current release, normalize team compositions independent of party order, and retain faint evidence across the complete episode for trustworthy promotion review.

Release 12.1.39

12.1.39

No runs loaded

Stage-10 training integrity

  • Replace random and progressive Stage-10 rebasing with a deterministic, worker-sharded ladder cycle that audits every fixed phase once and reserves twelve canonical clean starts for the ten-rehearsal promotion gate.
  • Start each new bounded release from its starter's promoted Stage-9 model, while preserving current-release checkpoint resume after interruption, and cap the catch/training roster at three members for attainable team diversity.
  • Keep Brock's final waypoint active until a trainer battle actually begins, prevent off-phase map discoveries from renewing the progress lease, and normalize ordinary stalled returns without weakening Nuzlocke penalties.

Evaluation integrity

  • Keep held Stage-10 tracks authoritative for clean-start canary selection and scope canary summaries to Stage 10, starter, evaluator release, and frozen model release.
  • Apply the ten-attempt canary stop gate per starter so one branch cannot stop or clear another, and report bounded evaluation holds without mislabeling them as zero-step semantic plateaus.

Release 12.1.38

12.1.38

No runs loaded

Recovery reliability

  • Stop the synchronized Center-PC opener on the first actionable storage-menu frame, preventing an extra confirmation from selecting an empty Withdraw box before the macro can choose Deposit.
  • Cover the corrected transition with both a deterministic menu-boundary test and a real PyBoy replay of the captured post-faint Pewter recovery state.

Supervision integrity

  • Hold a finished Stage-10 cohort for evaluation only after it reaches the configured training budget, so an unexpected under-budget clean exit remains eligible for supervised continuation.
  • Quarantine the bounded 12.1.37 cohort after its telemetry exposed the PC menu defect, retaining its atomic terminal checkpoints for diagnosis only.

Release 12.1.37

12.1.37

No runs loaded

Execution provenance

  • Create and rotate run/execution identity for local-only training as well as cloud-backed runs, so every terminal checkpoint manifest is attributable to one concrete execution without enabling external uploads.
  • Start the final bounded cohort from the unchanged Stage-9 checkpoints after an emulator-backed 12.1.36 shutdown smoke verified both atomic model ZIPs; the identity-less smoke artifacts remain ineligible for evaluation.

Release 12.1.36

12.1.36

No runs loaded

Checkpoint durability

  • Install the shutdown guard before emulator or model startup, immediately unwind blocked vector-environment reads on `SIGINT` or `SIGTERM`, and defer repeated signals until checkpoint commit finishes.
  • Stage and validate both `latest.zip` archives before atomic replacement, then publish terminal health with their exact paths, sizes, timestamps, and SHA-256 identities as the final execution commit.

Evaluation integrity

  • Require the paired evaluator's fresh model to exactly match the current terminal checkpoint manifest and reject missing, stale, active, malformed, or identity-mismatched executions before creating evaluation output.
  • Start a fresh release-scoped cohort from the unchanged Stage-9 checkpoints; the short multi-execution 12.1.35 artifacts remain quarantined.

Release 12.1.35

12.1.35

No runs loaded

Recovery reliability

  • Route low-HP Oak's Lab replay states through a collision-verified Pallet and Route 1 corridor to Viridian Center, and bound failed PC deposit macros so a broken menu interaction cannot pin a training worker indefinitely.
  • Save the current model and terminal health after either `SIGINT` or `SIGTERM`, preserving a trustworthy checkpoint when the supervisor stops a stale or plateaued cohort.

Evaluation and supervision

  • Revalidate existing Stage-10 promotions against the current static-start, release, candidate-sidecar, and distinct-team contract before advancing.
  • Hold completed bounded Stage-10 cohorts for evaluation instead of silently relaunching them, and require a stable terminal current-release checkpoint with no active trainer before the paired evaluator can freeze it.
  • Start a fresh release-scoped cohort so interrupted 12.1.34 checkpoints and failure evidence cannot satisfy the repaired evaluation or promotion gates.

Release 12.1.34

12.1.34

Recovery reliability

  • Advance only framebuffer-verified post-battle dialogue before recovery movement, without allowing generic collision handling to start NPC text.
  • Route standard Pokemon Center PC traffic through the collision-verified row-5 aisle in both directions, avoiding the blocked row-7 furniture.
  • Start a fresh release-scoped cohort so failed recovery evidence from the earlier routes cannot satisfy replay, evaluation, or promotion gates.

Release 12.1.33

12.1.33

Reward integrity

  • Clamp every failed or invalid Stage-10 episode's complete return below zero, even when accumulated route and Gym rewards exceed its terminal penalty.
  • Start a fresh release-scoped cohort so checkpoints and promotion evidence produced before the cumulative-return fix cannot satisfy the repaired gate.

Release 12.1.32

12.1.32

Training recovery

  • Prevent Route 1 recovery from starting fresh NPC dialogue, verify both failure hot spots with real PyBoy save states, and cover the complete low-HP trip to a successful Viridian Center heal.
  • Replace the monolithic opening rehearsal with a bounded, late-heavy state ladder, explicit Stage-10 waypoint rebasing, Gym-interior guidance, and readiness-aligned catch/train/flee controls for each starter.
  • Remove repeatable battle, level, flee, and healing reward proxies, quarantine older replay captures, and reserve promotion authority for canonical clean Route 1 starts.

Monitoring

  • Checkpoint and hold Stage-10 tracks after a bounded semantic plateau or a current-release 0/30 clean-start canary, while reporting parcel, inventory, party, catch, waypoint, badge, and legal-success frontiers.

Release 12.1.31

12.1.31

Training reliability

  • Keep the advanced level cap active between a verified Gym Leader defeat and the delayed badge-memory update, so legal Brock experience cannot terminate an otherwise successful attempt.
  • Synchronize ordinary and forced battle switches against exact menu RAM, accept verified direct send-outs, and retry one dropped party-menu opening.
  • Concentrate Boulder Badge replay on clean post-Junior-Trainer heals, avoid fallback attacks when any party member has reached the badge cap, and expand stuck Route 1 recovery through prompt clearing and the full detour sequence.

Monitoring

  • Restrict clean-opening Brock stall alerts to opening-rehearsal canary attempts, while leaving general canary status and failure summaries intact.

Release 12.1.30

12.1.30

Training reliability

  • Replace the Viridian south-entry recovery shortcut rejected by live smoke testing with the collision-verified corridor around the Pokemon Center wall.

Monitoring

  • Bind Hall-of-Fame success to the configured terminal canary target and infer it from legacy records only when eight-badge evidence is present.

Release 12.1.29

12.1.29

Training reliability

  • Let a solo under-level starter keep ordinary battle control instead of entering a return-to-finisher rotation with no finisher, and add the calibrated Route 1 retreat plus a south-entry Viridian Center branch required by low-HP recovery.
  • Guide Boulder training from the Pewter Gym entrance through the Junior Trainer to Brock, flee irrelevant wild battles, and reject replay snapshots above the legal first-Gym entry cap.
  • Prefer the supervisor's live active-stage assignment over stale historical progress baselines when selecting clean-start canary checkpoints.

Monitoring

  • Report clean-opening Brock and Hall-of-Fame canary success separately, drive the opening-stall alert from the Brock milestone, and retain structured recovery and rotation details on terminal attempts.

Release 12.1.28

12.1.28

Training reliability

  • Require every party to enter a Gym Leader battle at or below the current badge-indexed level cap, then activate the next Gym cap during the identified Leader battle so experience earned against the Leader remains legal.
  • When a threatened active Pokémon has no healthy switch target during an identified Gym Leader battle, return control to the policy while preserving faint retirement, forced switching, whiteout, and promotion legality checks.

Monitoring

  • Alert on repeated current-release Brock failures, the final automatic budget extension, and a rolling zero-success clean-start canary instead of treating raw step and episode movement as sufficient progress.
  • Refresh the Gym Leader level-cap policy from the authoritative Nuzlocke template when materializing starter-only Pewter and Boulder stages.

Release 12.1.26

12.1.26

Training reliability

  • Track trainer victories separately from wild wins, require a legal Center heal after the Pewter Junior Trainer, and replay stable fully healed post-trainer checkpoints so later episodes can concentrate on Brock.
  • Add calibrated Pewter Gym exit branches and a collision-safe city-to-Center corridor, accept both verified nurse-facing coordinates, and bound missing, stuck, or non-progressing recovery routes with a safe checkpoint restart.

Telemetry and dashboard

  • Preserve terminal wild and trainer battle context—including full trainer rosters—and publish it through episode logs and clean-start canary results.
  • Show each run's endpoint tile, terminal battle type, and opponent party in the atlas run inspector, while retaining useful party HP in canary details.

Release 12.1.25

12.1.25

Training reliability

  • Reset stage-local waypoint state when entering Pewter recovery, the Brock battle, or a clean opening rehearsal, and prevent those stages from replaying snapshots whose waypoint indexes belong to a different objective.

Release 12.1.24

12.1.24

Evaluation

  • Give clean-start canary attempts a rolling progress lease beyond the original 60,000-step opening horizon, retain a 500,000-step hard cap, and continue after Brock with the existing full-game map progression toward Hall of Fame.
  • Preserve the checkpoint's calibrated pre-Brock waypoint observations and report Boulder Badge and progress-lease milestones separately from the final full-game canary objective.
  • Lock the badge-indexed level-cap schedule to each next Pokemon Red Gym Leader's highest-level Pokemon and cover all eight Gym transitions.
  • Report evaluator release and frozen-checkpoint model provenance separately, marking ambiguous legacy attempts as unreported instead of relabeling old weights.

Dashboard

  • Publish completed clean-start canary evaluations with explicit canary labels, provenance, and filtering, without mixing their sampled movement into the training heatmap.
  • Render every observed Pewter interior with tile-aligned map art, connected placement, and authoritative bounds that discard map-transition coordinates.
  • Drive Nuxt map rendering and asset staging from the generated background registry so newly mapped areas cannot silently lose their artwork on the web.

Release 12.1.23

12.1.23

Recovery integrity

  • Compose map-keyed and legacy-list Center additions over the verified standard registry, preserving future-city overworld routes during profile application.
  • Make registered layout routes authoritative inside Centers, reject malformed runtime routes safely, and prevent item macros from targeting retired party members even when stale game bytes show restored HP or status.
  • Allow config-only materialization to refresh future stages before predecessor promotion, so dormant direct-resume configs retain the current recovery rules.
  • Evaluate readiness and badge gates across the current surviving roster rather than requiring the original party size after a mandatory death deposit.

Release 12.1.22

12.1.22

Nuzlocke integrity

  • Make survivor continuation a runtime invariant even for stale materialized configs: a single faint is retired and recovery-gated, while only a full party whiteout is terminal.
  • Apply one shared survivor-recovery profile to generic curriculum stages, starter tracks, and Brock races so future materialization cannot silently disable forced switching, PC deposit, healing, or verified route recovery.
  • Validate normalized Center profiles, targets, and interior route contracts at environment startup; registered layouts are authoritative and uncalibrated custom interiors fail closed instead of inheriting standard coordinates.

Release 12.1.21

12.1.21

Nuzlocke integrity

  • Treat an individual faint as permanent retirement instead of ending the run, force a living replacement, and require exact-identity PC deposit before ordinary progress or checkpoint promotion resumes. Only a full-party whiteout now ends an attempt for fainting.
  • Synchronize RUN against the verified battle command menu and Gen I result markers, rechecking party safety after a denied escape so transitional inputs cannot accidentally choose FIGHT or sacrifice a weakened lead.
  • Normalize all eleven verified standard Pokémon Centers through one reusable layout profile with entrance-to-PC and post-deposit PC-to-nurse routing; custom interiors must declare calibrated profiles explicitly.

Release 12.1.20

12.1.20

Training

  • Keep under-ready reach-Pewter replay states in the verified Forest training corridor, rank party-floor gains ahead of route depth, and renew the progress lease when a catch or level gain advances the roster.
  • Return a switch-trained party member to a healthy finisher before attacking, synchronize both persistent Gen I party-menu cursors from RAM, suppress failed rotation loops for the encounter, and bound failed recovery with a safe escape instead of spending the episode on status moves.
  • Gate the Forest exit and Pewter Gym entrance on the configured roster, readiness, and recovery requirements; Brock remains a separate stage after a clean Pewter Center handoff.

Reliability

  • Yield recovery navigation to the policy outside verified corridors, ignore status-only moves when evaluating battle PP, scope the Pewter recovery gate to the Gym door, and remove the route lease from recovery and Brock stages.
  • Align Boulder Badge promotion on one clean legal win across the trainer, supervisor, and status paths.

Release 12.1.19

12.1.19

No runs loaded

Orchestration

  • Report fresh health-backed starter allocations when process enumeration is unavailable, and fail closed on stale unverified health rather than claiming a dead trainer is active or risking a duplicate launch.

Training

  • Safely train a newly caught Forest party member to the existing party-wide readiness floor during reach-Pewter replay, then resume fleeing optional encounters once the full party is ready.

Release 12.1.18

12.1.18

Orchestration

  • Add a five-minute cron guardian for the inference-only clean-start canary, with singleton process checks, current-release heartbeat verification, clean release handoff, and canary release state in the regular status report.

Training

  • Reorder the party from the overworld when its lead approaches the current badge level cap, choosing the weakest safe replacement and verifying the exact Pokémon identity swap before continuing.
  • Keep starter level, move, promotion, and replay checks independent of party order so a deliberately rotated lead remains eligible for the curriculum.

Telemetry

  • Record a catch on the first sampled action after reset while excluding starter pickups, gifts, and PC withdrawals from catch-location events.
  • Count verified field lead rotations and failures, force each attempt into the step trace, and expose its cap, source, target, and identity details in run health telemetry.

Release 12.1.17

12.1.17

No runs loaded

Training integrity

  • Preserve the encounter species and location through delayed and full-party capture confirmation, preventing a later field read from losing or misattributing the catch.
  • Require the Boulder Badge in opening-rehearsal states so downstream runs cannot satisfy the rehearsal gate from a pre-Brock snapshot.

Evaluation

  • Add a frozen-checkpoint clean-start canary that rotates across all starter branches without training or promotion authority and reports generalization progress separately from advanced-state curriculum results.

Telemetry

  • Record each confirmed capture as its own episode event with species, encounter area, tile coordinates, and step so capture origins stay accurate even when the episode ends elsewhere or a full party sends the Pokémon to the PC.

Web atlas

  • Add the episode quilt and Kanto ghost map for exploring run outcomes and route coverage across the starter branches.
  • Show the caught area and tile in the Agent Pokédex and run details, while clearly labeling older captures whose exact location was never recorded.

Release 12.1.16

12.1.16

Orchestration

  • Refresh the shared Stage 5 model/state pair from a deterministic validated current-ruleset starter branch once all three capture candidates are present. This prevents a legacy aggregate artifact from crashing every battle-readiness launch before emulator environments are allocated.

Release 12.1.15

12.1.15

Nuzlocke integrity

  • Register an encounter area atomically before any catch, flee, switch, or attack macro, enforce Oak's Parcel at both the action and inventory layers, and preserve first-frame catches without misclassifying the caught species as a duplicate.
  • Retire blanket Viridian Forest ledger reopening. New runs catch the first legal Forest encounter normally, while pre-12.1.15 replay captures and promotions remain on disk but are excluded from new ruleset runs.
  • Reject ambiguous or illegal catch-training snapshots, including pre-Parcel, no-Ball, already-spent, fainted, and incompletely described encounters.
  • Match the strict web ruleset across curricula and the Brock race: any faint ends the attempt, and promotion cannot carry a death ledger forward.
  • Isolate Brock-race training captures by starter and stamp replay sidecars and promotions with the release/ruleset version used to create them.
  • Scope supervisor relaunch and extension limits to the active release/ruleset so exhausted historical checkpoints cannot block required legal retraining.

Web atlas

  • Clarify that an empty model-filtered Pokédex means that release recorded no new catch event; inherited party ledgers are not misattributed as new catches.

Release 12.1.14

12.1.14

No runs loaded

Training

  • Catch the first eligible non-duplicate Viridian Forest encounter during reach-Pewter navigation, while continuing to flee duplicates and already spent encounters.
  • Reopen Forest encounter ledgers polluted by the earlier forced-flee policy until a third species is caught, and retain three-member parties in progressive replay.

Release 12.1.13

12.1.13

No runs loaded

Training

  • Defer replay-state capture until menus, dialogue, fades, and battle transitions have settled; sanitize inherited reset states and exclude legacy transitional captures from progressive replay.
  • Restore the verified Gen I PKMN command path for safety switches.
  • Prevent recovery detours from immediately reversing route progress, and cancel replayed overlays before retrying Pokémon Center positioning.

Release 12.1.12

12.1.12

No runs loaded

Training

  • Give reach-Pewter attempts both a 12,000-step initial floor and rolling meaningful-progress renewal window, allowing longer Viridian Forest traversals.

Release 12.1.11

12.1.11

No runs loaded

Training

  • Correct the Gen I safety-switch command path to open PKMN instead of ITEM, and verify the selected party member actually became the battle mon.
  • Make reach-Pewter workers flee optional wild encounters before safety switching, preventing avoidable faints and pre-Brock level-cap violations.
  • Replay clean, level-capped route frontiers by ordered-waypoint progress so Forest navigation can continue from earned deep states instead of restarting from Route 1 every episode.
  • Add a southbound Viridian Forest recovery corridor for weakened or poisoned parties returning to the Viridian Pokémon Center.

Release 12.1.10

12.1.10

No runs loaded

Training

  • Resume reach-Pewter navigation from the Viridian waypoint after a center heal, preventing Route 2 and Forest checkpoints from being replayed in the wrong map.
  • Constrain non-objective battle controls to a bag-safe attack macro through the Brock and opening-rehearsal stages, while preserving legal catch macros.
  • Treat battles that finish during a safety-switch macro as safely resolved instead of terminating the episode as a failed switch.

Telemetry

  • Separate center-positioning retries from real nurse-dialogue heal attempts and failures.
  • Publish current-release environmental termination counters and reset runtime safety counters across releases so historical incidents no longer appear as active regressions.
  • Clarify party-level and legal-objective labels in authoritative status.

Release 12.1.9

12.1.9

No runs loaded

Training

  • Replace the fixed 30,000-step reach-Pewter horizon with a progress-aware lease: meaningful route progress renews the attempt up to a 60,000-step hard cap, while workers that stop advancing are recycled after 5,000 stale steps.
  • Keep stall-penalty clocks separate from true progress clocks so periodic penalties can no longer prevent configured early termination.
  • Add a deterministic Route 22-to-Viridian recovery handoff and ordered waypoints so Charmander's legal readiness promotion can rejoin the Route 2 path instead of wandering off-route.

Telemetry

  • Record progress-lease expirations as `progress_stalled`, including the last meaningful milestone, rolling deadline, and stale duration.

Release 12.1.8

12.1.8

No runs loaded

Training

  • Correct the Gen I battle-menu ITEM slot guard so spent-route and duplicate encounters cannot consume Poké Balls through raw menu inputs, while legal first non-duplicate encounters remain catchable.

Telemetry

  • Break environmental episode endings into stable reason codes for field safety, failed battle switches, early aborts, exploration stalls, battle limits, blackouts, and unknown terminal conditions; retain concise diagnostic details and cumulative counts in episode, health, summary, atlas, supervisor, and scheduled status data.

Release 12.1.7

12.1.7

No runs loaded

Training

  • Extend battle-readiness episodes from 30,000 to 40,000 steps so party-level objectives have enough time to complete.
  • Let untreated overworld poison resolve as a faint instead of preemptively aborting the episode, while continuing to use an Antidote when available.
  • End battle-readiness episodes on the first faint under Nuzlocke permadeath so the policy receives the configured faint penalty and a clear terminal signal.

Dashboard

  • Migrate the generated run dashboard to a component-based Nuxt 4 application while retaining Tailwind, Cloudflare Workers static-asset deployment, and a production Docker image.
  • Move build and deployment ownership to Cloudflare Workers Builds and remove the duplicate GitHub Actions deployment.
  • Make the scheduled atlas refresh a direct JSON-only command with schema validation before it commits and triggers the Nuxt deployment workflow.
  • Stop tracking local runtime logs, ROM data, raw walkthrough imports, and generated Python package metadata; compact the published atlas JSON.
  • Select the newest recorded model version when the web atlas first loads, while keeping an “All model versions” scope available.
  • Let the endpoint inspector switch from a focused marker back to all endpoints and temporarily show attempts from every model version without changing the current model scope for the rest of the atlas.
  • Add cursor-centered wheel, trackpad, pinch, double-click, and keyboard zoom to the connected world atlas while preserving drag-to-pan navigation.
  • Open any listed attempt in an accessible run-details modal, including runs referenced by endpoint results and Pokédex catch history.
  • Publish this version changelog in the web atlas.
  • Replace the route-by-route world overview with one connected, zoomable atlas and add missing Route 22 and Viridian Forest backdrops.

Release 12.1.6

12.1.6

No runs loaded

Training

  • Synchronized field-Potion use to Pokémon Red's live Start, bag, item-action, and party cursors; verified exactly one Potion is consumed and the menu returns to the overworld.
  • Recorded deterministic low-HP return routes from Route 2 and Viridian City to the Viridian Pokémon Center, including exact nurse positioning and full-party heal verification.

Telemetry

  • Added a dedicated healing event stream and persistent counters for Center attempts and failures, item heals, Potions consumed, and recovery-route starts and arrivals.

Release 12.1.5

12.1.5

No runs loaded

Training

  • Added low-HP and poison recovery navigation, legal field healing, and the Route 1 Potion gift to the battle-readiness route.
  • Hardened forced-switch handling and battle-menu synchronization so a fainted active party member does not strand an otherwise legal attempt.
  • Added release-aware supervisor handoffs and checkpoint-safe restarts for battle-readiness workers that stop making measurable progress.

Telemetry

  • Reported recovery, progressive replay, healing, and supervisor handoff state in run health and status output.

Release 12.1.4

12.1.4

No runs loaded

Training

  • Corrected battle-readiness waypoint restoration so Center-started attempts resume the intended route instead of inheriting a completed route position.
  • Restarted the configured waypoint cycle after a successful Pokémon Center heal.

Release 12.1.3

12.1.3

No runs loaded

Training

  • Added progressive replay from validated intermediate battle-readiness states to expose the policy to more of the Route 2 and Forest approach.
  • Rejected replay states that were unsafe, unadvanced, or incompatible with the active party and Nuzlocke constraints.

Release 12.1.2

12.1.2

No runs loaded

Training

  • Made PPO minibatches compatible with the supervisor’s 3, 4, 5, and 9 environment allocations.
  • Preserved checkpoint/resume continuity as the shared nine-environment pool was redistributed between starter tracks.

Release 12.1.1

12.1.1

No runs loaded

Orchestration

  • Tightened per-starter process ownership, checkpoint handoffs, and stage-major lockstep for Bulbasaur, Charmander, and Squirtle.
  • Improved promotion and resume bookkeeping for repeated battle-readiness extensions.

Release 12.1.0

12.1.0

No runs loaded

Orchestration

  • Added the starter-track supervisor, structured health reporting, and Brock race tooling.
  • Added stale-observation recovery and safer starter lifecycle management.

Training

  • Continued each starter independently through battle readiness while preserving validated curriculum promotions.

Release 12.0.1

12.0.1

No runs loaded

Training

  • Added per-starter curriculum runners and a clean promotion path through the Viridian story stages.
  • Expanded progression-safety coverage for supervisor, environment, reward, logging, and release behavior.

Dashboard

  • Added richer map backdrops and improved run-atlas episode telemetry.