ADSP-BF518C++98 embeddedFloat training + Q15 inferenceA/B persistenceGrow ↔ ShrinkSignature-driven policyv0.25 CONSOLIDATED R3

otTINN

Autonomous Neural Lifecycle on a resource-constrained DSP. A bounded, persistent and auditable system that can detect degradation, retrain, change neural topology, compare quality against computational cost, promote or rollback a Candidate, simplify itself after sustained stability, and remember failed simplification targets. R3 formalizes a multidimensional Model Performance Signature as the policy interface while preserving the R2 BF518 deployment-latency normalization.

Design thesis. Adaptation is useful only if authority remains explicit. R3 separates measurement, model identity, performance signature, policy, persistence and role transitions. It deliberately avoids an aggregate “fitness score”, opaque ranking and uncontrolled architecture search.
BF518 hardware correction retained from R2. ACTIVE and CANDIDATE cannot simultaneously receive equivalent scarce fast-memory residency. Shadow execution therefore remains conservative, but Candidate policy latency is measured separately in an “as-if-ACTIVE” deployment condition. R3 carries this forward unchanged and records the normalized deployment view inside the Candidate performance signature.
R3 production consolidation. R3 integrates SIMMEM1, reboot/resume hardening, otFILE v2.0.2 safe line input and SIGNATURE2. The clean production endurance campaign has now completed on the real BF518: 12 h 35 min 10 s, 9,568,406 samples, 797 full world cycles + 4,406 samples, 2 promotions, 0 rejects, 0 rollbacks and 32 training checkpoints. Neural thresholds and promotion semantics remained unchanged.

Contents

1. Architecture overview

Numerical core

1–4

hidden layers; float forward/backprop, exact checkpoint/resume, generic Q15 inference.

Lifecycle roles

A / C / P

ACTIVE, CANDIDATE and PREVIOUS are explicit roles; ACTIVE remains authoritative until promotion.

Decision model

Signature → policy

Independent gates + meaningful improvement + Pareto/dominance reasoning consume multidimensional model signatures; no scalar score.

Embedded discipline

4 KiB

BF518 stack limit drives placement of persistent/large objects outside L1 stack.

Data / Sensorsscenario or plant Neural Corefloat + Q15 Monitorhealth / drift Autonomyone next action LifecycleACTIVE / CANDIDATE /PREVIOUS A/B Storagecatalog + models +endurance state

2. Module contracts

otNeuralNetwork

Owns numerical mechanics: topology, float inference/training, checkpoint representation, Q15 preparation/inference and warm-start modes. It does not decide whether a model deserves authority.

otNeuralMonitor

Consumes inputs and optional observed error. It classifies health from LEARNING to RETRAIN_REQUIRED. Input drift alone can warn, but error evidence is required for strong degradation when feedback exists.

otNeuralModelLifecycle

Owns model roles, objective reports and Model Performance Signatures. R3 policy decisions consume ACTIVE/CANDIDATE signatures while role transitions remain explicit.

otNeuralAutonomy

Control-plane state machine. It permits exactly one next lifecycle action and implements bounded grow/shrink policies. It owns no model payload or dataset.

otNeuralLifecycleStorage

Reboot-safe A/B persistence. State v6 carries lifecycle, simplification and outcome-memory data; boot coherence reconciles catalog roles and nextModelId before execution continues.

otNeuralTraining

Slice-based, resumable Candidate training. Checkpoint/resume is preferred over monolithic embedded training so adaptation can survive reset and fit idle windows.

otNeuralScenario

Deterministic synthetic environment used by endurance validation. Seed and cycle are logged so anomalies can be reconstructed.

main.cpp

Hardware-facing orchestration: SD, UART, timers, RTC, datasets, endurance counters, telemetry and long-run scheduling.

otFILE v2.0.2

Embedded file helper with failed-open leak fix and safe freadline(): caller-supplied buffer, guaranteed NUL termination, CR/LF handling, truncation detection and no per-line heap allocation.

3. Model lifecycle

ACTIVE Degradationretrain / grow Long stabilitysimplify / shrink CANDIDATE Train → Validate → Q15→ Shadow → Dominance REJECT → keep ACTIVE PROMOTE → Probationaccept or rollback
Authority invariant. ACTIVE remains authoritative until an explicit promotion transaction. During shadow, CANDIDATE has zero actuation authority. PREVIOUS is retained through probation so a rollback has a known-good target.

4. Architecture growth: bounded escalation

Growth is not neural architecture search. The planner has a finite deterministic sequence, a maximum of four hidden layers, width 96 and a 32,768-parameter envelope. Resource failures can teach an empirical ceiling so the next attempt may skip an infeasible wide topology and try a deeper-but-smaller one.

256 → 28 → 10
       │
       ├── WIDEN-PRESERVE → 256 → 48 → 10
       ├── further bounded width steps
       └── depth escalation → 256 → 48 → 24 → 10 → ... up to 4 hidden layers

Candidate initialization

5. Complexity-aware promotion without a fitness score

R3 preserves the existing decision semantics but changes the formal interface. The policy no longer consumes scattered validation/complexity values directly in the production path; it consumes one multidimensional signature for ACTIVE and one for CANDIDATE.

measurements
   ↓
ACTIVE Signature     CANDIDATE Signature
        \               /
         \             /
          → comparison ←
          → meaningful improvement
          → Pareto / dominance
          → independent hard gates
          → shadow authority gate
                    ↓
              FINAL DECISION

Hard gates

Deployability and resource envelopes remain independent: Q15 required, absolute memory/parameter/MAC/latency bounds and shadow authority rules.

Meaningful gain

Accuracy and mean-error improvements are judged independently. Tiny differences do not buy complexity.

Dominance

Pareto-like comparison can favor an equally capable but cheaper model. A bounded trade-off path allows real quality gain to justify limited extra cost.

Explainability

Every decision retains a stable reason code and explicit quality/resource deltas. SIGNATURE2 changes the interface, not the meaning.

BF518 latency normalization remains explicit

During normal shadow execution ACTIVE is prepared first and retains first claim on scarce fast Q15 memory. CANDIDATE is prepared second and may be MIXED or NORMAL. That shadow arrangement is correct for safety but is not representative of Candidate deployment after promotion.

shadow execution
ACTIVE      → first Q15 residency priority
CANDIDATE   → remaining residency, often slower

policy deployment view
ACTIVE      → measured ACTIVE latency
CANDIDATE   → temporarily prepared alone, measured as-if-ACTIVE
               then normal ACTIVE-first / CANDIDATE-second state is restored
TOPPA HARDWARE BF518. The 2× relative-cost rule was never relaxed. Only the physical condition under which Candidate deployment latency is measured was corrected. In R3 that normalized deployment view is carried into the Candidate signature, while shadow latency remains diagnostic.

6. Model Performance Signature

R3 formalizes the concept introduced during the R2 campaign: a model is not assigned a single efficiency or fitness score. It is represented by a multidimensional Model Performance Signature. The signature is the boundary between measurement and policy.

MODELACTIVE or CANDIDATE MODEL PERFORMANCE SIGNATURE IdentitymodelId · revision · topology Qualityfloat/Q15 accuracy · mean error Complexityparameters · MACs · depth · width Resourcesmemory · residency · deployment latency Lifecycleage · stable samples · applicability POLICYno aggregate score

Identity

modelId, revision and topology identify exactly which individual is being compared.

Quality

Float and Q15 accuracy/mean-error metrics remain explicit; quantized deployment is not allowed to hide behind float-only quality.

Complexity

Parameters, MACs, hidden-layer count and maximum hidden width describe representational/computational size.

Resources

Float/Q15 bytes, workspace, Q15 residency and deployment latency describe target cost. Candidate residency is normalized to its as-if-ACTIVE deployment condition.

Lifecycle

ACTIVE age and stable dwell are carried only when meaningful; non-applicable lifecycle fields are explicitly distinguished from a real value of zero.

No scalar collapse

No fitness, weightedScore or overallEfficiency is generated. Trade-offs remain visible and auditable.

SIGNATURE1 → SIGNATURE2

SIGNATURE1 first made the structure descriptive and proved that BF518 Candidate resource fields could represent deployment rather than shadow placement. SIGNATURE2 then made ACTIVE/CANDIDATE signatures the production policy inputs. Legacy report-based functions remain as regression oracles, but the BF518 endurance decision path is signature-driven.

BF518 smoke evidenceObserved resultMeaning
SIGNATURE1 deployment normalizationCandidate 256-24-10: shadow=MIXED, deployment=FAST, normalized=1, latency ≈263.52 µs, authority=0The signature describes post-promotion deployment resources rather than penalized shadow residency.
SIGNATURE2 policy interfacesignatureDecisionInput=1, dominance valid, gates valid, final decision valid; observed reason REJECT_ACTIVE_DOMINATESThe real BF518 policy consumed the two signatures and reached a normal final reason without giving Candidate authority.
Policy-equivalence regressionOld path == signature path for REJECT_RELATIVE_COST, PROMOTE_BOUNDED_TRADEOFF, REJECT_ACTIVE_DOMINATES, PROMOTE_PARETOThe interface changed; decision semantics did not.

7. Autonomous simplification and outcome memory

v0.25 closes the opposite side of adaptation: after sustained stability the system may propose a smaller Candidate. R3 retains the conservative trigger and adds SIMMEM1, a compact memory preventing repeated deterministic retries of a simplification already shown to fail from the same ACTIVE baseline.

stable ACTIVE
   ↓  stable ≥ 24,000 + age ≥ 48,000 + cooldown ≥ 48,000 + saving ≥ 5%
planSimplerTopology()
   ↓
PRUNE-COPY
   ↓
normal train → validate → Q15 → shadow → signature policy
   ├── PROMOTE → probation → smaller ACTIVE
   └── REJECT/ROLLBACK
          ↓
   SIMMEM1 remembers parent ACTIVE + target topology + outcome
          ↓
   identical retry blocked until ACTIVE changes
Preemption rule. Simplification is optimization, never emergency behavior. If monitor health rises above NOTICE the attempt is aborted; RETRAIN_REQUIRED immediately outranks shrink.

What the controlled Calm Tail test established

StepTopology / parametersObserved outcome
Initial accepted baseline256-28-10 / 7,450Reference model after R2 adaptation.
First shrink level256-24-10 / 6,386One early reject, then a later promotion and probation acceptance under sustained calm.
Second shrink level256-16-10 / 4,258Promoted and probation accepted. This is ≈42.8% fewer parameters than 7,450.
Next proposed level256-8-10 / 2,130Repeatedly rejected before SIMMEM1; after SIMMEM1 a reject is remembered and identical retries are blocked.

The experiment therefore found a current simplification floor empirically: 256-16-10 remained adequate while 256-8-10 did not. The floor was not hard-coded as a winner; it emerged from the ordinary validation/promotion policy.

SIMMEM1 validation. In the final targeted run, ACTIVE 1010 rejected 256-8-10 once. After more than 331,000 additional samples—almost seven 48k cooldowns—the status still reported attempts=1 and simplifyMemory[parent=1010 target=256-8-10 outcome=REJECT]. The redundant retry loop was eliminated.

8. Persistence, reboot and interrupted-write recovery

R3 treats reboot recovery as part of the lifecycle, not as an afterthought. The persistent layer uses A/B endurance state plus durable model catalogs because an unattended embedded target must survive reset at awkward points.

otFILE v2.0.2

Persistence testing exposed a low-level line-input defect: the older fgetline() path could leave text without deterministic NUL termination at EOL, causing valid A/B records to be rejected depending on adjacent memory contents. R3 fixes the primitive rather than patching each caller.

otFILE::freadline(buffer, size, file, '\n', stripEol=true, &truncated)

properties:
  caller-supplied buffer
  guaranteed NUL termination
  CR/LF handling
  EOF distinction
  truncation detection + remainder consumption
  no per-line heap allocation

The endurance reader uses a fixed SDRAM buffer and rejects truncated records deterministically. Later BF518 boots recovered both A and B as valid, selected the newest sequence and continued from the real catalog/state pair.

Observed recovery evidence. After the fix, boots repeatedly reported both A/B copies valid with increasing sequences (for example 1239/1240), correct newest-copy selection, catalog restore and coherent nextModelId.

9. BF518 memory strategy

MEM_L1_DATA_A

Only genuinely hot historical/core data. Persistent lifecycle objects are not casually moved here.

FAST_DATA_2 / L2

Control-plane state: autonomy controller, compact reports and small persistent control structures.

SDRAM_DATA

Large/cold objects: lifecycle storage, model instances, datasets, training/Q15 workspaces, restore/catalog structures.

FAST_CODE_2

Cold lifecycle/planning/regression functions that should not consume scarce L1 code space.

Non-negotiable constraint: BF518 application stack is 4 KiB. Large local arrays/models/workspaces are forbidden; the design explicitly uses static/global/heap/SDRAM/L2 placement.

10. Why these parameters?

This is the section most likely to be challenged in a review. Values are therefore classified by origin. Implementation-derived values come from what has actually been built and validated; project resource bounds deliberately cap autonomy; engineering heuristics are conservative initial choices and are candidates for revision from long-run evidence.

ParameterCurrent valueWhy this valueIf lowerIf higherOrigin
World cycle12,000 samplesOne complete deterministic scenario cycle; provides a natural sample-based time unit independent of RTC.Shorter cycles make regime changes/frequency of adaptation less representative.Longer cycles slow every long-run observation and adaptation episode.Derived from scenario design
Telemetry period250 samples48 snapshots per world cycle: enough to reconstruct transitions while keeping UART/SD traffic modest.More I/O and larger logs; can perturb timing.Coarser forensic timeline; short transitions may be less visible.Engineering trade-off
State save / log fsync1,000 samples12 recovery points per world cycle, with immediate fsync on critical lifecycle events.More SD wear and I/O overhead.More non-critical progress can be lost after power loss.Engineering trade-off
Training epochs8Enough repeated exposure for a Candidate update while keeping each adaptation bounded on BF518.Faster but may underfit a new regime.Longer adaptation latency and energy cost; more risk of over-specializing a short window.Initial heuristic, validated operationally
Training block8 samplesSmall work slice keeps training interruptible/resumable and friendly to idle/standby operation.Higher loop/checkpoint overhead.Longer blocking slices and less responsive control-plane behavior.Embedded scheduling choice
Training checkpoint128 trained samples16 blocks between checkpoints; limits lost work without writing SD every block.More SD writes and latency.More retraining work can be lost on reset.Engineering trade-off
Probation GOOD dwell1,000 samplesRequires sustained post-promotion health before PREVIOUS is retired.A bad promotion can become irreversible sooner.Keeps PREVIOUS alive longer and delays cleanup.Conservative lifecycle choice
Generic retry cooldown2,000 samplesPrevents immediate repeated adaptation attempts after a failure while remaining much shorter than simplification cooldown.Can thrash on persistent borderline conditions.Recovery to a genuinely changed environment becomes slower.Hysteresis choice
Monitor reference500 samplesNon-trivial baseline but only ~4.2% of a 12k world cycle, so startup is not dominated by calibration.Reference statistics become noisier.Longer LEARNING phase; slower startup/reboot recovery.Engineering trade-off
Monitor evaluation period32 samplesReduces decision churn and provides the unit for the persistent-DEGRADED escape.More sensitive to short fluctuations and higher control overhead.Slower reaction to real degradation.Engineering trade-off
Recent EW alpha0.020Recent profile spans roughly tens-to-a-hundred samples: quicker than the reference, slower than single-sample noise.More sluggish recent estimate.More reactive/noisy estimate.Signal-filtering heuristic
Input drift NOTICE / WARNING0.20 / 0.30Input distribution change is advisory; it can warn but cannot alone declare the model wrong.More nuisance warnings.May miss early covariate shift.Initial heuristic
Error ratios1.35 / 1.75 / 2.75 / 5.0Staged bands around learned reference error, separating advisory, warning, degraded and retrain states.Earlier retraining and more churn.Longer exposure to degraded performance before action.Initial heuristic; long-run tuning candidate
Rise / fall confirmations2 / 6Fast escalation, slower recovery: explicit hysteresis prevents state flicker.If rise=1, single spikes can escalate; if fall lower, recovery can chatter.Too many confirmations delay legitimate transitions.Hysteresis design
Validation samples100 minPrevents promotion decisions from being based on a handful of examples.Higher variance in quality estimates.Longer adaptation latency and dataset requirement.Conservative minimum evidence
Shadow samples100 minCandidate must coexist with ACTIVE long enough to observe behavior before authority changes.Less evidence before promotion.Promotion becomes slower.Conservative minimum evidence
Allowed accuracy drop0.10 percentage pointsACTIVE is already proven, so ordinary promotion gates tolerate only a very small regression.May reject useful candidates due to sample noise.Permits visibly worse candidates to proceed.Conservative safety gate
Allowed mean-error increase2%Independent regression-style guardrail, aligned with the meaningful-improvement scale.More false rejects around noise.Allows more quality loss.Conservative safety gate
Q15 requiredYesQ15 is the target execution path; a float-only Candidate is not deployable on the intended runtime.N/AIf disabled, lifecycle could promote a model that cannot run in production form.Target-derived
Q15 latency ceiling1,000 usSimple 1 ms project envelope; keeps runtime cost bounded and auditable.Rejects more capable but slower models.Allows candidates that may encroach on real-time budget.Project envelope, not hardware maximum
Max hidden layers4Matches the generic core implementation and validated checkpoint/Q15 paths.Reduces representational flexibility.Requires new core/storage/Q15 validation; not just a policy change.Implementation-derived
Max hidden width96Finite architecture envelope; wide enough for staged growth while preventing unbounded NAS.Can saturate sooner on hard problems.Higher RAM/MAC cost and larger search space.Bounded-search design
Width quantum8 neuronsCoarse deterministic steps avoid dozens of near-equivalent topologies while permitting gradual growth/shrink.Finer search but more stages/churn.Larger jumps can overshoot the adequate complexity band.Engineering heuristic
Max parameters32,768Hard, auditable architecture/resource envelope shared by planner and complexity gate.Earlier saturation.Higher memory/training/inference cost.Project resource bound
Max MACs / inference32,768Independent compute envelope aligned with the parameter bound; prevents a topology from being cheap in storage but excessive in work.Earlier rejection of compute-heavy shapes.Higher latency/energy exposure.Project resource bound
Meaningful accuracy gain+0.25 ppLarger than the 0.01 pp comparison dead-band, so tiny numerical/sample differences are not treated as a reason to buy complexity.More promotions for marginal gains.May reject useful small gains.Initial heuristic; long-run validation candidate
Meaningful error reduction2% relativeIndependent continuous-quality signal; same threshold for float and Q15 avoids a weaker quantized standard.Marginal improvements justify complexity more often.Requires stronger improvement before complexity can grow.Initial heuristic
Relative cost ceiling+100% (2x ACTIVE)Covers the bounded v0.23 escalation envelope without inventing a weighted fitness penalty. In R2 latency ratios use residency-normalized deployment latency, not shadow latency.Rejects larger but potentially useful candidates.Can accept large cost increases for modest-but-meaningful gains.Derived from bounded growth policy
Comparison dead-bands0.01 pp accuracy; 1e-6 error; 1% latencySeparate noise/tie handling from meaningful improvement thresholds; timing gets a small jitter allowance.More ties become directional comparisons.Real differences can be hidden as ties.Numerical/timing stability choice
Simplification min hidden width8Keeps a non-trivial representation and mirrors the 8-neuron architecture quantum.Can over-prune capacity.Leaves more potentially redundant capacity.Conservative simplification floor
Simplification stable dwell24,000 samplesTwo complete world cycles must remain GOOD/NOTICE before a shrink attempt.More frequent shrink attempts; higher accordion risk.Slower convergence toward a smaller adequate model.Conservative long-run heuristic
ACTIVE age holdoff48,000 samplesFour cycles after promotion: model must prove long-lived stability well beyond probation before being asked to shrink.Can prune a newly adapted model too soon.Delays resource recovery.Strong hysteresis heuristic
Simplification cooldown48,000 samplesFour cycles after a shrink attempt prevents rapid retry and grow↔shrink oscillation.Higher retry churn.Very slow retry after a transient failed shrink.Strong hysteresis heuristic
Minimum simplification saving5% parametersLifecycle risk/training/checkpoint cost should buy a visible resource reduction; tiny savings are ignored.More attempts for negligible wins.May miss safe incremental reductions.Engineering heuristic
Promotion latency sourceR2 as-if-ACTIVE deployment measurementBF518 fast memory is too small to give ACTIVE and CANDIDATE equivalent residency simultaneously. Promotion must compare the cost the Candidate would have after promotion, while shadow latency remains a separate operational diagnostic.Using shadow latency falsely penalizes the Candidate because ACTIVE already owns the fastest memory.A purely theoretical estimate would hide target-specific placement effects; the present patch measures the real hardware.Hardware-constrained workaround
Important interpretation. Numbers such as 24k/48k dwell, 5% saving, +0.25 pp and 2% error reduction are not claimed as universal optima. They are intentionally visible, auditable starting points. The long-run dataset should tell us whether they are too permissive, too conservative or appropriately separated.

11. Validation evidence: accelerated regressions plus real BF518 campaigns

CapabilityEvidenceWhy it matters
Generic deep core1–4 hidden layers; train/checkpoint/Q15 regression PASSTopology changes are runtime capability, not metadata-only.
GrowthWiden/depth-change deterministic regression PASSArchitecture escalation survives checkpoint and lifecycle transitions.
Legacy vs SIGNATURE2 policyFour deterministic reason paths exactly equivalent: relative-cost reject, bounded-tradeoff promote, active-dominates reject, Pareto promoteSignature-driven policy is a refactor of the interface, not a new decision algorithm.
R1 long endurance≈15 h 45 min; 1,176,367 ticks; 6,563 checkpoints; 411 Candidate lifecyclesMechanical stability was good enough to expose a system-level BF518 residency bias invisible in short tests.
R2 latency correctionSame-topology retrained Candidates could finally promote; probation acceptedThe hardware normalization corrected the physical comparison without relaxing policy thresholds.
R2 post-fix endurance≈6 h 12 min; 4,710,503 samples; ~392 world cycles; two promotions, two probation accepts, no rollbackThe post-patch lifecycle remained stable across millions of samples and hundreds of deterministic perturbation cycles.
Real simplification7,450 → 6,386 → 4,258 parameters promoted; 2,130-parameter target rejectedGrow/shrink is no longer only an accelerated regression path; the target executed real autonomous reductions.
SIMMEM1Rejected 256-8-10 remembered for parent 1010; no retry across >331k subsequent samplesCooldown no longer turns a deterministic failed shrink into periodic churn.
Reboot/resumeRecovered PAUSED Candidate resumes real TrainingSession; nextModelId reconciled against restored rolesInterrupted adaptation does not silently freeze or reuse model IDs.
otFILE 2.0.2A/B v6 records read with safe freadline(); later boots show both copies valid and newest selectedRecovery is deterministic rather than dependent on adjacent memory contents.
SIGNATURE1 BF518 smokeCandidate shadow=MIXED, deployment=FAST, normalized=1, authority=0Signature resource fields represent deployment, not penalized shadow placement.
SIGNATURE2 BF518 smokesignatureDecisionInput=1, dominance/gates valid, final reason REJECT_ACTIVE_DOMINATES, authority=0The real target decision engine now consumes signatures end-to-end.
R3 status. A new clean production endurance campaign has now been started from a fresh SD containing only the dataset/bootstrap model. Its results are intentionally not summarized here yet; they will be added only after the run is complete and the log is analyzed.

12. Non-blocking live status

UART0 remains deliberately minimal. ? is read-only and immediately returns to the loop; ESC performs clean checkpoint/commit/stop.

[STATUS] tick=... world=... monitor=... phase=... heap=...
[STATUS] generation(modelId)=... revision=... topology=... params=... nextModelId=...
[STATUS] roles candidate=... previous=... arch[fail=... next=... cap=...]
[STATUS] simplify[mode=... stable=... activeAge=... sinceAttempt=... attempts=... promote=... reject=... rollback=...]
[STATUS] training[state=... epoch=... sample=... trained=... checkpointSeq=...]
[STATUS] simplifyMemory[parent=... target=... outcome=...]

The query performs no SD write, fsync, allocation or lifecycle transition. The additional training and simplification-memory lines proved useful during reboot/resume and SIMMEM1 validation because they expose whether a Candidate is genuinely progressing and whether a failed shrink target is being remembered.

Operational principle. Observability is additive. The status interface reports state; it does not repair or alter it.

13. Long-run evidence: R1 → R2 → R3

The endurance campaign evolved from discovering a hardware measurement bias to validating autonomous shrink, persistence hardening and finally a signature-driven policy interface.

R1 — mechanically stable, policy input physically biased

Duration

≈15 h 45 min

1,176,367 ticks across 99 synthetic-world configurations.

Candidate lifecycles

411

410 completed rejects.

Relative-cost rejects

370

369 were same-topology 256-28-10 CLONE candidates.

Core lesson

~3.18×

Typical Candidate shadow latency inflation from Q15 residency, not topology.

R1 proved that the mechanics and persistence could run long enough to reveal a deeper problem: ACTIVE deployment latency and Candidate shadow latency were being compared under non-equivalent BF518 memory placement.

R2 — corrected deployment comparison, then long survival

R2 kept shadow execution unchanged but measured Candidate latency in an as-if-ACTIVE maintenance condition. The first same-topology retrained Candidates then promoted and survived probation, as predicted by the diagnosis rather than by any relaxed threshold.

R2 observationResultInterpretation
Post-patch endurance≈6 h 12 min; 4,710,503 samples; ~392 world cyclesMillions of samples under deterministic gain/offset/noise/target perturbations without mechanical degradation.
Accepted generationsTwo promotions, both probation-accepted; no rollbackThe promotion door became usable without becoming permissive.
Long-lived accepted modelAfter acceptance, monitor spent approximately 97.8% of observed time in WARNING and ~2.2% in NOTICE, without returning to DEGRADED/RETRAINThe model remained sufficient under pressure rather than chasing constant GOOD state.
Simplification in normal enduranceMaximum stable dwell observed ≈4,464 samples versus required 24,000The aggressive world did not naturally exercise the shrink branch; this motivated a controlled Calm Tail experiment without changing policy thresholds.

Controlled Calm Tail — exercising the branch, not tuning the policy

A temporary laboratory scenario set gain=1, offset=0, noise=0 and target drift off after the aggressive prefix. The simplification thresholds stayed unchanged. This produced real autonomous shrink transactions: 256-28-10 → 256-24-10 → 256-16-10, while 256-8-10 was rejected.

Important distinction. The test did not prove that 24,000 stable samples is universally optimal. It proved that the existing simplification path, PRUNE-COPY, promotion/probation and rejection logic work when the required stability condition is actually present.

SIMMEM1 and persistence hardening

Repeated 256-8-10 rejects exposed a new lifecycle issue: cooldown alone delayed, but did not prevent, deterministic retries of the same failed shrink. SIMMEM1 now remembers the failed parent→target pair until ACTIVE changes. Reboot testing then exposed and fixed training-resume, nextModelId reconciliation and low-level line-input defects; these changes became BOOTFIX3 + otFILE v2.0.2.

SIGNATURE2 — policy interface validated on the real target

The final R2 development stage formalized the multidimensional signature and then routed production decisions through it. A BF518 one-shot smoke measured ACTIVE latency ≈217.0 µs and a diagnostic 256-24-10 Candidate at ≈263.51 µs deployment latency; the Candidate was MIXED in shadow but FAST in deployment, normalized=1, authority zero, and the signature-driven policy reached the final reason REJECT_ACTIVE_DOMINATES.

R3 — clean production campaign completed

R3 freezes the validated mechanics into one production baseline: normal aggressive world, SIMMEM1, BOOTFIX3, otFILE v2.0.2 and SIGNATURE2. The clean run was executed from a fresh SD so its genealogy and logs were not contaminated by Calm Tail or smoke-test history.

R3 observationResultInterpretation
Endurance duration12 h 35 min 10 s; 9,568,406 samples; 797 full world cycles + 4,406 samplesLong real-BF518 execution completed with a clean stop.
Model genealogy100/r1 → 1000/r1 → 1001/r1; topology remained 256-28-10Two retraining generations were promoted with PROMOTE_PARETO and both completed probation.
Lifecycle outcomes2 promotions, 0 rejects, 0 rollbacks, 32 training checkpoints, 0 architecture failuresNo orphan roles, model-ID reuse or topology oscillation was observed.
Long-lived ACTIVEModel 1001/r1 remained ACTIVE for 9,544,856 samples after promotionThe accepted model survived almost the entire campaign without further retraining.
Monitor residenceWARNING 97.490%; NOTICE 2.453%; DEGRADED 0.019%; RETRAIN 0.017%; LEARNING 0.016%; GOOD 0.005%The aggressive world kept the model under continuous pressure without forcing repeated retraining.
Simplification170 evaluations, 0 attempts; maximum stable dwell 752 vs required 24,000The normal aggressive world never sustained the stability condition required to enter the shrink branch.
HeapSteady plateau 131,789,408 B; training low 131,467,904 B; probation 131,744,128 B; final 131,789,408 BNo monotonic decline or unrecovered generation-to-generation memory loss was observed.
A/B persistenceA: valid v6 seq 9605; B: valid v6 seq 9606; reboot selected BBoth final state checksums were valid and the newest valid slot was selected correctly.
SIGNATURE2 deployment normalizationCandidate 1000: shadow ≈916.37 µs, deployment 288.31 µs; Candidate 1001: shadow ≈913.89 µs, deployment 285.77 µsPolicy consumed deployment-normalized FAST latency rather than penalized MIXED shadow latency.
R3 endurance result: PASS. The clean campaign completed with stable heap, coherent A/B persistence, monotonic model IDs, two accepted promotions, no rollback and no observed SIGNATURE2 contradiction. Under the aggressive synthetic world, NORA remained almost continuously in WARNING; this explains why autonomous simplification did not trigger in this run without requiring any change to the validated thresholds.

14. Limits, interpretation and next physical step

From otTINN to NORA. The software engine remains otTINN; the dedicated BF518/FPGA module is NEURA; the planned embodied adaptive entity is NORA — Neural Organism for Real-world Adaptation. R3 is now mature enough that the next major scientific step is not another synthetic feature: it is transferring the same explicit lifecycle, signature and authority rules to real sensors, actuators, timing and energy constraints while keeping hard real-time control isolated from slower adaptation.
What the design does claim: model adaptation, resource reasoning, persistence and authority changes are explicit, bounded, testable and reconstructable from logs.