id: 0aa966bc036d4db3965dc5e54037e2b9
parent_id: 818f3e59dd5c4b98baf2bdd3d885a460
item_type: 1
item_id: ec1518e71453490f8fe5ab75cff7180c
item_updated_time: 1786952699248
title_diff: "[]"
body_diff: "[{\"diffs\":[[0,\"-17 \"],[-1,\"morning \"],[0,\"— Eq\"]],\"start1\":31,\"start2\":31,\"length1\":16,\"length2\":8},{\"diffs\":[[0,\"uityNet \"],[-1,\"squash\"],[1,\"parked after\"],[0,\" root-ca\"]],\"start1\":39,\"start2\":39,\"length1\":22,\"length2\":28},{\"diffs\":[[0,\"use \"],[-1,\"hunt\"],[1,\"depth\"],[0,\")\\\n\\\n>\"]],\"start1\":67,\"start2\":67,\"length1\":12,\"length2\":13},{\"diffs\":[[0,\"et: \"],[-1,\"label-semantics bug found+fixed (huge win), multiway gate PASSED, one optimization pathology left.\\\n\\\n## Night summary (v7→v14)\\\n\\\n**SOLVED — the big one (commit 678080e)**: label semantics. Gen3 hs = CURRENT strength (beat-all now), labels were showdown equity. Relabeled 4.86M/7.39M records offline via stored opp_holes (`relabel_equity`). Multiway shadow MAE 0.096→\"],[1,\"investigated to the bottom, multiway gate passed once, calibration unsolved after 15 runs — **PARKED** per timebox. Post-mortem below.\\\n\\\n## EquityNet final status: PARKED\\\n\\\n**What works (proven):**\\\n- Architecture, data pipeline, labels: multiway hs MAE hit \"],[0,\"**0.04\"],[-1,\"5\"],[1,\"48\"],[0,\" (gate \"],[1,\"0.045 \"],[0,\"PASS)**\"],[-1,\".\\\n\\\n**SOLVED**: v8 sigmoid-collapse (ternary 87%-zero labels + L1 → saturation) → identity head + MSE. candle AdamW default weight_decay=0.01-on-biases found and zeroed. Model-size load inference from safetensors header.\\\n\\\n**OPEN — output squash\"],[1,\" on v10 (label-fit selection); potentials excellent (ppot 0.033, nutpot 0.028-0.04)\\\n- The label-semantics fix (relabel to current-strength, commit 678080e) — the single biggest win, MW MAE 0.096→0.045\\\n- Hand ORDERING in every model: street/n/suit gradients all correct\\\n\\\n**What doesn't (unsolved):**\\\n- **Output-scale calibration\"],[0,\"**: \"]],\"start1\":119,\"start2\":119,\"length1\":636,\"length2\":618},{\"diffs\":[[0,\"els \"],[-1,\"predic\"],[1,\"lodge a\"],[0,\"t \"],[-1,\"~\"],[0,\"0.3\"],[1,\"-0.5\"],[0,\"x th\"]],\"start1\":744,\"start2\":744,\"length1\":20,\"length2\":24},{\"diffs\":[[0,\"an (\"],[-1,\"diag: mean pred 0.08-0.15 vs label 0.32; even own TRAINING data). Narrowed by elimination:\\\n- NOT architecture/loss/optimizer: synthetic constant-0.5 → 0.482 ✓; 512 real HU samples memorize at lr 3e-3 in 400 steps ✓\\\n- NOT weight decay, NOT bias init (v12/v13 still squash\"],[1,\"slope 0.4-0.7). Diagnosed against: sigmoid saturation (fixed), weight decay on biases (fix\"],[0,\"ed),\"],[-1,\" NOT\"],[0,\" cap\"]],\"start1\":784,\"start2\":784,\"length1\":286,\"length2\":102},{\"diffs\":[[0,\"city\"],[-1,\" (v10 big same), NOT pos-class weighting (v11 worse)\\\n- **LR is the lever**: v14 (lr 2e-3 vs 5e-4) doubled the output level (diag 0.146, val slope 0.575→0.677) and was STILL\"],[1,\", class weighting, LR (5e-4 frozen / 2e-3 chaotic-50% / 3e-3 diverges), bias init, warm-restart, grad\"],[0,\" cli\"],[-1,\"mb\"],[1,\"pp\"],[0,\"ing \"],[-1,\"when the cosine expired (val improved every batch of epoch 6)\\\n- Next: v15 = lr 3e-3, 12 epochs, possibly warm restart if it plateaus. The sanity test succeeded at exactly 3e-3.\\\n\\\n## Current model ranking (shadow, fish field)\\\n| model | MW hs | nutpot | slope | note |\\\n|---|---|---|---|---|\\\n| v10 | **0.0448 PASS** | 0.041 | 0.575 | big, 20ep, lr 5e-4 |\\\n| v14 | 0.0468 | 0.072 | 0.467 | lr 2e-3, 6ep — les\"],[1,\"(prevents divergence but oscillates), checkpoint selection (v16's calibrated epochs 13-16 slope 0.96-1.06 were DISCARDED by MAE-based selection — fixed, but the recipe doesn't reproduce: v17/v19/v20 all failed to stabilize)\\\n- **Linear rescale**: fitted on corpus labels — fails on deployment distribution (label mean 0.32 vs field MC mean 0.057). Fitted on deployment (shadow log): MW MAE 0.061-0.068 — near-gate, but constant\"],[0,\"s \"],[-1,\"f\"],[0,\"ar\"],[-1,\" along |\\\n| v9 | 0.0463 | \"],[1,\"e distribution-specific (fish vs TAG differ). HU unusable regardless (\"],[0,\"0.\"],[-1,\"0\"],[0,\"28\"],[-1,\" | 0.634 | base, 12ep |\\\n\\\n## Morning options\\\n1. **v15 run** (lr 3e-3, 12ep, ~2h) — the evidence says this finishes the escape; then shadow → if slope ≥0.9 → A/B + sanity + adoption decision\\\n2. If v15 still short: warm-restart v15 from v14 checkpoint at constant 1e-3 (trainer lacks resume — 20-line addition)\\\n3. Fallback deployment posture: v10 as MULTIWAY-ONLY replacement (its MW g\"],[1,\").\\\n\\\n**Root assessment**: heavily-imbalanced (~87% zero) ternary regression is a genuinely hard optimization for this MLP+AdamW setup. The 512-sample memorization vs 7M-sample squash gap suggests the fresh-data regime starves the rare positive classes' gradient signal — a curriculum/2-stage or BCE-with-logits redesign might fix it, but that's a NEW research iteration, not a bug fix.\\\n\\\n**Assets kept:**\\\n- Corpus (7.25M relabeled, relabel_equity bin), recorder, model, trainer (warm-restart, clipping, slope-gated checkpoints), diag bin (mean-pred-by-n), shadow tooling — all committed\\\n- equity_v14 = best stable model (slope 0.677); v16 recipe = the one th\"],[0,\"at\"],[-1,\"e\"],[0,\" p\"],[-1,\"asses; HU/slope issues matter less on multiway where MC is the bottleneck) — gate NN replacement on n_opp>2, keep MC for HU. Could ship the sim-speedup before full calibra\"],[1,\"roduced calibrated epochs (unreproduced)\\\n- Shadow mode runs zero-risk on any sim (EQUITY_SHADOW_LOG)\\\n\\\n**Reactivation triggers**: (a) new training approach (BCE/curriculum), (b) MC replacement urgency for sim throughput, (c) corpus regrown with balanced labels (e.g. softer targets = true MC-vs-range values stored at collec\"],[0,\"tion\"],[1,\")\"],[0,\".\\\n\\\n## \"],[-1,\"Infra\\\n- eval_equity_diag bin (BIN=... model — mean pred vs label by n_opp) — the metric that localizes squash\\\n- All committed: 678080e, 1deaf4b, 1a1bcd1, 548ee48. Corpus train_full.bin (7.25M, relabeled).\\\n\\\n## Backlog unchanged: 7_6 · FutureActionNet · live fine-tune · call-path eq_cc · v18 · fmt debt\"],[1,\"Session results (live)\\\n- Day 1: +30.6M/36 hands (Pound It II 250K) ≈ +122 BB — lucky\\\n- Day 2: losses in smaller pots; today +26 BB/104 hands. **Sample sizes are noise at these levels** — improvement comes from the pipeline, not session results.\\\n- Big-pot spot checks: AK vs turned full house (pot-odds correct), AA vs rivered 777-full 3-way (cooler). No defects found in tracked hands.\\\n\\\n## Backlog (next up, in order)\\\n1. **7_6** sequence-aware capping (offline likelihood fit from corpus — infra now exists)\\\n2. **FutureActionNet** (#4) — drops into equity_vs_continue; corpus infra reusable\\\n3. **Live RangeNet fine-tune** — Torn profiles accumulating since Aug 15\\\n4. Call-path eq_cc (knob-gated)\\\n5. RangeNet v18 Tier-2\\\n6. Repo-wide fmt debt\\\n\\\n## Key notes\\\nIdeas `f3c42f44` | Roadmap `47031623` | Sanity `d632cea3` | Archive `283e98fe`\"]],\"start1\":887,\"start2\":887,\"length1\":1492,\"length2\":2446}]"
metadata_diff: {"new":{},"deleted":[]}
encryption_cipher_text: 
encryption_applied: 0
updated_time: 2026-08-17T07:45:44.955Z
created_time: 2026-08-17T07:45:44.955Z
is_locked: 0
type_: 13