id: 760c8acac2614551b60e383461f794f8
parent_id: 39213e0f0bbe4d48805e9b9cbdaddb66
item_type: 1
item_id: ec1518e71453490f8fe5ab75cff7180c
item_updated_time: 1785522276226
title_diff: "[]"
body_diff: "[{\"diffs\":[[0,\"07-31)\\\n\\\n\"],[1,\"> Gen 5 = G49 declared. Pure NN (alpha=1.0). Gen 6 = RL self-play (future).\\\n\\\n\"],[0,\"## Curre\"]],\"start1\":29,\"start2\":29,\"length1\":16,\"length2\":93},{\"diffs\":[[0,\"RangeNet\"],[1,\" (pure NN, alpha=1.0)\"],[0,\"**\\\n\\\n### \"]],\"start1\":156,\"start2\":156,\"length1\":16,\"length2\":37},{\"diffs\":[[0,\"s): \"],[-1,\"Collect\"],[1,\"~1.7M/4M samples with 81-dim features. Roll\"],[0,\"ing \"],[1,\"t\"],[0,\"ra\"],[-1,\"nge data with expanded 81-dim features. ETA ~6h\"],[1,\"ining at 1M milestones.\\\n- **range_v14_b1** trained (1M samples, 2.86% top-1, 19.0% top-10). Batch 2 triggers at 2M\"],[0,\".\\\n- **\"],[-1,\"G49 s\"],[1,\"S\"],[0,\"anit\"]],\"start1\":245,\"start2\":245,\"length1\":79,\"length2\":179},{\"diffs\":[[0,\"ty check\"],[1,\" (pure NN)\"],[0,\"**: Gen2\"]],\"start1\":423,\"start2\":423,\"length1\":16,\"length2\":26},{\"diffs\":[[0,\"en2 \"],[-1,\"-0.1\"],[1,\"✓ +0.5\"],[0,\", Gen3 \"],[-1,\"+0.1, G48 +1.2 BB/100. TAG/Nit gates running.\"],[1,\"✗ -0.1 (marginal), G48/TAG/Nit running.\\\n\\\n### Sanity Check Status (pure NN, range_v14_b1, combo-count fix)\\\n| Gate | Result | Notes |\\\n|------|--------|-------|\\\n| Harrington | 17/18 pass | Harr 4-5: raises turn with 99, book says check |\\\n| Self-play audit | ✓ passed | No guardrail overrides, no illegal raises |\\\n| Gen2 | ✓ +0.5 BB/100 | Up from -0.1 with blend |\\\n| Gen3 | ✗ -0.1 BB/100 | Marginal, needs more data |\\\n| G48 | running | NN inference is slow (CPU bottleneck) |\\\n| TAG/Nit | pending | |\"],[0,\"\\\n\\\n##\"]],\"start1\":446,\"start2\":446,\"length1\":64,\"length2\":516},{\"diffs\":[[0,\"teps\\\n1. \"],[1,\"**\"],[0,\"Wait for\"]],\"start1\":970,\"start2\":970,\"length1\":16,\"length2\":18},{\"diffs\":[[0,\"for \"],[-1,\"equity_v3 collection to complete\\\n2. Merge range samples: `cat equity_v3/table_*/range_data.jsonl > range_merged.jsonl`\\\n3. Train range_v14: `train_gen5_range range_merged.jsonl 30 512 1e-4 models/range_v14.safetensors`\\\n4. Update G49 config: add `range_model_path = \\\"models/range_v14.safetensors\\\"`\\\n5. Alpha sweep: test alpha_max values (0.5, 0.6, 0.7, 0.8)\\\n6. Full sanity check before live promotion\"],[1,\"sanity check to complete** (G48 + TAG + Nit gates)\\\n2. **If passes**: Deploy to live — update `cash_nl_g49_live.toml` with range model path\\\n3. **If Gen3 fails**: Marginal — may need more data or parameter tuning\\\n4. **Let collection reach 4M** → train final range_v14 model\\\n5. **Gen 5 tuning**: faster NN inference, better training data, formula parameter sweep with NN ranges\\\n6. **Gen 6 (future)**: RL self-play for continuous improvement\\\n\\\n## Key Decisions (2026-07-31)\\\n\\\n### Pure NN (alpha=1.0) — DECIDED\\\n- Sweep: pure NN (+138 BB/100 seed 42) > blend (+80.5) > heuristic (-0.5)\\\n- NN is strictly better including cold start (0 hands → predicts from default stats + game context)\\\n- Alpha configurable via TOML: `range_alpha_min`, `range_alpha_max` (both 1.0)\\\n- Heuristic range builder retained for Gen 3/4, unused in G49 path\\\n\\\n### Combo-Count Fix — CRITICAL\\\n- `probs_to_hand_range()` in `range_predictor.rs` was assigning full type probability to each combo\\\n- Inflated offsuit combos 3× vs suited (e.g., AKo has 12 combos, AKs has 4 — old code gave same per-combo weight)\\\n- Fixed: 2-pass — count remaining combos per type, then assign `type_prob / count` per combo\\\n- Impact: turned G49 blend from -16.4 → +15.1 BB/100\"],[0,\"\\\n\\\n##\"]],\"start1\":985,\"start2\":985,\"length1\":405,\"length2\":1223},{\"diffs\":[[0,\"sh\\\n\\\n\"],[-1,\"### Updated Configs\\\n- cash_nl_g49.toml: Removed stale range_v12 reference\\\n- cash_nl_g49_live.toml: Fixed nonexistent range_v8 reference\\\n- gen5_equity_collect.toml: Removed stale action model references, simplified to formula-only\\\n- eval_9max.sh, eval_mixed_field.sh: Updated default bot refs\\\n- sanity_check.sh: Updated example ref\\\n\\\n\"],[0,\"## R\"]],\"start1\":3348,\"start2\":3348,\"length1\":340,\"length2\":8},{\"diffs\":[[0,\"recorder.rs`\"],[1,\" — used by both training and inference.\"]],\"start1\":4338,\"start2\":4338,\"length1\":12,\"length2\":51}]"
metadata_diff: {"new":{},"deleted":[]}
encryption_cipher_text: 
encryption_applied: 0
updated_time: 2026-07-31T18:29:42.387Z
created_time: 2026-07-31T18:29:42.387Z
is_locked: 0
type_: 13