id: 8d8b45d89fb04e1496082f9743d2c147
parent_id: 1cc83c702d7b4370872237a4cc61324f
item_type: 1
item_id: ec1518e71453490f8fe5ab75cff7180c
item_updated_time: 1786269012435
title_diff: "[]"
body_diff: "[{\"diffs\":[[0,\" (2026-0\"],[-1,\"7-31\"],[1,\"8-09\"],[0,\")\\\n\\\n> Gen\"]],\"start1\":22,\"start2\":22,\"length1\":20,\"length2\":20},{\"diffs\":[[0,\"5 = \"],[-1,\"g50 declared. Pure NN (alpha=1.0). Gen 6 = RL self-play (future)\"],[1,\"G48 formula + RangeNet v17. Gen 6 PreflopNet/PPO abandoned. Live config updated to v17\"],[0,\".\\\n\\\n#\"]],\"start1\":43,\"start2\":43,\"length1\":72,\"length2\":94},{\"diffs\":[[0,\"angeNet \"],[1,\"v17 \"],[0,\"(pure NN\"]],\"start1\":179,\"start2\":179,\"length1\":16,\"length2\":20},{\"diffs\":[[0,\"### \"],[-1,\"Active Processes\"],[1,\"Completed\"],[0,\"\\\n- *\"]],\"start1\":239,\"start2\":239,\"length1\":24,\"length2\":17},{\"diffs\":[[0,\"equity_v\"],[-1,\"3\"],[1,\"5_2m4\"],[0,\" collect\"]],\"start1\":257,\"start2\":257,\"length1\":17,\"length2\":21},{\"diffs\":[[0,\"* (8\"],[-1,\" tables): ~1.7M/4M samples with 81-dim features. Rolling training at 1M milestones.\\\n- **range_v14_b1** trained (1M samples, 2.86% top-1, 19.0% top-10). Batch 2 triggers at 2M.\\\n- **Sanity check (pure NN)**: Gen2 ✓ +0.5, Gen3 ✗ -0.1 (marginal), G48/TAG/Nit running.\\\n\\\n### Sanity Check Status (pure NN, range_v14_b1, combo-count fix\"],[1,\"5-dim features, 19.2M samples, 2.4M hands) ✅\\\n- **range_v17 trained** (85-dim, 4.35% top-1, 24.2% top-10) ✅ **Production model**\\\n- **Sanity check passed** (all gates — 53 Harrington, self-play audit, performance vs field) ✅\\\n- **Live config updated** to range_v17 (`cash_nl_g50_live.toml`) ✅\\\n- **Parallel eval script** (`/tmp/run_eval_parallel.sh`, MAX_PARALLEL=15) ✅\\\n- **NN opponent configs** (`g48_{fish,whale,tag,nit,station,lag,maniac,reg}_nn.toml`) ✅\\\n- **PreflopNet trained** (99.38% eval accuracy, 142K model) — **DEPLOYED -644 BB/100** ❌\\\n- **PPO Phase 1** (value head corr=0.10, passes gate) — **Too weak signal** ❌\\\n- **PPO Phase 2** (off-policy v2, on-policy v3) — **No improvement over G60** ❌\\\n- **Gen 6 abandoned** — PreflopNet architecture insufficient ❌\\\n\\\n### Sanity Check (range_v17, 10 seeds × 5K hands\"],[0,\")\\\n| \"]],\"start1\":282,\"start2\":282,\"length1\":336,\"length2\":821},{\"diffs\":[[0,\"t | \"],[-1,\"Notes\"],[1,\"BB/100\"],[0,\" |\\\n|\"]],\"start1\":1115,\"start2\":1115,\"length1\":13,\"length2\":14},{\"diffs\":[[0,\"----|-------\"],[1,\"-\"],[0,\"|\\\n| Harringt\"]],\"start1\":1140,\"start2\":1140,\"length1\":24,\"length2\":25},{\"diffs\":[[0,\"gton | 1\"],[-1,\"7\"],[1,\"8\"],[0,\"/18 pass\"]],\"start1\":1163,\"start2\":1163,\"length1\":17,\"length2\":17},{\"diffs\":[[0,\"s | \"],[-1,\"Harr 4-5: raises turn with 99, book says check |\\\n| Self-play audit | ✓ passed | No guardrail overrides, no illegal raises |\\\n| Gen2\"],[1,\"— |\\\n| Self-play audit | ✓ passed | — |\\\n| G48_nn | ✓ +3.73 chips/hand | +18.7 |\\\n| TAG_nn | ✓ | +1.8 |\\\n| Nit_nn\"],[0,\" | ✓\"],[1,\" |\"],[0,\" +0.\"],[-1,\"5 BB/100 | Up from -0.1 with blend |\\\n| Gen3 | ✗ -0.1 BB/100 | Marginal, needs more data |\\\n| G48 | running | NN inference is slow (CPU bottleneck) |\\\n| TAG/Nit | pending | |\\\n\\\n### Next Steps\\\n1. **Wait for sanity check to complete** (G48 + TAG + Nit gates)\\\n2. **If passes**: Deploy to live — `cash_nl_g50_live.toml` already configured with live settings\\\n3. **If Gen3 fails**: Marginal — may need more data or parameter tuning\\\n4. **Let collection reach 4M** → train final range_v14 model\\\n5. **Gen 5 tuning**: faster NN inference, better training data, formula parameter sweep with NN \"],[1,\"1 |\\\n| Mixed (vs field) | ✓ | +0.3 |\\\n| Fish | ✓ | +4.5 |\\\n| Whale | ✓ | +4.3 |\\\n\\\n### Gen 6 Eval (PreflopNet/PPO, 10 seeds × 5K hands)\\\n| Bot | Avg chips/hand | BB/100 |\\\n|-----|---------------|--------|\\\n| G48_nn (baseline) | +3.73 | +18.7 |\\\n| G60 (PreflopNet imitation) | -128.75 | -643.8 |\\\n| G62 (PPO v2 off-policy) | -128.75 | -643.8 |\\\n| G63 (PPO v3 on-policy) | -129.45 | -647.3 |\\\n\\\nAll Gen6 bots catastrophically bad. PPO made zero improvement. G47 formula remains preflop.\\\n\\\n### Next Steps\\\n1. **Play live with G48_nn + range_v17** — already deployed\\\n2. **RangeNet improvement**: feature quality (showdown history, positional stats, pot-odds context), live opponent data collection\\\n3. **Consider NN postflop sizing** — smaller problem than preflop, better leve\"],[0,\"ra\"],[-1,\"n\"],[0,\"ge\"],[-1,\"s\\\n6\"],[1,\"\\\n4\"],[0,\". **\"]],\"start1\":1179,\"start2\":1179,\"length1\":733,\"length2\":890},{\"diffs\":[[0,\"n 6 \"],[-1,\"(future)**: RL self-play for continuous improvement\\\n\\\n## Config Rename\"],[1,\"revisit only if**: redesigned with 50+ features + continuous sizing\\\n\\\n## Sim Speed Analysi\"],[0,\"s (2\"]],\"start1\":2071,\"start2\":2071,\"length1\":77,\"length2\":97},{\"diffs\":[[0,\"26-0\"],[-1,\"7-31\"],[1,\"8-09\"],[0,\")\\\n\\\n| \"],[-1,\"Old | New | Purpose\"],[1,\"Setup | Speed | Notes\"],[0,\" |\\\n|\"]],\"start1\":2169,\"start2\":2169,\"length1\":36,\"length2\":38},{\"diffs\":[[0,\"tes |\\\n|-----\"],[-1,\"|\"],[1,\"--|--\"],[0,\"-----|------\"]],\"start1\":2200,\"start2\":2200,\"length1\":25,\"length2\":29},{\"diffs\":[[0,\"----\"],[-1,\"--\"],[0,\"|\\\n| \"],[-1,\"`cash_nl_g49.toml` | `cash_nl_g50.toml` | Gen 5 sim config |\\\n| `cash_nl_g49_live.toml` | `cash_nl_g50_live.toml` | Gen 5 live config |\\\n| `cash_nl_g49_nn.toml` | deleted | Was a duplicate |\\\n| `live_bots.toml` | cleaned up | Only g50_live + g48 fallback (removed all dead config references) |\\\n\\\n## Key Decisions (2026-07-31)\\\n\\\n### Pure NN (alpha=1.0) — DECIDED\\\n- Sweep: pure NN (+138 BB/100 seed 42) > blend (+80.5) > heuristic (-0.5)\\\n- NN is strictly better including cold start (0 hands → predicts from default stats + game context)\\\n- Alpha configurable via TOML: `range_alpha_min`, `range_alpha_max` (both 1.0)\\\n- Heuristic range builder retained for Gen 3/4, unused in G50 path\\\n\\\n### Combo-Count Fix — CRITICAL\\\n- `probs_to_hand_range()` in `range_predictor.rs` was assigning full type probability to each combo\\\n- Inflated offsuit combos 3× vs suited (e.g., AKo has 12 combos, AKs has 4 — old code gave same per-combo weight)\\\n- Fixed: 2-pass — count remaining combos per type, then assign `type_prob / count` per combo\\\n- Impact: turned blend from -16.4 → +15.1 BB/100\\\n\\\n## Cleanup Performed (2026-07-31)\\\n\\\n### Deleted Data (~380G freed)\\\n- all_combined_v8/v9/v10/v11.jsonl (213G)\\\n- teacher_postflop.jsonl (34G)\\\n- range_consistent_postflop.jsonl (12G)\\\n- single_v1/, selfplay_v1/, selfplay_v2/, combined_v2/, combined_v3/ (104G)\\\n- equity_v1_old_actions/, equity_v2/ (17G)\\\n\\\n### Deleted Models (all invalid — old action space or old range dims)\\\n- All gen5_postflop_v*.safetensors (old action space)\\\n- All gen5_preflop_v*.safetensors (old action space)\\\n- gen5_postflop_eq1-6.safetensors (failed equity experiment)\\\n- equity_postflop_v1.safetensors (redundant with G48 MC)\\\n- fold_eq_postflop_v1.safetensors (overfit, insufficient data)\\\n- \"],[1,\"9 Gen2 (formula) | 133 h/s | Baseline |\\\n| 9 Tag_nn | 1300 h/s | Tight opponents = few multiway pots |\\\n| 9 Fish_nn | 4.8 h/s | 27× slower — loose = 5-7 way flops |\\\n| G63 + 8 mixed G48_nn | 9.9 h/s | Mixed = many multiway postflop situations |\\\n\\\n**Bottleneck**: Multiway postflop computation (G48 formula per active player per street), NOT NN inference.\\\n- RangeNet forward pass: ~0.1ms (fast)\\\n- MC budget (1K vs 10K): zero effect on speed\\\n- Loose opponents create exponentially more postflop decisions\\\n\\\n## Parallel Eval Script\\\n\\\n`/tmp/run_eval_parallel.sh <bot_type> <name_tag> <seeds...>`\\\n- MAX_PARALLEL=15 (32 cores available)\\\n- Generates per-seed TOML configs dynamically\\\n- 10 seeds × 5K hands in ~11 min (was 83 min sequential)\\\n\\\n## Range Model Files\\\n- `models/\"],[0,\"range_v1\"],[-1,\"0-v13\"],[1,\"7\"],[0,\".saf\"]],\"start1\":2226,\"start2\":2226,\"length1\":1752,\"length2\":781},{\"diffs\":[[0,\"sors\"],[-1,\" (57-dim, superseded by 81-dim)\\\n\\\n### Deleted Configs\\\n- gen5_play.toml, gen5_v7_eval.toml, gen5_combined_eval.toml (dropped approaches)\\\n- gen5_selfplay.toml, gen5_collect_g48.toml (outdated)\\\n- cash_nl_g49_nn.toml (duplicate of g49/g50)\\\n- All old Gen 1-4 configs (g4, g45, g47_live, normal, loose, tight, etc.)\\\n\\\n### Deleted Scripts\\\n- All old pipeline scripts (loop_v11, multiloop, neutral, retrain, single_pipeline)\\\n- train_equity_postflop.sh, train_range_consistent_postflop.sh, train_teacher_postflop.sh\\\n- gen5_selfplay_collect.sh, gen5_collect_g48.sh, train_after_selfplay_v2.sh\"],[1,\"` — Production (85-dim, 19.2M samples, 4.35% top-1, 24.2% top-10)\\\n- `models/preflop_v1.safetensors` — PreflopNet 7-class (142K, abandoned)\\\n- `models/preflop_ppo_v1/v2/v3.safetensors` — PPO models (178K each, abandoned)\"],[0,\"\\\n\\\n##\"]],\"start1\":3011,\"start2\":3011,\"length1\":587,\"length2\":226},{\"diffs\":[[0,\"cture (8\"],[-1,\"1\"],[1,\"5\"],[0,\" dims)\\\n\\\n\"]],\"start1\":3259,\"start2\":3259,\"length1\":17,\"length2\":17},{\"diffs\":[[0,\"\\\n\\\n**\"],[-1,\"49\"],[1,\"53\"],[0,\" bas\"]],\"start1\":3274,\"start2\":3274,\"length1\":10,\"length2\":10},{\"diffs\":[[0,\"5): \"],[-1,\"ln(\"],[0,\"pot_bb\"],[-1,\")\"],[0,\", \"],[-1,\"ln(\"],[0,\"stack_bb\"],[-1,\")\"],[0,\", nu\"]],\"start1\":3541,\"start2\":3541,\"length1\":32,\"length2\":24},{\"diffs\":[[0,\"_bet\"],[-1,\" (did player face aggression before acting?)\\\n- first_action, last_action (action codes)\\\n- aggression_ratio (raises / (raises + calls))\"],[1,\", first_action, last_action, aggression_ratio\"],[0,\"\\\n\\\nSh\"]],\"start1\":4014,\"start2\":4014,\"length1\":142,\"length2\":53}]"
metadata_diff: {"new":{},"deleted":[]}
encryption_cipher_text: 
encryption_applied: 0
updated_time: 2026-08-09T09:53:33.180Z
created_time: 2026-08-09T09:53:33.180Z
is_locked: 0
type_: 13