id: b76fcf8c4cbc433e8f09277e3971b4a9
parent_id: 1dd41f6ce3264ce0a44c459b76ca1762
item_type: 1
item_id: 16ee20ba74264eb2a2bb0a20a9cea103
item_updated_time: 1785407084671
title_diff: "[]"
body_diff: "[{\"diffs\":[[0,\"\\\n## \"],[-1,\"Iteration History\\\n\\\n| Model | Postflop | Preflop | Range | Date | vs Gen2 | vs Gen3 | vs G48 | vs Mixed | Harrington | Notes |\\\n|-------|----------|---------|-------|------|---------|---------|--------|----------|------------|-------|\\\n| CF v6 (buggy) | — | — | — | 07-22 | — | **-321** | — | — | — | Buggy counterfactual EVs |\\\n| CF v7 (fixed) | — | — | — | 07-22 | — | **-209** | — | — | — | No aggressive masking |\\\n| CF v7 + mask | — | — | — | 07-22 | — | **-20.5** | — | — | — | Aggressive masking added |\\\n| Q-reg v1 | — | — | — | 07-22 | — | **-47.2** | — | — | — | G48 teacher data only |\\\n| **Q-reg v2 (prod v6/v10)** | v6 | v10 | none | 07-22 | **+23.7** | **+12.1** | — | — | 48/53 | Warm-start CF v7 + combined data. OLD test setup (5-bot). \"],[1,\"⚠️ CRITICAL: All results before 2026-07-30 12:25 are INVALID\\\n\\\nThe `holdem_bots` binary was compiled WITHOUT `gen5_nn`/`gen5_cuda`. The `ModelWeights` \\\"stub\\\" is literally an empty placeholder (`_placeholder: ()`). Without `gen5_nn`, `model_infer()` always returns `None`, so ALL decisions fell back to the G48 teacher formula. The NN never fired once.\\\n\\\n**Symptom**: Identical self-play results for different model configs (both fall back to same teacher).\\\n\\\n**Root cause**: `cargo test --release -p hand_replay` (Gate 1) recompiles holdem_bots as a dependency without `gen5_cuda`, corrupting shared build artifacts.\\\n\\\n**Fix**: Sanity check script now verifies binary has \\\"Using CUDA\\\" string before running, rebuilds if missing.\\\n\\\n## Iteration History (INVALID — all G48 teacher, not NN)\\\n\\\n| Model | Date | vs Gen2 | vs Gen3 | vs G48 | vs TAG | Notes |\\\n|-------|------|---------|---------|--------|--------|-------\"],[0,\"|\\\n| \"]],\"start1\":119,\"start2\":119,\"length1\":754,\"length2\":916},{\"diffs\":[[0,\"e-fix) |\"],[-1,\" v8 | v12 | v11 |\"],[0,\" 07-29 |\"]],\"start1\":1045,\"start2\":1045,\"length1\":33,\"length2\":16},{\"diffs\":[[0,\"9 | \"],[-1,\"**\"],[0,\"-14.4\"],[-1,\"**\"],[0,\" | \"],[-1,\"**\"],[0,\"-12.6\"],[-1,\"**\"],[0,\" | +\"]],\"start1\":1058,\"start2\":1058,\"length1\":29,\"length2\":21},{\"diffs\":[[0,\"0 | \"],[-1,\"0* | 52/53 | From scratch, self-play only, min_raise bug\"],[1,\"— | G48 teacher, not NN\"],[0,\" |\\\n|\"]],\"start1\":1081,\"start2\":1081,\"length1\":64,\"length2\":31},{\"diffs\":[[0,\"x) |\"],[-1,\" v8 | v12 | v11 |\"],[0,\" 07-\"]],\"start1\":1128,\"start2\":1128,\"length1\":25,\"length2\":8},{\"diffs\":[[0,\"0 | \"],[-1,\"**\"],[0,\"-0.1\"],[-1,\"**\"],[0,\" | \"],[-1,\"**+\"],[0,\"0.0\"],[-1,\"**\"],[0,\" | \"],[-1,\"pending | pending | 52/53 | min_raise fix → break-even. Train-inference mismatch. |\\\n| prod\"],[1,\"+37.2 | +37.3 | G48 teacher, not NN |\\\n|\"],[0,\" v6/\"]],\"start1\":1137,\"start2\":1137,\"length1\":120,\"length2\":60},{\"diffs\":[[0,\"x) |\"],[-1,\" v6 | v10 | none |\"],[0,\" 07-\"]],\"start1\":1209,\"start2\":1209,\"length1\":26,\"length2\":8},{\"diffs\":[[0,\"0 | \"],[-1,\"**\"],[0,\"-0.1\"],[-1,\"**\"],[0,\" | \"],[-1,\"**\"],[0,\"0.0\"],[-1,\"**\"],[0,\" | \"],[-1,\"pending | pending | 52/53 | Also break-even! Bug fix exposed train-inference mismatch. |\\\n| v9/v13 (warm-start, bug data) | v9 | v13 | v11 | 07-30 | pending | pending | pending | pending | pending | Warm-start from v6/v10 on combined_v2. Still bug data. |\\\n\\\n## Critical Insight: Train-Inference Mismatch (2026-07-30)\\\n\\\nThe min_raise fix eliminated 556 forced-call conversions/1000h, but ALL existing models were trained on bug-affected data. The models learned policies assuming their raises would be silently converted to calls. Now that raises actually execute, the models play lines they never learned. This is why both v6/v10 (production) and v8/v12 dropped to break-even.\\\n\\\n**Solution**: Re-collect self-play data WITHOUT the bug (selfplay_v2, in progress). Then retrain all models on clean\"],[1,\"+37.3 | +37.1 | G48 teacher, not NN |\\\n\\\n## Pre-bug History (from 07-22, OLD test setup — 5-bot tables)\\\n\\\n| Model | vs Gen2 | vs Gen3 | Notes |\\\n|-------|---------|---------|-------|\\\n| CF v6 (buggy) | — | -321 | Buggy EVs |\\\n| CF v7 + mask | — | -20.5 | First working aggressive masking |\\\n| Q-reg v1 | — | -47.2 | G48 teacher data only |\\\n| Q-reg v2 (=v6 prod) | +23.7 | +12.1 | Warm-start + combined\"],[0,\" data\"],[-1,\".\"],[1,\" |\"],[0,\"\\\n\\\n## C\"],[-1,\"orrelation vs Play Quality\"],[1,\"urrent Models\"],[0,\"\\\n\\\n| \"]],\"start1\":1218,\"start2\":1218,\"length1\":858,\"length2\":441},{\"diffs\":[[0,\"r | \"],[-1,\"vs Gen3 BB/100\"],[1,\"Status\"],[0,\" |\\\n|\"]],\"start1\":1670,\"start2\":1670,\"length1\":22,\"length2\":14},{\"diffs\":[[0,\"----\"],[-1,\"--------|\\\n| Q-reg v1\"],[1,\"|\\\n| postflop v6\"],[0,\" | 0.\"],[-1,\"605 | -47.2 |\\\n| Q-reg v2 (=v6) | 0.494 | +12.1 (old setup) / 0.0 (new setup, post-fix\"],[1,\"533 | Production (best in play historically\"],[0,\") |\\\n\"]],\"start1\":1703,\"start2\":1703,\"length1\":118,\"length2\":71},{\"diffs\":[[0,\"1 | \"],[-1,\"+0.0 (post-fix)\"],[1,\"From scratch, Q-reg\"],[0,\" |\\\n|\"]],\"start1\":1794,\"start2\":1794,\"length1\":23,\"length2\":27},{\"diffs\":[[0,\"0.594 | \"],[-1,\"pending\"],[1,\"Warm-start from v6, combined_v2 data\"],[0,\" |\\\n| pre\"]],\"start1\":1836,\"start2\":1836,\"length1\":23,\"length2\":52},{\"diffs\":[[0,\"0.904 | \"],[-1,\"—\"],[1,\"Production\"],[0,\" |\\\n| pre\"]],\"start1\":1899,\"start2\":1899,\"length1\":17,\"length2\":26},{\"diffs\":[[0,\"0.438 | \"],[-1,\"—\"],[1,\"From scratch, Q-reg\"],[0,\" |\\\n| pre\"]],\"start1\":1936,\"start2\":1936,\"length1\":17,\"length2\":35},{\"diffs\":[[0,\"5 | \"],[-1,\"pending |\\\n\\\n## Recipe for Clean Iteration\\\n\\\n1. \"],[1,\"Warm-start from v10, combined_v2 data |\\\n| range_v10 | 3.4% top-1 / 20.9% top-10 | Best range model |\\\n\\\n## Next Steps\\\n\\\n1. **Re-run sanity checks** with correct binary (gen5_nn + CUDA)\\\n2. **\"],[0,\"Wait\"]],\"start1\":1986,\"start2\":1986,\"length1\":53,\"length2\":195},{\"diffs\":[[0,\"fplay_v2\"],[1,\"** clean\"],[0,\" collect\"]],\"start1\":2189,\"start2\":2189,\"length1\":16,\"length2\":24},{\"diffs\":[[0,\"on (\"],[-1,\"clean, no min_raise bug)\\\n2. Combine: si\"],[1,\"now running with real NN)\\\n3. **Train ra\"],[0,\"ng\"],[-1,\"l\"],[0,\"e_v1\"],[-1,\" teacher (2.2M, clean) + selfplay_v2 (~2.3M, clean) = ~4.5M clean\\\n3. Warm-start from v9/v13 (best available initialization)\\\n4. Q-regression\"],[1,\"2** from clean range data\\\n4. **Train postflop v10 + preflop v14**: warm-start from v9/v13, clean combined data\"],[0,\", --\"]],\"start1\":2214,\"start2\":2214,\"length1\":193,\"length2\":163},{\"diffs\":[[0,\"e_v1\"],[-1,\"1\"],[1,\"2\"],[0,\"\\\n5. \"],[-1,\"Run s\"],[1,\"**S\"],[0,\"anit\"]],\"start1\":2393,\"start2\":2393,\"length1\":18,\"length2\":16},{\"diffs\":[[0,\"heck\"],[-1,\", r\"],[1,\" new models**\\\n6. **R\"],[0,\"ecord \"],[1,\"REAL \"],[0,\"resu\"]],\"start1\":2412,\"start2\":2412,\"length1\":17,\"length2\":39},{\"diffs\":[[0,\"lts here\"],[1,\"**\"],[0,\"\\\n\\\n## Key\"]],\"start1\":2451,\"start2\":2451,\"length1\":16,\"length2\":18},{\"diffs\":[[0,\"sons\\\n\\\n1.\"],[1,\" **Binary verification is critical**: Always check `strings holdem_bots | grep \\\"Using CUDA\\\"` before running\\\n2.\"],[0,\" **Warm-\"]],\"start1\":2473,\"start2\":2473,\"length1\":16,\"length2\":126},{\"diffs\":[[0,\"ible\"],[-1,\". The breakthrough came from warm-starting.\\\n2\"],[1,\"\\\n3\"],[0,\". **\"]],\"start1\":2653,\"start2\":2653,\"length1\":53,\"length2\":10},{\"diffs\":[[0,\"lone\"],[-1,\".\\\n3\"],[1,\"\\\n4\"],[0,\". **\"],[-1,\"Aggressive masking at deep stacks is necessary**: Without it, the model over-values big bets.\\\n4\"],[1,\"min_raise bug fixed**: 556 forced-call conversions → 0. Preflop mask now respects dynamic min_raise.\\\n5\"],[0,\". **\"]],\"start1\":2741,\"start2\":2741,\"length1\":110,\"length2\":116},{\"diffs\":[[0,\"orr \"],[-1,\"does not mean better play.\\\n5. **min_raise bug caused 556 forced-call conversions/1000 hands**: Preflop mask ignored engine's dynamic min_raise.\\\n6. **Train-inference mismatch**: Fixing the bug exposed that all models trained on bug data. Must re-collect and retrain.\"],[1,\"≠ better play\\\n6. **Train-inference mismatch**: All bug-era models trained on data where raises were silently converted to calls\"]],\"start1\":2901,\"start2\":2901,\"length1\":269,\"length2\":131}]"
metadata_diff: {"new":{},"deleted":[]}
encryption_cipher_text: 
encryption_applied: 0
updated_time: 2026-07-30T10:29:27.478Z
created_time: 2026-07-30T10:29:27.478Z
is_locked: 0
type_: 13