id: 6014d2d75c21482aa3617929af79f3a0
parent_id: a26e3cbaa2f44bde90c0ce900b6efd2b
item_type: 1
item_id: ec1518e71453490f8fe5ab75cff7180c
item_updated_time: 1787102268832
title_diff: "[]"
body_diff: "[{\"diffs\":[[0,\"08-1\"],[-1,\"7 night — real-logic oracle added)\\\n\\\n> **Live: G51.** Overnight chain: fitting collection → curve MLE fit → oracle-scoring collection (all autonomous).\\\n\\\n## Tonight's additions\\\n- **Real-logic oracle data** (c47d8a4): StrategyBot::act records each bot's OWN chosen action + personality name per (hand, hero, street). With realized holdings revealed, the corpus now supports the EMPIRICAL oracle P(action | holding, config) — real per-config bot logic, replacing the hand-fit proxy curves in the overlap score.\\\n- **User's overlap score, implemented** (fit_likelihood --score): binarized log-score, weight added when the holding would take the observed line, subtract\"],[1,\"9 morning — g53 temperature verdict: PARK)\\\n\\\n> **Live: G51 (unchanged, healthy).** Range-temperature candidate g53 fully evaluated overnight: **parked** — the offline calibration gain does not survive the composite production pipeline.\\\n\\\n## g53 evaluation (complete)\\\n\\\n| Gate | Result |\\\n|---|---|\\\n| Harrington | 52/52+1ign pass |\\\n| KK replay | REGRESSED — river hs 0.29→0.46, overplay Raise returns (temp partially undoes corrector) |\\\n| Calibration audit (917K, temp VERIFIED flowing: Gen5T HU mc 0.243 vs baseline 0.314) | **Mean bias NOT fixed**: overall 0.059 vs realized 0.152 (unchanged); HU slope 0.79→0.83 only; **nutpot destroyed 0.018→0.000** |\\\n| A/B vs g51 (200K hands) | −0.21 BB/100 dead heat; fish field +2.29 vs +2.19 control — play-neutral |\\\n\\\n**Why the offline gain didn't transfer**: the corrector runs AFTER the temperature and re-concentrates on call lines (its job) — netting out the softening exactly where it matters. The measured t≈0.55 was fitted on prior-only E[beat]; production hs passes through temp→corrector→MC. Composite systems resist single-knob fixes.\\\n\\\n## Incidents found & fixed this session (all committed)\\\n1. **Comma-spliced TOML** in collector configs (`key = 1.0, key2 = true` on one line = invalid TOML) — silently skipped (warning buried), stale same-named config satisfi\"],[0,\"ed \"],[1,\"b\"],[0,\"ot\"],[-1,\"herwise. v17 prior scores −0.60 (mass on hands tha\"],[1,\"_type → **the first \\\"g53 audit\\\" was actually a g51 re-run**. Fixed + unique collector names + glob-exclusion.\\\n2. **Registration failures now FATAL** in holdem_bots main (d2f9287) — immediately caugh\"],[0,\"t \"],[1,\"t\"],[0,\"wo\"],[-1,\"uldn't play these lines); softening monotonically improves it (t=1 → −0.15). Direction validates: the prior is miscalibrated exactly as the corrected sweep showed.\\\n- Chain: equity_v7_fit collection → fit (corrected binary) → **\"],[1,\" more broken configs (sim.toml parsed as bot config → skip; v6/_nc collector splices → repaired).\\\n3. Relabel nohup killed by shell timeout → restarted via background_process.\\\n\\\n## Where the range bias stands (cumulative knowledge)\\\n- Bias = RangeNet prior overconcentration (measured 3 ways: audit, temperature bracketing, overlap score −0.60)\\\n- Corrector exonerated (on/off identical); NOT fixable post-hoc by temperature (g53 proved this)\\\n- **The fix is v18 retrain**: calibration-aware training on the new corpora (multi-hand-per-context distributions from \"],[0,\"equity_\"],[1,\"v7_fit/\"],[0,\"v8_o\"]],\"start1\":29,\"start2\":29,\"length1\":963,\"length2\":2097},{\"diffs\":[[0,\"acle\"],[-1,\"** collection (own_action data) for\"],[1,\"), model selection by log-score/calibration (NOT top-k), plus\"],[0,\" the\"]],\"start1\":2127,\"start2\":2127,\"length1\":43,\"length2\":69},{\"diffs\":[[0,\"ical\"],[-1,\"-\"],[1,\" \"],[0,\"oracle \"],[-1,\"scor\"],[1,\"table as behavioral ground\"],[0,\"ing\"],[-1,\".\"],[0,\"\\\n\\\n## \"],[-1,\"Morning plan\\\n1. fit_log.txt → fitted curves + diagnostics\\\n2. Build empirical oracle table from equity_v8_oracle: P(own_action | beat(h), street, size, name-class); re-run overlap score with it as match(h)\\\n3. g53 candidates: fitted curves / global temperature ~0.55 / oracle-scored variant — full gates (KK, 7_6, Harrington, sanity, A/B, calibration re-audit)\\\n\\\n## Measurement state (corrected, 91a4fc2)\\\n- HU: realized hero-best 0.451; prior-weighted 0.349; ATC 0.527; match at t≈0.55-0.60 (slope 0.86)\\\n- Bias = prior overconcentration (RangeNet), not conditional structure (that was the pairs-only enumeration artifact), not the corrector, not MC\"],[1,\"Assets built this session\\\n- range_temp knob (probs_to_hand_range_temp) — kept, off by default, for future experiments\\\n- g53 audit corpus (917K, temp-flowing, relabeled) — documents the composite-pipeline effect\\\n- Empirical oracle table (models/oracle/oracle_table_v1.json, 5M decisions) + build_oracle_table bin\\\n- fit_likelihood --score / --score-oracle overlap scoring modes\\\n- Fatal-registration sim hardening\\\n\\\n## Recommended next (decision for user)\\\n1. **RangeNet v18 retrain** (the real fix): train on equity_v7_fit corpus (context-matched range data exists there via pre_ranges? — no: use the range_recorder data from the same runs) with (a) label smoothing/temperature-scaled targets, (b) checkpoint selection by calibration slope/log-score, (c) mixed-table data. ~2-3 days.\\\n2. Oracle-in-decisions (Tier 1): swap corrector curves for oracle-table lookups (per-config likelihoods) — improves eq_cc and corrector grounding without retraining.\\\n3. 7_6 still open; FutureActionNet data ready (oracle corpus).\\\n\\\n## Live\\\nG51 on persistent server; session results variance-dominated (day1 +122BB, day2 negative, A8-vs-AJ cooler verified standard).\"],[0,\"\\\n\\\n##\"]],\"start1\":2202,\"start2\":2202,\"length1\":674,\"length2\":1193}]"
metadata_diff: {"new":{},"deleted":[]}
encryption_cipher_text: 
encryption_applied: 0
updated_time: 2026-08-19T01:25:59.718Z
created_time: 2026-08-19T01:25:59.718Z
is_locked: 0
type_: 13