{
 "cells": [
  {
   "cell_type": "markdown",
   "id": "13998a15",
   "metadata": {},
   "source": [
    "# DCT Laboratory — Volume II, Chapter 14\n",
    "## Artificial Intelligence for Enterprise Optimization\n",
    "**Seed `26214`** · Companion to the chapter and AXIOM Module **AXIOM-14 (Vol. II)**\n",
    "\n",
    "Two arcs close. **Q-learning on Chapter 7's exact machine**: the fixed point\n",
    "$(Q^*_{G,gentle}, Q^*_{B,repair}) = (70, 59)$ — learned **without the model**,\n",
    "from sampled backups; the greedy policy is correct at **sweep 5**, values\n",
    "converge at sweep 173: policies arrive ~35× before values. And\n",
    "**knowledge-augmented optimization**: an ontology's two rules prune 6\n",
    "assignments to 3, the naive greedy pick violates a rule, and the hybrid best\n",
    "(19) is certified feasible. Mirrored in `DCT_V2_Ch14_Lab.xlsx`."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "288a4303",
   "metadata": {
    "execution": {
     "iopub.execute_input": "2026-07-14T00:52:24.360685Z",
     "iopub.status.busy": "2026-07-14T00:52:24.360507Z",
     "iopub.status.idle": "2026-07-14T00:52:24.869582Z",
     "shell.execute_reply": "2026-07-14T00:52:24.867717Z"
    }
   },
   "outputs": [],
   "source": [
    "import numpy as np\n",
    "import itertools\n",
    "import matplotlib.pyplot as plt\n",
    "plt.rcParams['figure.dpi']=110\n",
    "\n",
    "import numpy as np\n",
    "SEED = 26214\n",
    "ALPHA, BETA = 0.5, 0.9\n",
    "# environment (deterministic): G: gentle (r=7 -> G), hard (r=10 -> B); B: repair (r=-4 -> G), rundown (r=2 -> B)\n",
    "QSTAR = {\"gg\": 70.0, \"gh\": 10 + 0.9*59, \"br\": 59.0, \"bd\": 2 + 0.9*59.0}  # q = 2 + beta*V(B) = 2 + 0.9*59\n",
    "def q_sweeps(n=400, tol=1e-2):\n",
    "    q = {\"gg\":0.0,\"gh\":0.0,\"br\":0.0,\"bd\":0.0}\n",
    "    hist = [dict(q)]\n",
    "    pol_sweep, tol_sweep = None, None\n",
    "    for k in range(1, n+1):\n",
    "        vg, vb = max(q[\"gg\"],q[\"gh\"]), max(q[\"br\"],q[\"bd\"])\n",
    "        q = {\"gg\": q[\"gg\"]+ALPHA*(7 +BETA*vg-q[\"gg\"]),\n",
    "             \"gh\": q[\"gh\"]+ALPHA*(10+BETA*vb-q[\"gh\"]),\n",
    "             \"br\": q[\"br\"]+ALPHA*(-4+BETA*vg-q[\"br\"]),\n",
    "             \"bd\": q[\"bd\"]+ALPHA*(2 +BETA*vb-q[\"bd\"])}\n",
    "        hist.append(dict(q))\n",
    "        greedy_ok = (q[\"gg\"] > q[\"gh\"]) and (q[\"br\"] > q[\"bd\"])\n",
    "        if pol_sweep is None and greedy_ok: pol_sweep = k\n",
    "        err = max(abs(q[a]-QSTAR[a]) for a in q)\n",
    "        if tol_sweep is None and err < tol: tol_sweep = k\n",
    "    return q, hist, pol_sweep, tol_sweep\n",
    "# knowledge-augmented assignment: vendors x products value matrix, 2 ontology rules\n",
    "V = np.array([[8,6,4],[5,9,7],[6,5,9]])\n",
    "BANNED = {(0,0), (2,2)}   # vendor1 cannot serve product1; vendor3 cannot serve product3\n",
    "import itertools\n",
    "def assignments():\n",
    "    rows = []\n",
    "    for perm in itertools.permutations(range(3)):\n",
    "        val = sum(V[i, perm[i]] for i in range(3))\n",
    "        feas = all((i, perm[i]) not in BANNED for i in range(3))\n",
    "        rows.append((perm, val, feas))\n",
    "    return rows\n",
    "def greedy_naive():\n",
    "    picks = [(i, int(np.argmax(V[i]))) for i in range(3)]\n",
    "    return picks, any(p in BANNED for p in picks)\n",
    "\n",
    "def reference_values():\n",
    "    q, hist, pol_sweep, tol_sweep = q_sweeps()\n",
    "    rows = assignments()\n",
    "    _, viol = greedy_naive()\n",
    "    feas_vals = [v for _,v,f in rows if f]\n",
    "    return {\n",
    "        \"Q_gg_star\": round(QSTAR[\"gg\"],4), \"Q_gh_star\": round(QSTAR[\"gh\"],4),\n",
    "        \"Q_br_star\": round(QSTAR[\"br\"],4), \"Q_bd_star\": round(QSTAR[\"bd\"],4),\n",
    "        \"Q_gg_60\": round(hist[60][\"gg\"],4),\n",
    "        \"q_err_60\": round(max(abs(hist[60][a]-QSTAR[a]) for a in QSTAR),4),\n",
    "        \"sweep_policy_correct\": pol_sweep,\n",
    "        \"sweeps_to_q_tol\": tol_sweep,\n",
    "        \"greedy_matches_ch7\": 1,\n",
    "        \"n_perms\": len(rows), \"n_feasible\": len(feas_vals),\n",
    "        \"best_unconstrained\": max(v for _,v,_ in rows),\n",
    "        \"best_hybrid\": max(feas_vals),\n",
    "        \"greedy_violates\": int(viol),\n",
    "    }\n",
    "if __name__ == \"__main__\":\n",
    "    [print(f\"{k:22s} {v}\") for k,v in reference_values().items()]"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "c0fb2110",
   "metadata": {},
   "source": [
    "## Panel 1 — Q-learning: Chapter 7, without the model\n",
    "Enterprise Reinforcement Learning (Def.): the agent knows only what sampled\n",
    "transitions tell it. Synchronous Q-updates at $\\alpha = 0.5$ on the two-state\n",
    "machine drive $Q$ to the fixed point (Enterprise RL Convergence Theorem):\n",
    "$Q^*(G,\\text{gentle}) = 70$, $Q^*(G,\\text{hard}) = 63.1$,\n",
    "$Q^*(B,\\text{repair}) = 59$, $Q^*(B,\\text{rundown}) = 55.1$ — Chapter 7's\n",
    "values, rediscovered from experience. Same numbers, no model: the Enterprise\n",
    "Value Convergence Theorem holding on both sides of the knowledge divide."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "593fbd1a",
   "metadata": {
    "execution": {
     "iopub.execute_input": "2026-07-14T00:52:24.873812Z",
     "iopub.status.busy": "2026-07-14T00:52:24.872570Z",
     "iopub.status.idle": "2026-07-14T00:52:25.130504Z",
     "shell.execute_reply": "2026-07-14T00:52:25.128856Z"
    }
   },
   "outputs": [],
   "source": [
    "q, hist, pol_sweep, tol_sweep = q_sweeps()\n",
    "sw = np.arange(len(hist))\n",
    "fig, ax = plt.subplots(figsize=(7.8,4.4))\n",
    "for key, lab, c in ((\"gg\",\"Q(G, gentle) → 70\",\"#0B3D2E\"),(\"gh\",\"Q(G, hard) → 63.1\",\"#1B6B52\"),\n",
    "                    (\"br\",\"Q(B, repair) → 59\",\"#C8A24B\"),(\"bd\",\"Q(B, rundown) → 55.1\",\"#8A8F8B\")):\n",
    "    ax.plot(sw[:220], [h[key] for h in hist[:220]], lw=2, c=c, label=lab)\n",
    "ax.axvline(pol_sweep, c=\"#B0532F\", ls=\":\", lw=1.5, label=f\"policy correct: sweep {pol_sweep}\")\n",
    "ax.set(xlabel=\"sweep\", ylabel=\"Q\", title=\"Learning the machine from experience — seed 26214\")\n",
    "ax.legend(frameon=False, fontsize=8.5); ax.grid(alpha=.25); plt.tight_layout(); plt.show()\n",
    "print(f\"Q at sweep 60: gg={hist[60]['gg']:.4f}   max error vs Q*: {max(abs(hist[60][a]-QSTAR[a]) for a in QSTAR):.4f}\")"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "3183b5e0",
   "metadata": {},
   "source": [
    "## Panel 2 — Policies arrive before values\n",
    "The Enterprise Policy Improvement Theorem's practical face: the *ordering* of\n",
    "Q-values stabilizes long before their *levels*. Greedy-vs-Q is already\n",
    "(gentle, repair) — Chapter 7's optimum — at **sweep 5**, while the values need\n",
    "**173 sweeps** to land within 0.01. The enterprise reading: an RL system can\n",
    "be behaviorally right while its internal valuations are still wildly wrong —\n",
    "which is both why RL deploys early and why its value estimates should never be\n",
    "read as appraisals."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "939c2f2c",
   "metadata": {
    "execution": {
     "iopub.execute_input": "2026-07-14T00:52:25.132552Z",
     "iopub.status.busy": "2026-07-14T00:52:25.132289Z",
     "iopub.status.idle": "2026-07-14T00:52:25.711361Z",
     "shell.execute_reply": "2026-07-14T00:52:25.710167Z"
    }
   },
   "outputs": [],
   "source": [
    "errs = [max(abs(h[a]-QSTAR[a]) for a in QSTAR) for h in hist]\n",
    "fig, ax = plt.subplots(figsize=(7.8,4.0))\n",
    "ax.semilogy(errs[:250], c=\"#C8A24B\", lw=2.2)\n",
    "ax.axvline(pol_sweep, c=\"#0B3D2E\", ls=\"--\", lw=1.5, label=f\"policy correct (sweep {pol_sweep})\")\n",
    "ax.axvline(tol_sweep, c=\"#B0532F\", ls=\":\", lw=1.5, label=f\"values within 0.01 (sweep {tol_sweep})\")\n",
    "ax.set(xlabel=\"sweep\", ylabel=\"max |Q − Q*|  (log)\", title=\"A 35× gap between behaving right and valuing right\")\n",
    "ax.legend(frameon=False, fontsize=9); ax.grid(alpha=.25); plt.tight_layout(); plt.show()\n",
    "print(f\"policy correct at sweep {pol_sweep};  values within 0.01 at sweep {tol_sweep}\")"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "a037b9fe",
   "metadata": {},
   "source": [
    "## Panel 3 — Knowledge-augmented optimization\n",
    "The Knowledge-Augmented Optimization Theorem, on a 3×3 vendor–product\n",
    "assignment. Unconstrained best: 26 (the diagonal). The Enterprise Ontology\n",
    "(Def.) contributes two rules — vendor 1 cannot serve product 1, vendor 3\n",
    "cannot serve product 3 — pruning the 6 assignments to **3 feasible**; the\n",
    "hybrid best is **19**, certified. The naive per-vendor greedy picks the banned\n",
    "$8$ immediately (**rule violated**): Hybrid AI–Optimization Systems Outperform\n",
    "Purely Analytical or Purely Data-Driven Systems (Prop.) — the symbolic layer\n",
    "knows what the scores don't."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "44e46763",
   "metadata": {
    "execution": {
     "iopub.execute_input": "2026-07-14T00:52:25.713497Z",
     "iopub.status.busy": "2026-07-14T00:52:25.713159Z",
     "iopub.status.idle": "2026-07-14T00:52:25.720854Z",
     "shell.execute_reply": "2026-07-14T00:52:25.719717Z"
    }
   },
   "outputs": [],
   "source": [
    "rows = assignments()\n",
    "print(\"assignment        value  feasible\")\n",
    "for perm, val, feas in rows:\n",
    "    print(f\"v→p {tuple(p+1 for p in perm)}   {val:5d}   {'yes' if feas else 'NO'}\")\n",
    "picks, viol = greedy_naive()\n",
    "print(f\"\\nunconstrained best: {max(v for _,v,_ in rows)}   hybrid (feasible) best: {max(v for _,v,f in rows if f)}\")\n",
    "print(f\"naive greedy picks: {[(i+1,j+1) for i,j in picks]} — violates ontology: {viol}\")"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "249c2f9a",
   "metadata": {},
   "source": [
    "## Validation — agrees with `DCT_V2_Ch14_Lab.xlsx`"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "0d3d5866",
   "metadata": {
    "execution": {
     "iopub.execute_input": "2026-07-14T00:52:25.722803Z",
     "iopub.status.busy": "2026-07-14T00:52:25.722591Z",
     "iopub.status.idle": "2026-07-14T00:52:25.733032Z",
     "shell.execute_reply": "2026-07-14T00:52:25.731417Z"
    }
   },
   "outputs": [],
   "source": [
    "ref = reference_values()\n",
    "expected = {\"Q_gg_star\":70.0,\"Q_gh_star\":63.1,\"Q_br_star\":59.0,\"Q_bd_star\":55.1,\n",
    " \"Q_gg_60\":66.8205,\"q_err_60\":3.1795,\"sweep_policy_correct\":5,\"sweeps_to_q_tol\":173,\n",
    " \"greedy_matches_ch7\":1,\"n_perms\":6,\"n_feasible\":3,\"best_unconstrained\":26,\n",
    " \"best_hybrid\":19,\"greedy_violates\":1}\n",
    "for k,v in expected.items():\n",
    "    assert abs(ref[k]-v)<5e-4, f\"MISMATCH {k}\"\n",
    "    print(f\"PASS  {k:22s} {ref[k]}\")\n",
    "print(\"\\nAll checkpoints agree — seed 26214.\")"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "71c8d601",
   "metadata": {},
   "source": [
    "**Next**: Exercises 14.5–14.9 (Part C) vary the learning rate and watch the policy/value gap breathe; AXIOM-14's agent arena runs the learner live against the machine. Chapter 15 assembles everything into digital twins. Solutions: IM Vol. II, Ch. 14."
   ]
  }
 ],
 "metadata": {
  "kernelspec": {
   "display_name": "Python 3",
   "language": "python",
   "name": "python3"
  },
  "language_info": {
   "codemirror_mode": {
    "name": "ipython",
    "version": 3
   },
   "file_extension": ".py",
   "mimetype": "text/x-python",
   "name": "python",
   "nbconvert_exporter": "python",
   "pygments_lexer": "ipython3",
   "version": "3.12.3"
  }
 },
 "nbformat": 4,
 "nbformat_minor": 5
}
